Commit Graph

17 Commits

Author SHA1 Message Date
yoniebans 2f25c07a2c perf(install-e2e): 60-minute job caps; 35-minute updater wait
The slowest green leg ever recorded is 29 minutes; every cap hit in the
suite's history was a hang, never work. Caps were linux 75 / macos 120 /
windows 240, so a wedged leg burned up to 4 hours of runner time to
report what its log showed in the first minutes. 60 minutes covers the
slowest leg plus cold-cache variance, and every driver-internal bound
(dmg install 45m, AHK 50m, updater wait) still fires before the job cap
in any single-hang scenario, keeping failure diagnostics specific.

The detached-updater wait drops 90m -> 35m on the same evidence: a
working updater finishes far inside 35m; a wedged one never finishes at
any bound, and the longer wait only delayed the report by an hour.
2026-09-01 17:34:13 +02:00
yoniebans bf75c52ba2 docs(install-e2e): retire stale TODO prose now every driver arm exists
The workflow input descriptions and the skips README still declared
open-app-update and the Setup.exe re-run as driver TODOs; both run now.
Skips have exactly two causes and the prose names them: no OS entry
point for the pair, or the starting release predates the surface. The
chart's TODO label itself stays until the n/a relabel lands with the
known-broken-OLD gate work.
2026-09-01 17:11:42 +02:00
ethernet cd667debfa fix(install-e2e): windows transcripts were empty; player gets #zip= hash + one player per run
Windows transcripts were ZERO bytes: ts-prefix.ps1 formatted with
{0:D2}, but Floor() returns a double and the D specifier is
integer-only - it threw per line, and under the driver's relaxed EAP
every line errored into the void. {0:00} fixes it (custom numeric
format works on doubles). Reproduced the exact pipeline locally
(empty file + Format specifier invalid), verified the fix produces
prefixed merged stdout+stderr with exit code intact. That is also
why the log timeline never auto-synced: there was nothing in the
files to sync.

The GitHub artifact URL 307s to /suites/... server-side and strips
the ?zip= query param. The player now reads the zip URL from a
#zip= HASH param (client-side, survives the redirect) with ?zip=
as fallback; the hash path was verified in a real browser against a
real leg zip (auto-fetch + boot).

Per ethie's design, one player artifact for the whole run: new
leg-player job uploads playback.html (archive:false) before the
matrix legs, the report job needs it, and each ran cell gets TWO
links - 📼 to the player with #zip=<that leg's logs zip> and ⬇️ to
the raw zip. Per-leg player uploads removed from all three run
workflows.
2026-08-13 22:51:44 -04:00
ethernet e138cb555d feat(install-e2e): hook the leg player into the results table
Each leg uploads playback.html as a single-file artifact (archive:
false) before the driver runs, so it exists even on failure. The
results chart now links every leg that RAN (pass or fail, not skip)
to its player with ?zip= pointing at that leg's logs artifact.

Leg<->artifact mapping: the generator mints a leg_id per matrix entry
(sanitized matrix name, exported legId()), every run workflow names
its artifacts install-e2e-{player,logs}-<leg-id>, and the report job
feeds the run's artifact name->id list to the results renderer, which
rebuilds the leg id from the parsed job name. GitHub does not link
jobs to artifacts, so the deterministic name is the join key.

Empirical finding: GitHub artifact downloads are auth-gated (the
download URL 307s to /suites/... which is 404 anonymous), so a
locally-opened player page cannot fetch the zip cross-origin. The
player now degrades gracefully: ?zip= fetch failure renders a real
download link for the zip (a normal click carries the user's session)
plus a drag-and-drop / file-picker path, and no-param opens as a pure
drop target. Verified in a real browser against a real artifact URL.

Verified: generator emits leg_id, results renderer emits
✅/❌ [📼](...?zip=...) only on ran cells, npm run check PASS,
install tests 36/36, strict tsc PASS, actionlint x4 PASS.
2026-08-12 15:50:46 -04:00
ethernet 05ffab9d18 feat(install-e2e): playback.html leg player - zip in, video + time-synced logs
A static single-file player (tests/install/e2e-assets/playback.html):
?zip=<artifact zip url> unzips in-browser (JSZip), plays the screen
recording with a timer pinned top-left, and renders every *.log with
video<->log sync: the video follows the driver's transcript, clicking
a log line seeks the video. A sync-offset slider aligns the recording
start (ffmpeg comes up first) with the driver's relative clock.

Sync axis: drivers now prefix every transcript line with [+MM:SS]
relative to driver start (ts-prefix.sh / ts-prefix.ps1, pipe-safe
under pipefail / relaxed EAP). Browsers cannot play Matroska, so each
leg remuxes recording.mkv -> recording.mp4 (-c copy, no re-encode)
before the artifact upload, on all three OSes.

Verified end-to-end in a real browser against a generated artifact
zip: zip load, mp4 playback, timer, tab switching, follow-sync at
t=6/t=12, click-to-seek, autoplay policy (expected NotAllowedError on
synthetic play; real clicks fine).

Also fixes the shim fail message's dead variable ( ->
observed_git_url) in both posix drivers.
2026-08-12 15:28:08 -04:00
ethernet 0f903e14a3 test(install-e2e): hermes-desktop-app-update goes live on the script driver
Playwright must own the spawn (it needs the inspection pipe), but
hermes desktop is not just build+launch - stamp checks, integrity
gates, sandbox fixups, and a constructed child environment. So the
driver intercepts the product's own launch: a sitecustomize.py on
PYTHONPATH (opt-in via HERMES_E2E_CAPTURE_LAUNCH) wraps subprocess.run,
captures argv/cwd/env at the spawn site, and fakes success instead of
spawning; launch-from-spec.mjs then _electron.launch-es exactly that
spec and clicks Settings -> About -> Update now. Completion is product
state, not a Playwright event: the handoff result file or the checkout
reaching the expected sha (source installs write no result file).

Ships with the driver, so it works unchanged on every sampled OLD ref
- no product flag, no pre-flag fallback split. Both launch shapes are
matched (npm exec electron / packaged exe under apps/desktop/release);
npm BUILD calls pass through untouched. Exit 0 without a capture fails
the leg: a version that never reached its launch must not pass.

Probe-the-probe: scripts/launch_capture_probe.sh runs control rows
(no opt-in, non-launch argv) and both treatment shapes - all green
locally. Gate flips on the shared run workflow for linux/macos;
windows adopts the same path with the driver restructuring.
2026-08-12 04:20:56 -04:00
ethernet cdaf0cf091 ci(install-e2e): one screen-recording mechanism on every runner, Xvfb for headless linux
The composite action .github/actions/e2e-screen-record owns setup and
lifecycle on all three OSes: ffmpeg via apt/brew-verify/winget+cache,
capture via x11grab/gdigrab/avfoundation, mkv at 15fps stopped by 'q'
on live stdin with kill fallback. Linux runners have no display, so
start brings up a dedicated Xvfb :99 and exports DISPLAY - one display
serves both the recorder and any app a later step launches.

Recording moves out of the GUI driver into workflow infrastructure -
that is what makes it uniform - and a missing ffmpeg or a zero-frame
file now FAILS the leg instead of skipping silently: the graceful-skip
path is how the windows leg shipped no recording.mkv while green.

Lifecycle proven locally: start against lavfi testsrc, q-stop, ffprobe
duration check (record-start.sh/record-stop.sh under nix ffmpeg).
2026-08-12 04:15:24 -04:00
ethernet 239523414e test(install-e2e): installer-script+desktop is its own install and update method
The one-liner with its desktop stage opted in (--include-desktop /
-IncludeDesktop) is a real install kind, distinct on both sides:
on windows the stage builds Hermes.exe AND registers Start Menu /
Desktop shortcuts - a second path to a hand-launchable app - while
on linux/macos it builds into the checkout and registers no OS
entry point.

Declared on every OS and driven by both script drivers: the drivers
pass the flag through (hard failure if the ref predates it - the
tag-has-desktop gate already skips pre-desktop tags upstream) and
assert the built app exists under apps/desktop/release afterwards.
The run-workflow gates run +desktop pairs only on desktop-bearing
tags; app-update pairs from +desktop installs stay declared TODOs.
2026-08-12 04:01:16 -04:00
ethernet 1af2093663 test(install-e2e): split app-update into open-app-update + hermes-desktop-app-update
The desktop app has two launch paths, so app-update becomes two
methods. open-app-update starts the app from the OS entry point the
desktop installer created (the installed exe / the .app), so it exists
only where a desktop installer does. hermes-desktop-app-update starts
the app via hermes desktop, which every install method provides on
every OS that ships the desktop app - on linux it is the only app
surface, since no desktop installer or packaged artifact exists there.

Both variants are desktop-surface methods on every OS, so the
tag_has_desktop annotation moves from windows-only to every matrix
entry, install-e2e-run.yml grows the input, and the plan chart marks
pre-desktop cells on all OSes.

The windows GUI arm's implemented pair renames to open-app-update;
every other new combination is a declared TODO that natively skips.
2026-08-12 03:47:27 -04:00
ethernet cde3b4ff09 fix some names 2026-08-11 23:00:42 -04:00
ethernet ea4cd375f8 ci(install-e2e): retire the bubblewrap sandbox - git redirect everywhere, macos legs live
The fake Internet (bubblewrap + slirp4netns + MITM proxy +
upload-pack shim, 883 lines across dev-sandbox.sh, stage2-run.sh,
proxy.py, ssh-shim.sh, openssl.cnf, install-update-e2e.sh) existed to
isolate install.sh's network. The GIT_CONFIG_GLOBAL insteadOf redirect
the windows driver introduced does the same job with a gitconfig file
and works on any OS, so:

* install-e2e-run.yml now runs tests/install/installer-script-e2e.sh
  directly on the bare runner - no sandbox deps, no userns sysctls -
  and takes a runner input;
* the macos matrix calls the SAME workflow on macos-latest, deleting
  install-e2e-macos-run.yml: installer-script -> installer-script /
  hermes-update flip from grey to live, app-update pairs stay TODO
  inside the shared gate;
* install.sh is no longer curl'd through a fake CA - each leg runs
  the copy from the ref a user of that version actually executed;
* scripts/dev-sandbox.sh becomes the minimal isolation sandbox from
  ab6b9492f (separate HERMES_HOME / Electron userData / app name,
  same CLI surface: --persistent, --from, --delete), keeping its
  .hermes-sandbox dir name so gitignore and docs hold;
* nix/sandbox.nix drops the bwrap/proxy closure and keeps only the
  Electron runtime LD_LIBRARY_PATH the desktop app needs.

Verified: nix build .#sandbox + smoke run (isolated HERMES_HOME
created, ephemeral cleanup), shellcheck/bash -n on both scripts,
actionlint on all three workflows, and the new driver ran the full
v0.20.2 -> HEAD hermes-update pass locally before this commit.
2026-08-11 21:20:42 -04:00
ethernet 6d7ba86963 ci(install-e2e): collapse the method vocabulary - 3 install ids, install+2 update ids
Per review the unions were overcomplicated. Install methods are now
just: installer-script (the platform one-liner - curl | bash on
linux/macos, irm | iex on windows), desktop-installer, and
packaged-app (declared, unused). Update methods are every install
method (re-run it over the existing install) plus hermes-update and
app-update. desktop-installer-rerun, desktop-app, curl-bash, and
irm-iex are gone as ids; the windows driver's ValidateSet, switch
arms, and both run workflows' gates renamed to match. tsc --checkJs
clean; generator output re-verified (4 linux / 16 windows / 6 macos
legs for 2 tags).
2026-08-11 17:14:17 -04:00
ethernet 1c60bdc30d ci(install-e2e): back to the generator - workflows stay generic, names carry everything
Revert the hardcoded 16-job experiment: the combination spec belongs
in scripts/sandbox/generate-e2e-matrix.mjs (restored), not copy-pasted
YAML blocks. What survives from the experiment:

* leg names carry everything - 'os: install -> update (tag -> HEAD)' -
  generated per entry, since slash-joined names are all the graph
  renders;
* pick-releases annotates each tag ({ref, desktop}) and the generator
  threads tag_has_desktop onto windows entries, so the windows run
  workflow still gates pre-desktop tags without a probe job;
* the per-OS run workflows are untouched: single job, static 'e2e'
  name, native skip gates own all capability knowledge.

install-e2e.yml is one generate job + three per-OS matrix fanouts.
Generator shape (4/16/12 legs for 2 tags), annotation threading, and
all six error paths verified; all four workflows pass actionlint;
driver parses clean pure-ASCII.
2026-08-11 16:54:54 -04:00
ethernet 83e9f883d8 ci(install-e2e): per-leg names are just the transition; tag capability annotated at pick time
Graph polish + one structural simplification, after the first render
of the combo-box layout:

* Leg names: every combination job's display name is now
  '${{ matrix.tag.ref }} -> HEAD' - the box title (job id) already
  carries os+methods, so repeating them per leg was noise. The inner
  job renders as a short static 'e2e' tail (dynamic names render
  unexpanded on skipped jobs, so it must stay static).

* The windows probe job is gone: pick-releases now annotates each
  picked tag with whether its tree ships apps/desktop
  ({ref, desktop} objects in the matrix), and the windows run
  workflow gates on the new tag-has-desktop boolean input directly.
  One tree listing at pick time replaces N probe jobs, and the
  'probe tag' noise disappears from the graph.

Annotation loop verified against the real tag set (pre/post-desktop
split lands exactly at the app's introduction); 16-combo inventory
re-asserted; all four workflows pass actionlint.
2026-08-11 16:43:26 -04:00
ethernet 6515f8132a ci(install-e2e): hardcode combination jobs in the primary workflow - one matrix box per combo
GitHub only draws matrix boxes for the PRIMARY workflow's matrices;
everything inside a called workflow flattens into slash-joined names.
The generator + per-tag sub-workflow therefore bought no structure in
the graph and hid the support matrix in a script.

Invert it: install-e2e.yml now declares one job per {os,
install-method -> update-method} combination (2 linux + 8 windows +
6 macos - same 16 the generator produced, verified by inventory
before/after), each a matrix over the picked release tags. The graph
now renders one titled box per combination whose legs read
'... from vX' - the tag axis inside the combo axis. The per-OS run
workflows are unchanged: they own capability knowledge and natively
skip unimplemented method pairs and pre-desktop tags.

install-e2e-tag.yml and generate-e2e-matrix.mjs are deleted; adding a
method is now adding one job block here, implementing one is flipping
the run workflow's gate.
2026-08-11 16:35:03 -04:00
ethernet 6d94678e62 ci(install-e2e): jobs own their skips - generator is pure expansion
Remove all capability knowledge from the combination generator: no
IMPLEMENTED table, no skipped matrix, no per-OS special cases. It now
only declares and expands - every {os, install-method, update-method}
combination is dispatched to its OS's run workflow, and each run
workflow natively skips (grey, job-level if on the method inputs) the
pairs its driver cannot run yet:

* install-e2e-run.yml gains install-method/update-method inputs,
  gates on the supported pairs (curl-bash -> hermes-update/curl-bash),
  and maps the method id to the sandbox script's --route internally;
* install-e2e-macos-run.yml is new - all pairs skip until a macOS
  driver exists, and implementing one flips its job-level if;
* install-e2e-windows-run.yml already worked this way;
* install-e2e-skip.yml is deleted - nothing special-cases macOS
  anymore, so the tag workflow is three identical OS fanouts.

Structure is now uniformly matrix(tag) -> matrix(combination) ->
run-or-skip, with capability knowledge living only next to each
driver. Generator shape/route filters/error paths re-verified; all
five workflows pass actionlint.
2026-08-11 16:17:48 -04:00
ethernet 36cb5ae553 ci: test updating from sampled release tags, on tag + every 12h
Wires tests/install/install-update-e2e.sh into CI as a reusable workflow plus a
caller that fans out over real releases, because that is the question users care
about: can someone on a version they actually installed get to this commit?

install-e2e-run.yml takes `route` and `install-ref`, so the combinations that
matter are expressible without duplicating runner setup. Each leg is independent
-- its own runner, its own sandbox, its own install, nothing shared or rewound.

The starting versions are chosen at runtime by scripts/sandbox/pick-release-tags.sh:
newest, oldest, and an evenly spaced spread between (5 by default). Choosing at
runtime rather than hardcoding keeps the matrix honest -- a pinned list stops
covering the newest release the day after it ships, and pins an "oldest" long
after anyone still runs it. Newest catches "did the last release break
updating?", oldest is the longest upgrade jump still possible, and the spread
samples the migrations in between (config-schema bumps, venv layout changes,
dependency floors). Tags are read from the checkout with `git tag --list`, not
`git ls-remote`: the job has the repository already, so this needs no network,
works offline and on a fork, and takes 8ms. The repo is derived from the
script's own resolved path rather than $PWD, so a copy cannot silently report a
different checkout's tags. The pick-releases job takes the checkout that suits
it -- blob:none filter, sparse-checkout of just that script, and fetch-tags,
since tags are the entire input and the default shallow checkout has none.

Triggers match the shape of the work:

  * every 12 hours, so upstream drift (a new uv, a Node bump, a PyPI change)
    surfaces on a schedule instead of in someone's review cycle;
  * on release tags, the moment the set of versions users can update FROM
    changes and the moment a broken updater would strand them;
  * manually, with the route and the sample size as inputs.

Not on pull_request: a leg is ~9 minutes of real toolchain installation and the
matrix multiplies it. fail-fast is off so one broken release does not mask the
others, and max-parallel caps the fan-out so a run does not hammer the runners
or PyPI. The tag list is resolved once and shared by both route matrices, so the
two routes cover the same versions.

Artifact names include the sanitized install-ref, since a matrix runs the
reusable workflow several times per route and same-named artifacts collide; that
name is built in a step because Actions expressions have no string-replace
function. The name step runs with `if: always()`, since a failing leg is exactly
when its logs are wanted.

.gitignore covers .hermes-sandbox-e2e*/ rather than the bare directory: the
per-route sandbox trees (-update, -installer) fell outside it, so the sandbox
made the worktree dirty and dev-sandbox reacted by snapshotting the working copy
into a fresh fake-main commit on every invocation.
2026-08-04 17:36:26 -04:00