Playwright must own the spawn (it needs the inspection pipe), but
hermes desktop is not just build+launch - stamp checks, integrity
gates, sandbox fixups, and a constructed child environment. So the
driver intercepts the product's own launch: a sitecustomize.py on
PYTHONPATH (opt-in via HERMES_E2E_CAPTURE_LAUNCH) wraps subprocess.run,
captures argv/cwd/env at the spawn site, and fakes success instead of
spawning; launch-from-spec.mjs then _electron.launch-es exactly that
spec and clicks Settings -> About -> Update now. Completion is product
state, not a Playwright event: the handoff result file or the checkout
reaching the expected sha (source installs write no result file).
Ships with the driver, so it works unchanged on every sampled OLD ref
- no product flag, no pre-flag fallback split. Both launch shapes are
matched (npm exec electron / packaged exe under apps/desktop/release);
npm BUILD calls pass through untouched. Exit 0 without a capture fails
the leg: a version that never reached its launch must not pass.
Probe-the-probe: scripts/launch_capture_probe.sh runs control rows
(no opt-in, non-launch argv) and both treatment shapes - all green
locally. Gate flips on the shared run workflow for linux/macos;
windows adopts the same path with the driver restructuring.
The composite action .github/actions/e2e-screen-record owns setup and
lifecycle on all three OSes: ffmpeg via apt/brew-verify/winget+cache,
capture via x11grab/gdigrab/avfoundation, mkv at 15fps stopped by 'q'
on live stdin with kill fallback. Linux runners have no display, so
start brings up a dedicated Xvfb :99 and exports DISPLAY - one display
serves both the recorder and any app a later step launches.
Recording moves out of the GUI driver into workflow infrastructure -
that is what makes it uniform - and a missing ffmpeg or a zero-frame
file now FAILS the leg instead of skipping silently: the graceful-skip
path is how the windows leg shipped no recording.mkv while green.
Lifecycle proven locally: start against lavfi testsrc, q-stop, ffprobe
duration check (record-start.sh/record-stop.sh under nix ffmpeg).
The one-liner with its desktop stage opted in (--include-desktop /
-IncludeDesktop) is a real install kind, distinct on both sides:
on windows the stage builds Hermes.exe AND registers Start Menu /
Desktop shortcuts - a second path to a hand-launchable app - while
on linux/macos it builds into the checkout and registers no OS
entry point.
Declared on every OS and driven by both script drivers: the drivers
pass the flag through (hard failure if the ref predates it - the
tag-has-desktop gate already skips pre-desktop tags upstream) and
assert the built app exists under apps/desktop/release afterwards.
The run-workflow gates run +desktop pairs only on desktop-bearing
tags; app-update pairs from +desktop installs stay declared TODOs.
The desktop app has two launch paths, so app-update becomes two
methods. open-app-update starts the app from the OS entry point the
desktop installer created (the installed exe / the .app), so it exists
only where a desktop installer does. hermes-desktop-app-update starts
the app via hermes desktop, which every install method provides on
every OS that ships the desktop app - on linux it is the only app
surface, since no desktop installer or packaged artifact exists there.
Both variants are desktop-surface methods on every OS, so the
tag_has_desktop annotation moves from windows-only to every matrix
entry, install-e2e-run.yml grows the input, and the plan chart marks
pre-desktop cells on all OSes.
The windows GUI arm's implemented pair renames to open-app-update;
every other new combination is a declared TODO that natively skips.
Each script-driver leg now proves the installed CLI can build the
desktop app, after the install phase and again after the update.
--build-only runs the full desktop pipeline and stops before the
launch - the same call hermes update makes. Old releases that predate
the flag skip the phase after a --help probe of the installed binary.
Actually launching the app is a TODO: it needs the spawn-interception
launcher and, on linux runners, a virtual display.
Also adds tests/install/README.md describing how the test family
works: the four layers, the git-redirect isolation, the phases, the
probe-do-not-assume rule for old versions, skips, triggers, artifacts.
All three drivers (installer-script-e2e.sh, windows-installer-script
-e2e.ps1, windows-desktop-gui-e2e.ps1) now emit the complete
install/update transcript into the job log wrapped in ::group::/
::endgroup:: - collapsed by default, one click to expand, win or
lose. Replaces the tail-50-only-on-failure pattern: a green
install's transcript is how you diagnose the leg that fails next,
and the artifact download was the only way to see it before. The
GUI driver's bootstrap-installer.log / desktop-update-handoff.log
tails become full folded dumps too.
First real run of windows-installer-script-e2e.ps1 died on
'Cloning into ...': git clone writes progress to stderr, and under
EAP=Stop the outer PowerShell wraps a child's stderr into a
terminating NativeCommandError. Drop to EAP=Continue around every
native invocation that redirects with *> (install.ps1 run, hermes
--version, hermes update) and judge by exit code alone - the same
prevEap pattern the GUI driver already uses.
Run 31553293275: the desktop leg from v2026.6.19 failed at
'@playwright/test resolvable from installed apps/desktop' - that
release predates the dependency, and newer releases move it around
via workspace hoisting, so resolving it from the installed tree made
the leg's tooling a function of the version under test.
Install a pinned @playwright/test (param, default 1.58.2 - the repo
lockfile's current version) into a scratch driver dir with the
managed node/npm every run and launch the driver from there. The
driver talks to the app over Playwright's inspection pipe, so its
Playwright is independent of the app; every OLD ref now runs the
exact same driver stack.
The install.ps1 sibling of installer-script-e2e.sh: stage serve.git
(main parked at OLD, GIT_CONFIG_GLOBAL insteadOf redirect - NOT env
config, which install.ps1 clobbers), run the install.ps1 shipped AT
the OLD ref headless (-SkipSetup -HermesHome/-InstallDir explicit
because the oldest tags predate the HERMES_HOME env override;
-NonInteractive probed from the ref's own script text), assert the
checkout + venv hermes.exe, advance served main, update via
hermes-update (--yes probed) or HEAD's install.ps1, assert HEAD.
install-e2e-windows-run.yml grows a second job for the arm: the
installer-script x {hermes-update, installer-script} pairs flip from
grey to live, app-update from a script install stays a declared TODO.
PS 5.1-safe pure ASCII.
Stage logic verified behaviorally under pwsh (redirect resolves the
canonical URL to serve.git at OLD, per-ref install.ps1 extraction
parses, advance lands HEAD); the install legs themselves need a real
Windows runner - dispatched next.
The fake Internet (bubblewrap + slirp4netns + MITM proxy +
upload-pack shim, 883 lines across dev-sandbox.sh, stage2-run.sh,
proxy.py, ssh-shim.sh, openssl.cnf, install-update-e2e.sh) existed to
isolate install.sh's network. The GIT_CONFIG_GLOBAL insteadOf redirect
the windows driver introduced does the same job with a gitconfig file
and works on any OS, so:
* install-e2e-run.yml now runs tests/install/installer-script-e2e.sh
directly on the bare runner - no sandbox deps, no userns sysctls -
and takes a runner input;
* the macos matrix calls the SAME workflow on macos-latest, deleting
install-e2e-macos-run.yml: installer-script -> installer-script /
hermes-update flip from grey to live, app-update pairs stay TODO
inside the shared gate;
* install.sh is no longer curl'd through a fake CA - each leg runs
the copy from the ref a user of that version actually executed;
* scripts/dev-sandbox.sh becomes the minimal isolation sandbox from
ab6b9492f (separate HERMES_HOME / Electron userData / app name,
same CLI surface: --persistent, --from, --delete), keeping its
.hermes-sandbox dir name so gitignore and docs hold;
* nix/sandbox.nix drops the bwrap/proxy closure and keeps only the
Electron runtime LD_LIBRARY_PATH the desktop app needs.
Verified: nix build .#sandbox + smoke run (isolated HERMES_HOME
created, ephemeral cleanup), shellcheck/bash -n on both scripts,
actionlint on all three workflows, and the new driver ran the full
v0.20.2 -> HEAD hermes-update pass locally before this commit.
The POSIX sibling of windows-desktop-gui-e2e.ps1, sharing its staging
trick: bare-clone the checkout to serve.git, park main at OLD, point
every git process at it with url.<file://serve.git>.insteadOf in a
driver-owned GIT_CONFIG_GLOBAL. The installer and updater run
byte-for-byte against their real URLs; no bwrap, no MITM proxy, no
TLS interception - a disposable CI runner IS the sandbox, so the
same driver can run on macos-latest unchanged.
install.sh is not curl'd: the install leg runs the copy shipped AT
the OLD ref (what a user who installed then actually executed), the
installer-script update leg runs HEAD's copy (what the website
serves at update time). Flags are probed per-ref (--skip-browser is
newer than sampled tags); HOME is isolated because old installers
hardcode ~/.hermes; .skip_upstream_prompt suppresses the updater's
fork prompt on the file:// origin; the dirty-tree guard checks
tracked files only (-uno) since untracked files cannot leak into a
bare clone.
Verified locally end-to-end: v0.20.2 installed via its own
install.sh (uv, managed Python, Node, venv; hermes --version OK),
served main advanced, hermes update landed the checkout on HEAD
with a working hermes. bash -n + shellcheck clean.
Per review the unions were overcomplicated. Install methods are now
just: installer-script (the platform one-liner - curl | bash on
linux/macos, irm | iex on windows), desktop-installer, and
packaged-app (declared, unused). Update methods are every install
method (re-run it over the existing install) plus hermes-update and
app-update. desktop-installer-rerun, desktop-app, curl-bash, and
irm-iex are gone as ids; the windows driver's ValidateSet, switch
arms, and both run workflows' gates renamed to match. tsc --checkJs
clean; generator output re-verified (4 linux / 16 windows / 6 macos
legs for 2 tags).
Revert the hardcoded 16-job experiment: the combination spec belongs
in scripts/sandbox/generate-e2e-matrix.mjs (restored), not copy-pasted
YAML blocks. What survives from the experiment:
* leg names carry everything - 'os: install -> update (tag -> HEAD)' -
generated per entry, since slash-joined names are all the graph
renders;
* pick-releases annotates each tag ({ref, desktop}) and the generator
threads tag_has_desktop onto windows entries, so the windows run
workflow still gates pre-desktop tags without a probe job;
* the per-OS run workflows are untouched: single job, static 'e2e'
name, native skip gates own all capability knowledge.
install-e2e.yml is one generate job + three per-OS matrix fanouts.
Generator shape (4/16/12 legs for 2 tags), annotation threading, and
all six error paths verified; all four workflows pass actionlint;
driver parses clean pure-ASCII.
GitHub only draws matrix boxes for the PRIMARY workflow's matrices;
everything inside a called workflow flattens into slash-joined names.
The generator + per-tag sub-workflow therefore bought no structure in
the graph and hid the support matrix in a script.
Invert it: install-e2e.yml now declares one job per {os,
install-method -> update-method} combination (2 linux + 8 windows +
6 macos - same 16 the generator produced, verified by inventory
before/after), each a matrix over the picked release tags. The graph
now renders one titled box per combination whose legs read
'... from vX' - the tag axis inside the combo axis. The per-OS run
workflows are unchanged: they own capability knowledge and natively
skip unimplemented method pairs and pre-desktop tags.
install-e2e-tag.yml and generate-e2e-matrix.mjs are deleted; adding a
method is now adding one job block here, implementing one is flipping
the run workflow's gate.
Two structural changes to the combination fanout:
1. Tags become the OUTER axis, as a sub-graph per starting version:
install-e2e.yml fans a plain matrix over the picked tags into a new
per-tag reusable workflow (install-e2e-tag.yml), which runs the
combination generator for that one tag and fans out one job per
{os, install-method, update-method}. The Actions graph now reads
'from vX -> windows: install -> update' per leg. Nothing is
hardcoded in the workflows: the tag workflow calls the generator
itself.
2. Native skips move to the point that owns the capability knowledge:
macOS combos (no driving workflow exists) grey out in the tag
workflow via install-e2e-skip.yml, untouched by the tag axis; ALL
windows combos dispatch to install-e2e-windows-run.yml, which takes
install-method/update-method inputs and natively skips the pairs
its driver cannot run yet - so implementing a windows method is a
change in the run workflow + driver only. The driver's -Route ids
now match the generator's method ids verbatim.
Generator output shape, route filters, and all error paths re-verified
locally; all four workflows pass actionlint; driver re-parses clean
pure-ASCII.
Bring back the continuous recording the retired AHK-only driver had,
alongside the 3s frame captures: gdigrab 15fps to proof/<phase>/
recording.mkv, started before the installer/app launches and q-stopped
in the finally block win or lose. mkv stays playable when the process
dies unfinalized; skip gracefully when ffmpeg is not on PATH (it ships
on windows-latest). Lifecycle (start, null-skip, graceful q stop, clean
exit, playable output) verified locally with a lavfi source.
Run 31520559900: stage passed the whole multiline tag list to
rev-parse ('Filename too long'). In
(Invoke-Git @(...) -split pattern | ...) PowerShell parses -split as
ANOTHER ARGUMENT to Invoke-Git, not as an operator on its result - the
function got '-split' and the regex appended to its array and rev-parse
received every tag at once. Split via an intermediate variable instead.
Verified locally: pwsh picks [v0.20.2] and resolves it to a single
commit.
Run 31520267702 died in 3s: 'Missing an argument for parameter
InstallRef'. powershell.exe -File drops a "" argument from the command
line entirely, so the parameter binder saw -InstallRef followed by
-SetupExeUrl. Default both the workflow input and the script parameter
to 'auto' (= newest release tag) instead of empty.
Run 31519103491 failed the 'update genuinely available' assert with
the installer landing on HEAD itself. The staging assumed the website
exe installs a baked release pin, but the bootstrap log shows
Pin { commit: None, branch: main } - the published installer installs
whatever main serves, and serve.git's main was parked at HEAD.
Stage the way the linux axis does: park served main at OLD
(-InstallRef, default newest release tag; threaded through the
reusable workflow as install-ref) for the install phase, assert the
install lands exactly there, then advance main to HEAD in the update
phase - an update becomes available the same way it does for a real
user. allowAnySHA1InWant stays as belt-and-braces for installer builds
that DO bake a pin.
Restructure tek's two-job desktop-windows-e2e.yml into the shape the
linux axis already has: install-e2e.yml keeps its update/installer
routes untouched and windows-desktop returns as a route in the same
family, calling a reusable install-e2e-windows-run.yml.
Behind that route is now ONLY the real user flow - the headless
contract job (install.ps1 at HEAD~1, desktop-update.ps1 -NoUi,
BASE/CURRENT/NEXT ref dance) is gone, along with its driver. Every leg
goes through a surface a user touches: website Hermes-Setup.exe headed
with AutoHotkey clicking Install -> Launch, then the installed
Hermes.exe under Playwright's Electron driver clicking Settings ->
About -> 'Update now', through the detached hand-off to a relaunched
window asserted on HEAD.
The driver drops the synthetic-NEXT staging with the contract job:
serve.git just serves HEAD as main and OLD is the release pin baked
into the website exe - the literal starting point of every real GUI
user, same philosophy as the linux axis's release-tag matrix. The
-Route parameter (desktop today) declares the future update mechanisms
as arms: 'update' (hermes update from the installed venv) and
'installer' (re-run the bootstrap exe) raise until implemented, so the
workflow surface is stable when they land.
The cherry-picked desktop-windows-e2e.yml covers everything the
install-e2e-windows-run.yml axis did and more: the contract job drives
the same desktop-update.ps1 hand-off (plus a CURRENT->NEXT forward
leg), and the GUI job replaces AHK-only driving with the full real
user flow - website Hermes-Setup.exe, clicked Install/Launch, then
Playwright clicking Settings -> About -> 'Update now' in the packaged
app, through the detached hand-off to a relaunched window.
Remove the superseded workflow, its driver, and the AHK/button assets
under tests/install/windows/ (the GUI job's e2e-assets carry the
re-captured templates), and drop the windows-desktop route from
install-e2e.yml's dispatch options.
The real-user-flow job passed end to end (run 31492613931, 10m57s):
website Hermes-Setup.exe installed headed (AHK Install+Launch, real app
window), then TWO GUI updates driven by real Settings -> About ->
'Update now' clicks, each carried through the detached hand-off to a
relaunched desktop on the target commit. Every assertion green on both
legs (marker cleanup, checkout on target sha, working hermes, relaunch).
Two finishing touches:
* Foreground the relaunched Hermes window before the 99-relaunched proof
screenshot — the full-desktop grab is z-order dependent and one run
caught VS Code on top. The relaunch ASSERT already passed on the
process signal; this is purely to make the proof image show Hermes.
* Remove the temporary branch push trigger used for pre-merge validation;
back to main + nightly + release tags + manual dispatch only.
Attempt 10 ran 1h40m and the diagnosis is precise: the GUI update hung on
a bare input() in hermes update's _sync_with_upstream_if_needed. Our
serve.git origin is a file:// URL, so _is_fork() is true and the updater
asks 'Add official repo as upstream? [Y/n]' via raw input(). When the
Desktop spawns the hand-off through 'cmd start /min' that child has a real
but EMPTY console, so input() blocks forever (no EOF, no keystroke). The
contract job spawns the hand-off with inherited non-interactive stdin, so
input() hits EOF and defaults immediately -- which is why it never hung.
The proof chain confirmed everything else worked: backend exited, venv
unlocked, git pull found the commit and applied it; the process then just
sat in input(). update.log was never created because the hang is BEFORE
the desktop-build step.
Fix: create HERMES_HOME/.skip_upstream_prompt after install -- the
product's own 'don't ask about upstream' marker (_should_skip_upstream_
prompt). Real GUI users install from the official github origin where
_is_fork() is false and this prompt never fires, so this only neutralizes
a staging artifact of the file:// serve repo, not real behavior.
(Noted for a separate product follow-up: hermes update --gateway should
route this input() through _gateway_prompt like its other prompts, so a
fork-origin GUI update can't hang even without the marker.)
Attempt 9 drove the full real update through the hand-off: marker
detected, desktop exited, hermes update fetched from serve.git, found the
commit, pulled, restored -- all correct. It then timed out because the
updater legitimately runs LONG here: the website release we install
(v0.20.0) is weeks of main behind CURRENT, so the update pulls a large
diff AND does a full Electron desktop rebuild (vite + electron-builder)
plus uv sync. The contract job's BASE->CURRENT is a 1-commit tests-only
diff that skips the rebuild, which is why it finishes in ~1 min; the GUI
job's release->CURRENT does not.
* wait window 40 -> 90 min per leg; job timeout 180 -> 240 min
* tail logs/update.log during the wait so the desktop-rebuild phase is
visible in CI output instead of tens of minutes of silence (the rebuild
streams there, not to the handoff log)
The CURRENT->NEXT leg stays fast (NEXT is a same-tree child of CURRENT,
no rebuild), so total stays well within 240 min.
Attempt 8 drove the ENTIRE GUI update click-path successfully: onboarding
dismissed, Settings opened, About opened, Update now clicked, updating
overlay shown. The hand-off log proves the real update then ran: desktop
(pid 8880) exited, venv unlocked, 'hermes update --yes --gateway --force
--branch main' fetched from serve.git, found 1 new commit, pulled, and
restored. Everything worked.
The only failure was the driver waiting on Playwright's app 'close'
event, which doesn't fire reliably when the Electron app self-quits for
the hand-off. Switch to the authoritative signal: poll for the
HERMES_HOME/.hermes-update-in-progress marker (or the result JSON, or a
genuine window-gone), which the hand-off writes ~4s after the click. The
PowerShell driver still owns asserting the OUTCOME (target sha, marker
cleanup, working hermes, relaunched app) after the driver returns.
Attempt 7 got the whole way into the GUI update leg: the installed
Electron app launched under Playwright, booted, composer attached, first
screenshot captured. It then couldn't find the settings gear -- the
ERROR screenshot showed why: a fresh install with no CONFIGURED provider
(the seeded .env key isn't read as model.provider) shows the onboarding
card ('Let''s get you setup with Hermes Agent'), which covers the shell
and its settings gear.
The update path needs no provider, so the driver now clicks 'I'll choose
a provider later' (with skip fallbacks) to dismiss onboarding and reach
the shell before looking for the gear. Harmless no-op when onboarding
isn't shown. Gear (aria-label 'Open settings') and About nav ('About')
selectors already match the real components.
Install + GUI update leg now reached (attempt 6): full install passes,
first update leg begins. It tripped a preflight assert checking
apps/desktop/node_modules/@playwright/test — but the root npm ci HOISTS
workspace devDependencies to the REPO-ROOT node_modules, so that path is
empty by design. Node's own resolution walks up from apps/desktop and
finds it (which is exactly how the copied-in drive-update.cjs will load
it), so assert via 'node -e require.resolve(...)' from apps/desktop
instead of a hardcoded nested path.
Windows PowerShell 5.1 reads .ps1 without a BOM under the legacy OEM
codepage, mis-decoding UTF-8 bytes. Em-dashes/box-drawing survived in
comments through attempts 2-4, but the previous commit added an em-dash
INSIDE a double-quoted Write-Host string — the misdecode there ate the
quote boundary and cascaded into a whole-file parse failure at the Stage
step ('Unexpected token', 'string is missing the terminator').
scripts/install.ps1 documents this exact constraint ('pure ASCII for PS
5.1 parser compatibility'). Strip all non-ASCII from the .ps1 and .ahk
files (em-dash->--, arrows->->, box-drawing->-). drive-update.cjs keeps
UTF-8 (Node decodes it natively). Both PowerShell files parse clean.
The full GUI install flow now works end-to-end (attempt 4 proof: Install
clicked, bootstrap complete, Launch clicked, real Hermes.exe window
appeared 1024x720, installer exited, 5 Hermes processes running). The
only failure was an over-strict staging assertion.
The website Hermes-Setup.exe pins a main release commit. On a real
push-to-main run CURRENT is main's tip, so that pin is its ancestor and
the check holds. On a diverged feature branch CURRENT is a branch commit
the release pin is not an ancestor of — a legitimate topology, not a bug.
The update leg resets the checkout to serve.git's main ref (= CURRENT)
regardless of ancestry and asserts it lands there, which is the actual
forward-update proof. Downgrade the ancestor check to an informational
note so branch validation can exercise the update legs.
Attempt 3's proof frames showed the install SUCCEEDED end-to-end
(bootstrap complete, installer self-copied to HERMES_HOME) and the
window advanced to 'HERMES IS READY' with a [ LAUNCH ] button at the
same centered CTA spot the [ INSTALL ] button occupied — screen (511,454)
inside window x=64 y=34 w=896 h=659.
Two Launch-step bugs, both fixed from that evidence:
* launch-button.png was the stale #68183 template and never matched the
restyled '[ LAUNCH ]' button. Re-captured from the live frame.
* the window-relative fallback used fy=0.59, clicking y=422 — above the
real button. Correct fraction is (454-34)/659 = 0.637. With the
template now matching, the fallback is belt-and-braces anyway.
Install click, completion detection, and the app-window wait were all
already correct in attempt 3; only the Launch click missed.
Attempt 2's proof frames showed two bugs, both now fixed from the live
evidence:
1. The #68183 install-button.png predated the installer UI restyle to the
'[ INSTALL ]' bracket look, so the template never matched and we fell
through to the position fallback. Re-captured install-button.png from a
real CI desktop frame (the actual rendered button).
2. The fallback then clicked the WRONG spot: ahk_exe's first WinGetPos
matched a hidden 16x16 helper window ('Window found at w=16 h=16' in
ahk.log), and BTN_FY=0.87 aimed below the real button anyway. The
button center measured at ~(0.50, 0.59) of the ~full-screen window.
Rewrite:
* WaitForRealWindow() skips phantom/hidden matches (requires w>400,h>300)
and returns the true rect; the installer window is then activated before
any click.
* Install-finished is now driven primarily by the authoritative
'bootstrap complete' line in bootstrap-installer.log (matches
BootstrapEvent::Complete), with the Launch template as a secondary
signal and a window-relative fallback click.
* Fallback clicks use the corrected (0.50, 0.59) window fraction.
Frame-0005 of the proof capture showed the exact failure: 'Unhandled
error: (6) The handle is invalid' rendered over the installer within
seconds of launch. AutoHotkey started via Start-Process has no console,
so FileAppend to '*' (stdout) throws — and the throw fired inside Log(),
killing the script before it clicked anything. The installer then sat
untouched at the INSTALL screen for 50 minutes.
* Log() now try-wraps the stdout write (file log is the real record)
* Install/Launch clicks fall back to the button's relative window
position when the #68183-era PNG templates don't match the restyled
UI ('[ INSTALL ]' bracket style visible in the same frame)
* install-finished has a second signal: 'bootstrap complete' in
bootstrap-installer.log (read with write-sharing), so a template miss
can't strand the wait
* driver passes the bootstrap log path as arg 3
Second job on the Windows E2E workflow covering the surfaces a user
actually touches, per Teknium's requirement:
* INSTALL: downloads the production Hermes-Setup.exe from
hermes-assets.nousresearch.com, launches it HEADED, and AutoHotkey
clicks Install -> waits -> clicks Launch (button templates + ImageSearch
approach from @ethernet8023's #68183, retargeted by process name and
extended to exercise the Launch hand-off). The real Electron Hermes.exe
window must appear.
* UPDATE x2: the installed Hermes.exe is launched under Playwright's
Electron driver and the test CLICKS Settings -> About -> Update now.
The production hand-off chain runs untouched: app quits, detached
updater (repo script or staged binary) runs hermes update, rebuilds
the desktop, relaunches Hermes.exe. Asserts: target sha, marker
cleanup, result JSON when the script path wrote one, working hermes,
and the RELAUNCHED app window. Leg 1 -> CURRENT, leg 2 -> synthetic
NEXT.
Proof artifacts: per-step renderer screenshots (booted app, settings,
About panel, update-available, updating overlay), full-desktop frames
every 3s across the whole run, ahk.log, bootstrap-installer.log,
desktop-update-handoff.log — uploaded on success AND failure.
The website exe runs exactly as shipped (its own pinned install.ps1,
its baked release-pin commit); the only environmental deltas are the
serve.git URL redirect, uploadpack.allowAnySHA1InWant for the commit
pin fetch, and a placeholder provider key so the update legs meet the
app shell instead of onboarding.
The contract job from the previous commits is unchanged and independent
— it remains the rollback position if the GUI job proves flaky.
First CI run's install leg cloned real GitHub main instead of the staged
BASE (caught by the HEAD-at-BASE assert): install.ps1 sets
GIT_CONFIG_COUNT=1 / windows.appendAtomically itself, silently clobbering
the driver's env-config insteadOf rewrites. A driver-owned gitconfig file
selected via GIT_CONFIG_GLOBAL survives that (and install.ps1's own
--global writes land harmlessly in the same file). Verified locally by
cloning with the clobber vars set: clone lands on staged BASE.
Every commit on main now proves, on a real Windows machine, that:
1. the PRIOR commit (HEAD~1) installs from scratch through its own
scripts/install.ps1 (-IncludeDesktop: uv, managed Python, Node,
venv, packaged Electron Hermes.exe),
2. that install updates TO this commit through the real Desktop GUI
update path (scripts/desktop-update.ps1, the exact hand-off the
Update button spawns -- fail-closed gates, marker lifecycle,
hermes update, result JSON), and
3. this commit updates FORWARD to a synthetic next commit, proving
the updater code shipping in this commit is not the one that
strands users when the next commit lands.
Staging: the driver bare-clones the checkout into serve.git and
redirects the canonical GitHub URLs at it with git insteadOf env
config, then advances the served main ref BASE -> CURRENT -> NEXT
between legs. Installer and updater run byte-for-byte unmodified.
Supersedes the AutoHotkey pixel-driving approach (#68183): the GUI
Update button's entire effect is spawning desktop-update.ps1 with
documented flags, so driving that contract directly tests the same
production code deterministically.
Run 31457301901 proved both halves of the stderr fix and then hung
anyway, in uv pip install -e .[all] (process table caught it live):
- The SQLite-repair uv sync that deadlocked run 2 now streams its
whole package list and completes in ~80s WITH debug tracing on -
managed_uv.py is imported lazily after the git reset, so even this
old base ran the fixed copy.
- The .[all] install runs _run_install_with_heartbeat from main.py,
which was imported when hermes update STARTED - the v0.20.1 copy,
which pipes uv's stderr undrained. RUST_LOG=uv=debug guaranteed
>64KB of stderr, so the driver's own diagnostic manufactured the
deadlock. Comments in both fixed call sites now state the real
import-time reach of each arm.
This also explains run 1 failing WITHOUT tracing: the pipe budget is
cumulative across every child of the update sharing it. The old-code
sync burned ~30KB of it on the package list; the .[all] leg finished
the job. With the sync leg now on stdout, the old .[all] leg's
natural output should fit - which is the real-world story too: old
bases hang or survive on stderr luck, new bases are safe by
construction.
The first full AHK run (31449642122) got all the way through install,
verify, and the desktop hand-off's git leg (fetch from the fake remote,
reset to target - the proxying works), then sat 65 minutes inside
'Updating Python dependencies' with zero uv output until the job
timeout cancelled it. Cancellation kills any chance of a post-mortem:
no process table, no partial log.
Run the hand-off through Start-Process with a driver-owned 45-minute
deadline (the same budget the staged-exe branch already gets), poll-tail
its log into the console, and on deadline dump the live uv/python/git
process table plus stderr tail before killing the tree - so a hang
diagnoses itself instead of burning another silent 75 minutes.
RUST_LOG=uv=debug is scoped to the update leg so uv says what it is
doing (or waiting on: cache lock, network, resolution).
install-button.png is now cut from run 31449192962's LOSSLESS
welcome-screen.png (blue-text extents x465-558 y449-459 plus 3px margin,
verified complete '[ INSTALL ]' with no foreign pixels) instead of the
H.264 recording whose chroma subsampling made every video-sourced crop
miss the live screen. Tolerance drops from *60 to *20 accordingly.
PixelSearch is deleted outright: color-hunting matched the blue title
(run 31446691812) and the progress view's stage text (run 31447405319)
before it ever matched the button. Click-landed detection reuses the
same ImageSearch (button visible = not clicked). The TEMP
stop-after-screenshot exit is removed - the AHK path is live again.
The ImageSearch reference must be cropped from a lossless capture of the
real screen: the ffmpeg recording is H.264/yuv420p and its chroma
subsampling shifts glyph pixels enough that a video-sourced crop never
matches live rendering. Taken before the helper starts so no tooltip or
click marker contaminates it; lands in the log-dir artifact.
Run 31447405319 is the big win and the bug in one log: uv seeding
worked, the Install click LANDED, and 6 stages ran (clone off the fake
remote via SSH rewrite, venv, all Python dependencies) - then the AHK
loop, still hunting for 'the button', PixelSearch-matched the PROGRESS
view's own blue stage text (left column, x~106-115), decided the UI
'did not advance' 10 times, threw, and the driver killed a healthy
install mid-node-deps.
Fixes, all sourced from that recording: narrow the scan band to the
center third so the left-column stage text can never match; verify a
click by the blue vanishing AT THE CLICK POINT (a 24px box) instead of
anywhere in the band; and never throw from the click loop - the
authoritative failure signal is the driver's 'bootstrap FAILED' log
abort, and the marker deadline caps a wedged UI.
Run 31447045981: the click landed (manifest received, stages ran) but
Stage-Uv failed with 'uv installed but not found at ...\bin\uv.exe'.
GitHub windows runners ship uv preinstalled WITH an astral install
receipt; astral's cargo-dist installer then updates the receipt's
location in place and ignores UV_INSTALL_DIR, so Install-Uv's managed
copy never appears. Seed HERMES_HOME\bin\uv.exe from the runner's uv
before launching the installer - Install-Uv short-circuits on an
existing managed uv, and 'user already has a managed uv' is a
legitimate install state, not a bypass.
Also abort the run the moment the tailed bootstrap log says
'bootstrap FAILED': the failure screen waits on a human Retry, and the
AHK helper would otherwise idle out its whole 25-minute marker
deadline (and its blue-text retry loop hammers the Retry button,
re-running doomed installs - observed in run 7).
Run 31446691812: every attempt logged 'blue text at 220, 330' - exactly
the 45%-height scan boundary, which lands inside the HERMES AGENT title
(title bottom ~47% of window height; button ~62%, measured from the run
2/5 recordings). The click-landed check then correctly reported no
advance, ten times. Raise the boundary to 55%, between the two.
Run 31446292343 disproved the z-order theory: the recording shows the
installer frontmost, red click-marker dots painting on it, and the button
rendered - yet ImageSearch missed on all 5 attempts. The reference crop is
the problem: it came from an H.264/yuv420p recording whose chroma
subsampling smears glyph edges. Diffing the crop against run 5's OWN
recording of the same screen gives max 8 shades/channel (matches easily),
so the crop is video-faithful but not screen-faithful, and no tolerance
fixes that reliably.
Keep ImageSearch as the first try, but fall back to PixelSearch for the
button text's saturated blue (~0x3B82F6, variation 90) in the window's
lower half - the only blue there (the title sits in the upper third).
Verify the click landed by the blue vanishing (the progress view replaces
the button); retry up to 10 times.
Run 31445907233's recording shows the runner session's maximized console
covering the installer for the whole run: WinWait matches by title
regardless of z-order, but ImageSearch reads screen pixels, so the Install
button was never visible to it. WinActivate + WinMoveTop before every
attempt; run 31443096241 already proved the same reference crop renders
match-ably when the window is frontmost.
Run 31445244722's recording shows the published Hermes-Setup.exe renders
'[ INSTALL ]' as flat blue text on off-white - nothing like the solid-blue
'Install Hermes ->' reference from the dev-build era, so ImageSearch never
matched. Replace the reference with a crop of the real button taken from
that recording (tolerance *60 to absorb H.264 drift, click retried across
animation frames), and drop launch-button.png entirely: completion now
polls the installer's own bootstrap-complete marker
(.hermes-bootstrap-complete, see paths.rs likely_bootstrap_marker), which
cannot go stale with a UI restyle.
AutoHotkey64 is a GUI-subsystem exe: spawned without -NoNewWindow it has
no console, FileAppend('*') throws '(6) The handle is invalid' on the first
Log call, and OnError's own Log rethrows inside the handler - the script
hangs with the error tooltip painted over the installer and Install is
never clicked (confirmed from the run 31443096241 screen recording; ahk.log
was never created because the stdout write preceded the file write).
Wrap the stdout append in try (the log file is the record) and spawn the
helper with -NoNewWindow so its live lines reach the job log.
windows sibling of install-e2e-run.yml. no bubblewrap on windows, so the
git proxying is git's own transport rewrite: an isolated GIT_CONFIG_GLOBAL
with multi-valued url.<file://fake.git>.insteadOf for both hardcoded repo
URLs, so the published Hermes-Setup.exe's install.ps1 clone, hermes update's
fetch, and the desktop's ls-remote all land on a local bare repo whose main
the driver controls - installer and updater run verbatim.
one run: seed fake.git from the checkout, force fake main to the newest
release tag, drive the real published bootstrap installer with AutoHotkey
(GUI, no headless mode), promote fake main to HEAD, then apply the desktop
app's builtin update route (scripts/desktop-update.ps1 -NoUi when the
installed base ships it, staged hermes-setup.exe --update otherwise) and
assert HEAD == target with a working hermes.
TODO routes: bare hermes update, and re-running the bootstrap installer
over the existing checkout.
Nothing covered the update path, which is the worst thing to break: a broken
updater strands users on the version that cannot fix itself. `hermes update`
alone is ~2000 lines (hermes_cli/update_cmd.py) and had no end-to-end test.
tests/install/install-update-e2e.sh installs a genuine earlier Hermes through
the real one-liner (curl -fsSL https://…/install.sh | bash, served by
dev-sandbox's MITM proxy at the canonical URL, cloning "github.com" through the
upload-pack shim), which really installs uv, a managed Python, Node and the
venv. It then applies ONE update route and requires the checkout to land on this
commit with `hermes --version` still working -- so a pass means the venv and
entry point survived, not merely that git moved.
One route per run, each on a sandbox built from scratch. Sharing one install
across routes -- or rewinding with `git reset --hard` between them -- leaves the
second route running against a tree the first already updated (same venv, same
console script, same __pycache__), which is not the state any real user is in: a
route could pass only because its predecessor did the work, and a failure in the
first left the second exercising something undefined.
--install-ref chooses what to install first, so this covers "update from an
older release", not just from the tip. Installer flags are probed against the
target rather than assumed, because releases from months back predate flags
current Hermes takes for granted: --skip-browser is read out of that ref's own
install.sh, and `--yes` is asked of the installed `hermes update --help` (the
update subcommand has lived in main.py, subcommands/update.py and update_cmd.py
across the tags we sample, so a static parse rots silently -- and did). Without
those probes, old releases die on "Unknown option: --skip-browser" and
"unrecognized arguments: --yes" before doing any work.
Installer output is streamed through tee rather than captured: a real install of
uv, Python, Node and the venv IS the substance of this test, so it belongs in
the job log, not only in an artifact. pipefail keeps the installer's exit status
rather than tee's, so a failed install cannot look like a pass. The sandbox's
own proxy log is printed in full on failure, since a rejected TLS handshake
explains a failure that otherwise reads as a bare `curl: (35)`.
Deliberately reuses dev-sandbox rather than adding a second harness. An earlier
draft rewrote install.sh's hardcoded URLs with insteadOf and ran it against the
host; that tested the installer LESS faithfully (bash install.sh instead of the
real one-liner, host libs instead of a clean machine, ssh disabled to keep a
failed rewrite from reaching real GitHub) while duplicating a fake Internet we
already have.
Shell, not pytest, so scripts/run_tests.sh and run_tests_parallel.py stay
untouched: a pytest version needed an entry in the former's `env -i` credential
allowlist and a _SKIP_PARTS exclusion in the latter, and every meaningful line
was a command run inside the sandbox anyway.
Two guards, both earned during bring-up. It prefers the `sandbox` wrapper and
falls back to the raw script only when bwrap is on PATH (under Nix the wrapper
supplies the PATH and DEV_SANDBOX_* vars, so the bare script exits 127). And it
refuses to run on a dirty worktree: every dev-sandbox invocation re-derives fake
main from the working copy, so uncommitted changes move the update target
between the call that installs and the call that verifies -- a failure that
looks like a broken updater but is a moving reference.