Proved on windows-latest that a locked profile blocks (no kill/hang), the
approved close terminates Chrome + releases the lock, and snapshot then copies a
valid DB — and autoclose-off blocks with quit guidance. Per policy proof
workflows never land on main. Product + portable unit tests remain.
Refines the Windows path per three requirements:
1. Only when the toggle is set — closing is offered only if
browser.real_profile_autoclose is on.
2. Blocked when locked — snapshot_real_profile NEVER kills; a locked profile
always returns the [profile-locked] signal and the copy is refused. A later
attempt that is still locked blocks again (no loop, no auto-kill).
3. Ask approval to close — closing is an explicit, user-approved step:
(new CLI subcommand) runs
close_browser_holding_profile only when the agent has the user's OK. The
locked error tells the agent to ask first, then run it, then retry.
- browser_connect: snapshot blocks with _PROFILE_LOCKED_PREFIX (autoclose-armed
message offers the close; off message says fully-quit); no in-snapshot kill.
- main.py: subcommand (identity+binding-verified
tree kill via close_browser_holding_profile); added to _BUILTIN_SUBCOMMANDS.
- browser_tool: surfaces the locked signal + the exact approved-close command.
- Docs/config: toggle arms + agent asks + blocked-if-still-locked.
Tests: snapshot blocks-not-kills with autoclose on AND off; process matcher
identity/binding. 73 real-profile tests pass. Windows live E2E (proof): locked
blocks fast without killing → approved close terminates Chrome → snapshot then
copies a valid DB; autoclose-off blocks with quit guidance.
Branch-only evidence — proved on windows-latest that consented auto-close
terminates a running Chrome, releases the lock, and produces a valid profile
copy (and that autoclose-off fails fast, not hangs). Per policy proof workflows
never land on main. Product fix + portable unit tests remain in
hermes_cli/browser_connect.py and tests/tools/test_browser_real_profile.py.
Live Windows CI proved copy-while-running is impossible (Chrome opens the cookie
DB deny-all). So to make Windows actually WORK — not just fail cleanly — add
opt-in auto-close: browser.real_profile_autoclose (default false). When the
profile is locked and consent is on, snapshot_real_profile terminates the
browser process tree bound to THAT user-data-dir (psutil, identity+binding
verified like the daemon reaper — browser binary AND this exact --user-data-dir
in cmdline, fail-closed on ambiguity), waits for the lock to release, then
snapshots. Destructive (loses unsaved tabs) so it's off by default and the agent
asks first; the fail-fast message names the option. No effect on POSIX.
- close_browser_holding_profile: graceful terminate → kill → poll until the
cookie DB is openable again (bounded); reports relaunch/tray failure clearly.
- _processes_holding_profile: identity+binding matcher (never kills an
unrelated same-name process on a different dir).
- Config key + docs admonition.
Tests: autoclose closes-then-snapshots, autoclose-failure-reports, fail-fast
names the option, process-matcher identity/binding. 74 real-profile tests pass.
Windows live E2E (PROOF workflow, reverted before merge): autoclose-off fails
fast <30s; autoclose-on terminates real Chrome, lock releases, valid cookie DB
copied.
The windows-latest proof E2E and its live/diagnostic tests were branch-only
evidence (they proved the deny-all lock + fast-fail contract on a real runner).
Per policy proof workflows never land on main. The product fix (fast lock
probe + fail-fast message) and its portable unit tests remain in
tests/tools/test_browser_real_profile.py.
Prior run hung 24min in the product-path test (snapshot_real_profile against a
locked profile blocks on Windows — itself a finding). Diagnostic now runs FIRST
(each strategy internally bounded, reports fast), live test second under a
faulthandler 150s dump-and-die so a hang can't burn the job. Job timeout 12min.
Previous runs were auto-cancelling each other (ref-scoped group + cancel-in-progress). Per-sha group lets each proof run finish so the diagnostic actually reports.
Adds a Windows-live diagnostic that, against a cookie DB held by a running
Chrome, reports which read strategy succeeds: shutil, open-rb, sqlite mode=ro,
sqlite immutable=1, sqlite ro+nolock, raw win32 CreateFile with full share
flags. This tells us empirically whether any in-process read path exists
(immutable=1 / share-all open) before reaching for VSS/admin. Fails-closed test
marked xfail while the real behavior is derived from the diagnostic.
One-shot windows-latest E2E: launches real Chrome on a user-data-dir so it holds
the cookie DB with a Windows share lock, asserts a RAW copy fails (WinError 32
precondition — else skip, no vacuous green), then asserts _copy_auth_file copies
it via SQLite online-backup and the result is a readable Cookies DB with the
cookies table.
This proves the Windows 'file in use' fix on a real runner — the coverage the
Linux lanes cannot provide. PROOF branch evidence only: this workflow + test are
reverted before merge and must never land on main.
The heal decision is extracted into should_heal_self_marker_refusal()
so the contract is testable: heal ONLY on exit 2 + a marker naming this
process. Five tests pin it — self-owned heals, foreign owner (real live
sibling process) never heals, missing/garbage marker never heals,
non-exit-2 never heals, and the full acquire -> refuse -> drop-claim ->
retry-precondition lifecycle with a real UpdateMarkerGuard.
windows-rust-e2e.yml mirrors the wine2e pattern: fires only on
wine2e-rust/** pushes, runs the crate's cargo test --lib on
windows-latest (the shipping platform). The permanent Linux lane stays
authoritative for the unix-gated pipe-drain fixtures.
`run_tests.sh` defaults to twice the core count, and the value this branch
started with came from a rule of thumb of 1.5x cores plus a measurement on a
16-core machine. A sweep on the real runner disagrees with both.
Run 32549672063 on the 96-core runner (EPYC 7763, 377GB) timed the whole suite
at six worker counts, two repetitions for each. A warmup run came first, and
retries were off:
workers x cores rep 1 rep 2 mean
48 0.5x 138s 139s 138s
96 1.0x 127s 126s 126s <- fastest
144 1.5x 130s 134s 132s
192 2.0x 132s 133s 132s
240 2.5x 140s 139s 140s
288 3.0x 143s 142s 142s
One worker for each core wins. Both repetitions agree on the order.
The shape is the more useful result. The range is 126s to 142s across a 6x
range of worker counts. The suite has sufficient concurrency at this machine
size, so nothing above the core count buys anything. The remaining time
belongs to the slowest individual files and to the setup. A future gain must
come from those, and not from this number.
The sweep ran from a temporary workflow that this branch does not keep.
Every Linux lane that does real work ran on a 4-core `ubuntu-latest`. The
Python suite and the JS checks were split into many small jobs to make that
size usable. Each split job repeated the full setup. In most of the JS jobs
the repeated setup cost more than the work.
The work lanes move to larger runners. Then the splits that existed only to
make small runners usable go away.
Python tests: 12 slices become 1 job on a 96-core runner. Slicing cost a
matrix job, a duration cache, a per-slice artifact and a merge job. 96 cores
clear the floor that the slowest single test file sets, which is about 82s. A
second slice divides work that is already at that floor, and adds a second
setup. Duration data from run 32522943054 gives the numbers behind this: 3178
files, 11645s in series.
The worker count is explicit, because `run_tests.sh` defaults to twice the
core count. A later commit sets it from a measurement on this hardware.
JS checks: 14 jobs become 1. The matrix paid about 371s of repeated setup to
spread about 612s of work. One larger runner installs one time. The three UI
shard scripts and `run-ui-shard.mjs` are therefore removed, because the
unsharded `test:ui` covers the same tests.
The unit of parallel work inside that job is a CHECK, and not a workspace.
apps/desktop is most of the payload, and its own `check` is a serial && chain.
A spread across workspaces alone therefore leaves that chain as the long pole.
A package that declares `check:*` sub-scripts gives one unit for each
sub-script. That is the same selection rule the matrix used.
The loop lives in `.github/scripts/run-workspace-checks.mjs`, so the same
sequence runs on a laptop. It runs 11 units together, buffers the output of
each one, and fails at the end with the full list. Children that share one
stdout interleave their lines and make a failure hard to read.
`npm run --ws check` stops at the first workspace that fails.
`check:test:plugins` joins the desktop `check` script. The matrix prefers
`check:*` sub-scripts over the plain `check` script, so `check:test:plugins`
ran only as its own leg. Without this change the merge drops that suite and
the job stays green.
node_modules is cached on the lockfile, and `npm ci` is skipped on an exact
hit. The `cache: npm` option of `setup-node` caches only the ~/.npm tarball
cache, which leaves the extract and the postinstalls to pay again.
The arm64 image build stays on a native arm64 runner. A build of linux/arm64
on an x64 host uses emulation.
The docker test lane caps its workers at the core count. Each of those tests
drives a container, so the docker daemon sets the limit and not the processor.
`.github/actionlint.yaml` declares the runner labels. actionlint knows the
GitHub-hosted labels only, and an undeclared label reads as an error that
hides the real findings.
The `detect` job checks out one file through a sparse checkout, and its
timeout drops to 1 minute. It reads
`scripts/ci/classify_changes.py` and nothing else.
Verification:
- actionlint reports 9 findings across all workflows. An unmodified HEAD with
the same config reports the same 9. This change adds none.
- A wrong label still fails. actionlint reports `ubuntu-latest-32-cor` and
`ubuntu-latest-32-arm-cores`.
- Every changed workflow parses, and `name` parses as a string.
- A replay of the `save-durations` merge step against a three-artifact layout
returns all 3178 entries.
- An expansion of the npm script graph gives the same leaf commands for the
parallel units and for a plain `npm run check`, in both directions. Against
the 13-leg matrix the count is 13 to 11, and the whole difference is the
three UI shards that collapse into one unsharded `check:test:ui`.
- `--list` reports the 11 units, and a full local run completes and reports
the time of each unit.
- The runner labels cannot be verified here. The first real run is the test.
On-demand workflow (fires only on wine2e/** pushes, never on PRs/main)
that runs a live venv-holder E2E on windows-latest: real spawned
processes with Hermes argv shapes, real detection/classification/
message code against the live process table. Tests pin CORRECT behavior
for the cluster issues (#90778 mislabeling, #78089 long-path exemption,
#87594 ancestor-exclusion, #81774 serve premise), so unfixed bugs fail
on the runner — empirical premise-check before the consolidation fix.
apps/bootstrap-installer/.gitignore excludes src-tauri/Cargo.lock — a
create-tauri-app scaffold default nobody revisited. With nothing tracked,
`--locked` fails outright ("cannot create the lock file ... because
--locked was passed") and the cache key hashed an absent file.
Keyed on Cargo.toml instead. The underlying gap — a signed installer that
re-resolves its whole dependency graph on every build, in a repo whose
pinning policy is otherwise strict — is noted in the workflow and left
for its own change rather than widening this one.
The lane shipped dead. `classify_changes.py` emitted `rust`, the composite
action re-exported it, and ci.yaml's `rust-tests` job gated on
`needs.detect.outputs.rust` — but the `detect` job never declared that
output, so the expression was the empty string and the job reported
"skipping" on the very PR that added it. GitHub does not error on a
reference to an output a job never declared, so nothing went red.
Adds the missing line plus the invariant that catches the whole class:
every `needs.detect.outputs.X` referenced by a job's `if` must be
declared by `detect`. Verified it fails with the line removed.
The related check — every lane reaching the composite action — is
separate on purpose: nix.yml and docker.yml own their triggers and
re-export different subsets, so `docker` and `nix` are legitimately not
ci.yaml detect outputs.
Nothing in CI compiled this crate. `.rs` lives under `apps/`, so the
change classifier matched a Rust edit as `frontend` and ran the
TypeScript matrix, which cannot notice a Rust error — the crate's 58 unit
tests had never executed once, and neither would the pipe-drain tests in
the previous commit.
Adds a `rust` lane and a Linux `cargo test --lib` job. Linux on purpose:
the pipe-drain fixtures need a real process tree whose grandchild
inherits the parent's stdout and are `#[cfg(unix)]`, so a Windows runner
would compile them out and report green over zero coverage. The Windows
half of that contract is `-SelfTestPipeDrain` on the existing Windows
lane.
The workflow owns its triggers and ci.yml does not call it. A
reusable-workflow call holds the caller run in progress for the full
build, and GitHub refuses `gh run rerun` on a run that is still in
progress. A separate run reruns and cancels on its own.
The job restores /nix/store from the GitHub Actions cache and saves from
main only. A cache that a PR writes is visible to that PR alone, so a
save there spends the quota of the repository and helps no later run.
The docker.yml gate held its own copy of the build formula, in shell.
classify_changes.py now owns a derived docker lane, and the nix lane in
the next commit derives from the same file. Two formulas in two
languages drift apart, and one Python function with tests does not.
Review findings on the caching commit:
- ~/.cache/electron was dead weight: with npm ci skipped on an exact
cache hit, the download cache is never read (electron's unpacked
binary lives in node_modules/electron/dist, inside the cached tree);
it only inflated every saved archive by ~110MB.
- 'npm i -g npm@12' ran unconditionally in all 14 matrix jobs
(~5-15s each); now a no-op when the bundled npm is already 12.x,
which also keeps the installed major aligned with the npm12
cache-key tag.
yaml + actionlint pass.
Every job in the js-tests matrix (~10 jobs/run, 13 after the UI-suite
sharding) runs a full 'npm ci' that deletes and re-extracts the entire
workspace node_modules and reruns all postinstalls — including the
Electron binary fetch (~100MB) — because setup-node's 'cache: npm' only
caches the ~/.npm tarball cache.
Cache the installed tree itself with actions/cache (the SHA-pinned
v4.2.4 already used by e2e-desktop.yml), keyed on the exact lockfile
hash, and skip 'npm ci' on a hit:
- key includes runner.os + node26 + npm12 so a toolchain bump never
reuses a stale tree
- NO restore-keys: a partial hit would leave a stale tree ('npm ci'
skipped means nothing would repair it), so anything but an exact
lockfile match reinstalls from scratch
- distinct keys for the discovery job (--ignore-scripts tree) and the
check jobs (with-scripts tree + ~/.cache/electron), which differ in
postinstall artifacts
Measured from run 31783969717: the npm-ci step is 30-45s per check job.
On warm cache this drops to a few seconds of restore, saving roughly
5-8 runner-minutes per PR run and ~1GB of registry traffic, and taking
~35s off every job on the merge-gate critical path.
Three separate reds on main. Two are fixed here; the third needs no code.
1. tests/gateway/test_multiplex_busy_input_mode.py (blocks every merge)
Fails "Python tests / Run tests slice 5/12" and therefore "All required
checks pass". Semantic merge conflict between two PRs merged ~1h apart:
a31be480 fix(gateway): respect routed profile busy modes (added the test)
c8f235a1 feat(gateway): allow selective multiplex profile serving (added the gate)
c8f235a1 taught _profile_name_for_source to reject a route whose target
profile is not in the served set (profiles_to_serve). Each PR was green on
its own base; neither ran against the other's merge result.
The test asserts a route to profile "research" resolves to that profile's
busy mode, but never patches profiles_to_serve — so it reads the runner's
REAL on-disk profiles. "research" is not among them, the route is rejected
before the busy-mode snapshot is consulted, and the assertion gets the
gateway default:
WARNING gateway.run: Rejecting profile route 'research-chat':
target profile 'research' is not served
AssertionError: assert 'interrupt' == 'steer'
Patch profiles_to_serve for the assertion — the same seam every sibling
test in tests/gateway/test_profile_resolution.py already patches
(test_route_inside_allowlist_resolves, test_route_outside_allowlist_rejects).
This also removes an ambient-state dependency: the test previously passed
or failed based on which profiles happened to exist on the machine running
it. Verified passing under an empty HERMES_HOME.
Test-only. The serving gate from c8f235a1 is correct and left intact.
2. Skills-index workflows: local action used without actions/checkout
check-freshness has failed on all 12 of its last 12 scheduled runs:
##[error]Can't find 'action.yml', 'action.yaml' or 'Dockerfile' under
'.../.github/actions/get-app-token'. Did you forget to run
actions/checkout before running your local action?
./.github/actions/get-app-token is a LOCAL composite action and cannot
resolve without the repo on disk. skills-index-freshness.yml had no
checkout step at all. The step is gated on `status != 'ok'`, so the
watchdog broke exactly when it was supposed to file its issue — the live
index is currently 521.4h stale (limit 26h) and nobody was told.
An audit of all workflows for this bug class found one more instance:
skills-index.yml's `trigger-deploy` job, which re-triggers the docs deploy
so a refreshed index reaches the live site. Its sibling `build-index` job
checks out; this one did not. That is plausibly why the index went stale
in the first place. Both are fixed; the audit now reports zero remaining
jobs that use a local action without a prior checkout.
Pinned to the same actions/checkout SHA used by the other 35 call sites.
3. "Publish inline E2E evidence" — no fix needed
Failed once at 13:33Z on a transient TLS error reaching api.github.com
("certificate is not valid for any names") while installing a gh
extension. The last 25 runs of that workflow are 25/25 success. Infra
blip, not a code defect.
The publisher read the PR number from the CI run's pull_requests
payload. GitHub keeps that payload empty for fork runs, so the job
printed 'No pull request is associated' and stopped on every fork PR.
Resolve the PR from the run's head owner, branch, and SHA instead.
The SHA match skips runs that a newer push superseded.
A fork PR also has no CI review comment, because the live poller
skips forks. The publisher now logs this and exits clean instead of
raising; the evidence stays in the workflow artifact.
Each test slice uploads an artifact with the same file name,
test_durations.json. The save-durations job downloaded the 12
artifacts with merge-multiple, so all extractions wrote to one
path in parallel. This caused two faults:
- A race between two extractions wrote two JSON documents into
one file. The merge step then failed with 'JSONDecodeError:
Extra data' (run 31382130252).
- On green runs, the last write erased the other 11 slices. The
merged cache held ~230 of ~2760 file durations.
Remove merge-multiple so each artifact extracts into its own
directory, and point the glob at durations/*/test_durations.json.
A local merge of the 12 real artifacts from the failed run gives
2761 durations.
The poller job set GITHUB_RUN_ID in env: to point at the CI run.
The Actions runner sets the GITHUB_* defaults itself and ignores
the override. Thus the poller read its own run id and watched
itself. Its own run stays in_progress while the poller runs, so
runs_all_completed() was never true. The comment froze at
'waiting for jobs to start' and the job burned its full 3000s
timeout on every PR.
Rename the variable to CI_RUN_ID. Also drop the GITHUB_REPOSITORY
override — it was a no-op for the same reason, and the runner
default already holds the correct value.
The dep-version-gate ruleset requires a team review for package
manifests, eslint configs, and workflow files. If the autofix patch
contains one of these files, the bot PR waits for that review and
auto-merge stops. The patch step now excludes them, so a bot PR
never gates itself. The eslint check in typecheck.yml still reports
their lint errors.
The requested trigger fires when GitHub creates the run. A run from a
first-time contributor waits in action_required, and the poller then
polls a run that never starts until its timeout. The in_progress
trigger fires when the run starts, and it also fires on a re-run.
The concurrency group now contains the head repository. Fork PRs
frequently share a branch name, and two PRs must not cancel the
poller of each other.
`shell: bash` runs the step with -e injected, and `set -uo pipefail` does
not clear it. A non-zero pytest exit killed the script before `status=$?`,
so the -eq 5 branch and its ::error message never ran. The job still failed
red, but the diagnostic that names the cause never printed.
the markers from the previous commit skip off-host. without a host to
run them on, every marked test is a silent skip. this commit adds the
hosts.
- tests-os.yml runs -m macos_only on macos-latest and -m windows_only
on windows-latest. ci.yml requires both lanes in all-checks-pass.
- a lane fails on pytest exit code 5 (zero tests selected). a renamed
marker cannot produce a green job that ran nothing.
- each lane repeats 'not integration' because a command-line -m
replaces the addopts filter.
- scripts/ci/list_os_marked_tests.py selects which files each lane
imports. -m filters after collection, and collection imports every
module. without this helper, one unrelated ImportError on the
foreign host fails a job whose own tests passed. the helper exits
non-zero when a marker matches no file, and writes bytes with
explicit lf so windows crlf translation cannot corrupt the bash
file list. it has its own tests in tests/ci/.
- the local runner now reports the skipped count and prints a note:
macos_only/windows_only tests were skipped on this host, and this
ci lane runs them. a green local run on linux no longer reads as
coverage of the other hosts.
- the runner default job count is now #cpu, not #cpu*2.
The CI run stayed in progress until its last job ended. Two advisory jobs
set that time: the review-comment poller (40 minutes) and the Docker image
build (45 minutes). Neither job was required to merge.
GitHub refuses `gh run rerun` on a run that is in progress. Thus a reviewer
who added the `ci-reviewed` label had to wait for the two slow jobs, and
label-rerun.yml carried a 2100-second wait loop for this reason. The fast
required jobs were ready long before.
Each slow job now runs in its own workflow:
- docker.yml owns its `pull_request` trigger and does its own change
detection. The new `detect` job runs the same composite action with the
same condition that ci.yml applied, so a tests-only PR still skips the
build. The `workflow_call` trigger is gone.
- ci-review-comment.yml starts on `workflow_run` when CI starts. It reads
the workflow and the scripts from the default branch, which is the trust
boundary that the old job got from its `ref: default_branch` checkout.
The poller reads job results through the API, so it can report on a run
that it does not belong to. `WATCH_WORKFLOWS` names sibling workflows for
the same commit, and `select_watched_runs` keeps the newest run for each
name. Thus the comment still shows the Docker results. The list is
newline-separated, because a workflow name can contain a comma.
The poller always exits 0 now. It reports on the CI run from a different
run, so a failed CI job is not a failure of the poller. The CI run has its
own gate for that.
Also correct a parse error in label-rerun.yml. STATUS came from the already
truncated RUN_ID, so its value was the run id and never "completed". Thus
the wait branch always ran.
ci.yml no longer needs `packages: write`, because the image build has left.
Pages serves exactly the newest artifact, so every push-triggered deploy
deleted the previous build's content-hashed JS/CSS while edge caches
(max-age=300, stale-while-revalidate=3600) kept serving HTML that
referenced them. With deploys landing every ~15-30 min, docs pages spent
most of the day pointing at 404'd bundles — search (pure client JS) was
the loudest casualty.
Fix: keep a rolling 14-day pool of hashed assets (en + zh-Hans) in the
Actions cache and union-merge it into each deploy artifact, current
build authoritative on collision (cp --update=none). Stale HTML and
already-open tabs now keep resolving across any number of deploys.
The Docusaurus build step (198s) is 79% of the docs-site-checks job
wall time and is the CI critical path. The site has two locales (en +
zh-Hans, ~700 pages total); building only the default locale in PR
checks halves the build time.
Follows the same pattern Docusaurus uses internally: a build:fast
script that runs `docusaurus build --locale en` for CI/preview builds,
while the full bilingual build runs only in deploy-site.yml for
production deploys.
deploy-site.yml is unchanged — it still runs `npm run build` (all
locales) on push-to-main and release.