Commit Graph

241 Commits

Author SHA1 Message Date
Gille 3e1629009a fix(install): reject incompatible system npm 2026-08-28 05:05:36 -07:00
Teknium 547f4c952d chore: remove single-use Windows live-E2E proof workflow (run 33157529194 green) 2026-08-28 02:59:17 -07:00
Teknium 2a4278209f ci: single-use Windows live-E2E proof workflow [proof do-not-merge] 2026-08-28 02:59:17 -07:00
Teknium 824f7e081a chore: remove Windows real-profile PROOF workflow + live test
Proved on windows-latest that a locked profile blocks (no kill/hang), the
approved close terminates Chrome + releases the lock, and snapshot then copies a
valid DB — and autoclose-off blocks with quit guidance. Per policy proof
workflows never land on main. Product + portable unit tests remain.
2026-08-26 19:25:33 -07:00
Teknium e4451ec6e5 feat(browser): close-with-approval flow for Windows real-profile (toggle arms, agent asks, blocked if still locked) [proof do-not-merge]
Refines the Windows path per three requirements:
1. Only when the toggle is set — closing is offered only if
   browser.real_profile_autoclose is on.
2. Blocked when locked — snapshot_real_profile NEVER kills; a locked profile
   always returns the [profile-locked] signal and the copy is refused. A later
   attempt that is still locked blocks again (no loop, no auto-kill).
3. Ask approval to close — closing is an explicit, user-approved step:
    (new CLI subcommand) runs
   close_browser_holding_profile only when the agent has the user's OK. The
   locked error tells the agent to ask first, then run it, then retry.

- browser_connect: snapshot blocks with _PROFILE_LOCKED_PREFIX (autoclose-armed
  message offers the close; off message says fully-quit); no in-snapshot kill.
- main.py:  subcommand (identity+binding-verified
  tree kill via close_browser_holding_profile); added to _BUILTIN_SUBCOMMANDS.
- browser_tool: surfaces the locked signal + the exact approved-close command.
- Docs/config: toggle arms + agent asks + blocked-if-still-locked.

Tests: snapshot blocks-not-kills with autoclose on AND off; process matcher
identity/binding. 73 real-profile tests pass. Windows live E2E (proof): locked
blocks fast without killing → approved close terminates Chrome → snapshot then
copies a valid DB; autoclose-off blocks with quit guidance.
2026-08-26 19:25:33 -07:00
Teknium 00d5632249 chore: remove Windows real-profile PROOF workflow + live test
Branch-only evidence — proved on windows-latest that consented auto-close
terminates a running Chrome, releases the lock, and produces a valid profile
copy (and that autoclose-off fails fast, not hangs). Per policy proof workflows
never land on main. Product fix + portable unit tests remain in
hermes_cli/browser_connect.py and tests/tools/test_browser_real_profile.py.
2026-08-26 19:25:33 -07:00
Teknium 9e9e1b2245 feat(browser): consented auto-close of a running browser for Windows real-profile [proof workflow do-not-merge]
Live Windows CI proved copy-while-running is impossible (Chrome opens the cookie
DB deny-all). So to make Windows actually WORK — not just fail cleanly — add
opt-in auto-close: browser.real_profile_autoclose (default false). When the
profile is locked and consent is on, snapshot_real_profile terminates the
browser process tree bound to THAT user-data-dir (psutil, identity+binding
verified like the daemon reaper — browser binary AND this exact --user-data-dir
in cmdline, fail-closed on ambiguity), waits for the lock to release, then
snapshots. Destructive (loses unsaved tabs) so it's off by default and the agent
asks first; the fail-fast message names the option. No effect on POSIX.

- close_browser_holding_profile: graceful terminate → kill → poll until the
  cookie DB is openable again (bounded); reports relaunch/tray failure clearly.
- _processes_holding_profile: identity+binding matcher (never kills an
  unrelated same-name process on a different dir).
- Config key + docs admonition.

Tests: autoclose closes-then-snapshots, autoclose-failure-reports, fail-fast
names the option, process-matcher identity/binding. 74 real-profile tests pass.

Windows live E2E (PROOF workflow, reverted before merge): autoclose-off fails
fast <30s; autoclose-on terminates real Chrome, lock releases, valid cookie DB
copied.
2026-08-26 19:25:33 -07:00
Teknium b73f78714a chore: remove Windows real-profile PROOF workflow + live tests
The windows-latest proof E2E and its live/diagnostic tests were branch-only
evidence (they proved the deny-all lock + fast-fail contract on a real runner).
Per policy proof workflows never land on main. The product fix (fast lock
probe + fail-fast message) and its portable unit tests remain in
tests/tools/test_browser_real_profile.py.
2026-08-26 19:25:33 -07:00
Teknium 52772de402 test(ci): run Windows diagnostic first + hard-bound the live test [do-not-merge]
Prior run hung 24min in the product-path test (snapshot_real_profile against a
locked profile blocks on Windows — itself a finding). Diagnostic now runs FIRST
(each strategy internally bounded, reports fast), live test second under a
faulthandler 150s dump-and-die so a hang can't burn the job. Job timeout 12min.
2026-08-26 19:25:33 -07:00
Teknium 92b95f16ce test(ci): proof workflow uses per-sha concurrency, no cancel-in-progress [do-not-merge]
Previous runs were auto-cancelling each other (ref-scoped group + cancel-in-progress). Per-sha group lets each proof run finish so the diagnostic actually reports.
2026-08-26 19:25:33 -07:00
Teknium bd504bee6d test(browser): PROOF diag — probe which read strategy beats Chrome's Windows lock [do-not-merge]
Adds a Windows-live diagnostic that, against a cookie DB held by a running
Chrome, reports which read strategy succeeds: shutil, open-rb, sqlite mode=ro,
sqlite immutable=1, sqlite ro+nolock, raw win32 CreateFile with full share
flags. This tells us empirically whether any in-process read path exists
(immutable=1 / share-all open) before reaching for VSS/admin. Fails-closed test
marked xfail while the real behavior is derived from the diagnostic.
2026-08-26 19:25:33 -07:00
Teknium 2ecb18c4a7 test(browser): PROOF — Windows live E2E for locked-DB real-profile copy [do-not-merge]
One-shot windows-latest E2E: launches real Chrome on a user-data-dir so it holds
the cookie DB with a Windows share lock, asserts a RAW copy fails (WinError 32
precondition — else skip, no vacuous green), then asserts _copy_auth_file copies
it via SQLite online-backup and the result is a readable Cookies DB with the
cookies table.

This proves the Windows 'file in use' fix on a real runner — the coverage the
Linux lanes cannot provide. PROOF branch evidence only: this workflow + test are
reverted before merge and must never land on main.
2026-08-26 19:25:33 -07:00
Teknium cb45a813c3 chore: drop the on-demand Windows Rust lane from the PR — wine2e workflows run from the pushed branch and never need to live on main 2026-08-25 21:59:19 -07:00
Teknium bc4ebcf3b0 test(bootstrap): pin the exit-2 self-marker heal contract + on-demand windows-latest Rust lane
The heal decision is extracted into should_heal_self_marker_refusal()
so the contract is testable: heal ONLY on exit 2 + a marker naming this
process. Five tests pin it — self-owned heals, foreign owner (real live
sibling process) never heals, missing/garbage marker never heals,
non-exit-2 never heals, and the full acquire -> refuse -> drop-claim ->
retry-precondition lifecycle with a real UpdateMarkerGuard.

windows-rust-e2e.yml mirrors the wine2e pattern: fires only on
wine2e-rust/** pushes, runs the crate's cargo test --lib on
windows-latest (the shipping platform). The permanent Linux lane stays
authoritative for the unix-gated pipe-drain fixtures.
2026-08-25 21:59:19 -07:00
ethernet 0012dd1e0c perf(ci): set python test workers to one for each core, from measurement
`run_tests.sh` defaults to twice the core count, and the value this branch
started with came from a rule of thumb of 1.5x cores plus a measurement on a
16-core machine. A sweep on the real runner disagrees with both.

Run 32549672063 on the 96-core runner (EPYC 7763, 377GB) timed the whole suite
at six worker counts, two repetitions for each. A warmup run came first, and
retries were off:

    workers   x cores   rep 1   rep 2   mean
       48       0.5x     138s    139s   138s
       96       1.0x     127s    126s   126s   <- fastest
      144       1.5x     130s    134s   132s
      192       2.0x     132s    133s   132s
      240       2.5x     140s    139s   140s
      288       3.0x     143s    142s   142s

One worker for each core wins. Both repetitions agree on the order.

The shape is the more useful result. The range is 126s to 142s across a 6x
range of worker counts. The suite has sufficient concurrency at this machine
size, so nothing above the core count buys anything. The remaining time
belongs to the slowest individual files and to the setup. A future gain must
come from those, and not from this number.

The sweep ran from a temporary workflow that this branch does not keep.
2026-08-22 02:25:12 -04:00
ethernet 10f99bc15e ci: run the work lanes on larger runners and merge the split jobs
Every Linux lane that does real work ran on a 4-core `ubuntu-latest`. The
Python suite and the JS checks were split into many small jobs to make that
size usable. Each split job repeated the full setup. In most of the JS jobs
the repeated setup cost more than the work.

The work lanes move to larger runners. Then the splits that existed only to
make small runners usable go away.

Python tests: 12 slices become 1 job on a 96-core runner. Slicing cost a
matrix job, a duration cache, a per-slice artifact and a merge job. 96 cores
clear the floor that the slowest single test file sets, which is about 82s. A
second slice divides work that is already at that floor, and adds a second
setup. Duration data from run 32522943054 gives the numbers behind this: 3178
files, 11645s in series.

The worker count is explicit, because `run_tests.sh` defaults to twice the
core count. A later commit sets it from a measurement on this hardware.

JS checks: 14 jobs become 1. The matrix paid about 371s of repeated setup to
spread about 612s of work. One larger runner installs one time. The three UI
shard scripts and `run-ui-shard.mjs` are therefore removed, because the
unsharded `test:ui` covers the same tests.

The unit of parallel work inside that job is a CHECK, and not a workspace.
apps/desktop is most of the payload, and its own `check` is a serial && chain.
A spread across workspaces alone therefore leaves that chain as the long pole.
A package that declares `check:*` sub-scripts gives one unit for each
sub-script. That is the same selection rule the matrix used.

The loop lives in `.github/scripts/run-workspace-checks.mjs`, so the same
sequence runs on a laptop. It runs 11 units together, buffers the output of
each one, and fails at the end with the full list. Children that share one
stdout interleave their lines and make a failure hard to read.
`npm run --ws check` stops at the first workspace that fails.

`check:test:plugins` joins the desktop `check` script. The matrix prefers
`check:*` sub-scripts over the plain `check` script, so `check:test:plugins`
ran only as its own leg. Without this change the merge drops that suite and
the job stays green.

node_modules is cached on the lockfile, and `npm ci` is skipped on an exact
hit. The `cache: npm` option of `setup-node` caches only the ~/.npm tarball
cache, which leaves the extract and the postinstalls to pay again.

The arm64 image build stays on a native arm64 runner. A build of linux/arm64
on an x64 host uses emulation.

The docker test lane caps its workers at the core count. Each of those tests
drives a container, so the docker daemon sets the limit and not the processor.

`.github/actionlint.yaml` declares the runner labels. actionlint knows the
GitHub-hosted labels only, and an undeclared label reads as an error that
hides the real findings.

The `detect` job checks out one file through a sparse checkout, and its
timeout drops to 1 minute. It reads
`scripts/ci/classify_changes.py` and nothing else.

Verification:
- actionlint reports 9 findings across all workflows. An unmodified HEAD with
  the same config reports the same 9. This change adds none.
- A wrong label still fails. actionlint reports `ubuntu-latest-32-cor` and
  `ubuntu-latest-32-arm-cores`.
- Every changed workflow parses, and `name` parses as a string.
- A replay of the `save-durations` merge step against a three-artifact layout
  returns all 3178 entries.
- An expansion of the npm script graph gives the same leaf commands for the
  parallel units and for a plain `npm run check`, in both directions. Against
  the 13-leg matrix the count is 13 to 11, and the whole difference is the
  three UI shards that collapse into one unsharded `check:test:ui`.
- `--list` reports the 11 units, and a full local run completes and reports
  the time of each unit.
- The runner labels cannot be verified here. The first real run is the test.
2026-08-22 02:25:12 -04:00
Teknium aefcf4d10a ci(windows-venv-e2e): drop --timeout (pytest-timeout not in dev-only sync) 2026-08-21 19:11:55 -07:00
Teknium f9aed7d7f6 test(windows): on-demand live venv-holder E2E lane + probe suite (#91277)
On-demand workflow (fires only on wine2e/** pushes, never on PRs/main)
that runs a live venv-holder E2E on windows-latest: real spawned
processes with Hermes argv shapes, real detection/classification/
message code against the live process table. Tests pin CORRECT behavior
for the cluster issues (#90778 mislabeling, #78089 long-path exemption,
#87594 ancestor-exclusion, #81774 serve premise), so unfixed bugs fail
on the runner — empirical premise-check before the consolidation fix.
2026-08-21 19:11:55 -07:00
Brooklyn Nicholson 4d2c546a53 fix(ci): drop --locked, the installer crate has no tracked lockfile
apps/bootstrap-installer/.gitignore excludes src-tauri/Cargo.lock — a
create-tauri-app scaffold default nobody revisited. With nothing tracked,
`--locked` fails outright ("cannot create the lock file ... because
--locked was passed") and the cache key hashed an absent file.

Keyed on Cargo.toml instead. The underlying gap — a signed installer that
re-resolves its whole dependency graph on every build, in a repo whose
pinning policy is otherwise strict — is noted in the workflow and left
for its own change rather than widening this one.
2026-08-20 12:03:00 -05:00
Brooklyn Nicholson dd1e5b723d fix(ci): declare the rust lane on the detect job, and test that wiring
The lane shipped dead. `classify_changes.py` emitted `rust`, the composite
action re-exported it, and ci.yaml's `rust-tests` job gated on
`needs.detect.outputs.rust` — but the `detect` job never declared that
output, so the expression was the empty string and the job reported
"skipping" on the very PR that added it. GitHub does not error on a
reference to an output a job never declared, so nothing went red.

Adds the missing line plus the invariant that catches the whole class:
every `needs.detect.outputs.X` referenced by a job's `if` must be
declared by `detect`. Verified it fails with the line removed.

The related check — every lane reaching the composite action — is
separate on purpose: nix.yml and docker.yml own their triggers and
re-export different subsets, so `docker` and `nix` are legitimately not
ci.yaml detect outputs.
2026-08-20 11:46:30 -05:00
Brooklyn Nicholson 5caea5e501 ci: run cargo test for the bootstrap installer
Nothing in CI compiled this crate. `.rs` lives under `apps/`, so the
change classifier matched a Rust edit as `frontend` and ran the
TypeScript matrix, which cannot notice a Rust error — the crate's 58 unit
tests had never executed once, and neither would the pipe-drain tests in
the previous commit.

Adds a `rust` lane and a Linux `cargo test --lib` job. Linux on purpose:
the pipe-drain fixtures need a real process tree whose grandchild
inherits the parent's stdout and are `#[cfg(unix)]`, so a Windows runner
would compile them out and report green over zero coverage. The Windows
half of that contract is `-SelfTestPipeDrain` on the existing Windows
lane.
2026-08-20 11:40:49 -05:00
Teknium dc77f2c87f Revert "ci: force workflow re-parse after zero-job dispatch window"
This reverts commit ab173e26d2.
2026-08-19 02:15:02 -07:00
Teknium 43e67f1f0e ci: move orchestrator to ci.yaml — ci.yml workflow identity is wedged (0-job startup_failure on every dispatch; identical content dispatches fine under a new path, proven by probe PR #89894) 2026-08-19 02:13:28 -07:00
Teknium ab173e26d2 ci: force workflow re-parse after zero-job dispatch window 2026-08-19 02:07:57 -07:00
Teknium 9162ea6db1 Revert "chore(ci): touch ci.yml to bust poisoned workflow-parse cache"
This reverts commit 0f73adb74f.
2026-08-19 01:22:51 -07:00
Teknium 0f73adb74f chore(ci): touch ci.yml to bust poisoned workflow-parse cache 2026-08-19 01:22:49 -07:00
ethernet 00c3872882 feat(ci): add nix flake check as unrequired job
The workflow owns its triggers and ci.yml does not call it. A
reusable-workflow call holds the caller run in progress for the full
build, and GitHub refuses `gh run rerun` on a run that is still in
progress. A separate run reruns and cancels on its own.

The job restores /nix/store from the GitHub Actions cache and saves from
main only. A cache that a PR writes is visible to that PR alone, so a
save there spends the quota of the repository and helps no later run.
2026-08-18 20:42:06 -04:00
ethernet 1dbe469276 refactor(ci): hoist docker detect-changes into the .py file
The docker.yml gate held its own copy of the build formula, in shell.
classify_changes.py now owns a derived docker lane, and the nix lane in
the next commit derives from the same file. Two formulas in two
languages drift apart, and one Python function with tests does not.
2026-08-18 20:42:06 -04:00
Teknium f30ea0eb9a Revert "chore: nudge ci.yml blob to bust poisoned workflow-parse cache"
This reverts commit 824409fbe511ba6c872cff2c972ce62a82ee5c4a.
2026-08-17 20:39:02 -07:00
Teknium f6667e7bf9 chore: nudge ci.yml blob to bust poisoned workflow-parse cache 2026-08-17 20:39:02 -07:00
Teknium 7e77d9bd71 Revert "ci: cache-bust workflow re-parse"
This reverts commit b363038c1d39805062f484806f8acf80ffc4d770.
2026-08-17 02:40:59 -07:00
Teknium 077a7d536a ci: cache-bust workflow re-parse 2026-08-17 02:40:59 -07:00
Teknium 1ed94d2452 Revert "ci: touch ci.yml to force workflow re-parse"
This reverts commit a7ba91e2e58d337c899eef2760dcc67f44d208c5.
2026-08-17 02:11:27 -07:00
Teknium e9e3291e71 ci: touch ci.yml to force workflow re-parse 2026-08-17 02:11:27 -07:00
kshitij de0abc0617 ci(js-tests): drop dead electron download-cache path, skip redundant npm upgrade
Review findings on the caching commit:

- ~/.cache/electron was dead weight: with npm ci skipped on an exact
  cache hit, the download cache is never read (electron's unpacked
  binary lives in node_modules/electron/dist, inside the cached tree);
  it only inflated every saved archive by ~110MB.
- 'npm i -g npm@12' ran unconditionally in all 14 matrix jobs
  (~5-15s each); now a no-op when the bundled npm is already 12.x,
  which also keeps the installed major aligned with the npm12
  cache-key tag.

yaml + actionlint pass.
2026-08-14 12:14:03 -04:00
kshitij f56a9a1185 ci(js-tests): cache the installed node_modules tree, not just the npm tarball cache
Every job in the js-tests matrix (~10 jobs/run, 13 after the UI-suite
sharding) runs a full 'npm ci' that deletes and re-extracts the entire
workspace node_modules and reruns all postinstalls — including the
Electron binary fetch (~100MB) — because setup-node's 'cache: npm' only
caches the ~/.npm tarball cache.

Cache the installed tree itself with actions/cache (the SHA-pinned
v4.2.4 already used by e2e-desktop.yml), keyed on the exact lockfile
hash, and skip 'npm ci' on a hit:

- key includes runner.os + node26 + npm12 so a toolchain bump never
  reuses a stale tree
- NO restore-keys: a partial hit would leave a stale tree ('npm ci'
  skipped means nothing would repair it), so anything but an exact
  lockfile match reinstalls from scratch
- distinct keys for the discovery job (--ignore-scripts tree) and the
  check jobs (with-scripts tree + ~/.cache/electron), which differ in
  postinstall artifacts

Measured from run 31783969717: the npm-ci step is 30-45s per check job.
On warm cache this drops to a few seconds of restore, saving roughly
5-8 runner-minutes per PR run and ~1GB of registry traffic, and taking
~35s off every job on the merge-gate critical path.
2026-08-14 12:14:03 -04:00
copilot-swe-agent[bot] 0b50c8e48f ci: wrap uv python install in retry action on OS test lane
Co-authored-by: OutThisLife <770929+OutThisLife@users.noreply.github.com>
2026-08-13 13:18:49 -05:00
brooklyn! 9eab7a4473 fix(ci): stop running uv lock --check on PRs that can't touch the lockfile (#84675) 2026-08-12 13:30:01 -05:00
kshitij f51aa6a9b5 fix(ci): repair red main — busy-mode test + missing checkout in skills-index workflows
Three separate reds on main. Two are fixed here; the third needs no code.

1. tests/gateway/test_multiplex_busy_input_mode.py (blocks every merge)

Fails "Python tests / Run tests slice 5/12" and therefore "All required
checks pass". Semantic merge conflict between two PRs merged ~1h apart:

  a31be480 fix(gateway): respect routed profile busy modes             (added the test)
  c8f235a1 feat(gateway): allow selective multiplex profile serving    (added the gate)

c8f235a1 taught _profile_name_for_source to reject a route whose target
profile is not in the served set (profiles_to_serve). Each PR was green on
its own base; neither ran against the other's merge result.

The test asserts a route to profile "research" resolves to that profile's
busy mode, but never patches profiles_to_serve — so it reads the runner's
REAL on-disk profiles. "research" is not among them, the route is rejected
before the busy-mode snapshot is consulted, and the assertion gets the
gateway default:

  WARNING gateway.run: Rejecting profile route 'research-chat':
                       target profile 'research' is not served
  AssertionError: assert 'interrupt' == 'steer'

Patch profiles_to_serve for the assertion — the same seam every sibling
test in tests/gateway/test_profile_resolution.py already patches
(test_route_inside_allowlist_resolves, test_route_outside_allowlist_rejects).

This also removes an ambient-state dependency: the test previously passed
or failed based on which profiles happened to exist on the machine running
it. Verified passing under an empty HERMES_HOME.

Test-only. The serving gate from c8f235a1 is correct and left intact.

2. Skills-index workflows: local action used without actions/checkout

check-freshness has failed on all 12 of its last 12 scheduled runs:

  ##[error]Can't find 'action.yml', 'action.yaml' or 'Dockerfile' under
  '.../.github/actions/get-app-token'. Did you forget to run
  actions/checkout before running your local action?

./.github/actions/get-app-token is a LOCAL composite action and cannot
resolve without the repo on disk. skills-index-freshness.yml had no
checkout step at all. The step is gated on `status != 'ok'`, so the
watchdog broke exactly when it was supposed to file its issue — the live
index is currently 521.4h stale (limit 26h) and nobody was told.

An audit of all workflows for this bug class found one more instance:
skills-index.yml's `trigger-deploy` job, which re-triggers the docs deploy
so a refreshed index reaches the live site. Its sibling `build-index` job
checks out; this one did not. That is plausibly why the index went stale
in the first place. Both are fixed; the audit now reports zero remaining
jobs that use a local action without a prior checkout.

Pinned to the same actions/checkout SHA used by the other 35 call sites.

3. "Publish inline E2E evidence" — no fix needed

Failed once at 13:33Z on a transient TLS error reaching api.github.com
("certificate is not valid for any names") while installing a gh
extension. The last 25 runs of that workflow are 25/25 success. Infra
blip, not a code defect.
2026-08-11 21:26:28 +05:30
ethernet 3ee1bb323e fix(ci): resolve fork PRs in the E2E evidence publisher
The publisher read the PR number from the CI run's pull_requests
payload. GitHub keeps that payload empty for fork runs, so the job
printed 'No pull request is associated' and stopped on every fork PR.

Resolve the PR from the run's head owner, branch, and SHA instead.
The SHA match skips runs that a newer push superseded.

A fork PR also has no CI review comment, because the live poller
skips forks. The publisher now logs this and exits clean instead of
raising; the evidence stays in the workflow artifact.
2026-08-11 03:53:52 -04:00
ethernet 03da1606bc fix(ci): merge all duration slices, not one
Each test slice uploads an artifact with the same file name,
test_durations.json. The save-durations job downloaded the 12
artifacts with merge-multiple, so all extractions wrote to one
path in parallel. This caused two faults:

- A race between two extractions wrote two JSON documents into
  one file. The merge step then failed with 'JSONDecodeError:
  Extra data' (run 31382130252).
- On green runs, the last write erased the other 11 slices. The
  merged cache held ~230 of ~2760 file durations.

Remove merge-multiple so each artifact extracts into its own
directory, and point the glob at durations/*/test_durations.json.
A local merge of the 12 real artifacts from the failed run gives
2761 durations.
2026-08-10 14:33:36 -04:00
ethernet d5ddd442d7 fix(ci): review comment poller deadlocked on its own run
The poller job set GITHUB_RUN_ID in env: to point at the CI run.
The Actions runner sets the GITHUB_* defaults itself and ignores
the override. Thus the poller read its own run id and watched
itself. Its own run stays in_progress while the poller runs, so
runs_all_completed() was never true. The comment froze at
'waiting for jobs to start' and the job burned its full 3000s
timeout on every PR.

Rename the variable to CI_RUN_ID. Also drop the GITHUB_REPOSITORY
override — it was a no-op for the same reason, and the runner
default already holds the correct value.
2026-08-10 13:48:21 -04:00
ethernet adecaf8086 fix(ci): unbuffer live comment poller output 2026-08-10 02:36:55 -04:00
ethernet 12299ca54c fix(ci): keep review-gated files out of the js-autofix patch
The dep-version-gate ruleset requires a team review for package
manifests, eslint configs, and workflow files. If the autofix patch
contains one of these files, the bot PR waits for that review and
auto-merge stops. The patch step now excludes them, so a bot PR
never gates itself. The eslint check in typecheck.yml still reports
their lint errors.
2026-08-10 02:36:55 -04:00
ethernet ee0f060a7d fix(ci): start the poller on in_progress, key concurrency per repo
The requested trigger fires when GitHub creates the run. A run from a
first-time contributor waits in action_required, and the poller then
polls a run that never starts until its timeout. The in_progress
trigger fires when the run starts, and it also fires on a re-run.

The concurrency group now contains the head repository. Fork PRs
frequently share a branch name, and two PRs must not cancel the
poller of each other.
2026-08-10 02:36:55 -04:00
ethernet a298dbcfe3 ci: print the zero-selection diagnostic instead of dying first
`shell: bash` runs the step with -e injected, and `set -uo pipefail` does
not clear it. A non-zero pytest exit killed the script before `status=$?`,
so the -eq 5 branch and its ::error message never ran. The job still failed
red, but the diagnostic that names the cause never printed.
2026-08-09 22:09:49 -04:00
ethernet 641d254db4 ci: add macos and windows test lanes for the os-marked tests
the markers from the previous commit skip off-host. without a host to
run them on, every marked test is a silent skip. this commit adds the
hosts.

- tests-os.yml runs -m macos_only on macos-latest and -m windows_only
  on windows-latest. ci.yml requires both lanes in all-checks-pass.
- a lane fails on pytest exit code 5 (zero tests selected). a renamed
  marker cannot produce a green job that ran nothing.
- each lane repeats 'not integration' because a command-line -m
  replaces the addopts filter.
- scripts/ci/list_os_marked_tests.py selects which files each lane
  imports. -m filters after collection, and collection imports every
  module. without this helper, one unrelated ImportError on the
  foreign host fails a job whose own tests passed. the helper exits
  non-zero when a marker matches no file, and writes bytes with
  explicit lf so windows crlf translation cannot corrupt the bash
  file list. it has its own tests in tests/ci/.
- the local runner now reports the skipped count and prints a note:
  macos_only/windows_only tests were skipped on this host, and this
  ci lane runs them. a green local run on linux no longer reads as
  coverage of the other hosts.
- the runner default job count is now #cpu, not #cpu*2.
2026-08-09 22:09:49 -04:00
ethernet 7aecab56db ci: move the review comment and the image build out of the CI run
The CI run stayed in progress until its last job ended. Two advisory jobs
set that time: the review-comment poller (40 minutes) and the Docker image
build (45 minutes). Neither job was required to merge.

GitHub refuses `gh run rerun` on a run that is in progress. Thus a reviewer
who added the `ci-reviewed` label had to wait for the two slow jobs, and
label-rerun.yml carried a 2100-second wait loop for this reason. The fast
required jobs were ready long before.

Each slow job now runs in its own workflow:

- docker.yml owns its `pull_request` trigger and does its own change
  detection. The new `detect` job runs the same composite action with the
  same condition that ci.yml applied, so a tests-only PR still skips the
  build. The `workflow_call` trigger is gone.
- ci-review-comment.yml starts on `workflow_run` when CI starts. It reads
  the workflow and the scripts from the default branch, which is the trust
  boundary that the old job got from its `ref: default_branch` checkout.

The poller reads job results through the API, so it can report on a run
that it does not belong to. `WATCH_WORKFLOWS` names sibling workflows for
the same commit, and `select_watched_runs` keeps the newest run for each
name. Thus the comment still shows the Docker results. The list is
newline-separated, because a workflow name can contain a comma.

The poller always exits 0 now. It reports on the CI run from a different
run, so a failed CI job is not a failure of the poller. The CI run has its
own gate for that.

Also correct a parse error in label-rerun.yml. STATUS came from the already
truncated RUN_ID, so its value was the run id and never "completed". Thus
the wait branch always ran.

ci.yml no longer needs `packages: write`, because the image build has left.
2026-08-09 15:40:04 -04:00
ethernet 0e2b5f3835 chore: remove old plan files 2026-08-09 15:17:01 -04:00
Teknium 3d7dda4cf4 fix(docs): retain prior builds' hashed assets across Pages deploys
CI / Installer tests (push) Has been cancelled
CI / Detect affected areas (push) Has been cancelled
CI / Python tests (push) Has been cancelled
CI / Python lints (push) Has been cancelled
CI / JS & TS checks (push) Has been cancelled
CI / Desktop E2E (push) Has been cancelled
CI / Docs Site (push) Has been cancelled
CI / Deny unrelated histories (push) Has been cancelled
CI / Check contributors (push) Has been cancelled
CI / Check uv.lock (push) Has been cancelled
CI / Check no committed infographics (push) Has been cancelled
CI / package-lock.json diff (push) Has been cancelled
CI / Lint Docker scripts (push) Has been cancelled
CI / Build&Test Docker image (push) Has been cancelled
CI / Supply-chain scan (push) Has been cancelled
CI / Review label gate (push) Has been cancelled
CI / OSV scan (push) Has been cancelled
CI / CI review comment (live) (push) Has been cancelled
CI / All required checks pass (push) Has been cancelled
CI / CI timing report (push) Has been cancelled
Deploy Site / deploy-vercel (push) Has been cancelled
Deploy Site / deploy-docs (push) Has been cancelled
Docker Build, Test, and Publish / build (amd64, type=gha,scope=docker-amd64, type=gha,mode=max,scope=docker-amd64, linux/amd64, ubuntu-latest) (push) Has been cancelled
Docker Build, Test, and Publish / build (arm64, type=gha,scope=docker-arm64, type=gha,mode=max,scope=docker-arm64, linux/arm64, ubuntu-24.04-arm) (push) Has been cancelled
Docker Build, Test, and Publish / publish (amd64, type=gha,scope=docker-amd64, type=gha,mode=max,scope=docker-amd64, linux/amd64, ubuntu-latest) (push) Has been cancelled
Docker Build, Test, and Publish / publish (arm64, type=gha,scope=docker-arm64, type=gha,mode=max,scope=docker-arm64, linux/arm64, ubuntu-24.04-arm) (push) Has been cancelled
Docker Build, Test, and Publish / merge (push) Has been cancelled
auto-fix lint issues & formatting / Generate eslint --fix patch (push) Has been cancelled
auto-fix lint issues & formatting / Apply patch (push) Has been cancelled
Pages serves exactly the newest artifact, so every push-triggered deploy
deleted the previous build's content-hashed JS/CSS while edge caches
(max-age=300, stale-while-revalidate=3600) kept serving HTML that
referenced them. With deploys landing every ~15-30 min, docs pages spent
most of the day pointing at 404'd bundles — search (pure client JS) was
the loudest casualty.

Fix: keep a rolling 14-day pool of hashed assets (en + zh-Hans) in the
Actions cache and union-merge it into each deploy artifact, current
build authoritative on collision (cp --update=none). Stale HTML and
already-open tabs now keep resolving across any number of deploys.
2026-08-08 06:04:36 -07:00