cf60ebbdfd264c3fc067b9e3230df8752c50d2ec
13 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
10f99bc15e |
ci: run the work lanes on larger runners and merge the split jobs
Every Linux lane that does real work ran on a 4-core `ubuntu-latest`. The Python suite and the JS checks were split into many small jobs to make that size usable. Each split job repeated the full setup. In most of the JS jobs the repeated setup cost more than the work. The work lanes move to larger runners. Then the splits that existed only to make small runners usable go away. Python tests: 12 slices become 1 job on a 96-core runner. Slicing cost a matrix job, a duration cache, a per-slice artifact and a merge job. 96 cores clear the floor that the slowest single test file sets, which is about 82s. A second slice divides work that is already at that floor, and adds a second setup. Duration data from run 32522943054 gives the numbers behind this: 3178 files, 11645s in series. The worker count is explicit, because `run_tests.sh` defaults to twice the core count. A later commit sets it from a measurement on this hardware. JS checks: 14 jobs become 1. The matrix paid about 371s of repeated setup to spread about 612s of work. One larger runner installs one time. The three UI shard scripts and `run-ui-shard.mjs` are therefore removed, because the unsharded `test:ui` covers the same tests. The unit of parallel work inside that job is a CHECK, and not a workspace. apps/desktop is most of the payload, and its own `check` is a serial && chain. A spread across workspaces alone therefore leaves that chain as the long pole. A package that declares `check:*` sub-scripts gives one unit for each sub-script. That is the same selection rule the matrix used. The loop lives in `.github/scripts/run-workspace-checks.mjs`, so the same sequence runs on a laptop. It runs 11 units together, buffers the output of each one, and fails at the end with the full list. Children that share one stdout interleave their lines and make a failure hard to read. `npm run --ws check` stops at the first workspace that fails. `check:test:plugins` joins the desktop `check` script. The matrix prefers `check:*` sub-scripts over the plain `check` script, so `check:test:plugins` ran only as its own leg. Without this change the merge drops that suite and the job stays green. node_modules is cached on the lockfile, and `npm ci` is skipped on an exact hit. The `cache: npm` option of `setup-node` caches only the ~/.npm tarball cache, which leaves the extract and the postinstalls to pay again. The arm64 image build stays on a native arm64 runner. A build of linux/arm64 on an x64 host uses emulation. The docker test lane caps its workers at the core count. Each of those tests drives a container, so the docker daemon sets the limit and not the processor. `.github/actionlint.yaml` declares the runner labels. actionlint knows the GitHub-hosted labels only, and an undeclared label reads as an error that hides the real findings. The `detect` job checks out one file through a sparse checkout, and its timeout drops to 1 minute. It reads `scripts/ci/classify_changes.py` and nothing else. Verification: - actionlint reports 9 findings across all workflows. An unmodified HEAD with the same config reports the same 9. This change adds none. - A wrong label still fails. actionlint reports `ubuntu-latest-32-cor` and `ubuntu-latest-32-arm-cores`. - Every changed workflow parses, and `name` parses as a string. - A replay of the `save-durations` merge step against a three-artifact layout returns all 3178 entries. - An expansion of the npm script graph gives the same leaf commands for the parallel units and for a plain `npm run check`, in both directions. Against the 13-leg matrix the count is 13 to 11, and the whole difference is the three UI shards that collapse into one unsharded `check:test:ui`. - `--list` reports the 11 units, and a full local run completes and reports the time of each unit. - The runner labels cannot be verified here. The first real run is the test. |
||
|
|
de0abc0617 |
ci(js-tests): drop dead electron download-cache path, skip redundant npm upgrade
Review findings on the caching commit: - ~/.cache/electron was dead weight: with npm ci skipped on an exact cache hit, the download cache is never read (electron's unpacked binary lives in node_modules/electron/dist, inside the cached tree); it only inflated every saved archive by ~110MB. - 'npm i -g npm@12' ran unconditionally in all 14 matrix jobs (~5-15s each); now a no-op when the bundled npm is already 12.x, which also keeps the installed major aligned with the npm12 cache-key tag. yaml + actionlint pass. |
||
|
|
f56a9a1185 |
ci(js-tests): cache the installed node_modules tree, not just the npm tarball cache
Every job in the js-tests matrix (~10 jobs/run, 13 after the UI-suite
sharding) runs a full 'npm ci' that deletes and re-extracts the entire
workspace node_modules and reruns all postinstalls — including the
Electron binary fetch (~100MB) — because setup-node's 'cache: npm' only
caches the ~/.npm tarball cache.
Cache the installed tree itself with actions/cache (the SHA-pinned
v4.2.4 already used by e2e-desktop.yml), keyed on the exact lockfile
hash, and skip 'npm ci' on a hit:
- key includes runner.os + node26 + npm12 so a toolchain bump never
reuses a stale tree
- NO restore-keys: a partial hit would leave a stale tree ('npm ci'
skipped means nothing would repair it), so anything but an exact
lockfile match reinstalls from scratch
- distinct keys for the discovery job (--ignore-scripts tree) and the
check jobs (with-scripts tree + ~/.cache/electron), which differ in
postinstall artifacts
Measured from run 31783969717: the npm-ci step is 30-45s per check job.
On warm cache this drops to a few seconds of restore, saving roughly
5-8 runner-minutes per PR run and ~1GB of registry traffic, and taking
~35s off every job on the merge-gate critical path.
|
||
|
|
713a983e4a |
feat(runtime)!: require Node 26 across all installers, heal, and upgrade paths
Hermes now pins its toolchain to Node 26 everywhere. Every path that installs, accepts, heals, or upgrades a Node runtime moves from the old 22-default / `^20.19 || >=22.12` floor to a single rule: Node >=26. Installers: - scripts/install.sh — NODE_VERSION=26; node_satisfies_build() collapses the two-branch Vite floor to `major >= 26`; user-facing messages updated. - scripts/install.ps1 — $NodeVersion=26; Test-NodeVersionOk likewise; winget fallback switches OpenJS.NodeJS.LTS -> OpenJS.NodeJS (26 is Current, not LTS — the LTS manifest would reinstall a too-old Node). - Dockerfile — node_source stage node:22-bookworm-slim -> node:26 (digest pinned, amd64 sha256:9e6f...bf73). - nix/ was already on nodejs_26 (lib.nix, npm-12-0-2.nix); the checks.nix wrapper check ratchets from `>= 20` to `>= 26`. Heal/upgrade paths: - scripts/lib/node-bootstrap.sh — HERMES_NODE_TARGET_MAJOR default 22->26 and HERMES_NODE_MIN_VERSION default 20->26, so heal_managed_node, _nb_install_bundled_node, and the fnm/proto/nvm/brew rungs all target 26 and stop accepting an on-PATH Node below it. Both remain env-overridable. - hermes_constants.py — _HERMES_NODE_TARGET_MAJOR fallback 22->26, which drives the Windows heal path's latest-v26.x download. Version gates: - package.json engines.node >=20 -> >=26; apps/desktop engines `^20.19.0 || >=22.12.0` -> `>=26.0.0`. - CI setup-node: all five workflows 22 -> 26. - Docs describing Hermes's own toolchain updated (windows-native, docker, acp, nix-setup, contributing). Skill docs describing third-party tools' own requirements are untouched. Termux still installs via `pkg install nodejs` best-effort (nodejs.org ships no Android tarballs); that path was never version-gated. Verified: bash -n on both shell scripts, PowerShell AST parse of install.ps1, latest-v26.x index resolves (node-v26.5.1), and the install test suite — 18 tests across the 5 install/runtime test files — passes. |
||
|
|
f88ed6c717 | fix: fix @nousresearch/ui version, update to npm 12 | ||
|
|
5be99b6fce |
ci(js-tests): split check into parallel matrix shards per workspace (#70252)
Every npm workspace package now defines check:* scripts (check:unit, check:lint, check:bundle, check:typecheck, etc.) that fan out to separate matrix runners in CI. The check umbrella script chains all shards for local dev. The matrix discovery in the workspaces job queries npm workspaces, finds check:* scripts (in package.json insertion order), falls back to check when none exist, and emits an include matrix. No hardcoded package names — the workflow is fully auto-derived from workspace metadata. Previously every package ran a single check script on one worker, and the fix step (lint:fix + prettier) ran as a separate CI step with special-cased run_fix gating to avoid running on every shard. Now that lint is just another check:lint shard, the run_fix field and the fix step are gone entirely — lint runs in its own runner like everything else. |
||
|
|
597615ade4 |
fix(ci): make tests, workflows, and attribution reliable under load (#66373)
* feat(attribution): conflict-free contributor mappings via contributors/emails/ directory
The AUTHOR_MAP dict in scripts/release.py was a merge-conflict magnet:
every concurrent salvage PR appended entries to the same lines of the
same file, so parallel PRs re-conflicted on every merge to main.
New system: one file per email under contributors/emails/ — filename is
the commit-author email, first non-comment line is the GitHub login.
File additions never conflict, so any number of PRs can add mappings
concurrently.
- scripts/release.py: AUTHOR_MAP is now LEGACY_AUTHOR_MAP (frozen)
merged with the directory at import time (directory wins). All
existing consumers (resolve_author, contributor_audit.py) unchanged.
- scripts/add_contributor.py: idempotent CLI to add a mapping; refuses
conflicting reassignments (incl. against the legacy map), validates
email/login shapes.
- contributor-check.yml: attribution gate now accepts a mapping file OR
a legacy entry; failure message prints the exact add_contributor
command. Also auto-resolves bare <login>@users.noreply.github.com
emails is intentionally NOT added (kept id+login form only, matching
previous behavior).
- contributor_audit.py: guidance now points at add_contributor.py.
- tests/scripts/test_contributor_map.py: 12 tests covering loader,
merge precedence, CLI idempotency/conflict/validation, subprocess E2E.
* feat(ci): one-shot per-file flake retry in the parallel test runner
A failing test FILE is re-run once in a fresh subprocess. Pass-on-retry
counts as green but is loudly reported in a '⚠ FLAKY' summary section
(with both attempts' output preserved) so the flake gets fixed instead
of eating a full-run rerun. Deterministic failures fail both attempts —
regressions cannot be laundered green.
- --file-retries N / HERMES_TEST_FILE_RETRIES (default 1, 0 disables)
- E2E verified: simulated first-run-fail flake goes green with banner;
deterministic failure still exits 1; retries=0 restores old behavior.
This converts the dominant CI failure mode (one timing-sensitive test
flaking a 4600-test shard, requiring a manual 10-minute rerun and an
agent triage loop) into a self-healing retry that costs one file's
runtime.
* test(approval): loosen wall-clock perf bounds 0.15s -> 2.0s
These guard against catastrophic regex backtracking (seconds-to-minutes
class), but 0.15s is within scheduler-stall noise on loaded shared CI
runners — test_max_accepted_separator_free_input_is_fast failed a CI
shard this week on runner load alone. 2.0s still catches the regression
class with zero flake surface.
* fix(ci): job timeouts everywhere + retries on all network installs
Reliability pass over every workflow:
- timeout-minutes on all 21 jobs that lacked one (a hung job previously
burned the 6-hour default runner budget)
- ./.github/actions/retry wrapped around every network-fetching install
that lacked it: pip installs (deploy-site, skills-index), npm ci
(deploy-site website, upload_to_pypi web + ui-tui), uv sync (docker
test deps). Deterministic build steps (npm run build) deliberately
NOT retried — split into separate steps so a real build failure fails
fast instead of retrying 3x.
* docs(agents): document the file-retry flake policy
* fix(ci): curl retries on deploy hook + skills-index probe
* fix(ci): kill the remaining transient-failure classes in workflows + Dockerfile
From the workflow reliability audit:
- tests.yml: duration-cache restore had NO restore-keys while saves use
run_id-suffixed keys — the cache never matched once, so LPT slicing
always ran blind and unbalanced slices pushed heavy files toward the
per-file timeout. One-line restore-keys fixes slice balancing.
- Label gates (lint ci-reviewed, supply-chain mcp-catalog-reviewed):
'gh pr view || true' turned an API blip into 'label absent' → false
BLOCKING failure. Now 3x retry, and API failure is reported as an API
failure instead of a missing label.
- detect-changes action: compare API retried before failing open (was
silently running all lanes on any blip).
- uv-lockfile-check: 'uv lock --check' resolves against PyPI — retried
so registry blips don't read as 'lockfile stale'.
- docker.yml merge job: imagetools create retried (Docker Hub eventual
consistency on just-pushed digests).
- Dockerfile: apt-get Acquire::Retries=3; s6-overlay ADDs converted to
curl --retry 3 (ADD cannot retry; checksums still enforced); npm
--fetch-retries=5; playwright chromium fetch retried 3x.
- Advisory artifact uploads (per-slice durations, ci-timings report)
get continue-on-error so an artifact-service blip can't fail a green
test slice.
* fix(tests): kill the two root-cause flakes — leaking pre-warm timer + env-dependent provider list
- test_tui_gateway_server.py: session.create / non-eager session.resume
arm a 50ms threading.Timer (_schedule_agent_build) that outlives its
test and fires into the NEXT test's _make_agent mock, racily
corrupting captured state (the recurring session_resume shard
failures). Replaced the per-test whack-a-mole stub with a module-wide
autouse fixture; the 3 worker-lifecycle tests that genuinely need the
deferred build opt back in via @pytest.mark.real_agent_prewarm (new
marker in pyproject).
- test_api_key_providers.py: PROVIDER_ENV_VARS is now derived from the
live PROVIDER_REGISTRY instead of a hand-list that had drifted
(missing HF_TOKEN / DEEPINFRA_API_KEY) — resolve_provider('auto')
tests failed on any machine with HF_TOKEN exported. E2E-verified with
HF_TOKEN/DEEPINFRA_API_KEY set: 42/42 pass.
* test: de-flake 30 timing-sensitive test files for loaded CI runners
Root-cause fixes from the flake audit (session-DB mining + repo sweep):
Event-based sync instead of sleep-sync:
- title_generator: mock sets threading.Event, wait(10) replaces
sleep(0.3) hoping the daemon thread got scheduled
- docker zombie_reaping / profile_gateway: poll-for-state helpers
replace fixed 1-3s sleeps (s6 transitions + SIGCHLD reaping are async)
- process_registry tree test: select()-bounded readline replaces an
unbounded blocking read (parent wedge now fails THIS test with a clear
message instead of an opaque rc=124 file kill); SIGTERM grace 1s->2s
(the 1s partition window mid-interpreter-startup is how a child PID
escaped the live-system guard in CI)
Timeout raises (loaded 8-way-sliced runners see ~5s scheduling floors;
all of these complete in ms-to-1s when healthy so the raises cost
nothing on green runs):
- subprocess/thread waits <= 2s raised to 10-15s across mcp_tool,
mcp_circuit_breaker, mcp_reconnect_retry_reset, mcp_parked_self_probe,
mcp_cancelled_error_propagation, registry, clarify_gateway, interrupt,
voice_cli_integration, docker_environment, session_store_lock_io,
planned_stop_watcher, cli_interrupt_subagent, thread_scoped_output
(joins now also assert not is_alive() so stragglers fail loudly)
- wall-clock discrimination ceilings loosened where the guarded hang is
10x larger: local_background_child_hang 4s->10s, interrupt_cleanup
setup 5s->20s + pgid-exit 30s->60s, mcp_stability grandchild spinup
5s->15s, protocol/gil-starvation fast-handler 0.5s->2s,
iso_certify_seam 1.5s->5s, wait_for_mcp_discovery 0.1s->1s
- narrow assertion windows widened: honcho first-turn wait 0.4..0.65 ->
0.25..2.0 (property is bounded-not-hung, not an exact wall-clock);
compression fork-lock TTL 1s->3s (12 refresh chances per lease);
compression-lock expiry margins symmetric (ttl 0.05->0.5, sleep 1.0)
- telegram hung-DNS bound 1.0->1.4 (fake hang is 1.5s — must stay under)
* fix(tests): repair indentation from de-flake batch edit
* fix(tests): harden env isolation and replace remaining sleep-sync races
The full 42k-test run and complete npm check surfaced three more classes:
- Environment isolation: local ~/.honcho defaultHost and SSH_* variables
leaked into Python/TUI tests. Pin the default Honcho host in the
hermetic fixture, isolate the one fallback test from ~/.honcho, and
blank SSH_* around terminalSetup tests. This flipped 20 false failures
back to deterministic behavior on developer machines.
- Background-thread sleep-sync: Honcho async writer tests patched
time.sleep globally, then busy-polled with that same mocked sleep. Under
full-suite load the poller could starve the writer. Each test now waits
on an Event emitted by the exact flush/retry transition; 30/30 passed
under 15-way contention.
- Desktop streaming: the test slept 80ms and assumed a 500ms timer could
not fire before its assertion. A loaded runner descheduled the test for
>500ms and both chunks arrived. Producer controls now gate second-chunk
and completion transitions explicitly.
Also make file-retry observability complete: a self-healed flaky file now
prints BOTH attempts' full output in the FLAKY summary. Two behavioral
runner tests prove pass-on-retry is green+loud+traceback-preserving, while
a deterministic failure remains red.
* refactor(ci): use gh bot pat, better retries
refactor(ci): use retry action for PR label fetch
the retry action now captures stdout as a step output, so it can serve
double duty: retry + output capture for commands like 'gh pr view' whose
result must be consumed by later steps.
Retry action gains:
- 'stdout' output (heredoc-delimited to preserve newlines)
- tee to temp file so stdout still streams to the job log
- step id 'retry' for output reference
Both lint.yml and supply-chain-audit.yml now use the retry action
directly with 'command: gh pr view ...' and read
steps.<id>.outputs.stdout.
ci: use AUTOFIX_BOT_PAT for all gh CLI / GitHub API auth
Replace secrets.GITHUB_TOKEN and github.token with
secrets.AUTOFIX_BOT_PAT across all workflows and composite actions
that use the gh CLI or GitHub API. The PAT has consistent permissions
across fork PRs (where GITHUB_TOKEN is read-only), avoids API rate
limit sharing with the default token, and is already used by
js-autofix.yml for the same reasons.
19 sites swapped across 9 files:
- lint.yml (3): label fetch, comment post/edit, comment update
- supply-chain-audit.yml (5): scan, critical comment, unbounded dep
comment, label fetch, mcp-catalog comment
- lockfile-diff.yml (1): PR comment post/update
- skills-index-freshness.yml (1): issue creation on degraded probe
- skills-index.yml (2): index build, trigger deploy workflow
- upload_to_pypi.yml (2): release view poll, release upload
- ci.yml (1): timings report
- deploy-site.yml (2): skills index crawl
- detect-changes/action.yml (1): compare API call
---------
Co-authored-by: ethernet <arilotter@gmail.com>
|
||
|
|
3bfa6001f7 |
fix(js ci): don't ignore native deps anymore
we need em for desktop :) |
||
|
|
2179d5e8af |
ci: add eslint lint matrix to js-tests.yml
Add a 'lint' job to the JS tests workflow that runs 'eslint --fix' across all discovered npm workspaces (same matrix as the check job). Fixable issues auto-correct and don't block; eslint exits non-zero only when un-fixable errors remain. Also fix duplicate 'needs: workspaces' in the check job. |
||
|
|
7577834206 |
fix(ci): fail closed when workspace matrix discovery produces empty list
The set-matrix step wrote the npm workspace query result directly to $GITHUB_OUTPUT. If discovery ever produced [], the matrix would expand to zero check jobs, leaving the reusable workflow green without running any JS/TS checks. Now the step validates the result is a non-empty array before emitting it, and exits 1 with a GitHub annotation if it's empty or jq failed. |
||
|
|
c3aa81be2f | feat(ci): load npm workspaces from package.json | ||
|
|
6800ec9d66 | change(ci/desktop): move desktop app build into check job | ||
|
|
7b3f3047ab |
feat(ci): run JS tests in CI, add npm run check in ws root
|