Commit Graph

2024 Commits

Author SHA1 Message Date
Bryan Bednarski 8afd98ef2a refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski 0a079b946f fix(relay): retain legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski e8644e05a3 refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
rob-maron 0b588cb3a4 MCP CIMD auth 2026-08-18 20:03:18 -07:00
ethernet 00c3872882 feat(ci): add nix flake check as unrequired job
The workflow owns its triggers and ci.yml does not call it. A
reusable-workflow call holds the caller run in progress for the full
build, and GitHub refuses `gh run rerun` on a run that is still in
progress. A separate run reruns and cancels on its own.

The job restores /nix/store from the GitHub Actions cache and saves from
main only. A cache that a PR writes is visible to that PR alone, so a
save there spends the quota of the repository and helps no later run.
2026-08-18 20:42:06 -04:00
ethernet 1dbe469276 refactor(ci): hoist docker detect-changes into the .py file
The docker.yml gate held its own copy of the build formula, in shell.
classify_changes.py now owns a derived docker lane, and the nix lane in
the next commit derives from the same file. Two formulas in two
languages drift apart, and one Python function with tests does not.
2026-08-18 20:42:06 -04:00
f-trycua 5b010f448f fix(computer-use): verify Windows driver repair 2026-08-16 11:34:40 -07:00
Francesco Bonacci 81af2ef013 fix(computer-use): reconcile existing cua-driver installs 2026-08-16 11:34:40 -07:00
Yasushi Fukutake b44f956dff fix(desktop-update): put --daemonized ahead of ORIGINAL_ARGS in posix.sh re-exec
Appending --daemonized after ORIGINAL_ARGS put it past the `--`
relaunch-args separator on Linux, so it was absorbed into
RELAUNCH_ARGS instead of being parsed as a flag. HANDOFF_DAEMONIZED
never got set, so the one-shot self-detach block re-fired on every
re-exec -- an unbounded self-exec loop (thousands of iterations/sec,
100%+ CPU, argv growing until execve fails with E2BIG) whenever
relaunch args were present, which is the normal invocation shape on
Linux.

Fixes #86957

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-16 06:25:21 -07:00
RelaxJonh eac1f65340 fix(install): validate --commit SHA and fail hard on fetch/checkout errors (#87268)
Three problems with install.sh --commit:

1. No validation: non-hex or too-short arguments passed through to git,
   producing misleading errors.

2. Fetch failure swallowed by || true: abbreviated SHAs are refused by
   GitHub's server ("couldn't find remote ref"), but the error was
   silently ignored.

3. Checkout failure not checked: git checkout --detach with a missing
   object produces a misleading "does not take a path argument" error
   and the install continues unpinned, exiting 0.

Fix:
- Validate --commit is a 7-40 hex string up front
- Remove || true from fetch; fail with actionable message directing
  users to full 40-char SHAs
- Check git checkout --detach result and fail hard on error

Fixes #87268
2026-08-16 01:45:10 -07:00
RelaxJonh f887819421 fix(install): capture npm output on failure for diagnosable errors (#87340)
Both install-blocking npm install call sites (browser tools and TUI) ran
with --silent and no output capture, so failures printed only a generic
error message with no npm diagnostics.

Apply the same pattern used by the camofox install path: redirect npm
output to a temp file and replay it on failure, so users can see the
actual error (EBADENGINE, ETARGET, network timeout, registry 5xx, etc.).

Fixes #87340
2026-08-16 01:44:45 -07:00
Teknium 3af56c2203 fix(install.ps1): surface uv installer errors and add GitHub + existing-uv fallbacks (#69216)
Install-Uv piped the astral installer's entire output to Out-Null, so any
real failure (proxy block, AV quarantine, permissions) surfaced only as the
generic "uv installed but not found" message, and astral.sh was the sole
install source even though corporate proxies commonly block it while the
byte-identical GitHub releases installer downloads fine.

Three-rung ladder, all inside Install-Uv:
1. astral.sh installer with output captured via Tee-Object.
2. GitHub releases installer mirror (same UV_INSTALL_DIR).
3. Salvage an existing uv.exe (Get-Command uv, or the astral default
   %USERPROFILE%\.local\bin\uv.exe) by copying it into $HermesHome\bin so
   the managed-first invariant holds.

On total failure, print the last 15 lines of captured installer output plus
the existing manual-install pointer.

Reported by @BitBernd; proxy diagnosis by @gakugaku; Out-Null suppression
first identified by @webtecnica in #69366.

Closes #69216
2026-08-14 22:33:58 -07:00
konsisumer 4b0c1031db fix(desktop-update): wait for rebuilt executable before relaunch 2026-08-14 22:23:00 -07:00
Teknium aa5a960675 fix(installer): hold the venv rollback source through dependency validation (#83149)
Review finding on PR #83194 (egilewski): Install-Venv committed the venv
transaction as soon as the replacement had a working interpreter, deleting
the parked previous venv. Install-Dependencies is a separate later stage
(a separate process under the stage-per-process bootstrap) and every
dependency tier or the baseline-import gate can still fail after that
point - a failed update could still leave Hermes and the blocker probe
unusable with no rollback source.

Now:
- Install-Venv records the parked backup in venv.pending-backup instead
  of deleting it, and excludes it from the venv.stale.* sweep.
- Install-Dependencies wraps the dependency tiers + baseline-import gate
  in the transaction: Restore-VenvBackup on failure (parks the failed
  replacement as venv.failed.*, renames the previous venv back), and
  Complete-VenvTransaction only after the imports prove the replacement
  usable.
- Source-contract regression tests for the boundary
  (tests/test_install_ps1_venv_transaction_boundary.py).
2026-08-14 21:58:09 -07:00
HexLab98 18b442cdeb fix(install): abort Windows venv recreate when rename-aside fails
When Rename-Item on the live venv is denied, do not fall back to an
in-place Remove-Item that can gut site-packages and leave no rollback.
Also mark venv-blocker probe failures with probe_failed so they cannot
be read as a clear scan (#83149).
2026-08-14 21:58:09 -07:00
JoaoMarcos44 11268e7e6a fix(installer): make Windows venv recreation transactional (#83149) 2026-08-14 21:58:09 -07:00
Teknium bc36d7d6c8 fix(deps): exempt no-upload-date and exact-pinned packages from exclude-newer bricking
The relative exclude-newer = "14 days" cutoff bricks installs whenever the
resolver cannot see (or accept) a package's upload date:

- defusedxml / python-olm / unpaddedbase64 (#80387, #79434): ancient frozen
  releases (2021-2023) whose upload dates are often absent from mirror
  indexes and stale uv HTTP caches. uv then filters them entirely
  ("there are no versions of defusedxml"), breaking [youtube]/[wecom]/
  [matrix] resolution and daily `uv sync --locked` runs.

- setuptools / pillow / mcp (#78227, #75992, #76020): exact-pinned deps.
  When the pinned version's upload date is invisible, the resolver filters
  the ONLY acceptable candidate — setuptools==83.0.0 in
  [build-system].requires meant the project could not even be built from a
  git checkout on released v0.20.0. Exempting an exact pin costs nothing:
  the version cannot float without a reviewed pin bump.

Changes:
- pyproject.toml: add all six to the existing exclude-newer-package
  whitelist, with rationale comments per class.
- uv.lock: regenerated; diff is the whitelist metadata only (verified
  zero version drift, still 249 packages).
- tests/test_packaging_metadata.py: new standing guard
  test_build_system_requires_exempt_from_exclude_newer — every
  [build-system].requires package must be whitelisted while a relative
  exclude-newer cutoff is configured. Verified both directions (fails
  when setuptools is removed from the whitelist).
- scripts/install.sh: fix the stale tier-name comparison ("all (with
  RL/matrix extras)" vs actual "all") that mislabeled every successful
  Tier-1 install as a fallback-tier install (#79434 bonus finding).

Verification: uv lock --check green on uv 0.11.19 and 0.12.5;
uv sync --extra all --locked green; uv pip install -e '.[all]' resolves;
whitelist mechanism A/B-proven on a minimal project (unsatisfiable ->
resolves; build-requires variant: uv build fails -> succeeds).

Reported-by: MichaelClawHub (#80387), liujianqiu (#79434), maxonliu (#78227)
2026-08-14 21:50:48 -07:00
Alvis cf8b505531 fix(desktop-update): make posix hand-off survive Electron quit teardown (macOS)
The Desktop-spawned hand-off consistently died during Electron's quit
teardown on macOS: the orchestrator process group was terminated right
after `running: hermes update ...`, so no exit code, result file, bundle
swap, or relaunch ever happened, and the loopback shim window surfaced
the death as ERR_CONNECTION_REFUSED or "Aw, Snap!" error code 15
(reproductions in #66753).

- Re-exec the orchestrator through a one-shot setsid child and let the
  direct Electron child exit immediately; the real orchestrator is owned
  by launchd (PPID 1), outside Electron's teardown, same marker/result
  protocol.
- Hold TERM ignored across the `hermes update` invocation and
  log-and-ignore the single teardown TERM that can still arrive after
  the desktop PID dies (durable SIGNAL breadcrumb for diagnosis).
- Delay start_ui until the desktop PID is gone plus 1s so the shim
  server/window are never born inside the teardown window.
- Run both UI processes in their own sessions; keep SIGTERM/SIGHUP
  ignored in the shim server and stop it with SIGKILL, so a stray TERM
  can no longer leave the progress window on a dead loopback URL while
  the update continues.

Verified on a production git install (macOS arm64, Darwin 27.0,
v0.20.1): six consecutive Desktop-triggered/production-shape updates
completed end-to-end including a full desktop rebuild + codesign; the
shim survived a deliberately injected TERM+HUP mid-update and a full
`hermes desktop --force-build --build-only` running alongside it.

Fixes the macOS reproductions in #66753.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-14 21:48:22 -07:00
Christopher 9b3823d08b fix(install): provision Node 26 so managed npm satisfies engines
Node 22 ships npm 11.16.0, which engines.npm rejects (11.10–11.16
ignore min-release-age-exclude). Fresh Hermes-managed installs then
fail npm ci with EBADENGINE. Node 26 ships 11.17.0.
2026-08-14 21:48:14 -07:00
David Metcalfe d6a711c783 fix(win): CIM-based PS in-use check and same-volume staging swap
Adopts two improvements from #81586 (kshitijk4poor), with one correction:

- Test-ManagedNodeInUse now queries Win32_Process (ExecutablePath +
  CommandLine substring) instead of Get-Process .Path: a cmd.exe wrapper
  running npm.cmd from the tree reports its own exe in System32, so the
  tree path shows up only in the command line. Win32_Process.CommandLine
  works on Windows PowerShell 5.1 and 7+, unlike the Get-Process
  .CommandLine ETS property (7.4+ only); a single CIM query also beats a
  per-process property access loop. This closes the gap flagged by the
  triage bot on #81500 (Update-ManagedNpm's only in-use protection is
  this pre-check).
- Test-Node stages the extracted tree to a sibling node.new-* before the
  swap, so the final swap is a same-volume rename (atomic) instead of a
  cross-volume Move-Item (copy+delete, non-atomic).

The mtime-touch ordering from #81586 was NOT adopted: touching the backup
only after the swap succeeds leaves it at its old (long-lived-tree) mtime
for the whole swap window, which the age-gated litter sweep (st_mtime <
cutoff) removes — reopening the concurrent-sweep race the touch exists to
close. The touch stays immediately after the backup rename.

Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
2026-08-14 21:35:30 -07:00
David Metcalfe e4d0e4c3d8 fix(win): never rewrite the in-use managed Node tree (#80926)
The Hermes-managed Node tree at %HERMES_HOME%\node is destructively
rewritten while the desktop app's Node processes execute from it:
the Node-26 heal did shutil.rmtree + move, the EBADENGINE repair ran
npm install --global --prefix into the tree, and install.ps1's
Test-Node did Remove-Item + Move-Item. Windows rejects those writes
with PermissionError: [WinError 5] on npm.cmd.

- _heal_managed_node_windows: stage the fully-downloaded tree in a
  sibling node.new-* dir, then rename-swap (live tree -> node.old-*,
  staged -> node). The live tree is never deleted before its
  replacement is ready, so an interrupted heal cannot gut it; a
  refused rename is the OS-level in-use signal and defers (returns
  None) instead of forcing the write.
- heal_hermes_managed_node: an in-use deferral does not record the
  once-per-process attempt, so the heal retries once the tree is free.
- managed_node_tree_in_use: cheap psutil pre-check (Windows only) that
  avoids pointless 30-50MB re-downloads in long-lived processes.
- upgrade_managed_npm: defer the in-place npm self-upgrade while the
  tree is in use, with a notice.
- install.ps1: Test-ManagedNodeInUse guard around Update-ManagedNpm and
  the Test-Node install branch, which now rename-swaps instead of
  delete-then-move.

An in-use-but-outdated tree keeps serving the old runnable Node (old
Node beats no Node), and every npm resolution re-evaluates the heal, so
the upgrade applies automatically on the next update with the app
closed.
2026-08-14 21:35:30 -07:00
Teknium 20e5d51bea fix(browser): managed-first browser-use CLI resolution
Everything Browser Use is now managed by Hermes: the canonical binary
is the one install_cli() provisions into HERMES_HOME/bin, and every
resolution and provisioning site prefers it.

- _find_cli(): probe order flipped to managed bin -> PATH ->
  user-level tool dir (then uvx across the same order). A user's own
  uv tool install can no longer shadow the Hermes-managed copy with a
  drifted version; side installs only matter when we have nothing.
- install_cli(): a browser-use on PATH no longer short-circuits the
  install — only the managed copy does, so selecting any backend
  provisions the copy Hermes controls and updates.
- _ensure_browser_use_cli() (hermes tools): drops its own PATH check
  and always delegates to install_cli(), the single owner of the
  managed-copy policy.
- install.sh / install.ps1: same short-circuit fix — only
  HERMES_HOME/bin/browser-use counts as installed.

Follow-up to #86240 and #86320: with every non-Camofox backend
selection installing the CLI, managed-first closes the remaining
version-drift/shadowing class instead of guarding single sites.

Tests updated to pin the new contract: managed beats PATH and
user-local; PATH install does not satisfy install_cli; helper always
delegates. E2E-verified precedence chain with real files and a real
degraded-PATH install attempt.
2026-08-14 13:51:07 -07:00
Sascha 5bfb7ee42f fix(desktop-update): drive the Windows hand-off through the venv python, not the hermes.exe shim
`uv pip install -e .` has to replace the console-script shims, so
_quarantine_running_hermes_exe must first rename the running hermes.exe out
of the way. That rename fails whenever any child process spawned from that
hermes.exe is still alive: on Windows a child inherits a handle on the parent
image. It is the inherited handle, not the trampoline, that pins the file --
killing the child makes the identical rename succeed, and the shim flavour
(uv trampoline vs distlib launcher) makes no difference.

The updater spawns such children itself (npx cache warm, memory-provider
refresh -- hindsight-api runs as a daemon with --idle-timeout 300 and outlives
the step that started it), so this presents as a race rather than a hard
failure: the same hand-off succeeds on one run and dies on the next. Step 2's
shim-unlock preflight cannot catch it, because the shim genuinely is unlocked
at that moment; the pinning child appears later, during the update.

When the rename loses that race, _schedule_replace_on_reboot is the last
resort -- and MOVEFILE_DELAY_UNTIL_REBOOT writes to HKLM, so it needs
elevation. A Desktop-driven update is not elevated, so it returns
ERROR_ACCESS_DENIED, `uv pip install -e .` exits 2, and the ZIP fallback
repeats the identical sequence. The desktop build stage is then never reached
while the pre-build clean has already removed apps/desktop/release, leaving an
install whose Start Menu shortcut points at a Hermes.exe that no longer exists.

Running the same code as `python.exe -m hermes_cli.main update` puts the
inherited handles on python.exe, which uv never has to replace.

posix.sh is deliberately untouched: unlinking a running executable is legal
there, so the equivalent call is harmless.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 15:48:31 -05:00
AlexMnrs ca1b3f8705 fix(windows): avoid locale-sensitive update timestamps 2026-08-14 15:41:38 -05:00
ethernet cd667debfa fix(install-e2e): windows transcripts were empty; player gets #zip= hash + one player per run
Windows transcripts were ZERO bytes: ts-prefix.ps1 formatted with
{0:D2}, but Floor() returns a double and the D specifier is
integer-only - it threw per line, and under the driver's relaxed EAP
every line errored into the void. {0:00} fixes it (custom numeric
format works on doubles). Reproduced the exact pipeline locally
(empty file + Format specifier invalid), verified the fix produces
prefixed merged stdout+stderr with exit code intact. That is also
why the log timeline never auto-synced: there was nothing in the
files to sync.

The GitHub artifact URL 307s to /suites/... server-side and strips
the ?zip= query param. The player now reads the zip URL from a
#zip= HASH param (client-side, survives the redirect) with ?zip=
as fallback; the hash path was verified in a real browser against a
real leg zip (auto-fetch + boot).

Per ethie's design, one player artifact for the whole run: new
leg-player job uploads playback.html (archive:false) before the
matrix legs, the report job needs it, and each ran cell gets TWO
links - 📼 to the player with #zip=<that leg's logs zip> and ⬇️ to
the raw zip. Per-leg player uploads removed from all three run
workflows.
2026-08-13 22:51:44 -04:00
ethernet 08d9f27773 fix(install-e2e): results chart 📼 links against real artifact names
Debugged against run 31635036702 (real jobs + artifacts replayed
through the renderer). Two naming realities the renderer ignored:

1. upload-artifact with archive:false IGNORES the name: input and
   registers the artifact under the FILE's basename - every leg's
   player artifact is called 'playback.html' (all the same blob). The
   renderer now uses any 'playback.html' artifact for the player half
   of the 📼 link instead of install-e2e-player-<leg_id>.
2. The posix arms append -<sha> to the logs artifact name at upload
   (install-e2e-logs-<leg_id>-<sha>), while windows does not. Match
   by prefix instead of exact name.

Also fixed the pass/fail summary counters, which compared cells with
=== against '&#x2705;' - the appended reel link made every ran cell
count as neither passed nor failed (the summary said 0 passed while
the table was full of checks).

Replayed against the real run: 31 passed, 25 failed, 149 skipped,
24 📼 links on ran cells (both outcomes), none on skips.
2026-08-13 14:41:04 -04:00
brooklyn! 6a198f8a12 fix(install): a failed Node dependency install now fails the install instead of printing success (#85537)
* fix(install): fail when Node dependencies cannot install (#85297)

The POSIX installer converted root and TUI npm failures into warnings, then
printed a dependency-success message and reached the installation-complete
banner with a zero exit status. This left consumers with no usable
node_modules while reporting success.

Treat both required npm installs as fatal: log an error, restore tracked
lockfile churn, return status 1, and propagate the failure from the monolithic
and node-deps stage callers. Successful installs, Termux and missing-Node
skips, missing-manifest skips, and optional Playwright/Browser Use/Computer
Use best-effort behavior remain unchanged. The fix is limited to the POSIX
installer; the PowerShell installer is outside this issue's scope.

Focused and adjacent installer tests passed (32), with bash syntax,
py_compile, and diff checks clean. The broader installer family had 90 passes,
one unrelated pre-existing failure, and two skips; the full suite was
environment-limited by missing dependencies. CodeRabbit, iterative deep
security/compatibility reviews, and final confidence security/compatibility
reviews were clean against the final diff.

Fixes #85297

* fix(install): require npm alongside node in check_node (#77003)

A stray `node` symlink without a sibling `npm` (leftover from a node
version manager) made check_node report "Node.js found"; every later
npm install then failed and the desktop build died with an opaque
"Node.js / npm unavailable". Node now only counts as found when npm
resolves on the same PATH, with an explicit "stray node symlink?" branch
that falls through to the Hermes-managed Node (which bundles npm).

The overlapping success-log honesty half of the original PR is subsumed
by the previous commit, which makes a failed npm install fatal rather
than conditionally-logged; the behavioral tests there cover it, so this
commit keeps only the check_node PATH-gate assertions.

Fixes #77003.

Co-authored-by: criptogus <criptogus@users.noreply.github.com>

---------

Co-authored-by: Eugeniusz Gilewski <egilewski@egilewski.com>
Co-authored-by: CriptoGus <128640021+criptogus@users.noreply.github.com>
Co-authored-by: criptogus <criptogus@users.noreply.github.com>
2026-08-13 18:39:54 +00:00
brooklyn! d753957e8a fix(install): Windows setup no longer hangs forever on Node.js dependencies (#85529)
* fix(install): time-box the Windows node-deps stage so a stalled npm or Playwright install can't hang setup forever

scripts/install.sh has bounded this same work with run_with_timeout
"$NODE_DEPS_TIMEOUT" (600s default) since #39219, but install.ps1 never got
the guard: Install-NodeDeps ran both `npm install` and `npx playwright
install chromium` unbounded. A stalled registry fetch or a wedged Chromium
archive extraction (#76222, #84614) froze the installer indefinitely -- one
user left it running 12+ hours overnight before asking for help.

Route both invocations through _Invoke-NativeWithTimeout: cmd.exe launches
the native command with its output merged to a log, the parent polls with a
wall-clock deadline and tails new log lines to the console each tick (the
live progress that makes a 3-minute download distinguishable from a hang),
and on timeout taskkill /T /F kills the real process tree and returns 124 --
the same convention as coreutils timeout and bash's run_with_timeout.
Wait-Job was rejected for this: jobs swallow live output and Stop-Job leaves
the npm child running. Windows PowerShell 5.1-safe throughout.

Timeouts surface as a warning with the log path, a note that re-running the
installer resumes (stages are idempotent), and the NODE_DEPS_TIMEOUT env
override for slow links -- mirroring bash.

Fixes #76222.
Closes #84614.
Supersedes #76303.

Co-authored-by: JonthanaHanh <JonthanaHanh@users.noreply.github.com>

* fix(installer): roll stage timers over to hours so an overnight stall doesn't read as "744 hours"

formatElapsed rendered a running stage as m:ss with unbounded minutes: a
node-deps stage left hanging overnight showed "744:38", which the user who
reported the hang understandably read as 744 hours. formatDuration
(completed stages) had the same unbounded-minutes shape.

Move both formatters into src/lib/format.ts (pure, no React) and add the
hour rollover: h:mm:ss live, "Xh Ym" completed. tests-js pins the shapes,
including 744m38s -> 12:24:38.

---------

Co-authored-by: JonthanaHanh <JonthanaHanh@users.noreply.github.com>
2026-08-13 13:38:20 -05:00
Teknium 7060ac7bed feat(computer-use): provision cua-driver at install time and on toolset enable
Choosing Computer Use should be a config flip, not a hunt for
'hermes computer-use install'. Three provisioning rungs:

- install.sh / install.ps1 pre-install cua-driver (best-effort,
  non-fatal, time-boxed at 660s above the upstream installer's 600s
  lock window; --skip-computer-use / -SkipComputerUse to opt out;
  Termux and unwritable-/Applications skipped cleanly)
- PUT /api/tools/toolsets/{name} (dashboard + desktop toggle) spawns
  the background 'hermes tools post-setup cua_driver' action when the
  toolset is enabled while the binary is missing — previously the
  toggle 'saved' but the tool never appeared in the schema because
  check_computer_use_requirements() couldn't find the binary
- hermes tools interactive flow already installed via
  _toolset_needs_configuration_prompt/_POST_SETUP_INSTALLED (unchanged)

Docs: computer-use.md enabling section rewritten around the new flow;
installation.md documents --skip-computer-use.
2026-08-13 02:44:48 -07:00
Zak B. Elep 793f0b3ff1 fix(install): stop npm-installing agent-browser eagerly in install.sh/install.ps1
ensure_browser() (install.sh) and Install-AgentBrowser (install.ps1)
are reached only via the explicit --ensure browser / -Ensure browser
on-demand mode, itself only triggered by an actual browser-tool call's
lazy-install fallback or `hermes acp --setup-browser`. agent-browser
already resolves via npx in that same fallback before ever reaching
these scripts, so eagerly npm-installing a second, separately
version-pinned copy here was redundant and an extra credential/
supply-chain surface for a path npx already covers. Chromium
acquisition for this on-demand path is now deferred entirely to
_maybe_autoinstall_chromium's existing lazy fallback. camofox's
install and system-browser detection/configuration are unaffected.
install.ps1 also drops the now-dead -SkipChromium switch, confirmed
unused at its one call site.
2026-08-13 02:38:28 -07:00
ethernet e138cb555d feat(install-e2e): hook the leg player into the results table
Each leg uploads playback.html as a single-file artifact (archive:
false) before the driver runs, so it exists even on failure. The
results chart now links every leg that RAN (pass or fail, not skip)
to its player with ?zip= pointing at that leg's logs artifact.

Leg<->artifact mapping: the generator mints a leg_id per matrix entry
(sanitized matrix name, exported legId()), every run workflow names
its artifacts install-e2e-{player,logs}-<leg-id>, and the report job
feeds the run's artifact name->id list to the results renderer, which
rebuilds the leg id from the parsed job name. GitHub does not link
jobs to artifacts, so the deterministic name is the join key.

Empirical finding: GitHub artifact downloads are auth-gated (the
download URL 307s to /suites/... which is 404 anonymous), so a
locally-opened player page cannot fetch the zip cross-origin. The
player now degrades gracefully: ?zip= fetch failure renders a real
download link for the zip (a normal click carries the user's session)
plus a drag-and-drop / file-picker path, and no-param opens as a pure
drop target. Verified in a real browser against a real artifact URL.

Verified: generator emits leg_id, results renderer emits
✅/❌ [📼](...?zip=...) only on ran cells, npm run check PASS,
install tests 36/36, strict tsc PASS, actionlint x4 PASS.
2026-08-12 15:50:46 -04:00
brooklyn! 9eab7a4473 fix(ci): stop running uv lock --check on PRs that can't touch the lockfile (#84675) 2026-08-12 13:30:01 -05:00
ethernet 4578b5d0a6 ci(install-e2e): result chart says WHY a cell skipped
Skipped cells split into their reason: pre-desktop (a desktop-surface
method against a tag that predates apps/desktop) vs TODO (declared,
no driver arm yet). The report job passes pick-releases' annotated
tags into --format results; the shared methodNeedsDesktop() is the
same predicate the plan chart uses, so plan and results agree about
what pre-desktop means. Without --tags the renderer keeps the flat
skip label (backward compatible).

Verified against run 31579084845 real job list: 82 legs, 44 skips
labeled correctly.
2026-08-12 10:40:57 -04:00
Teknium f20d16fbf1 fix(windows): SSH ControlMaster gating + stop hijacking the user's python (#84452)
* fix(windows): SSH ControlMaster gating + stop hijacking the user's python

Two Windows environment-integrity fixes:

1. tools/environments/ssh.py (#73927): Windows OpenSSH has no
   Unix-domain-socket ControlMaster support, so unconditionally passing
   ControlPath/ControlMaster/ControlPersist failed EVERY tool call on a
   Windows-hosted ssh terminal backend with 'getsockname failed: Not a
   socket'. Gate the three multiplexing options behind a module-level
   _SSH_MULTIPLEX = (os.name != 'nt'); the scp upload path is gated the
   same way. On Windows the backend now works without connection pooling
   (each command a fresh connection); POSIX behavior is unchanged. The
   teardown 'ssh -O exit' is naturally inert because the socket never
   exists on Windows.

2. scripts/install.ps1 (#83797): the installer put the whole
   venv\Scripts directory on the user PATH, which contains python.exe /
   pythonw.exe / pip.exe and so silently hijacked the 'python' command in
   every terminal on the machine — unrelated projects started resolving
   python to Hermes' runtime interpreter. Now copy only the launchers
   (hermes.exe, hermes-acp.exe) into a dedicated $InstallDir\bin and put
   THAT on PATH. Existing installs are migrated: the legacy venv\Scripts
   entry is stripped from the user PATH on the next install/update. The
   new bin dir is under $InstallDir (…\hermes-agent), which the uninstall
   PATH sweep already matches via its \hermes-agent marker.

Updated the stale hermes_cli/update_cmd.py docstring that described the
old venv\Scripts-on-PATH layout.

Tests: SSH ControlMaster gating pinned both directions (multiplex on →
flags present; off → absent but BatchMode/StrictHostKeyChecking retained).
install.ps1 parses clean via the PowerShell AST parser.

* docs: update windows-native install docs for the bin\ launcher layout

CI (test_windows_native_docs) pins the docs and installer to the same
PATH layout. The #83797 fix moved the PATH entry from venv\Scripts to a
dedicated $InstallDir\bin holding only the hermes launchers, so update
the Windows-native guide to match: PATH-after-install section, the
install-steps list, the directory-layout table, the Get-Command
verification line, and the 'command not found' pitfall. Test now asserts
the bin\ layout and guards against a regression back to venv\Scripts on
PATH.

* fix: keep install.ps1 pure ASCII (PowerShell 5.1 codepage safety)

The two comments I added in the #83797 PATH-hijack fix used em-dashes,
tripping tests/test_install_ps1_ascii_only.py — Windows PowerShell 5.1
reads a BOM-less .ps1 in the system ANSI codepage (not UTF-8), so a
non-ASCII byte can misdecode into a stray quote and desync the parser
(issues #66994/#67000). Replace the em-dashes with ASCII '--'.
2026-08-12 02:56:33 -07:00
ethernet 2b39b885d6 test(install-e2e): macos desktop-installer arm - the published dmg, driven for real
macos gains the desktop-installer@latest install method: the website's
Hermes-Setup.dmg (verified live), mounted with hdiutil and its app
binary run DIRECTLY - an open-launched app inherits none of the git
redirect env, so direct exec is what keeps the isolation honest while
staying the same binary and first-launch flow.

install-e2e-macos-run.yml takes the windows shape: one workflow, one
inner job per driver arm, native skips. Arm 1 delegates script installs
to the shared OS-agnostic run workflow; arm 2 stages, installs from the
dmg, and drives both app-update methods through launch-from-spec.mjs -
open-app-update launches the installed .app (the double-click surface,
env via Playwright), hermes-desktop-app-update captures the product's
own hermes desktop spawn. Both end on sha asserts, never version
strings.
2026-08-12 04:36:03 -04:00
ethernet adf7d55f4b ci(install-e2e): windows composes install x update - one driver, one job
windows-desktop-gui-e2e.ps1 and windows-installer-script-e2e.ps1 fold
into tests/install/windows-e2e.ps1 with orthogonal -InstallMethod and
-Route axes: the install phase dispatches on one, the update phase on
the other, and shared workroot state carries how OLD landed - so any
implemented update method can follow any implemented install method.
Implementing a new pair is now a driver function plus a gate edit,
never a new job.

The run workflow collapses to ONE inner job whose if: is the
implemented-pairs table. Newly cheap pairs go live with the merge:
  desktop-installer@latest -> hermes-update / installer-script /
    installer-script+desktop / hermes-desktop-app-update
  installer-script(+desktop) -> hermes-desktop-app-update
  installer-script+desktop -> open-app-update (the -IncludeDesktop
    install registers real Start Menu / Desktop shortcuts)
Only desktop-installer@latest as an UPDATE method stays a declared
TODO. scripts/windows_e2e_harness.ps1 executes the parse/parameter/
dispatch checks under pwsh before any Windows runner spins up.
2026-08-12 04:30:12 -04:00
ethernet 0f903e14a3 test(install-e2e): hermes-desktop-app-update goes live on the script driver
Playwright must own the spawn (it needs the inspection pipe), but
hermes desktop is not just build+launch - stamp checks, integrity
gates, sandbox fixups, and a constructed child environment. So the
driver intercepts the product's own launch: a sitecustomize.py on
PYTHONPATH (opt-in via HERMES_E2E_CAPTURE_LAUNCH) wraps subprocess.run,
captures argv/cwd/env at the spawn site, and fakes success instead of
spawning; launch-from-spec.mjs then _electron.launch-es exactly that
spec and clicks Settings -> About -> Update now. Completion is product
state, not a Playwright event: the handoff result file or the checkout
reaching the expected sha (source installs write no result file).

Ships with the driver, so it works unchanged on every sampled OLD ref
- no product flag, no pre-flag fallback split. Both launch shapes are
matched (npm exec electron / packaged exe under apps/desktop/release);
npm BUILD calls pass through untouched. Exit 0 without a capture fails
the leg: a version that never reached its launch must not pass.

Probe-the-probe: scripts/launch_capture_probe.sh runs control rows
(no opt-in, non-launch argv) and both treatment shapes - all green
locally. Gate flips on the shared run workflow for linux/macos;
windows adopts the same path with the driver restructuring.
2026-08-12 04:20:56 -04:00
ethernet 239523414e test(install-e2e): installer-script+desktop is its own install and update method
The one-liner with its desktop stage opted in (--include-desktop /
-IncludeDesktop) is a real install kind, distinct on both sides:
on windows the stage builds Hermes.exe AND registers Start Menu /
Desktop shortcuts - a second path to a hand-launchable app - while
on linux/macos it builds into the checkout and registers no OS
entry point.

Declared on every OS and driven by both script drivers: the drivers
pass the flag through (hard failure if the ref predates it - the
tag-has-desktop gate already skips pre-desktop tags upstream) and
assert the built app exists under apps/desktop/release afterwards.
The run-workflow gates run +desktop pairs only on desktop-bearing
tags; app-update pairs from +desktop installs stay declared TODOs.
2026-08-12 04:01:16 -04:00
ethernet 1af2093663 test(install-e2e): split app-update into open-app-update + hermes-desktop-app-update
The desktop app has two launch paths, so app-update becomes two
methods. open-app-update starts the app from the OS entry point the
desktop installer created (the installed exe / the .app), so it exists
only where a desktop installer does. hermes-desktop-app-update starts
the app via hermes desktop, which every install method provides on
every OS that ships the desktop app - on linux it is the only app
surface, since no desktop installer or packaged artifact exists there.

Both variants are desktop-surface methods on every OS, so the
tag_has_desktop annotation moves from windows-only to every matrix
entry, install-e2e-run.yml grows the input, and the plan chart marks
pre-desktop cells on all OSes.

The windows GUI arm's implemented pair renames to open-app-update;
every other new combination is a declared TODO that natively skips.
2026-08-12 03:47:27 -04:00
ethernet 3cd6636471 fix(install-e2e): stray paren broke the generator - every node invocation died 2026-08-11 23:24:00 -04:00
ethernet 8e6f6d863c ci(install-e2e): result chart on the run summary - conclusions per combination x tag
generate-e2e-matrix.mjs grows --format results: reads the run's own
job list as NDJSON {name, conclusion} on stdin (per-leg conclusions
are NOT reachable through needs - a matrix job collapses to one
aggregate result) and re-renders the plan chart with each cell's
outcome. Legs are recognized by the exact name shape buildMatrices
mints, so unrelated jobs fall out; duplicate leg names (one windows
job per driver arm, only one runs) merge by significance - real
outcomes beat skips, failures beat successes. A final report job
(if: always, needs all three OS jobs) appends the chart to its step
summary via gh api with the default token.

Verified against two real runs: 31536931863 renders 11 passed / 0
failed / 54 skipped all-green; 31557865241 (the pre-EAP-fix run)
renders its 4 real failures + cancellations over the sibling arm's
skips.
2026-08-11 23:19:45 -04:00
ethernet db969ce696 ci(install-e2e): result chart on the run summary - conclusions per combination x tag
generate-e2e-matrix.mjs grows --format results: reads the run's own
job list as NDJSON {name, conclusion} on stdin (per-leg conclusions
are NOT reachable through needs - a matrix job collapses to one
aggregate result) and re-renders the plan chart with each cell's
outcome. Legs are recognized by the exact name shape buildMatrices
mints, so unrelated jobs fall out; duplicate leg names (one windows
job per driver arm, only one runs) merge by significance - real
outcomes beat skips, failures beat successes. A final report job
(if: always, needs all three OS jobs) appends the chart to its step
summary via gh api with the default token.

Verified against two real runs: 31536931863 renders 11 passed / 0
failed / 54 skipped all-green; 31557865241 (the pre-EAP-fix run)
renders its 4 real failures + cancellations over the sibling arm's
skips.
2026-08-11 23:16:09 -04:00
ethernet 80a0a198dd ci(install-e2e): plan chart on the run summary - combination x tag markdown table
generate-e2e-matrix.mjs grows --format markdown: one row per
{os, install -> update} combination, one column per starting tag,
appended to GITHUB_STEP_SUMMARY by the expand job. Cells mark
dispatched legs; run-vs-grey stays the run workflows' call, so the
only special cell is pre-desktop (the one annotation the plan owns).
JSON mode unchanged.
2026-08-11 23:06:42 -04:00
brooklyn! ed0e707914 Merge pull request #83634 from NousResearch/bb/handoff-window
Detached update hand-off on every OS: quit → hermes update → reopen, with one dumb shim window
2026-08-11 20:32:12 -05:00
ethernet ea4cd375f8 ci(install-e2e): retire the bubblewrap sandbox - git redirect everywhere, macos legs live
The fake Internet (bubblewrap + slirp4netns + MITM proxy +
upload-pack shim, 883 lines across dev-sandbox.sh, stage2-run.sh,
proxy.py, ssh-shim.sh, openssl.cnf, install-update-e2e.sh) existed to
isolate install.sh's network. The GIT_CONFIG_GLOBAL insteadOf redirect
the windows driver introduced does the same job with a gitconfig file
and works on any OS, so:

* install-e2e-run.yml now runs tests/install/installer-script-e2e.sh
  directly on the bare runner - no sandbox deps, no userns sysctls -
  and takes a runner input;
* the macos matrix calls the SAME workflow on macos-latest, deleting
  install-e2e-macos-run.yml: installer-script -> installer-script /
  hermes-update flip from grey to live, app-update pairs stay TODO
  inside the shared gate;
* install.sh is no longer curl'd through a fake CA - each leg runs
  the copy from the ref a user of that version actually executed;
* scripts/dev-sandbox.sh becomes the minimal isolation sandbox from
  ab6b9492f (separate HERMES_HOME / Electron userData / app name,
  same CLI surface: --persistent, --from, --delete), keeping its
  .hermes-sandbox dir name so gitignore and docs hold;
* nix/sandbox.nix drops the bwrap/proxy closure and keeps only the
  Electron runtime LD_LIBRARY_PATH the desktop app needs.

Verified: nix build .#sandbox + smoke run (isolated HERMES_HOME
created, ephemeral cleanup), shellcheck/bash -n on both scripts,
actionlint on all three workflows, and the new driver ran the full
v0.20.2 -> HEAD hermes-update pass locally before this commit.
2026-08-11 21:20:42 -04:00
Teknium 10e9da6f2d fix: ASCII-only install.ps1 comment; allow-list install_cli's uv PATH fallback
- install.ps1 must stay pure ASCII (PowerShell 5.1 ANSI code-page
  decoding, #66994/#67000): em-dash -> '--'
- tests/test_managed_runtime_resolution.py: install_cli()'s
  shutil.which('uv') is a reviewed fallback AFTER ensure_uv() misses
2026-08-11 17:06:15 -05:00
Teknium baa6b2e34d feat(browser): auto-install the Browser Use CLI instead of silently downgrading
The Browser Use CLI became the default browser backend, but nothing
provisioned it: users without uv/uvx (field report from DongyangHe on
macOS) silently fell back to the built-in browser tools with no notice.

- install_cli() in tools/browser_use_cli.py: uv tool install browser-use
  via the managed uv (bootstrapped on demand), linked into
  $HERMES_HOME/bin (UV_TOOL_BIN_DIR)
- _find_cli() now also probes $HERMES_HOME/bin for browser-use/uvx —
  Hermes' managed uv is not on the user's PATH
- hermes tools post_setup actually installs (Camofox standard) instead
  of printing instructions
- install.sh / install.ps1 provision the CLI at install time
  (best-effort, non-fatal, honors --skip-browser)
- CLI startup shows a one-line notice (24h rate-limited) when the
  default backend downgraded to the built-in tools
2026-08-11 17:06:15 -05:00
ethernet 6d7ba86963 ci(install-e2e): collapse the method vocabulary - 3 install ids, install+2 update ids
Per review the unions were overcomplicated. Install methods are now
just: installer-script (the platform one-liner - curl | bash on
linux/macos, irm | iex on windows), desktop-installer, and
packaged-app (declared, unused). Update methods are every install
method (re-run it over the existing install) plus hermes-update and
app-update. desktop-installer-rerun, desktop-app, curl-bash, and
irm-iex are gone as ids; the windows driver's ValidateSet, switch
arms, and both run workflows' gates renamed to match. tsc --checkJs
clean; generator output re-verified (4 linux / 16 windows / 6 macos
legs for 2 tags).
2026-08-11 17:14:17 -04:00
ethernet 17dea59026 ci(install-e2e): generator drops runtime validation for jsdoc type unions
Per review: the method/version vocabulary is now closed TYPE unions
(@ts-check + jsdoc typedefs - InstallerVersion, InstallMethod,
UpdateMethod - checked with tsc --checkJs, which rejects a SPEC entry
outside the unions; verified by corrupting a copy: 6 errors) instead
of runtime KNOWN_METHODS/ALLOWED_VERSIONS sets. validateEntry and
routeWants are deleted with all the paranoia: the generator always
emits every OS matrix and the dispatch route filter moved to plain
job-level ifs in install-e2e.yml, where the OS jobs already live.
secondUpdate is typed never[] so declaring one is a type error until
a leg implements it. Anything types cannot catch is self-evident on
the next CI run.
2026-08-11 17:11:07 -04:00
ethernet 1c60bdc30d ci(install-e2e): back to the generator - workflows stay generic, names carry everything
Revert the hardcoded 16-job experiment: the combination spec belongs
in scripts/sandbox/generate-e2e-matrix.mjs (restored), not copy-pasted
YAML blocks. What survives from the experiment:

* leg names carry everything - 'os: install -> update (tag -> HEAD)' -
  generated per entry, since slash-joined names are all the graph
  renders;
* pick-releases annotates each tag ({ref, desktop}) and the generator
  threads tag_has_desktop onto windows entries, so the windows run
  workflow still gates pre-desktop tags without a probe job;
* the per-OS run workflows are untouched: single job, static 'e2e'
  name, native skip gates own all capability knowledge.

install-e2e.yml is one generate job + three per-OS matrix fanouts.
Generator shape (4/16/12 legs for 2 tags), annotation threading, and
all six error paths verified; all four workflows pass actionlint;
driver parses clean pure-ASCII.
2026-08-11 16:54:54 -04:00