A manually-launched `hermes serve --host <ip>` powering a remote Desktop
was invisible to the entire update pipeline: not in the runtime
inventory, a permanent exit-2 dead-end at the Windows venv-holder guard,
and — when anything killed it — never relaunched, stranding the remote
client on a dead endpoint (#63206). Serve backends were also visible to
`hermes dashboard --stop` but hidden from `--status` (#81564's
asymmetry), so operators could kill what they couldn't see.
Built on the spawn ledger (positive identity, never argv guessing):
- process_identity.py: LedgerEntry gains structured host/port/profile
(backward-compatible — readers .get()); register_self accepts detail=;
argv capture widened 6→10 tokens so profiled launches survive.
- web_server.py: serve/dashboard registration moved AFTER the bind and
now records the ACTUAL bound host/port/profile.
- update_inventory.py: serve/dashboard collector reading the ledger —
manual backends inventory as supervisor=manual-serve with
restart_via=respawn-argv; Desktop-owned ones (live recorded spawner)
as desktop. Plan/receipts/fleet matrix see them for free.
- update_cmd.py: new venv-guard rung — manual serve/dashboard holders
are stopped for the update and relaunched via an idempotent atexit
token built from structured identity (same contract as the gateway
pause/resume); receipts record serve_pause/serve_relaunch.
Desktop-owned backends keep the refusal (the app respawns what we
kill).
- dashboard_procs.py: the process scan is augmented with live ledger
rows, so profiled launches (`hermes --profile p serve ...`) that match
no substring pattern are finally visible to kill/respawn.
- main.py: `--status` now lists serve-mode backends too, tagged [serve]
— closing the #81564 status/stop asymmetry.
Salvage note: detection deliberately does NOT reuse #70742's psutil
cmdline-pattern scan (the argv-guessing class this campaign retires);
its resume-token lifecycle (atexit + idempotent flag) and don't-replay
guard shaped the relaunch contract here — credit @Tranquil-Flow.
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
Reverts the interpreter-anchor halves of #95131 and #95478 (the anchor
module, its doctor check, and the update-time refresh). On real Macs the
anchored real-file copy of the uv interpreter dies in dyld: its LC_RPATH
(@executable_path/../lib) resolves into venv/lib/, which holds no
libpython — bricking EVERY hermes command including update and doctor
(#95425), and the re-pointed python3 aliases lost the stdlib
(ModuleNotFoundError: encodings, #95541). Linux CI could not catch this:
the fixture interpreters were one-byte fakes with no dynamic linking.
Kept: managed_uv._macos_sign_managed_python (#82529, @notkisk) — the
identifier-DR signing of repair generations is independent of the anchor
and unaffected by the dyld issue (it signs binaries IN PLACE in their
store, where their rpath is valid).
Added: doctor's check_macos_tcc_anchor_removed() heals venvs the anchor
already converted — restores bin/python to a symlink at the recorded
source (the anchor's own marker file) and re-points aliases; prints the
manual one-liner if the heal itself fails. Users whose CLI is fully
bricked can run the workaround from #95425 directly.
Re-land criteria: a dylib-complete anchor design (bundle libpython or
rewrite LC_RPATH), verified on macOS hardware BEFORE merge. Credit to
@kim-miram (#95358), @kokhlo (#95476), @zengzheqing (#95551) for the
forward-fix diagnoses that mapped the failure, and to the #95425/#95541
reporters.
The #91080 harness narrowed the #90950 corruption class to 'something
unlinks/replaces state.db or its -wal/-shm while another process holds
them'. The restore paths are exactly that actor:
- restore_quick_snapshot's unlink+move fallback replaced the inode and
deleted sidecars under a live holder -> deleted-fd split brain.
- update auto-restore copy2'd a snapshot over a possibly-live DB ->
the live writer's next checkpoint writes wrong-offset pages (the
page-1 compression_locks clobber from the report).
Add _foreign_db_holder_pids(), a /proc fd scan that counts holders of
the DB and its sidecars (including already-deleted generations), and
refuse the destructive replacement while any holder exists. The
backup-API path (safe under live connections) remains the primary
route for /snapshot restore.
Addresses review feedback on the regression test. The test previously parsed
the update_cmd.py AST to assert that each auto-restore call site cleared the
destination's sidecars before copying. That bound the fix to source text rather
than behaviour, and would break on unrelated refactors.
Extract _restore_state_db_from_snapshot(state_path, snap_state), which performs
the clear -> copy -> verify sequence as one unit and returns whether the
restored file passes its integrity check. Both auto-restore paths now call it,
so the ordering is guaranteed by construction instead of by inspection, and the
two byte-identical blocks collapse to a single call each.
The regression test now exercises that helper directly against a database that
still owns a hot WAL: removing the clear from inside the helper fails it with
201 rows where 400 were expected, so the guard remains bound to behaviour.
Also covers the two failure modes the callers already handle: a snapshot that
does not survive the copy returns False, and a missing snapshot raises OSError.
The post-update integrity guard (#68474) restores state.db from a pre-update
quick snapshot with a plain shutil.copy2, at both auto-restore sites: the
ZIP-update path in _update_via_zip and the git-pull path in _cmd_update_impl.
The snapshot image is produced by backup._safe_copy_db through sqlite3.backup(),
so it is already checkpointed and owns no WAL. That is precisely why
backup._EXCLUDED_SUFFIXES refuses to ship -wal/-shm/-journal inside a snapshot:
"shipping the live WAL / shared-memory / rollback-journal alongside would pair a
fresh snapshot with stale sidecar state and produce a torn restore on the next
open." The backup side excludes sidecars for that reason; the restore side never
cleared the destination's.
copy2 replaces only the main database file. A state.db-wal belonging to the old,
corrupt database survives the copy and is replayed over the fresh image on the
next open. The restored file then passes PRAGMA integrity_check while serving
the discarded database's contents, so _restored_ok reports valid and the CLI
prints "Auto-restored from snapshot" over data the user has lost. The first
subsequent checkpoint folds the stale WAL in permanently.
A hot -wal is reachable at exactly this moment: a second Hermes holder the
updater's drain did not stop, or the crash that corrupted state.db in the first
place, which is the very trigger for this code path.
Clearing the destination's sidecars is safe here specifically -- they belong to a
database the caller has already declared corrupt and is about to discard.
Contrast preflight_db_writability, which correctly refuses to delete a live WAL.
Reproduced against real SQLite: restoring a 400-row snapshot over a database
with a hot WAL yields 0 of the 400 rows, all 201 visible rows coming from the
old WAL, with integrity_check reporting ok.
Review feedback (AI review on #86391):
- guard _macos_desktop_dr subprocess.run against TimeoutExpired/FileNotFoundError
so a hanging codesign degrades to the unreadable-DR warning, never crashing
the doctor run (matches the file's existing subprocess guard pattern)
- select the desktop bundle by newest-mtime across release/mac-*/Hermes.app,
matching _desktop_packaged_executable, instead of a fixed arch order
- note the cdhash-match proxy assumption at the classification site
- document why /Applications/Hermes.app (Hermes-Setup launcher,
com.nousresearch.hermes.setup, certificate-anchored) is deliberately not probed
- extend the repair hint to cover per-service resets
- regression tests: codesign timeout and missing-codesign paths
TCC keys permission grants to the app's code-signing requirement. Grants
made to pre-#73681 builds carry a cdhash-pinned requirement that no
longer matches the rebuilt bundle, so macOS re-prompts on every capture
even though the System Settings toggle shows ON — and the modern prompt
has no Allow button, so users cannot complete the one-time re-grant.
- hermes doctor: new check_macos_tcc_grants() reports the desktop
bundle's DR class (cdhash-pinned → grants reset on every update;
identifier-pinned → stable) and prints the exact stale-grant repair
(tccutil reset, toggle ON, fully quit & relaunch).
- hermes update: after a successful update on macOS with a desktop app
installed, print the one-line stale-grant guidance.
- docs: desktop.md no longer claims grants persist 'out of the box';
documents the one-time re-grant for pre-fix grants.
Closes#86385
On Windows, the pre-update concurrent-instance gate aborted with exit 2
whenever ANY other process held the venv hermes.exe shim — including the
gateway itself, which _pause_windows_gateways_for_update() stops moments
later and the post-update restart phase brings back. Users with a running
gateway were forced into a manual taskkill dance before every update.
The gate now filters gateway runtimes out of the abort list and proceeds
when nothing else is concurrent. Classification delegates to
_is_pausable_gateway -> gateway.status.looks_like_gateway_command_line
(the canonical shlex-tokenized, profile-selector-aware matcher shared by
the Desktop preflight exemption and the venv-holder guard fallback), so
the gate's exemption and the pause machinery cannot drift apart. Anything
not positively identified as a gateway — REPLs, dashboard, Desktop
backend children, gateway MANAGEMENT commands like 'gateway status',
unreadable cmdlines — still aborts exactly as before, and the abort
message now lists only the PIDs that are actually the user's problem.
Surgical reapply of PR #37039 by @damadorPL onto current main (the gate
moved from hermes_cli/main.py to hermes_cli/update_cmd.py in the main.py
decomposition); his substring classifier was replaced with the canonical
matcher, which also fixes the 'hermes gateway status' misclassification
flagged in review.
Co-authored-by: Hermes <hermes@nousresearch.com>
Port of @jeff-mettel's fix onto the post-#91378/#92902 fleet-restart
shape. The current-profile restart was gated on `launchctl list <label>`
exiting 0 - a booted-out job (plist present, definition deregistered:
crashed helper, manual bootout, failed prior update) fails that check,
so the branch silently skipped: no restart, no message, KeepAlive unable
to revive a definition launchd no longer knows, update printing
'Update complete!' with the gateway down. `launchctl list` is also
session-scoped and unreliable as a loaded/unloaded classifier.
- _restart_launchd_gateway_after_update() (his extraction, adapted):
plist-exists is the ONLY gate; launchd_restart() owns the
bootout/bootstrap/kickstart ladder for every plist-present state;
every failure path is loud and names the manual recovery command.
The gate-error 'except: pass' (the second silent variant) now counts
the label failed and tells the operator.
- Success still requires the #92902 supervision verify (fresh
supervised PID), composing his fix with the returned-is-not-supervised
guard.
- His regression suite adapted to the (restarted, failed) contract; the
old 'unregistered -> left alone' pinning test FLIPPED - it pinned the
bug.
A/B: his suite + the flipped test red on merge-base product code
(silent skip live), green at head. No macOS CI lane exists; field
evidence is #74973's reproductions plus the launchctl print output
shapes pinned in the suite.
The #93410 guard keyed on (restarted_services or killed_pids), which never
fires on Windows: _pause_windows_gateways_for_update /
_resume_windows_gateways_after_update populate neither list, so a healthy
resumed Windows gateway still yielded zero fleet rows and exit 0.
Hoist the decision into _fleet_probe_expected_runtimes(), keyed on every
pre-update liveness signal:
- restarted_services / killed_pids (POSIX restart bookkeeping)
- _pre_restart_gateway_pids non-empty or None (unreadable pre-state,
same fail-closed contract as _restart_phase_failure_is_incomplete, #78574)
- pre-update plan inventoried >=1 runtime
- Windows pause/resume token carries profiles or unmapped entries
Gate the 2.0s settle sleep on the same condition so a resumed Windows
gateway gets its settle window before the probe. The guard keys only on
zero-rows-despite-expected-runtimes; non-empty snapshots (including
'unknown'-state rows) are still judged solely by print_fleet_version_matrix.
Regression tests cover: empty snapshot + plan runtimes -> incomplete;
empty snapshot + genuinely idle -> success; Windows-resume token path ->
fail-closed + settle sleep wiring.
Builds on RelaxJonh's #93410. Fixes#93406
collect_fleet_versions() swallows every probe exception via
logger.debug() and returns whatever accumulated — which can be an
empty list. print_fleet_version_matrix([]) returns False (no rows
to report), so the update exits 0 with "success" even though no
gateway was actually verified.
After the restart phase touches live gateways (restarted_services or
killed_pids is truthy), an empty fleet snapshot means verification
failed, not that everything is healthy. Treat it as incomplete so
the receipt records "partial" and the exit code is 1.
Fixes#93406
The shared checkout serves every profile, but hermes update migrated
only the active profile's config.yaml. Siblings kept their old
_config_version until their (correctly restarted, post-#91378) gateway
hit a config shape the new code couldn't read — the last unabsorbed
substance from the Phase-2 restart-swarm audit (#20438 earliest, 2026
field repro on #79048: sibling at v33 vs v37).
_migrate_sibling_profile_configs(): per sibling home, scope config
reads/writes via the context-local HERMES_HOME override (ContextVar —
never os.environ), check version, run the NON-INTERACTIVE safe
migration; prompt-requiring settings stay for the profile's own next
interactive session (same contract as gateway-mode). Broken profiles
are skipped without blocking the sweep; override always reset.
Sabotage-verified; live E2E in a fresh process with real drifted
config files: v12→v38 and v25→v38 on disk, provider preserved, the
documented #81946 personality-reset migration correctly applied to
siblings too, never-configured profile untouched, active home
untouched, second run idempotent.
The policy table was observational: restart_via was a display string and
the four platform restart branches re-discovered their own targets, so a
runtime the plan saw could be missed with zero signal (the #88654 class,
structurally).
- update_inventory: restart_via becomes a machine-readable mechanism id
(systemd|launchd|desktop|manual) — THE policy table as data; display
derived via describe_restart_mechanism. match_runtime_outcomes()
reconciles every planned runtime against the restart phase's
bookkeeping (restarted/stopped/failed/unaccounted);
report_unaccounted_runtimes() is the silent-miss tripwire.
- update_cmd: after the restart phase, the plan is reconciled; outcomes
land in the receipt (runtime_outcomes); any unaccounted runtime
escalates exactly like a STALE/DOWN fleet row (exit 1).
Sabotage-verified (reconciliation forced to 'restarted' fails the
tripwire tests); live E2E on this host's real fleet: the real
systemd-supervised gateway classified with a machine id, reported
unaccounted when the bookkeeping omits it, clean when accounted.
On macOS, `hermes update` printed "Update complete!" and exited 0 while the
ai.hermes.gateway LaunchAgent sat deregistered for 36 minutes (#88848).
_restart_macos_launchd_gateways already disagrees with itself about what
"restarted" means. Sibling profiles are only appended to restarted_services
once _wait_for_launchd_service_pid confirms launchd is running the job on a
fresh pid. The invoking profile was appended on "launchd_restart() did not
raise" alone.
That is a weaker claim than it looks. launchd_restart() returns as soon as the
restart has been REQUESTED: the _request_gateway_self_restart branch hands the
work to the running gateway and returns immediately, and a plist reload is
handed to a detached helper. Both are asynchronous, so a helper that dies
before its first bootstrap, or a `launchctl bootstrap` that exits 0 without
registering (measured by the reporter on macOS 26.6.1), were both invisible to
the caller. The systemd branch of the same phase has never drawn that
inference: it polls _wait_for_service_active before recording the unit.
Verification is domain-agnostic via a new
gateway.wait_for_launchd_gateway_supervision, NOT _wait_for_launchd_service_pid.
The sibling helper needs an explicit domain, and the invoking profile's gate
deliberately avoids a domain locate because it fails on macOS-26 hosts whose
per-user domains reject service management even though launchd_restart() owns
that fallback. The new helper judges by a live supervised pid rather than an
exit code (the predicate _launchctl_label_supervising_process already existed;
this only adds the wait), and returns True immediately when the detached
fallback marker is present, because a gateway running unsupervised there is the
designed state and not the silent failure this guards against.
A label that restarts but is never supervised now lands in
failed_or_stale_units, which sets gateway_fleet_restart_incomplete and makes
the update exit non-zero instead of reporting success over a gateway that is
down.
Tests: 12 in tests/hermes_cli/test_update_launchd_restart_verification.py, with
no platform gate, driving the real _restart_macos_launchd_gateways through
mocked launchctl outcomes. Reverting the verification to an unconditional
append fails 2 of them, including the #88848 regression case.
tests/hermes_cli/test_update_launchd_fleet_restart.py::_fleet stubs the new
verifier so its 27 existing cases keep asserting on routing rather than on a
real launchctl probe; unstubbed, each case would poll the full supervision
budget.
Fixes#88654.
After an in-place update, the manual-gateway leg of the restart phase did
this for every profile-mapped gateway:
restart_mode = _prepare_profile_gateway_update_restart(proc.profile, pid)
if restart_mode is None:
continue
A None means no relaunch could be armed. The bare continue skipped the
drain and the stop, and the unmapped sweep immediately below skips any
pid already in profile_processes, so the process was never killed and
never counted into the "Stopped N manual gateway process(es)" summary.
The gateway kept running with its pre-update modules resident while the
new code sat on disk, and every lazy import from that point mixed
versions:
cannot import name '_MAX_TOOL_ERROR_CHARS' from 'tools.registry'
with no operator signal of any kind.
Two changes.
_prepare_profile_gateway_update_restart now falls back to replaying the
process's own captured command line when the profile-derived relaunch
cannot be armed. launch_detached_gateway_restart_by_cmdline already
exists for exactly this case and documents itself as the companion for
gateways with no profile mapping; the Windows post-update path already
uses it the same way. The argv is captured a few lines earlier for the
external-supervisor check, so the fallback costs nothing extra. The
external-supervisor branch still short-circuits first, because replaying
argv there would escape the manager and race its replacement process.
When neither mechanism can arm a relaunch, the update path no longer
falls through silently. It says so, naming the profile and pid, and hands
the process to the existing unmapped sweep so it is stopped and reported
through the established "Restart manually: hermes gateway run" contract.
Leaving it running was the actual harm: a gateway on stale modules fails
every lazy import for as long as it lives.
The salvaged commit called managed_python_env() at the git-path sync
without an in-scope import (UnboundLocalError on every git update — CI
red). The repair test pinned the raw {**os.environ, VIRTUAL_ENV} dict, a
change-detector on exactly the construction #83914 replaces; it now
asserts the managed-env contract.
The salvaged fix covered the git-path sync; the same raw-os.environ
construction existed at the main update path and the interrupted-install
recovery path. All three now build their uv env via managed_python_env()
(#83914 class — same bug, all sites).
A/B-proven with real uv: poisoned UV_PYTHON/UV_SYSTEM_PYTHON steers the
merge-base construction into the hijacker's interpreter (VERDICT:
HIJACKED); the managed construction installs into the install's venv
(VERDICT: ISOLATED). Compose-checked with #92824's stale-VIRTUAL_ENV pin:
isolation + pin together install into the running interpreter on the
site-packages shape.
Address review feedback:
- Add two unit tests asserting the update's uv_env contract: third-party
UV_PYTHON_INSTALL_DIR is dropped, managed pins (UV_MANAGED_PYTHON=1,
UV_NO_CONFIG=1) are set, VIRTUAL_ENV points at this install's venv, and
the managed store stays under .hermes-runtime.
- Drop the inline dated comment in favor of intent description.
uv respects UV_PYTHON_INSTALL_DIR from the process environment. When a
third-party app (e.g. WorkBuddy) sets a User-level UV_PYTHON_INSTALL_DIR,
the update's uv pip install can target the wrong interpreter and fail
installing extras, leaving the venv entry-point shims missing. Use the
official managed_python_env() isolation (drops VIRTUAL_ENV/PYTHONPATH/
UV_PYTHON, forces UV_PYTHON_INSTALL_DIR to .hermes-runtime/python,
UV_NO_CONFIG=1) and then point VIRTUAL_ENV at this install's venv.
On a fork, `hermes update` compares HEAD against origin/main, and only then
syncs the fork from upstream — inside the `commit_count == 0` branch, which
returns immediately afterwards. So an update that pulls hundreds of commits
from upstream prints "Already up to date!" and skips everything the
post-update path does, including the dependency sync and the gateway restart.
Observed on a fork-based deployment: 1654 commits pulled, "Already up to
date!", and the launchd gateway left running. It then held pre-update modules
in memory while lazily importing post-update ones, and failed later with an
AttributeError for a method that plainly exists on disk — a mixed runtime that
looks nothing like an update problem. Correlating every run in update.log, a
restart happened on exactly the runs that pulled upstream *without* also
claiming to be up to date, and never once they started co-occurring.
Decide before the branch: capture HEAD, sync, and if HEAD moved, set
commit_count from the range so the normal post-update path runs. The pull that
follows is a no-op (the sync updates origin too); reaching the restart is the
point. commit_count is floored at 1 — HEAD moving *is* the update, so a failed
or zero count query must not send us back down the early return.
steps still being skipped afterwards.
Refs #73108
Phase-1 verification gap (#91277, found auditing our own landed matrix
against the mapped issues): collect_fleet_versions only listed gateways
with a LIVE pid, so 'restart stopped it and nothing came back' produced
NO row at all — the exact silent-failure shape the matrix exists to
catch (#88848/#74973 class) passed with exit 0.
- collect_fleet_versions(pre_restart_pids=...): a dead pid becomes a
'down' row only when it was alive at update start AND its runtime
status still claims a running state. Rollout-safe: no snapshot (old
callers), clean stops, startup failures, and stale records from
long-dead gateways keep the historical no-row behavior.
- print_fleet_version_matrix escalates on down rows like stale ones
(exit 1) with the per-profile restart remediation.
- cmd_update passes its existing pre-restart PID snapshot.
Sabotage-verified (reverting the membership check fails the new test);
live-verified with a real spawned-then-killed process producing the
DOWN row and matrix escalation.
On top of @686f6c61's premise-corrected #76745:
- _looks_like_desktop_control_plane now uses the parser-derived
_hermes_holder_subcommand instead of substring matching — the
#90778/#91869 class ('-m dashboard chat' and 'kanban --preserve-cache'
argv no longer read as control planes). Regression test added,
sabotage-verified (reverting to substrings fails it).
- Live E2E (this host, real processes + real spawn ledger): live
supervised serve owns lifecycle; killed spawner (orphan) does not;
dead serve entry excluded; empty ledger does not.
- Live Windows E2E for the wine2e lane: real self-registered ledger
entry suppresses the actual cold-start plan; dead serve restores it;
holder-scan fallback rung proves the token classifier live.
Co-authored-by: 686f6c61 <github@00b.tech>
Vestigial autostart is not proof the user wants a standalone gateway
run. When Desktop currently supervises this install's control plane,
the updater must not spawn a competing messaging daemon. Serve is not
treated as gateway-equivalent.
The #87331 remaining half: when hermes.exe (or a sibling shim) could not
be renamed aside, the updater printed a warning and ran the installer
anyway — which died partway on the same locks and stranded the venv
between versions.
- _run_quarantined_install gains strict_quarantine: any shim whose
rename failed every retry aborts BEFORE the install command runs
(successful renames rolled back), raising ShimQuarantineError.
- The update dependency sync passes strict_quarantine=True. The update
boundary turns the error into a refusal: defer via the
update-incomplete marker, exit 2 (recorded as refused by the receipt
net), never ZIP-fallback. Post-sync repair installs keep warn-and-try
(their venv is already mutated; refusing buys nothing).
- The recovery installer (_install_repair._run_install_cmd) is strict
unconditionally: marker survives, next launch retries after the
holder exits.
- Live Windows E2E for the wine2e lane: a real child holds hermes.exe
without FILE_SHARE_DELETE (the exact field lock shape), strict path
refuses with zero installer invocations, releases roll back, and the
same path proceeds once the holder exits.
Sabotage-verified: reverting the strict wiring makes both fail-closed
tests fail.
PR #92092 fixed the same vanished-launcher bug by restoring copies into
the legacy in-checkout hermes-agent\bin from the update tail. That
location is what this branch removes: untracked files there are swept
by the update autostash on every cycle (restore/sweep treadmill, plus a
parked stash entry per update under --keep-stash), and unconditional
exe copies break on relocatable venvs ('uv trampoline failed to
canonicalize script path'). This branch's managed-binary-dir layout
supersedes both mechanisms, so the merge resolves to it:
- drop _sync_windows_cli_launchers and its _ensure_acp_launcher call
(Windows staging/repair lives in ensure_windows_bin_launchers at
process start and migrate_windows_bin_path in the update tail);
_ensure_acp_launcher is a Windows no-op again
- keep #92092's genuinely better installer semantics: staging stays in
a dedicated Install-HermesCommandLaunchers function that throws
BEFORE any PATH mutation when the required launcher cannot be staged
and verified -- previously Set-PathVariable could put an empty dir on
PATH and still print 'hermes command ready'. Reworked for this
branch's layout: caller passes the destination ($HermesHome\bin),
launcher form follows the venv (exe copy vs .cmd delegator), and the
verify step accepts either form
- rework #92092's AST-lifted PowerShell test for the new function
signature, keeping its fail-before-PATH-mutation assertions and
adding relocatable-venv form-selection coverage
- drop tests/hermes_cli/test_windows_cli_launcher_repair.py (pinned the
superseded in-checkout mechanism; equivalent and broader coverage
lives in tests/hermes_cli/test_ensure_windows_bin_launchers.py)
The installer staged the hermes/hermes-acp launcher copies at
hermes-agent\bin -- inside the git working tree -- and put that dir on
the user PATH (#84452). The update command's pre-pull autostash
(git stash push --include-untracked) swept those untracked, unignored
copies off disk, and once the desktop updater stopped re-applying
stashes (--keep-stash, 5dd221d442) nothing restored them: `hermes`
stopped resolving in every new terminal on every desktop-updated
install.
Move the canonical launcher home to the managed binary dir
(%LOCALAPPDATA%\hermes\bin, next to the managed uv) -- outside the
checkout, where no git operation can ever touch it. The dir is
per-machine and shared by every profile, so all anchoring uses
get_default_hermes_root(), never HERMES_HOME (which points inside
profiles\<name> under `hermes -p`).
The copy design also had a second latent break: managed-uv rebuilds
create relocatable venvs, and a relocatable venv's exe trampoline
resolves relative to its own location -- a copy outside venv\Scripts
dies with 'uv trampoline failed to canonicalize script path'. Launcher
form now depends on the venv (lockstep in install.ps1 and
_install_repair.py): exe copy for normal venvs, a .cmd delegator
invoking the in-venv exe by absolute path for relocatable ones. Either
form counts as present, so pre-rebuild exe copies are left alone.
Delivery to the existing fleet, per cohort:
- already-broken installs cannot run the CLI, so an import-time heal in
hermes_cli.main (ensure_windows_bin_launchers) re-stages missing
launchers when the desktop app spawns its backend -- the one channel
that still reaches them. Gates fail toward inaction: canonical dir
only for the managed clone, legacy hermes-agent\bin only while the
user PATH still resolves through it (some pre-managed-uv installs
have no hermes\bin PATH entry; the legacy re-stage is what fixes
those). Staging-name + os.replace keeps concurrent process starts
from tearing a launcher; the helper never raises.
- healthy old-layout installs migrate in the update tail
(migrate_windows_bin_path): stage canonical launchers, verify them
BEFORE touching the registry, prepend hermes\bin to the user PATH,
strip the legacy entries (hermes-agent\bin and venv\Scripts, #83797),
preserving REG_EXPAND_SZ and raw %VARS%. The legacy dir's files stay
on purpose -- configs holding absolute launcher paths keep working;
only the sweepable PATH resolution route goes.
- fresh installs get the new layout from install.ps1 directly.
/bin/ is gitignored so the one update that DELIVERS this fix cannot
sweep pre-migration launchers a final time under the old rules; the
gitignore line, the legacy re-stage branch, and the update-tail call
are transition machinery with a named expiry once the fleet has
migrated.
Also rewrites _ensure_acp_launcher's stale Windows paragraph to match
(raw docstring fixes its invalid \S escape) and updates the Windows
native docs to the new layout, with a docs<->installer parity test.
The #70337/#87331 win-unpacked wipe half, from PR #70477 by @JonthanaHanh
(reimplemented against the two-phase staged swap that postdates that
branch — the live release/ dir is grafted into the staged apps copy
BEFORE the atomic commit, so preservation rides the same rollback
machinery instead of a post-hoc copy).
Co-authored-by: JonthanaHanh <92574114+JonthanaHanh@users.noreply.github.com>
Surgical reapply of PR #87878 (@kshitijk4poor's salvage of #87327 by
@liruixinch) onto current main — the receipt-boundary and summary
changes from this session made the original commits conflict.
- ZIP fallback now keys on git ACTUALLY having failed
(_should_zip_fallback_on_update_error): a dependency-install failure
after a successful pull can't be fixed by re-downloading source and
would clobber the tree (#87331 cascade trigger, #87304).
- _abort_zip_update_if_dirty_tree: refuse to overlay a dirty checkout
(-uall so user gitconfig can't blind the guard) + pre-swap TOCTOU
re-check with our own staging artifacts filtered (#91962, #87304).
- Failure-stage naming (_format_update_failure_stage) + stderr tail so
'Git update failed' stops mislabeling pip/uv failures.
- Receipt finalize preserved on the no-fallback failure path.
Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
Co-authored-by: liruixinch <liruixinch@outlook.com>
Review on #91869 (@andrexibiza): the handwritten value_flags subset
misparsed '--reasoning high serve' as subcommand 'high' and
'-m dashboard serve' as 'dashboard' — recreating the wrong-hint class.
_holder_value_flags() now introspects build_top_level_parser() (every
option with nargs != 0, plus the pre-argparse profile selectors), with
a static fallback for broken-tree updates, --flag=value handled.
Regressions for --reasoning/-m/-t/--model=/-c per review.
De-flake test_goal_resume_restart: the fixture only set the HERMES_HOME
env var, but get_hermes_home() prefers the context-local override — an
override leaked by any earlier test in the xdist worker pointed the
goals DB at a dead tmp dir and resume enqueued nothing (the CI-only
red). Fixture now pins the override via set/reset_hermes_home_override.
Mechanism proven both ways: env-only fixture cannot beat a leaked
override; pinned fixture immune.
#90778: _hermes_holder_subcommand() — token-based parse of the actual
Hermes subcommand (profile selectors skipped, flags never matched), so
'hermes dashboard' stops being labeled as the Desktop backend and
'--preserve-cache' stops matching 'serve'. Unknown argv gets no hint
instead of a wrong one.
#87594: ancestor-exclusion in _detect_venv_python_processes and
_venv_launcher_ancestors now carves out GATEWAY ancestors (canonical
looks_like_gateway_command_line): when /update runs as the gateway's
child, the gateway stays visible to the scan so the pause machinery can
stop it, while shells/terminals/own-venv ancestry stay excluded.
15 cross-platform classifier tests; live Windows E2E suite is the
acceptance gate on this branch.
The in-function import made _time local to all of _cmd_update_impl, so
the orphan-backend reap path (which runs earlier in the function) hit
UnboundLocalError before the import line executed. The module-level
'import time as _time' at the top of update_cmd.py already covers the
divergence-merge safety tag.
Review feedback on #89507: in-place merging suits a branch that tracks the
target with a small patch set, but a long-lived feature branch (a PR branch
hundreds of commits deep) does not want an update-driven merge commit
written into its history. Reported against a checkout carrying 819 unmerged
commits.
--switch-branch routes the unmerged case to the switch path instead: the
checkout moves to the update target and updates there, and the branch is
left byte-identical — no merge, no commit, nothing written to it. The tree
is known clean on that path (the guard checks dirty before cherry), so a
dirty tree still gets the loud skip, unchanged.
Opt-in: without the flag the default remains the in-place update, which is
what keeps a small-patch-set branch's running code current.
Tests: the flag switches and leaves the branch tip byte-identical; the
default without it still updates in place. The first fails if the flag's
branch is severed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The parked-branch guard (8ce8ffd429) distinguishes checkouts by what the
branch carries, then treats both non-clean cases the same: a stale
fully-merged leftover is switched back to the target (correct), but a
branch with unmerged commits — a branch someone is actually working on —
gets CODE UPDATE SKIPPED and exit 1. For anyone running a maintained
custom branch on top of main, every update now refuses, and the guidance
('checkout main') abandons their branch.
The guard's own reason codes already separate the cases, so use them:
- fully merged -> switch back to the target (unchanged)
- unmerged:N -> update the branch IN PLACE: fetch, then bring
origin/<target> into the checkout. Fast-forward
when possible; on divergence, a true merge behind
a pre-update safety tag, stopping cleanly on
conflict. The checkout never moves; local commits
survive; the running code advances.
- dirty/unverifiable/opted out -> skip loudly (unchanged)
The post-pull success gate learns that an in-place update legitimately
ends on a non-target branch: origin/<target> was merged INTO the checkout,
so refusing to claim success there would fail every update that did
exactly the right thing.
Guard tests updated: the unmerged case now asserts the in-place outcome —
target code arrives (b.txt from c3), the branch's own commit survives, and
HEAD never moves. 18/18 guard tests, 20/20 with the diverged-update suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A clean checkout parked on a feature branch now always switches to the
update target. Unmerged commits are safe on the branch (git checkout
never discards committed work) and get a loud 'kept' notice naming the
branch, count, and the checkout command to resume the work. Previously
the update hard-skipped with exit 1 — a dead end for the desktop update
button, gateway /update, and cron, which have no way to resolve a skip.
Dirty trees (uncommitted changes) still skip loudly, and the
updates.auto_switch_parked_branch: false opt-out still pins the branch.
The code swap and gateway fleet restart touch all profiles, but the
pre-update quick snapshot photographed only the invoking profile's home
— siblings had no snapshot for the post-update safety nets or manual
restore to draw on.
- backup.py: create_pre_update_snapshots_all_profiles() — the SAME
snapshot set, per-file 1GiB cap, and keep policy as the invoking
profile (no partial tier, no new restore-coherence class), each into
the sibling's own state-snapshots/; restore_cron_jobs_all_profiles()
runs the #34600 cron-loss safety net per profile against its OWN
snapshot (same-generation by construction).
- update_cmd.py: sibling snapshots taken right after the invoking
profile's (best-effort, receipt-recorded); post-update cron restore
extended to every sibling.
- Docs: updating.md pre-update snapshot step now states the per-profile
behavior and the file-loss-recovery vs rollback contract.
- 9 unit tests + E2E (real files: sibling snapshot on disk, clobbered
jobs.json restored 7/7 from the sibling's own snapshot, keep=1 prune).
Phase 2 core slice of #91277: the updater now knows WHAT it is operating
on before it mutates anything.
- hermes_cli/update_inventory.py (new): side-effect-free runtime
inventory — install kind via detect_install_method (git / docker / nix
/ apt, updatable-in-place or not, with the correct external update
command for image/package-managed installs), all profiles, every live
gateway with its supervisor (systemd / launchd / manual via the
fleet-wide _get_service_pids), running code_sha/code_version from the
#91283 gateway_state.json stamps, and the restart mechanism each
runtime will get.
- hermes update --plan: prints the plan and exits; runs BEFORE the
docker/nix refusal gates so image-managed installs get a useful
'not updatable in place + right command' report instead of a bare
refusal. Read-only, safe on a live fleet.
- Every real update run now records the pre-update plan in its receipt
('plan' key) and prints a one-line fleet summary, so post-mortems can
compare what the update SAW against what it did.
- Docs: updating.md (--plan section + receipts/fleet-check section),
cli-commands.md (flag row + receipts behavior bullet).
- 11 tests: two-profile fleet classification, docker not-in-place,
dead-PID exclusion, PID-file fallback dedupe, all-probes-fail
never-raises, JSON round-trip for the receipt, print output shapes,
receipt integration.
The macOS branch of the update's fleet-restart step only restarted the
invoking profile's LaunchAgent. Sibling ai.hermes.gateway-<profile>
services kept pre-update modules cached in sys.modules and died on their
next agent turn (ImportError on new lazy imports, or TypeError/
AttributeError with garbled tracebacks on wider version gaps). The
systemd branch already iterates every hermes-gateway* unit; this brings
launchd to parity:
- _restart_macos_launchd_gateways(): the invoking profile keeps the
existing launchd_restart() path; every other gateway of this install
is drained via SIGUSR1 (same as systemd siblings), then hard-
kickstarted unless KeepAlive already respawned it, then verified on a
fresh PID. TimeoutExpired is isolated per label (#68523 parity) and
counts toward failed_or_stale_units — including timeouts during
liveness discovery, which must not read as "unloaded".
- Install-scoped fleet enumeration: launchd_gateway_labels_for_install()
derives labels from THIS install's profiles (get_default_hermes_root),
not by globbing the shared per-user ~/Library/LaunchAgents — a
sandboxed HERMES_HOME (tests, capture sandboxes, side-by-side
installs) must never enumerate, let alone restart, another install's
fleet. This also keeps the hermetic test suite blind to a dev
machine's real gateways.
- Domain-explicit sibling handling via _locate_launchd_gateway_service():
liveness, kickstart, and fresh-PID verification all use the domain the
service was actually located in (gui/<uid> vs user/<uid> probed per
label via `launchctl print`). This addresses the #41403 review defect:
the process-wide _launchd_domain() cache resolves the current profile's
domain and must never be reused for a sibling. _launchd_domain() itself
becomes a thin caching wrapper; behavior unchanged.
- _get_service_pids(all_profiles=...): the update path's manual-process
sweep excludes every gateway service PID (mirror of the systemd
hermes-gateway* pattern) so it cannot mistake a freshly respawned
sibling service for a stale manual gateway. Default-scope callers
(gateway status, cron checks, stop_profile_gateway's orphan reaper —
which kills what it is fed) keep the current-profile-only contract.
- _warn_incomplete_gateway_fleet_restart() prints launchctl recovery
hints for launchd labels alongside the systemctl ones.
Supersedes and completes #41403, addressing its review feedback
(per-label domain resolution + mocked regression tests).
Co-authored-by: David Neyra <vyr.agent@vyrgs.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Replace ps -A eww with ps -Aww: the BSD e flag is illegal
on macOS/BSD ps, making the fallback silently return [] on every macOS
machine. The matcher only needs argv (not env vars), so e is
unnecessary. -ww keeps unlimited-width output on both BSD and
procps ps.
- Add all_profiles parameter to _get_service_pids(). When True
on macOS, enumerate every ai.hermes.gateway* launchd agent across
profiles via bare launchctl list instead of only the current
profile's label. This prevents the update sweep from misclassifying
sibling-profile launchd gateways as manual processes (#73626).
- Thread all_profiles through find_gateway_pids() to
_get_service_pids().
- Update two _get_service_pids() call sites in update_cmd.py to
pass all_profiles=True so the update fleet sweep excludes every
service-managed gateway across all profiles.
- Add TestPsFallbackBsdCompat: verifies ps argv uses -Aww
not -A eww, and that pid=,command= output columns are present.
- Add TestGetServicePidsAllProfiles: verifies default scope uses
launchctl list <label>, all_profiles uses bare launchctl list
with prefix filtering, handles empty/broken output gracefully, and
preserves systemd behavior.
Tranquil-Flow
get_code_identity() shelled 'git rev-parse HEAD', which broke two tightly
mocked test suites (sequenced subprocess.run side effects in the
head-moved gate, call-count asserts in the Windows taskkill test) and
added process-spawn cost to gateway runtime-status writes.
_resolve_git_head_sha() now reads HEAD/refs/packed-refs directly,
handling regular checkouts and worktree/submodule pointer files.
Also: skip the 2s fleet settle wait when the restart phase touched no
gateways, and hoist killed_pids init outside the restart try-block.
Phase 1 of the fleet-update reliability plan (#91277): the updater now
proves its outcome instead of assuming it.
- hermes_cli/build_info.py: get_code_identity() — process-cached code
identity (git sha for source installs, baked .hermes_build_sha for
Docker images, pyproject version).
- gateway/status.py: every runtime-status write stamps the writer's
code_sha/code_version into gateway_state.json, so a running gateway's
actual code generation is observable from disk.
- hermes_cli/update_receipt.py (new): machine-readable receipt of each
update run (steps, skips with reasons, gateway restart outcome, fleet
snapshot) under ~/.hermes/logs/update_receipts/ with a latest.json
pointer for the dashboard/desktop; plus collect_fleet_versions() /
print_fleet_version_matrix() comparing every live profile gateway
against the freshly updated checkout.
- hermes_cli/update_cmd.py: wires receipt begin/steps/finalize into the
git, ZIP, and hard-failure paths; after the restart phase, prints the
fleet version matrix and escalates provably-stale gateways into the
existing gateway_fleet_restart_incomplete exit-1 contract. Pre-stamp
gateways report 'unknown' and never fail the update (no false
positives during rollout).
Silent-failure classes made visible: #88848, #74973, #85753, #81193.
Mixed-version fleet classes made loud: #88654, #69754, #77553, #56717.
`uv pip install -e .` never audits an editable target. It reinstalls on every
invocation and rewrites the console-script shims each time, which is the only
reason `hermes update` has to quarantine the running `hermes.exe` on Windows —
and a quarantine that loses its race is the whole `os error 32` family.
Gate the reinstall on whether the pull actually touched a file that defines the
install. It's safe to skip because the editable finder is pinned to a static
module list (`py-modules` + `packages.find.include`), so the one source-only
change that could stale it — a new top-level module or package — cannot land
without a `pyproject.toml` diff. Dependencies and `[project.scripts]` live
there too, and new submodules inside an already-mapped package resolve through
the real directory.
The predicate fails closed: no pre-pull SHA, an unresolvable one, or a failed
`git diff` all reinstall as before. On the skip path the two verifiers that
normally run inside the install run directly, so a wrong skip self-heals into a
real install rather than leaving an unchecked venv.
This is the pattern the file already uses everywhere else — `_tui_need_npm_install`
diffs node_modules against package-lock.json, and the desktop build is gated on a
content hash so `hermes update` "will skip if nothing actually changed". The
Python editable install was the one path with no such gate.
`hermes update` runs in the pre-pull interpreter. The auto-restart phase
imports freshly-pulled gateway source, which resolves sibling imports
against the OLD sys.modules cache — so any update where an already-cached
module gained a new export ImportErrored the whole phase and left the
gateway serving pre-update code (2026-08-20 field failure: new gateway.py
needs cli_output.line_input, cached cli_output predates it).
Class fix replacing the per-symptom _UPDATE_RUNTIME_RELOAD_MODULES
approach: _purge_stale_hermes_modules() evicts every cached module under
the Hermes package prefixes (hermes_cli/gateway/tools/tui_gateway/agent)
right before the restart phase, so later lazy imports rebuild a
self-consistent module graph from the updated checkout. The updater's own
executing modules are exempt (purging them buys nothing; reload-in-place
is the unsafe op, and we never reload). Root-segment check spares
prefix-lookalike packages. Best-effort, never raises.
5 new tests incl. an end-to-end repro of the field failure shape
(stale module missing symbol -> ImportError -> purge -> import resolves).
The desktop updater ran `hermes update --yes`, which auto-restored any
uncommitted source-tree edits onto the freshly updated checkout. On dirty
from-source installs this silently carried local modifications across every
update and could break the rebuilt app (field report: Windows update handoff
leaving the app 'crashed').
New `hermes update --keep-stash`: local changes are still autostashed so the
update can proceed, but are never re-applied — they stay parked in git stash
with printed recovery guidance. Both desktop handoff scripts (windows.ps1,
posix.sh) now pass it, probing `update --help` first so older installed
backends without the flag keep working. Failure paths are unchanged (stash
preserved, no restore); updates.non_interactive_local_changes: discard still
wins.
Tests: park/restore/failure-path coverage incl. a sabotage-verified
regression test; docs updated.
Every long-lived Hermes process is now positively identifiable so reapers
never have to guess lineage from PPID archaeology or cmdline shape:
- hermes_cli/process_identity.py (new): HERMES_SPAWN tag build/parse,
spawn-ledger.json self-registration keyed on (pid, create_time) — PID
reuse cannot forge the pair — with #89298-style corrupt-file quarantine,
and a kill-on-close job-object self-attach (BREAKAWAY_OK preserved for
the existing CREATE_BREAKAWAY_FROM_JOB escape hatches).
- serve/dashboard (web_server.py) and the gateway entry point register
themselves at startup and attach to the job; Desktop legacy
HERMES_PARENT_PID/winms marker reused as spawner identity so lineage
works with every Desktop version.
- Desktop stamps HERMES_SPAWN on backend spawns (parent-process-identity.ts).
- hermes update gets a positive-identity rung ahead of the heuristic ones:
_ledger_reapable_backend_pids reaps holders the ledger PROVES are orphaned
backends (purpose reapable + recorded spawner provably dead) in ANY update
context. Ledger-unknown holders fall through to the existing rungs.
22 new tests, sabotage-verified.