The #70337/#87331 win-unpacked wipe half, from PR #70477 by @JonthanaHanh
(reimplemented against the two-phase staged swap that postdates that
branch — the live release/ dir is grafted into the staged apps copy
BEFORE the atomic commit, so preservation rides the same rollback
machinery instead of a post-hoc copy).
Co-authored-by: JonthanaHanh <92574114+JonthanaHanh@users.noreply.github.com>
Surgical reapply of PR #87878 (@kshitijk4poor's salvage of #87327 by
@liruixinch) onto current main — the receipt-boundary and summary
changes from this session made the original commits conflict.
- ZIP fallback now keys on git ACTUALLY having failed
(_should_zip_fallback_on_update_error): a dependency-install failure
after a successful pull can't be fixed by re-downloading source and
would clobber the tree (#87331 cascade trigger, #87304).
- _abort_zip_update_if_dirty_tree: refuse to overlay a dirty checkout
(-uall so user gitconfig can't blind the guard) + pre-swap TOCTOU
re-check with our own staging artifacts filtered (#91962, #87304).
- Failure-stage naming (_format_update_failure_stage) + stderr tail so
'Git update failed' stops mislabeling pip/uv failures.
- Receipt finalize preserved on the no-fallback failure path.
Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
Co-authored-by: liruixinch <liruixinch@outlook.com>
Review on #91869 (@andrexibiza): the handwritten value_flags subset
misparsed '--reasoning high serve' as subcommand 'high' and
'-m dashboard serve' as 'dashboard' — recreating the wrong-hint class.
_holder_value_flags() now introspects build_top_level_parser() (every
option with nargs != 0, plus the pre-argparse profile selectors), with
a static fallback for broken-tree updates, --flag=value handled.
Regressions for --reasoning/-m/-t/--model=/-c per review.
De-flake test_goal_resume_restart: the fixture only set the HERMES_HOME
env var, but get_hermes_home() prefers the context-local override — an
override leaked by any earlier test in the xdist worker pointed the
goals DB at a dead tmp dir and resume enqueued nothing (the CI-only
red). Fixture now pins the override via set/reset_hermes_home_override.
Mechanism proven both ways: env-only fixture cannot beat a leaked
override; pinned fixture immune.
#90778: _hermes_holder_subcommand() — token-based parse of the actual
Hermes subcommand (profile selectors skipped, flags never matched), so
'hermes dashboard' stops being labeled as the Desktop backend and
'--preserve-cache' stops matching 'serve'. Unknown argv gets no hint
instead of a wrong one.
#87594: ancestor-exclusion in _detect_venv_python_processes and
_venv_launcher_ancestors now carves out GATEWAY ancestors (canonical
looks_like_gateway_command_line): when /update runs as the gateway's
child, the gateway stays visible to the scan so the pause machinery can
stop it, while shells/terminals/own-venv ancestry stay excluded.
15 cross-platform classifier tests; live Windows E2E suite is the
acceptance gate on this branch.
The in-function import made _time local to all of _cmd_update_impl, so
the orphan-backend reap path (which runs earlier in the function) hit
UnboundLocalError before the import line executed. The module-level
'import time as _time' at the top of update_cmd.py already covers the
divergence-merge safety tag.
Review feedback on #89507: in-place merging suits a branch that tracks the
target with a small patch set, but a long-lived feature branch (a PR branch
hundreds of commits deep) does not want an update-driven merge commit
written into its history. Reported against a checkout carrying 819 unmerged
commits.
--switch-branch routes the unmerged case to the switch path instead: the
checkout moves to the update target and updates there, and the branch is
left byte-identical — no merge, no commit, nothing written to it. The tree
is known clean on that path (the guard checks dirty before cherry), so a
dirty tree still gets the loud skip, unchanged.
Opt-in: without the flag the default remains the in-place update, which is
what keeps a small-patch-set branch's running code current.
Tests: the flag switches and leaves the branch tip byte-identical; the
default without it still updates in place. The first fails if the flag's
branch is severed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The parked-branch guard (8ce8ffd429) distinguishes checkouts by what the
branch carries, then treats both non-clean cases the same: a stale
fully-merged leftover is switched back to the target (correct), but a
branch with unmerged commits — a branch someone is actually working on —
gets CODE UPDATE SKIPPED and exit 1. For anyone running a maintained
custom branch on top of main, every update now refuses, and the guidance
('checkout main') abandons their branch.
The guard's own reason codes already separate the cases, so use them:
- fully merged -> switch back to the target (unchanged)
- unmerged:N -> update the branch IN PLACE: fetch, then bring
origin/<target> into the checkout. Fast-forward
when possible; on divergence, a true merge behind
a pre-update safety tag, stopping cleanly on
conflict. The checkout never moves; local commits
survive; the running code advances.
- dirty/unverifiable/opted out -> skip loudly (unchanged)
The post-pull success gate learns that an in-place update legitimately
ends on a non-target branch: origin/<target> was merged INTO the checkout,
so refusing to claim success there would fail every update that did
exactly the right thing.
Guard tests updated: the unmerged case now asserts the in-place outcome —
target code arrives (b.txt from c3), the branch's own commit survives, and
HEAD never moves. 18/18 guard tests, 20/20 with the diverged-update suite.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A clean checkout parked on a feature branch now always switches to the
update target. Unmerged commits are safe on the branch (git checkout
never discards committed work) and get a loud 'kept' notice naming the
branch, count, and the checkout command to resume the work. Previously
the update hard-skipped with exit 1 — a dead end for the desktop update
button, gateway /update, and cron, which have no way to resolve a skip.
Dirty trees (uncommitted changes) still skip loudly, and the
updates.auto_switch_parked_branch: false opt-out still pins the branch.
The code swap and gateway fleet restart touch all profiles, but the
pre-update quick snapshot photographed only the invoking profile's home
— siblings had no snapshot for the post-update safety nets or manual
restore to draw on.
- backup.py: create_pre_update_snapshots_all_profiles() — the SAME
snapshot set, per-file 1GiB cap, and keep policy as the invoking
profile (no partial tier, no new restore-coherence class), each into
the sibling's own state-snapshots/; restore_cron_jobs_all_profiles()
runs the #34600 cron-loss safety net per profile against its OWN
snapshot (same-generation by construction).
- update_cmd.py: sibling snapshots taken right after the invoking
profile's (best-effort, receipt-recorded); post-update cron restore
extended to every sibling.
- Docs: updating.md pre-update snapshot step now states the per-profile
behavior and the file-loss-recovery vs rollback contract.
- 9 unit tests + E2E (real files: sibling snapshot on disk, clobbered
jobs.json restored 7/7 from the sibling's own snapshot, keep=1 prune).
Phase 2 core slice of #91277: the updater now knows WHAT it is operating
on before it mutates anything.
- hermes_cli/update_inventory.py (new): side-effect-free runtime
inventory — install kind via detect_install_method (git / docker / nix
/ apt, updatable-in-place or not, with the correct external update
command for image/package-managed installs), all profiles, every live
gateway with its supervisor (systemd / launchd / manual via the
fleet-wide _get_service_pids), running code_sha/code_version from the
#91283 gateway_state.json stamps, and the restart mechanism each
runtime will get.
- hermes update --plan: prints the plan and exits; runs BEFORE the
docker/nix refusal gates so image-managed installs get a useful
'not updatable in place + right command' report instead of a bare
refusal. Read-only, safe on a live fleet.
- Every real update run now records the pre-update plan in its receipt
('plan' key) and prints a one-line fleet summary, so post-mortems can
compare what the update SAW against what it did.
- Docs: updating.md (--plan section + receipts/fleet-check section),
cli-commands.md (flag row + receipts behavior bullet).
- 11 tests: two-profile fleet classification, docker not-in-place,
dead-PID exclusion, PID-file fallback dedupe, all-probes-fail
never-raises, JSON round-trip for the receipt, print output shapes,
receipt integration.
The macOS branch of the update's fleet-restart step only restarted the
invoking profile's LaunchAgent. Sibling ai.hermes.gateway-<profile>
services kept pre-update modules cached in sys.modules and died on their
next agent turn (ImportError on new lazy imports, or TypeError/
AttributeError with garbled tracebacks on wider version gaps). The
systemd branch already iterates every hermes-gateway* unit; this brings
launchd to parity:
- _restart_macos_launchd_gateways(): the invoking profile keeps the
existing launchd_restart() path; every other gateway of this install
is drained via SIGUSR1 (same as systemd siblings), then hard-
kickstarted unless KeepAlive already respawned it, then verified on a
fresh PID. TimeoutExpired is isolated per label (#68523 parity) and
counts toward failed_or_stale_units — including timeouts during
liveness discovery, which must not read as "unloaded".
- Install-scoped fleet enumeration: launchd_gateway_labels_for_install()
derives labels from THIS install's profiles (get_default_hermes_root),
not by globbing the shared per-user ~/Library/LaunchAgents — a
sandboxed HERMES_HOME (tests, capture sandboxes, side-by-side
installs) must never enumerate, let alone restart, another install's
fleet. This also keeps the hermetic test suite blind to a dev
machine's real gateways.
- Domain-explicit sibling handling via _locate_launchd_gateway_service():
liveness, kickstart, and fresh-PID verification all use the domain the
service was actually located in (gui/<uid> vs user/<uid> probed per
label via `launchctl print`). This addresses the #41403 review defect:
the process-wide _launchd_domain() cache resolves the current profile's
domain and must never be reused for a sibling. _launchd_domain() itself
becomes a thin caching wrapper; behavior unchanged.
- _get_service_pids(all_profiles=...): the update path's manual-process
sweep excludes every gateway service PID (mirror of the systemd
hermes-gateway* pattern) so it cannot mistake a freshly respawned
sibling service for a stale manual gateway. Default-scope callers
(gateway status, cron checks, stop_profile_gateway's orphan reaper —
which kills what it is fed) keep the current-profile-only contract.
- _warn_incomplete_gateway_fleet_restart() prints launchctl recovery
hints for launchd labels alongside the systemctl ones.
Supersedes and completes #41403, addressing its review feedback
(per-label domain resolution + mocked regression tests).
Co-authored-by: David Neyra <vyr.agent@vyrgs.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Replace ps -A eww with ps -Aww: the BSD e flag is illegal
on macOS/BSD ps, making the fallback silently return [] on every macOS
machine. The matcher only needs argv (not env vars), so e is
unnecessary. -ww keeps unlimited-width output on both BSD and
procps ps.
- Add all_profiles parameter to _get_service_pids(). When True
on macOS, enumerate every ai.hermes.gateway* launchd agent across
profiles via bare launchctl list instead of only the current
profile's label. This prevents the update sweep from misclassifying
sibling-profile launchd gateways as manual processes (#73626).
- Thread all_profiles through find_gateway_pids() to
_get_service_pids().
- Update two _get_service_pids() call sites in update_cmd.py to
pass all_profiles=True so the update fleet sweep excludes every
service-managed gateway across all profiles.
- Add TestPsFallbackBsdCompat: verifies ps argv uses -Aww
not -A eww, and that pid=,command= output columns are present.
- Add TestGetServicePidsAllProfiles: verifies default scope uses
launchctl list <label>, all_profiles uses bare launchctl list
with prefix filtering, handles empty/broken output gracefully, and
preserves systemd behavior.
Tranquil-Flow
get_code_identity() shelled 'git rev-parse HEAD', which broke two tightly
mocked test suites (sequenced subprocess.run side effects in the
head-moved gate, call-count asserts in the Windows taskkill test) and
added process-spawn cost to gateway runtime-status writes.
_resolve_git_head_sha() now reads HEAD/refs/packed-refs directly,
handling regular checkouts and worktree/submodule pointer files.
Also: skip the 2s fleet settle wait when the restart phase touched no
gateways, and hoist killed_pids init outside the restart try-block.
Phase 1 of the fleet-update reliability plan (#91277): the updater now
proves its outcome instead of assuming it.
- hermes_cli/build_info.py: get_code_identity() — process-cached code
identity (git sha for source installs, baked .hermes_build_sha for
Docker images, pyproject version).
- gateway/status.py: every runtime-status write stamps the writer's
code_sha/code_version into gateway_state.json, so a running gateway's
actual code generation is observable from disk.
- hermes_cli/update_receipt.py (new): machine-readable receipt of each
update run (steps, skips with reasons, gateway restart outcome, fleet
snapshot) under ~/.hermes/logs/update_receipts/ with a latest.json
pointer for the dashboard/desktop; plus collect_fleet_versions() /
print_fleet_version_matrix() comparing every live profile gateway
against the freshly updated checkout.
- hermes_cli/update_cmd.py: wires receipt begin/steps/finalize into the
git, ZIP, and hard-failure paths; after the restart phase, prints the
fleet version matrix and escalates provably-stale gateways into the
existing gateway_fleet_restart_incomplete exit-1 contract. Pre-stamp
gateways report 'unknown' and never fail the update (no false
positives during rollout).
Silent-failure classes made visible: #88848, #74973, #85753, #81193.
Mixed-version fleet classes made loud: #88654, #69754, #77553, #56717.
`uv pip install -e .` never audits an editable target. It reinstalls on every
invocation and rewrites the console-script shims each time, which is the only
reason `hermes update` has to quarantine the running `hermes.exe` on Windows —
and a quarantine that loses its race is the whole `os error 32` family.
Gate the reinstall on whether the pull actually touched a file that defines the
install. It's safe to skip because the editable finder is pinned to a static
module list (`py-modules` + `packages.find.include`), so the one source-only
change that could stale it — a new top-level module or package — cannot land
without a `pyproject.toml` diff. Dependencies and `[project.scripts]` live
there too, and new submodules inside an already-mapped package resolve through
the real directory.
The predicate fails closed: no pre-pull SHA, an unresolvable one, or a failed
`git diff` all reinstall as before. On the skip path the two verifiers that
normally run inside the install run directly, so a wrong skip self-heals into a
real install rather than leaving an unchecked venv.
This is the pattern the file already uses everywhere else — `_tui_need_npm_install`
diffs node_modules against package-lock.json, and the desktop build is gated on a
content hash so `hermes update` "will skip if nothing actually changed". The
Python editable install was the one path with no such gate.
`hermes update` runs in the pre-pull interpreter. The auto-restart phase
imports freshly-pulled gateway source, which resolves sibling imports
against the OLD sys.modules cache — so any update where an already-cached
module gained a new export ImportErrored the whole phase and left the
gateway serving pre-update code (2026-08-20 field failure: new gateway.py
needs cli_output.line_input, cached cli_output predates it).
Class fix replacing the per-symptom _UPDATE_RUNTIME_RELOAD_MODULES
approach: _purge_stale_hermes_modules() evicts every cached module under
the Hermes package prefixes (hermes_cli/gateway/tools/tui_gateway/agent)
right before the restart phase, so later lazy imports rebuild a
self-consistent module graph from the updated checkout. The updater's own
executing modules are exempt (purging them buys nothing; reload-in-place
is the unsafe op, and we never reload). Root-segment check spares
prefix-lookalike packages. Best-effort, never raises.
5 new tests incl. an end-to-end repro of the field failure shape
(stale module missing symbol -> ImportError -> purge -> import resolves).
The desktop updater ran `hermes update --yes`, which auto-restored any
uncommitted source-tree edits onto the freshly updated checkout. On dirty
from-source installs this silently carried local modifications across every
update and could break the rebuilt app (field report: Windows update handoff
leaving the app 'crashed').
New `hermes update --keep-stash`: local changes are still autostashed so the
update can proceed, but are never re-applied — they stay parked in git stash
with printed recovery guidance. Both desktop handoff scripts (windows.ps1,
posix.sh) now pass it, probing `update --help` first so older installed
backends without the flag keep working. Failure paths are unchanged (stash
preserved, no restore); updates.non_interactive_local_changes: discard still
wins.
Tests: park/restore/failure-path coverage incl. a sabotage-verified
regression test; docs updated.
Every long-lived Hermes process is now positively identifiable so reapers
never have to guess lineage from PPID archaeology or cmdline shape:
- hermes_cli/process_identity.py (new): HERMES_SPAWN tag build/parse,
spawn-ledger.json self-registration keyed on (pid, create_time) — PID
reuse cannot forge the pair — with #89298-style corrupt-file quarantine,
and a kill-on-close job-object self-attach (BREAKAWAY_OK preserved for
the existing CREATE_BREAKAWAY_FROM_JOB escape hatches).
- serve/dashboard (web_server.py) and the gateway entry point register
themselves at startup and attach to the job; Desktop legacy
HERMES_PARENT_PID/winms marker reused as spawner identity so lineage
works with every Desktop version.
- Desktop stamps HERMES_SPAWN on backend spawns (parent-process-identity.ts).
- hermes update gets a positive-identity rung ahead of the heuristic ones:
_ledger_reapable_backend_pids reaps holders the ledger PROVES are orphaned
backends (purpose reapable + recorded spawner provably dead) in ANY update
context. Ledger-unknown holders fall through to the existing rungs.
22 new tests, sabotage-verified.
The no-live-shim probe defaulted to True (proceed with the reap) when
_venv_scripts_dir() returned None or the concurrent-instance detection
raised. Flip the default and the except-arm to False so an unverifiable
shim state keeps the updater refusing, matching the fail-closed contract
stated in the PR. Also fix the docstring: the hand-off gate lives in the
caller, not a function parameter.
Field incident (2026-08-20): a Windows Desktop update hand-off
(update --yes --gateway --force) left a swarm of per-profile serve
backends (mr-tester, probe-inherit, turqoise, clippy, maroon, …) holding
cryptography/_rust.pyd. Some still had a live parent (the tearing-down
Electron process, or the venv launcher->worker two-hop chain mid-exit),
so the strict orphan-only reap (_orphaned_desktop_backend_pids, which
bails the instant ANY holder has a live parent) disqualified the whole
set and the venv-holder guard dead-ended. The user saw a ~12-minute hang,
force-closed, and the half-done state stranded bot sessions.
New rung: _handoff_reapable_backend_pids reaps surviving Hermes
serve/dashboard backends from this venv — live parent or not — but ONLY
in the hand-off context the caller gates on: args.gateway AND the
update-incomplete marker present AND no live hermes.exe shim. In that
window nothing legitimate supervises or respawns a serve backend (the
Desktop tree-kills its backends and parks any relaunch behind the marker,
#50238), so a surviving backend is a leak, not a race. A non-backend
holder (operator REPL, stray script) still disqualifies the whole set;
psutil-unavailable returns None (keep refusing). Wired as the final rung
before the existing dead-end, after the orphan-only reap.
`hermes update` on Windows detached on every run, including the
`Already up to date!` no-op that never touches the venv. emozilla hit
the visible half: the shim exits, PowerShell takes the console back,
and a child prints the result under a fresh prompt — it reads as a
frozen update. The invisible half is worse: the hand-off sat ahead of
the fetch, so it also carried off the stash and branch-switch
questions, which #90205 then had to answer by closing stdin. Nobody
who mods Hermes got asked about their local changes again.
The shim lock is real and the child is still required — a launcher
holds venv\Scripts\hermes.exe open without FILE_SHARE_DELETE for the
whole command, so the quarantine rename is refused and uv fails with
os error 32. A parent that waits deadlocks against the handle it is
itself holding, and Windows has no exec to escape with.
But that lock only binds one step. Move the hand-off to the dependency
sync boundary, beside the native-module deferral that solves the same
"this process holds a file the sync must replace" problem — and for
the reason that placement already exists (#86735: a preflight ahead of
the fetch re-bricked the flow it was meant to protect). Everything
before the sync now runs foreground in the user's console: the
preflight, the stash question, the branch switch, git pull. An
up-to-date run never hands off at all.
Deferring to the next launch cannot substitute here the way it does
for a mapped .pyd: every future `hermes` launch is also the shim, so
the marker would defer forever. The child re-runs the update to keep
the node/web/lazy-refresh tail, and takes the sync it was spawned for
rather than the up-to-date early return.
A GitHub-side HTTP 429 during 'hermes update' printed only
'Failed to fetch updates from origin.' — and the curl
'unable to access ... returned error: 429' shape even matched the
network-error branch, blaming the user's connection for a GitHub
outage.
- new _classify_fetch_failure(): 429/rate-limit -> 'GitHub is rate
limiting requests or having an outage — try again in 5 minutes';
5xx -> outage message with githubstatus.com; ordered BEFORE the
generic 'unable to access' network check
- both fetch-failure sites (update apply + --check) now share the
classifier via _print_fetch_failure(), and both always print the
first raw stderr line so the wire error stays diagnosable
- tests: classifier matrix + E2E against a live local HTTP server
returning 429 through real git
Fixes#89287
hermes update treated a failed desktop pack as non-fatal and still printed
✓ Update complete!, so Windows users kept running an old Hermes.exe after a
"successful" update. Withhold the success banner, surface the stale app in
the summary, and write .update_exit_code=1 for gateway watchers.
Supersedes #88359, #87984.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Hermes is an agent for one person. The credentials, the memory, the
sessions and the cron jobs all belong to that person. But the only
declarative path was a NixOS system service. Issue #9056 asks for the
user-level equivalent. 25 public Nix configurations already write one by
hand, and several of them copy nix/nixosModules.nix and edit the systemd
part.
This module is not a second copy of that file. The code that both modules
share moves into nix/moduleCommon.nix:
- the options
- the renderers for config.yaml, .env and the documents
- the activation body
- the command lines of the processes
nixosModules.nix keeps only the parts that need root. Those parts are the
service user, stateDir, addToSystemPackages, container mode and tmpfiles.
The file goes from 1008 lines to 666.
`services.hermes-agent` is now the same option set on both modules. A
NixOS example works on Home Manager without a change, and an option added
one time appears on both.
The Home Manager module is different only where it must be. It uses
systemd.user.services on Linux and launchd.agents on Darwin. It uses
home.activation and not system.activationScripts. It sets HERMES_HOME
directly, with the default ~/.hermes, so an existing directory continues
to work. It uses the modes 0600 and 0700, because the state has one user
and does not need the group-shared umask of the NixOS module. It does not
support container mode, which needs root and the Docker socket.
The change also makes four corrections that apply to both modules:
- backend.mode runs `hermes serve` or `hermes dashboard`. Both modules
had only the gateway. But Hermes Desktop and the web dashboard connect
to a different process, so six of the configurations in public repos
add a second unit by hand. serve and dashboard are one entry point with
one flag of difference, and you can run only one of them. Thus the
option is an enum. The NixOS module asserts against container mode with
a backend, and does not make a unit that cannot start.
- hermesHomeFiles installs files into HERMES_HOME. The `documents` option
installs into the working directory, which is correct for AGENTS.md but
wrong for SOUL.md and memories/. Hermes reads those files from
HERMES_HOME, in agent/prompt_builder.py:2095. A SOUL.md in `documents`
made a workspace file that Hermes never loaded as the identity. The
documentation said this in prose, but two directory diagrams showed the
opposite. This change corrects both. A key in either option can now
contain subdirectories.
- `documents` needs an explicit `workingDirectory`. The default of that
option is bad on both modules. It is the home directory of the user on
Home Manager, and ${stateDir}/workspace on NixOS. A user who declares
workspace files without a directory therefore gets a place that the
user did not select. The place is also different on each module. The
modules now refuse that combination.
The test is on the priority of the option and not on its value. An
option that nothing sets keeps the priority of its own default, and
each definition from a user is stronger. Thus a directory with the same
text as the default still counts as a selection, and so does a
mkDefault. A comparison of values detects neither case.
- Each activation writes .env again from a base in the Nix store, and
does not add to the file that exists. Thus a second activation cannot
put the same secret in the file two times, and a removed
environmentFile goes away. environmentFiles keeps the type `listOf
str` and not `path`, so Nix cannot copy a sops-nix or agenix path into
the Nix store, which all users can read.
- HERMES_MANAGED and the .managed marker now hold the name of the system
that manages the install. Thus a refusal says "managed by home-manager"
and not "managed by NixOS", and `hermes update` gives the Nix guidance
for both shapes. The CLI does not print a rebuild command for each
system. It names the owner, and the user knows their own tool. A bare
`true` and an empty marker still mean NixOS, so this does not change an
existing install.
Verification. Six new checks, all built:
nixos-module evaluates the module with evalModules and the
NixOS module list. It asserts both units, one
HERMES_HOME, and that the module refuses
container mode with a backend.
home-manager-module evaluates the module with the
homeManagerConfiguration function of
home-manager. The process assertions run against
systemd units on Linux and launchd agents on
Darwin.
module-option-parity asserts that each shared option is on both
modules, and that the two exclusion lists name
only options that exist.
env-file-assembly runs the real .env script and checks the
contents, the mode, that a second run gives the
same bytes, and that a removed file goes away.
workspace-files-need-a-directory
checks that the module refuses `documents`
without a directory, and accepts a directory
that has the same text as the default.
service-argv runs each command line that the modules build
through the real parser of the CLI, with one
sentinel flag added, and requires that argparse
refuses only the sentinel.
`nix flake check` passes, with 21 checks in total.
The CLI branches that treat an install as a Nix install move to one
helper, is_nix_install_method. Four call sites in main.py, web_server.py,
update_cmd.py and doctor.py tested the literal set {"nix", "nixos"}, and
each one missed home-manager. recommended_update_command asks the managed
state before the code-scoped stamp again, because a managed install can
carry a stale stamp that names an update path the managed guard refuses.
The metrics contract gets a home-manager bucket, so a Home Manager
install does not report as unknown.
Each check was mutation-probed. 22 faults were injected, and the checks
caught all 22:
- a lost --no-open
- a backend that runs the gateway
- an overwritten config.yaml
- documents in the wrong directory
- a different HERMES_HOME on the two processes
- a lost HERMES_HOME export
- a missing backend unit
- a removed assertion
- an .env file that grows at each activation
- an install that reports NixOS
- an empty .managed marker
- an option on the NixOS module only
- a stale entry in an exclusion list
- a renamed subcommand
- an unknown flag
- the workspace-files assertion always passes
- the assertion compares values instead of priorities
- an off-by-one that lets an untouched default through
- the assertion also fires for hermesHomeFiles
- a mkDefault no longer counts as a selection
- the Home Manager module stops wiring the assertion
- the NixOS module stops wiring the assertion
The 16 Python tests in tests/hermes_cli/test_managed_install_shapes.py
were probed the same way. 8 faults were injected and 8 were caught.
These tests fail on this tree. They fail in the same way on the stashed
HEAD, and they have no relation to Nix:
- test_git_probe_tree_kill.py (2 tests)
- test_update_import_guard.py (1 test)
- test_telegram_media_read_timeout.py (2 tests)
- test_teams.py (a collection error)
Closes#9056
# Conflicts:
# hermes_cli/main.py
# hermes_cli/update_cmd.py
# hermes_cli/web_server.py
Live incident 2026-08-17: the source checkout was parked on a stale feature
branch (claude-code-inspired/local-terminal-memory-limit, days behind main),
left there by earlier tooling. 'hermes update' autostashed, refreshed lazy
backends, synced skills, and printed '✓ Code updated!' / '✓ Update complete!'
while the checkout stayed on the stale branch with none of main's new code.
Two sessions burned time on 'the fix is missing' confusion.
- Parked-branch guard: auto-switch back to the update target ONLY when the
parked branch is clean and fully merged (git cherry origin/<target> shows
nothing unmerged); the checkout then STAYS on the target instead of being
re-parked. Otherwise: loud CODE UPDATE SKIPPED block naming the branch,
behind-count, and resolution commands; exit 1; branch untouched.
- The up-to-date (commit_count == 0) path no longer switches back to a
fully-merged parked branch either.
- Post-pull gate additionally refuses to print '✓ Code updated!' when HEAD
ends up attached to a non-target branch.
- Summary lines now carry the actual branch + HEAD short-sha:
'✓ Update complete! [main @ 30fcf9580]' — drift visible at a glance.
- New config toggle updates.auto_switch_parked_branch (default true).
- Real-git-fixture regression tests (init/clone/branch, no subprocess
mocks): clean+merged auto-switch, dirty skip, unmerged skip, cherry-picked
equivalence, config opt-out, unverifiable ref, on-main fast path,
up-to-date no-repark, summary branch/sha assertions.
'✓ Update complete!' now shows what the update actually delivered:
'✓ Update complete! (v0.19.4 → v0.20.0)' when the pyproject version
changed, '(v0.20.0)' when commits landed within one release, and the
plain message when the version cannot be read. Reads the on-disk
pyproject.toml (not importlib.metadata, which still describes the old
install after a pull). Applied to both the git and Windows-ZIP paths.
Mirror the strict unit-name shape from the hermes-serve gate (review on
PR #83595) on the gateway side too: the discovery gate and the SIGUSR1
eligibility helper now accept only `hermes-gateway.service` or the
`hermes-gateway-<profile>` family, so a near-prefix unit like
`hermes-gatewayd.service` can neither enter the restart path nor be sent
a SIGUSR1 it does not handle.
Review on #83595 flagged two service-lifecycle gaps in the hermes-serve
restart support:
- The unit-name gate accepted anything starting with "hermes-serve",
which also matched the unrelated hermes-server.service. Require the
exact base unit or the hyphenated profile family instead.
- The fleet-restart loop and _finish_dashboard_update_cleanup() could
both restart the same hermes-serve unit — the loop restarts it
directly, then cleanup's PID scan finds the fresh process and
restarts its owning unit again. Thread the fleet loop's restarted
unit names through to _kill_stale_dashboard_processes() so it skips
units already handled.
hermes update discovered and restarted hermes-gateway* systemd units but
never looked for hermes-serve* — the Desktop app's backend — so it kept
running stale pre-update code until the user restarted it by hand (#83438).
Extend the systemd unit discovery/restart loop to also match hermes-serve*
units. They don't wire SIGUSR1 to a graceful drain (only gateway/run.py
does), so restart eligibility for the graceful path is now gated on unit
name via a small, directly-tested helper; hermes-serve units fall straight
to the existing blunt systemctl restart path, matching the workaround the
issue already documents.
Widen PR #87757 to cover the ZIP path: _update_via_zip() also calls
_finish_dashboard_update_cleanup() but never runs _reload_config_modules,
so the Windows git-broken fallback would still crash with the same
ImportError (cannot import name 'bounded_probe_run' from the stale cached
hermes_cli._subprocess_compat).
- new _reload_process_scan_modules() called inside
_finish_dashboard_update_cleanup itself, so every current and future
call site is covered; reloads dependency-first
(_subprocess_compat, then dashboard_procs)
- reload failures log at warning (a miss surfaces seconds later as an
ImportError in the same process)
- regression tests: reload-before-kill ordering, node-failure skip,
stale-module symbol restoration (the exact #87134 boundary state),
nonfatal reload failure, and the #87757 reload-list contract
hermes update runs in the PRE-pull Python process. After git pull updates
source files on disk, modules already in sys.modules still hold the OLD
code. The existing _reload_config_modules() reloaded only config modules,
but the post-update dashboard cleanup path (_finish_dashboard_update_cleanup
-> _scan_dashboard_processes) imports hermes_cli._subprocess_compat lazily;
a new symbol added there (e.g. bounded_probe_run) is invisible to the
cached module object, causing ImportError during the cleanup step.
Extend the reload list to include hermes_cli._subprocess_compat and
hermes_cli.dashboard_procs so the cleanup uses freshly-pulled code.
Two related failure modes after a crashed/interrupted fetch on a shallow
clone (git clone --depth 1 installs):
1. STALE LOCK WEDGES EVERY FETCH. A killed fetch can leave .git/shallow.lock
behind; every later 'git fetch' then fails with 'Unable to create
.../shallow.lock: File exists'. 'hermes update --check' reported a hard
fetch failure, and the passive banner check swallowed the exception and
compared stale refs. Add hermes_cli.gitlock.clear_stale_git_locks(), a
guarded sweep (age + git-process check so a live fetch is never yanked)
wired into the check path, the apply path, and the banner's passive check.
2. SHALLOW TIP-SHA COMPARE FALSE-POSITIVES. On a shallow clone the check
cannot count commits, so it compares tip SHAs. Local cherry-picks on top
of the remote tip (e.g. re-applied local patches) make HEAD differ from
origin/main even though HEAD already contains it — a false 'update
available' banner. Add hermes_cli.gitlock.is_ancestor_of_head() and use
'git merge-base --is-ancestor' in the CLI check and banner paths before
reporting an update. Mirror in the desktop (update-count.ts gains an
isAncestor input; main.ts probes merge-base --is-ancestor).
Tests: tests/test_gitlock.py (9) covering stale/young/no-lock/no-repo sweeps
and ancestry true/false; update-count.test.ts +3 for the isAncestor path.
A gateway whose asyncio event loop is stalled (e.g. an in-loop
compression pass, #72707) cannot process SIGTERM/SIGUSR1 shutdown.
The updater's drain wait then burned the full 180s budget, warned
"Gateway PID X still running after 180.0s — restart may fail", and
`hermes update` could deadlock behind the wedged process — the user
cannot update their way out of the stall.
Fix: before any drain wait, read the loop-liveness heartbeat file the
gateway rewrites every 30s (#66892). Classification:
- alive (fresh heartbeat): busy-but-alive loop — take the normal
graceful drain, honoring the in-flight cron drain floor (#86684).
- wedged (heartbeat for this PID stale >90s = 3 missed beats): the
loop is provably dead; drain is pointless. Bounded escalation:
SIGTERM + 5s grace, then SIGKILL + 5s wait, then proceed (~10s
worst case, far under the 180s drain budget).
- unknown (missing/corrupt file, PID mismatch): never escalate on
ambiguity — full drain path.
Wired into launchd_restart, systemd_restart, and both updater
gateway-shutdown sites (systemd unit drain + manual profile
gateways). The probe is a local stat + JSON read (well inside the
10s query tier of the subprocess timeout tiering).
The cron drain floor from #86684 is bypassed ONLY when the loop is
provably dead — a merely busy gateway still refreshes its heartbeat
and keeps the full drain budget.
Root cause of the loop stall itself (compression blocking the loop)
is #72707 territory and deliberately out of scope here.
Fixes#81642
The #86687 self-lock preflight fired on every Windows `hermes update`:
bitwarden.py's module-level cryptography import (fixed in #86782 /
#86826-class change) meant cryptography._rust was ALWAYS mapped by the
time the preflight ran, so the update exited 2 before even fetching and
looped forever — including the Desktop in-app update (#86780).
Two structural fixes so the guard can never re-brick the flow it protects:
1. Version-gated detection: _detect_self_loaded_native_modules() now
consults _dependency_sync_would_rewrite(dist) — installed version vs
the on-disk pyproject pins (base deps + all extras, env markers
honored). A loaded module whose distribution the sync will not touch
is no lock risk and is not reported. Unknown → fail closed.
2. Relocated deferral: the check no longer runs pre-fetch. It runs via
_abort_dependency_sync_if_self_locked() immediately before each venv
rewrite (git-path dep sync, ZIP-path dep sync, current-checkout venv
repair) — AFTER the code swap. A deferral now leaves the user on NEW
code with only the dependency install pending (completed by the next
launch's marker recovery), instead of stranding them on the old
checkout in an exit-2 loop.
PyYAML's _yaml extension (loaded by every CLI process) joins the
registry — with version gating it is now safe to list.
Tests: version-gate unit coverage (no-change skip, stale pin, missing
dist, extras, markers, fail-closed None), deferral wiring (marker +
gateway resume + exit 2), placement guards (no detector call pre-fetch;
guard present at git/ZIP sync), and subprocess-verified import hygiene
(import hermes_cli.main and the update --check dispatch never load
cryptography._rust).
Follow-up to #86687 (Halldrix's #83590 salvage — the preflight's intent
stands as defence-in-depth; this makes it fire only when true).
Fixes#86735Fixes#86780Fixes#86781
Two gaps left every Windows git-checkout install unable to recover from
the exact failure state #83569 reports:
1. Self-lock detection. _detect_venv_python_processes() always excludes
the calling process by design — a CLI hermes update IS the venv python.
An updater that had already imported a native venv extension (the
canonical one being cryptography.hazmat.bindings._rust, mapped while
hermes_cli.main resolved external secret sources) passed every
preflight and then died mid-sync with os error 5 when uv tried to
rewrite the mapped .pyd, stranding the venv half-updated. A new
preflight now refuses the sync before touching the checkout, writes
the update-incomplete marker so the next fresh launch completes the
install, and exits 2. Verified on a live Windows 11 host: after
importing hermes_cli.main, tasklist /m _rust.pyd shows the .pyd mapped
in the caller, and a peer process cannot open it read-write
(Permission denied) — while a rename succeeds, matching how uv/pip
actually fail (truncate+write, not rename).
2. Early-recovery install path. _early_recovery._run_repair_install used
sys.executable -m pip unconditionally. Windows git checkouts install
on a uv-managed base interpreter (python-build-standalone), whose
EXTERNALLY-MANAGED marker makes plain pip abort with
externally-managed-environment — the repair no-oped and the venv
stayed broken. The repair now detects the PEP 668 marker, prefers
uv pip install with VIRTUAL_ENV pointed at the project venv, and
falls back to pip --break-system-packages when no uv binary exists.
Both fixes ship with subprocess/unit regressions (sabotage-verified):
the new tests fail on pre-fix code and pass with it. Complements #77517,
which keeps the updater from importing cryptography in the first place;
this PR is the defence-in-depth when any future path loads it anyway.
Fixes#83569
_cold_start_windows_gateway_after_update() printed the success line off a
successful Popen return alone, which only proves CreateProcess succeeded,
not that the child survived. On Windows, a job object denying
CREATE_BREAKAWAY_FROM_JOB hard-kills the child during updater teardown
before it logs anything, yet the updater still printed "Starting Windows
gateway after update (PID ...)" — leaving Telegram/Discord/etc. offline
with no indication anything failed (#84185).
Route the success report through gateway_windows._report_gateway_start(),
the same post-spawn liveness poll every other _spawn_detached() caller
already uses, so a dead child is reported as a failure with a
manual-recovery hint instead of a false success.
Review follow-up (#78574): the aborted-restart handler only flagged the fleet
stale when the post-failure survivor probe was None or non-empty. A positive
empty probe was treated as proof-of-safety — but `[]` is only safe when
nothing was running before the phase. If a gateway was discovered, stopped
(SIGTERM/drain), and its replacement never came back, the probe is empty at
exactly that unsafe moment and the update reported success — the fail-open
contract this fix exists to close.
Snapshot the pre-restart gateway PIDs before any stop/drain and route the
handler decision through a pure _restart_phase_failure_is_incomplete() helper
that fails closed on an empty survivor set whenever a gateway existed
pre-restart (or the pre-state could not be read). Add decision-level regression
tests covering the stopped-without-replacement gap, unknown pre-state, and the
truly-no-gateway positive control.
The gateway auto-restart phase in `hermes update` was wrapped in a blanket
`except Exception` that only logged at debug level. When the phase raised
early — e.g. importing `hermes_cli.gateway` from the freshly pulled checkout
inside a process that already loaded pre-update modules — every drain and
restart line vanished from the update output, the update printed
"Update complete!" and exited 0, and the still-running gateway kept serving
pre-update modules against replaced source files. The next Telegram turn died
with `ImportError: cannot import name 'is_trivial_prompt'`.
The handler now probes for surviving gateway processes and, unless it can
positively prove none are running, prints the cause plus a manual recovery
command and marks the fleet restart incomplete — which exits nonzero and
writes the gateway-mode exit-code marker, matching the existing
failed-or-stale-unit path.
Fixes#78574
A detached/pinned checkout can report 'N new commit(s)' against origin,
run the ff-only merge successfully, and still sit on the old commit
afterward (the branch-switch step re-detaches to the raw SHA). Before
this guard 'hermes update' printed '✓ Code updated!' and reinstalled
deps + rebuilt the desktop app against the stale tree - no error, no
warning, 'hermes doctor' healthy.
Compare pre-pull and post-pull HEAD; if they match, fail loudly with a
reattach hint instead of claiming success.
The hermes update APPLY path still ran an unconditional
rev-list --count HEAD..origin/<branch> — on a depth-1 installer checkout
that walks the truncated graph and reports the entire remote ancestry
(#53479's 'Found 9980 new commit(s)' on Windows 11). The zero/nonzero gate
stays (a 0 count is trustworthy on any graph); when the count is positive
on a shallow repo, recover the real number via the GitHub compare API
(added in PR #86257) and print count-free wording when that fails.
ahead_by==0 (local-ahead) falls through to the up-to-date path.
Completes the class fix from PR #86257 on its last remaining site.
The honesty half (no fabricated counts) leaves shallow installs permanently
count-less. The compare API knows the full graph regardless of local clone
depth: GET /repos/<o>/<r>/compare/<current>...<target> returns ahead_by —
exactly the behind count the shallow boundary lost.
- hermes_cli/banner.py: _github_compare_behind() (bounded, unauthenticated,
best-effort); wired into _check_via_rev and the shallow branch of
_check_via_local_git. ahead_by==0 with differing tips = local-ahead => 0.
- hermes_cli/update_cmd.py: hermes update --check shallow path prints the
exact count when recoverable, presence-only wording otherwise.
- apps/desktop/electron/update-count.ts: compareApiUrl() +
parseCompareBehindCount() pure helpers; main.ts fetches the count when
resolveBehindCount() returns null, and the SSH-official passive path stops
fabricating behind:1 (uses compare API + updateAvailable flag).
- apps/desktop/src/lib/version-status.ts: updateAvailable now applies to the
client target too, so a shallow desktop install shows '(update)' instead of
nothing (or the old frozen '(+1)').
Fixes#84591; CLI siblings of #78253 / #53479 behavior.
E2E: live compare API returned 61/62 for real 61/62-commit gaps and 0 for the
reversed (local-ahead) pair; real shallow-clone fixture (depth-1 clone +
depth-1 fetch, merge-base broken) recovers the exact count with the API and
falls back to the honest sentinel offline.
A failed npm install during `hermes update` prints "Fix npm and re-run
`hermes update`" -- but re-running on a current checkout hit the
"Already up to date!" early return before the Node refresh, so the
repair advice could never work and node_modules stayed stale forever
(#77211).
The commit_count == 0 path now runs the Node refresh through
_repair_node_deps_on_current_checkout. _update_node_dependencies
self-gates on the lockfile hash, which is only recorded after a
SUCCESSFUL npm install (and re-trips when node_modules is missing or
the web toolchain never landed), so healthy installs pay one hash
check and nothing else; a previously failed install actually repairs.
A clean refresh pairs with the web build like every other call site;
a failed one surfaces the fix-npm hint instead of "Already up to
date!".
Fixes#77211.
Co-authored-by: RelaxJonh <RelaxJonh@users.noreply.github.com>
Co-authored-by: JonthanaHanh <JonthanaHanh@users.noreply.github.com>
Root package.json still owns devDependencies (the shared ESLint flat
config every workspace's eslint.config.mjs imports) even though
agent-browser and @streamdown/math were already removed from root
dependencies. The scoped `npm ci --workspace ui-tui --workspace web`
prunes them the same way it used to prune those; --include-workspace-root
protects them without reintroducing apps/desktop into the install.