Commit Graph

135 Commits

Author SHA1 Message Date
Zheqing Zeng aa72df4b42 fix(macos): re-land dylib-complete TCC interpreter anchor
The first landing (#95131/#95478, reverted in #95563) copied the
uv-store interpreter into venv/bin/python so TCC grants would stick
to a stable path. On real Macs that copy bricked every hermes command
two ways: dynamically-linked builds died in dyld because
@executable_path/../lib/libpython resolved into venv/lib/ (#95425),
and alias symlinks to the copy made CPython getpath lose the venv
prefix (#95541, ModuleNotFoundError: encodings).

Re-land:

- Keep the signed real-file copy of bin/python (identifier-pinned
  via _macos_sign_managed_python).
- Materialize python3 / python3.N as real-file copies, never
  symlinks. Copies boot on every build we could reproduce and keep
  the TCC identity.
- Hardlink store libpython* into venv/lib/ when present (copy across
  devices). Existing LC_RPATH already points there.
- Pre-install boot gate: launch the staged copy, demand encodings
  plus the venv prefix, abort and leave the live venv untouched
  on failure.

Doctor reports/installs the new anchor (the revert-era heal is
removed). Update refreshes it after a successful code swap. Tests
cover layout, idempotence, predecessor-symlink repair, libpython
hardlink, boot-gate refusal, and a macos_only real-interpreter E2E.

Closes #95596.
2026-08-28 09:05:19 +05:30
kshitijk4poor f3cbb262c1 fix(update): valid --ignored=matching mode; rename-only path split; shared preserve constant
Review corrections on the first draft (caught by /simplify-code before
merge — the PR was disarmed for these):

- BLOCKER: --ignored=all is not a valid git mode (git exits 128 'Invalid
  ignored mode'); with it, every ZIP update was refused as 'could not
  check the working tree'. The mocked tests could not see this — a new
  real-git test creates an actual repo + .gitignore and asserts the guard
  runs clean, blocks on an ignored user file, and exempts ignored
  preserved entries. --ignored=matching also reports an ignored dir as
  one line instead of enumerating its contents.
- FAIL-OPEN HOLE: the ' -> ' two-path split now applies only to R/C
  rename/copy status codes. Porcelain v1 does not quote plain filenames
  with spaces, so an ignored file literally named 'venv -> node_modules'
  parsed as two preserved tops and slipped past the guard into the
  destructive swap.
- _update_via_zip's swap loop now consumes _ZIP_PRESERVED_TOP_LEVEL
  instead of a comment-synced duplicate set (change-detector test added).
2026-08-27 21:07:34 +05:30
joaomarcos e64db76982 fix(update): gitignored user files also block the ZIP overlay
Carried from #87392 (closed as superseded — its core guard landed via the
#87327 salvage chain): the dirty-tree check now passes --ignored=all, so a
gitignored-but-real user file (logs, scratch files, local data) blocks the
destructive ZIP overlay too. The ZIP path's own preserved top-level entries
(venv, node_modules, .git, .env — gitignored on every normal install) are
exempted so they don't become a false refusal.

Credit: @JoaoMarcos44, whose #87392 included this hardening.
2026-08-27 21:07:34 +05:30
Cursor Agent 8246c4f92a fix(cli): repair interrupted update fleet restart
An interrupted hermes update after git pull advanced HEAD never
restarted running gateways, and the next update said "Already up to
date" and skipped the fleet. Persist a HERMES_HOME fleet_restart_pending
marker after HEAD moves, clear it only when restart completes (or
nothing was running), and catch up on the next hermes update even when
git is current — also when latest.json records a stale runtime SHA.

Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>
2026-08-26 21:38:20 -07:00
Teknium 71823be9da fix(update): stop counting the Windows resume token as a fleet runtime
Fixes #93406 (residual). _fleet_probe_expected_runtimes counted the
_windows_gateway_resume pause/resume token (profiles/unmapped entries)
as an 'expected fleet rows' signal. The token is pause/resume
bookkeeping, not a runtime inventory, and its entries have no rows
collect_fleet_versions() can return: unmapped Scheduled-Task gateways
never publish gateway_state.json, and a resumed profile gateway
relaunches detached and may not republish within the probe window. So
every Windows update that paused a gateway set _fleet_rows_expected,
the verification loop silently waited out its polling window (~14 min
wall clock with the retry loop on user reports), printed 'Fleet version
check returned no rows', and exited 1 for an update that succeeded.

Expected-runtimes now keys only on row-capable signals: restart-phase
bookkeeping, the pre-restart PID snapshot, and the pre-update plan
inventory -- which already cover any genuinely live pre-update Windows
gateway.

Counterfactual proof: tests/hermes_cli/test_update_fleet_probe_resume_token.py
fails on the pre-fix predicate (token-only => True) and passes with the
fix; the row-capable signals are pinned unchanged.
2026-08-26 18:23:33 -07:00
fred0m 53057f2bc4 fix(update): run config migration on the 'Already up to date' repair path (#91360)
A failed update attempt can pull fresh code onto disk and then die before
the config-migration block (e.g. a PyPI timeout during the dependency
sync). The desktop hand-off retries; the retry takes the commit_count == 0
branch, repairs deps, prints 'Already up to date!' and returns early -
skipping _run_config_check_fresh / migrate_config entirely. The fresh
code (requiring a newer _config_version) then refuses to start against
the old config until 'hermes doctor --fix' is run.

Fix: _maybe_migrate_config_on_current() mirrors the version_bump_only
handling (silent, non-interactive) and is called on both repair-path
completion points before claiming success.

Also: scripts/desktop-update/posix.sh no longer retries when the update
was deliberately SKIPPED (checkout parked on a non-target branch) -, the
retry is deterministic and only wastes time. Uses a dedicated non-
colliding exit code (8) and an honest message instead of 'Update failed'.

New tests: tests/hermes_cli/test_update_config_migration_on_current.py
(5 cases: migrate-when-behind, noop-current, noop-ahead, warning re-
surface, silent check failure).
2026-08-26 17:42:50 -07:00
loulanyue bf5ff51078 fix(update): check and apply config migrations on current checkout / retry paths (#91360)
When an update was interrupted or failed mid-install (e.g. dependency install
timeout) after pulling new code, the subsequent update run takes the
'commit_count == 0' path and early-returned without checking or migrating
the configuration. Fresh code requiring a newer config version would fail to
boot on the next run.

Extract _check_and_apply_config_migration and invoke it across all update
completion paths (normal update, current checkout / node repair, and python
dependency repair).
2026-08-26 17:42:50 -07:00
Casey 790e1eb6bd fix(update): pause SCM-supervised Windows gateway services before venv mutation
On Windows installs where the gateway runs as an SCM service (WinSW,
NSSM, sc.exe create), the existing pause machinery kills the gateway
process directly — and the service wrapper's failure ladder resurrects
it within seconds, re-taking the venv file locks mid-update. The update
then dies partway through dependency sync with access-denied errors.

This extends _pause_windows_gateways_for_update() to detect when a
gateway's process tree is owned by a running SCM service, and to stop
the SERVICE through sc.exe instead of killing the child:

- gateway/status.py: expose service-ownership discovery for gateway
  runtimes (find_windows_gateway_services maps validated gateway PIDs
  through process ancestry to running SCM service PIDs, with
  create-time identity checks against PID reuse).
- hermes_cli/update_cmd.py: stop verified services via sc.exe before
  venv mutation and restart them afterward. Stops wait for a stable
  SCM 'stopped' state AND for the original descendant processes to
  exit (service 'Stopped' is not proof the child released its
  handles). Failure to prove ownership, stop a service, or restart it
  fails closed; rollback restores attempted services, and rollback
  failures are surfaced rather than swallowed.
- Fail-closed throughout: unreadable identities, ambiguous ancestry,
  or a service that will not reach a stable state abort the update
  before any file mutation.

Complements #37039 (gateway-only concurrent instances no longer abort):
that fix lets the update proceed past the gate; this one makes the
pause actually stick when the gateway is service-supervised.

Note: tests/gateway/test_status.py::TestReadProcessCmdlinePsFallback::
test_ps_fallback_when_proc_unavailable fails on Windows on current main
before this change as well (POSIX ps fallback asserted on a platform
without it); all other touched suites pass (155 passed, 5 skipped).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 16:45:31 -07:00
Teknium b2c011364e fix(update): conservative outcomes + serve-ledger coverage for fresh restart recovery
Salvage adjustments to PR #94392 per review:

- Narrow the supervisor claim to the systemd-VERIFIED path only. The fresh
  recovery child now probes 'systemctl --user is-active' after each relaunch;
  only an observed-active systemd unit is reported 'verified'. A relaunch that
  merely exited 0 is labelled 'relaunch_attempted', never counts as supervisor
  coverage, and never clears gateway_fleet_restart_incomplete.
- Serve-owned runtimes (serve/dashboard entries from the spawn ledger, per the
  update_inventory serve collector) are no longer silently skipped: the
  recovery pass records them (and manual gateways) as skipped-with-reason in
  the recovery result and the persisted update receipt.
- Receipt fresh_recovery persists the conservative vocabulary
  (requested/verified/relaunch_attempted/failed/skipped); 'succeeded' is gone.
- Added an end-to-end test that drives the real recovery module in a genuinely
  fresh interpreter (sitecustomize shim intercepts the grandchild
  'gateway restart' and systemctl probes).
2026-08-26 16:45:26 -07:00
joaomarcos ccdd7f41ed fix(update): persist per-profile recovery outcomes 2026-08-26 16:45:26 -07:00
joaomarcos f0045c5383 fix(update): verify fresh restart recovery results 2026-08-26 16:45:26 -07:00
joaomarcos 5609ccbece fix(update): recover aborted gateway restart in a fresh process 2026-08-26 16:45:26 -07:00
Teknium df3d41ee67 fix(update): sweep aborted-fetch tmp_pack debris before it corrupts the pack directory (#93732)
Every git fetch that dies mid-transfer (timeout, HTTP 429, dropped
line) strands a tmp_pack_* file in .git/objects/pack, and git never
cleans them. The banner's background update check is the main generator
on flaky lines — several aborted fetches a day — and the reporter's
install accumulated hundreds of files / 6.0 GB over 9 days until the
pack directory corrupted outright and every update check hung or
failed permanently.

clear_stale_tmp_packs() in gitlock.py sweeps tmp_pack_/tmp_idx_/
tmp_rev_/tmp_mtimes_ debris with the exact safety contract the lock
sweep already uses: only files past the 10-minute age floor, never
while any git process runs, never raises, real pack-*.pack/.idx files
untouchable by construction (prefix match). Wired into all three
fetch-adjacent sites: _cmd_update_check, the update apply path, and
the banner's passive check (generator = janitor).

Live E2E: 300 aged tmp_pack files (the reported scale-shape) swept
from a real repo; an in-flight fresh tmp and ancient real packs
survived; fsck clean and a real fetch round-trip succeeded after.
2026-08-26 16:45:22 -07:00
pierrenode de2a9de788 fix(update): feed Windows gateway relaunch outcome into fleet reconciliation
#91277 Phase 2's plan-vs-execution reconciliation (match_runtime_outcomes)
cross-checks every runtime collect_runtime_inventory() saw against
restarted_services / relaunched_profiles / externally_supervised_profiles /
killed_pids — the systemd/launchd restart phase's bookkeeping. That
inventory is cross-platform (control-socket / PID-file based), so it
includes Windows gateways too, but Windows's own pause/resume mechanism
(_pause_windows_gateways_for_update / _resume_windows_gateways_after_update)
never wrote into any of that bookkeeping.

Result: a Windows gateway that was correctly stopped and relaunched by
_resume_windows_gateways_after_update was still classified "unaccounted" by
the reconciliation (the plan saw it and no bookkeeping mentions it) —
report_unaccounted_runtimes() escalates that into sys.exit(1), and in
gateway_mode also writes ".update_exit_code"="1". Every successful
`hermes update` on Windows with a running gateway reported itself as
failed, unconditionally (the sys.exit(1) is not gated to gateway_mode).

_resume_windows_gateways_after_update now records the profiles it
successfully relaunched onto the resume token; _cmd_update_impl merges
that into the shared relaunched_profiles list right before reconciliation
runs. A profile whose relaunch genuinely fails is deliberately left off
the list, so it still surfaces as unaccounted — Windows has no watcher to
recover a failed relaunch, so that escalation is the correct signal.

Regression tests exercise _resume_windows_gateways_after_update directly
(records successes, omits failures) and reproduce the reconciliation-level
bug end to end: the same plan row resolves "unaccounted" without the merge
and "restarted" with it. Mutation-verified: with the fix reverted, three of
the four new tests fail (KeyError on the token / wrong outcome).
2026-08-26 16:14:27 -07:00
AlexGabbia b3e477f304 fix(update): wait for resumed Windows gateway before failing fleet check
The post-update fleet version check slept 2s and probed once. On Windows the
resume path relaunches the gateway detached, and it needs ~10s to boot (the
Telegram polling reconnect) before it stamps gateway_state.json or answers the
control socket. That race reported "no rows" for a healthy resume, exited 1,
and triggered a full retry that re-killed the gateway the first attempt had
just started — leaving it down and surfacing "Update failed (exit 1)".

Poll a bounded window (up to 30s) for the resumed gateway to publish its
identity, and only treat a persistently empty snapshot as verification
failure. The fail-closed contract from #93406 is preserved: a gateway that
genuinely never comes back still exits 1.
2026-08-26 15:05:32 -07:00
Teknium 4860978115 feat(update): image/package-managed installs refuse in-place updates through one shared gate (#91277 Phase 3)
Every surface that can start an in-place mutation — hermes update
(apply), update --check, and the dashboard's update endpoint — now
routes through evaluate_update_admission(): the baked image-provenance
marker first (authoritative; a bind-mounted checkout inside a container
looks like git to the heuristics while the filesystem is an immutable
image), then the pre-existing docker/nix/apt heuristics verbatim.

A refusal prints the real update command for the deployment kind,
records a 'refused' receipt (fleet tooling sees 'not updatable in
place, use <cmd>' instead of a silent non-update), and exits 2 on CLI
surfaces — distinct from exit-1 errors. The dashboard response keeps
the per-kind error codes its UI already keys on. collect_runtime
inventory()'s updatable_in_place also honors the marker, so --plan and
receipts report image-managed truthfully even with a bind-mounted
checkout.

Live E2E (real hermes update subprocesses, real marker file): apply and
--check both refuse exit-2 with docker-pull guidance, receipts land as
refused/image-marker, an in-place corrupted marker still refuses
(fail-closed), removing the marker admits the git checkout.
2026-08-26 11:41:04 -07:00
Teknium 03537d69dc feat(gateway): updaters pause gateways over the control socket instead of tree-killing them (#92091 step 2)
Windows updates forced a choice between 'gateway survives' and 'update
proceeds': the pause machinery's only tools were the planned-stop marker
poll and the force-kill ladder, so a mid-turn gateway was tree-killed and
its active turn lost. Step 2 of the socket migration adds the
pause-for-update verb: the updater ASKS the gateway to drain in-flight
turns and exit cleanly — releasing every venv file handle on the way out
— through the same request_restart(via_service=True) drain path SIGUSR1
and service restarts already use.

- gateway/run.py: pause-for-update verb handler registered on the
  existing control server; marshals onto the loop thread, ACKs with
  {pausing, already_stopping, pid, drain_timeout}.
- gateway/control_socket.py: pause_gateway_for_update() client — None on
  no-answer (older gateway / no socket), so every caller keeps the
  legacy path when the verb is missing.
- update_cmd.py (_pause_windows_gateways_for_update): socket-first ask
  per mapped profile gateway before the drain wait; positive ACKs extend
  the wait to the gateway's own declared drain budget (+ teardown grace)
  so a mid-turn gateway isn't force-killed at the end of a too-short
  local default. Marker write + force-kill ladder retained verbatim as
  the fallback.

Live E2E: real gateway process (isolated HERMES_HOME), real socket:
identify -> pause ACK {pausing: true} -> gateway drained and exited on
its own (rc=75, zero signals) -> dead-gateway re-ask returns None.
A step-1 gateway without the verb answers ok:false -> client None ->
legacy path (pinned by test).
2026-08-26 09:59:17 -07:00
Teknium 27385e586b feat(update): network-bound serve backends survive hermes update on their recorded endpoints (#63206)
A manually-launched `hermes serve --host <ip>` powering a remote Desktop
was invisible to the entire update pipeline: not in the runtime
inventory, a permanent exit-2 dead-end at the Windows venv-holder guard,
and — when anything killed it — never relaunched, stranding the remote
client on a dead endpoint (#63206). Serve backends were also visible to
`hermes dashboard --stop` but hidden from `--status` (#81564's
asymmetry), so operators could kill what they couldn't see.

Built on the spawn ledger (positive identity, never argv guessing):

- process_identity.py: LedgerEntry gains structured host/port/profile
  (backward-compatible — readers .get()); register_self accepts detail=;
  argv capture widened 6→10 tokens so profiled launches survive.
- web_server.py: serve/dashboard registration moved AFTER the bind and
  now records the ACTUAL bound host/port/profile.
- update_inventory.py: serve/dashboard collector reading the ledger —
  manual backends inventory as supervisor=manual-serve with
  restart_via=respawn-argv; Desktop-owned ones (live recorded spawner)
  as desktop. Plan/receipts/fleet matrix see them for free.
- update_cmd.py: new venv-guard rung — manual serve/dashboard holders
  are stopped for the update and relaunched via an idempotent atexit
  token built from structured identity (same contract as the gateway
  pause/resume); receipts record serve_pause/serve_relaunch.
  Desktop-owned backends keep the refusal (the app respawns what we
  kill).
- dashboard_procs.py: the process scan is augmented with live ledger
  rows, so profiled launches (`hermes --profile p serve ...`) that match
  no substring pattern are finally visible to kill/respawn.
- main.py: `--status` now lists serve-mode backends too, tagged [serve]
  — closing the #81564 status/stop asymmetry.

Salvage note: detection deliberately does NOT reuse #70742's psutil
cmdline-pattern scan (the argv-guessing class this campaign retires);
its resume-token lifecycle (atexit + idempotent flag) and don't-replay
guard shaped the relaunch contract here — credit @Tranquil-Flow.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-08-26 07:57:04 -07:00
Teknium 2f9e187001 revert(macos): remove the TCC interpreter anchor — anchored copies could not load libpython
Reverts the interpreter-anchor halves of #95131 and #95478 (the anchor
module, its doctor check, and the update-time refresh). On real Macs the
anchored real-file copy of the uv interpreter dies in dyld: its LC_RPATH
(@executable_path/../lib) resolves into venv/lib/, which holds no
libpython — bricking EVERY hermes command including update and doctor
(#95425), and the re-pointed python3 aliases lost the stdlib
(ModuleNotFoundError: encodings, #95541). Linux CI could not catch this:
the fixture interpreters were one-byte fakes with no dynamic linking.

Kept: managed_uv._macos_sign_managed_python (#82529, @notkisk) — the
identifier-DR signing of repair generations is independent of the anchor
and unaffected by the dyld issue (it signs binaries IN PLACE in their
store, where their rpath is valid).

Added: doctor's check_macos_tcc_anchor_removed() heals venvs the anchor
already converted — restores bin/python to a symlink at the recorded
source (the anchor's own marker file) and re-points aliases; prints the
manual one-liner if the heal itself fails. Users whose CLI is fully
bricked can run the workaround from #95425 directly.

Re-land criteria: a dylib-complete anchor design (bundle libpython or
rewrite LC_RPATH), verified on macOS hardware BEFORE merge. Credit to
@kim-miram (#95358), @kokhlo (#95476), @zengzheqing (#95551) for the
forward-fix diagnoses that mapped the failure, and to the #95425/#95541
reporters.
2026-08-26 06:50:53 -07:00
Teknium 7125e839d3 fix(state): fail closed when a live process still holds state.db during destructive restore (#90950)
The #91080 harness narrowed the #90950 corruption class to 'something
unlinks/replaces state.db or its -wal/-shm while another process holds
them'. The restore paths are exactly that actor:

- restore_quick_snapshot's unlink+move fallback replaced the inode and
  deleted sidecars under a live holder -> deleted-fd split brain.
- update auto-restore copy2'd a snapshot over a possibly-live DB ->
  the live writer's next checkpoint writes wrong-offset pages (the
  page-1 compression_locks clobber from the report).

Add _foreign_db_holder_pids(), a /proc fd scan that counts holders of
the DB and its sidecars (including already-deleted generations), and
refuse the destructive replacement while any holder exists. The
backup-API path (safe under live connections) remains the primary
route for /snapshot restore.
2026-08-26 06:28:43 -07:00
briandevans 33dce0eb7e refactor(update): fold the auto-restore sequence into a shared helper
Addresses review feedback on the regression test. The test previously parsed
the update_cmd.py AST to assert that each auto-restore call site cleared the
destination's sidecars before copying. That bound the fix to source text rather
than behaviour, and would break on unrelated refactors.

Extract _restore_state_db_from_snapshot(state_path, snap_state), which performs
the clear -> copy -> verify sequence as one unit and returns whether the
restored file passes its integrity check. Both auto-restore paths now call it,
so the ordering is guaranteed by construction instead of by inspection, and the
two byte-identical blocks collapse to a single call each.

The regression test now exercises that helper directly against a database that
still owns a hot WAL: removing the clear from inside the helper fails it with
201 rows where 400 were expected, so the guard remains bound to behaviour.

Also covers the two failure modes the callers already handle: a snapshot that
does not survive the copy returns False, and a missing snapshot raises OSError.
2026-08-26 06:28:43 -07:00
briandevans 86d719067b fix(update): clear stale SQLite sidecars before auto-restoring state.db
The post-update integrity guard (#68474) restores state.db from a pre-update
quick snapshot with a plain shutil.copy2, at both auto-restore sites: the
ZIP-update path in _update_via_zip and the git-pull path in _cmd_update_impl.

The snapshot image is produced by backup._safe_copy_db through sqlite3.backup(),
so it is already checkpointed and owns no WAL. That is precisely why
backup._EXCLUDED_SUFFIXES refuses to ship -wal/-shm/-journal inside a snapshot:
"shipping the live WAL / shared-memory / rollback-journal alongside would pair a
fresh snapshot with stale sidecar state and produce a torn restore on the next
open." The backup side excludes sidecars for that reason; the restore side never
cleared the destination's.

copy2 replaces only the main database file. A state.db-wal belonging to the old,
corrupt database survives the copy and is replayed over the fresh image on the
next open. The restored file then passes PRAGMA integrity_check while serving
the discarded database's contents, so _restored_ok reports valid and the CLI
prints "Auto-restored from snapshot" over data the user has lost. The first
subsequent checkpoint folds the stale WAL in permanently.

A hot -wal is reachable at exactly this moment: a second Hermes holder the
updater's drain did not stop, or the crash that corrupted state.db in the first
place, which is the very trigger for this code path.

Clearing the destination's sidecars is safe here specifically -- they belong to a
database the caller has already declared corrupt and is about to discard.
Contrast preflight_db_writability, which correctly refuses to delete a live WAL.

Reproduced against real SQLite: restoring a 400-row snapshot over a database
with a hot WAL yields 0 of the 400 rows, all 201 visible rows coming from the
old WAL, with integrity_check reporting ok.
2026-08-26 06:28:43 -07:00
ethernet 34c5fcb2a2 fix: update on macos referenced nonexisting variable 2026-08-26 00:24:26 -07:00
webtecnica cc5ff96f9e fix(macos): stable TCC anchor for uv-managed python interpreter (#85345) 2026-08-25 22:10:12 -07:00
David Metcalfe 2d37ed056e fix(macos): harden TCC check against codesign timeouts, clarify scope
Review feedback (AI review on #86391):
- guard _macos_desktop_dr subprocess.run against TimeoutExpired/FileNotFoundError
  so a hanging codesign degrades to the unreadable-DR warning, never crashing
  the doctor run (matches the file's existing subprocess guard pattern)
- select the desktop bundle by newest-mtime across release/mac-*/Hermes.app,
  matching _desktop_packaged_executable, instead of a fixed arch order
- note the cdhash-match proxy assumption at the classification site
- document why /Applications/Hermes.app (Hermes-Setup launcher,
  com.nousresearch.hermes.setup, certificate-anchored) is deliberately not probed
- extend the repair hint to cover per-service resets
- regression tests: codesign timeout and missing-codesign paths
2026-08-25 21:59:36 -07:00
David Metcalfe 36c1755065 fix(macos): detect stale TCC grants and guide one-time re-grant
TCC keys permission grants to the app's code-signing requirement. Grants
made to pre-#73681 builds carry a cdhash-pinned requirement that no
longer matches the rebuilt bundle, so macOS re-prompts on every capture
even though the System Settings toggle shows ON — and the modern prompt
has no Allow button, so users cannot complete the one-time re-grant.

- hermes doctor: new check_macos_tcc_grants() reports the desktop
  bundle's DR class (cdhash-pinned → grants reset on every update;
  identifier-pinned → stable) and prints the exact stale-grant repair
  (tccutil reset, toggle ON, fully quit & relaunch).
- hermes update: after a successful update on macOS with a desktop app
  installed, print the one-line stale-grant guidance.
- docs: desktop.md no longer claims grants persist 'out of the box';
  documents the one-time re-grant for pre-fix grants.

Closes #86385
2026-08-25 21:59:36 -07:00
Royalaid 0c23bf19af fix(update): defer interactive CUA installs on Windows 2026-08-25 16:53:01 -07:00
Krzysztof Radzikowski 36f1423411 fix(update): gateway-only concurrent instances no longer abort hermes update (#37039)
On Windows, the pre-update concurrent-instance gate aborted with exit 2
whenever ANY other process held the venv hermes.exe shim — including the
gateway itself, which _pause_windows_gateways_for_update() stops moments
later and the post-update restart phase brings back. Users with a running
gateway were forced into a manual taskkill dance before every update.

The gate now filters gateway runtimes out of the abort list and proceeds
when nothing else is concurrent. Classification delegates to
_is_pausable_gateway -> gateway.status.looks_like_gateway_command_line
(the canonical shlex-tokenized, profile-selector-aware matcher shared by
the Desktop preflight exemption and the venv-holder guard fallback), so
the gate's exemption and the pause machinery cannot drift apart. Anything
not positively identified as a gateway — REPLs, dashboard, Desktop
backend children, gateway MANAGEMENT commands like 'gateway status',
unreadable cmdlines — still aborts exactly as before, and the abort
message now lists only the PIDs that are actually the user's problem.

Surgical reapply of PR #37039 by @damadorPL onto current main (the gate
moved from hermes_cli/main.py to hermes_cli/update_cmd.py in the main.py
decomposition); his substring classifier was replaced with the canonical
matcher, which also fixes the 'hermes gateway status' misclassification
flagged in review.

Co-authored-by: Hermes <hermes@nousresearch.com>
2026-08-25 14:48:51 -07:00
Jeff Mettel bee489d563 fix(update): restart a booted-out launchd gateway instead of silently skipping it (#74973, salvage #75021)
Port of @jeff-mettel's fix onto the post-#91378/#92902 fleet-restart
shape. The current-profile restart was gated on `launchctl list <label>`
exiting 0 - a booted-out job (plist present, definition deregistered:
crashed helper, manual bootout, failed prior update) fails that check,
so the branch silently skipped: no restart, no message, KeepAlive unable
to revive a definition launchd no longer knows, update printing
'Update complete!' with the gateway down. `launchctl list` is also
session-scoped and unreliable as a loaded/unloaded classifier.

- _restart_launchd_gateway_after_update() (his extraction, adapted):
  plist-exists is the ONLY gate; launchd_restart() owns the
  bootout/bootstrap/kickstart ladder for every plist-present state;
  every failure path is loud and names the manual recovery command.
  The gate-error 'except: pass' (the second silent variant) now counts
  the label failed and tells the operator.
- Success still requires the #92902 supervision verify (fresh
  supervised PID), composing his fix with the returned-is-not-supervised
  guard.
- His regression suite adapted to the (restarted, failed) contract; the
  old 'unregistered -> left alone' pinning test FLIPPED - it pinned the
  bug.

A/B: his suite + the flipped test red on merge-base product code
(silent skip live), green at head. No macOS CI lane exists; field
evidence is #74973's reproductions plus the launchctl print output
shapes pinned in the suite.
2026-08-25 14:10:10 -07:00
Teknium 93bf6f7225 fix(cli): fail closed on empty fleet probe across all pre-update liveness signals (#93406)
The #93410 guard keyed on (restarted_services or killed_pids), which never
fires on Windows: _pause_windows_gateways_for_update /
_resume_windows_gateways_after_update populate neither list, so a healthy
resumed Windows gateway still yielded zero fleet rows and exit 0.

Hoist the decision into _fleet_probe_expected_runtimes(), keyed on every
pre-update liveness signal:
- restarted_services / killed_pids (POSIX restart bookkeeping)
- _pre_restart_gateway_pids non-empty or None (unreadable pre-state,
  same fail-closed contract as _restart_phase_failure_is_incomplete, #78574)
- pre-update plan inventoried >=1 runtime
- Windows pause/resume token carries profiles or unmapped entries

Gate the 2.0s settle sleep on the same condition so a resumed Windows
gateway gets its settle window before the probe. The guard keys only on
zero-rows-despite-expected-runtimes; non-empty snapshots (including
'unknown'-state rows) are still judged solely by print_fleet_version_matrix.

Regression tests cover: empty snapshot + plan runtimes -> incomplete;
empty snapshot + genuinely idle -> success; Windows-resume token path ->
fail-closed + settle sleep wiring.

Builds on RelaxJonh's #93410. Fixes #93406
2026-08-24 03:21:18 -07:00
RelaxJonh d74bbb9bd9 fix(cli): treat empty fleet probe as incomplete when gateways were restarted (#93406)
collect_fleet_versions() swallows every probe exception via
logger.debug() and returns whatever accumulated — which can be an
empty list.  print_fleet_version_matrix([]) returns False (no rows
to report), so the update exits 0 with "success" even though no
gateway was actually verified.

After the restart phase touches live gateways (restarted_services or
killed_pids is truthy), an empty fleet snapshot means verification
failed, not that everything is healthy.  Treat it as incomplete so
the receipt records "partial" and the exit code is 1.

Fixes #93406
2026-08-24 03:21:18 -07:00
Teknium 706f33d424 feat(update): sibling profiles' configs migrate with the fleet — no more silent version drift (#20438/#54926/#79048)
The shared checkout serves every profile, but hermes update migrated
only the active profile's config.yaml. Siblings kept their old
_config_version until their (correctly restarted, post-#91378) gateway
hit a config shape the new code couldn't read — the last unabsorbed
substance from the Phase-2 restart-swarm audit (#20438 earliest, 2026
field repro on #79048: sibling at v33 vs v37).

_migrate_sibling_profile_configs(): per sibling home, scope config
reads/writes via the context-local HERMES_HOME override (ContextVar —
never os.environ), check version, run the NON-INTERACTIVE safe
migration; prompt-requiring settings stay for the profile's own next
interactive session (same contract as gateway-mode). Broken profiles
are skipped without blocking the sweep; override always reset.

Sabotage-verified; live E2E in a fresh process with real drifted
config files: v12→v38 and v25→v38 on disk, provider preserved, the
documented #81946 personality-reset migration correctly applied to
siblings too, never-configured profile untouched, active home
untouched, second run idempotent.
2026-08-23 05:12:59 -07:00
Teknium 18b7fc82b6 feat(update): the plan is now the restart worklist — every planned runtime must be accounted for (#91277 Phase 2)
The policy table was observational: restart_via was a display string and
the four platform restart branches re-discovered their own targets, so a
runtime the plan saw could be missed with zero signal (the #88654 class,
structurally).

- update_inventory: restart_via becomes a machine-readable mechanism id
  (systemd|launchd|desktop|manual) — THE policy table as data; display
  derived via describe_restart_mechanism. match_runtime_outcomes()
  reconciles every planned runtime against the restart phase's
  bookkeeping (restarted/stopped/failed/unaccounted);
  report_unaccounted_runtimes() is the silent-miss tripwire.
- update_cmd: after the restart phase, the plan is reconciled; outcomes
  land in the receipt (runtime_outcomes); any unaccounted runtime
  escalates exactly like a STALE/DOWN fleet row (exit 1).

Sabotage-verified (reconciliation forced to 'restarted' fails the
tripwire tests); live E2E on this host's real fleet: the real
systemd-supervised gateway classified with a machine id, reported
unaccounted when the bookkeeping omits it, clean when accounted.
2026-08-23 04:45:29 -07:00
Jack Lau 1bf93660f1 fix(update): verify launchd is supervising the gateway after a restart
On macOS, `hermes update` printed "Update complete!" and exited 0 while the
ai.hermes.gateway LaunchAgent sat deregistered for 36 minutes (#88848).

_restart_macos_launchd_gateways already disagrees with itself about what
"restarted" means. Sibling profiles are only appended to restarted_services
once _wait_for_launchd_service_pid confirms launchd is running the job on a
fresh pid. The invoking profile was appended on "launchd_restart() did not
raise" alone.

That is a weaker claim than it looks. launchd_restart() returns as soon as the
restart has been REQUESTED: the _request_gateway_self_restart branch hands the
work to the running gateway and returns immediately, and a plist reload is
handed to a detached helper. Both are asynchronous, so a helper that dies
before its first bootstrap, or a `launchctl bootstrap` that exits 0 without
registering (measured by the reporter on macOS 26.6.1), were both invisible to
the caller. The systemd branch of the same phase has never drawn that
inference: it polls _wait_for_service_active before recording the unit.

Verification is domain-agnostic via a new
gateway.wait_for_launchd_gateway_supervision, NOT _wait_for_launchd_service_pid.
The sibling helper needs an explicit domain, and the invoking profile's gate
deliberately avoids a domain locate because it fails on macOS-26 hosts whose
per-user domains reject service management even though launchd_restart() owns
that fallback. The new helper judges by a live supervised pid rather than an
exit code (the predicate _launchctl_label_supervising_process already existed;
this only adds the wait), and returns True immediately when the detached
fallback marker is present, because a gateway running unsupervised there is the
designed state and not the silent failure this guards against.

A label that restarts but is never supervised now lands in
failed_or_stale_units, which sets gateway_fleet_restart_incomplete and makes
the update exit non-zero instead of reporting success over a gateway that is
down.

Tests: 12 in tests/hermes_cli/test_update_launchd_restart_verification.py, with
no platform gate, driving the real _restart_macos_launchd_gateways through
mocked launchctl outcomes. Reverting the verification to an unconditional
append fails 2 of them, including the #88848 regression case.

tests/hermes_cli/test_update_launchd_fleet_restart.py::_fleet stubs the new
verifier so its 27 existing cases keep asserting on routing rather than on a
real launchctl probe; unstubbed, each case would poll the full supervision
budget.
2026-08-23 04:45:29 -07:00
Jack Lau dfcef70061 fix(update): stop a gateway we cannot relaunch instead of leaving it on stale code
Fixes #88654.

After an in-place update, the manual-gateway leg of the restart phase did
this for every profile-mapped gateway:

    restart_mode = _prepare_profile_gateway_update_restart(proc.profile, pid)
    if restart_mode is None:
        continue

A None means no relaunch could be armed. The bare continue skipped the
drain and the stop, and the unmapped sweep immediately below skips any
pid already in profile_processes, so the process was never killed and
never counted into the "Stopped N manual gateway process(es)" summary.
The gateway kept running with its pre-update modules resident while the
new code sat on disk, and every lazy import from that point mixed
versions:

    cannot import name '_MAX_TOOL_ERROR_CHARS' from 'tools.registry'

with no operator signal of any kind.

Two changes.

_prepare_profile_gateway_update_restart now falls back to replaying the
process's own captured command line when the profile-derived relaunch
cannot be armed. launch_detached_gateway_restart_by_cmdline already
exists for exactly this case and documents itself as the companion for
gateways with no profile mapping; the Windows post-update path already
uses it the same way. The argv is captured a few lines earlier for the
external-supervisor check, so the fallback costs nothing extra. The
external-supervisor branch still short-circuits first, because replaying
argv there would escape the manager and race its replacement process.

When neither mechanism can arm a relaunch, the update path no longer
falls through silently. It says so, naming the profile and pid, and hands
the process to the existing unmapped sweep so it is stopped and reported
through the established "Restart manually: hermes gateway run" contract.
Leaving it running was the actual harm: a gateway on stale modules fails
every lazy import for as long as it lives.
2026-08-23 04:25:18 -07:00
Teknium 0c14f060db fix: import managed_python_env at the git-path site; assert the managed-env contract in the repair test
The salvaged commit called managed_python_env() at the git-path sync
without an in-scope import (UnboundLocalError on every git update — CI
red). The repair test pinned the raw {**os.environ, VIRTUAL_ENV} dict, a
change-detector on exactly the construction #83914 replaces; it now
asserts the managed-env contract.
2026-08-23 03:55:14 -07:00
Teknium fbfdb9312b fix(update): widen UV-env isolation to the sibling dependency-sync sites
The salvaged fix covered the git-path sync; the same raw-os.environ
construction existed at the main update path and the interrupted-install
recovery path. All three now build their uv env via managed_python_env()
(#83914 class — same bug, all sites).

A/B-proven with real uv: poisoned UV_PYTHON/UV_SYSTEM_PYTHON steers the
merge-base construction into the hijacker's interpreter (VERDICT:
HIJACKED); the managed construction installs into the install's venv
(VERDICT: ISOLATED). Compose-checked with #92824's stale-VIRTUAL_ENV pin:
isolation + pin together install into the running interpreter on the
site-packages shape.
2026-08-23 03:55:14 -07:00
suntech-wang 6ce145f38f test(update): lock managed uv-env isolation regression
Address review feedback:
- Add two unit tests asserting the update's uv_env contract: third-party
  UV_PYTHON_INSTALL_DIR is dropped, managed pins (UV_MANAGED_PYTHON=1,
  UV_NO_CONFIG=1) are set, VIRTUAL_ENV points at this install's venv, and
  the managed store stays under .hermes-runtime.
- Drop the inline dated comment in favor of intent description.
2026-08-23 03:55:14 -07:00
suntech-wang 08f5a0a98b fix(update): isolate pip install from third-party UV env vars
uv respects UV_PYTHON_INSTALL_DIR from the process environment. When a
third-party app (e.g. WorkBuddy) sets a User-level UV_PYTHON_INSTALL_DIR,
the update's uv pip install can target the wrong interpreter and fail
installing extras, leaving the venv entry-point shims missing. Use the
official managed_python_env() isolation (drops VIRTUAL_ENV/PYTHONPATH/
UV_PYTHON, forces UV_PYTHON_INSTALL_DIR to .hermes-runtime/python,
UV_NO_CONFIG=1) and then point VIRTUAL_ENV at this install's venv.
2026-08-23 03:55:14 -07:00
Franci Penov e366df6889 fix(cli): treat a fork's upstream sync as an update
On a fork, `hermes update` compares HEAD against origin/main, and only then
syncs the fork from upstream — inside the `commit_count == 0` branch, which
returns immediately afterwards. So an update that pulls hundreds of commits
from upstream prints "Already up to date!" and skips everything the
post-update path does, including the dependency sync and the gateway restart.

Observed on a fork-based deployment: 1654 commits pulled, "Already up to
date!", and the launchd gateway left running. It then held pre-update modules
in memory while lazily importing post-update ones, and failed later with an
AttributeError for a method that plainly exists on disk — a mixed runtime that
looks nothing like an update problem. Correlating every run in update.log, a
restart happened on exactly the runs that pulled upstream *without* also
claiming to be up to date, and never once they started co-occurring.

Decide before the branch: capture HEAD, sync, and if HEAD moved, set
commit_count from the range so the normal post-update path runs. The pull that
follows is a no-op (the sync updates origin too); reaching the restart is the
point. commit_count is floored at 1 — HEAD moving *is* the update, so a failed
or zero count query must not send us back down the early return.

steps still being skipped afterwards.

Refs #73108
2026-08-23 00:19:46 -07:00
Teknium 1684877868 fix(update): a gateway killed by the restart phase and never replaced now fails the fleet check (DOWN row)
Phase-1 verification gap (#91277, found auditing our own landed matrix
against the mapped issues): collect_fleet_versions only listed gateways
with a LIVE pid, so 'restart stopped it and nothing came back' produced
NO row at all — the exact silent-failure shape the matrix exists to
catch (#88848/#74973 class) passed with exit 0.

- collect_fleet_versions(pre_restart_pids=...): a dead pid becomes a
  'down' row only when it was alive at update start AND its runtime
  status still claims a running state. Rollout-safe: no snapshot (old
  callers), clean stops, startup failures, and stale records from
  long-dead gateways keep the historical no-row behavior.
- print_fleet_version_matrix escalates on down rows like stale ones
  (exit 1) with the per-profile restart remediation.
- cmd_update passes its existing pre-restart PID snapshot.

Sabotage-verified (reverting the membership check fails the new test);
live-verified with a real spawned-then-killed process producing the
DOWN row and matrix escalation.
2026-08-22 23:46:06 -07:00
Teknium f4067774aa fix(update): token-based control-plane classifier + live E2E for the Desktop-lifecycle cold-start skip (#76129 salvage follow-up)
On top of @686f6c61's premise-corrected #76745:

- _looks_like_desktop_control_plane now uses the parser-derived
  _hermes_holder_subcommand instead of substring matching — the
  #90778/#91869 class ('-m dashboard chat' and 'kanban --preserve-cache'
  argv no longer read as control planes). Regression test added,
  sabotage-verified (reverting to substrings fails it).
- Live E2E (this host, real processes + real spawn ledger): live
  supervised serve owns lifecycle; killed spawner (orphan) does not;
  dead serve entry excluded; empty ledger does not.
- Live Windows E2E for the wine2e lane: real self-registered ledger
  entry suppresses the actual cold-start plan; dead serve restores it;
  holder-scan fallback rung proves the token classifier live.

Co-authored-by: 686f6c61 <github@00b.tech>
2026-08-22 21:36:44 -07:00
686f6c61 4ccc4b6931 fix(update): skip Windows gateway cold-start when Desktop owns lifecycle
Vestigial autostart is not proof the user wants a standalone gateway
run. When Desktop currently supervises this install's control plane,
the updater must not spawn a competing messaging daemon. Serve is not
treated as gateway-equivalent.
2026-08-22 21:36:44 -07:00
Teknium c9c44d0df9 Merge pull request #92636 from NousResearch/fix/windows-launcher-managed-bin
fix(windows): stage hermes launchers in the managed binary dir, not the git checkout
2026-08-22 20:58:00 -07:00
Teknium 0c435f4601 fix(update): reword refusal message — footgun linter matched prose 'venv open (' as bare open() 2026-08-22 19:30:10 -07:00
Teknium 83864c0b5d fix(update): a contended venv is never mutated — failed shim quarantine now refuses instead of warning (#87331)
The #87331 remaining half: when hermes.exe (or a sibling shim) could not
be renamed aside, the updater printed a warning and ran the installer
anyway — which died partway on the same locks and stranded the venv
between versions.

- _run_quarantined_install gains strict_quarantine: any shim whose
  rename failed every retry aborts BEFORE the install command runs
  (successful renames rolled back), raising ShimQuarantineError.
- The update dependency sync passes strict_quarantine=True. The update
  boundary turns the error into a refusal: defer via the
  update-incomplete marker, exit 2 (recorded as refused by the receipt
  net), never ZIP-fallback. Post-sync repair installs keep warn-and-try
  (their venv is already mutated; refusing buys nothing).
- The recovery installer (_install_repair._run_install_cmd) is strict
  unconditionally: marker survives, next launch retries after the
  holder exits.
- Live Windows E2E for the wine2e lane: a real child holds hermes.exe
  without FILE_SHARE_DELETE (the exact field lock shape), strict path
  refuses with zero installer invocations, releases roll back, and the
  same path proceeds once the holder exits.

Sabotage-verified: reverting the strict wiring makes both fail-closed
tests fail.
2026-08-22 19:30:10 -07:00
emozilla fe95ed3930 Merge origin/main: reconcile with PR #92092 (in-checkout launcher restore)
PR #92092 fixed the same vanished-launcher bug by restoring copies into
the legacy in-checkout hermes-agent\bin from the update tail. That
location is what this branch removes: untracked files there are swept
by the update autostash on every cycle (restore/sweep treadmill, plus a
parked stash entry per update under --keep-stash), and unconditional
exe copies break on relocatable venvs ('uv trampoline failed to
canonicalize script path'). This branch's managed-binary-dir layout
supersedes both mechanisms, so the merge resolves to it:

- drop _sync_windows_cli_launchers and its _ensure_acp_launcher call
  (Windows staging/repair lives in ensure_windows_bin_launchers at
  process start and migrate_windows_bin_path in the update tail);
  _ensure_acp_launcher is a Windows no-op again
- keep #92092's genuinely better installer semantics: staging stays in
  a dedicated Install-HermesCommandLaunchers function that throws
  BEFORE any PATH mutation when the required launcher cannot be staged
  and verified -- previously Set-PathVariable could put an empty dir on
  PATH and still print 'hermes command ready'. Reworked for this
  branch's layout: caller passes the destination ($HermesHome\bin),
  launcher form follows the venv (exe copy vs .cmd delegator), and the
  verify step accepts either form
- rework #92092's AST-lifted PowerShell test for the new function
  signature, keeping its fail-before-PATH-mutation assertions and
  adding relocatable-venv form-selection coverage
- drop tests/hermes_cli/test_windows_cli_launcher_repair.py (pinned the
  superseded in-checkout mechanism; equivalent and broader coverage
  lives in tests/hermes_cli/test_ensure_windows_bin_launchers.py)
2026-08-22 22:16:38 -04:00
emozilla 679e9cd294 fix(windows): stage hermes launchers in the managed binary dir, not the git checkout
The installer staged the hermes/hermes-acp launcher copies at
hermes-agent\bin -- inside the git working tree -- and put that dir on
the user PATH (#84452). The update command's pre-pull autostash
(git stash push --include-untracked) swept those untracked, unignored
copies off disk, and once the desktop updater stopped re-applying
stashes (--keep-stash, 5dd221d442) nothing restored them: `hermes`
stopped resolving in every new terminal on every desktop-updated
install.

Move the canonical launcher home to the managed binary dir
(%LOCALAPPDATA%\hermes\bin, next to the managed uv) -- outside the
checkout, where no git operation can ever touch it. The dir is
per-machine and shared by every profile, so all anchoring uses
get_default_hermes_root(), never HERMES_HOME (which points inside
profiles\<name> under `hermes -p`).

The copy design also had a second latent break: managed-uv rebuilds
create relocatable venvs, and a relocatable venv's exe trampoline
resolves relative to its own location -- a copy outside venv\Scripts
dies with 'uv trampoline failed to canonicalize script path'. Launcher
form now depends on the venv (lockstep in install.ps1 and
_install_repair.py): exe copy for normal venvs, a .cmd delegator
invoking the in-venv exe by absolute path for relocatable ones. Either
form counts as present, so pre-rebuild exe copies are left alone.

Delivery to the existing fleet, per cohort:

- already-broken installs cannot run the CLI, so an import-time heal in
  hermes_cli.main (ensure_windows_bin_launchers) re-stages missing
  launchers when the desktop app spawns its backend -- the one channel
  that still reaches them. Gates fail toward inaction: canonical dir
  only for the managed clone, legacy hermes-agent\bin only while the
  user PATH still resolves through it (some pre-managed-uv installs
  have no hermes\bin PATH entry; the legacy re-stage is what fixes
  those). Staging-name + os.replace keeps concurrent process starts
  from tearing a launcher; the helper never raises.
- healthy old-layout installs migrate in the update tail
  (migrate_windows_bin_path): stage canonical launchers, verify them
  BEFORE touching the registry, prepend hermes\bin to the user PATH,
  strip the legacy entries (hermes-agent\bin and venv\Scripts, #83797),
  preserving REG_EXPAND_SZ and raw %VARS%. The legacy dir's files stay
  on purpose -- configs holding absolute launcher paths keep working;
  only the sweepable PATH resolution route goes.
- fresh installs get the new layout from install.ps1 directly.

/bin/ is gitignored so the one update that DELIVERS this fix cannot
sweep pre-migration launchers a final time under the old rules; the
gitignore line, the legacy re-stage branch, and the update-tail call
are transition machinery with a named expiry once the fleet has
migrated.

Also rewrites _ensure_acp_launcher's stale Windows paragraph to match
(raw docstring fixes its invalid \S escape) and updates the Windows
native docs to the new layout, with a docs<->installer parity test.
2026-08-22 13:38:56 -04:00
Gille 9782275b2a fix(windows): restore dedicated CLI launchers on update 2026-08-22 00:05:51 -07:00
JonthanaHanh 01c14ad7f3 fix(update): ZIP swap preserves the built desktop app (apps/desktop/release)
The #70337/#87331 win-unpacked wipe half, from PR #70477 by @JonthanaHanh
(reimplemented against the two-phase staged swap that postdates that
branch — the live release/ dir is grafted into the staged apps copy
BEFORE the atomic commit, so preservation rides the same rollback
machinery instead of a post-hoc copy).

Co-authored-by: JonthanaHanh <92574114+JonthanaHanh@users.noreply.github.com>
2026-08-21 22:08:16 -07:00