Commit Graph

164 Commits

Author SHA1 Message Date
salch-cred 49a71c9727 fix(update): break the cron-update three-way restart deadlock (#100179)
When hermes-auto-update runs \hermes update\ from cron, the update
process lives INSIDE the gateway's own process tree. Waiting for that
gateway to exit is a circular wait:

  gateway  waits on all in-flight work units (#77184 don't-amputate)
    -> cron agent session waits on the \hermes update\ process to exit
      -> \hermes update\ waits on the gateway to exit  [back to A]

The wedged-loop probe (#81642) cannot break it: the cron session posts
activity every ~180s (process-tool poll return), so it is 'actively
waiting forever' and never marked wedged. The gateway logs
'Restart deferred: waiting on 1 active work unit(s)' every 30s until the
1800s force-drain cap amputates its own updater's session — reported as
a 5+ minute hang with gateway_state.json stuck at draining +
restart_requested (v0.21.0, main @ 530aa7b10f).

Fix (the issue's recommended option 1): at both drain sites in
update_cmd.py — systemd (line ~9862) and the bare-process/launchd path
(line ~10203) — check \_is_pid_ancestor_of_current_process(pid)\ before
drain-waiting. When the target gateway IS an ancestor, use
\_request_gateway_self_restart\ (SIGUSR1, no exit-wait) and return: the
gateway's own restart flow completes normally once this process, and
therefore the cron work unit holding it, exits.

Both helpers already exist in hermes_cli/gateway.py (277-304) and
\_request_gateway_self_restart\ already refuses non-ancestor PIDs, so a
normal out-of-tree \hermes update\ keeps its full drain semantics
(including the #86684 cron floor) untouched.

Tests (tests/hermes_cli/test_update_cron_deadlock_guard.py, 6):
- own PID / parent PID are ancestors; 0 and negative are not
- self-restart refuses a non-ancestor PID [linux]
- ancestor path sends SIGUSR1 and NEVER calls _wait_for_pid_exit
  (the deadlock witness — a wait there is the bug) [linux]
- non-ancestor path still drain-waits with the given budget [linux]
Existing graceful/sigusr1/restart tests pass unchanged (9 passed).

Fixes #100179
2026-09-01 08:34:51 -07:00
Justin Wilson fb18fedf29 fix(updater): rebuild desktop on Windows hand-off repair path
The HERMES_UPDATE_REEXEC child and the current-checkout Node repair
path printed success without calling _rebuild_desktop_after_update.
A failed rebuild now withholds the success banner the same way the
commits-pulled path does.

Fixes #97343
2026-09-01 08:27:58 -07:00
JoaoMarcos44 86b50fb43a fix(update): back up HEAD to a rescue ref before orphan-history reset
On orphan divergence (no common ancestor with origin/<branch>, #87694),
`hermes update`'s ff-only fallback went straight to `reset --hard`,
silently discarding the entire local commit graph with no recovery path.

Probe `git merge-base HEAD origin/<branch>` before the reset; when no
common ancestor exists, park the pre-pull SHA under
refs/hermes-update-backups/orphan-<branch>-<utc-ts>-<sha12> via a single
`git update-ref`. Ordinary divergence (ancestor exists) is byte-for-byte
unchanged. The update-ref return code is checked so the user is never
told a backup exists when the write failed.

Bounded growth (size-analysis mandate): a rescue ref pins every object
reachable from the parked commit — in the incident shape that includes a
full working-tree snapshot which can be multi-GB. _prune_orphan_rescue_refs
enforces two limits on every orphan incident: keep at most 10 refs
(count cap) and expire any ref older than 30 days (age expiry, parsed
from the ref-name timestamp). The user-facing message states when the
backup expires.

Tests: orphan backup, honest failure messaging, count-cap prune,
age expiry, unparseable-name safety, ordinary-divergence regression
guard, update-ref sabotage (non-fatal), missing pre-pull SHA, reset
failure persistence, real-git merge-base anchor, and a real-git
end-to-end prune test proving pruned refs unpin objects for gc.

Fixes #87694
Salvaged from #87745 with expiry mitigation added.
2026-09-01 07:01:12 -07:00
joaomarcos 27cd0ff4e8 fix(update): keep serve-unit recovery identity scope-qualified
Review on #96235: discovery distinguished `(scope, unit)`, but the skip
payload and the reported outcomes reduced that to the bare service name.
`user/hermes-serve.service` and `system/hermes-serve.service` are two
different processes, so a single unqualified token could suppress recovery
of both: if the user-scope unit was already settled when the restart phase
aborted, the stale system-scope unit was never restarted and nothing
downstream reported it.

Scope now travels with the unit end to end:

- the in-process systemd loop records a scope-qualified twin of
  `restarted_services` (`restarted_scoped_units`) while the bare-name list
  keeps its existing vocabulary for the fleet probe and the receipt;
- the recovery payload carries `{"scope", "unit"}` objects, and the child
  keys discovery, skips, outcomes and accounting by `(scope, base)`;
- `verified` / `failed` — and therefore the receipt and the completion
  predicate — report `user/hermes-serve`, never a bare name;
- an entry with no scope (a payload written by a pre-update interpreter)
  stays unqualified and is read as scope-agnostic, and an unrecognized
  scope drops the skip rather than honouring it: dropping a skip can only
  cost one more restart-and-verify, honouring an unreadable one can leave
  a stale generation running.

Also from review: the survivor probe compared PIDs alone while the plan
discarded the process incarnation, so a new serve that reused the planned
PID read as the pre-update survivor. The inventory now records the ledger's
`create_time` in the serve/dashboard runtime detail and the probe compares
`(pid, create_time)`, still failing closed when either side has none.

Finally, abort recovery moves out of the update monolith into
`hermes_cli/update_abort_recovery.py` (417 lines) with `update_cmd`
re-exporting the names `hermes_cli.main` and the update flow address.
`update_cmd.py` ends up 75 lines smaller than the PR's base commit instead
of 249 lines larger.

Tests: dual-scope same-name regressions in both directions, proof that no
systemctl verb reaches an already-settled scope, per-scope outcomes, the
legacy unqualified shape, the qualified payload shape, scope-qualified
completion accounting, PID-reuse vs. same-incarnation survivors, and the
inventory carrying `create_time`.

Refs #92145

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YQ9oCBKgAMHSGG8CLEHLMC
2026-09-01 07:00:54 -07:00
joaomarcos f9fec169b5 fix(update): recover hermes-serve units after an aborted restart phase
The fresh-process recovery boundary added for #92145 only reaches gateway
profiles. `hermes serve` -- the runtime that hosts `tui_gateway.server`,
and the process the original report saw failing every chat turn -- is not a
gateway profile, so no `gateway restart` command can reach it and the
gateway-only `collect_fleet_versions` read-back cannot see it either.

The spawn-ledger collector classifies serve/dashboard runtimes purely by
spawner liveness, and a systemd-launched `hermes serve` sets neither
HERMES_SPAWN nor HERMES_PARENT_PID, so it is recorded as `manual-serve`
and the recovery partition skips it as unrecoverable. The result is an
update that clears its incomplete flag on gateway coverage alone while a
live serve process keeps serving the pre-update module graph.

- restart active `hermes-serve*` systemd units from the fresh child,
  enumerated from systemd rather than from the misclassifying inventory,
  and verify a changed MainPID on an active unit before claiming coverage;
- report any pre-update serve/dashboard process that is still the same
  process, and never kill one -- a manual or Desktop-owned serve has no
  relaunch authority;
- require every runtime family, not just the gateway leg, before a
  fresh-process recovery may clear the incomplete flag;
- persist serve-unit outcomes and surviving runtimes in the update receipt.
2026-09-01 07:00:54 -07:00
Teknium 0f16e413e6 fix(update): run pending fleet-restart catchup before the runtime-verification exit gate
The #95294/#91277 fleet contract requires the pending-restart check to
execute on every already-up-to-date pass; the verification exit(1) now
fires only after the catchup runs, so a vulnerable runtime demotes the
outcome to partial without stranding the fleet on stale code.
2026-08-31 12:58:14 -07:00
Teknium d6bf89de40 fix(update): treat an unprobeable post-update interpreter as non-blocking
Verification-features-must-not-false-positive-on-rollout rule: only a
POSITIVE vulnerable SQLite probe demotes success to partial. A dev
checkout without a venv (or a failed probe subprocess) keeps the
success banner and the fleet-restart catchup path — the CI-red
test_update_fleet_restart_pending already-up-to-date trio pinned this.
2026-08-31 12:58:14 -07:00
fangliquanflq 9caf31394f fix(update): verify repaired current-checkout runtime 2026-08-31 12:58:14 -07:00
fangliquanflq 4f19808915 fix(update): fail current-checkout runtime verification 2026-08-31 12:58:14 -07:00
fangliquanflq 7e22725bb4 fix(update): verify SQLite runtime remediation 2026-08-31 12:58:14 -07:00
SayHell0W0rld 6cccc2ef7e fix(update): locate PortableGit under the shared root, not profile home
Review feedback on #88136 (monerostar): a profile-scoped `hermes update`
sets HERMES_HOME to <root>/profiles/<name>, but the Hermes-managed
PortableGit tree lives under the SHARED root (<root>/git/...). The locator
checked get_hermes_home() only, so a broken trampoline during a
profile-scoped update was not swapped and fell through to ZIP.

Extract _portable_git_candidates() (shared root first, profile home as
fallback) and add a regression test for the profile layout.
2026-08-31 12:21:46 -07:00
SayHell0W0rld 0bab9ff9e7 fix(update): self-heal broken Git-for-Windows trampoline on Windows
A Git-for-Windows trampoline launcher (bin\git.exe / cmd\git.exe shim,
~46KB) that fails to re-exec the real git-core binary refuses every git
call with a "BUG (fork bomb)" guard instead of running it (#87876).

Detect the trampoline up front via `git --version`, locate a real git
binary (Git for Windows or Hermes-managed PortableGit locations), and
rebuild the git command with it so fetch/pull/checkout keep working with
a real git instead of degrading to the ZIP fallback. When no real binary
can be found, leave the command untouched so the existing fetch-failure
handler still falls back to the ZIP path on Windows (#88046).
2026-08-31 12:21:46 -07:00
andyst-dev 4d3e1e4d13 fix(desktop): prevent venv scan timeout on busy Windows hosts 2026-08-31 12:00:33 -07:00
Teknium ce942ab125 fix(update): fingerprint orphan backends from the classification psutil handle
The orphan-backend classifier fingerprinted candidates via
gateway.status.get_process_start_time, which prefers /proc/<pid>/stat —
the HOST process table, in clock ticks. Under the fake-psutil test harness
(and any containerized run where the PID number happens to exist on the
host) that returns the WRONG process's fingerprint in the WRONG units,
while pid_is_hermes verifies via psutil centiseconds at kill time: the
guard would then refuse every legitimate reap. Read create_time() from the
same psutil handle used for classification, quantized exactly like
gateway.status does on Windows, so the fingerprint round-trips.

Also covers the Windows-lane sibling: test_uses_netstat_and_taskkill_on_windows
now pins the guarded call path, plus a new refusal test for a non-bridge
listener PID (#89614 class).
2026-08-31 10:41:54 -07:00
Ayush Nangia c923b53913 fix(update): refuse gateway ancestor tree-kill on Windows 2026-08-31 10:41:54 -07:00
burak33bb ed6d5fc803 fix(windows): require process identity before taskkill 2026-08-31 10:41:54 -07:00
gebilaowang404 cdd063528f fix(hermes_cli): fail-closed PID-ownership guard before Windows taskkill
Guard every Windows `taskkill /PID` against stale/recycled PIDs
(#89614: 8x 0xEF blue screens; a rebooted PID can be svchost.exe).

Adopted the community patch by AlexMnrs (commit 0162465): shared
psutil-based (pid, create_time) guard reusing the repo's existing
get_process_start_time machinery:
- fail closed on invalid/unknown/recycled identities (0/-1/None/bool/non-int)
- capture identity at discovery, re-validate at kill time
- all three sites through pid_is_hermes; taskkill stays hidden

Sites: _subprocess_compat.kill_process_tree,
dashboard_procs._kill_stale_dashboard_processes (win32),
update_cmd._stop_process_trees.

Refs #90471, #89614

Co-authored-by: Alex Monrás <AlexMnrs@users.noreply.github.com>
2026-08-31 10:41:54 -07:00
fangliquanflq 915ec169a0 fix(update): verify failed restore cleanup 2026-08-31 10:04:56 -07:00
fangliquanflq 53cf38c3b8 fix(update): fail closed on incomplete restore checks 2026-08-31 10:04:56 -07:00
fangliquanflq 3e4dc5f2c3 fix(update): preserve unknown restore cleanup state 2026-08-31 10:04:56 -07:00
fangliquanflq 8236b51878 fix(update): authenticate import health markers 2026-08-31 10:04:56 -07:00
fangliquanflq 6c608e2f59 fix(update): reject terminated import probes 2026-08-31 10:04:56 -07:00
fangliquanflq 68f9681acb fix(update): capture terminating restored imports 2026-08-31 10:04:56 -07:00
fangliquanflq f6c3942957 fix(update): compare every restored module failure 2026-08-31 10:04:56 -07:00
fangliquanflq 7b958b3575 fix(update): detect restored import-time failures 2026-08-31 10:04:56 -07:00
fangliquanflq 716b10314a fix(update): reject unsafe stash restores 2026-08-31 10:04:56 -07:00
Teknium e49df40682 fix(update): refuse to mutate a venv containing foreign-owned files (#83529)
A venv ever touched by sudo pip / sudo hermes contains root-owned files
(classically site-packages/*.dist-info/INSTALLER). A later normal-user
'hermes update' pulls code fine, then 'uv pip install -e .' dies with
'Permission denied (os error 13)' mid-mutation — venv/bin/hermes already
deleted, CLI bricked.

Add a bounded, pure-stat ownership preflight (_venv_foreign_owned_paths)
that runs after the code pull and immediately before the dependency
install. If foreign-owned paths are found it refuses up front, names the
offending paths + owner uid, prints the exact recovery command
(sudo chown -R $(id -un): <root>), and confirms the venv is untouched.
Windows (no os.geteuid) and root skip entirely. Never raises, capped at
~2000 stat calls, no subprocess use (update tests mock subprocess.run).

Same refuse-before-mutate philosophy as the contended-venv gate (#87331).

Fixes #83529
Diagnosis and documented recovery by @eabase.
2026-08-31 10:04:40 -07:00
yoniebans be284cf5c0 fix(update): report when the official repo was not checked on the up-to-date path
Review follow-up on #97052 (helix4u): a fork with no upstream remote whose HEAD matches origin/main used to print plain "Already up to date!" under --yes even though official main was never consulted, so an unattended stale fork looked current. _sync_with_upstream_if_needed now returns whether the official upstream was actually checked, and the commit_count == 0 completion line says "Up to date with your fork (official repo not checked)." when it was not. Skip-as-decline semantics are unchanged: no prompt, no remote mutation, no decline marker. Caller-level regression test added for the fork + no-upstream + --yes + HEAD==origin/main path; helper tests now pin the return contract.
2026-08-28 13:33:56 -04:00
yoniebans b33fa127d2 fix(update): gate the fork-upstream prompt for --yes and non-tty runs
_sync_with_upstream_if_needed called bare input() with no assume_yes parameter and no tty check, so a fork checkout without an upstream remote wedged hermes update forever in any non-interactive context (CI, cron, the desktop updater hand-off): stdin stays open, EOFError never fires. Thread assume_yes and the gateway input_fn into the helper and skip the prompt as a decline under assume_yes or a non-tty stdio pair, without writing the decline marker or touching git remotes, so interactive runs still get asked later. Both call sites forward the interaction state; the config-migration and stash-restore prompts already carry this gate.

Closes #60240 (prompt half). Supersedes #78678, #92448, #92410.
Co-authored-by: BlackishGreen33 <BlackishGreen33@users.noreply.github.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Co-authored-by: jackulau <jackulau@users.noreply.github.com>
2026-08-28 13:33:56 -04:00
Zheqing Zeng aa72df4b42 fix(macos): re-land dylib-complete TCC interpreter anchor
The first landing (#95131/#95478, reverted in #95563) copied the
uv-store interpreter into venv/bin/python so TCC grants would stick
to a stable path. On real Macs that copy bricked every hermes command
two ways: dynamically-linked builds died in dyld because
@executable_path/../lib/libpython resolved into venv/lib/ (#95425),
and alias symlinks to the copy made CPython getpath lose the venv
prefix (#95541, ModuleNotFoundError: encodings).

Re-land:

- Keep the signed real-file copy of bin/python (identifier-pinned
  via _macos_sign_managed_python).
- Materialize python3 / python3.N as real-file copies, never
  symlinks. Copies boot on every build we could reproduce and keep
  the TCC identity.
- Hardlink store libpython* into venv/lib/ when present (copy across
  devices). Existing LC_RPATH already points there.
- Pre-install boot gate: launch the staged copy, demand encodings
  plus the venv prefix, abort and leave the live venv untouched
  on failure.

Doctor reports/installs the new anchor (the revert-era heal is
removed). Update refreshes it after a successful code swap. Tests
cover layout, idempotence, predecessor-symlink repair, libpython
hardlink, boot-gate refusal, and a macos_only real-interpreter E2E.

Closes #95596.
2026-08-28 09:05:19 +05:30
kshitijk4poor f3cbb262c1 fix(update): valid --ignored=matching mode; rename-only path split; shared preserve constant
Review corrections on the first draft (caught by /simplify-code before
merge — the PR was disarmed for these):

- BLOCKER: --ignored=all is not a valid git mode (git exits 128 'Invalid
  ignored mode'); with it, every ZIP update was refused as 'could not
  check the working tree'. The mocked tests could not see this — a new
  real-git test creates an actual repo + .gitignore and asserts the guard
  runs clean, blocks on an ignored user file, and exempts ignored
  preserved entries. --ignored=matching also reports an ignored dir as
  one line instead of enumerating its contents.
- FAIL-OPEN HOLE: the ' -> ' two-path split now applies only to R/C
  rename/copy status codes. Porcelain v1 does not quote plain filenames
  with spaces, so an ignored file literally named 'venv -> node_modules'
  parsed as two preserved tops and slipped past the guard into the
  destructive swap.
- _update_via_zip's swap loop now consumes _ZIP_PRESERVED_TOP_LEVEL
  instead of a comment-synced duplicate set (change-detector test added).
2026-08-27 21:07:34 +05:30
joaomarcos e64db76982 fix(update): gitignored user files also block the ZIP overlay
Carried from #87392 (closed as superseded — its core guard landed via the
#87327 salvage chain): the dirty-tree check now passes --ignored=all, so a
gitignored-but-real user file (logs, scratch files, local data) blocks the
destructive ZIP overlay too. The ZIP path's own preserved top-level entries
(venv, node_modules, .git, .env — gitignored on every normal install) are
exempted so they don't become a false refusal.

Credit: @JoaoMarcos44, whose #87392 included this hardening.
2026-08-27 21:07:34 +05:30
Cursor Agent 8246c4f92a fix(cli): repair interrupted update fleet restart
An interrupted hermes update after git pull advanced HEAD never
restarted running gateways, and the next update said "Already up to
date" and skipped the fleet. Persist a HERMES_HOME fleet_restart_pending
marker after HEAD moves, clear it only when restart completes (or
nothing was running), and catch up on the next hermes update even when
git is current — also when latest.json records a stale runtime SHA.

Co-authored-by: GokayAI <gokay-ai@users.noreply.github.com>
2026-08-26 21:38:20 -07:00
Teknium 71823be9da fix(update): stop counting the Windows resume token as a fleet runtime
Fixes #93406 (residual). _fleet_probe_expected_runtimes counted the
_windows_gateway_resume pause/resume token (profiles/unmapped entries)
as an 'expected fleet rows' signal. The token is pause/resume
bookkeeping, not a runtime inventory, and its entries have no rows
collect_fleet_versions() can return: unmapped Scheduled-Task gateways
never publish gateway_state.json, and a resumed profile gateway
relaunches detached and may not republish within the probe window. So
every Windows update that paused a gateway set _fleet_rows_expected,
the verification loop silently waited out its polling window (~14 min
wall clock with the retry loop on user reports), printed 'Fleet version
check returned no rows', and exited 1 for an update that succeeded.

Expected-runtimes now keys only on row-capable signals: restart-phase
bookkeeping, the pre-restart PID snapshot, and the pre-update plan
inventory -- which already cover any genuinely live pre-update Windows
gateway.

Counterfactual proof: tests/hermes_cli/test_update_fleet_probe_resume_token.py
fails on the pre-fix predicate (token-only => True) and passes with the
fix; the row-capable signals are pinned unchanged.
2026-08-26 18:23:33 -07:00
fred0m 53057f2bc4 fix(update): run config migration on the 'Already up to date' repair path (#91360)
A failed update attempt can pull fresh code onto disk and then die before
the config-migration block (e.g. a PyPI timeout during the dependency
sync). The desktop hand-off retries; the retry takes the commit_count == 0
branch, repairs deps, prints 'Already up to date!' and returns early -
skipping _run_config_check_fresh / migrate_config entirely. The fresh
code (requiring a newer _config_version) then refuses to start against
the old config until 'hermes doctor --fix' is run.

Fix: _maybe_migrate_config_on_current() mirrors the version_bump_only
handling (silent, non-interactive) and is called on both repair-path
completion points before claiming success.

Also: scripts/desktop-update/posix.sh no longer retries when the update
was deliberately SKIPPED (checkout parked on a non-target branch) -, the
retry is deterministic and only wastes time. Uses a dedicated non-
colliding exit code (8) and an honest message instead of 'Update failed'.

New tests: tests/hermes_cli/test_update_config_migration_on_current.py
(5 cases: migrate-when-behind, noop-current, noop-ahead, warning re-
surface, silent check failure).
2026-08-26 17:42:50 -07:00
loulanyue bf5ff51078 fix(update): check and apply config migrations on current checkout / retry paths (#91360)
When an update was interrupted or failed mid-install (e.g. dependency install
timeout) after pulling new code, the subsequent update run takes the
'commit_count == 0' path and early-returned without checking or migrating
the configuration. Fresh code requiring a newer config version would fail to
boot on the next run.

Extract _check_and_apply_config_migration and invoke it across all update
completion paths (normal update, current checkout / node repair, and python
dependency repair).
2026-08-26 17:42:50 -07:00
Casey 790e1eb6bd fix(update): pause SCM-supervised Windows gateway services before venv mutation
On Windows installs where the gateway runs as an SCM service (WinSW,
NSSM, sc.exe create), the existing pause machinery kills the gateway
process directly — and the service wrapper's failure ladder resurrects
it within seconds, re-taking the venv file locks mid-update. The update
then dies partway through dependency sync with access-denied errors.

This extends _pause_windows_gateways_for_update() to detect when a
gateway's process tree is owned by a running SCM service, and to stop
the SERVICE through sc.exe instead of killing the child:

- gateway/status.py: expose service-ownership discovery for gateway
  runtimes (find_windows_gateway_services maps validated gateway PIDs
  through process ancestry to running SCM service PIDs, with
  create-time identity checks against PID reuse).
- hermes_cli/update_cmd.py: stop verified services via sc.exe before
  venv mutation and restart them afterward. Stops wait for a stable
  SCM 'stopped' state AND for the original descendant processes to
  exit (service 'Stopped' is not proof the child released its
  handles). Failure to prove ownership, stop a service, or restart it
  fails closed; rollback restores attempted services, and rollback
  failures are surfaced rather than swallowed.
- Fail-closed throughout: unreadable identities, ambiguous ancestry,
  or a service that will not reach a stable state abort the update
  before any file mutation.

Complements #37039 (gateway-only concurrent instances no longer abort):
that fix lets the update proceed past the gate; this one makes the
pause actually stick when the gateway is service-supervised.

Note: tests/gateway/test_status.py::TestReadProcessCmdlinePsFallback::
test_ps_fallback_when_proc_unavailable fails on Windows on current main
before this change as well (POSIX ps fallback asserted on a platform
without it); all other touched suites pass (155 passed, 5 skipped).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-26 16:45:31 -07:00
Teknium b2c011364e fix(update): conservative outcomes + serve-ledger coverage for fresh restart recovery
Salvage adjustments to PR #94392 per review:

- Narrow the supervisor claim to the systemd-VERIFIED path only. The fresh
  recovery child now probes 'systemctl --user is-active' after each relaunch;
  only an observed-active systemd unit is reported 'verified'. A relaunch that
  merely exited 0 is labelled 'relaunch_attempted', never counts as supervisor
  coverage, and never clears gateway_fleet_restart_incomplete.
- Serve-owned runtimes (serve/dashboard entries from the spawn ledger, per the
  update_inventory serve collector) are no longer silently skipped: the
  recovery pass records them (and manual gateways) as skipped-with-reason in
  the recovery result and the persisted update receipt.
- Receipt fresh_recovery persists the conservative vocabulary
  (requested/verified/relaunch_attempted/failed/skipped); 'succeeded' is gone.
- Added an end-to-end test that drives the real recovery module in a genuinely
  fresh interpreter (sitecustomize shim intercepts the grandchild
  'gateway restart' and systemctl probes).
2026-08-26 16:45:26 -07:00
joaomarcos ccdd7f41ed fix(update): persist per-profile recovery outcomes 2026-08-26 16:45:26 -07:00
joaomarcos f0045c5383 fix(update): verify fresh restart recovery results 2026-08-26 16:45:26 -07:00
joaomarcos 5609ccbece fix(update): recover aborted gateway restart in a fresh process 2026-08-26 16:45:26 -07:00
Teknium df3d41ee67 fix(update): sweep aborted-fetch tmp_pack debris before it corrupts the pack directory (#93732)
Every git fetch that dies mid-transfer (timeout, HTTP 429, dropped
line) strands a tmp_pack_* file in .git/objects/pack, and git never
cleans them. The banner's background update check is the main generator
on flaky lines — several aborted fetches a day — and the reporter's
install accumulated hundreds of files / 6.0 GB over 9 days until the
pack directory corrupted outright and every update check hung or
failed permanently.

clear_stale_tmp_packs() in gitlock.py sweeps tmp_pack_/tmp_idx_/
tmp_rev_/tmp_mtimes_ debris with the exact safety contract the lock
sweep already uses: only files past the 10-minute age floor, never
while any git process runs, never raises, real pack-*.pack/.idx files
untouchable by construction (prefix match). Wired into all three
fetch-adjacent sites: _cmd_update_check, the update apply path, and
the banner's passive check (generator = janitor).

Live E2E: 300 aged tmp_pack files (the reported scale-shape) swept
from a real repo; an in-flight fresh tmp and ancient real packs
survived; fsck clean and a real fetch round-trip succeeded after.
2026-08-26 16:45:22 -07:00
pierrenode de2a9de788 fix(update): feed Windows gateway relaunch outcome into fleet reconciliation
#91277 Phase 2's plan-vs-execution reconciliation (match_runtime_outcomes)
cross-checks every runtime collect_runtime_inventory() saw against
restarted_services / relaunched_profiles / externally_supervised_profiles /
killed_pids — the systemd/launchd restart phase's bookkeeping. That
inventory is cross-platform (control-socket / PID-file based), so it
includes Windows gateways too, but Windows's own pause/resume mechanism
(_pause_windows_gateways_for_update / _resume_windows_gateways_after_update)
never wrote into any of that bookkeeping.

Result: a Windows gateway that was correctly stopped and relaunched by
_resume_windows_gateways_after_update was still classified "unaccounted" by
the reconciliation (the plan saw it and no bookkeeping mentions it) —
report_unaccounted_runtimes() escalates that into sys.exit(1), and in
gateway_mode also writes ".update_exit_code"="1". Every successful
`hermes update` on Windows with a running gateway reported itself as
failed, unconditionally (the sys.exit(1) is not gated to gateway_mode).

_resume_windows_gateways_after_update now records the profiles it
successfully relaunched onto the resume token; _cmd_update_impl merges
that into the shared relaunched_profiles list right before reconciliation
runs. A profile whose relaunch genuinely fails is deliberately left off
the list, so it still surfaces as unaccounted — Windows has no watcher to
recover a failed relaunch, so that escalation is the correct signal.

Regression tests exercise _resume_windows_gateways_after_update directly
(records successes, omits failures) and reproduce the reconciliation-level
bug end to end: the same plan row resolves "unaccounted" without the merge
and "restarted" with it. Mutation-verified: with the fix reverted, three of
the four new tests fail (KeyError on the token / wrong outcome).
2026-08-26 16:14:27 -07:00
AlexGabbia b3e477f304 fix(update): wait for resumed Windows gateway before failing fleet check
The post-update fleet version check slept 2s and probed once. On Windows the
resume path relaunches the gateway detached, and it needs ~10s to boot (the
Telegram polling reconnect) before it stamps gateway_state.json or answers the
control socket. That race reported "no rows" for a healthy resume, exited 1,
and triggered a full retry that re-killed the gateway the first attempt had
just started — leaving it down and surfacing "Update failed (exit 1)".

Poll a bounded window (up to 30s) for the resumed gateway to publish its
identity, and only treat a persistently empty snapshot as verification
failure. The fail-closed contract from #93406 is preserved: a gateway that
genuinely never comes back still exits 1.
2026-08-26 15:05:32 -07:00
Teknium 4860978115 feat(update): image/package-managed installs refuse in-place updates through one shared gate (#91277 Phase 3)
Every surface that can start an in-place mutation — hermes update
(apply), update --check, and the dashboard's update endpoint — now
routes through evaluate_update_admission(): the baked image-provenance
marker first (authoritative; a bind-mounted checkout inside a container
looks like git to the heuristics while the filesystem is an immutable
image), then the pre-existing docker/nix/apt heuristics verbatim.

A refusal prints the real update command for the deployment kind,
records a 'refused' receipt (fleet tooling sees 'not updatable in
place, use <cmd>' instead of a silent non-update), and exits 2 on CLI
surfaces — distinct from exit-1 errors. The dashboard response keeps
the per-kind error codes its UI already keys on. collect_runtime
inventory()'s updatable_in_place also honors the marker, so --plan and
receipts report image-managed truthfully even with a bind-mounted
checkout.

Live E2E (real hermes update subprocesses, real marker file): apply and
--check both refuse exit-2 with docker-pull guidance, receipts land as
refused/image-marker, an in-place corrupted marker still refuses
(fail-closed), removing the marker admits the git checkout.
2026-08-26 11:41:04 -07:00
Teknium 03537d69dc feat(gateway): updaters pause gateways over the control socket instead of tree-killing them (#92091 step 2)
Windows updates forced a choice between 'gateway survives' and 'update
proceeds': the pause machinery's only tools were the planned-stop marker
poll and the force-kill ladder, so a mid-turn gateway was tree-killed and
its active turn lost. Step 2 of the socket migration adds the
pause-for-update verb: the updater ASKS the gateway to drain in-flight
turns and exit cleanly — releasing every venv file handle on the way out
— through the same request_restart(via_service=True) drain path SIGUSR1
and service restarts already use.

- gateway/run.py: pause-for-update verb handler registered on the
  existing control server; marshals onto the loop thread, ACKs with
  {pausing, already_stopping, pid, drain_timeout}.
- gateway/control_socket.py: pause_gateway_for_update() client — None on
  no-answer (older gateway / no socket), so every caller keeps the
  legacy path when the verb is missing.
- update_cmd.py (_pause_windows_gateways_for_update): socket-first ask
  per mapped profile gateway before the drain wait; positive ACKs extend
  the wait to the gateway's own declared drain budget (+ teardown grace)
  so a mid-turn gateway isn't force-killed at the end of a too-short
  local default. Marker write + force-kill ladder retained verbatim as
  the fallback.

Live E2E: real gateway process (isolated HERMES_HOME), real socket:
identify -> pause ACK {pausing: true} -> gateway drained and exited on
its own (rc=75, zero signals) -> dead-gateway re-ask returns None.
A step-1 gateway without the verb answers ok:false -> client None ->
legacy path (pinned by test).
2026-08-26 09:59:17 -07:00
Teknium 27385e586b feat(update): network-bound serve backends survive hermes update on their recorded endpoints (#63206)
A manually-launched `hermes serve --host <ip>` powering a remote Desktop
was invisible to the entire update pipeline: not in the runtime
inventory, a permanent exit-2 dead-end at the Windows venv-holder guard,
and — when anything killed it — never relaunched, stranding the remote
client on a dead endpoint (#63206). Serve backends were also visible to
`hermes dashboard --stop` but hidden from `--status` (#81564's
asymmetry), so operators could kill what they couldn't see.

Built on the spawn ledger (positive identity, never argv guessing):

- process_identity.py: LedgerEntry gains structured host/port/profile
  (backward-compatible — readers .get()); register_self accepts detail=;
  argv capture widened 6→10 tokens so profiled launches survive.
- web_server.py: serve/dashboard registration moved AFTER the bind and
  now records the ACTUAL bound host/port/profile.
- update_inventory.py: serve/dashboard collector reading the ledger —
  manual backends inventory as supervisor=manual-serve with
  restart_via=respawn-argv; Desktop-owned ones (live recorded spawner)
  as desktop. Plan/receipts/fleet matrix see them for free.
- update_cmd.py: new venv-guard rung — manual serve/dashboard holders
  are stopped for the update and relaunched via an idempotent atexit
  token built from structured identity (same contract as the gateway
  pause/resume); receipts record serve_pause/serve_relaunch.
  Desktop-owned backends keep the refusal (the app respawns what we
  kill).
- dashboard_procs.py: the process scan is augmented with live ledger
  rows, so profiled launches (`hermes --profile p serve ...`) that match
  no substring pattern are finally visible to kill/respawn.
- main.py: `--status` now lists serve-mode backends too, tagged [serve]
  — closing the #81564 status/stop asymmetry.

Salvage note: detection deliberately does NOT reuse #70742's psutil
cmdline-pattern scan (the argv-guessing class this campaign retires);
its resume-token lifecycle (atexit + idempotent flag) and don't-replay
guard shaped the relaunch contract here — credit @Tranquil-Flow.

Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
2026-08-26 07:57:04 -07:00
Teknium 2f9e187001 revert(macos): remove the TCC interpreter anchor — anchored copies could not load libpython
Reverts the interpreter-anchor halves of #95131 and #95478 (the anchor
module, its doctor check, and the update-time refresh). On real Macs the
anchored real-file copy of the uv interpreter dies in dyld: its LC_RPATH
(@executable_path/../lib) resolves into venv/lib/, which holds no
libpython — bricking EVERY hermes command including update and doctor
(#95425), and the re-pointed python3 aliases lost the stdlib
(ModuleNotFoundError: encodings, #95541). Linux CI could not catch this:
the fixture interpreters were one-byte fakes with no dynamic linking.

Kept: managed_uv._macos_sign_managed_python (#82529, @notkisk) — the
identifier-DR signing of repair generations is independent of the anchor
and unaffected by the dyld issue (it signs binaries IN PLACE in their
store, where their rpath is valid).

Added: doctor's check_macos_tcc_anchor_removed() heals venvs the anchor
already converted — restores bin/python to a symlink at the recorded
source (the anchor's own marker file) and re-points aliases; prints the
manual one-liner if the heal itself fails. Users whose CLI is fully
bricked can run the workaround from #95425 directly.

Re-land criteria: a dylib-complete anchor design (bundle libpython or
rewrite LC_RPATH), verified on macOS hardware BEFORE merge. Credit to
@kim-miram (#95358), @kokhlo (#95476), @zengzheqing (#95551) for the
forward-fix diagnoses that mapped the failure, and to the #95425/#95541
reporters.
2026-08-26 06:50:53 -07:00
Teknium 7125e839d3 fix(state): fail closed when a live process still holds state.db during destructive restore (#90950)
The #91080 harness narrowed the #90950 corruption class to 'something
unlinks/replaces state.db or its -wal/-shm while another process holds
them'. The restore paths are exactly that actor:

- restore_quick_snapshot's unlink+move fallback replaced the inode and
  deleted sidecars under a live holder -> deleted-fd split brain.
- update auto-restore copy2'd a snapshot over a possibly-live DB ->
  the live writer's next checkpoint writes wrong-offset pages (the
  page-1 compression_locks clobber from the report).

Add _foreign_db_holder_pids(), a /proc fd scan that counts holders of
the DB and its sidecars (including already-deleted generations), and
refuse the destructive replacement while any holder exists. The
backup-API path (safe under live connections) remains the primary
route for /snapshot restore.
2026-08-26 06:28:43 -07:00
briandevans 33dce0eb7e refactor(update): fold the auto-restore sequence into a shared helper
Addresses review feedback on the regression test. The test previously parsed
the update_cmd.py AST to assert that each auto-restore call site cleared the
destination's sidecars before copying. That bound the fix to source text rather
than behaviour, and would break on unrelated refactors.

Extract _restore_state_db_from_snapshot(state_path, snap_state), which performs
the clear -> copy -> verify sequence as one unit and returns whether the
restored file passes its integrity check. Both auto-restore paths now call it,
so the ordering is guaranteed by construction instead of by inspection, and the
two byte-identical blocks collapse to a single call each.

The regression test now exercises that helper directly against a database that
still owns a hot WAL: removing the clear from inside the helper fails it with
201 rows where 400 were expected, so the guard remains bound to behaviour.

Also covers the two failure modes the callers already handle: a snapshot that
does not survive the copy returns False, and a missing snapshot raises OSError.
2026-08-26 06:28:43 -07:00