The fresh-process recovery boundary added for #92145 only reaches gateway
profiles. `hermes serve` -- the runtime that hosts `tui_gateway.server`,
and the process the original report saw failing every chat turn -- is not a
gateway profile, so no `gateway restart` command can reach it and the
gateway-only `collect_fleet_versions` read-back cannot see it either.
The spawn-ledger collector classifies serve/dashboard runtimes purely by
spawner liveness, and a systemd-launched `hermes serve` sets neither
HERMES_SPAWN nor HERMES_PARENT_PID, so it is recorded as `manual-serve`
and the recovery partition skips it as unrecoverable. The result is an
update that clears its incomplete flag on gateway coverage alone while a
live serve process keeps serving the pre-update module graph.
- restart active `hermes-serve*` systemd units from the fresh child,
enumerated from systemd rather than from the misclassifying inventory,
and verify a changed MainPID on an active unit before claiming coverage;
- report any pre-update serve/dashboard process that is still the same
process, and never kill one -- a manual or Desktop-owned serve has no
relaunch authority;
- require every runtime family, not just the gateway leg, before a
fresh-process recovery may clear the incomplete flag;
- persist serve-unit outcomes and surviving runtimes in the update receipt.
Salvage adjustments to PR #94392 per review:
- Narrow the supervisor claim to the systemd-VERIFIED path only. The fresh
recovery child now probes 'systemctl --user is-active' after each relaunch;
only an observed-active systemd unit is reported 'verified'. A relaunch that
merely exited 0 is labelled 'relaunch_attempted', never counts as supervisor
coverage, and never clears gateway_fleet_restart_incomplete.
- Serve-owned runtimes (serve/dashboard entries from the spawn ledger, per the
update_inventory serve collector) are no longer silently skipped: the
recovery pass records them (and manual gateways) as skipped-with-reason in
the recovery result and the persisted update receipt.
- Receipt fresh_recovery persists the conservative vocabulary
(requested/verified/relaunch_attempted/failed/skipped); 'succeeded' is gone.
- Added an end-to-end test that drives the real recovery module in a genuinely
fresh interpreter (sitecustomize shim intercepts the grandchild
'gateway restart' and systemctl probes).
Review on #91283: begin_update_receipt() fires early in _cmd_update_impl,
but finalization only existed on the success/ZIP/CalledProcessError
paths. Early sys.exit paths (Windows concurrent-instance preflight,
venv-holder refusal, head-pinned no-op, fetch failure) terminated with
the receipt started but never written — losing exactly the refused/
failed runs the receipt matters most for.
- update_receipt.py: finalize_pending_update_receipt(exit_code,
stop_reason) — boundary safety net; maps exit 2 → 'refused', other
non-zero → 'failed'; records exit_code + stop_reason. Exactly-once by
construction (singleton popped in finalize_update_receipt), so runs
the inner paths already finalized are untouched.
- main.py cmd_update: SystemExit/BaseException/else arms around
_cmd_update_impl persist any still-open receipt with the real exit
code, then re-raise unchanged. Future early exits are covered without
per-site finalize patches.
- 5 regression tests incl. end-to-end through the real cmd_update
wrapper (exit 2 preserved, outcome 'refused', stop reason recorded,
singleton cleared, exactly one receipt file).
Phase 1 of the fleet-update reliability plan (#91277): the updater now
proves its outcome instead of assuming it.
- hermes_cli/build_info.py: get_code_identity() — process-cached code
identity (git sha for source installs, baked .hermes_build_sha for
Docker images, pyproject version).
- gateway/status.py: every runtime-status write stamps the writer's
code_sha/code_version into gateway_state.json, so a running gateway's
actual code generation is observable from disk.
- hermes_cli/update_receipt.py (new): machine-readable receipt of each
update run (steps, skips with reasons, gateway restart outcome, fleet
snapshot) under ~/.hermes/logs/update_receipts/ with a latest.json
pointer for the dashboard/desktop; plus collect_fleet_versions() /
print_fleet_version_matrix() comparing every live profile gateway
against the freshly updated checkout.
- hermes_cli/update_cmd.py: wires receipt begin/steps/finalize into the
git, ZIP, and hard-failure paths; after the restart phase, prints the
fleet version matrix and escalates provably-stale gateways into the
existing gateway_fleet_restart_incomplete exit-1 contract. Pre-stamp
gateways report 'unknown' and never fail the update (no false
positives during rollout).
Silent-failure classes made visible: #88848, #74973, #85753, #81193.
Mixed-version fleet classes made loud: #88654, #69754, #77553, #56717.