Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
The five #105417 tests collapse into one positive (fleet current at expected_sha
discharges the marker) and one parametrized negative (stale row / empty probe /
unknown identity / checkout moved past the marker keep the warning). Both are red
on origin/main. collect_fleet_versions' docstring now names why a live PID alone
is not a `current` row (#110420): write_runtime_status re-stamps pid/code_sha for
whatever process writes it, so the state-file fallback must pass
live_gateway_pid_for_home like the inventory already does (#109680).
When the state-file writer cannot be verified as the profile gateway, null code_version with the SHA so the row does not display a self-reported identity claim.
Co-authored-by: Cursor <cursoragent@cursor.com>
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
The fresh-process recovery boundary added for #92145 only reaches gateway
profiles. `hermes serve` -- the runtime that hosts `tui_gateway.server`,
and the process the original report saw failing every chat turn -- is not a
gateway profile, so no `gateway restart` command can reach it and the
gateway-only `collect_fleet_versions` read-back cannot see it either.
The spawn-ledger collector classifies serve/dashboard runtimes purely by
spawner liveness, and a systemd-launched `hermes serve` sets neither
HERMES_SPAWN nor HERMES_PARENT_PID, so it is recorded as `manual-serve`
and the recovery partition skips it as unrecoverable. The result is an
update that clears its incomplete flag on gateway coverage alone while a
live serve process keeps serving the pre-update module graph.
- restart active `hermes-serve*` systemd units from the fresh child,
enumerated from systemd rather than from the misclassifying inventory,
and verify a changed MainPID on an active unit before claiming coverage;
- report any pre-update serve/dashboard process that is still the same
process, and never kill one -- a manual or Desktop-owned serve has no
relaunch authority;
- require every runtime family, not just the gateway leg, before a
fresh-process recovery may clear the incomplete flag;
- persist serve-unit outcomes and surviving runtimes in the update receipt.
Salvage adjustments to PR #94392 per review:
- Narrow the supervisor claim to the systemd-VERIFIED path only. The fresh
recovery child now probes 'systemctl --user is-active' after each relaunch;
only an observed-active systemd unit is reported 'verified'. A relaunch that
merely exited 0 is labelled 'relaunch_attempted', never counts as supervisor
coverage, and never clears gateway_fleet_restart_incomplete.
- Serve-owned runtimes (serve/dashboard entries from the spawn ledger, per the
update_inventory serve collector) are no longer silently skipped: the
recovery pass records them (and manual gateways) as skipped-with-reason in
the recovery result and the persisted update receipt.
- Receipt fresh_recovery persists the conservative vocabulary
(requested/verified/relaunch_attempted/failed/skipped); 'succeeded' is gone.
- Added an end-to-end test that drives the real recovery module in a genuinely
fresh interpreter (sitecustomize shim intercepts the grandchild
'gateway restart' and systemctl probes).
Addresses the two follow-up notes from review: document that
pre_restart_pids is a bare PID set (not (pid, start_time) pairs), so a
recycled PID from one gateway landing in another's stale record could
still mislabel it as down; and add a companion test asserting a
matching start_time still yields the live/current row.
collect_fleet_versions()'s gateway_state.json fallback path only checked
_pid_exists(pid) to decide whether a recorded gateway was still running.
On Windows, a paused gateway's PID can be recycled by an unrelated
process spawned during the update's own churn (npm/git/python
subprocesses) before the record is refreshed, so the dead gateway's
stale code_sha still gets compared against HEAD and reported STALE for a
PID that no longer belongs to it (#93258).
Switch to runtime_status_pid_is_live(), the existing (pid, start_time)
PID-reuse guard already used elsewhere in gateway/status.py, so a
recycled PID is treated the same as a dead one (DOWN row, or no row, per
the existing rollout-safety rules) instead of a false STALE.
Phase-1 verification gap (#91277, found auditing our own landed matrix
against the mapped issues): collect_fleet_versions only listed gateways
with a LIVE pid, so 'restart stopped it and nothing came back' produced
NO row at all — the exact silent-failure shape the matrix exists to
catch (#88848/#74973 class) passed with exit 0.
- collect_fleet_versions(pre_restart_pids=...): a dead pid becomes a
'down' row only when it was alive at update start AND its runtime
status still claims a running state. Rollout-safe: no snapshot (old
callers), clean stops, startup failures, and stale records from
long-dead gateways keep the historical no-row behavior.
- print_fleet_version_matrix escalates on down rows like stale ones
(exit 1) with the per-profile restart remediation.
- cmd_update passes its existing pre-restart PID snapshot.
Sabotage-verified (reverting the membership check fails the new test);
live-verified with a real spawned-then-killed process producing the
DOWN row and matrix escalation.
The gateway now creates a local control socket at startup (Unix domain
socket at $HERMES_HOME/gateway.sock with a pointer-file fallback for
long paths; named pipe on Windows) and answers versioned JSON verbs:
- identify: pid, profile, hermes_home, code_sha/code_version (#91283
stamps, now queryable live), self-declared supervisor kind, start_time
- status: the live runtime-status payload, answered by the process itself
Bound immediately after the PID-file O_EXCL claim (the moment the
process becomes the authoritative gateway for its HERMES_HOME), removed
on clean shutdown; a successor clears any stale socket on bind. Strictly
non-fatal: bind failure only means consumers use the old path.
Consumers migrated (observability only, scan layer demoted to fallback,
never deleted):
- collect_fleet_versions() (post-update fleet matrix): prefers a live
identify answer over gateway_state.json; entries carry source=socket
- collect_runtime_inventory() (hermes update --plan): prefers the
socket, and takes the gateway's own supervisor declaration instead of
inferring it from PID scans
Old gateways mid-upgrade, crashed processes, and bind failures behave
exactly as before. Never a TCP port; filesystem/pipe ACLs are the auth
boundary (0600 socket).
Part of #91277 (fleet-update reliability). Design: #92091.
Review on #91283: begin_update_receipt() fires early in _cmd_update_impl,
but finalization only existed on the success/ZIP/CalledProcessError
paths. Early sys.exit paths (Windows concurrent-instance preflight,
venv-holder refusal, head-pinned no-op, fetch failure) terminated with
the receipt started but never written — losing exactly the refused/
failed runs the receipt matters most for.
- update_receipt.py: finalize_pending_update_receipt(exit_code,
stop_reason) — boundary safety net; maps exit 2 → 'refused', other
non-zero → 'failed'; records exit_code + stop_reason. Exactly-once by
construction (singleton popped in finalize_update_receipt), so runs
the inner paths already finalized are untouched.
- main.py cmd_update: SystemExit/BaseException/else arms around
_cmd_update_impl persist any still-open receipt with the real exit
code, then re-raise unchanged. Future early exits are covered without
per-site finalize patches.
- 5 regression tests incl. end-to-end through the real cmd_update
wrapper (exit 2 preserved, outcome 'refused', stop reason recorded,
singleton cleared, exactly one receipt file).
Phase 1 of the fleet-update reliability plan (#91277): the updater now
proves its outcome instead of assuming it.
- hermes_cli/build_info.py: get_code_identity() — process-cached code
identity (git sha for source installs, baked .hermes_build_sha for
Docker images, pyproject version).
- gateway/status.py: every runtime-status write stamps the writer's
code_sha/code_version into gateway_state.json, so a running gateway's
actual code generation is observable from disk.
- hermes_cli/update_receipt.py (new): machine-readable receipt of each
update run (steps, skips with reasons, gateway restart outcome, fleet
snapshot) under ~/.hermes/logs/update_receipts/ with a latest.json
pointer for the dashboard/desktop; plus collect_fleet_versions() /
print_fleet_version_matrix() comparing every live profile gateway
against the freshly updated checkout.
- hermes_cli/update_cmd.py: wires receipt begin/steps/finalize into the
git, ZIP, and hard-failure paths; after the restart phase, prints the
fleet version matrix and escalates provably-stale gateways into the
existing gateway_fleet_restart_incomplete exit-1 contract. Pre-stamp
gateways report 'unknown' and never fail the update (no false
positives during rollout).
Silent-failure classes made visible: #88848, #74973, #85753, #81193.
Mixed-version fleet classes made loud: #88654, #69754, #77553, #56717.