18 Commits

Author SHA1 Message Date
teknium1 23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
teknium1 2fcb50c50e test(update): trim marker-discharge tests to two invariants; document the identity rule
The five #105417 tests collapse into one positive (fleet current at expected_sha
discharges the marker) and one parametrized negative (stale row / empty probe /
unknown identity / checkout moved past the marker keep the warning). Both are red
on origin/main. collect_fleet_versions' docstring now names why a live PID alone
is not a `current` row (#110420): write_runtime_status re-stamps pid/code_sha for
whatever process writes it, so the state-file fallback must pass
live_gateway_pid_for_home like the inventory already does (#109680).
2026-09-15 04:06:08 -07:00
KoNit-K dab6fe2933 fix(update): drop untrusted fleet fallback version
When the state-file writer cannot be verified as the profile gateway, null code_version with the SHA so the row does not display a self-reported identity claim.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-15 04:06:08 -07:00
KoNit-K 0eaa51674f fix(update): verify fleet state fallback gateway identity 2026-09-15 04:06:08 -07:00
tachyon-r 6e0fba4e58 refactor(update): share command-boundary receipt reason 2026-09-06 14:53:33 +05:30
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium d9d1858c50 refactor(hermes_cli/update): fold code-identity probes into one helper; fleet matrix row table 2026-09-02 23:59:44 -07:00
Teknium f36355859b refactor(hermes_cli/update): share profile-home + socket-identity probes between receipt and inventory; fold refusal builder; collapse try/except-pass 2026-09-02 21:49:26 -07:00
Teknium 2c7f5d12b5 refactor(hclib): cron/update/install lifecycle — cron status, update_* receipts and recovery, install repair, service manager 2026-09-02 14:45:25 -07:00
joaomarcos f9fec169b5 fix(update): recover hermes-serve units after an aborted restart phase
The fresh-process recovery boundary added for #92145 only reaches gateway
profiles. `hermes serve` -- the runtime that hosts `tui_gateway.server`,
and the process the original report saw failing every chat turn -- is not a
gateway profile, so no `gateway restart` command can reach it and the
gateway-only `collect_fleet_versions` read-back cannot see it either.

The spawn-ledger collector classifies serve/dashboard runtimes purely by
spawner liveness, and a systemd-launched `hermes serve` sets neither
HERMES_SPAWN nor HERMES_PARENT_PID, so it is recorded as `manual-serve`
and the recovery partition skips it as unrecoverable. The result is an
update that clears its incomplete flag on gateway coverage alone while a
live serve process keeps serving the pre-update module graph.

- restart active `hermes-serve*` systemd units from the fresh child,
  enumerated from systemd rather than from the misclassifying inventory,
  and verify a changed MainPID on an active unit before claiming coverage;
- report any pre-update serve/dashboard process that is still the same
  process, and never kill one -- a manual or Desktop-owned serve has no
  relaunch authority;
- require every runtime family, not just the gateway leg, before a
  fresh-process recovery may clear the incomplete flag;
- persist serve-unit outcomes and surviving runtimes in the update receipt.
2026-09-01 07:00:54 -07:00
Teknium b2c011364e fix(update): conservative outcomes + serve-ledger coverage for fresh restart recovery
Salvage adjustments to PR #94392 per review:

- Narrow the supervisor claim to the systemd-VERIFIED path only. The fresh
  recovery child now probes 'systemctl --user is-active' after each relaunch;
  only an observed-active systemd unit is reported 'verified'. A relaunch that
  merely exited 0 is labelled 'relaunch_attempted', never counts as supervisor
  coverage, and never clears gateway_fleet_restart_incomplete.
- Serve-owned runtimes (serve/dashboard entries from the spawn ledger, per the
  update_inventory serve collector) are no longer silently skipped: the
  recovery pass records them (and manual gateways) as skipped-with-reason in
  the recovery result and the persisted update receipt.
- Receipt fresh_recovery persists the conservative vocabulary
  (requested/verified/relaunch_attempted/failed/skipped); 'succeeded' is gone.
- Added an end-to-end test that drives the real recovery module in a genuinely
  fresh interpreter (sitecustomize shim intercepts the grandchild
  'gateway restart' and systemctl probes).
2026-08-26 16:45:26 -07:00
joaomarcos f9135c189c fix(update): persist fresh recovery outcome 2026-08-26 16:45:26 -07:00
chelsealong 2812d6121b fix(cli): note pre_restart_pids' per-PID data model gap, pin the matching-start_time path
Addresses the two follow-up notes from review: document that
pre_restart_pids is a bare PID set (not (pid, start_time) pairs), so a
recycled PID from one gateway landing in another's stale record could
still mislabel it as down; and add a companion test asserting a
matching start_time still yields the live/current row.
2026-08-26 16:14:32 -07:00
chelsealong 5d7ed70eef fix(cli): guard the post-update fleet check against PID reuse
collect_fleet_versions()'s gateway_state.json fallback path only checked
_pid_exists(pid) to decide whether a recorded gateway was still running.
On Windows, a paused gateway's PID can be recycled by an unrelated
process spawned during the update's own churn (npm/git/python
subprocesses) before the record is refreshed, so the dead gateway's
stale code_sha still gets compared against HEAD and reported STALE for a
PID that no longer belongs to it (#93258).

Switch to runtime_status_pid_is_live(), the existing (pid, start_time)
PID-reuse guard already used elsewhere in gateway/status.py, so a
recycled PID is treated the same as a dead one (DOWN row, or no row, per
the existing rollout-safety rules) instead of a false STALE.
2026-08-26 16:14:32 -07:00
Teknium 1684877868 fix(update): a gateway killed by the restart phase and never replaced now fails the fleet check (DOWN row)
Phase-1 verification gap (#91277, found auditing our own landed matrix
against the mapped issues): collect_fleet_versions only listed gateways
with a LIVE pid, so 'restart stopped it and nothing came back' produced
NO row at all — the exact silent-failure shape the matrix exists to
catch (#88848/#74973 class) passed with exit 0.

- collect_fleet_versions(pre_restart_pids=...): a dead pid becomes a
  'down' row only when it was alive at update start AND its runtime
  status still claims a running state. Rollout-safe: no snapshot (old
  callers), clean stops, startup failures, and stale records from
  long-dead gateways keep the historical no-row behavior.
- print_fleet_version_matrix escalates on down rows like stale ones
  (exit 1) with the per-profile restart remediation.
- cmd_update passes its existing pre-restart PID snapshot.

Sabotage-verified (reverting the membership check fails the new test);
live-verified with a real spawned-then-killed process producing the
DOWN row and matrix escalation.
2026-08-22 23:46:06 -07:00
Teknium 60b6269142 feat(gateway): gateway-owned control socket — identify/status verbs, fleet consumers prefer it over scans (#92091 step 1)
The gateway now creates a local control socket at startup (Unix domain
socket at $HERMES_HOME/gateway.sock with a pointer-file fallback for
long paths; named pipe on Windows) and answers versioned JSON verbs:

- identify: pid, profile, hermes_home, code_sha/code_version (#91283
  stamps, now queryable live), self-declared supervisor kind, start_time
- status: the live runtime-status payload, answered by the process itself

Bound immediately after the PID-file O_EXCL claim (the moment the
process becomes the authoritative gateway for its HERMES_HOME), removed
on clean shutdown; a successor clears any stale socket on bind. Strictly
non-fatal: bind failure only means consumers use the old path.

Consumers migrated (observability only, scan layer demoted to fallback,
never deleted):
- collect_fleet_versions() (post-update fleet matrix): prefers a live
  identify answer over gateway_state.json; entries carry source=socket
- collect_runtime_inventory() (hermes update --plan): prefers the
  socket, and takes the gateway's own supervisor declaration instead of
  inferring it from PID scans

Old gateways mid-upgrade, crashed processes, and bind failures behave
exactly as before. Never a TCP port; filesystem/pipe ACLs are the auth
boundary (0600 socket).

Part of #91277 (fleet-update reliability). Design: #92091.
2026-08-22 16:45:00 -07:00
Teknium 40b4a3bfe1 fix(update): every begun update receipt is now persisted — command-boundary finalization
Review on #91283: begin_update_receipt() fires early in _cmd_update_impl,
but finalization only existed on the success/ZIP/CalledProcessError
paths. Early sys.exit paths (Windows concurrent-instance preflight,
venv-holder refusal, head-pinned no-op, fetch failure) terminated with
the receipt started but never written — losing exactly the refused/
failed runs the receipt matters most for.

- update_receipt.py: finalize_pending_update_receipt(exit_code,
  stop_reason) — boundary safety net; maps exit 2 → 'refused', other
  non-zero → 'failed'; records exit_code + stop_reason. Exactly-once by
  construction (singleton popped in finalize_update_receipt), so runs
  the inner paths already finalized are untouched.
- main.py cmd_update: SystemExit/BaseException/else arms around
  _cmd_update_impl persist any still-open receipt with the real exit
  code, then re-raise unchanged. Future early exits are covered without
  per-site finalize patches.
- 5 regression tests incl. end-to-end through the real cmd_update
  wrapper (exit 2 preserved, outcome 'refused', stop reason recorded,
  singleton cleared, exactly one receipt file).
2026-08-21 03:44:55 -07:00
Teknium 1d74833d8d feat(update): structured update receipts + post-update fleet version verification
Phase 1 of the fleet-update reliability plan (#91277): the updater now
proves its outcome instead of assuming it.

- hermes_cli/build_info.py: get_code_identity() — process-cached code
  identity (git sha for source installs, baked .hermes_build_sha for
  Docker images, pyproject version).
- gateway/status.py: every runtime-status write stamps the writer's
  code_sha/code_version into gateway_state.json, so a running gateway's
  actual code generation is observable from disk.
- hermes_cli/update_receipt.py (new): machine-readable receipt of each
  update run (steps, skips with reasons, gateway restart outcome, fleet
  snapshot) under ~/.hermes/logs/update_receipts/ with a latest.json
  pointer for the dashboard/desktop; plus collect_fleet_versions() /
  print_fleet_version_matrix() comparing every live profile gateway
  against the freshly updated checkout.
- hermes_cli/update_cmd.py: wires receipt begin/steps/finalize into the
  git, ZIP, and hard-failure paths; after the restart phase, prints the
  fleet version matrix and escalates provably-stale gateways into the
  existing gateway_fleet_restart_incomplete exit-1 contract. Pre-stamp
  gateways report 'unknown' and never fail the update (no false
  positives during rollout).

Silent-failure classes made visible: #88848, #74973, #85753, #81193.
Mixed-version fleet classes made loud: #88654, #69754, #77553, #56717.
2026-08-20 21:54:25 -07:00