34 Commits

Author SHA1 Message Date
teknium1 2246c245f5 refactor(update): single defer flag for the deferred catch-up; trim tests; document the flag
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
2026-09-15 19:28:39 -07:00
Turgut Kural 7b27ea3639 fix(cli): allow hermes update without gateway restart for cron (rebased on upstream/main)
(cherry picked from commit e70f78e54a96f2e8037f7e385fc31563bbeaf392)
(cherry picked from commit 225f56ab29e91977f1fef2e743c50dd6e2e89b60)
2026-09-15 19:28:39 -07:00
teknium1 7c6b8c7a8d fix(update): Desktop-owned serve no longer fails the update or holds fleet_restart_pending
`hermes update` run while the Desktop app is open ended `partial`/exit 1 and re-armed
`fleet_restart_pending` on every run: `_gateway_recovery_partition` exempts a
`supervisor == "desktop"` serve from restart (`_DESKTOP_SERVE_SKIP_REASON` — it hosts the
live Desktop chats), while `match_runtime_outcomes` counted that same still-alive process
as `unaccounted` whenever the survivor probe found its pre-update incarnation in the ledger.
Nothing in the updater is allowed to discharge that obligation, so it could never finalize.

- update_inventory.match_runtime_outcomes: a Desktop-supervised serve/dashboard still alive
  reconciles as a new outcome `deferred` (handed back to its supervisor). "restarted" would
  be untrue — the process provably runs old code and the Desktop app does not respawn it after
  a terminal-side update. A gone one stays `restarted`; a manual/systemd survivor stays
  `unaccounted`.
- update_inventory.report_unaccounted_runtimes: prints `deferred` rows with the one remedy
  that exists (relaunch the Desktop app) without escalating; the `systemctl --user restart
  hermes-serve.service` hint is Linux-only now (it was shown on macOS too).
- update_abort_recovery: same class on the fresh-child recovery path — `_owed_stale_serve_rows`
  excludes Desktop-owned survivors from `_abort_recovery_is_complete` and the incomplete
  gate in update_cmd_fleet; they are still named by `_warn_stale_serve_runtimes` and kept
  in the receipt's `stale_runtimes`.
- tests: end-to-end (exit 0, receipt `success`, `runtime_outcomes` gateway=restarted /
  serve=deferred, marker cleared, relaunch hint printed) + abort-path predicate; both red on
  origin/main.

Fixes #111494
Supersedes #111499
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:30:32 -07:00
teknium1 f81c33cb6a fix(update): stop the fleet settle poll early when the restarted unit is dead
Salvage of #111385 (@JoaoMarcos44): the 30s -> 120s settle window is kept so a
slow host's gateway can publish its state stamp. This commit bounds the other
side of that trade: when every restarted systemd unit reports neither active
nor activating, the successor has died and nothing will ever publish, so the
poll fails closed at once instead of spending the full 120s. Unknown states
(no units, systemctl missing or slow) keep waiting. Also drops the test
assertion pinning the 2s poll cadence.
2026-09-15 15:10:41 -07:00
joaomarcos c99fee4e0a fix(update): wait for fleet state publication after supervised restart
A systemd unit can be active before the replacement gateway finishes
bootstrap and publishes gateway_state.json. The 30s poll in
_collect_fleet_snapshot() then returned [] with rows_expected=true,
causing _verify_fleet_after_update() to mark the restart incomplete and
exit(1) before clearing fleet_restart_pending. Every later CLI
startup/doctor therefore printed a false "did not restart" warning even
though the live gateway already served expected_sha.

Allow the default systemd startup budget plus publication slack by
extending the bounded settle window to 120s. Keeps fail-closed on
stale/down/empty after the deadline, only widens the window for slow
bootstraps (e.g. Raspberry Pi).

Fixes #111272
2026-09-15 15:10:41 -07:00
teknium1 b2577df807 fix(dashboard): accept command-scoped NOPASSWD sudo for system gateway actions
The dashboard's system-scope elevation gate hard-failed on a refused
`sudo -n true`, but the `hermes update` fleet restart it claims to
mirror treats that blanket probe as inconclusive and falls back to
probing the targeted command. A host with a sudoers entry scoped to
the hermes command (the hardened shape) therefore updated fine from
the CLI while the dashboard reported "passwordless sudo is
unavailable".

Factor the fleet's two-step probe into update_cmd_fleet
._sudo_noninteractive_ok and call it from both sites: the fleet keeps
its `reset-failed <unit>` fallback, the dashboard falls back to a
non-destructive `sudo -n -l -- <exact argv>` check before spawning.

Review finding: dashboard sudo gate diverged from the fleet posture it
borrowed (_needs_sudo only) and rejected targeted NOPASSWD sudoers.
2026-09-15 04:08:00 -07:00
teknium1 127214a66b fix: keep fleet-restart marker while a receipt-owed gateway is down
_marker_only_restart_obsolete cleared the marker when every row from
collect_fleet_versions() was current. At CLI startup the probe runs without
pre_restart_pids, so a gateway the restart phase stopped and never brought
back produces no row at all (dead-pid records are skipped, no DOWN
classification possible). With alpha current and beta down the marker was
discharged and beta's catch-up restart never happened.

Before clearing, also require every gateway identity latest.json owes
(plan.runtimes / fleet entries) to be covered by a current row; the
rows-only rule applies only when the receipt names no gateways. The
identity extraction is shared with _live_fleet_covers_receipt.

Review finding: marker discharged with a DOWN sibling gateway absent from the startup probe.
2026-09-15 04:06:08 -07:00
linmukong 8130274b37 fix(update): discharge fleet_restart_pending when the fleet provably serves expected_sha
_pending_fleet_restart_needed() returns True on marker existence alone. A
supervisor-level gateway restart (`systemctl --user restart`, launchctl, an ops
script) never passes through _clear_fleet_restart_pending_marker(), so a marker
written by a pulled update survives a restart that DID bring every live gateway
to the new code — and every later CLI call prints the interrupted-update warning
forever, a permanent false positive that trains operators to ignore it.

The comment on that branch is right that "an older receipt cannot discharge that
unknown obligation" — but the obligation is not unknown. The marker records its
own expected_sha, so it can be verified against the live fleet directly, with no
receipt involved. _live_fleet_covers_receipt() cannot substitute: it is
receipt-anchored and returns False when `owed` is empty, which is exactly the
marker-only case.

Hold the marker to the same evidence bar _live_fleet_covers_receipt() applies to
a receipt: at least one row, and every row a `current` gateway under a known
profile whose code_sha equals the marker's expected_sha, with the checkout HEAD
not moved past the marker.

Still keeps the marker (warns) on:
  - stale / down rows — the restart genuinely is owed
  - an all-`unknown` fleet — pre-code-identity gateways cannot prove currency
    (same conservatism as #88848/#74973)
  - a marker with no expected_sha — nothing to verify against
  - a newer pull that moved the checkout — it owns a fresh obligation
  - a probe that raises or answers empty — no proof either way

Discharging deletes the marker only after every row passes; the historical
receipt is left intact, since a supervisor restart is still not a successful
update.

Tests: discharge on a provably-current fleet; keep on stale, empty probe,
unknown identity, and newer-pull-moved-checkout (which asserts the fleet is
never probed once HEAD has moved).
2026-09-15 04:06:08 -07:00
teknium1 6a312fba54 feat(update): name the work a draining gateway is waiting on
`hermes update` printed "draining (up to 1875s)..." and then nothing for up
to 30 minutes while the gateway's in-band restart waited on in-flight work
(agent.restart_after_turn_timeout). Neither the updater nor the gateway log
said WHAT was being waited on, so a single long cron job read as a hung
update.

Gateway side: GatewayShutdownMixin._describe_active_work() enumerates each
unit the restart wait holds for — chat turns (session key, model, current
tool, elapsed), cron jobs (job id, elapsed, and the restart-safe external
worker pid when the run was handed off; cron/scheduler now records that pid
next to the running id), api/deferred runs by count. It is written to
gateway_state.json as `active_work` while the state is `draining` (cleared
otherwise) and appended to the 30s "Restart deferred" log line.

CLI side: hermes_cli/update_cmd_drain_report.py reads `active_work` and
prints a progress block every 30s during the SIGUSR1 exit wait — the
holder(s), their pids, elapsed time, seconds left before the forced
restart, and the config knob that caps the wait. Wired into the systemd,
launchd and manual gateway restart paths of `hermes update` and into
`hermes gateway restart`; `hermes gateway status` lists the same units
while draining. A pre-fix gateway (no `active_work` field) gets an explicit
"gateway did not report" line rather than silence.

Live A/B (real gateway, 90s no-agent cron job in flight, SIGUSR1 from the
caller): base = 79s of silence, no `active_work` in the state file; head =
the job named with pid/elapsed/remaining every interval, log line carries
the same detail.
2026-09-13 05:08:20 -07:00
Teknium 07a4ae016a feat(update): auto-migrate to one multiplexed gateway when unblocked
After the fleet restart is verified healthy, `hermes update` runs the
migration preflight on installs with >= 2 profiles and at least one
per-profile gateway. No blockers: migrate (same path as
`gateway migrate --multiplex --yes`, deterministic, never prompts).
Blockers: print them with their fixes and the one-liner, change nothing.
Skipped on the exit-1 (stale fleet) path and on single-profile installs.
2026-09-12 01:49:28 -07:00
Teknium 89c85b8466 fix(update): settle stale receipt warnings from matching live gateways
Reconcile receipt-only restart obligations at the shared warning/catch-up
predicate, requiring every historical runtime/profile identity to have a
current live gateway successor. Preserve missing and unknown obligations,
non-gateway identities, and independently authoritative pending markers.

Keep failed receipts unchanged instead of recording an unverified success.
Live isolated two-process A/B reproduces the warning on base and settles
it after the fix; stale, unknown, and missing-profile controls still warn.

Reported-by: duanzhiwei0315
Inspired-by: zengzheqing (#104295), RootZ3n (#100249)
2026-09-07 08:20:46 -07:00
Teknium 08f2c78d92 fix(update): cover catch-up restart clients with the unit budget 2026-09-07 08:20:09 -07:00
doryani-agent 2a980fbcbd fix(update): let systemd clients outwait legitimate unit transactions
Salvage the unit-budget implementation from #104745, replacing its test
matrix with two invariant tests and covering the sibling graceful start.
Keep unprivileged property reads, finite fallbacks, real manager errors,
and post-restart health verification.

Native disposable user unit: old client timed out after 15.03 seconds;
new client completed the same 16-second stop transaction in 16.13 seconds.
The unit stayed active with a new PID; missing-unit errors stayed errors.

Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
2026-09-07 08:20:09 -07:00
Teknium 7798241eab fix: retain pending fleet restarts until supervisors recover
Discover systemd targets before stopping old processes, restart even when
there are no gateway PIDs, and require successful scope listings plus active
verification. Pending launchd recovery also retains failures for inaccessible
listings and installed jobs without supervision. Keep existing PID cleanup
intact but before recovery so it cannot kill freshly verified workers.

Slim redo informed by #104274, #104283, and #104285.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-07 06:10:49 -07:00
kshitijk4poor a262b2e372 refactor(update): drop the tombstone comment at the old matcher site and a duplicated exactness leg 2026-09-06 20:47:52 +05:30
mengtanx b4b6235239 fix(update): credit launchd ai.hermes.gateway in fleet reconciliation (#103679)
The restart phase records macOS LaunchAgent labels (ai.hermes.gateway).
match_runtime_outcomes used a substring check for "hermes-gateway", so a
successful Desktop update on the default profile always tripped
"Planned runtimes the restart phase never touched" and exited 1.

Use the exact systemd/launchd/s6 matcher for both plan reconciliation
and abort-recovery so the two cannot drift.
2026-09-06 20:47:52 +05:30
kshitijk4poor bc1330eebc refactor(update): name the success invariant in _receipt_looks_unfinished; one predicate contract
The tail clause was correct only by ordering (exit_code != 0 there meant
exit_code is None). Name what it encodes: a stop_reason counts only when
nothing vouched for success. Same truth table. The three literal-dict tests on
the predicate collapse into one parametrized contract; the handoff-exit test
binds to COMMAND_BOUNDARY_STOP_REASON instead of re-spelling it.
2026-09-06 14:53:33 +05:30
kshitijk4poor dad698d88c test(update): pin the refused-receipt shape the stop_reason clause exists for
update_contract writes {"outcome": "refused", "stop_reason": <code>} with no
exit_code; that is the one production receipt where the stop_reason clause in
_receipt_looks_unfinished is load-bearing. The previous negative control used
exit_code=1, which the exit_code branch already catches. Docstring reworded:
a KeyboardInterrupt never lands on a success receipt (the boundary finalize is
a no-op once the inner path finalized).
2026-09-06 14:53:33 +05:30
tachyon-r cd27df3c2e fix(update): ignore successful receipt stop reason 2026-09-06 14:53:33 +05:30
Teknium 1cce7c6dd8 fix(update): unmapped gateway stops no longer fail the fleet check with "no rows"
A gateway with no profile mapping (or one whose relaunch could not be armed) is
SIGTERMed and listed under "Restart manually" — by design it has no successor and
publishes no fleet-matrix row. It still counted in ``killed_pids`` and the
pre-restart snapshot, so ``_fleet_probe_expected_runtimes`` demanded rows that
could not exist and a fully successful update exited 1 with "Fleet version check
returned no rows even though gateway runtimes were expected", leaving the
fleet_restart_pending marker behind and every later CLI start warning about it.

Track the unmapped stops on the restart outcome and subtract them from the
row-predicting signals (``fleet_probe_signals``); relaunched/systemd gateways
still predict rows exactly as before.
2026-09-05 00:24:40 -07:00
liuhao1024 b8e3c5c700 fix(update): scope fleet-probe runtime expectation to gateway-kind plan records
A dashboard-only runtime plan (gateway never started) made
_fleet_probe_expected_runtimes() return True from the unfiltered
'plan.runtimes is non-empty' check. collect_fleet_versions() reports
gateway identities only, so the probe waited for rows that cannot exist,
printed the incomplete-verification warning, and exited 1 after a
successful update (#97332).

Key the plan-derived expectation on kind == 'gateway' records — the same
row-capability rule already applied to the Windows resume token (#93406)
— and update the two tests that pinned the old object() placeholder so
they pin the runtime-kind distinction. Restart-phase, killed-PID, and
pre-restart-PID signals still fail closed unchanged.
2026-09-05 00:24:40 -07:00
Teknium 474eed838e review-fix(suppress-audit): patch_parser/skills_sync/update_cmd_fleet/api_server — restore BASE exception semantics 2026-09-03 09:56:39 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 4b9b9a441f refactor(hermes_cli/update): pack flat import/argument lists (AST-identical) 2026-09-02 22:14:59 -07:00
Teknium d9941ee4ce refactor(hermes_cli/update): reflow to 120 cols, drop blanks after local imports (AST-identical) 2026-09-02 21:45:19 -07:00
Teknium 90052abfb9 refactor(hermes_cli/update): reflow short multi-line calls and signatures (AST-identical) 2026-09-02 21:30:17 -07:00
Teknium e64942a018 refactor(hermes_cli/update): compact fleet restart bookkeeping into outcome dataclass; split maint backup/notice phases 2026-09-02 21:18:02 -07:00
Teknium ce45e216e7 refactor(update): promote nested _restart_one_systemd_gateway_unit closure to a module-level function (184 -> 70 LOC orchestrator) 2026-09-02 16:55:53 -07:00
Teknium 60041b787a refactor(update): join short multi-line statements onto one line (AST-identical, -416 lines) 2026-09-02 16:52:38 -07:00
Teknium 1aa9285312 refactor(update): hand-compact comments/docstrings in update_cmd.py, update_cmd_fleet.py, update_cmd_zip.py (AST-identical); rewrite stale module docstring 2026-09-02 16:45:44 -07:00
Teknium d46eb1f964 refactor(update): collapse 90 try/except-pass and try/except-logger.debug blocks into suppress()/_best_effort() (new leaf update_cmd_common.py) 2026-09-02 16:34:55 -07:00
Teknium e118abc61a refactor(update): lift manual-gateway restart, stuck-PID sweep and phase-abort recovery out of _restart_gateway_fleet_after_update (362 -> 154 LOC) 2026-09-02 16:23:31 -07:00
Teknium 4674f8b254 refactor(update): split zip/stash/config/deps/git/maint clusters out of update_cmd.py 2026-09-02 16:00:26 -07:00
Teknium 096826bf7d refactor(update): split gateway fleet restart/verify into update_cmd_fleet.py 2026-09-02 15:55:50 -07:00