Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
`hermes update` run while the Desktop app is open ended `partial`/exit 1 and re-armed
`fleet_restart_pending` on every run: `_gateway_recovery_partition` exempts a
`supervisor == "desktop"` serve from restart (`_DESKTOP_SERVE_SKIP_REASON` — it hosts the
live Desktop chats), while `match_runtime_outcomes` counted that same still-alive process
as `unaccounted` whenever the survivor probe found its pre-update incarnation in the ledger.
Nothing in the updater is allowed to discharge that obligation, so it could never finalize.
- update_inventory.match_runtime_outcomes: a Desktop-supervised serve/dashboard still alive
reconciles as a new outcome `deferred` (handed back to its supervisor). "restarted" would
be untrue — the process provably runs old code and the Desktop app does not respawn it after
a terminal-side update. A gone one stays `restarted`; a manual/systemd survivor stays
`unaccounted`.
- update_inventory.report_unaccounted_runtimes: prints `deferred` rows with the one remedy
that exists (relaunch the Desktop app) without escalating; the `systemctl --user restart
hermes-serve.service` hint is Linux-only now (it was shown on macOS too).
- update_abort_recovery: same class on the fresh-child recovery path — `_owed_stale_serve_rows`
excludes Desktop-owned survivors from `_abort_recovery_is_complete` and the incomplete
gate in update_cmd_fleet; they are still named by `_warn_stale_serve_runtimes` and kept
in the receipt's `stale_runtimes`.
- tests: end-to-end (exit 0, receipt `success`, `runtime_outcomes` gateway=restarted /
serve=deferred, marker cleared, relaunch hint printed) + abort-path predicate; both red on
origin/main.
Fixes#111494
Supersedes #111499
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Salvage of #111385 (@JoaoMarcos44): the 30s -> 120s settle window is kept so a
slow host's gateway can publish its state stamp. This commit bounds the other
side of that trade: when every restarted systemd unit reports neither active
nor activating, the successor has died and nothing will ever publish, so the
poll fails closed at once instead of spending the full 120s. Unknown states
(no units, systemctl missing or slow) keep waiting. Also drops the test
assertion pinning the 2s poll cadence.
A systemd unit can be active before the replacement gateway finishes
bootstrap and publishes gateway_state.json. The 30s poll in
_collect_fleet_snapshot() then returned [] with rows_expected=true,
causing _verify_fleet_after_update() to mark the restart incomplete and
exit(1) before clearing fleet_restart_pending. Every later CLI
startup/doctor therefore printed a false "did not restart" warning even
though the live gateway already served expected_sha.
Allow the default systemd startup budget plus publication slack by
extending the bounded settle window to 120s. Keeps fail-closed on
stale/down/empty after the deadline, only widens the window for slow
bootstraps (e.g. Raspberry Pi).
Fixes#111272
The dashboard's system-scope elevation gate hard-failed on a refused
`sudo -n true`, but the `hermes update` fleet restart it claims to
mirror treats that blanket probe as inconclusive and falls back to
probing the targeted command. A host with a sudoers entry scoped to
the hermes command (the hardened shape) therefore updated fine from
the CLI while the dashboard reported "passwordless sudo is
unavailable".
Factor the fleet's two-step probe into update_cmd_fleet
._sudo_noninteractive_ok and call it from both sites: the fleet keeps
its `reset-failed <unit>` fallback, the dashboard falls back to a
non-destructive `sudo -n -l -- <exact argv>` check before spawning.
Review finding: dashboard sudo gate diverged from the fleet posture it
borrowed (_needs_sudo only) and rejected targeted NOPASSWD sudoers.
_marker_only_restart_obsolete cleared the marker when every row from
collect_fleet_versions() was current. At CLI startup the probe runs without
pre_restart_pids, so a gateway the restart phase stopped and never brought
back produces no row at all (dead-pid records are skipped, no DOWN
classification possible). With alpha current and beta down the marker was
discharged and beta's catch-up restart never happened.
Before clearing, also require every gateway identity latest.json owes
(plan.runtimes / fleet entries) to be covered by a current row; the
rows-only rule applies only when the receipt names no gateways. The
identity extraction is shared with _live_fleet_covers_receipt.
Review finding: marker discharged with a DOWN sibling gateway absent from the startup probe.
_pending_fleet_restart_needed() returns True on marker existence alone. A
supervisor-level gateway restart (`systemctl --user restart`, launchctl, an ops
script) never passes through _clear_fleet_restart_pending_marker(), so a marker
written by a pulled update survives a restart that DID bring every live gateway
to the new code — and every later CLI call prints the interrupted-update warning
forever, a permanent false positive that trains operators to ignore it.
The comment on that branch is right that "an older receipt cannot discharge that
unknown obligation" — but the obligation is not unknown. The marker records its
own expected_sha, so it can be verified against the live fleet directly, with no
receipt involved. _live_fleet_covers_receipt() cannot substitute: it is
receipt-anchored and returns False when `owed` is empty, which is exactly the
marker-only case.
Hold the marker to the same evidence bar _live_fleet_covers_receipt() applies to
a receipt: at least one row, and every row a `current` gateway under a known
profile whose code_sha equals the marker's expected_sha, with the checkout HEAD
not moved past the marker.
Still keeps the marker (warns) on:
- stale / down rows — the restart genuinely is owed
- an all-`unknown` fleet — pre-code-identity gateways cannot prove currency
(same conservatism as #88848/#74973)
- a marker with no expected_sha — nothing to verify against
- a newer pull that moved the checkout — it owns a fresh obligation
- a probe that raises or answers empty — no proof either way
Discharging deletes the marker only after every row passes; the historical
receipt is left intact, since a supervisor restart is still not a successful
update.
Tests: discharge on a provably-current fleet; keep on stale, empty probe,
unknown identity, and newer-pull-moved-checkout (which asserts the fleet is
never probed once HEAD has moved).
`hermes update` printed "draining (up to 1875s)..." and then nothing for up
to 30 minutes while the gateway's in-band restart waited on in-flight work
(agent.restart_after_turn_timeout). Neither the updater nor the gateway log
said WHAT was being waited on, so a single long cron job read as a hung
update.
Gateway side: GatewayShutdownMixin._describe_active_work() enumerates each
unit the restart wait holds for — chat turns (session key, model, current
tool, elapsed), cron jobs (job id, elapsed, and the restart-safe external
worker pid when the run was handed off; cron/scheduler now records that pid
next to the running id), api/deferred runs by count. It is written to
gateway_state.json as `active_work` while the state is `draining` (cleared
otherwise) and appended to the 30s "Restart deferred" log line.
CLI side: hermes_cli/update_cmd_drain_report.py reads `active_work` and
prints a progress block every 30s during the SIGUSR1 exit wait — the
holder(s), their pids, elapsed time, seconds left before the forced
restart, and the config knob that caps the wait. Wired into the systemd,
launchd and manual gateway restart paths of `hermes update` and into
`hermes gateway restart`; `hermes gateway status` lists the same units
while draining. A pre-fix gateway (no `active_work` field) gets an explicit
"gateway did not report" line rather than silence.
Live A/B (real gateway, 90s no-agent cron job in flight, SIGUSR1 from the
caller): base = 79s of silence, no `active_work` in the state file; head =
the job named with pid/elapsed/remaining every interval, log line carries
the same detail.
After the fleet restart is verified healthy, `hermes update` runs the
migration preflight on installs with >= 2 profiles and at least one
per-profile gateway. No blockers: migrate (same path as
`gateway migrate --multiplex --yes`, deterministic, never prompts).
Blockers: print them with their fixes and the one-liner, change nothing.
Skipped on the exit-1 (stale fleet) path and on single-profile installs.
Reconcile receipt-only restart obligations at the shared warning/catch-up
predicate, requiring every historical runtime/profile identity to have a
current live gateway successor. Preserve missing and unknown obligations,
non-gateway identities, and independently authoritative pending markers.
Keep failed receipts unchanged instead of recording an unverified success.
Live isolated two-process A/B reproduces the warning on base and settles
it after the fix; stale, unknown, and missing-profile controls still warn.
Reported-by: duanzhiwei0315
Inspired-by: zengzheqing (#104295), RootZ3n (#100249)
Salvage the unit-budget implementation from #104745, replacing its test
matrix with two invariant tests and covering the sibling graceful start.
Keep unprivileged property reads, finite fallbacks, real manager errors,
and post-restart health verification.
Native disposable user unit: old client timed out after 15.03 seconds;
new client completed the same 16-second stop transaction in 16.13 seconds.
The unit stayed active with a new PID; missing-unit errors stayed errors.
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
Discover systemd targets before stopping old processes, restart even when
there are no gateway PIDs, and require successful scope listings plus active
verification. Pending launchd recovery also retains failures for inaccessible
listings and installed jobs without supervision. Keep existing PID cleanup
intact but before recovery so it cannot kill freshly verified workers.
Slim redo informed by #104274, #104283, and #104285.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
The restart phase records macOS LaunchAgent labels (ai.hermes.gateway).
match_runtime_outcomes used a substring check for "hermes-gateway", so a
successful Desktop update on the default profile always tripped
"Planned runtimes the restart phase never touched" and exited 1.
Use the exact systemd/launchd/s6 matcher for both plan reconciliation
and abort-recovery so the two cannot drift.
The tail clause was correct only by ordering (exit_code != 0 there meant
exit_code is None). Name what it encodes: a stop_reason counts only when
nothing vouched for success. Same truth table. The three literal-dict tests on
the predicate collapse into one parametrized contract; the handoff-exit test
binds to COMMAND_BOUNDARY_STOP_REASON instead of re-spelling it.
update_contract writes {"outcome": "refused", "stop_reason": <code>} with no
exit_code; that is the one production receipt where the stop_reason clause in
_receipt_looks_unfinished is load-bearing. The previous negative control used
exit_code=1, which the exit_code branch already catches. Docstring reworded:
a KeyboardInterrupt never lands on a success receipt (the boundary finalize is
a no-op once the inner path finalized).
A gateway with no profile mapping (or one whose relaunch could not be armed) is
SIGTERMed and listed under "Restart manually" — by design it has no successor and
publishes no fleet-matrix row. It still counted in ``killed_pids`` and the
pre-restart snapshot, so ``_fleet_probe_expected_runtimes`` demanded rows that
could not exist and a fully successful update exited 1 with "Fleet version check
returned no rows even though gateway runtimes were expected", leaving the
fleet_restart_pending marker behind and every later CLI start warning about it.
Track the unmapped stops on the restart outcome and subtract them from the
row-predicting signals (``fleet_probe_signals``); relaunched/systemd gateways
still predict rows exactly as before.
A dashboard-only runtime plan (gateway never started) made
_fleet_probe_expected_runtimes() return True from the unfiltered
'plan.runtimes is non-empty' check. collect_fleet_versions() reports
gateway identities only, so the probe waited for rows that cannot exist,
printed the incomplete-verification warning, and exited 1 after a
successful update (#97332).
Key the plan-derived expectation on kind == 'gateway' records — the same
row-capability rule already applied to the Windows resume token (#93406)
— and update the two tests that pinned the old object() placeholder so
they pin the runtime-kind distinction. Restart-phase, killed-PID, and
pre-restart-PID signals still fail closed unchanged.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.