`hermes update` printed "draining (up to 1875s)..." and then nothing for up
to 30 minutes while the gateway's in-band restart waited on in-flight work
(agent.restart_after_turn_timeout). Neither the updater nor the gateway log
said WHAT was being waited on, so a single long cron job read as a hung
update.
Gateway side: GatewayShutdownMixin._describe_active_work() enumerates each
unit the restart wait holds for — chat turns (session key, model, current
tool, elapsed), cron jobs (job id, elapsed, and the restart-safe external
worker pid when the run was handed off; cron/scheduler now records that pid
next to the running id), api/deferred runs by count. It is written to
gateway_state.json as `active_work` while the state is `draining` (cleared
otherwise) and appended to the 30s "Restart deferred" log line.
CLI side: hermes_cli/update_cmd_drain_report.py reads `active_work` and
prints a progress block every 30s during the SIGUSR1 exit wait — the
holder(s), their pids, elapsed time, seconds left before the forced
restart, and the config knob that caps the wait. Wired into the systemd,
launchd and manual gateway restart paths of `hermes update` and into
`hermes gateway restart`; `hermes gateway status` lists the same units
while draining. A pre-fix gateway (no `active_work` field) gets an explicit
"gateway did not report" line rather than silence.
Live A/B (real gateway, 90s no-agent cron job in flight, SIGUSR1 from the
caller): base = 79s of silence, no `active_work` in the state file; head =
the job named with pid/elapsed/remaining every interval, log line carries
the same detail.
After the fleet restart is verified healthy, `hermes update` runs the
migration preflight on installs with >= 2 profiles and at least one
per-profile gateway. No blockers: migrate (same path as
`gateway migrate --multiplex --yes`, deterministic, never prompts).
Blockers: print them with their fixes and the one-liner, change nothing.
Skipped on the exit-1 (stale fleet) path and on single-profile installs.
Reconcile receipt-only restart obligations at the shared warning/catch-up
predicate, requiring every historical runtime/profile identity to have a
current live gateway successor. Preserve missing and unknown obligations,
non-gateway identities, and independently authoritative pending markers.
Keep failed receipts unchanged instead of recording an unverified success.
Live isolated two-process A/B reproduces the warning on base and settles
it after the fix; stale, unknown, and missing-profile controls still warn.
Reported-by: duanzhiwei0315
Inspired-by: zengzheqing (#104295), RootZ3n (#100249)
Salvage the unit-budget implementation from #104745, replacing its test
matrix with two invariant tests and covering the sibling graceful start.
Keep unprivileged property reads, finite fallbacks, real manager errors,
and post-restart health verification.
Native disposable user unit: old client timed out after 15.03 seconds;
new client completed the same 16-second stop transaction in 16.13 seconds.
The unit stayed active with a new PID; missing-unit errors stayed errors.
Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
Discover systemd targets before stopping old processes, restart even when
there are no gateway PIDs, and require successful scope listings plus active
verification. Pending launchd recovery also retains failures for inaccessible
listings and installed jobs without supervision. Keep existing PID cleanup
intact but before recovery so it cannot kill freshly verified workers.
Slim redo informed by #104274, #104283, and #104285.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
The restart phase records macOS LaunchAgent labels (ai.hermes.gateway).
match_runtime_outcomes used a substring check for "hermes-gateway", so a
successful Desktop update on the default profile always tripped
"Planned runtimes the restart phase never touched" and exited 1.
Use the exact systemd/launchd/s6 matcher for both plan reconciliation
and abort-recovery so the two cannot drift.
The tail clause was correct only by ordering (exit_code != 0 there meant
exit_code is None). Name what it encodes: a stop_reason counts only when
nothing vouched for success. Same truth table. The three literal-dict tests on
the predicate collapse into one parametrized contract; the handoff-exit test
binds to COMMAND_BOUNDARY_STOP_REASON instead of re-spelling it.
update_contract writes {"outcome": "refused", "stop_reason": <code>} with no
exit_code; that is the one production receipt where the stop_reason clause in
_receipt_looks_unfinished is load-bearing. The previous negative control used
exit_code=1, which the exit_code branch already catches. Docstring reworded:
a KeyboardInterrupt never lands on a success receipt (the boundary finalize is
a no-op once the inner path finalized).
A gateway with no profile mapping (or one whose relaunch could not be armed) is
SIGTERMed and listed under "Restart manually" — by design it has no successor and
publishes no fleet-matrix row. It still counted in ``killed_pids`` and the
pre-restart snapshot, so ``_fleet_probe_expected_runtimes`` demanded rows that
could not exist and a fully successful update exited 1 with "Fleet version check
returned no rows even though gateway runtimes were expected", leaving the
fleet_restart_pending marker behind and every later CLI start warning about it.
Track the unmapped stops on the restart outcome and subtract them from the
row-predicting signals (``fleet_probe_signals``); relaunched/systemd gateways
still predict rows exactly as before.
A dashboard-only runtime plan (gateway never started) made
_fleet_probe_expected_runtimes() return True from the unfiltered
'plan.runtimes is non-empty' check. collect_fleet_versions() reports
gateway identities only, so the probe waited for rows that cannot exist,
printed the incomplete-verification warning, and exited 1 after a
successful update (#97332).
Key the plan-derived expectation on kind == 'gateway' records — the same
row-capability rule already applied to the Windows resume token (#93406)
— and update the two tests that pinned the old object() placeholder so
they pin the runtime-kind distinction. Restart-phase, killed-PID, and
pre-restart-PID signals still fail closed unchanged.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.