Commit Graph

26 Commits

Author SHA1 Message Date
teknium1 6a312fba54 feat(update): name the work a draining gateway is waiting on
`hermes update` printed "draining (up to 1875s)..." and then nothing for up
to 30 minutes while the gateway's in-band restart waited on in-flight work
(agent.restart_after_turn_timeout). Neither the updater nor the gateway log
said WHAT was being waited on, so a single long cron job read as a hung
update.

Gateway side: GatewayShutdownMixin._describe_active_work() enumerates each
unit the restart wait holds for — chat turns (session key, model, current
tool, elapsed), cron jobs (job id, elapsed, and the restart-safe external
worker pid when the run was handed off; cron/scheduler now records that pid
next to the running id), api/deferred runs by count. It is written to
gateway_state.json as `active_work` while the state is `draining` (cleared
otherwise) and appended to the 30s "Restart deferred" log line.

CLI side: hermes_cli/update_cmd_drain_report.py reads `active_work` and
prints a progress block every 30s during the SIGUSR1 exit wait — the
holder(s), their pids, elapsed time, seconds left before the forced
restart, and the config knob that caps the wait. Wired into the systemd,
launchd and manual gateway restart paths of `hermes update` and into
`hermes gateway restart`; `hermes gateway status` lists the same units
while draining. A pre-fix gateway (no `active_work` field) gets an explicit
"gateway did not report" line rather than silence.

Live A/B (real gateway, 90s no-agent cron job in flight, SIGUSR1 from the
caller): base = 79s of silence, no `active_work` in the state file; head =
the job named with pid/elapsed/remaining every interval, log line carries
the same detail.
2026-09-13 05:08:20 -07:00
Teknium 07a4ae016a feat(update): auto-migrate to one multiplexed gateway when unblocked
After the fleet restart is verified healthy, `hermes update` runs the
migration preflight on installs with >= 2 profiles and at least one
per-profile gateway. No blockers: migrate (same path as
`gateway migrate --multiplex --yes`, deterministic, never prompts).
Blockers: print them with their fixes and the one-liner, change nothing.
Skipped on the exit-1 (stale fleet) path and on single-profile installs.
2026-09-12 01:49:28 -07:00
Teknium 89c85b8466 fix(update): settle stale receipt warnings from matching live gateways
Reconcile receipt-only restart obligations at the shared warning/catch-up
predicate, requiring every historical runtime/profile identity to have a
current live gateway successor. Preserve missing and unknown obligations,
non-gateway identities, and independently authoritative pending markers.

Keep failed receipts unchanged instead of recording an unverified success.
Live isolated two-process A/B reproduces the warning on base and settles
it after the fix; stale, unknown, and missing-profile controls still warn.

Reported-by: duanzhiwei0315
Inspired-by: zengzheqing (#104295), RootZ3n (#100249)
2026-09-07 08:20:46 -07:00
Teknium 08f2c78d92 fix(update): cover catch-up restart clients with the unit budget 2026-09-07 08:20:09 -07:00
doryani-agent 2a980fbcbd fix(update): let systemd clients outwait legitimate unit transactions
Salvage the unit-budget implementation from #104745, replacing its test
matrix with two invariant tests and covering the sibling graceful start.
Keep unprivileged property reads, finite fallbacks, real manager errors,
and post-restart health verification.

Native disposable user unit: old client timed out after 15.03 seconds;
new client completed the same 16-second stop transaction in 16.13 seconds.
The unit stayed active with a new PID; missing-unit errors stayed errors.

Co-authored-by: Teknium <127238744+teknium1@users.noreply.github.com>
2026-09-07 08:20:09 -07:00
Teknium 7798241eab fix: retain pending fleet restarts until supervisors recover
Discover systemd targets before stopping old processes, restart even when
there are no gateway PIDs, and require successful scope listings plus active
verification. Pending launchd recovery also retains failures for inaccessible
listings and installed jobs without supervision. Keep existing PID cleanup
intact but before recovery so it cannot kill freshly verified workers.

Slim redo informed by #104274, #104283, and #104285.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-07 06:10:49 -07:00
kshitijk4poor a262b2e372 refactor(update): drop the tombstone comment at the old matcher site and a duplicated exactness leg 2026-09-06 20:47:52 +05:30
mengtanx b4b6235239 fix(update): credit launchd ai.hermes.gateway in fleet reconciliation (#103679)
The restart phase records macOS LaunchAgent labels (ai.hermes.gateway).
match_runtime_outcomes used a substring check for "hermes-gateway", so a
successful Desktop update on the default profile always tripped
"Planned runtimes the restart phase never touched" and exited 1.

Use the exact systemd/launchd/s6 matcher for both plan reconciliation
and abort-recovery so the two cannot drift.
2026-09-06 20:47:52 +05:30
kshitijk4poor bc1330eebc refactor(update): name the success invariant in _receipt_looks_unfinished; one predicate contract
The tail clause was correct only by ordering (exit_code != 0 there meant
exit_code is None). Name what it encodes: a stop_reason counts only when
nothing vouched for success. Same truth table. The three literal-dict tests on
the predicate collapse into one parametrized contract; the handoff-exit test
binds to COMMAND_BOUNDARY_STOP_REASON instead of re-spelling it.
2026-09-06 14:53:33 +05:30
kshitijk4poor dad698d88c test(update): pin the refused-receipt shape the stop_reason clause exists for
update_contract writes {"outcome": "refused", "stop_reason": <code>} with no
exit_code; that is the one production receipt where the stop_reason clause in
_receipt_looks_unfinished is load-bearing. The previous negative control used
exit_code=1, which the exit_code branch already catches. Docstring reworded:
a KeyboardInterrupt never lands on a success receipt (the boundary finalize is
a no-op once the inner path finalized).
2026-09-06 14:53:33 +05:30
tachyon-r cd27df3c2e fix(update): ignore successful receipt stop reason 2026-09-06 14:53:33 +05:30
Teknium 1cce7c6dd8 fix(update): unmapped gateway stops no longer fail the fleet check with "no rows"
A gateway with no profile mapping (or one whose relaunch could not be armed) is
SIGTERMed and listed under "Restart manually" — by design it has no successor and
publishes no fleet-matrix row. It still counted in ``killed_pids`` and the
pre-restart snapshot, so ``_fleet_probe_expected_runtimes`` demanded rows that
could not exist and a fully successful update exited 1 with "Fleet version check
returned no rows even though gateway runtimes were expected", leaving the
fleet_restart_pending marker behind and every later CLI start warning about it.

Track the unmapped stops on the restart outcome and subtract them from the
row-predicting signals (``fleet_probe_signals``); relaunched/systemd gateways
still predict rows exactly as before.
2026-09-05 00:24:40 -07:00
liuhao1024 b8e3c5c700 fix(update): scope fleet-probe runtime expectation to gateway-kind plan records
A dashboard-only runtime plan (gateway never started) made
_fleet_probe_expected_runtimes() return True from the unfiltered
'plan.runtimes is non-empty' check. collect_fleet_versions() reports
gateway identities only, so the probe waited for rows that cannot exist,
printed the incomplete-verification warning, and exited 1 after a
successful update (#97332).

Key the plan-derived expectation on kind == 'gateway' records — the same
row-capability rule already applied to the Windows resume token (#93406)
— and update the two tests that pinned the old object() placeholder so
they pin the runtime-kind distinction. Restart-phase, killed-PID, and
pre-restart-PID signals still fail closed unchanged.
2026-09-05 00:24:40 -07:00
Teknium 474eed838e review-fix(suppress-audit): patch_parser/skills_sync/update_cmd_fleet/api_server — restore BASE exception semantics 2026-09-03 09:56:39 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 4b9b9a441f refactor(hermes_cli/update): pack flat import/argument lists (AST-identical) 2026-09-02 22:14:59 -07:00
Teknium d9941ee4ce refactor(hermes_cli/update): reflow to 120 cols, drop blanks after local imports (AST-identical) 2026-09-02 21:45:19 -07:00
Teknium 90052abfb9 refactor(hermes_cli/update): reflow short multi-line calls and signatures (AST-identical) 2026-09-02 21:30:17 -07:00
Teknium e64942a018 refactor(hermes_cli/update): compact fleet restart bookkeeping into outcome dataclass; split maint backup/notice phases 2026-09-02 21:18:02 -07:00
Teknium ce45e216e7 refactor(update): promote nested _restart_one_systemd_gateway_unit closure to a module-level function (184 -> 70 LOC orchestrator) 2026-09-02 16:55:53 -07:00
Teknium 60041b787a refactor(update): join short multi-line statements onto one line (AST-identical, -416 lines) 2026-09-02 16:52:38 -07:00
Teknium 1aa9285312 refactor(update): hand-compact comments/docstrings in update_cmd.py, update_cmd_fleet.py, update_cmd_zip.py (AST-identical); rewrite stale module docstring 2026-09-02 16:45:44 -07:00
Teknium d46eb1f964 refactor(update): collapse 90 try/except-pass and try/except-logger.debug blocks into suppress()/_best_effort() (new leaf update_cmd_common.py) 2026-09-02 16:34:55 -07:00
Teknium e118abc61a refactor(update): lift manual-gateway restart, stuck-PID sweep and phase-abort recovery out of _restart_gateway_fleet_after_update (362 -> 154 LOC) 2026-09-02 16:23:31 -07:00
Teknium 4674f8b254 refactor(update): split zip/stash/config/deps/git/maint clusters out of update_cmd.py 2026-09-02 16:00:26 -07:00
Teknium 096826bf7d refactor(update): split gateway fleet restart/verify into update_cmd_fleet.py 2026-09-02 15:55:50 -07:00