Files
hermes-agent/hermes_cli
Teknium 8fdfe04371 fix(update): probe gateway loop liveness before drain; bounded escalation for wedged gateways
A gateway whose asyncio event loop is stalled (e.g. an in-loop
compression pass, #72707) cannot process SIGTERM/SIGUSR1 shutdown.
The updater's drain wait then burned the full 180s budget, warned
"Gateway PID X still running after 180.0s — restart may fail", and
`hermes update` could deadlock behind the wedged process — the user
cannot update their way out of the stall.

Fix: before any drain wait, read the loop-liveness heartbeat file the
gateway rewrites every 30s (#66892). Classification:

- alive (fresh heartbeat): busy-but-alive loop — take the normal
  graceful drain, honoring the in-flight cron drain floor (#86684).
- wedged (heartbeat for this PID stale >90s = 3 missed beats): the
  loop is provably dead; drain is pointless. Bounded escalation:
  SIGTERM + 5s grace, then SIGKILL + 5s wait, then proceed (~10s
  worst case, far under the 180s drain budget).
- unknown (missing/corrupt file, PID mismatch): never escalate on
  ambiguity — full drain path.

Wired into launchd_restart, systemd_restart, and both updater
gateway-shutdown sites (systemd unit drain + manual profile
gateways). The probe is a local stat + JSON read (well inside the
10s query tier of the subprocess timeout tiering).

The cron drain floor from #86684 is bypassed ONLY when the loop is
provably dead — a merely busy gateway still refreshes its heartbeat
and keeps the full drain budget.

Root cause of the loop stall itself (compression blocking the loop)
is #72707 territory and deliberately out of scope here.

Fixes #81642
2026-08-15 02:45:27 -07:00
..
2026-08-13 11:52:47 -07:00
…
2026-08-10 12:43:46 -07:00