A wall-clock step-back (NTP) can leave the marker's mtime in the future,
which would push the idle window out by the step size. Clamp with
min(mtime, now); _last_inbound_at has the same exposure but is at least
bounded by process uptime. Review comment on #100830.
Review finding: dashboard_client_last_seen() discarded the marker once it
was >= 45s old, so after the client disconnected the gateway fell back to
its own (much older) _last_inbound_at and suspended ~45-75s later, not
idle_timeout later as the PR claimed. The staging release leg had in fact
shown 46s.
The marker mtime is a timestamp of real inbound; is_idle already judges
recency. Drop the staleness cutoff entirely and always fold the raw mtime
into the inbound clock (max with _last_inbound_at). An old marker is
harmless: it is outside idle_timeout just like an old _last_inbound_at.
Also:
- touch the marker immediately after ws.accept(), before the ready/skin
setup, so a client waiting on a slow ready frame is still visible
- make the newer-message-wins test discriminating (timeout 10s, chat 5s,
marker 40s: choosing the marker would read idle)
- add < idle_timeout / >= idle_timeout / ancient-marker boundary tests
Re-validated live on hermes-agent-stg-test-6698: last client frame
03:18:11Z -> going dormant 03:20:17Z (126s = 120s timeout + watcher tick)
while the gateway's own inbound clock was 368s stale. Mutation check vs
origin/main files: run.py hunk reverted -> 4 fail, ws.py reverted -> 3
fail, "prefer marker over max" -> 1 fail.
The gateway's idle predicate only stamped _last_inbound_at for messaging
inbound, so a scale-to-zero instance suspended under an open desktop app /
web dashboard / TUI. The client's reconnect loop then re-poked the
Fly-proxied hostname, autostart resumed the box, and the instance flapped
suspend -> proxy-wake every ~60s (13 of 72 active opted-in prod instances
on 2026-09-02, with [PC05] connect timeouts visible to the user).
The dashboard runs in a separate process on hosted instances, so the
signal crosses over as a marker file under HERMES_HOME/state:
- tui_gateway/ws.py touches state/dashboard_clients.heartbeat on every
/api/ws connect and inbound frame (clients gateway.ping every 15s),
throttled to one write per 5s per process.
- gateway/scale_to_zero.py: dashboard_client_last_seen() reads the mtime;
a marker older than 45s (the clients' heartbeat deadline) is "client
gone", a missing marker is "no client" (not fail-awake, or nothing
would ever sleep), an unreadable marker fails awake.
- gateway/run.py folds that into seconds_since_last_inbound, so an
attached client gets exactly the same idle_timeout grace after it
disconnects as a chat message does. No new conjunct, _last_inbound_at
itself is not mutated.
Validated as a hot patch on hermes-agent-stg-test-6698 with a real
/api/ws client pinging every 15s: machine held awake for 6.5 min of zero
proxied traffic (was 2 min), then suspended within ~50s of the client
exiting. Tests exercise the real seams (temp HERMES_HOME, the runner's
_scale_to_zero_is_idle composition, tui_gateway.ws.handle_ws) and were
mutation-checked: reverting either half or dropping the staleness check
fails 2-3 of them.