Commit Graph

3 Commits

Author SHA1 Message Date
Ben Barclay 59786d280f fix(scale-to-zero): clamp dashboard marker mtime to now
A wall-clock step-back (NTP) can leave the marker's mtime in the future,
which would push the idle window out by the step size. Clamp with
min(mtime, now); _last_inbound_at has the same exposure but is at least
bounded by process uptime. Review comment on #100830.
2026-09-02 14:36:11 +10:00
Ben Barclay 7a4688d9ca fix(scale-to-zero): give a departed dashboard client the full idle_timeout grace
Review finding: dashboard_client_last_seen() discarded the marker once it
was >= 45s old, so after the client disconnected the gateway fell back to
its own (much older) _last_inbound_at and suspended ~45-75s later, not
idle_timeout later as the PR claimed. The staging release leg had in fact
shown 46s.

The marker mtime is a timestamp of real inbound; is_idle already judges
recency. Drop the staleness cutoff entirely and always fold the raw mtime
into the inbound clock (max with _last_inbound_at). An old marker is
harmless: it is outside idle_timeout just like an old _last_inbound_at.

Also:
- touch the marker immediately after ws.accept(), before the ready/skin
  setup, so a client waiting on a slow ready frame is still visible
- make the newer-message-wins test discriminating (timeout 10s, chat 5s,
  marker 40s: choosing the marker would read idle)
- add < idle_timeout / >= idle_timeout / ancient-marker boundary tests

Re-validated live on hermes-agent-stg-test-6698: last client frame
03:18:11Z -> going dormant 03:20:17Z (126s = 120s timeout + watcher tick)
while the gateway's own inbound clock was 368s stale. Mutation check vs
origin/main files: run.py hunk reverted -> 4 fail, ws.py reverted -> 3
fail, "prefer marker over max" -> 1 fail.
2026-09-02 13:21:38 +10:00
Ben Barclay 5838b2f9d8 fix(scale-to-zero): count an attached dashboard WS client as inbound activity
The gateway's idle predicate only stamped _last_inbound_at for messaging
inbound, so a scale-to-zero instance suspended under an open desktop app /
web dashboard / TUI. The client's reconnect loop then re-poked the
Fly-proxied hostname, autostart resumed the box, and the instance flapped
suspend -> proxy-wake every ~60s (13 of 72 active opted-in prod instances
on 2026-09-02, with [PC05] connect timeouts visible to the user).

The dashboard runs in a separate process on hosted instances, so the
signal crosses over as a marker file under HERMES_HOME/state:

- tui_gateway/ws.py touches state/dashboard_clients.heartbeat on every
  /api/ws connect and inbound frame (clients gateway.ping every 15s),
  throttled to one write per 5s per process.
- gateway/scale_to_zero.py: dashboard_client_last_seen() reads the mtime;
  a marker older than 45s (the clients' heartbeat deadline) is "client
  gone", a missing marker is "no client" (not fail-awake, or nothing
  would ever sleep), an unreadable marker fails awake.
- gateway/run.py folds that into seconds_since_last_inbound, so an
  attached client gets exactly the same idle_timeout grace after it
  disconnects as a chat message does. No new conjunct, _last_inbound_at
  itself is not mutated.

Validated as a hot patch on hermes-agent-stg-test-6698 with a real
/api/ws client pinging every 15s: machine held awake for 6.5 min of zero
proxied traffic (was 2 min), then suspended within ~50s of the client
exiting. Tests exercise the real seams (temp HERMES_HOME, the runner's
_scale_to_zero_is_idle composition, tui_gateway.ws.handle_ws) and were
mutation-checked: reverting either half or dropping the staleness check
fails 2-3 of them.
2026-09-02 12:27:59 +10:00