turn_in_flight read only the dashboard session table; an in-process cron run
never registers there, so the watchdog reported "no running turn" and exited
mid-job (tool calls then failed with "cannot schedule new futures after
interpreter shutdown", the execution was marked unknown, the slot lost). The
probe now also consults cron.scheduler.get_running_job_ids — the ledger the
gateway shutdown drain already uses.
Addresses #107485
serve --isolated is detached on purpose (setsid/nohup, PPID 1) so it survives the SSH
channel closing, and every teardown path lived on the client. A laptop that sleeps
mid-session (dark wake reconnects the tunnel, spawns a backend, sleeps again) therefore
left a new backend behind every cycle, each one an extra writer on state.db — the
multi-writer source behind the WAL corruption incidents.
Only for backends started with --ssh-session-token-file: an ASGI wrapper counts accepted
WebSocket sessions on every dashboard route without touching the handlers; a watchdog
requests a graceful uvicorn exit (WAL checkpoint, exit 0) once no client has been
connected for dashboard.ssh_isolated_idle_grace_s (default 15 min) AND no agent turn is
running. An unreadable turn state fails closed (the backend stays up). Loopback normally
disables the WS ping, but across a tunnel the local socket stays healthy while the far
end sleeps, so these backends keep a slow ping (60s / 10 min) that a GIL-holding turn
cannot trip.
Design, client-count/turn-probe/fail-closed shape and the tunnel-ping rationale from
#101678 by @StanleyStetson; this is the slim redo on the decomposed web server (no
exclusive home lock / handover protocol — the newcomer never needs to evict an idle
incumbent once idle incumbents exit on their own).