Windows updates forced a choice between 'gateway survives' and 'update
proceeds': the pause machinery's only tools were the planned-stop marker
poll and the force-kill ladder, so a mid-turn gateway was tree-killed and
its active turn lost. Step 2 of the socket migration adds the
pause-for-update verb: the updater ASKS the gateway to drain in-flight
turns and exit cleanly — releasing every venv file handle on the way out
— through the same request_restart(via_service=True) drain path SIGUSR1
and service restarts already use.
- gateway/run.py: pause-for-update verb handler registered on the
existing control server; marshals onto the loop thread, ACKs with
{pausing, already_stopping, pid, drain_timeout}.
- gateway/control_socket.py: pause_gateway_for_update() client — None on
no-answer (older gateway / no socket), so every caller keeps the
legacy path when the verb is missing.
- update_cmd.py (_pause_windows_gateways_for_update): socket-first ask
per mapped profile gateway before the drain wait; positive ACKs extend
the wait to the gateway's own declared drain budget (+ teardown grace)
so a mid-turn gateway isn't force-killed at the end of a too-short
local default. Marker write + force-kill ladder retained verbatim as
the fallback.
Live E2E: real gateway process (isolated HERMES_HOME), real socket:
identify -> pause ACK {pausing: true} -> gateway drained and exited on
its own (rc=75, zero signals) -> dead-gateway re-ask returns None.
A step-1 gateway without the verb answers ok:false -> client None ->
legacy path (pinned by test).