Pin that cron/subagent non-streaming calls receive a read timeout matching the stale budget, that an explicit timeout is left alone, and that force_close_tcp_sockets finds sockets on httpcore PoolRequest.connection and clears the socket timeout before shutdown without close().
Follow-up hardening on top of the salvaged #80809 watchdog, porting the
locked state-machine design from #75301 (credit: @Zeraphim):
- All lifecycle transitions (stale/cancelled/done) now happen under
request_client_lock. A user or monitor interrupt marks the request
'cancelled' so a racing stale timer can no longer misclassify the kill
as provider staleness and feed a false +1 into the #58962 cross-turn
circuit breaker.
- 'done' is set under the lock on completion, so a late timer callback
that lost the race to a successful response is inert instead of
leaving a spurious streak=1 behind the reset.
- Registration race closed: if the budget expires while the client is
still being constructed, _make_client aborts the freshly-registered
client and fails the call with a retryable TimeoutError instead of
opening a brand-new socket after the only watchdog already fired.
- _resolve_direct_stale_timeout now fails closed: a raising resolver
propagates (same as the worker path) instead of being swallowed into
an infinite budget that would silently disarm the watchdog and
reinstate the very hang #80759 is about.
4 new regression tests, each verified to fail against the pre-fix
watchdog implementation.
Co-authored-by: Zeraphim <diamantejc87@gmail.com>
Cron turns and delegated children are routed onto direct_api_call, which
ran the request inline with no stale detector. The abort plumbing was
registered but nothing ever invoked it, so a provider that accepted the
request and then went silent — connection held open, zero bytes, no
error — hung the run until an external actor killed it, which also
orphaned the execution row. The httpx read timeout is not a usable bound
(1800s default, and this failure mode never trips it), and the job-level
inactivity monitor was observed not to fire.
Arm a watchdog timer on the same budget the interrupt worker's poll loop
uses, so these turns get exactly the patience every other non-streaming
request already gets. On expiry it only aborts the in-flight sockets
through the already-registered hook — it never issues a request, so the
inline / no-worker property that fixes the nested-pool deadlock is
preserved — bumps the cross-turn stale circuit breaker, and surfaces a
retryable TimeoutError so the outer loop reconnects on a fresh pool.
Fixes#80759