Commit Graph

3 Commits

Author SHA1 Message Date
HexLab98 c5788b5ea3 test(agent): cover inline hard timeout and in-flight pool-request abort
Pin that cron/subagent non-streaming calls receive a read timeout matching the stale budget, that an explicit timeout is left alone, and that force_close_tcp_sockets finds sockets on httpcore PoolRequest.connection and clears the socket timeout before shutdown without close().
2026-08-14 21:19:55 -07:00
kshitij 1e5b507440 fix(cron): move watchdog state under the request lock; fail closed on resolver errors
Follow-up hardening on top of the salvaged #80809 watchdog, porting the
locked state-machine design from #75301 (credit: @Zeraphim):

- All lifecycle transitions (stale/cancelled/done) now happen under
  request_client_lock. A user or monitor interrupt marks the request
  'cancelled' so a racing stale timer can no longer misclassify the kill
  as provider staleness and feed a false +1 into the #58962 cross-turn
  circuit breaker.
- 'done' is set under the lock on completion, so a late timer callback
  that lost the race to a successful response is inert instead of
  leaving a spurious streak=1 behind the reset.
- Registration race closed: if the budget expires while the client is
  still being constructed, _make_client aborts the freshly-registered
  client and fails the call with a retryable TimeoutError instead of
  opening a brand-new socket after the only watchdog already fired.
- _resolve_direct_stale_timeout now fails closed: a raising resolver
  propagates (same as the worker path) instead of being swallowed into
  an infinite budget that would silently disarm the watchdog and
  reinstate the very hang #80759 is about.

4 new regression tests, each verified to fail against the pre-fix
watchdog implementation.

Co-authored-by: Zeraphim <diamantejc87@gmail.com>
2026-08-07 15:32:37 +05:30
HexLab98 cb066a971b fix(cron): bound the inline non-streaming call with a stale watchdog
Cron turns and delegated children are routed onto direct_api_call, which
ran the request inline with no stale detector. The abort plumbing was
registered but nothing ever invoked it, so a provider that accepted the
request and then went silent — connection held open, zero bytes, no
error — hung the run until an external actor killed it, which also
orphaned the execution row. The httpx read timeout is not a usable bound
(1800s default, and this failure mode never trips it), and the job-level
inactivity monitor was observed not to fire.

Arm a watchdog timer on the same budget the interrupt worker's poll loop
uses, so these turns get exactly the patience every other non-streaming
request already gets. On expiry it only aborts the in-flight sockets
through the already-registered hook — it never issues a request, so the
inline / no-worker property that fixes the nested-pool deadlock is
preserved — bumps the cross-turn stale circuit breaker, and surfaces a
retryable TimeoutError so the outer loop reconnects on a fresh pool.

Fixes #80759
2026-08-07 15:32:37 +05:30