Files
hermes-agent/tests
finn763 2355d593f3 fix(agent): stale-killed stream unwedges its reader and reconnects (#110769)
The stale-stream monitor aborted a wedged provider stream only via
force_close_tcp_sockets() -> shutdown(SHUT_RDWR). That is best-effort: a
parked body read is not unblocked on every platform (Windows keeps the
pending recv parked) and the sweep can miss the socket. The worker then
stayed blocked in the provider read, so the retry loop never retried; the
monitor re-killed every stale interval and the call only ended at the
byte-read timeout, far past the stale budget - the reported
"No response from provider for 180-240s ... Reconnecting" loop ending in
"The model server is not responding".

- _kill_stale_stream now also closes the killed attempt's own provider
  response (identity-guarded self._attempt_stream_response), which is what
  actually unblocks a parked reader; a racing retry's fresh response is
  never touched.
- an abort-induced httpx.ReadError counts as a transient connection error,
  so the aborted attempt reconnects instead of ending the turn.

Reproduced with a local SSE server: before, the worker stayed parked and no
second request was issued (recovery only at the byte-read timeout); after,
the kill unblocks the reader at the stale budget and the retry lands.
2026-09-15 12:46:27 +05:30
..
…