f50b5bb0fa
The Codex auxiliary Responses adapter enforced a single absolute deadline (300s floor for compression). A dead stream held the entire budget before fallback ran, and repeated compression attempts stacked those waits into 20+ minute 'Summarizing thread' stalls (masoria debug bundle, Aug 31 2026). Meanwhile a healthy-but-slow reasoning summary was killed at the same absolute deadline even while producing tokens. Replace the absolute kill with progress-aware deadlines: - 60s no-progress window for the first substantive payload AND between payloads; keepalive/lifecycle frames do not re-arm (mirrors the commit-fence gating, #96707) - a live stream re-arms per token and is bounded only by _aux_stream_total_ceiling() (max(600s, 4x configured timeout)), the same backstop the streamed chat.completions path already uses - the compression critical-path retry gate now distinguishes failure cost: a cheap first-token no-progress failure retries the same provider once; mid-stream stalls and ceiling hits still skip straight to provider fallback (#54465 semantics preserved) Live A/B (real OpenAI SDK against a local SSE server, real adapter): dead keepalive-only stream: main waits the full budget; fixed fails over at the window. Slow-but-alive stream (tokens past the configured timeout): main kills it mid-generation; fixed completes.