b828624479
A gateway restart kills every MCP stdio subprocess. An agent session that outlives the restart still holds a handle to the dead child, so its next tool call fails in 0.00s -- before anything reaches the network -- while the subprocess is respawned seconds later. Cron runs spanning a restart lose tool calls silently. The #81995/#95626 machinery already detects the dead child and signals a reconnect; it just never waits for it, so the caller eats the failure. Both fast-fail sites now raise _StdioChildExited, and the handler respawns the transport and retries the call once before any error reaches the model. Retrying here cannot hot-cycle respawns: the handler never spawns anything. It sets _reconnect_event (one signal per call, as before) and waits for the server task to publish a fresh session, so spawn frequency stays governed by run()'s rapid-drop budget (#62212). The retry is single-shot -- a child that dies again immediately reports and stops, and a genuinely broken server still parks with its tools deregistered. The error text no longer claims a timeout. "failing the call fast instead of waiting 300s" described a healthy remote backend as a timing problem and sent an afternoon's investigation into the wrong system. Verified on macOS against a real stdio subprocess, not only unit tests: - SIGKILL the child of a live session (what a restart does to it), then call again: 0.00s error before, 0.51s success after. - Child that exits on every tool call: 6 spawns across 8 calls, budget exhausted, parked, tools deregistered -- no respawn loop. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>