Commit Graph

3 Commits

Author SHA1 Message Date
NATHAN Menkin b828624479 fix(mcp): respawn and retry once when a stdio child died
A gateway restart kills every MCP stdio subprocess. An agent session that
outlives the restart still holds a handle to the dead child, so its next
tool call fails in 0.00s -- before anything reaches the network -- while
the subprocess is respawned seconds later. Cron runs spanning a restart
lose tool calls silently.

The #81995/#95626 machinery already detects the dead child and signals a
reconnect; it just never waits for it, so the caller eats the failure.
Both fast-fail sites now raise _StdioChildExited, and the handler respawns
the transport and retries the call once before any error reaches the model.

Retrying here cannot hot-cycle respawns: the handler never spawns anything.
It sets _reconnect_event (one signal per call, as before) and waits for the
server task to publish a fresh session, so spawn frequency stays governed by
run()'s rapid-drop budget (#62212). The retry is single-shot -- a child that
dies again immediately reports and stops, and a genuinely broken server
still parks with its tools deregistered.

The error text no longer claims a timeout. "failing the call fast instead of
waiting 300s" described a healthy remote backend as a timing problem and
sent an afternoon's investigation into the wrong system.

Verified on macOS against a real stdio subprocess, not only unit tests:
- SIGKILL the child of a live session (what a restart does to it), then
  call again: 0.00s error before, 0.51s success after.
- Child that exits on every tool call: 6 spawns across 8 calls, budget
  exhausted, parked, tools deregistered -- no respawn loop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 23:28:39 -07:00
kshitij 8966b0a700 review: tighten pre-call gate comment; drop redundant _ReadyAdapter test stub
Both from the simplify pass: the comment kept only the ownership-relevant
rationale (incl. the no-double-bump note); the test's _ReadyAdapter was a
verbatim delegate around threading.Event — the exercised paths only call
is_set/clear/set, so the bare Event is behaviorally identical.
2026-08-27 21:36:35 +05:30
kshitij 8aae2ea539 fix(mcp): signal reconnect from the mid-call fast-fail site too + regression tests
Widen the contributor's pre-call reconnect signal to the sibling site:
when the stdio subprocess dies mid-RPC the watcher race fast-fails, but
nothing cleared server.session, so the server stayed dead until the idle
keepalive probe noticed. Signal the reconnect there as well.

Also drop the explicit _bump_server_error at the pre-call gate: the
returned error payload already flows through the handler's JSON parse,
which bumps the breaker once — the explicit bump would double-count.

Two regression tests pin both sites (reconnect signaled exactly once,
no RPC attempted on a dead transport, single breaker bump).
2026-08-27 21:36:35 +05:30