Files
hermes-agent/tools
NATHAN Menkin b828624479 fix(mcp): respawn and retry once when a stdio child died
A gateway restart kills every MCP stdio subprocess. An agent session that
outlives the restart still holds a handle to the dead child, so its next
tool call fails in 0.00s -- before anything reaches the network -- while
the subprocess is respawned seconds later. Cron runs spanning a restart
lose tool calls silently.

The #81995/#95626 machinery already detects the dead child and signals a
reconnect; it just never waits for it, so the caller eats the failure.
Both fast-fail sites now raise _StdioChildExited, and the handler respawns
the transport and retries the call once before any error reaches the model.

Retrying here cannot hot-cycle respawns: the handler never spawns anything.
It sets _reconnect_event (one signal per call, as before) and waits for the
server task to publish a fresh session, so spawn frequency stays governed by
run()'s rapid-drop budget (#62212). The retry is single-shot -- a child that
dies again immediately reports and stops, and a genuinely broken server
still parks with its tools deregistered.

The error text no longer claims a timeout. "failing the call fast instead of
waiting 300s" described a healthy remote backend as a timing problem and
sent an afternoon's investigation into the wrong system.

Verified on macOS against a real stdio subprocess, not only unit tests:
- SIGKILL the child of a live session (what a restart does to it), then
  call again: 0.00s error before, 0.51s success after.
- Child that exits on every tool call: 6 spawns across 8 calls, budget
  exhausted, parked, tools deregistered -- no respawn loop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 23:28:39 -07:00
..