2 Commits

Author SHA1 Message Date
Teknium f6938b37f3 simplify(compat): terminal/file/environments — drop 42 re-exports/aliases, repoint 20 callers + 41 test files
tools/terminal_tool.py: drop 30 pure re-export names (lifecycle/config/backends/
sudo/guards/result/interrupt/utils/_DockerEnvironment/is_managed_tool_gateway_ready)
and the noqa-F401 comments on the 25 names the facade itself uses. Sibling modules
(terminal_tool_backends/_result/_sudo/_lifecycle, environments/base, process_registry)
that read removed names through the facade now import from the defining module.
tools/environments/base.py: drop 11 re-exports (base_output/base_session_env/
path_utils) and the BaseEnvironment.stop() compat alias (no in-tree caller; the
lifecycle hasattr(env, 'stop') fallback stays for third-party envs).
tools/environments/docker.py: drop 1 re-export + the re-export comment.
Callers/tests repointed to tools.terminal_tool_{lifecycle,backends,sudo,config,
guards,result}, tools.interrupt, tools.environments.{base_output,base_session_env,
path_utils}.
2026-09-03 13:29:55 -07:00
rainbowgits 1f6836cd81 fix(mcp): bound stdio initialize handshake to stop subprocess/FD leak
A stdio MCP server that never completes `initialize` (e.g. emits a
non-JSON-RPC frame and then blocks on stdin) leaks a child process plus its
stdio pipes/pidfd on every discovery-retry cycle — unbounded, until the
gateway hits EMFILE and every new open()/spawn fails (#59349).

Root cause (confirmed by instrumenting the live repro, and different from the
issue's own hypothesis): the spawned child IS captured in `new_pids`, so the
report's "new_pids empty at finally" guess is not it. The real cause is that
`session.initialize()` hangs forever on the garbage stream. `connect_timeout`
only bounds the caller's `.result()` wait on the foreground thread — it does
NOT cancel the `_run_stdio` coroutine on the background MCP loop. So the
coroutine is stuck at `await session.initialize()` permanently, its cleanup
`finally` never runs, the child is never reaped, and it stays invisible to the
orphan-reaper (whose `_orphan_stdio_pids` set never gets populated).

Fix: wrap `session.initialize()` in `asyncio.wait_for(..., connect_timeout)`
so a stalled handshake fails instead of hanging. The TimeoutError unwinds
through the SDK context managers (closing the child's stdin -> EOF -> exit)
and lets the existing `finally` reap any straggler. Cross-platform — no
signals/pgid/proc.

Scope: stdio only. The HTTP path has the same `await session.initialize()`
shape but spawns no subprocess (so it can't cause this leak) and already has
httpx transport timeouts.

Verified: the reporter's repro goes from unbounded growth to draining to zero;
added a hermetic regression test (fake transport whose `initialize()` hangs,
asserts the connect is bounded by connect_timeout) that fails on the pre-fix
code and passes on the fix; 566 existing MCP tests pass; ruff clean.

Repro confirmed on macOS (pipe FDs); the Linux-specific pidfd growth in the
report should be equivalent — the reporter offered to validate on Linux.

Closes #59349
2026-07-07 15:16:00 -07:00