- deadline: _result()/_abandon() replace 7 BoundedResult constructions and 2
cancel+callback sites; timers handled as a list; dead raise_if_timed_out removed.
- file_safety: retired classify_cross_profile_target (0 refs; get_cross_profile_warning
stub kept for external callers), _home_and_resolved/_mirror_warning shared by
the sandbox/container mirror guards; _find_sandbox_mirror_segments inlined.
- estop: _hermes_home/_canonical_root now the file_safety helpers;
_reset_log_state_for_tests inlined into its only test.
- subagent_lifecycle: _validate_request driven by _UNSUPPORTED_REQUEST_FIELDS table.
- turn_liveness: dead start()/_abort_message removed, _emit_warning shared.
- process_bootstrap: _enable_happy_eyeballs reused by the client variant.
- Comment/docstring compaction across the remaining leaf modules.
A hung terminal wait on the loop thread silently disabled asyncio deadlines
and let cron jobs idle thousands of seconds past HERMES_CRON_TIMEOUT. Drive
the wait from run_bounded_sync (sliced Event.wait, kill-on-timeout) and
move the cron inactivity monitor onto a daemon thread with the same kernel
timeout primitive. Copy the caller ContextVar scope and activity callback
onto the wait worker so profile secrets, session id, and heartbeats survive
the thread hop (#94285).
agent/deadline.py defined SuspectableBackend twice: the Phase 3a Protocol
(sync ensure_healthy(self) -> bool) and, further down the same module, an
unrelated concrete class with the same name (async
ensure_healthy(self, timeout=5.0)) added later by the MCP Phase 3b adopter.
Since Python executes class statements top-to-bottom, the second definition
silently shadowed the first at module scope.
Nothing in the tree imports or subclasses either by name today — the MCP
adopter duck-types the same-shaped contract directly on its own connection
class rather than referencing agent.deadline.SuspectableBackend — so this
caused no live behavior change. But it left the wrong (and differently
shaped) class resolvable under that name for the next Phase 3b adopter that
does import it for a type hint.
Record correction: the previous commit's message says the async flavor
offloads mark_suspect via asyncio.to_thread — it does NOT (and must not).
The mark is deliberately inline on the event loop: running it
synchronously guarantees mark-happens-before-BoundedResult-return and
mark-before-on_abandon-cleanup (cleanup is ensure_future'd and cannot
start until the next loop tick). An offloaded mark would race both.
The trade-off is that a slow adopter mark_suspect would block the loop
(measured: a 2s mark stalls every coroutine for 2.003s), so the adopter
contract is now explicit in the Protocol docstring and at the async call
site: mark_suspect must be cheap, non-blocking, lock-free; expensive
recycle work belongs in ensure_healthy.
New pins so the negotiated semantics can't silently regress:
- test_sync_mark_happens_before_on_timeout (the review-round ordering)
- test_async_mark_happens_before_on_abandon_cleanup (the scheduling
invariant an offloaded mark would break)
- test_sync_completion_never_marks_backend (sync counterpart of the
async completion test)
- mark_suspect runs BEFORE owner cleanup in both flavors (the reason
describes the state at timeout; a recycling cleanup never poisons the
healed replacement)
- the async flavor offloads the mark off the event loop
(asyncio.to_thread), matching how owner cleanup is scheduled
- the protocol documents the synchronous-cheap contract for adopters
- the windows-footgun annotation stays on its matched killpg line
Phase 3a of the #85125 unified-deadline plan. run_bounded_async and
run_bounded_sync accept backend= and call mark_suspect(label +
timeout) exactly once on timeout, never on completion. The layer
fails open: backends without the protocol (incremental Phase 3b
adoption) and raising mark_suspect implementations can never weaken
the deadline bound or corrupt the BoundedResult.
Fixes the four poisoned-connection classes (#81051, #77765, #84132,
#81995) with the SuspectableBackend cheap-mark/lazy-verify contract:
- mark_suspect/ensure_healthy protocol (agent/deadline.py): noticing a
poisoned state never does I/O; the NEXT caller pays once for a health
probe that clears the suspicion or forces a reconnect. A single
teardown-vs-keepalive race or auth-lock corruption can no longer park
a connection permanently — park stays reserved for genuinely
exhausted reconnect budgets.
- keepalive failure marks the connection suspect before requesting
reconnect; the next tool call probes and recycles if unhealthy.
- auth-classified permanent failures on a previously-proven session get
a suspect+reconnect path instead of an immediate park.
- fast-fail (#81995): stdio child pids are tracked at spawn and an
in-flight RPC races a child-watcher task, so a dead subprocess fails
the call immediately with a retryable timeout instead of riding out
the full 300s. Deliberate teardown/reconnect also fails in-flight
calls now instead of leaving them attached to a dying transport.
Dispatch-boundary hardening for test doubles: stubbed sessions
(MagicMock/non-awaitable call_tool, absent child-watcher) fall back to
the exact pre-change inline-await semantics, so only real transports
gain the race guard.
Salvage credit: in-flight approach from #73377 (@luijoc, wedged
transport recovery) and #48069 (@arminanton, keepalive/in-flight
interaction); both PRs' bases predate main's current park/reconnect
architecture, so this is a fresh implementation of their contracts.
Tests: tests/tools/ -k mcp = 639 passed (was 22 new failures during
development; final tree zero).
The os.killpg call sits below an early 'if sys.platform == win32: return'
so it can never execute on Windows; the scanner is line-based and needs
the inline marker.
- run_bounded_async: cancel + abandon the inner task when the CALLER is
cancelled (leak the telegram original also had)
- kill_process_tree: check taskkill exit code (Windows contract parity),
suppress console flash via windows_hide_flags, and sweep a psutil
descendant snapshot taken before signalling — reaches grandchildren in
their own setsid sessions and the non-group-leader case (#71148 class)
- resolve_timeout: reject bool (YAML true would become a 1s deadline) and
NaN config values with fall-through instead of resolving unbounded
- BoundedResult: kw_only to prevent positional transposition
- tests: real clamped-value time_t regression proof, own-session
descendant kill, external-cancellation task cleanup, bool/NaN config
fall-through; pin already-dead-pid contract
One shared foundation for the timeout/hang backlog instead of per-incident
site-local fixes:
- agent/deadline.py: run_bounded_async (thread-timer deadline that survives
a blocked event loop, generalizing the telegram adapter primitive),
run_bounded_sync, clamp_timeout (kills the #83220 time_t OverflowError
class at the boundary), resolve_timeout (config.yaml timeouts: section >
legacy env bridge > default), kill_process_tree (whole-tree termination
for the #71148 orphan class), DeadlineExpired (our deadline, mechanically
distinct from provider timeouts).
- tool_executor._resolve_concurrent_tool_timeout migrates onto the resolver;
exact legacy env-var contract preserved (default 420, 0 disables).
- timeouts: accepted as a known config root; documented in
cli-config.yaml.example.
Pure addition otherwise — no behavior change, no new env vars, no cache
impact. Later phases (#85125) migrate tool-execution, MCP, and subprocess
call sites onto these primitives.