Commit Graph

10 Commits

Author SHA1 Message Date
Teknium 624a752e79 refactor(agent/runtime): deadline/file_safety/estop/lifecycle leaf modules — dedupe result plumbing, drop dead guard code
- deadline: _result()/_abandon() replace 7 BoundedResult constructions and 2
  cancel+callback sites; timers handled as a list; dead raise_if_timed_out removed.
- file_safety: retired classify_cross_profile_target (0 refs; get_cross_profile_warning
  stub kept for external callers), _home_and_resolved/_mirror_warning shared by
  the sandbox/container mirror guards; _find_sandbox_mirror_segments inlined.
- estop: _hermes_home/_canonical_root now the file_safety helpers;
  _reset_log_state_for_tests inlined into its only test.
- subagent_lifecycle: _validate_request driven by _UNSUPPORTED_REQUEST_FIELDS table.
- turn_liveness: dead start()/_abort_message removed, _emit_warning shared.
- process_bootstrap: _enable_happy_eyeballs reused by the client variant.
- Comment/docstring compaction across the remaining leaf modules.
2026-09-02 13:29:48 -07:00
HexLab98 85bc25c949 fix(terminal): bound env.execute wait so a wedged poll cannot disable every timer
A hung terminal wait on the loop thread silently disabled asyncio deadlines
and let cron jobs idle thousands of seconds past HERMES_CRON_TIMEOUT. Drive
the wait from run_bounded_sync (sliced Event.wait, kill-on-timeout) and
move the cron inactivity monitor onto a daemon thread with the same kernel
timeout primitive. Copy the caller ContextVar scope and activity callback
onto the wait worker so profile secrets, session id, and heartbeats survive
the thread hop (#94285).
2026-08-31 10:42:39 -07:00
nftpoetrist 08b4875f4a fix(deadline): remove the dead second SuspectableBackend class shadowing the Phase 3a Protocol
agent/deadline.py defined SuspectableBackend twice: the Phase 3a Protocol
(sync ensure_healthy(self) -> bool) and, further down the same module, an
unrelated concrete class with the same name (async
ensure_healthy(self, timeout=5.0)) added later by the MCP Phase 3b adopter.
Since Python executes class statements top-to-bottom, the second definition
silently shadowed the first at module scope.

Nothing in the tree imports or subclasses either by name today — the MCP
adopter duck-types the same-shaped contract directly on its own connection
class rather than referencing agent.deadline.SuspectableBackend — so this
caused no live behavior change. But it left the wrong (and differently
shaped) class resolvable under that name for the next Phase 3b adopter that
does import it for a type hint.
2026-08-27 17:22:02 +05:30
kshitijk4poor dcfdc8deec fix(deadline): document the inline-mark contract; pin the ordering invariants (Phase 3a salvage round)
Record correction: the previous commit's message says the async flavor
offloads mark_suspect via asyncio.to_thread — it does NOT (and must not).
The mark is deliberately inline on the event loop: running it
synchronously guarantees mark-happens-before-BoundedResult-return and
mark-before-on_abandon-cleanup (cleanup is ensure_future'd and cannot
start until the next loop tick). An offloaded mark would race both.
The trade-off is that a slow adopter mark_suspect would block the loop
(measured: a 2s mark stalls every coroutine for 2.003s), so the adopter
contract is now explicit in the Protocol docstring and at the async call
site: mark_suspect must be cheap, non-blocking, lock-free; expensive
recycle work belongs in ensure_healthy.

New pins so the negotiated semantics can't silently regress:
- test_sync_mark_happens_before_on_timeout (the review-round ordering)
- test_async_mark_happens_before_on_abandon_cleanup (the scheduling
  invariant an offloaded mark would break)
- test_sync_completion_never_marks_backend (sync counterpart of the
  async completion test)
2026-08-26 17:47:22 +05:30
Ayush Nangia 9ee2744097 fix(deadline): Phase 3a review round — mark ordering, loop offload, annotation
- mark_suspect runs BEFORE owner cleanup in both flavors (the reason
  describes the state at timeout; a recycling cleanup never poisons the
  healed replacement)
- the async flavor offloads the mark off the event loop
  (asyncio.to_thread), matching how owner cleanup is scheduled
- the protocol documents the synchronous-cheap contract for adopters
- the windows-footgun annotation stays on its matched killpg line
2026-08-26 17:47:22 +05:30
Ayush Nangia 7a3aaf0143 feat(deadline): SuspectableBackend protocol — mark timed-out backends suspect
Phase 3a of the #85125 unified-deadline plan. run_bounded_async and
run_bounded_sync accept backend= and call mark_suspect(label +
timeout) exactly once on timeout, never on completion. The layer
fails open: backends without the protocol (incremental Phase 3b
adoption) and raising mark_suspect implementations can never weaken
the deadline bound or corrupt the BoundedResult.
2026-08-26 17:47:22 +05:30
kshitijk4poor 2f33833de8 fix(mcp): recover poisoned connections + fail fast on dead stdio transports (#85125 3b)
Fixes the four poisoned-connection classes (#81051, #77765, #84132,
#81995) with the SuspectableBackend cheap-mark/lazy-verify contract:

- mark_suspect/ensure_healthy protocol (agent/deadline.py): noticing a
  poisoned state never does I/O; the NEXT caller pays once for a health
  probe that clears the suspicion or forces a reconnect. A single
  teardown-vs-keepalive race or auth-lock corruption can no longer park
  a connection permanently — park stays reserved for genuinely
  exhausted reconnect budgets.
- keepalive failure marks the connection suspect before requesting
  reconnect; the next tool call probes and recycles if unhealthy.
- auth-classified permanent failures on a previously-proven session get
  a suspect+reconnect path instead of an immediate park.
- fast-fail (#81995): stdio child pids are tracked at spawn and an
  in-flight RPC races a child-watcher task, so a dead subprocess fails
  the call immediately with a retryable timeout instead of riding out
  the full 300s. Deliberate teardown/reconnect also fails in-flight
  calls now instead of leaving them attached to a dying transport.

Dispatch-boundary hardening for test doubles: stubbed sessions
(MagicMock/non-awaitable call_tool, absent child-watcher) fall back to
the exact pre-change inline-await semantics, so only real transports
gain the race guard.

Salvage credit: in-flight approach from #73377 (@luijoc, wedged
transport recovery) and #48069 (@arminanton, keepalive/in-flight
interaction); both PRs' bases predate main's current park/reconnect
architecture, so this is a fresh implementation of their contracts.

Tests: tests/tools/ -k mcp = 639 passed (was 22 new failures during
development; final tree zero).
2026-08-25 01:34:55 +05:30
kshitij 8b387962ac fix(agent): suppress windows-footgun lint on POSIX-only killpg branch
The os.killpg call sits below an early 'if sys.platform == win32: return'
so it can never execute on Windows; the scanner is line-based and needs
the inline marker.
2026-08-13 13:53:39 +05:30
kshitij 8ac9ff18ae fix(agent): harden deadline layer per self-review
- run_bounded_async: cancel + abandon the inner task when the CALLER is
  cancelled (leak the telegram original also had)
- kill_process_tree: check taskkill exit code (Windows contract parity),
  suppress console flash via windows_hide_flags, and sweep a psutil
  descendant snapshot taken before signalling — reaches grandchildren in
  their own setsid sessions and the non-group-leader case (#71148 class)
- resolve_timeout: reject bool (YAML true would become a 1s deadline) and
  NaN config values with fall-through instead of resolving unbounded
- BoundedResult: kw_only to prevent positional transposition
- tests: real clamped-value time_t regression proof, own-session
  descendant kill, external-cancellation task cleanup, bool/NaN config
  fall-through; pin already-dead-pid contract
2026-08-13 13:46:52 +05:30
kshitij 083f8a6071 feat(agent): unified deadline layer — bounded execution primitive + timeout resolver (#85125 Phase 1)
One shared foundation for the timeout/hang backlog instead of per-incident
site-local fixes:

- agent/deadline.py: run_bounded_async (thread-timer deadline that survives
  a blocked event loop, generalizing the telegram adapter primitive),
  run_bounded_sync, clamp_timeout (kills the #83220 time_t OverflowError
  class at the boundary), resolve_timeout (config.yaml timeouts: section >
  legacy env bridge > default), kill_process_tree (whole-tree termination
  for the #71148 orphan class), DeadlineExpired (our deadline, mechanically
  distinct from provider timeouts).
- tool_executor._resolve_concurrent_tool_timeout migrates onto the resolver;
  exact legacy env-var contract preserved (default 420, 0 disables).
- timeouts: accepted as a known config root; documented in
  cli-config.yaml.example.

Pure addition otherwise — no behavior change, no new env vars, no cache
impact. Later phases (#85125) migrate tool-execution, MCP, and subprocess
call sites onto these primitives.
2026-08-13 13:33:22 +05:30