Commit Graph

13 Commits

Author SHA1 Message Date
Teknium 62f353eed5 refactor(gateway): iteration-progress formatter lives in agent/session_activity, tests trimmed to the invariant bar
The salvaged helper was appended to the gateway/run.py facade; new behaviour
belongs in a topical sibling. agent/session_activity.py already owns the
activity-snapshot contract the three render sites read from, so the formatter
moves there as format_iteration_progress and the three call sites import it at
module level instead of late-importing the facade.

Tests trimmed to the salvage bar (<= 2 invariant tests per fix): one
parameterized contract on the helper (unbounded/None hide the ceiling, a real
budget keeps N/M) plus the contributor's end-to-end busy-ack test through the
real render path. The per-site heartbeat/timeout tests exercised the same
helper through mocks and are dropped; evals/gateway_status_render/
iteration_ceiling_ab.py drives all three real render sites with a real AIAgent
for before/after evidence (3 sentinel leaks on main -> 0).

Related: #103109 (same fix, same target module; credited in the PR),
#102845 (third render site, closed by its author in favour of #102817).
2026-09-06 05:35:12 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 63f1950b4f refactor(agent/E_session): by-hand docstring/comment compaction across 9 files; transcript_repair row lookup helper; activity snapshot literal 2026-09-02 22:23:45 -07:00
Teknium 334b72256a refactor(agent/E_session): contextlib.suppress for swallow-only try/except 2026-09-02 19:27:20 -07:00
Teknium 2a71fd4e03 refactor(agent/replay_cleanup,trace_upload,transcript_repair,session_activity,trajectory): collapse duplicate branches, compact docstrings 2026-09-02 18:31:56 -07:00
Teknium 55864960bd refactor(agent/E_session): AST-identical line hugging of wrapped signatures/conditions 2026-09-02 18:27:42 -07:00
Teknium 624a752e79 refactor(agent/runtime): deadline/file_safety/estop/lifecycle leaf modules — dedupe result plumbing, drop dead guard code
- deadline: _result()/_abandon() replace 7 BoundedResult constructions and 2
  cancel+callback sites; timers handled as a list; dead raise_if_timed_out removed.
- file_safety: retired classify_cross_profile_target (0 refs; get_cross_profile_warning
  stub kept for external callers), _home_and_resolved/_mirror_warning shared by
  the sandbox/container mirror guards; _find_sandbox_mirror_segments inlined.
- estop: _hermes_home/_canonical_root now the file_safety helpers;
  _reset_log_state_for_tests inlined into its only test.
- subagent_lifecycle: _validate_request driven by _UNSUPPORTED_REQUEST_FIELDS table.
- turn_liveness: dead start()/_abort_message removed, _emit_warning shared.
- process_bootstrap: _enable_happy_eyeballs reused by the client variant.
- Comment/docstring compaction across the remaining leaf modules.
2026-09-02 13:29:48 -07:00
Machan-Army 31b974dfb7 fix(gateway): split turn-hold expiry from idle-timeout failure path
Introduce HygieneTurnHoldExceeded exception so turn-hold budget expiry
no longer collapses into the generic asyncio.TimeoutError handler.

- Add HygieneTurnHoldExceeded exception (availability boundary, not a failure)
- Add dedicated handler: stamps AGENT_COMPRESSION_TURNHOLD provenance,
  sends deferral notice, does NOT increment failure cooldown
- Preserve #87011 contract: idle timeout still sends 'no output' message
  and takes failure path
- Add behavior witnesses: turn-hold ≠ idle timeout, cooldown untouched

Fixes semantic boundary collapse flagged in PR #90845 review.
2026-08-29 13:02:13 +05:30
Teknium 024a58ddea refactor(agent): pin session activity heartbeat cadence + harden best-effort write
Heartbeat write discipline for the durable SessionDB activity projection:

- Pin the cadence in a named constant
  (SESSION_ACTIVITY_HEARTBEAT_MIN_INTERVAL_SECONDS = 60s, contract >= 30s,
  deliberately config-independent so no compression.*/agent.* setting can
  turn the heartbeat into a high-frequency writer on the contended
  SessionDB write path).
- The write already rides the standard _execute_write patience path via
  SessionDB.touch_session_activity — verified, now documented in the
  docstring.
- Best-effort hardening: a failed heartbeat write never raises into the
  agent loop; the bare 'pass' becomes an explicit debug log with traceback.
- Tests: direct proof that a heartbeat DB failure doesn't propagate,
  the cadence constant is pinned >= 30s, and the rate limiter keys off
  the shared constant (boundary tested on both sides of the window).
2026-08-02 16:16:36 -07:00
fangliquanflq 240148b440 fix(agent): force-persist compression completed past SessionDB rate limit
{id: #72016}
2026-08-02 16:16:36 -07:00
fangliquanflq c2088efe9e feat(gateway): session activity watchdog, stall notify, compress timeout (#72424)
Three mechanisms to detect and notify when gateway sessions stall silently:

1. Mid-turn activity heartbeats stamped to SessionDB so hermes sessions list
   and hermes status show progress during long turns without new message rows.

2. Stall watchdog: when a busy session has pending inbound and the shared
   activity clock is idle past agent.session_stall_timeout (default 300),
   log a WARNING and notify the user once to try /new. Notify-only; does
   not kill the turn.

3. Compaction timeout: fenceless compress_context callers get a progress-aware
   host budget (compression.context_timeout_seconds default 120 idle,
   compression.context_total_ceiling_seconds default 600 ceiling). On timeout,
   cancel via commit fence, skip compaction without dropping messages, and
   continue the turn.

Closes #72016 (slices 1-3; slice 4 cumulative SSE stream-retry deadline
remains a follow-up).

Cherry-picked from PR #72424 by @fangliquanflq.
2026-08-02 16:16:36 -07:00
kshitijk4poor 2c1809e6ca revert: PR #72817 — session activity watchdog, stall notify, compress timeout
Reverting #72817 (salvage of #72424) pending further review.
All 4 commits reverted: feat, refactor, chore (contributor map), CI fix.
2026-07-28 00:15:00 +05:00
fangliquanflq cfb206fe2e feat(gateway): session activity watchdog, stall notify, compress timeout (#72424)
Three mechanisms to detect and notify when gateway sessions stall silently:

1. Mid-turn activity heartbeats stamped to SessionDB so hermes sessions list
   and hermes status show progress during long turns without new message rows.

2. Stall watchdog: when a busy session has pending inbound and the shared
   activity clock is idle past agent.session_stall_timeout (default 300),
   log a WARNING and notify the user once to try /new. Notify-only; does
   not kill the turn.

3. Compaction timeout: fenceless compress_context callers get a progress-aware
   host budget (compression.context_timeout_seconds default 120 idle,
   compression.context_total_ceiling_seconds default 600 ceiling). On timeout,
   cancel via commit fence, skip compaction without dropping messages, and
   continue the turn.

Closes #72016 (slices 1-3; slice 4 cumulative SSE stream-retry deadline
remains a follow-up).

Cherry-picked from PR #72424 by @fangliquanflq.
2026-07-28 00:44:02 +05:30