Commit Graph

13 Commits

Author SHA1 Message Date
Teknium 25d954c2cf fix(guardrails): hard stops catch replays, never legitimate iteration
Before turning hard stops on for unattended platforms, make sure they cannot
cut off normal work:

- Edit -> re-run is progress. A successful mutating call (write_file/patch,
  a green terminal/execute_code, browser actions, job/message/cron/memory/
  skill mutations) marks progress for every failing signature still being
  counted this turn; the next identical retry restarts its streak instead
  of accumulating toward exact_failure_block_after. A pure replay never
  mutates anything between attempts, so it is still blocked at 5.
- Distinct red commands are diagnosis. For FAILURE_TOLERANT_TOOL_NAMES
  (terminal, execute_code, process pollers, browser_navigate, web_extract)
  same_tool_failure_halt_after warns but never halts.
- subagent and api_server keep the warn-only default: both are supervised
  task loops with a live parent/client and do real edit -> re-run work.

Live A/B (real AIAgent platform=telegram, real patch+terminal, 8 rounds of
patch -> red check -> patch ...):
  unmitigated branch: HALTED at round 6 (repeated_exact_failure_block)
  this commit:        COMPLETED all 8 rounds, final answer delivered
Loop shapes still stopped: identical failing read_file 8 calls,
identical successful terminal 5 calls (vs 602 on main).
Six new tests pin these flows; all fail on the unmitigated version.
2026-09-02 00:26:57 -07:00
Teknium 76648a7faf fix(guardrails): identical-call streaks hard-stop any tool on unattended platforms
Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.

- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
  streak raises a halt (identical_call_streak_halt) at
  hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
  stay exempt; a changed result resets the streak; warning-only sessions are
  unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
  every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
  for pollers, or when results change.

Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
  identical failing read_file   main: 602 API calls, budget exhausted
                                branch: 8 calls, repeated_exact_failure_block
  identical successful terminal main: 602 API calls, budget exhausted
                                branch: 5 calls, identical_call_streak_halt
2026-09-02 00:26:57 -07:00
benbenwyb cd2d3089fb fix(agent): guard repeated skill reads
Treat skill_view and skills_list as idempotent read-only tools so the existing no-progress guardrail can warn or block repeated identical skill loads. This prevents large skill outputs from being re-added to the context in tool loops.

Add regression coverage for repeated skill_view results under hard-stop guardrails.
2026-09-02 00:26:57 -07:00
João Vitor Cunha 384fc4bf83 fix(guardrails): preserve interactive platform defaults 2026-09-02 00:26:57 -07:00
João Vitor Cunha ee2147f9e6 fix: hard stop tool loops on non-interactive platforms 2026-09-02 00:26:57 -07:00
Teknium 6b81590c55 test: prune low-value tests suite-wide (wave 1) — 46,820 → 28,106 test functions
Systematic prune per AGENTS.md test policy, one pass over every major
test tree (gateway, hermes_cli, tools, agent, run_agent, plugins, cli,
cron, tui_gateway, honcho/openviking, root-level):

- DELETE: source-reading tests (read_text/getsource on prod files),
  change-detector tests (exact catalog counts, model-name snapshots,
  config version literals), mock-echo tests (assert a mock returns what
  it was told), assertion-free/trivial tests, near-duplicate
  parametrizations (boundaries + one representative kept), async/sync
  twin duplicates, cosmetic within-file variations.
- KEEP (mandatory): security/redaction/approval guards, message-role
  alternation invariants, prompt-caching/deterministic-call-id
  invariants, issue-number regression tests (deduped), E2E tests.
- 6 test files deleted outright (script-style/no-assert or fully
  redundant); conftest.py, fakes/, fixtures/ untouched.
- tests/acp/conftest.py added: autouse fixture stubs the live
  models.dev/GitHub/Copilot/Anthropic inventory fetches that ACP server
  tests performed on every session create — test_server.py 147s → 3.4s,
  and the tests are now genuinely hermetic.
- Sleep-based slowness shrunk where safe (codex_ttfb_watchdog,
  compression_concurrent_fork, etc.); no wall-clock assertion tightened.

Verification: full hermetic suite via scripts/run_tests.sh —
2439 files, 31,130 tests passed, 0 failed, 0 flaky retries, 315s wall
(baseline: 583s wall, 13,564s subprocess CPU).
2026-07-29 13:10:23 -07:00
teknium1 cb06017b1d refactor(guardrails): make runaway-loop caps per-turn, not session-total
Per Teknium: the caps should bound a single agent loop, not accumulate
over the whole session. Rename SessionCapConfig -> LoopCapConfig and the
config section session_caps -> loop_caps; move the counters into
reset_for_turn (invoked per turn via turn_context) so each turn starts
with a fresh budget; retune defaults 200 -> 50 (a single turn issuing 50
web searches / spawning 50 subagents is already pathological). Block
codes session_*_cap -> loop_*_cap and messages updated to drop the
/new-resets-the-budget guidance (irrelevant now that it resets per turn).
Tests flipped: the old persists-across-turn-resets assertion becomes
resets-each-turn.
2026-07-26 21:27:45 -07:00
teknium1 b68787ad25 Inspired by Claude Code: session-wide runaway-loop caps for web_search and delegate_task
Add per-session lifetime caps on web_search calls and subagent spawns
(defaults 200/200, matching Claude Code v2.1.212). Unlike the existing
per-turn tool-loop guardrails, these count over the whole session and
reset only when a fresh agent is built (/new, /clear). Hitting a cap
blocks the offending call and halts the turn cleanly.

- agent/tool_guardrails.py: SessionCapConfig + session counters on the
  controller (in __init__, not reset_for_turn, so they persist across
  turns). before_call() enforces caps first, independent of
  hard_stop_enabled. delegate_task batches count each task.
- hermes_cli/config.py: tool_loop_guardrails.session_caps defaults.
- docs + tests (unit + E2E validated against a real AIAgent).
2026-07-26 21:27:45 -07:00
Juniper Bevensee fb0217c656 fix(agent): tolerate lone UTF-16 surrogates in tool-guardrail hashing
Tool results scraped from the web/social platforms can carry unpaired
UTF-16 surrogates (e.g. half of a mathematical-bold character pair).
_sha256() did a strict utf-8 encode, which raises UnicodeEncodeError on
that input and took down the whole conversation loop — the hash only
needs deterministic bytes, not valid UTF-8, so encode with
surrogatepass instead.
2026-07-18 02:08:39 -07:00
ooovenenoso d759a67c0f fix: add recovery hints to loop guard warnings 2026-05-19 00:12:12 -07:00
GodsBoy da0ddbf88a fix: classify landed file mutations with diagnostics 2026-05-13 06:46:23 -07:00
Mind-Dragon 0704589ceb fix(agent): make tool loop guardrails warning-first 2026-04-30 20:43:15 -07:00
Mind-Dragon 58b89965c8 fix(agent): add tool-call loop guardrails 2026-04-30 20:43:15 -07:00