Commit Graph

330 Commits

Author SHA1 Message Date
Teknium b3e41128ff refactor(agent): share predicates, truncation and the timeout ladder
- User-turn predicates reuse _is_context_summary_message (single detection path).
- _serialize_for_summary hoists the shared head/tail truncation.
- __init__ initialises per-session state via _reset_session_compaction_state.
- Timeout cooldown ladder (60/300/900s) is one module helper used by both
  record_timeout_failure and the summary failure path; module-level so a real
  method bound onto a stub compressor still exercises it.
2026-09-02 13:29:38 -07:00
Teknium b2717f4f4c refactor(agent): restore rationale comments lost in the docstring compaction
Hunk-by-hunk audit of the compaction commit found 14 dropped WHY/invariant
statements (route-pin scope, persistence-marker sweep failure mode, stamp
asymmetry, reasoning-replay boundary, size-class parity with preflight,
small-context floor cap, tail-anchor ordering rules, defer-baseline backstops).
Each is restored as a compact 1-3 line comment at its anchor.
2026-09-02 13:29:38 -07:00
Teknium ea090ccb62 refactor(agent): unify durable-state helpers, session reset and tail-cut walks
- Nine near-identical durable-row getters/setters -> _durable_read/_durable_write
  (also used for parent-lineage lookups in on_session_start).
- on_session_reset/on_session_end share _reset_session_compaction_state.
- _find_tail_cut_by_tokens: the two backward token walks become one _walk_back
  (randomized parity vs. the previous version: 1500 transcripts identical).
- _strip_context_summary_handoff_message: one _unwrapped() helper for all paths
  (parity: 1296 content shapes identical).
- ruff SIM102/SIM108/SIM114 collapses.
2026-09-02 13:29:38 -07:00
Teknium 4a64d42f9b refactor(agent): extract micro-compaction into agent/micro_compaction.py
The 15 rolling micro-compaction methods (cursor resolution, exchange
serialization, one-exchange summarization, defrag, splice, DB sync, telemetry)
move verbatim into MicroCompactionMixin, mixed into ContextCompressor
(MRO: ContextCompressor -> MicroCompactionMixin -> ContextEngine). Origin
constants and call_llm are imported lazily so tests patching
agent.context_compressor.X still apply; logger name parity is kept.
2026-09-02 13:29:38 -07:00
Teknium 64576a34f8 refactor(agent): decompose _generate_summary and the compress() abort path
- _generate_summary(): prompt assembly lifted into _build_summary_prompt, the
  exception path into _on_summary_failure; the two identical main-model retry
  branches are merged behind a reason lookup.
- compress(): terminal summary-failure abort classes are table-driven via
  _TERMINAL_SUMMARY_FAILURES (same precedence, log text and telemetry class;
  flag reads stay lazy so __new__-built test doubles keep working).
2026-09-02 13:29:38 -07:00
Teknium 8b84f8debe refactor(agent): dispatch-table the tool-result summarizer
_summarize_tool_result_unguarded routed on tool_name through ~20 if-branches; each
branch is now a small _sum_* function in _TOOL_RESULT_SUMMARIZERS with the generic
fallback unchanged. Output verified byte-identical against the previous version
across 1288 name/args/content combinations.
2026-09-02 13:29:38 -07:00
Teknium cfb234c54c refactor(agent): extract _scan_window_handoffs from compress(); drop unused imports
The in-transcript handoff rehydration scan (previous-summary rebuild, zero-user
provenance, merged-handoff unwrap, tail_start advance) moves into
ContextCompressor._scan_window_handoffs returning a _HandoffScan dataclass;
compress() keeps the rollback snapshot semantics unchanged.
2026-09-02 13:29:38 -07:00
Teknium 12be763de0 refactor(agent): compact context_compressor comments/docstrings to invariants
Hand-reviewed compaction: non-obvious rules, ordering constraints, invariants,
failure modes and WHY are kept in 1-3 line form; narrative, issue-number history
and code restatements are dropped.
2026-09-02 13:29:37 -07:00
Teknium d1efa0d78d fix(compression): provider-proven overflow gets one real compaction attempt while the failure cooldown is armed
After one failed/stalled summary attempt arms the 60/300/900s compression-
failure cooldown, a provider context_length_exceeded rejection entered the
reactive overflow branch in conversation_loop, which called _compress_context
without force. Since #97488 the cooldown gate returns the soft "temporarily
paused, retry in a moment" deferral instead of exhaustion, so every turn
deferred until the cooldown lapsed, and the next failure extended the ladder:
long-running sessions wedged with no automatic recovery (#100661, four sessions
lost).

Thread a narrow `bypass_cooldown` kwarg from the three provider-proven overflow
call sites (generic overflow, 413, output-cap recovery) through
AIAgent._compress_context -> compress_context -> ContextCompressor.compress ->
_generate_summary. It skips ONLY the summary-failure cooldown check at each gate.
Unlike force=True it does not clear the cooldown, does not skip the feasibility /
anti-thrash breakers, and a failed attempt records its cooldown normally. The
attempt is bounded by the existing compression_attempts/max_compression_attempts
budget, so there is no retry loop. The preflight threshold gate is unchanged:
ordinary over-threshold pressure still honors the cooldown (#11529).

Engines whose _automatic_compression_blocked()/compress() predate the kwarg
(plugins, test doubles) are called with the legacy signature.

Tests: cooldown armed + bypass_cooldown -> summarizer invoked and transcript
compacted; ordinary pass still deferred. Docs note the cooldown/overflow
contract in the developer guide.

Fixes #100661
Closes #97766 (overflow-force idea; the bundled continuation changes were not taken)

Co-authored-by: sgtworkman <178342791+sgtworkman@users.noreply.github.com>
2026-09-02 05:33:22 -07:00
Teknium 238b6c1ab9 fix(compression): persist the anti-thrash recovery deadline so gateway agent rebuilds cannot block a session forever
The #14694 recovery clock (`_anti_thrash_recovery_deadline`) was a
process-local `time.monotonic()` value zeroed in `bind_session_state()`.
The gateway rebuilds the AIAgent (and its ContextCompressor) on every
cache eviction, so each fresh compressor bound to a durably tripped
session row (#69872) re-armed a full 300s window and the half-open probe
never fired — a long messaging conversation above the threshold stayed
blocked permanently.

Persist the deadline as a wall-clock epoch in a new
`sessions.compression_recovery_deadline REAL` column (declarative column
reconciliation; SCHEMA_VERSION 26 -> 27) with
`SessionDB.get/set_compression_recovery_deadline`. The compressor loads it
in `bind_session_state()` and writes it on change only via
`_set_anti_thrash_recovery_deadline()`. A fresh compressor with no stored
deadline still starts a full window blocked (#54923 restart contract); one
that loads an armed deadline resumes that window. Backward clock jumps are
bounded to one window. The 300s window is unchanged.

Minimal salvage of #100185 (the probe-lease/fencing state machine and
model_config-blob storage were not carried).

Refs #100185
Co-authored-by: Komzpa <me@komzpa.net>
2026-09-02 04:14:10 -07:00
Teknium c5c9aa8d44 fix(gateway): hygiene compaction keeps the profile secret scope under multiplexing
Session-hygiene compaction ran _compress_context on a bare
loop.run_in_executor(None, ...) worker. Under gateway.multiplex_profiles the
profile secret scope and HERMES_HOME override are ContextVars installed by
the per-turn _profile_runtime_scope, and a bare worker starts with an empty
Context — so the summary model's get_secret(<PROVIDER>_API_KEY) failed
closed with UnscopedSecretError on EVERY hygiene pass and compaction
silently degraded to a lossy truncation (#100849 debug bundle:
'Failed to generate context summary: get_secret(SURPLUS_API_KEY) called
with no profile secret scope active').

- gateway/run.py: run both hygiene executor hops (detached-agent path and
  codex app-server path) inside copy_context().run, keeping the default
  executor so a fence-cancelled hung summary never occupies a gateway
  agent-work slot.
- agent/context_compressor.py: UnscopedSecretError is a missing-credential
  class failure — abort and preserve the session instead of dropping the
  middle window for a placeholder summary (same carve-out as 401/402/403).
- tools/daemon_pool.py: correct the salvaged docstrings — stdlib
  ThreadPoolExecutor only propagates contextvars from 3.14; nothing is
  stripped from the bundled runtime.
- tests: hygiene worker inherits caller ContextVars (fails on bare
  run_in_executor); UnscopedSecretError classified as access failure.

Live A/B (real get_secret in a run_in_executor worker, multiplex on, profile
.env scope installed): main -> UnscopedSecretError; fixed -> scoped value.
2026-09-01 22:28:52 -07:00
Teknium bd7cdd7c53 Merge origin/main into core-tool-deferral (resolve show_tip test seam onto the check_tips_enabled gate) 2026-09-01 21:49:14 -07:00
kshitijk4poor 29d4c0ebfd refactor(compression): extract preflight seed predicate onto ContextCompressor
Address review follow-ups on the seed fix:

- Move the 'seed only from the 0 state' guard from an inline block in
  build_turn_context() into
  ContextCompressor.maybe_seed_preflight_display_tokens(), co-locating
  the predicate with the rest of the speculative-seed lifecycle
  (snapshot_preflight_display_tokens /
  rollback_interrupted_preflight_display_tokens). Callers now use the
  method via a getattr guard so test doubles and external context
  engines without it are unaffected.
- Rewrite the TestPreflightSentinelGuard docstring, which still
  described the old >=0 guard ('treats any negative value as no real
  usage yet'); the ==0 policy protects ALL non-zero readings.
- Drop the _seed mirror-helper: the tests now call the real production
  method on the compressor fixture, eliminating mirror-drift risk (the
  helper comment had already drifted once).
- Note the accepted trade-off (partial-usage providers pin the meter
  low until their next report) in the method docstring.
2026-09-01 14:33:23 +05:30
Brooklyn Nicholson 5c6dbe22c3 fix(compression): take reasoning_content when the summarizer leaves content empty
Local and thinking backends (DeepSeek, Qwen, Kimi) often return a usable
summary in reasoning fields. Treat that as the summary instead of burning
another 100s+ empty-content retry. Leave the wire max_tokens omit intact.

Co-authored-by: Chris DePuy <chris@650group.com>
Co-authored-by: chenhm <chenhm@yuancheng.local>
2026-08-31 20:43:00 -05:00
theo 8207862212 fix(compression): stop timeout paths from blocking retries 2026-08-31 13:00:33 -07:00
Darafei Praliaskouski 4252aecc2e fix(agent): cap compaction threshold floor at 85% of the context window
The MINIMUM_CONTEXT_LENGTH floor in _compute_threshold_tokens only
degraded to the 85% trigger when it met or exceeded the effective
window exactly (#14690). Near-minimum windows slipped through: at
context_length=65536 the threshold passed through at 64,000 — 97.7%
of the window, ~1.5K tokens of output room — so pre-API compaction
effectively could not fire.

Providers that silently truncate over-window prompts instead of
rejecting them (e.g. ollama's OpenAI-compatible /v1 endpoint) never
deliver the reactive context-overflow backstop either. Observed live
on a 65,536-token local model: the session rode into the window
ceiling and each length-continuation retry re-sent a window-filling
prompt (65,120 -> 65,273 prompt tokens, 263 output tokens of room)
until the turn died with "Response remained truncated after 4
continuation attempts" — every retry paying a full multi-minute
prefill.

Cap the floored threshold at _MIN_CTX_TRIGGER_RATIO (85%) of the
effective input budget whenever the floor is the binding term. An
explicit threshold_percent above 85% is user intent and stays
uncapped; windows where the floor lands at/below the cap are
unchanged.
2026-08-31 12:19:29 -07:00
Teknium d8f8a07ee3 fix(compression): truncated summaries no longer become compaction checkpoints (port of earendil-works/pi#7048)
A summarization response with finish_reason == "length" contains PARTIAL
text — the generation stopped on the output-token cap mid-summary.
Previously all compressor summarization sites accepted such responses as
complete: the cut-off text replaced the real middle turns AND was fed back
into every subsequent iterative-update prompt, compounding the loss across
compactions.

Guards added at all four summarization sites (whole bug class):
- _generate_summary: length stop raises, gets the existing one-shot
  main-model fallback (a larger output budget may finish the summary), and
  on terminal failure ABORTS compression preserving the session unchanged
  (new _last_summary_truncated_failure flag, same class as empty-content).
- _micro_summarize_one: partial rolling-summary merge is discarded; the
  exchange stays unabsorbed for a later pass.
- _build_chunk_digests: partial lean digest degrades to the
  recover-via-session_search placeholder.
- trajectory_compressor (sync + async): length stop raises into the
  existing retry/backoff loop.

_response_finish_reason() reads dict- and object-shaped responses and
returns "" when the provider omits the field, so proxies that never send
finish_reason are unaffected.

Ported from earendil-works/pi commit 97fa14e39 (pi#7048), adapted to
hermes' abort-preserving compression failure machinery.

Tests: tests/agent/test_compressor_truncated_summary_guard.py (12 tests;
sabotage-verified — disabling the guards fails 4).
2026-08-31 10:39:55 -07:00
BrunoBza 9d9d9194d4 fix(state): archive carried-forward compaction tail as rewind rows (#86366)
archive_and_compact() soft-archives every active row with compacted=1 and
then re-inserts compacted_messages as fresh live rows. When the
compressor's protected tail rides inside that list verbatim - which is
the normal batch-compaction shape ([summary] + tail) - the tail's
ORIGINALS end up stored twice per compaction: (active=0, compacted=1)
next to their live clones. search_messages() recalls both flags without
DISTINCT, so every carried-forward message came back once per compaction
(measured up to 4 identical hits) and was mislabeled to users and the
agent as archived "summarized away" content.

Add an optional tail_count parameter: the last tail_count archived rows
are superseded byte-identical duplicates, stamped rewind-style
(active=0, compacted=0, hidden from recall) instead of compacted=1.

Callers:
- batch in-place compaction counts the compressor-tagged tail dicts
  (_COMPACTION_TAIL_MARKER set by compress() on every carried-forward
  message);
- micro-compaction splices [prefix, marker, suffix] - everything except
  the single marker row is carried forward, so tail_count=len-1;
- proactive tool-result pruning rewrites content in place (not verbatim),
  keeping the historical archive-everything behavior.

Fixes #86366
2026-08-31 10:02:22 -07:00
Teknium 452f6b7de2 fix(compression): route-aware stale-thinking charge parity between compaction trigger and tail walks (#84371)
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).

Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.

Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
2026-08-30 20:40:43 -07:00
Teknium 6101f52ba4 Merge remote-tracking branch 'origin/main' into core-tool-deferral 2026-08-30 19:47:30 -07:00
Teknium a6549922b8 fix(compression): stamp durable backoff rows with strategy and failure kind (#96775 #97488)
record_timeout_failure() now persists
'backoff:<failure_kind>:strategy=<tail_mode>' into the state.db
cooldown row (sessions.compression_failure_cooldown_until +
compression_failure_error), so a failed/stalled/cancelled attempt's
identity survives gateway restarts and the rebuilt compressor makes the
same skip decision via bind_session_state()/get_active_compression_
failure_cooldown(refresh=True). Host callers pass ceiling_exhausted /
stalled; the stall-interrupt path passes stall_interrupted. A
successful compression still clears the row.
2026-08-30 19:46:21 -07:00
HexLab98 027b339e79 fix(compression): persist stall-interrupted backoff on pre-commit cancel (#96775)
An explicit /stop after the summary stream has already gone idle restored the original transcript but left no durable cooldown, so the next automatic turn re-entered the same stalled strategy. Record a stall-specific failure on that path only, merge with any longer live deadline, and keep an ordinary early /stop cooldown-neutral.
2026-08-30 19:46:21 -07:00
fangliquanflq 8d567ccd22 fix(compression): stop work at the total deadline 2026-08-30 19:46:21 -07:00
Teknium 1f2bd9e763 fix(compression): stamp _DB_PERSISTED_MARKER after in-place batch compaction commit (#98450)
compress() returns marker-swept copies (_strip_persistence_markers, #57491);
the in-place branch committed them via archive_and_compact() but never
stamped the persistence marker, so the next _persist_session ->
_flush_messages_to_session_db_unlocked walk re-INSERTed the whole
post-compaction transcript (live set regrew ~58K -> ~512K tokens).

Centralize the post-commit contract in a shared helper,
stamp_db_persisted_markers(), used by all three archive_and_compact
callers: the in-place batch commit (previously missing), the
micro-compaction sync, and the proactive tool-result prune.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-08-30 19:46:11 -07:00
Teknium 4f22543509 fix(compression): lean compaction makes exactly one auxiliary request per attempt
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.

- The detailed session log is folded into the single summary request: the
  lean prompt template gains a '## Detailed Session Log (oldest first)'
  section carrying the digest prompt's HARD RULES (identifiers verbatim,
  dense bullets, transcript-is-data). Output guidance grows by
  _LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
  summary budget — the old worst case (28 x 1,400 digest tokens) was spread
  across many requests and mostly re-covered tool noise; a single dense
  4K-token log inside one response preserves the load-bearing record while
  staying well inside one aux response (the summary call still sends no hard
  max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
  whole region (_sample_summary_input: 8 proportionally spaced slices,
  oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
  anchored to the newest end) instead of head+tail truncated, so session-log
  coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
  session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
  _LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
  _LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
  sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
  attempt_summary_route_kwargs — no remaining callers; the single-use
  summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
  section lands in the summary; oversized regions sampled with elision
  markers, never a second request; anchor index + recovery footer present).
  Sabotage-verified: restoring a second call_llm makes the call-count test
  fail. Docs and the compaction eval wording updated to stop claiming
  per-chunk calls.

Fixes #96603.
2026-08-30 09:03:57 -07:00
Teknium e16ad33a9d feat(tool-search): core-tool deferral — curated 19-tool set behind the bridge by default; renames todo_list/cronjob_manage/process_manage/gui_tour/show_tip with legacy aliases (13.4K -> 6.9K desktop schemas, -49%) 2026-08-29 08:26:24 -07:00
isheng c6a426e9ad refactor(sanitizer): extract shared _classify_tool_call_orphans to eliminate drift
sanitize_api_messages (agent_runtime_helpers) and
_sanitize_tool_pairs (context_compressor) both collected
tool-call IDs and classified orphans with near-identical logic
that had already drifted: the canonical sanitizer added dedup
(#58350), but the compressor's copy did not.

Extract the shared orphan-detection logic into
_classify_tool_call_orphans(messages) in agent_runtime_helpers.
Both call sites now delegate to it, preserving their divergent
remediation strategies (insert-stubs vs strip-orphans) while
ensuring id-resolution rules and dedup stay in sync.

Closes #58357
2026-08-28 06:32:48 -07:00
srojk34 e024bf75ce fix(compression): strip whitespace from tool_call_id in _sanitize_tool_pairs
_sanitize_tool_pairs() in ContextCompressor compared raw tool_call_id
strings without stripping whitespace, the same bug fa3ab2ffd just fixed
in agent_runtime_helpers.py / run_agent.py (_get_tool_call_id_static +
sanitize_api_messages). ContextCompressor has its own near-identical
reimplementation of the pair-repair logic that was left unpatched.

When assistant-side and result-side IDs diverge only in surrounding
whitespace, the compressor misclassifies valid results as orphaned and
replaces them with [Result unavailable] stubs — silent data loss on
every compression cycle that touches such pairs.

Apply the same .strip() fix to all three sites:
- _get_tool_call_id (extracts IDs from assistant tool_calls)
- result_call_ids accumulation loop
- orphaned_results filter predicate

Closes the sibling gap of fa3ab2ffd / #42405.
2026-08-28 06:32:48 -07:00
Teknium 0dce46feb7 fix(compressor): widen compaction-time image aging to first-message and envelope shapes
Widen #90001's compaction-time strip to cover the gaps #89965 identified,
applied at compaction only per the cache ruling (request-time eviction
changes the per-call prefix and breaks prompt caching; compaction is the
one sanctioned cache break):

- Rule 1b: the opening attachment (anchor == 0) ages out once a newer
  tool-result image supersedes it. The reported session opened with a
  ~200KB poster that previously survived every compaction. The row keeps
  a non-empty text placeholder, so the zero-user-turn guard (#58753) and
  role alternation are untouched.
- Native {_multimodal: True, content: [...]} dict envelopes now both
  anchor (newest is kept) and strip (older collapse to their
  text_summary via _strip_images_from_tool_msg, which also drops the
  stale api_content sidecar per #97125's drop_stale_api_content).
- All three wire shapes (Chat Completions image_url, Responses
  input_image, Anthropic-native image) were already matched by
  _IMAGE_PART_TYPES; tests now pin each shape explicitly, plus
  determinism (double-run is a no-op returning the same object).

Refs #89938, #89965
2026-08-28 06:32:43 -07:00
Jack Lau b81b599d50 fix(compressor): age out stale tool-result images during compaction
_strip_historical_media anchors on the newest image-bearing USER message and
returns the list untouched when that anchor is index 0 or does not exist. A
session whose images arrive from tools rather than attachments therefore has
nothing to be "before": twenty vision_analyze results keep multi-MB of base64
in every request body, the provider answers 413, and the 413 handler's
recovery compaction lands right back in this function and frees nothing. The
reporter saw seven compactions in thirteen minutes, all below 200K tokens.

Age tool-result images on their own timeline: keep the newest one, since that
is the image the model is reasoning about, and strip every older one wherever
it sits, including inside the protected tail. The tail exists to preserve
conversational continuity, not to pin bytes the model has already moved past.

User-message images keep today's treatment exactly. The user anchor is
checked first, so a tool result that is the newest of its kind but still sits
before that anchor is stripped as it always has been, and the anchor message
itself is still kept byte-for-byte - test_compressor_zero_user_guard depends
on that.

Refs #89938
2026-08-28 06:32:43 -07:00
kshitijk4poor 61cd299c6e fix(compression): attempt-generation ownership for overlapping stall-fallback attempts
Follow-up to #96634 (stall-fallback retry, #78981) addressing
donovan-yohan's post-merge adversarial review. The stall path detaches a
timed-out primary worker (fence cancel wins; future stays on the pool)
and immediately runs the fallback against the SAME ContextCompressor,
creating two verified races:

1. Late-primary snapshot restore: the detached primary's unwind called
   _restore_compressor_attempt_state with the PRIMARY's pre-attempt
   snapshot. Landing after the fallback's commit it rolled
   _previous_summary/cooldown/provenance/telemetry back to pre-primary
   values, silently discarding fallback-owned state.
2. Shared _compression_cancelled_check: the late primary's `finally`
   cleared the callback the fallback had just installed, so the
   fallback's F4 cancellation consult read None.

Fix: a monotonic per-compressor attempt generation claimed under one
module lock (_claim_compressor_attempt). Snapshot restores carry their
claiming generation and no-op when stale; the cancelled-check set/clear
moves into owner-stamped helpers (_install_compression_cancelled_check /
_clear_compression_cancelled_check_if_owner) so only the installing
attempt can clear it. The commit fence keeps owning COMMIT admission;
the generation owns compressor-ATTRIBUTE writes — two boundaries.
Legacy callers (attempt_generation=None) and slotted third-party
compressors (generation 0) keep the historical unconditional behavior.

Secondary review items:
- Lean chunk digests during a stall-fallback retry now follow the
  summary onto the pinned healthy route: take_pinned_summary_route()
  echoes the consumed route into a context-local
  _SUMMARY_ROUTE_CONSUMED, and _build_chunk_digests passes
  attempt_summary_route_kwargs() (non-consuming) to call_llm. The pin's
  single-use contract for the SUMMARY call is unchanged — the
  main-model retry still never re-issues the pinned route.
- Worker re-run repeating pre-compression callbacks: documented as an
  accepted limitation on _retry_compression_on_fallback_chain
  (built-ins idempotent; resuming mid-pipeline would couple the retry
  to host callback ordering).

Tests (tests/agent/test_compression_attempt_ownership.py, 10 cases):
deterministic interleavings for both races (late-primary restore
no-ops + preserves fallback state; stale finally cannot clear the
fallback's callback), legacy/slotted compatibility, digest route
follow + context-locality of the consumed echo. Mutation-checked:
reverting only the two prod files to origin/main fails the suite;
restored stack green (21 passed incl. the original #78981 suite).

The one red in the wider sweep
(test_silence_cannot_approach_double_idle_timeout) is pre-existing on
clean origin/main — verified independently.
2026-08-28 12:52:53 +05:30
Mike DeMott dd5481aa40 refactor: centralize fast compression controls 2026-08-28 12:38:49 +05:30
Mike DeMott 372c4cdfce fix(compression): certify the effective fast route 2026-08-28 12:38:49 +05:30
Mike DeMott 213ae08e7a perf(compression): add guarded fast summary lane 2026-08-28 12:38:49 +05:30
Shaun Eccles 6151e59d65 fix(compression): surface an unpublished stall-fallback fence at WARNING
Review follow-up for #95433. When the host's fence factory is absent or
raises, the retry runs on a private CompressionCommitFence() that
hard-interrupt admission never reads — /stop would serialize against the
aborted attempt's fence instead of the retry's commit boundary. Promote
the factory-failure swallow from debug to warning and warn when the retry
has no published fence, so the control-plane degradation is visible rather
than silent.

Also documents the pin's coverage (single _generate_summary call; the
lean-mode chunk digests are a separate unpinned call path).
2026-08-28 02:29:27 +05:30
Shaun Eccles 2c6938dc3a fix(compression): retry a stalled summary on the fallback chain (#78981)
A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.

The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
2026-08-28 02:29:27 +05:30
lesseradmin 1341dfbd12 fix(compaction): exclude operational notifications from tail anchor and auto-focus (#92703)
Kanban/background completion wakes persist as role=user rows typed with
display_kind="internal_notification" (the synthetic-wake path in run.py).
The model-payload builder already strips display_kind before the request
and is_user_originated_turn already ignores it, but two compaction scans
still treated those rows as real user turns:

- _is_actionable_user_turn (tail anchor) only checked role/content, so a
  notification became the protected 'last user turn' the compressor keeps.
- _derive_auto_focus_topic only skipped synthetic compression turns, so
  operational notices leaked into the compact focus hint.

Both now exclude display_kind-typed rows, mirroring the existing
is_user_originated_turn exclusion. No schema change; cache- and
role-alternation-safe.

Behavior-contract tests feed 1,000 operational notifications around one
human turn and assert they never anchor the tail, become the auto-focus
source, or count as actionable user turns.

Fixes #92703
2026-08-26 21:38:20 -07:00
Teknium 6e5413844e feat(compression): lean tail retention is the default — compaction keeps 10-25K verbatim, not 100-240K
The legacy tail budget scales as threshold×target_ratio, which was designed
around 128K windows at a 50% trigger (~13K tail). On modern big-window
models with raised thresholds it silently hoards: a 1M-window session at
threshold 0.85 keeps a 170K-token verbatim tail (255K soft ceiling) out of
EVERY compaction, so a 540K manual /compress lands at ~290K and every
subsequent turn re-ships the hoard. Nobody chooses this; it is an artifact
of the formula outside its design envelope.

Lean mode (#87326, compaction-v2) was built for exactly this and its recall
was validated in the before/after eval (evals/compaction/results/): clamped
2.5%-of-window tail (10K floor / 25K cap), continuity carried by the
upgraded summary (digests, anchor index, verbatim user messages,
session_search recovery pointers). This flips the DEFAULT to lean; explicit
'tail_mode: legacy' in config keeps the old behavior exactly.

Also fixes a latent bug the flip exposed: update_model() re-assigned the
LEGACY formula directly when recomputing budgets, silently reverting a lean
compressor to the hoard on every mid-session model switch. The recompute
now routes through the mode-aware tail_token_budget property (regression
test included).

Surfaces: context_compressor.py defaults + getattr fallbacks, agent_init
parse default, DEFAULT_CONFIG, gateway _CACHE_BUSTING_CONFIG_KEYS gains
compression.tail_mode (mode changes now evict cached gateway agents like
target_ratio changes do), user + developer docs. Tests: 3 new default
contracts, legacy tests pinned explicitly, feasibility-skip scenario pinned
to legacy (under lean its payloads correctly become compressible).

E2E counterfactual (real imports, 1M window @ 0.85):
  main default:  legacy, tail 170,000 (ceiling 255,000)
  head default:  lean,   tail  25,000 (ceiling  37,500)
  head legacy:   170,000 (opt-out intact)
  update_model to 400K: 10,000 (lean preserved across switch)
2026-08-26 07:16:04 -07:00
kshitijk4poor 4ba2608524 fix(compressor): widen empty-content abort to sibling no-response shapes + snapshot state field
Follow-up to PR #94531 salvage:
- classify the auxiliary boundary's terminal 'None response' /
  'invalid response' errors (#7264) into the same empty-content abort
  carve-out so those shapes also preserve the session (#94459's wider
  classification, sibling shapes from #94448)
- register _last_summary_empty_content_failure in
  _COMPRESSOR_ATTEMPT_STATE_FIELDS so pre-commit hard-cancel rollback
  restores the flag (conversation_compression snapshot allow-list)
- tests: cooldown re-entry keeps aborting; both sibling shapes abort
- attribution: map zhangyswx@163.com -> YusenZhang0601
2026-08-26 13:02:27 +05:30
TonyRainforest fa210e5a96 fix(compressor): abort compression on empty-content provider degradation to prevent context loss (#94448)
When an auxiliary or main summarizer LLM returns an HTTP 200 with an empty or whitespace-only response (e.g., degraded provider/channel), abort compression and preserve the full conversation context rather than falling through to the destructive static-fallback branch that drops the middle window.

- Track _last_summary_empty_content_failure across _generate_summary() and compress()
- Attempt fallback to the main model when an aux model returns empty content
- Abort compression and preserve all messages intact if no valid summary can be generated
- Record summary_empty_content_failure in telemetry and log actionable diagnostic guidance
- Add comprehensive unit tests in tests/agent/test_context_compressor.py

Fixes #94448
2026-08-26 13:02:27 +05:30
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
Aniruddha Adak f778c0d941 fix(compression): structural no-ops defer retries instead of striking the breaker
Fixes #93022. A short session (protection window >= transcript) hits the
"insufficient messages" / "no compressible window" branches twice and
permanently trips the anti-thrash breaker, even though nothing was
eligible to compress - compression was never attempted, so there is
nothing "ineffective" to score. The session then rides past the
threshold with no compaction possible (recovery probes only soften,
not fix, the misclassification).

Distinguish "nothing eligible right now" from "attempted and
underperformed":

- New transient _structural_no_op_backoff_until (in-memory, 300s)
  armed by _record_structural_no_op() at the three structural no-op
  sites: insufficient_messages, no_compressible_window,
  empty_post_handoff_window. No strikes accumulate; auto-compaction
  resumes on its own once the backoff lapses or the transcript outgrows
  the protection window.
- The backoff gates should_compress via
  _automatic_compression_blocked_locally and surfaces in
  _compression_block_reason as "structural_backoff:<seconds>".
- #40803's frozen-CLI guarantee is preserved: a transcript that can
  never shrink retries at most once per backoff window instead of
  every turn.
- force=True (/compress) clears an active backoff before attempting;
  record_completed_compaction() lifts it - both prove the transcript
  is compressible/being worked.
- Genuine attempted-but-underperformed verdicts still strike the
  durable ineffective counter unchanged.

Tests: new tests/agent/test_context_compressor_structural_backoff.py;
updated the two tests that asserted the old strike-on-noop behavior.
2026-08-23 18:27:07 -07:00
Teknium a2a43f7e82 fix(agent): widen composite-id alias matching to the compressor; unify variant policy owners (#63000)
Follow-up on top of the salvaged #93335:

- context_compressor._sanitize_tool_pairs now expands alias spellings on
  the RESULT side too (tool_result_id_variants), so a composite
  call|item-keyed result pairs with its split-field tool_call instead of
  being dropped and its call stripped.
- The compressor's _tool_call_id_variants staticmethod and
  agent_runtime_helpers' module-level _tool_call_id_variants are now thin
  forwarders to agent.message_sanitization.tool_call_id_variants — one
  policy owner for alias expansion, so the pre-call sanitizer, repair
  pass, dedup pass, and compression sanitizer can never drift apart.
- Preserved the #91768 SDK-object tolerance in repair pass 1 (the
  shared helper handles non-dict tool_calls via getattr; the salvaged
  commit's isinstance-dict guard was dropped in the merge resolution).

New regression tests: composite-keyed results through
sanitize_api_messages (both directions) and _sanitize_tool_pairs, with
negative controls. Sabotage-verified: compressor test fails with raw
tool_call_id tracking.
2026-08-23 18:24:43 -07:00
joaomarcos 5496d5995a fix(agent): preserve tool results across ID variants
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
2026-08-23 18:24:43 -07:00
srojk34 52fb5081cc fix(compression): register both id/call_id variants in _sanitize_tool_pairs
_sanitize_tool_pairs() matched tool_call/tool_result pairs using a
single-value call_id||id precedence per tool_call (_get_tool_call_id).
In the Codex Responses API format an assistant tool_call carries both a
distinct id (fc_...) and call_id (call_...); a tool result's
tool_call_id may be keyed on either depending on which code path built
it. Whenever a genuinely matching pair used the field the precedence
didn't pick, the sanitizer misclassified it as orphaned on BOTH sides:
it dropped the valid tool result AND stripped the tool_call from the
assistant message, even though neither was orphaned.

Live-verified before the fix: {"id": "fc_777", "call_id": "call_777"}
+ a tool result with tool_call_id="fc_777" (a valid pair) was fully
removed by current main.

Register both id and call_id as valid match keys via a new
_tool_call_id_variants() helper (a set per tool_call, not a single
value), matching #58168's fix for repair_message_sequence's known-id
set today. A tool_call now survives if ANY of its id variants has a
matching result, which is not vulnerable to precedence order at all
(unlike swapping which field is checked first, which only trades which
sub-case is broken).

Note on #56425 (open, unreviewed): that PR touches this same function
for the same underlying issue (#55626) by swapping the call_id||id
precedence to id||call_id. That fixes the specific case where a result
matches `id` but not the reverse case (a result matching `call_id`
while `id` is also present) -- the precedence-swap approach cannot fix
the class, only relocate which sub-case is broken. This fix instead
mirrors the already-merged #58168 pattern (register the superset of
both ids as valid matches), which has no such blind spot. Adds 2
regression tests: the previously-mismatched case, and a negative
control confirming genuine orphans are still stripped alongside a
valid dual-id pair in the same window.
2026-08-23 17:01:20 -07:00
kshitijk4poor b7544dba01 fix(compression): share one image-strip policy across demote and retire passes
The demote pass (pass 2) and the retire pass (3.5, #92783) each carried
their own copy of the two image-strip branches. The copies had already
diverged: the retire pass dropped the stale api_content sidecar on
rewrite, the demote pass did not — leaving an exact-wire sidecar that
replay could use to resend the pre-strip image bytes.

Extract _strip_images_from_tool_msg as the single policy owner; both
passes now use it, closing the sidecar gap in the demote path.
2026-08-23 15:03:29 +05:30
kshitijk4poor dff84f1890 fix(browser): cap browser_vision native embeds for history reuse
browser_vision's native fast path base64-encoded screenshots at full
resolution and baked them into the tool result uncapped — the exact
sibling of the vision_analyze path #92699 fixed. Apply the same
proactive 256KB/1568px resize before the embed enters reusable history.

Fail-open by design: without Pillow the resize helper falls back to raw
bytes and the compressor's keep-newest pass still retires stale embeds.

Sibling-gap follow-up for the #92725 salvage; the shared-cap approach
mirrors the policy-owner idea from #92748.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-08-23 13:27:04 +05:30
HexLab98 7ff2fe8bc9 fix(compression): retire stale vision tool images in the protected tail
Images locked in protect_last_n never shrank, so compression savings
stayed under 10% and anti-thrash disabled further compaction. Keep the
newest three tool-result screenshots live for follow-up QA and replace
older native embeds with placeholders.
2026-08-23 13:27:04 +05:30
poisdahl ebae0064a2 fix(compression): preserve live assistant carriers after refresh 2026-08-22 16:48:57 +02:00
poisdahl a5b326a471 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821
# Conflicts:
#	tests/agent/test_reference_handoff_active_turn.py
2026-08-22 16:47:39 +02:00