Commit Graph

300 Commits

Author SHA1 Message Date
kshitijk4poor 61cd299c6e fix(compression): attempt-generation ownership for overlapping stall-fallback attempts
Follow-up to #96634 (stall-fallback retry, #78981) addressing
donovan-yohan's post-merge adversarial review. The stall path detaches a
timed-out primary worker (fence cancel wins; future stays on the pool)
and immediately runs the fallback against the SAME ContextCompressor,
creating two verified races:

1. Late-primary snapshot restore: the detached primary's unwind called
   _restore_compressor_attempt_state with the PRIMARY's pre-attempt
   snapshot. Landing after the fallback's commit it rolled
   _previous_summary/cooldown/provenance/telemetry back to pre-primary
   values, silently discarding fallback-owned state.
2. Shared _compression_cancelled_check: the late primary's `finally`
   cleared the callback the fallback had just installed, so the
   fallback's F4 cancellation consult read None.

Fix: a monotonic per-compressor attempt generation claimed under one
module lock (_claim_compressor_attempt). Snapshot restores carry their
claiming generation and no-op when stale; the cancelled-check set/clear
moves into owner-stamped helpers (_install_compression_cancelled_check /
_clear_compression_cancelled_check_if_owner) so only the installing
attempt can clear it. The commit fence keeps owning COMMIT admission;
the generation owns compressor-ATTRIBUTE writes — two boundaries.
Legacy callers (attempt_generation=None) and slotted third-party
compressors (generation 0) keep the historical unconditional behavior.

Secondary review items:
- Lean chunk digests during a stall-fallback retry now follow the
  summary onto the pinned healthy route: take_pinned_summary_route()
  echoes the consumed route into a context-local
  _SUMMARY_ROUTE_CONSUMED, and _build_chunk_digests passes
  attempt_summary_route_kwargs() (non-consuming) to call_llm. The pin's
  single-use contract for the SUMMARY call is unchanged — the
  main-model retry still never re-issues the pinned route.
- Worker re-run repeating pre-compression callbacks: documented as an
  accepted limitation on _retry_compression_on_fallback_chain
  (built-ins idempotent; resuming mid-pipeline would couple the retry
  to host callback ordering).

Tests (tests/agent/test_compression_attempt_ownership.py, 10 cases):
deterministic interleavings for both races (late-primary restore
no-ops + preserves fallback state; stale finally cannot clear the
fallback's callback), legacy/slotted compatibility, digest route
follow + context-locality of the consumed echo. Mutation-checked:
reverting only the two prod files to origin/main fails the suite;
restored stack green (21 passed incl. the original #78981 suite).

The one red in the wider sweep
(test_silence_cannot_approach_double_idle_timeout) is pre-existing on
clean origin/main — verified independently.
2026-08-28 12:52:53 +05:30
Mike DeMott dd5481aa40 refactor: centralize fast compression controls 2026-08-28 12:38:49 +05:30
Mike DeMott 372c4cdfce fix(compression): certify the effective fast route 2026-08-28 12:38:49 +05:30
Mike DeMott 213ae08e7a perf(compression): add guarded fast summary lane 2026-08-28 12:38:49 +05:30
Shaun Eccles 6151e59d65 fix(compression): surface an unpublished stall-fallback fence at WARNING
Review follow-up for #95433. When the host's fence factory is absent or
raises, the retry runs on a private CompressionCommitFence() that
hard-interrupt admission never reads — /stop would serialize against the
aborted attempt's fence instead of the retry's commit boundary. Promote
the factory-failure swallow from debug to warning and warn when the retry
has no published fence, so the control-plane degradation is visible rather
than silent.

Also documents the pin's coverage (single _generate_summary call; the
lean-mode chunk digests are a separate unpinned call path).
2026-08-28 02:29:27 +05:30
Shaun Eccles 2c6938dc3a fix(compression): retry a stalled summary on the fallback chain (#78981)
A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.

The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
2026-08-28 02:29:27 +05:30
lesseradmin 1341dfbd12 fix(compaction): exclude operational notifications from tail anchor and auto-focus (#92703)
Kanban/background completion wakes persist as role=user rows typed with
display_kind="internal_notification" (the synthetic-wake path in run.py).
The model-payload builder already strips display_kind before the request
and is_user_originated_turn already ignores it, but two compaction scans
still treated those rows as real user turns:

- _is_actionable_user_turn (tail anchor) only checked role/content, so a
  notification became the protected 'last user turn' the compressor keeps.
- _derive_auto_focus_topic only skipped synthetic compression turns, so
  operational notices leaked into the compact focus hint.

Both now exclude display_kind-typed rows, mirroring the existing
is_user_originated_turn exclusion. No schema change; cache- and
role-alternation-safe.

Behavior-contract tests feed 1,000 operational notifications around one
human turn and assert they never anchor the tail, become the auto-focus
source, or count as actionable user turns.

Fixes #92703
2026-08-26 21:38:20 -07:00
Teknium 6e5413844e feat(compression): lean tail retention is the default — compaction keeps 10-25K verbatim, not 100-240K
The legacy tail budget scales as threshold×target_ratio, which was designed
around 128K windows at a 50% trigger (~13K tail). On modern big-window
models with raised thresholds it silently hoards: a 1M-window session at
threshold 0.85 keeps a 170K-token verbatim tail (255K soft ceiling) out of
EVERY compaction, so a 540K manual /compress lands at ~290K and every
subsequent turn re-ships the hoard. Nobody chooses this; it is an artifact
of the formula outside its design envelope.

Lean mode (#87326, compaction-v2) was built for exactly this and its recall
was validated in the before/after eval (evals/compaction/results/): clamped
2.5%-of-window tail (10K floor / 25K cap), continuity carried by the
upgraded summary (digests, anchor index, verbatim user messages,
session_search recovery pointers). This flips the DEFAULT to lean; explicit
'tail_mode: legacy' in config keeps the old behavior exactly.

Also fixes a latent bug the flip exposed: update_model() re-assigned the
LEGACY formula directly when recomputing budgets, silently reverting a lean
compressor to the hoard on every mid-session model switch. The recompute
now routes through the mode-aware tail_token_budget property (regression
test included).

Surfaces: context_compressor.py defaults + getattr fallbacks, agent_init
parse default, DEFAULT_CONFIG, gateway _CACHE_BUSTING_CONFIG_KEYS gains
compression.tail_mode (mode changes now evict cached gateway agents like
target_ratio changes do), user + developer docs. Tests: 3 new default
contracts, legacy tests pinned explicitly, feasibility-skip scenario pinned
to legacy (under lean its payloads correctly become compressible).

E2E counterfactual (real imports, 1M window @ 0.85):
  main default:  legacy, tail 170,000 (ceiling 255,000)
  head default:  lean,   tail  25,000 (ceiling  37,500)
  head legacy:   170,000 (opt-out intact)
  update_model to 400K: 10,000 (lean preserved across switch)
2026-08-26 07:16:04 -07:00
kshitijk4poor 4ba2608524 fix(compressor): widen empty-content abort to sibling no-response shapes + snapshot state field
Follow-up to PR #94531 salvage:
- classify the auxiliary boundary's terminal 'None response' /
  'invalid response' errors (#7264) into the same empty-content abort
  carve-out so those shapes also preserve the session (#94459's wider
  classification, sibling shapes from #94448)
- register _last_summary_empty_content_failure in
  _COMPRESSOR_ATTEMPT_STATE_FIELDS so pre-commit hard-cancel rollback
  restores the flag (conversation_compression snapshot allow-list)
- tests: cooldown re-entry keeps aborting; both sibling shapes abort
- attribution: map zhangyswx@163.com -> YusenZhang0601
2026-08-26 13:02:27 +05:30
TonyRainforest fa210e5a96 fix(compressor): abort compression on empty-content provider degradation to prevent context loss (#94448)
When an auxiliary or main summarizer LLM returns an HTTP 200 with an empty or whitespace-only response (e.g., degraded provider/channel), abort compression and preserve the full conversation context rather than falling through to the destructive static-fallback branch that drops the middle window.

- Track _last_summary_empty_content_failure across _generate_summary() and compress()
- Attempt fallback to the main model when an aux model returns empty content
- Abort compression and preserve all messages intact if no valid summary can be generated
- Record summary_empty_content_failure in telemetry and log actionable diagnostic guidance
- Add comprehensive unit tests in tests/agent/test_context_compressor.py

Fixes #94448
2026-08-26 13:02:27 +05:30
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
Aniruddha Adak f778c0d941 fix(compression): structural no-ops defer retries instead of striking the breaker
Fixes #93022. A short session (protection window >= transcript) hits the
"insufficient messages" / "no compressible window" branches twice and
permanently trips the anti-thrash breaker, even though nothing was
eligible to compress - compression was never attempted, so there is
nothing "ineffective" to score. The session then rides past the
threshold with no compaction possible (recovery probes only soften,
not fix, the misclassification).

Distinguish "nothing eligible right now" from "attempted and
underperformed":

- New transient _structural_no_op_backoff_until (in-memory, 300s)
  armed by _record_structural_no_op() at the three structural no-op
  sites: insufficient_messages, no_compressible_window,
  empty_post_handoff_window. No strikes accumulate; auto-compaction
  resumes on its own once the backoff lapses or the transcript outgrows
  the protection window.
- The backoff gates should_compress via
  _automatic_compression_blocked_locally and surfaces in
  _compression_block_reason as "structural_backoff:<seconds>".
- #40803's frozen-CLI guarantee is preserved: a transcript that can
  never shrink retries at most once per backoff window instead of
  every turn.
- force=True (/compress) clears an active backoff before attempting;
  record_completed_compaction() lifts it - both prove the transcript
  is compressible/being worked.
- Genuine attempted-but-underperformed verdicts still strike the
  durable ineffective counter unchanged.

Tests: new tests/agent/test_context_compressor_structural_backoff.py;
updated the two tests that asserted the old strike-on-noop behavior.
2026-08-23 18:27:07 -07:00
Teknium a2a43f7e82 fix(agent): widen composite-id alias matching to the compressor; unify variant policy owners (#63000)
Follow-up on top of the salvaged #93335:

- context_compressor._sanitize_tool_pairs now expands alias spellings on
  the RESULT side too (tool_result_id_variants), so a composite
  call|item-keyed result pairs with its split-field tool_call instead of
  being dropped and its call stripped.
- The compressor's _tool_call_id_variants staticmethod and
  agent_runtime_helpers' module-level _tool_call_id_variants are now thin
  forwarders to agent.message_sanitization.tool_call_id_variants — one
  policy owner for alias expansion, so the pre-call sanitizer, repair
  pass, dedup pass, and compression sanitizer can never drift apart.
- Preserved the #91768 SDK-object tolerance in repair pass 1 (the
  shared helper handles non-dict tool_calls via getattr; the salvaged
  commit's isinstance-dict guard was dropped in the merge resolution).

New regression tests: composite-keyed results through
sanitize_api_messages (both directions) and _sanitize_tool_pairs, with
negative controls. Sabotage-verified: compressor test fails with raw
tool_call_id tracking.
2026-08-23 18:24:43 -07:00
joaomarcos 5496d5995a fix(agent): preserve tool results across ID variants
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
2026-08-23 18:24:43 -07:00
srojk34 52fb5081cc fix(compression): register both id/call_id variants in _sanitize_tool_pairs
_sanitize_tool_pairs() matched tool_call/tool_result pairs using a
single-value call_id||id precedence per tool_call (_get_tool_call_id).
In the Codex Responses API format an assistant tool_call carries both a
distinct id (fc_...) and call_id (call_...); a tool result's
tool_call_id may be keyed on either depending on which code path built
it. Whenever a genuinely matching pair used the field the precedence
didn't pick, the sanitizer misclassified it as orphaned on BOTH sides:
it dropped the valid tool result AND stripped the tool_call from the
assistant message, even though neither was orphaned.

Live-verified before the fix: {"id": "fc_777", "call_id": "call_777"}
+ a tool result with tool_call_id="fc_777" (a valid pair) was fully
removed by current main.

Register both id and call_id as valid match keys via a new
_tool_call_id_variants() helper (a set per tool_call, not a single
value), matching #58168's fix for repair_message_sequence's known-id
set today. A tool_call now survives if ANY of its id variants has a
matching result, which is not vulnerable to precedence order at all
(unlike swapping which field is checked first, which only trades which
sub-case is broken).

Note on #56425 (open, unreviewed): that PR touches this same function
for the same underlying issue (#55626) by swapping the call_id||id
precedence to id||call_id. That fixes the specific case where a result
matches `id` but not the reverse case (a result matching `call_id`
while `id` is also present) -- the precedence-swap approach cannot fix
the class, only relocate which sub-case is broken. This fix instead
mirrors the already-merged #58168 pattern (register the superset of
both ids as valid matches), which has no such blind spot. Adds 2
regression tests: the previously-mismatched case, and a negative
control confirming genuine orphans are still stripped alongside a
valid dual-id pair in the same window.
2026-08-23 17:01:20 -07:00
kshitijk4poor b7544dba01 fix(compression): share one image-strip policy across demote and retire passes
The demote pass (pass 2) and the retire pass (3.5, #92783) each carried
their own copy of the two image-strip branches. The copies had already
diverged: the retire pass dropped the stale api_content sidecar on
rewrite, the demote pass did not — leaving an exact-wire sidecar that
replay could use to resend the pre-strip image bytes.

Extract _strip_images_from_tool_msg as the single policy owner; both
passes now use it, closing the sidecar gap in the demote path.
2026-08-23 15:03:29 +05:30
kshitijk4poor dff84f1890 fix(browser): cap browser_vision native embeds for history reuse
browser_vision's native fast path base64-encoded screenshots at full
resolution and baked them into the tool result uncapped — the exact
sibling of the vision_analyze path #92699 fixed. Apply the same
proactive 256KB/1568px resize before the embed enters reusable history.

Fail-open by design: without Pillow the resize helper falls back to raw
bytes and the compressor's keep-newest pass still retires stale embeds.

Sibling-gap follow-up for the #92725 salvage; the shared-cap approach
mirrors the policy-owner idea from #92748.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-08-23 13:27:04 +05:30
HexLab98 7ff2fe8bc9 fix(compression): retire stale vision tool images in the protected tail
Images locked in protect_last_n never shrank, so compression savings
stayed under 10% and anti-thrash disabled further compaction. Keep the
newest three tool-result screenshots live for follow-up QA and replace
older native embeds with placeholders.
2026-08-23 13:27:04 +05:30
poisdahl ebae0064a2 fix(compression): preserve live assistant carriers after refresh 2026-08-22 16:48:57 +02:00
poisdahl a5b326a471 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821
# Conflicts:
#	tests/agent/test_reference_handoff_active_turn.py
2026-08-22 16:47:39 +02:00
abundantbeing 6b7aee2f80 fix(agent): preserve live merged tool-call carriers
Identify a completed merged assistant handoff from the carrier's own stop state instead of an unrelated adjacent history row. Keep carriers with pending tool calls actionable so compaction cannot abort a live tool chain.
2026-08-21 21:52:55 +05:30
abundantbeing fdf0111410 fix(agent): guard merged assistant compaction handoffs
Treat a merged assistant-role summary carrier as the driving reference handoff when it immediately follows a completed assistant stop. Its preserved prose and stale tool_calls are assistant continuity, not a fresh live user request.

Keep legitimate in-flight behavior unchanged when there is no completed stop, a real user turn follows, or a distinct later assistant tool-call row continues the loop.

Extends the #80622 active-turn guard for the merged-carrier shape reported under #42768.
2026-08-21 21:52:55 +05:30
poisdahl abf87e7248 Merge current main into composite-carrier fix 2026-08-21 15:56:45 +02:00
kshitijk4poor 0596ccdeb3 fix(compression): salvage follow-up — todo snapshot last-resort, reuse prune helpers
Review follow-up on the salvaged #90353:
- Todo snapshot (+ coupled pruned-skill reload notice, 7a16840add) is now
  reduced only as a LAST resort after reasoning/tool/summary shrink ops,
  and the reload notice survives even then.
- Reuse existing helpers/constants instead of re-hardcoding:
  _PRUNED_TOOL_PLACEHOLDER, _PRUNE_MIN_CHARS, _NEWEST_TURN_ONLY_BUDGET_KEYS,
  and _prune_stale_reasoning_replay (codex sidecar shrink, #71058 boundary).
- Assistant-role messages without the summary metadata key are no longer
  truncatable by the summary-cap heuristic.
- Caller passes budget so the estimator runs 3x, not 5x, per would-grow pass.
2026-08-20 11:36:54 +05:30
MindDragonLabs fb96247eaf fix(compression): salvage grown candidates before refusal 2026-08-20 11:36:54 +05:30
liuhao1024 62016a1b0a fix(compression): count a would-grow refusal as an ineffective strike
The anti-growth guard correctly refuses to persist a compressed
candidate larger than the original, but the rejection was never
recorded by the anti-thrashing breaker: _ineffective_compression_count
stayed at zero, the latch never tripped, and automatic compression
retried the SAME unchanged transcript on every turn - same summary
request, same refusal, same user-facing warning (#88568).

Add ContextCompressor.record_rejected_compaction(): one persisted
ineffective strike, without arming post-compaction real-usage
verification (nothing was committed) and without touching the
fallback-summary streak (no summary was accepted). The would-grow
abort path in conversation_compression calls it before returning the
original transcript. Two refusals latch the normal breaker, manual
/compress keeps bypassing it (force=True), and the existing recovery
window still allows one probe later.

Fixes #88568
2026-08-20 11:36:54 +05:30
Teknium 7bf66ec39b fix(compression): /compress refusal no longer reports a successful rewrite 2026-08-19 22:41:05 -07:00
poisdahl 7ca1987459 Merge upstream main into PR 81234 2026-08-16 12:20:50 +02:00
Christopher 5602f04a6e fix(agent): clear in-memory cooldown when hygiene overwrites the shared row
A later hygiene idle-timeout write can replace an aux-model cooldown on
the shared column. Drop the in-memory timer on that refresh so the
in-agent compressor is not still blocked after the DB row is hygiene.
2026-08-16 02:01:41 -07:00
Christopher 2cb8381963 fix(agent): do not let hygiene idle timeouts block in-agent compression
Session hygiene persists compression_failure_cooldown_until after a
30s no-progress watchdog so the pre-agent pass can skip. The
in-conversation compressor read the same column and then refused to
run even though its own budget is sufficient.

Ignore hygiene idle-timeout errors on the in-agent path. Real
aux-model faults such as rate limits still block.

Fixes #86972
2026-08-16 02:01:41 -07:00
Teknium 31ca1200ef feat(compression): field-proven summarizer prompt upgrades
- anti-injection rule in preamble (gemini-cli state_snapshot pattern)
- verbatim security-constraint preservation in Constraints & Preferences
  (claude-code rule)
- Errors & Fixes section with user-correction quoting (claude-code sections
  4/  + CompInt user-feedback emphasis)
2026-08-15 17:01:22 -07:00
Teknium c4bbb14e52 feat(compression): mechanical anchor index + region-scoping tripwire
- _build_anchor_index(): regex-harvests PR/issue numbers, SHAs, branches,
  file paths, error strings, handles, URLs from the compacted region into a
  bounded indexed summary section. LLM-free, so needle identifiers cannot be
  paraphrased away (the GUI-lineage failure class: 10/15 verbatim-or-nothing
  golds). Doubles as session_search query-anchor map.
- evals/compaction/test_region_scoping.py: sentinel tripwire proving the
  summarizer input carries ONLY the compacted region (head/tail sentinels
  never reach the serialized turns body) in both legacy and lean modes.
2026-08-15 17:01:22 -07:00
Teknium 7a82457ede feat(compression): digest noise filter + FTS5 recovery sim + digest-aware query hints
- _digest_worthy() drops no-signal tool rows before chunking (GUI-lineage
  digests were starving on tool-noise)
- eval recovery sim now uses in-memory SQLite FTS5 + BM25 (production
  session_search engine) instead of term-frequency scoring
- recovery query writer sees the digest section (front of context) so it can
  mine anchor identifiers
2026-08-15 17:01:22 -07:00
Teknium 8fe9025abd feat(compression): lean tail mode + recovery-aware eval arm
Lean mode (tail_mode='lean', default stays 'legacy'):
- tail budget = clamp(2.5% of window, 10K, 25K) instead of 0.20*window
- stale tail tool results demoted to session_search recovery stubs
- chunked identifier-preserving digests of the compacted region (map-reduce,
  pristine pre-prune tool contents)
- verbatim user messages embedded in summary (codex retention-by-role rule)
- deterministic session_search recovery footer

Eval: policies matrix gains lean + a '+recovery' arm giving the answerer one
simulated session_search round-trip against the archived region.
2026-08-15 17:01:22 -07:00
poisdahl 3f075d41dd Merge remote-tracking branch 'origin/main' into codex/81234-live-main-final
# Conflicts:
#	tui_gateway/methods_prompt.py
#	tui_gateway/server.py
2026-08-15 12:05:14 +02:00
HexLab98 fc100f4b3b fix(compression): scan full window for handoffs before cross-session discard
A degenerate compress_end can hide an in-window handoff past the cut; the
#57835 guard then cleared a valid same-session _previous_summary (#83248).
2026-08-14 20:52:16 -07:00
fangliquanflq 8cf9e8a61b fix(compression): ignore background process notifications 2026-08-14 01:08:13 -07:00
poisdahl 3e5e4c5d20 fix(agent): preserve live turns in compaction carriers 2026-08-13 21:38:30 +02:00
kshitij 4f675cf2fc fix: double-paren bug in dropped-tools prefix + add empty-response nudge sibling
Fix: _LENGTH_CONTINUATION_DROPPED_TOOLS_PREFIX ended with '(' but
_get_continuation_prompt still had f'({tool_list})', producing
'((write_file)' instead of '(write_file)'. Removed the '(' from
the prefix constant — the parenthesis belongs in the interpolation.

Widened: promoted the empty-response nudge (line 6993,
'You just executed tool calls but returned an empty response...')
to _EMPTY_TOOL_RESPONSE_NUDGE constant and added it to the
classifier's recognition set. Same bug class — its
_empty_recovery_synthetic metadata flag doesn't survive SessionDB
projection either.

Test: added parametrize case for the empty-response nudge (7→8 cases).
E2E: verified byte-for-byte string equivalence for all nudge constants.
2026-08-09 16:20:48 +05:30
pierrenode 45cd93fb5b fix(agent): recognize the retry loop's other synthetic nudges during compaction
aed114a69 taught _is_synthetic_compression_user_turn to recognize the
max-iteration nudge as ephemeral runtime scaffolding rather than a human
turn, since its role="user" metadata flag doesn't survive SessionDB
projection and a crash/interrupt mid-turn can persist it durably — becoming
the compaction anchor / auto-focus topic in place of the real task.

conversation_loop.py's retry loop appends several more role="user" rows
with the exact same "ephemeral, metadata-tag-only" shape, none of them
recognized by the classifier:

- The three _get_continuation_prompt variants (length-continuation nudge,
  tagged _length_continuation_nudge) — two fixed strings plus a third that
  interpolates the dropped-tool-call list.
- _CODEX_INCOMPLETE_NUDGE (codex/responses reasoning-only retry).
- The codex ack-continuation nudge (acknowledgment-only reply re-prompt).
- The dropped-tool-call nudge (tagged _dropped_toolcall_nudge) — persisted
  across up to 3 consecutive retries before the finalization pop-loop
  strips it; an interrupt/crash before that pop can persist it same as the
  max-iteration case.

Promote the previously-inline nudge strings to named module-level constants
in conversation_loop.py (single source of truth for both construction and
recognition), then extend the classifier to recognize all of them — exact
match for the five fixed-content nudges, a stable-prefix check for the
dropped-tool-call continuation variant (its tool list is interpolated so it
can't be exact-matched, same treatment TODO_INJECTION_HEADER already gets).
Imported lazily inside the classifier to avoid a module-load-order cycle —
conversation_loop.py already imports FROM context_compressor.py at call
time for the same reason.
2026-08-09 16:20:48 +05:30
Teknium 212e84176d fix(compression): charge stale thinking to the tail budget only on the newest assistant turn (#73624)
Generic thinking fields (reasoning / reasoning_content + the
reasoning_details text charge) are replayed for at most the NEWEST
assistant turn on every transport: Anthropic strips all-but-newest at
convert time, Bedrock Converse never replays thinking, and strict
chat-completions providers reject or one-space-pad the field. The tail
budget walks charged them on every message anyway, spending 19-24% of
the budget (per the issue's 1,025-message measurement) on bytes that
provably never reach the wire — so the tail cut landed early and each
compaction discarded more real transcript than configured.

_estimate_msg_budget_tokens now partitions the replay keys:

* _ALWAYS_REPLAYED_BUDGET_KEYS (codex_reasoning_items,
  codex_message_items) — charged unconditionally. These ride the wire
  on every retained turn (#55572), and codex_reasoning_items now also
  carries native server-side compaction checkpoints (#81747).
* _NEWEST_TURN_ONLY_BUDGET_KEYS (reasoning, reasoning_content) + the
  reasoning_details text charge — charged only for the newest assistant
  turn via charge_stale_thinking, resolved by the three budget walks
  (tail cut, raw-budget re-walk, proactive-prune boundary).

Default stays the conservative full charge for callers without
turn-position context. A partition invariant test pins that any future
_REPLAY_BUDGET_KEYS entry must be classified into exactly one class.

Direction credit: #73669 (@x7peeps) and #73730 (@webtecnica) both
attacked this; the keep_open reviews asked for provider/API-mode-aware
accounting that keeps Codex carriers charged — this implements that
shape.
2026-08-08 17:42:34 -07:00
Teknium e00965a7e8 fix(compression): correct prune boundary + exempt native compaction checkpoints
Two corrections on top of the #71077 base (the whole bug class):

1. Turn boundary = last USER message, not last assistant message. A Codex
   turn spans several assistant messages (assistant+tool_calls -> tool ->
   ... -> final assistant) whose reasoning items must replay together; the
   last-assistant boundary would strip reasoning mid-chain from the active
   turn (the gap flagged in PR #71077 review).

2. type="compaction" checkpoints (native server-side compaction, PR #81747)
   are exempt: they carry already-pruned history, not per-turn reasoning.
   Pruning filters items instead of popping the sidecar key.

Sibling site fixed in the same class: the Codex incomplete-continuation
dedup path blind-overwrote codex_reasoning_items on visually-duplicate
interim messages, which would drop the only copy of a checkpoint captured
on the earlier response. Extracted merge_interim_reasoning_items() into
agent/native_compaction.py; newer reasoning wins, prior checkpoints are
preserved unless the newer payload carries its own.
2026-08-08 14:09:41 -07:00
webtecnica adf9549cdd fix(compression): prune stale codex_reasoning_items during compaction (#71058) 2026-08-08 14:09:41 -07:00
kshitij d81f2f49ea refactor(compression): fold simplify findings — dedup floor constant, drift-guard test
- Wire the Pass-1 dedup floor (len < 200) to the shared _PRUNE_MIN_CHARS
  constant it was already documented as matching, and use the constant in
  the remaining test literal.
- Restructure the clarify 'resolved' computation (is_answer_shaped +
  sentinel check) instead of compute-then-flip.
- Add a live producer->recognizer drift guard: the REAL oneshot no-user
  callback's output must be recognized as a sentinel, so producer wording
  drift fails a test instead of silently reintroducing false attribution.
- Document the any()-poisoning semantic for multi-select sentinel lists.
2026-08-08 14:27:31 +05:30
kshitij 39056e8de4 fix(compression): filter clarify non-response sentinels; share prune floor constant
Follow-up to the salvaged #81244 commits:

- Timeout/no-user clarify callbacks (CLI timeout, gateway timeout and
  delivery failure, oneshot no-user) embed sentinel prose as
  user_response; quoting those as '[clarify] user responded: ...' would
  be false attribution. Route them to the generic summary path.
- Extract the shared _PRUNE_MIN_CHARS = 200 floor (prune default +
  proactive clamp) and cap the clarify summary at _PRUNE_MIN_CHARS - 1,
  removing the knife-edge equality the summary's survival depended on
  and keeping it out of the >=200-char dedup pass.
- Tests: 4 sentinel shapes + multi-select sentinel; mutation-checked.
2026-08-08 14:27:31 +05:30
Crypto Intern 3090e9e871 fix: reject forged clarify summaries 2026-08-08 14:27:31 +05:30
Crypto Intern 6433d5723f fix(compression): make clarify summaries UTF-8 safe 2026-08-08 14:27:31 +05:30
Crypto Intern d6511aecb6 fix(compression): preserve clarify responses 2026-08-08 14:27:31 +05:30
PRATHAMESH75 aed114a69b fix(agent): treat max-iteration nudge as synthetic during compaction
handle_max_iterations() appends its runtime summary request as a plain
role="user" row, which SessionDB persists verbatim. On later compaction the
synthetic-turn filters only recognized compaction summaries, continuation
rows, and todo snapshots, so the nudge could be selected as the latest
actionable user turn — becoming the task snapshot / auto-focus input and
getting summarized as "User asked: ...", demoting the real human task.

Metadata flags do not survive SessionDB projection (the reason the existing
markers are content-based), so recognition must key off stable content.
Extract the nudge into a shared MAX_ITERATIONS_SUMMARY_REQUEST constant and
teach _is_synthetic_compression_user_turn() to recognize it, mirroring the
continuation/todo markers. Every _is_actionable_user_turn call site already
pairs the synthetic guard, so the single recognizer change covers anchor
selection, auto-focus, and real-user-turn detection.

Fixes #78580
2026-08-08 14:24:39 +05:30
kshitij fecba5afcc refactor(agent): fold simplify findings — DB picker parity, single scan, canonical strip delegation
Review-pass follow-ups (three parallel reviewers, findings verified):

- hermes_state_search.py list_recent_user_messages now drops legacy
  standalone compaction handoffs in the decode loop (SQL can't see them:
  durable role=user, no display_kind). Closes the /undo N pairing skew
  where the in-memory count (new predicate) and the DB soft-delete pick
  (old predicate) targeted different turns on legacy sessions. Fetches
  with headroom so the requested limit is still honored. 3 new tests,
  mutation-checked (no-op'ing the skip fails 2/3).
- _should_skip_model_call_for_reference_handoff: single drive-check scan
  (was two — once inside the restore helper, once after); the restore
  helper no longer re-scans and its return value now decides the verdict.
- _final_response_from_messages replaced by the _HANDOFF_SKIP_FINAL_RESPONSE
  constant it always returned (parameter was unused).
- _handoff_carries_live_user_content delegates to the canonical
  _strip_context_summary_handoff_message — also fixes the edge where a
  merged-shaped row with an EMPTY preserved prior tail was wrongly
  treated as carrying live content.
- Site-level guard test for rollback.restore with a legacy handoff row
  (predicate-in-context, complements the unit tests).
2026-08-07 19:44:35 +05:30