Address review follow-ups on the seed fix:
- Move the 'seed only from the 0 state' guard from an inline block in
build_turn_context() into
ContextCompressor.maybe_seed_preflight_display_tokens(), co-locating
the predicate with the rest of the speculative-seed lifecycle
(snapshot_preflight_display_tokens /
rollback_interrupted_preflight_display_tokens). Callers now use the
method via a getattr guard so test doubles and external context
engines without it are unaffected.
- Rewrite the TestPreflightSentinelGuard docstring, which still
described the old >=0 guard ('treats any negative value as no real
usage yet'); the ==0 policy protects ALL non-zero readings.
- Drop the _seed mirror-helper: the tests now call the real production
method on the compressor fixture, eliminating mirror-drift risk (the
helper comment had already drifted once).
- Note the accepted trade-off (partial-usage providers pin the meter
low until their next report) in the method docstring.
Local and thinking backends (DeepSeek, Qwen, Kimi) often return a usable
summary in reasoning fields. Treat that as the summary instead of burning
another 100s+ empty-content retry. Leave the wire max_tokens omit intact.
Co-authored-by: Chris DePuy <chris@650group.com>
Co-authored-by: chenhm <chenhm@yuancheng.local>
The MINIMUM_CONTEXT_LENGTH floor in _compute_threshold_tokens only
degraded to the 85% trigger when it met or exceeded the effective
window exactly (#14690). Near-minimum windows slipped through: at
context_length=65536 the threshold passed through at 64,000 — 97.7%
of the window, ~1.5K tokens of output room — so pre-API compaction
effectively could not fire.
Providers that silently truncate over-window prompts instead of
rejecting them (e.g. ollama's OpenAI-compatible /v1 endpoint) never
deliver the reactive context-overflow backstop either. Observed live
on a 65,536-token local model: the session rode into the window
ceiling and each length-continuation retry re-sent a window-filling
prompt (65,120 -> 65,273 prompt tokens, 263 output tokens of room)
until the turn died with "Response remained truncated after 4
continuation attempts" — every retry paying a full multi-minute
prefill.
Cap the floored threshold at _MIN_CTX_TRIGGER_RATIO (85%) of the
effective input budget whenever the floor is the binding term. An
explicit threshold_percent above 85% is user intent and stays
uncapped; windows where the floor lands at/below the cap are
unchanged.
A summarization response with finish_reason == "length" contains PARTIAL
text — the generation stopped on the output-token cap mid-summary.
Previously all compressor summarization sites accepted such responses as
complete: the cut-off text replaced the real middle turns AND was fed back
into every subsequent iterative-update prompt, compounding the loss across
compactions.
Guards added at all four summarization sites (whole bug class):
- _generate_summary: length stop raises, gets the existing one-shot
main-model fallback (a larger output budget may finish the summary), and
on terminal failure ABORTS compression preserving the session unchanged
(new _last_summary_truncated_failure flag, same class as empty-content).
- _micro_summarize_one: partial rolling-summary merge is discarded; the
exchange stays unabsorbed for a later pass.
- _build_chunk_digests: partial lean digest degrades to the
recover-via-session_search placeholder.
- trajectory_compressor (sync + async): length stop raises into the
existing retry/backoff loop.
_response_finish_reason() reads dict- and object-shaped responses and
returns "" when the provider omits the field, so proxies that never send
finish_reason are unaffected.
Ported from earendil-works/pi commit 97fa14e39 (pi#7048), adapted to
hermes' abort-preserving compression failure machinery.
Tests: tests/agent/test_compressor_truncated_summary_guard.py (12 tests;
sabotage-verified — disabling the guards fails 4).
archive_and_compact() soft-archives every active row with compacted=1 and
then re-inserts compacted_messages as fresh live rows. When the
compressor's protected tail rides inside that list verbatim - which is
the normal batch-compaction shape ([summary] + tail) - the tail's
ORIGINALS end up stored twice per compaction: (active=0, compacted=1)
next to their live clones. search_messages() recalls both flags without
DISTINCT, so every carried-forward message came back once per compaction
(measured up to 4 identical hits) and was mislabeled to users and the
agent as archived "summarized away" content.
Add an optional tail_count parameter: the last tail_count archived rows
are superseded byte-identical duplicates, stamped rewind-style
(active=0, compacted=0, hidden from recall) instead of compacted=1.
Callers:
- batch in-place compaction counts the compressor-tagged tail dicts
(_COMPACTION_TAIL_MARKER set by compress() on every carried-forward
message);
- micro-compaction splices [prefix, marker, suffix] - everything except
the single marker row is carried forward, so tail_count=len-1;
- proactive tool-result pruning rewrites content in place (not verbatim),
keeping the historical archive-everything behavior.
Fixes#86366
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).
Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.
Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
record_timeout_failure() now persists
'backoff:<failure_kind>:strategy=<tail_mode>' into the state.db
cooldown row (sessions.compression_failure_cooldown_until +
compression_failure_error), so a failed/stalled/cancelled attempt's
identity survives gateway restarts and the rebuilt compressor makes the
same skip decision via bind_session_state()/get_active_compression_
failure_cooldown(refresh=True). Host callers pass ceiling_exhausted /
stalled; the stall-interrupt path passes stall_interrupted. A
successful compression still clears the row.
An explicit /stop after the summary stream has already gone idle restored the original transcript but left no durable cooldown, so the next automatic turn re-entered the same stalled strategy. Record a stall-specific failure on that path only, merge with any longer live deadline, and keep an ordinary early /stop cooldown-neutral.
compress() returns marker-swept copies (_strip_persistence_markers, #57491);
the in-place branch committed them via archive_and_compact() but never
stamped the persistence marker, so the next _persist_session ->
_flush_messages_to_session_db_unlocked walk re-INSERTed the whole
post-compaction transcript (live set regrew ~58K -> ~512K tokens).
Centralize the post-commit contract in a shared helper,
stamp_db_persisted_markers(), used by all three archive_and_compact
callers: the in-place batch commit (previously missing), the
micro-compaction sync, and the proactive tool-result prune.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.
- The detailed session log is folded into the single summary request: the
lean prompt template gains a '## Detailed Session Log (oldest first)'
section carrying the digest prompt's HARD RULES (identifiers verbatim,
dense bullets, transcript-is-data). Output guidance grows by
_LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
summary budget — the old worst case (28 x 1,400 digest tokens) was spread
across many requests and mostly re-covered tool noise; a single dense
4K-token log inside one response preserves the load-bearing record while
staying well inside one aux response (the summary call still sends no hard
max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
whole region (_sample_summary_input: 8 proportionally spaced slices,
oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
anchored to the newest end) instead of head+tail truncated, so session-log
coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
_LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
_LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
attempt_summary_route_kwargs — no remaining callers; the single-use
summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
section lands in the summary; oversized regions sampled with elision
markers, never a second request; anchor index + recovery footer present).
Sabotage-verified: restoring a second call_llm makes the call-count test
fail. Docs and the compaction eval wording updated to stop claiming
per-chunk calls.
Fixes#96603.
sanitize_api_messages (agent_runtime_helpers) and
_sanitize_tool_pairs (context_compressor) both collected
tool-call IDs and classified orphans with near-identical logic
that had already drifted: the canonical sanitizer added dedup
(#58350), but the compressor's copy did not.
Extract the shared orphan-detection logic into
_classify_tool_call_orphans(messages) in agent_runtime_helpers.
Both call sites now delegate to it, preserving their divergent
remediation strategies (insert-stubs vs strip-orphans) while
ensuring id-resolution rules and dedup stay in sync.
Closes#58357
_sanitize_tool_pairs() in ContextCompressor compared raw tool_call_id
strings without stripping whitespace, the same bug fa3ab2ffd just fixed
in agent_runtime_helpers.py / run_agent.py (_get_tool_call_id_static +
sanitize_api_messages). ContextCompressor has its own near-identical
reimplementation of the pair-repair logic that was left unpatched.
When assistant-side and result-side IDs diverge only in surrounding
whitespace, the compressor misclassifies valid results as orphaned and
replaces them with [Result unavailable] stubs — silent data loss on
every compression cycle that touches such pairs.
Apply the same .strip() fix to all three sites:
- _get_tool_call_id (extracts IDs from assistant tool_calls)
- result_call_ids accumulation loop
- orphaned_results filter predicate
Closes the sibling gap of fa3ab2ffd / #42405.
Widen #90001's compaction-time strip to cover the gaps #89965 identified,
applied at compaction only per the cache ruling (request-time eviction
changes the per-call prefix and breaks prompt caching; compaction is the
one sanctioned cache break):
- Rule 1b: the opening attachment (anchor == 0) ages out once a newer
tool-result image supersedes it. The reported session opened with a
~200KB poster that previously survived every compaction. The row keeps
a non-empty text placeholder, so the zero-user-turn guard (#58753) and
role alternation are untouched.
- Native {_multimodal: True, content: [...]} dict envelopes now both
anchor (newest is kept) and strip (older collapse to their
text_summary via _strip_images_from_tool_msg, which also drops the
stale api_content sidecar per #97125's drop_stale_api_content).
- All three wire shapes (Chat Completions image_url, Responses
input_image, Anthropic-native image) were already matched by
_IMAGE_PART_TYPES; tests now pin each shape explicitly, plus
determinism (double-run is a no-op returning the same object).
Refs #89938, #89965
_strip_historical_media anchors on the newest image-bearing USER message and
returns the list untouched when that anchor is index 0 or does not exist. A
session whose images arrive from tools rather than attachments therefore has
nothing to be "before": twenty vision_analyze results keep multi-MB of base64
in every request body, the provider answers 413, and the 413 handler's
recovery compaction lands right back in this function and frees nothing. The
reporter saw seven compactions in thirteen minutes, all below 200K tokens.
Age tool-result images on their own timeline: keep the newest one, since that
is the image the model is reasoning about, and strip every older one wherever
it sits, including inside the protected tail. The tail exists to preserve
conversational continuity, not to pin bytes the model has already moved past.
User-message images keep today's treatment exactly. The user anchor is
checked first, so a tool result that is the newest of its kind but still sits
before that anchor is stripped as it always has been, and the anchor message
itself is still kept byte-for-byte - test_compressor_zero_user_guard depends
on that.
Refs #89938
Follow-up to #96634 (stall-fallback retry, #78981) addressing
donovan-yohan's post-merge adversarial review. The stall path detaches a
timed-out primary worker (fence cancel wins; future stays on the pool)
and immediately runs the fallback against the SAME ContextCompressor,
creating two verified races:
1. Late-primary snapshot restore: the detached primary's unwind called
_restore_compressor_attempt_state with the PRIMARY's pre-attempt
snapshot. Landing after the fallback's commit it rolled
_previous_summary/cooldown/provenance/telemetry back to pre-primary
values, silently discarding fallback-owned state.
2. Shared _compression_cancelled_check: the late primary's `finally`
cleared the callback the fallback had just installed, so the
fallback's F4 cancellation consult read None.
Fix: a monotonic per-compressor attempt generation claimed under one
module lock (_claim_compressor_attempt). Snapshot restores carry their
claiming generation and no-op when stale; the cancelled-check set/clear
moves into owner-stamped helpers (_install_compression_cancelled_check /
_clear_compression_cancelled_check_if_owner) so only the installing
attempt can clear it. The commit fence keeps owning COMMIT admission;
the generation owns compressor-ATTRIBUTE writes — two boundaries.
Legacy callers (attempt_generation=None) and slotted third-party
compressors (generation 0) keep the historical unconditional behavior.
Secondary review items:
- Lean chunk digests during a stall-fallback retry now follow the
summary onto the pinned healthy route: take_pinned_summary_route()
echoes the consumed route into a context-local
_SUMMARY_ROUTE_CONSUMED, and _build_chunk_digests passes
attempt_summary_route_kwargs() (non-consuming) to call_llm. The pin's
single-use contract for the SUMMARY call is unchanged — the
main-model retry still never re-issues the pinned route.
- Worker re-run repeating pre-compression callbacks: documented as an
accepted limitation on _retry_compression_on_fallback_chain
(built-ins idempotent; resuming mid-pipeline would couple the retry
to host callback ordering).
Tests (tests/agent/test_compression_attempt_ownership.py, 10 cases):
deterministic interleavings for both races (late-primary restore
no-ops + preserves fallback state; stale finally cannot clear the
fallback's callback), legacy/slotted compatibility, digest route
follow + context-locality of the consumed echo. Mutation-checked:
reverting only the two prod files to origin/main fails the suite;
restored stack green (21 passed incl. the original #78981 suite).
The one red in the wider sweep
(test_silence_cannot_approach_double_idle_timeout) is pre-existing on
clean origin/main — verified independently.
Review follow-up for #95433. When the host's fence factory is absent or
raises, the retry runs on a private CompressionCommitFence() that
hard-interrupt admission never reads — /stop would serialize against the
aborted attempt's fence instead of the retry's commit boundary. Promote
the factory-failure swallow from debug to warning and warn when the retry
has no published fence, so the control-plane degradation is visible rather
than silent.
Also documents the pin's coverage (single _generate_summary call; the
lean-mode chunk digests are a separate unpinned call path).
A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.
The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
Kanban/background completion wakes persist as role=user rows typed with
display_kind="internal_notification" (the synthetic-wake path in run.py).
The model-payload builder already strips display_kind before the request
and is_user_originated_turn already ignores it, but two compaction scans
still treated those rows as real user turns:
- _is_actionable_user_turn (tail anchor) only checked role/content, so a
notification became the protected 'last user turn' the compressor keeps.
- _derive_auto_focus_topic only skipped synthetic compression turns, so
operational notices leaked into the compact focus hint.
Both now exclude display_kind-typed rows, mirroring the existing
is_user_originated_turn exclusion. No schema change; cache- and
role-alternation-safe.
Behavior-contract tests feed 1,000 operational notifications around one
human turn and assert they never anchor the tail, become the auto-focus
source, or count as actionable user turns.
Fixes#92703
The legacy tail budget scales as threshold×target_ratio, which was designed
around 128K windows at a 50% trigger (~13K tail). On modern big-window
models with raised thresholds it silently hoards: a 1M-window session at
threshold 0.85 keeps a 170K-token verbatim tail (255K soft ceiling) out of
EVERY compaction, so a 540K manual /compress lands at ~290K and every
subsequent turn re-ships the hoard. Nobody chooses this; it is an artifact
of the formula outside its design envelope.
Lean mode (#87326, compaction-v2) was built for exactly this and its recall
was validated in the before/after eval (evals/compaction/results/): clamped
2.5%-of-window tail (10K floor / 25K cap), continuity carried by the
upgraded summary (digests, anchor index, verbatim user messages,
session_search recovery pointers). This flips the DEFAULT to lean; explicit
'tail_mode: legacy' in config keeps the old behavior exactly.
Also fixes a latent bug the flip exposed: update_model() re-assigned the
LEGACY formula directly when recomputing budgets, silently reverting a lean
compressor to the hoard on every mid-session model switch. The recompute
now routes through the mode-aware tail_token_budget property (regression
test included).
Surfaces: context_compressor.py defaults + getattr fallbacks, agent_init
parse default, DEFAULT_CONFIG, gateway _CACHE_BUSTING_CONFIG_KEYS gains
compression.tail_mode (mode changes now evict cached gateway agents like
target_ratio changes do), user + developer docs. Tests: 3 new default
contracts, legacy tests pinned explicitly, feasibility-skip scenario pinned
to legacy (under lean its payloads correctly become compressible).
E2E counterfactual (real imports, 1M window @ 0.85):
main default: legacy, tail 170,000 (ceiling 255,000)
head default: lean, tail 25,000 (ceiling 37,500)
head legacy: 170,000 (opt-out intact)
update_model to 400K: 10,000 (lean preserved across switch)
Follow-up to PR #94531 salvage:
- classify the auxiliary boundary's terminal 'None response' /
'invalid response' errors (#7264) into the same empty-content abort
carve-out so those shapes also preserve the session (#94459's wider
classification, sibling shapes from #94448)
- register _last_summary_empty_content_failure in
_COMPRESSOR_ATTEMPT_STATE_FIELDS so pre-commit hard-cancel rollback
restores the flag (conversation_compression snapshot allow-list)
- tests: cooldown re-entry keeps aborting; both sibling shapes abort
- attribution: map zhangyswx@163.com -> YusenZhang0601
When an auxiliary or main summarizer LLM returns an HTTP 200 with an empty or whitespace-only response (e.g., degraded provider/channel), abort compression and preserve the full conversation context rather than falling through to the destructive static-fallback branch that drops the middle window.
- Track _last_summary_empty_content_failure across _generate_summary() and compress()
- Attempt fallback to the main model when an aux model returns empty content
- Abort compression and preserve all messages intact if no valid summary can be generated
- Record summary_empty_content_failure in telemetry and log actionable diagnostic guidance
- Add comprehensive unit tests in tests/agent/test_context_compressor.py
Fixes#94448
Fixes#93022. A short session (protection window >= transcript) hits the
"insufficient messages" / "no compressible window" branches twice and
permanently trips the anti-thrash breaker, even though nothing was
eligible to compress - compression was never attempted, so there is
nothing "ineffective" to score. The session then rides past the
threshold with no compaction possible (recovery probes only soften,
not fix, the misclassification).
Distinguish "nothing eligible right now" from "attempted and
underperformed":
- New transient _structural_no_op_backoff_until (in-memory, 300s)
armed by _record_structural_no_op() at the three structural no-op
sites: insufficient_messages, no_compressible_window,
empty_post_handoff_window. No strikes accumulate; auto-compaction
resumes on its own once the backoff lapses or the transcript outgrows
the protection window.
- The backoff gates should_compress via
_automatic_compression_blocked_locally and surfaces in
_compression_block_reason as "structural_backoff:<seconds>".
- #40803's frozen-CLI guarantee is preserved: a transcript that can
never shrink retries at most once per backoff window instead of
every turn.
- force=True (/compress) clears an active backoff before attempting;
record_completed_compaction() lifts it - both prove the transcript
is compressible/being worked.
- Genuine attempted-but-underperformed verdicts still strike the
durable ineffective counter unchanged.
Tests: new tests/agent/test_context_compressor_structural_backoff.py;
updated the two tests that asserted the old strike-on-noop behavior.
Follow-up on top of the salvaged #93335:
- context_compressor._sanitize_tool_pairs now expands alias spellings on
the RESULT side too (tool_result_id_variants), so a composite
call|item-keyed result pairs with its split-field tool_call instead of
being dropped and its call stripped.
- The compressor's _tool_call_id_variants staticmethod and
agent_runtime_helpers' module-level _tool_call_id_variants are now thin
forwarders to agent.message_sanitization.tool_call_id_variants — one
policy owner for alias expansion, so the pre-call sanitizer, repair
pass, dedup pass, and compression sanitizer can never drift apart.
- Preserved the #91768 SDK-object tolerance in repair pass 1 (the
shared helper handles non-dict tool_calls via getattr; the salvaged
commit's isinstance-dict guard was dropped in the merge resolution).
New regression tests: composite-keyed results through
sanitize_api_messages (both directions) and _sanitize_tool_pairs, with
negative controls. Sabotage-verified: compressor test fails with raw
tool_call_id tracking.
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
_sanitize_tool_pairs() matched tool_call/tool_result pairs using a
single-value call_id||id precedence per tool_call (_get_tool_call_id).
In the Codex Responses API format an assistant tool_call carries both a
distinct id (fc_...) and call_id (call_...); a tool result's
tool_call_id may be keyed on either depending on which code path built
it. Whenever a genuinely matching pair used the field the precedence
didn't pick, the sanitizer misclassified it as orphaned on BOTH sides:
it dropped the valid tool result AND stripped the tool_call from the
assistant message, even though neither was orphaned.
Live-verified before the fix: {"id": "fc_777", "call_id": "call_777"}
+ a tool result with tool_call_id="fc_777" (a valid pair) was fully
removed by current main.
Register both id and call_id as valid match keys via a new
_tool_call_id_variants() helper (a set per tool_call, not a single
value), matching #58168's fix for repair_message_sequence's known-id
set today. A tool_call now survives if ANY of its id variants has a
matching result, which is not vulnerable to precedence order at all
(unlike swapping which field is checked first, which only trades which
sub-case is broken).
Note on #56425 (open, unreviewed): that PR touches this same function
for the same underlying issue (#55626) by swapping the call_id||id
precedence to id||call_id. That fixes the specific case where a result
matches `id` but not the reverse case (a result matching `call_id`
while `id` is also present) -- the precedence-swap approach cannot fix
the class, only relocate which sub-case is broken. This fix instead
mirrors the already-merged #58168 pattern (register the superset of
both ids as valid matches), which has no such blind spot. Adds 2
regression tests: the previously-mismatched case, and a negative
control confirming genuine orphans are still stripped alongside a
valid dual-id pair in the same window.
The demote pass (pass 2) and the retire pass (3.5, #92783) each carried
their own copy of the two image-strip branches. The copies had already
diverged: the retire pass dropped the stale api_content sidecar on
rewrite, the demote pass did not — leaving an exact-wire sidecar that
replay could use to resend the pre-strip image bytes.
Extract _strip_images_from_tool_msg as the single policy owner; both
passes now use it, closing the sidecar gap in the demote path.
browser_vision's native fast path base64-encoded screenshots at full
resolution and baked them into the tool result uncapped — the exact
sibling of the vision_analyze path #92699 fixed. Apply the same
proactive 256KB/1568px resize before the embed enters reusable history.
Fail-open by design: without Pillow the resize helper falls back to raw
bytes and the compressor's keep-newest pass still retires stale embeds.
Sibling-gap follow-up for the #92725 salvage; the shared-cap approach
mirrors the policy-owner idea from #92748.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Images locked in protect_last_n never shrank, so compression savings
stayed under 10% and anti-thrash disabled further compaction. Keep the
newest three tool-result screenshots live for follow-up QA and replace
older native embeds with placeholders.
Identify a completed merged assistant handoff from the carrier's own stop state instead of an unrelated adjacent history row. Keep carriers with pending tool calls actionable so compaction cannot abort a live tool chain.
Treat a merged assistant-role summary carrier as the driving reference handoff when it immediately follows a completed assistant stop. Its preserved prose and stale tool_calls are assistant continuity, not a fresh live user request.
Keep legitimate in-flight behavior unchanged when there is no completed stop, a real user turn follows, or a distinct later assistant tool-call row continues the loop.
Extends the #80622 active-turn guard for the merged-carrier shape reported under #42768.
Review follow-up on the salvaged #90353:
- Todo snapshot (+ coupled pruned-skill reload notice, 7a16840add) is now
reduced only as a LAST resort after reasoning/tool/summary shrink ops,
and the reload notice survives even then.
- Reuse existing helpers/constants instead of re-hardcoding:
_PRUNED_TOOL_PLACEHOLDER, _PRUNE_MIN_CHARS, _NEWEST_TURN_ONLY_BUDGET_KEYS,
and _prune_stale_reasoning_replay (codex sidecar shrink, #71058 boundary).
- Assistant-role messages without the summary metadata key are no longer
truncatable by the summary-cap heuristic.
- Caller passes budget so the estimator runs 3x, not 5x, per would-grow pass.
The anti-growth guard correctly refuses to persist a compressed
candidate larger than the original, but the rejection was never
recorded by the anti-thrashing breaker: _ineffective_compression_count
stayed at zero, the latch never tripped, and automatic compression
retried the SAME unchanged transcript on every turn - same summary
request, same refusal, same user-facing warning (#88568).
Add ContextCompressor.record_rejected_compaction(): one persisted
ineffective strike, without arming post-compaction real-usage
verification (nothing was committed) and without touching the
fallback-summary streak (no summary was accepted). The would-grow
abort path in conversation_compression calls it before returning the
original transcript. Two refusals latch the normal breaker, manual
/compress keeps bypassing it (force=True), and the existing recovery
window still allows one probe later.
Fixes#88568
A later hygiene idle-timeout write can replace an aux-model cooldown on
the shared column. Drop the in-memory timer on that refresh so the
in-agent compressor is not still blocked after the DB row is hygiene.
Session hygiene persists compression_failure_cooldown_until after a
30s no-progress watchdog so the pre-agent pass can skip. The
in-conversation compressor read the same column and then refused to
run even though its own budget is sufficient.
Ignore hygiene idle-timeout errors on the in-agent path. Real
aux-model faults such as rate limits still block.
Fixes#86972
- _build_anchor_index(): regex-harvests PR/issue numbers, SHAs, branches,
file paths, error strings, handles, URLs from the compacted region into a
bounded indexed summary section. LLM-free, so needle identifiers cannot be
paraphrased away (the GUI-lineage failure class: 10/15 verbatim-or-nothing
golds). Doubles as session_search query-anchor map.
- evals/compaction/test_region_scoping.py: sentinel tripwire proving the
summarizer input carries ONLY the compacted region (head/tail sentinels
never reach the serialized turns body) in both legacy and lean modes.
- _digest_worthy() drops no-signal tool rows before chunking (GUI-lineage
digests were starving on tool-noise)
- eval recovery sim now uses in-memory SQLite FTS5 + BM25 (production
session_search engine) instead of term-frequency scoring
- recovery query writer sees the digest section (front of context) so it can
mine anchor identifiers