Follow-up to #96634 (stall-fallback retry, #78981) addressing
donovan-yohan's post-merge adversarial review. The stall path detaches a
timed-out primary worker (fence cancel wins; future stays on the pool)
and immediately runs the fallback against the SAME ContextCompressor,
creating two verified races:
1. Late-primary snapshot restore: the detached primary's unwind called
_restore_compressor_attempt_state with the PRIMARY's pre-attempt
snapshot. Landing after the fallback's commit it rolled
_previous_summary/cooldown/provenance/telemetry back to pre-primary
values, silently discarding fallback-owned state.
2. Shared _compression_cancelled_check: the late primary's `finally`
cleared the callback the fallback had just installed, so the
fallback's F4 cancellation consult read None.
Fix: a monotonic per-compressor attempt generation claimed under one
module lock (_claim_compressor_attempt). Snapshot restores carry their
claiming generation and no-op when stale; the cancelled-check set/clear
moves into owner-stamped helpers (_install_compression_cancelled_check /
_clear_compression_cancelled_check_if_owner) so only the installing
attempt can clear it. The commit fence keeps owning COMMIT admission;
the generation owns compressor-ATTRIBUTE writes — two boundaries.
Legacy callers (attempt_generation=None) and slotted third-party
compressors (generation 0) keep the historical unconditional behavior.
Secondary review items:
- Lean chunk digests during a stall-fallback retry now follow the
summary onto the pinned healthy route: take_pinned_summary_route()
echoes the consumed route into a context-local
_SUMMARY_ROUTE_CONSUMED, and _build_chunk_digests passes
attempt_summary_route_kwargs() (non-consuming) to call_llm. The pin's
single-use contract for the SUMMARY call is unchanged — the
main-model retry still never re-issues the pinned route.
- Worker re-run repeating pre-compression callbacks: documented as an
accepted limitation on _retry_compression_on_fallback_chain
(built-ins idempotent; resuming mid-pipeline would couple the retry
to host callback ordering).
Tests (tests/agent/test_compression_attempt_ownership.py, 10 cases):
deterministic interleavings for both races (late-primary restore
no-ops + preserves fallback state; stale finally cannot clear the
fallback's callback), legacy/slotted compatibility, digest route
follow + context-locality of the consumed echo. Mutation-checked:
reverting only the two prod files to origin/main fails the suite;
restored stack green (21 passed incl. the original #78981 suite).
The one red in the wider sweep
(test_silence_cannot_approach_double_idle_timeout) is pre-existing on
clean origin/main — verified independently.
Review follow-up for #95433. When the host's fence factory is absent or
raises, the retry runs on a private CompressionCommitFence() that
hard-interrupt admission never reads — /stop would serialize against the
aborted attempt's fence instead of the retry's commit boundary. Promote
the factory-failure swallow from debug to warning and warn when the retry
has no published fence, so the control-plane degradation is visible rather
than silent.
Also documents the pin's coverage (single _generate_summary call; the
lean-mode chunk digests are a separate unpinned call path).
A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.
The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
Kanban/background completion wakes persist as role=user rows typed with
display_kind="internal_notification" (the synthetic-wake path in run.py).
The model-payload builder already strips display_kind before the request
and is_user_originated_turn already ignores it, but two compaction scans
still treated those rows as real user turns:
- _is_actionable_user_turn (tail anchor) only checked role/content, so a
notification became the protected 'last user turn' the compressor keeps.
- _derive_auto_focus_topic only skipped synthetic compression turns, so
operational notices leaked into the compact focus hint.
Both now exclude display_kind-typed rows, mirroring the existing
is_user_originated_turn exclusion. No schema change; cache- and
role-alternation-safe.
Behavior-contract tests feed 1,000 operational notifications around one
human turn and assert they never anchor the tail, become the auto-focus
source, or count as actionable user turns.
Fixes#92703
The legacy tail budget scales as threshold×target_ratio, which was designed
around 128K windows at a 50% trigger (~13K tail). On modern big-window
models with raised thresholds it silently hoards: a 1M-window session at
threshold 0.85 keeps a 170K-token verbatim tail (255K soft ceiling) out of
EVERY compaction, so a 540K manual /compress lands at ~290K and every
subsequent turn re-ships the hoard. Nobody chooses this; it is an artifact
of the formula outside its design envelope.
Lean mode (#87326, compaction-v2) was built for exactly this and its recall
was validated in the before/after eval (evals/compaction/results/): clamped
2.5%-of-window tail (10K floor / 25K cap), continuity carried by the
upgraded summary (digests, anchor index, verbatim user messages,
session_search recovery pointers). This flips the DEFAULT to lean; explicit
'tail_mode: legacy' in config keeps the old behavior exactly.
Also fixes a latent bug the flip exposed: update_model() re-assigned the
LEGACY formula directly when recomputing budgets, silently reverting a lean
compressor to the hoard on every mid-session model switch. The recompute
now routes through the mode-aware tail_token_budget property (regression
test included).
Surfaces: context_compressor.py defaults + getattr fallbacks, agent_init
parse default, DEFAULT_CONFIG, gateway _CACHE_BUSTING_CONFIG_KEYS gains
compression.tail_mode (mode changes now evict cached gateway agents like
target_ratio changes do), user + developer docs. Tests: 3 new default
contracts, legacy tests pinned explicitly, feasibility-skip scenario pinned
to legacy (under lean its payloads correctly become compressible).
E2E counterfactual (real imports, 1M window @ 0.85):
main default: legacy, tail 170,000 (ceiling 255,000)
head default: lean, tail 25,000 (ceiling 37,500)
head legacy: 170,000 (opt-out intact)
update_model to 400K: 10,000 (lean preserved across switch)
Follow-up to PR #94531 salvage:
- classify the auxiliary boundary's terminal 'None response' /
'invalid response' errors (#7264) into the same empty-content abort
carve-out so those shapes also preserve the session (#94459's wider
classification, sibling shapes from #94448)
- register _last_summary_empty_content_failure in
_COMPRESSOR_ATTEMPT_STATE_FIELDS so pre-commit hard-cancel rollback
restores the flag (conversation_compression snapshot allow-list)
- tests: cooldown re-entry keeps aborting; both sibling shapes abort
- attribution: map zhangyswx@163.com -> YusenZhang0601
When an auxiliary or main summarizer LLM returns an HTTP 200 with an empty or whitespace-only response (e.g., degraded provider/channel), abort compression and preserve the full conversation context rather than falling through to the destructive static-fallback branch that drops the middle window.
- Track _last_summary_empty_content_failure across _generate_summary() and compress()
- Attempt fallback to the main model when an aux model returns empty content
- Abort compression and preserve all messages intact if no valid summary can be generated
- Record summary_empty_content_failure in telemetry and log actionable diagnostic guidance
- Add comprehensive unit tests in tests/agent/test_context_compressor.py
Fixes#94448
Fixes#93022. A short session (protection window >= transcript) hits the
"insufficient messages" / "no compressible window" branches twice and
permanently trips the anti-thrash breaker, even though nothing was
eligible to compress - compression was never attempted, so there is
nothing "ineffective" to score. The session then rides past the
threshold with no compaction possible (recovery probes only soften,
not fix, the misclassification).
Distinguish "nothing eligible right now" from "attempted and
underperformed":
- New transient _structural_no_op_backoff_until (in-memory, 300s)
armed by _record_structural_no_op() at the three structural no-op
sites: insufficient_messages, no_compressible_window,
empty_post_handoff_window. No strikes accumulate; auto-compaction
resumes on its own once the backoff lapses or the transcript outgrows
the protection window.
- The backoff gates should_compress via
_automatic_compression_blocked_locally and surfaces in
_compression_block_reason as "structural_backoff:<seconds>".
- #40803's frozen-CLI guarantee is preserved: a transcript that can
never shrink retries at most once per backoff window instead of
every turn.
- force=True (/compress) clears an active backoff before attempting;
record_completed_compaction() lifts it - both prove the transcript
is compressible/being worked.
- Genuine attempted-but-underperformed verdicts still strike the
durable ineffective counter unchanged.
Tests: new tests/agent/test_context_compressor_structural_backoff.py;
updated the two tests that asserted the old strike-on-noop behavior.
Follow-up on top of the salvaged #93335:
- context_compressor._sanitize_tool_pairs now expands alias spellings on
the RESULT side too (tool_result_id_variants), so a composite
call|item-keyed result pairs with its split-field tool_call instead of
being dropped and its call stripped.
- The compressor's _tool_call_id_variants staticmethod and
agent_runtime_helpers' module-level _tool_call_id_variants are now thin
forwarders to agent.message_sanitization.tool_call_id_variants — one
policy owner for alias expansion, so the pre-call sanitizer, repair
pass, dedup pass, and compression sanitizer can never drift apart.
- Preserved the #91768 SDK-object tolerance in repair pass 1 (the
shared helper handles non-dict tool_calls via getattr; the salvaged
commit's isinstance-dict guard was dropped in the merge resolution).
New regression tests: composite-keyed results through
sanitize_api_messages (both directions) and _sanitize_tool_pairs, with
negative controls. Sabotage-verified: compressor test fails with raw
tool_call_id tracking.
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
_sanitize_tool_pairs() matched tool_call/tool_result pairs using a
single-value call_id||id precedence per tool_call (_get_tool_call_id).
In the Codex Responses API format an assistant tool_call carries both a
distinct id (fc_...) and call_id (call_...); a tool result's
tool_call_id may be keyed on either depending on which code path built
it. Whenever a genuinely matching pair used the field the precedence
didn't pick, the sanitizer misclassified it as orphaned on BOTH sides:
it dropped the valid tool result AND stripped the tool_call from the
assistant message, even though neither was orphaned.
Live-verified before the fix: {"id": "fc_777", "call_id": "call_777"}
+ a tool result with tool_call_id="fc_777" (a valid pair) was fully
removed by current main.
Register both id and call_id as valid match keys via a new
_tool_call_id_variants() helper (a set per tool_call, not a single
value), matching #58168's fix for repair_message_sequence's known-id
set today. A tool_call now survives if ANY of its id variants has a
matching result, which is not vulnerable to precedence order at all
(unlike swapping which field is checked first, which only trades which
sub-case is broken).
Note on #56425 (open, unreviewed): that PR touches this same function
for the same underlying issue (#55626) by swapping the call_id||id
precedence to id||call_id. That fixes the specific case where a result
matches `id` but not the reverse case (a result matching `call_id`
while `id` is also present) -- the precedence-swap approach cannot fix
the class, only relocate which sub-case is broken. This fix instead
mirrors the already-merged #58168 pattern (register the superset of
both ids as valid matches), which has no such blind spot. Adds 2
regression tests: the previously-mismatched case, and a negative
control confirming genuine orphans are still stripped alongside a
valid dual-id pair in the same window.
The demote pass (pass 2) and the retire pass (3.5, #92783) each carried
their own copy of the two image-strip branches. The copies had already
diverged: the retire pass dropped the stale api_content sidecar on
rewrite, the demote pass did not — leaving an exact-wire sidecar that
replay could use to resend the pre-strip image bytes.
Extract _strip_images_from_tool_msg as the single policy owner; both
passes now use it, closing the sidecar gap in the demote path.
browser_vision's native fast path base64-encoded screenshots at full
resolution and baked them into the tool result uncapped — the exact
sibling of the vision_analyze path #92699 fixed. Apply the same
proactive 256KB/1568px resize before the embed enters reusable history.
Fail-open by design: without Pillow the resize helper falls back to raw
bytes and the compressor's keep-newest pass still retires stale embeds.
Sibling-gap follow-up for the #92725 salvage; the shared-cap approach
mirrors the policy-owner idea from #92748.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Images locked in protect_last_n never shrank, so compression savings
stayed under 10% and anti-thrash disabled further compaction. Keep the
newest three tool-result screenshots live for follow-up QA and replace
older native embeds with placeholders.
Identify a completed merged assistant handoff from the carrier's own stop state instead of an unrelated adjacent history row. Keep carriers with pending tool calls actionable so compaction cannot abort a live tool chain.
Treat a merged assistant-role summary carrier as the driving reference handoff when it immediately follows a completed assistant stop. Its preserved prose and stale tool_calls are assistant continuity, not a fresh live user request.
Keep legitimate in-flight behavior unchanged when there is no completed stop, a real user turn follows, or a distinct later assistant tool-call row continues the loop.
Extends the #80622 active-turn guard for the merged-carrier shape reported under #42768.
Review follow-up on the salvaged #90353:
- Todo snapshot (+ coupled pruned-skill reload notice, 7a16840add) is now
reduced only as a LAST resort after reasoning/tool/summary shrink ops,
and the reload notice survives even then.
- Reuse existing helpers/constants instead of re-hardcoding:
_PRUNED_TOOL_PLACEHOLDER, _PRUNE_MIN_CHARS, _NEWEST_TURN_ONLY_BUDGET_KEYS,
and _prune_stale_reasoning_replay (codex sidecar shrink, #71058 boundary).
- Assistant-role messages without the summary metadata key are no longer
truncatable by the summary-cap heuristic.
- Caller passes budget so the estimator runs 3x, not 5x, per would-grow pass.
The anti-growth guard correctly refuses to persist a compressed
candidate larger than the original, but the rejection was never
recorded by the anti-thrashing breaker: _ineffective_compression_count
stayed at zero, the latch never tripped, and automatic compression
retried the SAME unchanged transcript on every turn - same summary
request, same refusal, same user-facing warning (#88568).
Add ContextCompressor.record_rejected_compaction(): one persisted
ineffective strike, without arming post-compaction real-usage
verification (nothing was committed) and without touching the
fallback-summary streak (no summary was accepted). The would-grow
abort path in conversation_compression calls it before returning the
original transcript. Two refusals latch the normal breaker, manual
/compress keeps bypassing it (force=True), and the existing recovery
window still allows one probe later.
Fixes#88568
A later hygiene idle-timeout write can replace an aux-model cooldown on
the shared column. Drop the in-memory timer on that refresh so the
in-agent compressor is not still blocked after the DB row is hygiene.
Session hygiene persists compression_failure_cooldown_until after a
30s no-progress watchdog so the pre-agent pass can skip. The
in-conversation compressor read the same column and then refused to
run even though its own budget is sufficient.
Ignore hygiene idle-timeout errors on the in-agent path. Real
aux-model faults such as rate limits still block.
Fixes#86972
- _build_anchor_index(): regex-harvests PR/issue numbers, SHAs, branches,
file paths, error strings, handles, URLs from the compacted region into a
bounded indexed summary section. LLM-free, so needle identifiers cannot be
paraphrased away (the GUI-lineage failure class: 10/15 verbatim-or-nothing
golds). Doubles as session_search query-anchor map.
- evals/compaction/test_region_scoping.py: sentinel tripwire proving the
summarizer input carries ONLY the compacted region (head/tail sentinels
never reach the serialized turns body) in both legacy and lean modes.
- _digest_worthy() drops no-signal tool rows before chunking (GUI-lineage
digests were starving on tool-noise)
- eval recovery sim now uses in-memory SQLite FTS5 + BM25 (production
session_search engine) instead of term-frequency scoring
- recovery query writer sees the digest section (front of context) so it can
mine anchor identifiers
Fix: _LENGTH_CONTINUATION_DROPPED_TOOLS_PREFIX ended with '(' but
_get_continuation_prompt still had f'({tool_list})', producing
'((write_file)' instead of '(write_file)'. Removed the '(' from
the prefix constant — the parenthesis belongs in the interpolation.
Widened: promoted the empty-response nudge (line 6993,
'You just executed tool calls but returned an empty response...')
to _EMPTY_TOOL_RESPONSE_NUDGE constant and added it to the
classifier's recognition set. Same bug class — its
_empty_recovery_synthetic metadata flag doesn't survive SessionDB
projection either.
Test: added parametrize case for the empty-response nudge (7→8 cases).
E2E: verified byte-for-byte string equivalence for all nudge constants.
aed114a69 taught _is_synthetic_compression_user_turn to recognize the
max-iteration nudge as ephemeral runtime scaffolding rather than a human
turn, since its role="user" metadata flag doesn't survive SessionDB
projection and a crash/interrupt mid-turn can persist it durably — becoming
the compaction anchor / auto-focus topic in place of the real task.
conversation_loop.py's retry loop appends several more role="user" rows
with the exact same "ephemeral, metadata-tag-only" shape, none of them
recognized by the classifier:
- The three _get_continuation_prompt variants (length-continuation nudge,
tagged _length_continuation_nudge) — two fixed strings plus a third that
interpolates the dropped-tool-call list.
- _CODEX_INCOMPLETE_NUDGE (codex/responses reasoning-only retry).
- The codex ack-continuation nudge (acknowledgment-only reply re-prompt).
- The dropped-tool-call nudge (tagged _dropped_toolcall_nudge) — persisted
across up to 3 consecutive retries before the finalization pop-loop
strips it; an interrupt/crash before that pop can persist it same as the
max-iteration case.
Promote the previously-inline nudge strings to named module-level constants
in conversation_loop.py (single source of truth for both construction and
recognition), then extend the classifier to recognize all of them — exact
match for the five fixed-content nudges, a stable-prefix check for the
dropped-tool-call continuation variant (its tool list is interpolated so it
can't be exact-matched, same treatment TODO_INJECTION_HEADER already gets).
Imported lazily inside the classifier to avoid a module-load-order cycle —
conversation_loop.py already imports FROM context_compressor.py at call
time for the same reason.
Generic thinking fields (reasoning / reasoning_content + the
reasoning_details text charge) are replayed for at most the NEWEST
assistant turn on every transport: Anthropic strips all-but-newest at
convert time, Bedrock Converse never replays thinking, and strict
chat-completions providers reject or one-space-pad the field. The tail
budget walks charged them on every message anyway, spending 19-24% of
the budget (per the issue's 1,025-message measurement) on bytes that
provably never reach the wire — so the tail cut landed early and each
compaction discarded more real transcript than configured.
_estimate_msg_budget_tokens now partitions the replay keys:
* _ALWAYS_REPLAYED_BUDGET_KEYS (codex_reasoning_items,
codex_message_items) — charged unconditionally. These ride the wire
on every retained turn (#55572), and codex_reasoning_items now also
carries native server-side compaction checkpoints (#81747).
* _NEWEST_TURN_ONLY_BUDGET_KEYS (reasoning, reasoning_content) + the
reasoning_details text charge — charged only for the newest assistant
turn via charge_stale_thinking, resolved by the three budget walks
(tail cut, raw-budget re-walk, proactive-prune boundary).
Default stays the conservative full charge for callers without
turn-position context. A partition invariant test pins that any future
_REPLAY_BUDGET_KEYS entry must be classified into exactly one class.
Direction credit: #73669 (@x7peeps) and #73730 (@webtecnica) both
attacked this; the keep_open reviews asked for provider/API-mode-aware
accounting that keeps Codex carriers charged — this implements that
shape.
Two corrections on top of the #71077 base (the whole bug class):
1. Turn boundary = last USER message, not last assistant message. A Codex
turn spans several assistant messages (assistant+tool_calls -> tool ->
... -> final assistant) whose reasoning items must replay together; the
last-assistant boundary would strip reasoning mid-chain from the active
turn (the gap flagged in PR #71077 review).
2. type="compaction" checkpoints (native server-side compaction, PR #81747)
are exempt: they carry already-pruned history, not per-turn reasoning.
Pruning filters items instead of popping the sidecar key.
Sibling site fixed in the same class: the Codex incomplete-continuation
dedup path blind-overwrote codex_reasoning_items on visually-duplicate
interim messages, which would drop the only copy of a checkpoint captured
on the earlier response. Extracted merge_interim_reasoning_items() into
agent/native_compaction.py; newer reasoning wins, prior checkpoints are
preserved unless the newer payload carries its own.
- Wire the Pass-1 dedup floor (len < 200) to the shared _PRUNE_MIN_CHARS
constant it was already documented as matching, and use the constant in
the remaining test literal.
- Restructure the clarify 'resolved' computation (is_answer_shaped +
sentinel check) instead of compute-then-flip.
- Add a live producer->recognizer drift guard: the REAL oneshot no-user
callback's output must be recognized as a sentinel, so producer wording
drift fails a test instead of silently reintroducing false attribution.
- Document the any()-poisoning semantic for multi-select sentinel lists.
Follow-up to the salvaged #81244 commits:
- Timeout/no-user clarify callbacks (CLI timeout, gateway timeout and
delivery failure, oneshot no-user) embed sentinel prose as
user_response; quoting those as '[clarify] user responded: ...' would
be false attribution. Route them to the generic summary path.
- Extract the shared _PRUNE_MIN_CHARS = 200 floor (prune default +
proactive clamp) and cap the clarify summary at _PRUNE_MIN_CHARS - 1,
removing the knife-edge equality the summary's survival depended on
and keeping it out of the >=200-char dedup pass.
- Tests: 4 sentinel shapes + multi-select sentinel; mutation-checked.
handle_max_iterations() appends its runtime summary request as a plain
role="user" row, which SessionDB persists verbatim. On later compaction the
synthetic-turn filters only recognized compaction summaries, continuation
rows, and todo snapshots, so the nudge could be selected as the latest
actionable user turn — becoming the task snapshot / auto-focus input and
getting summarized as "User asked: ...", demoting the real human task.
Metadata flags do not survive SessionDB projection (the reason the existing
markers are content-based), so recognition must key off stable content.
Extract the nudge into a shared MAX_ITERATIONS_SUMMARY_REQUEST constant and
teach _is_synthetic_compression_user_turn() to recognize it, mirroring the
continuation/todo markers. Every _is_actionable_user_turn call site already
pairs the synthetic guard, so the single recognizer change covers anchor
selection, auto-focus, and real-user-turn detection.
Fixes#78580
Review-pass follow-ups (three parallel reviewers, findings verified):
- hermes_state_search.py list_recent_user_messages now drops legacy
standalone compaction handoffs in the decode loop (SQL can't see them:
durable role=user, no display_kind). Closes the /undo N pairing skew
where the in-memory count (new predicate) and the DB soft-delete pick
(old predicate) targeted different turns on legacy sessions. Fetches
with headroom so the requested limit is still honored. 3 new tests,
mutation-checked (no-op'ing the skip fails 2/3).
- _should_skip_model_call_for_reference_handoff: single drive-check scan
(was two — once inside the restore helper, once after); the restore
helper no longer re-scans and its return value now decides the verdict.
- _final_response_from_messages replaced by the _HANDOFF_SKIP_FINAL_RESPONSE
constant it always returned (parameter was unused).
- _handoff_carries_live_user_content delegates to the canonical
_strip_context_summary_handoff_message — also fixes the edge where a
merged-shaped row with an EMPTY preserved prior tail was wrongly
treated as carrying live content.
- Site-level guard test for rollback.restore with a legacy handoff row
(predicate-in-context, complements the unit tests).