Follow-up to PR #94531 salvage:
- classify the auxiliary boundary's terminal 'None response' /
'invalid response' errors (#7264) into the same empty-content abort
carve-out so those shapes also preserve the session (#94459's wider
classification, sibling shapes from #94448)
- register _last_summary_empty_content_failure in
_COMPRESSOR_ATTEMPT_STATE_FIELDS so pre-commit hard-cancel rollback
restores the flag (conversation_compression snapshot allow-list)
- tests: cooldown re-entry keeps aborting; both sibling shapes abort
- attribution: map zhangyswx@163.com -> YusenZhang0601
When an auxiliary or main summarizer LLM returns an HTTP 200 with an empty or whitespace-only response (e.g., degraded provider/channel), abort compression and preserve the full conversation context rather than falling through to the destructive static-fallback branch that drops the middle window.
- Track _last_summary_empty_content_failure across _generate_summary() and compress()
- Attempt fallback to the main model when an aux model returns empty content
- Abort compression and preserve all messages intact if no valid summary can be generated
- Record summary_empty_content_failure in telemetry and log actionable diagnostic guidance
- Add comprehensive unit tests in tests/agent/test_context_compressor.py
Fixes#94448
Fixes#93022. A short session (protection window >= transcript) hits the
"insufficient messages" / "no compressible window" branches twice and
permanently trips the anti-thrash breaker, even though nothing was
eligible to compress - compression was never attempted, so there is
nothing "ineffective" to score. The session then rides past the
threshold with no compaction possible (recovery probes only soften,
not fix, the misclassification).
Distinguish "nothing eligible right now" from "attempted and
underperformed":
- New transient _structural_no_op_backoff_until (in-memory, 300s)
armed by _record_structural_no_op() at the three structural no-op
sites: insufficient_messages, no_compressible_window,
empty_post_handoff_window. No strikes accumulate; auto-compaction
resumes on its own once the backoff lapses or the transcript outgrows
the protection window.
- The backoff gates should_compress via
_automatic_compression_blocked_locally and surfaces in
_compression_block_reason as "structural_backoff:<seconds>".
- #40803's frozen-CLI guarantee is preserved: a transcript that can
never shrink retries at most once per backoff window instead of
every turn.
- force=True (/compress) clears an active backoff before attempting;
record_completed_compaction() lifts it - both prove the transcript
is compressible/being worked.
- Genuine attempted-but-underperformed verdicts still strike the
durable ineffective counter unchanged.
Tests: new tests/agent/test_context_compressor_structural_backoff.py;
updated the two tests that asserted the old strike-on-noop behavior.
Follow-up on top of the salvaged #93335:
- context_compressor._sanitize_tool_pairs now expands alias spellings on
the RESULT side too (tool_result_id_variants), so a composite
call|item-keyed result pairs with its split-field tool_call instead of
being dropped and its call stripped.
- The compressor's _tool_call_id_variants staticmethod and
agent_runtime_helpers' module-level _tool_call_id_variants are now thin
forwarders to agent.message_sanitization.tool_call_id_variants — one
policy owner for alias expansion, so the pre-call sanitizer, repair
pass, dedup pass, and compression sanitizer can never drift apart.
- Preserved the #91768 SDK-object tolerance in repair pass 1 (the
shared helper handles non-dict tool_calls via getattr; the salvaged
commit's isinstance-dict guard was dropped in the merge resolution).
New regression tests: composite-keyed results through
sanitize_api_messages (both directions) and _sanitize_tool_pairs, with
negative controls. Sabotage-verified: compressor test fails with raw
tool_call_id tracking.
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
_sanitize_tool_pairs() matched tool_call/tool_result pairs using a
single-value call_id||id precedence per tool_call (_get_tool_call_id).
In the Codex Responses API format an assistant tool_call carries both a
distinct id (fc_...) and call_id (call_...); a tool result's
tool_call_id may be keyed on either depending on which code path built
it. Whenever a genuinely matching pair used the field the precedence
didn't pick, the sanitizer misclassified it as orphaned on BOTH sides:
it dropped the valid tool result AND stripped the tool_call from the
assistant message, even though neither was orphaned.
Live-verified before the fix: {"id": "fc_777", "call_id": "call_777"}
+ a tool result with tool_call_id="fc_777" (a valid pair) was fully
removed by current main.
Register both id and call_id as valid match keys via a new
_tool_call_id_variants() helper (a set per tool_call, not a single
value), matching #58168's fix for repair_message_sequence's known-id
set today. A tool_call now survives if ANY of its id variants has a
matching result, which is not vulnerable to precedence order at all
(unlike swapping which field is checked first, which only trades which
sub-case is broken).
Note on #56425 (open, unreviewed): that PR touches this same function
for the same underlying issue (#55626) by swapping the call_id||id
precedence to id||call_id. That fixes the specific case where a result
matches `id` but not the reverse case (a result matching `call_id`
while `id` is also present) -- the precedence-swap approach cannot fix
the class, only relocate which sub-case is broken. This fix instead
mirrors the already-merged #58168 pattern (register the superset of
both ids as valid matches), which has no such blind spot. Adds 2
regression tests: the previously-mismatched case, and a negative
control confirming genuine orphans are still stripped alongside a
valid dual-id pair in the same window.
The demote pass (pass 2) and the retire pass (3.5, #92783) each carried
their own copy of the two image-strip branches. The copies had already
diverged: the retire pass dropped the stale api_content sidecar on
rewrite, the demote pass did not — leaving an exact-wire sidecar that
replay could use to resend the pre-strip image bytes.
Extract _strip_images_from_tool_msg as the single policy owner; both
passes now use it, closing the sidecar gap in the demote path.
browser_vision's native fast path base64-encoded screenshots at full
resolution and baked them into the tool result uncapped — the exact
sibling of the vision_analyze path #92699 fixed. Apply the same
proactive 256KB/1568px resize before the embed enters reusable history.
Fail-open by design: without Pillow the resize helper falls back to raw
bytes and the compressor's keep-newest pass still retires stale embeds.
Sibling-gap follow-up for the #92725 salvage; the shared-cap approach
mirrors the policy-owner idea from #92748.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Images locked in protect_last_n never shrank, so compression savings
stayed under 10% and anti-thrash disabled further compaction. Keep the
newest three tool-result screenshots live for follow-up QA and replace
older native embeds with placeholders.
Identify a completed merged assistant handoff from the carrier's own stop state instead of an unrelated adjacent history row. Keep carriers with pending tool calls actionable so compaction cannot abort a live tool chain.
Treat a merged assistant-role summary carrier as the driving reference handoff when it immediately follows a completed assistant stop. Its preserved prose and stale tool_calls are assistant continuity, not a fresh live user request.
Keep legitimate in-flight behavior unchanged when there is no completed stop, a real user turn follows, or a distinct later assistant tool-call row continues the loop.
Extends the #80622 active-turn guard for the merged-carrier shape reported under #42768.
Review follow-up on the salvaged #90353:
- Todo snapshot (+ coupled pruned-skill reload notice, 7a16840add) is now
reduced only as a LAST resort after reasoning/tool/summary shrink ops,
and the reload notice survives even then.
- Reuse existing helpers/constants instead of re-hardcoding:
_PRUNED_TOOL_PLACEHOLDER, _PRUNE_MIN_CHARS, _NEWEST_TURN_ONLY_BUDGET_KEYS,
and _prune_stale_reasoning_replay (codex sidecar shrink, #71058 boundary).
- Assistant-role messages without the summary metadata key are no longer
truncatable by the summary-cap heuristic.
- Caller passes budget so the estimator runs 3x, not 5x, per would-grow pass.
The anti-growth guard correctly refuses to persist a compressed
candidate larger than the original, but the rejection was never
recorded by the anti-thrashing breaker: _ineffective_compression_count
stayed at zero, the latch never tripped, and automatic compression
retried the SAME unchanged transcript on every turn - same summary
request, same refusal, same user-facing warning (#88568).
Add ContextCompressor.record_rejected_compaction(): one persisted
ineffective strike, without arming post-compaction real-usage
verification (nothing was committed) and without touching the
fallback-summary streak (no summary was accepted). The would-grow
abort path in conversation_compression calls it before returning the
original transcript. Two refusals latch the normal breaker, manual
/compress keeps bypassing it (force=True), and the existing recovery
window still allows one probe later.
Fixes#88568
A later hygiene idle-timeout write can replace an aux-model cooldown on
the shared column. Drop the in-memory timer on that refresh so the
in-agent compressor is not still blocked after the DB row is hygiene.
Session hygiene persists compression_failure_cooldown_until after a
30s no-progress watchdog so the pre-agent pass can skip. The
in-conversation compressor read the same column and then refused to
run even though its own budget is sufficient.
Ignore hygiene idle-timeout errors on the in-agent path. Real
aux-model faults such as rate limits still block.
Fixes#86972
- _build_anchor_index(): regex-harvests PR/issue numbers, SHAs, branches,
file paths, error strings, handles, URLs from the compacted region into a
bounded indexed summary section. LLM-free, so needle identifiers cannot be
paraphrased away (the GUI-lineage failure class: 10/15 verbatim-or-nothing
golds). Doubles as session_search query-anchor map.
- evals/compaction/test_region_scoping.py: sentinel tripwire proving the
summarizer input carries ONLY the compacted region (head/tail sentinels
never reach the serialized turns body) in both legacy and lean modes.
- _digest_worthy() drops no-signal tool rows before chunking (GUI-lineage
digests were starving on tool-noise)
- eval recovery sim now uses in-memory SQLite FTS5 + BM25 (production
session_search engine) instead of term-frequency scoring
- recovery query writer sees the digest section (front of context) so it can
mine anchor identifiers
Fix: _LENGTH_CONTINUATION_DROPPED_TOOLS_PREFIX ended with '(' but
_get_continuation_prompt still had f'({tool_list})', producing
'((write_file)' instead of '(write_file)'. Removed the '(' from
the prefix constant — the parenthesis belongs in the interpolation.
Widened: promoted the empty-response nudge (line 6993,
'You just executed tool calls but returned an empty response...')
to _EMPTY_TOOL_RESPONSE_NUDGE constant and added it to the
classifier's recognition set. Same bug class — its
_empty_recovery_synthetic metadata flag doesn't survive SessionDB
projection either.
Test: added parametrize case for the empty-response nudge (7→8 cases).
E2E: verified byte-for-byte string equivalence for all nudge constants.
aed114a69 taught _is_synthetic_compression_user_turn to recognize the
max-iteration nudge as ephemeral runtime scaffolding rather than a human
turn, since its role="user" metadata flag doesn't survive SessionDB
projection and a crash/interrupt mid-turn can persist it durably — becoming
the compaction anchor / auto-focus topic in place of the real task.
conversation_loop.py's retry loop appends several more role="user" rows
with the exact same "ephemeral, metadata-tag-only" shape, none of them
recognized by the classifier:
- The three _get_continuation_prompt variants (length-continuation nudge,
tagged _length_continuation_nudge) — two fixed strings plus a third that
interpolates the dropped-tool-call list.
- _CODEX_INCOMPLETE_NUDGE (codex/responses reasoning-only retry).
- The codex ack-continuation nudge (acknowledgment-only reply re-prompt).
- The dropped-tool-call nudge (tagged _dropped_toolcall_nudge) — persisted
across up to 3 consecutive retries before the finalization pop-loop
strips it; an interrupt/crash before that pop can persist it same as the
max-iteration case.
Promote the previously-inline nudge strings to named module-level constants
in conversation_loop.py (single source of truth for both construction and
recognition), then extend the classifier to recognize all of them — exact
match for the five fixed-content nudges, a stable-prefix check for the
dropped-tool-call continuation variant (its tool list is interpolated so it
can't be exact-matched, same treatment TODO_INJECTION_HEADER already gets).
Imported lazily inside the classifier to avoid a module-load-order cycle —
conversation_loop.py already imports FROM context_compressor.py at call
time for the same reason.
Generic thinking fields (reasoning / reasoning_content + the
reasoning_details text charge) are replayed for at most the NEWEST
assistant turn on every transport: Anthropic strips all-but-newest at
convert time, Bedrock Converse never replays thinking, and strict
chat-completions providers reject or one-space-pad the field. The tail
budget walks charged them on every message anyway, spending 19-24% of
the budget (per the issue's 1,025-message measurement) on bytes that
provably never reach the wire — so the tail cut landed early and each
compaction discarded more real transcript than configured.
_estimate_msg_budget_tokens now partitions the replay keys:
* _ALWAYS_REPLAYED_BUDGET_KEYS (codex_reasoning_items,
codex_message_items) — charged unconditionally. These ride the wire
on every retained turn (#55572), and codex_reasoning_items now also
carries native server-side compaction checkpoints (#81747).
* _NEWEST_TURN_ONLY_BUDGET_KEYS (reasoning, reasoning_content) + the
reasoning_details text charge — charged only for the newest assistant
turn via charge_stale_thinking, resolved by the three budget walks
(tail cut, raw-budget re-walk, proactive-prune boundary).
Default stays the conservative full charge for callers without
turn-position context. A partition invariant test pins that any future
_REPLAY_BUDGET_KEYS entry must be classified into exactly one class.
Direction credit: #73669 (@x7peeps) and #73730 (@webtecnica) both
attacked this; the keep_open reviews asked for provider/API-mode-aware
accounting that keeps Codex carriers charged — this implements that
shape.
Two corrections on top of the #71077 base (the whole bug class):
1. Turn boundary = last USER message, not last assistant message. A Codex
turn spans several assistant messages (assistant+tool_calls -> tool ->
... -> final assistant) whose reasoning items must replay together; the
last-assistant boundary would strip reasoning mid-chain from the active
turn (the gap flagged in PR #71077 review).
2. type="compaction" checkpoints (native server-side compaction, PR #81747)
are exempt: they carry already-pruned history, not per-turn reasoning.
Pruning filters items instead of popping the sidecar key.
Sibling site fixed in the same class: the Codex incomplete-continuation
dedup path blind-overwrote codex_reasoning_items on visually-duplicate
interim messages, which would drop the only copy of a checkpoint captured
on the earlier response. Extracted merge_interim_reasoning_items() into
agent/native_compaction.py; newer reasoning wins, prior checkpoints are
preserved unless the newer payload carries its own.
- Wire the Pass-1 dedup floor (len < 200) to the shared _PRUNE_MIN_CHARS
constant it was already documented as matching, and use the constant in
the remaining test literal.
- Restructure the clarify 'resolved' computation (is_answer_shaped +
sentinel check) instead of compute-then-flip.
- Add a live producer->recognizer drift guard: the REAL oneshot no-user
callback's output must be recognized as a sentinel, so producer wording
drift fails a test instead of silently reintroducing false attribution.
- Document the any()-poisoning semantic for multi-select sentinel lists.
Follow-up to the salvaged #81244 commits:
- Timeout/no-user clarify callbacks (CLI timeout, gateway timeout and
delivery failure, oneshot no-user) embed sentinel prose as
user_response; quoting those as '[clarify] user responded: ...' would
be false attribution. Route them to the generic summary path.
- Extract the shared _PRUNE_MIN_CHARS = 200 floor (prune default +
proactive clamp) and cap the clarify summary at _PRUNE_MIN_CHARS - 1,
removing the knife-edge equality the summary's survival depended on
and keeping it out of the >=200-char dedup pass.
- Tests: 4 sentinel shapes + multi-select sentinel; mutation-checked.
handle_max_iterations() appends its runtime summary request as a plain
role="user" row, which SessionDB persists verbatim. On later compaction the
synthetic-turn filters only recognized compaction summaries, continuation
rows, and todo snapshots, so the nudge could be selected as the latest
actionable user turn — becoming the task snapshot / auto-focus input and
getting summarized as "User asked: ...", demoting the real human task.
Metadata flags do not survive SessionDB projection (the reason the existing
markers are content-based), so recognition must key off stable content.
Extract the nudge into a shared MAX_ITERATIONS_SUMMARY_REQUEST constant and
teach _is_synthetic_compression_user_turn() to recognize it, mirroring the
continuation/todo markers. Every _is_actionable_user_turn call site already
pairs the synthetic guard, so the single recognizer change covers anchor
selection, auto-focus, and real-user-turn detection.
Fixes#78580
Review-pass follow-ups (three parallel reviewers, findings verified):
- hermes_state_search.py list_recent_user_messages now drops legacy
standalone compaction handoffs in the decode loop (SQL can't see them:
durable role=user, no display_kind). Closes the /undo N pairing skew
where the in-memory count (new predicate) and the DB soft-delete pick
(old predicate) targeted different turns on legacy sessions. Fetches
with headroom so the requested limit is still honored. 3 new tests,
mutation-checked (no-op'ing the skip fails 2/3).
- _should_skip_model_call_for_reference_handoff: single drive-check scan
(was two — once inside the restore helper, once after); the restore
helper no longer re-scans and its return value now decides the verdict.
- _final_response_from_messages replaced by the _HANDOFF_SKIP_FINAL_RESPONSE
constant it always returned (parameter was unused).
- _handoff_carries_live_user_content delegates to the canonical
_strip_context_summary_handoff_message — also fixes the edge where a
merged-shaped row with an EMPTY preserved prior tail was wrongly
treated as carrying live content.
- Site-level guard test for rollback.restore with a legacy handoff row
(predicate-in-context, complements the unit tests).
Follow-ups on top of the salvaged #80696 fix (review findings):
- Sibling sites: rollback.restore, gateway /retry, CLI /retry and /undo N,
and both CLI resume turn counters now use is_user_originated_turn so
legacy-persisted standalone handoffs (durable role=user, no display_kind)
can never be truncation targets or counted as user turns (#80622
suggested regression 4, dispatcher-wide).
- Site-1 guard: hoist the api_call_count decrement + iteration-budget
refund above the break so a skipped turn no longer leaks a budget unit
and finalize_turn logs the true call count (matches the ollama early-exit
and the site-2 sibling).
- Site-2 guard: run the handoff guard BEFORE reanchoring so a restored
user ask is what the anchor lands on, not a stale pre-restore index.
- SUMMARY_PREFIX: add the mid-tool-loop carve-out the code-side guard
already implements, so a literal-minded model doesn't halt an in-flight
exchange after in-place compaction.
- Skip path returns a short compaction status instead of replaying the
previous turn's answer (finalize_turn would append it as a fresh
assistant row — duplicate prose in transcript and delivery).
After a completed assistant stop, a standalone CONTEXT COMPACTION handoff
could occupy the sole user slot and resume stale Historical Task Snapshot
work with no new human ask. Guard post-compaction continues, hide
standalone handoffs from session dispatch, and harden SUMMARY_PREFIX for
the empty-after-handoff case (#80622).
Two docstring corrections on top of the salvaged #80997 fix — behavior
unchanged, both verified against the code:
- The 'rough growth over-counts every content class' claim is false for
Cyrillic/Greek/Thai/Arabic (chars/4 vs ~2-3 chars/token on o200k):
growth there can under-count up to ~2x (#62605's direction). Document
the real backstops instead: an at/over-threshold real reading clears
the baseline (post-response gate fires on real usage within one call)
and the provider overflow handler compacts reactively.
- Document the two measurement bases (turn-prologue raw messages vs the
loop's fully assembled request that seeds the baseline) and why the
prologue's smaller basis can only OVER-defer — the loop's pre-API
pressure check re-runs the projection with the aligned basis before
every provider call, so a prologue over-defer never skips a needed
compaction.
The rough preflight estimate intentionally overestimates, but not by a
fixed margin: CJK text is counted at ~1.7x its o200k cost and
Responses-mode reasoning replay blobs at several times their billed
cost. Heavy sessions show rough estimates 2-3x real usage and compact
at 35-55% of the real window, stalling turns for minutes and discarding
detail (churn), because the defer guard only tolerated 5% rough growth
and sessions that never compressed had no baseline at all.
Pair every request's rough estimate (note_request_rough_estimate,
recorded in the conversation loop right after the pressure estimate)
with the provider's real prompt_tokens in update_from_response(), then
defer preflight while projected real usage — last real + rough growth
since that reading — stays under the threshold. Rough growth is itself
an overestimate of real growth, so the projection is an upper bound and
deferring below the threshold is safe; the provider's context-overflow
handler remains the backstop.
The baseline no longer ratchets on defer: it is refreshed by the
response pairing, and advancing it without a matching real reading
would shrink apparent growth and defer on stale data.
Review follow-up: after the walk-back widening, the exempted assistant
is often not the final message, and the partial-batch shape is
byte-identical to a settled-but-malformed orphan — so say 'presumed
pending' and document WHY presuming is safe (sanitize_api_messages
step 2 stubs any genuinely unanswered call pre-API on every path).
Widen #79293's trailing-in-flight guard from 'last message is assistant'
to 'last non-tool message is assistant': a multi-call batch snapshotted
between the executor's per-result appends looks like
[..., assistant(c1,c2,c3), tool(c1)] — c2/c3 are pending, not orphaned,
but the tail-only guard missed that shape and stripped them (same silent
result loss as the original bug, via concurrent /compress or the gateway
hygiene pass).
Preserving is safe on both shapes: the pre-API chokepoint
(sanitize_api_messages step 2) injects stub results for any call that
genuinely never gets an answer, while stripping a live call silently
loses its late result.
test_sanitizer_strips_orphaned_keeps_valid's mixed valid/orphan shape
moves mid-list — at the tail it is byte-identical to a live partial
batch and the sanitizer now correctly presumes in-flight there.
New regression test fails without the walk-back (c2/c3 stripped),
passes with it.
Tool_executor.py appends role=tool results AFTER running each call. When
context compression fires mid-chain, the trailing assistant(tool_calls)
message is a pending request whose result has not yet been appended.
_sanitize_tool_pairs previously stripped it as an 'orphan', so when the
executor later appended the real result, repair_message_sequence dropped
it as unmatched and the completed side effect (and final synthesis) was
lost. Preserve the trailing in-flight call verbatim; only genuinely
orphaned calls in the discarded region are stripped.
Adds regression tests: three unit tests for _sanitize_tool_pairs plus an
end-to-end test reproducing compression -> side-effect completion ->
result-returned flow. Confirmed failing on pre-fix code, passing with
the fix.
Three review follow-ups on the salvaged #79286 commit:
- update_model() zeroed the in-memory prune runway but left the durable
model_config copy stale, breaking the method's own durable-sync
discipline (the strike reset three lines above keeps its durable copy
in sync). A restart after a model switch resurrected a runway
computed under the old model's trigger sizes. New
_clear_durable_proactive_prune_rearm() removes the persisted key via
patch_session_model_config() without touching the transcript.
- The archive_and_compact capability check ran AFTER the expensive
3-pass prune scan, so a duck-typed session store lacking the method
paid the full scan on every eligible iteration forever with pruning
permanently no-opping. Hoist it above the scan (all in-tree stores
pass a real SessionDB; this only affects third-party stores).
- _load_proactive_prune_rearm_tokens now uses the shared
get_session_model_config_value() accessor instead of inlining a 5th
copy of the model_config JSON parse, matching its sibling loaders'
typed-accessor pattern.
Also documents why the rotation-publish-failure branch restores only
the runway field rather than the full attempt snapshot.
Tests: model-switch durable clear, patch_session_model_config
merge/delete/no-op, and a guard proving incapable stores skip the scan.