Commit Graph

190 Commits

Author SHA1 Message Date
Teknium 3a2eceabc8 refactor(agent/compression): remove dead memory-containment helpers; fold codex skip-log ladder
- delete _cached_prompt_reflects_builtin_memory + _builtin_memory_prompt_snapshot
  (zero callers since the commit site moved to byte-equality; the only test
  reference asserts their ABSENCE from the commit window)
- delete CompressionExecutorSaturatedError (never raised or caught anywhere)
- _compress_context_via_codex_app_server: three near-identical skip branches
  collapse into one skip_reason + single log line (same message text)
2026-09-02 13:29:37 -07:00
Teknium 7e3cef7d87 refactor(agent/compression): extract candidate rejection ladder from compress_context
_candidate_rejected(): compressor abort, no-progress, empty transcript and
superseded-attempt checks in their original order; caller handles the
shared release + unchanged-return. Aborted-branch try/finally collapsed
(the finally only released the lease, which the caller now does).
2026-09-02 13:29:37 -07:00
Teknium 0ea4c7963a refactor(agent/compression): extract post-commit boundary tail of compress_context
- _finish_compaction_boundary(): parent label clear, context-engine + memory
  provider notifications, compression-count warning, session:compress hook,
  usage re-arm, file/skill dedup reset (zero back-refs; returns rough estimate)
- _warn_summary_or_aux_fallback(): once-only summary/aux-model warnings
compress_context 794 -> 653 lines; bodies moved verbatim.
2026-09-02 13:29:37 -07:00
Teknium 01d16451f7 refactor(agent/compression): extract commit-phase helpers out of compress_context
- _fold_todo_snapshot(): stale-snapshot strip + live todo fold into the tail
- _rebuild_system_prompt_at_boundary(): tool refresh + byte-equal keep-prompt
- _salvage_or_refuse_grown_transcript(): commit-site anti-growth guard
- _publish_rotated_compaction(): parent flush, child publish, id re-point,
  goal/heartbeat/loop/title carry-over
- old_session_id is now an explicit Optional local instead of locals().get()
compress_context 1181 -> 794 lines; bodies moved verbatim.
2026-09-02 13:29:37 -07:00
Teknium cbd709124f refactor(agent/compression): extract pre-summary phases of compress_context
- _adopt_grown_durable_parent(): rotation-only durable-snapshot adoption
- _pre_compress_memory_context(): provider on_pre_compress / checkpoint gate
- _resolve_compress_call(): compress() + signature-filtered kwargs
- _run_summary_dispatch(): progress hook, stream deadline, interrupt guard
compress_context 1359 -> 1181 lines; bodies moved verbatim.
2026-09-02 13:29:37 -07:00
Teknium 42c41dfcff refactor(agent/compression): extract lease acquisition and lifecycle out of compress_context
- _CompactionLifecycle: the one-shot terminal status edge (was a nonlocal closure)
- _CompressionLease: holder/watermark/ttl + refresher, holder-only release,
  fence lock-setup bracket, release() (was 4 nested closures + 9 locals)
- _acquire_compression_lease(): the 190-line lock acquisition ladder
- _adopt_if_parent_rotated(): late-contender adoption check
compress_context 1767 -> 1359 lines; bodies moved verbatim.
2026-09-02 13:29:37 -07:00
Teknium 852f2a02d8 refactor(agent/compression): unify repeated abort-path snippets in compress_context
- _existing_system_prompt: 21 copies of cached-or-rebuild prompt
- _emit_aborted_attempt_telemetry: 11 copies of aborted/aborted telemetry
- _restore_messages_snapshot: 4 copies of deepcopy rollback
- _restore_prune_rearm_tokens: 3 copies of prune-runway restore
2026-09-02 13:29:37 -07:00
Teknium 682d48c48a refactor(agent/compression): compact comments and docstrings in conversation_compression (AST-identical) 2026-09-02 13:29:37 -07:00
Teknium d1efa0d78d fix(compression): provider-proven overflow gets one real compaction attempt while the failure cooldown is armed
After one failed/stalled summary attempt arms the 60/300/900s compression-
failure cooldown, a provider context_length_exceeded rejection entered the
reactive overflow branch in conversation_loop, which called _compress_context
without force. Since #97488 the cooldown gate returns the soft "temporarily
paused, retry in a moment" deferral instead of exhaustion, so every turn
deferred until the cooldown lapsed, and the next failure extended the ladder:
long-running sessions wedged with no automatic recovery (#100661, four sessions
lost).

Thread a narrow `bypass_cooldown` kwarg from the three provider-proven overflow
call sites (generic overflow, 413, output-cap recovery) through
AIAgent._compress_context -> compress_context -> ContextCompressor.compress ->
_generate_summary. It skips ONLY the summary-failure cooldown check at each gate.
Unlike force=True it does not clear the cooldown, does not skip the feasibility /
anti-thrash breakers, and a failed attempt records its cooldown normally. The
attempt is bounded by the existing compression_attempts/max_compression_attempts
budget, so there is no retry loop. The preflight threshold gate is unchanged:
ordinary over-threshold pressure still honors the cooldown (#11529).

Engines whose _automatic_compression_blocked()/compress() predate the kwarg
(plugins, test doubles) are called with the legacy signature.

Tests: cooldown armed + bypass_cooldown -> summarizer invoked and transcript
compacted; ordinary pass still deferred. Docs note the cooldown/overflow
contract in the developer guide.

Fixes #100661
Closes #97766 (overflow-force idea; the bundled continuation changes were not taken)

Co-authored-by: sgtworkman <178342791+sgtworkman@users.noreply.github.com>
2026-09-02 05:33:22 -07:00
Teknium bc71b8bc95 fix(compression): anchor on the LAST intent row — newer user turn outranks older steer (#100053 follow-up)
Follow-up to the salvaged #100114 commit. Its two-pass anchor selection
scanned steers first and real user rows second, so a transcript shaped
[user A, tool(steer B), ..., user C] anchored the already-consumed steer B
over the newer real request C — the same replay class the PR set out to
fix. Replace it with one reversed positional scan that picks whichever
intent-bearing row is last (real role=user or steer-bearing role=tool),
and make the compressed-transcript steer check count only role=tool rows
(the only place the runtime delivers a steer), so a summary quoting the
marker cannot masquerade as live intent.

Adds S1/S2/S3 regression tests (steer dropped by compaction, steer
surviving in tail, newer user turn after steer) plus alternation and
use-exactly-once assertions.
2026-09-02 04:12:12 -07:00
finn763 f40333a80e fix(agent): preserve busy steer during compression and avoid replaying historical user request
Compression with display.busy_input_mode: steer embeds the follow-up
as an out-of-band marker inside the latest role=tool result. The
post-compression user-turn preservation path only classified
non-scaffolding role=user rows as real intent, so a compressed
transcript that contained no role=user row would discard the steer
and clone an older historical role=user message as the new active
turn, re-activating a previously consumed request.

Fix _ensure_compressed_has_user_turn to (1) treat a compressed
transcript that already carries a steer marker as having user intent,
and (2) prioritize the latest steer payload from the original
transcript over historical user cloning, inserting it as a proper
role=user turn via _insert_real_user_anchor. This preserves the
actual current intent exactly once and never turns history into new
input.

Closes #100053
2026-09-02 04:12:12 -07:00
Teknium 9de9d7613c fix(compression): keep hygiene turn-hold worker's commit admission so thinking-model summaries are adopted, not burned
The 10s hygiene_max_turn_hold_seconds budget (#92318) releases the arriving
user turn while the summary model is still streaming. For thinking summary
models (DeepSeek-V4-Flash etc.) whose reasoning prefix alone exceeds 10s,
the abandonment path ALWAYS cancelled the commit fence — 100% of the summary
attempt (including the full thinking prefix) was discarded on every turn,
permanently disabling auto-compression while paying the summary model 10s
of thinking per turn, and the flat 60s retry-after then blocked the
agent-side preflight from a fresh chance.

Structural fix (maintainer-chosen direction in #97963): decouple the turn
from the compression instead of holding the turn longer or making the hold
progress-aware (which would reintroduce the #90845 frozen-turn bug):

- CompressionCommitFence gains mark_commit_watermark_fenced() /
  commit_watermark_fenced; compress_context marks the fence right after
  capturing get_active_message_watermark() under the durable compression
  lock (#75316/#87484) — the property that makes a LATE commit safe: rows
  appended after compression start survive both commit paths verbatim as
  cloned concurrent tail (archive_and_compact watermark= and
  publish_compression_child watermark/watermark_ceiling).
- gateway hygiene turn-hold handler: when the fence is watermark-fenced,
  the detached worker (already kept alive via
  _defer_agent_cleanup_until_future_done) KEEPS its commit admission; the
  user's turn proceeds on the uncompressed transcript at the same 10s
  budget, and the summary is adopted at the worker's own watermark-fenced
  commit boundary. Unfenced workers are cancelled exactly as before —
  never worse than the status quo.
- No retry-after is armed while the kept-admission attempt runs (it would
  block preflight adoption via the same-session cooldown); re-attempt
  spacing is covered by the durable compression lock
  (_session_has_compression_in_flight). If the worker ends WITHOUT
  committing, a done-callback restores the flat non-escalating 60s
  retry-after; a successful adoption resets the hygiene failure streak.
  The streak never advances for a deferral either way.
- Docs: configuration.md hygiene_max_turn_hold_seconds one-liner updated
  to describe deferred adoption and the thinking-model case;
  config_defaults.py comment updated. Knob stays config.yaml-only.

Invariants preserved:
- 10s user-latency cap stays hard (#90845/#92318):
  test_session_hygiene_turn_hold_budget_abandons_streaming_wait passes
  UNMODIFIED (its worker is not watermark-fenced, so it pins the cancel
  path through the public surface).
- Stale-clobber impossible: adoption only rides commits bounded by the
  start watermark; the fence still gates admission and unfenced/late
  results are discarded.

New regression tests (tests/gateway/test_session_hygiene_turnhold_adoption.py):
- watermark-fenced worker keeps admission, late summary is committed,
  turn still released at the budget, no cooldown while running,
  streak reset on adoption;
- kept-admission worker that ends without committing restores the flat
  turn-hold retry-after (<=120s, names turn-hold, streak untouched);
- unfenced worker still cancelled and discarded (status quo).
Sabotage-verified: disabling the keep-admission branch fails the two new
adoption tests and leaves the unfenced-cancel test green.

Fixes #97963
2026-09-01 23:56:23 -07:00
joaomarcos 904e5bb572 fix(compression): stop the summary stream at the host's own deadline
CompressionCommitFence.set_total_ceiling_seconds documents its deadline as
"shared by the host and worker", but only the host ever read it. The worker's
streamed summary bounds itself with _aux_stream_total_ceiling() instead —
max(600, 4 * aux_timeout) — which is >= the host's total ceiling for every
configured timeout AND starts counting later (after pool admission,
_serialize_for_summary, prompt build and TTFT). A stream that outlives its
abandoned host is therefore not an edge case; it is the guaranteed outcome of
every total-ceiling timeout.

8207862212 closed the first half: a cancelled fence now releases the
compression owner, freeing its pool slot and session lease. Its own comment
leaves the second half open — the isolated provider daemon that holds the
socket keeps streaming "until the auxiliary stream's longer absolute ceiling
expires". With the #99692 reporter's auxiliary.compression.timeout: 600 that
is 2400s of an orphaned ~500K-token summary the fence is already guaranteed to
refuse, and because the session never shrank, every following turn stacks a
fresh orphan on top of the last.

Publish the fence's deadline as an absolute monotonic instant
(CompressionCommitFence.deadline_monotonic) and give the auxiliary layer the
return leg it was missing: aux_stream_deadline() installs it thread-locally,
_ChatStreamAccumulator.feed() stops the stream once it passes, and
_run_protected_sync_provider_call propagates it onto the provider daemon
(thread-locals do not cross that boundary, so an owner-thread-only install
would be inert on exactly the path large-session compression takes).

Absolute, not relative: the deadline is unaffected by however long dispatch and
TTFT took before the accumulator was constructed. Checked as well as — not
instead of — the existing ceiling, so every caller without a host deadline is
byte-for-byte unchanged, and the "timed out" phrasing keeps _is_timeout_error
classification identical to a request timeout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EKrRS7LVgyHf2WQkEahSwu
2026-09-01 23:56:06 -07:00
Teknium c0495c6bce fix(cli): context meter no longer sawtooths on reasoning models — show durable transcript, not last-request replay
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.

- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
  provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
  delta estimate (stale reasoning excluded on all but the newest assistant
  message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
  figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.

Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
2026-09-01 15:34:03 -07:00
joaomarcos ecdbcef7af fix(compression): roll the live transcript back when an in-place compaction commit fails (#99477)
`archive_and_compact()` is atomic: when it raises, every pre-compaction row is
still `active = 1` and the compacted set was never inserted. The rotation branch
already rolled the live transcript back to `messages_before_compression` in that
case, but the in-place branch — the default (`compression_in_place` defaults to
True) — did not, so `compress_context()` handed the caller the uncommitted
compacted list.

That list is marker-swept by `_strip_persistence_markers` (#57491) and the
post-commit `stamp_db_persisted_markers` (#98450) never ran, so the next
append-only flush treated the whole compacted transcript as new and INSERTed it
on top of the rows it was supposed to replace. The active set then held the
summary AND the turns it summarized: the next resume reloaded both, the token
count went up, preflight fired again, and every failed attempt appended another
copy of the protected head plus tail.

The in-place rollback mirrors the rotation branch and is gated on
`split_status != "in_place_committed"`, which is assigned on the statement
immediately after the atomic commit returns, so a committed compaction can never
be rolled back into a mismatch of the opposite sign.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EKrRS7LVgyHf2WQkEahSwu
2026-09-01 09:28:31 -07:00
theo 8207862212 fix(compression): stop timeout paths from blocking retries 2026-08-31 13:00:33 -07:00
fangliquanflq 53c0df6de9 fix(agent): stop compression retries after host timeout (#98722)
Salvaged from #98741, composed on top of the merged #98424 preflight
fail-closed boundary. A host-ceiling compression timeout is now a typed,
thread-safe outcome consumed by every automatic caller:

- conversation_compression.py: threading.local + per-agent lock timeout
  state (mark/reset/read helpers) upgrading #98424's simple attribute
  where overlapping automatic/manual compression entrypoints matter;
  the _last_compression_timed_out attribute stays as compat mirror.
- conversation_loop.py: the mid-turn pre-API pass and the provider
  overflow (413/400 context_length_exceeded) recovery path end the turn
  with the typed compression_exhausted recovery contract instead of
  re-sending the unchanged oversized request and re-entering compression
  in the same turn.
- run_agent.py/turn_context.py: forwarder resets the typed state per
  attempt; the #98424 turn-start check reads it through the typed helper.

Tests: thread-safety/atomicity of the state helpers, overflow-recovery
non-re-entry, and typed terminal result.
2026-08-31 12:36:02 -07:00
Teknium d8f8a07ee3 fix(compression): truncated summaries no longer become compaction checkpoints (port of earendil-works/pi#7048)
A summarization response with finish_reason == "length" contains PARTIAL
text — the generation stopped on the output-token cap mid-summary.
Previously all compressor summarization sites accepted such responses as
complete: the cut-off text replaced the real middle turns AND was fed back
into every subsequent iterative-update prompt, compounding the loss across
compactions.

Guards added at all four summarization sites (whole bug class):
- _generate_summary: length stop raises, gets the existing one-shot
  main-model fallback (a larger output budget may finish the summary), and
  on terminal failure ABORTS compression preserving the session unchanged
  (new _last_summary_truncated_failure flag, same class as empty-content).
- _micro_summarize_one: partial rolling-summary merge is discarded; the
  exchange stays unabsorbed for a later pass.
- _build_chunk_digests: partial lean digest degrades to the
  recover-via-session_search placeholder.
- trajectory_compressor (sync + async): length stop raises into the
  existing retry/backoff loop.

_response_finish_reason() reads dict- and object-shaped responses and
returns "" when the provider omits the field, so proxies that never send
finish_reason are unaffected.

Ported from earendil-works/pi commit 97fa14e39 (pi#7048), adapted to
hermes' abort-preserving compression failure machinery.

Tests: tests/agent/test_compressor_truncated_summary_guard.py (12 tests;
sabotage-verified — disabling the guards fails 4).
2026-08-31 10:39:55 -07:00
Teknium 0c24b3a12c fix(compression): pop the tail tags before the anti-growth estimate; count against the final list 2026-08-31 10:02:22 -07:00
BrunoBza 9d9d9194d4 fix(state): archive carried-forward compaction tail as rewind rows (#86366)
archive_and_compact() soft-archives every active row with compacted=1 and
then re-inserts compacted_messages as fresh live rows. When the
compressor's protected tail rides inside that list verbatim - which is
the normal batch-compaction shape ([summary] + tail) - the tail's
ORIGINALS end up stored twice per compaction: (active=0, compacted=1)
next to their live clones. search_messages() recalls both flags without
DISTINCT, so every carried-forward message came back once per compaction
(measured up to 4 identical hits) and was mislabeled to users and the
agent as archived "summarized away" content.

Add an optional tail_count parameter: the last tail_count archived rows
are superseded byte-identical duplicates, stamped rewind-style
(active=0, compacted=0, hidden from recall) instead of compacted=1.

Callers:
- batch in-place compaction counts the compressor-tagged tail dicts
  (_COMPACTION_TAIL_MARKER set by compress() on every carried-forward
  message);
- micro-compaction splices [prefix, marker, suffix] - everything except
  the single marker row is carried forward, so tail_count=len-1;
- proactive tool-result pruning rewrites content in place (not verbatim),
  keeping the historical archive-everything behavior.

Fixes #86366
2026-08-31 10:02:22 -07:00
Teknium c0667439ec fix(compression): rotation heals stale automatic ended_at stamps instead of wedging (#88197)
TUI server shutdown stamps ended_at/end_reason='tui_shutdown' on sessions
whose agent keeps running; every rotation then aborts at
publish_compression_child's liveness check forever (the #88197 wedge; the
amplification half was fixed by #88411).

Class fix: is_automatic_end_reason() in hermes_state_common owns the
"accidental infrastructure cleanup vs deliberate boundary" taxonomy.
publish_compression_child clears automatic stamps in its own transaction
and proceeds (parent re-closes with its TRUE boundary,
end_reason='compression'); the #88411 pre-flush guard no longer aborts on
stamps the publish can heal. Deliberate boundaries (compression,
session_reset, explicit close) still fail closed at both sites.

TEST REPIN (deliberate contract change):
test_ended_parent_aborts_before_the_prepublish_flush pinned
"tui_shutdown stamp => rotation aborts and parent must not grow" — the
abort it required IS the #88197 wedge. Repinned as two tests:
- test_automatic_stamp_no_longer_wedges_rotation: automatic stamp =>
  rotation COMMITS (no abort loop, so no growth-by-abort is possible);
- test_deliberately_ended_parent_aborts_before_the_prepublish_flush:
  session_reset (deliberate boundary) => still aborts BEFORE the #47202
  flush, preserving #88411's no-growth contract where an abort remains
  correct.
The class invariant "no aborted rotation grows the parent" holds
everywhere: automatic stamps no longer produce aborts, deliberate
boundaries still abort pre-flush.
2026-08-31 09:58:11 -07:00
HexLab98 3a542bbef4 fix(tui): show status while idle/auto compaction runs
Idle and preflight compaction arrived as lifecycle status without the
"Compacting context" marker, so TUI never entered a compacting state.
Re-tag those lines and freeze the busy FaceTicker on "compacting" for
the whole pause instead of restoring "running…" after 4s.
2026-08-31 09:57:04 -07:00
kshitijk4poor 3aee290899 refactor(compression): name the split-failure cooldown; drop duplicate tests
Review folds from the formal gate battery:
- _SPLIT_FAILURE_COOLDOWN_SECONDS = 60 replaces the bare literal, with a
  comment pinning WHY it is the timeout ladder's first rung (transient
  lease/DB condition) rather than the 600s summary-provider cooldown.
- publish_compression_child docstring now states the compression_lock_holder
  condition on the refresh guard.
- Dropped 2 of 3 extracted unit tests as duplicates of existing coverage in
  test_compression_rotation_state.py / test_context_compressor.py; kept the
  force-bypass test (only site pinning that behavior for split failures) and
  the E2E test (now asserting the named constant).
2026-08-31 14:09:42 +05:30
kshitijk4poor dc71fb37e1 fix(compression): guard the split-failure cooldown call like sibling strikes
The cooldown is recorded inside the split-failure except handler; a raw
call on a stub/partial compressor would replace the real split error with
an AttributeError. Mirror the try/except-debug convention of the adjacent
record_rejected_compaction call.
2026-08-31 14:09:42 +05:30
VVV 087cc49a26 fix(compression): refresh lease in-transaction before publish; arm cooldown on split failure
Two narrow repairs for #97948 symptom B (large-session rotation aborts with
'Compression lease lost before publication' / session_split_failed, then the
next turn re-runs the identical doomed compression):

1. publish_compression_child gains require_lease_refresh: the lease is
   extended inside the same transaction as the expiry check (same conn, no
   TOCTOU), giving a worker whose refresher thread died from transient DB
   failures one final chance to keep its completed work.

2. A failed compression split now records a 60s failure cooldown, so the
   next turn cannot immediately re-trigger the same compression.

Salvaged from #98137 (author: vsd2807). The timeout-reconciliation half of
that PR is NOT carried: it has a blocking review (runtime sid vs persisted
session_key, one-shot check cannot observe a 6-minute commit, no identity
projection) and needs a redesign.
2026-08-31 14:09:42 +05:30
Teknium 452f6b7de2 fix(compression): route-aware stale-thinking charge parity between compaction trigger and tail walks (#84371)
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).

Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.

Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
2026-08-30 20:40:43 -07:00
Teknium ad925a08da fix(compression): scope worker-teardown grace to the total-ceiling path
The bounded-grace join only applies where the overlap hazard lives: a
total-ceiling expiry over a still-streaming worker (#97488). The
idle-stall path keeps its prompt detachment so the stall-fallback retry
preserves the #76354 S3 latency contract (silence never approaches 2x
the idle budget); its late unwind stays safe behind the fence poison
and attempt-generation supersession.
2026-08-30 19:46:21 -07:00
Teknium 892f756b8c fix(compression): tear down cancelled workers with bounded grace and discard superseded attempts (#97488)
A ceiling/idle-timeout host now joins its fence-cancelled worker for a
bounded grace before returning. A cooperative worker (which polls the
poison fence between provider phases) is reaped, proving quiescence, so
the durable lease releases normally. An uninterruptible worker is
orphaned behind the poison fence: its late result is discarded, and on
the total-ceiling path the holder-qualified lease stays retained until
it exits so no new attempt can overlap the unchanged session.

Supersession: a late candidate from an attempt whose compressor
generation was claimed by a newer attempt is discarded before the
commit boundary (failure_class=attempt_superseded), never committed
over newer state — covering fenceless callers the fence poison cannot
see.
2026-08-30 19:46:21 -07:00
HexLab98 027b339e79 fix(compression): persist stall-interrupted backoff on pre-commit cancel (#96775)
An explicit /stop after the summary stream has already gone idle restored the original transcript but left no durable cooldown, so the next automatic turn re-entered the same stalled strategy. Record a stall-specific failure on that path only, merge with any longer live deadline, and keep an ordinary early /stop cooldown-neutral.
2026-08-30 19:46:21 -07:00
fangliquanflq de4155b1bb fix(compression): report total ceiling expiry accurately 2026-08-30 19:46:21 -07:00
fangliquanflq 0bcad4519e fix(compression): preserve fallback before worker start 2026-08-30 19:46:21 -07:00
fangliquanflq 8d567ccd22 fix(compression): stop work at the total deadline 2026-08-30 19:46:21 -07:00
Teknium 1f2bd9e763 fix(compression): stamp _DB_PERSISTED_MARKER after in-place batch compaction commit (#98450)
compress() returns marker-swept copies (_strip_persistence_markers, #57491);
the in-place branch committed them via archive_and_compact() but never
stamped the persistence marker, so the next _persist_session ->
_flush_messages_to_session_db_unlocked walk re-INSERTed the whole
post-compaction transcript (live set regrew ~58K -> ~512K tokens).

Centralize the post-commit contract in a shared helper,
stamp_db_persisted_markers(), used by all three archive_and_compact
callers: the in-place batch commit (previously missing), the
micro-compaction sync, and the proactive tool-result prune.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-08-30 19:46:11 -07:00
Teknium 514707ff3e feat(compaction): system prompt always rebuilds at the commit boundary — updates finally reach long-lived sessions (#98426)
* feat(compaction): always rebuild the system prompt at the commit boundary — keep-prompt now gated on byte equality of the LIVE builder output; plugin sections re-render with fail-open to last good bytes

* feat(clock): 'Conversation started' resolves through the session-lineage ROOT — a compacted/rotated session keeps its original birth date (Bot Mode forever-chats know when they were first born)

* test: retire old-contract pins — plugin sections re-render at invalidate (freeze stays restore-only), commit boundary always runs the live builder, byte-equal keep preserves object identity
2026-08-30 01:20:27 -07:00
Screaming Sun 989e48cc85 fix(compression): retire completed todo snapshots safely 2026-08-29 20:39:04 -07:00
ericmaddox 1762d3788c fix(compression): strip embedded stale todo snapshot from list message content 2026-08-29 20:39:04 -07:00
Hyusein Leshov 835a913ffd fix(compression): arm the failure cooldown when codex compaction fails
Closes #75364.

`_compress_context_via_codex_app_server` returns the transcript unchanged
when the codex thread reports `interrupted` or `error`. The session is
therefore still above threshold, and nothing records that the attempt
failed — so the next turn retries immediately, and keeps retrying for as
long as the condition persists.

Every other compression path arms the shared failure cooldown, records an
ineffective-compression strike, or both. This path records neither:

* `_hygiene_compression_failure_cooldowns` is set only on
  `asyncio.TimeoutError`, or behind `_last_compress_aborted`, which is
  assigned exclusively in `context_compressor.py` on the Hermes summarizer
  path.
* `compression_ineffective_count` lives in `ContextCompressor`, and this
  path returns before any compressor bookkeeping runs.

`compress_context` already documents the rule this path was missing —
"Every automatic entrypoint must honor compressor-owned cooldown and
breaker state" — but the codex branch dispatches above that block and
returns from inside it.

`result.interrupted` needs no unusual configuration to occur: an ordinary
user message arriving mid-compaction sets it (see
`codex_app_server_session.py`, which produces the "compact turn
interrupted" string). Observed in production on a Discord gateway session
at ~315k tokens against a 258k window, where compaction was attempted on
essentially every turn for ~70 minutes; the session's
`compression_ineffective_count` was still 0 afterwards.

This reuses the existing cooldown rather than adding a new mechanism:

* arm `_record_compression_failure_cooldown` with the existing
  `_SUMMARY_FAILURE_COOLDOWN_SECONDS` when compaction returns
  interrupted/error;
* honor an active cooldown on entry, matching the Hermes path.

`force=True` bypasses both, so an explicit /compress is never braked by a
failure it did not cause, and a successful compaction arms nothing.
2026-08-29 22:29:28 +05:30
Teknium d3a1c46510 feat(agent): context size anchors on provider-reported usage — estimation shrinks to the last turn
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.

- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
  with a structural base-message identity check that fails closed on any
  transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
  site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
  pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
  compaction (codex_runtime), session reset/switch (run_agent), plus the
  fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
  estimation fallback (first request of a session).
2026-08-28 07:51:31 -07:00
Teknium c30ac90a92 feat(compaction): rebuild dynamic tool schemas at the compaction commit boundary — forever-sessions finally pick up config changes (#97073) 2026-08-28 04:01:05 -07:00
kshitijk4poor e078b2fe7c fix(compression): close the restore TOCTOU; fold review findings
Post-review hardening on the attempt-ownership commit:

- Write-time re-validation: the entry staleness check in
  _restore_compressor_attempt_state runs before the durable-cooldown DB
  I/O, so a fallback could claim the compressor in that window and the
  stale setattr loop would still clobber its state. The in-memory writes
  now re-validate AND execute under _COMPRESSOR_ATTEMPT_LOCK — the same
  lock claims are taken under. The DB rollback stays outside the lock
  (safe: the dangerous direction requires a prior claim, which the entry
  check rejects). Both the quality reviewer and the lead's independent
  pre-verification converged on this window.
  New deterministic test: TestMidRestoreClaimRace (claim injected
  between entry check and write via instrumented deepcopy).

- Documented gen-0 semantics on _claim_compressor_attempt: per-compressor
  all-or-nothing, never mixed with gen>0 on one instance (reviewer
  finding 2, verified unreachable — comment hardens against future
  confusion).

Dropped after verification (reviewer finding 3): resetting
_SUMMARY_ROUTE_CONSUMED on pin_summary_route exit — the echo lives in
the worker thread's COPIED context (propagate_context_to_thread) and
dies with it; a probe confirmed the next attempt's context is clean.
Resetting it would break digests running after the with-block.
2026-08-28 12:52:53 +05:30
kshitijk4poor 61cd299c6e fix(compression): attempt-generation ownership for overlapping stall-fallback attempts
Follow-up to #96634 (stall-fallback retry, #78981) addressing
donovan-yohan's post-merge adversarial review. The stall path detaches a
timed-out primary worker (fence cancel wins; future stays on the pool)
and immediately runs the fallback against the SAME ContextCompressor,
creating two verified races:

1. Late-primary snapshot restore: the detached primary's unwind called
   _restore_compressor_attempt_state with the PRIMARY's pre-attempt
   snapshot. Landing after the fallback's commit it rolled
   _previous_summary/cooldown/provenance/telemetry back to pre-primary
   values, silently discarding fallback-owned state.
2. Shared _compression_cancelled_check: the late primary's `finally`
   cleared the callback the fallback had just installed, so the
   fallback's F4 cancellation consult read None.

Fix: a monotonic per-compressor attempt generation claimed under one
module lock (_claim_compressor_attempt). Snapshot restores carry their
claiming generation and no-op when stale; the cancelled-check set/clear
moves into owner-stamped helpers (_install_compression_cancelled_check /
_clear_compression_cancelled_check_if_owner) so only the installing
attempt can clear it. The commit fence keeps owning COMMIT admission;
the generation owns compressor-ATTRIBUTE writes — two boundaries.
Legacy callers (attempt_generation=None) and slotted third-party
compressors (generation 0) keep the historical unconditional behavior.

Secondary review items:
- Lean chunk digests during a stall-fallback retry now follow the
  summary onto the pinned healthy route: take_pinned_summary_route()
  echoes the consumed route into a context-local
  _SUMMARY_ROUTE_CONSUMED, and _build_chunk_digests passes
  attempt_summary_route_kwargs() (non-consuming) to call_llm. The pin's
  single-use contract for the SUMMARY call is unchanged — the
  main-model retry still never re-issues the pinned route.
- Worker re-run repeating pre-compression callbacks: documented as an
  accepted limitation on _retry_compression_on_fallback_chain
  (built-ins idempotent; resuming mid-pipeline would couple the retry
  to host callback ordering).

Tests (tests/agent/test_compression_attempt_ownership.py, 10 cases):
deterministic interleavings for both races (late-primary restore
no-ops + preserves fallback state; stale finally cannot clear the
fallback's callback), legacy/slotted compatibility, digest route
follow + context-locality of the consumed echo. Mutation-checked:
reverting only the two prod files to origin/main fails the suite;
restored stack green (21 passed incl. the original #78981 suite).

The one red in the wider sweep
(test_silence_cannot_approach_double_idle_timeout) is pre-existing on
clean origin/main — verified independently.
2026-08-28 12:52:53 +05:30
Mike DeMott 213ae08e7a perf(compression): add guarded fast summary lane 2026-08-28 12:38:49 +05:30
kshitijk4poor 80ab7d2b1c fix(compression): dedupe current-turn rows when rotation splits the session mid-turn
When context-compression rotation fires mid-turn, the current user
message was persisted twice into the child session. Root cause: dedup
used id()-seeded sets of copies instead of markers on the live objects.

Replace with _DB_PERSISTED_MARKER-based dedup as the sole authority:
- _ensure_compressed_has_user_turn returns CompressedUserTurnOutcome
- After publish_compression_child succeeds, stamp the live anchor-source
  row (not a drifted index) with _DB_PERSISTED_MARKER
- _sync_persisted_markers mirrors stamps from result to live lists by
  scoped identity (handles direct-path, adoption divergence, _session_messages)
- Remove _flushed_db_message_ids from rotation commit path (markers replace it)
- Unconditional (loud) imports — no silent fallback

Salvage of #94996 by @fedosis, rebased on top of #95433 (stall-fallback,
already merged). Both conversation_compression.py and run_agent.py are
built from origin/main + #94996's diff applied on top, preserving the
force_terminal refactor and _publish_new_fence from #95433.

Credit: @fedosis original PR #94996.
2026-08-28 02:47:34 +05:30
kshitijk4poor ba7df6a4bf refactor: extract shared _coerce_positive_timeout helper (#95433)
The timeout validation (isinstance(raw, (int, float)) and not
isinstance(raw, bool) and raw > 0 → float(raw)) was duplicated
between auxiliary_client.py:_fallback_entry_timeout and
conversation_compression.py:resolve_compression_fallback_route.
Extracted into _coerce_positive_timeout to prevent drift if the
config schema evolves.
2026-08-28 02:29:27 +05:30
Shaun Eccles 6151e59d65 fix(compression): surface an unpublished stall-fallback fence at WARNING
Review follow-up for #95433. When the host's fence factory is absent or
raises, the retry runs on a private CompressionCommitFence() that
hard-interrupt admission never reads — /stop would serialize against the
aborted attempt's fence instead of the retry's commit boundary. Promote
the factory-failure swallow from debug to warning and warn when the retry
has no published fence, so the control-plane degradation is visible rather
than silent.

Also documents the pin's coverage (single _generate_summary call; the
lean-mode chunk digests are a separate unpinned call path).
2026-08-28 02:29:27 +05:30
Shaun Eccles 2c6938dc3a fix(compression): retry a stalled summary on the fallback chain (#78981)
A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.

The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
2026-08-28 02:29:27 +05:30
kshitijk4poor c693c772b5 refactor(compression): fold review follow-ups on #71488 salvage
- Reuse the existing _commit_status variable for the terminal-edge gate
  instead of the parallel _compaction_succeeded boolean (derived state).
- Give the commit_fence_cancelled abort the same force_terminal=True
  terminal edge as the lock-contended abort, and reword the closure
  comment that overstated the lock contender as 'the one exception'.
- Inline the codex app-server path's lifecycle closure: after gating on
  success it reduced to a single success-site emit, so the scaffolding
  (done-flag + closure + two no-op failure-path calls) was dead.
2026-08-27 12:42:38 +05:30
Uttkarsh Tiwari 36ae32524c fix(compression): preserve terminal lifecycle for lock skips 2026-08-27 12:42:38 +05:30
Uttkarsh Tiwari 7a21bfe68a fix(compression): suppress duplicate completion notices 2026-08-27 12:42:38 +05:30
kshitijk4poor 4ba2608524 fix(compressor): widen empty-content abort to sibling no-response shapes + snapshot state field
Follow-up to PR #94531 salvage:
- classify the auxiliary boundary's terminal 'None response' /
  'invalid response' errors (#7264) into the same empty-content abort
  carve-out so those shapes also preserve the session (#94459's wider
  classification, sibling shapes from #94448)
- register _last_summary_empty_content_failure in
  _COMPRESSOR_ATTEMPT_STATE_FIELDS so pre-commit hard-cancel rollback
  restores the flag (conversation_compression snapshot allow-list)
- tests: cooldown re-entry keeps aborting; both sibling shapes abort
- attribution: map zhangyswx@163.com -> YusenZhang0601
2026-08-26 13:02:27 +05:30