Commit Graph

28271 Commits

Author SHA1 Message Date
Teknium 64cc87e668 fix(compression): keep estimate seam positional-compatible for monkeypatched estimators
Test seams and plugin engines monkeypatch estimate_messages_tokens_rough with (messages)-only signatures; route callers only pass the charge_stale_thinking kwarg on the False path.
2026-08-30 20:40:43 -07:00
Teknium 452f6b7de2 fix(compression): route-aware stale-thinking charge parity between compaction trigger and tail walks (#84371)
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).

Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.

Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
2026-08-30 20:40:43 -07:00
Teknium 5a134383fe fix: failed subagents now surface a clean error to the user (CLI + gateway)
A delegate_task child that died (provider 404/400, timeout, crash)
previously vanished silently: the child's conversation loop returns
failed=True with the error summary in final_response, which the
classifier treated as usable output -> status 'completed'. And even
correctly-failed children only reached the parent MODEL — platforms
with tool_progress off (Telegram/Slack defaults) never showed the
human anything.

- delegate_tool: result.failed now forces status 'failed' (with the
  error carried on the entry); new shared format_subagent_failure_line()
  renders one clean human-readable line (traceback -> exception message,
  length-capped); CLI tree + batch ✗ lines now include the reason.
- gateway TurnRunner.progress_callback: subagent.complete events with a
  terminal failure status deliver that line via _deliver_platform_notice
  BEFORE all progress-queue gates; tool_progress_callback is now always
  attached (body gates each event class itself).
- tests: failed-flag classification regression + notice rendering suite.
- docs: Failure Visibility section in delegation docs.
2026-08-30 20:40:14 -07:00
Teknium e730deedd1 feat(bot-mode): Group Chats survive the authority gateway dying — log replication and fenced takeover
Every participant gateway can now keep a durable copy of a hosted room's
ordered log and continue the room when its authority host is gone:

- gateway/hosted_room_replicas.py: replica store in root state.db.
  ingest_page() persists authority-stamped groups.log pages idempotently,
  refusing sequence gaps and authority-epoch regressions. promote_replica()
  continues the room locally at epoch+1 with a lineage-proving
  authority.claimed event; the stale owner is fenced everywhere the claim
  replicates. demote_room() lets a returning stale authority fence itself
  (authority.lost) upon observing a newer epoch, killing split-brain writes.
- tui_gateway/methods_groups.py: groups.replicate / groups.replica_state /
  groups.promote / groups.demote RPC surface. Promotion requires
  confirm=true — storage decides HOW takeover is atomic and provable, the
  caller (user action now, lease/quorum driver later) decides WHEN it is
  safe, matching the boundary blessed on #97681.

Validation: 20 new tests incl. a full failover round-trip (A hosts, B
replicates incrementally, A dies, B promotes with complete history, A
returns demoted and fenced); 69 total across the hosted-rooms area; E2E
with two real gateway stores and real install identities.
2026-08-30 20:39:58 -07:00
Teknium ff3835a630 fix(gateway): compact the live codex thread instead of no-op mirror rewrites (#73503)
On the codex_app_server runtime the model's real working context is the
app-server's server-side thread: CodexAppServerSession is constructed with
no history and each turn submits only the new user message
(agent/codex_runtime.py), so Hermes' transcript is a mirror that is never
replayed into a thread. Every out-of-turn compression call site (gateway
session hygiene, gateway /compress) built a DETACHED agent whose
_codex_session was None, so the codex route bailed at its "no active codex
thread" guard and returned the transcript unchanged ("compressed 150 ->
150 msgs") — and hygiene's finally-clause then evicted the cached live
agent, destroying the only real context: the next turn spawned an empty
thread while Hermes still mirrored a full history.

Fix, per the documented compression.codex_app_server_auto contract:

* Session hygiene now routes codex_app_server sessions to
  run_codex_hygiene_compaction(): in 'hermes' mode it compacts the LIVE
  cached agent's thread via thread/compact/start (through the existing
  codex route in _compress_context) and KEEPS that agent cached; 'native'
  and 'off' skip cleanly with no eviction and no local fallback. A wedged
  compaction records the persistent failure cooldown; success resets the
  hygiene failure streak.
* Gateway /compress detects the codex_app_server runtime before building
  a temporary compression agent and compacts the live thread with
  force=True instead (a manual compress is an explicit user decision in
  every mode). No live thread -> honest "nothing to compact" reply
  instead of a mirror rewrite plus eviction.
* No mode ever runs the local transcript compressor on this runtime:
  rewriting the mirror cannot shrink the thread, so the #73715-style
  local fallback (including its force=True leak into native/off) is
  deliberately not adopted.

Diagnosis of the mode-gate/no-thread deadlock builds on PR #73715.

Closes #73503

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-08-30 19:52:43 -07:00
Teknium 6101f52ba4 Merge remote-tracking branch 'origin/main' into core-tool-deferral 2026-08-30 19:47:30 -07:00
Teknium ed3562bbbc test(compression): drop moot digest-loop tests; match stamped backoff errors
PR #98628 removed _build_chunk_digests, so the two lean chunk-digest
cancellation tests reintroduced by the #97512 cherry-pick target a
deleted mechanism — removed. The #96775 stall-interrupt assertions now
match the stall_interrupted marker inside the strategy/kind-stamped
durable error instead of assuming it is the prefix.
2026-08-30 19:46:21 -07:00
Teknium ad925a08da fix(compression): scope worker-teardown grace to the total-ceiling path
The bounded-grace join only applies where the overlap hazard lives: a
total-ceiling expiry over a still-streaming worker (#97488). The
idle-stall path keeps its prompt detachment so the stall-fallback retry
preserves the #76354 S3 latency contract (silence never approaches 2x
the idle budget); its late unwind stays safe behind the fence poison
and attempt-generation supersession.
2026-08-30 19:46:21 -07:00
Teknium 74f9c5d7e6 test(compression): pin attempt-lifecycle contracts (#97488 #96775)
Sabotage-verified regression tests: bounded-grace worker teardown on
ceiling (cooperative join + uninterruptible orphan with retained
lease), durable strategy/kind-stamped backoff that survives a simulated
gateway restart against a real temp SessionDB, success clearing the
backoff, superseded-attempt late results discarded, and the
transient-block signal (type-pinned against MagicMock agents).
2026-08-30 19:46:21 -07:00
Teknium 19a59e9c93 fix(compression): transiently-blocked no-op is a soft defer, never exhaustion (#97488)
compress_context() now publishes agent._compression_blocked_transient
(reason string) when an automatic pass no-ops because a timed guard —
summary-failure cooldown or structural backoff — is active, with a
clear skip log line. The overflow-recovery and preflight loops in
conversation_loop treat that signal like the #69870 lock-skip: refund
the attempt and end the turn as compression_deferred instead of
counting the no-op toward compression_exhausted, which auto-resets
(wipes) the session at the gateway. Fixes the false auto-reset where a
real context_length_exceeded arrived while the host-timeout cooldown
was still active. The permanent 'ineffective' breaker intentionally
does not set the signal so genuinely incompressible sessions can still
exhaust.
2026-08-30 19:46:21 -07:00
Teknium a6549922b8 fix(compression): stamp durable backoff rows with strategy and failure kind (#96775 #97488)
record_timeout_failure() now persists
'backoff:<failure_kind>:strategy=<tail_mode>' into the state.db
cooldown row (sessions.compression_failure_cooldown_until +
compression_failure_error), so a failed/stalled/cancelled attempt's
identity survives gateway restarts and the rebuilt compressor makes the
same skip decision via bind_session_state()/get_active_compression_
failure_cooldown(refresh=True). Host callers pass ceiling_exhausted /
stalled; the stall-interrupt path passes stall_interrupted. A
successful compression still clears the row.
2026-08-30 19:46:21 -07:00
Teknium 892f756b8c fix(compression): tear down cancelled workers with bounded grace and discard superseded attempts (#97488)
A ceiling/idle-timeout host now joins its fence-cancelled worker for a
bounded grace before returning. A cooperative worker (which polls the
poison fence between provider phases) is reaped, proving quiescence, so
the durable lease releases normally. An uninterruptible worker is
orphaned behind the poison fence: its late result is discarded, and on
the total-ceiling path the holder-qualified lease stays retained until
it exits so no new attempt can overlap the unchanged session.

Supersession: a late candidate from an attempt whose compressor
generation was claimed by a newer attempt is discarded before the
commit boundary (failure_class=attempt_superseded), never committed
over newer state — covering fenceless callers the fence poison cannot
see.
2026-08-30 19:46:21 -07:00
HexLab98 f132f34ab8 test(compression): cover stall-interrupt vs early /stop backoff (#96775)
Pin both AuxiliaryExplicitCancellation and commit-fence cancellation, keep early /stop cooldown-neutral, merge with a longer live deadline, and prove force=/compress still bypasses the automatic brake.
2026-08-30 19:46:21 -07:00
HexLab98 027b339e79 fix(compression): persist stall-interrupted backoff on pre-commit cancel (#96775)
An explicit /stop after the summary stream has already gone idle restored the original transcript but left no durable cooldown, so the next automatic turn re-entered the same stalled strategy. Record a stall-specific failure on that path only, merge with any longer live deadline, and keep an ordinary early /stop cooldown-neutral.
2026-08-30 19:46:21 -07:00
fangliquanflq de4155b1bb fix(compression): report total ceiling expiry accurately 2026-08-30 19:46:21 -07:00
fangliquanflq c992c6de4f test(compression): isolate deadline timing assertions 2026-08-30 19:46:21 -07:00
fangliquanflq 0bcad4519e fix(compression): preserve fallback before worker start 2026-08-30 19:46:21 -07:00
fangliquanflq 14f40c0d6a test(compression): isolate commit overrun scheduling 2026-08-30 19:46:21 -07:00
fangliquanflq c0787c8ec1 fix(compression): release idle-timeout lease promptly 2026-08-30 19:46:21 -07:00
fangliquanflq 8d567ccd22 fix(compression): stop work at the total deadline 2026-08-30 19:46:21 -07:00
Teknium cc4b5ba1fc feat(bot-mode): replay pages carry authority lineage; pin byte-bounded replay
Follow-ups on top of the salvaged #97712 foundation:

- read_events() pages now include the room's authority stamp
  (authority.gateway_id + authority.epoch), so a replicating participant
  can persist lineage with every page and a future takeover layer can
  fence stale authorities from replayed state alone.
- New regression test proves the replay page bound counts UTF-8 BYTES,
  not characters: sabotaging LENGTH(CAST(.. AS BLOB)) back to
  LENGTH(TEXT) and dropping the .encode('utf-8') guard fails the test;
  the pre-existing multibyte test passed under that sabotage.
- _raise_room_not_found typed NoReturn so narrowing survives closures.
2026-08-30 19:46:14 -07:00
David Dudok de Wit cbc67b939f feat(bot-mode): add durable Group Chat authority and replay 2026-08-30 19:46:14 -07:00
Teknium 1f2bd9e763 fix(compression): stamp _DB_PERSISTED_MARKER after in-place batch compaction commit (#98450)
compress() returns marker-swept copies (_strip_persistence_markers, #57491);
the in-place branch committed them via archive_and_compact() but never
stamped the persistence marker, so the next _persist_session ->
_flush_messages_to_session_db_unlocked walk re-INSERTed the whole
post-compaction transcript (live set regrew ~58K -> ~512K tokens).

Centralize the post-commit contract in a shared helper,
stamp_db_persisted_markers(), used by all three archive_and_compact
callers: the in-place batch commit (previously missing), the
micro-compaction sync, and the proactive tool-result prune.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-08-30 19:46:11 -07:00
Teknium d63f996a75 feat(photon): read-receipt toggle, receipt-type alias, docs
Follow-ups on top of #98964's cherry-pick:
- PHOTON_READ_RECEIPTS env toggle (default true) so users can keep
  messages at Delivered; declared in plugin.yaml optional_env
- adapter drops both 'read' and 'read_receipt' content types (alias
  coverage from #91759 by @mooserini) + regression test
- docs: photon.md feature note + environment-variables.md row
2026-08-30 18:37:50 -07:00
Zihan Huang 9744fc0c99 feat(photon): support iMessage read receipts 2026-08-30 18:37:50 -07:00
Teknium 4f22543509 fix(compression): lean compaction makes exactly one auxiliary request per attempt
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.

- The detailed session log is folded into the single summary request: the
  lean prompt template gains a '## Detailed Session Log (oldest first)'
  section carrying the digest prompt's HARD RULES (identifiers verbatim,
  dense bullets, transcript-is-data). Output guidance grows by
  _LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
  summary budget — the old worst case (28 x 1,400 digest tokens) was spread
  across many requests and mostly re-covered tool noise; a single dense
  4K-token log inside one response preserves the load-bearing record while
  staying well inside one aux response (the summary call still sends no hard
  max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
  whole region (_sample_summary_input: 8 proportionally spaced slices,
  oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
  anchored to the newest end) instead of head+tail truncated, so session-log
  coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
  session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
  _LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
  _LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
  sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
  attempt_summary_route_kwargs — no remaining callers; the single-use
  summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
  section lands in the summary; oversized regions sampled with elision
  markers, never a second request; anchor index + recovery footer present).
  Sabotage-verified: restoring a second call_llm makes the call-count test
  fail. Docs and the compaction eval wording updated to stop claiming
  per-chunk calls.

Fixes #96603.
2026-08-30 09:03:57 -07:00
kshitijk4poor 5cc1369fa2 refactor(skills): clarify _find_skill docstring, narrow _local_root except
Simplify-code pass findings:
- The docstring claimed 'Matching bare-name suffix' but the fast path
  matches the exact directory name (parent.name) — reworded to say
  what the code actually does.
- _local_root() swallowed every Exception silently; narrowed to
  OSError (what resolve() raises) with a logger.debug breadcrumb so a
  recurring resolve failure is diagnosable instead of degrading every
  categorized lookup to a silent not-found.
2026-08-30 20:04:40 +05:30
kshitijk4poor 70220e529a refactor(skills): short-circuit bare-name match before resolve machinery
Review feedback (kokhlo): the categorized-name match ran
resolve().relative_to() for every SKILL.md in the walk even when the
bare-name branch already matched — 50+ resolve calls per invocation on
a bare-name lookup in a large profile.

Restructure so the bare directory-name check stays first and the
resolve/relative_to machinery only runs when the lookup name actually
contains a path separator. The skills root is resolved once, lazily,
only when a categorized lookup happens at all. Also compare the
relative path via as_posix() so 'category/skill' lookups work on
Windows, where str(Path) renders backslashes.
2026-08-30 20:04:40 +05:30
kshitijk4poor 70370e089c fix(skills): skill_view directory file_path + skill_manage categorized name resolution
Two agent-facing errors that recur constantly in optimization audit
logs (thousands of occurrences over five months):

1. skill_view(name, file_path='references') returned a raw
   '[Errno 21] Is a directory' OS error. The local-skill branch gated
   on target_file.exists(); a directory passes exists(), fell through
   to read_text(), and raised. The plugin-skill sibling branch already
   gated on is_file() — this aligns the local branch so a directory
   request gets the same helpful not-found payload with
   available_files listing instead of an OS error.

2. skill_manage rejected categorized names ('category/skill-name')
   with 'not found in active profile'. _find_skill matched only the
   bare directory name, while skill_view's own ambiguity hint tells
   the caller to use exactly the categorized form — every call that
   followed the hint failed. _find_skill now also matches the full
   relative path of the skill dir, giving skill_manage resolution
   parity with skill_view across edit/patch/delete/write_file/
   remove_file.

Both fixes are covered by regression tests that fail on main.
2026-08-30 20:04:40 +05:30
Teknium dce2ecb8a9 test: cron ContextVar-masking test now uses a chat platform
test_explicit_blank_masks_leaked_cron_env_for_gateway_classification
used platform=api_server as an arbitrary gateway platform; api_server
is now intentionally excluded from gateway approval contexts
(unattended class). Switch to telegram — the test's subject is the
blank-cron-ContextVar masking, not platform policy.
2026-08-30 07:04:16 -07:00
Teknium ef71f2cad8 fix(approval): widen webhook exclusion to all unattended platforms, deny by default
Builds on liuhao1024's webhook exclusion (#37317): instead of falling
through to auto-approve, unattended programmatic platforms (webhook,
msgraph_webhook, api_server) now resolve approval decisions instantly
via approvals.unattended_mode (default deny), mirroring cron_mode.

- _UNATTENDED_APPROVAL_PLATFORMS set + _is_unattended_platform_approval_context()
- approvals.unattended_mode config key (deny | approve), default deny
- Deny branches in _run_approval_gate, check_all_command_guards (with
  tirith parity), and check_execute_code_guard (#87509 sibling site)
- Docs: security.md; config_defaults.py comment + default

Fixes #37284. Also fixes the api_server half of #87509.
2026-08-30 07:04:16 -07:00
liuhao1024 73f8fb74e0 fix(approval): exclude webhook sessions from gateway approval context
Webhook sessions trigger the gateway approval branch because
HERMES_SESSION_PLATFORM is set, but the webhook adapter has no
send_exec_approval and no way to receive /approve replies.  This
blocks the session for the full approval timeout (60-300 s) with
no human who can resolve it.

Fix: _is_gateway_approval_context() now returns False when the
session platform is 'webhook', falling through to the non-interactive
path (auto-approve with warning, or deny if cron).

Regression tests added for webhook, non-webhook gateway, cron, and
no-platform scenarios.
2026-08-30 07:04:16 -07:00
Teknium 66666f6e2e fix(gateway): bust agent cache on remaining compaction-routing config keys
Follow-up to imsuperseller's #96740 (cherry-picked as the previous commit).
Widens _CACHE_BUSTING_CONFIG_KEYS with the other construction-baked
compaction-routing settings that had the same stale-cache shape:
compression.in_place, checkpoint_required, micro_compact,
micro_compact_every_n_turns, micro_compact_defrag_threshold_tokens.
Without these, a messaging-gateway session cached before a config edit
keeps the old compaction routing forever.

Not added (reported instead): abort_on_summary_failure, max_attempts,
protect_first_n, codex_gpt55_autoraise_notice, idle_compact_after_seconds
— behavior-tuning rather than routing, left for a deliberate pass.
2026-08-30 05:16:19 -07:00
imsuperseller 77f5de6272 fix(compression): hot-apply native compaction settings 2026-08-30 05:16:19 -07:00
Teknium 8a8aa850f1 fix(delegate): inherit endpoint-scoped capability map only on the parent's exact route
Same-class follow-up to #94036/#97292: a subagent spawned on the parent's
exact provider+base_url inherits the trusted-proxy capability map
(openai_native_compaction), so it keeps native compaction instead of
silently falling back to local summarization. Any provider- or
endpoint-changing delegation override stays DEFAULT-DENY, matching the
/model switch posture.
2026-08-30 05:16:10 -07:00
Teknium 1b2aaae1f8 chore: map steveonjava contributor email (PR #94036/#97292 salvage) 2026-08-30 05:16:10 -07:00
Stephen Chin f245765a6d fix(gateway): preserve capabilities on model switches
Carry normalized provider capabilities through /model results and session overrides so a live model switch does not wait for gateway rehydration.
2026-08-30 05:16:10 -07:00
Stephen Chin 80044bf385 fix(gateway): propagate trusted proxy capabilities
Forward normalized custom-provider capabilities on the default gateway path so native compaction does not depend on session rehydration. Document the content trust boundary and cover both lookup and gateway resolution.
2026-08-30 05:16:10 -07:00
Stephen Chin c9b9b5e6c7 fix(gateway): preserve native compaction capability on resume 2026-08-30 05:16:10 -07:00
Stephen Chin 48a4201f40 fix(compaction): preserve switch compatibility fixtures
Keep model-switch callers compatible with result objects created before runtime_capabilities was added, and do not roll back minimal agents that lack optional LM Studio helpers. Preserve rollback for real helper failures.
2026-08-30 05:16:10 -07:00
Stephen Chin 5247a6f07f fix(compaction): clarify runtime capability state
Use a distinct runtime_capabilities field on agents, preserve compatibility with earlier snapshots, and resolve the canonical direct OpenAI endpoint when a cross-provider switch omits base_url. Keep ambiguous proxy routes fail-closed.
2026-08-30 05:16:10 -07:00
Stephen Chin 903c36b6d4 fix(compaction): resolve capability from effective switch URL 2026-08-30 05:16:10 -07:00
Stephen Chin 08c7879ca1 fix(compaction): preserve native capability across runtime switches
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
2026-08-30 05:16:10 -07:00
Teknium 0af3d62e7b chore: map ijnotion@pm.me -> james47kjv (PR #98008 salvage) 2026-08-30 05:16:02 -07:00
james47kjv 80764b6d39 fix(codex): nudge the second continuation of a compaction-only turn
gpt-5.6 on the Codex backend answers a large turn with a server-side
`compaction` checkpoint and no message. The checkpoint rides the
`codex_reasoning_items` sidecar, so the interim assistant message looks
"replayable" and `interim_replayable` suppresses the continuation nudge.

But replayable is not the same as different. A checkpoint carries no
answer and no new instruction, and a replayed checkpoint makes
`prune_pre_checkpoint_items` drop every pre-checkpoint item. Measured on
a real 262-message session: the wire collapses from 489 items to 12 —
all 186 `function_call` / `function_call_output` pairs deleted — and
ends on an empty assistant turn. The model has nothing to answer, so it
returns another empty response; the next attempt sends the same bytes
(the provider's prefix cache reports 99-100% on the repeats) and returns
the same nothing. Three attempts later the turn dies with "Codex
response remained incomplete after 3 continuation attempts" and the
whole turn's work is lost.

Keep the first continuation bare — the model often just needs another
turn, and nudging immediately would cut multi-phase work short. Once
that bare retry has also come back incomplete, it is proven not to work
for this turn, so every remaining attempt carries the nudge.
2026-08-30 05:16:02 -07:00
Eric Maddox 52637fee63 test(native-compaction): cover image-only retention with an interleaved assistant message
Folded from PR #98345 (@ericmaddox): the one scenario its suite covered
that #91557's did not — an assistant message between the image-only user
message and the checkpoint, asserting post-prune ordering.
2026-08-30 05:15:54 -07:00
Andrex Ibiza, MBA 532b2d8874 fix(native-compaction): retain image-only user content
Preserve valid normalized input_image user messages across native-compaction checkpoints at bounded one-token retention cost. Keep text extraction text-only, reject malformed or unknown multipart placeholders, and prove the production adapter path without claiming unsupported input_file behavior.

Republish the identical source tree after an unrelated nondeterministic focus-redraw test failure; this commit contains no source delta from the previously verified object.

Refs #90976 and #91477.
2026-08-30 05:15:54 -07:00
Teknium be92703788 fix(compression): keep tool-schema tokens in the unanchored fallback estimate
The sibling-site widening replaced estimate_request_tokens_rough with estimate_messages_tokens_rough as the generic fallback feeding the route-aware wrapper, dropping the 20-30K token tool-schema envelope (#14695 class) and shifting the pinned mid-turn retry comparison. Restore the tools-inclusive figure as the fallback.
2026-08-30 05:15:46 -07:00
Teknium c49e2a496e fix(compression): route-aware pruned estimate at remaining pressure sibling sites
Follow-up to the mid-turn pre-API guard fix (#96995 / #97602): sweep the
remaining call sites that derive automatic compression pressure from a
generic estimate over the assembled durable history, which on a compacted
native-Codex session overstates the wire payload by orders of magnitude.

- agent/turn_context.py idle-triggered compaction: use
  _preflight_request_tokens (anchor -> native pruned -> generic) instead
  of the raw generic request estimate, so resuming a compacted codex
  session after an idle gap does not fire a compaction the next request
  never needed.
- agent/turn_context.py uncompressed-session overflow-warn RE-ARM: match
  the warn site's route-aware figure so the dedup re-arms correctly on
  native sessions.
- agent/conversation_loop.py post-response should_compress fallback
  (last_prompt_tokens==0, i.e. no provider usage after a disconnect or
  gateway restart — the unanchored case in #97602's repro): route through
  _midturn_request_pressure_tokens instead of the generic figure.

Left alone deliberately: provider-proven overflow recovery paths (413 /
context-length errors — the provider already proved the request does not
fit, figures there only arm recovery and score progress), compression
progress before/after pairs (relative deltas on the same scale), manual
/compress display estimates (gateway/CLI/ACP feedback, not automatic
triggers), MoA advisor budget trimming (not a codex-native wire payload),
and context_compressor internals (measure local durable-history shrink).
2026-08-30 05:15:46 -07:00
liuhao1024 222fda3d6b fix(compression): use the route-aware pruned estimate in the mid-turn pre-API guard
The #96155 fix (#96644) made the turn-prologue preflight estimate the
checkpoint-pruned native Responses payload, but the independent mid-turn
pre-API pressure guard in conversation_loop still estimated the full
assembled durable history. On a compacted native-Codex session the
generic figure overstates the wire by orders of magnitude (the issue's
deterministic probe: 1,037,241 generic vs 6,036 pruned, 171x), so the
guard false-tripped a 600-second local compression the main request
never needed — the live sequence shows the actual request then fit at
164k input tokens against a 765k threshold (#96995).

Extract the guard's pressure figure into _midturn_request_pressure_tokens
and mirror the turn-prologue: when native Responses compaction is proven
eligible, use estimate_native_responses_preflight_tokens (system prompt
and tools included, checkpoint-pruned); otherwise keep the generic
message+tools figure. Passing the assembled api_messages alongside
effective_system counts the system prompt exactly once — the estimator's
converter skips system-role rows and adds the prompt separately.
total_chars (verbose log proxy) and the non-codex paths are unchanged.

Fixes #96995
2026-08-30 05:15:46 -07:00