Test seams and plugin engines monkeypatch estimate_messages_tokens_rough with (messages)-only signatures; route callers only pass the charge_stale_thinking kwarg on the False path.
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).
Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.
Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
A delegate_task child that died (provider 404/400, timeout, crash)
previously vanished silently: the child's conversation loop returns
failed=True with the error summary in final_response, which the
classifier treated as usable output -> status 'completed'. And even
correctly-failed children only reached the parent MODEL — platforms
with tool_progress off (Telegram/Slack defaults) never showed the
human anything.
- delegate_tool: result.failed now forces status 'failed' (with the
error carried on the entry); new shared format_subagent_failure_line()
renders one clean human-readable line (traceback -> exception message,
length-capped); CLI tree + batch ✗ lines now include the reason.
- gateway TurnRunner.progress_callback: subagent.complete events with a
terminal failure status deliver that line via _deliver_platform_notice
BEFORE all progress-queue gates; tool_progress_callback is now always
attached (body gates each event class itself).
- tests: failed-flag classification regression + notice rendering suite.
- docs: Failure Visibility section in delegation docs.
Every participant gateway can now keep a durable copy of a hosted room's
ordered log and continue the room when its authority host is gone:
- gateway/hosted_room_replicas.py: replica store in root state.db.
ingest_page() persists authority-stamped groups.log pages idempotently,
refusing sequence gaps and authority-epoch regressions. promote_replica()
continues the room locally at epoch+1 with a lineage-proving
authority.claimed event; the stale owner is fenced everywhere the claim
replicates. demote_room() lets a returning stale authority fence itself
(authority.lost) upon observing a newer epoch, killing split-brain writes.
- tui_gateway/methods_groups.py: groups.replicate / groups.replica_state /
groups.promote / groups.demote RPC surface. Promotion requires
confirm=true — storage decides HOW takeover is atomic and provable, the
caller (user action now, lease/quorum driver later) decides WHEN it is
safe, matching the boundary blessed on #97681.
Validation: 20 new tests incl. a full failover round-trip (A hosts, B
replicates incrementally, A dies, B promotes with complete history, A
returns demoted and fenced); 69 total across the hosted-rooms area; E2E
with two real gateway stores and real install identities.
On the codex_app_server runtime the model's real working context is the
app-server's server-side thread: CodexAppServerSession is constructed with
no history and each turn submits only the new user message
(agent/codex_runtime.py), so Hermes' transcript is a mirror that is never
replayed into a thread. Every out-of-turn compression call site (gateway
session hygiene, gateway /compress) built a DETACHED agent whose
_codex_session was None, so the codex route bailed at its "no active codex
thread" guard and returned the transcript unchanged ("compressed 150 ->
150 msgs") — and hygiene's finally-clause then evicted the cached live
agent, destroying the only real context: the next turn spawned an empty
thread while Hermes still mirrored a full history.
Fix, per the documented compression.codex_app_server_auto contract:
* Session hygiene now routes codex_app_server sessions to
run_codex_hygiene_compaction(): in 'hermes' mode it compacts the LIVE
cached agent's thread via thread/compact/start (through the existing
codex route in _compress_context) and KEEPS that agent cached; 'native'
and 'off' skip cleanly with no eviction and no local fallback. A wedged
compaction records the persistent failure cooldown; success resets the
hygiene failure streak.
* Gateway /compress detects the codex_app_server runtime before building
a temporary compression agent and compacts the live thread with
force=True instead (a manual compress is an explicit user decision in
every mode). No live thread -> honest "nothing to compact" reply
instead of a mirror rewrite plus eviction.
* No mode ever runs the local transcript compressor on this runtime:
rewriting the mirror cannot shrink the thread, so the #73715-style
local fallback (including its force=True leak into native/off) is
deliberately not adopted.
Diagnosis of the mode-gate/no-thread deadlock builds on PR #73715.
Closes#73503
Co-authored-by: webtecnica <webtecnica@gmail.com>
PR #98628 removed _build_chunk_digests, so the two lean chunk-digest
cancellation tests reintroduced by the #97512 cherry-pick target a
deleted mechanism — removed. The #96775 stall-interrupt assertions now
match the stall_interrupted marker inside the strategy/kind-stamped
durable error instead of assuming it is the prefix.
The bounded-grace join only applies where the overlap hazard lives: a
total-ceiling expiry over a still-streaming worker (#97488). The
idle-stall path keeps its prompt detachment so the stall-fallback retry
preserves the #76354 S3 latency contract (silence never approaches 2x
the idle budget); its late unwind stays safe behind the fence poison
and attempt-generation supersession.
Sabotage-verified regression tests: bounded-grace worker teardown on
ceiling (cooperative join + uninterruptible orphan with retained
lease), durable strategy/kind-stamped backoff that survives a simulated
gateway restart against a real temp SessionDB, success clearing the
backoff, superseded-attempt late results discarded, and the
transient-block signal (type-pinned against MagicMock agents).
compress_context() now publishes agent._compression_blocked_transient
(reason string) when an automatic pass no-ops because a timed guard —
summary-failure cooldown or structural backoff — is active, with a
clear skip log line. The overflow-recovery and preflight loops in
conversation_loop treat that signal like the #69870 lock-skip: refund
the attempt and end the turn as compression_deferred instead of
counting the no-op toward compression_exhausted, which auto-resets
(wipes) the session at the gateway. Fixes the false auto-reset where a
real context_length_exceeded arrived while the host-timeout cooldown
was still active. The permanent 'ineffective' breaker intentionally
does not set the signal so genuinely incompressible sessions can still
exhaust.
record_timeout_failure() now persists
'backoff:<failure_kind>:strategy=<tail_mode>' into the state.db
cooldown row (sessions.compression_failure_cooldown_until +
compression_failure_error), so a failed/stalled/cancelled attempt's
identity survives gateway restarts and the rebuilt compressor makes the
same skip decision via bind_session_state()/get_active_compression_
failure_cooldown(refresh=True). Host callers pass ceiling_exhausted /
stalled; the stall-interrupt path passes stall_interrupted. A
successful compression still clears the row.
A ceiling/idle-timeout host now joins its fence-cancelled worker for a
bounded grace before returning. A cooperative worker (which polls the
poison fence between provider phases) is reaped, proving quiescence, so
the durable lease releases normally. An uninterruptible worker is
orphaned behind the poison fence: its late result is discarded, and on
the total-ceiling path the holder-qualified lease stays retained until
it exits so no new attempt can overlap the unchanged session.
Supersession: a late candidate from an attempt whose compressor
generation was claimed by a newer attempt is discarded before the
commit boundary (failure_class=attempt_superseded), never committed
over newer state — covering fenceless callers the fence poison cannot
see.
Pin both AuxiliaryExplicitCancellation and commit-fence cancellation, keep early /stop cooldown-neutral, merge with a longer live deadline, and prove force=/compress still bypasses the automatic brake.
An explicit /stop after the summary stream has already gone idle restored the original transcript but left no durable cooldown, so the next automatic turn re-entered the same stalled strategy. Record a stall-specific failure on that path only, merge with any longer live deadline, and keep an ordinary early /stop cooldown-neutral.
Follow-ups on top of the salvaged #97712 foundation:
- read_events() pages now include the room's authority stamp
(authority.gateway_id + authority.epoch), so a replicating participant
can persist lineage with every page and a future takeover layer can
fence stale authorities from replayed state alone.
- New regression test proves the replay page bound counts UTF-8 BYTES,
not characters: sabotaging LENGTH(CAST(.. AS BLOB)) back to
LENGTH(TEXT) and dropping the .encode('utf-8') guard fails the test;
the pre-existing multibyte test passed under that sabotage.
- _raise_room_not_found typed NoReturn so narrowing survives closures.
compress() returns marker-swept copies (_strip_persistence_markers, #57491);
the in-place branch committed them via archive_and_compact() but never
stamped the persistence marker, so the next _persist_session ->
_flush_messages_to_session_db_unlocked walk re-INSERTed the whole
post-compaction transcript (live set regrew ~58K -> ~512K tokens).
Centralize the post-commit contract in a shared helper,
stamp_db_persisted_markers(), used by all three archive_and_compact
callers: the in-place batch commit (previously missing), the
micro-compaction sync, and the proactive tool-result prune.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Follow-ups on top of #98964's cherry-pick:
- PHOTON_READ_RECEIPTS env toggle (default true) so users can keep
messages at Delivered; declared in plugin.yaml optional_env
- adapter drops both 'read' and 'read_receipt' content types (alias
coverage from #91759 by @mooserini) + regression test
- docs: photon.md feature note + environment-variables.md row
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.
- The detailed session log is folded into the single summary request: the
lean prompt template gains a '## Detailed Session Log (oldest first)'
section carrying the digest prompt's HARD RULES (identifiers verbatim,
dense bullets, transcript-is-data). Output guidance grows by
_LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
summary budget — the old worst case (28 x 1,400 digest tokens) was spread
across many requests and mostly re-covered tool noise; a single dense
4K-token log inside one response preserves the load-bearing record while
staying well inside one aux response (the summary call still sends no hard
max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
whole region (_sample_summary_input: 8 proportionally spaced slices,
oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
anchored to the newest end) instead of head+tail truncated, so session-log
coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
_LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
_LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
attempt_summary_route_kwargs — no remaining callers; the single-use
summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
section lands in the summary; oversized regions sampled with elision
markers, never a second request; anchor index + recovery footer present).
Sabotage-verified: restoring a second call_llm makes the call-count test
fail. Docs and the compaction eval wording updated to stop claiming
per-chunk calls.
Fixes#96603.
Simplify-code pass findings:
- The docstring claimed 'Matching bare-name suffix' but the fast path
matches the exact directory name (parent.name) — reworded to say
what the code actually does.
- _local_root() swallowed every Exception silently; narrowed to
OSError (what resolve() raises) with a logger.debug breadcrumb so a
recurring resolve failure is diagnosable instead of degrading every
categorized lookup to a silent not-found.
Review feedback (kokhlo): the categorized-name match ran
resolve().relative_to() for every SKILL.md in the walk even when the
bare-name branch already matched — 50+ resolve calls per invocation on
a bare-name lookup in a large profile.
Restructure so the bare directory-name check stays first and the
resolve/relative_to machinery only runs when the lookup name actually
contains a path separator. The skills root is resolved once, lazily,
only when a categorized lookup happens at all. Also compare the
relative path via as_posix() so 'category/skill' lookups work on
Windows, where str(Path) renders backslashes.
Two agent-facing errors that recur constantly in optimization audit
logs (thousands of occurrences over five months):
1. skill_view(name, file_path='references') returned a raw
'[Errno 21] Is a directory' OS error. The local-skill branch gated
on target_file.exists(); a directory passes exists(), fell through
to read_text(), and raised. The plugin-skill sibling branch already
gated on is_file() — this aligns the local branch so a directory
request gets the same helpful not-found payload with
available_files listing instead of an OS error.
2. skill_manage rejected categorized names ('category/skill-name')
with 'not found in active profile'. _find_skill matched only the
bare directory name, while skill_view's own ambiguity hint tells
the caller to use exactly the categorized form — every call that
followed the hint failed. _find_skill now also matches the full
relative path of the skill dir, giving skill_manage resolution
parity with skill_view across edit/patch/delete/write_file/
remove_file.
Both fixes are covered by regression tests that fail on main.
test_explicit_blank_masks_leaked_cron_env_for_gateway_classification
used platform=api_server as an arbitrary gateway platform; api_server
is now intentionally excluded from gateway approval contexts
(unattended class). Switch to telegram — the test's subject is the
blank-cron-ContextVar masking, not platform policy.
Webhook sessions trigger the gateway approval branch because
HERMES_SESSION_PLATFORM is set, but the webhook adapter has no
send_exec_approval and no way to receive /approve replies. This
blocks the session for the full approval timeout (60-300 s) with
no human who can resolve it.
Fix: _is_gateway_approval_context() now returns False when the
session platform is 'webhook', falling through to the non-interactive
path (auto-approve with warning, or deny if cron).
Regression tests added for webhook, non-webhook gateway, cron, and
no-platform scenarios.
Follow-up to imsuperseller's #96740 (cherry-picked as the previous commit).
Widens _CACHE_BUSTING_CONFIG_KEYS with the other construction-baked
compaction-routing settings that had the same stale-cache shape:
compression.in_place, checkpoint_required, micro_compact,
micro_compact_every_n_turns, micro_compact_defrag_threshold_tokens.
Without these, a messaging-gateway session cached before a config edit
keeps the old compaction routing forever.
Not added (reported instead): abort_on_summary_failure, max_attempts,
protect_first_n, codex_gpt55_autoraise_notice, idle_compact_after_seconds
— behavior-tuning rather than routing, left for a deliberate pass.
Same-class follow-up to #94036/#97292: a subagent spawned on the parent's
exact provider+base_url inherits the trusted-proxy capability map
(openai_native_compaction), so it keeps native compaction instead of
silently falling back to local summarization. Any provider- or
endpoint-changing delegation override stays DEFAULT-DENY, matching the
/model switch posture.
Forward normalized custom-provider capabilities on the default gateway path so native compaction does not depend on session rehydration. Document the content trust boundary and cover both lookup and gateway resolution.
Keep model-switch callers compatible with result objects created before runtime_capabilities was added, and do not roll back minimal agents that lack optional LM Studio helpers. Preserve rollback for real helper failures.
Use a distinct runtime_capabilities field on agents, preserve compatibility with earlier snapshots, and resolve the canonical direct OpenAI endpoint when a cross-provider switch omits base_url. Keep ambiguous proxy routes fail-closed.
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
gpt-5.6 on the Codex backend answers a large turn with a server-side
`compaction` checkpoint and no message. The checkpoint rides the
`codex_reasoning_items` sidecar, so the interim assistant message looks
"replayable" and `interim_replayable` suppresses the continuation nudge.
But replayable is not the same as different. A checkpoint carries no
answer and no new instruction, and a replayed checkpoint makes
`prune_pre_checkpoint_items` drop every pre-checkpoint item. Measured on
a real 262-message session: the wire collapses from 489 items to 12 —
all 186 `function_call` / `function_call_output` pairs deleted — and
ends on an empty assistant turn. The model has nothing to answer, so it
returns another empty response; the next attempt sends the same bytes
(the provider's prefix cache reports 99-100% on the repeats) and returns
the same nothing. Three attempts later the turn dies with "Codex
response remained incomplete after 3 continuation attempts" and the
whole turn's work is lost.
Keep the first continuation bare — the model often just needs another
turn, and nudging immediately would cut multi-phase work short. Once
that bare retry has also come back incomplete, it is proven not to work
for this turn, so every remaining attempt carries the nudge.
Folded from PR #98345 (@ericmaddox): the one scenario its suite covered
that #91557's did not — an assistant message between the image-only user
message and the checkpoint, asserting post-prune ordering.
Preserve valid normalized input_image user messages across native-compaction checkpoints at bounded one-token retention cost. Keep text extraction text-only, reject malformed or unknown multipart placeholders, and prove the production adapter path without claiming unsupported input_file behavior.
Republish the identical source tree after an unrelated nondeterministic focus-redraw test failure; this commit contains no source delta from the previously verified object.
Refs #90976 and #91477.
The sibling-site widening replaced estimate_request_tokens_rough with estimate_messages_tokens_rough as the generic fallback feeding the route-aware wrapper, dropping the 20-30K token tool-schema envelope (#14695 class) and shifting the pinned mid-turn retry comparison. Restore the tools-inclusive figure as the fallback.
Follow-up to the mid-turn pre-API guard fix (#96995 / #97602): sweep the
remaining call sites that derive automatic compression pressure from a
generic estimate over the assembled durable history, which on a compacted
native-Codex session overstates the wire payload by orders of magnitude.
- agent/turn_context.py idle-triggered compaction: use
_preflight_request_tokens (anchor -> native pruned -> generic) instead
of the raw generic request estimate, so resuming a compacted codex
session after an idle gap does not fire a compaction the next request
never needed.
- agent/turn_context.py uncompressed-session overflow-warn RE-ARM: match
the warn site's route-aware figure so the dedup re-arms correctly on
native sessions.
- agent/conversation_loop.py post-response should_compress fallback
(last_prompt_tokens==0, i.e. no provider usage after a disconnect or
gateway restart — the unanchored case in #97602's repro): route through
_midturn_request_pressure_tokens instead of the generic figure.
Left alone deliberately: provider-proven overflow recovery paths (413 /
context-length errors — the provider already proved the request does not
fit, figures there only arm recovery and score progress), compression
progress before/after pairs (relative deltas on the same scale), manual
/compress display estimates (gateway/CLI/ACP feedback, not automatic
triggers), MoA advisor budget trimming (not a codex-native wire payload),
and context_compressor internals (measure local durable-history shrink).
The #96155 fix (#96644) made the turn-prologue preflight estimate the
checkpoint-pruned native Responses payload, but the independent mid-turn
pre-API pressure guard in conversation_loop still estimated the full
assembled durable history. On a compacted native-Codex session the
generic figure overstates the wire by orders of magnitude (the issue's
deterministic probe: 1,037,241 generic vs 6,036 pruned, 171x), so the
guard false-tripped a 600-second local compression the main request
never needed — the live sequence shows the actual request then fit at
164k input tokens against a 765k threshold (#96995).
Extract the guard's pressure figure into _midturn_request_pressure_tokens
and mirror the turn-prologue: when native Responses compaction is proven
eligible, use estimate_native_responses_preflight_tokens (system prompt
and tools included, checkpoint-pruned); otherwise keep the generic
message+tools figure. Passing the assembled api_messages alongside
effective_system counts the system prompt exactly once — the estimator's
converter skips system-role rows and adds the prompt separately.
total_chars (verbose log proxy) and the non-codex paths are unchanged.
Fixes#96995