A 413 is a byte-size error, but the recovery loop scored compression
progress with estimate_messages_tokens_rough, which deliberately prices
every image at a flat per-image token cost (so screenshots don't trigger
premature compaction). When the payload is image-dominated that check can
never pass: in the reporting session two vision_analyze results were
5,627,202 bytes (96.6% of the request body) but ~3K of the ~80K token
estimate, so every attempt reported no_progress, the budget burned, and
the session wedged permanently with 'max compression attempts (3)
reached' at 13% context usage.
Post-#97160, the 413 path already routes into compaction and compaction's
historical-media aging genuinely frees the image bytes — but the
token-scored yardstick could not see the megabytes it freed. Add
serialized_messages_bytes() (exact serialized payload size, measured
identically before and after each pass — a measurement, not an estimate)
and score the 413 progress check with it. Tokens remain for status
display only; the context-overflow branch keeps its token yardstick,
because that error IS a token-budget error.
Images are never evicted from live history outside compaction (cache
invariant); the original strip-from-history mechanism in this PR was
superseded by #97160's compaction-time aging and is dropped in salvage.
Salvaged from #88960. Fixes#47339.
sanitize_api_messages (agent_runtime_helpers) and
_sanitize_tool_pairs (context_compressor) both collected
tool-call IDs and classified orphans with near-identical logic
that had already drifted: the canonical sanitizer added dedup
(#58350), but the compressor's copy did not.
Extract the shared orphan-detection logic into
_classify_tool_call_orphans(messages) in agent_runtime_helpers.
Both call sites now delegate to it, preserving their divergent
remediation strategies (insert-stubs vs strip-orphans) while
ensuring id-resolution rules and dedup stay in sync.
Closes#58357
Rebased onto current main to drop the empty-tool_calls fix (already on
main via #86654, cherry-picked from #77944 with @webtecnica's
authorship). This PR now carries only the two fixes unique to it:
1. A pre-existing api_content sidecar left stale on the consecutive-
assistant merge. The sidecar takes priority over content at
API-build time, so a merge could silently discard its own freshly
concatenated content on the next call. Only dropped when the merge
actually changes the resulting value (wz-heng, #78063 review) --
content_rewritten compares before/after value, not just whether an
assignment branch fired, so a falsy new_content (e.g. "") that
strips to nothing no longer trips a spurious sidecar drop.
2. sanitize_api_messages never flagged a tool result with a missing/
empty tool_call_id -- its orphan-detection set only ever collected
truthy ids, so an unpaired result with no id passed the final
chokepoint untouched.
Addresses teknium1's rebase request and wz-heng's review findings on
_sanitize_tool_pairs() in ContextCompressor compared raw tool_call_id
strings without stripping whitespace, the same bug fa3ab2ffd just fixed
in agent_runtime_helpers.py / run_agent.py (_get_tool_call_id_static +
sanitize_api_messages). ContextCompressor has its own near-identical
reimplementation of the pair-repair logic that was left unpatched.
When assistant-side and result-side IDs diverge only in surrounding
whitespace, the compressor misclassifies valid results as orphaned and
replaces them with [Result unavailable] stubs — silent data loss on
every compression cycle that touches such pairs.
Apply the same .strip() fix to all three sites:
- _get_tool_call_id (extracts IDs from assistant tool_calls)
- result_call_ids accumulation loop
- orphaned_results filter predicate
Closes the sibling gap of fa3ab2ffd / #42405.
Widen #90001's compaction-time strip to cover the gaps #89965 identified,
applied at compaction only per the cache ruling (request-time eviction
changes the per-call prefix and breaks prompt caching; compaction is the
one sanctioned cache break):
- Rule 1b: the opening attachment (anchor == 0) ages out once a newer
tool-result image supersedes it. The reported session opened with a
~200KB poster that previously survived every compaction. The row keeps
a non-empty text placeholder, so the zero-user-turn guard (#58753) and
role alternation are untouched.
- Native {_multimodal: True, content: [...]} dict envelopes now both
anchor (newest is kept) and strip (older collapse to their
text_summary via _strip_images_from_tool_msg, which also drops the
stale api_content sidecar per #97125's drop_stale_api_content).
- All three wire shapes (Chat Completions image_url, Responses
input_image, Anthropic-native image) were already matched by
_IMAGE_PART_TYPES; tests now pin each shape explicitly, plus
determinism (double-run is a no-op returning the same object).
Refs #89938, #89965
_strip_historical_media anchors on the newest image-bearing USER message and
returns the list untouched when that anchor is index 0 or does not exist. A
session whose images arrive from tools rather than attachments therefore has
nothing to be "before": twenty vision_analyze results keep multi-MB of base64
in every request body, the provider answers 413, and the 413 handler's
recovery compaction lands right back in this function and frees nothing. The
reporter saw seven compactions in thirteen minutes, all below 200K tokens.
Age tool-result images on their own timeline: keep the newest one, since that
is the image the model is reasoning about, and strip every older one wherever
it sits, including inside the protected tail. The tail exists to preserve
conversational continuity, not to pin bytes the model has already moved past.
User-message images keep today's treatment exactly. The user anchor is
checked first, so a tool result that is the newest of its kind but still sits
before that anchor is stripped as it always has been, and the anchor message
itself is still kept byte-for-byte - test_compressor_zero_user_guard depends
on that.
Refs #89938
_strip_images_from_messages() deleted any non-tool message whose content
became empty after image removal. An assistant message whose content was
entirely images but which carried tool_calls was therefore dropped,
orphaning its paired tool responses — providers reject the next request
with unmatched tool_call_id errors (HTTP 400). Replace such messages
with the plaintext placeholder instead, exactly like tool-role messages.
Adds a regression test covering the assistant + tool_calls +
image-only-content case.
Closes#40463
`api_content` is the byte-stability sidecar from #67274: it holds the exact
bytes previously sent for a message, and every turn substitutes it back into
`content` when building `api_messages`. `drop_stale_api_content` exists so a
content rewrite cannot be replayed from it — its own docstring states the
contract, and names the historical image strip as one of the callers:
Replaying the pre-rewrite sidecar would resend exactly what the rewrite
removed, so it must be dropped — the cost is one cache boundary miss,
never wrong content.
`_strip_images_from_messages` never drops it. The image-rejection recovery in
`conversation_loop` runs it over the persistent history, not just the per-call
copy:
agent._vision_supported = False
_imgs_removed = _strip_images_from_messages(messages) # history
if isinstance(api_messages, list):
_strip_images_from_messages(api_messages)
and `api_messages` are copies (`api_msg = msg.copy()`), so the history message
keeps its sidecar. The strip is therefore undone on the very next turn.
Reproduced with the real functions:
history content after strip : [{'type': 'text', 'text': 'look'}]
sidecar still present : True
NEXT TURN sends : 'look<IMAGE BYTES SENT LAST TURN>'
This is worse than a one-turn glitch, because the recovery cannot fire again:
it is gated on `getattr(agent, "_vision_supported", True)` and just set that
False. So on every subsequent turn the sidecar re-injects the images, the
text-only endpoint rejects them again, and the branch that would strip them is
disabled — the session stays wedged on a 4xx it already knew how to fix.
Drop the sidecar on each message the strip rewrites, inside the function so
every caller is covered. Messages with no images keep theirs, so only the
rewritten message pays a cache boundary — the tradeoff the invariant
prescribes. The two sibling recovery paths, `_sanitize_messages_surrogates`
and `_sanitize_messages_non_ascii`, are already safe: both walk every string
field on the message and so scrub the sidecar in passing. This one only
touches `content`.
tests/run_agent/test_image_rejection_fallback.py: new
TestStripImagesDropsStaleApiContent — the rewritten message loses its sidecar,
the next turn does not resend the stripped images, the tool-placeholder rewrite
is covered too, and untouched messages keep their sidecar. All four fail on
main. 53 passed across the image-rejection and api_content-sidecar suites; 307
passed across the sanitization/image/sidecar/replay agent tests (8 failures in
test_image_routing.py / test_save_url_image.py are pre-existing and fail
identically on clean main).
* refactor(video_generate): capability-gated dynamic schema — 6 optional args render only when the active provider/model honors them; fleet capability declarations + declaration<->implementation contract tests
* fix(video_gen): H3/Grok/Happy-Horse/Gemini audio is ALWAYS-ON native, not absent — new audio_native family key + audio_always_on capability surfaces as description line (maintainer catch)
* test(video_gen): duration-span test pins the active-model contract — resolved family's real window, short families not inflated, union fallback still spans 30s
MiniMax's Anthropic-compatible endpoint rejects an oversized native image
part with "media exceeds size limit: max 10485760 bytes (2013)" — no
occurrence of the word "image", so none of _IMAGE_TOO_LARGE_PATTERNS
matched. The 400 fell through to _REQUEST_VALIDATION_PATTERNS (the body
is type: invalid_request_error) and classified as format_error /
non-retryable.
That skipped the image-shrink recovery in conversation_loop, which is
gated on FailoverReason.image_too_large. Because the oversized part is
already baked into history as a tool_result image block, and the context
compressor rewrites text but not image data, every later turn re-sent the
same bytes and failed identically — the session stayed dead until the
user forked it.
Match on the "media" fragment, mirroring the existing "image exceeds"
entry so reworded vendor variants are caught too. A non-image media
rejection routed here is safe: the shrink pass finds no image parts,
returns False, and the caller surfaces the original error unchanged.
Fixes#76039
Treat the ChatGPT Codex invalid image-data 400 as an image rejection so Hermes strips image parts and retries text-only instead of aborting the session. Add coverage for the exact error wording.
Truncated or corrupt image bytes baked into immutable conversation history
get re-sent on every retry. Kimi/Moonshot reject them with HTTP 400
'prepare image failed ... failed to decode image: invalid or unsupported
image format', which was missing from _IMAGE_REJECTION_PHRASES, so the
turn exhausted retries and wedged the session instead of stripping the
images and recovering.
Adds the phrase to the recovery list plus a regression test mirroring the
exact Kimi error body. Complements PR #76896 (proactive full-decode
validation in vision_tools) with reactive recovery for already-poisoned
sessions. Fixes#76884.
The observed wire error — 'Downloaded response does not contain a valid JPG, PNG, WebP, or ICO image.' — has no match in _IMAGE_CORRUPT_PATTERNS, so it falls through to the non-retryable 400 handler and the session replays the same image parts into the same 400 until /new.
Adds the full observed sentence to the pattern list. Deliberately not the shorter prefixes: a bare 'downloaded response does not contain a valid' also matches non-image 400s, and a negative test now pins that a downloaded-response certificate 400 keeps falling through to the existing handler. Covered on both the 400 path and the message-only path.
Reported by ryuhaneul in #69078.
The image_corrupt recovery stripped images from canonical messages, not
just the retry payload — a transient provider rejection (xAI 'Invalid
PNG image.') permanently erased history, breaking the copy-on-write
contract (e762a5a473). Strip only the per-call api_messages copy (its
rows are shallow copies; the strip replaces content instead of mutating
the shared parts list) and pin history isolation with two regressions
(#69104 sweeper review).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Review (Sol xhigh) on the prior commit found a P1: the generic strip-
and-retry fallback ("any non-retryable 400 with image parts present
strips and retries") was too blunt. It couldn't tell an actual
image-corruption 400 apart from an unrelated one — bad tool schema,
unsupported parameter, billing, content policy — that merely happened
to carry image parts in the request. Any of those would silently erase
vision history and retry the still-invalid request, degrading sessions
that were never bricked in the first place. That's worse than the bug
it was meant to fix.
Revert the generic fallback (agent/conversation_loop.py). Keep only
the classifier-routed path: FailoverReason.image_corrupt +
_IMAGE_CORRUPT_PATTERNS, checked before _IMAGE_TOO_LARGE_PATTERNS
because shrinking corrupt bytes can't repair them. Corrupt-image
wordings still route to strip-and-retry; everything else falls through
to normal (non-retryable) handling as before. Add xAI's second wire
wording for the same corruption class ("base64 string of provided
image cannot be decoded", returned on unaligned truncation vs "Invalid
PNG image." on aligned truncation) and a compound-message test pinning
that image_corrupt wins when a body matches both pattern lists.
Drop TurnRetryState.stripped_images_this_turn. It's unnecessary now
that only one branch is left: the branch already only retries when
_strip_images_from_messages reports it removed something, and that
helper strips every image part from the request in one pass — so a
second corrupt-image hit on the retried (now text-only) request has
nothing left to strip and falls through on its own. No separate
one-shot flag needed.
Add a run_conversation integration test at the sequenced-provider
layer: corrupt 400 on attempt 1, strip, retry succeeds on attempt 2,
with explicit before/after assertions on the outgoing image_url part.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
The permanent-brick class in #69078: xAI returns 'Invalid PNG image'
when a re-serialized image part in replayed history becomes
undecodable. The existing image-error patterns cover only Anthropic
'exceeds max dimension' wordings and 'model does not support images'
strings, so the classifier lands on a generic non-retryable 400 and
neither the shrink path nor the strip path fires. Every subsequent
turn (even bare text) fails identically because the poison stays in
history — the session is permanently wedged until deleted.
Two recovery layers, deliberately separate:
- Semantic split: new FailoverReason.image_corrupt with
_IMAGE_CORRUPT_PATTERNS ('invalid png image' / 'invalid jpeg image'),
checked BEFORE _IMAGE_TOO_LARGE_PATTERNS in both _classify_400 and
_classify_by_message. Corrupt bytes route to strip-and-retry, never
to the shrink path (shrinking corrupt bytes cannot help).
- Generic fallback: any non-retryable 400 whose outgoing messages
still contain image parts gets one strip-and-retry via the existing
_strip_images_from_messages helper, guarded by a new
stripped_images_this_turn one-shot flag on TurnRetryState. This
un-bricks the session for any current or future provider wording
without adding another pattern list to maintain.
Item 3 from the report (multimodal-part integrity across FTS
persistence + compaction handoff) is a separate investigation and
remains follow-up work.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
Live on both providers (verified 2026-08-28 against openrouter.ai/api/v1/models
and inference-api.nousresearch.com/v1/models) but absent from both curated
picker lists. Adds the entry directly below qwen3.8-max per newest-first
family ordering, an explicit 1M DEFAULT_CONTEXT_LENGTHS entry (new family
slug would otherwise fall through to the generic qwen 131072 catch-all —
same class as #69881), and regenerates model-catalog.json.
Scoped rollout: only the named providers touched. Pricing snapshot skipped
(both routes bill via official_models_api live pricing). Reasoning floor
already fires via the qwen3 prefix entry (180s, verified).
_call_llm_impl applied auxiliary_max_tokens_param whenever
fast_compression_cap was non-None — but _compression_fast_lane_controls
passes an explicit caller max_tokens straight through, so a compression
call that set its own cap had the param force-injected onto providers
where _build_call_kwargs deliberately omits it (ZAI vision hard-400s on
max_tokens; GPT-5/Copilot require max_completion_tokens). Pre-fast-lane
main omitted the param for that exact call shape (verified via
subprocess pinned to the pre-PR base).
Gate the forced param on 'max_tokens is None' so it applies only to caps
the certified lane itself produced — the same guard the fallback path
already uses.
Regression test pins the pre-PR wire shape. Mutation-checked.
_run_protected_sync_provider_call propagates the forward-progress hook to
its daemon worker but not the new _aux_dispatch/_aux_provider_response
timing hooks (both threading.local). When compression takes the protected
path — the common case, since the summary call runs under
aux_interrupt_protection with a hard-cancel source — provider_dispatch_ms
and time_to_first_progress_ms were silently absent from telemetry.
Also collapse the two byte-identical save/restore context managers
(aux_progress_hook, _aux_timing_hook) onto one _aux_thread_local_hook
implementation so the propagation semantics can never drift between the
progress and timing slots.
Regression test drives _run_protected_sync_provider_call with both timing
hooks installed and asserts the worker-thread notifies reach them.
Mutation-checked (reverting the propagation fails the new test).
Post-review hardening on the attempt-ownership commit:
- Write-time re-validation: the entry staleness check in
_restore_compressor_attempt_state runs before the durable-cooldown DB
I/O, so a fallback could claim the compressor in that window and the
stale setattr loop would still clobber its state. The in-memory writes
now re-validate AND execute under _COMPRESSOR_ATTEMPT_LOCK — the same
lock claims are taken under. The DB rollback stays outside the lock
(safe: the dangerous direction requires a prior claim, which the entry
check rejects). Both the quality reviewer and the lead's independent
pre-verification converged on this window.
New deterministic test: TestMidRestoreClaimRace (claim injected
between entry check and write via instrumented deepcopy).
- Documented gen-0 semantics on _claim_compressor_attempt: per-compressor
all-or-nothing, never mixed with gen>0 on one instance (reviewer
finding 2, verified unreachable — comment hardens against future
confusion).
Dropped after verification (reviewer finding 3): resetting
_SUMMARY_ROUTE_CONSUMED on pin_summary_route exit — the echo lives in
the worker thread's COPIED context (propagate_context_to_thread) and
dies with it; a probe confirmed the next attempt's context is clean.
Resetting it would break digests running after the with-block.
Follow-up to #96634 (stall-fallback retry, #78981) addressing
donovan-yohan's post-merge adversarial review. The stall path detaches a
timed-out primary worker (fence cancel wins; future stays on the pool)
and immediately runs the fallback against the SAME ContextCompressor,
creating two verified races:
1. Late-primary snapshot restore: the detached primary's unwind called
_restore_compressor_attempt_state with the PRIMARY's pre-attempt
snapshot. Landing after the fallback's commit it rolled
_previous_summary/cooldown/provenance/telemetry back to pre-primary
values, silently discarding fallback-owned state.
2. Shared _compression_cancelled_check: the late primary's `finally`
cleared the callback the fallback had just installed, so the
fallback's F4 cancellation consult read None.
Fix: a monotonic per-compressor attempt generation claimed under one
module lock (_claim_compressor_attempt). Snapshot restores carry their
claiming generation and no-op when stale; the cancelled-check set/clear
moves into owner-stamped helpers (_install_compression_cancelled_check /
_clear_compression_cancelled_check_if_owner) so only the installing
attempt can clear it. The commit fence keeps owning COMMIT admission;
the generation owns compressor-ATTRIBUTE writes — two boundaries.
Legacy callers (attempt_generation=None) and slotted third-party
compressors (generation 0) keep the historical unconditional behavior.
Secondary review items:
- Lean chunk digests during a stall-fallback retry now follow the
summary onto the pinned healthy route: take_pinned_summary_route()
echoes the consumed route into a context-local
_SUMMARY_ROUTE_CONSUMED, and _build_chunk_digests passes
attempt_summary_route_kwargs() (non-consuming) to call_llm. The pin's
single-use contract for the SUMMARY call is unchanged — the
main-model retry still never re-issues the pinned route.
- Worker re-run repeating pre-compression callbacks: documented as an
accepted limitation on _retry_compression_on_fallback_chain
(built-ins idempotent; resuming mid-pipeline would couple the retry
to host callback ordering).
Tests (tests/agent/test_compression_attempt_ownership.py, 10 cases):
deterministic interleavings for both races (late-primary restore
no-ops + preserves fallback state; stale finally cannot clear the
fallback's callback), legacy/slotted compatibility, digest route
follow + context-locality of the consumed echo. Mutation-checked:
reverting only the two prod files to origin/main fails the suite;
restored stack green (21 passed incl. the original #78981 suite).
The one red in the wider sweep
(test_silence_cannot_approach_double_idle_timeout) is pre-existing on
clean origin/main — verified independently.
resolve_compression_fast_lane and _compression_config_claims_fast_lane
each hand-parsed the same four config fields (provider, model,
reasoning_effort, max_output_tokens) with copy-pasted normalization and
int-coercion. Extract _fast_lane_config_fields() as the single source of
truth for both.
This also fixes a real inconsistency the duplication hid: certification
checked the literal string 'none' while _get_task_extra_body routes
reasoning_effort through parse_reasoning_effort, which treats 'false',
'disabled', and YAML boolean false as disabled too. A user writing
reasoning_effort: false got reasoning disabled but silently lost the
fast-lane cap. Certification now delegates to parse_reasoning_effort so
the two predicates can never disagree.
Regression test: every disabled-spelling certifies; empty/real efforts
do not. Mutation-checked (reverting to the literal check fails the new
test).
_compression_fast_lane_controls unconditionally did body = dict(extra_body)
before the early-return guard. Move the copy below the guard so non-
compression auxiliary calls (vision, title_generation, etc.) return the
original extra_body reference without a shallow copy.
Follow-up to the #96217 salvage: the codex/xai/github route checks were
re-implemented inline at four sites (codex_responses_adapter helpers,
chat_completion_helpers kwargs build, _is_openai_codex_backend, the
run_agent silent-reject hint). Consolidate them into
classify_responses_route() / ResponsesRouteFlags in
codex_responses_adapter and migrate every site — backend-identity
predicate class (#22548/#70893/#59561/#72468).
Host checks use exact-host-or-subdomain semantics, never substring
matching.
Automatic preflight used the full durable transcript even when the
Codex Responses request would prune around a native compaction
checkpoint. That false-triggered a 600s local summary against history
the main request never sent. Estimate the converted, checkpoint-pruned
payload when native compaction is eligible, and keep the generic
estimate as the conservative fallback.
`_fast_model_from_catalog` treats the catalog's keys as a source of ids,
scanning them for a cheap model to use for side tasks like titling. Two
problems for Nous.
The credential lookup goes through `resolve_api_key_provider_credentials`,
which raises for Nous because it is OAuth. The read then went out
anonymous and came back with the full catalog rather than the one the org
may reach, so a policy-hidden model could be selected and then refused at
request time with `model_blocked_by_org_policy`.
Fall back to the Nous credential resolver when the api-key path raises,
and narrow the resulting ids by the org policy the same way the pickers'
lists are narrowed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`fetch_models_with_pricing` checked its cache above the point where the
Authorization header is built, and keyed that cache on the base URL alone.
Whichever read of a given base URL landed first in a process therefore
answered every later read, whatever key it passed — a non-empty result is
held for the life of the process.
That is wrong for any endpoint whose answer depends on who is asking. The
Nous inference gateway filters `GET /v1/models` by the caller's org model
policy, so an anonymous read landing first makes a later authenticated read
return the full, unfiltered catalog without a request going out.
Separate the URL root from the cache key and fold auth state into the
latter. Only whether a key was supplied participates, never its value, so
no secret reaches the key.
`credits_tracker` peeked into the private `_pricing_cache` and duplicated
the key shape to do it; it now calls `peek_cached_pricing`, which owns both
the /v1-suffix normalization and the preference for the authenticated
catalog.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
When context-compression rotation fires mid-turn, the current user
message was persisted twice into the child session. Root cause: dedup
used id()-seeded sets of copies instead of markers on the live objects.
Replace with _DB_PERSISTED_MARKER-based dedup as the sole authority:
- _ensure_compressed_has_user_turn returns CompressedUserTurnOutcome
- After publish_compression_child succeeds, stamp the live anchor-source
row (not a drifted index) with _DB_PERSISTED_MARKER
- _sync_persisted_markers mirrors stamps from result to live lists by
scoped identity (handles direct-path, adoption divergence, _session_messages)
- Remove _flushed_db_message_ids from rotation commit path (markers replace it)
- Unconditional (loud) imports — no silent fallback
Salvage of #94996 by @fedosis, rebased on top of #95433 (stall-fallback,
already merged). Both conversation_compression.py and run_agent.py are
built from origin/main + #94996's diff applied on top, preserving the
force_terminal refactor and _publish_new_fence from #95433.
Credit: @fedosis original PR #94996.
The timeout validation (isinstance(raw, (int, float)) and not
isinstance(raw, bool) and raw > 0 → float(raw)) was duplicated
between auxiliary_client.py:_fallback_entry_timeout and
conversation_compression.py:resolve_compression_fallback_route.
Extracted into _coerce_positive_timeout to prevent drift if the
config schema evolves.
Review follow-up for #95433. When the host's fence factory is absent or
raises, the retry runs on a private CompressionCommitFence() that
hard-interrupt admission never reads — /stop would serialize against the
aborted attempt's fence instead of the retry's commit boundary. Promote
the factory-failure swallow from debug to warning and warn when the retry
has no published fence, so the control-plane degradation is visible rather
than silent.
Also documents the pin's coverage (single _generate_summary call; the
lean-mode chunk digests are a separate unpinned call path).
A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.
The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
_translate_tool_result_to_gemini called _coerce_content_to_text unconditionally,
silently dropping image_url parts from multimodal tool results (e.g. vision_analyze
responses). Gemini 3.x supports a functionResponse.parts field for embedding
inlineData images directly inside the function response; Gemini 2.x does not.
Thread is_gemini3 through _build_gemini_contents → _translate_tool_result_to_gemini
and gate image embedding on _gemini_major_version >= 3. Reuses the existing
_extract_multimodal_parts helper (no duplicate code). Non-3.x path unchanged.
Original PR #32352 by @hbentel, salvaged onto current main.
Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
agent/deadline.py defined SuspectableBackend twice: the Phase 3a Protocol
(sync ensure_healthy(self) -> bool) and, further down the same module, an
unrelated concrete class with the same name (async
ensure_healthy(self, timeout=5.0)) added later by the MCP Phase 3b adopter.
Since Python executes class statements top-to-bottom, the second definition
silently shadowed the first at module scope.
Nothing in the tree imports or subclasses either by name today — the MCP
adopter duck-types the same-shaped contract directly on its own connection
class rather than referencing agent.deadline.SuspectableBackend — so this
caused no live behavior change. But it left the wrong (and differently
shaped) class resolvable under that name for the next Phase 3b adopter that
does import it for a type hint.
- move minimax/minimax-m3:free into the Free tier section (house
convention: :free SKUs group together, matching glm-5.2:free and the
nemotron :free entries) and regenerate model-catalog.json
- add Inkling family context length (1,048,576 — OpenRouter live
metadata, 2026-08-27) to DEFAULT_CONTEXT_LENGTHS; new family slug
otherwise fell through to no entry
- add Inkling to the reasoning stale-timeout floor table (300s tier,
same as Grok reasoning / Ox Alpha; OpenRouter marks the family as
reasoning-capable)
- widen the floor matcher's right-anchor separator class to include
':' so OpenRouter SKU suffixes (:free/:batch/:nitro) inherit the
family floor — inkling:free previously missed the inkling entry
- regression tests for the inkling floor + ':' separator
- Reuse the existing _commit_status variable for the terminal-edge gate
instead of the parallel _compaction_succeeded boolean (derived state).
- Give the commit_fence_cancelled abort the same force_terminal=True
terminal edge as the lock-contended abort, and reword the closure
comment that overstated the lock contender as 'the one exception'.
- Inline the codex app-server path's lifecycle closure: after gating on
success it reduced to a single success-site emit, so the scaffolding
(done-flag + closure + two no-op failure-path calls) was dead.