Commit Graph

3725 Commits

Author SHA1 Message Date
Brian b855f86bc8 fix(agent): 413 recovery measures bytes, not token estimates
A 413 is a byte-size error, but the recovery loop scored compression
progress with estimate_messages_tokens_rough, which deliberately prices
every image at a flat per-image token cost (so screenshots don't trigger
premature compaction). When the payload is image-dominated that check can
never pass: in the reporting session two vision_analyze results were
5,627,202 bytes (96.6% of the request body) but ~3K of the ~80K token
estimate, so every attempt reported no_progress, the budget burned, and
the session wedged permanently with 'max compression attempts (3)
reached' at 13% context usage.

Post-#97160, the 413 path already routes into compaction and compaction's
historical-media aging genuinely frees the image bytes — but the
token-scored yardstick could not see the megabytes it freed. Add
serialized_messages_bytes() (exact serialized payload size, measured
identically before and after each pass — a measurement, not an estimate)
and score the 413 progress check with it. Tokens remain for status
display only; the context-overflow branch keeps its token yardstick,
because that error IS a token-budget error.

Images are never evicted from live history outside compaction (cache
invariant); the original strip-from-history mechanism in this PR was
superseded by #97160's compaction-time aging and is dropped in salvage.

Salvaged from #88960. Fixes #47339.
2026-08-28 07:51:16 -07:00
Teknium eff97a8a05 refactor(profiles): retire the cross-profile write guard — profiles are not isolated (maintainer decision); mirror lost-write guards (#32049) survive; patch/write_file schemas drop cross_profile (-83 tok/call) (#97165) 2026-08-28 06:36:22 -07:00
Hermes Agent 225fa13bd3 fix(sanitizer): drop duplicated legacy _classify_tool_call_orphans left by cherry-pick auto-merge 2026-08-28 06:32:48 -07:00
isheng c6a426e9ad refactor(sanitizer): extract shared _classify_tool_call_orphans to eliminate drift
sanitize_api_messages (agent_runtime_helpers) and
_sanitize_tool_pairs (context_compressor) both collected
tool-call IDs and classified orphans with near-identical logic
that had already drifted: the canonical sanitizer added dedup
(#58350), but the compressor's copy did not.

Extract the shared orphan-detection logic into
_classify_tool_call_orphans(messages) in agent_runtime_helpers.
Both call sites now delegate to it, preserving their divergent
remediation strategies (insert-stubs vs strip-orphans) while
ensuring id-resolution rules and dedup stay in sync.

Closes #58357
2026-08-28 06:32:48 -07:00
joaomarcos f0ac2c8f12 fix(agent): drop stale api_content sidecar and unpaired tool results
Rebased onto current main to drop the empty-tool_calls fix (already on
main via #86654, cherry-picked from #77944 with @webtecnica's
authorship). This PR now carries only the two fixes unique to it:

1. A pre-existing api_content sidecar left stale on the consecutive-
   assistant merge. The sidecar takes priority over content at
   API-build time, so a merge could silently discard its own freshly
   concatenated content on the next call. Only dropped when the merge
   actually changes the resulting value (wz-heng, #78063 review) --
   content_rewritten compares before/after value, not just whether an
   assignment branch fired, so a falsy new_content (e.g. "") that
   strips to nothing no longer trips a spurious sidecar drop.

2. sanitize_api_messages never flagged a tool result with a missing/
   empty tool_call_id -- its orphan-detection set only ever collected
   truthy ids, so an unpaired result with no id passed the final
   chokepoint untouched.

Addresses teknium1's rebase request and wz-heng's review findings on
2026-08-28 06:32:48 -07:00
srojk34 e024bf75ce fix(compression): strip whitespace from tool_call_id in _sanitize_tool_pairs
_sanitize_tool_pairs() in ContextCompressor compared raw tool_call_id
strings without stripping whitespace, the same bug fa3ab2ffd just fixed
in agent_runtime_helpers.py / run_agent.py (_get_tool_call_id_static +
sanitize_api_messages). ContextCompressor has its own near-identical
reimplementation of the pair-repair logic that was left unpatched.

When assistant-side and result-side IDs diverge only in surrounding
whitespace, the compressor misclassifies valid results as orphaned and
replaces them with [Result unavailable] stubs — silent data loss on
every compression cycle that touches such pairs.

Apply the same .strip() fix to all three sites:
- _get_tool_call_id (extracts IDs from assistant tool_calls)
- result_call_ids accumulation loop
- orphaned_results filter predicate

Closes the sibling gap of fa3ab2ffd / #42405.
2026-08-28 06:32:48 -07:00
Teknium 0dce46feb7 fix(compressor): widen compaction-time image aging to first-message and envelope shapes
Widen #90001's compaction-time strip to cover the gaps #89965 identified,
applied at compaction only per the cache ruling (request-time eviction
changes the per-call prefix and breaks prompt caching; compaction is the
one sanctioned cache break):

- Rule 1b: the opening attachment (anchor == 0) ages out once a newer
  tool-result image supersedes it. The reported session opened with a
  ~200KB poster that previously survived every compaction. The row keeps
  a non-empty text placeholder, so the zero-user-turn guard (#58753) and
  role alternation are untouched.
- Native {_multimodal: True, content: [...]} dict envelopes now both
  anchor (newest is kept) and strip (older collapse to their
  text_summary via _strip_images_from_tool_msg, which also drops the
  stale api_content sidecar per #97125's drop_stale_api_content).
- All three wire shapes (Chat Completions image_url, Responses
  input_image, Anthropic-native image) were already matched by
  _IMAGE_PART_TYPES; tests now pin each shape explicitly, plus
  determinism (double-run is a no-op returning the same object).

Refs #89938, #89965
2026-08-28 06:32:43 -07:00
Jack Lau b81b599d50 fix(compressor): age out stale tool-result images during compaction
_strip_historical_media anchors on the newest image-bearing USER message and
returns the list untouched when that anchor is index 0 or does not exist. A
session whose images arrive from tools rather than attachments therefore has
nothing to be "before": twenty vision_analyze results keep multi-MB of base64
in every request body, the provider answers 413, and the 413 handler's
recovery compaction lands right back in this function and frees nothing. The
reporter saw seven compactions in thirteen minutes, all below 200K tokens.

Age tool-result images on their own timeline: keep the newest one, since that
is the image the model is reasoning about, and strip every older one wherever
it sits, including inside the protected tail. The tail exists to preserve
conversational continuity, not to pin bytes the model has already moved past.

User-message images keep today's treatment exactly. The user anchor is
checked first, so a tool result that is the newest of its kind but still sits
before that anchor is stripped as it always has been, and the anchor message
itself is still kept byte-for-byte - test_compressor_zero_user_guard depends
on that.

Refs #89938
2026-08-28 06:32:43 -07:00
Teknium a641644f12 fix(paths): display_hermes_home renders POSIX separators on Windows — kills ~/AppData\Local\hermes chimeras in tool schemas and user-facing messages (#97137) 2026-08-28 05:45:00 -07:00
Hakan Baysal cb8027afed fix(sanitize): preserve assistant messages with tool_calls when stripping images
_strip_images_from_messages() deleted any non-tool message whose content
became empty after image removal. An assistant message whose content was
entirely images but which carried tool_calls was therefore dropped,
orphaning its paired tool responses — providers reject the next request
with unmatched tool_call_id errors (HTTP 400). Replace such messages
with the plaintext placeholder instead, exactly like tool-role messages.

Adds a regression test covering the assistant + tool_calls +
image-only-content case.

Closes #40463
2026-08-28 05:17:26 -07:00
Frowtek e1762bd30b fix(agent): drop the api_content sidecar when stripping images from history
`api_content` is the byte-stability sidecar from #67274: it holds the exact
bytes previously sent for a message, and every turn substitutes it back into
`content` when building `api_messages`. `drop_stale_api_content` exists so a
content rewrite cannot be replayed from it — its own docstring states the
contract, and names the historical image strip as one of the callers:

    Replaying the pre-rewrite sidecar would resend exactly what the rewrite
    removed, so it must be dropped — the cost is one cache boundary miss,
    never wrong content.

`_strip_images_from_messages` never drops it. The image-rejection recovery in
`conversation_loop` runs it over the persistent history, not just the per-call
copy:

    agent._vision_supported = False
    _imgs_removed = _strip_images_from_messages(messages)      # history
    if isinstance(api_messages, list):
        _strip_images_from_messages(api_messages)

and `api_messages` are copies (`api_msg = msg.copy()`), so the history message
keeps its sidecar. The strip is therefore undone on the very next turn.

Reproduced with the real functions:

    history content after strip : [{'type': 'text', 'text': 'look'}]
    sidecar still present       : True
    NEXT TURN sends             : 'look<IMAGE BYTES SENT LAST TURN>'

This is worse than a one-turn glitch, because the recovery cannot fire again:
it is gated on `getattr(agent, "_vision_supported", True)` and just set that
False. So on every subsequent turn the sidecar re-injects the images, the
text-only endpoint rejects them again, and the branch that would strip them is
disabled — the session stays wedged on a 4xx it already knew how to fix.

Drop the sidecar on each message the strip rewrites, inside the function so
every caller is covered. Messages with no images keep theirs, so only the
rewritten message pays a cache boundary — the tradeoff the invariant
prescribes. The two sibling recovery paths, `_sanitize_messages_surrogates`
and `_sanitize_messages_non_ascii`, are already safe: both walk every string
field on the message and so scrub the sidecar in passing. This one only
touches `content`.

tests/run_agent/test_image_rejection_fallback.py: new
TestStripImagesDropsStaleApiContent — the rewritten message loses its sidecar,
the next turn does not resend the stripped images, the tool-placeholder rewrite
is covered too, and untouched messages keep their sidecar. All four fail on
main. 53 passed across the image-rejection and api_content-sidecar suites; 307
passed across the sanitization/image/sidecar/replay agent tests (8 failures in
test_image_routing.py / test_save_url_image.py are pre-existing and fail
identically on clean main).
2026-08-28 05:17:21 -07:00
Teknium 536adb35c4 refactor(video_generate): capability-gated dynamic schema (~814 → 458/377 tok/call) (#97095)
* refactor(video_generate): capability-gated dynamic schema — 6 optional args render only when the active provider/model honors them; fleet capability declarations + declaration<->implementation contract tests

* fix(video_gen): H3/Grok/Happy-Horse/Gemini audio is ALWAYS-ON native, not absent — new audio_native family key + audio_always_on capability surfaces as description line (maintainer catch)

* test(video_gen): duration-span test pins the active-model contract — resolved family's real window, short families not inflated, union fallback still spans 30s
2026-08-28 05:10:33 -07:00
lEWFkRAD b80b9d8271 fix(codex): preserve assistant image slots in replay 2026-08-28 04:58:14 -07:00
lEWFkRAD 8de45940fb fix(codex): drop assistant images from Responses replay
Fixes #96816
2026-08-28 04:58:14 -07:00
Koduri Mahesh Bhushan Chowdary b3f4f50771 fix(agent): classify "media exceeds size limit" as image_too_large
MiniMax's Anthropic-compatible endpoint rejects an oversized native image
part with "media exceeds size limit: max 10485760 bytes (2013)" — no
occurrence of the word "image", so none of _IMAGE_TOO_LARGE_PATTERNS
matched. The 400 fell through to _REQUEST_VALIDATION_PATTERNS (the body
is type: invalid_request_error) and classified as format_error /
non-retryable.

That skipped the image-shrink recovery in conversation_loop, which is
gated on FailoverReason.image_too_large. Because the oversized part is
already baked into history as a tool_result image block, and the context
compressor rewrites text but not image data, every later turn re-sent the
same bytes and failed identically — the session stayed dead until the
user forked it.

Match on the "media" fragment, mirroring the existing "image exceeds"
entry so reworded vendor variants are caught too. A non-image media
rejection routed here is safe: the shrink pass finds no image parts,
returns False, and the caller surfaces the original error unchanged.

Fixes #76039
2026-08-28 04:58:06 -07:00
yoma 98a84783c7 fix(vision): recover from generic image content rejection 2026-08-28 04:57:58 -07:00
Teknium c30ac90a92 feat(compaction): rebuild dynamic tool schemas at the compaction commit boundary — forever-sessions finally pick up config changes (#97073) 2026-08-28 04:01:05 -07:00
Al Cooke 0241619068 fix: retry text-only on Codex invalid image data errors
Treat the ChatGPT Codex invalid image-data 400 as an image rejection so Hermes strips image parts and retries text-only instead of aborting the session. Add coverage for the exact error wording.
2026-08-28 03:46:24 -07:00
fkdls112 cd72689e03 fix(agent): strip images on Kimi/Moonshot 'failed to decode image' 400
Truncated or corrupt image bytes baked into immutable conversation history
get re-sent on every retry. Kimi/Moonshot reject them with HTTP 400
'prepare image failed ... failed to decode image: invalid or unsupported
image format', which was missing from _IMAGE_REJECTION_PHRASES, so the
turn exhausted retries and wedged the session instead of stripping the
images and recovering.

Adds the phrase to the recovery list plus a regression test mirroring the
exact Kimi error body. Complements PR #76896 (proactive full-decode
validation in vision_tools) with reactive recovery for already-poisoned
sessions. Fixes #76884.
2026-08-28 03:46:24 -07:00
Sora-bluesky 564113572e fix(agent): classify xAI's downloaded-response wording as a corrupt image
The observed wire error — 'Downloaded response does not contain a valid JPG, PNG, WebP, or ICO image.' — has no match in _IMAGE_CORRUPT_PATTERNS, so it falls through to the non-retryable 400 handler and the session replays the same image parts into the same 400 until /new.

Adds the full observed sentence to the pattern list. Deliberately not the shorter prefixes: a bare 'downloaded response does not contain a valid' also matches non-image 400s, and a negative test now pins that a downloaded-response certificate 400 keeps falling through to the existing handler. Covered on both the 400 path and the message-only path.

Reported by ryuhaneul in #69078.
2026-08-28 03:46:24 -07:00
Sora-bluesky a3177d0570 fix(agent): keep canonical history intact during image-corrupt retry
The image_corrupt recovery stripped images from canonical messages, not
just the retry payload — a transient provider rejection (xAI 'Invalid
PNG image.') permanently erased history, breaking the copy-on-write
contract (e762a5a473). Strip only the per-call api_messages copy (its
rows are shallow copies; the strip replaces content instead of mutating
the shared parts list) and pin history isolation with two regressions
(#69104 sweeper review).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-28 03:46:24 -07:00
Sora-bluesky d61411d131 fix(agent): narrow #69078 image-corrupt recovery to the classifier route
Review (Sol xhigh) on the prior commit found a P1: the generic strip-
and-retry fallback ("any non-retryable 400 with image parts present
strips and retries") was too blunt. It couldn't tell an actual
image-corruption 400 apart from an unrelated one — bad tool schema,
unsupported parameter, billing, content policy — that merely happened
to carry image parts in the request. Any of those would silently erase
vision history and retry the still-invalid request, degrading sessions
that were never bricked in the first place. That's worse than the bug
it was meant to fix.

Revert the generic fallback (agent/conversation_loop.py). Keep only
the classifier-routed path: FailoverReason.image_corrupt +
_IMAGE_CORRUPT_PATTERNS, checked before _IMAGE_TOO_LARGE_PATTERNS
because shrinking corrupt bytes can't repair them. Corrupt-image
wordings still route to strip-and-retry; everything else falls through
to normal (non-retryable) handling as before. Add xAI's second wire
wording for the same corruption class ("base64 string of provided
image cannot be decoded", returned on unaligned truncation vs "Invalid
PNG image." on aligned truncation) and a compound-message test pinning
that image_corrupt wins when a body matches both pattern lists.

Drop TurnRetryState.stripped_images_this_turn. It's unnecessary now
that only one branch is left: the branch already only retries when
_strip_images_from_messages reports it removed something, and that
helper strips every image part from the request in one pass — so a
second corrupt-image hit on the retried (now text-only) request has
nothing left to strip and falls through on its own. No separate
one-shot flag needed.

Add a run_conversation integration test at the sequenced-provider
layer: corrupt 400 on attempt 1, strip, retry succeeds on attempt 2,
with explicit before/after assertions on the outgoing image_url part.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
2026-08-28 03:46:24 -07:00
Sora-bluesky 8aeb3f6ee3 fix(agent): un-brick sessions on non-retryable 400s that carry image parts
The permanent-brick class in #69078: xAI returns 'Invalid PNG image'
when a re-serialized image part in replayed history becomes
undecodable. The existing image-error patterns cover only Anthropic
'exceeds max dimension' wordings and 'model does not support images'
strings, so the classifier lands on a generic non-retryable 400 and
neither the shrink path nor the strip path fires. Every subsequent
turn (even bare text) fails identically because the poison stays in
history — the session is permanently wedged until deleted.

Two recovery layers, deliberately separate:

- Semantic split: new FailoverReason.image_corrupt with
  _IMAGE_CORRUPT_PATTERNS ('invalid png image' / 'invalid jpeg image'),
  checked BEFORE _IMAGE_TOO_LARGE_PATTERNS in both _classify_400 and
  _classify_by_message. Corrupt bytes route to strip-and-retry, never
  to the shrink path (shrinking corrupt bytes cannot help).
- Generic fallback: any non-retryable 400 whose outgoing messages
  still contain image parts gets one strip-and-retry via the existing
  _strip_images_from_messages helper, guarded by a new
  stripped_images_this_turn one-shot flag on TurnRetryState. This
  un-bricks the session for any current or future provider wording
  without adding another pattern list to maintain.

Item 3 from the report (multimodal-part integrity across FTS
persistence + compaction handoff) is a separate investigation and
remains follow-up work.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
2026-08-28 03:46:24 -07:00
Teknium 48d2528066 feat(models): qwen3.8-flash now selectable on OpenRouter and Nous portal
Live on both providers (verified 2026-08-28 against openrouter.ai/api/v1/models
and inference-api.nousresearch.com/v1/models) but absent from both curated
picker lists. Adds the entry directly below qwen3.8-max per newest-first
family ordering, an explicit 1M DEFAULT_CONTEXT_LENGTHS entry (new family
slug would otherwise fall through to the generic qwen 131072 catch-all —
same class as #69881), and regenerates model-catalog.json.

Scoped rollout: only the named providers touched. Pricing snapshot skipped
(both routes bill via official_models_api live pricing). Reasoning floor
already fires via the qwen3 prefix entry (180s, verified).
2026-08-28 01:09:58 -07:00
kshitij 3f315e46fe Merge pull request #96963 from kshitijk4poor/refactor/fast-lane-consolidation
fix(compression): fast-lane follow-up — certification parity, worker-thread telemetry, caller-cap wire shape
2026-08-28 13:11:03 +05:30
kshitijk4poor 1564a9748f fix(compression): don't force a wire cap for explicit caller max_tokens
_call_llm_impl applied auxiliary_max_tokens_param whenever
fast_compression_cap was non-None — but _compression_fast_lane_controls
passes an explicit caller max_tokens straight through, so a compression
call that set its own cap had the param force-injected onto providers
where _build_call_kwargs deliberately omits it (ZAI vision hard-400s on
max_tokens; GPT-5/Copilot require max_completion_tokens). Pre-fast-lane
main omitted the param for that exact call shape (verified via
subprocess pinned to the pre-PR base).

Gate the forced param on 'max_tokens is None' so it applies only to caps
the certified lane itself produced — the same guard the fallback path
already uses.

Regression test pins the pre-PR wire shape. Mutation-checked.
2026-08-28 13:05:06 +05:30
kshitijk4poor d24e6a34d2 fix(compression): propagate timing hooks to the protected-call worker
_run_protected_sync_provider_call propagates the forward-progress hook to
its daemon worker but not the new _aux_dispatch/_aux_provider_response
timing hooks (both threading.local). When compression takes the protected
path — the common case, since the summary call runs under
aux_interrupt_protection with a hard-cancel source — provider_dispatch_ms
and time_to_first_progress_ms were silently absent from telemetry.

Also collapse the two byte-identical save/restore context managers
(aux_progress_hook, _aux_timing_hook) onto one _aux_thread_local_hook
implementation so the propagation semantics can never drift between the
progress and timing slots.

Regression test drives _run_protected_sync_provider_call with both timing
hooks installed and asserts the worker-thread notifies reach them.
Mutation-checked (reverting the propagation fails the new test).
2026-08-28 12:57:06 +05:30
kshitijk4poor e078b2fe7c fix(compression): close the restore TOCTOU; fold review findings
Post-review hardening on the attempt-ownership commit:

- Write-time re-validation: the entry staleness check in
  _restore_compressor_attempt_state runs before the durable-cooldown DB
  I/O, so a fallback could claim the compressor in that window and the
  stale setattr loop would still clobber its state. The in-memory writes
  now re-validate AND execute under _COMPRESSOR_ATTEMPT_LOCK — the same
  lock claims are taken under. The DB rollback stays outside the lock
  (safe: the dangerous direction requires a prior claim, which the entry
  check rejects). Both the quality reviewer and the lead's independent
  pre-verification converged on this window.
  New deterministic test: TestMidRestoreClaimRace (claim injected
  between entry check and write via instrumented deepcopy).

- Documented gen-0 semantics on _claim_compressor_attempt: per-compressor
  all-or-nothing, never mixed with gen>0 on one instance (reviewer
  finding 2, verified unreachable — comment hardens against future
  confusion).

Dropped after verification (reviewer finding 3): resetting
_SUMMARY_ROUTE_CONSUMED on pin_summary_route exit — the echo lives in
the worker thread's COPIED context (propagate_context_to_thread) and
dies with it; a probe confirmed the next attempt's context is clean.
Resetting it would break digests running after the with-block.
2026-08-28 12:52:53 +05:30
kshitijk4poor 61cd299c6e fix(compression): attempt-generation ownership for overlapping stall-fallback attempts
Follow-up to #96634 (stall-fallback retry, #78981) addressing
donovan-yohan's post-merge adversarial review. The stall path detaches a
timed-out primary worker (fence cancel wins; future stays on the pool)
and immediately runs the fallback against the SAME ContextCompressor,
creating two verified races:

1. Late-primary snapshot restore: the detached primary's unwind called
   _restore_compressor_attempt_state with the PRIMARY's pre-attempt
   snapshot. Landing after the fallback's commit it rolled
   _previous_summary/cooldown/provenance/telemetry back to pre-primary
   values, silently discarding fallback-owned state.
2. Shared _compression_cancelled_check: the late primary's `finally`
   cleared the callback the fallback had just installed, so the
   fallback's F4 cancellation consult read None.

Fix: a monotonic per-compressor attempt generation claimed under one
module lock (_claim_compressor_attempt). Snapshot restores carry their
claiming generation and no-op when stale; the cancelled-check set/clear
moves into owner-stamped helpers (_install_compression_cancelled_check /
_clear_compression_cancelled_check_if_owner) so only the installing
attempt can clear it. The commit fence keeps owning COMMIT admission;
the generation owns compressor-ATTRIBUTE writes — two boundaries.
Legacy callers (attempt_generation=None) and slotted third-party
compressors (generation 0) keep the historical unconditional behavior.

Secondary review items:
- Lean chunk digests during a stall-fallback retry now follow the
  summary onto the pinned healthy route: take_pinned_summary_route()
  echoes the consumed route into a context-local
  _SUMMARY_ROUTE_CONSUMED, and _build_chunk_digests passes
  attempt_summary_route_kwargs() (non-consuming) to call_llm. The pin's
  single-use contract for the SUMMARY call is unchanged — the
  main-model retry still never re-issues the pinned route.
- Worker re-run repeating pre-compression callbacks: documented as an
  accepted limitation on _retry_compression_on_fallback_chain
  (built-ins idempotent; resuming mid-pipeline would couple the retry
  to host callback ordering).

Tests (tests/agent/test_compression_attempt_ownership.py, 10 cases):
deterministic interleavings for both races (late-primary restore
no-ops + preserves fallback state; stale finally cannot clear the
fallback's callback), legacy/slotted compatibility, digest route
follow + context-locality of the consumed echo. Mutation-checked:
reverting only the two prod files to origin/main fails the suite;
restored stack green (21 passed incl. the original #78981 suite).

The one red in the wider sweep
(test_silence_cannot_approach_double_idle_timeout) is pre-existing on
clean origin/main — verified independently.
2026-08-28 12:52:53 +05:30
kshitijk4poor d20ca3bc80 refactor(compression): consolidate fast-lane certification onto one predicate
resolve_compression_fast_lane and _compression_config_claims_fast_lane
each hand-parsed the same four config fields (provider, model,
reasoning_effort, max_output_tokens) with copy-pasted normalization and
int-coercion. Extract _fast_lane_config_fields() as the single source of
truth for both.

This also fixes a real inconsistency the duplication hid: certification
checked the literal string 'none' while _get_task_extra_body routes
reasoning_effort through parse_reasoning_effort, which treats 'false',
'disabled', and YAML boolean false as disabled too. A user writing
reasoning_effort: false got reasoning disabled but silently lost the
fast-lane cap. Certification now delegates to parse_reasoning_effort so
the two predicates can never disagree.

Regression test: every disabled-spelling certifies; empty/real efforts
do not. Mutation-checked (reverting to the literal check fails the new
test).
2026-08-28 12:48:40 +05:30
kshitijk4poor 699bfcd026 perf: skip dict copy for non-compression auxiliary calls
_compression_fast_lane_controls unconditionally did body = dict(extra_body)
before the early-return guard. Move the copy below the guard so non-
compression auxiliary calls (vision, title_generation, etc.) return the
original extra_body reference without a shallow copy.
2026-08-28 12:38:49 +05:30
Mike DeMott dd5481aa40 refactor: centralize fast compression controls 2026-08-28 12:38:49 +05:30
Mike DeMott 3581983459 fix(compression): reject boolean fast caps 2026-08-28 12:38:49 +05:30
Mike DeMott 7568dd551b fix(compression): contain drifted fast controls 2026-08-28 12:38:49 +05:30
Mike DeMott 372c4cdfce fix(compression): certify the effective fast route 2026-08-28 12:38:49 +05:30
Mike DeMott 213ae08e7a perf(compression): add guarded fast summary lane 2026-08-28 12:38:49 +05:30
Teknium 420e156bf3 refactor(agent): single owner for Responses route predicates
Follow-up to the #96217 salvage: the codex/xai/github route checks were
re-implemented inline at four sites (codex_responses_adapter helpers,
chat_completion_helpers kwargs build, _is_openai_codex_backend, the
run_agent silent-reject hint). Consolidate them into
classify_responses_route() / ResponsesRouteFlags in
codex_responses_adapter and migrate every site — backend-identity
predicate class (#22548/#70893/#59561/#72468).

Host checks use exact-host-or-subdomain semantics, never substring
matching.
2026-08-27 18:56:21 -07:00
686f6c61 e6b4f3750b fix(agent): count native Responses preflight against pruned wire
Automatic preflight used the full durable transcript even when the
Codex Responses request would prune around a native compaction
checkpoint. That false-triggered a 600s local summary against history
the main request never sent. Estimate the converted, checkpoint-pruned
payload when native compaction is eligible, and keep the generic
estimate as the conservative fallback.
2026-08-27 18:56:21 -07:00
Mariano Nicolini 35e0d15861 fix(aux): read the Nous fast-model catalog with credentials, and filter it
`_fast_model_from_catalog` treats the catalog's keys as a source of ids,
scanning them for a cheap model to use for side tasks like titling. Two
problems for Nous.

The credential lookup goes through `resolve_api_key_provider_credentials`,
which raises for Nous because it is OAuth. The read then went out
anonymous and came back with the full catalog rather than the one the org
may reach, so a policy-hidden model could be selected and then refused at
request time with `model_blocked_by_org_policy`.

Fall back to the Nous credential resolver when the api-key path raises,
and narrow the resulting ids by the org policy the same way the pickers'
lists are narrowed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:14:22 -03:00
Mariano Nicolini 4caeb02735 fix(models): key the pricing cache on auth state, not just the base URL
`fetch_models_with_pricing` checked its cache above the point where the
Authorization header is built, and keyed that cache on the base URL alone.
Whichever read of a given base URL landed first in a process therefore
answered every later read, whatever key it passed — a non-empty result is
held for the life of the process.

That is wrong for any endpoint whose answer depends on who is asking. The
Nous inference gateway filters `GET /v1/models` by the caller's org model
policy, so an anonymous read landing first makes a later authenticated read
return the full, unfiltered catalog without a request going out.

Separate the URL root from the cache key and fold auth state into the
latter. Only whether a key was supplied participates, never its value, so
no secret reaches the key.

`credits_tracker` peeked into the private `_pricing_cache` and duplicated
the key shape to do it; it now calls `peek_cached_pricing`, which owns both
the /v1-suffix normalization and the preference for the authenticated
catalog.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:12:58 -03:00
kshitijk4poor 80ab7d2b1c fix(compression): dedupe current-turn rows when rotation splits the session mid-turn
When context-compression rotation fires mid-turn, the current user
message was persisted twice into the child session. Root cause: dedup
used id()-seeded sets of copies instead of markers on the live objects.

Replace with _DB_PERSISTED_MARKER-based dedup as the sole authority:
- _ensure_compressed_has_user_turn returns CompressedUserTurnOutcome
- After publish_compression_child succeeds, stamp the live anchor-source
  row (not a drifted index) with _DB_PERSISTED_MARKER
- _sync_persisted_markers mirrors stamps from result to live lists by
  scoped identity (handles direct-path, adoption divergence, _session_messages)
- Remove _flushed_db_message_ids from rotation commit path (markers replace it)
- Unconditional (loud) imports — no silent fallback

Salvage of #94996 by @fedosis, rebased on top of #95433 (stall-fallback,
already merged). Both conversation_compression.py and run_agent.py are
built from origin/main + #94996's diff applied on top, preserving the
force_terminal refactor and _publish_new_fence from #95433.

Credit: @fedosis original PR #94996.
2026-08-28 02:47:34 +05:30
kshitijk4poor ba7df6a4bf refactor: extract shared _coerce_positive_timeout helper (#95433)
The timeout validation (isinstance(raw, (int, float)) and not
isinstance(raw, bool) and raw > 0 → float(raw)) was duplicated
between auxiliary_client.py:_fallback_entry_timeout and
conversation_compression.py:resolve_compression_fallback_route.
Extracted into _coerce_positive_timeout to prevent drift if the
config schema evolves.
2026-08-28 02:29:27 +05:30
Shaun Eccles 6151e59d65 fix(compression): surface an unpublished stall-fallback fence at WARNING
Review follow-up for #95433. When the host's fence factory is absent or
raises, the retry runs on a private CompressionCommitFence() that
hard-interrupt admission never reads — /stop would serialize against the
aborted attempt's fence instead of the retry's commit boundary. Promote
the factory-failure swallow from debug to warning and warn when the retry
has no published fence, so the control-plane degradation is visible rather
than silent.

Also documents the pin's coverage (single _generate_summary call; the
lean-mode chunk digests are a separate unpinned call path).
2026-08-28 02:29:27 +05:30
Shaun Eccles 2c6938dc3a fix(compression): retry a stalled summary on the fallback chain (#78981)
A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.

The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
2026-08-28 02:29:27 +05:30
hbentel 2673d5f5bd fix(gemini): embed images in Gemini 3.x functionResponse.parts for multimodal tool results
_translate_tool_result_to_gemini called _coerce_content_to_text unconditionally,
silently dropping image_url parts from multimodal tool results (e.g. vision_analyze
responses). Gemini 3.x supports a functionResponse.parts field for embedding
inlineData images directly inside the function response; Gemini 2.x does not.

Thread is_gemini3 through _build_gemini_contents → _translate_tool_result_to_gemini
and gate image embedding on _gemini_major_version >= 3. Reuses the existing
_extract_multimodal_parts helper (no duplicate code). Non-3.x path unchanged.

Original PR #32352 by @hbentel, salvaged onto current main.

Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
2026-08-27 21:07:21 +05:30
nftpoetrist 08b4875f4a fix(deadline): remove the dead second SuspectableBackend class shadowing the Phase 3a Protocol
agent/deadline.py defined SuspectableBackend twice: the Phase 3a Protocol
(sync ensure_healthy(self) -> bool) and, further down the same module, an
unrelated concrete class with the same name (async
ensure_healthy(self, timeout=5.0)) added later by the MCP Phase 3b adopter.
Since Python executes class statements top-to-bottom, the second definition
silently shadowed the first at module scope.

Nothing in the tree imports or subclasses either by name today — the MCP
adopter duck-types the same-shaped contract directly on its own connection
class rather than referencing agent.deadline.SuspectableBackend — so this
caused no live behavior change. But it left the wrong (and differently
shaped) class resolvable under that name for the next Phase 3b adopter that
does import it for a type hint.
2026-08-27 17:22:02 +05:30
fangliquanflq a65ad15636 fix(agent): honor explicit free OpenRouter models 2026-08-27 04:37:36 -07:00
Teknium 9b44273c05 fix: follow-up for salvaged PRs #93250 + #96234
- move minimax/minimax-m3:free into the Free tier section (house
  convention: :free SKUs group together, matching glm-5.2:free and the
  nemotron :free entries) and regenerate model-catalog.json
- add Inkling family context length (1,048,576 — OpenRouter live
  metadata, 2026-08-27) to DEFAULT_CONTEXT_LENGTHS; new family slug
  otherwise fell through to no entry
- add Inkling to the reasoning stale-timeout floor table (300s tier,
  same as Grok reasoning / Ox Alpha; OpenRouter marks the family as
  reasoning-capable)
- widen the floor matcher's right-anchor separator class to include
  ':' so OpenRouter SKU suffixes (:free/:batch/:nitro) inherit the
  family floor — inkling:free previously missed the inkling entry
- regression tests for the inkling floor + ':' separator
2026-08-27 03:24:20 -07:00
kshitij f45477b4a8 Merge pull request #86412 from kshitijk4poor/salvage/83225-overflow-clamp
fix(approval): oversized approvals.timeout crashes parallel tool batches — clamp at config read (salvage #83225/#83298, #85125 2b)
2026-08-27 14:26:38 +05:30
kshitijk4poor c693c772b5 refactor(compression): fold review follow-ups on #71488 salvage
- Reuse the existing _commit_status variable for the terminal-edge gate
  instead of the parallel _compaction_succeeded boolean (derived state).
- Give the commit_fence_cancelled abort the same force_terminal=True
  terminal edge as the lock-contended abort, and reword the closure
  comment that overstated the lock contender as 'the one exception'.
- Inline the codex app-server path's lifecycle closure: after gating on
  success it reduced to a single success-site emit, so the scaffolding
  (done-flag + closure + two no-op failure-path calls) was dead.
2026-08-27 12:42:38 +05:30