Commit Graph

3725 Commits

Author SHA1 Message Date
Teknium c9e2a46df6 fix(pricing): support Gemini context-tiered rates in pricing snapshot (#93469)
The pricing snapshot could only express flat per-million rates, so
gemini-3.1-pro sessions with prompts over 200k tokens under-counted
input 2x ($2 vs $4/M) and output 1.5x ($12 vs $18/M).

- Add optional tier fields to PricingEntry: tier_threshold_tokens,
  input/output/cache_read_cost_per_million_above (None = flat, falls
  back to base rate per-field).
- estimate_usage_cost selects the above-threshold rates for the WHOLE
  request once usage.prompt_tokens (input + cache read + cache write)
  exceeds the threshold, matching Google's billing semantics.
- Populate gemini-3.1-pro (4.00/18.00/0.40 above 200k; alias
  gemini-3.1-pro-preview inherits) and gemini-2.5-pro (2.50/15.00
  above 200k).
- Flat entries are untouched: no threshold means no behavior change.

Reported and tier-field shape designed by @tornike14 (#93469).

Tests: below/at threshold unchanged, above-threshold tiered whole-request
pricing, cache-read tier rate and base-rate fallback, preview alias,
flat entries unaffected.
2026-08-24 03:23:07 -07:00
Kyzcreig 4d729e4b31 fix(model-metadata): local ctx probe must not read max_tokens as the context window
The two local-server context probes in _query_local_context_length read
data.get("max_tokens") as a context-window candidate. On an
OpenAI-compatible /v1/models passthrough max_tokens is the max OUTPUT
tokens, so a 1M-context model advertising a 128K output cap resolves to
128000 and auto-compaction fires ~7x early.

Route both branches through the module's own key vocabulary
(_CONTEXT_LENGTH_KEYS), which already classifies max_tokens as a
_MAX_COMPLETION_KEYS entry.
2026-08-24 03:20:57 -07:00
carlotestor 394f0f0902 fix(context): prefer max_input_tokens over max_tokens for Anthropic proxies
Local /v1/models probes treated Anthropic `max_tokens` (max output) as the
context window when `max_model_len`/`context_length` were absent. Anthropic
and Anthropic-compatible reverse proxies expose both:

  max_input_tokens = context window (e.g. 1M for claude-fable-5)
  max_tokens       = max output     (e.g. 128k)

That under-reported windows (1M → 128k), persisted the wrong value into
context_length_cache.yaml, and fired compression at ~96k (75% of 128k).

Route model objects through a shared helper that prefers input-window keys
via _extract_context_length, and only falls back to max_tokens when no
input-window field is present.
2026-08-24 03:20:57 -07:00
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
kchernev 10a070bd49 fix: route codex payloads around the SDK's GIL-holding request transform (#93650)
responses.create re-walks the entire request body against the
ResponseCreateParams union graph client-side while holding the GIL.
#93650 documents that walk wedging for 12+ hours on a ~1.4 MB
conversation, starving every other thread including the TTFB/stale
watchdogs — and no socket kill can unblock a pre-network hang.

Hermes payloads are JSON round-trips and already wire format, so the
bulk fields (input, tools) are now routed through extra_body, which the
SDK merges into the JSON body after the transform. Guarded by a
plain-JSON check (anything else keeps the typed path) and a
HERMES_CODEX_SDK_TRANSFORM=1 escape hatch. Applied to both the primary
stream path and the auxiliary adapter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 15:29:45 +05:30
kshitijk4poor f93b350711 fix: align cache-policy pre-gate identity with the capability matcher
Follow-ups on top of the salvaged #92785 commit:

- Pre-gate now matches base URLs via normalize_route_base_url and
  provider ids via custom_provider_aliases, mirroring the semantics of
  get_custom_provider_model_capability. The raw string comparison
  silently dropped declarations whose config spelling differed only by
  host case or trailing slash (proven empirically: …/v1/ vs …/v1 with a
  non-matching provider name returned (False, False) despite an explicit
  prompt_caching: true).
- get_provider(..., allow_network=False) in the early-init/stub branch:
  the policy runs per request destination (MoA aggregator, auxiliary
  replans via blank_cache_policy_stub, early agent init) and a cold
  models.dev cache triggered a measured ~450 ms foreground registry
  fetch from the send path. A catalog miss degrades to the conservative
  side.
- Debug-log the previously silent provider-lookup exception fallback.
- Tests: _make_agent defaults _custom_providers=[] (post-init reality;
  keeps built-in-route tests off the catalog/config fallback), the two
  early-init tests delete the attr explicitly, and three regression
  tests pin the URL-drift, spaced-legacy-name, and no-network contracts
  (all three fail on the unfixed commit).
2026-08-24 14:58:42 +05:30
Blood Shot 0204e4898e fix(agent): normalize custom provider route identity 2026-08-24 14:58:42 +05:30
Blood Shot 0a3b7efec5 fix(agent): honor prompt_caching for custom providers
Apply explicit per-model prompt_caching capabilities to custom
chat-completions routes, rather than limiting them to recognized providers,
hosts, or model families.

Keep undeclared routes conservative, derive the marker layout from the wire
transport, and leave Responses and Bedrock caching paths unchanged.
2026-08-24 14:58:42 +05:30
Gille 21b92d2687 fix(agent): bypass response cache for empty retries 2026-08-23 23:31:54 -07:00
Teknium 0c3a507535 fix(classifier): 429 quota walls route to billing across providers; reset signals stay rate-limited
Consolidates the 429-quota-classifier cluster on top of the merged #93419
Anthropic core. Three independent contributor findings salvaged into one
coherent change to the single 429 branch:

- Broaden the 429 usage-limit check from the narrow 'usage limit' string to
  the full _USAGE_LIMIT_PATTERNS ('quota', 'limit exceeded', 'key limit
  exceeded') and add _BILLING_PATTERNS detection on 429 ('insufficient
  credits' wrapped in a 429 instead of 402), guarded by a _RATE_LIMIT_PATTERNS
  exclusion so an explicit 'Rate limit exceeded' never promotes to
  non-retryable billing. (credit @Pluviobyte, #39441 — earliest submitter)
- Add 'resets in' to the transient signals: Codex's 'Weekly usage limit
  reached. Resets in 6hr 29min.' wrongly read as terminal billing because
  main only had 'reset in' (no substring match). (credit @LeonSGP43, #63021)
- Add 'reset after' / 'available in' / 'per minute' / 'per second' transient
  signals. (credit @jtstothard, #74785)

Supersedes #65633 (defective branch placement, no tests). The aux-client
path already covers these shapes (_is_payment_error catches weekly/quota
walls; _is_rate_limit_error treats 'resets in' as transient), so no change
there.

Tests: 6 new cases (generic quota wall, insufficient-credits 429, rate-limit
guard, Codex resets-in, extra transient phrases). Guard sabotage-verified.

Co-authored-by: Pluviobyte <Pluviobyte@users.noreply.github.com>
Co-authored-by: LeonSGP43 <LeonSGP43@users.noreply.github.com>
Co-authored-by: jtstothard <jtstothard@users.noreply.github.com>
2026-08-23 20:02:07 -07:00
fangliquanflq 37411f349a fix(auth): rotate credentials for named custom providers after 401/429
Salvage of #93214 (5 commits squashed onto current main; agent_runtime_helpers.py
diverged since the PR base and was 3-way reapplied). The credential-rotation
guard in recover_with_credential_pool and both restore_primary_runtime paths
only tolerated the custom-naming split when the agent carried the literal label
'custom', so a named custom provider (agent.provider='gemini-no-filter', pool
'custom:gemini-no-filter') tripped the mismatch guard and skipped rotation on
every 401/429. Now all three guard sites use the canonical
credential_pool_matches_provider boundary predicate + resolve_runtime_pool_key,
which recognizes configured named-custom aliases and validates endpoints.

Fixes #93188.
2026-08-23 20:01:18 -07:00
Teknium 7526bd39a8 feat: every subagent's prompt embeds the workspace's project context files
Widened from /review to the class: _build_child_system_prompt now runs
the parent's resolved workspace_path through
agent.prompt_builder.build_context_files_prompt (same discovery/
priority/caps as the main system prompt: .hermes.md > AGENTS.md chain >
CLAUDE.md > .cursorrules; SOUL.md skipped) and embeds the result as
binding conventions. All delegate_task children get it — reviewer
included — since children are built with skip_context_files=True and
previously worked in repos without the repo's own conventions.

The review-engine-local load_workspace_context duplicate is removed;
the reviewer inherits the block via the shared child prompt path.
workspace_path comes only from explicit sources (_resolve_workspace_hint
— TERMINAL_CWD / agent cwd hints, never bare getcwd), so the #64590
install-tree-fallback guard concern doesn't apply.

Tests moved to pin the generalized path (real-filesystem AGENTS.md via
_build_child_system_prompt, empty/no-workspace negatives, reviewer E2E
through start_review). Docs: subagent-context section + /review flow
(en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium 23fb949f2c feat: /review briefing embeds the workspace's project context files
load_workspace_context() resolves the parent's workspace via the same
_resolve_workspace_hint used for child prompts (explicit sources only —
TERMINAL_CWD / agent cwd hints, never a bare getcwd fallback, so the
#64590 install-tree-leak guard concern doesn't apply) and runs it
through agent.prompt_builder.build_context_files_prompt — the exact
discovery/priority/cap logic the main system prompt uses (.hermes.md >
AGENTS.md chain > CLAUDE.md > .cursorrules; SOUL.md skipped). The
result is embedded in the reviewer briefing as binding review
standards. Subagents are built with skip_context_files=True, so without
this the reviewer judged repo work without the repo's own conventions.

5 new tests incl. real-filesystem AGENTS.md discovery through the real
loader. Docs updated (en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium 22381edc11 feat: /review briefing carries the parent's loaded skills
The reviewer subagent now inherits the primary agent's working skill
context: collect_parent_loaded_skills() gathers launch-preloaded skills
(from the activation notes in ephemeral_system_prompt) and mid-session
skill_view loads (from assistant tool_calls in history), deduped and
capped at 8, and the briefing instructs the reviewer to skill_view each
and treat their conventions as binding for the assessment.

Reference-file reads (file_path=...) don't count as loads; full-skill
injection was rejected as too costly (a single dev skill can be 40KB+).

Docs: delegation.md /review flow updated (en + zh-Hans).
2026-08-23 19:04:37 -07:00
Teknium 580060ffd8 fix: reuse first-observed sequence when announced items land via output_item.done
Follow-up to salvaged PR #92767 (review round 2 P1): the .done path
allocated a fresh tail sequence even for items announced earlier via
output_item.added, so a mixed announced/pending stream without
output_index values reordered the calls ([B, A] instead of [A, B]).
First-observed ordering metadata is now recorded for every announced
item and reused at .done; a fresh sequence is allocated only for
genuinely unannounced items. The .done event's own output_index wins
when present, with the announced index as fallback.

Regressions: two announced calls without indices where the first later
receives .done; an announced non-function item preceding a pending call.
2026-08-23 19:02:30 -07:00
cxxCoolStar 4f3ae189a3 fix(agent): harden pending Responses tool call settlement 2026-08-23 19:02:30 -07:00
cxxCoolStar 720344cfba fix(codex): settle pending Responses tool calls when output_item.done is omitted
Backends that omit per-item done events on a successful completion
(anomalyco/opencode#37159) caused an announced function call to be
silently dropped: the turn ended with output == [] and the tool never
executed. Track calls announced via output_item.added, accumulate
argument deltas, and settle still-pending calls from accumulated state
at a successful terminal event. output_item.done stays authoritative.
Mirrors anomalyco/opencode#43575.
2026-08-23 19:02:30 -07:00
fangliquanflq 654d537088 fix(agent): honor structured quota reset signals 2026-08-23 18:43:12 -07:00
fangliquanflq c2090ba6b4 fix(desktop): distinguish provider quota exhaustion 2026-08-23 18:43:12 -07:00
Aniruddha Adak f778c0d941 fix(compression): structural no-ops defer retries instead of striking the breaker
Fixes #93022. A short session (protection window >= transcript) hits the
"insufficient messages" / "no compressible window" branches twice and
permanently trips the anti-thrash breaker, even though nothing was
eligible to compress - compression was never attempted, so there is
nothing "ineffective" to score. The session then rides past the
threshold with no compaction possible (recovery probes only soften,
not fix, the misclassification).

Distinguish "nothing eligible right now" from "attempted and
underperformed":

- New transient _structural_no_op_backoff_until (in-memory, 300s)
  armed by _record_structural_no_op() at the three structural no-op
  sites: insufficient_messages, no_compressible_window,
  empty_post_handoff_window. No strikes accumulate; auto-compaction
  resumes on its own once the backoff lapses or the transcript outgrows
  the protection window.
- The backoff gates should_compress via
  _automatic_compression_blocked_locally and surfaces in
  _compression_block_reason as "structural_backoff:<seconds>".
- #40803's frozen-CLI guarantee is preserved: a transcript that can
  never shrink retries at most once per backoff window instead of
  every turn.
- force=True (/compress) clears an active backoff before attempting;
  record_completed_compaction() lifts it - both prove the transcript
  is compressible/being worked.
- Genuine attempted-but-underperformed verdicts still strike the
  durable ineffective counter unchanged.

Tests: new tests/agent/test_context_compressor_structural_backoff.py;
updated the two tests that asserted the old strike-on-noop behavior.
2026-08-23 18:27:07 -07:00
Aintworth ce51f535d3 test(gemini): cover nested/list/non-pointer ref cases; document false-positive tolerance
Address review feedback:
- Add tests for deeply-nested $ref (recursion), top-level JSON array
  (already wrapped, no 400 path), and $ref without '#/' prefix (stays
  structured).
- Document the deliberate structural (false-positive-tolerant) detection and
  its O(n) cost in the helper docstring.
2026-08-23 18:27:04 -07:00
Aintworth 03477166f9 fix(gemini): wrap schema-bearing tool results as opaque text
Gemini 3 resolves JSON-Schema $ref/$defs pointers inside a
functionResponse.response payload and rejects unknown references with
HTTP 400 INVALID_ARGUMENT ('referenced name #/$defs/...' does not match
a display_name; see vercel/ai#14369).

tool_describe (and any tool whose result is itself a JSON Schema) returns
schema text that previously went back as a structured response, tripping
Gemini's pointer resolution. Detect such results with a $ref-pointer scan
and wrap them as opaque text instead.

Adds regression tests for the wrap path and the unchanged structured path.
2026-08-23 18:27:04 -07:00
Finn763 74e6885f0d fix(review): fail-closed compressor detachment + warm-cache first request (#93057 review)
Adversarial-review fixes for the #93057 snapshot-compaction PR:

- Fail-closed detachment: only re-enable compression after
  bind_session_state successfully severs the engine's parent binding.
  A failed rebind keeps the historical compression_enabled=False
  behavior and warns, instead of running compaction against a
  compressor still bound to the parent's SessionDB (#38727 re-open).
- Warm-cache parity: defer both compression gates (turn-prologue
  preflight + pre-API pressure check) until the fork's first provider
  response, so the first request replays the full snapshot as the
  intended cached read and compaction applies from the second request
  on — matching the documented budget mental model.
- Tests: regression for the rebind-failure fail-closed path (red on
  pre-fix code) and the existing threshold-crossing test reworked to a
  two-request review asserting the warm first request + compacted
  second request. 116 tests green across all touched suites; ruff
  clean.
2026-08-23 18:25:19 -07:00
Finn763 4202a508fd fix(review): bound same-model background review replay
Detach the review fork's compressor from the parent SessionDB/session_id
and re-enable in-memory-only compaction for oversized snapshots, instead
of the historical compression_enabled=False guard that left the fork's
replayed transcript unbounded (350k-384k input tokens per request, 1.49M
total across one 8-request review). Add an aggregate input-token budget
(auxiliary.background_review.max_input_tokens, default 600k) so repeated
tool calls cannot recreate an unbounded transcript; the tool loop stops
before the provider call that would cross it.

Closes #93057
2026-08-23 18:25:19 -07:00
fangliquanflq 2033f4cc34 fix(agent): separate cancellation diagnostics from tool output 2026-08-23 18:25:19 -07:00
fangliquanflq c1c0efa375 fix(code-exec): preserve interrupt cancellation source 2026-08-23 18:25:19 -07:00
Teknium a2a43f7e82 fix(agent): widen composite-id alias matching to the compressor; unify variant policy owners (#63000)
Follow-up on top of the salvaged #93335:

- context_compressor._sanitize_tool_pairs now expands alias spellings on
  the RESULT side too (tool_result_id_variants), so a composite
  call|item-keyed result pairs with its split-field tool_call instead of
  being dropped and its call stripped.
- The compressor's _tool_call_id_variants staticmethod and
  agent_runtime_helpers' module-level _tool_call_id_variants are now thin
  forwarders to agent.message_sanitization.tool_call_id_variants — one
  policy owner for alias expansion, so the pre-call sanitizer, repair
  pass, dedup pass, and compression sanitizer can never drift apart.
- Preserved the #91768 SDK-object tolerance in repair pass 1 (the
  shared helper handles non-dict tool_calls via getattr; the salvaged
  commit's isinstance-dict guard was dropped in the merge resolution).

New regression tests: composite-keyed results through
sanitize_api_messages (both directions) and _sanitize_tool_pairs, with
negative controls. Sabotage-verified: compressor test fails with raw
tool_call_id tracking.
2026-08-23 18:24:43 -07:00
joaomarcos 5496d5995a fix(agent): preserve tool results across ID variants
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
2026-08-23 18:24:43 -07:00
liuhao1024 51239e8e2a fix(vision): forward the API key to the server-type probe and cache failed verdicts
The image-routing vision path calls detect_local_server_type without
the provider's API key. Against a remote API-keyed endpoint (sglang /
vLLM with --api-key) every leg of the 5-request probe waterfall came
back 401 — and because a failed verdict was never written to the
in-memory cache (only positive verdicts were), the waterfall re-ran on
EVERY image-bearing turn (#89863: 51 detail-less busy-acks observed in
one Slack channel while the probe sprayed the user's own server).

Two changes:

- image_routing._should_probe_ollama_vision now takes the API key and
  forwards it; a new _resolve_inference_api_key mirrors
  _resolve_inference_base_url's resolution order (runtime value,
  model.api_key, providers blocks) so the key always matches the URL
  being probed.

- detect_local_server_type caches a None verdict in memory with a short
  failure TTL (5 min, vs 1h for positives) so the next turn is served
  from the negative entry instead of re-running the waterfall — while
  a transient failure (server starting, key being fixed) recovers in
  minutes. Negative verdicts are deliberately not written to the
  cross-process disk cache.
2026-08-23 17:47:50 -07:00
JinUltimate 4ca993c746 fix(image_routing): stop fingerprint-probing remote OpenAI-compatible endpoints
Fixes #89863. With a custom: provider pointing at a remote, API-keyed
endpoint (sglang/vLLM/OpenAI-compat), every image turn triggered a 5-request
probe waterfall without Authorization, spraying 401s at the backend.

Two fixes:

1. _should_probe_ollama_vision now takes api_key and forwards it to
   detect_local_server_type so keyed local servers don't 401.

2. When provider != 'ollama', remote endpoints (per is_local_endpoint) are
   rejected early — server-fingerprint probing is only valid for local
   boxes. Non-Ollama remotes expose Ollama-compat endpoints that can
   misidentify and trigger unnecessary /api/show probes.

_lookup_supports_vision resolves the runtime api_key via
_runtime_main_value and forwards it to both helpers. New test class
TestShouldProbeOllamaVision covers the contract in both directions.
2026-08-23 17:47:50 -07:00
honor2030 fe483de4d3 fix(agent): keep max-iteration warnings out of quiet stdout
Route the max-iterations diagnostic through logging when quiet_mode is active so automation wrappers keep stdout machine-readable.

Add a regression test covering quiet max-iteration summary handling.
2026-08-23 17:45:58 -07:00
Teknium 12395e57b4 feat: /review command — independent reviewer subagent on every surface
/review takes the last 10 chat messages plus optional instructions,
spawns a full-privilege background subagent (the async delegation
rail) that investigates the referenced work (PR, code, docs), and its
complete review re-enters the spawning session as a normal
async-delegation completion the primary agent can act on.

- agent/review_engine.py: shared engine (snapshot, briefing,
  auxiliary.review credential resolution, dispatch, note formatting)
- tools/delegate_tool.py: internal credentials_cfg per-call override
  (never model-facing) resolved through the same credential system as
  delegation.provider pins
- auxiliary.review config block (provider/model/base_url/api_key/
  api_mode); provider auto + empty model = inherit the main model
- Surfaces: CLI process_command, gateway run.py dispatch +
  slash_commands handler (binds the approval session key so the
  completion routes back), TUI/Desktop live dispatch in
  tui_gateway/server.py, CommandDef registry (+Slack /hermes-only cap)
- Docs: delegation.md section + slash-commands.md (both tables)
- Tests: 15 engine tests (sabotage-verified: credentials_cfg and
  dispatch tests fail without the fix), 4 gateway handler tests
  through the real async rail
2026-08-23 17:38:38 -07:00
Teknium faa2399e2b fix(agent): make the pre-call dedup pass variant-aware; widen batch regression coverage (#93251)
Follow-up on top of the salvaged cluster: sanitize_api_messages step 3
(duplicate tool_call_id dedup) still tracked only the coalesced
(call_id||id) value in outstanding_call_ids, so after step 2's
variant-aware matching preserved a result keyed on the OTHER id variant,
step 3 deleted it as answering no outstanding call — whole parallel
batches of real results vanished with no stub at all (#93251's total-loss
mode). Track the full variant set per call and consume all siblings when
answered, preserving #58327 duplicate protection and llama.cpp
constant-id re-arm semantics.

Also aligns the #58287 compressor test with the in-flight tool chain
protection (#79278) that landed after that PR was opened: a trailing
user turn keeps the negative-control assistant message out of the
protected trailing window.

New regression tests: divergent-id batch survival through the dedup
pass, sibling-id replay still dropped, constant-id re-arm preserved.
Sabotage-verified: tests fail with the old single-id tracking.
2026-08-23 17:01:20 -07:00
joaomarcos 36b4da5489 fix: repair_message_sequence drops tool results for SDK tool_call objects
The tool_call id-matching pass in repair_message_sequence only read
`.get("id"/"call_id")` on plain dicts, skipping non-dict tool_calls
entirely (`if not isinstance(tc, dict): continue`). Host-fed and
pre-serialization histories can carry unserialized SDK tool_call
objects (e.g. `ChatCompletionMessageToolCall`) instead of dicts, which
left `known_tool_ids` empty for that assistant turn. The following
`tool` message — a legitimate result already produced by executing the
tool — was then misclassified as an orphan and silently dropped,
corrupting the persisted conversation history and leaving the
assistant's tool_calls unanswered (itself a trigger for HTTP 400 on
strict providers).

Fix: extract id/call_id via getattr() for non-dict entries too,
mirroring AIAgent._get_tool_call_id_static's existing dict-or-object
tolerance, instead of skipping them.
2026-08-23 17:01:20 -07:00
Frowtek b9a62f6590 fix(agent): consume every tool_call id variant when pairing tool results
`repair_message_sequence` registers BOTH `id` and `call_id` for each
assistant tool_call, because a matching tool result may be keyed on either
depending on which path built it (#58168). The duplicate guard added for
dropped rather than replayed.

Those two behaviours don't compose: a Codex/Responses tool_call registers
two DIFFERENT ids (`fc_...` and `call_...`), but only the id the first
result referenced is discarded. Its sibling stays in `known_tool_ids`, so a
duplicate result keyed on that sibling still matches and is kept — two tool
messages replayed for one call, which is exactly the HTTP 400 on strict
providers the consume step exists to prevent.

Duplicates of this kind come from the retry / crash / session-resume glitch
the guard was written for; the id-variant split just lets them slip past it.

Track each registered id back to its tool_call's full variant set and
discard all of them on a match. Results keyed on either variant are still
accepted (no false orphaning), and two parallel Codex calls answered via
different variants both survive.

Adds regression tests for the sibling-keyed duplicate and for the
two-calls/mixed-keys case that must NOT be affected.
2026-08-23 17:01:20 -07:00
srojk34 52fb5081cc fix(compression): register both id/call_id variants in _sanitize_tool_pairs
_sanitize_tool_pairs() matched tool_call/tool_result pairs using a
single-value call_id||id precedence per tool_call (_get_tool_call_id).
In the Codex Responses API format an assistant tool_call carries both a
distinct id (fc_...) and call_id (call_...); a tool result's
tool_call_id may be keyed on either depending on which code path built
it. Whenever a genuinely matching pair used the field the precedence
didn't pick, the sanitizer misclassified it as orphaned on BOTH sides:
it dropped the valid tool result AND stripped the tool_call from the
assistant message, even though neither was orphaned.

Live-verified before the fix: {"id": "fc_777", "call_id": "call_777"}
+ a tool result with tool_call_id="fc_777" (a valid pair) was fully
removed by current main.

Register both id and call_id as valid match keys via a new
_tool_call_id_variants() helper (a set per tool_call, not a single
value), matching #58168's fix for repair_message_sequence's known-id
set today. A tool_call now survives if ANY of its id variants has a
matching result, which is not vulnerable to precedence order at all
(unlike swapping which field is checked first, which only trades which
sub-case is broken).

Note on #56425 (open, unreviewed): that PR touches this same function
for the same underlying issue (#55626) by swapping the call_id||id
precedence to id||call_id. That fixes the specific case where a result
matches `id` but not the reverse case (a result matching `call_id`
while `id` is also present) -- the precedence-swap approach cannot fix
the class, only relocate which sub-case is broken. This fix instead
mirrors the already-merged #58168 pattern (register the superset of
both ids as valid matches), which has no such blind spot. Adds 2
regression tests: the previously-mismatched case, and a negative
control confirming genuine orphans are still stripped alongside a
valid dual-id pair in the same window.
2026-08-23 17:01:20 -07:00
Bartok9 1a83b1e588 fix(agent): keep tool results keyed on a tool_call's id variant (#55626)
Register every id variant (call_id AND id) of each assistant tool_call in
sanitize_api_messages so a tool result keyed on either variant is treated
as paired. Previously only the coalesced (call_id||id) value was
registered, so Responses-style tool_calls carrying divergent id (fc_...)
and call_id (call_...) had their real results dropped as orphans and
replaced with '[Result unavailable]' stubs.

Cherry-picked from PR #56148 (unrelated busy_ack_templates files dropped
per the author's own follow-up commit).
2026-08-23 17:01:20 -07:00
kshitijk4poor b8cd00f968 fix: set failed=True for repeated_outer_errors exit + drop append_message
Follow-up to @BrunoBza's #93062:

1. Set failed=True only for the new repeated_outer_errors exit reason.
   Previously the error exit left failed=False, so finalize_turn reported
   completed=True for a turn that actually failed — incorrect.

2. Don't append_message the assistant response at the break. A thinking-
   prefill or interim assistant may already be the tail, and appending
   would create assistant→assistant role-alternation violation.
   finalize_turn (lines 341-353) handles this safely by checking
   _tail_role != 'assistant' before appending.

3. Update test to assert failed=True and completed=False for the
   repeated_outer_errors exit.
2026-08-24 04:47:07 +05:30
BrunoBza 56e7fd2adf fix(loop): bound outer-loop error retries per turn instead of relying on max_iterations (#92450)
The outer conversation-loop except handler only left the loop on a
local-processing error or when api_call_count >= max_iterations - 1.
With the turn budget now unlimited by default (sys.maxsize), a
permanent failure that escaped the inner retry/fallback machinery
retried forever: ~64 retries/s, one core pegged, and the rotated
agent.log history overwritten within minutes.

Bound the loop with a small per-turn cap on total escaping exceptions
(_MAX_OUTER_LOOP_ERRORS = 8, scaled down by a tiny explicit
max_iterations so a manually bounded budget still governs). The legacy
local-processing and near-limit exits are byte-identical; a new
'repeated_outer_errors' exit reason gets a user-facing explanation.

The inner retry/fallback layer owns transient API recovery and
terminates on its own, so only exceptions that escape it reach this
cap - a successful turn is unaffected.

Fixes #92450
2026-08-24 04:47:07 +05:30
Teknium 26cf456598 refactor: reconcile with #93269 — one shared shutdown predicate, keep both guard sites
PR #93269 (kshitijk4poor) landed the outer-handler break for the same
symptom while this branch was in flight. Keep his guard (it covers
shutdown errors from local post-processing and does the resume-hint +
best-effort persist) and keep this branch's inner-retry-handler return
(it fires BEFORE the ⚠️ retry trace, credential rotation, and fallback
attempts that the outer handler never sees). Point his
_is_interpreter_shutdown_error at tools/interpreter_shutdown.py so the
class has exactly one text-matching site, preserving his RuntimeError
type gate and all 7 of his tests.
2026-08-23 16:04:41 -07:00
Teknium 9ea7fe9938 fix: quitting the CLI no longer spams shutdown-race API errors onto the shell
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.

Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
  (matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
  predicates now delegate to the shared home (tool_executor previously
  matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
  shutdown signal and abandons the turn — one log warning, no print,
  no traceback, no debug dump, no retry; outer handler gets the same
  guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
  (set by the background-review fork) instead of bypassing it.

Refs #55924 #58720 (same class in cron delivery), adjacent to #90683.
2026-08-23 16:04:41 -07:00
kshitijk4poor e63786fc23 fix(agent): break immediately on interpreter shutdown in conversation loop
When the Python interpreter begins teardown (user closes hermes, SIGTERM,
OOM-kill), every executor-backed operation raises 'cannot schedule new
futures after interpreter shutdown'. The outer except handler in
run_conversation caught this error but did not recognize it as fatal —
it kept retrying (API calls #4, #5, #6) until max_iterations, each time
hitting the same dead executor and printing another traceback.

The fix adds an early check: if sys.is_finalizing() or the error matches
the 'cannot schedule new futures' pattern, break immediately with a clean
interpreter_shutdown exit reason instead of retrying. The codebase already
had this pattern in cron/scheduler.py and agent/tool_executor.py — the
conversation loop just wasn't using it.
2026-08-24 04:14:26 +05:30
Teknium 933c209e96 fix(compression): -900k Codex variants keep the global 50% threshold; 85% autoraise stays on 272K base slugs
The 85% compaction autoraise exists to stop wasting the small advertised
272K Codex window. -900k large-context picker variants (#92797) run at
~900K, where the global compression.threshold (default 50%, ~450K) is the
right behavior — autoraising them to 85% (~765K) would delay compaction
far past what the user configured.

- _is_codex_gpt54_or_gpt55() excludes valid -900k variants, so both the
  85% override and the one-time autoraise notice skip those sessions.
- Base slugs are unchanged: 272K window + 85% autoraise.
- Tests: variant/base threshold pairs incl. namespaced ids; docs note in
  the -900k section.
2026-08-23 02:54:30 -07:00
kshitijk4poor b7544dba01 fix(compression): share one image-strip policy across demote and retire passes
The demote pass (pass 2) and the retire pass (3.5, #92783) each carried
their own copy of the two image-strip branches. The copies had already
diverged: the retire pass dropped the stale api_content sidecar on
rewrite, the demote pass did not — leaving an exact-wire sidecar that
replay could use to resend the pre-strip image bytes.

Extract _strip_images_from_tool_msg as the single policy owner; both
passes now use it, closing the sidecar gap in the demote path.
2026-08-23 15:03:29 +05:30
Teknium 7b89e17774 fix(model): single exact eligibility predicate for -900k variants; reject ineligible aliases
Review findings on #92797 (@100yenadmin):
- is_codex_900k_base() is now the single source of truth used by picker
  synthesis, context resolution, /model validation, and wire stripping.
  Eligibility is an exact table (sol/terra/luna, gpt-5.4, daybreak alias)
  plus date-shaped 5.6 snapshots — family-prefix matching removed, so
  non-routable -pro slugs and unknown descendants never gain variants.
- strip_codex_context_variant_suffix() strips conditionally: ineligible
  aliases (gpt-5.5-900k) are returned unchanged and fail honestly at the
  API instead of silently running as the base model at 272K.
- validate_requested_model() rejects ineligible *-900k aliases before the
  hidden-slug soft-accept, and accepts valid variants missing from a
  stale catalog without letting the typo auto-corrector eat the suffix.
- Codex context resolver drops vendor/ namespaces, so
  openai/gpt-5.6-sol-900k resolves to 900K like the bare id.
- Table-driven regression covering eligible bases/snapshots/namespaced
  ids and rejected -pro/-mini/5.5/unknown aliases, asserting context AND
  wire model.
2026-08-23 02:14:35 -07:00
Teknium 63a9c26fbe feat(model): Codex GPT slugs default back to 272K; explicit -900k picker variants opt into the verified large window
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the
live-verified 900K burned through subscription usage for users who never
asked for the larger window (bigger window = more input tokens per request).

- Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the
  advertised 272K again — the cheaper limit is the default.
- The model picker synthesizes explicit <slug>-900k variants (e.g.
  gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into
  the 900K window. Slugs that genuinely enforce 272K (gpt-5.5,
  gpt-5.4-mini) get no variant.
- The -900k suffix is Hermes-side only: stripped before the model id hits
  the wire (main transport + auxiliary Responses adapter), and pricing
  aliases the variants onto the base entries.
- Docs: new opt-in section in context-compression-and-caching.md.
2026-08-23 02:14:35 -07:00
Kshitij Kapoor 72ae855da5 refactor: consolidate _db_persisted literal + marker-insensitive no-op check (#92231 follow-up)
Review-pass follow-up on the load-time durability stamp:

- hermes_state.py: import the marker from agent.context_compressor instead
  of a third synced literal (hermes_state already imports agent.* at module
  level; only run_agent is circular). Old comment claimed otherwise.
- agent/turn_finalizer.py: replace the raw "_db_persisted" string at the
  fill-empty-tail pop site with the shared constant (was outside the drift
  guard).
- agent/conversation_compression.py: the no-op progress check now falls back
  to a marker-insensitive comparison (_strip_marker_for_comparison). Loaded
  rows are stamped at materialization time while compress() output is
  marker-swept, so a semantically-identical no-op copy on a cold-resumed
  session would previously compare unequal and take the progress branch.
  Raw == still runs first so engine-returned list subclasses keep their
  __eq__ semantics.
- test_marker_constant_in_sync extended to turn_finalizer + identity
  assertions; new test_noop_progress_check_is_marker_insensitive
  (mutation-checked: fails when the helper is neutered).
2026-08-23 13:34:32 +05:30
kshitijk4poor dff84f1890 fix(browser): cap browser_vision native embeds for history reuse
browser_vision's native fast path base64-encoded screenshots at full
resolution and baked them into the tool result uncapped — the exact
sibling of the vision_analyze path #92699 fixed. Apply the same
proactive 256KB/1568px resize before the embed enters reusable history.

Fail-open by design: without Pillow the resize helper falls back to raw
bytes and the compressor's keep-newest pass still retires stale embeds.

Sibling-gap follow-up for the #92725 salvage; the shared-cap approach
mirrors the policy-owner idea from #92748.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-08-23 13:27:04 +05:30
HexLab98 7ff2fe8bc9 fix(compression): retire stale vision tool images in the protected tail
Images locked in protect_last_n never shrank, so compression savings
stayed under 10% and anti-thrash disabled further compaction. Keep the
newest three tool-result screenshots live for follow-up QA and replace
older native embeds with placeholders.
2026-08-23 13:27:04 +05:30
Josh Holt 5f0a8f8739 fix(picker): harden keyless provider gate logging and credentials path validation 2026-08-22 23:25:43 -07:00