Commit Graph

229 Commits

Author SHA1 Message Date
Teknium 3e8977086f refactor(agent): compact Anthropic stream + retry-loop rationale comments/docstrings (AST-identical) 2026-09-02 13:29:36 -07:00
Teknium 9aee90335b refactor(agent): compact Bedrock and chat_completions stream-loop rationale comments (AST-identical) 2026-09-02 13:29:36 -07:00
Teknium d224e22507 refactor(agent): compact non-streaming call-path rationale comments/docstrings to their invariants (AST-identical) 2026-09-02 13:29:36 -07:00
Teknium 89a5ce12dd refactor(agent): build_api_kwargs — shared chat kwargs for profile/legacy paths; xAI tool_search alias rewrite helper 2026-09-02 13:29:36 -07:00
Teknium c5c6d7e566 refactor(agent): lift non-streaming stale/codex watchdog resolution into _resolve_nonstream_watchdogs 2026-09-02 13:29:36 -07:00
Teknium 3c80d4fe77 refactor(agent): compact build_assistant_message rationale comments to their invariants 2026-09-02 13:29:36 -07:00
Teknium 24b6b0e439 refactor(agent): compact streaming-path rationale comments to their invariants; drop no-op stale-kill if/else 2026-09-02 13:29:36 -07:00
Teknium 0e1875e859 refactor(agent): try_activate_fallback — lift api_mode hint/resolution and credential-pool rebinding into helpers 2026-09-02 13:29:36 -07:00
Teknium 3f497cf3f8 refactor(agent): interruptible_api_call — shared watchdog abort + post-kill worker wait helpers 2026-09-02 13:29:36 -07:00
Teknium 3e2ab5b382 refactor(agent): shared relay stream identity/metadata builders for the three streaming wires 2026-09-02 13:29:36 -07:00
Teknium ce9efccbaa refactor(agent): extract _ToolCallAccumulator from the chat_completions stream loop 2026-09-02 13:29:36 -07:00
Teknium 2d93a8b127 refactor(agent): share SSE connection-drop phrase check and codex silent-hang hint lookup 2026-09-02 13:29:36 -07:00
Teknium 3dc2403c45 refactor(agent): handle_max_iterations — single per-wire summary attempt with one retry pass; drop dead summary_request literal 2026-09-02 13:29:36 -07:00
Teknium e5e491ef5a refactor(agent): unify per-request client registry into _RequestClientRegistry (streaming + non-streaming) 2026-09-02 13:29:36 -07:00
Teknium ce276cd864 refactor(agent): decompose interruptible_streaming_api_call into _StreamingCall + per-wire branch functions
Closures -> _StreamingCall methods (shared worker/monitor state on the
instance), codex passthrough and Bedrock Converse branches -> free
functions, _stream_final_text/_emit_stream_* -> module helpers. Bodies
are AST-identical modulo the self.<name> renames (verified by script);
stream event handling and wire shapes unchanged.
2026-09-02 13:29:36 -07:00
Teknium c7e2e0b779 feat(fast): bounded /fast auto|cold windows behind one route-aware gate
Adds two bounded fast modes on top of the static /fast toggle, default OFF:

- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
  window; requests inside it carry the provider fast param, later tool-loop
  requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
  user/assistant/tool history).

agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.

resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.

Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.

Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes #64785, #74730.

Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
2026-09-02 05:33:13 -07:00
teknium1 c83ea9bed7 test(agent): pin the reasoning-off continuation to exactly one request; document its prompt-cache cost
The one-shot reasoning-off retry changes a request parameter that is part
of the provider cache key on config-sensitive providers (Anthropic renders
thinking/effort into the prompt; OpenAI lists reasoning.effort as
prefix-affecting), so that request is a deliberate single cache miss.
Pin the bound: the request AFTER it must carry the configured reasoning
again and the system prompt must be byte-identical across the whole retry
sequence. Sabotage-verified (sticky flag -> test fails on request 3).
Docstring on _consume_ephemeral_reasoning_off states the cost honestly.
2026-09-02 00:55:42 -07:00
AlexGabbia fb76fb0526 fix(agent): thinking-only length truncations no longer wedge continuations
GLM-5.3-flash on ollama-cloud with reasoning_effort=high can spend the ENTIRE
output cap on reasoning delivered in a separate field and return
finish_reason=length with no visible content (verified live: max_tokens=4096,
completion_tokens=4096, content empty).

The length-continuation path handled that shape badly:
  1. the empty response was appended as an interim assistant fragment,
     poisoning the transcript until the pre-call sanitizer healed it
     (observed 3+ healings per turn on the reporting user's session);
  2. every continuation re-ran with thinking ON, re-deriving the whole
     thinking budget against a growing context, so 4 attempts still produced
     nothing and the turn died with 'Response remains truncated after 4
     continuation attempts'.

Now:
  - interim assistant fragments with no visible content are never appended
    (whichever way they got empty);
  - a thinking-only truncation sets a one-shot reasoning-off override that
    build_api_kwargs consumes for the next request, so the continuation
    writes the answer instead of re-thinking it;
  - the ceiling exit clears a pending override and, when every fragment was
    empty, returns an actionable final_response instead of an invisible None.
2026-09-02 00:55:42 -07:00
AgentLinker cac9db7caf fix(codex): retired stream requests must not synthesize a completed response
When a watchdog (TTFB / stream-idle / stale-call) force-closes a Codex
Responses request, the worker thread can still be draining SSE frames.
`_consume_codex_event_stream` returns `status=terminal_status`, which defaults
to `"completed"`, and its only truncation guard is
`if not saw_terminal and not output`. A mid-stream kill leaves
`saw_terminal=False` but `output`/text non-empty, so the partial text came back
as a `finish_reason=stop` response and got persisted as a finished assistant
turn — a long reply just stops mid-sentence with no error surfaced.

Observed as a long generation dying at `1. Create (6/6)` and never emitting its
end marker, with the truncated text already stored in state.db.

Fix: publish a per-request retirement token so the worker can tell it has been
retired.

- `agent/chat_completion_helpers.py`: `interruptible_api_call` installs
  `agent._active_codex_stream_request_token` before handing off to the worker
  (codex_responses only) and clears it at all four kill sites plus the worker's
  own `finally`. Retirement is cleared BEFORE `_close_request_client_once`,
  which can raise — every other call site wraps it in try/except, and a leaked
  token would let a later worker mistake itself for the owning attempt. The
  request-local `_codex_request_retired` mirror also swallows the transport
  error our own force-close causes, so the worker's local error cannot replace
  the watchdog's retryable TimeoutError (same split as `_request_cancelled`).
- `agent/codex_runtime.py`: `run_codex_stream` captures the token and raises
  `TimeoutError` from `interrupt_check` when it no longer owns the request —
  raising rather than breaking, because a break returns the partial `final`.
  The four stream callbacks also drop post-retirement frames so an abandoned
  attempt cannot stream tokens into the live turn's bubble (the gateway caches
  AIAgent instances per session).

`TimeoutError` is not an httpx / ConnectionError / RuntimeError subclass, so it
passes through the four `except` clauses around the consume call untouched.
No token installed (auxiliary callers such as `handle_max_iterations` drive
`_run_codex_stream` directly) means every check passes — behavior unchanged.

Tests: 5 new cases. Retirement raises instead of returning partial output;
post-retirement deltas stop reaching callbacks; the no-token path keeps its
existing terminal-frame tolerance; the watchdog installs and clears the token;
non-codex api_modes install nothing. A `_LazyCreateStream` helper is needed
because `_FakeCreateStream` materializes events in __init__, which would run
the retirement side effect before consumption starts.
2026-09-01 22:14:06 -07:00
Teknium c5b99a3ee5 fix(agent): delegated children and cron turns stream again — inline, no worker
should_use_direct_api_call() contexts (gateway cron turns #62151, delegate_task
children #60203) were short-circuited onto the NON-streaming wire because the
interrupt worker wedges inside their nested thread pools. That dropped every
liveness property streaming provides: edge proxies kill the silent POST
(z.ai HTTP 524 — three retries later the child dies as "max_iterations"), and
the non-stream stale watchdog cannot tell a reasoning model's thinking phase
from a hung provider, so children die at exactly stale_timeout (#100260).

Keep those contexts on interruptible_streaming_api_call. The request now runs
INLINE on the conversation thread (no worker → the deadlock class stays
closed) while the existing poll loop — 30s heartbeat, stale-stream detector,
cross-thread interrupt abort — moves onto a monitor thread that only ever
aborts sockets, never dispatches (same shape as direct_api_call's watchdog
timer). Interactive sessions are unchanged: worker + poll loop as before.

should_use_direct_api_call() itself is untouched; only what it routes to.

Live A/B (real SSE server, real AIAgent.run_conversation):
  before: subagent/cron wire stream=None, request on conversation thread
  after:  subagent/cron wire stream=True, request on conversation thread
          cli unchanged (stream=True, spawned worker)
  inline stale detector kills a one-chunk-then-silence stream at budget;
  AIAgent.interrupt() from another thread unwinds the inline stream in 0.6s.

Co-authored-by: Expri-commits <184641533+Expri-commits@users.noreply.github.com>
2026-09-01 21:42:19 -07:00
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
rainbowgits 622883bad7 fix(agent): accept marker-only finish_reason after stream supersession
A superseded writer was fencing the payload-empty terminal chunk, so
completed streams were mislabeled as mid-stream drops.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:27:06 -07:00
Teknium 93591eccb5 fix(proxy): harden SSE DONE tracker — spec multi-line data joins, truthy lastOne, EOF-write guard
Follow-ups on the salvaged cluster:
- sse_done.py: dispatch SSE events at blank-line boundaries and join
  consecutive data: lines per the SSE spec (a split JSON event no longer
  reads as two malformed fragments that disable synthesis)
- accept integer/string-truthy lastOne sentinels (1 / "true") in both the
  proxy tracker and the agent stream reader
- server.py: guard the [DONE] append against client hangup at EOF and
  widen the interrupt tuple with OSError
- contributor email mappings for loulanyue and jon-nielsen
2026-09-01 10:12:21 -07:00
Jon Nielsen d304422b3d fix(streaming): extract finish_reason/usage before content-shape continues
vLLM >= 0.1.dev20051 merges finish_reason into the final content chunk.
When the SSE-echo guard is engaged at that moment (GLM-family tokenizers
emit standalone ':' / ' id' tokens mid-prose), the guard's content-shape
continue paths swallow the terminal chunk and finish_reason is never
captured, so a complete stream is misclassified as a mid-stream drop and
retried.

Extract finish_reason/usage at the top of the chunk loop body, before any
content-shape continue; the late tail-side extraction becomes redundant.

Addresses a second root cause of #94614 (the consume-gate fence is
covered by #94625; usage-side classification by #91376).
2026-09-01 10:12:21 -07:00
loulanyue 66d42e0dba fix(stream): do not misclassify stream with final usage chunk as mid-stream drop (#91373)
When stream_options={'include_usage': True} is requested, OpenAI-compliant
providers (e.g. vLLM, OpenAI, DeepSeek) emit a final usage-only chunk with
empty choices (choices=[]) and no finish_reason.

If the preceding text chunks did not explicitly set finish_reason, the
check in _call_chat_completions evaluated _text_only_dropped_no_finish to
True and returned a partial-stream stub with finish_reason='length'. The
conversation loop then assumed the connection was cut off and injected a
spurious continuation nudge, causing the model to rewrite the full answer.

Require usage_obj is None in _text_only_dropped_no_finish so streams that
delivered valid usage metadata complete cleanly with finish_reason='stop'.
2026-09-01 10:12:21 -07:00
rainbowgits ce7f805869 fix(proxy): append SSE [DONE] when Nous streams omit the sentinel
Complete Portal streams can finish with finish_reason/lastOne and clean
EOF without data: [DONE], which strict OpenAI clients treat as truncation.
Normalize at the hermes proxy boundary after clean EOF only.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-01 10:12:21 -07:00
overtoneblue e17276c7b4 agent: forward first stream chunk timestamp to post_api_request hook
interruptible_streaming_api_call already records first_chunk_at in its
per-attempt stream diagnostics (agent.stream_diag) for failure telemetry,
but the value was dropped on the success path. Stash it on the agent at
stream completion and forward it as first_chunk_at in the existing
post_api_request plugin-hook payload, alongside started_at/ended_at.

Consumers (observability plugins, shell hooks) can now derive TTFB
(first_chunk_at - started_at) and true generation throughput
(output_tokens / (ended_at - first_chunk_at)) without any new
instrumentation in the hot path.

Backward compatible: existing hook subscribers ignore unknown kwargs.
2026-09-01 08:30:45 -07:00
CoolStar aaca343110 fix(agent): preserve Bedrock redacted reasoning replay 2026-09-01 08:30:26 -07:00
Teknium 7beb0de676 fix: follow-up for salvaged PR #71904 — source raw inbound id, strip persistence key from the wire
- Thread event.message_id (raw inbound id) as TurnContext.inbound_message_id
  instead of reusing event_message_id, which is the reply/thread anchor and
  can be the replied-to message on Slack/Mattermost/Buzz or None for
  Telegram topics.
- Add platform_message_id to the schema-foreign strip sets in
  ChatCompletionsTransport.convert_messages and the summary path so strict
  providers never see the persistence-only key.
2026-08-31 12:46:32 -07:00
Teknium b7ebe6456f fix(xai): request-local alias provenance + collision-safe wire aliasing
Hardens the two #95003 alias carriers per review feedback on #95019/#95011:

- _alias_reserved_tools / _rename_tool_search_bridge_for_xai now return the
  alias map THIS request emitted; the transport stashes it
  (_last_wire_aliases) and normalize_response reverses ONLY those aliases.
  A real user/plugin/MCP tool named hermes_tool_search is never silently
  dispatched as tool_search when no alias was sent.
- Collision safety: if a real tool already occupies the alias name, the
  bridge takes hermes_tool_search_2/_3 — no duplicate wire declarations.
- Legacy static reverse map retained only for normalize-only call sites
  that never built a request on the transport instance.
- chat_completion_helpers resets provenance per request so stale maps from
  a prior request can't leak into the next response's dispatch.

Refs #95003
2026-08-31 10:09:04 -07:00
liuhao1024 5e2f8b9865 fix(xai): alias the reserved tool_search bridge name on chat completions
xAI's chat-completions API reserves the function name tool_search for
its native server-side tool and rejects the whole request when the
client Tool Search bridge declares it (HTTP 400 'The function name
tool_search is reserved for the tool_search tool', #95003) — Grok
providers were unusable whenever the bridge assembled into the payload
(default tools.tool_search: auto). Mirror the web_search treatment in
transports/codex.py: rename the bridge's wire declaration to
hermes_tool_search for xAI targets (deep-copied first, #27907 lesson)
and map the alias back to tool_search in normalize_response so dispatch
is unchanged. Alias matches the Codex-side fix for the same class
(#83122).
2026-08-31 10:09:04 -07:00
Stephen Chin 5247a6f07f fix(compaction): clarify runtime capability state
Use a distinct runtime_capabilities field on agents, preserve compatibility with earlier snapshots, and resolve the canonical direct OpenAI endpoint when a cross-provider switch omits base_url. Keep ambiguous proxy routes fail-closed.
2026-08-30 05:16:10 -07:00
Stephen Chin 08c7879ca1 fix(compaction): preserve native capability across runtime switches
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
2026-08-30 05:16:10 -07:00
Teknium 3b3ad958d7 fix(runtime): key-scoped fallback extra_body re-resolution + request_overrides in switch_model snapshot
Follow-up hardening on the two cherry-picked contributor commits:

- try_activate_fallback: replace the blanket request_overrides.pop('extra_body')
  with KEY-SCOPED removal — only keys the OLD provider's custom_providers
  entry contributed (value unchanged since the init-time merge) are dropped.
  Caller/profile-provided extra_body keys survive the swap, matching the
  caller-over-provider precedence in agent_init._merge_custom_provider_extra_body.
  The fallback provider's own extra_body is then merged back in.
- switch_model: the live _primary_runtime snapshot it rebuilds now carries
  request_overrides, so a post-switch transport recovery or fallback restore
  reinstates the switched-to identity's overrides instead of dropping them.
- Tests: activation-level stale-key removal + caller-override preservation
  (test_provider_fallback.py), switch-then-recover / switch-then-restore
  (test_primary_runtime_restore.py).

Cache-safety: none of these paths mutate past context or rebuild the system
prompt — only outbound request kwargs change.

Fixes #75091
2026-08-29 19:13:12 -07:00
RelaxJonh 91d60d2f9e fix(fallback): re-resolve extra_body when activating fallback provider (#75091)
`try_activate_fallback()` re-resolved `reasoning_config` for the new
fallback provider (fix for #21256), but never re-resolved `extra_body`.
The primary provider's `extra_body` (e.g. `reasoning_effort: "none"`)
rode along onto the fallback provider, which is a different API that
may reject those fields.

Example: primary has `extra_body: {reasoning_effort: "none"}`, fallback
is OpenRouter. After failover, every request to OpenRouter carries both
the stray top-level `reasoning_effort` AND the nested `reasoning` object,
and OpenRouter rejects the pair:
  HTTP 400: "reasoning_effort" and "reasoning.effort" are both provided

The fallback is dead precisely when it is needed.

Fix: after swapping provider/model/base_url, clear the primary's
extra_body from request_overrides, then re-resolve from the fallback
provider's config using the existing _merge_custom_provider_extra_body
helper.  Same pattern as the reasoning_config re-resolution above.
2026-08-29 19:13:12 -07:00
Jakub Wolniewicz 23bae43cfa fix(agent): normalize list-shaped streaming content deltas 2026-08-29 12:45:43 +05:30
joaomarcos 88a78ecc96 fix(bedrock): recover from server-side cachePoint rejections per placement
Bedrock's cachePoint rules are per-model-family AND per-field. Amazon Nova
accepts a cachePoint block in `system` and `messages` but rejects it inside
`toolConfig.tools`, failing the whole request with

    ValidationException: Malformed input request: #/toolConfig/tools/18:
    extraneous key [cachePoint] is not permitted

so every tool-enabled Nova turn fails, with no retry path and no way for the
user to turn cache markers off (#97281).

The adapter decided placement from one static allowlist that answers only
"does this model cache at all", never "in which section". Any family whose
placement rules differ breaks 100% of turns until someone edits the table and
ships a release — the same maintenance trap the `_NON_TOOL_CALLING_PATTERNS`
comment already admits to ("if a model fails with a tool-related
ValidationException, add it here").

Make Bedrock's own verdict authoritative alongside the table: classify the
rejection by the JSON pointer AWS returns, drop the marker for that one
section, retry the request once, and remember the verdict for the rest of the
process so later turns are built clean. The other sections keep their cache
markers, so Nova still gets system/messages caching instead of losing prompt
caching wholesale. This mirrors the module's existing self-heal idiom
(`is_streaming_access_denied_error` → non-streaming `converse()`).

Applied at all four boto3 call sites: `call_converse`, `call_converse_stream`,
and both Bedrock dispatch sites in `chat_completion_helpers` (the streaming
one is the path in the report). A rejection with no marker to strip returns
None so the caller re-raises instead of looping.

Fixes #97281

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gcoy6nLTg5R6FHHhjcLZEC
2026-08-29 11:33:05 +05:30
Teknium 420e156bf3 refactor(agent): single owner for Responses route predicates
Follow-up to the #96217 salvage: the codex/xai/github route checks were
re-implemented inline at four sites (codex_responses_adapter helpers,
chat_completion_helpers kwargs build, _is_openai_codex_backend, the
run_agent silent-reject hint). Consolidate them into
classify_responses_route() / ResponsesRouteFlags in
codex_responses_adapter and migrate every site — backend-identity
predicate class (#22548/#70893/#59561/#72468).

Host checks use exact-host-or-subdomain semantics, never substring
matching.
2026-08-27 18:56:21 -07:00
Ailirag 5908c577f9 fix(fallback): surface provider transitions and primary recovery 2026-08-25 12:12:08 +05:30
honor2030 fe483de4d3 fix(agent): keep max-iteration warnings out of quiet stdout
Route the max-iterations diagnostic through logging when quiet_mode is active so automation wrappers keep stdout machine-readable.

Add a regression test covering quiet max-iteration summary handling.
2026-08-23 17:45:58 -07:00
Yingliang Zhang 73243b0d2e feat(config): per-provider reasoning_echo opt-in for custom providers
Add model.reasoning_echo (default false) and per-fallback-entry
reasoning_echo to preserve assistant reasoning_content when
replaying history to custom providers and OpenAI-compatible gateways
that proxy thinking-mode models (Kimi K3, GLM-5.2, DeepSeek, etc.)
but are not matched by the built-in host-based _REASONING_ECHO_RULES.

The flag is per-active-provider, not a global toggle:
- Primary: read from model.reasoning_echo at init and switch_model
- Fallback: set by try_activate_fallback from the fallback entry
- Restore: restore_primary_runtime copies the switch_model snapshot

Unlike PR #76019 global agent.reasoning_echo toggle, the
per-provider flag travels with the active provider — falling back to
a strict provider (Mistral, Groq, Cerebras) correctly strips
reasoning_content even when the primary had the flag enabled,
because the flag is False for the strict fallback.

Complements PR #27361 (dynamic detection) which fires after the first
API response; this PR covers turn-1 and history-replay-on-fresh-session
where dynamic detection has not fired yet.

Closes #76018
Refs: #27297, #27361, #76019

Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
2026-08-20 00:03:50 +05:30
Bryan Bednarski a51337ecee Merge main into fix-openai-sparse-response-objects
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-17 18:58:48 -07:00
fangliquanflq a6f405314f fix(agent): cover provider wait teardown paths 2026-08-16 01:55:00 -07:00
fangliquanflq 0fd059e745 fix(agent): preserve stalled-provider escalation 2026-08-16 01:55:00 -07:00
Tuck ca77157639 fix: create message_metadata module and update chat_completion_helpers 2026-08-15 01:04:19 -07:00
kshitij e3fab0437e refactor(cache): never-raising scope resolver shared by both call sites
/simplify-code finding: turn_context evaluated resolve_prompt_cache_scope()
inside set_runtime_main's argument list under the umbrella try/except — a
resolution failure would silently skip the ENTIRE runtime binding
(provider/model/base_url/api_key/session_id for all aux calls that turn),
not just the cache scope.

- prompt_cache_scope: add resolve_prompt_cache_scope_safe() (never raises,
  returns None on failure/empty).
- turn_context: resolve the scope into a local via the safe variant BEFORE
  the set_runtime_main call, so a failure can only lose the scope.
- chat_completion_helpers: _prompt_cache_scope_for_agent delegates to the
  shared safe variant (guarded import retained).
- tests: +1 (hostile-property agent -> None; normal/empty passthrough).
2026-08-15 11:09:56 +05:30
kshitij 96cdf19a0b refactor(cache): fold self-review findings on the rotation-scope fix
- prompt_cache_scope: memo key now includes DB presence (a lazily attached
  _session_db re-resolves instead of staying pinned to the physical id);
  _persist_disabled agents (background-review forks that never get a DB row)
  memoize the fallback instead of re-querying the lineage per API call;
  module docstring cross-references get_conversation_root and why the two
  lineage resolvers must not be deduplicated.
- chat_completion_helpers: hoist the triplicated
  _prompt_cache_scope_for_agent(agent) call to a single local above the
  OpenAI-wire dispatch (after the anthropic/bedrock early returns, which
  don't use prompt_cache_key).
- codex transport docstring: x-client-request-id mirrors the derived body
  key, not the raw scope id.
- turn_context comment: acknowledge the first-turn pre-persist fallback.
- tests: +2 (persist-disabled memoization; lazy DB attach re-resolution).
2026-08-15 11:09:56 +05:30
kshitij cee2446222 fix(cache): keep prompt_cache_key warm across compression session rotation
Legacy compaction mode (compression.in_place: false) rotates the physical
session_id mid-conversation. The prompt-cache scope introduced in #79161 was
derived from that physical id, so every rotation moved the same conversation
into a fresh cache bucket - the prompt cache went cold at every rotation
boundary (#79017).

Fix: resolve a rotation-stable logical scope - the compression-lineage ROOT
of the current session (SessionDB.get_compression_lineage, fork-aware
post-#79193) - once per turn, memoized per transcript segment, and prefer it
over the physical session_id at every prompt_cache_key derivation site:

- agent/prompt_cache_scope.py (new): resolve_prompt_cache_scope(agent) -
  lineage-root walk with per-segment memo; falls back to the physical id
  when no DB is attached or the walk fails, degrading to pre-fix behavior.
- transports/codex.py: build_kwargs accepts cache_scope_id and prefers it
  for the body prompt_cache_key, the xAI x-grok-conv-id header, and the
  Codex x-client-request-id routing header. The Codex session_id header
  keeps the raw physical id (transcript identity, #57012 contract).
- transports/chat_completions.py: _add_prompt_cache_key accepts
  cache_scope_id with the same precedence.
- chat_completion_helpers.py: build_api_kwargs threads the resolved scope
  into all three build_kwargs call sites (codex, profile, legacy).
- auxiliary_client.py: set_runtime_main carries cache_scope; the aux
  Responses cache-key site prefers it over the physical session_id.
- turn_context.py: resolves the scope once per turn and threads it through
  set_runtime_main (no DB walk on the per-API-call hot path).

Scope semantics preserved from #79161: /new starts a fresh scope (new
lineage), /branch children, delegate subagents, and tool children stay
isolated (explicit-fork exclusion in get_compression_lineage), unrelated
sessions keep distinct buckets, and cron per-fire timestamps still
normalize via _cache_scope_from_session_id.

Default installs compact in place (session_id never rotates), so they hit
the memo and produce byte-identical keys to before.

Fixes #79017
2026-08-15 11:09:56 +05:30
Teknium 0cbc4ce83b fix(streaming): gate the interrupt worker join on live Relay managed execution
The unconditional 2s join before InterruptedError delayed interrupt
detection when Relay managed execution was not active (CI:
tests/run_agent/test_interrupt_propagation.py — detection took 2.34s
against a <1.0s budget, because the mocked worker sleeps 5s and there
is no Relay scope to unwind).

Extract the join into _join_worker_for_relay_teardown(), which no-ops
unless a Relay runtime exists AND managed execution consumers are
registered — the only case where an orphaned physical scope can corrupt
the LIFO stack (#81521). Applied at all three interrupt sites
(streaming, non-streaming, Bedrock streaming). The regression test now
simulates a live runtime so the join path stays covered.
2026-08-14 22:11:30 -07:00
Teknium 3537ef9d01 fix(streaming): widen #81521 interrupt join to sibling paths and use version-correct Relay top accessor
Follow-up to HexLab98's salvaged commits:

- Apply the same bounded worker join before raising InterruptedError at
  the two sibling interrupt sites that share the raise-without-join
  shape: the non-streaming API poll loop and the Bedrock streaming poll
  loop. Both workers run Relay-managed physical attempts, so raising
  immediately allowed turn teardown to race a still-open physical scope
  exactly as in the streaming path.

- Address the #81601 review finding (egilewski): the pinned nemo-relay
  binding's get_scope_stack() returns a native ScopeStack object which
  scope.pop rejects with TypeError, so the orphan drain never drained
  under the real binding. current_top() now prefers the version-correct
  scope.get_handle() accessor and falls back to the old list-unwrap for
  fake/legacy shapes. Handle comparisons go through same_handle(),
  comparing by uuid, because native ScopeHandle instances do not
  implement value equality.

- Add a real-binding regression test that reproduces the orphaned-scope
  session close against the pinned native wheel (skips where the native
  binding is unavailable), alongside the existing fake-based coverage.
2026-08-14 22:11:30 -07:00