Closures -> _StreamingCall methods (shared worker/monitor state on the
instance), codex passthrough and Bedrock Converse branches -> free
functions, _stream_final_text/_emit_stream_* -> module helpers. Bodies
are AST-identical modulo the self.<name> renames (verified by script);
stream event handling and wire shapes unchanged.
Adds two bounded fast modes on top of the static /fast toggle, default OFF:
- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
window; requests inside it carry the provider fast param, later tool-loop
requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
user/assistant/tool history).
agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.
resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.
Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.
Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes#64785, #74730.
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
The one-shot reasoning-off retry changes a request parameter that is part
of the provider cache key on config-sensitive providers (Anthropic renders
thinking/effort into the prompt; OpenAI lists reasoning.effort as
prefix-affecting), so that request is a deliberate single cache miss.
Pin the bound: the request AFTER it must carry the configured reasoning
again and the system prompt must be byte-identical across the whole retry
sequence. Sabotage-verified (sticky flag -> test fails on request 3).
Docstring on _consume_ephemeral_reasoning_off states the cost honestly.
GLM-5.3-flash on ollama-cloud with reasoning_effort=high can spend the ENTIRE
output cap on reasoning delivered in a separate field and return
finish_reason=length with no visible content (verified live: max_tokens=4096,
completion_tokens=4096, content empty).
The length-continuation path handled that shape badly:
1. the empty response was appended as an interim assistant fragment,
poisoning the transcript until the pre-call sanitizer healed it
(observed 3+ healings per turn on the reporting user's session);
2. every continuation re-ran with thinking ON, re-deriving the whole
thinking budget against a growing context, so 4 attempts still produced
nothing and the turn died with 'Response remains truncated after 4
continuation attempts'.
Now:
- interim assistant fragments with no visible content are never appended
(whichever way they got empty);
- a thinking-only truncation sets a one-shot reasoning-off override that
build_api_kwargs consumes for the next request, so the continuation
writes the answer instead of re-thinking it;
- the ceiling exit clears a pending override and, when every fragment was
empty, returns an actionable final_response instead of an invisible None.
When a watchdog (TTFB / stream-idle / stale-call) force-closes a Codex
Responses request, the worker thread can still be draining SSE frames.
`_consume_codex_event_stream` returns `status=terminal_status`, which defaults
to `"completed"`, and its only truncation guard is
`if not saw_terminal and not output`. A mid-stream kill leaves
`saw_terminal=False` but `output`/text non-empty, so the partial text came back
as a `finish_reason=stop` response and got persisted as a finished assistant
turn — a long reply just stops mid-sentence with no error surfaced.
Observed as a long generation dying at `1. Create (6/6)` and never emitting its
end marker, with the truncated text already stored in state.db.
Fix: publish a per-request retirement token so the worker can tell it has been
retired.
- `agent/chat_completion_helpers.py`: `interruptible_api_call` installs
`agent._active_codex_stream_request_token` before handing off to the worker
(codex_responses only) and clears it at all four kill sites plus the worker's
own `finally`. Retirement is cleared BEFORE `_close_request_client_once`,
which can raise — every other call site wraps it in try/except, and a leaked
token would let a later worker mistake itself for the owning attempt. The
request-local `_codex_request_retired` mirror also swallows the transport
error our own force-close causes, so the worker's local error cannot replace
the watchdog's retryable TimeoutError (same split as `_request_cancelled`).
- `agent/codex_runtime.py`: `run_codex_stream` captures the token and raises
`TimeoutError` from `interrupt_check` when it no longer owns the request —
raising rather than breaking, because a break returns the partial `final`.
The four stream callbacks also drop post-retirement frames so an abandoned
attempt cannot stream tokens into the live turn's bubble (the gateway caches
AIAgent instances per session).
`TimeoutError` is not an httpx / ConnectionError / RuntimeError subclass, so it
passes through the four `except` clauses around the consume call untouched.
No token installed (auxiliary callers such as `handle_max_iterations` drive
`_run_codex_stream` directly) means every check passes — behavior unchanged.
Tests: 5 new cases. Retirement raises instead of returning partial output;
post-retirement deltas stop reaching callbacks; the no-token path keeps its
existing terminal-frame tolerance; the watchdog installs and clears the token;
non-codex api_modes install nothing. A `_LazyCreateStream` helper is needed
because `_FakeCreateStream` materializes events in __init__, which would run
the retirement side effect before consumption starts.
should_use_direct_api_call() contexts (gateway cron turns #62151, delegate_task
children #60203) were short-circuited onto the NON-streaming wire because the
interrupt worker wedges inside their nested thread pools. That dropped every
liveness property streaming provides: edge proxies kill the silent POST
(z.ai HTTP 524 — three retries later the child dies as "max_iterations"), and
the non-stream stale watchdog cannot tell a reasoning model's thinking phase
from a hung provider, so children die at exactly stale_timeout (#100260).
Keep those contexts on interruptible_streaming_api_call. The request now runs
INLINE on the conversation thread (no worker → the deadlock class stays
closed) while the existing poll loop — 30s heartbeat, stale-stream detector,
cross-thread interrupt abort — moves onto a monitor thread that only ever
aborts sockets, never dispatches (same shape as direct_api_call's watchdog
timer). Interactive sessions are unchanged: worker + poll loop as before.
should_use_direct_api_call() itself is untouched; only what it routes to.
Live A/B (real SSE server, real AIAgent.run_conversation):
before: subagent/cron wire stream=None, request on conversation thread
after: subagent/cron wire stream=True, request on conversation thread
cli unchanged (stream=True, spawned worker)
inline stale detector kills a one-chunk-then-silence stream at budget;
AIAgent.interrupt() from another thread unwinds the inline stream in 0.6s.
Co-authored-by: Expri-commits <184641533+Expri-commits@users.noreply.github.com>
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.
Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
by context window
- derived recommendation: quality-ranked picks gated by a predicted
decode-speed floor, bandwidth-aware on unified memory; the decision
table is pinned as a test (pick AND reason per memory class), and the
Recommended badge explains its pick in a tooltip fed by the resolver's
actual branch
- engine install + model download with resumable split parts, cumulative
plan-level progress, and staged-model integrity (a split GGUF counts
only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
progress relayed over SSE, abandoned-request cleanup
Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
engine, download the recommended model, boot) plus per-model download/
activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
send instead of wedging the session
Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
A superseded writer was fencing the payload-empty terminal chunk, so
completed streams were mislabeled as mid-stream drops.
Co-authored-by: Cursor <cursoragent@cursor.com>
Follow-ups on the salvaged cluster:
- sse_done.py: dispatch SSE events at blank-line boundaries and join
consecutive data: lines per the SSE spec (a split JSON event no longer
reads as two malformed fragments that disable synthesis)
- accept integer/string-truthy lastOne sentinels (1 / "true") in both the
proxy tracker and the agent stream reader
- server.py: guard the [DONE] append against client hangup at EOF and
widen the interrupt tuple with OSError
- contributor email mappings for loulanyue and jon-nielsen
vLLM >= 0.1.dev20051 merges finish_reason into the final content chunk.
When the SSE-echo guard is engaged at that moment (GLM-family tokenizers
emit standalone ':' / ' id' tokens mid-prose), the guard's content-shape
continue paths swallow the terminal chunk and finish_reason is never
captured, so a complete stream is misclassified as a mid-stream drop and
retried.
Extract finish_reason/usage at the top of the chunk loop body, before any
content-shape continue; the late tail-side extraction becomes redundant.
Addresses a second root cause of #94614 (the consume-gate fence is
covered by #94625; usage-side classification by #91376).
When stream_options={'include_usage': True} is requested, OpenAI-compliant
providers (e.g. vLLM, OpenAI, DeepSeek) emit a final usage-only chunk with
empty choices (choices=[]) and no finish_reason.
If the preceding text chunks did not explicitly set finish_reason, the
check in _call_chat_completions evaluated _text_only_dropped_no_finish to
True and returned a partial-stream stub with finish_reason='length'. The
conversation loop then assumed the connection was cut off and injected a
spurious continuation nudge, causing the model to rewrite the full answer.
Require usage_obj is None in _text_only_dropped_no_finish so streams that
delivered valid usage metadata complete cleanly with finish_reason='stop'.
Complete Portal streams can finish with finish_reason/lastOne and clean
EOF without data: [DONE], which strict OpenAI clients treat as truncation.
Normalize at the hermes proxy boundary after clean EOF only.
Co-authored-by: Cursor <cursoragent@cursor.com>
interruptible_streaming_api_call already records first_chunk_at in its
per-attempt stream diagnostics (agent.stream_diag) for failure telemetry,
but the value was dropped on the success path. Stash it on the agent at
stream completion and forward it as first_chunk_at in the existing
post_api_request plugin-hook payload, alongside started_at/ended_at.
Consumers (observability plugins, shell hooks) can now derive TTFB
(first_chunk_at - started_at) and true generation throughput
(output_tokens / (ended_at - first_chunk_at)) without any new
instrumentation in the hot path.
Backward compatible: existing hook subscribers ignore unknown kwargs.
- Thread event.message_id (raw inbound id) as TurnContext.inbound_message_id
instead of reusing event_message_id, which is the reply/thread anchor and
can be the replied-to message on Slack/Mattermost/Buzz or None for
Telegram topics.
- Add platform_message_id to the schema-foreign strip sets in
ChatCompletionsTransport.convert_messages and the summary path so strict
providers never see the persistence-only key.
Hardens the two #95003 alias carriers per review feedback on #95019/#95011:
- _alias_reserved_tools / _rename_tool_search_bridge_for_xai now return the
alias map THIS request emitted; the transport stashes it
(_last_wire_aliases) and normalize_response reverses ONLY those aliases.
A real user/plugin/MCP tool named hermes_tool_search is never silently
dispatched as tool_search when no alias was sent.
- Collision safety: if a real tool already occupies the alias name, the
bridge takes hermes_tool_search_2/_3 — no duplicate wire declarations.
- Legacy static reverse map retained only for normalize-only call sites
that never built a request on the transport instance.
- chat_completion_helpers resets provenance per request so stale maps from
a prior request can't leak into the next response's dispatch.
Refs #95003
xAI's chat-completions API reserves the function name tool_search for
its native server-side tool and rejects the whole request when the
client Tool Search bridge declares it (HTTP 400 'The function name
tool_search is reserved for the tool_search tool', #95003) — Grok
providers were unusable whenever the bridge assembled into the payload
(default tools.tool_search: auto). Mirror the web_search treatment in
transports/codex.py: rename the bridge's wire declaration to
hermes_tool_search for xAI targets (deep-copied first, #27907 lesson)
and map the alias back to tool_search in normalize_response so dispatch
is unchanged. Alias matches the Codex-side fix for the same class
(#83122).
Use a distinct runtime_capabilities field on agents, preserve compatibility with earlier snapshots, and resolve the canonical direct OpenAI endpoint when a cross-provider switch omits base_url. Keep ambiguous proxy routes fail-closed.
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
Follow-up hardening on the two cherry-picked contributor commits:
- try_activate_fallback: replace the blanket request_overrides.pop('extra_body')
with KEY-SCOPED removal — only keys the OLD provider's custom_providers
entry contributed (value unchanged since the init-time merge) are dropped.
Caller/profile-provided extra_body keys survive the swap, matching the
caller-over-provider precedence in agent_init._merge_custom_provider_extra_body.
The fallback provider's own extra_body is then merged back in.
- switch_model: the live _primary_runtime snapshot it rebuilds now carries
request_overrides, so a post-switch transport recovery or fallback restore
reinstates the switched-to identity's overrides instead of dropping them.
- Tests: activation-level stale-key removal + caller-override preservation
(test_provider_fallback.py), switch-then-recover / switch-then-restore
(test_primary_runtime_restore.py).
Cache-safety: none of these paths mutate past context or rebuild the system
prompt — only outbound request kwargs change.
Fixes#75091
`try_activate_fallback()` re-resolved `reasoning_config` for the new
fallback provider (fix for #21256), but never re-resolved `extra_body`.
The primary provider's `extra_body` (e.g. `reasoning_effort: "none"`)
rode along onto the fallback provider, which is a different API that
may reject those fields.
Example: primary has `extra_body: {reasoning_effort: "none"}`, fallback
is OpenRouter. After failover, every request to OpenRouter carries both
the stray top-level `reasoning_effort` AND the nested `reasoning` object,
and OpenRouter rejects the pair:
HTTP 400: "reasoning_effort" and "reasoning.effort" are both provided
The fallback is dead precisely when it is needed.
Fix: after swapping provider/model/base_url, clear the primary's
extra_body from request_overrides, then re-resolve from the fallback
provider's config using the existing _merge_custom_provider_extra_body
helper. Same pattern as the reasoning_config re-resolution above.
Bedrock's cachePoint rules are per-model-family AND per-field. Amazon Nova
accepts a cachePoint block in `system` and `messages` but rejects it inside
`toolConfig.tools`, failing the whole request with
ValidationException: Malformed input request: #/toolConfig/tools/18:
extraneous key [cachePoint] is not permitted
so every tool-enabled Nova turn fails, with no retry path and no way for the
user to turn cache markers off (#97281).
The adapter decided placement from one static allowlist that answers only
"does this model cache at all", never "in which section". Any family whose
placement rules differ breaks 100% of turns until someone edits the table and
ships a release — the same maintenance trap the `_NON_TOOL_CALLING_PATTERNS`
comment already admits to ("if a model fails with a tool-related
ValidationException, add it here").
Make Bedrock's own verdict authoritative alongside the table: classify the
rejection by the JSON pointer AWS returns, drop the marker for that one
section, retry the request once, and remember the verdict for the rest of the
process so later turns are built clean. The other sections keep their cache
markers, so Nova still gets system/messages caching instead of losing prompt
caching wholesale. This mirrors the module's existing self-heal idiom
(`is_streaming_access_denied_error` → non-streaming `converse()`).
Applied at all four boto3 call sites: `call_converse`, `call_converse_stream`,
and both Bedrock dispatch sites in `chat_completion_helpers` (the streaming
one is the path in the report). A rejection with no marker to strip returns
None so the caller re-raises instead of looping.
Fixes#97281
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gcoy6nLTg5R6FHHhjcLZEC
Follow-up to the #96217 salvage: the codex/xai/github route checks were
re-implemented inline at four sites (codex_responses_adapter helpers,
chat_completion_helpers kwargs build, _is_openai_codex_backend, the
run_agent silent-reject hint). Consolidate them into
classify_responses_route() / ResponsesRouteFlags in
codex_responses_adapter and migrate every site — backend-identity
predicate class (#22548/#70893/#59561/#72468).
Host checks use exact-host-or-subdomain semantics, never substring
matching.
Route the max-iterations diagnostic through logging when quiet_mode is active so automation wrappers keep stdout machine-readable.
Add a regression test covering quiet max-iteration summary handling.
Add model.reasoning_echo (default false) and per-fallback-entry
reasoning_echo to preserve assistant reasoning_content when
replaying history to custom providers and OpenAI-compatible gateways
that proxy thinking-mode models (Kimi K3, GLM-5.2, DeepSeek, etc.)
but are not matched by the built-in host-based _REASONING_ECHO_RULES.
The flag is per-active-provider, not a global toggle:
- Primary: read from model.reasoning_echo at init and switch_model
- Fallback: set by try_activate_fallback from the fallback entry
- Restore: restore_primary_runtime copies the switch_model snapshot
Unlike PR #76019 global agent.reasoning_echo toggle, the
per-provider flag travels with the active provider — falling back to
a strict provider (Mistral, Groq, Cerebras) correctly strips
reasoning_content even when the primary had the flag enabled,
because the flag is False for the strict fallback.
Complements PR #27361 (dynamic detection) which fires after the first
API response; this PR covers turn-1 and history-replay-on-fresh-session
where dynamic detection has not fired yet.
Closes#76018
Refs: #27297, #27361, #76019
Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
/simplify-code finding: turn_context evaluated resolve_prompt_cache_scope()
inside set_runtime_main's argument list under the umbrella try/except — a
resolution failure would silently skip the ENTIRE runtime binding
(provider/model/base_url/api_key/session_id for all aux calls that turn),
not just the cache scope.
- prompt_cache_scope: add resolve_prompt_cache_scope_safe() (never raises,
returns None on failure/empty).
- turn_context: resolve the scope into a local via the safe variant BEFORE
the set_runtime_main call, so a failure can only lose the scope.
- chat_completion_helpers: _prompt_cache_scope_for_agent delegates to the
shared safe variant (guarded import retained).
- tests: +1 (hostile-property agent -> None; normal/empty passthrough).
- prompt_cache_scope: memo key now includes DB presence (a lazily attached
_session_db re-resolves instead of staying pinned to the physical id);
_persist_disabled agents (background-review forks that never get a DB row)
memoize the fallback instead of re-querying the lineage per API call;
module docstring cross-references get_conversation_root and why the two
lineage resolvers must not be deduplicated.
- chat_completion_helpers: hoist the triplicated
_prompt_cache_scope_for_agent(agent) call to a single local above the
OpenAI-wire dispatch (after the anthropic/bedrock early returns, which
don't use prompt_cache_key).
- codex transport docstring: x-client-request-id mirrors the derived body
key, not the raw scope id.
- turn_context comment: acknowledge the first-turn pre-persist fallback.
- tests: +2 (persist-disabled memoization; lazy DB attach re-resolution).
Legacy compaction mode (compression.in_place: false) rotates the physical
session_id mid-conversation. The prompt-cache scope introduced in #79161 was
derived from that physical id, so every rotation moved the same conversation
into a fresh cache bucket - the prompt cache went cold at every rotation
boundary (#79017).
Fix: resolve a rotation-stable logical scope - the compression-lineage ROOT
of the current session (SessionDB.get_compression_lineage, fork-aware
post-#79193) - once per turn, memoized per transcript segment, and prefer it
over the physical session_id at every prompt_cache_key derivation site:
- agent/prompt_cache_scope.py (new): resolve_prompt_cache_scope(agent) -
lineage-root walk with per-segment memo; falls back to the physical id
when no DB is attached or the walk fails, degrading to pre-fix behavior.
- transports/codex.py: build_kwargs accepts cache_scope_id and prefers it
for the body prompt_cache_key, the xAI x-grok-conv-id header, and the
Codex x-client-request-id routing header. The Codex session_id header
keeps the raw physical id (transcript identity, #57012 contract).
- transports/chat_completions.py: _add_prompt_cache_key accepts
cache_scope_id with the same precedence.
- chat_completion_helpers.py: build_api_kwargs threads the resolved scope
into all three build_kwargs call sites (codex, profile, legacy).
- auxiliary_client.py: set_runtime_main carries cache_scope; the aux
Responses cache-key site prefers it over the physical session_id.
- turn_context.py: resolves the scope once per turn and threads it through
set_runtime_main (no DB walk on the per-API-call hot path).
Scope semantics preserved from #79161: /new starts a fresh scope (new
lineage), /branch children, delegate subagents, and tool children stay
isolated (explicit-fork exclusion in get_compression_lineage), unrelated
sessions keep distinct buckets, and cron per-fire timestamps still
normalize via _cache_scope_from_session_id.
Default installs compact in place (session_id never rotates), so they hit
the memo and produce byte-identical keys to before.
Fixes#79017
The unconditional 2s join before InterruptedError delayed interrupt
detection when Relay managed execution was not active (CI:
tests/run_agent/test_interrupt_propagation.py — detection took 2.34s
against a <1.0s budget, because the mocked worker sleeps 5s and there
is no Relay scope to unwind).
Extract the join into _join_worker_for_relay_teardown(), which no-ops
unless a Relay runtime exists AND managed execution consumers are
registered — the only case where an orphaned physical scope can corrupt
the LIFO stack (#81521). Applied at all three interrupt sites
(streaming, non-streaming, Bedrock streaming). The regression test now
simulates a live runtime so the join path stays covered.
Follow-up to HexLab98's salvaged commits:
- Apply the same bounded worker join before raising InterruptedError at
the two sibling interrupt sites that share the raise-without-join
shape: the non-streaming API poll loop and the Bedrock streaming poll
loop. Both workers run Relay-managed physical attempts, so raising
immediately allowed turn teardown to race a still-open physical scope
exactly as in the streaming path.
- Address the #81601 review finding (egilewski): the pinned nemo-relay
binding's get_scope_stack() returns a native ScopeStack object which
scope.pop rejects with TypeError, so the orphan drain never drained
under the real binding. current_top() now prefers the version-correct
scope.get_handle() accessor and falls back to the old list-unwrap for
fake/legacy shapes. Handle comparisons go through same_handle(),
comparing by uuid, because native ScopeHandle instances do not
implement value equality.
- Add a real-binding regression test that reproduces the orphaned-scope
session close against the pinned native wheel (skips where the native
binding is unavailable), alongside the existing fake-based coverage.