Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).
On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
Zen endpoint routing incl. muse->responses); keyless predicate extended
with unsuffixed free slugs (big-pickle); free runtime pins EVERY
opencode-free model keyless; curated catalog replaces the models.dev
cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
invariant (every curated model must satisfy the keyless predicate)
E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.
The free tier requires a real account API key and throttles third-party
clients by User-Agent:
- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
unconditionally (the stale keyless-tier assumption), and credential-pool
exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
message.
Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
Builds on @Lesnak1's #85619 (issue #85589):
- New opencode_provider_family() single-owner predicate in
hermes_cli/models.py — resolves built-in AND custom family providers
(opencode-go-bridge, OpenCode-Zen-Custom, ...) case-insensitively.
Migrated all 8 inlined family checks (models.py x3, runtime_provider.py
x4 from the salvaged commits) plus 4 sibling sites the PR missed:
cli.py api_mode sync, agent_runtime_helpers.py double-/v1 guard,
model_normalize.py flat-namespace strip, model_switch.py base_url
normalization.
- Responses transport: alias OpenCode-reserved function names
(web_search, search_files -> hermes_*) on the wire and map them back on
dispatch — same pattern as the xAI web_search collision fix. Matches
family providers and any base_url on opencode.ai. Fixes the HTTP 400
'custom function name X is reserved' half of #85589.
- Tests: custom-provider routing assertions + 5 new transport alias tests.
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.
Four pieces, and they only make sense together:
· an in-page engine that inventories what is interactable and performs the
verb, injected as source because it has to run inside the guest page;
· the preview.act.request bridge from the gateway into the pane;
· drive_preview, for acting: elements, click, type, scroll, press, and the
pane's own back/forward/reload;
· annotate_preview, for marking without acting.
Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.
Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.
Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
Composio eval traces showed Hermes wasting turns re-issuing identical tool
calls (same tool, same args, same result — 3x/4x in one run) and ending
turns by announcing an action it never took. Two conservative, config-gated
guards (agent.stall_guards, default true):
- Identical-call loop breaker: ToolCallGuardrailController.observe_identical_call
tracks the consecutive streak of (tool, canonical args, result-hash); on
the 3rd identical call a compact one-line notice is appended to that tool
RESULT at construction time (cache-safe — tool results are append-only).
Never blocks the call. Pollers (process, *_get_result, *_poll) are exempt
via STALL_GUARD_REPEATABLE_TOOLS. Streak resets on any different call,
changed result, or new turn. Observed on the raw result before the
tool-loop warning suffix so its changing count can't defeat matching.
- Said-continue-but-stopped recovery: trailing_continue_intent() detects a
short reply ENDING on an announced next action ('Let me now…', 'I will
now…', 'Next, I…'); the conversation loop feeds it into the EXISTING
intent-ack continuation path (same interim-assistant + user-nudge
mechanism, same codex_ack_continuations cap of 2), preserving message
alternation — no parallel recovery machinery.
Config: agent.stall_guards in DEFAULT_CONFIG; docs in configuration.md;
unit tests for streak/allowlist/reset/gate and detector pos/neg cases.
The identical 6-line try/except block for reading model.reasoning_echo
from config appeared in both agent_init.py (init) and
agent_runtime_helpers.py (switch_model). Extracted into
AIAgent._read_reasoning_echo_from_config() static method — net -1 LOC.
Address review feedback on PR #76503:
1. Init-time primary snapshot (agent_init.py:2756) was missing
reasoning_echo_flag — after fallback recovery the flag was
restored as False even when model.reasoning_echo: true was set.
2. Switch transaction snapshot (agent_runtime_helpers.py:2284) was
missing _reasoning_echo_flag — a failed client rebuild during
switch_model would leave the old provider with the new provider
echo policy.
Both omissions now fixed. No test regressions (56 passed).
Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
Add model.reasoning_echo (default false) and per-fallback-entry
reasoning_echo to preserve assistant reasoning_content when
replaying history to custom providers and OpenAI-compatible gateways
that proxy thinking-mode models (Kimi K3, GLM-5.2, DeepSeek, etc.)
but are not matched by the built-in host-based _REASONING_ECHO_RULES.
The flag is per-active-provider, not a global toggle:
- Primary: read from model.reasoning_echo at init and switch_model
- Fallback: set by try_activate_fallback from the fallback entry
- Restore: restore_primary_runtime copies the switch_model snapshot
Unlike PR #76019 global agent.reasoning_echo toggle, the
per-provider flag travels with the active provider — falling back to
a strict provider (Mistral, Groq, Cerebras) correctly strips
reasoning_content even when the primary had the flag enabled,
because the flag is False for the strict fallback.
Complements PR #27361 (dynamic detection) which fires after the first
API response; this PR covers turn-1 and history-replay-on-fresh-session
where dynamic detection has not fired yet.
Closes#76018
Refs: #27297, #27361, #76019
Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
One generic tool in the desktop_ui toolset: discover what is on screen,
highlight an element with narration, or hand the user a paged tour. No tour
content lives in the code — the agent authors each one live, which is what
makes 'how does this work?' answerable as a walkthrough instead of a wall
of text.
Rides the existing blocking-prompt bridge (tour.request/.respond) like
read_preview, so it works on every connection topology.
One clarify.request carries the question list (qid, question, choices,
multi_select per entry). clarify.respond gains an optional question_id:
each respond locks one answer, a repeat respond overwrites it, and the
batch resolves when every question is locked. A respond without
question_id keeps its existing meaning (cancel the whole prompt).
Locked answers survive the deadline: a timed-out batch returns the
partial answer map with a timed_out flag instead of an empty string.
The reconnect replay snapshot also carries the locked answers, so a
reattached client restores its per-question state.
Both agent-side clarify dispatch sites forward the questions arg.
The #58327 dedup passes treat a repeated tool_call_id as garbage from a
retry/crash/resume glitch and drop it. That assumes tool_call_id is
globally unique, which it is not: llama.cpp emits a single constant id
for every tool call it ever returns (verified — three separate
completions from one server all carried the same id).
Under a seen-once-drop-forever rule, the SECOND legitimate tool result
of such a session looks like a duplicate and is deleted. From the second
tool call onward the model never sees any result: it announces its next
action, the turn ends, and the task is left unfinished. Bisected to
dba585c17 over a 2258-commit range; reproduced live on v0.19.0 (1/6 runs
completed a 4-step file task, vs 20/20 on the last release before that
commit, same model and server).
Key off OUTSTANDING calls instead of every id ever seen. Both original
protections are preserved: a replayed result still answers no pending
call and is still dropped, and duplicate tool_calls sharing an id within
one assistant message are still collapsed. A genuine new call that
reuses the id re-arms it first.
repair_message_sequence needs no change — it already resets its id set
per assistant message, so only the final pre-API pass mis-fires.
Live result after the fix: 8/8 runs complete, 17-26s each (was 1/6 with
runs hitting a 150s ceiling).
Self-review follow-up, caught by benchmarking the previous commit.
Widening the custom-provider capability-lookup gate to `is_anthropic_wire or
_is_litellm_route(...)` made EVERY chat_completions route with a litellm-ish
provider/host enter the lookup, including non-Claude models that the grant
branch below can never match. Measured on a route with no config.yaml
(the uncached worst case) that was ~7.5us -> ~1528us per evaluation.
Narrowed the gate to the exact condition the LiteLLM branch grants on
(chat_completions + Claude + litellm route), computed once into a local and
reused by the branch itself so the predicate no longer runs twice.
Measured with a realistic config.yaml present (mtime cache warm), vs
origin/main:
live-agent policy 20.6us -> 61.7us
destination planning 219.3us -> 347.7us
Sub-millisecond and scoped to the routes that actually opted in. The
earlier 1.5ms figures were a tempdir artifact: load_config_readonly's
mtime cache cannot engage when no config.yaml exists, which is never true
of a real install. Non-LiteLLM and non-Claude routes are unaffected
(openrouter Claude measured flat at ~7.9us).
Tests: 82 passed across the policy and TTL-propagation modules.
Self-review follow-up. The previous commit fixed substring matching on the
HOST but left the provider-id side as a bare substring, so a user-named
provider like `custom:notlitellm` or `mylitellmthing` still matched and was
handed Anthropic markers — the same bug class, half-fixed.
Both signals now match `litellm` as a whole delimited token via a shared
helper. Real spellings (`litellm`, `custom:litellm`, `litellm-router`, and
the already-lowercased `LiteLLM`) still match; lookalikes no longer do.
Tests: 71 passed. Adds lookalike-provider and real-spelling guards; both
new guards mutation-checked. Differential matrix over 2688 configs vs
origin/main: 60 changes, every one a Claude model on a genuine LiteLLM
route getting the envelope layout, zero pre-existing routes altered.
Follow-up to the salvaged LiteLLM cache grant. The grant itself is right;
four things about how it was scoped were not.
1. Layout. The branch returned the native inner-block layout
(use_native_layout=True) on api_mode == "chat_completions". That layout
writes a TOP-LEVEL msg["cache_control"] on role:tool and empty-content
messages and depends on the Anthropic adapter to relocate it into the
block — but that adapter only runs for api_mode == "anthropic_messages"
(agent/transports/anthropic.py registers there), and the
chat_completions transport does no relocation. Measured on a 3-tool-turn
transcript: 2 of the 4 available breakpoints landed on markers the
provider never sees. Worse, when LiteLLM itself relocates a top-level
marker for an OpenRouter-backed Claude route
(OpenrouterConfig._move_cache_control_to_content), the marker lands on
an empty assistant turn and produces a cache_control-marked empty text
block — the HTTP 400 "text content blocks must contain" shape already
guarded in agent/anthropic_adapter.py (#69512). Switched to the envelope
layout, matching every other OpenAI-wire grant in this function:
4 of 4 breakpoints honored, zero empty blocks.
2. Host matching. `"litellm" in base_url_hostname(...)` is the substring
false-positive class base_url_hostname's own docstring warns against; it
granted Anthropic markers to notlitellm.example.com,
foolitellmbar.example and friends. Replaced with a label-token match in
a named helper, so "litellm" must be a whole dot- or hyphen-delimited
token. All three of the original test hosts still match; a "litellm"
path segment on an unrelated host still does not.
3. Transport gate. `not is_anthropic_wire` also swept in codex_responses,
bedrock_converse and codex_app_server. Gated on
api_mode == "chat_completions" explicitly.
4. Operator override. The grant is inferred from a provider/host name, but
the custom-provider capability lookup was gated on is_anthropic_wire, so
an explicit `prompt_caching: false` for the route+model was honored on
/v1/messages and silently ignored on /v1/chat/completions. The lookup
now also runs for a LiteLLM route, and its layout follows the transport
rather than the declaration (an explicit `true` must not promote a
chat_completions request to the native layout).
Tests: 64 passed. Adds the wire-shape contract the original matrix was
missing (asserts no breakpoint sits on the message envelope, rather than
only checking the returned tuple), plus lookalike-host, other-transport,
and both operator-override directions. All five guards mutation-checked —
reverting each fix turns the corresponding test red.
anthropic_prompt_cache_policy() only granted Anthropic cache_control
markers to LiteLLM over the native Anthropic wire
(api_mode == "anthropic_messages"). A LiteLLM deployment exposing the
OpenAI-compatible surface instead (/v1/chat/completions, /v1/messages
-> 404) matched no grant branch and fell through to (False, False): no
cache_control injected, the system prompt sent as a plain string, and
the provider serving zero cache hits -- the entire prompt re-billed at
full price on every turn. Silent: no error, no warning, usage simply
shows 100% uncached input forever.
Add one branch after the is_anthropic_wire/is_claude case that grants
caching to Claude-family models on a LiteLLM endpoint regardless of
wire, with the native inner-block layout. Same failure class already
documented in-function for Qwen/DashScope.
Design:
- Gated on the Claude family only (is_claude); a Gemini/GPT/Qwen route
through the same proxy must not receive markers (they may reject the
cache_control block format -- cf. the DeepSeek/OpenCode exclusion).
- Matches on provider string OR base_url host, since provider naming
varies per install (litellm, custom:litellm, or a bare custom alias
pointed at a LiteLLM host).
- prompt_caching.cache_ttl: false still wins (the _cache_disabled early
return is untouched).
- Generic strict OpenAI-wire custom providers (e.g. Fireworks) remain
excluded -- verified by the existing over-reach regression test.
Tests: adds TestLiteLLMOpenAIWire covering the grant (several model
spellings x provider/host signals), no-over-reach (non-Claude on the
same proxy get nothing; operator disable wins), and adjacent behavior
(LiteLLM in Anthropic proxy mode still native layout). Full module:
43 passed.
Closes#84506. Original diagnosis, patch design, and measurements by
@ottosulin.
Adds a `modify` response type to pre_tool_call hooks so a hook can
transform tool arguments before the tool executes, instead of repairing
results afterwards via post_tool_call.
- hermes_cli/plugins.py: _dispatch_pre_tool_call_hooks() fires hooks once
and returns (block_message, modified_args); modify directives
shallow-merge into an accumulated dict built from the original args.
- agent/shell_hooks.py: _parse_response() accepts both the canonical
{"action": "modify", "args": {...}} and Claude Code-compatible
{"decision": "modify", "tool_input": {...}} wire formats.
- model_tools.py, agent/tool_executor.py, agent/agent_runtime_helpers.py:
dispatch sites migrated; modified args applied before execution.
- Docs + 10 new tests (merge semantics, precedence, block interplay).
Salvaged from PR #28953. Best fix for #18988.
Promote _split_model_config_default to hermes_cli/config.py as the single
shared helper for flattening dict-valued model.default/model.model config.
All 8 defense-in-depth sites now route through it instead of inlining
their own isinstance checks with inconsistent key orders.
Changes:
- Add split_model_config_default() to hermes_cli/config.py (public)
- cli.py: _split_model_config_default delegates to shared helper
- Fix key extraction order: agent_runtime_helpers.py was reversed
(default->model); now consistent (model->default) across all sites
- Remove provider-as-model-name fallback from main.py, oneshot.py,
model_tools.py, cli.py — provider is a routing key, not a model ID
- Add 'name' to _normalize_root_model_keys flattening loop and
_has_nested_default detection to cover the deprecated model.name alias
Tests: 118 passed + 1 skipped (cli_init, managed_scope, config).
E2E: 31/31 passed (config chokepoint, managed scope, crash site,
negative cases, edge cases).
A dict-valued model.default (e.g. {provider:..., model:...}) in config.yaml
was leaking into agent.model and crashing the agent at init:
AttributeError: 'dict' object has no attribute 'lower'
agent/agent_runtime_helpers.py: anthropic_prompt_cache_policy
This manifested on the Telegram gateway as an infinite reset loop: every
turn built an agent with model=dict, crashed during init, the gateway
treated the failed turn as a session needing reset, and /reset rebuilt the
agent and crashed again.
Coerce dict -> string at every model-resolution entry point so the value
is normalized once and never reaches a .lower() call as a dict:
- agent/agent_runtime_helpers.py: anthropic_prompt_cache_policy (the crash site)
- agent/agent_init.py: configured default model resolution
- cli.py: CLI config model + _normalize_model_for_provider
- hermes_cli/main.py: _has_any_provider_configured
- hermes_cli/oneshot.py: _run_agent model resolution
- hermes_cli/runtime_provider.py: _get_model_config default handling
- model_tools.py: _resolve_active_context_length
Review follow-up (egilewski): the previous commit only hedged the guidance
text; the exact Anthropic 400 was still classified, persisted, and surfaced
as confirmed billing exhaustion. Carry the ambiguity all the way through:
- agent/error_classifier.py: 'out of extra usage' matches on the 400 and
status-less paths now attach error_context {billing_unverified,
possible_content_filter}. Reason stays FailoverReason.billing (rotation +
fallback remain the right recovery either way); ClassifiedError grows a
billing_unverified property.
- agent/credential_pool.py: new FAILURE_REASON_BILLING_UNVERIFIED. An
unverified billing exhaustion gets the short transient cooldown instead of
the one-hour bench, regardless of pool size: a content-filter rejection
leaves the credential healthy and fails identically on every key, and the
hour-long sole-credential latch is what replayed the stored error and made
real fixes look ineffective. A true 402 keeps the full bench. The marker
persists with the entry so a restart cannot upgrade it back to a bench.
- agent/agent_runtime_helpers.py + run_agent.py: recover_with_credential_pool
threads billing_unverified and hands the pool 'billing_unverified' as the
persisted failure_reason.
- agent/conversation_loop.py: the fallback-switch status, max-retries status,
terminal label, and both structured terminal results hedge when the verdict
is unverified. New _billing_terminal_label + _billing_failure_result build
the returned terminal response in one place; the result dict now carries
billing_unverified and the billing_block gains 'unverified': true. The
confirmed-billing path (a real 402 or an API-key credit depletion) keeps
the original assertive wording, so the caveat no longer dilutes it.
Regression tests: classifier marking (400 + status-less + unambiguous-body
negative), pool cooldown TTLs + persistence round-trip, pool failure_reason
plumbing, and the returned terminal response for both unverified and
confirmed verdicts.
Note: tests/agent/test_credential_pool_routing.py::TestFailureAttribution::
test_unmatched_key_does_not_retry_only_pool_entry fails identically on
current main without this change (pre-existing, unrelated).
When agent.provider == "moa", the MoAClient facade *is* the client - there is
no real OpenAI wire endpoint behind the moa://local placeholder. Client rebuilds
(_replace_primary_openai_client: stream-retry pool cleanup, credential rotation,
dead-connection cleanup, fallback+restore) go through create_openai_client and
produce a native OpenAI client while provider stays "moa". The next primary
call then either raises a `_moa_prepared_request` TypeError (#78382) or, when
_client_kwargs carry an unrelated relay base_url, leaks the request to a foreign
gateway (observed as HTTP 503 "group ... no available channel" from an
unrelated new-api relay right after an aggregator empty-stream retry).
Fix: in create_openai_client, when provider is "moa", return
build_moa_facade(agent, model) instead of a native client. This covers every
rebuild entry point. The three already-fixed call sites
(restore_primary_runtime, try_recover_primary_transport, switch_model) assign
the facade directly and do not go through create_openai_client, so they are
unaffected.
Closes#78382
The dedup pass in sanitize_api_messages (introduced by #58327) can
produce an empty tool_calls array when all tool_call_ids in a message
are duplicates of earlier messages in a long conversation history.
DeepSeek v4 and newer OpenAI reject empty tool_calls with HTTP 400:
'Invalid messages[N].tool_calls: empty array'.
When kept_tcs is empty after dedup, drop the tool_calls key entirely
instead of writing tool_calls: [].
Fixes#64335
The keepalive httpx client uses read=None, and stranger-thread abort cannot close FDs, so a DeepSeek stall on the cron inline path waited until TCP died — hours past the 600s watchdog. Inject a per-call read timeout matching the stale budget, walk in-flight pool requests, and clear the socket timeout before shutdown without releasing the FD.
Port of the bug class from earendil-works/pi#7933 (DeepSeek base-URL
detection matched by raw substring, missing case variants and matching
lookalike URLs). Hermes had the same class at five sites:
- cli_agent_setup_mixin.py: keyless-custom-endpoint detection treated any
URL containing the OpenRouter host substring (path segment, lookalike
domain) as OpenRouter, and missed case variants of the real host.
- models.py validate_requested_model: same substring check for routing an
openrouter provider with a custom base_url to the custom catalog.
- runtime_provider.py: local-endpoint autodetect matched the string
localhost anywhere in the URL, including remote hostnames containing it.
- gateway/run.py: /status endpoint display, same local-host substring.
- agent_runtime_helpers.py: Nous Portal cache-layout detection matched
the nousresearch substring anywhere in the URL.
All sites now use the existing base_url_host_matches / base_url_hostname
helpers (exact host or subdomain, case-insensitive). Regression tests
proven to fail against the old predicates.
- /model switch now refreshes agent._custom_providers from the config
loaded during the switch before re-evaluating cache policy — a
prompt_caching flag added to config.yaml after session start was
invisible to a mid-session switch (policy read the stale init-time
snapshot while context_length resolution used the live list).
- Production-path test: real config.yaml in the modern providers: dict
shape through the real loader chain, exercising the init-order fallback
(no _custom_providers attr) for both the fable opt-in and the opus
explicit opt-out.
- Pin operator kill-switch precedence: _cache_disabled (prompt_caching.
cache_ttl falsy) beats an explicit per-model prompt_caching: true.
- Log (debug) instead of silently swallowing capability-lookup failures in
anthropic_prompt_cache_policy — a swallowed failure would otherwise
downgrade an explicit prompt_caching: true to (False, False) with zero
trace. Matches the sibling MoA branch's logger.debug style.
- Use load_config_readonly() for the None-fallback in
get_custom_provider_model_capability: the helper only reads, and the
fallback fires on the blank-stub paths (agent init before
_custom_providers is assigned, MoA/auxiliary destination planning), so
skip the ~135us defensive deepcopy per call.
- Add route-isolation regression tests at both levels (config helper +
agent policy): a prompt_caching declaration for one provider route must
never apply to another route with the same model name. Mutation-checked:
both tests fail when the URL match is disabled.
Simplify-pass follow-ups on the salvage stack (all guard tests re-run,
mutation-checked):
1. conversation_loop.py: moved `_preflight_compression_blocked = False`
from 9 per-site copies into the restart_with_rebuilt_messages handler
(its single consumer). Besides removing the 9 duplicated blocks, this
fixes a 10th pre-existing retry-loop site (content-filter stall
failover, #32421) that set the flag and broke WITHOUT clearing the
preflight block — a content-filter failover previously restarted with
preflight compression still blocked against the fallback's smaller
window, the same #84733 bug class. The outer-loop empty-response site
keeps its own clear (it never passes through the handler). New AST
guard test_restart_handler_clears_preflight_block pins the hoisted
clear (mutation-checked).
2. agent_runtime_helpers.py: extracted _raw_cache_ttl_from_config() —
prompt_caching_disabled_from_config and configured_cache_ttl were
verbatim copies of the same config read. Added VALID_CACHE_TTLS.
3. prompt_caching.py: added is_qwen_model() next to
ALIBABA_FAMILY_PROVIDERS; effective_cache_ttl and
anthropic_prompt_cache_policy now share both the family set and the
qwen predicate — neither can desync.
4. Guard-test hardening: assert every _try_activate_fallback reference
is a direct `if agent._try_activate_fallback(...):` site, so a future
`activated = ...` form can't silently escape the restart-discipline
guard.
Follow-ups on the salvaged #84782 (webtecnica):
1. conversation_loop.py: the empty-response fallback site sits directly
in the OUTER iteration loop, not the retry loop. The salvaged commit's
`break` there exited the conversation loop and ended the turn without
ever calling the just-activated fallback (caught by CI:
test_empty_response_triggers_fallback_provider). Restored `continue`
(which already re-runs the pre-API preflight at the top of the next
outer iteration) while keeping the `_preflight_compression_blocked`
reset. The other 9 sites are inside the retry loop, where `break` to
the restart_with_rebuilt_messages handler is correct.
2. test_prompt_cache_ttl_propagation.py: made the AST guard loop-aware —
retry-loop sites must break, outer-loop sites must continue (the old
assertion pinned the bug in (1)). Mutation-checked both directions.
3. test_failover_identity.py: added `model` to the SimpleNamespace agent
fixture — _redecorate_prompt_cache_for_provider now reads agent.model
for the per-destination TTL clamp (2 CI failures).
4. prompt_caching.py / agent_runtime_helpers.py: single source of truth
for the alibaba-family provider set — ALIBABA_FAMILY_PROVIDERS lives
in prompt_caching and anthropic_prompt_cache_policy imports it, so the
cache-policy opt-in and the TTL clamp can never desync.
5. auxiliary_client.py: threaded the configured tier into
_replan_synchronous_cache_sections via new configured_cache_ttl()
(no live agent on that path) — the aux half of #84733's report also
stopped regressing 1h to 5m. Guarded by
TestAuxFallbackReplanThreadsTtl (mutation-checked).
6. Dropped the redundant `or "5m"` at the two threaded call sites —
effective_cache_ttl already resolves None to "5m", and the `or`
masked the cache-disabled (None) semantics.
New desktop_ui tool: the agent proposes an MCP server (install/enable/
authorize + a one-line reason) and blocks on mcp.setup.request until the
renderer's consent card answers mcp.setup.respond with the outcome
(installed/enabled/authorized/declined/unanswered/error). Same lifecycle
as clarify: 10-min timeout, allow_expired late answers, tool lifecycle
events forced on so the card mounts even with tool progress off. Desktop
prompt hint steers the model to the tool instead of hand-editing config;
every other surface keeps the schema out and is pointed at hermes mcp
install.
Follow-up fixes on top of the salvaged #83678 commit:
1. Hoist the MiniMax-M3 marker exclusion ABOVE the native-Anthropic
early return. provider="anthropic" pointed at a MiniMax /anthropic
proxy is a supported override (_anthropic_base_url_override_ok), and
the is_native_anthropic branch matched on provider alone — returning
(True, True) before the M3 exclusion was reached. Two regression
tests pin the proxy route (M3 off, M2.7 still on).
2. Reuse the existing _model_name_suggests_minimax_m3() helper from
agent/model_metadata.py instead of a second inline substring copy.
3. Drop the debug kwarg on normalize_usage() — it had zero production
callers and duplicated standard logging level gating. The
cache-observability line is now a plain logger.debug scoped to
MiniMax providers on the Anthropic wire only, so the "+128 floor"
note can no longer appear for native Anthropic where it is false.
Tests updated accordingly (MiniMax logs, native Anthropic does not).
MiniMax-M3 ships server-side automatic prefix caching on the
Anthropic-compatible endpoint (content-keyed, no marker needed —
see platform.minimax.io/docs/api-reference/text-prompt-caching).
cache_control markers are NOT on its explicit-cache support list
(which covers only M2.7/M2.5/M2.1/M2).
Emitting markers on M3:
- wasted serialization overhead
- risked perturbing the server-side prefix hash
- gave users a false sense of explicit-cache savings (the
cache_read_input_tokens field carries a +128 constant floor
and cache_creation_input_tokens is always 0 for M3)
Also add an opt-in debug=True parameter to normalize_usage() that
emits a debug-level log line carrying the observable cache fields.
This is the only reliable cache signal for M3 — off by default,
debug-level, scoped to the anthropic_messages wire, so production
callers see no impact.
Pin both changes with 8 new tests:
- 4 M3 tests covering provider, host, and custom-provider paths
- 1 regression guard ensuring M2.x caching is unaffected
- 3 observability tests (off-by-default, on-with-M3, on-with-Claude)
Verified end-to-end against api.minimaxi.com/anthropic/v1/messages
with MiniMax-M3[1m]: identical system prompt hit-rate with and
without markers; cache_read field is unreliable (128 floor),
input_tokens drop (8467 -> 1) is the real hit signal.
Desktop-gated (desktop_ui toolset) metadata-only window awareness: the agent
can ask which application window sits directly behind the Hermes window
(app, title, bounds — never pixels). Rides the same blocking bridge as
read_terminal: the gateway emits window.read.request and the renderer
answers window.read.respond.
Review follow-up (W1): the pre-send transcript sanitizer
(agent_runtime_helpers.sanitize_tool_call_arguments) runs on the
PERSISTED messages list before every api_messages build and rewrites any
json.loads-failing argument string to "{}" in the transcript, prepending
a corruption marker to the paired tool result. That in-transcript repair
is deliberate (the stored turn must be replayable next call), but it
destroys the model's original bytes — for a truncated write_file call
those bytes are the user's streamed file content (#80498), and they
previously survived only as an 80-char log preview.
Until a sidecar-preservation design exists, make the bytes recoverable:
both destruction sites (the transcript sanitizer's WARNING and
_repair_tool_call_arguments' unrepairable-path WARNING) now log the full
original argument string bounded at 100KB instead of 80 chars. Corrupted
calls are rare; an oversized WARNING is a fair price for the only copy
of real user content.
* feat(agent): read_preview — the desktop-gated tool that reads the in-app browser
The agent could open the preview pane (open_preview) and read the embedded
terminal (read_terminal), but the browser it had just opened was a black box —
'what does this page say?' had no answer. read_preview mirrors read_terminal
end to end: HERMES_DESKTOP-gated via check_fn (zero schema footprint outside
the GUI), dispatched through the same agent callback pattern, windowed with
start/count so a long page pages instead of flooding context.
* feat(gateway): preview.read blocking bridge
Same lifecycle as terminal.read: the tool blocks on preview.read.request, the
renderer answers preview.read.respond (allow_expired — a slow page extraction
losing the 45s race must not surface a raw 4009), and a timeout emits
preview.read.expire so late answers resolve quietly.
* feat(desktop): the renderer serializes the active preview tab for the agent
preview-reader.ts is the preview analog of the terminal's buffer registry: the
URL pane registers a page reader (webview executeJavaScript → title + visible
innerText) keyed by tab id; readActivePreview resolves the ACTIVE tab, windows
the text (24k cap per read), and answers file/artifact tabs with identity plus
a note pointing at the tool that reads that content directly. The gateway
event handler answers preview.read.request beside terminal.read.request.
The sole-credential cooldown sized the bench from the raw HTTP status, but
403 is overloaded: error_classifier maps OpenRouter's "key limit exceeded"
and xAI's spending-limit block to FailoverReason.billing, while an edge
throttle with the same status is transient. Only 402 was excluded from the
short cooldown, so a spent account on a single key retried every 60 seconds
and re-failed forever.
Thread the classified reason from recover_with_credential_pool through
mark_exhausted_and_rotate to _exhausted_ttl. Billing keeps the full bench
regardless of status; everything else transient still recovers in 60s. The
verdict is stored on the entry (_EXTRA_KEYS, so it persists to auth.json) —
without that a restart would re-read a bare 403 and downgrade the bench.
Tests: sole billing-403 stays benched, survives reload, unclassified 403
still recovers; call-site coverage that the reason actually reaches the pool.
Three existing kwargs assertions updated for the new argument.
Replace the fixed 60-second cooldown with exponential backoff:
30min → 1h → 2h → 4h cap.
The counter is reset by restore_primary_runtime on successful
primary-provider recovery, so the backoff is strictly for
consecutive failures within a single degradation window.
Closes#29702
restore_primary_runtime retries the primary every turn once the 60s
transient cooldown clears. For subscription-window limits (Claude
Pro/Max 5h windows, Codex weekly caps) the reset is hours or days away,
so every retry is a guaranteed failure costing two provider switches
and two prompt-cache invalidations per turn.
Add CredentialPool.next_available_at() (earliest reset across exhausted
entries; None when available now or no reset info) and gate the restore
on it: skip while the primary's pool says nobody can serve, restore on
the first turn after the reset elapses. Fail-open: any gate error or
missing reset info falls through to the existing per-turn retry, so
recovery can never be later than today. Cross-provider fallbacks
consult the PRIMARY's pool (not the attached fallback pool), reusing
the loaded pool for the existing rebind to keep auth reads at one per
restore.
OpenCode Zen's relay rejects the Anthropic-style content block format
that cache markers produce (content becomes a block array instead of a
plain string), causing HTTP 400 with "content must be string, not block
array" for DeepSeek models.
Reverts the DeepSeek addition from commit 6b6435a874 while preserving
the Qwen/Alibaba caching path which continues to work.
Fixes#77217
Simplify-pass follow-up on the #69653 salvage: the original code built
these patterns from a name loop; hand-expanding them into 10 literals
lost that single source. Tag-name tuples restore it (adding a 6th
reasoning tag is now a one-place change), and the gnarly named-function
pattern regained a pointer to its step-1c rationale. Byte-equivalence
of every rebuilt pattern verified programmatically (alternation-order
neutrality probed: the \b and > anchors make order irrelevant).
strip_think_blocks passed the same response-scrubbing strings through
re's pattern dispatcher on every response. Skills Guard repeated the
same work for 121 patterns against every scanned line.
Compile the existing expressions once and reuse Pattern.sub/search. Keep
each generic tool-call tag in its own paired expression so mismatched
openers retain their payload while existing stray-closer cleanup remains
unchanged.
Part of #33208
Salvaged from #32713 by @ErnestHysa.
Co-authored-by: ErnestHysa <takis312@hotmail.com>
Follow-ups from review of #76113:
- Extract cache_ttl_means_disabled() as the single disable-synonym
predicate; agent_init and prompt_caching_disabled_from_config both use
it so the two detection sites can no longer drift (drift would recreate
the #76085 bug class).
- Mirror _run_reference's not-None injection guard in
aggregate_moa_context (stamping None was a harmless no-op copy).
- Replace a vacuous trailing test assertion with the intended
input-non-mutation check; drop a stray blank line.
- Add a predicate-parity regression test (unknown TTL values keep
caching enabled, matching historical agent_init semantics).
Absorb the useful deltas from the parallel #76121 approach: a single
blank_cache_policy_stub factory so _cache_disabled cannot be left off
hand-rolled SimpleNamespaces, and pin the live agent disable onto MoA
advisor fan-out and one-shot aggregate_moa_context decoration so those
paths track conversation state rather than a fresh config re-read.
Keeps the earlier tri-state prepared-aggregator no-agent fix. Adds
factory and synthesis/advisor regressions.
Coordinates with #76121 / #76085.
Co-authored-by: JoaoMarcos44 <87440198+JoaoMarcos44@users.noreply.github.com>
Blank SimpleNamespace stubs used by MoA decoration and
plan_cache_sections_for_destination never set _cache_disabled, so
anthropic_prompt_cache_policy re-injected cache_control markers after
operators turned caching off. Stamp the disable onto those stubs from
an explicit flag or the live config, and pass the agent flag from the
MoA aggregator path.
Fixes#76085
Three copies of the same logic landed with #76032:
- MoA's _call_prepared_aggregator and auxiliary_client's
_replan_synchronous_cache_sections both implemented stub → policy →
strip → plan for a resolved destination. Extract
plan_cache_sections_for_destination() into agent_runtime_helpers (which
already owns the policy functions) and route both through it. Also
removes a redundant full-transcript deepcopy+strip per request (the
caller pre-stripped what build_prompt_cache_plan strips again).
- The fallback_chain[N] label regex + chain-entry lookup lived in
_fallback_entry_timeout AND _fallback_destination. Extract
_fallback_chain_entry() and reuse.
MoA's cache-plan failure log is promoted debug → warning: the call-block
site skips MoA, so this block is the aggregator's only decoration path —
a silent failure ships an undecorated request (the 0%-cache MoA bug class).
Behavior-preserving; 195 targeted tests green.
Follow-up to #76032 (#20880).
Setting prompt_caching.cache_ttl to a falsy value (false, null, off,
disabled, no, none) now fully disables prompt caching instead of
being silently ignored.
The disable propagates through anthropic_prompt_cache_policy() (early
return when _cache_disabled flag is set) and restore_primary_runtime()
(override after snapshot restore), so it survives /model switches and
fallback re-derivation — the gap that caused #56105 to be reverted in
#56126.
Salvage of #33555 by @BB-light, with model-switch/fallback survival
gap fixed on top.
Co-authored-by: BB-light <BB-light@users.noreply.github.com>