Follow-up to the #99622 salvage:
- agent/transports/chat_completions.py: the legacy (no provider profile)
chat_completions path always re-emitted extra_body.reasoning with
enabled=True, so both reasoning_effort: none and the one-shot
length-continuation override went out as {enabled: true, effort: none}.
Honor enabled=False / effort=none the way the profile path does.
- agent/conversation_loop.py: reset agent._ephemeral_reasoning_off at
turn start so a flag armed by an interrupted/errored turn can never
strip thinking from the next turn's first request.
- User-facing hints now name the real slash command (/reasoning); the
/thinkon//thinkoff commands do not exist.
- tests: wire-level regression (continuation request carries
reasoning.enabled=false) and a stale-flag turn-scope test.
Anthropic subscription OAuth (claude_code credential) misroutes Hermes
sessions carrying the session_search or memory toolset into the
extra-usage lane, surfacing as HTTP 400 "You're out of extra usage" on
a valid subscription. Live-verified via the
anthropic-ratelimit-unified-representative-claim response header
(deterministic lane oracle, no dependency on the laggy usage counter):
tool schemas are innocent, the trigger is three specific system-prompt
sentences (session_search recall + two skill_manage sentences),
required jointly — breaking any one clears the classifier.
Two independent layers:
1. OAuth wire alias (anthropic_adapter.py, transports/anthropic.py):
session_search -> chat_history_lookup, memory -> context_notes in
tool name, description, and (session_search only) system-prompt
prose, with wire-collision guarding and a normalize_response
reverse-map that keeps GH-25255 registered-tool precedence. Also
routes named tool_choice through the same normalizer, closing a gap
where a forced tool_choice would leak the raw trigger string and
stop matching tools[].
2. Prompt-preserving reword (prompt_builder.py): rewords the two
triggering SKILLS_GUIDANCE sentences while keeping the same
meaning, still naming skill_manage, and leaving the Skill Safety
Rule section untouched. Applies to all auth paths since it's a
prompt-copy change, not a wire-level transform.
Two layers rather than one because the three-sentence AND-condition
means a classifier tightening could start firing on either remaining
leg alone.
Fixes#65365
- Thread event.message_id (raw inbound id) as TurnContext.inbound_message_id
instead of reusing event_message_id, which is the reply/thread anchor and
can be the replied-to message on Slack/Mattermost/Buzz or None for
Telegram topics.
- Add platform_message_id to the schema-foreign strip sets in
ChatCompletionsTransport.convert_messages and the summary path so strict
providers never see the persistence-only key.
Hardens the two #95003 alias carriers per review feedback on #95019/#95011:
- _alias_reserved_tools / _rename_tool_search_bridge_for_xai now return the
alias map THIS request emitted; the transport stashes it
(_last_wire_aliases) and normalize_response reverses ONLY those aliases.
A real user/plugin/MCP tool named hermes_tool_search is never silently
dispatched as tool_search when no alias was sent.
- Collision safety: if a real tool already occupies the alias name, the
bridge takes hermes_tool_search_2/_3 — no duplicate wire declarations.
- Legacy static reverse map retained only for normalize-only call sites
that never built a request on the transport instance.
- chat_completion_helpers resets provenance per request so stale maps from
a prior request can't leak into the next response's dispatch.
Refs #95003
xAI's chat-completions API reserves the function name tool_search for
its native server-side tool and rejects the whole request when the
client Tool Search bridge declares it (HTTP 400 'The function name
tool_search is reserved for the tool_search tool', #95003) — Grok
providers were unusable whenever the bridge assembled into the payload
(default tools.tool_search: auto). Mirror the web_search treatment in
transports/codex.py: rename the bridge's wire declaration to
hermes_tool_search for xAI targets (deep-copied first, #27907 lesson)
and map the alias back to tool_search in normalize_response so dispatch
is unchanged. Alias matches the Codex-side fix for the same class
(#83122).
xAI reserves the function name `tool_search` for Grok's native
server-side Tool Search and rejects the client declaration outright:
HTTP 400 {"code":"invalid-argument","error":"The function name
tool_search is reserved for the tool_search tool"}
Hermes's progressive-disclosure bridge registers exactly that literal
(`TOOL_SEARCH_NAME` in tools/tool_search.py) and assembly is not
provider gated, so with the default `tools.tool_search.enabled: auto`
every grok turn fails the moment the catalog crosses the threshold —
mid-session, which reads to the user as a session reset.
Same treatment as the two collisions already handled on this
transport (xAI `web_search` #48108, OpenCode reserved names #85589):
alias to `hermes_tool_search` on the wire in build_kwargs, map back in
normalize_response so Hermes dispatch and the bridge contract are
untouched. `tool_describe` / `tool_call` are not reserved by xAI and
are left alone.
Folds the per-provider rename helpers into one `_alias_reserved_tools`
owner parameterized by the reserved-name tuple, and extends the
existing `_RESERVED_ALIAS_TO_NAME` reverse map so the dispatch-side
un-aliasing needs no new branch.
Scope note: this covers the Responses transport, which is where every
api.x.ai route lands by default (`_fallback_api_mode` maps api.x.ai →
codex_responses, and the xai provider profile declares it). An xAI
model forced onto `api_mode: chat_completions` would still hit the
400; that path has no provider-specific tool rewriting today and would
need the symmetric hook in agent/transports/chat_completions.py. Happy
to add it here if you'd rather have both in one change.
Tests: new TestXaiReservedToolSearchAlias covering the wire alias,
non-xAI backends keeping the canonical name, composition with the
native web_search swap, and the normalize_response round trip.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012vLaAmnsdii3Gm9jMDs5gw
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
ring entry removed from keyless_mcp, legacy backend set / credential
ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
nous_subscription surfaces. The tvly- redaction pattern stays --
legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
environment-variables, tools-reference, web-dashboard, provider
plugin dev guide.
Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
Review findings on salvaged #93548: _warm_efforts_async now returns
early under PYTEST_CURRENT_TEST (matching the canonical OpenRouter caps
warmer) so a test that forgets to monkeypatch it can't fire live HTTP
when RAMP_ROUTER_API_KEY is set; the codex transport's fail-open
except in _profile_declared_efforts logs at debug instead of silently
swallowing profile-hook bugs.
Addresses the automated review on this PR:
- _profile_declared_efforts falls back from provider name to the
endpoint's host (via model_metadata's URL->provider map), so a named
custom provider pointed at api.router.com — which the host mandate
already routes onto this transport — gets the catalog clamp instead
of the default vocabulary and a Router 400.
- _parse_efforts validates catalog levels against EFFORT_LADDER at
ingest, logging and dropping unrecognized tiers; a model whose whole
vocabulary is unrecognized stays out of the map (transport defaults)
instead of passing the requested effort through unclamped.
- fetch_models dedupes ids while preserving Router's deliberate listing
order.
- plugin.yaml credits the human contributor per repo convention.
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:
- plugins/model-providers/router/: RouterProfile plugin —
api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
runtime_provider._detect_api_mode_for_url: api.router.com ->
codex_responses. The host is Responses-only — POST /v1/chat/completions
does not exist and 404s — so this is a genuine host mandate (exact
hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
hook (tri-state: None=defer, ()=model takes no reasoning params,
tuple=clamp set). Router validates reasoning.effort per model and
returns HTTP 400 invalid-argument on levels outside the model's
published vocabulary, and 400 unsupported_parameter when a
non-reasoning model receives any reasoning field (both verified live).
The profile answers from a cached copy of the catalog's
router.capabilities.reasoning block: cache-only on the hot path,
seeded for free by fetch_models(), disk-mirrored across processes
(/cache/router_catalog.json), background-warmed when cold
— same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
the generic effort-clamp branch (xai/actual/github branches untouched;
profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
rejection, profile registration + auth auto-registry wiring, catalog
parsing, and transport clamp/suppression/fallback paths.
Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.
Port from anomalyco/opencode#44571: OpenAI caps prompt_cache_key at 64
chars (DeepSeek and Zai inherit the same limit via their OpenAI-compatible
APIs) and rejects longer values with HTTP 400. The Responses transport
already bounds keys via _bounded_prompt_cache_key, but the Chat
Completions transport passed caller-supplied keys (request_overrides,
top-level or extra_body) through unmodified on both the profile and
legacy kwargs paths.
Bound caller keys with the same pck_<sha256[:24]> hash shape codex.py
uses so both transports behave identically; blank keys are dropped
instead of sent empty. Hermes-generated keys were already safe
(content-addressed pck_ hashes).
The reuse reviewer found a third copy of the fc_->call_ synthesis in
agent/transports/codex.py _pair_ids (item_id[3:] spelling, which is why
the len('fc_') grep missed it). All three sites now share
_canonical_call_id_from_fc(), keeping the pairing invariant in one place.
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the
live-verified 900K burned through subscription usage for users who never
asked for the larger window (bigger window = more input tokens per request).
- Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the
advertised 272K again — the cheaper limit is the default.
- The model picker synthesizes explicit <slug>-900k variants (e.g.
gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into
the 900K window. Slugs that genuinely enforce 272K (gpt-5.5,
gpt-5.4-mini) get no variant.
- The -900k suffix is Hermes-side only: stripped before the model id hits
the wire (main transport + auxiliary Responses adapter), and pricing
aliases the variants onto the base entries.
- Docs: new opt-in section in context-compression-and-caching.md.
Builds on @Lesnak1's #85619 (issue #85589):
- New opencode_provider_family() single-owner predicate in
hermes_cli/models.py — resolves built-in AND custom family providers
(opencode-go-bridge, OpenCode-Zen-Custom, ...) case-insensitively.
Migrated all 8 inlined family checks (models.py x3, runtime_provider.py
x4 from the salvaged commits) plus 4 sibling sites the PR missed:
cli.py api_mode sync, agent_runtime_helpers.py double-/v1 guard,
model_normalize.py flat-namespace strip, model_switch.py base_url
normalization.
- Responses transport: alias OpenCode-reserved function names
(web_search, search_files -> hermes_*) on the wire and map them back on
dispatch — same pattern as the xAI web_search collision fix. Matches
family providers and any base_url on opencode.ai. Fixes the HTTP 400
'custom function name X is reserved' half of #85589.
- Tests: custom-provider routing assertions + 5 new transport alias tests.
Live probes against api.openai.com/v1/responses (Aug 2026):
- gpt-5.6: accepts none/low/medium/high/xhigh/max; rejects minimal, ultra
- gpt-5.5: accepts none/low/medium/high/xhigh; rejects max ('Unsupported
value'), minimal, ultra
So #68365's premise was half right: 'max' does 400 — but only on pre-5.6
models; blanket-clamping max->xhigh on gpt-5.6 (its fix) would have capped
the one model that supports max. The declared-vocabulary design absorbs
this as data: codex_supported_efforts(model) picks CODEX_GPT56_EFFORTS or
CODEX_LEGACY_EFFORTS, and the shared clamp does the rest. Both the main
Codex transport and the auxiliary client's Responses path use it.
Wire outcomes: ultra -> max on gpt-5.6, ultra/max -> xhigh on gpt-5.5/o5,
minimal -> low everywhere.
The #89503/#70058/#74295/#87279 bug class kept regenerating because every
transport and provider profile hand-rolled its own effort translation map
(9 sites, 4 distinct policies). New agent/reasoning_effort.py is the single
source of truth:
- EFFORT_LADDER: canonical low->high ordering (superset check against
VALID_REASONING_EFFORTS pinned by test)
- clamp_effort(): one policy — supported passes verbatim, otherwise nearest
WEAKER supported level (never escalate, never invert the ladder), floor
when nothing weaker, 'none' never a degradation target, declared
vendor-documented overrides win, bespoke names pass through
- declared wire vocabularies as data: OpenAI-compat, Codex Responses,
xAI (4.6/legacy), Actual relays, Kimi K3/K2, TokenHub, GLM-5.2,
DeepSeek V4, Ollama Cloud, Meta, Solar
Converted sites (all behavior-preserving except noted):
- chat_completions chokepoint, Kimi + TokenHub paths
- codex transport (backend branches now pick a declared set)
- auxiliary_client Responses path
- hermes_cli.models clamp_reasoning_effort_to_supported -> thin wrapper
- plugins: kimi-coding, zai, opencode-zen, deepseek, ollama-cloud,
meta-ai, upstage, custom (copilot already routes via the wrapper)
Behavior fixes the shared policy surfaces:
- ollama-cloud/opencode-go 'minimal' now degrades to 'low' instead of
being dropped (drop left the server default = MORE thinking than asked)
New tests: ladder contract (every configurable level is clamped by every
declared wire set; monotonicity across the full ladder for every set).
The chat_completions chokepoint fix (ultra->max for every model,
cherry-picked from #89509) has siblings with the same bug shape:
- codex.py: ultra->max was gated on gpt-5.6 only; now baseline for all
Responses-API models (backend-specific branches still override).
- Kimi top-level reasoning_effort: K3 accepts low/high/max only —
'medium' and upper-ladder levels were dropped to the medium default
(400s on K3, ladder inversion on K2). Full ladder mapped per family,
mirroring the kimi-coding plugin's K3 map.
- TokenHub: 'minimal' fell through to the 'high' default (asked least,
got most); full ladder now mapped onto low/medium/high.
- auxiliary_client Responses path: ultra->max alongside the existing
minimal->low clamp.
- custom provider plugin: ultra capped at max instead of forwarded
verbatim to GLM/vLLM/SGLang backends that reject it.
- copilot plugin: ad-hoc downgrade rules replaced with the shared
clamp_reasoning_effort_to_supported ladder walk so ultra/max resolve
to the strongest supported level instead of medium (#74295).
Sabotage-verified: new sibling-site tests fail 6/10 without the fixes.
Hermes' internal effort vocabulary extends the wire set with ultra
(documented by /reasoning as none..xhigh|max|ultra). OpenAI-compatible
wires — OpenRouter chief among them — accept exactly
max|xhigh|high|medium|low|minimal|none and reject the extension with
HTTP 400, so an ultra configured while the default model was Anthropic
worked (the Anthropic adapter maps its own levels) but leaked
untranslated the moment a per-job override pinned a non-Anthropic
model, failing every call for that job.
The wire-compat chokepoint for this transport previously mapped
ultra to max only for gpt-5.6; generalize the cap to every model.
- hermes_cli/providers.host_mandated_api_mode: add exact-hostname clause for
api.meta.ai → codex_responses (measured 0% cache on /chat/completions vs
93-99% on /responses with retention); update docstring.
- hermes_cli/runtime_provider._detect_api_mode_for_url: mirror clause for
api.meta.ai (exact hostname, #32243) to keep runtime resolver in lockstep.
- agent/agent_init: call host_mandated_api_mode early in api_mode cascade
(after explicit api_mode wins, before provider-name specials) via lazy
import; single source of truth, preserves user override.
- agent/transports/codex._default_prompt_cache_retention_for_request: return
24h for api.meta.ai unconditionally; build_kwargs setdefault preserves
override; Bedrock branch untouched.
- cli-config.yaml.example: add commented providers.meta example (api_mode
auto-detected).
- website/docs/developer-guide/adding-providers.md: list Meta alongside
Codex/xAI as codex_responses native provider with retention note.
- tests: add hermetic behavior-contract suites for mandate, retention,
content-addressed prompt_cache_key, reasoning passthrough, AIAgent init,
usage cache reporting, model-switch override, and config roundtrip; extend
test_model_switch_openai_api_mode with meta cases.
mcp 2.0.0 implements MCP revision 2026-07-28 and makes three breaking
changes Hermes sits on top of: `mcp.server.fastmcp` is gone, every model
field is renamed to snake_case (camelCase survives only as a
serialization alias, which pydantic does not expose to attribute
access), and the SDK's own HTTP stack moved from `httpx` to `httpx2`.
Bump the pin across the dev/mcp/computer-use extras and port the tree:
- `mcp_serve.py` and `agent/transports/hermes_tools_mcp_server.py` move
from `FastMCP` to `mcp.server.MCPServer`, which has the same
decorator/add_tool surface. The hermes-tools server already
synthesised `__signature__` from Hermes' JSON Schema, which is exactly
what 2.0's `add_tool` reads.
- SDK model reads go through `mcp_field(obj, snake, camel)`, which reads
both spellings. A single-spelling read fails *silently* on the other
generation — empty tool schemas, dropped structured content, tool
results vanishing from sampling conversations — and `mcp` is an
optional extra users install at their own version.
- `sdk_httpx()` resolves the httpx flavour from the SDK's own transport
module, so objects handed to `streamable_http_client`, the `sse_client`
factory, and the OAuth metadata helpers come from the module the
installed SDK actually imports.
- HTTP support is gated on either streamable-HTTP entry point, not just
the deprecated alias 2.0 removed.
- OAuth: `OAuthClientProvider` lost its `timeout` argument (the
configured `oauth.timeout` now bounds the callback waiter's own poll
loop, where the browser round-trip was always awaited), and
`callback_handler` must return `AuthorizationCodeResult` rather than a
tuple. 2.0 also validates the RFC 9207 `iss` parameter, so the
callback handler and paste fallback capture it.
`mcp`/`mcp-types` 2.0.0 are inside the 14-day `exclude-newer` window, so
two narrow `exclude-newer-package` entries unblock `uv lock`, annotated
for removal on or after 2026-08-11. `httpx2` needs no exemption: 2.7.0 is
already outside the window and satisfies mcp's floor.
Refs #69931
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Gemini bills thought tokens against maxOutputTokens/max_tokens, so a
global 4096 cap can be fully consumed by thinking on the first
request, leaving zero content tokens and aborting after 4
continuations. When thinking is enabled, raise the effective output
cap to the 65,535 ceiling on both the native and chat-completions
paths.
Refs #83915
- prompt_cache_scope: memo key now includes DB presence (a lazily attached
_session_db re-resolves instead of staying pinned to the physical id);
_persist_disabled agents (background-review forks that never get a DB row)
memoize the fallback instead of re-querying the lineage per API call;
module docstring cross-references get_conversation_root and why the two
lineage resolvers must not be deduplicated.
- chat_completion_helpers: hoist the triplicated
_prompt_cache_scope_for_agent(agent) call to a single local above the
OpenAI-wire dispatch (after the anthropic/bedrock early returns, which
don't use prompt_cache_key).
- codex transport docstring: x-client-request-id mirrors the derived body
key, not the raw scope id.
- turn_context comment: acknowledge the first-turn pre-persist fallback.
- tests: +2 (persist-disabled memoization; lazy DB attach re-resolution).
Legacy compaction mode (compression.in_place: false) rotates the physical
session_id mid-conversation. The prompt-cache scope introduced in #79161 was
derived from that physical id, so every rotation moved the same conversation
into a fresh cache bucket - the prompt cache went cold at every rotation
boundary (#79017).
Fix: resolve a rotation-stable logical scope - the compression-lineage ROOT
of the current session (SessionDB.get_compression_lineage, fork-aware
post-#79193) - once per turn, memoized per transcript segment, and prefer it
over the physical session_id at every prompt_cache_key derivation site:
- agent/prompt_cache_scope.py (new): resolve_prompt_cache_scope(agent) -
lineage-root walk with per-segment memo; falls back to the physical id
when no DB is attached or the walk fails, degrading to pre-fix behavior.
- transports/codex.py: build_kwargs accepts cache_scope_id and prefers it
for the body prompt_cache_key, the xAI x-grok-conv-id header, and the
Codex x-client-request-id routing header. The Codex session_id header
keeps the raw physical id (transcript identity, #57012 contract).
- transports/chat_completions.py: _add_prompt_cache_key accepts
cache_scope_id with the same precedence.
- chat_completion_helpers.py: build_api_kwargs threads the resolved scope
into all three build_kwargs call sites (codex, profile, legacy).
- auxiliary_client.py: set_runtime_main carries cache_scope; the aux
Responses cache-key site prefers it over the physical session_id.
- turn_context.py: resolves the scope once per turn and threads it through
set_runtime_main (no DB walk on the per-API-call hot path).
Scope semantics preserved from #79161: /new starts a fresh scope (new
lineage), /branch children, delegate subagents, and tool children stay
isolated (explicit-fork exclusion in get_compression_lineage), unrelated
sessions keep distinct buckets, and cron per-fire timestamps still
normalize via _cache_scope_from_session_id.
Default installs compact in place (session_id never rotates), so they hit
the memo and produce byte-identical keys to before.
Fixes#79017
Strict OpenAI-compatible providers (onerouter / Qwen, DeepSeek v4) reject
an assistant message carrying tool_calls: [] (or null) with HTTP 400
'Empty tool_calls is not supported in message.'
The pre-API sanitizer in agent_runtime_helpers.sanitize_api_messages already
drops these on the conversation_loop path, but auxiliary / custom-provider
routes that bypass that sanitizer can still reach the wire with an invalid
empty array and abort the whole session (non-retryable 400).
Normalize at the transport layer too: detect an empty-list / null
tool_calls on assistant messages, strip the key on the per-call copy (never
mutate the stored history), and keep real tool_calls untouched. Includes
unit tests covering empty-list, null, real-call preservation, mixed batches,
user-role non-mutation, copy-on-write, and cross-provider parity.
Follow-up to #58755.
A captured native-compaction checkpoint lives in the persisted
codex_reasoning_items sidecar, but the wire restructure that follows it
(prune_pre_checkpoint_items) ran unconditionally: the native gate only
decided whether context_management went into the request, and no signal
from it ever reached _chat_messages_to_responses_input.
So a single checkpoint kept deleting every pre-checkpoint item from all
later requests — after a mid-session swap out of the gpt-5.6 family,
after compression.enabled: false, after the rejection kill switch, and
after a session resume that reloads the sidecar from state.db. The model
receiving the opaque blob was no longer the one able to decode it, and
nothing was logged.
Thread a single native_compaction_eligible boolean, derived from the same
value that gates the context_management field, into the converter. When
ineligible: do not replay type: "compaction" items and do not prune. Safe
because native compaction never truncates Hermes' local history, so the
fallback still carries the full conversation.
All Responses call sites are covered: build_kwargs and convert_messages
derive the flag via _native_compaction_active, the auxiliary/compression
client is explicitly ineligible, and the converter defaults to False
(pre-feature wire) so future call sites are safe by construction.
Fixes#85914
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Azure Foundry's OpenAI-compatible Responses surface rejects the post-tool
follow-up payload with HTTP 400 `invalid_payload` when a replayed encrypted
`reasoning` item is sent alongside `function_call` / `function_call_output`.
The initial function-call request and ordinary multi-turn continuity are both
accepted, so the failure only appears after the first tool executes.
Detect the Foundry endpoint in `ResponsesApiTransport.build_kwargs` and drop
only the encrypted reasoning replay on that follow-up turn, leaving
function_call / function_call_output continuity intact.
Salvage of #59981, rebuilt on current main. Same root cause and fix direction
as the original, which was correct; this version resolves three defects:
- No `chat_completion_helpers.py` change. main already forwards `provider`
and `base_url` to the Responses transport, so the original's re-added
arguments produced `SyntaxError: keyword argument repeated: provider` on
merge. Dropping the hunk removed the syntax error and the conflict.
- Host matching uses `utils.base_url_host_matches`, not a substring test.
`".services.ai.azure.com" in base_url` also matches URLs carrying the
domain in a path or query segment, which would silently disable reasoning
replay on an unrelated provider.
- The post-tool predicate tests the trailing messages, not the whole history.
Scanning for any tool call plus any tool result made it sticky: one tool
call early in a conversation suppressed reasoning on every later turn.
- Tool calls pair on `call_id` as well as `id`. Responses histories carry the
function call id in `call_id` while `id` holds the response item id
(`fc_...`). Identity is resolved via the converter's own
`_split_responses_tool_id`, covering composite `"call_x|fc_y"` ids and bare
`fc_` ids on both sides of the pairing.
Tests: 27 cases across the transport and the live `build_api_kwargs` bridge,
including six parametrized tool-call id shapes, non-Foundry host lookalikes,
the sticky-history guard, parallel tool results, and an unpaired tool result.
Each guard was confirmed to catch its defect by reverting the fix.
Verified with `scripts/run_tests.sh tests/agent/ tests/run_agent/`:
532 files, 5602 tests passed, 0 failed.
Not verified against a live Azure Foundry endpoint — no credentials. The
original HTTP 400 reproduction and post-fix Foundry Project / Azure Container
Apps harness runs are @AshuJoshi's, from #59981. This change is verified at
the payload-construction layer only.
Closes#59981.
Co-authored-by: Ashu Joshi <AshuJoshi@users.noreply.github.com>
Add a non-terminal "review" status so a worker that finished implementation
can hand off for human review without abusing kanban_block. The old
kanban_block(reason="review-required: ...") convention routed the handoff
through the unblock-loop breaker, so a normal review -> changes -> review
cycle was falsely escalated to triage.
- kanban_db: request_review (running/ready -> review, non-block, emits
review_requested), reopen_review_task (review -> ready/todo, review_reopened),
complete_task accepts review -> done, and a review_dispatch gate (default off,
shared by the dispatcher loop and the gateway health probe).
- kanban_request_review worker tool + `request-review` / `reopen-review` CLI
verbs; tool wired through toolsets, EXPOSED_TOOLS, _POLISHED_TOOLS.
- Gateway notifier wakes the origin subscriber on review_requested and
block_loop_detected; the subscription survives until done/archived, so every
review cycle re-notifies.
- Dashboard PATCH + bulk route the review transitions (request_review /
reopen_review_task) and render the review column.
- goals.py goal-loop and KANBAN_GUIDANCE recognize review as a terminator.
- Docs (reference tables, user guide, AGENTS.md, zh-Hans mirrors) + tests.
needs_input / failed are unchanged: they still route through kanban_block,
still count toward block_recurrences, and still escalate to triage.
After a partial update (stash restore overwriting providers/base.py with
an older version), the NousProfile singleton was instantiated from a
ProviderProfile class that predates the supports_prompt_cache_key field
(added in f4fb23f3d). Accessing profile.supports_prompt_cache_key raised
AttributeError, crashing every API call with:
'NousProfile' object has no attribute 'supports_prompt_cache_key'
Use getattr(profile, 'supports_prompt_cache_key', False) so a stale
profile degrades to 'no prompt cache key' instead of crashing.
The Anthropic SDK's streaming accumulator builds ParsedMessage snapshots
whose ParsedTextBlock content doesn't match the generic union pydantic
expects, so model_dump() on stream events (message_stop) emits
PydanticSerializationUnexpectedValue UserWarnings straight into the
user's CLI output mid-response.
Pass warnings=False at every helper that dumps arbitrary SDK models
(relay_llm/_jsonable, relay_tools/_jsonable, anthropic_adapter
_to_plain_data, run_agent _hook_jsonable, chat_completion_helpers
extra_content/reasoning_details sites, chat_completions transport),
with a TypeError fallback for duck-typed model_dump implementations.
Adds regression tests including a precondition test that proves the
fixture still trips the warning without suppression.
Opt-in via compression.codex_responses_native (default: false). When enabled,
gpt-5.6-family models on the direct OpenAI API (api.openai.com) or a ChatGPT
Codex subscription send context_management=[{type: compaction,
compact_threshold: N}] on Responses requests. OpenAI compacts server-side and
returns an encrypted compaction output item; Hermes captures it into the
existing codex_reasoning_items sidecar and replays it on later turns in place
of the pruned history — inheriting persistence, session replay, the
cross-issuer guard, and the encrypted-replay kill switch with zero new state.
Scope is deliberately hard-gated (agent/native_compaction.py, re-checked per
request): gpt-5.6 family only — gpt-5.1/5.2 fail server-side on the field
(HTTP 500 / stream stall, no structured rejection; live-verified) — and
direct OpenAI/Codex routes only; xAI, GitHub/Copilot, OpenRouter, relays,
and local servers never see the field.
Hermes' local compression stays armed as the fallback owner: the native
threshold is clamped ~8K tokens below the local trigger so the server
compacts first, and a structured provider rejection of context_management
disables native compaction for the session and retries without it
(one-shot guard in TurnRetryState).
Live-verified E2E on api.openai.com/gpt-5.6: server compaction fired at a
4K threshold, checkpoints captured and replayed, recall preserved across
3 turns; gpt-5.1 with the flag enabled stays clean (field never sent).
Direction credit: PR #76950 by @laryhorb explored native Responses
compaction; this is a minimal reimplementation on current main.
Cherry-picked from PR #78959 by @JoaoMarcos44 with authorship preserved.
Follow-up: hoist _cache_scope_from_session_id(session_id) to a local in
build_kwargs so it's computed once instead of 4 times per call.
Closes#78941. Closes#79012. Closes#79013. Closes#79014. Closes#79015.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Drop the manual web.search_backend / web.backend config-reading block
that duplicated _read_config_key in web_search_registry.py. The function
now delegates directly to get_active_search_provider() (which reads the
same config keys via the registry's canonical resolver) and falls back
to _get_search_backend() only when the registry has no providers loaded.
Also updates the TestXaiWebSearchBackendPreference tests to monkeypatch
the registry instead of load_config_readonly, and adds two new tests for
the legacy fallback path (no provider registered -> _get_search_backend).
When Grok runs on xAI Responses, only swap to native server-side
web_search when the active/configured backend is xai. For Firecrawl
and other Hermes providers, keep client dispatch under a renamed wire
tool so Grok cannot hijack web_search and ignore user config.
Review follow-up on the #56798 salvage: the gate shipped fully dormant
(no provider profile sets supports_prompt_cache_key, no production
caller passes it, and no plain 'openai' profile exists to set it on) —
AGENTS.md rejects dead code wired in without E2E proof.
Activate the one endpoint where the field is first-class: exact-host
api.openai.com (OpenAI documents prompt_cache_key; GPT-5.6+ docs
recommend it for cache routing). Deliberately NOT substring matching —
Azure/OpenAI-compat endpoints may reject unknown fields and stay
opt-in via the flag. 4 new tests (imply + 3 spoof/proxy/Azure
negatives); mutation-checked (substring-weakened host check fails the
spoof tests).
When an approval prompt expired without a response, every CLI-side path
collapsed the timeout into the same 'deny' choice as an explicit user
refusal, so the agent was told the user denied the action when the user
simply never answered. The gateway wait already distinguished the two
('timed out without user response... Silence is not consent.'); this
brings the CLI/TUI/ACP surfaces to parity.
- prompt_dangerous_approval(): input()-path expiry now returns a distinct
'timeout' choice (still fail-closed).
- cli.py _approval_callback + hermes_cli/callbacks.py approval_callback:
deadline expiry returns 'timeout' instead of 'deny'.
- check_all_command_guards / _run_approval_gate CLI tails: 'timeout' maps
to outcome='timeout' with a 'timed out without user response... Silence
is not consent.' BLOCKED message (matching the gateway wording);
explicit deny keeps outcome='denied' and gains user_consent=False for
shape parity.
- computer_use: 'timeout' verdict threads through the CLI adapter and
yields a 'prompt timed out — the user did not respond' error instead of
'denied by user'.
- ACP permissions bridge: FutureTimeout returns 'timeout' (other failures
still 'deny'); elicitation maps 'timeout' to 'cancel' like the gateway's
unresolved outcome; codex wire mapping documents deny/timeout→decline.
- write_approval already treats unknown choices as 'stage, not drop', so
a timeout now stages the memory write instead of silently refusing it.
Every timeout path remains fail-closed — the action never runs; only the
classification reported to the agent changes.