573 Commits

Author SHA1 Message Date
kshitijk4poor 5d07f7fe66 test(agent): fold the aux create_client() seam tests to two contracts
Two parametrized invariant tests (native client sync/async; None or raising
profile falls back to the standard client) replace four near-duplicates, and
the hook's kwargs are pinned exactly to the mapping openai.OpenAI would have
received. The fixture isolates both provider registries through monkeypatch
instead of a hand-rolled snapshot/restore, drops the HERMES_HOME override that
tests/conftest.py already provides, and replaces the tuple-truthiness lambda
with a plain function. Call-site comment trimmed to the ordering WHY.
2026-09-16 11:57:56 +05:30
liuhao1024 4cd2eb013c fix(agent): honor ProviderProfile.create_client() for api_key aux routes
The auxiliary api_key branch built openai.OpenAI directly, so an
out-of-tree provider registered with auth_type="api_key" lost its
native transport for auxiliary tasks even though the main-agent path
(_provider_supplied_client) and the external_process branch honor the
same hook. Consult the profile's create_client() before the built-in
gemini/OpenAI ladder; None (the default) falls through untouched, and
a raising profile is logged and skipped (#112384).
2026-09-16 11:57:56 +05:30
kshitijk4poor a12b3c7aa3 refactor(agent): import the transform bypass from its defining module, no re-export shim
Internal moves get no compat aliases (root AGENTS.md); codex_runtime and
auxiliary_client import bypass_sdk_request_transform from agent.sdk_transform_bypass.
2026-09-15 19:23:53 -07:00
teknium1 dc7e52874e docs(agent): record why the Z.AI vision default is glm-5.3-flash
Pinned vision ids rot silently (glm-5v-turbo was retired from the Coding Plan endpoints
while still valid on pay-as-you-go), so name the data behind the new pin — the only
image-capable GLM id on every Z.AI surface — and why the pin cannot simply be dropped in
favour of ProviderProfile.default_vision_model() (ZaiProfile returns None, which would route
vision to a text-only chat model).

Co-authored-by: POWERFULMOVES <142271328+POWERFULMOVES@users.noreply.github.com>
2026-09-15 18:23:55 -07:00
KoNit-K c93f2e1d59 fix(agent): update zai vision fallback 2026-09-15 18:23:55 -07:00
teknium1 5435ac8cc4 fix(aux): clamp extra_body.reasoning effort too so auxiliary.<task>.reasoning_effort ultra never reaches the wire 2026-09-15 18:20:13 -07:00
teknium1 9e45a90488 fix: clamp aux reasoning effort once before profile projection
Follow-up to the cherry-picked #112019 (@KoNit-K): the clamp in the generic
``extra_body.reasoning`` fallback only covered providers WITHOUT a
reasoning-aware profile. On the profile path (OpenRouter/Nous slots used as
MoA aggregator or aux model) ``_project_provider_profile`` received the raw
config and the OpenRouter profile passes ``ultra`` through whenever the
catalog vocabulary is cold, so the 400 from #112010 survived there.

Move the clamp up to ``_build_call_kwargs`` so both the profile projection
and the fallback see a wire-level effort — the same entry clamp the main
transport applies in ``_reasoning_config_for_model`` (#89503). The shared
policy lives once in ``agent.reasoning_effort.clamp_reasoning_config``; the
transport delegates to it instead of carrying its own copy.

Offline kwargs probe (issue's exact call): before
``extra_body.reasoning == {'enabled': True, 'effort': 'ultra'}`` on nous and
openrouter aux/MoA routes; after ``'effort': 'max'`` on every route,
``high`` verbatim and ``{'enabled': False}`` unchanged.
2026-09-15 18:20:13 -07:00
KoNit-K c02db64077 fix(agent): clamp auxiliary ultra reasoning effort 2026-09-15 18:20:13 -07:00
teknium1 f55d4f6747 fix(aux): bare-custom AuthError yields no custom endpoint instead of a stale env OPENAI_BASE_URL 2026-09-15 18:19:50 -07:00
teknium1 9336fb11cd fix: honour HERMES_CODEX_BASE_URL on the raw Codex client too
The raw_codex branch of _resolve_openai_codex_branch (main agent built without
explicit creds via agent_init._routed_client_kwargs, and the mid-turn fallback
chain) still hardcoded the official Codex endpoint while the pooled/aux/singleton
paths honoured the override, so a proxy user's main agent silently bypassed it.
Hoist the profile-scoped read into _codex_base_url_override and use it in both
builders; the Cloudflare identity headers follow the resolved base_url.

Review finding: raw_codex client ignored HERMES_CODEX_BASE_URL while pooled/aux/singleton honoured it.
2026-09-15 04:52:47 -07:00
teknium1 70ff4863d7 fix(aux): read HERMES_CODEX_BASE_URL through the profile secret scope
The aux Codex client already resolves its API-key env vars via _scoped_key_env
so a multiplexed profile never borrows a sibling's value; the endpoint override
follows the same rule instead of a bare os.getenv.
2026-09-15 04:52:47 -07:00
joaomarcos b62bb2a3d5 fix(codex): honor HERMES_CODEX_BASE_URL on pooled and aux codex client resolution 2026-09-15 04:52:47 -07:00
teknium1 d4e0df1e8d fix: keep the OpenCode keyless Authorization blank on the async aux client
_to_async_client rebuilds default_headers from scratch and dropped the blank
Authorization set by _create_openai_client, so every async aux call (the path
aux tasks actually take) for a free-tier OpenCode model still shipped the
keyless placeholder as a bearer and 401'd. Re-apply the same keyless policy on
the async twin.

Review finding: fourth client builder (_to_async_client) missed the keyless header policy.
2026-09-15 04:51:15 -07:00
teknium1 27faf9ced2 fix: surface the narrowed error when the aux recovery ladder is exhausted
With no fallback_chain (or an exhausted one) the ladder tail returned the
_RERAISE_ORIGINAL sentinel and call_llm's bare `raise` re-raised the ORIGINAL
error. After a 401 that the refresh rung already healed, the retry's actionable
"404 requires available credits" was hidden behind "401 Unauthorized", steering
the user to re-login instead of top-up. Raise the narrowed `first_err` from the
ladder tail instead and drop the sentinel; the pool-rotation rung shares the
same tail so it is covered by the same change.

Review finding: exhausted no-chain path re-raised the healed 401 instead of the retry's 404 credits error.
2026-09-15 04:49:01 -07:00
Synero 1f88ebf3ec test(agent): pin the auth-refresh retry boundary and the pool-gate split
Review feedback on the fall-through fix: the accept boundary and the
retry-fails-with-a-connection-error path through the credential rung were
reachable but untested, and two comments described behavior the code does
not have. Tests and comments only, no behavior change.
2026-09-15 04:49:01 -07:00
Nacho 64a4687153 fix(agent): fall through to the configured chain when an auth-refresh retry fails
Both auth-refresh rungs in the auxiliary recovery ladder performed their
post-refresh retry with a bare `yield` inside a `return` statement, so an
exception from that retry escaped `_aux_recovery_ladder` instead of being
converted into `(None, err)` by `_rung()`. The ladder therefore never reached
`_ladder_provider_fallback` and the configured
`auxiliary.<task>.fallback_chain` was silently skipped.

Observable case: a stale Nous runtime token makes the auxiliary request return
401; the ladder refreshes the credential and retries; when the retried model is
out of credits the resulting 404 escapes the ladder and the task fails (context
compression degrades to truncation) even though a healthy fallback_chain is
configured. With a fresh token the same 404 takes the payment path, which
already falls through correctly.

Guard both rungs with `_rung()` - the helper the sibling payment rung and the
credential-pool rung below already use - so a failed retry resumes the ladder.
2026-09-15 04:49:01 -07:00
moxian 9c7de107f2 fix(auxiliary): retry without response_format when a provider rejects the object-form json_schema by shape
_is_structured_output_rejection matched several phrasings for a provider refusing the
structured-output field, but not gateways that validate the body with a strict pydantic
model and reject the OBJECT-form json_schema by shape:

  HTTP 422: {"detail":[{"loc":["body","response_format","json_schema"],
             "msg":"str type expected","type":"type_error.str"}]}

422 already passes the status check; only the phrasing was missing, so the
one-retry-without-response_format rung never engaged and the auxiliary task failed
hard (title generation left 'HTTP 422' in the session title). Treat the shape error as
a rejection: the field is what the provider refuses and the remedy is identical.

Fixes #110631
2026-09-15 04:46:36 -07:00
teknium1 9219f56e44 fix(auxiliary): recognise Bedrock's "doesn't support"/"is deprecated" sampling-param rejections
_is_unsupported_parameter_error matched "does not support" but not the contraction Bedrock
Converse returns for xAI Grok ("This model doesn't support the temperature field"), nor the
inference-profile Claude wording ("`temperature` is deprecated for this model"), so the aux
ladder's retry-without-temperature rung never fired and the ValidationException surfaced on
title/vision/compression calls (#111043, second half).
2026-09-15 04:45:33 -07:00
teknium1 88f2844d46 fix(agent): cover the remaining refusal-only surfaces and fold the tests
- codex_runtime._CODEX_PROGRESS_DELTA_TYPES gains response.refusal.delta so the
  stream watchdog sees progress on a refusal-only stream instead of timing it
  out as idle.
- auxiliary_client._parse_codex_final_response reads type=refusal content
  parts; without it an aux refusal-only turn parsed to content=None and hit the
  empty-response path the main loop was just taught to avoid.
- tests: parametrize test_streamed_refusal_accumulated (refusal-only /
  alongside-content) so there is one test per surface; drop upstream product
  references from docstrings (credit stays in the PR body); pass encoding= to
  the read_text calls flagged by the Windows footgun scanner.
- docs: fallback-providers notes that a streamed refusal is a terminal
  content_filter result, not an empty response to retry.
2026-09-15 03:39:07 -07:00
kshitijk4poor 14a93d67eb refactor(codex): take the aux adapter's xAI flag from the shared route classifier
`_CodexCompletionsAdapter._build_responses_kwargs` already asks
`classify_responses_route` for the GitHub flag but hand-rolled its own
x.ai host match for `is_xai`. Two classifiers for one route drift; use
`route.is_xai_responses` so the aux path agrees with the main transport.
`is_copilot` (used for `is_github_responses` replay stripping) is left
as-is.
2026-09-15 10:49:19 +05:30
kshitijk4poor 16986c4bff fix(codex): canonicalise the custom-endpoint issuer kind
The openai SDK appends a trailing slash to `client.base_url`, so the aux
adapter stamped `other:https://h/v1/` while the main transport stamped
`other:https://h/v1`. On custom Responses endpoints every aux call
(compression, flush_memories) therefore dropped all main-minted reasoning
items as "foreign".

`_classify_responses_issuer` now strips whitespace and trailing slashes and
lowercases scheme+netloc before stamping. The aux adapter also derives its
route flags from `classify_responses_route` — the single owner of the
codex/xai/github predicates — instead of an inline chatgpt.com host check,
and reuses the same flags for the effort clamp.
2026-09-15 10:49:19 +05:30
Fangliquan 51ebdff570 fix(codex): scope encrypted-reasoning replay to the issuing model
Encrypted reasoning blobs are sealed to the model that minted them, not
only to the endpoint. Switching models on the same custom Responses
endpoint therefore replayed blobs the new model cannot decrypt and the
turn failed with HTTP 400.

Stamp captured reasoning items with `_issuer_model` (the canonical wire
model) alongside `_issuer_kind`, and replay an item only when both the
issuer kind and the model match the current request. Endpoint-stamped
legacy items without model provenance are dropped once the current
model is known (fail closed); ordinary assistant text stays replayable.
The transport threads the effective wire model (request_overrides win)
into conversion and normalization; the auxiliary Codex adapter stamps
and filters against its own model rather than the main agent's. The
400 classifier also recognises the custom-endpoint wording
"encrypted content could not be decrypted or parsed" so recovery strips
the replay state instead of aborting.

Hand-grafted from #95849 (final head d9cf6bcc08) onto current main; the
middleware-model-rewrite half is intentionally left out.

Closes #95834
2026-09-15 10:49:19 +05:30
Steve Ahlstrom a714ff9823 fix(aux): stop probe stubs poisoning the client cache
`_store_cached_client()` refuses an `_AuxProbeClientStub`, but
`_get_cached_client()` assigns to `_client_cache` directly and so never
reaches that guard. check_fns resolve through this path inside
`aux_probe_mode()` during tool-schema assembly, and the cache key carries no
probe/runtime distinction — so the stored stub is returned to the next real
caller sharing that key, which dies on attribute access with
`_AuxProbeClientStub used as a real client (attribute 'chat')`.

The `async_mode` field in the cache key is what kept this latent: the probe
caches the sync variant, so async consumers (`analyze_image`) miss the entry
and build a real client, while sync consumers (`browser_vision`) hit the
poisoned one and fail on every call.

Observed against a local OpenAI-compatible vision endpoint, where
`browser_vision` failed every call with that RuntimeError while
`analyze_image` against the same provider worked.

Guard the inline store the same way `_store_cached_client()` does, and
return the stub to the probe caller without caching it.

The existing `test_probe_stub_never_cached` pins the invariant only on
`_store_cached_client()`, which is why the unguarded door went unnoticed;
the new test exercises `_get_cached_client()` and fails without this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-13 11:22:08 -07:00
kshitijk4poor aa40c1d21b fix(auxiliary): wrap the Messages adapter on the profile's declared api_mode
resolve_provider_client("commandcode-anthropic") with no explicit api_mode (a bare
``auxiliary.<task>.provider`` entry) built a plain OpenAI client: _wrap_transport
only consulted req.api_mode and URL heuristics, and api.commandcode.ai/provider/v1
matches none. With _reasoning_config now emitted for that profile, an unwrapped
client would TypeError on the unknown kwarg; before, the OpenAI-wire call simply
misfired against a Messages endpoint. Fall back to the registered profile's
api_mode so the wrap and the reasoning gate agree.
2026-09-13 19:46:52 +05:30
kshitijk4poor 5dd8fa1daa fix(auxiliary): anthropic_messages profiles keep /reasoning reachable on aux calls
commandcode-anthropic (api_mode=anthropic_messages, OpenAI-shaped base URL) lost
thinking control on compression/title/vision calls after #109530: once its class
overrides build_api_kwargs_extras, _project_provider_profile marks the profile as
handling reasoning and drops the generic extra_body.reasoning fallback that the
Anthropic adapter used to read — and the _reasoning_config gate only fired for
provider=anthropic, Portal /v1/messages ids, /anthropic URLs and MiniMax.

Carry the profile's api_mode on the projection and include it in that gate, so
any anthropic_messages profile reaches the adapter regardless of URL shape.
Chat-completions profiles are unchanged.

Before: _build_call_kwargs("commandcode-anthropic", claude-haiku, {"enabled": False})
        -> no _reasoning_config, no extra_body.reasoning; adapter defaults thinking.
After:  -> _reasoning_config={"enabled": False}.
2026-09-13 19:46:52 +05:30
Keane Yan d1a13ae244 fix(auxiliary): normalize model on auto cache miss 2026-09-12 08:28:49 -07:00
Teknium 4b8c01f691 fix(multiplex): key tool-side and agent-side memos by profile home
Camofox VNC one-shot, computer-use aux-vision verdict, tirith binary path, MCP
discovery lock path, remote-backend probe text, learned image token costs,
auxiliary per-task semaphores and the custom-endpoint /models memo all held one
profile's config-derived value for the whole process. The skill-sync debounce
Timer ran with empty ContextVars, so a secondary's write pushed as the launch
profile (and cancelled its pending push).

Each memo is now keyed by hermes_home_key() (or credential fingerprint for the
per-key catalog) under an override; the timer is per home and runs its callback
inside the scheduling turn's copied context. Unscoped slots are unchanged.
2026-09-12 01:35:05 -07:00
Teknium 819517fbac fix(gateway): routed profiles get their own max_turns, fallback chain, hooks, aux auth and media policy
One multiplexed gateway process serves every profile, but several per-turn
reads still went through state frozen from the LAUNCH profile:

- `_current_max_iterations` re-bridged `agent.max_turns`/`sessions.*` from the
  module constant `_hermes_home` into one process-wide HERMES_MAX_ITERATIONS,
  so every secondary ran with the default profile's turn budget. A routed turn
  (HERMES_HOME override) now resolves `agent.max_turns` from its own config.
- `_refresh_fallback_model` read `_hermes_home/config.yaml` into one runner-wide
  slot, so secondaries fell back through the default's provider/model with their
  own keys. It now reads the active gateway home and keeps a last-known-good
  chain per home.
- `_load_prefill_messages` resolved relative paths against the launch home.
- `agent/auxiliary_client._AUTH_JSON_PATH` was an import-time constant, so a
  secondary's compression/title/vision calls authenticated to Nous with the
  default profile's token when it had no pool entry. Resolved per call via
  `hermes_cli.auth._auth_file_path()` (patched constant still wins in tests).
- `gateway/hooks.HOOKS_DIR` was frozen at import and one `HookRegistry` was
  loaded outside any profile scope, so secondaries' `hooks/` never ran and the
  default profile's handlers received every profile's messages, responses and
  user ids. `HOOKS_DIR` now resolves per call (salvaged from #56508) and the
  runner holds one registry per served home, picked from the active scope at
  emit time and front-loaded under each secondary's startup scope.
- Shell-hook subprocesses inherited the launch `os.environ` (default HERMES_HOME
  and the default profile's secrets). They now get the routed HERMES_HOME via
  `build_subprocess_env`, scrubbed under multiplexing, and the stdin payload
  carries `profile` so one script can tell which profile fired it.
- Media-delivery policy (`gateway.strict`, `media_delivery_allow_dirs`,
  `trust_recent_files*`) was bridged once into env at startup and read from env
  per delivery; under a HERMES_HOME override the validator now reads the routed
  profile's config. Single-profile runs keep the env-bridge contract.

Audit: /tmp/mux_audit F3, F4, F6 (auth.json half), F7, F12 (media). Live repro
(temp HERMES_HOME A with profiles/B): before, B saw max_iterations 7,
fallback A/fallback, TOKEN_A, A's hooks, strict=A; after, all B's values.
2026-09-11 15:44:00 -07:00
Teknium a9838c2100 fix(multiplex): tool and memory-provider env reads stay inside the routed profile
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's .env; a
secondary profile's values exist only in the per-turn secret scope. Every reader
below still read os.environ/os.getenv at call time, so a secondary profile's turn
silently used the default profile's value.

Credentials (F6): FIRECRAWL_API_KEY (read_file hosted OCR), OPENVIKING_API_KEY,
mem0-OSS OPENAI_API_KEY, MODAL_TOKEN_ID/SECRET and BROWSER_USE_API_KEY presence
gates, and the xAI video plugin's os.getenv("XAI_API_KEY") fallback AFTER the
scoped resolver had already missed — the exact fallback-after-miss shape
gateway/AGENTS.md forbids. Deleted, not re-scoped: the resolver is the scope.

Identity / tenant (F7): MEM0_USER_ID/AGENT_ID/HOST/MODE, SUPERMEMORY_CONTAINER_TAG,
RETAINDB_PROJECT, OPENVIKING_ACCOUNT/USER/AGENT (and the whole layered() env
read), HINDSIGHT_BANK_ID/MODE/retain shaping, HERMES_HONCHO_HOST. A raw read
put a secondary profile's memories into the default profile's account/bank/
project/tenant and recalled them back into the default's turns. Each now uses
get_secret with the provider's own per-profile default on a miss.

Endpoints (F8): OPENAI_BASE_URL (aux custom runtime + direct-alias expansion),
XAI_BASE_URL/HERMES_XAI_BASE_URL (aux OAuth), NOUS_INFERENCE_BASE_URL (#65941,
both the aux builder and hermes_cli.auth_nous._nous_inference_env_override),
GATEWAY_PROXY_URL (same UnscopedSecretError-only fallback shape as
GATEWAY_PROXY_KEY three lines below), FIRECRAWL_API_URL, BROWSERBASE_BASE_URL,
SUPERMEMORY/RETAINDB/HONCHO/HINDSIGHT URLs. The keys beside them were already
scoped, so a secondary's key was sent to the default profile's proxy or host.

Targets / display (F11): WEIXIN_HOME_CHANNEL (message posted into the default's
chat), HERMES_LANGUAGE, and agent/i18n's process-wide lru_cache of
display.language — now keyed by HERMES_HOME.

Outbound webhooks: hooks.outbound[].secret_env resolved from os.environ while
the gateway registers each profile's targets inside that profile's scope, so a
secondary's deliveries were signed with the default's secret or left unsigned.

Agent-cache eviction: _spawn_release_thread started a bare threading.Thread, so
commit_memory_session -> provider on_session_end ran with an EMPTY context. The
thread now runs copy_context() and, for the unscoped housekeeping sweep, enters
the owning profile's _profile_runtime_scope resolved from the session key
(agent:<profile>:...). The pressure batch does the same per key.

session_search (#82903): agent/inline_tool_executors.py::_session_search
forwarded every schema argument except `profile`, so a gateway agent could
never select a named profile's store. Forwarded; the ownership-scoping design
in #87779/#87847 is a separate design call and is not attempted here.

Live repro (/tmp/mux_audit/fix-tool-memory-reads/repro.py): 28 FAIL on
origin/main -> 0 FAIL with this change; 10 new invariant tests red on base.

Fixes #82903
Fixes #65941
Fixes #99121
Addresses #87779
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
Co-authored-by: Michael Versluis (Berry) <michael@wve.nl>
2026-09-11 15:26:46 -07:00
Teknium 2b4deeb32b fix(auth): key sibling per-process credential memos by profile home under multiplex
Same class as the resolve_nous_access_token memo: three more process-wide
memos carried a credential resolved under one profile's HERMES_HOME override
into another profile's turn for their TTL.

- hermes_cli/nous_billing.py::_token_cache (30s (token, base) memo for the
  charge poll loop) was a single unkeyed slot -> dict keyed by
  hermes_home_key(); invalidate_cached_token() clears the dict.
- agent/moa_loop.py::_runtime_cache carried api_key/base_url/api_mode keyed
  only (provider, model) for 5 min -> (hermes_home_key(), provider, model).
- agent/auxiliary_client.py::_client_cache_key had no profile component, so
  callers that omit api_key (pool / Nous auth.json paths) could be handed a
  client built with another profile's bearer -> hermes_home_key() leads the key.

WHY hermes_home_key(): it reads the per-turn HERMES_HOME override the
multiplex gateway sets (falling back to the env var), and it is
symlink-stable, so the memo key is exactly the credential home the
resolution itself read from. Profiles stay independent islands; the
default-profile process env never leaks into a secondary's turn.

Tests: one invariant per site, proven red on origin/main.
2026-09-10 18:11:25 -07:00
Siddharth Balyan 4bdd64b334 The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract

The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an
anonymous account exactly one model on its own host, refuses everything else with a structured
429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and
tells a signed-in account that still asks for `nous/welcome` what to switch to in an
`x-nous-model-switch` header. Four client-side gaps against that contract:

- Auxiliary calls were refused on every session. The auxiliary client asked the welcome host
  for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free`
  before each fallback. On the welcome host it now uses `nous/welcome` (its backing model
  covers auxiliary work) and skips Nous for vision, which the welcome model does not take.

- The structured 429 body was never read. The classifier now parses `reason` /
  `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are
  non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited`
  are rate limits that honour `retry_after` and never rotate the free tier's only credential.
  The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and
  fall back instead of retrying or re-exchanging. The terminal paths say what happened and
  name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal).

- The `x-nous-model-switch` header was ignored. The chat-completions transport records it
  beside the rate-limit and credits headers; the next call moves the session, and the config
  default when it still names `nous/welcome`, to the backing model the gateway named.

- A guest fell back to the paid host. With `inference_base_url` absent from the exchange or
  outside the host allowlist, routing defaulted to inference-api, where every request is a
  400. A guest now defaults to the welcome literal at the exchange, in the shared store's
  shape, and in effective routing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8)

* feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use

A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided
setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier
picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime
rung now sets it up there instead of failing "not logged in", so the guided chat no longer races
the root profile's first-run mint.

nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever
nothing else is configured; "on-request" mints only when the free tier is asked for by name
(nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers —
the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the
connector token path — still adopt what the shared store holds, so every profile follows the one
identity the guided setup created, but never create one on their own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c)
(cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165)

* feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit

Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity
(nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is
configured. "explicit" means Hermes never creates one on its own: the only creator is the new
provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup
on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and
before the guided chat exists — so the identity lands in the root store every profile reads
through and is there before any session asks for nous/welcome. That closes the race against the
backend's own setup, and makes "only when the setup-bot flow is used" literally true.

The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome
(the hermes model row, --provider nous), which treated a model name as intent and was broader
than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not
logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to
sign in from. Implicit callers still adopt an identity the shared store holds, and a retired
credential is replaced (a continuation, not a creation).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0)

* fix(auth): remove the nous.guest_setup knob; the free tier is created on first use

`nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its
default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`,
unknown values read as `auto`, and under `explicit` a CLI-only install could never get an
identity, which contradicts the first-run contract (first command mints, then chats).

The mint race the knob accompanied is already benign: every caller takes the profile lock then
the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup
win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which
stays. `nous.guest` remains the only free-tier policy.

Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through
`ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call
sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config
entries. The three policy tests that hold regardless of the knob are kept under
`TestExplicitProvision`; the two that only tested the knob are deleted.

(cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261)

* fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent

When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway
serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch`
moves the live session to that model and moves `config.yaml`'s default off the alias in the
same step. The messaging gateway's fallback-eviction check compares the agent's model with the
config default and evicts on any mismatch that is not a /model override, so when the config
write did not land (unreadable config, lock) the cached agent was evicted once per turn, and
prompt caching with it.

`apply_model_switch` now stamps the alias it moved the session off on the agent, and
`_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as
deliberate, beside the existing /model override case. The check takes the agent and the config
model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them.

(cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954)

* fix(auth): the free tier outranks implicit host credentials in provider resolution

On a fresh install with a leftover ~/.aws profile, resolve_provider("auto")
reached the Bedrock rung before the free-tier rung, so the first turn ran on
Bedrock and failed 403 while the free tier was still being minted in the
background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s,
three retries, no answer; the next process then switched to nous/welcome.

The free-tier rung now sits directly above the Bedrock chain: when nous.guest
is on, an existing free-tier identity answers, else a blocking mint runs, and
only then does the boto chain get a say. Everything above is unchanged and
still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter
pool, a logged-in active_provider. nous.guest: false skips the rung, and a
failed mint still falls through to Bedrock and the no-provider guidance.

Tests: six precedence cases (identity present, fresh mint, free tier off, env
key still wins, sign-in still wins, failed mint falls through). The opt-out
test now neutralizes the AWS chain like the precedence tests do; on a machine
with ~/.aws it was failing for the same reason as the bug.

Live after the fix, same Mac, AWS credentials visible, isolated shared store:
identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in
11 s.

(cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a)

* fix(auth): review follow-ups for the free-tier rung (NS-829)

- tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches
  the free tier off; its contract is the boto chain, and the free tier now
  sits above it.
- gateway/run_notifications.py: the free-tier startup line reads auth.json
  before consulting the resolver, so a gateway boot on a machine with AWS
  credentials never mints or refreshes over the network.
- hermes_cli/anon_auth.py: module docstring says where the free tier sits in
  the ladder instead of "the ladder is untouched".
- tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of
  six (parametrized ladder cases; a failed mint that returns None or raises
  falls through to Bedrock).

scripts/run_tests.sh on the five affected files: 147 passed, 0 failed.

(cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11)

* feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone

The free tier is pre-GA. Until GA it must not exist for anyone who did not
ask for it: no identity minted, no portal traffic, no free-tier copy on any
surface. One environment variable now decides that, and one function reads it.

`guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly
"1"; only then does `nous.guest` (the user's off switch) get consulted. Every
free-tier site already funnels through `guest_enabled()`, so the gate closes
minting, routing, connector entitlement, status lines and the picker row in
one place. With the variable unset, `resolve_provider("auto")` on a fresh
install raises `no_provider_configured` exactly as upstream does.

`HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the
gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted
identities as a side effect of provider resolution, and `_has_any_provider_
configured` read them ahead of every other check, making the CLI a second
reader of a flag that must have exactly one. `_forced_new_done` and the
`force` parameter of `_reconcile_and_provision` go with them.

Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in
b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847).

Not a user preference: the variable is never written to config.yaml or .env
and never shown in setup. It is deleted at GA together with its comment in
anon_auth.py. This is a deliberate, temporary exception to the "no new
HERMES_* env vars for non-secret config" rule.

Tests: fixtures set the gate instead of deleting the old lever; one new
invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "",
"0", "true" and "new" all leave the tier off with zero portal calls, red on the
previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the
feature.

* feat(auth): the free-tier identity is created in one place, at boot; every other site is a read

Before this commit eight sites could create a Nous free-tier identity as a
side effect of something else: resolving a provider, the CLI's first-run
check, the CLI's session setup (in the background beside an own key), a
connector bearer read, the desktop polling `free_tier.status`, the sign-in
precondition, the desktop's `free_tier.provision`, and the dead-credential
re-mint. A poll could mint. Provider resolution could hit the network. Two
of them raced each other on a fresh install.

Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator.
`hermes serve` runs it on a daemon thread from `_lifespan` beside the other
background boots; `cmd_chat` runs it synchronously before the first-run
guard. It inventories credentials first (`resolve_provider("auto",
skip_free_tier=True)`: what would carry inference if the free tier did not
exist), creates the identity only when `guest_enabled()`, resolves inference,
records a `SetupRecord` in process memory and broadcasts ONE `setup.ready`
event. It runs on every boot; only the mint is gated.

`ensure_portal_identity` now requires `explicit=True` and raises otherwise.
Its callers are the bootstrap, the desktop's `free_tier.provision` (the
explicit retry when the boot could not create the identity) and the two
dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`,
`managed_tool_gateway._replace_dead_guest_token`). The background thread
path and `provision_free_tier` are deleted with their last callers.

Reads that used to mint and now only read: `auth.py::resolve_provider`
rung 7 (an existing identity still outranks the Bedrock chain, NS-829
ordering kept), `main.py::_has_any_provider_configured`,
`cli_agent_setup_mixin._ensure_runtime_credentials`,
`managed_tool_gateway.read_nous_access_token` (no identity -> None),
`anon_sign_in.run_sign_in` (no identity -> Unavailable),
`methods_free_tier` `free_tier.status`.

`setup.status` answers from the record for the launch profile, blocking up
to 8 s while the bootstrap is in flight so a client's first poll lands after
the identity exists rather than racing it; a named profile, or a process
that never ran the bootstrap, keeps today's live probe. The record's fields
ride along additively (`ready`, `free_tier`, `other_providers`,
`inference_provider`).

Identity and inference are decoupled (NS-845 Q1.3): the mint sets
`active_provider="nous"` only when the inventory found nothing else usable
(`_mint_locked(carries_inference=)`); an adopted account always does. A token
refresh no longer re-elects the provider it refreshed
(`_save_provider_state_to_source` writes credentials, not the user's
choice) — that write was how an own-key install ended up on the free tier
after the first connector call.

Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check),
bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint),
62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and
`provision_free_tier`), and a04b05260c (blocking mint in the resolver).
Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847.

Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps
inference; reads never reach the portal; a refused mint is memoised),
`free_tier.status` fails loudly if it ever calls the creator, the resolver
stub fails loudly if resolution ever mints, `setup.status` reads the record,
`skip_free_tier` proves the inventory question. The three sign-in tests for
the deleted pre-mint collapse into one (`no identity -> Unavailable, zero
portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and
off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI
matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20.

* fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup"

A free-tier identity carries $0 by design, so the portal seed reports
`paid_access=False` for it. `is_free_tier_model` did not know the welcome
host, read that as a depleted account, and every free-tier turn ended with
the credits-depleted notice telling the user to top up an account they do
not have.

Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host
(`anon_auth.route_is_welcome_host`) is the free tier. The host is the
evidence, not the model name: the paid inference host can serve
`nous/welcome` to a named account and that account's depletion is real, so
`("nous/welcome", <inference host>)` stays False. Local data only, like the
three rules above it.

Restores the two contracts dropped by hermes-magic 674e11d1eaa (the
prototype line ran without unit tests): the welcome host is free without
any pricing evidence; the model name alone is not. The first is red without
this fix.

* fix(copy): free-tier text stops promising a connector transfer and never names the config key

Sign-in copy on every surface said "Sign in to keep your connectors" and
ended with "Your connectors are kept." The transfer registry that would
make that true is empty (NS-821): nothing carries over today. The copy now
says what signing in does give ("unlock more models and tools") and the
completion line names the account, not a transfer. The docs page loses the
"connectors carry over" paragraph for the same reason.

The picker's off-state line exposed `nous.guest: false` and the word
"guest"; user copy names the free tier only (R-USR-1).

The docs page gains the pre-rollout note: until GA nothing on it happens
without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and
"replaced on next use" sentences now describe the boot bootstrap.

zh is a strict locale: the `freeTier` block was English placeholder text
copied from `en`; it is now Chinese. `connectorsKept` is renamed
`completedBody` since it no longer talks about connectors.

* feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn

The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1"
as on. Until now nothing in the desktop set it, so a packaged app could
never turn the free tier on, and a backend spawned by the app could
disagree with the app about whether the tier was live.

`electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled`
is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has
`--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch
into a module constant. `desktopBackendSpawnEnv` wraps every backend env
as the outermost call and writes the flag LAST, as "1" or an explicit "0",
so no earlier spread (`process.env`, `backend.env`) can resurrect a stray
value from the parent shell.

Stamped onto all three spawn sites: the primary `serve` spawn, the pooled
per-profile spawn, and the remote SSH `exec env ...` command (which gains
` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and
the backend probes are not backend spawns and do not get it: a
`hermes --tui` typed in the pane must not mint.

The renderer learns the same fact read-only through the existing
`hermes:launch-flags` sync IPC (`guestOnboarding`) and preload
(`window.hermesDesktop.guestOnboardingEnabled`).

Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding`
maps to it in main). Two invariant tests on the pure helpers: only "1" or
the argv flag enables; the spawn env carries "1"/"0" as the last word and
preserves every other key.

* feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll

The backend's boot bootstrap now announces `setup.ready` once, after it has
created (or refused) the free-tier identity and resolved the inference
route. The renderer used to discover both by polling `setup.status`,
`setup.runtime_check` and `free_tier.status` every 60 s from
`useStatusSnapshot`; a fresh install's chip, notice strip and onboarding
overlay could sit stale for up to a minute after boot, and three RPCs a
minute per window kept asking a question whose answer changes only at
boundaries the backend already announces.

`handleLifecycleEvent` routes `setup.ready` (active source only, like
`skin.changed`) to `notifySetupReady()`, a one-shot tick atom in
`live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens
to it and runs one readiness round at once (`setup.status` +
`setup.runtime_check` + `free_tier.status`). The readiness legs also run
once on open and on return from another app, as today. The 60 s tick keeps
only `getStatus()`.

`SetupStatusSnapshot` types the record's additive fields (`ready`,
`free_tier`, `other_providers`, `inference_provider`); readiness semantics
are unchanged and still key on `provider_configured` + `runtime_check`.

Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one
refresh from the active source and none from another; the snapshot hook's
contract is three legs on open, one leg on the tick.

* fix(cli): the banner names the free tier's model instead of "no model configured"

The welcome banner prints before credentials resolve, so on a fresh install
`model` is empty and the banner said, in red, "no model configured — run
/model or hermes setup". Under the free tier that is false: the route is
already known from local state (identity on disk, tier on), and the first
message will run on `nous/welcome`.

`_banner_left_lines` now asks the route the same question when `model` is
empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous
Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the
banner's 'no model configured' line reads the resolved route").

Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`;
gate off -> the red line, zero portal calls.

* fix(aux): vision on the free tier uses nous/welcome too

The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the
backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the
repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the
welcome host would have sent every image step past the free tier for no reason, so the
auxiliary client pins the route's one model for every lane. A backing model that takes
no images answers with the upstream's own error, which the ladder handles as it always has.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415)

* fix(gateway): hermes gateway run is a boot owner of the free tier too

Rung 5 made every demand-time free-tier site a read: resolve_provider,
the connector token, the /login precondition. That is only correct if
every process that can reach those sites ran the bootstrap first. The
CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone
messaging gateway did not. A fresh HERMES_HOME with the gate on and
`hermes gateway run` reached provider resolution with no identity to
consume, and /login returned Unavailable. Reported by @andrexibiza on
#107697 (P1).

GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an
executor thread right after startup recovery and BEFORE any adapter
connects, so a fast first DM cannot arrive with nothing to resolve. It
is its own step, not part of the turn-machinery warm-up: the warm-up is
an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0);
the bootstrap is correctness and must always run. With the gate unset it
is a local inventory and no network.

Live, real GatewayRunner.start against a fake portal in a fresh home:
  gate on   -> 1 create, identity persisted, resolve_runtime_provider=nous,
               /login precondition sees the identity
  gate off  -> 0 portal calls, no identity, no_provider_configured
Before the fix the gate-on row was identical to the gate-off row.

Test: the bootstrap seam runs before _start_prefilter_platforms and
delegates to the one creator. Red on 5554eb6993 (no seam), green here.

---------

Co-authored-by: Robin Fernandes <robin@soal.org>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 03:45:33 +05:30
Justin Bennington 8a6b5b67a7 fix(providers): block Actual at Responses send sites (E-1047) 2026-09-10 15:13:02 -04:00
Justin Bennington 8b2b83a906 fix(providers): pin all Actual routes to chat completions (E-1047) 2026-09-10 14:56:55 -04:00
Justin Bennington d7b0a72c2a fix(providers): route Actual through chat completions (E-1047) 2026-09-10 14:56:55 -04:00
Teknium ee35a4624f fix(routing): address review — session runtime decides, quarantine continues the chain, unclassified ids break ownership
Three defects found in review of the first head (@ehz0ah):

* The mid-request hop read the PERSISTED provider from disk. A live
  `/model xai-oauth` session over `model.provider: auto` therefore still
  reached the discovery chain and billed Nous. _try_payment_fallback now
  takes the route's main_runtime snapshot; disk is the fallback only when
  no session runtime exists.

* After a configured fallback was quarantined mid-request (401, refresh
  failed), the second pass went straight to the discovery chain, which
  the new gate refuses — so later CONFIGURED entries never ran and the
  original error was re-raised. The second pass now re-walks the task
  chain and main chain (the quarantined entry is unhealthy and skipped)
  before discovery.

* current_provider_owns_vendor dropped ids detect_vendor could not
  classify, so Bedrock's 15-id catalog (14 unclassified `us.anthropic…`)
  looked exclusively DeepSeek and `/model deepseek-v4-pro` stuck on
  Bedrock. An unclassified id now counts as evidence of a multi-vendor
  catalog: ownership requires every id to classify to the one vendor.
2026-09-10 09:11:00 -07:00
Teknium 7e0b5cd235 fix(auxiliary): auto never bills a provider the user did not select
With a main provider selected, an unusable main route (expired xAI/Codex
OAuth token, 401/402/429 mid-session) fell through the built-in discovery
chain (OpenRouter -> Nous -> custom -> api-key) and quietly ran every
compression, title and memory-flush call on whichever OTHER account was
still logged in. Reported as "using Grok on my Premium+ sub, my Nous
Portal balance kept draining" — the chat visibly stayed on Grok while the
side tasks were billed elsewhere, and re-logging into X did not help
because the aux side never consulted the selected provider.

The discovery chain is now reserved for installs with no selected main
provider (`model.provider: auto` / unset). Otherwise the ladder is
main -> auxiliary.<task>.fallback_chain -> fallback_providers -> refuse
with a warning naming the dead provider and the fix. Both entry points
gate on the same predicate: the resolve-time route and the mid-request
payment/auth hop (_try_payment_fallback).

Existing chain tests that asserted the hop now pin `provider=auto`, the
one case where discovery is still the contract.
2026-09-10 09:11:00 -07:00
Teknium edac5f6378 fix(auxiliary): origin-scoped key shield, task overrides over named providers, negotiation-only stream fallback
Follow-up to #106866 from the post-merge audits (@ehz0ah, @victor-kyriazakos):

- _recoverable_pool_provider: the "endpoint mismatch, not a dead key" shield now (a) compares the full
  origin via base_url_origin() (scheme+host+port — a different port or an HTTPS->HTTP downgrade is a
  different trust boundary) and (b) applies only when the rejected client carried the SESSION's key, so
  an independently owned auxiliary pool keeps rotating at its own configured origin.
- _resolve_named_custom_branch: explicit/per-task base_url and api_key compose OVER the providers.<name>
  entry field-by-field instead of the entry silently replacing them (URL-only, key-only, both, neither).
- _acreate_with_progress: plain-create fallback only for a rejected stream NEGOTIATION (mirrors the sync
  wrapper); a mid-stream failure propagates to the classified recovery ladder instead of re-sending the
  whole prompt non-streaming.

Tests (all red on main): foreign origin x {host, port, scheme} shields the session key; same origin and
independent aux pool still rotate; named-provider override matrix through resolve_provider_client on a
real temp HERMES_HOME; async content-then-error is not retried non-streaming.
2026-09-09 16:21:36 -07:00
Teknium 2db0c7a2d8 fix(auxiliary): compression sticks to the session's OpenAI endpoint; foreign-host rejections don't kill the key
`auxiliary.<task>.provider: openai` was rewritten to custom + api.openai.com/v1 unconditionally, and the
api-key discovery rung used the registry default endpoint, so proxy/gateway users (OPENAI_BASE_URL or
providers.openai) had compression hop to the public endpoint with a proxy-issued key -> 401 -> the
key/pool was quarantined and the session sat over the compression threshold.

- _expand_direct_api_alias: a providers.openai entry keeps its name (named-custom branch applies its
  base_url/key); otherwise OPENAI_BASE_URL wins over the public default.
- _resolve_api_key_provider: when the bound session runtime is that provider, use its endpoint + key.
- _recoverable_pool_provider: a rejection at a host other than the session's configured endpoint for
  the same provider is an endpoint mismatch, not a dead key -> no rotation / unhealthy mark on the pool.
2026-09-09 14:06:12 -07:00
Teknium e6b890e584 fix(auxiliary): async progress-hook wrapper + route the async primary through it (#98466)
Complete the seam default from the salvaged commit: _relay_async_completion needs an async twin of
_create_with_progress, and the async primary attempt previously streamed only for stream-only
providers, so a hooked async compression call ticked the watchdog zero times.

Adds one invariant test: with a hook installed, BOTH relay defaults stream and tick per chunk;
without a hook they are byte-identical plain creates. Red on origin/main.
2026-09-09 14:06:12 -07:00
Konstantin Khlopkov fba6110cde fix(agent): default relay completions to the progress-hook wrapper 2026-09-09 14:06:12 -07:00
kshitijk4poor 170a73637d fix(openai): drop prompt_cache_options from Astra requests — not an SDK kwarg, 30m is the server default
Every direct-API (api.openai.com) Astra request raised
``TypeError: Responses.create() got an unexpected keyword argument 'prompt_cache_options'``
before reaching the network: openai 2.24.0's Responses.create has no such parameter and no
**kwargs, and neither send path relocates it into extra_body. The PR's tests stopped at
build_kwargs/preflight so the SDK boundary was never crossed.

OpenAI's prompt-caching guide states ``prompt_cache_options.ttl`` accepts only ``30m`` and that
``30m`` is the default, so the field carried no information: sending nothing yields the same
cache lifetime. The sanitizer now only removes what the API rejects (none/minimal effort,
sampling/logprob knobs, the pre-5.6 ``prompt_cache_retention``) and never adds a field, which
also keeps the request body byte-stable for the cache prefix.

Also: none/minimal→low no longer needs a bespoke {"", "none", "disabled", "off"} set —
``clamp_effort`` against CODEX_ASTRA_EFFORTS already resolves to the floor (``low``); and the
auxiliary adapter derives ``is_codex_backend`` from ``classify_responses_route`` (the declared
single owner of that predicate) instead of re-implementing the host test inline.

Tests reshaped to contracts: the two proxy/subdomain cases collapse into one parametrised
"exact host only" test asserting effort and temperature pass through untouched.
2026-09-07 21:43:54 +05:30
Eva 5a82258626 fix(openai): keep Astra rules on eligible routes
(cherry picked from commit f92fb9d9964e6e8ea118035c548aae812054f504)
2026-09-07 21:43:54 +05:30
Eva c990017482 test(openai): close Astra baseline review gaps
(cherry picked from commit 8c27b7c9316ba675e4d9ebcc7a659754f37f18ef)
2026-09-07 21:43:54 +05:30
Teknium 37fb7adfd6 fix: structured reasoning no longer breaks chat consumers
Normalize incoming reasoning at the shared heading boundary and completed
extraction, and flatten auxiliary content and reasoning before accumulation.
Reuse the existing text flattener with no implicit fragment separators.

Combine the earliest related work from zsuroy (#85791), the diagnosis and
patch from 2025hcsmile2010-hue (#104711, #104848), and completed extraction
work from liuhao1024 (#104717) as a slim redo, not a verbatim cherry-pick.

Two invariant tests exercise the real SDK and local HTTP fixture across
main streaming, Relay collection, auxiliary sync/async and completed output.
The standalone matrix improves from 32/84 to 84/84, preserving answers.

Co-authored-by: suroy <suroy@qq.com>
Co-authored-by: 2025hcsmile2010-hue <2025hcsmile2010@gmail.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-07 08:24:54 -07:00
Teknium bdf45abd7c fix: prefer owned Anthropic grants and bind auxiliary refresh to request credentials
Port owned-before-borrowed resolution from #104624, crediting the root cause in #104622. Include the synchronous and asynchronous auxiliary fallback recovery sites: forward the failed request key so an unrelated borrowed login never owns that refresh. Local-wire probes preserve the borrowed file and exchange only the owned grant. Broader validation remains queued; do not treat this commit as ready.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>

Co-authored-by: d-bow-dev <24577047+dwb1991@users.noreply.github.com>
2026-09-07 08:07:26 -07:00
Teknium cb1a42d33b refactor: isolate custom health identity and trim redundant alias tests 2026-09-07 07:06:58 -07:00
fangliquanflq 84bbf176a3 fix(agent): preserve route URL path identity 2026-09-07 07:06:58 -07:00
fangliquanflq e3ff3e78d1 fix(agent): quarantine failed fallback destination 2026-09-07 07:06:58 -07:00
fangliquanflq e033bbaa3c fix(agent): quarantine bare custom aliases by endpoint 2026-09-07 07:06:58 -07:00
fangliquanflq 3fcbf13ddb fix(agent): scope named custom health by endpoint 2026-09-07 07:06:58 -07:00