Commit Graph

297 Commits

Author SHA1 Message Date
kshitijk4poor 476a45f4f3 test(agent): one real-transport test for the codex summary tool controls
Replace the two stubbed tests with a single test that drives the real
CodexTransport.build_kwargs (the fixture agent already carries a tool), so
the test binds the summary body actually sent, not a hand-written dict. It
covers both the first attempt and the empty-summary retry; the previous
pair asserted the same three keys twice and reused one dict across
attempts, so the second-attempt check passed trivially after the first pop.

Add the WHY comment on the pops: the transport emits tools, tool_choice and
parallel_tool_calls as one block, and strict Responses backends 400 on the
controls without tools.
2026-09-12 22:40:28 +05:30
Julien Talbot 16695bfa07 fix(codex): strip tool controls from summary calls
(cherry picked from commit b058c25740b4a4bfa40ffbe2ce7bcddb9fe930dd)

Co-authored-by: Matthieu Talbot <1246794+MartyLake@users.noreply.github.com>
2026-09-12 22:40:28 +05:30
teknium1 10864f394d fix(agent): stream_options compatibility retry does not consume the transient budget
Review finding on #109143: `_handle_stream_error` returned True for the
stream_options rejection without checking whether `_call` had another
iteration. With HERMES_STREAM_RETRIES=0, or after the transient budget was
spent, the loop ended with neither a response nor an error set and the call
returned None instead of raising.

The compatibility retry now extends the loop by one attempt exactly once
(`_compat_retries`); the transient budget is untouched. Test pinned with
HERMES_STREAM_RETRIES=0 (red on the previous head).
2026-09-12 08:25:40 -07:00
DavidMetcalfe 6527af2286 fix(agent): retry once without stream_options when an endpoint rejects it (HTTP 400/422)
Azure AI Foundry serverless (MaaS) endpoints validate the request body
strictly and reject `stream_options: {"include_usage": true}` with 422
`extra_forbidden`. Hermes sent the field on every streaming call, so the
agent was unusable against that endpoint family and the fallback chain
failed too.

When a 400/422 names `stream_options` as an extra/unsupported field and no
delta has been delivered yet, re-open the stream without the field and
remember the rejection on the agent (`_stream_options_unsupported`) so later
turns skip it up front. Streaming itself stays on — this is not the
"stream not supported" case. Usage accounting for such endpoints falls back
to the estimator, as it already does for native Gemini.

Salvage of PR #53271 by @DavidMetcalfe, reshaped onto the `_StreamingCall`
retry loop with per-agent state instead of a process-wide host set; one
end-to-end invariant test through `_interruptible_streaming_api_call`.

Fixes #9705
2026-09-12 08:25:40 -07:00
Siddharth Balyan 4bdd64b334 The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract

The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an
anonymous account exactly one model on its own host, refuses everything else with a structured
429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and
tells a signed-in account that still asks for `nous/welcome` what to switch to in an
`x-nous-model-switch` header. Four client-side gaps against that contract:

- Auxiliary calls were refused on every session. The auxiliary client asked the welcome host
  for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free`
  before each fallback. On the welcome host it now uses `nous/welcome` (its backing model
  covers auxiliary work) and skips Nous for vision, which the welcome model does not take.

- The structured 429 body was never read. The classifier now parses `reason` /
  `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are
  non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited`
  are rate limits that honour `retry_after` and never rotate the free tier's only credential.
  The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and
  fall back instead of retrying or re-exchanging. The terminal paths say what happened and
  name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal).

- The `x-nous-model-switch` header was ignored. The chat-completions transport records it
  beside the rate-limit and credits headers; the next call moves the session, and the config
  default when it still names `nous/welcome`, to the backing model the gateway named.

- A guest fell back to the paid host. With `inference_base_url` absent from the exchange or
  outside the host allowlist, routing defaulted to inference-api, where every request is a
  400. A guest now defaults to the welcome literal at the exchange, in the shared store's
  shape, and in effective routing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8)

* feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use

A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided
setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier
picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime
rung now sets it up there instead of failing "not logged in", so the guided chat no longer races
the root profile's first-run mint.

nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever
nothing else is configured; "on-request" mints only when the free tier is asked for by name
(nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers —
the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the
connector token path — still adopt what the shared store holds, so every profile follows the one
identity the guided setup created, but never create one on their own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c)
(cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165)

* feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit

Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity
(nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is
configured. "explicit" means Hermes never creates one on its own: the only creator is the new
provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup
on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and
before the guided chat exists — so the identity lands in the root store every profile reads
through and is there before any session asks for nous/welcome. That closes the race against the
backend's own setup, and makes "only when the setup-bot flow is used" literally true.

The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome
(the hermes model row, --provider nous), which treated a model name as intent and was broader
than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not
logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to
sign in from. Implicit callers still adopt an identity the shared store holds, and a retired
credential is replaced (a continuation, not a creation).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0)

* fix(auth): remove the nous.guest_setup knob; the free tier is created on first use

`nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its
default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`,
unknown values read as `auto`, and under `explicit` a CLI-only install could never get an
identity, which contradicts the first-run contract (first command mints, then chats).

The mint race the knob accompanied is already benign: every caller takes the profile lock then
the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup
win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which
stays. `nous.guest` remains the only free-tier policy.

Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through
`ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call
sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config
entries. The three policy tests that hold regardless of the knob are kept under
`TestExplicitProvision`; the two that only tested the knob are deleted.

(cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261)

* fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent

When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway
serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch`
moves the live session to that model and moves `config.yaml`'s default off the alias in the
same step. The messaging gateway's fallback-eviction check compares the agent's model with the
config default and evicts on any mismatch that is not a /model override, so when the config
write did not land (unreadable config, lock) the cached agent was evicted once per turn, and
prompt caching with it.

`apply_model_switch` now stamps the alias it moved the session off on the agent, and
`_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as
deliberate, beside the existing /model override case. The check takes the agent and the config
model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them.

(cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954)

* fix(auth): the free tier outranks implicit host credentials in provider resolution

On a fresh install with a leftover ~/.aws profile, resolve_provider("auto")
reached the Bedrock rung before the free-tier rung, so the first turn ran on
Bedrock and failed 403 while the free tier was still being minted in the
background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s,
three retries, no answer; the next process then switched to nous/welcome.

The free-tier rung now sits directly above the Bedrock chain: when nous.guest
is on, an existing free-tier identity answers, else a blocking mint runs, and
only then does the boto chain get a say. Everything above is unchanged and
still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter
pool, a logged-in active_provider. nous.guest: false skips the rung, and a
failed mint still falls through to Bedrock and the no-provider guidance.

Tests: six precedence cases (identity present, fresh mint, free tier off, env
key still wins, sign-in still wins, failed mint falls through). The opt-out
test now neutralizes the AWS chain like the precedence tests do; on a machine
with ~/.aws it was failing for the same reason as the bug.

Live after the fix, same Mac, AWS credentials visible, isolated shared store:
identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in
11 s.

(cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a)

* fix(auth): review follow-ups for the free-tier rung (NS-829)

- tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches
  the free tier off; its contract is the boto chain, and the free tier now
  sits above it.
- gateway/run_notifications.py: the free-tier startup line reads auth.json
  before consulting the resolver, so a gateway boot on a machine with AWS
  credentials never mints or refreshes over the network.
- hermes_cli/anon_auth.py: module docstring says where the free tier sits in
  the ladder instead of "the ladder is untouched".
- tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of
  six (parametrized ladder cases; a failed mint that returns None or raises
  falls through to Bedrock).

scripts/run_tests.sh on the five affected files: 147 passed, 0 failed.

(cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11)

* feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone

The free tier is pre-GA. Until GA it must not exist for anyone who did not
ask for it: no identity minted, no portal traffic, no free-tier copy on any
surface. One environment variable now decides that, and one function reads it.

`guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly
"1"; only then does `nous.guest` (the user's off switch) get consulted. Every
free-tier site already funnels through `guest_enabled()`, so the gate closes
minting, routing, connector entitlement, status lines and the picker row in
one place. With the variable unset, `resolve_provider("auto")` on a fresh
install raises `no_provider_configured` exactly as upstream does.

`HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the
gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted
identities as a side effect of provider resolution, and `_has_any_provider_
configured` read them ahead of every other check, making the CLI a second
reader of a flag that must have exactly one. `_forced_new_done` and the
`force` parameter of `_reconcile_and_provision` go with them.

Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in
b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847).

Not a user preference: the variable is never written to config.yaml or .env
and never shown in setup. It is deleted at GA together with its comment in
anon_auth.py. This is a deliberate, temporary exception to the "no new
HERMES_* env vars for non-secret config" rule.

Tests: fixtures set the gate instead of deleting the old lever; one new
invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "",
"0", "true" and "new" all leave the tier off with zero portal calls, red on the
previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the
feature.

* feat(auth): the free-tier identity is created in one place, at boot; every other site is a read

Before this commit eight sites could create a Nous free-tier identity as a
side effect of something else: resolving a provider, the CLI's first-run
check, the CLI's session setup (in the background beside an own key), a
connector bearer read, the desktop polling `free_tier.status`, the sign-in
precondition, the desktop's `free_tier.provision`, and the dead-credential
re-mint. A poll could mint. Provider resolution could hit the network. Two
of them raced each other on a fresh install.

Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator.
`hermes serve` runs it on a daemon thread from `_lifespan` beside the other
background boots; `cmd_chat` runs it synchronously before the first-run
guard. It inventories credentials first (`resolve_provider("auto",
skip_free_tier=True)`: what would carry inference if the free tier did not
exist), creates the identity only when `guest_enabled()`, resolves inference,
records a `SetupRecord` in process memory and broadcasts ONE `setup.ready`
event. It runs on every boot; only the mint is gated.

`ensure_portal_identity` now requires `explicit=True` and raises otherwise.
Its callers are the bootstrap, the desktop's `free_tier.provision` (the
explicit retry when the boot could not create the identity) and the two
dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`,
`managed_tool_gateway._replace_dead_guest_token`). The background thread
path and `provision_free_tier` are deleted with their last callers.

Reads that used to mint and now only read: `auth.py::resolve_provider`
rung 7 (an existing identity still outranks the Bedrock chain, NS-829
ordering kept), `main.py::_has_any_provider_configured`,
`cli_agent_setup_mixin._ensure_runtime_credentials`,
`managed_tool_gateway.read_nous_access_token` (no identity -> None),
`anon_sign_in.run_sign_in` (no identity -> Unavailable),
`methods_free_tier` `free_tier.status`.

`setup.status` answers from the record for the launch profile, blocking up
to 8 s while the bootstrap is in flight so a client's first poll lands after
the identity exists rather than racing it; a named profile, or a process
that never ran the bootstrap, keeps today's live probe. The record's fields
ride along additively (`ready`, `free_tier`, `other_providers`,
`inference_provider`).

Identity and inference are decoupled (NS-845 Q1.3): the mint sets
`active_provider="nous"` only when the inventory found nothing else usable
(`_mint_locked(carries_inference=)`); an adopted account always does. A token
refresh no longer re-elects the provider it refreshed
(`_save_provider_state_to_source` writes credentials, not the user's
choice) — that write was how an own-key install ended up on the free tier
after the first connector call.

Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check),
bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint),
62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and
`provision_free_tier`), and a04b05260c (blocking mint in the resolver).
Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847.

Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps
inference; reads never reach the portal; a refused mint is memoised),
`free_tier.status` fails loudly if it ever calls the creator, the resolver
stub fails loudly if resolution ever mints, `setup.status` reads the record,
`skip_free_tier` proves the inventory question. The three sign-in tests for
the deleted pre-mint collapse into one (`no identity -> Unavailable, zero
portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and
off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI
matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20.

* fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup"

A free-tier identity carries $0 by design, so the portal seed reports
`paid_access=False` for it. `is_free_tier_model` did not know the welcome
host, read that as a depleted account, and every free-tier turn ended with
the credits-depleted notice telling the user to top up an account they do
not have.

Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host
(`anon_auth.route_is_welcome_host`) is the free tier. The host is the
evidence, not the model name: the paid inference host can serve
`nous/welcome` to a named account and that account's depletion is real, so
`("nous/welcome", <inference host>)` stays False. Local data only, like the
three rules above it.

Restores the two contracts dropped by hermes-magic 674e11d1eaa (the
prototype line ran without unit tests): the welcome host is free without
any pricing evidence; the model name alone is not. The first is red without
this fix.

* fix(copy): free-tier text stops promising a connector transfer and never names the config key

Sign-in copy on every surface said "Sign in to keep your connectors" and
ended with "Your connectors are kept." The transfer registry that would
make that true is empty (NS-821): nothing carries over today. The copy now
says what signing in does give ("unlock more models and tools") and the
completion line names the account, not a transfer. The docs page loses the
"connectors carry over" paragraph for the same reason.

The picker's off-state line exposed `nous.guest: false` and the word
"guest"; user copy names the free tier only (R-USR-1).

The docs page gains the pre-rollout note: until GA nothing on it happens
without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and
"replaced on next use" sentences now describe the boot bootstrap.

zh is a strict locale: the `freeTier` block was English placeholder text
copied from `en`; it is now Chinese. `connectorsKept` is renamed
`completedBody` since it no longer talks about connectors.

* feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn

The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1"
as on. Until now nothing in the desktop set it, so a packaged app could
never turn the free tier on, and a backend spawned by the app could
disagree with the app about whether the tier was live.

`electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled`
is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has
`--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch
into a module constant. `desktopBackendSpawnEnv` wraps every backend env
as the outermost call and writes the flag LAST, as "1" or an explicit "0",
so no earlier spread (`process.env`, `backend.env`) can resurrect a stray
value from the parent shell.

Stamped onto all three spawn sites: the primary `serve` spawn, the pooled
per-profile spawn, and the remote SSH `exec env ...` command (which gains
` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and
the backend probes are not backend spawns and do not get it: a
`hermes --tui` typed in the pane must not mint.

The renderer learns the same fact read-only through the existing
`hermes:launch-flags` sync IPC (`guestOnboarding`) and preload
(`window.hermesDesktop.guestOnboardingEnabled`).

Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding`
maps to it in main). Two invariant tests on the pure helpers: only "1" or
the argv flag enables; the spawn env carries "1"/"0" as the last word and
preserves every other key.

* feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll

The backend's boot bootstrap now announces `setup.ready` once, after it has
created (or refused) the free-tier identity and resolved the inference
route. The renderer used to discover both by polling `setup.status`,
`setup.runtime_check` and `free_tier.status` every 60 s from
`useStatusSnapshot`; a fresh install's chip, notice strip and onboarding
overlay could sit stale for up to a minute after boot, and three RPCs a
minute per window kept asking a question whose answer changes only at
boundaries the backend already announces.

`handleLifecycleEvent` routes `setup.ready` (active source only, like
`skin.changed`) to `notifySetupReady()`, a one-shot tick atom in
`live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens
to it and runs one readiness round at once (`setup.status` +
`setup.runtime_check` + `free_tier.status`). The readiness legs also run
once on open and on return from another app, as today. The 60 s tick keeps
only `getStatus()`.

`SetupStatusSnapshot` types the record's additive fields (`ready`,
`free_tier`, `other_providers`, `inference_provider`); readiness semantics
are unchanged and still key on `provider_configured` + `runtime_check`.

Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one
refresh from the active source and none from another; the snapshot hook's
contract is three legs on open, one leg on the tick.

* fix(cli): the banner names the free tier's model instead of "no model configured"

The welcome banner prints before credentials resolve, so on a fresh install
`model` is empty and the banner said, in red, "no model configured — run
/model or hermes setup". Under the free tier that is false: the route is
already known from local state (identity on disk, tier on), and the first
message will run on `nous/welcome`.

`_banner_left_lines` now asks the route the same question when `model` is
empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous
Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the
banner's 'no model configured' line reads the resolved route").

Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`;
gate off -> the red line, zero portal calls.

* fix(aux): vision on the free tier uses nous/welcome too

The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the
backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the
repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the
welcome host would have sent every image step past the free tier for no reason, so the
auxiliary client pins the route's one model for every lane. A backing model that takes
no images answers with the upstream's own error, which the ladder handles as it always has.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415)

* fix(gateway): hermes gateway run is a boot owner of the free tier too

Rung 5 made every demand-time free-tier site a read: resolve_provider,
the connector token, the /login precondition. That is only correct if
every process that can reach those sites ran the bootstrap first. The
CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone
messaging gateway did not. A fresh HERMES_HOME with the gate on and
`hermes gateway run` reached provider resolution with no identity to
consume, and /login returned Unavailable. Reported by @andrexibiza on
#107697 (P1).

GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an
executor thread right after startup recovery and BEFORE any adapter
connects, so a fast first DM cannot arrive with nothing to resolve. It
is its own step, not part of the turn-machinery warm-up: the warm-up is
an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0);
the bootstrap is correctness and must always run. With the gate unset it
is a local inventory and no network.

Live, real GatewayRunner.start against a fake portal in a fresh home:
  gate on   -> 1 create, identity persisted, resolve_runtime_provider=nous,
               /login precondition sees the identity
  gate off  -> 0 portal calls, no identity, no_provider_configured
Before the fix the gate-on row was identical to the gate-off row.

Test: the bootstrap seam runs before _start_prefilter_platforms and
delegates to the one creator. Red on 5554eb6993 (no seam), green here.

---------

Co-authored-by: Robin Fernandes <robin@soal.org>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 03:45:33 +05:30
Justin Bennington 8b2b83a906 fix(providers): pin all Actual routes to chat completions (E-1047) 2026-09-10 14:56:55 -04:00
teknium1 6c89cb032e refactor(agent): move the entitlement rejection marker into agent/fallback_cooldown
The two facades post-#102117 (chat_completion_helpers, agent_runtime_helpers) must not grow
behaviour; the marker/predicate belong with the sibling that already owns primary-cooldown
state shared by the fallback walk and restore_primary_runtime. Also drops the two
try/except-return-False wrappers around pool.entries() and normalize_model_for_provider
(neither raises for a live pool / the never-raising normalizer).

Tests trimmed to the two invariants (#106475): the marker fires only on the single-credential
Codex entitlement 400 and the walk skips the rejected slug in either form; restore is gated on
a rejected primary and still restores an unrelated one.
2026-09-09 11:29:56 -07:00
liuhao1024 808c5b30aa fix(agent): fail closed when Codex account-model entitlement 400s exhaust the chain
A Codex ChatGPT-account 400 ('The X model is not supported when using Codex
with a ChatGPT account.') names the model, so with a single credential the
slug is dead for that account. The fallback walk still re-selected it and
restore_primary_runtime switched back to the primary at the start of every
turn, announcing an unverified 'Primary model restored' — the two warnings
alternated forever with zero delivered answers (#106475).

Record the rejected (provider, model) pair on the non-retryable client-error
path (only when no multi-credential pool exists — rotation covers that case,
#71970), skip rejected entries during the fallback walk, and gate
restore_primary_runtime on the primary's slug so the session fails closed
with the terminal entitlement error instead of oscillating. Fixes #106475.

(cherry picked from commit 471435b10288f15387b2549a8b02551d82c0f670)
2026-09-09 11:29:56 -07:00
kshitijk4poor 9e6c4100cb fix(agent): close the interrupted tool tail on the overflow terminal
The overflow-terminal path ends the turn without reaching finalize_turn, so
a transcript that overflowed right after a tool batch ended on a raw tool
result; strict providers reject the next user turn (tool -> user). Close it
with the same final text, mirroring the truncated-tool-call terminal above.

Also: classify once before either log so an overflow no longer emits a
"so the loop can continue" WARNING followed by the contradicting "NOT
seeding" one; reset the stale-streak breaker once for both branches; drop
the "compression could not recover it" wording (this path never reached
compression); trim the test file to the three tests that bind behaviour
(stream -> terminal stub; 413 stays non-terminal; the terminal ends the
turn, closes the tool tail, carries compression_exhausted).
2026-09-09 17:45:12 +05:30
ca-shrimp 3b0e81459b fix(agent): carry compression_exhausted bit; scope overflow terminal to context_overflow
Address review P1s (andrexibiza) on #106266:

1. The overflow-terminal exit in recover_from_truncation now forwards the
   #98722 typed compression_exhausted bit (partial_result/end_turn gained the
   flag) so the gateway resets/moves future input to a clean session instead
   of leaving the bloated durable session authoritative for the next turn.

2. _overflow_terminal is scoped to FailoverReason.context_overflow ONLY.
   payload_too_large (413) has its own byte-scored recovery owner
   (turn_overflow._recover_payload_too_large, #88960/#47339) that must not be
   bypassed; a post-delta 413 keeps its normal continuation stub. Regression
   covers both lanes.

Tests: unit asserts result compression_exhausted=True on the marker; a
real streamed partial hitting a 413 payload-too-large error keeps content and
is not terminal. 50 streaming/continuation/gateway regressions pass.
2026-09-09 17:45:12 +05:30
ca-shrimp ece584b8f5 fix(agent): don't seed continuation stub after a context-overflow stream death
When a stream delivered text and then died on a context-overflow /
payload-too-large error, the partial content (often tens of KB) was seeded as
a length-continuation stub, growing the transcript monotonically. In a
session whose transcript cannot be compressed back under budget
(protect_last_n covers everything -> no_progress, or the summary would
itself be larger -> would_grow), every later request is larger than the one
that just failed — an unrecoverable loop where the user sees a 30+ minute
fake hang and the only remedy is killing the session (#106260).

classify_api_error already labels these errors context_overflow /
payload_too_large (should_compress=True). _partial_stream_stub now returns
an EMPTY stub marked _overflow_terminal for that class instead of seeding
the recovered text, and recover_from_truncation treats the marker as
terminal: the turn ends via the recovery contract with a clear message
(start /new) and the transcript is not polluted with the partial.

Normal partials (network stall, output-cap truncation, tool-call drops) are
unchanged — only the overflow error class changes behavior.

Tests: stub marker + empty content; a real streamed partial hitting a
'maximum context length' error returns the terminal stub; recover_from_
truncation ends the turn (no fragment/nudge appended) on the marker while a
normal stub still runs the continuation path. 67 streaming/continuation
regressions pass.
2026-09-09 17:45:12 +05:30
Teknium fc70d05bb1 fix: arm primary cooldown once while walking fallback candidates 2026-09-07 07:08:25 -07:00
Teknium 68b16aa7e3 fix: retain exhausted-chain cooldown classification after extraction 2026-09-07 07:08:25 -07:00
Teknium fb2f66f586 fix: describe remaining retry eligibility without promising recovery 2026-09-07 07:08:25 -07:00
Halldrix ae4f777c69 fix(agent): surface armed rate-limit cooldown in fallback notice (#104120)
_arm_rate_limit_cooldown now returns the armed backoff seconds so
try_activate_fallback can append them to the user-facing notice
(Primary retried in ~N min/h) instead of discarding the one number
that decides whether the user waits or re-plans. Duration, never
wall-clock: the stored deadline is monotonic-based. Non-rate-limit
reasons and chain-switches from an active fallback arm nothing, so
no suffix is printed there.
2026-09-07 07:08:25 -07:00
Teknium fd3565deec fix: remove dedicated user-facing output cap controls 2026-09-07 06:15:43 -07:00
Teknium d0081d9597 refactor: colocate fallback handoff with client lifecycle 2026-09-07 06:07:16 -07:00
Teknium ba395af50f refactor(agent): isolate streaming wait monitor phase 2026-09-07 02:37:13 -07:00
Teknium 47d73afeb1 refactor(agent): isolate non-stream request polling 2026-09-07 02:37:13 -07:00
liuhao1024 3c9c225425 fix(agent): tolerate schema-valued "type" keys in the watchdog payload estimator (#104793)
The image-part membership test in _payload_chars() ran on every dict in
the request, so a tool JSON Schema with a parameter named "type" (its
value is the sub-schema dict) or a multi-type value like
["string", "null"] raised TypeError: unhashable type before the
request reached the provider. Guard the membership test with
isinstance(str): structured "type" values are payload data, not
content parts, and keep the legacy walk for them.
2026-09-07 00:49:51 -07:00
crdesign8 922c0d670c fix(agent): stale-call watchdog prices images at the learned per-image cost, not base64 length (salvage #76471)
estimate_request_context_tokens (the stale-stream / non-stream watchdog and Codex TTFB floor
estimator) did len(str(payload)) // 4 over the wire payload, so one native screenshot read as
~100K tokens and selected the 1200s giant-conversation floor while the provider's real prompt
was ~2K (#63871, #76411). Image content parts on both wire shapes (Chat image_url / Responses
input_image) now cost the per-image price learned from provider usage (agent/image_token_cost,
bound per turn; worker threads inherit it via copy_context); base64-looking STRINGS stay text
and tool schemas mentioning image are not images.

A/B, one 300KB screenshot: main estimate 100,545 -> codex floor 1200s / stale 300s; now 2,025
-> floor 0s / stale 180s.

Design and boundary (structured parts only, tools/instructions opaque, plain data-URL strings
stay text) from #76471 by @crdesign8; re-authored to reuse the learned cost instead of a second
fixed constant, and trimmed to 2 invariant tests.
2026-09-06 22:34:20 -07:00
sylvainCDA 693641aa8b fix(chat-completions): strip name from tool-result messages for strict providers
The Chat Completions schema has no `name` field on `role: tool` (only on
the long-removed `role: function`), but Hermes carries the tool name over
onto the result message. Permissive providers ignore it; strict ones
(aki.io) reject the whole payload with `contains item with unknown key
name`, which breaks every tool call in the session.

Follow-up to the review on #51365:

- The strip now goes through the copy-on-write `mutable_msg()` path in
  `convert_messages()`, preserving the identity/copy-on-write contract
  instead of mutating `msg` in place.
- `handle_max_iterations()` hand-builds its summary payload and calls
  `chat.completions.create()` directly, bypassing the transport, so it
  leaked `name` even with the transport fixed. It now mirrors the same
  role-qualified removal, next to the existing tool_name/codex_*/timestamp
  strips.

The removal is role-qualified: `name` stays on user/assistant messages,
where it is schema-valid.

Regression coverage on both paths; both tests fail without the fix.
2026-09-07 02:50:45 +05:30
Teknium 6d2645e64e feat(logging): the API call line carries cache write count, the provider's response id, and the serving upstream
Diagnosing the 1,393-agent run's cache misses took a state.db join and a live
probe because none of the three were on the one line we log per call:

  write=<n>    cache_creation tokens; a write costs 50x a read, so this is the
               money, and "read stuck, write large" on consecutive lines is a
               routing miss visible without a probe
  id=<id>      the provider's response id (Anthropic msg_..., OpenRouter/Nous
               gen-...): what a provider needs to look a request up
  upstream=<n> who actually served, when the route reports it (OpenRouter's
               `provider`); how we learned GMI was not serving

Fields are appended to the existing line and omitted when absent, so every
existing parser (evals/postmortem, the two cache_prefix probes) keeps
matching; the forensics parser reads them when present.

The streamed chat path fabricated `id="stream-<uuid>"` and dropped the
chunks' id/provider on the floor; it now keeps the first chunk's id and
provider, falling back to the fabricated id only when the stream never sent
one. Nothing depended on the prefix.

Live (Nous, Fable 5.1, both wires): chat -> `write=3759
id=gen-1788727882-... upstream=Anthropic`; native -> `id=gen-1788727891-...`
(the Anthropic Message object's id was already real). Tests: fields present
and omitted, prefix unchanged; forensics parser reads new and old lines.
2026-09-06 13:52:15 -07:00
kshitijk4poor 3ac671dbba refactor(agent): drop the write-only agent-level codex watchdog timestamps
With a request-local watchdog state on every codex request, agent._codex_stream_last_event_ts / _last_progress_ts had no remaining reader; the state-None snapshot only runs for non-codex requests where no watchdog consults it.
2026-09-07 00:41:09 +05:30
kshitijk4poor 1ad8b5dd4a refactor(agent): trim the watchdog phase split to what the tests pin
- Explicit-vs-implicit idle timeout detection uses env_float with a sentinel
  instead of a hand-rolled parse (unset and unparseable both mean implicit).
- _on_event writes the agent-level timestamps only when there is no
  request-local watchdog state; with state present they were dual bookkeeping
  nobody read (the snapshot prefers the state).
- _abort_request keeps main's close-then-retire order: the reorder had no
  test teeth and is a separate concern from #90449.
2026-09-07 00:41:09 +05:30
PeaceMaker-best 463d4bcd15 fix(agent): defer codex idle watchdog until progress
Signed-off-by: PeaceMaker-best <221849497+PeaceMaker-best@users.noreply.github.com>
2026-09-07 00:41:09 +05:30
Benjamin Brumbaugh cd71ee0708 fix(compression): defer local preflight after native checkpoint
A native Responses compaction checkpoint is opaque ciphertext; the rough
preflight estimator counts it as text (5.17M chars -> ~1.29M tokens against
a 204K trigger) and fires local compression on a request whose real prompt is
~116K. Arm the existing one-response real-usage latch when a replayable
checkpoint is captured (build_assistant_message) or restored into a fresh
agent (_hydrate_from_history), honor it in the post-tool gate and idle
compaction, and require non-empty encrypted_content for a checkpoint.

Squash of the author's source commits from #100642 (0e3c234ea0, 771e1b3365,
bb1505a119) plus the fdf140c81d test refresh, re-based onto current main by
patch application. Source delta is byte-identical to the PR head d6ce3e236d.

Fixes #100611
2026-09-06 09:09:00 -07:00
Teknium 12871bd01e feat(openrouter): per-model provider_routing.models.<id> overrides
`provider_routing.models.<model-id>` now takes the same only/ignore/order/sort/
require_parameters/data_collection keys and overlays the flat provider_routing
values whenever the agent is on that model. Resolution lives in the one
chokepoint every request path already uses (_provider_preferences_for_agent),
so CLI, gateway, TUI/Desktop, cron, /model switches, fallback activation and
delegated children on another model all honour it with no per-surface plumbing.
Matching is spelling-tolerant, sharing _canonical_model_variants with
agent.reasoning_overrides.

The OpenRouter profile's speed-tier pin no longer overwrites an explicit user
`only` on the BASE gpt-6-astra slug: the pin exists to keep default routing off
flex/fast, and a user pin is the stronger intent (only: [openai] stays [openai]
instead of becoming [openai, azure, azure/us]). Tier slugs (-fast/-flex) keep
owning `only`.

Live A/B (config only: {gpt-6-astra: [openai], claude-fable-5.1: [anthropic]}):
main sent {"sort":"price"} for fable and OpenRouter served it from Azure; with
this change it sends {"only":["anthropic"],"sort":"price"} and Anthropic serves it.

Schema proposed in #24495 (samplesabotage) and #100711 (Artemonim); this is a
slim chokepoint implementation of that design.

Co-authored-by: samplesabotage <samplesabotage@users.noreply.github.com>
2026-09-06 02:18:14 -07:00
RelaxJonh 4ee2c78324 fix(agent): Kimi Code fallbacks use the Anthropic Messages wire (#77256)
`try_activate_fallback` derived the fallback api_mode with its own host
heuristic (api.anthropic.com / trailing `/anthropic`) that lagged the
primary path's `_detect_api_mode_for_url()`: a `fallback_providers` entry
pointing at `api.kimi.com/coding` activated but POSTed OpenAI-shaped
`/chat/completions`, which that endpoint 404s, so every fallback turn
failed after the chain had correctly switched.

Route `_is_anthropic_wire_url` through `providers.host_mandated_api_mode`
— the single Messages-only host table the primary path already uses — so
the fallback path can never drift from it again. Exact-hostname match
rejects `api.kimi.com.attacker.test`.

Salvaged from #77304 by @RelaxJonh (original detection branch).
2026-09-05 15:54:14 +05:30
kshitijk4poor 74de0fd4fe fix(relay): build the chat-completions Relay response from collector-observed chunks
Relay invokes its finalizer as soon as the provider stream ends, concurrently with
Hermes' consumer thread, so any finalizer that reads the consumer loop's closures
races the last chunk. #103104 moved `usage` onto the collector; the same race still
truncated the final tool-call `arguments` delta, dropped the last content delta, and
downgraded `finish_reason` to "stop" in Relay's annotated response (reproduced ~50%
of runs). Every other Relay integration (Anthropic, Bedrock, Codex) already
rebuilds from collector-observed chunks; this makes chat_completions match via
`_RelayChatAccumulator` and removes the single-field nonlocal stopgap.

Test: the ordering harness is shared, and a second invariant test pins the
tool-call/finish_reason case (fails on the stopgap, passes here). Hermes' own
returned response was never affected.
2026-09-05 10:54:19 +05:30
mnajafian-nv 2cf5e68c5d fix(relay): retain usage from trailing chat completion chunks
Signed-off-by: mnajafian-nv <mnajafian@nvidia.com>
2026-09-05 10:54:19 +05:30
Teknium 912a497ce2 simplify(compat): tools/browser_tool — repoint run_agent cleanup_browser + chat_completion_helpers _is_headed_mode to the defining browser_tool_* siblings 2026-09-03 14:17:11 -07:00
Teknium 98c140bc4b simplify(compat): code_execution_tool/environments.local — drop 36 re-exports, repoint 5 callers + 14 test files 2026-09-03 13:24:04 -07:00
Teknium d179f28307 simplify(compat): anthropic_adapter — drop 30 re-exports + 1 alias, repoint 22 caller files (32 sites), 38 test files (~125 sites) 2026-09-03 13:16:47 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 022785a541 Merge origin/main (63279301bc): reasoning-mandatory 400 recovery folded into turn_recovery/error_classifier/models_reasoning_caps 2026-09-03 05:36:21 -07:00
Teknium 63279301bc fix(reasoning): retry after a mandatory-reasoning 400 resends the user's own effort
The retry must land on the same provider cache key as every prior request
in the session. Discard only the one-shot continuation disable and send
agent.reasoning_config verbatim; a config that is itself a disable is
omitted (that session never sent anything else, so nothing warm is lost).

Live: user effort=high, ephemeral disable → 400 → retry carries
{enabled: true, effort: high}.
2026-09-03 05:20:11 -07:00
Teknium f6bd1633f7 fix(reasoning): GLM-5.3 on Nous/OpenRouter no longer 400s when thinking is disabled
Reasoning-mandatory routes answer reasoning: {enabled: false} with HTTP 400
"Reasoning is mandatory for this endpoint and cannot be disabled". Hermes
sends that disable for /reasoning none, agent.reasoning_effort: none, and the
one-shot thinking-exhaustion continuation override (which GLM-5.3-flash
triggers on its own). The Nous profile's catalog guard swallows the disable
only when its per-process capability cache already says mandatory; a gateway
that warmed the cache before the route flipped kept sending it, and the 400
was classified as a non-retryable format_error that aborted the turn.

- error_classifier: new reasoning_mandatory reason (retryable, no fallback,
  no compression), matched before the request-validation branch.
- conversation_loop: one-shot recovery — set agent._reasoning_disable_rejected,
  queue a catalog refresh for the provider, retry.
- chat_completion_helpers: _reasoning_config_for_wire drops every disable
  (configured or ephemeral) once the route has rejected one.
- hermes_cli/models: refresh_reasoning_caps_async(provider) forces a
  background re-fetch of the Nous/OpenRouter catalog.
- openrouter profile: omit a disable when the catalog marks the route
  mandatory (parity with the Nous profile).

Live: z-ai/glm-5.3-flash on the Portal with a poisoned mandatory:false cache.
Before: turn aborted with the 400. After: one retry, thinking stays on, turn
completes.
2026-09-03 05:20:11 -07:00
Teknium 0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium 24a6b8d202 refactor(agent/chat_completion_helpers): fold silent-hang hint and exhausted-cooldown helpers 2026-09-02 23:35:58 -07:00
Teknium b2b0e34038 refactor(agent/chat_completion_helpers): fold fallback api-mode/credential branches and load-notice early returns 2026-09-02 23:34:14 -07:00
Teknium 21e50b7249 refactor(agent/chat_completion_helpers): share cloud stale-timeout derivation, Bedrock first-delta wrapper, summary text normalizer 2026-09-02 23:30:52 -07:00
Teknium 185a0c9037 refactor(agent/chat_completion_helpers): AST-identical — drop blank lines after in-function imports 2026-09-02 23:23:41 -07:00
Teknium e5758d1593 refactor(agent/chat_completion_helpers): unify Bedrock converse/converse_stream recovery, fold role ladder, reflow comments 2026-09-02 23:22:23 -07:00
fangliquan 390c27c6db fix(agent): accumulate streamed tool arguments linearly 2026-09-03 11:43:00 +05:30
Teknium b2ab212a57 refactor(agent/chat_completion_helpers): compact narrative docstrings (keep every invariant/why) 2026-09-02 23:11:00 -07:00
Teknium a341b7cc27 refactor(agent/chat_completion_helpers): move direct_api_call heartbeat/stale-timer lifecycle onto _InlineRequest (direct_api_call 103 -> 60 LOC) 2026-09-02 23:07:06 -07:00
Teknium afaf61c521 refactor(agent/chat_completion_helpers): share stream start/end bracket, lift chat-stream closures to methods, unify writer fence 2026-09-02 22:59:27 -07:00
Teknium 6ee2ed773e refactor(agent/chat_completion_helpers): collapse single-statement try/pass to suppress, fold single-use locals, share stale-base resolver 2026-09-02 22:47:34 -07:00
Teknium 10c023101c refactor(agent/chat_completion_helpers): AST-identical repack of wrapped calls/collections (3947 -> 3773 LOC) 2026-09-02 22:28:13 -07:00