Commit Graph

4669 Commits

Author SHA1 Message Date
Teknium 2b4deeb32b fix(auth): key sibling per-process credential memos by profile home under multiplex
Same class as the resolve_nous_access_token memo: three more process-wide
memos carried a credential resolved under one profile's HERMES_HOME override
into another profile's turn for their TTL.

- hermes_cli/nous_billing.py::_token_cache (30s (token, base) memo for the
  charge poll loop) was a single unkeyed slot -> dict keyed by
  hermes_home_key(); invalidate_cached_token() clears the dict.
- agent/moa_loop.py::_runtime_cache carried api_key/base_url/api_mode keyed
  only (provider, model) for 5 min -> (hermes_home_key(), provider, model).
- agent/auxiliary_client.py::_client_cache_key had no profile component, so
  callers that omit api_key (pool / Nous auth.json paths) could be handed a
  client built with another profile's bearer -> hermes_home_key() leads the key.

WHY hermes_home_key(): it reads the per-turn HERMES_HOME override the
multiplex gateway sets (falling back to the env var), and it is
symlink-stable, so the memo key is exactly the credential home the
resolution itself read from. Profiles stay independent islands; the
default-profile process env never leaks into a secondary's turn.

Tests: one invariant per site, proven red on origin/main.
2026-09-10 18:11:25 -07:00
Teknium 942973ae90 fix: every Bedrock client rebuild lands on the startup wire (Claude SDK, Converse region, guardrails)
Follow-up to 564aef2946 (Mantle SigV4 on /model). The same class of bug covered the
other two Bedrock wires and two more rebuild paths:

- Claude on Bedrock (anthropic_messages): startup builds an AnthropicBedrock SDK
  client (SigV4 via boto3). /model, fallback-to-Bedrock and fallback restore built a
  plain Anthropic client with api_key="aws-sdk" against bedrock-runtime → 401/403.
- Converse models (bedrock_converse: Nova, DeepSeek, Llama): only agent_init set
  _bedrock_region / _bedrock_guardrail_config. After a rebuild the transport fell
  back to us-east-1 and guardrail_config=None, so eu-/ap- users hit the wrong region
  and configured Guardrails silently dropped. switch_model also built a pointless
  OpenAI client against bedrock-runtime.

Introduce bedrock_adapter.bind_bedrock_runtime(agent, base_url, api_mode) as the one
place that puts an agent on a non-Mantle Bedrock wire, and call it from agent_init
(replacing the two duplicated bodies), _build_switched_client, _rebuild_primary_client
and _swap_fallback_clients. try_recover_primary_transport had a hand-copied version of
_rebuild_primary_client's ladder (and would have hit the same gap); it now calls the
shared helper. The region/guardrail parsing that lived in agent_init and
runtime_provider_backends is now bedrock_region_from_runtime_url /
bedrock_guardrail_config in the adapter.

Live probe on origin/main across switch_model, restore_primary_runtime and
_swap_fallback_clients for both wires: 0/6 correct before, 6/6 after.
2026-09-10 18:08:16 -07:00
Siddharth Balyan 4bdd64b334 The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract

The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an
anonymous account exactly one model on its own host, refuses everything else with a structured
429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and
tells a signed-in account that still asks for `nous/welcome` what to switch to in an
`x-nous-model-switch` header. Four client-side gaps against that contract:

- Auxiliary calls were refused on every session. The auxiliary client asked the welcome host
  for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free`
  before each fallback. On the welcome host it now uses `nous/welcome` (its backing model
  covers auxiliary work) and skips Nous for vision, which the welcome model does not take.

- The structured 429 body was never read. The classifier now parses `reason` /
  `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are
  non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited`
  are rate limits that honour `retry_after` and never rotate the free tier's only credential.
  The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and
  fall back instead of retrying or re-exchanging. The terminal paths say what happened and
  name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal).

- The `x-nous-model-switch` header was ignored. The chat-completions transport records it
  beside the rate-limit and credits headers; the next call moves the session, and the config
  default when it still names `nous/welcome`, to the backing model the gateway named.

- A guest fell back to the paid host. With `inference_base_url` absent from the exchange or
  outside the host allowlist, routing defaulted to inference-api, where every request is a
  400. A guest now defaults to the welcome literal at the exchange, in the shared store's
  shape, and in effective routing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8)

* feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use

A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided
setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier
picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime
rung now sets it up there instead of failing "not logged in", so the guided chat no longer races
the root profile's first-run mint.

nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever
nothing else is configured; "on-request" mints only when the free tier is asked for by name
(nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers —
the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the
connector token path — still adopt what the shared store holds, so every profile follows the one
identity the guided setup created, but never create one on their own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c)
(cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165)

* feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit

Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity
(nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is
configured. "explicit" means Hermes never creates one on its own: the only creator is the new
provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup
on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and
before the guided chat exists — so the identity lands in the root store every profile reads
through and is there before any session asks for nous/welcome. That closes the race against the
backend's own setup, and makes "only when the setup-bot flow is used" literally true.

The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome
(the hermes model row, --provider nous), which treated a model name as intent and was broader
than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not
logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to
sign in from. Implicit callers still adopt an identity the shared store holds, and a retired
credential is replaced (a continuation, not a creation).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0)

* fix(auth): remove the nous.guest_setup knob; the free tier is created on first use

`nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its
default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`,
unknown values read as `auto`, and under `explicit` a CLI-only install could never get an
identity, which contradicts the first-run contract (first command mints, then chats).

The mint race the knob accompanied is already benign: every caller takes the profile lock then
the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup
win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which
stays. `nous.guest` remains the only free-tier policy.

Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through
`ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call
sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config
entries. The three policy tests that hold regardless of the knob are kept under
`TestExplicitProvision`; the two that only tested the knob are deleted.

(cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261)

* fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent

When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway
serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch`
moves the live session to that model and moves `config.yaml`'s default off the alias in the
same step. The messaging gateway's fallback-eviction check compares the agent's model with the
config default and evicts on any mismatch that is not a /model override, so when the config
write did not land (unreadable config, lock) the cached agent was evicted once per turn, and
prompt caching with it.

`apply_model_switch` now stamps the alias it moved the session off on the agent, and
`_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as
deliberate, beside the existing /model override case. The check takes the agent and the config
model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them.

(cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954)

* fix(auth): the free tier outranks implicit host credentials in provider resolution

On a fresh install with a leftover ~/.aws profile, resolve_provider("auto")
reached the Bedrock rung before the free-tier rung, so the first turn ran on
Bedrock and failed 403 while the free tier was still being minted in the
background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s,
three retries, no answer; the next process then switched to nous/welcome.

The free-tier rung now sits directly above the Bedrock chain: when nous.guest
is on, an existing free-tier identity answers, else a blocking mint runs, and
only then does the boto chain get a say. Everything above is unchanged and
still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter
pool, a logged-in active_provider. nous.guest: false skips the rung, and a
failed mint still falls through to Bedrock and the no-provider guidance.

Tests: six precedence cases (identity present, fresh mint, free tier off, env
key still wins, sign-in still wins, failed mint falls through). The opt-out
test now neutralizes the AWS chain like the precedence tests do; on a machine
with ~/.aws it was failing for the same reason as the bug.

Live after the fix, same Mac, AWS credentials visible, isolated shared store:
identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in
11 s.

(cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a)

* fix(auth): review follow-ups for the free-tier rung (NS-829)

- tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches
  the free tier off; its contract is the boto chain, and the free tier now
  sits above it.
- gateway/run_notifications.py: the free-tier startup line reads auth.json
  before consulting the resolver, so a gateway boot on a machine with AWS
  credentials never mints or refreshes over the network.
- hermes_cli/anon_auth.py: module docstring says where the free tier sits in
  the ladder instead of "the ladder is untouched".
- tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of
  six (parametrized ladder cases; a failed mint that returns None or raises
  falls through to Bedrock).

scripts/run_tests.sh on the five affected files: 147 passed, 0 failed.

(cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11)

* feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone

The free tier is pre-GA. Until GA it must not exist for anyone who did not
ask for it: no identity minted, no portal traffic, no free-tier copy on any
surface. One environment variable now decides that, and one function reads it.

`guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly
"1"; only then does `nous.guest` (the user's off switch) get consulted. Every
free-tier site already funnels through `guest_enabled()`, so the gate closes
minting, routing, connector entitlement, status lines and the picker row in
one place. With the variable unset, `resolve_provider("auto")` on a fresh
install raises `no_provider_configured` exactly as upstream does.

`HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the
gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted
identities as a side effect of provider resolution, and `_has_any_provider_
configured` read them ahead of every other check, making the CLI a second
reader of a flag that must have exactly one. `_forced_new_done` and the
`force` parameter of `_reconcile_and_provision` go with them.

Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in
b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847).

Not a user preference: the variable is never written to config.yaml or .env
and never shown in setup. It is deleted at GA together with its comment in
anon_auth.py. This is a deliberate, temporary exception to the "no new
HERMES_* env vars for non-secret config" rule.

Tests: fixtures set the gate instead of deleting the old lever; one new
invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "",
"0", "true" and "new" all leave the tier off with zero portal calls, red on the
previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the
feature.

* feat(auth): the free-tier identity is created in one place, at boot; every other site is a read

Before this commit eight sites could create a Nous free-tier identity as a
side effect of something else: resolving a provider, the CLI's first-run
check, the CLI's session setup (in the background beside an own key), a
connector bearer read, the desktop polling `free_tier.status`, the sign-in
precondition, the desktop's `free_tier.provision`, and the dead-credential
re-mint. A poll could mint. Provider resolution could hit the network. Two
of them raced each other on a fresh install.

Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator.
`hermes serve` runs it on a daemon thread from `_lifespan` beside the other
background boots; `cmd_chat` runs it synchronously before the first-run
guard. It inventories credentials first (`resolve_provider("auto",
skip_free_tier=True)`: what would carry inference if the free tier did not
exist), creates the identity only when `guest_enabled()`, resolves inference,
records a `SetupRecord` in process memory and broadcasts ONE `setup.ready`
event. It runs on every boot; only the mint is gated.

`ensure_portal_identity` now requires `explicit=True` and raises otherwise.
Its callers are the bootstrap, the desktop's `free_tier.provision` (the
explicit retry when the boot could not create the identity) and the two
dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`,
`managed_tool_gateway._replace_dead_guest_token`). The background thread
path and `provision_free_tier` are deleted with their last callers.

Reads that used to mint and now only read: `auth.py::resolve_provider`
rung 7 (an existing identity still outranks the Bedrock chain, NS-829
ordering kept), `main.py::_has_any_provider_configured`,
`cli_agent_setup_mixin._ensure_runtime_credentials`,
`managed_tool_gateway.read_nous_access_token` (no identity -> None),
`anon_sign_in.run_sign_in` (no identity -> Unavailable),
`methods_free_tier` `free_tier.status`.

`setup.status` answers from the record for the launch profile, blocking up
to 8 s while the bootstrap is in flight so a client's first poll lands after
the identity exists rather than racing it; a named profile, or a process
that never ran the bootstrap, keeps today's live probe. The record's fields
ride along additively (`ready`, `free_tier`, `other_providers`,
`inference_provider`).

Identity and inference are decoupled (NS-845 Q1.3): the mint sets
`active_provider="nous"` only when the inventory found nothing else usable
(`_mint_locked(carries_inference=)`); an adopted account always does. A token
refresh no longer re-elects the provider it refreshed
(`_save_provider_state_to_source` writes credentials, not the user's
choice) — that write was how an own-key install ended up on the free tier
after the first connector call.

Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check),
bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint),
62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and
`provision_free_tier`), and a04b05260c (blocking mint in the resolver).
Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847.

Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps
inference; reads never reach the portal; a refused mint is memoised),
`free_tier.status` fails loudly if it ever calls the creator, the resolver
stub fails loudly if resolution ever mints, `setup.status` reads the record,
`skip_free_tier` proves the inventory question. The three sign-in tests for
the deleted pre-mint collapse into one (`no identity -> Unavailable, zero
portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and
off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI
matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20.

* fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup"

A free-tier identity carries $0 by design, so the portal seed reports
`paid_access=False` for it. `is_free_tier_model` did not know the welcome
host, read that as a depleted account, and every free-tier turn ended with
the credits-depleted notice telling the user to top up an account they do
not have.

Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host
(`anon_auth.route_is_welcome_host`) is the free tier. The host is the
evidence, not the model name: the paid inference host can serve
`nous/welcome` to a named account and that account's depletion is real, so
`("nous/welcome", <inference host>)` stays False. Local data only, like the
three rules above it.

Restores the two contracts dropped by hermes-magic 674e11d1eaa (the
prototype line ran without unit tests): the welcome host is free without
any pricing evidence; the model name alone is not. The first is red without
this fix.

* fix(copy): free-tier text stops promising a connector transfer and never names the config key

Sign-in copy on every surface said "Sign in to keep your connectors" and
ended with "Your connectors are kept." The transfer registry that would
make that true is empty (NS-821): nothing carries over today. The copy now
says what signing in does give ("unlock more models and tools") and the
completion line names the account, not a transfer. The docs page loses the
"connectors carry over" paragraph for the same reason.

The picker's off-state line exposed `nous.guest: false` and the word
"guest"; user copy names the free tier only (R-USR-1).

The docs page gains the pre-rollout note: until GA nothing on it happens
without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and
"replaced on next use" sentences now describe the boot bootstrap.

zh is a strict locale: the `freeTier` block was English placeholder text
copied from `en`; it is now Chinese. `connectorsKept` is renamed
`completedBody` since it no longer talks about connectors.

* feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn

The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1"
as on. Until now nothing in the desktop set it, so a packaged app could
never turn the free tier on, and a backend spawned by the app could
disagree with the app about whether the tier was live.

`electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled`
is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has
`--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch
into a module constant. `desktopBackendSpawnEnv` wraps every backend env
as the outermost call and writes the flag LAST, as "1" or an explicit "0",
so no earlier spread (`process.env`, `backend.env`) can resurrect a stray
value from the parent shell.

Stamped onto all three spawn sites: the primary `serve` spawn, the pooled
per-profile spawn, and the remote SSH `exec env ...` command (which gains
` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and
the backend probes are not backend spawns and do not get it: a
`hermes --tui` typed in the pane must not mint.

The renderer learns the same fact read-only through the existing
`hermes:launch-flags` sync IPC (`guestOnboarding`) and preload
(`window.hermesDesktop.guestOnboardingEnabled`).

Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding`
maps to it in main). Two invariant tests on the pure helpers: only "1" or
the argv flag enables; the spawn env carries "1"/"0" as the last word and
preserves every other key.

* feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll

The backend's boot bootstrap now announces `setup.ready` once, after it has
created (or refused) the free-tier identity and resolved the inference
route. The renderer used to discover both by polling `setup.status`,
`setup.runtime_check` and `free_tier.status` every 60 s from
`useStatusSnapshot`; a fresh install's chip, notice strip and onboarding
overlay could sit stale for up to a minute after boot, and three RPCs a
minute per window kept asking a question whose answer changes only at
boundaries the backend already announces.

`handleLifecycleEvent` routes `setup.ready` (active source only, like
`skin.changed`) to `notifySetupReady()`, a one-shot tick atom in
`live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens
to it and runs one readiness round at once (`setup.status` +
`setup.runtime_check` + `free_tier.status`). The readiness legs also run
once on open and on return from another app, as today. The 60 s tick keeps
only `getStatus()`.

`SetupStatusSnapshot` types the record's additive fields (`ready`,
`free_tier`, `other_providers`, `inference_provider`); readiness semantics
are unchanged and still key on `provider_configured` + `runtime_check`.

Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one
refresh from the active source and none from another; the snapshot hook's
contract is three legs on open, one leg on the tick.

* fix(cli): the banner names the free tier's model instead of "no model configured"

The welcome banner prints before credentials resolve, so on a fresh install
`model` is empty and the banner said, in red, "no model configured — run
/model or hermes setup". Under the free tier that is false: the route is
already known from local state (identity on disk, tier on), and the first
message will run on `nous/welcome`.

`_banner_left_lines` now asks the route the same question when `model` is
empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous
Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the
banner's 'no model configured' line reads the resolved route").

Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`;
gate off -> the red line, zero portal calls.

* fix(aux): vision on the free tier uses nous/welcome too

The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the
backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the
repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the
welcome host would have sent every image step past the free tier for no reason, so the
auxiliary client pins the route's one model for every lane. A backing model that takes
no images answers with the upstream's own error, which the ladder handles as it always has.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415)

* fix(gateway): hermes gateway run is a boot owner of the free tier too

Rung 5 made every demand-time free-tier site a read: resolve_provider,
the connector token, the /login precondition. That is only correct if
every process that can reach those sites ran the bootstrap first. The
CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone
messaging gateway did not. A fresh HERMES_HOME with the gate on and
`hermes gateway run` reached provider resolution with no identity to
consume, and /login returned Unavailable. Reported by @andrexibiza on
#107697 (P1).

GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an
executor thread right after startup recovery and BEFORE any adapter
connects, so a fast first DM cannot arrive with nothing to resolve. It
is its own step, not part of the turn-machinery warm-up: the warm-up is
an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0);
the bootstrap is correctness and must always run. With the gate unset it
is a local inventory and no network.

Live, real GatewayRunner.start against a fake portal in a fresh home:
  gate on   -> 1 create, identity persisted, resolve_runtime_provider=nous,
               /login precondition sees the identity
  gate off  -> 0 portal calls, no identity, no_provider_configured
Before the fix the gate-on row was identical to the gate-off row.

Test: the bootstrap seam runs before _start_prefilter_platforms and
delegates to the one creator. Red on 5554eb6993 (no seam), green here.

---------

Co-authored-by: Robin Fernandes <robin@soal.org>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 03:45:33 +05:30
Siddharth Balyan cbcf7b72f7 feat(gateway): sign in with a Nous account from a chat (/login), one shared sign-in flow (#105261)
* refactor(auth): one sign-in flow behind SignInState, rendered by the CLI and the desktop

* feat(gateway): /signin signs the free tier into a Nous account from a DM

* feat(cli): chat surfaces name /signin as the sign-in verb

* fix(auth): review follow-ups for the shared sign-in flow and /signin

* fix(i18n): carry the /status free-tier line in every locale catalog

* refactor(cli): the chat sign-in command is /login

* fix(auth): durable override cleanup in the /login sweep, and the sign-in flow in its own modules
2026-09-11 03:45:33 +05:30
Siddharth Balyan a2db110ccc feat(auth): Nous free tier: free inference and connectors out of the box, one command to sign in (#105258)
* feat(auth): Nous free tier core: anonymous identity minted on first use, welcome inference, shared-store scoping

A fresh install with no provider sets up a free Nous identity (anonymous auth method of the nous
provider) instead of forcing the setup wizard. The identity is persisted through the same path a
real login uses, so the resolver ladder is unchanged. Two seams differ: token acquisition
re-exchanges the anon credential (no refresh token), and routing pins the welcome host's single
model nous/welcome. One identity per shared store; nous.guest: false turns the free tier off.

* test(auth): free tier core contracts: lifecycle, resolver precedence, exchange seam, model pin

* docs(user-guide): free tier and signing in

New page explaining what a fresh install gets before any key or sign-in
(free inference on nous/welcome plus connectors), how the free tier
coexists with a user's own API key, how to sign in with hermes auth
upgrade and keep connectors, how to turn the free tier off with
nous.guest, what hermes logout does in each state, a troubleshooting
table, and a plain privacy note. Wired into the Using Hermes sidebar.

* fix(auth): logout leaves the free tier alone and clears the shared store for a real Nous account

Logging out of the free tier is a no-op: it is not a login, so nothing is cleared and the user
is told they were never signed in. Logging out of a real Nous account now also clears the
cross-profile store, so a profile logout is not silently re-adopted on the next boot.

* fix(model): switching off the free tier points at signing in, never hops providers

* Name the free tier in the gateway startup notice and tell explicit-provider installs about it once

* Render the Nous free tier as free tier on auth status, auth list, hermes status and portal info, short-circuit billing copy for it, and skip the keepalive when there is no refresh token

* fix(auth): free tier is set up where nothing is configured: resolver last rung and first-run check

Both the provider resolver's terminal rung and the CLI first-run check now try to set up the free
tier before declaring nothing configured. On a fresh install the first command lands in chat on
nous/welcome; a failed setup still falls through to the existing guidance.

* Add hermes auth upgrade: sign the free tier into a Nous account while keeping its connectors

The device-code flow runs as usual, with a promotion intent registered on the portal between the
code request and the token poll so the account that approves the code inherits the free tier's
connectors. The promotion status decides the outcome: only a completed one is followed by the
token grant, which is persisted over the free-tier singleton and the shared store. Declined,
superseded, retired and busy outcomes each print their own plain copy, and a retired identity is
cleared so the next use sets up a fresh one. User-facing text never names the free tier's internals.

* Show the Nous free tier as one picker row with nous/welcome and hide it when nous.guest is off

* fix(auth): upgrade opens the consent page for this sign-in; one mint attempt per process; forced free tier wins the first-run check

The browser leg of hermes auth upgrade now prints and opens the promotion claim URL with the
claim code, not the generic device page. A failed mint is attempted once per process so several
bootstrap sites cannot hit a closed gate or a 429 twice; a retired credential resets that so
re-minting still happens. HERMES_FORCE_GUEST is honoured ahead of the first-run provider check.

* fix(auth): pin the welcome model on the selected route, not on profile state; background setup retries after a failure

A credential-pool entry can select a paid Nous key while the profile singleton is still the free
tier. The model pin now keys on the resolved endpoint (welcome host) in agent init and /model, and
the pin in model normalization is removed since it had no route to look at. A failed background
identity setup releases its latch so a later attempt in the same process can try again.

* fix(auth): decide the Nous model together with the route on every credential-pool swap

The credential pool can move a Nous agent between the welcome host and the portal host after
init. One helper, pin_model_for_route, now runs at init and inside every pool swap, so the
welcome host always carries nous/welcome and a paid endpoint always keeps the caller's model.

* fix(auth): apply the route model policy on every wire mode during a pool swap; release the setup latch if the thread cannot start

* fix(auth): free-tier lifecycle takes profile then shared lock, reconciles with the shared store, persists the mint before exchanging, and clears only the identity that died

The shared store is the identity of record for a Hermes root: a profile holding a stale free-tier
identity adopts a sibling's newer sign-in instead of keeping the guest, and never overwrites the
shared account. Locks are taken in the documented order (profile, then shared). A minted credential
is stored as soon as create succeeds, so a rate-limited or timed-out exchange does not lose it and
trigger a second mint. Retiring a dead credential removes only that credential from both stores.
Guest exchange uses the resolver's canonical portal URL.

* fix(auth): a credential rotation never rewrites the conversation model; connectors honour the off switch and replace a retired free-tier credential

The welcome host serves one model, so a rotation onto it is refused for any conversation on another
model instead of silently switching that conversation to nous/welcome (the model pin applies only
when a route is first chosen). The connector token path now treats the free tier as absent when
nous.guest is false, including cached tokens, and shares the one dead-credential rule with
inference: a retired identity is replaced once rather than returning its stale token.

* fix(auth): plain login never imports the free tier as OAuth credentials; the gateway startup line reads persisted state only

A free-tier identity in the shared store is not an OAuth credential to offer for import; a real
sign-in replaces it. The gateway's startup notice now answers provider precedence from persisted
state (no token refresh at boot), so an expired free-tier token cannot stall the online message.
2026-09-11 03:45:31 +05:30
SHL0MS 2ddeba9e17 Merge pull request #99523 from somewheresy/justin/e-1047-route-hermes-actual-provider-through-chat-completions-with
fix(providers): use chat completions for all Actual routes
2026-09-10 17:25:53 -04:00
Teknium 564aef2946 fix: /model onto Bedrock Mantle keeps SigV4 auth instead of 401ing
Startup (agent_init._init_openai_client) ran configure_bedrock_openai_client_kwargs,
so the aws-sdk sentinel became a SigV4-signing httpx client. Every later client
rebuild — switch_model, fallback restore, credential rotation, request-scoped
clients — went through create_openai_client with bare {api_key, base_url} kwargs,
so the OpenAI SDK sent "Authorization: Bearer aws-sdk" and Mantle answered 401
"Invalid bearer token". Symptom: `hermes --provider bedrock --model
openai.gpt-5.6-terra` works, `/model openai.gpt-5.6-terra` inside the CLI/TUI fails.

Install SigV4 in create_openai_client itself (the single chokepoint every primary
OpenAI-wire client passes through) whenever the base_url is a Mantle host, so all
rebuild paths inherit the fix rather than each remembering to call the adapter.
A real AWS_BEARER_TOKEN_BEDROCK key is left alone (the adapter only rewrites the
aws-sdk / no-key-required placeholders).

Live repro (local HTTP sink capturing the Authorization header after switch_model):
before "Bearer aws-sdk", after "AWS4-HMAC-SHA256 Credential=...".
2026-09-10 12:18:44 -07:00
Justin Bennington 8a6b5b67a7 fix(providers): block Actual at Responses send sites (E-1047) 2026-09-10 15:13:02 -04:00
Justin Bennington 4135933ec9 fix(providers): preserve Actual routing during startup and key reload (E-1047) 2026-09-10 15:01:17 -04:00
Justin Bennington 8b2b83a906 fix(providers): pin all Actual routes to chat completions (E-1047) 2026-09-10 14:56:55 -04:00
Justin Bennington d7b0a72c2a fix(providers): route Actual through chat completions (E-1047) 2026-09-10 14:56:55 -04:00
Teknium f98cb00a8b fix(vault): 2FA review follow-ups — per-digit fill only for an unmistakable maxlength=1 widget; honour otpauth digits/period/algorithm
Reviewer findings on #107585:
- build_otp_fills split any >=4 code-like controls into digits. A page with
  promo/zip/referral 'code' inputs next to the real OTP box would have had a
  digit sprayed across unrelated fields. Split now requires exactly len(code)
  controls that are all maxlength=1, same form, adjacent in DOM order
  (inspection JS exports maxLength); anything else fills ONE field, the
  best-scoring one. Verified on the real Browser Use stack: 6-box widget gets
  one digit each; scattered page fills only the one-time-code input.
- normalize_otp_secret dropped digits/period/algorithm from otpauth:// URIs,
  so an 8-digit or SHA-256 authenticator would get wrong codes. Non-default
  parameters are now stored as seed|digits|period|algo and honoured (RFC 6238
  SHA-256 8-digit vector added); hotp:// is rejected explicitly.
- rebase on main (prompts.ts conflict) + prettier.
2026-09-10 11:48:01 -07:00
Teknium d9ca9c974d feat(vault): two-factor codes — automatic from a saved authenticator key, otherwise asked for in the user's UI
Follow-up to #106480. Sites that ask for a code after the password stopped
the agent cold: the login classifier excludes one-time-code fields on
purpose (a password must never land in an OTP box) and there was no tool
for the second step, so the only move was to ask in chat.

browser_vault_enter_code
  Fills the one-time code the current page asks for. Two sources, same
  invariant as passwords (the code goes to the page over the supervisor
  socket and never enters model context):
  - a TOTP seed on the login: local vault `otp_secret` (RFC 6238, stdlib,
    verified against the RFC test vectors), 1Password `op item get --otp`,
    Bitwarden `bw get totp`. Nobody is asked.
  - no seed: the surface prompts "Verification code for {site}"; the user
    types what their phone/email/app shows. Enter on empty / Skip declines
    and the tool returns code_declined ("do not ask again this turn").
  no_code_field tells the model the site wants a passkey / hardware key /
  app approval: hand it to the user's device and wait for navigation.
  Per-digit OTP boxes (maxlength=1 pattern) get one digit each in DOM order.

Surfaces
  CLI: sudo-style panel, code shown as typed (not a secret worth masking,
  typos must be visible), Enter submits, ESC/empty skips.
  Desktop: "Verification code for {site}" card via vault.code.request /
  vault.code.respond (gateway), owner-routed like the other vault prompts.
  Settings → Passwords & Logins: optional "Authenticator key" field on the
  add form (base32 or otpauth:// link); items with one show a "2FA auto"
  badge. `hermes vault add` asks for the same optional key.
  browser_vault_fill's result now says what to do next ("if the site asks
  for a verification code, call browser_vault_enter_code with this handle").
  Six locales.

Verified live (real model, local 2FA site that checks the TOTP; CLI PTY):
  A. login saved with authenticator key → signed in through 2FA, zero
     prompts, code/password absent from the transcript
  B. login without key → code panel → user types code → signed in
  C. panel dismissed → agent stops and explains, never asks in chat
Unit: RFC 6238 vectors, seed normalisation, mint-without-asking, per-digit
spread, decline, no-code-field; Desktop card test (owner routing, trim, Skip).
2026-09-10 11:48:01 -07:00
Teknium a42388b3a9 feat(agent): a2a_key names a bot author's turns
Dropped from the #103888 salvage for lack of a consumer; the honcho a2a
session key is that consumer.
2026-09-10 10:45:57 -07:00
Teknium 98d11c95f4 feat(vault): zero-setup UX — save a login on the page that needs it, managers auto-detected, one "Passwords & Logins" surface
Nobody should have to learn `hermes vault add` or find a toggle before "log into GitHub" works.

- browser_vault_save_login: when the agent reaches a sign-in page with no saved login it asks the user
  on THEIR surface (CLI two-step panel on the sudo modal: identifier shown, password masked; Desktop
  card with labelled Email/username + Password fields). The answer goes to the encrypted vault bound to
  the page origin and is filled at once; the model gets back only the handle and identifier. Declining
  returns save_declined; headless sessions get prompt_unavailable. Never a password in chat.
- Vault tools ride with the browser toolset (check_browser_requirements) instead of appearing only once
  the vault has items — an empty vault is exactly when save_login is needed. browser_vault_list hints
  at it when empty.
- 1Password / Bitwarden are login sources as soon as their CLI is installed; `vault.<name>.enabled`
  is opt-OUT only. Settings shows Detected/Locked/Unlocked/Off/Not detected with a switch only for
  installed managers; `hermes vault sources` reports detection, `--disable`/`--enable` flip the opt-out.
- Desktop nav/page renamed "Passwords & Logins"; empty state tells the user they do not need to add
  anything; all five locales updated. Docs rewritten from "how it works" to "say log into X".
- New per-thread SaveLoginPrompt callback (agent/vault_backends/unlock.py) installed beside the unlock
  prompt on every CLI site and the gateway bridge (vault.save_login.request/respond/expire), propagated
  to worker threads via tools.thread_context.

Live: CLI PTY (real model, packaged Chromium, local login server) — panel shown, identifier + masked
password typed, server received the correct password, password absent from terminal transcript and
from every file under HERMES_HOME outside vault/. Native Electron (headless, isolated HOME/HERMES_HOME,
own Vite + CDP port) — card shown, "Save & sign in", server received the password, Settings lists the
saved item, password absent from the rendered UI.
2026-09-10 10:35:07 -07:00
Teknium c8998c3957 feat(vault): fill payment cards and addresses at checkout, cards behind a confirm prompt
payment and address items could be stored (CLI wizard, Desktop dialog) but nothing could
fill them: a dead surface holding real card numbers. browser_vault_fill now handles all
three kinds through the same origin-bound, supervisor-only, redacted path:

- classify_checkout_control / select_checkout_fills map WHATWG autocomplete tokens
  (cc-number, cc-exp[-month|-year], cc-csc, address-line1/2, address-level1/2, postal-code,
  country-name) with label/name heuristics as backup; a combined "MM/YY" control gets
  exp_month+exp_year and suppresses the split fills; inspection now covers <select>
  (country, state, expiry month) and the fill script picks an option by value or text.
- Every payment fill goes through request_elicitation_consent (gateway button round-trip
  or CLI panel) before a byte is written; declined → payment_declined, headless sessions
  are refused. A prompt injection that reaches a checkout can ask, not spend. Card values
  join the redaction registry like passwords; the result lists targeted field tokens only.
- Origin is now required for every kind (CLI wizard asks; Desktop dialog always shows the
  field) because a card without a bound origin is unfillable.
- The tool descriptions, docs and CLI copy drop "Phase 1 / login only".

Live (evals/vault_fill_live_e2e.py, real browser_exec + packaged Chromium): decline writes
nothing; accept fills card/expiry/CVC on the /checkout tab, leaves the email box and the
country <select> untouched, and neither the card number nor the CVC appears in any result.

browser_vault_tool also: focuses the tab on the bound origin holding the right form before the
origin pre-check (focus_page from the previous commit); tool descriptions say "the browser's
input tool" (rewritten per session by model_tools); _check_vault_available is registered
uncached because its answer is per profile (vault dir + config) and the probe is a file stat.
2026-09-10 10:35:07 -07:00
Teknium 27ca97e469 fix(vault): cross-process store lock + fsync; profile-scoped, bounded redaction registry
Two review findings on #96988/#106480 that were still open:

- VaultStore serialized read-modify-write with a threading.Lock only. The Desktop gateway,
  a CLI `hermes vault add` and a TUI slash worker are separate processes writing the same
  vault.json.enc, so two adds could drop each other's items, and the temp file was renamed
  into place without fsync (a crash between rename and the next sync loses the vault).
  Writes now take an flock/msvcrt lock on <vault>/.vault.lock and fsync file + directory.
- The vault redaction registry was one process-global, unbounded set. Under gateway
  multiplexing profile A's passwords scrubbed profile B's browser output (and confirmed to
  B that those bytes exist). It is now keyed by profile home, capped at the 64 most recent
  values per profile, and clearable (clear_vault_redaction_values).

Also canonicalizes payment/address payloads (PAYMENT_FIELDS / ADDRESS_FIELDS, required
fields enforced, stray keys dropped) so the checkout fill in the next commit never has to
guess a user's ad-hoc field names; the Desktop dialog already wrote these names.
2026-09-10 10:35:07 -07:00
Teknium dedc99ec6c fix(vault): ownership and race findings from the second independent review
Manager tokens: lock generation fence (a Lock acknowledged while `bw unlock`
/ `op signin` is still running discards the late token); tokens record the
unlocking gateway session and are released when THAT session ends, not when
any sibling session in the profile is torn down.

1Password: OP_CONNECT_HOST/TOKEN come from the profile's scoped secret store
like the service token (Connect outranks a service token inside op), never
from the launch environment.

Vault RPCs bind params.profile (home + secret scope) so a shared remote
backend serving several profiles locks/lists/unlocks the requested one;
unknown profile → RPC error, not a crash.

Fill target: inspection stamps are `<nonce>:<index>`; a fill resolves only
its own inspection's stamps, so an interleaved second inspection can no
longer redirect A's password into a newly mounted field (real Chrome: 0
filled, both fields empty).

Desktop Settings: every RPC goes through the owner profile's socket
(requestGatewayForProfile), query keys carry (connection, profile), an owner
change closes dialogs and wipes drafts (a master password typed for A is
never submitted to B; a late list from A never paints under B), and vault.add
secrets travel in a ref consumed by the mutationFn instead of mutation
variables. Three owner-routing invariant tests on the real component.

Docs/PR body: session-scoped release, lock-race semantics, bw --passwordenv.
2026-09-10 10:35:07 -07:00
Teknium 5aa9e7f033 fix(vault): independent-review findings — vendor contracts, profile scope, transport, target binding
Bitwarden unlock now uses the CLI's documented non-interactive channel:
`bw unlock --raw --nointeraction --passwordenv VAR`, VAR set on the child
environment only (bw 2026.x rejects a piped password with "Master password
is required"). Verified against the real published binary.

Manager session tokens are keyed by (profile home, backend): a Desktop
gateway hosting several profiles can no longer reuse or lock another
profile's session. Status probes (`vault.sources`, is_unlocked) no longer
refresh the idle TTL; only real manager calls do. Gateway session teardown
locks the profile's managers (a per-session unlock ends with the session).

1Password service-account token comes from the profile-scoped secret store
(get_secret), not ambient os.environ.

`vault.source.set` no longer references a module constant (bind_module
rebinding dropped it → NameError on every Settings toggle).

Fill target binding: inspection stamps each input with a per-inspection
slot attribute; the fill resolves by stamp and requires type=password, then
strips every stamp. A DOM reflow between inspect and fill can no longer
redirect the password into a text field (reproduced in real Chrome before,
0 filled after).

Redaction boundary: no 4-char floor, CR/LF-normalized form registered
(what a text input actually stores), JSON object KEYS scrubbed in both
browser redactors; longest value first. Docs now state the real trust
model: accidental-disclosure protection, not an execution sandbox.

Desktop: the mid-turn card sends the master password through the owning
session's socket (requestForOwnedSession), never the ambient foreground
gateway; `vault.unlock.expire` clears a stale card; Settings keeps the
master password out of react-query mutation variables (ref consumed by the
mutationFn). One renderer invariant test for the routing.
2026-09-10 10:35:07 -07:00
Teknium e9f7d44f3f fix(vault): carry the unlock prompt onto tool worker threads; honour binary_path in install checks
Live Desktop repro: the fixture model called browser_vault_unlock and got
unlock_unavailable although the renderer was interactive. tool_executor runs
handlers on a propagated worker thread; thread_context only copied the
approval and sudo thread-local callbacks, so the unlock prompt registered by
_wire_callbacks was invisible there and can_prompt_here() said nobody could
answer. The callback table in thread_context now lists every per-thread
prompt (approval, sudo, vault unlock) so a new one cannot silently drop off
worker threads again. After the fix the same turn shows the masked card and
completes.

vault.sources / `hermes vault sources` report a manager as installed when
its configured binary_path exists, not only when it is on PATH.
2026-09-10 10:35:07 -07:00
Teknium 8e7e9be217 feat(vault): sign in with 1Password or Bitwarden logins, unlocked per session
The browser vault now draws from three login sources behind one handle
shape: the local encrypted vault (vault_…), 1Password Login items (op:…)
and Bitwarden Password Manager logins (bw:…). browser_vault_list aggregates
metadata across them; browser_vault_fill routes by prefix and resolves the
password at fill time only, through the manager CLI.

External managers are locked until the user unlocks them for the current
session. The new browser_vault_unlock tool (and the fill path, implicitly)
asks the surface to show a masked master-password prompt — CLI panel
(reuses the sudo panel state), TUI/Desktop via a vault.unlock.request
blocking card. The password goes to `op signin --raw` / `bw unlock --raw`
on stdin, never argv or env; only the session token is kept, in memory,
with a 30-minute idle TTL, cleared on session close or `vault.lock`.

Headless contexts (cron, webhook, api_server, -q) can never prompt: the
manager is reported as locked with unlock=unavailable_in_this_session and
fill refuses — the same posture approvals take where nobody can answer.

Config: vault.onepassword / vault.bitwarden {enabled, binary_path, …};
a 1Password service-account token skips the prompt for headless use.
RPC: vault.sources, vault.source.set, vault.unlock, vault.lock for Settings.

Tests (2, real subprocess against a fake bw; each proven red by sabotage):
headless never prompts or spawns; unlock feeds stdin only, token never
enters os.environ, fill routes by prefix and the password only reaches the
fill script.
2026-09-10 10:35:07 -07:00
Teknium cde0ad0edd feat: agent signs into sites from an encrypted local vault (CLI, browser fill, Desktop Settings)
Consolidated re-apply of #96988 onto current main. Ported from
Merit-Systems/OpenInstinct (MIT) opaque-handle autofill design: the model
sees vault handles + login metadata, the password is resolved and filled
server-side over the supervised CDP socket, and filled values are scrubbed
from every browser tool result by an unconditional redaction registry.

Rebase adaptations to the Sep-2026 facade/sibling layout:
- toolsets: one _HERMES_CORE_TOOLS entry (the browser toolset derives from it)
- hermes_cli/main.py: vault parser registered via the subcommand owner table
- file_safety: vault/ joins the _READ_DENIED_DIRS credential-dir table
- redact: registry scrub runs before the redact_secrets early-return
- browser_vault_tool: _run_browser_command now lives in browser_tool_session
2026-09-10 10:35:07 -07:00
Teknium 94f77dfa0d fix(agent): thread turn_author through conversation_loop.run_conversation
The facade forwarded turn_author= to the loop's public entry point, which did not
declare it: every real AIAgent.run_conversation() turn raised TypeError (16 CI
failures across provider, sidecar, cron and finite-chat suites). The PR's tests
only exercised build_turn_context directly, so the missing hop was invisible.
Adds one facade-through-loop test that goes red when the kwarg is dropped.
2026-09-10 10:27:07 -07:00
Teknium ff90eb28ff test(bot-mode): trim the author and loop-guard suites to their invariants
Keep one or two behaviour tests per seam (author reset on a cached agent, forged
_turn_author refused, guard trips and cools, single charge on the busy path) and
drop the parser/setting enumerations. a2a_key goes with them: nothing in this PR
reads it; the honcho follow-up that does can bring it back with its consumer.
2026-09-10 10:27:07 -07:00
Erosika 091ac34ba9 fix(bot-mode): qualify the peer-dm author id with the sender's hostname
message_agent sent a bare bot:<profile> id through hermes peer dm, so a remote coder and the recipient's own coder shared one author id. The peer branch now sends bot:<hostname>/<profile>, with the hostname cleaned like any author field and slashes dropped. The direct local path keeps the bare id.
2026-09-10 10:27:07 -07:00
Erosika d3c8bacbe3 fix(bot-mode): qualify relayed authors with the sender's connection id
A relayed DM stamped bot:<profile> on the recipient turn, so an ops profile on another machine and the local ops profile shared one author id. The Desktop now forwards from_connection with each bot_relay.deliver, and delivery_turn_author builds bot:<connection>/<profile> for it while the Desktop's own gateway ("local") keeps the bare id. An api author object accepts an optional origin string that yields the same shape.
2026-09-10 10:27:07 -07:00
Erosika b41b036e64 fix(memory): pass on_turn_start kwargs only to providers that accept them
MemoryManager.on_turn_start forwarded the author kwargs to every provider and _each_provider swallowed the TypeError, so a provider with the two-positional on_turn_start(n, text) stopped running. The kwargs are now filtered against the provider's signature the way _provider_sync_accepts filters sync_turn.
2026-09-10 10:27:07 -07:00
Erosika 55b3ea0b11 fix(gateway): count each bot message once in the loop guard and consume the author variable
The Telegram adapter asks the authorization check before dispatch, the ingress gate asks it
again, and the busy path asks a third time. Each call counted one loop-guard event, so a
Telegram bot tripped the budget after a third of the configured messages. The verdict now only
refuses a chat that is cooling down. The ingress gate counts an admitted bot message once.

`parse_turn_author` treats only booleans, integers and the strings true/1/yes as a bot flag,
and returns None for an author with neither id nor name. Names keep format characters and
non-breaking spaces so emoji sequences survive. The quiet one-shot pops HERMES_TURN_AUTHOR
before the turn so tool subprocesses do not inherit it. `max_events` must be a whole positive
number. Issue numbers move out of code comments.
2026-09-10 10:27:07 -07:00
Erosika 8969511209 feat(memory): carry the turn's author to sync_turn
on_turn_start already received the author trio. sync_turn did not, so a
provider that wanted to write the turn under its author had to stash state
between the two hooks. sync_turn now takes turn_author as a keyword-only
argument, and MemoryManager sends it only to providers whose signature
accepts it, so existing providers keep working unchanged.

build_turn_context resets the author on the agent at the start of every turn
so a cached gateway agent never carries a bot author into the next human
turn. agent/turn_author.py holds the parsing and the HERMES_TURN_AUTHOR
carrier.

MemoryProvider.identity_signature() is a new optional hook: the identity
values a provider writes under, declared by the provider itself, for the
gateway's agent cache to key on.
2026-09-10 10:27:07 -07:00
Erosika 70b1ff6930 feat(memory): carry the turn's author into the memory-provider contract
`on_turn_start` documents a per-turn kwargs channel — "kwargs may include:
remaining_tokens, model, platform, tool_count" — and `MemoryManager`
forwards whatever it receives. Its only caller passed nothing, so a memory
provider had no way to learn who wrote the turn it was being told about.

Providers that key durable state on identity resolve one identity when the
session is created. A shared session does not work that way: threads are
shared by default (`thread_sessions_per_user` is False), so alice, bob, and
another agent all write turns into a session whose peer is whoever spoke
first. The gateway's answer today is the `[name]` prefix it prepends to the
message text, which the model reads and a provider cannot.

`turn_author` now travels from the gateway through `run_conversation` into
`build_turn_context`, which forwards `author_id`, `author_name`, and
`author_is_bot` to every provider. It stops there — the trio never reaches
the model, and providers that ignore the kwargs are unaffected.

The bot flag is sent on every transport, not only shared sessions: a
provider deciding whether a turn may write to durable memory needs it in a
DM too.

`SessionSource.is_bot` is only as good as its producers. `build_source`
defaults it to False and 3 of 32 adapter call sites pass it, so most
platforms still report every author as human. Populating the rest is
follow-up work; nothing here depends on the flag being right yet.
2026-09-10 10:27:07 -07:00
Teknium 073c57872a feat(models): DeepSeek V4.1 Flash on the Nous Portal and OpenRouter pickers
Add deepseek/deepseek-v4.1-flash to OPENROUTER_MODELS (Nous list derives from it),
regenerate the docs manifest, and give the slug its own 1M context entry and 600s
reasoning-stale floor — the longest-key-first scan otherwise lands the new slug on
the 128K `deepseek` catch-all and no floor. Live probed on both routes: echoed
model matches, usage.cost billed.
2026-09-10 09:22:31 -07:00
Teknium ee35a4624f fix(routing): address review — session runtime decides, quarantine continues the chain, unclassified ids break ownership
Three defects found in review of the first head (@ehz0ah):

* The mid-request hop read the PERSISTED provider from disk. A live
  `/model xai-oauth` session over `model.provider: auto` therefore still
  reached the discovery chain and billed Nous. _try_payment_fallback now
  takes the route's main_runtime snapshot; disk is the fallback only when
  no session runtime exists.

* After a configured fallback was quarantined mid-request (401, refresh
  failed), the second pass went straight to the discovery chain, which
  the new gate refuses — so later CONFIGURED entries never ran and the
  original error was re-raised. The second pass now re-walks the task
  chain and main chain (the quarantined entry is unhealthy and skipped)
  before discovery.

* current_provider_owns_vendor dropped ids detect_vendor could not
  classify, so Bedrock's 15-id catalog (14 unclassified `us.anthropic…`)
  looked exclusively DeepSeek and `/model deepseek-v4-pro` stuck on
  Bedrock. An unclassified id now counts as evidence of a multi-vendor
  catalog: ownership requires every id to classify to the one vendor.
2026-09-10 09:11:00 -07:00
Teknium 7e0b5cd235 fix(auxiliary): auto never bills a provider the user did not select
With a main provider selected, an unusable main route (expired xAI/Codex
OAuth token, 401/402/429 mid-session) fell through the built-in discovery
chain (OpenRouter -> Nous -> custom -> api-key) and quietly ran every
compression, title and memory-flush call on whichever OTHER account was
still logged in. Reported as "using Grok on my Premium+ sub, my Nous
Portal balance kept draining" — the chat visibly stayed on Grok while the
side tasks were billed elsewhere, and re-logging into X did not help
because the aux side never consulted the selected provider.

The discovery chain is now reserved for installs with no selected main
provider (`model.provider: auto` / unset). Otherwise the ladder is
main -> auxiliary.<task>.fallback_chain -> fallback_providers -> refuse
with a warning naming the dead provider and the fix. Both entry points
gate on the same predicate: the resolve-time route and the mid-request
payment/auth hop (_try_payment_fallback).

Existing chain tests that asserted the hop now pin `provider=auto`, the
one case where discovery is still the contract.
2026-09-10 09:11:00 -07:00
fangliquanflq fa75692211 docs(agent): clarify detached fork lifecycle boundary 2026-09-10 13:23:00 +02:00
fangliquanflq 77f2ce853f fix(agent): suppress detached fork session-end hooks 2026-09-10 13:23:00 +02:00
fangliquanflq feb03196a2 fix(agent): isolate detached forks from lifecycle hooks 2026-09-10 13:23:00 +02:00
Teknium a0749d583a refactor(sessions): trim #106543 to the host-side heal and two invariants
The publish-side taxonomy (_end_stamp_class, _compression_parent_obstacle,
compression_parent_deliberately_ended) and the agent-guard delegation produced
exactly main's verdict -- every non-automatic stamp still fails closed -- so
they only changed an error string. Inline the "explicit close with no
continuation" test into reopen_if_explicitly_closed(), restore main's publish
branch and agent guard untouched, and keep two tests: the field shape rotates
after the host clears the stamp; boundary/compression/automatic stamps and a
session already claimed for teardown are never cleared.
2026-09-10 03:04:51 -07:00
Totoro-qaq 9cd1453f26 fix(sessions): a stale explicit-close stamp no longer makes a session permanently uncompressible
Consolidated from PR #106543 (5 commits, final tree d2c4d908) by @Totoro-qaq.
publish_compression_child() fails closed on any non-automatic end stamp and
end_session() is first-stamp-wins, so a stale tui_close on a session the TUI
still routes turned every rotation into "compute the summary, then discard it"
(#106459). The host that still routes the session clears the stale explicit
close via SessionDB.reopen_if_explicitly_closed() before the turn starts;
publication never heals explicit closes. Review probes by @ehz0ah.
2026-09-10 03:04:51 -07:00
Teknium aeecb110f8 fix(deepseek): deepseek-flash is the canonical Flash id; retired names fold onto it
DeepSeek retired deepseek-v4-flash on 2026-09-10 (V4.1-Flash release); the API's
model name is now `deepseek-flash` and /v1/models lists only it. Hermes still
folded every non-V-series name onto deepseek-v4-flash, so `/model deepseek-flash`
on the DeepSeek provider was rewritten, then the validator "auto-corrected" it
back against the live listing: "Auto-corrected deepseek-v4-flash -> deepseek-flash"
on every switch.

Retired aliases (deepseek-chat / -reasoner and other fuzzy names) now fold onto
deepseek-flash; the curated catalog, profile fallback list, aux default, goal-judge
hint and pricing snapshot (2026-09-10 off-peak USD) follow the docs. Dated
deepseek-v4-* ids still pass through untouched.

Builds on YipTszkwan's #107126 (earliest fix in the cluster).
2026-09-10 02:44:26 -07:00
YipTszkwan 8435a3ae00 fix(deepseek): recognise the version-less deepseek-flash model id
DeepSeek's 2026-09 Flash refresh introduced a version-less canonical id:
GET /v1/models now returns `deepseek-flash` (alongside `deepseek-v4-pro`), the
API accepts it directly, and the older `deepseek-v4-flash` is server-side
aliased onto it. Every DeepSeek model-id gate in Hermes keys off the
`deepseek-v<N>` prefix, so the new id silently missed all four:

* DeepSeekProfile.build_api_kwargs_extras classified it as non-thinking and
  omitted `extra_body.thinking`. The server then defaults to thinking-on, so
  the user's thinking toggle and `reasoning_effort` were quietly ignored.
* `_normalize_for_deepseek` folded it onto `deepseek-v4-flash` (it misses the
  V-series regex), so the id a user picked never reached the wire and the
  config stored a different model than the picker advertised.
* `DEFAULT_CONTEXT_LENGTHS` fell through to the 128K `deepseek` catch-all
  instead of the real 1M window, capping the model at an eighth of its
  context before compaction kicked in.
* `_REASONING_STALE_TIMEOUT_FLOORS` had no entry, leaving the stale-stream
  detector at its 180s default instead of the 600s reasoning-model floor.

Verified live against api.deepseek.com: `deepseek-flash` answers 200 with
`model: deepseek-flash`, accepts image input (the refresh folds vision into
the Flash model), and the in-between id `deepseek-v4.1-flash` is rejected
with "The supported API model names are deepseek-flash, deepseek-v4-pro".

Adds the id to all four gates plus regression coverage for each site.
2026-09-10 02:44:26 -07:00
Teknium 4baf4269ff fix(metadata): GLM context windows from the live catalog; hyphenated slugs, entry-level overrides, aux inherits main pin
Every GLM slug that missed a DEFAULT_CONTEXT_LENGTHS key fell to the "glm" 202,752 catch-all, so
compression fired at ~20% of the real window and the aux feasibility check auto-lowered the session
threshold to that number (the 202,752 in the Coatue compression report).

- Catalog: GLM entries from Nous + OpenRouter /v1/models (2026-09-09): 5.3 / 5.3-flash 1,310,720
  (:batch/:US 1,048,576), 5.2 1,048,576, 5 / 5.1 / 4.7 / 4.6 204,800, *-turbo / 4.7-flash 202,752.
- _catalog_key_matches: version separators normalised on both sides so relay slugs like z-ai-glm-5-3
  hit glm-5.3 (#97398; approach from PR #97412 by @shellybotmoyer, re-implemented on current main).
- get_custom_provider_context_length: entry-level custom_providers[].context_length backs every model the
  entry serves when no per-model override exists, so /model switch stops dropping it (#98387; from
  PR #98396 by @liuhao1024, re-implemented on the config_providers sibling).
- check_compression_model_feasibility: when aux compression is the main model on the main route, reuse the
  main model's resolved window instead of re-resolving without its pin (#89500 mechanism 2, #45519).

Closes #97398, #98387, #89500, #87825, #97595, #97820 (suffix strip already on main; catalog values now match).
2026-09-10 01:08:18 -07:00
Teknium 97ca90f184 fix(desktop): expired OAuth grant shows a one-click 'Sign in again' instead of a retryable Provider error
A rejected OAuth token (HTTP 401 'User not found' from Nous Portal, Codex,
xAI…) reached the desktop error card as 'Provider error' with Retry as the
first action, which just replays the same dead credential.

Backend: nonretryable_client_error_result dropped failure_reason /
failure_retryable, so error_surface classified every non-retryable 4xx as a
retryable provider failure. It now stamps the classifier verdict like the
max-retries path, and auth-layer descriptors carry auth_kind
(oauth|api_key, derived from the provider catalog tab) + provider_label.

Desktop: an auth/oauth surface renders 'Authentication error', explains that
the <provider> sign-in expired/was revoked, and offers 'Sign in to <provider>
again' which launches that provider's existing onboarding OAuth flow scoped
to the failed session's gateway profile. Retry stays as the follow-up click.
Re-login to the provider already in use keeps the current model instead of
swapping in the recommended default.
2026-09-09 16:30:24 -07:00
Teknium edac5f6378 fix(auxiliary): origin-scoped key shield, task overrides over named providers, negotiation-only stream fallback
Follow-up to #106866 from the post-merge audits (@ehz0ah, @victor-kyriazakos):

- _recoverable_pool_provider: the "endpoint mismatch, not a dead key" shield now (a) compares the full
  origin via base_url_origin() (scheme+host+port — a different port or an HTTPS->HTTP downgrade is a
  different trust boundary) and (b) applies only when the rejected client carried the SESSION's key, so
  an independently owned auxiliary pool keeps rotating at its own configured origin.
- _resolve_named_custom_branch: explicit/per-task base_url and api_key compose OVER the providers.<name>
  entry field-by-field instead of the entry silently replacing them (URL-only, key-only, both, neither).
- _acreate_with_progress: plain-create fallback only for a rejected stream NEGOTIATION (mirrors the sync
  wrapper); a mid-stream failure propagates to the classified recovery ladder instead of re-sending the
  whole prompt non-streaming.

Tests (all red on main): foreign origin x {host, port, scheme} shields the session key; same origin and
independent aux pool still rotate; named-provider override matrix through resolve_provider_client on a
real temp HERMES_HOME; async content-then-error is not retried non-streaming.
2026-09-09 16:21:36 -07:00
Teknium 2db0c7a2d8 fix(auxiliary): compression sticks to the session's OpenAI endpoint; foreign-host rejections don't kill the key
`auxiliary.<task>.provider: openai` was rewritten to custom + api.openai.com/v1 unconditionally, and the
api-key discovery rung used the registry default endpoint, so proxy/gateway users (OPENAI_BASE_URL or
providers.openai) had compression hop to the public endpoint with a proxy-issued key -> 401 -> the
key/pool was quarantined and the session sat over the compression threshold.

- _expand_direct_api_alias: a providers.openai entry keeps its name (named-custom branch applies its
  base_url/key); otherwise OPENAI_BASE_URL wins over the public default.
- _resolve_api_key_provider: when the bound session runtime is that provider, use its endpoint + key.
- _recoverable_pool_provider: a rejection at a host other than the session's configured endpoint for
  the same provider is an endpoint mismatch, not a dead key -> no rotation / unhealthy mark on the pool.
2026-09-09 14:06:12 -07:00
Teknium e6b890e584 fix(auxiliary): async progress-hook wrapper + route the async primary through it (#98466)
Complete the seam default from the salvaged commit: _relay_async_completion needs an async twin of
_create_with_progress, and the async primary attempt previously streamed only for stream-only
providers, so a hooked async compression call ticked the watchdog zero times.

Adds one invariant test: with a hook installed, BOTH relay defaults stream and tick per chunk;
without a hook they are byte-identical plain creates. Red on origin/main.
2026-09-09 14:06:12 -07:00
Konstantin Khlopkov fba6110cde fix(agent): default relay completions to the progress-hook wrapper 2026-09-09 14:06:12 -07:00
Siddharth Balyan b4d04eb8fd Connector tools (Gmail, Linear, Notion, ...) are searchable and callable through tool_search for signed-in Nous users (#106842)
* feat: add session-scoped connector access for onboarding

* fix(connectors): availability is the config flag AND the portal entitlement — no free-tier leg

The port carried a third availability leg from hermes-magic: a stored guest
(free-tier) identity short-circuits the managed-tool entitlement check. That
leg reads hermes_cli.anon_auth, which does not exist on hermes-agent main, so
connectors_available() raised ImportError inside its fail-closed try and the
whole connector surface was silently dark on a plain upstream checkout.

On this tree availability is the two-leg AND the design started with:
tools.connectors.enabled AND managed_nous_tools_enabled(). The free-tier leg
is a hermes-magic concern and belongs in hermes-magic's own delta over this
branch, next to the identity it depends on. Its integration test goes with it.

* docs(tool-search): connectors section — remote tools through the bridge

The squashed port carried the code but not the user-facing docs. Restores the
Connectors section of the Tool Search page and the connector-gateway host /
CONNECTOR_GATEWAY_URL override on the Tool Gateway page, updated for the
manage_connections tool and the pure-connector batch rule.

* fix(tool-search): connector tools rank with local tools in one pass instead of taking leftover slots

dispatch_tool_search ran BM25 over the local catalog, filled `limit` slots,
then appended connector hits only into slots left empty. On a 300-tool
catalog no slot was ever empty, so with Gmail and Google Calendar connected
"send gmail email" returned five betterstack tools and zero connector tools.

The gateway's hits for a query now become catalog entries (connector name,
slug words, description as the search text) and join the local catalog for
that query's BM25 pass. One ranking, one rarest-token admission rule for both
sources, `limit` as the total per query. The merge loop and the separate
record builder for connector hits are gone; `_shared_tool_record` serves both
sources.

The gateway search timeout rises from 8 s to 30 s. One request with six
use_cases measured 7 s, so 8 s sat on the edge and cut real answers off; the
failure path is unchanged (local-only results, no error to the model).

Live, 311 local tools + gateway, before -> after:
  "send gmail email":           5 betterstack tools -> gmail SEND_EMAIL, CREATE_EMAIL_DRAFT
  "read google calendar events": 5 betterstack tools -> googlecalendar EVENTS_LIST_ALL_CALENDARS
  "linear create issue", "betterstack incident": unchanged
Benchmark (25 labelled queries): connector recall 0.09 -> 0.82, precision@5
0.18 -> 0.59, false positives on absent intents 17 -> 2.

* refactor(tool-search): connector leg into tools/connector_search.py

tools/tool_search.py is a facade. The connector leg (gateway hits as catalog
entries for tool_search, remote schemas for tool_describe, the
connections_in_scope gate) was appended to it by the port. It now lives in
its own sibling, tools/connector_search.py, and the facade imports the three
entry points: connections_in_scope, connector_entries_by_group,
remote_schemas_for.

No behaviour change. The tool_describe remote block became
remote_schemas_for(names, current_tool_defs, connector_describe) with the
same inputs, the same silent-degradation contract and the same injection
seam the tests already use.

* fix(tool-search): at most 7 queries per call, the gateway's search limit

One tool_search call sends all its queries to the connector gateway as one
search request. The gateway answers 7 use_cases per request and returns
HTTP 502 for 8 or more (measured 2026-09-09, re-measured with one-word
use_cases: it is a count limit, not a size limit). With the client cap at
10, a model sending 8 to 10 queries lost every connector hit for that call
and saw local-only results with no error.

The shared constant splits: _MAX_QUERIES_PER_CALL = 7 for search,
_MAX_DESCRIBE_NAMES_PER_CALL = 10 for describe, which has no remote count
limit. Eight or more queries now get the existing "too many queries" retry
hint before any request is made. No chunking: one call, one request.

* fix(tool-search): the model is told that connectors__ names are manage_connections accounts

tool_search results carry names like connectors__gmail__CREATE_EMAIL_DRAFT and
manage_connections is the tool that checks and connects those accounts, but
nothing told the model the two are the same thing. A model that hit
CONNECTION_REQUIRED had to infer the fix on its own.

The tool_search description gains one sentence making the link, added at
assembly only when manage_connections is in the session's tools. Signed out
or with connectors off the tool is absent and the description is unchanged,
so it never names a tool the model cannot call. This follows the existing
rule for cross-tool references (tools/AGENTS.md): they are added dynamically
from the session's actual tool set, never hardcoded in a schema.

Tool defs are fixed for the life of a conversation, so the description is
byte-stable per conversation; this is a one-time prefix change.

Live, real get_tool_definitions() against a signed-in home: sentence present.
Same home with auth.json removed: manage_connections absent, sentence absent.

* fix(connectors): /stop halts a connector batch before the next remote call

dispatch_connector_batch runs every remote entry of a tool_call batch in
sequence. The executor only checks the interrupt flag between tools, and
the whole batch is one tool to it, so a /stop landing during entry 1 of
20 still sent the other 19 to the gateway.

The loop now reads tools.interrupt.is_interrupted before each dispatch.
Once set, it stops calling handle_function_call and fills every unstarted
slot with the loop's existing error-slot shape, code INTERRUPTED and the
message "Stopped by the user before this call was made.", so the result
envelope stays valid and the counts stay honest. Entries already
dispatched keep their real results.

Test: three connector calls where the fake client sets the interrupt on
the first execute. The client sees exactly one call and slots 2 and 3
carry INTERRUPTED. Red on the base branch, green with the fix.

* test(connections): schema assertions become dispatch contracts

test_schema_documents_wait_and_its_timeout froze description fragments
("REQUIRED", "can NOT disconnect", "Nous Portal"). A wording edit fails
it while a real regression (a disconnect that reaches the gateway) does
not. That is a snapshot of prose, not a behaviour contract.

Delete it. The requirement that wait needs connectors is already covered
by test_wait_requires_connectors. The user-only disconnect boundary is
now asserted as behaviour: action disconnect with a connector returns an
error and the fake client records no call. That replaces the earlier
de-authenticate test, which only checked that the word "dashboard"
appeared in the error text.

Test count in the file goes from 26 to 25.

* docs(tool-search): connector batches are one gateway request per entry

The user guide said a connector batch travels as one gateway request. It
does not: model_tools_connectors.dispatch_connector_batch re-enters core
dispatch per entry, and each entry becomes its own execute request in
bridge._run_remote (plus at most one literal-slug retry when the gateway
reports TOOL_NOT_FOUND under the conventional slug). The docstrings in
tools/tool_gateway/bridge.py and tools/tool_gateway/__init__.py still
described the abandoned V1 plan and claimed nothing outside the package
imports it.

Rewrite those sentences to match the code: one request per entry, in
input order, dispatched from model_tools_connectors.py, with the per-entry
approval and interrupt behaviour that motivated the split. The guide also
still showed the single-call shape tool_call(name, arguments); both
places now show the `calls: [{name, arguments}]` array the schema
advertises and note that a single local call is an array of one.

Docs only, no test.

* fix(tools): the between-turns refresh never rewrites the bridge tools

The per-turn MCP refresh folds a fresh tool snapshot into the live array
with preserve_prefix: order and membership stay, but a name present in both
takes the fresh schema. That is right for ordinary tools, whose schema is a
constant. tool_search is the one tool whose description is derived from the
session: the deferred-tool count, the embedded listing, and, on this branch,
whether manage_connections was present. A late MCP server or one failed
portal lookup (manage_connections' check_fn fails closed) changed those bytes
on the next turn, and every byte after tool_search in the cached prefix was
re-prefilled. The array also contradicted itself in that case: the flapping
manage_connections was carried forward while the description lost its hint.

The bridge entries now keep the bytes they were built with for the life of
the conversation. Nothing is lost: tool_search reads the live catalog at
dispatch, so tools that arrived late are still found; connector availability
is checked at dispatch too. The compaction-boundary rebuild (content_aware,
the one sanctioned cache break) still refreshes the description.

Consequence: connector exposure in the prompt is decided once, at agent
build, by whether the user was signed in then. That is the intended
contract.

* refactor(tool-search): normalize_tool_call_entries lives with the other argument validation

The port appended the tool_call argument parser to the tool_search facade.
The family already has tools/tool_search_validation.py for exactly this
work (schema validation of deferred call arguments), so the parser moves
there and the facade imports it. No behaviour change; the one test that
imported it now imports from the defining module.

* refactor(connectors): delete the unused batch dispatcher; _run_remote becomes run_remote

bridge.dispatch_calls and its helpers (_dispatch_calls_inner, _run_pre_dispatch,
_run_local, _error_slot, _maybe_parse_json) and the LocalDispatch / PreDispatch
seams had no production caller. Connector dispatch runs through
model_tools_connectors: dispatch_connector_batch re-enters handle_function_call
once per entry, so scope, hook, approval and middleware policy fire against each
composed name inside core dispatch, and dispatch_connector_call hands the single
planned entry to the bridge's transport function. Only tests called the batch
dispatcher, and they exercised policy seams that production never wires.

The transport function is the module's real entry point, so it drops the
underscore: _run_remote becomes run_remote, body unchanged. The module
docstring now describes the two legs that exist (availability with D32 silent
degradation, and run_remote) instead of the injected seams. Imports that only
the deleted code used are gone; merge.py is untouched because every export
still has a caller.

Tests that drove dispatch_calls are deleted where they covered the removed
seams (pre_dispatch blocks and rewrites, local_dispatch classification, mixed
batches). The literal-slug fallback, the per-entry transport failure, and the
hook rewrite reaching the gateway request body are re-targeted at
handle_function_call('tool_call', ...) with the fake client swapped in at
bridge._default_client_factory, the same seam test_connector_dispatch_policy
uses. Each re-targeted test fails when the retry is disabled in run_remote.

* fix(connectors): search keeps the twin a colliding name reaches, and says so

format_connector_name strips the toolkit prefix, so GMAIL_FETCH_PROFILE and a
literal FETCH_PROFILE on gmail both compose to connectors__gmail__FETCH_PROFILE.
describe and execute decode that name to the prefixed slug first, so the
literal twin is unreachable under it. If a vendor ever shipped both, search
could describe the literal under a name that runs the prefixed tool.

Search is the one place that sees both twins in one response. It now keeps
the twin the name reaches and drops the other with a WARNING that names both
slugs, whichever the gateway listed first. Short names stay; no marker, no
per-process map, no change to describe or execute. No such pair exists in the
live catalog today; the guard turns a silent alias into a logged one.
2026-09-10 02:21:16 +05:30
Erosika 267a6b79c8 fix(codex): gate the newest-only reasoning replay on its own Azure predicate
_is_azure_foundry_responses matches the azure-foundry provider or the
project-scoped services.ai.azure.com host. #101243 proposes narrowing it
to that host alone. The multi-item rejection this branch handles happens
on resource-level *.openai.azure.com hosts too, so the pruning gate now
uses _is_azure_responses: the provider, or either Azure host.

Refs #105369
2026-09-09 13:01:48 -07:00
Erosika 2de15babf9 fix(codex): replay only the newest sealed reasoning item on Azure Foundry
The first request of every turn after the first replays the encrypted
reasoning item of every earlier response. Azure Foundry accepts one such
item and rejects two or more with HTTP 400 "Conflicting authenticated
continuation identities". In-turn requests never hit this because
_is_post_tool_replay already suppresses replay on Azure when the request
ends on a tool result.

On Azure Foundry, build_kwargs now keeps codex_reasoning_items only on the
newest assistant row that has any. include still requests
reasoning.encrypted_content, compaction checkpoints stay on every row, and
the canonical messages are not mutated. Other Responses endpoints are
unchanged.

Rebuilding a failed session from state.db through the transport and
sending it live: unpatched, 5 reasoning items, 400; patched, 1 item, 200.

Fixes #105369
2026-09-09 13:01:48 -07:00
Erosika 5db0df28de fix(classifier): treat Azure Foundry's continuation-identity 400 as a rejected reasoning replay
Azure Foundry (gpt-6-astra) answers a request that replays encrypted
reasoning from more than one prior response with HTTP 400 "Conflicting
authenticated continuation identities", code invalid_value. The classifier
mapped that to a non-retryable client error, so the turn aborted and every
later turn in the session failed the same way.

Map the message to invalid_encrypted_content. The existing recovery in
turn_recovery then strips the replay state and retries once.

Refs #105369
2026-09-09 13:01:48 -07:00