33 Commits

Author SHA1 Message Date
Robin Fernandes 2a94ca80e7 fix(free-tier): review round 2 — route-gate the allowance verdict, keep policy/billing 403s, pool the provision RPC, guard the retry race
Should-fix
- _is_genuine_nous_rate_limit: the structured rate_limited verdict counts only
  on the welcome host; a paid-host 429 keeps main's exhausted-bucket rule.
- _nous_welcome_tier: the route-keyed dark-tier 403 applies only to a 403 that
  matches neither the content-policy nor the billing patterns, so a safety
  refusal or billing wall on the welcome host keeps its own recovery.
- free_tier.provision joins _LONG_HANDLERS (a forced mint + lock waits +
  re-inventory no longer block the RPC reader).
- retry_bootstrap_mint: under the lock, a build that found no identity never
  overwrites a record that has one (the loop racing the user's click).

Simplifications from the review
- _raise_for_anon_status is a (status, error) table; retryable derives from
  ANON_TERMINAL_CODES once (a bare 401 on sign-up now rides the ladder
  instead of dying for the process).
- classify_mint_exception is public and pure; the hand-built failure dict in
  free_tier.provision is gone (the memo is the one source).
- SetupRecord carries the memo payload as one `failure` dict instead of three
  unpacked fields.
- _welcome_surface_kind is a closed table with a "refused" default;
  _welcome_outage_copy excludes the classifier's `unknown` catch-all.
- FREE_TIER_RATE_LIMIT_CHAT is CARD + the sign-in tail, not a slice.
- Copy tests assert the contract (model named, tail present/absent) instead
  of freezing whole sentences.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 2e3b8ecc82 fix(free-tier): the desktop card body leaves the "To sign in" tail off; the button is the door
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 16def7b8cc fix(free-tier): the desktop renders a free-tier refusal as its own card, not an OAuth re-login
A welcome-tier 403 classifies as auth_permanent, so the desktop's error
surface mapped it to "Your Nous Portal sign-in expired" with a Nous Portal
re-login button — the chat sentence never reached the user. Terminal results
on the free route now carry a structured free_tier block (kind + the chat
sentence); agent/error_surface.py turns it into a free_tier_<kind> code on
the provider layer with the sentence as `message`. The desktop gives those
codes their own titles, shows the backend sentence as the body, and offers
"Sign in with a Nous account" (the free-tier dialog) instead of the OAuth
re-login, with Retry only where a later send can succeed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 59fad62a40 fix(free-tier): review follow-ups — read the classifier's context, never replace a locked identity, re-inventory on retry
Correctness
- The welcome-tier recovery hooks (model_not_free move, wrong-host heal) and
  the long-wait rate-limit check read the turn's extract_api_error_context()
  dict, which never carries welcome_refusal / welcome_route. They now read
  classified.error_context, where _nous_welcome_tier parks them; the guard
  records the classifier's reset_at. Tests drive the real classifier and the
  real extractor so the two-context boundary is exercised.
- The connector path caught every AnonCredentialDead and re-minted; a locked
  account (anon_account_locked) is now retired without replacement, matching
  the inference resolver.
- A background bootstrap retry reused the boot-time provider inventory; it
  re-inventories, so a provider connected during the cooldown keeps
  inference.
- The desktop's setup.ready listener only refreshes an untouched picker
  (oauth mode, no local endpoint, idle flow) and re-checks after the
  readiness round, so an API-key form opened meanwhile is never dismissed.
- /__log on the rehearsal server sent its response while holding the state
  lock that _send re-acquires; the log is copied out first.

Reductions
- One shared FakePortal / install_portal (tests/hermes_cli/anon_portal.py)
  behind both free-tier fixtures, with a single httpx.Client transport seam.
- The rehearsal server's static inference answers are a table; dead
  scaffolding (REAL_PAID_URL, claim_codes, the no-op dead_once branch,
  extra_headers) removed.
- Setup-notice copy is a code-to-key map; its test uses real codes (the old
  loop built nonexistent ones and only exercised the fallback).
- The ineffective FreeTierErrorCode union is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes 51e39af967 feat(free-tier): ruled behaviour for every welcome-api failure, with friendly copy and a fault-injecting rehearsal server
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.

Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
  temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
  into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
  honours the server's wait, climbs a short ladder when the service is
  unreachable, never retries terminal codes, and yields to the user's own
  retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
  background loop retries transient failures and re-announces setup.ready.
  setup.status and free_tier.status expose the block; free_tier.provision is
  the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
  the route); model_not_free moves onto the gateway's alternate once;
  anon_on_paid_host re-reads the route once; a long rate_limited refusal
  trips the cross-session guard; a locked account is retired but never
  replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
  retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
  service is off" (what is unavailable is using Hermes without signing in,
  and signing in is free), no jargon, spoken waits.

Desktop
- A setup-failure notice above the provider picker: one sentence per code,
  a retry when the backend says one can work, the sign-in pointer only when
  the account service answered at all. The overlay re-checks readiness on
  setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.

Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
  real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
  (dev-only, env-only) lets the route rules treat it as the welcome host.
  Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
teknium1 23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
brooklyn! 4f1966edac feat(desktop): connector cards that wait for the sign-in, and a guided first launch that holds together (#108292)
* feat(desktop): give Button a loading prop that swaps label for spinner without layout shift

The label stays in the box, invisible, and the spinner is absolutely
centred over it, so a Connect or Approve button keeps its width while it
works instead of collapsing to a spinner. The approval bar had the same
thrash and moves onto it.

* refactor(desktop): one consent card for connectors and MCP setup

McpSetupTool rendered its own copy of the connector card's markup. It now
renders ConnectorCard for the pending question and ConnectorSummary once
settled, and the card gains what MCP needed: keyboard accelerators, a
source line, a question heading. The card also gets an avatar variant
(40px mark in the left gutter, text and buttons on one column) and a
collapseWhenSettled switch so a connector can stay a full card with a
green Connected pill in the action slot while MCP keeps its one-line
summary. Brand marks for Gmail, Calendar, Drive, Discord, Telegram and
Spotify; Slack via Tabler because simple-icons dropped the mark.

* feat(desktop): connector card drives the agent through manage_connections wait

The offer used to end in a Continue in chat button, and the agent, seeing
an unconnected status, would improvise around the app. Now the card does
what the TUI does. Clicking Connect opens the browser and sends one hidden
line telling the agent to park in manage_connections action=wait for that
slug and to never call connect again (a second link cancels the one being
signed into). Not now sends its own line. A hidden request that lands
while the turn is busy steers it, or queues if the turn just ended.

Which call owns the live card changes too: consecutive calls naming the
same apps are one exchange (connect, the wait, the status that follows),
and the first of the last exchange is the card, so the agent's wait no
longer demotes the card mid-authorization and mints a fresh one below it.
A targeted ask renders one or two bare cards; only a real catalog gets the
header, search and refresh.

* feat(desktop): onboarding connects apps in chat and keeps tasks finishable without them

The welcome chat knew connectors only as preferences to pick and wire up
later, so asked to connect Gmail it invented a Settings page that does not
exist. Both scripts now carry one rule set: status once, one batched
connect for every app named, the card is the ask so write a line and end
the turn, never route around a declined app with another client or
credential. The build handoff checks real connection status instead of
asserting none are connected, and the first task must be finishable, not
free of, the apps they picked. The connectors card explains what
connecting means and reports the count on its Continue button.

* fix(tools): resolve the Nous identity for share_auth profiles in the connector gate

A profile created with share_auth has no auth.json of its own and signs
in through the root store. Every other credential reader falls back to
the global root; the connector gate read HERMES_HOME/auth.json directly,
saw nothing, and stripped manage_connections from the profile's tool
list, so the welcome chat's agent truthfully reported the tool missing.
The gate now goes through get_provider_auth_state.

* fix(agent): name a provider retry backoff on the live status line

The retry status is buffered and replays only when every retry fails, so
during a 60s backoff after a 5xx the user saw a bare spinner. Right after
a connector sign-in landed this read as the agent going silent. The
backoff now also rewrites the live wait notice, which the desktop already
renders in the thread status row; it is transient and clears on recovery.

* test(desktop): connector rehearsal launcher and flagged connector spec

connector-rehearsal.mjs starts the real desktop and backend under a fresh
HERMES_HOME with no copied credentials, a fixed Vite port and CDP on 9344,
so the onboarding connector flow can be driven end to end by hand or from
outside. The Playwright spec covers the flagged connector step.

* fix(desktop): send the agent back into wait when the user keeps waiting after a timeout

The card's Keep waiting re-entered the poll but the agent's own wait had
timed out too and nothing told it to go back in, so it would start
talking mid-authorization. keepWaiting now fires onWaiting like connect
does. Tests also pin that an expired or revoked grant asks the gateway
for reconnect, not connect.

* style(desktop): blank lines in connector-flow test per lint

* feat(desktop): HERMES_SKIP_INTRO=1 / --skip-intro skips the first-run film

The intro is a one-time reveal, so anyone rehearsing the guided chat behind
it sits through it on every fresh HERMES_HOME. The flag rides the existing
launch-flags path (main → preload → renderer) next to guestOnboarding and
only gates isIntroRevealEnabled; the backend never sees it. The rehearsal
launcher sets it.

* fix(desktop): onboarding card Continue stays Done after the transcript rebuilds

The card kept its Done flag in component state. The hidden submit and the
turn-end hydrate both rebuild the message list, so the card remounted with
the flag false and Continue came back live, letting a step be answered
twice. The committed steps now live with the other onboarding answers,
keyed by step, and the first-build chip pick rides the same store.
remember_onboarding projects by key, so the new field never reaches USER.md.

* fix(desktop): no provider picker or free-tier chip over the guided first launch

Two sign-in surfaces leaked into the guide. A credential probe on the
setup profile (a free-tier token mid refresh, a session before its runtime
settled) hit requestDesktopOnboarding and dropped the provider picker over
the chat the user was in; and the statusbar free-tier chip sat there
offering a second sign-in the whole time. Both now yield while the gate
phase is cinematic, guided or handoff. The free tier is the provider for
those phases, and the guide offers sign-in on its own ready screen.

* fix(desktop): onboarding connector picks are real catalog slugs

The picker offered Spotify, GitHub and Stripe, none of which the deployed
connector catalog carries, and spelled Calendar and Drive with hyphens the
gateway does not use. A pick the build chat could not honour ended as
"Spotify isn't in the connector list" after the user had been told to
expect it. The list is now twelve slugs from the live status catalog,
spelled as the gateway spells them; GitHub is out (the terminal has git
and gh), chat channels stay on Messaging. Marks for the new entries; the
Google marks answer both spellings. The build runbook offers the picked
connections in its first turn rather than after the work is underway.

* fix(desktop): the free-tier ready screen never interrupts the guided chat

A readiness round fires when the layout pick assembles the window, and it
raised the free-tier ready screen over the conversation: the user was
dropped into the main app, dismissed it, and came back to a card they had
already answered. The guide is the introduction. The ready screen now
yields while the gate is cinematic, guided or handoff, and the notice is
acked the moment the guided chat takes the screen, not only when the film
does, so a skipped film no longer leaves it pending.

* feat(desktop): tour options that lead to building, and a fork that follows the tour

"Just the basics" and "Show me around" read as a click-through with no
exit; "I'll figure it out" read as declining help. Now Quick tour, Show me
everything, and Skip, let's build something. The script also folds the
fork into the same turn as the tour, so when the user closes the overlay
the next ask is already waiting instead of a transcript that ends on the
tour call.

* feat(desktop): the onboarding connector picker reads the live catalog

A hardcoded list, however carefully copied from today's catalog, is the
next drift. The picker now asks connectors.list through the same
session-owned RPC the connector cards use and offers exactly what the
gateway carries: a curated lead order puts the everyday apps first, chat
channels stay on Messaging, everything else is reachable by search. The
picks are gateway slugs, handed straight to manage_connections. No
catalog (toolset off, gateway unreachable) ends the step honestly with
Skip instead of inventing apps.

* test(desktop): the guided first launch never forces a sign-in

The acceptance criterion the guided onboarding was built to, as a test:
while the gate is cinematic, guided or handoff, the provider picker does
not open and a credential warning is dropped rather than deferred to the
next send. Outside the guide the picker opens as before. Red against the
tree before the guards landed (6 of 9).

* fix(desktop): a relaunch mid-guide resumes the guide, in the guide's shape

Closing the app during the guided first launch and reopening it booted the
normal shell around the persisted solo layout: the connecting splash, the
stock composer and model picker, a small window whose sidebars would not
open, while the gate still read guided. The gate now queues a kickoff for
the guided phase too (the kickoff adopts the existing guide chat by title),
takes the solo shape before the gateway opens rather than after, and the
connecting overlay yields to the guide's own opening. A typed reply in the
composer now closes an ask card and the first-build chips the same way a
click does; the layout card's Continue comes back Done.

* style(desktop): one answeredAfter helper for the ask card and first-build chips

* fix(desktop): the guide takes its shape on the tick the film ends, not after the window shows

Between the film and the greeting the full-size shell painted for a beat:
finishIntroReveal showed the main window, then the kickoff shrank it once
the setup profile answered. The listener on the intro's hidden edge now
takes the guide's shape (solo layout + small centred window) synchronously,
so the window is already the guide when it is shown. One takeGuideShape
owns the pair; kickoff and the boot gate call it idempotently.

* style(desktop): the 'nothing connects yet' line reads first on the connectors card
2026-09-12 10:00:06 +05:30
Siddharth Balyan 4bdd64b334 The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract

The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an
anonymous account exactly one model on its own host, refuses everything else with a structured
429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and
tells a signed-in account that still asks for `nous/welcome` what to switch to in an
`x-nous-model-switch` header. Four client-side gaps against that contract:

- Auxiliary calls were refused on every session. The auxiliary client asked the welcome host
  for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free`
  before each fallback. On the welcome host it now uses `nous/welcome` (its backing model
  covers auxiliary work) and skips Nous for vision, which the welcome model does not take.

- The structured 429 body was never read. The classifier now parses `reason` /
  `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are
  non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited`
  are rate limits that honour `retry_after` and never rotate the free tier's only credential.
  The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and
  fall back instead of retrying or re-exchanging. The terminal paths say what happened and
  name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal).

- The `x-nous-model-switch` header was ignored. The chat-completions transport records it
  beside the rate-limit and credits headers; the next call moves the session, and the config
  default when it still names `nous/welcome`, to the backing model the gateway named.

- A guest fell back to the paid host. With `inference_base_url` absent from the exchange or
  outside the host allowlist, routing defaulted to inference-api, where every request is a
  400. A guest now defaults to the welcome literal at the exchange, in the shared store's
  shape, and in effective routing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8)

* feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use

A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided
setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier
picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime
rung now sets it up there instead of failing "not logged in", so the guided chat no longer races
the root profile's first-run mint.

nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever
nothing else is configured; "on-request" mints only when the free tier is asked for by name
(nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers —
the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the
connector token path — still adopt what the shared store holds, so every profile follows the one
identity the guided setup created, but never create one on their own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c)
(cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165)

* feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit

Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity
(nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is
configured. "explicit" means Hermes never creates one on its own: the only creator is the new
provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup
on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and
before the guided chat exists — so the identity lands in the root store every profile reads
through and is there before any session asks for nous/welcome. That closes the race against the
backend's own setup, and makes "only when the setup-bot flow is used" literally true.

The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome
(the hermes model row, --provider nous), which treated a model name as intent and was broader
than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not
logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to
sign in from. Implicit callers still adopt an identity the shared store holds, and a retired
credential is replaced (a continuation, not a creation).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0)

* fix(auth): remove the nous.guest_setup knob; the free tier is created on first use

`nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its
default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`,
unknown values read as `auto`, and under `explicit` a CLI-only install could never get an
identity, which contradicts the first-run contract (first command mints, then chats).

The mint race the knob accompanied is already benign: every caller takes the profile lock then
the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup
win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which
stays. `nous.guest` remains the only free-tier policy.

Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through
`ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call
sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config
entries. The three policy tests that hold regardless of the knob are kept under
`TestExplicitProvision`; the two that only tested the knob are deleted.

(cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261)

* fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent

When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway
serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch`
moves the live session to that model and moves `config.yaml`'s default off the alias in the
same step. The messaging gateway's fallback-eviction check compares the agent's model with the
config default and evicts on any mismatch that is not a /model override, so when the config
write did not land (unreadable config, lock) the cached agent was evicted once per turn, and
prompt caching with it.

`apply_model_switch` now stamps the alias it moved the session off on the agent, and
`_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as
deliberate, beside the existing /model override case. The check takes the agent and the config
model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them.

(cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954)

* fix(auth): the free tier outranks implicit host credentials in provider resolution

On a fresh install with a leftover ~/.aws profile, resolve_provider("auto")
reached the Bedrock rung before the free-tier rung, so the first turn ran on
Bedrock and failed 403 while the free tier was still being minted in the
background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s,
three retries, no answer; the next process then switched to nous/welcome.

The free-tier rung now sits directly above the Bedrock chain: when nous.guest
is on, an existing free-tier identity answers, else a blocking mint runs, and
only then does the boto chain get a say. Everything above is unchanged and
still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter
pool, a logged-in active_provider. nous.guest: false skips the rung, and a
failed mint still falls through to Bedrock and the no-provider guidance.

Tests: six precedence cases (identity present, fresh mint, free tier off, env
key still wins, sign-in still wins, failed mint falls through). The opt-out
test now neutralizes the AWS chain like the precedence tests do; on a machine
with ~/.aws it was failing for the same reason as the bug.

Live after the fix, same Mac, AWS credentials visible, isolated shared store:
identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in
11 s.

(cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a)

* fix(auth): review follow-ups for the free-tier rung (NS-829)

- tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches
  the free tier off; its contract is the boto chain, and the free tier now
  sits above it.
- gateway/run_notifications.py: the free-tier startup line reads auth.json
  before consulting the resolver, so a gateway boot on a machine with AWS
  credentials never mints or refreshes over the network.
- hermes_cli/anon_auth.py: module docstring says where the free tier sits in
  the ladder instead of "the ladder is untouched".
- tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of
  six (parametrized ladder cases; a failed mint that returns None or raises
  falls through to Bedrock).

scripts/run_tests.sh on the five affected files: 147 passed, 0 failed.

(cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11)

* feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone

The free tier is pre-GA. Until GA it must not exist for anyone who did not
ask for it: no identity minted, no portal traffic, no free-tier copy on any
surface. One environment variable now decides that, and one function reads it.

`guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly
"1"; only then does `nous.guest` (the user's off switch) get consulted. Every
free-tier site already funnels through `guest_enabled()`, so the gate closes
minting, routing, connector entitlement, status lines and the picker row in
one place. With the variable unset, `resolve_provider("auto")` on a fresh
install raises `no_provider_configured` exactly as upstream does.

`HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the
gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted
identities as a side effect of provider resolution, and `_has_any_provider_
configured` read them ahead of every other check, making the CLI a second
reader of a flag that must have exactly one. `_forced_new_done` and the
`force` parameter of `_reconcile_and_provision` go with them.

Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in
b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847).

Not a user preference: the variable is never written to config.yaml or .env
and never shown in setup. It is deleted at GA together with its comment in
anon_auth.py. This is a deliberate, temporary exception to the "no new
HERMES_* env vars for non-secret config" rule.

Tests: fixtures set the gate instead of deleting the old lever; one new
invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "",
"0", "true" and "new" all leave the tier off with zero portal calls, red on the
previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the
feature.

* feat(auth): the free-tier identity is created in one place, at boot; every other site is a read

Before this commit eight sites could create a Nous free-tier identity as a
side effect of something else: resolving a provider, the CLI's first-run
check, the CLI's session setup (in the background beside an own key), a
connector bearer read, the desktop polling `free_tier.status`, the sign-in
precondition, the desktop's `free_tier.provision`, and the dead-credential
re-mint. A poll could mint. Provider resolution could hit the network. Two
of them raced each other on a fresh install.

Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator.
`hermes serve` runs it on a daemon thread from `_lifespan` beside the other
background boots; `cmd_chat` runs it synchronously before the first-run
guard. It inventories credentials first (`resolve_provider("auto",
skip_free_tier=True)`: what would carry inference if the free tier did not
exist), creates the identity only when `guest_enabled()`, resolves inference,
records a `SetupRecord` in process memory and broadcasts ONE `setup.ready`
event. It runs on every boot; only the mint is gated.

`ensure_portal_identity` now requires `explicit=True` and raises otherwise.
Its callers are the bootstrap, the desktop's `free_tier.provision` (the
explicit retry when the boot could not create the identity) and the two
dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`,
`managed_tool_gateway._replace_dead_guest_token`). The background thread
path and `provision_free_tier` are deleted with their last callers.

Reads that used to mint and now only read: `auth.py::resolve_provider`
rung 7 (an existing identity still outranks the Bedrock chain, NS-829
ordering kept), `main.py::_has_any_provider_configured`,
`cli_agent_setup_mixin._ensure_runtime_credentials`,
`managed_tool_gateway.read_nous_access_token` (no identity -> None),
`anon_sign_in.run_sign_in` (no identity -> Unavailable),
`methods_free_tier` `free_tier.status`.

`setup.status` answers from the record for the launch profile, blocking up
to 8 s while the bootstrap is in flight so a client's first poll lands after
the identity exists rather than racing it; a named profile, or a process
that never ran the bootstrap, keeps today's live probe. The record's fields
ride along additively (`ready`, `free_tier`, `other_providers`,
`inference_provider`).

Identity and inference are decoupled (NS-845 Q1.3): the mint sets
`active_provider="nous"` only when the inventory found nothing else usable
(`_mint_locked(carries_inference=)`); an adopted account always does. A token
refresh no longer re-elects the provider it refreshed
(`_save_provider_state_to_source` writes credentials, not the user's
choice) — that write was how an own-key install ended up on the free tier
after the first connector call.

Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check),
bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint),
62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and
`provision_free_tier`), and a04b05260c (blocking mint in the resolver).
Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847.

Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps
inference; reads never reach the portal; a refused mint is memoised),
`free_tier.status` fails loudly if it ever calls the creator, the resolver
stub fails loudly if resolution ever mints, `setup.status` reads the record,
`skip_free_tier` proves the inventory question. The three sign-in tests for
the deleted pre-mint collapse into one (`no identity -> Unavailable, zero
portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and
off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI
matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20.

* fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup"

A free-tier identity carries $0 by design, so the portal seed reports
`paid_access=False` for it. `is_free_tier_model` did not know the welcome
host, read that as a depleted account, and every free-tier turn ended with
the credits-depleted notice telling the user to top up an account they do
not have.

Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host
(`anon_auth.route_is_welcome_host`) is the free tier. The host is the
evidence, not the model name: the paid inference host can serve
`nous/welcome` to a named account and that account's depletion is real, so
`("nous/welcome", <inference host>)` stays False. Local data only, like the
three rules above it.

Restores the two contracts dropped by hermes-magic 674e11d1eaa (the
prototype line ran without unit tests): the welcome host is free without
any pricing evidence; the model name alone is not. The first is red without
this fix.

* fix(copy): free-tier text stops promising a connector transfer and never names the config key

Sign-in copy on every surface said "Sign in to keep your connectors" and
ended with "Your connectors are kept." The transfer registry that would
make that true is empty (NS-821): nothing carries over today. The copy now
says what signing in does give ("unlock more models and tools") and the
completion line names the account, not a transfer. The docs page loses the
"connectors carry over" paragraph for the same reason.

The picker's off-state line exposed `nous.guest: false` and the word
"guest"; user copy names the free tier only (R-USR-1).

The docs page gains the pre-rollout note: until GA nothing on it happens
without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and
"replaced on next use" sentences now describe the boot bootstrap.

zh is a strict locale: the `freeTier` block was English placeholder text
copied from `en`; it is now Chinese. `connectorsKept` is renamed
`completedBody` since it no longer talks about connectors.

* feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn

The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1"
as on. Until now nothing in the desktop set it, so a packaged app could
never turn the free tier on, and a backend spawned by the app could
disagree with the app about whether the tier was live.

`electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled`
is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has
`--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch
into a module constant. `desktopBackendSpawnEnv` wraps every backend env
as the outermost call and writes the flag LAST, as "1" or an explicit "0",
so no earlier spread (`process.env`, `backend.env`) can resurrect a stray
value from the parent shell.

Stamped onto all three spawn sites: the primary `serve` spawn, the pooled
per-profile spawn, and the remote SSH `exec env ...` command (which gains
` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and
the backend probes are not backend spawns and do not get it: a
`hermes --tui` typed in the pane must not mint.

The renderer learns the same fact read-only through the existing
`hermes:launch-flags` sync IPC (`guestOnboarding`) and preload
(`window.hermesDesktop.guestOnboardingEnabled`).

Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding`
maps to it in main). Two invariant tests on the pure helpers: only "1" or
the argv flag enables; the spawn env carries "1"/"0" as the last word and
preserves every other key.

* feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll

The backend's boot bootstrap now announces `setup.ready` once, after it has
created (or refused) the free-tier identity and resolved the inference
route. The renderer used to discover both by polling `setup.status`,
`setup.runtime_check` and `free_tier.status` every 60 s from
`useStatusSnapshot`; a fresh install's chip, notice strip and onboarding
overlay could sit stale for up to a minute after boot, and three RPCs a
minute per window kept asking a question whose answer changes only at
boundaries the backend already announces.

`handleLifecycleEvent` routes `setup.ready` (active source only, like
`skin.changed`) to `notifySetupReady()`, a one-shot tick atom in
`live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens
to it and runs one readiness round at once (`setup.status` +
`setup.runtime_check` + `free_tier.status`). The readiness legs also run
once on open and on return from another app, as today. The 60 s tick keeps
only `getStatus()`.

`SetupStatusSnapshot` types the record's additive fields (`ready`,
`free_tier`, `other_providers`, `inference_provider`); readiness semantics
are unchanged and still key on `provider_configured` + `runtime_check`.

Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one
refresh from the active source and none from another; the snapshot hook's
contract is three legs on open, one leg on the tick.

* fix(cli): the banner names the free tier's model instead of "no model configured"

The welcome banner prints before credentials resolve, so on a fresh install
`model` is empty and the banner said, in red, "no model configured — run
/model or hermes setup". Under the free tier that is false: the route is
already known from local state (identity on disk, tier on), and the first
message will run on `nous/welcome`.

`_banner_left_lines` now asks the route the same question when `model` is
empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous
Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the
banner's 'no model configured' line reads the resolved route").

Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`;
gate off -> the red line, zero portal calls.

* fix(aux): vision on the free tier uses nous/welcome too

The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the
backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the
repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the
welcome host would have sent every image step past the free tier for no reason, so the
auxiliary client pins the route's one model for every lane. A backing model that takes
no images answers with the upstream's own error, which the ladder handles as it always has.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415)

* fix(gateway): hermes gateway run is a boot owner of the free tier too

Rung 5 made every demand-time free-tier site a read: resolve_provider,
the connector token, the /login precondition. That is only correct if
every process that can reach those sites ran the bootstrap first. The
CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone
messaging gateway did not. A fresh HERMES_HOME with the gate on and
`hermes gateway run` reached provider resolution with no identity to
consume, and /login returned Unavailable. Reported by @andrexibiza on
#107697 (P1).

GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an
executor thread right after startup recovery and BEFORE any adapter
connects, so a fast first DM cannot arrive with nothing to resolve. It
is its own step, not part of the turn-machinery warm-up: the warm-up is
an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0);
the bootstrap is correctness and must always run. With the gate unset it
is a local inventory and no network.

Live, real GatewayRunner.start against a fake portal in a fresh home:
  gate on   -> 1 create, identity persisted, resolve_runtime_provider=nous,
               /login precondition sees the identity
  gate off  -> 0 portal calls, no identity, no_provider_configured
Before the fix the gate-on row was identical to the gate-off row.

Test: the bootstrap seam runs before _start_prefilter_platforms and
delegates to the one creator. Red on 5554eb6993 (no seam), green here.

---------

Co-authored-by: Robin Fernandes <robin@soal.org>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 03:45:33 +05:30
Teknium 97ca90f184 fix(desktop): expired OAuth grant shows a one-click 'Sign in again' instead of a retryable Provider error
A rejected OAuth token (HTTP 401 'User not found' from Nous Portal, Codex,
xAI…) reached the desktop error card as 'Provider error' with Retry as the
first action, which just replays the same dead credential.

Backend: nonretryable_client_error_result dropped failure_reason /
failure_retryable, so error_surface classified every non-retryable 4xx as a
retryable provider failure. It now stamps the classifier verdict like the
max-retries path, and auth-layer descriptors carry auth_kind
(oauth|api_key, derived from the provider catalog tab) + provider_label.

Desktop: an auth/oauth surface renders 'Authentication error', explains that
the <provider> sign-in expired/was revoked, and offers 'Sign in to <provider>
again' which launches that provider's existing onboarding OAuth flow scoped
to the failed session's gateway profile. Retry stays as the follow-up click.
Re-login to the provider already in use keeps the current model instead of
swapping in the recommended default.
2026-09-09 16:30:24 -07:00
liuhao1024 cabe430a25 fix(agent): derive tile-aware shrink cap from Codex patch-budget 400s
The classifier fix routes Codex patch-budget 400s into the shrink
recovery, but _image_error_max_dimension still returned None for the
Codex wording, so the recovery fell back to the 8000 px default cap
and skipped images between ~5542 and 8000 px that already exceed the
30000-tile budget — burning the single shrink retry without shrinking
anything. Parse the reported patch limit and convert it to a per-side
pixel cap of isqrt(limit)*32 (5536 px for 30000 patches), which keeps
a square image under the budget.
2026-09-09 12:06:53 -07:00
kshitijk4poor 44e52e7524 fix(agent): normalize zero-cooldown semantics, dedupe test driver
/simplify-code pass on the Retry-After salvage stack:

- compute_error_backoff now decides "no usable cooldown" exactly once:
  a parsed 0.0 (retry-after: 0, or an HTTP-date in the past, which the
  shared parser clamps to 0) is treated as absent instead of falling
  through an accidental falsy check — prevents a hot-loop retry and
  keeps the sentinel semantics uniform (is None / is not None at all
  four sites).
- Comment corrected: the sibling unwrap lives in
  extract_api_error_context, not _extract_rate_limit_context.
- Test driver deduped: one _retryable_error factory + one _drive_once
  shared by the 429 and 524 tests; added the over-cap (3600 → 600) and
  no-cooldown fallback rows requested in the #103722 review.

Mutation checks: nested-unwrap neutralized → nested row red; cap
removed → over-cap row red; restored → all 298 green.
2026-09-06 21:37:47 +05:30
kshitijk4poor c6ef075613 fix(agent): parse nested retry_after bodies and emit long 5xx cooldowns
Salvage follow-ups on #88236 (krunkosaurus):

- Some providers nest the cooldown as body["error"]["retry_after"] (the
  same unwrap _extract_rate_limit_context already uses); only the top-level
  shape was read, so those errors silently fell back to jittered backoff.
- A 5xx Retry-After can reach the 600s cap; that wait was buffered
  (replayed only on terminal failure), leaving the user silent for
  minutes. Long provider cooldowns now emit immediately, mirroring the
  zai_coding_overload_long path. Jittered waits keep the old buffering.

Test widened with the nested-body parametrize row; proven red when the
unwrap is neutralized.
2026-09-06 21:37:47 +05:30
Mauvis Ledford a227484906 fix(agent): honor Retry-After on retryable 5xx
Retryable server errors can carry provider cooldowns just like 429 responses. Parse Retry-After from response headers or structured error bodies before falling back to jittered backoff, and cover HTTP 524 behavior with runtime regression tests.
2026-09-06 21:37:47 +05:30
Teknium 12871bd01e feat(openrouter): per-model provider_routing.models.<id> overrides
`provider_routing.models.<model-id>` now takes the same only/ignore/order/sort/
require_parameters/data_collection keys and overlays the flat provider_routing
values whenever the agent is on that model. Resolution lives in the one
chokepoint every request path already uses (_provider_preferences_for_agent),
so CLI, gateway, TUI/Desktop, cron, /model switches, fallback activation and
delegated children on another model all honour it with no per-surface plumbing.
Matching is spelling-tolerant, sharing _canonical_model_variants with
agent.reasoning_overrides.

The OpenRouter profile's speed-tier pin no longer overwrites an explicit user
`only` on the BASE gpt-6-astra slug: the pin exists to keep default routing off
flex/fast, and a user pin is the stronger intent (only: [openai] stays [openai]
instead of becoming [openai, azure, azure/us]). Tier slugs (-fast/-flex) keep
owning `only`.

Live A/B (config only: {gpt-6-astra: [openai], claude-fable-5.1: [anthropic]}):
main sent {"sort":"price"} for fable and OpenRouter served it from Azure; with
this change it sends {"only":["anthropic"],"sort":"price"} and Anthropic serves it.

Schema proposed in #24495 (samplesabotage) and #100711 (Artemonim); this is a
slim chokepoint implementation of that design.

Co-authored-by: samplesabotage <samplesabotage@users.noreply.github.com>
2026-09-06 02:18:14 -07:00
Teknium 89fbd5d4d3 simplify(compat): conversation_loop — drop 9 re-exports, repoint 7 callers, 35 test sites 2026-09-03 13:12:23 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 022785a541 Merge origin/main (63279301bc): reasoning-mandatory 400 recovery folded into turn_recovery/error_classifier/models_reasoning_caps 2026-09-03 05:36:21 -07:00
Teknium 2f4534eaaa refactor(agent/turn): _NONRETRYABLE_LABELS table shared by terminal status + fallback announce; docstring trims 2026-09-02 19:39:27 -07:00
Teknium de02f3f8bf refactor(agent/turn_recovery): lift long-context tier cap into _cap_long_context_tier 2026-09-02 19:28:38 -07:00
Teknium 011849540e refactor(agent/turn_recovery): 401 gate hoisted in _refresh_credentials_after_401; flatten retry-after parse 2026-09-02 19:22:17 -07:00
Teknium 4f59c57d2d refactor(agent/turn_recovery): _blines buffered-trace helper, _failure_hint_for table, inline display_hermes_home 2026-09-02 19:13:57 -07:00
Teknium cf59e5ad1f refactor(agent/turn_recovery): lift format-recovery strips into _recover_format_errors 2026-09-02 19:02:02 -07:00
Teknium b925acb5a9 refactor(agent/turn_recovery): abort_turn_on_interrupt shared by backoff sleep and handle_api_error 2026-09-02 18:55:11 -07:00
Teknium 200f63ef1e refactor(agent/turn_recovery): lift per-provider 401 credential refresh into _refresh_credentials_after_401 2026-09-02 18:49:06 -07:00
Teknium 74bc031281 refactor(agent/turn_recovery): _vlines/_plines print helpers, lift unicode/401-diagnostic/auth-guidance phases, dedupe failed-turn dicts, flatten validate_response_shape 2026-09-02 18:29:20 -07:00
Teknium c94ced6225 refactor(agent/turn): AST-neutral bracket/signature packing across r3-08 slice 2026-09-02 18:01:27 -07:00
Teknium 22cba7dd41 refactor(agent): extract classified-error routing (compaction gate, long-context tier, eager/auth fallback, Nous 429) into turn_recovery.route_classified_error 2026-09-02 13:30:23 -07:00
Teknium df42f37e09 refactor(agent): extract response-shape validation, invalid-response diagnostics and Codex incomplete continuation 2026-09-02 13:30:19 -07:00
Teknium 8b958e1b0f refactor(agent): unify interruptible backoff sleeps and fallback-restart arming; extract compute_error_backoff 2026-09-02 13:30:18 -07:00
Teknium 5543e7ae0f refactor(agent): resume — extract API-error attempt logging into turn_recovery.log_api_error_attempt (verified partial work) 2026-09-02 13:30:17 -07:00
Teknium 25b165add5 refactor(agent): extract terminal API-failure result builders into agent/turn_recovery.py 2026-09-02 13:30:16 -07:00
Teknium fefc471deb refactor(agent): extract post-classification one-shot recovery chain into agent/turn_recovery.py 2026-09-02 13:30:15 -07:00
Teknium 79469656e7 refactor(agent): extract pre-classification API-error recovery into agent/turn_recovery.py 2026-09-02 13:29:47 -07:00