An external-process provider is an agent CLI Hermes drives over stdio rather
than an HTTP endpoint. Three things about it were spelled out for one vendor,
and each was a hard stop for any other:
* ``resolve_provider()`` gates on ``PROVIDER_REGISTRY``. Its auto-extend from
``providers/`` covered api-key providers only, so an external-process profile
never entered it and ``hermes -m <that provider>`` died with "Unknown
provider" before a client was ever built.
* ``resolve_runtime_provider()`` keyed the external-process branch on the
literal ``"copilot-acp"``, so anything else silently fell through to the
OpenRouter default instead of its own runtime.
* ``resolve_external_process_provider_credentials()`` hardcoded the binary
(``copilot``), the argv (``--acp --stdio``), the env var names and the
placeholder api_key — so a third-party provider would have been handed
another vendor's CLI.
Now the profile carries what only the provider knows — ``process_command``,
``process_args``, ``process_command_env_vars``, ``process_args_env_var`` — and
the three core paths key on ``auth_type == "external_process"`` instead of a
name. copilot-acp's values move into its profile verbatim, so
``HERMES_COPILOT_ACP_COMMAND`` / ``COPILOT_CLI_PATH`` /
``HERMES_COPILOT_ACP_ARGS`` and its ``copilot-acp`` api_key placeholder behave
exactly as before; the new tests assert that alongside the out-of-tree case at
every step.
The error for a missing binary now names the provider and its own env override
instead of telling every user to install GitHub Copilot CLI.
Co-Authored-By: Junie <junie@jetbrains.com>
N processes sharing one Nous OAuth pool entry hit the hourly expiry
together; each force-refreshed, each rotation invalidated the token a
sibling had just adopted, and processes that lost the auth-store flock
race had their only entry benched ("matched no nous entry ... pool size
0") — ~120 sessions surfaced 401 'out of funds' on Sep 2 2026.
- resolve_nous_runtime_credentials(stale_access_token=): under the store
lock, skip the refresh POST when the on-disk token differs from the one
that failed and is usable (a peer already rotated) — adopt instead.
- credential_pool nous path: adopt a peer-rotated key after the pre-sync,
pass the failed bearer through, and treat a lock TimeoutError as
'retry later', never as an exhausted credential.
- Live 120-process stampede harness: 41 refreshes/9 unrecovered -> 1
refresh/0 unrecovered.
Two follow-up gaps found by actually running 'copilot login' end-to-end:
1. The CLI (without an OS keychain) stores its token in
~/.copilot/config.json under copilotTokens — a JSONC file with
//-comment header lines. Add it as an auth-evidence source in
_external_process_auth_evidence(), parsed comment-tolerantly and
counting only a non-empty copilotTokens map (config.json exists after
first launch even when logged out).
2. The desktop chat picker requests explicit_only rows, and
_filter_explicit_provider_rows() dropped copilot-acp because a CLI
login leaves no trace in active_provider, model.provider, or env vars
— exactly the Anthropic-OAuth carve-out case. Keep external_process
rows when their CLI credentials are verified (auth_verified), while
still dropping ambient executable-on-PATH-only rows so the filter's
narrower contract holds.
Net effect: after 'copilot login', copilot-acp appears in the desktop
picker and the Accounts card reads signed in; a machine with only the
binary installed keeps today's hidden-until-configured behavior.
get_auth_status() special-cased the literal slug 'copilot-acp'; any other
external_process provider (the pending kiro/devin/junie ACP backends) fell
through to {'logged_in': False}. Dispatch on
PROVIDER_REGISTRY[target].auth_type == 'external_process' instead so the
whole class gets a real status.
get_external_process_provider_status() equated 'logged_in' with 'the
executable resolves', which says nothing about whether the Copilot CLI is
actually signed in. Add auth_verified/auth_source: positive-only evidence
from supported env tokens (validated via copilot_auth, classic ghp_* PATs
excluded) or known on-disk GitHub Copilot credential stores. No evidence
means unknown — never presented as signed out, because the CLI may keep its
session in an OS keychain. Deliberately subprocess-free to avoid re-creating
the gh-auth-token cold-start stall (#60800).
The clone-strip and root-write-through in the previous commit stop NEW forks
but leave installs that forked before upgrading in the broken state: each
profile keeps its own copy of the root grant, whichever profile rotated last
holds the only live refresh token, and root plus every sibling still hit
invalid_grant on their next refresh. The PR body asked those users to
re-auth at root and hand-edit profiles/*/auth.json; this makes it automatic.
`heal_forked_single_use_oauth_grants(provider)` (hermes_cli/auth.py) runs at
the top of a profile's `load_pool()` for SINGLE_USE_REFRESH_POOL_PROVIDERS.
Under the profile lock then the root lock it matches each profile OAuth row
to its root counterpart by lineage — same pool id (preserved by both fork
paths), same JWT account identity, same token material, else same provider +
same client (Anthropic pkce grants carry no claims) — keeps the copy with the
freshest rotation (`expires_at_ms` / `last_refresh` / JWT exp), writes it into
ROOT when root's is older, and strips the profile copy (pool rows, the
`providers.<id>` device-code block for Codex/xAI, and a profile-local
`.anthropic_oauth.json`) so the profile borrows root from then on. Root's
singleton and its hermes_pkce row are kept in step so root's own re-seed
cannot resurrect the spent pair.
Guarantees: idempotent (mtime-keyed clean mark skips the locked scan on the
per-call hot path); one INFO line per healed profile; API-key rows untouched;
a row with no root counterpart (root lost its grant, or an independent
account whose claims differ) is never deleted; only the two auth.json files
the root fallback already reads are touched — no environ/secret-scope reads.
`hermes auth list` / `hermes auth status <provider>` print the heal note.
Live repro (real imports, temp root + forge/atlas each holding a pre-fix
verbatim copy, forge already rotated RT0->RT1 into its own file, fake
single-use token endpoint): before — atlas None, forge AT2 (only in forge),
root None; server log 4x REUSE of spent RT0. After — forge's load heals to
root and rotates there, atlas and root select AT2, profiles/*/auth.json hold
no anthropic rows, server log exactly one ROTATE and zero REUSE.
Anthropic / Codex / xAI OAuth refresh tokens are single-use: a grant copied
into a second auth.json is one credential with two owners, and the first
profile to refresh it revokes the pair for every sibling (invalid_grant /
refresh_token_reused). Two code paths forked grants that way:
1. `hermes profile create --clone-all` and the dashboard/TUI
`mirror_credentials` flow copied auth.json (+ .anthropic_oauth.json)
verbatim. Both now run `strip_cloned_single_use_oauth_grants()`, which
drops OAuth rows for SINGLE_USE_REFRESH_POOL_PROVIDERS, the matching
`providers.<id>` device-code blocks, and the PKCE singleton file; API
keys are still copied. The clone reads the root grant through the
existing credential-pool root fallback.
2. A named profile with no local rows BORROWS the root grant via
`read_credential_pool()`'s fallback, but every persist
(`CredentialPool._persist`, `load_pool` reseed, `remove_index`) wrote the
rows into the profile's own auth.json — materializing a fork on the first
rotation. `persist_pool_entries()` now routes borrowed single-use rows
back to the root store (update-only, under the root lock; never falls
back to a local copy). A borrowed `hermes_pkce` rotation commits its
singleton to the root `.anthropic_oauth.json`, the borrower never prunes
root-seeded rows it cannot see the backing file for, and
`hermes -p <profile> auth add` persists only the profile's own rows.
Live repro (real imports, temp root + profiles, fake single-use token
endpoint): before — first profile rotation RT0->RT1 in profile only; root
and sibling then hit `invalid_grant`, `resolve_anthropic_token()` -> None.
After — rotation lands in root; root and both siblings select AT1, no reuse.
Direction per Teknium: stop cloning OAuth into profiles (ONE grant at root,
children inherit via context) rather than making clones survive. Supersedes
the clone-strip/root-write-through half of #100389 and the init-refresh idea
in #100703 (an expired-but-refreshable row already refreshes on select()).
Closes#100339
Co-authored-by: HexLab98 <liruixinch@outlook.com>
#94388 (salvage of #70007) added RFC 8305 IPv6/IPv4 connection racing for
the direct synchronous chatgpt.com/backend-api/codex chat transport only.
Per the #13834 residual list, the auxiliary Codex paths were still serial:
- hermes_cli/auth.py Codex OAuth clients (token refresh at
auth.openai.com/oauth/token, device-code login, token exchange, usage
probe) each built plain httpx.Client()s — on broken-but-advertised IPv6
every connect eats the full timeout per AAAA before IPv4 is tried, so
auth fails where the official Codex CLI (which races) works.
- The async transport (async_mode=True in build_keepalive_http_client)
had no explicit racing wired.
Changes:
- agent/process_bootstrap.py: add enable_happy_eyeballs_on_client() —
installs the existing _HappyEyeballsSyncBackend on a ready-built sync
httpx.Client's direct transports (default transport + mounts), skipping
proxy-backed pools (HTTPProxy/SOCKSProxy: TCP connect goes to the proxy
host, out of scope). Export it.
- hermes_cli/auth.py: add _codex_http_client() wrapper and use it for the
five Codex OAuth/probe endpoints. Best-effort: falls back to default
serial behavior if the backend can't be installed.
- Async transport: verified httpcore's AnyIOBackend already implements
RFC 8305 natively via anyio.connect_tcp(happy_eyeballs_delay=0.25) —
no custom backend needed. Documented in build_keepalive_http_client and
pinned by tests (contract test on the anyio signature + a live
regression test where a blackholed 100::1 IPv6 addr hangs and local
IPv4 wins in ~250ms instead of the serial connect timeout).
network.force_ipv4 is unaffected: it patches socket.getaddrinfo below
all these layers and keeps working as the interim workaround.
Refs #13834; follows #94388 (9cce8725).
Three kills at the shared chokepoints:
1. resolve_provider() now REFUSES env-key/pool auto-adoption of openrouter
while the active config.yaml is corrupt (AuthError code=corrupt_config).
A broken config falls back to DEFAULT_CONFIG, so tier-2 found no
model.provider and tier-3/4 silently adopted the PAID openrouter provider
against the user's real (unparseable) intent. New probe:
hermes_cli.config.get_active_config_parse_failure(), recorded in the
existing _warn_config_parse_failure() funnel keyed by (mtime_ns, size) —
a fixed file clears the block immediately. Explicit provider requests
are untouched.
2. auxiliary lane built-in OpenRouter fallback model is now a :free SKU
(nvidia/nemotron-3-ultra-550b-a55b:free) instead of the paid
google/gemini-3.6-flash. User-configured auxiliary.openrouter_model is
honored untouched (paid-lane warning retained).
3. env->pool ingestion of OPENROUTER_API_KEY now logs a WARNING (once per
process per provider) when a credential is newly ingested — ingestion
itself stays allowed.
Fixes#81952 (silent-paid-default half; sibling PR covers the
non-interactive fail-closed guard).
Rescuing an empty list after partitioning put paid models back into a
free-tier user's selectable list, and the dashboard could pick one as the
silent default. Narrowing first also drops the separate unavailable-list
filter.
Dashboard PKCE login reused the code_verifier as the OAuth state (leaking
it and disabling CSRF validation) and never checked state on callback --
the same class of bug already fixed for the CLI flow. Credential-pool
refresh excluded "anthropic" from the cross-process lock Codex/xAI already
get, so concurrent Hermes processes racing a single-use refresh token could
leave the loser stuck exhausted with no recovery for hermes_pkce/dashboard
sources. The dashboard OAuth save also never cleared a stale
ANTHROPIC_API_KEY, which resolve_anthropic_token() prioritizes over the
OAuth pool entry by design -- so a leftover key silently kept billing
pay-per-token after a Claude Pro/Max login.
A concurrency stress test written to validate the refresh-race fix under
load surfaced a fifth, unrelated bug: _auth_store_lock()'s Windows
lock-file "ensure content" write was unguarded and could raise an uncaught
PermissionError under real contention -- affecting every single-use-token
provider sharing that lock, not just Anthropic.
Fixes#87887, #87888, #87889.
The fallback also ran on unavailable_models, which is legitimately empty on a
paid tier, filling the picker with the whole reachable set. Make it opt-in.
The gateway omits a policy-blocked model from `/v1/models` rather than
marking it, so after the preceding commits a restricted model is simply
absent from the pickers. That reads as "Hermes does not support this"
instead of "your organization disallows it".
Show one line when the org is governed, in the two flows where a user
picks a model. It enumerates nothing: model policy is an allowlist, so an
org admitting a handful of models blocks the whole rest of the catalog,
and graying hundreds of rows would be a worse UI than omitting them.
Driven by the `policy_present` claim, which is tri-state — the line shows
only when it is explicitly true, because an absent claim means an older
mint rather than an unrestricted org. The claim is stamped at mint time,
so the line can lag a policy change by up to the access token's lifetime.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four surfaces list Nous models, and none of them was filtered. All four
seed from the docs-hosted curated manifest and union the Portal's
`recommended-models` endpoint; neither source is authenticated, so org
policy had no effect on the model a user picks — which is the model they
then use. The Portal endpoint compounds it, serving one globally
CDN-cached payload for the whole platform, invalidated only by admin
pricing edits and never by a policy change, so it can put a hidden model
straight back into a list.
Narrow all four against the authenticated catalog:
- `_login_nous`, which chooses the model the session starts on
- `_model_flow_nous`, the `hermes model` picker
- `list_authenticated_providers`, the `/model` picker
- `/api/model/recommended-default`, dashboard onboarding
The list stays curated and curated-ordered — the policy set only ever
subtracts. Replacing a list with the catalog's keys would swap a curated
agentic list for a large alphabetical dump of vendor-prefixed models,
which is the regression the picker's nous branch already exists to avoid.
The `/model` picker's filter sits outside the try that wraps the Portal
union, so a Portal outage still yields a policy-filtered curated list.
`_login_nous` and `_model_flow_nous` also narrow their unavailable lists,
so a policy-hidden model is not offered as a free-tier upsell either.
For an org with no policy — the common case — the filter is a no-op and
every list is what it was.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
OpenRouter and Nous already list z-ai/glm-5.3-flash (#95621). The
native z.ai picker, OpenCode Go/Zen fallbacks, setup wizard, and
Coding Plan probes did not. Context still resolves through the
existing glm-5.3 1M key.
The docstring and all four caller comments still said the helper
refuses only / and top-level directories. Since #93757 it also refuses
the entire hermes-agent install tree. Bring the docstring and the
comments at the four credential-write call sites in line with the
actual behavior so future changes are not misled by a stale safety
description.
Follow-up to #93757.
A malformed OPENROUTER_API_KEY in ~/.hermes/.env (truncated paste, wrong
provider's key) passed has_usable_secret's length/placeholder check and was
returned by _resolve_api_key_provider_secret before the credential-pool
fallback was ever reached, producing opaque '401 Missing Authentication
header' errors even when a valid pool entry existed (#93593).
- Add KNOWN_PROVIDER_KEY_PREFIXES (openrouter: sk-or-) and skip env values
that mismatch a declared prefix, logging a WARNING naming the env var and
expected prefix, then continuing to the next env var / pool fallback.
- Iterate credential-pool entries (peek first, then entries()) instead of
only peek(), so one malformed pool entry doesn't block a valid one.
- Providers without a declared prefix are fail-open: unknown key formats
are never rejected. Valid env keys still win over the pool (precedence
unchanged).
Fixes#93593
Address review feedback: the gate reused has_vertex_credentials(), which also
returns True for an ambient GOOGLE_APPLICATION_CREDENTIALS path. That var is
commonly set globally for unrelated GCP work, so a user who never configured
Hermes for Vertex would see it in the explicit-only picker and could spend
against those credentials — weakening the explicit-configuration guarantee the
gate documents (mirrors the existing _IMPLICIT_ENV_VARS carve-out).
Add has_explicit_vertex_config() in agent/vertex_adapter.py that checks only
Hermes-scoped signals — VERTEX_PROJECT_ID / vertex.project_id (project
override) or a resolvable VERTEX_CREDENTIALS_PATH — and NOT
GOOGLE_APPLICATION_CREDENTIALS. Route the auth gate through it.
Adds a regression test asserting an ambient GOOGLE_APPLICATION_CREDENTIALS
path alone does not mark Vertex explicit, and updates the existing test to
drive the real config signal instead of mocking has_vertex_credentials().
is_provider_explicitly_configured() only checked provider env vars for
auth_type="api_key" providers. Bedrock is registered with
auth_type="aws_sdk" and an empty api_key_env_vars tuple, so a user who
sets AWS_BEARER_TOKEN_BEDROCK (or an AWS_ACCESS_KEY_ID +
AWS_SECRET_ACCESS_KEY pair) in .env was never counted as having
explicitly configured the provider.
Symptom: the desktop model picker (and any consumer of
build_models_payload(explicit_only=True)) silently hid the Bedrock row
even though list_authenticated_providers had discovered credentials and
built a full model list for it. Reproduced on main:
build_models_payload(ctx, explicit_only=False)
-> ['moa', 'nous', 'bedrock', ...] # row exists, 132 models
build_models_payload(ctx, explicit_only=True)
-> ['nous', ...] # bedrock filtered out
Fix: aws_sdk-type providers now count as explicitly configured when
Bedrock-relevant env credentials are present. Deliberately env-var-only:
ambient sources (AWS_PROFILE / SSO, EC2 IMDS, container credentials)
still do NOT auto-surface, consistent with the gate's purpose (#56974)
and with a lone AWS_ACCESS_KEY_ID (no secret) not counting.
Tests: six behavior-contract cases in test_auth_provider_gate.py
covering bearer token, key pair, lone key id, ambient AWS_PROFILE,
no-credential baseline, and non-leakage into other providers.
Free ($0/$0) Nous Portal models sat with a blank discount column and no
sale star (stealth/ox-alpha, upstage/solar-pro4:free), reading as missing
data next to the -20% sale rows. compute_sale_discount now returns a flat
100% for free models; was_* raws pass through only when the gateway served
a pricing.original, so natively-free models render bare '-100%' with no
fabricated 'was ?/?'. CLI picker star follows on_sale automatically;
inventory feed carries discount_percent=100 to Desktop, whose FREE badge
row now renders the amber -100% pill beside it.
GLM-5.3 is live on api.z.ai (coding plan endpoint) but had no entries in
Hermes, so it silently fell back to the generic 202K GLM context —
triggering premature context compression on a 1M-window model.
- model_metadata: 'glm-5.3': 1_048_576 (same base model as 5.2; 1M
context / 128K max output per docs.z.ai/guides/llm/glm-5.3, verified
2026-08-14)
- auth: add glm-5.3 to coding-plan probe lists (global + CN)
- models: add glm-5.3 to picker/model lists (6 sites)
- zai provider: reasoning_effort mapping covers glm-5.3 (accepted live
by the endpoint, HTTP 200)
A keyless provider has no credential to lack, but every auth-gated
surface treated 'no key' as 'not authenticated', so opencode-free was
invisible in /model, provider:model listing, and the desktop model
pickers unless a user had unrelated OpenCode env vars set.
One policy, three gates, all derived from the HermesOverlay keyless
flag (#91358):
- auth.py get_api_key_provider_status: keyless providers report
configured/logged_in=True with key_source 'keyless' — flows through
get_auth_status to every status consumer (hermes status, dashboards,
list_available_providers).
- model_switch.py list_authenticated_providers: keyless overlay rows
get has_creds=True before any env/pool/auth-store checks — this is
the source for /model, the TUI picker, and the desktop
/api/model/options payload.
- inventory.py explicit-only filter (desktop chat pickers): keyless
providers are kept — there is nothing to 'explicitly configure', and
hiding a zero-setup provider defeats its purpose.
E2E (temp HERMES_HOME, all keys stripped): get_auth_status logged_in,
list_available_providers authenticated, picker row with 6 models,
desktop payload default AND explicit_only both include the provider,
and the full switch pipeline (parse free:x-preview-f-free →
switch_model) resolves to the keyless runtime. 4 new tests.
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).
On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
Zen endpoint routing incl. muse->responses); keyless predicate extended
with unsuffixed free slugs (big-pickle); free runtime pins EVERY
opencode-free model keyless; curated catalog replaces the models.dev
cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
invariant (every curated model must satisfy the keyless predicate)
E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.
The free tier requires a real account API key and throttles third-party
clients by User-Agent:
- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
unconditionally (the stale keyless-tier assumption), and credential-pool
exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
message.
Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
Widens #90327's line_input() to the whole bug class: all 46 bare input()
free-text prompt sites across the setup wizards (model_setup_flows, setup,
config, gateway, auth, auth_commands, plugins_cmd, skills_hub, bundles,
setup_whatsapp_cloud, main) now route through line_input(), and the shared
cli_output.prompt() / setup.prompt() helpers do too — so every CLI wizard
gets cursor editing, not just the custom model prompt.
Redirected stdin and missing prompt_toolkit keep builtin input() behavior.
E2E: real PTY with raw escape bytes through line_input, cli_output.prompt,
and setup.prompt (arrows + Ctrl+A/E edit correctly); redirected-stdin
fallback verified.
The previous except Exception silently fell back to os.getenv if the
_scoped_key_env import ever failed — under multiplex that is exactly the
fail-open this PR removes (secondary profiles would regress to 'No LLM
provider configured' with zero trace). Catch only ImportError, log a
WARNING naming the consequence, and let any other failure propagate.
Also replaces the lambda fallback with a named nested function.
resolve_provider's auto path read provider API keys with bare
os.getenv — under multiplex a secondary profile's keys live only in its
secret scope, so auto-detection found nothing and every secondary
profile with model.provider: auto failed with 'No LLM provider
configured' at agent init (reproduced on a live 7-profile gateway).
Route both env-key reads (the OPENAI/OPENROUTER tier and the
PROVIDER_REGISTRY loop) through _scoped_key_env, the scope-aware helper
auxiliary_client already uses: secret scope wins under multiplex,
UnscopedSecretError falls back to os.environ (default-profile/CLI
paths unchanged). Same bug class as #86905.
Verified in a gateway-accurate simulation (hermes_home_override +
profile scope): resolve_provider('auto') now returns the secondary
profile's own provider (deepseek) instead of erroring.
get_default_hermes_root() resolves HERMES_HOME against the platform
native home (~80us of path resolution) on EVERY call and is called at
31+ sites — every _load_global_auth_store() (per provider row in the
/model picker), kanban, backup, gateway, update. Its result depends
only on (HERMES_HOME, native home), so memoise it keyed on those two
inputs, compared for free on each call (freshness-correct even if a
test or plugin mutates HERMES_HOME mid-process).
_load_global_auth_store() re-read + re-parsed the global auth.json on
every call; read_credential_pool() -> load_pool() runs it once per
provider row in the /model picker even when the profile has entries and
the global fallback never fires. Memoise keyed on the global auth
file's path+mtime (same pattern as _nous_auth_status_cache); the store
only changes when a global-scope auth write touches the file.
Measured (profile mode, 30-provider global store): get_default_hermes_root
81us -> 10us; _load_global_auth_store 128us -> 66us; load_pool 165us ->
137us per call — ~2ms saved per /model picker render (20 provider rows).
Regression tests: hermes_constants memo pin (no path resolution on
repeat calls, HERMES_HOME change forces a fresh resolution); global-store
memo pins (store read once across repeats, mtime bump re-reads once,
absent store stays cheap).
(cherry picked from commit be348f32e5bd7479c26fabb652de549fb9c8a1e1)
On an Anthropic subscription OAuth credential, every request failed with
HTTP 400 "You're out of extra usage. Add more at claude.ai/settings/usage".
That is not a billing condition: Anthropic's server-side content filter rejects
the first sentence of Hermes' own built-in SKILLS_GUIDANCE prompt, and the
rejection is surfaced with a billing-shaped message. Because the message points
at the usage settings page, it reliably sends people to buy quota they do not
need — the reporter lost three debugging sessions to it.
Bisected against the live API with the real 71,721-char assembled prompt: the
first SKILLS_GUIDANCE sentence alone reproduces the 400 and removing it alone
clears it. Size was ruled out (20 KB of unrelated filler returns 200) and so was
the system[0] identity gate (that returns 429, a different failure).
Three changes, all serving the same outcome — a subscription user can no longer
be misdirected by this 400:
- agent/prompt_builder.py: reword the triggering sentence to the phrasing the
reporter verified returns 200. Meaning, the skill_manage reference, and the
## Skill Safety Rule block are all preserved. The reword is empirically
validated rather than understood, so a comment records the bisect and warns
that any rewrite must be re-verified against an OAuth token, not an API key.
- agent/conversation_loop.py: the Anthropic branch of the billing guidance no
longer asserts exhaustion as fact. It hedges the opening line, names the
content-filter alternative, and gives the operator a way to tell the two apart
(if the usage page still shows quota, suspect a content rejection). It also
points at `hermes auth reset anthropic`, because the credential exhaustion
latch replays the stored error for ~60 min without issuing a request — which
makes a real fix look like it did not work.
- hermes_cli/auth.py: document that CLAUDE_CODE_OAUTH_TOKEN is an OAuth token,
not an API key, despite auth_type="api_key". It stays in api_key_env_vars
because that tuple doubles as the credential-discovery list; removing it would
stop Hermes finding a `claude setup-token` credential at all.
Docs updated to match the reworded prompt.
Fixes#82154
Adds hermes_cli/model_selection_guards.py: a single evaluation point that
runs every selection guard (cost + the new data-policy guard) and returns
the warnings that fired. All seven model-selection surfaces (CLI picker,
cli.py TUI modal, gateway typed /model, dashboard web_server, TUI gateway,
Telegram and Discord pickers) now call the registry instead of importing
model_cost_guard directly — so the data-training-tier warning from
PR #81416 fires everywhere at once, and future guards need zero surface
wiring.
Guard modules keep their public APIs; existing mock patch points
(hermes_cli.model_cost_guard.expensive_model_warning) remain valid.
muse-spark-1.2-contributor is heavily discounted BECAUSE Meta uses your
prompts and completions to train future models. Selecting it for the price
without realising the data trade-off is a footgun.
Add hermes_cli/model_data_policy_guard.py (mirrors model_cost_guard):
data_training_warning(model_id, provider, base_url) -> DataTrainingWarning|None,
driven by a vendor-agnostic rule table. The status is not machine-readable on
/v1/models or models.dev, so the v1 rule keys on the documented '-contributor'
model id (fires regardless of provider, so it also covers custom/gateway
routes). Message mirrors Meta's pricing-doc language and figures
(https://dev.meta.ai/docs/pricing-rate-limits/).
Wire it into the CLI model picker's confirm flow (auth.py) as a [y/N]
disclosure, chained after the expensive-model cost guard. Fires only on the
contributor tier; silent on muse-spark-1.1/1.2 and all other models.