Review findings on #92797 (@100yenadmin):
- is_codex_900k_base() is now the single source of truth used by picker
synthesis, context resolution, /model validation, and wire stripping.
Eligibility is an exact table (sol/terra/luna, gpt-5.4, daybreak alias)
plus date-shaped 5.6 snapshots — family-prefix matching removed, so
non-routable -pro slugs and unknown descendants never gain variants.
- strip_codex_context_variant_suffix() strips conditionally: ineligible
aliases (gpt-5.5-900k) are returned unchanged and fail honestly at the
API instead of silently running as the base model at 272K.
- validate_requested_model() rejects ineligible *-900k aliases before the
hidden-slug soft-accept, and accepts valid variants missing from a
stale catalog without letting the typo auto-corrector eat the suffix.
- Codex context resolver drops vendor/ namespaces, so
openai/gpt-5.6-sol-900k resolves to 900K like the bare id.
- Table-driven regression covering eligible bases/snapshots/namespaced
ids and rejected -pro/-mini/5.5/unknown aliases, asserting context AND
wire model.
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the
live-verified 900K burned through subscription usage for users who never
asked for the larger window (bigger window = more input tokens per request).
- Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the
advertised 272K again — the cheaper limit is the default.
- The model picker synthesizes explicit <slug>-900k variants (e.g.
gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into
the 900K window. Slugs that genuinely enforce 272K (gpt-5.5,
gpt-5.4-mini) get no variant.
- The -900k suffix is Hermes-side only: stripped before the model id hits
the wire (main transport + auxiliary Responses adapter), and pricing
aliases the variants onto the base entries.
- Docs: new opt-in section in context-compression-and-caching.md.
Follow-up structural pass on the review fix:
- Runtime provider, auxiliary resolution, model validation
(hermes_cli/models.py), live discovery (bedrock_model_ids_or_none),
and the Mantle URL/SigV4 fallbacks all resolve their region through
resolve_bedrock_runtime_region() — one canonical implementation of the
config-first priority instead of three hand-rolled copies.
- agent_init: drop the 'if "client_kwargs" in locals()' guard by
initializing client_kwargs unconditionally at the top of the else
branch; the Mantle kwargs hook is a documented no-op for non-Mantle
base URLs.
GPT-5.6 Sol, Terra, and Luna went GA on Amazon Bedrock on 2026-07-13.
Like GPT-5.5, they are served exclusively from the Bedrock Mantle
OpenAI-compatible Responses endpoint (the model cards list
bedrock-runtime/Converse as unsupported), so they ride the allowlist
routing introduced for GPT-5.5:
- Add openai.gpt-5.6-{sol,terra,luna} to BEDROCK_OPENAI_RESPONSES_MODEL_IDS
so runtime resolution, auxiliary calls, and MoA slots all take the
SigV4/bearer Mantle Responses path.
- Surface the family in the curated Bedrock picker list.
- Record the 272K context window from the AWS model cards for all four
Mantle OpenAI models (previously fell back to the 128K default).
- Generalize picker tests from the hardcoded single-model checks to the
BEDROCK_OPENAI_RESPONSES_MODEL_IDS allowlist so future Mantle model
additions do not require test surgery; add routing, picker, and
context-length coverage for the 5.6 family.
Docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-cards-openai.html
Route Bedrock-hosted OpenAI GPT-5.5 through the Bedrock Mantle OpenAI Responses endpoint with SigV4 request signing. Keep native Bedrock Converse and Claude Bedrock routing unchanged, and add picker/runtime regression coverage.
Free ($0/$0) Nous Portal models sat with a blank discount column and no
sale star (stealth/ox-alpha, upstage/solar-pro4:free), reading as missing
data next to the -20% sale rows. compute_sale_discount now returns a flat
100% for free models; was_* raws pass through only when the gateway served
a pricing.original, so natively-free models render bare '-100%' with no
fabricated 'was ?/?'. CLI picker star follows on_sale automatically;
inventory feed carries discount_percent=100 to Desktop, whose FREE badge
row now renders the amber -100% pill beside it.
Follow-up to the salvaged GLM-5.3 support commit: drop z-ai/glm-5.1 from
both curated lists per Teknium's direction (glm-5.2 keeps the 'default'
tag), and regenerate the docs manifest. glm-5.1 remains available via
live discovery and on out-of-scope surfaces (zai plugin, setup defaults,
opencode-go) — named leftovers, not silently swept.
GLM-5.3 is live on api.z.ai (coding plan endpoint) but had no entries in
Hermes, so it silently fell back to the generic 202K GLM context —
triggering premature context compression on a 1M-window model.
- model_metadata: 'glm-5.3': 1_048_576 (same base model as 5.2; 1M
context / 128K max output per docs.z.ai/guides/llm/glm-5.3, verified
2026-08-14)
- auth: add glm-5.3 to coding-plan probe lists (global + CN)
- models: add glm-5.3 to picker/model lists (6 sites)
- zai provider: reasoning_effort mapping covers glm-5.3 (accepted live
by the endpoint, HTTP 200)
Third surface for the Ox Alpha stealth reasoning model (after the
OpenCode Zen rollout in #91250 and the OpenRouter listing in #91284).
Adds stealth/ox-alpha to the curated Nous list and regenerates the docs
manifest. Free on the portal ($0/$0), 1M context, 131K max output —
verified against the live inference-api.nousresearch.com/v1/models.
Provider-agnostic metadata already resolves via the bare ox-alpha slug
(DEFAULT_CONTEXT_LENGTHS 1,048,576; reasoning_timeouts 300s floor), and
the nous route bills via official_models_api, so no pricing snapshot is
needed.
First actioned report from the overhauled model-catalog-scout cron
(2026-08-21 validation run), every item re-verified live before edit:
Delisted (gone from live catalogs):
- opencode-zen curated: claude-opus-4-1, qwen3.7-max, qwen3.7-plus
(absent from live zen /v1/models; qwen3.7 family remains on Go)
- OPENROUTER_MODELS free section: poolside/laguna-m.1:free (rotated to
s-2.1/xs-2.1), tencent/hy3:free, inclusionai/ring-2.6-1t:free
Added (present + verified in live catalogs):
- OpenRouter free: z-ai/glm-5.2:free (256K), poolside/laguna-s-2.1:free
+ laguna-xs-2.1:free (262K), nvidia/nemotron-3.5-lightning:free (1M)
- opencode-go curated: ox-alpha-free (Go-subscription twin of the Zen
keyless Ox Alpha; keyed — Go relay 401s anonymous requests)
Metadata:
- DEFAULT_CONTEXT_LENGTHS: laguna-s-2.1/xs-2.1 262144;
nemotron-3.5-lightning 1M (overrides the generic 131K nemotron entry);
glm-5.2:free 256K (the free variant is capped below the 1M paid entry)
Keyless-heal hardening (the real find):
- opencode_zen_free_runtime now gates the zen/go→keyless heal on
MEMBERSHIP in the verified opencode-free catalog, not the -free
suffix — ox-alpha-free is a KEYED Go model despite its suffix, and
suffix-based healing would have routed it to a Zen relay that
doesn't serve it (verified: zen 401s 'not supported', go 401s
'Missing API key'). New regression test pins this.
Fixture sweep: tencent/hy3:free catalog assertion updated (delisted
slug); nous-route fixtures using hy3:free as incidental model names
left alone (self-consistent mocks). model-catalog.json regenerated.
Live verification (2026-08-21): big-pickle and mimo-v2.5-free return 429
FreeUsageLimitError for ANY User-Agent except the opencode CLI's own
'opencode/latest' — same IP, no cooldown effect, while the other six free
models serve our honest HermesAgent UA freely. Hermes sends deliberate
attribution headers and does not impersonate other clients, so these two
models are broken for our users by policy on OpenCode's side; delist them
rather than ship dead picker entries.
- opencode-free catalog: 8 -> 6 models (both curated lists)
- plugin default_aux_model: big-pickle -> laguna-s-2.1-free (fastest
non-gated free model)
- keyless predicate keeps big-pickle (it IS free-tier; correct routing if
a user enters it manually — they get the relay's own 429, not our 401)
E2E: picker shows 6, laguna aux default completes a keyless agent turn.
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).
On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
Zen endpoint routing incl. muse->responses); keyless predicate extended
with unsuffixed free slugs (big-pickle); free runtime pins EVERY
opencode-free model keyless; curated catalog replaces the models.dev
cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
invariant (every curated model must satisfy the keyless predicate)
E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.
The free tier requires a real account API key and throttles third-party
clients by User-Agent:
- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
unconditionally (the stale keyless-tier assumption), and credential-pool
exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
message.
Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
Adds OpenRouter's free "Ox Alpha" stealth reasoning model
(stealth/ox-alpha) to the OpenRouter fallback snapshot, plus the
provider-agnostic metadata it needs:
- OPENROUTER_MODELS: free-tier entry (1M ctx)
- DEFAULT_CONTEXT_LENGTHS: ox-alpha -> 1,048,576 (verified against
OpenRouter live /api/v1/models; without this the slug fell through
to no match)
- reasoning_timeouts.py: 300s stale floor for ox-alpha and the
OpenCode Zen twin slug x-preview-f-free (reasoning model,
long-horizon agentic work per its model card)
- model-catalog.json regenerated
Pricing snapshot skipped: openrouter bills via official_models_api
(live pricing; model is free anyway).
The Zen relay serves *-free models (x-preview-f-free / Ox Alpha) ONLY
anonymously: any Authorization bearer it doesn't recognize is a 401
'Invalid API key' — including our no-key-required placeholder and valid
OpenCode GO subscription keys. The Go relay doesn't serve the free tier
at all ('Model x is not supported'). So the free model failed for every
Hermes user: keyless setups got the placeholder bearer, and OpenCode
subscribers sent a Go key to a relay that rejects it.
Fix (class-wide for all 8 current *-free Zen slugs, not just Ox Alpha):
- hermes_cli/models.py: is_opencode_zen_free_model / opencode_zen_free_runtime
/ opencode_zen_free_headers — one shared policy: free slugs pin to the
Zen relay with a keyless placeholder and an empty Authorization header
that overrides the OpenAI SDK's 'Bearer <key>'.
- runtime_provider.py: free slugs route through the keyless runtime before
the credential-pool/explicit/api_key paths (no key required; Go
selections heal to Zen). Paid models still fail closed without a key.
- agent_init.py + auxiliary_client.py: the placeholder key swaps in the
empty-Authorization headers at both client-build chokepoints.
Verified live (2026-08-21): anonymous chat/completions 200 incl. tools,
streaming, parallel; bad bearer 401; full E2E AIAgent turn with a real
terminal tool round-trip completes keyless under both opencode-zen and
opencode-go providers. Sabotage run: routing tests fail without the fix.
Builds on @Lesnak1's #85619 (issue #85589):
- New opencode_provider_family() single-owner predicate in
hermes_cli/models.py — resolves built-in AND custom family providers
(opencode-go-bridge, OpenCode-Zen-Custom, ...) case-insensitively.
Migrated all 8 inlined family checks (models.py x3, runtime_provider.py
x4 from the salvaged commits) plus 4 sibling sites the PR missed:
cli.py api_mode sync, agent_runtime_helpers.py double-/v1 guard,
model_normalize.py flat-namespace strip, model_switch.py base_url
normalization.
- Responses transport: alias OpenCode-reserved function names
(web_search, search_files -> hermes_*) on the wire and map them back on
dispatch — same pattern as the xAI web_search collision fix. Matches
family providers and any base_url on opencode.ai. Fixes the HTTP 400
'custom function name X is reserved' half of #85589.
- Tests: custom-provider routing assertions + 5 new transport alias tests.
- Add x-preview-f-free (Ox Alpha: free, 1M context, ZDR) plus all newly
listed Zen models (gpt-5.6 sol/terra/luna, claude-opus-5, gemini-3.7/3.6
flash + lite, grok-4.6/4.5, muse-spark-1.2, kimi-k3, qwen3.7-max,
hy3-free, laguna-s-2.1-free, nemotron-3.5-lightning-free,
muse-spark-1.2-contributor-free) and Go models (gpt-5.6-luna, grok-4.5,
glm-5.3, qwen3.8-max, hy3, hy3-preview, muse-spark-1.2-contributor).
- Drop delisted north-mini-code-free from Zen.
- Route grok-* on Zen and Go through /v1/responses per the published
endpoint tables (grok-4.6/4.5/build-0.1 on Zen, grok-4.5 on Go).
- 1M context fallback for x-preview-f (Ox Alpha).
- Refresh hermes setup provider samples for both providers.
Catalogs verified against live GET /zen/v1/models and /zen/go/v1/models
plus https://opencode.ai/docs/zen/ and /docs/go/ endpoint tables (2026-08-20).
OpenCode Go and Zen serve muse-spark* only on /v1/responses.
Hermes was sending /chat/completions, which returns HTTP 503
with an empty assistant message. Match the published endpoint
table and the existing gpt-* routing.
- Route muse-spark* to codex_responses on opencode-go and opencode-zen
- Add regression assertions next to the gpt-5.6-luna cases
Portal reasoning capabilities were held only in memory, so a process that
had not yet fetched them answered "unknown" — and on that answer the Nous
profile drops the disable rather than risk a 400. A short-lived process
(`hermes -p`, a cron job, a freshly booted gateway) is always in that
state, so every one of those runs silently ignored "thinking off" and
billed the user for reasoning they had turned off.
The parsed catalog is now mirrored to `cache/reasoning_caps.json`, keyed
by the URL it came from, and hydrated on a cold lookup without touching
the network. Every picker and pricing fetch already pulls that same
document, so they seed the mirror for free.
The catalog URL itself now resolves through the same ladder as the rest
of the Nous catalog reads (`NOUS_INFERENCE_BASE_URL` → credential base →
production) instead of being pinned to production, which had a staging
profile deciding the reasoning-mandatory question from prod's answers.
Keying the mirror by URL keeps those deployments apart.
The Portal serves OpenRouter's catalog schema, so the existing parser and
cache-only tri-state contract carry over unchanged. Only the HTTP fetch is
generalized across the two catalogs; each keeps its own cache because they
list different models.
The Portal 403s a catalog read with no User-Agent, so the shared fetch now
sends one.
The #89503/#70058/#74295/#87279 bug class kept regenerating because every
transport and provider profile hand-rolled its own effort translation map
(9 sites, 4 distinct policies). New agent/reasoning_effort.py is the single
source of truth:
- EFFORT_LADDER: canonical low->high ordering (superset check against
VALID_REASONING_EFFORTS pinned by test)
- clamp_effort(): one policy — supported passes verbatim, otherwise nearest
WEAKER supported level (never escalate, never invert the ladder), floor
when nothing weaker, 'none' never a degradation target, declared
vendor-documented overrides win, bespoke names pass through
- declared wire vocabularies as data: OpenAI-compat, Codex Responses,
xAI (4.6/legacy), Actual relays, Kimi K3/K2, TokenHub, GLM-5.2,
DeepSeek V4, Ollama Cloud, Meta, Solar
Converted sites (all behavior-preserving except noted):
- chat_completions chokepoint, Kimi + TokenHub paths
- codex transport (backend branches now pick a declared set)
- auxiliary_client Responses path
- hermes_cli.models clamp_reasoning_effort_to_supported -> thin wrapper
- plugins: kimi-coding, zai, opencode-zen, deepseek, ollama-cloud,
meta-ai, upstage, custom (copilot already routes via the wrapper)
Behavior fixes the shared policy surfaces:
- ollama-cloud/opencode-go 'minimal' now degrades to 'low' instead of
being dropped (drop left the server default = MORE thinking than asked)
New tests: ladder contract (every configurable level is clamped by every
declared wire set; monotonicity across the full ladder for every set).
PR #67934 marked auto-discovered catalogs by writing two sentinel keys
INSIDE the user-facing ``models`` mapping of custom provider entries:
``__discovered_model_catalog__`` (written by
_save_discovered_models_to_config) and ``__explicit_model_allowlist__``
(injected by _normalize_custom_provider_entry). Every consumer of that
mapping — pickers, selectors, gateway/agent readers, and the user's own
config.yaml — had to know to filter those keys, and any site that
didn't listed them as phantom model IDs (``__discovered_model_catalog__``
showing up as a selectable "model"). The v11→v12 config migration and
the ACP session-state test caught exactly that leak on main.
Replace the in-mapping sentinels with a single entry-level flag:
- ``models_discovered: true`` now sits next to ``models``/``base_url``
on the provider entry; the models mapping stays a clean
``{model_id: metadata}`` dict with no reserved keys.
- _save_discovered_models_to_config writes the new shape and refreshes
catalogs it previously discovered (entry-level flag or legacy
sentinel) instead of treating them as user-curated metadata.
- _normalize_custom_provider_entry no longer injects
``__explicit_model_allowlist__``; a dict-shaped models mapping counts
as an explicit allowlist exactly when the entry is NOT marked
models_discovered.
- _models_config_is_allowlist takes the discovered flag as a parameter
(new helper _entry_models_discovered resolves it, including the
legacy in-mapping sentinel); all call sites updated
(model_switch.py, model_setup_flows.py, acp_adapter/server.py).
- Backward compat, no config version bump: configs written by a
pre-fix Hermes (sentinels inside models) still read correctly —
``__discovered_model_catalog__: true`` is treated as
models_discovered, both sentinel keys are stripped from model
listings, and the next discovery save migrates the entry to the
clean shape. Covered by a new regression test.
Also restore ``except Exception:`` on the pre-existing guards this PR
had narrowed to specific exception tuples (the resolve_runtime_provider
fallback in switch_model, the picker discovery/cache guards in
list_authenticated_providers, _get_model_config_dict, and
_credential_fingerprint). Those guards were intentionally broad on
main — a failed resolution or probe must degrade to the fallback path,
never crash the model switch. Guards the PR introduced for its own new
probe code keep their authored tuples.
The ACP new_session payload also goes back to
probe_current_custom_provider=False, matching the contract main's
test_new_session_returns_authenticated_cross_provider_model_state pins
(session opens must not block on live-probing the current custom
endpoint).
Surface meta/muse-spark-1.2 in the Hermes model selector (CLI, desktop,
gateway) via the curated OpenRouter list and regenerated model-catalog.json.
The model is live on OpenRouter with tool calling; the picker intentionally
does not show the full OpenRouter catalogue.
When the 1h provider_models_cache.json TTL lapses, the model picker
serially fetches /v1/models for each authenticated provider. With 10+
providers this stacks to 15-30s of blocking before the picker renders.
Add a parallel prefetch step before the serial picker build loops:
- _collect_authed_provider_slugs(): lightweight credential pre-scan
that mirrors sections 1/2/2b without fetching model lists
- _prefetch_provider_models_parallel(): ThreadPoolExecutor-based
concurrent fetch of stale/missing cache entries (max 8 workers)
- update_provider_cache_entry(): thread-safe single-entry cache writer
with threading.Lock to prevent concurrent write races
Guardrails:
- Skipped when <=3 authed providers (overhead not worth it)
- Skipped when refresh=True (serial path force-refreshes)
- Exception-isolated (falls back to serial path on any failure)
- No behavioral change (same model lists, same picker output)
Closes#80413
(cherry picked from commit 89dddd6cb5d53d73278e0518c375fb5b878e5c6b)
Port of the bug class from earendil-works/pi#7933 (DeepSeek base-URL
detection matched by raw substring, missing case variants and matching
lookalike URLs). Hermes had the same class at five sites:
- cli_agent_setup_mixin.py: keyless-custom-endpoint detection treated any
URL containing the OpenRouter host substring (path segment, lookalike
domain) as OpenRouter, and missed case variants of the real host.
- models.py validate_requested_model: same substring check for routing an
openrouter provider with a custom base_url to the custom catalog.
- runtime_provider.py: local-endpoint autodetect matched the string
localhost anywhere in the URL, including remote hostnames containing it.
- gateway/run.py: /status endpoint display, same local-host substring.
- agent_runtime_helpers.py: Nous Portal cache-layout detection matched
the nousresearch substring anywhere in the URL.
All sites now use the existing base_url_host_matches / base_url_hostname
helpers (exact host or subdomain, case-insensitive). Regression tests
proven to fail against the old predicates.
Per-provider ssl_ca_cert / ssl_verify reached the httpx chat client and the
auxiliary clients (#56681), but the endpoint discovery and pricing probes did
not. Both probe families resolved TLS from process-wide env vars only:
- the requests-based metadata/pricing probe
(agent/model_metadata.py::_resolve_requests_verify)
- the urllib-based /models catalog probe
(hermes_cli/models.py::probe_api_models)
A custom endpoint whose chain verifies against the provider's configured
bundle, but not the process SSL_CERT_FILE, then logged a spurious
CERTIFICATE_VERIFY_FAILED on every probe even though the chat path worked.
Pointing a global CA env var at the bundle fixes it but changes verification
for every provider, defeating the point of a per-provider setting.
This threads the selected provider's TLS settings into both probe paths,
reusing get_custom_provider_tls_settings so there is no second precedence
chain:
- _resolve_requests_verify(base_url) looks up the provider's ssl_verify /
ssl_ca_cert before falling back to the env vars. Callers with no base_url
keep the exact env-only behavior.
- probe_api_models builds an ssl.SSLContext from the provider settings and
passes it through open_credentialed_url, which gains an ssl_context seam on
the cloned secure opener. Unmatched or public endpoints pass None and keep
urllib's default policy.
Tests: tests/agent/test_custom_provider_ca_probes.py covers both probe
families (provider CA, ssl_verify:false, unmatched, missing file, config
lookup failure) plus end-to-end assertions that the resolved verify value and
SSLContext actually reach the request seam. Verified against the neighboring
metadata, pricing, TLS, and urllib-security suites (266 tests) with no
regressions.
Stop freezing the xAI/xAI-OAuth catalog at import so /model and setup
pick up new Grok IDs after the models.dev cache refreshes. Put xai and
xai-oauth on the shared picker-time models.dev merge path and pin
grok-4.6 as the default headline model.
Swaps the Google flash entry in the OpenRouter and Nous Portal curated
lists to the newly released gemini-3.7-flash (half the price of
3.6-flash: $0.375/M in, $1.875/M out per OpenRouter live metadata;
served on both endpoints, verified live). Also updates the OpenRouter
plugin fallback_models mirror and regenerates model-catalog.json.
Scoped to the two named providers: vertex/gemini/gmi curated lists and
aux defaults still carry 3.6-flash.
Renames the openai-codex provider's display label across the CLI
(hermes model picker, provider labels), the dashboard OAuth accounts
catalog, and the Desktop onboarding + settings provider pickers.
Slug, aliases, and auth flows are unchanged.
OpenRouter's /v1/models entries advertise reasoning capability
(supported_parameters + reasoning.mandatory/supported_efforts). Use that
metadata as the primary gate in _supports_reasoning_extra_body instead of
the hand-maintained vendor-prefix allowlist, which went stale one vendor at
a time (nvidia/ missing -> #75386). Also clamp the requested effort to the
nearest LOWER catalog-supported level in the OpenRouter profile so ultra/max
against a high-capped route no longer 4xxes.
Cache-only on the hot path: capabilities parse for free out of the existing
fetch_openrouter_models() payload, a background warmer covers cold starts,
and unknown models/offline catalogs fall back to the static prefix list
unchanged.
The fast-model picker reads /v1/models to find the small model a provider
currently serves, and it asked anonymously. Most of those endpoints need a
key, so the fetch 401'd and the empty result read as "this provider has no
small model" — the picker fell back to its curated list and never noticed.
Worse, a failed fetch cached its empty result forever, so one bad moment
during startup disabled live model discovery for the life of the process,
and the processes that read this run for weeks. Give the failure an expiry
and pass the provider's credentials.
The bare family rungs (-mini, -flash, haiku) also picked whichever id
sorted first, which is the oldest generation a provider still serves:
gpt-3.5-mini over gpt-5.4-mini, claude-3-haiku over claude-haiku-4.5.
Compare the digit runs as numbers so the rung meant to keep us current
does.
* fix(model-picker): serve cached custom-provider catalog on no-probe opens
#58183 stopped GUI picker opens from live-probing saved custom
OpenAI-compatible endpoints so a stopped local server could not stall the
picker. It gated the whole discovery block, not just the network call, so
`cached_fetch_api_models()` was skipped too — and with it the catalog an
earlier probe had already written to `provider_models_cache.json`.
A custom endpoint that is not the current provider therefore renders only
the models named in its config entry. A local server with 8 models loaded
shows the 1 model that was saved when the provider was first added, on
every picker open, while an explicit Refresh shows all 8.
Add `cache_only` to `cached_fetch_api_models()`: answer from disk within
the existing stale-serve window, never fetch, never revalidate off-thread,
return None on a miss. Split the three call sites in
`list_authenticated_providers()` into what the user's config permits
(`discover_models`, an explicit `models:` allowlist) and how we may obtain
it, so suppressing the probe now downgrades to a cached read instead of
skipping discovery outright. `discover_models: false` still pins, and a
cache hit no longer writes back to config since the probe that populated
it already did.
The latency win stands: a cold cache is a miss, so picker opens against
offline endpoints still make zero network calls.
* test(model-picker): pin the cached-catalog contract for no-probe opens
Cover both halves of the invariant, since fixing either one alone
reintroduces a bug the other guards against.
`cache_only` on `cached_fetch_api_models()`: a fresh entry and an entry
past its TTL but inside the stale-serve window both serve; an entry beyond
that window, an empty cache, rotated credentials, `force_refresh`, and a
missing base_url are all misses — and none of them fetch or spawn a
background revalidation.
`list_authenticated_providers()` on the GUI path: a non-current endpoint
with a warm cache reports its full catalog across all three provider
shapes (`custom_providers`, `providers:`, bare `provider: custom`) with no
live fetch attempted. A cold cache keeps the configured list and still
makes no network call, which is the #58183 guarantee. `discover_models:
false` keeps pinning, and a cache hit does not write back to config.
* fix: persist discovered custom-provider models in the hermes model flow
The `hermes model` named-custom-provider flow (_model_flow_named_custom)
probes the endpoint and shows the full catalog, but never persists it to the
entry's `models:` list. No-probe surfaces (dashboard, desktop, ACP) call
build_models_payload(..., probe_custom_providers=False) and only render the
configured `models:` list, so a provider added via `hermes model` collapses
to the single `model:` default everywhere except the CLI. OpenAI-compatible
providers added via a probing picker already benefit from
_save_discovered_models_to_config; the CLI flow did not.
Persist the live catalog after a successful probe, mirroring the picker path
in model_switch.py. A failed save is non-fatal.
* fix(model-picker): stop an auto-saved catalog pinning a keyless endpoint
The cached-catalog read added for no-probe picker opens still sat behind
the no-key discovery gate, so it never reached the shape that motivated
it: a keyless local model server.
`bool(api_key) or not has_explicit_models` is a network-cost gate. It
exists so Hermes does not probe an endpoint it cannot authenticate to
when that endpoint already declares its catalog (5f00f36ba, 1039e90b5).
Reading a catalog an earlier probe already paid for costs nothing, so
the gate belongs on the probe, not on discovery as a whole.
Left on the discovery side it re-pins the endpoint it was meant to
spare. A successful probe calls `_save_discovered_models_to_config()`,
which writes a plain list into `models:` — exactly the shape
`_models_config_is_allowlist()` reads back as an explicit user
allowlist. A keyless server therefore froze on the catalog of its first
probe and could never widen again, which is the "lineup changes after
config was written" case. f66319097 already carved the dict shape out of
this trap for the same reason; the list shape is the other door into it.
Move the clause to `_probe_live` at both custom-endpoint sites. Probe
suppression is unchanged — verified byte-identical to main across the
keyed/keyless x declared/undeclared matrix — and `discover_models: false`
remains the documented way to pin a catalog.
* test(model-picker): cover the keyless auto-save pinning trap
Three tests around the gate move, each failing on the code before it:
- a keyless endpoint carrying an auto-saved `models:` list still reads
its full cached catalog
- the same row, cold cache and probing enabled, still makes zero live
fetches — the network-cost gate the clause exists for
- an end-to-end round trip: persist a probe result via
`_save_discovered_models_to_config()`, reload it, and assert the shape
we wrote does not read back as a user pin
The round-trip test guards the whole chain rather than one branch, so a
future change that makes the saved shape look like an intentional
allowlist fails here even if the gate logic is refactored.
* fix(model-picker): key the custom-endpoint model cache by api_mode
`cached_fetch_api_models()` fingerprints entries with `api_mode`, but no
call site in `list_authenticated_providers()` passed it, so every custom
row resolved to the `api_mode=None` fingerprint. Two rows sharing a
base_url and credential but differing by `api_mode` are deliberately
distinct picker rows — it is part of `group_key` at both sites — yet they
collapsed onto one cache entry.
That was latent while probing was the only way to fill a row: a mismatched
entry was overwritten by the row's own live fetch. Serving that entry
without a probe makes it visible, so an `anthropic_messages` row could
render the catalog an OpenAI-mode row cached against the same URL. The
wire protocols differ (`x-api-key` + `anthropic-version` vs
`Authorization: Bearer`), so those catalogs are not interchangeable.
Persist `api_mode` on the group at both grouping sites — it is already
part of `group_key`, so it is constant across the group — and pass it
into the cache read. Section 3b (bare `provider: custom`) has no
`api_mode` in scope and already reads with the empty-credential
fingerprint, so it is unchanged.
Reported by Copilot review on #81973.
---------
Co-authored-by: xxxigm <tuancanhnguyen706@gmail.com>
Co-authored-by: Navlem <114683850+Navlem@users.noreply.github.com>
OpenCode Go serves GPT 5.6 Luna only via the Responses API per its
published endpoint table (https://opencode.ai/docs/go/#endpoints), but
opencode_model_api_mode() had no gpt- case in the Go branch, sending
Luna to /v1/chat/completions. The relay's shim streams full text but
never emits a finish_reason chunk, so every complete answer is
classified as a mid-stream drop and each turn fails with 'Response
remained truncated after 4 continuation attempts'.
Mirror the Zen branch: gpt- on Go -> codex_responses. Base URL needs
no change (normalize_opencode_base_url already keeps /v1 for
codex_responses). Extend test_opencode_go_api_modes_match_docs with
the Luna assertions.
Surfaced during the post-merge review pass on our own #81113 follow-up:
cached_fetch_api_models gained _cache_entry_valid (numeric-'at'
validation) but its sibling cached_provider_model_ids still did
float(entry.get('at', 0)), which raises ValueError/TypeError on a
hand-edited or corrupted provider_models_cache.json row and propagates
uncaught into the /model picker call sites. Same fix, same helper:
corrupt rows are now a cache miss (live fetch), never an exception.
Both wrappers now share the identical validity predicate, closing the
divergence the 'mirrors' docstring promised away.
Also two test nits from the same review: unused OrderedDict import
dropped and the drain-order assertion strengthened to pin LRU-first
FIFO order in tests/gateway/test_agent_cache_pressure.py.
Mutation-checked: restoring the raising float() form makes the new
corrupt-at tests fail.
- Give cached_fetch_api_models the same stale-while-revalidate tier as
cached_provider_model_ids: TTL-expired entries within the 7d window are
served instantly while a background refresh rewrites the cache —
without this, every /model open an hour into the session re-blocked on
the live probe (#72762's stall class, deferred).
- Generalize _spawn_swr_refresh(cache_key, refresh_fn) so non-slug
custom:<base_url> keys reuse the same inflight-dedupe scaffolding;
slug behavior unchanged (default refresh_fn preserved).
- Convert the missed sibling site: acp_adapter/server.py
_named_custom_provider_catalogs() live-probed every custom_providers
row's /v1/models per ACP catalog build.
- Extract _cache_entry_valid() (the fp/models predicate existed 4x) and
validate 'at' is numeric so hand-edited/corrupt cache JSON degrades to
a live fetch instead of raising through the picker's blanket except.
- Flatten the dead api_mode conditional (fetch_api_models declares
api_mode=None; branch was behaviorally inert).
- Tests: 4 new guards (stale-serve, stale-window cutoff, generalized SWR
write-through, corrupt-at degradation) — stale-serve and corrupt-at
mutation-checked; 2 existing tests updated for the new behavior.
Custom OpenAI-compatible endpoints (named custom_providers rows, bare
provider: custom, and per-endpoint-map entries) called fetch_api_models()
directly at three call sites in model_switch.py, with no disk cache — unlike
first-class providers, which go through cached_provider_model_ids(). Every
plain /model open live-probed the active custom endpoint's /v1/models,
regardless of how recently it had already been probed.
Adds cached_fetch_api_models() in hermes_cli/models.py: a TTL disk-cache
wrapper keyed on custom:<base_url> (custom endpoints have no
PROVIDER_REGISTRY slug to key on) and fingerprinted on api_key/api_mode/
headers, with the same stale-beats-nothing fallback policy as
cached_provider_model_ids(). Routes all three probe call sites through it.
Since prewarm_picker_cache_async() already calls list_authenticated_providers()
with probe_custom_providers defaulting True, this also fixes the endpoint
being warmed on boot (populating the disk cache) instead of that work being
discarded on every open — any custom endpoint (an LLM gateway, a
self-hosted vLLM/SGLang server, etc.), not just one specific provider.
Fixes#72762. Salvaged from #72810 per review feedback: extracts just the
verified custom-endpoint cache fix with real cache-contract test coverage
(hit/stale/rotation/refresh/fallback), leaving the credential-pool and
Copilot-token-exchange costs described in the issue for separate follow-up.