Internal moves get no re-export aliases: models_validate reads
hermes_constants.openrouter_variant_base directly and the private
_OPENROUTER_VARIANT_SUFFIXES/_openrouter_variant_base shims are dropped.
`:nitro`, `:floor`, `:exacto`, and `:online` are request-time routing
modifiers, not catalog models — OpenRouter's /models lists only the base
id, and a variant runs the same model with the same context window.
`get_model_context_length()` keyed every lookup on the full suffixed id,
so each one missed and the resolver fell through to a generic family
default or the 256K fallback:
openai/gpt-5.5:nitro -> 256K (real 1.05M)
x-ai/grok-4.6:nitro -> 131K (generic "grok" catch-all)
anthropic/claude-opus-4.6:nitro-> 200K (generic "claude" catch-all)
The window silently shrank, triggering early compression and a wrong
/usage readout. f14059fa fixed the sibling half of this bug class in
/model validation; this fixes the metadata half.
Strip a recognized variant suffix for LOOKUP only, keeping the suffixed
id on the wire so the routing opt-in survives. Applied after the explicit
config overrides (steps 0b/0c) so a user-pinned value still wins, and
before every cache/catalog lookup. Gated on the request actually routing
through OpenRouter, so a local Ollama `model:tag` is untouched.
`:free`/`:batch`/`:thinking` are deliberately excluded — those ARE
distinct catalog SKUs with their own windows, so stripping them would
report the wrong number.
The suffix set and base-id split move to hermes_constants (import-safe,
dependency-free) so the metadata layer shares one definition with
hermes_cli.models instead of duplicating it.
Add deepseek/deepseek-v4.1-flash to OPENROUTER_MODELS (Nous list derives from it),
regenerate the docs manifest, and give the slug its own 1M context entry and 600s
reasoning-stale floor — the longest-key-first scan otherwise lands the new slug on
the 128K `deepseek` catch-all and no floor. Live probed on both routes: echoed
model matches, usage.cost billed.
DeepSeek retired deepseek-v4-flash on 2026-09-10 (V4.1-Flash release); the API's
model name is now `deepseek-flash` and /v1/models lists only it. Hermes still
folded every non-V-series name onto deepseek-v4-flash, so `/model deepseek-flash`
on the DeepSeek provider was rewritten, then the validator "auto-corrected" it
back against the live listing: "Auto-corrected deepseek-v4-flash -> deepseek-flash"
on every switch.
Retired aliases (deepseek-chat / -reasoner and other fuzzy names) now fold onto
deepseek-flash; the curated catalog, profile fallback list, aux default, goal-judge
hint and pricing snapshot (2026-09-10 off-peak USD) follow the docs. Dated
deepseek-v4-* ids still pass through untouched.
Builds on YipTszkwan's #107126 (earliest fix in the cluster).
The Go relay (GET /zen/go/v1/models) delisted ox-alpha-free 2026-09-09, but
_PROVIDER_MODELS["opencode-go"] still carried it. _profile_live_catalog merges
the curated floor into the live list (live-first for opencode-go), so the
model picker kept offering a model that now 401s — the same failure class as
#95914 (opencode-free / x-preview-f-free).
Remove it from the floor and add two regression tests: an end-to-end merge
test through provider_model_ids with the real floor (fails if a stale floor
resurrects it) and a floor-pin test asserting the known-delisted model stays
out of the offline fallback.
(cherry picked from commit 091fd85865caa09e828928f93ee985dae7e9580d)
OpenRouter and Nous already ship these ids, but the native Anthropic
curated list still stopped at Fable 5 / Sonnet 5. Live /v1/models often
lags or 401s on subscription tokens, so the picker fell back to that
stale list and hid models that already work when addressed directly.
- Generalize Gemini 3 thinking config model prefix match to gemini-3*
- Add pricing snapshot entries for gemini-3.7-flash and gemini-3.8-flash
- Add model identifiers to Google, OpenRouter, and Vertex CLI lists
- Add test coverage for gemini-3.8-flash thinkingLevel configuration
The bare deepseek/deepseek-v4-flash slug is the pre-snapshot release; both
aggregators now carry deepseek-v4-flash-0731 (and Nous exposes the rolling
~deepseek/deepseek-v4-flash-latest alias). Keeping both rows in the curated
picker just duplicates the flash tier.
- OPENROUTER_MODELS (derives the nous list): drop deepseek/deepseek-v4-flash
- model-catalog.json regenerated
Deliberately KEPT: alibaba-token-plan / opencode-go / commandcode / deepseek
direct plugin fallback_models + default_aux_model (bare id is the wire slug
those providers serve), DEFAULT_CONTEXT_LENGTHS / reasoning floor / pricing
snapshot entries (manually-typed id still behaves), and the
deepseek-chat/deepseek-reasoner -> deepseek-v4-flash alias normalization.
Alibaba shipped qwen3.8-max-0902 (alias qwen3.8-max-2026-09-02), the upgraded
snapshot of Qwen3.8-Max: 1M context / 131K output, $2 in / $6 out. Both
aggregators now serve it and dropped the bare slug from their catalogs
(verified 2026-09-02: OpenRouter /v1/models + a 200 completion echoing the
slug; Nous Portal /v1/models).
- OPENROUTER_MODELS (feeds the nous list too): qwen/qwen3.8-max -> qwen/qwen3.8-max-0902
- _ALIBABA_TOKEN_PLAN_MODELS: qwen3.8-max-preview -> qwen3.8-max-0902
- test_empty_model_fallback fixture follows the nous catalog slug
- model-catalog.json regenerated
No new metadata: DEFAULT_CONTEXT_LENGTHS substring-matches qwen3.8-max (1M),
reasoning floor prefix fires (180s), both aggregator routes bill via
official_models_api. alibaba/alibaba-cn/opencode-go/setup.py left unchanged
(out of scope).
Six slugs land in the nous and openrouter curated lists, above the gpt-5.6 line:
openai/gpt-6-astra{,-fast,-flex} and openai/gpt-6-astra-pro{,-fast,-flex}.
Nous Portal serves the tiers as distinct slugs (verified live: each echoes its id, service_tier
default/priority/flex, cost 1x/2x/0.5x). OpenRouter serves them as ENDPOINTS of the base model
(tags openai/fast, openai/flex) and silently routes an unknown suffix to the standard tier at
standard price, so the OpenRouter profile rewrites a tier slug to its base wire model and pins
provider.only to that tier's endpoints (OPENROUTER_ENDPOINT_PINS). The base slug is pinned to
openai/azure/azure-us so default routing never lands on a flex or fast endpoint.
Provider-agnostic metadata: one DEFAULT_CONTEXT_LENGTHS entry (gpt-6-astra: 1,050,000, live on
OpenRouter for both models; substring-matches -pro and the tier suffixes). Pricing is skipped:
both routes bill via official_models_api. Reasoning floor not added (no evidence of long thinks).