Commit Graph

17 Commits

Author SHA1 Message Date
Teknium 30e4673776 refactor(models): import openrouter_variant_base from its defining module
Internal moves get no re-export aliases: models_validate reads
hermes_constants.openrouter_variant_base directly and the private
_OPENROUTER_VARIANT_SUFFIXES/_openrouter_variant_base shims are dropped.
2026-09-11 16:50:20 -07:00
jakobdylanc 284ee03df1 fix(models): resolve context length for OpenRouter :nitro/:floor routing variants
`:nitro`, `:floor`, `:exacto`, and `:online` are request-time routing
modifiers, not catalog models — OpenRouter's /models lists only the base
id, and a variant runs the same model with the same context window.

`get_model_context_length()` keyed every lookup on the full suffixed id,
so each one missed and the resolver fell through to a generic family
default or the 256K fallback:

  openai/gpt-5.5:nitro           -> 256K   (real 1.05M)
  x-ai/grok-4.6:nitro            -> 131K   (generic "grok" catch-all)
  anthropic/claude-opus-4.6:nitro-> 200K   (generic "claude" catch-all)

The window silently shrank, triggering early compression and a wrong
/usage readout. f14059fa fixed the sibling half of this bug class in
/model validation; this fixes the metadata half.

Strip a recognized variant suffix for LOOKUP only, keeping the suffixed
id on the wire so the routing opt-in survives. Applied after the explicit
config overrides (steps 0b/0c) so a user-pinned value still wins, and
before every cache/catalog lookup. Gated on the request actually routing
through OpenRouter, so a local Ollama `model:tag` is untouched.

`:free`/`:batch`/`:thinking` are deliberately excluded — those ARE
distinct catalog SKUs with their own windows, so stripping them would
report the wrong number.

The suffix set and base-id split move to hermes_constants (import-safe,
dependency-free) so the metadata layer shares one definition with
hermes_cli.models instead of duplicating it.
2026-09-11 16:50:20 -07:00
Teknium 073c57872a feat(models): DeepSeek V4.1 Flash on the Nous Portal and OpenRouter pickers
Add deepseek/deepseek-v4.1-flash to OPENROUTER_MODELS (Nous list derives from it),
regenerate the docs manifest, and give the slug its own 1M context entry and 600s
reasoning-stale floor — the longest-key-first scan otherwise lands the new slug on
the 128K `deepseek` catch-all and no floor. Live probed on both routes: echoed
model matches, usage.cost billed.
2026-09-10 09:22:31 -07:00
Teknium aeecb110f8 fix(deepseek): deepseek-flash is the canonical Flash id; retired names fold onto it
DeepSeek retired deepseek-v4-flash on 2026-09-10 (V4.1-Flash release); the API's
model name is now `deepseek-flash` and /v1/models lists only it. Hermes still
folded every non-V-series name onto deepseek-v4-flash, so `/model deepseek-flash`
on the DeepSeek provider was rewritten, then the validator "auto-corrected" it
back against the live listing: "Auto-corrected deepseek-v4-flash -> deepseek-flash"
on every switch.

Retired aliases (deepseek-chat / -reasoner and other fuzzy names) now fold onto
deepseek-flash; the curated catalog, profile fallback list, aux default, goal-judge
hint and pricing snapshot (2026-09-10 off-peak USD) follow the docs. Dated
deepseek-v4-* ids still pass through untouched.

Builds on YipTszkwan's #107126 (earliest fix in the cluster).
2026-09-10 02:44:26 -07:00
ten82e b2cae54035 fix(hermes_cli): drop delisted ox-alpha-free from opencode-go curated floor
The Go relay (GET /zen/go/v1/models) delisted ox-alpha-free 2026-09-09, but
_PROVIDER_MODELS["opencode-go"] still carried it. _profile_live_catalog merges
the curated floor into the live list (live-first for opencode-go), so the
model picker kept offering a model that now 401s — the same failure class as
#95914 (opencode-free / x-preview-f-free).

Remove it from the floor and add two regression tests: an end-to-end merge
test through provider_model_ids with the real floor (fails if a stale floor
resurrects it) and a floor-pin test asserting the known-delisted model stays
out of the offline fallback.

(cherry picked from commit 091fd85865caa09e828928f93ee985dae7e9580d)
2026-09-09 11:51:00 -07:00
xxxigm f91a78e9c4 fix(models): list Opus 5 and Fable 5.1 on the native Anthropic picker
OpenRouter and Nous already ship these ids, but the native Anthropic
curated list still stopped at Fable 5 / Sonnet 5. Live /v1/models often
lags or 401s on subscription tokens, so the picker fell back to that
stale list and hid models that already work when addressed directly.
2026-09-09 23:40:59 +05:30
tylman b0adce1cbf feat(models): add gemini-3.7-flash and gemini-3.8-flash support
- Generalize Gemini 3 thinking config model prefix match to gemini-3*
- Add pricing snapshot entries for gemini-3.7-flash and gemini-3.8-flash
- Add model identifiers to Google, OpenRouter, and Vertex CLI lists
- Add test coverage for gemini-3.8-flash thinkingLevel configuration
2026-09-06 05:42:22 -07:00
Teknium 21bcd0ce6c chore(models): OpenRouter and Nous pickers drop the undated deepseek-v4-flash in favor of the 0731 snapshot
The bare deepseek/deepseek-v4-flash slug is the pre-snapshot release; both
aggregators now carry deepseek-v4-flash-0731 (and Nous exposes the rolling
~deepseek/deepseek-v4-flash-latest alias). Keeping both rows in the curated
picker just duplicates the flash tier.

- OPENROUTER_MODELS (derives the nous list): drop deepseek/deepseek-v4-flash
- model-catalog.json regenerated

Deliberately KEPT: alibaba-token-plan / opencode-go / commandcode / deepseek
direct plugin fallback_models + default_aux_model (bare id is the wire slug
those providers serve), DEFAULT_CONTEXT_LENGTHS / reasoning floor / pricing
snapshot entries (manually-typed id still behaves), and the
deepseek-chat/deepseek-reasoner -> deepseek-v4-flash alias normalization.
2026-09-05 06:20:45 -07:00
Teknium d20a8e4475 feat(models): OpenRouter, Nous Portal, and Alibaba Token Plan pickers carry Qwen3.8-Max-0902
Alibaba shipped qwen3.8-max-0902 (alias qwen3.8-max-2026-09-02), the upgraded
snapshot of Qwen3.8-Max: 1M context / 131K output, $2 in / $6 out. Both
aggregators now serve it and dropped the bare slug from their catalogs
(verified 2026-09-02: OpenRouter /v1/models + a 200 completion echoing the
slug; Nous Portal /v1/models).

- OPENROUTER_MODELS (feeds the nous list too): qwen/qwen3.8-max -> qwen/qwen3.8-max-0902
- _ALIBABA_TOKEN_PLAN_MODELS: qwen3.8-max-preview -> qwen3.8-max-0902
- test_empty_model_fallback fixture follows the nous catalog slug
- model-catalog.json regenerated

No new metadata: DEFAULT_CONTEXT_LENGTHS substring-matches qwen3.8-max (1M),
reasoning floor prefix fires (180s), both aggregator routes bill via
official_models_api. alibaba/alibaba-cn/opencode-go/setup.py left unchanged
(out of scope).
2026-09-05 02:31:06 -07:00
Teknium f159e581c7 feat(models): add GPT-6 Astra + Astra Pro with fast/flex speed tiers to Nous Portal and OpenRouter
Six slugs land in the nous and openrouter curated lists, above the gpt-5.6 line:
openai/gpt-6-astra{,-fast,-flex} and openai/gpt-6-astra-pro{,-fast,-flex}.

Nous Portal serves the tiers as distinct slugs (verified live: each echoes its id, service_tier
default/priority/flex, cost 1x/2x/0.5x). OpenRouter serves them as ENDPOINTS of the base model
(tags openai/fast, openai/flex) and silently routes an unknown suffix to the standard tier at
standard price, so the OpenRouter profile rewrites a tier slug to its base wire model and pins
provider.only to that tier's endpoints (OPENROUTER_ENDPOINT_PINS). The base slug is pinned to
openai/azure/azure-us so default routing never lands on a flex or fast endpoint.

Provider-agnostic metadata: one DEFAULT_CONTEXT_LENGTHS entry (gpt-6-astra: 1,050,000, live on
OpenRouter for both models; substring-matches -pro and the tier suffixes). Pricing is skipped:
both routes bill via official_models_api. Reasoning floor not added (no evidence of long thinks).
2026-09-04 15:58:59 -07:00
Teknium 7b8c11bcf7 simplify(compat): models — drop 52 re-exports from hermes_cli.models, repoint 16 callers + 41 test files 2026-09-03 13:48:49 -07:00
Teknium 0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium c64a1fbd6c refactor(hermes_cli): repack literal tables, one-line trivial branches, share _drop_authorization 2026-09-02 21:57:37 -07:00
Teknium 8507b731bd refactor(hermes_cli): compact static catalog tables; derive nous/copilot/-cn twins from shared lists (values identical) 2026-09-02 20:21:32 -07:00
Teknium 64cf496ca5 refactor(models): collapse deepinfra/custom-cache/provider-list helpers; module docstrings 2026-09-02 16:24:19 -07:00
Teknium 96d8883986 refactor(models): collapse copilot catalog filter, probe/dedupe helpers, static-detect ladders; compact comments 2026-09-02 16:12:56 -07:00
Teknium cb5af57623 refactor(models): dict-dispatch provider_model_ids; extract static catalog tables and unified reasoning-caps cache into models_catalog_static / models_reasoning_caps 2026-09-02 15:53:59 -07:00