484 Commits

Author SHA1 Message Date
Teknium 93a3c40260 refactor(hermes_cli): models cluster — AST-neutral layout compaction (hug brackets, pack hanging lists) 2026-09-02 20:51:30 -07:00
Teknium c349485896 refactor(hermes_cli): models.py — unify cache row writes, route catalog GETs through _get_json, compact cache/copilot/opencode/deepinfra docs 2026-09-02 20:32:27 -07:00
Teknium 75ee8c00b6 refactor(hermes_cli): models.py — drop dead re-exports, unify JSON GET/merge/zero-price helpers, compact docstrings 2026-09-02 20:20:47 -07:00
Teknium 291c33040c refactor(models): fold credential-file mtime fingerprinting into one helper 2026-09-02 16:38:16 -07:00
Teknium 64cf496ca5 refactor(models): collapse deepinfra/custom-cache/provider-list helpers; module docstrings 2026-09-02 16:24:19 -07:00
Teknium 975783c31e refactor(models): table-drive opencode api-mode routing; collapse copilot/anthropic catalog helpers 2026-09-02 16:20:36 -07:00
Teknium 96d8883986 refactor(models): collapse copilot catalog filter, probe/dedupe helpers, static-detect ladders; compact comments 2026-09-02 16:12:56 -07:00
Teknium ade1c5a746 refactor(models): extract models_local + models_pricing; unify json disk-cache I/O, portal recommendation union, live-catalog index, copilot token exchange, probe/cache-entry builders 2026-09-02 16:06:33 -07:00
Teknium cb5af57623 refactor(models): dict-dispatch provider_model_ids; extract static catalog tables and unified reasoning-caps cache into models_catalog_static / models_reasoning_caps 2026-09-02 15:53:59 -07:00
Teknium 120aa68922 refactor(models): extract validate_requested_model into models_validate with shared verdict/match helpers; drop dead _OPENCODE_KEYLESS_EXTRA_SLUGS 2026-09-02 15:34:04 -07:00
Teknium 3ffd44acd3 refactor(hclib): models/runtime — models, inventory, runtime_provider, provider catalog, local_runtime, banner 2026-09-02 14:44:48 -07:00
Teknium 0ee98eda52 feat(models): add google/gemini-3.8-flash to nous + openrouter catalogs
Slots above gemini-3.7-flash (kept) in OPENROUTER_MODELS and
_PROVIDER_MODELS["nous"]; openrouter plugin fallback_models bumped
3.7 -> 3.8; model-catalog.json regenerated.

Verified live with test completions on both Nous Portal and OpenRouter
(model echo + billed). Same 1,048,576 window / 65,536 output / pricing
as 3.7-flash, so provider-agnostic metadata resolves via the existing
gemini entries and both routes bill live (official_models_api) — no
pricing snapshot needed.

Scoped to the two named providers: vertex/gemini/kilocode/gmi curated
lists, setup.py samples, and aux defaults untouched.
2026-09-02 09:46:15 -07:00
unsupportedpastels 5487222658 fix(models): live Copilot catalog for CLI-login users; unbreak picker-flag-empty catalogs
Three gaps between the copilot-acp picker row and what the user's
subscription actually serves (reported: picker showed the stale curated
list while the Copilot CLI offered Sonnet 5 / Opus 5 / GPT-5.6):

1. _resolve_copilot_catalog_api_key() never looked at the Copilot CLI's
   own token store (~/.copilot/config.json copilotTokens). A user whose
   only credential is 'copilot login' got no catalog key, the live fetch
   401'd, and copilot-acp silently fell back to the stale curated list.
   Add it as resolution source 3, JSONC-tolerant, with each candidate
   validated and exchanged like pool entries.

2. The existing credential-pool branch unpacked exchange_copilot_token()
   into two names, but it returns (api_token, expires_at, base_url) —
   the ValueError was swallowed by the enclosing except, disabling that
   entire resolution path. Latent since the base_url return was added.

3. GitHub now returns model_picker_enabled: false for EVERY model on
   some accounts/token types, so honoring the flag rejected the whole
   live catalog. Treat the flag as a display hint: when it empties the
   result, refilter without it (chat/endpoint checks still exclude
   embeddings and non-chat rows).

Verified live: catalog resolves 44 models for a copilot-login-only
account, matching the CLI's own picker (claude-sonnet-5, claude-opus-5,
gpt-5.6-sol/terra, gemini, kimi).
2026-09-02 20:51:07 +05:30
Teknium c7e2e0b779 feat(fast): bounded /fast auto|cold windows behind one route-aware gate
Adds two bounded fast modes on top of the static /fast toggle, default OFF:

- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
  window; requests inside it carry the provider fast param, later tool-loop
  requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
  user/assistant/tool history).

agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.

resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.

Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.

Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes #64785, #74730.

Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
2026-09-02 05:33:13 -07:00
Aldo 6ff9426d22 fix(anthropic): track the current fast-mode model matrix (Opus 4.8 / Opus 5)
The speed=fast allowlist still gates on Opus 4.6, but the fast-mode
matrix has changed twice since it was written (verified against the
live docs, platform.claude.com/docs/en/build-with-claude/fast-mode):

- Opus 4.8 and Opus 5 SUPPORT fast mode (research preview, Claude API
  only — not Bedrock/Vertex/Foundry).
- Opus 4.6 LOST fast mode on 2026-06-29. The parameter does not error:
  requests silently run at standard speed and bill standard rates
  (usage.speed: 'standard'). Today's allowlist therefore shows 4.6
  users a fast toggle that does nothing, while denying it to the two
  models that actually support it.
- Opus 4.7 never had it and hard-400s (unchanged).
- Dedicated '…-fast' model ids (OpenRouter's claude-opus-4.8-fast)
  select fast inference via the model field and are explicitly
  excluded from the param gate.

Both gates move in lock-step as before: the adapter param gate
(agent.anthropic_adapter._supports_fast_mode) and the CLI toggle gate
(hermes_cli.models._is_anthropic_fast_model). Docstrings now record
the history in both directions so the next matrix change has context.

## How to test

scripts/run_tests.sh tests/agent/test_anthropic_adapter.py tests/cli/test_fast_command.py -- -q

113 tests pass. The updated predicate/matrix tests fail against the
previous allowlist (verified by stashing the source changes). Tested
on Linux (aarch64).
2026-09-02 05:33:13 -07:00
Pedro Fontana b3576a29c3 Merge pull request #97354 from NousResearch/fix/nous-org-model-policy
fix(nous): honour the org model policy in the model pickers
2026-09-01 18:20:18 -03:00
Mariano Nicolini 3c4e84c166 fix(models): peek past expired and superseded pricing entries 2026-09-01 16:35:47 -03:00
liuhao1024 f82d2f1301 fix(models): scope prefix routing to user-configured providers only 2026-09-01 12:07:35 -07:00
liuhao1024 4033f3fc5f fix(models): honor vendor/model prefix and dict model.aliases in provider detection (#87189) 2026-09-01 12:07:35 -07:00
Teknium 9f069a1175 feat(models): add anthropic/claude-fable-5.1 to OpenRouter and Nous catalogs
Curated picker lists (OPENROUTER_MODELS + _PROVIDER_MODELS['nous']) gain
claude-fable-5.1 above claude-fable-5 per newest-first ordering; manifest
regenerated via scripts/build_model_catalog.py.

Provider-agnostic metadata verified as already resolving for the 5.1 slug
(no new entries needed): DEFAULT_CONTEXT_LENGTHS fuzzy-matches the
claude-fable-5 prefix (1,000,000), reasoning stale-timeout floor fires
(600s), and both routes bill via official_models_api (live pricing, no
snapshot entry required).
2026-09-01 11:31:58 -07:00
Mariano Nicolini 79972c6781 fix(models): expire the Nous catalog so policy changes land
A cached catalog was held for the life of the process, so a long-lived
gateway or desktop kept offering models the org had since blocked until
restart. Opt-in TTL — other providers keep no-expiry caching.
2026-08-31 16:18:07 -03:00
Mariano Nicolini e89f0087b4 fix(models): key the pricing cache per credential, not per auth state 2026-08-31 15:29:53 -03:00
Teknium c3948e6602 refactor(models): hoist preset suffix re-attachment into one helper
Follow-up to salvaged PR #89129: both auto-correct sites now call
_with_preset_suffix() so a future correction path can't forget to
re-attach the @preset/<slug> routing suffix.
2026-08-31 11:19:26 -07:00
mbac 3e912874bf fix(models): support OpenRouter preset references 2026-08-31 11:19:26 -07:00
Teknium 2215fb0e35 fix(providers): mirror new Qwen Cloud models onto alibaba-cn
Follow-up to the #87808 salvage: the domestic alibaba-cn picker list
gets the same five additions (same DashScope catalog, per models.dev).
2026-08-29 19:12:19 -07:00
icocode 04ef14e31f fix(providers): add missing Qwen Cloud (alibaba) models — qwen3.8-max, qwen3.6-flash, glm-5.2, deepseek-v4-pro/flash-0731 2026-08-29 19:12:19 -07:00
Teknium 93b6cf2ef8 feat(providers): curated picker lists for the Alibaba CN variants
Follow-up to the #77848 salvage: mirror the curated model lists onto
the alibaba-cn / alibaba-coding-plan-cn / alibaba-token-plan-cn
profiles registered in #73345, and add all Alibaba variants to the
qwen provider group so they appear in the drill-down picker.
2026-08-29 18:34:30 -07:00
MumuTW 46076b2d6b feat(providers): curated model list for alibaba-token-plan picker
Token Plan (Personal Edition) model catalog for hermes model /
provider pickers, verified against a live Token Plan subscription
(2026-08-03). Provider profiles landed separately in #73345; this
carries the picker-list half of #77848.

Co-salvaged-from: PR #77848
2026-08-29 18:34:30 -07:00
kshitijk4poor 4209d371aa refactor(models): reuse _extract_model_name in the Portal-recommendation validation tier
The inline set-comprehension re-implemented modelName extraction that
_extract_model_name() already provides (and that both
union_with_portal_free/paid_recommendations already use). Beyond the
duplication, the inline str(entry.get("modelName", "")) stringified
non-string values — a malformed Portal entry with modelName 5 would have
produced a garbage "5" match that discard("") does not filter. The helper
isinstance-checks and returns None for those, so routing through it makes
the validation tier semantically identical to the union helpers.

Adds test_non_string_model_name_entries_ignored locking the behavior
(mutation-checked: fails on the raw-stringify form, passes on the helper).
2026-08-29 22:41:55 +05:30
ygd58 0ffad55e09 fix(models): accept live Nous Portal recommendations in /model validation
Fixes #71312 (duplicate #71313).

When selecting a model via the Telegram /model picker (or any other
messaging-platform slash command, since they all share
validate_requested_model() through gateway/slash_commands.py ->
model_switch.switch_model()), a model available via Nous Portal's live
/api/nous/recommended-models endpoint but not yet in the hardcoded
curated catalog (_PROVIDER_MODELS["nous"]) was rejected with "was not
found in this provider's model listing" -- even though the exact same
model works fine via `hermes chat -m <model> --provider nous`.

Root cause: `hermes chat` merges Portal recommendations into its model
list via union_with_portal_free_recommendations() /
union_with_portal_paid_recommendations() at model-list build time
(hermes_cli/auth.py, web_server.py, model_setup_flows.py,
model_switch.py), so the model already appears "known" by the time
validation runs for that path. validate_requested_model() itself,
which every per-message /model command goes through, only checked the
live /v1/models listing and the curated catalog (_model_in_provider_catalog) --
never the Portal recommendations feed -- so a model that exists only
in Portal Recommendations was rejected on that path specifically.

Fix: add a Nous-specific fallback tier in validate_requested_model(),
checked after the curated-catalog fallback and before the final
rejection, reading the same fetch_nous_recommended_models() feed
(free + paid tiers) the CLI union helpers already use. Scoped to
provider == "nous" only; short-circuits before the network call when
an earlier tier already accepted the model; fails closed (rejects,
doesn't crash) if the Portal feed is unreachable.

Reported two issues filed 3 minutes apart with identical content by
the same author (#71312, #71313) -- commented on #71313 marking it a
duplicate of #71312 and pointing to this fix (could not close it
directly, no admin rights on the repo from this token).

6/6 new tests pass in TestValidateRequestedModelNousPortalRecommendations;
95/95 in the full tests/hermes_cli/test_model_validation.py file;
87/87 in tests/hermes_cli/test_models.py (unaffected, confirmed).
2026-08-29 22:41:55 +05:30
simonweng 0fb5cab0d4 feat:add hy4-preview model and tokenplan provider 2026-08-29 20:51:17 +05:30
amrrs 13bad590f3 feat(providers): add Nebius Token Factory provider 2026-08-29 20:39:44 +05:30
Teknium e2037a6c71 fix: follow-up for salvaged PR #95943
- Widen opencode_zen_free_runtime healing to the union of the static floor,
  the in-process live memo, and the SWR disk cache — a newly-live free model
  now heals opencode-go/zen selections without a release (sibling site the
  original PR missed).
- Memoize _fetch_opencode_free_models() in-process (5 min, negative caching
  included) so direct provider_model_ids() validation callers don't each
  block on a network round-trip or timeout.
- Drop delisted x-preview-f-free from the offline floor and setup.py sample
  list (offline fallback must not offer a model that 401s); add the newly
  live deepseek-v4-flash-free / mimo-v2.5-free to setup.py.
- Update stale test fixtures to a live exemplar; add regression tests for
  memoization, negative caching, and union healing; docs note in providers.md.
2026-08-29 05:02:54 -07:00
Jackal991 d9d6112aa7 fix(opencode): revalidate keyless opencode-free catalog live against the Zen relay
opencode-free (keyless) models were served exclusively from a hardcoded
in-repo snapshot (_PROVIDER_MODELS["opencode-free"]). The SWR disk cache
only revalidated AUTHED providers — its entries were keyed by a credential
fingerprint, which keyless providers have none of — so the catalog never
refreshed against GET /zen/v1/models. When the relay delisted a free model
(e.g. x-preview-f-free, 2026-08-26) the picker kept offering it and
selecting it 401'd: "Model x-preview-f-free is not supported".

Now provider_model_ids("opencode-free") fetches the live /zen/v1/models
catalog anonymously, filters it to the anonymous-servable free tier
(excluding KEYED suffix-fakes like Go's ox-alpha-free), and falls back to
the curated static floor only when the live fetch fails or is empty. The
keyless provider gets a stable disk-cache fingerprint so the picker's SWR
path serves stale immediately while refreshing off-thread — the same
behavior authed providers already get.

Regression tests prove the fix: the delisted/newly-live model assertions
fail when the live-fetch wiring is reverted.

Closes #95914
2026-08-29 05:02:54 -07:00
rob-maron f7c79efbac add tencent/hy4-preview to model pickers 2026-08-28 19:53:06 -07:00
Mariano Nicolini 705a10850d refactor(nous): trim comments and drop unused code 2026-08-28 17:00:26 -03:00
Mariano Nicolini 4d482ed344 refactor(nous): trim comments and drop an unused field 2026-08-28 15:39:23 -03:00
Mariano Nicolini da3c2435e2 fix(nous): only rescue an empty list where emptiness means "filtered out"
The fallback also ran on unavailable_models, which is legitimately empty on a
paid tier, filling the picker with the whole reachable set. Make it opt-in.
2026-08-28 13:34:27 -03:00
Mariano Nicolini 04647f15c8 fix(nous): only fall back to the reachable set when the overlap is empty
Surfacing allowed models the curated list lacks was gated on the size of the
reachable set alone. A jurisdiction or provider policy leaves few enough
models to pass that cap, so it appended the remainder — pushing non-curated
alphabetical ids into a picker that shows a curated order on purpose, and
making the list long enough that the non-curses fallback's input prompt
scrolled off screen and read as a hang.

Gate on the intersection instead. The fallback exists for an allowlist that
names nothing curated, which is the empty-overlap case; a policy that merely
narrows the catalog keeps the curated overlap and needs no help. The size cap
stays as a guard on that one path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 12:39:29 -03:00
Mariano Nicolini 117e7fef88 fix(nous): surface allowed models the curated list does not carry
An org allowlist can name a model the docs-hosted curated manifest has
never heard of. Intersecting the curated list against the reachable set
then produced an empty picker — "No models available for Nous Portal after
filtering" — which is strictly worse than showing an unfiltered list,
because the one model the org may actually use is the one that got dropped.

When the reachable set is small enough to be a human-authored allowlist,
append whatever it admits that the curated list is missing, after the
curated entries so their order survives.

Bounded by size, which is what separates the two kinds of policy: an
allowlist is small, while a provider-only policy leaves the whole catalog
reachable and appending it would bury the curated order. Past the cap the
intersection stands alone and the picker's custom-model entry remains the
way to reach anything omitted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 12:24:55 -03:00
Teknium 48d2528066 feat(models): qwen3.8-flash now selectable on OpenRouter and Nous portal
Live on both providers (verified 2026-08-28 against openrouter.ai/api/v1/models
and inference-api.nousresearch.com/v1/models) but absent from both curated
picker lists. Adds the entry directly below qwen3.8-max per newest-first
family ordering, an explicit 1M DEFAULT_CONTEXT_LENGTHS entry (new family
slug would otherwise fall through to the generic qwen 131072 catch-all —
same class as #69881), and regenerates model-catalog.json.

Scoped rollout: only the named providers touched. Pricing snapshot skipped
(both routes bill via official_models_api live pricing). Reasoning floor
already fires via the qwen3 prefix entry (180s, verified).
2026-08-28 01:09:58 -07:00
Mariano Nicolini c248d5356c feat(nous): read the org model policy and expose it as a list filter
A Nous team admin can restrict which models and which serving providers
their org may use. The inference gateway applies that policy to
`GET /v1/models`, omitting blocked rows with no marker field, so the keys
of an authenticated catalog read are the reachable set.

Add the two pieces the pickers need:

`nous_policy_present()` reads the `policy_present` claim off the OAuth
access token, which costs no request. `/api/oauth/account` does not carry
the claim, so this reads the token rather than going through
`get_nous_portal_account_info`. The claim is tri-state — absent means an
older mint, which is not the same as "no policy" and must not be reported
as one.

`nous_policy_allowed_ids()` turns the authenticated pricing response into
that set, reusing the cache entry a caller asking for pricing already
populates rather than issuing a second round trip. It returns None —
"leave the list alone" — for an org with no policy, for an anonymous read
whose catalog is unfiltered, and for an empty read, each of which would
otherwise narrow a list on evidence that cannot support it.

`restrict_to_nous_policy()` applies the set while preserving the caller's
order, and keeps a `:free` sibling whose base model is reachable. The
gateway admits a row when any of its requestable ids passes and treats
anything unknown as a keep, on the grounds that over-listing costs a 403
from the authoritative gate while hiding a row the gate would serve is
unrecoverable from the client. This mirrors that.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:13:13 -03:00
Mariano Nicolini 4caeb02735 fix(models): key the pricing cache on auth state, not just the base URL
`fetch_models_with_pricing` checked its cache above the point where the
Authorization header is built, and keyed that cache on the base URL alone.
Whichever read of a given base URL landed first in a process therefore
answered every later read, whatever key it passed — a non-empty result is
held for the life of the process.

That is wrong for any endpoint whose answer depends on who is asking. The
Nous inference gateway filters `GET /v1/models` by the caller's org model
policy, so an anonymous read landing first makes a later authenticated read
return the full, unfiltered catalog without a request going out.

Separate the URL root from the cache key and fold auth state into the
latter. Only whether a key was supplied participates, never its value, so
no secret reaches the key.

`credits_tracker` peeked into the private `_pricing_cache` and duplicated
the key shape to do it; it now calls `peek_cached_pricing`, which owns both
the /v1-suffix normalization and the preference for the authenticated
catalog.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 19:12:58 -03:00
Teknium ab80a38514 fix(models): delist openrouter/elephant-alpha (no longer served by OpenRouter)
Live /api/v1/models probe (2026-08-27) confirms the id is gone from the
catalog, so the curated picker entry was a dead pick. Manifest
regenerated. No provider-agnostic metadata existed for the slug.

Delist credit: @orouge97 flagged this in PR #80036.
2026-08-27 04:37:31 -07:00
Adolanium a9611f3c6f feat(models): add GLM-5.3-Flash to z.ai and OpenCode Go pickers
OpenRouter and Nous already list z-ai/glm-5.3-flash (#95621). The
native z.ai picker, OpenCode Go/Zen fallbacks, setup wizard, and
Coding Plan probes did not. Context still resolves through the
existing glm-5.3 1M key.
2026-08-27 04:14:31 -07:00
Teknium 9b44273c05 fix: follow-up for salvaged PRs #93250 + #96234
- move minimax/minimax-m3:free into the Free tier section (house
  convention: :free SKUs group together, matching glm-5.2:free and the
  nemotron :free entries) and regenerate model-catalog.json
- add Inkling family context length (1,048,576 — OpenRouter live
  metadata, 2026-08-27) to DEFAULT_CONTEXT_LENGTHS; new family slug
  otherwise fell through to no entry
- add Inkling to the reasoning stale-timeout floor table (300s tier,
  same as Grok reasoning / Ox Alpha; OpenRouter marks the family as
  reasoning-capable)
- widen the floor matcher's right-anchor separator class to include
  ':' so OpenRouter SKU suffixes (:free/:batch/:nitro) inherit the
  family floor — inkling:free previously missed the inkling entry
- regression tests for the inkling floor + ':' separator
2026-08-27 03:24:20 -07:00
Eric Zhang 51df117def feat(models): add Inkling free models to OpenRouter catalog 2026-08-27 03:24:20 -07:00
kshitij 6607f70673 feat(models): add minimax/minimax-m3:free to OpenRouter picker (#96234)
The OpenRouter model picker builds its list from a curated set of model
IDs, then filters against OpenRouter's live catalog. minimax/minimax-m3:free
exists on OpenRouter (free tier, 1M context, tool-calling) but was missing
from both the in-repo fallback list and the remote catalog manifest.

Add it to OPENROUTER_MODELS and website/static/api/model-catalog.json so
the free variant surfaces in the picker alongside the paid one.
2026-08-27 09:34:48 +00:00
Teknium 64424a16a2 feat(models): add z-ai/glm-5.3-flash to OpenRouter and Nous Portal catalogs
Slots below glm-5.3, above glm-5.2 in both curated lists; regenerates
model-catalog.json. No new metadata entries needed: context resolves via
the existing glm-5.3 fuzzy key (1,048,576 — matches OpenRouter live), and
both routes bill via official_models_api (live pricing).
2026-08-26 07:52:56 -07:00
Teknium f277637bde chore: remove stealth/ox-alpha from OpenRouter and Nous Portal model catalogs
Drops the retired Ox Alpha stealth preview from both curated picker lists
and regenerates website/static/api/model-catalog.json. Metadata entries
(context window, reasoning timeout) and the generic stealth/ free-tier
policy are left intact so manually-entered ids still behave.
2026-08-26 07:25:03 -07:00