Commit Graph

405 Commits

Author SHA1 Message Date
Teknium c3948e6602 refactor(models): hoist preset suffix re-attachment into one helper
Follow-up to salvaged PR #89129: both auto-correct sites now call
_with_preset_suffix() so a future correction path can't forget to
re-attach the @preset/<slug> routing suffix.
2026-08-31 11:19:26 -07:00
mbac 3e912874bf fix(models): support OpenRouter preset references 2026-08-31 11:19:26 -07:00
Teknium 2215fb0e35 fix(providers): mirror new Qwen Cloud models onto alibaba-cn
Follow-up to the #87808 salvage: the domestic alibaba-cn picker list
gets the same five additions (same DashScope catalog, per models.dev).
2026-08-29 19:12:19 -07:00
icocode 04ef14e31f fix(providers): add missing Qwen Cloud (alibaba) models — qwen3.8-max, qwen3.6-flash, glm-5.2, deepseek-v4-pro/flash-0731 2026-08-29 19:12:19 -07:00
Teknium 93b6cf2ef8 feat(providers): curated picker lists for the Alibaba CN variants
Follow-up to the #77848 salvage: mirror the curated model lists onto
the alibaba-cn / alibaba-coding-plan-cn / alibaba-token-plan-cn
profiles registered in #73345, and add all Alibaba variants to the
qwen provider group so they appear in the drill-down picker.
2026-08-29 18:34:30 -07:00
MumuTW 46076b2d6b feat(providers): curated model list for alibaba-token-plan picker
Token Plan (Personal Edition) model catalog for hermes model /
provider pickers, verified against a live Token Plan subscription
(2026-08-03). Provider profiles landed separately in #73345; this
carries the picker-list half of #77848.

Co-salvaged-from: PR #77848
2026-08-29 18:34:30 -07:00
kshitijk4poor 4209d371aa refactor(models): reuse _extract_model_name in the Portal-recommendation validation tier
The inline set-comprehension re-implemented modelName extraction that
_extract_model_name() already provides (and that both
union_with_portal_free/paid_recommendations already use). Beyond the
duplication, the inline str(entry.get("modelName", "")) stringified
non-string values — a malformed Portal entry with modelName 5 would have
produced a garbage "5" match that discard("") does not filter. The helper
isinstance-checks and returns None for those, so routing through it makes
the validation tier semantically identical to the union helpers.

Adds test_non_string_model_name_entries_ignored locking the behavior
(mutation-checked: fails on the raw-stringify form, passes on the helper).
2026-08-29 22:41:55 +05:30
ygd58 0ffad55e09 fix(models): accept live Nous Portal recommendations in /model validation
Fixes #71312 (duplicate #71313).

When selecting a model via the Telegram /model picker (or any other
messaging-platform slash command, since they all share
validate_requested_model() through gateway/slash_commands.py ->
model_switch.switch_model()), a model available via Nous Portal's live
/api/nous/recommended-models endpoint but not yet in the hardcoded
curated catalog (_PROVIDER_MODELS["nous"]) was rejected with "was not
found in this provider's model listing" -- even though the exact same
model works fine via `hermes chat -m <model> --provider nous`.

Root cause: `hermes chat` merges Portal recommendations into its model
list via union_with_portal_free_recommendations() /
union_with_portal_paid_recommendations() at model-list build time
(hermes_cli/auth.py, web_server.py, model_setup_flows.py,
model_switch.py), so the model already appears "known" by the time
validation runs for that path. validate_requested_model() itself,
which every per-message /model command goes through, only checked the
live /v1/models listing and the curated catalog (_model_in_provider_catalog) --
never the Portal recommendations feed -- so a model that exists only
in Portal Recommendations was rejected on that path specifically.

Fix: add a Nous-specific fallback tier in validate_requested_model(),
checked after the curated-catalog fallback and before the final
rejection, reading the same fetch_nous_recommended_models() feed
(free + paid tiers) the CLI union helpers already use. Scoped to
provider == "nous" only; short-circuits before the network call when
an earlier tier already accepted the model; fails closed (rejects,
doesn't crash) if the Portal feed is unreachable.

Reported two issues filed 3 minutes apart with identical content by
the same author (#71312, #71313) -- commented on #71313 marking it a
duplicate of #71312 and pointing to this fix (could not close it
directly, no admin rights on the repo from this token).

6/6 new tests pass in TestValidateRequestedModelNousPortalRecommendations;
95/95 in the full tests/hermes_cli/test_model_validation.py file;
87/87 in tests/hermes_cli/test_models.py (unaffected, confirmed).
2026-08-29 22:41:55 +05:30
simonweng 0fb5cab0d4 feat:add hy4-preview model and tokenplan provider 2026-08-29 20:51:17 +05:30
amrrs 13bad590f3 feat(providers): add Nebius Token Factory provider 2026-08-29 20:39:44 +05:30
Teknium e2037a6c71 fix: follow-up for salvaged PR #95943
- Widen opencode_zen_free_runtime healing to the union of the static floor,
  the in-process live memo, and the SWR disk cache — a newly-live free model
  now heals opencode-go/zen selections without a release (sibling site the
  original PR missed).
- Memoize _fetch_opencode_free_models() in-process (5 min, negative caching
  included) so direct provider_model_ids() validation callers don't each
  block on a network round-trip or timeout.
- Drop delisted x-preview-f-free from the offline floor and setup.py sample
  list (offline fallback must not offer a model that 401s); add the newly
  live deepseek-v4-flash-free / mimo-v2.5-free to setup.py.
- Update stale test fixtures to a live exemplar; add regression tests for
  memoization, negative caching, and union healing; docs note in providers.md.
2026-08-29 05:02:54 -07:00
Jackal991 d9d6112aa7 fix(opencode): revalidate keyless opencode-free catalog live against the Zen relay
opencode-free (keyless) models were served exclusively from a hardcoded
in-repo snapshot (_PROVIDER_MODELS["opencode-free"]). The SWR disk cache
only revalidated AUTHED providers — its entries were keyed by a credential
fingerprint, which keyless providers have none of — so the catalog never
refreshed against GET /zen/v1/models. When the relay delisted a free model
(e.g. x-preview-f-free, 2026-08-26) the picker kept offering it and
selecting it 401'd: "Model x-preview-f-free is not supported".

Now provider_model_ids("opencode-free") fetches the live /zen/v1/models
catalog anonymously, filters it to the anonymous-servable free tier
(excluding KEYED suffix-fakes like Go's ox-alpha-free), and falls back to
the curated static floor only when the live fetch fails or is empty. The
keyless provider gets a stable disk-cache fingerprint so the picker's SWR
path serves stale immediately while refreshing off-thread — the same
behavior authed providers already get.

Regression tests prove the fix: the delisted/newly-live model assertions
fail when the live-fetch wiring is reverted.

Closes #95914
2026-08-29 05:02:54 -07:00
rob-maron f7c79efbac add tencent/hy4-preview to model pickers 2026-08-28 19:53:06 -07:00
Teknium 48d2528066 feat(models): qwen3.8-flash now selectable on OpenRouter and Nous portal
Live on both providers (verified 2026-08-28 against openrouter.ai/api/v1/models
and inference-api.nousresearch.com/v1/models) but absent from both curated
picker lists. Adds the entry directly below qwen3.8-max per newest-first
family ordering, an explicit 1M DEFAULT_CONTEXT_LENGTHS entry (new family
slug would otherwise fall through to the generic qwen 131072 catch-all —
same class as #69881), and regenerates model-catalog.json.

Scoped rollout: only the named providers touched. Pricing snapshot skipped
(both routes bill via official_models_api live pricing). Reasoning floor
already fires via the qwen3 prefix entry (180s, verified).
2026-08-28 01:09:58 -07:00
Teknium ab80a38514 fix(models): delist openrouter/elephant-alpha (no longer served by OpenRouter)
Live /api/v1/models probe (2026-08-27) confirms the id is gone from the
catalog, so the curated picker entry was a dead pick. Manifest
regenerated. No provider-agnostic metadata existed for the slug.

Delist credit: @orouge97 flagged this in PR #80036.
2026-08-27 04:37:31 -07:00
Adolanium a9611f3c6f feat(models): add GLM-5.3-Flash to z.ai and OpenCode Go pickers
OpenRouter and Nous already list z-ai/glm-5.3-flash (#95621). The
native z.ai picker, OpenCode Go/Zen fallbacks, setup wizard, and
Coding Plan probes did not. Context still resolves through the
existing glm-5.3 1M key.
2026-08-27 04:14:31 -07:00
Teknium 9b44273c05 fix: follow-up for salvaged PRs #93250 + #96234
- move minimax/minimax-m3:free into the Free tier section (house
  convention: :free SKUs group together, matching glm-5.2:free and the
  nemotron :free entries) and regenerate model-catalog.json
- add Inkling family context length (1,048,576 — OpenRouter live
  metadata, 2026-08-27) to DEFAULT_CONTEXT_LENGTHS; new family slug
  otherwise fell through to no entry
- add Inkling to the reasoning stale-timeout floor table (300s tier,
  same as Grok reasoning / Ox Alpha; OpenRouter marks the family as
  reasoning-capable)
- widen the floor matcher's right-anchor separator class to include
  ':' so OpenRouter SKU suffixes (:free/:batch/:nitro) inherit the
  family floor — inkling:free previously missed the inkling entry
- regression tests for the inkling floor + ':' separator
2026-08-27 03:24:20 -07:00
Eric Zhang 51df117def feat(models): add Inkling free models to OpenRouter catalog 2026-08-27 03:24:20 -07:00
kshitij 6607f70673 feat(models): add minimax/minimax-m3:free to OpenRouter picker (#96234)
The OpenRouter model picker builds its list from a curated set of model
IDs, then filters against OpenRouter's live catalog. minimax/minimax-m3:free
exists on OpenRouter (free tier, 1M context, tool-calling) but was missing
from both the in-repo fallback list and the remote catalog manifest.

Add it to OPENROUTER_MODELS and website/static/api/model-catalog.json so
the free variant surfaces in the picker alongside the paid one.
2026-08-27 09:34:48 +00:00
Teknium 64424a16a2 feat(models): add z-ai/glm-5.3-flash to OpenRouter and Nous Portal catalogs
Slots below glm-5.3, above glm-5.2 in both curated lists; regenerates
model-catalog.json. No new metadata entries needed: context resolves via
the existing glm-5.3 fuzzy key (1,048,576 — matches OpenRouter live), and
both routes bill via official_models_api (live pricing).
2026-08-26 07:52:56 -07:00
Teknium f277637bde chore: remove stealth/ox-alpha from OpenRouter and Nous Portal model catalogs
Drops the retired Ox Alpha stealth preview from both curated picker lists
and regenerates website/static/api/model-catalog.json. Metadata entries
(context window, reasoning timeout) and the generic stealth/ free-tier
policy are left intact so manually-entered ids still behave.
2026-08-26 07:25:03 -07:00
Teknium 03c97d984b Revert "chore: remove stealth/ox-alpha from OpenRouter and Nous Portal model catalogs"
This reverts commit 8e1bc8d342.
2026-08-26 07:04:42 -07:00
Teknium 8e1bc8d342 chore: remove stealth/ox-alpha from OpenRouter and Nous Portal model catalogs
Drops the retired Ox Alpha stealth preview from both curated picker lists
and regenerates website/static/api/model-catalog.json. Metadata entries
(context window, reasoning timeout) and the generic stealth/ free-tier
policy are left intact so manually-entered ids still behave.
2026-08-26 07:03:53 -07:00
Teknium f14059fad2 fix(models): OpenRouter :nitro/:floor routing variants no longer rejected by /model validation
OpenRouter's :nitro, :floor, :exacto, and :online suffixes are request-time
routing modifiers valid on any model id — /models lists only the base model.
validate_requested_model() compared the full suffixed id against the listing,
so a valid variant was either rejected outright or fuzzy-auto-corrected to
the base id, silently stripping the user's routing opt-in.

Now, for OpenRouter only, a recognized variant suffix validates the BASE id
against the live listing (and the curated-catalog soft-accept and static-
catalog fallback paths) while preserving the suffixed id for persistence and
API requests — checked BEFORE fuzzy correction. :free/:batch/:thinking
remain direct catalog SKUs and keep exact-match semantics; unknown suffixes
and unknown bases are still rejected.

Reported by JEB (Jakob's Hermes Agent) via Discord.
2026-08-24 13:36:23 -07:00
Teknium 7b89e17774 fix(model): single exact eligibility predicate for -900k variants; reject ineligible aliases
Review findings on #92797 (@100yenadmin):
- is_codex_900k_base() is now the single source of truth used by picker
  synthesis, context resolution, /model validation, and wire stripping.
  Eligibility is an exact table (sol/terra/luna, gpt-5.4, daybreak alias)
  plus date-shaped 5.6 snapshots — family-prefix matching removed, so
  non-routable -pro slugs and unknown descendants never gain variants.
- strip_codex_context_variant_suffix() strips conditionally: ineligible
  aliases (gpt-5.5-900k) are returned unchanged and fail honestly at the
  API instead of silently running as the base model at 272K.
- validate_requested_model() rejects ineligible *-900k aliases before the
  hidden-slug soft-accept, and accepts valid variants missing from a
  stale catalog without letting the typo auto-corrector eat the suffix.
- Codex context resolver drops vendor/ namespaces, so
  openai/gpt-5.6-sol-900k resolves to 900K like the bare id.
- Table-driven regression covering eligible bases/snapshots/namespaced
  ids and rejected -pro/-mini/5.5/unknown aliases, asserting context AND
  wire model.
2026-08-23 02:14:35 -07:00
Teknium 63a9c26fbe feat(model): Codex GPT slugs default back to 272K; explicit -900k picker variants opt into the verified large window
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the
live-verified 900K burned through subscription usage for users who never
asked for the larger window (bigger window = more input tokens per request).

- Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the
  advertised 272K again — the cheaper limit is the default.
- The model picker synthesizes explicit <slug>-900k variants (e.g.
  gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into
  the 900K window. Slugs that genuinely enforce 272K (gpt-5.5,
  gpt-5.4-mini) get no variant.
- The -900k suffix is Hermes-side only: stripped before the model id hits
  the wire (main transport + auxiliary Responses adapter), and pricing
  aliases the variants onto the base entries.
- Docs: new opt-in section in context-compression-and-caching.md.
2026-08-23 02:14:35 -07:00
Kevin Yin 9ce46a0968 fix(models): bind Anthropic pool key to its endpoint 2026-08-22 23:25:29 -07:00
Kevin Yin 8ede2e1472 fix(models): discover Anthropic pool API keys 2026-08-22 23:25:29 -07:00
Vinay Shah 5cb7b521cf refactor(bedrock): make resolve_bedrock_runtime_region the single region chokepoint
Follow-up structural pass on the review fix:

- Runtime provider, auxiliary resolution, model validation
  (hermes_cli/models.py), live discovery (bedrock_model_ids_or_none),
  and the Mantle URL/SigV4 fallbacks all resolve their region through
  resolve_bedrock_runtime_region() — one canonical implementation of the
  config-first priority instead of three hand-rolled copies.
- agent_init: drop the 'if "client_kwargs" in locals()' guard by
  initializing client_kwargs unconditionally at the top of the else
  branch; the Mantle kwargs hook is a documented no-op for non-Mantle
  base URLs.
2026-08-21 15:02:29 -07:00
Vinay Shah 16476fad10 feat(bedrock): add OpenAI GPT-5.6 family (Sol/Terra/Luna) to Mantle Responses routing
GPT-5.6 Sol, Terra, and Luna went GA on Amazon Bedrock on 2026-07-13.
Like GPT-5.5, they are served exclusively from the Bedrock Mantle
OpenAI-compatible Responses endpoint (the model cards list
bedrock-runtime/Converse as unsupported), so they ride the allowlist
routing introduced for GPT-5.5:

- Add openai.gpt-5.6-{sol,terra,luna} to BEDROCK_OPENAI_RESPONSES_MODEL_IDS
  so runtime resolution, auxiliary calls, and MoA slots all take the
  SigV4/bearer Mantle Responses path.
- Surface the family in the curated Bedrock picker list.
- Record the 272K context window from the AWS model cards for all four
  Mantle OpenAI models (previously fell back to the 128K default).
- Generalize picker tests from the hardcoded single-model checks to the
  BEDROCK_OPENAI_RESPONSES_MODEL_IDS allowlist so future Mantle model
  additions do not require test surgery; add routing, picker, and
  context-length coverage for the 5.6 family.

Docs: https://docs.aws.amazon.com/bedrock/latest/userguide/model-cards-openai.html
2026-08-21 15:02:29 -07:00
Nathaniel Branscum e5b96fcb10 feat(bedrock): support OpenAI Responses models
Route Bedrock-hosted OpenAI GPT-5.5 through the Bedrock Mantle OpenAI Responses endpoint with SigV4 request signing. Keep native Bedrock Converse and Claude Bedrock routing unchanged, and add picker/runtime regression coverage.
2026-08-21 15:02:29 -07:00
Teknium bd93a5f316 feat(models): free models show star + -100% in the model picker discount column
Free ($0/$0) Nous Portal models sat with a blank discount column and no
sale star (stealth/ox-alpha, upstage/solar-pro4:free), reading as missing
data next to the -20% sale rows. compute_sale_discount now returns a flat
100% for free models; was_* raws pass through only when the gateway served
a pricing.original, so natively-free models render bare '-100%' with no
fabricated 'was ?/?'. CLI picker star follows on_sale automatically;
inventory feed carries discount_percent=100 to Desktop, whose FREE badge
row now renders the amber -100% pill beside it.
2026-08-21 14:38:41 -07:00
Teknium 907da145b6 feat(models): glm-5.3 replaces glm-5.1 in the OpenRouter and Nous Portal catalogs
Follow-up to the salvaged GLM-5.3 support commit: drop z-ai/glm-5.1 from
both curated lists per Teknium's direction (glm-5.2 keeps the 'default'
tag), and regenerate the docs manifest. glm-5.1 remains available via
live discovery and on out-of-scope surfaces (zai plugin, setup defaults,
opencode-go) — named leftovers, not silently swept.
2026-08-21 14:38:32 -07:00
openclaw 01d8562fce fix(zai): add GLM-5.3 support — 1M context window, model lists, reasoning_effort
GLM-5.3 is live on api.z.ai (coding plan endpoint) but had no entries in
Hermes, so it silently fell back to the generic 202K GLM context —
triggering premature context compression on a 1M-window model.

- model_metadata: 'glm-5.3': 1_048_576 (same base model as 5.2; 1M
  context / 128K max output per docs.z.ai/guides/llm/glm-5.3, verified
  2026-08-14)
- auth: add glm-5.3 to coding-plan probe lists (global + CN)
- models: add glm-5.3 to picker/model lists (6 sites)
- zai provider: reasoning_effort mapping covers glm-5.3 (accepted live
  by the endpoint, HTTP 200)
2026-08-21 14:38:32 -07:00
Teknium 67af79d7e1 feat(models): stealth/ox-alpha free model in the Nous Portal catalog
Third surface for the Ox Alpha stealth reasoning model (after the
OpenCode Zen rollout in #91250 and the OpenRouter listing in #91284).
Adds stealth/ox-alpha to the curated Nous list and regenerates the docs
manifest. Free on the portal ($0/$0), 1M context, 131K max output —
verified against the live inference-api.nousresearch.com/v1/models.

Provider-agnostic metadata already resolves via the bare ox-alpha slug
(DEFAULT_CONTEXT_LENGTHS 1,048,576; reasoning_timeouts 300s floor), and
the nous route bills via official_models_api, so no pricing snapshot is
needed.
2026-08-21 13:06:42 -07:00
Teknium 624723130b fix: catalog drift sync — dead free slugs out, live OpenRouter free models in, ox-alpha-free on Go
First actioned report from the overhauled model-catalog-scout cron
(2026-08-21 validation run), every item re-verified live before edit:

Delisted (gone from live catalogs):
- opencode-zen curated: claude-opus-4-1, qwen3.7-max, qwen3.7-plus
  (absent from live zen /v1/models; qwen3.7 family remains on Go)
- OPENROUTER_MODELS free section: poolside/laguna-m.1:free (rotated to
  s-2.1/xs-2.1), tencent/hy3:free, inclusionai/ring-2.6-1t:free

Added (present + verified in live catalogs):
- OpenRouter free: z-ai/glm-5.2:free (256K), poolside/laguna-s-2.1:free
  + laguna-xs-2.1:free (262K), nvidia/nemotron-3.5-lightning:free (1M)
- opencode-go curated: ox-alpha-free (Go-subscription twin of the Zen
  keyless Ox Alpha; keyed — Go relay 401s anonymous requests)

Metadata:
- DEFAULT_CONTEXT_LENGTHS: laguna-s-2.1/xs-2.1 262144;
  nemotron-3.5-lightning 1M (overrides the generic 131K nemotron entry);
  glm-5.2:free 256K (the free variant is capped below the 1M paid entry)

Keyless-heal hardening (the real find):
- opencode_zen_free_runtime now gates the zen/go→keyless heal on
  MEMBERSHIP in the verified opencode-free catalog, not the -free
  suffix — ox-alpha-free is a KEYED Go model despite its suffix, and
  suffix-based healing would have routed it to a Zen relay that
  doesn't serve it (verified: zen 401s 'not supported', go 401s
  'Missing API key'). New regression test pins this.

Fixture sweep: tencent/hy3:free catalog assertion updated (delisted
slug); nous-route fixtures using hy3:free as incidental model names
left alone (self-consistent mocks). model-catalog.json regenerated.
2026-08-21 04:44:15 -07:00
Teknium 30ccd01ba2 fix: delist UA-gated free models from opencode-free — big-pickle and mimo-v2.5-free 429 every non-opencode client
Live verification (2026-08-21): big-pickle and mimo-v2.5-free return 429
FreeUsageLimitError for ANY User-Agent except the opencode CLI's own
'opencode/latest' — same IP, no cooldown effect, while the other six free
models serve our honest HermesAgent UA freely. Hermes sends deliberate
attribution headers and does not impersonate other clients, so these two
models are broken for our users by policy on OpenCode's side; delist them
rather than ship dead picker entries.

- opencode-free catalog: 8 -> 6 models (both curated lists)
- plugin default_aux_model: big-pickle -> laguna-s-2.1-free (fastest
  non-gated free model)
- keyless predicate keeps big-pickle (it IS free-tier; correct routing if
  a user enters it manually — they get the relay's own 429, not our 401)

E2E: picker shows 6, laguna aux default completes a keyless agent turn.
2026-08-21 04:00:16 -07:00
Teknium ca06b87689 feat: opencode-free is fully keyless — no env var, no account, anonymous wire
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).

On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
  never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
  Zen endpoint routing incl. muse->responses); keyless predicate extended
  with unsuffixed free slugs (big-pickle); free runtime pins EVERY
  opencode-free model keyless; curated catalog replaces the models.dev
  cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
  there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
  wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
  invariant (every curated model must satisfy the keyless predicate)

E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
2026-08-21 00:24:32 -07:00
Rudraksh Chahal 28a9b6c565 feat(providers): add OpenCode Free provider with keyed auth and opencode User-Agent
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.

The free tier requires a real account API key and throttles third-party
clients by User-Agent:

- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
  and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
  Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
  unconditionally (the stale keyless-tier assumption), and credential-pool
  exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
  message.

Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
2026-08-21 00:24:32 -07:00
Teknium c01cd26f95 feat: stealth/ox-alpha free model in the OpenRouter catalog
Adds OpenRouter's free "Ox Alpha" stealth reasoning model
(stealth/ox-alpha) to the OpenRouter fallback snapshot, plus the
provider-agnostic metadata it needs:

- OPENROUTER_MODELS: free-tier entry (1M ctx)
- DEFAULT_CONTEXT_LENGTHS: ox-alpha -> 1,048,576 (verified against
  OpenRouter live /api/v1/models; without this the slug fell through
  to no match)
- reasoning_timeouts.py: 300s stale floor for ox-alpha and the
  OpenCode Zen twin slug x-preview-f-free (reasoning model,
  long-horizon agentic work per its model card)
- model-catalog.json regenerated

Pricing snapshot skipped: openrouter bills via official_models_api
(live pricing; model is free anyway).
2026-08-20 23:12:38 -07:00
Teknium 1017a56274 fix: OpenCode Zen free-tier models (Ox Alpha) work keyless — any bearer 401s them
The Zen relay serves *-free models (x-preview-f-free / Ox Alpha) ONLY
anonymously: any Authorization bearer it doesn't recognize is a 401
'Invalid API key' — including our no-key-required placeholder and valid
OpenCode GO subscription keys. The Go relay doesn't serve the free tier
at all ('Model x is not supported'). So the free model failed for every
Hermes user: keyless setups got the placeholder bearer, and OpenCode
subscribers sent a Go key to a relay that rejects it.

Fix (class-wide for all 8 current *-free Zen slugs, not just Ox Alpha):
- hermes_cli/models.py: is_opencode_zen_free_model / opencode_zen_free_runtime
  / opencode_zen_free_headers — one shared policy: free slugs pin to the
  Zen relay with a keyless placeholder and an empty Authorization header
  that overrides the OpenAI SDK's 'Bearer <key>'.
- runtime_provider.py: free slugs route through the keyless runtime before
  the credential-pool/explicit/api_key paths (no key required; Go
  selections heal to Zen). Paid models still fail closed without a key.
- agent_init.py + auxiliary_client.py: the placeholder key swaps in the
  empty-Authorization headers at both client-build chokepoints.

Verified live (2026-08-21): anonymous chat/completions 200 incl. tools,
streaming, parallel; bad bearer 401; full E2E AIAgent turn with a real
terminal tool round-trip completes keyless under both opencode-zen and
opencode-go providers. Sabotage run: routing tests fail without the fix.
2026-08-20 23:05:57 -07:00
Teknium a83c3915a3 fix(opencode): family-wide provider predicate + reserved tool-name aliases for custom opencode-* providers
Builds on @Lesnak1's #85619 (issue #85589):

- New opencode_provider_family() single-owner predicate in
  hermes_cli/models.py — resolves built-in AND custom family providers
  (opencode-go-bridge, OpenCode-Zen-Custom, ...) case-insensitively.
  Migrated all 8 inlined family checks (models.py x3, runtime_provider.py
  x4 from the salvaged commits) plus 4 sibling sites the PR missed:
  cli.py api_mode sync, agent_runtime_helpers.py double-/v1 guard,
  model_normalize.py flat-namespace strip, model_switch.py base_url
  normalization.
- Responses transport: alias OpenCode-reserved function names
  (web_search, search_files -> hermes_*) on the wire and map them back on
  dispatch — same pattern as the xAI web_search collision fix. Matches
  family providers and any base_url on opencode.ai. Fixes the HTTP 400
  'custom function name X is reserved' half of #85589.
- Tests: custom-provider routing assertions + 5 new transport alias tests.
2026-08-20 20:21:12 -07:00
Lesnak1 12a3a35521 fix(providers): normalize OpenCode provider matching to case-insensitive and symmetric Zen/Go support 2026-08-20 20:21:12 -07:00
Lesnak1 f9cd51eb66 fix(providers): support custom opencode-go-* provider routing and grok models 2026-08-20 20:21:12 -07:00
Teknium 19c6a11924 feat(opencode): sync Zen/Go catalogs — Ox Alpha stealth model, Grok routing, new free tier
- Add x-preview-f-free (Ox Alpha: free, 1M context, ZDR) plus all newly
  listed Zen models (gpt-5.6 sol/terra/luna, claude-opus-5, gemini-3.7/3.6
  flash + lite, grok-4.6/4.5, muse-spark-1.2, kimi-k3, qwen3.7-max,
  hy3-free, laguna-s-2.1-free, nemotron-3.5-lightning-free,
  muse-spark-1.2-contributor-free) and Go models (gpt-5.6-luna, grok-4.5,
  glm-5.3, qwen3.8-max, hy3, hy3-preview, muse-spark-1.2-contributor).
- Drop delisted north-mini-code-free from Zen.
- Route grok-* on Zen and Go through /v1/responses per the published
  endpoint tables (grok-4.6/4.5/build-0.1 on Zen, grok-4.5 on Go).
- 1M context fallback for x-preview-f (Ox Alpha).
- Refresh hermes setup provider samples for both providers.

Catalogs verified against live GET /zen/v1/models and /zen/go/v1/models
plus https://opencode.ai/docs/zen/ and /docs/go/ endpoint tables (2026-08-20).
2026-08-20 19:41:03 -07:00
Abdulkadir Ateş 5969ea1558 fix(opencode): route Muse Spark through the Responses API
OpenCode Go and Zen serve muse-spark* only on /v1/responses.
Hermes was sending /chat/completions, which returns HTTP 503
with an empty assistant message. Match the published endpoint
table and the existing gpt-* routing.

- Route muse-spark* to codex_responses on opencode-go and opencode-zen
- Add regression assertions next to the gpt-5.6-luna cases
2026-08-20 19:41:03 -07:00
Brooklyn Nicholson 9c0cd1add2 fix(nous): make "thinking off" stick on a cold start
Portal reasoning capabilities were held only in memory, so a process that
had not yet fetched them answered "unknown" — and on that answer the Nous
profile drops the disable rather than risk a 400. A short-lived process
(`hermes -p`, a cron job, a freshly booted gateway) is always in that
state, so every one of those runs silently ignored "thinking off" and
billed the user for reasoning they had turned off.

The parsed catalog is now mirrored to `cache/reasoning_caps.json`, keyed
by the URL it came from, and hydrated on a cold lookup without touching
the network. Every picker and pricing fetch already pulls that same
document, so they seed the mirror for free.

The catalog URL itself now resolves through the same ladder as the rest
of the Nous catalog reads (`NOUS_INFERENCE_BASE_URL` → credential base →
production) instead of being pinned to production, which had a staging
profile deciding the reasoning-mandatory question from prod's answers.
Keying the mirror by URL keeps those deployments apart.
2026-08-19 23:28:14 -05:00
Brooklyn Nicholson 608fa9c7af feat(models): read Nous Portal reasoning capabilities from its catalog
The Portal serves OpenRouter's catalog schema, so the existing parser and
cache-only tri-state contract carry over unchanged. Only the HTTP fetch is
generalized across the two catalogs; each keeps its own cache because they
list different models.

The Portal 403s a catalog read with no User-Agent, so the shared fetch now
sends one.
2026-08-19 22:14:56 -05:00
Teknium f7d90c9410 refactor: single canonical reasoning-effort vocabulary ends the per-vendor clamp drift
The #89503/#70058/#74295/#87279 bug class kept regenerating because every
transport and provider profile hand-rolled its own effort translation map
(9 sites, 4 distinct policies). New agent/reasoning_effort.py is the single
source of truth:

- EFFORT_LADDER: canonical low->high ordering (superset check against
  VALID_REASONING_EFFORTS pinned by test)
- clamp_effort(): one policy — supported passes verbatim, otherwise nearest
  WEAKER supported level (never escalate, never invert the ladder), floor
  when nothing weaker, 'none' never a degradation target, declared
  vendor-documented overrides win, bespoke names pass through
- declared wire vocabularies as data: OpenAI-compat, Codex Responses,
  xAI (4.6/legacy), Actual relays, Kimi K3/K2, TokenHub, GLM-5.2,
  DeepSeek V4, Ollama Cloud, Meta, Solar

Converted sites (all behavior-preserving except noted):
- chat_completions chokepoint, Kimi + TokenHub paths
- codex transport (backend branches now pick a declared set)
- auxiliary_client Responses path
- hermes_cli.models clamp_reasoning_effort_to_supported -> thin wrapper
- plugins: kimi-coding, zai, opencode-zen, deepseek, ollama-cloud,
  meta-ai, upstage, custom (copilot already routes via the wrapper)

Behavior fixes the shared policy surfaces:
- ollama-cloud/opencode-go 'minimal' now degrades to 'low' instead of
  being dropped (drop left the server default = MORE thinking than asked)

New tests: ladder contract (every configurable level is clamped by every
declared wire set; monotonicity across the full ladder for every set).
2026-08-19 19:29:10 -07:00
Jeffrey Quesnelle b5455fdd16 Merge pull request #84645 from rroverin/feat/add-nemotron-lightning-35-model
feat(models): add Nemotron 3.5 Lightning 30B-A3B to NVIDIA NIM picker
2026-08-19 11:17:24 -04:00