Commit Graph

14 Commits

Author SHA1 Message Date
Teknium d4d04098a5 fix: Ox Alpha reasoning effort reaches the wire clamped — shared across zen and free providers
Widens the salvaged #91323 fix (@vinsew): the effort vocabulary moves to
agent.reasoning_effort (OX_ALPHA_EFFORTS/OVERRIDES, the declared-policy
home every other model vocabulary lives in), and the translation is
shared between the opencode-zen profile and the keyless opencode-free
profile — Ox Alpha is reachable through both, and the free profile
previously dropped effort entirely.

Live-verified: medium clamps to low (raw medium 400s: 'This model always
engages in thinking... use low, high, or max'), xhigh rounds to max, and
full agent turns with effort=medium complete on BOTH providers.
2026-08-21 03:04:06 -07:00
vinsew 54227416ce fix(opencode): send Ox Alpha reasoning effort through Zen
OpenCode documents x-preview-f-free as accepting low, high, and max reasoning effort on its Zen Chat Completions endpoint. Hermes previously resolved the user's per-model override to max but the plain Zen provider profile discarded it, so successful calls silently ran at the server default.

Introduce an OpenCodeZenProfile scoped only to x-preview-f-free. It forwards the normalized top-level reasoning_effort, maps xhigh to max, preserves server defaults when unset or disabled, and leaves every other Zen model untouched.

Add profile and full transport tests that prove max reaches the outgoing request and that non-target models are unaffected. Also correct the nanoid security-pin comment to match the already-locked 3.3.18 release.
2026-08-21 03:04:06 -07:00
Teknium 19c6a11924 feat(opencode): sync Zen/Go catalogs — Ox Alpha stealth model, Grok routing, new free tier
- Add x-preview-f-free (Ox Alpha: free, 1M context, ZDR) plus all newly
  listed Zen models (gpt-5.6 sol/terra/luna, claude-opus-5, gemini-3.7/3.6
  flash + lite, grok-4.6/4.5, muse-spark-1.2, kimi-k3, qwen3.7-max,
  hy3-free, laguna-s-2.1-free, nemotron-3.5-lightning-free,
  muse-spark-1.2-contributor-free) and Go models (gpt-5.6-luna, grok-4.5,
  glm-5.3, qwen3.8-max, hy3, hy3-preview, muse-spark-1.2-contributor).
- Drop delisted north-mini-code-free from Zen.
- Route grok-* on Zen and Go through /v1/responses per the published
  endpoint tables (grok-4.6/4.5/build-0.1 on Zen, grok-4.5 on Go).
- 1M context fallback for x-preview-f (Ox Alpha).
- Refresh hermes setup provider samples for both providers.

Catalogs verified against live GET /zen/v1/models and /zen/go/v1/models
plus https://opencode.ai/docs/zen/ and /docs/go/ endpoint tables (2026-08-20).
2026-08-20 19:41:03 -07:00
Abdulkadir Ateş 5969ea1558 fix(opencode): route Muse Spark through the Responses API
OpenCode Go and Zen serve muse-spark* only on /v1/responses.
Hermes was sending /chat/completions, which returns HTTP 503
with an empty assistant message. Match the published endpoint
table and the existing gpt-* routing.

- Route muse-spark* to codex_responses on opencode-go and opencode-zen
- Add regression assertions next to the gpt-5.6-luna cases
2026-08-20 19:41:03 -07:00
Teknium f7d90c9410 refactor: single canonical reasoning-effort vocabulary ends the per-vendor clamp drift
The #89503/#70058/#74295/#87279 bug class kept regenerating because every
transport and provider profile hand-rolled its own effort translation map
(9 sites, 4 distinct policies). New agent/reasoning_effort.py is the single
source of truth:

- EFFORT_LADDER: canonical low->high ordering (superset check against
  VALID_REASONING_EFFORTS pinned by test)
- clamp_effort(): one policy — supported passes verbatim, otherwise nearest
  WEAKER supported level (never escalate, never invert the ladder), floor
  when nothing weaker, 'none' never a degradation target, declared
  vendor-documented overrides win, bespoke names pass through
- declared wire vocabularies as data: OpenAI-compat, Codex Responses,
  xAI (4.6/legacy), Actual relays, Kimi K3/K2, TokenHub, GLM-5.2,
  DeepSeek V4, Ollama Cloud, Meta, Solar

Converted sites (all behavior-preserving except noted):
- chat_completions chokepoint, Kimi + TokenHub paths
- codex transport (backend branches now pick a declared set)
- auxiliary_client Responses path
- hermes_cli.models clamp_reasoning_effort_to_supported -> thin wrapper
- plugins: kimi-coding, zai, opencode-zen, deepseek, ollama-cloud,
  meta-ai, upstage, custom (copilot already routes via the wrapper)

Behavior fixes the shared policy surfaces:
- ollama-cloud/opencode-go 'minimal' now degrades to 'low' instead of
  being dropped (drop left the server default = MORE thinking than asked)

New tests: ladder contract (every configurable level is clamped by every
declared wire set; monotonicity across the full ladder for every set).
2026-08-19 19:29:10 -07:00
Adolanium 7fa084f58e fix: send Hermes Agent attribution headers to OpenCode Zen and Go
OpenCode identifies clients by request headers, the same way OpenRouter
does. Our opencode-zen and opencode-go profiles never set any, so every
request went out with the OpenAI SDK default "OpenAI/Python x.y.z"
User-Agent and OpenCode had no way to tell the traffic was Hermes Agent.

Two changes:

- Add HTTP-Referer, X-Title, and a HermesAgent User-Agent to both
  OpenCode profiles through profile.default_headers, the same path
  Fireworks uses. This covers chat_completions, codex_responses,
  auxiliary clients, model switches, and the models catalog fetch.
- Merge the same headers in build_anthropic_client for opencode.ai
  base URLs. The Anthropic Messages route (Claude on Zen, MiniMax and
  Qwen on Go) builds its client there and never sees profile headers.

Verified against the live Go relay with a real key. Both wire formats
return HTTP 200 and the requests now carry X-Title "Hermes Agent",
HTTP-Referer, and User-Agent HermesAgent/0.20.0.
2026-08-13 02:03:40 -07:00
kshitij b3344502f8 chore: map Axmr1 email + update stale Go routing docstring
- Add contributors/emails/Axmr1@users.noreply.github.com for CI
  attribution check (bare noreply format needs explicit mapping)
- Update opencode-zen plugin docstring: Go routing now includes
  GPT → codex_responses and Qwen → anthropic_messages (was stale,
  only listed MiniMax and GLM/Kimi)
2026-08-08 14:47:11 +05:30
Teknium 7550c594ce feat(reasoning): add max and ultra effort levels (#62650) 2026-07-12 00:26:49 -07:00
Teknium a6079dd350 feat(providers): GLM-5.2 native reasoning_effort controls (#58884)
Port from Kilo-Org/kilocode#11555: GLM-5.2 exposes a native
reasoning_effort knob with two enabled levels (high / max) on its
OpenAI-compatible endpoints. Previously the zai profile (direct Z.AI
/api/paas/v4) used the base ProviderProfile and emitted nothing, and the
OpenCode Go profile only handled Kimi K2 / DeepSeek — so a user's effort
preference for GLM-5.2 was silently dropped on both routes.

- zai: ZaiProfile maps effort onto high/max (xhigh/max -> max, lower -> high)
- opencode-go: same mapping for GLM-5.2, alongside existing Kimi/DeepSeek
- alias spellings recognized (glm-5.2 / glm-5-2 / glm-5p2, vendor-prefixed)
- disabled / no effort leaves the server default untouched
2026-07-05 13:48:01 -07:00
teknium1 03392b67d6 fix(opencode-go): gate thinking when reasoning_effort set to avoid HTTP 400
Salvaged from #40429; re-verified on main, tightened, tested.

Co-authored-by: jimjsong <jimjsong@users.noreply.github.com>
2026-06-07 01:24:29 -07:00
teknium1 8cf6b3da9d fix(opencode-go): cap mimo-v2.5-pro max_tokens at 131072
The opencode-go relay defaults max_tokens to 262144 when none is sent,
but Xiami mimo-v2.5-pro only supports 131072 completion tokens — every
request 400s with "max_tokens is too large: 262144" before the agent
can do anything.

Add a get_max_tokens(model) hook on ProviderProfile (default returns
default_max_tokens) so profiles fronting multiple upstreams can vary
the cap per-model. Wire chat_completions transport through the hook.
Override on OpenCodeGoProfile with mimo-v2.5-pro=131072.

Only mimo-v2.5-pro is capped — other opencode-go models (kimi, glm,
qwen, minimax, other mimo variants) unchanged.
2026-05-28 20:49:53 -07:00
teknium1 70aaa774be fix(opencode-go): emit Kimi reasoning_effort, match KimiProfile shape
The Kimi K2 branch added in the prior commit only emitted extra_body.thinking
and dropped reasoning_effort entirely. KimiProfile (api.moonshot.ai/v1) sends
both fields, and OpenCode Go proxies to the same Moonshot backend. Mirror that
shape on the Go path so /reasoning effort actually reaches Kimi.

- low/medium/high pass through verbatim
- xhigh/max clamp to high (Moonshot's max supported value)
- minimal / unknown effort → omit reasoning_effort, keep thinking on
- disabled / no config → unchanged
- DeepSeek branch unchanged
2026-05-23 02:20:28 -07:00
Harish Kukreja 3589960e03 fix(provider): expose OpenCode Go reasoning controls 2026-05-23 02:20:28 -07:00
Teknium 9022804d78 feat(providers): make all 33 providers pluggable under plugins/model-providers/
Every provider profile is now a self-contained plugin under
plugins/model-providers/<name>/, mirroring the plugins/platforms/
pattern established for IRC and Teams. The ProviderProfile ABC
stays in providers/; the per-provider profile data moves out.

- plugins/model-providers/<name>/__init__.py calls register_provider()
- plugins/model-providers/<name>/plugin.yaml declares kind: model-provider
- providers/__init__.py._discover_providers() lazily scans bundled plugins
  then $HERMES_HOME/plugins/model-providers/<name>/ (user override path)
- User plugins with the same name override bundled ones (last-writer-wins
  in register_provider)
- Legacy providers/<name>.py layout still supported for back-compat with
  out-of-tree editable installs
- Hermes PluginManager: new kind=model-provider; skipped like memory
  plugins (providers/ discovery owns them); standalone plugins with
  register_provider+ProviderProfile in their __init__.py auto-coerce to
  this kind (same heuristic as memory providers)
- skip_names extended to include 'model-providers' so the general
  PluginManager doesn't double-scan the category
- 4 new tests in tests/providers/test_plugin_discovery.py covering
  bundled discovery, user override, and general-loader isolation
- Docs updated: website/docs/developer-guide/adding-providers.md,
  provider-runtime.md, providers/README.md, plugins/model-providers/README.md

No API break: auth.py / config.py / doctor.py / models.py / runtime_provider.py /
model_metadata.py / auxiliary_client.py / chat_completions.py / run_agent.py
all still consume providers via get_provider_profile() / list_providers() —
they just now see plugin-discovered entries instead of pkgutil-iterated ones.

Third parties can now drop a single directory into
~/.hermes/plugins/model-providers/<name>/ to add or override an inference
provider without touching the repo.
2026-05-05 13:40:01 -07:00