Commit Graph

7 Commits

Author SHA1 Message Date
Teknium 77743eac8a refactor(agent/models): compact pricing snapshot, billing/subscription views, reasoning helpers
- usage_pricing: _snap() builder for official-docs pricing entries (table values identical, verified by dump), shared source/version dicts, drop dead DEFAULT_PRICING
- models_dev: _registry_models/_iter_model_entries/_extract_limit helpers replace repeated registry walking; drop dead ModelInfo.format_cost
- billing_view/subscription_view: OrgRoleCapability mixin replaces duplicated is_admin/can_change_plan; shared fetch_portal_state/parse_org_fields
- reasoning_effort/timeouts/summaries, thinking_timeout_guidance, portal_tags: dispatch tables and compacted comment essays; drop dead CODEX_RESPONSES_EFFORTS alias and _match_any
2026-09-02 13:52:51 -07:00
kshitijk4poor b954547e72 fix(nebius): route effort through canonical clamp_effort — hand-rolled map inverted the ladder
Review finding on salvaged #28253: the hand-rolled mapping sent
ultra -> medium while xhigh -> high (stronger request, weaker wire
value). Declare NEBIUS_EFFORTS in agent/reasoning_effort.py and use
clamp_effort like the zai/kimi/tokenhub call sites; disable detection
stays ahead of the clamp since clamp_effort('none', ...) returns the
floor, not off. Adds a monotonicity regression test.
2026-08-29 20:39:44 +05:30
Teknium 30f9955a44 fix(zai): GLM-5.3 low/medium reasoning effort reaches the wire instead of clamping to high
GLM-5.3 accepts a graded low/medium/high/max reasoning_effort scale
(verified live in #91789: monotonic reasoning-token scaling, no 400s),
but the effort mapper reused GLM-5.2's two-level vocabulary, silently
rewriting low/medium to high. Adds GLM53_EFFORTS/GLM53_OVERRIDES and a
per-model vocabulary pick in the zai plugin; 5.2 keeps its high/max
clamp. Closes #91789. Also covers the gap noted when closing #86947
(credit @santhanakrishnan-d and @terje1965 for the graded-scale finding).
2026-08-21 14:58:37 -07:00
Teknium d4d04098a5 fix: Ox Alpha reasoning effort reaches the wire clamped — shared across zen and free providers
Widens the salvaged #91323 fix (@vinsew): the effort vocabulary moves to
agent.reasoning_effort (OX_ALPHA_EFFORTS/OVERRIDES, the declared-policy
home every other model vocabulary lives in), and the translation is
shared between the opencode-zen profile and the keyless opencode-free
profile — Ox Alpha is reachable through both, and the free profile
previously dropped effort entirely.

Live-verified: medium clamps to low (raw medium 400s: 'This model always
engages in thinking... use low, high, or max'), xhigh rounds to max, and
full agent turns with effort=medium complete on BOTH providers.
2026-08-21 03:04:06 -07:00
Teknium 4999af5cc1 fix: K3 plan-variant slugs (k3-256k) now get K3's effort vocabulary
kimi_supported_efforts() used exact/prefix matching and missed Kimi
Coding plan variants like k3-256k, which fell back to the K2-era
low/medium/high set and mistranslated efforts on a K3 wire. Replaced
with the boundary-token regex from #76427 (credit @ruizanthony), which
matches k3/k3-256k/kimi-k3* without matching kimi-k2.6 or mk3000.
2026-08-19 22:56:04 -07:00
Teknium d0573880a8 fix: Codex Responses effort vocabulary is now per-model — gpt-5.5 no longer 400s on 'max' (#68365 confirmed live)
Live probes against api.openai.com/v1/responses (Aug 2026):
- gpt-5.6: accepts none/low/medium/high/xhigh/max; rejects minimal, ultra
- gpt-5.5: accepts none/low/medium/high/xhigh; rejects max ('Unsupported
  value'), minimal, ultra

So #68365's premise was half right: 'max' does 400 — but only on pre-5.6
models; blanket-clamping max->xhigh on gpt-5.6 (its fix) would have capped
the one model that supports max. The declared-vocabulary design absorbs
this as data: codex_supported_efforts(model) picks CODEX_GPT56_EFFORTS or
CODEX_LEGACY_EFFORTS, and the shared clamp does the rest. Both the main
Codex transport and the auxiliary client's Responses path use it.

Wire outcomes: ultra -> max on gpt-5.6, ultra/max -> xhigh on gpt-5.5/o5,
minimal -> low everywhere.
2026-08-19 22:37:46 -07:00
Teknium f7d90c9410 refactor: single canonical reasoning-effort vocabulary ends the per-vendor clamp drift
The #89503/#70058/#74295/#87279 bug class kept regenerating because every
transport and provider profile hand-rolled its own effort translation map
(9 sites, 4 distinct policies). New agent/reasoning_effort.py is the single
source of truth:

- EFFORT_LADDER: canonical low->high ordering (superset check against
  VALID_REASONING_EFFORTS pinned by test)
- clamp_effort(): one policy — supported passes verbatim, otherwise nearest
  WEAKER supported level (never escalate, never invert the ladder), floor
  when nothing weaker, 'none' never a degradation target, declared
  vendor-documented overrides win, bespoke names pass through
- declared wire vocabularies as data: OpenAI-compat, Codex Responses,
  xAI (4.6/legacy), Actual relays, Kimi K3/K2, TokenHub, GLM-5.2,
  DeepSeek V4, Ollama Cloud, Meta, Solar

Converted sites (all behavior-preserving except noted):
- chat_completions chokepoint, Kimi + TokenHub paths
- codex transport (backend branches now pick a declared set)
- auxiliary_client Responses path
- hermes_cli.models clamp_reasoning_effort_to_supported -> thin wrapper
- plugins: kimi-coding, zai, opencode-zen, deepseek, ollama-cloud,
  meta-ai, upstage, custom (copilot already routes via the wrapper)

Behavior fixes the shared policy surfaces:
- ollama-cloud/opencode-go 'minimal' now degrades to 'low' instead of
  being dropped (drop left the server default = MORE thinking than asked)

New tests: ladder contract (every configurable level is clamped by every
declared wire set; monotonicity across the full ladder for every set).
2026-08-19 19:29:10 -07:00