294 Commits

Author SHA1 Message Date
teknium1 d54ae02606 fix(pricing): custom-provider /models prices already per-million are no longer inflated 1e6x
`_extract_pricing`'s generic path copied catalog values verbatim, while
usage_pricing unconditionally applies OpenRouter's per-token convention and
multiplies by 1e6. A provider quoting USD per 1M tokens (Neosantara 0.6/M,
Crof cost.input 0.04/M) or declaring `unit: per_1m_tokens` therefore priced
at $600,000/M and corrupted estimated_cost_usd in state.db and every cost
report summing across providers.

Normalize at the producer, where Novita/DeepInfra unit handling already
lives: an explicit `unit` beside the rates wins (per_token / per_1k_tokens /
per_1m_tokens); without one, a token rate at or above $0.001 per token
($1,000/MTok - no real model) can only be a per-million quote. Output keeps
the per-token-string contract, so the consumer is untouched; `request` fees
and per-token catalogs pass through unchanged.

The $0.001/token magnitude threshold is the one proposed in #34263 by
@Bartok9 (earliest fix); #112036 by @kvnloo proposed the same heuristic at
the consumer.

Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-15 18:20:51 -07:00
Teknium 225d53f953 Port from cline/cline#12876: classify Anthropic output-cap errors 2026-09-15 03:54:58 -07:00
Teknium 2aad4035a5 Port pattern from zed-industries/zed#62729: request ungated Codex model catalog (clean-room)
The ChatGPT Codex models endpoint interprets client_version as a Codex
CLI compatibility version and filters out any model whose
minimal_client_version is newer than the value sent. Hermes hardcoded
client_version=1.0.0 at both catalog request sites, so model visibility
was accidentally coupled to a version scheme Hermes doesn't follow —
future models gated behind a higher minimal version would silently
vanish from the account catalog.

The backend accepts the exact sentinel 0.0.0 as an ungated request
returning the complete account catalog (verified live: 0.0.0 and
current versions return identical model sets today, while omitting the
parameter is HTTP 400 and out-of-sequence values like 0.0.1 return no
models). Both request sites (hermes_cli/codex_models.py and the
context-length probe in agent/model_metadata.py) now share one
CODEX_UNGATED_CLIENT_VERSION constant.

Clean-room port of the observed behavior in zed-industries/zed#62729;
no GPL code translated.
2026-09-13 20:50:10 -07:00
teknium1 3f59b5c594 refactor(agent): /context breakdown and native-compaction retention use the canonical token estimator
context_breakdown._chars_to_tokens and native_compaction._approx_tokens did raw
chars//4, under-counting CJK/Cyrillic by 2-4x next to the conversation slice
that already used estimate_tokens_rough — the /context pie chart mixed two
estimators. Both now call the canonical. The four private `= 4` ratio
constants import one CHARS_PER_TOKEN from agent/model_metadata.py. Estimates
only feed UI and budgets; no prompt or message bytes change.
2026-09-13 05:09:43 -07:00
teknium1 b309a713a8 fix(model-metadata): trim dated docstring, pin output-cap vs context-limit invariant
Follow-up to the #106769 pick: the 8-line dated bugfix narrative in
parse_context_limit_from_error becomes a 2-line WHY, and one invariant
test pins the contract: an output-cap message parses to None as a
context limit, 16384 as the available output tokens, and classifies as
an output-cap error, while a genuine "maximum context length" message
still parses. Red on origin/main (returned 16384 as the context limit).
2026-09-12 07:46:36 -07:00
Francesco Carucci 9ee9e920c6 patch(model-metadata): stop parsing output-limit errors as context limits
Switchyard's "max_tokens cannot exceed the configured model output limit
of 16384" matched the generic "limit ... of N" pattern and was cached as
the CONTEXT window. switchyard/fallback's real context is >117K, so every
later request was needlessly clamped to 16K. Bail out when a message
mentions an output limit and never says "context", and add an explicit
pattern so it is classified as an output-cap error instead. Genuine
context messages always say "context", so they are unaffected.

(cherry picked from commit 981923b31ea9f68534d1f8c239a69042beca731f)
2026-09-12 07:46:36 -07:00
Teknium 4b8c01f691 fix(multiplex): key tool-side and agent-side memos by profile home
Camofox VNC one-shot, computer-use aux-vision verdict, tirith binary path, MCP
discovery lock path, remote-backend probe text, learned image token costs,
auxiliary per-task semaphores and the custom-endpoint /models memo all held one
profile's config-derived value for the whole process. The skill-sync debounce
Timer ran with empty ContextVars, so a secondary's write pushed as the launch
profile (and cancelled its pending push).

Each memo is now keyed by hermes_home_key() (or credential fingerprint for the
per-key catalog) under an override; the timer is per home and runs its callback
inside the scheduling turn's copied context. Unscoped slots are unchanged.
2026-09-12 01:35:05 -07:00
jakobdylanc 284ee03df1 fix(models): resolve context length for OpenRouter :nitro/:floor routing variants
`:nitro`, `:floor`, `:exacto`, and `:online` are request-time routing
modifiers, not catalog models — OpenRouter's /models lists only the base
id, and a variant runs the same model with the same context window.

`get_model_context_length()` keyed every lookup on the full suffixed id,
so each one missed and the resolver fell through to a generic family
default or the 256K fallback:

  openai/gpt-5.5:nitro           -> 256K   (real 1.05M)
  x-ai/grok-4.6:nitro            -> 131K   (generic "grok" catch-all)
  anthropic/claude-opus-4.6:nitro-> 200K   (generic "claude" catch-all)

The window silently shrank, triggering early compression and a wrong
/usage readout. f14059fa fixed the sibling half of this bug class in
/model validation; this fixes the metadata half.

Strip a recognized variant suffix for LOOKUP only, keeping the suffixed
id on the wire so the routing opt-in survives. Applied after the explicit
config overrides (steps 0b/0c) so a user-pinned value still wins, and
before every cache/catalog lookup. Gated on the request actually routing
through OpenRouter, so a local Ollama `model:tag` is untouched.

`:free`/`:batch`/`:thinking` are deliberately excluded — those ARE
distinct catalog SKUs with their own windows, so stripping them would
report the wrong number.

The suffix set and base-id split move to hermes_constants (import-safe,
dependency-free) so the metadata layer shares one definition with
hermes_cli.models instead of duplicating it.
2026-09-11 16:50:20 -07:00
Teknium 759024bdff fix(deepseek): honor Flash 1M window leftovers and native vision
Users on native DeepSeek were told to pin model.context_length and
model.supports_vision in config.yaml. That is the wrong layer: the 1M
window is already in DEFAULT_CONTEXT_LENGTHS, and a global
supports_vision pin would also mark text-only deepseek-v4-pro as
multimodal.

Two catalog gaps still produced the reported symptoms:

- A leftover context_length_cache.yaml entry of 128K (the old
  ``deepseek`` catch-all) outlived the 1M catalog keys because
  deepseek-flash was missing from _PRE_CATALOG_STALE_KEYS.
- When models.dev is empty/cold, Flash has no capability record, so
  image routing falls through to lossy text. Vendor docs (2026-09-10)
  mark deepseek-flash as vision-capable and deepseek-v4-pro as not.

Discard those 128K leftovers, fill Flash (and retired Flash aliases)
via _BUILTIN_MODEL_METADATA, and leave Pro catalog-only.
2026-09-11 12:09:41 -07:00
Teknium 073c57872a feat(models): DeepSeek V4.1 Flash on the Nous Portal and OpenRouter pickers
Add deepseek/deepseek-v4.1-flash to OPENROUTER_MODELS (Nous list derives from it),
regenerate the docs manifest, and give the slug its own 1M context entry and 600s
reasoning-stale floor — the longest-key-first scan otherwise lands the new slug on
the 128K `deepseek` catch-all and no floor. Live probed on both routes: echoed
model matches, usage.cost billed.
2026-09-10 09:22:31 -07:00
YipTszkwan 8435a3ae00 fix(deepseek): recognise the version-less deepseek-flash model id
DeepSeek's 2026-09 Flash refresh introduced a version-less canonical id:
GET /v1/models now returns `deepseek-flash` (alongside `deepseek-v4-pro`), the
API accepts it directly, and the older `deepseek-v4-flash` is server-side
aliased onto it. Every DeepSeek model-id gate in Hermes keys off the
`deepseek-v<N>` prefix, so the new id silently missed all four:

* DeepSeekProfile.build_api_kwargs_extras classified it as non-thinking and
  omitted `extra_body.thinking`. The server then defaults to thinking-on, so
  the user's thinking toggle and `reasoning_effort` were quietly ignored.
* `_normalize_for_deepseek` folded it onto `deepseek-v4-flash` (it misses the
  V-series regex), so the id a user picked never reached the wire and the
  config stored a different model than the picker advertised.
* `DEFAULT_CONTEXT_LENGTHS` fell through to the 128K `deepseek` catch-all
  instead of the real 1M window, capping the model at an eighth of its
  context before compaction kicked in.
* `_REASONING_STALE_TIMEOUT_FLOORS` had no entry, leaving the stale-stream
  detector at its 180s default instead of the 600s reasoning-model floor.

Verified live against api.deepseek.com: `deepseek-flash` answers 200 with
`model: deepseek-flash`, accepts image input (the refresh folds vision into
the Flash model), and the in-between id `deepseek-v4.1-flash` is rejected
with "The supported API model names are deepseek-flash, deepseek-v4-pro".

Adds the id to all four gates plus regression coverage for each site.
2026-09-10 02:44:26 -07:00
Teknium 4baf4269ff fix(metadata): GLM context windows from the live catalog; hyphenated slugs, entry-level overrides, aux inherits main pin
Every GLM slug that missed a DEFAULT_CONTEXT_LENGTHS key fell to the "glm" 202,752 catch-all, so
compression fired at ~20% of the real window and the aux feasibility check auto-lowered the session
threshold to that number (the 202,752 in the Coatue compression report).

- Catalog: GLM entries from Nous + OpenRouter /v1/models (2026-09-09): 5.3 / 5.3-flash 1,310,720
  (:batch/:US 1,048,576), 5.2 1,048,576, 5 / 5.1 / 4.7 / 4.6 204,800, *-turbo / 4.7-flash 202,752.
- _catalog_key_matches: version separators normalised on both sides so relay slugs like z-ai-glm-5-3
  hit glm-5.3 (#97398; approach from PR #97412 by @shellybotmoyer, re-implemented on current main).
- get_custom_provider_context_length: entry-level custom_providers[].context_length backs every model the
  entry serves when no per-model override exists, so /model switch stops dropping it (#98387; from
  PR #98396 by @liuhao1024, re-implemented on the config_providers sibling).
- check_compression_model_feasibility: when aux compression is the main model on the main route, reuse the
  main model's resolved window instead of re-resolving without its pin (#89500 mechanism 2, #45519).

Closes #97398, #98387, #89500, #87825, #97595, #97820 (suffix strip already on main; catalog values now match).
2026-09-10 01:08:18 -07:00
teknium1 8d93081971 fix(desktop): stop flagging local/LAN auxiliary pins as stale
An aux task pinned to a private endpoint via `base_url` (a home Ollama
box at `byron.local`, a LAN IP, localhost) is the intended per-task
endpoint feature and can never bill a provider. The Settings → Model
banner still counted it as "still run on openai" forever and offered
"Reset all to main", which would wipe the working local setup; the
post-switch `stale_aux` report had the same blind spot; and the aux row
never showed the `base_url` the backend already sends, so the pin was
indistinguishable from a paid-provider pin.

- `GET /api/model/auxiliary` now stamps each task with `local_endpoint`,
  the verdict of the one canonical classifier
  (`agent/model_metadata.py::is_local_endpoint`) — no TS mirror of the
  private-range rules, so frontend and runtime cannot drift.
- Desktop: the persistent banner filter is the pure
  `staleAuxAssignments()` and skips `local_endpoint` pins; the pinned row
  appends ` · <base_url>` when one is set.
- `_stale_aux_pins` (post-switch report) skips local pins the same way.
- `is_local_endpoint`: `*.local` (RFC 6762 mDNS) now counts as local, and
  IPv6 literals no longer ride the "no dots ⇒ unqualified host" rule, so
  a global-scope address (`2607:f8b0::1`) is not local while `::1`,
  ULA and link-local still are via the `ipaddress` scope checks.

Slim redo of #106236 (@webtecnica) and #106234 (@huklaa), which fixed the
same symptom with a client-side classifier copy; the bug class, row
display and mDNS/IPv6 classifier corrections are theirs.

Refs #106228

Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-09 10:33:00 -07:00
Michael Steuer 17c7be1485 feat(models): preserve live-verified Astra 900K opt-in from #103132
Retain the two context-variant metadata additions by Michael Steuer. Keep the dedicated Astra reasoning contract already present in #103057 rather than replacing it with the GPT-5.6 vocabulary.

(cherry picked from commit add3a4fa6f31ea3f5fdad701a840d4f8530eb30e)
(cherry picked from commit 1672f12c260260ea38ed636528c02aee734d62b1)
2026-09-07 21:43:54 +05:30
Eva 3825d25191 fix(openai): cap Astra Codex OAuth fallback
(cherry picked from commit bb4156c3881cbaf20736f0fcfd6c3cc5c861a6fb)
2026-09-07 21:43:54 +05:30
Teknium 7d44fe9c74 fix: capability probes send minted credentials instead of callable representations
Extend #104477 to the native thinking, vision, metadata, and local header paths identified by #87641. Materialize only at probe boundaries; leave the chat callable and cache ownership untouched. Local-wire A/B: thinking and vision show requests change from 403 to 200, while static credentials and callable chat retain success. Target suites queued.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-07 08:08:04 -07:00
Teknium be58c276ee feat(compression): per-image token cost learned from the provider's own usage (#70328, supersedes #70463)
A flat per-image constant (1500 in the trigger estimator, 1600 in the tail-budget walk) is wrong in
both directions: a screenshot costs ~1,100 tokens on one provider and 4,000+ on a local mmproj
model. In a GUI loop on a 64K window the estimate sat at ~20K while the real prompt passed 80K,
so compaction never fired and the provider rejected every request (#70328).

The provider prices every image exactly on the request that carries it, so the cost is
observable from usage alone, with no vendor formula: with a fresh usage anchor, the residual
between the next real prompt_tokens and anchor + text-only delta is the price of the N images
that delta introduced.

- agent/image_token_cost.py: calibrate_from_usage() runs in record_response_usage before the new
  anchor is captured; the learned value (EMA, plausibility-banded) is kept per model@host in
  ~/.hermes/cache/image_token_costs.json and bound per turn through a ContextVar.
- estimate_messages_tokens_rough, _content_length_for_budget (tail walk) and gateway hygiene all
  read the same bound value, so trigger and walk agree; the per-message memo now caches text
  tokens and image COUNT so a recalibration re-prices cached rows.
- One flat default (1500) remains only until the first vision turn; the duplicate 1600 is gone.

evals/token_accounting/ab_image_cost_calibration.py (real AIAgent, fake provider pricing images
at 4,000, one screenshot per turn, 64K window): main learns nothing (1500) and the tail walk
under-prices its own protected tail by 56.5%; this branch learns 4,374 after one vision turn
and the walk's error is +8.5%.

Reporter and first-fix credit: @JonthanaHanh (#70328, #70463).
2026-09-06 14:19:42 -07:00
Teknium 562e6e4824 fix(compression): opaque encrypted_content costs 0 in every local token estimate
Codex Responses reasoning and compaction items carry ciphertext the provider prices by its own
token count, never by bytes; a single native compaction checkpoint is ~5M chars, which the
bytes/4 estimator turned into ~1.29M "tokens" against a 204K threshold (#100611). #104192 deferred
that decision for one request; this removes the mis-pricing at the source so the preflight
estimator and the tail-budget walk agree (a mismatched size class protects blob-heavy rows as
"small" and compaction re-fires). Only real usage prices these items, and the usage anchor carries
that price forward.

evals/native_compaction/ab_checkpoint_preflight.py: preflight estimate after checkpoint
1,292,413 -> 58 rough tokens; the over-threshold negative arm still compresses.
2026-09-06 13:21:17 -07:00
686f6c61 c0aaa238f6 feat(compression): usage anchor survives DB reloads and process restarts (salvage #99585)
The usage anchor (real usage.prompt_tokens + delta estimate of what was appended since)
identified the priced transcript by id() of the last message, so it was None on EVERY
gateway turn (history is re-read from the DB each turn) and in every fresh process
(--resume, desktop per-turn serve). Those are exactly the surfaces where the bytes/4
estimate then fired local compression against payloads the provider priced far under
threshold (#99421, #104462).

- agent/usage_anchor.py owns the anchor: content fingerprint instead of id(), persisted on
  the session row (model_config._usage_anchor) via set_usage_anchor(), restored on the first
  resumed turn while the durable transcript still matches, cleared with the row on
  compaction / codex-native rewrite / session reset.
- Callers repointed from model_metadata (the compat table follows).

Design and persistence slot from #99585 by @686f6c61; re-authored against the Sep 2026
layout (the branch predates the model_metadata / agent_init split).
2026-09-06 13:21:17 -07:00
Teknium f159e581c7 feat(models): add GPT-6 Astra + Astra Pro with fast/flex speed tiers to Nous Portal and OpenRouter
Six slugs land in the nous and openrouter curated lists, above the gpt-5.6 line:
openai/gpt-6-astra{,-fast,-flex} and openai/gpt-6-astra-pro{,-fast,-flex}.

Nous Portal serves the tiers as distinct slugs (verified live: each echoes its id, service_tier
default/priority/flex, cost 1x/2x/0.5x). OpenRouter serves them as ENDPOINTS of the base model
(tags openai/fast, openai/flex) and silently routes an unknown suffix to the standard tier at
standard price, so the OpenRouter profile rewrites a tier slug to its base wire model and pins
provider.only to that tier's endpoints (OPENROUTER_ENDPOINT_PINS). The base slug is pinned to
openai/azure/azure-us so default routing never lands on a flex or fast endpoint.

Provider-agnostic metadata: one DEFAULT_CONTEXT_LENGTHS entry (gpt-6-astra: 1,050,000, live on
OpenRouter for both models; substring-matches -pro and the tier suffixes). Pricing is skipped:
both routes bill via official_models_api. Reasoning floor not added (no evidence of long thinks).
2026-09-04 15:58:59 -07:00
Teknium 7a33369e81 simplify(compat): interrupt — drop _ThreadAwareEventProxy/_interrupt_event legacy alias, repoint 2 test files
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
2026-09-03 14:00:59 -07:00
Teknium c93ace77c2 simplify(compat): config/runtime_provider/plugins/commands/secrets_cli/kanban — drop 96 re-exports (incl. PEP 562 facades) + 3 aliases (get_pre_tool_call_directive/_block_message, get_telegram_handler_factories), repoint 56 callers + 50 test files 2026-09-03 14:00:17 -07:00
Teknium e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium 0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Steve-prog001 779aecb62b fix(context): resolve commandcode models via live /models
commandcode (api.commandcode.ai) exposes authoritative
context_length via /models (muse-spark 1M, etc.) but as a
known provider it skipped the custom-endpoint probe at step 2
and has no models.dev entry, so every model fell through to the
256K DEFAULT_FALLBACK. Add a provider-aware branch mirroring
gmi/nous to resolve via _resolve_endpoint_context_length.

Fixes GOAT docs vs status-bar mismatch: muse-spark 1M was shown
as 256K.
2026-09-03 00:58:16 -07:00
GTHell bb8f4afa46 fix(context): add muse-spark 1M fallback (zen/GO SG showed 256k)
Muse Spark 1.2 family (api.meta.ai) ships 1M context (models.dev
opencode/muse-spark-1.2 = 1048576, meta/muse-spark-1.2 = 1048576).

Zen/GO SG /v1/models only returns id (no limit.context), and
models.dev lookup via opencode was missing a hardcoded fallback, so
get_model_context_length fell back to DEFAULT_FALLBACK_CONTEXT=256k.
Banner showed Context: 256,000 for both zen and router-sg lanes.

Add longest-prefix entries 'muse-spark' and 'muse' = 1_048_576 so
all variants (1.1, 1.2, contributor, contributor-free) resolve to 1M
without network.
2026-09-03 00:58:16 -07:00
mr-r0b0t cfa7e72c9e fix(models): correct contributor guard, 1M context, docs for muse-spark-1.3
- model_data_policy_guard: name the triggering -contributor model instead
  of hardcoded 1.2; per-version verified price tables (1.3 standard
  $1.25/$4.25 via OpenRouter live metadata; cached figures 1.2-only)
- model_metadata: muse-spark-1.3 + muse-spark family at 1048576 (OpenRouter
  verified 2026-09-02) with pre-catalog stale-cache keys so 256K-fallback
  sessions self-heal
- docs: contributor-tier notes cover 1.2 + 1.3
- tests: 1.3 guard regression, muse stale-cache guard, live-catalog mirror
  gains 1.3-contributor-free (confirmed on live relay)

143 tests pass (guard, selection guards, opencode catalog, model_metadata);
ruff clean.
2026-09-03 00:57:55 -07:00
Teknium f42fc3c5c7 refactor(agent): compact token estimation helpers (shadow, fingerprint, tools cache) 2026-09-02 21:49:03 -07:00
Teknium 11e0f3862d refactor(agent): table-drive cache validation drops and sourced provider resolvers in get_model_context_length 2026-09-02 21:44:44 -07:00
Teknium ea8b3250f4 refactor(agent): dispatch local context probes by server type; unify Ollama /api/show calls 2026-09-02 21:36:24 -07:00
Teknium 30ffd52f84 refactor(agent): compact probe/cache helpers and payload parsers in model_metadata 2026-09-02 21:27:46 -07:00
Teknium 1e8e2a7328 refactor(agent): model_metadata — restore _is_known_provider_base_url (test-patched seam) 2026-09-02 19:41:04 -07:00
Teknium 0e5c0323fd refactor(agent): model_metadata/models_dev — fold codex variant helpers, compact tables and comments, squeeze blanks 2026-09-02 19:18:35 -07:00
Teknium 6a4582ce33 refactor(agent): model_metadata — inline zero-ref wrappers, pack data tables, compact docstrings/comments by hand 2026-09-02 19:03:38 -07:00
Teknium 22df91d65d refactor(agent): model_metadata — table-drive server-type waterfall, unify cache-drop logging, compact tables/signatures 2026-09-02 18:49:56 -07:00
Teknium d45621c032 refactor(agent): model_metadata — extract get_model_context_length phases, unify probe memo/blackhole/ollama helpers 2026-09-02 18:43:54 -07:00
kshitijk4poor a1d5a976b3 fix(agent): never floor an anchored pressure figure; keep the estimator total on lone surrogates
Follow-ups from review of the two salvaged #87490 commits:

- _pressure_with_real_floor now applies only on the rough fallback branch.
  A valid usage anchor is provider-exact and wins as-is: on MoA turns the
  anchor deliberately uses the pre-fold aggregator usage while
  last_real_prompt_tokens holds the folded figure, so flooring the anchored
  value would re-add fan-out tokens the anchor exists to exclude. Docstring
  rewritten to describe the real path split (anchor since d3a1c46510).
- estimate_tokens_rough: encode with errors="replace". main's estimator
  never raised; text.encode() on a lone surrogate (routine in tool output,
  see message_sanitization) raised UnicodeEncodeError and would abort a
  turn where main produced a slightly-off number.
- Record the cl100k/o200k/Qwen2.5 calibration for the bytes/4 rule.
- tests: accented Latin within +10% of the ASCII rule; mixed Cyrillic/ASCII
  counts ASCII at one byte; lone surrogates don't raise; anchored pressure
  is never floored (wiring shape).
2026-09-03 03:09:06 +05:30
Darafei Praliaskouski f07b6ff426 fix(agent): count non-CJK sparse text by UTF-8 bytes in the rough estimator
The ~4 chars/token rule is calibrated for ASCII; Cyrillic, Greek, Arabic and
similar 2-byte scripts tokenize at ~2-3 chars/token, so chars/4 under-counts
them ~2x and the pre-flight pressure figure trails real usage by tens of
percent on non-English sessions. Counting UTF-8 BYTES at ~4/token uses the
encoding width itself as the corrective: ASCII is unchanged (1 byte/char),
2-byte scripts count at chars/2, and the CJK dense path keeps its explicit
~1 token/char rule with the sparse remainder byte-counted. The ASCII
isascii() O(1) fast path is preserved; the non-ASCII paths add a single
C-level encode over text that was already being regex-scanned.

Complements the last-real-prompt floor: the floor catches sessions that are
already at the ceiling, this keeps the estimate from lagging in the first
place.
2026-09-03 03:09:06 +05:30
Teknium 2273f2bd20 refactor(agent/models): unify model_metadata TTL disk memos
_ttl_memo_get/_ttl_memo_put back both the local-probe and endpoint-metadata
disk caches (file layout, key order and log strings verified identical).
2026-09-02 13:52:52 -07:00
Teknium 684e8322ff refactor(agent/models): extract LM Studio and llama.cpp /props branches from fetch_endpoint_model_metadata
_lmstudio_native_models and _apply_llamacpp_props; the router-mode child probe
shares one /v1/props -> /props fallback helper. HTTP call sequence verified
identical old vs new across 10 fake-endpoint scenarios.
2026-09-02 13:52:52 -07:00
Teknium 771045655b refactor(agent/models): table-driven credits header parsing, unified account-usage fetchers
- credits_tracker: _MICROS_FIELDS/_USD_FIELDS/_BOOL_FIELDS drive parse_credits_headers;
  _sticky_notice() replaces repeated AgentNotice literals; collapsed defensive layers
- account_usage: _USAGE_FETCHERS dispatch table replaces provider if/elif; shared
  fetch/parse helpers across the codex/anthropic/openrouter fetchers; drop dead
  _resolve_codex_usage_url
- billing_links, fast_mode: compacted docstrings, collapsed if/return chains
- model_metadata: fold effective_provider inference into one expression
Notice keys, header names, URLs, log strings and dataclass defaults unchanged.
2026-09-02 13:52:52 -07:00
Teknium d465b45f38 refactor(agent/models): compact model_metadata comment essays to their invariants
Every non-obvious rule, ordering, failure mode and WHY is kept (1-3 lines);
issue numbers, dates and narrative dropped. Data tables unchanged.
2026-09-02 13:52:51 -07:00
Teknium 4a65e98f50 refactor(agent/models): extract get_model_context_length steps into helpers
_validate_cached_context_length (step 1), _resolve_bedrock_context_length (1b),
_resolve_custom_endpoint_context_length (2-3); resolution order and every log
message unchanged. Compact the per-step comment essays to their invariants.
2026-09-02 13:52:51 -07:00
Teknium f3a46eb5d0 refactor(agent/models): dedupe model_metadata probe helpers
- _server_root / _ollama_show_context / _longest_key_match / _probe_local_context_length /
  _endpoint_model_entry replace 5 copies of the same local-probe and lookup bodies
- detect_local_server_type waterfall as an ordered (name, probe) table
- output-cap error classification via phrase-group tables shared by
  is_output_cap_error and parse_available_output_tokens_from_error
- endpoint-scoped context overrides as a data table
- drop dead _fetch_codex_oauth_context_lengths, _resolve_codex_oauth_context_length,
  _estimate_message_chars (zero refs repo-wide)
Verified with a differential harness (old vs new module, ~700 pure-function probes, 0 mismatches).
2026-09-02 13:52:51 -07:00
Teknium 32fe129324 perf(bot-mode): cold DM hops skip the live /models probe; relay replies land within 250ms
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.

- model_metadata: memoize successful remote /models probes on disk
  (cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
  in-memory cache, so authority semantics are unchanged (reconciliation
  still lands within 5 minutes) but the answer is shared across
  processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
  250ms instead of every 2s — up to 2s of dead air on every relayed reply.

Nothing here changes turn ordering: DMs and group rounds stay serial.

Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
2026-09-02 03:42:01 -07:00
Teknium c0495c6bce fix(cli): context meter no longer sawtooths on reasoning models — show durable transcript, not last-request replay
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.

- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
  provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
  delta estimate (stale reasoning excluded on all but the newest assistant
  message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
  figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.

Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
2026-09-01 15:34:03 -07:00
emozilla 43e67d872f feat: local models — managed llama.cpp runtime with one-click desktop setup
Run models locally as a first-class provider. The CLI grows a managed
llama.cpp runtime (engine install, model download, server supervision);
the desktop app grows the full setup and management story on top of it.
GUI surfaces ship behind the desktop --local launch flag (hermes desktop
--local, or the flag on the packaged app); backend routes and the CLI
are always live.

Runtime (hermes_cli/local_runtime/):
- curated GGUF catalog with per-machine variant selection: hardware
  probe (VRAM/RAM/UMA), fit planning with spill accounting, quant choice
  by context window
- derived recommendation: quality-ranked picks gated by a predicted
  decode-speed floor, bandwidth-aware on unified memory; the decision
  table is pinned as a test (pick AND reason per memory class), and the
  Recommended badge explains its pick in a tooltip fed by the resolver's
  actual branch
- engine install + model download with resumable split parts, cumulative
  plan-level progress, and staged-model integrity (a split GGUF counts
  only when every part is present)
- server supervision: spawn/adopt/stop, router mode with per-model load
  progress relayed over SSE, abandoned-request cleanup

Desktop:
- Settings -> Providers -> Local models: one-click quickstart (install
  engine, download the recommended model, boot) plus per-model download/
  activate/eject, fit-ranked catalog with context pills
- model pickers (composer dropdown + Cmd+K) show staged local models,
  in-flight downloads as live progress rows, and load-into-memory bars
- local-setup campaign tip for eligible hardware; System resources
  statusbar widget (GPU/VRAM/RAM); in-chat load progress during sends
- friendly dead-server errors, and failed agent builds retry on the next
  send instead of wedging the session

Co-developed with NVIDIA field feedback on RTX 5090 and DGX Spark.
2026-09-01 16:01:53 -04:00
Teknium 58f5b1e277 fix(model_metadata): parse Google's 'supports up to N' context-limit phrasing
Google Gemini/Gemma overflow errors read 'Unable to submit request because
the input token count is 32825 but model only supports up to 32768'.
parse_context_limit_from_error had no pattern for the 'supports up to N'
phrasing, so overflow recovery kept the wrong window and burned its retry
attempts instead of recalibrating to the provider-reported limit.

Add the anchored pattern (limit follows 'supports up to'; the larger input
count before it is never captured) plus regression tests covering the exact
message and the get_context_length_from_provider_error recalibration path.

Reported by @Artemonim in #57275 (residual claim 5).
2026-08-31 12:20:02 -07:00
Teknium 64cc87e668 fix(compression): keep estimate seam positional-compatible for monkeypatched estimators
Test seams and plugin engines monkeypatch estimate_messages_tokens_rough with (messages)-only signatures; route callers only pass the charge_stale_thinking kwarg on the False path.
2026-08-30 20:40:43 -07:00
Teknium 452f6b7de2 fix(compression): route-aware stale-thinking charge parity between compaction trigger and tail walks (#84371)
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).

Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.

Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
2026-08-30 20:40:43 -07:00