Commit Graph

96 Commits

Author SHA1 Message Date
teknium1 b309a713a8 fix(model-metadata): trim dated docstring, pin output-cap vs context-limit invariant
Follow-up to the #106769 pick: the 8-line dated bugfix narrative in
parse_context_limit_from_error becomes a 2-line WHY, and one invariant
test pins the contract: an output-cap message parses to None as a
context limit, 16384 as the available output tokens, and classifies as
an output-cap error, while a genuine "maximum context length" message
still parses. Red on origin/main (returned 16384 as the context limit).
2026-09-12 07:46:36 -07:00
Teknium 3a2cfb7871 test(models): trim routing-variant tests to invariants
One relation test per consumer file: a routed id resolves to its base's
metadata while a real :free SKU keeps its own window.
2026-09-11 16:50:20 -07:00
jakobdylanc 284ee03df1 fix(models): resolve context length for OpenRouter :nitro/:floor routing variants
`:nitro`, `:floor`, `:exacto`, and `:online` are request-time routing
modifiers, not catalog models — OpenRouter's /models lists only the base
id, and a variant runs the same model with the same context window.

`get_model_context_length()` keyed every lookup on the full suffixed id,
so each one missed and the resolver fell through to a generic family
default or the 256K fallback:

  openai/gpt-5.5:nitro           -> 256K   (real 1.05M)
  x-ai/grok-4.6:nitro            -> 131K   (generic "grok" catch-all)
  anthropic/claude-opus-4.6:nitro-> 200K   (generic "claude" catch-all)

The window silently shrank, triggering early compression and a wrong
/usage readout. f14059fa fixed the sibling half of this bug class in
/model validation; this fixes the metadata half.

Strip a recognized variant suffix for LOOKUP only, keeping the suffixed
id on the wire so the routing opt-in survives. Applied after the explicit
config overrides (steps 0b/0c) so a user-pinned value still wins, and
before every cache/catalog lookup. Gated on the request actually routing
through OpenRouter, so a local Ollama `model:tag` is untouched.

`:free`/`:batch`/`:thinking` are deliberately excluded — those ARE
distinct catalog SKUs with their own windows, so stripping them would
report the wrong number.

The suffix set and base-id split move to hermes_constants (import-safe,
dependency-free) so the metadata layer shares one definition with
hermes_cli.models instead of duplicating it.
2026-09-11 16:50:20 -07:00
Teknium 759024bdff fix(deepseek): honor Flash 1M window leftovers and native vision
Users on native DeepSeek were told to pin model.context_length and
model.supports_vision in config.yaml. That is the wrong layer: the 1M
window is already in DEFAULT_CONTEXT_LENGTHS, and a global
supports_vision pin would also mark text-only deepseek-v4-pro as
multimodal.

Two catalog gaps still produced the reported symptoms:

- A leftover context_length_cache.yaml entry of 128K (the old
  ``deepseek`` catch-all) outlived the 1M catalog keys because
  deepseek-flash was missing from _PRE_CATALOG_STALE_KEYS.
- When models.dev is empty/cold, Flash has no capability record, so
  image routing falls through to lossy text. Vendor docs (2026-09-10)
  mark deepseek-flash as vision-capable and deepseek-v4-pro as not.

Discard those 128K leftovers, fill Flash (and retired Flash aliases)
via _BUILTIN_MODEL_METADATA, and leave Pro catalog-only.
2026-09-11 12:09:41 -07:00
YipTszkwan 8435a3ae00 fix(deepseek): recognise the version-less deepseek-flash model id
DeepSeek's 2026-09 Flash refresh introduced a version-less canonical id:
GET /v1/models now returns `deepseek-flash` (alongside `deepseek-v4-pro`), the
API accepts it directly, and the older `deepseek-v4-flash` is server-side
aliased onto it. Every DeepSeek model-id gate in Hermes keys off the
`deepseek-v<N>` prefix, so the new id silently missed all four:

* DeepSeekProfile.build_api_kwargs_extras classified it as non-thinking and
  omitted `extra_body.thinking`. The server then defaults to thinking-on, so
  the user's thinking toggle and `reasoning_effort` were quietly ignored.
* `_normalize_for_deepseek` folded it onto `deepseek-v4-flash` (it misses the
  V-series regex), so the id a user picked never reached the wire and the
  config stored a different model than the picker advertised.
* `DEFAULT_CONTEXT_LENGTHS` fell through to the 128K `deepseek` catch-all
  instead of the real 1M window, capping the model at an eighth of its
  context before compaction kicked in.
* `_REASONING_STALE_TIMEOUT_FLOORS` had no entry, leaving the stale-stream
  detector at its 180s default instead of the 600s reasoning-model floor.

Verified live against api.deepseek.com: `deepseek-flash` answers 200 with
`model: deepseek-flash`, accepts image input (the refresh folds vision into
the Flash model), and the in-between id `deepseek-v4.1-flash` is rejected
with "The supported API model names are deepseek-flash, deepseek-v4-pro".

Adds the id to all four gates plus regression coverage for each site.
2026-09-10 02:44:26 -07:00
Teknium 4baf4269ff fix(metadata): GLM context windows from the live catalog; hyphenated slugs, entry-level overrides, aux inherits main pin
Every GLM slug that missed a DEFAULT_CONTEXT_LENGTHS key fell to the "glm" 202,752 catch-all, so
compression fired at ~20% of the real window and the aux feasibility check auto-lowered the session
threshold to that number (the 202,752 in the Coatue compression report).

- Catalog: GLM entries from Nous + OpenRouter /v1/models (2026-09-09): 5.3 / 5.3-flash 1,310,720
  (:batch/:US 1,048,576), 5.2 1,048,576, 5 / 5.1 / 4.7 / 4.6 204,800, *-turbo / 4.7-flash 202,752.
- _catalog_key_matches: version separators normalised on both sides so relay slugs like z-ai-glm-5-3
  hit glm-5.3 (#97398; approach from PR #97412 by @shellybotmoyer, re-implemented on current main).
- get_custom_provider_context_length: entry-level custom_providers[].context_length backs every model the
  entry serves when no per-model override exists, so /model switch stops dropping it (#98387; from
  PR #98396 by @liuhao1024, re-implemented on the config_providers sibling).
- check_compression_model_feasibility: when aux compression is the main model on the main route, reuse the
  main model's resolved window instead of re-resolving without its pin (#89500 mechanism 2, #45519).

Closes #97398, #98387, #89500, #87825, #97595, #97820 (suffix strip already on main; catalog values now match).
2026-09-10 01:08:18 -07:00
Teknium 2a95791992 simplify(compat): run_agent/model_tools/toolsets/acp/providers — drop 42 re-exports/aliases, repoint 15 callers + 99 test files
run_agent.py: delete the `# noqa: F401` re-export block (agent.process_bootstrap
OpenAI/_SafeWriter/_get_proxy_*, model_tools get_tool_definitions/
handle_function_call/check_toolset_requirements, FailoverReason,
_qwen_portal_headers/_routermint_headers, session_persistence names,
estimate_request_tokens_rough, ContextCompressor + friends, jittered_backoff,
prompt_builder names, message_sanitization names, tool_dispatch_helpers
names) — 41 names run_agent never used itself — and the `_STREAM_DIAG_HEADERS`
back-compat class alias (no in-tree reader). run_agent now imports only what
it uses (get_toolset_for_tool, is_local_endpoint, coalesce/uniquify tool-call
ids, cleanup_vm/get_active_env from terminal_tool_lifecycle).

agent/*: `_ra().X` late-binds that only reached a re-export now import the
defining module directly (agent_runtime_helpers -> process_bootstrap.OpenAI,
model_tools.handle_function_call, session_persistence._safe_session_filename_component;
agent_init -> model_tools.get_tool_definitions/check_toolset_requirements,
_lazy_headers("agent.client_lifecycle", ...) for qwen/routermint;
system_prompt -> agent.prompt_builder / model_tools directly, dropping its
own _ra() shim and the `_r` parameter threading). `_ra()` stays for
run_agent-resident names (logger, AIAgent, _hermes_home, _set_interrupt, ...).

toolsets.py: remove resolve_multiple_toolsets (shim-only, restored by
34abf954bd); tests/test_toolsets.py pins the same union behavior via
resolve_toolset over each name.

providers/__init__.py: drop the OMIT_TEMPERATURE re-export (no callers via the
package); ProviderProfile stays because __init__ uses it for annotations —
2 tests repointed to providers.base.

agent/iteration_budget.py: drop the "run_agent re-exports the class"
docstring pointer; 4 tests import IterationBudget from its home.

model_tools.py (arg_coercion names), agent/tool_executor.py, and
hermes_cli/cli_session_mixin.py repoints landed via a sibling commit on this
shared worktree.

Callers repointed: gateway/run.py, hermes_cli/cli_chat_turn_mixin.py,
hermes_cli/cli_tui_mixin.py, tui_gateway/session_workdir.py,
agent/transports/codex.py (one-line imports) + comment pointers in
tools/file_state.py, tools/schema_sanitizer.py, scripts/tool_search_livetest.py.
Tests: patch("run_agent.X") / monkeypatch.setattr(run_agent, "X") /
`from run_agent import X` -> defining module across 99 test files.
2026-09-03 13:28:22 -07:00
Teknium 4359af7705 fix(models_dev): alias opencode-free to the Zen "opencode" catalog; pin Muse Spark 1M invariant
opencode-free had no PROVIDER_TO_MODELS_DEV entry, so every models.dev
lookup on the free tier missed and Muse Spark fell to the 256K default.
The free tier is served by the Zen relay (hermes_cli/models.py:
"opencode-free is Zen-hosted"), and models.dev's "opencode" provider is
the catalog that lists muse-spark-1.2 / -1.2-contributor-free /
-1.3-contributor-free at 1,048,576 — so the alias is "opencode", not
"opencode-go" (Go's catalog carries only the paid -contributor SKUs).

Missing alias identified by @Steve-prog001 in #101905.

Tests: one parametrized offline invariant (models.dev + live /models
mocked away) asserting 1,048,576 on opencode-free / opencode-go /
meta-ai / commandcode — fails on main, passes here — plus the alias pin.
2026-09-03 00:58:16 -07:00
mr-r0b0t cfa7e72c9e fix(models): correct contributor guard, 1M context, docs for muse-spark-1.3
- model_data_policy_guard: name the triggering -contributor model instead
  of hardcoded 1.2; per-version verified price tables (1.3 standard
  $1.25/$4.25 via OpenRouter live metadata; cached figures 1.2-only)
- model_metadata: muse-spark-1.3 + muse-spark family at 1048576 (OpenRouter
  verified 2026-09-02) with pre-catalog stale-cache keys so 256K-fallback
  sessions self-heal
- docs: contributor-tier notes cover 1.2 + 1.3
- tests: 1.3 guard regression, muse stale-cache guard, live-catalog mirror
  gains 1.3-contributor-free (confirmed on live relay)

143 tests pass (guard, selection guards, opencode catalog, model_metadata);
ruff clean.
2026-09-03 00:57:55 -07:00
Teknium 32fe129324 perf(bot-mode): cold DM hops skip the live /models probe; relay replies land within 250ms
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.

- model_metadata: memoize successful remote /models probes on disk
  (cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
  in-memory cache, so authority semantics are unchanged (reconciliation
  still lands within 5 minutes) but the answer is shared across
  processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
  250ms instead of every 2s — up to 2s of dead air on every relayed reply.

Nothing here changes turn ordering: DMs and group rounds stay serial.

Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
2026-09-02 03:42:01 -07:00
Teknium 58f5b1e277 fix(model_metadata): parse Google's 'supports up to N' context-limit phrasing
Google Gemini/Gemma overflow errors read 'Unable to submit request because
the input token count is 32825 but model only supports up to 32768'.
parse_context_limit_from_error had no pattern for the 'supports up to N'
phrasing, so overflow recovery kept the wrong window and burned its retry
attempts instead of recalibrating to the provider-reported limit.

Add the anchored pattern (limit follows 'supports up to'; the larger input
count before it is never captured) plus regression tests covering the exact
message and the get_context_length_from_provider_error recalibration path.

Reported by @Artemonim in #57275 (residual claim 5).
2026-08-31 12:20:02 -07:00
Teknium 7b89e17774 fix(model): single exact eligibility predicate for -900k variants; reject ineligible aliases
Review findings on #92797 (@100yenadmin):
- is_codex_900k_base() is now the single source of truth used by picker
  synthesis, context resolution, /model validation, and wire stripping.
  Eligibility is an exact table (sol/terra/luna, gpt-5.4, daybreak alias)
  plus date-shaped 5.6 snapshots — family-prefix matching removed, so
  non-routable -pro slugs and unknown descendants never gain variants.
- strip_codex_context_variant_suffix() strips conditionally: ineligible
  aliases (gpt-5.5-900k) are returned unchanged and fail honestly at the
  API instead of silently running as the base model at 272K.
- validate_requested_model() rejects ineligible *-900k aliases before the
  hidden-slug soft-accept, and accepts valid variants missing from a
  stale catalog without letting the typo auto-corrector eat the suffix.
- Codex context resolver drops vendor/ namespaces, so
  openai/gpt-5.6-sol-900k resolves to 900K like the bare id.
- Table-driven regression covering eligible bases/snapshots/namespaced
  ids and rejected -pro/-mini/5.5/unknown aliases, asserting context AND
  wire model.
2026-08-23 02:14:35 -07:00
Teknium 63a9c26fbe feat(model): Codex GPT slugs default back to 272K; explicit -900k picker variants opt into the verified large window
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the
live-verified 900K burned through subscription usage for users who never
asked for the larger window (bigger window = more input tokens per request).

- Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the
  advertised 272K again — the cheaper limit is the default.
- The model picker synthesizes explicit <slug>-900k variants (e.g.
  gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into
  the 900K window. Slugs that genuinely enforce 272K (gpt-5.5,
  gpt-5.4-mini) get no variant.
- The -900k suffix is Hermes-side only: stripped before the model id hits
  the wire (main transport + auxiliary Responses adapter), and pricing
  aliases the variants onto the base entries.
- Docs: new opt-in section in context-compression-and-caching.md.
2026-08-23 02:14:35 -07:00
unsupportedpastels f8e5949f61 fix(model_metadata): add Daybreak Codex 900K context 2026-08-21 13:19:07 -07:00
Teknium bab7be3ca7 feat: raise Codex OAuth context to 900K for gpt-5.6 family and gpt-5.4 (subscription 1M rollout)
OpenAI enabled the large-context window for ChatGPT-subscription Codex
accounts (announced by @thsottiaux Aug 16 2026; previously API-key-only).
Live re-probe the same day: 911,276 input tokens completed OK on
gpt-5.6-sol; ~925K+ rejected with context_length_exceeded (1.05M window
minus reserved output headroom). terra, luna, and gpt-5.4 all completed
900,026 tokens OK. The Codex catalog still advertises 272K, so the
stale-advertisement override from #87981 is the right lever — this just
raises its value 350K -> 900K.

gpt-5.5 and gpt-5.4-mini still enforce 272K live (rejected 500K) and
remain excluded. Override semantics unchanged: fires only on an
exactly-272,000 advertisement; any live catalog change is trusted
verbatim.
2026-08-16 18:31:54 -07:00
Teknium 5229975438 feat: raise Codex OAuth context to live-verified 350K for gpt-5.6 family and gpt-5.4
The Codex /models catalog advertises 272K for the gpt-5.6 (sol/terra/luna)
and gpt-5.4 slugs, but the backend actually accepts ~371K input tokens
(verified live against chatgpt.com/backend-api/codex/responses, Aug 16 2026:
~371K completed OK on all four slugs; ~382K+ rejected with
context_length_exceeded). 350K keeps ~22K margin under the observed ~372K
enforcement.

The bump applies ONLY when the resolved value is exactly the known-stale
272,000 advertisement — any other advertised value (higher or lower) is
trusted as a real server-side change, so a future catalog correction
deactivates the override automatically. gpt-5.5 and gpt-5.4-mini both
genuinely enforce 272K (rejected 360K live) and are excluded.
2026-08-16 15:43:09 -07:00
coe0718 7f7aefe5cb fix: restore complete message timestamp coverage 2026-08-15 01:04:19 -07:00
luoxiao6645 b09e1daa84 fix(agent): reject stale 32k metadata for MiniMax 2026-08-13 11:12:05 -07:00
sasquatch9818 6def7ce1df fix(models): write context-length cache atomically
save_context_length() and _invalidate_cached_context_length() did an
unguarded read-modify-write into $HERMES_HOME/context_length_cache.yaml.
The plain `open(path, "w")` truncates the file before the dump runs. If
the process is killed mid-dump, the file is left empty or partial. The
next _load_context_cache() swallows the YAML error and returns {} —
silently wiping every persisted context length. A concurrent process
reading between truncate and dump-complete also sees a torn file.

After the cache is lost, every model re-probes the network, and when a
probe fails it falls back to the generic 256K default — so a user on a
1M-window model ends up with a wrong, short context window.

Hermes routinely runs several processes against one shared $HERMES_HOME
(a cron agent plus an interactive session, multiple gateway sessions),
so this is hit in normal use.

Switch both writers to the existing utils.atomic_yaml_write helper
(temp file + fsync + os.replace, symlink- and mode-preserving). The real
file is only ever swapped from a fully written temp file, so an
interrupted write leaves the previous cache intact and readers never see
a partial file. Matches the atomic-write pattern already used for
auth.json, config.yaml, and other persisted state.

Makes the persistent model context-length cache write crash-safe. The
old non-atomic write could truncate or wipe the entire cache on an
interrupted or concurrent write, which then forces models onto the wrong
fallback context window. The fix routes both cache writers through the
repo's atomic temp-file + os.replace helper.

N/A

- [x] 🐛 Bug fix (non-breaking change that fixes an issue)
- [ ] ✨ New feature (non-breaking change that adds functionality)
- [ ] 🔒 Security fix
- [ ] 📝 Documentation update
- [ ] ✅ Tests (adding or improving test coverage)
- [ ] ♻️ Refactor (no behavior change)
- [ ] 🎯 New skill (bundled or hub)

- `agent/model_metadata.py`: `save_context_length()` and
  `_invalidate_cached_context_length()` now write via
  `utils.atomic_yaml_write` instead of a truncating `open(path, "w")`.
  Added the `atomic_yaml_write` import.
- `tests/agent/test_model_metadata.py`: added
  `test_write_failure_leaves_existing_cache_intact` — simulates a crash
  during the atomic swap and asserts the existing cache survives
  byte-for-byte with no stray temp file.

1. `pytest tests/agent/test_model_metadata.py -q` — 98 pass, including
   the new crash-safety test.
2. The new test seeds a valid cache, forces the swap step to raise, and
   confirms the file is not truncated and no `.cache_*.tmp` is left.
3. `ruff check agent/model_metadata.py` passes.

- [x] I've read the Contributing Guide
- [x] My commit messages follow Conventional Commits (`fix(scope):`, etc.)
- [x] I searched for existing PRs to make sure this isn't a duplicate
- [x] My PR contains **only** changes related to this fix
- [x] I've run the affected tests (`pytest tests/agent/test_model_metadata.py -q`) and they pass
- [x] I've added tests for my changes
- [x] I've tested on my platform: macOS 15 (Darwin 25.5)

- [x] I've updated relevant documentation (README, `docs/`, docstrings) — or N/A
- [x] I've updated `cli-config.yaml.example` if I added/changed config keys — or N/A
- [x] I've updated `CONTRIBUTING.md` or `AGENTS.md` if I changed architecture or workflows — or N/A
- [x] I've considered cross-platform impact (Windows, macOS) — the helper uses os.replace, which is atomic on both
- [x] I've updated tool descriptions/schemas if I changed tool behavior — or N/A
2026-08-13 11:08:26 -07:00
Teknium d3a8be4630 test: regression coverage for the non-positive context-cache guard
Follow-up to the salvaged #25812 — the original PR shipped without tests.
2026-08-13 11:05:49 -07:00
whirmill 4a6d3640b9 fix(agent): default context lookup for empty model IDs
An empty/blank model id reaching get_model_context_length() can't be
meaningfully resolved — and it's worse than a miss: the endpoint
metadata fuzzy matcher ('model in key or key in model') is vacuously
true for "", so it matches an ARBITRARY catalog entry from the live
/v1/models response and returns whatever context length that entry
happens to have, persisting it under a junk '@<base_url>' cache key.

This started failing CI on main when the Nous portal catalog changed:
tests/run_agent/test_primary_runtime_restore.py constructs agents with
model='' against the live portal URL, the arbitrary match now lands on
a 32K entry, and init_agent raises the 64K-floor ValueError
(test_allowed_for_nous_anthropic_messages, red on every PR's slice).

Guard early: a blank model id falls back to DEFAULT_FALLBACK_CONTEXT
immediately, before any cache write or network probe.

Salvaged from #65515 by @whirmill (rebased onto current main; the
guard now sits after the malformed-base_url normalization added since,
and carries an explanatory comment for the fuzzy-match footgun).

Fixes the red slice on #85444, #85452 and every other open PR.

Co-authored-by: whirmill <5079591+whirmill@users.noreply.github.com>
2026-08-13 23:19:07 +05:30
Teknium 91e550b0cf fix(model_metadata): generalize pre-catalog stale context-cache guard
Replaces the per-model _model_name_suggests_grok_4_3/_grok_4_6/
_minimax_m3 stale-cache predicates with one generic
_stale_pre_catalog_cache_entry() guard driven by
_PRE_CATALOG_STALE_KEYS. A cached context length is dropped when the
model resolves (longest-key-first, same as step 8) to a listed catalog
key and the cached value is at or below what the old resolution path
could have produced (largest shorter matching catch-all, or the 256K
fallback).

Also covers qwen3.6-plus, grok-4-fast, and grok-4.20 (the models
PR #37684 requested guards for), absorbing that PR.

_model_name_suggests_minimax_m3 is kept for its two non-cache callers
(models.dev underreport guard, cache-control gating in
agent_runtime_helpers).
2026-08-13 10:21:50 -07:00
Julientalbot 53ad7794e5 fix(xai): drop stale 256K grok-4.6 context cache
docs.x.ai (2026-08-12): grok-4.6 is the flagship, 500K context.
Live GET /v1/models lists grok-4.6 at context_length 500000
(no grok-4.6-latest alias).

#84661 landed the catalog. Main already lists native grok-4.6
on the xAI picker. This is only the leftover cache guard
(same pattern as grok-4.3): pre-catalog builds persisted the
grok-4 catch-all (256K).
2026-08-13 10:21:50 -07:00
Teknium 1a796a1247 fix(model_metadata): never fuzzy-match an empty model name against endpoint catalogs
'' is a substring of every catalog key, so _resolve_endpoint_context_length
with an empty model name "matched" whatever the endpoint listed first —
on the Nous portal that is currently a 32K embedding model, which poisoned
the resolved context length and made AIAgent init fail the 64K minimum.
This is what turned tests/run_agent/test_primary_runtime_restore.py::
TestTryRecoverPrimaryTransport::test_allowed_for_nous_anthropic_messages
red on every PR (CI slice 7/12) after the portal catalog reordered.

Single-model endpoints still resolve with an empty name (unambiguous);
non-empty names keep the substring fuzzy match.
2026-08-13 10:15:12 -07:00
Teknium 357b97eda6 test(model-metadata): use explicit fixture encodings 2026-08-09 01:51:12 -07:00
Teknium d143bf7a3b fix(model-metadata): resolve provider prefixes from live registry 2026-08-09 01:51:12 -07:00
Oliver Mee 19e51d2cca fix(model-metadata): auto-extend provider prefixes from registered profiles
_PROVIDER_PREFIXES was a hand-maintained frozenset, so providers that ship
as plugins (bundled like fireworks, or user plugins under
$HERMES_HOME/plugins/model-providers/) were never recognised as
provider: prefixes in model strings, and metadata/context-window lookups
received the unstripped string. Mirror the _URL_TO_PROVIDER auto-extend
that already sits below it: add each registered profile's name and
aliases after discovery. The _OLLAMA_TAG_PATTERN guard keeps model:tag
strings intact.

Fixes #66106
2026-08-09 01:51:12 -07:00
Josh Tsai 013779924f fix(agent): fail fast on custom-provider /models auth errors
- Short-circuit the candidate waterfall on HTTP 401/403: an auth wall
  proves the endpoint family exists, so probing the alternate URL just
  doubles the wasted wait (the reported endpoint takes ~10s to return
  401 without a key).
- Stream the probe so 4xx never downloads a slow error body; responses
  are closed on every exit path.
- Regression tests: single-call assertion on 401/403 (fails on main),
  negative-cache reuse, 404 waterfall preserved, no .json() on 4xx.

Fixes #69905

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 22:47:48 +05:30
kshitijk4poor e078c8c6ef fix: widen fallback warning to the sibling custom-endpoint 256K path
Review pass 2 (reuse reviewer HIGH): the step-3b probe-down fallback for
custom/local endpoints returns the same silent 256K default but only
logged at INFO - invisible by default, and it is the MORE common path
for small local models (the exact users the warning exists for).

Extract _warn_context_length_fallback() (deduped per model+base_url)
and call it from both fallback sites, per the fix-the-whole-bug-class
rule. Regression test drives the custom-endpoint path and fails without
the widening (mutation-checked).
2026-08-01 15:05:05 +05:30
kshitijk4poor 4c2d0c7fd8 refactor: dedupe fallback warning per model, drive pool-cleanup tests through real run()
Review follow-up:
- Warn once per (model, base_url) at the step-9 fallback via a module-level
  dedup set (established _WARNED_* idiom). The fallback result is
  deliberately never cached, so the un-deduped warning fired on every
  resolution - e.g. once per gateway message via the session-hygiene path.
- Replace the three inline-mock pool-cleanup tests (which reproduced the
  try/except block against a MagicMock and passed even with the production
  code reverted) with a parametrized test that drives the real
  BatchRunner.run() with a patched Pool; drop the CPython stdlib
  signature change-detector test.
- Add a once-per-model warning regression test; clean up dead imports.

All tests verified to fail against pre-PR batch_runner.py/model_metadata.py
and pass with the fix (mutation check).
2026-08-01 15:05:05 +05:30
kshitijk4poor a1ff62a139 fix: context-length fallback logging, batch trajectory durability, pool cleanup
Salvage of #6629 by aaronlab (kshitijk4poor reworked against current main).

Three concerns from the original PR, reworked to address review feedback:

1. Context-length fallback diagnostic (agent/model_metadata.py):
   get_model_context_length() silently returned 256K when all 9 detection
   methods failed. Users with small-context models (8K, 32K) would get 256K
   silently, causing hard-to-debug API context-length errors. Added a
   warning log at the step 9 fallback with model name, base_url, and the
   correct config override hint (model.context_length, not context_length).
   The token-estimation ceiling-division fix from the original PR already
   landed on main (5c2ecdec) with CJK handling — not duplicated here.

2. Fsync for batch trajectory writes (batch_runner.py):
   Trajectory entries were written without flush/fsync, but the checkpoint
   immediately marked them as completed. A crash between write and disk
   sync would leave the checkpoint claiming completion with no trajectory
   data on disk. Added flush() + os.fsync() before checkpoint update.

3. Pool cleanup on interruption (batch_runner.py):
   Ctrl+C during pool.imap_unordered() relied on context manager cleanup
   which can hang. Added explicit pool.terminate() + pool.join() for both
   KeyboardInterrupt and Exception paths. The original PR used
   pool.join(timeout=10) which is invalid — CPython's Pool.join() takes
   no timeout parameter. Fixed to use pool.join() without arguments.

Tests:
  - test_warning_emitted_on_fallback: verifies warning fires at step 9
  - test_no_warning_when_cached: verifies no false warning when cache hits
  - test_trajectory_entry_is_synced_to_disk: verifies os.fsync is called
  - test_pool_terminate_called_on_exception: verifies cleanup on RuntimeError
  - test_pool_terminate_called_on_keyboard_interrupt: verifies cleanup on Ctrl+C
  - test_pool_join_called_without_timeout: verifies no timeout arg to join()
  - test_real_pool_join_accepts_no_timeout: integration check on CPython API

Co-authored-by: Aaron Lab <aaronlab@users.noreply.github.com>
2026-08-01 15:05:05 +05:30
Israel Lot 5c45d9c208 fix(agent): mirror substitute_api_content's guard in the estimator shadow
Review follow-up on #75102. The shadow substituted the sidecar whenever
the ``api_content`` key was merely PRESENT, but the wire only substitutes
a non-empty string sidecar on a user/assistant row (see
``turn_context.substitute_api_content``). For any other shape the sidecar
is popped and discarded while the clean ``content`` is sent -- so the
shadow dropped real content from the estimate and UNDERcounted, the
dangerous direction: compaction fires too late and the turn dies on a
hard context-length error instead of merely compressing early.

Gate the substitution on the same predicate, and cover the divergent
shapes (None, empty string, int, list, non-user/assistant role) with a
test that fails against the unconditional version.

Also rename the image test: it never carried a sidecar, so it was not
testing what its name claimed. It is a non-regression pin on the flat
per-image accounting that moved into ``_wire_message_shadow()``, and is
now named for that.
2026-08-01 11:10:15 +05:30
Israel Lot e3bc517034 fix(agent): stop double-counting api_content in the token estimator
`api_content` is a SUBSTITUTE for `content`, not an addition to it.
`turn_context.substitute_api_content()` pops the sidecar and overwrites
`content` at every API-bound message-build site (the `api_messages` build
in `conversation_loop`, the max-iterations summary in
`chat_completion_helpers`, the chat-completions transport), so exactly one
of the two is ever sent to the provider.

The preflight estimator counted both, because both `_estimate_message_chars`
and `_estimate_message_tokens_without_images` walked every key of the
persisted dict with a single-entry denylist (`_anthropic_content_blocks`).
Any message whose sidecar differs from its clean stored content was counted
twice — exactly 2.00x on a 40KB sidecar.

The sidecar exists to keep the provider prompt-cache prefix byte-stable, so
it is written on precisely the long, cache-pinned messages where the
doubling hurts most. Because `estimate_messages_tokens_rough()` also feeds
the compaction threshold via `context_compressor` and `conversation_loop`,
the inflated estimate makes compression fire on phantom bytes.

Fix: substitute rather than sum, mirroring the wire. The two estimator
helpers had drifted into near-identical copies of the same shadow-building
loop, so this factors the shared logic into `_wire_message_shadow()` and
fixes the class once instead of patching one site and leaving the other.

Image accounting is unchanged: base64 payloads are still replaced with a
placeholder and charged at the flat `_count_image_tokens` rate, and the
`_multimodal` text_summary path is preserved.

Tests: three cases in `TestEstimateMessagesTokensRough` — sidecar equal to
content is counted once, a sidecar that DIFFERS is still counted (a lower
bound, so it fails if the field were dropped rather than substituted, which
would undercount the real request), and a sidecar cannot smuggle raw base64
past the flat image rate.

Verified on Linux (Python 3.11): 53 passed in
tests/agent/test_model_metadata.py, 57 passed with
tests/agent/test_context_breakdown.py, 656 passed / 3 skipped across the
compression/context/token/estimate/prune surface of tests/agent.
Mutation-tested: reverting the substitution fails the new equality test.
`scripts/check-windows-footguns.py` is not applicable — no file I/O,
process management, terminal handling, subprocesses, or signals.
2026-08-01 11:10:15 +05:30
Gille 29eac371d1 fix(context): persist NVIDIA DeepSeek endpoint limit 2026-07-31 22:31:22 -07:00
Teknium 6b81590c55 test: prune low-value tests suite-wide (wave 1) — 46,820 → 28,106 test functions
Systematic prune per AGENTS.md test policy, one pass over every major
test tree (gateway, hermes_cli, tools, agent, run_agent, plugins, cli,
cron, tui_gateway, honcho/openviking, root-level):

- DELETE: source-reading tests (read_text/getsource on prod files),
  change-detector tests (exact catalog counts, model-name snapshots,
  config version literals), mock-echo tests (assert a mock returns what
  it was told), assertion-free/trivial tests, near-duplicate
  parametrizations (boundaries + one representative kept), async/sync
  twin duplicates, cosmetic within-file variations.
- KEEP (mandatory): security/redaction/approval guards, message-role
  alternation invariants, prompt-caching/deterministic-call-id
  invariants, issue-number regression tests (deduped), E2E tests.
- 6 test files deleted outright (script-style/no-assert or fully
  redundant); conftest.py, fakes/, fixtures/ untouched.
- tests/acp/conftest.py added: autouse fixture stubs the live
  models.dev/GitHub/Copilot/Anthropic inventory fetches that ACP server
  tests performed on every session create — test_server.py 147s → 3.4s,
  and the tests are now genuinely hermetic.
- Sleep-based slowness shrunk where safe (codex_ttfb_watchdog,
  compression_concurrent_fork, etc.); no wall-clock assertion tightened.

Verification: full hermetic suite via scripts/run_tests.sh —
2439 files, 31,130 tests passed, 0 failed, 0 flaky retries, 315s wall
(baseline: 583s wall, 13,564s subprocess CPU).
2026-07-29 13:10:23 -07:00
atakan g 713982a8f8 test(model-metadata): cover local Ollama fallback 2026-07-28 14:18:18 -07:00
atakan g 0c2d9aee0b fix(model-metadata): prefer local Ollama num_ctx 2026-07-28 14:18:18 -07:00
wjq990112 78312c192d fix(moa): preserve custom provider context metadata
Preserve compatible custom provider metadata through MoA aggregator context resolution and cover the resolver and compressor paths.
2026-07-23 11:21:04 -07:00
Teknium ea0fd393db perf(compression): gate CJK-aware token estimation behind an ASCII fast path
The salvaged estimator ran a per-character Python loop on every
estimate_tokens_rough() call — a ~28,000,000x slowdown vs (len+3)//4 on a
1MB ASCII tool output (measured ~3.0s per call). Gate it:

- str.isascii() O(1) fast path keeps pure-ASCII text bit-identical to the
  classic (len+3)//4 rule at ~1.3x baseline cost (0.23us vs 0.17us per
  1MB call).
- Non-ASCII text counts dense CJK chars via a compiled character-class
  regex in C (len(text) - len(re.sub(''))): ~352ms/1MB hangul vs ~2.1s
  for the per-char loop.
- Non-ASCII-but-non-CJK text (accents, Cyrillic, emoji) keeps the classic
  rule.

Also: parity tests against the per-char reference implementation, and
updated two stale expectations that encoded the old behavior (CJK now
counted ~1 token/char; short string content now ceil-divided instead of
floored to 0). The continuity test now detects merged-into-tail summaries
via _is_context_summary_content.
2026-07-22 06:57:22 -07:00
sbe27 0c0ec18d6f test(context): cover Codex context rollback 2026-07-21 04:29:34 -07:00
sbe27 60afc290a8 fix(context): scope Codex catalogue cache by credential 2026-07-21 04:29:34 -07:00
sbe27 9a34cc91a5 test(context): document Codex cache persistence coverage 2026-07-21 04:29:34 -07:00
sbe27 8a0701ca48 fix(context): revalidate Codex OAuth context windows 2026-07-21 04:29:34 -07:00
Teknium 25eafd7d71 fix(models): complete kimi-k3 rollout across Kimi-direct catalog surfaces
Follow-up widening for salvaged PRs #67115, #67685, #67620:

- _PROVIDER_MODELS: add kimi-k3 atop kimi-coding / moonshot / opencode-go
  curated lists (kimi-coding-cn covered by cherry-picked #67620)
- setup.py _DEFAULT_PROVIDER_MODELS: kimi-k3 for kimi-coding(-cn) + opencode-go
- model_metadata: align DEFAULT_CONTEXT_LENGTHS kimi-k3 entry to 1,048,576
  (matches endpoint-scoped override, models.dev, and OpenRouter live metadata)
- anthropic_adapter: classify the bare Coding Plan slug 'k3' (and k3.x/k3-*)
  as Kimi family so adaptive thinking applies on proxied endpoints
- moonshot_schema: is_moonshot_model matches bare 'k3' so tool-schema
  sanitization runs on the chat-completions path
- contributor mappings for githubespresso407, datachainsystems, Punyko8

Tests: 582 passed across 11 targeted files; hermetic E2E verifies picker
order (kimi-k3 first), no dupes, and 1M context resolution.
2026-07-20 08:47:55 -07:00
datachainsystems 54c39c0301 fix: add Kimi K3 1M context window to DEFAULT_CONTEXT_LENGTHS
Kimi K3 ships with a 1M-token context window (verified against
platform.kimi.ai/docs/overview) but was falling through to the generic
'kimi': 262144 catch-all. Added 'kimi-k3': 1_000_000 before the catch-all
so longest-key-first substring matching resolves K3 to 1M while older
Kimi models still hit the 256K default.

Added matching test_kimi_k3_context_1m test covering native,
vendor-prefixed (kimi/, moonshotai/), and older model fallback.
2026-07-20 08:47:55 -07:00
githubespresso407 77aa026ca6 Resolve kimi-k3 context length to 1M on canonical Kimi Coding endpoints
Kimi Coding serves K3 under the bare slug 'k3', but users can also
configure or select the public-facing aliases 'kimi-k3' and
'kimi-k3-cot'. The endpoint-scoped 1M context window was only keyed
on the bare 'k3' slug, so selecting 'kimi-k3' fell through to the
generic 'kimi' catch-all (262k).

Extend the guard in _endpoint_scoped_context_length to also recognize
'kimi-k3' and 'kimi-k3-cot', while keeping the endpoint check that
limits the 1M value to https://api.kimi.com/coding (legacy Moonshot
endpoints still fall back to 262k). Update the existing test to cover
all three aliases.

Fixes: context window limited to 262k when using kimi-k3 via kimi-coding.
2026-07-20 08:47:55 -07:00
Avi Fenesh 77ba81f75f fix(bedrock): add Fable + Claude 4.6/4.7/4.8 1M entries to context table, drop stale cached values
BEDROCK_CONTEXT_LENGTHS was missing entries for current 1M-context Claude
models, and the resolution path in get_model_context_length() short-circuits
to that table (step 1b) before DEFAULT_CONTEXT_LENGTHS is ever consulted, so
the catalog's correct values could never apply on Bedrock:

- claude-fable-5 (no entry at all) fell through to
  BEDROCK_DEFAULT_CONTEXT_LENGTH and reported 128K for a 1M model.
- opus-4-7 / opus-4-8 substring-matched the generic 'anthropic.claude-opus-4'
  key and reported 200K.
- opus-4-6 / sonnet-4-6 had explicit 200K entries predating their 1M windows.

The practical symptom: the agent compresses context prematurely (at ~128K or
~200K of a 1M window) on every Bedrock-hosted current Claude model.

Fixing the table alone is not enough for existing installs: a previously
persisted 128K/200K value in the context-length cache wins at step 1 and
masks the corrected table forever. Step 1 now reconciles Bedrock-context
cache hits against the static table (the table is authoritative for Bedrock
— there is no live probe to reconcile against), invalidating stale entries
so existing users converge to the right window without manual cache surgery.

Tests cover the new table entries (incl. inference-profile and versioned ID
forms), the 128K-default regression for Fable, the stale-cache invalidation
path, and that pre-4.6 models keep their 200K entries.
2026-07-20 05:38:55 -07:00
amanning3390 311a5b0a55 feat(kimi): discover K3 on coding endpoint 2026-07-16 13:33:02 -07:00
teknium1 881a9520e3 test: regression coverage for null context_lengths key (#47135) 2026-07-09 19:57:43 -07:00
kshitijk4poor 55dbc3ffb5 fix(model_metadata): bound the tools-token estimate cache
Follow-up to the salvaged str(tools) fix. The id()-keyed
_TOOLS_TOKENS_CACHE had no eviction, so a long-lived gateway/desktop
backend could accumulate an unbounded number of stale entries as it
builds transient tool lists. Cap it at 256 with oldest-first eviction
(insertion-ordered dict) and add a regression test asserting the cache
never exceeds the cap.
2026-07-09 12:34:11 +05:30