Google Gemini/Gemma overflow errors read 'Unable to submit request because
the input token count is 32825 but model only supports up to 32768'.
parse_context_limit_from_error had no pattern for the 'supports up to N'
phrasing, so overflow recovery kept the wrong window and burned its retry
attempts instead of recalibrating to the provider-reported limit.
Add the anchored pattern (limit follows 'supports up to'; the larger input
count before it is never captured) plus regression tests covering the exact
message and the get_context_length_from_provider_error recalibration path.
Reported by @Artemonim in #57275 (residual claim 5).
Test seams and plugin engines monkeypatch estimate_messages_tokens_rough with (messages)-only signatures; route callers only pass the charge_stale_thinking kwarg on the False path.
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).
Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.
Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
Follow-up for salvaged PR #96939: main already added hy4-preview at
1_048_576 (f7c79efbac); the cherry-picked duplicate key later in the
dict silently overrode it with 1024000.
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.
- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
with a structural base-message identity check that fails closed on any
transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
compaction (codex_runtime), session reset/switch (run_agent), plus the
fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
estimation fallback (first request of a session).
Live on both providers (verified 2026-08-28 against openrouter.ai/api/v1/models
and inference-api.nousresearch.com/v1/models) but absent from both curated
picker lists. Adds the entry directly below qwen3.8-max per newest-first
family ordering, an explicit 1M DEFAULT_CONTEXT_LENGTHS entry (new family
slug would otherwise fall through to the generic qwen 131072 catch-all —
same class as #69881), and regenerates model-catalog.json.
Scoped rollout: only the named providers touched. Pricing snapshot skipped
(both routes bill via official_models_api live pricing). Reasoning floor
already fires via the qwen3 prefix entry (180s, verified).
- move minimax/minimax-m3:free into the Free tier section (house
convention: :free SKUs group together, matching glm-5.2:free and the
nemotron :free entries) and regenerate model-catalog.json
- add Inkling family context length (1,048,576 — OpenRouter live
metadata, 2026-08-27) to DEFAULT_CONTEXT_LENGTHS; new family slug
otherwise fell through to no entry
- add Inkling to the reasoning stale-timeout floor table (300s tier,
same as Grok reasoning / Ox Alpha; OpenRouter marks the family as
reasoning-capable)
- widen the floor matcher's right-anchor separator class to include
':' so OpenRouter SKU suffixes (:free/:batch/:nitro) inherit the
family floor — inkling:free previously missed the inkling entry
- regression tests for the inkling floor + ':' separator
The two local-server context probes in _query_local_context_length read
data.get("max_tokens") as a context-window candidate. On an
OpenAI-compatible /v1/models passthrough max_tokens is the max OUTPUT
tokens, so a 1M-context model advertising a 128K output cap resolves to
128000 and auto-compaction fires ~7x early.
Route both branches through the module's own key vocabulary
(_CONTEXT_LENGTH_KEYS), which already classifies max_tokens as a
_MAX_COMPLETION_KEYS entry.
Local /v1/models probes treated Anthropic `max_tokens` (max output) as the
context window when `max_model_len`/`context_length` were absent. Anthropic
and Anthropic-compatible reverse proxies expose both:
max_input_tokens = context window (e.g. 1M for claude-fable-5)
max_tokens = max output (e.g. 128k)
That under-reported windows (1M → 128k), persisted the wrong value into
context_length_cache.yaml, and fired compression at ~96k (75% of 128k).
Route model objects through a shared helper that prefers input-window keys
via _extract_context_length, and only falls back to max_tokens when no
input-window field is present.
The image-routing vision path calls detect_local_server_type without
the provider's API key. Against a remote API-keyed endpoint (sglang /
vLLM with --api-key) every leg of the 5-request probe waterfall came
back 401 — and because a failed verdict was never written to the
in-memory cache (only positive verdicts were), the waterfall re-ran on
EVERY image-bearing turn (#89863: 51 detail-less busy-acks observed in
one Slack channel while the probe sprayed the user's own server).
Two changes:
- image_routing._should_probe_ollama_vision now takes the API key and
forwards it; a new _resolve_inference_api_key mirrors
_resolve_inference_base_url's resolution order (runtime value,
model.api_key, providers blocks) so the key always matches the URL
being probed.
- detect_local_server_type caches a None verdict in memory with a short
failure TTL (5 min, vs 1h for positives) so the next turn is served
from the negative entry instead of re-running the waterfall — while
a transient failure (server starting, key being fixed) recovers in
minutes. Negative verdicts are deliberately not written to the
cross-process disk cache.
Review findings on #92797 (@100yenadmin):
- is_codex_900k_base() is now the single source of truth used by picker
synthesis, context resolution, /model validation, and wire stripping.
Eligibility is an exact table (sol/terra/luna, gpt-5.4, daybreak alias)
plus date-shaped 5.6 snapshots — family-prefix matching removed, so
non-routable -pro slugs and unknown descendants never gain variants.
- strip_codex_context_variant_suffix() strips conditionally: ineligible
aliases (gpt-5.5-900k) are returned unchanged and fail honestly at the
API instead of silently running as the base model at 272K.
- validate_requested_model() rejects ineligible *-900k aliases before the
hidden-slug soft-accept, and accepts valid variants missing from a
stale catalog without letting the typo auto-corrector eat the suffix.
- Codex context resolver drops vendor/ namespaces, so
openai/gpt-5.6-sol-900k resolves to 900K like the bare id.
- Table-driven regression covering eligible bases/snapshots/namespaced
ids and rejected -pro/-mini/5.5/unknown aliases, asserting context AND
wire model.
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the
live-verified 900K burned through subscription usage for users who never
asked for the larger window (bigger window = more input tokens per request).
- Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the
advertised 272K again — the cheaper limit is the default.
- The model picker synthesizes explicit <slug>-900k variants (e.g.
gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into
the 900K window. Slugs that genuinely enforce 272K (gpt-5.5,
gpt-5.4-mini) get no variant.
- The -900k suffix is Hermes-side only: stripped before the model id hits
the wire (main transport + auxiliary Responses adapter), and pricing
aliases the variants onto the base entries.
- Docs: new opt-in section in context-compression-and-caching.md.
GLM-5.3 is live on api.z.ai (coding plan endpoint) but had no entries in
Hermes, so it silently fell back to the generic 202K GLM context —
triggering premature context compression on a 1M-window model.
- model_metadata: 'glm-5.3': 1_048_576 (same base model as 5.2; 1M
context / 128K max output per docs.z.ai/guides/llm/glm-5.3, verified
2026-08-14)
- auth: add glm-5.3 to coding-plan probe lists (global + CN)
- models: add glm-5.3 to picker/model lists (6 sites)
- zai provider: reasoning_effort mapping covers glm-5.3 (accepted live
by the endpoint, HTTP 200)
First actioned report from the overhauled model-catalog-scout cron
(2026-08-21 validation run), every item re-verified live before edit:
Delisted (gone from live catalogs):
- opencode-zen curated: claude-opus-4-1, qwen3.7-max, qwen3.7-plus
(absent from live zen /v1/models; qwen3.7 family remains on Go)
- OPENROUTER_MODELS free section: poolside/laguna-m.1:free (rotated to
s-2.1/xs-2.1), tencent/hy3:free, inclusionai/ring-2.6-1t:free
Added (present + verified in live catalogs):
- OpenRouter free: z-ai/glm-5.2:free (256K), poolside/laguna-s-2.1:free
+ laguna-xs-2.1:free (262K), nvidia/nemotron-3.5-lightning:free (1M)
- opencode-go curated: ox-alpha-free (Go-subscription twin of the Zen
keyless Ox Alpha; keyed — Go relay 401s anonymous requests)
Metadata:
- DEFAULT_CONTEXT_LENGTHS: laguna-s-2.1/xs-2.1 262144;
nemotron-3.5-lightning 1M (overrides the generic 131K nemotron entry);
glm-5.2:free 256K (the free variant is capped below the 1M paid entry)
Keyless-heal hardening (the real find):
- opencode_zen_free_runtime now gates the zen/go→keyless heal on
MEMBERSHIP in the verified opencode-free catalog, not the -free
suffix — ox-alpha-free is a KEYED Go model despite its suffix, and
suffix-based healing would have routed it to a Zen relay that
doesn't serve it (verified: zen 401s 'not supported', go 401s
'Missing API key'). New regression test pins this.
Fixture sweep: tencent/hy3:free catalog assertion updated (delisted
slug); nous-route fixtures using hy3:free as incidental model names
left alone (self-consistent mocks). model-catalog.json regenerated.
Adds OpenRouter's free "Ox Alpha" stealth reasoning model
(stealth/ox-alpha) to the OpenRouter fallback snapshot, plus the
provider-agnostic metadata it needs:
- OPENROUTER_MODELS: free-tier entry (1M ctx)
- DEFAULT_CONTEXT_LENGTHS: ox-alpha -> 1,048,576 (verified against
OpenRouter live /api/v1/models; without this the slug fell through
to no match)
- reasoning_timeouts.py: 300s stale floor for ox-alpha and the
OpenCode Zen twin slug x-preview-f-free (reasoning model,
long-horizon agentic work per its model card)
- model-catalog.json regenerated
Pricing snapshot skipped: openrouter bills via official_models_api
(live pricing; model is free anyway).
- Add x-preview-f-free (Ox Alpha: free, 1M context, ZDR) plus all newly
listed Zen models (gpt-5.6 sol/terra/luna, claude-opus-5, gemini-3.7/3.6
flash + lite, grok-4.6/4.5, muse-spark-1.2, kimi-k3, qwen3.7-max,
hy3-free, laguna-s-2.1-free, nemotron-3.5-lightning-free,
muse-spark-1.2-contributor-free) and Go models (gpt-5.6-luna, grok-4.5,
glm-5.3, qwen3.8-max, hy3, hy3-preview, muse-spark-1.2-contributor).
- Drop delisted north-mini-code-free from Zen.
- Route grok-* on Zen and Go through /v1/responses per the published
endpoint tables (grok-4.6/4.5/build-0.1 on Zen, grok-4.5 on Go).
- 1M context fallback for x-preview-f (Ox Alpha).
- Refresh hermes setup provider samples for both providers.
Catalogs verified against live GET /zen/v1/models and /zen/go/v1/models
plus https://opencode.ai/docs/zen/ and /docs/go/ endpoint tables (2026-08-20).
Recognizes the DeepSeek/OpenAI-compatible relay wording
max_tokens (98304) exceeds model's maximum output tokens (65536)
in both parse_available_output_tokens_from_error (returns the cap) and
is_output_cap_error (keeps the 400 out of the compression death-loop).
Salvaged from PR #72283; the conversation_loop early-clamp block was
dropped in favor of routing through the existing output-cap handler
(follow-up commit).
Fixes the retry loop that spins forever when a vLLM server rejects a
request for having a max_tokens too big for what is left of the context
window.
The catch is that vLLM does not tell you how big your prompt actually is
in that situation. It works the number backwards from the constraint it
just failed, so you get:
"requested 65536 output tokens and your prompt contains at least
36865 input tokens, for a total of at least 102401 tokens"
That 36865 is just window + 1 - requested, and the total is always
exactly window + 1. Subtracting it from the window hands back
requested - 1 every single time, whatever the real prompt size is.
parse_available_output_tokens_from_error believed it and returned
requested - 1. conversation_loop then takes off its 64 token safety
margin and retries, which walks the cap down 65 tokens at a time while
the reported input walks up by the same 65:
65536 -> 65471 -> 65406 -> 65341
Three attempts is the default budget, so the session gives up with
"Context length exceeded" having closed 195 tokens of a roughly 28000
token gap. Compression cannot save it either, because the input was
never the problem, which is why the compressor keeps refusing with
"summary would have GROWN".
This is also what is behind the unexplained "input-token drift" in
issue #61761. The input is not drifting. It is a derived number, and it
moves because we moved max_tokens.
So when that shape shows up (the "at least" wording, plus a budget that
works out to exactly requested - 1), halve the requested cap instead. It
is still guaranteed to sit under whatever was just rejected, and it
converges on the first retry: 65536 -> 32768, which next to a real 36865
token prompt comes to 69633 against a 102400 window.
Nothing else moves. A measured input is still trusted, and a genuine
input overflow still returns None so the caller falls through to
compression the way it always did.
The existing test asserted the bogus 65535, so it is updated. Added
tests for the measured input path, and for the retry actually
converging.
OpenAI enabled the large-context window for ChatGPT-subscription Codex
accounts (announced by @thsottiaux Aug 16 2026; previously API-key-only).
Live re-probe the same day: 911,276 input tokens completed OK on
gpt-5.6-sol; ~925K+ rejected with context_length_exceeded (1.05M window
minus reserved output headroom). terra, luna, and gpt-5.4 all completed
900,026 tokens OK. The Codex catalog still advertises 272K, so the
stale-advertisement override from #87981 is the right lever — this just
raises its value 350K -> 900K.
gpt-5.5 and gpt-5.4-mini still enforce 272K live (rejected 500K) and
remain excluded. Override semantics unchanged: fires only on an
exactly-272,000 advertisement; any live catalog change is trusted
verbatim.
The Codex /models catalog advertises 272K for the gpt-5.6 (sol/terra/luna)
and gpt-5.4 slugs, but the backend actually accepts ~371K input tokens
(verified live against chatgpt.com/backend-api/codex/responses, Aug 16 2026:
~371K completed OK on all four slugs; ~382K+ rejected with
context_length_exceeded). 350K keeps ~22K margin under the observed ~372K
enforcement.
The bump applies ONLY when the resolved value is exactly the known-stale
272,000 advertisement — any other advertised value (higher or lower) is
trusted as a real server-side change, so a future catalog correction
deactivates the override automatically. gpt-5.5 and gpt-5.4-mini both
genuinely enforce 272K (rejected 360K live) and are excluded.
Per-provider ssl_ca_cert / ssl_verify reached the httpx chat client and the
auxiliary clients (#56681), but the endpoint discovery and pricing probes did
not. Both probe families resolved TLS from process-wide env vars only:
- the requests-based metadata/pricing probe
(agent/model_metadata.py::_resolve_requests_verify)
- the urllib-based /models catalog probe
(hermes_cli/models.py::probe_api_models)
A custom endpoint whose chain verifies against the provider's configured
bundle, but not the process SSL_CERT_FILE, then logged a spurious
CERTIFICATE_VERIFY_FAILED on every probe even though the chat path worked.
Pointing a global CA env var at the bundle fixes it but changes verification
for every provider, defeating the point of a per-provider setting.
This threads the selected provider's TLS settings into both probe paths,
reusing get_custom_provider_tls_settings so there is no second precedence
chain:
- _resolve_requests_verify(base_url) looks up the provider's ssl_verify /
ssl_ca_cert before falling back to the env vars. Callers with no base_url
keep the exact env-only behavior.
- probe_api_models builds an ssl.SSLContext from the provider settings and
passes it through open_credentialed_url, which gains an ssl_context seam on
the cloned secure opener. Unmatched or public endpoints pass None and keep
urllib's default policy.
Tests: tests/agent/test_custom_provider_ca_probes.py covers both probe
families (provider CA, ssl_verify:false, unmatched, missing file, config
lookup failure) plus end-to-end assertions that the resolved verify value and
SSLContext actually reach the request seam. Verified against the neighboring
metadata, pricing, TLS, and urllib-security suites (266 tests) with no
regressions.
Review follow-ups on the model_overrides feature:
- ONE canonical override schema everywhere. get_model_info previously
merged the override dict raw into the models.dev catalog shape
({**raw, **override}), so the documented context_window/supports_*
keys silently did nothing on that path (cost guard, inventory) while
working in capabilities/context paths — same config key, two
incompatible schemas. Overrides are now translated into the catalog
shape at the get_model_info boundary (_override_to_catalog_shape),
and sub-dicts (limit, modalities) are MERGED, not clobbered — an
override setting only context_window no longer wipes the catalog's
limit.output.
- _default is now a FILL-GAP default, not an override: it applies only
to models the catalog does not know (the #8731/#84482 self-unblock
path) and never displaces catalog data. A
_default: {context_window: 128000} can no longer clamp every model
of a provider. Explicit per-provider+model entries keep their
win-over-catalog semantics.
- Early-chain _override_context_window (model_metadata step 0b) is
explicit-only, so a _default can never preempt custom_providers
per-model settings or live probes; fill-gap defaults apply at the
lookup_models_dev_context catalog-miss boundary (step 5f) instead.
This fixes the precedence inversion where a provider/global _default
silently overrode an explicit per-endpoint per-model context_length.
- Provider keys accept BOTH id spaces (Hermes id and models.dev id:
copilot/github-copilot both work) and model ids match
case-insensitively, mirroring catalog lookup.
- Malformed override values (context_window: '512k') log a one-shot
warning instead of being silently swallowed.
- DEFAULT_CONFIG comment: removed the false family/dated-snapshot
inheritance claim, documented the recognized field list, fill-gap
semantics, and the id-space rule.
- Tests: rewritten for the new contracts (fill-gap invariants,
dual-id-space keys, sub-dict merge preservation, one-shot warning);
added a real-config-yaml e2e plumbing test (mutation-checked: fails
when the config key wiring is broken).
Add a unified model_overrides config section that lets users manually
declare context_window, max_output_tokens, capabilities, cost, and
family for any provider+model — winning over models.dev, OpenRouter, and
hardcoded defaults.
Resolution order (first hit wins):
1. model_overrides.<provider>.<model_id> (per-provider+model)
2. model_overrides.<provider>._default (per-provider default)
3. model_overrides._default (global default)
4. Normal catalog resolution
Key subtlety: an unknown model id (not in the
catalog) derives base metadata from sensible defaults before patching,
so overriding a model the catalog doesn't know yet is the supported
self-unblock path. This is exactly the #84482 scenario (Upstage
solar-pro4/syn-pro wrong context) and the #8731 scenario (custom/local
models with manual capability declaration).
Wired into:
- get_model_capabilities() — patches capability fields; unknown models
get safe defaults (tools on, vision/reasoning off) before patching
- lookup_models_dev_context() — context_window override, checked before
catalog lookup so it works even for providers not in PROVIDER_TO_MODELS_DEV
- get_model_info() — merges override dict onto catalog entry (shallow
merge); for unknown models, the override is the sole source of metadata
- get_model_context_length() — step 0b in the resolution pipeline,
before custom_providers (0c) and before any network probe
Config example:
model_overrides:
upstage:
solar-pro4:
context_window: 524288
syn-pro:
context_window: 65536
custom:my-local-vllm:
my-llava-model:
context_window: 8192
supports_vision: true
supports_reasoning: false
supports_tools: true
_default:
context_window: 128000
Fixes#8731Fixes#84482
Refs #47247
save_context_length() and _invalidate_cached_context_length() did an
unguarded read-modify-write into $HERMES_HOME/context_length_cache.yaml.
The plain `open(path, "w")` truncates the file before the dump runs. If
the process is killed mid-dump, the file is left empty or partial. The
next _load_context_cache() swallows the YAML error and returns {} —
silently wiping every persisted context length. A concurrent process
reading between truncate and dump-complete also sees a torn file.
After the cache is lost, every model re-probes the network, and when a
probe fails it falls back to the generic 256K default — so a user on a
1M-window model ends up with a wrong, short context window.
Hermes routinely runs several processes against one shared $HERMES_HOME
(a cron agent plus an interactive session, multiple gateway sessions),
so this is hit in normal use.
Switch both writers to the existing utils.atomic_yaml_write helper
(temp file + fsync + os.replace, symlink- and mode-preserving). The real
file is only ever swapped from a fully written temp file, so an
interrupted write leaves the previous cache intact and readers never see
a partial file. Matches the atomic-write pattern already used for
auth.json, config.yaml, and other persisted state.
Makes the persistent model context-length cache write crash-safe. The
old non-atomic write could truncate or wipe the entire cache on an
interrupted or concurrent write, which then forces models onto the wrong
fallback context window. The fix routes both cache writers through the
repo's atomic temp-file + os.replace helper.
N/A
- [x] 🐛 Bug fix (non-breaking change that fixes an issue)
- [ ] ✨ New feature (non-breaking change that adds functionality)
- [ ] 🔒 Security fix
- [ ] 📝 Documentation update
- [ ] ✅ Tests (adding or improving test coverage)
- [ ] ♻️ Refactor (no behavior change)
- [ ] 🎯 New skill (bundled or hub)
- `agent/model_metadata.py`: `save_context_length()` and
`_invalidate_cached_context_length()` now write via
`utils.atomic_yaml_write` instead of a truncating `open(path, "w")`.
Added the `atomic_yaml_write` import.
- `tests/agent/test_model_metadata.py`: added
`test_write_failure_leaves_existing_cache_intact` — simulates a crash
during the atomic swap and asserts the existing cache survives
byte-for-byte with no stray temp file.
1. `pytest tests/agent/test_model_metadata.py -q` — 98 pass, including
the new crash-safety test.
2. The new test seeds a valid cache, forces the swap step to raise, and
confirms the file is not truncated and no `.cache_*.tmp` is left.
3. `ruff check agent/model_metadata.py` passes.
- [x] I've read the Contributing Guide
- [x] My commit messages follow Conventional Commits (`fix(scope):`, etc.)
- [x] I searched for existing PRs to make sure this isn't a duplicate
- [x] My PR contains **only** changes related to this fix
- [x] I've run the affected tests (`pytest tests/agent/test_model_metadata.py -q`) and they pass
- [x] I've added tests for my changes
- [x] I've tested on my platform: macOS 15 (Darwin 25.5)
- [x] I've updated relevant documentation (README, `docs/`, docstrings) — or N/A
- [x] I've updated `cli-config.yaml.example` if I added/changed config keys — or N/A
- [x] I've updated `CONTRIBUTING.md` or `AGENTS.md` if I changed architecture or workflows — or N/A
- [x] I've considered cross-platform impact (Windows, macOS) — the helper uses os.replace, which is atomic on both
- [x] I've updated tool descriptions/schemas if I changed tool behavior — or N/A
Reapply the non-positive context-length guards onto the post-history-replacement
mainline without carrying any stale branch history. save_context_length() now
refuses to persist length <= 0 (keeping upstream's normalized _context_cache_key),
and get_model_context_length() drops non-positive cache hits at the head of the
invalidation chain (Codex/Kimi/MiniMax/Grok branches become elif) so a poisoned
entry re-resolves instead of short-circuiting to 0.
Refresh of PR #25812; original head d62ed5eb92f057d8c707ba937b44f168f2df0677.
An empty/blank model id reaching get_model_context_length() can't be
meaningfully resolved — and it's worse than a miss: the endpoint
metadata fuzzy matcher ('model in key or key in model') is vacuously
true for "", so it matches an ARBITRARY catalog entry from the live
/v1/models response and returns whatever context length that entry
happens to have, persisting it under a junk '@<base_url>' cache key.
This started failing CI on main when the Nous portal catalog changed:
tests/run_agent/test_primary_runtime_restore.py constructs agents with
model='' against the live portal URL, the arbitrary match now lands on
a 32K entry, and init_agent raises the 64K-floor ValueError
(test_allowed_for_nous_anthropic_messages, red on every PR's slice).
Guard early: a blank model id falls back to DEFAULT_FALLBACK_CONTEXT
immediately, before any cache write or network probe.
Salvaged from #65515 by @whirmill (rebased onto current main; the
guard now sits after the malformed-base_url normalization added since,
and carries an explanatory comment for the fuzzy-match footgun).
Fixes the red slice on #85444, #85452 and every other open PR.
Co-authored-by: whirmill <5079591+whirmill@users.noreply.github.com>
Replaces the per-model _model_name_suggests_grok_4_3/_grok_4_6/
_minimax_m3 stale-cache predicates with one generic
_stale_pre_catalog_cache_entry() guard driven by
_PRE_CATALOG_STALE_KEYS. A cached context length is dropped when the
model resolves (longest-key-first, same as step 8) to a listed catalog
key and the cached value is at or below what the old resolution path
could have produced (largest shorter matching catch-all, or the 256K
fallback).
Also covers qwen3.6-plus, grok-4-fast, and grok-4.20 (the models
PR #37684 requested guards for), absorbing that PR.
_model_name_suggests_minimax_m3 is kept for its two non-cache callers
(models.dev underreport guard, cache-control gating in
agent_runtime_helpers).
docs.x.ai (2026-08-12): grok-4.6 is the flagship, 500K context.
Live GET /v1/models lists grok-4.6 at context_length 500000
(no grok-4.6-latest alias).
#84661 landed the catalog. Main already lists native grok-4.6
on the xAI picker. This is only the leftover cache guard
(same pattern as grok-4.3): pre-catalog builds persisted the
grok-4 catch-all (256K).
'' is a substring of every catalog key, so _resolve_endpoint_context_length
with an empty model name "matched" whatever the endpoint listed first —
on the Nous portal that is currently a 32K embedding model, which poisoned
the resolved context length and made AIAgent init fail the 64K minimum.
This is what turned tests/run_agent/test_primary_runtime_restore.py::
TestTryRecoverPrimaryTransport::test_allowed_for_nous_anthropic_messages
red on every PR (CI slice 7/12) after the portal catalog reordered.
Single-model endpoints still resolve with an empty name (unambiguous);
non-empty names keep the substring fuzzy match.
_PROVIDER_PREFIXES was a hand-maintained frozenset, so providers that ship
as plugins (bundled like fireworks, or user plugins under
$HERMES_HOME/plugins/model-providers/) were never recognised as
provider: prefixes in model strings, and metadata/context-window lookups
received the unstripped string. Mirror the _URL_TO_PROVIDER auto-extend
that already sits below it: add each registered profile's name and
aliases after discovery. The _OLLAMA_TAG_PATTERN guard keeps model:tag
strings intact.
Fixes#66106
Qwen3.8 Max is live on both OpenRouter and the Nous portal
(qwen/qwen3.8-max, 1M context, 131K max output). Per the
newest-max-replaces-last-max convention, it takes qwen3.7-max's slot
in both curated lists.
- hermes_cli/models.py: OPENROUTER_MODELS + _PROVIDER_MODELS[nous]
swap qwen/qwen3.7-max -> qwen/qwen3.8-max
- agent/model_metadata.py: DEFAULT_CONTEXT_LENGTHS entry for
qwen3.8-max at 1,000,000 (verified against OpenRouter live
metadata and Nous /v1/models 2026-08-03)
- tests/test_empty_model_fallback.py: swap incidental catalog fixture
to the surviving slug
- website/static/api/model-catalog.json: regenerated
Pricing snapshot skipped: both routes bill via official_models_api
(live pricing), verified with resolve_billing_route. Reasoning
timeout floor already covered by the qwen3 prefix (180s).
Salvage of #71282 (Fixes#71281): a routable-but-dead endpoint (corp
LAN address while off-VPN) blackholes TCP SYNs, so every probe in the
model-metadata waterfall waits out its full connect timeout — 20+
seconds of stall per startup across detect_local_server_type,
fetch_endpoint_model_metadata, and the per-model probes.
A module-level blackhole cache keyed on host:port is populated when
any probe observes a ConnectTimeout (httpx or requests; read timeouts
deliberately excluded — an accepted connection is not a blackhole) and
consulted at the top of each guarded function. 30s TTL: long enough to
collapse one startup burst, short enough that VPN recovery is picked
up without a restart. Guard ordering: blackhole check -> disk L2 ->
HTTP waterfall, and a blackholed leg aborts the remaining legs instead
of letting each stall in turn.
Squash of the PR's two real commits (the branch's merge commits made
it un-rebase-merge-able; content verified identical via merge-tree).
CI slice 3/7 failures: run_conversation tests pass MagicMock base_urls
through the metadata probe path; re.sub raised TypeError where the old
code let non-strings flow through. Preserve that contract.
fetch_endpoint_model_metadata's generic (non-LM-Studio) /models fetch and
its llama.cpp /v1/props context-length follow-up built request URLs
straight from the unrewritten candidate, unlike every other local-probe
site. Both retained the multi-second dual-stack IPv6 connect penalty
that _localhost_to_ipv4() exists to skip (measured on macOS: localhost
32.9ms vs 127.0.0.1 0.1ms on a dead port; ~2s on Windows). normalized
stays the cache key so caching behavior is unchanged; only the outbound
request target is rewritten.
Re-derived from PR #61528 onto current main (original no longer applied
cleanly).
- Short-circuit the candidate waterfall on HTTP 401/403: an auth wall
proves the endpoint family exists, so probing the alternate URL just
doubles the wasted wait (the reported endpoint takes ~10s to return
401 without a key).
- Stream the probe so 4xx never downloads a slow error body; responses
are closed on every exit path.
- Regression tests: single-call assertion on 401/403 (fails on main),
negative-cache reuse, 404 waterfall preserved, no .json() on 4xx.
Fixes#69905
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Review pass 2 (reuse reviewer HIGH): the step-3b probe-down fallback for
custom/local endpoints returns the same silent 256K default but only
logged at INFO - invisible by default, and it is the MORE common path
for small local models (the exact users the warning exists for).
Extract _warn_context_length_fallback() (deduped per model+base_url)
and call it from both fallback sites, per the fix-the-whole-bug-class
rule. Regression test drives the custom-endpoint path and fails without
the widening (mutation-checked).
Review follow-up:
- Warn once per (model, base_url) at the step-9 fallback via a module-level
dedup set (established _WARNED_* idiom). The fallback result is
deliberately never cached, so the un-deduped warning fired on every
resolution - e.g. once per gateway message via the session-hygiene path.
- Replace the three inline-mock pool-cleanup tests (which reproduced the
try/except block against a MagicMock and passed even with the production
code reverted) with a parametrized test that drives the real
BatchRunner.run() with a patched Pool; drop the CPython stdlib
signature change-detector test.
- Add a once-per-model warning regression test; clean up dead imports.
All tests verified to fail against pre-PR batch_runner.py/model_metadata.py
and pass with the fix (mutation check).
Salvage of #6629 by aaronlab (kshitijk4poor reworked against current main).
Three concerns from the original PR, reworked to address review feedback:
1. Context-length fallback diagnostic (agent/model_metadata.py):
get_model_context_length() silently returned 256K when all 9 detection
methods failed. Users with small-context models (8K, 32K) would get 256K
silently, causing hard-to-debug API context-length errors. Added a
warning log at the step 9 fallback with model name, base_url, and the
correct config override hint (model.context_length, not context_length).
The token-estimation ceiling-division fix from the original PR already
landed on main (5c2ecdec) with CJK handling — not duplicated here.
2. Fsync for batch trajectory writes (batch_runner.py):
Trajectory entries were written without flush/fsync, but the checkpoint
immediately marked them as completed. A crash between write and disk
sync would leave the checkpoint claiming completion with no trajectory
data on disk. Added flush() + os.fsync() before checkpoint update.
3. Pool cleanup on interruption (batch_runner.py):
Ctrl+C during pool.imap_unordered() relied on context manager cleanup
which can hang. Added explicit pool.terminate() + pool.join() for both
KeyboardInterrupt and Exception paths. The original PR used
pool.join(timeout=10) which is invalid — CPython's Pool.join() takes
no timeout parameter. Fixed to use pool.join() without arguments.
Tests:
- test_warning_emitted_on_fallback: verifies warning fires at step 9
- test_no_warning_when_cached: verifies no false warning when cache hits
- test_trajectory_entry_is_synced_to_disk: verifies os.fsync is called
- test_pool_terminate_called_on_exception: verifies cleanup on RuntimeError
- test_pool_terminate_called_on_keyboard_interrupt: verifies cleanup on Ctrl+C
- test_pool_join_called_without_timeout: verifies no timeout arg to join()
- test_real_pool_join_accepts_no_timeout: integration check on CPython API
Co-authored-by: Aaron Lab <aaronlab@users.noreply.github.com>
The reasoning_details field (OpenRouter/Anthropic thinking blocks +
opaque cryptographic signature blobs) inflates the rough token estimate
by ~4x. Providers do not bill these envelope bytes as prompt tokens.
In a measured Kimi K3 session, reasoning_details held 2,124K chars
vs 281K chars of actual thinking text. The estimator reported ~533K
tokens when real prompt_tokens was ~140K — triggering compression at
~27% of the configured threshold.
Fix: skip reasoning_details in both _estimate_message_chars and
_estimate_message_tokens_without_images, alongside the existing
_anthropic_content_blocks exclusion.
Fixes#73298
Review follow-up on #75102. The shadow substituted the sidecar whenever
the ``api_content`` key was merely PRESENT, but the wire only substitutes
a non-empty string sidecar on a user/assistant row (see
``turn_context.substitute_api_content``). For any other shape the sidecar
is popped and discarded while the clean ``content`` is sent -- so the
shadow dropped real content from the estimate and UNDERcounted, the
dangerous direction: compaction fires too late and the turn dies on a
hard context-length error instead of merely compressing early.
Gate the substitution on the same predicate, and cover the divergent
shapes (None, empty string, int, list, non-user/assistant role) with a
test that fails against the unconditional version.
Also rename the image test: it never carried a sidecar, so it was not
testing what its name claimed. It is a non-regression pin on the flat
per-image accounting that moved into ``_wire_message_shadow()``, and is
now named for that.