Follow-up to the #106769 pick: the 8-line dated bugfix narrative in
parse_context_limit_from_error becomes a 2-line WHY, and one invariant
test pins the contract: an output-cap message parses to None as a
context limit, 16384 as the available output tokens, and classifies as
an output-cap error, while a genuine "maximum context length" message
still parses. Red on origin/main (returned 16384 as the context limit).
`:nitro`, `:floor`, `:exacto`, and `:online` are request-time routing
modifiers, not catalog models — OpenRouter's /models lists only the base
id, and a variant runs the same model with the same context window.
`get_model_context_length()` keyed every lookup on the full suffixed id,
so each one missed and the resolver fell through to a generic family
default or the 256K fallback:
openai/gpt-5.5:nitro -> 256K (real 1.05M)
x-ai/grok-4.6:nitro -> 131K (generic "grok" catch-all)
anthropic/claude-opus-4.6:nitro-> 200K (generic "claude" catch-all)
The window silently shrank, triggering early compression and a wrong
/usage readout. f14059fa fixed the sibling half of this bug class in
/model validation; this fixes the metadata half.
Strip a recognized variant suffix for LOOKUP only, keeping the suffixed
id on the wire so the routing opt-in survives. Applied after the explicit
config overrides (steps 0b/0c) so a user-pinned value still wins, and
before every cache/catalog lookup. Gated on the request actually routing
through OpenRouter, so a local Ollama `model:tag` is untouched.
`:free`/`:batch`/`:thinking` are deliberately excluded — those ARE
distinct catalog SKUs with their own windows, so stripping them would
report the wrong number.
The suffix set and base-id split move to hermes_constants (import-safe,
dependency-free) so the metadata layer shares one definition with
hermes_cli.models instead of duplicating it.
Users on native DeepSeek were told to pin model.context_length and
model.supports_vision in config.yaml. That is the wrong layer: the 1M
window is already in DEFAULT_CONTEXT_LENGTHS, and a global
supports_vision pin would also mark text-only deepseek-v4-pro as
multimodal.
Two catalog gaps still produced the reported symptoms:
- A leftover context_length_cache.yaml entry of 128K (the old
``deepseek`` catch-all) outlived the 1M catalog keys because
deepseek-flash was missing from _PRE_CATALOG_STALE_KEYS.
- When models.dev is empty/cold, Flash has no capability record, so
image routing falls through to lossy text. Vendor docs (2026-09-10)
mark deepseek-flash as vision-capable and deepseek-v4-pro as not.
Discard those 128K leftovers, fill Flash (and retired Flash aliases)
via _BUILTIN_MODEL_METADATA, and leave Pro catalog-only.
DeepSeek's 2026-09 Flash refresh introduced a version-less canonical id:
GET /v1/models now returns `deepseek-flash` (alongside `deepseek-v4-pro`), the
API accepts it directly, and the older `deepseek-v4-flash` is server-side
aliased onto it. Every DeepSeek model-id gate in Hermes keys off the
`deepseek-v<N>` prefix, so the new id silently missed all four:
* DeepSeekProfile.build_api_kwargs_extras classified it as non-thinking and
omitted `extra_body.thinking`. The server then defaults to thinking-on, so
the user's thinking toggle and `reasoning_effort` were quietly ignored.
* `_normalize_for_deepseek` folded it onto `deepseek-v4-flash` (it misses the
V-series regex), so the id a user picked never reached the wire and the
config stored a different model than the picker advertised.
* `DEFAULT_CONTEXT_LENGTHS` fell through to the 128K `deepseek` catch-all
instead of the real 1M window, capping the model at an eighth of its
context before compaction kicked in.
* `_REASONING_STALE_TIMEOUT_FLOORS` had no entry, leaving the stale-stream
detector at its 180s default instead of the 600s reasoning-model floor.
Verified live against api.deepseek.com: `deepseek-flash` answers 200 with
`model: deepseek-flash`, accepts image input (the refresh folds vision into
the Flash model), and the in-between id `deepseek-v4.1-flash` is rejected
with "The supported API model names are deepseek-flash, deepseek-v4-pro".
Adds the id to all four gates plus regression coverage for each site.
Every GLM slug that missed a DEFAULT_CONTEXT_LENGTHS key fell to the "glm" 202,752 catch-all, so
compression fired at ~20% of the real window and the aux feasibility check auto-lowered the session
threshold to that number (the 202,752 in the Coatue compression report).
- Catalog: GLM entries from Nous + OpenRouter /v1/models (2026-09-09): 5.3 / 5.3-flash 1,310,720
(:batch/:US 1,048,576), 5.2 1,048,576, 5 / 5.1 / 4.7 / 4.6 204,800, *-turbo / 4.7-flash 202,752.
- _catalog_key_matches: version separators normalised on both sides so relay slugs like z-ai-glm-5-3
hit glm-5.3 (#97398; approach from PR #97412 by @shellybotmoyer, re-implemented on current main).
- get_custom_provider_context_length: entry-level custom_providers[].context_length backs every model the
entry serves when no per-model override exists, so /model switch stops dropping it (#98387; from
PR #98396 by @liuhao1024, re-implemented on the config_providers sibling).
- check_compression_model_feasibility: when aux compression is the main model on the main route, reuse the
main model's resolved window instead of re-resolving without its pin (#89500 mechanism 2, #45519).
Closes#97398, #98387, #89500, #87825, #97595, #97820 (suffix strip already on main; catalog values now match).
run_agent.py: delete the `# noqa: F401` re-export block (agent.process_bootstrap
OpenAI/_SafeWriter/_get_proxy_*, model_tools get_tool_definitions/
handle_function_call/check_toolset_requirements, FailoverReason,
_qwen_portal_headers/_routermint_headers, session_persistence names,
estimate_request_tokens_rough, ContextCompressor + friends, jittered_backoff,
prompt_builder names, message_sanitization names, tool_dispatch_helpers
names) — 41 names run_agent never used itself — and the `_STREAM_DIAG_HEADERS`
back-compat class alias (no in-tree reader). run_agent now imports only what
it uses (get_toolset_for_tool, is_local_endpoint, coalesce/uniquify tool-call
ids, cleanup_vm/get_active_env from terminal_tool_lifecycle).
agent/*: `_ra().X` late-binds that only reached a re-export now import the
defining module directly (agent_runtime_helpers -> process_bootstrap.OpenAI,
model_tools.handle_function_call, session_persistence._safe_session_filename_component;
agent_init -> model_tools.get_tool_definitions/check_toolset_requirements,
_lazy_headers("agent.client_lifecycle", ...) for qwen/routermint;
system_prompt -> agent.prompt_builder / model_tools directly, dropping its
own _ra() shim and the `_r` parameter threading). `_ra()` stays for
run_agent-resident names (logger, AIAgent, _hermes_home, _set_interrupt, ...).
toolsets.py: remove resolve_multiple_toolsets (shim-only, restored by
34abf954bd); tests/test_toolsets.py pins the same union behavior via
resolve_toolset over each name.
providers/__init__.py: drop the OMIT_TEMPERATURE re-export (no callers via the
package); ProviderProfile stays because __init__ uses it for annotations —
2 tests repointed to providers.base.
agent/iteration_budget.py: drop the "run_agent re-exports the class"
docstring pointer; 4 tests import IterationBudget from its home.
model_tools.py (arg_coercion names), agent/tool_executor.py, and
hermes_cli/cli_session_mixin.py repoints landed via a sibling commit on this
shared worktree.
Callers repointed: gateway/run.py, hermes_cli/cli_chat_turn_mixin.py,
hermes_cli/cli_tui_mixin.py, tui_gateway/session_workdir.py,
agent/transports/codex.py (one-line imports) + comment pointers in
tools/file_state.py, tools/schema_sanitizer.py, scripts/tool_search_livetest.py.
Tests: patch("run_agent.X") / monkeypatch.setattr(run_agent, "X") /
`from run_agent import X` -> defining module across 99 test files.
opencode-free had no PROVIDER_TO_MODELS_DEV entry, so every models.dev
lookup on the free tier missed and Muse Spark fell to the 256K default.
The free tier is served by the Zen relay (hermes_cli/models.py:
"opencode-free is Zen-hosted"), and models.dev's "opencode" provider is
the catalog that lists muse-spark-1.2 / -1.2-contributor-free /
-1.3-contributor-free at 1,048,576 — so the alias is "opencode", not
"opencode-go" (Go's catalog carries only the paid -contributor SKUs).
Missing alias identified by @Steve-prog001 in #101905.
Tests: one parametrized offline invariant (models.dev + live /models
mocked away) asserting 1,048,576 on opencode-free / opencode-go /
meta-ai / commandcode — fails on main, passes here — plus the alias pin.
Every bot-to-bot DM is a fresh `hermes -p <bot> chat -Q` process, so it
pays agent startup on each hop. Profiling one hop showed the single
largest controllable cost was a live GET /models against the provider on
EVERY launch (0.3-0.6s normally, up to the 15s probe timeout on a slow
endpoint) — the in-memory endpoint-metadata cache is per process and the
Nous persistent context cache is bypassed by design so the portal stays
authoritative.
- model_metadata: memoize successful remote /models probes on disk
(cache/endpoint_model_metadata.json) with the SAME 300s TTL as the
in-memory cache, so authority semantics are unchanged (reconciliation
still lands within 5 minutes) but the answer is shared across
processes. Local endpoints are never memoized (LM Studio reloads).
- bot_relay: the cross-machine reply waiter polls the reply file every
250ms instead of every 2s — up to 2s of dead air on every relayed reply.
Nothing here changes turn ordering: DMs and group rounds stay serial.
Live (polis-hermes bot, spawn -> first API request, cold, 5-6 runs):
main median 1.23s (one 20.8s outlier = probe stall) -> 0.96s, no stalls.
Google Gemini/Gemma overflow errors read 'Unable to submit request because
the input token count is 32825 but model only supports up to 32768'.
parse_context_limit_from_error had no pattern for the 'supports up to N'
phrasing, so overflow recovery kept the wrong window and burned its retry
attempts instead of recalibrating to the provider-reported limit.
Add the anchored pattern (limit follows 'supports up to'; the larger input
count before it is never captured) plus regression tests covering the exact
message and the get_context_length_from_provider_error recalibration path.
Reported by @Artemonim in #57275 (residual claim 5).
Review findings on #92797 (@100yenadmin):
- is_codex_900k_base() is now the single source of truth used by picker
synthesis, context resolution, /model validation, and wire stripping.
Eligibility is an exact table (sol/terra/luna, gpt-5.4, daybreak alias)
plus date-shaped 5.6 snapshots — family-prefix matching removed, so
non-routable -pro slugs and unknown descendants never gain variants.
- strip_codex_context_variant_suffix() strips conditionally: ineligible
aliases (gpt-5.5-900k) are returned unchanged and fail honestly at the
API instead of silently running as the base model at 272K.
- validate_requested_model() rejects ineligible *-900k aliases before the
hidden-slug soft-accept, and accepts valid variants missing from a
stale catalog without letting the typo auto-corrector eat the suffix.
- Codex context resolver drops vendor/ namespaces, so
openai/gpt-5.6-sol-900k resolves to 900K like the bare id.
- Table-driven regression covering eligible bases/snapshots/namespaced
ids and rejected -pro/-mini/5.5/unknown aliases, asserting context AND
wire model.
The Aug 16 change that auto-raised gpt-5.4/5.6 Codex OAuth context to the
live-verified 900K burned through subscription usage for users who never
asked for the larger window (bigger window = more input tokens per request).
- Base Codex slugs (gpt-5.6-sol/terra/luna, gpt-5.4) now resolve to the
advertised 272K again — the cheaper limit is the default.
- The model picker synthesizes explicit <slug>-900k variants (e.g.
gpt-5.6-sol-900k) for every live-verified slug; selecting one opts into
the 900K window. Slugs that genuinely enforce 272K (gpt-5.5,
gpt-5.4-mini) get no variant.
- The -900k suffix is Hermes-side only: stripped before the model id hits
the wire (main transport + auxiliary Responses adapter), and pricing
aliases the variants onto the base entries.
- Docs: new opt-in section in context-compression-and-caching.md.
OpenAI enabled the large-context window for ChatGPT-subscription Codex
accounts (announced by @thsottiaux Aug 16 2026; previously API-key-only).
Live re-probe the same day: 911,276 input tokens completed OK on
gpt-5.6-sol; ~925K+ rejected with context_length_exceeded (1.05M window
minus reserved output headroom). terra, luna, and gpt-5.4 all completed
900,026 tokens OK. The Codex catalog still advertises 272K, so the
stale-advertisement override from #87981 is the right lever — this just
raises its value 350K -> 900K.
gpt-5.5 and gpt-5.4-mini still enforce 272K live (rejected 500K) and
remain excluded. Override semantics unchanged: fires only on an
exactly-272,000 advertisement; any live catalog change is trusted
verbatim.
The Codex /models catalog advertises 272K for the gpt-5.6 (sol/terra/luna)
and gpt-5.4 slugs, but the backend actually accepts ~371K input tokens
(verified live against chatgpt.com/backend-api/codex/responses, Aug 16 2026:
~371K completed OK on all four slugs; ~382K+ rejected with
context_length_exceeded). 350K keeps ~22K margin under the observed ~372K
enforcement.
The bump applies ONLY when the resolved value is exactly the known-stale
272,000 advertisement — any other advertised value (higher or lower) is
trusted as a real server-side change, so a future catalog correction
deactivates the override automatically. gpt-5.5 and gpt-5.4-mini both
genuinely enforce 272K (rejected 360K live) and are excluded.
save_context_length() and _invalidate_cached_context_length() did an
unguarded read-modify-write into $HERMES_HOME/context_length_cache.yaml.
The plain `open(path, "w")` truncates the file before the dump runs. If
the process is killed mid-dump, the file is left empty or partial. The
next _load_context_cache() swallows the YAML error and returns {} —
silently wiping every persisted context length. A concurrent process
reading between truncate and dump-complete also sees a torn file.
After the cache is lost, every model re-probes the network, and when a
probe fails it falls back to the generic 256K default — so a user on a
1M-window model ends up with a wrong, short context window.
Hermes routinely runs several processes against one shared $HERMES_HOME
(a cron agent plus an interactive session, multiple gateway sessions),
so this is hit in normal use.
Switch both writers to the existing utils.atomic_yaml_write helper
(temp file + fsync + os.replace, symlink- and mode-preserving). The real
file is only ever swapped from a fully written temp file, so an
interrupted write leaves the previous cache intact and readers never see
a partial file. Matches the atomic-write pattern already used for
auth.json, config.yaml, and other persisted state.
Makes the persistent model context-length cache write crash-safe. The
old non-atomic write could truncate or wipe the entire cache on an
interrupted or concurrent write, which then forces models onto the wrong
fallback context window. The fix routes both cache writers through the
repo's atomic temp-file + os.replace helper.
N/A
- [x] 🐛 Bug fix (non-breaking change that fixes an issue)
- [ ] ✨ New feature (non-breaking change that adds functionality)
- [ ] 🔒 Security fix
- [ ] 📝 Documentation update
- [ ] ✅ Tests (adding or improving test coverage)
- [ ] ♻️ Refactor (no behavior change)
- [ ] 🎯 New skill (bundled or hub)
- `agent/model_metadata.py`: `save_context_length()` and
`_invalidate_cached_context_length()` now write via
`utils.atomic_yaml_write` instead of a truncating `open(path, "w")`.
Added the `atomic_yaml_write` import.
- `tests/agent/test_model_metadata.py`: added
`test_write_failure_leaves_existing_cache_intact` — simulates a crash
during the atomic swap and asserts the existing cache survives
byte-for-byte with no stray temp file.
1. `pytest tests/agent/test_model_metadata.py -q` — 98 pass, including
the new crash-safety test.
2. The new test seeds a valid cache, forces the swap step to raise, and
confirms the file is not truncated and no `.cache_*.tmp` is left.
3. `ruff check agent/model_metadata.py` passes.
- [x] I've read the Contributing Guide
- [x] My commit messages follow Conventional Commits (`fix(scope):`, etc.)
- [x] I searched for existing PRs to make sure this isn't a duplicate
- [x] My PR contains **only** changes related to this fix
- [x] I've run the affected tests (`pytest tests/agent/test_model_metadata.py -q`) and they pass
- [x] I've added tests for my changes
- [x] I've tested on my platform: macOS 15 (Darwin 25.5)
- [x] I've updated relevant documentation (README, `docs/`, docstrings) — or N/A
- [x] I've updated `cli-config.yaml.example` if I added/changed config keys — or N/A
- [x] I've updated `CONTRIBUTING.md` or `AGENTS.md` if I changed architecture or workflows — or N/A
- [x] I've considered cross-platform impact (Windows, macOS) — the helper uses os.replace, which is atomic on both
- [x] I've updated tool descriptions/schemas if I changed tool behavior — or N/A
An empty/blank model id reaching get_model_context_length() can't be
meaningfully resolved — and it's worse than a miss: the endpoint
metadata fuzzy matcher ('model in key or key in model') is vacuously
true for "", so it matches an ARBITRARY catalog entry from the live
/v1/models response and returns whatever context length that entry
happens to have, persisting it under a junk '@<base_url>' cache key.
This started failing CI on main when the Nous portal catalog changed:
tests/run_agent/test_primary_runtime_restore.py constructs agents with
model='' against the live portal URL, the arbitrary match now lands on
a 32K entry, and init_agent raises the 64K-floor ValueError
(test_allowed_for_nous_anthropic_messages, red on every PR's slice).
Guard early: a blank model id falls back to DEFAULT_FALLBACK_CONTEXT
immediately, before any cache write or network probe.
Salvaged from #65515 by @whirmill (rebased onto current main; the
guard now sits after the malformed-base_url normalization added since,
and carries an explanatory comment for the fuzzy-match footgun).
Fixes the red slice on #85444, #85452 and every other open PR.
Co-authored-by: whirmill <5079591+whirmill@users.noreply.github.com>
Replaces the per-model _model_name_suggests_grok_4_3/_grok_4_6/
_minimax_m3 stale-cache predicates with one generic
_stale_pre_catalog_cache_entry() guard driven by
_PRE_CATALOG_STALE_KEYS. A cached context length is dropped when the
model resolves (longest-key-first, same as step 8) to a listed catalog
key and the cached value is at or below what the old resolution path
could have produced (largest shorter matching catch-all, or the 256K
fallback).
Also covers qwen3.6-plus, grok-4-fast, and grok-4.20 (the models
PR #37684 requested guards for), absorbing that PR.
_model_name_suggests_minimax_m3 is kept for its two non-cache callers
(models.dev underreport guard, cache-control gating in
agent_runtime_helpers).
docs.x.ai (2026-08-12): grok-4.6 is the flagship, 500K context.
Live GET /v1/models lists grok-4.6 at context_length 500000
(no grok-4.6-latest alias).
#84661 landed the catalog. Main already lists native grok-4.6
on the xAI picker. This is only the leftover cache guard
(same pattern as grok-4.3): pre-catalog builds persisted the
grok-4 catch-all (256K).
'' is a substring of every catalog key, so _resolve_endpoint_context_length
with an empty model name "matched" whatever the endpoint listed first —
on the Nous portal that is currently a 32K embedding model, which poisoned
the resolved context length and made AIAgent init fail the 64K minimum.
This is what turned tests/run_agent/test_primary_runtime_restore.py::
TestTryRecoverPrimaryTransport::test_allowed_for_nous_anthropic_messages
red on every PR (CI slice 7/12) after the portal catalog reordered.
Single-model endpoints still resolve with an empty name (unambiguous);
non-empty names keep the substring fuzzy match.
_PROVIDER_PREFIXES was a hand-maintained frozenset, so providers that ship
as plugins (bundled like fireworks, or user plugins under
$HERMES_HOME/plugins/model-providers/) were never recognised as
provider: prefixes in model strings, and metadata/context-window lookups
received the unstripped string. Mirror the _URL_TO_PROVIDER auto-extend
that already sits below it: add each registered profile's name and
aliases after discovery. The _OLLAMA_TAG_PATTERN guard keeps model:tag
strings intact.
Fixes#66106
- Short-circuit the candidate waterfall on HTTP 401/403: an auth wall
proves the endpoint family exists, so probing the alternate URL just
doubles the wasted wait (the reported endpoint takes ~10s to return
401 without a key).
- Stream the probe so 4xx never downloads a slow error body; responses
are closed on every exit path.
- Regression tests: single-call assertion on 401/403 (fails on main),
negative-cache reuse, 404 waterfall preserved, no .json() on 4xx.
Fixes#69905
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Review pass 2 (reuse reviewer HIGH): the step-3b probe-down fallback for
custom/local endpoints returns the same silent 256K default but only
logged at INFO - invisible by default, and it is the MORE common path
for small local models (the exact users the warning exists for).
Extract _warn_context_length_fallback() (deduped per model+base_url)
and call it from both fallback sites, per the fix-the-whole-bug-class
rule. Regression test drives the custom-endpoint path and fails without
the widening (mutation-checked).
Review follow-up:
- Warn once per (model, base_url) at the step-9 fallback via a module-level
dedup set (established _WARNED_* idiom). The fallback result is
deliberately never cached, so the un-deduped warning fired on every
resolution - e.g. once per gateway message via the session-hygiene path.
- Replace the three inline-mock pool-cleanup tests (which reproduced the
try/except block against a MagicMock and passed even with the production
code reverted) with a parametrized test that drives the real
BatchRunner.run() with a patched Pool; drop the CPython stdlib
signature change-detector test.
- Add a once-per-model warning regression test; clean up dead imports.
All tests verified to fail against pre-PR batch_runner.py/model_metadata.py
and pass with the fix (mutation check).
Salvage of #6629 by aaronlab (kshitijk4poor reworked against current main).
Three concerns from the original PR, reworked to address review feedback:
1. Context-length fallback diagnostic (agent/model_metadata.py):
get_model_context_length() silently returned 256K when all 9 detection
methods failed. Users with small-context models (8K, 32K) would get 256K
silently, causing hard-to-debug API context-length errors. Added a
warning log at the step 9 fallback with model name, base_url, and the
correct config override hint (model.context_length, not context_length).
The token-estimation ceiling-division fix from the original PR already
landed on main (5c2ecdec) with CJK handling — not duplicated here.
2. Fsync for batch trajectory writes (batch_runner.py):
Trajectory entries were written without flush/fsync, but the checkpoint
immediately marked them as completed. A crash between write and disk
sync would leave the checkpoint claiming completion with no trajectory
data on disk. Added flush() + os.fsync() before checkpoint update.
3. Pool cleanup on interruption (batch_runner.py):
Ctrl+C during pool.imap_unordered() relied on context manager cleanup
which can hang. Added explicit pool.terminate() + pool.join() for both
KeyboardInterrupt and Exception paths. The original PR used
pool.join(timeout=10) which is invalid — CPython's Pool.join() takes
no timeout parameter. Fixed to use pool.join() without arguments.
Tests:
- test_warning_emitted_on_fallback: verifies warning fires at step 9
- test_no_warning_when_cached: verifies no false warning when cache hits
- test_trajectory_entry_is_synced_to_disk: verifies os.fsync is called
- test_pool_terminate_called_on_exception: verifies cleanup on RuntimeError
- test_pool_terminate_called_on_keyboard_interrupt: verifies cleanup on Ctrl+C
- test_pool_join_called_without_timeout: verifies no timeout arg to join()
- test_real_pool_join_accepts_no_timeout: integration check on CPython API
Co-authored-by: Aaron Lab <aaronlab@users.noreply.github.com>
Review follow-up on #75102. The shadow substituted the sidecar whenever
the ``api_content`` key was merely PRESENT, but the wire only substitutes
a non-empty string sidecar on a user/assistant row (see
``turn_context.substitute_api_content``). For any other shape the sidecar
is popped and discarded while the clean ``content`` is sent -- so the
shadow dropped real content from the estimate and UNDERcounted, the
dangerous direction: compaction fires too late and the turn dies on a
hard context-length error instead of merely compressing early.
Gate the substitution on the same predicate, and cover the divergent
shapes (None, empty string, int, list, non-user/assistant role) with a
test that fails against the unconditional version.
Also rename the image test: it never carried a sidecar, so it was not
testing what its name claimed. It is a non-regression pin on the flat
per-image accounting that moved into ``_wire_message_shadow()``, and is
now named for that.
`api_content` is a SUBSTITUTE for `content`, not an addition to it.
`turn_context.substitute_api_content()` pops the sidecar and overwrites
`content` at every API-bound message-build site (the `api_messages` build
in `conversation_loop`, the max-iterations summary in
`chat_completion_helpers`, the chat-completions transport), so exactly one
of the two is ever sent to the provider.
The preflight estimator counted both, because both `_estimate_message_chars`
and `_estimate_message_tokens_without_images` walked every key of the
persisted dict with a single-entry denylist (`_anthropic_content_blocks`).
Any message whose sidecar differs from its clean stored content was counted
twice — exactly 2.00x on a 40KB sidecar.
The sidecar exists to keep the provider prompt-cache prefix byte-stable, so
it is written on precisely the long, cache-pinned messages where the
doubling hurts most. Because `estimate_messages_tokens_rough()` also feeds
the compaction threshold via `context_compressor` and `conversation_loop`,
the inflated estimate makes compression fire on phantom bytes.
Fix: substitute rather than sum, mirroring the wire. The two estimator
helpers had drifted into near-identical copies of the same shadow-building
loop, so this factors the shared logic into `_wire_message_shadow()` and
fixes the class once instead of patching one site and leaving the other.
Image accounting is unchanged: base64 payloads are still replaced with a
placeholder and charged at the flat `_count_image_tokens` rate, and the
`_multimodal` text_summary path is preserved.
Tests: three cases in `TestEstimateMessagesTokensRough` — sidecar equal to
content is counted once, a sidecar that DIFFERS is still counted (a lower
bound, so it fails if the field were dropped rather than substituted, which
would undercount the real request), and a sidecar cannot smuggle raw base64
past the flat image rate.
Verified on Linux (Python 3.11): 53 passed in
tests/agent/test_model_metadata.py, 57 passed with
tests/agent/test_context_breakdown.py, 656 passed / 3 skipped across the
compression/context/token/estimate/prune surface of tests/agent.
Mutation-tested: reverting the substitution fails the new equality test.
`scripts/check-windows-footguns.py` is not applicable — no file I/O,
process management, terminal handling, subprocesses, or signals.
The salvaged estimator ran a per-character Python loop on every
estimate_tokens_rough() call — a ~28,000,000x slowdown vs (len+3)//4 on a
1MB ASCII tool output (measured ~3.0s per call). Gate it:
- str.isascii() O(1) fast path keeps pure-ASCII text bit-identical to the
classic (len+3)//4 rule at ~1.3x baseline cost (0.23us vs 0.17us per
1MB call).
- Non-ASCII text counts dense CJK chars via a compiled character-class
regex in C (len(text) - len(re.sub(''))): ~352ms/1MB hangul vs ~2.1s
for the per-char loop.
- Non-ASCII-but-non-CJK text (accents, Cyrillic, emoji) keeps the classic
rule.
Also: parity tests against the per-char reference implementation, and
updated two stale expectations that encoded the old behavior (CJK now
counted ~1 token/char; short string content now ceil-divided instead of
floored to 0). The continuity test now detects merged-into-tail summaries
via _is_context_summary_content.
Follow-up widening for salvaged PRs #67115, #67685, #67620:
- _PROVIDER_MODELS: add kimi-k3 atop kimi-coding / moonshot / opencode-go
curated lists (kimi-coding-cn covered by cherry-picked #67620)
- setup.py _DEFAULT_PROVIDER_MODELS: kimi-k3 for kimi-coding(-cn) + opencode-go
- model_metadata: align DEFAULT_CONTEXT_LENGTHS kimi-k3 entry to 1,048,576
(matches endpoint-scoped override, models.dev, and OpenRouter live metadata)
- anthropic_adapter: classify the bare Coding Plan slug 'k3' (and k3.x/k3-*)
as Kimi family so adaptive thinking applies on proxied endpoints
- moonshot_schema: is_moonshot_model matches bare 'k3' so tool-schema
sanitization runs on the chat-completions path
- contributor mappings for githubespresso407, datachainsystems, Punyko8
Tests: 582 passed across 11 targeted files; hermetic E2E verifies picker
order (kimi-k3 first), no dupes, and 1M context resolution.
Kimi K3 ships with a 1M-token context window (verified against
platform.kimi.ai/docs/overview) but was falling through to the generic
'kimi': 262144 catch-all. Added 'kimi-k3': 1_000_000 before the catch-all
so longest-key-first substring matching resolves K3 to 1M while older
Kimi models still hit the 256K default.
Added matching test_kimi_k3_context_1m test covering native,
vendor-prefixed (kimi/, moonshotai/), and older model fallback.
Kimi Coding serves K3 under the bare slug 'k3', but users can also
configure or select the public-facing aliases 'kimi-k3' and
'kimi-k3-cot'. The endpoint-scoped 1M context window was only keyed
on the bare 'k3' slug, so selecting 'kimi-k3' fell through to the
generic 'kimi' catch-all (262k).
Extend the guard in _endpoint_scoped_context_length to also recognize
'kimi-k3' and 'kimi-k3-cot', while keeping the endpoint check that
limits the 1M value to https://api.kimi.com/coding (legacy Moonshot
endpoints still fall back to 262k). Update the existing test to cover
all three aliases.
Fixes: context window limited to 262k when using kimi-k3 via kimi-coding.
BEDROCK_CONTEXT_LENGTHS was missing entries for current 1M-context Claude
models, and the resolution path in get_model_context_length() short-circuits
to that table (step 1b) before DEFAULT_CONTEXT_LENGTHS is ever consulted, so
the catalog's correct values could never apply on Bedrock:
- claude-fable-5 (no entry at all) fell through to
BEDROCK_DEFAULT_CONTEXT_LENGTH and reported 128K for a 1M model.
- opus-4-7 / opus-4-8 substring-matched the generic 'anthropic.claude-opus-4'
key and reported 200K.
- opus-4-6 / sonnet-4-6 had explicit 200K entries predating their 1M windows.
The practical symptom: the agent compresses context prematurely (at ~128K or
~200K of a 1M window) on every Bedrock-hosted current Claude model.
Fixing the table alone is not enough for existing installs: a previously
persisted 128K/200K value in the context-length cache wins at step 1 and
masks the corrected table forever. Step 1 now reconciles Bedrock-context
cache hits against the static table (the table is authoritative for Bedrock
— there is no live probe to reconcile against), invalidating stale entries
so existing users converge to the right window without manual cache surgery.
Tests cover the new table entries (incl. inference-profile and versioned ID
forms), the 128K-default regression for Fable, the stale-cache invalidation
path, and that pre-4.6 models keep their 200K entries.
Follow-up to the salvaged str(tools) fix. The id()-keyed
_TOOLS_TOKENS_CACHE had no eviction, so a long-lived gateway/desktop
backend could accumulate an unbounded number of stale entries as it
builds transient tool lists. Cap it at 256 with oldest-first eviction
(insertion-ordered dict) and add a regression test asserting the cache
never exceeds the cap.