test_every_marker_emitting_call_site_goes_through_the_central_clamp read
production .py files with a regex and pinned an exact call-site count --
the change-detector / reads-source-in-tests antipattern AGENTS.md bans
outright. Replaced with a test that drives the real
plan_cache_sections_for_destination fan-in and asserts the emitted wire
markers: ttl=1h on the measured opencode-go route, ttl stripped on the
unmeasured opencode route. Mutation-checked: emptying
MEASURED_1H_PROVIDERS turns it red.
`effective_cache_ttl` evaluates the generic `is_qwen_model` clamp before any
route-level allowance, so a configured `prompt_caching.cache_ttl: 1h` is
silently rewritten to `5m` for every Qwen model on opencode-go. Reproduced on
this commit's parent, no provider traffic:
effective_cache_ttl('1h', 'opencode-go', 'qwen3.7-plus') -> '5m'
_build_marker('5m') -> {'type': 'ephemeral'}
A repair for this exists historically -- payload d6b33faae1, merged as
a43fe4918d -- but neither commit is reachable from current main
(`git merge-base --is-ancestor` returns non-zero for both), so the regression
is live on this lineage.
Restores the precedence fix: MEASURED_1H_PROVIDERS (an allow-list holding only
opencode-go, the one route measured with a delayed read past five minutes) is
consulted ahead of the generic Qwen clamp, with NO_1H_TIER_MODELS nested inside
it for models measured to ignore the tier even on a capable route.
opencode-go stays in ALIBABA_FAMILY_PROVIDERS. That set is also the
cache-marker-layout opt-in read by
agent_runtime_helpers.anthropic_prompt_cache_policy, so narrowing it would turn
working five-minute caching into *no* caching rather than extending the window.
The two sets are kept separate on purpose and a test pins the separation.
One deliberate divergence from the historical payload, found by independent
review: that patch checked NO_1H_TIER_MODELS globally, ahead of the provider
gate. MiniMax on its own Anthropic-compatible endpoint is a separate and
genuinely cache-eligible route, and the global check regressed its configured
1h to 5m off the back of an opencode-go observation -- an unrelated-provider
change this repair must not make. Measured:
base effective_cache_ttl('1h', 'minimax', 'MiniMax-M2.5') -> '1h'
payload -> '5m'
here -> '1h'
Scope note on the evidence: the delayed-read run covered qwen3.8-max and
glm-5.2. The rule is keyed on the route, not the model, so the deployed
qwen3.7-plus is covered by it but has never itself been measured. This change
restores *sending* the requested 1h marker; it does not establish that the
provider honours 1h retention. Those stay separate claims, and the provider
labels every write `ephemeral_5m_input_tokens` whatever ttl was requested, so
only a delayed read past five minutes with no intervening call can settle it.
Tests: 11 new cases covering the deployed model, marker shape, precedence
against the generic clamp, cache eligibility surviving the repair, negative
controls for unrelated routes, provider case normalization, and a closed-set
audit of every marker-emitting call site. Ten mutations were applied to scratch
copies and each turned the suite red, including hoisting the generic clamp back
above the allowance, dropping opencode-go from ALIBABA_FAMILY_PROVIDERS,
un-nesting the model denial, and adding a new unclamped sender.
Known risk, not discharged here: this marker shape has never been sent on the
Go relay. #77217 records the sibling Zen relay returning HTTP 400 on an
unexpected marker shape, so field validation must check HTTP status before it
looks at any cache counter.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QV9Ld4f5ZTqdZWncpfZcQh
It exercises build_prompt_cache_plan's direct_native_tool_cache fallback,
not repeated apply on pre-decorated input, so it belongs with the other
plan-layout tests rather than in TestApplyIdempotency.
Follow-up to the #90972 salvage:
- strip loop: copy.deepcopy(msg) -> dict(msg). strip_anthropic_cache_control
is copy-on-write on content parts by contract (pops the top-level key,
rebuilds content lists/part dicts fresh), so a shallow top-level copy
preserves the caller-non-mutation guarantee — verified for all four
marker shapes — and removes the redundant second deepcopy the re-mark
path paid on already-decorated input. Docstring updated to match.
- tests: moved the surviving idempotency tests into
tests/agent/test_prompt_caching.py (where this module's tests live) as
TestApplyIdempotency; dropped the three tests that duplicated existing
coverage (dynamic_tool_accounting ~= TestPromptCachePlan::
test_copies_sections_and_keeps_canonical_tools_plain which already
asserts == 4; can_carry_marker_envelope_vs_native ~= TestCanCarryMarker;
never_exceeds_four_markers subsumed by the idempotency test).
- exact-count assertions per review: idempotency fixture pins == 4,
no-tools fallback pins == 3 (marker loss can no longer masquerade as
safety); added the one new _can_carry_marker assertion (native=True
empty assistant) to TestCanCarryMarker.
- new part-level stale-marker mutation guard (the other detection branch,
where part-dict aliasing is the risk); fails on pre-fix base with
marker accumulation (9 > 4), passes with the fix.
_apply_static_prefix_marker duplicated _apply_system_cache_markers' split
logic minus its empty-suffix guard: when the stored system prompt equals
the static prefix exactly, the tool-cache plan emitted a two-part split
with a trailing empty text block — HTTP 400 on native Anthropic. Fold the
tool-cache layout into the existing helper via mark_suffix /
fallback_to_whole flags; the empty-suffix case now marks the prompt as
one whole block. Behavior-parity verified against the merged planner for
every non-empty-suffix shape.
Follow-up to #76032 (#20880).
Self-review found the deepcopy->shallow-copy change in
apply_anthropic_cache_control (#57046) had no test pinning the
"prompt caching is sacred / never mutate the caller's list" invariant.
The old test_returns_deep_copy only exercised the single marked message,
never an un-marked shared reference, so a regression that deep-copied too
little would pass the whole suite.
Add two tests: (1) caller list + every element left byte-identical after
the call, un-marked middle messages returned as shared references, marked
messages fresh copies, and mutating a returned marked message does not leak
upstream; (2) structural byte-equivalence vs a reference full-deepcopy
implementation across both native_anthropic modes and two TTLs.
Mutation-verified: neutering the per-message deepcopy makes test (1) fail.
strip_anthropic_cache_control flattened ANY pure-text multi-part content
list with a separator-less join. Decoration only ever produces a single
text part or the 2-part [static, volatile] system split; organic
multi-part text (merged user turns, imported transcripts) got word-jammed
and parts carrying extra keys (citations) were silently dropped — on the
common no-failover path, since redecoration runs on every attempt.
Restrict the flatten to the exact decoration-produced shapes and make
marker removal copy-on-write on part dicts (the per-call message copy is
shallow, so parts alias the persistent history).
Add strip_anthropic_cache_control coverage and policy-change cases
(cache-off→on, on→off, native→envelope, MoA guidance outside marker)
that TestSyncFailoverPreservesCacheDecoration did not exercise.
The conversation loop normalizes message text right before the API call so
the request prefix is byte-identical across turns -- the stated reason is
KV cache reuse on local inference servers and better cache hit rates on
cloud providers. Cache breakpoints were injected *before* that pass, which
defeats it.
`_apply_cache_marker` rewrites a plain-string `content` into a
`[{"type": "text", ...}]` block. The normalization pass is guarded on
`isinstance(content, str)`, so every message that just got marked is
silently skipped by it and keeps its raw leading/trailing whitespace. A
message is only marked while it sits in the last-3 window, so:
turn N in the window -> marked, content "file1\nfile2\n"
turn N+1 rolled out -> plain, content "file1\nfile2"
The same logical message is sent with different bytes on consecutive
turns. The prefix stops matching at that position -- which is inside the
span the breakpoints were placed to protect -- so the reusable prefix
collapses back toward the system breakpoint on every turn. Tool results
carry a trailing newline almost by default (any shell command output), so
this is the common case, not an edge case.
Move the injection below every message mutation. Besides fixing the
whitespace divergence this stops breakpoints from being spent on messages
that the orphan sweep or the thinking-only drop is about to remove or
merge away -- a marker on a dropped message is a wasted breakpoint out of
the four available.
Nothing between the old and new call sites reads `cache_control`, and the
mutators now see the plain-string shapes they were written against.
Trim the salvaged commit to its two still-valid conversions:
- ChatCompletionsTransport.convert_messages (copy-on-write sanitize)
- QwenProfile.prepare_messages (copy-on-write normalize + cache_control)
Dropped from the original PR:
- agent/prompt_caching.py selective-copy: superseded by #57229 which
already rewrites apply_anthropic_cache_control on current main.
- Gemma extra_content narrowing (_model_consumes_thought_signature
'gemini or gemma' -> 'gemini' only) + its two tests: unrelated
behavior change reverting deliberate e8c3ac2f5; belongs in its own
PR with its own justification if pursued.
Conflict resolution: preserved main's newer timestamp-stripping
(#47868) inside the copy-on-write path.
Follow-up to the salvaged #57845 fix. _can_carry_marker used
any(isinstance(part, dict)) but _apply_cache_marker only marks the LAST
content part, so a list whose last element is a non-dict passed the carrier
gate yet received no marker — wasting one of the four breakpoints. Tighten
the predicate to require content[-1] to be a dict (mirroring the apply
logic) and add a regression test. Flagged by a 3-agent review.
Remove unused imports (F401) and duplicate/shadowed import
redefinitions (F811) across the codebase using ruff's safe
autofixes. No behavioral changes -- imports only.
- ~1400 safe autofixes applied across 644 files (net -1072 lines)
- __init__.py re-exports preserved (excluded from F401 removal so
public re-export surfaces stay intact)
- Re-exports that are imported or monkeypatched by tests but look
unused in their defining module are kept with explicit # noqa:
F401 (gateway/run.py load_dotenv; run_agent re-exports from
agent.message_sanitization, agent.context_compressor,
agent.retry_utils, agent.prompt_builder, agent.process_bootstrap,
agent.codex_responses_adapter)
- Unsafe F841 (unused-variable) fixes deliberately skipped -- those
can change behavior when the RHS has side effects
- ruff lints remain disabled in pyproject.toml (only PLW1514 is
selected); this is a one-time cleanup, not a config change
Verification:
- python -m compileall: clean
- pytest --collect-only: all 27161 tests collect (zero import errors)
- core entry points import clean (run_agent, model_tools, cli,
toolsets, hermes_state, batch_runner, gateway)
- static scan: every name any test imports directly from an edited
module still resolves
The long-lived prefix-cache layout split the system prompt into stable/
context/volatile blocks and re-derived them on every API call. The
volatile tier (timestamp + memory snapshot + USER profile) ticks per
turn, so the system message bytes mutated mid-conversation and broke
upstream prompt caches (OpenRouter, Nous Portal, Anthropic).
Diagnosed via live wire-format diffing: an 8-turn conversation showed
OLD layout flipping system block[1] sha mid-session at the minute
boundary, dropping cached_tokens to 0 on that turn (cumulative
66.6% vs 83.3% for the single-block layout). Hermes invariant:
history (system + all but the last 1-2 messages) must be static.
Fix: drop the long-lived layout entirely. Single layout everywhere —
system_and_3 with one cached system string built once on first turn,
replayed verbatim on every subsequent turn. Loses cross-session 1h
prefix caching for Claude (the feature that motivated the split), but
within-session caching now actually works on every provider.
Removed:
- run_agent.py: _use_long_lived_prefix_cache flag, _long_lived_cache_ttl,
_supports_long_lived_anthropic_cache method, the long-lived branch in
run_conversation, mark_tools_for_long_lived_cache call site
- agent/prompt_caching.py: apply_anthropic_cache_control_long_lived,
mark_tools_for_long_lived_cache, _mark_system_stable_block helper
- hermes_cli/config.py: prompt_caching.long_lived_prefix and
prompt_caching.long_lived_ttl config keys
- tests/agent/test_prompt_caching_live.py (entire file)
- tests/agent/test_prompt_caching.py: TestMarkToolsForLongLivedCache,
TestApplyAnthropicCacheControlLongLived
- tests/run_agent/test_anthropic_prompt_cache_policy.py:
TestSupportsLongLivedAnthropicCache
Targeted tests: 62/62 pass.
Cuts input cost for first-turn Claude requests by ~85-90% on subsequent
sessions within an hour. Tools array (~13k tokens for default toolset) +
stable system prefix (~5-8k tokens) get a 1h cache_control marker; the
volatile suffix (memory, USER profile, timestamp, session id) sits in a
separate non-cached block at the end so it doesn't poison the cross-session
prefix when it changes.
Provider gate: Claude on native Anthropic (incl. OAuth subscription),
OpenRouter, and Nous Portal (which proxies to OpenRouter). All other
providers keep today's system_and_3 layout unchanged.
Layout (4 cache_control breakpoints, Anthropic max):
1. tools[-1] -> 1h (cross-session)
2. system content[0] -> 1h (cross-session, stable prefix)
3. messages[-2] -> 5m (within-session rolling)
4. messages[-1] -> 5m (within-session rolling)
Within-session rolling shrinks from 3 messages to 2 to free the breakpoint
budget. On Claude with realistic tool loadouts the long-lived tier carries
the bulk of cross-session value anyway.
System prompt is now always assembled cache-friendly: stable identity /
guidance / skills / platform hints first, then session-stable context
files (AGENTS.md, .cursorrules), then per-call volatile content. Old
single-string callers see the same logical content (same join order),
just reordered so volatile lives at the end.
Config knobs (defaults shown):
prompt_caching:
cache_ttl: "5m" # rolling-window TTL (unchanged)
long_lived_prefix: true # opt-out switch
long_lived_ttl: "1h" # cross-session prefix TTL
Live E2E (tests/agent/test_prompt_caching_live.py, gated on
OPENROUTER_API_KEY) on anthropic/claude-haiku-4.5 with default toolset:
Call 1 (cold): cache_write=13,415 cache_read=0
Call 2 (NEW agent + msg): cache_write=391 cache_read=13,025
Cross-session reuse: 97.09%
Implementation:
* agent/prompt_caching.py: new apply_anthropic_cache_control_long_lived()
+ mark_tools_for_long_lived_cache(); existing apply_anthropic_cache_control()
preserved verbatim for the fallback path.
* agent/anthropic_adapter.py: convert_tools_to_anthropic() now forwards
cache_control onto each Anthropic-format tool dict.
* run_agent.py: _build_system_prompt_parts() returns the 3-tier dict;
_build_system_prompt() joins them (backward compatible).
_supports_long_lived_anthropic_cache() policy added next to the existing
_anthropic_prompt_cache_policy() (which now also recognises Nous Portal
Claude — pre-existing gap fixed in passing).
_build_api_kwargs() resolves tools_for_api once and propagates the
marker through all four build paths (anthropic_messages, bedrock,
codex_responses, profile/legacy chat completions).
Long-lived flag plumbed into the runtime snapshot/restore + model-switch
+ fallback-promotion paths.
Tests:
* tests/agent/test_prompt_caching.py: +8 tests (TestMarkToolsForLongLivedCache,
TestApplyAnthropicCacheControlLongLived).
* tests/run_agent/test_anthropic_prompt_cache_policy.py: +9 tests
(TestSupportsLongLivedAnthropicCache matrix across 8 endpoint classes
+ a fallback-target case).
* tests/agent/test_prompt_caching_live.py: new live E2E (skipif when
OPENROUTER_API_KEY is unset; runs outside the hermetic suite).
* Targeted suites: 327/327 pass (caching/adapter/policy/builder).
* tests/agent/ + tests/run_agent/: 3992 pass, 17 skip, 1 pre-existing
flake (test_async_httpx_del_neuter::test_same_key_replaces_stale_loop_entry,
verified failing on pristine origin/main).
On the native Anthropic Messages API path, convert_messages_to_anthropic()
moves top-level cache_control on role:tool messages inside the tool_result
block. On OpenRouter (chat_completions), no such conversion happens — the
unexpected top-level field causes a silent hang on the second tool call.
Add native_anthropic parameter to _apply_cache_marker() and
apply_anthropic_cache_control(). When False (OpenRouter), role:tool messages
are skipped entirely. When True (native Anthropic), existing behaviour is
preserved.
Fixes#2362
Cover model_tools, toolset_distributions, context_compressor,
prompt_caching, cronjob_tools, session_search, process_registry,
and cron/scheduler with 127 new test cases.