Drop the bulk test additions from the original PR; keep only mandatory
picker-assertion adaptations (Mantle IDs join the discovery lists), one
allowlist routing test covering all four Mantle model IDs, the 272K
context check, and the two review-mandated auxiliary regressions
(config-region-beats-env for the Mantle path, aux Responses client).
Preserve the Bedrock provider identity for MoA reference and aggregator slots so Bedrock OpenAI Responses models use the aws_sdk/SigV4 runtime instead of being downgraded to a generic custom endpoint. Add regression coverage for Bedrock GPT-5.5 MoA slots.
prune_pre_checkpoint_items() had a hardcoded role=='user' filter that
discarded all non-user messages before a checkpoint — including Hermes'
own compression summaries (role='assistant'), causing total context amnesia
about past conversation summaries.
The fix:
- _is_summary_item delegates to the canonical
agent.context_compressor.is_compaction_summary_message provenance check
(not an ad-hoc heuristic)
- Summaries are retained whole (never byte-sliced) within a 32k token budget
- Idempotent across repeated checkpoints (dedup by identical text)
- _chat_messages_to_responses_input threads item_sources (raw chat messages)
through to the pruner, so it can read summary content directly from the
source when the Responses conversion shape is lossy (tool-result carrier
becomes function_call_output, or stale codex_message_items replay shadows
merged content)
Fixes#90975.
Salvage of #90976 by @JoaoMarcos44.
Change fail-closed behavior to proceed-with-warning when a background
review does not acknowledge cancellation within the bounded deadline.
The review is non-critical self-improvement work and must never block
a user-facing turn (#84423). Keep the off-thread interrupt to ensure
a broken abort path cannot stall the bounded wait.
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).
On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
Zen endpoint routing incl. muse->responses); keyless predicate extended
with unsuffixed free slugs (big-pickle); free runtime pins EVERY
opencode-free model keyless; curated catalog replaces the models.dev
cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
invariant (every curated model must satisfy the keyless predicate)
E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.
The free tier requires a real account API key and throttles third-party
clients by User-Agent:
- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
unconditionally (the stale keyless-tier assumption), and credential-pool
exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
message.
Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
Follow-up to the #90972 salvage:
- strip loop: copy.deepcopy(msg) -> dict(msg). strip_anthropic_cache_control
is copy-on-write on content parts by contract (pops the top-level key,
rebuilds content lists/part dicts fresh), so a shallow top-level copy
preserves the caller-non-mutation guarantee — verified for all four
marker shapes — and removes the redundant second deepcopy the re-mark
path paid on already-decorated input. Docstring updated to match.
- tests: moved the surviving idempotency tests into
tests/agent/test_prompt_caching.py (where this module's tests live) as
TestApplyIdempotency; dropped the three tests that duplicated existing
coverage (dynamic_tool_accounting ~= TestPromptCachePlan::
test_copies_sections_and_keeps_canonical_tools_plain which already
asserts == 4; can_carry_marker_envelope_vs_native ~= TestCanCarryMarker;
never_exceeds_four_markers subsumed by the idempotency test).
- exact-count assertions per review: idempotency fixture pins == 4,
no-tools fallback pins == 3 (marker loss can no longer masquerade as
safety); added the one new _can_carry_marker assertion (native=True
empty assistant) to TestCanCarryMarker.
- new part-level stale-marker mutation guard (the other detection branch,
where part-dict aliasing is the risk); fails on pre-fix base with
marker accumulation (9 > 4), passes with the fix.
apply_anthropic_cache_control never stripped pre-existing cache_control
markers before placing new ones, so calling it twice (or handing it
messages a prior call already marked) accumulated markers past
Anthropic's 4-breakpoint limit and produced HTTP 400
'cache_control can only be specified up to 4 times'.
Strip any pre-existing markers from per-message copies before marking,
mirroring the strip-then-mark pattern build_prompt_cache_plan already
uses. Only messages that already carry a marker pay the copy cost; the
copy-on-write contract (caller-owned messages are never mutated) is
preserved. Repeated calls now converge to byte-identical output.
Salvaged from #90972 by @JoaoMarcos44 (net diff of the PR's commit
stack, intermediate reverts collapsed).
Related: #90971
Normalize malformed memory config during initialization and bind per-target write permissions to the session MemoryStore so direct and staged writes cannot update a disabled built-in store.
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.
Four pieces, and they only make sense together:
· an in-page engine that inventories what is interactable and performs the
verb, injected as source because it has to run inside the guest page;
· the preview.act.request bridge from the gateway into the pane;
· drive_preview, for acting: elements, click, type, scroll, press, and the
pane's own back/forward/reload;
· annotate_preview, for marking without acting.
Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.
Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.
Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
Salvage follow-up for #72283: instead of a second pre-retry clamp block
(which bypassed the #55546 clamp+compress path and broke its three
regression tests), parse the output cap ONCE at classification time and:
- exempt parseable wrapped output-cap 429s from the eager rate-limit
provider fallback (a deterministic request-shape failure that failover
cannot fix but the clamp fixes in one retry), and
- widen is_context_length_error so they reach the SAME #55546
clamp+compress recovery as plain output-cap 400s.
Adds both #72283 regression scenarios plus an ordering guard proving a
NON-EMPTY fallback chain does not consume the wrapped 429 (fallback
slot unspent, model unchanged). 119 fallback/rate-limit tests green.
Review follow-up on the salvaged #90353:
- Todo snapshot (+ coupled pruned-skill reload notice, 7a16840add) is now
reduced only as a LAST resort after reasoning/tool/summary shrink ops,
and the reload notice survives even then.
- Reuse existing helpers/constants instead of re-hardcoding:
_PRUNED_TOOL_PLACEHOLDER, _PRUNE_MIN_CHARS, _NEWEST_TURN_ONLY_BUDGET_KEYS,
and _prune_stale_reasoning_replay (codex sidecar shrink, #71058 boundary).
- Assistant-role messages without the summary metadata key are no longer
truncatable by the summary-cap heuristic.
- Caller passes budget so the estimator runs 3x, not 5x, per would-grow pass.
With memory_enabled: false but user_profile_enabled: true, the memory tool
stays (it backs USER.md) but the full MEMORY_GUIDANCE told the model to save
notes to a MEMORY.md store that does not exist. Split the guidance: a
profile-only block is injected for that configuration, directing writes to
target='user' only.
Walks the real resolution chain -- config.yaml on a temp HERMES_HOME ->
check_memory_requirements -> get_tool_definitions -- rather than mocking
the availability check, since the bug was in how the flags reach the
schema. Covers both flags off, either one alone, no config file at all,
and a config read that raises (must fail open).
Also asserts the external provider's tools survive with the built-in tool
gone, so the fix cannot regress into taking Hindsight/Mem0 down with it,
while disabled_toolsets keeps its documented "hide everything" meaning.
The existing MEMORY_GUIDANCE test built a skip_memory agent whose flags
were both false, so it was asserting the old tool-presence-only behavior;
it now states its precondition and gains the false-case mirror.
Un-fences OPENAI_MODEL_EXECUTION_GUIDANCE from the gpt/codex/grok substring
check and gives it its own injection gate, independent of
tool_use_enforcement, controlled by config.yaml `agent.execution_guidance`
(auto/true/false/list — same semantics as tool_use_enforcement). The "auto"
list (EXECUTION_GUIDANCE_MODELS) now also covers deepseek, kimi, qwen, glm,
minimax, mimo, and mistral.
Composio agentic-eval traces showed Hermes+DeepSeek/Kimi failing where
competitors passed: financial math done in prose, no read-back after
external writes, malformed identifiers "repaired", completeness claimed
despite count mismatches. The discipline block existed but those models
never received it.
The block is extended with compact clauses distilled from that analysis:
- external-write read-back (tool-call success is not task success; internal
file edits already confirmed by the tool are not re-verified)
- count reconciliation (declared totals/has_more are hard assertions)
- literal preservation (never normalize identifiers that fail a stated
format; lookup success does not validate a malformed token)
- retry-differently (empty/partial/suspiciously narrow results get a
broader retry before concluding)
- completion gated on verification (done = every named acceptance
criterion verified, never a plausible subset)
The todo tool description now encourages enumeration-as-checklist for
"all N items" tasks and gates completed status on verified work, never
intent.
Guidance is chosen once at session start keyed on model name, so the
system prompt stays byte-stable for the life of a conversation.
Supersedes/absorbs prior contributor proposals: #20588, #35087, #41874
(MiMo), #53847 (GLM tool-calls-as-text stall).
Co-authored-by: Mat-London <56627804+Mat-London@users.noreply.github.com>
Co-authored-by: intelac <8803887+intelac@users.noreply.github.com>
Co-authored-by: 6ylqq <51219463+6ylqq@users.noreply.github.com>
Co-authored-by: tauros1983 <267660491+tauros1983@users.noreply.github.com>
Review polish from the 3-angle pass on the final stack:
- The warning now says WHICH shape leaked (top-level, extra_body, or both).
Relay injects top-level while request_overrides typically inject via
extra_body, so the shape identifies the offending middleware when
debugging.
- Fold the 'always returns a fresh mapping' assertion into the parametrized
real-endpoint test (the caller mutates the result with stream=True, so the
copy contract is load-bearing on no-drop paths too) and drop the
SimpleNamespace stub test it strictly subsumes. The nested-preserve stub
stays: the parametrized test only exercises top-level retention.
The wire guard only removed the top-level prompt_cache_retention kwarg, but
the OpenAI SDK merges extra_body into the outgoing JSON body, so a nested
extra_body.prompt_cache_retention reaches chatgpt.com/backend-api/codex just
the same and still triggers the non-retryable HTTP 400. Both injection
vectors are real and probe-verified: the Relay overlay's 'key not in
baseline' arm admits an interceptor-added extra_body, and
request_overrides={'extra_body': {...}} lands verbatim in build_kwargs
output.
Close the gap in the same helper: strip the nested field too (copy-on-write,
never mutating the caller's mapping), drop extra_body entirely when it
empties, and log the same warning. Compatible endpoints keep nested
retention untouched.
Mutation-verified: removing the extra_body leg fails both new nested tests.
Reported by egilewski's review on #89969.
The salvaged compatibility test stubs `_is_codex_backend=lambda: False` on a
SimpleNamespace, so it proves the helper honors its own boolean but not that
the boolean is right for any real endpoint. A predicate change that widened
the drop onto retention-supporting hosts would keep it green.
Adds a parametrized test that builds a real AIAgent per base URL and asserts
the drop only fires for chatgpt.com/backend-api/codex, while api.meta.ai,
bedrock-mantle.*.api.aws, api.openai.com and a same-host/different-path
backend keep their supported 24h value. Also asserts prompt_cache_key
survives untouched on every endpoint, since retention and cache-key routing
are independent and the guard must not disturb caching.
Verified non-vacuous: relaxing the guard's condition to drop on every
endpoint fails 4 of the 6 cases (Meta, Bedrock, OpenAI, non-codex path).
Drive-by on the guard itself: drop the dead `None` default on the `pop` that
is already gated by an `in` check, and record why the predicate is resolved
via getattr -- run_codex_stream is driven with lightweight stand-in agents
that lack `_is_codex_backend`, so a bare call would raise AttributeError.
Salvaged from #73811 per the consolidation triage on #76503, adapted to this
PR's per-active-provider model.reasoning_echo design.
Existing reasoning_echo tests hand-set _reasoning_echo_flag; none drives the
real config path init_agent uses:
load_config_readonly().get("model").get("reasoning_echo")
which is wrapped in `except Exception: False`, so a broken read would silently
disable the feature untested. This adds a temp-HERMES_HOME test that resolves a
named custom provider via the real resolve_runtime_provider (asserting
provider == "custom"), materializes the flag from a real config file, and
checks reasoning_content is preserved with the flag on and stripped with it off.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ab7R4NizLNbNyQZShtciKg
Ten behavior tests for target discovery, stable-selector ordering, step
paging, recovery hints, and the self-containment contract the preview
injection depends on. Docs cover data-tour markup and the curated-tour
entry point alongside the tool itself.
The salvaged writer-side fix stamps api_content on NEW hidden redirect
placeholders, but rows persisted before it (content="" + display_kind=hidden,
no sidecar) would keep re-triggering repair_empty_non_final_messages on every
call forever. Substitute [response interrupted] on the wire copy at the
api_content/display_kind projection stage so legacy sessions converge too.
Never the interrupt scaffold (#81841). Durable transcript untouched.
Regression tests drive run_conversation end-to-end with a spied sanitizer:
the projection must leave the sanitizer nothing to heal (its per-turn warning
spam is the bug), verified failing via sabotage run against the writer-only
fix.
Projection-side approach credit: @JoaoMarcos44 (PR #88996).
Bot-mode interrupted member turns with no visible assistant text persisted an
empty assistant row (content="" + display_kind="hidden"). The pre-call
sanitizer repair_empty_non_final_messages() re-healed that row on every later
call (wire copy only), so the loop never converged (#88955).
Stamp api_content="[response interrupted]" (the canonical
_INTERRUPTED_PLACEHOLDER) on the hidden placeholder instead. display_kind is
stripped before sanitization, but api_content is projected back into content
for historical assistant rows, so the provider sees a non-empty neutral turn
and the sanitizer stops touching the row — while the durable transcript stays
hidden and empty. Uses the neutral interruption text, never the
_INTERRUPTED_SCAFFOLD_MARKER, which replaying as assistant text caused #81841.
Adds regression coverage proving (A) the placeholder carries the replay
sidecar, (B) two consecutive projections converge without sanitizer healing,
(C) the sanitizer still repairs genuinely-empty unmarked assistants.
Refs #88955
The steer-survives-budget tests pinned the inline 'Truncated:' fallback
shape, which only occurred because env=None persistence was broken.
Now that host-side spillover succeeds, budget enforcement produces a
<persisted-output> block instead. Assert the actual contract — the
oversized payload was replaced (persisted OR truncated) — via a shared
helper, not which replacement shape was used.
Follow-up to the salvaged #70734 fix:
- test_sanitize_dedup_drops_tool_calls_key_when_all_removed encoded the old
global-uniqueness assumption (its second assistant call reused the id AFTER
the first call was answered, which is now a legitimate new call). The
replayed call now precedes the result, making it a true duplicate of a
still-outstanding call, preserving the intended empty-tool_calls key-drop
coverage from #64335.
- New test: Hermes' own deterministic local counter ids repeating across
turns (the #76632 scenario) survive sanitization.
- New test: the 50-step constant-id field repro from #70724 (Kimi K3 /
llama.cpp) — stock main kept 1/50 tool results, now 50/50.
The #58327 dedup passes treat a repeated tool_call_id as garbage from a
retry/crash/resume glitch and drop it. That assumes tool_call_id is
globally unique, which it is not: llama.cpp emits a single constant id
for every tool call it ever returns (verified — three separate
completions from one server all carried the same id).
Under a seen-once-drop-forever rule, the SECOND legitimate tool result
of such a session looks like a duplicate and is deleted. From the second
tool call onward the model never sees any result: it announces its next
action, the turn ends, and the task is left unfinished. Bisected to
dba585c17 over a 2258-commit range; reproduced live on v0.19.0 (1/6 runs
completed a 4-step file task, vs 20/20 on the last release before that
commit, same model and server).
Key off OUTSTANDING calls instead of every id ever seen. Both original
protections are preserved: a replayed result still answers no pending
call and is still dropped, and duplicate tool_calls sharing an id within
one assistant message are still collapsed. A genuine new call that
reuses the id re-arms it first.
repair_message_sequence needs no change — it already resets its id set
per assistant message, so only the final pre-API pass mis-fires.
Live result after the fix: 8/8 runs complete, 17-26s each (was 1/6 with
runs hitting a 150s ceiling).
Two data-loss bugs reported by users:
1. /handoff CLI→gateway race (#88234): After /handoff completed, CLI
cleanup called finalize_session on the session the gateway just
reopened. This set end_reason on a row the gateway was actively
writing to, causing the handoff leg to vanish from session history
and breaking session_search recall. Fix: add _handed_off_session_ids
module-level set (mirrors _single_query_finalize_attempted_session_ids
pattern). _handle_handoff_command registers the session_id on
completion; _should_emit_cleanup_session_finalize and
_emit_interrupted_session_end check it before firing.
2. state.db corruption silent failure (#88235): When SessionDB init
failed at gateway startup, the error stayed in logs — messages
flowed but nothing was persisted, with no user-visible indication.
Fix: store _session_db_init_error on GatewayRunner, broadcast a
recovery-guidance message to all home channels via
_send_session_db_warning_notifications() after the gateway connects.
Also improved the 'corrupt' persistence cause wording in
_format_turn_completion_explanation to include the full recovery
path (hermes doctor --fix, sqlite3 .recover, backups).
Tests: 6 new tests for handoff cleanup race, 3 for corruption wording.
All existing CLI/turn-completion tests pass.
Per project policy, .env / HERMES_* env vars are reserved for
credentials; behavioural settings belong in config.yaml. Replaces
HERMES_DETERMINISTIC_EMPTY_GUARD and
HERMES_EMPTY_RETRY_COST_THRESHOLD_USD with an additive
agent.empty_response_guard section:
agent:
empty_response_guard:
enabled: true # false = legacy fixed 3-retry behaviour
cost_threshold_usd: 0.25 # per-attempt cost that halves the budget
- hermes_cli/config_defaults.py: new documented subsection under agent
(additive key, no config-version bump needed).
- agent/empty_response_guard.py: resolve_guard_settings() maps the
section to (enabled, threshold) with fail-open tolerance for
malformed values; guard_enabled()/_cost_threshold_usd() now read the
init-resolved agent attributes instead of os.environ.
- agent/agent_init.py: resolves the section once at init into
agent._empty_guard_enabled / agent._empty_guard_cost_threshold_usd,
following the existing tool_use_enforcement extraction pattern.
- Tests updated to config-attr injection; new TestResolveGuardSettings
covering malformed sections, YAML string booleans, bad thresholds,
and a DEFAULT_CONFIG sync check; new integration test proving
enabled:false restores the legacy 1+3-call behaviour.
Requested by isak-ialogics on PR #75115.
Every empty-response retry re-sends the full conversation input at full
price. On large contexts a single turn that produces no visible output
could bill the user several dollars across the 3-retry + fallback-chain
walk (reported: ~$2.33 for one empty answer on a ~26K-token session).
Signaled refusals (finish_reason=content_filter, Anthropic refusal
stop_reason, guardrail interventions) are already terminal today and
never reach this loop. The uncovered class is *unsignaled* refusals:
the provider returns 200 with zero output tokens and a generic finish
reason. Those are deterministic — resending the identical prompt
reproduces the same empty — so burning the remaining retry budget only
multiplies the charge.
New agent/empty_response_guard.py, two independent guards, both failing
OPEN to today's behaviour:
- Deterministic-empty detection: two consecutive empty attempts with
usage present, output_tokens == 0 (reasoning tokens count as output),
and identical (model, provider, finish_reason) skip the remaining
retries and go straight to the fallback chain — a different model may
well answer. Missing usage, nonzero output, or any signature change
keeps the full budget.
- Cost-aware retry budget: when one attempt's estimated input cost
exceeds HERMES_EMPTY_RETRY_COST_THRESHOLD_USD (default $0.25), the
empty-retry budget drops 3 -> 1 for that streak. Unknown pricing or
included/subscription routes are untouched.
At exhaustion the status trace now includes the estimated cost of the
empty attempts so the charge is at least explained in-session.
Streak state lives on the agent and self-clears whenever
_empty_content_retries resets to 0, transparently honouring every
existing reset site (turn start, tool success, compaction, fallback
activation) without touching them.
Set HERMES_DETERMINISTIC_EMPTY_GUARD=0 to disable both guards.
Tests: tests/agent/test_empty_response_guard.py (26 unit tests) plus
two loop-level integration tests in tests/run_agent/test_run_agent.py
proving the api_call reduction and the fail-open path.
Refs NS-503.
'database disk image is malformed' contains the word 'disk', so
classify_persistence_error bucketed SQLITE_CORRUPT / SQLITE_NOTADB
failures as 'disk' and the turn-completion explainer told users to
free disk space for a structurally damaged state.db (the #77386-family
misdiagnosis, reproduced in the v0.20.0 malformed-DB incident report).
- hermes_state: new 'corrupt' bucket in PERSISTENCE_ERROR_CAUSES,
matched via _DB_CORRUPTION_MARKERS BEFORE the locked/disk buckets
- run_agent: explainer text for 'corrupt' points at hermes doctor and
explicitly says freeing space will not help
- cron explainer-variant suppression picks the new variant up
automatically (it iterates PERSISTENCE_ERROR_CAUSES)
Follow-up to #87400: drop the max_iterations and prompt_file knobs from
auxiliary.background_review. The aux model routing (provider/model/
base_url/...) predates #87400 and stays; the enabled switch and the
usage telemetry stay. The fork's iteration budget returns to the
historical hardcoded 16.
Persist fork token usage under session_model_usage task=background_review,
emit a per-fork completion log line, and expose enabled/max_iterations/
prompt_file so operators can see and bound the automatic review cost.
Address review feedback: load auxiliary.background_review once per spawn,
classify completion logs by summarize action prefixes, treat explicit
api_call_count=None as the documented default of 1, and WARNING on the
fail-open enabled-gate path.
Review follow-up. Three coverage gaps in the LiteLLM matrix:
- The operator opt-out on a litellm-named provider pointed at an OpenRouter
host. That route previously took the OpenRouter branch and ignored an
explicit per-model `prompt_caching: false`; it is the only cell in the
differential matrix where the salvage REMOVES caching, so pin it as
intended rather than leaving it to be read as a regression.
- Signal precedence: an explicitly litellm-named provider grants even on a
lookalike host, because the provider id is an independent signal and only
the host-derived signal is token-gated. Intentional, now documented.
- A hyphen-delimited host label (`my-litellm-gw.internal.example.com`),
which the token matcher handles but nothing exercised.
Traded the redundant `claude-3-7-sonnet` parametrize cell for the new host
case, so the matrix covers more shapes with the same cell count.
Tests: 83 passed. All three production fixes re-mutation-checked against
the final stack.
Self-review follow-up. The previous commit fixed substring matching on the
HOST but left the provider-id side as a bare substring, so a user-named
provider like `custom:notlitellm` or `mylitellmthing` still matched and was
handed Anthropic markers — the same bug class, half-fixed.
Both signals now match `litellm` as a whole delimited token via a shared
helper. Real spellings (`litellm`, `custom:litellm`, `litellm-router`, and
the already-lowercased `LiteLLM`) still match; lookalikes no longer do.
Tests: 71 passed. Adds lookalike-provider and real-spelling guards; both
new guards mutation-checked. Differential matrix over 2688 configs vs
origin/main: 60 changes, every one a Claude model on a genuine LiteLLM
route getting the envelope layout, zero pre-existing routes altered.
Follow-up to the salvaged LiteLLM cache grant. The grant itself is right;
four things about how it was scoped were not.
1. Layout. The branch returned the native inner-block layout
(use_native_layout=True) on api_mode == "chat_completions". That layout
writes a TOP-LEVEL msg["cache_control"] on role:tool and empty-content
messages and depends on the Anthropic adapter to relocate it into the
block — but that adapter only runs for api_mode == "anthropic_messages"
(agent/transports/anthropic.py registers there), and the
chat_completions transport does no relocation. Measured on a 3-tool-turn
transcript: 2 of the 4 available breakpoints landed on markers the
provider never sees. Worse, when LiteLLM itself relocates a top-level
marker for an OpenRouter-backed Claude route
(OpenrouterConfig._move_cache_control_to_content), the marker lands on
an empty assistant turn and produces a cache_control-marked empty text
block — the HTTP 400 "text content blocks must contain" shape already
guarded in agent/anthropic_adapter.py (#69512). Switched to the envelope
layout, matching every other OpenAI-wire grant in this function:
4 of 4 breakpoints honored, zero empty blocks.
2. Host matching. `"litellm" in base_url_hostname(...)` is the substring
false-positive class base_url_hostname's own docstring warns against; it
granted Anthropic markers to notlitellm.example.com,
foolitellmbar.example and friends. Replaced with a label-token match in
a named helper, so "litellm" must be a whole dot- or hyphen-delimited
token. All three of the original test hosts still match; a "litellm"
path segment on an unrelated host still does not.
3. Transport gate. `not is_anthropic_wire` also swept in codex_responses,
bedrock_converse and codex_app_server. Gated on
api_mode == "chat_completions" explicitly.
4. Operator override. The grant is inferred from a provider/host name, but
the custom-provider capability lookup was gated on is_anthropic_wire, so
an explicit `prompt_caching: false` for the route+model was honored on
/v1/messages and silently ignored on /v1/chat/completions. The lookup
now also runs for a LiteLLM route, and its layout follows the transport
rather than the declaration (an explicit `true` must not promote a
chat_completions request to the native layout).
Tests: 64 passed. Adds the wire-shape contract the original matrix was
missing (asserts no breakpoint sits on the message envelope, rather than
only checking the returned tuple), plus lookalike-host, other-transport,
and both operator-override directions. All five guards mutation-checked —
reverting each fix turns the corresponding test red.
anthropic_prompt_cache_policy() only granted Anthropic cache_control
markers to LiteLLM over the native Anthropic wire
(api_mode == "anthropic_messages"). A LiteLLM deployment exposing the
OpenAI-compatible surface instead (/v1/chat/completions, /v1/messages
-> 404) matched no grant branch and fell through to (False, False): no
cache_control injected, the system prompt sent as a plain string, and
the provider serving zero cache hits -- the entire prompt re-billed at
full price on every turn. Silent: no error, no warning, usage simply
shows 100% uncached input forever.
Add one branch after the is_anthropic_wire/is_claude case that grants
caching to Claude-family models on a LiteLLM endpoint regardless of
wire, with the native inner-block layout. Same failure class already
documented in-function for Qwen/DashScope.
Design:
- Gated on the Claude family only (is_claude); a Gemini/GPT/Qwen route
through the same proxy must not receive markers (they may reject the
cache_control block format -- cf. the DeepSeek/OpenCode exclusion).
- Matches on provider string OR base_url host, since provider naming
varies per install (litellm, custom:litellm, or a bare custom alias
pointed at a LiteLLM host).
- prompt_caching.cache_ttl: false still wins (the _cache_disabled early
return is untouched).
- Generic strict OpenAI-wire custom providers (e.g. Fireworks) remain
excluded -- verified by the existing over-reach regression test.
Tests: adds TestLiteLLMOpenAIWire covering the grant (several model
spellings x provider/host signals), no-over-reach (non-Claude on the
same proxy get nothing; operator disable wins), and adjacent behavior
(LiteLLM in Anthropic proxy mode still native layout). Full module:
43 passed.
Closes#84506. Original diagnosis, patch design, and measurements by
@ottosulin.
The "Conversation started:" line carried a bare date (%A, %B %d, %Y). Tools
that accept instants -- nutrition, calendar and similar MCP servers -- reject
naive datetimes and require an explicit UTC offset, so the model had to infer
EST vs EDT from the date alone. Near a DST boundary that is a coin flip, and a
wrong guess does not error: it silently writes the record onto the wrong day.
Append the IANA zone (when configured), the zone abbreviation and the UTC
offset, e.g.:
Conversation started: Saturday, August 15, 2026 (America/New_York, EDT, UTC-04:00)
get_timezone() returns None when no timezone is configured; in that case the
line falls back to the abbreviation and offset of the server-local (still
tz-aware) time, so behaviour is unchanged for users who never set one:
Conversation started: Saturday, August 15, 2026 (EDT, UTC-04:00)
Daily byte-stability is preserved -- the property the date-only format exists
to protect (PR #20451). Zone name, abbreviation and offset are all constant for
the whole day; they shift only at a DST transition, where a change is correct.
The static-prefix reconstruction guard in _restore_plugin_sections matches on
"\n\nConversation started:" and is unaffected by a suffix after the date.
test_datetime_is_date_only_not_minute_precision used `re.search(r":\d{2}")`
over the whole line as a proxy for "no time-of-day". A UTC offset also matches
that pattern, so the check now applies to the date portion (everything before
the zone parenthetical) and the invariant is tightened rather than relaxed:
- test_datetime_includes_utc_offset asserts the offset is present
- test_datetime_line_is_stable_across_rebuilds asserts two rebuilds in the
same day produce a byte-identical line
Fixes#87403
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>