Third in the series. The gateway rebuild path (previous two commits)
carries a custom provider's `request_overrides` (`extra_body`, e.g.
`chat_template_kwargs`) into the agent, but the *in-place* live switch used
by the TUI dashboard and the CLI — `agent.switch_model()` ->
`agent_runtime_helpers.switch_model()` — swapped
model/provider/base_url/api_key without ever updating `request_overrides`.
So a `/model` switch to a thinking-enabled custom provider in the TUI/CLI
kept the previous provider's `extra_body`.
`switch_model()` now re-derives the switched-to provider's
`request_overrides` (via `_get_named_custom_provider`) and applies it in
place, preserving non-provider overrides (`service_tier`/`speed` from
`/fast`). Logic factored into `_apply_switched_provider_request_overrides`
for testability.
Adds tests/agent/test_switch_model_request_overrides.py.
The consumed-but-uncommitted rotation verdict was process-local
(_SPENT_ROTATION_FINGERPRINTS), while the credential it protects is
explicitly cross-process: ~/.claude/.credentials.json is shared by every
Hermes profile and process. A fresh interpreter could lease the stale
access token or re-POST the already-spent single-use refresh token and
burn the credential family into invalid_grant.
- Persist non-secret one-way fingerprints to a sidecar registry next to
the shared singleton source (claude_code / hermes_pkce), written under
the same path-keyed cross-process lock that serializes refreshes.
- Consult the sidecar in the pool resolver, the pool refresh path, and
the direct claude_code resolver/refresh before leasing or POSTing.
- Two-process regression: A rotates and loses the commit; B (fresh
interpreter, empty local registry) must neither lease the stale pair
nor POST the spent refresh token. Plus a no-verdict control.
Closes the remaining P1 from the exact-head review of f228439b on
PR #87891.
Two runtime blockers from the exact-head review of c057ef5.
1. A sanitized `claude_code` pool row was treated as token authority.
`claude_code` is a borrowed source: it is absent from the owned-source
allowlist, so `sanitize_borrowed_credential_payload` strips `access_token`
and `refresh_token` before the row reaches `auth.json`. `load_pool()`
re-hydrates the live pair from the singleton on every load, which is what
makes `~/.claude/.credentials.json` — not the pool store — authoritative
for this source.
`_sync_anthropic_entry_from_pool_store()` re-read that persisted row during
refresh. Being token-less, it "differed" from the live entry, so it was
adopted as a rotation performed by another process: `_refresh_entry()`
replaced a usable credential with an empty one and returned it before
`_claude_code_credentials_lock()` and the authoritative re-read were ever
entered. The empty OAuth entry then stayed selectable, because the
empty-runtime-key guard in `_available_entries()` covered API-key rows only.
Repairs: the pool-store sync refuses borrowed sources outright (plus a
defensive refusal of any token-less row, for future sources that sanitize on
write); the `claude_code` branch of `_refresh_entry()` now runs before the
generic adopt-and-return shortcut, so the path-keyed lock and the
authoritative re-read are always entered before deciding to POST or adopt;
and an OAuth entry with no access token is never leased.
2. A failed commit still fell through to the same spent credential.
`_refresh_oauth_token()` correctly returns None when the refresh POST
rotated the single-use token but the replacement could not be committed.
That verdict did not survive the caller: `resolve_anthropic_token()`
continued to `_resolve_anthropic_pool_token()`, which enumerates read-only
(`clear_expired=False, refresh=False`) over a pool that `load_pool()` had
just re-seeded from the unchanged singleton — so the pair whose refresh half
was already spent came back as a healthy token, and
`_refresh_provider_credentials("anthropic")` reported success and evicted
its cached clients.
Repair: every commit-failure path records the consumed pre-rotation pair as
non-reversible fingerprints (bounded, process-local), and both the Claude
Code file resolver and the pool resolver refuse a credential whose
fingerprint is on that list. `_refresh_provider_credentials("anthropic")`
consequently returns False when the spent family is the only credential,
while genuinely independent pool credentials stay eligible.
Coverage: `test_anthropic_borrowed_row_authority.py` starts from
`load_pool()` reading an actually persisted, actually sanitized row, forces
a refresh, and asserts the full pair survives with exactly one POST and one
commit, that the shared-file lock is entered, and that no empty OAuth entry
can be leased. `test_anthropic_spent_rotation_verdict.py` takes the full
resolver path: successful POST plus failed commit must make
`resolve_anthropic_token()` return None, make
`_refresh_provider_credentials("anthropic")` return False, and keep the
spent fingerprint out of every lease — with a control proving a successful
commit quarantines nothing and an independent credential still resolving.
Five of the seven new borrowed-row tests fail on the previous head, and the
three resolution tests fail with the verdict disabled.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gcoy6nLTg5R6FHHhjcLZEC
`agent/anthropic_adapter.py` was 3,423 lines and this PR adds another auth
boundary to it. Split along the seams that were already there, so the
credential surface this PR changes has a single owner instead of being
interleaved with request building:
- `agent/anthropic_endpoints.py` (258) — base-URL/endpoint-family predicates.
Pure functions over a URL string, which is what lets both of the modules
below depend on it without a cycle.
- `agent/anthropic_message_convert.py` (1,225) — OpenAI-style to Anthropic
Messages payload conversion: model ids, tool schemas, content/thinking
blocks, tool_use pairing, cache_control, screenshot eviction, blank-block
scrubbing.
- `agent/anthropic_credentials.py` (910) — credential sources, the OAuth
flows, and the refresh commit (`CredentialPersistError` and both singleton
writers).
- `agent/anthropic_adapter.py` (1,215) — client construction and the Messages
API call, re-exporting every name from the three modules above so existing
`from agent.anthropic_adapter import ...` imports keep resolving. The
re-export surface was diffed against the pre-split module: nothing dropped.
Call sites that read a moved name through the adapter's namespace at runtime
(`credential_pool._refresh_entry_impl`, `auxiliary_client`) now import it from
the defining module, so there is one patchable seam rather than two bindings
that can disagree. The tests that monkeypatched those seams were retargeted to
match; no assertion was changed.
No behavior change.
Anthropic OAuth refresh tokens are single-use: the POST that returns a new
pair invalidates the one that was sent. The replacement therefore only
becomes real once it reaches its authoritative store -
~/.claude/.credentials.json for claude_code entries,
~/.hermes/.anthropic_oauth.json for hermes_pkce ones. Both writers caught
OSError/IOError, logged at debug level and returned nothing, so no caller
could tell a durable commit from a failed one.
That let a refresh spend the only refresh token, report success, and leave
the consumed pre-rotation pair on disk. _seed_from_singletons() re-reads
those files on every load_pool(), so the next process seeded the spent pair
back over the fresh pool row and the following refresh replayed a consumed
token (invalid_grant / refresh_token_reused) - exactly the failure this PR
set out to remove.
- _write_claude_code_credentials() and _write_hermes_oauth_credentials()
now raise CredentialPersistError instead of swallowing the write error.
- _refresh_oauth_token() treats a failed commit as a failed refresh and
returns None rather than handing back an access token whose refresh half
was lost.
- _refresh_entry_impl() fails closed on both the primary and the recovery
path: the rotated pair is never marked, persisted or returned, and the
entry is quarantined DEAD with a credential_persist_failed reason so it
leaves rotation and surfaces as an explicit re-auth instead of a silent
fallback to another provider. The retry path now commits to the singleton
before persisting the pool row.
- _upsert_entry() no longer treats re-seeding a borrowed source as a
rotation. Borrowed rows (claude_code, env-backed) are written to auth.json
without their secret, so comparing the re-seeded token against the empty
stored value reported a rotation on every load and cleared the DEAD state
the previous process had just written - resurrecting the quarantined,
already-consumed credential on restart. It now compares the incoming
token against the row's secret_fingerprint.
Adds failure-injection coverage for both writers, the direct resolver, the
claude_code and hermes_pkce pool paths and the retry path, each asserting
that a reload cannot bring the pre-refresh pair back as a usable credential.
Add a cross-process lock over the shared ~/.claude/.credentials.json file
so concurrent Hermes processes racing a claude_code-sourced Anthropic
refresh resync instead of losing the update (mirrors the existing
per-profile auth-store lock, kept as the outer lock per the documented
lock-ordering invariant).
Remove the dashboard-triggered Anthropic PKCE OAuth flow entirely rather
than continue patching it: an unattended HTTP endpoint minting Claude
Pro/Max subscription tokens outside Anthropic's own client sits on the
wrong side of Anthropic's OAuth usage policy. The provider catalog entry
is now flow == "external", pointing at `hermes auth add anthropic`
(terminal PKCE, unaffected, out of scope). Drop the now-dead PKCE
functions/constants and the tests that exercised only that removed code.
Dashboard PKCE login reused the code_verifier as the OAuth state (leaking
it and disabling CSRF validation) and never checked state on callback --
the same class of bug already fixed for the CLI flow. Credential-pool
refresh excluded "anthropic" from the cross-process lock Codex/xAI already
get, so concurrent Hermes processes racing a single-use refresh token could
leave the loser stuck exhausted with no recovery for hermes_pkce/dashboard
sources. The dashboard OAuth save also never cleared a stale
ANTHROPIC_API_KEY, which resolve_anthropic_token() prioritizes over the
OAuth pool entry by design -- so a leftover key silently kept billing
pay-per-token after a Claude Pro/Max login.
A concurrency stress test written to validate the refresh-race fix under
load surfaced a fifth, unrelated bug: _auth_store_lock()'s Windows
lock-file "ensure content" write was unguarded and could raise an uncaught
PermissionError under real contention -- affecting every single-use-token
provider sharing that lock, not just Anthropic.
Fixes#87887, #87888, #87889.
The rename sweep in the base commit missed the sibling-test blast radius
(18 red files on CI). Three classes, all fixed:
1. Stale old names in tests (todo/cronjob/process/tour/tip) — updated to
todo_list/cronjob_manage/process_manage/gui_tour/show_tip at every
registry.get_entry/dispatch/coerce/preview/allowlist call site, plus
the coding-brief sentence in agent/coding_context.py now names
todo_list (and its gating test).
2. Missed rename in production: AGENT_RUNTIME_POST_HOOK_TOOL_NAMES still
held 'tour' — post-hook ownership would have double-emitted for
gui_tour via the bridge path.
3. Tests pinning pre-deferral assembly (blank-slate surface, modal
sandbox resolution, desktop diet, HUD note) now pin their ACTUAL
contract under the legacy defer:[] override, or assert on granted
tool names instead of visible schemas.
Also fixes a pre-existing ordering flake surfaced by the sweep:
test_holds_exactly_the_gui_affordances depended on whether an earlier
test had imported apply_layout_tool (registry-registered, not in the
static desktop_ui list) — now forces discovery and pins the full set.
649 tests green locally across all touched files, both orderings.
The chat_completions and transport-parity Mistral tests pinned the
same branch with different reasoning_config shapes ({effort:none} vs
{enabled:False, effort:none}) — behaviorally identical since effort
short-circuits first, but the drift reads as a semantic difference.
Align both to the explicit form. Also scope the port-guard comment to
the try/except shape it actually shares with hermes_cli/models.py.
The initial /btw implementation (#97937) answered from a rendered
plain-text transcript digest — truncated context, cold-written tokens on
every question. Teknium's call: reuse the self-improvement review fork
instead, which keeps the entire prompt cache stable for the fork and
gives it the complete conversation for very cheap.
- agent/background_review.py: extract the review-fork construction into
build_cache_parity_fork() — same runtime/credentials as the parent,
byte-identical system prompt / tools[] / reasoning config on the
same-model path, shared session_id for prefix warmth, full persistence
detachment (no state.db writes, no rotation, no external memory,
in-place-only compaction). The review thread now calls the helper;
behavior unchanged (full review test suite green).
- agent/side_question.py: /btw prefers the fork when a live parent
AIAgent exists — replays the untruncated snapshot as warm cache reads,
denies every tool at dispatch via an empty thread whitelist (tools[]
stays byte-identical for cache parity), attributes usage to the parent,
and trims a mid-turn snapshot tail so role alternation holds. The
one-shot digest remains as fallback (no live agent = cold cache anyway,
and any fork failure degrades gracefully).
- CLI passes self.agent, TUI passes the session agent, gateway looks up
the chat's cached agent (parity with how turns reuse it).
Live-verified: /btw on the worktree runs the fork path (agent.log shows
the side question as a forked conversation turn on the parent session_id
with the full history replayed), answers correctly from context.
Addresses the automated review on this PR:
- _profile_declared_efforts falls back from provider name to the
endpoint's host (via model_metadata's URL->provider map), so a named
custom provider pointed at api.router.com — which the host mandate
already routes onto this transport — gets the catalog clamp instead
of the default vocabulary and a Router 400.
- _parse_efforts validates catalog levels against EFFORT_LADDER at
ingest, logging and dropping unrecognized tiers; a model whose whole
vocabulary is unrecognized stays out of the map (transport defaults)
instead of passing the requested effort through unclamped.
- fetch_models dedupes ids while preserving Router's deliberate listing
order.
- plugin.yaml credits the human contributor per repo convention.
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:
- plugins/model-providers/router/: RouterProfile plugin —
api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
runtime_provider._detect_api_mode_for_url: api.router.com ->
codex_responses. The host is Responses-only — POST /v1/chat/completions
does not exist and 404s — so this is a genuine host mandate (exact
hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
hook (tri-state: None=defer, ()=model takes no reasoning params,
tuple=clamp set). Router validates reasoning.effort per model and
returns HTTP 400 invalid-argument on levels outside the model's
published vocabulary, and 400 unsupported_parameter when a
non-reasoning model receives any reasoning field (both verified live).
The profile answers from a cached copy of the catalog's
router.capabilities.reasoning block: cache-only on the hot path,
seeded for free by fetch_models(), disk-mirrored across processes
(/cache/router_catalog.json), background-warmed when cold
— same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
the generic effort-clamp branch (xai/actual/github branches untouched;
profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
rejection, profile registration + auth auto-registry wiring, catalog
parsing, and transport clamp/suppression/fallback paths.
Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.
/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.
Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
* test(system_prompt): cover session-start anchoring + fix Windows-portable expectation
- Add TestSessionStartLike unit tests for _session_start_like(): session-id
embedded timestamp, session_start fallback, now fallback, non-matching id.
- Add a build-level regression: a session started Jan 1 must still render
'Conversation started: Thursday, January 01' when the prompt is rebuilt
on Jan 2 (the rebuild-drift bug).
- test_coding_prompt_preserves_legacy_workspace_order hardcoded '/hermes'
while production renders str(Path('/hermes')) — backslash on Windows made
the suite fail on Windows (CI runs Linux, so it was never caught). Build
the expectation via str(Path()) to match production on every platform.
* fix(system_prompt): anchor 'Conversation started' to the real session start
The timestamp line stamped hermes_time.now() at system-prompt build time.
The prompt is rebuilt on compression, fresh-agent gateway turns, and
resume-without-stored-prompt, so the date silently advanced to whatever
day the prompt was last rebuilt — a chat that started on Wednesday read
as 'Conversation started: Thursday' after a Thursday-morning resume,
contradicting the fresh per-turn time hint.
Resolve the true start via _session_start_like(): the timestamp embedded
in the session id (YYYYMMDD_HHMMSS_..., immutable for the session life)
-> agent.session_start -> now() only as last resort. Box-local stamps are
attached to the box's local zone then converted to the rendered zone so
the date is consistent with the per-turn clock. The line stays date-only
and is now byte-stable for the whole session (never moves on rebuild),
preserving prefix-cache KV. The zone suffix and _bot_chat_timeless_prompt
behaviour are untouched.
* feat(system_prompt): two-line conversation clock — anchored start (salvaged #96224, credit @bobaba76) + as-of-last-rebuild date for multi-day sessions
---------
Co-authored-by: bobaba76 <79245850+bobaba76@users.noreply.github.com>
- Widen opencode_zen_free_runtime healing to the union of the static floor,
the in-process live memo, and the SWR disk cache — a newly-live free model
now heals opencode-go/zen selections without a release (sibling site the
original PR missed).
- Memoize _fetch_opencode_free_models() in-process (5 min, negative caching
included) so direct provider_model_ids() validation callers don't each
block on a network round-trip or timeout.
- Drop delisted x-preview-f-free from the offline floor and setup.py sample
list (offline fallback must not offer a model that 401s); add the newly
live deepseek-v4-flash-free / mimo-v2.5-free to setup.py.
- Update stale test fixtures to a live exemplar; add regression tests for
memoization, negative caching, and union healing; docs note in providers.md.
* refactor(prompt): diet the memory/skills guidance block — schema-taught curricula removed, form rule + pruning contract kept (537 -> 223 tok in the combined block)
* polish: literal check-glyphs in source; memory capacity posture — save proactively, replace/consolidate when full
* refactor: single spine for memory/profile guidance — form rule + capacity posture written once, variants differ only in opening frame
* refactor: ONE memory-guidance builder — frame adapts to enabled stores, body written once, positive posture leads (maintainer direction)
* fix wording: memory is loaded per SESSION, not injected per turn (maintainer correction)
Final /simplify-code pass on the full diff: after the content-digest
rework, _creds_cache_key had become a wrapper with zero production
callers — get_vertex_credentials inlined the same read/except dance —
kept alive only by its own tests. One helper now owns the
(bytes, cache_key) resolution for all three cases (ADC sentinel,
readable file, unreadable fallback); prod calls it, tests target it.
No behavior change: 10/10 green, same mutation-check results.
test_every_marker_emitting_call_site_goes_through_the_central_clamp read
production .py files with a regex and pinned an exact call-site count --
the change-detector / reads-source-in-tests antipattern AGENTS.md bans
outright. Replaced with a test that drives the real
plan_cache_sections_for_destination fan-in and asserts the emitted wire
markers: ttl=1h on the measured opencode-go route, ttl stripped on the
unmeasured opencode route. Mutation-checked: emptying
MEASURED_1H_PROVIDERS turns it red.
`effective_cache_ttl` evaluates the generic `is_qwen_model` clamp before any
route-level allowance, so a configured `prompt_caching.cache_ttl: 1h` is
silently rewritten to `5m` for every Qwen model on opencode-go. Reproduced on
this commit's parent, no provider traffic:
effective_cache_ttl('1h', 'opencode-go', 'qwen3.7-plus') -> '5m'
_build_marker('5m') -> {'type': 'ephemeral'}
A repair for this exists historically -- payload d6b33faae1, merged as
a43fe4918d -- but neither commit is reachable from current main
(`git merge-base --is-ancestor` returns non-zero for both), so the regression
is live on this lineage.
Restores the precedence fix: MEASURED_1H_PROVIDERS (an allow-list holding only
opencode-go, the one route measured with a delayed read past five minutes) is
consulted ahead of the generic Qwen clamp, with NO_1H_TIER_MODELS nested inside
it for models measured to ignore the tier even on a capable route.
opencode-go stays in ALIBABA_FAMILY_PROVIDERS. That set is also the
cache-marker-layout opt-in read by
agent_runtime_helpers.anthropic_prompt_cache_policy, so narrowing it would turn
working five-minute caching into *no* caching rather than extending the window.
The two sets are kept separate on purpose and a test pins the separation.
One deliberate divergence from the historical payload, found by independent
review: that patch checked NO_1H_TIER_MODELS globally, ahead of the provider
gate. MiniMax on its own Anthropic-compatible endpoint is a separate and
genuinely cache-eligible route, and the global check regressed its configured
1h to 5m off the back of an opencode-go observation -- an unrelated-provider
change this repair must not make. Measured:
base effective_cache_ttl('1h', 'minimax', 'MiniMax-M2.5') -> '1h'
payload -> '5m'
here -> '1h'
Scope note on the evidence: the delayed-read run covered qwen3.8-max and
glm-5.2. The rule is keyed on the route, not the model, so the deployed
qwen3.7-plus is covered by it but has never itself been measured. This change
restores *sending* the requested 1h marker; it does not establish that the
provider honours 1h retention. Those stay separate claims, and the provider
labels every write `ephemeral_5m_input_tokens` whatever ttl was requested, so
only a delayed read past five minutes with no intervening call can settle it.
Tests: 11 new cases covering the deployed model, marker shape, precedence
against the generic clamp, cache eligibility surviving the repair, negative
controls for unrelated routes, provider case normalization, and a closed-set
audit of every marker-emitting call site. Ten mutations were applied to scratch
copies and each turned the suite red, including hoisting the generic clamp back
above the allowance, dropping opencode-go from ALIBABA_FAMILY_PROVIDERS,
un-nesting the model denial, and adding a new unclamped sender.
Known risk, not discharged here: this marker shape has never been sent on the
Go relay. #77217 records the sibling Zen relay returning HTTP 400 on an
unexpected marker shape, so field validation must check HTTP status before it
looks at any cache counter.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QV9Ld4f5ZTqdZWncpfZcQh
Security review on #97701 (unsupportedpastels) found and reproduced a
metadata-signature collision: atomic replacement that preserves size
and mtime (deployment tools that restore metadata; equal-length JSON)
rotates the private key under an IDENTICAL (path, mtime_ns, size)
signature, so the cache kept serving the old identity — exactly the
failure the PR set out to close.
The stat idiom is right for config caches (guards a parse, mtime
collisions are harmless). It is wrong for a credential cache (guards
an identity). The key is now (path, sha256(content)):
- _read_sa_file() reads the file ONCE per call, returning both the
bytes and the digest key; on a miss the credentials are built from
that same snapshot via from_service_account_info — closing the
stat->read TOCTOU the reviewer also flagged (key and credentials can
never describe different bytes).
- Read failure degrades to the bare-path key and the SDK's own file
read: byte-for-byte the pre-signature behavior.
- Cost: one read + sha256 of a ~2KB JSON per cache probe, noise next
to the OAuth token mint the cache exists to avoid.
New regression test per the review: equal-length content swap with
mtime restored and atomic replace — asserts a NEW cache key and a NEW
Credentials object (fails on the stat-keyed version, reproducing the
reviewer's cache_keys_equal=True). Fake google-auth harness gains
from_service_account_info.
Suite 10/10; ruff green.
The /simplify-code reviewer caught a real regression in the signature-
keyed cache commit: the ADC-failure fallback still compared
`cache_key == "__adc__"` (string), but keys are tuples now — the
comparison is always False, silently disabling the retry that picks up
a service-account file added after startup. The lead's pre-verification
had checked sentinel collision, failure-path pop, and TOCTOU, but
missed this consumer of the OLD key shape.
Guard now tests the actual condition (`not resolved_path` — this
attempt was ADC) instead of a key literal, so it can't rot again if
the key shape changes. New regression test drives the full path: ADC
raises, SA file appears on re-resolution, retry succeeds — fails on
the tuple-comparison version AND on any future key-shape change that
breaks the guard.
Pattern-D fix (stale cache after out-of-band change): the Vertex
credentials cache was keyed on the service-account file PATH alone, so
rotating the file on disk (key revoked and re-issued, new identity)
kept serving tokens minted from the OLD Credentials object for the
life of the process. Operators rotate compromised keys precisely when
they most need the new identity to take effect.
The cache key is now the file's (path, mtime_ns, size) signature — the
established idiom (agent/skill_utils.py:414, hermes_cli/config.py:3343,
and the shape #89792 applies to model overrides). Rotation bumps the
signature, forcing one re-read; the superseded entry for the same path
is evicted on insert so the cache stays bounded at one Credentials per
file. ADC keeps a stable sentinel key ("__adc__",) and its existing
expiry/refresh handling; a stat failure degrades to the bare-path key,
i.e. exactly the pre-signature behavior.
Tests (tests/agent/test_vertex_adapter.py):
- rotation invalidates: rewrite + mtime bump -> new key, new
Credentials object, old entry evicted (fails on main: main serves
the first identity's object after rotation)
- stat failure falls back to bare-path key, never raises
- ADC sentinel stable across None/empty resolved paths
Suite 8/8 green; ruff green.
LiteLLM OpenAI->Anthropic translation copies tool-message content parts
verbatim, so the envelope-layout part-level cache_control landed at
tool_result.content[0] - a placement the Anthropic Messages schema rejects
with a non-retryable HTTP 400 that killed the whole turn (any tool-using
cron/session on a LiteLLM-fronted Anthropic route).
New envelope_tool_part_cache_markers_supported() predicate (keyed on the
existing _is_litellm_route token matcher) threads a tool_part_markers flag
through build_prompt_cache_plan / apply_anthropic_cache_control and all
four decoration sites (main loop x2, destination replan, MoA). On LiteLLM
routes role:tool messages carry no markers and the breakpoint budget
reallocates to the nearest eligible message; OpenRouter/Nous Portal keep
the part-level form they honor, native Anthropic layout unchanged.
- Preserve streamed assistant text in Desktop UI when message.complete delivers empty text.
- Prevent destructive hydration in Desktop useMessageStream over rendered text on empty completion.
- Recover stream buffer in finalize_turn when final_response is empty on healthy turns.
- Unify in-place blank assistant repair, watermark clone resolution, non-blank concurrent winner adoption, and batch row appends into a single atomic guarded SessionDB transaction.
- Synchronize canonical committed content to live in-memory messages dicts and preserve all-or-nothing rollback semantics on persistence failure.
Squash of the three commits on PR #96768 (net diff is tests-only: the
mid-series production hunk in agent/transports/codex.py was reverted
within the PR after review). Pins the cache-scope isolation invariant
for hosts that mint one physical session per response, plus the
system-prompt write-path lifecycle under a per-response session (#96570).
Bedrock's cachePoint rules are per-model-family AND per-field. Amazon Nova
accepts a cachePoint block in `system` and `messages` but rejects it inside
`toolConfig.tools`, failing the whole request with
ValidationException: Malformed input request: #/toolConfig/tools/18:
extraneous key [cachePoint] is not permitted
so every tool-enabled Nova turn fails, with no retry path and no way for the
user to turn cache markers off (#97281).
The adapter decided placement from one static allowlist that answers only
"does this model cache at all", never "in which section". Any family whose
placement rules differ breaks 100% of turns until someone edits the table and
ships a release — the same maintenance trap the `_NON_TOOL_CALLING_PATTERNS`
comment already admits to ("if a model fails with a tool-related
ValidationException, add it here").
Make Bedrock's own verdict authoritative alongside the table: classify the
rejection by the JSON pointer AWS returns, drop the marker for that one
section, retry the request once, and remember the verdict for the rest of the
process so later turns are built clean. The other sections keep their cache
markers, so Nova still gets system/messages caching instead of losing prompt
caching wholesale. This mirrors the module's existing self-heal idiom
(`is_streaming_access_denied_error` → non-streaming `converse()`).
Applied at all four boto3 call sites: `call_converse`, `call_converse_stream`,
and both Bedrock dispatch sites in `chat_completion_helpers` (the streaming
one is the path in the report). A rejection with no marker to strip returns
None so the caller re-raises instead of looping.
Fixes#97281
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gcoy6nLTg5R6FHHhjcLZEC
Port from anomalyco/opencode#44571: OpenAI caps prompt_cache_key at 64
chars (DeepSeek and Zai inherit the same limit via their OpenAI-compatible
APIs) and rejects longer values with HTTP 400. The Responses transport
already bounds keys via _bounded_prompt_cache_key, but the Chat
Completions transport passed caller-supplied keys (request_overrides,
top-level or extra_body) through unmodified on both the profile and
legacy kwargs paths.
Bound caller keys with the same pck_<sha256[:24]> hash shape codex.py
uses so both transports behave identically; blank keys are dropped
instead of sent empty. Hermes-generated keys were already safe
(content-addressed pck_ hashes).
- Anthropic classifier counts signature_delta and citations_delta payloads
(content-bearing delta types the transport emits — relay_llm.py handles
both) so signed-thinking/cited-text generation keeps ticking the fence.
- Fix the stale 'per streamed event' comment above the Anthropic
on_stream_event lambda (left over from the conflict resolution).
- Rename test_completed_response_without_stream_payload_does_not_tick to
test_completed_response_ticks_only_terminal_signals — the old name
contradicted its own assertion (dispatch + shim ticks are expected).
Follow-up for the salvaged substantive-progress fix, adapted to main's
three-hook architecture:
- Keep the dispatch tick and the completed-response shim tick (both
deliberate on main; one-shot terminal events that cannot defeat an
inactivity timeout) — update the two PR assertions accordingly.
- Pin the end-to-end bug: keepalive/empty-role chunks must leave
CompressionCommitFence stale, a substantive token must refresh it.
- Pin the fast-lane telemetry contract (#96945/#96963):
time_to_first_progress_ms records on the first frame of any kind via
the split-out _notify_aux_timing_response helper.
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.
- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
with a structural base-message identity check that fails closed on any
transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
compaction (codex_runtime), session reset/switch (run_agent), plus the
fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
estimation fallback (first request of a session).
A 413 is a byte-size error, but the recovery loop scored compression
progress with estimate_messages_tokens_rough, which deliberately prices
every image at a flat per-image token cost (so screenshots don't trigger
premature compaction). When the payload is image-dominated that check can
never pass: in the reporting session two vision_analyze results were
5,627,202 bytes (96.6% of the request body) but ~3K of the ~80K token
estimate, so every attempt reported no_progress, the budget burned, and
the session wedged permanently with 'max compression attempts (3)
reached' at 13% context usage.
Post-#97160, the 413 path already routes into compaction and compaction's
historical-media aging genuinely frees the image bytes — but the
token-scored yardstick could not see the megabytes it freed. Add
serialized_messages_bytes() (exact serialized payload size, measured
identically before and after each pass — a measurement, not an estimate)
and score the 413 progress check with it. Tokens remain for status
display only; the context-overflow branch keeps its token yardstick,
because that error IS a token-budget error.
Images are never evicted from live history outside compaction (cache
invariant); the original strip-from-history mechanism in this PR was
superseded by #97160's compaction-time aging and is dropped in salvage.
Salvaged from #88960. Fixes#47339.
_sanitize_tool_pairs() in ContextCompressor compared raw tool_call_id
strings without stripping whitespace, the same bug fa3ab2ffd just fixed
in agent_runtime_helpers.py / run_agent.py (_get_tool_call_id_static +
sanitize_api_messages). ContextCompressor has its own near-identical
reimplementation of the pair-repair logic that was left unpatched.
When assistant-side and result-side IDs diverge only in surrounding
whitespace, the compressor misclassifies valid results as orphaned and
replaces them with [Result unavailable] stubs — silent data loss on
every compression cycle that touches such pairs.
Apply the same .strip() fix to all three sites:
- _get_tool_call_id (extracts IDs from assistant tool_calls)
- result_call_ids accumulation loop
- orphaned_results filter predicate
Closes the sibling gap of fa3ab2ffd / #42405.
Widen #90001's compaction-time strip to cover the gaps #89965 identified,
applied at compaction only per the cache ruling (request-time eviction
changes the per-call prefix and breaks prompt caching; compaction is the
one sanctioned cache break):
- Rule 1b: the opening attachment (anchor == 0) ages out once a newer
tool-result image supersedes it. The reported session opened with a
~200KB poster that previously survived every compaction. The row keeps
a non-empty text placeholder, so the zero-user-turn guard (#58753) and
role alternation are untouched.
- Native {_multimodal: True, content: [...]} dict envelopes now both
anchor (newest is kept) and strip (older collapse to their
text_summary via _strip_images_from_tool_msg, which also drops the
stale api_content sidecar per #97125's drop_stale_api_content).
- All three wire shapes (Chat Completions image_url, Responses
input_image, Anthropic-native image) were already matched by
_IMAGE_PART_TYPES; tests now pin each shape explicitly, plus
determinism (double-run is a no-op returning the same object).
Refs #89938, #89965
_strip_historical_media anchors on the newest image-bearing USER message and
returns the list untouched when that anchor is index 0 or does not exist. A
session whose images arrive from tools rather than attachments therefore has
nothing to be "before": twenty vision_analyze results keep multi-MB of base64
in every request body, the provider answers 413, and the 413 handler's
recovery compaction lands right back in this function and frees nothing. The
reporter saw seven compactions in thirteen minutes, all below 200K tokens.
Age tool-result images on their own timeline: keep the newest one, since that
is the image the model is reasoning about, and strip every older one wherever
it sits, including inside the protected tail. The tail exists to preserve
conversational continuity, not to pin bytes the model has already moved past.
User-message images keep today's treatment exactly. The user anchor is
checked first, so a tool result that is the newest of its kind but still sits
before that anchor is stripped as it always has been, and the anchor message
itself is still kept byte-for-byte - test_compressor_zero_user_guard depends
on that.
Refs #89938
Review (Sol xhigh) on the prior commit found a P1: the generic strip-
and-retry fallback ("any non-retryable 400 with image parts present
strips and retries") was too blunt. It couldn't tell an actual
image-corruption 400 apart from an unrelated one — bad tool schema,
unsupported parameter, billing, content policy — that merely happened
to carry image parts in the request. Any of those would silently erase
vision history and retry the still-invalid request, degrading sessions
that were never bricked in the first place. That's worse than the bug
it was meant to fix.
Revert the generic fallback (agent/conversation_loop.py). Keep only
the classifier-routed path: FailoverReason.image_corrupt +
_IMAGE_CORRUPT_PATTERNS, checked before _IMAGE_TOO_LARGE_PATTERNS
because shrinking corrupt bytes can't repair them. Corrupt-image
wordings still route to strip-and-retry; everything else falls through
to normal (non-retryable) handling as before. Add xAI's second wire
wording for the same corruption class ("base64 string of provided
image cannot be decoded", returned on unaligned truncation vs "Invalid
PNG image." on aligned truncation) and a compound-message test pinning
that image_corrupt wins when a body matches both pattern lists.
Drop TurnRetryState.stripped_images_this_turn. It's unnecessary now
that only one branch is left: the branch already only retries when
_strip_images_from_messages reports it removed something, and that
helper strips every image part from the request in one pass — so a
second corrupt-image hit on the retried (now text-only) request has
nothing left to strip and falls through on its own. No separate
one-shot flag needed.
Add a run_conversation integration test at the sequenced-provider
layer: corrupt 400 on attempt 1, strip, retry succeeds on attempt 2,
with explicit before/after assertions on the outgoing image_url part.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
The permanent-brick class in #69078: xAI returns 'Invalid PNG image'
when a re-serialized image part in replayed history becomes
undecodable. The existing image-error patterns cover only Anthropic
'exceeds max dimension' wordings and 'model does not support images'
strings, so the classifier lands on a generic non-retryable 400 and
neither the shrink path nor the strip path fires. Every subsequent
turn (even bare text) fails identically because the poison stays in
history — the session is permanently wedged until deleted.
Two recovery layers, deliberately separate:
- Semantic split: new FailoverReason.image_corrupt with
_IMAGE_CORRUPT_PATTERNS ('invalid png image' / 'invalid jpeg image'),
checked BEFORE _IMAGE_TOO_LARGE_PATTERNS in both _classify_400 and
_classify_by_message. Corrupt bytes route to strip-and-retry, never
to the shrink path (shrinking corrupt bytes cannot help).
- Generic fallback: any non-retryable 400 whose outgoing messages
still contain image parts gets one strip-and-retry via the existing
_strip_images_from_messages helper, guarded by a new
stripped_images_this_turn one-shot flag on TurnRetryState. This
un-bricks the session for any current or future provider wording
without adding another pattern list to maintain.
Item 3 from the report (multimodal-part integrity across FTS
persistence + compaction handoff) is a separate investigation and
remains follow-up work.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: paultaki <paultaki@users.noreply.github.com>
_call_llm_impl applied auxiliary_max_tokens_param whenever
fast_compression_cap was non-None — but _compression_fast_lane_controls
passes an explicit caller max_tokens straight through, so a compression
call that set its own cap had the param force-injected onto providers
where _build_call_kwargs deliberately omits it (ZAI vision hard-400s on
max_tokens; GPT-5/Copilot require max_completion_tokens). Pre-fast-lane
main omitted the param for that exact call shape (verified via
subprocess pinned to the pre-PR base).
Gate the forced param on 'max_tokens is None' so it applies only to caps
the certified lane itself produced — the same guard the fallback path
already uses.
Regression test pins the pre-PR wire shape. Mutation-checked.
_run_protected_sync_provider_call propagates the forward-progress hook to
its daemon worker but not the new _aux_dispatch/_aux_provider_response
timing hooks (both threading.local). When compression takes the protected
path — the common case, since the summary call runs under
aux_interrupt_protection with a hard-cancel source — provider_dispatch_ms
and time_to_first_progress_ms were silently absent from telemetry.
Also collapse the two byte-identical save/restore context managers
(aux_progress_hook, _aux_timing_hook) onto one _aux_thread_local_hook
implementation so the propagation semantics can never drift between the
progress and timing slots.
Regression test drives _run_protected_sync_provider_call with both timing
hooks installed and asserts the worker-thread notifies reach them.
Mutation-checked (reverting the propagation fails the new test).