The #58327 dedup passes treat a repeated tool_call_id as garbage from a
retry/crash/resume glitch and drop it. That assumes tool_call_id is
globally unique, which it is not: llama.cpp emits a single constant id
for every tool call it ever returns (verified — three separate
completions from one server all carried the same id).
Under a seen-once-drop-forever rule, the SECOND legitimate tool result
of such a session looks like a duplicate and is deleted. From the second
tool call onward the model never sees any result: it announces its next
action, the turn ends, and the task is left unfinished. Bisected to
dba585c17 over a 2258-commit range; reproduced live on v0.19.0 (1/6 runs
completed a 4-step file task, vs 20/20 on the last release before that
commit, same model and server).
Key off OUTSTANDING calls instead of every id ever seen. Both original
protections are preserved: a replayed result still answers no pending
call and is still dropped, and duplicate tool_calls sharing an id within
one assistant message are still collapsed. A genuine new call that
reuses the id re-arms it first.
repair_message_sequence needs no change — it already resets its id set
per assistant message, so only the final pre-API pass mis-fires.
Live result after the fix: 8/8 runs complete, 17-26s each (was 1/6 with
runs hitting a 150s ceiling).
mcp 2.0.0 implements MCP revision 2026-07-28 and makes three breaking
changes Hermes sits on top of: `mcp.server.fastmcp` is gone, every model
field is renamed to snake_case (camelCase survives only as a
serialization alias, which pydantic does not expose to attribute
access), and the SDK's own HTTP stack moved from `httpx` to `httpx2`.
Bump the pin across the dev/mcp/computer-use extras and port the tree:
- `mcp_serve.py` and `agent/transports/hermes_tools_mcp_server.py` move
from `FastMCP` to `mcp.server.MCPServer`, which has the same
decorator/add_tool surface. The hermes-tools server already
synthesised `__signature__` from Hermes' JSON Schema, which is exactly
what 2.0's `add_tool` reads.
- SDK model reads go through `mcp_field(obj, snake, camel)`, which reads
both spellings. A single-spelling read fails *silently* on the other
generation — empty tool schemas, dropped structured content, tool
results vanishing from sampling conversations — and `mcp` is an
optional extra users install at their own version.
- `sdk_httpx()` resolves the httpx flavour from the SDK's own transport
module, so objects handed to `streamable_http_client`, the `sse_client`
factory, and the OAuth metadata helpers come from the module the
installed SDK actually imports.
- HTTP support is gated on either streamable-HTTP entry point, not just
the deprecated alias 2.0 removed.
- OAuth: `OAuthClientProvider` lost its `timeout` argument (the
configured `oauth.timeout` now bounds the callback waiter's own poll
loop, where the browser round-trip was always awaited), and
`callback_handler` must return `AuthorizationCodeResult` rather than a
tuple. 2.0 also validates the RFC 9207 `iss` parameter, so the
callback handler and paste fallback capture it.
`mcp`/`mcp-types` 2.0.0 are inside the 14-day `exclude-newer` window, so
two narrow `exclude-newer-package` entries unblock `uv lock`, annotated
for removal on or after 2026-08-11. `httpx2` needs no exemption: 2.7.0 is
already outside the window and satisfies mcp's floor.
Refs #69931
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A tiny title model that ignores the 3-7 word titling task and answers
the user's first message instead used to have its whole reply stored
(truncated at 80 chars) as the session title. Truncating an assistant
blob still leaves an assistant blob — generate_title now rejects output
over 12 words and returns None, letting maybe_auto_title retry on the
next exchange. The 80-char truncation remains for genuine-but-wordy
titles that pass the word bound.
Gemini 3+ models require explicit tool call IDs on functionCall /
functionResponse parts in replayed history; without them parallel tool
calls can be rejected or mispaired. The native adapter now:
- threads the model id into request building and includes ids for
Gemini >= 3 (version-gated: 2.x rejects unexpected id fields)
- preserves provider-returned functionCall.id on both non-streaming
and streaming responses instead of always minting a random one
AGENTS.override.md now takes priority over AGENTS.md in both startup
project-context loading (prompt_builder) and progressive subdirectory
hint discovery (subdirectory_hints). Lets developers keep a personal,
typically-gitignored override next to committed project instructions
without editing the tracked file.
Port from anomalyco/opencode#40707: connection-establishment and DNS
failure messages wrapped in generic exceptions (RuntimeError from local
shims, MCP bridges, SDKs re-raising without chaining) fell through to
FailoverReason.unknown, which misses the retry loop's eager transport
fallback — the full retry budget burned against a dead endpoint before
provider fallback.
New _CONNECTION_MESSAGE_PATTERNS (connect refused, no route, network
unreachable, DNS phrasings across Python/glibc/macOS/Node, fetch failed,
Envoy upstream connect error) classify as retryable timeout via
_classify_by_message, mirroring _TIMEOUT_MESSAGE_PATTERNS. Mid-stream
disconnect strings are deliberately excluded — they keep their
_SERVER_DISCONNECT_PATTERNS routing (large-session compression).
Tracker #79686 P3. Every skill mutation — curator, agent, or user — now
appends one entry to the append-only JSONL ledger at
~/.hermes/skills/.curator_ledger.jsonl, with per-file before/after
manifests whose contents are stored content-addressed (sha256-deduped)
under ~/.hermes/.curator_backups/blobs/.
- tools/skill_ledger.py: append/list/get, blob store, actor derivation
(curator|agent|user), single-entry rollback that takes a pre-rollback
safety entry first and FAILS CLOSED when that capture fails (consistent
with the whole-run tarball rollback hardening from #63366). Path
containment check so a hand-edited ledger can't write outside
HERMES_HOME.
- Hooked all three choke points: skill_manage() dispatch (all actors,
delete intent recorded via absorbed_into/archived evidence),
archive_skill()/restore_skill(), and curator auto-transitions (tagged
actor=curator via a ContextVar override).
- Ledger failures never block the mutation — telemetry, not a gate.
Config gate skills.ledger (default true).
- hermes curator ledger [--skill NAME] [--limit N] and
hermes curator rollback <entry-id> (whole-tree snapshot rollback
unchanged).
- Optional TTL purge of skills/.archive/: curator.archive_ttl_days
(default 0 = never) + explicit hermes curator purge, recorded in the
ledger with before-blobs so purges stay recoverable.
- Docs: curator.md sections on the ledger, single-edit rollback, and
archive TTL purge.
Curator invariants unchanged: only created_by:agent skills auto-transition,
never hard-delete autonomously, pinned exempt; foreground user deletes stay
hard-delete (and are now recoverable via the ledger).
Closes#45778, #50875. Tests adapted from #50261 by @yu-xin-c.
Per project policy, .env / HERMES_* env vars are reserved for
credentials; behavioural settings belong in config.yaml. Replaces
HERMES_DETERMINISTIC_EMPTY_GUARD and
HERMES_EMPTY_RETRY_COST_THRESHOLD_USD with an additive
agent.empty_response_guard section:
agent:
empty_response_guard:
enabled: true # false = legacy fixed 3-retry behaviour
cost_threshold_usd: 0.25 # per-attempt cost that halves the budget
- hermes_cli/config_defaults.py: new documented subsection under agent
(additive key, no config-version bump needed).
- agent/empty_response_guard.py: resolve_guard_settings() maps the
section to (enabled, threshold) with fail-open tolerance for
malformed values; guard_enabled()/_cost_threshold_usd() now read the
init-resolved agent attributes instead of os.environ.
- agent/agent_init.py: resolves the section once at init into
agent._empty_guard_enabled / agent._empty_guard_cost_threshold_usd,
following the existing tool_use_enforcement extraction pattern.
- Tests updated to config-attr injection; new TestResolveGuardSettings
covering malformed sections, YAML string booleans, bad thresholds,
and a DEFAULT_CONFIG sync check; new integration test proving
enabled:false restores the legacy 1+3-call behaviour.
Requested by isak-ialogics on PR #75115.
Every empty-response retry re-sends the full conversation input at full
price. On large contexts a single turn that produces no visible output
could bill the user several dollars across the 3-retry + fallback-chain
walk (reported: ~$2.33 for one empty answer on a ~26K-token session).
Signaled refusals (finish_reason=content_filter, Anthropic refusal
stop_reason, guardrail interventions) are already terminal today and
never reach this loop. The uncovered class is *unsignaled* refusals:
the provider returns 200 with zero output tokens and a generic finish
reason. Those are deterministic — resending the identical prompt
reproduces the same empty — so burning the remaining retry budget only
multiplies the charge.
New agent/empty_response_guard.py, two independent guards, both failing
OPEN to today's behaviour:
- Deterministic-empty detection: two consecutive empty attempts with
usage present, output_tokens == 0 (reasoning tokens count as output),
and identical (model, provider, finish_reason) skip the remaining
retries and go straight to the fallback chain — a different model may
well answer. Missing usage, nonzero output, or any signature change
keeps the full budget.
- Cost-aware retry budget: when one attempt's estimated input cost
exceeds HERMES_EMPTY_RETRY_COST_THRESHOLD_USD (default $0.25), the
empty-retry budget drops 3 -> 1 for that streak. Unknown pricing or
included/subscription routes are untouched.
At exhaustion the status trace now includes the estimated cost of the
empty attempts so the charge is at least explained in-session.
Streak state lives on the agent and self-clears whenever
_empty_content_retries resets to 0, transparently honouring every
existing reset site (turn start, tool success, compaction, fallback
activation) without touching them.
Set HERMES_DETERMINISTIC_EMPTY_GUARD=0 to disable both guards.
Tests: tests/agent/test_empty_response_guard.py (26 unit tests) plus
two loop-level integration tests in tests/run_agent/test_run_agent.py
proving the api_call reduction and the fail-open path.
Refs NS-503.
Switching Desktop profiles mid-session changes HERMES_HOME but not the
platform scope, so get_skill_commands() kept serving the previous
profile's skill list. A skill only available under the new profile then
looked like a cache miss to callers such as slash.exec, which fall
through to the slash_worker dead path (#88023).
OpenAI enabled the large-context window for ChatGPT-subscription Codex
accounts (announced by @thsottiaux Aug 16 2026; previously API-key-only).
Live re-probe the same day: 911,276 input tokens completed OK on
gpt-5.6-sol; ~925K+ rejected with context_length_exceeded (1.05M window
minus reserved output headroom). terra, luna, and gpt-5.4 all completed
900,026 tokens OK. The Codex catalog still advertises 272K, so the
stale-advertisement override from #87981 is the right lever — this just
raises its value 350K -> 900K.
gpt-5.5 and gpt-5.4-mini still enforce 272K live (rejected 500K) and
remain excluded. Override semantics unchanged: fires only on an
exactly-272,000 advertisement; any live catalog change is trusted
verbatim.
Bot Chats created before the epoch mechanism persisted prompts with no
protocol section and no stamp — the staleness check only fires on
stamped prompts, so pre-existing bots would never learn to message
teammates. stored_bot_chat_prompt_needs_upgrade() migrates them: one
rebuild, title-gated to Bot Chat, only when the probe would actually
emit a section (SOUL-append legacies and unmanaged installs are left
alone — rebuilding those would loop). The rebuilt prompt carries the
stamp, so the upgrade can never re-fire.
E2E v3b through the real restore path: legacy Bot Chat upgraded once
then verbatim-reused; legacy regular sessions byte-untouched.
tests/agent/ 4648/4648.
Bot Chats break the "new sessions come often" assumption behind
build-once system prompts: capability edits used to sit invisible until
/new or compression, and the frozen birth date became misinformation.
- tools/bot_mode_probe.py: capability_fingerprint() hashes the profile's
capability surface (disabled skills, toolset pins, MCP config, SOUL.md,
installed skills, Bot-Mode roster); Bot Chat prompts embed the 12-hex
epoch stamp
- agent/conversation_loop.py restore path: stored Bot Chat prompt whose
epoch mismatches disk → ONE rebuild (through a cleared skills-prompt
cache so new installs appear), persisted so the next turn reuses the
new bytes verbatim. Prompts without a stamp — every non-Bot-Chat
session — never take the branch; probe failure fails closed to reuse
- agent/system_prompt.py: Bot Chat prompts are timeless — the
"Conversation started:" date is dropped (timezone kept); no ticking
fields in an eternal session
- tui_gateway: _sync_bot_capabilities at turn start rebuilds the live
agent (tool definitions are construction-baked) when the fingerprint
moves, same session id/history, with a user-visible notice
Cache stance: this is the /model exception applied to capabilities — a
loud, user-initiated, once-per-change prefix break. Unchanged state
hashes identically and stored bytes are reused verbatim (E2E-proven).
Validation: 9 probe unit tests incl. per-axis fingerprint changes;
E2E v3 against the real restore path (fresh build → verbatim reuse →
skill install → single refresh w/ new skill in index → verbatim reuse;
regular sessions dated, unstamped, never refreshed); tests/agent/
4647/4647.
Live desktop E2E caught a write-ordering bug the automated E2E missed:
tui_gateway applies pending_title to state.db AFTER the first turn, but
the system prompt builds at turn START — the DB-title gate saw nothing
and the Bot Chat was cached protocol-less forever. The gateway now
hands the agent its intended title at construction and the gate checks
the hint first, DB second (CLI/messaging-gateway paths unchanged).
Live-verified on the running desktop: fresh bot's Bot Chat persisted
with the protocol section, handle, and roster in its system prompt;
regular sessions and SOUL.md untouched.
Per review: the protocol belongs only in official Bot Mode interactions,
not every session on a managed install. The prompt builder now injects
the section only when the agent's session row is titled "Bot Chat"
(BOT_CHAT_TITLE, matching the desktop's createCanonicalChat pin and the
`hermes -p <bot> chat -c "Bot Chat"` resume target). Regular sessions
never carry it; the desktop composer middleware owns @mention sends.
Title is read once at first prompt build and the rendered prompt is
cached + DB-restored — cache-safe. E2E against the real AIAgent +
SessionDB: absent in an untitled session, present in Bot Chat,
byte-stable across rebuilds, absent after retitle, absent with the
flag off. Overhead unchanged (~916B, Bot Chat sessions only).
Replaces the plugin-side SOUL.md protocol append: on Bot-Mode-managed
installs (any profile carrying ui_meta['hermes-bots']) the prompt builder
injects the "Messaging other agents" section into every session of every
profile — including headless `hermes -p <bot> chat` sessions a teammate
starts — so bot handoffs work without mutating user-authored SOUL files.
- tools/bot_mode_probe.py: silent-when-unmanaged probe, cached per
(process, home), keyed off the agent's OWN home (not ambient
HERMES_HOME); silent when SOUL.md already carries the legacy section
- agent/system_prompt.py + agent_init.py + config_defaults.py: wired as
agent.bot_mode_protocol (default True), stable tier, byte-stable
across rebuilds (E2E-verified against the real build_system_prompt)
- tui_gateway profiles.list gains bot_mode_protocol capability flag;
the bundled plugin gates ALL SOUL protocol writes on it (backfill,
composeSoul, Edit save) — older gateways keep the SOUL-append path
- overhead: ~916 bytes, only on Bot-Mode installs; zero elsewhere
Supersedes the SOUL backfill half of Hermes-Bot-Mode#99 (credit
@kaduxo — the handle fix, `hermes profile list` correction, and
idempotent-append guards from that PR ship in the bundled plugin).
The Codex /models catalog advertises 272K for the gpt-5.6 (sol/terra/luna)
and gpt-5.4 slugs, but the backend actually accepts ~371K input tokens
(verified live against chatgpt.com/backend-api/codex/responses, Aug 16 2026:
~371K completed OK on all four slugs; ~382K+ rejected with
context_length_exceeded). 350K keeps ~22K margin under the observed ~372K
enforcement.
The bump applies ONLY when the resolved value is exactly the known-stale
272,000 advertisement — any other advertised value (higher or lower) is
trusted as a real server-side change, so a future catalog correction
deactivates the override automatically. gpt-5.5 and gpt-5.4-mini both
genuinely enforce 272K (rejected 360K live) and are excluded.
The sequential tool path only noticed a user interrupt after the running
tool returned: with the deadline disabled it ran the tool inline (fully
blocking), and with a deadline it waited in 5s slices without ever
checking agent._interrupt_requested. Any tool without cooperative
is_interrupted() polling (image_generate, tts, transcription, skills
sync, ...) held the whole turn hostage — the reported symptom was a
redirect queued ~40s behind a FAL image generation + upscale pass.
Executor backstop (class fix, covers ALL tools):
- _run_sequential_tool_execution_middleware always dispatches on the
daemon worker (timeout None no longer means inline blocking) and polls
the interrupt flag every 1s.
- On interrupt: 3s cooperative grace (mirrors the concurrent path), then
synthesize a cancelled tool result (_ToolCancelledResult), emit the
terminal post_tool_call with status=cancelled, and abandon the worker.
- _ToolCancelledResult suppresses downstream post-hook double emission
exactly like _ToolTimeoutResult, so an abandoned worker finishing late
cannot report success for a cancelled call.
- clarify (interactive, _NEVER_PARALLEL_TOOLS) keeps the inline path —
it owns its own human wait.
Cooperative layer in the reported offender:
- image_generation_tool: blind handler.get() (generation + Clarity
upscale) replaced with _wait_fal_result(), which polls is_interrupted()
in 0.5s slices and raises ImageGenerationInterrupted immediately.
- _upscale_image propagates the interrupt instead of swallowing it into
the "upscale failed, use original" fallback.
Message alternation is preserved: the cancelled result is a normal tool
result for the call_id. Sabotage-verified: with the old wait loop
restored, the new tests fail (tool blocks full runtime); with the fix
they pass in ~4s.
'database disk image is malformed' contains the word 'disk', so
classify_persistence_error bucketed SQLITE_CORRUPT / SQLITE_NOTADB
failures as 'disk' and the turn-completion explainer told users to
free disk space for a structurally damaged state.db (the #77386-family
misdiagnosis, reproduced in the v0.20.0 malformed-DB incident report).
- hermes_state: new 'corrupt' bucket in PERSISTENCE_ERROR_CAUSES,
matched via _DB_CORRUPTION_MARKERS BEFORE the locked/disk buckets
- run_agent: explainer text for 'corrupt' points at hermes doctor and
explicitly says freeing space will not help
- cron explainer-variant suppression picks the new variant up
automatically (it iterates PERSISTENCE_ERROR_CAUSES)
Follow-up to #87400: drop the max_iterations and prompt_file knobs from
auxiliary.background_review. The aux model routing (provider/model/
base_url/...) predates #87400 and stays; the enabled switch and the
usage telemetry stay. The fork's iteration budget returns to the
historical hardcoded 16.
Persist fork token usage under session_model_usage task=background_review,
emit a per-fork completion log line, and expose enabled/max_iterations/
prompt_file so operators can see and bound the automatic review cost.
Address review feedback: load auxiliary.background_review once per spawn,
classify completion logs by summarize action prefixes, treat explicit
api_call_count=None as the documented default of 1, and WARNING on the
fail-open enabled-gate path.
Salvage hardening on top of #87308 (thanks @Dudeman456):
- Tri-state verdict: inconclusive probes (binary missing, --help
failed/timed out) return None and fall through to the normal spawn
path, preserving the established 'Could not start Copilot ACP
command' error instead of masking it. This also fixes the two
test_copilot_acp_client HOME-env regressions that went red on the
PR: their mocked-Popen path was intercepted by the new unmocked
subprocess.run probe.
- Cache definitive verdicts per binary path so CLIs that DO support
--acp pay the ~50ms --help cost once per process, not per prompt.
- Skip the probe entirely when custom ACP args don't include --acp.
- Fix the help-text regex: the old pattern never matched '[--acp]'
(leading '[' is neither start-of-string nor whitespace) and \b
after 'p' matched '--acpfoo'.
- Hermeticity: stub subprocess.run in the two HOME-env tests; add 6
probe-specific tests (fast-fail, fall-through, caching, skip).
CopilotACPClient unconditionally passes [self._acp_command] +
self._acp_args (default ['--acp', '--stdio']) to subprocess.Popen.
When the resolved CLI doesn't accept --acp (e.g. Claude Code
v2.1.233, where 'claude --acp --stdio' exits 1 with
'error: unknown option') the subprocess dies in ~250ms with the
error on stderr, but the parent ACP loop has no fast-fail for this
shape and waits the full child_timeout_seconds (default 600s,
observed 109s+ before user interruption) for stdout that never
arrives.
Add _acp_supported() that probes the CLI's --help output for the
--acp flag in ~50ms, then call it at the top of _run_prompt before
any spawn happens. When the probe fails, raise a RuntimeError that
names the unsupported flag, lists the expected fix (install
@github/copilot late 2025+, or set HERMES_COPILOT_ACP_*), and
returns control to the caller in ~280ms instead of hanging the
delegate_task parent for hundreds of seconds.
Measured locally against Claude Code v2.1.233:
- Before: delegate_task acp_command=claude hangs 109s+ then
returns tokens={input:0, output:0}.
- After: delegate_task acp_command=claude raises RuntimeError
in 280ms with a clear actionable message.
This does NOT change behavior for supported CLIs (the new
@github/copilot ships with --acp) — the probe returns True and
the spawn proceeds unchanged.
Refs the bundled claude-review-delegate skill which already
documents this class of transport-mismatch pitfall for users
who call 'claude -p' directly; this fix closes the same gap for
the delegate_task MCP path.
`hermes config set` and JSON-mode editor saves store lists as quoted
strings (e.g. '["skill-a","skill-b"]' or "['memory']"). Both disable
filters treated such a string as a single name, so curated disable
lists silently filtered nothing with zero diagnostics.
Add parse_config_string_list() in agent.skill_utils and use it in
_normalize_string_set (skills.disabled / platform_disabled) and at
every agent.disabled_toolsets read site: tools_config resolve +
reconcile, CLI, gateway agent construction (both sites), cron
scheduler, and prompt_size. A scalar string still names a single
entry (#13026); malformed JSON falls back to the single-name
behavior instead of raising.
Fixes#86661
Self-review follow-up, caught by benchmarking the previous commit.
Widening the custom-provider capability-lookup gate to `is_anthropic_wire or
_is_litellm_route(...)` made EVERY chat_completions route with a litellm-ish
provider/host enter the lookup, including non-Claude models that the grant
branch below can never match. Measured on a route with no config.yaml
(the uncached worst case) that was ~7.5us -> ~1528us per evaluation.
Narrowed the gate to the exact condition the LiteLLM branch grants on
(chat_completions + Claude + litellm route), computed once into a local and
reused by the branch itself so the predicate no longer runs twice.
Measured with a realistic config.yaml present (mtime cache warm), vs
origin/main:
live-agent policy 20.6us -> 61.7us
destination planning 219.3us -> 347.7us
Sub-millisecond and scoped to the routes that actually opted in. The
earlier 1.5ms figures were a tempdir artifact: load_config_readonly's
mtime cache cannot engage when no config.yaml exists, which is never true
of a real install. Non-LiteLLM and non-Claude routes are unaffected
(openrouter Claude measured flat at ~7.9us).
Tests: 82 passed across the policy and TTL-propagation modules.
Self-review follow-up. The previous commit fixed substring matching on the
HOST but left the provider-id side as a bare substring, so a user-named
provider like `custom:notlitellm` or `mylitellmthing` still matched and was
handed Anthropic markers — the same bug class, half-fixed.
Both signals now match `litellm` as a whole delimited token via a shared
helper. Real spellings (`litellm`, `custom:litellm`, `litellm-router`, and
the already-lowercased `LiteLLM`) still match; lookalikes no longer do.
Tests: 71 passed. Adds lookalike-provider and real-spelling guards; both
new guards mutation-checked. Differential matrix over 2688 configs vs
origin/main: 60 changes, every one a Claude model on a genuine LiteLLM
route getting the envelope layout, zero pre-existing routes altered.
Follow-up to the salvaged LiteLLM cache grant. The grant itself is right;
four things about how it was scoped were not.
1. Layout. The branch returned the native inner-block layout
(use_native_layout=True) on api_mode == "chat_completions". That layout
writes a TOP-LEVEL msg["cache_control"] on role:tool and empty-content
messages and depends on the Anthropic adapter to relocate it into the
block — but that adapter only runs for api_mode == "anthropic_messages"
(agent/transports/anthropic.py registers there), and the
chat_completions transport does no relocation. Measured on a 3-tool-turn
transcript: 2 of the 4 available breakpoints landed on markers the
provider never sees. Worse, when LiteLLM itself relocates a top-level
marker for an OpenRouter-backed Claude route
(OpenrouterConfig._move_cache_control_to_content), the marker lands on
an empty assistant turn and produces a cache_control-marked empty text
block — the HTTP 400 "text content blocks must contain" shape already
guarded in agent/anthropic_adapter.py (#69512). Switched to the envelope
layout, matching every other OpenAI-wire grant in this function:
4 of 4 breakpoints honored, zero empty blocks.
2. Host matching. `"litellm" in base_url_hostname(...)` is the substring
false-positive class base_url_hostname's own docstring warns against; it
granted Anthropic markers to notlitellm.example.com,
foolitellmbar.example and friends. Replaced with a label-token match in
a named helper, so "litellm" must be a whole dot- or hyphen-delimited
token. All three of the original test hosts still match; a "litellm"
path segment on an unrelated host still does not.
3. Transport gate. `not is_anthropic_wire` also swept in codex_responses,
bedrock_converse and codex_app_server. Gated on
api_mode == "chat_completions" explicitly.
4. Operator override. The grant is inferred from a provider/host name, but
the custom-provider capability lookup was gated on is_anthropic_wire, so
an explicit `prompt_caching: false` for the route+model was honored on
/v1/messages and silently ignored on /v1/chat/completions. The lookup
now also runs for a LiteLLM route, and its layout follows the transport
rather than the declaration (an explicit `true` must not promote a
chat_completions request to the native layout).
Tests: 64 passed. Adds the wire-shape contract the original matrix was
missing (asserts no breakpoint sits on the message envelope, rather than
only checking the returned tuple), plus lookalike-host, other-transport,
and both operator-override directions. All five guards mutation-checked —
reverting each fix turns the corresponding test red.
anthropic_prompt_cache_policy() only granted Anthropic cache_control
markers to LiteLLM over the native Anthropic wire
(api_mode == "anthropic_messages"). A LiteLLM deployment exposing the
OpenAI-compatible surface instead (/v1/chat/completions, /v1/messages
-> 404) matched no grant branch and fell through to (False, False): no
cache_control injected, the system prompt sent as a plain string, and
the provider serving zero cache hits -- the entire prompt re-billed at
full price on every turn. Silent: no error, no warning, usage simply
shows 100% uncached input forever.
Add one branch after the is_anthropic_wire/is_claude case that grants
caching to Claude-family models on a LiteLLM endpoint regardless of
wire, with the native inner-block layout. Same failure class already
documented in-function for Qwen/DashScope.
Design:
- Gated on the Claude family only (is_claude); a Gemini/GPT/Qwen route
through the same proxy must not receive markers (they may reject the
cache_control block format -- cf. the DeepSeek/OpenCode exclusion).
- Matches on provider string OR base_url host, since provider naming
varies per install (litellm, custom:litellm, or a bare custom alias
pointed at a LiteLLM host).
- prompt_caching.cache_ttl: false still wins (the _cache_disabled early
return is untouched).
- Generic strict OpenAI-wire custom providers (e.g. Fireworks) remain
excluded -- verified by the existing over-reach regression test.
Tests: adds TestLiteLLMOpenAIWire covering the grant (several model
spellings x provider/host signals), no-over-reach (non-Claude on the
same proxy get nothing; operator disable wins), and adjacent behavior
(LiteLLM in Anthropic proxy mode still native layout). Full module:
43 passed.
Closes#84506. Original diagnosis, patch design, and measurements by
@ottosulin.
A later hygiene idle-timeout write can replace an aux-model cooldown on
the shared column. Drop the in-memory timer on that refresh so the
in-agent compressor is not still blocked after the DB row is hygiene.
Session hygiene persists compression_failure_cooldown_until after a
30s no-progress watchdog so the pre-agent pass can skip. The
in-conversation compressor read the same column and then refused to
run even though its own budget is sufficient.
Ignore hygiene idle-timeout errors on the in-agent path. Real
aux-model faults such as rate limits still block.
Fixes#86972
Gemini bills thought tokens against maxOutputTokens/max_tokens, so a
global 4096 cap can be fully consumed by thinking on the first
request, leaving zero content tokens and aborting after 4
continuations. When thinking is enabled, raise the effective output
cap to the 65,535 ceiling on both the native and chat-completions
paths.
Refs #83915
The "Conversation started:" line carried a bare date (%A, %B %d, %Y). Tools
that accept instants -- nutrition, calendar and similar MCP servers -- reject
naive datetimes and require an explicit UTC offset, so the model had to infer
EST vs EDT from the date alone. Near a DST boundary that is a coin flip, and a
wrong guess does not error: it silently writes the record onto the wrong day.
Append the IANA zone (when configured), the zone abbreviation and the UTC
offset, e.g.:
Conversation started: Saturday, August 15, 2026 (America/New_York, EDT, UTC-04:00)
get_timezone() returns None when no timezone is configured; in that case the
line falls back to the abbreviation and offset of the server-local (still
tz-aware) time, so behaviour is unchanged for users who never set one:
Conversation started: Saturday, August 15, 2026 (EDT, UTC-04:00)
Daily byte-stability is preserved -- the property the date-only format exists
to protect (PR #20451). Zone name, abbreviation and offset are all constant for
the whole day; they shift only at a DST transition, where a change is correct.
The static-prefix reconstruction guard in _restore_plugin_sections matches on
"\n\nConversation started:" and is unaffected by a suffix after the date.
test_datetime_is_date_only_not_minute_precision used `re.search(r":\d{2}")`
over the whole line as a proxy for "no time-of-day". A UTC offset also matches
that pattern, so the check now applies to the date portion (everything before
the zone parenthetical) and the invariant is tightened rather than relaxed:
- test_datetime_includes_utc_offset asserts the offset is present
- test_datetime_line_is_stable_across_rebuilds asserts two rebuilds in the
same day produce a byte-identical line
Fixes#87403
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds a bounded fallback path in turn_finalizer.py that always records
a terminal timed_out outcome via _record_task_failure (CAS receipt path)
when the iteration budget is exhausted, regardless of whether the normal
fallback paths (interrupted/failed/anomalous exit_reason) were eligible.
Previously, a kanban worker whose budget was exhausted but whose turn was
interrupted, failed, or exited with an anomalous reason would silently
leave its task in an ambiguous lifecycle state — the dispatcher would
eventually detect it as a crashed or protocol-violation worker, but the
failure was not bounded and could take a full tick cycle to reconcile.
The CAS invariant in _end_run (WHERE ended_at IS NULL) guarantees
idempotence: if another path already closed the run, the call is a no-op.
Extracted the inline kanban-budget-exhausted recording into a shared
helper function (_record_kanban_budget_exhausted) used by both the
existing iteration_limit_fallback path and the new bounded fallback path.
Closes#87096
CI caught two rotation-path regressions from the unbounded clone: the #47202
pre-publish flush writes the rotator's OWN input transcript to the parent
(above the start-watermark), and the clone was duplicating it into the child
alongside the handoff. publish_compression_child gains watermark_ceiling —
the MAX(id) captured immediately BEFORE that flush — so only rows in
(watermark, ceiling] (genuinely foreign concurrent appends) clone across.
Ceiling capture failure falls back to no tail preservation (historical
behavior) rather than risking duplication. Ceiling-exclusion test added.
CI caught the sibling site the in-place fix missed: legacy (non-in-place)
compression rotates via publish_compression_child, where a mid-summary
append previously stranded in the closed parent. Same watermark + pure-SQL
column clone as archive_and_compact, with session_id rewritten to the child.
Lineage-guard test flipped to pin the appends-flow-freely contract; rotation
watermark tests added (tail follows the child; None = historical behavior).
Redesign of the #75316 class (supersedes the approach in PR #87307).
Root cause family: the compression lock fenced ORDINARY transcript appends
for the whole slow provider-summary call. Turns died as
session_persistence_failed whenever a message overlapped a compression
(#74568, #77386, #75083), stale dead-PID locks blocked writes for the full
TTL, and the busy-wait mitigation (#75264) was an order of magnitude shorter
than real summaries. Separately, the commit archived from a pre-call
snapshot, so rows appended mid-compression were swept into the archive.
Design: the commit transaction is already exclusive — no lock phases needed.
1. Appends never check compression_locks. The lock's only job is stopping
two compressions colliding; it keeps that job. The whole stale-lock /
busy-wait symptom family dies as a class.
2. Watermark captured in the DB at compression start
(get_active_message_watermark = MAX(id) of active rows) — not from
in-memory message dicts, which carry no row ids in production.
3. archive_and_compact(watermark=, lock_holder=): one transaction verifies
the holder still owns an unexpired lease (a reclaimed lease cannot
publish a stale compaction), archives the snapshot, inserts the compacted
set, and re-sequences the concurrent tail (id > watermark) via a
pure-SQL column clone — every column except id survives byte-exact
(api_content, platform_message_id, reasoning sidecars, token counts),
FTS triggers index the clones naturally, originals stay archived and
recoverable. watermark=None preserves the historical behavior.
Removed: the append-side compression fence in _check_transcript_write_guards
(with rationale note), making the _COMPRESSION_BUSY_WAIT_S retry lane
unreachable from append paths (kept for other callers).
Tests: 12 new (watermark contract, column-exact clone, commit fence incl.
lease-lost/expired/rollback failure injection, append-vs-commit race);
busy-retry suite flipped to pin the new contract; sabotage-verified (5 fail
with the watermark disabled, 12 pass restored); E2E through the real
compress_context seam with a mid-summary append landing and surviving.
Adds a `modify` response type to pre_tool_call hooks so a hook can
transform tool arguments before the tool executes, instead of repairing
results afterwards via post_tool_call.
- hermes_cli/plugins.py: _dispatch_pre_tool_call_hooks() fires hooks once
and returns (block_message, modified_args); modify directives
shallow-merge into an accumulated dict built from the original args.
- agent/shell_hooks.py: _parse_response() accepts both the canonical
{"action": "modify", "args": {...}} and Claude Code-compatible
{"decision": "modify", "tool_input": {...}} wire formats.
- model_tools.py, agent/tool_executor.py, agent/agent_runtime_helpers.py:
dispatch sites migrated; modified args applied before execution.
- Docs + 10 new tests (merge semantics, precedence, block interplay).
Salvaged from PR #28953. Best fix for #18988.
#87326 shipped the lean-compaction capability on the compressor; this adds
the config.yaml surface (compression.tail_mode: legacy|lean, default
legacy), DEFAULT_CONFIG entry, and docs on both the dev-guide compression
page and the user-guide configuration page, with zh-Hans parity.
- _build_anchor_index(): regex-harvests PR/issue numbers, SHAs, branches,
file paths, error strings, handles, URLs from the compacted region into a
bounded indexed summary section. LLM-free, so needle identifiers cannot be
paraphrased away (the GUI-lineage failure class: 10/15 verbatim-or-nothing
golds). Doubles as session_search query-anchor map.
- evals/compaction/test_region_scoping.py: sentinel tripwire proving the
summarizer input carries ONLY the compacted region (head/tail sentinels
never reach the serialized turns body) in both legacy and lean modes.
- _digest_worthy() drops no-signal tool rows before chunking (GUI-lineage
digests were starving on tool-noise)
- eval recovery sim now uses in-memory SQLite FTS5 + BM25 (production
session_search engine) instead of term-frequency scoring
- recovery query writer sees the digest section (front of context) so it can
mine anchor identifiers