archive_and_compact() soft-archives every active row with compacted=1 and
then re-inserts compacted_messages as fresh live rows. When the
compressor's protected tail rides inside that list verbatim - which is
the normal batch-compaction shape ([summary] + tail) - the tail's
ORIGINALS end up stored twice per compaction: (active=0, compacted=1)
next to their live clones. search_messages() recalls both flags without
DISTINCT, so every carried-forward message came back once per compaction
(measured up to 4 identical hits) and was mislabeled to users and the
agent as archived "summarized away" content.
Add an optional tail_count parameter: the last tail_count archived rows
are superseded byte-identical duplicates, stamped rewind-style
(active=0, compacted=0, hidden from recall) instead of compacted=1.
Callers:
- batch in-place compaction counts the compressor-tagged tail dicts
(_COMPACTION_TAIL_MARKER set by compress() on every carried-forward
message);
- micro-compaction splices [prefix, marker, suffix] - everything except
the single marker row is carried forward, so tail_count=len-1;
- proactive tool-result pruning rewrites content in place (not verbatim),
keeping the historical archive-everything behavior.
Fixes#86366
TUI server shutdown stamps ended_at/end_reason='tui_shutdown' on sessions
whose agent keeps running; every rotation then aborts at
publish_compression_child's liveness check forever (the #88197 wedge; the
amplification half was fixed by #88411).
Class fix: is_automatic_end_reason() in hermes_state_common owns the
"accidental infrastructure cleanup vs deliberate boundary" taxonomy.
publish_compression_child clears automatic stamps in its own transaction
and proceeds (parent re-closes with its TRUE boundary,
end_reason='compression'); the #88411 pre-flush guard no longer aborts on
stamps the publish can heal. Deliberate boundaries (compression,
session_reset, explicit close) still fail closed at both sites.
TEST REPIN (deliberate contract change):
test_ended_parent_aborts_before_the_prepublish_flush pinned
"tui_shutdown stamp => rotation aborts and parent must not grow" — the
abort it required IS the #88197 wedge. Repinned as two tests:
- test_automatic_stamp_no_longer_wedges_rotation: automatic stamp =>
rotation COMMITS (no abort loop, so no growth-by-abort is possible);
- test_deliberately_ended_parent_aborts_before_the_prepublish_flush:
session_reset (deliberate boundary) => still aborts BEFORE the #47202
flush, preserving #88411's no-growth contract where an abort remains
correct.
The class invariant "no aborted rotation grows the parent" holds
everywhere: automatic stamps no longer produce aborts, deliberate
boundaries still abort pre-flush.
Hardening on top of @Soju06's forwarding fix: v2 providers written against
the original docs example (def on_pre_compress(self, messages)) must not
TypeError when the host forwards require_checkpoint — inspect the signature
and fall back to the legacy call shape. Docs example updated to advertise
the keyword.
MemoryManager.on_pre_compress() detects checkpoint API v2 providers,
selects the normalized evidence list for them, and re-raises their
failures under require_checkpoint — but it never tells the provider
that a checkpoint is required: the call passes only the messages.
A v2 provider therefore runs in its default best-effort mode, swallows
durable-write failures, and returns normally; the host then treats the
checkpoint as succeeded and lossy compression proceeds. With
compression.checkpoint_required: true this silently defeats the
guarantee the option exists to provide.
Forward require_checkpoint only to providers advertising the requested
checkpoint API version. Legacy providers keep the strict one-argument
on_pre_compress(self, messages) contract, so bundled v1 providers
(honcho, mem0, supermemory, ...) are unaffected.
Regression tests cover required and best-effort signaling, legacy
signature compatibility, and required-mode failure propagation.
Idle and preflight compaction arrived as lifecycle status without the
"Compacting context" marker, so TUI never entered a compacting state.
Re-tag those lines and freeze the busy FaceTicker on "compacting" for
the whole pause instead of restoring "running…" after 4s.
Review folds from the formal gate battery:
- _SPLIT_FAILURE_COOLDOWN_SECONDS = 60 replaces the bare literal, with a
comment pinning WHY it is the timeout ladder's first rung (transient
lease/DB condition) rather than the 600s summary-provider cooldown.
- publish_compression_child docstring now states the compression_lock_holder
condition on the refresh guard.
- Dropped 2 of 3 extracted unit tests as duplicates of existing coverage in
test_compression_rotation_state.py / test_context_compressor.py; kept the
force-bypass test (only site pinning that behavior for split failures) and
the E2E test (now asserting the named constant).
The cooldown is recorded inside the split-failure except handler; a raw
call on a stub/partial compressor would replace the real split error with
an AttributeError. Mirror the try/except-debug convention of the adjacent
record_rejected_compaction call.
Two narrow repairs for #97948 symptom B (large-session rotation aborts with
'Compression lease lost before publication' / session_split_failed, then the
next turn re-runs the identical doomed compression):
1. publish_compression_child gains require_lease_refresh: the lease is
extended inside the same transaction as the expiry check (same conn, no
TOCTOU), giving a worker whose refresher thread died from transient DB
failures one final chance to keep its completed work.
2. A failed compression split now records a 60s failure cooldown, so the
next turn cannot immediately re-trigger the same compression.
Salvaged from #98137 (author: vsd2807). The timeout-reconciliation half of
that PR is NOT carried: it has a blocking review (runtime sid vs persisted
session_key, one-shot check cannot observe a 6-minute commit, no identity
projection) and needs a redesign.
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
ring entry removed from keyless_mcp, legacy backend set / credential
ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
nous_subscription surfaces. The tvly- redaction pattern stays --
legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
environment-variables, tools-reference, web-dashboard, provider
plugin dev guide.
Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
Test seams and plugin engines monkeypatch estimate_messages_tokens_rough with (messages)-only signatures; route callers only pass the charge_stale_thinking kwarg on the False path.
The preflight trigger charged reasoning/reasoning_content on every assistant message while the tail-budget walks charged newest-turn-only (#73624), so reasoning-heavy codex_responses sessions fired compaction forever while the walk protected everything (middle_window_tokens=0, no_progress every turn, each attempt a full aux summarization).
Wire truth: the codex_responses input builder never ships the text thinking keys (encrypted codex_reasoning_items carry the chain and were already charged unconditionally by both sides), so the trigger overcounted reality; echo-back chat-completions families (DeepSeek/Kimi/MiMo thinking mode) replay stored reasoning_content on every turn, so there the walk undercounted. New single wire-truth predicate message_sanitization.stale_thinking_reaches_wire() now drives BOTH sides: trigger estimates exclude stale thinking on non-echo routes; tail/prune walks charge it on echo routes.
Also: reasoning/reasoning_content double-count fixed in both estimators (wire ships at most one; +53% overcount vs provider prompt_tokens per issue comment), and the commit-layer no_progress path now arms the structural no-op backoff so an unchanged-transcript compaction cannot re-fire every turn (defense in depth; overlaps the #96775 re-entry class).
The bounded-grace join only applies where the overlap hazard lives: a
total-ceiling expiry over a still-streaming worker (#97488). The
idle-stall path keeps its prompt detachment so the stall-fallback retry
preserves the #76354 S3 latency contract (silence never approaches 2x
the idle budget); its late unwind stays safe behind the fence poison
and attempt-generation supersession.
compress_context() now publishes agent._compression_blocked_transient
(reason string) when an automatic pass no-ops because a timed guard —
summary-failure cooldown or structural backoff — is active, with a
clear skip log line. The overflow-recovery and preflight loops in
conversation_loop treat that signal like the #69870 lock-skip: refund
the attempt and end the turn as compression_deferred instead of
counting the no-op toward compression_exhausted, which auto-resets
(wipes) the session at the gateway. Fixes the false auto-reset where a
real context_length_exceeded arrived while the host-timeout cooldown
was still active. The permanent 'ineffective' breaker intentionally
does not set the signal so genuinely incompressible sessions can still
exhaust.
record_timeout_failure() now persists
'backoff:<failure_kind>:strategy=<tail_mode>' into the state.db
cooldown row (sessions.compression_failure_cooldown_until +
compression_failure_error), so a failed/stalled/cancelled attempt's
identity survives gateway restarts and the rebuilt compressor makes the
same skip decision via bind_session_state()/get_active_compression_
failure_cooldown(refresh=True). Host callers pass ceiling_exhausted /
stalled; the stall-interrupt path passes stall_interrupted. A
successful compression still clears the row.
A ceiling/idle-timeout host now joins its fence-cancelled worker for a
bounded grace before returning. A cooperative worker (which polls the
poison fence between provider phases) is reaped, proving quiescence, so
the durable lease releases normally. An uninterruptible worker is
orphaned behind the poison fence: its late result is discarded, and on
the total-ceiling path the holder-qualified lease stays retained until
it exits so no new attempt can overlap the unchanged session.
Supersession: a late candidate from an attempt whose compressor
generation was claimed by a newer attempt is discarded before the
commit boundary (failure_class=attempt_superseded), never committed
over newer state — covering fenceless callers the fence poison cannot
see.
An explicit /stop after the summary stream has already gone idle restored the original transcript but left no durable cooldown, so the next automatic turn re-entered the same stalled strategy. Record a stall-specific failure on that path only, merge with any longer live deadline, and keep an ordinary early /stop cooldown-neutral.
compress() returns marker-swept copies (_strip_persistence_markers, #57491);
the in-place branch committed them via archive_and_compact() but never
stamped the persistence marker, so the next _persist_session ->
_flush_messages_to_session_db_unlocked walk re-INSERTed the whole
post-compaction transcript (live set regrew ~58K -> ~512K tokens).
Centralize the post-commit contract in a shared helper,
stamp_db_persisted_markers(), used by all three archive_and_compact
callers: the in-place batch commit (previously missing), the
micro-compaction sync, and the proactive tool-result prune.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.
- The detailed session log is folded into the single summary request: the
lean prompt template gains a '## Detailed Session Log (oldest first)'
section carrying the digest prompt's HARD RULES (identifiers verbatim,
dense bullets, transcript-is-data). Output guidance grows by
_LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
summary budget — the old worst case (28 x 1,400 digest tokens) was spread
across many requests and mostly re-covered tool noise; a single dense
4K-token log inside one response preserves the load-bearing record while
staying well inside one aux response (the summary call still sends no hard
max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
whole region (_sample_summary_input: 8 proportionally spaced slices,
oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
anchored to the newest end) instead of head+tail truncated, so session-log
coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
_LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
_LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
attempt_summary_route_kwargs — no remaining callers; the single-use
summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
section lands in the summary; oversized regions sampled with elision
markers, never a second request; anchor index + recovery footer present).
Sabotage-verified: restoring a second call_llm makes the call-count test
fail. Docs and the compaction eval wording updated to stop claiming
per-chunk calls.
Fixes#96603.
Keep model-switch callers compatible with result objects created before runtime_capabilities was added, and do not roll back minimal agents that lack optional LM Studio helpers. Preserve rollback for real helper failures.
Use a distinct runtime_capabilities field on agents, preserve compatibility with earlier snapshots, and resolve the canonical direct OpenAI endpoint when a cross-provider switch omits base_url. Keep ambiguous proxy routes fail-closed.
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
gpt-5.6 on the Codex backend answers a large turn with a server-side
`compaction` checkpoint and no message. The checkpoint rides the
`codex_reasoning_items` sidecar, so the interim assistant message looks
"replayable" and `interim_replayable` suppresses the continuation nudge.
But replayable is not the same as different. A checkpoint carries no
answer and no new instruction, and a replayed checkpoint makes
`prune_pre_checkpoint_items` drop every pre-checkpoint item. Measured on
a real 262-message session: the wire collapses from 489 items to 12 —
all 186 `function_call` / `function_call_output` pairs deleted — and
ends on an empty assistant turn. The model has nothing to answer, so it
returns another empty response; the next attempt sends the same bytes
(the provider's prefix cache reports 99-100% on the repeats) and returns
the same nothing. Three attempts later the turn dies with "Codex
response remained incomplete after 3 continuation attempts" and the
whole turn's work is lost.
Keep the first continuation bare — the model often just needs another
turn, and nudging immediately would cut multi-phase work short. Once
that bare retry has also come back incomplete, it is proven not to work
for this turn, so every remaining attempt carries the nudge.
Preserve valid normalized input_image user messages across native-compaction checkpoints at bounded one-token retention cost. Keep text extraction text-only, reject malformed or unknown multipart placeholders, and prove the production adapter path without claiming unsupported input_file behavior.
Republish the identical source tree after an unrelated nondeterministic focus-redraw test failure; this commit contains no source delta from the previously verified object.
Refs #90976 and #91477.
The sibling-site widening replaced estimate_request_tokens_rough with estimate_messages_tokens_rough as the generic fallback feeding the route-aware wrapper, dropping the 20-30K token tool-schema envelope (#14695 class) and shifting the pinned mid-turn retry comparison. Restore the tools-inclusive figure as the fallback.
Follow-up to the mid-turn pre-API guard fix (#96995 / #97602): sweep the
remaining call sites that derive automatic compression pressure from a
generic estimate over the assembled durable history, which on a compacted
native-Codex session overstates the wire payload by orders of magnitude.
- agent/turn_context.py idle-triggered compaction: use
_preflight_request_tokens (anchor -> native pruned -> generic) instead
of the raw generic request estimate, so resuming a compacted codex
session after an idle gap does not fire a compaction the next request
never needed.
- agent/turn_context.py uncompressed-session overflow-warn RE-ARM: match
the warn site's route-aware figure so the dedup re-arms correctly on
native sessions.
- agent/conversation_loop.py post-response should_compress fallback
(last_prompt_tokens==0, i.e. no provider usage after a disconnect or
gateway restart — the unanchored case in #97602's repro): route through
_midturn_request_pressure_tokens instead of the generic figure.
Left alone deliberately: provider-proven overflow recovery paths (413 /
context-length errors — the provider already proved the request does not
fit, figures there only arm recovery and score progress), compression
progress before/after pairs (relative deltas on the same scale), manual
/compress display estimates (gateway/CLI/ACP feedback, not automatic
triggers), MoA advisor budget trimming (not a codex-native wire payload),
and context_compressor internals (measure local durable-history shrink).
The #96155 fix (#96644) made the turn-prologue preflight estimate the
checkpoint-pruned native Responses payload, but the independent mid-turn
pre-API pressure guard in conversation_loop still estimated the full
assembled durable history. On a compacted native-Codex session the
generic figure overstates the wire by orders of magnitude (the issue's
deterministic probe: 1,037,241 generic vs 6,036 pruned, 171x), so the
guard false-tripped a 600-second local compression the main request
never needed — the live sequence shows the actual request then fit at
164k input tokens against a 765k threshold (#96995).
Extract the guard's pressure figure into _midturn_request_pressure_tokens
and mirror the turn-prologue: when native Responses compaction is proven
eligible, use estimate_native_responses_preflight_tokens (system prompt
and tools included, checkpoint-pruned); otherwise keep the generic
message+tools figure. Passing the assembled api_messages alongside
effective_system counts the system prompt exactly once — the estimator's
converter skips system-role rows and adds the prompt separately.
total_chars (verbose log proxy) and the non-codex paths are unchanged.
Fixes#96995
* refactor(skills): shipped-set slim — 15 skills to optional, github six-way merge, pdf absorbs OCR+nano-pdf, channel-gated teams pipeline
Maintainer-directed shipped-skills curation (skills index 1,900 -> ~1,400
tok/call on desktop; every session pays the index, so this is a per-call
diet on all installs):
- optional-skills moves (installable via skills hub, history preserved):
creative comfyui/ascii-art/excalidraw/pretext/sketch/touchdesigner-mcp;
ALL of mlops (huggingface-hub, llama-cpp, serving-llms-vllm,
weights-and-biases, evaluating-llms-harness — subcategory structure
kept); research-paper-writing (55 supporting files, 17.3K-tok load);
openhue; blogwatcher (first taught the cronjob monitor-field watch
pattern + web_extract instead of pre-cron manual workflows)
- DELETED session-librarian (Aug-12 'inspired by Perplexity Computer'
port, never maintainer-intended; session_search covers discovery)
- github: six skills (auth, issues, pr-workflow, issue-to-pr,
code-review, repo-management) merged into ONE software-development/
github skill — routing body + complete per-workflow references;
benbarclay authorship credited; codebase-inspection rides along;
discipline pins from test_github_issue_to_pr_skill.py preserved
against the reference body in the new test_github_skill.py
- pdf absorbs ocr-and-documents + nano-pdf as references/ + scripts
(extract_pymupdf, extract_marker converted to the argparse house
standard its contract test enforces)
- NEW session_platforms frontmatter gate (metadata.hermes): hides a
skill from the index on gateway channels it is not for; fail-open on
unknown platform; teams-meeting-pipeline gated to [teams, cron]
- blocked-page-recovery: research -> new web category; trigger-first
description ('Use when a fetch fails: 403/429, paywall, WAF, bot
wall.') so the model actually reaches for it on blocked fetches
- docs regenerated via generate-skill-docs.py (195 pages); related_skills
swept repo-wide; tests: 1672 passed (2 openclaw failures pre-existing
on clean main, Windows-local)
* chore: ignore .skills_prompt_snapshot.json (local index cache, accidentally committed)
The supporting-files block emitted every file twice per line
(relative -> absolute), duplicating an identical directory prefix
hundreds of times on reference-heavy skills. The absolute base is
already stated once in the [Skill directory: ...] header and the
footer example, so each line now carries only the relative path.
On hermes-agent-dev (462 supporting files) the activation message
drops from 42,145 to 29,042 o200k tokens (-13,103, -31%) — paid on
every session that preloads or invokes the skill.
* feat(compaction): always rebuild the system prompt at the commit boundary — keep-prompt now gated on byte equality of the LIVE builder output; plugin sections re-render with fail-open to last good bytes
* feat(clock): 'Conversation started' resolves through the session-lineage ROOT — a compacted/rotated session keeps its original birth date (Bot Mode forever-chats know when they were first born)
* test: retire old-contract pins — plugin sections re-render at invalidate (freeze stays restore-only), commit boundary always runs the live builder, byte-equal keep preserves object identity
The bundled plan skill's auto-generated slash command fell off the capped
Telegram/Discord command menus for most installs (skills are the only tier
trimmed at the platform caps, alphabetically — 'plan' sat past the cutoff at
index 57 of 82 bundled skills). Converting it to a first-class CommandDef
gives it a guaranteed core-tier menu slot on every platform.
- agent/plan_prompt.py: build_plan_prompt() — plan-mode rules + authoring
craft distilled from the retired skill; prompt-injection pattern like
/learn and /init (no engine, no model-tool footprint, cache-safe).
- CLI: _handle_plan_command mixin handler (pending-input injection).
- Gateway: /plan branch rewrites event.text and falls through (role
alternation preserved).
- TUI: command.dispatch branch ('plan' was already in
_PENDING_INPUT_COMMANDS).
- Removed skills/software-development/plan/ + docs pages (EN + zh-Hans),
catalog rows, sidebar entry, related_skills references.
- PROTECTED_BUILTIN_SKILLS is now empty (mechanism kept); dependent
curator/usage tests moved to monkeypatched sentinels.
Salvages #67292 by @webtecnica (credit: first /plan command submission,
issue #67264); reworked from inline planning prompt to the prompt-injection
pattern with workspace-saved plans. Closes#67264, closes#36821 (empty
/plan infers task from conversation context).
Follow-up hardening on the two cherry-picked contributor commits:
- try_activate_fallback: replace the blanket request_overrides.pop('extra_body')
with KEY-SCOPED removal — only keys the OLD provider's custom_providers
entry contributed (value unchanged since the init-time merge) are dropped.
Caller/profile-provided extra_body keys survive the swap, matching the
caller-over-provider precedence in agent_init._merge_custom_provider_extra_body.
The fallback provider's own extra_body is then merged back in.
- switch_model: the live _primary_runtime snapshot it rebuilds now carries
request_overrides, so a post-switch transport recovery or fallback restore
reinstates the switched-to identity's overrides instead of dropping them.
- Tests: activation-level stale-key removal + caller-override preservation
(test_provider_fallback.py), switch-then-recover / switch-then-restore
(test_primary_runtime_restore.py).
Cache-safety: none of these paths mutate past context or rebuild the system
prompt — only outbound request kwargs change.
Fixes#75091
`try_activate_fallback()` re-resolved `reasoning_config` for the new
fallback provider (fix for #21256), but never re-resolved `extra_body`.
The primary provider's `extra_body` (e.g. `reasoning_effort: "none"`)
rode along onto the fallback provider, which is a different API that
may reject those fields.
Example: primary has `extra_body: {reasoning_effort: "none"}`, fallback
is OpenRouter. After failover, every request to OpenRouter carries both
the stray top-level `reasoning_effort` AND the nested `reasoning` object,
and OpenRouter rejects the pair:
HTTP 400: "reasoning_effort" and "reasoning.effort" are both provided
The fallback is dead precisely when it is needed.
Fix: after swapping provider/model/base_url, clear the primary's
extra_body from request_overrides, then re-resolve from the fallback
provider's config using the existing _merge_custom_provider_extra_body
helper. Same pattern as the reasoning_config re-resolution above.
Addresses the hermes-sweeper review on #53765. The in-place /model switch
helper (_apply_switched_provider_request_overrides) derived a custom
provider's extra_body by provider *name* only, while build-time matching in
agent_init._merge_custom_provider_extra_body matches by provider key, base_url,
AND model. So a different model selected at the same named endpoint could
inherit an extra_body configured for another model.
Reuse the shared agent_init._custom_provider_extra_body_for_agent matcher
(provider key + base_url + model), sourcing custom_providers from the
init-time agent._custom_providers cache (fresh-load fallback if absent). A
stale extra_body is always cleared when no entry matches; non-provider
overrides (service_tier / speed from /fast) are preserved.
Tests: add nonmatching-model and endpoint-mismatch regressions; update the
existing switch tests onto the model/base_url-aware matcher.
Third in the series. The gateway rebuild path (previous two commits)
carries a custom provider's `request_overrides` (`extra_body`, e.g.
`chat_template_kwargs`) into the agent, but the *in-place* live switch used
by the TUI dashboard and the CLI — `agent.switch_model()` ->
`agent_runtime_helpers.switch_model()` — swapped
model/provider/base_url/api_key without ever updating `request_overrides`.
So a `/model` switch to a thinking-enabled custom provider in the TUI/CLI
kept the previous provider's `extra_body`.
`switch_model()` now re-derives the switched-to provider's
`request_overrides` (via `_get_named_custom_provider`) and applies it in
place, preserving non-provider overrides (`service_tier`/`speed` from
`/fast`). Logic factored into `_apply_switched_provider_request_overrides`
for testability.
Adds tests/agent/test_switch_model_request_overrides.py.
Profiles can now grant narrowly scoped tools to the background review runtime whitelist while unrelated tools remain denied. Document the configuration and cover it with a real-config regression test.
Agent: codex
- Add rolling status bar metrics:
- cache hit ratio (◈) delta since model/compression reset
(hit = cache_read / prompt, verified against live logs)
- avg latency (◷) and throughput (↑ t/s) over last 10 API calls
(deque in agent, displayed in wide bar only)
- Add display.tui_statusbar_fields list to filter segments:
model, ctx, ctx_bar, cache_hit, latency, tps, compressions,
bg_tasks, bg_processes, bg_subagents, goal, duration, prompt,
idle, focus, yolo, stash, battery, title
Missing/null -> all enabled (backward compat). Unknown keys ignored.
Title gated via right-align; stash/battery also gated.
- Wide bar (≥76 cols) respects fields, narrow/medium filtered,
overflow trim preserved. Battery also respects display.battery.
No private data; mock data in tests.
Test: pytest tests/cli/test_cli_status_bar.py etc. 68 passed,
check-windows-footguns clean.
The consumed-but-uncommitted rotation verdict was process-local
(_SPENT_ROTATION_FINGERPRINTS), while the credential it protects is
explicitly cross-process: ~/.claude/.credentials.json is shared by every
Hermes profile and process. A fresh interpreter could lease the stale
access token or re-POST the already-spent single-use refresh token and
burn the credential family into invalid_grant.
- Persist non-secret one-way fingerprints to a sidecar registry next to
the shared singleton source (claude_code / hermes_pkce), written under
the same path-keyed cross-process lock that serializes refreshes.
- Consult the sidecar in the pool resolver, the pool refresh path, and
the direct claude_code resolver/refresh before leasing or POSTing.
- Two-process regression: A rotates and loses the commit; B (fresh
interpreter, empty local registry) must neither lease the stale pair
nor POST the spent refresh token. Plus a no-verdict control.
Closes the remaining P1 from the exact-head review of f228439b on
PR #87891.