Maintainer ruling: no new slash command for this. `display.vim_mode: true`
in config.yaml enables vi keybindings in the composer at startup; the
NORMAL/INSERT/REPLACE status-bar label stays. Removes the CommandDef, the
handler, its dispatch-table entry and slash-command docs; documents the key
under Display Settings.
MiniMax Code CLI 0.3.1 added a status-line segment showing the git branch
for the current workspace. Hermes' status bar had no repo-awareness field.
Adds `git_branch` to display.status_bar.fields (opt-in only — the default
set never probes the filesystem). Reads .git/HEAD directly with a 5s
per-directory TTL cache (no subprocess per repaint); follows gitdir:
pointer files so worktrees/submodules resolve their private HEAD; a
detached HEAD renders the abbreviated commit.
Inspired by MiniMax Code CLI 0.3.1 changelog (agent.minimax.io/docs/changelog).
The in-process last-known-good (the codex#31188 port) only protects a
long-running gateway. A CLI restart or `hermes config get` against broken
YAML fell through to DEFAULT_CONFIG and silently dropped every override,
including approvals.deny (#102945).
Every successful parse now leaves a `good` copy in backups/config/ through
the existing bounded, byte-deduped backup_config(); the fallback reads the
newest one, runs it through the normal canonicalize/expand/managed-overlay
pipeline, and says so on stderr. The broken config.yaml is never modified.
The backup is the raw file, so ${VAR} templates stay templates on disk.
Redo of #61796 (which added a config.validated.yaml sibling and
re-validated the whole config on every load) on top of the backups/config/
ruling from bf53ff00a7.
The "one gateway for all profiles" section had drifted from the runtime. Each
claim was re-verified at its defining symbol on current main and rewritten to
the behaviour users will actually see:
- named-profile guard: only `gateway run` refuses (exit 78 / EX_CONFIG,
systemd RestartPreventExitStatus; launchd KeepAlive still retries);
`start`/`install` do not refuse themselves and the message lands in the
service log; `--force` is a `run`-only flag
(hermes_cli/gateway.py::_guard_named_profile_under_multiplexer, _cmd_start)
- same (platform, token) in two profiles: the duplicate adapter is parked as
fatal/duplicate_credential and the gateway keeps running — it was described
as a fail-fast startup error (gateway/run_adapters.py::_refuse_duplicate_claim)
- status surfaces: one gateway_state.json under the default home with
`<profile>:<platform>` entries + served_profiles; nothing is written under a
secondary home (the page claimed a per-profile runtime_status.json), and
`hermes status` does not list served profiles — `gateway list`,
`-p X gateway status` and /api/status do (gateway/status.py::write_runtime_status,
hermes_cli/status.py::_render_gateway, hermes_cli/gateway.py::_cmd_status)
- API_SERVER_KEY in a secondary .env auto-enables api_server and trips the
port-binding skip; document the `enabled: false` pin
(gateway/config_env.py::_api_server / _enable_from_env)
- allowlist: it is a start-time snapshot, and the Desktop backend's cron
ticker enumerates every local profile regardless of it
(hermes_cli/web_server.py::_start_desktop_cron_ticker)
- routed-profile cron via the shared bot: only when the profile has no live
adapter of its own, and a route carrying guild_id never matches a cron
target because delivery matches on chat_id/thread_id only
(cron/scheduler_provider.py::tick_adapters_for,
cron/scheduler_preflight.py::SharedRouteAdapters.get)
- add a "What is isolated per profile" table (credentials, authorization,
endpoints, media denylist, MCP child env, outbound egress, session
namespace, logs, terminal) describing behaviour, not PR numbers
- configuration.md: the ${VAR} scoping paragraph now says where it applies
and links to the table
The Python passive check (`banner.check_for_updates`, used by the CLI
banner, `hermes --tui`, every `tui_gateway` spawn and the dashboard's
/api/hermes/update/check) also ran `git fetch origin main` on every cache
miss, and never cached an inconclusive result so a flaky line retried on
every start. Same GitHub complaint, same fix:
- remote tip via GET /repos/{slug}/commits/main (vnd.github.sha), local tip
via rev-parse, exact count + changelog via the compare API when they
differ. HTTPS `ls-remote` remains only as the fallback when the API is
unreachable or the origin isn't on GitHub.
- cache TTL 6h -> 24h, failures cached 1h; the cache is keyed on HEAD so
`hermes update` invalidates it immediately.
- the dashboard's "what's changed" list comes from the memoized compare
payload (`upstream_commits_behind`) instead of `git log HEAD..origin/main`,
which was stale without a fetch.
Tests rewritten to the new contract: passive checks must not run
`git fetch`/`ls-remote` for a GitHub origin; the daily cache invalidates
when HEAD moves and re-asks after the failure window.
Four writers each dropped their own uniquely-named copy of config.yaml next to
the real file and none of them ever deleted anything: hermes setup
(config.yaml.bak.YYYYMMDD_HHMMSS, one per run even with no change), the
corrupt-YAML snapshot (config.yaml.corrupt.<ts>.bak), hermes migrate xai
(config.yaml.bak-pre-migrate-xai-<ts>) and the Docker boot migration
(config.yaml.bak-<ts>, .env.bak-<ts>). A home dir accumulated a dozen variants
with no way to tell which mattered.
hermes_cli/config_backups.py::backup_config is now the single writer:
backups/config/config.yaml.<reason>.<YYYYMMDD-HHMMSS>, skipped when the newest
copy for that reason is byte-identical, rotated to the newest five per reason.
backups/ is already excluded from full backups so nothing nests. Legacy
siblings written by the old schemes are moved into the dir on first use;
hand-named copies (config.yaml.bak-my-note) are left alone.
Live: three `hermes setup --non-interactive` runs against an unchanged config
went from three .bak files in HERMES_HOME to one pre-setup copy under
backups/config/; repeated loads of broken YAML produce one corrupt copy
instead of one per process (deduped by content).
The ledger in verification_evidence.db exists only to feed the verify-on-stop
guard, but the recorder kept running on every foreground terminal command and
every file edit after #53552 turned the guard off by default. Users who never
opted in still accumulated a multi-MB database (7 MB / 4.6k rows on one install).
Every ledger entry point (record_terminal_result, record_verify_run,
mark_workspace_edited, verification_status) now checks verify_on_stop_enabled()
first and returns without opening or creating the database when the guard is
off. verification_status reports {"status": "disabled"} in that case; no client
consumes the verification.status RPC yet, so nothing downstream changes.
Existing ledger tests pin HERMES_VERIFY_ON_STOP=1 since they exercise the ledger
itself; the new test proves the off path never creates the file (red on base).
Adopted from PR #80421 with the author's explicit go-ahead on #80450
('Please proceed!'): config_defaults entry, cli-config.yaml.example
block, and user-guide docs for the delegation-scoped fallback chain.
Co-authored-by: Andrex Ibiza, MBA <84248988+andrexibiza@users.noreply.github.com>
`_on_tool_progress` bailed on the tool-progress gate before dispatching
`subagent.*`, so a Desktop/TUI user who hid tool-call chrome also lost the
subagent rows in the status stack and spawn tree. Subagent lifecycle is
application state (like `todo.updated`, clarify and MCP consent cards,
which already bypass the gate); the gate now applies only to the optional
progress chrome (reasoning previews, MoA rows, tool.generating).
Give dispatcher-owned workers a tool-capable reporting opportunity before the
hard iteration cap, without accepting arbitrary diffs or weakening failure
counting. Add opt-in per-turn iteration checkpoints for ordinary agents.
Persist checkpoint text with the fresh tool result, never rewrite cached rows.
Salvages the opt-in ratio and per-turn reset implementation from #104683;
credits the earlier default-off signpost proposal in #92438.
Local fixture wire A/B: Kanban ready/1 failure -> done/0; deliberately stuck
workers still reach blocked/2 after two runs. Default-off control unchanged.
Targeted and affected-directory suites queued behind campaign test lock.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: C. Michael Gibbs <252231331+MikeGibbsOnyx@users.noreply.github.com>
The real closed-WebSocket resident variant still wedges after the missing-timer repair. Re-enter existing transport cleanup before rearming orphan timers, preserving viewer transfer and the reconnect/delegation fence rather than deleting registry rows from a stale snapshot. Extend the same invariant to both dead-transport shapes. Live class investigation informed by #104710; no direct-vouch reclaim machinery imported.
Salvage #104704 (e9423d2d0bbe3e795c5eaccb86a913f1d95444ba, 5fa2b98fc833fc4e1ee7f1aaa7eb45cdd3cd7cde). Reuse guarded orphan teardown instead of deleting ownership fences. Real two-backend WebSocket probe reproduces the missing-timer wedge on base and proves reconnect/delegation protection and recovery. Add reusable probe and user documentation.
Use request-local stream silence for the waiting notice, preserving quiet
activity heartbeats and all existing watchdog policies. Distinguish a stream
that stopped from a request with no response, and clear this request's notice
on the next poll when events resume. Existing fresh first-event retry phases
also reset the display; recovery deadlines explicitly use total call elapsed.
Add two invariant tests (eight cases), proven red on main, plus EN/ZH docs.
Local SDK SSE through classic CLI callbacks in a PTY verifies active reasoning,
true silence, and an already-visible warning clearing on resumed reasoning.
Related: #92657 addresses repeated waiting notices; its phase deduplication
still labels active streams as no response and is not incorporated here.
Slimmed after review: the 200K default is dropped. Children compact at the same
0.50 x window ratio trigger as their parent (500K on a 1M model). Reasons:
- the run this came from happened at 0.85 (850K); main was already at 0.50, so
the real delta against main was 500K -> 200K, not 850K -> 200K;
- a replay of the run's 22,489 logged calls (evals/postmortem, cap sweep) put
200K-400K caps within 5% of each other in cost once cache prefixes are intact,
because the write price dominates and the cap only trims read volume;
- every compaction is a chance to lose detail, and the accuracy side was never
measured; at 500K a 1M child compacts roughly never.
What stays: the reviewer's finding that the value was coerced, not validated
(YAML true -> int 1 -> a one-token trigger; "200k" -> silently off). Values are
validated: int >= 16000 enables the cap, 0/false/null/unset = off, anything else
is warned and ignored. Docs and config comment restated accordingly.
A delegate_task child inherits the compression threshold as a RATIO of the
model window. On a 1M-window model at the run's configured 0.85 that is an
850K-token trigger: in the 1,393-agent refactor run 1,373 of 1,375 children
never compressed once, 62% of all API calls carried >150K of context, and the
calls above 200K carried ~$10.9k of the $19.3k bill (58% of it cache WRITES,
i.e. re-sending a 300-800K prefix on every call). A sawtooth replay of the
logged calls with a 200K cap / 65K floor cuts context spend by ~49% (~$7k).
Children are brief-driven and disposable; they re-read their brief and the
files they touch, so a large window buys them little. New
delegation.compression_threshold_tokens (default 200000) is applied to the
child's ContextCompressor right after construction as the lower of it and
any global compression.threshold_tokens; the parent's own trigger is
untouched. 0 disables the subagent-specific cap. The compressor applies
threshold_tokens_cap on first window resolution, so this is byte-equivalent
to the user having set compression.threshold_tokens for the child.
Live through the real spawn path (_build_child_agent, real imports, temp
HERMES_HOME, 1M-window model): main child trigger 500,000 / branch 200,000;
parent 500,000 on both.
Tests (3): default caps a 1M child at 200K; the cap is the lower of the
delegation and global values and never raises a small-window child's
trigger; 0 disables and an already-resolved trigger is re-clamped.
Docs: delegation.md, configuration.md.
Adds plugins/web/perplexity — a keyed-only WebSearchProvider over httpx:
- search: POST https://api.perplexity.ai/search (documented Search API),
search_context_size=low so `snippet` stays description-sized;
results[].snippet -> description, max_results capped at the API's 20.
- extract: POST /sdk/content/snippets — the query-relevant page-excerpt
route behind `pplx content snippets` (the CLI's `content fetch` is
deprecated upstream). web_extract has no query, so the URLs' path words
serve as the relevance query; per-URL `error` entries survive a 200.
- Wired into the same touchpoints as the other keyed vendors: legacy
backend set + credential ladder + availability probe (web_tools),
registry preference walk, OPTIONAL_ENV_VARS, `hermes config`/status/
dump key lists, nous_subscription direct-credential detection, setup
summary, test conftests, docs.
Not a keyless-ring member (Perplexity has no anonymous tier). Related
closed PRs #9192 / #23981 / #45225 predate the plugin ABC.
On agent.tool_use_enforcement/execution_guidance "auto", muse-spark-* was in
neither model tuple, so it received only the universal finish-the-job block,
answered in prose with 0 tool calls, and the turn closed on finish_reason=stop.
Add "muse" to both tuples; Claude and every other family are unchanged.
Co-authored-by: Edder Talmor <talmoredder@gmail.com>
- tests/hermes_cli/test_terminal_notify.py: OSC 9 body emitted+sanitized
only when bell flag on; Warp payload only under a supported Warp build.
- configuration.md display section: document the notification behavior
of bell_on_prompt / bell_on_complete.
- contributors/emails: glitchbunny0 (#58957), harshmoney123 (#100805).
- display.bell_on_approval (default false): same BEL mechanism as
bell_on_complete, rings when a dangerous-command approval prompt
opens (_approval_callback / approval.request event). Complements
bell_on_clarify from the previous commit.
- fix(ui-tui): eslint curly error in useConfigSync.applyDisplay
(if without braces) that failed the CI JS & TS checks job.
Same BEL mechanism as display.bell_on_complete (\a / \x07), gated by
display.bell_on_clarify (default false). CLI rings in _clarify_callback
and _clarify_callback_batch before _paint_now(); TUI rings on
clarify.request when bellOnClarify && stdout.isTTY. Docs in
cli-config.yaml.example and website/docs/user-guide/configuration.md.
Adds two bounded fast modes on top of the static /fast toggle, default OFF:
- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
window; requests inside it carry the provider fast param, later tool-loop
requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
user/assistant/tool history).
agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.
resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.
Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.
Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes#64785, #74730.
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
Before turning hard stops on for unattended platforms, make sure they cannot
cut off normal work:
- Edit -> re-run is progress. A successful mutating call (write_file/patch,
a green terminal/execute_code, browser actions, job/message/cron/memory/
skill mutations) marks progress for every failing signature still being
counted this turn; the next identical retry restarts its streak instead
of accumulating toward exact_failure_block_after. A pure replay never
mutates anything between attempts, so it is still blocked at 5.
- Distinct red commands are diagnosis. For FAILURE_TOLERANT_TOOL_NAMES
(terminal, execute_code, process pollers, browser_navigate, web_extract)
same_tool_failure_halt_after warns but never halts.
- subagent and api_server keep the warn-only default: both are supervised
task loops with a live parent/client and do real edit -> re-run work.
Live A/B (real AIAgent platform=telegram, real patch+terminal, 8 rounds of
patch -> red check -> patch ...):
unmitigated branch: HALTED at round 6 (repeated_exact_failure_block)
this commit: COMPLETED all 8 rounds, final answer delivered
Loop shapes still stopped: identical failing read_file 8 calls,
identical successful terminal 5 calls (vs 602 on main).
Six new tests pin these flows; all fail on the unmitigated version.
Widen the salvaged #49189 hard-stop default so it covers the loop shape in
the #100849 debug bundle and #89069: a model replaying the same SUCCESSFUL
call (terminal, skill_view, memory) with a byte-identical result. The
per-turn idempotent_no_progress block only tracks IDEMPOTENT_TOOL_NAMES, so
those loops ran until the iteration budget (600 calls, ~40 min) with only a
notice appended.
- agent/tool_guardrails.py: observe_call's tool-agnostic consecutive-identical
streak raises a halt (identical_call_streak_halt) at
hard_stop_after.idempotent_no_progress when hard stops are active. Pollers
stay exempt; a changed result resets the streak; warning-only sessions are
unchanged.
- run_agent.py: surface that halt from _append_guardrail_observation like
every other guardrail halt (appends guidance, ends the turn).
- hermes_cli/config_defaults.py: declare non_interactive_hard_stop_enabled.
- docs: configuration.md describes the streak hard-stop.
- tests: streak halts terminal under hard_stop; never under soft mode,
for pollers, or when results change.
Live A/B (real AIAgent platform=telegram, mocked client replaying one call):
identical failing read_file main: 602 API calls, budget exhausted
branch: 8 calls, repeated_exact_failure_block
identical successful terminal main: 602 API calls, budget exhausted
branch: 5 calls, identical_call_streak_halt
Follow-up to @zoser69's #78111 cherry-pick:
- lift the redirect into _redirect_platform_display_key() and apply it
BEFORE _validate_config_key / type coercion, so the unknown-key hint and
the string-vs-bool coercion both see the canonical path
- widen to the sibling surfaces: config get resolves the canonical key
(previously echoed the dead top-level value — the misleading half of the
report) and config unset removes the canonical leaf
- regression tests: get mirrors gateway resolve_display_setting, unset
removes the redirected leaf, note printed, helper touches ONLY
OVERRIDEABLE_KEYS (connection keys / 4-segment / already-canonical
paths untouched)
- docs: configuration.md per-platform section names the canonical CLI
path and the accepted shorthand
The 10s hygiene_max_turn_hold_seconds budget (#92318) releases the arriving
user turn while the summary model is still streaming. For thinking summary
models (DeepSeek-V4-Flash etc.) whose reasoning prefix alone exceeds 10s,
the abandonment path ALWAYS cancelled the commit fence — 100% of the summary
attempt (including the full thinking prefix) was discarded on every turn,
permanently disabling auto-compression while paying the summary model 10s
of thinking per turn, and the flat 60s retry-after then blocked the
agent-side preflight from a fresh chance.
Structural fix (maintainer-chosen direction in #97963): decouple the turn
from the compression instead of holding the turn longer or making the hold
progress-aware (which would reintroduce the #90845 frozen-turn bug):
- CompressionCommitFence gains mark_commit_watermark_fenced() /
commit_watermark_fenced; compress_context marks the fence right after
capturing get_active_message_watermark() under the durable compression
lock (#75316/#87484) — the property that makes a LATE commit safe: rows
appended after compression start survive both commit paths verbatim as
cloned concurrent tail (archive_and_compact watermark= and
publish_compression_child watermark/watermark_ceiling).
- gateway hygiene turn-hold handler: when the fence is watermark-fenced,
the detached worker (already kept alive via
_defer_agent_cleanup_until_future_done) KEEPS its commit admission; the
user's turn proceeds on the uncompressed transcript at the same 10s
budget, and the summary is adopted at the worker's own watermark-fenced
commit boundary. Unfenced workers are cancelled exactly as before —
never worse than the status quo.
- No retry-after is armed while the kept-admission attempt runs (it would
block preflight adoption via the same-session cooldown); re-attempt
spacing is covered by the durable compression lock
(_session_has_compression_in_flight). If the worker ends WITHOUT
committing, a done-callback restores the flat non-escalating 60s
retry-after; a successful adoption resets the hygiene failure streak.
The streak never advances for a deferral either way.
- Docs: configuration.md hygiene_max_turn_hold_seconds one-liner updated
to describe deferred adoption and the thinking-model case;
config_defaults.py comment updated. Knob stays config.yaml-only.
Invariants preserved:
- 10s user-latency cap stays hard (#90845/#92318):
test_session_hygiene_turn_hold_budget_abandons_streaming_wait passes
UNMODIFIED (its worker is not watermark-fenced, so it pins the cancel
path through the public surface).
- Stale-clobber impossible: adoption only rides commits bounded by the
start watermark; the fence still gates admission and unfenced/late
results are discarded.
New regression tests (tests/gateway/test_session_hygiene_turnhold_adoption.py):
- watermark-fenced worker keeps admission, late summary is committed,
turn still released at the budget, no cooldown while running,
streak reset on adoption;
- kept-admission worker that ends without committing restores the flat
turn-hold retry-after (<=120s, names turn-hold, streak untouched);
- unfenced worker still cancelled and discarded (status quo).
Sabotage-verified: disabling the keep-admission branch fails the two new
adoption tests and leaves the unfenced-cancel test green.
Fixes#97963
should_use_direct_api_call() contexts (gateway cron turns #62151, delegate_task
children #60203) were short-circuited onto the NON-streaming wire because the
interrupt worker wedges inside their nested thread pools. That dropped every
liveness property streaming provides: edge proxies kill the silent POST
(z.ai HTTP 524 — three retries later the child dies as "max_iterations"), and
the non-stream stale watchdog cannot tell a reasoning model's thinking phase
from a hung provider, so children die at exactly stale_timeout (#100260).
Keep those contexts on interruptible_streaming_api_call. The request now runs
INLINE on the conversation thread (no worker → the deadlock class stays
closed) while the existing poll loop — 30s heartbeat, stale-stream detector,
cross-thread interrupt abort — moves onto a monitor thread that only ever
aborts sockets, never dispatches (same shape as direct_api_call's watchdog
timer). Interactive sessions are unchanged: worker + poll loop as before.
should_use_direct_api_call() itself is untouched; only what it routes to.
Live A/B (real SSE server, real AIAgent.run_conversation):
before: subagent/cron wire stream=None, request on conversation thread
after: subagent/cron wire stream=True, request on conversation thread
cli unchanged (stream=True, spawned worker)
inline stale detector kills a one-chunk-then-silence stream at budget;
AIAgent.interrupt() from another thread unwinds the inline stream in 0.6s.
Co-authored-by: Expri-commits <184641533+Expri-commits@users.noreply.github.com>
The 20s ws-orphan grace (14b50f5edd) interrupts a RUNNING turn whenever
the client is absent past the grace window — killing healthy long turns
on deliberate client absence (desktop closed, PC asleep, mobile
backgrounded, Electron tab-switch throttling, desktop update/relaunch).
The reaper now interrupts a detached running turn ONLY when BOTH the
client is absent past the grace AND the turn's activity clock is stale
(seconds_since_activity >= dashboard.ws_orphan_activity_stale_s,
default 600s — matching agent.turn_liveness.timeout_s semantics from
PR #99758). A detached-but-actively-producing turn keeps running to
completion (the sentinel transport already buffers detached emits);
a detached AND activity-stale turn is interrupted/reaped as today.
Non-running orphaned sessions keep current behavior. Reuses the
existing AIAgent.get_activity_summary() clock — no parallel tracker
(rejected in PR #4864).
Fixes#98028Fixes#100325
Closes the #95663 round-8 review blocker (false settlement before
commit veto): the pre-commit surface (`_surface_stall`) logged
"Force-aborting the turn and stopping lease renewal" and warned the
user "aborting it so the session can recover" BEFORE `_commit_abort`
could veto — so a turn that resumed during the warning window (or an
exceptional interrupt path that declines fail-closed) was reported as
force-aborted with lease stopped while it actually continued running.
- Split the surface: `_surface_stall` is now observational only ("no
progress for Ns; attempting recovery"), and the definitive
aborted/lease-stopped settlement moves to a new
`_surface_committed_abort` that runs only after `_commit_abort`
succeeds and the turn lease is deactivated.
- Rate-limit repeated pre-commit surfaces per observed generation: a
turn whose aborts keep declining no longer re-logs an ERROR and
re-warns the user every poll interval.
- Add the committed-path regression test
(`test_watchdog_publishes_definitive_settlement_only_after_commit`)
and extend the declined-path witness
(`...resumes_during_warning`) to assert no committed-abort or
definitive pre-commit claim appears when the abort is vetoed. Both
fail on the pre-fix tree (mutation-checked).
- Document the `_interrupt_turn` lease-loss asymmetry (fires
unconditionally, no generation claim — losing the lease means the
process no longer owns the session).
- Trim review-round archaeology from comments/docstrings (keep the
WHY, drop the round numbering), and drop the dead
`cancel_event` compat note from the test fence.
- Document `agent.turn_liveness` in the configuration guide.
On top of PR #95663 by Finn763 (cherry-picked with authorship
preserved).
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
ring entry removed from keyless_mcp, legacy backend set / credential
ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
nous_subscription surfaces. The tvly- redaction pattern stays --
legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
environment-variables, tools-reference, web-dashboard, provider
plugin dev guide.
Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
The lean tail mode's per-chunk digest loop (_build_chunk_digests) issued up
to 28 extra call_llm requests sequentially per compaction attempt. With lean
now the default (#95571), users on slow auxiliary routes hit 7-11 minute
compactions (#96603). Remove the loop entirely: a lean compaction attempt now
makes EXACTLY ONE auxiliary LLM request — the main summary call.
- The detailed session log is folded into the single summary request: the
lean prompt template gains a '## Detailed Session Log (oldest first)'
section carrying the digest prompt's HARD RULES (identifiers verbatim,
dense bullets, transcript-is-data). Output guidance grows by
_LEAN_SESSION_LOG_BUDGET_TOKENS = 4,000 tokens on top of the scaled
summary budget — the old worst case (28 x 1,400 digest tokens) was spread
across many requests and mostly re-covered tool noise; a single dense
4K-token log inside one response preserves the load-bearing record while
staying well inside one aux response (the summary call still sends no hard
max_tokens, so no provider cap can truncate it mid-section).
- Input sizing: oversized regions (500K+ chars) are EVEN-SAMPLED across the
whole region (_sample_summary_input: 8 proportionally spaced slices,
oldest-to-newest, explicit '[... N chars elided ...]' markers, last slice
anchored to the newest end) instead of head+tail truncated, so session-log
coverage stays uniform. Legacy mode keeps _bound_summary_input unchanged.
- The LLM-free anchor index still runs over the FULL region, and the
session_search recovery footer is unchanged.
- Dead code removed: _build_chunk_digests, _LEAN_DIGEST_* constants,
_LEAN_DIGEST_PROMPT, _serialize_turns_for_digest, _digest_worthy,
_LOW_SIGNAL_TOOL_RE, the _lean_pristine_tools snapshot, and the
sibling-call route echo (_SUMMARY_ROUTE_CONSUMED /
attempt_summary_route_kwargs — no remaining callers; the single-use
summary pin semantics are unchanged).
- Tests pin the new contract (exactly one call_llm in lean mode; session-log
section lands in the summary; oversized regions sampled with elision
markers, never a second request; anchor index + recovery footer present).
Sabotage-verified: restoring a second call_llm makes the call-count test
fail. Docs and the compaction eval wording updated to stop claiming
per-chunk calls.
Fixes#96603.
Completes the #90953 salvage on post-#98237 main:
- New _merge_request_overrides helper defines the precedence contract:
explicit delegation.request_overrides merges OVER runtime/parent-derived
overrides — explicit top-level keys win; extra_body is deep-merged one
level so runtime extra_body keys survive unless redefined. Inputs are
copy.deepcopy'd so transport-side mutation can't leak into config or the
provider runtime cache.
- Direct base_url branch: explicit key now merges over the #98237
provider-alongside-base_url runtime overrides instead of being a separate
return shape; max_output_tokens preserved.
- Named-provider branch and parent-inherit branch now honor the key too, so
delegation.request_overrides never silently no-ops.
- _build_child_agent honors override_request_overrides whenever set
(previously only when override_provider was set), enabling the inherit
branch's merged value to reach the child.
- DEFAULT_CONFIG: delegation.request_overrides entry with comment.
- Tests: expanded tests/tools/test_delegate_request_overrides.py — deep-copy
proofs, explicit-over-runtime precedence on the provider-alongside-base_url
path, named-provider branch, inherit branch, and merge-helper unit tests.
- Docs: configuration.md delegation section + features/delegation.md document
the key, precedence, and example YAML (OpenRouter extra_body.provider.sort).
Extends PR #98250's classic-CLI status-bar upgrades to the Ink TUI:
- tui_gateway/server.py _get_usage() now emits cache_hit_pct,
avg_latency_s, avg_tps (reads the same per-call deque history from
agent/conversation_loop.py; keys omitted when no data — Codex
app-server has no latency, zero cache reads show no %)
- StatusRule renders the three read-outs as width-budgeted tail
segments (breakpoints 96/104/110 cols, lowest priority — they shed
first on narrow terminals)
- display.status_bar.fields (the SAME key the classic CLI honors)
filters TUI segments too: cache_hit, latency, tps, duration,
compressions, bg_tasks, bg_subagents, voice, battery, title,
context_pct, context_detail
- values ride the existing usage payload/ticker; constants between
events so the usage==last dedup keeps suppressing repaints
- 3 new server tests, 5 new TUI tests; full ui-tui suite 1727 green
Follow-up to the cherry-picked #41909/#92696 + #39760 + #97970 cluster:
- single field-key namespace (display.status_bar.fields) instead of the
second tui_statusbar_fields list; cache_hit/latency/tps/stash/battery/
title join the existing key set
- cache-hit % prefers the baseline-delta regime (resets on model switch
and compression) and hides on zero cache reads instead of alarming 0%
- latency/tps segments added to the styled fragment renderer too
- docs updated in website/docs/user-guide/configuration.md
- 7 new tests: rolling latency/t/s, NaN/negative guard, field filtering,
baseline resets, title badge gating
Allow users to control which fields appear in the interactive CLI status
bar via display.status_bar.fields in config.yaml.
Available fields: model, context_pct, context_detail, compressions,
bg_tasks, bg_processes, duration, prompt_elapsed, yolo, total_tokens.
When the list is empty (default), all fields are shown as before.
The field order is fixed (model always first); the config controls
visibility only. Narrow terminals (<76 cols) automatically drop
context_detail regardless of config.
total_tokens is opt-in only (not shown by default) to avoid width
overflow in the prompt_toolkit fragment renderer.
Closes#41909