A goal quality gate is a shell command persisted by `/goal gate add` and
later executed with `subprocess.run(shell=True)` at every goal turn
boundary (run_gate), with no approval prompt. On the messaging gateway,
slash access is backward-compatible: with no `allow_admin_from` list
configured (the default), every *allowed* chat user is treated as
unrestricted. So an allowed but non-admin remote sender could add an
arbitrary shell command and get authenticated RCE as the Hermes process
account.
Gate ONLY the shell-creating operation (`gate add`) behind the existing
fail-closed explicit-admin check (`_resume_caller_is_admin`, the same one
that guards cross-origin `/resume`). `gate list` / `remove` / `clear`
stay open so a non-admin can still inspect and recover. CLI/TUI/Desktop
`gate add` is local (the user already has host access) and is unchanged.
Salvaged from #91677 by @unsupportedpastels — narrowed to the single
shell-creating op and reusing the existing admin helper rather than
renaming it. Contributor's regression tests preserved.
Co-authored-by: unsupportedpastels <theoldwizard123@pm.me>
Voice state was keyed `<platform>:<chat_id>` with no profile namespace,
so two bots in one Discord channel shared one /voice mode; every
`_voice_input_callback` was the bare `_handle_voice_channel_input`, which
(like `_handle_voice_timeout_cleanup` and the /voice slash handler) always
picked `self.adapters[DISCORD]` — a secondary profile's voice transcripts
were dispatched through the default profile's bot.
- `_voice_key(platform, chat_id, profile=None)`: named profiles get a
`<profile>:` prefix; default keeps the legacy shape (persisted state valid).
- `_voice_key_for_source` keys by the transport-OWNING profile
(`_adapter_profile_for_source`), matching what `_sync_voice_mode_state_to_adapter`
now restores per adapter via `_owner_profile`.
- `_bind_voice_input_callback` binds the capturing adapter into the
transcript handler (functools.partial); used at primary connect,
primary reconnect, /voice channel join, and `_configure_profile_adapter`.
- `_handle_voice_timeout_cleanup` takes the adapter it was bound to.
- /voice, join, leave and `_should_send_voice_reply` resolve the adapter via
`_adapter_for_source` (fail-closed) instead of `self.adapters[platform]`.
- #84872: `_start_one_profile_adapters` now calls
`_sync_voice_mode_state_to_adapter` on secondary INITIAL connect, as the
primary path and both reconnect paths already did.
Co-authored-by: davidxyuan <124700534+davidxyuan@users.noreply.github.com>
GatewayRunner.__init__ snapshotted _ephemeral_system_prompt once from the
launch profile's config and _get_system_prompt_for_channel returned that
string for every source, so under multiplex a routed profile's
display.personality / agent.system_prompt never injected (#89161), and
/personality from any chat rewrote the one process-global attribute for
everyone.
Drop the snapshot: _get_system_prompt_for_channel now calls
_load_ephemeral_system_prompt() (env var, then
resolve_ephemeral_system_prompt_from_config(_load_gateway_runtime_config()))
on each call. Its caller run_sync already runs inside
_profile_runtime_scope, so the routed profile's config.yaml is what gets
read; single-profile hot-edits of the personality also take effect on the
next turn instead of requiring a restart. /personality only persists via
persist_personality() (get_hermes_home()/config.yaml = the routed profile)
and no longer touches in-memory state.
Fixes#89161
Co-authored-by: worlldz <101180447+worlldz@users.noreply.github.com>
Slash dispatch already runs inside _profile_runtime_scope under multiplex,
but _save_gateway_config_key (/reasoning --global, /fast, show/hide),
/memory approval, /skills approval, /verbose and /footer built their write
path from the module constant gateway.run._hermes_home — the launch home —
so a routed profile's toggles landed in the default profile's config.yaml
while the reads (via _gateway_config_home()) saw the routed one.
Resolve the write path through _gateway_config_home() at all five sites so
reads and writes agree. Single-profile gateways never install the override
and keep resolving the launch home.
Fixes#87939Fixes#75684
Co-authored-by: Bao <nnqbao@gmail.com>
Sibling sites of the bare loop.run_in_executor(None, …) class fixed for
/insights, /debug and /goal draft: the worker started with an empty
context, so get_hermes_home()-relative reads (skills.external_dirs,
disabled skills, the reviewer subagent's home and secret scope) resolved
the launch home instead of the routed profile under multiplex.
The multiplexed inbound handler wraps every message in _profile_runtime_scope,
which installs the routed profile's HERMES_HOME override and its secret scope
as contextvars. A bare loop.run_in_executor(None, fn) starts the worker with an
EMPTY context, so neither reaches the blocking work.
GatewaySlashCommandsMixin already knows this -- /compress goes through
_run_in_executor_with_context and the call site says why. Three siblings in the
same file still used the bare hop:
/insights SessionDB() with no explicit path resolves get_hermes_home() at
call time (_default_db_path), so the worker opened the DEFAULT
profile's state.db. Under multiplexing the command reported
another profile's conversations, session counts and sources to
this profile's user.
/debug collects that home's logs/config and uploads them to a public
paste, so it published the default profile's diagnostics from
another profile's chat.
/goal draft calls the auxiliary LLM, whose provider/credential resolution
reads the profile secret scope -- unscoped it falls back to
process-global os.environ, which under multiplexing may hold a
different profile's keys.
Route all three through _run_in_executor_with_context.
/reload-skills is deliberately left alone: tools.skills_tool binds SKILLS_DIR
at import time, so it does not follow the contextvar either way. Fixing that
needs the module-global retarget web_server._profile_scope performs under a
lock, which is a different change from context propagation.
Single-profile gateways never enter the scope, so their behaviour is unchanged.
A multiplexed Hermes process (gateway.multiplex_profiles, unified
dashboard/TUI, or cron) serves several profiles at once, but terminal.*
resolved through process-global TERMINAL_* env vars bridged ONCE at
startup from the launch profile (gateway/run.py ~2700-2760) plus the
one-shot _ensure_terminal_env_bridged() guard. Every routed profile
therefore inherited the launch profile's backend, cwd, docker volumes,
SSH target and shared-container key: a local profile ran inside another
profile's docker sandbox (or a docker profile escaped to the host), and a
container labeled profile A carried profile B's RW bind mounts.
Fix: an authoritative per-profile terminal policy seam, mirroring
agent/secret_scope.py:
- tools/terminal_scope.py: ContextVar holding the routed profile's
COMPLETE effective TERMINAL_* policy (defined defaults <- profile .env
TERMINAL_* <- config.yaml terminal:). While bound, terminal_env()
resolves ONLY from it - an omitted key yields the defined default,
never os.environ. Unreadable/malformed policy installs a refusal
scope; terminal_tool / execute_code refuse instead of running under
ambient launch-process policy (fail closed).
- Installed at every in-process profile boundary: gateway
_profile_runtime_scope, tui_gateway session/build/turn scopes, cron
per-job fire. The unscoped single-process path is byte-identical.
- Every terminal.* consumer reads through the scope: terminal_tool
(_get_env_config, _resolve_container_task_id shared key, orphan
reaper lifetime, degraded mode), gateway/platforms/base.py docker
media translation (volumes, shared key, persistence), runtime_cwd /
agent_init / skill_utils / code_execution_tool / file_tools cwd
anchors, prompt_builder / browser_tool / env_probe backend checks,
gateway footer, @-refs and slash-command cwd. env_probe resolves the
backend in the caller's context, since the probe worker thread does
not inherit the ContextVar.
Salvage of #99225 onto current main: adds the three ambient reads the PR
missed (tools/file_tools.py TERMINAL_CWD, tools/browser_tool.py and
tools/env_probe.py TERMINAL_ENV; shape from #79117) and trims the test
module to the leak matrix driven through the real gateway boundary,
omitted-key defaults, refusal, and boundary reset.
Fixes#68559Fixes#94200Fixes#101132Fixes#95470
Co-authored-by: x7peeps <9640837+x7peeps@users.noreply.github.com>
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: ExitMaster <292490062+ExitMaster@users.noreply.github.com>
Adds two bounded fast modes on top of the static /fast toggle, default OFF:
- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
window; requests inside it carry the provider fast param, later tool-loop
requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
user/assistant/tool history).
agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.
resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.
Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.
Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes#64785, #74730.
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
Move the structural clone from the four call sites (auto review, codex
runtime, CLI /refine, gateway /refine) into AIAgent._spawn_background_review,
which every review path — immediate, idle-queue deferred, requeued — passes
through. Callers can no longer forget it, and the private helper is no longer
imported across hermes_cli/ and gateway/ package boundaries.
Tests now bind the real chokepoint (capturing at _spawn_background_review_now)
so they still fail if the clone is removed.
Widen #100802 to the two explicit review entry points. The CLI and gateway
/refine handlers built their own snapshot with a shallow list(), which
aliases the nested tool_calls/content containers of the live history. The
review fork sanitizes its transcript in place (sanitize_tool_call_arguments
rewrites function["arguments"]), so a /refine could rewrite the parent's
persisted transcript exactly like the automatic review could (#100795).
Both sites now use _clone_background_review_messages, the same structural
clone the automatic review uses. Regression tests drive the real handlers
and assert the snapshot shares no containers with the live transcript.
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).
Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.
- acquire(path): same resolved path returns the same instance (one
writer connection, one lock, one token-writer thread) for every
long-lived in-process caller (gateway runner, SessionStore, per-agent
lazy recall, cron per-job, mirror, channel_directory, slash_commands,
shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
lifecycle, so one caller's close can never tear down a writer other
callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
RETIRES the live generation (never lent again) but keeps it alive for
existing holders; release is object-keyed so holders of the old
generation drain it independently of the new one. The old
generation's own write path still fails with the typed
StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
generation (live + retired) as the final safety net.
CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.
References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
A non-admin '/sessions all' or '/resume --all' silently downgraded to
chat-scoped listing with zero feedback, which reads as 'my session
vanished' (community Telegram report). Both surfaces now append a notice
that cross-chat listing requires a configured admin. Follows up the
salvaged current-session '(current)' marker (PR #68556, fixes#68547):
sibling tests updated to pin the new contract, new i18n key
gateway.resume.all_requires_admin added to all 17 locales, docs updated.
On the codex_app_server runtime the model's real working context is the
app-server's server-side thread: CodexAppServerSession is constructed with
no history and each turn submits only the new user message
(agent/codex_runtime.py), so Hermes' transcript is a mirror that is never
replayed into a thread. Every out-of-turn compression call site (gateway
session hygiene, gateway /compress) built a DETACHED agent whose
_codex_session was None, so the codex route bailed at its "no active codex
thread" guard and returned the transcript unchanged ("compressed 150 ->
150 msgs") — and hygiene's finally-clause then evicted the cached live
agent, destroying the only real context: the next turn spawned an empty
thread while Hermes still mirrored a full history.
Fix, per the documented compression.codex_app_server_auto contract:
* Session hygiene now routes codex_app_server sessions to
run_codex_hygiene_compaction(): in 'hermes' mode it compacts the LIVE
cached agent's thread via thread/compact/start (through the existing
codex route in _compress_context) and KEEPS that agent cached; 'native'
and 'off' skip cleanly with no eviction and no local fallback. A wedged
compaction records the persistent failure cooldown; success resets the
hygiene failure streak.
* Gateway /compress detects the codex_app_server runtime before building
a temporary compression agent and compacts the live thread with
force=True instead (a manual compress is an explicit user decision in
every mode). No live thread -> honest "nothing to compact" reply
instead of a mirror rewrite plus eviction.
* No mode ever runs the local transcript compressor on this runtime:
rewriting the mirror cannot shrink the thread, so the #73715-style
local fallback (including its force=True leak into native/off) is
deliberately not adopted.
Diagnosis of the mode-gate/no-thread deadlock builds on PR #73715.
Closes#73503
Co-authored-by: webtecnica <webtecnica@gmail.com>
Keep model-switch callers compatible with result objects created before runtime_capabilities was added, and do not roll back minimal agents that lack optional LM Studio helpers. Preserve rollback for real helper failures.
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
Carry provider-derived request_overrides through runtime resolution,
fallback projection, session /model state, restart rehydration, and
turn-route merge so named custom providers keep extra_body and related
overrides.
The initial /btw implementation (#97937) answered from a rendered
plain-text transcript digest — truncated context, cold-written tokens on
every question. Teknium's call: reuse the self-improvement review fork
instead, which keeps the entire prompt cache stable for the fork and
gives it the complete conversation for very cheap.
- agent/background_review.py: extract the review-fork construction into
build_cache_parity_fork() — same runtime/credentials as the parent,
byte-identical system prompt / tools[] / reasoning config on the
same-model path, shared session_id for prefix warmth, full persistence
detachment (no state.db writes, no rotation, no external memory,
in-place-only compaction). The review thread now calls the helper;
behavior unchanged (full review test suite green).
- agent/side_question.py: /btw prefers the fork when a live parent
AIAgent exists — replays the untruncated snapshot as warm cache reads,
denies every tool at dispatch via an empty thread whitelist (tools[]
stays byte-identical for cache parity), attributes usage to the parent,
and trims a mid-turn snapshot tail so role alternation holds. The
one-shot digest remains as fallback (no live agent = cold cache anyway,
and any fork failure degrades gracefully).
- CLI passes self.agent, TUI passes the session agent, gateway looks up
the chat's cached agent (parity with how turns reuse it).
Live-verified: /btw on the worktree runs the fork path (agent.log shows
the side question as a forked conversation turn on the parent session_id
with the full history replayed), answers correctly from context.
Review finding (quality pass): on a single-profile gateway,
_handle_busy_command set _busy_input_mode but left _busy_text_mode
stale, so the adapter refresh a line later re-read the old value —
/busy queue persisted to config but live text messages kept
interrupting until restart. The profile path already re-derives both
from the fresh config; the non-profile path now does the same via
_load_busy_text_mode() (busy_input_mode is the source of truth,
run.py:9877). Regression assertion added to test_set_mode_persists;
verified red without the production fix.
Teknium review items:
1. Parse with event.get_command_args() instead of raw event.text
(matching _handle_fast_command pattern at line 2842)
2. Add mocked persistence tests for queue/steer/interrupt setter
success, save-failure, and exception branches (5 new tests)
Removes cli_only=True from /busy CommandDef and adds gateway
handler with subcommand dispatch (status/queue/steer/interrupt).
Applied on top of latest upstream/main while preserving original
commit intent from PR #18366.
Also adds smoke tests for the gateway /busy command handler.
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.
/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.
Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
Pattern-A architectural fix: blocking calls inside async functions freeze
the gateway/uvicorn event loop for every adapter, timer, and health check.
Known incidents: 17-minute getaddrinfo freeze (#91912 class), 10s restart
freeze in start_gateway (#36163).
Fixes at the four unguarded core sites:
- gateway/platforms/webhook.py: `gh pr comment` subprocess (30s timeout)
now runs via asyncio.to_thread — a webhook delivery no longer freezes
every other platform for the duration of a network call.
- gateway/run.py start_gateway --replace: two time.sleep() waits (10s +
5s worst case) become await asyncio.sleep() (re-lands #36163 at current
line numbers, credit AhmetArif0).
- gateway/slash_commands.py /save: session render + file write move off
the loop (scales with transcript size).
- hermes_cli/web_server.py voice TTS: multi-MB audio file read + unlink
move off the loop.
Prevention gate so the bug class cannot re-enter:
- pyproject.toml [tool.ruff.lint] select gains ASYNC210/220/221/251
(blocking HTTP / Popen / subprocess.run / time.sleep in async def).
These run in the existing blocking `ruff check .` CI job.
- Frozen ratchet baseline in per-file-ignores for the remaining legacy
sites (detached restart watchers; router sweep in flight via #84376;
two platform adapters), each documented for burn-down. New files or
new violations fail CI immediately.
- tests/** keeps the relaxation (deliberate sleeps in fixtures).
Verification:
- ruff check . green on this branch; sabotage file with time.sleep +
subprocess.run in async def fails the gate with 2 errors.
- New behavioral test test_webhook_offloop_delivery.py asserts loop
liveness DURING delivery (ticker coroutine): 1 tick on the old
blocking code (fails), 21 ticks off-loop (passes).
- 50 webhook/replace gateway tests + 26 save/export tests pass.
Co-authored-by: AhmetArif0 <147827411+AhmetArif0@users.noreply.github.com>
- The direct adapter.send() confirmation path in /approve and /deny is only
needed on native-streaming platforms (WeCom) where the reply stream is
already finalized; other platforms keep the return-text contract (fixes
4 approve/deny regression tests, guarded with 'is not True' against
MagicMock auto-attributes).
- Remove a stray [DEBUG] logger.info left in _deliver_media_from_response.
- website/docs wecom.md: replace the 'does not stream' notes with the native
msgtype:stream behavior and document the stream keepalive extra keys.
Implement native reply streaming for the WeCom (企业微信) adapter over the
long-connection "msgtype: stream" transport, so a reply renders as a single
live-updating typing bubble instead of one final block. Aligns with the
official wecom-openclaw-plugin streaming behavior.
Includes the machinery intrinsic to native streaming on WeCom:
- Transport: seed frame (<think></think>) opens the typing bubble, intermediate
frames update it, a finalize frame closes it; native-streaming adapters are
let past the edit-only gate. Fire-and-forget intermediate frames (WeCom
long-connection mode has no documented edit-rate limit); an adapter-level
frame cap is retained. (Early builds gated frames behind a char throttle;
removed in favor of fire-and-forget + identity dedup.)
- Per-turn isolation: each turn owns a unique turn_id; concurrent messages are
isolated via (chat_id, turn_id)-keyed state. Dual-lane priority queue
(control vs normal) plus a per-chat token bucket to stay under WeCom's rate
limit (errcode 846607).
- Dedup-safe delivery + ack-race handling: deliver-once contract (a frame is
delivered the moment it is emitted; failures logged, not re-sent; delivery
marked once per turn), per-req_id reply queue with ack tracking, and the
timeout-inversion / orphan-queue race fixes. Robust fallback on 846608 /
846609 / errcode 6000 / passive-reply timeout via proactive send.
- Interaction boundaries: finalize + reset before approval/clarify prompts so
the prompt is the last thing on screen and never traps a lingering bubble;
eager re-seed after a clarify answer so the typing bubble reappears instantly.
- Stream-level keepalive: optional periodic finish=false frame + finalize-time
stream-age guard to refresh WeCom's ~6-minute reply-stream window on long
turns (mitigates 846604 / 846608). Off by default; tunable via config.yaml.
- Tool-progress folded into the same native-stream bubble instead of separate
messages; image+text double-callback merged into one turn.
Tests cover the streaming lifecycle, per-turn isolation, duplicate-send / ack
timing, approval + clarify boundaries, eager re-seed, and tool-progress.
Folds review findings: surface failed_deletes in CLI and gateway
/rollback output (new gateway.rollback.failed_deletes locale key, 17
locales), emit skipped_oversize on the nothing-to-restore early return
too, document all three report keys in the restore() docstring, and pin
the failed_deletes contract from both sides in tests.
Follow-up to the salvaged #95207 fix, completing the misreport bug class:
- restore() now also drops delete_targets whose unlink failed (OSError
swallowed) from restored_files — the sibling of the kept-oversize
misreport the salvaged fix closed.
- /rollback output in the CLI (cli_commands_mixin) and gateway
(slash_commands + gateway.rollback.kept_oversize locale key in all 17
catalogs) now tells the user which files were kept because the size
cap excluded them from every checkpoint; previously the file was
correctly preserved but the user got no notice it was not reverted.
- Regression test for the failed-unlink misreport.
Addresses the review on #93996:
- gateway: hygiene and manual /compress load the memory provider only when
compression.checkpoint_required is enabled (skip_memory=not required).
The historical fast path — no provider init, no best-effort hook — is
back for everyone who did not opt in, so default behavior is truly
unchanged.
- conversation_compression: assistant messages carrying both prose and
tool_calls keep their prose in the checkpoint evidence (the tool_calls
payload is stripped, the original message is not mutated); pure
tool-call wrappers without prose are still dropped.
- tests: legacy-database regression proving the _compressed_summary column
is added by the declarative _reconcile_columns() path on a plain reopen
(no version-gated migration needed — append_message works right after),
plus coverage for the prose-preserving filter.
- docs: providers must implement idempotent, content-keyed checkpoint
writes — a fail-closed block means the next attempt re-runs
on_pre_compress over largely the same transcript.
Refs #93986
Context compression is intentionally lossy. Deployments that archive
transcript evidence to an external durable store before compaction had no
way to guarantee the archive actually happened: MemoryManager.on_pre_compress
swallows provider failures by design, so a failed archive silently degraded
into data loss.
This adds an opt-in, provider-agnostic checkpoint contract:
- memory_provider: PRE_COMPRESS_CHECKPOINT_API_VERSION = 1; providers opt in
by advertising pre_compress_checkpoint_api_version. Version 0 keeps the
historical best-effort hook semantics.
- memory_manager: supports_pre_compress_checkpoint() capability probe;
on_pre_compress(require_checkpoint=True) propagates checkpoint-provider
failures and raises when no capable provider completed the checkpoint.
- conversation_compression: new compression.checkpoint_required config key
(default false, documented in cli-config.yaml.example). When enabled,
compaction fails closed with BLOCKED_MISSING_PREREQUISITE (the
uncompressed transcript is preserved) unless a checkpoint-capable provider
confirms the durable checkpoint. Providers receive normalized direct
user/assistant evidence: tool rows, system messages, tool-call wrappers,
and prior compaction summaries are filtered host-side into one stable
contract. codex_app_server compaction is rejected under the gate because
it exposes no truthful pre-compaction transcript boundary.
- hermes_state: persistent _compressed_summary column (declarative schema
migration via _reconcile_columns) so summary provenance survives process
restarts; only the resume model history carries the marker, keeping
get_messages_as_conversation on its existing contract.
- gateway: the lossy hygiene/auto-compact paths load the memory provider
(skip_memory=False) so a required checkpoint also guards those rewrites.
The gate arms only on an explicit boolean True (bare-MagicMock agents in
existing tests have truthy auto-attributes). Default behavior is unchanged:
checkpoint_required=false preserves best-effort semantics for all existing
providers. Contract tests, including a restart round-trip of the summary
marker, in tests/agent/test_pre_compress_checkpoint_contract.py.
Refs #93986
/review takes the last 10 chat messages plus optional instructions,
spawns a full-privilege background subagent (the async delegation
rail) that investigates the referenced work (PR, code, docs), and its
complete review re-enters the spawning session as a normal
async-delegation completion the primary agent can act on.
- agent/review_engine.py: shared engine (snapshot, briefing,
auxiliary.review credential resolution, dispatch, note formatting)
- tools/delegate_tool.py: internal credentials_cfg per-call override
(never model-facing) resolved through the same credential system as
delegation.provider pins
- auxiliary.review config block (provider/model/base_url/api_key/
api_mode); provider auto + empty model = inherit the main model
- Surfaces: CLI process_command, gateway run.py dispatch +
slash_commands handler (binds the approval session key so the
completion routes back), TUI/Desktop live dispatch in
tui_gateway/server.py, CommandDef registry (+Slack /hermes-only cap)
- Docs: delegation.md section + slash-commands.md (both tables)
- Tests: 15 engine tests (sabotage-verified: credentials_cfg and
dispatch tests fail without the fix), 4 gateway handler tests
through the real async rail