Context compression is intentionally lossy. Deployments that archive
transcript evidence to an external durable store before compaction had no
way to guarantee the archive actually happened: MemoryManager.on_pre_compress
swallows provider failures by design, so a failed archive silently degraded
into data loss.
This adds an opt-in, provider-agnostic checkpoint contract:
- memory_provider: PRE_COMPRESS_CHECKPOINT_API_VERSION = 1; providers opt in
by advertising pre_compress_checkpoint_api_version. Version 0 keeps the
historical best-effort hook semantics.
- memory_manager: supports_pre_compress_checkpoint() capability probe;
on_pre_compress(require_checkpoint=True) propagates checkpoint-provider
failures and raises when no capable provider completed the checkpoint.
- conversation_compression: new compression.checkpoint_required config key
(default false, documented in cli-config.yaml.example). When enabled,
compaction fails closed with BLOCKED_MISSING_PREREQUISITE (the
uncompressed transcript is preserved) unless a checkpoint-capable provider
confirms the durable checkpoint. Providers receive normalized direct
user/assistant evidence: tool rows, system messages, tool-call wrappers,
and prior compaction summaries are filtered host-side into one stable
contract. codex_app_server compaction is rejected under the gate because
it exposes no truthful pre-compaction transcript boundary.
- hermes_state: persistent _compressed_summary column (declarative schema
migration via _reconcile_columns) so summary provenance survives process
restarts; only the resume model history carries the marker, keeping
get_messages_as_conversation on its existing contract.
- gateway: the lossy hygiene/auto-compact paths load the memory provider
(skip_memory=False) so a required checkpoint also guards those rewrites.
The gate arms only on an explicit boolean True (bare-MagicMock agents in
existing tests have truthy auto-attributes). Default behavior is unchanged:
checkpoint_required=false preserves best-effort semantics for all existing
providers. Contract tests, including a restart round-trip of the summary
marker, in tests/agent/test_pre_compress_checkpoint_contract.py.
Refs #93986
Widens the cherry-picked reasoning-callback fix to the whole leak class
(#93220):
- quiet branch also neutralizes tool_progress_callback,
tool_start_callback, tool_complete_callback (inline diff rendering via
render_edit_diff_with_delta was gated by NEITHER quiet_mode nor
tool_progress_mode) and syncs agent.tool_progress_mode='off'.
- _should_emit_quiet_tool_messages() returns False under
suppress_status_output: with callbacks neutralized, the quiet-mode
KawaiiSpinner fallback printed '[tool]'/'[done]' lines into captured
stdout. Also covers oneshot.py and background-review forks, which set
the same flag and expect strict silence.
E2E (isolated HERMES_HOME, live model, write_file turn): base leaks
'┊ review diff' + full SVG source into stdout; head emits exactly the
final response. Regression tests pin the quiet-branch statements and the
gate (sabotage-verified).
Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
The outer conversation-loop except handler only left the loop on a
local-processing error or when api_call_count >= max_iterations - 1.
With the turn budget now unlimited by default (sys.maxsize), a
permanent failure that escaped the inner retry/fallback machinery
retried forever: ~64 retries/s, one core pegged, and the rotated
agent.log history overwritten within minutes.
Bound the loop with a small per-turn cap on total escaping exceptions
(_MAX_OUTER_LOOP_ERRORS = 8, scaled down by a tiny explicit
max_iterations so a manually bounded budget still governs). The legacy
local-processing and near-limit exits are byte-identical; a new
'repeated_outer_errors' exit reason gets a user-facing explanation.
The inner retry/fallback layer owns transient API recovery and
terminates on its own, so only exceptions that escape it reach this
cap - a successful turn is unaffected.
Fixes#92450
Reviewer point on #92539: nothing documented that an in-place content
mutation of a stamped dict must pop _DB_PERSISTED_MARKER (and invalidate
the bounded flush-scan prefix) or the DB silently goes stale. Both
existing mutators (turn_finalizer fill-empty-tail, context_compressor
micro-compaction defrag) already follow the contract; this states it at
the constant so the next one does too.
- credits_tracker: trim inline comment block (duplicated docstring) and
correct its safety claim - a paid model under stealth/ would fail
closed (suppressed banner), not open; state the trade-off honestly.
- run_agent: update stale call-site comment to mention stealth/ prefix.
- auxiliary_client: widen sibling free-SKU detector _is_free_model to
recognize stealth/ prefix (same bug class as #91843: free_only=true
wrongly skipped the OpenRouter fallback and the paid-lane warning
fired spuriously for stealth models).
- tests: bind the new sibling behavior (stealth/ox-alpha free,
my-stealth/model not).
Change fail-closed behavior to proceed-with-warning when a background
review does not acknowledge cancellation within the bounded deadline.
The review is non-critical self-improvement work and must never block
a user-facing turn (#84423). Keep the off-thread interrupt to ensure
a broken abort path cannot stall the bounded wait.
Builds on @fattchris resolve_turn_limit salvage (#67696): flips the default
from a numeric cap to unlimited across all construction paths (CLI, agent_init,
run_agent subagents), adds inf/infinity/null to the unlimited spellings, and
sets DEFAULT_CONFIG agent.max_turns to null. The turn cap caused more problems
than it solved (silent mid-task truncation).
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.
Four pieces, and they only make sense together:
· an in-page engine that inventories what is interactable and performs the
verb, injected as source because it has to run inside the guest page;
· the preview.act.request bridge from the gateway into the pane;
· drive_preview, for acting: elements, click, type, scroll, press, and the
pane's own back/forward/reload;
· annotate_preview, for marking without acting.
Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.
Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.
Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
When compression is explicitly disabled (compression.enabled: false), conversations can grow past the model's context window across hundreds of messages (e.g., 824 messages / 460K+ tokens in #89297). Serializing massive JSON payloads repeatedly under memory-constrained environments leads to swap thrashing (STAT=U) and unhandled provider errors.
Add a pre-flight uncompressed context overflow guardrail in build_turn_context and a deduped _warn_uncompressed_context_overflow method on AIAgent to alert users to run /compact or enable compression before unmanageable payloads freeze the process.
Composio eval traces showed Hermes wasting turns re-issuing identical tool
calls (same tool, same args, same result — 3x/4x in one run) and ending
turns by announcing an action it never took. Two conservative, config-gated
guards (agent.stall_guards, default true):
- Identical-call loop breaker: ToolCallGuardrailController.observe_identical_call
tracks the consecutive streak of (tool, canonical args, result-hash); on
the 3rd identical call a compact one-line notice is appended to that tool
RESULT at construction time (cache-safe — tool results are append-only).
Never blocks the call. Pollers (process, *_get_result, *_poll) are exempt
via STALL_GUARD_REPEATABLE_TOOLS. Streak resets on any different call,
changed result, or new turn. Observed on the raw result before the
tool-loop warning suffix so its changing count can't defeat matching.
- Said-continue-but-stopped recovery: trailing_continue_intent() detects a
short reply ENDING on an announced next action ('Let me now…', 'I will
now…', 'Next, I…'); the conversation loop feeds it into the EXISTING
intent-ack continuation path (same interim-assistant + user-nudge
mechanism, same codex_ack_continuations cap of 2), preserving message
alternation — no parallel recovery machinery.
Config: agent.stall_guards in DEFAULT_CONFIG; docs in configuration.md;
unit tests for streak/allowlist/reset/gate and detector pos/neg cases.
The identical 6-line try/except block for reading model.reasoning_echo
from config appeared in both agent_init.py (init) and
agent_runtime_helpers.py (switch_model). Extracted into
AIAgent._read_reasoning_echo_from_config() static method — net -1 LOC.
Add model.reasoning_echo (default false) and per-fallback-entry
reasoning_echo to preserve assistant reasoning_content when
replaying history to custom providers and OpenAI-compatible gateways
that proxy thinking-mode models (Kimi K3, GLM-5.2, DeepSeek, etc.)
but are not matched by the built-in host-based _REASONING_ECHO_RULES.
The flag is per-active-provider, not a global toggle:
- Primary: read from model.reasoning_echo at init and switch_model
- Fallback: set by try_activate_fallback from the fallback entry
- Restore: restore_primary_runtime copies the switch_model snapshot
Unlike PR #76019 global agent.reasoning_echo toggle, the
per-provider flag travels with the active provider — falling back to
a strict provider (Mistral, Groq, Cerebras) correctly strips
reasoning_content even when the primary had the flag enabled,
because the flag is False for the strict fallback.
Complements PR #27361 (dynamic detection) which fires after the first
API response; this PR covers turn-1 and history-replay-on-fresh-session
where dynamic detection has not fired yet.
Closes#76018
Refs: #27297, #27361, #76019
Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
One generic tool in the desktop_ui toolset: discover what is on screen,
highlight an element with narration, or hand the user a paged tour. No tour
content lives in the code — the agent authors each one live, which is what
makes 'how does this work?' answerable as a walkthrough instead of a wall
of text.
Rides the existing blocking-prompt bridge (tour.request/.respond) like
read_preview, so it works on every connection topology.
Two data-loss bugs reported by users:
1. /handoff CLI→gateway race (#88234): After /handoff completed, CLI
cleanup called finalize_session on the session the gateway just
reopened. This set end_reason on a row the gateway was actively
writing to, causing the handoff leg to vanish from session history
and breaking session_search recall. Fix: add _handed_off_session_ids
module-level set (mirrors _single_query_finalize_attempted_session_ids
pattern). _handle_handoff_command registers the session_id on
completion; _should_emit_cleanup_session_finalize and
_emit_interrupted_session_end check it before firing.
2. state.db corruption silent failure (#88235): When SessionDB init
failed at gateway startup, the error stayed in logs — messages
flowed but nothing was persisted, with no user-visible indication.
Fix: store _session_db_init_error on GatewayRunner, broadcast a
recovery-guidance message to all home channels via
_send_session_db_warning_notifications() after the gateway connects.
Also improved the 'corrupt' persistence cause wording in
_format_turn_completion_explanation to include the full recovery
path (hermes doctor --fix, sqlite3 .recover, backups).
Tests: 6 new tests for handoff cleanup race, 3 for corruption wording.
All existing CLI/turn-completion tests pass.
Copilot CLI 1.0.78 reworked /rewind to restore only the files the agent
changed, 'skipping any file whose contents no longer match what Copilot
last wrote'. This ports that protection to Hermes checkpoints:
- tools/checkpoint_manager.py: per-project agent-write ledger
(sha256 of every landed write_file/patch), safe_restore_plan()
classifier, and restore(safe=True) that reverts only agent-authored
changes, deletes agent-created files, and preserves user hand-edits.
Empty ledger (pre-existing stores) falls back to the classic full
restore.
- run_agent.py: feed the ledger from _record_file_mutation_result on
every landed mutation (zero new hooks; rides the existing verifier).
- CLI + gateway /rollback: safe mode is the default; --all/--force
restores everything; skipped files are reported with a hint.
- 17 locales: new gateway.rollback.kept_user_edits key.
- Docs: checkpoints-and-rollback.md updated.
- Tests: 7 new cases incl. user-edit preservation, post-agent user
tweaks, agent-created file removal, empty-ledger fallback.
'database disk image is malformed' contains the word 'disk', so
classify_persistence_error bucketed SQLITE_CORRUPT / SQLITE_NOTADB
failures as 'disk' and the turn-completion explainer told users to
free disk space for a structurally damaged state.db (the #77386-family
misdiagnosis, reproduced in the v0.20.0 malformed-DB incident report).
- hermes_state: new 'corrupt' bucket in PERSISTENCE_ERROR_CAUSES,
matched via _DB_CORRUPTION_MARKERS BEFORE the locked/disk buckets
- run_agent: explainer text for 'corrupt' points at hermes doctor and
explicitly says freeing space will not help
- cron explainer-variant suppression picks the new variant up
automatically (it iterates PERSISTENCE_ERROR_CAUSES)
Follow-up to #87400: drop the max_iterations and prompt_file knobs from
auxiliary.background_review. The aux model routing (provider/model/
base_url/...) predates #87400 and stays; the enabled switch and the
usage telemetry stay. The fork's iteration budget returns to the
historical hardcoded 16.
Persist fork token usage under session_model_usage task=background_review,
emit a per-fork completion log line, and expose enabled/max_iterations/
prompt_file so operators can see and bound the automatic review cost.
Address review feedback: load auxiliary.background_review once per spawn,
classify completion logs by summarize action prefixes, treat explicit
api_call_count=None as the documented default of 1, and WARNING on the
fail-open enabled-gate path.
The post-turn background review fork (`agent/background_review.py`) inherits
the parent agent's live runtime by default. That is a cost win when the parent
IS the main chat model (warm prompt cache, cheap), but the fork also fires
inside delegation subagents, where it inherits the *subagent's* model. When a
subagent runs a premium delegation model, the review silently replays the whole
conversation and emits skill/memory-update output at premium rates, with
nothing in the log or config flagging it (#85859).
Subagents are already barred from writing shared MEMORY.md
(`DELEGATE_BLOCKED_TOOLS`) and are spawned with `skip_memory=True`, so an
automatic review here has little to persist. Guard `_spawn_background_review`
(the single choke point both the turn-finalizer and codex-runtime callers pass
through) to return early when `_delegate_depth > 0`. An explicit `/refine`
(`focus` set) is a deliberate user request and still runs; the top-level path
is unchanged.
Fixes#85859
_file_mutation_verifier_enabled and _turn_completion_explainer_enabled
re-read config.yaml on every call via load_config() (~1ms deepcopy per
call). finalize_turn runs these gates at the end of every turn, so each
turn paid two redundant config deepcopies. The sibling
_credits_notices_enabled already caches on self; mirror that pattern.
The env-var override stays authoritative and uncached, so runtime flips
still work. Config flips now apply on the next session, matching the
documented sibling semantics.
(cherry picked from commit f2d0e00d71ba35fa3a53cd41076261e3c5feb45d)
_create_request_anthropic_client() built a fresh anthropic.Anthropic
client (and httpx pool) on every single LLM call, and
_close_request_anthropic_client() always fully closed it right after
- unlike the OpenAI-wire path, which caches and reuses one warm
client across sequential calls via a single-slot cache keyed on the
effective client kwargs.
Add the same single-slot cache to the Anthropic-wire path: keyed on
credentials, base URL/Bedrock region, per-model timeout, and the
1M-beta flag; in_use guards concurrent calls from sharing one pool's
close/abort lifecycle; poisoned marks a cross-thread-aborted slot so
the owner-thread close discards it; reuse only on request_complete /
stream_request_complete (the same _REQUEST_CLIENT_REUSE_REASONS the
OpenAI path already uses). Wires a teardown hook into
release_clients()/close() mirroring _close_cached_request_openai_client.
Fixes #HPA-02
(cherry picked from commit 37f90df15593e6ded0390f827f8e5604bf0acc86)
Review follow-up (egilewski): the previous commit only hedged the guidance
text; the exact Anthropic 400 was still classified, persisted, and surfaced
as confirmed billing exhaustion. Carry the ambiguity all the way through:
- agent/error_classifier.py: 'out of extra usage' matches on the 400 and
status-less paths now attach error_context {billing_unverified,
possible_content_filter}. Reason stays FailoverReason.billing (rotation +
fallback remain the right recovery either way); ClassifiedError grows a
billing_unverified property.
- agent/credential_pool.py: new FAILURE_REASON_BILLING_UNVERIFIED. An
unverified billing exhaustion gets the short transient cooldown instead of
the one-hour bench, regardless of pool size: a content-filter rejection
leaves the credential healthy and fails identically on every key, and the
hour-long sole-credential latch is what replayed the stored error and made
real fixes look ineffective. A true 402 keeps the full bench. The marker
persists with the entry so a restart cannot upgrade it back to a bench.
- agent/agent_runtime_helpers.py + run_agent.py: recover_with_credential_pool
threads billing_unverified and hands the pool 'billing_unverified' as the
persisted failure_reason.
- agent/conversation_loop.py: the fallback-switch status, max-retries status,
terminal label, and both structured terminal results hedge when the verdict
is unverified. New _billing_terminal_label + _billing_failure_result build
the returned terminal response in one place; the result dict now carries
billing_unverified and the billing_block gains 'unverified': true. The
confirmed-billing path (a real 402 or an API-key credit depletion) keeps
the original assertive wording, so the caveat no longer dilutes it.
Regression tests: classifier marking (400 + status-less + unambiguous-body
negative), pool cooldown TTLs + persistence round-trip, pool failure_reason
plumbing, and the returned terminal response for both unverified and
confirmed verdicts.
Note: tests/agent/test_credential_pool_routing.py::TestFailureAttribution::
test_unmatched_key_does_not_retry_only_pool_entry fails identically on
current main without this change (pre-existing, unrelated).
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).
- Close partially initialized SessionDB connections on every constructor
exception path via a finally block guarded by an initialization-complete
flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
thread, close on worker exit, reject new enqueues after shutdown starts,
and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
API disconnect failures, shutdown recovery, RetainDB late enqueue, and
foreign-loop async clients.
Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A turn writing against a session already closed by compression died with
session_persistence_failed and a misleading "this is often a full disk"
dialog, even though the store was healthy and a live continuation existed
(#82001). Depth-1 recovery (find_live_compression_child) could not resolve
lineages with >=2 compression hops (root -> mid -> tip), reproduced
independently on two- and three-hop chains.
- run_agent.py flush chokepoint: on CompressionSessionClosedError, resolve
tip = db.get_compression_tip(old_id) (canonical bounded transitive walk),
adopt only when tip != old_id AND the tip row is live, retry the flush
exactly once (adoption budget); otherwise fail closed.
- gateway/session.py append_to_transcript: replace the depth-1 live-child
lookup with the same tip + liveness contract, so gateway transcript
reroutes follow full chains.
- agent/conversation_compression.py _adopt_live_compression_child: turn-start
recovery preflight now resolves via get_compression_tip with the same
liveness check, closing the last depth-1 consumer in this family.
- classify_persistence_error: new "compression_closed" bucket; the turn-end
explanation names compression rotation and tells the client to refresh the
session id instead of blaming a full disk.
Tests: depth-1 adoption, multi-hop chain adoption (agent + gateway), fail
closed with no continuation / stale-closed (ws_orphan_reap) tip, exactly-once
adoption budget, and error-wording guards (compression-closed never mentions
disk; real disk failures keep disk guidance).
Closes#82001
Co-authored-by: Al3xand3r1987 <125030427+Al3xand3r1987@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
Three follow-up fixes to the cross-process turn lease:
1. Move _clear_durable_turn_lease_interrupt() to after the refresher
thread join in the outer finally. A refresher firing between the
inner stop and the join could set an interrupt that survives the
clear and poisons the next turn on a cached agent.
2. Set self._interrupt_message in the except branch of _interrupt_turn
so _clear_durable_turn_lease_interrupt can match and clear it even
when self.interrupt() itself raised.
3. Gate the conversation_history reload behind _lease_waited so an
immediate acquisition (no contention) does not replace the
in-memory history and cause an unnecessary prompt cache miss.