Commit Graph

1330 Commits

Author SHA1 Message Date
Jan-Stefan Janetzky 1104ffe0b9 feat(memory): opt-in fail-closed pre-compress checkpoint contract (API v1)
Context compression is intentionally lossy. Deployments that archive
transcript evidence to an external durable store before compaction had no
way to guarantee the archive actually happened: MemoryManager.on_pre_compress
swallows provider failures by design, so a failed archive silently degraded
into data loss.

This adds an opt-in, provider-agnostic checkpoint contract:

- memory_provider: PRE_COMPRESS_CHECKPOINT_API_VERSION = 1; providers opt in
  by advertising pre_compress_checkpoint_api_version. Version 0 keeps the
  historical best-effort hook semantics.
- memory_manager: supports_pre_compress_checkpoint() capability probe;
  on_pre_compress(require_checkpoint=True) propagates checkpoint-provider
  failures and raises when no capable provider completed the checkpoint.
- conversation_compression: new compression.checkpoint_required config key
  (default false, documented in cli-config.yaml.example). When enabled,
  compaction fails closed with BLOCKED_MISSING_PREREQUISITE (the
  uncompressed transcript is preserved) unless a checkpoint-capable provider
  confirms the durable checkpoint. Providers receive normalized direct
  user/assistant evidence: tool rows, system messages, tool-call wrappers,
  and prior compaction summaries are filtered host-side into one stable
  contract. codex_app_server compaction is rejected under the gate because
  it exposes no truthful pre-compaction transcript boundary.
- hermes_state: persistent _compressed_summary column (declarative schema
  migration via _reconcile_columns) so summary provenance survives process
  restarts; only the resume model history carries the marker, keeping
  get_messages_as_conversation on its existing contract.
- gateway: the lossy hygiene/auto-compact paths load the memory provider
  (skip_memory=False) so a required checkpoint also guards those rewrites.

The gate arms only on an explicit boolean True (bare-MagicMock agents in
existing tests have truthy auto-attributes). Default behavior is unchanged:
checkpoint_required=false preserves best-effort semantics for all existing
providers. Contract tests, including a restart round-trip of the summary
marker, in tests/agent/test_pre_compress_checkpoint_contract.py.

Refs #93986
2026-08-25 03:55:55 -07:00
Ailirag 5908c577f9 fix(fallback): surface provider transitions and primary recovery 2026-08-25 12:12:08 +05:30
Nathan Shan d934bbd4d5 fix(agent): race Codex IPv6 and IPv4 connections
- Add RFC 8305-style staggered address attempts for synchronous ChatGPT Codex requests.
- Share the keepalive client builder across primary and auxiliary model paths.
- Cover blackholed IPv6 fallback, provider scoping, and existing proxy and TLS behavior.
2026-08-24 21:46:05 -07:00
Vignesh Ramesh a0795acc83 fix(codex): identify Hermes requests 2026-08-24 11:25:04 -07:00
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
fangliquanflq 2033f4cc34 fix(agent): separate cancellation diagnostics from tool output 2026-08-23 18:25:19 -07:00
fangliquanflq c1c0efa375 fix(code-exec): preserve interrupt cancellation source 2026-08-23 18:25:19 -07:00
Teknium 637716755c fix(cli): -Q stdout carries only the final response — no tool diffs, spinner lines, or reasoning
Widens the cherry-picked reasoning-callback fix to the whole leak class
(#93220):

- quiet branch also neutralizes tool_progress_callback,
  tool_start_callback, tool_complete_callback (inline diff rendering via
  render_edit_diff_with_delta was gated by NEITHER quiet_mode nor
  tool_progress_mode) and syncs agent.tool_progress_mode='off'.
- _should_emit_quiet_tool_messages() returns False under
  suppress_status_output: with callbacks neutralized, the quiet-mode
  KawaiiSpinner fallback printed '[tool]'/'[done]' lines into captured
  stdout. Also covers oneshot.py and background-review forks, which set
  the same flag and expect strict silence.

E2E (isolated HERMES_HOME, live model, write_file turn): base leaks
'┊ review diff' + full SVG source into stdout; head emits exactly the
final response. Regression tests pin the quiet-branch statements and the
gate (sabotage-verified).

Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
2026-08-23 17:01:31 -07:00
BrunoBza 56e7fd2adf fix(loop): bound outer-loop error retries per turn instead of relying on max_iterations (#92450)
The outer conversation-loop except handler only left the loop on a
local-processing error or when api_call_count >= max_iterations - 1.
With the turn budget now unlimited by default (sys.maxsize), a
permanent failure that escaped the inner retry/fallback machinery
retried forever: ~64 retries/s, one core pegged, and the rotated
agent.log history overwritten within minutes.

Bound the loop with a small per-turn cap on total escaping exceptions
(_MAX_OUTER_LOOP_ERRORS = 8, scaled down by a tiny explicit
max_iterations so a manually bounded budget still governs). The legacy
local-processing and near-limit exits are byte-identical; a new
'repeated_outer_errors' exit reason gets a user-facing explanation.

The inner retry/fallback layer owns transient API recovery and
terminates on its own, so only exceptions that escape it reach this
cap - a successful turn is unaffected.

Fixes #92450
2026-08-24 04:47:07 +05:30
Kshitij Kapoor 183e53656e docs: make the mutate-then-persist marker contract explicit (#92231 review)
Reviewer point on #92539: nothing documented that an in-place content
mutation of a stamped dict must pop _DB_PERSISTED_MARKER (and invalidate
the bounded flush-scan prefix) or the DB silently goes stale. Both
existing mutators (turn_finalizer fill-empty-tail, context_compressor
micro-compaction defrag) already follow the contract; this states it at
the constant so the next one does too.
2026-08-23 13:34:32 +05:30
poisdahl a5b326a471 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821
# Conflicts:
#	tests/agent/test_reference_handoff_active_turn.py
2026-08-22 16:47:39 +02:00
kshitijk4poor 8a949659c3 refactor(credits): fold review findings for stealth free-tier fix
- credits_tracker: trim inline comment block (duplicated docstring) and
  correct its safety claim - a paid model under stealth/ would fail
  closed (suppressed banner), not open; state the trade-off honestly.
- run_agent: update stale call-site comment to mention stealth/ prefix.
- auxiliary_client: widen sibling free-SKU detector _is_free_model to
  recognize stealth/ prefix (same bug class as #91843: free_only=true
  wrongly skipped the OpenRouter fallback and the paid-lane warning
  fired spuriously for stealth models).
- tests: bind the new sibling behavior (stealth/ox-alpha free,
  my-stealth/model not).
2026-08-22 04:20:48 +05:30
poisdahl 13fcf2fe38 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821 2026-08-21 16:02:59 +02:00
poisdahl abf87e7248 Merge current main into composite-carrier fix 2026-08-21 15:56:45 +02:00
kshitijk4poor b883756b79 fix: foreground priority for background review cancel timeout
Change fail-closed behavior to proceed-with-warning when a background
review does not acknowledge cancellation within the bounded deadline.
The review is non-critical self-improvement work and must never block
a user-facing turn (#84423). Keep the off-thread interrupt to ensure
a broken abort path cannot stall the bounded wait.
2026-08-21 16:12:57 +05:30
qixuancao 37da0d4d50 fix(agent): synchronize background review cancellation 2026-08-21 16:12:57 +05:30
Teknium c32119b12c feat(config): default agent.max_turns to unlimited; accept inf/infinity/null spellings
Builds on @fattchris resolve_turn_limit salvage (#67696): flips the default
from a numeric cap to unlimited across all construction paths (CLI, agent_init,
run_agent subagents), adds inf/infinity/null to the unlimited spellings, and
sets DEFAULT_CONFIG agent.max_turns to null. The turn cap caused more problems
than it solved (silent mid-task truncation).
2026-08-20 04:50:39 -07:00
Brooklyn Nicholson c57581cd0d feat(tools): drive_preview and annotate_preview — the agent can use the page it opened
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.

Four pieces, and they only make sense together:

  · an in-page engine that inventories what is interactable and performs the
    verb, injected as source because it has to run inside the guest page;
  · the preview.act.request bridge from the gateway into the pane;
  · drive_preview, for acting: elements, click, type, scroll, press, and the
    pane's own back/forward/reload;
  · annotate_preview, for marking without acting.

Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.

Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.

Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
2026-08-20 05:26:37 -05:00
Teknium 761990b780 feat: identical re-calls enter context as reference stubs, not duplicate payloads 2026-08-20 00:16:22 -07:00
joaomarcos db5d5dffea fix(agent): guard against uncompressed session overflow when compression is disabled (#89297)
When compression is explicitly disabled (compression.enabled: false), conversations can grow past the model's context window across hundreds of messages (e.g., 824 messages / 460K+ tokens in #89297). Serializing massive JSON payloads repeatedly under memory-constrained environments leads to swap thrashing (STAT=U) and unhandled provider errors.

Add a pre-flight uncompressed context overflow guardrail in build_turn_context and a deduped _warn_uncompressed_context_overflow method on AIAgent to alert users to run /compact or enable compression before unmanageable payloads freeze the process.
2026-08-20 12:35:20 +05:30
Teknium 449471c334 feat: runtime stall guards — identical-call loop breaker and continue-intent recovery (agent.stall_guards)
Composio eval traces showed Hermes wasting turns re-issuing identical tool
calls (same tool, same args, same result — 3x/4x in one run) and ending
turns by announcing an action it never took. Two conservative, config-gated
guards (agent.stall_guards, default true):

- Identical-call loop breaker: ToolCallGuardrailController.observe_identical_call
  tracks the consecutive streak of (tool, canonical args, result-hash); on
  the 3rd identical call a compact one-line notice is appended to that tool
  RESULT at construction time (cache-safe — tool results are append-only).
  Never blocks the call. Pollers (process, *_get_result, *_poll) are exempt
  via STALL_GUARD_REPEATABLE_TOOLS. Streak resets on any different call,
  changed result, or new turn. Observed on the raw result before the
  tool-loop warning suffix so its changing count can't defeat matching.

- Said-continue-but-stopped recovery: trailing_continue_intent() detects a
  short reply ENDING on an announced next action ('Let me now…', 'I will
  now…', 'Next, I…'); the conversation loop feeds it into the EXISTING
  intent-ack continuation path (same interim-assistant + user-nudge
  mechanism, same codex_ack_continuations cap of 2), preserving message
  alternation — no parallel recovery machinery.

Config: agent.stall_guards in DEFAULT_CONFIG; docs in configuration.md;
unit tests for streak/allowlist/reset/gate and detector pos/neg cases.
2026-08-19 16:34:21 -07:00
Teknium 803397ecc3 feat: wall-clock run budget — wrap-up injection at 80% and deadline-scaled stale timeouts (agent.run_budget_seconds / --run-budget) 2026-08-19 16:32:17 -07:00
kshitij kapoor b2057c1685 refactor: extract duplicated load_config_readonly try/except into helper
The identical 6-line try/except block for reading model.reasoning_echo
from config appeared in both agent_init.py (init) and
agent_runtime_helpers.py (switch_model). Extracted into
AIAgent._read_reasoning_echo_from_config() static method — net -1 LOC.
2026-08-20 00:03:50 +05:30
Yingliang Zhang 73243b0d2e feat(config): per-provider reasoning_echo opt-in for custom providers
Add model.reasoning_echo (default false) and per-fallback-entry
reasoning_echo to preserve assistant reasoning_content when
replaying history to custom providers and OpenAI-compatible gateways
that proxy thinking-mode models (Kimi K3, GLM-5.2, DeepSeek, etc.)
but are not matched by the built-in host-based _REASONING_ECHO_RULES.

The flag is per-active-provider, not a global toggle:
- Primary: read from model.reasoning_echo at init and switch_model
- Fallback: set by try_activate_fallback from the fallback entry
- Restore: restore_primary_runtime copies the switch_model snapshot

Unlike PR #76019 global agent.reasoning_echo toggle, the
per-provider flag travels with the active provider — falling back to
a strict provider (Mistral, Groq, Cerebras) correctly strips
reasoning_content even when the primary had the flag enabled,
because the flag is False for the strict fallback.

Complements PR #27361 (dynamic detection) which fires after the first
API response; this PR covers turn-1 and history-replay-on-fresh-session
where dynamic detection has not fired yet.

Closes #76018
Refs: #27297, #27361, #76019

Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
2026-08-20 00:03:50 +05:30
Brooklyn Nicholson 23d88c2b0e feat(tools): tour — let the agent walk a user through the UI
One generic tool in the desktop_ui toolset: discover what is on screen,
highlight an element with narration, or hand the user a paged tour. No tour
content lives in the code — the agent authors each one live, which is what
makes 'how does this work?' answerable as a walkthrough instead of a wall
of text.

Rides the existing blocking-prompt bridge (tour.request/.respond) like
read_preview, so it works on every connection topology.
2026-08-19 00:52:54 -05:00
kshitij 979ca57a50 Merge pull request #88244 from kshitijk4poor/fix/handoff-cleanup-race
fix: prevent handoff leg data loss + surface state.db corruption to users
2026-08-17 17:41:44 +05:30
kshitij 59f302fef9 fix: prevent handoff leg data loss + surface state.db corruption to users
Two data-loss bugs reported by users:

1. /handoff CLI→gateway race (#88234): After /handoff completed, CLI
   cleanup called finalize_session on the session the gateway just
   reopened. This set end_reason on a row the gateway was actively
   writing to, causing the handoff leg to vanish from session history
   and breaking session_search recall. Fix: add _handed_off_session_ids
   module-level set (mirrors _single_query_finalize_attempted_session_ids
   pattern). _handle_handoff_command registers the session_id on
   completion; _should_emit_cleanup_session_finalize and
   _emit_interrupted_session_end check it before firing.

2. state.db corruption silent failure (#88235): When SessionDB init
   failed at gateway startup, the error stayed in logs — messages
   flowed but nothing was persisted, with no user-visible indication.
   Fix: store _session_db_init_error on GatewayRunner, broadcast a
   recovery-guidance message to all home channels via
   _send_session_db_warning_notifications() after the gateway connects.
   Also improved the 'corrupt' persistence cause wording in
   _format_turn_completion_explanation to include the full recovery
   path (hermes doctor --fix, sqlite3 .recover, backups).

Tests: 6 new tests for handoff cleanup race, 3 for corruption wording.
All existing CLI/turn-completion tests pass.
2026-08-17 13:19:58 +05:30
Teknium bd4b709258 Inspired by Copilot CLI: /rollback keeps user hand-edits by default
Copilot CLI 1.0.78 reworked /rewind to restore only the files the agent
changed, 'skipping any file whose contents no longer match what Copilot
last wrote'. This ports that protection to Hermes checkpoints:

- tools/checkpoint_manager.py: per-project agent-write ledger
  (sha256 of every landed write_file/patch), safe_restore_plan()
  classifier, and restore(safe=True) that reverts only agent-authored
  changes, deletes agent-created files, and preserves user hand-edits.
  Empty ledger (pre-existing stores) falls back to the classic full
  restore.
- run_agent.py: feed the ledger from _record_file_mutation_result on
  every landed mutation (zero new hooks; rides the existing verifier).
- CLI + gateway /rollback: safe mode is the default; --all/--force
  restores everything; skipped files are reported with a hint.
- 17 locales: new gateway.rollback.kept_user_edits key.
- Docs: checkpoints-and-rollback.md updated.
- Tests: 7 new cases incl. user-edit preservation, post-agent user
  tweaks, agent-created file removal, empty-ledger fallback.
2026-08-16 22:08:47 -07:00
Teknium 250232ff91 Revert "fix(agent): harden canonical tool call deduplication"
This reverts commit 8fc4189edd.
2026-08-16 10:51:11 -07:00
Teknium 587ad8748f Revert "fix(agent): preserve local reasoning timeout opt-out"
This reverts commit 26b2b47593.
2026-08-16 10:51:11 -07:00
Teknium 06b9141109 fix(state): classify structural DB corruption as its own persistence cause
'database disk image is malformed' contains the word 'disk', so
classify_persistence_error bucketed SQLITE_CORRUPT / SQLITE_NOTADB
failures as 'disk' and the turn-completion explainer told users to
free disk space for a structurally damaged state.db (the #77386-family
misdiagnosis, reproduced in the v0.20.0 malformed-DB incident report).

- hermes_state: new 'corrupt' bucket in PERSISTENCE_ERROR_CAUSES,
  matched via _DB_CORRUPTION_MARKERS BEFORE the locked/disk buckets
- run_agent: explainer text for 'corrupt' points at hermes doctor and
  explicitly says freeing space will not help
- cron explainer-variant suppression picks the new variant up
  automatically (it iterates PERSISTENCE_ERROR_CAUSES)
2026-08-16 10:34:23 -07:00
Teknium d709d29f19 fix(agent): trim background_review to the enabled switch
Follow-up to #87400: drop the max_iterations and prompt_file knobs from
auxiliary.background_review. The aux model routing (provider/model/
base_url/...) predates #87400 and stays; the enabled switch and the
usage telemetry stay. The fork's iteration budget returns to the
historical hardcoded 16.
2026-08-16 10:27:52 -07:00
Ojas Sharma 7095e23eb2 fix(agent): attribute background-review usage and add cost controls
Persist fork token usage under session_model_usage task=background_review,
emit a per-fork completion log line, and expose enabled/max_iterations/
prompt_file so operators can see and bound the automatic review cost.

Address review feedback: load auxiliary.background_review once per spawn,
classify completion logs by summarize action prefixes, treat explicit
api_call_count=None as the documented default of 1, and WARNING on the
fail-open enabled-gate path.
2026-08-16 06:38:38 -07:00
poisdahl 7ca1987459 Merge upstream main into PR 81234 2026-08-16 12:20:50 +02:00
fangliquanflq 8fc4189edd fix(agent): harden canonical tool call deduplication 2026-08-16 01:59:08 -07:00
fangliquanflq a55d29d9fb fix(agent): canonicalize duplicate tool call arguments 2026-08-16 01:59:08 -07:00
fangliquanflq 26b2b47593 fix(agent): preserve local reasoning timeout opt-out 2026-08-16 01:53:19 -07:00
poisdahl fcea2175e2 Merge remote-tracking branch 'origin/main' into codex/81234-live-main-final
# Conflicts:
#	tests/test_tui_gateway_server.py
#	tui_gateway/methods_prompt.py
#	tui_gateway/methods_tools.py
2026-08-15 23:02:54 +02:00
fangliquanflq 79b7d969d3 fix(gateway): reject unstamped durable ordinal rewinds 2026-08-15 13:36:25 -07:00
poisdahl d65b7d0760 Merge remote-tracking branch 'origin/main' into codex/81234-live-main-final 2026-08-15 12:05:28 +02:00
poisdahl 3f075d41dd Merge remote-tracking branch 'origin/main' into codex/81234-live-main-final
# Conflicts:
#	tui_gateway/methods_prompt.py
#	tui_gateway/server.py
2026-08-15 12:05:14 +02:00
PRATHAMESH75 ece3bc7d52 fix(agent): skip the automatic background review inside delegation subagents
The post-turn background review fork (`agent/background_review.py`) inherits
the parent agent's live runtime by default. That is a cost win when the parent
IS the main chat model (warm prompt cache, cheap), but the fork also fires
inside delegation subagents, where it inherits the *subagent's* model. When a
subagent runs a premium delegation model, the review silently replays the whole
conversation and emits skill/memory-update output at premium rates, with
nothing in the log or config flagging it (#85859).

Subagents are already barred from writing shared MEMORY.md
(`DELEGATE_BLOCKED_TOOLS`) and are spawned with `skip_memory=True`, so an
automatic review here has little to persist. Guard `_spawn_background_review`
(the single choke point both the turn-finalizer and codex-runtime callers pass
through) to return early when `_delegate_depth > 0`. An explicit `/refine`
(`focus` set) is a deliberate user request and still runs; the top-level path
is unchanged.

Fixes #85859
2026-08-15 02:44:38 -07:00
spfcraze e25cafc83e perf(agent): cache per-turn display-flag config reads
_file_mutation_verifier_enabled and _turn_completion_explainer_enabled
re-read config.yaml on every call via load_config() (~1ms deepcopy per
call). finalize_turn runs these gates at the end of every turn, so each
turn paid two redundant config deepcopies. The sibling
_credits_notices_enabled already caches on self; mirror that pattern.

The env-var override stays authoritative and uncached, so runtime flips
still work. Config flips now apply on the next session, matching the
documented sibling semantics.

(cherry picked from commit f2d0e00d71ba35fa3a53cd41076261e3c5feb45d)
2026-08-15 00:36:03 -07:00
joaomarcos 2162d583b1 perf(run-agent): reuse the Anthropic request-local client instead of rebuilding it per call
_create_request_anthropic_client() built a fresh anthropic.Anthropic
client (and httpx pool) on every single LLM call, and
_close_request_anthropic_client() always fully closed it right after
- unlike the OpenAI-wire path, which caches and reuses one warm
client across sequential calls via a single-slot cache keyed on the
effective client kwargs.

Add the same single-slot cache to the Anthropic-wire path: keyed on
credentials, base URL/Bedrock region, per-model timeout, and the
1M-beta flag; in_use guards concurrent calls from sharing one pool's
close/abort lifecycle; poisoned marks a cross-thread-aborted slot so
the owner-thread close discards it; reuse only on request_complete /
stream_request_complete (the same _REQUEST_CLIENT_REUSE_REASONS the
OpenAI path already uses). Wires a teardown hook into
release_clients()/close() mirroring _close_cached_request_openai_client.

Fixes #HPA-02

(cherry picked from commit 37f90df15593e6ded0390f827f8e5604bf0acc86)
2026-08-15 00:34:29 -07:00
Teknium 4b7b2b0049 fix: widen base-URL hostname identity class to remaining substring sites
Follow-up to #85737, which migrated five provider-identity sites onto
utils.base_url_host_matches()/base_url_hostname(). This completes the class
sweep (never-patch-predicates: one owner, every site) and folds in the two
open contributor PRs attacking individual sites:

- agent/auxiliary_client.py ZAI/Kimi OpenAI-wire rewrite (PR #85715,
  pierrenode): 'bigmodel'/'api.z.ai'/'api.kimi.com' substring checks
  rewrote proxy paths containing those markers.
- hermes_cli/runtime_provider.py Azure endpoint detection (PR #74721,
  RelaxJonh, issue #74312): 'azure.com' substring picked the Azure key
  for non-Azure hosts whose path contained the text.
- run_agent.py: _is_azure_openai_url, _is_copilot_url, Anthropic
  credential-refresh azure guard, _anthropic_preserve_dots host
  allowlist, OpenRouter/mistral reasoning gates.
- agent/chat_completion_helpers.py: nousresearch / nvidia detection.
- agent/conversation_loop.py: GitHub Models 413 hint.
- agent/usage_pricing.py: localhost billing-route detection.
- hermes_cli/model_switch.py: api.openai.com catalog fallback and
  localhost custom-provider detection.
- cli.py: local-model autodetect and Ollama/LM Studio context-length
  hints (port-anchored instead of '11434' in URL).
- tools/mcp_oauth.py: Figma remote-MCP detection.
- tools/skills_hub.py: raw.githubusercontent.com source-URL check.

Regression tests extend tests/hermes_cli/test_base_url_host_identity.py
(azure/copilot/dotted-model/figma proxy-path + lookalike cases) and
tests/agent/test_minimax_auxiliary_url.py (ZAI/Kimi path false positives).

Closes #74312. Salvages #85715 and #74721 with authorship preserved.
2026-08-14 22:04:16 -07:00
Jack Lau f57209bc9f fix(agent): carry the ambiguity of Anthropic's 'out of extra usage' 400 through classification, cooldown, and terminal surfaces
Review follow-up (egilewski): the previous commit only hedged the guidance
text; the exact Anthropic 400 was still classified, persisted, and surfaced
as confirmed billing exhaustion. Carry the ambiguity all the way through:

- agent/error_classifier.py: 'out of extra usage' matches on the 400 and
  status-less paths now attach error_context {billing_unverified,
  possible_content_filter}. Reason stays FailoverReason.billing (rotation +
  fallback remain the right recovery either way); ClassifiedError grows a
  billing_unverified property.

- agent/credential_pool.py: new FAILURE_REASON_BILLING_UNVERIFIED. An
  unverified billing exhaustion gets the short transient cooldown instead of
  the one-hour bench, regardless of pool size: a content-filter rejection
  leaves the credential healthy and fails identically on every key, and the
  hour-long sole-credential latch is what replayed the stored error and made
  real fixes look ineffective. A true 402 keeps the full bench. The marker
  persists with the entry so a restart cannot upgrade it back to a bench.

- agent/agent_runtime_helpers.py + run_agent.py: recover_with_credential_pool
  threads billing_unverified and hands the pool 'billing_unverified' as the
  persisted failure_reason.

- agent/conversation_loop.py: the fallback-switch status, max-retries status,
  terminal label, and both structured terminal results hedge when the verdict
  is unverified. New _billing_terminal_label + _billing_failure_result build
  the returned terminal response in one place; the result dict now carries
  billing_unverified and the billing_block gains 'unverified': true. The
  confirmed-billing path (a real 402 or an API-key credit depletion) keeps
  the original assertive wording, so the caveat no longer dilutes it.

Regression tests: classifier marking (400 + status-less + unambiguous-body
negative), pool cooldown TTLs + persistence round-trip, pool failure_reason
plumbing, and the returned terminal response for both unverified and
confirmed verdicts.

Note: tests/agent/test_credential_pool_routing.py::TestFailureAttribution::
test_unmatched_key_does_not_retry_only_pool_entry fails identically on
current main without this change (pre-existing, unrelated).
2026-08-14 21:54:56 -07:00
joaomarcos 39e480c051 fix(state): close leaked SessionDB connections on exception paths (#83226)
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).

- Close partially initialized SessionDB connections on every constructor
  exception path via a finally block guarded by an initialization-complete
  flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
  CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
  thread, close on worker exit, reject new enqueues after shutdown starts,
  and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
  API disconnect failures, shutdown recovery, RetainDB late enqueue, and
  foreign-loop async clients.

Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-14 21:41:26 -07:00
Teknium 8dc5608a78 fix(compression): adopt live continuation tip at flush across multi-hop chains
A turn writing against a session already closed by compression died with
session_persistence_failed and a misleading "this is often a full disk"
dialog, even though the store was healthy and a live continuation existed
(#82001). Depth-1 recovery (find_live_compression_child) could not resolve
lineages with >=2 compression hops (root -> mid -> tip), reproduced
independently on two- and three-hop chains.

- run_agent.py flush chokepoint: on CompressionSessionClosedError, resolve
  tip = db.get_compression_tip(old_id) (canonical bounded transitive walk),
  adopt only when tip != old_id AND the tip row is live, retry the flush
  exactly once (adoption budget); otherwise fail closed.
- gateway/session.py append_to_transcript: replace the depth-1 live-child
  lookup with the same tip + liveness contract, so gateway transcript
  reroutes follow full chains.
- agent/conversation_compression.py _adopt_live_compression_child: turn-start
  recovery preflight now resolves via get_compression_tip with the same
  liveness check, closing the last depth-1 consumer in this family.
- classify_persistence_error: new "compression_closed" bucket; the turn-end
  explanation names compression rotation and tells the client to refresh the
  session id instead of blaming a full disk.

Tests: depth-1 adoption, multi-hop chain adoption (agent + gateway), fail
closed with no continuation / stale-closed (ws_orphan_reap) tip, exactly-once
adoption budget, and error-wording guards (compression-closed never mentions
disk; real disk failures keep disk guidance).

Closes #82001

Co-authored-by: Al3xand3r1987 <125030427+Al3xand3r1987@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
2026-08-14 21:39:44 -07:00
burak33bb 453e6d8b95 fix(moa): preserve facade across client rebuilds 2026-08-14 21:36:41 -07:00
kshitij d729019e46 fix(sessions): harden turn lease refresher and skip unnecessary reloads
Three follow-up fixes to the cross-process turn lease:

1. Move _clear_durable_turn_lease_interrupt() to after the refresher
   thread join in the outer finally. A refresher firing between the
   inner stop and the join could set an interrupt that survives the
   clear and poisons the next turn on a cached agent.

2. Set self._interrupt_message in the except branch of _interrupt_turn
   so _clear_durable_turn_lease_interrupt can match and clear it even
   when self.interrupt() itself raised.

3. Gate the conversation_history reload behind _lease_waited so an
   immediate acquisition (no contention) does not replace the
   in-memory history and cause an unnecessary prompt cache miss.
2026-08-15 03:22:16 +05:30