Commit Graph

1306 Commits

Author SHA1 Message Date
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
kshitijk4poor 45bd48ff67 fix: initialize affinity_token before the turn-lease early returns
The turn-lease timeout/interrupt paths return from inside the try block
before set_affinity_scope() runs; the finally then read an unassigned
local -> UnboundLocalError. This was the cause of the 4 red
cross-process lease tests on PR #97158's CI.

(cherry picked from commit 05a3c8a4caa59831b01d13f3950e61385aadbfa5)
2026-09-01 02:14:35 -07:00
joaomarcos 65672e3a93 fix(cache): honor the host-declared conversation key on the affinity-key path
Every conversation-affinity hint Hermes sends is derived from the PHYSICAL
session id: prompt_cache_key on both OpenAI-wire transports, OpenRouter's and
Nous Portal's sticky session_id, and xAI's x-grok-conv-id. A host that mints
one physical session per RESPONSE re-keys all four on every reply, so the
conversation never lands back on the routing bucket it just warmed (#96811).

Two hosts do exactly that. Hermes Studio's group chat mints
gc_run_<room>_<profile>_<name>_<uuid4hex> per reply and destroys it after,
and POST /v1/responses with client-managed history mints str(uuid4()) per
request — while parsing X-Hermes-Session-Key one screen earlier and handing
it to the agent.

Hermes must not infer the logical conversation from the id's syntax: that
rule merges independent client-supplied ids and Studio members truncated past
its 96-character boundary (the #79017 failure class). It does not have to.
gateway_session_key is already the "stable per-chat key" built by
gateway.session.build_session_key from that header, and branching
deliberately does not key off it. The affinity path simply never consulted it.

- agent/prompt_cache_scope.py: declared_conversation_scope() resolves the key
  into gwk_<sha256[:24]> and outranks the lineage walk (it is stable across
  rotation AND across per-response ids). Hashed because, unlike a session id,
  the key embeds platform/chat/user identifiers and leaves the process
  verbatim as a sticky id and as x-grok-conv-id.
- agent/portal_tags.py: a separate ambient scope for ROUTING, published only
  when a host declared one. The providers read the attribution id when it is
  unset, so delegate trees keep sharing their parent's sticky key and every
  host that keeps one id per conversation is byte-identical to before.
- hermes_state.py: is_explicit_fork_child() — the public view of the marker
  rules that keep /branch children, delegate subagents and tool children off
  their parent's chat key. Background-review forks clone the live runtime, so
  _persist_disabled excludes them for the same reason (#79161).

Refs #96570
Fixes #96811
2026-09-01 02:14:35 -07:00
kshitijk4poor f20bbfa40d fix: a declined liveness abort must not cancel a pending compression
Closes the #99758 review P1 (andrexibiza): with a generation claim in
play, `interrupt()` called `_admit_hard_cancel()` BEFORE the claim was
validated at the final mutation edge, and the production
`CompressionCommitFence.cancel_before_commit()` irreversibly sets
`_cancelled = True` whenever no commit has started. So a watchdog abort
that ultimately DECLINED (real progress landed in the window, claim went
stale) had already killed the recovered turn's legitimate pending
compression commit — `begin_commit()` refuses a cancelled fence forever.
Generation authority covered interrupt publication but not the
compression-fence mutation that preceded it.

Split hard-cancel admission into two halves:

- `_wait_for_compression_commit()` runs pre-claim and is NON-mutating:
  it only blocks when `commit_in_flight` is true (the started-commit
  branch of the production fence waits for `finish_commit` without
  cancelling), so the interrupt still publishes only after an in-flight
  SessionDB mutation has finished — exactly as before.
- `_cancel_pending_compression_commit()` runs AFTER
  `_consume_claim_and_publish_first_state()` survives, so the
  destructive pending-commit cancellation can never outlive a stale
  claim. If a commit crossed its boundary in between, it is no longer
  fence-cancellable and completes on its own.

Regression coverage (both use the real `CompressionCommitFence`):

- `test_declined_abort_does_not_cancel_pending_compression_commit`:
  parks the interrupt at the claim-reservation release, lands real
  progress (G+1), lets the interrupt decline, then proves
  `fence.begin_commit()` still admits. Red on the pre-fix tree
  (mutation-checked: the fence was left cancelled).
- `test_declined_abort_parks_and_leaves_fence_operational`: the
  in-flight-commit window variant — activity lands while the interrupt
  waits on a started commit; the interrupt declines and a fresh
  `begin_commit()` still admits afterwards.
- The round-6 witness (`...resumes_inside_interrupt_publication`) now
  models an in-flight commit (`commit_in_flight = True`) so its park
  point stays inside the pre-claim wait, matching the new admission
  shape.

Also updates the `interrupt()` docstring for the deferred destructive
cancellation.
2026-09-01 03:19:59 +05:30
kshitijk4poor c394b005fc fix: publish watchdog settlement only after the abort commits
Closes the #95663 round-8 review blocker (false settlement before
commit veto): the pre-commit surface (`_surface_stall`) logged
"Force-aborting the turn and stopping lease renewal" and warned the
user "aborting it so the session can recover" BEFORE `_commit_abort`
could veto — so a turn that resumed during the warning window (or an
exceptional interrupt path that declines fail-closed) was reported as
force-aborted with lease stopped while it actually continued running.

- Split the surface: `_surface_stall` is now observational only ("no
  progress for Ns; attempting recovery"), and the definitive
  aborted/lease-stopped settlement moves to a new
  `_surface_committed_abort` that runs only after `_commit_abort`
  succeeds and the turn lease is deactivated.
- Rate-limit repeated pre-commit surfaces per observed generation: a
  turn whose aborts keep declining no longer re-logs an ERROR and
  re-warns the user every poll interval.
- Add the committed-path regression test
  (`test_watchdog_publishes_definitive_settlement_only_after_commit`)
  and extend the declined-path witness
  (`...resumes_during_warning`) to assert no committed-abort or
  definitive pre-commit claim appears when the abort is vetoed. Both
  fail on the pre-fix tree (mutation-checked).
- Document the `_interrupt_turn` lease-loss asymmetry (fires
  unconditionally, no generation claim — losing the lease means the
  process no longer owns the session).
- Trim review-round archaeology from comments/docstrings (keep the
  WHY, drop the round numbering), and drop the dead
  `cancel_event` compat note from the test fence.
- Document `agent.turn_liveness` in the configuration guide.

On top of PR #95663 by Finn763 (cherry-picked with authorship
preserved).
2026-09-01 03:19:59 +05:30
Finn763 0fe7abe37a fix(agent): surface silent turn stalls with a bounded turn-liveness watchdog (#95548, #95663)
Add a turn-liveness watchdog keyed to the agent activity clock: a turn
that stalls mid-flight while the durable lease keeps renewing is logged
loudly, surfaced to the UI, force-interrupted, and — when the hard
interrupt cannot unwind the wedge — lease renewal is stopped so
stale-turn cleanup can reclaim the session.

Race safety (rounds 3/4/6 of the #95663 review, all folded into this
squashed commit):
- AIAgent.interrupt(require_generation=G) re-validates the generation
  claim at the last instant before the hammer; a stale claim abandons
  the abort and the turn continues.
- The claim is reserved under the activity lock, invalidated by any real
  progress in _touch_activity(), and consumed immediately before the
  first observable interrupt publication; exceptional paths fail closed.
- Claim consumption and the first interrupt publication are atomic
  inside one _liveness_activity_lock() critical section; unclaimed
  interrupts publish lock-free so AIAgent stand-ins without the liveness
  seam keep working (CI 33096454629 regression, fixed here).

Deterministic race regressions (written red-first) in
tests/run_agent/test_turn_liveness_watchdog.py cover the
post-revalidation window, the consume-to-publication window, the
exceptional path, and atomic claim consumption.

Round-7 rebuild: single squashed commit on current origin/main; the
former four-commit lineage (241f8e484..299122558 on merge base
6defe7eb6c) no longer exists, so no surviving commit carries a red
exact-object CI record, and no empty CI-trigger commit was added.
2026-09-01 03:19:59 +05:30
rainbowgits 71256dfd01 fix(state): fail loudly when state.db is replaced under a live process
Detect same-inode cp via a generation stamp, halt FTS repair, and divert
unwritten transcripts to sessions/<id>.jsonl plus the gateway pending spool.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-31 14:02:55 -07:00
Kyzcreig c1bd0511cf fix(gateway): persist the platform message id on every user turn 2026-08-31 12:46:32 -07:00
fangliquanflq 53c0df6de9 fix(agent): stop compression retries after host timeout (#98722)
Salvaged from #98741, composed on top of the merged #98424 preflight
fail-closed boundary. A host-ceiling compression timeout is now a typed,
thread-safe outcome consumed by every automatic caller:

- conversation_compression.py: threading.local + per-agent lock timeout
  state (mark/reset/read helpers) upgrading #98424's simple attribute
  where overlapping automatic/manual compression entrypoints matter;
  the _last_compression_timed_out attribute stays as compat mirror.
- conversation_loop.py: the mid-turn pre-API pass and the provider
  overflow (413/400 context_length_exceeded) recovery path end the turn
  with the typed compression_exhausted recovery contract instead of
  re-sending the unchanged oversized request and re-entering compression
  in the same turn.
- run_agent.py/turn_context.py: forwarder resets the typed state per
  attempt; the #98424 turn-start check reads it through the typed helper.

Tests: thread-safety/atomicity of the state helpers, overflow-recovery
non-re-entry, and typed terminal result.
2026-08-31 12:36:02 -07:00
Dev 532ee1a199 fix(gateway): persist gateway routing identity on lazily-created session rows
When the default/global state.db is corrupt at gateway startup,
SessionStore degrades (_db=None) and record_gateway_session_peer never
self-heals. Under multiplexed profile routes the AIAgent's lazy
_ensure_db_session was then the ONLY durable write for the session, and
it created the row identity-less (session_key/chat_id/chat_type/
thread_id/user_id/origin_json all NULL) — unrecoverable by
find_latest_gateway_session_for_peer, so Telegram chats forgot prior
turns.

Persist the routing identity the agent already carries into
create_session; plain CLI sessions keep the old identity-less shape.

Salvaged leg 1 of #88804; the transcript-recovery leg is covered by the
scope-aware session DB resolution already on main.
2026-08-31 12:32:33 -07:00
anhtahaylove de49e1ef05 fix(context): fail closed when preflight compression stalls 2026-08-31 11:18:53 -07:00
fangliquanflq fd1d8271db fix(cron): isolate lazy imports from stale modules 2026-08-31 09:58:51 -07:00
Teknium 6874b99d49 fix(sessions): stamp launch-profile name on new session rows instead of NULL
Sessions created on the launch/default profile were persisted with
profile_name = NULL by all three writers (run_agent._ensure_db_session
None'd out 'default'; the desktop backend's _ensure_session_db_row and
session.branch passed None when no profile_home override was set).

NULL used to mean 'launch profile' by convention, but the desktop now
keys sessions by (profile, id), filters the sidebar by profile scope,
and resolves @session:<profile>/<id> deep links by profile match — a
NULL row matches nothing, so sessions created around a profile switch
vanished from the sidebar and their deep links could not be opened
(#99222). The #94724 one-shot legacy-owner backfill stamps literal
'default' onto old NULL rows, so writers minting NEW NULL rows after
that backfill ran recreated the exact state it exists to repair.

Stamp the real profile name at creation time in all three writers.
E2E-verified against a temp HERMES_HOME: both the desktop create path
and the agent path now persist profile_name='default'.

Fixes #99222
2026-08-31 05:56:18 -07:00
Teknium a6549922b8 fix(compression): stamp durable backoff rows with strategy and failure kind (#96775 #97488)
record_timeout_failure() now persists
'backoff:<failure_kind>:strategy=<tail_mode>' into the state.db
cooldown row (sessions.compression_failure_cooldown_until +
compression_failure_error), so a failed/stalled/cancelled attempt's
identity survives gateway restarts and the rebuilt compressor makes the
same skip decision via bind_session_state()/get_active_compression_
failure_cooldown(refresh=True). Host callers pass ceiling_exhausted /
stalled; the stall-interrupt path passes stall_interrupted. A
successful compression still clears the row.
2026-08-30 19:46:21 -07:00
fangliquanflq de4155b1bb fix(compression): report total ceiling expiry accurately 2026-08-30 19:46:21 -07:00
Stephen Chin c9b9b5e6c7 fix(gateway): preserve native compaction capability on resume 2026-08-30 05:16:10 -07:00
Stephen Chin 08c7879ca1 fix(compaction): preserve native capability across runtime switches
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
2026-08-30 05:16:10 -07:00
itsflownium 393af4a310 fix(todo): live task state via revisioned snapshots and a dedicated todo.updated event
Salvaged from PR #97815 by @itsflownium, slimmed to the schema-free core:
- TodoStore gains a monotonic in-memory revision; the todo tool result
  returns it so clients can reject stale updates
- tui_gateway emits a dedicated todo.updated full-snapshot event that
  bypasses optional tool-progress display settings
- session resume/activate responses attach the authoritative todo
  snapshot; renderer restores it with revision arbitration
- desktop store tracks per-session revisions and rejects regressions

The session_todo_state DB table from the original PR is intentionally
dropped: canonical todo tool results already persist in conversation
history, so resume paths derive the snapshot from the stored transcript
instead of a parallel store.
2026-08-29 18:40:51 -07:00
StanleyStetson 24e54b55f5 fix(desktop): preserve streamed assistant text and unify atomic persistence (#95514)
- Preserve streamed assistant text in Desktop UI when message.complete delivers empty text.

- Prevent destructive hydration in Desktop useMessageStream over rendered text on empty completion.

- Recover stream buffer in finalize_turn when final_response is empty on healthy turns.

- Unify in-place blank assistant repair, watermark clone resolution, non-blank concurrent winner adoption, and batch row appends into a single atomic guarded SessionDB transaction.

- Synchronize canonical committed content to live in-memory messages dicts and preserve all-or-nothing rollback semantics on persistence failure.
2026-08-29 11:33:05 +05:30
rob-maron f7c79efbac add tencent/hy4-preview to model pickers 2026-08-28 19:53:06 -07:00
Teknium d3a1c46510 feat(agent): context size anchors on provider-reported usage — estimation shrinks to the last turn
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.

- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
  with a structural base-message identity check that fails closed on any
  transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
  site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
  pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
  compaction (codex_runtime), session reset/switch (run_agent), plus the
  fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
  estimation fallback (first request of a session).
2026-08-28 07:51:31 -07:00
yoma 98a84783c7 fix(vision): recover from generic image content rejection 2026-08-28 04:57:58 -07:00
Teknium 420e156bf3 refactor(agent): single owner for Responses route predicates
Follow-up to the #96217 salvage: the codex/xai/github route checks were
re-implemented inline at four sites (codex_responses_adapter helpers,
chat_completion_helpers kwargs build, _is_openai_codex_backend, the
run_agent silent-reject hint). Consolidate them into
classify_responses_route() / ResponsesRouteFlags in
codex_responses_adapter and migrate every site — backend-identity
predicate class (#22548/#70893/#59561/#72468).

Host checks use exact-host-or-subdomain semantics, never substring
matching.
2026-08-27 18:56:21 -07:00
kshitijk4poor 80ab7d2b1c fix(compression): dedupe current-turn rows when rotation splits the session mid-turn
When context-compression rotation fires mid-turn, the current user
message was persisted twice into the child session. Root cause: dedup
used id()-seeded sets of copies instead of markers on the live objects.

Replace with _DB_PERSISTED_MARKER-based dedup as the sole authority:
- _ensure_compressed_has_user_turn returns CompressedUserTurnOutcome
- After publish_compression_child succeeds, stamp the live anchor-source
  row (not a drifted index) with _DB_PERSISTED_MARKER
- _sync_persisted_markers mirrors stamps from result to live lists by
  scoped identity (handles direct-path, adoption divergence, _session_messages)
- Remove _flushed_db_message_ids from rotation commit path (markers replace it)
- Unconditional (loud) imports — no silent fallback

Salvage of #94996 by @fedosis, rebased on top of #95433 (stall-fallback,
already merged). Both conversation_compression.py and run_agent.py are
built from origin/main + #94996's diff applied on top, preserving the
force_terminal refactor and _publish_new_fence from #95433.

Credit: @fedosis original PR #94996.
2026-08-28 02:47:34 +05:30
Shaun Eccles 2c6938dc3a fix(compression): retry a stalled summary on the fallback chain (#78981)
A stalled compression summary never raises, so the auxiliary client's
exception-path fallback is unreachable from it. When the progress-aware
timeout aborts a stalled worker, re-run the summary once pinned to the
first auxiliary.compression.fallback_chain entry before degrading to
continue-without-compression.

The pin is a single-use ContextVar consumed by the context compressor's
summary call, so it cannot leak into the detached stalled worker or the
compressor's own main-model retry. A fresh fence is minted through the
host factory so a /stop during the retry still admits against the live
commit boundary.
2026-08-28 02:29:27 +05:30
Teknium 4032a15ad0 refactor(prompt): remove the ~1.2K-token Nous Subscription block from the system prompt (#95005)
* refactor(prompt): remove the Nous Subscription block from the system prompt (~1.2K tokens/call)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-08-25 12:59:29 -07:00
Jan-Stefan Janetzky 1104ffe0b9 feat(memory): opt-in fail-closed pre-compress checkpoint contract (API v1)
Context compression is intentionally lossy. Deployments that archive
transcript evidence to an external durable store before compaction had no
way to guarantee the archive actually happened: MemoryManager.on_pre_compress
swallows provider failures by design, so a failed archive silently degraded
into data loss.

This adds an opt-in, provider-agnostic checkpoint contract:

- memory_provider: PRE_COMPRESS_CHECKPOINT_API_VERSION = 1; providers opt in
  by advertising pre_compress_checkpoint_api_version. Version 0 keeps the
  historical best-effort hook semantics.
- memory_manager: supports_pre_compress_checkpoint() capability probe;
  on_pre_compress(require_checkpoint=True) propagates checkpoint-provider
  failures and raises when no capable provider completed the checkpoint.
- conversation_compression: new compression.checkpoint_required config key
  (default false, documented in cli-config.yaml.example). When enabled,
  compaction fails closed with BLOCKED_MISSING_PREREQUISITE (the
  uncompressed transcript is preserved) unless a checkpoint-capable provider
  confirms the durable checkpoint. Providers receive normalized direct
  user/assistant evidence: tool rows, system messages, tool-call wrappers,
  and prior compaction summaries are filtered host-side into one stable
  contract. codex_app_server compaction is rejected under the gate because
  it exposes no truthful pre-compaction transcript boundary.
- hermes_state: persistent _compressed_summary column (declarative schema
  migration via _reconcile_columns) so summary provenance survives process
  restarts; only the resume model history carries the marker, keeping
  get_messages_as_conversation on its existing contract.
- gateway: the lossy hygiene/auto-compact paths load the memory provider
  (skip_memory=False) so a required checkpoint also guards those rewrites.

The gate arms only on an explicit boolean True (bare-MagicMock agents in
existing tests have truthy auto-attributes). Default behavior is unchanged:
checkpoint_required=false preserves best-effort semantics for all existing
providers. Contract tests, including a restart round-trip of the summary
marker, in tests/agent/test_pre_compress_checkpoint_contract.py.

Refs #93986
2026-08-25 03:55:55 -07:00
Ailirag 5908c577f9 fix(fallback): surface provider transitions and primary recovery 2026-08-25 12:12:08 +05:30
Nathan Shan d934bbd4d5 fix(agent): race Codex IPv6 and IPv4 connections
- Add RFC 8305-style staggered address attempts for synchronous ChatGPT Codex requests.
- Share the keepalive client builder across primary and auxiliary model paths.
- Cover blackholed IPv6 fallback, provider scoping, and existing proxy and TLS behavior.
2026-08-24 21:46:05 -07:00
Vignesh Ramesh a0795acc83 fix(codex): identify Hermes requests 2026-08-24 11:25:04 -07:00
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
fangliquanflq 2033f4cc34 fix(agent): separate cancellation diagnostics from tool output 2026-08-23 18:25:19 -07:00
fangliquanflq c1c0efa375 fix(code-exec): preserve interrupt cancellation source 2026-08-23 18:25:19 -07:00
Teknium 637716755c fix(cli): -Q stdout carries only the final response — no tool diffs, spinner lines, or reasoning
Widens the cherry-picked reasoning-callback fix to the whole leak class
(#93220):

- quiet branch also neutralizes tool_progress_callback,
  tool_start_callback, tool_complete_callback (inline diff rendering via
  render_edit_diff_with_delta was gated by NEITHER quiet_mode nor
  tool_progress_mode) and syncs agent.tool_progress_mode='off'.
- _should_emit_quiet_tool_messages() returns False under
  suppress_status_output: with callbacks neutralized, the quiet-mode
  KawaiiSpinner fallback printed '[tool]'/'[done]' lines into captured
  stdout. Also covers oneshot.py and background-review forks, which set
  the same flag and expect strict silence.

E2E (isolated HERMES_HOME, live model, write_file turn): base leaks
'┊ review diff' + full SVG source into stdout; head emits exactly the
final response. Regression tests pin the quiet-branch statements and the
gate (sabotage-verified).

Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
2026-08-23 17:01:31 -07:00
BrunoBza 56e7fd2adf fix(loop): bound outer-loop error retries per turn instead of relying on max_iterations (#92450)
The outer conversation-loop except handler only left the loop on a
local-processing error or when api_call_count >= max_iterations - 1.
With the turn budget now unlimited by default (sys.maxsize), a
permanent failure that escaped the inner retry/fallback machinery
retried forever: ~64 retries/s, one core pegged, and the rotated
agent.log history overwritten within minutes.

Bound the loop with a small per-turn cap on total escaping exceptions
(_MAX_OUTER_LOOP_ERRORS = 8, scaled down by a tiny explicit
max_iterations so a manually bounded budget still governs). The legacy
local-processing and near-limit exits are byte-identical; a new
'repeated_outer_errors' exit reason gets a user-facing explanation.

The inner retry/fallback layer owns transient API recovery and
terminates on its own, so only exceptions that escape it reach this
cap - a successful turn is unaffected.

Fixes #92450
2026-08-24 04:47:07 +05:30
Kshitij Kapoor 183e53656e docs: make the mutate-then-persist marker contract explicit (#92231 review)
Reviewer point on #92539: nothing documented that an in-place content
mutation of a stamped dict must pop _DB_PERSISTED_MARKER (and invalidate
the bounded flush-scan prefix) or the DB silently goes stale. Both
existing mutators (turn_finalizer fill-empty-tail, context_compressor
micro-compaction defrag) already follow the contract; this states it at
the constant so the next one does too.
2026-08-23 13:34:32 +05:30
poisdahl a5b326a471 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821
# Conflicts:
#	tests/agent/test_reference_handoff_active_turn.py
2026-08-22 16:47:39 +02:00
kshitijk4poor 8a949659c3 refactor(credits): fold review findings for stealth free-tier fix
- credits_tracker: trim inline comment block (duplicated docstring) and
  correct its safety claim - a paid model under stealth/ would fail
  closed (suppressed banner), not open; state the trade-off honestly.
- run_agent: update stale call-site comment to mention stealth/ prefix.
- auxiliary_client: widen sibling free-SKU detector _is_free_model to
  recognize stealth/ prefix (same bug class as #91843: free_only=true
  wrongly skipped the OpenRouter fallback and the paid-lane warning
  fired spuriously for stealth models).
- tests: bind the new sibling behavior (stealth/ox-alpha free,
  my-stealth/model not).
2026-08-22 04:20:48 +05:30
poisdahl 13fcf2fe38 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821 2026-08-21 16:02:59 +02:00
poisdahl abf87e7248 Merge current main into composite-carrier fix 2026-08-21 15:56:45 +02:00
kshitijk4poor b883756b79 fix: foreground priority for background review cancel timeout
Change fail-closed behavior to proceed-with-warning when a background
review does not acknowledge cancellation within the bounded deadline.
The review is non-critical self-improvement work and must never block
a user-facing turn (#84423). Keep the off-thread interrupt to ensure
a broken abort path cannot stall the bounded wait.
2026-08-21 16:12:57 +05:30
qixuancao 37da0d4d50 fix(agent): synchronize background review cancellation 2026-08-21 16:12:57 +05:30
Teknium c32119b12c feat(config): default agent.max_turns to unlimited; accept inf/infinity/null spellings
Builds on @fattchris resolve_turn_limit salvage (#67696): flips the default
from a numeric cap to unlimited across all construction paths (CLI, agent_init,
run_agent subagents), adds inf/infinity/null to the unlimited spellings, and
sets DEFAULT_CONFIG agent.max_turns to null. The turn cap caused more problems
than it solved (silent mid-task truncation).
2026-08-20 04:50:39 -07:00
Brooklyn Nicholson c57581cd0d feat(tools): drive_preview and annotate_preview — the agent can use the page it opened
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.

Four pieces, and they only make sense together:

  · an in-page engine that inventories what is interactable and performs the
    verb, injected as source because it has to run inside the guest page;
  · the preview.act.request bridge from the gateway into the pane;
  · drive_preview, for acting: elements, click, type, scroll, press, and the
    pane's own back/forward/reload;
  · annotate_preview, for marking without acting.

Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.

Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.

Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
2026-08-20 05:26:37 -05:00
Teknium 761990b780 feat: identical re-calls enter context as reference stubs, not duplicate payloads 2026-08-20 00:16:22 -07:00
joaomarcos db5d5dffea fix(agent): guard against uncompressed session overflow when compression is disabled (#89297)
When compression is explicitly disabled (compression.enabled: false), conversations can grow past the model's context window across hundreds of messages (e.g., 824 messages / 460K+ tokens in #89297). Serializing massive JSON payloads repeatedly under memory-constrained environments leads to swap thrashing (STAT=U) and unhandled provider errors.

Add a pre-flight uncompressed context overflow guardrail in build_turn_context and a deduped _warn_uncompressed_context_overflow method on AIAgent to alert users to run /compact or enable compression before unmanageable payloads freeze the process.
2026-08-20 12:35:20 +05:30
Teknium 449471c334 feat: runtime stall guards — identical-call loop breaker and continue-intent recovery (agent.stall_guards)
Composio eval traces showed Hermes wasting turns re-issuing identical tool
calls (same tool, same args, same result — 3x/4x in one run) and ending
turns by announcing an action it never took. Two conservative, config-gated
guards (agent.stall_guards, default true):

- Identical-call loop breaker: ToolCallGuardrailController.observe_identical_call
  tracks the consecutive streak of (tool, canonical args, result-hash); on
  the 3rd identical call a compact one-line notice is appended to that tool
  RESULT at construction time (cache-safe — tool results are append-only).
  Never blocks the call. Pollers (process, *_get_result, *_poll) are exempt
  via STALL_GUARD_REPEATABLE_TOOLS. Streak resets on any different call,
  changed result, or new turn. Observed on the raw result before the
  tool-loop warning suffix so its changing count can't defeat matching.

- Said-continue-but-stopped recovery: trailing_continue_intent() detects a
  short reply ENDING on an announced next action ('Let me now…', 'I will
  now…', 'Next, I…'); the conversation loop feeds it into the EXISTING
  intent-ack continuation path (same interim-assistant + user-nudge
  mechanism, same codex_ack_continuations cap of 2), preserving message
  alternation — no parallel recovery machinery.

Config: agent.stall_guards in DEFAULT_CONFIG; docs in configuration.md;
unit tests for streak/allowlist/reset/gate and detector pos/neg cases.
2026-08-19 16:34:21 -07:00
Teknium 803397ecc3 feat: wall-clock run budget — wrap-up injection at 80% and deadline-scaled stale timeouts (agent.run_budget_seconds / --run-budget) 2026-08-19 16:32:17 -07:00
kshitij kapoor b2057c1685 refactor: extract duplicated load_config_readonly try/except into helper
The identical 6-line try/except block for reading model.reasoning_echo
from config appeared in both agent_init.py (init) and
agent_runtime_helpers.py (switch_model). Extracted into
AIAgent._read_reasoning_echo_from_config() static method — net -1 LOC.
2026-08-20 00:03:50 +05:30
Yingliang Zhang 73243b0d2e feat(config): per-provider reasoning_echo opt-in for custom providers
Add model.reasoning_echo (default false) and per-fallback-entry
reasoning_echo to preserve assistant reasoning_content when
replaying history to custom providers and OpenAI-compatible gateways
that proxy thinking-mode models (Kimi K3, GLM-5.2, DeepSeek, etc.)
but are not matched by the built-in host-based _REASONING_ECHO_RULES.

The flag is per-active-provider, not a global toggle:
- Primary: read from model.reasoning_echo at init and switch_model
- Fallback: set by try_activate_fallback from the fallback entry
- Restore: restore_primary_runtime copies the switch_model snapshot

Unlike PR #76019 global agent.reasoning_echo toggle, the
per-provider flag travels with the active provider — falling back to
a strict provider (Mistral, Groq, Cerebras) correctly strips
reasoning_content even when the primary had the flag enabled,
because the flag is False for the strict fallback.

Complements PR #27361 (dynamic detection) which fires after the first
API response; this PR covers turn-1 and history-replay-on-fresh-session
where dynamic detection has not fired yet.

Closes #76018
Refs: #27297, #27361, #76019

Signed-off-by: Yingliang Zhang <zhangyingliang@outlook.com>
2026-08-20 00:03:50 +05:30