Commit Graph

1056 Commits

Author SHA1 Message Date
Teknium ddd8065232 test: regression coverage for first_chunk_at in post_api_request hook
Covers the reviewer-requested cases for #98555:
- successful streamed response emits the latest attempt's non-null
  first_chunk_at (started_at <= first_chunk_at <= ended_at)
- non-streamed, failed-stream, and partial-stream-stub paths emit None
- a stale timestamp from a prior API call cannot leak into the next
  call's post_api_request payload (per-attempt reset in the loop)
2026-09-01 08:30:45 -07:00
kshitijk4poor db339f0051 fix(state): consolidate gateway SessionDB writers via process-wide shared registry
A gateway process opened state.db from ~12 call sites, each minting its
own writer connection, self._lock, close-time WAL checkpoint, and
token-writer thread. With N independent writers on one WAL file, one
connection's close-time checkpoint could race another's growth — the
lost/reordered-page-write signature across 11+ incidents (#90837).

Adds hermes_state_registry.py: a process-wide, per-path, refcounted
shared registry owning the writer boundary.

- acquire(path): same resolved path returns the same instance (one
  writer connection, one lock, one token-writer thread) for every
  long-lived in-process caller (gateway runner, SessionStore, per-agent
  lazy recall, cron per-job, mirror, channel_directory, slash_commands,
  shutdown_flush, session_search, react_to_message, delegate, mcp_serve,
  auto_archive, tui_gateway).
- close() on a shared instance is a NO-OP — the registry owns the
  lifecycle, so one caller's close can never tear down a writer other
  callers still hold.
- Generation-aware retirement on inode change: a replaced state.db
  RETIRES the live generation (never lent again) but keeps it alive for
  existing holders; release is object-keyed so holders of the old
  generation drain it independently of the new one. The old
  generation's own write path still fails with the typed
  StateDbReplacedError (existing protection, unchanged).
- Replacement-open failure leaves NO registry entry for the path —
  the next acquire retries fresh, never hands out a closed stale object.
- All teardown runs OUTSIDE the registry lock: a final release's WAL
  checkpoint can never stall acquisition for every state.db.
- close_shared_session_dbs() at gateway shutdown drains every
  generation (live + retired) as the final safety net.

CLI one-shots, recovery flows, and read-only cross-profile opens keep
using SessionDB() directly with their own close() — only long-lived
in-process sites route through the registry.

References #90837 (root-cause tracker stays open: the #10 EOF signature
and the WAL-lifecycle A/B verdict remain under investigation there).
2026-09-01 20:55:35 +05:30
kshitijk4poor 2755d3dd27 test: give the base FakeReviewAgent the release_clients cleanup seam
The base fake still stubbed the OLD cleanup surface (shutdown_memory_provider
/ close). Production now calls release_clients(); on fakes without it the
AttributeError is swallowed by the cleanup's except-Exception, so those
tests silently stopped exercising the cleanup path. The inner recording
fake in the renamed test keeps its close() stub deliberately — it's the
tripwire proving close() is never called on the shared session.
2026-09-01 16:27:13 +05:30
konsisumer fcd5bbb3fb fix(agent): preserve foreground resources after background review 2026-09-01 16:27:13 +05:30
Yong Li f41ed09b51 fix(gemini): strip call ids on insert, name the realignments in the log
Review follow-ups:

- Strip the tool_call id when populating the call-name map so it matches the
  stripped lookup (and pass 1's result_call_ids). A padded id previously
  skipped realignment silently.
- Log which names were rewritten, not just how many.
- Note in the comment that a result whose assistant call frame was pruned is
  already dropped by the orphan pass, so it cannot reach the provider with a
  stale name; cover that with a test.
- Rename test_sanitize_leaves_matching_and_unpaired_tool_result_names_alone,
  which only ever exercised the matching case, and add the padded-id case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 14:56:30 +05:30
Yong Li 2e9435d2f6 fix(gemini): echo bridged tool_call name on the OpenAI-compatible path
Google matches functionResponse.name against functionCall.name and rejects
a mismatch with HTTP 400 INVALID_ARGUMENT. #72089 fixed this for the native
Gemini adapter, where _translate_tool_result_to_gemini() now prefers
tool_name_by_call_id over the result message's internal name.

Requests that reach Gemini through an OpenAI-compatible gateway (OpenRouter,
Vertex/LiteLLM proxies) never run that translation, so they still put the
unwrapped internal tool name on the wire: the model calls the tool_search
bridge tool `tool_call`, make_tool_result_message() labels the result
`mcp__strava__get_recent_activities`, and the next turn 400s with a bare
"Provider returned error". The bad pair stays in the transcript, so every
later request in that session fails too.

Hold the same invariant at the final pre-API chokepoint instead of in the
OpenAI-compat serializer: Gemini arrives under many model strings and base
URLs, so sniffing for "is this really Google?" is unreliable, while every
other provider either ignores the field or already agrees with the call
name. Only a name that is present and disagrees is rewritten, so clean
transcripts still pass through byte-identical for prompt caching, and the
rewrite lands on the per-call copy so the stored trajectory keeps the real
tool name for the session DB and UI. No-op for the native Gemini path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 14:56:30 +05:30
VJ Pixel 3571118218 fix(agent): run init-time fallback for ANY exhausted primary provider
The #17929 init-time fallback block was nested inside the
'_explicit not in {auto, openrouter, custom}' guard, so an exhausted
openrouter credential pool skipped fallback_providers entirely and
AIAgent.__init__ raised 'No LLM provider configured' — surfacing on
Telegram only as the generic 'unexpected error' message.

2026-08-23 outage (~22:09-23:50, 60 occurrences incl. a cron job):
single-entry openrouter pool hit daily free-tier quota; chain had a
healthy local Ollama entry that never got tried because provider was
the default 'openrouter'.

Un-nest the block so any primary without usable credentials walks the
chain before failing; providers explicitly chosen by name keep their
dedicated missing-key diagnostic when both primary AND chain fail.
Regression tests cover the openrouter-exhausted path both with and
without a usable chain entry.
2026-09-01 02:00:05 -07:00
kshitijk4poor cc0931d235 test(agent): loosen brittle error/final_response equality to substring
result['error'] and result['final_response'] are independently settable
keys that only coincidentally share _COMPRESSION_TIMEOUT_FINAL_RESPONSE
today; assert the actionable substring instead so a benign prefix or
rewording does not break the terminal-contract test.
2026-09-01 03:48:40 +05:30
kshitijk4poor 1e8f6a0491 fix(agent): surface preflight compression timeout as typed result, not generic error
When the turn-start fail-closed boundary (#98424) raises
PreflightCompressionTimedOut, the exception escaped run_conversation to
the surfaces' generic exception handlers. The gateway deliberately never
exposes raw exception text, so users saw 'Sorry, I encountered an
unexpected error... Try again or use /reset' instead of the boundary's
actionable guidance, and the compression_exhausted clean-session
recovery contract (#9893/#35809) never engaged.

Catch it at the build_turn_context callsite and convert it into the
same typed recovery dict the in-loop timeout consumers return
(salvaged #98741 / PR #99710): failed=True, partial=True,
compression_exhausted=True, turn_exit_reason=context_compression_timeout,
with the actionable message in final_response and error.

Regression test proves the exception no longer escapes and the typed
contract fields survive to the caller (mutation-checked: test fails on
main without the handler).
2026-09-01 03:48:40 +05:30
kshitijk4poor f20bbfa40d fix: a declined liveness abort must not cancel a pending compression
Closes the #99758 review P1 (andrexibiza): with a generation claim in
play, `interrupt()` called `_admit_hard_cancel()` BEFORE the claim was
validated at the final mutation edge, and the production
`CompressionCommitFence.cancel_before_commit()` irreversibly sets
`_cancelled = True` whenever no commit has started. So a watchdog abort
that ultimately DECLINED (real progress landed in the window, claim went
stale) had already killed the recovered turn's legitimate pending
compression commit — `begin_commit()` refuses a cancelled fence forever.
Generation authority covered interrupt publication but not the
compression-fence mutation that preceded it.

Split hard-cancel admission into two halves:

- `_wait_for_compression_commit()` runs pre-claim and is NON-mutating:
  it only blocks when `commit_in_flight` is true (the started-commit
  branch of the production fence waits for `finish_commit` without
  cancelling), so the interrupt still publishes only after an in-flight
  SessionDB mutation has finished — exactly as before.
- `_cancel_pending_compression_commit()` runs AFTER
  `_consume_claim_and_publish_first_state()` survives, so the
  destructive pending-commit cancellation can never outlive a stale
  claim. If a commit crossed its boundary in between, it is no longer
  fence-cancellable and completes on its own.

Regression coverage (both use the real `CompressionCommitFence`):

- `test_declined_abort_does_not_cancel_pending_compression_commit`:
  parks the interrupt at the claim-reservation release, lands real
  progress (G+1), lets the interrupt decline, then proves
  `fence.begin_commit()` still admits. Red on the pre-fix tree
  (mutation-checked: the fence was left cancelled).
- `test_declined_abort_parks_and_leaves_fence_operational`: the
  in-flight-commit window variant — activity lands while the interrupt
  waits on a started commit; the interrupt declines and a fresh
  `begin_commit()` still admits afterwards.
- The round-6 witness (`...resumes_inside_interrupt_publication`) now
  models an in-flight commit (`commit_in_flight = True`) so its park
  point stays inside the pre-claim wait, matching the new admission
  shape.

Also updates the `interrupt()` docstring for the deferred destructive
cancellation.
2026-09-01 03:19:59 +05:30
kshitijk4poor c394b005fc fix: publish watchdog settlement only after the abort commits
Closes the #95663 round-8 review blocker (false settlement before
commit veto): the pre-commit surface (`_surface_stall`) logged
"Force-aborting the turn and stopping lease renewal" and warned the
user "aborting it so the session can recover" BEFORE `_commit_abort`
could veto — so a turn that resumed during the warning window (or an
exceptional interrupt path that declines fail-closed) was reported as
force-aborted with lease stopped while it actually continued running.

- Split the surface: `_surface_stall` is now observational only ("no
  progress for Ns; attempting recovery"), and the definitive
  aborted/lease-stopped settlement moves to a new
  `_surface_committed_abort` that runs only after `_commit_abort`
  succeeds and the turn lease is deactivated.
- Rate-limit repeated pre-commit surfaces per observed generation: a
  turn whose aborts keep declining no longer re-logs an ERROR and
  re-warns the user every poll interval.
- Add the committed-path regression test
  (`test_watchdog_publishes_definitive_settlement_only_after_commit`)
  and extend the declined-path witness
  (`...resumes_during_warning`) to assert no committed-abort or
  definitive pre-commit claim appears when the abort is vetoed. Both
  fail on the pre-fix tree (mutation-checked).
- Document the `_interrupt_turn` lease-loss asymmetry (fires
  unconditionally, no generation claim — losing the lease means the
  process no longer owns the session).
- Trim review-round archaeology from comments/docstrings (keep the
  WHY, drop the round numbering), and drop the dead
  `cancel_event` compat note from the test fence.
- Document `agent.turn_liveness` in the configuration guide.

On top of PR #95663 by Finn763 (cherry-picked with authorship
preserved).
2026-09-01 03:19:59 +05:30
Finn763 0fe7abe37a fix(agent): surface silent turn stalls with a bounded turn-liveness watchdog (#95548, #95663)
Add a turn-liveness watchdog keyed to the agent activity clock: a turn
that stalls mid-flight while the durable lease keeps renewing is logged
loudly, surfaced to the UI, force-interrupted, and — when the hard
interrupt cannot unwind the wedge — lease renewal is stopped so
stale-turn cleanup can reclaim the session.

Race safety (rounds 3/4/6 of the #95663 review, all folded into this
squashed commit):
- AIAgent.interrupt(require_generation=G) re-validates the generation
  claim at the last instant before the hammer; a stale claim abandons
  the abort and the turn continues.
- The claim is reserved under the activity lock, invalidated by any real
  progress in _touch_activity(), and consumed immediately before the
  first observable interrupt publication; exceptional paths fail closed.
- Claim consumption and the first interrupt publication are atomic
  inside one _liveness_activity_lock() critical section; unclaimed
  interrupts publish lock-free so AIAgent stand-ins without the liveness
  seam keep working (CI 33096454629 regression, fixed here).

Deterministic race regressions (written red-first) in
tests/run_agent/test_turn_liveness_watchdog.py cover the
post-revalidation window, the consume-to-publication window, the
exceptional path, and atomic claim consumption.

Round-7 rebuild: single squashed commit on current origin/main; the
former four-commit lineage (241f8e484..299122558 on merge base
6defe7eb6c) no longer exists, so no surviving commit carries a red
exact-object CI record, and no empty CI-trigger commit was added.
2026-09-01 03:19:59 +05:30
rainbowgits 71256dfd01 fix(state): fail loudly when state.db is replaced under a live process
Detect same-inode cp via a generation stamp, halt FTS repair, and divert
unwritten transcripts to sessions/<id>.jsonl plus the gateway pending spool.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-31 14:02:55 -07:00
Teknium fb9b2c893f feat(agent): escalate repeated transcript-sanitiser heals with a one-time user notice (#96870)
Builds the escalation layer on top of HexLab98's heal-log windowing
(salvaged from PR #96916):

- Per-session heal counters (heal events + messages healed) tracked by the
  repair path in agent_runtime_helpers.py, session totals preserved across
  10-minute log windows.
- Threshold escalation: after N heals in a session window (default 3,
  configurable via agent.sanitizer_heal_escalation_threshold in
  config.yaml, 0 = off) log ONE ERROR carrying session id + heal pattern
  (events/messages/window/threshold), then stay quiet.
- ONE-TIME out-of-band user notice queued at the threshold and delivered by
  the conversation loop through _emit_warning (status callback -> gateway
  status message / CLI print). Never injected into conversation context or
  the wire copy: prompt caching, role alternation, and durable history are
  untouched. Never re-arms on a new window; scoped per session.
- Counters visible in diagnostics: get_sanitizer_heal_stats() rendered in
  the /debug share // hermes debug report, and the config key surfaced in
  hermes dump overrides. errors.log carries the ERROR line for `hermes logs
  errors`.
2026-08-31 13:11:41 -07:00
HexLab98 a779f527fa test(agent): cover empty-transcript projection fill and heal-log escalation (#96870) 2026-08-31 13:11:41 -07:00
fangliquanflq 53c0df6de9 fix(agent): stop compression retries after host timeout (#98722)
Salvaged from #98741, composed on top of the merged #98424 preflight
fail-closed boundary. A host-ceiling compression timeout is now a typed,
thread-safe outcome consumed by every automatic caller:

- conversation_compression.py: threading.local + per-agent lock timeout
  state (mark/reset/read helpers) upgrading #98424's simple attribute
  where overlapping automatic/manual compression entrypoints matter;
  the _last_compression_timed_out attribute stays as compat mirror.
- conversation_loop.py: the mid-turn pre-API pass and the provider
  overflow (413/400 context_length_exceeded) recovery path end the turn
  with the typed compression_exhausted recovery contract instead of
  re-sending the unchanged oversized request and re-entering compression
  in the same turn.
- run_agent.py/turn_context.py: forwarder resets the typed state per
  attempt; the #98424 turn-start check reads it through the typed helper.

Tests: thread-safety/atomicity of the state helpers, overflow-recovery
non-re-entry, and typed terminal result.
2026-08-31 12:36:02 -07:00
Teknium 7cefa87ea7 fix(agent_init): reserve Gemini's default maxOutputTokens in the compressor when max_tokens is unset
The native generateContent adapter never runs uncapped: when
model.max_tokens is unset it sends maxOutputTokens=65,535
(GEMINI_DEFAULT_MAX_OUTPUT_TOKENS) because Gemini treats an omitted cap
as a low internal default. The context compressor's trigger is
pct×(window − max_tokens), and constructing it with max_tokens=None
reserved 0 — so on a 128K Gemma window the trigger landed at 98,304
while the real safe input budget was 65,537, and the provider 400'd
before compaction fired.

Live repro (real imports, temp HERMES_HOME, native Gemini base_url,
window=131072, max_tokens unset):
  before: compressor.max_tokens=None, threshold_tokens=98304,
          wire maxOutputTokens=65535 → trigger ABOVE the safe budget
  after:  compressor.max_tokens=65535, threshold_tokens=64000 → below it

Scoped to the native Gemini wiring (provider names + native base_url via
is_native_gemini_base_url; the /openai compat endpoint is excluded). The
generic provider-default reservation gap remains tracked in #63839.

Reported by @Artemonim in #57275 (residual claim 4).
2026-08-31 12:22:55 -07:00
Teknium 1f2bd9e763 fix(compression): stamp _DB_PERSISTED_MARKER after in-place batch compaction commit (#98450)
compress() returns marker-swept copies (_strip_persistence_markers, #57491);
the in-place branch committed them via archive_and_compact() but never
stamped the persistence marker, so the next _persist_session ->
_flush_messages_to_session_db_unlocked walk re-INSERTed the whole
post-compaction transcript (live set regrew ~58K -> ~512K tokens).

Centralize the post-commit contract in a shared helper,
stamp_db_persisted_markers(), used by all three archive_and_compact
callers: the in-place batch commit (previously missing), the
micro-compaction sync, and the proactive tool-result prune.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-08-30 19:46:11 -07:00
Stephen Chin c9b9b5e6c7 fix(gateway): preserve native compaction capability on resume 2026-08-30 05:16:10 -07:00
Stephen Chin 5247a6f07f fix(compaction): clarify runtime capability state
Use a distinct runtime_capabilities field on agents, preserve compatibility with earlier snapshots, and resolve the canonical direct OpenAI endpoint when a cross-provider switch omits base_url. Keep ambiguous proxy routes fail-closed.
2026-08-30 05:16:10 -07:00
Stephen Chin 903c36b6d4 fix(compaction): resolve capability from effective switch URL 2026-08-30 05:16:10 -07:00
Stephen Chin 08c7879ca1 fix(compaction): preserve native capability across runtime switches
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
2026-08-30 05:16:10 -07:00
james47kjv 80764b6d39 fix(codex): nudge the second continuation of a compaction-only turn
gpt-5.6 on the Codex backend answers a large turn with a server-side
`compaction` checkpoint and no message. The checkpoint rides the
`codex_reasoning_items` sidecar, so the interim assistant message looks
"replayable" and `interim_replayable` suppresses the continuation nudge.

But replayable is not the same as different. A checkpoint carries no
answer and no new instruction, and a replayed checkpoint makes
`prune_pre_checkpoint_items` drop every pre-checkpoint item. Measured on
a real 262-message session: the wire collapses from 489 items to 12 —
all 186 `function_call` / `function_call_output` pairs deleted — and
ends on an empty assistant turn. The model has nothing to answer, so it
returns another empty response; the next attempt sends the same bytes
(the provider's prefix cache reports 99-100% on the repeats) and returns
the same nothing. Three attempts later the turn dies with "Codex
response remained incomplete after 3 continuation attempts" and the
whole turn's work is lost.

Keep the first continuation bare — the model often just needs another
turn, and nudging immediately would cut multi-phase work short. Once
that bare retry has also come back incomplete, it is proven not to work
for this turn, so every remaining attempt carries the nudge.
2026-08-30 05:16:02 -07:00
Eric Maddox 52637fee63 test(native-compaction): cover image-only retention with an interleaved assistant message
Folded from PR #98345 (@ericmaddox): the one scenario its suite covered
that #91557's did not — an assistant message between the image-only user
message and the checkpoint, asserting post-prune ordering.
2026-08-30 05:15:54 -07:00
Andrex Ibiza, MBA 532b2d8874 fix(native-compaction): retain image-only user content
Preserve valid normalized input_image user messages across native-compaction checkpoints at bounded one-token retention cost. Keep text extraction text-only, reject malformed or unknown multipart placeholders, and prove the production adapter path without claiming unsupported input_file behavior.

Republish the identical source tree after an unrelated nondeterministic focus-redraw test failure; this commit contains no source delta from the previously verified object.

Refs #90976 and #91477.
2026-08-30 05:15:54 -07:00
Marco Fernstaedt a2af8405d1 fix(compression): derive native threshold from local trigger 2026-08-30 05:15:37 -07:00
Teknium 514707ff3e feat(compaction): system prompt always rebuilds at the commit boundary — updates finally reach long-lived sessions (#98426)
* feat(compaction): always rebuild the system prompt at the commit boundary — keep-prompt now gated on byte equality of the LIVE builder output; plugin sections re-render with fail-open to last good bytes

* feat(clock): 'Conversation started' resolves through the session-lineage ROOT — a compacted/rotated session keeps its original birth date (Bot Mode forever-chats know when they were first born)

* test: retire old-contract pins — plugin sections re-render at invalidate (freeze stays restore-only), commit boundary always runs the live builder, byte-equal keep preserves object identity
2026-08-30 01:20:27 -07:00
Teknium 3b3ad958d7 fix(runtime): key-scoped fallback extra_body re-resolution + request_overrides in switch_model snapshot
Follow-up hardening on the two cherry-picked contributor commits:

- try_activate_fallback: replace the blanket request_overrides.pop('extra_body')
  with KEY-SCOPED removal — only keys the OLD provider's custom_providers
  entry contributed (value unchanged since the init-time merge) are dropped.
  Caller/profile-provided extra_body keys survive the swap, matching the
  caller-over-provider precedence in agent_init._merge_custom_provider_extra_body.
  The fallback provider's own extra_body is then merged back in.
- switch_model: the live _primary_runtime snapshot it rebuilds now carries
  request_overrides, so a post-switch transport recovery or fallback restore
  reinstates the switched-to identity's overrides instead of dropping them.
- Tests: activation-level stale-key removal + caller-override preservation
  (test_provider_fallback.py), switch-then-recover / switch-then-restore
  (test_primary_runtime_restore.py).

Cache-safety: none of these paths mutate past context or rebuild the system
prompt — only outbound request kwargs change.

Fixes #75091
2026-08-29 19:13:12 -07:00
Adam Fortuna 131501229a fix(runtime): restore request_overrides after transport recovery
Include request_overrides in primary runtime snapshots so transport recovery restores request-level model parameters.
2026-08-29 19:13:12 -07:00
Brin Shadewater c8bbde7770 feat: allow configured background review tools
Profiles can now grant narrowly scoped tools to the background review runtime whitelist while unrelated tools remain denied. Document the configuration and cover it with a real-config regression test.

Agent: codex
2026-08-29 19:10:06 -07:00
itsflownium 393af4a310 fix(todo): live task state via revisioned snapshots and a dedicated todo.updated event
Salvaged from PR #97815 by @itsflownium, slimmed to the schema-free core:
- TodoStore gains a monotonic in-memory revision; the todo tool result
  returns it so clients can reject stale updates
- tui_gateway emits a dedicated todo.updated full-snapshot event that
  bypasses optional tool-progress display settings
- session resume/activate responses attach the authoritative todo
  snapshot; renderer restores it with revision arbitration
- desktop store tracks per-session revisions and rejects regressions

The session_todo_state DB table from the original PR is intentionally
dropped: canonical todo tool results already persist in conversation
history, so resume paths derive the snapshot from the stored transcript
instead of a parallel store.
2026-08-29 18:40:51 -07:00
Teknium 1ee30352ca fix: background review can now read skills before patching — denial storm ended, cache parity intact (#61521, #39996)
The self-improvement review fork advertises the parent's full tool schema
(deliberate — tools[] must stay byte-identical for prompt-cache parity)
but denied everything except memory/skill tools at dispatch. Models
naturally reach for read_file to inspect a SKILL.md before patching, got
denied, then attempted a blind skill_manage patch which the
read-before-write guard correctly refused. One deployment logged ~142
denials + ~204 refusals over 2 days: the self-improvement loop ran
continuously but almost never landed a skill patch.

Fix is dispatch-side ONLY — zero request-body change, cache untouched:

- Whitelist read_file + search_files on the review fork (reads are
  side-effect-free). Write tools (write_file/patch/terminal) stay denied:
  autonomous maintenance must go through skill_manage's validation.
- read_file now registers full reads with the review fork's
  read-before-write guard (same as skill_view), so the natural
  read_file -> skill_manage(patch) sequence lands. Partial reads
  (offset>1 / truncated) don't count. No-op outside review forks.
- Self-correcting deny message: names skill_view/skill_manage/memory as
  substitutes so one denial redirects the model instead of a storm
  (the actionable half of #61521's proposal 2).

Rejects #39997's alternative (narrow the advertised schema on local
endpoints): local backends have KV/prefix caches too, and re-prefilling
a large snapshot is most expensive exactly there.

Live A/B (real dispatch path, isolated HERMES_HOME): on main,
read_file DENIED -> patch REFUSED (read-before-write); on this branch,
read_file OK -> patch LANDED. tools[] identical in both.
2026-08-29 18:33:41 -07:00
Hyusein Leshov 835a913ffd fix(compression): arm the failure cooldown when codex compaction fails
Closes #75364.

`_compress_context_via_codex_app_server` returns the transcript unchanged
when the codex thread reports `interrupted` or `error`. The session is
therefore still above threshold, and nothing records that the attempt
failed — so the next turn retries immediately, and keeps retrying for as
long as the condition persists.

Every other compression path arms the shared failure cooldown, records an
ineffective-compression strike, or both. This path records neither:

* `_hygiene_compression_failure_cooldowns` is set only on
  `asyncio.TimeoutError`, or behind `_last_compress_aborted`, which is
  assigned exclusively in `context_compressor.py` on the Hermes summarizer
  path.
* `compression_ineffective_count` lives in `ContextCompressor`, and this
  path returns before any compressor bookkeeping runs.

`compress_context` already documents the rule this path was missing —
"Every automatic entrypoint must honor compressor-owned cooldown and
breaker state" — but the codex branch dispatches above that block and
returns from inside it.

`result.interrupted` needs no unusual configuration to occur: an ordinary
user message arriving mid-compaction sets it (see
`codex_app_server_session.py`, which produces the "compact turn
interrupted" string). Observed in production on a Discord gateway session
at ~315k tokens against a 258k window, where compaction was attempted on
essentially every turn for ~70 minutes; the session's
`compression_ineffective_count` was still 0 afterwards.

This reuses the existing cooldown rather than adding a new mechanism:

* arm `_record_compression_failure_cooldown` with the existing
  `_SUMMARY_FAILURE_COOLDOWN_SECONDS` when compaction returns
  interrupted/error;
* honor an active cooldown on entry, matching the Hermes path.

`force=True` bypasses both, so an explicit /compress is never braked by a
failure it did not cause, and a successful compaction arms nothing.
2026-08-29 22:29:28 +05:30
Teknium e387cbc0aa refactor(prompt): platform-hint diet — 7 heavies compressed, −657 tok across the map (facts probe-pinned) (#97899)
* refactor(prompt): platform-hint diet — shared _MEDIA_NATIVE spine; seven heavies compressed with every verified fact intact (3,175 -> ~2,520 map total, -657)

* refactor(prompt): steer-channel note diet 225 -> 155 — marker is self-describing since its own provenance+replay clauses; prompt keeps only anti-lookalike + authority + latest-results scope (#40240/#76805 archaeology in comment)
2026-08-29 06:37:06 -07:00
Jakub Wolniewicz 23bae43cfa fix(agent): normalize list-shaped streaming content deltas 2026-08-29 12:45:43 +05:30
Teknium 217ab2f8df refactor(desktop-tools): consolidate preview + project, diet the desktop_ui suite (3,861 → 2,293 tok/call, −41%) (#97659)
* refactor(desktop-tools): consolidate preview(open/close/read) + project(create/switch/list), diet the desktop_ui suite — 3,861 -> 2,293 tok/call on desktop sessions (-41%)

* rename: preview -> desktop_preview, project -> desktop_project — namespace desktop-app tools against MCP/plugin name collisions

* test: sync remaining old-name pins — per-file registration import, GUI_TOOLS set, post-hook case read_preview -> desktop_preview action=read
2026-08-28 23:10:01 -07:00
StanleyStetson 24e54b55f5 fix(desktop): preserve streamed assistant text and unify atomic persistence (#95514)
- Preserve streamed assistant text in Desktop UI when message.complete delivers empty text.

- Prevent destructive hydration in Desktop useMessageStream over rendered text on empty completion.

- Recover stream buffer in finalize_turn when final_response is empty on healthy turns.

- Unify in-place blank assistant repair, watermark clone resolution, non-blank concurrent winner adoption, and batch row appends into a single atomic guarded SessionDB transaction.

- Synchronize canonical committed content to live in-memory messages dicts and preserve all-or-nothing rollback semantics on persistence failure.
2026-08-29 11:33:05 +05:30
Teknium 306db2776c test(codex): mid-turn compaction fixtures report realistic anchored usage
The usage anchor (#97206) now trusts provider-reported usage. These two
tests simulated a tool-heavy near-overflow turn while the shared fixture
reported a 12-token prompt — the anchored pressure check honestly
concluded no pressure. Give the scenario 18K anchored prompt tokens so
the tests pin the same compaction decision they always did.
2026-08-28 07:51:31 -07:00
Teknium d3a1c46510 feat(agent): context size anchors on provider-reported usage — estimation shrinks to the last turn
Every provider response carries usage.prompt_tokens — exact ground truth
for the full request (system prompt + tool schemas + history). Context-size
checks now anchor on the last main-loop response's usage and estimate only
the messages appended since, instead of re-estimating the whole history
with chars/4 heuristics and flat 1500-token image costs. The estimate error
window shrinks from the entire conversation to one turn and self-corrects
at every response.

- agent/model_metadata.py: capture_usage_anchor() / anchored_context_tokens()
  with a structural base-message identity check that fails closed on any
  transcript rewrite.
- agent/conversation_loop.py: anchor captured at the single main-loop usage
  site (MoA uses pre-fold aggregator usage; advisor/aux calls never anchor);
  pre-API pressure check prefers the anchor.
- agent/turn_context.py: preflight compression estimate prefers the anchor.
- agent/context_breakdown.py: /context display prefers the anchor.
- Invalidation: compaction rewrite (conversation_compression), codex native
  compaction (codex_runtime), session reset/switch (run_agent), plus the
  fail-closed structural check for splices/micro-compaction.
- Usage-less responses keep the previous anchor; no anchor -> pure
  estimation fallback (first request of a session).
2026-08-28 07:51:31 -07:00
fedebyes 93f4dc7561 fix: make positional prune variant-aware; add replayed-call regression tests
Pass 2 of repair_message_sequence matched results only by id/call_id,
pruning calls answered through response_item_id or composite bridge
ids. Use the shared variant helpers (tool_call_id_variants /
tool_result_id_variants) so the unified alias policy applies
(#55626/#63000/#93251).

The positional sanitizer pass changes the crash/resume duplicate shape:
an interrupted first occurrence is now stubbed instead of deduped, so
the replayed call survives with its own immediate result. Update the
#64335 empty-key test to the new semantics and add regression tests for
the #94704 acceptance shape (historical-result + replayed-call +
fresh-call) and the production interrupted-turn shape (session
7d57a602b83d).
2026-08-28 07:51:23 -07:00
Tiberiu Danciu c7761573f5 fix: prune positionally unanswered tool_calls before API send
DeepSeek v4 rejects a payload where an assistant message carries a
tool_call whose tool result does not follow it immediately (HTTP 400
"An assistant message with 'tool_calls' must be followed by tool
messages responding to each 'tool_call_id'"). Context compression can
displace a tool result past a user turn; the result then lands ~100
messages away from its declaring assistant message.

Two gaps let the poisoned shape reach the wire (reproduced from the
production request dump of session 4d8727cbcf04, replayed through both
functions):

1. repair_message_sequence Pass 1 drops the displaced tool RESULT as
   stray but leaves the declaring assistant message carrying the now
   unanswered tool_call (with empty content) in the durable history.
2. sanitize_api_messages stubbed only globally-absent result ids: the
   displaced result still exists in the transcript, so the id survives
   the set-subtraction, no stub is injected, and the payload 400s.

Fix both layers so every path is order-independent:

- repair_message_sequence: new Pass 2 prunes tool_calls that have no
  result in the immediately-following tool run (matching on id or
  call_id, same superset rule as Pass 1). If pruning empties the turn
  (no content/reasoning left), the whole message is dropped rather than
  sending an empty assistant message. Codex interim turns are exempt,
  as in Pass 0.
- sanitize_api_messages: the orphan/stub logic is rewritten as a single
  rolling positional walk that drops results not immediately following
  their declaring assistant (including results appearing BEFORE their
  call) and injects stub results for positionally-uncovered calls even
  when a mispositioned result exists elsewhere.

Adds six regression tests: repair pruning, whole-turn drop when pruned
calls were the only payload, valid-pair negative control, positional
stub injection, result-before-call orphan drop, and a fully-paired
transcript negative control.
2026-08-28 07:51:23 -07:00
isheng 87cff9d4c1 test(sanitizer): add unit tests for _classify_tool_call_orphans 2026-08-28 06:32:48 -07:00
joaomarcos f0ac2c8f12 fix(agent): drop stale api_content sidecar and unpaired tool results
Rebased onto current main to drop the empty-tool_calls fix (already on
main via #86654, cherry-picked from #77944 with @webtecnica's
authorship). This PR now carries only the two fixes unique to it:

1. A pre-existing api_content sidecar left stale on the consecutive-
   assistant merge. The sidecar takes priority over content at
   API-build time, so a merge could silently discard its own freshly
   concatenated content on the next call. Only dropped when the merge
   actually changes the resulting value (wz-heng, #78063 review) --
   content_rewritten compares before/after value, not just whether an
   assignment branch fired, so a falsy new_content (e.g. "") that
   strips to nothing no longer trips a spurious sidecar drop.

2. sanitize_api_messages never flagged a tool result with a missing/
   empty tool_call_id -- its orphan-detection set only ever collected
   truthy ids, so an unpaired result with no id passed the final
   chokepoint untouched.

Addresses teknium1's rebase request and wz-heng's review findings on
2026-08-28 06:32:48 -07:00
Hakan Baysal cb8027afed fix(sanitize): preserve assistant messages with tool_calls when stripping images
_strip_images_from_messages() deleted any non-tool message whose content
became empty after image removal. An assistant message whose content was
entirely images but which carried tool_calls was therefore dropped,
orphaning its paired tool responses — providers reject the next request
with unmatched tool_call_id errors (HTTP 400). Replace such messages
with the plaintext placeholder instead, exactly like tool-role messages.

Adds a regression test covering the assistant + tool_calls +
image-only-content case.

Closes #40463
2026-08-28 05:17:26 -07:00
Frowtek e1762bd30b fix(agent): drop the api_content sidecar when stripping images from history
`api_content` is the byte-stability sidecar from #67274: it holds the exact
bytes previously sent for a message, and every turn substitutes it back into
`content` when building `api_messages`. `drop_stale_api_content` exists so a
content rewrite cannot be replayed from it — its own docstring states the
contract, and names the historical image strip as one of the callers:

    Replaying the pre-rewrite sidecar would resend exactly what the rewrite
    removed, so it must be dropped — the cost is one cache boundary miss,
    never wrong content.

`_strip_images_from_messages` never drops it. The image-rejection recovery in
`conversation_loop` runs it over the persistent history, not just the per-call
copy:

    agent._vision_supported = False
    _imgs_removed = _strip_images_from_messages(messages)      # history
    if isinstance(api_messages, list):
        _strip_images_from_messages(api_messages)

and `api_messages` are copies (`api_msg = msg.copy()`), so the history message
keeps its sidecar. The strip is therefore undone on the very next turn.

Reproduced with the real functions:

    history content after strip : [{'type': 'text', 'text': 'look'}]
    sidecar still present       : True
    NEXT TURN sends             : 'look<IMAGE BYTES SENT LAST TURN>'

This is worse than a one-turn glitch, because the recovery cannot fire again:
it is gated on `getattr(agent, "_vision_supported", True)` and just set that
False. So on every subsequent turn the sidecar re-injects the images, the
text-only endpoint rejects them again, and the branch that would strip them is
disabled — the session stays wedged on a 4xx it already knew how to fix.

Drop the sidecar on each message the strip rewrites, inside the function so
every caller is covered. Messages with no images keep theirs, so only the
rewritten message pays a cache boundary — the tradeoff the invariant
prescribes. The two sibling recovery paths, `_sanitize_messages_surrogates`
and `_sanitize_messages_non_ascii`, are already safe: both walk every string
field on the message and so scrub the sidecar in passing. This one only
touches `content`.

tests/run_agent/test_image_rejection_fallback.py: new
TestStripImagesDropsStaleApiContent — the rewritten message loses its sidecar,
the next turn does not resend the stripped images, the tool-placeholder rewrite
is covered too, and untouched messages keep their sidecar. All four fail on
main. 53 passed across the image-rejection and api_content-sidecar suites; 307
passed across the sanitization/image/sidecar/replay agent tests (8 failures in
test_image_routing.py / test_save_url_image.py are pre-existing and fail
identically on clean main).
2026-08-28 05:17:21 -07:00
Koduri Mahesh Bhushan Chowdary b3f4f50771 fix(agent): classify "media exceeds size limit" as image_too_large
MiniMax's Anthropic-compatible endpoint rejects an oversized native image
part with "media exceeds size limit: max 10485760 bytes (2013)" — no
occurrence of the word "image", so none of _IMAGE_TOO_LARGE_PATTERNS
matched. The 400 fell through to _REQUEST_VALIDATION_PATTERNS (the body
is type: invalid_request_error) and classified as format_error /
non-retryable.

That skipped the image-shrink recovery in conversation_loop, which is
gated on FailoverReason.image_too_large. Because the oversized part is
already baked into history as a tool_result image block, and the context
compressor rewrites text but not image data, every later turn re-sent the
same bytes and failed identically — the session stayed dead until the
user forked it.

Match on the "media" fragment, mirroring the existing "image exceeds"
entry so reworded vendor variants are caught too. A non-image media
rejection routed here is safe: the shrink pass finds no image parts,
returns False, and the caller surfaces the original error unchanged.

Fixes #76039
2026-08-28 04:58:06 -07:00
yoma 98a84783c7 fix(vision): recover from generic image content rejection 2026-08-28 04:57:58 -07:00
Al Cooke 0241619068 fix: retry text-only on Codex invalid image data errors
Treat the ChatGPT Codex invalid image-data 400 as an image rejection so Hermes strips image parts and retries text-only instead of aborting the session. Add coverage for the exact error wording.
2026-08-28 03:46:24 -07:00
fkdls112 cd72689e03 fix(agent): strip images on Kimi/Moonshot 'failed to decode image' 400
Truncated or corrupt image bytes baked into immutable conversation history
get re-sent on every retry. Kimi/Moonshot reject them with HTTP 400
'prepare image failed ... failed to decode image: invalid or unsupported
image format', which was missing from _IMAGE_REJECTION_PHRASES, so the
turn exhausted retries and wedged the session instead of stripping the
images and recovering.

Adds the phrase to the recovery list plus a regression test mirroring the
exact Kimi error body. Complements PR #76896 (proactive full-decode
validation in vision_tools) with reactive recovery for already-poisoned
sessions. Fixes #76884.
2026-08-28 03:46:24 -07:00
Sora-bluesky 564113572e fix(agent): classify xAI's downloaded-response wording as a corrupt image
The observed wire error — 'Downloaded response does not contain a valid JPG, PNG, WebP, or ICO image.' — has no match in _IMAGE_CORRUPT_PATTERNS, so it falls through to the non-retryable 400 handler and the session replays the same image parts into the same 400 until /new.

Adds the full observed sentence to the pattern list. Deliberately not the shorter prefixes: a bare 'downloaded response does not contain a valid' also matches non-image 400s, and a negative test now pins that a downloaded-response certificate 400 keeps falling through to the existing handler. Covered on both the 400 path and the message-only path.

Reported by ryuhaneul in #69078.
2026-08-28 03:46:24 -07:00