Commit Graph

4613 Commits

Author SHA1 Message Date
KeyArgo 43ecad8fe6 fix(agent): classify Novita 'server overload' 429 as overloaded
Novita returns HTTP 429 with message 'server overload, please try again
later' and error type 'server_overload' when its server is genuinely busy.
Neither phrase was in _OVERLOADED_PATTERNS, so the 429 fell through to
_V_RATE_LIMIT and set should_fallback=True + should_rotate_credential=True —
rotating the credential / falling back early instead of retrying the same
key. Add 'server overload' and 'server_overload' to the overload tuple so
this reaches the existing _V_OVERLOADED verdict (retryable, no rotation).
Closes #106205
2026-09-09 10:47:49 -07:00
teknium1 8d93081971 fix(desktop): stop flagging local/LAN auxiliary pins as stale
An aux task pinned to a private endpoint via `base_url` (a home Ollama
box at `byron.local`, a LAN IP, localhost) is the intended per-task
endpoint feature and can never bill a provider. The Settings → Model
banner still counted it as "still run on openai" forever and offered
"Reset all to main", which would wipe the working local setup; the
post-switch `stale_aux` report had the same blind spot; and the aux row
never showed the `base_url` the backend already sends, so the pin was
indistinguishable from a paid-provider pin.

- `GET /api/model/auxiliary` now stamps each task with `local_endpoint`,
  the verdict of the one canonical classifier
  (`agent/model_metadata.py::is_local_endpoint`) — no TS mirror of the
  private-range rules, so frontend and runtime cannot drift.
- Desktop: the persistent banner filter is the pure
  `staleAuxAssignments()` and skips `local_endpoint` pins; the pinned row
  appends ` · <base_url>` when one is set.
- `_stale_aux_pins` (post-switch report) skips local pins the same way.
- `is_local_endpoint`: `*.local` (RFC 6762 mDNS) now counts as local, and
  IPv6 literals no longer ride the "no dots ⇒ unqualified host" rule, so
  a global-scope address (`2607:f8b0::1`) is not local while `::1`,
  ULA and link-local still are via the `ipaddress` scope checks.

Slim redo of #106236 (@webtecnica) and #106234 (@huklaa), which fixed the
same symptom with a client-side classifier copy; the bug class, row
display and mDNS/IPv6 classifier corrections are theirs.

Refs #106228

Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-09 10:33:00 -07:00
Teknium 69c999f8a7 fix(agent): keep the tripping correction and explain the restart-bound exit
When the redirect cap trips, the correction that cancelled the final attempt is
still sitting in _pending_redirect; finalize_turn's clear_interrupt() would drop
it silently. Drain it into the steer slot so it rides result["pending_steer"] and
becomes the next user turn on every surface that already honours that key.

Both new exit reasons get a turn-completion explanation so the user sees why the
turn stopped instead of an empty reply.
2026-09-09 09:51:31 -07:00
yoyodine-industries e9312da68b fix(agent): bound redirect/rebuilt restart refunds so a runaway turn can't hold the session lease
The redirect and rebuilt-for-fallback restart paths in apply_retry_restarts
refund the iteration budget and re-issue the iteration with no per-turn
bound. A redirect/interrupt that keeps re-arming the flag refunds forever,
so the turn loop never exits and the durable session turn lease is held
indefinitely (concurrent processes block up to LEASE_WAIT_SECONDS).

Add a per-turn restart_count accumulator (threaded through _run_phase like
the other loop locals) and break out once it exceeds max_retries, matching
the bound the compression path already has.
2026-09-09 09:51:31 -07:00
gaoanze888 8af248042c fix(compression): preserve batch clarify answers in summarize pass
_sum_clarify only extracted the top-level ``user_response`` key, so batch clarify
results (questions=[...] -> responses[].user_response) fell through to the generic
placeholder and the summarizer never saw the user's answer/permission decision.

Closes #106077.
2026-09-09 09:21:55 -07:00
kshitijk4poor 9e6c4100cb fix(agent): close the interrupted tool tail on the overflow terminal
The overflow-terminal path ends the turn without reaching finalize_turn, so
a transcript that overflowed right after a tool batch ended on a raw tool
result; strict providers reject the next user turn (tool -> user). Close it
with the same final text, mirroring the truncated-tool-call terminal above.

Also: classify once before either log so an overflow no longer emits a
"so the loop can continue" WARNING followed by the contradicting "NOT
seeding" one; reset the stale-streak breaker once for both branches; drop
the "compression could not recover it" wording (this path never reached
compression); trim the test file to the three tests that bind behaviour
(stream -> terminal stub; 413 stays non-terminal; the terminal ends the
turn, closes the tool tail, carries compression_exhausted).
2026-09-09 17:45:12 +05:30
ca-shrimp 3b0e81459b fix(agent): carry compression_exhausted bit; scope overflow terminal to context_overflow
Address review P1s (andrexibiza) on #106266:

1. The overflow-terminal exit in recover_from_truncation now forwards the
   #98722 typed compression_exhausted bit (partial_result/end_turn gained the
   flag) so the gateway resets/moves future input to a clean session instead
   of leaving the bloated durable session authoritative for the next turn.

2. _overflow_terminal is scoped to FailoverReason.context_overflow ONLY.
   payload_too_large (413) has its own byte-scored recovery owner
   (turn_overflow._recover_payload_too_large, #88960/#47339) that must not be
   bypassed; a post-delta 413 keeps its normal continuation stub. Regression
   covers both lanes.

Tests: unit asserts result compression_exhausted=True on the marker; a
real streamed partial hitting a 413 payload-too-large error keeps content and
is not terminal. 50 streaming/continuation/gateway regressions pass.
2026-09-09 17:45:12 +05:30
ca-shrimp ece584b8f5 fix(agent): don't seed continuation stub after a context-overflow stream death
When a stream delivered text and then died on a context-overflow /
payload-too-large error, the partial content (often tens of KB) was seeded as
a length-continuation stub, growing the transcript monotonically. In a
session whose transcript cannot be compressed back under budget
(protect_last_n covers everything -> no_progress, or the summary would
itself be larger -> would_grow), every later request is larger than the one
that just failed — an unrecoverable loop where the user sees a 30+ minute
fake hang and the only remedy is killing the session (#106260).

classify_api_error already labels these errors context_overflow /
payload_too_large (should_compress=True). _partial_stream_stub now returns
an EMPTY stub marked _overflow_terminal for that class instead of seeding
the recovered text, and recover_from_truncation treats the marker as
terminal: the turn ends via the recovery contract with a clear message
(start /new) and the transcript is not polluted with the partial.

Normal partials (network stall, output-cap truncation, tool-call drops) are
unchanged — only the overflow error class changes behavior.

Tests: stub marker + empty content; a real streamed partial hitting a
'maximum context length' error returns the terminal stub; recover_from_
truncation ends the turn (no fragment/nudge appended) on the marker while a
normal stub still runs the continuation path. 67 streaming/continuation
regressions pass.
2026-09-09 17:45:12 +05:30
nftpoetrist 511633be90 fix(agent): stop the run-budget wrap-up notice from mutating a persisted tool row
_maybe_inject_run_budget_wrapup() appends its wrap-up notice to the newest
role:"tool" message in place, with no _DB_PERSISTED_MARKER check. Its sibling,
_maybe_inject_iteration_budget_warning(), got exactly this guard added in the
same recent saga (turn_iteration_prep.py), with the comment "an older turn may
already be cached."

The reachability is structural, not an edge case: _maybe_inject_run_budget_wrapup
is only ever called from prepare_iteration(), at the START of the next iteration
-- strictly after tool_executor.py's _flush_session_db_after_tool_progress has
already flushed and marked the previous iteration's tool row persisted. So every
successful injection was mutating an already-persisted row: the wire request for
that turn carried the notice, but the durable transcript never did, diverging
replay from the live bytes and invalidating the provider's prompt-cache prefix
from that row onward.

Fix:
- Add the same _DB_PERSISTED_MARKER guard to _maybe_inject_run_budget_wrapup,
  scoped to the specific tool row the reversed scan lands on (not just
  messages[-1], since this function -- unlike its sibling -- scans backward for
  the newest tool row rather than only checking the tail).
- Wire _maybe_inject_run_budget_wrapup into _flush_session_db_after_tool_progress
  (pre-flush), mirroring exactly how _maybe_inject_iteration_budget_warning is
  wired in both places. Without this, the guard alone would make the notice stop
  firing in the common case, since prepare_iteration's call site almost always
  hits an already-persisted row -- the pre-flush call site is what actually lets
  it land in durable bytes.

Verified empirically: read the real call graph (tool_executor.py's three
_flush_session_db_after_tool_progress call sites cover every tool-completion
path) to confirm the guard's premise, then added an end-to-end test using a real
AIAgent + SessionDB that flushes and checks the persisted row for the notice
text. Mutation-verified: reverting the two production files drops exactly the 2
new/updated assertions (28 pass, 2 fail); reapplying restores green (30 passed).
Also ran the sibling iteration-budget-warning and /steer suites (71 passed) to
check for interaction regressions -- none.
2026-09-09 17:04:30 +05:30
Teknium e1838c5b5a fix: trim opencode-go 422 salvage to the invariant set
Drop the bare "tool.content" pattern (any 400 mentioning tool.content in a
non-list context would be sent through the image-strip path) and the profile
flag snapshot test; the behaviour tests (classifier verdict + proactive
downgrade) already pin the contract.
2026-09-09 03:52:47 -07:00
ericmaddox bee840bc8c fix(providers,agent): handle strict-string tool message validation and 422 on opencode-go (fixes #104731)
- Declare `supports_vision_tool_messages=False` and `supports_vision=True` on `opencode_go` provider profile in `plugins/model-providers/opencode-zen/__init__.py`
- Route HTTP 422 errors through `_IMAGE_TOOL_RULES` and add `tool.content.str`, `tool.content`, and `input should be a valid string` patterns to `_MULTIMODAL_TOOL_CONTENT_PATTERNS` in `agent/error_classifier.py`
- Add unit tests for OpenCode Go proactive tool result downgrade, HTTP 422 Console Go classification, and profile capability contract in `tests/run_agent/test_multimodal_tool_content_recovery.py` and `tests/plugins/model_providers/test_opencode_go_profile.py`
2026-09-09 03:52:47 -07:00
Teknium 06dc51d62d fix: verification evidence ledger is inert while verify_on_stop is off
The ledger in verification_evidence.db exists only to feed the verify-on-stop
guard, but the recorder kept running on every foreground terminal command and
every file edit after #53552 turned the guard off by default. Users who never
opted in still accumulated a multi-MB database (7 MB / 4.6k rows on one install).

Every ledger entry point (record_terminal_result, record_verify_run,
mark_workspace_edited, verification_status) now checks verify_on_stop_enabled()
first and returns without opening or creating the database when the guard is
off. verification_status reports {"status": "disabled"} in that case; no client
consumes the verification.status RPC yet, so nothing downstream changes.

Existing ledger tests pin HERMES_VERIFY_ON_STOP=1 since they exercise the ledger
itself; the new test proves the off path never creates the file (red on base).
2026-09-09 02:35:41 -07:00
kshitijk4poor 26f4a674e0 fix(agent): a /steer row is human input for every user-turn predicate
Follow-up to #106317. Typing the steer row (display_kind="steer") for the renderer and the
alternation-repair guard collided with the convention that any display_kind on a user row means
scaffolding: is_user_originated_turn / _is_actionable_user_turn / split_user_originated_turn
returned False for it (tail anchoring, auto-focus, dispatcher views, resume counts) while
_is_real_user_message returned True (anchor restoration) — the two predicate families disagreed
on the same row, and list_recent_user_messages (/undo, /rewind) skipped it in SQL. A steer
carries full user authority; the steer kind is now whitelisted in all four.

Also: the pre-API drain's requeue tail reuses _requeue_pending_steer instead of a copy; the TUI
history projection compares against STEER_DISPLAY_KIND; the steer() docstring describes the row.
2026-09-09 13:08:25 +05:30
kshitijk4poor 91433c8466 fix(loop): the turn-boundary export skips preflight-timeout envelopes and stops re-anchoring the persist index
Follow-up to #106312. _preflight_timeout_result carries the prior history without this turn's
user row (#7100); with a repeated prompt ("continue") the verbatim scan resolved to the
historical copy and exported it as this turn's proven boundary — the exact relabeling the export
exists to prevent. Nothing is exported for that envelope now.

The trailing `agent._persist_user_message_idx = idx` ran after finalize_turn had already flushed
the transcript, so it never influenced a persist and the next turn reset it: dead state, removed.
2026-09-09 12:55:43 +05:30
kshitijk4poor 7dc796463d fix(agent): a persisted /steer row survives the next prompt's alternation repair; typed for history
Both steer sites now build the row through one helper, prompt_builder.steer_user_row:
a role:user row with display_kind="steer" and no leading blank lines. The alternation
repair (_merge_consecutive_users) skips a steer-typed prev row, so a run that ended
right after a steered batch (Ctrl-C, interrupt) does not get the next real prompt
merged INTO the already-persisted steer row — which would have rewritten it in place
and re-broken live≠replay parity, the exact class this PR fixes.

TUI/desktop history projects the steer row as the user's own words instead of the
model-facing marker wrapper; 'steer' joins the display_kind union. The compression
anchor scan keeps its tool-row branch for transcripts persisted before this change and
its docstring says so.
2026-09-09 12:21:28 +05:30
kshitijk4poor 4d0cec9a7d fix(agent): the pre-API-call /steer drain also stops smearing the persisted tool row
Second site of the same bug class #104444 fixes in apply_pending_steer_to_tool_results:
_inject_steer_into_newest_tool_result (the drain that runs when a /steer lands during an
API call) mutated the newest role:tool row in place. That row was already flushed
append-only, so the replayed history diverged from the live request bytes at the
injection point and broke the prompt cache exactly like the post-batch path.

Deliver it the same way: a standalone user row inserted right after the newest tool
result (not yet persisted, so the next flush writes it to the transcript). Restash when
there is no tool row yet, unchanged. Stale comments claiming steer lands "in the newest
tool result" and agent/AGENTS.md's alternation rule now describe the real shape.
2026-09-09 12:21:28 +05:30
Albert.Zhou d24810483d fix(agent): persist /steer as a standalone user message
`apply_pending_steer_to_tool_results` used to smear the steer text onto
the last `role:tool` message's content. That tool row had already been
flushed to the session store and carries `_DB_PERSISTED_MARKER`; the
append-only persistence never rewrites it, so the replayable transcript
diverged from the live request bytes at the injection point — resumed
sessions (surface switch / process restart / background-review close)
missed the provider prompt cache (75-85% hit) and the user's mid-run
instructions were never part of the durable history.

The steer is now emitted as a standalone `role:user` message (marker
text preserved):
- role alternation stays legal: assistant(tool_calls) -> tool -> user is
  the documented 'user jumped in mid-run' pattern that
  `repair_message_sequence` deliberately keeps;
- the appended dict carries no `_DB_PERSISTED_MARKER`, so the next
  `_flush_messages_to_session_db` writes it to the session store —
  transcript bytes and replayed history finally agree, and the steer
  becomes searchable/retrievable like any other user message;
- the no-tool-result fallback (interrupt) still requeues the steer, which
  the caller then delivers as a normal next-turn user message.

Tests: TestSteerInjection updated for the new shape plus a persistability
assertion (no marker => flushable); tool-batch-segmentation malformed
scenario updated. steer + segmentation suites: 67 passed, 1 skipped.
2026-09-09 12:21:28 +05:30
kshitijk4poor 1f6718b0ff test(agent): prove the re-anchor through prepare_iteration; reuse the compaction _reanchor
The salvaged regression test exercised only repair_message_sequence and
reanchor_current_turn_user_idx — pre-existing helpers — so reverting the fix left it
green. It now drives prepare_iteration on a real AIAgent with adjacent user rows and
asserts the returned index addresses this turn's row and mirrors into
_persist_user_message_idx (red without the re-anchor: IndexError).

Both re-anchor sites (repair and compression restart) call
turn_context_compaction._reanchor instead of inlining "reanchor + mirror", so they
cannot drift. The export tests fold into one parametrized invariant plus the
run_conversation envelope test; the WHAT-restating comment shrinks to the WHY.
2026-09-09 12:20:04 +05:30
Felipe Portavales 37f42713ef feat(loop): export {turn_id, current_turn_user_idx} on every result envelope
Hosts that settle their own transcript by index (hermes-webui) cannot prove which
row of result["messages"] is the current user turn once this loop rewrote history
(alternation repair, compaction, post-turn micro-compaction): the instance-side
_persist_user_message_idx predates those rewrites, and a text match relabels an
identical historical prompt and claims its old answer. Only the producer can
assert the coordinate against the exact list it returns.

run_conversation now wraps the turn (_run_conversation_turn) and stamps the pair
through export_current_turn_boundary on every envelope that leaves the loop
(success, partial/error, interrupt, retry-exhausted, tool-limit, preflight
timeout, codex runtime), computed on the final messages after finalize_turn and
micro-compaction. The pair is exported only when the addressed row is this turn's
user message verbatim (reanchor's last-match rule); a rewritten row exports
nothing so hosts fail closed. The final index is mirrored into
_persist_user_message_idx for the persist override.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013gp366ijf39n4UUtJhZuMh
2026-09-09 12:20:04 +05:30
Felipe Portavales fe21a4d2f0 fix(loop): re-anchor current_turn_user_idx after the alternation repair merges rows
prepare_iteration() runs repair_message_sequence_with_cursor() before each API
call; the repair merges adjacent user rows in place (after a compaction, the
role=user summary sits next to the protected first user message). The loop's
current_turn_user_idx was recorded at turn start, so after a merge it points
past the current user row: the per-turn context injection (prefetch/plugin
context) silently misses it, and hosts that settle the transcript by this index
(hermes-webui) write the current user turn to the FRONT of the context —
rewriting the prompt's leading messages every turn (0% prefix-cache hits at
200K+ tokens, ~100 s re-prefill per turn) and duplicating the user's question.

The in-loop compression restart path already re-anchors; do the same after a
repair that changed the list: reanchor_current_turn_user_idx (last user row
carrying this turn's text), return the index through the IterationPrep verdict
so the loop state picks it up, and mirror it into agent._persist_user_message_idx,
which hosts read when the result carries no index. The new phase parameters
default to None so direct callers keep their signature.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013gp366ijf39n4UUtJhZuMh
2026-09-09 12:20:04 +05:30
kshitijk4poor e333113871 fix(review): keep /refine under the background_review origin; attendedness is its own flag
The salvaged commit forked an explicit /refine under a new "refine_review" origin so
the memory delete gate would not treat it as unattended. But is_background_review()
is the key for every other review guard — skill_manager_guards (curator-owned-only,
read-before-write), skill_manager_tool (archive instead of rmtree), skill_ledger
actor, write_approval staging, the [auto] tag — so a /refine fork silently escaped
all of them.

Carry attendedness separately: the fork keeps origin "background_review" and sets
_review_attended; turn_context binds it beside the origin ContextVar; the memory
gate keys on the new is_unattended_review(). Also run the gate AFTER
_validate_single_op / the operations list check, as memory_tool's own docstring
requires, so an invalid replace is rejected now rather than staged and failed at
approve time.
2026-09-09 12:19:13 +05:30
liuhao1024 c0714575c3 fix(review): distinguish explicit /refine from unattended reviews and surface staged consolidations
Review follow-up on #105944 (#105921):

- explicit /refine forks now run under the refine_review write origin
  (explicit flows from the CLI/gateway handlers through
  _spawn_background_review_now and spawn_background_review_thread down
  to build_cache_parity_fork), so a user-requested review keeps the
  full memory operation set; only automatic reviews stay behind the
  unattended delete gate.
- the unattended delete gate now stages the denied replace/remove (or
  whole batch) into the pending store instead of dropping it: the
  fork's own review summary is never published, so a plain denial lost
  the consolidation request with no surfacing path. The staged proposal
  carries a proposal_staged marker that summarize surfaces as an action
  line, and a staging failure still fails closed to a plain denial.
- regression tests: explicit-path origin pass-through, refine_review
  keeping replace working, near-limit denial end to end (add rejected
  by budget -> replace staged -> proposal surfaces, store unchanged).
2026-09-09 12:19:13 +05:30
liuhao1024 1571f502a9 fix(agent): scope background review memory access to its trigger (#105921)
The review fork's tool whitelist granted the whole memory toolset
whenever the profile had memory enabled, regardless of which nudge
fired, so a skill-nudge fork held remove/replace on MEMORY.md it was
never asked to use; combined with the memory tool's near-limit
'consolidate now' hint, an unattended fork deleted standing rules with
no user in the loop.

- Pass review_memory from spawn_background_review_thread through
  _run_review_in_thread/_run_review_fork into _review_tool_whitelist;
  a skill-only review no longer gets the memory tool at all.
- Fail-closed operation gate in memory_tool: a background-review fork
  may add, never replace/remove (single or in a batch) — consolidation
  decisions reach a human via the review summary instead.
- Keep the deny/prompt wording in sync with the whitelist so a
  memory-less review doesn't advertise memory.
2026-09-09 12:19:13 +05:30
kshitijk4poor 2369606f98 refactor(agent): one _requeue for the three heappush sites; timing test asserts ordering, not a 0.3 s bound
Fold the identical heappush(...) into PeriodicScheduler._requeue; notify() instead of
notify_all() now that the scheduler thread is the only condition waiter; drop the
PR-history paragraph from the module docstring (the commit carries it).

Tests: the blocked-sibling test asserted `sibling_ran.wait(0.30)` — a wall-clock bound
under the repo's ≥2 s flake floor; it now asserts the sibling fired while the blocker
still held its worker. The worker-start-failure fake keys on this scheduler's own
_run_callback rather than the global thread-name prefix so a leaked handle on _DEFAULT
cannot consume the single injected failure. The base-green no-overlap test is dropped
(it does not prove the fix).
2026-09-09 12:17:57 +05:30
finn763 1562b87d5d fix(agent): isolate periodic scheduler callbacks from blocking siblings (#102574) 2026-09-09 12:17:57 +05:30
Tim Kaufmann b7e4712029 fix(agent): classify local-inference memory-ceiling rejections as overloaded
oMLX/MLX prefill memory-guard rejections name an allocation peak in BYTES but
close with "Reduce context length", so _CONTEXT_OVERFLOW_PATTERNS claims them
and the turn enters the compress-and-shrink loop. Compression cannot lower a
prefill peak — the prompt is usually far below the window — so it burns the
compression budget, re-hits the wedged server on every attempt and ends in
"Cannot compress further" plus a destructive session reset.

Classify them as `overloaded` instead: retry with backoff, no compression, no
session reset (mirrors 503/529 recovery).

The guard runs before the overflow check AND before the usage-limit
disambiguation. The second ordering matters more than it looks: "memory limit
exceeded" contains "limit exceeded", so a status-less rejection — a proxy that
flattened the body — is currently classified as `billing` and rotates a
healthy credential.

Sites covered:
  - _OVERFLOW_AS_5XX_RULES, which _400_TAIL_RULES extends → 400, 500, 502,
    503, 529
  - _MESSAGE_HEAD_RULES for the status-less path (ahead of usage-limit)
  - _ERROR_CODE_VERDICTS for the structured oMLX codes
  - _classify_400, because _by_status runs before _by_error_code, so a body
    whose wording a proxy stripped would otherwise fall through to
    format_error

Every pattern names memory/allocation in bytes, never a token or window count,
so the list stays disjoint from _CONTEXT_OVERFLOW_PATTERNS. Both oMLX wordings
are kept: 0.5.6 says "Prefill would require ~13.87 GB peak", 0.5.7 reworded it
to "predicted peak would require/exceed" and both are in the field. A test
pins that a genuine window overflow still compresses.

Refs #52261. Supersedes #52289, which predates the classifier rewrite and can
no longer be merged.
2026-09-09 12:17:29 +05:30
kshitijk4poor 586a7831d2 refactor(background-review): collapse the tool-surface copy to the agent_init shape; 2 tests
agent.tools is always a list (agent_init assigns it from
get_tool_definitions) and every entry is a well-formed function schema,
so the isinstance ladder over parent/entry/function/name guarded shapes
that cannot reach this helper. Use the same two lines agent_init uses;
`or []` keeps the empty-surface contract from the previous commit.
Docstring cut to the WHY (the between-turn refresh note described the
other guard). Tests trimmed to the two invariants: inherited tools
survive the compaction-boundary refresh (deep-copy isolation folded in),
and an empty parent surface is copied and frozen. Literal sentinel
asserts replaced with the constant.
2026-09-09 10:32:49 +05:30
0xAlyDev 06e59f2815 fix(background-review): inherit and freeze empty parent tools list for cache parity (#103579)
Copy and freeze review_agent._tool_snapshot_generation even when parent.tools is an empty list ([]). Previously, the truthiness check (\
ot parent_tools\) caused an empty parent tool surface to be skipped, allowing newly available late MCP or plugin tools to be retained on the review fork and leaving its snapshot generation unfrozen. This broke the byte-parity contract when no tools were active on the parent.

Returning early only when parent_tools is not an instance of list or tuple guarantees that an empty tool snapshot is faithfully inherited and frozen. Adds dedicated regression test test_unrouted_review_fork_inherits_empty_tool_surface.
2026-09-09 10:32:49 +05:30
0xAlyDev 0d72e07e67 fix(background-review): freeze review fork tool snapshot generation against compaction refresh
Freezes review_agent._tool_snapshot_generation to _FROZEN_TOOL_SNAPSHOT_GENERATION
(2_147_483_647) when inheriting the parent tool surface for same-model cache parity.

When in-place compaction boundaries trigger refresh_agent_mcp_tools(content_aware=True),
the staleness guard in _publish_tool_snapshot refuses the rebuild (snapshot_generation < published_gen),
preventing agent.tools from being reconstructed from the raw registry and preserving
inherited memory-provider and late tools across compaction boundaries (#103579).

Adds unit regression test verifying tool preservation across content_aware refresh.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-09 10:32:49 +05:30
0xAlyDev 6a97436e29 fix(agent): inherit parent's full tool surface on review fork for cache parity (#103579)
Ensure unrouted background_review forks inherit the parent's full advertised
tools[] surface. Without this, skip_memory=True caused memory-provider tools
(e.g. fact_store/fact_feedback) and dynamically injected plugin/late MCP tools
to be omitted from the fork's tools array, breaking byte-exact prefix-cache parity
and incurring full cold-read costs on providers where tools are part of the cache key.
Inheriting the full parent tools array preserves complete prefix cache parity
while execution dispatch remains strictly bounded by the thread tool whitelist.
2026-09-09 10:32:49 +05:30
kshitijk4poor 63c6f9bf14 simplify(agent): sidecar backfill — drop the hasattr guard and the duplicated row-id predicate; tests 7→6
_session_db is always a SessionDB (agent_init / delegate_tool), so the
"fail closed on a store wrapper" hasattr was defense around code that
cannot fail; the store's own guard binds the value into SQL, so the
prologue only needs the sibling idiom isinstance(_row_id, int) that
session_persistence and transcript_repair already use. The positional
hazard is explained once, on set_latest_user_api_content. The in-place
compaction test duplicated test_api_content_sidecar's
test_inplace_compaction_backfills_sidecar_into_db verbatim (its row_id
parameter was never varied); dropped, as was the positional-helper tail
of test_older_identical_row_is_untouched already covered there.
2026-09-09 10:32:01 +05:30
kshitijk4poor 73e3547ffd refactor(agent): one durable-row rule for the flush and the sidecar stamp; trim tests
The turn-start stamp had grown its own copy of the "what does the current
user row hold" rule (persist override = clean transcript, live content =
wire bytes = sidecar when they differ) that _db_flush_row already
implements. Two copies drift; extract durable_user_row_content() in
session_persistence and call it from both.

Also: reuse _persist_lock() instead of a third open-coded lock/nullcontext
ladder; drop the hasattr guard on set_latest_user_api_content (it predates
this fix and exists on every SessionDB); cut the comment to the WHY;
trim the new test file from 18 cases to the 7 invariants (real close
flush E2E, repeated-"ok" positional protection, API-only pre-flushed
turn, normal path writes nothing, compaction keeps positional, store
guards). Still 3 red / 4 green when agent/turn_context.py is swapped
for main's copy.
2026-09-09 10:32:01 +05:30
kshitijk4poor 4126b144bb fix(agent): read the sidecar row id under the session persist lock
_stamp_api_content_sidecar read _row_id without holding
_session_persist_lock. A close/early flush holds that lock while it
commits the row and only afterwards writes _row_id back onto the live
dict; a stamp that ran in between saw no id and skipped the backfill,
the flush finished with api_content = NULL and marked the message
persisted, and the turn-start persist skipped it — the row kept the
wrong bytes with no writer left to fix it.

Run the _row_id read and the DB backfill under the (re-entrant) lock,
re-checking _row_id after acquiring it. Race reported by @ehz0ah on

Co-authored-by: sal <141555468+salch-cred@users.noreply.github.com>
#102411; same fix shape as @salch-cred's follow-up on #103721.
2026-09-09 10:32:01 +05:30
joaomarcos bc16c32c05 fix(agent): row-addressed api_content backfill for pre-persisted user turns (#102194)
The api_content sidecar ('persist what you send') preserves prompt-cache
stability across turn boundaries by persisting the exact API-bound bytes
(including memory-manager prefetch, plugin injections, and API-only notes)
and substituting them on replay.

When a user turn was already materialized in the database before the
sidecar could be composed (in-place preflight compaction or a close/early
flush racing the prologue on the CLI path), the turn-start crash persist
marker-skips that message. Previously, the backfill was gated strictly on
in-place compaction (_preflight_compressed and _last_compaction_in_place),
so racing CLI flushes left api_content = NULL in SQLite and broke prompt
caching on subsequent turns (#102194).

Positional approaches (such as #102239 and #102286) using LIMIT 1 on the
newest active user row are unsafe: repeated common inputs ('ok', 'yes',
'continue') cause the backfill to match and overwrite the PREVIOUS turn's
row with the new turn's sidecar, corrupting history and breaking cache parity.

Resolve all landing blockers and review feedback from #102411:

1. Bounded state owner (Sahilvishnaliya):
   Add SessionDB.set_message_api_content(session_id, row_id, content, api_content)
   to SessionMessagesMixin in hermes_state_messages.py instead of growing
   hermes_state.py. Update set_latest_user_api_content docstring with durable
   warning on the positional hazard.

2. API-only turns & durable content selection (ehz0ah):
   When a pre-flushed clean input has an API-only difference (e.g. voice
   prefix or model-switch note):
   - Retain the differing API-facing bytes as api_content even when no
     new memory or plugin context was injected.
   - Derive the durable content guard using _override_replaces_content so
     the SQL 'content IS ?' guard matches the clean override text stored
     in the DB row rather than the restored wire text.

3. Turn prologue gating (_row_id) & fail-closed store duck-typing (ehz0ah):
   In agent/turn_context.py::_stamp_api_content_sidecar: check _row_id on
   the live user dict (stamped by _insert_message_rows and synced by
   sync_flushed_message_markers). If valid (positive int, not bool), address
   by exact ID. Do NOT fall back to positional matching when a row ID is
   present: if an external or custom wrapper lacks set_message_api_content,
   fail closed and skip rather than corrupting a neighbouring row. If absent
   but in-place compacted, fall back to positional update. On normal turns,
   skip the backfill entirely (single atomic INSERT).

4. Real lifecycle test coverage (salch-cred, ehz0ah):
   Comprehensive tests in tests/agent/test_api_content_row_addressed_backfill.py
   covering store guards, surrogate scrubbing, gate non-arming, older identical
   row protection, real close-flush row_id synchronization, API-only clean
   override preservation with exact wire replay, and duck-typed store fail-closed
   verification when set_message_api_content is absent.

Fixes #102194.
Closes #102411.
2026-09-09 10:32:01 +05:30
kshitijk4poor defdf64790 simplify(agent): surface switch — reuse flatten_message_text / agent_tool_names / one runtime-boundary split
- _transcript_row_texts re-implemented agent.message_content.flatten_message_text
  and the api_content sidecar rule; the note can only land on a user row,
  so the transcript scan now skips assistant/tool rows (the bulk of the bytes).
- Three sites computed "names of agent.tools"; tools.mcp_tool_agent gains
  agent_tool_names() used by the switch note and conversation_loop, which
  also stops importing the private _def_name across modules. The name list
  is only captured when a switch was announced.
- split_runtime_boundary() is the single owner of the runtime-block
  rpartition/END check for both identity_line_value and
  _stored_prompt_matches_runtime.
- platform_surface_hint was a public alias of _platform_hint; the function is
  now platform_hint (its docstring pointed at the pre-move module).
- consume_gateway_turn_context_notes and consume_surface_switch_note share
  _pop_turn_note so the two one-shot channels have identical semantics.
- platform check hoisted above the transcript scan.
2026-09-09 10:31:26 +05:30
kshitijk4poor 6cc177a76c refactor(agent): surface-switch note lives in its own sibling; skip it where no sidecar exists
Move the six surface-switch helpers out of the conversation_loop facade
into agent/surface_switch.py (AGENTS.md: new behaviour goes in a topical
sibling), and fold the review findings on #104494:

- MoA and codex_app_server turns never stamp the api_content sidecar, so
  the staged note could not be read back from the transcript and was
  re-sent on every turn after a switch. Those modes now skip the note
  (stored prompt still reused).
- The announced surface was parsed with split(".") — a plugin platform
  with a dot in its name would never compare equal and re-stage the note
  every turn. The note now closes the name with a fixed terminator.
- One identity-line parser (identity_line_value) shared by
  _stored_prompt_matches_runtime and the switch detector instead of two
  copies of the runtime-boundary/rpartition logic; tool names via the
  existing tools.mcp_tool_agent._def_name; the transcript scan is bounded
  to the last 200 rows (it ran every turn over the whole history).
- consume_surface_switch_note reduced to a plain pop; developer-guide
  prompt-assembly.md updated (Platform is no longer an identity field);
  17 new tests trimmed to 10 (same-shape pin/retire variants folded).

Restoring Platform as an identity field still turns 5 tests red.
2026-09-09 10:31:26 +05:30
joaomarcos 4e7a49d182 fix(agent): retire stale surface notes on bot-chat refresh, isolate platform from decoys
When Bot Chat capability refresh rebuilds the system prompt for the current
surface, call _stage_surface_switch_note() so any earlier switch note sitting in
the transcript is retired instead of overriding the rebuilt prompt.

Also isolate _stored_prompt_platform() to parse only the authoritative identity
portion before '# Hermes runtime environment' (with legacy fallback for prompts
without the boundary), preventing embedder prose or HERMES_ENVIRONMENT_HINT decoys
from shadowing the real platform and falsely suppressing surface switch announcements.

Credit to @ehz0ah, who identified both correctness gaps on current main and
verified the regression scenarios.
2026-09-09 10:31:26 +05:30
joaomarcos 4a96311503 fix(agent): hold the tools pin through a surface switch, name what it carried
The announcing turn used to skip the tools freeze and re-persist the array the new
surface had just built. That is the one mutation this fix cannot afford: tools[] is
serialized ahead of the system prompt, so rebuilding it moves the request at token 0
and re-prefills everything behind it — the exact cost #104414 measured (1% cache hit
on a 220K session), spent on the very turn the fix exists to make cheap. On a
`desktop -> tui` switch with a configured toolset selection (`_gui_surface_toolsets`
gives desktop `desktop_ui`, the TUI nothing), skipping the pin dropped ~a dozen tools
and bought back the whole miss.

The pin now holds. `_merge_preserving_prefix` still appends what the new surface
brought, so a `tui -> desktop` switch pays a break no freeze could have avoided, and
the tools it carries FORWARD are named at the end of the surface note instead of being
silently advertised: a `focus_pane` a terminal turn can only answer with
`tool_error("desktop only")` now reads as unavailable rather than as live capability.
The toolset converges at the next real rebuild boundary, where the break is already
paid.

Credit to @StanleyStetson, who caught that the tool array is evaluated ahead of the
system prompt and that the bypass reintroduced the miss this PR is about.
2026-09-09 10:31:26 +05:30
joaomarcos a020050c8c fix(agent): the surface note must not outlive its own truth
Two holes the first cut left open, both created by the note itself.

Switching BACK to the surface the prompt was built for (desktop -> tui -> desktop) left the
`Platform:` trailer agreeing with the runtime, so nothing was staged — while the newest note
in the transcript still told the model it was on tui. And a rebuild for an unrelated reason
(a model switch) refreshed the prompt but not that note, leaving the same contradiction from
the other side.

Compare the runtime surface against what the model was last TOLD — the newest surface note
when one exists, else the prompt's own trailer — and stage from the rebuild path too. The
full surface guidance rides along only when the prompt itself is out of date; when the prompt
already describes the current surface the note just retires the stale one and points at it.
2026-09-09 10:31:26 +05:30
joaomarcos aa40dbe765 fix(agent): skip the tools freeze once on a surface switch, not for the session
The first cut gated the saved tool_names pin on "the surface drifted", which stays true
for as long as the stored prompt names the old surface — i.e. until the next compaction.
On the gateway path, where a fresh AIAgent is built per turn, that left the tools freeze
off for every remaining turn, so a check_fn that flaps could reorder `tools[]` and break
the tool cache block on its own.

Gate it on the turn that actually ANNOUNCES the switch instead, and persist the fresh,
toolset-correct names there. The next turn's row already holds this surface's tools, so
the pin resumes immediately: skipped once, not disabled.
2026-09-09 10:31:26 +05:30
joaomarcos 80d6bda144 fix(agent): a surface switch must not re-prefill the whole request (#104414)
`_stored_prompt_matches_runtime` treated `Platform` as a runtime-identity field, so
answering a live session from another surface — desktop -> TUI, or a resume after a
dashboard restart whose chat is a PTY TUI child — declared the stored prompt stale and
rebuilt it. The system prompt is the first thing in the request, so changing any byte of
it moves the first divergent byte to the head of a 220K-token request and the entire
conversation behind it re-prefills: a session that was hitting 240000/240287 came back
at 1536/219861.

The guard was not wrong about correctness — a desktop-built prompt on a terminal session
advertises inline widgets and a MEDIA: channel the TUI does not have — but the surface is
advisory metadata about the renderer, not a cache domain. Model/provider and cwd drift
change what the prompt should SAY; the surface changes only one paragraph.

Reuse the stored bytes across a surface switch and correct the paragraph where it costs
nothing to cache: `_stage_surface_switch_note` stages a one-shot note carrying the CURRENT
surface's guidance on the same per-turn user-message channel the gateway's must-deliver
notes use. It lands after the cached prefix and is stamped into the byte-stable
`api_content` sidecar, so later turns replay it instead of re-prefilling, and the prompt
converges at the next compaction — a boundary that already breaks the cache.

The saved tool_names prefix is not pinned across a switch: the tool registry is
process-global, so `_merge_preserving_prefix` would carry a saved-but-unloaded tool
forward (under `coding_context: focus` desktop gets a desktop_ui toolset the TUI cannot
run). On the same surface the tools freeze is untouched.
2026-09-09 10:31:26 +05:30
Teknium 78afbc3c37 fix(desktop): load property card guidance only on demand 2026-09-08 13:39:01 -07:00
Teknium 19cd839d54 fix(compression): keep lean tails lean after auxiliary feasibility
Lowering the session trigger must not replace the window-relative lean
selection budget with threshold times target_ratio. Invalidate the lean
cache through the existing property while preserving explicit legacy and
external-engine fallback behavior.

Narrow adaptation of the aux-sync diagnosis and invariants in #93576,
without adding a required recalibration method to context engines.
Related: #95681, #93576

Co-authored-by: Turgut Kural <58116817+TurgutKural@users.noreply.github.com>
2026-09-08 13:37:01 -07:00
Teknium 9d661c2c92 fix(prompt): keep memory guidance within available tools 2026-09-08 13:36:05 -07:00
Teknium 9e048186a9 fix: make review workers recognizable in live subagent viewers 2026-09-08 04:38:47 -07:00
Teknium 520e63661c fix: keep command-auth model discovery lazy across config and setup 2026-09-07 21:22:49 -07:00
Hayden Moulds c111ede3e5 fix(picker): resolve key_cmd credentials for model discovery
`key_cmd` (#86891) authenticates a provider with a SHORT-LIVED bearer minted
by a command — SSO/OIDC brokers, cloud IAM, internal auth proxies. The
request path has honoured it since it landed, but the picker resolved probe
credentials from `api_key`/`key_env` ONLY, so a key_cmd provider probed
`/v1/models` with an EMPTY key.

Against an authenticated endpoint the probe 401s, discovery returns nothing,
and the provider falls back to its single configured default model. The
picker shows ONE model, indistinguishable from an endpoint that genuinely
serves one — while inference keeps working, because that path mints
correctly. Reproduced against a LiteLLM gateway behind Entra OIDC: 0 models
discovered with an empty key, 26 with the minted token.

Both picker probe sites already funnel through `_entry_credentials()`, so
the fix lands in one place: it now reports a `cmd:<key_cmd>` identity, and
each site falls back to `resolve_probe_token()` after api_key/key_env. An
explicit static key still wins, so existing configs are unaffected.

The identity is keyed on the COMMAND, never the minted token: the token
rotates on every refresh, so keying on its value would change the group
fingerprint constantly and force a re-probe on every open. Two entries on
one URL with different helpers still get distinct rows.

`resolve_probe_token()` lives in agent.command_token_source, which already
owns key_cmd minting, and shares the CommandTokenSource cache with the
request path — a cache read, not a fresh sign-in. Fail-closed: a helper
needing an interactive sign-in degrades to today's empty-key behaviour
rather than taking down every other provider's row.

`_model_flow_named_custom` (the `hermes model` setup flow) is the sibling
path — it builds its own `Authorization: Bearer` from the same incomplete
resolution — and is fixed the same way, with one ordering constraint: the
value persisted to config.yaml is computed BEFORE the mint, so a short-lived
bearer can never be written back to shadow the key_cmd meant to re-mint it.

Tests drive the real code paths and assert on the credential each probe
receives rather than on function source, so a semantics-preserving refactor
does not fail them. Verified they fail with the fix reverted.
2026-09-07 21:22:49 -07:00
Teknium bbcf1ee180 fix: preserve native Gemini union constraints
Complete the type-array normalization salvaged from #55643: stringify mixed
union enum metadata, preserve existing anyOf constraints, and keep array
items and object properties/required on the corresponding typed branches.

Exercise real native request serialization over loopback and Google SDK
validation with a scalar control; no live Google credentials were available.
2026-09-07 21:10:50 -07:00
Max Freedom Pollard 6a04ea67c0 fix(gemini): collapse array-typed tool schemas instead of crashing translation
The enum-compatibility check evaluated `[...] in {...}` on an array `type`,
raising TypeError: unhashable type: 'list' and aborting translation of the whole
tool catalog rather than the one offending tool.

Addresses both review points:

- Reuses tools.schema_sanitizer._normalize_type_array instead of picking the
  first non-null member, so a real union becomes an anyOf of single-type
  branches and no branch is dropped.
- Derives the type outside the key loop and sets nullable after it, so the
  flag implied by "null" in the array beats an input nullable: false whichever
  key the producer emitted first. Both orders are pinned by a parametrized test.
2026-09-07 21:10:50 -07:00
Teknium 5280fe9987 fix: cron and local DMs reach an open Desktop Bot Chat
Route local producers to durable owner ingress before attempting the unowned
CLI lane. Preserve per-run/per-message IDs and receipt-first retry handling;
never fall back after ambiguous admission. Report cron admission as queued,
not completed or failed, in job status, the execution ledger and CLI/tool UX.

Native isolated Electron validation reproduces SESSION_NOT_OWNED on main for
both idle and busy owners. Fixed owner consumes idle cron, busy cron, local
DM and mounted-chat cron exactly once, keeps its lease, yields to queued
human input, and preserves the prior model-request prefix and tool schema.
Inference alone used a deterministic loopback wire stub; no paid model call.
2026-09-07 16:48:29 -07:00