Commit Graph

25 Commits

Author SHA1 Message Date
Yong Li f41ed09b51 fix(gemini): strip call ids on insert, name the realignments in the log
Review follow-ups:

- Strip the tool_call id when populating the call-name map so it matches the
  stripped lookup (and pass 1's result_call_ids). A padded id previously
  skipped realignment silently.
- Log which names were rewritten, not just how many.
- Note in the comment that a result whose assistant call frame was pruned is
  already dropped by the orphan pass, so it cannot reach the provider with a
  stale name; cover that with a test.
- Rename test_sanitize_leaves_matching_and_unpaired_tool_result_names_alone,
  which only ever exercised the matching case, and add the padded-id case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 14:56:30 +05:30
Yong Li 2e9435d2f6 fix(gemini): echo bridged tool_call name on the OpenAI-compatible path
Google matches functionResponse.name against functionCall.name and rejects
a mismatch with HTTP 400 INVALID_ARGUMENT. #72089 fixed this for the native
Gemini adapter, where _translate_tool_result_to_gemini() now prefers
tool_name_by_call_id over the result message's internal name.

Requests that reach Gemini through an OpenAI-compatible gateway (OpenRouter,
Vertex/LiteLLM proxies) never run that translation, so they still put the
unwrapped internal tool name on the wire: the model calls the tool_search
bridge tool `tool_call`, make_tool_result_message() labels the result
`mcp__strava__get_recent_activities`, and the next turn 400s with a bare
"Provider returned error". The bad pair stays in the transcript, so every
later request in that session fails too.

Hold the same invariant at the final pre-API chokepoint instead of in the
OpenAI-compat serializer: Gemini arrives under many model strings and base
URLs, so sniffing for "is this really Google?" is unreliable, while every
other provider either ignores the field or already agrees with the call
name. Only a name that is present and disagrees is rewritten, so clean
transcripts still pass through byte-identical for prompt caching, and the
rewrite lands on the per-call copy so the stored trajectory keeps the real
tool name for the session DB and UI. No-op for the native Gemini path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-01 14:56:30 +05:30
fedebyes 93f4dc7561 fix: make positional prune variant-aware; add replayed-call regression tests
Pass 2 of repair_message_sequence matched results only by id/call_id,
pruning calls answered through response_item_id or composite bridge
ids. Use the shared variant helpers (tool_call_id_variants /
tool_result_id_variants) so the unified alias policy applies
(#55626/#63000/#93251).

The positional sanitizer pass changes the crash/resume duplicate shape:
an interrupted first occurrence is now stubbed instead of deduped, so
the replayed call survives with its own immediate result. Update the
#64335 empty-key test to the new semantics and add regression tests for
the #94704 acceptance shape (historical-result + replayed-call +
fresh-call) and the production interrupted-turn shape (session
7d57a602b83d).
2026-08-28 07:51:23 -07:00
Tiberiu Danciu c7761573f5 fix: prune positionally unanswered tool_calls before API send
DeepSeek v4 rejects a payload where an assistant message carries a
tool_call whose tool result does not follow it immediately (HTTP 400
"An assistant message with 'tool_calls' must be followed by tool
messages responding to each 'tool_call_id'"). Context compression can
displace a tool result past a user turn; the result then lands ~100
messages away from its declaring assistant message.

Two gaps let the poisoned shape reach the wire (reproduced from the
production request dump of session 4d8727cbcf04, replayed through both
functions):

1. repair_message_sequence Pass 1 drops the displaced tool RESULT as
   stray but leaves the declaring assistant message carrying the now
   unanswered tool_call (with empty content) in the durable history.
2. sanitize_api_messages stubbed only globally-absent result ids: the
   displaced result still exists in the transcript, so the id survives
   the set-subtraction, no stub is injected, and the payload 400s.

Fix both layers so every path is order-independent:

- repair_message_sequence: new Pass 2 prunes tool_calls that have no
  result in the immediately-following tool run (matching on id or
  call_id, same superset rule as Pass 1). If pruning empties the turn
  (no content/reasoning left), the whole message is dropped rather than
  sending an empty assistant message. Codex interim turns are exempt,
  as in Pass 0.
- sanitize_api_messages: the orphan/stub logic is rewritten as a single
  rolling positional walk that drops results not immediately following
  their declaring assistant (including results appearing BEFORE their
  call) and injects stub results for positionally-uncovered calls even
  when a mispositioned result exists elsewhere.

Adds six regression tests: repair pruning, whole-turn drop when pruned
calls were the only payload, valid-pair negative control, positional
stub injection, result-before-call orphan drop, and a fully-paired
transcript negative control.
2026-08-28 07:51:23 -07:00
isheng 87cff9d4c1 test(sanitizer): add unit tests for _classify_tool_call_orphans 2026-08-28 06:32:48 -07:00
joaomarcos f0ac2c8f12 fix(agent): drop stale api_content sidecar and unpaired tool results
Rebased onto current main to drop the empty-tool_calls fix (already on
main via #86654, cherry-picked from #77944 with @webtecnica's
authorship). This PR now carries only the two fixes unique to it:

1. A pre-existing api_content sidecar left stale on the consecutive-
   assistant merge. The sidecar takes priority over content at
   API-build time, so a merge could silently discard its own freshly
   concatenated content on the next call. Only dropped when the merge
   actually changes the resulting value (wz-heng, #78063 review) --
   content_rewritten compares before/after value, not just whether an
   assignment branch fired, so a falsy new_content (e.g. "") that
   strips to nothing no longer trips a spurious sidecar drop.

2. sanitize_api_messages never flagged a tool result with a missing/
   empty tool_call_id -- its orphan-detection set only ever collected
   truthy ids, so an unpaired result with no id passed the final
   chokepoint untouched.

Addresses teknium1's rebase request and wz-heng's review findings on
2026-08-28 06:32:48 -07:00
Teknium a2a43f7e82 fix(agent): widen composite-id alias matching to the compressor; unify variant policy owners (#63000)
Follow-up on top of the salvaged #93335:

- context_compressor._sanitize_tool_pairs now expands alias spellings on
  the RESULT side too (tool_result_id_variants), so a composite
  call|item-keyed result pairs with its split-field tool_call instead of
  being dropped and its call stripped.
- The compressor's _tool_call_id_variants staticmethod and
  agent_runtime_helpers' module-level _tool_call_id_variants are now thin
  forwarders to agent.message_sanitization.tool_call_id_variants — one
  policy owner for alias expansion, so the pre-call sanitizer, repair
  pass, dedup pass, and compression sanitizer can never drift apart.
- Preserved the #91768 SDK-object tolerance in repair pass 1 (the
  shared helper handles non-dict tool_calls via getattr; the salvaged
  commit's isinstance-dict guard was dropped in the merge resolution).

New regression tests: composite-keyed results through
sanitize_api_messages (both directions) and _sanitize_tool_pairs, with
negative controls. Sabotage-verified: compressor test fails with raw
tool_call_id tracking.
2026-08-23 18:24:43 -07:00
joaomarcos 5496d5995a fix(agent): preserve tool results across ID variants
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
2026-08-23 18:24:43 -07:00
Teknium faa2399e2b fix(agent): make the pre-call dedup pass variant-aware; widen batch regression coverage (#93251)
Follow-up on top of the salvaged cluster: sanitize_api_messages step 3
(duplicate tool_call_id dedup) still tracked only the coalesced
(call_id||id) value in outstanding_call_ids, so after step 2's
variant-aware matching preserved a result keyed on the OTHER id variant,
step 3 deleted it as answering no outstanding call — whole parallel
batches of real results vanished with no stub at all (#93251's total-loss
mode). Track the full variant set per call and consume all siblings when
answered, preserving #58327 duplicate protection and llama.cpp
constant-id re-arm semantics.

Also aligns the #58287 compressor test with the in-flight tool chain
protection (#79278) that landed after that PR was opened: a trailing
user turn keeps the negative-control assistant message out of the
protected trailing window.

New regression tests: divergent-id batch survival through the dedup
pass, sibling-id replay still dropped, constant-id re-arm preserved.
Sabotage-verified: tests fail with the old single-id tracking.
2026-08-23 17:01:20 -07:00
joaomarcos 36b4da5489 fix: repair_message_sequence drops tool results for SDK tool_call objects
The tool_call id-matching pass in repair_message_sequence only read
`.get("id"/"call_id")` on plain dicts, skipping non-dict tool_calls
entirely (`if not isinstance(tc, dict): continue`). Host-fed and
pre-serialization histories can carry unserialized SDK tool_call
objects (e.g. `ChatCompletionMessageToolCall`) instead of dicts, which
left `known_tool_ids` empty for that assistant turn. The following
`tool` message — a legitimate result already produced by executing the
tool — was then misclassified as an orphan and silently dropped,
corrupting the persisted conversation history and leaving the
assistant's tool_calls unanswered (itself a trigger for HTTP 400 on
strict providers).

Fix: extract id/call_id via getattr() for non-dict entries too,
mirroring AIAgent._get_tool_call_id_static's existing dict-or-object
tolerance, instead of skipping them.
2026-08-23 17:01:20 -07:00
Frowtek b9a62f6590 fix(agent): consume every tool_call id variant when pairing tool results
`repair_message_sequence` registers BOTH `id` and `call_id` for each
assistant tool_call, because a matching tool result may be keyed on either
depending on which path built it (#58168). The duplicate guard added for
dropped rather than replayed.

Those two behaviours don't compose: a Codex/Responses tool_call registers
two DIFFERENT ids (`fc_...` and `call_...`), but only the id the first
result referenced is discarded. Its sibling stays in `known_tool_ids`, so a
duplicate result keyed on that sibling still matches and is kept — two tool
messages replayed for one call, which is exactly the HTTP 400 on strict
providers the consume step exists to prevent.

Duplicates of this kind come from the retry / crash / session-resume glitch
the guard was written for; the id-variant split just lets them slip past it.

Track each registered id back to its tool_call's full variant set and
discard all of them on a match. Results keyed on either variant are still
accepted (no false orphaning), and two parallel Codex calls answered via
different variants both survive.

Adds regression tests for the sibling-keyed duplicate and for the
two-calls/mixed-keys case that must NOT be affected.
2026-08-23 17:01:20 -07:00
Teknium df68cc1c1a test: adapt stale dedup test to outstanding-call semantics, add cross-turn coverage
Follow-up to the salvaged #70734 fix:

- test_sanitize_dedup_drops_tool_calls_key_when_all_removed encoded the old
  global-uniqueness assumption (its second assistant call reused the id AFTER
  the first call was answered, which is now a legitimate new call). The
  replayed call now precedes the result, making it a true duplicate of a
  still-outstanding call, preserving the intended empty-tool_calls key-drop
  coverage from #64335.
- New test: Hermes' own deterministic local counter ids repeating across
  turns (the #76632 scenario) survive sanitization.
- New test: the 50-step constant-id field repro from #70724 (Kimi K3 /
  llama.cpp) — stock main kept 1/50 tool results, now 50/50.
2026-08-17 03:25:07 -07:00
flaviovargasbrandao 0b8fd04bea fix(agent): keep tool results when a server reuses one tool_call_id
The #58327 dedup passes treat a repeated tool_call_id as garbage from a
retry/crash/resume glitch and drop it. That assumes tool_call_id is
globally unique, which it is not: llama.cpp emits a single constant id
for every tool call it ever returns (verified — three separate
completions from one server all carried the same id).

Under a seen-once-drop-forever rule, the SECOND legitimate tool result
of such a session looks like a duplicate and is deleted. From the second
tool call onward the model never sees any result: it announces its next
action, the turn ends, and the task is left unfinished. Bisected to
dba585c17 over a 2258-commit range; reproduced live on v0.19.0 (1/6 runs
completed a 4-step file task, vs 20/20 on the last release before that
commit, same model and server).

Key off OUTSTANDING calls instead of every id ever seen. Both original
protections are preserved: a replayed result still answers no pending
call and is still dropped, and duplicate tool_calls sharing an id within
one assistant message are still collapsed. A genuine new call that
reuses the id re-arms it first.

repair_message_sequence needs no change — it already resets its id set
per assistant message, so only the final pre-API pass mis-fires.

Live result after the fix: 8/8 runs complete, 17-26s each (was 1/6 with
runs hitting a 150s ceiling).
2026-08-17 03:25:07 -07:00
webtecnica f316f7d086 fix(session): drop empty tool_calls in repair_message_sequence (#77921) 2026-08-14 21:25:48 -07:00
liuhao1024 b2453b5894 fix(sanitize): drop tool_calls key when dedup removes all calls
The dedup pass in sanitize_api_messages (introduced by #58327) can
produce an empty tool_calls array when all tool_call_ids in a message
are duplicates of earlier messages in a long conversation history.

DeepSeek v4 and newer OpenAI reject empty tool_calls with HTTP 400:
'Invalid messages[N].tool_calls: empty array'.

When kept_tcs is empty after dedup, drop the tool_calls key entirely
instead of writing tool_calls: [].

Fixes #64335
2026-08-14 21:25:48 -07:00
devsart95 06ae5b6faa perf(state): batch the turn flush into one SQLite transaction
Re-derivation of #23254 (@devsart95) on today's flush loop. The turn
flush in _flush_messages_to_session_db wrote one BEGIN IMMEDIATE
transaction per message row; a typical agent turn (user + assistant +
tool results) paid 3-8 transactions -- and, off WAL (the default on
macOS while the WAL-reset guard is active), 3-8 fsyncs -- per turn.

Adds SessionDB.append_messages_batch: same row shape as append_message
(shared _prepare_message_row serializer + _MESSAGE_INSERT_SQL column
list, so the two writers cannot drift), same compression-lock and
compression-closed guards, one aggregated session-counter UPDATE, one
transaction for the whole batch. Row serialization stays outside the
write lock.

The flush loop now collects the turn's new rows and writes them in one
call. All-or-nothing pairs exactly with the persisted-marker stamping:
on failure no rows landed and no markers were stamped, so the next
flush re-writes the whole tail (same recovery contract as before,
minus the partial-prefix case that could double-count).

Measured (same harness, 5-message turn, journal_mode=DELETE,
synchronous=FULL): 2.32ms -> 0.83ms median per turn flush (64% faster,
5 fsyncs -> 1). On WAL the win is smaller but the atomicity fix holds.
2026-08-03 20:43:38 +05:30
Teknium 6b81590c55 test: prune low-value tests suite-wide (wave 1) — 46,820 → 28,106 test functions
Systematic prune per AGENTS.md test policy, one pass over every major
test tree (gateway, hermes_cli, tools, agent, run_agent, plugins, cli,
cron, tui_gateway, honcho/openviking, root-level):

- DELETE: source-reading tests (read_text/getsource on prod files),
  change-detector tests (exact catalog counts, model-name snapshots,
  config version literals), mock-echo tests (assert a mock returns what
  it was told), assertion-free/trivial tests, near-duplicate
  parametrizations (boundaries + one representative kept), async/sync
  twin duplicates, cosmetic within-file variations.
- KEEP (mandatory): security/redaction/approval guards, message-role
  alternation invariants, prompt-caching/deterministic-call-id
  invariants, issue-number regression tests (deduped), E2E tests.
- 6 test files deleted outright (script-style/no-assert or fully
  redundant); conftest.py, fakes/, fixtures/ untouched.
- tests/acp/conftest.py added: autouse fixture stubs the live
  models.dev/GitHub/Copilot/Anthropic inventory fetches that ACP server
  tests performed on every session create — test_server.py 147s → 3.4s,
  and the tests are now genuinely hermetic.
- Sleep-based slowness shrunk where safe (codex_ttfb_watchdog,
  compression_concurrent_fork, etc.); no wall-clock assertion tightened.

Verification: full hermetic suite via scripts/run_tests.sh —
2439 files, 31,130 tests passed, 0 failed, 0 flaky retries, 315s wall
(baseline: 583s wall, 13,564s subprocess CPU).
2026-07-29 13:10:23 -07:00
Hermes Agent 725c7ba534 refactor: single owner for empty-content wire repair (class fix)
The concept 'never send a turn that strict wire validation rejects as
empty' was forked across four sites, each with its own predicate and its
own blind spots:

1. build_assistant_message write-time ' ' pad — broke codex commentary
   turns (content:'' is a designed state), and a DB-side pad can't
   survive _rows_to_conversation's whitespace strip anyway. REMOVED.
2. conversation_loop send-time ' ' pad — main-loop only (summary path
   uncovered), ordering-fragile (had to run after whitespace
   normalization), assistant-only. REMOVED.
3. stream-stub '[response interrupted]' substitution — defeated the
   loop's empty-stub guard (the stub no longer looked empty, entered
   history, and the placeholder leaked into the stitched final
   response via truncated_response_parts). REMOVED.
4. repair_empty_non_final_messages in sanitize_api_messages — the
   unconditional pre-send chokepoint shared by the main loop AND the
   summary path, covers user and assistant turns, non-final only,
   copy-on-write. This is now the SINGLE OWNER.

The owner's payload predicate (_msg_has_payload) is extended to treat
codex_message_items / codex_reasoning_items as payload, so
designed-empty codex commentary turns are never rewritten on any
api_mode — the failure shape that broke site 1 in CI is encoded in the
owner, not special-cased at a call site.

Tests updated to pin the new contracts: builder stores textless turns
as-is; the empty stream stub stays recognizably empty for the loop
guard; poisoned resumed histories are repaired to the placeholder at
the send boundary; codex item carriers are never rewritten.
Sabotage-verified: unwiring the owner fails 3 regression tests.
2026-07-27 20:27:56 -07:00
Aaron Weiker df45811198 fix(errors): self-heal empty-content non-final messages before send
Third layer of the empty-stub fix: full self-recovery. A poisoned transcript
(empty assistant stub or empty user turn already persisted before the write-
time guard, or fed in from a host history) previously 400'd every subsequent
request until it scrolled out — needing a manual DB edit + gateway restart.

sanitize_api_messages() (the unconditional pre-send chokepoint) now repairs
empty non-final messages on the per-call copy by substituting a minimal
'[response interrupted]' placeholder, so the session recovers itself IN MEMORY
on the very next send. The final message is left untouched (empty final
assistant is legal); stored history is never mutated; reasoning-only and
tool_call turns are preserved (negative controls).

Tests: production-shape repro (tool -> empty assistant -> user), empty-user
case, non-destructive guarantee, and negative controls. RED verified by
disabling the wire-in.
2026-07-27 20:27:56 -07:00
xxxigm 9080c8b4fc test(agent): cover empty tool_calls array stripping in sanitizer (#58755)
Adds regression coverage for the DeepSeek v4 HTTP 400 fix:
- empty ``tool_calls: []`` is dropped, content preserved
- malformed non-list ``tool_calls`` is dropped
- stripping is non-destructive to the caller's persisted dicts
- populated tool_calls arrays survive untouched (negative control)
2026-07-05 21:36:37 -07:00
kshitijk4poor dba585c179 fix(agent): deduplicate tool_call_id across the pre-API sanitizers (#58327)
Strict providers (DeepSeek) reject a payload where the same tool_call_id
appears more than once with HTTP 400 'Duplicate value for tool_call_id'.
The issue was filed as an 'orphaned tool message' compression bug, but the
pasted error is a DUPLICATE tool_call_id — orphans are already handled on
main; duplicates were not. Reproduced live on main: both shapes leaked
through repair_message_sequence and sanitize_api_messages.

Two chokepoints, two shapes:
- repair_message_sequence: consume the id from known_tool_ids on first
  match so a SECOND tool result reusing it falls into the drop branch
  (duplicate tool-result shape). This is @Robinlovelace's kernel from
  #55436 (applied manually — that PR was ~800 commits stale and bundled
  an unrelated duplicate-DB-write change for #860, which is dropped here).
- sanitize_api_messages (final pre-API pass): add a dedup pass covering
  BOTH (a) duplicate tool_calls sharing an id WITHIN one assistant message
  (the message[6] shape) and (b) later tool result messages reusing an
  already-seen id. #55436 covered neither of these at this chokepoint.

Tests: duplicate-tool-result dedup at both functions, duplicate-assistant-
tool_call-id collapse, and a negative control proving distinct ids are
never dropped (no over-dedup).

Credit: @Robinlovelace (#55436) for the repair_message_sequence dedup kernel.
Closes #58327.
2026-07-04 21:14:37 +05:30
kshitijk4poor 88f2c0caf6 fix(agent): match tool results on call_id||id in pre-request repair (#58168)
repair_message_sequence Pass 1 registered only tc.get("id") when building
the set of known assistant tool_call ids, then matched tool results against
it by tool_call_id. In the Codex Responses format an assistant tool_call
carries both id (fc_...) and a distinct call_id (call_...); a tool result's
tool_call_id may be keyed on either depending on which builder produced it.
Registering only id made a valid tool result whose tool_call_id matched
call_id look orphaned, so the pass dropped it and left the assistant
tool_call unanswered -- producing HTTP 400 on strict providers (DeepSeek,
Kimi): 'Messages with role tool must be a response to a preceding message
with tool_calls'. Long-running sessions that persisted such a sequence were
permanently broken, re-sending the orphan every turn.

Register both id and call_id for each assistant tool_call so a result
matching either key is recognized, consistent with
AIAgent._get_tool_call_id_static and the compressor's _sanitize_tool_pairs.
Apply the same call_id||id precedence to the corrupted-args sanitizer's
existing-result scan / stub insertion, which had the identical mismatch.

Adds 3 regression tests covering the codex id!=call_id case (match on
call_id, match on only call_id, match on id when both present).
2026-07-04 15:15:33 +05:30
Teknium cbe397ef45 fix(agent): merge consecutive assistant messages before API replay (#29148, #49147) (#55603)
* fix(agent): merge consecutive assistant messages in repair_message_sequence

Strict OpenAI-compatible providers (DeepSeek v4, Moonshot/Kimi) reject a
replayed history where an assistant message carrying tool_calls is
immediately followed by another assistant message instead of its tool
results — HTTP 400 'An assistant message with tool_calls must be
followed by tool messages...'.

repair_message_sequence (the defensive belt run before every API call)
fixed orphan-tool and consecutive-user shapes but never merged
consecutive assistant messages. Adds a Pass 0 that collapses adjacent
assistant turns into one — union of tool_calls, concatenated content,
carried reasoning_content — covering both reported shapes:
  - parallel tool calls split across two assistant turns (#29148)
  - content-only assistant followed by tool_calls-only assistant (#49147)

A tool result or user turn between two assistants blocks the merge
(distinct, valid rounds). Runs before Pass 1 so the merged union of
tool_call ids is known to the orphan-tool filter.

Closes #29148, #49147.
Co-authored-by: Bartok9 <danielrpike9@gmail.com>
Co-authored-by: woaini30050 <woaini30050@users.noreply.github.com>
Co-authored-by: weidzhou <weidzhou@users.noreply.github.com>

* fix(agent): exempt codex Responses interim turns from assistant merge

The Pass 0 consecutive-assistant merge collapsed codex_responses interim
turns, which legitimately stay separate — each carries its own encrypted
continuation state (codex_reasoning_items / codex_message_items) that
must replay verbatim. Skip the merge when either side is a codex interim
(has codex_reasoning_items / codex_message_items / finish_reason=='incomplete').

Fixes the slice-2 regression in test_run_agent_codex_responses.py
(test_duplicate_detection_distinguishes_different_codex_{reasoning,message_items}).

---------

Co-authored-by: Bartok9 <danielrpike9@gmail.com>
Co-authored-by: woaini30050 <woaini30050@users.noreply.github.com>
Co-authored-by: weidzhou <weidzhou@users.noreply.github.com>
2026-06-30 04:22:56 -07:00
Teknium 8905ee6b8a fix(agent): rewind flush cursor exactly when repair compacts before the cursor
Follow-up to the #44837 clamp: a min() clamp only fixes cursor overshoot
past the new end of the list. When repair_message_sequence drops/merges
messages at indexes below the cursor, the clamp leaves the cursor pointing
past unflushed rows and the turn-end flush silently skips them.

Extract repair_message_sequence_with_cursor(): snapshot the flushed prefix
by object identity before repair, then recompute the cursor as the count
of surviving flushed messages. Falls back to the clamp when no snapshot is
available. Keeps the safety guard in _flush_messages_to_session_db.

Adds targeted tests for overshoot, before-cursor compaction, no-repair,
bare-agent, and the flush guard.
2026-06-12 16:29:01 -07:00
Teknium 812ce0b987 fix(run_agent): break permanent empty-response loop from orphan tool-tail (#21385)
When empty-response terminal scaffolding fires on a tool-result turn,
_drop_trailing_empty_response_scaffolding left the live history ending at
a bare 'tool' message. The next user input then landed as [...tool, user],
a protocol-invalid sequence that OpenRouter/Opus and other providers
silently fail on (returns empty content). That retriggered the empty-retry
recovery every turn, and recovery flags never hit SQLite (no column for
them), so history kept looking broken on every reload.

Two fixes:

1. Scaffolding strip rewinds the orphan assistant(tool_calls)+tool pair
   after popping sentinels. Only fires when scaffolding flags were
   actually present, so mid-iteration tool loops are untouched.

2. _repair_message_sequence runs right before every API call as a
   defensive belt: drops stray tool messages with unknown tool_call_ids,
   merges consecutive user messages so no user input is lost. Does NOT
   rewind assistant(tool_calls)+tool+user — that pattern is valid when
   the user redirected before the model got its continuation turn.

Repro: session 20260507_044111_fa7e65. Opus-4.7/OpenRouter returned
content-less response after a 42KB execute_code output, nudge+retry
chain exhausted (no fallback configured), terminal sentinel appended,
scaffolding stripped leaving bare tool tail, user typed 'wtf happened..'
and landed as tool→user violation. Every subsequent turn collapsed in
<50ms with the same 3-retry empty chain because the API request itself
was malformed.

Verified live via HTTP mock: pre-fix reproduced 5 api_calls/0.15s exit
'empty_response_exhausted'; post-fix 1 api_call/0.10s exit
'text_response(finish_reason=stop)'. Three-turn session flows cleanly
through the scenario. Full run_agent suite: 1242 passed (0 regressions,
2 pre-existing concurrent_interrupt failures unrelated).
2026-05-07 08:35:10 -07:00