Review follow-ups:
- Strip the tool_call id when populating the call-name map so it matches the
stripped lookup (and pass 1's result_call_ids). A padded id previously
skipped realignment silently.
- Log which names were rewritten, not just how many.
- Note in the comment that a result whose assistant call frame was pruned is
already dropped by the orphan pass, so it cannot reach the provider with a
stale name; cover that with a test.
- Rename test_sanitize_leaves_matching_and_unpaired_tool_result_names_alone,
which only ever exercised the matching case, and add the padded-id case.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Google matches functionResponse.name against functionCall.name and rejects
a mismatch with HTTP 400 INVALID_ARGUMENT. #72089 fixed this for the native
Gemini adapter, where _translate_tool_result_to_gemini() now prefers
tool_name_by_call_id over the result message's internal name.
Requests that reach Gemini through an OpenAI-compatible gateway (OpenRouter,
Vertex/LiteLLM proxies) never run that translation, so they still put the
unwrapped internal tool name on the wire: the model calls the tool_search
bridge tool `tool_call`, make_tool_result_message() labels the result
`mcp__strava__get_recent_activities`, and the next turn 400s with a bare
"Provider returned error". The bad pair stays in the transcript, so every
later request in that session fails too.
Hold the same invariant at the final pre-API chokepoint instead of in the
OpenAI-compat serializer: Gemini arrives under many model strings and base
URLs, so sniffing for "is this really Google?" is unreliable, while every
other provider either ignores the field or already agrees with the call
name. Only a name that is present and disagrees is rewritten, so clean
transcripts still pass through byte-identical for prompt caching, and the
rewrite lands on the per-call copy so the stored trajectory keeps the real
tool name for the session DB and UI. No-op for the native Gemini path.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Pass 2 of repair_message_sequence matched results only by id/call_id,
pruning calls answered through response_item_id or composite bridge
ids. Use the shared variant helpers (tool_call_id_variants /
tool_result_id_variants) so the unified alias policy applies
(#55626/#63000/#93251).
The positional sanitizer pass changes the crash/resume duplicate shape:
an interrupted first occurrence is now stubbed instead of deduped, so
the replayed call survives with its own immediate result. Update the
#64335 empty-key test to the new semantics and add regression tests for
the #94704 acceptance shape (historical-result + replayed-call +
fresh-call) and the production interrupted-turn shape (session
7d57a602b83d).
DeepSeek v4 rejects a payload where an assistant message carries a
tool_call whose tool result does not follow it immediately (HTTP 400
"An assistant message with 'tool_calls' must be followed by tool
messages responding to each 'tool_call_id'"). Context compression can
displace a tool result past a user turn; the result then lands ~100
messages away from its declaring assistant message.
Two gaps let the poisoned shape reach the wire (reproduced from the
production request dump of session 4d8727cbcf04, replayed through both
functions):
1. repair_message_sequence Pass 1 drops the displaced tool RESULT as
stray but leaves the declaring assistant message carrying the now
unanswered tool_call (with empty content) in the durable history.
2. sanitize_api_messages stubbed only globally-absent result ids: the
displaced result still exists in the transcript, so the id survives
the set-subtraction, no stub is injected, and the payload 400s.
Fix both layers so every path is order-independent:
- repair_message_sequence: new Pass 2 prunes tool_calls that have no
result in the immediately-following tool run (matching on id or
call_id, same superset rule as Pass 1). If pruning empties the turn
(no content/reasoning left), the whole message is dropped rather than
sending an empty assistant message. Codex interim turns are exempt,
as in Pass 0.
- sanitize_api_messages: the orphan/stub logic is rewritten as a single
rolling positional walk that drops results not immediately following
their declaring assistant (including results appearing BEFORE their
call) and injects stub results for positionally-uncovered calls even
when a mispositioned result exists elsewhere.
Adds six regression tests: repair pruning, whole-turn drop when pruned
calls were the only payload, valid-pair negative control, positional
stub injection, result-before-call orphan drop, and a fully-paired
transcript negative control.
Rebased onto current main to drop the empty-tool_calls fix (already on
main via #86654, cherry-picked from #77944 with @webtecnica's
authorship). This PR now carries only the two fixes unique to it:
1. A pre-existing api_content sidecar left stale on the consecutive-
assistant merge. The sidecar takes priority over content at
API-build time, so a merge could silently discard its own freshly
concatenated content on the next call. Only dropped when the merge
actually changes the resulting value (wz-heng, #78063 review) --
content_rewritten compares before/after value, not just whether an
assignment branch fired, so a falsy new_content (e.g. "") that
strips to nothing no longer trips a spurious sidecar drop.
2. sanitize_api_messages never flagged a tool result with a missing/
empty tool_call_id -- its orphan-detection set only ever collected
truthy ids, so an unpaired result with no id passed the final
chokepoint untouched.
Addresses teknium1's rebase request and wz-heng's review findings on
Follow-up on top of the salvaged #93335:
- context_compressor._sanitize_tool_pairs now expands alias spellings on
the RESULT side too (tool_result_id_variants), so a composite
call|item-keyed result pairs with its split-field tool_call instead of
being dropped and its call stripped.
- The compressor's _tool_call_id_variants staticmethod and
agent_runtime_helpers' module-level _tool_call_id_variants are now thin
forwarders to agent.message_sanitization.tool_call_id_variants — one
policy owner for alias expansion, so the pre-call sanitizer, repair
pass, dedup pass, and compression sanitizer can never drift apart.
- Preserved the #91768 SDK-object tolerance in repair pass 1 (the
shared helper handles non-dict tool_calls via getattr; the salvaged
commit's isinstance-dict guard was dropped in the merge resolution).
New regression tests: composite-keyed results through
sanitize_api_messages (both directions) and _sanitize_tool_pairs, with
negative controls. Sabotage-verified: compressor test fails with raw
tool_call_id tracking.
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
Follow-up on top of the salvaged cluster: sanitize_api_messages step 3
(duplicate tool_call_id dedup) still tracked only the coalesced
(call_id||id) value in outstanding_call_ids, so after step 2's
variant-aware matching preserved a result keyed on the OTHER id variant,
step 3 deleted it as answering no outstanding call — whole parallel
batches of real results vanished with no stub at all (#93251's total-loss
mode). Track the full variant set per call and consume all siblings when
answered, preserving #58327 duplicate protection and llama.cpp
constant-id re-arm semantics.
Also aligns the #58287 compressor test with the in-flight tool chain
protection (#79278) that landed after that PR was opened: a trailing
user turn keeps the negative-control assistant message out of the
protected trailing window.
New regression tests: divergent-id batch survival through the dedup
pass, sibling-id replay still dropped, constant-id re-arm preserved.
Sabotage-verified: tests fail with the old single-id tracking.
The tool_call id-matching pass in repair_message_sequence only read
`.get("id"/"call_id")` on plain dicts, skipping non-dict tool_calls
entirely (`if not isinstance(tc, dict): continue`). Host-fed and
pre-serialization histories can carry unserialized SDK tool_call
objects (e.g. `ChatCompletionMessageToolCall`) instead of dicts, which
left `known_tool_ids` empty for that assistant turn. The following
`tool` message — a legitimate result already produced by executing the
tool — was then misclassified as an orphan and silently dropped,
corrupting the persisted conversation history and leaving the
assistant's tool_calls unanswered (itself a trigger for HTTP 400 on
strict providers).
Fix: extract id/call_id via getattr() for non-dict entries too,
mirroring AIAgent._get_tool_call_id_static's existing dict-or-object
tolerance, instead of skipping them.
`repair_message_sequence` registers BOTH `id` and `call_id` for each
assistant tool_call, because a matching tool result may be keyed on either
depending on which path built it (#58168). The duplicate guard added for
dropped rather than replayed.
Those two behaviours don't compose: a Codex/Responses tool_call registers
two DIFFERENT ids (`fc_...` and `call_...`), but only the id the first
result referenced is discarded. Its sibling stays in `known_tool_ids`, so a
duplicate result keyed on that sibling still matches and is kept — two tool
messages replayed for one call, which is exactly the HTTP 400 on strict
providers the consume step exists to prevent.
Duplicates of this kind come from the retry / crash / session-resume glitch
the guard was written for; the id-variant split just lets them slip past it.
Track each registered id back to its tool_call's full variant set and
discard all of them on a match. Results keyed on either variant are still
accepted (no false orphaning), and two parallel Codex calls answered via
different variants both survive.
Adds regression tests for the sibling-keyed duplicate and for the
two-calls/mixed-keys case that must NOT be affected.
Follow-up to the salvaged #70734 fix:
- test_sanitize_dedup_drops_tool_calls_key_when_all_removed encoded the old
global-uniqueness assumption (its second assistant call reused the id AFTER
the first call was answered, which is now a legitimate new call). The
replayed call now precedes the result, making it a true duplicate of a
still-outstanding call, preserving the intended empty-tool_calls key-drop
coverage from #64335.
- New test: Hermes' own deterministic local counter ids repeating across
turns (the #76632 scenario) survive sanitization.
- New test: the 50-step constant-id field repro from #70724 (Kimi K3 /
llama.cpp) — stock main kept 1/50 tool results, now 50/50.
The #58327 dedup passes treat a repeated tool_call_id as garbage from a
retry/crash/resume glitch and drop it. That assumes tool_call_id is
globally unique, which it is not: llama.cpp emits a single constant id
for every tool call it ever returns (verified — three separate
completions from one server all carried the same id).
Under a seen-once-drop-forever rule, the SECOND legitimate tool result
of such a session looks like a duplicate and is deleted. From the second
tool call onward the model never sees any result: it announces its next
action, the turn ends, and the task is left unfinished. Bisected to
dba585c17 over a 2258-commit range; reproduced live on v0.19.0 (1/6 runs
completed a 4-step file task, vs 20/20 on the last release before that
commit, same model and server).
Key off OUTSTANDING calls instead of every id ever seen. Both original
protections are preserved: a replayed result still answers no pending
call and is still dropped, and duplicate tool_calls sharing an id within
one assistant message are still collapsed. A genuine new call that
reuses the id re-arms it first.
repair_message_sequence needs no change — it already resets its id set
per assistant message, so only the final pre-API pass mis-fires.
Live result after the fix: 8/8 runs complete, 17-26s each (was 1/6 with
runs hitting a 150s ceiling).
The dedup pass in sanitize_api_messages (introduced by #58327) can
produce an empty tool_calls array when all tool_call_ids in a message
are duplicates of earlier messages in a long conversation history.
DeepSeek v4 and newer OpenAI reject empty tool_calls with HTTP 400:
'Invalid messages[N].tool_calls: empty array'.
When kept_tcs is empty after dedup, drop the tool_calls key entirely
instead of writing tool_calls: [].
Fixes#64335
Re-derivation of #23254 (@devsart95) on today's flush loop. The turn
flush in _flush_messages_to_session_db wrote one BEGIN IMMEDIATE
transaction per message row; a typical agent turn (user + assistant +
tool results) paid 3-8 transactions -- and, off WAL (the default on
macOS while the WAL-reset guard is active), 3-8 fsyncs -- per turn.
Adds SessionDB.append_messages_batch: same row shape as append_message
(shared _prepare_message_row serializer + _MESSAGE_INSERT_SQL column
list, so the two writers cannot drift), same compression-lock and
compression-closed guards, one aggregated session-counter UPDATE, one
transaction for the whole batch. Row serialization stays outside the
write lock.
The flush loop now collects the turn's new rows and writes them in one
call. All-or-nothing pairs exactly with the persisted-marker stamping:
on failure no rows landed and no markers were stamped, so the next
flush re-writes the whole tail (same recovery contract as before,
minus the partial-prefix case that could double-count).
Measured (same harness, 5-message turn, journal_mode=DELETE,
synchronous=FULL): 2.32ms -> 0.83ms median per turn flush (64% faster,
5 fsyncs -> 1). On WAL the win is smaller but the atomicity fix holds.
The concept 'never send a turn that strict wire validation rejects as
empty' was forked across four sites, each with its own predicate and its
own blind spots:
1. build_assistant_message write-time ' ' pad — broke codex commentary
turns (content:'' is a designed state), and a DB-side pad can't
survive _rows_to_conversation's whitespace strip anyway. REMOVED.
2. conversation_loop send-time ' ' pad — main-loop only (summary path
uncovered), ordering-fragile (had to run after whitespace
normalization), assistant-only. REMOVED.
3. stream-stub '[response interrupted]' substitution — defeated the
loop's empty-stub guard (the stub no longer looked empty, entered
history, and the placeholder leaked into the stitched final
response via truncated_response_parts). REMOVED.
4. repair_empty_non_final_messages in sanitize_api_messages — the
unconditional pre-send chokepoint shared by the main loop AND the
summary path, covers user and assistant turns, non-final only,
copy-on-write. This is now the SINGLE OWNER.
The owner's payload predicate (_msg_has_payload) is extended to treat
codex_message_items / codex_reasoning_items as payload, so
designed-empty codex commentary turns are never rewritten on any
api_mode — the failure shape that broke site 1 in CI is encoded in the
owner, not special-cased at a call site.
Tests updated to pin the new contracts: builder stores textless turns
as-is; the empty stream stub stays recognizably empty for the loop
guard; poisoned resumed histories are repaired to the placeholder at
the send boundary; codex item carriers are never rewritten.
Sabotage-verified: unwiring the owner fails 3 regression tests.
Third layer of the empty-stub fix: full self-recovery. A poisoned transcript
(empty assistant stub or empty user turn already persisted before the write-
time guard, or fed in from a host history) previously 400'd every subsequent
request until it scrolled out — needing a manual DB edit + gateway restart.
sanitize_api_messages() (the unconditional pre-send chokepoint) now repairs
empty non-final messages on the per-call copy by substituting a minimal
'[response interrupted]' placeholder, so the session recovers itself IN MEMORY
on the very next send. The final message is left untouched (empty final
assistant is legal); stored history is never mutated; reasoning-only and
tool_call turns are preserved (negative controls).
Tests: production-shape repro (tool -> empty assistant -> user), empty-user
case, non-destructive guarantee, and negative controls. RED verified by
disabling the wire-in.
Strict providers (DeepSeek) reject a payload where the same tool_call_id
appears more than once with HTTP 400 'Duplicate value for tool_call_id'.
The issue was filed as an 'orphaned tool message' compression bug, but the
pasted error is a DUPLICATE tool_call_id — orphans are already handled on
main; duplicates were not. Reproduced live on main: both shapes leaked
through repair_message_sequence and sanitize_api_messages.
Two chokepoints, two shapes:
- repair_message_sequence: consume the id from known_tool_ids on first
match so a SECOND tool result reusing it falls into the drop branch
(duplicate tool-result shape). This is @Robinlovelace's kernel from
#55436 (applied manually — that PR was ~800 commits stale and bundled
an unrelated duplicate-DB-write change for #860, which is dropped here).
- sanitize_api_messages (final pre-API pass): add a dedup pass covering
BOTH (a) duplicate tool_calls sharing an id WITHIN one assistant message
(the message[6] shape) and (b) later tool result messages reusing an
already-seen id. #55436 covered neither of these at this chokepoint.
Tests: duplicate-tool-result dedup at both functions, duplicate-assistant-
tool_call-id collapse, and a negative control proving distinct ids are
never dropped (no over-dedup).
Credit: @Robinlovelace (#55436) for the repair_message_sequence dedup kernel.
Closes#58327.
repair_message_sequence Pass 1 registered only tc.get("id") when building
the set of known assistant tool_call ids, then matched tool results against
it by tool_call_id. In the Codex Responses format an assistant tool_call
carries both id (fc_...) and a distinct call_id (call_...); a tool result's
tool_call_id may be keyed on either depending on which builder produced it.
Registering only id made a valid tool result whose tool_call_id matched
call_id look orphaned, so the pass dropped it and left the assistant
tool_call unanswered -- producing HTTP 400 on strict providers (DeepSeek,
Kimi): 'Messages with role tool must be a response to a preceding message
with tool_calls'. Long-running sessions that persisted such a sequence were
permanently broken, re-sending the orphan every turn.
Register both id and call_id for each assistant tool_call so a result
matching either key is recognized, consistent with
AIAgent._get_tool_call_id_static and the compressor's _sanitize_tool_pairs.
Apply the same call_id||id precedence to the corrupted-args sanitizer's
existing-result scan / stub insertion, which had the identical mismatch.
Adds 3 regression tests covering the codex id!=call_id case (match on
call_id, match on only call_id, match on id when both present).
* fix(agent): merge consecutive assistant messages in repair_message_sequence
Strict OpenAI-compatible providers (DeepSeek v4, Moonshot/Kimi) reject a
replayed history where an assistant message carrying tool_calls is
immediately followed by another assistant message instead of its tool
results — HTTP 400 'An assistant message with tool_calls must be
followed by tool messages...'.
repair_message_sequence (the defensive belt run before every API call)
fixed orphan-tool and consecutive-user shapes but never merged
consecutive assistant messages. Adds a Pass 0 that collapses adjacent
assistant turns into one — union of tool_calls, concatenated content,
carried reasoning_content — covering both reported shapes:
- parallel tool calls split across two assistant turns (#29148)
- content-only assistant followed by tool_calls-only assistant (#49147)
A tool result or user turn between two assistants blocks the merge
(distinct, valid rounds). Runs before Pass 1 so the merged union of
tool_call ids is known to the orphan-tool filter.
Closes#29148, #49147.
Co-authored-by: Bartok9 <danielrpike9@gmail.com>
Co-authored-by: woaini30050 <woaini30050@users.noreply.github.com>
Co-authored-by: weidzhou <weidzhou@users.noreply.github.com>
* fix(agent): exempt codex Responses interim turns from assistant merge
The Pass 0 consecutive-assistant merge collapsed codex_responses interim
turns, which legitimately stay separate — each carries its own encrypted
continuation state (codex_reasoning_items / codex_message_items) that
must replay verbatim. Skip the merge when either side is a codex interim
(has codex_reasoning_items / codex_message_items / finish_reason=='incomplete').
Fixes the slice-2 regression in test_run_agent_codex_responses.py
(test_duplicate_detection_distinguishes_different_codex_{reasoning,message_items}).
---------
Co-authored-by: Bartok9 <danielrpike9@gmail.com>
Co-authored-by: woaini30050 <woaini30050@users.noreply.github.com>
Co-authored-by: weidzhou <weidzhou@users.noreply.github.com>
Follow-up to the #44837 clamp: a min() clamp only fixes cursor overshoot
past the new end of the list. When repair_message_sequence drops/merges
messages at indexes below the cursor, the clamp leaves the cursor pointing
past unflushed rows and the turn-end flush silently skips them.
Extract repair_message_sequence_with_cursor(): snapshot the flushed prefix
by object identity before repair, then recompute the cursor as the count
of surviving flushed messages. Falls back to the clamp when no snapshot is
available. Keeps the safety guard in _flush_messages_to_session_db.
Adds targeted tests for overshoot, before-cursor compaction, no-repair,
bare-agent, and the flush guard.
When empty-response terminal scaffolding fires on a tool-result turn,
_drop_trailing_empty_response_scaffolding left the live history ending at
a bare 'tool' message. The next user input then landed as [...tool, user],
a protocol-invalid sequence that OpenRouter/Opus and other providers
silently fail on (returns empty content). That retriggered the empty-retry
recovery every turn, and recovery flags never hit SQLite (no column for
them), so history kept looking broken on every reload.
Two fixes:
1. Scaffolding strip rewinds the orphan assistant(tool_calls)+tool pair
after popping sentinels. Only fires when scaffolding flags were
actually present, so mid-iteration tool loops are untouched.
2. _repair_message_sequence runs right before every API call as a
defensive belt: drops stray tool messages with unknown tool_call_ids,
merges consecutive user messages so no user input is lost. Does NOT
rewind assistant(tool_calls)+tool+user — that pattern is valid when
the user redirected before the model got its continuation turn.
Repro: session 20260507_044111_fa7e65. Opus-4.7/OpenRouter returned
content-less response after a 42KB execute_code output, nudge+retry
chain exhausted (no fallback configured), terminal sentinel appended,
scaffolding stripped leaving bare tool tail, user typed 'wtf happened..'
and landed as tool→user violation. Every subsequent turn collapsed in
<50ms with the same 3-retry empty chain because the API request itself
was malformed.
Verified live via HTTP mock: pre-fix reproduced 5 api_calls/0.15s exit
'empty_response_exhausted'; post-fix 1 api_call/0.10s exit
'text_response(finish_reason=stop)'. Three-turn session flows cleanly
through the scenario. Full run_agent suite: 1242 passed (0 regressions,
2 pre-existing concurrent_interrupt failures unrelated).