Commit Graph

1002 Commits

Author SHA1 Message Date
kshitijk4poor 80ab7d2b1c fix(compression): dedupe current-turn rows when rotation splits the session mid-turn
When context-compression rotation fires mid-turn, the current user
message was persisted twice into the child session. Root cause: dedup
used id()-seeded sets of copies instead of markers on the live objects.

Replace with _DB_PERSISTED_MARKER-based dedup as the sole authority:
- _ensure_compressed_has_user_turn returns CompressedUserTurnOutcome
- After publish_compression_child succeeds, stamp the live anchor-source
  row (not a drifted index) with _DB_PERSISTED_MARKER
- _sync_persisted_markers mirrors stamps from result to live lists by
  scoped identity (handles direct-path, adoption divergence, _session_messages)
- Remove _flushed_db_message_ids from rotation commit path (markers replace it)
- Unconditional (loud) imports — no silent fallback

Salvage of #94996 by @fedosis, rebased on top of #95433 (stall-fallback,
already merged). Both conversation_compression.py and run_agent.py are
built from origin/main + #94996's diff applied on top, preserving the
force_terminal refactor and _publish_new_fence from #95433.

Credit: @fedosis original PR #94996.
2026-08-28 02:47:34 +05:30
Uttkarsh Tiwari 36ae32524c fix(compression): preserve terminal lifecycle for lock skips 2026-08-27 12:42:38 +05:30
Uttkarsh Tiwari 7a21bfe68a fix(compression): suppress duplicate completion notices 2026-08-27 12:42:38 +05:30
Teknium f4df86fe1a test: adapt summary-continuity + rotation-flush fixtures to the lean default
Continuity tests pin tail_mode=legacy (they assert the raw LLM text
terminates the stored summary; lean's verbatim-user appendix follows it by
design and the contract under test is mode-independent). The #57491
rotation fixture grows 200→2000 chars/message: at ~2.5K total tokens the
old fixture fit entirely inside lean's 10K tail floor, so the no-growth
guard correctly refused the rotation the test exercises.
2026-08-26 07:16:04 -07:00
Teknium 4032a15ad0 refactor(prompt): remove the ~1.2K-token Nous Subscription block from the system prompt (#95005)
* refactor(prompt): remove the Nous Subscription block from the system prompt (~1.2K tokens/call)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-08-25 12:59:29 -07:00
Teknium 9e551d2931 refactor(memory): renumber checkpoint API — v1 is the implicit historical contract, v2 opts into fail-closed checkpoints
Per review: existing providers should not be retroactively re-versioned or
handed a changed payload. Version 1 is now the implicit historical
on_pre_compress() contract (best-effort, raw message list) that every
pre-existing provider is already on; the fail-closed checkpoint contract
becomes version 2. MemoryManager routes the raw transcript to v1 providers
unchanged and hands the host-normalized evidence list only to v2+ checkpoint
providers, so the plugin surface contract for shipped providers is
byte-identical with the gate off.
2026-08-25 03:55:55 -07:00
Teknium a70d2ffce5 fix(background-review): teach review prompts the enforced read-before-write handshake
The skill_manage guard (added in #55906) refuses any patch/edit of an
existing SKILL.md, or overwrite/removal of an existing support file,
unless the exact target was loaded via skill_view during the review.
Neither _SKILL_REVIEW_PROMPT nor _COMBINED_REVIEW_PROMPT ever mentioned
this, so models routinely issued the write without the pre-read, got
refused, and burned review iterations (#62397).

Both prompts now carry a Read-before-write section scoped to the
guard's actual contract: existing targets only, exact-path pre-read for
support files, transcript quotes don't count, new skills/new support
files exempt, and a bounded one-view-one-retry recovery instead of a
loop. Direction follows #60331 by @kkwills13 with the scope corrections
requested in review (existing-target-only wording, no delete claim,
bounded retry, contract tests for both prompt variants).

Fixes #62397.
2026-08-25 00:18:35 -07:00
Ailirag 5908c577f9 fix(fallback): surface provider transitions and primary recovery 2026-08-25 12:12:08 +05:30
kshitij 6e534df114 Merge pull request #93773 from kshitijk4poor/fix/codex-sdk-transform-bypass-93650-v2
fix: route codex payloads around the SDK's GIL-holding request transform (#93650)
2026-08-24 16:11:18 +05:30
Teknium 1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
kshitijk4poor 9975544101 style: ruff format test_codex_sdk_transform_bypass.py 2026-08-24 15:30:30 +05:30
kchernev 10a070bd49 fix: route codex payloads around the SDK's GIL-holding request transform (#93650)
responses.create re-walks the entire request body against the
ResponseCreateParams union graph client-side while holding the GIL.
#93650 documents that walk wedging for 12+ hours on a ~1.4 MB
conversation, starving every other thread including the TTFB/stale
watchdogs — and no socket kill can unblock a pre-network hang.

Hermes payloads are JSON round-trips and already wire format, so the
bulk fields (input, tools) are now routed through extra_body, which the
SDK merges into the JSON body after the transform. Guarded by a
plain-JSON check (anything else keeps the typed path) and a
HERMES_CODEX_SDK_TRANSFORM=1 escape hatch. Applied to both the primary
stream path and the auxiliary adapter.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-24 15:29:45 +05:30
kshitijk4poor f93b350711 fix: align cache-policy pre-gate identity with the capability matcher
Follow-ups on top of the salvaged #92785 commit:

- Pre-gate now matches base URLs via normalize_route_base_url and
  provider ids via custom_provider_aliases, mirroring the semantics of
  get_custom_provider_model_capability. The raw string comparison
  silently dropped declarations whose config spelling differed only by
  host case or trailing slash (proven empirically: …/v1/ vs …/v1 with a
  non-matching provider name returned (False, False) despite an explicit
  prompt_caching: true).
- get_provider(..., allow_network=False) in the early-init/stub branch:
  the policy runs per request destination (MoA aggregator, auxiliary
  replans via blank_cache_policy_stub, early agent init) and a cold
  models.dev cache triggered a measured ~450 ms foreground registry
  fetch from the send path. A catalog miss degrades to the conservative
  side.
- Debug-log the previously silent provider-lookup exception fallback.
- Tests: _make_agent defaults _custom_providers=[] (post-init reality;
  keeps built-in-route tests off the catalog/config fallback), the two
  early-init tests delete the attr explicitly, and three regression
  tests pin the URL-drift, spaced-legacy-name, and no-network contracts
  (all three fail on the unfixed commit).
2026-08-24 14:58:42 +05:30
Blood Shot 0204e4898e fix(agent): normalize custom provider route identity 2026-08-24 14:58:42 +05:30
Blood Shot 0a3b7efec5 fix(agent): honor prompt_caching for custom providers
Apply explicit per-model prompt_caching capabilities to custom
chat-completions routes, rather than limiting them to recognized providers,
hosts, or model families.

Keep undeclared routes conservative, derive the marker layout from the wire
transport, and leave Responses and Bedrock caching paths unchanged.
2026-08-24 14:58:42 +05:30
Gille 21b92d2687 fix(agent): bypass response cache for empty retries 2026-08-23 23:31:54 -07:00
fangliquanflq 37411f349a fix(auth): rotate credentials for named custom providers after 401/429
Salvage of #93214 (5 commits squashed onto current main; agent_runtime_helpers.py
diverged since the PR base and was 3-way reapplied). The credential-rotation
guard in recover_with_credential_pool and both restore_primary_runtime paths
only tolerated the custom-naming split when the agent carried the literal label
'custom', so a named custom provider (agent.provider='gemini-no-filter', pool
'custom:gemini-no-filter') tripped the mismatch guard and skipped rotation on
every 401/429. Now all three guard sites use the canonical
credential_pool_matches_provider boundary predicate + resolve_runtime_pool_key,
which recognizes configured named-custom aliases and validates endpoints.

Fixes #93188.
2026-08-23 20:01:18 -07:00
aniruddhaadak80 cd6c088928 test(compression): align no-op strike tests with structural backoff (#93022)
Two suites still encoded the pre-#93093 contract that the three
structural no-op branches (insufficient_messages, no_compressible_window,
empty_post_handoff_window) increment _ineffective_compression_count:

- tests/agent/test_compaction_anti_thrash.py::
  TestMinimumMessagesBranch::test_too_few_messages_records_an_ineffective_pass
- tests/run_agent/test_infinite_compaction_loop.py::
  TestCompressNoOpRegistersIneffective::{test_no_op_increments_counter,
  test_two_no_ops_block_should_compress}

Structural no-ops are transcript-shape facts, not evidence of an
incompressible floor, so they now arm _structural_no_op_backoff_until
and leave the strike counter untouched. Update the tests to pin the new
contract (count unchanged, backoff armed via time.monotonic(),
should_compress blocked while it holds) and rename accordingly. The
outcome contract of test_two_no_ops_block_should_compress is preserved:
repeated no-ops still block further automatic compression.
2026-08-23 18:27:07 -07:00
Finn763 4202a508fd fix(review): bound same-model background review replay
Detach the review fork's compressor from the parent SessionDB/session_id
and re-enable in-memory-only compaction for oversized snapshots, instead
of the historical compression_enabled=False guard that left the fork's
replayed transcript unbounded (350k-384k input tokens per request, 1.49M
total across one 8-request review). Add an aggregate input-token budget
(auxiliary.background_review.max_input_tokens, default 600k) so repeated
tool calls cannot recreate an unbounded transcript; the tool loop stops
before the provider call that would cross it.

Closes #93057
2026-08-23 18:25:19 -07:00
fangliquanflq 2033f4cc34 fix(agent): separate cancellation diagnostics from tool output 2026-08-23 18:25:19 -07:00
fangliquanflq c1c0efa375 fix(code-exec): preserve interrupt cancellation source 2026-08-23 18:25:19 -07:00
Teknium a2a43f7e82 fix(agent): widen composite-id alias matching to the compressor; unify variant policy owners (#63000)
Follow-up on top of the salvaged #93335:

- context_compressor._sanitize_tool_pairs now expands alias spellings on
  the RESULT side too (tool_result_id_variants), so a composite
  call|item-keyed result pairs with its split-field tool_call instead of
  being dropped and its call stripped.
- The compressor's _tool_call_id_variants staticmethod and
  agent_runtime_helpers' module-level _tool_call_id_variants are now thin
  forwarders to agent.message_sanitization.tool_call_id_variants — one
  policy owner for alias expansion, so the pre-call sanitizer, repair
  pass, dedup pass, and compression sanitizer can never drift apart.
- Preserved the #91768 SDK-object tolerance in repair pass 1 (the
  shared helper handles non-dict tool_calls via getattr; the salvaged
  commit's isinstance-dict guard was dropped in the merge resolution).

New regression tests: composite-keyed results through
sanitize_api_messages (both directions) and _sanitize_tool_pairs, with
negative controls. Sabotage-verified: compressor test fails with raw
tool_call_id tracking.
2026-08-23 18:24:43 -07:00
joaomarcos 5496d5995a fix(agent): preserve tool results across ID variants
Match Responses/Codex tool-call aliases across execution, repair, sanitization, replay, and duplicate handling so valid parallel results are not replaced by unavailable stubs.\n\nFixes #93251
2026-08-23 18:24:43 -07:00
honor2030 fe483de4d3 fix(agent): keep max-iteration warnings out of quiet stdout
Route the max-iterations diagnostic through logging when quiet_mode is active so automation wrappers keep stdout machine-readable.

Add a regression test covering quiet max-iteration summary handling.
2026-08-23 17:45:58 -07:00
Teknium faa2399e2b fix(agent): make the pre-call dedup pass variant-aware; widen batch regression coverage (#93251)
Follow-up on top of the salvaged cluster: sanitize_api_messages step 3
(duplicate tool_call_id dedup) still tracked only the coalesced
(call_id||id) value in outstanding_call_ids, so after step 2's
variant-aware matching preserved a result keyed on the OTHER id variant,
step 3 deleted it as answering no outstanding call — whole parallel
batches of real results vanished with no stub at all (#93251's total-loss
mode). Track the full variant set per call and consume all siblings when
answered, preserving #58327 duplicate protection and llama.cpp
constant-id re-arm semantics.

Also aligns the #58287 compressor test with the in-flight tool chain
protection (#79278) that landed after that PR was opened: a trailing
user turn keeps the negative-control assistant message out of the
protected trailing window.

New regression tests: divergent-id batch survival through the dedup
pass, sibling-id replay still dropped, constant-id re-arm preserved.
Sabotage-verified: tests fail with the old single-id tracking.
2026-08-23 17:01:20 -07:00
joaomarcos 36b4da5489 fix: repair_message_sequence drops tool results for SDK tool_call objects
The tool_call id-matching pass in repair_message_sequence only read
`.get("id"/"call_id")` on plain dicts, skipping non-dict tool_calls
entirely (`if not isinstance(tc, dict): continue`). Host-fed and
pre-serialization histories can carry unserialized SDK tool_call
objects (e.g. `ChatCompletionMessageToolCall`) instead of dicts, which
left `known_tool_ids` empty for that assistant turn. The following
`tool` message — a legitimate result already produced by executing the
tool — was then misclassified as an orphan and silently dropped,
corrupting the persisted conversation history and leaving the
assistant's tool_calls unanswered (itself a trigger for HTTP 400 on
strict providers).

Fix: extract id/call_id via getattr() for non-dict entries too,
mirroring AIAgent._get_tool_call_id_static's existing dict-or-object
tolerance, instead of skipping them.
2026-08-23 17:01:20 -07:00
Frowtek b9a62f6590 fix(agent): consume every tool_call id variant when pairing tool results
`repair_message_sequence` registers BOTH `id` and `call_id` for each
assistant tool_call, because a matching tool result may be keyed on either
depending on which path built it (#58168). The duplicate guard added for
dropped rather than replayed.

Those two behaviours don't compose: a Codex/Responses tool_call registers
two DIFFERENT ids (`fc_...` and `call_...`), but only the id the first
result referenced is discarded. Its sibling stays in `known_tool_ids`, so a
duplicate result keyed on that sibling still matches and is kept — two tool
messages replayed for one call, which is exactly the HTTP 400 on strict
providers the consume step exists to prevent.

Duplicates of this kind come from the retry / crash / session-resume glitch
the guard was written for; the id-variant split just lets them slip past it.

Track each registered id back to its tool_call's full variant set and
discard all of them on a match. Results keyed on either variant are still
accepted (no false orphaning), and two parallel Codex calls answered via
different variants both survive.

Adds regression tests for the sibling-keyed duplicate and for the
two-calls/mixed-keys case that must NOT be affected.
2026-08-23 17:01:20 -07:00
Bartok9 1a83b1e588 fix(agent): keep tool results keyed on a tool_call's id variant (#55626)
Register every id variant (call_id AND id) of each assistant tool_call in
sanitize_api_messages so a tool result keyed on either variant is treated
as paired. Previously only the coalesced (call_id||id) value was
registered, so Responses-style tool_calls carrying divergent id (fc_...)
and call_id (call_...) had their real results dropped as orphans and
replaced with '[Result unavailable]' stubs.

Cherry-picked from PR #56148 (unrelated busy_ack_templates files dropped
per the author's own follow-up commit).
2026-08-23 17:01:20 -07:00
kshitijk4poor b8cd00f968 fix: set failed=True for repeated_outer_errors exit + drop append_message
Follow-up to @BrunoBza's #93062:

1. Set failed=True only for the new repeated_outer_errors exit reason.
   Previously the error exit left failed=False, so finalize_turn reported
   completed=True for a turn that actually failed — incorrect.

2. Don't append_message the assistant response at the break. A thinking-
   prefill or interim assistant may already be the tail, and appending
   would create assistant→assistant role-alternation violation.
   finalize_turn (lines 341-353) handles this safely by checking
   _tail_role != 'assistant' before appending.

3. Update test to assert failed=True and completed=False for the
   repeated_outer_errors exit.
2026-08-24 04:47:07 +05:30
BrunoBza 56e7fd2adf fix(loop): bound outer-loop error retries per turn instead of relying on max_iterations (#92450)
The outer conversation-loop except handler only left the loop on a
local-processing error or when api_call_count >= max_iterations - 1.
With the turn budget now unlimited by default (sys.maxsize), a
permanent failure that escaped the inner retry/fallback machinery
retried forever: ~64 retries/s, one core pegged, and the rotated
agent.log history overwritten within minutes.

Bound the loop with a small per-turn cap on total escaping exceptions
(_MAX_OUTER_LOOP_ERRORS = 8, scaled down by a tiny explicit
max_iterations so a manually bounded budget still governs). The legacy
local-processing and near-limit exits are byte-identical; a new
'repeated_outer_errors' exit reason gets a user-facing explanation.

The inner retry/fallback layer owns transient API recovery and
terminates on its own, so only exceptions that escape it reach this
cap - a successful turn is unaffected.

Fixes #92450
2026-08-24 04:47:07 +05:30
Teknium 9ea7fe9938 fix: quitting the CLI no longer spams shutdown-race API errors onto the shell
When the TUI exits while the post-turn background review fork is still
mid-request, every further API attempt raises 'cannot schedule new
futures after interpreter shutdown'. The conversation loop treated this
as a retryable API error: un-gated ❌ prints leaked onto the user's
shell AFTER the TUI exited (call #4, #5, #6...) and the loop retried a
doomed request until the interpreter froze the thread.

Fix the class, not the site:
- tools/interpreter_shutdown.py: single shared shutdown predicate
  (matches both CPython message variants + sys.is_finalizing()).
- cron/scheduler.py, agent/tool_executor.py: existing per-site
  predicates now delegate to the shared home (tool_executor previously
  matched only the fuller variant).
- agent/conversation_loop.py: inner retry handler recognizes the
  shutdown signal and abandons the turn — one log warning, no print,
  no traceback, no debug dump, no retry; outer handler gets the same
  guard for shutdown errors raised outside the API call.
- The outer handler's bare print() now honors suppress_status_output
  (set by the background-review fork) instead of bypassing it.

Refs #55924 #58720 (same class in cron delivery), adjacent to #90683.
2026-08-23 16:04:41 -07:00
Teknium fd760435c6 test(bedrock): make the botocore stub windows airtight — kills the vendored-import flake
CI flake mechanism (PR #92617 red, reproduced standalone): tests plant
fake botocore modules via patch.dict; when the REAL botocore.exceptions
is first imported in an interpreter state where a fake parent is (or
was) installed, its 'from botocore.vendored import requests' resolves
against a module with no __path__ and every exception test in the worker
dies with "No module named 'botocore.vendored'" — ordering-dependent,
so green locally, red in CI workers.

Defenses (both, in depth):
- test_bedrock_adapter.py pre-imports the real botocore.exceptions at
  module scope, before any test can stub sys.modules — later imports are
  cache hits that can never re-execute the vendored import under a
  poisoned parent. Proven standalone: fake-parent repro fails without
  the pre-import, succeeds with it.
- autouse _boto_sys_modules_hygiene fixtures in all three files that
  plant fake boto* modules (adapter, integration, model-picker):
  snapshot every boto* sys.modules entry before each test, evict+restore
  after — no stub window can leak state into a later test regardless of
  worker ordering.
- importorskip targets botocore.exceptions (the module the tests
  actually need) instead of bare botocore, so a torn install skips
  instead of erroring.

148/148 across the four affected suites.
2026-08-22 19:30:10 -07:00
Teknium 42dd219f46 test: trim salvage of #65076 to a lean regression set
Drop the bulk test additions from the original PR; keep only mandatory
picker-assertion adaptations (Mantle IDs join the discovery lists), one
allowlist routing test covering all four Mantle model IDs, the 272K
context check, and the two review-mandated auxiliary regressions
(config-region-beats-env for the Mantle path, aux Responses client).
2026-08-21 15:02:29 -07:00
Nathaniel Branscum e57d55fc7a fix(moa): keep Bedrock slots on provider runtime
Preserve the Bedrock provider identity for MoA reference and aggregator slots so Bedrock OpenAI Responses models use the aws_sdk/SigV4 runtime instead of being downgraded to a generic custom endpoint. Add regression coverage for Bedrock GPT-5.5 MoA slots.
2026-08-21 15:02:29 -07:00
poisdahl 13fcf2fe38 Merge remote-tracking branch 'origin/main' into agent/81234-merge-20260821 2026-08-21 16:02:59 +02:00
poisdahl abf87e7248 Merge current main into composite-carrier fix 2026-08-21 15:56:45 +02:00
JoaoMarcos44 fb27614add fix(native_compaction): preserve compression summary messages during pre-checkpoint pruning
prune_pre_checkpoint_items() had a hardcoded role=='user' filter that
discarded all non-user messages before a checkpoint — including Hermes'
own compression summaries (role='assistant'), causing total context amnesia
about past conversation summaries.

The fix:
- _is_summary_item delegates to the canonical
  agent.context_compressor.is_compaction_summary_message provenance check
  (not an ad-hoc heuristic)
- Summaries are retained whole (never byte-sliced) within a 32k token budget
- Idempotent across repeated checkpoints (dedup by identical text)
- _chat_messages_to_responses_input threads item_sources (raw chat messages)
  through to the pruner, so it can read summary content directly from the
  source when the Responses conversion shape is lossy (tool-result carrier
  becomes function_call_output, or stale codex_message_items replay shadows
  merged content)

Fixes #90975.

Salvage of #90976 by @JoaoMarcos44.
2026-08-21 17:24:22 +05:30
kshitijk4poor b883756b79 fix: foreground priority for background review cancel timeout
Change fail-closed behavior to proceed-with-warning when a background
review does not acknowledge cancellation within the bounded deadline.
The review is non-critical self-improvement work and must never block
a user-facing turn (#84423). Keep the off-thread interrupt to ensure
a broken abort path cannot stall the bounded wait.
2026-08-21 16:12:57 +05:30
qixuancao 1b92a94962 refactor(agent): simplify background review run state 2026-08-21 16:12:57 +05:30
qixuancao 37da0d4d50 fix(agent): synchronize background review cancellation 2026-08-21 16:12:57 +05:30
Teknium ca06b87689 feat: opencode-free is fully keyless — no env var, no account, anonymous wire
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).

On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
  never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
  Zen endpoint routing incl. muse->responses); keyless predicate extended
  with unsuffixed free slugs (big-pickle); free runtime pins EVERY
  opencode-free model keyless; curated catalog replaces the models.dev
  cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
  there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
  wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
  invariant (every curated model must satisfy the keyless predicate)

E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
2026-08-21 00:24:32 -07:00
Rudraksh Chahal 28a9b6c565 feat(providers): add OpenCode Free provider with keyed auth and opencode User-Agent
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.

The free tier requires a real account API key and throttles third-party
clients by User-Agent:

- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
  and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
  Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
  unconditionally (the stale keyless-tier assumption), and credential-pool
  exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
  message.

Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
2026-08-21 00:24:32 -07:00
kshitijk4poor c26357ad6a refactor(prompt_caching): shallow strip copy, exact-count guards, dedupe idempotency tests
Follow-up to the #90972 salvage:

- strip loop: copy.deepcopy(msg) -> dict(msg). strip_anthropic_cache_control
  is copy-on-write on content parts by contract (pops the top-level key,
  rebuilds content lists/part dicts fresh), so a shallow top-level copy
  preserves the caller-non-mutation guarantee — verified for all four
  marker shapes — and removes the redundant second deepcopy the re-mark
  path paid on already-decorated input. Docstring updated to match.
- tests: moved the surviving idempotency tests into
  tests/agent/test_prompt_caching.py (where this module's tests live) as
  TestApplyIdempotency; dropped the three tests that duplicated existing
  coverage (dynamic_tool_accounting ~= TestPromptCachePlan::
  test_copies_sections_and_keeps_canonical_tools_plain which already
  asserts == 4; can_carry_marker_envelope_vs_native ~= TestCanCarryMarker;
  never_exceeds_four_markers subsumed by the idempotency test).
- exact-count assertions per review: idempotency fixture pins == 4,
  no-tools fallback pins == 3 (marker loss can no longer masquerade as
  safety); added the one new _can_carry_marker assertion (native=True
  empty assistant) to TestCanCarryMarker.
- new part-level stale-marker mutation guard (the other detection branch,
  where part-dict aliasing is the risk); fails on pre-fix base with
  marker accumulation (9 > 4), passes with the fix.
2026-08-21 12:23:47 +05:30
joaomarcos 0fc52b055f fix(prompt_caching): make apply_anthropic_cache_control idempotent on pre-decorated input
apply_anthropic_cache_control never stripped pre-existing cache_control
markers before placing new ones, so calling it twice (or handing it
messages a prior call already marked) accumulated markers past
Anthropic's 4-breakpoint limit and produced HTTP 400
'cache_control can only be specified up to 4 times'.

Strip any pre-existing markers from per-message copies before marking,
mirroring the strip-then-mark pattern build_prompt_cache_plan already
uses. Only messages that already carry a marker pay the copy cost; the
copy-on-write contract (caller-owned messages are never mutated) is
preserved. Repeated calls now converge to byte-identical output.

Salvaged from #90972 by @JoaoMarcos44 (net diff of the PR's commit
stack, intermediate reverts collapsed).

Related: #90971
2026-08-21 12:23:47 +05:30
kshitijk4poor 2cf7b36e11 fix(memory): enforce independent built-in store permissions
Normalize malformed memory config during initialization and bind per-target write permissions to the session MemoryStore so direct and staged writes cannot update a disabled built-in store.
2026-08-20 20:20:23 -07:00
Brooklyn Nicholson c57581cd0d feat(tools): drive_preview and annotate_preview — the agent can use the page it opened
The in-app browser was a one-way mirror. open_preview put a page in the pane
and read_preview read its text back, but nothing could touch it. A click meant
falling back to the browser_* tools, which drive a separate Chromium the user
cannot see — so "log into this and pull my invoices" happened in a different
browser from the one on screen, with none of the sessions the user is already
signed into.

Four pieces, and they only make sense together:

  · an in-page engine that inventories what is interactable and performs the
    verb, injected as source because it has to run inside the guest page;
  · the preview.act.request bridge from the gateway into the pane;
  · drive_preview, for acting: elements, click, type, scroll, press, and the
    pane's own back/forward/reload;
  · annotate_preview, for marking without acting.

Those last two started as one tool doing two unrelated jobs. Leaving a mark is
not an action — it outlives the turn that drew it — so it gets its own verb,
and the interaction verb gets a name that says what it does.

Gating is the existing surface rule: desktop_ui folds in on session
source: 'desktop', and the bridge refuses to act for a background session, so a
turn running behind the user's back cannot reach into the page they are working
in.

Two details worth a reviewer's attention. Typing assigns through the
prototype's value setter, because React shadows value with its own accessor and
ignores an input event whose value it believes it already wrote — a plain
el.value = … types into a field that snaps back on the next render. And
clicking replays the pointer/mouse pair before activation, because frameworks
bind to mousedown as often as to click.
2026-08-20 05:26:37 -05:00
Teknium c2f5d2da21 test: vary marathon-turn fixture args — identical calls now legitimately dedupe to stubs 2026-08-20 00:16:22 -07:00
kshitijk4poor b7e12decc6 fix(agent): route relay-wrapped output-cap 429s into the output-cap handler
Salvage follow-up for #72283: instead of a second pre-retry clamp block
(which bypassed the #55546 clamp+compress path and broke its three
regression tests), parse the output cap ONCE at classification time and:
- exempt parseable wrapped output-cap 429s from the eager rate-limit
  provider fallback (a deterministic request-shape failure that failover
  cannot fix but the clamp fixes in one retry), and
- widen is_context_length_error so they reach the SAME #55546
  clamp+compress recovery as plain output-cap 400s.

Adds both #72283 regression scenarios plus an ordering guard proving a
NON-EMPTY fallback chain does not consume the wrapped 429 (fallback
slot unspent, model unchanged). 119 fallback/rate-limit tests green.
2026-08-20 11:37:01 +05:30
kshitijk4poor 0596ccdeb3 fix(compression): salvage follow-up — todo snapshot last-resort, reuse prune helpers
Review follow-up on the salvaged #90353:
- Todo snapshot (+ coupled pruned-skill reload notice, 7a16840add) is now
  reduced only as a LAST resort after reasoning/tool/summary shrink ops,
  and the reload notice survives even then.
- Reuse existing helpers/constants instead of re-hardcoding:
  _PRUNED_TOOL_PLACEHOLDER, _PRUNE_MIN_CHARS, _NEWEST_TURN_ONLY_BUDGET_KEYS,
  and _prune_stale_reasoning_replay (codex sidecar shrink, #71058 boundary).
- Assistant-role messages without the summary metadata key are no longer
  truncatable by the summary-cap heuristic.
- Caller passes budget so the estimator runs 3x, not 5x, per would-grow pass.
2026-08-20 11:36:54 +05:30
MindDragonLabs fb96247eaf fix(compression): salvage grown candidates before refusal 2026-08-20 11:36:54 +05:30