Skip checkpoint evaluation when a turn is interrupted so cancellation rows
remain durable without urging continued execution. The existing minimal-agent
interrupt regression also avoids dereferencing an absent iteration budget.
Consolidate the warning coverage into two invariants, including real SQLite
readback and dispatcher/child scope controls. Cold-start tool availability
between construction cases to model independent worker processes. Place ratio
normalization beside the existing iteration budget instead of growing init.
Real cancelled-tool A/B against current main, draft, and fix: three cancelled
rows and zero writes on all arms; persisted checkpoint notices 0 / 1 / 0.
Repeated scripted HTTP/SQLite loop A/B preserves completion opportunity,
ordinary default-off behavior, and blocked/two-failure exhaustion behavior.
Local targeted run initially passed 15 cases with one fixture cache-isolation
failure; corrected target and inherited affected suites remain queued behind
the campaign lock. This commit is not a CI-green or merge-ready claim.
Give dispatcher-owned workers a tool-capable reporting opportunity before the
hard iteration cap, without accepting arbitrary diffs or weakening failure
counting. Add opt-in per-turn iteration checkpoints for ordinary agents.
Persist checkpoint text with the fresh tool result, never rewrite cached rows.
Salvages the opt-in ratio and per-turn reset implementation from #104683;
credits the earlier default-off signpost proposal in #92438.
Local fixture wire A/B: Kanban ready/1 failure -> done/0; deliberately stuck
workers still reach blocked/2 after two runs. Default-off control unchanged.
Targeted and affected-directory suites queued behind campaign test lock.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: C. Michael Gibbs <252231331+MikeGibbsOnyx@users.noreply.github.com>
Normalize incoming reasoning at the shared heading boundary and completed
extraction, and flatten auxiliary content and reasoning before accumulation.
Reuse the existing text flattener with no implicit fragment separators.
Combine the earliest related work from zsuroy (#85791), the diagnosis and
patch from 2025hcsmile2010-hue (#104711, #104848), and completed extraction
work from liuhao1024 (#104717) as a slim redo, not a verbatim cherry-pick.
Two invariant tests exercise the real SDK and local HTTP fixture across
main streaming, Relay collection, auxiliary sync/async and completed output.
The standalone matrix improves from 32/84 to 84/84, preserving answers.
Co-authored-by: suroy <suroy@qq.com>
Co-authored-by: 2025hcsmile2010-hue <2025hcsmile2010@gmail.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Port the exact submitted-wire-text ownership boundary from #93546 onto
current topical runtime code. Do not add the candidate's mocked-result
fallback or storage-level content deduplication. Preserve later distinct
and identical user events, separate identical accepted turns, and keyless
inputs. Add two regression invariants and offline subprocess-wire A/B.
Local wire A/B: 4/8 control matrix passing on base, 8/8 after.
Broader tests queued behind campaign lock; not ready for merge.
Refs #104653
Original diagnosis: @gitszabolcs (#38254)
Original implementation: #43127, submitted by @vashkartik
Focused salvage and wire-text correction: @fancyboi999 (#93546)
Current-main carry-forward considered: #104698
Co-authored-by: Xinmin Zeng <135568692+fancyboi999@users.noreply.github.com>
Co-authored-by: VECTOR <vector.hq@outlook.com>
Extend #104477 to the native thinking, vision, metadata, and local header paths identified by #87641. Materialize only at probe boundaries; leave the chat callable and cache ownership untouched. Local-wire A/B: thinking and vision show requests change from 403 to 200, while static credentials and callable chat retain success. Target suites queued.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Two test-only corrections for the failures on run 34105317301:
- tests/agent/test_anthropic_adapter.py::TestResolveAnthropicToken: the new
skip_borrowed branch reads ``entry.source``. ``PooledCredential.source`` is a
required dataclass field that ``from_dict`` always materializes (defaults to
SOURCE_MANUAL) and ``_available_entries`` returns only PooledCredential, so a
production entry can never lack it. The three SimpleNamespace doubles were the
incomplete side; build them via ``PooledCredential.from_dict`` instead of
duck-typing production with getattr.
- tests/agent/test_auxiliary_client.py::test_stale_anthropic_fallback_refreshes_and_retries:
the PR itself now passes ``failed_api_key=<client.api_key>`` into
``_refresh_provider_credentials`` so an unrelated borrowed login never owns the
refresh; the assertion still expected the bare ``("anthropic")`` call. Give the
stale client an explicit api_key and assert the request-bound call. Main's
auxiliary changes since the PR base (ebe4e7bb44, b40998bc3c..cb1a42d33b) did not
move this call.
Port owned-before-borrowed resolution from #104624, crediting the root cause in #104622. Include the synchronous and asynchronous auxiliary fallback recovery sites: forward the failed request key so an unrelated borrowed login never owns that refresh. Local-wire probes preserve the borrowed file and exchange only the owned grant. Broader validation remains queued; do not treat this commit as ready.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: d-bow-dev <24577047+dwb1991@users.noreply.github.com>
The memory_tool schema advertises new_text as an alias for content, and
memory_tool resolves it when content is None. But the table-driven inline
executor's arg_specs (agent/inline_tool_executors.py) did not list new_text,
so _call_tool's allowlist silently dropped it: a replace call using the
documented alias reached memory_tool with both fields None and failed with
"content is required for 'replace' action." — even though the caller
supplied the value. Forward new_text alongside content/old_text so the
documented alias fires and content still wins when both are set, matching
what the batch path (op.get("content") or op.get("new_text")) already
accepts.
Regression coverage from PR #104214. Live inherited-compress probe reproduces the same failure on main; serial unit runner lock is busy.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Slim salvage of #104688: place project context before workspace state and
keep cwd outside the stable prefix. Put runtime hints behind a final
renderer-owned boundary so quoted operator, memory, plugin and embedder
examples cannot override the persisted runtime cwd or identity fields.
Retain legacy unmarked prompt validation, add two invariant tests and a
credential-free real-AIAgent/git-worktree replay harness. No provider
cache-hit or billing measurements are claimed.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Co-authored-by: HexLab98 <liruixinch@outlook.com>
Track raw task identities across an agent's turns and match them against
process owner_task_id during close. Session IDs and shared terminal keys
are not process ownership, so the old bulk cleanup missed delegated work.
Preserve parent/sibling processes and consume teardown notifications.
Move task-resource cleanup into the lifecycle mixin, add real-process
isolation regressions, and document background process lifetime.
Use request-local stream silence for the waiting notice, preserving quiet
activity heartbeats and all existing watchdog policies. Distinguish a stream
that stopped from a request with no response, and clear this request's notice
on the next poll when events resume. Existing fresh first-event retry phases
also reset the display; recovery deadlines explicitly use total call elapsed.
Add two invariant tests (eight cases), proven red on main, plus EN/ZH docs.
Local SDK SSE through classic CLI callbacks in a PTY verifies active reasoning,
true silence, and an already-visible warning clearing on resumed reasoning.
Related: #92657 addresses repeated waiting notices; its phase deduplication
still labels active streams as no response and is not incorporated here.
The image-part membership test in _payload_chars() ran on every dict in
the request, so a tool JSON Schema with a parameter named "type" (its
value is the sub-schema dict) or a multi-type value like
["string", "null"] raised TypeError: unhashable type before the
request reached the provider. Guard the membership test with
isinstance(str): structured "type" values are payload data, not
content parts, and keep the legacy walk for them.
estimate_request_context_tokens (the stale-stream / non-stream watchdog and Codex TTFB floor
estimator) did len(str(payload)) // 4 over the wire payload, so one native screenshot read as
~100K tokens and selected the 1200s giant-conversation floor while the provider's real prompt
was ~2K (#63871, #76411). Image content parts on both wire shapes (Chat image_url / Responses
input_image) now cost the per-image price learned from provider usage (agent/image_token_cost,
bound per turn; worker threads inherit it via copy_context); base64-looking STRINGS stay text
and tool schemas mentioning image are not images.
A/B, one 300KB screenshot: main estimate 100,545 -> codex floor 1200s / stale 300s; now 2,025
-> floor 0s / stale 180s.
Design and boundary (structured parts only, tools/instructions opaque, plain data-URL strings
stay text) from #76471 by @crdesign8; re-authored to reuse the learned cost instead of a second
fixed constant, and trimmed to 2 invariant tests.