feat: agent-teams part D - async expert dispatch mechanism (#391)

* fix(deps): pin openrouter below 0.11 to avoid SSE stream regressions (#373)

* fix(openrouter): address SSE stream leak by closing response iterator

* refactor(openrouter): pass through SDK args in SSE leak patch

* test(openrouter): make SSE leak tests version-agnostic across SDK generations

* fix(openrouter): remove SSE stream leak patch and update dependencies

* fix: repair interrupted tool call history (#366)

* fix: repair interrupted tool call history

Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths.

* fix: repair malformed tool calls and dedupe repair warnings

Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted
threads with syntactically invalid tool calls get synthesized error
results and are accepted by strict providers.

Preserve the originating tool call's name in the synthesized ToolMessage,
and deduplicate repair warnings per unique tool-call id via a warned set
owned by the middleware instance, since the middleware rewrites the
request but not thread state.

Document the middleware's scope versus deepagents' PatchToolCallsMiddleware
(orphan ToolMessage dropping and mid-run coverage).

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Update README.md

* Update README.md

* Update README.zh-CN.md

* fix: scrub host path from skill_manager output and guard batch install (#377)

* fix: scrub host path from skill_manager output and guard batch install

* test: tighten install leak guards to catch host path in either tier

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: log missing async-subagent tools at DEBUG, not WARNING (#378)

* fix: log missing async-subagent tools at DEBUG, not WARNING

* fix: distinguish load_subagents callers via async_swap_pending flag

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: default reasoning context for codex proxy Responses API (#380)

* fix: set langgraph and codex proxy runtime defaults

* fix: address runtime default review feedback

* fix: drop langgraph dev env defaults per maintainer review

langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment,
so the reported KeyError cannot come from this flow; the env defaults added
here were unnecessary. Scope the PR back to the codex proxy reasoning
context fix only.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* release: v0.2.4 (#389)

* fix: add support for new Anthropic models and enhance adaptive thinking tests

* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models

* fix: update version to v0.2.4 in badges, README, and project files

* fix: update Star History chart links in README and README.zh-CN

* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog

* fix: update wechat group image in assets

* refactor(runtime): centralize async bridges under an owned runtime (#376)

* feat(runtime): add application-scoped async runtime

* refactor(cli): use owned runtime for session stats

* refactor(onboard): use the owned async runtime

* docs(runtime): record async bridge ownership

* refactor(middleware): keep sync fallback synchronous

* refactor(mcp): load tools on an owned runtime

* refactor(cli): share owned runtime across entry points

* refactor(channels): make inbound sync bridge explicit

* refactor(stream): run Rich streaming on owned runtime

* chore(runtime): remove nest-asyncio dependency

* refactor(asyncio): require active loops in async code

* docs(runtime): document final event loop ownership

* fix(stream): cancel stalled owned streams

* fix(cli): recover cleanly from stream cancellation

* fix(runtime): drain executor work before shutdown

* fix(runtime): terminate cancelled shell process trees

* fix(models): let fallback bypass selector failures

* fix(cli): reset interrupt handling between turns

* docs: rm implementation spec

* fix(serve): cancel active turns during shutdown

* fix(runtime): protect settlement from waiter cancellation

* fix(backends): reject empty shell commands

* fix(runtime): terminate descendants after shell exit

* fix(mcp): keep standalone discovery off channel loop

* fix(cli): own and settle interactive prompt cancellation

* fix(serve): keep channel sends off runtime loop

* fix(stream): scope cancel context to iterator steps

* refactor(serve): require the owned async runtime

* fix(channels): keep interactive sends off runtime loop

* fix(selector): surface fallback without log spam

* test(runtime): normalize Windows shell marker

* fix(cli): serialize interactive session turns

* fix(shell): bound output drain after termination

* fix(ui): do not retry owned runtime failures

* fix(shell): allow signal-safe registry reentry

* fix(shell): avoid terminating reused process ids

* fix(channels): preserve streaming send order

* fix(cli): report runtime shutdown timeouts cleanly

* fix(mcp): guide async callers to async loader

* docs(runtime): clarify reserved async bridge APIs

* fix(runtime): bound code interpreter cleanup

* test(shell): use active Python for drain regression

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* feat: add payload-aware EvoAsyncSubAgentMiddleware

* feat: register expert_container_async graph for async expert dispatch

* feat: fold installed expert skills into async subagent registry

* fix: accept 'async' as valid default_dispatch value

* feat: dispatch-aware ActiveTeamMiddleware cue (task vs start_async_task)

* fix: drop future annotations in expert_async_subagent so ToolRuntime injects

* feat: surface output_path and skill_name to expert container as runtime cue

* fix: extend AsyncWatcher client cache with expert specs for completion nudge

* feat: teach main agent the async-expert return envelope shape

* feat: propagate cfg.model to expert-async runs.create via ClientCacheProxy

* chore: guard AsyncWatcher client-cache extension against upstream rename

* test: cover output_path runtime-context tail block and wrong-type guard

* docs: drop out-of-repo notes/ ref from expert_container_async module doc

* fix: warn on unrecognized default_dispatch frontmatter value

* fix: reject empty-body expert skills on async dispatch to match sync policy

* docs: explain why expert container includes general-purpose subagent

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385)

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL

* fix: keep parent env authoritative over workspace .env for mapped keys

* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins

* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS

* fix: merge .env via dotenv_values to close empty-value and RMW-race edges

* chore: align docstrings after .env-merge rework

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Add Requesty as an LLM provider (#346)

* Add Requesty as an LLM provider

* Address review: Requesty prompt caching, model ordering, key validation

- Declare Anthropic-style prompt caching for Requesty Claude models by
  default (mirroring the OpenRouter behavior), with an opt-out flag
  EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
  provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
  overrides native/OpenRouter models for names it shares with them
  (the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
  invalid/missing key (public catalog), so it cannot validate a key. Use a
  minimal authenticated /v1/chat/completions request instead (200 = valid,
  403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).

* Validate Requesty key against auth layer, not a specific model

The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix(llm): filter unnamed tool calls (#390)

* fix(llm): filter unnamed tool calls

* test(llm): cover tool call sanitization branches

* fix(llm): repair unnamed tool calls in middleware

---------

Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix(middleware): mount tool-history repair on sync subagents and harden raw tool-call vetting (#393)

* feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation

* fix(tests): add test for dropping non-list raw tool calls in repair_tool_history

* Add Atlas Cloud LLM provider (#388)

* Add Atlas Cloud LLM provider

* Add Atlas Cloud onboarding support

* fix(validators): update atlascloud key validation to handle insufficient balance case

---------

Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix: prepend EvoAsyncSubAgentMiddleware for prefix cache stability

* docs: clarify list_dispatchable_experts covers both dispatch shapes

* fix: guard async expert fold-in against reserved-name collisions

* fix: honest advertising surfaces for async expert dispatch

* fix: compose expert persona into base-stack system_message

* fix: drop payload from start_async_task, inject skill_name by construction

* feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395

- Introduced TodoListMiddleware to the middleware stack for better task management.
- Updated HITL interrupt configuration to include 'delete' operations requiring approval.
- Implemented error handling for delete operations in read-only and memory backends.
- Enhanced approval prompt formatting to display file paths for delete actions.
- Added tests to ensure delete operations are correctly blocked or prompted for approval.
- Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality.

* revert: drop skill_manager from sync expert-container tool_registry

* revert: drop skill_manager from async expert-container tools

* fix: drop removed ASYNC_TASK_SYSTEM_PROMPT import for deepagents 0.7.0

* fix: mock list_dispatchable_experts in single-cue test for CI

* chore: drop stale output_path from async container graph docstring

* fix: forward configurable_extra through owned-runtime and HITL re-invocations

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: Sanjay Santhanam <51058514+Sanjays2402@users.noreply.github.com>
Co-authored-by: Yougang Lyu <82445958+youganglyu@users.noreply.github.com>
Co-authored-by: houren Antony <2212222@mail.nankai.edu.cn>
Co-authored-by: dinos <dinospk1999@gmail.com>
Co-authored-by: Thibault Jaigu <84420566+Thibaultjaigu@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
Co-authored-by: nightcityblade <jackchen@haloailabs.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nb213 <binyangzhu000@gmail.com>
Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
This commit is contained in:
jfilipiuk
2026-07-31 16:13:30 +02:00
committed by Xi Zhang
parent 3cda9894c7
commit bd2464423a
20 changed files with 2597 additions and 144 deletions
+109 -10
View File
@@ -549,6 +549,105 @@ def _maybe_swap_async_subagents(
return out
def _route_async_specs_through_evo_middleware(
subs: list, base_middleware: list, *, cfg=None
) -> list:
"""Move ``AsyncSubAgent`` specs from ``subs`` into ``EvoAsyncSubAgentMiddleware``.
Deepagents' ``create_deep_agent`` auto-composes the vanilla
``AsyncSubAgentMiddleware`` when it sees ``graph_id``-carrying entries
in ``subagents=``. We need our payload-aware subclass to handle those
(see ``EvoScientist/middleware/expert_async_subagent.py`` for the
upstream-workaround rationale). To prevent the auto-composition and
route all async dispatch through our subclass, we strip AsyncSubAgent
specs from ``subs`` here and hand them to our middleware.
Also folds in ``AsyncSubAgent`` specs for installed
``default_dispatch: async`` expert skills — all pointing at the shared
``expert-container-async`` graph, marked ``is_expert=True`` so the
middleware requires a payload with ``skill_name``.
Returns:
``subs`` with ``graph_id``-carrying entries removed. Safe to pass
as ``create_deep_agent(subagents=...)`` — the async-auto-compose
branch is skipped for empty async lists.
"""
from .middleware.async_watcher import AsyncWatcherMiddleware
from .middleware.expert_async_subagent import EvoAsyncSubAgentMiddleware
from .subagents.expert_container_async import build_expert_async_subagent_specs
cfg = cfg if cfg is not None else _ensure_config()
async_specs = [s for s in subs if "graph_id" in s]
sync_subs = [s for s in subs if "graph_id" not in s]
expert_specs = build_expert_async_subagent_specs(cfg=cfg)
async_specs.extend(expert_specs)
if async_specs:
# ``_maybe_swap_async_subagents`` installs the model-passthrough patch
# only when the yaml-async spec list is non-empty. An expert-only setup
# (no ``writing-agent`` / ``data-analysis-agent`` / ``scheduler`` in
# yaml) would otherwise miss the patch entirely, so we install it here
# too. Idempotent — the shared ``_model_passthrough_patched`` flag
# guards against double-patching.
from .llm.patches import _patch_deepagents_model_passthrough
_patch_deepagents_model_passthrough()
# Prepend rather than append so the ``## Async subagents`` prompt
# section stays in the stable prefix. Appending pushes it past the
# volatile memory tail, invalidating the cached prefix on every
# memory change.
base_middleware.insert(
0, EvoAsyncSubAgentMiddleware(async_subagents=async_specs)
)
# Extend AsyncWatcherMiddleware's client cache with expert specs so
# start_async_task launches for experts spawn a completion watcher —
# otherwise the watcher's ``get_async(agent_name)`` KeyErrors on the
# expert name, no notification is enqueued, and the main agent never
# learns the task finished. ``_maybe_swap_async_subagents`` above only
# populates the watcher with YAML-defined async subagents (writing-agent,
# data-analysis-agent, scheduler); this hook folds in the experts too.
if expert_specs:
watcher = next(
(m for m in base_middleware if isinstance(m, AsyncWatcherMiddleware)),
None,
)
if watcher is not None:
# The mutation reaches through two layers of private state:
# ``AsyncWatcherMiddleware._clients`` (our own) and
# ``_ClientCache._agents`` (upstream deepagents). If upstream ever
# renames ``_agents`` or wraps it in an immutable snapshot, the
# ``.update(...)`` below silently lands on nothing — expert
# completion nudges then stop firing without a diagnostic surface.
# Convert that silent-drop into a grep-able error line and bail
# out of the extension path; expert dispatches still work, just
# without completion notifications until upstream drift is fixed.
if not hasattr(watcher._clients, "_agents"):
logging.getLogger(__name__).error(
"AsyncWatcherMiddleware._clients has no `_agents` slot — "
"deepagents internal renamed; expert completion "
"notifications will not fire until the extension hook is "
"updated to the new attribute name."
)
return sync_subs
watcher._clients._agents.update({s["name"]: s for s in expert_specs})
else:
# No YAML async subagents were registered, so ``_maybe_swap`` did
# not install the watcher. Install it now so experts still get
# completion notifications.
from .cli import async_notifier
base_middleware.append(
AsyncWatcherMiddleware(
{s["name"]: s for s in expert_specs},
notifier=async_notifier,
)
)
return sync_subs
def _build_base_kwargs(
base_backend, base_middleware, *, cfg=None, chat_model=None, workspace_dir=None
):
@@ -557,12 +656,7 @@ def _build_base_kwargs(
from .utils import load_subagents
cfg = cfg if cfg is not None else _ensure_config()
# `skill_manager` is registered here in addition to `base_tools` because
# expert subagents resolve their default toolset from `tool_registry` (see
# `_DEFAULT_EXPERT_TOOLS` in expert_container.py). Without this entry the
# tool silently misses from every expert sub-agent — e.g. idea-brainstorm
# can't run its `paper-navigator` precondition check.
tool_registry = {"think_tool": think_tool, "skill_manager": skill_manager}
tool_registry = {"think_tool": think_tool}
if os.environ.get("TAVILY_API_KEY"):
tool_registry["tavily_search"] = tavily_search
base_tools = [think_tool, skill_manager]
@@ -582,6 +676,10 @@ def _build_base_kwargs(
subs, workspace_dir=workspace_dir, cfg=cfg, chat_model=chat_model
)
subs = _maybe_swap_async_subagents(subs, base_middleware, cfg=cfg)
# Route AsyncSubAgent specs (both standard and expert) through
# EvoAsyncSubAgentMiddleware so the payload-aware start_async_task tool
# replaces upstream's non-parameterisable one.
subs = _route_async_specs_through_evo_middleware(subs, base_middleware, cfg=cfg)
return {
"name": "EvoScientist",
"model": chat_model if chat_model is not None else _ensure_chat_model(),
@@ -634,10 +732,7 @@ def load_mcp_and_build_kwargs(
workspace_dir=workspace_dir,
)
# Match `_build_base_kwargs`: register `skill_manager` in the registry so
# expert subagents (which resolve tools via `_DEFAULT_EXPERT_TOOLS` from
# `expert_container.py`) actually get it.
tool_registry = {"think_tool": think_tool, "skill_manager": skill_manager}
tool_registry = {"think_tool": think_tool}
if os.environ.get("TAVILY_API_KEY"):
tool_registry["tavily_search"] = tavily_search
base_tools = [think_tool, skill_manager]
@@ -672,6 +767,10 @@ def load_mcp_and_build_kwargs(
# Swap selected sub-agents to AsyncSubAgent (must happen AFTER MCP injection
# since async sub-agents are remote graphs that load their own tools).
subs = _maybe_swap_async_subagents(subs, base_middleware, cfg=cfg)
# Mirror the base path: route AsyncSubAgent specs through
# EvoAsyncSubAgentMiddleware so the payload-aware start_async_task tool
# is the one composed into the main agent.
subs = _route_async_specs_through_evo_middleware(subs, base_middleware, cfg=cfg)
return {
"name": "EvoScientist",