feat: agent-teams part D - async expert dispatch mechanism (#391)
* fix(deps): pin openrouter below 0.11 to avoid SSE stream regressions (#373) * fix(openrouter): address SSE stream leak by closing response iterator * refactor(openrouter): pass through SDK args in SSE leak patch * test(openrouter): make SSE leak tests version-agnostic across SDK generations * fix(openrouter): remove SSE stream leak patch and update dependencies * fix: repair interrupted tool call history (#366) * fix: repair interrupted tool call history Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths. * fix: repair malformed tool calls and dedupe repair warnings Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted threads with syntactically invalid tool calls get synthesized error results and are accepted by strict providers. Preserve the originating tool call's name in the synthesized ToolMessage, and deduplicate repair warnings per unique tool-call id via a warned set owned by the middleware instance, since the middleware rewrites the request but not thread state. Document the middleware's scope versus deepagents' PatchToolCallsMiddleware (orphan ToolMessage dropping and mid-run coverage). --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * Update README.md * Update README.md * Update README.zh-CN.md * fix: scrub host path from skill_manager output and guard batch install (#377) * fix: scrub host path from skill_manager output and guard batch install * test: tighten install leak guards to catch host path in either tier --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * fix: log missing async-subagent tools at DEBUG, not WARNING (#378) * fix: log missing async-subagent tools at DEBUG, not WARNING * fix: distinguish load_subagents callers via async_swap_pending flag --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * fix: default reasoning context for codex proxy Responses API (#380) * fix: set langgraph and codex proxy runtime defaults * fix: address runtime default review feedback * fix: drop langgraph dev env defaults per maintainer review langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment, so the reported KeyError cannot come from this flow; the env defaults added here were unnecessary. Scope the PR back to the codex proxy reasoning context fix only. --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * release: v0.2.4 (#389) * fix: add support for new Anthropic models and enhance adaptive thinking tests * fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models * fix: update version to v0.2.4 in badges, README, and project files * fix: update Star History chart links in README and README.zh-CN * fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog * fix: update wechat group image in assets * refactor(runtime): centralize async bridges under an owned runtime (#376) * feat(runtime): add application-scoped async runtime * refactor(cli): use owned runtime for session stats * refactor(onboard): use the owned async runtime * docs(runtime): record async bridge ownership * refactor(middleware): keep sync fallback synchronous * refactor(mcp): load tools on an owned runtime * refactor(cli): share owned runtime across entry points * refactor(channels): make inbound sync bridge explicit * refactor(stream): run Rich streaming on owned runtime * chore(runtime): remove nest-asyncio dependency * refactor(asyncio): require active loops in async code * docs(runtime): document final event loop ownership * fix(stream): cancel stalled owned streams * fix(cli): recover cleanly from stream cancellation * fix(runtime): drain executor work before shutdown * fix(runtime): terminate cancelled shell process trees * fix(models): let fallback bypass selector failures * fix(cli): reset interrupt handling between turns * docs: rm implementation spec * fix(serve): cancel active turns during shutdown * fix(runtime): protect settlement from waiter cancellation * fix(backends): reject empty shell commands * fix(runtime): terminate descendants after shell exit * fix(mcp): keep standalone discovery off channel loop * fix(cli): own and settle interactive prompt cancellation * fix(serve): keep channel sends off runtime loop * fix(stream): scope cancel context to iterator steps * refactor(serve): require the owned async runtime * fix(channels): keep interactive sends off runtime loop * fix(selector): surface fallback without log spam * test(runtime): normalize Windows shell marker * fix(cli): serialize interactive session turns * fix(shell): bound output drain after termination * fix(ui): do not retry owned runtime failures * fix(shell): allow signal-safe registry reentry * fix(shell): avoid terminating reused process ids * fix(channels): preserve streaming send order * fix(cli): report runtime shutdown timeouts cleanly * fix(mcp): guide async callers to async loader * docs(runtime): clarify reserved async bridge APIs * fix(runtime): bound code interpreter cleanup * test(shell): use active Python for drain regression --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * feat: add payload-aware EvoAsyncSubAgentMiddleware * feat: register expert_container_async graph for async expert dispatch * feat: fold installed expert skills into async subagent registry * fix: accept 'async' as valid default_dispatch value * feat: dispatch-aware ActiveTeamMiddleware cue (task vs start_async_task) * fix: drop future annotations in expert_async_subagent so ToolRuntime injects * feat: surface output_path and skill_name to expert container as runtime cue * fix: extend AsyncWatcher client cache with expert specs for completion nudge * feat: teach main agent the async-expert return envelope shape * feat: propagate cfg.model to expert-async runs.create via ClientCacheProxy * chore: guard AsyncWatcher client-cache extension against upstream rename * test: cover output_path runtime-context tail block and wrong-type guard * docs: drop out-of-repo notes/ ref from expert_container_async module doc * fix: warn on unrecognized default_dispatch frontmatter value * fix: reject empty-body expert skills on async dispatch to match sync policy * docs: explain why expert container includes general-purpose subagent * fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385) * fix: propagate langgraph dev bind port into subprocess env for self-loop URL * fix: keep parent env authoritative over workspace .env for mapped keys * fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins * fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS * fix: merge .env via dotenv_values to close empty-value and RMW-race edges * chore: align docstrings after .env-merge rework --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * Add Requesty as an LLM provider (#346) * Add Requesty as an LLM provider * Address review: Requesty prompt caching, model ordering, key validation - Declare Anthropic-style prompt caching for Requesty Claude models by default (mirroring the OpenRouter behavior), with an opt-out flag EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed provider, so the caching check now uses the original provider name. - Move the Requesty model entries above OpenRouter so Requesty no longer overrides native/OpenRouter models for names it shares with them (the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry. - Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an invalid/missing key (public catalog), so it cannot validate a key. Use a minimal authenticated /v1/chat/completions request instead (200 = valid, 403 = invalid), verified against the live endpoint. - Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip). * Validate Requesty key against auth layer, not a specific model The onboarding validator probed /v1/chat/completions with a hardcoded real model (openai/gpt-4o-mini), which tied key validation to that model staying available upstream. The router resolves auth before the model, so probe a deliberately nonexistent sentinel model (requesty/auth-preflight) instead: a valid key yields 404 (model-not-found, auth passed), an invalid key yields 401/403, and 429/5xx stay inconclusive so a transient outage does not reject a good key. Add unit tests covering each case. --------- Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com> * fix(llm): filter unnamed tool calls (#390) * fix(llm): filter unnamed tool calls * test(llm): cover tool call sanitization branches * fix(llm): repair unnamed tool calls in middleware --------- Co-authored-by: nightcityblade <nightcityblade@gmail.com> Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * fix(middleware): mount tool-history repair on sync subagents and harden raw tool-call vetting (#393) * feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation * fix(tests): add test for dropping non-list raw tool calls in repair_tool_history * Add Atlas Cloud LLM provider (#388) * Add Atlas Cloud LLM provider * Add Atlas Cloud onboarding support * fix(validators): update atlascloud key validation to handle insufficient balance case --------- Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com> Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com> * fix: prepend EvoAsyncSubAgentMiddleware for prefix cache stability * docs: clarify list_dispatchable_experts covers both dispatch shapes * fix: guard async expert fold-in against reserved-name collisions * fix: honest advertising surfaces for async expert dispatch * fix: compose expert persona into base-stack system_message * fix: drop payload from start_async_task, inject skill_name by construction * feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395 - Introduced TodoListMiddleware to the middleware stack for better task management. - Updated HITL interrupt configuration to include 'delete' operations requiring approval. - Implemented error handling for delete operations in read-only and memory backends. - Enhanced approval prompt formatting to display file paths for delete actions. - Added tests to ensure delete operations are correctly blocked or prompted for approval. - Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality. * revert: drop skill_manager from sync expert-container tool_registry * revert: drop skill_manager from async expert-container tools * fix: drop removed ASYNC_TASK_SYSTEM_PROMPT import for deepagents 0.7.0 * fix: mock list_dispatchable_experts in single-cue test for CI * chore: drop stale output_path from async container graph docstring * fix: forward configurable_extra through owned-runtime and HITL re-invocations --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> Co-authored-by: Sanjay Santhanam <51058514+Sanjays2402@users.noreply.github.com> Co-authored-by: Yougang Lyu <82445958+youganglyu@users.noreply.github.com> Co-authored-by: houren Antony <2212222@mail.nankai.edu.cn> Co-authored-by: dinos <dinospk1999@gmail.com> Co-authored-by: Thibault Jaigu <84420566+Thibaultjaigu@users.noreply.github.com> Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com> Co-authored-by: nightcityblade <jackchen@haloailabs.com> Co-authored-by: nightcityblade <nightcityblade@gmail.com> Co-authored-by: nb213 <binyangzhu000@gmail.com> Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
This commit is contained in:
+109
-10
@@ -549,6 +549,105 @@ def _maybe_swap_async_subagents(
|
||||
return out
|
||||
|
||||
|
||||
def _route_async_specs_through_evo_middleware(
|
||||
subs: list, base_middleware: list, *, cfg=None
|
||||
) -> list:
|
||||
"""Move ``AsyncSubAgent`` specs from ``subs`` into ``EvoAsyncSubAgentMiddleware``.
|
||||
|
||||
Deepagents' ``create_deep_agent`` auto-composes the vanilla
|
||||
``AsyncSubAgentMiddleware`` when it sees ``graph_id``-carrying entries
|
||||
in ``subagents=``. We need our payload-aware subclass to handle those
|
||||
(see ``EvoScientist/middleware/expert_async_subagent.py`` for the
|
||||
upstream-workaround rationale). To prevent the auto-composition and
|
||||
route all async dispatch through our subclass, we strip AsyncSubAgent
|
||||
specs from ``subs`` here and hand them to our middleware.
|
||||
|
||||
Also folds in ``AsyncSubAgent`` specs for installed
|
||||
``default_dispatch: async`` expert skills — all pointing at the shared
|
||||
``expert-container-async`` graph, marked ``is_expert=True`` so the
|
||||
middleware requires a payload with ``skill_name``.
|
||||
|
||||
Returns:
|
||||
``subs`` with ``graph_id``-carrying entries removed. Safe to pass
|
||||
as ``create_deep_agent(subagents=...)`` — the async-auto-compose
|
||||
branch is skipped for empty async lists.
|
||||
"""
|
||||
from .middleware.async_watcher import AsyncWatcherMiddleware
|
||||
from .middleware.expert_async_subagent import EvoAsyncSubAgentMiddleware
|
||||
from .subagents.expert_container_async import build_expert_async_subagent_specs
|
||||
|
||||
cfg = cfg if cfg is not None else _ensure_config()
|
||||
|
||||
async_specs = [s for s in subs if "graph_id" in s]
|
||||
sync_subs = [s for s in subs if "graph_id" not in s]
|
||||
expert_specs = build_expert_async_subagent_specs(cfg=cfg)
|
||||
async_specs.extend(expert_specs)
|
||||
|
||||
if async_specs:
|
||||
# ``_maybe_swap_async_subagents`` installs the model-passthrough patch
|
||||
# only when the yaml-async spec list is non-empty. An expert-only setup
|
||||
# (no ``writing-agent`` / ``data-analysis-agent`` / ``scheduler`` in
|
||||
# yaml) would otherwise miss the patch entirely, so we install it here
|
||||
# too. Idempotent — the shared ``_model_passthrough_patched`` flag
|
||||
# guards against double-patching.
|
||||
from .llm.patches import _patch_deepagents_model_passthrough
|
||||
|
||||
_patch_deepagents_model_passthrough()
|
||||
|
||||
# Prepend rather than append so the ``## Async subagents`` prompt
|
||||
# section stays in the stable prefix. Appending pushes it past the
|
||||
# volatile memory tail, invalidating the cached prefix on every
|
||||
# memory change.
|
||||
base_middleware.insert(
|
||||
0, EvoAsyncSubAgentMiddleware(async_subagents=async_specs)
|
||||
)
|
||||
|
||||
# Extend AsyncWatcherMiddleware's client cache with expert specs so
|
||||
# start_async_task launches for experts spawn a completion watcher —
|
||||
# otherwise the watcher's ``get_async(agent_name)`` KeyErrors on the
|
||||
# expert name, no notification is enqueued, and the main agent never
|
||||
# learns the task finished. ``_maybe_swap_async_subagents`` above only
|
||||
# populates the watcher with YAML-defined async subagents (writing-agent,
|
||||
# data-analysis-agent, scheduler); this hook folds in the experts too.
|
||||
if expert_specs:
|
||||
watcher = next(
|
||||
(m for m in base_middleware if isinstance(m, AsyncWatcherMiddleware)),
|
||||
None,
|
||||
)
|
||||
if watcher is not None:
|
||||
# The mutation reaches through two layers of private state:
|
||||
# ``AsyncWatcherMiddleware._clients`` (our own) and
|
||||
# ``_ClientCache._agents`` (upstream deepagents). If upstream ever
|
||||
# renames ``_agents`` or wraps it in an immutable snapshot, the
|
||||
# ``.update(...)`` below silently lands on nothing — expert
|
||||
# completion nudges then stop firing without a diagnostic surface.
|
||||
# Convert that silent-drop into a grep-able error line and bail
|
||||
# out of the extension path; expert dispatches still work, just
|
||||
# without completion notifications until upstream drift is fixed.
|
||||
if not hasattr(watcher._clients, "_agents"):
|
||||
logging.getLogger(__name__).error(
|
||||
"AsyncWatcherMiddleware._clients has no `_agents` slot — "
|
||||
"deepagents internal renamed; expert completion "
|
||||
"notifications will not fire until the extension hook is "
|
||||
"updated to the new attribute name."
|
||||
)
|
||||
return sync_subs
|
||||
watcher._clients._agents.update({s["name"]: s for s in expert_specs})
|
||||
else:
|
||||
# No YAML async subagents were registered, so ``_maybe_swap`` did
|
||||
# not install the watcher. Install it now so experts still get
|
||||
# completion notifications.
|
||||
from .cli import async_notifier
|
||||
|
||||
base_middleware.append(
|
||||
AsyncWatcherMiddleware(
|
||||
{s["name"]: s for s in expert_specs},
|
||||
notifier=async_notifier,
|
||||
)
|
||||
)
|
||||
return sync_subs
|
||||
|
||||
|
||||
def _build_base_kwargs(
|
||||
base_backend, base_middleware, *, cfg=None, chat_model=None, workspace_dir=None
|
||||
):
|
||||
@@ -557,12 +656,7 @@ def _build_base_kwargs(
|
||||
from .utils import load_subagents
|
||||
|
||||
cfg = cfg if cfg is not None else _ensure_config()
|
||||
# `skill_manager` is registered here in addition to `base_tools` because
|
||||
# expert subagents resolve their default toolset from `tool_registry` (see
|
||||
# `_DEFAULT_EXPERT_TOOLS` in expert_container.py). Without this entry the
|
||||
# tool silently misses from every expert sub-agent — e.g. idea-brainstorm
|
||||
# can't run its `paper-navigator` precondition check.
|
||||
tool_registry = {"think_tool": think_tool, "skill_manager": skill_manager}
|
||||
tool_registry = {"think_tool": think_tool}
|
||||
if os.environ.get("TAVILY_API_KEY"):
|
||||
tool_registry["tavily_search"] = tavily_search
|
||||
base_tools = [think_tool, skill_manager]
|
||||
@@ -582,6 +676,10 @@ def _build_base_kwargs(
|
||||
subs, workspace_dir=workspace_dir, cfg=cfg, chat_model=chat_model
|
||||
)
|
||||
subs = _maybe_swap_async_subagents(subs, base_middleware, cfg=cfg)
|
||||
# Route AsyncSubAgent specs (both standard and expert) through
|
||||
# EvoAsyncSubAgentMiddleware so the payload-aware start_async_task tool
|
||||
# replaces upstream's non-parameterisable one.
|
||||
subs = _route_async_specs_through_evo_middleware(subs, base_middleware, cfg=cfg)
|
||||
return {
|
||||
"name": "EvoScientist",
|
||||
"model": chat_model if chat_model is not None else _ensure_chat_model(),
|
||||
@@ -634,10 +732,7 @@ def load_mcp_and_build_kwargs(
|
||||
workspace_dir=workspace_dir,
|
||||
)
|
||||
|
||||
# Match `_build_base_kwargs`: register `skill_manager` in the registry so
|
||||
# expert subagents (which resolve tools via `_DEFAULT_EXPERT_TOOLS` from
|
||||
# `expert_container.py`) actually get it.
|
||||
tool_registry = {"think_tool": think_tool, "skill_manager": skill_manager}
|
||||
tool_registry = {"think_tool": think_tool}
|
||||
if os.environ.get("TAVILY_API_KEY"):
|
||||
tool_registry["tavily_search"] = tavily_search
|
||||
base_tools = [think_tool, skill_manager]
|
||||
@@ -672,6 +767,10 @@ def load_mcp_and_build_kwargs(
|
||||
# Swap selected sub-agents to AsyncSubAgent (must happen AFTER MCP injection
|
||||
# since async sub-agents are remote graphs that load their own tools).
|
||||
subs = _maybe_swap_async_subagents(subs, base_middleware, cfg=cfg)
|
||||
# Mirror the base path: route AsyncSubAgent specs through
|
||||
# EvoAsyncSubAgentMiddleware so the payload-aware start_async_task tool
|
||||
# is the one composed into the main agent.
|
||||
subs = _route_async_specs_through_evo_middleware(subs, base_middleware, cfg=cfg)
|
||||
|
||||
return {
|
||||
"name": "EvoScientist",
|
||||
|
||||
Reference in New Issue
Block a user