Files
EvoScientist-Multi/tests/test_expert_async_subagent.py
T
jfilipiuk bd2464423a feat: agent-teams part D - async expert dispatch mechanism (#391)
* fix(deps): pin openrouter below 0.11 to avoid SSE stream regressions (#373)

* fix(openrouter): address SSE stream leak by closing response iterator

* refactor(openrouter): pass through SDK args in SSE leak patch

* test(openrouter): make SSE leak tests version-agnostic across SDK generations

* fix(openrouter): remove SSE stream leak patch and update dependencies

* fix: repair interrupted tool call history (#366)

* fix: repair interrupted tool call history

Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths.

* fix: repair malformed tool calls and dedupe repair warnings

Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted
threads with syntactically invalid tool calls get synthesized error
results and are accepted by strict providers.

Preserve the originating tool call's name in the synthesized ToolMessage,
and deduplicate repair warnings per unique tool-call id via a warned set
owned by the middleware instance, since the middleware rewrites the
request but not thread state.

Document the middleware's scope versus deepagents' PatchToolCallsMiddleware
(orphan ToolMessage dropping and mid-run coverage).

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Update README.md

* Update README.md

* Update README.zh-CN.md

* fix: scrub host path from skill_manager output and guard batch install (#377)

* fix: scrub host path from skill_manager output and guard batch install

* test: tighten install leak guards to catch host path in either tier

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: log missing async-subagent tools at DEBUG, not WARNING (#378)

* fix: log missing async-subagent tools at DEBUG, not WARNING

* fix: distinguish load_subagents callers via async_swap_pending flag

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: default reasoning context for codex proxy Responses API (#380)

* fix: set langgraph and codex proxy runtime defaults

* fix: address runtime default review feedback

* fix: drop langgraph dev env defaults per maintainer review

langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment,
so the reported KeyError cannot come from this flow; the env defaults added
here were unnecessary. Scope the PR back to the codex proxy reasoning
context fix only.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* release: v0.2.4 (#389)

* fix: add support for new Anthropic models and enhance adaptive thinking tests

* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models

* fix: update version to v0.2.4 in badges, README, and project files

* fix: update Star History chart links in README and README.zh-CN

* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog

* fix: update wechat group image in assets

* refactor(runtime): centralize async bridges under an owned runtime (#376)

* feat(runtime): add application-scoped async runtime

* refactor(cli): use owned runtime for session stats

* refactor(onboard): use the owned async runtime

* docs(runtime): record async bridge ownership

* refactor(middleware): keep sync fallback synchronous

* refactor(mcp): load tools on an owned runtime

* refactor(cli): share owned runtime across entry points

* refactor(channels): make inbound sync bridge explicit

* refactor(stream): run Rich streaming on owned runtime

* chore(runtime): remove nest-asyncio dependency

* refactor(asyncio): require active loops in async code

* docs(runtime): document final event loop ownership

* fix(stream): cancel stalled owned streams

* fix(cli): recover cleanly from stream cancellation

* fix(runtime): drain executor work before shutdown

* fix(runtime): terminate cancelled shell process trees

* fix(models): let fallback bypass selector failures

* fix(cli): reset interrupt handling between turns

* docs: rm implementation spec

* fix(serve): cancel active turns during shutdown

* fix(runtime): protect settlement from waiter cancellation

* fix(backends): reject empty shell commands

* fix(runtime): terminate descendants after shell exit

* fix(mcp): keep standalone discovery off channel loop

* fix(cli): own and settle interactive prompt cancellation

* fix(serve): keep channel sends off runtime loop

* fix(stream): scope cancel context to iterator steps

* refactor(serve): require the owned async runtime

* fix(channels): keep interactive sends off runtime loop

* fix(selector): surface fallback without log spam

* test(runtime): normalize Windows shell marker

* fix(cli): serialize interactive session turns

* fix(shell): bound output drain after termination

* fix(ui): do not retry owned runtime failures

* fix(shell): allow signal-safe registry reentry

* fix(shell): avoid terminating reused process ids

* fix(channels): preserve streaming send order

* fix(cli): report runtime shutdown timeouts cleanly

* fix(mcp): guide async callers to async loader

* docs(runtime): clarify reserved async bridge APIs

* fix(runtime): bound code interpreter cleanup

* test(shell): use active Python for drain regression

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* feat: add payload-aware EvoAsyncSubAgentMiddleware

* feat: register expert_container_async graph for async expert dispatch

* feat: fold installed expert skills into async subagent registry

* fix: accept 'async' as valid default_dispatch value

* feat: dispatch-aware ActiveTeamMiddleware cue (task vs start_async_task)

* fix: drop future annotations in expert_async_subagent so ToolRuntime injects

* feat: surface output_path and skill_name to expert container as runtime cue

* fix: extend AsyncWatcher client cache with expert specs for completion nudge

* feat: teach main agent the async-expert return envelope shape

* feat: propagate cfg.model to expert-async runs.create via ClientCacheProxy

* chore: guard AsyncWatcher client-cache extension against upstream rename

* test: cover output_path runtime-context tail block and wrong-type guard

* docs: drop out-of-repo notes/ ref from expert_container_async module doc

* fix: warn on unrecognized default_dispatch frontmatter value

* fix: reject empty-body expert skills on async dispatch to match sync policy

* docs: explain why expert container includes general-purpose subagent

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385)

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL

* fix: keep parent env authoritative over workspace .env for mapped keys

* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins

* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS

* fix: merge .env via dotenv_values to close empty-value and RMW-race edges

* chore: align docstrings after .env-merge rework

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Add Requesty as an LLM provider (#346)

* Add Requesty as an LLM provider

* Address review: Requesty prompt caching, model ordering, key validation

- Declare Anthropic-style prompt caching for Requesty Claude models by
  default (mirroring the OpenRouter behavior), with an opt-out flag
  EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
  provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
  overrides native/OpenRouter models for names it shares with them
  (the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
  invalid/missing key (public catalog), so it cannot validate a key. Use a
  minimal authenticated /v1/chat/completions request instead (200 = valid,
  403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).

* Validate Requesty key against auth layer, not a specific model

The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix(llm): filter unnamed tool calls (#390)

* fix(llm): filter unnamed tool calls

* test(llm): cover tool call sanitization branches

* fix(llm): repair unnamed tool calls in middleware

---------

Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix(middleware): mount tool-history repair on sync subagents and harden raw tool-call vetting (#393)

* feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation

* fix(tests): add test for dropping non-list raw tool calls in repair_tool_history

* Add Atlas Cloud LLM provider (#388)

* Add Atlas Cloud LLM provider

* Add Atlas Cloud onboarding support

* fix(validators): update atlascloud key validation to handle insufficient balance case

---------

Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix: prepend EvoAsyncSubAgentMiddleware for prefix cache stability

* docs: clarify list_dispatchable_experts covers both dispatch shapes

* fix: guard async expert fold-in against reserved-name collisions

* fix: honest advertising surfaces for async expert dispatch

* fix: compose expert persona into base-stack system_message

* fix: drop payload from start_async_task, inject skill_name by construction

* feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395

- Introduced TodoListMiddleware to the middleware stack for better task management.
- Updated HITL interrupt configuration to include 'delete' operations requiring approval.
- Implemented error handling for delete operations in read-only and memory backends.
- Enhanced approval prompt formatting to display file paths for delete actions.
- Added tests to ensure delete operations are correctly blocked or prompted for approval.
- Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality.

* revert: drop skill_manager from sync expert-container tool_registry

* revert: drop skill_manager from async expert-container tools

* fix: drop removed ASYNC_TASK_SYSTEM_PROMPT import for deepagents 0.7.0

* fix: mock list_dispatchable_experts in single-cue test for CI

* chore: drop stale output_path from async container graph docstring

* fix: forward configurable_extra through owned-runtime and HITL re-invocations

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: Sanjay Santhanam <51058514+Sanjays2402@users.noreply.github.com>
Co-authored-by: Yougang Lyu <82445958+youganglyu@users.noreply.github.com>
Co-authored-by: houren Antony <2212222@mail.nankai.edu.cn>
Co-authored-by: dinos <dinospk1999@gmail.com>
Co-authored-by: Thibault Jaigu <84420566+Thibaultjaigu@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
Co-authored-by: nightcityblade <jackchen@haloailabs.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nb213 <binyangzhu000@gmail.com>
Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
2026-08-07 17:07:01 +01:00

341 lines
14 KiB
Python

"""Tests for the skill-name-injecting AsyncSubAgentMiddleware subclass."""
from __future__ import annotations
from types import SimpleNamespace
from unittest.mock import AsyncMock, MagicMock, patch
import pytest
from EvoScientist.middleware.expert_async_subagent import (
EvoAsyncSubAgentMiddleware,
_build_run_input,
)
class _TestPayloadValidationRemoved:
"""Placeholder — the ``_payload_validation_error`` helper was deleted
when ``payload`` was dropped from the tool schema (PR #391 review, X-4).
The seven tests that lived here (``TestPayloadValidation``) no longer
apply: subagent_type is validated by ``_validate_agent_type``,
``skill_name`` is injected by construction, and no other user-supplied
fields reach ``client.runs.create(input=...)``. See
``TestBuildRunInput`` below and ``TestStartToolInvocation`` for the
replacement coverage.
"""
# =============================================================================
# _build_run_input — the shared input-dict factory
# =============================================================================
class TestBuildRunInput:
"""``skill_name`` is injected for expert specs, absent for standard specs.
The description always lands in ``messages`` verbatim — no LLM-authored
key can overwrite it (was the pre-fix bug when ``payload`` was in scope).
"""
def test_expert_spec_injects_skill_name(self):
spec = {"name": "e", "graph_id": "g", "is_expert": True}
result = _build_run_input(spec, "literature-review", "write a survey")
assert result == {
"messages": [{"role": "user", "content": "write a survey"}],
"skill_name": "literature-review",
}
def test_standard_spec_matches_upstream_shape(self):
"""Standard specs (writing-agent, scheduler, ...) reach ``runs.create``
with the upstream single-key shape — no ``skill_name`` injected."""
spec = {"name": "writing-agent", "graph_id": "writing_agent"}
result = _build_run_input(spec, "writing-agent", "hi")
assert result == {"messages": [{"role": "user", "content": "hi"}]}
def test_is_expert_false_treated_as_standard(self):
"""Explicit ``is_expert=False`` matches the default (absent) behaviour."""
spec = {"name": "std", "graph_id": "writing_agent", "is_expert": False}
result = _build_run_input(spec, "std", "hi")
assert result == {"messages": [{"role": "user", "content": "hi"}]}
def test_description_lands_verbatim(self):
"""Regression guard against the pre-fix bug where an LLM-authored
``payload`` could overwrite ``messages`` — description now travels
through a channel the LLM cannot corrupt."""
spec = {"name": "e", "graph_id": "g", "is_expert": True}
result = _build_run_input(
spec, "e", "write to ./artifacts/e/foo.md a summary of X"
)
assert result["messages"][0]["content"] == (
"write to ./artifacts/e/foo.md a summary of X"
)
# =============================================================================
# EvoAsyncSubAgentMiddleware — end-to-end tool invocation
# =============================================================================
def _standard_spec():
return {
"name": "writing-agent",
"description": "std writer",
"graph_id": "writing_agent",
}
def _expert_spec():
return {
"name": "literature-review",
"description": "expert lit review",
"graph_id": "expert_container",
"is_expert": True,
}
class TestMiddlewareConstruction:
def test_middleware_has_five_tools(self):
mw = EvoAsyncSubAgentMiddleware(async_subagents=[_standard_spec()])
names = [t.name for t in mw.tools]
assert set(names) == {
"start_async_task",
"check_async_task",
"update_async_task",
"cancel_async_task",
"list_async_tasks",
}
def test_start_tool_schema_matches_upstream(self):
"""The tool signature returned to upstream's exact shape when
``payload`` was dropped — schema is now ``deepagents``'s
``StartAsyncTaskSchema``."""
from deepagents.middleware.async_subagents import StartAsyncTaskSchema
mw = EvoAsyncSubAgentMiddleware(async_subagents=[_standard_spec()])
start = next(t for t in mw.tools if t.name == "start_async_task")
assert start.args_schema is StartAsyncTaskSchema
def test_construction_rejects_empty_subagents(self):
with pytest.raises(ValueError, match="At least one async subagent"):
EvoAsyncSubAgentMiddleware(async_subagents=[])
def test_construction_rejects_duplicate_names(self):
with pytest.raises(ValueError, match="Duplicate"):
EvoAsyncSubAgentMiddleware(
async_subagents=[_standard_spec(), _standard_spec()]
)
def _fake_sync_client():
client = MagicMock()
client.threads.create.return_value = {"thread_id": "task-abc"}
client.runs.create.return_value = {"run_id": "run-xyz"}
return client
def _fake_async_client():
client = MagicMock()
client.threads.create = AsyncMock(return_value={"thread_id": "task-abc"})
client.runs.create = AsyncMock(return_value={"run_id": "run-xyz"})
return client
class TestStartToolInvocation:
"""Direct invocation of the start tool's sync function.
Mocks ``_ClientCache.get_sync`` so we can assert on the ``input`` dict
handed to ``runs.create`` without any real network round-trip.
"""
def test_start_injects_skill_name_for_expert_spec(self):
"""The middleware sets ``input_dict['skill_name'] = subagent_type``
by construction — the shared container graph resolves the right
persona without a payload dict crossing the LLM channel."""
mw = EvoAsyncSubAgentMiddleware(async_subagents=[_expert_spec()])
start = next(t for t in mw.tools if t.name == "start_async_task")
client = _fake_sync_client()
with patch(
"EvoScientist.middleware.expert_async_subagent._ClientCache.get_sync",
return_value=client,
):
result = start.func(
description="write to ./artifacts/literature-review/attn.md a survey on X",
subagent_type="literature-review",
runtime=SimpleNamespace(tool_call_id="tc1"),
)
client.runs.create.assert_called_once()
kwargs = client.runs.create.call_args.kwargs
assert kwargs["assistant_id"] == "expert_container"
assert kwargs["input"]["messages"] == [
{
"role": "user",
"content": (
"write to ./artifacts/literature-review/attn.md a survey on X"
),
}
]
assert kwargs["input"]["skill_name"] == "literature-review"
assert "payload" not in kwargs["input"]
assert "output_path" not in kwargs["input"]
# Return value stamps the task into async_tasks state.
assert "async_tasks" in result.update
assert "task-abc" in result.update["async_tasks"]
def test_start_injects_cfg_model_into_configurable(self):
"""cfg.model / cfg.provider land in ``config.configurable`` on every
``runs.create`` so the deployed graph re-resolves its chat model per
run instead of using whatever was baked at container-build time.
Without this the ``/model`` CLI switch silently doesn't propagate to
expert launches.
"""
from EvoScientist.config.settings import EvoScientistConfig
mw = EvoAsyncSubAgentMiddleware(async_subagents=[_expert_spec()])
start = next(t for t in mw.tools if t.name == "start_async_task")
client = _fake_sync_client()
fake_cfg = EvoScientistConfig(model="test-model-abc", provider="test-provider")
with (
patch(
"EvoScientist.middleware.expert_async_subagent._ClientCache.get_sync",
return_value=client,
),
patch("EvoScientist.EvoScientist._ensure_config", return_value=fake_cfg),
):
start.func(
description="w",
subagent_type="literature-review",
runtime=SimpleNamespace(tool_call_id="tc1"),
)
kwargs = client.runs.create.call_args.kwargs
assert "config" in kwargs
configurable = kwargs["config"]["configurable"]
assert configurable["model"] == "test-model-abc"
assert configurable["model_provider"] == "test-provider"
def test_start_standard_spec_matches_upstream_input_shape(self):
"""Standard subagents (writing-agent, scheduler, ...) reach
``runs.create`` with the upstream single-key ``messages`` shape."""
mw = EvoAsyncSubAgentMiddleware(async_subagents=[_standard_spec()])
start = next(t for t in mw.tools if t.name == "start_async_task")
client = _fake_sync_client()
with patch(
"EvoScientist.middleware.expert_async_subagent._ClientCache.get_sync",
return_value=client,
):
start.func(
description="hi",
subagent_type="writing-agent",
runtime=SimpleNamespace(tool_call_id="tc1"),
)
kwargs = client.runs.create.call_args.kwargs
assert kwargs["input"] == {"messages": [{"role": "user", "content": "hi"}]}
def test_start_unknown_subagent_returns_error(self):
mw = EvoAsyncSubAgentMiddleware(async_subagents=[_standard_spec()])
start = next(t for t in mw.tools if t.name == "start_async_task")
result = start.func(
description="hi",
subagent_type="does-not-exist",
runtime=SimpleNamespace(tool_call_id="tc1"),
)
assert isinstance(result, str)
assert "Unknown async subagent type" in result
class TestAstartToolInvocation:
"""Mirror ``TestStartToolInvocation`` against ``astart_async_task`` — the
coroutine langgraph_api actually runs in production. Pre-fix zero
coverage: X-iZhang flagged that a fix applied only to the sync body
would leave tests green and production broken."""
@pytest.mark.asyncio
async def test_astart_injects_skill_name_for_expert_spec(self):
mw = EvoAsyncSubAgentMiddleware(async_subagents=[_expert_spec()])
start = next(t for t in mw.tools if t.name == "start_async_task")
client = _fake_async_client()
with patch(
"EvoScientist.middleware.expert_async_subagent._ClientCache.get_async",
return_value=client,
):
result = await start.coroutine(
description="write to ./artifacts/literature-review/attn.md a survey on X",
subagent_type="literature-review",
runtime=SimpleNamespace(tool_call_id="tc1"),
)
client.runs.create.assert_awaited_once()
kwargs = client.runs.create.await_args.kwargs
assert kwargs["assistant_id"] == "expert_container"
assert kwargs["input"]["skill_name"] == "literature-review"
assert kwargs["input"]["messages"][0]["content"].startswith(
"write to ./artifacts/literature-review/attn.md"
)
assert "payload" not in kwargs["input"]
assert "async_tasks" in result.update
assert "task-abc" in result.update["async_tasks"]
@pytest.mark.asyncio
async def test_astart_injects_cfg_model_into_configurable(self):
from EvoScientist.config.settings import EvoScientistConfig
mw = EvoAsyncSubAgentMiddleware(async_subagents=[_expert_spec()])
start = next(t for t in mw.tools if t.name == "start_async_task")
client = _fake_async_client()
fake_cfg = EvoScientistConfig(model="test-model-abc", provider="test-provider")
with (
patch(
"EvoScientist.middleware.expert_async_subagent._ClientCache.get_async",
return_value=client,
),
patch("EvoScientist.EvoScientist._ensure_config", return_value=fake_cfg),
):
await start.coroutine(
description="w",
subagent_type="literature-review",
runtime=SimpleNamespace(tool_call_id="tc1"),
)
kwargs = client.runs.create.await_args.kwargs
assert "config" in kwargs
configurable = kwargs["config"]["configurable"]
assert configurable["model"] == "test-model-abc"
assert configurable["model_provider"] == "test-provider"
@pytest.mark.asyncio
async def test_astart_standard_spec_matches_upstream_input_shape(self):
mw = EvoAsyncSubAgentMiddleware(async_subagents=[_standard_spec()])
start = next(t for t in mw.tools if t.name == "start_async_task")
client = _fake_async_client()
with patch(
"EvoScientist.middleware.expert_async_subagent._ClientCache.get_async",
return_value=client,
):
await start.coroutine(
description="hi",
subagent_type="writing-agent",
runtime=SimpleNamespace(tool_call_id="tc1"),
)
kwargs = client.runs.create.await_args.kwargs
assert kwargs["input"] == {"messages": [{"role": "user", "content": "hi"}]}
@pytest.mark.asyncio
async def test_astart_unknown_subagent_returns_error(self):
mw = EvoAsyncSubAgentMiddleware(async_subagents=[_standard_spec()])
start = next(t for t in mw.tools if t.name == "start_async_task")
result = await start.coroutine(
description="hi",
subagent_type="does-not-exist",
runtime=SimpleNamespace(tool_call_id="tc1"),
)
assert isinstance(result, str)
assert "Unknown async subagent type" in result