Files
EvoScientist-Multi/tests/test_active_team_middleware.py
T
jfilipiuk bd2464423a feat: agent-teams part D - async expert dispatch mechanism (#391)
* fix(deps): pin openrouter below 0.11 to avoid SSE stream regressions (#373)

* fix(openrouter): address SSE stream leak by closing response iterator

* refactor(openrouter): pass through SDK args in SSE leak patch

* test(openrouter): make SSE leak tests version-agnostic across SDK generations

* fix(openrouter): remove SSE stream leak patch and update dependencies

* fix: repair interrupted tool call history (#366)

* fix: repair interrupted tool call history

Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths.

* fix: repair malformed tool calls and dedupe repair warnings

Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted
threads with syntactically invalid tool calls get synthesized error
results and are accepted by strict providers.

Preserve the originating tool call's name in the synthesized ToolMessage,
and deduplicate repair warnings per unique tool-call id via a warned set
owned by the middleware instance, since the middleware rewrites the
request but not thread state.

Document the middleware's scope versus deepagents' PatchToolCallsMiddleware
(orphan ToolMessage dropping and mid-run coverage).

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Update README.md

* Update README.md

* Update README.zh-CN.md

* fix: scrub host path from skill_manager output and guard batch install (#377)

* fix: scrub host path from skill_manager output and guard batch install

* test: tighten install leak guards to catch host path in either tier

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: log missing async-subagent tools at DEBUG, not WARNING (#378)

* fix: log missing async-subagent tools at DEBUG, not WARNING

* fix: distinguish load_subagents callers via async_swap_pending flag

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: default reasoning context for codex proxy Responses API (#380)

* fix: set langgraph and codex proxy runtime defaults

* fix: address runtime default review feedback

* fix: drop langgraph dev env defaults per maintainer review

langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment,
so the reported KeyError cannot come from this flow; the env defaults added
here were unnecessary. Scope the PR back to the codex proxy reasoning
context fix only.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* release: v0.2.4 (#389)

* fix: add support for new Anthropic models and enhance adaptive thinking tests

* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models

* fix: update version to v0.2.4 in badges, README, and project files

* fix: update Star History chart links in README and README.zh-CN

* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog

* fix: update wechat group image in assets

* refactor(runtime): centralize async bridges under an owned runtime (#376)

* feat(runtime): add application-scoped async runtime

* refactor(cli): use owned runtime for session stats

* refactor(onboard): use the owned async runtime

* docs(runtime): record async bridge ownership

* refactor(middleware): keep sync fallback synchronous

* refactor(mcp): load tools on an owned runtime

* refactor(cli): share owned runtime across entry points

* refactor(channels): make inbound sync bridge explicit

* refactor(stream): run Rich streaming on owned runtime

* chore(runtime): remove nest-asyncio dependency

* refactor(asyncio): require active loops in async code

* docs(runtime): document final event loop ownership

* fix(stream): cancel stalled owned streams

* fix(cli): recover cleanly from stream cancellation

* fix(runtime): drain executor work before shutdown

* fix(runtime): terminate cancelled shell process trees

* fix(models): let fallback bypass selector failures

* fix(cli): reset interrupt handling between turns

* docs: rm implementation spec

* fix(serve): cancel active turns during shutdown

* fix(runtime): protect settlement from waiter cancellation

* fix(backends): reject empty shell commands

* fix(runtime): terminate descendants after shell exit

* fix(mcp): keep standalone discovery off channel loop

* fix(cli): own and settle interactive prompt cancellation

* fix(serve): keep channel sends off runtime loop

* fix(stream): scope cancel context to iterator steps

* refactor(serve): require the owned async runtime

* fix(channels): keep interactive sends off runtime loop

* fix(selector): surface fallback without log spam

* test(runtime): normalize Windows shell marker

* fix(cli): serialize interactive session turns

* fix(shell): bound output drain after termination

* fix(ui): do not retry owned runtime failures

* fix(shell): allow signal-safe registry reentry

* fix(shell): avoid terminating reused process ids

* fix(channels): preserve streaming send order

* fix(cli): report runtime shutdown timeouts cleanly

* fix(mcp): guide async callers to async loader

* docs(runtime): clarify reserved async bridge APIs

* fix(runtime): bound code interpreter cleanup

* test(shell): use active Python for drain regression

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* feat: add payload-aware EvoAsyncSubAgentMiddleware

* feat: register expert_container_async graph for async expert dispatch

* feat: fold installed expert skills into async subagent registry

* fix: accept 'async' as valid default_dispatch value

* feat: dispatch-aware ActiveTeamMiddleware cue (task vs start_async_task)

* fix: drop future annotations in expert_async_subagent so ToolRuntime injects

* feat: surface output_path and skill_name to expert container as runtime cue

* fix: extend AsyncWatcher client cache with expert specs for completion nudge

* feat: teach main agent the async-expert return envelope shape

* feat: propagate cfg.model to expert-async runs.create via ClientCacheProxy

* chore: guard AsyncWatcher client-cache extension against upstream rename

* test: cover output_path runtime-context tail block and wrong-type guard

* docs: drop out-of-repo notes/ ref from expert_container_async module doc

* fix: warn on unrecognized default_dispatch frontmatter value

* fix: reject empty-body expert skills on async dispatch to match sync policy

* docs: explain why expert container includes general-purpose subagent

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385)

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL

* fix: keep parent env authoritative over workspace .env for mapped keys

* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins

* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS

* fix: merge .env via dotenv_values to close empty-value and RMW-race edges

* chore: align docstrings after .env-merge rework

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Add Requesty as an LLM provider (#346)

* Add Requesty as an LLM provider

* Address review: Requesty prompt caching, model ordering, key validation

- Declare Anthropic-style prompt caching for Requesty Claude models by
  default (mirroring the OpenRouter behavior), with an opt-out flag
  EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
  provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
  overrides native/OpenRouter models for names it shares with them
  (the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
  invalid/missing key (public catalog), so it cannot validate a key. Use a
  minimal authenticated /v1/chat/completions request instead (200 = valid,
  403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).

* Validate Requesty key against auth layer, not a specific model

The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix(llm): filter unnamed tool calls (#390)

* fix(llm): filter unnamed tool calls

* test(llm): cover tool call sanitization branches

* fix(llm): repair unnamed tool calls in middleware

---------

Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix(middleware): mount tool-history repair on sync subagents and harden raw tool-call vetting (#393)

* feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation

* fix(tests): add test for dropping non-list raw tool calls in repair_tool_history

* Add Atlas Cloud LLM provider (#388)

* Add Atlas Cloud LLM provider

* Add Atlas Cloud onboarding support

* fix(validators): update atlascloud key validation to handle insufficient balance case

---------

Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix: prepend EvoAsyncSubAgentMiddleware for prefix cache stability

* docs: clarify list_dispatchable_experts covers both dispatch shapes

* fix: guard async expert fold-in against reserved-name collisions

* fix: honest advertising surfaces for async expert dispatch

* fix: compose expert persona into base-stack system_message

* fix: drop payload from start_async_task, inject skill_name by construction

* feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395

- Introduced TodoListMiddleware to the middleware stack for better task management.
- Updated HITL interrupt configuration to include 'delete' operations requiring approval.
- Implemented error handling for delete operations in read-only and memory backends.
- Enhanced approval prompt formatting to display file paths for delete actions.
- Added tests to ensure delete operations are correctly blocked or prompted for approval.
- Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality.

* revert: drop skill_manager from sync expert-container tool_registry

* revert: drop skill_manager from async expert-container tools

* fix: drop removed ASYNC_TASK_SYSTEM_PROMPT import for deepagents 0.7.0

* fix: mock list_dispatchable_experts in single-cue test for CI

* chore: drop stale output_path from async container graph docstring

* fix: forward configurable_extra through owned-runtime and HITL re-invocations

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: Sanjay Santhanam <51058514+Sanjays2402@users.noreply.github.com>
Co-authored-by: Yougang Lyu <82445958+youganglyu@users.noreply.github.com>
Co-authored-by: houren Antony <2212222@mail.nankai.edu.cn>
Co-authored-by: dinos <dinospk1999@gmail.com>
Co-authored-by: Thibault Jaigu <84420566+Thibaultjaigu@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
Co-authored-by: nightcityblade <jackchen@haloailabs.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nb213 <binyangzhu000@gmail.com>
Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
2026-08-07 17:07:01 +01:00

348 lines
13 KiB
Python

"""Tests for EvoScientist.middleware.active_team."""
from __future__ import annotations
from types import SimpleNamespace
from unittest.mock import MagicMock, patch
from langchain_core.messages import SystemMessage
from EvoScientist.middleware.active_team import (
ActiveTeamMiddleware,
_read_active_teams,
create_active_team_middleware,
)
def _request():
"""A minimal ModelRequest stand-in supporting the fields the middleware
reads (`system_message`) and the `.override(**kwargs)` mutator."""
request = SimpleNamespace(
state={},
runtime=object(),
system_message=SystemMessage(content="base system"),
)
request.override = lambda **kwargs: SimpleNamespace(
**{
"state": request.state,
"runtime": request.runtime,
"system_message": kwargs.get("system_message", request.system_message),
}
)
return request
def _system_text(modified) -> str:
system_message = modified.system_message
assert system_message is not None
return str(system_message.content)
def _mock_config():
cfg = MagicMock()
cfg.enable_ask_user = False
cfg.auto_mode = False
cfg.auto_approve = False
cfg.model_fallbacks = None
cfg.auxiliary_model = ""
cfg.auxiliary_provider = ""
cfg.code_interpreter_timeout = 60
cfg.code_interpreter_max_result_chars = 6000
return cfg
# ---- unit tests: _read_active_teams behavior --------------------------------
@patch("langgraph.config.get_config")
def test_read_active_teams_returns_list_when_present(mock_get_config):
mock_get_config.return_value = {
"configurable": {"active_teams": ["idea-brainstorm"]},
}
assert _read_active_teams() == ["idea-brainstorm"]
@patch("langgraph.config.get_config")
def test_read_active_teams_returns_empty_when_configurable_missing(mock_get_config):
mock_get_config.return_value = {}
assert _read_active_teams() == []
@patch("langgraph.config.get_config")
def test_read_active_teams_returns_empty_when_active_teams_missing(mock_get_config):
mock_get_config.return_value = {"configurable": {"other_field": "x"}}
assert _read_active_teams() == []
@patch("langgraph.config.get_config")
def test_read_active_teams_returns_empty_when_value_not_list(mock_get_config):
"""WebUI mistakenly sends a scalar instead of a list; must not crash."""
mock_get_config.return_value = {
"configurable": {"active_teams": "idea-brainstorm"},
}
assert _read_active_teams() == []
@patch("langgraph.config.get_config")
def test_read_active_teams_filters_non_string_entries(mock_get_config):
mock_get_config.return_value = {
"configurable": {
"active_teams": ["idea-brainstorm", None, 42, "", "lit-review"]
},
}
assert _read_active_teams() == ["idea-brainstorm", "lit-review"]
@patch("langgraph.config.get_config", side_effect=RuntimeError("outside context"))
def test_read_active_teams_returns_empty_outside_runnable_context(mock_get_config):
assert _read_active_teams() == []
# ---- unit tests: middleware behavior ---------------------------------------
@patch("langgraph.config.get_config")
def test_middleware_no_op_when_active_teams_absent(mock_get_config):
mock_get_config.return_value = {"configurable": {}}
middleware = ActiveTeamMiddleware()
request = _request()
modified = middleware.modify_request(request)
# No override applied: original request returned as-is.
assert modified is request
@patch("langgraph.config.get_config")
def test_middleware_no_op_when_active_teams_empty_list(mock_get_config):
mock_get_config.return_value = {"configurable": {"active_teams": []}}
middleware = ActiveTeamMiddleware()
request = _request()
modified = middleware.modify_request(request)
assert modified is request
def _mock_expert(name: str, dispatch: str) -> MagicMock:
"""Build a MagicMock ``SkillInfo`` with the given dispatch shape.
``name`` on ``MagicMock`` must be set via attribute assignment; passing
``name=`` to the constructor names the mock instance itself.
"""
info = MagicMock(default_dispatch=dispatch)
info.name = name
return info
@patch("EvoScientist.subagents.expert_container.list_dispatchable_experts")
@patch("langgraph.config.get_config")
def test_middleware_appends_single_expert_cue(mock_get_config, mock_dispatchable):
mock_get_config.return_value = {
"configurable": {"active_teams": ["idea-brainstorm"]},
}
mock_dispatchable.return_value = [_mock_expert("idea-brainstorm", "sync")]
middleware = ActiveTeamMiddleware()
modified = middleware.modify_request(_request())
text = _system_text(modified)
assert "<active_expert>" in text
assert "`idea-brainstorm`" in text
assert "Consult it via `task(" in text
assert "base system" in text # original preserved
@patch("EvoScientist.subagents.expert_container.list_dispatchable_experts")
@patch("langgraph.config.get_config")
def test_middleware_appends_multi_expert_cue(mock_get_config, mock_dispatchable):
mock_get_config.return_value = {
"configurable": {"active_teams": ["idea-brainstorm", "literature-review"]},
}
mock_dispatchable.return_value = [
_mock_expert("idea-brainstorm", "sync"),
_mock_expert("literature-review", "sync"),
]
middleware = ActiveTeamMiddleware()
modified = middleware.modify_request(_request())
text = _system_text(modified)
assert "<active_experts>" in text
assert "`idea-brainstorm`" in text
assert "`literature-review`" in text
# Multi-cue: header names both experts, then per-expert dispatch lines follow.
assert "The user has invited the following experts" in text
assert "Per-expert dispatch" in text
assert "base system" in text
@patch("EvoScientist.subagents.expert_container.list_dispatchable_experts")
@patch("langgraph.config.get_config")
def test_middleware_omits_cue_for_undispatchable_names(
mock_get_config, mock_dispatchable
):
"""Names not in ``list_dispatchable_experts`` are dropped from the cue.
Covers uninstalled experts, empty-body experts, name collisions, and
async-declared experts when async dispatch is unavailable — anything
the model would find missing at dispatch time.
"""
mock_get_config.return_value = {
"configurable": {"active_teams": ["nonexistent-expert"]},
}
mock_dispatchable.return_value = [] # nothing dispatchable
request = _request()
middleware = ActiveTeamMiddleware()
modified = middleware.modify_request(request)
# No cue appended — modify_request returns the original request untouched.
assert modified is request
@patch("EvoScientist.subagents.expert_container.list_dispatchable_experts")
@patch("langgraph.config.get_config")
def test_middleware_uses_start_async_task_cue_for_async_dispatch(
mock_get_config, mock_dispatchable
):
"""An expert declared ``default_dispatch: async`` gets the async cue.
Only reaches the cue when async dispatch is actually registered — the
honest-surface filter in ``list_dispatchable_experts`` drops
async-declared experts otherwise.
"""
mock_get_config.return_value = {
"configurable": {"active_teams": ["literature-review"]},
}
mock_dispatchable.return_value = [_mock_expert("literature-review", "async")]
middleware = ActiveTeamMiddleware()
modified = middleware.modify_request(_request())
text = _system_text(modified)
assert "<active_expert>" in text
assert "start_async_task(" in text
assert "subagent_type: 'literature-review'" in text
# Post-X-4: no payload dict. The cue instructs the main agent to embed
# the desired output path directly in the description string.
assert "payload" not in text
assert "output path" in text.lower() or "output_path" in text
assert "check_async_task" in text
# Sync cue must NOT be advertised for async experts.
assert "Consult it via `task(" not in text
@patch("EvoScientist.subagents.expert_container.list_dispatchable_experts")
@patch("langgraph.config.get_config")
def test_middleware_uses_task_cue_for_sync_dispatch(mock_get_config, mock_dispatchable):
"""Sync-dispatched experts get the ``task()`` cue."""
mock_get_config.return_value = {
"configurable": {"active_teams": ["idea-brainstorm"]},
}
mock_dispatchable.return_value = [_mock_expert("idea-brainstorm", "sync")]
middleware = ActiveTeamMiddleware()
modified = middleware.modify_request(_request())
text = _system_text(modified)
assert "Consult it via `task(" in text
assert "runs synchronously" in text
# No async-specific fragments for a sync expert.
assert "start_async_task(" not in text
assert "output_path" not in text
@patch("EvoScientist.subagents.expert_container.list_dispatchable_experts")
@patch("langgraph.config.get_config")
def test_middleware_multi_mixed_dispatch(mock_get_config, mock_dispatchable):
"""When both sync and async experts are active, each gets its own cue."""
mock_get_config.return_value = {
"configurable": {"active_teams": ["idea-brainstorm", "literature-review"]},
}
mock_dispatchable.return_value = [
_mock_expert("idea-brainstorm", "sync"),
_mock_expert("literature-review", "async"),
]
middleware = ActiveTeamMiddleware()
modified = middleware.modify_request(_request())
text = _system_text(modified)
# Both cue shapes appear once each in the per-expert block.
assert text.count("`task(") == 1
assert text.count("start_async_task(") == 1
assert "`idea-brainstorm`:" in text
assert "`literature-review`:" in text
@patch("EvoScientist.subagents.expert_container.list_dispatchable_experts")
@patch("langgraph.config.get_config")
def test_middleware_drops_invited_expert_that_is_not_dispatchable(
mock_get_config, mock_dispatchable
):
"""An async-declared expert stays invited across a config change, but
when async dispatch turns unavailable it drops out of
``list_dispatchable_experts``. The cue must not mention it — otherwise
the model is told to reach for a tool that either doesn't exist or
doesn't list the expert."""
mock_get_config.return_value = {
"configurable": {
"active_teams": ["idea-brainstorm", "literature-review"],
},
}
# literature-review invited but not dispatchable this turn.
mock_dispatchable.return_value = [_mock_expert("idea-brainstorm", "sync")]
middleware = ActiveTeamMiddleware()
modified = middleware.modify_request(_request())
text = _system_text(modified)
# Single-cue shape (only one expert survived the filter).
assert "<active_expert>" in text
assert "`idea-brainstorm`" in text
assert "literature-review" not in text
assert "start_async_task(" not in text
@patch("langgraph.config.get_config", side_effect=RuntimeError("outside context"))
def test_middleware_no_op_outside_runnable_context(mock_get_config):
middleware = ActiveTeamMiddleware()
request = _request()
modified = middleware.modify_request(request)
assert modified is request
# ---- composition tests: _get_default_middleware ----------------------------
@patch(
"EvoScientist.middleware.create_tool_selector_middleware",
return_value=[MagicMock(), MagicMock()],
)
@patch("EvoScientist.EvoScientist._ensure_chat_model")
@patch("EvoScientist.EvoScientist._ensure_config")
def test_default_middleware_includes_active_team_for_main_agent(
mock_config, mock_model, mock_tool_selector
):
mock_config.return_value = _mock_config()
mock_model.return_value = MagicMock(profile={"max_input_tokens": 200_000})
from EvoScientist.EvoScientist import _get_default_middleware
middleware = _get_default_middleware()
assert any(isinstance(m, ActiveTeamMiddleware) for m in middleware)
@patch(
"EvoScientist.middleware.create_tool_selector_middleware",
return_value=[MagicMock(), MagicMock()],
)
@patch("EvoScientist.EvoScientist._ensure_chat_model")
@patch("EvoScientist.EvoScientist._ensure_config")
def test_default_middleware_excludes_active_team_for_async_subagent(
mock_config, mock_model, mock_tool_selector
):
mock_config.return_value = _mock_config()
mock_model.return_value = MagicMock(profile={"max_input_tokens": 200_000})
from EvoScientist.EvoScientist import _get_default_middleware
middleware = _get_default_middleware(for_async_subagent=True)
assert not any(isinstance(m, ActiveTeamMiddleware) for m in middleware)
# ---- factory --------------------------------------------------------------
def test_factory_returns_middleware_instance():
assert isinstance(create_active_team_middleware(), ActiveTeamMiddleware)