01845f4311
* chore: add pytest-asyncio in auto mode * test: migrate channel and stream tests to native async Convert run_async() wrapper tests to plain 'async def test_*' under pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a coroutine awaited at every call site. * test: migrate command and model/middleware tests to native async Convert run_async() wrappers (import, alias, and fixture forms) to plain 'async def test_*'. Multi-call tests merge onto one loop as sequential awaits; none asserted on loop identity. * test: migrate TUI, notifier, gateway, and session tests to native async TUI/notifier/gateway files convert run_async wrappers to plain async tests. test_sessions.py's unittest.TestCase classes move to unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async methods on plain TestCase; converting blindly would have made ~70 tests silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget in test_tui_widgets.py drops its TestCase base for the same reason. * test: replace direct asyncio.run() calls with native async tests Convert tests that called asyncio.run() (directly or via a local _run helper) to plain 'async def test_*'; delete the local helpers. * test: drop undeclared anyio markers and delete run_async helper The @pytest.mark.anyio tests relied on anyio being a transitive dep of httpx; auto-mode pytest-asyncio collects them natively. run_async() and its fixture are unreferenced after the migration, so remove them — pytest-asyncio's per-test loop teardown covers the pending-task cancellation the helper existed for (verified: full suite runs with no 'Event loop is closed' errors or destroyed-task warnings). * test: add autouse fixture for watcher cleanup * refactor: remove redundant hasattr calls * refactor: add typed middleware event sink and thread through assembly Add MiddlewareEventSink protocol + NoOpSink in middleware/events.py with a documented any-thread non-blocking contract (contract test uses a deliberately-slow fake sink). Thread an optional `events` parameter through create_cli_agent -> _get_default_middleware -> tool selector / model fallback constructors; subagent stacks are always forced to NoOpSink. * refactor: inject a notifier port into async-watcher and background middleware Add public pre_cancel_watcher() and enqueue_task_notification() to cli/async_notifier.py and a small NotifierPort protocol (middleware/notifier.py) that the module satisfies structurally. AsyncWatcherMiddleware and BackgroundExecutionMiddleware now receive the port by constructor injection at the composition root, deleting the lazy 'from ..cli import async_notifier' imports and the private _watcher_by_thread / _enqueue pokes. * refactor: invert tool-selection ownership onto a frontend event sink The adaptive tool selector now reports on_tool_selection_started / on_tool_selection / on_tool_selection_ended to the injected sink instead of writing four process-global module variables. The frontend sink (stream/sink.py FrontendEventSink) owns the selected/total/active state with consume-once + dedup-vs-last-emitted semantics; stream/tool_selection.py reads that sink object (a ToolSelectionView) rather than reaching into tool_selector's globals. Deleted: the 4 module globals, the cross-module mutations in tool_selection.py, the track_stream_selection flag, the now-vestigial _ToolSelectionTrackerMiddleware, reset_tool_selection_state_for_tests, and the autouse conftest fixture. The sink is threaded from the two interactive frontends through create_runtime_gateways -> LocalGraphGateway (read side) and _load_agent -> create_cli_agent (write side); subagent / headless stacks get NoOpSink. * refactor: route model-fallback narration through the injected event sink Delete the _ui_emit_fn / set_ui_emit module global and the ..stream.console import from model_fallback.py. The fallback middleware now reports through its injected sink: the fallback transition via the structured on_model_fallback (the frontend formats the '-> Falling back to ...' line), and the surrounding narration (primary-failure header, per-attempt outcome, exhaustion, non-fallbackable rejection) via emit_fallback_notice, preserving the exact user-facing text. The TUI binds its _append_system as the sink's fallback display where it used to call set_ui_emit (cleared on exit); the Rich CLI's sink prints to the console. _try_fallbacks / _guard_and_fallback take the sink. * refactor: declare events on the GraphGateway protocol Both gateway implementations now carry an explicit events attribute (LangGraphServerGateway holds None — no frontend renders middleware events across the HTTP boundary), so the four call sites use plain attribute access instead of getattr probing an implicit contract. * refactor: bind fallback display via the closure-scoped concrete sink The App methods used gateway.events (typed as the read-side view) and hasattr-probed for the concrete FrontendEventSink API. The enclosing factory creates that sink two hundred lines up — close over it directly: no probing, fully typed, and it becomes a constructor parameter naturally when the App class is hoisted out of the factory. * fix: end tool selection before fallback handler * fix: keep fallback display errors non-fatal * fix: preserve selector suppression for default streams * fix: restore fallback notice console display * refactor: consolidate fallback narration events * refactor: clean middleware event sink plumbing * fix: type gateway session events * refactor: make all event protocols runtime-checkable MiddlewareEventSink already carried @runtime_checkable (the stream binding guard isinstance-checks it); ToolSelectionView and SessionEvents now match, so mirroring that pattern against any of the three protocols works instead of raising TypeError. * fix(cli): close QuickJS workers after one-shot failures * fix(cli): honor no-thinking in final output * fix(channels): report failed startup accurately * fix(channels): make Telegram cleanup idempotent * fix(tui): skip command sync during exit * fix(channels): preserve startup state during retries * refactor(channels): share pending startup status * refactor(cli): expose channel startup snapshot * fix(tui): move channel startup off event loop * test(channels): release retry gate on assertion failure --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
319 lines
12 KiB
Python
319 lines
12 KiB
Python
"""Tests for ``EvoScientist.subagents._factory.build_async_subagent_graph``.
|
|
|
|
Pins the integration contract that the factory must request middleware
|
|
in async-safe mode (``for_async_subagent=True``). Without this, a future
|
|
refactor that drops the keyword argument would silently re-introduce
|
|
``AskUserMiddleware`` into the deployed graph and reproduce the
|
|
``interrupt()``-based deadlock the flag was added to prevent.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from unittest.mock import MagicMock, patch
|
|
|
|
from EvoScientist.config import MemoryObservationWriter
|
|
from EvoScientist.memory import MemorySourceType
|
|
|
|
|
|
def _single_middleware(subagent: dict, class_name: str):
|
|
matches = [m for m in subagent["middleware"] if type(m).__name__ == class_name]
|
|
assert len(matches) == 1
|
|
return matches[0]
|
|
|
|
|
|
def _assert_subagent_memory_middleware(subagent: dict, *, source_agent: str) -> None:
|
|
memory_middleware = _single_middleware(subagent, "EvoMemoryMiddleware")
|
|
lifecycle_middleware = _single_middleware(
|
|
subagent,
|
|
"EvoMemoryLifecycleMiddleware",
|
|
)
|
|
|
|
assert [tool.name for tool in memory_middleware.tools] == [
|
|
"search_observations",
|
|
"read_memory",
|
|
"record_observation",
|
|
]
|
|
assert lifecycle_middleware._source_type == MemorySourceType.SUBAGENT
|
|
assert lifecycle_middleware._source_agent == source_agent
|
|
assert lifecycle_middleware._project_id == memory_middleware.project_id
|
|
|
|
|
|
@patch("deepagents.create_deep_agent")
|
|
@patch("EvoScientist.EvoScientist._load_mcp_tools_cached", return_value={})
|
|
@patch("EvoScientist.EvoScientist._get_default_middleware", return_value=[])
|
|
@patch("EvoScientist.EvoScientist._get_default_backend")
|
|
@patch("EvoScientist.EvoScientist._ensure_chat_model")
|
|
@patch("EvoScientist.utils.load_subagents")
|
|
@patch("EvoScientist.config.apply_config_to_env")
|
|
@patch("EvoScientist.config.get_effective_config")
|
|
def test_factory_requests_async_safe_middleware(
|
|
mock_get_cfg,
|
|
mock_apply_env,
|
|
mock_load_subs,
|
|
mock_chat,
|
|
mock_backend,
|
|
mock_get_mw,
|
|
mock_mcp,
|
|
mock_create,
|
|
):
|
|
"""``build_async_subagent_graph`` must call ``_get_default_middleware``
|
|
with ``for_async_subagent=True``.
|
|
|
|
The bare argument call would silently include ``AskUserMiddleware`` in
|
|
the deployed graph, which deadlocks via ``interrupt()`` (no UI in the
|
|
langgraph dev subprocess to resume the interrupt).
|
|
"""
|
|
# Minimal config stub so factory's `cfg.recursion_limit` access works.
|
|
cfg = MagicMock()
|
|
cfg.recursion_limit = 1_000_000
|
|
cfg.memory_profile_enabled = True
|
|
cfg.memory_observations_enabled = True
|
|
cfg.memory_observation_writer = MemoryObservationWriter.ALL
|
|
cfg.memory_workers_enabled = True
|
|
mock_get_cfg.return_value = cfg
|
|
# Factory looks up the requested name in the loaded subagent specs;
|
|
# any matching name is fine.
|
|
mock_load_subs.return_value = [
|
|
{
|
|
"name": "writing-agent",
|
|
"system_prompt": "",
|
|
"tools": [],
|
|
"skills": None,
|
|
}
|
|
]
|
|
# ``create_deep_agent(...).with_config({...})`` chain — return something
|
|
# chainable so the factory's terminal ``.with_config(...)`` doesn't blow up.
|
|
mock_create.return_value.with_config.return_value = MagicMock()
|
|
|
|
from EvoScientist.subagents._factory import build_async_subagent_graph
|
|
|
|
build_async_subagent_graph("writing-agent")
|
|
|
|
# The contract: factory MUST pass async-safe mode and the source agent name.
|
|
mock_get_mw.assert_called_once_with(
|
|
for_async_subagent=True,
|
|
memory_source_agent="writing-agent",
|
|
)
|
|
subagents = mock_create.call_args.kwargs["subagents"]
|
|
assert subagents[0]["name"] == "general-purpose"
|
|
_assert_subagent_memory_middleware(
|
|
subagents[0],
|
|
source_agent="general-purpose",
|
|
)
|
|
|
|
|
|
@patch("EvoScientist.EvoScientist._ensure_chat_model")
|
|
def test_inject_subagent_adds_memory_middleware(mock_model, tmp_path):
|
|
mock_model.return_value = MagicMock(profile={"max_input_tokens": 200_000})
|
|
|
|
from EvoScientist.EvoScientist import _inject_subagent_middleware
|
|
|
|
workspace = tmp_path / "workspace"
|
|
workspace.mkdir()
|
|
subs = [{"name": "test-agent"}]
|
|
|
|
_inject_subagent_middleware(subs, workspace_dir=workspace)
|
|
|
|
_assert_subagent_memory_middleware(subs[0], source_agent="test-agent")
|
|
|
|
|
|
@patch("EvoScientist.EvoScientist._ensure_chat_model")
|
|
@patch("EvoScientist.EvoScientist._ensure_config")
|
|
def test_inject_subagent_omits_memory_middleware_when_memory_disabled(
|
|
mock_config, mock_model, tmp_path
|
|
):
|
|
mock_model.return_value = MagicMock(profile={"max_input_tokens": 200_000})
|
|
cfg = MagicMock()
|
|
cfg.memory_profile_enabled = False
|
|
cfg.memory_observations_enabled = False
|
|
cfg.memory_observation_writer = MemoryObservationWriter.ALL
|
|
cfg.memory_workers_enabled = True
|
|
cfg.auxiliary_model = ""
|
|
cfg.auxiliary_provider = ""
|
|
mock_config.return_value = cfg
|
|
|
|
from EvoScientist.EvoScientist import _inject_subagent_middleware
|
|
|
|
workspace = tmp_path / "workspace"
|
|
workspace.mkdir()
|
|
subs = [{"name": "test-agent"}]
|
|
|
|
_inject_subagent_middleware(subs, workspace_dir=workspace)
|
|
|
|
assert not [
|
|
m
|
|
for m in subs[0]["middleware"]
|
|
if type(m).__name__ in {"EvoMemoryMiddleware", "EvoMemoryLifecycleMiddleware"}
|
|
]
|
|
|
|
|
|
@patch("EvoScientist.EvoScientist._ensure_chat_model")
|
|
@patch("EvoScientist.EvoScientist._ensure_config")
|
|
def test_inject_subagent_worker_only_observation_writer_keeps_live_tool_off(
|
|
mock_config, mock_model, tmp_path
|
|
):
|
|
mock_model.return_value = MagicMock(profile={"max_input_tokens": 200_000})
|
|
cfg = MagicMock()
|
|
cfg.memory_profile_enabled = False
|
|
cfg.memory_observations_enabled = True
|
|
cfg.memory_observation_writer = MemoryObservationWriter.WORKER
|
|
cfg.memory_workers_enabled = True
|
|
cfg.auxiliary_model = ""
|
|
cfg.auxiliary_provider = ""
|
|
mock_config.return_value = cfg
|
|
|
|
from EvoScientist.EvoScientist import _inject_subagent_middleware
|
|
|
|
workspace = tmp_path / "workspace"
|
|
workspace.mkdir()
|
|
subs = [{"name": "test-agent"}]
|
|
|
|
_inject_subagent_middleware(subs, workspace_dir=workspace)
|
|
|
|
memory_middleware = _single_middleware(subs[0], "EvoMemoryMiddleware")
|
|
lifecycle_middleware = _single_middleware(
|
|
subs[0],
|
|
"EvoMemoryLifecycleMiddleware",
|
|
)
|
|
assert [tool.name for tool in memory_middleware.tools] == [
|
|
"search_observations",
|
|
"read_memory",
|
|
]
|
|
assert lifecycle_middleware._source_type == MemorySourceType.SUBAGENT
|
|
|
|
|
|
@patch(
|
|
"EvoScientist.middleware.create_tool_selector_middleware",
|
|
return_value=[MagicMock()],
|
|
)
|
|
@patch("EvoScientist.EvoScientist._ensure_chat_model")
|
|
@patch("EvoScientist.EvoScientist._ensure_config")
|
|
def test_all_observation_writer_schedules_turn_worker_without_profile_memory(
|
|
mock_config, mock_chat, mock_tool_selector
|
|
):
|
|
cfg = MagicMock()
|
|
cfg.enable_ask_user = False
|
|
cfg.auto_mode = False
|
|
cfg.auto_approve = False
|
|
cfg.model_fallbacks = None
|
|
cfg.memory_profile_enabled = False
|
|
cfg.memory_observations_enabled = True
|
|
cfg.memory_observation_writer = MemoryObservationWriter.ALL
|
|
cfg.memory_workers_enabled = True
|
|
cfg.auxiliary_model = ""
|
|
cfg.auxiliary_provider = ""
|
|
mock_config.return_value = cfg
|
|
mock_chat.return_value = MagicMock(profile={"max_input_tokens": 200_000})
|
|
|
|
from EvoScientist.EvoScientist import _get_default_middleware
|
|
|
|
middleware = _get_default_middleware()
|
|
memory_middleware = next(
|
|
m for m in middleware if type(m).__name__ == "EvoMemoryMiddleware"
|
|
)
|
|
|
|
assert [tool.name for tool in memory_middleware.tools] == [
|
|
"search_observations",
|
|
"read_memory",
|
|
"record_observation",
|
|
]
|
|
lifecycle_middleware = next(
|
|
m for m in middleware if type(m).__name__ == "EvoMemoryLifecycleMiddleware"
|
|
)
|
|
assert lifecycle_middleware._source_type == MemorySourceType.TURN
|
|
|
|
|
|
# ---------------------------------------------------------------------------
|
|
# Direct behavior test for ``_get_default_middleware`` filter
|
|
# ---------------------------------------------------------------------------
|
|
#
|
|
# The factory test above pins the *contract* (factory passes the flag).
|
|
# This test pins the *behavior* (the flag actually excludes
|
|
# AskUserMiddleware), so a future refactor that renames the flag or
|
|
# restructures the middleware list cannot silently re-introduce the
|
|
# interrupt-based deadlock.
|
|
|
|
|
|
@patch(
|
|
"EvoScientist.middleware.create_tool_selector_middleware",
|
|
return_value=[MagicMock()],
|
|
)
|
|
@patch("EvoScientist.EvoScientist._ensure_chat_model")
|
|
@patch("EvoScientist.EvoScientist._ensure_config")
|
|
def test_async_subagent_mode_filters_ask_user(
|
|
mock_config, mock_chat, mock_tool_selector
|
|
):
|
|
"""``_get_default_middleware(for_async_subagent=True)`` must drop
|
|
``AskUserMiddleware`` even when ``enable_ask_user`` is on.
|
|
|
|
Without mocking the middleware list itself: we let the real list be
|
|
constructed and assert ``AskUserMiddleware`` is absent. Mocks here
|
|
cover only the heavy dependencies (chat model, tool-selector) that
|
|
the middleware list builder pulls in transitively.
|
|
"""
|
|
cfg = MagicMock()
|
|
cfg.enable_ask_user = True # would normally include AskUserMiddleware
|
|
cfg.auto_mode = False
|
|
cfg.auto_approve = False
|
|
cfg.model_fallbacks = None
|
|
cfg.memory_profile_enabled = True
|
|
cfg.memory_observations_enabled = True
|
|
cfg.memory_observation_writer = MemoryObservationWriter.ALL
|
|
cfg.memory_workers_enabled = True
|
|
cfg.auxiliary_model = ""
|
|
cfg.auxiliary_provider = ""
|
|
mock_config.return_value = cfg
|
|
mock_chat.return_value = MagicMock(profile={"max_input_tokens": 200_000})
|
|
|
|
from EvoScientist.EvoScientist import _get_default_middleware
|
|
from EvoScientist.middleware.ask_user import AskUserMiddleware
|
|
|
|
# CLI / in-process path includes AskUserMiddleware …
|
|
cli_mw = _get_default_middleware()
|
|
assert any(isinstance(m, AskUserMiddleware) for m in cli_mw), (
|
|
"Sanity check: with enable_ask_user=True and CLI mode, "
|
|
"AskUserMiddleware should be present."
|
|
)
|
|
|
|
# … but the async-subagent path filters it out.
|
|
async_mw = _get_default_middleware(for_async_subagent=True)
|
|
assert not any(isinstance(m, AskUserMiddleware) for m in async_mw), (
|
|
"AskUserMiddleware leaked into async sub-agent middleware — its "
|
|
"interrupt() call deadlocks the deployed graph (no UI to resume)."
|
|
)
|
|
|
|
|
|
@patch(
|
|
"EvoScientist.middleware.create_tool_selector_middleware",
|
|
return_value=[MagicMock()],
|
|
)
|
|
@patch("EvoScientist.EvoScientist._ensure_chat_model")
|
|
@patch("EvoScientist.EvoScientist._ensure_config")
|
|
def test_async_subagent_disables_tool_selector_stream_tracking(
|
|
mock_config, mock_chat, mock_tool_selector
|
|
):
|
|
"""Async subagents still select tools, but must not drive main-agent UI state."""
|
|
cfg = MagicMock()
|
|
cfg.enable_ask_user = False
|
|
cfg.auto_mode = False
|
|
cfg.auto_approve = False
|
|
cfg.model_fallbacks = None
|
|
cfg.memory_profile_enabled = True
|
|
cfg.memory_observations_enabled = True
|
|
cfg.memory_observation_writer = MemoryObservationWriter.ALL
|
|
cfg.memory_workers_enabled = True
|
|
cfg.auxiliary_model = ""
|
|
cfg.auxiliary_provider = ""
|
|
mock_config.return_value = cfg
|
|
mock_chat.return_value = MagicMock(profile={"max_input_tokens": 200_000})
|
|
|
|
from EvoScientist.EvoScientist import _get_default_middleware
|
|
from EvoScientist.middleware.events import NoOpSink
|
|
|
|
_get_default_middleware(for_async_subagent=True)
|
|
|
|
# Async subagents still select tools, but are wired to the silent NoOpSink
|
|
# so they never drive the main-agent tool-selection widget.
|
|
mock_tool_selector.assert_called_once()
|
|
assert isinstance(mock_tool_selector.call_args.kwargs["events"], NoOpSink)
|