01845f4311
* chore: add pytest-asyncio in auto mode * test: migrate channel and stream tests to native async Convert run_async() wrapper tests to plain 'async def test_*' under pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a coroutine awaited at every call site. * test: migrate command and model/middleware tests to native async Convert run_async() wrappers (import, alias, and fixture forms) to plain 'async def test_*'. Multi-call tests merge onto one loop as sequential awaits; none asserted on loop identity. * test: migrate TUI, notifier, gateway, and session tests to native async TUI/notifier/gateway files convert run_async wrappers to plain async tests. test_sessions.py's unittest.TestCase classes move to unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async methods on plain TestCase; converting blindly would have made ~70 tests silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget in test_tui_widgets.py drops its TestCase base for the same reason. * test: replace direct asyncio.run() calls with native async tests Convert tests that called asyncio.run() (directly or via a local _run helper) to plain 'async def test_*'; delete the local helpers. * test: drop undeclared anyio markers and delete run_async helper The @pytest.mark.anyio tests relied on anyio being a transitive dep of httpx; auto-mode pytest-asyncio collects them natively. run_async() and its fixture are unreferenced after the migration, so remove them — pytest-asyncio's per-test loop teardown covers the pending-task cancellation the helper existed for (verified: full suite runs with no 'Event loop is closed' errors or destroyed-task warnings). * test: add autouse fixture for watcher cleanup * refactor: remove redundant hasattr calls * refactor: add typed middleware event sink and thread through assembly Add MiddlewareEventSink protocol + NoOpSink in middleware/events.py with a documented any-thread non-blocking contract (contract test uses a deliberately-slow fake sink). Thread an optional `events` parameter through create_cli_agent -> _get_default_middleware -> tool selector / model fallback constructors; subagent stacks are always forced to NoOpSink. * refactor: inject a notifier port into async-watcher and background middleware Add public pre_cancel_watcher() and enqueue_task_notification() to cli/async_notifier.py and a small NotifierPort protocol (middleware/notifier.py) that the module satisfies structurally. AsyncWatcherMiddleware and BackgroundExecutionMiddleware now receive the port by constructor injection at the composition root, deleting the lazy 'from ..cli import async_notifier' imports and the private _watcher_by_thread / _enqueue pokes. * refactor: invert tool-selection ownership onto a frontend event sink The adaptive tool selector now reports on_tool_selection_started / on_tool_selection / on_tool_selection_ended to the injected sink instead of writing four process-global module variables. The frontend sink (stream/sink.py FrontendEventSink) owns the selected/total/active state with consume-once + dedup-vs-last-emitted semantics; stream/tool_selection.py reads that sink object (a ToolSelectionView) rather than reaching into tool_selector's globals. Deleted: the 4 module globals, the cross-module mutations in tool_selection.py, the track_stream_selection flag, the now-vestigial _ToolSelectionTrackerMiddleware, reset_tool_selection_state_for_tests, and the autouse conftest fixture. The sink is threaded from the two interactive frontends through create_runtime_gateways -> LocalGraphGateway (read side) and _load_agent -> create_cli_agent (write side); subagent / headless stacks get NoOpSink. * refactor: route model-fallback narration through the injected event sink Delete the _ui_emit_fn / set_ui_emit module global and the ..stream.console import from model_fallback.py. The fallback middleware now reports through its injected sink: the fallback transition via the structured on_model_fallback (the frontend formats the '-> Falling back to ...' line), and the surrounding narration (primary-failure header, per-attempt outcome, exhaustion, non-fallbackable rejection) via emit_fallback_notice, preserving the exact user-facing text. The TUI binds its _append_system as the sink's fallback display where it used to call set_ui_emit (cleared on exit); the Rich CLI's sink prints to the console. _try_fallbacks / _guard_and_fallback take the sink. * refactor: declare events on the GraphGateway protocol Both gateway implementations now carry an explicit events attribute (LangGraphServerGateway holds None — no frontend renders middleware events across the HTTP boundary), so the four call sites use plain attribute access instead of getattr probing an implicit contract. * refactor: bind fallback display via the closure-scoped concrete sink The App methods used gateway.events (typed as the read-side view) and hasattr-probed for the concrete FrontendEventSink API. The enclosing factory creates that sink two hundred lines up — close over it directly: no probing, fully typed, and it becomes a constructor parameter naturally when the App class is hoisted out of the factory. * fix: end tool selection before fallback handler * fix: keep fallback display errors non-fatal * fix: preserve selector suppression for default streams * fix: restore fallback notice console display * refactor: consolidate fallback narration events * refactor: clean middleware event sink plumbing * fix: type gateway session events * refactor: make all event protocols runtime-checkable MiddlewareEventSink already carried @runtime_checkable (the stream binding guard isinstance-checks it); ToolSelectionView and SessionEvents now match, so mirroring that pattern against any of the three protocols works instead of raising TypeError. * fix(cli): close QuickJS workers after one-shot failures * fix(cli): honor no-thinking in final output * fix(channels): report failed startup accurately * fix(channels): make Telegram cleanup idempotent * fix(tui): skip command sync during exit * fix(channels): preserve startup state during retries * refactor(channels): share pending startup status * refactor(cli): expose channel startup snapshot * fix(tui): move channel startup off event loop * test(channels): release retry gate on assertion failure --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
172 lines
5.6 KiB
Python
172 lines
5.6 KiB
Python
"""Shared fixtures for EvoScientist tests."""
|
|
|
|
from pathlib import Path
|
|
|
|
import pytest
|
|
|
|
_NONEXISTENT_DOTENV = str(Path(__file__).with_name(".pytest-dotenv-does-not-exist"))
|
|
|
|
|
|
@pytest.fixture
|
|
def sample_tool_call():
|
|
"""A minimal tool call dict."""
|
|
return {"id": "tc_001", "name": "execute", "args": {"command": "ls -la"}}
|
|
|
|
|
|
@pytest.fixture
|
|
def sample_tool_result():
|
|
"""A minimal tool result dict."""
|
|
return {
|
|
"id": "tc_001",
|
|
"name": "execute",
|
|
"content": "[OK] file1.py file2.py",
|
|
"success": True,
|
|
}
|
|
|
|
|
|
@pytest.fixture
|
|
def sample_events():
|
|
"""A sequence of stream event dicts covering common types."""
|
|
return [
|
|
{"type": "thinking", "content": "Let me think..."},
|
|
{"type": "text", "content": "Here is the answer."},
|
|
{
|
|
"type": "tool_call",
|
|
"id": "tc_001",
|
|
"name": "execute",
|
|
"args": {"command": "ls"},
|
|
},
|
|
{
|
|
"type": "tool_result",
|
|
"id": "tc_001",
|
|
"name": "execute",
|
|
"content": "[OK] done",
|
|
"success": True,
|
|
},
|
|
{
|
|
"type": "subagent_start",
|
|
"name": "research-agent",
|
|
"description": "Find papers",
|
|
"instance_id": "task:research",
|
|
"tool_call_id": "tc_task_001",
|
|
},
|
|
{
|
|
"type": "subagent_tool_call",
|
|
"subagent": "research-agent",
|
|
"instance_id": "task:research",
|
|
"name": "tavily_search",
|
|
"args": {"query": "test"},
|
|
"id": "tc_sa_001",
|
|
},
|
|
{
|
|
"type": "subagent_tool_result",
|
|
"subagent": "research-agent",
|
|
"instance_id": "task:research",
|
|
"name": "tavily_search",
|
|
"content": "Results...",
|
|
"success": True,
|
|
"id": "tc_sa_001",
|
|
},
|
|
{
|
|
"type": "subagent_end",
|
|
"name": "research-agent",
|
|
"instance_id": "task:research",
|
|
},
|
|
{"type": "done", "response": "Here is the answer."},
|
|
]
|
|
|
|
|
|
@pytest.fixture
|
|
def tmp_workspace(tmp_path):
|
|
"""Provide a temporary workspace directory path."""
|
|
ws = tmp_path / "workspace"
|
|
ws.mkdir()
|
|
return str(ws)
|
|
|
|
|
|
@pytest.fixture
|
|
def runtime_paths(tmp_path, monkeypatch):
|
|
"""Isolate ``langgraph_dev.manager.RUNTIME`` under a temp directory.
|
|
|
|
Replaces the module-level ``RUNTIME`` with a fully temp-rooted bundle
|
|
so every path (``pid_dir``, ``pid_file``, ``log_file``,
|
|
``workspace_sidecar``, ``lock_file``) is contained under ``tmp_path``.
|
|
|
|
Tests that need a variant of a single field can still call
|
|
``dataclasses.replace(runtime_paths, log_file=…)`` etc. — the
|
|
baseline is already isolated, so forgetting a field just keeps it
|
|
under ``tmp_path``, never ``~/.config/evoscientist``.
|
|
"""
|
|
from EvoScientist.langgraph_dev import manager
|
|
|
|
runtime = manager.LanggraphRuntimePaths.for_directory(tmp_path / "runtime")
|
|
monkeypatch.setattr(manager, "RUNTIME", runtime)
|
|
return runtime
|
|
|
|
|
|
# Capture deepagents tool factories at conftest load time — BEFORE any test
|
|
# imports EvoScientist, which can trigger ``_patch_deepagents_model_passthrough``
|
|
# during agent construction. Once captured here, the ``restore_model_passthrough_patch``
|
|
# fixture has a stable "truly unpatched" baseline to reset to between tests, even
|
|
# if upstream code paths apply the patch as a side effect.
|
|
try:
|
|
from deepagents.middleware import async_subagents as _ds_async_subagents
|
|
|
|
_DEEPAGENTS_ORIGINAL_BUILD_START = _ds_async_subagents._build_start_tool
|
|
_DEEPAGENTS_ORIGINAL_BUILD_UPDATE = _ds_async_subagents._build_update_tool
|
|
except Exception:
|
|
_ds_async_subagents = None
|
|
_DEEPAGENTS_ORIGINAL_BUILD_START = None
|
|
_DEEPAGENTS_ORIGINAL_BUILD_UPDATE = None
|
|
|
|
|
|
@pytest.fixture
|
|
def restore_model_passthrough_patch():
|
|
"""Reset deepagents internals + ``_model_passthrough_patched`` to unpatched.
|
|
|
|
The model-passthrough patch wraps ``deepagents.middleware.async_subagents``
|
|
module-level functions in place. The originals are captured at conftest
|
|
load time (above) so this fixture can always start each test from a
|
|
known-unpatched state regardless of what other tests / agent fixtures
|
|
did to the module before.
|
|
"""
|
|
from EvoScientist.llm import patches as patches_mod
|
|
|
|
if _ds_async_subagents is None:
|
|
# deepagents not importable — fixture is a no-op (the patch fn itself
|
|
# returns early in that case).
|
|
yield
|
|
return
|
|
|
|
def _reset() -> None:
|
|
_ds_async_subagents._build_start_tool = _DEEPAGENTS_ORIGINAL_BUILD_START
|
|
_ds_async_subagents._build_update_tool = _DEEPAGENTS_ORIGINAL_BUILD_UPDATE
|
|
patches_mod._model_passthrough_patched = False
|
|
|
|
_reset()
|
|
try:
|
|
yield
|
|
finally:
|
|
_reset()
|
|
|
|
|
|
@pytest.fixture(autouse=True)
|
|
def _isolate_dotenv(monkeypatch):
|
|
"""Keep the developer's real .env out of the test environment.
|
|
|
|
``get_effective_config`` runs ``load_dotenv(find_dotenv(usecwd=True),
|
|
override=True)``, so any test that loads config injects the repo's
|
|
real .env into ``os.environ`` for the rest of the pytest process.
|
|
An empty-valued line like ``MINIMAX_BASE_URL=`` then makes
|
|
``os.environ.get(key, default)`` return "" instead of the default,
|
|
breaking unrelated tests later in the run (see issue #322).
|
|
|
|
Pointing ``find_dotenv`` at a fixed path that does not exist makes
|
|
``load_dotenv`` a no-op without creating a temporary directory for
|
|
every test.
|
|
"""
|
|
monkeypatch.setattr(
|
|
"EvoScientist.config.settings.find_dotenv",
|
|
lambda *args, **kwargs: _NONEXISTENT_DOTENV,
|
|
)
|