Files
EvoScientist-Multi/tests/test_background_middleware.py
T
dinos 01845f4311 refactor: route middleware display events through an injected event sink (#343)
* chore: add pytest-asyncio in auto mode

* test: migrate channel and stream tests to native async

Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.

* test: migrate command and model/middleware tests to native async

Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.

* test: migrate TUI, notifier, gateway, and session tests to native async

TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.

* test: replace direct asyncio.run() calls with native async tests

Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.

* test: drop undeclared anyio markers and delete run_async helper

The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).

* test: add autouse fixture for watcher cleanup

* refactor: remove redundant hasattr calls

* refactor: add typed middleware event sink and thread through assembly

Add MiddlewareEventSink protocol + NoOpSink in middleware/events.py
with a documented any-thread non-blocking contract (contract test uses a
deliberately-slow fake sink). Thread an optional `events` parameter
through create_cli_agent -> _get_default_middleware -> tool selector /
model fallback constructors; subagent stacks are always forced to
NoOpSink.

* refactor: inject a notifier port into async-watcher and background middleware

Add public pre_cancel_watcher() and enqueue_task_notification() to
cli/async_notifier.py and a small NotifierPort protocol
(middleware/notifier.py) that the module satisfies structurally.
AsyncWatcherMiddleware and BackgroundExecutionMiddleware now receive the
port by constructor injection at the composition root, deleting the lazy
'from ..cli import async_notifier' imports and the private
_watcher_by_thread / _enqueue pokes.

* refactor: invert tool-selection ownership onto a frontend event sink

The adaptive tool selector now reports on_tool_selection_started /
on_tool_selection / on_tool_selection_ended to the injected sink instead
of writing four process-global module variables. The frontend sink
(stream/sink.py FrontendEventSink) owns the selected/total/active state
with consume-once + dedup-vs-last-emitted semantics;
stream/tool_selection.py reads that sink object (a ToolSelectionView)
rather than reaching into tool_selector's globals.

Deleted: the 4 module globals, the cross-module mutations in
tool_selection.py, the track_stream_selection flag, the now-vestigial
_ToolSelectionTrackerMiddleware, reset_tool_selection_state_for_tests,
and the autouse conftest fixture. The sink is threaded from the two
interactive frontends through create_runtime_gateways ->
LocalGraphGateway (read side) and _load_agent -> create_cli_agent (write
side); subagent / headless stacks get NoOpSink.

* refactor: route model-fallback narration through the injected event sink

Delete the _ui_emit_fn / set_ui_emit module global and the
..stream.console import from model_fallback.py. The fallback middleware
now reports through its injected sink: the fallback transition via the
structured on_model_fallback (the frontend formats the '-> Falling back
to ...' line), and the surrounding narration (primary-failure header,
per-attempt outcome, exhaustion, non-fallbackable rejection) via
emit_fallback_notice, preserving the exact user-facing text. The TUI
binds its _append_system as the sink's fallback display where it used to
call set_ui_emit (cleared on exit); the Rich CLI's sink prints to the
console. _try_fallbacks / _guard_and_fallback take the sink.

* refactor: declare events on the GraphGateway protocol

Both gateway implementations now carry an explicit events attribute
(LangGraphServerGateway holds None — no frontend renders middleware
events across the HTTP boundary), so the four call sites use plain
attribute access instead of getattr probing an implicit contract.

* refactor: bind fallback display via the closure-scoped concrete sink

The App methods used gateway.events (typed as the read-side view) and
hasattr-probed for the concrete FrontendEventSink API. The enclosing
factory creates that sink two hundred lines up — close over it directly:
no probing, fully typed, and it becomes a constructor parameter
naturally when the App class is hoisted out of the factory.

* fix: end tool selection before fallback handler

* fix: keep fallback display errors non-fatal

* fix: preserve selector suppression for default streams

* fix: restore fallback notice console display

* refactor: consolidate fallback narration events

* refactor: clean middleware event sink plumbing

* fix: type gateway session events

* refactor: make all event protocols runtime-checkable

MiddlewareEventSink already carried @runtime_checkable (the stream
binding guard isinstance-checks it); ToolSelectionView and SessionEvents
now match, so mirroring that pattern against any of the three protocols
works instead of raising TypeError.

* fix(cli): close QuickJS workers after one-shot failures

* fix(cli): honor no-thinking in final output

* fix(channels): report failed startup accurately

* fix(channels): make Telegram cleanup idempotent

* fix(tui): skip command sync during exit

* fix(channels): preserve startup state during retries

* refactor(channels): share pending startup status

* refactor(cli): expose channel startup snapshot

* fix(tui): move channel startup off event loop

* test(channels): release retry gate on assertion failure

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-14 22:34:17 +00:00

330 lines
12 KiB
Python

"""Tests for BackgroundExecutionMiddleware and its tools."""
import sys
import time
import pytest
from EvoScientist import background as bg
from EvoScientist.cli import async_notifier
from EvoScientist.middleware.background import (
BackgroundExecutionMiddleware,
_make_run_in_background,
check_process,
list_processes,
stop_process,
)
def _run_bg(*, dangerous: bool = False, notifier=async_notifier):
"""Build the injected ``run_in_background`` tool for direct-invoke tests."""
return _make_run_in_background(notifier, dangerous)
def _sleep_cmd(seconds: int) -> str:
"""Cross-platform command that sleeps for *seconds* and exits 0."""
if sys.platform == "win32":
return f"ping -n {seconds + 1} 127.0.0.1 > nul"
return f"sleep {seconds}"
def _true_cmd() -> str:
"""Cross-platform command that exits 0 immediately."""
if sys.platform == "win32":
return "cmd /c exit /b 0"
return "true"
def _wait_until(predicate, timeout=4.0, interval=0.05):
"""Poll ``predicate`` until true or ``timeout`` — avoids flaky fixed sleeps on slow CI."""
deadline = time.time() + timeout
while time.time() < deadline:
if predicate():
return True
time.sleep(interval)
return False
@pytest.fixture(autouse=True)
def _clean_registry():
from EvoScientist.cli import async_notifier
bg._PROCESSES.clear()
async_notifier.drain_notifications(None)
yield
for proc in list(bg._PROCESSES.values()):
try:
proc.popen.kill()
except Exception:
pass
bg._PROCESSES.clear()
async_notifier.drain_notifications(None)
def test_middleware_registers_four_tools():
mw = BackgroundExecutionMiddleware(async_notifier)
names = {t.name for t in mw.tools}
assert names == {
"run_in_background",
"check_process",
"stop_process",
"list_processes",
}
def test_no_job_in_tool_names():
"""Naming ADR: the word 'job' must not appear in the tool surface."""
mw = BackgroundExecutionMiddleware(async_notifier)
assert not any("job" in t.name.lower() for t in mw.tools)
def test_run_rejects_dangerous_command_without_launching(monkeypatch):
launched = {"called": False}
def _spy(*args, **kwargs):
launched["called"] = True
return "should-not-happen"
monkeypatch.setattr(bg, "launch", _spy)
out = _run_bg().invoke({"command": "sudo rm -rf /"})
assert launched["called"] is False
assert "blocked" in out.lower()
def test_run_launches_valid_command(tmp_path, monkeypatch):
# Pin the workspace cwd to a temp dir so the launch is isolated.
monkeypatch.setattr("EvoScientist.paths.resolve_virtual_path", lambda _vp: tmp_path)
out = _run_bg().invoke({"command": "echo ok", "name": "demo"})
assert "Started background process" in out
assert "check_process" in out
assert len(bg._PROCESSES) == 1
def test_run_applies_virtual_path_rewriting(tmp_path, monkeypatch):
"""run_in_background must rewrite virtual paths like execute (shared preprocessing)."""
monkeypatch.setattr("EvoScientist.paths.resolve_virtual_path", lambda _vp: tmp_path)
captured = {}
def _spy(command, cwd, name=None, *, origin_thread_id=None, on_exit=None):
captured["command"] = command
return "pidX"
monkeypatch.setattr(bg, "launch", _spy)
_run_bg().invoke({"command": "python /train.py"})
# virtual absolute path -> workspace-relative, same as execute would produce
assert captured["command"] == "python ./train.py"
def test_run_dangerous_allows_real_path_no_rewrite(tmp_path, monkeypatch):
"""In dangerous mode, background commands keep real absolute paths (parity with execute)."""
monkeypatch.setattr("EvoScientist.paths.resolve_virtual_path", lambda _vp: tmp_path)
captured = {}
def _spy(command, cwd, name=None, *, origin_thread_id=None, on_exit=None):
captured["command"] = command
return "pidX"
monkeypatch.setattr(bg, "launch", _spy)
# Absolute path + traversal would be BLOCKED in normal mode; allowed here.
out = _run_bg(dangerous=True).invoke({"command": "cat /etc/hosts && cat ../x"})
assert "blocked" not in out.lower()
assert captured["command"] == "cat /etc/hosts && cat ../x" # no ./ rewrite
# Advertised log path is the real path, not the virtual /.bg_processes/.
assert f"{tmp_path}/.bg_processes/" in out
assert "Output -> /.bg_processes/" not in out
def test_run_dangerous_still_blocks_privileged_command(tmp_path, monkeypatch):
"""Dangerous mode must NOT relax the privileged-command blocklist."""
monkeypatch.setattr("EvoScientist.paths.resolve_virtual_path", lambda _vp: tmp_path)
launched = {"called": False}
def _spy(*args, **kwargs):
launched["called"] = True
return "should-not-happen"
monkeypatch.setattr(bg, "launch", _spy)
out = _run_bg(dangerous=True).invoke({"command": "sudo rm x"})
assert launched["called"] is False
assert "blocked" in out.lower()
def test_run_enqueues_completion_notification(tmp_path, monkeypatch):
"""A finished background process enqueues a shell completion notification."""
from EvoScientist.cli import async_notifier
monkeypatch.setattr("EvoScientist.paths.resolve_virtual_path", lambda _vp: tmp_path)
_run_bg().invoke({"command": _true_cmd(), "name": "quick"})
# drain consumes, so accumulate across polls until the watcher's on_exit enqueues.
notifs = []
deadline = time.time() + 4.0
while time.time() < deadline:
notifs.extend(async_notifier.drain_notifications(None))
if any(n.kind == "bg-process" for n in notifs):
break
time.sleep(0.05)
assert any(n.kind == "bg-process" and n.status == "success" for n in notifs)
def test_origin_thread_id_reads_runtime_config():
"""thread_id is read from runtime.config['configurable'] (graph-injected)."""
from types import SimpleNamespace
from EvoScientist.middleware.background import _origin_thread_id
runtime = SimpleNamespace(config={"configurable": {"thread_id": "T-7"}})
assert _origin_thread_id(runtime) == "T-7"
assert _origin_thread_id(None) is None # direct .invoke() / no runtime
def test_notify_done_routes_to_origin_thread(tmp_path):
"""_notify_done enqueues the completion notification to the launching thread."""
from EvoScientist.cli import async_notifier
from EvoScientist.middleware.background import _notify_done
pid = bg.launch(_true_cmd(), str(tmp_path)) # no on_exit -> no auto-notify here
assert _wait_until(lambda: bg._PROCESSES[pid].finished_ts is not None)
_notify_done(bg._PROCESSES[pid], "T-123", async_notifier)
routed = async_notifier.drain_notifications("T-123")
assert any(n.task_id == pid and n.origin_cli_thread_id == "T-123" for n in routed)
def test_stopped_process_suppresses_notification(tmp_path, monkeypatch):
"""A user-stopped process must NOT emit a completion notification."""
from EvoScientist.cli import async_notifier
monkeypatch.setattr("EvoScientist.paths.resolve_virtual_path", lambda _vp: tmp_path)
_run_bg().invoke({"command": _sleep_cmd(600)})
(pid,) = list(bg._PROCESSES.keys())
stop_process.invoke({"process_id": pid})
# Wait until the watcher observed the exit — it would have enqueued here if the
# process weren't user-stopped. _notify_done is a no-op for stopped processes.
assert _wait_until(lambda: bg._PROCESSES[pid].finished_ts is not None)
notifs = async_notifier.drain_notifications(None)
assert not any(n.task_id == pid for n in notifs)
def test_checked_after_exit_dedups_notification(tmp_path):
"""Agent checking a finished process suppresses its completion notification."""
from EvoScientist.cli.async_notifier import (
AsyncTaskNotification,
dedup_notifications,
)
pid = bg.launch(_true_cmd(), str(tmp_path))
assert _wait_until(lambda: bg._PROCESSES[pid].finished_ts is not None)
bg.status(pid) # agent checks AFTER exit
assert bg.was_observed_done(pid) is True
n = AsyncTaskNotification(
task_id=pid,
agent_name="x",
status="success",
received_at="t",
kind="bg-process",
)
assert dedup_notifications([n], {}) == [] # deduped
def test_not_checked_after_exit_keeps_notification(tmp_path):
"""A finished process the agent never checked still notifies."""
from EvoScientist.cli.async_notifier import (
AsyncTaskNotification,
dedup_notifications,
)
pid = bg.launch(_true_cmd(), str(tmp_path))
assert _wait_until(
lambda: bg._PROCESSES[pid].finished_ts is not None
) # exit, but do NOT check
assert bg.was_observed_done(pid) is False
n = AsyncTaskNotification(
task_id=pid,
agent_name="x",
status="success",
received_at="t",
kind="bg-process",
)
assert dedup_notifications([n], {}) == [n] # survives
def test_shell_notification_renders_own_background_frame():
"""Shell notifications render under '✦ Background ✦', not 'Agent Teams'."""
from EvoScientist.cli.async_notifier import (
AsyncTaskNotification,
format_notification_lines,
)
n = AsyncTaskNotification(
task_id="fe60ce9c",
agent_name="test-20s",
status="success",
received_at="",
prompt="python train.py",
kind="bg-process",
)
lines = format_notification_lines([n])
top, body = lines[0][0], lines[1][0]
assert "Background" in top
assert "Agent Teams" not in top
assert "test-20s" in body
assert "Cmd:" in body
def test_mixed_notifications_render_two_frames():
"""A mixed batch shows both an Agent Teams frame and a Background frame."""
from EvoScientist.cli.async_notifier import (
AsyncTaskNotification,
format_notification_lines,
)
task = AsyncTaskNotification("t1", "writing-agent", "success", "", "")
shell = AsyncTaskNotification("p1", "demo", "success", "", "", kind="bg-process")
blob = "\n".join(t for t, _ in format_notification_lines([task, shell]))
assert "Agent Teams" in blob
assert "Background" in blob
def test_shell_notification_hints_check_process():
"""format_batch_message points shell processes to check_process, not check_async_task."""
from EvoScientist.cli.async_notifier import (
AsyncTaskNotification,
format_batch_message,
)
n = AsyncTaskNotification(
task_id="ab12",
agent_name="demo",
status="success",
received_at="x",
kind="bg-process",
)
msg = format_batch_message([n])
assert "check_process" in msg
assert "check_async_task" not in msg # shell-only batch -> no sub-agent hint
def test_check_and_list_route_to_manager(tmp_path, monkeypatch):
monkeypatch.setattr("EvoScientist.paths.resolve_virtual_path", lambda _vp: tmp_path)
_run_bg().invoke({"command": _sleep_cmd(1)})
(pid,) = bg._PROCESSES.keys()
assert pid in check_process.invoke({"process_id": pid})
assert pid in list_processes.invoke({})
assert "Stopped" in stop_process.invoke(
{"process_id": pid}
) or "finished" in stop_process.invoke({"process_id": pid})
def test_list_processes_forwards_all_threads(monkeypatch):
"""The all_threads tool arg is forwarded to background.list_all(include_all=...)."""
captured = {}
def _spy(thread_id=None, *, include_all=False):
captured["include_all"] = include_all
return "ok"
monkeypatch.setattr(bg, "list_all", _spy)
list_processes.invoke({"all_threads": True})
assert captured["include_all"] is True
list_processes.invoke({})
assert captured["include_all"] is False