Files
EvoScientist-Multi/tests/test_agent_loader.py
T
dinos 65828e9666 perf(cli): cut startup latency and defer MCP loading to the background (#171)
* perf(cli): cut startup time of `evosci --help` from ~2.2s to ~0.3s

Module-level imports were eagerly pulling in langchain.chat_models (with
the whole anthropic/openai/google stack), langgraph, textual, and
prompt_toolkit on every invocation — even for `--help` or `config list`.

Defer those with PEP 562 `__getattr__`, using `lazy_loader.attach` (SPEC-1,
the scientific-python standard) where it's a clean attach pattern:

- `EvoScientist/llm/__init__.py`: attach `models` lazily so importing
  `context_window` from this package no longer drags in langchain.
- `EvoScientist/stream/__init__.py`: attach display/events lazily; split
  the shared Rich `Console` singleton into a new lightweight
  `stream/console.py` so callers that only need `console` skip the
  `stream.events` → `langchain_core.messages` chain.
- `EvoScientist/cli/__init__.py`: hand-rolled `__getattr__` (reaches into
  `..stream.state`, which `lazy_loader` doesn't cover) so `commands` and
  `app` are the only eager loads.
- `EvoScientist/cli/commands.py`: move `cmd_interactive`/`cmd_run` to
  in-function imports so prompt_toolkit + textual only load when the
  interactive path actually runs.
- `EvoScientist/cli/_constants.py`: read `AGENT_NAME` on demand so
  `build_metadata` doesn't eagerly import `sessions` (langgraph/aiosqlite).

Adds `lazy-loader>=0.5` as a dependency.

* feat(cli): defer MCP tool loading with live per-server progress

The CLI was blocking ~5 s on MCP tool enumeration before the first
prompt appeared. Move the agent construction off the event loop and
surface per-server progress so the user can interact immediately and see
what's happening.

MCP client:
- Add an `on_progress` callback to `load_mcp_tools` / `aload_mcp_tools`
  / `_load_tools` emitting `start` / `success` / `error` events per
  server.
- Fan connection attempts out with `asyncio.gather` so latency no longer
  scales linearly with server count; cap simultaneous attempts at
  `_MAX_CONCURRENT_CONNECTIONS` (8) via a semaphore so a big stdio fleet
  doesn't spawn every subprocess at once.

Agent wiring:
- Plumb `on_mcp_progress` through `create_cli_agent` / `_load_agent` /
  `load_mcp_and_build_kwargs` so CLI and TUI can plug in collectors.

CLI (`cmd_interactive`):
- Run `_load_agent` in a background thread via `asyncio.to_thread`; the
  prompt and banner render immediately.
- `_await_agent_ready()` awaits the task before each agent-using site
  (first turn, channel messages, `/channel`, `/compact`). Raises if
  called without a prior `_start_agent_load` instead of silently
  reloading without the SQLite checkpointer.
- Pre-prime the progress dict from `load_mcp_config()` so the
  bottom-toolbar's `N/M` denominator is stable from the first render.
- Wrap `session.prompt_async` in `patch_stdout(raw=True)` so
  `console.print` from the worker-thread progress callback lands cleanly
  above the prompt as inline chat messages instead of stomping the
  prompt cursor.

TUI (`EvoTextualInteractiveApp`):
- Same background load + `_await_agent_ready()` gates on every
  `self._agent` read.
- New `MCPLoaderWidget` mounted at the top of `#input-shell` shows a
  header with `N/M` and one live row per server (spinner → ✓ / ✗ with
  tool count or error detail). On completion:
  - all-clean loads auto-dismiss ~2.5 s later;
  - cache hits (no events ever fired) dismiss immediately rather than
    flashing a misleading "0/N loaded";
  - failures keep the widget mounted so the user can read the errors.
  - `dismissed` property lets the app clear its ref so late events from
    slow servers become no-ops. The error branch of `_on_agent_loaded`
    also settles the widget so a load failure can't leave the spinner
    animating forever.
- Chat input is `disabled` while MCP resolves — no placeholder hack, no
"waiting…" system message.

Shared:
- Hoist braille spinner frames to `status_bar.SPINNER_FRAMES` and import
  them in the TUI widget so CLI and TUI animate in sync.

Tests:
- Extend `test_agent_mcp_cache` fakes to accept the new `on_progress`
  kwarg.
- New `TestLoadToolsProgressCallback` in `test_mcp_client` exercises the
  event sequence for success/failure/mixed fleets, verifies a buggy
  callback doesn't break the load, and asserts the semaphore caps
  in-flight connections.

* style: ruff

* chore: update uv.lock

* chore: uv.lock

* fix: coderabbit issues

* style: fmt

* fix: move _await_agent_ready inside try block

* fix(tui): auto-dismiss MCP loader widget on failure

The widget was designed to stay mounted on failure so the user could
read error detail, but since it's pinned above the input it never went
away in practice — just permanent banner clutter.

Auto-dismiss on failure too, with a longer grace (12s vs 2.5s) so the
error summary stays readable.

* fix: address second coderabbit pass

- Channel handlers (CLI + TUI): catch agent-load failures so the
  channel request doesn't hang; CLI moves `_await_agent_ready()`
  inside the existing try/except, TUI catches explicitly and calls
  `_set_channel_response` with the error.

- Stale background loads: `prev.cancel()` only stops the asyncio
  wrapper, not the thread running `_load_agent`. Added a generation
  token (`agent_load_id` / `self._agent_load_id`) and gated both
  progress and completion callbacks on it so a superseded load can't
  clobber the current session's state or UI.

- TUI prompt lifecycle: added `_agent_load_pending()` and gated the
  `_process_channel_message` / `_handle_command` finally blocks on it
  so `/new` or `/resume` invoked from a command keeps the prompt
  disabled until the fresh load settles.

- TUI readiness failures: `_run_turn` and `_handle_command` now
  catch exceptions from `_await_agent_ready()` and surface a
  "Agent failed to load: …" system message instead of letting the
  exception escape into Textual's traceback panel.

* refactor(cli): share background agent loader between CLI and TUI

The CLI and TUI were carrying near-identical copies of the same
background-load state machine: the `agent_task`, the `agent_load_id`
generation token, the gated progress/completion callbacks, and the
per-server progress dict. Every CodeRabbit finding on that lifecycle
had to be fixed in both files.

Extract it into `cli/_agent_loader.py`:

- `MCPProgressTracker` — owns the `server -> (state, detail)` dict;
  exposes `prime`, `record`, `snapshot`, `totals`.

- `BackgroundAgentLoader` — owns `agent`, the in-flight task, and the
  generation token. Exposes `start(**loader_kwargs)`, `await_ready()`,
  `is_pending`. Internally gates all progress/completion callbacks by
  generation so a superseded load can't clobber the current session.
  UI-specific rendering plugs in via `on_progress` / `on_success` /
  `on_failure` callbacks.

Both surfaces now just wire their UI hooks; the loader file holds no
Rich / prompt_toolkit / Textual dependencies. Net -345 lines from
`interactive.py` + `tui_interactive.py`; +20 unit tests pinning the
lifecycle (generation filtering, cache-hit short-circuit, failure
reset, progress ordering).

* refactor(cli): make _on_done the sole authority for agent state transitions

await_ready no longer sets self.agent — it just awaits the task and
reads what _on_done already wrote. Eliminates the dual-write overlap
(asyncio guarantees done-callbacks fire in registration order).

* fix(tui): let users type during MCP load, only block on send

Remove prompt-disabling during background agent load — the TUI now
matches the CLI approach where the input stays enabled and only gates
on await_ready() at submit time. The MCPLoaderWidget still provides
visual feedback that loading is in progress.

* fix(loader): preserve real load error on await_ready; dedup failure message

CodeRabbit flagged two issues with the new loader:

1. After a failed load, `_on_done` nulled `self._task`, so the next
   `await_ready()` hit the "before start()" branch and the CLI wrapper
   remapped it to a misleading "checkpointer not available" message —
   losing the real exception (bad MCP config, network, etc.).

   Keep `_task` set on failure so `await_ready` re-raises the real
   exception. Added `needs_restart` so TUI's auto-retry check stays a
   one-liner and doesn't need to reach into task internals.

2. TUI reported each load failure twice: once from
   `_on_agent_load_failure` (the done-callback) and once from each
   caller of `_await_agent_ready` (`_run_turn`,
   `_process_channel_message`, `_handle_command`) catching the re-raise.

   `_on_agent_load_failure` is now the sole local reporter; callers
   just handle control flow (return cleanly, set channel response to
   unblock remote).

* fix(cli): wire /model handler through the agent loader

The /model command from main (merged via f1f0d7c) still reached for
`state["agent"]` (CLI) and `self._agent` (TUI) — both removed by the
background-loader refactor. CLI raised KeyError on first invocation;
TUI raised AttributeError. Writes to the old fields also had no effect
because every other code path now reads from `agent_loader.agent`, so
the model switch would have silently failed.

Route everything through the loader: `await _await_agent_ready()` up
front so /model doesn't race with the initial background load, build
the `CommandContext` with the current agent, and sync `ctx.agent` back
into `agent_loader.agent` (plus channel globals) when the command
replaces it.

* fix(cli): isolate progress callback, capture awaited agent, gate by requires_agent

Three CodeRabbit findings on the loader + command dispatch path:

- Wrap ``_on_progress`` in try/except inside the loader's gated wrapper
  so a buggy UI adapter can't bubble into ``loader_fn`` and fail the
  whole background load. The MCP client already protects this, but
  defence-in-depth keeps the loader self-contained.

- In CLI channel + main-loop streaming, capture the agent returned by
  ``_await_agent_ready()`` and pass that into ``run_streaming`` rather
  than reading ``agent_loader.agent`` after a subsequent ``await``.
  A concurrent ``/new``/``/resume``/``/model`` could have swapped it.

- Add ``requires_agent: ClassVar[bool] = False`` to ``Command`` and
  mark ``/compact``, ``/model``, ``/channel`` as ``True``. TUI dispatch
  sites (channel and keyboard) now check ``cmd_manager.resolve(...)``
  and only wait for readiness when the command actually needs the
  agent. ``/mcp add``, ``/skills``, ``/new`` etc. no longer deadlock
  behind a failing MCP load they are meant to fix.

* fix(cli): guard sync-back, subcommand-aware gating, /model adopt-path

Three CodeRabbit findings on command dispatch:

- ``_handle_command`` unconditionally synced ``ctx.agent`` back into
  ``agent_loader``.  For non-agent commands ``ctx.agent`` is ``None``,
  so ``/threads`` / ``/mcp`` / ``/skills`` (etc.) could clobber a valid
  loaded agent — and rebind channel globals to ``None``.  Guard the
  sync on ``ctx.agent is not None``.

- ``/channel status`` and ``/channel stop`` don't touch ``ctx.agent``
  but the class-level ``requires_agent = True`` blocked them behind
  agent readiness.  Added ``Command.needs_agent(args)`` (defaults to
  ``requires_agent``) so ``/channel`` can override with subcommand
  awareness; kept the class flag for the common case.

- ``/model`` builds a new agent from scratch, it never reads the
  existing one — gating it on readiness meant a broken provider
  blocked the command that would fix it.  Flipped it to
  ``requires_agent = False`` and added ``BackgroundAgentLoader.adopt``
  so the UI can seat the replacement and supersede any in-flight
  load (the generation token keeps a late completion from clobbering
  the adopted agent).

Bonus cleanup: ``CommandManager.resolve`` now returns
``(command, args)`` so callers can invoke ``needs_agent`` without
re-implementing ``shlex`` parsing.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-22 18:22:33 +02:00

361 lines
12 KiB
Python

"""Tests for ``cli/_agent_loader``."""
from __future__ import annotations
import asyncio
import pytest
from EvoScientist.cli._agent_loader import BackgroundAgentLoader, MCPProgressTracker
# ──────────────────────────────────────────────────────────────────────
# MCPProgressTracker
# ──────────────────────────────────────────────────────────────────────
class TestMCPProgressTracker:
def test_prime_empty_when_no_config(self, monkeypatch):
import EvoScientist.mcp as mcp_pkg
monkeypatch.setattr(mcp_pkg, "load_mcp_config", lambda: {})
t = MCPProgressTracker()
t.prime()
assert t.progress == {}
def test_prime_seeds_pending_entries(self, monkeypatch):
import EvoScientist.mcp as mcp_pkg
monkeypatch.setattr(mcp_pkg, "load_mcp_config", lambda: {"a": {}, "b": {}})
t = MCPProgressTracker()
t.prime()
assert t.progress == {"a": ("pending", ""), "b": ("pending", "")}
def test_prime_swallows_config_errors(self, monkeypatch):
import EvoScientist.mcp as mcp_pkg
def _boom():
raise RuntimeError("config broken")
monkeypatch.setattr(mcp_pkg, "load_mcp_config", _boom)
t = MCPProgressTracker()
t.prime()
assert t.progress == {}
def test_record_maps_events(self):
t = MCPProgressTracker()
assert t.record("start", "srv", "") == "pending"
assert t.record("success", "srv", "5") == "ok"
assert t.record("error", "srv", "timeout") == "error"
assert t.record("bogus", "srv", "") is None
assert t.progress == {"srv": ("error", "timeout")}
def test_start_does_not_overwrite_existing_state(self):
t = MCPProgressTracker()
t.record("success", "srv", "3")
t.record("start", "srv", "")
assert t.progress["srv"] == ("ok", "3")
def test_snapshot_is_independent_copy(self):
t = MCPProgressTracker()
t.record("success", "srv", "1")
snap = t.snapshot()
t.record("error", "srv", "oops")
assert snap == [("ok", "1")]
def test_totals(self):
t = MCPProgressTracker()
t.record("start", "a", "")
t.record("success", "b", "1")
t.record("error", "c", "boom")
done, total = t.totals()
assert (done, total) == (2, 3)
# ──────────────────────────────────────────────────────────────────────
# BackgroundAgentLoader
# ──────────────────────────────────────────────────────────────────────
def _make_loader_fn(agent_value="AGENT", fail_with=None, capture=None):
"""Build a sync loader that records ``on_mcp_progress`` + kwargs."""
def _loader(*, on_mcp_progress=None, **kwargs):
if capture is not None:
capture["on_mcp_progress"] = on_mcp_progress
capture.setdefault("kwargs", []).append(kwargs)
if on_mcp_progress is not None:
on_mcp_progress("start", "srv", "")
on_mcp_progress("success", "srv", "1")
if fail_with is not None:
raise fail_with
return agent_value
return _loader
def _run(coro):
return asyncio.run(coro)
class TestBackgroundAgentLoaderStart:
def test_start_creates_task_and_forwards_kwargs(self):
captured: dict = {}
loader = BackgroundAgentLoader(_make_loader_fn(capture=captured))
async def _go():
loader.start(workspace_dir="/ws", checkpointer="CK")
assert loader.task is not None
assert loader.is_pending
await loader.await_ready()
_run(_go())
assert captured["kwargs"][0] == {"workspace_dir": "/ws", "checkpointer": "CK"}
def test_start_bumps_load_id(self):
loader = BackgroundAgentLoader(_make_loader_fn())
async def _go():
assert loader._load_id == 0
loader.start()
assert loader._load_id == 1
loader.start()
assert loader._load_id == 2
await loader.await_ready()
_run(_go())
def test_start_cancels_in_flight_prior_task(self):
import time
def _blocking(*, on_mcp_progress=None):
time.sleep(0.05)
return "LATE"
async def _go():
loader = BackgroundAgentLoader(_blocking)
loader.start()
first_task = loader.task
# Supersede immediately; asyncio.to_thread wrapper gets cancelled.
loader._loader_fn = _make_loader_fn("FRESH")
loader.start()
agent = await loader.await_ready()
assert agent == "FRESH"
# Let the first thread drain so its done callback (gated) fires.
await asyncio.sleep(0.1)
assert first_task.cancelled() or first_task.done()
_run(_go())
class TestBackgroundAgentLoaderCallbacks:
def test_progress_hook_sees_events_in_order(self):
events: list[tuple[str, str, str]] = []
loader = BackgroundAgentLoader(
_make_loader_fn(capture={}),
on_progress=lambda e, s, d: events.append((e, s, d)),
)
async def _go():
loader.start()
await loader.await_ready()
_run(_go())
assert events == [("start", "srv", ""), ("success", "srv", "1")]
def test_stale_progress_events_are_dropped(self):
"""A progress event fired after a newer `start` must not reach the hook."""
import time
seen: list[str] = []
# Loader 1 sleeps so its progress event fires AFTER load 2 starts.
def slow_loader(*, on_mcp_progress=None):
time.sleep(0.08)
if on_mcp_progress is not None:
on_mcp_progress("success", "from-slow", "1")
return "slow-agent"
def fast_loader(*, on_mcp_progress=None):
if on_mcp_progress is not None:
on_mcp_progress("success", "from-fast", "1")
return "fast-agent"
loader = BackgroundAgentLoader(
slow_loader, on_progress=lambda e, s, d: seen.append(s)
)
async def _go():
loader.start()
# Supersede before the slow thread's event fires.
await asyncio.sleep(0.01)
loader._loader_fn = fast_loader
loader.start()
await loader.await_ready()
# Let the superseded thread finish (its event is gated out).
await asyncio.sleep(0.1)
_run(_go())
assert "from-fast" in seen
assert "from-slow" not in seen
def test_success_callback_fires_on_completion(self):
got = []
loader = BackgroundAgentLoader(
_make_loader_fn("MY_AGENT"),
on_success=lambda a: got.append(a),
)
async def _go():
loader.start()
await loader.await_ready()
await asyncio.sleep(0) # let done-callback run
_run(_go())
assert got == ["MY_AGENT"]
def test_failure_callback_fires_on_error(self):
err = RuntimeError("load failed")
got_failures = []
got_successes = []
loader = BackgroundAgentLoader(
_make_loader_fn(fail_with=err),
on_success=lambda a: got_successes.append(a),
on_failure=lambda e: got_failures.append(e),
)
async def _go():
loader.start()
with pytest.raises(RuntimeError, match="load failed"):
await loader.await_ready()
await asyncio.sleep(0)
_run(_go())
assert got_failures == [err]
assert got_successes == []
class TestBackgroundAgentLoaderAwaitReady:
def test_returns_cached_agent_without_reawaiting(self):
captured: dict = {}
loader = BackgroundAgentLoader(_make_loader_fn("A", capture=captured))
async def _go():
loader.start()
assert await loader.await_ready() == "A"
assert await loader.await_ready() == "A"
_run(_go())
assert len(captured["kwargs"]) == 1
def test_raises_if_started_not_called(self):
loader = BackgroundAgentLoader(_make_loader_fn())
async def _go():
with pytest.raises(RuntimeError, match="before start"):
await loader.await_ready()
_run(_go())
def test_reraises_real_error_on_subsequent_awaits(self):
"""After a failure, ``await_ready`` must keep raising the real exception —
not the "before start()" sentinel — until ``start`` is called again."""
def _fail(*, on_mcp_progress=None):
raise RuntimeError("bad MCP config")
loader = BackgroundAgentLoader(_fail)
async def _go():
loader.start()
with pytest.raises(RuntimeError, match="bad MCP config"):
await loader.await_ready()
with pytest.raises(RuntimeError, match="bad MCP config"):
await loader.await_ready()
_run(_go())
def test_needs_restart_flags_failed_load_for_retry(self):
calls = {"n": 0}
def flaky(*, on_mcp_progress=None):
calls["n"] += 1
if calls["n"] == 1:
raise RuntimeError("first attempt failed")
return "SECOND"
loader = BackgroundAgentLoader(flaky)
async def _go():
assert loader.needs_restart # never started
loader.start()
with pytest.raises(RuntimeError):
await loader.await_ready()
assert loader.needs_restart # failed, caller may retry
loader.start()
assert await loader.await_ready() == "SECOND"
assert not loader.needs_restart # success → no retry
_run(_go())
class TestBackgroundAgentLoaderAdopt:
def test_adopt_seats_external_agent(self):
loader = BackgroundAgentLoader(_make_loader_fn())
loader.adopt("EXTERNAL")
assert loader.agent == "EXTERNAL"
assert not loader.is_pending
def test_adopt_supersedes_in_flight_load(self):
"""A late background completion must not overwrite an adopted agent."""
import time
def _slow(*, on_mcp_progress=None):
time.sleep(0.08)
return "FROM_BACKGROUND"
loader = BackgroundAgentLoader(_slow)
async def _go():
loader.start()
await asyncio.sleep(0.01)
loader.adopt("FROM_MODEL")
# Give the background thread time to finish and fire its
# done-callback; the generation token should make it a no-op.
await asyncio.sleep(0.1)
assert loader.agent == "FROM_MODEL"
_run(_go())
class TestBackgroundAgentLoaderIsPending:
def test_false_before_start(self):
loader = BackgroundAgentLoader(_make_loader_fn())
assert not loader.is_pending
def test_false_after_completion(self):
loader = BackgroundAgentLoader(_make_loader_fn())
async def _go():
loader.start()
await loader.await_ready()
_run(_go())
assert not loader.is_pending
def test_true_between_start_and_completion(self):
import time
def _wait_loader(*, on_mcp_progress=None):
time.sleep(0.05)
return "ok"
loader = BackgroundAgentLoader(_wait_loader)
async def _go():
loader.start()
assert loader.is_pending
await loader.await_ready()
assert not loader.is_pending
_run(_go())