Files
EvoScientist-Multi/tests/test_experts_command.py
T
jfilipiuk bd2464423a feat: agent-teams part D - async expert dispatch mechanism (#391)
* fix(deps): pin openrouter below 0.11 to avoid SSE stream regressions (#373)

* fix(openrouter): address SSE stream leak by closing response iterator

* refactor(openrouter): pass through SDK args in SSE leak patch

* test(openrouter): make SSE leak tests version-agnostic across SDK generations

* fix(openrouter): remove SSE stream leak patch and update dependencies

* fix: repair interrupted tool call history (#366)

* fix: repair interrupted tool call history

Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths.

* fix: repair malformed tool calls and dedupe repair warnings

Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted
threads with syntactically invalid tool calls get synthesized error
results and are accepted by strict providers.

Preserve the originating tool call's name in the synthesized ToolMessage,
and deduplicate repair warnings per unique tool-call id via a warned set
owned by the middleware instance, since the middleware rewrites the
request but not thread state.

Document the middleware's scope versus deepagents' PatchToolCallsMiddleware
(orphan ToolMessage dropping and mid-run coverage).

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Update README.md

* Update README.md

* Update README.zh-CN.md

* fix: scrub host path from skill_manager output and guard batch install (#377)

* fix: scrub host path from skill_manager output and guard batch install

* test: tighten install leak guards to catch host path in either tier

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: log missing async-subagent tools at DEBUG, not WARNING (#378)

* fix: log missing async-subagent tools at DEBUG, not WARNING

* fix: distinguish load_subagents callers via async_swap_pending flag

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: default reasoning context for codex proxy Responses API (#380)

* fix: set langgraph and codex proxy runtime defaults

* fix: address runtime default review feedback

* fix: drop langgraph dev env defaults per maintainer review

langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment,
so the reported KeyError cannot come from this flow; the env defaults added
here were unnecessary. Scope the PR back to the codex proxy reasoning
context fix only.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* release: v0.2.4 (#389)

* fix: add support for new Anthropic models and enhance adaptive thinking tests

* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models

* fix: update version to v0.2.4 in badges, README, and project files

* fix: update Star History chart links in README and README.zh-CN

* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog

* fix: update wechat group image in assets

* refactor(runtime): centralize async bridges under an owned runtime (#376)

* feat(runtime): add application-scoped async runtime

* refactor(cli): use owned runtime for session stats

* refactor(onboard): use the owned async runtime

* docs(runtime): record async bridge ownership

* refactor(middleware): keep sync fallback synchronous

* refactor(mcp): load tools on an owned runtime

* refactor(cli): share owned runtime across entry points

* refactor(channels): make inbound sync bridge explicit

* refactor(stream): run Rich streaming on owned runtime

* chore(runtime): remove nest-asyncio dependency

* refactor(asyncio): require active loops in async code

* docs(runtime): document final event loop ownership

* fix(stream): cancel stalled owned streams

* fix(cli): recover cleanly from stream cancellation

* fix(runtime): drain executor work before shutdown

* fix(runtime): terminate cancelled shell process trees

* fix(models): let fallback bypass selector failures

* fix(cli): reset interrupt handling between turns

* docs: rm implementation spec

* fix(serve): cancel active turns during shutdown

* fix(runtime): protect settlement from waiter cancellation

* fix(backends): reject empty shell commands

* fix(runtime): terminate descendants after shell exit

* fix(mcp): keep standalone discovery off channel loop

* fix(cli): own and settle interactive prompt cancellation

* fix(serve): keep channel sends off runtime loop

* fix(stream): scope cancel context to iterator steps

* refactor(serve): require the owned async runtime

* fix(channels): keep interactive sends off runtime loop

* fix(selector): surface fallback without log spam

* test(runtime): normalize Windows shell marker

* fix(cli): serialize interactive session turns

* fix(shell): bound output drain after termination

* fix(ui): do not retry owned runtime failures

* fix(shell): allow signal-safe registry reentry

* fix(shell): avoid terminating reused process ids

* fix(channels): preserve streaming send order

* fix(cli): report runtime shutdown timeouts cleanly

* fix(mcp): guide async callers to async loader

* docs(runtime): clarify reserved async bridge APIs

* fix(runtime): bound code interpreter cleanup

* test(shell): use active Python for drain regression

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* feat: add payload-aware EvoAsyncSubAgentMiddleware

* feat: register expert_container_async graph for async expert dispatch

* feat: fold installed expert skills into async subagent registry

* fix: accept 'async' as valid default_dispatch value

* feat: dispatch-aware ActiveTeamMiddleware cue (task vs start_async_task)

* fix: drop future annotations in expert_async_subagent so ToolRuntime injects

* feat: surface output_path and skill_name to expert container as runtime cue

* fix: extend AsyncWatcher client cache with expert specs for completion nudge

* feat: teach main agent the async-expert return envelope shape

* feat: propagate cfg.model to expert-async runs.create via ClientCacheProxy

* chore: guard AsyncWatcher client-cache extension against upstream rename

* test: cover output_path runtime-context tail block and wrong-type guard

* docs: drop out-of-repo notes/ ref from expert_container_async module doc

* fix: warn on unrecognized default_dispatch frontmatter value

* fix: reject empty-body expert skills on async dispatch to match sync policy

* docs: explain why expert container includes general-purpose subagent

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385)

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL

* fix: keep parent env authoritative over workspace .env for mapped keys

* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins

* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS

* fix: merge .env via dotenv_values to close empty-value and RMW-race edges

* chore: align docstrings after .env-merge rework

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Add Requesty as an LLM provider (#346)

* Add Requesty as an LLM provider

* Address review: Requesty prompt caching, model ordering, key validation

- Declare Anthropic-style prompt caching for Requesty Claude models by
  default (mirroring the OpenRouter behavior), with an opt-out flag
  EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
  provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
  overrides native/OpenRouter models for names it shares with them
  (the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
  invalid/missing key (public catalog), so it cannot validate a key. Use a
  minimal authenticated /v1/chat/completions request instead (200 = valid,
  403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).

* Validate Requesty key against auth layer, not a specific model

The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix(llm): filter unnamed tool calls (#390)

* fix(llm): filter unnamed tool calls

* test(llm): cover tool call sanitization branches

* fix(llm): repair unnamed tool calls in middleware

---------

Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix(middleware): mount tool-history repair on sync subagents and harden raw tool-call vetting (#393)

* feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation

* fix(tests): add test for dropping non-list raw tool calls in repair_tool_history

* Add Atlas Cloud LLM provider (#388)

* Add Atlas Cloud LLM provider

* Add Atlas Cloud onboarding support

* fix(validators): update atlascloud key validation to handle insufficient balance case

---------

Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix: prepend EvoAsyncSubAgentMiddleware for prefix cache stability

* docs: clarify list_dispatchable_experts covers both dispatch shapes

* fix: guard async expert fold-in against reserved-name collisions

* fix: honest advertising surfaces for async expert dispatch

* fix: compose expert persona into base-stack system_message

* fix: drop payload from start_async_task, inject skill_name by construction

* feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395

- Introduced TodoListMiddleware to the middleware stack for better task management.
- Updated HITL interrupt configuration to include 'delete' operations requiring approval.
- Implemented error handling for delete operations in read-only and memory backends.
- Enhanced approval prompt formatting to display file paths for delete actions.
- Added tests to ensure delete operations are correctly blocked or prompted for approval.
- Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality.

* revert: drop skill_manager from sync expert-container tool_registry

* revert: drop skill_manager from async expert-container tools

* fix: drop removed ASYNC_TASK_SYSTEM_PROMPT import for deepagents 0.7.0

* fix: mock list_dispatchable_experts in single-cue test for CI

* chore: drop stale output_path from async container graph docstring

* fix: forward configurable_extra through owned-runtime and HITL re-invocations

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: Sanjay Santhanam <51058514+Sanjays2402@users.noreply.github.com>
Co-authored-by: Yougang Lyu <82445958+youganglyu@users.noreply.github.com>
Co-authored-by: houren Antony <2212222@mail.nankai.edu.cn>
Co-authored-by: dinos <dinospk1999@gmail.com>
Co-authored-by: Thibault Jaigu <84420566+Thibaultjaigu@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
Co-authored-by: nightcityblade <jackchen@haloailabs.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nb213 <binyangzhu000@gmail.com>
Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
2026-08-07 17:07:01 +01:00

261 lines
10 KiB
Python

"""Unit tests for /experts and /expert slash commands."""
from __future__ import annotations
from dataclasses import dataclass, field
from typing import Any
from unittest.mock import patch
import pytest
from EvoScientist.commands.base import ChannelRuntime, CommandContext
from EvoScientist.commands.implementation.experts import (
ExpertCommand,
ExpertsCommand,
invalidate_experts_cache,
)
@pytest.fixture(autouse=True)
def _bust_experts_cache_between_tests():
"""The dispatchable-experts cache in ``experts.py`` is module-level; without
resetting it, a test that patches ``list_expert_skills`` sees the previous
test's fakes.
"""
invalidate_experts_cache()
yield
invalidate_experts_cache()
class _FakeUI:
"""Minimal CommandUI capturing outputs for assertion."""
supports_interactive = False
def __init__(self) -> None:
self.lines: list[tuple[str, str]] = []
self.mounted: list[Any] = []
def append_system(self, text: str, style: str = "dim") -> None:
self.lines.append((text, style))
def mount_renderable(self, renderable: Any) -> None:
self.mounted.append(renderable)
@dataclass
class _FakeSkillInfo:
"""Enough of ``SkillInfo`` for the commands to render."""
name: str
description: str = ""
role: str = ""
default_dispatch: str = ""
type: str = "expert"
tags: list[str] = field(default_factory=list)
source: str = "builtin"
# Non-empty by default so the fake passes the empty-body filter in
# ``list_dispatchable_experts``. Tests that specifically want to
# exercise the empty-body reject path pass ``body=""``.
body: str = "persona"
def _make_ctx(active_teams: list[str] | None = None) -> tuple[CommandContext, _FakeUI]:
ui = _FakeUI()
runtime = ChannelRuntime()
if active_teams:
runtime.active_teams = list(active_teams)
ctx = CommandContext(
agent=None,
thread_id="t1",
ui=ui,
channel_runtime=runtime,
)
return ctx, ui
class TestExpertsList:
async def test_lists_installed_experts_in_table(self):
ctx, ui = _make_ctx()
with patch(
"EvoScientist.tools.skills_manager.list_expert_skills",
return_value=[
_FakeSkillInfo(
name="idea-brainstorm",
role="Research idea brainstormer",
default_dispatch="sync",
),
],
):
await ExpertsCommand().execute(ctx, args=[])
# A Rich Table was mounted, and the no-experts-invited hint appeared.
assert len(ui.mounted) == 1
assert any("No experts invited" in text for text, _ in ui.lines)
async def test_empty_list_prints_help_hint(self):
ctx, ui = _make_ctx()
with patch(
"EvoScientist.tools.skills_manager.list_expert_skills",
return_value=[],
):
await ExpertsCommand().execute(ctx, args=[])
assert any("No expert skills installed" in text for text, _ in ui.lines)
assert not ui.mounted
async def test_active_expert_marked_in_table(self):
ctx, ui = _make_ctx(active_teams=["idea-brainstorm"])
with patch(
"EvoScientist.tools.skills_manager.list_expert_skills",
return_value=[
_FakeSkillInfo(
name="idea-brainstorm",
role="Research idea brainstormer",
default_dispatch="sync",
),
],
):
await ExpertsCommand().execute(ctx, args=[])
assert any("Active: idea-brainstorm" in text for text, _ in ui.lines)
class TestExpertToggle:
async def test_missing_arg_prints_usage(self):
ctx, ui = _make_ctx()
await ExpertCommand().execute(ctx, args=[])
assert any("Usage:" in text for text, _ in ui.lines)
async def test_unknown_expert_errors(self):
ctx, ui = _make_ctx()
with patch(
"EvoScientist.tools.skills_manager.list_expert_skills",
return_value=[_FakeSkillInfo(name="idea-brainstorm")],
):
await ExpertCommand().execute(ctx, args=["not-an-expert"])
assert any(
"No expert skill named 'not-an-expert'" in text for text, _ in ui.lines
)
assert ctx.channel_runtime.active_teams == []
async def test_async_expert_refused_with_reason_when_async_unavailable(self):
"""When an installed expert declares ``default_dispatch: async`` but
async dispatch is unavailable, ``/expert`` must refuse with the specific
reason — not the empty-body / name-collision default — so the user
knows to enable ``enable_async_subagents`` or start langgraph dev.
Reviewer thread on PR #391."""
ctx, ui = _make_ctx()
async_expert = _FakeSkillInfo(
name="literature-review", default_dispatch="async"
)
with (
patch(
"EvoScientist.tools.skills_manager.list_expert_skills",
return_value=[async_expert],
),
patch(
"EvoScientist.subagents.expert_container.is_async_dispatch_available",
return_value=False,
),
):
await ExpertCommand().execute(ctx, args=["literature-review"])
assert any("async dispatch is unavailable" in text for text, _ in ui.lines), (
f"expected honest async-unavailable message, got: {ui.lines}"
)
assert ctx.channel_runtime.active_teams == []
async def test_invite_adds_to_active_teams(self):
ctx, ui = _make_ctx()
with patch(
"EvoScientist.tools.skills_manager.list_expert_skills",
return_value=[_FakeSkillInfo(name="idea-brainstorm")],
):
await ExpertCommand().execute(ctx, args=["idea-brainstorm"])
assert ctx.channel_runtime.active_teams == ["idea-brainstorm"]
assert any("Invited expert: idea-brainstorm" in text for text, _ in ui.lines)
async def test_toggle_dismisses_when_already_invited(self):
ctx, ui = _make_ctx(active_teams=["idea-brainstorm"])
with patch(
"EvoScientist.tools.skills_manager.list_expert_skills",
return_value=[_FakeSkillInfo(name="idea-brainstorm")],
):
await ExpertCommand().execute(ctx, args=["idea-brainstorm"])
assert ctx.channel_runtime.active_teams == []
assert any("Dismissed expert: idea-brainstorm" in text for text, _ in ui.lines)
async def test_clear_dismisses_all(self):
ctx, ui = _make_ctx(active_teams=["idea-brainstorm", "second"])
await ExpertCommand().execute(ctx, args=["clear"])
assert ctx.channel_runtime.active_teams == []
assert any(
"Dismissed experts: idea-brainstorm, second" in text for text, _ in ui.lines
)
async def test_clear_on_empty_list_reports_nothing_to_do(self):
ctx, ui = _make_ctx()
await ExpertCommand().execute(ctx, args=["clear"])
assert ctx.channel_runtime.active_teams == []
assert any("No experts invited" in text for text, _ in ui.lines)
async def test_no_channel_runtime_prints_warning(self):
ui = _FakeUI()
ctx = CommandContext(agent=None, thread_id="t1", ui=ui, channel_runtime=None)
await ExpertCommand().execute(ctx, args=["idea-brainstorm"])
assert any("/expert requires a session runtime" in text for text, _ in ui.lines)
class TestExpertCompletions:
"""``ExpertCommand.get_completions`` mixes dynamic expert names with the
static ``clear`` subcommand. Regression coverage for the three fixes on
PR #371: exact-match suppression, past-first-arg guard, and
case-insensitive matching.
"""
def _patched_experts(self, *names: str):
# Patch ``list_dispatchable_experts`` directly (not the underlying
# ``list_expert_skills``) so the test does not depend on the shipped
# yaml sub-agent set — the reserved-name filter would otherwise
# silently reject a fake whose name collides with a future yaml
# sub-agent.
return patch(
"EvoScientist.subagents.expert_container.list_dispatchable_experts",
return_value=[_FakeSkillInfo(name=n) for n in names],
)
def test_lists_installed_experts_and_clear(self):
cmd = ExpertCommand()
with self._patched_experts("smoke-test-sync-expert", "smoke-test-alt-expert"):
completions = cmd.get_completions([""])
names = {name for name, _ in completions}
assert names == {"smoke-test-sync-expert", "smoke-test-alt-expert", "clear"}
def test_case_insensitive_prefix_match(self):
# Skill dir names sometimes have uppercase; completion typed
# lowercase must still surface them.
cmd = ExpertCommand()
with self._patched_experts("Smoke-Test-Case-Expert", "smoke-test-sync-expert"):
completions = cmd.get_completions(["smoke-test-c"])
names = {name for name, _ in completions}
assert names == {"Smoke-Test-Case-Expert"}
def test_exact_match_hides_popup_same_case(self):
cmd = ExpertCommand()
with self._patched_experts("smoke-test-sync-expert"):
completions = cmd.get_completions(["smoke-test-sync-expert"])
assert completions == []
def test_exact_match_hides_popup_different_case(self):
# Case-insensitive exact-match suppression: typing the name in a
# different case than the skill dir still fully completes it and
# hides the popup.
cmd = ExpertCommand()
with self._patched_experts("Smoke-Test-Case-Expert"):
completions = cmd.get_completions(["smoke-test-case-expert"])
assert completions == []
def test_past_first_arg_returns_empty(self):
# /expert takes a single positional. Trailing space -> tokens == ["n", ""].
cmd = ExpertCommand()
with self._patched_experts("smoke-test-sync-expert"):
assert cmd.get_completions(["smoke-test-sync-expert", ""]) == []
assert cmd.get_completions(["smoke-test-sync-expert", "foo"]) == []