Files
EvoScientist/tests/test_serve_agent_holder.py
T
dinos 92d95dee68 feat(memory): add observation memory lifecycle (#259)
* feat(memory): add observation memory lifecycle

Add file-backed observation memory with deterministic markdown records,
structured record_observation tooling, startup indexing, and
profile/observation prompt guidance.

Launch post-turn and post-subagent EvoMemory workers through LangGraph
dev so completed runs can update profile memory, save durable
observations, and write subagent execution summaries without blocking
the active agent.

Wire memory middleware into the main agent, subagents, async graphs, TUI
status reporting, worker activity accounting, and observation-aware
research prompts, with regression coverage for storage, lifecycle
scheduling, graph registration, status display, and stream reset
behavior.

* fix(cli): sync background agent server on resume

Resume flows now need to keep the LangGraph dev background server
aligned with the active workspace even when async subagents are
disabled. EvoMemory workers use that server too, so gating resume-time
sync on enable_async_subagents could leave workers pinned to the launch
workspace after resuming a thread from another workspace.

Run workspace sync unconditionally for Rich CLI and Textual resume
paths, while preserving WorkspaceMismatchError handling so failed sync
aborts the resume before mutating the active thread or workspace.

Propagate aborted resume callbacks through the command UI so
channel-issued /resume commands do not send false success or history
output. Channel slash dispatch now treats CommandManager-caught command
errors as command errors and skips completion hooks for those failed
commands.

Add regression coverage for disabled async subagents, callback aborts,
and channel command error reporting.

* fix(cli): prepare serve resume workspace before adopting

Load the resumed workspace agent and sync the background server as a
single pre-adoption step. Restore the previous active workspace if
preparation fails so serve mode keeps using the old session
consistently.

* fix(memory): untrack abandoned worker status watches

Stop treating watcher shutdown as confirmed worker completion. Terminal
worker statuses still count memory deltas, while poll failures or
watcher setup failures now remove the active run without crediting
partial outputs.

* fix(cli): report channel command failures accurately

Treat command_error as a None sentinel so empty error strings still
fail, and let TUI resumes continue only on non-mismatch
background-server sync failures while reporting degraded mode.

* fix(stream): clear memory counters for resume streams

Reset completed-memory counters for every new agent stream, including
Command-based HITL and resume streams, so saved-memory indicators do not
leak across turns.

* docs(tools): make observation recording guidance conditional

Clarify that agents should call record_observation only when the
observation tool is available, preserving the existing durability and
usefulness criteria.

* feat(config): add controls for profile and observation memory

Add config flags for profile memory, observation memory, observation
writer placement, and background memory workers.

Wire the controls through main agents, subagents, EvoMemory middleware,
and memory lifecycle workers so observation writes can be assigned to
the live agent, subagent worker, both, or neither. Keep turn memory
workers profile-only and make prompts reflect the available observation
read/write paths. Skip langgraph dev startup when neither async
subagents nor memory workers need the background server.

Add coverage for config parsing, prompt gating, middleware wiring, and
worker tool availability.

* test(cli): include memory defaults in serve config stubs

* fix(memory): offload async worker launch blocking calls

Run the langgraph-dev health check and memory-output snapshot in worker
threads from the async EvoMemory launcher so it does not block the event
loop.

* chore(memory): harden turn worker subagent guardrail

* chore(memory): refresh profile context per request

* fix(memory): offload async profile file reads

* fix(memory): offload async worker completion accounting
2026-06-05 15:11:20 +01:00

532 lines
17 KiB
Python

"""Tests for the serve-mode ``on_cmd_completed`` hook factory.
Regression coverage for the follow-up to issue #181 — `/model` invoked
over a channel in ``EvoSci serve`` must swap the running agent for
subsequent messages, not silently keep the stale one the while-loop
captured at startup.
"""
from unittest.mock import AsyncMock, MagicMock, patch
import pytest
from EvoScientist.cli.channel import (
ChannelMessage,
_register_channel_request,
)
from EvoScientist.cli.commands import (
_make_serve_cmd_completed_hook,
_make_serve_handle_session_resume_cb,
_make_serve_start_new_session_cb,
_serve_process_message,
)
from EvoScientist.commands.base import ChannelRuntime
from tests.conftest import run_async as _run
def test_hook_updates_holder_on_agent_swap():
"""``/model`` mutates ``ctx.agent`` to a new handle — the hook must
push that handle into the shared holder so the outer poll loop sees
it on the next message."""
holder = {"agent": "original-agent"}
hook = _make_serve_cmd_completed_hook(holder)
ctx = MagicMock()
ctx.agent = "new-agent"
cmd = MagicMock()
cmd.name = "/model"
_run(hook(ctx, "original-agent", cmd))
assert holder["agent"] == "new-agent"
def test_hook_syncs_channel_runtime():
"""Other readers (the bus) look at ``ChannelRuntime.agent``; the
hook keeps the runtime in sync with the holder update."""
holder = {"agent": "original-agent", "thread_id": "t"}
runtime = ChannelRuntime(agent="original-agent", thread_id="t")
hook = _make_serve_cmd_completed_hook(holder, runtime)
ctx = MagicMock()
ctx.agent = "new-agent"
# Pin ctx.thread_id explicitly — a bare MagicMock would let the
# hook's getattr fall through to a fresh MagicMock attribute and
# silently mutate runtime.thread_id, hiding regressions.
ctx.thread_id = "t"
cmd = MagicMock()
cmd.name = "/model"
_run(hook(ctx, "original-agent", cmd))
assert runtime.agent == "new-agent"
assert runtime.thread_id == "t"
def test_hook_noop_when_agent_unchanged():
"""Commands like ``/evoskills`` don't touch ``ctx.agent`` — the
holder must stay put."""
holder = {"agent": "original-agent"}
hook = _make_serve_cmd_completed_hook(holder)
ctx = MagicMock()
ctx.agent = "original-agent" # no swap
cmd = MagicMock()
cmd.name = "/evoskills"
_run(hook(ctx, "original-agent", cmd))
assert holder["agent"] == "original-agent"
def test_hook_noop_when_ctx_agent_is_none():
"""Guard against commands that reset ``ctx.agent`` to ``None`` —
we never want to write ``None`` into the holder."""
holder = {"agent": "original-agent"}
hook = _make_serve_cmd_completed_hook(holder)
ctx = MagicMock()
ctx.agent = None
cmd = MagicMock()
cmd.name = "/whatever"
_run(hook(ctx, "original-agent", cmd))
assert holder["agent"] == "original-agent"
def test_hook_updates_thread_id_on_resume():
"""``/resume`` mutates ``ctx.thread_id`` — the hook must push the
new id into the holder so the outer poll loop runs subsequent
messages on the resumed thread."""
holder = {"agent": "a", "thread_id": "original-tid"}
hook = _make_serve_cmd_completed_hook(holder)
ctx = MagicMock()
ctx.agent = "a" # no agent swap
ctx.thread_id = "new-tid"
ctx.workspace_dir = None
cmd = MagicMock()
cmd.name = "/resume"
_run(hook(ctx, "a", cmd))
assert holder["thread_id"] == "new-tid"
def test_hook_updates_workspace_dir_on_resume():
"""`/resume` can restore a different workspace; serve must reload for it."""
cfg = object()
holder = {
"agent": "old-agent",
"thread_id": "original-tid",
"workspace_dir": "/old-ws",
"config": cfg,
}
hook = _make_serve_cmd_completed_hook(holder, config=cfg)
ctx = MagicMock()
ctx.agent = "old-agent"
ctx.thread_id = "new-tid"
ctx.workspace_dir = "/restored-ws"
cmd = MagicMock()
cmd.name = "/resume"
with (
patch(
"EvoScientist.cli.commands._sync_background_agent_server_workspace",
new=AsyncMock(),
) as sync_server,
patch(
"EvoScientist.cli.commands._load_agent",
return_value="reloaded-agent",
) as load_agent,
):
_run(hook(ctx, "old-agent", cmd))
sync_server.assert_awaited_once_with(cfg, workspace_dir="/restored-ws")
load_agent.assert_called_once_with(workspace_dir="/restored-ws", config=cfg)
assert holder["workspace_dir"] == "/restored-ws"
assert holder["agent"] == "reloaded-agent"
def test_hook_syncs_channel_runtime_thread_id():
"""The bus reads ``ChannelRuntime.thread_id``; hook must sync it
alongside the holder update."""
holder = {"agent": "a", "thread_id": "original-tid"}
runtime = ChannelRuntime(agent="a", thread_id="original-tid")
hook = _make_serve_cmd_completed_hook(holder, runtime)
ctx = MagicMock()
ctx.agent = "a"
ctx.thread_id = "new-tid"
ctx.workspace_dir = None
cmd = MagicMock()
cmd.name = "/resume"
_run(hook(ctx, "a", cmd))
assert runtime.thread_id == "new-tid"
def test_hook_noop_when_thread_id_unchanged():
"""Most commands don't touch thread_id — holder stays put."""
holder = {"agent": "a", "thread_id": "same-tid"}
hook = _make_serve_cmd_completed_hook(holder)
ctx = MagicMock()
ctx.agent = "a"
ctx.thread_id = "same-tid"
cmd = MagicMock()
cmd.name = "/evoskills"
_run(hook(ctx, "a", cmd))
assert holder["thread_id"] == "same-tid"
def test_hook_skips_resume_warning_when_thread_unchanged():
"""Bare ``/resume`` with no argument prints usage but leaves
``ctx.thread_id`` unchanged — the in-memory-state warning must NOT
fire because no resume actually happened."""
holder = {"agent": "a", "thread_id": "original-tid"}
hook = _make_serve_cmd_completed_hook(holder)
ctx = MagicMock()
ctx.agent = "a"
ctx.thread_id = "original-tid" # unchanged — bare /resume case
ctx.workspace_dir = None
cmd = MagicMock()
cmd.name = "/resume"
_run(hook(ctx, "a", cmd))
ctx.ui.append_system.assert_not_called()
ctx.ui.flush.assert_not_called()
def test_hook_emits_resume_warning_when_thread_changed():
"""``/resume <tid>`` that actually changes thread_id must surface
the in-memory-state warning via ``ctx.ui``."""
holder = {"agent": "a", "thread_id": "original-tid"}
hook = _make_serve_cmd_completed_hook(holder)
ctx = MagicMock()
# Mock out async flush so the test can synchronously run the hook.
ctx.ui.flush = AsyncMock()
ctx.agent = "a"
ctx.thread_id = "abc12345-resumed-tid"
ctx.workspace_dir = None
cmd = MagicMock()
cmd.name = "/resume"
_run(hook(ctx, "a", cmd))
ctx.ui.append_system.assert_called_once()
warn_text, warn_kwargs = (
ctx.ui.append_system.call_args.args,
ctx.ui.append_system.call_args.kwargs,
)
assert "in-memory state" in warn_text[0]
assert "abc12345" in warn_text[0]
assert warn_kwargs.get("style") == "yellow"
ctx.ui.flush.assert_awaited_once()
def test_start_new_session_cb_rotates_thread_id():
"""``/new`` via channel calls this callback — must generate a new
thread id, push into holder, and sync the channel runtime."""
holder = {"agent": "a", "thread_id": "old-tid"}
runtime = ChannelRuntime(agent="a", thread_id="old-tid")
with patch(
"EvoScientist.sessions.generate_thread_id",
return_value="freshly-generated-tid",
):
cb = _make_serve_start_new_session_cb(holder, runtime)
cb()
assert holder["thread_id"] == "freshly-generated-tid"
assert runtime.thread_id == "freshly-generated-tid"
def test_start_new_session_cb_leaves_agent_alone():
"""``/new`` rotates thread only — agent handle must stay put
(serve's agent is a single pre-loaded instance, not per-thread)."""
holder = {"agent": "a", "thread_id": "old-tid"}
with patch(
"EvoScientist.sessions.generate_thread_id",
return_value="new-tid",
):
cb = _make_serve_start_new_session_cb(holder)
cb()
assert holder["agent"] == "a"
def test_serve_resume_callback_syncs_reloads_and_adopts_workspace():
cfg = object()
holder = {
"agent": "old-agent",
"thread_id": "old-tid",
"workspace_dir": "/old-ws",
"config": cfg,
}
runtime = ChannelRuntime(agent="old-agent", thread_id="old-tid")
cb = _make_serve_handle_session_resume_cb(holder, runtime, config=cfg)
call_order: list[str] = []
def _load_agent(**_kwargs):
call_order.append("load")
return "reloaded-agent"
async def _sync_server(*_args, **_kwargs):
call_order.append("sync")
with (
patch(
"EvoScientist.cli.commands._sync_background_agent_server_workspace",
new=AsyncMock(side_effect=_sync_server),
) as sync_server,
patch(
"EvoScientist.cli.commands._load_agent",
side_effect=_load_agent,
) as load_agent,
):
_run(cb("new-tid", "/new-ws"))
sync_server.assert_awaited_once_with(cfg, workspace_dir="/new-ws")
load_agent.assert_called_once_with(workspace_dir="/new-ws", config=cfg)
assert call_order == ["load", "sync"]
assert holder["thread_id"] == "new-tid"
assert holder["workspace_dir"] == "/new-ws"
assert holder["agent"] == "reloaded-agent"
assert runtime.thread_id == "new-tid"
assert runtime.agent == "reloaded-agent"
def test_hook_emits_resume_warning_after_resume_callback_adopts_thread():
cfg = object()
holder = {
"agent": "old-agent",
"thread_id": "old-tid",
"workspace_dir": "/old-ws",
"config": cfg,
}
runtime = ChannelRuntime(agent="old-agent", thread_id="old-tid")
cb = _make_serve_handle_session_resume_cb(holder, runtime, config=cfg)
with (
patch(
"EvoScientist.cli.commands._sync_background_agent_server_workspace",
new=AsyncMock(),
),
patch(
"EvoScientist.cli.commands._load_agent",
return_value="reloaded-agent",
),
):
_run(cb("abc12345-resumed-tid", "/new-ws"))
hook = _make_serve_cmd_completed_hook(holder, runtime, config=cfg)
ctx = MagicMock()
ctx.ui.flush = AsyncMock()
ctx.agent = "reloaded-agent"
ctx.thread_id = "abc12345-resumed-tid"
ctx.workspace_dir = "/new-ws"
cmd = MagicMock()
cmd.name = "/resume"
_run(hook(ctx, "reloaded-agent", cmd))
ctx.ui.append_system.assert_called_once()
assert "in-memory state" in ctx.ui.append_system.call_args.args[0]
ctx.ui.flush.assert_awaited_once()
def test_serve_resume_callback_preserves_state_when_sync_fails():
cfg = object()
holder = {
"agent": "old-agent",
"thread_id": "old-tid",
"workspace_dir": "/old-ws",
"config": cfg,
}
runtime = ChannelRuntime(agent="old-agent", thread_id="old-tid")
cb = _make_serve_handle_session_resume_cb(holder, runtime, config=cfg)
with (
patch(
"EvoScientist.cli.commands._sync_background_agent_server_workspace",
new=AsyncMock(side_effect=RuntimeError("workspace conflict")),
),
patch(
"EvoScientist.cli.commands._load_agent",
return_value="loaded-but-not-adopted",
) as load_agent,
patch("EvoScientist.cli.commands.set_active_workspace") as set_active,
pytest.raises(RuntimeError, match="workspace conflict"),
):
_run(cb("new-tid", "/new-ws"))
load_agent.assert_called_once_with(workspace_dir="/new-ws", config=cfg)
set_active.assert_called_once_with("/old-ws")
assert "loaded-but-not-adopted" not in holder.values()
assert "_resume_warning_thread_id" not in holder
assert holder == {
"agent": "old-agent",
"thread_id": "old-tid",
"workspace_dir": "/old-ws",
"config": cfg,
}
assert runtime.agent == "old-agent"
assert runtime.thread_id == "old-tid"
def test_serve_resume_callback_load_failure_does_not_sync_or_adopt():
cfg = object()
holder = {
"agent": "old-agent",
"thread_id": "old-tid",
"workspace_dir": "/old-ws",
"config": cfg,
}
runtime = ChannelRuntime(agent="old-agent", thread_id="old-tid")
cb = _make_serve_handle_session_resume_cb(holder, runtime, config=cfg)
with (
patch(
"EvoScientist.cli.commands._load_agent",
side_effect=RuntimeError("load failed"),
) as load_agent,
patch("EvoScientist.cli.commands.set_active_workspace") as set_active,
patch(
"EvoScientist.cli.commands._sync_background_agent_server_workspace",
new=AsyncMock(),
) as sync_server,
pytest.raises(RuntimeError, match="load failed"),
):
_run(cb("new-tid", "/new-ws"))
load_agent.assert_called_once_with(workspace_dir="/new-ws", config=cfg)
set_active.assert_called_once_with("/old-ws")
sync_server.assert_not_awaited()
assert "_resume_warning_thread_id" not in holder
assert holder == {
"agent": "old-agent",
"thread_id": "old-tid",
"workspace_dir": "/old-ws",
"config": cfg,
}
assert runtime.agent == "old-agent"
assert runtime.thread_id == "old-tid"
def test_hook_handles_both_agent_and_thread_swap():
"""Edge case: a command that changes both (hypothetical). Both
updates must land in the holder."""
holder = {"agent": "old-agent", "thread_id": "old-tid"}
hook = _make_serve_cmd_completed_hook(holder)
ctx = MagicMock()
ctx.agent = "new-agent"
ctx.thread_id = "new-tid"
cmd = MagicMock()
_run(hook(ctx, "old-agent", cmd))
assert holder["agent"] == "new-agent"
assert holder["thread_id"] == "new-tid"
def test_serve_process_message_reports_slash_dispatch_error_without_fallback():
"""Defensive: if ``dispatch_channel_slash_command`` ever leaks an
exception past its own wrapper, ``_serve_process_message`` must set
one error response and not fall through to ``run_streaming``.
"""
msg = ChannelMessage(
msg_id="msg-1",
content="/evoskills core",
sender="channel-user",
channel_type="imessage",
metadata={},
channel_ref=None,
bus_ref=None,
chat_id="channel-user",
message_id="ts-1",
)
holder = {"agent": "agent", "thread_id": "tid"}
with (
patch(
"EvoScientist.cli.commands.dispatch_channel_slash_command",
new=AsyncMock(side_effect=RuntimeError("slash broke")),
),
patch("EvoScientist.cli.commands._set_channel_response") as mock_set_resp,
patch("EvoScientist.cli.tui_runtime.run_streaming") as mock_run_streaming,
):
_register_channel_request(msg)
_serve_process_message(
msg,
agent_holder=holder,
model="model",
workspace_dir="/tmp",
show_thinking=False,
)
mock_set_resp.assert_called_once_with("msg-1", "Command error: slash broke")
mock_run_streaming.assert_not_called()
def test_serve_process_message_uses_runtime_workspace_from_holder():
"""After `/resume`, serve should use the adopted workspace, not startup ws."""
msg = ChannelMessage(
msg_id="msg-2",
content="hello",
sender="channel-user",
channel_type="imessage",
metadata={},
channel_ref=None,
bus_ref=None,
chat_id="channel-user",
message_id="ts-2",
)
holder = {
"agent": "agent",
"thread_id": "tid",
"workspace_dir": "/restored-workspace",
}
captured: dict[str, str] = {}
async def _fake_dispatch(*args, **kwargs):
captured["slash_workspace"] = kwargs["workspace_dir"]
return False
def _fake_build_metadata(workspace_dir: str, _model: str | None):
captured["meta_workspace"] = workspace_dir
return {}
with (
patch(
"EvoScientist.cli.commands.dispatch_channel_slash_command",
new=AsyncMock(side_effect=_fake_dispatch),
),
patch(
"EvoScientist.cli.commands.build_metadata",
side_effect=_fake_build_metadata,
),
patch("EvoScientist.cli.tui_runtime.run_streaming", return_value="ok"),
):
_register_channel_request(msg)
_serve_process_message(
msg,
agent_holder=holder,
model="model",
workspace_dir="/startup-workspace",
show_thinking=False,
)
assert captured["slash_workspace"] == "/restored-workspace"
assert captured["meta_workspace"] == "/restored-workspace"