Files
EvoScientist/tests/conftest.py
T
houren Antony 2dc1e227eb fix(langgraph-dev): rotate langgraph_dev.log when it exceeds 50MB (#270)
* fix(langgraph-dev): rotate langgraph_dev.log when it exceeds 50MB

``_LOG_FILE`` (``~/.config/evoscientist/langgraph_dev.log``) was
opened in ``start_langgraph_dev`` with plain ``"ab"`` and never
rotated, so it grew unbounded over weeks/months of heavy use —
especially when chatty MCP servers spawned by langgraph dev
filled it, or when failure paths produced stack traces.

Implement the recommended option 1 from #209: filesize-based
rollover. When the active log exceeds 50MB on the next
``start_langgraph_dev`` invocation, rename it to
``langgraph_dev.log.1`` (overwriting any existing backup) via
``os.replace`` and start fresh. Single-backup policy keeps the
disk footprint bounded at roughly 2x threshold.

Rotation is best-effort: ``_rotate_log_if_needed`` logs and
swallows OSError so a permission error or racing rename can't
block langgraph dev from starting. The next ``start`` invocation
will try again — worst case the log grows for one more session.

Options 2 (timestamped per-session + 7-day sweep) and 3
(``RotatingFileHandler`` + pipe) are explicitly NOT done — option
1 is simplest, no async machinery, matches the issue's
recommendation.

Closes #209

* test(langgraph-dev): redirect _PID_DIR in rotate integration test

Address CodeRabbit review comment on #270: the
``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open``
test patched only ``_LOG_FILE`` to a tmp path, but
``start_langgraph_dev`` also calls ``_PID_DIR.mkdir(...)`` as part
of its prelude, which would create a real directory under
``~/.config/evoscientist/`` on a dev machine. Redirect
``_PID_DIR`` to ``tmp_path / "pids"`` too so the test stays
fully isolated. Add a final assertion that ``pid_dir.is_dir()``
holds, proving the function reached past the mkdir call.

* refactor(langgraph-dev): bundle runtime paths into LanggraphRuntimePaths

@din0s review follow-up on #270: the previous test isolation patched
only ``_LOG_FILE`` (and after a second round, ``_PID_DIR``), but
``start_langgraph_dev`` still touches 5 distinct on-disk paths. Patching
any subset of those still leaves the others pointing at the user's real
``~/.config/evoscientist/`` — exactly the case that produced the
"Port 6174 cannot be bound after waiting 60s" symptom on the
reviewer's machine.

Replace the five free-floating module-level constants
(``_PID_DIR`` / ``_PID_FILE`` / ``_LOG_FILE`` / ``_WORKSPACE_SIDECAR``
/ ``_FILE_LOCK_PATH``) with a single ``LanggraphRuntimePaths`` frozen
dataclass exposed as a module-level ``RUNTIME`` instance. Production
code accesses ``RUNTIME.pid_file`` etc.; tests can now substitute the
*whole* bundle in one assignment:

    monkeypatch.setattr(
        manager, "RUNTIME",
        manager.LanggraphRuntimePaths.for_directory(tmp_path / "runtime"),
    )

The classmethod ``for_directory(pid_dir)`` builds an isolated bundle
rooted at a single dir, so the test author doesn't spell out every
path field. Tests that only care about one field (e.g. pid_file
during the stale-process kill path) use
``dataclasses.replace(manager.RUNTIME, pid_file=X)`` — frozen
dataclass-friendly, no need to enumerate the other four fields.

The dataclass's docstring records the migration rationale (the old
five-name layout invited inconsistent patches).

External callers of the old constants updated:
- ``EvoScientist/deploy/server.py`` and ``webui.py`` now import
  ``RUNTIME`` and use ``RUNTIME.log_file`` for the on-screen log
  path hint. The other imports they had (``_DEFAULT_PORT``,
  ``_is_port_occupied``, ``_read_workspace_sidecar``) are still
  module-level functions/values, untouched.

Test updates:
- ``tests/test_langgraph_manager.py``: ``patch.object(manager, "_XXX",
  X)`` patterns now go through ``dataclasses.replace(manager.RUNTIME,
  xxx=X)``; the ``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open``
  test (from the previous #270 review iteration) uses
  ``for_directory`` for one-shot isolation.
- ``tests/test_langgraph_dev_workspace_sidecar.py``: each test now
  goes through a tiny ``_isolated_runtime(monkeypatch, tmp_path)``
  helper that calls ``for_directory``.
- ``tests/test_langgraph_dev_deploy_mode.py``: same ``for_directory``
  swap.

No production behavior change. All ``langgraph_dev``-side tests
(``test_langgraph_manager.py`` 26/26, ``test_langgraph_dev_workspace_sidecar.py``
14/14, ``test_langgraph_dev_deploy_mode.py`` 14/14, ``test_cli_deploy.py``
18/18 — which indirectly exercises deploy/server.py and deploy/webui.py
imports) pass. Full-project test count unchanged from baseline; the
remaining 22 Windows-only pre-existing failures (test_background
``os.killpg``, test_file_mentions tilde, mcp_client ``shutil.which``,
test_sessions 8.3 short path) are documented as out-of-scope for #207.

* style: apply ruff format to langgraph_dev test + module files

CI lint check on #270 failed:

  Run ruff format --check .
  Would reformat: EvoScientist/langgraph_dev/manager.py
  Would reformat: tests/test_langgraph_manager.py

Plus two test files touched by the prior consolidation commit that
``ruff format`` hadn't seen yet:

  tests/test_langgraph_dev_deploy_mode.py
  tests/test_langgraph_dev_workspace_sidecar.py

Just formatting. No logic change. All 75 refactor-related tests pass.

* fix(test): use for_directory for full path isolation + patch _can_bind_port to skip real socket ops

Two fixes for TestStartLanggraphDevRotatesLog:

1. Replace dataclasses.replace(manager.RUNTIME, ...) with
   LanggraphRuntimePaths.for_directory(pid_dir) so pid_file,
   workspace_sidecar, and lock_file are also temp-rooted
   (prevents leak to ~/.config/evoscientist/).

2. Monkeypatch _can_bind_port to always return True so the
   bind-poll loop in _wait_for_port_bindable passes immediately
   without touching real sockets (fixes 60s timeout on machines
   where port 6174 is already in use).

* fix: cross-platform compatibility for Windows CI runners

- background.py: replace POSIX-only os.killpg/os.getpgid with
  cross-platform _kill_process_tree() helper. On Windows falls back
  to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX
  keeps existing os.killpg logic.

- test_backends.py: replace mkdir -p shell execution in
  test_literal_workspace_path_replaced with preprocessing-boundary
  assertion (patch LocalShellBackend.execute, capture command,
  assert workspace path was rewritten to ./). Avoids POSIX-only
  mkdir -p on Windows runners.

- test_file_mentions.py: monkeypatch USERPROFILE on Windows so
  ntpath.expanduser() resolves ~ to tmp_path even when HOME is
  unset on CI runners.

* refactor(test): add runtime_paths fixture to isolate manager.RUNTIME

Adds a reusable fixture that monkeypatches manager.RUNTIME to a
temp-rooted LanggraphRuntimePaths.for_directory(). Tests that need
specific fields can still dataclasses.replace(runtime_paths, ...)
but the baseline is always temp-isolated, preventing leaks to
~/.config/evoscientist/.

Updated test_langgraph_dev_deploy_mode.py, test_langgraph_dev_workspace_sidecar.py,
and test_langgraph_manager.py to use the fixture, consolidating sequential
lock_file + pid_dir patches into single dataclasses.replace calls.

* Revert "fix: cross-platform compatibility for Windows CI runners"

This reverts commit eb025d24af32e195a982cd40f6d70dba885c4019.

* style: ruff format conftest.py

* fix: address review issues in log-rotation + runtime paths

- Use for_directory(tmp_path/pids) as base in ensure_langgraph_dev tests
  so pid_file/log_file are co-located with pid_dir, not split across paths
- Remove unused runtime_paths param from test_no_existing_file_is_noop
- Replace manager.RUNTIME with runtime_paths in two sidecar tests
- Use for_directory(DEFAULT_PID_DIR) instead of explicit construction
- Fix stale _LOG_FILE reference in TestRotateLogIfNeeded docstring

* style: ruff format test files

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-09 15:55:44 +01:00

175 lines
5.5 KiB
Python

"""Shared fixtures for EvoScientist tests."""
import asyncio
import pytest
def run_async(coro):
"""Run an async coroutine safely, cancelling pending tasks before closing.
This prevents 'Event loop is closed' errors from asyncio.Queue cleanup
when tasks are still waiting on Queue.get() at teardown time.
"""
loop = asyncio.new_event_loop()
try:
return loop.run_until_complete(coro)
finally:
# Cancel all pending tasks so Queue getters don't raise on close
pending = asyncio.all_tasks(loop)
for task in pending:
task.cancel()
if pending:
loop.run_until_complete(asyncio.gather(*pending, return_exceptions=True))
loop.run_until_complete(loop.shutdown_asyncgens())
loop.close()
@pytest.fixture(name="run_async")
def run_async_fixture():
"""Pytest fixture that exposes run_async as a callable for test functions."""
return run_async
@pytest.fixture
def sample_tool_call():
"""A minimal tool call dict."""
return {"id": "tc_001", "name": "execute", "args": {"command": "ls -la"}}
@pytest.fixture
def sample_tool_result():
"""A minimal tool result dict."""
return {
"id": "tc_001",
"name": "execute",
"content": "[OK] file1.py file2.py",
"success": True,
}
@pytest.fixture
def sample_events():
"""A sequence of stream event dicts covering common types."""
return [
{"type": "thinking", "content": "Let me think..."},
{"type": "text", "content": "Here is the answer."},
{
"type": "tool_call",
"id": "tc_001",
"name": "execute",
"args": {"command": "ls"},
},
{
"type": "tool_result",
"id": "tc_001",
"name": "execute",
"content": "[OK] done",
"success": True,
},
{
"type": "subagent_start",
"name": "research-agent",
"description": "Find papers",
"instance_id": "task:research",
"tool_call_id": "tc_task_001",
},
{
"type": "subagent_tool_call",
"subagent": "research-agent",
"instance_id": "task:research",
"name": "tavily_search",
"args": {"query": "test"},
"id": "tc_sa_001",
},
{
"type": "subagent_tool_result",
"subagent": "research-agent",
"instance_id": "task:research",
"name": "tavily_search",
"content": "Results...",
"success": True,
"id": "tc_sa_001",
},
{
"type": "subagent_end",
"name": "research-agent",
"instance_id": "task:research",
},
{"type": "done", "response": "Here is the answer."},
]
@pytest.fixture
def tmp_workspace(tmp_path):
"""Provide a temporary workspace directory path."""
ws = tmp_path / "workspace"
ws.mkdir()
return str(ws)
@pytest.fixture
def runtime_paths(tmp_path, monkeypatch):
"""Isolate ``langgraph_dev.manager.RUNTIME`` under a temp directory.
Replaces the module-level ``RUNTIME`` with a fully temp-rooted bundle
so every path (``pid_dir``, ``pid_file``, ``log_file``,
``workspace_sidecar``, ``lock_file``) is contained under ``tmp_path``.
Tests that need a variant of a single field can still call
``dataclasses.replace(runtime_paths, log_file=…)`` etc. — the
baseline is already isolated, so forgetting a field just keeps it
under ``tmp_path``, never ``~/.config/evoscientist``.
"""
from EvoScientist.langgraph_dev import manager
runtime = manager.LanggraphRuntimePaths.for_directory(tmp_path / "runtime")
monkeypatch.setattr(manager, "RUNTIME", runtime)
return runtime
# Capture deepagents tool factories at conftest load time — BEFORE any test
# imports EvoScientist, which can trigger ``_patch_deepagents_model_passthrough``
# during agent construction. Once captured here, the ``restore_model_passthrough_patch``
# fixture has a stable "truly unpatched" baseline to reset to between tests, even
# if upstream code paths apply the patch as a side effect.
try:
from deepagents.middleware import async_subagents as _ds_async_subagents
_DEEPAGENTS_ORIGINAL_BUILD_START = _ds_async_subagents._build_start_tool
_DEEPAGENTS_ORIGINAL_BUILD_UPDATE = _ds_async_subagents._build_update_tool
except Exception:
_ds_async_subagents = None
_DEEPAGENTS_ORIGINAL_BUILD_START = None
_DEEPAGENTS_ORIGINAL_BUILD_UPDATE = None
@pytest.fixture
def restore_model_passthrough_patch():
"""Reset deepagents internals + ``_model_passthrough_patched`` to unpatched.
The model-passthrough patch wraps ``deepagents.middleware.async_subagents``
module-level functions in place. The originals are captured at conftest
load time (above) so this fixture can always start each test from a
known-unpatched state regardless of what other tests / agent fixtures
did to the module before.
"""
from EvoScientist.llm import patches as patches_mod
if _ds_async_subagents is None:
# deepagents not importable — fixture is a no-op (the patch fn itself
# returns early in that case).
yield
return
def _reset() -> None:
_ds_async_subagents._build_start_tool = _DEEPAGENTS_ORIGINAL_BUILD_START
_ds_async_subagents._build_update_tool = _DEEPAGENTS_ORIGINAL_BUILD_UPDATE
patches_mod._model_passthrough_patched = False
_reset()
try:
yield
finally:
_reset()