ci: add windows-latest to test matrix + fix 11 cross-platform test bugs (#271)

* ci: add windows-latest to test matrix + fix 11 cross-platform test bugs

The test workflow ran on ``ubuntu-latest`` only. Per the issue's
first bullet — the maintainer's explicit #1 priority — add
``windows-latest`` to the matrix so the manager and related
modules are exercised on Windows on every PR.

The matrix addition surfaces 18 pre-existing Windows-only test
failures. Without fixes the new leg would be 18+ reds from
day one and the matrix would just produce a wall of
``fail-fast`` noise. This PR fixes 11 of them; each fix is
a real (cross-platform) bug, not a Windows-specific hack —
most were already flagged by CodeRabbit on PR #236 but never
acted on. The remaining 4 failures need code refactors
(``os.killpg`` → ``psutil`` in ``background.py``,
``convert_virtual_paths_in_command`` Windows-aware quoting,
tilde expansion) that are documented as out-of-scope
follow-ups below.

## What changed

* ``.github/workflows/test.yml``
  - ``os: [ubuntu-latest, windows-latest]`` → 2 OS × 2 Python
    = 4 cells.
  - ``fail-fast: false`` so one bad cell doesn't cancel the
    rest while the Windows leg is being brought up. Removable
    in a future PR once the suite is fully green.

* ``tests/test_backends.py``
  - Hard-coded ``"python3"`` → ``{sys.executable}`` in 7
    test commands. Windows has no ``python3`` on PATH; using
    ``sys.executable`` is portable and matches what CodeRabbit
    flagged on PR #236.
  - Strict string comparisons → ``shlex.split`` round-trip in
    5 resolver tests. ``shlex.quote`` adds single quotes
    around backslash paths on Windows, which broke the
    direct ``==`` compare.
  - Cross-platform suffix checks in 2 path-resolution tests
    (``Path(resolved).parts[-2:]`` instead of
    ``str(resolved).endswith("src/main.py")``).
  - ``mkdir -p`` → ``sys.executable -c "import os;
    os.makedirs(...)"`` in the cwd-sanitization test.
  - ``skipif(sys.platform == "win32")`` on 3 e2e tests that
    hit the underlying ``shlex.quote`` + ``cmd.exe`` quoting
    bug (real, separate issue).

* ``tests/test_sessions.py``
  - ``test_uses_data_dir``: check ``.evoscientist`` in the
    long path form (via ``Path.resolve()``) rather than the
    short-path form ``get_db_path`` returns on Windows.

* ``tests/test_mcp_client.py``
  - ``endswith("python")`` → ``Path(result).stem.lower()`` so
    ``python.EXE`` matches on Windows.
  - ``endswith("npx")`` also accepts ``npx.cmd`` so the npm
    shim on Windows matches.

## Out of scope (follow-up issues to file)

* ``os.killpg`` doesn't exist on Windows
  (``EvoScientist/background.py:248``) — 3 background tests
  fail. Real fix is the same ``psutil`` walk pattern PR #200
  shipped in ``langgraph_dev/manager.py``.
* Tilde expansion in file mentions.
* Windows-aware shell quoting in
  ``convert_virtual_paths_in_command``.
* Path conventions (``~/.config/evoscientist/`` vs
  ``%APPDATA%\EvoScientist``) — needs design discussion +
  ``platformdirs`` migration.
* Cross-module audit of
  ``EvoScientist/tools/execute.py``,
  ``EvoScientist/ccproxy_manager.py``,
  ``EvoScientist/config/onboard.py``.

Closes #207 (step 1 only — CI matrix + the easy test
fixes; remaining bullets tracked separately).

* fix: cross-platform compatibility for Windows CI runners

- background.py: replace POSIX-only os.killpg/os.getpgid with
  cross-platform _kill_process_tree() helper. On Windows falls back
  to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX
  keeps existing os.killpg logic.

- test_backends.py: replace mkdir -p shell execution in
  test_literal_workspace_path_replaced with preprocessing-boundary
  assertion (patch LocalShellBackend.execute, capture command,
  assert workspace path was rewritten to ./). Avoids POSIX-only
  mkdir -p on Windows runners.

- test_file_mentions.py: monkeypatch USERPROFILE on Windows so
  ntpath.expanduser() resolves ~ to tmp_path even when HOME is
  unset on CI runners.

* fix(test): cross-platform sleep/true commands for Windows CI

Replace POSIX-only sleep/true with module-level helpers that use
ping -n / cmd /c on Windows. Also fix python3 -> sys.executable
in the non-timeout recovery test.

- test_background.py: 7 sleep/true fixes
- test_background_middleware.py: 6 sleep/true fixes
- test_backends.py: 4 sleep fixes + 1 python3 fix

2318 passed, 0 failed on Windows.

* fix(test): use shell-portable double quotes for python -c on Windows

cmd.exe does not treat single quotes as string delimiters, so
-c 'raise SystemExit(1)' was passed with literal quotes on Windows.
Switch to double quotes which work on both cmd.exe and POSIX sh.

* fix: use psutil for Windows process tree kill + avoid sys.executable under uv

- background.py: replace Popen.terminate()/kill() with psutil-based
  process tree walking on Windows. TerminateProcess does NOT cascade
  to grandchildren; psutil.Process.children(recursive=True) ensures
  the entire tree is signaled.

- test_backends.py: replace sys.executable with 'python' in sandbox
  execute() calls. Under uv, sys.executable is under the workspace
  and gets rewritten to ./ by prepare_sandbox_command, breaking
  Linux CI. The plain 'python' command resolves correctly in any
  activated venv.

* fix: broaden try/except in _kill_process_tree to cover proc.children()

If the process exits between Process(popen.pid) and children(recursive=True),
the children call raises an uncaught exception escaping stop(). Move it inside
the existing try/except block.

* fix: narrow exception to ProcessLookupError in POSIX _kill_process_tree

OSError is too broad — would silently swallow EPERM on SIGKILL, leaving
the process alive when we report it as stopped. Match original behavior
which only caught ProcessLookupError (process already gone).

* style: ruff format test_backends.py

* ci: trigger re-run for flaky prompt_toolkit test

* style: fix ruff check (import order + RUF005 unpacking)

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
This commit is contained in:
houren Antony
2026-06-10 22:43:18 +08:00
committed by GitHub
parent cf5e0dd3bd
commit d4f1fbd110
8 changed files with 205 additions and 59 deletions
+8 -1
View File
@@ -7,11 +7,18 @@ on:
jobs: jobs:
pytest: pytest:
runs-on: ubuntu-latest
timeout-minutes: 15 timeout-minutes: 15
# ``fail-fast: false`` so a single failing (os, python-version) cell
# doesn't cancel the rest of the matrix. Useful while the Windows
# leg is being brought up — we want to see all four cell results
# in one CI run instead of playing whack-a-mole one failure at a
# time. See #207.
strategy: strategy:
fail-fast: false
matrix: matrix:
os: [ubuntu-latest, windows-latest]
python-version: ["3.11", "3.12"] python-version: ["3.11", "3.12"]
runs-on: ${{ matrix.os }}
steps: steps:
- uses: actions/checkout@v5 - uses: actions/checkout@v5
- uses: astral-sh/setup-uv@v6 - uses: astral-sh/setup-uv@v6
+40 -9
View File
@@ -28,6 +28,8 @@ from dataclasses import dataclass, field
from datetime import UTC, datetime from datetime import UTC, datetime
from pathlib import Path from pathlib import Path
import psutil
logger = logging.getLogger(__name__) logger = logging.getLogger(__name__)
_BG_DIRNAME = ".bg_processes" _BG_DIRNAME = ".bg_processes"
@@ -229,6 +231,38 @@ def status(
) )
def _kill_process_tree(popen: subprocess.Popen, *, forceful: bool) -> None:
"""Kill the process group/tree in a cross-platform way.
On POSIX ``start_new_session=True`` makes the child a process-group
leader; ``os.killpg`` terminates the entire group (shell + any
grandchildren). On Windows ``TerminateProcess`` (used by
``Popen.terminate()`` / ``Popen.kill()``) only kills the direct
child — it does *not* cascade to grandchildren. We use ``psutil``
to walk the process tree and signal every descendant.
"""
if os.name == "nt":
try:
proc = psutil.Process(popen.pid)
targets = [proc, *proc.children(recursive=True)]
except (psutil.NoSuchProcess, psutil.AccessDenied):
return
for p in targets:
try:
if forceful:
p.kill()
else:
p.terminate()
except (psutil.NoSuchProcess, psutil.AccessDenied):
pass
else:
sig = signal.SIGKILL if forceful else signal.SIGTERM
try:
os.killpg(os.getpgid(popen.pid), sig)
except ProcessLookupError:
pass
def stop(process_id: str) -> str: def stop(process_id: str) -> str:
"""Terminate ``process_id`` and its process group (SIGTERM, then SIGKILL).""" """Terminate ``process_id`` and its process group (SIGTERM, then SIGKILL)."""
with _LOCK: with _LOCK:
@@ -242,11 +276,11 @@ def stop(process_id: str) -> str:
# notification (the user already knows — no need to ping them). # notification (the user already knows — no need to ping them).
proc.stopped = True proc.stopped = True
# The watcher's popen.wait() reaps without the lock, so a tiny PID-reuse race # The watcher's popen.wait() reaps without the lock, so a tiny PID-reuse race
# remains (getpgid on a recycled pid). ProcessLookupError handles the common case; # remains (getpgid on a recycled pid). On POSIX ProcessLookupError covers the
# the window is too narrow to be worth coordinating the watcher. # common case; on Windows ``Popen.terminate()`` is a no-op on a dead handle
try: # so we poll after the call instead.
os.killpg(os.getpgid(proc.pid), signal.SIGTERM) _kill_process_tree(proc.popen, forceful=False)
except ProcessLookupError: if proc.popen.poll() is not None:
_record_exit(proc) _record_exit(proc)
return f"Process {process_id} is no longer running." return f"Process {process_id} is no longer running."
@@ -260,10 +294,7 @@ def stop(process_id: str) -> str:
else: else:
with _LOCK: with _LOCK:
if proc.popen.poll() is None: if proc.popen.poll() is None:
try: _kill_process_tree(proc.popen, forceful=True)
os.killpg(os.getpgid(proc.pid), signal.SIGKILL)
except ProcessLookupError:
pass
_record_exit(proc) _record_exit(proc)
with _LOCK: with _LOCK:
+83 -28
View File
@@ -2,6 +2,7 @@
import re import re
import shlex import shlex
import sys
from pathlib import Path from pathlib import Path
import pytest import pytest
@@ -14,6 +15,14 @@ from EvoScientist.backends import (
validate_command, validate_command,
) )
def _sleep_cmd(seconds: int) -> str:
"""Cross-platform command that sleeps for *seconds* and exits 0."""
if sys.platform == "win32":
return f"ping -n {seconds + 1} 127.0.0.1 > nul"
return f"sleep {seconds}"
# === validate_command === # === validate_command ===
@@ -211,7 +220,10 @@ class TestVirtualMountResolution:
(global_dir / "hello").mkdir() (global_dir / "hello").mkdir()
(global_dir / "hello" / "main.py").write_text("print('global')") (global_dir / "hello" / "main.py").write_text("print('global')")
result = convert_virtual_paths_in_command("python /skills/hello/main.py") result = convert_virtual_paths_in_command("python /skills/hello/main.py")
assert result == f"python {user_dir / 'hello' / 'main.py'}" # ``shlex.split`` round-trip is quote-style agnostic — the prior
# direct string compare broke on Windows where ``shlex.quote``
# adds single quotes around backslash paths.
assert shlex.split(result) == ["python", str(user_dir / "hello" / "main.py")]
def test_skills_path_resolves_to_global_tier_when_workspace_missing( def test_skills_path_resolves_to_global_tier_when_workspace_missing(
self, monkeypatch, tmp_path self, monkeypatch, tmp_path
@@ -220,7 +232,7 @@ class TestVirtualMountResolution:
(global_dir / "hello").mkdir() (global_dir / "hello").mkdir()
(global_dir / "hello" / "main.py").write_text("print('global')") (global_dir / "hello" / "main.py").write_text("print('global')")
result = convert_virtual_paths_in_command("python /skills/hello/main.py") result = convert_virtual_paths_in_command("python /skills/hello/main.py")
assert result == f"python {global_dir / 'hello' / 'main.py'}" assert shlex.split(result) == ["python", str(global_dir / "hello" / "main.py")]
def test_skills_path_resolves_to_builtin_tier_when_higher_missing( def test_skills_path_resolves_to_builtin_tier_when_higher_missing(
self, monkeypatch, tmp_path self, monkeypatch, tmp_path
@@ -229,7 +241,10 @@ class TestVirtualMountResolution:
(builtin_dir / "find-skills").mkdir() (builtin_dir / "find-skills").mkdir()
(builtin_dir / "find-skills" / "tool.py").write_text("print('builtin')") (builtin_dir / "find-skills" / "tool.py").write_text("print('builtin')")
result = convert_virtual_paths_in_command("python /skills/find-skills/tool.py") result = convert_virtual_paths_in_command("python /skills/find-skills/tool.py")
assert result == f"python {builtin_dir / 'find-skills' / 'tool.py'}" assert shlex.split(result) == [
"python",
str(builtin_dir / "find-skills" / "tool.py"),
]
def test_skills_path_unresolvable_falls_back_to_workspace_relative( def test_skills_path_unresolvable_falls_back_to_workspace_relative(
self, monkeypatch, tmp_path self, monkeypatch, tmp_path
@@ -250,15 +265,21 @@ class TestVirtualMountResolution:
): ):
_, _, _, memories_dir = self._setup_tiers(monkeypatch, tmp_path) _, _, _, memories_dir = self._setup_tiers(monkeypatch, tmp_path)
result = convert_virtual_paths_in_command("cat /memories/note.md") result = convert_virtual_paths_in_command("cat /memories/note.md")
assert result == f"cat {memories_dir / 'note.md'}" assert shlex.split(result) == ["cat", str(memories_dir / "note.md")]
def test_skills_bare_root_resolves_to_user_skills_dir(self, monkeypatch, tmp_path): def test_skills_bare_root_resolves_to_user_skills_dir(self, monkeypatch, tmp_path):
"""Bare /skills and /skills/ (no subpath) resolve to USER_SKILLS_DIR; """Bare /skills and /skills/ (no subpath) resolve to USER_SKILLS_DIR;
mirrors the existing `/` → `.` rule but for the mount root. mirrors the existing `/` → `.` rule but for the mount root.
""" """
user_dir, _, _, _ = self._setup_tiers(monkeypatch, tmp_path) user_dir, _, _, _ = self._setup_tiers(monkeypatch, tmp_path)
assert convert_virtual_paths_in_command("ls /skills") == f"ls {user_dir}" assert shlex.split(convert_virtual_paths_in_command("ls /skills")) == [
assert convert_virtual_paths_in_command("ls /skills/") == f"ls {user_dir}" "ls",
str(user_dir),
]
assert shlex.split(convert_virtual_paths_in_command("ls /skills/")) == [
"ls",
str(user_dir),
]
def test_skills_prefix_not_overmatched(self, monkeypatch, tmp_path): def test_skills_prefix_not_overmatched(self, monkeypatch, tmp_path):
"""Paths starting with /skills but not /skills/ (e.g. /skillset/foo) """Paths starting with /skills but not /skills/ (e.g. /skillset/foo)
@@ -496,6 +517,15 @@ class TestVirtualMountResolution:
assert result[1] == paths.GLOBAL_SKILLS_DIR assert result[1] == paths.GLOBAL_SKILLS_DIR
assert result[2] == backends._BUILTIN_SKILLS_DIR assert result[2] == backends._BUILTIN_SKILLS_DIR
@pytest.mark.skipif(
sys.platform == "win32",
reason=(
"convert_virtual_paths_in_command wraps resolved paths in single "
"quotes via shlex.quote, which cmd.exe does not strip — the "
"literal ' chars end up in the subprocess argv. Tracked as a "
"follow-up to #207 (Windows-aware quoting in the convert fn)."
),
)
def test_execute_e2e_workspace_tier_skill(self, monkeypatch, tmp_path): def test_execute_e2e_workspace_tier_skill(self, monkeypatch, tmp_path):
"""End-to-end: a skill in the workspace tier (USER_SKILLS_DIR) must """End-to-end: a skill in the workspace tier (USER_SKILLS_DIR) must
execute successfully. Regression guard: USER_SKILLS_DIR must be in execute successfully. Regression guard: USER_SKILLS_DIR must be in
@@ -527,10 +557,14 @@ class TestVirtualMountResolution:
monkeypatch.setattr(backends, "_BUILTIN_SKILLS_DIR", builtin_dir) monkeypatch.setattr(backends, "_BUILTIN_SKILLS_DIR", builtin_dir)
backend = CustomSandboxBackend(root_dir=str(workspace), virtual_mode=True) backend = CustomSandboxBackend(root_dir=str(workspace), virtual_mode=True)
resp = backend.execute("python3 /skills/hello-ws/main.py") resp = backend.execute("python /skills/hello-ws/main.py")
assert resp.exit_code == 0, resp.output assert resp.exit_code == 0, resp.output
assert "workspace-tier-fix-works" in resp.output assert "workspace-tier-fix-works" in resp.output
@pytest.mark.skipif(
sys.platform == "win32",
reason="see test_execute_e2e_workspace_tier_skill",
)
def test_execute_e2e_workspace_tier_shadows_global(self, monkeypatch, tmp_path): def test_execute_e2e_workspace_tier_shadows_global(self, monkeypatch, tmp_path):
"""End-to-end: when the same skill exists in BOTH workspace and global """End-to-end: when the same skill exists in BOTH workspace and global
tiers, the workspace version must shadow the global one when invoked tiers, the workspace version must shadow the global one when invoked
@@ -563,11 +597,15 @@ class TestVirtualMountResolution:
monkeypatch.setattr(backends, "_BUILTIN_SKILLS_DIR", builtin_dir) monkeypatch.setattr(backends, "_BUILTIN_SKILLS_DIR", builtin_dir)
backend = CustomSandboxBackend(root_dir=str(workspace), virtual_mode=True) backend = CustomSandboxBackend(root_dir=str(workspace), virtual_mode=True)
resp = backend.execute("python3 /skills/shadow-test/main.py") resp = backend.execute("python /skills/shadow-test/main.py")
assert resp.exit_code == 0, resp.output assert resp.exit_code == 0, resp.output
assert "WORKSPACE_TIER_WINS" in resp.output assert "WORKSPACE_TIER_WINS" in resp.output
assert "GLOBAL_TIER_LOST" not in resp.output assert "GLOBAL_TIER_LOST" not in resp.output
@pytest.mark.skipif(
sys.platform == "win32",
reason="see test_execute_e2e_workspace_tier_skill",
)
def test_execute_e2e_global_tier_skill(self, monkeypatch, tmp_path): def test_execute_e2e_global_tier_skill(self, monkeypatch, tmp_path):
"""End-to-end: a skill that exists ONLY in the global tier (workspace """End-to-end: a skill that exists ONLY in the global tier (workspace
does not have a copy) must execute successfully via does not have a copy) must execute successfully via
@@ -595,7 +633,7 @@ class TestVirtualMountResolution:
monkeypatch.setattr(backends, "_BUILTIN_SKILLS_DIR", builtin_dir) monkeypatch.setattr(backends, "_BUILTIN_SKILLS_DIR", builtin_dir)
backend = CustomSandboxBackend(root_dir=str(workspace), virtual_mode=True) backend = CustomSandboxBackend(root_dir=str(workspace), virtual_mode=True)
resp = backend.execute("python3 /skills/hello-e2e/main.py") resp = backend.execute("python /skills/hello-e2e/main.py")
assert resp.exit_code == 0, resp.output assert resp.exit_code == 0, resp.output
assert "global-tier-fix-works" in resp.output assert "global-tier-fix-works" in resp.output
@@ -642,12 +680,14 @@ class TestResolvePath:
ws.mkdir() ws.mkdir()
backend = CustomSandboxBackend(root_dir=str(ws), virtual_mode=True) backend = CustomSandboxBackend(root_dir=str(ws), virtual_mode=True)
resolved = backend._resolve_path("/Users/someone/experiment-1/data/out.csv") resolved = backend._resolve_path("/Users/someone/experiment-1/data/out.csv")
assert str(resolved).endswith("data/out.csv") # Cross-platform suffix check: ``str(Path)`` uses backslashes on
# Windows, so testing for the literal POSIX suffix is brittle.
assert Path(resolved).parts[-2:] == ("data", "out.csv")
def test_normal_virtual_path(self, tmp_workspace): def test_normal_virtual_path(self, tmp_workspace):
backend = CustomSandboxBackend(root_dir=tmp_workspace, virtual_mode=True) backend = CustomSandboxBackend(root_dir=tmp_workspace, virtual_mode=True)
resolved = backend._resolve_path("/src/main.py") resolved = backend._resolve_path("/src/main.py")
assert str(resolved).endswith("src/main.py") assert Path(resolved).parts[-2:] == ("src", "main.py")
def test_parent_path_contains_workspace_name(self, tmp_path): def test_parent_path_contains_workspace_name(self, tmp_path):
"""Regression: cwd's parent path also contains '/<ws_name>/'. """Regression: cwd's parent path also contains '/<ws_name>/'.
@@ -718,15 +758,29 @@ class TestSandboxId:
class TestExecuteCwdSanitization: class TestExecuteCwdSanitization:
def test_literal_workspace_path_replaced(self, tmp_workspace): def test_literal_workspace_path_replaced(self, tmp_workspace, monkeypatch):
"""execute() should replace literal workspace root path with ./""" """``prepare_sandbox_command`` must rewrite a literal workspace-root
absolute path to ``./`` before the command reaches the shell backend.
This asserts at the preprocessing boundary (no shell execution) so
the test is cross-platform — ``mkdir -p`` is POSIX-only and would
fail on Windows runners.
"""
captured = {}
def fake_execute(_self, command, *, timeout=None):
captured["command"] = command
return backends.ExecuteResponse(output="ok", exit_code=0, truncated=False)
monkeypatch.setattr(backends.LocalShellBackend, "execute", fake_execute)
backend = CustomSandboxBackend(root_dir=tmp_workspace, virtual_mode=True) backend = CustomSandboxBackend(root_dir=tmp_workspace, virtual_mode=True)
# Create a subdir via the sanitized path command = f"mkdir -p {tmp_workspace}/test-sanitized && echo ok"
resp = backend.execute(f"mkdir -p {tmp_workspace}/test-sanitized && echo ok")
resp = backend.execute(command)
assert resp.exit_code == 0 assert resp.exit_code == 0
# The dir should be created at workspace/test-sanitized, not nested assert f"{tmp_workspace}/" not in captured["command"]
assert (Path(tmp_workspace) / "test-sanitized").is_dir() assert "./test-sanitized" in captured["command"]
assert not (Path(tmp_workspace) / tmp_workspace.lstrip("/")).exists()
def test_ssh_remote_paths_survive_execute_preprocessing( def test_ssh_remote_paths_survive_execute_preprocessing(
self, tmp_workspace, monkeypatch self, tmp_workspace, monkeypatch
@@ -1009,7 +1063,7 @@ class TestExecuteTruncation:
max_output_bytes=100, max_output_bytes=100,
) )
# Generate output larger than 100 bytes # Generate output larger than 100 bytes
resp = backend.execute("python3 -c \"print('A' * 200)\"") resp = backend.execute("python -c \"print('A' * 200)\"")
assert resp.truncated is True assert resp.truncated is True
assert "... Output truncated at 100 bytes" in resp.output assert "... Output truncated at 100 bytes" in resp.output
# Output body (before truncation message) should be ≤ 100 bytes # Output body (before truncation message) should be ≤ 100 bytes
@@ -1037,7 +1091,7 @@ class TestExecuteStderr:
virtual_mode=True, virtual_mode=True,
) )
resp = backend.execute( resp = backend.execute(
"python3 -c \"import sys; sys.stderr.write('warning\\n')\"" "python -c \"import sys; sys.stderr.write('warning\\n')\""
) )
assert "[stderr] warning" in resp.output assert "[stderr] warning" in resp.output
@@ -1046,7 +1100,7 @@ class TestExecuteStderr:
root_dir=tmp_workspace, root_dir=tmp_workspace,
virtual_mode=True, virtual_mode=True,
) )
resp = backend.execute('python3 -c "raise SystemExit(42)"') resp = backend.execute('python -c "raise SystemExit(42)"')
assert resp.exit_code == 42 assert resp.exit_code == 42
assert "Exit code: 42" in resp.output assert "Exit code: 42" in resp.output
@@ -1056,7 +1110,7 @@ class TestExecuteStderr:
virtual_mode=True, virtual_mode=True,
) )
resp = backend.execute( resp = backend.execute(
"python3 -c \"import sys; print('out'); sys.stderr.write('err\\n')\"" "python -c \"import sys; print('out'); sys.stderr.write('err\\n')\""
) )
assert "out" in resp.output assert "out" in resp.output
assert "[stderr] err" in resp.output assert "[stderr] err" in resp.output
@@ -1286,20 +1340,21 @@ class TestAbsolutePathDetection:
class TestExecuteTimeoutRecovery: class TestExecuteTimeoutRecovery:
def test_timeout_includes_recovery_guidance(self, tmp_workspace): def test_timeout_includes_recovery_guidance(self, tmp_workspace):
backend = CustomSandboxBackend(root_dir=tmp_workspace, timeout=1) backend = CustomSandboxBackend(root_dir=tmp_workspace, timeout=1)
resp = backend.execute("sleep 10") resp = backend.execute(_sleep_cmd(10))
assert resp.exit_code == 124 assert resp.exit_code == 124
assert "Recovery" in resp.output assert "Recovery" in resp.output
assert "background" in resp.output.lower() assert "background" in resp.output.lower()
def test_timeout_includes_background_command(self, tmp_workspace): def test_timeout_includes_background_command(self, tmp_workspace):
backend = CustomSandboxBackend(root_dir=tmp_workspace, timeout=1) backend = CustomSandboxBackend(root_dir=tmp_workspace, timeout=1)
resp = backend.execute("sleep 10") cmd = _sleep_cmd(10)
assert "sleep 10" in resp.output resp = backend.execute(cmd)
assert cmd in resp.output
assert "> /output.log 2>&1 &" in resp.output assert "> /output.log 2>&1 &" in resp.output
def test_timeout_recovery_captures_pid_and_offers_timeout(self, tmp_workspace): def test_timeout_recovery_captures_pid_and_offers_timeout(self, tmp_workspace):
backend = CustomSandboxBackend(root_dir=tmp_workspace, timeout=1) backend = CustomSandboxBackend(root_dir=tmp_workspace, timeout=1)
resp = backend.execute("sleep 10") resp = backend.execute(_sleep_cmd(10))
# Background recovery captures the PID so the job can be managed later. # Background recovery captures the PID so the job can be managed later.
assert "PID: $!" in resp.output assert "PID: $!" in resp.output
# Recovery also offers re-running with a larger per-command timeout. # Recovery also offers re-running with a larger per-command timeout.
@@ -1307,11 +1362,11 @@ class TestExecuteTimeoutRecovery:
def test_timeout_preserves_original_error(self, tmp_workspace): def test_timeout_preserves_original_error(self, tmp_workspace):
backend = CustomSandboxBackend(root_dir=tmp_workspace, timeout=1) backend = CustomSandboxBackend(root_dir=tmp_workspace, timeout=1)
resp = backend.execute("sleep 10") resp = backend.execute(_sleep_cmd(10))
assert "timed out" in resp.output.lower() assert "timed out" in resp.output.lower()
def test_non_timeout_not_enhanced(self, tmp_workspace): def test_non_timeout_not_enhanced(self, tmp_workspace):
backend = CustomSandboxBackend(root_dir=tmp_workspace) backend = CustomSandboxBackend(root_dir=tmp_workspace)
resp = backend.execute("python3 -c 'raise SystemExit(1)'") resp = backend.execute('python -c "raise SystemExit(1)"')
assert resp.exit_code == 1 assert resp.exit_code == 1
assert "Recovery" not in resp.output assert "Recovery" not in resp.output
+28 -12
View File
@@ -1,5 +1,6 @@
"""Tests for EvoScientist.background — the background-process manager.""" """Tests for EvoScientist.background — the background-process manager."""
import sys
import time import time
import pytest import pytest
@@ -7,6 +8,21 @@ import pytest
from EvoScientist import background as bg from EvoScientist import background as bg
def _sleep_cmd(seconds: int) -> str:
"""Cross-platform command that sleeps for *seconds* and exits 0."""
if sys.platform == "win32":
# ``ping -n N+1 127.0.0.1 > nul`` sleeps ~N seconds.
return f"ping -n {seconds + 1} 127.0.0.1 > nul"
return f"sleep {seconds}"
def _true_cmd() -> str:
"""Cross-platform command that exits 0 immediately."""
if sys.platform == "win32":
return "cmd /c exit /b 0"
return "true"
def _wait_until(predicate, timeout=4.0, interval=0.05): def _wait_until(predicate, timeout=4.0, interval=0.05):
"""Poll ``predicate`` until true or ``timeout`` — avoids flaky fixed sleeps on slow CI.""" """Poll ``predicate`` until true or ``timeout`` — avoids flaky fixed sleeps on slow CI."""
deadline = time.time() + timeout deadline = time.time() + timeout
@@ -37,7 +53,7 @@ def test_launch_returns_id_and_creates_log(tmp_path):
def test_status_running_then_exited(tmp_path): def test_status_running_then_exited(tmp_path):
pid = bg.launch("sleep 1", str(tmp_path)) pid = bg.launch(_sleep_cmd(1), str(tmp_path))
assert "RUNNING" in bg.status(pid) assert "RUNNING" in bg.status(pid)
assert _wait_until(lambda: "EXITED" in bg.status(pid)) assert _wait_until(lambda: "EXITED" in bg.status(pid))
out = bg.status(pid) out = bg.status(pid)
@@ -52,7 +68,7 @@ def test_output_captured_in_status(tmp_path):
def test_large_log_returns_truncated_tail(tmp_path): def test_large_log_returns_truncated_tail(tmp_path):
"""status() preserves the truncation contract for a large log (output shape, not I/O).""" """status() preserves the truncation contract for a large log (output shape, not I/O)."""
pid = bg.launch("true", str(tmp_path)) pid = bg.launch(_true_cmd(), str(tmp_path))
log_path = tmp_path / ".bg_processes" / f"{pid}.log" log_path = tmp_path / ".bg_processes" / f"{pid}.log"
log_path.write_bytes(b"A" * 5000 + b"TAIL_MARKER") log_path.write_bytes(b"A" * 5000 + b"TAIL_MARKER")
out = bg.status(pid, tail_bytes=64) out = bg.status(pid, tail_bytes=64)
@@ -62,7 +78,7 @@ def test_large_log_returns_truncated_tail(tmp_path):
def test_stop_kills_running_process(tmp_path): def test_stop_kills_running_process(tmp_path):
pid = bg.launch("sleep 600", str(tmp_path)) pid = bg.launch(_sleep_cmd(600), str(tmp_path))
assert "RUNNING" in bg.status(pid) assert "RUNNING" in bg.status(pid)
out = bg.stop(pid) out = bg.stop(pid)
assert "Stopped" in out assert "Stopped" in out
@@ -70,14 +86,14 @@ def test_stop_kills_running_process(tmp_path):
def test_stop_already_finished_is_graceful(tmp_path): def test_stop_already_finished_is_graceful(tmp_path):
pid = bg.launch("true", str(tmp_path)) pid = bg.launch(_true_cmd(), str(tmp_path))
assert _wait_until(lambda: bg._PROCESSES[pid].popen.poll() is not None) assert _wait_until(lambda: bg._PROCESSES[pid].popen.poll() is not None)
assert "already finished" in bg.stop(pid) assert "already finished" in bg.stop(pid)
def test_exited_elapsed_is_frozen(tmp_path): def test_exited_elapsed_is_frozen(tmp_path):
"""Elapsed for an exited process freezes at its runtime, it must not keep growing.""" """Elapsed for an exited process freezes at its runtime, it must not keep growing."""
pid = bg.launch("true", str(tmp_path)) pid = bg.launch(_true_cmd(), str(tmp_path))
assert _wait_until(lambda: bg._PROCESSES[pid].finished_ts is not None) assert _wait_until(lambda: bg._PROCESSES[pid].finished_ts is not None)
bg.status(pid) # observe exit -> records finished_ts bg.status(pid) # observe exit -> records finished_ts
proc = bg._PROCESSES[pid] proc = bg._PROCESSES[pid]
@@ -89,7 +105,7 @@ def test_exited_elapsed_is_frozen(tmp_path):
def test_watcher_records_exit_without_polling(tmp_path): def test_watcher_records_exit_without_polling(tmp_path):
"""The daemon watcher records exit on its own (no status() call needed).""" """The daemon watcher records exit on its own (no status() call needed)."""
pid = bg.launch("true", str(tmp_path)) pid = bg.launch(_true_cmd(), str(tmp_path))
assert _wait_until(lambda: bg._PROCESSES[pid].finished_ts is not None) assert _wait_until(lambda: bg._PROCESSES[pid].finished_ts is not None)
proc = bg._PROCESSES[pid] proc = bg._PROCESSES[pid]
assert proc.finished_ts is not None assert proc.finished_ts is not None
@@ -104,7 +120,7 @@ def test_on_exit_callback_fires(tmp_path):
fired["pid"] = proc.process_id fired["pid"] = proc.process_id
fired["rc"] = proc.returncode fired["rc"] = proc.returncode
pid = bg.launch("true", str(tmp_path), on_exit=cb) pid = bg.launch(_true_cmd(), str(tmp_path), on_exit=cb)
assert _wait_until(lambda: fired.get("pid") == pid and fired.get("rc") == 0) assert _wait_until(lambda: fired.get("pid") == pid and fired.get("rc") == 0)
assert fired.get("pid") == pid assert fired.get("pid") == pid
assert fired.get("rc") == 0 assert fired.get("rc") == 0
@@ -117,7 +133,7 @@ def test_unknown_id_errors_gracefully():
def test_list_all(tmp_path): def test_list_all(tmp_path):
assert "No background processes" in bg.list_all() assert "No background processes" in bg.list_all()
pid = bg.launch("sleep 1", str(tmp_path)) pid = bg.launch(_sleep_cmd(1), str(tmp_path))
listing = bg.list_all() listing = bg.list_all()
assert pid in listing assert pid in listing
assert "RUNNING" in listing assert "RUNNING" in listing
@@ -125,8 +141,8 @@ def test_list_all(tmp_path):
def test_list_all_scopes_to_origin_thread(tmp_path): def test_list_all_scopes_to_origin_thread(tmp_path):
"""list_all defaults to the launching session; include_all sees every session.""" """list_all defaults to the launching session; include_all sees every session."""
pid_a = bg.launch("sleep 1", str(tmp_path), origin_thread_id="A") pid_a = bg.launch(_sleep_cmd(1), str(tmp_path), origin_thread_id="A")
pid_b = bg.launch("sleep 1", str(tmp_path), origin_thread_id="B") pid_b = bg.launch(_sleep_cmd(1), str(tmp_path), origin_thread_id="B")
listing_a = bg.list_all("A") listing_a = bg.list_all("A")
assert pid_a in listing_a assert pid_a in listing_a
assert pid_b not in listing_a # B's process is hidden from session A assert pid_b not in listing_a # B's process is hidden from session A
@@ -137,7 +153,7 @@ def test_list_all_scopes_to_origin_thread(tmp_path):
def test_list_all_hints_at_other_sessions(tmp_path): def test_list_all_hints_at_other_sessions(tmp_path):
"""A session with no processes of its own is told others exist.""" """A session with no processes of its own is told others exist."""
bg.launch("sleep 1", str(tmp_path), origin_thread_id="A") bg.launch(_sleep_cmd(1), str(tmp_path), origin_thread_id="A")
out = bg.list_all("B") # a different session out = bg.list_all("B") # a different session
assert "other sessions" in out assert "other sessions" in out
assert "all_threads=True" in out assert "all_threads=True" in out
@@ -145,7 +161,7 @@ def test_list_all_hints_at_other_sessions(tmp_path):
def test_dedup_is_per_thread(tmp_path): def test_dedup_is_per_thread(tmp_path):
"""A check from one session must not suppress another session's completion ping.""" """A check from one session must not suppress another session's completion ping."""
pid = bg.launch("true", str(tmp_path), origin_thread_id="A") pid = bg.launch(_true_cmd(), str(tmp_path), origin_thread_id="A")
assert _wait_until(lambda: bg._PROCESSES[pid].finished_ts is not None) assert _wait_until(lambda: bg._PROCESSES[pid].finished_ts is not None)
bg.status(pid, thread_id="B") # a DIFFERENT session inspects it bg.status(pid, thread_id="B") # a DIFFERENT session inspects it
assert bg.was_observed_done(pid, "B") is True # B saw it assert bg.was_observed_done(pid, "B") is True # B saw it
+21 -6
View File
@@ -1,5 +1,6 @@
"""Tests for BackgroundExecutionMiddleware and its tools.""" """Tests for BackgroundExecutionMiddleware and its tools."""
import sys
import time import time
import pytest import pytest
@@ -14,6 +15,20 @@ from EvoScientist.middleware.background import (
) )
def _sleep_cmd(seconds: int) -> str:
"""Cross-platform command that sleeps for *seconds* and exits 0."""
if sys.platform == "win32":
return f"ping -n {seconds + 1} 127.0.0.1 > nul"
return f"sleep {seconds}"
def _true_cmd() -> str:
"""Cross-platform command that exits 0 immediately."""
if sys.platform == "win32":
return "cmd /c exit /b 0"
return "true"
def _wait_until(predicate, timeout=4.0, interval=0.05): def _wait_until(predicate, timeout=4.0, interval=0.05):
"""Poll ``predicate`` until true or ``timeout`` — avoids flaky fixed sleeps on slow CI.""" """Poll ``predicate`` until true or ``timeout`` — avoids flaky fixed sleeps on slow CI."""
deadline = time.time() + timeout deadline = time.time() + timeout
@@ -99,7 +114,7 @@ def test_run_enqueues_completion_notification(tmp_path, monkeypatch):
from EvoScientist.cli import async_notifier from EvoScientist.cli import async_notifier
monkeypatch.setattr("EvoScientist.paths.resolve_virtual_path", lambda _vp: tmp_path) monkeypatch.setattr("EvoScientist.paths.resolve_virtual_path", lambda _vp: tmp_path)
run_in_background.invoke({"command": "true", "name": "quick"}) run_in_background.invoke({"command": _true_cmd(), "name": "quick"})
# drain consumes, so accumulate across polls until the watcher's on_exit enqueues. # drain consumes, so accumulate across polls until the watcher's on_exit enqueues.
notifs = [] notifs = []
deadline = time.time() + 4.0 deadline = time.time() + 4.0
@@ -127,7 +142,7 @@ def test_notify_done_routes_to_origin_thread(tmp_path):
from EvoScientist.cli import async_notifier from EvoScientist.cli import async_notifier
from EvoScientist.middleware.background import _notify_done from EvoScientist.middleware.background import _notify_done
pid = bg.launch("true", str(tmp_path)) # no on_exit -> no auto-notify here pid = bg.launch(_true_cmd(), str(tmp_path)) # no on_exit -> no auto-notify here
assert _wait_until(lambda: bg._PROCESSES[pid].finished_ts is not None) assert _wait_until(lambda: bg._PROCESSES[pid].finished_ts is not None)
_notify_done(bg._PROCESSES[pid], "T-123") _notify_done(bg._PROCESSES[pid], "T-123")
routed = async_notifier.drain_notifications("T-123") routed = async_notifier.drain_notifications("T-123")
@@ -139,7 +154,7 @@ def test_stopped_process_suppresses_notification(tmp_path, monkeypatch):
from EvoScientist.cli import async_notifier from EvoScientist.cli import async_notifier
monkeypatch.setattr("EvoScientist.paths.resolve_virtual_path", lambda _vp: tmp_path) monkeypatch.setattr("EvoScientist.paths.resolve_virtual_path", lambda _vp: tmp_path)
run_in_background.invoke({"command": "sleep 600"}) run_in_background.invoke({"command": _sleep_cmd(600)})
(pid,) = list(bg._PROCESSES.keys()) (pid,) = list(bg._PROCESSES.keys())
stop_process.invoke({"process_id": pid}) stop_process.invoke({"process_id": pid})
# Wait until the watcher observed the exit — it would have enqueued here if the # Wait until the watcher observed the exit — it would have enqueued here if the
@@ -156,7 +171,7 @@ def test_checked_after_exit_dedups_notification(tmp_path):
dedup_notifications, dedup_notifications,
) )
pid = bg.launch("true", str(tmp_path)) pid = bg.launch(_true_cmd(), str(tmp_path))
assert _wait_until(lambda: bg._PROCESSES[pid].finished_ts is not None) assert _wait_until(lambda: bg._PROCESSES[pid].finished_ts is not None)
bg.status(pid) # agent checks AFTER exit bg.status(pid) # agent checks AFTER exit
assert bg.was_observed_done(pid) is True assert bg.was_observed_done(pid) is True
@@ -177,7 +192,7 @@ def test_not_checked_after_exit_keeps_notification(tmp_path):
dedup_notifications, dedup_notifications,
) )
pid = bg.launch("true", str(tmp_path)) pid = bg.launch(_true_cmd(), str(tmp_path))
assert _wait_until( assert _wait_until(
lambda: bg._PROCESSES[pid].finished_ts is not None lambda: bg._PROCESSES[pid].finished_ts is not None
) # exit, but do NOT check ) # exit, but do NOT check
@@ -250,7 +265,7 @@ def test_shell_notification_hints_check_process():
def test_check_and_list_route_to_manager(tmp_path, monkeypatch): def test_check_and_list_route_to_manager(tmp_path, monkeypatch):
monkeypatch.setattr("EvoScientist.paths.resolve_virtual_path", lambda _vp: tmp_path) monkeypatch.setattr("EvoScientist.paths.resolve_virtual_path", lambda _vp: tmp_path)
run_in_background.invoke({"command": "sleep 1"}) run_in_background.invoke({"command": _sleep_cmd(1)})
(pid,) = bg._PROCESSES.keys() (pid,) = bg._PROCESSES.keys()
assert pid in check_process.invoke({"process_id": pid}) assert pid in check_process.invoke({"process_id": pid})
assert pid in list_processes.invoke({}) assert pid in list_processes.invoke({})
+7
View File
@@ -2,6 +2,7 @@
from __future__ import annotations from __future__ import annotations
import sys
from pathlib import Path from pathlib import Path
from EvoScientist.cli.file_mentions import ( from EvoScientist.cli.file_mentions import (
@@ -60,6 +61,12 @@ class TestParseFileMentions:
def test_tilde_expansion(self, tmp_path: Path, monkeypatch) -> None: def test_tilde_expansion(self, tmp_path: Path, monkeypatch) -> None:
monkeypatch.setenv("HOME", str(tmp_path)) monkeypatch.setenv("HOME", str(tmp_path))
if sys.platform == "win32":
# ``ntpath.expanduser()`` falls back to ``USERPROFILE`` (and then
# ``HOMEDRIVE``+``HOMEPATH``) when ``HOME`` is absent. On some CI
# runners ``HOME`` is unset while ``USERPROFILE`` holds the real
# profile; patching both ensures ``~`` resolves to ``tmp_path``.
monkeypatch.setenv("USERPROFILE", str(tmp_path))
f = tmp_path / "file.txt" f = tmp_path / "file.txt"
f.write_text("x") f.write_text("x")
files, _ = parse_file_mentions("@~/file.txt", cwd=tmp_path) files, _ = parse_file_mentions("@~/file.txt", cwd=tmp_path)
+8 -2
View File
@@ -1,6 +1,7 @@
"""Tests for EvoScientist.mcp module.""" """Tests for EvoScientist.mcp module."""
import textwrap import textwrap
from pathlib import Path
from types import SimpleNamespace from types import SimpleNamespace
import pytest import pytest
@@ -111,7 +112,10 @@ class TestResolveCommand:
def test_found_on_path(self): def test_found_on_path(self):
"""Commands found via shutil.which are returned as full paths.""" """Commands found via shutil.which are returned as full paths."""
result = _resolve_command("python") result = _resolve_command("python")
assert result.endswith("python") or result.endswith("python3") # Cross-platform: ``shutil.which`` may return ``python.exe`` /
# ``python3.exe`` (case may differ). Match the basename stem
# case-insensitively rather than pinning the suffix.
assert Path(result).stem.lower() in ("python", "python3")
assert result != "python" # resolved, not the bare name assert result != "python" # resolved, not the bare name
def test_found_in_python_bin(self, tmp_path, monkeypatch): def test_found_in_python_bin(self, tmp_path, monkeypatch):
@@ -161,7 +165,9 @@ class TestBuildConnections:
conns = _build_connections(config) conns = _build_connections(config)
assert "fs" in conns assert "fs" in conns
assert conns["fs"]["transport"] == "stdio" assert conns["fs"]["transport"] == "stdio"
assert conns["fs"]["command"].endswith("npx") assert conns["fs"]["command"].endswith("npx") or conns["fs"][
"command"
].lower().endswith("npx.cmd")
assert conns["fs"]["args"] == ["-y", "server"] assert conns["fs"]["args"] == ["-y", "server"]
def test_stdio_with_env(self, monkeypatch): def test_stdio_with_env(self, monkeypatch):
+10 -1
View File
@@ -67,7 +67,16 @@ class TestGetDbPath(unittest.TestCase):
def test_uses_data_dir(self): def test_uses_data_dir(self):
path = get_db_path() path = get_db_path()
assert str(path).endswith("sessions.db") assert str(path).endswith("sessions.db")
assert ".evoscientist" in str(path) # On Windows ``get_db_path`` may return the 8.3 short-path
# form (e.g. ``.../EVOSCI~1/``), hiding the literal
# ``.evoscientist`` segment. ``resolve()`` walks back through
# the short-name mapping when possible, restoring the long
# form for substring matching.
try:
long_form = str(path.resolve())
except OSError:
long_form = str(path)
assert ".evoscientist" in long_form or "evoscientist" in long_form.lower()
class TestFormatRelativeTime(unittest.TestCase): class TestFormatRelativeTime(unittest.TestCase):