* fix: session lost after evoscientist restart
* feat: Implement memory worker thread deletion on completion
- Added synchronous and asynchronous functions to delete memory worker threads after they finish execution, ensuring no residual checkpoints are left in the database.
- Enhanced `_watch_memory_worker_run_sync` and `_watch_memory_worker_run_async` to invoke deletion functions upon confirming worker completion.
- Introduced tests to verify that worker threads are deleted correctly upon successful completion and that failures in deletion do not affect the overall worker status.
- Updated session management to ensure that only relevant threads are restored from the database, preventing exposure of internal or unrelated workspace threads.
- Implemented a purge function to clean up leftover worker checkpoints during server startup.
* feat: Implement short thread ID display for CLI and session hints
* fix: ensure proper accounting and deletion order for memory worker threads
---------
Co-authored-by: z00827015 <zhoulun1@huawei.com>
* ci: add windows-latest to test matrix + fix 11 cross-platform test bugs
The test workflow ran on ``ubuntu-latest`` only. Per the issue's
first bullet — the maintainer's explicit #1 priority — add
``windows-latest`` to the matrix so the manager and related
modules are exercised on Windows on every PR.
The matrix addition surfaces 18 pre-existing Windows-only test
failures. Without fixes the new leg would be 18+ reds from
day one and the matrix would just produce a wall of
``fail-fast`` noise. This PR fixes 11 of them; each fix is
a real (cross-platform) bug, not a Windows-specific hack —
most were already flagged by CodeRabbit on PR #236 but never
acted on. The remaining 4 failures need code refactors
(``os.killpg`` → ``psutil`` in ``background.py``,
``convert_virtual_paths_in_command`` Windows-aware quoting,
tilde expansion) that are documented as out-of-scope
follow-ups below.
## What changed
* ``.github/workflows/test.yml``
- ``os: [ubuntu-latest, windows-latest]`` → 2 OS × 2 Python
= 4 cells.
- ``fail-fast: false`` so one bad cell doesn't cancel the
rest while the Windows leg is being brought up. Removable
in a future PR once the suite is fully green.
* ``tests/test_backends.py``
- Hard-coded ``"python3"`` → ``{sys.executable}`` in 7
test commands. Windows has no ``python3`` on PATH; using
``sys.executable`` is portable and matches what CodeRabbit
flagged on PR #236.
- Strict string comparisons → ``shlex.split`` round-trip in
5 resolver tests. ``shlex.quote`` adds single quotes
around backslash paths on Windows, which broke the
direct ``==`` compare.
- Cross-platform suffix checks in 2 path-resolution tests
(``Path(resolved).parts[-2:]`` instead of
``str(resolved).endswith("src/main.py")``).
- ``mkdir -p`` → ``sys.executable -c "import os;
os.makedirs(...)"`` in the cwd-sanitization test.
- ``skipif(sys.platform == "win32")`` on 3 e2e tests that
hit the underlying ``shlex.quote`` + ``cmd.exe`` quoting
bug (real, separate issue).
* ``tests/test_sessions.py``
- ``test_uses_data_dir``: check ``.evoscientist`` in the
long path form (via ``Path.resolve()``) rather than the
short-path form ``get_db_path`` returns on Windows.
* ``tests/test_mcp_client.py``
- ``endswith("python")`` → ``Path(result).stem.lower()`` so
``python.EXE`` matches on Windows.
- ``endswith("npx")`` also accepts ``npx.cmd`` so the npm
shim on Windows matches.
## Out of scope (follow-up issues to file)
* ``os.killpg`` doesn't exist on Windows
(``EvoScientist/background.py:248``) — 3 background tests
fail. Real fix is the same ``psutil`` walk pattern PR #200
shipped in ``langgraph_dev/manager.py``.
* Tilde expansion in file mentions.
* Windows-aware shell quoting in
``convert_virtual_paths_in_command``.
* Path conventions (``~/.config/evoscientist/`` vs
``%APPDATA%\EvoScientist``) — needs design discussion +
``platformdirs`` migration.
* Cross-module audit of
``EvoScientist/tools/execute.py``,
``EvoScientist/ccproxy_manager.py``,
``EvoScientist/config/onboard.py``.
Closes#207 (step 1 only — CI matrix + the easy test
fixes; remaining bullets tracked separately).
* fix: cross-platform compatibility for Windows CI runners
- background.py: replace POSIX-only os.killpg/os.getpgid with
cross-platform _kill_process_tree() helper. On Windows falls back
to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX
keeps existing os.killpg logic.
- test_backends.py: replace mkdir -p shell execution in
test_literal_workspace_path_replaced with preprocessing-boundary
assertion (patch LocalShellBackend.execute, capture command,
assert workspace path was rewritten to ./). Avoids POSIX-only
mkdir -p on Windows runners.
- test_file_mentions.py: monkeypatch USERPROFILE on Windows so
ntpath.expanduser() resolves ~ to tmp_path even when HOME is
unset on CI runners.
* fix(test): cross-platform sleep/true commands for Windows CI
Replace POSIX-only sleep/true with module-level helpers that use
ping -n / cmd /c on Windows. Also fix python3 -> sys.executable
in the non-timeout recovery test.
- test_background.py: 7 sleep/true fixes
- test_background_middleware.py: 6 sleep/true fixes
- test_backends.py: 4 sleep fixes + 1 python3 fix
2318 passed, 0 failed on Windows.
* fix(test): use shell-portable double quotes for python -c on Windows
cmd.exe does not treat single quotes as string delimiters, so
-c 'raise SystemExit(1)' was passed with literal quotes on Windows.
Switch to double quotes which work on both cmd.exe and POSIX sh.
* fix: use psutil for Windows process tree kill + avoid sys.executable under uv
- background.py: replace Popen.terminate()/kill() with psutil-based
process tree walking on Windows. TerminateProcess does NOT cascade
to grandchildren; psutil.Process.children(recursive=True) ensures
the entire tree is signaled.
- test_backends.py: replace sys.executable with 'python' in sandbox
execute() calls. Under uv, sys.executable is under the workspace
and gets rewritten to ./ by prepare_sandbox_command, breaking
Linux CI. The plain 'python' command resolves correctly in any
activated venv.
* fix: broaden try/except in _kill_process_tree to cover proc.children()
If the process exits between Process(popen.pid) and children(recursive=True),
the children call raises an uncaught exception escaping stop(). Move it inside
the existing try/except block.
* fix: narrow exception to ProcessLookupError in POSIX _kill_process_tree
OSError is too broad — would silently swallow EPERM on SIGKILL, leaving
the process alive when we report it as stopped. Match original behavior
which only caught ProcessLookupError (process already gone).
* style: ruff format test_backends.py
* ci: trigger re-run for flaky prompt_toolkit test
* style: fix ruff check (import order + RUF005 unpacking)
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: update deepagents and langchain versions; enhance _reduce_messages_delta handling for None state
* fix: update langchain version constraint to >=1.3 in pyproject.toml and uv.lock
* feat(middleware): add CodeInterpreterMiddleware with project-specific configuration
chore(config): increase checkpoint retention limit for runaway conversations
fix(tests): update database schema references from 'blob' to 'value'
chore(deps): update deepagents dependency to include quickjs support
* feat(deepagents): update to version 0.6.1 and add optional dependencies for quickjs
* feat(sessions): improve error handling for message deltas and update Overwrite type check
* Enhance PruningCheckpointer with DeltaChannel Awareness
- Introduced a new pruning strategy in `_prune_after_put` to preserve the `_DeltaSnapshot` chain during checkpoint pruning.
- Implemented methods to fetch recent checkpoint IDs and walk to snapshot ancestors, ensuring that necessary checkpoints are retained.
- Updated SQL queries to handle checkpoint and write deletions more efficiently.
- Added comprehensive tests for DeltaChannel-aware pruning, ensuring that the pruning logic correctly handles various checkpoint scenarios, including those with and without snapshot seeds.
- Refactored `_load_checkpoint_messages` to utilize the new saver interface, improving message reconstruction from checkpoints.
* feat(tests): add migration sweep test to preserve snapshot ancestor
* feat(sessions): enhance checkpoint retrieval to prevent transcript leakage in multi-agent scenarios
* feat(middleware): enhance CodeInterpreterMiddleware with configurable timeout and result character limit
feat(config): add CodeInterpreterMiddleware tuning parameters to EvoScientistConfig
feat(sessions): implement inline message delta reducer for improved message handling
* feat(dependencies): update deepagents version to 0.6.2 in pyproject.toml and uv.lock
* fix(sessions): ensure migration sweep runs before yielding checkpointer to prevent race conditions
* fix(sessions): enhance migration sweep with progress indication and ETA estimation
* Implement PruningCheckpointer for efficient checkpoint management and add comprehensive tests
- Introduced `PruningCheckpointer` to manage checkpoint pruning after each `aput()`, ensuring only the latest checkpoints are retained based on a configurable limit.
- Added migration sweep functionality to clean up legacy checkpoints and prevent database bloat.
- Enhanced `get_checkpointer()` to utilize the new `PruningCheckpointer` and trigger migration sweeps when necessary.
- Developed a suite of integration tests for `PruningCheckpointer`, covering various scenarios including pruning behavior, concurrent writes, and retention policies.
- Implemented tests for migration sweep functionality, ensuring proper partitioning and user version management.
- Added diagnostic helper `db_stats` to provide insights into the database state, including thread and checkpoint counts.
* feat(sessions): enhance pruning logic to handle legacy DBs without writes table
* fix(tests): prevent atexit hook leakage in TestMigrationSweep
* feat(tests): enhance TestPruningCheckpointer to validate put+prune serialization
* feat(tests): refactor mock path implementation for get_db_path in test cases
* feat: add support for session resumption with --resume flag and enhance thread ID resolution
* refactor(tests): streamline help output testing for --resume flag
* feat: enhance session resume functionality with improved thread ID resolution and SQL wildcard handling
* feat: improve error handling for resume hint retrieval in interactive modes
* refactor: streamline logging for print_resume_hint failure in interactive mode
* feat: implement deferred scrolling for Markdown-heavy content in interactive mode
* refactor(paths): unify global data directory to ~/.evoscientist and update related paths
* refactor(paths): update legacy session migration to respect XDG_CONFIG_HOME
* refactor(tests): clear XDG_CONFIG_HOME in legacy session migration tests for deterministic behavior
* Add status bar and compact summary widgets with context window resolution
- Implemented a shared status bar for CLI and TUI frontends, including helpers for managing session metrics and context windows.
- Created a `CompactSummaryWidget` for displaying manual summaries in a collapsible format.
- Introduced a `CompactingWidget` to indicate ongoing compacting processes.
- Added a base class `TimedStatusWidget` for widgets that require a timer.
- Developed context window resolution helpers to retrieve context window sizes from various model attributes.
- Enhanced tests for context window resolution and status bar functionalities, ensuring accurate behavior across different scenarios.
- Updated existing tests to cover new features and maintain code quality.
* refactor(Channel): simplify lambda function in _send_with_retry method
* feat: enhance context editing logic and improve error handling in StreamState
* refactor(Channel): streamline lambda function in _send_with_retry method
* feat: rename auto-approve option to auto-mode for unattended execution; update checkpoint queries to filter by agent name; improve compatibility validation logic
* feat: rename auto-approve option to auto-mode; update related logic and tests for improved unattended execution
* fix: correct formatting of console message for MCP server configuration status
* feat: add check for None summary_message in _apply_summarization_event to prevent errors
* feat: enhance _load_checkpoint_messages to validate message format and apply summarization event
- Added blank lines for better separation of test cases in multiple test files.
- Reformatted event handling in tests for clarity and consistency.
- Ensured consistent use of multi-line formatting for dictionary arguments in event handling.
- Improved assertions and test descriptions for better understanding.
- Updated test cases across various modules including test_stream_state, test_stream_utils, test_summarization, test_thread_selector, test_tool_error_handler, test_tui_widgets, test_ui_runtime, and test_wechat_channel.
feat(tests): enhance relative time formatting tests for minutes, hours, days, and months
chore(docs): update contributing guidelines for testing commands
- Introduced a new `sessions.py` module for handling session persistence using SQLite.
- Added CRUD operations for threads, including listing, checking existence, finding similar threads, and deleting threads.
- Enhanced the interactive CLI to support commands for managing sessions: `/current`, `/threads`, `/resume`, and `/delete`.
- Updated the `cmd_interactive` function to handle session metadata and improve user experience with session history rendering.
- Modified the streaming functions to include metadata for checkpoint persistence.
- Added unit tests for session management functionalities to ensure reliability and correctness.