main
12 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
01845f4311 |
refactor: route middleware display events through an injected event sink (#343)
* chore: add pytest-asyncio in auto mode * test: migrate channel and stream tests to native async Convert run_async() wrapper tests to plain 'async def test_*' under pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a coroutine awaited at every call site. * test: migrate command and model/middleware tests to native async Convert run_async() wrappers (import, alias, and fixture forms) to plain 'async def test_*'. Multi-call tests merge onto one loop as sequential awaits; none asserted on loop identity. * test: migrate TUI, notifier, gateway, and session tests to native async TUI/notifier/gateway files convert run_async wrappers to plain async tests. test_sessions.py's unittest.TestCase classes move to unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async methods on plain TestCase; converting blindly would have made ~70 tests silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget in test_tui_widgets.py drops its TestCase base for the same reason. * test: replace direct asyncio.run() calls with native async tests Convert tests that called asyncio.run() (directly or via a local _run helper) to plain 'async def test_*'; delete the local helpers. * test: drop undeclared anyio markers and delete run_async helper The @pytest.mark.anyio tests relied on anyio being a transitive dep of httpx; auto-mode pytest-asyncio collects them natively. run_async() and its fixture are unreferenced after the migration, so remove them — pytest-asyncio's per-test loop teardown covers the pending-task cancellation the helper existed for (verified: full suite runs with no 'Event loop is closed' errors or destroyed-task warnings). * test: add autouse fixture for watcher cleanup * refactor: remove redundant hasattr calls * refactor: add typed middleware event sink and thread through assembly Add MiddlewareEventSink protocol + NoOpSink in middleware/events.py with a documented any-thread non-blocking contract (contract test uses a deliberately-slow fake sink). Thread an optional `events` parameter through create_cli_agent -> _get_default_middleware -> tool selector / model fallback constructors; subagent stacks are always forced to NoOpSink. * refactor: inject a notifier port into async-watcher and background middleware Add public pre_cancel_watcher() and enqueue_task_notification() to cli/async_notifier.py and a small NotifierPort protocol (middleware/notifier.py) that the module satisfies structurally. AsyncWatcherMiddleware and BackgroundExecutionMiddleware now receive the port by constructor injection at the composition root, deleting the lazy 'from ..cli import async_notifier' imports and the private _watcher_by_thread / _enqueue pokes. * refactor: invert tool-selection ownership onto a frontend event sink The adaptive tool selector now reports on_tool_selection_started / on_tool_selection / on_tool_selection_ended to the injected sink instead of writing four process-global module variables. The frontend sink (stream/sink.py FrontendEventSink) owns the selected/total/active state with consume-once + dedup-vs-last-emitted semantics; stream/tool_selection.py reads that sink object (a ToolSelectionView) rather than reaching into tool_selector's globals. Deleted: the 4 module globals, the cross-module mutations in tool_selection.py, the track_stream_selection flag, the now-vestigial _ToolSelectionTrackerMiddleware, reset_tool_selection_state_for_tests, and the autouse conftest fixture. The sink is threaded from the two interactive frontends through create_runtime_gateways -> LocalGraphGateway (read side) and _load_agent -> create_cli_agent (write side); subagent / headless stacks get NoOpSink. * refactor: route model-fallback narration through the injected event sink Delete the _ui_emit_fn / set_ui_emit module global and the ..stream.console import from model_fallback.py. The fallback middleware now reports through its injected sink: the fallback transition via the structured on_model_fallback (the frontend formats the '-> Falling back to ...' line), and the surrounding narration (primary-failure header, per-attempt outcome, exhaustion, non-fallbackable rejection) via emit_fallback_notice, preserving the exact user-facing text. The TUI binds its _append_system as the sink's fallback display where it used to call set_ui_emit (cleared on exit); the Rich CLI's sink prints to the console. _try_fallbacks / _guard_and_fallback take the sink. * refactor: declare events on the GraphGateway protocol Both gateway implementations now carry an explicit events attribute (LangGraphServerGateway holds None — no frontend renders middleware events across the HTTP boundary), so the four call sites use plain attribute access instead of getattr probing an implicit contract. * refactor: bind fallback display via the closure-scoped concrete sink The App methods used gateway.events (typed as the read-side view) and hasattr-probed for the concrete FrontendEventSink API. The enclosing factory creates that sink two hundred lines up — close over it directly: no probing, fully typed, and it becomes a constructor parameter naturally when the App class is hoisted out of the factory. * fix: end tool selection before fallback handler * fix: keep fallback display errors non-fatal * fix: preserve selector suppression for default streams * fix: restore fallback notice console display * refactor: consolidate fallback narration events * refactor: clean middleware event sink plumbing * fix: type gateway session events * refactor: make all event protocols runtime-checkable MiddlewareEventSink already carried @runtime_checkable (the stream binding guard isinstance-checks it); ToolSelectionView and SessionEvents now match, so mirroring that pattern against any of the three protocols works instead of raising TypeError. * fix(cli): close QuickJS workers after one-shot failures * fix(cli): honor no-thinking in final output * fix(channels): report failed startup accurately * fix(channels): make Telegram cleanup idempotent * fix(tui): skip command sync during exit * fix(channels): preserve startup state during retries * refactor(channels): share pending startup status * refactor(cli): expose channel startup snapshot * fix(tui): move channel startup off event loop * test(channels): release retry gate on assertion failure --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
f3e65a446f |
fix(tests): isolate the developer's real .env from the test suite (#329)
get_effective_config() runs load_dotenv(find_dotenv(usecwd=True), override=True), so any test that loads config injected the repo's real .env into os.environ for the rest of the pytest process. An empty-valued line like MINIMAX_BASE_URL= then made os.environ.get(key, default) return '' instead of the default, failing the MiniMax routing tests in full-suite runs while they passed in isolation. Generalizes the find_dotenv redirect that test_config.py's temp_config_dir fixture already applied locally into a suite-wide autouse fixture, pointing at a never-created path so tests writing their own tmp_path/.env cannot collide with it. Adds a regression test reproducing the leak. Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
690b903f85 |
test: standardize async tests on pytest-asyncio auto mode (#338)
* chore: add pytest-asyncio in auto mode * test: migrate channel and stream tests to native async Convert run_async() wrapper tests to plain 'async def test_*' under pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a coroutine awaited at every call site. * test: migrate command and model/middleware tests to native async Convert run_async() wrappers (import, alias, and fixture forms) to plain 'async def test_*'. Multi-call tests merge onto one loop as sequential awaits; none asserted on loop identity. * test: migrate TUI, notifier, gateway, and session tests to native async TUI/notifier/gateway files convert run_async wrappers to plain async tests. test_sessions.py's unittest.TestCase classes move to unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async methods on plain TestCase; converting blindly would have made ~70 tests silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget in test_tui_widgets.py drops its TestCase base for the same reason. * test: replace direct asyncio.run() calls with native async tests Convert tests that called asyncio.run() (directly or via a local _run helper) to plain 'async def test_*'; delete the local helpers. * test: drop undeclared anyio markers and delete run_async helper The @pytest.mark.anyio tests relied on anyio being a transitive dep of httpx; auto-mode pytest-asyncio collects them natively. run_async() and its fixture are unreferenced after the migration, so remove them — pytest-asyncio's per-test loop teardown covers the pending-task cancellation the helper existed for (verified: full suite runs with no 'Event loop is closed' errors or destroyed-task warnings). |
||
|
|
b1dccf17ea |
fix(tool-selector): memory tools & state for main agent (#305)
* fix(tool-selector): always include memory tools * fix(tool-selector): only track & show state for main agent * docs: update docstring |
||
|
|
2dc1e227eb |
fix(langgraph-dev): rotate langgraph_dev.log when it exceeds 50MB (#270)
* fix(langgraph-dev): rotate langgraph_dev.log when it exceeds 50MB ``_LOG_FILE`` (``~/.config/evoscientist/langgraph_dev.log``) was opened in ``start_langgraph_dev`` with plain ``"ab"`` and never rotated, so it grew unbounded over weeks/months of heavy use — especially when chatty MCP servers spawned by langgraph dev filled it, or when failure paths produced stack traces. Implement the recommended option 1 from #209: filesize-based rollover. When the active log exceeds 50MB on the next ``start_langgraph_dev`` invocation, rename it to ``langgraph_dev.log.1`` (overwriting any existing backup) via ``os.replace`` and start fresh. Single-backup policy keeps the disk footprint bounded at roughly 2x threshold. Rotation is best-effort: ``_rotate_log_if_needed`` logs and swallows OSError so a permission error or racing rename can't block langgraph dev from starting. The next ``start`` invocation will try again — worst case the log grows for one more session. Options 2 (timestamped per-session + 7-day sweep) and 3 (``RotatingFileHandler`` + pipe) are explicitly NOT done — option 1 is simplest, no async machinery, matches the issue's recommendation. Closes #209 * test(langgraph-dev): redirect _PID_DIR in rotate integration test Address CodeRabbit review comment on #270: the ``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open`` test patched only ``_LOG_FILE`` to a tmp path, but ``start_langgraph_dev`` also calls ``_PID_DIR.mkdir(...)`` as part of its prelude, which would create a real directory under ``~/.config/evoscientist/`` on a dev machine. Redirect ``_PID_DIR`` to ``tmp_path / "pids"`` too so the test stays fully isolated. Add a final assertion that ``pid_dir.is_dir()`` holds, proving the function reached past the mkdir call. * refactor(langgraph-dev): bundle runtime paths into LanggraphRuntimePaths @din0s review follow-up on #270: the previous test isolation patched only ``_LOG_FILE`` (and after a second round, ``_PID_DIR``), but ``start_langgraph_dev`` still touches 5 distinct on-disk paths. Patching any subset of those still leaves the others pointing at the user's real ``~/.config/evoscientist/`` — exactly the case that produced the "Port 6174 cannot be bound after waiting 60s" symptom on the reviewer's machine. Replace the five free-floating module-level constants (``_PID_DIR`` / ``_PID_FILE`` / ``_LOG_FILE`` / ``_WORKSPACE_SIDECAR`` / ``_FILE_LOCK_PATH``) with a single ``LanggraphRuntimePaths`` frozen dataclass exposed as a module-level ``RUNTIME`` instance. Production code accesses ``RUNTIME.pid_file`` etc.; tests can now substitute the *whole* bundle in one assignment: monkeypatch.setattr( manager, "RUNTIME", manager.LanggraphRuntimePaths.for_directory(tmp_path / "runtime"), ) The classmethod ``for_directory(pid_dir)`` builds an isolated bundle rooted at a single dir, so the test author doesn't spell out every path field. Tests that only care about one field (e.g. pid_file during the stale-process kill path) use ``dataclasses.replace(manager.RUNTIME, pid_file=X)`` — frozen dataclass-friendly, no need to enumerate the other four fields. The dataclass's docstring records the migration rationale (the old five-name layout invited inconsistent patches). External callers of the old constants updated: - ``EvoScientist/deploy/server.py`` and ``webui.py`` now import ``RUNTIME`` and use ``RUNTIME.log_file`` for the on-screen log path hint. The other imports they had (``_DEFAULT_PORT``, ``_is_port_occupied``, ``_read_workspace_sidecar``) are still module-level functions/values, untouched. Test updates: - ``tests/test_langgraph_manager.py``: ``patch.object(manager, "_XXX", X)`` patterns now go through ``dataclasses.replace(manager.RUNTIME, xxx=X)``; the ``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open`` test (from the previous #270 review iteration) uses ``for_directory`` for one-shot isolation. - ``tests/test_langgraph_dev_workspace_sidecar.py``: each test now goes through a tiny ``_isolated_runtime(monkeypatch, tmp_path)`` helper that calls ``for_directory``. - ``tests/test_langgraph_dev_deploy_mode.py``: same ``for_directory`` swap. No production behavior change. All ``langgraph_dev``-side tests (``test_langgraph_manager.py`` 26/26, ``test_langgraph_dev_workspace_sidecar.py`` 14/14, ``test_langgraph_dev_deploy_mode.py`` 14/14, ``test_cli_deploy.py`` 18/18 — which indirectly exercises deploy/server.py and deploy/webui.py imports) pass. Full-project test count unchanged from baseline; the remaining 22 Windows-only pre-existing failures (test_background ``os.killpg``, test_file_mentions tilde, mcp_client ``shutil.which``, test_sessions 8.3 short path) are documented as out-of-scope for #207. * style: apply ruff format to langgraph_dev test + module files CI lint check on #270 failed: Run ruff format --check . Would reformat: EvoScientist/langgraph_dev/manager.py Would reformat: tests/test_langgraph_manager.py Plus two test files touched by the prior consolidation commit that ``ruff format`` hadn't seen yet: tests/test_langgraph_dev_deploy_mode.py tests/test_langgraph_dev_workspace_sidecar.py Just formatting. No logic change. All 75 refactor-related tests pass. * fix(test): use for_directory for full path isolation + patch _can_bind_port to skip real socket ops Two fixes for TestStartLanggraphDevRotatesLog: 1. Replace dataclasses.replace(manager.RUNTIME, ...) with LanggraphRuntimePaths.for_directory(pid_dir) so pid_file, workspace_sidecar, and lock_file are also temp-rooted (prevents leak to ~/.config/evoscientist/). 2. Monkeypatch _can_bind_port to always return True so the bind-poll loop in _wait_for_port_bindable passes immediately without touching real sockets (fixes 60s timeout on machines where port 6174 is already in use). * fix: cross-platform compatibility for Windows CI runners - background.py: replace POSIX-only os.killpg/os.getpgid with cross-platform _kill_process_tree() helper. On Windows falls back to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX keeps existing os.killpg logic. - test_backends.py: replace mkdir -p shell execution in test_literal_workspace_path_replaced with preprocessing-boundary assertion (patch LocalShellBackend.execute, capture command, assert workspace path was rewritten to ./). Avoids POSIX-only mkdir -p on Windows runners. - test_file_mentions.py: monkeypatch USERPROFILE on Windows so ntpath.expanduser() resolves ~ to tmp_path even when HOME is unset on CI runners. * refactor(test): add runtime_paths fixture to isolate manager.RUNTIME Adds a reusable fixture that monkeypatches manager.RUNTIME to a temp-rooted LanggraphRuntimePaths.for_directory(). Tests that need specific fields can still dataclasses.replace(runtime_paths, ...) but the baseline is always temp-isolated, preventing leaks to ~/.config/evoscientist/. Updated test_langgraph_dev_deploy_mode.py, test_langgraph_dev_workspace_sidecar.py, and test_langgraph_manager.py to use the fixture, consolidating sequential lock_file + pid_dir patches into single dataclasses.replace calls. * Revert "fix: cross-platform compatibility for Windows CI runners" This reverts commit eb025d24af32e195a982cd40f6d70dba885c4019. * style: ruff format conftest.py * fix: address review issues in log-rotation + runtime paths - Use for_directory(tmp_path/pids) as base in ensure_langgraph_dev tests so pid_file/log_file are co-located with pid_dir, not split across paths - Remove unused runtime_paths param from test_no_existing_file_is_noop - Replace manager.RUNTIME with runtime_paths in two sidecar tests - Use for_directory(DEFAULT_PID_DIR) instead of explicit construction - Fix stale _LOG_FILE reference in TestRotateLogIfNeeded docstring * style: ruff format test files --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
8bb1d6c0e3 |
refactor(stream): langgraph streaming v3 (#268)
* refactor(stream): langgraph streaming v3 * fix: address CR comments * chore(stream): add success field to state --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
c407d2e20f |
Fix/async subagent model switch (#217)
* feat(middleware): add ConfigurableModelMiddleware for dynamic model resolution - Introduced ConfigurableModelMiddleware to resolve chat models from RunnableConfig.configurable on each call. - Updated middleware initialization to include ConfigurableModelMiddleware. - Enhanced context editing middleware tests to verify presence of ConfigurableModelMiddleware. - Implemented tests for ConfigurableModelMiddleware to ensure correct model overriding and caching behavior. - Added tests for deepagents model-passthrough patch to verify configuration injection in async tasks. * feat(async-subagent): update middleware handling to prevent deadlocks in async sub-agents * style: Refactor code formatting for improved readability in patches and test files * refactor: streamline middleware construction and improve async handling in ConfigurableModelMiddleware * fix: remove unused request parameter from _read_model_override function * refactor: improve async handling in _ClientProxy and enhance logging in ConfigurableModelMiddleware test: add behavior test to ensure AskUserMiddleware is excluded in async subagent mode |
||
|
|
80f1f4fa0f |
feat: Implement async sub-agent auto-notification system (#214)
* feat: Implement async sub-agent auto-notification system - Added async notifier functionality to handle notifications for sub-agents reaching terminal states. - Introduced `AsyncTaskNotification` dataclass for structured notification data. - Implemented `watch_run_and_notify` to monitor agent runs and enqueue notifications. - Created `spawn_watcher` to manage watcher tasks and ensure proper cancellation of previous watchers. - Developed `consume_notifications` to process notifications, deduplicate them, and format messages for LLM. - Added tests for notification handling, including draining, deduplication, and formatting. - Patched deepagents to integrate the new watcher functionality into start and update tools. * Enhance async notifier with per-thread notification routing and error handling - Introduced `origin_cli_thread_id` to `AsyncTaskNotification` for routing notifications back to the originating CLI session. - Implemented per-thread notification queues to handle notifications based on the originating thread. - Updated `has_pending_notifications` and `drain_notifications` to respect thread-specific queues. - Enhanced `watch_run_and_notify` to detect in-band error events from the SSE stream and handle clean exits. - Modified tests to verify the new notification routing behavior and ensure proper handling of notifications across threads. - Added a fixture to restore the async watcher patch state in tests to prevent state leakage. - Updated deepagents patching to capture the main agent's CLI thread ID for notification routing. * feat: Enhance async notifier with thread-specific watcher management and notification filtering * test: Enhance notification draining logic for cleaner test setup * refactor: Remove summary field from AsyncTaskNotification and update related tests * feat: Enhance async notification handling with target thread ID support * Refactor async notifier and middleware for improved task management - Removed the no-op shutdown watcher loop from async_notifier.py as it is no longer needed. - Updated watch_run_and_notify to clarify notification handling and race conditions. - Cleaned up shutdown handling in commands.py, interactive.py, and tui_interactive.py by removing obsolete shutdown watcher calls. - Deleted the deepagents async watcher patch from patches.py, transitioning to a new middleware approach. - Introduced AsyncWatcherMiddleware to handle async task notifications directly during tool calls. - Updated tests to validate the new middleware functionality and ensure proper watcher spawning and cancellation. - Enhanced test coverage for async watcher middleware, including edge cases and error handling. * feat(tests): add fixture to reset notifier state before each test |
||
|
|
c5a4d559a2 |
Refactor test cases for improved readability and consistency
- Added blank lines for better separation of test cases in multiple test files. - Reformatted event handling in tests for clarity and consistency. - Ensured consistent use of multi-line formatting for dictionary arguments in event handling. - Improved assertions and test descriptions for better understanding. - Updated test cases across various modules including test_stream_state, test_stream_utils, test_summarization, test_thread_selector, test_tool_error_handler, test_tui_widgets, test_ui_runtime, and test_wechat_channel. |
||
|
|
2dd331be02 | feat(tests): refactor async coroutine handling with shared run_async function | ||
|
|
e725fb08e2 | add CI workflows | ||
|
|
19548016d9 | EvoScientist Initial |