477481f9d97358d12b56e573220f8c3bb7e7fca2
339 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
477481f9d9 |
feat(backends): use _platform_quote for Windows cmd.exe compatibility (#280)
* feat(backends): use _platform_quote for Windows cmd.exe compatibility Resolves the 3 skipped E2E tests in test_backends.py that exercised the /skills/... mount path. The path-rewriter was wrapping resolved absolute paths via shlex.quote (POSIX single-quote style); cmd.exe doesn't strip single quotes, so the literal ' characters ended up in the subprocess argv and the python script failed to find its file. Replace the 3 shlex.quote call sites in _resolve_virtual_mount_path with _platform_quote, a thin platform dispatcher: - POSIX: shlex.quote (unchanged) - Windows: _cmd_quote uses cmd.exe-compatible double-quote wrapping and properly escapes embedded " and percent signs Adds: - backends.py: _is_windows, _cmd_quote, _platform_quote (~40 lines) - test_backends.py: 6 TestPlatformQuote unit tests + _split_cmd cross-platform tokenizer helper to replace shlex.split in the 8 sites that tokenize convert_virtual_paths_in_command results (POSIX shlex strips backslashes from bare Windows paths, which broke the 5 TestVirtualMountResolution assertions on Windows) Removes: - 3 @pytest.mark.skipif(sys.platform == "win32") markers on the E2E tests for /skills/... mount resolution Refs #274. * fix: escape % as %% in _cmd_quote instead of relying on double-quoting cmd.exe expands %VAR% before processing quotes, so double-quoting cannot neutralize percent signs. Escape bare % as %% (the cmd.exe idiom for a literal percent) before any other quoting logic. Also updates _cmd_quote docstring and _resolve_virtual_mount_path docstring to reflect the actual quoting strategy. * style: fix ruff format (single → double quotes) * fix: treat % as regular char in _cmd_quote, document limitation %% escaping only collapses in .bat/.cmd files, not via cmd /c. Since virtual-mount paths should never contain % in practice, simpler to leave % alone and document the caveat. |
||
|
|
76972449c7 |
feat(cli): multi-stage slash command completions with subcommand awareness (Phase 1 of #82) (#273)
* feat(cli): multi-stage slash command completions with subcommand awareness Phase 1 of #82 — subcommand and argument awareness in completions. - commands/base.py: add SubCommand dataclass and subcommands/category ClassVars to the Command ABC. Each SubCommand has name, description, and optional arguments. - commands/manager.py: add get_subcommands() and list_subcommands() methods to expose subcommand metadata for completion rendering. - commands/implementation/mcp.py: declare 6 subcommands (list, config, add, edit, remove, install). - commands/implementation/model_fallback.py: declare 6 subcommands (list, add, remove, clear, save, help). - commands/implementation/channel.py: declare 2 subcommands (status, stop). - cli/tui_interactive.py: rewrite on_text_area_changed slash-completion branch. When the user types a command name + trailing space and the command has subcommands, show subcommand completions instead of hiding the popup. Filter subcommands by typed prefix in multi-token input. - commands/implementation/general.py: /help now lists subcommands below each command that declares them. Tests: 8 new tests covering SubCommand creation, CommandManager subcommand lookup, and cross-command verification. 2293 passed baseline, no regressions. * fix: subcommand completion preserves prefix + prompt_toolkit + tests - _apply_selected_completion: preserve '/mcp ' prefix when completing subcommands via _comp_is_subcommand flag - SlashCommandCompleter (Rich CLI): add subcommand completion support - Fix trailing-space bug: rstrip prefix before top-level matching - test_tui_widgets.py: update stub on_input_changed to match multi-stage logic; add 5 new subcommand tests - test_command_manager.py: 8 tests for SubCommand + CommandManager 28 passed, 0 failed. * style: ruff format tui_interactive.py + test_tui_widgets.py * style: fix RUF012 ClassVar annotation on subcommands lists * fix: sync test stub, add len>=3 guard, remove exact-match hide - Sync test stub on_input_changed with real TUI code (remove exact-match hide for subcommands, add len(parts)>=3 guard) - Update test_input_changed_exact_subcommand_hides -> shows_confirmation - Add test_input_changed_three_parts_hides - Remove unused category ClassVar (din0s: what is this for) * refactor(commands): extract shared completion engine Per din0s feedback: one shared completion engine (commands/_completion_engine.py) that parses text + cursor once, returns structured CompletionCandidate objects with replace_start/replace_end ranges. - SlashCommandCompleter (Rich CLI): thin adapter, delegates to engine - on_text_area_changed (TUI): thin adapter, delegates to engine - _apply_selected_completion: uses candidate.replace_start/replace_end instead of _comp_is_subcommand flag - Tests: engine tested directly (10 new tests), stub methods updated 29 passed, 0 failed. * style: ruff format * fix: preserve text after cursor when applying completion CodeRabbit: replace_start only cuts from start to cursor, dropping any suffix after the cursor. Use replace_start + replace_end to correctly splice the replacement while preserving trailing text. * fix(cli): repair slash-command completion (TUI crash, subcommand bugs, sort) Apology + context: the previous push shipped a TUI-breaking change (the new shared engine assumed ``event.text_area.cursor_position`` existed, but ``ChatTextArea`` / ``Changed`` don't expose it). User caught the crash on ``/``; fixing that surfaced two more bugs in the engine that din0s had already flagged. This commit addresses all of them and drops a piece of dead stub code. ## Bug fixes 1. **TUI crash on ``/``** (``tui_interactive.py:2335``) ``event.cursor_position`` doesn't exist on the ``Changed`` event, and ``ChatTextArea`` (Textual ``TextArea`` subclass) doesn't expose ``cursor_position`` either. Pass ``len(event.text_area.text)`` instead — the user types at the end of the input in practice. 2. **Subcommand trailing-space duplication** (``_completion_engine.py``) Typing ``/mcp a `` + Tab produced ``/mcp aadd``. The engine included the trailing space in ``replace_end``; the TUI apply unconditionally appended ``" "``, producing double-space output. Fix: ``replace_end`` excludes the trailing space; the TUI apply checks ``current[replace_end:].startswith(" ")`` and skips the separator when the suffix already has one. 3. **Subcommand exact-match confirmation noise** (``_completion_engine.py``) Typing ``/mcp list`` + Tab re-inserted ``list`` and the popup kept showing the same subcommand. Add a guard mirroring the top-level exact-match rule: when the only subcommand match is the prefix itself (no trailing space), return ``empty``. 4. **Alphabetical sort dropped in CLI** (``cli/interactive.py``) The new completer iterated ``result.candidates`` in registration order. Re-add ``sorted(result.candidates, key=lambda c: c.text)``. Same sort added to the TUI for consistency. ## Cleanup - Drop the dead ``on_input_changed`` method from the ``_StubApp`` test stub (0 call sites) plus the unused ``_slash_commands`` / ``_subcommands`` locals that fed it. This addresses din0s's comment about the stub duplicating real TUI logic — the inlined copy is no longer needed since the real completer now routes through the shared engine. ## Tests - ``test_engine_exact_subcommand_shows_confirmation`` → renamed to ``test_engine_exact_subcommand_hides`` to match new behavior. - New: ``test_engine_subcommand_trailing_space_excludes_space_from_range`` and ``test_engine_subcommand_trailing_space_apply_does_not_double_space``. - All 97 tests in ``test_tui_widgets.py`` pass. - ``ruff check`` / ``ruff format`` clean. - Local TUI smoke: ``/`` (no crash, top-level popup), ``/mcp `` (subcommand popup), ``/mcp a `` + Tab → ``/mcp add ``. Refs the din0s review comments on PR #273. CLI path tests and the ``category`` ClassVar follow-up are deferred to a separate PR (the former is a test-suite addition; the latter is already absent from ``base.py`` on the current branch). * fix: address remaining review items (help duplication, CLI tests, stub sync, docstrings) - mcp.py: auto-generate help text from subcommands ClassVar (#1) - tests/test_cli_completion.py: add 9 CLI completer tests (#2c) - test_tui_widgets.py: sync _apply_selected_completion stub with real code (#4) - mcp.py + interactive.py: add docstrings to key functions (#8) * fix: hide completions on exact subcommand match regardless of trailing space Remove the ot has_trailing_space guard from the exact-subcommand check. Previously /mcp list (with trailing space) would still return candidates, causing Tab to oscillate between adding and removing the trailing whitespace. Now the engine hides whenever the subcommand is an exact match, same as the top-level rule. Added test_engine_exact_subcommand_with_trailing_space_hides to cover the scenario din0s flagged. * refactor: use StrEnum for CompletionResult.kind Replace plain str with CompletionKind(StrEnum) for type safety. Backward-compatible with existing string comparisons. * fix: normalize @file completion tuples to CompletionCandidate complete_file_mention() returns list[tuple[str, str]] but the TUI rendering/apply code expects objects with .text/.description. Wrap tuples in CompletionCandidate to prevent AttributeError crash. --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
8f9dfa159e |
fix(backends): rewrite quoted virtual paths containing whitespace (#269)
* fix(backends): rewrite quoted virtual paths containing whitespace
The `convert_virtual_paths_in_command` regex
`(?<=\s)/[^\s;|&<>'"`]*` stopped at the first whitespace or quote,
so:
- `python "/skills/my skill/main.py"` was left completely
unchanged (the `(?<=\s)` lookbehind failed after the opening
`"`), and the shell then broke the inner unquoted path at the
embedded space.
- `python /skills/my skill/main.py` was truncated to
`python ./skills/my skill/main.py` (only `/skills/my` rewritten).
Replace the regex with `shlex.shlex(command, posix=True,
punctuation_chars=";|&<>")` so quoted regions stay whole, then
splice the rewrite back into the original command — extending the
splice span to include any matching quote chars around the path so
the fresh `shlex.quote` of the replacement isn't double-wrapped.
`_resolve_virtual_mount_path` now returns the unquoted path; the
caller owns shell-quoting, which avoids the previous
`shlex.quote` inside original `"…"` leaving literal `'` chars in
the argument value.
Unquoted paths with embedded whitespace remain a known limitation
(shlex has no way to know the user meant one path) — the
workaround of avoiding spaces in skill directory names still
applies, as flagged in the original issue.
Closes #237
* fix(backends): backslash-escaped paths, multi-path per token, subshell paths
- Fix backslash-escape handling: use unescape before rewriting
- Fix re.search→re.finditer: all /-paths in a token are rewritten
- Keep ( ) and backticks inside word tokens so paths spanning
\ or wrapped in backticks are matched correctly
- Add _try_rewrite helper with URL detection and unescape logic
- Add 10 contract tests pinning the din0s review cases
* fix(backends): restore () and backtick as shell operators for validate_command
- Restore ( ) and backtick to the operator set in _shell_token_spans.
Removing them caused a security regression: commands like (sudo ls)
would not detect sudo as a blocked command because (sudo became one
word token. With operators restored, validate_command correctly
catches blocked commands inside subshells and command substitutions.
- Fix _value_span_to_raw_span: the 'quoted' flag from the tokenizer
means the token *contains* a quoted segment (not necessarily starts
with a quote). Replace raw[0] assumption with a forward scan for
the first quote char, consuming unquoted prefix chars 1:1.
- Update test_system_path_with_shell_expansion: paths are now
partially rewritten because () are operators. Test updated to
reflect this known limitation (security > path rewriting).
* fix(test): cross-platform compatibility for pre-existing Windows failures
- python3 -> python in execute() calls (python is on PATH in any activated venv)
- sleep 10 -> _sleep_cmd(10) cross-platform helper
- str().endswith() -> Path().parts assertions (backslash-safe on Windows)
- shlex.quote exact-match assertions -> 'in' assertions (Windows quotes paths differently)
- mkdir -p E2E test -> preprocessor boundary test
- Skip 3 E2E tests on Windows: shlex.quote produces POSIX quoting incompatible with cmd.exe
141 passed, 3 skipped on Windows.
* fix: update docstring + strengthen shell-expansion test assertion
- Fix _value_span_to_raw_span docstring: no longer assumes raw[0] is
the opening quote, scans forward for first quote char
- Strengthen test_system_path_with_shell_expansion: verify
./workspace/notes is rewritten, not just notes in result
* style: ruff format backends.py + test_backends.py
* refactor(backends): simplify quoted virtual path rewriting
Replace 500+ line shlex tokenizer with 12-line pre-process step. Match quoted args via regex, unescape, rewrite via _rewrite_quoted_path, substitute with shlex.quote. 133 passed, 3 skipped.
* fix: guard bare absolute paths from double-rewrite by post-process regex
On POSIX, shlex.quote returns bare paths (e.g. /tmp/memories/note.md).
The pre-process substitutes these into the command, then the post-process
regex re-matches and incorrectly rewrites them.
Fix: _guard_bare_absolute wraps bare /-paths in single quotes so the
post-process regex''s character class stops at the quote char.
* style: ruff format
* fix(backends): narrow pre-process to exclude system-prefixed paths
Only rewrite quoted paths that are NOT known system prefixes.
* fix: narrow quoted-path pre-process to virtual mounts only
Only rewrite quoted /... paths that resolve to actual virtual mounts (/skills/..., /memories/...) or workspace-prefixed system paths. Remove catch-all that incorrectly rewrote bare paths like echo /hi.
Addresses din0s review feedback on #269.
* docs: update docstring for narrower quoted-path rewrite scope
|
||
|
|
05a5e5f8a0 |
fix: session lost after evoscientist restart (supersedes #278) (#279)
* fix: session lost after evoscientist restart * feat: Implement memory worker thread deletion on completion - Added synchronous and asynchronous functions to delete memory worker threads after they finish execution, ensuring no residual checkpoints are left in the database. - Enhanced `_watch_memory_worker_run_sync` and `_watch_memory_worker_run_async` to invoke deletion functions upon confirming worker completion. - Introduced tests to verify that worker threads are deleted correctly upon successful completion and that failures in deletion do not affect the overall worker status. - Updated session management to ensure that only relevant threads are restored from the database, preventing exposure of internal or unrelated workspace threads. - Implemented a purge function to clean up leftover worker checkpoints during server startup. * feat: Implement short thread ID display for CLI and session hints * fix: ensure proper accounting and deletion order for memory worker threads --------- Co-authored-by: z00827015 <zhoulun1@huawei.com> |
||
|
|
526c571b10 |
feat(tunnel): add Cloudflare tunnel support for EvoSci deploy and update documentation
|
||
|
|
49f23560fd | chore: update version to v0.1.5 | ||
|
|
c02be519f6 |
feat(dangerous-mode): implement real-filesystem access with safety ch… (#276)
* feat(dangerous-mode): implement real-filesystem access with safety checks - Introduced a 'dangerous mode' allowing the agent to operate on the real filesystem. - Updated command validation to bypass path confinement while enforcing a blocklist for privileged commands. - Added warnings and guidelines for users when operating in dangerous mode. - Enhanced configuration to support dangerous mode and ensure it implies auto-approval. - Updated tests to verify the behavior of commands and configurations in dangerous mode. * feat(dangerous-mode): enhance logging and environment management for dangerous mode * feat(dangerous-mode): improve handling of dangerous mode with environment flags and enhance test isolation |
||
|
|
d4f1fbd110 |
ci: add windows-latest to test matrix + fix 11 cross-platform test bugs (#271)
* ci: add windows-latest to test matrix + fix 11 cross-platform test bugs The test workflow ran on ``ubuntu-latest`` only. Per the issue's first bullet — the maintainer's explicit #1 priority — add ``windows-latest`` to the matrix so the manager and related modules are exercised on Windows on every PR. The matrix addition surfaces 18 pre-existing Windows-only test failures. Without fixes the new leg would be 18+ reds from day one and the matrix would just produce a wall of ``fail-fast`` noise. This PR fixes 11 of them; each fix is a real (cross-platform) bug, not a Windows-specific hack — most were already flagged by CodeRabbit on PR #236 but never acted on. The remaining 4 failures need code refactors (``os.killpg`` → ``psutil`` in ``background.py``, ``convert_virtual_paths_in_command`` Windows-aware quoting, tilde expansion) that are documented as out-of-scope follow-ups below. ## What changed * ``.github/workflows/test.yml`` - ``os: [ubuntu-latest, windows-latest]`` → 2 OS × 2 Python = 4 cells. - ``fail-fast: false`` so one bad cell doesn't cancel the rest while the Windows leg is being brought up. Removable in a future PR once the suite is fully green. * ``tests/test_backends.py`` - Hard-coded ``"python3"`` → ``{sys.executable}`` in 7 test commands. Windows has no ``python3`` on PATH; using ``sys.executable`` is portable and matches what CodeRabbit flagged on PR #236. - Strict string comparisons → ``shlex.split`` round-trip in 5 resolver tests. ``shlex.quote`` adds single quotes around backslash paths on Windows, which broke the direct ``==`` compare. - Cross-platform suffix checks in 2 path-resolution tests (``Path(resolved).parts[-2:]`` instead of ``str(resolved).endswith("src/main.py")``). - ``mkdir -p`` → ``sys.executable -c "import os; os.makedirs(...)"`` in the cwd-sanitization test. - ``skipif(sys.platform == "win32")`` on 3 e2e tests that hit the underlying ``shlex.quote`` + ``cmd.exe`` quoting bug (real, separate issue). * ``tests/test_sessions.py`` - ``test_uses_data_dir``: check ``.evoscientist`` in the long path form (via ``Path.resolve()``) rather than the short-path form ``get_db_path`` returns on Windows. * ``tests/test_mcp_client.py`` - ``endswith("python")`` → ``Path(result).stem.lower()`` so ``python.EXE`` matches on Windows. - ``endswith("npx")`` also accepts ``npx.cmd`` so the npm shim on Windows matches. ## Out of scope (follow-up issues to file) * ``os.killpg`` doesn't exist on Windows (``EvoScientist/background.py:248``) — 3 background tests fail. Real fix is the same ``psutil`` walk pattern PR #200 shipped in ``langgraph_dev/manager.py``. * Tilde expansion in file mentions. * Windows-aware shell quoting in ``convert_virtual_paths_in_command``. * Path conventions (``~/.config/evoscientist/`` vs ``%APPDATA%\EvoScientist``) — needs design discussion + ``platformdirs`` migration. * Cross-module audit of ``EvoScientist/tools/execute.py``, ``EvoScientist/ccproxy_manager.py``, ``EvoScientist/config/onboard.py``. Closes #207 (step 1 only — CI matrix + the easy test fixes; remaining bullets tracked separately). * fix: cross-platform compatibility for Windows CI runners - background.py: replace POSIX-only os.killpg/os.getpgid with cross-platform _kill_process_tree() helper. On Windows falls back to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX keeps existing os.killpg logic. - test_backends.py: replace mkdir -p shell execution in test_literal_workspace_path_replaced with preprocessing-boundary assertion (patch LocalShellBackend.execute, capture command, assert workspace path was rewritten to ./). Avoids POSIX-only mkdir -p on Windows runners. - test_file_mentions.py: monkeypatch USERPROFILE on Windows so ntpath.expanduser() resolves ~ to tmp_path even when HOME is unset on CI runners. * fix(test): cross-platform sleep/true commands for Windows CI Replace POSIX-only sleep/true with module-level helpers that use ping -n / cmd /c on Windows. Also fix python3 -> sys.executable in the non-timeout recovery test. - test_background.py: 7 sleep/true fixes - test_background_middleware.py: 6 sleep/true fixes - test_backends.py: 4 sleep fixes + 1 python3 fix 2318 passed, 0 failed on Windows. * fix(test): use shell-portable double quotes for python -c on Windows cmd.exe does not treat single quotes as string delimiters, so -c 'raise SystemExit(1)' was passed with literal quotes on Windows. Switch to double quotes which work on both cmd.exe and POSIX sh. * fix: use psutil for Windows process tree kill + avoid sys.executable under uv - background.py: replace Popen.terminate()/kill() with psutil-based process tree walking on Windows. TerminateProcess does NOT cascade to grandchildren; psutil.Process.children(recursive=True) ensures the entire tree is signaled. - test_backends.py: replace sys.executable with 'python' in sandbox execute() calls. Under uv, sys.executable is under the workspace and gets rewritten to ./ by prepare_sandbox_command, breaking Linux CI. The plain 'python' command resolves correctly in any activated venv. * fix: broaden try/except in _kill_process_tree to cover proc.children() If the process exits between Process(popen.pid) and children(recursive=True), the children call raises an uncaught exception escaping stop(). Move it inside the existing try/except block. * fix: narrow exception to ProcessLookupError in POSIX _kill_process_tree OSError is too broad — would silently swallow EPERM on SIGKILL, leaving the process alive when we report it as stopped. Match original behavior which only caught ProcessLookupError (process already gone). * style: ruff format test_backends.py * ci: trigger re-run for flaky prompt_toolkit test * style: fix ruff check (import order + RUF005 unpacking) --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
cf5e0dd3bd | feat(models): add support for 'claude-fable-5' mode | ||
|
|
2dc1e227eb |
fix(langgraph-dev): rotate langgraph_dev.log when it exceeds 50MB (#270)
* fix(langgraph-dev): rotate langgraph_dev.log when it exceeds 50MB ``_LOG_FILE`` (``~/.config/evoscientist/langgraph_dev.log``) was opened in ``start_langgraph_dev`` with plain ``"ab"`` and never rotated, so it grew unbounded over weeks/months of heavy use — especially when chatty MCP servers spawned by langgraph dev filled it, or when failure paths produced stack traces. Implement the recommended option 1 from #209: filesize-based rollover. When the active log exceeds 50MB on the next ``start_langgraph_dev`` invocation, rename it to ``langgraph_dev.log.1`` (overwriting any existing backup) via ``os.replace`` and start fresh. Single-backup policy keeps the disk footprint bounded at roughly 2x threshold. Rotation is best-effort: ``_rotate_log_if_needed`` logs and swallows OSError so a permission error or racing rename can't block langgraph dev from starting. The next ``start`` invocation will try again — worst case the log grows for one more session. Options 2 (timestamped per-session + 7-day sweep) and 3 (``RotatingFileHandler`` + pipe) are explicitly NOT done — option 1 is simplest, no async machinery, matches the issue's recommendation. Closes #209 * test(langgraph-dev): redirect _PID_DIR in rotate integration test Address CodeRabbit review comment on #270: the ``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open`` test patched only ``_LOG_FILE`` to a tmp path, but ``start_langgraph_dev`` also calls ``_PID_DIR.mkdir(...)`` as part of its prelude, which would create a real directory under ``~/.config/evoscientist/`` on a dev machine. Redirect ``_PID_DIR`` to ``tmp_path / "pids"`` too so the test stays fully isolated. Add a final assertion that ``pid_dir.is_dir()`` holds, proving the function reached past the mkdir call. * refactor(langgraph-dev): bundle runtime paths into LanggraphRuntimePaths @din0s review follow-up on #270: the previous test isolation patched only ``_LOG_FILE`` (and after a second round, ``_PID_DIR``), but ``start_langgraph_dev`` still touches 5 distinct on-disk paths. Patching any subset of those still leaves the others pointing at the user's real ``~/.config/evoscientist/`` — exactly the case that produced the "Port 6174 cannot be bound after waiting 60s" symptom on the reviewer's machine. Replace the five free-floating module-level constants (``_PID_DIR`` / ``_PID_FILE`` / ``_LOG_FILE`` / ``_WORKSPACE_SIDECAR`` / ``_FILE_LOCK_PATH``) with a single ``LanggraphRuntimePaths`` frozen dataclass exposed as a module-level ``RUNTIME`` instance. Production code accesses ``RUNTIME.pid_file`` etc.; tests can now substitute the *whole* bundle in one assignment: monkeypatch.setattr( manager, "RUNTIME", manager.LanggraphRuntimePaths.for_directory(tmp_path / "runtime"), ) The classmethod ``for_directory(pid_dir)`` builds an isolated bundle rooted at a single dir, so the test author doesn't spell out every path field. Tests that only care about one field (e.g. pid_file during the stale-process kill path) use ``dataclasses.replace(manager.RUNTIME, pid_file=X)`` — frozen dataclass-friendly, no need to enumerate the other four fields. The dataclass's docstring records the migration rationale (the old five-name layout invited inconsistent patches). External callers of the old constants updated: - ``EvoScientist/deploy/server.py`` and ``webui.py`` now import ``RUNTIME`` and use ``RUNTIME.log_file`` for the on-screen log path hint. The other imports they had (``_DEFAULT_PORT``, ``_is_port_occupied``, ``_read_workspace_sidecar``) are still module-level functions/values, untouched. Test updates: - ``tests/test_langgraph_manager.py``: ``patch.object(manager, "_XXX", X)`` patterns now go through ``dataclasses.replace(manager.RUNTIME, xxx=X)``; the ``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open`` test (from the previous #270 review iteration) uses ``for_directory`` for one-shot isolation. - ``tests/test_langgraph_dev_workspace_sidecar.py``: each test now goes through a tiny ``_isolated_runtime(monkeypatch, tmp_path)`` helper that calls ``for_directory``. - ``tests/test_langgraph_dev_deploy_mode.py``: same ``for_directory`` swap. No production behavior change. All ``langgraph_dev``-side tests (``test_langgraph_manager.py`` 26/26, ``test_langgraph_dev_workspace_sidecar.py`` 14/14, ``test_langgraph_dev_deploy_mode.py`` 14/14, ``test_cli_deploy.py`` 18/18 — which indirectly exercises deploy/server.py and deploy/webui.py imports) pass. Full-project test count unchanged from baseline; the remaining 22 Windows-only pre-existing failures (test_background ``os.killpg``, test_file_mentions tilde, mcp_client ``shutil.which``, test_sessions 8.3 short path) are documented as out-of-scope for #207. * style: apply ruff format to langgraph_dev test + module files CI lint check on #270 failed: Run ruff format --check . Would reformat: EvoScientist/langgraph_dev/manager.py Would reformat: tests/test_langgraph_manager.py Plus two test files touched by the prior consolidation commit that ``ruff format`` hadn't seen yet: tests/test_langgraph_dev_deploy_mode.py tests/test_langgraph_dev_workspace_sidecar.py Just formatting. No logic change. All 75 refactor-related tests pass. * fix(test): use for_directory for full path isolation + patch _can_bind_port to skip real socket ops Two fixes for TestStartLanggraphDevRotatesLog: 1. Replace dataclasses.replace(manager.RUNTIME, ...) with LanggraphRuntimePaths.for_directory(pid_dir) so pid_file, workspace_sidecar, and lock_file are also temp-rooted (prevents leak to ~/.config/evoscientist/). 2. Monkeypatch _can_bind_port to always return True so the bind-poll loop in _wait_for_port_bindable passes immediately without touching real sockets (fixes 60s timeout on machines where port 6174 is already in use). * fix: cross-platform compatibility for Windows CI runners - background.py: replace POSIX-only os.killpg/os.getpgid with cross-platform _kill_process_tree() helper. On Windows falls back to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX keeps existing os.killpg logic. - test_backends.py: replace mkdir -p shell execution in test_literal_workspace_path_replaced with preprocessing-boundary assertion (patch LocalShellBackend.execute, capture command, assert workspace path was rewritten to ./). Avoids POSIX-only mkdir -p on Windows runners. - test_file_mentions.py: monkeypatch USERPROFILE on Windows so ntpath.expanduser() resolves ~ to tmp_path even when HOME is unset on CI runners. * refactor(test): add runtime_paths fixture to isolate manager.RUNTIME Adds a reusable fixture that monkeypatches manager.RUNTIME to a temp-rooted LanggraphRuntimePaths.for_directory(). Tests that need specific fields can still dataclasses.replace(runtime_paths, ...) but the baseline is always temp-isolated, preventing leaks to ~/.config/evoscientist/. Updated test_langgraph_dev_deploy_mode.py, test_langgraph_dev_workspace_sidecar.py, and test_langgraph_manager.py to use the fixture, consolidating sequential lock_file + pid_dir patches into single dataclasses.replace calls. * Revert "fix: cross-platform compatibility for Windows CI runners" This reverts commit eb025d24af32e195a982cd40f6d70dba885c4019. * style: ruff format conftest.py * fix: address review issues in log-rotation + runtime paths - Use for_directory(tmp_path/pids) as base in ensure_langgraph_dev tests so pid_file/log_file are co-located with pid_dir, not split across paths - Remove unused runtime_paths param from test_no_existing_file_is_noop - Replace manager.RUNTIME with runtime_paths in two sidecar tests - Use for_directory(DEFAULT_PID_DIR) instead of explicit construction - Fix stale _LOG_FILE reference in TestRotateLogIfNeeded docstring * style: ruff format test files --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
cd2baa9588 |
feat(models): add opt-in prompt caching support for anthropic via openrouter (#272)
* feat(models): add opt-in prompt caching support for anthropic via openrouter * chore: don't coerce model_kwargs to dict |
||
|
|
4b6a969df2 |
refactor(agent): make create_cli_agent(config=, chat_model=) pure (#267)
* refactor(agent): make create_cli_agent(config=, chat_model=) pure Re-applies the #183 purity refactor on top of the observation-memory lifecycle that landed in #259, integrating the two cleanly. create_cli_agent gains a pure path: when both `config` and `chat_model` are passed it builds the agent entirely from locals and writes none of the cached module globals (`_config`, `_chat_model`, `_chat_model_key`, `_EvoScientist_agent`). `/model` commits the switch via `set_active_config` / `set_chat_model_instance` only after a successful build, so a failed rebuild leaves the session on the original model (replaces the old snapshot/restore rollback). Supporting changes: - Extract `set_active_config` (write-half of `_ensure_config`), `_apply_env_from_config`, `_build_chat_model`, and `set_chat_model_instance`. - Thread `cfg` / `chat_model` through `_get_default_middleware`, `_build_base_kwargs`, `load_mcp_and_build_kwargs`, `_maybe_swap_async_subagents`, and `_inject_subagent_middleware` so the pure path never falls back to the global-writing `_ensure_config()` / `_ensure_chat_model()`. - Integrate with #259's memory middleware: subagent context-editing middleware binds the threaded `chat_model`, and the configured system prompt / memory controls read the threaded `cfg` (new threading vs the original #183, required because #259 made these paths read config). - Consolidate `cfg` resolution to one `cfg if cfg is not None else _ensure_config()` at the top of each kwargs builder, matching the pattern already used in the other config-aware helpers. * fix(agent): keep pure tool selector off global cache * fix(model): apply config switch in place to preserve reference integrity --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com> |
||
|
|
8bb1d6c0e3 |
refactor(stream): langgraph streaming v3 (#268)
* refactor(stream): langgraph streaming v3 * fix: address CR comments * chore(stream): add success field to state --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
ac052bbb4c |
feat: free-scrolling (#262)
* feat: free-scrolling * fix: anchor * fix: textual private vars |
||
|
|
b4ffb35d71 | feat(langgraph): add --no-reload option to start_langgraph_dev | ||
|
|
63969b596d |
Release/v0.1.4 (#266)
* feat(middleware): reposition code interpreter middleware in the stack * feat(models): add qwen3.7-plus model entry and update context window comment * feat(models): add qwen3.7-max and qwen3.7-plus model entries for DashScope * feat(auxiliary): implement auxiliary model support for background tasks and tool selection - Added auxiliary model configuration to EvoScientistConfig. - Introduced _ensure_auxiliary_chat_model function to manage auxiliary model instances. - Updated onboarding steps to include auxiliary model selection. - Modified middleware to route tool selection to the auxiliary model when applicable. - Enhanced tests to cover auxiliary model functionality and configuration. * feat(steps): update UI backend selection options and descriptions * Refactor code structure for improved readability and maintainability * feat(patches): implement OpenRouter response reasoning item stripping to prevent multi-turn errors * feat: update version to v0.1.4 in badges, README, and pyproject.toml; adjust skill counts in steps.py * feat(config): add auxiliary model and provider environment variables to test setup |
||
|
|
3563c1d94f | Update source for Scientific Skills in steps.py (#265) | ||
|
|
92d95dee68 |
feat(memory): add observation memory lifecycle (#259)
* feat(memory): add observation memory lifecycle Add file-backed observation memory with deterministic markdown records, structured record_observation tooling, startup indexing, and profile/observation prompt guidance. Launch post-turn and post-subagent EvoMemory workers through LangGraph dev so completed runs can update profile memory, save durable observations, and write subagent execution summaries without blocking the active agent. Wire memory middleware into the main agent, subagents, async graphs, TUI status reporting, worker activity accounting, and observation-aware research prompts, with regression coverage for storage, lifecycle scheduling, graph registration, status display, and stream reset behavior. * fix(cli): sync background agent server on resume Resume flows now need to keep the LangGraph dev background server aligned with the active workspace even when async subagents are disabled. EvoMemory workers use that server too, so gating resume-time sync on enable_async_subagents could leave workers pinned to the launch workspace after resuming a thread from another workspace. Run workspace sync unconditionally for Rich CLI and Textual resume paths, while preserving WorkspaceMismatchError handling so failed sync aborts the resume before mutating the active thread or workspace. Propagate aborted resume callbacks through the command UI so channel-issued /resume commands do not send false success or history output. Channel slash dispatch now treats CommandManager-caught command errors as command errors and skips completion hooks for those failed commands. Add regression coverage for disabled async subagents, callback aborts, and channel command error reporting. * fix(cli): prepare serve resume workspace before adopting Load the resumed workspace agent and sync the background server as a single pre-adoption step. Restore the previous active workspace if preparation fails so serve mode keeps using the old session consistently. * fix(memory): untrack abandoned worker status watches Stop treating watcher shutdown as confirmed worker completion. Terminal worker statuses still count memory deltas, while poll failures or watcher setup failures now remove the active run without crediting partial outputs. * fix(cli): report channel command failures accurately Treat command_error as a None sentinel so empty error strings still fail, and let TUI resumes continue only on non-mismatch background-server sync failures while reporting degraded mode. * fix(stream): clear memory counters for resume streams Reset completed-memory counters for every new agent stream, including Command-based HITL and resume streams, so saved-memory indicators do not leak across turns. * docs(tools): make observation recording guidance conditional Clarify that agents should call record_observation only when the observation tool is available, preserving the existing durability and usefulness criteria. * feat(config): add controls for profile and observation memory Add config flags for profile memory, observation memory, observation writer placement, and background memory workers. Wire the controls through main agents, subagents, EvoMemory middleware, and memory lifecycle workers so observation writes can be assigned to the live agent, subagent worker, both, or neither. Keep turn memory workers profile-only and make prompts reflect the available observation read/write paths. Skip langgraph dev startup when neither async subagents nor memory workers need the background server. Add coverage for config parsing, prompt gating, middleware wiring, and worker tool availability. * test(cli): include memory defaults in serve config stubs * fix(memory): offload async worker launch blocking calls Run the langgraph-dev health check and memory-output snapshot in worker threads from the async EvoMemory launcher so it does not block the event loop. * chore(memory): harden turn worker subagent guardrail * chore(memory): refresh profile context per request * fix(memory): offload async profile file reads * fix(memory): offload async worker completion accounting |
||
|
|
cedb3aa744 |
Fix ssh remote path handling in sandbox (#242)
* Fix ssh remote path handling in sandbox * Address ssh remote command review feedback * Format backend files with ruff * Narrow SSH remote command handling * Narrow SSH preprocessing to single-quoted remote args * Address remaining SSH preprocessing review feedback * Tighten SSH executable recognition * Recognize only literal ssh wrapper --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
9cffe9d457 |
Enhance multimodal handling in LLM model (#256)
* Enhance multimodal handling in LLM model - Updated `_flatten_message_content` to preserve media blocks (images, files) while flattening text content. - Introduced `_sanitize_messages` to manage media hoisting for tool messages, ensuring compatibility with OpenAI APIs. - Modified `_patch_openai_compat_content` to accommodate new media handling logic, including retry mechanisms for media errors. - Added comprehensive tests for media preservation, including various scenarios with images, files, and unsupported media types. * fix: preserve order of text and media blocks in message flattening * test: add tests for _strip_media_types to ensure position preservation and deduplication |
||
|
|
d348076f40 | Add runtime context middleware (#255) | ||
|
|
9285c6dad8 |
Migrate memory middleware to profile files (#253)
* feat(memory): migrate to profile memory files * chore(stream): read profile headings from templates * fix(display): keep assistant responses if response_text has started * fix(memory): do not treat failed bootstraps as profile creation * chore(memory): unlink blank legacy memory * fix(memory): resolve project_id once * fix(memory): preserve unreadable profile files * chore(tui): render streamed narration inline with tool timeline Update the TUI streaming timeline so assistant text emitted before or between tool calls is rendered inline where it occurs, rather than being kept as a single answer bubble above or below the tools. If the model begins an assistant response and then emits another tool call, the provisional response is converted into inline narration before that tool. The final assistant message then renders only the remaining response suffix, avoiding duplicate text in the completed transcript. Stop/cancel handling now preserves any active inline narration, appends the visible stopped marker only to the remaining displayed segment, and still returns the full normalized stopped response for channel callers. Completed tools continue to collapse while long runs are active, but expand again when the turn reaches a final state so the completed transcript shows the full tool timeline. * fix(stream): preserve narration around tool timelines Keep assistant narration attached to the tool call that follows it instead of folding all streamed text into the final answer block. Track narrated response segments in stream state, render them before their corresponding regular or task tool entries, and keep final answers limited to the response suffix that has not already been shown inline. Preserve narration across normal completion, stop/error final frames, sub-agent task calls, and collapsed live tool summaries. Add regression coverage for pending tools, completed tools, sub-agent task delegations, collapsed completed/running tool summaries, and final stop frames. * fix(tui): finalize inline narration transitions * test(memory): use canonical project id helper |
||
|
|
d53bfa35c5 | feat: update MiniMax model entries and context window for M3 variant | ||
|
|
fbd1d709ca |
feat: add WebUI mode support with related configuration and onboarding (#252)
* feat: add WebUI mode support with related configuration and onboarding steps * feat: enhance WebUI port configuration to prevent conflicts with backend port * feat: add support for fresh interactive session detection in WebUI |
||
|
|
3ce6523faf |
fix: update deepagents and langchain versions (#251)
* fix: update deepagents and langchain versions; enhance _reduce_messages_delta handling for None state * fix: update langchain version constraint to >=1.3 in pyproject.toml and uv.lock |
||
|
|
a13904185d |
Feat/sandbox execute timeout (#243)
* feat: implement configurable sandbox execute timeout and enhance recovery instructions * feat: add background process management tools and middleware for sandbox execution * feat: enhance background process management with completion notifications and deduplication * feat: enhance sandbox execution timeout validation and update related messages * feat: enhance background process management with thread-specific completion notifications and HITL approval handling * test: assert completion notification waits for process finish timestamp |
||
|
|
2364e6b130 |
fix(cli): forward async-notifier replies back to originating channel (#244)
* fix(cli): forward async-notifier replies back to originating channel When PR #214's auto-notifier fires a synthetic agent turn after a channel-originated conversation, the synthesized response only rendered to the local CLI/TUI — the channel user (iMessage etc.) saw nothing and had to manually re-prompt to find out what happened. Adds a per-thread channel-origin registry in cli/channel.py and wires the three notifier paths (Rich CLI / TUI / serve) to publish the final response back via bus.publish_outbound when the originating thread was started by a channel turn. Publish is fire-and-forget (scheduled on the bus loop + done-callback for failure logging) so the notifier turn doesn't block on the asyncio / textual event loop. The registry is cleared on /new and /resume rotation so stale entries don't accumulate. * fix(cli): address review feedback on channel-origin forwarding Follow-up to the review on #244 (din0s, X-iZhang): - Guard the /resume origin cleanup on a real thread change in Rich CLI and TUI (serve mode already did via thread_changed). Resuming the already-active thread no longer wipes its still-live origin, which would otherwise silently drop a later async-notifier forward — the exact gap this PR closes. - Re-bind the now-current thread to its channel after a channel-issued /new or /resume slash command (which rotates the thread inside the dispatch), so notifier turns on the rotated thread still forward. - Guard the publish done-callback against a cancelled future, whose .exception() raises CancelledError (rather than returning it) on bus-loop teardown, so the intended warning still logs. - Mirror the normal reply path's manager.record_message(channel, "sent") for forwarded notifications so per-channel stats stay accurate. - Print the closing "[channel: Replied to ...]" line in all three notifier paths (Rich CLI / TUI / serve) when a forward actually happened, so the forwarded block reads as terminated on screen. Adds test_publish_records_sent_metric. ruff clean; notification-origin suite (10) + related channel/CLI/serve suites (728) pass. * fix(cli): store sender information separately from chat_id in channel origin --------- Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com> |
||
|
|
721a03c25b | feat: update model version from claude-sonnet-4-5 to claude-sonnet-4-6 and related adjustments | ||
|
|
f75bfcda51 |
Add onboarding wizard with style and validation components (#241)
* Add onboarding wizard with style and validation components - Introduced `style.py` for shared visual elements used in the onboarding wizard. - Created `validators.py` for input validation, including integer and choice validators, and API key validation functions for various providers. - Implemented `wizard.py` as the entry point for the onboarding process, managing user prompts and configuration steps. - Added progress rendering and autosave functionality to enhance user experience during the onboarding process. * feat(onboarding): enhance validation and configuration for onboarding wizard - Added validation for UI backends, workspace modes, and providers in the onboarding command. - Updated channel definitions to include secret field handling for sensitive tokens. - Improved user prompts for required fields, ensuring sensitive data is masked. - Introduced constants for valid providers, UI backends, and workspace modes to maintain consistency. - Implemented tests to ensure alignment between constants and interactive choices in onboarding steps. * feat(onboarding): improve WeChat account ID prompt and validation for newly enabled channels * feat(onboarding): enhance WeChat backend credential prompts and validation * feat(onboarding): refine WeChat backend credential prompts for wecom and wechatmp * Refactor onboarding package for improved structure and clarity - Simplified the onboarding package by removing unnecessary re-exports and consolidating public API to only include `run_onboard`. - Updated `install_back_keys` to `install_navigation_keys` for clarity and consistency in the prompter module. - Enhanced the `NonInteractivePrompter` class to support strict mode, allowing for better handling of non-interactive prompts. - Adjusted the onboarding steps to utilize the new navigation keys installation method. - Improved the `run_onboard` function to handle section implications based on user flags, enhancing the onboarding experience. - Updated tests to reflect changes in imports and ensure compatibility with the new structure. * feat(onboarding): enhance validation logic for non-interactive prompts * refactor(onboarding): streamline onboarding module structure and enhance validation error handling * refactor(onboarding): enhance config revert logic to preserve original file state * refactor(onboarding): enhance tavily key validation and error handling in onboarding process |
||
|
|
b9ad694467 | fix: resolve path correctly when workspace name appears in parent path | ||
|
|
d2283397a4 |
feat(feishu): scan-to-create QR onboarding + silence unsubscribed WS events (#239)
* feat(feishu): scan-to-create QR onboarding flow
Add a device-code flow against accounts.feishu.cn/oauth/v1/app/registration
that lets users scan a terminal QR code with Feishu / Lark mobile to
auto-create a PersonalAgent bot app with the required IM permissions
pre-attached. The poll endpoint returns app_id + app_secret, which the
onboarding wizard then writes into the channel config — no manual app
creation on open.feishu.cn required.
- channels/feishu/onboard.py: qr_register() public entry, init/begin/poll
helpers, QR rendering via the soft qrcode dep, automatic feishu↔lark
domain switch based on the scanning user's tenant_brand, and a
best-effort bot probe to surface the bot name in the wizard
- channels/feishu/__init__.py: re-export qr_register (mirrors qq)
- config/onboard.py: offer "Scan QR code (recommended) / Enter manually"
in the Feishu branch, ask for region (feishu vs lark), then call
qr_register and populate feishu_app_id / feishu_app_secret /
feishu_domain; add qrcode>=7.4 to the feishu pip extras
* fix(feishu): silently absorb unsubscribed WebSocket events
Feishu auto-subscribes PersonalAgent apps to many event types
(im.message.reaction.created_v1, message.read_v1, message.recalled_v1,
chat.member.*, ...) that EvoScientist doesn't register handlers for.
Without intervention, lark-oapi's dispatcher raises EventException
("processor not found, type: ..."), the WS client logs it at ERROR and
replies HTTP 500 on the frame, and Feishu marks the event as failed
and retries it.
The problem is amplified by _send_ack_reaction: every inbound message
triggers our own reaction, which Feishu echoes back as
reaction.created_v1, creating a continuous ERROR-log feedback loop and
pointless retries.
Wrap EventDispatcherHandler._do_without_validation after build() to
swallow "processor not found" EventExceptions (debug log + return None)
while letting all other errors propagate. Failure-safe: if lark-oapi's
internal API changes the wrapper degrades to the prior behavior rather
than breaking the channel.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
|
||
|
|
7959495a13 |
feat(deploy): add EvoSci deploy subcommand (#228)
* feat(deploy): implement standalone LangGraph server and CLI command for deployment * feat(deploy): enhance port validation and environment variable management for deployment * Refactor langgraph dev deployment and introduce workspace sidecar protocol - Updated the deployment mode handling in `server.py` to use a single environment variable `EVOSCIENTIST_DEPLOY_MODE` with values `full` and `stripped`. - Enhanced the `manager.py` to implement a workspace fingerprint sidecar, allowing cross-process reuse of langgraph dev instances while ensuring workspace consistency. - Introduced functions to write and read the workspace sidecar, with error handling for missing or corrupt data. - Added tests for the workspace sidecar functionality, including validation of the JSON schema and ensuring proper error handling for workspace mismatches. - Updated existing tests to reflect changes in deployment mode handling and added new tests for signal handling during shutdown. - Ensured that cleanup routines remove the workspace sidecar alongside the PID file during shutdown. * fix(langgraph): improve workspace sidecar checks for process ownership and stale handles |
||
|
|
331056cdc8 |
feat(middleware): upgrade deepagents 0.5.7 → 0.6.2 (#231)
* feat(middleware): add CodeInterpreterMiddleware with project-specific configuration chore(config): increase checkpoint retention limit for runaway conversations fix(tests): update database schema references from 'blob' to 'value' chore(deps): update deepagents dependency to include quickjs support * feat(deepagents): update to version 0.6.1 and add optional dependencies for quickjs * feat(sessions): improve error handling for message deltas and update Overwrite type check * Enhance PruningCheckpointer with DeltaChannel Awareness - Introduced a new pruning strategy in `_prune_after_put` to preserve the `_DeltaSnapshot` chain during checkpoint pruning. - Implemented methods to fetch recent checkpoint IDs and walk to snapshot ancestors, ensuring that necessary checkpoints are retained. - Updated SQL queries to handle checkpoint and write deletions more efficiently. - Added comprehensive tests for DeltaChannel-aware pruning, ensuring that the pruning logic correctly handles various checkpoint scenarios, including those with and without snapshot seeds. - Refactored `_load_checkpoint_messages` to utilize the new saver interface, improving message reconstruction from checkpoints. * feat(tests): add migration sweep test to preserve snapshot ancestor * feat(sessions): enhance checkpoint retrieval to prevent transcript leakage in multi-agent scenarios * feat(middleware): enhance CodeInterpreterMiddleware with configurable timeout and result character limit feat(config): add CodeInterpreterMiddleware tuning parameters to EvoScientistConfig feat(sessions): implement inline message delta reducer for improved message handling * feat(dependencies): update deepagents version to 0.6.2 in pyproject.toml and uv.lock |
||
|
|
385f9756c1 |
feat(backends): implement tier-aware virtual mount resolution for ski… (#236)
* feat(backends): implement tier-aware virtual mount resolution for skills and memories * test: add end-to-end test for workspace tier shadowing global tier in CustomSandboxBackend * feat(backends): enhance virtual mount resolution for skills and memories with tier paths and quoting * fix(tests): update Python command in virtual mount resolution tests to use python3 |
||
|
|
7f1aa3b0f6 |
fix(cli): handle spaces in @file mentions (#234)
* fix(cli): handle spaces in @file mentions The @file parser truncated at the first space, so dragging or pasting a filename like `@PREPING_ Building Agent.pdf` only matched `@PREPING_` and warned "file not found". Now supports `@"..."` / `@'...'` quoted form for explicit paths, plus a greedy expansion fallback that walks across whitespace until an existing file resolves (bounded by newlines, the next `@`, and a 20-token cap). Autocomplete also returns quoted mentions for any candidate containing a space. * style: apply ruff format to file_mentions |
||
|
|
a4c9c779c9 |
feat: status and elapsed time indicator (#218)
* feat: status and elapsed time indicator * test: add tests for tui-status * fix: move to enum+switch, change phase calculation * feat: remove 'done' phase |
||
|
|
4b0c91190a |
feat(llm): add dashscope-code provider for Alibaba Coding Plan keys (#225)
* feat(llm): add dashscope-code provider for Alibaba Coding Plan keys Alibaba Cloud Bailian "Coding Plan" subscription keys (sk-sp-*) route through a separate endpoint (coding.dashscope.aliyuncs.com/v1) that the standard `dashscope` provider can't reach. Add a sibling provider entry matching the zhipu/zhipu-code and moonshot/kimi-coding precedents, with its own validator (the coding endpoint returns 404 on /models, so probe via chat.completions instead). Closes #224 * fix(llm): keep dashscope as default provider for qwen3-coder shortcut The MODELS dict is built from _MODEL_ENTRIES via a last-write-wins dict comprehension. The initial commit listed dashscope-code AFTER dashscope, which silently flipped the bare `get_chat_model("qwen3-coder")` shortcut to the coding endpoint — breaking standard sk-* keys. Reorder to match the zhipu-code / zhipu precedent: coding endpoint first, general endpoint last so the general endpoint wins the collision and remains the default for the shared "qwen3-coder" short name. |
||
|
|
8fe774b056 |
Feat/qq interactive buttons (#220)
* feat(qq): add inline keyboard buttons for C2C HITL approval
QQ Bot supports inline buttons via `markdown + keyboard` payloads. Clicks
arrive as `interaction_create` events through the existing botpy
WebSocket gateway — no extra subscription needed beyond enabling the
`interaction` intent. Group-scope clicks are out of scope here (DM only).
Send path
- `_build_qq_keyboard(buttons)` mirrors the Feishu helper, mapping the
generic `{text, value, type}` shape to QQ's `{render_data, action}`
with action.type=1 (callback). One button per row for mobile clarity.
- `_send_chunk` extracts `metadata["buttons"]` and threads a `keyboard`
payload into `_post_markdown_message` for C2C only.
- Markdown→plain fallback can't carry a keyboard, so when buttons were
attached the fallback content gets a textual `Reply: 1=Approve, …`
hint built from the button list. `_parse_approval_reply` accepts
the same values typed manually, so the user is never stuck.
Receive path
- `on_interaction_create` is registered on the bot class.
- `_on_interaction` extracts `data.resolved.button_data`, builds an
InboundMessage, runs it through inbound middleware (Dedup suppresses
retry callbacks), and publishes directly to the bus — bypassing the
per-sender debounce buffer so the click value isn't merged with any
text typed in the same window.
- Always ACKs via `api.on_interaction_result(id, 0)` in `finally` so
QQ doesn't show the button as "expired", even if middleware drops
the click or something throws downstream.
`QQ.inline_buttons=True`; `_approval_prompt_metadata` now auto-attaches
the Approve/Reject/Approve-all button row for QQ HITL prompts.
* fix(qq): button-value coercion, ACK timing, HITL consumer wiring
Fixes 6 bugs found in the inline-keyboard commit and consolidates the
button helpers so the keyboard builder, plain-text fallback hint, and
interaction handler share one coercion path.
- Plain-text fallback no longer crashes on non-string `value` (e.g.
`{"text": "OK", "value": 42}`). Extracted `_normalize_button` is now
the single place that resolves `(label, value)` and coerces non-strings.
- `metadata["button_value"]` is the coerced string instead of the raw
payload, matching `content` and downstream string comparisons.
- `_on_interaction` ACKs first, before publishing to the bus, so the
QQ button UI never shows "expired" if middleware is slow.
- Wire `_approval_prompt_metadata` + `_format_approval_prompt(with_buttons=)`
into `InboundConsumer._stream_with_hitl` and `cli.channel.channel_hitl_prompt`
so the QQ `inline_buttons=True` capability is actually used end-to-end
(HITL prompts auto-attach Approve/Reject/Approve-all buttons when the
channel advertises the capability).
- Trim contradictory `_QQ_DEFAULT_PERMISSION` comment.
- Fix `test_group_interaction_ignored` docstring (ACK runs first now,
not in `finally` after a `return`).
Tests: `_normalize_button` covered indirectly via existing keyboard tests;
new regressions for non-string fallback hint, ACK-on-handler-throw, and
string-coerced `button_value` metadata.
* refactor(qq): slim button helpers and explicit has_buttons flag
Inline single-use _button_hint and the _QQ_BUTTON_STYLE/_QQ_DEFAULT_PERMISSION
constants in qq/channel.py; tighten _on_interaction (drop unreachable
"[button click]" sentinel and unused triggering_message_id metadata; collapse
"if resolved else" ternaries via `or ""`).
Replace the metadata round-trip ("buttons" in metadata) used to detect button
support in consumer.py and cli/channel.py with an explicit has_buttons bool
threaded through both the prompt formatter and metadata builder.
Apply ruff format to the previously unformatted blocks introduced earlier on
this branch so CI lint passes.
* feat(qq): send post-decision confirmation after HITL approval
Send a visible confirmation message ("✅ 已批准" / "❌ 已拒绝") right after
the user resolves a HITL approval — QQ Bot has no message-recall or edit API
for C2C, so a follow-up message is the only way to give the click/reply
strong feedback.
Bus consumer (consumer.py): only sends the confirmation when the user
actually responded (event was set), to avoid pretending the user approved
when the request really timed out and auto-approved.
CLI HITL prompt (cli/channel.py): mirrors the same set of confirmation
strings. Timeout and unrecognized-reply paths keep their existing English
text since their semantics differ (auto-reject vs auto-approve, plus a
hint about the unparsed input).
|
||
|
|
c407d2e20f |
Fix/async subagent model switch (#217)
* feat(middleware): add ConfigurableModelMiddleware for dynamic model resolution - Introduced ConfigurableModelMiddleware to resolve chat models from RunnableConfig.configurable on each call. - Updated middleware initialization to include ConfigurableModelMiddleware. - Enhanced context editing middleware tests to verify presence of ConfigurableModelMiddleware. - Implemented tests for ConfigurableModelMiddleware to ensure correct model overriding and caching behavior. - Added tests for deepagents model-passthrough patch to verify configuration injection in async tasks. * feat(async-subagent): update middleware handling to prevent deadlocks in async sub-agents * style: Refactor code formatting for improved readability in patches and test files * refactor: streamline middleware construction and improve async handling in ConfigurableModelMiddleware * fix: remove unused request parameter from _read_model_override function * refactor: improve async handling in _ClientProxy and enhance logging in ConfigurableModelMiddleware test: add behavior test to ensure AskUserMiddleware is excluded in async subagent mode |
||
|
|
89b0ecdbf3 |
feat(qq): add QR-code scan-to-configure onboarding for QQ Bot (#213)
* feat(qq): add QR-code scan-to-configure onboarding for QQ Bot Adds a `qr_register()` flow that drives q.qq.com's create_bind_task / poll_bind_result APIs so the wizard can auto-fill `qq_app_id` and `qq_app_secret` after the developer scans a QR code with a bound QQ account, falling back to manual entry on failure or cancel. - channels/qq/crypto.py: AES-256-GCM helpers for decrypting the bot's client_secret returned by poll_bind_result. - channels/qq/onboard.py: portal API client + polling loop. - channels/qq/__init__.py: re-export `qr_register`. - config/onboard.py: QQ branch in `_step_channels` that offers "Scan QR code" vs "Enter manually", and skips the manual prompt loop when a scan succeeded. * style(qq): fix ruff lint errors in onboard.py Move `import os` to the top-level import block (E402), drop the legacy `typing.Optional`/`typing.Tuple` imports (UP035), and use the PEP 585/604 builtin generics (`tuple[...]`, `X | None`) for the few annotations that still referenced them (UP006/UP045). No behavior change. * fix(qq): harden QR onboard error paths and declare scan deps Address review feedback on PR #213: - Declare cryptography>=41.0 and qrcode>=7.4 in [qq]/[all-channels] extras and in _CHANNEL_PIP_DEPS so the scan flow no longer fails with an opaque ImportError on a fresh `evoscientist[qq]` install. - Polling loop logs each _poll_bind_result failure and aborts after 5 consecutive errors instead of silently spinning until the 600s timeout, restoring the documented Raises: RuntimeError contract. - Wrap decrypt_secret in try/except so failures honor the None-on-failure contract instead of letting exceptions escape. - Preflight `import cryptography` in the scan branch and offer install or fall back to manual entry. * style: ruff format collapse two over-wrapped log/console lines |
||
|
|
4e04ac5b72 |
fix: Improve watcher logic to prevent false-positive notifications on… (#216)
* fix: Improve watcher logic to prevent false-positive notifications on clean stream exits * fix: Update watcher logic to drop notifications on persistent runs.get failures * fix: Refactor test for watcher persistent failure notification handling * fix: Enhance watcher test to validate all notification queues are empty after reconnect budget exhaustion |
||
|
|
80f1f4fa0f |
feat: Implement async sub-agent auto-notification system (#214)
* feat: Implement async sub-agent auto-notification system - Added async notifier functionality to handle notifications for sub-agents reaching terminal states. - Introduced `AsyncTaskNotification` dataclass for structured notification data. - Implemented `watch_run_and_notify` to monitor agent runs and enqueue notifications. - Created `spawn_watcher` to manage watcher tasks and ensure proper cancellation of previous watchers. - Developed `consume_notifications` to process notifications, deduplicate them, and format messages for LLM. - Added tests for notification handling, including draining, deduplication, and formatting. - Patched deepagents to integrate the new watcher functionality into start and update tools. * Enhance async notifier with per-thread notification routing and error handling - Introduced `origin_cli_thread_id` to `AsyncTaskNotification` for routing notifications back to the originating CLI session. - Implemented per-thread notification queues to handle notifications based on the originating thread. - Updated `has_pending_notifications` and `drain_notifications` to respect thread-specific queues. - Enhanced `watch_run_and_notify` to detect in-band error events from the SSE stream and handle clean exits. - Modified tests to verify the new notification routing behavior and ensure proper handling of notifications across threads. - Added a fixture to restore the async watcher patch state in tests to prevent state leakage. - Updated deepagents patching to capture the main agent's CLI thread ID for notification routing. * feat: Enhance async notifier with thread-specific watcher management and notification filtering * test: Enhance notification draining logic for cleaner test setup * refactor: Remove summary field from AsyncTaskNotification and update related tests * feat: Enhance async notification handling with target thread ID support * Refactor async notifier and middleware for improved task management - Removed the no-op shutdown watcher loop from async_notifier.py as it is no longer needed. - Updated watch_run_and_notify to clarify notification handling and race conditions. - Cleaned up shutdown handling in commands.py, interactive.py, and tui_interactive.py by removing obsolete shutdown watcher calls. - Deleted the deepagents async watcher patch from patches.py, transitioning to a new middleware approach. - Introduced AsyncWatcherMiddleware to handle async task notifications directly during tool calls. - Updated tests to validate the new middleware functionality and ensure proper watcher spawning and cancellation. - Enhanced test coverage for async watcher middleware, including edge cases and error handling. * feat(tests): add fixture to reset notifier state before each test |
||
|
|
692dc491ac |
# feat(wechat): add personal-WeChat (iLink) backend with QR login (#212)
* feat(wechat): add personal-WeChat (iLink) backend with QR-code login Adds a third WeChat backend alongside WeCom and Official Account: ``personal`` rides Tencent's iLink Bot long-poll gateway so a personal WeChat account can act as a bot. Credentials are obtained via QR-code scan and persisted under ``DATA_DIR/wechat_personal/accounts/``. - channels/wechat/personal.py: WeixinPersonalChannel + qr_login. - channels/wechat/crypto.py: aes128_ecb_decrypt + parse_ilink_aes_key for the iLink CDN media protocol. - channels/wechat/probe.py: validate_wechat_personal credential probe. - channels/wechat/serve.py: --backend personal CLI + --qr-login flow. - channels/wechat/__init__.py: factory dispatch on wechat_backend; pull in the new dependencies in the docstring. - config/settings.py: wechat_personal_* fields. - config/onboard.py: WeChat-backend picker + QR-scan flow in the wizard + personal-backend probe in _probe_channel. - pyproject.toml / uv.lock: add qrcode + certifi to wechat & all-channels extras (aiohttp was already pulled in transitively). * fix(wechat): address ruff failures and CodeRabbit review on personal-WeChat PR - personal.py: drop unused imports (`field`, `PollingMixin`); replace `asyncio.TimeoutError` with builtin; hold references to background `asyncio.create_task` results so they aren't GC'd; wire `dm_policy` through `_process_message` (disabled/allowlist) so `wechat_personal_dm_policy` actually takes effect for DMs. - onboard.py: import-check gate now validates the full WeChat dependency set (aiohttp, qrcode, Crypto, certifi) instead of only aiohttp; mask `WeCom Secret` and `MP App Secret` prompts via `questionary.password`; derive the QR-login hint path from `_account_dir()` instead of the hard-coded `~/.evoscientist/...`; stop copying the QR-login token into the main config (already persisted per-account on disk — copying broadens secret exposure and risks staleness). - pyproject.toml: allow Chinese full-width punctuation in `allowed-confusables` for user-facing CN messages. * style(wechat): apply ruff format `ruff format --check` was failing CI on three files (one pre-existing in `__init__.py` plus formatter-driven line-merges in the files touched by the previous fix commit). Ran `ruff format` to bring them in line; both `ruff check` and `ruff format --check` now pass. |
||
|
|
9e51ec6fdd |
feat(cmd): add /model-fallback command (#196)
* feat(cmd): add /model-fallback command * fix: apply feedback * fix: lock usage with _fallback_chain * fix: apply feedback * fix: apply feedback * feat: add tests * fix: tests * Update EvoScientist/middleware/model_fallback.py Co-authored-by: dinos <dinospk1999@gmail.com> --------- Co-authored-by: dinos <dinospk1999@gmail.com> |
||
|
|
22a65b640d |
refactor(channels): remove dead MessageBus dispatcher (#205)
* refactor(channels): remove dead MessageBus dispatcher Outbound routing has two implementations: ``MessageBus.dispatch_outbound`` (subscriber-based) and ``ChannelManager._dispatch_outbound`` (registry lookup). Only the latter is ever started in production — the former is reachable solely from tests, yet both consume from the same ``bus.outbound`` queue. If anyone followed the bus's own API surface they would silently steal messages from the real dispatcher. Drop the unused machinery to leave a single, obvious outbound path: - ``MessageBus.subscribe_outbound`` / ``dispatch_outbound`` / ``stop`` - ``_running`` flag and ``_outbound_subscribers`` map - ``OutboundCallback`` type alias - The lone ``bus.stop()`` call in ``cli/channel.py`` (was no-op) - Four tests covering the removed code paths * test(channels): drop empty MessageBus stubs after dispatcher removal --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
d7c0eec0e9 | chore: update dependencies and remove unused OpenRouter patch (#211) | ||
|
|
f41584e10b |
Refactor sub-agent architecture and introduce async support (#200)
* Refactor sub-agent architecture and introduce async support - Removed the legacy subagent.yaml file and replaced it with individual YAML files for each sub-agent in the subagents directory. - Updated the load_subagents function to support both directory and single file layouts for loading sub-agent configurations. - Added new langgraph_dev module for managing async sub-agent lifecycle and deployment. - Created graphs for async sub-agents (writing-agent, data-analysis-agent) and updated langgraph.json for deployment. - Introduced new sub-agent definitions for planner, research, debug, code, and writing agents with appropriate system prompts and configurations. - Enhanced package data inclusion in pyproject.toml to accommodate new sub-agent YAML files. * Refactor code for improved readability by consolidating conditional statements and formatting * feat: enhance async sub-agent support with workspace synchronization and user feedback - Added console status messages during async sub-agent server startup and workspace synchronization to improve user experience. - Implemented a new WorkspaceSyncWidget for live feedback during workspace sync operations. - Updated onboarding to reject occupied ports and ensure proper workspace handling for async sub-agents. - Introduced locking mechanisms to manage concurrent access to langgraph dev processes and workspace states. * feat: add async sub-agent configuration and server management functions * feat: improve port occupation handling and log file management in start_langgraph_dev * feat: enhance async sub-agent handling and introduce comprehensive tests - Updated `_maybe_swap_async_subagents` to improve async sub-agent management, ensuring internal flags are stripped before handoff. - Enhanced port management in `onboard.py` to allow reuse of occupied ports if already running by the same service. - Introduced file locking in `manager.py` to prevent race conditions during concurrent CLI invocations. - Added new tests for async sub-agent swapping and langgraph manager functionalities to ensure reliability and correctness. - Updated dependencies in `pyproject.toml` to include `psutil` and `filelock`. * fix(docs): clarify sub-agent configuration in README * test(manager): isolate _PID_DIR + tighten reuse-path assertion Addresses CodeRabbit review on tests/test_langgraph_manager.py: - Patch _PID_DIR to tmp_path so the FileLock setup in ensure_langgraph_dev doesn't mkdir the user's real ~/.config/evoscientist/ dir as a test side-effect. - Tighten "result is None or hasattr(result, 'poll')" to a strict "result is None" — the reuse path returns None unconditionally, so the OR clause was hiding potential regressions. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(manager): clean up stale PID file when unrelated process reuses PID * feat(tests): add validation tests for async flag in load_subagents * fix(load_subagents): restrict to .yaml files and clarify configuration handling * fix(load_subagents): improve error handling for non-dict specifications in YAML * feat(onboard): add "LangGraph Port" step to onboarding process * feat(langgraph): add concurrency configuration for langgraph dev workers * feat(async-subagents): enhance MCP tool routing for async sub-agents * fix(manager): update exception handling for connection errors and prevent zombie processes --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
ab3787f86a |
fix(markdown): ensure proper spacing for ATX headings in Markdown ren… (#201)
* fix(markdown): ensure proper spacing for ATX headings in Markdown rendering * fix(markdown): improve docstrings and tests for heading spacing functionality |
||
|
|
29e7fd383d |
fix/cli hitl (#202)
* feat(display): enhance approval prompt with questionary for better navigation * feat(cli): add HumanInTheLoopMiddleware for user approval in main agent |
||
|
|
73928a2d78 |
fix(onboard): detect installed skill packs via install manifest (#199)
* fix(onboard): detect installed skill packs via install manifest
Onboarding's _step_skills only inspected USER_SKILLS_DIR and matched
recommended entries by directory-name hint, so a pack like
EvoScientist/EvoSkills@skills (which explodes into paper-writing/,
evo-memory/, etc. under GLOBAL_SKILLS_DIR) was never detected and kept
appearing as not-yet-installed.
skills_manager now writes a per-tier .installed.yaml mapping skill
directory name -> original install source on every install, removes the
entry on uninstall, and exposes installed_sources(). _step_skills checks
both tiers and treats a recommended source as installed when present in
any manifest -- so packs are recognized regardless of how their child
dirs are named.
* fix(onboard): write install manifest atomically
Stage to a sibling temp file, fsync, then os.replace into place. A crash
mid-write can no longer leave a half-written .installed.yaml behind,
which would otherwise wipe out pack detection until the next reinstall.
* style: fmt
* fix(onboard): catch decode errors when loading install manifest
read_text() can raise UnicodeDecodeError on a hand-edited or corrupt .installed.yaml; pin encoding="utf-8" and add UnicodeError to the except clause so the function honors its "returns {} on any error" contract.
|