* refactor(onboard): shared flow for ccproxy providers
* feat(onboard): support oauth configuration for auxiliary models
* fix(onboard): reuse main model auth for same-provider auxiliary
* fix(onboard): reconcile oauth providers
* feat(cli): add --output-format stream-json for headless clients
Emit EvoScientist's native event stream as line-delimited JSON on stdout
in single-shot (-p) mode, with all human output redirected to stderr so
stdout stays pure JSONL. Intended as the integration surface for
programmatic clients (e.g. an agent runtime) that drive EvoSci headlessly.
- stream/json_sink.py: write_events_as_json + stream_json sink, plus
redirect_console_to_stderr helper for stdout purity
- cli/interactive.py: cmd_run gains output_format; stream-json branch runs
the sink instead of the Rich renderer
- cli/commands.py: --output-format option + validation (stream-json
requires -p; value must be text|stream-json)
- docs/stream-json.md: event-schema contract + example transcript
- tests: json sink serialization, CLI dispatch, console redirect, validation
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(cli): honor explicit --no-auto-mode over config in stream-json
Address CodeRabbit review (discussion_r3514041123): the auto-mode override
block only wrote to cli_overrides when the resolved value was True, so an
explicit --no-auto-mode silently fell back to a config that enables
auto-mode -- breaking "explicit flags always win" and leaving stream-json
running unattended despite the warning. Write auto_mode=False when the flag
is explicitly False. Add regression tests that capture the overrides.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(tui): keep welcome banner at top after /new
PR #262 replaced scroll_end() with anchor() for free-scrolling.
When /new clears a long anchored conversation, the anchor kept
the viewport pinned to the (now empty) bottom, producing a
negative scroll_y and pushing the welcome banner out of view.
Reset the anchor and scroll to the top in clear_chat(), and
restore the follow/new-content flags so the fresh session starts
correctly.
Closes#301
* fix(tui): suppress anchor when chat content fits viewport
The previous fix for #301 only handled the /new path. din0s reported
that the banner still dropped to the bottom after a normal short turn
(user types 'hi', agent replies) — i.e. whenever the conversation
fit in the viewport. Root cause is in Textual's compositor
(textual._compositor): when a widget is anchored, scroll_y is
recomputed via set_reactive, which bypasses the validator. If the
anchored widget's content is shorter than the viewport, scroll_y
goes negative on the next layout pass and the welcome banner is
pushed below the visible region.
PR #262 made _stream_with_widgets re-engage the anchor at the end of
every turn via _anchor_chat, so the bug surfaced on any short reply
that fit in the viewport. Markdown re-renders, status-bar updates,
or any subsequent mount would then trip the compositor.
Fix in three places:
* _anchor_chat: only engage the anchor when max_scroll_y > 0;
otherwise release and scroll_home so the banner stays at the top.
* streaming anchor loop: if content shrinks below the viewport
mid-stream (e.g. loading widget removed), release the anchor
instead of leaving _anchored=True for the compositor to trip on.
* clear_chat: keep the unconditional reset (children are removed
asynchronously so a max_scroll_y check would be stale) but
document why.
Adds two regressions:
* test_short_turn_keeps_banner_at_top_after_layout_refresh — the
exact scenario din0s tested; fails with scroll_y=-10 on the
previous code, passes with the fix.
* test_long_turn_keeps_viewport_pinned_to_bottom — guards against
regressing free-scrolling for overflowing conversations.
Manually verified: 'hi' -> reply (banner stays at top) -> /new
(banner at top) -> another turn (banner stays at top).
* test(tui): address review feedback on banner-position regressions
- extract `_release_anchor_and_pin_top` helper for the 3-line
`anchor(False) + scroll_home(...)` pattern repeated in
`clear_chat`, `_anchor_chat`, and the streaming loop
- replace `pytest.skip` in `_capture_app` with a hard `RuntimeError`
so a broken capture never silently passes
- drop the redundant `load_agent` and `create_session_workspace`
monkeypatches (the factory is given those as parameters, so the
module-level symbols never run; added a comment explaining why)
- add a defensive `_FakeChannelRuntime` patch for symmetry with the
other module-level fakes
- drop the local `_run` helper and use the `run_async` fixture from
`conftest.py` (its teardown is better)
* test(tui): replace _FakeChannelRuntime with _auto_start_channel no-op
The _FakeChannelRuntime patch was ineffective because ChannelRuntime is
just a dataclass — the real channel manager still started via
_auto_start_channel, leaving pending tasks and non-hermetic test state.
Per review feedback, stub _auto_start_channel directly instead.
* feat: expose model registry at GET /api/models
* fix: include Ollama models in /api/models endpoint
* fix: honor env vars override in /api/models endpoint
* fix: offload get_effective_config to thread to satisfy blockbuster
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat: add scheduler functionality with cron-style task management
- Implemented a new scheduler subagent to automate recurring tasks using cron expressions.
- Enhanced the subagent factory to include the skill manager and auxiliary chat model for the scheduler.
- Created a YAML configuration for the scheduler with a detailed system prompt and toolset.
- Updated README files to include documentation on scheduled tasks and usage examples.
- Added tests for the scheduler, including command execution, scheduling tools, and middleware integration.
- Introduced new dependencies for timezone handling and ensured compatibility in the project configuration.
* fix(async-notifier): ensure fallback hint is used for unknown notification kinds
* feat: enhance scheduling functionality and improve system message handling
- Ensure 'task' is excluded from the default PTC allowlist to prevent ValueError in langchain-quickjs >=0.3.
- Verify that essential async dispatch tools remain in the allowlist.
- Confirm that the live quickjs filter accepts the default allowlist even with a 'task' tool present.
- Test the creation of the code_interpreter middleware to ensure it builds correctly.
* feat(gateway): graph gateway protocol
* refactor(cli): wire gateway in cli/tui
* refactor(gateway): centralize runtime gateway init
* chore(gateway): restrict RunRequest message type
* feat(gateway): add langgraph server gateway
* chore(cli): tighten serve runtime state typing
* refactor(cli): route async task state reads through graph gateway
* refactor(gateway): support graph targets in server gateway
* refactor(cli): route session commands through graph gateway
* refactor(cli): fold thread store under graph gateway
* refactor(gateway): route graph state access through gateway
* refactor(channels): wire graph gateway
* refactor(memory): preserve graph threads for cloning
* feat(gateway): add thread cloning
* fix(tui): pass effective workspace for thread creation
* chore(memory): add workspare dir to memory worker metadata
* fix(sessions): filter preloaded UUID registy entries by the current scope
* test(fakes): use https
* refactor(consumer): consolidate imports
* fix(stream): optional summarization event
* fix(gateway): resolve abbreviated thread IDs by search
* fix(gateway): page server thread listings
* fix(gateway): emit pending interrupt events
* style: fmt
* feat(gateway): persist workspace_dir & model in thread metadata
* fix(gateway): page server thread prefix resolution
* fix(gateway): expose server thread list metadata
* refactor: add back type def
* refactor: tighten types
* revert: add back worker thread deletion
The worker thread forking changes are out of scope for now, so to
maintain parity with the existing behavior we'll leave this intact.
* fix(gateway): apply compaction to server thread history
* refactor(stream): restore direct summary replay suppression
* fix(gateway): preserve compaction state and server stream output
* fix(gateway): close local stream generator on cancellation
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Typing `/model` and pressing Enter did nothing in the TUI; the picker
only opened via `/model --save` or `/model <name>`. The completion popup
matched both `/model` and `/model-fallback` by prefix, so the
exact-match-hide guard (which required a single match) never fired. With
the popup still visible, the TUI's Enter handler completed the text
instead of submitting the command, so it never executed.
Treat the typed prefix as an exact match whenever it equals any matched
command name, not only when it is the sole match. This hides the popup
on a complete command name so Enter submits it, even when a longer
command shares the prefix.
* feat(backends): use _platform_quote for Windows cmd.exe compatibility
Resolves the 3 skipped E2E tests in test_backends.py that exercised
the /skills/... mount path. The path-rewriter was wrapping resolved
absolute paths via shlex.quote (POSIX single-quote style); cmd.exe
doesn't strip single quotes, so the literal ' characters ended up in
the subprocess argv and the python script failed to find its file.
Replace the 3 shlex.quote call sites in _resolve_virtual_mount_path
with _platform_quote, a thin platform dispatcher:
- POSIX: shlex.quote (unchanged)
- Windows: _cmd_quote uses cmd.exe-compatible double-quote wrapping
and properly escapes embedded " and percent signs
Adds:
- backends.py: _is_windows, _cmd_quote, _platform_quote (~40 lines)
- test_backends.py: 6 TestPlatformQuote unit tests + _split_cmd
cross-platform tokenizer helper to replace shlex.split in the 8
sites that tokenize convert_virtual_paths_in_command results
(POSIX shlex strips backslashes from bare Windows paths, which
broke the 5 TestVirtualMountResolution assertions on Windows)
Removes:
- 3 @pytest.mark.skipif(sys.platform == "win32") markers on the
E2E tests for /skills/... mount resolution
Refs #274.
* fix: escape % as %% in _cmd_quote instead of relying on double-quoting
cmd.exe expands %VAR% before processing quotes, so double-quoting
cannot neutralize percent signs. Escape bare % as %% (the cmd.exe
idiom for a literal percent) before any other quoting logic.
Also updates _cmd_quote docstring and _resolve_virtual_mount_path
docstring to reflect the actual quoting strategy.
* style: fix ruff format (single → double quotes)
* fix: treat % as regular char in _cmd_quote, document limitation
%% escaping only collapses in .bat/.cmd files, not via cmd /c.
Since virtual-mount paths should never contain % in practice,
simpler to leave % alone and document the caveat.
* feat(cli): multi-stage slash command completions with subcommand awareness
Phase 1 of #82 — subcommand and argument awareness in completions.
- commands/base.py: add SubCommand dataclass and subcommands/category
ClassVars to the Command ABC. Each SubCommand has name, description,
and optional arguments.
- commands/manager.py: add get_subcommands() and list_subcommands()
methods to expose subcommand metadata for completion rendering.
- commands/implementation/mcp.py: declare 6 subcommands (list, config,
add, edit, remove, install).
- commands/implementation/model_fallback.py: declare 6 subcommands
(list, add, remove, clear, save, help).
- commands/implementation/channel.py: declare 2 subcommands
(status, stop).
- cli/tui_interactive.py: rewrite on_text_area_changed slash-completion
branch. When the user types a command name + trailing space and the
command has subcommands, show subcommand completions instead of hiding
the popup. Filter subcommands by typed prefix in multi-token input.
- commands/implementation/general.py: /help now lists subcommands
below each command that declares them.
Tests: 8 new tests covering SubCommand creation, CommandManager
subcommand lookup, and cross-command verification.
2293 passed baseline, no regressions.
* fix: subcommand completion preserves prefix + prompt_toolkit + tests
- _apply_selected_completion: preserve '/mcp ' prefix when completing
subcommands via _comp_is_subcommand flag
- SlashCommandCompleter (Rich CLI): add subcommand completion support
- Fix trailing-space bug: rstrip prefix before top-level matching
- test_tui_widgets.py: update stub on_input_changed to match multi-stage
logic; add 5 new subcommand tests
- test_command_manager.py: 8 tests for SubCommand + CommandManager
28 passed, 0 failed.
* style: ruff format tui_interactive.py + test_tui_widgets.py
* style: fix RUF012 ClassVar annotation on subcommands lists
* fix: sync test stub, add len>=3 guard, remove exact-match hide
- Sync test stub on_input_changed with real TUI code (remove exact-match
hide for subcommands, add len(parts)>=3 guard)
- Update test_input_changed_exact_subcommand_hides -> shows_confirmation
- Add test_input_changed_three_parts_hides
- Remove unused category ClassVar (din0s: what is this for)
* refactor(commands): extract shared completion engine
Per din0s feedback: one shared completion engine (commands/_completion_engine.py)
that parses text + cursor once, returns structured CompletionCandidate objects
with replace_start/replace_end ranges.
- SlashCommandCompleter (Rich CLI): thin adapter, delegates to engine
- on_text_area_changed (TUI): thin adapter, delegates to engine
- _apply_selected_completion: uses candidate.replace_start/replace_end instead
of _comp_is_subcommand flag
- Tests: engine tested directly (10 new tests), stub methods updated
29 passed, 0 failed.
* style: ruff format
* fix: preserve text after cursor when applying completion
CodeRabbit: replace_start only cuts from start to cursor,
dropping any suffix after the cursor. Use replace_start + replace_end
to correctly splice the replacement while preserving trailing text.
* fix(cli): repair slash-command completion (TUI crash, subcommand bugs, sort)
Apology + context: the previous push shipped a TUI-breaking change
(the new shared engine assumed ``event.text_area.cursor_position``
existed, but ``ChatTextArea`` / ``Changed`` don't expose it). User
caught the crash on ``/``; fixing that surfaced two more bugs in
the engine that din0s had already flagged. This commit addresses
all of them and drops a piece of dead stub code.
## Bug fixes
1. **TUI crash on ``/``** (``tui_interactive.py:2335``)
``event.cursor_position`` doesn't exist on the ``Changed`` event,
and ``ChatTextArea`` (Textual ``TextArea`` subclass) doesn't expose
``cursor_position`` either. Pass ``len(event.text_area.text)``
instead — the user types at the end of the input in practice.
2. **Subcommand trailing-space duplication** (``_completion_engine.py``)
Typing ``/mcp a `` + Tab produced ``/mcp aadd``. The engine
included the trailing space in ``replace_end``; the TUI apply
unconditionally appended ``" "``, producing double-space output.
Fix: ``replace_end`` excludes the trailing space; the TUI apply
checks ``current[replace_end:].startswith(" ")`` and skips the
separator when the suffix already has one.
3. **Subcommand exact-match confirmation noise** (``_completion_engine.py``)
Typing ``/mcp list`` + Tab re-inserted ``list`` and the popup
kept showing the same subcommand. Add a guard mirroring the
top-level exact-match rule: when the only subcommand match is
the prefix itself (no trailing space), return ``empty``.
4. **Alphabetical sort dropped in CLI** (``cli/interactive.py``)
The new completer iterated ``result.candidates`` in registration
order. Re-add ``sorted(result.candidates, key=lambda c: c.text)``.
Same sort added to the TUI for consistency.
## Cleanup
- Drop the dead ``on_input_changed`` method from the ``_StubApp``
test stub (0 call sites) plus the unused ``_slash_commands`` /
``_subcommands`` locals that fed it. This addresses din0s's
comment about the stub duplicating real TUI logic — the inlined
copy is no longer needed since the real completer now routes
through the shared engine.
## Tests
- ``test_engine_exact_subcommand_shows_confirmation`` → renamed to
``test_engine_exact_subcommand_hides`` to match new behavior.
- New: ``test_engine_subcommand_trailing_space_excludes_space_from_range``
and ``test_engine_subcommand_trailing_space_apply_does_not_double_space``.
- All 97 tests in ``test_tui_widgets.py`` pass.
- ``ruff check`` / ``ruff format`` clean.
- Local TUI smoke: ``/`` (no crash, top-level popup), ``/mcp ``
(subcommand popup), ``/mcp a `` + Tab → ``/mcp add ``.
Refs the din0s review comments on PR #273. CLI path tests and the
``category`` ClassVar follow-up are deferred to a separate PR (the
former is a test-suite addition; the latter is already absent from
``base.py`` on the current branch).
* fix: address remaining review items (help duplication, CLI tests, stub sync, docstrings)
- mcp.py: auto-generate help text from subcommands ClassVar (#1)
- tests/test_cli_completion.py: add 9 CLI completer tests (#2c)
- test_tui_widgets.py: sync _apply_selected_completion stub with real code (#4)
- mcp.py + interactive.py: add docstrings to key functions (#8)
* fix: hide completions on exact subcommand match regardless of trailing space
Remove the
ot has_trailing_space guard from the exact-subcommand
check. Previously /mcp list (with trailing space) would still
return candidates, causing Tab to oscillate between adding and removing
the trailing whitespace. Now the engine hides whenever the subcommand
is an exact match, same as the top-level rule.
Added test_engine_exact_subcommand_with_trailing_space_hides to cover
the scenario din0s flagged.
* refactor: use StrEnum for CompletionResult.kind
Replace plain str with CompletionKind(StrEnum) for type safety.
Backward-compatible with existing string comparisons.
* fix: normalize @file completion tuples to CompletionCandidate
complete_file_mention() returns list[tuple[str, str]] but the TUI
rendering/apply code expects objects with .text/.description.
Wrap tuples in CompletionCandidate to prevent AttributeError crash.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix(backends): rewrite quoted virtual paths containing whitespace
The `convert_virtual_paths_in_command` regex
`(?<=\s)/[^\s;|&<>'"`]*` stopped at the first whitespace or quote,
so:
- `python "/skills/my skill/main.py"` was left completely
unchanged (the `(?<=\s)` lookbehind failed after the opening
`"`), and the shell then broke the inner unquoted path at the
embedded space.
- `python /skills/my skill/main.py` was truncated to
`python ./skills/my skill/main.py` (only `/skills/my` rewritten).
Replace the regex with `shlex.shlex(command, posix=True,
punctuation_chars=";|&<>")` so quoted regions stay whole, then
splice the rewrite back into the original command — extending the
splice span to include any matching quote chars around the path so
the fresh `shlex.quote` of the replacement isn't double-wrapped.
`_resolve_virtual_mount_path` now returns the unquoted path; the
caller owns shell-quoting, which avoids the previous
`shlex.quote` inside original `"…"` leaving literal `'` chars in
the argument value.
Unquoted paths with embedded whitespace remain a known limitation
(shlex has no way to know the user meant one path) — the
workaround of avoiding spaces in skill directory names still
applies, as flagged in the original issue.
Closes#237
* fix(backends): backslash-escaped paths, multi-path per token, subshell paths
- Fix backslash-escape handling: use unescape before rewriting
- Fix re.search→re.finditer: all /-paths in a token are rewritten
- Keep ( ) and backticks inside word tokens so paths spanning
\ or wrapped in backticks are matched correctly
- Add _try_rewrite helper with URL detection and unescape logic
- Add 10 contract tests pinning the din0s review cases
* fix(backends): restore () and backtick as shell operators for validate_command
- Restore ( ) and backtick to the operator set in _shell_token_spans.
Removing them caused a security regression: commands like (sudo ls)
would not detect sudo as a blocked command because (sudo became one
word token. With operators restored, validate_command correctly
catches blocked commands inside subshells and command substitutions.
- Fix _value_span_to_raw_span: the 'quoted' flag from the tokenizer
means the token *contains* a quoted segment (not necessarily starts
with a quote). Replace raw[0] assumption with a forward scan for
the first quote char, consuming unquoted prefix chars 1:1.
- Update test_system_path_with_shell_expansion: paths are now
partially rewritten because () are operators. Test updated to
reflect this known limitation (security > path rewriting).
* fix(test): cross-platform compatibility for pre-existing Windows failures
- python3 -> python in execute() calls (python is on PATH in any activated venv)
- sleep 10 -> _sleep_cmd(10) cross-platform helper
- str().endswith() -> Path().parts assertions (backslash-safe on Windows)
- shlex.quote exact-match assertions -> 'in' assertions (Windows quotes paths differently)
- mkdir -p E2E test -> preprocessor boundary test
- Skip 3 E2E tests on Windows: shlex.quote produces POSIX quoting incompatible with cmd.exe
141 passed, 3 skipped on Windows.
* fix: update docstring + strengthen shell-expansion test assertion
- Fix _value_span_to_raw_span docstring: no longer assumes raw[0] is
the opening quote, scans forward for first quote char
- Strengthen test_system_path_with_shell_expansion: verify
./workspace/notes is rewritten, not just notes in result
* style: ruff format backends.py + test_backends.py
* refactor(backends): simplify quoted virtual path rewriting
Replace 500+ line shlex tokenizer with 12-line pre-process step. Match quoted args via regex, unescape, rewrite via _rewrite_quoted_path, substitute with shlex.quote. 133 passed, 3 skipped.
* fix: guard bare absolute paths from double-rewrite by post-process regex
On POSIX, shlex.quote returns bare paths (e.g. /tmp/memories/note.md).
The pre-process substitutes these into the command, then the post-process
regex re-matches and incorrectly rewrites them.
Fix: _guard_bare_absolute wraps bare /-paths in single quotes so the
post-process regex''s character class stops at the quote char.
* style: ruff format
* fix(backends): narrow pre-process to exclude system-prefixed paths
Only rewrite quoted paths that are NOT known system prefixes.
* fix: narrow quoted-path pre-process to virtual mounts only
Only rewrite quoted /... paths that resolve to actual virtual mounts (/skills/..., /memories/...) or workspace-prefixed system paths. Remove catch-all that incorrectly rewrote bare paths like echo /hi.
Addresses din0s review feedback on #269.
* docs: update docstring for narrower quoted-path rewrite scope
* fix: session lost after evoscientist restart
* feat: Implement memory worker thread deletion on completion
- Added synchronous and asynchronous functions to delete memory worker threads after they finish execution, ensuring no residual checkpoints are left in the database.
- Enhanced `_watch_memory_worker_run_sync` and `_watch_memory_worker_run_async` to invoke deletion functions upon confirming worker completion.
- Introduced tests to verify that worker threads are deleted correctly upon successful completion and that failures in deletion do not affect the overall worker status.
- Updated session management to ensure that only relevant threads are restored from the database, preventing exposure of internal or unrelated workspace threads.
- Implemented a purge function to clean up leftover worker checkpoints during server startup.
* feat: Implement short thread ID display for CLI and session hints
* fix: ensure proper accounting and deletion order for memory worker threads
---------
Co-authored-by: z00827015 <zhoulun1@huawei.com>
* feat(dangerous-mode): implement real-filesystem access with safety checks
- Introduced a 'dangerous mode' allowing the agent to operate on the real filesystem.
- Updated command validation to bypass path confinement while enforcing a blocklist for privileged commands.
- Added warnings and guidelines for users when operating in dangerous mode.
- Enhanced configuration to support dangerous mode and ensure it implies auto-approval.
- Updated tests to verify the behavior of commands and configurations in dangerous mode.
* feat(dangerous-mode): enhance logging and environment management for dangerous mode
* feat(dangerous-mode): improve handling of dangerous mode with environment flags and enhance test isolation
* ci: add windows-latest to test matrix + fix 11 cross-platform test bugs
The test workflow ran on ``ubuntu-latest`` only. Per the issue's
first bullet — the maintainer's explicit #1 priority — add
``windows-latest`` to the matrix so the manager and related
modules are exercised on Windows on every PR.
The matrix addition surfaces 18 pre-existing Windows-only test
failures. Without fixes the new leg would be 18+ reds from
day one and the matrix would just produce a wall of
``fail-fast`` noise. This PR fixes 11 of them; each fix is
a real (cross-platform) bug, not a Windows-specific hack —
most were already flagged by CodeRabbit on PR #236 but never
acted on. The remaining 4 failures need code refactors
(``os.killpg`` → ``psutil`` in ``background.py``,
``convert_virtual_paths_in_command`` Windows-aware quoting,
tilde expansion) that are documented as out-of-scope
follow-ups below.
## What changed
* ``.github/workflows/test.yml``
- ``os: [ubuntu-latest, windows-latest]`` → 2 OS × 2 Python
= 4 cells.
- ``fail-fast: false`` so one bad cell doesn't cancel the
rest while the Windows leg is being brought up. Removable
in a future PR once the suite is fully green.
* ``tests/test_backends.py``
- Hard-coded ``"python3"`` → ``{sys.executable}`` in 7
test commands. Windows has no ``python3`` on PATH; using
``sys.executable`` is portable and matches what CodeRabbit
flagged on PR #236.
- Strict string comparisons → ``shlex.split`` round-trip in
5 resolver tests. ``shlex.quote`` adds single quotes
around backslash paths on Windows, which broke the
direct ``==`` compare.
- Cross-platform suffix checks in 2 path-resolution tests
(``Path(resolved).parts[-2:]`` instead of
``str(resolved).endswith("src/main.py")``).
- ``mkdir -p`` → ``sys.executable -c "import os;
os.makedirs(...)"`` in the cwd-sanitization test.
- ``skipif(sys.platform == "win32")`` on 3 e2e tests that
hit the underlying ``shlex.quote`` + ``cmd.exe`` quoting
bug (real, separate issue).
* ``tests/test_sessions.py``
- ``test_uses_data_dir``: check ``.evoscientist`` in the
long path form (via ``Path.resolve()``) rather than the
short-path form ``get_db_path`` returns on Windows.
* ``tests/test_mcp_client.py``
- ``endswith("python")`` → ``Path(result).stem.lower()`` so
``python.EXE`` matches on Windows.
- ``endswith("npx")`` also accepts ``npx.cmd`` so the npm
shim on Windows matches.
## Out of scope (follow-up issues to file)
* ``os.killpg`` doesn't exist on Windows
(``EvoScientist/background.py:248``) — 3 background tests
fail. Real fix is the same ``psutil`` walk pattern PR #200
shipped in ``langgraph_dev/manager.py``.
* Tilde expansion in file mentions.
* Windows-aware shell quoting in
``convert_virtual_paths_in_command``.
* Path conventions (``~/.config/evoscientist/`` vs
``%APPDATA%\EvoScientist``) — needs design discussion +
``platformdirs`` migration.
* Cross-module audit of
``EvoScientist/tools/execute.py``,
``EvoScientist/ccproxy_manager.py``,
``EvoScientist/config/onboard.py``.
Closes#207 (step 1 only — CI matrix + the easy test
fixes; remaining bullets tracked separately).
* fix: cross-platform compatibility for Windows CI runners
- background.py: replace POSIX-only os.killpg/os.getpgid with
cross-platform _kill_process_tree() helper. On Windows falls back
to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX
keeps existing os.killpg logic.
- test_backends.py: replace mkdir -p shell execution in
test_literal_workspace_path_replaced with preprocessing-boundary
assertion (patch LocalShellBackend.execute, capture command,
assert workspace path was rewritten to ./). Avoids POSIX-only
mkdir -p on Windows runners.
- test_file_mentions.py: monkeypatch USERPROFILE on Windows so
ntpath.expanduser() resolves ~ to tmp_path even when HOME is
unset on CI runners.
* fix(test): cross-platform sleep/true commands for Windows CI
Replace POSIX-only sleep/true with module-level helpers that use
ping -n / cmd /c on Windows. Also fix python3 -> sys.executable
in the non-timeout recovery test.
- test_background.py: 7 sleep/true fixes
- test_background_middleware.py: 6 sleep/true fixes
- test_backends.py: 4 sleep fixes + 1 python3 fix
2318 passed, 0 failed on Windows.
* fix(test): use shell-portable double quotes for python -c on Windows
cmd.exe does not treat single quotes as string delimiters, so
-c 'raise SystemExit(1)' was passed with literal quotes on Windows.
Switch to double quotes which work on both cmd.exe and POSIX sh.
* fix: use psutil for Windows process tree kill + avoid sys.executable under uv
- background.py: replace Popen.terminate()/kill() with psutil-based
process tree walking on Windows. TerminateProcess does NOT cascade
to grandchildren; psutil.Process.children(recursive=True) ensures
the entire tree is signaled.
- test_backends.py: replace sys.executable with 'python' in sandbox
execute() calls. Under uv, sys.executable is under the workspace
and gets rewritten to ./ by prepare_sandbox_command, breaking
Linux CI. The plain 'python' command resolves correctly in any
activated venv.
* fix: broaden try/except in _kill_process_tree to cover proc.children()
If the process exits between Process(popen.pid) and children(recursive=True),
the children call raises an uncaught exception escaping stop(). Move it inside
the existing try/except block.
* fix: narrow exception to ProcessLookupError in POSIX _kill_process_tree
OSError is too broad — would silently swallow EPERM on SIGKILL, leaving
the process alive when we report it as stopped. Match original behavior
which only caught ProcessLookupError (process already gone).
* style: ruff format test_backends.py
* ci: trigger re-run for flaky prompt_toolkit test
* style: fix ruff check (import order + RUF005 unpacking)
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix(langgraph-dev): rotate langgraph_dev.log when it exceeds 50MB
``_LOG_FILE`` (``~/.config/evoscientist/langgraph_dev.log``) was
opened in ``start_langgraph_dev`` with plain ``"ab"`` and never
rotated, so it grew unbounded over weeks/months of heavy use —
especially when chatty MCP servers spawned by langgraph dev
filled it, or when failure paths produced stack traces.
Implement the recommended option 1 from #209: filesize-based
rollover. When the active log exceeds 50MB on the next
``start_langgraph_dev`` invocation, rename it to
``langgraph_dev.log.1`` (overwriting any existing backup) via
``os.replace`` and start fresh. Single-backup policy keeps the
disk footprint bounded at roughly 2x threshold.
Rotation is best-effort: ``_rotate_log_if_needed`` logs and
swallows OSError so a permission error or racing rename can't
block langgraph dev from starting. The next ``start`` invocation
will try again — worst case the log grows for one more session.
Options 2 (timestamped per-session + 7-day sweep) and 3
(``RotatingFileHandler`` + pipe) are explicitly NOT done — option
1 is simplest, no async machinery, matches the issue's
recommendation.
Closes#209
* test(langgraph-dev): redirect _PID_DIR in rotate integration test
Address CodeRabbit review comment on #270: the
``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open``
test patched only ``_LOG_FILE`` to a tmp path, but
``start_langgraph_dev`` also calls ``_PID_DIR.mkdir(...)`` as part
of its prelude, which would create a real directory under
``~/.config/evoscientist/`` on a dev machine. Redirect
``_PID_DIR`` to ``tmp_path / "pids"`` too so the test stays
fully isolated. Add a final assertion that ``pid_dir.is_dir()``
holds, proving the function reached past the mkdir call.
* refactor(langgraph-dev): bundle runtime paths into LanggraphRuntimePaths
@din0s review follow-up on #270: the previous test isolation patched
only ``_LOG_FILE`` (and after a second round, ``_PID_DIR``), but
``start_langgraph_dev`` still touches 5 distinct on-disk paths. Patching
any subset of those still leaves the others pointing at the user's real
``~/.config/evoscientist/`` — exactly the case that produced the
"Port 6174 cannot be bound after waiting 60s" symptom on the
reviewer's machine.
Replace the five free-floating module-level constants
(``_PID_DIR`` / ``_PID_FILE`` / ``_LOG_FILE`` / ``_WORKSPACE_SIDECAR``
/ ``_FILE_LOCK_PATH``) with a single ``LanggraphRuntimePaths`` frozen
dataclass exposed as a module-level ``RUNTIME`` instance. Production
code accesses ``RUNTIME.pid_file`` etc.; tests can now substitute the
*whole* bundle in one assignment:
monkeypatch.setattr(
manager, "RUNTIME",
manager.LanggraphRuntimePaths.for_directory(tmp_path / "runtime"),
)
The classmethod ``for_directory(pid_dir)`` builds an isolated bundle
rooted at a single dir, so the test author doesn't spell out every
path field. Tests that only care about one field (e.g. pid_file
during the stale-process kill path) use
``dataclasses.replace(manager.RUNTIME, pid_file=X)`` — frozen
dataclass-friendly, no need to enumerate the other four fields.
The dataclass's docstring records the migration rationale (the old
five-name layout invited inconsistent patches).
External callers of the old constants updated:
- ``EvoScientist/deploy/server.py`` and ``webui.py`` now import
``RUNTIME`` and use ``RUNTIME.log_file`` for the on-screen log
path hint. The other imports they had (``_DEFAULT_PORT``,
``_is_port_occupied``, ``_read_workspace_sidecar``) are still
module-level functions/values, untouched.
Test updates:
- ``tests/test_langgraph_manager.py``: ``patch.object(manager, "_XXX",
X)`` patterns now go through ``dataclasses.replace(manager.RUNTIME,
xxx=X)``; the ``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open``
test (from the previous #270 review iteration) uses
``for_directory`` for one-shot isolation.
- ``tests/test_langgraph_dev_workspace_sidecar.py``: each test now
goes through a tiny ``_isolated_runtime(monkeypatch, tmp_path)``
helper that calls ``for_directory``.
- ``tests/test_langgraph_dev_deploy_mode.py``: same ``for_directory``
swap.
No production behavior change. All ``langgraph_dev``-side tests
(``test_langgraph_manager.py`` 26/26, ``test_langgraph_dev_workspace_sidecar.py``
14/14, ``test_langgraph_dev_deploy_mode.py`` 14/14, ``test_cli_deploy.py``
18/18 — which indirectly exercises deploy/server.py and deploy/webui.py
imports) pass. Full-project test count unchanged from baseline; the
remaining 22 Windows-only pre-existing failures (test_background
``os.killpg``, test_file_mentions tilde, mcp_client ``shutil.which``,
test_sessions 8.3 short path) are documented as out-of-scope for #207.
* style: apply ruff format to langgraph_dev test + module files
CI lint check on #270 failed:
Run ruff format --check .
Would reformat: EvoScientist/langgraph_dev/manager.py
Would reformat: tests/test_langgraph_manager.py
Plus two test files touched by the prior consolidation commit that
``ruff format`` hadn't seen yet:
tests/test_langgraph_dev_deploy_mode.py
tests/test_langgraph_dev_workspace_sidecar.py
Just formatting. No logic change. All 75 refactor-related tests pass.
* fix(test): use for_directory for full path isolation + patch _can_bind_port to skip real socket ops
Two fixes for TestStartLanggraphDevRotatesLog:
1. Replace dataclasses.replace(manager.RUNTIME, ...) with
LanggraphRuntimePaths.for_directory(pid_dir) so pid_file,
workspace_sidecar, and lock_file are also temp-rooted
(prevents leak to ~/.config/evoscientist/).
2. Monkeypatch _can_bind_port to always return True so the
bind-poll loop in _wait_for_port_bindable passes immediately
without touching real sockets (fixes 60s timeout on machines
where port 6174 is already in use).
* fix: cross-platform compatibility for Windows CI runners
- background.py: replace POSIX-only os.killpg/os.getpgid with
cross-platform _kill_process_tree() helper. On Windows falls back
to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX
keeps existing os.killpg logic.
- test_backends.py: replace mkdir -p shell execution in
test_literal_workspace_path_replaced with preprocessing-boundary
assertion (patch LocalShellBackend.execute, capture command,
assert workspace path was rewritten to ./). Avoids POSIX-only
mkdir -p on Windows runners.
- test_file_mentions.py: monkeypatch USERPROFILE on Windows so
ntpath.expanduser() resolves ~ to tmp_path even when HOME is
unset on CI runners.
* refactor(test): add runtime_paths fixture to isolate manager.RUNTIME
Adds a reusable fixture that monkeypatches manager.RUNTIME to a
temp-rooted LanggraphRuntimePaths.for_directory(). Tests that need
specific fields can still dataclasses.replace(runtime_paths, ...)
but the baseline is always temp-isolated, preventing leaks to
~/.config/evoscientist/.
Updated test_langgraph_dev_deploy_mode.py, test_langgraph_dev_workspace_sidecar.py,
and test_langgraph_manager.py to use the fixture, consolidating sequential
lock_file + pid_dir patches into single dataclasses.replace calls.
* Revert "fix: cross-platform compatibility for Windows CI runners"
This reverts commit eb025d24af32e195a982cd40f6d70dba885c4019.
* style: ruff format conftest.py
* fix: address review issues in log-rotation + runtime paths
- Use for_directory(tmp_path/pids) as base in ensure_langgraph_dev tests
so pid_file/log_file are co-located with pid_dir, not split across paths
- Remove unused runtime_paths param from test_no_existing_file_is_noop
- Replace manager.RUNTIME with runtime_paths in two sidecar tests
- Use for_directory(DEFAULT_PID_DIR) instead of explicit construction
- Fix stale _LOG_FILE reference in TestRotateLogIfNeeded docstring
* style: ruff format test files
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* refactor(agent): make create_cli_agent(config=, chat_model=) pure
Re-applies the #183 purity refactor on top of the observation-memory
lifecycle that landed in #259, integrating the two cleanly.
create_cli_agent gains a pure path: when both `config` and `chat_model`
are passed it builds the agent entirely from locals and writes none of
the cached module globals (`_config`, `_chat_model`, `_chat_model_key`,
`_EvoScientist_agent`). `/model` commits the switch via
`set_active_config` / `set_chat_model_instance` only after a successful
build, so a failed rebuild leaves the session on the original model
(replaces the old snapshot/restore rollback).
Supporting changes:
- Extract `set_active_config` (write-half of `_ensure_config`),
`_apply_env_from_config`, `_build_chat_model`, and
`set_chat_model_instance`.
- Thread `cfg` / `chat_model` through `_get_default_middleware`,
`_build_base_kwargs`, `load_mcp_and_build_kwargs`,
`_maybe_swap_async_subagents`, and `_inject_subagent_middleware` so the
pure path never falls back to the global-writing `_ensure_config()` /
`_ensure_chat_model()`.
- Integrate with #259's memory middleware: subagent context-editing
middleware binds the threaded `chat_model`, and the configured system
prompt / memory controls read the threaded `cfg` (new threading vs the
original #183, required because #259 made these paths read config).
- Consolidate `cfg` resolution to one `cfg if cfg is not None else
_ensure_config()` at the top of each kwargs builder, matching the
pattern already used in the other config-aware helpers.
* fix(agent): keep pure tool selector off global cache
* fix(model): apply config switch in place to preserve reference integrity
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* feat(middleware): reposition code interpreter middleware in the stack
* feat(models): add qwen3.7-plus model entry and update context window comment
* feat(models): add qwen3.7-max and qwen3.7-plus model entries for DashScope
* feat(auxiliary): implement auxiliary model support for background tasks and tool selection
- Added auxiliary model configuration to EvoScientistConfig.
- Introduced _ensure_auxiliary_chat_model function to manage auxiliary model instances.
- Updated onboarding steps to include auxiliary model selection.
- Modified middleware to route tool selection to the auxiliary model when applicable.
- Enhanced tests to cover auxiliary model functionality and configuration.
* feat(steps): update UI backend selection options and descriptions
* Refactor code structure for improved readability and maintainability
* feat(patches): implement OpenRouter response reasoning item stripping to prevent multi-turn errors
* feat: update version to v0.1.4 in badges, README, and pyproject.toml; adjust skill counts in steps.py
* feat(config): add auxiliary model and provider environment variables to test setup
* feat(memory): add observation memory lifecycle
Add file-backed observation memory with deterministic markdown records,
structured record_observation tooling, startup indexing, and
profile/observation prompt guidance.
Launch post-turn and post-subagent EvoMemory workers through LangGraph
dev so completed runs can update profile memory, save durable
observations, and write subagent execution summaries without blocking
the active agent.
Wire memory middleware into the main agent, subagents, async graphs, TUI
status reporting, worker activity accounting, and observation-aware
research prompts, with regression coverage for storage, lifecycle
scheduling, graph registration, status display, and stream reset
behavior.
* fix(cli): sync background agent server on resume
Resume flows now need to keep the LangGraph dev background server
aligned with the active workspace even when async subagents are
disabled. EvoMemory workers use that server too, so gating resume-time
sync on enable_async_subagents could leave workers pinned to the launch
workspace after resuming a thread from another workspace.
Run workspace sync unconditionally for Rich CLI and Textual resume
paths, while preserving WorkspaceMismatchError handling so failed sync
aborts the resume before mutating the active thread or workspace.
Propagate aborted resume callbacks through the command UI so
channel-issued /resume commands do not send false success or history
output. Channel slash dispatch now treats CommandManager-caught command
errors as command errors and skips completion hooks for those failed
commands.
Add regression coverage for disabled async subagents, callback aborts,
and channel command error reporting.
* fix(cli): prepare serve resume workspace before adopting
Load the resumed workspace agent and sync the background server as a
single pre-adoption step. Restore the previous active workspace if
preparation fails so serve mode keeps using the old session
consistently.
* fix(memory): untrack abandoned worker status watches
Stop treating watcher shutdown as confirmed worker completion. Terminal
worker statuses still count memory deltas, while poll failures or
watcher setup failures now remove the active run without crediting
partial outputs.
* fix(cli): report channel command failures accurately
Treat command_error as a None sentinel so empty error strings still
fail, and let TUI resumes continue only on non-mismatch
background-server sync failures while reporting degraded mode.
* fix(stream): clear memory counters for resume streams
Reset completed-memory counters for every new agent stream, including
Command-based HITL and resume streams, so saved-memory indicators do not
leak across turns.
* docs(tools): make observation recording guidance conditional
Clarify that agents should call record_observation only when the
observation tool is available, preserving the existing durability and
usefulness criteria.
* feat(config): add controls for profile and observation memory
Add config flags for profile memory, observation memory, observation
writer placement, and background memory workers.
Wire the controls through main agents, subagents, EvoMemory middleware,
and memory lifecycle workers so observation writes can be assigned to
the live agent, subagent worker, both, or neither. Keep turn memory
workers profile-only and make prompts reflect the available observation
read/write paths. Skip langgraph dev startup when neither async
subagents nor memory workers need the background server.
Add coverage for config parsing, prompt gating, middleware wiring, and
worker tool availability.
* test(cli): include memory defaults in serve config stubs
* fix(memory): offload async worker launch blocking calls
Run the langgraph-dev health check and memory-output snapshot in worker
threads from the async EvoMemory launcher so it does not block the event
loop.
* chore(memory): harden turn worker subagent guardrail
* chore(memory): refresh profile context per request
* fix(memory): offload async profile file reads
* fix(memory): offload async worker completion accounting
* Enhance multimodal handling in LLM model
- Updated `_flatten_message_content` to preserve media blocks (images, files) while flattening text content.
- Introduced `_sanitize_messages` to manage media hoisting for tool messages, ensuring compatibility with OpenAI APIs.
- Modified `_patch_openai_compat_content` to accommodate new media handling logic, including retry mechanisms for media errors.
- Added comprehensive tests for media preservation, including various scenarios with images, files, and unsupported media types.
* fix: preserve order of text and media blocks in message flattening
* test: add tests for _strip_media_types to ensure position preservation and deduplication
* feat(memory): migrate to profile memory files
* chore(stream): read profile headings from templates
* fix(display): keep assistant responses if response_text has started
* fix(memory): do not treat failed bootstraps as profile creation
* chore(memory): unlink blank legacy memory
* fix(memory): resolve project_id once
* fix(memory): preserve unreadable profile files
* chore(tui): render streamed narration inline with tool timeline
Update the TUI streaming timeline so assistant text emitted before or
between tool calls is rendered inline where it occurs, rather than being
kept as a single answer bubble above or below the tools.
If the model begins an assistant response and then emits another tool
call, the provisional response is converted into inline narration before
that tool. The final assistant message then renders only the remaining
response suffix, avoiding duplicate text in the completed transcript.
Stop/cancel handling now preserves any active inline narration, appends
the visible stopped marker only to the remaining displayed segment, and
still returns the full normalized stopped response for channel callers.
Completed tools continue to collapse while long runs are active, but
expand again when the turn reaches a final state so the completed
transcript shows the full tool timeline.
* fix(stream): preserve narration around tool timelines
Keep assistant narration attached to the tool call that follows it
instead of folding all streamed text into the final answer block.
Track narrated response segments in stream state, render them before
their corresponding regular or task tool entries, and keep final answers
limited to the response suffix that has not already been shown inline.
Preserve narration across normal completion, stop/error final frames,
sub-agent task calls, and collapsed live tool summaries.
Add regression coverage for pending tools, completed tools, sub-agent
task delegations, collapsed completed/running tool summaries, and final
stop frames.
* fix(tui): finalize inline narration transitions
* test(memory): use canonical project id helper
* feat: add WebUI mode support with related configuration and onboarding steps
* feat: enhance WebUI port configuration to prevent conflicts with backend port
* feat: add support for fresh interactive session detection in WebUI
* fix: update deepagents and langchain versions; enhance _reduce_messages_delta handling for None state
* fix: update langchain version constraint to >=1.3 in pyproject.toml and uv.lock
* feat: implement configurable sandbox execute timeout and enhance recovery instructions
* feat: add background process management tools and middleware for sandbox execution
* feat: enhance background process management with completion notifications and deduplication
* feat: enhance sandbox execution timeout validation and update related messages
* feat: enhance background process management with thread-specific completion notifications and HITL approval handling
* test: assert completion notification waits for process finish timestamp
* fix(cli): forward async-notifier replies back to originating channel
When PR #214's auto-notifier fires a synthetic agent turn after a
channel-originated conversation, the synthesized response only rendered
to the local CLI/TUI — the channel user (iMessage etc.) saw nothing
and had to manually re-prompt to find out what happened.
Adds a per-thread channel-origin registry in cli/channel.py and wires
the three notifier paths (Rich CLI / TUI / serve) to publish the final
response back via bus.publish_outbound when the originating thread was
started by a channel turn. Publish is fire-and-forget (scheduled on the
bus loop + done-callback for failure logging) so the notifier turn
doesn't block on the asyncio / textual event loop.
The registry is cleared on /new and /resume rotation so stale entries
don't accumulate.
* fix(cli): address review feedback on channel-origin forwarding
Follow-up to the review on #244 (din0s, X-iZhang):
- Guard the /resume origin cleanup on a real thread change in Rich CLI
and TUI (serve mode already did via thread_changed). Resuming the
already-active thread no longer wipes its still-live origin, which
would otherwise silently drop a later async-notifier forward — the
exact gap this PR closes.
- Re-bind the now-current thread to its channel after a channel-issued
/new or /resume slash command (which rotates the thread inside the
dispatch), so notifier turns on the rotated thread still forward.
- Guard the publish done-callback against a cancelled future, whose
.exception() raises CancelledError (rather than returning it) on
bus-loop teardown, so the intended warning still logs.
- Mirror the normal reply path's manager.record_message(channel, "sent")
for forwarded notifications so per-channel stats stay accurate.
- Print the closing "[channel: Replied to ...]" line in all three
notifier paths (Rich CLI / TUI / serve) when a forward actually
happened, so the forwarded block reads as terminated on screen.
Adds test_publish_records_sent_metric. ruff clean; notification-origin
suite (10) + related channel/CLI/serve suites (728) pass.
* fix(cli): store sender information separately from chat_id in channel origin
---------
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* Add onboarding wizard with style and validation components
- Introduced `style.py` for shared visual elements used in the onboarding wizard.
- Created `validators.py` for input validation, including integer and choice validators, and API key validation functions for various providers.
- Implemented `wizard.py` as the entry point for the onboarding process, managing user prompts and configuration steps.
- Added progress rendering and autosave functionality to enhance user experience during the onboarding process.
* feat(onboarding): enhance validation and configuration for onboarding wizard
- Added validation for UI backends, workspace modes, and providers in the onboarding command.
- Updated channel definitions to include secret field handling for sensitive tokens.
- Improved user prompts for required fields, ensuring sensitive data is masked.
- Introduced constants for valid providers, UI backends, and workspace modes to maintain consistency.
- Implemented tests to ensure alignment between constants and interactive choices in onboarding steps.
* feat(onboarding): improve WeChat account ID prompt and validation for newly enabled channels
* feat(onboarding): enhance WeChat backend credential prompts and validation
* feat(onboarding): refine WeChat backend credential prompts for wecom and wechatmp
* Refactor onboarding package for improved structure and clarity
- Simplified the onboarding package by removing unnecessary re-exports and consolidating public API to only include `run_onboard`.
- Updated `install_back_keys` to `install_navigation_keys` for clarity and consistency in the prompter module.
- Enhanced the `NonInteractivePrompter` class to support strict mode, allowing for better handling of non-interactive prompts.
- Adjusted the onboarding steps to utilize the new navigation keys installation method.
- Improved the `run_onboard` function to handle section implications based on user flags, enhancing the onboarding experience.
- Updated tests to reflect changes in imports and ensure compatibility with the new structure.
* feat(onboarding): enhance validation logic for non-interactive prompts
* refactor(onboarding): streamline onboarding module structure and enhance validation error handling
* refactor(onboarding): enhance config revert logic to preserve original file state
* refactor(onboarding): enhance tavily key validation and error handling in onboarding process