19 Commits

Author SHA1 Message Date
m4 470cf75722 merge: bring upstream v0.3.0 (72 commits) into Ai4Sci fork
Merged upstream/main (418abca, release v0.3.0) into our fork on a
dedicated branch. 21 conflicting files resolved; main worktree untouched.

Resolution policy and key decisions:
- Keep Ai4Sci runtime endpoints, durable dispatch, workspace scopes and
  the HITL/DynamicReview approval chain (approval path is product-critical).
- Adopt upstream model registry (llm/registry.py): our 136 model entries
  are a strict subset of upstream's 180, so dropping our inline table
  loses nothing and gains 44 new models.
- Adopt upstream native EvoChatDeepSeek; drop our obsolete
  _patch_deepseek_reasoning_passback monkey patch.
- Keep our six patches.py additions, ported onto upstream's new
  _OpenAICompatContent class: stable tool-call ids, tool-history
  sanitization, drop_reasoning_metadata, empty-SSE keepalive,
  extracted-document-text patch, _has_assistant_tool_protocol.
- Keep our skill-budget middleware path (skills=None) instead of passing
  skills through, to avoid double loading.
- Keep sanitized error labels (_safe_error_label) while adopting
  upstream's injected MiddlewareEventSink for fallback narration.
- Keep port 3076 and the LANGGRAPH_SERVER_URL override; adopt upstream's
  host/probe-host handling and CONFIG_DRIFT_SINCE_LAUNCH.
- Adopt upstream dependency stack: deepagents 0.7.6, langchain-quickjs
  0.3.7, langgraph-api 0.14; keep our extra deps (rfc8785, pillow,
  firecrawl-anydoc, nest-asyncio).
- Align call sites with upstream APIs: create_tool_selector_middleware
  now takes events= instead of track_stream_selection=.
2026-09-13 16:07:27 +08:00
houren Antony 932c934485 fix(mcp): give stdio subprocess a real stderr fd under redirected streams (#423)
* fix(mcp): give stdio subprocess a real stderr fd under redirected streams (#418)

On Windows the Textual TUI redirects sys.stderr to an in-memory capture
(textual.app._PrintCapture) whose fileno() returns -1. The MCP SDK forwards
that stderr to stdio server subprocesses via subprocess.Popen(stderr=...),
and Popen rejects the invalid handle with OSError: [Errno 9] Bad file
descriptor — so only stdio servers fail to load (HTTP/SSE are unaffected).

Wrap mcp.client.stdio.stdio_client so that, whenever the configured errlog
has no usable fileno, it falls back to sys.__stderr__ (or os.devnull in GUI
hosts). Idempotent, no-op when the SDK is absent, warns if the SDK renames
stdio_client. Adds 9 regression tests and a troubleshooting note.

* fix(mcp): validate live fd and close fallback errlog after stdio session

Address CodeRabbit review on #423:
- _stdio_errlog_is_usable now os.fstat()s the fd to reject closed streams
  that still report their former positive fileno (prevents a deferred
  [Errno 9] from subprocess.Popen).
- The stdio_client wrapper owns the devnull fallback it allocates and
  closes it once the session exits, so repeated MCP reloads no longer leak
  file descriptors. Caller-provided usable errlogs pass through untouched.
- Tests cover the closed-fd case, the fd-leak/closure invariant, and
  confirm langchain-mcp-adapters binds the patched stdio_client.

* fix(mcp): rebind adapter stdio_client, forward errlog by kw, harden tests

Address CodeRabbit round-2 review on #423:
- The patch now also rebinds langchain_mcp_adapters.sessions.stdio_client,
  which the adapter captures via a 'from' import at module load — so the
  wrapped function reaches the adapter regardless of import order.
- errlog is forwarded to the SDK by keyword (original(server, *args,
  errlog=errlog, **kwargs)) so a future SDK inserting a positional
  parameter before errlog can't mis-bind the fallback.
- The fallback stream is now allocated inside the async context manager,
  so it is closed on session exit even if the CM is constructed but never
  entered (narrower fd-leak path).
- test_closed_fd_rejected now reaches the os.fstat branch (stale positive
  fd stub) instead of the ValueError path; test_adapter_binds_patched_stdio_client
  documents and asserts the import-order-independent rebind.

* fix(mcp): close fallback errlog when stdio_client construction fails

Address CodeRabbit round-3 review on #423: move the original(server, *args,
errlog=errlog, **kwargs) construction inside the try block so a failure
during subprocess/client setup still reaches the finally and closes the
wrapper-owned os.devnull stream. Added test_fallback_closed_when_construction_fails
covering the path.

* refactor(mcp): track fallback ownership via (stream, opened_by_us)

Address din0s review on #423:
- _safe_stdio_errlog() now returns (stream, opened_by_us); the wrapper closes
  the fallback only when opened_by_us is True, instead of inferring ownership
  from needs_fallback + an identity check against sys.__stderr__. Simpler and
  less likely to regress.
- Removed dead try/finally in test_closed_fd_rejected.
- Added test_wrapped_stdio_client_swaps_explicit_bad_errlog covering the
  'not _stdio_errlog_is_usable(errlog)' branch (explicit bad errlog, not the
  default sentinel).
- Updated test_safe_errlog_returns_usable_stream for the tuple return.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-08-13 17:39:45 +00:00
dinos 8b1451cdda refactor(runtime): centralize async bridges under an owned runtime (#376)
* feat(runtime): add application-scoped async runtime

* refactor(cli): use owned runtime for session stats

* refactor(onboard): use the owned async runtime

* docs(runtime): record async bridge ownership

* refactor(middleware): keep sync fallback synchronous

* refactor(mcp): load tools on an owned runtime

* refactor(cli): share owned runtime across entry points

* refactor(channels): make inbound sync bridge explicit

* refactor(stream): run Rich streaming on owned runtime

* chore(runtime): remove nest-asyncio dependency

* refactor(asyncio): require active loops in async code

* docs(runtime): document final event loop ownership

* fix(stream): cancel stalled owned streams

* fix(cli): recover cleanly from stream cancellation

* fix(runtime): drain executor work before shutdown

* fix(runtime): terminate cancelled shell process trees

* fix(models): let fallback bypass selector failures

* fix(cli): reset interrupt handling between turns

* docs: rm implementation spec

* fix(serve): cancel active turns during shutdown

* fix(runtime): protect settlement from waiter cancellation

* fix(backends): reject empty shell commands

* fix(runtime): terminate descendants after shell exit

* fix(mcp): keep standalone discovery off channel loop

* fix(cli): own and settle interactive prompt cancellation

* fix(serve): keep channel sends off runtime loop

* fix(stream): scope cancel context to iterator steps

* refactor(serve): require the owned async runtime

* fix(channels): keep interactive sends off runtime loop

* fix(selector): surface fallback without log spam

* test(runtime): normalize Windows shell marker

* fix(cli): serialize interactive session turns

* fix(shell): bound output drain after termination

* fix(ui): do not retry owned runtime failures

* fix(shell): allow signal-safe registry reentry

* fix(shell): avoid terminating reused process ids

* fix(channels): preserve streaming send order

* fix(cli): report runtime shutdown timeouts cleanly

* fix(mcp): guide async callers to async loader

* docs(runtime): clarify reserved async bridge APIs

* fix(runtime): bound code interpreter cleanup

* test(shell): use active Python for drain regression

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-27 14:17:57 +01:00
m4 4fc74e7da7 EvoScientist Ai4Sci
Docker / build (push) Has been cancelled
Build / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
2026-07-14 22:07:14 +08:00
dinos 690b903f85 test: standardize async tests on pytest-asyncio auto mode (#338)
* chore: add pytest-asyncio in auto mode

* test: migrate channel and stream tests to native async

Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.

* test: migrate command and model/middleware tests to native async

Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.

* test: migrate TUI, notifier, gateway, and session tests to native async

TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.

* test: replace direct asyncio.run() calls with native async tests

Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.

* test: drop undeclared anyio markers and delete run_async helper

The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).
2026-07-08 18:37:48 +00:00
houren Antony d4f1fbd110 ci: add windows-latest to test matrix + fix 11 cross-platform test bugs (#271)
* ci: add windows-latest to test matrix + fix 11 cross-platform test bugs

The test workflow ran on ``ubuntu-latest`` only. Per the issue's
first bullet — the maintainer's explicit #1 priority — add
``windows-latest`` to the matrix so the manager and related
modules are exercised on Windows on every PR.

The matrix addition surfaces 18 pre-existing Windows-only test
failures. Without fixes the new leg would be 18+ reds from
day one and the matrix would just produce a wall of
``fail-fast`` noise. This PR fixes 11 of them; each fix is
a real (cross-platform) bug, not a Windows-specific hack —
most were already flagged by CodeRabbit on PR #236 but never
acted on. The remaining 4 failures need code refactors
(``os.killpg`` → ``psutil`` in ``background.py``,
``convert_virtual_paths_in_command`` Windows-aware quoting,
tilde expansion) that are documented as out-of-scope
follow-ups below.

## What changed

* ``.github/workflows/test.yml``
  - ``os: [ubuntu-latest, windows-latest]`` → 2 OS × 2 Python
    = 4 cells.
  - ``fail-fast: false`` so one bad cell doesn't cancel the
    rest while the Windows leg is being brought up. Removable
    in a future PR once the suite is fully green.

* ``tests/test_backends.py``
  - Hard-coded ``"python3"`` → ``{sys.executable}`` in 7
    test commands. Windows has no ``python3`` on PATH; using
    ``sys.executable`` is portable and matches what CodeRabbit
    flagged on PR #236.
  - Strict string comparisons → ``shlex.split`` round-trip in
    5 resolver tests. ``shlex.quote`` adds single quotes
    around backslash paths on Windows, which broke the
    direct ``==`` compare.
  - Cross-platform suffix checks in 2 path-resolution tests
    (``Path(resolved).parts[-2:]`` instead of
    ``str(resolved).endswith("src/main.py")``).
  - ``mkdir -p`` → ``sys.executable -c "import os;
    os.makedirs(...)"`` in the cwd-sanitization test.
  - ``skipif(sys.platform == "win32")`` on 3 e2e tests that
    hit the underlying ``shlex.quote`` + ``cmd.exe`` quoting
    bug (real, separate issue).

* ``tests/test_sessions.py``
  - ``test_uses_data_dir``: check ``.evoscientist`` in the
    long path form (via ``Path.resolve()``) rather than the
    short-path form ``get_db_path`` returns on Windows.

* ``tests/test_mcp_client.py``
  - ``endswith("python")`` → ``Path(result).stem.lower()`` so
    ``python.EXE`` matches on Windows.
  - ``endswith("npx")`` also accepts ``npx.cmd`` so the npm
    shim on Windows matches.

## Out of scope (follow-up issues to file)

* ``os.killpg`` doesn't exist on Windows
  (``EvoScientist/background.py:248``) — 3 background tests
  fail. Real fix is the same ``psutil`` walk pattern PR #200
  shipped in ``langgraph_dev/manager.py``.
* Tilde expansion in file mentions.
* Windows-aware shell quoting in
  ``convert_virtual_paths_in_command``.
* Path conventions (``~/.config/evoscientist/`` vs
  ``%APPDATA%\EvoScientist``) — needs design discussion +
  ``platformdirs`` migration.
* Cross-module audit of
  ``EvoScientist/tools/execute.py``,
  ``EvoScientist/ccproxy_manager.py``,
  ``EvoScientist/config/onboard.py``.

Closes #207 (step 1 only — CI matrix + the easy test
fixes; remaining bullets tracked separately).

* fix: cross-platform compatibility for Windows CI runners

- background.py: replace POSIX-only os.killpg/os.getpgid with
  cross-platform _kill_process_tree() helper. On Windows falls back
  to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX
  keeps existing os.killpg logic.

- test_backends.py: replace mkdir -p shell execution in
  test_literal_workspace_path_replaced with preprocessing-boundary
  assertion (patch LocalShellBackend.execute, capture command,
  assert workspace path was rewritten to ./). Avoids POSIX-only
  mkdir -p on Windows runners.

- test_file_mentions.py: monkeypatch USERPROFILE on Windows so
  ntpath.expanduser() resolves ~ to tmp_path even when HOME is
  unset on CI runners.

* fix(test): cross-platform sleep/true commands for Windows CI

Replace POSIX-only sleep/true with module-level helpers that use
ping -n / cmd /c on Windows. Also fix python3 -> sys.executable
in the non-timeout recovery test.

- test_background.py: 7 sleep/true fixes
- test_background_middleware.py: 6 sleep/true fixes
- test_backends.py: 4 sleep fixes + 1 python3 fix

2318 passed, 0 failed on Windows.

* fix(test): use shell-portable double quotes for python -c on Windows

cmd.exe does not treat single quotes as string delimiters, so
-c 'raise SystemExit(1)' was passed with literal quotes on Windows.
Switch to double quotes which work on both cmd.exe and POSIX sh.

* fix: use psutil for Windows process tree kill + avoid sys.executable under uv

- background.py: replace Popen.terminate()/kill() with psutil-based
  process tree walking on Windows. TerminateProcess does NOT cascade
  to grandchildren; psutil.Process.children(recursive=True) ensures
  the entire tree is signaled.

- test_backends.py: replace sys.executable with 'python' in sandbox
  execute() calls. Under uv, sys.executable is under the workspace
  and gets rewritten to ./ by prepare_sandbox_command, breaking
  Linux CI. The plain 'python' command resolves correctly in any
  activated venv.

* fix: broaden try/except in _kill_process_tree to cover proc.children()

If the process exits between Process(popen.pid) and children(recursive=True),
the children call raises an uncaught exception escaping stop(). Move it inside
the existing try/except block.

* fix: narrow exception to ProcessLookupError in POSIX _kill_process_tree

OSError is too broad — would silently swallow EPERM on SIGKILL, leaving
the process alive when we report it as stopped. Match original behavior
which only caught ProcessLookupError (process already gone).

* style: ruff format test_backends.py

* ci: trigger re-run for flaky prompt_toolkit test

* style: fix ruff check (import order + RUF005 unpacking)

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-10 15:43:18 +01:00
dinos 65828e9666 perf(cli): cut startup latency and defer MCP loading to the background (#171)
* perf(cli): cut startup time of `evosci --help` from ~2.2s to ~0.3s

Module-level imports were eagerly pulling in langchain.chat_models (with
the whole anthropic/openai/google stack), langgraph, textual, and
prompt_toolkit on every invocation — even for `--help` or `config list`.

Defer those with PEP 562 `__getattr__`, using `lazy_loader.attach` (SPEC-1,
the scientific-python standard) where it's a clean attach pattern:

- `EvoScientist/llm/__init__.py`: attach `models` lazily so importing
  `context_window` from this package no longer drags in langchain.
- `EvoScientist/stream/__init__.py`: attach display/events lazily; split
  the shared Rich `Console` singleton into a new lightweight
  `stream/console.py` so callers that only need `console` skip the
  `stream.events` → `langchain_core.messages` chain.
- `EvoScientist/cli/__init__.py`: hand-rolled `__getattr__` (reaches into
  `..stream.state`, which `lazy_loader` doesn't cover) so `commands` and
  `app` are the only eager loads.
- `EvoScientist/cli/commands.py`: move `cmd_interactive`/`cmd_run` to
  in-function imports so prompt_toolkit + textual only load when the
  interactive path actually runs.
- `EvoScientist/cli/_constants.py`: read `AGENT_NAME` on demand so
  `build_metadata` doesn't eagerly import `sessions` (langgraph/aiosqlite).

Adds `lazy-loader>=0.5` as a dependency.

* feat(cli): defer MCP tool loading with live per-server progress

The CLI was blocking ~5 s on MCP tool enumeration before the first
prompt appeared. Move the agent construction off the event loop and
surface per-server progress so the user can interact immediately and see
what's happening.

MCP client:
- Add an `on_progress` callback to `load_mcp_tools` / `aload_mcp_tools`
  / `_load_tools` emitting `start` / `success` / `error` events per
  server.
- Fan connection attempts out with `asyncio.gather` so latency no longer
  scales linearly with server count; cap simultaneous attempts at
  `_MAX_CONCURRENT_CONNECTIONS` (8) via a semaphore so a big stdio fleet
  doesn't spawn every subprocess at once.

Agent wiring:
- Plumb `on_mcp_progress` through `create_cli_agent` / `_load_agent` /
  `load_mcp_and_build_kwargs` so CLI and TUI can plug in collectors.

CLI (`cmd_interactive`):
- Run `_load_agent` in a background thread via `asyncio.to_thread`; the
  prompt and banner render immediately.
- `_await_agent_ready()` awaits the task before each agent-using site
  (first turn, channel messages, `/channel`, `/compact`). Raises if
  called without a prior `_start_agent_load` instead of silently
  reloading without the SQLite checkpointer.
- Pre-prime the progress dict from `load_mcp_config()` so the
  bottom-toolbar's `N/M` denominator is stable from the first render.
- Wrap `session.prompt_async` in `patch_stdout(raw=True)` so
  `console.print` from the worker-thread progress callback lands cleanly
  above the prompt as inline chat messages instead of stomping the
  prompt cursor.

TUI (`EvoTextualInteractiveApp`):
- Same background load + `_await_agent_ready()` gates on every
  `self._agent` read.
- New `MCPLoaderWidget` mounted at the top of `#input-shell` shows a
  header with `N/M` and one live row per server (spinner → ✓ / ✗ with
  tool count or error detail). On completion:
  - all-clean loads auto-dismiss ~2.5 s later;
  - cache hits (no events ever fired) dismiss immediately rather than
    flashing a misleading "0/N loaded";
  - failures keep the widget mounted so the user can read the errors.
  - `dismissed` property lets the app clear its ref so late events from
    slow servers become no-ops. The error branch of `_on_agent_loaded`
    also settles the widget so a load failure can't leave the spinner
    animating forever.
- Chat input is `disabled` while MCP resolves — no placeholder hack, no
"waiting…" system message.

Shared:
- Hoist braille spinner frames to `status_bar.SPINNER_FRAMES` and import
  them in the TUI widget so CLI and TUI animate in sync.

Tests:
- Extend `test_agent_mcp_cache` fakes to accept the new `on_progress`
  kwarg.
- New `TestLoadToolsProgressCallback` in `test_mcp_client` exercises the
  event sequence for success/failure/mixed fleets, verifies a buggy
  callback doesn't break the load, and asserts the semaphore caps
  in-flight connections.

* style: ruff

* chore: update uv.lock

* chore: uv.lock

* fix: coderabbit issues

* style: fmt

* fix: move _await_agent_ready inside try block

* fix(tui): auto-dismiss MCP loader widget on failure

The widget was designed to stay mounted on failure so the user could
read error detail, but since it's pinned above the input it never went
away in practice — just permanent banner clutter.

Auto-dismiss on failure too, with a longer grace (12s vs 2.5s) so the
error summary stays readable.

* fix: address second coderabbit pass

- Channel handlers (CLI + TUI): catch agent-load failures so the
  channel request doesn't hang; CLI moves `_await_agent_ready()`
  inside the existing try/except, TUI catches explicitly and calls
  `_set_channel_response` with the error.

- Stale background loads: `prev.cancel()` only stops the asyncio
  wrapper, not the thread running `_load_agent`. Added a generation
  token (`agent_load_id` / `self._agent_load_id`) and gated both
  progress and completion callbacks on it so a superseded load can't
  clobber the current session's state or UI.

- TUI prompt lifecycle: added `_agent_load_pending()` and gated the
  `_process_channel_message` / `_handle_command` finally blocks on it
  so `/new` or `/resume` invoked from a command keeps the prompt
  disabled until the fresh load settles.

- TUI readiness failures: `_run_turn` and `_handle_command` now
  catch exceptions from `_await_agent_ready()` and surface a
  "Agent failed to load: …" system message instead of letting the
  exception escape into Textual's traceback panel.

* refactor(cli): share background agent loader between CLI and TUI

The CLI and TUI were carrying near-identical copies of the same
background-load state machine: the `agent_task`, the `agent_load_id`
generation token, the gated progress/completion callbacks, and the
per-server progress dict. Every CodeRabbit finding on that lifecycle
had to be fixed in both files.

Extract it into `cli/_agent_loader.py`:

- `MCPProgressTracker` — owns the `server -> (state, detail)` dict;
  exposes `prime`, `record`, `snapshot`, `totals`.

- `BackgroundAgentLoader` — owns `agent`, the in-flight task, and the
  generation token. Exposes `start(**loader_kwargs)`, `await_ready()`,
  `is_pending`. Internally gates all progress/completion callbacks by
  generation so a superseded load can't clobber the current session.
  UI-specific rendering plugs in via `on_progress` / `on_success` /
  `on_failure` callbacks.

Both surfaces now just wire their UI hooks; the loader file holds no
Rich / prompt_toolkit / Textual dependencies. Net -345 lines from
`interactive.py` + `tui_interactive.py`; +20 unit tests pinning the
lifecycle (generation filtering, cache-hit short-circuit, failure
reset, progress ordering).

* refactor(cli): make _on_done the sole authority for agent state transitions

await_ready no longer sets self.agent — it just awaits the task and
reads what _on_done already wrote. Eliminates the dual-write overlap
(asyncio guarantees done-callbacks fire in registration order).

* fix(tui): let users type during MCP load, only block on send

Remove prompt-disabling during background agent load — the TUI now
matches the CLI approach where the input stays enabled and only gates
on await_ready() at submit time. The MCPLoaderWidget still provides
visual feedback that loading is in progress.

* fix(loader): preserve real load error on await_ready; dedup failure message

CodeRabbit flagged two issues with the new loader:

1. After a failed load, `_on_done` nulled `self._task`, so the next
   `await_ready()` hit the "before start()" branch and the CLI wrapper
   remapped it to a misleading "checkpointer not available" message —
   losing the real exception (bad MCP config, network, etc.).

   Keep `_task` set on failure so `await_ready` re-raises the real
   exception. Added `needs_restart` so TUI's auto-retry check stays a
   one-liner and doesn't need to reach into task internals.

2. TUI reported each load failure twice: once from
   `_on_agent_load_failure` (the done-callback) and once from each
   caller of `_await_agent_ready` (`_run_turn`,
   `_process_channel_message`, `_handle_command`) catching the re-raise.

   `_on_agent_load_failure` is now the sole local reporter; callers
   just handle control flow (return cleanly, set channel response to
   unblock remote).

* fix(cli): wire /model handler through the agent loader

The /model command from main (merged via f1f0d7c) still reached for
`state["agent"]` (CLI) and `self._agent` (TUI) — both removed by the
background-loader refactor. CLI raised KeyError on first invocation;
TUI raised AttributeError. Writes to the old fields also had no effect
because every other code path now reads from `agent_loader.agent`, so
the model switch would have silently failed.

Route everything through the loader: `await _await_agent_ready()` up
front so /model doesn't race with the initial background load, build
the `CommandContext` with the current agent, and sync `ctx.agent` back
into `agent_loader.agent` (plus channel globals) when the command
replaces it.

* fix(cli): isolate progress callback, capture awaited agent, gate by requires_agent

Three CodeRabbit findings on the loader + command dispatch path:

- Wrap ``_on_progress`` in try/except inside the loader's gated wrapper
  so a buggy UI adapter can't bubble into ``loader_fn`` and fail the
  whole background load. The MCP client already protects this, but
  defence-in-depth keeps the loader self-contained.

- In CLI channel + main-loop streaming, capture the agent returned by
  ``_await_agent_ready()`` and pass that into ``run_streaming`` rather
  than reading ``agent_loader.agent`` after a subsequent ``await``.
  A concurrent ``/new``/``/resume``/``/model`` could have swapped it.

- Add ``requires_agent: ClassVar[bool] = False`` to ``Command`` and
  mark ``/compact``, ``/model``, ``/channel`` as ``True``. TUI dispatch
  sites (channel and keyboard) now check ``cmd_manager.resolve(...)``
  and only wait for readiness when the command actually needs the
  agent. ``/mcp add``, ``/skills``, ``/new`` etc. no longer deadlock
  behind a failing MCP load they are meant to fix.

* fix(cli): guard sync-back, subcommand-aware gating, /model adopt-path

Three CodeRabbit findings on command dispatch:

- ``_handle_command`` unconditionally synced ``ctx.agent`` back into
  ``agent_loader``.  For non-agent commands ``ctx.agent`` is ``None``,
  so ``/threads`` / ``/mcp`` / ``/skills`` (etc.) could clobber a valid
  loaded agent — and rebind channel globals to ``None``.  Guard the
  sync on ``ctx.agent is not None``.

- ``/channel status`` and ``/channel stop`` don't touch ``ctx.agent``
  but the class-level ``requires_agent = True`` blocked them behind
  agent readiness.  Added ``Command.needs_agent(args)`` (defaults to
  ``requires_agent``) so ``/channel`` can override with subcommand
  awareness; kept the class flag for the common case.

- ``/model`` builds a new agent from scratch, it never reads the
  existing one — gating it on readiness meant a broken provider
  blocked the command that would fix it.  Flipped it to
  ``requires_agent = False`` and added ``BackgroundAgentLoader.adopt``
  so the UI can seat the replacement and supersede any in-flight
  load (the generation token keeps a late completion from clobbering
  the adopted agent).

Bonus cleanup: ``CommandManager.resolve`` now returns
``(command, args)`` so callers can invoke ``needs_agent`` without
re-implementing ``shlex`` parsing.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-22 18:22:33 +02:00
dinos 05f54334ba fix(mcp): stdio env passthrough + durable package installs (#169)
* fix(mcp): forward proxy and CA bundle env vars to stdio subprocesses

The MCP SDK's stdio transport inherits only a minimal allowlist (HOME,
PATH, USER, …) from the parent, stripping http_proxy/https_proxy and
SSL_CERT_FILE/REQUESTS_CA_BUNDLE/etc. Behind a proxy or with a custom CA
bundle, stdio MCP servers silently hang on outbound requests while the
same server over HTTP transport works. Auto-forward the proxy and cert
vars when present; user-configured env still takes precedence.

* fix(mcp): use `uv tool install` so MCP packages survive uv sync

Source installs previously used `uv pip install --python $VENV <pkg>`,
which lands in the evosci venv but is not recorded in pyproject.toml or
uv.lock. A subsequent `uv sync` (typical after `git pull`) reconciles
the venv to the lockfile and removes the MCP package, forcing users to
re-run onboard.

Prefer `uv tool install <pkg>` for the non-uv-tool install path: the
binary symlink in ~/.local/bin survives uv sync and evosci upgrades,
and the MCP server gets its own isolated env (no dep conflicts).
Verify the expected CLI entry point resolves afterward; if not (package
has no console-script), fall through to the old uv-pip path so
command-less packages still work.

The uv-tool-env path (`uv tool install evoscientist --with <pkg>`) is
unchanged — it was already durable via uv's receipt.

* fix(mcp): gate standalone uv tool install on verify_command

Previously `install_pip_package` would route every install through
`uv tool install <pkg>` when `verify_command` was None, returning
success as long as the uv subprocess exited 0. Library callers
(`evoscientist[oauth]`, `lark-oapi`, etc.) expect the package to land
in the active venv so they can import it — a standalone uv tool env
is not importable, so the import fails at the next line.

Gate the `uv tool install <pkg>` branch on `verify_command` being
set: that signals the caller wants a durable CLI binary, which is
what `uv tool install` produces. Library callers omit it and go
straight to the pip-install-into-venv path.

Also: log info messages on every fall-through so stale-binary and
entry-point-missing failure modes are debuggable, and document the
--with → standalone recovery path.

* fix(mcp): resolve MCP binaries to `uv tool dir --bin`, not `.venv/bin`

Under `uv run`, the project venv's `bin/` comes first on PATH, so
`shutil.which("arxiv-mcp-server")` returns a stale `.venv/bin/` copy
left over from an earlier install instead of the fresh symlink that
`uv tool install` just placed in `~/.local/bin`. The venv copy gets
written to mcp.yaml and is then wiped by the next `uv sync` — exactly
the failure mode the durability fix was meant to prevent.

Query `uv tool dir --bin` directly and prefer binaries found there
over `shutil.which`. Same change to the post-install verify in
`install_pip_package` so a venv shadow can't falsely short-circuit
the fallback.

* refactor(mcp): split install_pip_package into install_library + install_cli_tool

`verify_command` was doing double duty: naming the CLI binary to check
*and* signaling "this is a CLI install, use the standalone `uv tool
install` path." Callers routed library installs through the CLI branch
any time they forgot to pass it, and the resulting standalone uv tool
env wasn't importable from the active venv.

Separate the two use cases into distinct functions, each with one
install strategy per environment shape. Shared logic lives in private
`_install_with_uv_tool_env` / `_install_via_pip` helpers.

- install_library(pkg): uv-tool-env --with → pip. Never uses standalone
  `uv tool install <pkg>` (not importable from active venv).
- install_cli_tool(pkg, *, verify_command): uv-tool-env --with →
  standalone `uv tool install` → pip. `verify_command` is now required.

Callers pick the right function at the call site: registry.py picks
based on whether `entry.command` is set; onboard.py call sites all
install libraries.
2026-04-21 16:59:06 +01:00
dinos b9e809aeb6 fix: use uv tool install --with for durable MCP server installs (#125)
* fix: use `uv tool install --with` for durable MCP server installs (#121)

When EvoScientist is installed via `uv tool install`, MCP server packages
added during onboarding were installed with `uv pip install`, which is
not tracked by uv. Running `uv tool upgrade evoscientist` would recreate
the venv from scratch and silently wipe the MCP server binaries.

Now `install_pip_package()` detects uv tool environments and uses
`uv tool install <tool> --with <package>`, which records the dependency
in uv-receipt.toml so it survives upgrades. Existing --with packages
are read from the receipt and preserved.

Falls back to the old `uv pip install` path if the durable method fails.

* style: fmt

* fix: preserve requirement specs and normalize dedup in uv tool installs

Address review feedback: _uv_tool_existing_requirements() now returns
a dict mapping bare names to full PEP 508 specs (preserving extras and
version constraints from uv-receipt.toml). Dedup check uses
_bare_package_name() to normalize the incoming package argument before
comparing against receipt entries.
2026-04-02 10:51:54 +02:00
Xi Zhang fab5f85eee v0.0.4 (#93)
* feat(tui): enhance conversation history rendering and implement two-level thread hierarchy in picker

* feat(tui): improve conversation history display and enhance thread selection UI

* feat(file_mentions): implement @file mention parsing and completion for CLI and TUI

* feat(uv-tool): add compatibility checks and installation helpers for uv tool environments

* feat(dependencies): update package versions in uv.lock for compatibility and improvements

* feat(badges): update PyPI version to v0.0.4 in SVG assets and README files

* feat(tests): format code in TestUvToolCompat for improved readability
2026-03-24 18:13:42 +00:00
Jan Piotrowski 2fdf961fee chore: fix ruff linting and async patterns in tests/ 2026-03-19 17:04:05 +01:00
Jan Piotrowski 4a3d6c0318 chore: add ruff lint rules and turn on formatting 2026-03-19 17:04:02 +01:00
X-iZhang 7089d1c179 feat: implement _resolve_command function for command path resolution 2026-03-17 23:05:05 +00:00
X-iZhang c5a4d559a2 Refactor test cases for improved readability and consistency
- Added blank lines for better separation of test cases in multiple test files.
- Reformatted event handling in tests for clarity and consistency.
- Ensured consistent use of multi-line formatting for dictionary arguments in event handling.
- Improved assertions and test descriptions for better understanding.
- Updated test cases across various modules including test_stream_state, test_stream_utils, test_summarization, test_thread_selector, test_tool_error_handler, test_tui_widgets, test_ui_runtime, and test_wechat_channel.
2026-03-15 21:12:51 +00:00
Dinos Papakostas 97ac3e0401 feat(mcp): support wildcards in tool filtering 2026-02-14 00:06:08 +08:00
X-iZhang 8404a8e013 fix: update allowed_senders handling and improve logging configuration 2026-02-10 22:20:56 +00:00
X-iZhang 250a23682c update 2026-02-09 02:54:41 +00:00
X-iZhang 7b15b3b8a0 update 2026-02-07 01:53:46 +00:00
Dinos Papakostas e07bcf2a91 test(mcp): add tests for MCP client 2026-02-06 23:17:45 +08:00