dinos 65828e9666 perf(cli): cut startup latency and defer MCP loading to the background (#171)
* perf(cli): cut startup time of `evosci --help` from ~2.2s to ~0.3s

Module-level imports were eagerly pulling in langchain.chat_models (with
the whole anthropic/openai/google stack), langgraph, textual, and
prompt_toolkit on every invocation — even for `--help` or `config list`.

Defer those with PEP 562 `__getattr__`, using `lazy_loader.attach` (SPEC-1,
the scientific-python standard) where it's a clean attach pattern:

- `EvoScientist/llm/__init__.py`: attach `models` lazily so importing
  `context_window` from this package no longer drags in langchain.
- `EvoScientist/stream/__init__.py`: attach display/events lazily; split
  the shared Rich `Console` singleton into a new lightweight
  `stream/console.py` so callers that only need `console` skip the
  `stream.events` → `langchain_core.messages` chain.
- `EvoScientist/cli/__init__.py`: hand-rolled `__getattr__` (reaches into
  `..stream.state`, which `lazy_loader` doesn't cover) so `commands` and
  `app` are the only eager loads.
- `EvoScientist/cli/commands.py`: move `cmd_interactive`/`cmd_run` to
  in-function imports so prompt_toolkit + textual only load when the
  interactive path actually runs.
- `EvoScientist/cli/_constants.py`: read `AGENT_NAME` on demand so
  `build_metadata` doesn't eagerly import `sessions` (langgraph/aiosqlite).

Adds `lazy-loader>=0.5` as a dependency.

* feat(cli): defer MCP tool loading with live per-server progress

The CLI was blocking ~5 s on MCP tool enumeration before the first
prompt appeared. Move the agent construction off the event loop and
surface per-server progress so the user can interact immediately and see
what's happening.

MCP client:
- Add an `on_progress` callback to `load_mcp_tools` / `aload_mcp_tools`
  / `_load_tools` emitting `start` / `success` / `error` events per
  server.
- Fan connection attempts out with `asyncio.gather` so latency no longer
  scales linearly with server count; cap simultaneous attempts at
  `_MAX_CONCURRENT_CONNECTIONS` (8) via a semaphore so a big stdio fleet
  doesn't spawn every subprocess at once.

Agent wiring:
- Plumb `on_mcp_progress` through `create_cli_agent` / `_load_agent` /
  `load_mcp_and_build_kwargs` so CLI and TUI can plug in collectors.

CLI (`cmd_interactive`):
- Run `_load_agent` in a background thread via `asyncio.to_thread`; the
  prompt and banner render immediately.
- `_await_agent_ready()` awaits the task before each agent-using site
  (first turn, channel messages, `/channel`, `/compact`). Raises if
  called without a prior `_start_agent_load` instead of silently
  reloading without the SQLite checkpointer.
- Pre-prime the progress dict from `load_mcp_config()` so the
  bottom-toolbar's `N/M` denominator is stable from the first render.
- Wrap `session.prompt_async` in `patch_stdout(raw=True)` so
  `console.print` from the worker-thread progress callback lands cleanly
  above the prompt as inline chat messages instead of stomping the
  prompt cursor.

TUI (`EvoTextualInteractiveApp`):
- Same background load + `_await_agent_ready()` gates on every
  `self._agent` read.
- New `MCPLoaderWidget` mounted at the top of `#input-shell` shows a
  header with `N/M` and one live row per server (spinner → ✓ / ✗ with
  tool count or error detail). On completion:
  - all-clean loads auto-dismiss ~2.5 s later;
  - cache hits (no events ever fired) dismiss immediately rather than
    flashing a misleading "0/N loaded";
  - failures keep the widget mounted so the user can read the errors.
  - `dismissed` property lets the app clear its ref so late events from
    slow servers become no-ops. The error branch of `_on_agent_loaded`
    also settles the widget so a load failure can't leave the spinner
    animating forever.
- Chat input is `disabled` while MCP resolves — no placeholder hack, no
"waiting…" system message.

Shared:
- Hoist braille spinner frames to `status_bar.SPINNER_FRAMES` and import
  them in the TUI widget so CLI and TUI animate in sync.

Tests:
- Extend `test_agent_mcp_cache` fakes to accept the new `on_progress`
  kwarg.
- New `TestLoadToolsProgressCallback` in `test_mcp_client` exercises the
  event sequence for success/failure/mixed fleets, verifies a buggy
  callback doesn't break the load, and asserts the semaphore caps
  in-flight connections.

* style: ruff

* chore: update uv.lock

* chore: uv.lock

* fix: coderabbit issues

* style: fmt

* fix: move _await_agent_ready inside try block

* fix(tui): auto-dismiss MCP loader widget on failure

The widget was designed to stay mounted on failure so the user could
read error detail, but since it's pinned above the input it never went
away in practice — just permanent banner clutter.

Auto-dismiss on failure too, with a longer grace (12s vs 2.5s) so the
error summary stays readable.

* fix: address second coderabbit pass

- Channel handlers (CLI + TUI): catch agent-load failures so the
  channel request doesn't hang; CLI moves `_await_agent_ready()`
  inside the existing try/except, TUI catches explicitly and calls
  `_set_channel_response` with the error.

- Stale background loads: `prev.cancel()` only stops the asyncio
  wrapper, not the thread running `_load_agent`. Added a generation
  token (`agent_load_id` / `self._agent_load_id`) and gated both
  progress and completion callbacks on it so a superseded load can't
  clobber the current session's state or UI.

- TUI prompt lifecycle: added `_agent_load_pending()` and gated the
  `_process_channel_message` / `_handle_command` finally blocks on it
  so `/new` or `/resume` invoked from a command keeps the prompt
  disabled until the fresh load settles.

- TUI readiness failures: `_run_turn` and `_handle_command` now
  catch exceptions from `_await_agent_ready()` and surface a
  "Agent failed to load: …" system message instead of letting the
  exception escape into Textual's traceback panel.

* refactor(cli): share background agent loader between CLI and TUI

The CLI and TUI were carrying near-identical copies of the same
background-load state machine: the `agent_task`, the `agent_load_id`
generation token, the gated progress/completion callbacks, and the
per-server progress dict. Every CodeRabbit finding on that lifecycle
had to be fixed in both files.

Extract it into `cli/_agent_loader.py`:

- `MCPProgressTracker` — owns the `server -> (state, detail)` dict;
  exposes `prime`, `record`, `snapshot`, `totals`.

- `BackgroundAgentLoader` — owns `agent`, the in-flight task, and the
  generation token. Exposes `start(**loader_kwargs)`, `await_ready()`,
  `is_pending`. Internally gates all progress/completion callbacks by
  generation so a superseded load can't clobber the current session.
  UI-specific rendering plugs in via `on_progress` / `on_success` /
  `on_failure` callbacks.

Both surfaces now just wire their UI hooks; the loader file holds no
Rich / prompt_toolkit / Textual dependencies. Net -345 lines from
`interactive.py` + `tui_interactive.py`; +20 unit tests pinning the
lifecycle (generation filtering, cache-hit short-circuit, failure
reset, progress ordering).

* refactor(cli): make _on_done the sole authority for agent state transitions

await_ready no longer sets self.agent — it just awaits the task and
reads what _on_done already wrote. Eliminates the dual-write overlap
(asyncio guarantees done-callbacks fire in registration order).

* fix(tui): let users type during MCP load, only block on send

Remove prompt-disabling during background agent load — the TUI now
matches the CLI approach where the input stays enabled and only gates
on await_ready() at submit time. The MCPLoaderWidget still provides
visual feedback that loading is in progress.

* fix(loader): preserve real load error on await_ready; dedup failure message

CodeRabbit flagged two issues with the new loader:

1. After a failed load, `_on_done` nulled `self._task`, so the next
   `await_ready()` hit the "before start()" branch and the CLI wrapper
   remapped it to a misleading "checkpointer not available" message —
   losing the real exception (bad MCP config, network, etc.).

   Keep `_task` set on failure so `await_ready` re-raises the real
   exception. Added `needs_restart` so TUI's auto-retry check stays a
   one-liner and doesn't need to reach into task internals.

2. TUI reported each load failure twice: once from
   `_on_agent_load_failure` (the done-callback) and once from each
   caller of `_await_agent_ready` (`_run_turn`,
   `_process_channel_message`, `_handle_command`) catching the re-raise.

   `_on_agent_load_failure` is now the sole local reporter; callers
   just handle control flow (return cleanly, set channel response to
   unblock remote).

* fix(cli): wire /model handler through the agent loader

The /model command from main (merged via f1f0d7c) still reached for
`state["agent"]` (CLI) and `self._agent` (TUI) — both removed by the
background-loader refactor. CLI raised KeyError on first invocation;
TUI raised AttributeError. Writes to the old fields also had no effect
because every other code path now reads from `agent_loader.agent`, so
the model switch would have silently failed.

Route everything through the loader: `await _await_agent_ready()` up
front so /model doesn't race with the initial background load, build
the `CommandContext` with the current agent, and sync `ctx.agent` back
into `agent_loader.agent` (plus channel globals) when the command
replaces it.

* fix(cli): isolate progress callback, capture awaited agent, gate by requires_agent

Three CodeRabbit findings on the loader + command dispatch path:

- Wrap ``_on_progress`` in try/except inside the loader's gated wrapper
  so a buggy UI adapter can't bubble into ``loader_fn`` and fail the
  whole background load. The MCP client already protects this, but
  defence-in-depth keeps the loader self-contained.

- In CLI channel + main-loop streaming, capture the agent returned by
  ``_await_agent_ready()`` and pass that into ``run_streaming`` rather
  than reading ``agent_loader.agent`` after a subsequent ``await``.
  A concurrent ``/new``/``/resume``/``/model`` could have swapped it.

- Add ``requires_agent: ClassVar[bool] = False`` to ``Command`` and
  mark ``/compact``, ``/model``, ``/channel`` as ``True``. TUI dispatch
  sites (channel and keyboard) now check ``cmd_manager.resolve(...)``
  and only wait for readiness when the command actually needs the
  agent. ``/mcp add``, ``/skills``, ``/new`` etc. no longer deadlock
  behind a failing MCP load they are meant to fix.

* fix(cli): guard sync-back, subcommand-aware gating, /model adopt-path

Three CodeRabbit findings on command dispatch:

- ``_handle_command`` unconditionally synced ``ctx.agent`` back into
  ``agent_loader``.  For non-agent commands ``ctx.agent`` is ``None``,
  so ``/threads`` / ``/mcp`` / ``/skills`` (etc.) could clobber a valid
  loaded agent — and rebind channel globals to ``None``.  Guard the
  sync on ``ctx.agent is not None``.

- ``/channel status`` and ``/channel stop`` don't touch ``ctx.agent``
  but the class-level ``requires_agent = True`` blocked them behind
  agent readiness.  Added ``Command.needs_agent(args)`` (defaults to
  ``requires_agent``) so ``/channel`` can override with subcommand
  awareness; kept the class flag for the common case.

- ``/model`` builds a new agent from scratch, it never reads the
  existing one — gating it on readiness meant a broken provider
  blocked the command that would fix it.  Flipped it to
  ``requires_agent = False`` and added ``BackgroundAgentLoader.adopt``
  so the UI can seat the replacement and supersede any in-flight
  load (the generation token keeps a late completion from clobbering
  the adopted agent).

Bonus cleanup: ``CommandManager.resolve`` now returns
``(command, args)`` so callers can invoke ``needs_agent`` without
re-implementing ``shlex`` parsing.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-22 18:22:33 +02:00
2026-03-11 11:09:44 +00:00
2026-04-18 16:49:31 +01:00
2026-04-18 16:49:31 +01:00

EvoScientist Logo

Typing SVG

English | 简体中文

EvoScientist aims to harness vibe research by enabling self-evolving AI scientists that autonomously explore, generate insights, and iteratively improve. It is designed to be opinionated and ready to use out of the box, offering a living research system that grows alongside evolving agent skills, toolsets, and memory bases. Moving beyond traditional human-in-the-loop systems, EvoScientist adopts a human-on-the-loop paradigm, where AI acts as a research buddy that co-evolves with human researchers and internalizes scholarly taste and scientific judgment.

🏆 Awards & Recognition

ICAIS 2025 Awards
Best Paper & Appraisal Award
Best Paper
AI-Generated Best Paper
DeepResearch Bench II #1
#1 on DeepResearch Bench II

DeepResearch Bench #1
#1 on DeepResearch Bench
AstaBench Code & Execution #1
#1 on AstaBench Code & Execution
AstaBench Data Analysis #1
#1 on AstaBench Data Analysis

⚡ Unified Control, Different Surfaces

🖥️ CLI / TUI

📱 Mobile

✨ Features

  • 🤖 Multi-Agent Team — 6 sub-agents (plan, research, code, debug, analyze, write) working in concert.
  • 🧠 Persistent Memory — Context, preferences, and findings survive across sessions.
  • 🌐 Multi-Provider — Anthropic, OpenAI, Google, MiniMax, NVIDIA — one config to switch.
  • 📱 Multi-Channel — CLI as the hub; Telegram, Slack, Feishu, WeChat, and more — one agent session.
  • 🔬 Scientific Workflow — Intake → plan → execute → evaluate → write → verify.
  • 🔄 Code Generation Modes — More Effort (iterative refinement), continuously improving code quality.
  • ⚡ Adaptive Tools — Per-turn tool selection keeps only relevant tools visible, reducing noise.
  • ✂️ Context Editing — Dynamic system prompt rewriting based on conversation state.
  • 🔌 MCP & Skills — Plug in MCP servers or install skills from GitHub on the fly.

Tip

Looking for ready-to-use research skills? Check out EvoSkills — powered by EvoScientist's engine and installable skills, the entire end-to-end research lifecycle is covered out of the box. EvoSkills are also compatible with other CLI coding agents.

🔥 News

📖 Table of Contents

📦 Installation

Tip

Requires Python 3.11+ (< 3.14). We recommend uv or conda for dependency management and virtual environments.

🪛 Install uv (if you don't have it)
curl -LsSf https://astral.sh/uv/install.sh | sh

Quick Install

uv tool install EvoScientist

Note

To update an existing installation to the latest version, use uv tool upgrade:

uv tool upgrade EvoScientist

Or install into the current environment instead:

uv pip install EvoScientist

Latest from GitHub

To get the latest patches before a PyPI release:

uv pip install git+https://github.com/EvoScientist/EvoScientist.git

Development Install

git clone https://github.com/EvoScientist/EvoScientist.git
cd EvoScientist
uv sync --dev

enable pre-commit hooks:

uv run pre-commit install
Using conda
conda create -n EvoSci python=3.11 -y
conda activate EvoSci
pip install -e ".[dev]"
Using PyPi
pip install EvoScientist          # quick install
pip install -e ".[dev]"           # development install
Optional: Channel dependencies

Messaging channel integrations require extra dependencies. Install only what you need:

uv pip install "EvoScientist[telegram]"     # Telegram
uv pip install "EvoScientist[discord]"      # Discord
uv pip install "EvoScientist[slack]"        # Slack
uv pip install "EvoScientist[wechat]"       # WeChat
uv pip install "EvoScientist[qq]"           # QQ
uv pip install "EvoScientist[feishu]"       # Feishu
uv pip install "EvoScientist[all-channels]" # everything
Upgrade to the latest code base
git pull && uv sync --dev

🔝Back to top

🔑 Configuration

The easiest way to configure API keys is the interactive wizard:

EvoSci onboard

Tip

It walks you through provider selection, key validation, model choice, and workspace mode. Supports OAuth sign-in for CLI coding agent subscribers — no API key needed.

onboard

📟 Manual configuration via environment variables

Set at least one LLM provider key and (optionally) a search key:

# Pick one LLM provider
export ANTHROPIC_API_KEY="sk-..."   # Claude  — console.anthropic.com
export OPENAI_API_KEY="sk-..."      # GPT    — platform.openai.com
export GOOGLE_API_KEY="AI..."       # Gemini  — aistudio.google.com/api-keys
export MINIMAX_API_KEY="sk-..."     # MiniMax — platform.minimaxi.com (China, default) or platform.minimax.io (Global)
export MINIMAX_BASE_URL="https://api.minimax.io/anthropic"  # only needed for Global keys (default: https://api.minimaxi.com/anthropic)
export NVIDIA_API_KEY="nvapi-..."   # NIM    — build.nvidia.com

# Web search (optional)
export TAVILY_API_KEY="tvly-..."    # app.tavily.com

Or use EvoSci config set to persist keys in ~/.config/evoscientist/config.yaml.

Alternatively, copy the example .env file for project-level configuration:

cp .env.example .env  # then fill in your keys

⚠️ Never commit .env files with real keys. It is already in .gitignore.

🔝Back to top

⚡ Quick Start

EvoSci  # or EvoScientist — interactive mode (TUI by default)

demo

Run EvoSci -h for all CLI options.

cli help

Tip

Need to copy long outputs? Use --ui cli for classic mode where native terminal copy works freely. On macOS, iTerm2 users can also hold ⌥ Option while dragging to select, then ⌘+C.

Common examples
EvoSci                            # interactive mode (TUI by default)
EvoSci -p "your question"        # single-shot mode
EvoSci --workdir /path/to/project # open in a specific directory
EvoSci -m run                     # isolated per-session workspace
EvoSci --ui cli                   # classic CLI (lightweight)
EvoSci serve                      # headless mode — channels only, no interactive prompt
Action Approval

By default, shell commands (execute tool) require human approval before running. To skip approval prompts:

# Per-session: auto-approve via CLI flag
EvoSci --auto-approve
EvoSci -p "query" --auto-approve

# Persistent: set in config (applies to all future sessions)
EvoSci config set auto_approve true

# Or allow only specific command prefixes
EvoSci config set shell_allow_list "python,pip,pytest,ruff,git"

During a session you can also reply 3 (Approve all) at any approval prompt to auto-approve for the rest of that session.

Agent Questions

The agent can proactively ask you questions when it needs clarification (e.g., dataset choice, experiment direction). This is enabled by default. To disable:

# Persistent: set in config
EvoSci config set enable_ask_user false

# Re-enable
EvoSci config set enable_ask_user true
In-session commands
Command Description
/current Show current session info
/threads List recent sessions
/resume Resume a previous session
/delete Delete a saved session
/new Start a new session
/clear Clear chat history
/skills List installed skills
/install-skill <src> Add a skill from path or GitHub
/uninstall-skill <name> Remove an installed skill
/mcp Manage MCP servers
/channel Configure messaging channels
/help Show available commands
/exit Quit
Script Inference
from EvoScientist import EvoScientist_agent
from langchain_core.messages import HumanMessage
from EvoScientist.utils import format_messages

thread = {"configurable": {"thread_id": "1"}}
last_len = 0

for state in EvoScientist_agent.stream(
    {"messages": [HumanMessage(content="Hi?")]},
    config=thread,
    stream_mode="values",
):
    msgs = state["messages"]
    if len(msgs) > last_len:
        format_messages(msgs[last_len:])
        last_len = len(msgs)

🔝Back to top

🍪 Examples & Recipes

A curated collection of official examples, advanced usage patterns, and community-contributed recipes to help you get the most out of EvoScientist.

👉 Browse all examples & recipes

🔝Back to top

🔌 MCP Integration

Add external tools via MCP servers with a single command:

# Usage
EvoSci mcp add <name> <command> [-- args...]

# Example
EvoSci mcp add sequential-thinking npx -- -y @modelcontextprotocol/server-sequential-thinking

Tip

For command options, config fields, tool routing, wildcard filtering, and troubleshooting, see the MCP Integration Guide.

🔝Back to top

📱 Channels

Connect messaging platforms so they share the same agent session as the CLI:

# Usage
EvoSci channel setup <channel>

# Example
EvoSci channel setup telegram

Multiple channels can run concurrently — comma-separate names in the config:

channel_enabled: "telegram,slack,feishu,qq"

The channel can also be started interactively with /channel in the CLI session.

Tip

For per-channel setup guides, capability matrix, architecture details, and troubleshooting, see the Channel Integration Guide.

🔝Back to top

📚 Acknowledgments

This project builds upon the following outstanding open-source works:

  • LangChain — A framework for building agents and LLM-powered applications.
  • DeepAgents — The batteries-included agent harness.

We thank the authors for their valuable contributions to the open-source community.

🔝Back to top

🎯 ᯓ➤ Roadmap

Coming soon:

  • 🖥️ Full-screen TUI and classic CLI interfaces
  • 📻 EvoMemory v1.0 shipped
  • ⚒️ 200+ predefined skills built in
  • 🧩 Built-in research-lifecycle skills shipped
  • 👋 Human-in-the-loop action approval
  • 🦾 Agent-initiated human clarification
  • 📑 Technical report on the way
  • 🔐 OAuth sign-in (CLI coding agent subscribers)
  • 📺 Web app with workspace UI
  • 📹 Demo and tutorial in the works
  • 📊 Benchmark suite to be released
  • ⏰ Scheduled tasks for the core system planned

Stay tuned — more features are on the way!

🔝Back to top

🌍 Project Roles

Core Contributors

Xi Zhang
Xi Zhang
Yougang Lyu
Yougang Lyu
Dinos Papakostas
Dinos Papakostas
Yuyue Zhao
Yuyue Zhao
Ziheng Zhang
Ziheng Zhang
Xiaohui Yan
Xiaohui Yan

Contributors

Jan Piotrowski, Wiktor Cupiał, Jakub Kaliski, Jakub Filipiuk, Xinhao Yi, Shuyu Guo, Andreas Sauter, Wenxiang Hu, Jacopo Urbani, Zaiqiao Meng, Jun Luo, Lun Zhou

Xiaoyi DeepResearch Xiaoyi DeepResearch Team and the wider open-source community contribute to this project.

For any inquiries or collaboration opportunities, please contact: EvoScientist.ai@gmail.com

🔝Back to top

🤝 Contributing

EvoScientist Team

We welcome contributions from developers, researchers, and AI coding agents at all levels. Our Contributing Guidelines are designed for both humans and AI agents — covering architecture, patterns, extension guides, and code standards to help you contribute safely and effectively.

👥 Community Contributors

⚗️ Join the EvoScientist community to discuss AI-driven research, share experiment results, and help shape the future of automated scientific discovery.

  • Discord — Ask questions, share findings, and collaborate with researchers and developers in real-time.

  • WeChat — Connect with our Chinese-speaking research community.

    WeChat QR Code

Every contribution brings us one step closer to a future where AI accelerates scientific breakthroughs for all of humanity.

📈 Star History

Star History Chart

🔝Back to top

📝 Citation

If you find our paper and code useful in your research and applications, please cite using this BibTeX:

@article{evoscientist2026, 
  title={EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery}, 
  author={Yougang Lyu and Xi Zhang and Xinhao Yi and Yuyue Zhao and Shuyu Guo and Wenxiang Hu and Jan Piotrowski and Jakub Kaliski and Jacopo Urbani and Zaiqiao Meng and Lun Zhou and Xiaohui Yan}, 
  journal={arXiv preprint arXiv:2603.08127}, 
  year={2026} 
}

🔝Back to top

📜 License

This project is licensed under the Apache License 2.0 - see the LICENSE file for details.

🔝Back to top

S
Description
No description provided
Readme 70 MiB
v0.2.3 Latest
2026-08-13 11:03:15 +08:00
Languages
Python 98.8%
HTML 1.2%