befdb17e0b11e55d31ad2e11cb300f59a7af7987
14 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3cda9894c7 |
feat: agent-teams part C - expert selection UX (depends on part B) (#371)
* feat: bias main-agent delegation toward configurable.active_teams * feat: add /experts and /expert TUI commands for expert-skill summoning * feat: align expert-selection wording with WebUI (invite/dismiss) * chore: clear active_teams on /new, cleanup active_team.py comment * fix: use local append_to_system_message in ActiveTeamMiddleware * fix: prevent configurable_extra from overriding thread_id * fix: cache expert-skill lookup for /expert completions * fix: suppress /expert completions past first arg and on exact match * fix: invalidate /expert completion cache on skill install/uninstall * fix: refuse /expert invites for non-dispatchable expert skills * fix: fire /expert cache invalidation on every install_skill / uninstall_skill path * fix: propagate active_teams to Rich CLI and serve dispatch surfaces * fix: keep invited experts across channel shutdown * fix: match /expert completions case-insensitively |
||
|
|
8b1451cdda |
refactor(runtime): centralize async bridges under an owned runtime (#376)
* feat(runtime): add application-scoped async runtime * refactor(cli): use owned runtime for session stats * refactor(onboard): use the owned async runtime * docs(runtime): record async bridge ownership * refactor(middleware): keep sync fallback synchronous * refactor(mcp): load tools on an owned runtime * refactor(cli): share owned runtime across entry points * refactor(channels): make inbound sync bridge explicit * refactor(stream): run Rich streaming on owned runtime * chore(runtime): remove nest-asyncio dependency * refactor(asyncio): require active loops in async code * docs(runtime): document final event loop ownership * fix(stream): cancel stalled owned streams * fix(cli): recover cleanly from stream cancellation * fix(runtime): drain executor work before shutdown * fix(runtime): terminate cancelled shell process trees * fix(models): let fallback bypass selector failures * fix(cli): reset interrupt handling between turns * docs: rm implementation spec * fix(serve): cancel active turns during shutdown * fix(runtime): protect settlement from waiter cancellation * fix(backends): reject empty shell commands * fix(runtime): terminate descendants after shell exit * fix(mcp): keep standalone discovery off channel loop * fix(cli): own and settle interactive prompt cancellation * fix(serve): keep channel sends off runtime loop * fix(stream): scope cancel context to iterator steps * refactor(serve): require the owned async runtime * fix(channels): keep interactive sends off runtime loop * fix(selector): surface fallback without log spam * test(runtime): normalize Windows shell marker * fix(cli): serialize interactive session turns * fix(shell): bound output drain after termination * fix(ui): do not retry owned runtime failures * fix(shell): allow signal-safe registry reentry * fix(shell): avoid terminating reused process ids * fix(channels): preserve streaming send order * fix(cli): report runtime shutdown timeouts cleanly * fix(mcp): guide async callers to async loader * docs(runtime): clarify reserved async bridge APIs * fix(runtime): bound code interpreter cleanup * test(shell): use active Python for drain regression --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
bd307f3a11 |
refactor: LangGraph gateway layer for UI-agnostic graph and thread access (#295)
* feat(gateway): graph gateway protocol * refactor(cli): wire gateway in cli/tui * refactor(gateway): centralize runtime gateway init * chore(gateway): restrict RunRequest message type * feat(gateway): add langgraph server gateway * chore(cli): tighten serve runtime state typing * refactor(cli): route async task state reads through graph gateway * refactor(gateway): support graph targets in server gateway * refactor(cli): route session commands through graph gateway * refactor(cli): fold thread store under graph gateway * refactor(gateway): route graph state access through gateway * refactor(channels): wire graph gateway * refactor(memory): preserve graph threads for cloning * feat(gateway): add thread cloning * fix(tui): pass effective workspace for thread creation * chore(memory): add workspare dir to memory worker metadata * fix(sessions): filter preloaded UUID registy entries by the current scope * test(fakes): use https * refactor(consumer): consolidate imports * fix(stream): optional summarization event * fix(gateway): resolve abbreviated thread IDs by search * fix(gateway): page server thread listings * fix(gateway): emit pending interrupt events * style: fmt * feat(gateway): persist workspace_dir & model in thread metadata * fix(gateway): page server thread prefix resolution * fix(gateway): expose server thread list metadata * refactor: add back type def * refactor: tighten types * revert: add back worker thread deletion The worker thread forking changes are out of scope for now, so to maintain parity with the existing behavior we'll leave this intact. * fix(gateway): apply compaction to server thread history * refactor(stream): restore direct summary replay suppression * fix(gateway): preserve compaction state and server stream output * fix(gateway): close local stream generator on cancellation --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
fbd1d709ca |
feat: add WebUI mode support with related configuration and onboarding (#252)
* feat: add WebUI mode support with related configuration and onboarding steps * feat: enhance WebUI port configuration to prevent conflicts with backend port * feat: add support for fresh interactive session detection in WebUI |
||
|
|
da74c325d6 |
fix(channel): scope stop and restore resume history (#186)
* fix(channel): scope stop and restore resume history * refactor(channel): simplify stop and resume patch * Delete PR_MESSAGE.md * fix(channel): address review feedback * fix(channel): address remaining review bugs * fix(channel): clean up stopped request handling * fix(channel): preserve resolved replies and sync tui commands --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
65828e9666 |
perf(cli): cut startup latency and defer MCP loading to the background (#171)
* perf(cli): cut startup time of `evosci --help` from ~2.2s to ~0.3s
Module-level imports were eagerly pulling in langchain.chat_models (with
the whole anthropic/openai/google stack), langgraph, textual, and
prompt_toolkit on every invocation — even for `--help` or `config list`.
Defer those with PEP 562 `__getattr__`, using `lazy_loader.attach` (SPEC-1,
the scientific-python standard) where it's a clean attach pattern:
- `EvoScientist/llm/__init__.py`: attach `models` lazily so importing
`context_window` from this package no longer drags in langchain.
- `EvoScientist/stream/__init__.py`: attach display/events lazily; split
the shared Rich `Console` singleton into a new lightweight
`stream/console.py` so callers that only need `console` skip the
`stream.events` → `langchain_core.messages` chain.
- `EvoScientist/cli/__init__.py`: hand-rolled `__getattr__` (reaches into
`..stream.state`, which `lazy_loader` doesn't cover) so `commands` and
`app` are the only eager loads.
- `EvoScientist/cli/commands.py`: move `cmd_interactive`/`cmd_run` to
in-function imports so prompt_toolkit + textual only load when the
interactive path actually runs.
- `EvoScientist/cli/_constants.py`: read `AGENT_NAME` on demand so
`build_metadata` doesn't eagerly import `sessions` (langgraph/aiosqlite).
Adds `lazy-loader>=0.5` as a dependency.
* feat(cli): defer MCP tool loading with live per-server progress
The CLI was blocking ~5 s on MCP tool enumeration before the first
prompt appeared. Move the agent construction off the event loop and
surface per-server progress so the user can interact immediately and see
what's happening.
MCP client:
- Add an `on_progress` callback to `load_mcp_tools` / `aload_mcp_tools`
/ `_load_tools` emitting `start` / `success` / `error` events per
server.
- Fan connection attempts out with `asyncio.gather` so latency no longer
scales linearly with server count; cap simultaneous attempts at
`_MAX_CONCURRENT_CONNECTIONS` (8) via a semaphore so a big stdio fleet
doesn't spawn every subprocess at once.
Agent wiring:
- Plumb `on_mcp_progress` through `create_cli_agent` / `_load_agent` /
`load_mcp_and_build_kwargs` so CLI and TUI can plug in collectors.
CLI (`cmd_interactive`):
- Run `_load_agent` in a background thread via `asyncio.to_thread`; the
prompt and banner render immediately.
- `_await_agent_ready()` awaits the task before each agent-using site
(first turn, channel messages, `/channel`, `/compact`). Raises if
called without a prior `_start_agent_load` instead of silently
reloading without the SQLite checkpointer.
- Pre-prime the progress dict from `load_mcp_config()` so the
bottom-toolbar's `N/M` denominator is stable from the first render.
- Wrap `session.prompt_async` in `patch_stdout(raw=True)` so
`console.print` from the worker-thread progress callback lands cleanly
above the prompt as inline chat messages instead of stomping the
prompt cursor.
TUI (`EvoTextualInteractiveApp`):
- Same background load + `_await_agent_ready()` gates on every
`self._agent` read.
- New `MCPLoaderWidget` mounted at the top of `#input-shell` shows a
header with `N/M` and one live row per server (spinner → ✓ / ✗ with
tool count or error detail). On completion:
- all-clean loads auto-dismiss ~2.5 s later;
- cache hits (no events ever fired) dismiss immediately rather than
flashing a misleading "0/N loaded";
- failures keep the widget mounted so the user can read the errors.
- `dismissed` property lets the app clear its ref so late events from
slow servers become no-ops. The error branch of `_on_agent_loaded`
also settles the widget so a load failure can't leave the spinner
animating forever.
- Chat input is `disabled` while MCP resolves — no placeholder hack, no
"waiting…" system message.
Shared:
- Hoist braille spinner frames to `status_bar.SPINNER_FRAMES` and import
them in the TUI widget so CLI and TUI animate in sync.
Tests:
- Extend `test_agent_mcp_cache` fakes to accept the new `on_progress`
kwarg.
- New `TestLoadToolsProgressCallback` in `test_mcp_client` exercises the
event sequence for success/failure/mixed fleets, verifies a buggy
callback doesn't break the load, and asserts the semaphore caps
in-flight connections.
* style: ruff
* chore: update uv.lock
* chore: uv.lock
* fix: coderabbit issues
* style: fmt
* fix: move _await_agent_ready inside try block
* fix(tui): auto-dismiss MCP loader widget on failure
The widget was designed to stay mounted on failure so the user could
read error detail, but since it's pinned above the input it never went
away in practice — just permanent banner clutter.
Auto-dismiss on failure too, with a longer grace (12s vs 2.5s) so the
error summary stays readable.
* fix: address second coderabbit pass
- Channel handlers (CLI + TUI): catch agent-load failures so the
channel request doesn't hang; CLI moves `_await_agent_ready()`
inside the existing try/except, TUI catches explicitly and calls
`_set_channel_response` with the error.
- Stale background loads: `prev.cancel()` only stops the asyncio
wrapper, not the thread running `_load_agent`. Added a generation
token (`agent_load_id` / `self._agent_load_id`) and gated both
progress and completion callbacks on it so a superseded load can't
clobber the current session's state or UI.
- TUI prompt lifecycle: added `_agent_load_pending()` and gated the
`_process_channel_message` / `_handle_command` finally blocks on it
so `/new` or `/resume` invoked from a command keeps the prompt
disabled until the fresh load settles.
- TUI readiness failures: `_run_turn` and `_handle_command` now
catch exceptions from `_await_agent_ready()` and surface a
"Agent failed to load: …" system message instead of letting the
exception escape into Textual's traceback panel.
* refactor(cli): share background agent loader between CLI and TUI
The CLI and TUI were carrying near-identical copies of the same
background-load state machine: the `agent_task`, the `agent_load_id`
generation token, the gated progress/completion callbacks, and the
per-server progress dict. Every CodeRabbit finding on that lifecycle
had to be fixed in both files.
Extract it into `cli/_agent_loader.py`:
- `MCPProgressTracker` — owns the `server -> (state, detail)` dict;
exposes `prime`, `record`, `snapshot`, `totals`.
- `BackgroundAgentLoader` — owns `agent`, the in-flight task, and the
generation token. Exposes `start(**loader_kwargs)`, `await_ready()`,
`is_pending`. Internally gates all progress/completion callbacks by
generation so a superseded load can't clobber the current session.
UI-specific rendering plugs in via `on_progress` / `on_success` /
`on_failure` callbacks.
Both surfaces now just wire their UI hooks; the loader file holds no
Rich / prompt_toolkit / Textual dependencies. Net -345 lines from
`interactive.py` + `tui_interactive.py`; +20 unit tests pinning the
lifecycle (generation filtering, cache-hit short-circuit, failure
reset, progress ordering).
* refactor(cli): make _on_done the sole authority for agent state transitions
await_ready no longer sets self.agent — it just awaits the task and
reads what _on_done already wrote. Eliminates the dual-write overlap
(asyncio guarantees done-callbacks fire in registration order).
* fix(tui): let users type during MCP load, only block on send
Remove prompt-disabling during background agent load — the TUI now
matches the CLI approach where the input stays enabled and only gates
on await_ready() at submit time. The MCPLoaderWidget still provides
visual feedback that loading is in progress.
* fix(loader): preserve real load error on await_ready; dedup failure message
CodeRabbit flagged two issues with the new loader:
1. After a failed load, `_on_done` nulled `self._task`, so the next
`await_ready()` hit the "before start()" branch and the CLI wrapper
remapped it to a misleading "checkpointer not available" message —
losing the real exception (bad MCP config, network, etc.).
Keep `_task` set on failure so `await_ready` re-raises the real
exception. Added `needs_restart` so TUI's auto-retry check stays a
one-liner and doesn't need to reach into task internals.
2. TUI reported each load failure twice: once from
`_on_agent_load_failure` (the done-callback) and once from each
caller of `_await_agent_ready` (`_run_turn`,
`_process_channel_message`, `_handle_command`) catching the re-raise.
`_on_agent_load_failure` is now the sole local reporter; callers
just handle control flow (return cleanly, set channel response to
unblock remote).
* fix(cli): wire /model handler through the agent loader
The /model command from main (merged via f1f0d7c) still reached for
`state["agent"]` (CLI) and `self._agent` (TUI) — both removed by the
background-loader refactor. CLI raised KeyError on first invocation;
TUI raised AttributeError. Writes to the old fields also had no effect
because every other code path now reads from `agent_loader.agent`, so
the model switch would have silently failed.
Route everything through the loader: `await _await_agent_ready()` up
front so /model doesn't race with the initial background load, build
the `CommandContext` with the current agent, and sync `ctx.agent` back
into `agent_loader.agent` (plus channel globals) when the command
replaces it.
* fix(cli): isolate progress callback, capture awaited agent, gate by requires_agent
Three CodeRabbit findings on the loader + command dispatch path:
- Wrap ``_on_progress`` in try/except inside the loader's gated wrapper
so a buggy UI adapter can't bubble into ``loader_fn`` and fail the
whole background load. The MCP client already protects this, but
defence-in-depth keeps the loader self-contained.
- In CLI channel + main-loop streaming, capture the agent returned by
``_await_agent_ready()`` and pass that into ``run_streaming`` rather
than reading ``agent_loader.agent`` after a subsequent ``await``.
A concurrent ``/new``/``/resume``/``/model`` could have swapped it.
- Add ``requires_agent: ClassVar[bool] = False`` to ``Command`` and
mark ``/compact``, ``/model``, ``/channel`` as ``True``. TUI dispatch
sites (channel and keyboard) now check ``cmd_manager.resolve(...)``
and only wait for readiness when the command actually needs the
agent. ``/mcp add``, ``/skills``, ``/new`` etc. no longer deadlock
behind a failing MCP load they are meant to fix.
* fix(cli): guard sync-back, subcommand-aware gating, /model adopt-path
Three CodeRabbit findings on command dispatch:
- ``_handle_command`` unconditionally synced ``ctx.agent`` back into
``agent_loader``. For non-agent commands ``ctx.agent`` is ``None``,
so ``/threads`` / ``/mcp`` / ``/skills`` (etc.) could clobber a valid
loaded agent — and rebind channel globals to ``None``. Guard the
sync on ``ctx.agent is not None``.
- ``/channel status`` and ``/channel stop`` don't touch ``ctx.agent``
but the class-level ``requires_agent = True`` blocked them behind
agent readiness. Added ``Command.needs_agent(args)`` (defaults to
``requires_agent``) so ``/channel`` can override with subcommand
awareness; kept the class flag for the common case.
- ``/model`` builds a new agent from scratch, it never reads the
existing one — gating it on readiness meant a broken provider
blocked the command that would fix it. Flipped it to
``requires_agent = False`` and added ``BackgroundAgentLoader.adopt``
so the UI can seat the replacement and supersede any in-flight
load (the generation token keeps a late completion from clobbering
the adopted agent).
Bonus cleanup: ``CommandManager.resolve`` now returns
``(command, args)`` so callers can invoke ``needs_agent`` without
re-implementing ``shlex`` parsing.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
|
||
|
|
65db3a4fcd |
Add status bar and compact summary widgets with context window resolu… (#152)
* Add status bar and compact summary widgets with context window resolution - Implemented a shared status bar for CLI and TUI frontends, including helpers for managing session metrics and context windows. - Created a `CompactSummaryWidget` for displaying manual summaries in a collapsible format. - Introduced a `CompactingWidget` to indicate ongoing compacting processes. - Added a base class `TimedStatusWidget` for widgets that require a timer. - Developed context window resolution helpers to retrieve context window sizes from various model attributes. - Enhanced tests for context window resolution and status bar functionalities, ensuring accurate behavior across different scenarios. - Updated existing tests to cover new features and maintain code quality. * refactor(Channel): simplify lambda function in _send_with_retry method * feat: enhance context editing logic and improve error handling in StreamState * refactor(Channel): streamline lambda function in _send_with_retry method * feat: rename auto-approve option to auto-mode for unattended execution; update checkpoint queries to filter by agent name; improve compatibility validation logic * feat: rename auto-approve option to auto-mode; update related logic and tests for improved unattended execution * fix: correct formatting of console message for MCP server configuration status * feat: add check for None summary_message in _apply_summarization_event to prevent errors * feat: enhance _load_checkpoint_messages to validate message format and apply summarization event |
||
|
|
4a3d6c0318 | chore: add ruff lint rules and turn on formatting | ||
|
|
c5a4d559a2 |
Refactor test cases for improved readability and consistency
- Added blank lines for better separation of test cases in multiple test files. - Reformatted event handling in tests for clarity and consistency. - Ensured consistent use of multi-line formatting for dictionary arguments in event handling. - Improved assertions and test descriptions for better understanding. - Updated test cases across various modules including test_stream_state, test_stream_utils, test_summarization, test_thread_selector, test_tool_error_handler, test_tui_widgets, test_ui_runtime, and test_wechat_channel. |
||
|
|
c15c0dfdf0 |
feat: add ask_user middleware and interactive widget for user prompts
- Implemented `ask_user` middleware to facilitate agent-initiated questions during research workflows. - Created `AskUserWidget` for interactive user prompts, supporting both text and multiple choice questions. - Enhanced `StreamEventEmitter` to handle `ask_user` interrupts and updated event handling in `stream_agent_events`. - Added state management for pending `ask_user` events in `StreamState`. - Developed validation for question structures and parsing for responses. - Introduced unit tests covering middleware functionality, event handling, and widget behavior. |
||
|
|
c4de31c772 | fix(ui): map legacy UI backend values to current equivalents in normalization | ||
|
|
07e57e110b | feat(ui): update UI backend options to use 'cli' and 'tui' | ||
|
|
833fa0abfb |
Implement HITL (Human-in-the-Loop) approval mechanism
- Added HITL approval lifecycle to ToolCallWidget, including a new "rejected" status. - Introduced configuration options for auto-approval and shell command allow list in EvoScientistConfig. - Developed approval handling functions in display module to manage HITL interrupts and user decisions. - Created ApprovalWidget for user interaction during approval prompts. - Enhanced StreamEventEmitter to support interrupt events. - Updated StreamState to track pending interrupts. - Implemented tests for HITL functionality, including event structure, state handling, and approval logic. |
||
|
|
03f4134190 |
feat: add UserMessage widget and UI backend selection during onboarding
- Introduced UserMessage widget for displaying user input with a styled prompt. - Updated onboarding steps to include UI backend selection (Rich CLI or Textual TUI). - Modified EvoScientistConfig to store selected UI backend. - Enhanced configuration handling to support UI backend environment variable. - Updated README with new UI backend options and commands. - Added tests for new UI backend functionality and UserMessage widget. - Removed obsolete test files and ensured existing tests are updated accordingly. |