37 Commits

Author SHA1 Message Date
jfilipiuk be8c23861d feat: dispatch newly installed experts in the background without /new (#420)
* refactor: rename AGENTS.md to EXPERT.md per EvoSkills convention

* feat: resolve newly installed experts on start_async_task miss

* chore: reword the /expert invite hint to state the dispatch boundary

* docs: note the resolve-on-miss caller in build_expert_async_subagent_specs

* fix: thread the construction cfg through resolve-on-miss and symmetric setdefault

* fix: guard agent_map iteration against concurrent resolve-on-miss writes

* fix: warn once per broken expert on repeated resolve-on-miss walks

* fix: scope the /expert invite hint to newly installed experts

* docs: note live uninstalls as a resolve-on-miss limitation

* chore: isolate the warn-once collision test key from the route-specs suite
2026-09-10 22:47:35 +00:00
Xi Zhang a45563ea7f feat(scheduler): optional rubric acceptance checklist for scheduled tasks (#451)
Scheduled tasks gain an optional `rubric`: an acceptance checklist graded
after each run by deepagents' RubricMiddleware (LLM-as-a-judge on the
auxiliary model) with one revision retry. No `rubric` key = no-op.

- cron/schedule.py: create_schedule/run_now carry the rubric in the run
  input and cron metadata only when non-blank
- middleware/scheduler.py: schedule_task gains `rubric`; list marks graded rows
- commands/implementation/schedule.py: `/schedule add ... --rubric`,
  `/schedule run` forwards the stored rubric, list gets a Rubric column
- subagents/_factory.py: RubricMiddleware mounted last on the scheduler
  graph so a needs_revision jump skips the memory lifecycle until the
  accepted run; grader is read-only (ls + read_file, eviction off),
  bounded by a 12-call budget, and gets an explicit structured-output
  strategy on OpenRouter (Gemini JSON mode, Anthropic tool calling);
  warns at build for anthropic/claude-fable-5.1 via OpenRouter, which
  grades under neither strategy today
- tests: 17 new cases; fix two pre-existing fixture leaks (callable
  backend stub in test_hitl, import-under-patch in test_async_subagent_factory)
2026-09-05 07:04:26 +01:00
X-iZhang b40b6f784d feat: clarify comment in NewCommand to align with explicit expert clearing 2026-08-07 17:07:01 +01:00
X-iZhang aff63ccd65 feat: enhance expert invitation handling with case-insensitive matching and session management 2026-08-07 17:07:01 +01:00
jfilipiuk ab1a6b0062 feat: adopt the AGENTS.md expert-skill contract (#404)
* feat: detect expert skills by AGENTS.md presence

* feat: let the orchestrator choose the expert dispatch tool

* refactor: drop per-skill expert dispatch classification

* chore: remove duplicate TestSkillManager test classes left by rebase

* fix: clear legacy actor fields when AGENTS.md declares the expert

* refactor: move the expert prompt into ActiveTeamMiddleware
2026-08-07 17:07:01 +01:00
jfilipiuk bd2464423a feat: agent-teams part D - async expert dispatch mechanism (#391)
* fix(deps): pin openrouter below 0.11 to avoid SSE stream regressions (#373)

* fix(openrouter): address SSE stream leak by closing response iterator

* refactor(openrouter): pass through SDK args in SSE leak patch

* test(openrouter): make SSE leak tests version-agnostic across SDK generations

* fix(openrouter): remove SSE stream leak patch and update dependencies

* fix: repair interrupted tool call history (#366)

* fix: repair interrupted tool call history

Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths.

* fix: repair malformed tool calls and dedupe repair warnings

Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted
threads with syntactically invalid tool calls get synthesized error
results and are accepted by strict providers.

Preserve the originating tool call's name in the synthesized ToolMessage,
and deduplicate repair warnings per unique tool-call id via a warned set
owned by the middleware instance, since the middleware rewrites the
request but not thread state.

Document the middleware's scope versus deepagents' PatchToolCallsMiddleware
(orphan ToolMessage dropping and mid-run coverage).

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Update README.md

* Update README.md

* Update README.zh-CN.md

* fix: scrub host path from skill_manager output and guard batch install (#377)

* fix: scrub host path from skill_manager output and guard batch install

* test: tighten install leak guards to catch host path in either tier

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: log missing async-subagent tools at DEBUG, not WARNING (#378)

* fix: log missing async-subagent tools at DEBUG, not WARNING

* fix: distinguish load_subagents callers via async_swap_pending flag

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: default reasoning context for codex proxy Responses API (#380)

* fix: set langgraph and codex proxy runtime defaults

* fix: address runtime default review feedback

* fix: drop langgraph dev env defaults per maintainer review

langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment,
so the reported KeyError cannot come from this flow; the env defaults added
here were unnecessary. Scope the PR back to the codex proxy reasoning
context fix only.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* release: v0.2.4 (#389)

* fix: add support for new Anthropic models and enhance adaptive thinking tests

* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models

* fix: update version to v0.2.4 in badges, README, and project files

* fix: update Star History chart links in README and README.zh-CN

* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog

* fix: update wechat group image in assets

* refactor(runtime): centralize async bridges under an owned runtime (#376)

* feat(runtime): add application-scoped async runtime

* refactor(cli): use owned runtime for session stats

* refactor(onboard): use the owned async runtime

* docs(runtime): record async bridge ownership

* refactor(middleware): keep sync fallback synchronous

* refactor(mcp): load tools on an owned runtime

* refactor(cli): share owned runtime across entry points

* refactor(channels): make inbound sync bridge explicit

* refactor(stream): run Rich streaming on owned runtime

* chore(runtime): remove nest-asyncio dependency

* refactor(asyncio): require active loops in async code

* docs(runtime): document final event loop ownership

* fix(stream): cancel stalled owned streams

* fix(cli): recover cleanly from stream cancellation

* fix(runtime): drain executor work before shutdown

* fix(runtime): terminate cancelled shell process trees

* fix(models): let fallback bypass selector failures

* fix(cli): reset interrupt handling between turns

* docs: rm implementation spec

* fix(serve): cancel active turns during shutdown

* fix(runtime): protect settlement from waiter cancellation

* fix(backends): reject empty shell commands

* fix(runtime): terminate descendants after shell exit

* fix(mcp): keep standalone discovery off channel loop

* fix(cli): own and settle interactive prompt cancellation

* fix(serve): keep channel sends off runtime loop

* fix(stream): scope cancel context to iterator steps

* refactor(serve): require the owned async runtime

* fix(channels): keep interactive sends off runtime loop

* fix(selector): surface fallback without log spam

* test(runtime): normalize Windows shell marker

* fix(cli): serialize interactive session turns

* fix(shell): bound output drain after termination

* fix(ui): do not retry owned runtime failures

* fix(shell): allow signal-safe registry reentry

* fix(shell): avoid terminating reused process ids

* fix(channels): preserve streaming send order

* fix(cli): report runtime shutdown timeouts cleanly

* fix(mcp): guide async callers to async loader

* docs(runtime): clarify reserved async bridge APIs

* fix(runtime): bound code interpreter cleanup

* test(shell): use active Python for drain regression

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* feat: add payload-aware EvoAsyncSubAgentMiddleware

* feat: register expert_container_async graph for async expert dispatch

* feat: fold installed expert skills into async subagent registry

* fix: accept 'async' as valid default_dispatch value

* feat: dispatch-aware ActiveTeamMiddleware cue (task vs start_async_task)

* fix: drop future annotations in expert_async_subagent so ToolRuntime injects

* feat: surface output_path and skill_name to expert container as runtime cue

* fix: extend AsyncWatcher client cache with expert specs for completion nudge

* feat: teach main agent the async-expert return envelope shape

* feat: propagate cfg.model to expert-async runs.create via ClientCacheProxy

* chore: guard AsyncWatcher client-cache extension against upstream rename

* test: cover output_path runtime-context tail block and wrong-type guard

* docs: drop out-of-repo notes/ ref from expert_container_async module doc

* fix: warn on unrecognized default_dispatch frontmatter value

* fix: reject empty-body expert skills on async dispatch to match sync policy

* docs: explain why expert container includes general-purpose subagent

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385)

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL

* fix: keep parent env authoritative over workspace .env for mapped keys

* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins

* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS

* fix: merge .env via dotenv_values to close empty-value and RMW-race edges

* chore: align docstrings after .env-merge rework

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Add Requesty as an LLM provider (#346)

* Add Requesty as an LLM provider

* Address review: Requesty prompt caching, model ordering, key validation

- Declare Anthropic-style prompt caching for Requesty Claude models by
  default (mirroring the OpenRouter behavior), with an opt-out flag
  EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
  provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
  overrides native/OpenRouter models for names it shares with them
  (the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
  invalid/missing key (public catalog), so it cannot validate a key. Use a
  minimal authenticated /v1/chat/completions request instead (200 = valid,
  403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).

* Validate Requesty key against auth layer, not a specific model

The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix(llm): filter unnamed tool calls (#390)

* fix(llm): filter unnamed tool calls

* test(llm): cover tool call sanitization branches

* fix(llm): repair unnamed tool calls in middleware

---------

Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix(middleware): mount tool-history repair on sync subagents and harden raw tool-call vetting (#393)

* feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation

* fix(tests): add test for dropping non-list raw tool calls in repair_tool_history

* Add Atlas Cloud LLM provider (#388)

* Add Atlas Cloud LLM provider

* Add Atlas Cloud onboarding support

* fix(validators): update atlascloud key validation to handle insufficient balance case

---------

Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix: prepend EvoAsyncSubAgentMiddleware for prefix cache stability

* docs: clarify list_dispatchable_experts covers both dispatch shapes

* fix: guard async expert fold-in against reserved-name collisions

* fix: honest advertising surfaces for async expert dispatch

* fix: compose expert persona into base-stack system_message

* fix: drop payload from start_async_task, inject skill_name by construction

* feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395

- Introduced TodoListMiddleware to the middleware stack for better task management.
- Updated HITL interrupt configuration to include 'delete' operations requiring approval.
- Implemented error handling for delete operations in read-only and memory backends.
- Enhanced approval prompt formatting to display file paths for delete actions.
- Added tests to ensure delete operations are correctly blocked or prompted for approval.
- Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality.

* revert: drop skill_manager from sync expert-container tool_registry

* revert: drop skill_manager from async expert-container tools

* fix: drop removed ASYNC_TASK_SYSTEM_PROMPT import for deepagents 0.7.0

* fix: mock list_dispatchable_experts in single-cue test for CI

* chore: drop stale output_path from async container graph docstring

* fix: forward configurable_extra through owned-runtime and HITL re-invocations

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: Sanjay Santhanam <51058514+Sanjays2402@users.noreply.github.com>
Co-authored-by: Yougang Lyu <82445958+youganglyu@users.noreply.github.com>
Co-authored-by: houren Antony <2212222@mail.nankai.edu.cn>
Co-authored-by: dinos <dinospk1999@gmail.com>
Co-authored-by: Thibault Jaigu <84420566+Thibaultjaigu@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
Co-authored-by: nightcityblade <jackchen@haloailabs.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nb213 <binyangzhu000@gmail.com>
Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
2026-08-07 17:07:01 +01:00
jfilipiuk 3cda9894c7 feat: agent-teams part C - expert selection UX (depends on part B) (#371)
* feat: bias main-agent delegation toward configurable.active_teams

* feat: add /experts and /expert TUI commands for expert-skill summoning

* feat: align expert-selection wording with WebUI (invite/dismiss)

* chore: clear active_teams on /new, cleanup active_team.py comment

* fix: use local append_to_system_message in ActiveTeamMiddleware

* fix: prevent configurable_extra from overriding thread_id

* fix: cache expert-skill lookup for /expert completions

* fix: suppress /expert completions past first arg and on exact match

* fix: invalidate /expert completion cache on skill install/uninstall

* fix: refuse /expert invites for non-dispatchable expert skills

* fix: fire /expert cache invalidation on every install_skill / uninstall_skill path

* fix: propagate active_teams to Rich CLI and serve dispatch surfaces

* fix: keep invited experts across channel shutdown

* fix: match /expert completions case-insensitively
2026-08-07 17:07:01 +01:00
dinos 8b1451cdda refactor(runtime): centralize async bridges under an owned runtime (#376)
* feat(runtime): add application-scoped async runtime

* refactor(cli): use owned runtime for session stats

* refactor(onboard): use the owned async runtime

* docs(runtime): record async bridge ownership

* refactor(middleware): keep sync fallback synchronous

* refactor(mcp): load tools on an owned runtime

* refactor(cli): share owned runtime across entry points

* refactor(channels): make inbound sync bridge explicit

* refactor(stream): run Rich streaming on owned runtime

* chore(runtime): remove nest-asyncio dependency

* refactor(asyncio): require active loops in async code

* docs(runtime): document final event loop ownership

* fix(stream): cancel stalled owned streams

* fix(cli): recover cleanly from stream cancellation

* fix(runtime): drain executor work before shutdown

* fix(runtime): terminate cancelled shell process trees

* fix(models): let fallback bypass selector failures

* fix(cli): reset interrupt handling between turns

* docs: rm implementation spec

* fix(serve): cancel active turns during shutdown

* fix(runtime): protect settlement from waiter cancellation

* fix(backends): reject empty shell commands

* fix(runtime): terminate descendants after shell exit

* fix(mcp): keep standalone discovery off channel loop

* fix(cli): own and settle interactive prompt cancellation

* fix(serve): keep channel sends off runtime loop

* fix(stream): scope cancel context to iterator steps

* refactor(serve): require the owned async runtime

* fix(channels): keep interactive sends off runtime loop

* fix(selector): surface fallback without log spam

* test(runtime): normalize Windows shell marker

* fix(cli): serialize interactive session turns

* fix(shell): bound output drain after termination

* fix(ui): do not retry owned runtime failures

* fix(shell): allow signal-safe registry reentry

* fix(shell): avoid terminating reused process ids

* fix(channels): preserve streaming send order

* fix(cli): report runtime shutdown timeouts cleanly

* fix(mcp): guide async callers to async loader

* docs(runtime): clarify reserved async bridge APIs

* fix(runtime): bound code interpreter cleanup

* test(shell): use active Python for drain regression

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-27 14:17:57 +01:00
dinos 06a9511bdd fix(channels): telegram slash commands (#364) 2026-07-17 16:05:49 +01:00
dinos 01845f4311 refactor: route middleware display events through an injected event sink (#343)
* chore: add pytest-asyncio in auto mode

* test: migrate channel and stream tests to native async

Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.

* test: migrate command and model/middleware tests to native async

Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.

* test: migrate TUI, notifier, gateway, and session tests to native async

TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.

* test: replace direct asyncio.run() calls with native async tests

Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.

* test: drop undeclared anyio markers and delete run_async helper

The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).

* test: add autouse fixture for watcher cleanup

* refactor: remove redundant hasattr calls

* refactor: add typed middleware event sink and thread through assembly

Add MiddlewareEventSink protocol + NoOpSink in middleware/events.py
with a documented any-thread non-blocking contract (contract test uses a
deliberately-slow fake sink). Thread an optional `events` parameter
through create_cli_agent -> _get_default_middleware -> tool selector /
model fallback constructors; subagent stacks are always forced to
NoOpSink.

* refactor: inject a notifier port into async-watcher and background middleware

Add public pre_cancel_watcher() and enqueue_task_notification() to
cli/async_notifier.py and a small NotifierPort protocol
(middleware/notifier.py) that the module satisfies structurally.
AsyncWatcherMiddleware and BackgroundExecutionMiddleware now receive the
port by constructor injection at the composition root, deleting the lazy
'from ..cli import async_notifier' imports and the private
_watcher_by_thread / _enqueue pokes.

* refactor: invert tool-selection ownership onto a frontend event sink

The adaptive tool selector now reports on_tool_selection_started /
on_tool_selection / on_tool_selection_ended to the injected sink instead
of writing four process-global module variables. The frontend sink
(stream/sink.py FrontendEventSink) owns the selected/total/active state
with consume-once + dedup-vs-last-emitted semantics;
stream/tool_selection.py reads that sink object (a ToolSelectionView)
rather than reaching into tool_selector's globals.

Deleted: the 4 module globals, the cross-module mutations in
tool_selection.py, the track_stream_selection flag, the now-vestigial
_ToolSelectionTrackerMiddleware, reset_tool_selection_state_for_tests,
and the autouse conftest fixture. The sink is threaded from the two
interactive frontends through create_runtime_gateways ->
LocalGraphGateway (read side) and _load_agent -> create_cli_agent (write
side); subagent / headless stacks get NoOpSink.

* refactor: route model-fallback narration through the injected event sink

Delete the _ui_emit_fn / set_ui_emit module global and the
..stream.console import from model_fallback.py. The fallback middleware
now reports through its injected sink: the fallback transition via the
structured on_model_fallback (the frontend formats the '-> Falling back
to ...' line), and the surrounding narration (primary-failure header,
per-attempt outcome, exhaustion, non-fallbackable rejection) via
emit_fallback_notice, preserving the exact user-facing text. The TUI
binds its _append_system as the sink's fallback display where it used to
call set_ui_emit (cleared on exit); the Rich CLI's sink prints to the
console. _try_fallbacks / _guard_and_fallback take the sink.

* refactor: declare events on the GraphGateway protocol

Both gateway implementations now carry an explicit events attribute
(LangGraphServerGateway holds None — no frontend renders middleware
events across the HTTP boundary), so the four call sites use plain
attribute access instead of getattr probing an implicit contract.

* refactor: bind fallback display via the closure-scoped concrete sink

The App methods used gateway.events (typed as the read-side view) and
hasattr-probed for the concrete FrontendEventSink API. The enclosing
factory creates that sink two hundred lines up — close over it directly:
no probing, fully typed, and it becomes a constructor parameter
naturally when the App class is hoisted out of the factory.

* fix: end tool selection before fallback handler

* fix: keep fallback display errors non-fatal

* fix: preserve selector suppression for default streams

* fix: restore fallback notice console display

* refactor: consolidate fallback narration events

* refactor: clean middleware event sink plumbing

* fix: type gateway session events

* refactor: make all event protocols runtime-checkable

MiddlewareEventSink already carried @runtime_checkable (the stream
binding guard isinstance-checks it); ToolSelectionView and SessionEvents
now match, so mirroring that pattern against any of the three protocols
works instead of raising TypeError.

* fix(cli): close QuickJS workers after one-shot failures

* fix(cli): honor no-thinking in final output

* fix(channels): report failed startup accurately

* fix(channels): make Telegram cleanup idempotent

* fix(tui): skip command sync during exit

* fix(channels): preserve startup state during retries

* refactor(channels): share pending startup status

* refactor(cli): expose channel startup snapshot

* fix(tui): move channel startup off event loop

* test(channels): release retry gate on assertion failure

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-14 22:34:17 +00:00
Wiktor Cupiał 1d117ff277 feat: completion enchancements (#302)
* feat: completion enchancements

* fix: handle exception

* fix: duplicate view

* fix: remove deadcode

* fix tab
2026-07-05 05:14:42 +00:00
dinos 2b244888ec feat(memory): autoskills (#319) 2026-07-03 11:16:33 +02:00
Xi Zhang 7ccfe68f3f feat: add scheduler functionality with cron-style task management (#306)
* feat: add scheduler functionality with cron-style task management

- Implemented a new scheduler subagent to automate recurring tasks using cron expressions.
- Enhanced the subagent factory to include the skill manager and auxiliary chat model for the scheduler.
- Created a YAML configuration for the scheduler with a detailed system prompt and toolset.
- Updated README files to include documentation on scheduled tasks and usage examples.
- Added tests for the scheduler, including command execution, scheduling tools, and middleware integration.
- Introduced new dependencies for timezone handling and ensured compatibility in the project configuration.

* fix(async-notifier): ensure fallback hint is used for unknown notification kinds

* feat: enhance scheduling functionality and improve system message handling
2026-06-25 17:31:23 +01:00
dinos bd307f3a11 refactor: LangGraph gateway layer for UI-agnostic graph and thread access (#295)
* feat(gateway): graph gateway protocol

* refactor(cli): wire gateway in cli/tui

* refactor(gateway): centralize runtime gateway init

* chore(gateway): restrict RunRequest message type

* feat(gateway): add langgraph server gateway

* chore(cli): tighten serve runtime state typing

* refactor(cli): route async task state reads through graph gateway

* refactor(gateway): support graph targets in server gateway

* refactor(cli): route session commands through graph gateway

* refactor(cli): fold thread store under graph gateway

* refactor(gateway): route graph state access through gateway

* refactor(channels): wire graph gateway

* refactor(memory): preserve graph threads for cloning

* feat(gateway): add thread cloning

* fix(tui): pass effective workspace for thread creation

* chore(memory): add workspare dir to memory worker metadata

* fix(sessions): filter preloaded UUID registy entries by the current scope

* test(fakes): use https

* refactor(consumer): consolidate imports

* fix(stream): optional summarization event

* fix(gateway): resolve abbreviated thread IDs by search

* fix(gateway): page server thread listings

* fix(gateway): emit pending interrupt events

* style: fmt

* feat(gateway): persist workspace_dir & model in thread metadata

* fix(gateway): page server thread prefix resolution

* fix(gateway): expose server thread list metadata

* refactor: add back type def

* refactor: tighten types

* revert: add back worker thread deletion

The worker thread forking changes are out of scope for now, so to
maintain parity with the existing behavior we'll leave this intact.

* fix(gateway): apply compaction to server thread history

* refactor(stream): restore direct summary replay suppression

* fix(gateway): preserve compaction state and server stream output

* fix(gateway): close local stream generator on cancellation

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-22 13:54:14 +00:00
Ziheng Zhang 6593ac9b5b fix(cli): submit slash command on Enter when name prefixes another (#293) (#300)
Typing `/model` and pressing Enter did nothing in the TUI; the picker
only opened via `/model --save` or `/model <name>`. The completion popup
matched both `/model` and `/model-fallback` by prefix, so the
exact-match-hide guard (which required a single match) never fired. With
the popup still visible, the TUI's Enter handler completed the text
instead of submitting the command, so it never executed.

Treat the typed prefix as an exact match whenever it equals any matched
command name, not only when it is the sole match. This hides the popup
on a complete command name so Enter submits it, even when a longer
command shares the prefix.
2026-06-18 14:07:43 +01:00
houren Antony 76972449c7 feat(cli): multi-stage slash command completions with subcommand awareness (Phase 1 of #82) (#273)
* feat(cli): multi-stage slash command completions with subcommand awareness

Phase 1 of #82 — subcommand and argument awareness in completions.

- commands/base.py: add SubCommand dataclass and subcommands/category
  ClassVars to the Command ABC. Each SubCommand has name, description,
  and optional arguments.

- commands/manager.py: add get_subcommands() and list_subcommands()
  methods to expose subcommand metadata for completion rendering.

- commands/implementation/mcp.py: declare 6 subcommands (list, config,
  add, edit, remove, install).

- commands/implementation/model_fallback.py: declare 6 subcommands
  (list, add, remove, clear, save, help).

- commands/implementation/channel.py: declare 2 subcommands
  (status, stop).

- cli/tui_interactive.py: rewrite on_text_area_changed slash-completion
  branch. When the user types a command name + trailing space and the
  command has subcommands, show subcommand completions instead of hiding
  the popup. Filter subcommands by typed prefix in multi-token input.

- commands/implementation/general.py: /help now lists subcommands
  below each command that declares them.

Tests: 8 new tests covering SubCommand creation, CommandManager
subcommand lookup, and cross-command verification.
2293 passed baseline, no regressions.

* fix: subcommand completion preserves prefix + prompt_toolkit + tests

- _apply_selected_completion: preserve '/mcp ' prefix when completing
  subcommands via _comp_is_subcommand flag
- SlashCommandCompleter (Rich CLI): add subcommand completion support
- Fix trailing-space bug: rstrip prefix before top-level matching
- test_tui_widgets.py: update stub on_input_changed to match multi-stage
  logic; add 5 new subcommand tests
- test_command_manager.py: 8 tests for SubCommand + CommandManager

28 passed, 0 failed.

* style: ruff format tui_interactive.py + test_tui_widgets.py

* style: fix RUF012 ClassVar annotation on subcommands lists

* fix: sync test stub, add len>=3 guard, remove exact-match hide

- Sync test stub on_input_changed with real TUI code (remove exact-match
  hide for subcommands, add len(parts)>=3 guard)
- Update test_input_changed_exact_subcommand_hides -> shows_confirmation
- Add test_input_changed_three_parts_hides
- Remove unused category ClassVar (din0s: what is this for)

* refactor(commands): extract shared completion engine

Per din0s feedback: one shared completion engine (commands/_completion_engine.py)
that parses text + cursor once, returns structured CompletionCandidate objects
with replace_start/replace_end ranges.

- SlashCommandCompleter (Rich CLI): thin adapter, delegates to engine
- on_text_area_changed (TUI): thin adapter, delegates to engine
- _apply_selected_completion: uses candidate.replace_start/replace_end instead
  of _comp_is_subcommand flag
- Tests: engine tested directly (10 new tests), stub methods updated

29 passed, 0 failed.

* style: ruff format

* fix: preserve text after cursor when applying completion

CodeRabbit: replace_start only cuts from start to cursor,
dropping any suffix after the cursor. Use replace_start + replace_end
to correctly splice the replacement while preserving trailing text.

* fix(cli): repair slash-command completion (TUI crash, subcommand bugs, sort)

Apology + context: the previous push shipped a TUI-breaking change
(the new shared engine assumed ``event.text_area.cursor_position``
existed, but ``ChatTextArea`` / ``Changed`` don't expose it). User
caught the crash on ``/``; fixing that surfaced two more bugs in
the engine that din0s had already flagged. This commit addresses
all of them and drops a piece of dead stub code.

## Bug fixes

1. **TUI crash on ``/``** (``tui_interactive.py:2335``)
   ``event.cursor_position`` doesn't exist on the ``Changed`` event,
   and ``ChatTextArea`` (Textual ``TextArea`` subclass) doesn't expose
   ``cursor_position`` either. Pass ``len(event.text_area.text)``
   instead — the user types at the end of the input in practice.

2. **Subcommand trailing-space duplication** (``_completion_engine.py``)
   Typing ``/mcp a `` + Tab produced ``/mcp aadd``. The engine
   included the trailing space in ``replace_end``; the TUI apply
   unconditionally appended ``" "``, producing double-space output.
   Fix: ``replace_end`` excludes the trailing space; the TUI apply
   checks ``current[replace_end:].startswith(" ")`` and skips the
   separator when the suffix already has one.

3. **Subcommand exact-match confirmation noise** (``_completion_engine.py``)
   Typing ``/mcp list`` + Tab re-inserted ``list`` and the popup
   kept showing the same subcommand. Add a guard mirroring the
   top-level exact-match rule: when the only subcommand match is
   the prefix itself (no trailing space), return ``empty``.

4. **Alphabetical sort dropped in CLI** (``cli/interactive.py``)
   The new completer iterated ``result.candidates`` in registration
   order. Re-add ``sorted(result.candidates, key=lambda c: c.text)``.
   Same sort added to the TUI for consistency.

## Cleanup

- Drop the dead ``on_input_changed`` method from the ``_StubApp``
  test stub (0 call sites) plus the unused ``_slash_commands`` /
  ``_subcommands`` locals that fed it. This addresses din0s's
  comment about the stub duplicating real TUI logic — the inlined
  copy is no longer needed since the real completer now routes
  through the shared engine.

## Tests

- ``test_engine_exact_subcommand_shows_confirmation`` → renamed to
  ``test_engine_exact_subcommand_hides`` to match new behavior.
- New: ``test_engine_subcommand_trailing_space_excludes_space_from_range``
  and ``test_engine_subcommand_trailing_space_apply_does_not_double_space``.
- All 97 tests in ``test_tui_widgets.py`` pass.
- ``ruff check`` / ``ruff format`` clean.
- Local TUI smoke: ``/`` (no crash, top-level popup), ``/mcp ``
  (subcommand popup), ``/mcp a `` + Tab → ``/mcp add ``.

Refs the din0s review comments on PR #273. CLI path tests and the
``category`` ClassVar follow-up are deferred to a separate PR (the
former is a test-suite addition; the latter is already absent from
``base.py`` on the current branch).

* fix: address remaining review items (help duplication, CLI tests, stub sync, docstrings)

- mcp.py: auto-generate help text from subcommands ClassVar (#1)
- tests/test_cli_completion.py: add 9 CLI completer tests (#2c)
- test_tui_widgets.py: sync _apply_selected_completion stub with real code (#4)
- mcp.py + interactive.py: add docstrings to key functions (#8)

* fix: hide completions on exact subcommand match regardless of trailing space

Remove the
ot has_trailing_space guard from the exact-subcommand
check.  Previously /mcp list  (with trailing space) would still
return candidates, causing Tab to oscillate between adding and removing
the trailing whitespace.  Now the engine hides whenever the subcommand
is an exact match, same as the top-level rule.

Added test_engine_exact_subcommand_with_trailing_space_hides to cover
the scenario din0s flagged.

* refactor: use StrEnum for CompletionResult.kind

Replace plain str with CompletionKind(StrEnum) for type safety.
Backward-compatible with existing string comparisons.

* fix: normalize @file completion tuples to CompletionCandidate

complete_file_mention() returns list[tuple[str, str]] but the TUI
rendering/apply code expects objects with .text/.description.
Wrap tuples in CompletionCandidate to prevent AttributeError crash.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-13 17:06:23 +00:00
Xi Zhang 05a5e5f8a0 fix: session lost after evoscientist restart (supersedes #278) (#279)
* fix: session lost after evoscientist restart

* feat: Implement memory worker thread deletion on completion

- Added synchronous and asynchronous functions to delete memory worker threads after they finish execution, ensuring no residual checkpoints are left in the database.
- Enhanced `_watch_memory_worker_run_sync` and `_watch_memory_worker_run_async` to invoke deletion functions upon confirming worker completion.
- Introduced tests to verify that worker threads are deleted correctly upon successful completion and that failures in deletion do not affect the overall worker status.
- Updated session management to ensure that only relevant threads are restored from the database, preventing exposure of internal or unrelated workspace threads.
- Implemented a purge function to clean up leftover worker checkpoints during server startup.

* feat: Implement short thread ID display for CLI and session hints

* fix: ensure proper accounting and deletion order for memory worker threads

---------

Co-authored-by: z00827015 <zhoulun1@huawei.com>
2026-06-11 23:51:51 +01:00
Ziheng Zhang 4b6a969df2 refactor(agent): make create_cli_agent(config=, chat_model=) pure (#267)
* refactor(agent): make create_cli_agent(config=, chat_model=) pure

Re-applies the #183 purity refactor on top of the observation-memory
lifecycle that landed in #259, integrating the two cleanly.

create_cli_agent gains a pure path: when both `config` and `chat_model`
are passed it builds the agent entirely from locals and writes none of
the cached module globals (`_config`, `_chat_model`, `_chat_model_key`,
`_EvoScientist_agent`). `/model` commits the switch via
`set_active_config` / `set_chat_model_instance` only after a successful
build, so a failed rebuild leaves the session on the original model
(replaces the old snapshot/restore rollback).

Supporting changes:
- Extract `set_active_config` (write-half of `_ensure_config`),
  `_apply_env_from_config`, `_build_chat_model`, and
  `set_chat_model_instance`.
- Thread `cfg` / `chat_model` through `_get_default_middleware`,
  `_build_base_kwargs`, `load_mcp_and_build_kwargs`,
  `_maybe_swap_async_subagents`, and `_inject_subagent_middleware` so the
  pure path never falls back to the global-writing `_ensure_config()` /
  `_ensure_chat_model()`.
- Integrate with #259's memory middleware: subagent context-editing
  middleware binds the threaded `chat_model`, and the configured system
  prompt / memory controls read the threaded `cfg` (new threading vs the
  original #183, required because #259 made these paths read config).
- Consolidate `cfg` resolution to one `cfg if cfg is not None else
  _ensure_config()` at the top of each kwargs builder, matching the
  pattern already used in the other config-aware helpers.

* fix(agent): keep pure tool selector off global cache

* fix(model): apply config switch in place to preserve reference integrity

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-06-08 18:38:13 +01:00
dinos 92d95dee68 feat(memory): add observation memory lifecycle (#259)
* feat(memory): add observation memory lifecycle

Add file-backed observation memory with deterministic markdown records,
structured record_observation tooling, startup indexing, and
profile/observation prompt guidance.

Launch post-turn and post-subagent EvoMemory workers through LangGraph
dev so completed runs can update profile memory, save durable
observations, and write subagent execution summaries without blocking
the active agent.

Wire memory middleware into the main agent, subagents, async graphs, TUI
status reporting, worker activity accounting, and observation-aware
research prompts, with regression coverage for storage, lifecycle
scheduling, graph registration, status display, and stream reset
behavior.

* fix(cli): sync background agent server on resume

Resume flows now need to keep the LangGraph dev background server
aligned with the active workspace even when async subagents are
disabled. EvoMemory workers use that server too, so gating resume-time
sync on enable_async_subagents could leave workers pinned to the launch
workspace after resuming a thread from another workspace.

Run workspace sync unconditionally for Rich CLI and Textual resume
paths, while preserving WorkspaceMismatchError handling so failed sync
aborts the resume before mutating the active thread or workspace.

Propagate aborted resume callbacks through the command UI so
channel-issued /resume commands do not send false success or history
output. Channel slash dispatch now treats CommandManager-caught command
errors as command errors and skips completion hooks for those failed
commands.

Add regression coverage for disabled async subagents, callback aborts,
and channel command error reporting.

* fix(cli): prepare serve resume workspace before adopting

Load the resumed workspace agent and sync the background server as a
single pre-adoption step. Restore the previous active workspace if
preparation fails so serve mode keeps using the old session
consistently.

* fix(memory): untrack abandoned worker status watches

Stop treating watcher shutdown as confirmed worker completion. Terminal
worker statuses still count memory deltas, while poll failures or
watcher setup failures now remove the active run without crediting
partial outputs.

* fix(cli): report channel command failures accurately

Treat command_error as a None sentinel so empty error strings still
fail, and let TUI resumes continue only on non-mismatch
background-server sync failures while reporting degraded mode.

* fix(stream): clear memory counters for resume streams

Reset completed-memory counters for every new agent stream, including
Command-based HITL and resume streams, so saved-memory indicators do not
leak across turns.

* docs(tools): make observation recording guidance conditional

Clarify that agents should call record_observation only when the
observation tool is available, preserving the existing durability and
usefulness criteria.

* feat(config): add controls for profile and observation memory

Add config flags for profile memory, observation memory, observation
writer placement, and background memory workers.

Wire the controls through main agents, subagents, EvoMemory middleware,
and memory lifecycle workers so observation writes can be assigned to
the live agent, subagent worker, both, or neither. Keep turn memory
workers profile-only and make prompts reflect the available observation
read/write paths. Skip langgraph dev startup when neither async
subagents nor memory workers need the background server.

Add coverage for config parsing, prompt gating, middleware wiring, and
worker tool availability.

* test(cli): include memory defaults in serve config stubs

* fix(memory): offload async worker launch blocking calls

Run the langgraph-dev health check and memory-output snapshot in worker
threads from the async EvoMemory launcher so it does not block the event
loop.

* chore(memory): harden turn worker subagent guardrail

* chore(memory): refresh profile context per request

* fix(memory): offload async profile file reads

* fix(memory): offload async worker completion accounting
2026-06-05 15:11:20 +01:00
Wiktor Cupiał 9e51ec6fdd feat(cmd): add /model-fallback command (#196)
* feat(cmd): add /model-fallback command

* fix: apply feedback

* fix: lock usage with _fallback_chain

* fix: apply feedback

* fix: apply feedback

* feat: add tests

* fix: tests

* Update EvoScientist/middleware/model_fallback.py

Co-authored-by: dinos <dinospk1999@gmail.com>

---------

Co-authored-by: dinos <dinospk1999@gmail.com>
2026-05-07 15:35:26 +02:00
dinos 5c829942d7 refactor(cli): replace channel module globals with ChannelRuntime (#197)
* refactor(cli): replace channel module globals with ChannelRuntime

Removes _cli_agent / _cli_thread_id from EvoScientist/cli/channel.py
and threads a ChannelRuntime via CommandContext.channel_runtime so
/model and /channel rebind without poking module-level state.

* fix(cli): address coderabbit review

- _auto_start_channel: bind ChannelRuntime only after
  _start_channels_bus_mode succeeds, so a startup failure no longer
  leaves a stale binding pointing at channels that never started.
- _sync_tui_command_completion (TUI) and the Rich CLI command-completion
  paths: rebind the runtime on thread rotation, not just agent swap, so
  /new and /resume keep ChannelRuntime in sync with the running thread
  (matches the serve-mode hook contract).
- test_hook_syncs_channel_runtime: pin ctx.thread_id explicitly so a
  bare MagicMock attribute can't silently mutate runtime.thread_id.
- New regression test covering the rebind-on-thread-rotation contract.
2026-04-30 18:06:22 +01:00
Ziheng Zhang da74c325d6 fix(channel): scope stop and restore resume history (#186)
* fix(channel): scope stop and restore resume history

* refactor(channel): simplify stop and resume patch

* Delete PR_MESSAGE.md

* fix(channel): address review feedback

* fix(channel): address remaining review bugs

* fix(channel): clean up stopped request handling

* fix(channel): preserve resolved replies and sync tui commands

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-25 16:12:49 +01:00
Xi Zhang 558360b558 feat: Enhance ModelPickerWidget for Ollama integration (#187)
* feat: Enhance ModelPickerWidget for Ollama integration

- Implemented a sentinel row for "Custom Ollama model..." in ModelPickerWidget, allowing users to input arbitrary model names.
- Updated action handling in ModelPickerWidget to manage transitions between list and input modes.
- Added async model discovery for Ollama models, integrating with the /model command to fetch locally installed models.
- Created tests for Ollama model discovery and ModelPickerWidget behavior, ensuring proper functionality and user experience.
- Refactored validate_ollama_connection and discover_ollama_models for improved error handling and response management.

* fix: Simplify code by removing unnecessary line breaks in ModelPickerWidget and test cases

* fix: Restore globals on set_chat_model failure to prevent half-switched session

* fix: Improve error handling in ModelCommand by restoring globals on failure
2026-04-25 00:37:48 +01:00
Xi Zhang 3831198f19 fix(chat): resolve model switch lag by tracking model/provider key (#180)
* fix(chat): resolve model switch lag by tracking model/provider key for cache invalidation

* style: format set_chat_model function for improved readability

* fix(chat): improve model switching logic to prevent unnecessary cache rebuilds

* fix(model): ensure globals are restored on agent load failure to prevent model switch issues
2026-04-24 11:36:33 +01:00
Xi Zhang 92df3e1844 Refactor/cli command manager (#178)
* feat(cli): migrate command handling to CommandManager and enhance UI interactions

* feat(cli): add /clear and /help commands to enhance user experience

* feat(cli): implement /new and /resume commands with interactive session management

* Refactor MCP and Skills Command Handling

- Moved the interactive picker style to a centralized widget for consistency across MCP and Skills commands.
- Updated the MCP command to remove the old command dispatch logic, delegating to the new InstallMCPCommand.
- Enhanced the Skills command to utilize a new interactive picker for skill selection, improving user experience.
- Implemented cancellation handling in the picker to differentiate between user cancellations and empty selections.
- Added comprehensive tests for the new command structures and picker functionalities to ensure reliability.

* refactor(cli): streamline CommandManager dispatch and remove deprecated command set

* refactor(cli): enhance error handling and state management in ChannelCommand and RichCLICommandUI

* refactor(cli): update lifecycle callback terminology and improve async prompt handling in RichCLICommandUI

* refactor(cli): enhance SlashCommandCompleter to dynamically fetch workspace directory for autocompletion

* refactor(cli): unify quit handling in RichCLICommandUI with shared _stop helper

* refactor(cli): remove hardcoded slash commands and utilize command manager for dynamic completion

* refactor(cli): update MCP and skills command files for improved clarity and organization

* refactor(mcp_ui): remove unnecessary newline in _show_mcp_config function
2026-04-24 10:38:44 +01:00
dinos 65828e9666 perf(cli): cut startup latency and defer MCP loading to the background (#171)
* perf(cli): cut startup time of `evosci --help` from ~2.2s to ~0.3s

Module-level imports were eagerly pulling in langchain.chat_models (with
the whole anthropic/openai/google stack), langgraph, textual, and
prompt_toolkit on every invocation — even for `--help` or `config list`.

Defer those with PEP 562 `__getattr__`, using `lazy_loader.attach` (SPEC-1,
the scientific-python standard) where it's a clean attach pattern:

- `EvoScientist/llm/__init__.py`: attach `models` lazily so importing
  `context_window` from this package no longer drags in langchain.
- `EvoScientist/stream/__init__.py`: attach display/events lazily; split
  the shared Rich `Console` singleton into a new lightweight
  `stream/console.py` so callers that only need `console` skip the
  `stream.events` → `langchain_core.messages` chain.
- `EvoScientist/cli/__init__.py`: hand-rolled `__getattr__` (reaches into
  `..stream.state`, which `lazy_loader` doesn't cover) so `commands` and
  `app` are the only eager loads.
- `EvoScientist/cli/commands.py`: move `cmd_interactive`/`cmd_run` to
  in-function imports so prompt_toolkit + textual only load when the
  interactive path actually runs.
- `EvoScientist/cli/_constants.py`: read `AGENT_NAME` on demand so
  `build_metadata` doesn't eagerly import `sessions` (langgraph/aiosqlite).

Adds `lazy-loader>=0.5` as a dependency.

* feat(cli): defer MCP tool loading with live per-server progress

The CLI was blocking ~5 s on MCP tool enumeration before the first
prompt appeared. Move the agent construction off the event loop and
surface per-server progress so the user can interact immediately and see
what's happening.

MCP client:
- Add an `on_progress` callback to `load_mcp_tools` / `aload_mcp_tools`
  / `_load_tools` emitting `start` / `success` / `error` events per
  server.
- Fan connection attempts out with `asyncio.gather` so latency no longer
  scales linearly with server count; cap simultaneous attempts at
  `_MAX_CONCURRENT_CONNECTIONS` (8) via a semaphore so a big stdio fleet
  doesn't spawn every subprocess at once.

Agent wiring:
- Plumb `on_mcp_progress` through `create_cli_agent` / `_load_agent` /
  `load_mcp_and_build_kwargs` so CLI and TUI can plug in collectors.

CLI (`cmd_interactive`):
- Run `_load_agent` in a background thread via `asyncio.to_thread`; the
  prompt and banner render immediately.
- `_await_agent_ready()` awaits the task before each agent-using site
  (first turn, channel messages, `/channel`, `/compact`). Raises if
  called without a prior `_start_agent_load` instead of silently
  reloading without the SQLite checkpointer.
- Pre-prime the progress dict from `load_mcp_config()` so the
  bottom-toolbar's `N/M` denominator is stable from the first render.
- Wrap `session.prompt_async` in `patch_stdout(raw=True)` so
  `console.print` from the worker-thread progress callback lands cleanly
  above the prompt as inline chat messages instead of stomping the
  prompt cursor.

TUI (`EvoTextualInteractiveApp`):
- Same background load + `_await_agent_ready()` gates on every
  `self._agent` read.
- New `MCPLoaderWidget` mounted at the top of `#input-shell` shows a
  header with `N/M` and one live row per server (spinner → ✓ / ✗ with
  tool count or error detail). On completion:
  - all-clean loads auto-dismiss ~2.5 s later;
  - cache hits (no events ever fired) dismiss immediately rather than
    flashing a misleading "0/N loaded";
  - failures keep the widget mounted so the user can read the errors.
  - `dismissed` property lets the app clear its ref so late events from
    slow servers become no-ops. The error branch of `_on_agent_loaded`
    also settles the widget so a load failure can't leave the spinner
    animating forever.
- Chat input is `disabled` while MCP resolves — no placeholder hack, no
"waiting…" system message.

Shared:
- Hoist braille spinner frames to `status_bar.SPINNER_FRAMES` and import
  them in the TUI widget so CLI and TUI animate in sync.

Tests:
- Extend `test_agent_mcp_cache` fakes to accept the new `on_progress`
  kwarg.
- New `TestLoadToolsProgressCallback` in `test_mcp_client` exercises the
  event sequence for success/failure/mixed fleets, verifies a buggy
  callback doesn't break the load, and asserts the semaphore caps
  in-flight connections.

* style: ruff

* chore: update uv.lock

* chore: uv.lock

* fix: coderabbit issues

* style: fmt

* fix: move _await_agent_ready inside try block

* fix(tui): auto-dismiss MCP loader widget on failure

The widget was designed to stay mounted on failure so the user could
read error detail, but since it's pinned above the input it never went
away in practice — just permanent banner clutter.

Auto-dismiss on failure too, with a longer grace (12s vs 2.5s) so the
error summary stays readable.

* fix: address second coderabbit pass

- Channel handlers (CLI + TUI): catch agent-load failures so the
  channel request doesn't hang; CLI moves `_await_agent_ready()`
  inside the existing try/except, TUI catches explicitly and calls
  `_set_channel_response` with the error.

- Stale background loads: `prev.cancel()` only stops the asyncio
  wrapper, not the thread running `_load_agent`. Added a generation
  token (`agent_load_id` / `self._agent_load_id`) and gated both
  progress and completion callbacks on it so a superseded load can't
  clobber the current session's state or UI.

- TUI prompt lifecycle: added `_agent_load_pending()` and gated the
  `_process_channel_message` / `_handle_command` finally blocks on it
  so `/new` or `/resume` invoked from a command keeps the prompt
  disabled until the fresh load settles.

- TUI readiness failures: `_run_turn` and `_handle_command` now
  catch exceptions from `_await_agent_ready()` and surface a
  "Agent failed to load: …" system message instead of letting the
  exception escape into Textual's traceback panel.

* refactor(cli): share background agent loader between CLI and TUI

The CLI and TUI were carrying near-identical copies of the same
background-load state machine: the `agent_task`, the `agent_load_id`
generation token, the gated progress/completion callbacks, and the
per-server progress dict. Every CodeRabbit finding on that lifecycle
had to be fixed in both files.

Extract it into `cli/_agent_loader.py`:

- `MCPProgressTracker` — owns the `server -> (state, detail)` dict;
  exposes `prime`, `record`, `snapshot`, `totals`.

- `BackgroundAgentLoader` — owns `agent`, the in-flight task, and the
  generation token. Exposes `start(**loader_kwargs)`, `await_ready()`,
  `is_pending`. Internally gates all progress/completion callbacks by
  generation so a superseded load can't clobber the current session.
  UI-specific rendering plugs in via `on_progress` / `on_success` /
  `on_failure` callbacks.

Both surfaces now just wire their UI hooks; the loader file holds no
Rich / prompt_toolkit / Textual dependencies. Net -345 lines from
`interactive.py` + `tui_interactive.py`; +20 unit tests pinning the
lifecycle (generation filtering, cache-hit short-circuit, failure
reset, progress ordering).

* refactor(cli): make _on_done the sole authority for agent state transitions

await_ready no longer sets self.agent — it just awaits the task and
reads what _on_done already wrote. Eliminates the dual-write overlap
(asyncio guarantees done-callbacks fire in registration order).

* fix(tui): let users type during MCP load, only block on send

Remove prompt-disabling during background agent load — the TUI now
matches the CLI approach where the input stays enabled and only gates
on await_ready() at submit time. The MCPLoaderWidget still provides
visual feedback that loading is in progress.

* fix(loader): preserve real load error on await_ready; dedup failure message

CodeRabbit flagged two issues with the new loader:

1. After a failed load, `_on_done` nulled `self._task`, so the next
   `await_ready()` hit the "before start()" branch and the CLI wrapper
   remapped it to a misleading "checkpointer not available" message —
   losing the real exception (bad MCP config, network, etc.).

   Keep `_task` set on failure so `await_ready` re-raises the real
   exception. Added `needs_restart` so TUI's auto-retry check stays a
   one-liner and doesn't need to reach into task internals.

2. TUI reported each load failure twice: once from
   `_on_agent_load_failure` (the done-callback) and once from each
   caller of `_await_agent_ready` (`_run_turn`,
   `_process_channel_message`, `_handle_command`) catching the re-raise.

   `_on_agent_load_failure` is now the sole local reporter; callers
   just handle control flow (return cleanly, set channel response to
   unblock remote).

* fix(cli): wire /model handler through the agent loader

The /model command from main (merged via f1f0d7c) still reached for
`state["agent"]` (CLI) and `self._agent` (TUI) — both removed by the
background-loader refactor. CLI raised KeyError on first invocation;
TUI raised AttributeError. Writes to the old fields also had no effect
because every other code path now reads from `agent_loader.agent`, so
the model switch would have silently failed.

Route everything through the loader: `await _await_agent_ready()` up
front so /model doesn't race with the initial background load, build
the `CommandContext` with the current agent, and sync `ctx.agent` back
into `agent_loader.agent` (plus channel globals) when the command
replaces it.

* fix(cli): isolate progress callback, capture awaited agent, gate by requires_agent

Three CodeRabbit findings on the loader + command dispatch path:

- Wrap ``_on_progress`` in try/except inside the loader's gated wrapper
  so a buggy UI adapter can't bubble into ``loader_fn`` and fail the
  whole background load. The MCP client already protects this, but
  defence-in-depth keeps the loader self-contained.

- In CLI channel + main-loop streaming, capture the agent returned by
  ``_await_agent_ready()`` and pass that into ``run_streaming`` rather
  than reading ``agent_loader.agent`` after a subsequent ``await``.
  A concurrent ``/new``/``/resume``/``/model`` could have swapped it.

- Add ``requires_agent: ClassVar[bool] = False`` to ``Command`` and
  mark ``/compact``, ``/model``, ``/channel`` as ``True``. TUI dispatch
  sites (channel and keyboard) now check ``cmd_manager.resolve(...)``
  and only wait for readiness when the command actually needs the
  agent. ``/mcp add``, ``/skills``, ``/new`` etc. no longer deadlock
  behind a failing MCP load they are meant to fix.

* fix(cli): guard sync-back, subcommand-aware gating, /model adopt-path

Three CodeRabbit findings on command dispatch:

- ``_handle_command`` unconditionally synced ``ctx.agent`` back into
  ``agent_loader``.  For non-agent commands ``ctx.agent`` is ``None``,
  so ``/threads`` / ``/mcp`` / ``/skills`` (etc.) could clobber a valid
  loaded agent — and rebind channel globals to ``None``.  Guard the
  sync on ``ctx.agent is not None``.

- ``/channel status`` and ``/channel stop`` don't touch ``ctx.agent``
  but the class-level ``requires_agent = True`` blocked them behind
  agent readiness.  Added ``Command.needs_agent(args)`` (defaults to
  ``requires_agent``) so ``/channel`` can override with subcommand
  awareness; kept the class flag for the common case.

- ``/model`` builds a new agent from scratch, it never reads the
  existing one — gating it on readiness meant a broken provider
  blocked the command that would fix it.  Flipped it to
  ``requires_agent = False`` and added ``BackgroundAgentLoader.adopt``
  so the UI can seat the replacement and supersede any in-flight
  load (the generation token keeps a late completion from clobbering
  the adopted agent).

Bonus cleanup: ``CommandManager.resolve`` now returns
``(command, args)`` so callers can invoke ``needs_agent`` without
re-implementing ``shlex`` parsing.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-22 18:22:33 +02:00
Wiktor Cupiał 28f3e81b4c feat: add /model command for changing models inside TUI/CLI (#162)
* rebase main

* fix: apply pr comments

* fix: fix critical issue

* fix: ordering /model in cli mode

* feat: refactor /model command handling and add Rich CLI support

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-04-22 14:43:21 +01:00
Xi Zhang aa3dd00409 feat: add support for session resumption with --resume flag and enhan… (#170)
* feat: add support for session resumption with --resume flag and enhance thread ID resolution

* refactor(tests): streamline help output testing for --resume flag

* feat: enhance session resume functionality with improved thread ID resolution and SQL wildcard handling

* feat: improve error handling for resume hint retrieval in interactive modes

* refactor: streamline logging for print_resume_hint failure in interactive mode

* feat: implement deferred scrolling for Markdown-heavy content in interactive mode
2026-04-21 21:26:40 +01:00
Xi Zhang 210e8864f6 feat(memory): migrate MEMORY.md to global path & enhance ask-user prompts (#161)
* feat(prompt): enhance user interaction with multiple-choice and free-text questions

* refactor(paths): rename MEMORY_DIR to MEMORIES_DIR for consistency

* style(tests): format code for better readability in test cases

* feat(prompt): add validation for 'other' option in user prompt

* feat(prompt): refactor validation logic for user prompts and add skip option

* feat(style): refactor to use shared _PICKER_STYLE from interactive module
2026-04-16 15:33:45 +01:00
Xi Zhang 65db3a4fcd Add status bar and compact summary widgets with context window resolu… (#152)
* Add status bar and compact summary widgets with context window resolution

- Implemented a shared status bar for CLI and TUI frontends, including helpers for managing session metrics and context windows.
- Created a `CompactSummaryWidget` for displaying manual summaries in a collapsible format.
- Introduced a `CompactingWidget` to indicate ongoing compacting processes.
- Added a base class `TimedStatusWidget` for widgets that require a timer.
- Developed context window resolution helpers to retrieve context window sizes from various model attributes.
- Enhanced tests for context window resolution and status bar functionalities, ensuring accurate behavior across different scenarios.
- Updated existing tests to cover new features and maintain code quality.

* refactor(Channel): simplify lambda function in _send_with_retry method

* feat: enhance context editing logic and improve error handling in StreamState

* refactor(Channel): streamline lambda function in _send_with_retry method

* feat: rename auto-approve option to auto-mode for unattended execution; update checkpoint queries to filter by agent name; improve compatibility validation logic

* feat: rename auto-approve option to auto-mode; update related logic and tests for improved unattended execution

* fix: correct formatting of console message for MCP server configuration status

* feat: add check for None summary_message in _apply_summarization_event to prevent errors

* feat: enhance _load_checkpoint_messages to validate message format and apply summarization event
2026-04-12 17:47:37 +01:00
Xi Zhang ff15f515cc Release/v0.0.7 (#151)
* chore(assets): update wechat_group image file

* Refactor code structure for improved readability and maintainability

* feat(backends): enhance MergedReadOnlyBackend with improved ls, grep, and glob methods

* fix(docs): update WeChat QR code image link in README files

* feat(skills): enhance skill management to support global and workspace tiers

* style: apply ruff format to skills_cmd and commands/implementation/skills

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(skills): improve uninstall_skill to prevent removal of built-in skills

* fix(docs): update skill installation documentation for clarity on global and user directories

* fix(skills): enhance uninstall_skill to validate skill directory before removal

* fix(skills): improve error handling in install_skill and uninstall_skill for directory creation and validation

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-10 20:04:05 +01:00
Xi Zhang fe6a2c4b83 fix: resolve #93 review comments and close #95 (#92 regression fix included) (#96)
* feat: enhance file mention parsing with deduplication and warning handling

* feat: optimize ancestor grouping by improving path comparison efficiency

* fix: update middleware injection to use extend for better readability

* Refactor code structure for improved readability and maintainability

* feat: update README files to include AstaBench ranking and adjust award image layout

* update

* fix: deduplicate file mentions and improve warning message formatting
2026-03-25 17:01:46 +00:00
Xi Zhang 372a6272b7 fix(onboard): Windows npx detection & rename /install-skills to /evoskills (#79)
* fix: improve npx detection on Windows using shutil.which

* fix: rename /install-skills command to /evoskills for consistency
2026-03-20 10:42:56 +00:00
Jan Piotrowski 81316ffb2d chore: resolve all ruff linting and static analysis errors in EvoScientist/
- Fix RUF006: Implement background task tracking in Discord, iMessage, WeChat, and TUI to prevent premature GC of fire-and-forget tasks.
- Fix B904: Add explicit exception chaining (raise ... from) across all exception handlers.
- Fix RUF012: Annotate mutable class attributes with ClassVar for command arguments and media maps.
- Fix B008: Refactor Typer commands in cli/commands.py to use Annotated for argument and option defaults.
- Fix B023/B018: Resolve late-binding issues in lambdas and remove useless expressions.
- Fix syntax errors in retry.py docstrings and models.py lambda parameter ordering.
2026-03-19 17:04:03 +01:00
Jan Piotrowski 4a3d6c0318 chore: add ruff lint rules and turn on formatting 2026-03-19 17:04:02 +01:00
dinos 7ebaa2cec0 feat: implement command to browse and install MCP servers (#65)
* refactor(mcp): extract MCP server registry from onboard into shared module

Move _RECOMMENDED_MCP_SERVERS, _install_pip_package, and
_pip_install_hint from config/onboard.py into mcp/registry.py as a
shared MCPServerEntry dataclass and registry functions. This enables
reuse by the new /install-mcp command and marketplace integration.

* feat(mcp): add /install-mcp command for browsing and installing MCP servers

Interactive browser for MCP servers (built-in registry + EvoSkills
marketplace), supporting three modes:
- /install-mcp         — interactive tag filter + checkbox selection
- /install-mcp <name>  — direct install by name or tag pre-filter
- /install-mcp file.yaml — import servers from arbitrary YAML file

Also available as /mcp install and EvoSci mcp install. Includes TUI
browser widget (MCPBrowserWidget) mirroring the skill browser UX.

* chore(tavily): conditionally pass tavily_search when TAVILY_API_KEY is set

* fix(tui): Enter key detection for /install-mcp

- Fix message handler names: Textual converts MCPBrowserWidget to
  mcpbrowser_widget (not mcp_browser_widget), so Confirmed/Cancelled
  messages were never received by the app
- Distinguish empty selection from cancel in result handling

* chore(widgets): stop auto-advancing cursor on Space toggle in browser widgets

* style: linter

* refactor(mcp): simplify MCP registry to marketplace-only

Remove built-in server list and arbitrary YAML import — all server
definitions now come from the EvoSkills marketplace (mcp/*.yaml).
Onboarding filters by the `onboarding` tag instead of a hardcoded list.

* refactor(mcp): consolidate /install-mcp into /mcp install

Remove standalone /install-mcp command — use /mcp install as the
single entry point. CLI adapter now delegates logic to the shared
InstallMCPCommand class, keeping only the questionary UI layer.

* chore(cli): rm reference to yaml import
2026-03-19 12:45:32 +00:00
Jan Piotrowski f9734c7e67 feat: extract commands from TUI to commands module
- adjust discord max message length to 2000 (the actual limit)
- /install-skill will fallback to searching EvoSkills repository if path not found to allow for installing skills from channels
- remove discord config tests (it asserted @dataclass behaviour)
2026-03-18 16:59:50 +01:00