01845f43110ad444b7e2a61b920effdf7e719029
136 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
01845f4311 |
refactor: route middleware display events through an injected event sink (#343)
* chore: add pytest-asyncio in auto mode * test: migrate channel and stream tests to native async Convert run_async() wrapper tests to plain 'async def test_*' under pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a coroutine awaited at every call site. * test: migrate command and model/middleware tests to native async Convert run_async() wrappers (import, alias, and fixture forms) to plain 'async def test_*'. Multi-call tests merge onto one loop as sequential awaits; none asserted on loop identity. * test: migrate TUI, notifier, gateway, and session tests to native async TUI/notifier/gateway files convert run_async wrappers to plain async tests. test_sessions.py's unittest.TestCase classes move to unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async methods on plain TestCase; converting blindly would have made ~70 tests silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget in test_tui_widgets.py drops its TestCase base for the same reason. * test: replace direct asyncio.run() calls with native async tests Convert tests that called asyncio.run() (directly or via a local _run helper) to plain 'async def test_*'; delete the local helpers. * test: drop undeclared anyio markers and delete run_async helper The @pytest.mark.anyio tests relied on anyio being a transitive dep of httpx; auto-mode pytest-asyncio collects them natively. run_async() and its fixture are unreferenced after the migration, so remove them — pytest-asyncio's per-test loop teardown covers the pending-task cancellation the helper existed for (verified: full suite runs with no 'Event loop is closed' errors or destroyed-task warnings). * test: add autouse fixture for watcher cleanup * refactor: remove redundant hasattr calls * refactor: add typed middleware event sink and thread through assembly Add MiddlewareEventSink protocol + NoOpSink in middleware/events.py with a documented any-thread non-blocking contract (contract test uses a deliberately-slow fake sink). Thread an optional `events` parameter through create_cli_agent -> _get_default_middleware -> tool selector / model fallback constructors; subagent stacks are always forced to NoOpSink. * refactor: inject a notifier port into async-watcher and background middleware Add public pre_cancel_watcher() and enqueue_task_notification() to cli/async_notifier.py and a small NotifierPort protocol (middleware/notifier.py) that the module satisfies structurally. AsyncWatcherMiddleware and BackgroundExecutionMiddleware now receive the port by constructor injection at the composition root, deleting the lazy 'from ..cli import async_notifier' imports and the private _watcher_by_thread / _enqueue pokes. * refactor: invert tool-selection ownership onto a frontend event sink The adaptive tool selector now reports on_tool_selection_started / on_tool_selection / on_tool_selection_ended to the injected sink instead of writing four process-global module variables. The frontend sink (stream/sink.py FrontendEventSink) owns the selected/total/active state with consume-once + dedup-vs-last-emitted semantics; stream/tool_selection.py reads that sink object (a ToolSelectionView) rather than reaching into tool_selector's globals. Deleted: the 4 module globals, the cross-module mutations in tool_selection.py, the track_stream_selection flag, the now-vestigial _ToolSelectionTrackerMiddleware, reset_tool_selection_state_for_tests, and the autouse conftest fixture. The sink is threaded from the two interactive frontends through create_runtime_gateways -> LocalGraphGateway (read side) and _load_agent -> create_cli_agent (write side); subagent / headless stacks get NoOpSink. * refactor: route model-fallback narration through the injected event sink Delete the _ui_emit_fn / set_ui_emit module global and the ..stream.console import from model_fallback.py. The fallback middleware now reports through its injected sink: the fallback transition via the structured on_model_fallback (the frontend formats the '-> Falling back to ...' line), and the surrounding narration (primary-failure header, per-attempt outcome, exhaustion, non-fallbackable rejection) via emit_fallback_notice, preserving the exact user-facing text. The TUI binds its _append_system as the sink's fallback display where it used to call set_ui_emit (cleared on exit); the Rich CLI's sink prints to the console. _try_fallbacks / _guard_and_fallback take the sink. * refactor: declare events on the GraphGateway protocol Both gateway implementations now carry an explicit events attribute (LangGraphServerGateway holds None — no frontend renders middleware events across the HTTP boundary), so the four call sites use plain attribute access instead of getattr probing an implicit contract. * refactor: bind fallback display via the closure-scoped concrete sink The App methods used gateway.events (typed as the read-side view) and hasattr-probed for the concrete FrontendEventSink API. The enclosing factory creates that sink two hundred lines up — close over it directly: no probing, fully typed, and it becomes a constructor parameter naturally when the App class is hoisted out of the factory. * fix: end tool selection before fallback handler * fix: keep fallback display errors non-fatal * fix: preserve selector suppression for default streams * fix: restore fallback notice console display * refactor: consolidate fallback narration events * refactor: clean middleware event sink plumbing * fix: type gateway session events * refactor: make all event protocols runtime-checkable MiddlewareEventSink already carried @runtime_checkable (the stream binding guard isinstance-checks it); ToolSelectionView and SessionEvents now match, so mirroring that pattern against any of the three protocols works instead of raising TypeError. * fix(cli): close QuickJS workers after one-shot failures * fix(cli): honor no-thinking in final output * fix(channels): report failed startup accurately * fix(channels): make Telegram cleanup idempotent * fix(tui): skip command sync during exit * fix(channels): preserve startup state during retries * refactor(channels): share pending startup status * refactor(cli): expose channel startup snapshot * fix(tui): move channel startup off event loop * test(channels): release retry gate on assertion failure --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
db1abce8d8 |
refactor: extract a shared HITL/ask_user interaction engine (#342)
* chore: add pytest-asyncio in auto mode * test: migrate channel and stream tests to native async Convert run_async() wrapper tests to plain 'async def test_*' under pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a coroutine awaited at every call site. * test: migrate command and model/middleware tests to native async Convert run_async() wrappers (import, alias, and fixture forms) to plain 'async def test_*'. Multi-call tests merge onto one loop as sequential awaits; none asserted on loop identity. * test: migrate TUI, notifier, gateway, and session tests to native async TUI/notifier/gateway files convert run_async wrappers to plain async tests. test_sessions.py's unittest.TestCase classes move to unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async methods on plain TestCase; converting blindly would have made ~70 tests silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget in test_tui_widgets.py drops its TestCase base for the same reason. * test: replace direct asyncio.run() calls with native async tests Convert tests that called asyncio.run() (directly or via a local _run helper) to plain 'async def test_*'; delete the local helpers. * test: drop undeclared anyio markers and delete run_async helper The @pytest.mark.anyio tests relied on anyio being a transitive dep of httpx; auto-mode pytest-asyncio collects them natively. run_async() and its fixture are unreferenced after the migration, so remove them — pytest-asyncio's per-test loop teardown covers the pending-task cancellation the helper existed for (verified: full suite runs with no 'Event loop is closed' errors or destroyed-task warnings). * test: add autouse fixture for watcher cleanup * refactor: remove redundant hasattr calls * refactor: extract shared HITL/ask_user interaction grammar Extract prompt/question formatting, the reply grammar (approval letters, ask_user choice letters + the 'Other' sub-flow, stop-commands), the ApprovalPolicy (config auto-approve rule + session registry + session-key derivation), per-flow timeout constants, and the bilingual feedback strings into channels/interaction.py. Both drivers now point at the shared functions: this reverses cli/channel.py's imports of consumer privates and closes the /stop drift at the parsing layer (serve-mode ask_user now checks stop-commands before parsing an answer, matching the CLI path). * refactor: add interaction engine + registry; port InboundConsumer Introduce InteractionIO (transport adapter Protocol), PendingReplyRegistry (one asyncio-based reply router per process), and the engine coroutines resolve_ask_user / resolve_approval in channels/interaction.py. Port InboundConsumer onto them: a _ConsumerIO adapter over bus.publish_outbound + the registry, one ApprovalPolicy replacing the config/session auto-approve checks, and a single reply-interception point (registry.try_resolve) replacing the parallel ask_user/HITL pending dicts. _resolve_ask_user and the approval section of _stream_with_hitl are now thin engine calls. Behavior unification (serve mode): an unrecognized HITL reply now declines with the shared 'Unrecognized reply' notice instead of rejecting-and-refeeding as a fresh turn, and /stop mid-approval cancels cleanly — both via the shared parser. * refactor: port CLI channel bridge onto the interaction engine Replace the ~250-line parallel bodies of channel_ask_user_prompt / channel_hitl_prompt with thin bridges that run resolve_ask_user / resolve_approval on the bus loop via run_coroutine_threadsafe(...).result() (outer = engine per-flow timeout + slack, so the engine's own timeout fires first). The 15s send timeout moves into the _BridgeIO adapter. Delete the _pending_hitl / _hitl_lock / _hitl_auto_approve module globals and the _register_hitl_wait / _try_set_hitl_reply / _pop_hitl_reply helpers, absorbed by one bus-loop PendingReplyRegistry + one ApprovalPolicy. The bus consumer feeds the registry via try_resolve ahead of normal enqueue. * refactor: restore serve-mode refeed for unrecognized HITL replies Gate-review fix: the engine no longer decides transport policy for unparseable approval replies. resolve_approval now returns an ApprovalOutcome carrying unrecognized_reply (raw text) when parsing fails, sending no feedback itself; recognized reject keeps the sharedi rejection message. Consumer driver (serve mode) restores the pre-engine semantics: an unrecognized reply rejects the pending action, confirms with the rejection message, and the text is re-dispatched as a NEW agent turn — _stream_with_hitl returns the captured text and _handle_message starts the refeed turn only after the current one has released the chat lock (old fall-through ordering). CLI bridge keeps its old no-refeed path byte-for-byte: 'Unrecognized reply. Action rejected.' and decline. Tests: serve refeed pinned end-to-end (prompt → unrecognized text → rejection feedback → text reaches the stream path as a new turn), CLI no-refeed pinned (notice sent, nothing enqueued), engine test updated to assert the outcome struct with no engine-side feedback. * refactor: polish the interaction engine surface - English feedback strings (Approved / Rejected / auto-approving) - drop the consumer's backwards-compatible re-exports and both modules' private timeout aliases; callers use the canonical interaction names - replace byte-for-byte prompt goldens with structural format tests and assert feedback via the shared constants instead of string literals - strip audit/design shorthand (R1/R2/G3, stage numbers) from comments * fix: propagate pending reply task cancellation * fix: preserve reply context when refeeding HITL replies * fix: honor HITL session grants without bus loop * fix: bound bridge waits by send latency * fix: handle empty ask_user replies explicitly * chore: remove stale interaction helpers * fix: harden interaction engine reply edge cases Review follow-ups on the interaction engine: - normalize ask_user choices before .get(): the tool args come from model JSON and only presence is validated, so plain-string choices must render and parse instead of crashing the turn - treat only None as an approval timeout, so an empty/media-only reply flows through the unrecognized path and serve mode refeeds it with its preserved context - intercept prompt replies before _get_thread_id so a consumed reply cannot create an orphan graph thread or touch the sender-session LRU - close engine coroutines the bridge failed to schedule (no bus loop / scheduling error) to avoid never-awaited warnings - clear pending-response and channel-request state in the bridge test fixture --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
1d117ff277 |
feat: completion enchancements (#302)
* feat: completion enchancements * fix: handle exception * fix: duplicate view * fix: remove deadcode * fix tab |
||
|
|
bd54eaa0a4 |
feat(cli): add --output-format stream-json for headless clients (#309)
* feat(cli): add --output-format stream-json for headless clients Emit EvoScientist's native event stream as line-delimited JSON on stdout in single-shot (-p) mode, with all human output redirected to stderr so stdout stays pure JSONL. Intended as the integration surface for programmatic clients (e.g. an agent runtime) that drive EvoSci headlessly. - stream/json_sink.py: write_events_as_json + stream_json sink, plus redirect_console_to_stderr helper for stdout purity - cli/interactive.py: cmd_run gains output_format; stream-json branch runs the sink instead of the Rich renderer - cli/commands.py: --output-format option + validation (stream-json requires -p; value must be text|stream-json) - docs/stream-json.md: event-schema contract + example transcript - tests: json sink serialization, CLI dispatch, console redirect, validation Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(cli): honor explicit --no-auto-mode over config in stream-json Address CodeRabbit review (discussion_r3514041123): the auto-mode override block only wrote to cli_overrides when the resolved value was True, so an explicit --no-auto-mode silently fell back to a config that enables auto-mode -- breaking "explicit flags always win" and leaving stream-json running unattended despite the warning. Write auto_mode=False when the flag is explicitly False. Add regression tests that capture the overrides. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
2b244888ec | feat(memory): autoskills (#319) | ||
|
|
7568a6bc1f |
fix(tui): keep welcome banner at top after /new (#311)
* fix(tui): keep welcome banner at top after /new PR #262 replaced scroll_end() with anchor() for free-scrolling. When /new clears a long anchored conversation, the anchor kept the viewport pinned to the (now empty) bottom, producing a negative scroll_y and pushing the welcome banner out of view. Reset the anchor and scroll to the top in clear_chat(), and restore the follow/new-content flags so the fresh session starts correctly. Closes #301 * fix(tui): suppress anchor when chat content fits viewport The previous fix for #301 only handled the /new path. din0s reported that the banner still dropped to the bottom after a normal short turn (user types 'hi', agent replies) — i.e. whenever the conversation fit in the viewport. Root cause is in Textual's compositor (textual._compositor): when a widget is anchored, scroll_y is recomputed via set_reactive, which bypasses the validator. If the anchored widget's content is shorter than the viewport, scroll_y goes negative on the next layout pass and the welcome banner is pushed below the visible region. PR #262 made _stream_with_widgets re-engage the anchor at the end of every turn via _anchor_chat, so the bug surfaced on any short reply that fit in the viewport. Markdown re-renders, status-bar updates, or any subsequent mount would then trip the compositor. Fix in three places: * _anchor_chat: only engage the anchor when max_scroll_y > 0; otherwise release and scroll_home so the banner stays at the top. * streaming anchor loop: if content shrinks below the viewport mid-stream (e.g. loading widget removed), release the anchor instead of leaving _anchored=True for the compositor to trip on. * clear_chat: keep the unconditional reset (children are removed asynchronously so a max_scroll_y check would be stale) but document why. Adds two regressions: * test_short_turn_keeps_banner_at_top_after_layout_refresh — the exact scenario din0s tested; fails with scroll_y=-10 on the previous code, passes with the fix. * test_long_turn_keeps_viewport_pinned_to_bottom — guards against regressing free-scrolling for overflowing conversations. Manually verified: 'hi' -> reply (banner stays at top) -> /new (banner at top) -> another turn (banner stays at top). * test(tui): address review feedback on banner-position regressions - extract `_release_anchor_and_pin_top` helper for the 3-line `anchor(False) + scroll_home(...)` pattern repeated in `clear_chat`, `_anchor_chat`, and the streaming loop - replace `pytest.skip` in `_capture_app` with a hard `RuntimeError` so a broken capture never silently passes - drop the redundant `load_agent` and `create_session_workspace` monkeypatches (the factory is given those as parameters, so the module-level symbols never run; added a comment explaining why) - add a defensive `_FakeChannelRuntime` patch for symmetry with the other module-level fakes - drop the local `_run` helper and use the `run_async` fixture from `conftest.py` (its teardown is better) * test(tui): replace _FakeChannelRuntime with _auto_start_channel no-op The _FakeChannelRuntime patch was ineffective because ChannelRuntime is just a dataclass — the real channel manager still started via _auto_start_channel, leaving pending tasks and non-hermetic test state. Per review feedback, stub _auto_start_channel directly instead. |
||
|
|
f2f010a350 |
feat(memory): observation linking (#307)
* refactor(gateway): create module for launching async/bg agents * refactor(memory): refactor worker launch around source context & output deltas * refactor(gateway): generalize async/bg module * refactor(memory): revamp worker launching * feat(memory): add observation linking * test(memory): remove redundant test branches * fix(memory): make 'supersedes' relation directional * fix(memory): don't create empty project observation dirs * fix(memory): schedule direct observations for linking * fix(cli): wait for observation linker before shutdown * fix(memory): block arbitrary writes to /memories * fix(linker): remove `linked_by` attribute from frontmatter * refactor(linker): rename base relationship to `comlpements` * fix(cli): bump worker wait to 2m * feat(tools): catch malformed tool calls & retry * feat(status): add linking result to statusbar * fix(linker): don't launch linker when observations are disabled * fix(memory): use posix paths * fix(watcher): call abort hook on error status * fix(watcher): delete thread on failed run creation * fix(watcher): preserve url * fix(observation): record session_id, drop unused fields * fix(memory): reject unsupported worker source types * refactor(backends): shared memory backend builder * fix(scheduler): resolve linker inputs outside lock * fix(memory): dont launch workers / record observations without thread_id * feat(memory): include related observations in tool results * fix(memory): skip malformed observation frontmatter * revert(tools): drop tool error handling changes from this PR * fix(memory): serialize observation link writes * fix(memory): queue observations written by aborted workers * fix(memory): track observation linker launch handoff * fix(memory): resolve cross-project related observations * fix(status): avoid recounting reason-only link updates * fix(memory): avoid rereading file for content * fix(linker): use neutral prose for bidirectional reasons * test(memory): coverage for aborted/failed launches * test(memory): cleanup & helpers * feat(linker): add observations index hint |
||
|
|
7ccfe68f3f |
feat: add scheduler functionality with cron-style task management (#306)
* feat: add scheduler functionality with cron-style task management - Implemented a new scheduler subagent to automate recurring tasks using cron expressions. - Enhanced the subagent factory to include the skill manager and auxiliary chat model for the scheduler. - Created a YAML configuration for the scheduler with a detailed system prompt and toolset. - Updated README files to include documentation on scheduled tasks and usage examples. - Added tests for the scheduler, including command execution, scheduling tools, and middleware integration. - Introduced new dependencies for timezone handling and ensured compatibility in the project configuration. * fix(async-notifier): ensure fallback hint is used for unknown notification kinds * feat: enhance scheduling functionality and improve system message handling |
||
|
|
bd307f3a11 |
refactor: LangGraph gateway layer for UI-agnostic graph and thread access (#295)
* feat(gateway): graph gateway protocol * refactor(cli): wire gateway in cli/tui * refactor(gateway): centralize runtime gateway init * chore(gateway): restrict RunRequest message type * feat(gateway): add langgraph server gateway * chore(cli): tighten serve runtime state typing * refactor(cli): route async task state reads through graph gateway * refactor(gateway): support graph targets in server gateway * refactor(cli): route session commands through graph gateway * refactor(cli): fold thread store under graph gateway * refactor(gateway): route graph state access through gateway * refactor(channels): wire graph gateway * refactor(memory): preserve graph threads for cloning * feat(gateway): add thread cloning * fix(tui): pass effective workspace for thread creation * chore(memory): add workspare dir to memory worker metadata * fix(sessions): filter preloaded UUID registy entries by the current scope * test(fakes): use https * refactor(consumer): consolidate imports * fix(stream): optional summarization event * fix(gateway): resolve abbreviated thread IDs by search * fix(gateway): page server thread listings * fix(gateway): emit pending interrupt events * style: fmt * feat(gateway): persist workspace_dir & model in thread metadata * fix(gateway): page server thread prefix resolution * fix(gateway): expose server thread list metadata * refactor: add back type def * refactor: tighten types * revert: add back worker thread deletion The worker thread forking changes are out of scope for now, so to maintain parity with the existing behavior we'll leave this intact. * fix(gateway): apply compaction to server thread history * refactor(stream): restore direct summary replay suppression * fix(gateway): preserve compaction state and server stream output * fix(gateway): close local stream generator on cancellation --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
de3f588fbf |
fix: windows async MCP tool execution, restore graph state after interruptions (#290)
Co-authored-by: z00827015 <zhoulun1@huawei.com> |
||
|
|
f356de36a6 | feat: memory retrieval (#281) | ||
|
|
76972449c7 |
feat(cli): multi-stage slash command completions with subcommand awareness (Phase 1 of #82) (#273)
* feat(cli): multi-stage slash command completions with subcommand awareness Phase 1 of #82 — subcommand and argument awareness in completions. - commands/base.py: add SubCommand dataclass and subcommands/category ClassVars to the Command ABC. Each SubCommand has name, description, and optional arguments. - commands/manager.py: add get_subcommands() and list_subcommands() methods to expose subcommand metadata for completion rendering. - commands/implementation/mcp.py: declare 6 subcommands (list, config, add, edit, remove, install). - commands/implementation/model_fallback.py: declare 6 subcommands (list, add, remove, clear, save, help). - commands/implementation/channel.py: declare 2 subcommands (status, stop). - cli/tui_interactive.py: rewrite on_text_area_changed slash-completion branch. When the user types a command name + trailing space and the command has subcommands, show subcommand completions instead of hiding the popup. Filter subcommands by typed prefix in multi-token input. - commands/implementation/general.py: /help now lists subcommands below each command that declares them. Tests: 8 new tests covering SubCommand creation, CommandManager subcommand lookup, and cross-command verification. 2293 passed baseline, no regressions. * fix: subcommand completion preserves prefix + prompt_toolkit + tests - _apply_selected_completion: preserve '/mcp ' prefix when completing subcommands via _comp_is_subcommand flag - SlashCommandCompleter (Rich CLI): add subcommand completion support - Fix trailing-space bug: rstrip prefix before top-level matching - test_tui_widgets.py: update stub on_input_changed to match multi-stage logic; add 5 new subcommand tests - test_command_manager.py: 8 tests for SubCommand + CommandManager 28 passed, 0 failed. * style: ruff format tui_interactive.py + test_tui_widgets.py * style: fix RUF012 ClassVar annotation on subcommands lists * fix: sync test stub, add len>=3 guard, remove exact-match hide - Sync test stub on_input_changed with real TUI code (remove exact-match hide for subcommands, add len(parts)>=3 guard) - Update test_input_changed_exact_subcommand_hides -> shows_confirmation - Add test_input_changed_three_parts_hides - Remove unused category ClassVar (din0s: what is this for) * refactor(commands): extract shared completion engine Per din0s feedback: one shared completion engine (commands/_completion_engine.py) that parses text + cursor once, returns structured CompletionCandidate objects with replace_start/replace_end ranges. - SlashCommandCompleter (Rich CLI): thin adapter, delegates to engine - on_text_area_changed (TUI): thin adapter, delegates to engine - _apply_selected_completion: uses candidate.replace_start/replace_end instead of _comp_is_subcommand flag - Tests: engine tested directly (10 new tests), stub methods updated 29 passed, 0 failed. * style: ruff format * fix: preserve text after cursor when applying completion CodeRabbit: replace_start only cuts from start to cursor, dropping any suffix after the cursor. Use replace_start + replace_end to correctly splice the replacement while preserving trailing text. * fix(cli): repair slash-command completion (TUI crash, subcommand bugs, sort) Apology + context: the previous push shipped a TUI-breaking change (the new shared engine assumed ``event.text_area.cursor_position`` existed, but ``ChatTextArea`` / ``Changed`` don't expose it). User caught the crash on ``/``; fixing that surfaced two more bugs in the engine that din0s had already flagged. This commit addresses all of them and drops a piece of dead stub code. ## Bug fixes 1. **TUI crash on ``/``** (``tui_interactive.py:2335``) ``event.cursor_position`` doesn't exist on the ``Changed`` event, and ``ChatTextArea`` (Textual ``TextArea`` subclass) doesn't expose ``cursor_position`` either. Pass ``len(event.text_area.text)`` instead — the user types at the end of the input in practice. 2. **Subcommand trailing-space duplication** (``_completion_engine.py``) Typing ``/mcp a `` + Tab produced ``/mcp aadd``. The engine included the trailing space in ``replace_end``; the TUI apply unconditionally appended ``" "``, producing double-space output. Fix: ``replace_end`` excludes the trailing space; the TUI apply checks ``current[replace_end:].startswith(" ")`` and skips the separator when the suffix already has one. 3. **Subcommand exact-match confirmation noise** (``_completion_engine.py``) Typing ``/mcp list`` + Tab re-inserted ``list`` and the popup kept showing the same subcommand. Add a guard mirroring the top-level exact-match rule: when the only subcommand match is the prefix itself (no trailing space), return ``empty``. 4. **Alphabetical sort dropped in CLI** (``cli/interactive.py``) The new completer iterated ``result.candidates`` in registration order. Re-add ``sorted(result.candidates, key=lambda c: c.text)``. Same sort added to the TUI for consistency. ## Cleanup - Drop the dead ``on_input_changed`` method from the ``_StubApp`` test stub (0 call sites) plus the unused ``_slash_commands`` / ``_subcommands`` locals that fed it. This addresses din0s's comment about the stub duplicating real TUI logic — the inlined copy is no longer needed since the real completer now routes through the shared engine. ## Tests - ``test_engine_exact_subcommand_shows_confirmation`` → renamed to ``test_engine_exact_subcommand_hides`` to match new behavior. - New: ``test_engine_subcommand_trailing_space_excludes_space_from_range`` and ``test_engine_subcommand_trailing_space_apply_does_not_double_space``. - All 97 tests in ``test_tui_widgets.py`` pass. - ``ruff check`` / ``ruff format`` clean. - Local TUI smoke: ``/`` (no crash, top-level popup), ``/mcp `` (subcommand popup), ``/mcp a `` + Tab → ``/mcp add ``. Refs the din0s review comments on PR #273. CLI path tests and the ``category`` ClassVar follow-up are deferred to a separate PR (the former is a test-suite addition; the latter is already absent from ``base.py`` on the current branch). * fix: address remaining review items (help duplication, CLI tests, stub sync, docstrings) - mcp.py: auto-generate help text from subcommands ClassVar (#1) - tests/test_cli_completion.py: add 9 CLI completer tests (#2c) - test_tui_widgets.py: sync _apply_selected_completion stub with real code (#4) - mcp.py + interactive.py: add docstrings to key functions (#8) * fix: hide completions on exact subcommand match regardless of trailing space Remove the ot has_trailing_space guard from the exact-subcommand check. Previously /mcp list (with trailing space) would still return candidates, causing Tab to oscillate between adding and removing the trailing whitespace. Now the engine hides whenever the subcommand is an exact match, same as the top-level rule. Added test_engine_exact_subcommand_with_trailing_space_hides to cover the scenario din0s flagged. * refactor: use StrEnum for CompletionResult.kind Replace plain str with CompletionKind(StrEnum) for type safety. Backward-compatible with existing string comparisons. * fix: normalize @file completion tuples to CompletionCandidate complete_file_mention() returns list[tuple[str, str]] but the TUI rendering/apply code expects objects with .text/.description. Wrap tuples in CompletionCandidate to prevent AttributeError crash. --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
05a5e5f8a0 |
fix: session lost after evoscientist restart (supersedes #278) (#279)
* fix: session lost after evoscientist restart * feat: Implement memory worker thread deletion on completion - Added synchronous and asynchronous functions to delete memory worker threads after they finish execution, ensuring no residual checkpoints are left in the database. - Enhanced `_watch_memory_worker_run_sync` and `_watch_memory_worker_run_async` to invoke deletion functions upon confirming worker completion. - Introduced tests to verify that worker threads are deleted correctly upon successful completion and that failures in deletion do not affect the overall worker status. - Updated session management to ensure that only relevant threads are restored from the database, preventing exposure of internal or unrelated workspace threads. - Implemented a purge function to clean up leftover worker checkpoints during server startup. * feat: Implement short thread ID display for CLI and session hints * fix: ensure proper accounting and deletion order for memory worker threads --------- Co-authored-by: z00827015 <zhoulun1@huawei.com> |
||
|
|
c02be519f6 |
feat(dangerous-mode): implement real-filesystem access with safety ch… (#276)
* feat(dangerous-mode): implement real-filesystem access with safety checks - Introduced a 'dangerous mode' allowing the agent to operate on the real filesystem. - Updated command validation to bypass path confinement while enforcing a blocklist for privileged commands. - Added warnings and guidelines for users when operating in dangerous mode. - Enhanced configuration to support dangerous mode and ensure it implies auto-approval. - Updated tests to verify the behavior of commands and configurations in dangerous mode. * feat(dangerous-mode): enhance logging and environment management for dangerous mode * feat(dangerous-mode): improve handling of dangerous mode with environment flags and enhance test isolation |
||
|
|
4b6a969df2 |
refactor(agent): make create_cli_agent(config=, chat_model=) pure (#267)
* refactor(agent): make create_cli_agent(config=, chat_model=) pure Re-applies the #183 purity refactor on top of the observation-memory lifecycle that landed in #259, integrating the two cleanly. create_cli_agent gains a pure path: when both `config` and `chat_model` are passed it builds the agent entirely from locals and writes none of the cached module globals (`_config`, `_chat_model`, `_chat_model_key`, `_EvoScientist_agent`). `/model` commits the switch via `set_active_config` / `set_chat_model_instance` only after a successful build, so a failed rebuild leaves the session on the original model (replaces the old snapshot/restore rollback). Supporting changes: - Extract `set_active_config` (write-half of `_ensure_config`), `_apply_env_from_config`, `_build_chat_model`, and `set_chat_model_instance`. - Thread `cfg` / `chat_model` through `_get_default_middleware`, `_build_base_kwargs`, `load_mcp_and_build_kwargs`, `_maybe_swap_async_subagents`, and `_inject_subagent_middleware` so the pure path never falls back to the global-writing `_ensure_config()` / `_ensure_chat_model()`. - Integrate with #259's memory middleware: subagent context-editing middleware binds the threaded `chat_model`, and the configured system prompt / memory controls read the threaded `cfg` (new threading vs the original #183, required because #259 made these paths read config). - Consolidate `cfg` resolution to one `cfg if cfg is not None else _ensure_config()` at the top of each kwargs builder, matching the pattern already used in the other config-aware helpers. * fix(agent): keep pure tool selector off global cache * fix(model): apply config switch in place to preserve reference integrity --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com> |
||
|
|
8bb1d6c0e3 |
refactor(stream): langgraph streaming v3 (#268)
* refactor(stream): langgraph streaming v3 * fix: address CR comments * chore(stream): add success field to state --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
ac052bbb4c |
feat: free-scrolling (#262)
* feat: free-scrolling * fix: anchor * fix: textual private vars |
||
|
|
92d95dee68 |
feat(memory): add observation memory lifecycle (#259)
* feat(memory): add observation memory lifecycle Add file-backed observation memory with deterministic markdown records, structured record_observation tooling, startup indexing, and profile/observation prompt guidance. Launch post-turn and post-subagent EvoMemory workers through LangGraph dev so completed runs can update profile memory, save durable observations, and write subagent execution summaries without blocking the active agent. Wire memory middleware into the main agent, subagents, async graphs, TUI status reporting, worker activity accounting, and observation-aware research prompts, with regression coverage for storage, lifecycle scheduling, graph registration, status display, and stream reset behavior. * fix(cli): sync background agent server on resume Resume flows now need to keep the LangGraph dev background server aligned with the active workspace even when async subagents are disabled. EvoMemory workers use that server too, so gating resume-time sync on enable_async_subagents could leave workers pinned to the launch workspace after resuming a thread from another workspace. Run workspace sync unconditionally for Rich CLI and Textual resume paths, while preserving WorkspaceMismatchError handling so failed sync aborts the resume before mutating the active thread or workspace. Propagate aborted resume callbacks through the command UI so channel-issued /resume commands do not send false success or history output. Channel slash dispatch now treats CommandManager-caught command errors as command errors and skips completion hooks for those failed commands. Add regression coverage for disabled async subagents, callback aborts, and channel command error reporting. * fix(cli): prepare serve resume workspace before adopting Load the resumed workspace agent and sync the background server as a single pre-adoption step. Restore the previous active workspace if preparation fails so serve mode keeps using the old session consistently. * fix(memory): untrack abandoned worker status watches Stop treating watcher shutdown as confirmed worker completion. Terminal worker statuses still count memory deltas, while poll failures or watcher setup failures now remove the active run without crediting partial outputs. * fix(cli): report channel command failures accurately Treat command_error as a None sentinel so empty error strings still fail, and let TUI resumes continue only on non-mismatch background-server sync failures while reporting degraded mode. * fix(stream): clear memory counters for resume streams Reset completed-memory counters for every new agent stream, including Command-based HITL and resume streams, so saved-memory indicators do not leak across turns. * docs(tools): make observation recording guidance conditional Clarify that agents should call record_observation only when the observation tool is available, preserving the existing durability and usefulness criteria. * feat(config): add controls for profile and observation memory Add config flags for profile memory, observation memory, observation writer placement, and background memory workers. Wire the controls through main agents, subagents, EvoMemory middleware, and memory lifecycle workers so observation writes can be assigned to the live agent, subagent worker, both, or neither. Keep turn memory workers profile-only and make prompts reflect the available observation read/write paths. Skip langgraph dev startup when neither async subagents nor memory workers need the background server. Add coverage for config parsing, prompt gating, middleware wiring, and worker tool availability. * test(cli): include memory defaults in serve config stubs * fix(memory): offload async worker launch blocking calls Run the langgraph-dev health check and memory-output snapshot in worker threads from the async EvoMemory launcher so it does not block the event loop. * chore(memory): harden turn worker subagent guardrail * chore(memory): refresh profile context per request * fix(memory): offload async profile file reads * fix(memory): offload async worker completion accounting |
||
|
|
9285c6dad8 |
Migrate memory middleware to profile files (#253)
* feat(memory): migrate to profile memory files * chore(stream): read profile headings from templates * fix(display): keep assistant responses if response_text has started * fix(memory): do not treat failed bootstraps as profile creation * chore(memory): unlink blank legacy memory * fix(memory): resolve project_id once * fix(memory): preserve unreadable profile files * chore(tui): render streamed narration inline with tool timeline Update the TUI streaming timeline so assistant text emitted before or between tool calls is rendered inline where it occurs, rather than being kept as a single answer bubble above or below the tools. If the model begins an assistant response and then emits another tool call, the provisional response is converted into inline narration before that tool. The final assistant message then renders only the remaining response suffix, avoiding duplicate text in the completed transcript. Stop/cancel handling now preserves any active inline narration, appends the visible stopped marker only to the remaining displayed segment, and still returns the full normalized stopped response for channel callers. Completed tools continue to collapse while long runs are active, but expand again when the turn reaches a final state so the completed transcript shows the full tool timeline. * fix(stream): preserve narration around tool timelines Keep assistant narration attached to the tool call that follows it instead of folding all streamed text into the final answer block. Track narrated response segments in stream state, render them before their corresponding regular or task tool entries, and keep final answers limited to the response suffix that has not already been shown inline. Preserve narration across normal completion, stop/error final frames, sub-agent task calls, and collapsed live tool summaries. Add regression coverage for pending tools, completed tools, sub-agent task delegations, collapsed completed/running tool summaries, and final stop frames. * fix(tui): finalize inline narration transitions * test(memory): use canonical project id helper |
||
|
|
fbd1d709ca |
feat: add WebUI mode support with related configuration and onboarding (#252)
* feat: add WebUI mode support with related configuration and onboarding steps * feat: enhance WebUI port configuration to prevent conflicts with backend port * feat: add support for fresh interactive session detection in WebUI |
||
|
|
a13904185d |
Feat/sandbox execute timeout (#243)
* feat: implement configurable sandbox execute timeout and enhance recovery instructions * feat: add background process management tools and middleware for sandbox execution * feat: enhance background process management with completion notifications and deduplication * feat: enhance sandbox execution timeout validation and update related messages * feat: enhance background process management with thread-specific completion notifications and HITL approval handling * test: assert completion notification waits for process finish timestamp |
||
|
|
2364e6b130 |
fix(cli): forward async-notifier replies back to originating channel (#244)
* fix(cli): forward async-notifier replies back to originating channel When PR #214's auto-notifier fires a synthetic agent turn after a channel-originated conversation, the synthesized response only rendered to the local CLI/TUI — the channel user (iMessage etc.) saw nothing and had to manually re-prompt to find out what happened. Adds a per-thread channel-origin registry in cli/channel.py and wires the three notifier paths (Rich CLI / TUI / serve) to publish the final response back via bus.publish_outbound when the originating thread was started by a channel turn. Publish is fire-and-forget (scheduled on the bus loop + done-callback for failure logging) so the notifier turn doesn't block on the asyncio / textual event loop. The registry is cleared on /new and /resume rotation so stale entries don't accumulate. * fix(cli): address review feedback on channel-origin forwarding Follow-up to the review on #244 (din0s, X-iZhang): - Guard the /resume origin cleanup on a real thread change in Rich CLI and TUI (serve mode already did via thread_changed). Resuming the already-active thread no longer wipes its still-live origin, which would otherwise silently drop a later async-notifier forward — the exact gap this PR closes. - Re-bind the now-current thread to its channel after a channel-issued /new or /resume slash command (which rotates the thread inside the dispatch), so notifier turns on the rotated thread still forward. - Guard the publish done-callback against a cancelled future, whose .exception() raises CancelledError (rather than returning it) on bus-loop teardown, so the intended warning still logs. - Mirror the normal reply path's manager.record_message(channel, "sent") for forwarded notifications so per-channel stats stay accurate. - Print the closing "[channel: Replied to ...]" line in all three notifier paths (Rich CLI / TUI / serve) when a forward actually happened, so the forwarded block reads as terminated on screen. Adds test_publish_records_sent_metric. ruff clean; notification-origin suite (10) + related channel/CLI/serve suites (728) pass. * fix(cli): store sender information separately from chat_id in channel origin --------- Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com> |
||
|
|
721a03c25b | feat: update model version from claude-sonnet-4-5 to claude-sonnet-4-6 and related adjustments | ||
|
|
f75bfcda51 |
Add onboarding wizard with style and validation components (#241)
* Add onboarding wizard with style and validation components - Introduced `style.py` for shared visual elements used in the onboarding wizard. - Created `validators.py` for input validation, including integer and choice validators, and API key validation functions for various providers. - Implemented `wizard.py` as the entry point for the onboarding process, managing user prompts and configuration steps. - Added progress rendering and autosave functionality to enhance user experience during the onboarding process. * feat(onboarding): enhance validation and configuration for onboarding wizard - Added validation for UI backends, workspace modes, and providers in the onboarding command. - Updated channel definitions to include secret field handling for sensitive tokens. - Improved user prompts for required fields, ensuring sensitive data is masked. - Introduced constants for valid providers, UI backends, and workspace modes to maintain consistency. - Implemented tests to ensure alignment between constants and interactive choices in onboarding steps. * feat(onboarding): improve WeChat account ID prompt and validation for newly enabled channels * feat(onboarding): enhance WeChat backend credential prompts and validation * feat(onboarding): refine WeChat backend credential prompts for wecom and wechatmp * Refactor onboarding package for improved structure and clarity - Simplified the onboarding package by removing unnecessary re-exports and consolidating public API to only include `run_onboard`. - Updated `install_back_keys` to `install_navigation_keys` for clarity and consistency in the prompter module. - Enhanced the `NonInteractivePrompter` class to support strict mode, allowing for better handling of non-interactive prompts. - Adjusted the onboarding steps to utilize the new navigation keys installation method. - Improved the `run_onboard` function to handle section implications based on user flags, enhancing the onboarding experience. - Updated tests to reflect changes in imports and ensure compatibility with the new structure. * feat(onboarding): enhance validation logic for non-interactive prompts * refactor(onboarding): streamline onboarding module structure and enhance validation error handling * refactor(onboarding): enhance config revert logic to preserve original file state * refactor(onboarding): enhance tavily key validation and error handling in onboarding process |
||
|
|
7959495a13 |
feat(deploy): add EvoSci deploy subcommand (#228)
* feat(deploy): implement standalone LangGraph server and CLI command for deployment * feat(deploy): enhance port validation and environment variable management for deployment * Refactor langgraph dev deployment and introduce workspace sidecar protocol - Updated the deployment mode handling in `server.py` to use a single environment variable `EVOSCIENTIST_DEPLOY_MODE` with values `full` and `stripped`. - Enhanced the `manager.py` to implement a workspace fingerprint sidecar, allowing cross-process reuse of langgraph dev instances while ensuring workspace consistency. - Introduced functions to write and read the workspace sidecar, with error handling for missing or corrupt data. - Added tests for the workspace sidecar functionality, including validation of the JSON schema and ensuring proper error handling for workspace mismatches. - Updated existing tests to reflect changes in deployment mode handling and added new tests for signal handling during shutdown. - Ensured that cleanup routines remove the workspace sidecar alongside the PID file during shutdown. * fix(langgraph): improve workspace sidecar checks for process ownership and stale handles |
||
|
|
7f1aa3b0f6 |
fix(cli): handle spaces in @file mentions (#234)
* fix(cli): handle spaces in @file mentions The @file parser truncated at the first space, so dragging or pasting a filename like `@PREPING_ Building Agent.pdf` only matched `@PREPING_` and warned "file not found". Now supports `@"..."` / `@'...'` quoted form for explicit paths, plus a greedy expansion fallback that walks across whitespace until an existing file resolves (bounded by newlines, the next `@`, and a 20-token cap). Autocomplete also returns quoted mentions for any candidate containing a space. * style: apply ruff format to file_mentions |
||
|
|
a4c9c779c9 |
feat: status and elapsed time indicator (#218)
* feat: status and elapsed time indicator * test: add tests for tui-status * fix: move to enum+switch, change phase calculation * feat: remove 'done' phase |
||
|
|
8fe774b056 |
Feat/qq interactive buttons (#220)
* feat(qq): add inline keyboard buttons for C2C HITL approval
QQ Bot supports inline buttons via `markdown + keyboard` payloads. Clicks
arrive as `interaction_create` events through the existing botpy
WebSocket gateway — no extra subscription needed beyond enabling the
`interaction` intent. Group-scope clicks are out of scope here (DM only).
Send path
- `_build_qq_keyboard(buttons)` mirrors the Feishu helper, mapping the
generic `{text, value, type}` shape to QQ's `{render_data, action}`
with action.type=1 (callback). One button per row for mobile clarity.
- `_send_chunk` extracts `metadata["buttons"]` and threads a `keyboard`
payload into `_post_markdown_message` for C2C only.
- Markdown→plain fallback can't carry a keyboard, so when buttons were
attached the fallback content gets a textual `Reply: 1=Approve, …`
hint built from the button list. `_parse_approval_reply` accepts
the same values typed manually, so the user is never stuck.
Receive path
- `on_interaction_create` is registered on the bot class.
- `_on_interaction` extracts `data.resolved.button_data`, builds an
InboundMessage, runs it through inbound middleware (Dedup suppresses
retry callbacks), and publishes directly to the bus — bypassing the
per-sender debounce buffer so the click value isn't merged with any
text typed in the same window.
- Always ACKs via `api.on_interaction_result(id, 0)` in `finally` so
QQ doesn't show the button as "expired", even if middleware drops
the click or something throws downstream.
`QQ.inline_buttons=True`; `_approval_prompt_metadata` now auto-attaches
the Approve/Reject/Approve-all button row for QQ HITL prompts.
* fix(qq): button-value coercion, ACK timing, HITL consumer wiring
Fixes 6 bugs found in the inline-keyboard commit and consolidates the
button helpers so the keyboard builder, plain-text fallback hint, and
interaction handler share one coercion path.
- Plain-text fallback no longer crashes on non-string `value` (e.g.
`{"text": "OK", "value": 42}`). Extracted `_normalize_button` is now
the single place that resolves `(label, value)` and coerces non-strings.
- `metadata["button_value"]` is the coerced string instead of the raw
payload, matching `content` and downstream string comparisons.
- `_on_interaction` ACKs first, before publishing to the bus, so the
QQ button UI never shows "expired" if middleware is slow.
- Wire `_approval_prompt_metadata` + `_format_approval_prompt(with_buttons=)`
into `InboundConsumer._stream_with_hitl` and `cli.channel.channel_hitl_prompt`
so the QQ `inline_buttons=True` capability is actually used end-to-end
(HITL prompts auto-attach Approve/Reject/Approve-all buttons when the
channel advertises the capability).
- Trim contradictory `_QQ_DEFAULT_PERMISSION` comment.
- Fix `test_group_interaction_ignored` docstring (ACK runs first now,
not in `finally` after a `return`).
Tests: `_normalize_button` covered indirectly via existing keyboard tests;
new regressions for non-string fallback hint, ACK-on-handler-throw, and
string-coerced `button_value` metadata.
* refactor(qq): slim button helpers and explicit has_buttons flag
Inline single-use _button_hint and the _QQ_BUTTON_STYLE/_QQ_DEFAULT_PERMISSION
constants in qq/channel.py; tighten _on_interaction (drop unreachable
"[button click]" sentinel and unused triggering_message_id metadata; collapse
"if resolved else" ternaries via `or ""`).
Replace the metadata round-trip ("buttons" in metadata) used to detect button
support in consumer.py and cli/channel.py with an explicit has_buttons bool
threaded through both the prompt formatter and metadata builder.
Apply ruff format to the previously unformatted blocks introduced earlier on
this branch so CI lint passes.
* feat(qq): send post-decision confirmation after HITL approval
Send a visible confirmation message ("✅ 已批准" / "❌ 已拒绝") right after
the user resolves a HITL approval — QQ Bot has no message-recall or edit API
for C2C, so a follow-up message is the only way to give the click/reply
strong feedback.
Bus consumer (consumer.py): only sends the confirmation when the user
actually responded (event was set), to avoid pretending the user approved
when the request really timed out and auto-approved.
CLI HITL prompt (cli/channel.py): mirrors the same set of confirmation
strings. Timeout and unrecognized-reply paths keep their existing English
text since their semantics differ (auto-reject vs auto-approve, plus a
hint about the unparsed input).
|
||
|
|
4e04ac5b72 |
fix: Improve watcher logic to prevent false-positive notifications on… (#216)
* fix: Improve watcher logic to prevent false-positive notifications on clean stream exits * fix: Update watcher logic to drop notifications on persistent runs.get failures * fix: Refactor test for watcher persistent failure notification handling * fix: Enhance watcher test to validate all notification queues are empty after reconnect budget exhaustion |
||
|
|
80f1f4fa0f |
feat: Implement async sub-agent auto-notification system (#214)
* feat: Implement async sub-agent auto-notification system - Added async notifier functionality to handle notifications for sub-agents reaching terminal states. - Introduced `AsyncTaskNotification` dataclass for structured notification data. - Implemented `watch_run_and_notify` to monitor agent runs and enqueue notifications. - Created `spawn_watcher` to manage watcher tasks and ensure proper cancellation of previous watchers. - Developed `consume_notifications` to process notifications, deduplicate them, and format messages for LLM. - Added tests for notification handling, including draining, deduplication, and formatting. - Patched deepagents to integrate the new watcher functionality into start and update tools. * Enhance async notifier with per-thread notification routing and error handling - Introduced `origin_cli_thread_id` to `AsyncTaskNotification` for routing notifications back to the originating CLI session. - Implemented per-thread notification queues to handle notifications based on the originating thread. - Updated `has_pending_notifications` and `drain_notifications` to respect thread-specific queues. - Enhanced `watch_run_and_notify` to detect in-band error events from the SSE stream and handle clean exits. - Modified tests to verify the new notification routing behavior and ensure proper handling of notifications across threads. - Added a fixture to restore the async watcher patch state in tests to prevent state leakage. - Updated deepagents patching to capture the main agent's CLI thread ID for notification routing. * feat: Enhance async notifier with thread-specific watcher management and notification filtering * test: Enhance notification draining logic for cleaner test setup * refactor: Remove summary field from AsyncTaskNotification and update related tests * feat: Enhance async notification handling with target thread ID support * Refactor async notifier and middleware for improved task management - Removed the no-op shutdown watcher loop from async_notifier.py as it is no longer needed. - Updated watch_run_and_notify to clarify notification handling and race conditions. - Cleaned up shutdown handling in commands.py, interactive.py, and tui_interactive.py by removing obsolete shutdown watcher calls. - Deleted the deepagents async watcher patch from patches.py, transitioning to a new middleware approach. - Introduced AsyncWatcherMiddleware to handle async task notifications directly during tool calls. - Updated tests to validate the new middleware functionality and ensure proper watcher spawning and cancellation. - Enhanced test coverage for async watcher middleware, including edge cases and error handling. * feat(tests): add fixture to reset notifier state before each test |
||
|
|
9e51ec6fdd |
feat(cmd): add /model-fallback command (#196)
* feat(cmd): add /model-fallback command * fix: apply feedback * fix: lock usage with _fallback_chain * fix: apply feedback * fix: apply feedback * feat: add tests * fix: tests * Update EvoScientist/middleware/model_fallback.py Co-authored-by: dinos <dinospk1999@gmail.com> --------- Co-authored-by: dinos <dinospk1999@gmail.com> |
||
|
|
22a65b640d |
refactor(channels): remove dead MessageBus dispatcher (#205)
* refactor(channels): remove dead MessageBus dispatcher Outbound routing has two implementations: ``MessageBus.dispatch_outbound`` (subscriber-based) and ``ChannelManager._dispatch_outbound`` (registry lookup). Only the latter is ever started in production — the former is reachable solely from tests, yet both consume from the same ``bus.outbound`` queue. If anyone followed the bus's own API surface they would silently steal messages from the real dispatcher. Drop the unused machinery to leave a single, obvious outbound path: - ``MessageBus.subscribe_outbound`` / ``dispatch_outbound`` / ``stop`` - ``_running`` flag and ``_outbound_subscribers`` map - ``OutboundCallback`` type alias - The lone ``bus.stop()`` call in ``cli/channel.py`` (was no-op) - Four tests covering the removed code paths * test(channels): drop empty MessageBus stubs after dispatcher removal --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
f41584e10b |
Refactor sub-agent architecture and introduce async support (#200)
* Refactor sub-agent architecture and introduce async support - Removed the legacy subagent.yaml file and replaced it with individual YAML files for each sub-agent in the subagents directory. - Updated the load_subagents function to support both directory and single file layouts for loading sub-agent configurations. - Added new langgraph_dev module for managing async sub-agent lifecycle and deployment. - Created graphs for async sub-agents (writing-agent, data-analysis-agent) and updated langgraph.json for deployment. - Introduced new sub-agent definitions for planner, research, debug, code, and writing agents with appropriate system prompts and configurations. - Enhanced package data inclusion in pyproject.toml to accommodate new sub-agent YAML files. * Refactor code for improved readability by consolidating conditional statements and formatting * feat: enhance async sub-agent support with workspace synchronization and user feedback - Added console status messages during async sub-agent server startup and workspace synchronization to improve user experience. - Implemented a new WorkspaceSyncWidget for live feedback during workspace sync operations. - Updated onboarding to reject occupied ports and ensure proper workspace handling for async sub-agents. - Introduced locking mechanisms to manage concurrent access to langgraph dev processes and workspace states. * feat: add async sub-agent configuration and server management functions * feat: improve port occupation handling and log file management in start_langgraph_dev * feat: enhance async sub-agent handling and introduce comprehensive tests - Updated `_maybe_swap_async_subagents` to improve async sub-agent management, ensuring internal flags are stripped before handoff. - Enhanced port management in `onboard.py` to allow reuse of occupied ports if already running by the same service. - Introduced file locking in `manager.py` to prevent race conditions during concurrent CLI invocations. - Added new tests for async sub-agent swapping and langgraph manager functionalities to ensure reliability and correctness. - Updated dependencies in `pyproject.toml` to include `psutil` and `filelock`. * fix(docs): clarify sub-agent configuration in README * test(manager): isolate _PID_DIR + tighten reuse-path assertion Addresses CodeRabbit review on tests/test_langgraph_manager.py: - Patch _PID_DIR to tmp_path so the FileLock setup in ensure_langgraph_dev doesn't mkdir the user's real ~/.config/evoscientist/ dir as a test side-effect. - Tighten "result is None or hasattr(result, 'poll')" to a strict "result is None" — the reuse path returns None unconditionally, so the OR clause was hiding potential regressions. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(manager): clean up stale PID file when unrelated process reuses PID * feat(tests): add validation tests for async flag in load_subagents * fix(load_subagents): restrict to .yaml files and clarify configuration handling * fix(load_subagents): improve error handling for non-dict specifications in YAML * feat(onboard): add "LangGraph Port" step to onboarding process * feat(langgraph): add concurrency configuration for langgraph dev workers * feat(async-subagents): enhance MCP tool routing for async sub-agents * fix(manager): update exception handling for connection errors and prevent zombie processes --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
ab3787f86a |
fix(markdown): ensure proper spacing for ATX headings in Markdown ren… (#201)
* fix(markdown): ensure proper spacing for ATX headings in Markdown rendering * fix(markdown): improve docstrings and tests for heading spacing functionality |
||
|
|
5c829942d7 |
refactor(cli): replace channel module globals with ChannelRuntime (#197)
* refactor(cli): replace channel module globals with ChannelRuntime Removes _cli_agent / _cli_thread_id from EvoScientist/cli/channel.py and threads a ChannelRuntime via CommandContext.channel_runtime so /model and /channel rebind without poking module-level state. * fix(cli): address coderabbit review - _auto_start_channel: bind ChannelRuntime only after _start_channels_bus_mode succeeds, so a startup failure no longer leaves a stale binding pointing at channels that never started. - _sync_tui_command_completion (TUI) and the Rich CLI command-completion paths: rebind the runtime on thread rotation, not just agent swap, so /new and /resume keep ChannelRuntime in sync with the running thread (matches the serve-mode hook contract). - test_hook_syncs_channel_runtime: pin ctx.thread_id explicitly so a bare MagicMock attribute can't silently mutate runtime.thread_id. - New regression test covering the rebind-on-thread-rotation contract. |
||
|
|
50719ef256 |
Implement PruningCheckpointer for efficient checkpoint management and… (#194)
* Implement PruningCheckpointer for efficient checkpoint management and add comprehensive tests - Introduced `PruningCheckpointer` to manage checkpoint pruning after each `aput()`, ensuring only the latest checkpoints are retained based on a configurable limit. - Added migration sweep functionality to clean up legacy checkpoints and prevent database bloat. - Enhanced `get_checkpointer()` to utilize the new `PruningCheckpointer` and trigger migration sweeps when necessary. - Developed a suite of integration tests for `PruningCheckpointer`, covering various scenarios including pruning behavior, concurrent writes, and retention policies. - Implemented tests for migration sweep functionality, ensuring proper partitioning and user version management. - Added diagnostic helper `db_stats` to provide insights into the database state, including thread and checkpoint counts. * feat(sessions): enhance pruning logic to handle legacy DBs without writes table * fix(tests): prevent atexit hook leakage in TestMigrationSweep * feat(tests): enhance TestPruningCheckpointer to validate put+prune serialization * feat(tests): refactor mock path implementation for get_db_path in test cases |
||
|
|
da74c325d6 |
fix(channel): scope stop and restore resume history (#186)
* fix(channel): scope stop and restore resume history * refactor(channel): simplify stop and resume patch * Delete PR_MESSAGE.md * fix(channel): address review feedback * fix(channel): address remaining review bugs * fix(channel): clean up stopped request handling * fix(channel): preserve resolved replies and sync tui commands --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
27fee3256e | fix(clipboard): remove unnecessary check for selection end in copy_selection_to_clipboard | ||
|
|
558360b558 |
feat: Enhance ModelPickerWidget for Ollama integration (#187)
* feat: Enhance ModelPickerWidget for Ollama integration - Implemented a sentinel row for "Custom Ollama model..." in ModelPickerWidget, allowing users to input arbitrary model names. - Updated action handling in ModelPickerWidget to manage transitions between list and input modes. - Added async model discovery for Ollama models, integrating with the /model command to fetch locally installed models. - Created tests for Ollama model discovery and ModelPickerWidget behavior, ensuring proper functionality and user experience. - Refactored validate_ollama_connection and discover_ollama_models for improved error handling and response management. * fix: Simplify code by removing unnecessary line breaks in ModelPickerWidget and test cases * fix: Restore globals on set_chat_model failure to prevent half-switched session * fix: Improve error handling in ModelCommand by restoring globals on failure |
||
|
|
155c4eaa40 |
fix: text copy on remote sessions/legacy terminal emulators (#185)
* fix: text copy on remote sessions/legacy terminal emulators * fix: display warning only once |
||
|
|
c134a16e19 |
Fix/channel slash rich cli (#184)
* feat(cli): implement slash command dispatch for channel messages * fix(cli): make EvoSci serve exit on Ctrl+C and hot-swap /model (#181) * fix(cli): streamline debug logging and formatting in channel command handling * fix(cli): ensure proper handling of asyncio event loop in slash command processing * fix(cli): add error handling for unexpected exceptions in slash command dispatch * fix(cli): improve error messaging for slash command dispatch failures * fix(tests): enhance test setup by restoring channel globals and simplifying assertions * fix(cli): enhance slash command handling across UI surfaces and improve resume command warnings |
||
|
|
92df3e1844 |
Refactor/cli command manager (#178)
* feat(cli): migrate command handling to CommandManager and enhance UI interactions * feat(cli): add /clear and /help commands to enhance user experience * feat(cli): implement /new and /resume commands with interactive session management * Refactor MCP and Skills Command Handling - Moved the interactive picker style to a centralized widget for consistency across MCP and Skills commands. - Updated the MCP command to remove the old command dispatch logic, delegating to the new InstallMCPCommand. - Enhanced the Skills command to utilize a new interactive picker for skill selection, improving user experience. - Implemented cancellation handling in the picker to differentiate between user cancellations and empty selections. - Added comprehensive tests for the new command structures and picker functionalities to ensure reliability. * refactor(cli): streamline CommandManager dispatch and remove deprecated command set * refactor(cli): enhance error handling and state management in ChannelCommand and RichCLICommandUI * refactor(cli): update lifecycle callback terminology and improve async prompt handling in RichCLICommandUI * refactor(cli): enhance SlashCommandCompleter to dynamically fetch workspace directory for autocompletion * refactor(cli): unify quit handling in RichCLICommandUI with shared _stop helper * refactor(cli): remove hardcoded slash commands and utilize command manager for dynamic completion * refactor(cli): update MCP and skills command files for improved clarity and organization * refactor(mcp_ui): remove unnecessary newline in _show_mcp_config function |
||
|
|
65828e9666 |
perf(cli): cut startup latency and defer MCP loading to the background (#171)
* perf(cli): cut startup time of `evosci --help` from ~2.2s to ~0.3s
Module-level imports were eagerly pulling in langchain.chat_models (with
the whole anthropic/openai/google stack), langgraph, textual, and
prompt_toolkit on every invocation — even for `--help` or `config list`.
Defer those with PEP 562 `__getattr__`, using `lazy_loader.attach` (SPEC-1,
the scientific-python standard) where it's a clean attach pattern:
- `EvoScientist/llm/__init__.py`: attach `models` lazily so importing
`context_window` from this package no longer drags in langchain.
- `EvoScientist/stream/__init__.py`: attach display/events lazily; split
the shared Rich `Console` singleton into a new lightweight
`stream/console.py` so callers that only need `console` skip the
`stream.events` → `langchain_core.messages` chain.
- `EvoScientist/cli/__init__.py`: hand-rolled `__getattr__` (reaches into
`..stream.state`, which `lazy_loader` doesn't cover) so `commands` and
`app` are the only eager loads.
- `EvoScientist/cli/commands.py`: move `cmd_interactive`/`cmd_run` to
in-function imports so prompt_toolkit + textual only load when the
interactive path actually runs.
- `EvoScientist/cli/_constants.py`: read `AGENT_NAME` on demand so
`build_metadata` doesn't eagerly import `sessions` (langgraph/aiosqlite).
Adds `lazy-loader>=0.5` as a dependency.
* feat(cli): defer MCP tool loading with live per-server progress
The CLI was blocking ~5 s on MCP tool enumeration before the first
prompt appeared. Move the agent construction off the event loop and
surface per-server progress so the user can interact immediately and see
what's happening.
MCP client:
- Add an `on_progress` callback to `load_mcp_tools` / `aload_mcp_tools`
/ `_load_tools` emitting `start` / `success` / `error` events per
server.
- Fan connection attempts out with `asyncio.gather` so latency no longer
scales linearly with server count; cap simultaneous attempts at
`_MAX_CONCURRENT_CONNECTIONS` (8) via a semaphore so a big stdio fleet
doesn't spawn every subprocess at once.
Agent wiring:
- Plumb `on_mcp_progress` through `create_cli_agent` / `_load_agent` /
`load_mcp_and_build_kwargs` so CLI and TUI can plug in collectors.
CLI (`cmd_interactive`):
- Run `_load_agent` in a background thread via `asyncio.to_thread`; the
prompt and banner render immediately.
- `_await_agent_ready()` awaits the task before each agent-using site
(first turn, channel messages, `/channel`, `/compact`). Raises if
called without a prior `_start_agent_load` instead of silently
reloading without the SQLite checkpointer.
- Pre-prime the progress dict from `load_mcp_config()` so the
bottom-toolbar's `N/M` denominator is stable from the first render.
- Wrap `session.prompt_async` in `patch_stdout(raw=True)` so
`console.print` from the worker-thread progress callback lands cleanly
above the prompt as inline chat messages instead of stomping the
prompt cursor.
TUI (`EvoTextualInteractiveApp`):
- Same background load + `_await_agent_ready()` gates on every
`self._agent` read.
- New `MCPLoaderWidget` mounted at the top of `#input-shell` shows a
header with `N/M` and one live row per server (spinner → ✓ / ✗ with
tool count or error detail). On completion:
- all-clean loads auto-dismiss ~2.5 s later;
- cache hits (no events ever fired) dismiss immediately rather than
flashing a misleading "0/N loaded";
- failures keep the widget mounted so the user can read the errors.
- `dismissed` property lets the app clear its ref so late events from
slow servers become no-ops. The error branch of `_on_agent_loaded`
also settles the widget so a load failure can't leave the spinner
animating forever.
- Chat input is `disabled` while MCP resolves — no placeholder hack, no
"waiting…" system message.
Shared:
- Hoist braille spinner frames to `status_bar.SPINNER_FRAMES` and import
them in the TUI widget so CLI and TUI animate in sync.
Tests:
- Extend `test_agent_mcp_cache` fakes to accept the new `on_progress`
kwarg.
- New `TestLoadToolsProgressCallback` in `test_mcp_client` exercises the
event sequence for success/failure/mixed fleets, verifies a buggy
callback doesn't break the load, and asserts the semaphore caps
in-flight connections.
* style: ruff
* chore: update uv.lock
* chore: uv.lock
* fix: coderabbit issues
* style: fmt
* fix: move _await_agent_ready inside try block
* fix(tui): auto-dismiss MCP loader widget on failure
The widget was designed to stay mounted on failure so the user could
read error detail, but since it's pinned above the input it never went
away in practice — just permanent banner clutter.
Auto-dismiss on failure too, with a longer grace (12s vs 2.5s) so the
error summary stays readable.
* fix: address second coderabbit pass
- Channel handlers (CLI + TUI): catch agent-load failures so the
channel request doesn't hang; CLI moves `_await_agent_ready()`
inside the existing try/except, TUI catches explicitly and calls
`_set_channel_response` with the error.
- Stale background loads: `prev.cancel()` only stops the asyncio
wrapper, not the thread running `_load_agent`. Added a generation
token (`agent_load_id` / `self._agent_load_id`) and gated both
progress and completion callbacks on it so a superseded load can't
clobber the current session's state or UI.
- TUI prompt lifecycle: added `_agent_load_pending()` and gated the
`_process_channel_message` / `_handle_command` finally blocks on it
so `/new` or `/resume` invoked from a command keeps the prompt
disabled until the fresh load settles.
- TUI readiness failures: `_run_turn` and `_handle_command` now
catch exceptions from `_await_agent_ready()` and surface a
"Agent failed to load: …" system message instead of letting the
exception escape into Textual's traceback panel.
* refactor(cli): share background agent loader between CLI and TUI
The CLI and TUI were carrying near-identical copies of the same
background-load state machine: the `agent_task`, the `agent_load_id`
generation token, the gated progress/completion callbacks, and the
per-server progress dict. Every CodeRabbit finding on that lifecycle
had to be fixed in both files.
Extract it into `cli/_agent_loader.py`:
- `MCPProgressTracker` — owns the `server -> (state, detail)` dict;
exposes `prime`, `record`, `snapshot`, `totals`.
- `BackgroundAgentLoader` — owns `agent`, the in-flight task, and the
generation token. Exposes `start(**loader_kwargs)`, `await_ready()`,
`is_pending`. Internally gates all progress/completion callbacks by
generation so a superseded load can't clobber the current session.
UI-specific rendering plugs in via `on_progress` / `on_success` /
`on_failure` callbacks.
Both surfaces now just wire their UI hooks; the loader file holds no
Rich / prompt_toolkit / Textual dependencies. Net -345 lines from
`interactive.py` + `tui_interactive.py`; +20 unit tests pinning the
lifecycle (generation filtering, cache-hit short-circuit, failure
reset, progress ordering).
* refactor(cli): make _on_done the sole authority for agent state transitions
await_ready no longer sets self.agent — it just awaits the task and
reads what _on_done already wrote. Eliminates the dual-write overlap
(asyncio guarantees done-callbacks fire in registration order).
* fix(tui): let users type during MCP load, only block on send
Remove prompt-disabling during background agent load — the TUI now
matches the CLI approach where the input stays enabled and only gates
on await_ready() at submit time. The MCPLoaderWidget still provides
visual feedback that loading is in progress.
* fix(loader): preserve real load error on await_ready; dedup failure message
CodeRabbit flagged two issues with the new loader:
1. After a failed load, `_on_done` nulled `self._task`, so the next
`await_ready()` hit the "before start()" branch and the CLI wrapper
remapped it to a misleading "checkpointer not available" message —
losing the real exception (bad MCP config, network, etc.).
Keep `_task` set on failure so `await_ready` re-raises the real
exception. Added `needs_restart` so TUI's auto-retry check stays a
one-liner and doesn't need to reach into task internals.
2. TUI reported each load failure twice: once from
`_on_agent_load_failure` (the done-callback) and once from each
caller of `_await_agent_ready` (`_run_turn`,
`_process_channel_message`, `_handle_command`) catching the re-raise.
`_on_agent_load_failure` is now the sole local reporter; callers
just handle control flow (return cleanly, set channel response to
unblock remote).
* fix(cli): wire /model handler through the agent loader
The /model command from main (merged via f1f0d7c) still reached for
`state["agent"]` (CLI) and `self._agent` (TUI) — both removed by the
background-loader refactor. CLI raised KeyError on first invocation;
TUI raised AttributeError. Writes to the old fields also had no effect
because every other code path now reads from `agent_loader.agent`, so
the model switch would have silently failed.
Route everything through the loader: `await _await_agent_ready()` up
front so /model doesn't race with the initial background load, build
the `CommandContext` with the current agent, and sync `ctx.agent` back
into `agent_loader.agent` (plus channel globals) when the command
replaces it.
* fix(cli): isolate progress callback, capture awaited agent, gate by requires_agent
Three CodeRabbit findings on the loader + command dispatch path:
- Wrap ``_on_progress`` in try/except inside the loader's gated wrapper
so a buggy UI adapter can't bubble into ``loader_fn`` and fail the
whole background load. The MCP client already protects this, but
defence-in-depth keeps the loader self-contained.
- In CLI channel + main-loop streaming, capture the agent returned by
``_await_agent_ready()`` and pass that into ``run_streaming`` rather
than reading ``agent_loader.agent`` after a subsequent ``await``.
A concurrent ``/new``/``/resume``/``/model`` could have swapped it.
- Add ``requires_agent: ClassVar[bool] = False`` to ``Command`` and
mark ``/compact``, ``/model``, ``/channel`` as ``True``. TUI dispatch
sites (channel and keyboard) now check ``cmd_manager.resolve(...)``
and only wait for readiness when the command actually needs the
agent. ``/mcp add``, ``/skills``, ``/new`` etc. no longer deadlock
behind a failing MCP load they are meant to fix.
* fix(cli): guard sync-back, subcommand-aware gating, /model adopt-path
Three CodeRabbit findings on command dispatch:
- ``_handle_command`` unconditionally synced ``ctx.agent`` back into
``agent_loader``. For non-agent commands ``ctx.agent`` is ``None``,
so ``/threads`` / ``/mcp`` / ``/skills`` (etc.) could clobber a valid
loaded agent — and rebind channel globals to ``None``. Guard the
sync on ``ctx.agent is not None``.
- ``/channel status`` and ``/channel stop`` don't touch ``ctx.agent``
but the class-level ``requires_agent = True`` blocked them behind
agent readiness. Added ``Command.needs_agent(args)`` (defaults to
``requires_agent``) so ``/channel`` can override with subcommand
awareness; kept the class flag for the common case.
- ``/model`` builds a new agent from scratch, it never reads the
existing one — gating it on readiness meant a broken provider
blocked the command that would fix it. Flipped it to
``requires_agent = False`` and added ``BackgroundAgentLoader.adopt``
so the UI can seat the replacement and supersede any in-flight
load (the generation token keeps a late completion from clobbering
the adopted agent).
Bonus cleanup: ``CommandManager.resolve`` now returns
``(command, args)`` so callers can invoke ``needs_agent`` without
re-implementing ``shlex`` parsing.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
|
||
|
|
28f3e81b4c |
feat: add /model command for changing models inside TUI/CLI (#162)
* rebase main * fix: apply pr comments * fix: fix critical issue * fix: ordering /model in cli mode * feat: refactor /model command handling and add Rich CLI support --------- Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com> |
||
|
|
aa3dd00409 |
feat: add support for session resumption with --resume flag and enhan… (#170)
* feat: add support for session resumption with --resume flag and enhance thread ID resolution * refactor(tests): streamline help output testing for --resume flag * feat: enhance session resume functionality with improved thread ID resolution and SQL wildcard handling * feat: improve error handling for resume hint retrieval in interactive modes * refactor: streamline logging for print_resume_hint failure in interactive mode * feat: implement deferred scrolling for Markdown-heavy content in interactive mode |
||
|
|
06822f236c | feat: enhance tool result handling with tool_call_id for concurrent execution | ||
|
|
f4a3617646 |
refactor(paths): unify global data directory to ~/.evoscientist and u… (#164)
* refactor(paths): unify global data directory to ~/.evoscientist and update related paths * refactor(paths): update legacy session migration to respect XDG_CONFIG_HOME * refactor(tests): clear XDG_CONFIG_HOME in legacy session migration tests for deterministic behavior |
||
|
|
210e8864f6 |
feat(memory): migrate MEMORY.md to global path & enhance ask-user prompts (#161)
* feat(prompt): enhance user interaction with multiple-choice and free-text questions * refactor(paths): rename MEMORY_DIR to MEMORIES_DIR for consistency * style(tests): format code for better readability in test cases * feat(prompt): add validation for 'other' option in user prompt * feat(prompt): refactor validation logic for user prompts and add skip option * feat(style): refactor to use shared _PICKER_STYLE from interactive module |
||
|
|
65db3a4fcd |
Add status bar and compact summary widgets with context window resolu… (#152)
* Add status bar and compact summary widgets with context window resolution - Implemented a shared status bar for CLI and TUI frontends, including helpers for managing session metrics and context windows. - Created a `CompactSummaryWidget` for displaying manual summaries in a collapsible format. - Introduced a `CompactingWidget` to indicate ongoing compacting processes. - Added a base class `TimedStatusWidget` for widgets that require a timer. - Developed context window resolution helpers to retrieve context window sizes from various model attributes. - Enhanced tests for context window resolution and status bar functionalities, ensuring accurate behavior across different scenarios. - Updated existing tests to cover new features and maintain code quality. * refactor(Channel): simplify lambda function in _send_with_retry method * feat: enhance context editing logic and improve error handling in StreamState * refactor(Channel): streamline lambda function in _send_with_retry method * feat: rename auto-approve option to auto-mode for unattended execution; update checkpoint queries to filter by agent name; improve compatibility validation logic * feat: rename auto-approve option to auto-mode; update related logic and tests for improved unattended execution * fix: correct formatting of console message for MCP server configuration status * feat: add check for None summary_message in _apply_summarization_event to prevent errors * feat: enhance _load_checkpoint_messages to validate message format and apply summarization event |
||
|
|
ff15f515cc |
Release/v0.0.7 (#151)
* chore(assets): update wechat_group image file * Refactor code structure for improved readability and maintainability * feat(backends): enhance MergedReadOnlyBackend with improved ls, grep, and glob methods * fix(docs): update WeChat QR code image link in README files * feat(skills): enhance skill management to support global and workspace tiers * style: apply ruff format to skills_cmd and commands/implementation/skills Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(skills): improve uninstall_skill to prevent removal of built-in skills * fix(docs): update skill installation documentation for clarity on global and user directories * fix(skills): enhance uninstall_skill to validate skill directory before removal * fix(skills): improve error handling in install_skill and uninstall_skill for directory creation and validation --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> |