Commit Graph

619 Commits

Author SHA1 Message Date
Xi Zhang 3ce6523faf fix: update deepagents and langchain versions (#251)
* fix: update deepagents and langchain versions; enhance _reduce_messages_delta handling for None state

* fix: update langchain version constraint to >=1.3 in pyproject.toml and uv.lock
2026-05-31 22:12:13 +01:00
Xi Zhang a13904185d Feat/sandbox execute timeout (#243)
* feat: implement configurable sandbox execute timeout and enhance recovery instructions

* feat: add background process management tools and middleware for sandbox execution

* feat: enhance background process management with completion notifications and deduplication

* feat: enhance sandbox execution timeout validation and update related messages

* feat: enhance background process management with thread-specific completion notifications and HITL approval handling

* test: assert completion notification waits for process finish timestamp
2026-05-31 15:11:25 +01:00
Ziheng Zhang 2364e6b130 fix(cli): forward async-notifier replies back to originating channel (#244)
* fix(cli): forward async-notifier replies back to originating channel

When PR #214's auto-notifier fires a synthetic agent turn after a
channel-originated conversation, the synthesized response only rendered
to the local CLI/TUI — the channel user (iMessage etc.) saw nothing
and had to manually re-prompt to find out what happened.

Adds a per-thread channel-origin registry in cli/channel.py and wires
the three notifier paths (Rich CLI / TUI / serve) to publish the final
response back via bus.publish_outbound when the originating thread was
started by a channel turn. Publish is fire-and-forget (scheduled on the
bus loop + done-callback for failure logging) so the notifier turn
doesn't block on the asyncio / textual event loop.

The registry is cleared on /new and /resume rotation so stale entries
don't accumulate.

* fix(cli): address review feedback on channel-origin forwarding

Follow-up to the review on #244 (din0s, X-iZhang):

- Guard the /resume origin cleanup on a real thread change in Rich CLI
  and TUI (serve mode already did via thread_changed). Resuming the
  already-active thread no longer wipes its still-live origin, which
  would otherwise silently drop a later async-notifier forward — the
  exact gap this PR closes.
- Re-bind the now-current thread to its channel after a channel-issued
  /new or /resume slash command (which rotates the thread inside the
  dispatch), so notifier turns on the rotated thread still forward.
- Guard the publish done-callback against a cancelled future, whose
  .exception() raises CancelledError (rather than returning it) on
  bus-loop teardown, so the intended warning still logs.
- Mirror the normal reply path's manager.record_message(channel, "sent")
  for forwarded notifications so per-channel stats stay accurate.
- Print the closing "[channel: Replied to ...]" line in all three
  notifier paths (Rich CLI / TUI / serve) when a forward actually
  happened, so the forwarded block reads as terminated on screen.

Adds test_publish_records_sent_metric. ruff clean; notification-origin
suite (10) + related channel/CLI/serve suites (728) pass.

* fix(cli): store sender information separately from chat_id in channel origin

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-05-31 14:43:42 +01:00
X-iZhang 721a03c25b feat: update model version from claude-sonnet-4-5 to claude-sonnet-4-6 and related adjustments 2026-05-29 00:40:46 +01:00
Xi Zhang f75bfcda51 Add onboarding wizard with style and validation components (#241)
* Add onboarding wizard with style and validation components

- Introduced `style.py` for shared visual elements used in the onboarding wizard.
- Created `validators.py` for input validation, including integer and choice validators, and API key validation functions for various providers.
- Implemented `wizard.py` as the entry point for the onboarding process, managing user prompts and configuration steps.
- Added progress rendering and autosave functionality to enhance user experience during the onboarding process.

* feat(onboarding): enhance validation and configuration for onboarding wizard

- Added validation for UI backends, workspace modes, and providers in the onboarding command.
- Updated channel definitions to include secret field handling for sensitive tokens.
- Improved user prompts for required fields, ensuring sensitive data is masked.
- Introduced constants for valid providers, UI backends, and workspace modes to maintain consistency.
- Implemented tests to ensure alignment between constants and interactive choices in onboarding steps.

* feat(onboarding): improve WeChat account ID prompt and validation for newly enabled channels

* feat(onboarding): enhance WeChat backend credential prompts and validation

* feat(onboarding): refine WeChat backend credential prompts for wecom and wechatmp

* Refactor onboarding package for improved structure and clarity

- Simplified the onboarding package by removing unnecessary re-exports and consolidating public API to only include `run_onboard`.
- Updated `install_back_keys` to `install_navigation_keys` for clarity and consistency in the prompter module.
- Enhanced the `NonInteractivePrompter` class to support strict mode, allowing for better handling of non-interactive prompts.
- Adjusted the onboarding steps to utilize the new navigation keys installation method.
- Improved the `run_onboard` function to handle section implications based on user flags, enhancing the onboarding experience.
- Updated tests to reflect changes in imports and ensure compatibility with the new structure.

* feat(onboarding): enhance validation logic for non-interactive prompts

* refactor(onboarding): streamline onboarding module structure and enhance validation error handling

* refactor(onboarding): enhance config revert logic to preserve original file state

* refactor(onboarding): enhance tavily key validation and error handling in onboarding process
2026-05-28 12:42:49 +01:00
Xi Zhang b9ad694467 fix: resolve path correctly when workspace name appears in parent path 2026-05-23 12:37:21 +01:00
Ziheng Zhang d2283397a4 feat(feishu): scan-to-create QR onboarding + silence unsubscribed WS events (#239)
* feat(feishu): scan-to-create QR onboarding flow

Add a device-code flow against accounts.feishu.cn/oauth/v1/app/registration
that lets users scan a terminal QR code with Feishu / Lark mobile to
auto-create a PersonalAgent bot app with the required IM permissions
pre-attached. The poll endpoint returns app_id + app_secret, which the
onboarding wizard then writes into the channel config — no manual app
creation on open.feishu.cn required.

- channels/feishu/onboard.py: qr_register() public entry, init/begin/poll
  helpers, QR rendering via the soft qrcode dep, automatic feishu↔lark
  domain switch based on the scanning user's tenant_brand, and a
  best-effort bot probe to surface the bot name in the wizard
- channels/feishu/__init__.py: re-export qr_register (mirrors qq)
- config/onboard.py: offer "Scan QR code (recommended) / Enter manually"
  in the Feishu branch, ask for region (feishu vs lark), then call
  qr_register and populate feishu_app_id / feishu_app_secret /
  feishu_domain; add qrcode>=7.4 to the feishu pip extras

* fix(feishu): silently absorb unsubscribed WebSocket events

Feishu auto-subscribes PersonalAgent apps to many event types
(im.message.reaction.created_v1, message.read_v1, message.recalled_v1,
chat.member.*, ...) that EvoScientist doesn't register handlers for.
Without intervention, lark-oapi's dispatcher raises EventException
("processor not found, type: ..."), the WS client logs it at ERROR and
replies HTTP 500 on the frame, and Feishu marks the event as failed
and retries it.

The problem is amplified by _send_ack_reaction: every inbound message
triggers our own reaction, which Feishu echoes back as
reaction.created_v1, creating a continuous ERROR-log feedback loop and
pointless retries.

Wrap EventDispatcherHandler._do_without_validation after build() to
swallow "processor not found" EventExceptions (debug log + return None)
while letting all other errors propagate. Failure-safe: if lark-oapi's
internal API changes the wrapper degrades to the prior behavior rather
than breaking the channel.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-05-20 22:53:16 +08:00
dependabot[bot] b11932f91b chore(deps): bump idna in the uv group across 1 directory (#238)
Bumps the uv group with 1 update in the / directory: [idna](https://github.com/kjd/idna).


Updates `idna` from 3.13 to 3.15
- [Release notes](https://github.com/kjd/idna/releases)
- [Changelog](https://github.com/kjd/idna/blob/master/HISTORY.md)
- [Commits](https://github.com/kjd/idna/compare/v3.13...v3.15)

---
updated-dependencies:
- dependency-name: idna
  dependency-version: '3.15'
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-05-20 15:39:12 +01:00
Xi Zhang 7959495a13 feat(deploy): add EvoSci deploy subcommand (#228)
* feat(deploy): implement standalone LangGraph server and CLI command for deployment

* feat(deploy): enhance port validation and environment variable management for deployment

* Refactor langgraph dev deployment and introduce workspace sidecar protocol

- Updated the deployment mode handling in `server.py` to use a single environment variable `EVOSCIENTIST_DEPLOY_MODE` with values `full` and `stripped`.
- Enhanced the `manager.py` to implement a workspace fingerprint sidecar, allowing cross-process reuse of langgraph dev instances while ensuring workspace consistency.
- Introduced functions to write and read the workspace sidecar, with error handling for missing or corrupt data.
- Added tests for the workspace sidecar functionality, including validation of the JSON schema and ensuring proper error handling for workspace mismatches.
- Updated existing tests to reflect changes in deployment mode handling and added new tests for signal handling during shutdown.
- Ensured that cleanup routines remove the workspace sidecar alongside the PID file during shutdown.

* fix(langgraph): improve workspace sidecar checks for process ownership and stale handles
2026-05-20 11:06:35 +01:00
X-iZhang 9c7347eedb chore: update version to v0.1.1 in badges and pyproject.toml 2026-05-19 14:06:22 +01:00
Xi Zhang 331056cdc8 feat(middleware): upgrade deepagents 0.5.7 → 0.6.2 (#231)
* feat(middleware): add CodeInterpreterMiddleware with project-specific configuration

chore(config): increase checkpoint retention limit for runaway conversations

fix(tests): update database schema references from 'blob' to 'value'

chore(deps): update deepagents dependency to include quickjs support

* feat(deepagents): update to version 0.6.1 and add optional dependencies for quickjs

* feat(sessions): improve error handling for message deltas and update Overwrite type check

* Enhance PruningCheckpointer with DeltaChannel Awareness

- Introduced a new pruning strategy in `_prune_after_put` to preserve the `_DeltaSnapshot` chain during checkpoint pruning.
- Implemented methods to fetch recent checkpoint IDs and walk to snapshot ancestors, ensuring that necessary checkpoints are retained.
- Updated SQL queries to handle checkpoint and write deletions more efficiently.
- Added comprehensive tests for DeltaChannel-aware pruning, ensuring that the pruning logic correctly handles various checkpoint scenarios, including those with and without snapshot seeds.
- Refactored `_load_checkpoint_messages` to utilize the new saver interface, improving message reconstruction from checkpoints.

* feat(tests): add migration sweep test to preserve snapshot ancestor

* feat(sessions): enhance checkpoint retrieval to prevent transcript leakage in multi-agent scenarios

* feat(middleware): enhance CodeInterpreterMiddleware with configurable timeout and result character limit

feat(config): add CodeInterpreterMiddleware tuning parameters to EvoScientistConfig

feat(sessions): implement inline message delta reducer for improved message handling

* feat(dependencies): update deepagents version to 0.6.2 in pyproject.toml and uv.lock
2026-05-19 12:35:36 +01:00
Xi Zhang 385f9756c1 feat(backends): implement tier-aware virtual mount resolution for ski… (#236)
* feat(backends): implement tier-aware virtual mount resolution for skills and memories

* test: add end-to-end test for workspace tier shadowing global tier in CustomSandboxBackend

* feat(backends): enhance virtual mount resolution for skills and memories with tier paths and quoting

* fix(tests): update Python command in virtual mount resolution tests to use python3
2026-05-19 11:50:49 +01:00
Ziheng Zhang 7f1aa3b0f6 fix(cli): handle spaces in @file mentions (#234)
* fix(cli): handle spaces in @file mentions

The @file parser truncated at the first space, so dragging or pasting a
filename like `@PREPING_ Building Agent.pdf` only matched `@PREPING_`
and warned "file not found". Now supports `@"..."` / `@'...'` quoted
form for explicit paths, plus a greedy expansion fallback that walks
across whitespace until an existing file resolves (bounded by newlines,
the next `@`, and a 20-token cap). Autocomplete also returns quoted
mentions for any candidate containing a space.

* style: apply ruff format to file_mentions
2026-05-18 12:05:21 +01:00
X-iZhang 0d51406149 chore: update wechat_group image asset 2026-05-16 16:07:32 +01:00
Wiktor Cupiał a4c9c779c9 feat: status and elapsed time indicator (#218)
* feat: status and elapsed time indicator

* test: add tests for tui-status

* fix: move to enum+switch, change phase calculation

* feat: remove 'done' phase
2026-05-13 14:15:39 +01:00
dinos 4b0c91190a feat(llm): add dashscope-code provider for Alibaba Coding Plan keys (#225)
* feat(llm): add dashscope-code provider for Alibaba Coding Plan keys

Alibaba Cloud Bailian "Coding Plan" subscription keys (sk-sp-*) route
through a separate endpoint (coding.dashscope.aliyuncs.com/v1) that the
standard `dashscope` provider can't reach. Add a sibling provider entry
matching the zhipu/zhipu-code and moonshot/kimi-coding precedents, with
its own validator (the coding endpoint returns 404 on /models, so probe
via chat.completions instead).

Closes #224

* fix(llm): keep dashscope as default provider for qwen3-coder shortcut

The MODELS dict is built from _MODEL_ENTRIES via a last-write-wins dict
comprehension. The initial commit listed dashscope-code AFTER dashscope,
which silently flipped the bare `get_chat_model("qwen3-coder")` shortcut
to the coding endpoint — breaking standard sk-* keys.

Reorder to match the zhipu-code / zhipu precedent: coding endpoint first,
general endpoint last so the general endpoint wins the collision and
remains the default for the shared "qwen3-coder" short name.
2026-05-13 10:23:32 +01:00
dependabot[bot] 35ea2bfb52 chore(deps): bump urllib3 in the uv group across 1 directory (#222)
Bumps the uv group with 1 update in the / directory: [urllib3](https://github.com/urllib3/urllib3).


Updates `urllib3` from 2.6.3 to 2.7.0
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.6.3...2.7.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.7.0
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-12 08:53:08 +01:00
Ziheng Zhang 8fe774b056 Feat/qq interactive buttons (#220)
* feat(qq): add inline keyboard buttons for C2C HITL approval

QQ Bot supports inline buttons via `markdown + keyboard` payloads. Clicks
arrive as `interaction_create` events through the existing botpy
WebSocket gateway — no extra subscription needed beyond enabling the
`interaction` intent. Group-scope clicks are out of scope here (DM only).

Send path
- `_build_qq_keyboard(buttons)` mirrors the Feishu helper, mapping the
  generic `{text, value, type}` shape to QQ's `{render_data, action}`
  with action.type=1 (callback). One button per row for mobile clarity.
- `_send_chunk` extracts `metadata["buttons"]` and threads a `keyboard`
  payload into `_post_markdown_message` for C2C only.
- Markdown→plain fallback can't carry a keyboard, so when buttons were
  attached the fallback content gets a textual `Reply: 1=Approve, …`
  hint built from the button list. `_parse_approval_reply` accepts
  the same values typed manually, so the user is never stuck.

Receive path
- `on_interaction_create` is registered on the bot class.
- `_on_interaction` extracts `data.resolved.button_data`, builds an
  InboundMessage, runs it through inbound middleware (Dedup suppresses
  retry callbacks), and publishes directly to the bus — bypassing the
  per-sender debounce buffer so the click value isn't merged with any
  text typed in the same window.
- Always ACKs via `api.on_interaction_result(id, 0)` in `finally` so
  QQ doesn't show the button as "expired", even if middleware drops
  the click or something throws downstream.

`QQ.inline_buttons=True`; `_approval_prompt_metadata` now auto-attaches
the Approve/Reject/Approve-all button row for QQ HITL prompts.

* fix(qq): button-value coercion, ACK timing, HITL consumer wiring

Fixes 6 bugs found in the inline-keyboard commit and consolidates the
button helpers so the keyboard builder, plain-text fallback hint, and
interaction handler share one coercion path.

- Plain-text fallback no longer crashes on non-string `value` (e.g.
  `{"text": "OK", "value": 42}`).  Extracted `_normalize_button` is now
  the single place that resolves `(label, value)` and coerces non-strings.
- `metadata["button_value"]` is the coerced string instead of the raw
  payload, matching `content` and downstream string comparisons.
- `_on_interaction` ACKs first, before publishing to the bus, so the
  QQ button UI never shows "expired" if middleware is slow.
- Wire `_approval_prompt_metadata` + `_format_approval_prompt(with_buttons=)`
  into `InboundConsumer._stream_with_hitl` and `cli.channel.channel_hitl_prompt`
  so the QQ `inline_buttons=True` capability is actually used end-to-end
  (HITL prompts auto-attach Approve/Reject/Approve-all buttons when the
  channel advertises the capability).
- Trim contradictory `_QQ_DEFAULT_PERMISSION` comment.
- Fix `test_group_interaction_ignored` docstring (ACK runs first now,
  not in `finally` after a `return`).

Tests: `_normalize_button` covered indirectly via existing keyboard tests;
new regressions for non-string fallback hint, ACK-on-handler-throw, and
string-coerced `button_value` metadata.

* refactor(qq): slim button helpers and explicit has_buttons flag

Inline single-use _button_hint and the _QQ_BUTTON_STYLE/_QQ_DEFAULT_PERMISSION
constants in qq/channel.py; tighten _on_interaction (drop unreachable
"[button click]" sentinel and unused triggering_message_id metadata; collapse
"if resolved else" ternaries via `or ""`).

Replace the metadata round-trip ("buttons" in metadata) used to detect button
support in consumer.py and cli/channel.py with an explicit has_buttons bool
threaded through both the prompt formatter and metadata builder.

Apply ruff format to the previously unformatted blocks introduced earlier on
this branch so CI lint passes.

* feat(qq): send post-decision confirmation after HITL approval

Send a visible confirmation message ("✅ 已批准" / "❌ 已拒绝") right after
the user resolves a HITL approval — QQ Bot has no message-recall or edit API
for C2C, so a follow-up message is the only way to give the click/reply
strong feedback.

Bus consumer (consumer.py): only sends the confirmation when the user
actually responded (event was set), to avoid pretending the user approved
when the request really timed out and auto-approved.

CLI HITL prompt (cli/channel.py): mirrors the same set of confirmation
strings.  Timeout and unrecognized-reply paths keep their existing English
text since their semantics differ (auto-reject vs auto-approve, plus a
hint about the unparsed input).
2026-05-11 10:44:20 +01:00
dependabot[bot] ccb3083183 chore(deps): bump langchain-core in the uv group across 1 directory (#219)
Bumps the uv group with 1 update in the / directory: [langchain-core](https://github.com/langchain-ai/langchain).


Updates `langchain-core` from 1.3.2 to 1.3.3
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain-core==1.3.2...langchain-core==1.3.3)

---
updated-dependencies:
- dependency-name: langchain-core
  dependency-version: 1.3.3
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-09 16:50:25 +01:00
X-iZhang f00ee1b5cc chore: update version to v0.1.0 2026-05-08 22:57:30 +01:00
Xi Zhang c407d2e20f Fix/async subagent model switch (#217)
* feat(middleware): add ConfigurableModelMiddleware for dynamic model resolution

- Introduced ConfigurableModelMiddleware to resolve chat models from RunnableConfig.configurable on each call.
- Updated middleware initialization to include ConfigurableModelMiddleware.
- Enhanced context editing middleware tests to verify presence of ConfigurableModelMiddleware.
- Implemented tests for ConfigurableModelMiddleware to ensure correct model overriding and caching behavior.
- Added tests for deepagents model-passthrough patch to verify configuration injection in async tasks.

* feat(async-subagent): update middleware handling to prevent deadlocks in async sub-agents

* style: Refactor code formatting for improved readability in patches and test files

* refactor: streamline middleware construction and improve async handling in ConfigurableModelMiddleware

* fix: remove unused request parameter from _read_model_override function

* refactor: improve async handling in _ClientProxy and enhance logging in ConfigurableModelMiddleware
test: add behavior test to ensure AskUserMiddleware is excluded in async subagent mode
2026-05-08 22:52:24 +01:00
Ziheng Zhang 89b0ecdbf3 feat(qq): add QR-code scan-to-configure onboarding for QQ Bot (#213)
* feat(qq): add QR-code scan-to-configure onboarding for QQ Bot

Adds a `qr_register()` flow that drives q.qq.com's create_bind_task /
poll_bind_result APIs so the wizard can auto-fill `qq_app_id` and
`qq_app_secret` after the developer scans a QR code with a bound QQ
account, falling back to manual entry on failure or cancel.

- channels/qq/crypto.py: AES-256-GCM helpers for decrypting the bot's
  client_secret returned by poll_bind_result.
- channels/qq/onboard.py: portal API client + polling loop.
- channels/qq/__init__.py: re-export `qr_register`.
- config/onboard.py: QQ branch in `_step_channels` that offers
  "Scan QR code" vs "Enter manually", and skips the manual prompt
  loop when a scan succeeded.

* style(qq): fix ruff lint errors in onboard.py

Move `import os` to the top-level import block (E402), drop the legacy
`typing.Optional`/`typing.Tuple` imports (UP035), and use the PEP 585/604
builtin generics (`tuple[...]`, `X | None`) for the few annotations that
still referenced them (UP006/UP045). No behavior change.

* fix(qq): harden QR onboard error paths and declare scan deps

Address review feedback on PR #213:
- Declare cryptography>=41.0 and qrcode>=7.4 in [qq]/[all-channels]
  extras and in _CHANNEL_PIP_DEPS so the scan flow no longer fails
  with an opaque ImportError on a fresh `evoscientist[qq]` install.
- Polling loop logs each _poll_bind_result failure and aborts after
  5 consecutive errors instead of silently spinning until the 600s
  timeout, restoring the documented Raises: RuntimeError contract.
- Wrap decrypt_secret in try/except so failures honor the
  None-on-failure contract instead of letting exceptions escape.
- Preflight `import cryptography` in the scan branch and offer
  install or fall back to manual entry.

* style: ruff format collapse two over-wrapped log/console lines
2026-05-08 12:08:52 +01:00
Xi Zhang 4e04ac5b72 fix: Improve watcher logic to prevent false-positive notifications on… (#216)
* fix: Improve watcher logic to prevent false-positive notifications on clean stream exits

* fix: Update watcher logic to drop notifications on persistent runs.get failures

* fix: Refactor test for watcher persistent failure notification handling

* fix: Enhance watcher test to validate all notification queues are empty after reconnect budget exhaustion
2026-05-07 22:16:36 +01:00
Xi Zhang 80f1f4fa0f feat: Implement async sub-agent auto-notification system (#214)
* feat: Implement async sub-agent auto-notification system

- Added async notifier functionality to handle notifications for sub-agents reaching terminal states.
- Introduced `AsyncTaskNotification` dataclass for structured notification data.
- Implemented `watch_run_and_notify` to monitor agent runs and enqueue notifications.
- Created `spawn_watcher` to manage watcher tasks and ensure proper cancellation of previous watchers.
- Developed `consume_notifications` to process notifications, deduplicate them, and format messages for LLM.
- Added tests for notification handling, including draining, deduplication, and formatting.
- Patched deepagents to integrate the new watcher functionality into start and update tools.

* Enhance async notifier with per-thread notification routing and error handling

- Introduced `origin_cli_thread_id` to `AsyncTaskNotification` for routing notifications back to the originating CLI session.
- Implemented per-thread notification queues to handle notifications based on the originating thread.
- Updated `has_pending_notifications` and `drain_notifications` to respect thread-specific queues.
- Enhanced `watch_run_and_notify` to detect in-band error events from the SSE stream and handle clean exits.
- Modified tests to verify the new notification routing behavior and ensure proper handling of notifications across threads.
- Added a fixture to restore the async watcher patch state in tests to prevent state leakage.
- Updated deepagents patching to capture the main agent's CLI thread ID for notification routing.

* feat: Enhance async notifier with thread-specific watcher management and notification filtering

* test: Enhance notification draining logic for cleaner test setup

* refactor: Remove summary field from AsyncTaskNotification and update related tests

* feat: Enhance async notification handling with target thread ID support

* Refactor async notifier and middleware for improved task management

- Removed the no-op shutdown watcher loop from async_notifier.py as it is no longer needed.
- Updated watch_run_and_notify to clarify notification handling and race conditions.
- Cleaned up shutdown handling in commands.py, interactive.py, and tui_interactive.py by removing obsolete shutdown watcher calls.
- Deleted the deepagents async watcher patch from patches.py, transitioning to a new middleware approach.
- Introduced AsyncWatcherMiddleware to handle async task notifications directly during tool calls.
- Updated tests to validate the new middleware functionality and ensure proper watcher spawning and cancellation.
- Enhanced test coverage for async watcher middleware, including edge cases and error handling.

* feat(tests): add fixture to reset notifier state before each test
2026-05-07 16:03:02 +01:00
Ziheng Zhang 692dc491ac # feat(wechat): add personal-WeChat (iLink) backend with QR login (#212)
* feat(wechat): add personal-WeChat (iLink) backend with QR-code login

Adds a third WeChat backend alongside WeCom and Official Account:
``personal`` rides Tencent's iLink Bot long-poll gateway so a personal
WeChat account can act as a bot.  Credentials are obtained via QR-code
scan and persisted under ``DATA_DIR/wechat_personal/accounts/``.

- channels/wechat/personal.py: WeixinPersonalChannel + qr_login.
- channels/wechat/crypto.py: aes128_ecb_decrypt + parse_ilink_aes_key
  for the iLink CDN media protocol.
- channels/wechat/probe.py: validate_wechat_personal credential probe.
- channels/wechat/serve.py: --backend personal CLI + --qr-login flow.
- channels/wechat/__init__.py: factory dispatch on wechat_backend; pull
  in the new dependencies in the docstring.
- config/settings.py: wechat_personal_* fields.
- config/onboard.py: WeChat-backend picker + QR-scan flow in the wizard
  + personal-backend probe in _probe_channel.
- pyproject.toml / uv.lock: add qrcode + certifi to wechat & all-channels
  extras (aiohttp was already pulled in transitively).

* fix(wechat): address ruff failures and CodeRabbit review on personal-WeChat PR

- personal.py: drop unused imports (`field`, `PollingMixin`); replace
  `asyncio.TimeoutError` with builtin; hold references to background
  `asyncio.create_task` results so they aren't GC'd; wire `dm_policy`
  through `_process_message` (disabled/allowlist) so `wechat_personal_dm_policy`
  actually takes effect for DMs.
- onboard.py: import-check gate now validates the full WeChat dependency
  set (aiohttp, qrcode, Crypto, certifi) instead of only aiohttp; mask
  `WeCom Secret` and `MP App Secret` prompts via `questionary.password`;
  derive the QR-login hint path from `_account_dir()` instead of the
  hard-coded `~/.evoscientist/...`; stop copying the QR-login token into
  the main config (already persisted per-account on disk — copying broadens
  secret exposure and risks staleness).
- pyproject.toml: allow Chinese full-width punctuation in `allowed-confusables`
  for user-facing CN messages.

* style(wechat): apply ruff format

`ruff format --check` was failing CI on three files (one pre-existing in
`__init__.py` plus formatter-driven line-merges in the files touched by
the previous fix commit). Ran `ruff format` to bring them in line; both
`ruff check` and `ruff format --check` now pass.
2026-05-07 16:33:30 +02:00
Wiktor Cupiał 9e51ec6fdd feat(cmd): add /model-fallback command (#196)
* feat(cmd): add /model-fallback command

* fix: apply feedback

* fix: lock usage with _fallback_chain

* fix: apply feedback

* fix: apply feedback

* feat: add tests

* fix: tests

* Update EvoScientist/middleware/model_fallback.py

Co-authored-by: dinos <dinospk1999@gmail.com>

---------

Co-authored-by: dinos <dinospk1999@gmail.com>
2026-05-07 15:35:26 +02:00
Ziheng Zhang 22a65b640d refactor(channels): remove dead MessageBus dispatcher (#205)
* refactor(channels): remove dead MessageBus dispatcher

Outbound routing has two implementations: ``MessageBus.dispatch_outbound``
(subscriber-based) and ``ChannelManager._dispatch_outbound`` (registry
lookup).  Only the latter is ever started in production — the former
is reachable solely from tests, yet both consume from the same
``bus.outbound`` queue.  If anyone followed the bus's own API surface
they would silently steal messages from the real dispatcher.

Drop the unused machinery to leave a single, obvious outbound path:
- ``MessageBus.subscribe_outbound`` / ``dispatch_outbound`` / ``stop``
- ``_running`` flag and ``_outbound_subscribers`` map
- ``OutboundCallback`` type alias
- The lone ``bus.stop()`` call in ``cli/channel.py`` (was no-op)
- Four tests covering the removed code paths

* test(channels): drop empty MessageBus stubs after dispatcher removal

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-05-07 12:11:06 +01:00
Xi Zhang d7c0eec0e9 chore: update dependencies and remove unused OpenRouter patch (#211) 2026-05-05 18:51:00 +01:00
Xi Zhang f41584e10b Refactor sub-agent architecture and introduce async support (#200)
* Refactor sub-agent architecture and introduce async support

- Removed the legacy subagent.yaml file and replaced it with individual YAML files for each sub-agent in the subagents directory.
- Updated the load_subagents function to support both directory and single file layouts for loading sub-agent configurations.
- Added new langgraph_dev module for managing async sub-agent lifecycle and deployment.
- Created graphs for async sub-agents (writing-agent, data-analysis-agent) and updated langgraph.json for deployment.
- Introduced new sub-agent definitions for planner, research, debug, code, and writing agents with appropriate system prompts and configurations.
- Enhanced package data inclusion in pyproject.toml to accommodate new sub-agent YAML files.

* Refactor code for improved readability by consolidating conditional statements and formatting

* feat: enhance async sub-agent support with workspace synchronization and user feedback

- Added console status messages during async sub-agent server startup and workspace synchronization to improve user experience.
- Implemented a new WorkspaceSyncWidget for live feedback during workspace sync operations.
- Updated onboarding to reject occupied ports and ensure proper workspace handling for async sub-agents.
- Introduced locking mechanisms to manage concurrent access to langgraph dev processes and workspace states.

* feat: add async sub-agent configuration and server management functions

* feat: improve port occupation handling and log file management in start_langgraph_dev

* feat: enhance async sub-agent handling and introduce comprehensive tests

- Updated `_maybe_swap_async_subagents` to improve async sub-agent management, ensuring internal flags are stripped before handoff.
- Enhanced port management in `onboard.py` to allow reuse of occupied ports if already running by the same service.
- Introduced file locking in `manager.py` to prevent race conditions during concurrent CLI invocations.
- Added new tests for async sub-agent swapping and langgraph manager functionalities to ensure reliability and correctness.
- Updated dependencies in `pyproject.toml` to include `psutil` and `filelock`.

* fix(docs): clarify sub-agent configuration in README

* test(manager): isolate _PID_DIR + tighten reuse-path assertion

Addresses CodeRabbit review on tests/test_langgraph_manager.py:

- Patch _PID_DIR to tmp_path so the FileLock setup in
  ensure_langgraph_dev doesn't mkdir the user's real
  ~/.config/evoscientist/ dir as a test side-effect.
- Tighten "result is None or hasattr(result, 'poll')" to a strict
  "result is None" — the reuse path returns None unconditionally,
  so the OR clause was hiding potential regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(manager): clean up stale PID file when unrelated process reuses PID

* feat(tests): add validation tests for async flag in load_subagents

* fix(load_subagents): restrict to .yaml files and clarify configuration handling

* fix(load_subagents): improve error handling for non-dict specifications in YAML

* feat(onboard): add "LangGraph Port" step to onboarding process

* feat(langgraph): add concurrency configuration for langgraph dev workers

* feat(async-subagents): enhance MCP tool routing for async sub-agents

* fix(manager): update exception handling for connection errors and prevent zombie processes

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 12:45:50 +01:00
X-iZhang d5ffa00263 docs(README): enhance Docker section with installation details and usage instructions 2026-05-04 15:37:05 +01:00
X-iZhang a297445f92 fix(assets): update wechat_group image file 2026-05-04 10:51:33 +01:00
Xi Zhang ab3787f86a fix(markdown): ensure proper spacing for ATX headings in Markdown ren… (#201)
* fix(markdown): ensure proper spacing for ATX headings in Markdown rendering

* fix(markdown): improve docstrings and tests for heading spacing functionality
2026-05-02 00:59:30 +01:00
Xi Zhang 29e7fd383d fix/cli hitl (#202)
* feat(display): enhance approval prompt with questionary for better navigation

* feat(cli): add HumanInTheLoopMiddleware for user approval in main agent
2026-05-02 00:54:40 +01:00
dinos 0d8ac4f24b feat(docker): official image with all runtime deps pre-installed (#198)
* feat(docker): official image with all runtime deps pre-installed

Multi-stage build using uv for the EvoScientist core + all messaging-channel
extras, plus Node.js 24 LTS (for npx-based MCP servers) and uv (for runtime
Python MCP installs) in the runtime layer. Runs as non-root user evosci,
with workspace, app data, and config (XDG_CONFIG_HOME) all consolidated
under a single /home/evosci/.evoscientist volume so a single mount
persists everything across container restarts.

Includes a docker-compose.yml starter, a build/push GitHub Actions
workflow targeting ghcr.io with multi-arch (amd64/arm64) and PR-only
build verification, a .dockerignore, and a new Docker section in the
README documenting mounts, derivation recipes for the unbundled stt /
oauth / TinyTeX extras, and proxy/cert handling expectations.

* fix(docker): pin trixie base + drop redundant python image

Switch builder and runtime from `python:3.11-slim-bookworm` to a single
`ghcr.io/astral-sh/uv:python3.11-trixie-slim` base — trixie drops several
CRITICAL vulnerabilities that bookworm carries today, and reusing the uv
image for runtime eliminates the separate `COPY --from=…/uv` line.

* chore(docker): pin GitHub Actions to commit SHAs in workflow

Replace mutable major-version tags with full commit SHAs (with the
corresponding semver tag in a trailing comment) so a compromised /
retagged action release can't silently change what runs in the publish
pipeline.

* chore(deps): enable Dependabot version updates for Dockerfile pins

Adds a weekly `docker` ecosystem that watches the Dockerfile's `FROM` /
`COPY --from=` references — including the ARG-bound `BASE_IMAGE` and
`NODE_IMAGE` digests — and opens one grouped PR per cadence bumping
both the @sha256 digest and the trailing version comment. This keeps
the otherwise-frozen pins flowing with Debian point releases and
upstream patches.

* fix(docker): use nodejs alias stage so NODE_IMAGE ARG actually resolves

`COPY --from=${NODE_IMAGE}` left the dollar-curly literal at parse time
under buildkit 29.x — it expands ARGs in `FROM` but reads `--from=` as a
static stage/image name. Introduce a tiny `FROM ${NODE_IMAGE} AS nodejs`
alias and `COPY --from=nodejs …` against it, which preserves the
ARG-driven Dependabot updates without tripping the parser.

* fix(docker): harden venv ownership and PATH ordering

- Drop `--chown` on the `/opt/venv` COPY so the venv stays root-owned.
  The runtime user only needs read+execute (default Unix perms allow
  that); making it user-owned let the agent rewrite its own
  dependencies, which defeats the sandboxing premise. All persistent
  agent state already lives under /home/evosci/.evoscientist/.
- Reorder PATH so /opt/venv/bin precedes the user-writable
  UV_TOOL_BIN_DIR. Otherwise a stray binary dropped into the latter
  (e.g. via `uv tool install`) could shadow the canonical
  `evosci` / `python` / `pip` shipped with the image.

* docs: update README

* docs(docker): warn about non-root UID and `curl | sh` for derived images

- The image runs as `evosci` (UID 1000), so a host-side `./workspace`
  bind mount fails if the host user has a different UID — same gotcha
  that bites onboarding's `mcp.yaml` write. Add an !IMPORTANT block
  with the two practical fixes (`chown -R 1000:1000` once, or
  `--user "$(id -u):$(id -g)"` on each run).
- The TinyTeX derivation snippet pipes an unpinned remote installer
  into `sh`. Add a one-line pointer to fetching a pinned release
  tarball from `rstudio/tinytex-releases` for users who'd rather not
  trust the upstream script blindly. The official installer is kept
  as the default since that's what TinyTeX itself recommends.

* chore(docker): cancel in-flight workflow runs + flag iMessage as host-only

- Add `concurrency: cancel-in-progress: true` to the docker workflow
  so successive pushes on the same ref supersede the prior run rather
  than queueing in parallel — multi-arch buildx is the slowest job in
  CI, no point burning minutes on superseded builds.
- Spell out that the docker image installs the `all-chanels` extra and
  call out iMessage as a deliberate host-only exclusion: it requires
  the `imsg` CLI bridging to macOS's Messages.app, which no Linux
  container config can satisfy.
2026-05-01 13:24:14 +02:00
dinos 73928a2d78 fix(onboard): detect installed skill packs via install manifest (#199)
* fix(onboard): detect installed skill packs via install manifest

Onboarding's _step_skills only inspected USER_SKILLS_DIR and matched
recommended entries by directory-name hint, so a pack like
EvoScientist/EvoSkills@skills (which explodes into paper-writing/,
evo-memory/, etc. under GLOBAL_SKILLS_DIR) was never detected and kept
appearing as not-yet-installed.

skills_manager now writes a per-tier .installed.yaml mapping skill
directory name -> original install source on every install, removes the
entry on uninstall, and exposes installed_sources(). _step_skills checks
both tiers and treats a recommended source as installed when present in
any manifest -- so packs are recognized regardless of how their child
dirs are named.

* fix(onboard): write install manifest atomically

Stage to a sibling temp file, fsync, then os.replace into place. A crash
mid-write can no longer leave a half-written .installed.yaml behind,
which would otherwise wipe out pack detection until the next reinstall.

* style: fmt

* fix(onboard): catch decode errors when loading install manifest

read_text() can raise UnicodeDecodeError on a hand-edited or corrupt .installed.yaml; pin encoding="utf-8" and add UnicodeError to the except clause so the function honors its "returns {} on any error" contract.
2026-05-01 13:23:59 +02:00
X-iZhang 27a97c3097 Refactor code structure for improved readability and maintainability 2026-04-30 22:00:19 +01:00
dinos 5c829942d7 refactor(cli): replace channel module globals with ChannelRuntime (#197)
* refactor(cli): replace channel module globals with ChannelRuntime

Removes _cli_agent / _cli_thread_id from EvoScientist/cli/channel.py
and threads a ChannelRuntime via CommandContext.channel_runtime so
/model and /channel rebind without poking module-level state.

* fix(cli): address coderabbit review

- _auto_start_channel: bind ChannelRuntime only after
  _start_channels_bus_mode succeeds, so a startup failure no longer
  leaves a stale binding pointing at channels that never started.
- _sync_tui_command_completion (TUI) and the Rich CLI command-completion
  paths: rebind the runtime on thread rotation, not just agent swap, so
  /new and /resume keep ChannelRuntime in sync with the running thread
  (matches the serve-mode hook contract).
- test_hook_syncs_channel_runtime: pin ctx.thread_id explicitly so a
  bare MagicMock attribute can't silently mutate runtime.thread_id.
- New regression test covering the rebind-on-thread-rotation contract.
2026-04-30 18:06:22 +01:00
Xi Zhang 7cfec02416 Fix/sessions migration sweep race (#195)
* fix(sessions): ensure migration sweep runs before yielding checkpointer to prevent race conditions

* fix(sessions): enhance migration sweep with progress indication and ETA estimation
2026-04-29 02:06:25 +01:00
Xi Zhang 50719ef256 Implement PruningCheckpointer for efficient checkpoint management and… (#194)
* Implement PruningCheckpointer for efficient checkpoint management and add comprehensive tests

- Introduced `PruningCheckpointer` to manage checkpoint pruning after each `aput()`, ensuring only the latest checkpoints are retained based on a configurable limit.
- Added migration sweep functionality to clean up legacy checkpoints and prevent database bloat.
- Enhanced `get_checkpointer()` to utilize the new `PruningCheckpointer` and trigger migration sweeps when necessary.
- Developed a suite of integration tests for `PruningCheckpointer`, covering various scenarios including pruning behavior, concurrent writes, and retention policies.
- Implemented tests for migration sweep functionality, ensuring proper partitioning and user version management.
- Added diagnostic helper `db_stats` to provide insights into the database state, including thread and checkpoint counts.

* feat(sessions): enhance pruning logic to handle legacy DBs without writes table

* fix(tests): prevent atexit hook leakage in TestMigrationSweep

* feat(tests): enhance TestPruningCheckpointer to validate put+prune serialization

* feat(tests): refactor mock path implementation for get_db_path in test cases
2026-04-28 22:26:59 +02:00
Xi Zhang 56cc2fef85 fix(deepseek): add empty-string fallback for reasoning_content in cross-provider scenarios (#192) 2026-04-27 19:19:50 +01:00
Xi Zhang 52f8d3a3a5 Feat/llm context window patch table (#191)
* feat(context-window): add model context window patch table and apply function

* fix(tests): clean up formatting in context window tests

* fix(tests): update context window tests for Claude model exceptions
2026-04-27 18:28:57 +01:00
X-iZhang 26e3452ef6 chore: update version to v0.0.9 in badges, README, and project files 2026-04-26 15:16:12 +01:00
Xi Zhang 20c06d4897 feat(deepseek): implement reasoning_content passback for multi-turn s… (#190)
* feat(deepseek): implement reasoning_content passback for multi-turn scenarios

* fix(tests): ensure consistent import of EvoScientist.llm.patches in test cases

* fix(tests): streamline tool_calls formatting in TestPatchDeepseekReasoningPassback

* feat(patches): add reasoning_content capture and re-injection for DeepSeek assistant messages

* fix(deepseek): optimize reasoning_content extraction and assignment in passback

* fix(deepseek): refine reasoning_content handling in OpenAI capture patch
2026-04-26 14:53:08 +01:00
Xi Zhang e48bc1cb71 Refactor/system prompt structure (#189)
* feat: Refactor system prompt structure and enhance documentation for clarity

* refactor: Improve clarity and consistency in prompt documentation

* refactor(tests): Improve readability of first-person avoidance test assertion

* refactor: Remove redundant datetime imports and enhance prompt documentation
2026-04-26 11:08:08 +01:00
Ziheng Zhang da74c325d6 fix(channel): scope stop and restore resume history (#186)
* fix(channel): scope stop and restore resume history

* refactor(channel): simplify stop and resume patch

* Delete PR_MESSAGE.md

* fix(channel): address review feedback

* fix(channel): address remaining review bugs

* fix(channel): clean up stopped request handling

* fix(channel): preserve resolved replies and sync tui commands

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-25 16:12:49 +01:00
X-iZhang 27fee3256e fix(clipboard): remove unnecessary check for selection end in copy_selection_to_clipboard 2026-04-25 13:04:19 +01:00
Xi Zhang 558360b558 feat: Enhance ModelPickerWidget for Ollama integration (#187)
* feat: Enhance ModelPickerWidget for Ollama integration

- Implemented a sentinel row for "Custom Ollama model..." in ModelPickerWidget, allowing users to input arbitrary model names.
- Updated action handling in ModelPickerWidget to manage transitions between list and input modes.
- Added async model discovery for Ollama models, integrating with the /model command to fetch locally installed models.
- Created tests for Ollama model discovery and ModelPickerWidget behavior, ensuring proper functionality and user experience.
- Refactored validate_ollama_connection and discover_ollama_models for improved error handling and response management.

* fix: Simplify code by removing unnecessary line breaks in ModelPickerWidget and test cases

* fix: Restore globals on set_chat_model failure to prevent half-switched session

* fix: Improve error handling in ModelCommand by restoring globals on failure
2026-04-25 00:37:48 +01:00
Xi Zhang 49b03c36eb feat: add support for gpt-5.5 model in the LLM configuration and tests (#188) 2026-04-24 23:52:55 +01:00
Wiktor Cupiał 155c4eaa40 fix: text copy on remote sessions/legacy terminal emulators (#185)
* fix: text copy on remote sessions/legacy terminal emulators

* fix: display warning only once
2026-04-24 23:16:39 +02:00
Xi Zhang c134a16e19 Fix/channel slash rich cli (#184)
* feat(cli): implement slash command dispatch for channel messages

* fix(cli): make EvoSci serve exit on Ctrl+C and hot-swap /model (#181)

* fix(cli): streamline debug logging and formatting in channel command handling

* fix(cli): ensure proper handling of asyncio event loop in slash command processing

* fix(cli): add error handling for unexpected exceptions in slash command dispatch

* fix(cli): improve error messaging for slash command dispatch failures

* fix(tests): enhance test setup by restoring channel globals and simplifying assertions

* fix(cli): enhance slash command handling across UI surfaces and improve resume command warnings
2026-04-24 16:10:25 +01:00