Commit Graph

604 Commits

Author SHA1 Message Date
dinos 4b0c91190a feat(llm): add dashscope-code provider for Alibaba Coding Plan keys (#225)
* feat(llm): add dashscope-code provider for Alibaba Coding Plan keys

Alibaba Cloud Bailian "Coding Plan" subscription keys (sk-sp-*) route
through a separate endpoint (coding.dashscope.aliyuncs.com/v1) that the
standard `dashscope` provider can't reach. Add a sibling provider entry
matching the zhipu/zhipu-code and moonshot/kimi-coding precedents, with
its own validator (the coding endpoint returns 404 on /models, so probe
via chat.completions instead).

Closes #224

* fix(llm): keep dashscope as default provider for qwen3-coder shortcut

The MODELS dict is built from _MODEL_ENTRIES via a last-write-wins dict
comprehension. The initial commit listed dashscope-code AFTER dashscope,
which silently flipped the bare `get_chat_model("qwen3-coder")` shortcut
to the coding endpoint — breaking standard sk-* keys.

Reorder to match the zhipu-code / zhipu precedent: coding endpoint first,
general endpoint last so the general endpoint wins the collision and
remains the default for the shared "qwen3-coder" short name.
2026-05-13 10:23:32 +01:00
dependabot[bot] 35ea2bfb52 chore(deps): bump urllib3 in the uv group across 1 directory (#222)
Bumps the uv group with 1 update in the / directory: [urllib3](https://github.com/urllib3/urllib3).


Updates `urllib3` from 2.6.3 to 2.7.0
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.6.3...2.7.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.7.0
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-12 08:53:08 +01:00
Ziheng Zhang 8fe774b056 Feat/qq interactive buttons (#220)
* feat(qq): add inline keyboard buttons for C2C HITL approval

QQ Bot supports inline buttons via `markdown + keyboard` payloads. Clicks
arrive as `interaction_create` events through the existing botpy
WebSocket gateway — no extra subscription needed beyond enabling the
`interaction` intent. Group-scope clicks are out of scope here (DM only).

Send path
- `_build_qq_keyboard(buttons)` mirrors the Feishu helper, mapping the
  generic `{text, value, type}` shape to QQ's `{render_data, action}`
  with action.type=1 (callback). One button per row for mobile clarity.
- `_send_chunk` extracts `metadata["buttons"]` and threads a `keyboard`
  payload into `_post_markdown_message` for C2C only.
- Markdown→plain fallback can't carry a keyboard, so when buttons were
  attached the fallback content gets a textual `Reply: 1=Approve, …`
  hint built from the button list. `_parse_approval_reply` accepts
  the same values typed manually, so the user is never stuck.

Receive path
- `on_interaction_create` is registered on the bot class.
- `_on_interaction` extracts `data.resolved.button_data`, builds an
  InboundMessage, runs it through inbound middleware (Dedup suppresses
  retry callbacks), and publishes directly to the bus — bypassing the
  per-sender debounce buffer so the click value isn't merged with any
  text typed in the same window.
- Always ACKs via `api.on_interaction_result(id, 0)` in `finally` so
  QQ doesn't show the button as "expired", even if middleware drops
  the click or something throws downstream.

`QQ.inline_buttons=True`; `_approval_prompt_metadata` now auto-attaches
the Approve/Reject/Approve-all button row for QQ HITL prompts.

* fix(qq): button-value coercion, ACK timing, HITL consumer wiring

Fixes 6 bugs found in the inline-keyboard commit and consolidates the
button helpers so the keyboard builder, plain-text fallback hint, and
interaction handler share one coercion path.

- Plain-text fallback no longer crashes on non-string `value` (e.g.
  `{"text": "OK", "value": 42}`).  Extracted `_normalize_button` is now
  the single place that resolves `(label, value)` and coerces non-strings.
- `metadata["button_value"]` is the coerced string instead of the raw
  payload, matching `content` and downstream string comparisons.
- `_on_interaction` ACKs first, before publishing to the bus, so the
  QQ button UI never shows "expired" if middleware is slow.
- Wire `_approval_prompt_metadata` + `_format_approval_prompt(with_buttons=)`
  into `InboundConsumer._stream_with_hitl` and `cli.channel.channel_hitl_prompt`
  so the QQ `inline_buttons=True` capability is actually used end-to-end
  (HITL prompts auto-attach Approve/Reject/Approve-all buttons when the
  channel advertises the capability).
- Trim contradictory `_QQ_DEFAULT_PERMISSION` comment.
- Fix `test_group_interaction_ignored` docstring (ACK runs first now,
  not in `finally` after a `return`).

Tests: `_normalize_button` covered indirectly via existing keyboard tests;
new regressions for non-string fallback hint, ACK-on-handler-throw, and
string-coerced `button_value` metadata.

* refactor(qq): slim button helpers and explicit has_buttons flag

Inline single-use _button_hint and the _QQ_BUTTON_STYLE/_QQ_DEFAULT_PERMISSION
constants in qq/channel.py; tighten _on_interaction (drop unreachable
"[button click]" sentinel and unused triggering_message_id metadata; collapse
"if resolved else" ternaries via `or ""`).

Replace the metadata round-trip ("buttons" in metadata) used to detect button
support in consumer.py and cli/channel.py with an explicit has_buttons bool
threaded through both the prompt formatter and metadata builder.

Apply ruff format to the previously unformatted blocks introduced earlier on
this branch so CI lint passes.

* feat(qq): send post-decision confirmation after HITL approval

Send a visible confirmation message ("✅ 已批准" / "❌ 已拒绝") right after
the user resolves a HITL approval — QQ Bot has no message-recall or edit API
for C2C, so a follow-up message is the only way to give the click/reply
strong feedback.

Bus consumer (consumer.py): only sends the confirmation when the user
actually responded (event was set), to avoid pretending the user approved
when the request really timed out and auto-approved.

CLI HITL prompt (cli/channel.py): mirrors the same set of confirmation
strings.  Timeout and unrecognized-reply paths keep their existing English
text since their semantics differ (auto-reject vs auto-approve, plus a
hint about the unparsed input).
2026-05-11 10:44:20 +01:00
dependabot[bot] ccb3083183 chore(deps): bump langchain-core in the uv group across 1 directory (#219)
Bumps the uv group with 1 update in the / directory: [langchain-core](https://github.com/langchain-ai/langchain).


Updates `langchain-core` from 1.3.2 to 1.3.3
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain-core==1.3.2...langchain-core==1.3.3)

---
updated-dependencies:
- dependency-name: langchain-core
  dependency-version: 1.3.3
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-09 16:50:25 +01:00
X-iZhang f00ee1b5cc chore: update version to v0.1.0 2026-05-08 22:57:30 +01:00
Xi Zhang c407d2e20f Fix/async subagent model switch (#217)
* feat(middleware): add ConfigurableModelMiddleware for dynamic model resolution

- Introduced ConfigurableModelMiddleware to resolve chat models from RunnableConfig.configurable on each call.
- Updated middleware initialization to include ConfigurableModelMiddleware.
- Enhanced context editing middleware tests to verify presence of ConfigurableModelMiddleware.
- Implemented tests for ConfigurableModelMiddleware to ensure correct model overriding and caching behavior.
- Added tests for deepagents model-passthrough patch to verify configuration injection in async tasks.

* feat(async-subagent): update middleware handling to prevent deadlocks in async sub-agents

* style: Refactor code formatting for improved readability in patches and test files

* refactor: streamline middleware construction and improve async handling in ConfigurableModelMiddleware

* fix: remove unused request parameter from _read_model_override function

* refactor: improve async handling in _ClientProxy and enhance logging in ConfigurableModelMiddleware
test: add behavior test to ensure AskUserMiddleware is excluded in async subagent mode
2026-05-08 22:52:24 +01:00
Ziheng Zhang 89b0ecdbf3 feat(qq): add QR-code scan-to-configure onboarding for QQ Bot (#213)
* feat(qq): add QR-code scan-to-configure onboarding for QQ Bot

Adds a `qr_register()` flow that drives q.qq.com's create_bind_task /
poll_bind_result APIs so the wizard can auto-fill `qq_app_id` and
`qq_app_secret` after the developer scans a QR code with a bound QQ
account, falling back to manual entry on failure or cancel.

- channels/qq/crypto.py: AES-256-GCM helpers for decrypting the bot's
  client_secret returned by poll_bind_result.
- channels/qq/onboard.py: portal API client + polling loop.
- channels/qq/__init__.py: re-export `qr_register`.
- config/onboard.py: QQ branch in `_step_channels` that offers
  "Scan QR code" vs "Enter manually", and skips the manual prompt
  loop when a scan succeeded.

* style(qq): fix ruff lint errors in onboard.py

Move `import os` to the top-level import block (E402), drop the legacy
`typing.Optional`/`typing.Tuple` imports (UP035), and use the PEP 585/604
builtin generics (`tuple[...]`, `X | None`) for the few annotations that
still referenced them (UP006/UP045). No behavior change.

* fix(qq): harden QR onboard error paths and declare scan deps

Address review feedback on PR #213:
- Declare cryptography>=41.0 and qrcode>=7.4 in [qq]/[all-channels]
  extras and in _CHANNEL_PIP_DEPS so the scan flow no longer fails
  with an opaque ImportError on a fresh `evoscientist[qq]` install.
- Polling loop logs each _poll_bind_result failure and aborts after
  5 consecutive errors instead of silently spinning until the 600s
  timeout, restoring the documented Raises: RuntimeError contract.
- Wrap decrypt_secret in try/except so failures honor the
  None-on-failure contract instead of letting exceptions escape.
- Preflight `import cryptography` in the scan branch and offer
  install or fall back to manual entry.

* style: ruff format collapse two over-wrapped log/console lines
2026-05-08 12:08:52 +01:00
Xi Zhang 4e04ac5b72 fix: Improve watcher logic to prevent false-positive notifications on… (#216)
* fix: Improve watcher logic to prevent false-positive notifications on clean stream exits

* fix: Update watcher logic to drop notifications on persistent runs.get failures

* fix: Refactor test for watcher persistent failure notification handling

* fix: Enhance watcher test to validate all notification queues are empty after reconnect budget exhaustion
2026-05-07 22:16:36 +01:00
Xi Zhang 80f1f4fa0f feat: Implement async sub-agent auto-notification system (#214)
* feat: Implement async sub-agent auto-notification system

- Added async notifier functionality to handle notifications for sub-agents reaching terminal states.
- Introduced `AsyncTaskNotification` dataclass for structured notification data.
- Implemented `watch_run_and_notify` to monitor agent runs and enqueue notifications.
- Created `spawn_watcher` to manage watcher tasks and ensure proper cancellation of previous watchers.
- Developed `consume_notifications` to process notifications, deduplicate them, and format messages for LLM.
- Added tests for notification handling, including draining, deduplication, and formatting.
- Patched deepagents to integrate the new watcher functionality into start and update tools.

* Enhance async notifier with per-thread notification routing and error handling

- Introduced `origin_cli_thread_id` to `AsyncTaskNotification` for routing notifications back to the originating CLI session.
- Implemented per-thread notification queues to handle notifications based on the originating thread.
- Updated `has_pending_notifications` and `drain_notifications` to respect thread-specific queues.
- Enhanced `watch_run_and_notify` to detect in-band error events from the SSE stream and handle clean exits.
- Modified tests to verify the new notification routing behavior and ensure proper handling of notifications across threads.
- Added a fixture to restore the async watcher patch state in tests to prevent state leakage.
- Updated deepagents patching to capture the main agent's CLI thread ID for notification routing.

* feat: Enhance async notifier with thread-specific watcher management and notification filtering

* test: Enhance notification draining logic for cleaner test setup

* refactor: Remove summary field from AsyncTaskNotification and update related tests

* feat: Enhance async notification handling with target thread ID support

* Refactor async notifier and middleware for improved task management

- Removed the no-op shutdown watcher loop from async_notifier.py as it is no longer needed.
- Updated watch_run_and_notify to clarify notification handling and race conditions.
- Cleaned up shutdown handling in commands.py, interactive.py, and tui_interactive.py by removing obsolete shutdown watcher calls.
- Deleted the deepagents async watcher patch from patches.py, transitioning to a new middleware approach.
- Introduced AsyncWatcherMiddleware to handle async task notifications directly during tool calls.
- Updated tests to validate the new middleware functionality and ensure proper watcher spawning and cancellation.
- Enhanced test coverage for async watcher middleware, including edge cases and error handling.

* feat(tests): add fixture to reset notifier state before each test
2026-05-07 16:03:02 +01:00
Ziheng Zhang 692dc491ac # feat(wechat): add personal-WeChat (iLink) backend with QR login (#212)
* feat(wechat): add personal-WeChat (iLink) backend with QR-code login

Adds a third WeChat backend alongside WeCom and Official Account:
``personal`` rides Tencent's iLink Bot long-poll gateway so a personal
WeChat account can act as a bot.  Credentials are obtained via QR-code
scan and persisted under ``DATA_DIR/wechat_personal/accounts/``.

- channels/wechat/personal.py: WeixinPersonalChannel + qr_login.
- channels/wechat/crypto.py: aes128_ecb_decrypt + parse_ilink_aes_key
  for the iLink CDN media protocol.
- channels/wechat/probe.py: validate_wechat_personal credential probe.
- channels/wechat/serve.py: --backend personal CLI + --qr-login flow.
- channels/wechat/__init__.py: factory dispatch on wechat_backend; pull
  in the new dependencies in the docstring.
- config/settings.py: wechat_personal_* fields.
- config/onboard.py: WeChat-backend picker + QR-scan flow in the wizard
  + personal-backend probe in _probe_channel.
- pyproject.toml / uv.lock: add qrcode + certifi to wechat & all-channels
  extras (aiohttp was already pulled in transitively).

* fix(wechat): address ruff failures and CodeRabbit review on personal-WeChat PR

- personal.py: drop unused imports (`field`, `PollingMixin`); replace
  `asyncio.TimeoutError` with builtin; hold references to background
  `asyncio.create_task` results so they aren't GC'd; wire `dm_policy`
  through `_process_message` (disabled/allowlist) so `wechat_personal_dm_policy`
  actually takes effect for DMs.
- onboard.py: import-check gate now validates the full WeChat dependency
  set (aiohttp, qrcode, Crypto, certifi) instead of only aiohttp; mask
  `WeCom Secret` and `MP App Secret` prompts via `questionary.password`;
  derive the QR-login hint path from `_account_dir()` instead of the
  hard-coded `~/.evoscientist/...`; stop copying the QR-login token into
  the main config (already persisted per-account on disk — copying broadens
  secret exposure and risks staleness).
- pyproject.toml: allow Chinese full-width punctuation in `allowed-confusables`
  for user-facing CN messages.

* style(wechat): apply ruff format

`ruff format --check` was failing CI on three files (one pre-existing in
`__init__.py` plus formatter-driven line-merges in the files touched by
the previous fix commit). Ran `ruff format` to bring them in line; both
`ruff check` and `ruff format --check` now pass.
2026-05-07 16:33:30 +02:00
Wiktor Cupiał 9e51ec6fdd feat(cmd): add /model-fallback command (#196)
* feat(cmd): add /model-fallback command

* fix: apply feedback

* fix: lock usage with _fallback_chain

* fix: apply feedback

* fix: apply feedback

* feat: add tests

* fix: tests

* Update EvoScientist/middleware/model_fallback.py

Co-authored-by: dinos <dinospk1999@gmail.com>

---------

Co-authored-by: dinos <dinospk1999@gmail.com>
2026-05-07 15:35:26 +02:00
Ziheng Zhang 22a65b640d refactor(channels): remove dead MessageBus dispatcher (#205)
* refactor(channels): remove dead MessageBus dispatcher

Outbound routing has two implementations: ``MessageBus.dispatch_outbound``
(subscriber-based) and ``ChannelManager._dispatch_outbound`` (registry
lookup).  Only the latter is ever started in production — the former
is reachable solely from tests, yet both consume from the same
``bus.outbound`` queue.  If anyone followed the bus's own API surface
they would silently steal messages from the real dispatcher.

Drop the unused machinery to leave a single, obvious outbound path:
- ``MessageBus.subscribe_outbound`` / ``dispatch_outbound`` / ``stop``
- ``_running`` flag and ``_outbound_subscribers`` map
- ``OutboundCallback`` type alias
- The lone ``bus.stop()`` call in ``cli/channel.py`` (was no-op)
- Four tests covering the removed code paths

* test(channels): drop empty MessageBus stubs after dispatcher removal

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-05-07 12:11:06 +01:00
Xi Zhang d7c0eec0e9 chore: update dependencies and remove unused OpenRouter patch (#211) 2026-05-05 18:51:00 +01:00
Xi Zhang f41584e10b Refactor sub-agent architecture and introduce async support (#200)
* Refactor sub-agent architecture and introduce async support

- Removed the legacy subagent.yaml file and replaced it with individual YAML files for each sub-agent in the subagents directory.
- Updated the load_subagents function to support both directory and single file layouts for loading sub-agent configurations.
- Added new langgraph_dev module for managing async sub-agent lifecycle and deployment.
- Created graphs for async sub-agents (writing-agent, data-analysis-agent) and updated langgraph.json for deployment.
- Introduced new sub-agent definitions for planner, research, debug, code, and writing agents with appropriate system prompts and configurations.
- Enhanced package data inclusion in pyproject.toml to accommodate new sub-agent YAML files.

* Refactor code for improved readability by consolidating conditional statements and formatting

* feat: enhance async sub-agent support with workspace synchronization and user feedback

- Added console status messages during async sub-agent server startup and workspace synchronization to improve user experience.
- Implemented a new WorkspaceSyncWidget for live feedback during workspace sync operations.
- Updated onboarding to reject occupied ports and ensure proper workspace handling for async sub-agents.
- Introduced locking mechanisms to manage concurrent access to langgraph dev processes and workspace states.

* feat: add async sub-agent configuration and server management functions

* feat: improve port occupation handling and log file management in start_langgraph_dev

* feat: enhance async sub-agent handling and introduce comprehensive tests

- Updated `_maybe_swap_async_subagents` to improve async sub-agent management, ensuring internal flags are stripped before handoff.
- Enhanced port management in `onboard.py` to allow reuse of occupied ports if already running by the same service.
- Introduced file locking in `manager.py` to prevent race conditions during concurrent CLI invocations.
- Added new tests for async sub-agent swapping and langgraph manager functionalities to ensure reliability and correctness.
- Updated dependencies in `pyproject.toml` to include `psutil` and `filelock`.

* fix(docs): clarify sub-agent configuration in README

* test(manager): isolate _PID_DIR + tighten reuse-path assertion

Addresses CodeRabbit review on tests/test_langgraph_manager.py:

- Patch _PID_DIR to tmp_path so the FileLock setup in
  ensure_langgraph_dev doesn't mkdir the user's real
  ~/.config/evoscientist/ dir as a test side-effect.
- Tighten "result is None or hasattr(result, 'poll')" to a strict
  "result is None" — the reuse path returns None unconditionally,
  so the OR clause was hiding potential regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(manager): clean up stale PID file when unrelated process reuses PID

* feat(tests): add validation tests for async flag in load_subagents

* fix(load_subagents): restrict to .yaml files and clarify configuration handling

* fix(load_subagents): improve error handling for non-dict specifications in YAML

* feat(onboard): add "LangGraph Port" step to onboarding process

* feat(langgraph): add concurrency configuration for langgraph dev workers

* feat(async-subagents): enhance MCP tool routing for async sub-agents

* fix(manager): update exception handling for connection errors and prevent zombie processes

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 12:45:50 +01:00
X-iZhang d5ffa00263 docs(README): enhance Docker section with installation details and usage instructions 2026-05-04 15:37:05 +01:00
X-iZhang a297445f92 fix(assets): update wechat_group image file 2026-05-04 10:51:33 +01:00
Xi Zhang ab3787f86a fix(markdown): ensure proper spacing for ATX headings in Markdown ren… (#201)
* fix(markdown): ensure proper spacing for ATX headings in Markdown rendering

* fix(markdown): improve docstrings and tests for heading spacing functionality
2026-05-02 00:59:30 +01:00
Xi Zhang 29e7fd383d fix/cli hitl (#202)
* feat(display): enhance approval prompt with questionary for better navigation

* feat(cli): add HumanInTheLoopMiddleware for user approval in main agent
2026-05-02 00:54:40 +01:00
dinos 0d8ac4f24b feat(docker): official image with all runtime deps pre-installed (#198)
* feat(docker): official image with all runtime deps pre-installed

Multi-stage build using uv for the EvoScientist core + all messaging-channel
extras, plus Node.js 24 LTS (for npx-based MCP servers) and uv (for runtime
Python MCP installs) in the runtime layer. Runs as non-root user evosci,
with workspace, app data, and config (XDG_CONFIG_HOME) all consolidated
under a single /home/evosci/.evoscientist volume so a single mount
persists everything across container restarts.

Includes a docker-compose.yml starter, a build/push GitHub Actions
workflow targeting ghcr.io with multi-arch (amd64/arm64) and PR-only
build verification, a .dockerignore, and a new Docker section in the
README documenting mounts, derivation recipes for the unbundled stt /
oauth / TinyTeX extras, and proxy/cert handling expectations.

* fix(docker): pin trixie base + drop redundant python image

Switch builder and runtime from `python:3.11-slim-bookworm` to a single
`ghcr.io/astral-sh/uv:python3.11-trixie-slim` base — trixie drops several
CRITICAL vulnerabilities that bookworm carries today, and reusing the uv
image for runtime eliminates the separate `COPY --from=…/uv` line.

* chore(docker): pin GitHub Actions to commit SHAs in workflow

Replace mutable major-version tags with full commit SHAs (with the
corresponding semver tag in a trailing comment) so a compromised /
retagged action release can't silently change what runs in the publish
pipeline.

* chore(deps): enable Dependabot version updates for Dockerfile pins

Adds a weekly `docker` ecosystem that watches the Dockerfile's `FROM` /
`COPY --from=` references — including the ARG-bound `BASE_IMAGE` and
`NODE_IMAGE` digests — and opens one grouped PR per cadence bumping
both the @sha256 digest and the trailing version comment. This keeps
the otherwise-frozen pins flowing with Debian point releases and
upstream patches.

* fix(docker): use nodejs alias stage so NODE_IMAGE ARG actually resolves

`COPY --from=${NODE_IMAGE}` left the dollar-curly literal at parse time
under buildkit 29.x — it expands ARGs in `FROM` but reads `--from=` as a
static stage/image name. Introduce a tiny `FROM ${NODE_IMAGE} AS nodejs`
alias and `COPY --from=nodejs …` against it, which preserves the
ARG-driven Dependabot updates without tripping the parser.

* fix(docker): harden venv ownership and PATH ordering

- Drop `--chown` on the `/opt/venv` COPY so the venv stays root-owned.
  The runtime user only needs read+execute (default Unix perms allow
  that); making it user-owned let the agent rewrite its own
  dependencies, which defeats the sandboxing premise. All persistent
  agent state already lives under /home/evosci/.evoscientist/.
- Reorder PATH so /opt/venv/bin precedes the user-writable
  UV_TOOL_BIN_DIR. Otherwise a stray binary dropped into the latter
  (e.g. via `uv tool install`) could shadow the canonical
  `evosci` / `python` / `pip` shipped with the image.

* docs: update README

* docs(docker): warn about non-root UID and `curl | sh` for derived images

- The image runs as `evosci` (UID 1000), so a host-side `./workspace`
  bind mount fails if the host user has a different UID — same gotcha
  that bites onboarding's `mcp.yaml` write. Add an !IMPORTANT block
  with the two practical fixes (`chown -R 1000:1000` once, or
  `--user "$(id -u):$(id -g)"` on each run).
- The TinyTeX derivation snippet pipes an unpinned remote installer
  into `sh`. Add a one-line pointer to fetching a pinned release
  tarball from `rstudio/tinytex-releases` for users who'd rather not
  trust the upstream script blindly. The official installer is kept
  as the default since that's what TinyTeX itself recommends.

* chore(docker): cancel in-flight workflow runs + flag iMessage as host-only

- Add `concurrency: cancel-in-progress: true` to the docker workflow
  so successive pushes on the same ref supersede the prior run rather
  than queueing in parallel — multi-arch buildx is the slowest job in
  CI, no point burning minutes on superseded builds.
- Spell out that the docker image installs the `all-chanels` extra and
  call out iMessage as a deliberate host-only exclusion: it requires
  the `imsg` CLI bridging to macOS's Messages.app, which no Linux
  container config can satisfy.
2026-05-01 13:24:14 +02:00
dinos 73928a2d78 fix(onboard): detect installed skill packs via install manifest (#199)
* fix(onboard): detect installed skill packs via install manifest

Onboarding's _step_skills only inspected USER_SKILLS_DIR and matched
recommended entries by directory-name hint, so a pack like
EvoScientist/EvoSkills@skills (which explodes into paper-writing/,
evo-memory/, etc. under GLOBAL_SKILLS_DIR) was never detected and kept
appearing as not-yet-installed.

skills_manager now writes a per-tier .installed.yaml mapping skill
directory name -> original install source on every install, removes the
entry on uninstall, and exposes installed_sources(). _step_skills checks
both tiers and treats a recommended source as installed when present in
any manifest -- so packs are recognized regardless of how their child
dirs are named.

* fix(onboard): write install manifest atomically

Stage to a sibling temp file, fsync, then os.replace into place. A crash
mid-write can no longer leave a half-written .installed.yaml behind,
which would otherwise wipe out pack detection until the next reinstall.

* style: fmt

* fix(onboard): catch decode errors when loading install manifest

read_text() can raise UnicodeDecodeError on a hand-edited or corrupt .installed.yaml; pin encoding="utf-8" and add UnicodeError to the except clause so the function honors its "returns {} on any error" contract.
2026-05-01 13:23:59 +02:00
X-iZhang 27a97c3097 Refactor code structure for improved readability and maintainability 2026-04-30 22:00:19 +01:00
dinos 5c829942d7 refactor(cli): replace channel module globals with ChannelRuntime (#197)
* refactor(cli): replace channel module globals with ChannelRuntime

Removes _cli_agent / _cli_thread_id from EvoScientist/cli/channel.py
and threads a ChannelRuntime via CommandContext.channel_runtime so
/model and /channel rebind without poking module-level state.

* fix(cli): address coderabbit review

- _auto_start_channel: bind ChannelRuntime only after
  _start_channels_bus_mode succeeds, so a startup failure no longer
  leaves a stale binding pointing at channels that never started.
- _sync_tui_command_completion (TUI) and the Rich CLI command-completion
  paths: rebind the runtime on thread rotation, not just agent swap, so
  /new and /resume keep ChannelRuntime in sync with the running thread
  (matches the serve-mode hook contract).
- test_hook_syncs_channel_runtime: pin ctx.thread_id explicitly so a
  bare MagicMock attribute can't silently mutate runtime.thread_id.
- New regression test covering the rebind-on-thread-rotation contract.
2026-04-30 18:06:22 +01:00
Xi Zhang 7cfec02416 Fix/sessions migration sweep race (#195)
* fix(sessions): ensure migration sweep runs before yielding checkpointer to prevent race conditions

* fix(sessions): enhance migration sweep with progress indication and ETA estimation
2026-04-29 02:06:25 +01:00
Xi Zhang 50719ef256 Implement PruningCheckpointer for efficient checkpoint management and… (#194)
* Implement PruningCheckpointer for efficient checkpoint management and add comprehensive tests

- Introduced `PruningCheckpointer` to manage checkpoint pruning after each `aput()`, ensuring only the latest checkpoints are retained based on a configurable limit.
- Added migration sweep functionality to clean up legacy checkpoints and prevent database bloat.
- Enhanced `get_checkpointer()` to utilize the new `PruningCheckpointer` and trigger migration sweeps when necessary.
- Developed a suite of integration tests for `PruningCheckpointer`, covering various scenarios including pruning behavior, concurrent writes, and retention policies.
- Implemented tests for migration sweep functionality, ensuring proper partitioning and user version management.
- Added diagnostic helper `db_stats` to provide insights into the database state, including thread and checkpoint counts.

* feat(sessions): enhance pruning logic to handle legacy DBs without writes table

* fix(tests): prevent atexit hook leakage in TestMigrationSweep

* feat(tests): enhance TestPruningCheckpointer to validate put+prune serialization

* feat(tests): refactor mock path implementation for get_db_path in test cases
2026-04-28 22:26:59 +02:00
Xi Zhang 56cc2fef85 fix(deepseek): add empty-string fallback for reasoning_content in cross-provider scenarios (#192) 2026-04-27 19:19:50 +01:00
Xi Zhang 52f8d3a3a5 Feat/llm context window patch table (#191)
* feat(context-window): add model context window patch table and apply function

* fix(tests): clean up formatting in context window tests

* fix(tests): update context window tests for Claude model exceptions
2026-04-27 18:28:57 +01:00
X-iZhang 26e3452ef6 chore: update version to v0.0.9 in badges, README, and project files 2026-04-26 15:16:12 +01:00
Xi Zhang 20c06d4897 feat(deepseek): implement reasoning_content passback for multi-turn s… (#190)
* feat(deepseek): implement reasoning_content passback for multi-turn scenarios

* fix(tests): ensure consistent import of EvoScientist.llm.patches in test cases

* fix(tests): streamline tool_calls formatting in TestPatchDeepseekReasoningPassback

* feat(patches): add reasoning_content capture and re-injection for DeepSeek assistant messages

* fix(deepseek): optimize reasoning_content extraction and assignment in passback

* fix(deepseek): refine reasoning_content handling in OpenAI capture patch
2026-04-26 14:53:08 +01:00
Xi Zhang e48bc1cb71 Refactor/system prompt structure (#189)
* feat: Refactor system prompt structure and enhance documentation for clarity

* refactor: Improve clarity and consistency in prompt documentation

* refactor(tests): Improve readability of first-person avoidance test assertion

* refactor: Remove redundant datetime imports and enhance prompt documentation
2026-04-26 11:08:08 +01:00
Ziheng Zhang da74c325d6 fix(channel): scope stop and restore resume history (#186)
* fix(channel): scope stop and restore resume history

* refactor(channel): simplify stop and resume patch

* Delete PR_MESSAGE.md

* fix(channel): address review feedback

* fix(channel): address remaining review bugs

* fix(channel): clean up stopped request handling

* fix(channel): preserve resolved replies and sync tui commands

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-25 16:12:49 +01:00
X-iZhang 27fee3256e fix(clipboard): remove unnecessary check for selection end in copy_selection_to_clipboard 2026-04-25 13:04:19 +01:00
Xi Zhang 558360b558 feat: Enhance ModelPickerWidget for Ollama integration (#187)
* feat: Enhance ModelPickerWidget for Ollama integration

- Implemented a sentinel row for "Custom Ollama model..." in ModelPickerWidget, allowing users to input arbitrary model names.
- Updated action handling in ModelPickerWidget to manage transitions between list and input modes.
- Added async model discovery for Ollama models, integrating with the /model command to fetch locally installed models.
- Created tests for Ollama model discovery and ModelPickerWidget behavior, ensuring proper functionality and user experience.
- Refactored validate_ollama_connection and discover_ollama_models for improved error handling and response management.

* fix: Simplify code by removing unnecessary line breaks in ModelPickerWidget and test cases

* fix: Restore globals on set_chat_model failure to prevent half-switched session

* fix: Improve error handling in ModelCommand by restoring globals on failure
2026-04-25 00:37:48 +01:00
Xi Zhang 49b03c36eb feat: add support for gpt-5.5 model in the LLM configuration and tests (#188) 2026-04-24 23:52:55 +01:00
Wiktor Cupiał 155c4eaa40 fix: text copy on remote sessions/legacy terminal emulators (#185)
* fix: text copy on remote sessions/legacy terminal emulators

* fix: display warning only once
2026-04-24 23:16:39 +02:00
Xi Zhang c134a16e19 Fix/channel slash rich cli (#184)
* feat(cli): implement slash command dispatch for channel messages

* fix(cli): make EvoSci serve exit on Ctrl+C and hot-swap /model (#181)

* fix(cli): streamline debug logging and formatting in channel command handling

* fix(cli): ensure proper handling of asyncio event loop in slash command processing

* fix(cli): add error handling for unexpected exceptions in slash command dispatch

* fix(cli): improve error messaging for slash command dispatch failures

* fix(tests): enhance test setup by restoring channel globals and simplifying assertions

* fix(cli): enhance slash command handling across UI surfaces and improve resume command warnings
2026-04-24 16:10:25 +01:00
Xi Zhang 3831198f19 fix(chat): resolve model switch lag by tracking model/provider key (#180)
* fix(chat): resolve model switch lag by tracking model/provider key for cache invalidation

* style: format set_chat_model function for improved readability

* fix(chat): improve model switching logic to prevent unnecessary cache rebuilds

* fix(model): ensure globals are restored on agent load failure to prevent model switch issues
2026-04-24 11:36:33 +01:00
Xi Zhang 92df3e1844 Refactor/cli command manager (#178)
* feat(cli): migrate command handling to CommandManager and enhance UI interactions

* feat(cli): add /clear and /help commands to enhance user experience

* feat(cli): implement /new and /resume commands with interactive session management

* Refactor MCP and Skills Command Handling

- Moved the interactive picker style to a centralized widget for consistency across MCP and Skills commands.
- Updated the MCP command to remove the old command dispatch logic, delegating to the new InstallMCPCommand.
- Enhanced the Skills command to utilize a new interactive picker for skill selection, improving user experience.
- Implemented cancellation handling in the picker to differentiate between user cancellations and empty selections.
- Added comprehensive tests for the new command structures and picker functionalities to ensure reliability.

* refactor(cli): streamline CommandManager dispatch and remove deprecated command set

* refactor(cli): enhance error handling and state management in ChannelCommand and RichCLICommandUI

* refactor(cli): update lifecycle callback terminology and improve async prompt handling in RichCLICommandUI

* refactor(cli): enhance SlashCommandCompleter to dynamically fetch workspace directory for autocompletion

* refactor(cli): unify quit handling in RichCLICommandUI with shared _stop helper

* refactor(cli): remove hardcoded slash commands and utilize command manager for dynamic completion

* refactor(cli): update MCP and skills command files for improved clarity and organization

* refactor(mcp_ui): remove unnecessary newline in _show_mcp_config function
2026-04-24 10:38:44 +01:00
X-iZhang f704e3d761 update 2026-04-23 14:33:02 +01:00
dinos 65828e9666 perf(cli): cut startup latency and defer MCP loading to the background (#171)
* perf(cli): cut startup time of `evosci --help` from ~2.2s to ~0.3s

Module-level imports were eagerly pulling in langchain.chat_models (with
the whole anthropic/openai/google stack), langgraph, textual, and
prompt_toolkit on every invocation — even for `--help` or `config list`.

Defer those with PEP 562 `__getattr__`, using `lazy_loader.attach` (SPEC-1,
the scientific-python standard) where it's a clean attach pattern:

- `EvoScientist/llm/__init__.py`: attach `models` lazily so importing
  `context_window` from this package no longer drags in langchain.
- `EvoScientist/stream/__init__.py`: attach display/events lazily; split
  the shared Rich `Console` singleton into a new lightweight
  `stream/console.py` so callers that only need `console` skip the
  `stream.events` → `langchain_core.messages` chain.
- `EvoScientist/cli/__init__.py`: hand-rolled `__getattr__` (reaches into
  `..stream.state`, which `lazy_loader` doesn't cover) so `commands` and
  `app` are the only eager loads.
- `EvoScientist/cli/commands.py`: move `cmd_interactive`/`cmd_run` to
  in-function imports so prompt_toolkit + textual only load when the
  interactive path actually runs.
- `EvoScientist/cli/_constants.py`: read `AGENT_NAME` on demand so
  `build_metadata` doesn't eagerly import `sessions` (langgraph/aiosqlite).

Adds `lazy-loader>=0.5` as a dependency.

* feat(cli): defer MCP tool loading with live per-server progress

The CLI was blocking ~5 s on MCP tool enumeration before the first
prompt appeared. Move the agent construction off the event loop and
surface per-server progress so the user can interact immediately and see
what's happening.

MCP client:
- Add an `on_progress` callback to `load_mcp_tools` / `aload_mcp_tools`
  / `_load_tools` emitting `start` / `success` / `error` events per
  server.
- Fan connection attempts out with `asyncio.gather` so latency no longer
  scales linearly with server count; cap simultaneous attempts at
  `_MAX_CONCURRENT_CONNECTIONS` (8) via a semaphore so a big stdio fleet
  doesn't spawn every subprocess at once.

Agent wiring:
- Plumb `on_mcp_progress` through `create_cli_agent` / `_load_agent` /
  `load_mcp_and_build_kwargs` so CLI and TUI can plug in collectors.

CLI (`cmd_interactive`):
- Run `_load_agent` in a background thread via `asyncio.to_thread`; the
  prompt and banner render immediately.
- `_await_agent_ready()` awaits the task before each agent-using site
  (first turn, channel messages, `/channel`, `/compact`). Raises if
  called without a prior `_start_agent_load` instead of silently
  reloading without the SQLite checkpointer.
- Pre-prime the progress dict from `load_mcp_config()` so the
  bottom-toolbar's `N/M` denominator is stable from the first render.
- Wrap `session.prompt_async` in `patch_stdout(raw=True)` so
  `console.print` from the worker-thread progress callback lands cleanly
  above the prompt as inline chat messages instead of stomping the
  prompt cursor.

TUI (`EvoTextualInteractiveApp`):
- Same background load + `_await_agent_ready()` gates on every
  `self._agent` read.
- New `MCPLoaderWidget` mounted at the top of `#input-shell` shows a
  header with `N/M` and one live row per server (spinner → ✓ / ✗ with
  tool count or error detail). On completion:
  - all-clean loads auto-dismiss ~2.5 s later;
  - cache hits (no events ever fired) dismiss immediately rather than
    flashing a misleading "0/N loaded";
  - failures keep the widget mounted so the user can read the errors.
  - `dismissed` property lets the app clear its ref so late events from
    slow servers become no-ops. The error branch of `_on_agent_loaded`
    also settles the widget so a load failure can't leave the spinner
    animating forever.
- Chat input is `disabled` while MCP resolves — no placeholder hack, no
"waiting…" system message.

Shared:
- Hoist braille spinner frames to `status_bar.SPINNER_FRAMES` and import
  them in the TUI widget so CLI and TUI animate in sync.

Tests:
- Extend `test_agent_mcp_cache` fakes to accept the new `on_progress`
  kwarg.
- New `TestLoadToolsProgressCallback` in `test_mcp_client` exercises the
  event sequence for success/failure/mixed fleets, verifies a buggy
  callback doesn't break the load, and asserts the semaphore caps
  in-flight connections.

* style: ruff

* chore: update uv.lock

* chore: uv.lock

* fix: coderabbit issues

* style: fmt

* fix: move _await_agent_ready inside try block

* fix(tui): auto-dismiss MCP loader widget on failure

The widget was designed to stay mounted on failure so the user could
read error detail, but since it's pinned above the input it never went
away in practice — just permanent banner clutter.

Auto-dismiss on failure too, with a longer grace (12s vs 2.5s) so the
error summary stays readable.

* fix: address second coderabbit pass

- Channel handlers (CLI + TUI): catch agent-load failures so the
  channel request doesn't hang; CLI moves `_await_agent_ready()`
  inside the existing try/except, TUI catches explicitly and calls
  `_set_channel_response` with the error.

- Stale background loads: `prev.cancel()` only stops the asyncio
  wrapper, not the thread running `_load_agent`. Added a generation
  token (`agent_load_id` / `self._agent_load_id`) and gated both
  progress and completion callbacks on it so a superseded load can't
  clobber the current session's state or UI.

- TUI prompt lifecycle: added `_agent_load_pending()` and gated the
  `_process_channel_message` / `_handle_command` finally blocks on it
  so `/new` or `/resume` invoked from a command keeps the prompt
  disabled until the fresh load settles.

- TUI readiness failures: `_run_turn` and `_handle_command` now
  catch exceptions from `_await_agent_ready()` and surface a
  "Agent failed to load: …" system message instead of letting the
  exception escape into Textual's traceback panel.

* refactor(cli): share background agent loader between CLI and TUI

The CLI and TUI were carrying near-identical copies of the same
background-load state machine: the `agent_task`, the `agent_load_id`
generation token, the gated progress/completion callbacks, and the
per-server progress dict. Every CodeRabbit finding on that lifecycle
had to be fixed in both files.

Extract it into `cli/_agent_loader.py`:

- `MCPProgressTracker` — owns the `server -> (state, detail)` dict;
  exposes `prime`, `record`, `snapshot`, `totals`.

- `BackgroundAgentLoader` — owns `agent`, the in-flight task, and the
  generation token. Exposes `start(**loader_kwargs)`, `await_ready()`,
  `is_pending`. Internally gates all progress/completion callbacks by
  generation so a superseded load can't clobber the current session.
  UI-specific rendering plugs in via `on_progress` / `on_success` /
  `on_failure` callbacks.

Both surfaces now just wire their UI hooks; the loader file holds no
Rich / prompt_toolkit / Textual dependencies. Net -345 lines from
`interactive.py` + `tui_interactive.py`; +20 unit tests pinning the
lifecycle (generation filtering, cache-hit short-circuit, failure
reset, progress ordering).

* refactor(cli): make _on_done the sole authority for agent state transitions

await_ready no longer sets self.agent — it just awaits the task and
reads what _on_done already wrote. Eliminates the dual-write overlap
(asyncio guarantees done-callbacks fire in registration order).

* fix(tui): let users type during MCP load, only block on send

Remove prompt-disabling during background agent load — the TUI now
matches the CLI approach where the input stays enabled and only gates
on await_ready() at submit time. The MCPLoaderWidget still provides
visual feedback that loading is in progress.

* fix(loader): preserve real load error on await_ready; dedup failure message

CodeRabbit flagged two issues with the new loader:

1. After a failed load, `_on_done` nulled `self._task`, so the next
   `await_ready()` hit the "before start()" branch and the CLI wrapper
   remapped it to a misleading "checkpointer not available" message —
   losing the real exception (bad MCP config, network, etc.).

   Keep `_task` set on failure so `await_ready` re-raises the real
   exception. Added `needs_restart` so TUI's auto-retry check stays a
   one-liner and doesn't need to reach into task internals.

2. TUI reported each load failure twice: once from
   `_on_agent_load_failure` (the done-callback) and once from each
   caller of `_await_agent_ready` (`_run_turn`,
   `_process_channel_message`, `_handle_command`) catching the re-raise.

   `_on_agent_load_failure` is now the sole local reporter; callers
   just handle control flow (return cleanly, set channel response to
   unblock remote).

* fix(cli): wire /model handler through the agent loader

The /model command from main (merged via f1f0d7c) still reached for
`state["agent"]` (CLI) and `self._agent` (TUI) — both removed by the
background-loader refactor. CLI raised KeyError on first invocation;
TUI raised AttributeError. Writes to the old fields also had no effect
because every other code path now reads from `agent_loader.agent`, so
the model switch would have silently failed.

Route everything through the loader: `await _await_agent_ready()` up
front so /model doesn't race with the initial background load, build
the `CommandContext` with the current agent, and sync `ctx.agent` back
into `agent_loader.agent` (plus channel globals) when the command
replaces it.

* fix(cli): isolate progress callback, capture awaited agent, gate by requires_agent

Three CodeRabbit findings on the loader + command dispatch path:

- Wrap ``_on_progress`` in try/except inside the loader's gated wrapper
  so a buggy UI adapter can't bubble into ``loader_fn`` and fail the
  whole background load. The MCP client already protects this, but
  defence-in-depth keeps the loader self-contained.

- In CLI channel + main-loop streaming, capture the agent returned by
  ``_await_agent_ready()`` and pass that into ``run_streaming`` rather
  than reading ``agent_loader.agent`` after a subsequent ``await``.
  A concurrent ``/new``/``/resume``/``/model`` could have swapped it.

- Add ``requires_agent: ClassVar[bool] = False`` to ``Command`` and
  mark ``/compact``, ``/model``, ``/channel`` as ``True``. TUI dispatch
  sites (channel and keyboard) now check ``cmd_manager.resolve(...)``
  and only wait for readiness when the command actually needs the
  agent. ``/mcp add``, ``/skills``, ``/new`` etc. no longer deadlock
  behind a failing MCP load they are meant to fix.

* fix(cli): guard sync-back, subcommand-aware gating, /model adopt-path

Three CodeRabbit findings on command dispatch:

- ``_handle_command`` unconditionally synced ``ctx.agent`` back into
  ``agent_loader``.  For non-agent commands ``ctx.agent`` is ``None``,
  so ``/threads`` / ``/mcp`` / ``/skills`` (etc.) could clobber a valid
  loaded agent — and rebind channel globals to ``None``.  Guard the
  sync on ``ctx.agent is not None``.

- ``/channel status`` and ``/channel stop`` don't touch ``ctx.agent``
  but the class-level ``requires_agent = True`` blocked them behind
  agent readiness.  Added ``Command.needs_agent(args)`` (defaults to
  ``requires_agent``) so ``/channel`` can override with subcommand
  awareness; kept the class flag for the common case.

- ``/model`` builds a new agent from scratch, it never reads the
  existing one — gating it on readiness meant a broken provider
  blocked the command that would fix it.  Flipped it to
  ``requires_agent = False`` and added ``BackgroundAgentLoader.adopt``
  so the UI can seat the replacement and supersede any in-flight
  load (the generation token keeps a late completion from clobbering
  the adopted agent).

Bonus cleanup: ``CommandManager.resolve`` now returns
``(command, args)`` so callers can invoke ``needs_agent`` without
re-implementing ``shlex`` parsing.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-22 18:22:33 +02:00
Wiktor Cupiał 28f3e81b4c feat: add /model command for changing models inside TUI/CLI (#162)
* rebase main

* fix: apply pr comments

* fix: fix critical issue

* fix: ordering /model in cli mode

* feat: refactor /model command handling and add Rich CLI support

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-04-22 14:43:21 +01:00
Xi Zhang aa3dd00409 feat: add support for session resumption with --resume flag and enhan… (#170)
* feat: add support for session resumption with --resume flag and enhance thread ID resolution

* refactor(tests): streamline help output testing for --resume flag

* feat: enhance session resume functionality with improved thread ID resolution and SQL wildcard handling

* feat: improve error handling for resume hint retrieval in interactive modes

* refactor: streamline logging for print_resume_hint failure in interactive mode

* feat: implement deferred scrolling for Markdown-heavy content in interactive mode
2026-04-21 21:26:40 +01:00
dinos 05f54334ba fix(mcp): stdio env passthrough + durable package installs (#169)
* fix(mcp): forward proxy and CA bundle env vars to stdio subprocesses

The MCP SDK's stdio transport inherits only a minimal allowlist (HOME,
PATH, USER, …) from the parent, stripping http_proxy/https_proxy and
SSL_CERT_FILE/REQUESTS_CA_BUNDLE/etc. Behind a proxy or with a custom CA
bundle, stdio MCP servers silently hang on outbound requests while the
same server over HTTP transport works. Auto-forward the proxy and cert
vars when present; user-configured env still takes precedence.

* fix(mcp): use `uv tool install` so MCP packages survive uv sync

Source installs previously used `uv pip install --python $VENV <pkg>`,
which lands in the evosci venv but is not recorded in pyproject.toml or
uv.lock. A subsequent `uv sync` (typical after `git pull`) reconciles
the venv to the lockfile and removes the MCP package, forcing users to
re-run onboard.

Prefer `uv tool install <pkg>` for the non-uv-tool install path: the
binary symlink in ~/.local/bin survives uv sync and evosci upgrades,
and the MCP server gets its own isolated env (no dep conflicts).
Verify the expected CLI entry point resolves afterward; if not (package
has no console-script), fall through to the old uv-pip path so
command-less packages still work.

The uv-tool-env path (`uv tool install evoscientist --with <pkg>`) is
unchanged — it was already durable via uv's receipt.

* fix(mcp): gate standalone uv tool install on verify_command

Previously `install_pip_package` would route every install through
`uv tool install <pkg>` when `verify_command` was None, returning
success as long as the uv subprocess exited 0. Library callers
(`evoscientist[oauth]`, `lark-oapi`, etc.) expect the package to land
in the active venv so they can import it — a standalone uv tool env
is not importable, so the import fails at the next line.

Gate the `uv tool install <pkg>` branch on `verify_command` being
set: that signals the caller wants a durable CLI binary, which is
what `uv tool install` produces. Library callers omit it and go
straight to the pip-install-into-venv path.

Also: log info messages on every fall-through so stale-binary and
entry-point-missing failure modes are debuggable, and document the
--with → standalone recovery path.

* fix(mcp): resolve MCP binaries to `uv tool dir --bin`, not `.venv/bin`

Under `uv run`, the project venv's `bin/` comes first on PATH, so
`shutil.which("arxiv-mcp-server")` returns a stale `.venv/bin/` copy
left over from an earlier install instead of the fresh symlink that
`uv tool install` just placed in `~/.local/bin`. The venv copy gets
written to mcp.yaml and is then wiped by the next `uv sync` — exactly
the failure mode the durability fix was meant to prevent.

Query `uv tool dir --bin` directly and prefer binaries found there
over `shutil.which`. Same change to the post-install verify in
`install_pip_package` so a venv shadow can't falsely short-circuit
the fallback.

* refactor(mcp): split install_pip_package into install_library + install_cli_tool

`verify_command` was doing double duty: naming the CLI binary to check
*and* signaling "this is a CLI install, use the standalone `uv tool
install` path." Callers routed library installs through the CLI branch
any time they forgot to pass it, and the resulting standalone uv tool
env wasn't importable from the active venv.

Separate the two use cases into distinct functions, each with one
install strategy per environment shape. Shared logic lives in private
`_install_with_uv_tool_env` / `_install_via_pip` helpers.

- install_library(pkg): uv-tool-env --with → pip. Never uses standalone
  `uv tool install <pkg>` (not importable from active venv).
- install_cli_tool(pkg, *, verify_command): uv-tool-env --with →
  standalone `uv tool install` → pip. `verify_command` is now required.

Callers pick the right function at the call site: registry.py picks
based on whether `entry.command` is set; onboard.py call sites all
install libraries.
2026-04-21 16:59:06 +01:00
X-iZhang 1e4c011b7c fix(assets): update wechat_group image for improved clarity 2026-04-21 12:06:49 +01:00
Xi Zhang 06822f236c feat: enhance tool result handling with tool_call_id for concurrent execution 2026-04-19 23:15:20 +01:00
Ziheng Zhang bd501cce34 fix(channel/qq): deliver HITL approval prompts reliably (#166)
* fix(channel/qq): deliver HITL approval prompts reliably

QQ approval prompts were silently dropped when the markdown send hit
a QQ server-side error (e.g. template not configured, content audit)
because the fallback path only matched TypeError / specific string
patterns, and the plain-text retry reused the already-consumed
msg_seq which QQ then rejects as duplicate.

- Consume a fresh msg_seq for the plain-text fallback send
- Recognize QQ server error codes (304014/304023/304003/40034059)
  and CN fragments ("模版"/"审核") as markdown-fallback triggers
- Promote send failure logs from debug to warning/error with
  chat_id/msg_id/seq so real-world errors can be diagnosed
- Extend test_qq_channel with a server-error-code fallback case

* style(channel/qq): apply ruff formatter to approval-delivery fix

* Fix
2026-04-19 11:14:23 +01:00
Xi Zhang 58435dba52 Release/v0.0.8 (#167)
* chore(release): update version to v0.0.8 and dependencies in project files

* feat(models): add new model entries for Claude Opus 4-7 and update version handling

* Refactor code structure for improved readability and maintainability
2026-04-18 16:49:31 +01:00
Xi Zhang f4a3617646 refactor(paths): unify global data directory to ~/.evoscientist and u… (#164)
* refactor(paths): unify global data directory to ~/.evoscientist and update related paths

* refactor(paths): update legacy session migration to respect XDG_CONFIG_HOME

* refactor(tests): clear XDG_CONFIG_HOME in legacy session migration tests for deterministic behavior
2026-04-18 15:01:22 +01:00
Xi Zhang 7c6b6755f2 fix(docs): update survey literature and macOS deployment links for accuracy 2026-04-17 15:25:23 +01:00
Xi Zhang 4b2aaea49a fix(docs): update survey literature link to point to the correct GitHub path 2026-04-17 15:20:59 +01:00
Xi Zhang 1965dd8661 Add new asset images for survey literature examples
- Added model_selection.png to illustrate model selection process.
- Added prompt.png for visual representation of prompts used in surveys.
- Added skill_selection.png to depict skill selection criteria.
2026-04-17 15:18:06 +01:00