* fix: update deepagents and langchain versions; enhance _reduce_messages_delta handling for None state
* fix: update langchain version constraint to >=1.3 in pyproject.toml and uv.lock
* feat: implement configurable sandbox execute timeout and enhance recovery instructions
* feat: add background process management tools and middleware for sandbox execution
* feat: enhance background process management with completion notifications and deduplication
* feat: enhance sandbox execution timeout validation and update related messages
* feat: enhance background process management with thread-specific completion notifications and HITL approval handling
* test: assert completion notification waits for process finish timestamp
* fix(cli): forward async-notifier replies back to originating channel
When PR #214's auto-notifier fires a synthetic agent turn after a
channel-originated conversation, the synthesized response only rendered
to the local CLI/TUI — the channel user (iMessage etc.) saw nothing
and had to manually re-prompt to find out what happened.
Adds a per-thread channel-origin registry in cli/channel.py and wires
the three notifier paths (Rich CLI / TUI / serve) to publish the final
response back via bus.publish_outbound when the originating thread was
started by a channel turn. Publish is fire-and-forget (scheduled on the
bus loop + done-callback for failure logging) so the notifier turn
doesn't block on the asyncio / textual event loop.
The registry is cleared on /new and /resume rotation so stale entries
don't accumulate.
* fix(cli): address review feedback on channel-origin forwarding
Follow-up to the review on #244 (din0s, X-iZhang):
- Guard the /resume origin cleanup on a real thread change in Rich CLI
and TUI (serve mode already did via thread_changed). Resuming the
already-active thread no longer wipes its still-live origin, which
would otherwise silently drop a later async-notifier forward — the
exact gap this PR closes.
- Re-bind the now-current thread to its channel after a channel-issued
/new or /resume slash command (which rotates the thread inside the
dispatch), so notifier turns on the rotated thread still forward.
- Guard the publish done-callback against a cancelled future, whose
.exception() raises CancelledError (rather than returning it) on
bus-loop teardown, so the intended warning still logs.
- Mirror the normal reply path's manager.record_message(channel, "sent")
for forwarded notifications so per-channel stats stay accurate.
- Print the closing "[channel: Replied to ...]" line in all three
notifier paths (Rich CLI / TUI / serve) when a forward actually
happened, so the forwarded block reads as terminated on screen.
Adds test_publish_records_sent_metric. ruff clean; notification-origin
suite (10) + related channel/CLI/serve suites (728) pass.
* fix(cli): store sender information separately from chat_id in channel origin
---------
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* Add onboarding wizard with style and validation components
- Introduced `style.py` for shared visual elements used in the onboarding wizard.
- Created `validators.py` for input validation, including integer and choice validators, and API key validation functions for various providers.
- Implemented `wizard.py` as the entry point for the onboarding process, managing user prompts and configuration steps.
- Added progress rendering and autosave functionality to enhance user experience during the onboarding process.
* feat(onboarding): enhance validation and configuration for onboarding wizard
- Added validation for UI backends, workspace modes, and providers in the onboarding command.
- Updated channel definitions to include secret field handling for sensitive tokens.
- Improved user prompts for required fields, ensuring sensitive data is masked.
- Introduced constants for valid providers, UI backends, and workspace modes to maintain consistency.
- Implemented tests to ensure alignment between constants and interactive choices in onboarding steps.
* feat(onboarding): improve WeChat account ID prompt and validation for newly enabled channels
* feat(onboarding): enhance WeChat backend credential prompts and validation
* feat(onboarding): refine WeChat backend credential prompts for wecom and wechatmp
* Refactor onboarding package for improved structure and clarity
- Simplified the onboarding package by removing unnecessary re-exports and consolidating public API to only include `run_onboard`.
- Updated `install_back_keys` to `install_navigation_keys` for clarity and consistency in the prompter module.
- Enhanced the `NonInteractivePrompter` class to support strict mode, allowing for better handling of non-interactive prompts.
- Adjusted the onboarding steps to utilize the new navigation keys installation method.
- Improved the `run_onboard` function to handle section implications based on user flags, enhancing the onboarding experience.
- Updated tests to reflect changes in imports and ensure compatibility with the new structure.
* feat(onboarding): enhance validation logic for non-interactive prompts
* refactor(onboarding): streamline onboarding module structure and enhance validation error handling
* refactor(onboarding): enhance config revert logic to preserve original file state
* refactor(onboarding): enhance tavily key validation and error handling in onboarding process
* feat(feishu): scan-to-create QR onboarding flow
Add a device-code flow against accounts.feishu.cn/oauth/v1/app/registration
that lets users scan a terminal QR code with Feishu / Lark mobile to
auto-create a PersonalAgent bot app with the required IM permissions
pre-attached. The poll endpoint returns app_id + app_secret, which the
onboarding wizard then writes into the channel config — no manual app
creation on open.feishu.cn required.
- channels/feishu/onboard.py: qr_register() public entry, init/begin/poll
helpers, QR rendering via the soft qrcode dep, automatic feishu↔lark
domain switch based on the scanning user's tenant_brand, and a
best-effort bot probe to surface the bot name in the wizard
- channels/feishu/__init__.py: re-export qr_register (mirrors qq)
- config/onboard.py: offer "Scan QR code (recommended) / Enter manually"
in the Feishu branch, ask for region (feishu vs lark), then call
qr_register and populate feishu_app_id / feishu_app_secret /
feishu_domain; add qrcode>=7.4 to the feishu pip extras
* fix(feishu): silently absorb unsubscribed WebSocket events
Feishu auto-subscribes PersonalAgent apps to many event types
(im.message.reaction.created_v1, message.read_v1, message.recalled_v1,
chat.member.*, ...) that EvoScientist doesn't register handlers for.
Without intervention, lark-oapi's dispatcher raises EventException
("processor not found, type: ..."), the WS client logs it at ERROR and
replies HTTP 500 on the frame, and Feishu marks the event as failed
and retries it.
The problem is amplified by _send_ack_reaction: every inbound message
triggers our own reaction, which Feishu echoes back as
reaction.created_v1, creating a continuous ERROR-log feedback loop and
pointless retries.
Wrap EventDispatcherHandler._do_without_validation after build() to
swallow "processor not found" EventExceptions (debug log + return None)
while letting all other errors propagate. Failure-safe: if lark-oapi's
internal API changes the wrapper degrades to the prior behavior rather
than breaking the channel.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat(deploy): implement standalone LangGraph server and CLI command for deployment
* feat(deploy): enhance port validation and environment variable management for deployment
* Refactor langgraph dev deployment and introduce workspace sidecar protocol
- Updated the deployment mode handling in `server.py` to use a single environment variable `EVOSCIENTIST_DEPLOY_MODE` with values `full` and `stripped`.
- Enhanced the `manager.py` to implement a workspace fingerprint sidecar, allowing cross-process reuse of langgraph dev instances while ensuring workspace consistency.
- Introduced functions to write and read the workspace sidecar, with error handling for missing or corrupt data.
- Added tests for the workspace sidecar functionality, including validation of the JSON schema and ensuring proper error handling for workspace mismatches.
- Updated existing tests to reflect changes in deployment mode handling and added new tests for signal handling during shutdown.
- Ensured that cleanup routines remove the workspace sidecar alongside the PID file during shutdown.
* fix(langgraph): improve workspace sidecar checks for process ownership and stale handles
* feat(middleware): add CodeInterpreterMiddleware with project-specific configuration
chore(config): increase checkpoint retention limit for runaway conversations
fix(tests): update database schema references from 'blob' to 'value'
chore(deps): update deepagents dependency to include quickjs support
* feat(deepagents): update to version 0.6.1 and add optional dependencies for quickjs
* feat(sessions): improve error handling for message deltas and update Overwrite type check
* Enhance PruningCheckpointer with DeltaChannel Awareness
- Introduced a new pruning strategy in `_prune_after_put` to preserve the `_DeltaSnapshot` chain during checkpoint pruning.
- Implemented methods to fetch recent checkpoint IDs and walk to snapshot ancestors, ensuring that necessary checkpoints are retained.
- Updated SQL queries to handle checkpoint and write deletions more efficiently.
- Added comprehensive tests for DeltaChannel-aware pruning, ensuring that the pruning logic correctly handles various checkpoint scenarios, including those with and without snapshot seeds.
- Refactored `_load_checkpoint_messages` to utilize the new saver interface, improving message reconstruction from checkpoints.
* feat(tests): add migration sweep test to preserve snapshot ancestor
* feat(sessions): enhance checkpoint retrieval to prevent transcript leakage in multi-agent scenarios
* feat(middleware): enhance CodeInterpreterMiddleware with configurable timeout and result character limit
feat(config): add CodeInterpreterMiddleware tuning parameters to EvoScientistConfig
feat(sessions): implement inline message delta reducer for improved message handling
* feat(dependencies): update deepagents version to 0.6.2 in pyproject.toml and uv.lock
* feat(backends): implement tier-aware virtual mount resolution for skills and memories
* test: add end-to-end test for workspace tier shadowing global tier in CustomSandboxBackend
* feat(backends): enhance virtual mount resolution for skills and memories with tier paths and quoting
* fix(tests): update Python command in virtual mount resolution tests to use python3
* fix(cli): handle spaces in @file mentions
The @file parser truncated at the first space, so dragging or pasting a
filename like `@PREPING_ Building Agent.pdf` only matched `@PREPING_`
and warned "file not found". Now supports `@"..."` / `@'...'` quoted
form for explicit paths, plus a greedy expansion fallback that walks
across whitespace until an existing file resolves (bounded by newlines,
the next `@`, and a 20-token cap). Autocomplete also returns quoted
mentions for any candidate containing a space.
* style: apply ruff format to file_mentions
* feat(llm): add dashscope-code provider for Alibaba Coding Plan keys
Alibaba Cloud Bailian "Coding Plan" subscription keys (sk-sp-*) route
through a separate endpoint (coding.dashscope.aliyuncs.com/v1) that the
standard `dashscope` provider can't reach. Add a sibling provider entry
matching the zhipu/zhipu-code and moonshot/kimi-coding precedents, with
its own validator (the coding endpoint returns 404 on /models, so probe
via chat.completions instead).
Closes#224
* fix(llm): keep dashscope as default provider for qwen3-coder shortcut
The MODELS dict is built from _MODEL_ENTRIES via a last-write-wins dict
comprehension. The initial commit listed dashscope-code AFTER dashscope,
which silently flipped the bare `get_chat_model("qwen3-coder")` shortcut
to the coding endpoint — breaking standard sk-* keys.
Reorder to match the zhipu-code / zhipu precedent: coding endpoint first,
general endpoint last so the general endpoint wins the collision and
remains the default for the shared "qwen3-coder" short name.
* feat(qq): add inline keyboard buttons for C2C HITL approval
QQ Bot supports inline buttons via `markdown + keyboard` payloads. Clicks
arrive as `interaction_create` events through the existing botpy
WebSocket gateway — no extra subscription needed beyond enabling the
`interaction` intent. Group-scope clicks are out of scope here (DM only).
Send path
- `_build_qq_keyboard(buttons)` mirrors the Feishu helper, mapping the
generic `{text, value, type}` shape to QQ's `{render_data, action}`
with action.type=1 (callback). One button per row for mobile clarity.
- `_send_chunk` extracts `metadata["buttons"]` and threads a `keyboard`
payload into `_post_markdown_message` for C2C only.
- Markdown→plain fallback can't carry a keyboard, so when buttons were
attached the fallback content gets a textual `Reply: 1=Approve, …`
hint built from the button list. `_parse_approval_reply` accepts
the same values typed manually, so the user is never stuck.
Receive path
- `on_interaction_create` is registered on the bot class.
- `_on_interaction` extracts `data.resolved.button_data`, builds an
InboundMessage, runs it through inbound middleware (Dedup suppresses
retry callbacks), and publishes directly to the bus — bypassing the
per-sender debounce buffer so the click value isn't merged with any
text typed in the same window.
- Always ACKs via `api.on_interaction_result(id, 0)` in `finally` so
QQ doesn't show the button as "expired", even if middleware drops
the click or something throws downstream.
`QQ.inline_buttons=True`; `_approval_prompt_metadata` now auto-attaches
the Approve/Reject/Approve-all button row for QQ HITL prompts.
* fix(qq): button-value coercion, ACK timing, HITL consumer wiring
Fixes 6 bugs found in the inline-keyboard commit and consolidates the
button helpers so the keyboard builder, plain-text fallback hint, and
interaction handler share one coercion path.
- Plain-text fallback no longer crashes on non-string `value` (e.g.
`{"text": "OK", "value": 42}`). Extracted `_normalize_button` is now
the single place that resolves `(label, value)` and coerces non-strings.
- `metadata["button_value"]` is the coerced string instead of the raw
payload, matching `content` and downstream string comparisons.
- `_on_interaction` ACKs first, before publishing to the bus, so the
QQ button UI never shows "expired" if middleware is slow.
- Wire `_approval_prompt_metadata` + `_format_approval_prompt(with_buttons=)`
into `InboundConsumer._stream_with_hitl` and `cli.channel.channel_hitl_prompt`
so the QQ `inline_buttons=True` capability is actually used end-to-end
(HITL prompts auto-attach Approve/Reject/Approve-all buttons when the
channel advertises the capability).
- Trim contradictory `_QQ_DEFAULT_PERMISSION` comment.
- Fix `test_group_interaction_ignored` docstring (ACK runs first now,
not in `finally` after a `return`).
Tests: `_normalize_button` covered indirectly via existing keyboard tests;
new regressions for non-string fallback hint, ACK-on-handler-throw, and
string-coerced `button_value` metadata.
* refactor(qq): slim button helpers and explicit has_buttons flag
Inline single-use _button_hint and the _QQ_BUTTON_STYLE/_QQ_DEFAULT_PERMISSION
constants in qq/channel.py; tighten _on_interaction (drop unreachable
"[button click]" sentinel and unused triggering_message_id metadata; collapse
"if resolved else" ternaries via `or ""`).
Replace the metadata round-trip ("buttons" in metadata) used to detect button
support in consumer.py and cli/channel.py with an explicit has_buttons bool
threaded through both the prompt formatter and metadata builder.
Apply ruff format to the previously unformatted blocks introduced earlier on
this branch so CI lint passes.
* feat(qq): send post-decision confirmation after HITL approval
Send a visible confirmation message ("✅ 已批准" / "❌ 已拒绝") right after
the user resolves a HITL approval — QQ Bot has no message-recall or edit API
for C2C, so a follow-up message is the only way to give the click/reply
strong feedback.
Bus consumer (consumer.py): only sends the confirmation when the user
actually responded (event was set), to avoid pretending the user approved
when the request really timed out and auto-approved.
CLI HITL prompt (cli/channel.py): mirrors the same set of confirmation
strings. Timeout and unrecognized-reply paths keep their existing English
text since their semantics differ (auto-reject vs auto-approve, plus a
hint about the unparsed input).
* feat(middleware): add ConfigurableModelMiddleware for dynamic model resolution
- Introduced ConfigurableModelMiddleware to resolve chat models from RunnableConfig.configurable on each call.
- Updated middleware initialization to include ConfigurableModelMiddleware.
- Enhanced context editing middleware tests to verify presence of ConfigurableModelMiddleware.
- Implemented tests for ConfigurableModelMiddleware to ensure correct model overriding and caching behavior.
- Added tests for deepagents model-passthrough patch to verify configuration injection in async tasks.
* feat(async-subagent): update middleware handling to prevent deadlocks in async sub-agents
* style: Refactor code formatting for improved readability in patches and test files
* refactor: streamline middleware construction and improve async handling in ConfigurableModelMiddleware
* fix: remove unused request parameter from _read_model_override function
* refactor: improve async handling in _ClientProxy and enhance logging in ConfigurableModelMiddleware
test: add behavior test to ensure AskUserMiddleware is excluded in async subagent mode
* feat(qq): add QR-code scan-to-configure onboarding for QQ Bot
Adds a `qr_register()` flow that drives q.qq.com's create_bind_task /
poll_bind_result APIs so the wizard can auto-fill `qq_app_id` and
`qq_app_secret` after the developer scans a QR code with a bound QQ
account, falling back to manual entry on failure or cancel.
- channels/qq/crypto.py: AES-256-GCM helpers for decrypting the bot's
client_secret returned by poll_bind_result.
- channels/qq/onboard.py: portal API client + polling loop.
- channels/qq/__init__.py: re-export `qr_register`.
- config/onboard.py: QQ branch in `_step_channels` that offers
"Scan QR code" vs "Enter manually", and skips the manual prompt
loop when a scan succeeded.
* style(qq): fix ruff lint errors in onboard.py
Move `import os` to the top-level import block (E402), drop the legacy
`typing.Optional`/`typing.Tuple` imports (UP035), and use the PEP 585/604
builtin generics (`tuple[...]`, `X | None`) for the few annotations that
still referenced them (UP006/UP045). No behavior change.
* fix(qq): harden QR onboard error paths and declare scan deps
Address review feedback on PR #213:
- Declare cryptography>=41.0 and qrcode>=7.4 in [qq]/[all-channels]
extras and in _CHANNEL_PIP_DEPS so the scan flow no longer fails
with an opaque ImportError on a fresh `evoscientist[qq]` install.
- Polling loop logs each _poll_bind_result failure and aborts after
5 consecutive errors instead of silently spinning until the 600s
timeout, restoring the documented Raises: RuntimeError contract.
- Wrap decrypt_secret in try/except so failures honor the
None-on-failure contract instead of letting exceptions escape.
- Preflight `import cryptography` in the scan branch and offer
install or fall back to manual entry.
* style: ruff format collapse two over-wrapped log/console lines
* fix: Improve watcher logic to prevent false-positive notifications on clean stream exits
* fix: Update watcher logic to drop notifications on persistent runs.get failures
* fix: Refactor test for watcher persistent failure notification handling
* fix: Enhance watcher test to validate all notification queues are empty after reconnect budget exhaustion
* feat: Implement async sub-agent auto-notification system
- Added async notifier functionality to handle notifications for sub-agents reaching terminal states.
- Introduced `AsyncTaskNotification` dataclass for structured notification data.
- Implemented `watch_run_and_notify` to monitor agent runs and enqueue notifications.
- Created `spawn_watcher` to manage watcher tasks and ensure proper cancellation of previous watchers.
- Developed `consume_notifications` to process notifications, deduplicate them, and format messages for LLM.
- Added tests for notification handling, including draining, deduplication, and formatting.
- Patched deepagents to integrate the new watcher functionality into start and update tools.
* Enhance async notifier with per-thread notification routing and error handling
- Introduced `origin_cli_thread_id` to `AsyncTaskNotification` for routing notifications back to the originating CLI session.
- Implemented per-thread notification queues to handle notifications based on the originating thread.
- Updated `has_pending_notifications` and `drain_notifications` to respect thread-specific queues.
- Enhanced `watch_run_and_notify` to detect in-band error events from the SSE stream and handle clean exits.
- Modified tests to verify the new notification routing behavior and ensure proper handling of notifications across threads.
- Added a fixture to restore the async watcher patch state in tests to prevent state leakage.
- Updated deepagents patching to capture the main agent's CLI thread ID for notification routing.
* feat: Enhance async notifier with thread-specific watcher management and notification filtering
* test: Enhance notification draining logic for cleaner test setup
* refactor: Remove summary field from AsyncTaskNotification and update related tests
* feat: Enhance async notification handling with target thread ID support
* Refactor async notifier and middleware for improved task management
- Removed the no-op shutdown watcher loop from async_notifier.py as it is no longer needed.
- Updated watch_run_and_notify to clarify notification handling and race conditions.
- Cleaned up shutdown handling in commands.py, interactive.py, and tui_interactive.py by removing obsolete shutdown watcher calls.
- Deleted the deepagents async watcher patch from patches.py, transitioning to a new middleware approach.
- Introduced AsyncWatcherMiddleware to handle async task notifications directly during tool calls.
- Updated tests to validate the new middleware functionality and ensure proper watcher spawning and cancellation.
- Enhanced test coverage for async watcher middleware, including edge cases and error handling.
* feat(tests): add fixture to reset notifier state before each test
* feat(wechat): add personal-WeChat (iLink) backend with QR-code login
Adds a third WeChat backend alongside WeCom and Official Account:
``personal`` rides Tencent's iLink Bot long-poll gateway so a personal
WeChat account can act as a bot. Credentials are obtained via QR-code
scan and persisted under ``DATA_DIR/wechat_personal/accounts/``.
- channels/wechat/personal.py: WeixinPersonalChannel + qr_login.
- channels/wechat/crypto.py: aes128_ecb_decrypt + parse_ilink_aes_key
for the iLink CDN media protocol.
- channels/wechat/probe.py: validate_wechat_personal credential probe.
- channels/wechat/serve.py: --backend personal CLI + --qr-login flow.
- channels/wechat/__init__.py: factory dispatch on wechat_backend; pull
in the new dependencies in the docstring.
- config/settings.py: wechat_personal_* fields.
- config/onboard.py: WeChat-backend picker + QR-scan flow in the wizard
+ personal-backend probe in _probe_channel.
- pyproject.toml / uv.lock: add qrcode + certifi to wechat & all-channels
extras (aiohttp was already pulled in transitively).
* fix(wechat): address ruff failures and CodeRabbit review on personal-WeChat PR
- personal.py: drop unused imports (`field`, `PollingMixin`); replace
`asyncio.TimeoutError` with builtin; hold references to background
`asyncio.create_task` results so they aren't GC'd; wire `dm_policy`
through `_process_message` (disabled/allowlist) so `wechat_personal_dm_policy`
actually takes effect for DMs.
- onboard.py: import-check gate now validates the full WeChat dependency
set (aiohttp, qrcode, Crypto, certifi) instead of only aiohttp; mask
`WeCom Secret` and `MP App Secret` prompts via `questionary.password`;
derive the QR-login hint path from `_account_dir()` instead of the
hard-coded `~/.evoscientist/...`; stop copying the QR-login token into
the main config (already persisted per-account on disk — copying broadens
secret exposure and risks staleness).
- pyproject.toml: allow Chinese full-width punctuation in `allowed-confusables`
for user-facing CN messages.
* style(wechat): apply ruff format
`ruff format --check` was failing CI on three files (one pre-existing in
`__init__.py` plus formatter-driven line-merges in the files touched by
the previous fix commit). Ran `ruff format` to bring them in line; both
`ruff check` and `ruff format --check` now pass.
* refactor(channels): remove dead MessageBus dispatcher
Outbound routing has two implementations: ``MessageBus.dispatch_outbound``
(subscriber-based) and ``ChannelManager._dispatch_outbound`` (registry
lookup). Only the latter is ever started in production — the former
is reachable solely from tests, yet both consume from the same
``bus.outbound`` queue. If anyone followed the bus's own API surface
they would silently steal messages from the real dispatcher.
Drop the unused machinery to leave a single, obvious outbound path:
- ``MessageBus.subscribe_outbound`` / ``dispatch_outbound`` / ``stop``
- ``_running`` flag and ``_outbound_subscribers`` map
- ``OutboundCallback`` type alias
- The lone ``bus.stop()`` call in ``cli/channel.py`` (was no-op)
- Four tests covering the removed code paths
* test(channels): drop empty MessageBus stubs after dispatcher removal
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* Refactor sub-agent architecture and introduce async support
- Removed the legacy subagent.yaml file and replaced it with individual YAML files for each sub-agent in the subagents directory.
- Updated the load_subagents function to support both directory and single file layouts for loading sub-agent configurations.
- Added new langgraph_dev module for managing async sub-agent lifecycle and deployment.
- Created graphs for async sub-agents (writing-agent, data-analysis-agent) and updated langgraph.json for deployment.
- Introduced new sub-agent definitions for planner, research, debug, code, and writing agents with appropriate system prompts and configurations.
- Enhanced package data inclusion in pyproject.toml to accommodate new sub-agent YAML files.
* Refactor code for improved readability by consolidating conditional statements and formatting
* feat: enhance async sub-agent support with workspace synchronization and user feedback
- Added console status messages during async sub-agent server startup and workspace synchronization to improve user experience.
- Implemented a new WorkspaceSyncWidget for live feedback during workspace sync operations.
- Updated onboarding to reject occupied ports and ensure proper workspace handling for async sub-agents.
- Introduced locking mechanisms to manage concurrent access to langgraph dev processes and workspace states.
* feat: add async sub-agent configuration and server management functions
* feat: improve port occupation handling and log file management in start_langgraph_dev
* feat: enhance async sub-agent handling and introduce comprehensive tests
- Updated `_maybe_swap_async_subagents` to improve async sub-agent management, ensuring internal flags are stripped before handoff.
- Enhanced port management in `onboard.py` to allow reuse of occupied ports if already running by the same service.
- Introduced file locking in `manager.py` to prevent race conditions during concurrent CLI invocations.
- Added new tests for async sub-agent swapping and langgraph manager functionalities to ensure reliability and correctness.
- Updated dependencies in `pyproject.toml` to include `psutil` and `filelock`.
* fix(docs): clarify sub-agent configuration in README
* test(manager): isolate _PID_DIR + tighten reuse-path assertion
Addresses CodeRabbit review on tests/test_langgraph_manager.py:
- Patch _PID_DIR to tmp_path so the FileLock setup in
ensure_langgraph_dev doesn't mkdir the user's real
~/.config/evoscientist/ dir as a test side-effect.
- Tighten "result is None or hasattr(result, 'poll')" to a strict
"result is None" — the reuse path returns None unconditionally,
so the OR clause was hiding potential regressions.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(manager): clean up stale PID file when unrelated process reuses PID
* feat(tests): add validation tests for async flag in load_subagents
* fix(load_subagents): restrict to .yaml files and clarify configuration handling
* fix(load_subagents): improve error handling for non-dict specifications in YAML
* feat(onboard): add "LangGraph Port" step to onboarding process
* feat(langgraph): add concurrency configuration for langgraph dev workers
* feat(async-subagents): enhance MCP tool routing for async sub-agents
* fix(manager): update exception handling for connection errors and prevent zombie processes
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(display): enhance approval prompt with questionary for better navigation
* feat(cli): add HumanInTheLoopMiddleware for user approval in main agent
* feat(docker): official image with all runtime deps pre-installed
Multi-stage build using uv for the EvoScientist core + all messaging-channel
extras, plus Node.js 24 LTS (for npx-based MCP servers) and uv (for runtime
Python MCP installs) in the runtime layer. Runs as non-root user evosci,
with workspace, app data, and config (XDG_CONFIG_HOME) all consolidated
under a single /home/evosci/.evoscientist volume so a single mount
persists everything across container restarts.
Includes a docker-compose.yml starter, a build/push GitHub Actions
workflow targeting ghcr.io with multi-arch (amd64/arm64) and PR-only
build verification, a .dockerignore, and a new Docker section in the
README documenting mounts, derivation recipes for the unbundled stt /
oauth / TinyTeX extras, and proxy/cert handling expectations.
* fix(docker): pin trixie base + drop redundant python image
Switch builder and runtime from `python:3.11-slim-bookworm` to a single
`ghcr.io/astral-sh/uv:python3.11-trixie-slim` base — trixie drops several
CRITICAL vulnerabilities that bookworm carries today, and reusing the uv
image for runtime eliminates the separate `COPY --from=…/uv` line.
* chore(docker): pin GitHub Actions to commit SHAs in workflow
Replace mutable major-version tags with full commit SHAs (with the
corresponding semver tag in a trailing comment) so a compromised /
retagged action release can't silently change what runs in the publish
pipeline.
* chore(deps): enable Dependabot version updates for Dockerfile pins
Adds a weekly `docker` ecosystem that watches the Dockerfile's `FROM` /
`COPY --from=` references — including the ARG-bound `BASE_IMAGE` and
`NODE_IMAGE` digests — and opens one grouped PR per cadence bumping
both the @sha256 digest and the trailing version comment. This keeps
the otherwise-frozen pins flowing with Debian point releases and
upstream patches.
* fix(docker): use nodejs alias stage so NODE_IMAGE ARG actually resolves
`COPY --from=${NODE_IMAGE}` left the dollar-curly literal at parse time
under buildkit 29.x — it expands ARGs in `FROM` but reads `--from=` as a
static stage/image name. Introduce a tiny `FROM ${NODE_IMAGE} AS nodejs`
alias and `COPY --from=nodejs …` against it, which preserves the
ARG-driven Dependabot updates without tripping the parser.
* fix(docker): harden venv ownership and PATH ordering
- Drop `--chown` on the `/opt/venv` COPY so the venv stays root-owned.
The runtime user only needs read+execute (default Unix perms allow
that); making it user-owned let the agent rewrite its own
dependencies, which defeats the sandboxing premise. All persistent
agent state already lives under /home/evosci/.evoscientist/.
- Reorder PATH so /opt/venv/bin precedes the user-writable
UV_TOOL_BIN_DIR. Otherwise a stray binary dropped into the latter
(e.g. via `uv tool install`) could shadow the canonical
`evosci` / `python` / `pip` shipped with the image.
* docs: update README
* docs(docker): warn about non-root UID and `curl | sh` for derived images
- The image runs as `evosci` (UID 1000), so a host-side `./workspace`
bind mount fails if the host user has a different UID — same gotcha
that bites onboarding's `mcp.yaml` write. Add an !IMPORTANT block
with the two practical fixes (`chown -R 1000:1000` once, or
`--user "$(id -u):$(id -g)"` on each run).
- The TinyTeX derivation snippet pipes an unpinned remote installer
into `sh`. Add a one-line pointer to fetching a pinned release
tarball from `rstudio/tinytex-releases` for users who'd rather not
trust the upstream script blindly. The official installer is kept
as the default since that's what TinyTeX itself recommends.
* chore(docker): cancel in-flight workflow runs + flag iMessage as host-only
- Add `concurrency: cancel-in-progress: true` to the docker workflow
so successive pushes on the same ref supersede the prior run rather
than queueing in parallel — multi-arch buildx is the slowest job in
CI, no point burning minutes on superseded builds.
- Spell out that the docker image installs the `all-chanels` extra and
call out iMessage as a deliberate host-only exclusion: it requires
the `imsg` CLI bridging to macOS's Messages.app, which no Linux
container config can satisfy.
* fix(onboard): detect installed skill packs via install manifest
Onboarding's _step_skills only inspected USER_SKILLS_DIR and matched
recommended entries by directory-name hint, so a pack like
EvoScientist/EvoSkills@skills (which explodes into paper-writing/,
evo-memory/, etc. under GLOBAL_SKILLS_DIR) was never detected and kept
appearing as not-yet-installed.
skills_manager now writes a per-tier .installed.yaml mapping skill
directory name -> original install source on every install, removes the
entry on uninstall, and exposes installed_sources(). _step_skills checks
both tiers and treats a recommended source as installed when present in
any manifest -- so packs are recognized regardless of how their child
dirs are named.
* fix(onboard): write install manifest atomically
Stage to a sibling temp file, fsync, then os.replace into place. A crash
mid-write can no longer leave a half-written .installed.yaml behind,
which would otherwise wipe out pack detection until the next reinstall.
* style: fmt
* fix(onboard): catch decode errors when loading install manifest
read_text() can raise UnicodeDecodeError on a hand-edited or corrupt .installed.yaml; pin encoding="utf-8" and add UnicodeError to the except clause so the function honors its "returns {} on any error" contract.
* refactor(cli): replace channel module globals with ChannelRuntime
Removes _cli_agent / _cli_thread_id from EvoScientist/cli/channel.py
and threads a ChannelRuntime via CommandContext.channel_runtime so
/model and /channel rebind without poking module-level state.
* fix(cli): address coderabbit review
- _auto_start_channel: bind ChannelRuntime only after
_start_channels_bus_mode succeeds, so a startup failure no longer
leaves a stale binding pointing at channels that never started.
- _sync_tui_command_completion (TUI) and the Rich CLI command-completion
paths: rebind the runtime on thread rotation, not just agent swap, so
/new and /resume keep ChannelRuntime in sync with the running thread
(matches the serve-mode hook contract).
- test_hook_syncs_channel_runtime: pin ctx.thread_id explicitly so a
bare MagicMock attribute can't silently mutate runtime.thread_id.
- New regression test covering the rebind-on-thread-rotation contract.
* fix(sessions): ensure migration sweep runs before yielding checkpointer to prevent race conditions
* fix(sessions): enhance migration sweep with progress indication and ETA estimation
* Implement PruningCheckpointer for efficient checkpoint management and add comprehensive tests
- Introduced `PruningCheckpointer` to manage checkpoint pruning after each `aput()`, ensuring only the latest checkpoints are retained based on a configurable limit.
- Added migration sweep functionality to clean up legacy checkpoints and prevent database bloat.
- Enhanced `get_checkpointer()` to utilize the new `PruningCheckpointer` and trigger migration sweeps when necessary.
- Developed a suite of integration tests for `PruningCheckpointer`, covering various scenarios including pruning behavior, concurrent writes, and retention policies.
- Implemented tests for migration sweep functionality, ensuring proper partitioning and user version management.
- Added diagnostic helper `db_stats` to provide insights into the database state, including thread and checkpoint counts.
* feat(sessions): enhance pruning logic to handle legacy DBs without writes table
* fix(tests): prevent atexit hook leakage in TestMigrationSweep
* feat(tests): enhance TestPruningCheckpointer to validate put+prune serialization
* feat(tests): refactor mock path implementation for get_db_path in test cases
* feat(context-window): add model context window patch table and apply function
* fix(tests): clean up formatting in context window tests
* fix(tests): update context window tests for Claude model exceptions
* feat: Refactor system prompt structure and enhance documentation for clarity
* refactor: Improve clarity and consistency in prompt documentation
* refactor(tests): Improve readability of first-person avoidance test assertion
* refactor: Remove redundant datetime imports and enhance prompt documentation
* feat: Enhance ModelPickerWidget for Ollama integration
- Implemented a sentinel row for "Custom Ollama model..." in ModelPickerWidget, allowing users to input arbitrary model names.
- Updated action handling in ModelPickerWidget to manage transitions between list and input modes.
- Added async model discovery for Ollama models, integrating with the /model command to fetch locally installed models.
- Created tests for Ollama model discovery and ModelPickerWidget behavior, ensuring proper functionality and user experience.
- Refactored validate_ollama_connection and discover_ollama_models for improved error handling and response management.
* fix: Simplify code by removing unnecessary line breaks in ModelPickerWidget and test cases
* fix: Restore globals on set_chat_model failure to prevent half-switched session
* fix: Improve error handling in ModelCommand by restoring globals on failure