Commit Graph

482 Commits

Author SHA1 Message Date
Xi Zhang 7cfec02416 Fix/sessions migration sweep race (#195)
* fix(sessions): ensure migration sweep runs before yielding checkpointer to prevent race conditions

* fix(sessions): enhance migration sweep with progress indication and ETA estimation
2026-04-29 02:06:25 +01:00
Xi Zhang 50719ef256 Implement PruningCheckpointer for efficient checkpoint management and… (#194)
* Implement PruningCheckpointer for efficient checkpoint management and add comprehensive tests

- Introduced `PruningCheckpointer` to manage checkpoint pruning after each `aput()`, ensuring only the latest checkpoints are retained based on a configurable limit.
- Added migration sweep functionality to clean up legacy checkpoints and prevent database bloat.
- Enhanced `get_checkpointer()` to utilize the new `PruningCheckpointer` and trigger migration sweeps when necessary.
- Developed a suite of integration tests for `PruningCheckpointer`, covering various scenarios including pruning behavior, concurrent writes, and retention policies.
- Implemented tests for migration sweep functionality, ensuring proper partitioning and user version management.
- Added diagnostic helper `db_stats` to provide insights into the database state, including thread and checkpoint counts.

* feat(sessions): enhance pruning logic to handle legacy DBs without writes table

* fix(tests): prevent atexit hook leakage in TestMigrationSweep

* feat(tests): enhance TestPruningCheckpointer to validate put+prune serialization

* feat(tests): refactor mock path implementation for get_db_path in test cases
2026-04-28 22:26:59 +02:00
Xi Zhang 56cc2fef85 fix(deepseek): add empty-string fallback for reasoning_content in cross-provider scenarios (#192) 2026-04-27 19:19:50 +01:00
Xi Zhang 52f8d3a3a5 Feat/llm context window patch table (#191)
* feat(context-window): add model context window patch table and apply function

* fix(tests): clean up formatting in context window tests

* fix(tests): update context window tests for Claude model exceptions
2026-04-27 18:28:57 +01:00
X-iZhang 26e3452ef6 chore: update version to v0.0.9 in badges, README, and project files 2026-04-26 15:16:12 +01:00
Xi Zhang 20c06d4897 feat(deepseek): implement reasoning_content passback for multi-turn s… (#190)
* feat(deepseek): implement reasoning_content passback for multi-turn scenarios

* fix(tests): ensure consistent import of EvoScientist.llm.patches in test cases

* fix(tests): streamline tool_calls formatting in TestPatchDeepseekReasoningPassback

* feat(patches): add reasoning_content capture and re-injection for DeepSeek assistant messages

* fix(deepseek): optimize reasoning_content extraction and assignment in passback

* fix(deepseek): refine reasoning_content handling in OpenAI capture patch
2026-04-26 14:53:08 +01:00
Xi Zhang e48bc1cb71 Refactor/system prompt structure (#189)
* feat: Refactor system prompt structure and enhance documentation for clarity

* refactor: Improve clarity and consistency in prompt documentation

* refactor(tests): Improve readability of first-person avoidance test assertion

* refactor: Remove redundant datetime imports and enhance prompt documentation
2026-04-26 11:08:08 +01:00
Ziheng Zhang da74c325d6 fix(channel): scope stop and restore resume history (#186)
* fix(channel): scope stop and restore resume history

* refactor(channel): simplify stop and resume patch

* Delete PR_MESSAGE.md

* fix(channel): address review feedback

* fix(channel): address remaining review bugs

* fix(channel): clean up stopped request handling

* fix(channel): preserve resolved replies and sync tui commands

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-25 16:12:49 +01:00
X-iZhang 27fee3256e fix(clipboard): remove unnecessary check for selection end in copy_selection_to_clipboard 2026-04-25 13:04:19 +01:00
Xi Zhang 558360b558 feat: Enhance ModelPickerWidget for Ollama integration (#187)
* feat: Enhance ModelPickerWidget for Ollama integration

- Implemented a sentinel row for "Custom Ollama model..." in ModelPickerWidget, allowing users to input arbitrary model names.
- Updated action handling in ModelPickerWidget to manage transitions between list and input modes.
- Added async model discovery for Ollama models, integrating with the /model command to fetch locally installed models.
- Created tests for Ollama model discovery and ModelPickerWidget behavior, ensuring proper functionality and user experience.
- Refactored validate_ollama_connection and discover_ollama_models for improved error handling and response management.

* fix: Simplify code by removing unnecessary line breaks in ModelPickerWidget and test cases

* fix: Restore globals on set_chat_model failure to prevent half-switched session

* fix: Improve error handling in ModelCommand by restoring globals on failure
2026-04-25 00:37:48 +01:00
Xi Zhang 49b03c36eb feat: add support for gpt-5.5 model in the LLM configuration and tests (#188) 2026-04-24 23:52:55 +01:00
Wiktor Cupiał 155c4eaa40 fix: text copy on remote sessions/legacy terminal emulators (#185)
* fix: text copy on remote sessions/legacy terminal emulators

* fix: display warning only once
2026-04-24 23:16:39 +02:00
Xi Zhang c134a16e19 Fix/channel slash rich cli (#184)
* feat(cli): implement slash command dispatch for channel messages

* fix(cli): make EvoSci serve exit on Ctrl+C and hot-swap /model (#181)

* fix(cli): streamline debug logging and formatting in channel command handling

* fix(cli): ensure proper handling of asyncio event loop in slash command processing

* fix(cli): add error handling for unexpected exceptions in slash command dispatch

* fix(cli): improve error messaging for slash command dispatch failures

* fix(tests): enhance test setup by restoring channel globals and simplifying assertions

* fix(cli): enhance slash command handling across UI surfaces and improve resume command warnings
2026-04-24 16:10:25 +01:00
Xi Zhang 3831198f19 fix(chat): resolve model switch lag by tracking model/provider key (#180)
* fix(chat): resolve model switch lag by tracking model/provider key for cache invalidation

* style: format set_chat_model function for improved readability

* fix(chat): improve model switching logic to prevent unnecessary cache rebuilds

* fix(model): ensure globals are restored on agent load failure to prevent model switch issues
2026-04-24 11:36:33 +01:00
Xi Zhang 92df3e1844 Refactor/cli command manager (#178)
* feat(cli): migrate command handling to CommandManager and enhance UI interactions

* feat(cli): add /clear and /help commands to enhance user experience

* feat(cli): implement /new and /resume commands with interactive session management

* Refactor MCP and Skills Command Handling

- Moved the interactive picker style to a centralized widget for consistency across MCP and Skills commands.
- Updated the MCP command to remove the old command dispatch logic, delegating to the new InstallMCPCommand.
- Enhanced the Skills command to utilize a new interactive picker for skill selection, improving user experience.
- Implemented cancellation handling in the picker to differentiate between user cancellations and empty selections.
- Added comprehensive tests for the new command structures and picker functionalities to ensure reliability.

* refactor(cli): streamline CommandManager dispatch and remove deprecated command set

* refactor(cli): enhance error handling and state management in ChannelCommand and RichCLICommandUI

* refactor(cli): update lifecycle callback terminology and improve async prompt handling in RichCLICommandUI

* refactor(cli): enhance SlashCommandCompleter to dynamically fetch workspace directory for autocompletion

* refactor(cli): unify quit handling in RichCLICommandUI with shared _stop helper

* refactor(cli): remove hardcoded slash commands and utilize command manager for dynamic completion

* refactor(cli): update MCP and skills command files for improved clarity and organization

* refactor(mcp_ui): remove unnecessary newline in _show_mcp_config function
2026-04-24 10:38:44 +01:00
X-iZhang f704e3d761 update 2026-04-23 14:33:02 +01:00
dinos 65828e9666 perf(cli): cut startup latency and defer MCP loading to the background (#171)
* perf(cli): cut startup time of `evosci --help` from ~2.2s to ~0.3s

Module-level imports were eagerly pulling in langchain.chat_models (with
the whole anthropic/openai/google stack), langgraph, textual, and
prompt_toolkit on every invocation — even for `--help` or `config list`.

Defer those with PEP 562 `__getattr__`, using `lazy_loader.attach` (SPEC-1,
the scientific-python standard) where it's a clean attach pattern:

- `EvoScientist/llm/__init__.py`: attach `models` lazily so importing
  `context_window` from this package no longer drags in langchain.
- `EvoScientist/stream/__init__.py`: attach display/events lazily; split
  the shared Rich `Console` singleton into a new lightweight
  `stream/console.py` so callers that only need `console` skip the
  `stream.events` → `langchain_core.messages` chain.
- `EvoScientist/cli/__init__.py`: hand-rolled `__getattr__` (reaches into
  `..stream.state`, which `lazy_loader` doesn't cover) so `commands` and
  `app` are the only eager loads.
- `EvoScientist/cli/commands.py`: move `cmd_interactive`/`cmd_run` to
  in-function imports so prompt_toolkit + textual only load when the
  interactive path actually runs.
- `EvoScientist/cli/_constants.py`: read `AGENT_NAME` on demand so
  `build_metadata` doesn't eagerly import `sessions` (langgraph/aiosqlite).

Adds `lazy-loader>=0.5` as a dependency.

* feat(cli): defer MCP tool loading with live per-server progress

The CLI was blocking ~5 s on MCP tool enumeration before the first
prompt appeared. Move the agent construction off the event loop and
surface per-server progress so the user can interact immediately and see
what's happening.

MCP client:
- Add an `on_progress` callback to `load_mcp_tools` / `aload_mcp_tools`
  / `_load_tools` emitting `start` / `success` / `error` events per
  server.
- Fan connection attempts out with `asyncio.gather` so latency no longer
  scales linearly with server count; cap simultaneous attempts at
  `_MAX_CONCURRENT_CONNECTIONS` (8) via a semaphore so a big stdio fleet
  doesn't spawn every subprocess at once.

Agent wiring:
- Plumb `on_mcp_progress` through `create_cli_agent` / `_load_agent` /
  `load_mcp_and_build_kwargs` so CLI and TUI can plug in collectors.

CLI (`cmd_interactive`):
- Run `_load_agent` in a background thread via `asyncio.to_thread`; the
  prompt and banner render immediately.
- `_await_agent_ready()` awaits the task before each agent-using site
  (first turn, channel messages, `/channel`, `/compact`). Raises if
  called without a prior `_start_agent_load` instead of silently
  reloading without the SQLite checkpointer.
- Pre-prime the progress dict from `load_mcp_config()` so the
  bottom-toolbar's `N/M` denominator is stable from the first render.
- Wrap `session.prompt_async` in `patch_stdout(raw=True)` so
  `console.print` from the worker-thread progress callback lands cleanly
  above the prompt as inline chat messages instead of stomping the
  prompt cursor.

TUI (`EvoTextualInteractiveApp`):
- Same background load + `_await_agent_ready()` gates on every
  `self._agent` read.
- New `MCPLoaderWidget` mounted at the top of `#input-shell` shows a
  header with `N/M` and one live row per server (spinner → ✓ / ✗ with
  tool count or error detail). On completion:
  - all-clean loads auto-dismiss ~2.5 s later;
  - cache hits (no events ever fired) dismiss immediately rather than
    flashing a misleading "0/N loaded";
  - failures keep the widget mounted so the user can read the errors.
  - `dismissed` property lets the app clear its ref so late events from
    slow servers become no-ops. The error branch of `_on_agent_loaded`
    also settles the widget so a load failure can't leave the spinner
    animating forever.
- Chat input is `disabled` while MCP resolves — no placeholder hack, no
"waiting…" system message.

Shared:
- Hoist braille spinner frames to `status_bar.SPINNER_FRAMES` and import
  them in the TUI widget so CLI and TUI animate in sync.

Tests:
- Extend `test_agent_mcp_cache` fakes to accept the new `on_progress`
  kwarg.
- New `TestLoadToolsProgressCallback` in `test_mcp_client` exercises the
  event sequence for success/failure/mixed fleets, verifies a buggy
  callback doesn't break the load, and asserts the semaphore caps
  in-flight connections.

* style: ruff

* chore: update uv.lock

* chore: uv.lock

* fix: coderabbit issues

* style: fmt

* fix: move _await_agent_ready inside try block

* fix(tui): auto-dismiss MCP loader widget on failure

The widget was designed to stay mounted on failure so the user could
read error detail, but since it's pinned above the input it never went
away in practice — just permanent banner clutter.

Auto-dismiss on failure too, with a longer grace (12s vs 2.5s) so the
error summary stays readable.

* fix: address second coderabbit pass

- Channel handlers (CLI + TUI): catch agent-load failures so the
  channel request doesn't hang; CLI moves `_await_agent_ready()`
  inside the existing try/except, TUI catches explicitly and calls
  `_set_channel_response` with the error.

- Stale background loads: `prev.cancel()` only stops the asyncio
  wrapper, not the thread running `_load_agent`. Added a generation
  token (`agent_load_id` / `self._agent_load_id`) and gated both
  progress and completion callbacks on it so a superseded load can't
  clobber the current session's state or UI.

- TUI prompt lifecycle: added `_agent_load_pending()` and gated the
  `_process_channel_message` / `_handle_command` finally blocks on it
  so `/new` or `/resume` invoked from a command keeps the prompt
  disabled until the fresh load settles.

- TUI readiness failures: `_run_turn` and `_handle_command` now
  catch exceptions from `_await_agent_ready()` and surface a
  "Agent failed to load: …" system message instead of letting the
  exception escape into Textual's traceback panel.

* refactor(cli): share background agent loader between CLI and TUI

The CLI and TUI were carrying near-identical copies of the same
background-load state machine: the `agent_task`, the `agent_load_id`
generation token, the gated progress/completion callbacks, and the
per-server progress dict. Every CodeRabbit finding on that lifecycle
had to be fixed in both files.

Extract it into `cli/_agent_loader.py`:

- `MCPProgressTracker` — owns the `server -> (state, detail)` dict;
  exposes `prime`, `record`, `snapshot`, `totals`.

- `BackgroundAgentLoader` — owns `agent`, the in-flight task, and the
  generation token. Exposes `start(**loader_kwargs)`, `await_ready()`,
  `is_pending`. Internally gates all progress/completion callbacks by
  generation so a superseded load can't clobber the current session.
  UI-specific rendering plugs in via `on_progress` / `on_success` /
  `on_failure` callbacks.

Both surfaces now just wire their UI hooks; the loader file holds no
Rich / prompt_toolkit / Textual dependencies. Net -345 lines from
`interactive.py` + `tui_interactive.py`; +20 unit tests pinning the
lifecycle (generation filtering, cache-hit short-circuit, failure
reset, progress ordering).

* refactor(cli): make _on_done the sole authority for agent state transitions

await_ready no longer sets self.agent — it just awaits the task and
reads what _on_done already wrote. Eliminates the dual-write overlap
(asyncio guarantees done-callbacks fire in registration order).

* fix(tui): let users type during MCP load, only block on send

Remove prompt-disabling during background agent load — the TUI now
matches the CLI approach where the input stays enabled and only gates
on await_ready() at submit time. The MCPLoaderWidget still provides
visual feedback that loading is in progress.

* fix(loader): preserve real load error on await_ready; dedup failure message

CodeRabbit flagged two issues with the new loader:

1. After a failed load, `_on_done` nulled `self._task`, so the next
   `await_ready()` hit the "before start()" branch and the CLI wrapper
   remapped it to a misleading "checkpointer not available" message —
   losing the real exception (bad MCP config, network, etc.).

   Keep `_task` set on failure so `await_ready` re-raises the real
   exception. Added `needs_restart` so TUI's auto-retry check stays a
   one-liner and doesn't need to reach into task internals.

2. TUI reported each load failure twice: once from
   `_on_agent_load_failure` (the done-callback) and once from each
   caller of `_await_agent_ready` (`_run_turn`,
   `_process_channel_message`, `_handle_command`) catching the re-raise.

   `_on_agent_load_failure` is now the sole local reporter; callers
   just handle control flow (return cleanly, set channel response to
   unblock remote).

* fix(cli): wire /model handler through the agent loader

The /model command from main (merged via f1f0d7c) still reached for
`state["agent"]` (CLI) and `self._agent` (TUI) — both removed by the
background-loader refactor. CLI raised KeyError on first invocation;
TUI raised AttributeError. Writes to the old fields also had no effect
because every other code path now reads from `agent_loader.agent`, so
the model switch would have silently failed.

Route everything through the loader: `await _await_agent_ready()` up
front so /model doesn't race with the initial background load, build
the `CommandContext` with the current agent, and sync `ctx.agent` back
into `agent_loader.agent` (plus channel globals) when the command
replaces it.

* fix(cli): isolate progress callback, capture awaited agent, gate by requires_agent

Three CodeRabbit findings on the loader + command dispatch path:

- Wrap ``_on_progress`` in try/except inside the loader's gated wrapper
  so a buggy UI adapter can't bubble into ``loader_fn`` and fail the
  whole background load. The MCP client already protects this, but
  defence-in-depth keeps the loader self-contained.

- In CLI channel + main-loop streaming, capture the agent returned by
  ``_await_agent_ready()`` and pass that into ``run_streaming`` rather
  than reading ``agent_loader.agent`` after a subsequent ``await``.
  A concurrent ``/new``/``/resume``/``/model`` could have swapped it.

- Add ``requires_agent: ClassVar[bool] = False`` to ``Command`` and
  mark ``/compact``, ``/model``, ``/channel`` as ``True``. TUI dispatch
  sites (channel and keyboard) now check ``cmd_manager.resolve(...)``
  and only wait for readiness when the command actually needs the
  agent. ``/mcp add``, ``/skills``, ``/new`` etc. no longer deadlock
  behind a failing MCP load they are meant to fix.

* fix(cli): guard sync-back, subcommand-aware gating, /model adopt-path

Three CodeRabbit findings on command dispatch:

- ``_handle_command`` unconditionally synced ``ctx.agent`` back into
  ``agent_loader``.  For non-agent commands ``ctx.agent`` is ``None``,
  so ``/threads`` / ``/mcp`` / ``/skills`` (etc.) could clobber a valid
  loaded agent — and rebind channel globals to ``None``.  Guard the
  sync on ``ctx.agent is not None``.

- ``/channel status`` and ``/channel stop`` don't touch ``ctx.agent``
  but the class-level ``requires_agent = True`` blocked them behind
  agent readiness.  Added ``Command.needs_agent(args)`` (defaults to
  ``requires_agent``) so ``/channel`` can override with subcommand
  awareness; kept the class flag for the common case.

- ``/model`` builds a new agent from scratch, it never reads the
  existing one — gating it on readiness meant a broken provider
  blocked the command that would fix it.  Flipped it to
  ``requires_agent = False`` and added ``BackgroundAgentLoader.adopt``
  so the UI can seat the replacement and supersede any in-flight
  load (the generation token keeps a late completion from clobbering
  the adopted agent).

Bonus cleanup: ``CommandManager.resolve`` now returns
``(command, args)`` so callers can invoke ``needs_agent`` without
re-implementing ``shlex`` parsing.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-22 18:22:33 +02:00
Wiktor Cupiał 28f3e81b4c feat: add /model command for changing models inside TUI/CLI (#162)
* rebase main

* fix: apply pr comments

* fix: fix critical issue

* fix: ordering /model in cli mode

* feat: refactor /model command handling and add Rich CLI support

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-04-22 14:43:21 +01:00
Xi Zhang aa3dd00409 feat: add support for session resumption with --resume flag and enhan… (#170)
* feat: add support for session resumption with --resume flag and enhance thread ID resolution

* refactor(tests): streamline help output testing for --resume flag

* feat: enhance session resume functionality with improved thread ID resolution and SQL wildcard handling

* feat: improve error handling for resume hint retrieval in interactive modes

* refactor: streamline logging for print_resume_hint failure in interactive mode

* feat: implement deferred scrolling for Markdown-heavy content in interactive mode
2026-04-21 21:26:40 +01:00
dinos 05f54334ba fix(mcp): stdio env passthrough + durable package installs (#169)
* fix(mcp): forward proxy and CA bundle env vars to stdio subprocesses

The MCP SDK's stdio transport inherits only a minimal allowlist (HOME,
PATH, USER, …) from the parent, stripping http_proxy/https_proxy and
SSL_CERT_FILE/REQUESTS_CA_BUNDLE/etc. Behind a proxy or with a custom CA
bundle, stdio MCP servers silently hang on outbound requests while the
same server over HTTP transport works. Auto-forward the proxy and cert
vars when present; user-configured env still takes precedence.

* fix(mcp): use `uv tool install` so MCP packages survive uv sync

Source installs previously used `uv pip install --python $VENV <pkg>`,
which lands in the evosci venv but is not recorded in pyproject.toml or
uv.lock. A subsequent `uv sync` (typical after `git pull`) reconciles
the venv to the lockfile and removes the MCP package, forcing users to
re-run onboard.

Prefer `uv tool install <pkg>` for the non-uv-tool install path: the
binary symlink in ~/.local/bin survives uv sync and evosci upgrades,
and the MCP server gets its own isolated env (no dep conflicts).
Verify the expected CLI entry point resolves afterward; if not (package
has no console-script), fall through to the old uv-pip path so
command-less packages still work.

The uv-tool-env path (`uv tool install evoscientist --with <pkg>`) is
unchanged — it was already durable via uv's receipt.

* fix(mcp): gate standalone uv tool install on verify_command

Previously `install_pip_package` would route every install through
`uv tool install <pkg>` when `verify_command` was None, returning
success as long as the uv subprocess exited 0. Library callers
(`evoscientist[oauth]`, `lark-oapi`, etc.) expect the package to land
in the active venv so they can import it — a standalone uv tool env
is not importable, so the import fails at the next line.

Gate the `uv tool install <pkg>` branch on `verify_command` being
set: that signals the caller wants a durable CLI binary, which is
what `uv tool install` produces. Library callers omit it and go
straight to the pip-install-into-venv path.

Also: log info messages on every fall-through so stale-binary and
entry-point-missing failure modes are debuggable, and document the
--with → standalone recovery path.

* fix(mcp): resolve MCP binaries to `uv tool dir --bin`, not `.venv/bin`

Under `uv run`, the project venv's `bin/` comes first on PATH, so
`shutil.which("arxiv-mcp-server")` returns a stale `.venv/bin/` copy
left over from an earlier install instead of the fresh symlink that
`uv tool install` just placed in `~/.local/bin`. The venv copy gets
written to mcp.yaml and is then wiped by the next `uv sync` — exactly
the failure mode the durability fix was meant to prevent.

Query `uv tool dir --bin` directly and prefer binaries found there
over `shutil.which`. Same change to the post-install verify in
`install_pip_package` so a venv shadow can't falsely short-circuit
the fallback.

* refactor(mcp): split install_pip_package into install_library + install_cli_tool

`verify_command` was doing double duty: naming the CLI binary to check
*and* signaling "this is a CLI install, use the standalone `uv tool
install` path." Callers routed library installs through the CLI branch
any time they forgot to pass it, and the resulting standalone uv tool
env wasn't importable from the active venv.

Separate the two use cases into distinct functions, each with one
install strategy per environment shape. Shared logic lives in private
`_install_with_uv_tool_env` / `_install_via_pip` helpers.

- install_library(pkg): uv-tool-env --with → pip. Never uses standalone
  `uv tool install <pkg>` (not importable from active venv).
- install_cli_tool(pkg, *, verify_command): uv-tool-env --with →
  standalone `uv tool install` → pip. `verify_command` is now required.

Callers pick the right function at the call site: registry.py picks
based on whether `entry.command` is set; onboard.py call sites all
install libraries.
2026-04-21 16:59:06 +01:00
X-iZhang 1e4c011b7c fix(assets): update wechat_group image for improved clarity 2026-04-21 12:06:49 +01:00
Xi Zhang 06822f236c feat: enhance tool result handling with tool_call_id for concurrent execution 2026-04-19 23:15:20 +01:00
Ziheng Zhang bd501cce34 fix(channel/qq): deliver HITL approval prompts reliably (#166)
* fix(channel/qq): deliver HITL approval prompts reliably

QQ approval prompts were silently dropped when the markdown send hit
a QQ server-side error (e.g. template not configured, content audit)
because the fallback path only matched TypeError / specific string
patterns, and the plain-text retry reused the already-consumed
msg_seq which QQ then rejects as duplicate.

- Consume a fresh msg_seq for the plain-text fallback send
- Recognize QQ server error codes (304014/304023/304003/40034059)
  and CN fragments ("模版"/"审核") as markdown-fallback triggers
- Promote send failure logs from debug to warning/error with
  chat_id/msg_id/seq so real-world errors can be diagnosed
- Extend test_qq_channel with a server-error-code fallback case

* style(channel/qq): apply ruff formatter to approval-delivery fix

* Fix
2026-04-19 11:14:23 +01:00
Xi Zhang 58435dba52 Release/v0.0.8 (#167)
* chore(release): update version to v0.0.8 and dependencies in project files

* feat(models): add new model entries for Claude Opus 4-7 and update version handling

* Refactor code structure for improved readability and maintainability
2026-04-18 16:49:31 +01:00
Xi Zhang f4a3617646 refactor(paths): unify global data directory to ~/.evoscientist and u… (#164)
* refactor(paths): unify global data directory to ~/.evoscientist and update related paths

* refactor(paths): update legacy session migration to respect XDG_CONFIG_HOME

* refactor(tests): clear XDG_CONFIG_HOME in legacy session migration tests for deterministic behavior
2026-04-18 15:01:22 +01:00
Xi Zhang 7c6b6755f2 fix(docs): update survey literature and macOS deployment links for accuracy 2026-04-17 15:25:23 +01:00
Xi Zhang 4b2aaea49a fix(docs): update survey literature link to point to the correct GitHub path 2026-04-17 15:20:59 +01:00
Xi Zhang 1965dd8661 Add new asset images for survey literature examples
- Added model_selection.png to illustrate model selection process.
- Added prompt.png for visual representation of prompts used in surveys.
- Added skill_selection.png to depict skill selection criteria.
2026-04-17 15:18:06 +01:00
dependabot[bot] 03b4412c9a chore(deps): bump langchain-openai in the uv group across 1 directory (#163)
Bumps the uv group with 1 update in the / directory: [langchain-openai](https://github.com/langchain-ai/langchain).


Updates `langchain-openai` from 1.1.12 to 1.1.14
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain-openai==1.1.12...langchain-openai==1.1.14)

---
updated-dependencies:
- dependency-name: langchain-openai
  dependency-version: 1.1.14
  dependency-type: direct:production
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-17 10:27:38 +02:00
Xi Zhang 210e8864f6 feat(memory): migrate MEMORY.md to global path & enhance ask-user prompts (#161)
* feat(prompt): enhance user interaction with multiple-choice and free-text questions

* refactor(paths): rename MEMORY_DIR to MEMORIES_DIR for consistency

* style(tests): format code for better readability in test cases

* feat(prompt): add validation for 'other' option in user prompt

* feat(prompt): refactor validation logic for user prompts and add skip option

* feat(style): refactor to use shared _PICKER_STYLE from interactive module
2026-04-16 15:33:45 +01:00
dependabot[bot] 7f522cb4fe chore(deps): bump langsmith in the uv group across 1 directory (#160)
Bumps the uv group with 1 update in the / directory: [langsmith](https://github.com/langchain-ai/langsmith-sdk).


Updates `langsmith` from 0.7.30 to 0.7.31
- [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases)
- [Commits](https://github.com/langchain-ai/langsmith-sdk/compare/v0.7.30...v0.7.31)

---
updated-dependencies:
- dependency-name: langsmith
  dependency-version: 0.7.31
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-16 10:38:20 +01:00
Xi Zhang 0b7c162d1b fix(skill-manager): update skill source filtering to include workspace and global tiers (#159) 2026-04-15 15:29:11 +01:00
Ziheng Zhang a2d2ddc5a2 fix(channel): avoid replaying thinking after resume (#154)
* fix(channel): avoid replaying thinking after resume

* fix(channel): relay fresh thinking after resume

* Fix

* Fix

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-15 14:05:47 +01:00
dinos 2961e5ee88 fix(minimax): correct API endpoint and add region selection (#158) 2026-04-15 14:30:26 +02:00
Xi Zhang c4237fecb8 feat(backends): rename MergedReadOnlyBackend to MergedSkillsBackend a… (#157)
* feat(backends): rename MergedReadOnlyBackend to MergedSkillsBackend and update documentation for clarity

refactor(paths): simplify ensure_dirs function to create only memory directory eagerly

fix(prompts): update skills availability description for accuracy

test(paths): adjust test to reflect skills directory creation on demand

* refactor(tests): format assertion for skills directory existence in ensure_dirs test
2026-04-15 01:52:31 +01:00
X-iZhang 64721f5967 update 2026-04-13 20:57:27 +01:00
Xi Zhang 65db3a4fcd Add status bar and compact summary widgets with context window resolu… (#152)
* Add status bar and compact summary widgets with context window resolution

- Implemented a shared status bar for CLI and TUI frontends, including helpers for managing session metrics and context windows.
- Created a `CompactSummaryWidget` for displaying manual summaries in a collapsible format.
- Introduced a `CompactingWidget` to indicate ongoing compacting processes.
- Added a base class `TimedStatusWidget` for widgets that require a timer.
- Developed context window resolution helpers to retrieve context window sizes from various model attributes.
- Enhanced tests for context window resolution and status bar functionalities, ensuring accurate behavior across different scenarios.
- Updated existing tests to cover new features and maintain code quality.

* refactor(Channel): simplify lambda function in _send_with_retry method

* feat: enhance context editing logic and improve error handling in StreamState

* refactor(Channel): streamline lambda function in _send_with_retry method

* feat: rename auto-approve option to auto-mode for unattended execution; update checkpoint queries to filter by agent name; improve compatibility validation logic

* feat: rename auto-approve option to auto-mode; update related logic and tests for improved unattended execution

* fix: correct formatting of console message for MCP server configuration status

* feat: add check for None summary_message in _apply_summarization_event to prevent errors

* feat: enhance _load_checkpoint_messages to validate message format and apply summarization event
2026-04-12 17:47:37 +01:00
Xi Zhang ff15f515cc Release/v0.0.7 (#151)
* chore(assets): update wechat_group image file

* Refactor code structure for improved readability and maintainability

* feat(backends): enhance MergedReadOnlyBackend with improved ls, grep, and glob methods

* fix(docs): update WeChat QR code image link in README files

* feat(skills): enhance skill management to support global and workspace tiers

* style: apply ruff format to skills_cmd and commands/implementation/skills

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(skills): improve uninstall_skill to prevent removal of built-in skills

* fix(docs): update skill installation documentation for clarity on global and user directories

* fix(skills): enhance uninstall_skill to validate skill directory before removal

* fix(skills): improve error handling in install_skill and uninstall_skill for directory creation and validation

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-10 20:04:05 +01:00
Xi Zhang d5b982c980 fix(ccproxy): update Responses API handling and patch system role con… (#149)
* fix(ccproxy): update Responses API handling and patch system role conversion

* fix(ccproxy): streamline _agenerate method in system to developer patch

* fix(ccproxy): improve handling of None output in Codex compatibility patch
2026-04-09 12:51:05 +02:00
Ziheng Zhang 3e493233cf feat(channels): simplify debug tracing and add serve debug mode (#143)
* feat(channels): simplify debug tracing and add serve debug mode

* Fix

* fix(channels): remove serve loop patch
2026-04-09 12:47:48 +02:00
dependabot[bot] 011dffdabc chore(deps): bump the uv group across 1 directory with 2 updates (#148)
Bumps the uv group with 2 updates in the / directory: [cryptography](https://github.com/pyca/cryptography) and [langchain-core](https://github.com/langchain-ai/langchain).


Updates `cryptography` from 46.0.6 to 46.0.7
- [Changelog](https://github.com/pyca/cryptography/blob/main/CHANGELOG.rst)
- [Commits](https://github.com/pyca/cryptography/compare/46.0.6...46.0.7)

Updates `langchain-core` from 1.2.25 to 1.2.28
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain-core==1.2.25...langchain-core==1.2.28)

---
updated-dependencies:
- dependency-name: cryptography
  dependency-version: 46.0.7
  dependency-type: indirect
  dependency-group: uv
- dependency-name: langchain-core
  dependency-version: 1.2.28
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-09 00:09:19 +01:00
Xi Zhang 4f11de23a2 fix(llm): patch _stream/_astream for OpenAI-compatible content flattening (#147)
* fix(llm): patch _stream/_astream for OpenAI-compatible content flattening

_patch_openai_compat_content() only patched _generate/_agenerate but
EvoSci CLI uses streaming paths. This extends the content flattening
to _stream/_astream so strict OpenAI-compatible relays receive plain
string content during streaming calls.

Closes #142

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): use asyncio.run() instead of pytest-asyncio for CI compat

CI does not have pytest-asyncio installed, so async tests must use
asyncio.run() wrapper instead of @pytest.mark.asyncio decorator.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): use @pytest.mark.anyio for async tests (CI compat)

CI does not have pytest-asyncio. Use @pytest.mark.anyio consistent
with existing async tests in the project.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 20:24:13 +01:00
Peidong Yang 066f36fa13 feat: add Moonshot and Kimi Coding Plan as LLM providers (#128)
* feat: add Moonshot and Kimi Coding Plan as LLM providers

Add two new providers for Moonshot AI:
- `moonshot`: OpenAI-compatible direct API (api.moonshot.cn/v1) with
  kimi-k2.5, kimi-k2-thinking, moonshot-v1-auto/128k/32k/8k models
- `kimi-coding`: Anthropic-compatible Kimi Coding Plan endpoint
  (api.kimi.com/coding/) with User-Agent header for compatibility

Both providers disable thinking to avoid multi-turn tool calling
errors caused by LangChain dropping reasoning_content from history.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: update Moonshot thinking comment and add provider assertions

- Add clarifying comment for disabling thinking on all Moonshot models
- Add moonshot and kimi-coding assertions to test_entries_has_all_providers

* fix: exclude Moonshot and Kimi Coding from content patch

Tested and verified both APIs support standard list content format:
- Moonshot (OpenAI-compatible): supports list content, no patch needed
- Kimi Coding (Anthropic-compatible): supports list content, no patch needed

Only apply _patch_openai_compat_content to strict providers like DeepSeek.

* fix: set _original_provider in routed provider branches

Ensure _original_provider is set before provider is reassigned to
'openai' or 'anthropic', so the no-patch exclusion for Moonshot
and Kimi Coding works correctly.

* style: translate Moonshot comments to English

* style: translate comment to English to fix ruff lint error

* merge: resolve conflicts

* chore: revert uv.lock and translate Chinese comments to English

Revert unrelated uv.lock dependency changes and replace Chinese code
comments with English for codebase consistency per review feedback.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: fix ruff format for models.py

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: ypd <ypd@ypddeMac-mini.local>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: Xiaohui Yan <xhcloud@gmail.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-04-08 19:53:49 +01:00
JackyFan 632159261c Fix: Support XDG_CONFIG_HOME for sessions.db on Windows with non-ASCII usernames (#102)
* fix: support XDG_CONFIG_HOME for sessions.db on Windows

Fixes SQLite database opening failure on Windows systems with non-ASCII
usernames (e.g., Chinese characters). The get_db_path() function now
supports the XDG_CONFIG_HOME environment variable, consistent with
get_config_dir() in settings.py.

Closes #101

* fix: auto-resolve Windows Unicode path for sqlite3 via 8.3 short path

Refactor get_db_path() to reuse get_config_dir() (XDG_CONFIG_HOME
support) and add _to_short_path() helper that converts the config
directory to its Windows 8.3 short form via GetShortPathNameW. This
automatically resolves sqlite3 failures on Windows systems with
non-ASCII usernames (e.g., Chinese characters) without requiring
manual environment variable configuration.

The short-path conversion is best-effort: it targets the directory
(which exists after mkdir) rather than the db file, and falls back
gracefully on non-Windows, non-NTFS, or when 8.3 naming is disabled.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 18:50:42 +01:00
Ziheng Zhang 3647ff1afa fix: preserve QQ markdown formatting and newline rendering (#144)
* fix: preserve QQ markdown formatting

* fix: remove stale qq trace fallback hook

* fix: narrow qq markdown fallback handling

* style: format qq channel with ruff
2026-04-08 13:52:45 +01:00
Xiaohui Yan e89b71aa60 feat(cli): add --debug flag for verbose logging in serve mode (#141)
* feat(cli): add --debug flag for verbose logging in serve mode

* feat(cli): add log_level config field with priority over env var

Replace dead `debug` parameter in `main()` with a proper `log_level`
config field in EvoScientistConfig. Enables `EvoSci config set log_level
debug` with priority: config file > EVOSCIENTIST_LOG_LEVEL env var.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-08 13:23:10 +08:00
Ziheng Zhang 729cab11be Fix late channel response delivery after timeout (#140)
* Fix late channel response delivery after timeout

* Fix
2026-04-07 15:48:12 +01:00
X-iZhang 080a5c06f3 feat: enhance OpenRouter support with additional reasoning handling and model entries 2026-04-03 17:14:19 +01:00
X-iZhang 117750eda7 chore: update wechat group image asset 2026-04-03 15:44:40 +01:00
Xi Zhang 028dbe79d3 Release/v0.0.6 (#138)
* chore: clean up empty code change sections in the changes log

* feat: add adaptive tools and context editing features to README
2026-04-03 15:20:45 +01:00