Commit Graph

547 Commits

Author SHA1 Message Date
Xi Zhang c02be519f6 feat(dangerous-mode): implement real-filesystem access with safety ch… (#276)
* feat(dangerous-mode): implement real-filesystem access with safety checks

- Introduced a 'dangerous mode' allowing the agent to operate on the real filesystem.
- Updated command validation to bypass path confinement while enforcing a blocklist for privileged commands.
- Added warnings and guidelines for users when operating in dangerous mode.
- Enhanced configuration to support dangerous mode and ensure it implies auto-approval.
- Updated tests to verify the behavior of commands and configurations in dangerous mode.

* feat(dangerous-mode): enhance logging and environment management for dangerous mode

* feat(dangerous-mode): improve handling of dangerous mode with environment flags and enhance test isolation
2026-06-10 18:19:13 +01:00
X-iZhang fc05b5eed2 fix(tests): ensure _check_npx is True in _patch_all_questionary to prevent extra prompts on slow runners 2026-06-10 16:11:49 +01:00
X-iZhang c438966b28 fix(tests): mock _ensure_npx in TestStepSkills to prevent integration issues on headless Windows CI 2026-06-10 15:59:10 +01:00
houren Antony d4f1fbd110 ci: add windows-latest to test matrix + fix 11 cross-platform test bugs (#271)
* ci: add windows-latest to test matrix + fix 11 cross-platform test bugs

The test workflow ran on ``ubuntu-latest`` only. Per the issue's
first bullet — the maintainer's explicit #1 priority — add
``windows-latest`` to the matrix so the manager and related
modules are exercised on Windows on every PR.

The matrix addition surfaces 18 pre-existing Windows-only test
failures. Without fixes the new leg would be 18+ reds from
day one and the matrix would just produce a wall of
``fail-fast`` noise. This PR fixes 11 of them; each fix is
a real (cross-platform) bug, not a Windows-specific hack —
most were already flagged by CodeRabbit on PR #236 but never
acted on. The remaining 4 failures need code refactors
(``os.killpg`` → ``psutil`` in ``background.py``,
``convert_virtual_paths_in_command`` Windows-aware quoting,
tilde expansion) that are documented as out-of-scope
follow-ups below.

## What changed

* ``.github/workflows/test.yml``
  - ``os: [ubuntu-latest, windows-latest]`` → 2 OS × 2 Python
    = 4 cells.
  - ``fail-fast: false`` so one bad cell doesn't cancel the
    rest while the Windows leg is being brought up. Removable
    in a future PR once the suite is fully green.

* ``tests/test_backends.py``
  - Hard-coded ``"python3"`` → ``{sys.executable}`` in 7
    test commands. Windows has no ``python3`` on PATH; using
    ``sys.executable`` is portable and matches what CodeRabbit
    flagged on PR #236.
  - Strict string comparisons → ``shlex.split`` round-trip in
    5 resolver tests. ``shlex.quote`` adds single quotes
    around backslash paths on Windows, which broke the
    direct ``==`` compare.
  - Cross-platform suffix checks in 2 path-resolution tests
    (``Path(resolved).parts[-2:]`` instead of
    ``str(resolved).endswith("src/main.py")``).
  - ``mkdir -p`` → ``sys.executable -c "import os;
    os.makedirs(...)"`` in the cwd-sanitization test.
  - ``skipif(sys.platform == "win32")`` on 3 e2e tests that
    hit the underlying ``shlex.quote`` + ``cmd.exe`` quoting
    bug (real, separate issue).

* ``tests/test_sessions.py``
  - ``test_uses_data_dir``: check ``.evoscientist`` in the
    long path form (via ``Path.resolve()``) rather than the
    short-path form ``get_db_path`` returns on Windows.

* ``tests/test_mcp_client.py``
  - ``endswith("python")`` → ``Path(result).stem.lower()`` so
    ``python.EXE`` matches on Windows.
  - ``endswith("npx")`` also accepts ``npx.cmd`` so the npm
    shim on Windows matches.

## Out of scope (follow-up issues to file)

* ``os.killpg`` doesn't exist on Windows
  (``EvoScientist/background.py:248``) — 3 background tests
  fail. Real fix is the same ``psutil`` walk pattern PR #200
  shipped in ``langgraph_dev/manager.py``.
* Tilde expansion in file mentions.
* Windows-aware shell quoting in
  ``convert_virtual_paths_in_command``.
* Path conventions (``~/.config/evoscientist/`` vs
  ``%APPDATA%\EvoScientist``) — needs design discussion +
  ``platformdirs`` migration.
* Cross-module audit of
  ``EvoScientist/tools/execute.py``,
  ``EvoScientist/ccproxy_manager.py``,
  ``EvoScientist/config/onboard.py``.

Closes #207 (step 1 only — CI matrix + the easy test
fixes; remaining bullets tracked separately).

* fix: cross-platform compatibility for Windows CI runners

- background.py: replace POSIX-only os.killpg/os.getpgid with
  cross-platform _kill_process_tree() helper. On Windows falls back
  to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX
  keeps existing os.killpg logic.

- test_backends.py: replace mkdir -p shell execution in
  test_literal_workspace_path_replaced with preprocessing-boundary
  assertion (patch LocalShellBackend.execute, capture command,
  assert workspace path was rewritten to ./). Avoids POSIX-only
  mkdir -p on Windows runners.

- test_file_mentions.py: monkeypatch USERPROFILE on Windows so
  ntpath.expanduser() resolves ~ to tmp_path even when HOME is
  unset on CI runners.

* fix(test): cross-platform sleep/true commands for Windows CI

Replace POSIX-only sleep/true with module-level helpers that use
ping -n / cmd /c on Windows. Also fix python3 -> sys.executable
in the non-timeout recovery test.

- test_background.py: 7 sleep/true fixes
- test_background_middleware.py: 6 sleep/true fixes
- test_backends.py: 4 sleep fixes + 1 python3 fix

2318 passed, 0 failed on Windows.

* fix(test): use shell-portable double quotes for python -c on Windows

cmd.exe does not treat single quotes as string delimiters, so
-c 'raise SystemExit(1)' was passed with literal quotes on Windows.
Switch to double quotes which work on both cmd.exe and POSIX sh.

* fix: use psutil for Windows process tree kill + avoid sys.executable under uv

- background.py: replace Popen.terminate()/kill() with psutil-based
  process tree walking on Windows. TerminateProcess does NOT cascade
  to grandchildren; psutil.Process.children(recursive=True) ensures
  the entire tree is signaled.

- test_backends.py: replace sys.executable with 'python' in sandbox
  execute() calls. Under uv, sys.executable is under the workspace
  and gets rewritten to ./ by prepare_sandbox_command, breaking
  Linux CI. The plain 'python' command resolves correctly in any
  activated venv.

* fix: broaden try/except in _kill_process_tree to cover proc.children()

If the process exits between Process(popen.pid) and children(recursive=True),
the children call raises an uncaught exception escaping stop(). Move it inside
the existing try/except block.

* fix: narrow exception to ProcessLookupError in POSIX _kill_process_tree

OSError is too broad — would silently swallow EPERM on SIGKILL, leaving
the process alive when we report it as stopped. Match original behavior
which only caught ProcessLookupError (process already gone).

* style: ruff format test_backends.py

* ci: trigger re-run for flaky prompt_toolkit test

* style: fix ruff check (import order + RUF005 unpacking)

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-10 15:43:18 +01:00
X-iZhang cf5e0dd3bd feat(models): add support for 'claude-fable-5' mode 2026-06-10 00:51:29 +01:00
X-iZhang de3785346f feat(docs): update README to reflect Desktop WebUI changes and add demo video 2026-06-09 18:04:34 +01:00
houren Antony 2dc1e227eb fix(langgraph-dev): rotate langgraph_dev.log when it exceeds 50MB (#270)
* fix(langgraph-dev): rotate langgraph_dev.log when it exceeds 50MB

``_LOG_FILE`` (``~/.config/evoscientist/langgraph_dev.log``) was
opened in ``start_langgraph_dev`` with plain ``"ab"`` and never
rotated, so it grew unbounded over weeks/months of heavy use —
especially when chatty MCP servers spawned by langgraph dev
filled it, or when failure paths produced stack traces.

Implement the recommended option 1 from #209: filesize-based
rollover. When the active log exceeds 50MB on the next
``start_langgraph_dev`` invocation, rename it to
``langgraph_dev.log.1`` (overwriting any existing backup) via
``os.replace`` and start fresh. Single-backup policy keeps the
disk footprint bounded at roughly 2x threshold.

Rotation is best-effort: ``_rotate_log_if_needed`` logs and
swallows OSError so a permission error or racing rename can't
block langgraph dev from starting. The next ``start`` invocation
will try again — worst case the log grows for one more session.

Options 2 (timestamped per-session + 7-day sweep) and 3
(``RotatingFileHandler`` + pipe) are explicitly NOT done — option
1 is simplest, no async machinery, matches the issue's
recommendation.

Closes #209

* test(langgraph-dev): redirect _PID_DIR in rotate integration test

Address CodeRabbit review comment on #270: the
``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open``
test patched only ``_LOG_FILE`` to a tmp path, but
``start_langgraph_dev`` also calls ``_PID_DIR.mkdir(...)`` as part
of its prelude, which would create a real directory under
``~/.config/evoscientist/`` on a dev machine. Redirect
``_PID_DIR`` to ``tmp_path / "pids"`` too so the test stays
fully isolated. Add a final assertion that ``pid_dir.is_dir()``
holds, proving the function reached past the mkdir call.

* refactor(langgraph-dev): bundle runtime paths into LanggraphRuntimePaths

@din0s review follow-up on #270: the previous test isolation patched
only ``_LOG_FILE`` (and after a second round, ``_PID_DIR``), but
``start_langgraph_dev`` still touches 5 distinct on-disk paths. Patching
any subset of those still leaves the others pointing at the user's real
``~/.config/evoscientist/`` — exactly the case that produced the
"Port 6174 cannot be bound after waiting 60s" symptom on the
reviewer's machine.

Replace the five free-floating module-level constants
(``_PID_DIR`` / ``_PID_FILE`` / ``_LOG_FILE`` / ``_WORKSPACE_SIDECAR``
/ ``_FILE_LOCK_PATH``) with a single ``LanggraphRuntimePaths`` frozen
dataclass exposed as a module-level ``RUNTIME`` instance. Production
code accesses ``RUNTIME.pid_file`` etc.; tests can now substitute the
*whole* bundle in one assignment:

    monkeypatch.setattr(
        manager, "RUNTIME",
        manager.LanggraphRuntimePaths.for_directory(tmp_path / "runtime"),
    )

The classmethod ``for_directory(pid_dir)`` builds an isolated bundle
rooted at a single dir, so the test author doesn't spell out every
path field. Tests that only care about one field (e.g. pid_file
during the stale-process kill path) use
``dataclasses.replace(manager.RUNTIME, pid_file=X)`` — frozen
dataclass-friendly, no need to enumerate the other four fields.

The dataclass's docstring records the migration rationale (the old
five-name layout invited inconsistent patches).

External callers of the old constants updated:
- ``EvoScientist/deploy/server.py`` and ``webui.py`` now import
  ``RUNTIME`` and use ``RUNTIME.log_file`` for the on-screen log
  path hint. The other imports they had (``_DEFAULT_PORT``,
  ``_is_port_occupied``, ``_read_workspace_sidecar``) are still
  module-level functions/values, untouched.

Test updates:
- ``tests/test_langgraph_manager.py``: ``patch.object(manager, "_XXX",
  X)`` patterns now go through ``dataclasses.replace(manager.RUNTIME,
  xxx=X)``; the ``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open``
  test (from the previous #270 review iteration) uses
  ``for_directory`` for one-shot isolation.
- ``tests/test_langgraph_dev_workspace_sidecar.py``: each test now
  goes through a tiny ``_isolated_runtime(monkeypatch, tmp_path)``
  helper that calls ``for_directory``.
- ``tests/test_langgraph_dev_deploy_mode.py``: same ``for_directory``
  swap.

No production behavior change. All ``langgraph_dev``-side tests
(``test_langgraph_manager.py`` 26/26, ``test_langgraph_dev_workspace_sidecar.py``
14/14, ``test_langgraph_dev_deploy_mode.py`` 14/14, ``test_cli_deploy.py``
18/18 — which indirectly exercises deploy/server.py and deploy/webui.py
imports) pass. Full-project test count unchanged from baseline; the
remaining 22 Windows-only pre-existing failures (test_background
``os.killpg``, test_file_mentions tilde, mcp_client ``shutil.which``,
test_sessions 8.3 short path) are documented as out-of-scope for #207.

* style: apply ruff format to langgraph_dev test + module files

CI lint check on #270 failed:

  Run ruff format --check .
  Would reformat: EvoScientist/langgraph_dev/manager.py
  Would reformat: tests/test_langgraph_manager.py

Plus two test files touched by the prior consolidation commit that
``ruff format`` hadn't seen yet:

  tests/test_langgraph_dev_deploy_mode.py
  tests/test_langgraph_dev_workspace_sidecar.py

Just formatting. No logic change. All 75 refactor-related tests pass.

* fix(test): use for_directory for full path isolation + patch _can_bind_port to skip real socket ops

Two fixes for TestStartLanggraphDevRotatesLog:

1. Replace dataclasses.replace(manager.RUNTIME, ...) with
   LanggraphRuntimePaths.for_directory(pid_dir) so pid_file,
   workspace_sidecar, and lock_file are also temp-rooted
   (prevents leak to ~/.config/evoscientist/).

2. Monkeypatch _can_bind_port to always return True so the
   bind-poll loop in _wait_for_port_bindable passes immediately
   without touching real sockets (fixes 60s timeout on machines
   where port 6174 is already in use).

* fix: cross-platform compatibility for Windows CI runners

- background.py: replace POSIX-only os.killpg/os.getpgid with
  cross-platform _kill_process_tree() helper. On Windows falls back
  to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX
  keeps existing os.killpg logic.

- test_backends.py: replace mkdir -p shell execution in
  test_literal_workspace_path_replaced with preprocessing-boundary
  assertion (patch LocalShellBackend.execute, capture command,
  assert workspace path was rewritten to ./). Avoids POSIX-only
  mkdir -p on Windows runners.

- test_file_mentions.py: monkeypatch USERPROFILE on Windows so
  ntpath.expanduser() resolves ~ to tmp_path even when HOME is
  unset on CI runners.

* refactor(test): add runtime_paths fixture to isolate manager.RUNTIME

Adds a reusable fixture that monkeypatches manager.RUNTIME to a
temp-rooted LanggraphRuntimePaths.for_directory(). Tests that need
specific fields can still dataclasses.replace(runtime_paths, ...)
but the baseline is always temp-isolated, preventing leaks to
~/.config/evoscientist/.

Updated test_langgraph_dev_deploy_mode.py, test_langgraph_dev_workspace_sidecar.py,
and test_langgraph_manager.py to use the fixture, consolidating sequential
lock_file + pid_dir patches into single dataclasses.replace calls.

* Revert "fix: cross-platform compatibility for Windows CI runners"

This reverts commit eb025d24af32e195a982cd40f6d70dba885c4019.

* style: ruff format conftest.py

* fix: address review issues in log-rotation + runtime paths

- Use for_directory(tmp_path/pids) as base in ensure_langgraph_dev tests
  so pid_file/log_file are co-located with pid_dir, not split across paths
- Remove unused runtime_paths param from test_no_existing_file_is_noop
- Replace manager.RUNTIME with runtime_paths in two sidecar tests
- Use for_directory(DEFAULT_PID_DIR) instead of explicit construction
- Fix stale _LOG_FILE reference in TestRotateLogIfNeeded docstring

* style: ruff format test files

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-09 15:55:44 +01:00
dinos cd2baa9588 feat(models): add opt-in prompt caching support for anthropic via openrouter (#272)
* feat(models): add opt-in prompt caching support for anthropic via openrouter

* chore: don't coerce model_kwargs to dict
2026-06-09 15:13:49 +01:00
Ziheng Zhang 4b6a969df2 refactor(agent): make create_cli_agent(config=, chat_model=) pure (#267)
* refactor(agent): make create_cli_agent(config=, chat_model=) pure

Re-applies the #183 purity refactor on top of the observation-memory
lifecycle that landed in #259, integrating the two cleanly.

create_cli_agent gains a pure path: when both `config` and `chat_model`
are passed it builds the agent entirely from locals and writes none of
the cached module globals (`_config`, `_chat_model`, `_chat_model_key`,
`_EvoScientist_agent`). `/model` commits the switch via
`set_active_config` / `set_chat_model_instance` only after a successful
build, so a failed rebuild leaves the session on the original model
(replaces the old snapshot/restore rollback).

Supporting changes:
- Extract `set_active_config` (write-half of `_ensure_config`),
  `_apply_env_from_config`, `_build_chat_model`, and
  `set_chat_model_instance`.
- Thread `cfg` / `chat_model` through `_get_default_middleware`,
  `_build_base_kwargs`, `load_mcp_and_build_kwargs`,
  `_maybe_swap_async_subagents`, and `_inject_subagent_middleware` so the
  pure path never falls back to the global-writing `_ensure_config()` /
  `_ensure_chat_model()`.
- Integrate with #259's memory middleware: subagent context-editing
  middleware binds the threaded `chat_model`, and the configured system
  prompt / memory controls read the threaded `cfg` (new threading vs the
  original #183, required because #259 made these paths read config).
- Consolidate `cfg` resolution to one `cfg if cfg is not None else
  _ensure_config()` at the top of each kwargs builder, matching the
  pattern already used in the other config-aware helpers.

* fix(agent): keep pure tool selector off global cache

* fix(model): apply config switch in place to preserve reference integrity

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-06-08 18:38:13 +01:00
dinos 8bb1d6c0e3 refactor(stream): langgraph streaming v3 (#268)
* refactor(stream): langgraph streaming v3

* fix: address CR comments

* chore(stream): add success field to state

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-08 17:14:25 +01:00
Wiktor Cupiał ac052bbb4c feat: free-scrolling (#262)
* feat: free-scrolling

* fix: anchor

* fix: textual private vars
2026-06-08 16:23:46 +02:00
Xi Zhang b4ffb35d71 feat(langgraph): add --no-reload option to start_langgraph_dev 2026-06-08 01:35:02 +01:00
Xi Zhang 63969b596d Release/v0.1.4 (#266)
* feat(middleware): reposition code interpreter middleware in the stack

* feat(models): add qwen3.7-plus model entry and update context window comment

* feat(models): add qwen3.7-max and qwen3.7-plus model entries for DashScope

* feat(auxiliary): implement auxiliary model support for background tasks and tool selection

- Added auxiliary model configuration to EvoScientistConfig.
- Introduced _ensure_auxiliary_chat_model function to manage auxiliary model instances.
- Updated onboarding steps to include auxiliary model selection.
- Modified middleware to route tool selection to the auxiliary model when applicable.
- Enhanced tests to cover auxiliary model functionality and configuration.

* feat(steps): update UI backend selection options and descriptions

* Refactor code structure for improved readability and maintainability

* feat(patches): implement OpenRouter response reasoning item stripping to prevent multi-turn errors

* feat: update version to v0.1.4 in badges, README, and pyproject.toml; adjust skill counts in steps.py

* feat(config): add auxiliary model and provider environment variables to test setup
2026-06-07 00:52:59 +01:00
Eliot Drizzle 3563c1d94f Update source for Scientific Skills in steps.py (#265) 2026-06-06 14:53:14 +01:00
dinos 92d95dee68 feat(memory): add observation memory lifecycle (#259)
* feat(memory): add observation memory lifecycle

Add file-backed observation memory with deterministic markdown records,
structured record_observation tooling, startup indexing, and
profile/observation prompt guidance.

Launch post-turn and post-subagent EvoMemory workers through LangGraph
dev so completed runs can update profile memory, save durable
observations, and write subagent execution summaries without blocking
the active agent.

Wire memory middleware into the main agent, subagents, async graphs, TUI
status reporting, worker activity accounting, and observation-aware
research prompts, with regression coverage for storage, lifecycle
scheduling, graph registration, status display, and stream reset
behavior.

* fix(cli): sync background agent server on resume

Resume flows now need to keep the LangGraph dev background server
aligned with the active workspace even when async subagents are
disabled. EvoMemory workers use that server too, so gating resume-time
sync on enable_async_subagents could leave workers pinned to the launch
workspace after resuming a thread from another workspace.

Run workspace sync unconditionally for Rich CLI and Textual resume
paths, while preserving WorkspaceMismatchError handling so failed sync
aborts the resume before mutating the active thread or workspace.

Propagate aborted resume callbacks through the command UI so
channel-issued /resume commands do not send false success or history
output. Channel slash dispatch now treats CommandManager-caught command
errors as command errors and skips completion hooks for those failed
commands.

Add regression coverage for disabled async subagents, callback aborts,
and channel command error reporting.

* fix(cli): prepare serve resume workspace before adopting

Load the resumed workspace agent and sync the background server as a
single pre-adoption step. Restore the previous active workspace if
preparation fails so serve mode keeps using the old session
consistently.

* fix(memory): untrack abandoned worker status watches

Stop treating watcher shutdown as confirmed worker completion. Terminal
worker statuses still count memory deltas, while poll failures or
watcher setup failures now remove the active run without crediting
partial outputs.

* fix(cli): report channel command failures accurately

Treat command_error as a None sentinel so empty error strings still
fail, and let TUI resumes continue only on non-mismatch
background-server sync failures while reporting degraded mode.

* fix(stream): clear memory counters for resume streams

Reset completed-memory counters for every new agent stream, including
Command-based HITL and resume streams, so saved-memory indicators do not
leak across turns.

* docs(tools): make observation recording guidance conditional

Clarify that agents should call record_observation only when the
observation tool is available, preserving the existing durability and
usefulness criteria.

* feat(config): add controls for profile and observation memory

Add config flags for profile memory, observation memory, observation
writer placement, and background memory workers.

Wire the controls through main agents, subagents, EvoMemory middleware,
and memory lifecycle workers so observation writes can be assigned to
the live agent, subagent worker, both, or neither. Keep turn memory
workers profile-only and make prompts reflect the available observation
read/write paths. Skip langgraph dev startup when neither async
subagents nor memory workers need the background server.

Add coverage for config parsing, prompt gating, middleware wiring, and
worker tool availability.

* test(cli): include memory defaults in serve config stubs

* fix(memory): offload async worker launch blocking calls

Run the langgraph-dev health check and memory-output snapshot in worker
threads from the async EvoMemory launcher so it does not block the event
loop.

* chore(memory): harden turn worker subagent guardrail

* chore(memory): refresh profile context per request

* fix(memory): offload async profile file reads

* fix(memory): offload async worker completion accounting
2026-06-05 15:11:20 +01:00
dependabot[bot] a27e5230c7 chore(deps): bump starlette in the uv group across 1 directory (#261)
Bumps the uv group with 1 update in the / directory: [starlette](https://github.com/Kludex/starlette).


Updates `starlette` from 1.0.0 to 1.0.1
- [Release notes](https://github.com/Kludex/starlette/releases)
- [Changelog](https://github.com/Kludex/starlette/blob/main/docs/release-notes.md)
- [Commits](https://github.com/Kludex/starlette/compare/1.0.0...1.0.1)

---
updated-dependencies:
- dependency-name: starlette
  dependency-version: 1.0.1
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-04 20:05:08 +01:00
dependabot[bot] 5c8975bc28 chore(deps): bump aiohttp in the uv group across 1 directory (#260)
---
updated-dependencies:
- dependency-name: aiohttp
  dependency-version: 3.14.0
  dependency-type: direct:production
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-04 11:06:05 +02:00
Zhen-Yi Zhou cedb3aa744 Fix ssh remote path handling in sandbox (#242)
* Fix ssh remote path handling in sandbox

* Address ssh remote command review feedback

* Format backend files with ruff

* Narrow SSH remote command handling

* Narrow SSH preprocessing to single-quoted remote args

* Address remaining SSH preprocessing review feedback

* Tighten SSH executable recognition

* Recognize only literal ssh wrapper

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-03 17:55:27 +01:00
X-iZhang 044a85ccd2 Update ResearchClawBench ranking details in README files 2026-06-03 17:35:10 +01:00
Wanghan Xu 0198e50e7f Add ResearchClawBench ranking news (#257) 2026-06-03 14:59:47 +01:00
X-iZhang faea53be52 chore: bump version to v0.1.3 2026-06-03 01:37:16 +01:00
Xi Zhang 9cffe9d457 Enhance multimodal handling in LLM model (#256)
* Enhance multimodal handling in LLM model

- Updated `_flatten_message_content` to preserve media blocks (images, files) while flattening text content.
- Introduced `_sanitize_messages` to manage media hoisting for tool messages, ensuring compatibility with OpenAI APIs.
- Modified `_patch_openai_compat_content` to accommodate new media handling logic, including retry mechanisms for media errors.
- Added comprehensive tests for media preservation, including various scenarios with images, files, and unsupported media types.

* fix: preserve order of text and media blocks in message flattening

* test: add tests for _strip_media_types to ensure position preservation and deduplication
2026-06-03 01:06:46 +01:00
dinos d348076f40 Add runtime context middleware (#255) 2026-06-02 18:47:26 +01:00
dinos 9285c6dad8 Migrate memory middleware to profile files (#253)
* feat(memory): migrate to profile memory files

* chore(stream): read profile headings from templates

* fix(display): keep assistant responses if response_text has started

* fix(memory): do not treat failed bootstraps as profile creation

* chore(memory): unlink blank legacy memory

* fix(memory): resolve project_id once

* fix(memory): preserve unreadable profile files

* chore(tui): render streamed narration inline with tool timeline

Update the TUI streaming timeline so assistant text emitted before or
between tool calls is rendered inline where it occurs, rather than being
kept as a single answer bubble above or below the tools.

If the model begins an assistant response and then emits another tool
call, the provisional response is converted into inline narration before
that tool. The final assistant message then renders only the remaining
response suffix, avoiding duplicate text in the completed transcript.

Stop/cancel handling now preserves any active inline narration, appends
the visible stopped marker only to the remaining displayed segment, and
still returns the full normalized stopped response for channel callers.

Completed tools continue to collapse while long runs are active, but
expand again when the turn reaches a final state so the completed
transcript shows the full tool timeline.

* fix(stream): preserve narration around tool timelines

Keep assistant narration attached to the tool call that follows it
instead of folding all streamed text into the final answer block.

Track narrated response segments in stream state, render them before
their corresponding regular or task tool entries, and keep final answers
limited to the response suffix that has not already been shown inline.
Preserve narration across normal completion, stop/error final frames,
sub-agent task calls, and collapsed live tool summaries.

Add regression coverage for pending tools, completed tools, sub-agent
task delegations, collapsed completed/running tool summaries, and final
stop frames.

* fix(tui): finalize inline narration transitions

* test(memory): use canonical project id helper
2026-06-02 18:21:14 +01:00
X-iZhang b56234317a fix: restrict textual version to avoid CJK input issues on iTerm2 2026-06-02 11:58:12 +01:00
X-iZhang 565d9647ac chore: update version to v0.1.2 in project files and badges 2026-06-02 00:32:09 +01:00
X-iZhang d53bfa35c5 feat: update MiniMax model entries and context window for M3 variant 2026-06-02 00:00:52 +01:00
Xi Zhang fbd1d709ca feat: add WebUI mode support with related configuration and onboarding (#252)
* feat: add WebUI mode support with related configuration and onboarding steps

* feat: enhance WebUI port configuration to prevent conflicts with backend port

* feat: add support for fresh interactive session detection in WebUI
2026-06-01 12:01:08 +01:00
Xi Zhang 3ce6523faf fix: update deepagents and langchain versions (#251)
* fix: update deepagents and langchain versions; enhance _reduce_messages_delta handling for None state

* fix: update langchain version constraint to >=1.3 in pyproject.toml and uv.lock
2026-05-31 22:12:13 +01:00
Xi Zhang a13904185d Feat/sandbox execute timeout (#243)
* feat: implement configurable sandbox execute timeout and enhance recovery instructions

* feat: add background process management tools and middleware for sandbox execution

* feat: enhance background process management with completion notifications and deduplication

* feat: enhance sandbox execution timeout validation and update related messages

* feat: enhance background process management with thread-specific completion notifications and HITL approval handling

* test: assert completion notification waits for process finish timestamp
2026-05-31 15:11:25 +01:00
Ziheng Zhang 2364e6b130 fix(cli): forward async-notifier replies back to originating channel (#244)
* fix(cli): forward async-notifier replies back to originating channel

When PR #214's auto-notifier fires a synthetic agent turn after a
channel-originated conversation, the synthesized response only rendered
to the local CLI/TUI — the channel user (iMessage etc.) saw nothing
and had to manually re-prompt to find out what happened.

Adds a per-thread channel-origin registry in cli/channel.py and wires
the three notifier paths (Rich CLI / TUI / serve) to publish the final
response back via bus.publish_outbound when the originating thread was
started by a channel turn. Publish is fire-and-forget (scheduled on the
bus loop + done-callback for failure logging) so the notifier turn
doesn't block on the asyncio / textual event loop.

The registry is cleared on /new and /resume rotation so stale entries
don't accumulate.

* fix(cli): address review feedback on channel-origin forwarding

Follow-up to the review on #244 (din0s, X-iZhang):

- Guard the /resume origin cleanup on a real thread change in Rich CLI
  and TUI (serve mode already did via thread_changed). Resuming the
  already-active thread no longer wipes its still-live origin, which
  would otherwise silently drop a later async-notifier forward — the
  exact gap this PR closes.
- Re-bind the now-current thread to its channel after a channel-issued
  /new or /resume slash command (which rotates the thread inside the
  dispatch), so notifier turns on the rotated thread still forward.
- Guard the publish done-callback against a cancelled future, whose
  .exception() raises CancelledError (rather than returning it) on
  bus-loop teardown, so the intended warning still logs.
- Mirror the normal reply path's manager.record_message(channel, "sent")
  for forwarded notifications so per-channel stats stay accurate.
- Print the closing "[channel: Replied to ...]" line in all three
  notifier paths (Rich CLI / TUI / serve) when a forward actually
  happened, so the forwarded block reads as terminated on screen.

Adds test_publish_records_sent_metric. ruff clean; notification-origin
suite (10) + related channel/CLI/serve suites (728) pass.

* fix(cli): store sender information separately from chat_id in channel origin

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-05-31 14:43:42 +01:00
X-iZhang 721a03c25b feat: update model version from claude-sonnet-4-5 to claude-sonnet-4-6 and related adjustments 2026-05-29 00:40:46 +01:00
Xi Zhang f75bfcda51 Add onboarding wizard with style and validation components (#241)
* Add onboarding wizard with style and validation components

- Introduced `style.py` for shared visual elements used in the onboarding wizard.
- Created `validators.py` for input validation, including integer and choice validators, and API key validation functions for various providers.
- Implemented `wizard.py` as the entry point for the onboarding process, managing user prompts and configuration steps.
- Added progress rendering and autosave functionality to enhance user experience during the onboarding process.

* feat(onboarding): enhance validation and configuration for onboarding wizard

- Added validation for UI backends, workspace modes, and providers in the onboarding command.
- Updated channel definitions to include secret field handling for sensitive tokens.
- Improved user prompts for required fields, ensuring sensitive data is masked.
- Introduced constants for valid providers, UI backends, and workspace modes to maintain consistency.
- Implemented tests to ensure alignment between constants and interactive choices in onboarding steps.

* feat(onboarding): improve WeChat account ID prompt and validation for newly enabled channels

* feat(onboarding): enhance WeChat backend credential prompts and validation

* feat(onboarding): refine WeChat backend credential prompts for wecom and wechatmp

* Refactor onboarding package for improved structure and clarity

- Simplified the onboarding package by removing unnecessary re-exports and consolidating public API to only include `run_onboard`.
- Updated `install_back_keys` to `install_navigation_keys` for clarity and consistency in the prompter module.
- Enhanced the `NonInteractivePrompter` class to support strict mode, allowing for better handling of non-interactive prompts.
- Adjusted the onboarding steps to utilize the new navigation keys installation method.
- Improved the `run_onboard` function to handle section implications based on user flags, enhancing the onboarding experience.
- Updated tests to reflect changes in imports and ensure compatibility with the new structure.

* feat(onboarding): enhance validation logic for non-interactive prompts

* refactor(onboarding): streamline onboarding module structure and enhance validation error handling

* refactor(onboarding): enhance config revert logic to preserve original file state

* refactor(onboarding): enhance tavily key validation and error handling in onboarding process
2026-05-28 12:42:49 +01:00
Xi Zhang b9ad694467 fix: resolve path correctly when workspace name appears in parent path 2026-05-23 12:37:21 +01:00
Ziheng Zhang d2283397a4 feat(feishu): scan-to-create QR onboarding + silence unsubscribed WS events (#239)
* feat(feishu): scan-to-create QR onboarding flow

Add a device-code flow against accounts.feishu.cn/oauth/v1/app/registration
that lets users scan a terminal QR code with Feishu / Lark mobile to
auto-create a PersonalAgent bot app with the required IM permissions
pre-attached. The poll endpoint returns app_id + app_secret, which the
onboarding wizard then writes into the channel config — no manual app
creation on open.feishu.cn required.

- channels/feishu/onboard.py: qr_register() public entry, init/begin/poll
  helpers, QR rendering via the soft qrcode dep, automatic feishu↔lark
  domain switch based on the scanning user's tenant_brand, and a
  best-effort bot probe to surface the bot name in the wizard
- channels/feishu/__init__.py: re-export qr_register (mirrors qq)
- config/onboard.py: offer "Scan QR code (recommended) / Enter manually"
  in the Feishu branch, ask for region (feishu vs lark), then call
  qr_register and populate feishu_app_id / feishu_app_secret /
  feishu_domain; add qrcode>=7.4 to the feishu pip extras

* fix(feishu): silently absorb unsubscribed WebSocket events

Feishu auto-subscribes PersonalAgent apps to many event types
(im.message.reaction.created_v1, message.read_v1, message.recalled_v1,
chat.member.*, ...) that EvoScientist doesn't register handlers for.
Without intervention, lark-oapi's dispatcher raises EventException
("processor not found, type: ..."), the WS client logs it at ERROR and
replies HTTP 500 on the frame, and Feishu marks the event as failed
and retries it.

The problem is amplified by _send_ack_reaction: every inbound message
triggers our own reaction, which Feishu echoes back as
reaction.created_v1, creating a continuous ERROR-log feedback loop and
pointless retries.

Wrap EventDispatcherHandler._do_without_validation after build() to
swallow "processor not found" EventExceptions (debug log + return None)
while letting all other errors propagate. Failure-safe: if lark-oapi's
internal API changes the wrapper degrades to the prior behavior rather
than breaking the channel.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-05-20 22:53:16 +08:00
dependabot[bot] b11932f91b chore(deps): bump idna in the uv group across 1 directory (#238)
Bumps the uv group with 1 update in the / directory: [idna](https://github.com/kjd/idna).


Updates `idna` from 3.13 to 3.15
- [Release notes](https://github.com/kjd/idna/releases)
- [Changelog](https://github.com/kjd/idna/blob/master/HISTORY.md)
- [Commits](https://github.com/kjd/idna/compare/v3.13...v3.15)

---
updated-dependencies:
- dependency-name: idna
  dependency-version: '3.15'
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-05-20 15:39:12 +01:00
Xi Zhang 7959495a13 feat(deploy): add EvoSci deploy subcommand (#228)
* feat(deploy): implement standalone LangGraph server and CLI command for deployment

* feat(deploy): enhance port validation and environment variable management for deployment

* Refactor langgraph dev deployment and introduce workspace sidecar protocol

- Updated the deployment mode handling in `server.py` to use a single environment variable `EVOSCIENTIST_DEPLOY_MODE` with values `full` and `stripped`.
- Enhanced the `manager.py` to implement a workspace fingerprint sidecar, allowing cross-process reuse of langgraph dev instances while ensuring workspace consistency.
- Introduced functions to write and read the workspace sidecar, with error handling for missing or corrupt data.
- Added tests for the workspace sidecar functionality, including validation of the JSON schema and ensuring proper error handling for workspace mismatches.
- Updated existing tests to reflect changes in deployment mode handling and added new tests for signal handling during shutdown.
- Ensured that cleanup routines remove the workspace sidecar alongside the PID file during shutdown.

* fix(langgraph): improve workspace sidecar checks for process ownership and stale handles
2026-05-20 11:06:35 +01:00
X-iZhang 9c7347eedb chore: update version to v0.1.1 in badges and pyproject.toml 2026-05-19 14:06:22 +01:00
Xi Zhang 331056cdc8 feat(middleware): upgrade deepagents 0.5.7 → 0.6.2 (#231)
* feat(middleware): add CodeInterpreterMiddleware with project-specific configuration

chore(config): increase checkpoint retention limit for runaway conversations

fix(tests): update database schema references from 'blob' to 'value'

chore(deps): update deepagents dependency to include quickjs support

* feat(deepagents): update to version 0.6.1 and add optional dependencies for quickjs

* feat(sessions): improve error handling for message deltas and update Overwrite type check

* Enhance PruningCheckpointer with DeltaChannel Awareness

- Introduced a new pruning strategy in `_prune_after_put` to preserve the `_DeltaSnapshot` chain during checkpoint pruning.
- Implemented methods to fetch recent checkpoint IDs and walk to snapshot ancestors, ensuring that necessary checkpoints are retained.
- Updated SQL queries to handle checkpoint and write deletions more efficiently.
- Added comprehensive tests for DeltaChannel-aware pruning, ensuring that the pruning logic correctly handles various checkpoint scenarios, including those with and without snapshot seeds.
- Refactored `_load_checkpoint_messages` to utilize the new saver interface, improving message reconstruction from checkpoints.

* feat(tests): add migration sweep test to preserve snapshot ancestor

* feat(sessions): enhance checkpoint retrieval to prevent transcript leakage in multi-agent scenarios

* feat(middleware): enhance CodeInterpreterMiddleware with configurable timeout and result character limit

feat(config): add CodeInterpreterMiddleware tuning parameters to EvoScientistConfig

feat(sessions): implement inline message delta reducer for improved message handling

* feat(dependencies): update deepagents version to 0.6.2 in pyproject.toml and uv.lock
2026-05-19 12:35:36 +01:00
Xi Zhang 385f9756c1 feat(backends): implement tier-aware virtual mount resolution for ski… (#236)
* feat(backends): implement tier-aware virtual mount resolution for skills and memories

* test: add end-to-end test for workspace tier shadowing global tier in CustomSandboxBackend

* feat(backends): enhance virtual mount resolution for skills and memories with tier paths and quoting

* fix(tests): update Python command in virtual mount resolution tests to use python3
2026-05-19 11:50:49 +01:00
Ziheng Zhang 7f1aa3b0f6 fix(cli): handle spaces in @file mentions (#234)
* fix(cli): handle spaces in @file mentions

The @file parser truncated at the first space, so dragging or pasting a
filename like `@PREPING_ Building Agent.pdf` only matched `@PREPING_`
and warned "file not found". Now supports `@"..."` / `@'...'` quoted
form for explicit paths, plus a greedy expansion fallback that walks
across whitespace until an existing file resolves (bounded by newlines,
the next `@`, and a 20-token cap). Autocomplete also returns quoted
mentions for any candidate containing a space.

* style: apply ruff format to file_mentions
2026-05-18 12:05:21 +01:00
X-iZhang 0d51406149 chore: update wechat_group image asset 2026-05-16 16:07:32 +01:00
Wiktor Cupiał a4c9c779c9 feat: status and elapsed time indicator (#218)
* feat: status and elapsed time indicator

* test: add tests for tui-status

* fix: move to enum+switch, change phase calculation

* feat: remove 'done' phase
2026-05-13 14:15:39 +01:00
dinos 4b0c91190a feat(llm): add dashscope-code provider for Alibaba Coding Plan keys (#225)
* feat(llm): add dashscope-code provider for Alibaba Coding Plan keys

Alibaba Cloud Bailian "Coding Plan" subscription keys (sk-sp-*) route
through a separate endpoint (coding.dashscope.aliyuncs.com/v1) that the
standard `dashscope` provider can't reach. Add a sibling provider entry
matching the zhipu/zhipu-code and moonshot/kimi-coding precedents, with
its own validator (the coding endpoint returns 404 on /models, so probe
via chat.completions instead).

Closes #224

* fix(llm): keep dashscope as default provider for qwen3-coder shortcut

The MODELS dict is built from _MODEL_ENTRIES via a last-write-wins dict
comprehension. The initial commit listed dashscope-code AFTER dashscope,
which silently flipped the bare `get_chat_model("qwen3-coder")` shortcut
to the coding endpoint — breaking standard sk-* keys.

Reorder to match the zhipu-code / zhipu precedent: coding endpoint first,
general endpoint last so the general endpoint wins the collision and
remains the default for the shared "qwen3-coder" short name.
2026-05-13 10:23:32 +01:00
dependabot[bot] 35ea2bfb52 chore(deps): bump urllib3 in the uv group across 1 directory (#222)
Bumps the uv group with 1 update in the / directory: [urllib3](https://github.com/urllib3/urllib3).


Updates `urllib3` from 2.6.3 to 2.7.0
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.6.3...2.7.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.7.0
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-12 08:53:08 +01:00
Ziheng Zhang 8fe774b056 Feat/qq interactive buttons (#220)
* feat(qq): add inline keyboard buttons for C2C HITL approval

QQ Bot supports inline buttons via `markdown + keyboard` payloads. Clicks
arrive as `interaction_create` events through the existing botpy
WebSocket gateway — no extra subscription needed beyond enabling the
`interaction` intent. Group-scope clicks are out of scope here (DM only).

Send path
- `_build_qq_keyboard(buttons)` mirrors the Feishu helper, mapping the
  generic `{text, value, type}` shape to QQ's `{render_data, action}`
  with action.type=1 (callback). One button per row for mobile clarity.
- `_send_chunk` extracts `metadata["buttons"]` and threads a `keyboard`
  payload into `_post_markdown_message` for C2C only.
- Markdown→plain fallback can't carry a keyboard, so when buttons were
  attached the fallback content gets a textual `Reply: 1=Approve, …`
  hint built from the button list. `_parse_approval_reply` accepts
  the same values typed manually, so the user is never stuck.

Receive path
- `on_interaction_create` is registered on the bot class.
- `_on_interaction` extracts `data.resolved.button_data`, builds an
  InboundMessage, runs it through inbound middleware (Dedup suppresses
  retry callbacks), and publishes directly to the bus — bypassing the
  per-sender debounce buffer so the click value isn't merged with any
  text typed in the same window.
- Always ACKs via `api.on_interaction_result(id, 0)` in `finally` so
  QQ doesn't show the button as "expired", even if middleware drops
  the click or something throws downstream.

`QQ.inline_buttons=True`; `_approval_prompt_metadata` now auto-attaches
the Approve/Reject/Approve-all button row for QQ HITL prompts.

* fix(qq): button-value coercion, ACK timing, HITL consumer wiring

Fixes 6 bugs found in the inline-keyboard commit and consolidates the
button helpers so the keyboard builder, plain-text fallback hint, and
interaction handler share one coercion path.

- Plain-text fallback no longer crashes on non-string `value` (e.g.
  `{"text": "OK", "value": 42}`).  Extracted `_normalize_button` is now
  the single place that resolves `(label, value)` and coerces non-strings.
- `metadata["button_value"]` is the coerced string instead of the raw
  payload, matching `content` and downstream string comparisons.
- `_on_interaction` ACKs first, before publishing to the bus, so the
  QQ button UI never shows "expired" if middleware is slow.
- Wire `_approval_prompt_metadata` + `_format_approval_prompt(with_buttons=)`
  into `InboundConsumer._stream_with_hitl` and `cli.channel.channel_hitl_prompt`
  so the QQ `inline_buttons=True` capability is actually used end-to-end
  (HITL prompts auto-attach Approve/Reject/Approve-all buttons when the
  channel advertises the capability).
- Trim contradictory `_QQ_DEFAULT_PERMISSION` comment.
- Fix `test_group_interaction_ignored` docstring (ACK runs first now,
  not in `finally` after a `return`).

Tests: `_normalize_button` covered indirectly via existing keyboard tests;
new regressions for non-string fallback hint, ACK-on-handler-throw, and
string-coerced `button_value` metadata.

* refactor(qq): slim button helpers and explicit has_buttons flag

Inline single-use _button_hint and the _QQ_BUTTON_STYLE/_QQ_DEFAULT_PERMISSION
constants in qq/channel.py; tighten _on_interaction (drop unreachable
"[button click]" sentinel and unused triggering_message_id metadata; collapse
"if resolved else" ternaries via `or ""`).

Replace the metadata round-trip ("buttons" in metadata) used to detect button
support in consumer.py and cli/channel.py with an explicit has_buttons bool
threaded through both the prompt formatter and metadata builder.

Apply ruff format to the previously unformatted blocks introduced earlier on
this branch so CI lint passes.

* feat(qq): send post-decision confirmation after HITL approval

Send a visible confirmation message ("✅ 已批准" / "❌ 已拒绝") right after
the user resolves a HITL approval — QQ Bot has no message-recall or edit API
for C2C, so a follow-up message is the only way to give the click/reply
strong feedback.

Bus consumer (consumer.py): only sends the confirmation when the user
actually responded (event was set), to avoid pretending the user approved
when the request really timed out and auto-approved.

CLI HITL prompt (cli/channel.py): mirrors the same set of confirmation
strings.  Timeout and unrecognized-reply paths keep their existing English
text since their semantics differ (auto-reject vs auto-approve, plus a
hint about the unparsed input).
2026-05-11 10:44:20 +01:00
dependabot[bot] ccb3083183 chore(deps): bump langchain-core in the uv group across 1 directory (#219)
Bumps the uv group with 1 update in the / directory: [langchain-core](https://github.com/langchain-ai/langchain).


Updates `langchain-core` from 1.3.2 to 1.3.3
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain-core==1.3.2...langchain-core==1.3.3)

---
updated-dependencies:
- dependency-name: langchain-core
  dependency-version: 1.3.3
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-09 16:50:25 +01:00
X-iZhang f00ee1b5cc chore: update version to v0.1.0 2026-05-08 22:57:30 +01:00
Xi Zhang c407d2e20f Fix/async subagent model switch (#217)
* feat(middleware): add ConfigurableModelMiddleware for dynamic model resolution

- Introduced ConfigurableModelMiddleware to resolve chat models from RunnableConfig.configurable on each call.
- Updated middleware initialization to include ConfigurableModelMiddleware.
- Enhanced context editing middleware tests to verify presence of ConfigurableModelMiddleware.
- Implemented tests for ConfigurableModelMiddleware to ensure correct model overriding and caching behavior.
- Added tests for deepagents model-passthrough patch to verify configuration injection in async tasks.

* feat(async-subagent): update middleware handling to prevent deadlocks in async sub-agents

* style: Refactor code formatting for improved readability in patches and test files

* refactor: streamline middleware construction and improve async handling in ConfigurableModelMiddleware

* fix: remove unused request parameter from _read_model_override function

* refactor: improve async handling in _ClientProxy and enhance logging in ConfigurableModelMiddleware
test: add behavior test to ensure AskUserMiddleware is excluded in async subagent mode
2026-05-08 22:52:24 +01:00
Ziheng Zhang 89b0ecdbf3 feat(qq): add QR-code scan-to-configure onboarding for QQ Bot (#213)
* feat(qq): add QR-code scan-to-configure onboarding for QQ Bot

Adds a `qr_register()` flow that drives q.qq.com's create_bind_task /
poll_bind_result APIs so the wizard can auto-fill `qq_app_id` and
`qq_app_secret` after the developer scans a QR code with a bound QQ
account, falling back to manual entry on failure or cancel.

- channels/qq/crypto.py: AES-256-GCM helpers for decrypting the bot's
  client_secret returned by poll_bind_result.
- channels/qq/onboard.py: portal API client + polling loop.
- channels/qq/__init__.py: re-export `qr_register`.
- config/onboard.py: QQ branch in `_step_channels` that offers
  "Scan QR code" vs "Enter manually", and skips the manual prompt
  loop when a scan succeeded.

* style(qq): fix ruff lint errors in onboard.py

Move `import os` to the top-level import block (E402), drop the legacy
`typing.Optional`/`typing.Tuple` imports (UP035), and use the PEP 585/604
builtin generics (`tuple[...]`, `X | None`) for the few annotations that
still referenced them (UP006/UP045). No behavior change.

* fix(qq): harden QR onboard error paths and declare scan deps

Address review feedback on PR #213:
- Declare cryptography>=41.0 and qrcode>=7.4 in [qq]/[all-channels]
  extras and in _CHANNEL_PIP_DEPS so the scan flow no longer fails
  with an opaque ImportError on a fresh `evoscientist[qq]` install.
- Polling loop logs each _poll_bind_result failure and aborts after
  5 consecutive errors instead of silently spinning until the 600s
  timeout, restoring the documented Raises: RuntimeError contract.
- Wrap decrypt_secret in try/except so failures honor the
  None-on-failure contract instead of letting exceptions escape.
- Preflight `import cryptography` in the scan branch and offer
  install or fall back to manual entry.

* style: ruff format collapse two over-wrapped log/console lines
2026-05-08 12:08:52 +01:00