ddf337fc407bd21140d9511d4ffc713205c279e0
102 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ddf337fc40 |
审批暂停与继续:锚点恢复准入 + 不批准终止本轮(含 HITL 装配收敛与工作区范围)
Docker / build (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
Build / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
langgraph_dev/http.py - _compatible_checkpoint/_history_admission 支持带锚点准入:以承载 __interrupt__ 的检查点 为父状态读取并校验中断标识;payload 保留锚点作为父状态 - 不带锚点时行为与原来完全一致(兼容性只靠这一条) middleware/dynamic_review.py - _abort_requested/_finish_turn + hook_config(can_jump_to=["end"]):在库消费中断后结束本轮 (工具不执行、不再发起模型调用);无标记时完全休眠 随本批一并落定(早前改动):EvoScientist.py / workspace_scope.py 的 HITL 装配收敛与工作区范围, 以及对应测试调整。 |
||
|
|
2e5db60dc5 |
fix(merge): keep the Ai4Sci runtime working on upstream v0.3.0's dependency stack
The runtime (llm/runtime.py, stream/stop.py) is a separate development line whose stop adapter drives LangGraph internals. Upstream v0.3.0 bumps its dependencies for currency, and two exact-version assertions in that adapter turned the bump into a silent regression: stop ownership was refused, so cancellation/continuation runs never reached a terminal state. Resolved without touching the adapter's logic: - stream/stop.py: claim ownership by *capability* instead of an exact version string. The internals the adapter swaps (_graph_aiter / _pump_cond / _exhausted / _aborting / _anext_task / _mux) and the SQLite saver's connection lock are present and identical in langgraph 1.2.6 and 1.2.11, and langgraph-checkpoint-sqlite 3.1.1 exposes the same barrier as 3.0.3. A new patch release can no longer disable stop ownership by being newer; a release that really drops the internals still fails closed with CHECKPOINT_STOP_ADAPTER_UNSUPPORTED. - EvoScientist.py: supply TodoListMiddleware only when deepagents' own default chain lacks it. deepagents 0.7 dropped it (upstream adds one back); 0.6.x still ships it, and a second instance collides by name in langchain's create_agent. - backends.py: fall back to a shape-compatible DeleteResult when deepagents has no delete support, so upstream v0.3.0's delete refusals import and run on either line. - tests/test_backends.py: gate the delete-behaviour tests on the framework actually providing backend deletion instead of asserting a specific stack. Verified: 4216 passed / 33 skipped / 29 failed / 16 errors — every remaining failure is pre-existing on the untouched pre-merge tree except two (a google-stream cleanup-order assertion and one webui launcher test). |
||
|
|
4c338ed914 |
fix(merge): resolve integration gaps found by running the v0.3.0 test suite
Post-merge validation fixes (upstream v0.3.0 + Ai4Sci fork): - llm/patches.py: restore the two module-level patch calls the merge dropped (_patch_openai_empty_sse_keepalive, _patch_deepagents_extracted_document_text) and make _is_ccproxy_codex accept an explicit base_url/api_key so the invocation plan can classify an endpoint without mutating the process env. - llm/models.py: an explicit per-call plan now wins over EVOSCIENTIST_USE_RESPONSES_API (env is only a default), an explicit caller `reasoning` block survives an explicit use_responses_api=False, and the third-party (openrouter) default effort stays the fork's fixed `medium`. - EvoScientist.py: sub-agent stacks pass NO_OP_SINK as `events` instead of None. - middleware/error_normalization.py: platform-generated diagnostics (ModelOutputTruncatedError) keep their actionable text while provider SDK errors still get the canned redacted message. - pyproject.toml: hold google-genai 1.x (langchain-google-genai>=4.3.7,<4.4) because llm/gemini_interactions.py drives the 1.x Interactions API; this is also what deepagents 0.7.13 requires. - config/settings.py: restore upstream's use_responses_api config field. `reasoning_effort` stays deleted on purpose — Ai4Sci keeps reasoning an invocation-plan parameter, never a deployment-env override. - tests: align upstream tests that encode replaced behaviour (ccproxy responses-api context, reasoning-effort-overrides-env, fingerprint coverage) with the fork's contracts. |
||
|
|
470cf75722 |
merge: bring upstream v0.3.0 (72 commits) into Ai4Sci fork
Merged upstream/main (
|
||
|
|
ea99ce9f7e | feat(runtime): host execution registry, web checkpointer, stop control and sandbox cancellation | ||
|
|
be8c23861d |
feat: dispatch newly installed experts in the background without /new (#420)
* refactor: rename AGENTS.md to EXPERT.md per EvoSkills convention * feat: resolve newly installed experts on start_async_task miss * chore: reword the /expert invite hint to state the dispatch boundary * docs: note the resolve-on-miss caller in build_expert_async_subagent_specs * fix: thread the construction cfg through resolve-on-miss and symmetric setdefault * fix: guard agent_map iteration against concurrent resolve-on-miss writes * fix: warn once per broken expert on repeated resolve-on-miss walks * fix: scope the /expert invite hint to newly installed experts * docs: note live uninstalls as a resolve-on-miss limitation * chore: isolate the warn-once collision test key from the route-specs suite |
||
|
|
7dbb68d807 | feat(memory): first-contact profile bootstrap with USER_PROFILE frontmatter (#450) | ||
|
|
c683f6e739 |
feat: prepare EvoScientist 0.3.0
Docker / build (push) Has been cancelled
Build / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
Add bounded document ingestion, controlled web search, recoverable session support, subagent timeouts, and the native sandbox runtime contract. Unify package versioning and add release-focused regression coverage. |
||
|
|
7605136189 |
fix: pass cfg explicitly to web_full middleware
Docker / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
Build / build (push) Has been cancelled
|
||
|
|
b68fdf6d57 |
feat: disable agent observation writes in web deploy
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
8376f56ab4 |
feat: native sandbox execution, dynamic review middleware, and workspace files
Adds native sandbox execution runtime, dynamic review middleware, and workspace file handling, with supporting stream events, prompt, and scope registry changes plus architecture docs. |
||
|
|
5a581c78a2 |
feat: add scoped model runtime configuration
Build / build (push) Has been cancelled
Docker / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
Introduce provider, model, and invocation contracts with encrypted configuration persistence. Add web runtime fencing, route fallback, recovery middleware, workspace scoping, and comprehensive tests. |
||
|
|
1f3e8f57a3 |
fix: resolve subagent tools at execution decision (#381)
* fix: defer subagent tool resolution Signed-off-by: Aditya Datta <crazyme07071996@gmail.com> * fix: preserve inherited subagent tools * test: cover injected subagent tools * style: format subagent regression --------- Signed-off-by: Aditya Datta <crazyme07071996@gmail.com> |
||
|
|
ab1a6b0062 |
feat: adopt the AGENTS.md expert-skill contract (#404)
* feat: detect expert skills by AGENTS.md presence * feat: let the orchestrator choose the expert dispatch tool * refactor: drop per-skill expert dispatch classification * chore: remove duplicate TestSkillManager test classes left by rebase * fix: clear legacy actor fields when AGENTS.md declares the expert * refactor: move the expert prompt into ActiveTeamMiddleware |
||
|
|
bd2464423a |
feat: agent-teams part D - async expert dispatch mechanism (#391)
* fix(deps): pin openrouter below 0.11 to avoid SSE stream regressions (#373) * fix(openrouter): address SSE stream leak by closing response iterator * refactor(openrouter): pass through SDK args in SSE leak patch * test(openrouter): make SSE leak tests version-agnostic across SDK generations * fix(openrouter): remove SSE stream leak patch and update dependencies * fix: repair interrupted tool call history (#366) * fix: repair interrupted tool call history Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths. * fix: repair malformed tool calls and dedupe repair warnings Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted threads with syntactically invalid tool calls get synthesized error results and are accepted by strict providers. Preserve the originating tool call's name in the synthesized ToolMessage, and deduplicate repair warnings per unique tool-call id via a warned set owned by the middleware instance, since the middleware rewrites the request but not thread state. Document the middleware's scope versus deepagents' PatchToolCallsMiddleware (orphan ToolMessage dropping and mid-run coverage). --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * Update README.md * Update README.md * Update README.zh-CN.md * fix: scrub host path from skill_manager output and guard batch install (#377) * fix: scrub host path from skill_manager output and guard batch install * test: tighten install leak guards to catch host path in either tier --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * fix: log missing async-subagent tools at DEBUG, not WARNING (#378) * fix: log missing async-subagent tools at DEBUG, not WARNING * fix: distinguish load_subagents callers via async_swap_pending flag --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * fix: default reasoning context for codex proxy Responses API (#380) * fix: set langgraph and codex proxy runtime defaults * fix: address runtime default review feedback * fix: drop langgraph dev env defaults per maintainer review langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment, so the reported KeyError cannot come from this flow; the env defaults added here were unnecessary. Scope the PR back to the codex proxy reasoning context fix only. --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * release: v0.2.4 (#389) * fix: add support for new Anthropic models and enhance adaptive thinking tests * fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models * fix: update version to v0.2.4 in badges, README, and project files * fix: update Star History chart links in README and README.zh-CN * fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog * fix: update wechat group image in assets * refactor(runtime): centralize async bridges under an owned runtime (#376) * feat(runtime): add application-scoped async runtime * refactor(cli): use owned runtime for session stats * refactor(onboard): use the owned async runtime * docs(runtime): record async bridge ownership * refactor(middleware): keep sync fallback synchronous * refactor(mcp): load tools on an owned runtime * refactor(cli): share owned runtime across entry points * refactor(channels): make inbound sync bridge explicit * refactor(stream): run Rich streaming on owned runtime * chore(runtime): remove nest-asyncio dependency * refactor(asyncio): require active loops in async code * docs(runtime): document final event loop ownership * fix(stream): cancel stalled owned streams * fix(cli): recover cleanly from stream cancellation * fix(runtime): drain executor work before shutdown * fix(runtime): terminate cancelled shell process trees * fix(models): let fallback bypass selector failures * fix(cli): reset interrupt handling between turns * docs: rm implementation spec * fix(serve): cancel active turns during shutdown * fix(runtime): protect settlement from waiter cancellation * fix(backends): reject empty shell commands * fix(runtime): terminate descendants after shell exit * fix(mcp): keep standalone discovery off channel loop * fix(cli): own and settle interactive prompt cancellation * fix(serve): keep channel sends off runtime loop * fix(stream): scope cancel context to iterator steps * refactor(serve): require the owned async runtime * fix(channels): keep interactive sends off runtime loop * fix(selector): surface fallback without log spam * test(runtime): normalize Windows shell marker * fix(cli): serialize interactive session turns * fix(shell): bound output drain after termination * fix(ui): do not retry owned runtime failures * fix(shell): allow signal-safe registry reentry * fix(shell): avoid terminating reused process ids * fix(channels): preserve streaming send order * fix(cli): report runtime shutdown timeouts cleanly * fix(mcp): guide async callers to async loader * docs(runtime): clarify reserved async bridge APIs * fix(runtime): bound code interpreter cleanup * test(shell): use active Python for drain regression --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * feat: add payload-aware EvoAsyncSubAgentMiddleware * feat: register expert_container_async graph for async expert dispatch * feat: fold installed expert skills into async subagent registry * fix: accept 'async' as valid default_dispatch value * feat: dispatch-aware ActiveTeamMiddleware cue (task vs start_async_task) * fix: drop future annotations in expert_async_subagent so ToolRuntime injects * feat: surface output_path and skill_name to expert container as runtime cue * fix: extend AsyncWatcher client cache with expert specs for completion nudge * feat: teach main agent the async-expert return envelope shape * feat: propagate cfg.model to expert-async runs.create via ClientCacheProxy * chore: guard AsyncWatcher client-cache extension against upstream rename * test: cover output_path runtime-context tail block and wrong-type guard * docs: drop out-of-repo notes/ ref from expert_container_async module doc * fix: warn on unrecognized default_dispatch frontmatter value * fix: reject empty-body expert skills on async dispatch to match sync policy * docs: explain why expert container includes general-purpose subagent * fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385) * fix: propagate langgraph dev bind port into subprocess env for self-loop URL * fix: keep parent env authoritative over workspace .env for mapped keys * fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins * fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS * fix: merge .env via dotenv_values to close empty-value and RMW-race edges * chore: align docstrings after .env-merge rework --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * Add Requesty as an LLM provider (#346) * Add Requesty as an LLM provider * Address review: Requesty prompt caching, model ordering, key validation - Declare Anthropic-style prompt caching for Requesty Claude models by default (mirroring the OpenRouter behavior), with an opt-out flag EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed provider, so the caching check now uses the original provider name. - Move the Requesty model entries above OpenRouter so Requesty no longer overrides native/OpenRouter models for names it shares with them (the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry. - Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an invalid/missing key (public catalog), so it cannot validate a key. Use a minimal authenticated /v1/chat/completions request instead (200 = valid, 403 = invalid), verified against the live endpoint. - Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip). * Validate Requesty key against auth layer, not a specific model The onboarding validator probed /v1/chat/completions with a hardcoded real model (openai/gpt-4o-mini), which tied key validation to that model staying available upstream. The router resolves auth before the model, so probe a deliberately nonexistent sentinel model (requesty/auth-preflight) instead: a valid key yields 404 (model-not-found, auth passed), an invalid key yields 401/403, and 429/5xx stay inconclusive so a transient outage does not reject a good key. Add unit tests covering each case. --------- Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com> * fix(llm): filter unnamed tool calls (#390) * fix(llm): filter unnamed tool calls * test(llm): cover tool call sanitization branches * fix(llm): repair unnamed tool calls in middleware --------- Co-authored-by: nightcityblade <nightcityblade@gmail.com> Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> * fix(middleware): mount tool-history repair on sync subagents and harden raw tool-call vetting (#393) * feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation * fix(tests): add test for dropping non-list raw tool calls in repair_tool_history * Add Atlas Cloud LLM provider (#388) * Add Atlas Cloud LLM provider * Add Atlas Cloud onboarding support * fix(validators): update atlascloud key validation to handle insufficient balance case --------- Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com> Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com> * fix: prepend EvoAsyncSubAgentMiddleware for prefix cache stability * docs: clarify list_dispatchable_experts covers both dispatch shapes * fix: guard async expert fold-in against reserved-name collisions * fix: honest advertising surfaces for async expert dispatch * fix: compose expert persona into base-stack system_message * fix: drop payload from start_async_task, inject skill_name by construction * feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395 - Introduced TodoListMiddleware to the middleware stack for better task management. - Updated HITL interrupt configuration to include 'delete' operations requiring approval. - Implemented error handling for delete operations in read-only and memory backends. - Enhanced approval prompt formatting to display file paths for delete actions. - Added tests to ensure delete operations are correctly blocked or prompted for approval. - Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality. * revert: drop skill_manager from sync expert-container tool_registry * revert: drop skill_manager from async expert-container tools * fix: drop removed ASYNC_TASK_SYSTEM_PROMPT import for deepagents 0.7.0 * fix: mock list_dispatchable_experts in single-cue test for CI * chore: drop stale output_path from async container graph docstring * fix: forward configurable_extra through owned-runtime and HITL re-invocations --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> Co-authored-by: Sanjay Santhanam <51058514+Sanjays2402@users.noreply.github.com> Co-authored-by: Yougang Lyu <82445958+youganglyu@users.noreply.github.com> Co-authored-by: houren Antony <2212222@mail.nankai.edu.cn> Co-authored-by: dinos <dinospk1999@gmail.com> Co-authored-by: Thibault Jaigu <84420566+Thibaultjaigu@users.noreply.github.com> Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com> Co-authored-by: nightcityblade <jackchen@haloailabs.com> Co-authored-by: nightcityblade <nightcityblade@gmail.com> Co-authored-by: nb213 <binyangzhu000@gmail.com> Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com> |
||
|
|
3cda9894c7 |
feat: agent-teams part C - expert selection UX (depends on part B) (#371)
* feat: bias main-agent delegation toward configurable.active_teams * feat: add /experts and /expert TUI commands for expert-skill summoning * feat: align expert-selection wording with WebUI (invite/dismiss) * chore: clear active_teams on /new, cleanup active_team.py comment * fix: use local append_to_system_message in ActiveTeamMiddleware * fix: prevent configurable_extra from overriding thread_id * fix: cache expert-skill lookup for /expert completions * fix: suppress /expert completions past first arg and on exact match * fix: invalidate /expert completion cache on skill install/uninstall * fix: refuse /expert invites for non-dispatchable expert skills * fix: fire /expert cache invalidation on every install_skill / uninstall_skill path * fix: propagate active_teams to Rich CLI and serve dispatch surfaces * fix: keep invited experts across channel shutdown * fix: match /expert completions case-insensitively |
||
|
|
3a1dbf0a0f |
feat: agent-teams part B - expert-skill backend mechanism (#370)
* feat: add expert-skill schema and type filter to skill_manager * feat: fold installed expert skills into main-agent subagent registry * feat: add GET /api/teams listing expert skills for gallery * chore: cache SKILL.md body on SkillInfo, cleanup expert-container comments * fix: register skill_manager in expert-subagent tool_registry * fix: catch UnicodeDecodeError in expert-skill body loader * fix: guard expert subagent registration against name collisions * fix: skip expert registration when SKILL.md body is empty * fix: drop redundant str() guards on expert-skill frontmatter * fix: harden SKILL.md parsing on expert-registration hot path |
||
|
|
3c5cc831c0 |
Feat/configurable bind host (#402)
* feat: configurable bind host for WebUI and langgraph dev (refs #400) WebUI mode was only reachable from the machine running it: the front-end got no bind interface, and `start_langgraph_dev(...)` was called without a host, so both servers stayed on loopback with no way to widen them. Adds two config fields with deliberately different defaults: webui_host = 0.0.0.0 front-end serves the app shell, no secrets langgraph_dev_host = 127.0.0.1 unauthenticated API, agent can run shell The design hinges on separating bind address from client address. Only bind() uses the configured interface; every consumer that *connects* (health probes, occupancy checks, async sub-agent self-dispatch) goes through the new `_probe_host`, which maps a wildcard bind back to loopback and honors a pinned interface verbatim. `_can_bind_port` is the one exception and binds the literal host, since it must replicate the bind the server itself will attempt. - manager.py: `_probe_host`, `_is_loopback_host`, `_format_hostport`; host kwarg threaded through the probes and `start_langgraph_dev`, which now emits `--host` and propagates EVOSCIENTIST_LANGGRAPH_DEV_HOST to the subprocess - sdk.py: `langgraph_dev_url` tracks host as well as port; EvoScientist.py reuses it instead of an inline f-string - server.py: `--host` flag mirroring `--port`, plus a red PUBLIC BIND banner whenever the bind is not provably loopback - webui.py: forwards both hosts; the front-end is widened via HOSTNAME because @evoscientist/webui ships no --host flag — its bin launcher does `HOSTNAME: process.env.HOSTNAME || "127.0.0.1"`. The warning is gated on the backend host only, so the shipped front-end default doesn't print a banner on every launch Verified end to end against a live server: requesting 0.0.0.0 yields a socket listening on 0.0.0.0 with the health probe correctly resolved to 127.0.0.1, while the default still binds 127.0.0.1 only. Note: webui_host defaulting to 0.0.0.0 is a behavior change — upgrading users will find the front-end reachable from the LAN. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: default both bind hosts to 0.0.0.0, add --host and wizard host rendering (closes #400) Completes the remaining items from #400. - `langgraph_dev_host` now defaults to 0.0.0.0, matching `webui_host`. Remote WebUI use needs both anyway (the UI reaches the backend from the browser, not server-side), so a loopback backend default just meant every remote user hit a silently failing UI. `_DEFAULT_HOST` and sdk's `DEFAULT_LANGGRAPH_DEV_HOST` follow, so there is one story about where these servers listen. SECURITY: this exposes an unauthenticated API whose agent can run shell commands. The red PUBLIC BIND banner consequently fires on every launch while exposed — kept deliberately, since the exposure is real and the escape hatch (`--host 127.0.0.1` / `config set langgraph_dev_host`) is only discoverable if we say so. READMEs now lead with the warning and document the SSH-tunnel alternative. - `EvoSci --host <ip>` on the WebUI launch path, driving both servers. In WebUI mode they are two halves of one surface; moving only one leaves the UI loading but unable to reach the agent. Blank values are dropped rather than written as an empty override that would beat the config file. - Onboarding wizard no longer prints hard-coded `http://127.0.0.1:{port}` / `http://localhost:{port}` (steps.py:160, :223) — both render the configured bind through `_base_url` / `_format_hostport`, so a pinned interface is reported honestly and a wildcard still shows loopback. Verified against a live server: with no host argument at all, resolution through EvoScientistConfig yields a socket listening on 0.0.0.0, a client URL of http://127.0.0.1, and the warning gate returning True. Still open and tracked separately: the front-end takes its backend URL from browser input: `@evoscientist/webui` reads only HOSTNAME, PORT and EVOSCIENTIST_LANGGRAPH_DEV_PORT, so advertising a backend URL needs a change in that repo. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci: bump setup-uv v6 -> v9.0.0 to drop the deprecated node20 runtime GitHub now warns that setup-uv@v6 targets Node.js 20 and is being forced onto Node.js 24. v7.0.0 is the release that made that switch, so anything >= v7 clears the warning; v9.0.0 is current. Pinned to the full tag deliberately: setup-uv stopped publishing major and minor tags in v8.0.0 as supply-chain hardening, so `@v9` and `@v8` return 404 and would fail the job outright. Releases are immutable from v8 on, so the full tag is as tamper-proof as a SHA. Comment left in lint.yml because "simplifying" this back to `@v9` is an easy and CI-breaking mistake. actions/checkout@v5 is already node24 and needs no change. Note: v9.0.0 flips the `prune-cache` default to false (upstream did this to ease load on PyPI infrastructure). None of these workflows set it, so they follow the new default and Actions cache usage may grow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(cli): correct --host help text and warn on public bind in non-WebUI modes The --host help claimed "WebUI mode only", which is wrong in a way that matters for security. `--host` writes `langgraph_dev_host` unconditionally, and `_ensure_async_subagent_server` auto-starts that backend for tui / cli / serve as well — the langgraph dev server is shared across UI modes. So the flag narrows or widens the agent API in every mode, and only `webui_host` is actually WebUI-specific. Reported against cli/commands.py. The documentation error hid a real gap: the PUBLIC BIND banner lived only in deploy/server.py and deploy/webui.py, so a plain `EvoSci` session bound 0.0.0.0 with no runtime signal whatsoever — and `--help` is opt-in, so fixing the text alone would not surface it. Added the same banner to the shared CLI path, gated on `is_async_subagents_available()`: ensure_langgraph_dev fails soft (async degrades to in-process delegation), and warning about a bind that never happened would be worse than staying quiet. READMEs (EN + zh-CN) get the same correction — the warning block sat inside the Desktop WebUI section and read as WebUI-scoped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(deploy): strip the config-derived bind host, not just the CLI one `deploy()` only stripped the `--host` branch. When the flag was omitted, `getattr(config, "langgraph_dev_host", ...)` flowed unstripped into `_is_port_occupied`, `is_langgraph_dev_running`, `start_langgraph_dev` and the banner. `run_webui` already strips unconditionally; this aligns the two. Reachable because `deploy()` reads through `getattr` and is routinely handed duck-typed config objects (tests, embedders) that never run `EvoScientistConfig.__post_init__`, which is what normally normalizes these fields. Worst case was not just a bad bind: `_is_loopback_host(" 127.0.0.1 ")` is False, so a padded loopback value would print a false PUBLIC BIND warning while binding a string socket.bind() rejects outright — a security banner saying the opposite of the truth. Three regression tests added, each verified to fail against the old code. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * style: apply ruff format to the bind-host changes The Lint workflow runs both `ruff check` and `ruff format --check`; I had only been running the former locally, so five files landed unformatted and failed CI. Whitespace and line-wrapping only — no semantic change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(security): keep the langgraph dev backend on loopback by default The backend is an unauthenticated API whose agent can run shell commands, and it is auto-started in every UI mode (tui/cli/webui/serve/deploy) — so a 0.0.0.0 default put it on the network for users who never asked. Restore 127.0.0.1 as the default and make 0.0.0.0 an explicit opt-in. webui_host keeps its 0.0.0.0 default: the front-end serves the app shell only and holds no credentials. run_webui already prints a remote-backend hint when the front-end is exposed and the backend is not. Help text and both READMEs are reframed around widening rather than narrowing; the escape-hatch tests are inverted to assert the public-bind opt-in survives into argv. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
d249e320bd |
feat: unified HITL approval for sync + async sub-agents (closes #387) (#396)
* feat: implement guard for dangerous commands and enhance HITL interrupt handling * feat: enhance HITL approval mechanism and introduce session auto-approve decisions * feat: add refuse_delete option to backend and enhance async delete guards * feat: simplify delete method in CustomSandboxBackend and clarify adelete behavior * feat: enhance allow-list behavior for command resolution and add related tests |
||
|
|
f81a8b086e |
feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395
- Introduced TodoListMiddleware to the middleware stack for better task management. - Updated HITL interrupt configuration to include 'delete' operations requiring approval. - Implemented error handling for delete operations in read-only and memory backends. - Enhanced approval prompt formatting to display file paths for delete actions. - Added tests to ensure delete operations are correctly blocked or prompted for approval. - Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality. |
||
|
|
562ce0eb83 |
fix(middleware): mount tool-history repair on sync subagents and harden raw tool-call vetting (#393)
* feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation * fix(tests): add test for dropping non-list raw tool calls in repair_tool_history |
||
|
|
8b1451cdda |
refactor(runtime): centralize async bridges under an owned runtime (#376)
* feat(runtime): add application-scoped async runtime * refactor(cli): use owned runtime for session stats * refactor(onboard): use the owned async runtime * docs(runtime): record async bridge ownership * refactor(middleware): keep sync fallback synchronous * refactor(mcp): load tools on an owned runtime * refactor(cli): share owned runtime across entry points * refactor(channels): make inbound sync bridge explicit * refactor(stream): run Rich streaming on owned runtime * chore(runtime): remove nest-asyncio dependency * refactor(asyncio): require active loops in async code * docs(runtime): document final event loop ownership * fix(stream): cancel stalled owned streams * fix(cli): recover cleanly from stream cancellation * fix(runtime): drain executor work before shutdown * fix(runtime): terminate cancelled shell process trees * fix(models): let fallback bypass selector failures * fix(cli): reset interrupt handling between turns * docs: rm implementation spec * fix(serve): cancel active turns during shutdown * fix(runtime): protect settlement from waiter cancellation * fix(backends): reject empty shell commands * fix(runtime): terminate descendants after shell exit * fix(mcp): keep standalone discovery off channel loop * fix(cli): own and settle interactive prompt cancellation * fix(serve): keep channel sends off runtime loop * fix(stream): scope cancel context to iterator steps * refactor(serve): require the owned async runtime * fix(channels): keep interactive sends off runtime loop * fix(selector): surface fallback without log spam * test(runtime): normalize Windows shell marker * fix(cli): serialize interactive session turns * fix(shell): bound output drain after termination * fix(ui): do not retry owned runtime failures * fix(shell): allow signal-safe registry reentry * fix(shell): avoid terminating reused process ids * fix(channels): preserve streaming send order * fix(cli): report runtime shutdown timeouts cleanly * fix(mcp): guide async callers to async loader * docs(runtime): clarify reserved async bridge APIs * fix(runtime): bound code interpreter cleanup * test(shell): use active Python for drain regression --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
1863f0730c |
fix: log missing async-subagent tools at DEBUG, not WARNING (#378)
* fix: log missing async-subagent tools at DEBUG, not WARNING * fix: distinguish load_subagents callers via async_swap_pending flag --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
f802a49535 |
fix: repair interrupted tool call history (#366)
* fix: repair interrupted tool call history Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths. * fix: repair malformed tool calls and dedupe repair warnings Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted threads with syntactically invalid tool calls get synthesized error results and are accepted by strict providers. Preserve the originating tool call's name in the synthesized ToolMessage, and deduplicate repair warnings per unique tool-call id via a warned set owned by the middleware instance, since the middleware rewrites the request but not thread state. Document the middleware's scope versus deepagents' PatchToolCallsMiddleware (orphan ToolMessage dropping and mid-run coverage). --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
3ce5614254 | fix: harden tool-call protocol and fallback handling | ||
|
|
01845f4311 |
refactor: route middleware display events through an injected event sink (#343)
* chore: add pytest-asyncio in auto mode * test: migrate channel and stream tests to native async Convert run_async() wrapper tests to plain 'async def test_*' under pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a coroutine awaited at every call site. * test: migrate command and model/middleware tests to native async Convert run_async() wrappers (import, alias, and fixture forms) to plain 'async def test_*'. Multi-call tests merge onto one loop as sequential awaits; none asserted on loop identity. * test: migrate TUI, notifier, gateway, and session tests to native async TUI/notifier/gateway files convert run_async wrappers to plain async tests. test_sessions.py's unittest.TestCase classes move to unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async methods on plain TestCase; converting blindly would have made ~70 tests silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget in test_tui_widgets.py drops its TestCase base for the same reason. * test: replace direct asyncio.run() calls with native async tests Convert tests that called asyncio.run() (directly or via a local _run helper) to plain 'async def test_*'; delete the local helpers. * test: drop undeclared anyio markers and delete run_async helper The @pytest.mark.anyio tests relied on anyio being a transitive dep of httpx; auto-mode pytest-asyncio collects them natively. run_async() and its fixture are unreferenced after the migration, so remove them — pytest-asyncio's per-test loop teardown covers the pending-task cancellation the helper existed for (verified: full suite runs with no 'Event loop is closed' errors or destroyed-task warnings). * test: add autouse fixture for watcher cleanup * refactor: remove redundant hasattr calls * refactor: add typed middleware event sink and thread through assembly Add MiddlewareEventSink protocol + NoOpSink in middleware/events.py with a documented any-thread non-blocking contract (contract test uses a deliberately-slow fake sink). Thread an optional `events` parameter through create_cli_agent -> _get_default_middleware -> tool selector / model fallback constructors; subagent stacks are always forced to NoOpSink. * refactor: inject a notifier port into async-watcher and background middleware Add public pre_cancel_watcher() and enqueue_task_notification() to cli/async_notifier.py and a small NotifierPort protocol (middleware/notifier.py) that the module satisfies structurally. AsyncWatcherMiddleware and BackgroundExecutionMiddleware now receive the port by constructor injection at the composition root, deleting the lazy 'from ..cli import async_notifier' imports and the private _watcher_by_thread / _enqueue pokes. * refactor: invert tool-selection ownership onto a frontend event sink The adaptive tool selector now reports on_tool_selection_started / on_tool_selection / on_tool_selection_ended to the injected sink instead of writing four process-global module variables. The frontend sink (stream/sink.py FrontendEventSink) owns the selected/total/active state with consume-once + dedup-vs-last-emitted semantics; stream/tool_selection.py reads that sink object (a ToolSelectionView) rather than reaching into tool_selector's globals. Deleted: the 4 module globals, the cross-module mutations in tool_selection.py, the track_stream_selection flag, the now-vestigial _ToolSelectionTrackerMiddleware, reset_tool_selection_state_for_tests, and the autouse conftest fixture. The sink is threaded from the two interactive frontends through create_runtime_gateways -> LocalGraphGateway (read side) and _load_agent -> create_cli_agent (write side); subagent / headless stacks get NoOpSink. * refactor: route model-fallback narration through the injected event sink Delete the _ui_emit_fn / set_ui_emit module global and the ..stream.console import from model_fallback.py. The fallback middleware now reports through its injected sink: the fallback transition via the structured on_model_fallback (the frontend formats the '-> Falling back to ...' line), and the surrounding narration (primary-failure header, per-attempt outcome, exhaustion, non-fallbackable rejection) via emit_fallback_notice, preserving the exact user-facing text. The TUI binds its _append_system as the sink's fallback display where it used to call set_ui_emit (cleared on exit); the Rich CLI's sink prints to the console. _try_fallbacks / _guard_and_fallback take the sink. * refactor: declare events on the GraphGateway protocol Both gateway implementations now carry an explicit events attribute (LangGraphServerGateway holds None — no frontend renders middleware events across the HTTP boundary), so the four call sites use plain attribute access instead of getattr probing an implicit contract. * refactor: bind fallback display via the closure-scoped concrete sink The App methods used gateway.events (typed as the read-side view) and hasattr-probed for the concrete FrontendEventSink API. The enclosing factory creates that sink two hundred lines up — close over it directly: no probing, fully typed, and it becomes a constructor parameter naturally when the App class is hoisted out of the factory. * fix: end tool selection before fallback handler * fix: keep fallback display errors non-fatal * fix: preserve selector suppression for default streams * fix: restore fallback notice console display * refactor: consolidate fallback narration events * refactor: clean middleware event sink plumbing * fix: type gateway session events * refactor: make all event protocols runtime-checkable MiddlewareEventSink already carried @runtime_checkable (the stream binding guard isinstance-checks it); ToolSelectionView and SessionEvents now match, so mirroring that pattern against any of the three protocols works instead of raising TypeError. * fix(cli): close QuickJS workers after one-shot failures * fix(cli): honor no-thinking in final output * fix(channels): report failed startup accurately * fix(channels): make Telegram cleanup idempotent * fix(tui): skip command sync during exit * fix(channels): preserve startup state during retries * refactor(channels): share pending startup status * refactor(cli): expose channel startup snapshot * fix(tui): move channel startup off event loop * test(channels): release retry gate on assertion failure --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
4fc74e7da7 |
EvoScientist Ai4Sci
Docker / build (push) Has been cancelled
Build / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
|
||
|
|
88ac9f5ba1 |
fix: surface real exception class+message in SSE error events (#315)
* fix: surface real exception class+message in SSE error events * fix: tighten SSE error patch scope and key redaction * fix: redact base64-style secret suffixes fully * style: remove notes/ reference from the dosctring * fix: rebuild env cache on each error call * fix: route BaseException through serde.default on SSE/webhook paths * fix: distinguish routed providers by request URL host * feat: normalize provider-SDK exceptions via ErrorNormalizationMiddleware * refactor: drop json_dumpb dataclass-bypass wrappers, superseded by middleware * fix: guard _extract_host against SDK properties that raise * refactor: derive provider tag from ModelRequest.model, not the exception * refactor: drop serde.default patch and exception-based inference; ProviderStreamError.model_dump handles the emit * refactor: move envelope helpers from patches.py to errors.py * feat: extend ErrorNormalizationMiddleware coverage to every model-call path * chore: clean up review findings from middleware pivot * fix: pass through all langgraph.errors * fix: move langgraph.errors pass-through into _normalize * fix: pass through ContextOverflowError in _normalize |
||
|
|
f2f010a350 |
feat(memory): observation linking (#307)
* refactor(gateway): create module for launching async/bg agents * refactor(memory): refactor worker launch around source context & output deltas * refactor(gateway): generalize async/bg module * refactor(memory): revamp worker launching * feat(memory): add observation linking * test(memory): remove redundant test branches * fix(memory): make 'supersedes' relation directional * fix(memory): don't create empty project observation dirs * fix(memory): schedule direct observations for linking * fix(cli): wait for observation linker before shutdown * fix(memory): block arbitrary writes to /memories * fix(linker): remove `linked_by` attribute from frontmatter * refactor(linker): rename base relationship to `comlpements` * fix(cli): bump worker wait to 2m * feat(tools): catch malformed tool calls & retry * feat(status): add linking result to statusbar * fix(linker): don't launch linker when observations are disabled * fix(memory): use posix paths * fix(watcher): call abort hook on error status * fix(watcher): delete thread on failed run creation * fix(watcher): preserve url * fix(observation): record session_id, drop unused fields * fix(memory): reject unsupported worker source types * refactor(backends): shared memory backend builder * fix(scheduler): resolve linker inputs outside lock * fix(memory): dont launch workers / record observations without thread_id * feat(memory): include related observations in tool results * fix(memory): skip malformed observation frontmatter * revert(tools): drop tool error handling changes from this PR * fix(memory): serialize observation link writes * fix(memory): queue observations written by aborted workers * fix(memory): track observation linker launch handoff * fix(memory): resolve cross-project related observations * fix(status): avoid recounting reason-only link updates * fix(memory): avoid rereading file for content * fix(linker): use neutral prose for bidirectional reasons * test(memory): coverage for aborted/failed launches * test(memory): cleanup & helpers * feat(linker): add observations index hint |
||
|
|
7ccfe68f3f |
feat: add scheduler functionality with cron-style task management (#306)
* feat: add scheduler functionality with cron-style task management - Implemented a new scheduler subagent to automate recurring tasks using cron expressions. - Enhanced the subagent factory to include the skill manager and auxiliary chat model for the scheduler. - Created a YAML configuration for the scheduler with a detailed system prompt and toolset. - Updated README files to include documentation on scheduled tasks and usage examples. - Added tests for the scheduler, including command execution, scheduling tools, and middleware integration. - Introduced new dependencies for timezone handling and ensured compatibility in the project configuration. * fix(async-notifier): ensure fallback hint is used for unknown notification kinds * feat: enhance scheduling functionality and improve system message handling |
||
|
|
b1dccf17ea |
fix(tool-selector): memory tools & state for main agent (#305)
* fix(tool-selector): always include memory tools * fix(tool-selector): only track & show state for main agent * docs: update docstring |
||
|
|
bd307f3a11 |
refactor: LangGraph gateway layer for UI-agnostic graph and thread access (#295)
* feat(gateway): graph gateway protocol * refactor(cli): wire gateway in cli/tui * refactor(gateway): centralize runtime gateway init * chore(gateway): restrict RunRequest message type * feat(gateway): add langgraph server gateway * chore(cli): tighten serve runtime state typing * refactor(cli): route async task state reads through graph gateway * refactor(gateway): support graph targets in server gateway * refactor(cli): route session commands through graph gateway * refactor(cli): fold thread store under graph gateway * refactor(gateway): route graph state access through gateway * refactor(channels): wire graph gateway * refactor(memory): preserve graph threads for cloning * feat(gateway): add thread cloning * fix(tui): pass effective workspace for thread creation * chore(memory): add workspare dir to memory worker metadata * fix(sessions): filter preloaded UUID registy entries by the current scope * test(fakes): use https * refactor(consumer): consolidate imports * fix(stream): optional summarization event * fix(gateway): resolve abbreviated thread IDs by search * fix(gateway): page server thread listings * fix(gateway): emit pending interrupt events * style: fmt * feat(gateway): persist workspace_dir & model in thread metadata * fix(gateway): page server thread prefix resolution * fix(gateway): expose server thread list metadata * refactor: add back type def * refactor: tighten types * revert: add back worker thread deletion The worker thread forking changes are out of scope for now, so to maintain parity with the existing behavior we'll leave this intact. * fix(gateway): apply compaction to server thread history * refactor(stream): restore direct summary replay suppression * fix(gateway): preserve compaction state and server stream output * fix(gateway): close local stream generator on cancellation --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
f356de36a6 | feat: memory retrieval (#281) | ||
|
|
c02be519f6 |
feat(dangerous-mode): implement real-filesystem access with safety ch… (#276)
* feat(dangerous-mode): implement real-filesystem access with safety checks - Introduced a 'dangerous mode' allowing the agent to operate on the real filesystem. - Updated command validation to bypass path confinement while enforcing a blocklist for privileged commands. - Added warnings and guidelines for users when operating in dangerous mode. - Enhanced configuration to support dangerous mode and ensure it implies auto-approval. - Updated tests to verify the behavior of commands and configurations in dangerous mode. * feat(dangerous-mode): enhance logging and environment management for dangerous mode * feat(dangerous-mode): improve handling of dangerous mode with environment flags and enhance test isolation |
||
|
|
4b6a969df2 |
refactor(agent): make create_cli_agent(config=, chat_model=) pure (#267)
* refactor(agent): make create_cli_agent(config=, chat_model=) pure Re-applies the #183 purity refactor on top of the observation-memory lifecycle that landed in #259, integrating the two cleanly. create_cli_agent gains a pure path: when both `config` and `chat_model` are passed it builds the agent entirely from locals and writes none of the cached module globals (`_config`, `_chat_model`, `_chat_model_key`, `_EvoScientist_agent`). `/model` commits the switch via `set_active_config` / `set_chat_model_instance` only after a successful build, so a failed rebuild leaves the session on the original model (replaces the old snapshot/restore rollback). Supporting changes: - Extract `set_active_config` (write-half of `_ensure_config`), `_apply_env_from_config`, `_build_chat_model`, and `set_chat_model_instance`. - Thread `cfg` / `chat_model` through `_get_default_middleware`, `_build_base_kwargs`, `load_mcp_and_build_kwargs`, `_maybe_swap_async_subagents`, and `_inject_subagent_middleware` so the pure path never falls back to the global-writing `_ensure_config()` / `_ensure_chat_model()`. - Integrate with #259's memory middleware: subagent context-editing middleware binds the threaded `chat_model`, and the configured system prompt / memory controls read the threaded `cfg` (new threading vs the original #183, required because #259 made these paths read config). - Consolidate `cfg` resolution to one `cfg if cfg is not None else _ensure_config()` at the top of each kwargs builder, matching the pattern already used in the other config-aware helpers. * fix(agent): keep pure tool selector off global cache * fix(model): apply config switch in place to preserve reference integrity --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com> |
||
|
|
8bb1d6c0e3 |
refactor(stream): langgraph streaming v3 (#268)
* refactor(stream): langgraph streaming v3 * fix: address CR comments * chore(stream): add success field to state --------- Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com> |
||
|
|
63969b596d |
Release/v0.1.4 (#266)
* feat(middleware): reposition code interpreter middleware in the stack * feat(models): add qwen3.7-plus model entry and update context window comment * feat(models): add qwen3.7-max and qwen3.7-plus model entries for DashScope * feat(auxiliary): implement auxiliary model support for background tasks and tool selection - Added auxiliary model configuration to EvoScientistConfig. - Introduced _ensure_auxiliary_chat_model function to manage auxiliary model instances. - Updated onboarding steps to include auxiliary model selection. - Modified middleware to route tool selection to the auxiliary model when applicable. - Enhanced tests to cover auxiliary model functionality and configuration. * feat(steps): update UI backend selection options and descriptions * Refactor code structure for improved readability and maintainability * feat(patches): implement OpenRouter response reasoning item stripping to prevent multi-turn errors * feat: update version to v0.1.4 in badges, README, and pyproject.toml; adjust skill counts in steps.py * feat(config): add auxiliary model and provider environment variables to test setup |
||
|
|
92d95dee68 |
feat(memory): add observation memory lifecycle (#259)
* feat(memory): add observation memory lifecycle Add file-backed observation memory with deterministic markdown records, structured record_observation tooling, startup indexing, and profile/observation prompt guidance. Launch post-turn and post-subagent EvoMemory workers through LangGraph dev so completed runs can update profile memory, save durable observations, and write subagent execution summaries without blocking the active agent. Wire memory middleware into the main agent, subagents, async graphs, TUI status reporting, worker activity accounting, and observation-aware research prompts, with regression coverage for storage, lifecycle scheduling, graph registration, status display, and stream reset behavior. * fix(cli): sync background agent server on resume Resume flows now need to keep the LangGraph dev background server aligned with the active workspace even when async subagents are disabled. EvoMemory workers use that server too, so gating resume-time sync on enable_async_subagents could leave workers pinned to the launch workspace after resuming a thread from another workspace. Run workspace sync unconditionally for Rich CLI and Textual resume paths, while preserving WorkspaceMismatchError handling so failed sync aborts the resume before mutating the active thread or workspace. Propagate aborted resume callbacks through the command UI so channel-issued /resume commands do not send false success or history output. Channel slash dispatch now treats CommandManager-caught command errors as command errors and skips completion hooks for those failed commands. Add regression coverage for disabled async subagents, callback aborts, and channel command error reporting. * fix(cli): prepare serve resume workspace before adopting Load the resumed workspace agent and sync the background server as a single pre-adoption step. Restore the previous active workspace if preparation fails so serve mode keeps using the old session consistently. * fix(memory): untrack abandoned worker status watches Stop treating watcher shutdown as confirmed worker completion. Terminal worker statuses still count memory deltas, while poll failures or watcher setup failures now remove the active run without crediting partial outputs. * fix(cli): report channel command failures accurately Treat command_error as a None sentinel so empty error strings still fail, and let TUI resumes continue only on non-mismatch background-server sync failures while reporting degraded mode. * fix(stream): clear memory counters for resume streams Reset completed-memory counters for every new agent stream, including Command-based HITL and resume streams, so saved-memory indicators do not leak across turns. * docs(tools): make observation recording guidance conditional Clarify that agents should call record_observation only when the observation tool is available, preserving the existing durability and usefulness criteria. * feat(config): add controls for profile and observation memory Add config flags for profile memory, observation memory, observation writer placement, and background memory workers. Wire the controls through main agents, subagents, EvoMemory middleware, and memory lifecycle workers so observation writes can be assigned to the live agent, subagent worker, both, or neither. Keep turn memory workers profile-only and make prompts reflect the available observation read/write paths. Skip langgraph dev startup when neither async subagents nor memory workers need the background server. Add coverage for config parsing, prompt gating, middleware wiring, and worker tool availability. * test(cli): include memory defaults in serve config stubs * fix(memory): offload async worker launch blocking calls Run the langgraph-dev health check and memory-output snapshot in worker threads from the async EvoMemory launcher so it does not block the event loop. * chore(memory): harden turn worker subagent guardrail * chore(memory): refresh profile context per request * fix(memory): offload async profile file reads * fix(memory): offload async worker completion accounting |
||
|
|
d348076f40 | Add runtime context middleware (#255) | ||
|
|
9285c6dad8 |
Migrate memory middleware to profile files (#253)
* feat(memory): migrate to profile memory files * chore(stream): read profile headings from templates * fix(display): keep assistant responses if response_text has started * fix(memory): do not treat failed bootstraps as profile creation * chore(memory): unlink blank legacy memory * fix(memory): resolve project_id once * fix(memory): preserve unreadable profile files * chore(tui): render streamed narration inline with tool timeline Update the TUI streaming timeline so assistant text emitted before or between tool calls is rendered inline where it occurs, rather than being kept as a single answer bubble above or below the tools. If the model begins an assistant response and then emits another tool call, the provisional response is converted into inline narration before that tool. The final assistant message then renders only the remaining response suffix, avoiding duplicate text in the completed transcript. Stop/cancel handling now preserves any active inline narration, appends the visible stopped marker only to the remaining displayed segment, and still returns the full normalized stopped response for channel callers. Completed tools continue to collapse while long runs are active, but expand again when the turn reaches a final state so the completed transcript shows the full tool timeline. * fix(stream): preserve narration around tool timelines Keep assistant narration attached to the tool call that follows it instead of folding all streamed text into the final answer block. Track narrated response segments in stream state, render them before their corresponding regular or task tool entries, and keep final answers limited to the response suffix that has not already been shown inline. Preserve narration across normal completion, stop/error final frames, sub-agent task calls, and collapsed live tool summaries. Add regression coverage for pending tools, completed tools, sub-agent task delegations, collapsed completed/running tool summaries, and final stop frames. * fix(tui): finalize inline narration transitions * test(memory): use canonical project id helper |
||
|
|
a13904185d |
Feat/sandbox execute timeout (#243)
* feat: implement configurable sandbox execute timeout and enhance recovery instructions * feat: add background process management tools and middleware for sandbox execution * feat: enhance background process management with completion notifications and deduplication * feat: enhance sandbox execution timeout validation and update related messages * feat: enhance background process management with thread-specific completion notifications and HITL approval handling * test: assert completion notification waits for process finish timestamp |
||
|
|
7959495a13 |
feat(deploy): add EvoSci deploy subcommand (#228)
* feat(deploy): implement standalone LangGraph server and CLI command for deployment * feat(deploy): enhance port validation and environment variable management for deployment * Refactor langgraph dev deployment and introduce workspace sidecar protocol - Updated the deployment mode handling in `server.py` to use a single environment variable `EVOSCIENTIST_DEPLOY_MODE` with values `full` and `stripped`. - Enhanced the `manager.py` to implement a workspace fingerprint sidecar, allowing cross-process reuse of langgraph dev instances while ensuring workspace consistency. - Introduced functions to write and read the workspace sidecar, with error handling for missing or corrupt data. - Added tests for the workspace sidecar functionality, including validation of the JSON schema and ensuring proper error handling for workspace mismatches. - Updated existing tests to reflect changes in deployment mode handling and added new tests for signal handling during shutdown. - Ensured that cleanup routines remove the workspace sidecar alongside the PID file during shutdown. * fix(langgraph): improve workspace sidecar checks for process ownership and stale handles |
||
|
|
331056cdc8 |
feat(middleware): upgrade deepagents 0.5.7 → 0.6.2 (#231)
* feat(middleware): add CodeInterpreterMiddleware with project-specific configuration chore(config): increase checkpoint retention limit for runaway conversations fix(tests): update database schema references from 'blob' to 'value' chore(deps): update deepagents dependency to include quickjs support * feat(deepagents): update to version 0.6.1 and add optional dependencies for quickjs * feat(sessions): improve error handling for message deltas and update Overwrite type check * Enhance PruningCheckpointer with DeltaChannel Awareness - Introduced a new pruning strategy in `_prune_after_put` to preserve the `_DeltaSnapshot` chain during checkpoint pruning. - Implemented methods to fetch recent checkpoint IDs and walk to snapshot ancestors, ensuring that necessary checkpoints are retained. - Updated SQL queries to handle checkpoint and write deletions more efficiently. - Added comprehensive tests for DeltaChannel-aware pruning, ensuring that the pruning logic correctly handles various checkpoint scenarios, including those with and without snapshot seeds. - Refactored `_load_checkpoint_messages` to utilize the new saver interface, improving message reconstruction from checkpoints. * feat(tests): add migration sweep test to preserve snapshot ancestor * feat(sessions): enhance checkpoint retrieval to prevent transcript leakage in multi-agent scenarios * feat(middleware): enhance CodeInterpreterMiddleware with configurable timeout and result character limit feat(config): add CodeInterpreterMiddleware tuning parameters to EvoScientistConfig feat(sessions): implement inline message delta reducer for improved message handling * feat(dependencies): update deepagents version to 0.6.2 in pyproject.toml and uv.lock |
||
|
|
c407d2e20f |
Fix/async subagent model switch (#217)
* feat(middleware): add ConfigurableModelMiddleware for dynamic model resolution - Introduced ConfigurableModelMiddleware to resolve chat models from RunnableConfig.configurable on each call. - Updated middleware initialization to include ConfigurableModelMiddleware. - Enhanced context editing middleware tests to verify presence of ConfigurableModelMiddleware. - Implemented tests for ConfigurableModelMiddleware to ensure correct model overriding and caching behavior. - Added tests for deepagents model-passthrough patch to verify configuration injection in async tasks. * feat(async-subagent): update middleware handling to prevent deadlocks in async sub-agents * style: Refactor code formatting for improved readability in patches and test files * refactor: streamline middleware construction and improve async handling in ConfigurableModelMiddleware * fix: remove unused request parameter from _read_model_override function * refactor: improve async handling in _ClientProxy and enhance logging in ConfigurableModelMiddleware test: add behavior test to ensure AskUserMiddleware is excluded in async subagent mode |
||
|
|
80f1f4fa0f |
feat: Implement async sub-agent auto-notification system (#214)
* feat: Implement async sub-agent auto-notification system - Added async notifier functionality to handle notifications for sub-agents reaching terminal states. - Introduced `AsyncTaskNotification` dataclass for structured notification data. - Implemented `watch_run_and_notify` to monitor agent runs and enqueue notifications. - Created `spawn_watcher` to manage watcher tasks and ensure proper cancellation of previous watchers. - Developed `consume_notifications` to process notifications, deduplicate them, and format messages for LLM. - Added tests for notification handling, including draining, deduplication, and formatting. - Patched deepagents to integrate the new watcher functionality into start and update tools. * Enhance async notifier with per-thread notification routing and error handling - Introduced `origin_cli_thread_id` to `AsyncTaskNotification` for routing notifications back to the originating CLI session. - Implemented per-thread notification queues to handle notifications based on the originating thread. - Updated `has_pending_notifications` and `drain_notifications` to respect thread-specific queues. - Enhanced `watch_run_and_notify` to detect in-band error events from the SSE stream and handle clean exits. - Modified tests to verify the new notification routing behavior and ensure proper handling of notifications across threads. - Added a fixture to restore the async watcher patch state in tests to prevent state leakage. - Updated deepagents patching to capture the main agent's CLI thread ID for notification routing. * feat: Enhance async notifier with thread-specific watcher management and notification filtering * test: Enhance notification draining logic for cleaner test setup * refactor: Remove summary field from AsyncTaskNotification and update related tests * feat: Enhance async notification handling with target thread ID support * Refactor async notifier and middleware for improved task management - Removed the no-op shutdown watcher loop from async_notifier.py as it is no longer needed. - Updated watch_run_and_notify to clarify notification handling and race conditions. - Cleaned up shutdown handling in commands.py, interactive.py, and tui_interactive.py by removing obsolete shutdown watcher calls. - Deleted the deepagents async watcher patch from patches.py, transitioning to a new middleware approach. - Introduced AsyncWatcherMiddleware to handle async task notifications directly during tool calls. - Updated tests to validate the new middleware functionality and ensure proper watcher spawning and cancellation. - Enhanced test coverage for async watcher middleware, including edge cases and error handling. * feat(tests): add fixture to reset notifier state before each test |
||
|
|
9e51ec6fdd |
feat(cmd): add /model-fallback command (#196)
* feat(cmd): add /model-fallback command * fix: apply feedback * fix: lock usage with _fallback_chain * fix: apply feedback * fix: apply feedback * feat: add tests * fix: tests * Update EvoScientist/middleware/model_fallback.py Co-authored-by: dinos <dinospk1999@gmail.com> --------- Co-authored-by: dinos <dinospk1999@gmail.com> |
||
|
|
f41584e10b |
Refactor sub-agent architecture and introduce async support (#200)
* Refactor sub-agent architecture and introduce async support - Removed the legacy subagent.yaml file and replaced it with individual YAML files for each sub-agent in the subagents directory. - Updated the load_subagents function to support both directory and single file layouts for loading sub-agent configurations. - Added new langgraph_dev module for managing async sub-agent lifecycle and deployment. - Created graphs for async sub-agents (writing-agent, data-analysis-agent) and updated langgraph.json for deployment. - Introduced new sub-agent definitions for planner, research, debug, code, and writing agents with appropriate system prompts and configurations. - Enhanced package data inclusion in pyproject.toml to accommodate new sub-agent YAML files. * Refactor code for improved readability by consolidating conditional statements and formatting * feat: enhance async sub-agent support with workspace synchronization and user feedback - Added console status messages during async sub-agent server startup and workspace synchronization to improve user experience. - Implemented a new WorkspaceSyncWidget for live feedback during workspace sync operations. - Updated onboarding to reject occupied ports and ensure proper workspace handling for async sub-agents. - Introduced locking mechanisms to manage concurrent access to langgraph dev processes and workspace states. * feat: add async sub-agent configuration and server management functions * feat: improve port occupation handling and log file management in start_langgraph_dev * feat: enhance async sub-agent handling and introduce comprehensive tests - Updated `_maybe_swap_async_subagents` to improve async sub-agent management, ensuring internal flags are stripped before handoff. - Enhanced port management in `onboard.py` to allow reuse of occupied ports if already running by the same service. - Introduced file locking in `manager.py` to prevent race conditions during concurrent CLI invocations. - Added new tests for async sub-agent swapping and langgraph manager functionalities to ensure reliability and correctness. - Updated dependencies in `pyproject.toml` to include `psutil` and `filelock`. * fix(docs): clarify sub-agent configuration in README * test(manager): isolate _PID_DIR + tighten reuse-path assertion Addresses CodeRabbit review on tests/test_langgraph_manager.py: - Patch _PID_DIR to tmp_path so the FileLock setup in ensure_langgraph_dev doesn't mkdir the user's real ~/.config/evoscientist/ dir as a test side-effect. - Tighten "result is None or hasattr(result, 'poll')" to a strict "result is None" — the reuse path returns None unconditionally, so the OR clause was hiding potential regressions. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(manager): clean up stale PID file when unrelated process reuses PID * feat(tests): add validation tests for async flag in load_subagents * fix(load_subagents): restrict to .yaml files and clarify configuration handling * fix(load_subagents): improve error handling for non-dict specifications in YAML * feat(onboard): add "LangGraph Port" step to onboarding process * feat(langgraph): add concurrency configuration for langgraph dev workers * feat(async-subagents): enhance MCP tool routing for async sub-agents * fix(manager): update exception handling for connection errors and prevent zombie processes --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
29e7fd383d |
fix/cli hitl (#202)
* feat(display): enhance approval prompt with questionary for better navigation * feat(cli): add HumanInTheLoopMiddleware for user approval in main agent |
||
|
|
3831198f19 |
fix(chat): resolve model switch lag by tracking model/provider key (#180)
* fix(chat): resolve model switch lag by tracking model/provider key for cache invalidation * style: format set_chat_model function for improved readability * fix(chat): improve model switching logic to prevent unnecessary cache rebuilds * fix(model): ensure globals are restored on agent load failure to prevent model switch issues |
||
|
|
65828e9666 |
perf(cli): cut startup latency and defer MCP loading to the background (#171)
* perf(cli): cut startup time of `evosci --help` from ~2.2s to ~0.3s
Module-level imports were eagerly pulling in langchain.chat_models (with
the whole anthropic/openai/google stack), langgraph, textual, and
prompt_toolkit on every invocation — even for `--help` or `config list`.
Defer those with PEP 562 `__getattr__`, using `lazy_loader.attach` (SPEC-1,
the scientific-python standard) where it's a clean attach pattern:
- `EvoScientist/llm/__init__.py`: attach `models` lazily so importing
`context_window` from this package no longer drags in langchain.
- `EvoScientist/stream/__init__.py`: attach display/events lazily; split
the shared Rich `Console` singleton into a new lightweight
`stream/console.py` so callers that only need `console` skip the
`stream.events` → `langchain_core.messages` chain.
- `EvoScientist/cli/__init__.py`: hand-rolled `__getattr__` (reaches into
`..stream.state`, which `lazy_loader` doesn't cover) so `commands` and
`app` are the only eager loads.
- `EvoScientist/cli/commands.py`: move `cmd_interactive`/`cmd_run` to
in-function imports so prompt_toolkit + textual only load when the
interactive path actually runs.
- `EvoScientist/cli/_constants.py`: read `AGENT_NAME` on demand so
`build_metadata` doesn't eagerly import `sessions` (langgraph/aiosqlite).
Adds `lazy-loader>=0.5` as a dependency.
* feat(cli): defer MCP tool loading with live per-server progress
The CLI was blocking ~5 s on MCP tool enumeration before the first
prompt appeared. Move the agent construction off the event loop and
surface per-server progress so the user can interact immediately and see
what's happening.
MCP client:
- Add an `on_progress` callback to `load_mcp_tools` / `aload_mcp_tools`
/ `_load_tools` emitting `start` / `success` / `error` events per
server.
- Fan connection attempts out with `asyncio.gather` so latency no longer
scales linearly with server count; cap simultaneous attempts at
`_MAX_CONCURRENT_CONNECTIONS` (8) via a semaphore so a big stdio fleet
doesn't spawn every subprocess at once.
Agent wiring:
- Plumb `on_mcp_progress` through `create_cli_agent` / `_load_agent` /
`load_mcp_and_build_kwargs` so CLI and TUI can plug in collectors.
CLI (`cmd_interactive`):
- Run `_load_agent` in a background thread via `asyncio.to_thread`; the
prompt and banner render immediately.
- `_await_agent_ready()` awaits the task before each agent-using site
(first turn, channel messages, `/channel`, `/compact`). Raises if
called without a prior `_start_agent_load` instead of silently
reloading without the SQLite checkpointer.
- Pre-prime the progress dict from `load_mcp_config()` so the
bottom-toolbar's `N/M` denominator is stable from the first render.
- Wrap `session.prompt_async` in `patch_stdout(raw=True)` so
`console.print` from the worker-thread progress callback lands cleanly
above the prompt as inline chat messages instead of stomping the
prompt cursor.
TUI (`EvoTextualInteractiveApp`):
- Same background load + `_await_agent_ready()` gates on every
`self._agent` read.
- New `MCPLoaderWidget` mounted at the top of `#input-shell` shows a
header with `N/M` and one live row per server (spinner → ✓ / ✗ with
tool count or error detail). On completion:
- all-clean loads auto-dismiss ~2.5 s later;
- cache hits (no events ever fired) dismiss immediately rather than
flashing a misleading "0/N loaded";
- failures keep the widget mounted so the user can read the errors.
- `dismissed` property lets the app clear its ref so late events from
slow servers become no-ops. The error branch of `_on_agent_loaded`
also settles the widget so a load failure can't leave the spinner
animating forever.
- Chat input is `disabled` while MCP resolves — no placeholder hack, no
"waiting…" system message.
Shared:
- Hoist braille spinner frames to `status_bar.SPINNER_FRAMES` and import
them in the TUI widget so CLI and TUI animate in sync.
Tests:
- Extend `test_agent_mcp_cache` fakes to accept the new `on_progress`
kwarg.
- New `TestLoadToolsProgressCallback` in `test_mcp_client` exercises the
event sequence for success/failure/mixed fleets, verifies a buggy
callback doesn't break the load, and asserts the semaphore caps
in-flight connections.
* style: ruff
* chore: update uv.lock
* chore: uv.lock
* fix: coderabbit issues
* style: fmt
* fix: move _await_agent_ready inside try block
* fix(tui): auto-dismiss MCP loader widget on failure
The widget was designed to stay mounted on failure so the user could
read error detail, but since it's pinned above the input it never went
away in practice — just permanent banner clutter.
Auto-dismiss on failure too, with a longer grace (12s vs 2.5s) so the
error summary stays readable.
* fix: address second coderabbit pass
- Channel handlers (CLI + TUI): catch agent-load failures so the
channel request doesn't hang; CLI moves `_await_agent_ready()`
inside the existing try/except, TUI catches explicitly and calls
`_set_channel_response` with the error.
- Stale background loads: `prev.cancel()` only stops the asyncio
wrapper, not the thread running `_load_agent`. Added a generation
token (`agent_load_id` / `self._agent_load_id`) and gated both
progress and completion callbacks on it so a superseded load can't
clobber the current session's state or UI.
- TUI prompt lifecycle: added `_agent_load_pending()` and gated the
`_process_channel_message` / `_handle_command` finally blocks on it
so `/new` or `/resume` invoked from a command keeps the prompt
disabled until the fresh load settles.
- TUI readiness failures: `_run_turn` and `_handle_command` now
catch exceptions from `_await_agent_ready()` and surface a
"Agent failed to load: …" system message instead of letting the
exception escape into Textual's traceback panel.
* refactor(cli): share background agent loader between CLI and TUI
The CLI and TUI were carrying near-identical copies of the same
background-load state machine: the `agent_task`, the `agent_load_id`
generation token, the gated progress/completion callbacks, and the
per-server progress dict. Every CodeRabbit finding on that lifecycle
had to be fixed in both files.
Extract it into `cli/_agent_loader.py`:
- `MCPProgressTracker` — owns the `server -> (state, detail)` dict;
exposes `prime`, `record`, `snapshot`, `totals`.
- `BackgroundAgentLoader` — owns `agent`, the in-flight task, and the
generation token. Exposes `start(**loader_kwargs)`, `await_ready()`,
`is_pending`. Internally gates all progress/completion callbacks by
generation so a superseded load can't clobber the current session.
UI-specific rendering plugs in via `on_progress` / `on_success` /
`on_failure` callbacks.
Both surfaces now just wire their UI hooks; the loader file holds no
Rich / prompt_toolkit / Textual dependencies. Net -345 lines from
`interactive.py` + `tui_interactive.py`; +20 unit tests pinning the
lifecycle (generation filtering, cache-hit short-circuit, failure
reset, progress ordering).
* refactor(cli): make _on_done the sole authority for agent state transitions
await_ready no longer sets self.agent — it just awaits the task and
reads what _on_done already wrote. Eliminates the dual-write overlap
(asyncio guarantees done-callbacks fire in registration order).
* fix(tui): let users type during MCP load, only block on send
Remove prompt-disabling during background agent load — the TUI now
matches the CLI approach where the input stays enabled and only gates
on await_ready() at submit time. The MCPLoaderWidget still provides
visual feedback that loading is in progress.
* fix(loader): preserve real load error on await_ready; dedup failure message
CodeRabbit flagged two issues with the new loader:
1. After a failed load, `_on_done` nulled `self._task`, so the next
`await_ready()` hit the "before start()" branch and the CLI wrapper
remapped it to a misleading "checkpointer not available" message —
losing the real exception (bad MCP config, network, etc.).
Keep `_task` set on failure so `await_ready` re-raises the real
exception. Added `needs_restart` so TUI's auto-retry check stays a
one-liner and doesn't need to reach into task internals.
2. TUI reported each load failure twice: once from
`_on_agent_load_failure` (the done-callback) and once from each
caller of `_await_agent_ready` (`_run_turn`,
`_process_channel_message`, `_handle_command`) catching the re-raise.
`_on_agent_load_failure` is now the sole local reporter; callers
just handle control flow (return cleanly, set channel response to
unblock remote).
* fix(cli): wire /model handler through the agent loader
The /model command from main (merged via f1f0d7c) still reached for
`state["agent"]` (CLI) and `self._agent` (TUI) — both removed by the
background-loader refactor. CLI raised KeyError on first invocation;
TUI raised AttributeError. Writes to the old fields also had no effect
because every other code path now reads from `agent_loader.agent`, so
the model switch would have silently failed.
Route everything through the loader: `await _await_agent_ready()` up
front so /model doesn't race with the initial background load, build
the `CommandContext` with the current agent, and sync `ctx.agent` back
into `agent_loader.agent` (plus channel globals) when the command
replaces it.
* fix(cli): isolate progress callback, capture awaited agent, gate by requires_agent
Three CodeRabbit findings on the loader + command dispatch path:
- Wrap ``_on_progress`` in try/except inside the loader's gated wrapper
so a buggy UI adapter can't bubble into ``loader_fn`` and fail the
whole background load. The MCP client already protects this, but
defence-in-depth keeps the loader self-contained.
- In CLI channel + main-loop streaming, capture the agent returned by
``_await_agent_ready()`` and pass that into ``run_streaming`` rather
than reading ``agent_loader.agent`` after a subsequent ``await``.
A concurrent ``/new``/``/resume``/``/model`` could have swapped it.
- Add ``requires_agent: ClassVar[bool] = False`` to ``Command`` and
mark ``/compact``, ``/model``, ``/channel`` as ``True``. TUI dispatch
sites (channel and keyboard) now check ``cmd_manager.resolve(...)``
and only wait for readiness when the command actually needs the
agent. ``/mcp add``, ``/skills``, ``/new`` etc. no longer deadlock
behind a failing MCP load they are meant to fix.
* fix(cli): guard sync-back, subcommand-aware gating, /model adopt-path
Three CodeRabbit findings on command dispatch:
- ``_handle_command`` unconditionally synced ``ctx.agent`` back into
``agent_loader``. For non-agent commands ``ctx.agent`` is ``None``,
so ``/threads`` / ``/mcp`` / ``/skills`` (etc.) could clobber a valid
loaded agent — and rebind channel globals to ``None``. Guard the
sync on ``ctx.agent is not None``.
- ``/channel status`` and ``/channel stop`` don't touch ``ctx.agent``
but the class-level ``requires_agent = True`` blocked them behind
agent readiness. Added ``Command.needs_agent(args)`` (defaults to
``requires_agent``) so ``/channel`` can override with subcommand
awareness; kept the class flag for the common case.
- ``/model`` builds a new agent from scratch, it never reads the
existing one — gating it on readiness meant a broken provider
blocked the command that would fix it. Flipped it to
``requires_agent = False`` and added ``BackgroundAgentLoader.adopt``
so the UI can seat the replacement and supersede any in-flight
load (the generation token keeps a late completion from clobbering
the adopted agent).
Bonus cleanup: ``CommandManager.resolve`` now returns
``(command, args)`` so callers can invoke ``needs_agent`` without
re-implementing ``shlex`` parsing.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
|