115 Commits

Author SHA1 Message Date
m4 4c338ed914 fix(merge): resolve integration gaps found by running the v0.3.0 test suite
Post-merge validation fixes (upstream v0.3.0 + Ai4Sci fork):

- llm/patches.py: restore the two module-level patch calls the merge dropped
  (_patch_openai_empty_sse_keepalive, _patch_deepagents_extracted_document_text)
  and make _is_ccproxy_codex accept an explicit base_url/api_key so the
  invocation plan can classify an endpoint without mutating the process env.
- llm/models.py: an explicit per-call plan now wins over
  EVOSCIENTIST_USE_RESPONSES_API (env is only a default), an explicit caller
  `reasoning` block survives an explicit use_responses_api=False, and the
  third-party (openrouter) default effort stays the fork's fixed `medium`.
- EvoScientist.py: sub-agent stacks pass NO_OP_SINK as `events` instead of None.
- middleware/error_normalization.py: platform-generated diagnostics
  (ModelOutputTruncatedError) keep their actionable text while provider SDK
  errors still get the canned redacted message.
- pyproject.toml: hold google-genai 1.x (langchain-google-genai>=4.3.7,<4.4)
  because llm/gemini_interactions.py drives the 1.x Interactions API; this is
  also what deepagents 0.7.13 requires.
- config/settings.py: restore upstream's use_responses_api config field.
  `reasoning_effort` stays deleted on purpose — Ai4Sci keeps reasoning an
  invocation-plan parameter, never a deployment-env override.
- tests: align upstream tests that encode replaced behaviour (ccproxy
  responses-api context, reasoning-effort-overrides-env, fingerprint coverage)
  with the fork's contracts.
2026-09-13 16:57:03 +08:00
m4 470cf75722 merge: bring upstream v0.3.0 (72 commits) into Ai4Sci fork
Merged upstream/main (418abca, release v0.3.0) into our fork on a
dedicated branch. 21 conflicting files resolved; main worktree untouched.

Resolution policy and key decisions:
- Keep Ai4Sci runtime endpoints, durable dispatch, workspace scopes and
  the HITL/DynamicReview approval chain (approval path is product-critical).
- Adopt upstream model registry (llm/registry.py): our 136 model entries
  are a strict subset of upstream's 180, so dropping our inline table
  loses nothing and gains 44 new models.
- Adopt upstream native EvoChatDeepSeek; drop our obsolete
  _patch_deepseek_reasoning_passback monkey patch.
- Keep our six patches.py additions, ported onto upstream's new
  _OpenAICompatContent class: stable tool-call ids, tool-history
  sanitization, drop_reasoning_metadata, empty-SSE keepalive,
  extracted-document-text patch, _has_assistant_tool_protocol.
- Keep our skill-budget middleware path (skills=None) instead of passing
  skills through, to avoid double loading.
- Keep sanitized error labels (_safe_error_label) while adopting
  upstream's injected MiddlewareEventSink for fallback narration.
- Keep port 3076 and the LANGGRAPH_SERVER_URL override; adopt upstream's
  host/probe-host handling and CONFIG_DRIFT_SINCE_LAUNCH.
- Adopt upstream dependency stack: deepagents 0.7.6, langchain-quickjs
  0.3.7, langgraph-api 0.14; keep our extra deps (rfc8785, pillow,
  firecrawl-anydoc, nest-asyncio).
- Align call sites with upstream APIs: create_tool_selector_middleware
  now takes events= instead of track_stream_selection=.
2026-09-13 16:07:27 +08:00
Xi Zhang b36c19a22a fix(llm): honor openrouter_app_title only alongside a custom referer (#453)
OpenRouter keys app pages by HTTP-Referer; X-Title only renames that
page. A custom openrouter_app_title on the default referer therefore
renamed the shared EvoScientist app page for everyone. Force the default
title whenever the resolved referer is the default, silently, so usage
keeps being attributed to EvoScientist; a private fork still overrides
both together.
2026-09-07 17:22:18 +08:00
jfilipiuk 3d6cc959e3 fix: cut deployed-graph rebuild from 15s to <1s by deferring observation-index reads (#427)
* fix: defer eager observation-index build to first model call

* fix: add mtime-keyed cache to list_observation_documents to avoid re-parsing unchanged files

* fix: return a copy of the cached document list and strengthen the deletion test

* fix: copy cached document list on read and write to prevent caller mutations

* fix: split global and project cache to avoid duplicate parsing and cross-project invalidation

* fix: bump mtime explicitly in cache modification test for Windows NTFS resolution

* fix: bound project observation cache with LRU eviction

* fix: deduplicate path logic and strengthen cache typing

* fix: group cache tests under TestObservationCache with autouse fixture and fix f-string interpolation

* fix: reject non-positive observation cache cap in config validation

* fix: cache resolved observation docs and config cap to avoid repeated work

* fix: use st_mtime_ns and st_size in cache signature for NTFS reliability

* test: clear EVOSCIENTIST_MAX_CACHED_PROJECTS in test env cleanup fixtures

* test: cover per-file observation cache semantics

* fix: replace layered observation caches with per-file parse cache

* fix: serialize observation parse cache transactions
2026-08-21 06:49:32 +00:00
jax-novita bcee009917 Add Novita AI as an LLM provider (#422)
* Add Novita as an LLM provider

Registers Novita (novita.ai) as an OpenAI-routed provider, following the
same pattern as Requesty/Atlas Cloud/SiliconFlow: a base_url + API key env
var entry in _OPENAI_ROUTED_PROVIDERS, a handful of model registry entries
(DeepSeek/Qwen/GLM), onboarding wizard support (constants/steps/wizard/
helpers), a key validator using the auth-preflight sentinel pattern (Novita's
/v1/models endpoint returns the public catalog even for an invalid key, so
auth must be checked via a chat completion instead), and a host-to-provider
mapping entry for error attribution.

* Recommend Novita's current flagship models

The models listed for Novita were older ids that no longer reflect what
the platform leads with. Point the recommendations at the three current
flagships instead, each verified against api.novita.ai:

  moonshotai/kimi-k3              1M context, native vision
  zai-org/glm-5.2                 1M context, long-horizon agentic work
  deepseek/deepseek-v4-flash-0731 1M context, cheapest of the three

Context windows, output limits, input modalities and pricing were taken
from the live /openai/v1/models response rather than carried over.

* Keep branch CI workflow files unchanged (no workflow OAuth scope)

Co-authored-by: multica-agent <github@multica.ai>

* ci: restore workflow files to match main

---------

Co-authored-by: jax-novita <jax-novita@users.noreply.github.com>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Dinos Papakostas <dinospk1999@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-08-18 08:16:45 +00:00
Xi Zhang 1b324906fd feat(onboard): add NVIDIA BioNeMo Agent Toolkit to recommended skill packs (#431)
Add a "BioNeMo Skills" entry to the onboard Step 7 checkbox so users can
opt into the 31 life-science skills (protein folding, docking, generative
chemistry, genomics, protein design) with one selection.

The pack installs through the existing GitHub shorthand path — no installer
change. Source points at the toolkit's plugin skills directory
(plugins/bionemo-agent-toolkit/skills), which is the flat aggregate of all
31 skills; installing from the repo root would only reach 14 because the
remaining skills sit three levels deep.

Closes #333
2026-08-17 11:09:36 +01:00
m4 5a581c78a2 feat: add scoped model runtime configuration
Build / build (push) Has been cancelled
Docker / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
Introduce provider, model, and invocation contracts with encrypted configuration persistence. Add web runtime fencing, route fallback, recovery middleware, workspace scoping, and comprehensive tests.
2026-08-14 22:03:04 +08:00
Xi Zhang fb329e4aaa perf: cut startup latency — lazy import surface, langgraph dev keepalive, indexed thread listing (#407)
* feat: add thread metadata index for improved performance in thread listing

- Implemented a new SQLite index on the `checkpoints` table to optimize thread listing queries by indexing relevant metadata fields.
- Updated the `list_threads` function to ensure the index is created if it does not exist.
- Added a test to verify the creation of the metadata index during thread listing.

feat: enhance workspace sidecar management with owner tracking

- Modified the workspace sidecar to include `owner_pids` to track the current process owners.
- Updated tests to validate the new owner tracking functionality and ensure proper behavior when managing workspace sidecars.

chore: introduce model registry for streamlined model management

- Created a new `registry.py` file to maintain a comprehensive model registry, including model names, IDs, providers, and routing tables.
- Added functions to retrieve models by provider and list available models, enhancing the modularity and maintainability of model management.

* feat: enhance workspace sidecar management and improve thread metadata indexing

* fix(tests): ensure sidecar correctly registers owner with original workspace and pid

* refactor: simplify workspace sidecar management by removing owner tracking

* feat(server): add commands to manage background langgraph dev server

- Introduced `server_app` for managing the langgraph dev server with commands to check status and stop the server.
- Enhanced workspace sidecar management to include configuration fingerprint for drift detection.
- Updated deployment functions to handle server configuration and state more effectively.

* feat(server): enhance server status command to display PID with stale record warning

* feat(langgraph_dev): exclusion-set config fingerprint, webui keepalive, unified stop guidance

* fix(cli): platform-specific manual-stop hint; document keepalive endpoint-change limitation
2026-08-11 08:37:58 +01:00
Xi Zhang 0c21a01f6f fix: default WebUI bind host back to loopback (#412)
* fix: change default bind host to loopback for security across all components

* fix: update documentation and tests for loopback host configuration and security warnings
2026-08-07 17:13:57 +01:00
Ziheng Zhang 33979e5371 fix(llm): support Volcengine Coding model aliases (#411)
* fix(llm): support Volcengine Coding model aliases

* refactor(llm): add Volcengine Coding provider

* style: format Volcengine Coding test
2026-08-07 14:30:23 +08:00
Xiaohui Yan 3c5cc831c0 Feat/configurable bind host (#402)
* feat: configurable bind host for WebUI and langgraph dev (refs #400)

WebUI mode was only reachable from the machine running it: the front-end
got no bind interface, and `start_langgraph_dev(...)` was called without a
host, so both servers stayed on loopback with no way to widen them.

Adds two config fields with deliberately different defaults:

  webui_host        = 0.0.0.0    front-end serves the app shell, no secrets
  langgraph_dev_host = 127.0.0.1  unauthenticated API, agent can run shell

The design hinges on separating bind address from client address. Only
bind() uses the configured interface; every consumer that *connects*
(health probes, occupancy checks, async sub-agent self-dispatch) goes
through the new `_probe_host`, which maps a wildcard bind back to
loopback and honors a pinned interface verbatim. `_can_bind_port` is the
one exception and binds the literal host, since it must replicate the
bind the server itself will attempt.

  - manager.py: `_probe_host`, `_is_loopback_host`, `_format_hostport`;
    host kwarg threaded through the probes and `start_langgraph_dev`,
    which now emits `--host` and propagates
    EVOSCIENTIST_LANGGRAPH_DEV_HOST to the subprocess
  - sdk.py: `langgraph_dev_url` tracks host as well as port;
    EvoScientist.py reuses it instead of an inline f-string
  - server.py: `--host` flag mirroring `--port`, plus a red PUBLIC BIND
    banner whenever the bind is not provably loopback
  - webui.py: forwards both hosts; the front-end is widened via HOSTNAME
    because @evoscientist/webui ships no --host flag — its bin launcher
    does `HOSTNAME: process.env.HOSTNAME || "127.0.0.1"`. The warning is
    gated on the backend host only, so the shipped front-end default
    doesn't print a banner on every launch

Verified end to end against a live server: requesting 0.0.0.0 yields a
socket listening on 0.0.0.0 with the health probe correctly resolved to
127.0.0.1, while the default still binds 127.0.0.1 only.

Note: webui_host defaulting to 0.0.0.0 is a behavior change — upgrading
users will find the front-end reachable from the LAN.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: default both bind hosts to 0.0.0.0, add --host and wizard host rendering (closes #400)

Completes the remaining items from #400.

  - `langgraph_dev_host` now defaults to 0.0.0.0, matching `webui_host`.
    Remote WebUI use needs both anyway (the UI reaches the backend from the
    browser, not server-side), so a loopback backend default just meant every
    remote user hit a silently failing UI. `_DEFAULT_HOST` and sdk's
    `DEFAULT_LANGGRAPH_DEV_HOST` follow, so there is one story about where
    these servers listen.

    SECURITY: this exposes an unauthenticated API whose agent can run shell
    commands. The red PUBLIC BIND banner consequently fires on every launch
    while exposed — kept deliberately, since the exposure is real and the
    escape hatch (`--host 127.0.0.1` / `config set langgraph_dev_host`) is
    only discoverable if we say so. READMEs now lead with the warning and
    document the SSH-tunnel alternative.

  - `EvoSci --host <ip>` on the WebUI launch path, driving both servers. In
    WebUI mode they are two halves of one surface; moving only one leaves the
    UI loading but unable to reach the agent. Blank values are dropped rather
    than written as an empty override that would beat the config file.

  - Onboarding wizard no longer prints hard-coded `http://127.0.0.1:{port}` /
    `http://localhost:{port}` (steps.py:160, :223) — both render the
    configured bind through `_base_url` / `_format_hostport`, so a pinned
    interface is reported honestly and a wildcard still shows loopback.

Verified against a live server: with no host argument at all, resolution
through EvoScientistConfig yields a socket listening on 0.0.0.0, a client URL
of http://127.0.0.1, and the warning gate returning True.

Still open and tracked separately: the front-end takes its backend URL from
browser input: `@evoscientist/webui` reads only HOSTNAME, PORT and
EVOSCIENTIST_LANGGRAPH_DEV_PORT, so advertising a backend URL needs a change
in that repo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: bump setup-uv v6 -> v9.0.0 to drop the deprecated node20 runtime

GitHub now warns that setup-uv@v6 targets Node.js 20 and is being forced
onto Node.js 24. v7.0.0 is the release that made that switch, so anything
>= v7 clears the warning; v9.0.0 is current.

Pinned to the full tag deliberately: setup-uv stopped publishing major and
minor tags in v8.0.0 as supply-chain hardening, so `@v9` and `@v8` return
404 and would fail the job outright. Releases are immutable from v8 on, so
the full tag is as tamper-proof as a SHA. Comment left in lint.yml because
"simplifying" this back to `@v9` is an easy and CI-breaking mistake.

actions/checkout@v5 is already node24 and needs no change.

Note: v9.0.0 flips the `prune-cache` default to false (upstream did this to
ease load on PyPI infrastructure). None of these workflows set it, so they
follow the new default and Actions cache usage may grow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(cli): correct --host help text and warn on public bind in non-WebUI modes

The --host help claimed "WebUI mode only", which is wrong in a way that
matters for security. `--host` writes `langgraph_dev_host` unconditionally,
and `_ensure_async_subagent_server` auto-starts that backend for tui / cli /
serve as well — the langgraph dev server is shared across UI modes. So the
flag narrows or widens the agent API in every mode, and only `webui_host` is
actually WebUI-specific. Reported against cli/commands.py.

The documentation error hid a real gap: the PUBLIC BIND banner lived only in
deploy/server.py and deploy/webui.py, so a plain `EvoSci` session bound
0.0.0.0 with no runtime signal whatsoever — and `--help` is opt-in, so
fixing the text alone would not surface it. Added the same banner to the
shared CLI path, gated on `is_async_subagents_available()`: ensure_langgraph_dev
fails soft (async degrades to in-process delegation), and warning about a
bind that never happened would be worse than staying quiet.

READMEs (EN + zh-CN) get the same correction — the warning block sat inside
the Desktop WebUI section and read as WebUI-scoped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(deploy): strip the config-derived bind host, not just the CLI one

`deploy()` only stripped the `--host` branch. When the flag was omitted,
`getattr(config, "langgraph_dev_host", ...)` flowed unstripped into
`_is_port_occupied`, `is_langgraph_dev_running`, `start_langgraph_dev` and
the banner. `run_webui` already strips unconditionally; this aligns the two.

Reachable because `deploy()` reads through `getattr` and is routinely handed
duck-typed config objects (tests, embedders) that never run
`EvoScientistConfig.__post_init__`, which is what normally normalizes these
fields.

Worst case was not just a bad bind: `_is_loopback_host(" 127.0.0.1 ")` is
False, so a padded loopback value would print a false PUBLIC BIND warning
while binding a string socket.bind() rejects outright — a security banner
saying the opposite of the truth.

Three regression tests added, each verified to fail against the old code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* style: apply ruff format to the bind-host changes

The Lint workflow runs both `ruff check` and `ruff format --check`; I had
only been running the former locally, so five files landed unformatted and
failed CI. Whitespace and line-wrapping only — no semantic change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(security): keep the langgraph dev backend on loopback by default

The backend is an unauthenticated API whose agent can run shell commands,
and it is auto-started in every UI mode (tui/cli/webui/serve/deploy) — so a
0.0.0.0 default put it on the network for users who never asked. Restore
127.0.0.1 as the default and make 0.0.0.0 an explicit opt-in.

webui_host keeps its 0.0.0.0 default: the front-end serves the app shell
only and holds no credentials. run_webui already prints a remote-backend
hint when the front-end is exposed and the backend is not.

Help text and both READMEs are reframed around widening rather than
narrowing; the escape-hatch tests are inverted to assert the public-bind
opt-in survives into argv.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 16:15:49 +08:00
Xi Zhang d249e320bd feat: unified HITL approval for sync + async sub-agents (closes #387) (#396)
* feat: implement guard for dangerous commands and enhance HITL interrupt handling

* feat: enhance HITL approval mechanism and introduce session auto-approve decisions

* feat: add refuse_delete option to backend and enhance async delete guards

* feat: simplify delete method in CustomSandboxBackend and clarify adelete behavior

* feat: enhance allow-list behavior for command resolution and add related tests
2026-07-31 10:23:05 +01:00
Xi Zhang f81a8b086e feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395
- Introduced TodoListMiddleware to the middleware stack for better task management.
- Updated HITL interrupt configuration to include 'delete' operations requiring approval.
- Implemented error handling for delete operations in read-only and memory backends.
- Enhanced approval prompt formatting to display file paths for delete actions.
- Added tests to ensure delete operations are correctly blocked or prompted for approval.
- Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality.
2026-07-30 10:41:12 +01:00
nb213 4ddf7ebe52 Add Atlas Cloud LLM provider (#388)
* Add Atlas Cloud LLM provider

* Add Atlas Cloud onboarding support

* fix(validators): update atlascloud key validation to handle insufficient balance case

---------

Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-07-29 14:09:44 +01:00
Thibault Jaigu 8cac50e2ef Add Requesty as an LLM provider (#346)
* Add Requesty as an LLM provider

* Address review: Requesty prompt caching, model ordering, key validation

- Declare Anthropic-style prompt caching for Requesty Claude models by
  default (mirroring the OpenRouter behavior), with an opt-out flag
  EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
  provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
  overrides native/OpenRouter models for names it shares with them
  (the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
  invalid/missing key (public catalog), so it cannot validate a key. Use a
  minimal authenticated /v1/chat/completions request instead (200 = valid,
  403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).

* Validate Requesty key against auth layer, not a specific model

The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-07-28 15:35:41 +01:00
jfilipiuk 10c032450e fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385)
* fix: propagate langgraph dev bind port into subprocess env for self-loop URL

* fix: keep parent env authoritative over workspace .env for mapped keys

* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins

* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS

* fix: merge .env via dotenv_values to close empty-value and RMW-race edges

* chore: align docstrings after .env-merge rework

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-28 12:26:22 +00:00
dinos 8b1451cdda refactor(runtime): centralize async bridges under an owned runtime (#376)
* feat(runtime): add application-scoped async runtime

* refactor(cli): use owned runtime for session stats

* refactor(onboard): use the owned async runtime

* docs(runtime): record async bridge ownership

* refactor(middleware): keep sync fallback synchronous

* refactor(mcp): load tools on an owned runtime

* refactor(cli): share owned runtime across entry points

* refactor(channels): make inbound sync bridge explicit

* refactor(stream): run Rich streaming on owned runtime

* chore(runtime): remove nest-asyncio dependency

* refactor(asyncio): require active loops in async code

* docs(runtime): document final event loop ownership

* fix(stream): cancel stalled owned streams

* fix(cli): recover cleanly from stream cancellation

* fix(runtime): drain executor work before shutdown

* fix(runtime): terminate cancelled shell process trees

* fix(models): let fallback bypass selector failures

* fix(cli): reset interrupt handling between turns

* docs: rm implementation spec

* fix(serve): cancel active turns during shutdown

* fix(runtime): protect settlement from waiter cancellation

* fix(backends): reject empty shell commands

* fix(runtime): terminate descendants after shell exit

* fix(mcp): keep standalone discovery off channel loop

* fix(cli): own and settle interactive prompt cancellation

* fix(serve): keep channel sends off runtime loop

* fix(stream): scope cancel context to iterator steps

* refactor(serve): require the owned async runtime

* fix(channels): keep interactive sends off runtime loop

* fix(selector): surface fallback without log spam

* test(runtime): normalize Windows shell marker

* fix(cli): serialize interactive session turns

* fix(shell): bound output drain after termination

* fix(ui): do not retry owned runtime failures

* fix(shell): allow signal-safe registry reentry

* fix(shell): avoid terminating reused process ids

* fix(channels): preserve streaming send order

* fix(cli): report runtime shutdown timeouts cleanly

* fix(mcp): guide async callers to async loader

* docs(runtime): clarify reserved async bridge APIs

* fix(runtime): bound code interpreter cleanup

* test(shell): use active Python for drain regression

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-27 14:17:57 +01:00
m4 3ce5614254 fix: harden tool-call protocol and fallback handling 2026-07-19 12:05:56 +08:00
m4 4fc74e7da7 EvoScientist Ai4Sci
Docker / build (push) Has been cancelled
Build / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
2026-07-14 22:07:14 +08:00
Mani Saint-Victor da6ca38d53 fix(llm): respect reasoning_effort setting on native OpenAI path (#321)
* fix(llm): respect reasoning_effort setting on native OpenAI path

The native OpenAI provider path hardcoded reasoning effort to xhigh for
gpt-5.4/5.5/codex models, silently ignoring the user's reasoning_effort
config setting. The OpenRouter path already honors the
EVOSCIENTIST_REASONING_EFFORT env var that settings.py exports from that
setting; this applies the same lookup on the native path, falling back
to the previous defaults when unset.

Adds a regression test and isolates the existing xhigh test from the
env var.

* fix(llm): preserve model reasoning defaults

* fix(llm): preserve GPT-5.6 reasoning default

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 10:42:55 +00:00
Zixin Dong f72f7b93d5 feat(llm): add OpenRouter app attribution headers (#339) (#344)
* feat(llm): add OpenRouter app attribution headers (#339)

Attach EvoScientist app-attribution at the shared model-init layer so all
OpenRouter calls are credited to the project. langchain-openrouter maps
app_url/app_title/app_categories -> HTTP-Referer / X-Title /
X-OpenRouter-Categories. Applied only for the openrouter provider, via
setdefault so explicit caller kwargs win. Configurable through new
openrouter_http_referer / openrouter_app_title / openrouter_app_categories
settings and their EVOSCIENTIST_OPENROUTER_* env vars.

Closes #339

* refactor(llm): centralize OpenRouter attribution defaults + cap categories

Address PR #344 review:
- Define the app-attribution default constants once in config/settings.py
  (the config fields and llm/models.py both use them) instead of duplicating
  the literals across the two modules.
- Reduce the default categories to creative-writing,personal-agent and cap the
  sent list to OpenRouter's 2-per-request limit, warning when a configured list
  exceeds it, so extras are dropped predictably (and surfaced) here rather than
  being silently truncated server-side.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 11:08:59 +01:00
dinos d2452c54d5 Refactor onboarding OAuth flow for auxiliary models (#337)
* refactor(onboard): shared flow for ccproxy providers

* feat(onboard): support oauth configuration for auxiliary models

* fix(onboard): reuse main model auth for same-provider auxiliary

* fix(onboard): reconcile oauth providers
2026-07-08 18:28:44 +00:00
dinos a7b9e175c1 fix(config): set config.yaml permissions to 0x600 (#336) 2026-07-08 19:25:01 +01:00
renaissancefieldlite f086d77756 Fix UTF-8 config reads on Windows (#318)
* Fix UTF-8 config reads on Windows

* test: cover utf8 production loaders

* fix: read and write settings as utf8

* Apply ruff formatting
2026-07-05 05:10:24 +00:00
dinos 2b244888ec feat(memory): autoskills (#319) 2026-07-03 11:16:33 +02:00
Xi Zhang 7ccfe68f3f feat: add scheduler functionality with cron-style task management (#306)
* feat: add scheduler functionality with cron-style task management

- Implemented a new scheduler subagent to automate recurring tasks using cron expressions.
- Enhanced the subagent factory to include the skill manager and auxiliary chat model for the scheduler.
- Created a YAML configuration for the scheduler with a detailed system prompt and toolset.
- Updated README files to include documentation on scheduled tasks and usage examples.
- Added tests for the scheduler, including command execution, scheduling tools, and middleware integration.
- Introduced new dependencies for timezone handling and ensured compatibility in the project configuration.

* fix(async-notifier): ensure fallback hint is used for unknown notification kinds

* feat: enhance scheduling functionality and improve system message handling
2026-06-25 17:31:23 +01:00
dinos 6eb467e70b feat(llm): make OpenRouter Anthropic prompt cache opt-out (#299)
* feat(llm): make OpenRouter Anthropic prompt cache opt-out

* fix(llm): restore truthy env flag helper
2026-06-18 15:09:22 +02:00
dinos f356de36a6 feat: memory retrieval (#281) 2026-06-16 09:14:29 +02:00
X-iZhang 49f23560fd chore: update version to v0.1.5 2026-06-10 22:37:52 +01:00
Xi Zhang c02be519f6 feat(dangerous-mode): implement real-filesystem access with safety ch… (#276)
* feat(dangerous-mode): implement real-filesystem access with safety checks

- Introduced a 'dangerous mode' allowing the agent to operate on the real filesystem.
- Updated command validation to bypass path confinement while enforcing a blocklist for privileged commands.
- Added warnings and guidelines for users when operating in dangerous mode.
- Enhanced configuration to support dangerous mode and ensure it implies auto-approval.
- Updated tests to verify the behavior of commands and configurations in dangerous mode.

* feat(dangerous-mode): enhance logging and environment management for dangerous mode

* feat(dangerous-mode): improve handling of dangerous mode with environment flags and enhance test isolation
2026-06-10 18:19:13 +01:00
dinos cd2baa9588 feat(models): add opt-in prompt caching support for anthropic via openrouter (#272)
* feat(models): add opt-in prompt caching support for anthropic via openrouter

* chore: don't coerce model_kwargs to dict
2026-06-09 15:13:49 +01:00
Xi Zhang 63969b596d Release/v0.1.4 (#266)
* feat(middleware): reposition code interpreter middleware in the stack

* feat(models): add qwen3.7-plus model entry and update context window comment

* feat(models): add qwen3.7-max and qwen3.7-plus model entries for DashScope

* feat(auxiliary): implement auxiliary model support for background tasks and tool selection

- Added auxiliary model configuration to EvoScientistConfig.
- Introduced _ensure_auxiliary_chat_model function to manage auxiliary model instances.
- Updated onboarding steps to include auxiliary model selection.
- Modified middleware to route tool selection to the auxiliary model when applicable.
- Enhanced tests to cover auxiliary model functionality and configuration.

* feat(steps): update UI backend selection options and descriptions

* Refactor code structure for improved readability and maintainability

* feat(patches): implement OpenRouter response reasoning item stripping to prevent multi-turn errors

* feat: update version to v0.1.4 in badges, README, and pyproject.toml; adjust skill counts in steps.py

* feat(config): add auxiliary model and provider environment variables to test setup
2026-06-07 00:52:59 +01:00
Eliot Drizzle 3563c1d94f Update source for Scientific Skills in steps.py (#265) 2026-06-06 14:53:14 +01:00
dinos 92d95dee68 feat(memory): add observation memory lifecycle (#259)
* feat(memory): add observation memory lifecycle

Add file-backed observation memory with deterministic markdown records,
structured record_observation tooling, startup indexing, and
profile/observation prompt guidance.

Launch post-turn and post-subagent EvoMemory workers through LangGraph
dev so completed runs can update profile memory, save durable
observations, and write subagent execution summaries without blocking
the active agent.

Wire memory middleware into the main agent, subagents, async graphs, TUI
status reporting, worker activity accounting, and observation-aware
research prompts, with regression coverage for storage, lifecycle
scheduling, graph registration, status display, and stream reset
behavior.

* fix(cli): sync background agent server on resume

Resume flows now need to keep the LangGraph dev background server
aligned with the active workspace even when async subagents are
disabled. EvoMemory workers use that server too, so gating resume-time
sync on enable_async_subagents could leave workers pinned to the launch
workspace after resuming a thread from another workspace.

Run workspace sync unconditionally for Rich CLI and Textual resume
paths, while preserving WorkspaceMismatchError handling so failed sync
aborts the resume before mutating the active thread or workspace.

Propagate aborted resume callbacks through the command UI so
channel-issued /resume commands do not send false success or history
output. Channel slash dispatch now treats CommandManager-caught command
errors as command errors and skips completion hooks for those failed
commands.

Add regression coverage for disabled async subagents, callback aborts,
and channel command error reporting.

* fix(cli): prepare serve resume workspace before adopting

Load the resumed workspace agent and sync the background server as a
single pre-adoption step. Restore the previous active workspace if
preparation fails so serve mode keeps using the old session
consistently.

* fix(memory): untrack abandoned worker status watches

Stop treating watcher shutdown as confirmed worker completion. Terminal
worker statuses still count memory deltas, while poll failures or
watcher setup failures now remove the active run without crediting
partial outputs.

* fix(cli): report channel command failures accurately

Treat command_error as a None sentinel so empty error strings still
fail, and let TUI resumes continue only on non-mismatch
background-server sync failures while reporting degraded mode.

* fix(stream): clear memory counters for resume streams

Reset completed-memory counters for every new agent stream, including
Command-based HITL and resume streams, so saved-memory indicators do not
leak across turns.

* docs(tools): make observation recording guidance conditional

Clarify that agents should call record_observation only when the
observation tool is available, preserving the existing durability and
usefulness criteria.

* feat(config): add controls for profile and observation memory

Add config flags for profile memory, observation memory, observation
writer placement, and background memory workers.

Wire the controls through main agents, subagents, EvoMemory middleware,
and memory lifecycle workers so observation writes can be assigned to
the live agent, subagent worker, both, or neither. Keep turn memory
workers profile-only and make prompts reflect the available observation
read/write paths. Skip langgraph dev startup when neither async
subagents nor memory workers need the background server.

Add coverage for config parsing, prompt gating, middleware wiring, and
worker tool availability.

* test(cli): include memory defaults in serve config stubs

* fix(memory): offload async worker launch blocking calls

Run the langgraph-dev health check and memory-output snapshot in worker
threads from the async EvoMemory launcher so it does not block the event
loop.

* chore(memory): harden turn worker subagent guardrail

* chore(memory): refresh profile context per request

* fix(memory): offload async profile file reads

* fix(memory): offload async worker completion accounting
2026-06-05 15:11:20 +01:00
X-iZhang d53bfa35c5 feat: update MiniMax model entries and context window for M3 variant 2026-06-02 00:00:52 +01:00
Xi Zhang fbd1d709ca feat: add WebUI mode support with related configuration and onboarding (#252)
* feat: add WebUI mode support with related configuration and onboarding steps

* feat: enhance WebUI port configuration to prevent conflicts with backend port

* feat: add support for fresh interactive session detection in WebUI
2026-06-01 12:01:08 +01:00
Xi Zhang a13904185d Feat/sandbox execute timeout (#243)
* feat: implement configurable sandbox execute timeout and enhance recovery instructions

* feat: add background process management tools and middleware for sandbox execution

* feat: enhance background process management with completion notifications and deduplication

* feat: enhance sandbox execution timeout validation and update related messages

* feat: enhance background process management with thread-specific completion notifications and HITL approval handling

* test: assert completion notification waits for process finish timestamp
2026-05-31 15:11:25 +01:00
X-iZhang 721a03c25b feat: update model version from claude-sonnet-4-5 to claude-sonnet-4-6 and related adjustments 2026-05-29 00:40:46 +01:00
Xi Zhang f75bfcda51 Add onboarding wizard with style and validation components (#241)
* Add onboarding wizard with style and validation components

- Introduced `style.py` for shared visual elements used in the onboarding wizard.
- Created `validators.py` for input validation, including integer and choice validators, and API key validation functions for various providers.
- Implemented `wizard.py` as the entry point for the onboarding process, managing user prompts and configuration steps.
- Added progress rendering and autosave functionality to enhance user experience during the onboarding process.

* feat(onboarding): enhance validation and configuration for onboarding wizard

- Added validation for UI backends, workspace modes, and providers in the onboarding command.
- Updated channel definitions to include secret field handling for sensitive tokens.
- Improved user prompts for required fields, ensuring sensitive data is masked.
- Introduced constants for valid providers, UI backends, and workspace modes to maintain consistency.
- Implemented tests to ensure alignment between constants and interactive choices in onboarding steps.

* feat(onboarding): improve WeChat account ID prompt and validation for newly enabled channels

* feat(onboarding): enhance WeChat backend credential prompts and validation

* feat(onboarding): refine WeChat backend credential prompts for wecom and wechatmp

* Refactor onboarding package for improved structure and clarity

- Simplified the onboarding package by removing unnecessary re-exports and consolidating public API to only include `run_onboard`.
- Updated `install_back_keys` to `install_navigation_keys` for clarity and consistency in the prompter module.
- Enhanced the `NonInteractivePrompter` class to support strict mode, allowing for better handling of non-interactive prompts.
- Adjusted the onboarding steps to utilize the new navigation keys installation method.
- Improved the `run_onboard` function to handle section implications based on user flags, enhancing the onboarding experience.
- Updated tests to reflect changes in imports and ensure compatibility with the new structure.

* feat(onboarding): enhance validation logic for non-interactive prompts

* refactor(onboarding): streamline onboarding module structure and enhance validation error handling

* refactor(onboarding): enhance config revert logic to preserve original file state

* refactor(onboarding): enhance tavily key validation and error handling in onboarding process
2026-05-28 12:42:49 +01:00
Ziheng Zhang d2283397a4 feat(feishu): scan-to-create QR onboarding + silence unsubscribed WS events (#239)
* feat(feishu): scan-to-create QR onboarding flow

Add a device-code flow against accounts.feishu.cn/oauth/v1/app/registration
that lets users scan a terminal QR code with Feishu / Lark mobile to
auto-create a PersonalAgent bot app with the required IM permissions
pre-attached. The poll endpoint returns app_id + app_secret, which the
onboarding wizard then writes into the channel config — no manual app
creation on open.feishu.cn required.

- channels/feishu/onboard.py: qr_register() public entry, init/begin/poll
  helpers, QR rendering via the soft qrcode dep, automatic feishu↔lark
  domain switch based on the scanning user's tenant_brand, and a
  best-effort bot probe to surface the bot name in the wizard
- channels/feishu/__init__.py: re-export qr_register (mirrors qq)
- config/onboard.py: offer "Scan QR code (recommended) / Enter manually"
  in the Feishu branch, ask for region (feishu vs lark), then call
  qr_register and populate feishu_app_id / feishu_app_secret /
  feishu_domain; add qrcode>=7.4 to the feishu pip extras

* fix(feishu): silently absorb unsubscribed WebSocket events

Feishu auto-subscribes PersonalAgent apps to many event types
(im.message.reaction.created_v1, message.read_v1, message.recalled_v1,
chat.member.*, ...) that EvoScientist doesn't register handlers for.
Without intervention, lark-oapi's dispatcher raises EventException
("processor not found, type: ..."), the WS client logs it at ERROR and
replies HTTP 500 on the frame, and Feishu marks the event as failed
and retries it.

The problem is amplified by _send_ack_reaction: every inbound message
triggers our own reaction, which Feishu echoes back as
reaction.created_v1, creating a continuous ERROR-log feedback loop and
pointless retries.

Wrap EventDispatcherHandler._do_without_validation after build() to
swallow "processor not found" EventExceptions (debug log + return None)
while letting all other errors propagate. Failure-safe: if lark-oapi's
internal API changes the wrapper degrades to the prior behavior rather
than breaking the channel.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-05-20 22:53:16 +08:00
Xi Zhang 331056cdc8 feat(middleware): upgrade deepagents 0.5.7 → 0.6.2 (#231)
* feat(middleware): add CodeInterpreterMiddleware with project-specific configuration

chore(config): increase checkpoint retention limit for runaway conversations

fix(tests): update database schema references from 'blob' to 'value'

chore(deps): update deepagents dependency to include quickjs support

* feat(deepagents): update to version 0.6.1 and add optional dependencies for quickjs

* feat(sessions): improve error handling for message deltas and update Overwrite type check

* Enhance PruningCheckpointer with DeltaChannel Awareness

- Introduced a new pruning strategy in `_prune_after_put` to preserve the `_DeltaSnapshot` chain during checkpoint pruning.
- Implemented methods to fetch recent checkpoint IDs and walk to snapshot ancestors, ensuring that necessary checkpoints are retained.
- Updated SQL queries to handle checkpoint and write deletions more efficiently.
- Added comprehensive tests for DeltaChannel-aware pruning, ensuring that the pruning logic correctly handles various checkpoint scenarios, including those with and without snapshot seeds.
- Refactored `_load_checkpoint_messages` to utilize the new saver interface, improving message reconstruction from checkpoints.

* feat(tests): add migration sweep test to preserve snapshot ancestor

* feat(sessions): enhance checkpoint retrieval to prevent transcript leakage in multi-agent scenarios

* feat(middleware): enhance CodeInterpreterMiddleware with configurable timeout and result character limit

feat(config): add CodeInterpreterMiddleware tuning parameters to EvoScientistConfig

feat(sessions): implement inline message delta reducer for improved message handling

* feat(dependencies): update deepagents version to 0.6.2 in pyproject.toml and uv.lock
2026-05-19 12:35:36 +01:00
dinos 4b0c91190a feat(llm): add dashscope-code provider for Alibaba Coding Plan keys (#225)
* feat(llm): add dashscope-code provider for Alibaba Coding Plan keys

Alibaba Cloud Bailian "Coding Plan" subscription keys (sk-sp-*) route
through a separate endpoint (coding.dashscope.aliyuncs.com/v1) that the
standard `dashscope` provider can't reach. Add a sibling provider entry
matching the zhipu/zhipu-code and moonshot/kimi-coding precedents, with
its own validator (the coding endpoint returns 404 on /models, so probe
via chat.completions instead).

Closes #224

* fix(llm): keep dashscope as default provider for qwen3-coder shortcut

The MODELS dict is built from _MODEL_ENTRIES via a last-write-wins dict
comprehension. The initial commit listed dashscope-code AFTER dashscope,
which silently flipped the bare `get_chat_model("qwen3-coder")` shortcut
to the coding endpoint — breaking standard sk-* keys.

Reorder to match the zhipu-code / zhipu precedent: coding endpoint first,
general endpoint last so the general endpoint wins the collision and
remains the default for the shared "qwen3-coder" short name.
2026-05-13 10:23:32 +01:00
Ziheng Zhang 89b0ecdbf3 feat(qq): add QR-code scan-to-configure onboarding for QQ Bot (#213)
* feat(qq): add QR-code scan-to-configure onboarding for QQ Bot

Adds a `qr_register()` flow that drives q.qq.com's create_bind_task /
poll_bind_result APIs so the wizard can auto-fill `qq_app_id` and
`qq_app_secret` after the developer scans a QR code with a bound QQ
account, falling back to manual entry on failure or cancel.

- channels/qq/crypto.py: AES-256-GCM helpers for decrypting the bot's
  client_secret returned by poll_bind_result.
- channels/qq/onboard.py: portal API client + polling loop.
- channels/qq/__init__.py: re-export `qr_register`.
- config/onboard.py: QQ branch in `_step_channels` that offers
  "Scan QR code" vs "Enter manually", and skips the manual prompt
  loop when a scan succeeded.

* style(qq): fix ruff lint errors in onboard.py

Move `import os` to the top-level import block (E402), drop the legacy
`typing.Optional`/`typing.Tuple` imports (UP035), and use the PEP 585/604
builtin generics (`tuple[...]`, `X | None`) for the few annotations that
still referenced them (UP006/UP045). No behavior change.

* fix(qq): harden QR onboard error paths and declare scan deps

Address review feedback on PR #213:
- Declare cryptography>=41.0 and qrcode>=7.4 in [qq]/[all-channels]
  extras and in _CHANNEL_PIP_DEPS so the scan flow no longer fails
  with an opaque ImportError on a fresh `evoscientist[qq]` install.
- Polling loop logs each _poll_bind_result failure and aborts after
  5 consecutive errors instead of silently spinning until the 600s
  timeout, restoring the documented Raises: RuntimeError contract.
- Wrap decrypt_secret in try/except so failures honor the
  None-on-failure contract instead of letting exceptions escape.
- Preflight `import cryptography` in the scan branch and offer
  install or fall back to manual entry.

* style: ruff format collapse two over-wrapped log/console lines
2026-05-08 12:08:52 +01:00
Ziheng Zhang 692dc491ac # feat(wechat): add personal-WeChat (iLink) backend with QR login (#212)
* feat(wechat): add personal-WeChat (iLink) backend with QR-code login

Adds a third WeChat backend alongside WeCom and Official Account:
``personal`` rides Tencent's iLink Bot long-poll gateway so a personal
WeChat account can act as a bot.  Credentials are obtained via QR-code
scan and persisted under ``DATA_DIR/wechat_personal/accounts/``.

- channels/wechat/personal.py: WeixinPersonalChannel + qr_login.
- channels/wechat/crypto.py: aes128_ecb_decrypt + parse_ilink_aes_key
  for the iLink CDN media protocol.
- channels/wechat/probe.py: validate_wechat_personal credential probe.
- channels/wechat/serve.py: --backend personal CLI + --qr-login flow.
- channels/wechat/__init__.py: factory dispatch on wechat_backend; pull
  in the new dependencies in the docstring.
- config/settings.py: wechat_personal_* fields.
- config/onboard.py: WeChat-backend picker + QR-scan flow in the wizard
  + personal-backend probe in _probe_channel.
- pyproject.toml / uv.lock: add qrcode + certifi to wechat & all-channels
  extras (aiohttp was already pulled in transitively).

* fix(wechat): address ruff failures and CodeRabbit review on personal-WeChat PR

- personal.py: drop unused imports (`field`, `PollingMixin`); replace
  `asyncio.TimeoutError` with builtin; hold references to background
  `asyncio.create_task` results so they aren't GC'd; wire `dm_policy`
  through `_process_message` (disabled/allowlist) so `wechat_personal_dm_policy`
  actually takes effect for DMs.
- onboard.py: import-check gate now validates the full WeChat dependency
  set (aiohttp, qrcode, Crypto, certifi) instead of only aiohttp; mask
  `WeCom Secret` and `MP App Secret` prompts via `questionary.password`;
  derive the QR-login hint path from `_account_dir()` instead of the
  hard-coded `~/.evoscientist/...`; stop copying the QR-login token into
  the main config (already persisted per-account on disk — copying broadens
  secret exposure and risks staleness).
- pyproject.toml: allow Chinese full-width punctuation in `allowed-confusables`
  for user-facing CN messages.

* style(wechat): apply ruff format

`ruff format --check` was failing CI on three files (one pre-existing in
`__init__.py` plus formatter-driven line-merges in the files touched by
the previous fix commit). Ran `ruff format` to bring them in line; both
`ruff check` and `ruff format --check` now pass.
2026-05-07 16:33:30 +02:00
Wiktor Cupiał 9e51ec6fdd feat(cmd): add /model-fallback command (#196)
* feat(cmd): add /model-fallback command

* fix: apply feedback

* fix: lock usage with _fallback_chain

* fix: apply feedback

* fix: apply feedback

* feat: add tests

* fix: tests

* Update EvoScientist/middleware/model_fallback.py

Co-authored-by: dinos <dinospk1999@gmail.com>

---------

Co-authored-by: dinos <dinospk1999@gmail.com>
2026-05-07 15:35:26 +02:00
Xi Zhang f41584e10b Refactor sub-agent architecture and introduce async support (#200)
* Refactor sub-agent architecture and introduce async support

- Removed the legacy subagent.yaml file and replaced it with individual YAML files for each sub-agent in the subagents directory.
- Updated the load_subagents function to support both directory and single file layouts for loading sub-agent configurations.
- Added new langgraph_dev module for managing async sub-agent lifecycle and deployment.
- Created graphs for async sub-agents (writing-agent, data-analysis-agent) and updated langgraph.json for deployment.
- Introduced new sub-agent definitions for planner, research, debug, code, and writing agents with appropriate system prompts and configurations.
- Enhanced package data inclusion in pyproject.toml to accommodate new sub-agent YAML files.

* Refactor code for improved readability by consolidating conditional statements and formatting

* feat: enhance async sub-agent support with workspace synchronization and user feedback

- Added console status messages during async sub-agent server startup and workspace synchronization to improve user experience.
- Implemented a new WorkspaceSyncWidget for live feedback during workspace sync operations.
- Updated onboarding to reject occupied ports and ensure proper workspace handling for async sub-agents.
- Introduced locking mechanisms to manage concurrent access to langgraph dev processes and workspace states.

* feat: add async sub-agent configuration and server management functions

* feat: improve port occupation handling and log file management in start_langgraph_dev

* feat: enhance async sub-agent handling and introduce comprehensive tests

- Updated `_maybe_swap_async_subagents` to improve async sub-agent management, ensuring internal flags are stripped before handoff.
- Enhanced port management in `onboard.py` to allow reuse of occupied ports if already running by the same service.
- Introduced file locking in `manager.py` to prevent race conditions during concurrent CLI invocations.
- Added new tests for async sub-agent swapping and langgraph manager functionalities to ensure reliability and correctness.
- Updated dependencies in `pyproject.toml` to include `psutil` and `filelock`.

* fix(docs): clarify sub-agent configuration in README

* test(manager): isolate _PID_DIR + tighten reuse-path assertion

Addresses CodeRabbit review on tests/test_langgraph_manager.py:

- Patch _PID_DIR to tmp_path so the FileLock setup in
  ensure_langgraph_dev doesn't mkdir the user's real
  ~/.config/evoscientist/ dir as a test side-effect.
- Tighten "result is None or hasattr(result, 'poll')" to a strict
  "result is None" — the reuse path returns None unconditionally,
  so the OR clause was hiding potential regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(manager): clean up stale PID file when unrelated process reuses PID

* feat(tests): add validation tests for async flag in load_subagents

* fix(load_subagents): restrict to .yaml files and clarify configuration handling

* fix(load_subagents): improve error handling for non-dict specifications in YAML

* feat(onboard): add "LangGraph Port" step to onboarding process

* feat(langgraph): add concurrency configuration for langgraph dev workers

* feat(async-subagents): enhance MCP tool routing for async sub-agents

* fix(manager): update exception handling for connection errors and prevent zombie processes

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 12:45:50 +01:00
dinos 73928a2d78 fix(onboard): detect installed skill packs via install manifest (#199)
* fix(onboard): detect installed skill packs via install manifest

Onboarding's _step_skills only inspected USER_SKILLS_DIR and matched
recommended entries by directory-name hint, so a pack like
EvoScientist/EvoSkills@skills (which explodes into paper-writing/,
evo-memory/, etc. under GLOBAL_SKILLS_DIR) was never detected and kept
appearing as not-yet-installed.

skills_manager now writes a per-tier .installed.yaml mapping skill
directory name -> original install source on every install, removes the
entry on uninstall, and exposes installed_sources(). _step_skills checks
both tiers and treats a recommended source as installed when present in
any manifest -- so packs are recognized regardless of how their child
dirs are named.

* fix(onboard): write install manifest atomically

Stage to a sibling temp file, fsync, then os.replace into place. A crash
mid-write can no longer leave a half-written .installed.yaml behind,
which would otherwise wipe out pack detection until the next reinstall.

* style: fmt

* fix(onboard): catch decode errors when loading install manifest

read_text() can raise UnicodeDecodeError on a hand-edited or corrupt .installed.yaml; pin encoding="utf-8" and add UnicodeError to the except clause so the function honors its "returns {} on any error" contract.
2026-05-01 13:23:59 +02:00
Xi Zhang 50719ef256 Implement PruningCheckpointer for efficient checkpoint management and… (#194)
* Implement PruningCheckpointer for efficient checkpoint management and add comprehensive tests

- Introduced `PruningCheckpointer` to manage checkpoint pruning after each `aput()`, ensuring only the latest checkpoints are retained based on a configurable limit.
- Added migration sweep functionality to clean up legacy checkpoints and prevent database bloat.
- Enhanced `get_checkpointer()` to utilize the new `PruningCheckpointer` and trigger migration sweeps when necessary.
- Developed a suite of integration tests for `PruningCheckpointer`, covering various scenarios including pruning behavior, concurrent writes, and retention policies.
- Implemented tests for migration sweep functionality, ensuring proper partitioning and user version management.
- Added diagnostic helper `db_stats` to provide insights into the database state, including thread and checkpoint counts.

* feat(sessions): enhance pruning logic to handle legacy DBs without writes table

* fix(tests): prevent atexit hook leakage in TestMigrationSweep

* feat(tests): enhance TestPruningCheckpointer to validate put+prune serialization

* feat(tests): refactor mock path implementation for get_db_path in test cases
2026-04-28 22:26:59 +02:00
Xi Zhang 558360b558 feat: Enhance ModelPickerWidget for Ollama integration (#187)
* feat: Enhance ModelPickerWidget for Ollama integration

- Implemented a sentinel row for "Custom Ollama model..." in ModelPickerWidget, allowing users to input arbitrary model names.
- Updated action handling in ModelPickerWidget to manage transitions between list and input modes.
- Added async model discovery for Ollama models, integrating with the /model command to fetch locally installed models.
- Created tests for Ollama model discovery and ModelPickerWidget behavior, ensuring proper functionality and user experience.
- Refactored validate_ollama_connection and discover_ollama_models for improved error handling and response management.

* fix: Simplify code by removing unnecessary line breaks in ModelPickerWidget and test cases

* fix: Restore globals on set_chat_model failure to prevent half-switched session

* fix: Improve error handling in ModelCommand by restoring globals on failure
2026-04-25 00:37:48 +01:00
dinos 05f54334ba fix(mcp): stdio env passthrough + durable package installs (#169)
* fix(mcp): forward proxy and CA bundle env vars to stdio subprocesses

The MCP SDK's stdio transport inherits only a minimal allowlist (HOME,
PATH, USER, …) from the parent, stripping http_proxy/https_proxy and
SSL_CERT_FILE/REQUESTS_CA_BUNDLE/etc. Behind a proxy or with a custom CA
bundle, stdio MCP servers silently hang on outbound requests while the
same server over HTTP transport works. Auto-forward the proxy and cert
vars when present; user-configured env still takes precedence.

* fix(mcp): use `uv tool install` so MCP packages survive uv sync

Source installs previously used `uv pip install --python $VENV <pkg>`,
which lands in the evosci venv but is not recorded in pyproject.toml or
uv.lock. A subsequent `uv sync` (typical after `git pull`) reconciles
the venv to the lockfile and removes the MCP package, forcing users to
re-run onboard.

Prefer `uv tool install <pkg>` for the non-uv-tool install path: the
binary symlink in ~/.local/bin survives uv sync and evosci upgrades,
and the MCP server gets its own isolated env (no dep conflicts).
Verify the expected CLI entry point resolves afterward; if not (package
has no console-script), fall through to the old uv-pip path so
command-less packages still work.

The uv-tool-env path (`uv tool install evoscientist --with <pkg>`) is
unchanged — it was already durable via uv's receipt.

* fix(mcp): gate standalone uv tool install on verify_command

Previously `install_pip_package` would route every install through
`uv tool install <pkg>` when `verify_command` was None, returning
success as long as the uv subprocess exited 0. Library callers
(`evoscientist[oauth]`, `lark-oapi`, etc.) expect the package to land
in the active venv so they can import it — a standalone uv tool env
is not importable, so the import fails at the next line.

Gate the `uv tool install <pkg>` branch on `verify_command` being
set: that signals the caller wants a durable CLI binary, which is
what `uv tool install` produces. Library callers omit it and go
straight to the pip-install-into-venv path.

Also: log info messages on every fall-through so stale-binary and
entry-point-missing failure modes are debuggable, and document the
--with → standalone recovery path.

* fix(mcp): resolve MCP binaries to `uv tool dir --bin`, not `.venv/bin`

Under `uv run`, the project venv's `bin/` comes first on PATH, so
`shutil.which("arxiv-mcp-server")` returns a stale `.venv/bin/` copy
left over from an earlier install instead of the fresh symlink that
`uv tool install` just placed in `~/.local/bin`. The venv copy gets
written to mcp.yaml and is then wiped by the next `uv sync` — exactly
the failure mode the durability fix was meant to prevent.

Query `uv tool dir --bin` directly and prefer binaries found there
over `shutil.which`. Same change to the post-install verify in
`install_pip_package` so a venv shadow can't falsely short-circuit
the fallback.

* refactor(mcp): split install_pip_package into install_library + install_cli_tool

`verify_command` was doing double duty: naming the CLI binary to check
*and* signaling "this is a CLI install, use the standalone `uv tool
install` path." Callers routed library installs through the CLI branch
any time they forgot to pass it, and the resulting standalone uv tool
env wasn't importable from the active venv.

Separate the two use cases into distinct functions, each with one
install strategy per environment shape. Shared logic lives in private
`_install_with_uv_tool_env` / `_install_via_pip` helpers.

- install_library(pkg): uv-tool-env --with → pip. Never uses standalone
  `uv tool install <pkg>` (not importable from active venv).
- install_cli_tool(pkg, *, verify_command): uv-tool-env --with →
  standalone `uv tool install` → pip. `verify_command` is now required.

Callers pick the right function at the call site: registry.py picks
based on whether `entry.command` is set; onboard.py call sites all
install libraries.
2026-04-21 16:59:06 +01:00