* fix: defer eager observation-index build to first model call
* fix: add mtime-keyed cache to list_observation_documents to avoid re-parsing unchanged files
* fix: return a copy of the cached document list and strengthen the deletion test
* fix: copy cached document list on read and write to prevent caller mutations
* fix: split global and project cache to avoid duplicate parsing and cross-project invalidation
* fix: bump mtime explicitly in cache modification test for Windows NTFS resolution
* fix: bound project observation cache with LRU eviction
* fix: deduplicate path logic and strengthen cache typing
* fix: group cache tests under TestObservationCache with autouse fixture and fix f-string interpolation
* fix: reject non-positive observation cache cap in config validation
* fix: cache resolved observation docs and config cap to avoid repeated work
* fix: use st_mtime_ns and st_size in cache signature for NTFS reliability
* test: clear EVOSCIENTIST_MAX_CACHED_PROJECTS in test env cleanup fixtures
* test: cover per-file observation cache semantics
* fix: replace layered observation caches with per-file parse cache
* fix: serialize observation parse cache transactions
* docs: add subscription OAuth recipe (Claude + ChatGPT/Codex)
Standalone guide for running EvoScientist on Claude Pro/Max and ChatGPT
Plus/Pro subscriptions via ccproxy OAuth, previously only partially
covered inside the macOS deployment recipe.
Documents the two Codex-route pitfalls from #323 with their exact error
strings — ccproxy's default model mappings silently rewriting gpt-* to
gpt-5.3-codex, and the backend's client-identity gate ('requires a newer
version of Codex') — plus the manual ccproxy TOML fix for self-managed
instances, account-tier model availability, verification probes, and a
troubleshooting table. Adds the recipe to the docs index.
* docs: correct subscription OAuth guidance
* docs: remove stale Codex client version guidance
* docs: show the config section to add instead of a clobbering heredoc
* docs: correct reasoning_effort default behavior on the Codex route
* docs: drop private helper import from the Codex probe
* docs: warn that local ccproxy config files shadow the global one
---------
Co-authored-by: Dinos Papakostas <dinospk1999@gmail.com>
* Add Novita as an LLM provider
Registers Novita (novita.ai) as an OpenAI-routed provider, following the
same pattern as Requesty/Atlas Cloud/SiliconFlow: a base_url + API key env
var entry in _OPENAI_ROUTED_PROVIDERS, a handful of model registry entries
(DeepSeek/Qwen/GLM), onboarding wizard support (constants/steps/wizard/
helpers), a key validator using the auth-preflight sentinel pattern (Novita's
/v1/models endpoint returns the public catalog even for an invalid key, so
auth must be checked via a chat completion instead), and a host-to-provider
mapping entry for error attribution.
* Recommend Novita's current flagship models
The models listed for Novita were older ids that no longer reflect what
the platform leads with. Point the recommendations at the three current
flagships instead, each verified against api.novita.ai:
moonshotai/kimi-k3 1M context, native vision
zai-org/glm-5.2 1M context, long-horizon agentic work
deepseek/deepseek-v4-flash-0731 1M context, cheapest of the three
Context windows, output limits, input modalities and pricing were taken
from the live /openai/v1/models response rather than carried over.
* Keep branch CI workflow files unchanged (no workflow OAuth scope)
Co-authored-by: multica-agent <github@multica.ai>
* ci: restore workflow files to match main
---------
Co-authored-by: jax-novita <jax-novita@users.noreply.github.com>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Dinos Papakostas <dinospk1999@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Add a "BioNeMo Skills" entry to the onboard Step 7 checkbox so users can
opt into the 31 life-science skills (protein folding, docking, generative
chemistry, genomics, protein design) with one selection.
The pack installs through the existing GitHub shorthand path — no installer
change. Source points at the toolkit's plugin skills directory
(plugins/bionemo-agent-toolkit/skills), which is the flat aggregate of all
31 skills; installing from the repo root would only reach 14 because the
remaining skills sit three levels deep.
Closes#333
* fix: prevent session-emptying crash on resume commands
* docs: fix stale docstrings in goto=None crash tests and patch
- Correct checkpoint corruption claim: the error state replaces the
previous conversation state (messages: [], files: {}), it IS corrupted.
- Replace WebUI-specific language with UI-agnostic wording.
- Remove references to uncommitted local notes files.
- Add upstream issue reference (langchain-ai/langgraph#5656).
* fix: wrap _control_branch instead of reimplementing, use dataclasses.replace, add END routing and session preservation tests
* fix: rebind map_cmd on already-loaded consumers, rewrite crash-path tests to exercise __start__ via checkpoint deletion
Closes#392 (uncontroversial part).
WeChat (`_handle_message`) and Feishu (`_handle_event`) gated their
signature/decryption checks behind a condition the REQUEST controls:
- WeChat: `if encrypt and self._crypto:` -- a POST with no `<Encrypt>`
element took the false branch and reached `_safe_process_message`
without any verification, even when `encoding_aes_key` + `token` were
configured.
- Feishu: `if self.config.encrypt_key and "encrypt" in body:` -- a
plaintext body skipped decryption entirely and was processed directly.
Since the webhook port is the channel's only inbound boundary, an
attacker could POST forged plaintext and reach the agent, spoofing
`sender_id` / `FromUserName` (and, with an empty allowlist, passing the
sender gate).
Fix: when encryption is configured, an inbound POST MUST carry the
encrypted field (`<Encrypt>` / `encrypt`) -- otherwise it is rejected
with 403 and never reaches the agent. Plaintext mode (no encryption
configured) is unchanged, so existing plaintext deployments are not
affected. The remaining fail-closed question (what to do when
credentials are entirely unset) is left for the maintainers to decide
as the policy part of the issue.
Regression tests (9 new):
- WeChat: plaintext rejected / missing Encrypt rejected / bad signature
rejected / valid signature decrypts and processes / plaintext still
accepted when no crypto.
- Feishu: plaintext rejected / non-dict body rejected / encrypted body
decrypts and processes / plaintext still accepted when no encrypt_key.
93 tests in the two channel files pass; full suite 3045 passed, 13
skipped; ruff clean.
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix(mcp): give stdio subprocess a real stderr fd under redirected streams (#418)
On Windows the Textual TUI redirects sys.stderr to an in-memory capture
(textual.app._PrintCapture) whose fileno() returns -1. The MCP SDK forwards
that stderr to stdio server subprocesses via subprocess.Popen(stderr=...),
and Popen rejects the invalid handle with OSError: [Errno 9] Bad file
descriptor — so only stdio servers fail to load (HTTP/SSE are unaffected).
Wrap mcp.client.stdio.stdio_client so that, whenever the configured errlog
has no usable fileno, it falls back to sys.__stderr__ (or os.devnull in GUI
hosts). Idempotent, no-op when the SDK is absent, warns if the SDK renames
stdio_client. Adds 9 regression tests and a troubleshooting note.
* fix(mcp): validate live fd and close fallback errlog after stdio session
Address CodeRabbit review on #423:
- _stdio_errlog_is_usable now os.fstat()s the fd to reject closed streams
that still report their former positive fileno (prevents a deferred
[Errno 9] from subprocess.Popen).
- The stdio_client wrapper owns the devnull fallback it allocates and
closes it once the session exits, so repeated MCP reloads no longer leak
file descriptors. Caller-provided usable errlogs pass through untouched.
- Tests cover the closed-fd case, the fd-leak/closure invariant, and
confirm langchain-mcp-adapters binds the patched stdio_client.
* fix(mcp): rebind adapter stdio_client, forward errlog by kw, harden tests
Address CodeRabbit round-2 review on #423:
- The patch now also rebinds langchain_mcp_adapters.sessions.stdio_client,
which the adapter captures via a 'from' import at module load — so the
wrapped function reaches the adapter regardless of import order.
- errlog is forwarded to the SDK by keyword (original(server, *args,
errlog=errlog, **kwargs)) so a future SDK inserting a positional
parameter before errlog can't mis-bind the fallback.
- The fallback stream is now allocated inside the async context manager,
so it is closed on session exit even if the CM is constructed but never
entered (narrower fd-leak path).
- test_closed_fd_rejected now reaches the os.fstat branch (stale positive
fd stub) instead of the ValueError path; test_adapter_binds_patched_stdio_client
documents and asserts the import-order-independent rebind.
* fix(mcp): close fallback errlog when stdio_client construction fails
Address CodeRabbit round-3 review on #423: move the original(server, *args,
errlog=errlog, **kwargs) construction inside the try block so a failure
during subprocess/client setup still reaches the finally and closes the
wrapper-owned os.devnull stream. Added test_fallback_closed_when_construction_fails
covering the path.
* refactor(mcp): track fallback ownership via (stream, opened_by_us)
Address din0s review on #423:
- _safe_stdio_errlog() now returns (stream, opened_by_us); the wrapper closes
the fallback only when opened_by_us is True, instead of inferring ownership
from needs_fallback + an identity check against sys.__stderr__. Simpler and
less likely to regress.
- Removed dead try/finally in test_closed_fd_rejected.
- Added test_wrapped_stdio_client_swaps_explicit_bad_errlog covering the
'not _stdio_errlog_is_usable(errlog)' branch (explicit bad errlog, not the
default sentinel).
- Updated test_safe_errlog_returns_usable_stream for the tuple return.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat: add thread metadata index for improved performance in thread listing
- Implemented a new SQLite index on the `checkpoints` table to optimize thread listing queries by indexing relevant metadata fields.
- Updated the `list_threads` function to ensure the index is created if it does not exist.
- Added a test to verify the creation of the metadata index during thread listing.
feat: enhance workspace sidecar management with owner tracking
- Modified the workspace sidecar to include `owner_pids` to track the current process owners.
- Updated tests to validate the new owner tracking functionality and ensure proper behavior when managing workspace sidecars.
chore: introduce model registry for streamlined model management
- Created a new `registry.py` file to maintain a comprehensive model registry, including model names, IDs, providers, and routing tables.
- Added functions to retrieve models by provider and list available models, enhancing the modularity and maintainability of model management.
* feat: enhance workspace sidecar management and improve thread metadata indexing
* fix(tests): ensure sidecar correctly registers owner with original workspace and pid
* refactor: simplify workspace sidecar management by removing owner tracking
* feat(server): add commands to manage background langgraph dev server
- Introduced `server_app` for managing the langgraph dev server with commands to check status and stop the server.
- Enhanced workspace sidecar management to include configuration fingerprint for drift detection.
- Updated deployment functions to handle server configuration and state more effectively.
* feat(server): enhance server status command to display PID with stale record warning
* feat(langgraph_dev): exclusion-set config fingerprint, webui keepalive, unified stop guidance
* fix(cli): platform-specific manual-stop hint; document keepalive endpoint-change limitation
* ci: publish to PyPI via trusted publishing; build version images on release
* ci: pin publish actions to commit SHAs; extend version guard to docker and manual dispatch
* ci: disable setup-uv cache in the publish workflow
* fix: change default bind host to loopback for security across all components
* fix: update documentation and tests for loopback host configuration and security warnings
* feat: detect expert skills by AGENTS.md presence
* feat: let the orchestrator choose the expert dispatch tool
* refactor: drop per-skill expert dispatch classification
* chore: remove duplicate TestSkillManager test classes left by rebase
* fix: clear legacy actor fields when AGENTS.md declares the expert
* refactor: move the expert prompt into ActiveTeamMiddleware
* fix(deps): pin openrouter below 0.11 to avoid SSE stream regressions (#373)
* fix(openrouter): address SSE stream leak by closing response iterator
* refactor(openrouter): pass through SDK args in SSE leak patch
* test(openrouter): make SSE leak tests version-agnostic across SDK generations
* fix(openrouter): remove SSE stream leak patch and update dependencies
* fix: repair interrupted tool call history (#366)
* fix: repair interrupted tool call history
Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths.
* fix: repair malformed tool calls and dedupe repair warnings
Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted
threads with syntactically invalid tool calls get synthesized error
results and are accepted by strict providers.
Preserve the originating tool call's name in the synthesized ToolMessage,
and deduplicate repair warnings per unique tool-call id via a warned set
owned by the middleware instance, since the middleware rewrites the
request but not thread state.
Document the middleware's scope versus deepagents' PatchToolCallsMiddleware
(orphan ToolMessage dropping and mid-run coverage).
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* Update README.md
* Update README.md
* Update README.zh-CN.md
* fix: scrub host path from skill_manager output and guard batch install (#377)
* fix: scrub host path from skill_manager output and guard batch install
* test: tighten install leak guards to catch host path in either tier
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: log missing async-subagent tools at DEBUG, not WARNING (#378)
* fix: log missing async-subagent tools at DEBUG, not WARNING
* fix: distinguish load_subagents callers via async_swap_pending flag
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: default reasoning context for codex proxy Responses API (#380)
* fix: set langgraph and codex proxy runtime defaults
* fix: address runtime default review feedback
* fix: drop langgraph dev env defaults per maintainer review
langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment,
so the reported KeyError cannot come from this flow; the env defaults added
here were unnecessary. Scope the PR back to the codex proxy reasoning
context fix only.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* release: v0.2.4 (#389)
* fix: add support for new Anthropic models and enhance adaptive thinking tests
* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models
* fix: update version to v0.2.4 in badges, README, and project files
* fix: update Star History chart links in README and README.zh-CN
* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog
* fix: update wechat group image in assets
* refactor(runtime): centralize async bridges under an owned runtime (#376)
* feat(runtime): add application-scoped async runtime
* refactor(cli): use owned runtime for session stats
* refactor(onboard): use the owned async runtime
* docs(runtime): record async bridge ownership
* refactor(middleware): keep sync fallback synchronous
* refactor(mcp): load tools on an owned runtime
* refactor(cli): share owned runtime across entry points
* refactor(channels): make inbound sync bridge explicit
* refactor(stream): run Rich streaming on owned runtime
* chore(runtime): remove nest-asyncio dependency
* refactor(asyncio): require active loops in async code
* docs(runtime): document final event loop ownership
* fix(stream): cancel stalled owned streams
* fix(cli): recover cleanly from stream cancellation
* fix(runtime): drain executor work before shutdown
* fix(runtime): terminate cancelled shell process trees
* fix(models): let fallback bypass selector failures
* fix(cli): reset interrupt handling between turns
* docs: rm implementation spec
* fix(serve): cancel active turns during shutdown
* fix(runtime): protect settlement from waiter cancellation
* fix(backends): reject empty shell commands
* fix(runtime): terminate descendants after shell exit
* fix(mcp): keep standalone discovery off channel loop
* fix(cli): own and settle interactive prompt cancellation
* fix(serve): keep channel sends off runtime loop
* fix(stream): scope cancel context to iterator steps
* refactor(serve): require the owned async runtime
* fix(channels): keep interactive sends off runtime loop
* fix(selector): surface fallback without log spam
* test(runtime): normalize Windows shell marker
* fix(cli): serialize interactive session turns
* fix(shell): bound output drain after termination
* fix(ui): do not retry owned runtime failures
* fix(shell): allow signal-safe registry reentry
* fix(shell): avoid terminating reused process ids
* fix(channels): preserve streaming send order
* fix(cli): report runtime shutdown timeouts cleanly
* fix(mcp): guide async callers to async loader
* docs(runtime): clarify reserved async bridge APIs
* fix(runtime): bound code interpreter cleanup
* test(shell): use active Python for drain regression
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat: add payload-aware EvoAsyncSubAgentMiddleware
* feat: register expert_container_async graph for async expert dispatch
* feat: fold installed expert skills into async subagent registry
* fix: accept 'async' as valid default_dispatch value
* feat: dispatch-aware ActiveTeamMiddleware cue (task vs start_async_task)
* fix: drop future annotations in expert_async_subagent so ToolRuntime injects
* feat: surface output_path and skill_name to expert container as runtime cue
* fix: extend AsyncWatcher client cache with expert specs for completion nudge
* feat: teach main agent the async-expert return envelope shape
* feat: propagate cfg.model to expert-async runs.create via ClientCacheProxy
* chore: guard AsyncWatcher client-cache extension against upstream rename
* test: cover output_path runtime-context tail block and wrong-type guard
* docs: drop out-of-repo notes/ ref from expert_container_async module doc
* fix: warn on unrecognized default_dispatch frontmatter value
* fix: reject empty-body expert skills on async dispatch to match sync policy
* docs: explain why expert container includes general-purpose subagent
* fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385)
* fix: propagate langgraph dev bind port into subprocess env for self-loop URL
* fix: keep parent env authoritative over workspace .env for mapped keys
* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins
* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS
* fix: merge .env via dotenv_values to close empty-value and RMW-race edges
* chore: align docstrings after .env-merge rework
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* Add Requesty as an LLM provider (#346)
* Add Requesty as an LLM provider
* Address review: Requesty prompt caching, model ordering, key validation
- Declare Anthropic-style prompt caching for Requesty Claude models by
default (mirroring the OpenRouter behavior), with an opt-out flag
EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
overrides native/OpenRouter models for names it shares with them
(the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
invalid/missing key (public catalog), so it cannot validate a key. Use a
minimal authenticated /v1/chat/completions request instead (200 = valid,
403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).
* Validate Requesty key against auth layer, not a specific model
The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.
---------
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* fix(llm): filter unnamed tool calls (#390)
* fix(llm): filter unnamed tool calls
* test(llm): cover tool call sanitization branches
* fix(llm): repair unnamed tool calls in middleware
---------
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix(middleware): mount tool-history repair on sync subagents and harden raw tool-call vetting (#393)
* feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation
* fix(tests): add test for dropping non-list raw tool calls in repair_tool_history
* Add Atlas Cloud LLM provider (#388)
* Add Atlas Cloud LLM provider
* Add Atlas Cloud onboarding support
* fix(validators): update atlascloud key validation to handle insufficient balance case
---------
Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* fix: prepend EvoAsyncSubAgentMiddleware for prefix cache stability
* docs: clarify list_dispatchable_experts covers both dispatch shapes
* fix: guard async expert fold-in against reserved-name collisions
* fix: honest advertising surfaces for async expert dispatch
* fix: compose expert persona into base-stack system_message
* fix: drop payload from start_async_task, inject skill_name by construction
* feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395
- Introduced TodoListMiddleware to the middleware stack for better task management.
- Updated HITL interrupt configuration to include 'delete' operations requiring approval.
- Implemented error handling for delete operations in read-only and memory backends.
- Enhanced approval prompt formatting to display file paths for delete actions.
- Added tests to ensure delete operations are correctly blocked or prompted for approval.
- Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality.
* revert: drop skill_manager from sync expert-container tool_registry
* revert: drop skill_manager from async expert-container tools
* fix: drop removed ASYNC_TASK_SYSTEM_PROMPT import for deepagents 0.7.0
* fix: mock list_dispatchable_experts in single-cue test for CI
* chore: drop stale output_path from async container graph docstring
* fix: forward configurable_extra through owned-runtime and HITL re-invocations
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: Sanjay Santhanam <51058514+Sanjays2402@users.noreply.github.com>
Co-authored-by: Yougang Lyu <82445958+youganglyu@users.noreply.github.com>
Co-authored-by: houren Antony <2212222@mail.nankai.edu.cn>
Co-authored-by: dinos <dinospk1999@gmail.com>
Co-authored-by: Thibault Jaigu <84420566+Thibaultjaigu@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
Co-authored-by: nightcityblade <jackchen@haloailabs.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nb213 <binyangzhu000@gmail.com>
Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
* feat: add expert-skill schema and type filter to skill_manager
* feat: fold installed expert skills into main-agent subagent registry
* feat: add GET /api/teams listing expert skills for gallery
* chore: cache SKILL.md body on SkillInfo, cleanup expert-container comments
* fix: register skill_manager in expert-subagent tool_registry
* fix: catch UnicodeDecodeError in expert-skill body loader
* fix: guard expert subagent registration against name collisions
* fix: skip expert registration when SKILL.md body is empty
* fix: drop redundant str() guards on expert-skill frontmatter
* fix: harden SKILL.md parsing on expert-registration hot path
* feat: document sync / QuickJS panel / async dispatch modes in DELEGATION_STRATEGY
* feat: surface QuickJS panel dispatches in TUI
* fix: finalize running panel dispatches on turn cancellation
* fix: replace panel widget cancel path with public finalize API
* fix: drop dead isinstance guard on panel dispatch duration_ms
* fix: re-arm panel widget when a new dispatch arrives after finalize
* fix: stop panel timer when finalize_running finds zero running rows
* fix: register panel widget in cleanup dict before awaiting mount
* feat: configurable bind host for WebUI and langgraph dev (refs #400)
WebUI mode was only reachable from the machine running it: the front-end
got no bind interface, and `start_langgraph_dev(...)` was called without a
host, so both servers stayed on loopback with no way to widen them.
Adds two config fields with deliberately different defaults:
webui_host = 0.0.0.0 front-end serves the app shell, no secrets
langgraph_dev_host = 127.0.0.1 unauthenticated API, agent can run shell
The design hinges on separating bind address from client address. Only
bind() uses the configured interface; every consumer that *connects*
(health probes, occupancy checks, async sub-agent self-dispatch) goes
through the new `_probe_host`, which maps a wildcard bind back to
loopback and honors a pinned interface verbatim. `_can_bind_port` is the
one exception and binds the literal host, since it must replicate the
bind the server itself will attempt.
- manager.py: `_probe_host`, `_is_loopback_host`, `_format_hostport`;
host kwarg threaded through the probes and `start_langgraph_dev`,
which now emits `--host` and propagates
EVOSCIENTIST_LANGGRAPH_DEV_HOST to the subprocess
- sdk.py: `langgraph_dev_url` tracks host as well as port;
EvoScientist.py reuses it instead of an inline f-string
- server.py: `--host` flag mirroring `--port`, plus a red PUBLIC BIND
banner whenever the bind is not provably loopback
- webui.py: forwards both hosts; the front-end is widened via HOSTNAME
because @evoscientist/webui ships no --host flag — its bin launcher
does `HOSTNAME: process.env.HOSTNAME || "127.0.0.1"`. The warning is
gated on the backend host only, so the shipped front-end default
doesn't print a banner on every launch
Verified end to end against a live server: requesting 0.0.0.0 yields a
socket listening on 0.0.0.0 with the health probe correctly resolved to
127.0.0.1, while the default still binds 127.0.0.1 only.
Note: webui_host defaulting to 0.0.0.0 is a behavior change — upgrading
users will find the front-end reachable from the LAN.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat: default both bind hosts to 0.0.0.0, add --host and wizard host rendering (closes#400)
Completes the remaining items from #400.
- `langgraph_dev_host` now defaults to 0.0.0.0, matching `webui_host`.
Remote WebUI use needs both anyway (the UI reaches the backend from the
browser, not server-side), so a loopback backend default just meant every
remote user hit a silently failing UI. `_DEFAULT_HOST` and sdk's
`DEFAULT_LANGGRAPH_DEV_HOST` follow, so there is one story about where
these servers listen.
SECURITY: this exposes an unauthenticated API whose agent can run shell
commands. The red PUBLIC BIND banner consequently fires on every launch
while exposed — kept deliberately, since the exposure is real and the
escape hatch (`--host 127.0.0.1` / `config set langgraph_dev_host`) is
only discoverable if we say so. READMEs now lead with the warning and
document the SSH-tunnel alternative.
- `EvoSci --host <ip>` on the WebUI launch path, driving both servers. In
WebUI mode they are two halves of one surface; moving only one leaves the
UI loading but unable to reach the agent. Blank values are dropped rather
than written as an empty override that would beat the config file.
- Onboarding wizard no longer prints hard-coded `http://127.0.0.1:{port}` /
`http://localhost:{port}` (steps.py:160, :223) — both render the
configured bind through `_base_url` / `_format_hostport`, so a pinned
interface is reported honestly and a wildcard still shows loopback.
Verified against a live server: with no host argument at all, resolution
through EvoScientistConfig yields a socket listening on 0.0.0.0, a client URL
of http://127.0.0.1, and the warning gate returning True.
Still open and tracked separately: the front-end takes its backend URL from
browser input: `@evoscientist/webui` reads only HOSTNAME, PORT and
EVOSCIENTIST_LANGGRAPH_DEV_PORT, so advertising a backend URL needs a change
in that repo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* ci: bump setup-uv v6 -> v9.0.0 to drop the deprecated node20 runtime
GitHub now warns that setup-uv@v6 targets Node.js 20 and is being forced
onto Node.js 24. v7.0.0 is the release that made that switch, so anything
>= v7 clears the warning; v9.0.0 is current.
Pinned to the full tag deliberately: setup-uv stopped publishing major and
minor tags in v8.0.0 as supply-chain hardening, so `@v9` and `@v8` return
404 and would fail the job outright. Releases are immutable from v8 on, so
the full tag is as tamper-proof as a SHA. Comment left in lint.yml because
"simplifying" this back to `@v9` is an easy and CI-breaking mistake.
actions/checkout@v5 is already node24 and needs no change.
Note: v9.0.0 flips the `prune-cache` default to false (upstream did this to
ease load on PyPI infrastructure). None of these workflows set it, so they
follow the new default and Actions cache usage may grow.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(cli): correct --host help text and warn on public bind in non-WebUI modes
The --host help claimed "WebUI mode only", which is wrong in a way that
matters for security. `--host` writes `langgraph_dev_host` unconditionally,
and `_ensure_async_subagent_server` auto-starts that backend for tui / cli /
serve as well — the langgraph dev server is shared across UI modes. So the
flag narrows or widens the agent API in every mode, and only `webui_host` is
actually WebUI-specific. Reported against cli/commands.py.
The documentation error hid a real gap: the PUBLIC BIND banner lived only in
deploy/server.py and deploy/webui.py, so a plain `EvoSci` session bound
0.0.0.0 with no runtime signal whatsoever — and `--help` is opt-in, so
fixing the text alone would not surface it. Added the same banner to the
shared CLI path, gated on `is_async_subagents_available()`: ensure_langgraph_dev
fails soft (async degrades to in-process delegation), and warning about a
bind that never happened would be worse than staying quiet.
READMEs (EN + zh-CN) get the same correction — the warning block sat inside
the Desktop WebUI section and read as WebUI-scoped.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(deploy): strip the config-derived bind host, not just the CLI one
`deploy()` only stripped the `--host` branch. When the flag was omitted,
`getattr(config, "langgraph_dev_host", ...)` flowed unstripped into
`_is_port_occupied`, `is_langgraph_dev_running`, `start_langgraph_dev` and
the banner. `run_webui` already strips unconditionally; this aligns the two.
Reachable because `deploy()` reads through `getattr` and is routinely handed
duck-typed config objects (tests, embedders) that never run
`EvoScientistConfig.__post_init__`, which is what normally normalizes these
fields.
Worst case was not just a bad bind: `_is_loopback_host(" 127.0.0.1 ")` is
False, so a padded loopback value would print a false PUBLIC BIND warning
while binding a string socket.bind() rejects outright — a security banner
saying the opposite of the truth.
Three regression tests added, each verified to fail against the old code.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* style: apply ruff format to the bind-host changes
The Lint workflow runs both `ruff check` and `ruff format --check`; I had
only been running the former locally, so five files landed unformatted and
failed CI. Whitespace and line-wrapping only — no semantic change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(security): keep the langgraph dev backend on loopback by default
The backend is an unauthenticated API whose agent can run shell commands,
and it is auto-started in every UI mode (tui/cli/webui/serve/deploy) — so a
0.0.0.0 default put it on the network for users who never asked. Restore
127.0.0.1 as the default and make 0.0.0.0 an explicit opt-in.
webui_host keeps its 0.0.0.0 default: the front-end serves the app shell
only and holds no credentials. run_webui already prints a remote-backend
hint when the front-end is exposed and the backend is not.
Help text and both READMEs are reframed around widening rather than
narrowing; the escape-hatch tests are inverted to assert the public-bind
opt-in survives into argv.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(middleware): normalize blank tool_call_id before provider request
Closes#345.
Some streaming providers (notably Kimi and Zhipu on streamed tool calls)
occasionally emit tool calls whose id is an empty string, whitespace, or
None. Strict providers reject the next turn with `invalid tool_call_id`
(HTTP 400, code 3), freezing the thread after the very first tool round.
The existing repair logic skipped blank ids via a truthy check
(`if tool_call_id := call.get("id")`), so an AIMessage carrying a blank
id was passed through unchanged while its paired ToolMessage was dropped
as an orphan -- the next model call then 400'd.
This adds a pre-pass (`_normalize_blank_tool_call_ids`) before the main
repair loop that:
* assigns each blank-id tool call on an AIMessage a fresh `_repair_<uuid>` id
* pairs subsequent blank-id ToolMessages in arrival order (FIFO) so existing
exchanges stay paired
* lets unmatched blank calls fall through to the main loop, which synthesizes
an interrupted-result ToolMessage with the fresh id
* drops `additional_kwargs["tool_calls"]` on touched messages so
langchain-openai's serializer falls back to the now-valid parsed form
instead of preferring the raw form (which still carries the blank id)
* pre-adds the fresh ids to `warned` so repair stays silent on subsequent
model calls (the fresh ids are non-deterministic across calls)
Verified end-to-end: `_convert_message_to_dict` on the repaired history
puts no blank id on the wire and preserves AIMessage<->ToolMessage pairing.
* fix(middleware): scope blank-id FIFO per exchange; cover invalid_tool_calls
Addresses CodeRabbit review on PR #399.
(1) Critical -- pending_slots FIFO leak across exchanges
--------------------------------------------------------
The FIFO queue of fresh ids for blank-id calls was never closed at
non-ToolMessage boundaries, so a later exchange's blank ToolMessage
could be paired with a stale id from an earlier, already-interrupted
exchange. The main loop then dropped the real tool result as an orphan
and synthesized a fake interrupted result for the real call:
AIMessage1(blank A) -- interrupted
HumanMessage
AIMessage2(blank B)
ToolMessage(blank) -> popped A_new from FIFO front, not B_new
Fix: close pending_slots at every AIMessage / HumanMessage /
SystemMessage boundary via _close_unclaimed, mirroring close_pending()
in the main loop. Unclaimed slots are merged into `warned` so the
main loop's synthesized interrupted-result for that id stays silent
across model calls.
(2) Major -- invalid_tool_calls with blank id were silently dropped
-------------------------------------------------------------------
`any_changed` was set only inside the `tool_calls` loop, so a message
with a blank id only in `invalid_tool_calls` never entered the
model_copy update path and the blank id survived untouched.
langchain-openai's serializer puts `tool_calls + invalid_tool_calls`
on the wire when either parsed list is non-empty (it does NOT skip
invalid calls), so a blank id on an invalid call reaches the provider
just as readily as one on a valid call -- verified by direct
inspection of `_convert_message_to_dict`.
Fix: track `invalid_changed` separately and include it in the
update-path condition; also push invalid fresh ids into pending_slots
(valid-first ordering ensures a real ToolMessage for a valid blank
call never accidentally claims an invalid call's id).
Tests
-----
* test_pending_slots_scoped_per_exchange_not_global_fifo: the exact
cross-exchange leak scenario CodeRabbit described.
* test_normalizes_blank_id_in_invalid_tool_calls_only: the
invalid-only case that previously slipped through.
All 26 tests in test_tool_history_repair_middleware.py pass; full
suite 3044 passed, 13 skipped (Windows-compatible subset).
* fix(middleware): deterministic repair ids; don't pair invalid calls with orphan results
Addresses din0s review on PR #399.
(1) IDs are now deterministic, not uuid4
----------------------------------------
Each blank id is rewritten as `_repair_{msg_idx}_{v|i}{call_idx}` so the same
blank call gets the same id on every model call. The middleware re-runs on
every request but can only rewrite the outgoing request, not the thread
state, so random uuids made the wire payload unstable and grew the `warned`
set unboundedly. Deterministic ids let the main loop's existing `warned`-set
dedup suppress the synthesized-result warning from the second call on --
no pre-add hack needed. The `warned` parameter is therefore dropped from
`_normalize_blank_tool_call_ids` (and the `_close_unclaimed` helper removed
in favor of plain `pending_slots.clear()`).
(2) invalid_tool_calls fresh ids are no longer pushed to pending_slots
----------------------------------------------------------------------
Invalid calls are never executed by LangGraph (args can't be parsed), so no
real ToolMessage can claim their slot. Pushing it let an orphan blank
ToolMessage from some other call mis-pair with the invalid call, surfacing
the orphan's content under the invalid call's name. Without the push the
orphan keeps its blank id and the main loop drops it, which is what we want.
Tests
-----
* Updated `test_blank_id_repair_does_not_spam_warnings` to expect the
one-time-warning-then-silent pattern (matches
`test_warning_deduplicates_across_calls`) and asserts the deterministic id.
* New `test_invalid_blank_id_not_pushed_to_pending_slots` reproduces the
exact mis-pairing scenario (orphan blank result + invalid call) and
asserts the orphan content never leaks into a repaired tool result.
27 tests in test_tool_history_repair_middleware.py pass; ruff clean.
- Introduced TodoListMiddleware to the middleware stack for better task management.
- Updated HITL interrupt configuration to include 'delete' operations requiring approval.
- Implemented error handling for delete operations in read-only and memory backends.
- Enhanced approval prompt formatting to display file paths for delete actions.
- Added tests to ensure delete operations are correctly blocked or prompted for approval.
- Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality.
* feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation
* fix(tests): add test for dropping non-list raw tool calls in repair_tool_history
* Add Requesty as an LLM provider
* Address review: Requesty prompt caching, model ordering, key validation
- Declare Anthropic-style prompt caching for Requesty Claude models by
default (mirroring the OpenRouter behavior), with an opt-out flag
EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
overrides native/OpenRouter models for names it shares with them
(the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
invalid/missing key (public catalog), so it cannot validate a key. Use a
minimal authenticated /v1/chat/completions request instead (200 = valid,
403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).
* Validate Requesty key against auth layer, not a specific model
The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.
---------
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* fix: propagate langgraph dev bind port into subprocess env for self-loop URL
* fix: keep parent env authoritative over workspace .env for mapped keys
* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins
* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS
* fix: merge .env via dotenv_values to close empty-value and RMW-race edges
* chore: align docstrings after .env-merge rework
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: add support for new Anthropic models and enhance adaptive thinking tests
* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models
* fix: update version to v0.2.4 in badges, README, and project files
* fix: update Star History chart links in README and README.zh-CN
* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog
* fix: update wechat group image in assets
* fix: set langgraph and codex proxy runtime defaults
* fix: address runtime default review feedback
* fix: drop langgraph dev env defaults per maintainer review
langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment,
so the reported KeyError cannot come from this flow; the env defaults added
here were unnecessary. Scope the PR back to the codex proxy reasoning
context fix only.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: log missing async-subagent tools at DEBUG, not WARNING
* fix: distinguish load_subagents callers via async_swap_pending flag
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: scrub host path from skill_manager output and guard batch install
* test: tighten install leak guards to catch host path in either tier
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: repair interrupted tool call history
Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths.
* fix: repair malformed tool calls and dedupe repair warnings
Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted
threads with syntactically invalid tool calls get synthesized error
results and are accepted by strict providers.
Preserve the originating tool call's name in the synthesized ToolMessage,
and deduplicate repair warnings per unique tool-call id via a warned set
owned by the middleware instance, since the middleware rewrites the
request but not thread state.
Document the middleware's scope versus deepagents' PatchToolCallsMiddleware
(orphan ToolMessage dropping and mid-run coverage).
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat(context-window): add Kimi K3 model with 1M context window
* feat(openrouter): implement structured output for Kimi K3 and add 429 retry handling
* Refactor code structure for improved readability and maintainability