Commit Graph

646 Commits

Author SHA1 Message Date
houren Antony da15b70535 fix(channels): reject unsigned webhook POSTs on encryption-configured channels (#401)
Closes #392 (uncontroversial part).

WeChat (`_handle_message`) and Feishu (`_handle_event`) gated their
signature/decryption checks behind a condition the REQUEST controls:

- WeChat: `if encrypt and self._crypto:` -- a POST with no `<Encrypt>`
  element took the false branch and reached `_safe_process_message`
  without any verification, even when `encoding_aes_key` + `token` were
  configured.
- Feishu: `if self.config.encrypt_key and "encrypt" in body:` -- a
  plaintext body skipped decryption entirely and was processed directly.

Since the webhook port is the channel's only inbound boundary, an
attacker could POST forged plaintext and reach the agent, spoofing
`sender_id` / `FromUserName` (and, with an empty allowlist, passing the
sender gate).

Fix: when encryption is configured, an inbound POST MUST carry the
encrypted field (`<Encrypt>` / `encrypt`) -- otherwise it is rejected
with 403 and never reaches the agent. Plaintext mode (no encryption
configured) is unchanged, so existing plaintext deployments are not
affected. The remaining fail-closed question (what to do when
credentials are entirely unset) is left for the maintainers to decide
as the policy part of the issue.

Regression tests (9 new):
- WeChat: plaintext rejected / missing Encrypt rejected / bad signature
  rejected / valid signature decrypts and processes / plaintext still
  accepted when no crypto.
- Feishu: plaintext rejected / non-dict body rejected / encrypted body
  decrypts and processes / plaintext still accepted when no encrypt_key.

93 tests in the two channel files pass; full suite 3045 passed, 13
skipped; ruff clean.

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-08-14 17:25:57 +01:00
houren Antony 932c934485 fix(mcp): give stdio subprocess a real stderr fd under redirected streams (#423)
* fix(mcp): give stdio subprocess a real stderr fd under redirected streams (#418)

On Windows the Textual TUI redirects sys.stderr to an in-memory capture
(textual.app._PrintCapture) whose fileno() returns -1. The MCP SDK forwards
that stderr to stdio server subprocesses via subprocess.Popen(stderr=...),
and Popen rejects the invalid handle with OSError: [Errno 9] Bad file
descriptor — so only stdio servers fail to load (HTTP/SSE are unaffected).

Wrap mcp.client.stdio.stdio_client so that, whenever the configured errlog
has no usable fileno, it falls back to sys.__stderr__ (or os.devnull in GUI
hosts). Idempotent, no-op when the SDK is absent, warns if the SDK renames
stdio_client. Adds 9 regression tests and a troubleshooting note.

* fix(mcp): validate live fd and close fallback errlog after stdio session

Address CodeRabbit review on #423:
- _stdio_errlog_is_usable now os.fstat()s the fd to reject closed streams
  that still report their former positive fileno (prevents a deferred
  [Errno 9] from subprocess.Popen).
- The stdio_client wrapper owns the devnull fallback it allocates and
  closes it once the session exits, so repeated MCP reloads no longer leak
  file descriptors. Caller-provided usable errlogs pass through untouched.
- Tests cover the closed-fd case, the fd-leak/closure invariant, and
  confirm langchain-mcp-adapters binds the patched stdio_client.

* fix(mcp): rebind adapter stdio_client, forward errlog by kw, harden tests

Address CodeRabbit round-2 review on #423:
- The patch now also rebinds langchain_mcp_adapters.sessions.stdio_client,
  which the adapter captures via a 'from' import at module load — so the
  wrapped function reaches the adapter regardless of import order.
- errlog is forwarded to the SDK by keyword (original(server, *args,
  errlog=errlog, **kwargs)) so a future SDK inserting a positional
  parameter before errlog can't mis-bind the fallback.
- The fallback stream is now allocated inside the async context manager,
  so it is closed on session exit even if the CM is constructed but never
  entered (narrower fd-leak path).
- test_closed_fd_rejected now reaches the os.fstat branch (stale positive
  fd stub) instead of the ValueError path; test_adapter_binds_patched_stdio_client
  documents and asserts the import-order-independent rebind.

* fix(mcp): close fallback errlog when stdio_client construction fails

Address CodeRabbit round-3 review on #423: move the original(server, *args,
errlog=errlog, **kwargs) construction inside the try block so a failure
during subprocess/client setup still reaches the finally and closes the
wrapper-owned os.devnull stream. Added test_fallback_closed_when_construction_fails
covering the path.

* refactor(mcp): track fallback ownership via (stream, opened_by_us)

Address din0s review on #423:
- _safe_stdio_errlog() now returns (stream, opened_by_us); the wrapper closes
  the fallback only when opened_by_us is True, instead of inferring ownership
  from needs_fallback + an identity check against sys.__stderr__. Simpler and
  less likely to regress.
- Removed dead try/finally in test_closed_fd_rejected.
- Added test_wrapped_stdio_client_swaps_explicit_bad_errlog covering the
  'not _stdio_errlog_is_usable(errlog)' branch (explicit bad errlog, not the
  default sentinel).
- Updated test_safe_errlog_returns_usable_stream for the tuple return.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-08-13 17:39:45 +00:00
ADITYA 1f3e8f57a3 fix: resolve subagent tools at execution decision (#381)
* fix: defer subagent tool resolution

Signed-off-by: Aditya Datta <crazyme07071996@gmail.com>

* fix: preserve inherited subagent tools

* test: cover injected subagent tools

* style: format subagent regression

---------

Signed-off-by: Aditya Datta <crazyme07071996@gmail.com>
2026-08-11 14:13:04 +01:00
Xi Zhang fb329e4aaa perf: cut startup latency — lazy import surface, langgraph dev keepalive, indexed thread listing (#407)
* feat: add thread metadata index for improved performance in thread listing

- Implemented a new SQLite index on the `checkpoints` table to optimize thread listing queries by indexing relevant metadata fields.
- Updated the `list_threads` function to ensure the index is created if it does not exist.
- Added a test to verify the creation of the metadata index during thread listing.

feat: enhance workspace sidecar management with owner tracking

- Modified the workspace sidecar to include `owner_pids` to track the current process owners.
- Updated tests to validate the new owner tracking functionality and ensure proper behavior when managing workspace sidecars.

chore: introduce model registry for streamlined model management

- Created a new `registry.py` file to maintain a comprehensive model registry, including model names, IDs, providers, and routing tables.
- Added functions to retrieve models by provider and list available models, enhancing the modularity and maintainability of model management.

* feat: enhance workspace sidecar management and improve thread metadata indexing

* fix(tests): ensure sidecar correctly registers owner with original workspace and pid

* refactor: simplify workspace sidecar management by removing owner tracking

* feat(server): add commands to manage background langgraph dev server

- Introduced `server_app` for managing the langgraph dev server with commands to check status and stop the server.
- Enhanced workspace sidecar management to include configuration fingerprint for drift detection.
- Updated deployment functions to handle server configuration and state more effectively.

* feat(server): enhance server status command to display PID with stale record warning

* feat(langgraph_dev): exclusion-set config fingerprint, webui keepalive, unified stop guidance

* fix(cli): platform-specific manual-stop hint; document keepalive endpoint-change limitation
2026-08-11 08:37:58 +01:00
dependabot[bot] 3f45ebd6fa chore(deps): bump h2 in the uv group across 1 directory (#421)
Bumps the uv group with 1 update in the / directory: [h2](https://github.com/python-hyper/h2).


Updates `h2` from 4.3.0 to 4.4.1
- [Changelog](https://github.com/python-hyper/h2/blob/master/CHANGELOG.rst)
- [Commits](https://github.com/python-hyper/h2/compare/v4.3.0...v4.4.1)

---
updated-dependencies:
- dependency-name: h2
  dependency-version: 4.4.1
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-10 10:58:57 +01:00
dependabot[bot] a1be0d1437 chore(deps): bump langgraph-checkpoint-sqlite (#416)
Bumps the uv group with 1 update in the / directory: [langgraph-checkpoint-sqlite](https://github.com/langchain-ai/langgraph).


Updates `langgraph-checkpoint-sqlite` from 3.1.0 to 3.1.1
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/checkpointsqlite==3.1.0...checkpointsqlite==3.1.1)

---
updated-dependencies:
- dependency-name: langgraph-checkpoint-sqlite
  dependency-version: 3.1.1
  dependency-type: direct:production
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-08-10 08:13:14 +00:00
Xi Zhang e086f76da7 ci: publish to PyPI via trusted publishing; build version images on release (#417)
* ci: publish to PyPI via trusted publishing; build version images on release

* ci: pin publish actions to commit SHAs; extend version guard to docker and manual dispatch

* ci: disable setup-uv cache in the publish workflow
2026-08-09 17:31:17 +01:00
Xi Zhang 12adc62868 release: v0.2.6
* chore: update version to v0.2.6 in badges, README, and configuration files

* feat: add Qwen 3.8-Max model support and update changelog for v0.2.6
2026-08-07 18:27:19 +01:00
Xi Zhang 0c21a01f6f fix: default WebUI bind host back to loopback (#412)
* fix: change default bind host to loopback for security across all components

* fix: update documentation and tests for loopback host configuration and security warnings
2026-08-07 17:13:57 +01:00
X-iZhang b40b6f784d feat: clarify comment in NewCommand to align with explicit expert clearing 2026-08-07 17:07:01 +01:00
X-iZhang aff63ccd65 feat: enhance expert invitation handling with case-insensitive matching and session management 2026-08-07 17:07:01 +01:00
jfilipiuk ab1a6b0062 feat: adopt the AGENTS.md expert-skill contract (#404)
* feat: detect expert skills by AGENTS.md presence

* feat: let the orchestrator choose the expert dispatch tool

* refactor: drop per-skill expert dispatch classification

* chore: remove duplicate TestSkillManager test classes left by rebase

* fix: clear legacy actor fields when AGENTS.md declares the expert

* refactor: move the expert prompt into ActiveTeamMiddleware
2026-08-07 17:07:01 +01:00
jfilipiuk bd2464423a feat: agent-teams part D - async expert dispatch mechanism (#391)
* fix(deps): pin openrouter below 0.11 to avoid SSE stream regressions (#373)

* fix(openrouter): address SSE stream leak by closing response iterator

* refactor(openrouter): pass through SDK args in SSE leak patch

* test(openrouter): make SSE leak tests version-agnostic across SDK generations

* fix(openrouter): remove SSE stream leak patch and update dependencies

* fix: repair interrupted tool call history (#366)

* fix: repair interrupted tool call history

Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths.

* fix: repair malformed tool calls and dedupe repair warnings

Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted
threads with syntactically invalid tool calls get synthesized error
results and are accepted by strict providers.

Preserve the originating tool call's name in the synthesized ToolMessage,
and deduplicate repair warnings per unique tool-call id via a warned set
owned by the middleware instance, since the middleware rewrites the
request but not thread state.

Document the middleware's scope versus deepagents' PatchToolCallsMiddleware
(orphan ToolMessage dropping and mid-run coverage).

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Update README.md

* Update README.md

* Update README.zh-CN.md

* fix: scrub host path from skill_manager output and guard batch install (#377)

* fix: scrub host path from skill_manager output and guard batch install

* test: tighten install leak guards to catch host path in either tier

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: log missing async-subagent tools at DEBUG, not WARNING (#378)

* fix: log missing async-subagent tools at DEBUG, not WARNING

* fix: distinguish load_subagents callers via async_swap_pending flag

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: default reasoning context for codex proxy Responses API (#380)

* fix: set langgraph and codex proxy runtime defaults

* fix: address runtime default review feedback

* fix: drop langgraph dev env defaults per maintainer review

langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment,
so the reported KeyError cannot come from this flow; the env defaults added
here were unnecessary. Scope the PR back to the codex proxy reasoning
context fix only.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* release: v0.2.4 (#389)

* fix: add support for new Anthropic models and enhance adaptive thinking tests

* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models

* fix: update version to v0.2.4 in badges, README, and project files

* fix: update Star History chart links in README and README.zh-CN

* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog

* fix: update wechat group image in assets

* refactor(runtime): centralize async bridges under an owned runtime (#376)

* feat(runtime): add application-scoped async runtime

* refactor(cli): use owned runtime for session stats

* refactor(onboard): use the owned async runtime

* docs(runtime): record async bridge ownership

* refactor(middleware): keep sync fallback synchronous

* refactor(mcp): load tools on an owned runtime

* refactor(cli): share owned runtime across entry points

* refactor(channels): make inbound sync bridge explicit

* refactor(stream): run Rich streaming on owned runtime

* chore(runtime): remove nest-asyncio dependency

* refactor(asyncio): require active loops in async code

* docs(runtime): document final event loop ownership

* fix(stream): cancel stalled owned streams

* fix(cli): recover cleanly from stream cancellation

* fix(runtime): drain executor work before shutdown

* fix(runtime): terminate cancelled shell process trees

* fix(models): let fallback bypass selector failures

* fix(cli): reset interrupt handling between turns

* docs: rm implementation spec

* fix(serve): cancel active turns during shutdown

* fix(runtime): protect settlement from waiter cancellation

* fix(backends): reject empty shell commands

* fix(runtime): terminate descendants after shell exit

* fix(mcp): keep standalone discovery off channel loop

* fix(cli): own and settle interactive prompt cancellation

* fix(serve): keep channel sends off runtime loop

* fix(stream): scope cancel context to iterator steps

* refactor(serve): require the owned async runtime

* fix(channels): keep interactive sends off runtime loop

* fix(selector): surface fallback without log spam

* test(runtime): normalize Windows shell marker

* fix(cli): serialize interactive session turns

* fix(shell): bound output drain after termination

* fix(ui): do not retry owned runtime failures

* fix(shell): allow signal-safe registry reentry

* fix(shell): avoid terminating reused process ids

* fix(channels): preserve streaming send order

* fix(cli): report runtime shutdown timeouts cleanly

* fix(mcp): guide async callers to async loader

* docs(runtime): clarify reserved async bridge APIs

* fix(runtime): bound code interpreter cleanup

* test(shell): use active Python for drain regression

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* feat: add payload-aware EvoAsyncSubAgentMiddleware

* feat: register expert_container_async graph for async expert dispatch

* feat: fold installed expert skills into async subagent registry

* fix: accept 'async' as valid default_dispatch value

* feat: dispatch-aware ActiveTeamMiddleware cue (task vs start_async_task)

* fix: drop future annotations in expert_async_subagent so ToolRuntime injects

* feat: surface output_path and skill_name to expert container as runtime cue

* fix: extend AsyncWatcher client cache with expert specs for completion nudge

* feat: teach main agent the async-expert return envelope shape

* feat: propagate cfg.model to expert-async runs.create via ClientCacheProxy

* chore: guard AsyncWatcher client-cache extension against upstream rename

* test: cover output_path runtime-context tail block and wrong-type guard

* docs: drop out-of-repo notes/ ref from expert_container_async module doc

* fix: warn on unrecognized default_dispatch frontmatter value

* fix: reject empty-body expert skills on async dispatch to match sync policy

* docs: explain why expert container includes general-purpose subagent

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385)

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL

* fix: keep parent env authoritative over workspace .env for mapped keys

* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins

* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS

* fix: merge .env via dotenv_values to close empty-value and RMW-race edges

* chore: align docstrings after .env-merge rework

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Add Requesty as an LLM provider (#346)

* Add Requesty as an LLM provider

* Address review: Requesty prompt caching, model ordering, key validation

- Declare Anthropic-style prompt caching for Requesty Claude models by
  default (mirroring the OpenRouter behavior), with an opt-out flag
  EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
  provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
  overrides native/OpenRouter models for names it shares with them
  (the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
  invalid/missing key (public catalog), so it cannot validate a key. Use a
  minimal authenticated /v1/chat/completions request instead (200 = valid,
  403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).

* Validate Requesty key against auth layer, not a specific model

The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix(llm): filter unnamed tool calls (#390)

* fix(llm): filter unnamed tool calls

* test(llm): cover tool call sanitization branches

* fix(llm): repair unnamed tool calls in middleware

---------

Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix(middleware): mount tool-history repair on sync subagents and harden raw tool-call vetting (#393)

* feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation

* fix(tests): add test for dropping non-list raw tool calls in repair_tool_history

* Add Atlas Cloud LLM provider (#388)

* Add Atlas Cloud LLM provider

* Add Atlas Cloud onboarding support

* fix(validators): update atlascloud key validation to handle insufficient balance case

---------

Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix: prepend EvoAsyncSubAgentMiddleware for prefix cache stability

* docs: clarify list_dispatchable_experts covers both dispatch shapes

* fix: guard async expert fold-in against reserved-name collisions

* fix: honest advertising surfaces for async expert dispatch

* fix: compose expert persona into base-stack system_message

* fix: drop payload from start_async_task, inject skill_name by construction

* feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395

- Introduced TodoListMiddleware to the middleware stack for better task management.
- Updated HITL interrupt configuration to include 'delete' operations requiring approval.
- Implemented error handling for delete operations in read-only and memory backends.
- Enhanced approval prompt formatting to display file paths for delete actions.
- Added tests to ensure delete operations are correctly blocked or prompted for approval.
- Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality.

* revert: drop skill_manager from sync expert-container tool_registry

* revert: drop skill_manager from async expert-container tools

* fix: drop removed ASYNC_TASK_SYSTEM_PROMPT import for deepagents 0.7.0

* fix: mock list_dispatchable_experts in single-cue test for CI

* chore: drop stale output_path from async container graph docstring

* fix: forward configurable_extra through owned-runtime and HITL re-invocations

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: Sanjay Santhanam <51058514+Sanjays2402@users.noreply.github.com>
Co-authored-by: Yougang Lyu <82445958+youganglyu@users.noreply.github.com>
Co-authored-by: houren Antony <2212222@mail.nankai.edu.cn>
Co-authored-by: dinos <dinospk1999@gmail.com>
Co-authored-by: Thibault Jaigu <84420566+Thibaultjaigu@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
Co-authored-by: nightcityblade <jackchen@haloailabs.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nb213 <binyangzhu000@gmail.com>
Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
2026-08-07 17:07:01 +01:00
jfilipiuk 3cda9894c7 feat: agent-teams part C - expert selection UX (depends on part B) (#371)
* feat: bias main-agent delegation toward configurable.active_teams

* feat: add /experts and /expert TUI commands for expert-skill summoning

* feat: align expert-selection wording with WebUI (invite/dismiss)

* chore: clear active_teams on /new, cleanup active_team.py comment

* fix: use local append_to_system_message in ActiveTeamMiddleware

* fix: prevent configurable_extra from overriding thread_id

* fix: cache expert-skill lookup for /expert completions

* fix: suppress /expert completions past first arg and on exact match

* fix: invalidate /expert completion cache on skill install/uninstall

* fix: refuse /expert invites for non-dispatchable expert skills

* fix: fire /expert cache invalidation on every install_skill / uninstall_skill path

* fix: propagate active_teams to Rich CLI and serve dispatch surfaces

* fix: keep invited experts across channel shutdown

* fix: match /expert completions case-insensitively
2026-08-07 17:07:01 +01:00
jfilipiuk 3a1dbf0a0f feat: agent-teams part B - expert-skill backend mechanism (#370)
* feat: add expert-skill schema and type filter to skill_manager

* feat: fold installed expert skills into main-agent subagent registry

* feat: add GET /api/teams listing expert skills for gallery

* chore: cache SKILL.md body on SkillInfo, cleanup expert-container comments

* fix: register skill_manager in expert-subagent tool_registry

* fix: catch UnicodeDecodeError in expert-skill body loader

* fix: guard expert subagent registration against name collisions

* fix: skip expert registration when SKILL.md body is empty

* fix: drop redundant str() guards on expert-skill frontmatter

* fix: harden SKILL.md parsing on expert-registration hot path
2026-08-07 17:07:01 +01:00
jfilipiuk b5b01d50c2 feat: agent-teams part A - TUI panel visibility + DELEGATION_STRATEGY routing (#369)
* feat: document sync / QuickJS panel / async dispatch modes in DELEGATION_STRATEGY

* feat: surface QuickJS panel dispatches in TUI

* fix: finalize running panel dispatches on turn cancellation

* fix: replace panel widget cancel path with public finalize API

* fix: drop dead isinstance guard on panel dispatch duration_ms

* fix: re-arm panel widget when a new dispatch arrives after finalize

* fix: stop panel timer when finalize_running finds zero running rows

* fix: register panel widget in cleanup dict before awaiting mount
2026-08-07 17:07:01 +01:00
Ziheng Zhang 33979e5371 fix(llm): support Volcengine Coding model aliases (#411)
* fix(llm): support Volcengine Coding model aliases

* refactor(llm): add Volcengine Coding provider

* style: format Volcengine Coding test
2026-08-07 14:30:23 +08:00
Xiaohui Yan 3c5cc831c0 Feat/configurable bind host (#402)
* feat: configurable bind host for WebUI and langgraph dev (refs #400)

WebUI mode was only reachable from the machine running it: the front-end
got no bind interface, and `start_langgraph_dev(...)` was called without a
host, so both servers stayed on loopback with no way to widen them.

Adds two config fields with deliberately different defaults:

  webui_host        = 0.0.0.0    front-end serves the app shell, no secrets
  langgraph_dev_host = 127.0.0.1  unauthenticated API, agent can run shell

The design hinges on separating bind address from client address. Only
bind() uses the configured interface; every consumer that *connects*
(health probes, occupancy checks, async sub-agent self-dispatch) goes
through the new `_probe_host`, which maps a wildcard bind back to
loopback and honors a pinned interface verbatim. `_can_bind_port` is the
one exception and binds the literal host, since it must replicate the
bind the server itself will attempt.

  - manager.py: `_probe_host`, `_is_loopback_host`, `_format_hostport`;
    host kwarg threaded through the probes and `start_langgraph_dev`,
    which now emits `--host` and propagates
    EVOSCIENTIST_LANGGRAPH_DEV_HOST to the subprocess
  - sdk.py: `langgraph_dev_url` tracks host as well as port;
    EvoScientist.py reuses it instead of an inline f-string
  - server.py: `--host` flag mirroring `--port`, plus a red PUBLIC BIND
    banner whenever the bind is not provably loopback
  - webui.py: forwards both hosts; the front-end is widened via HOSTNAME
    because @evoscientist/webui ships no --host flag — its bin launcher
    does `HOSTNAME: process.env.HOSTNAME || "127.0.0.1"`. The warning is
    gated on the backend host only, so the shipped front-end default
    doesn't print a banner on every launch

Verified end to end against a live server: requesting 0.0.0.0 yields a
socket listening on 0.0.0.0 with the health probe correctly resolved to
127.0.0.1, while the default still binds 127.0.0.1 only.

Note: webui_host defaulting to 0.0.0.0 is a behavior change — upgrading
users will find the front-end reachable from the LAN.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: default both bind hosts to 0.0.0.0, add --host and wizard host rendering (closes #400)

Completes the remaining items from #400.

  - `langgraph_dev_host` now defaults to 0.0.0.0, matching `webui_host`.
    Remote WebUI use needs both anyway (the UI reaches the backend from the
    browser, not server-side), so a loopback backend default just meant every
    remote user hit a silently failing UI. `_DEFAULT_HOST` and sdk's
    `DEFAULT_LANGGRAPH_DEV_HOST` follow, so there is one story about where
    these servers listen.

    SECURITY: this exposes an unauthenticated API whose agent can run shell
    commands. The red PUBLIC BIND banner consequently fires on every launch
    while exposed — kept deliberately, since the exposure is real and the
    escape hatch (`--host 127.0.0.1` / `config set langgraph_dev_host`) is
    only discoverable if we say so. READMEs now lead with the warning and
    document the SSH-tunnel alternative.

  - `EvoSci --host <ip>` on the WebUI launch path, driving both servers. In
    WebUI mode they are two halves of one surface; moving only one leaves the
    UI loading but unable to reach the agent. Blank values are dropped rather
    than written as an empty override that would beat the config file.

  - Onboarding wizard no longer prints hard-coded `http://127.0.0.1:{port}` /
    `http://localhost:{port}` (steps.py:160, :223) — both render the
    configured bind through `_base_url` / `_format_hostport`, so a pinned
    interface is reported honestly and a wildcard still shows loopback.

Verified against a live server: with no host argument at all, resolution
through EvoScientistConfig yields a socket listening on 0.0.0.0, a client URL
of http://127.0.0.1, and the warning gate returning True.

Still open and tracked separately: the front-end takes its backend URL from
browser input: `@evoscientist/webui` reads only HOSTNAME, PORT and
EVOSCIENTIST_LANGGRAPH_DEV_PORT, so advertising a backend URL needs a change
in that repo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: bump setup-uv v6 -> v9.0.0 to drop the deprecated node20 runtime

GitHub now warns that setup-uv@v6 targets Node.js 20 and is being forced
onto Node.js 24. v7.0.0 is the release that made that switch, so anything
>= v7 clears the warning; v9.0.0 is current.

Pinned to the full tag deliberately: setup-uv stopped publishing major and
minor tags in v8.0.0 as supply-chain hardening, so `@v9` and `@v8` return
404 and would fail the job outright. Releases are immutable from v8 on, so
the full tag is as tamper-proof as a SHA. Comment left in lint.yml because
"simplifying" this back to `@v9` is an easy and CI-breaking mistake.

actions/checkout@v5 is already node24 and needs no change.

Note: v9.0.0 flips the `prune-cache` default to false (upstream did this to
ease load on PyPI infrastructure). None of these workflows set it, so they
follow the new default and Actions cache usage may grow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(cli): correct --host help text and warn on public bind in non-WebUI modes

The --host help claimed "WebUI mode only", which is wrong in a way that
matters for security. `--host` writes `langgraph_dev_host` unconditionally,
and `_ensure_async_subagent_server` auto-starts that backend for tui / cli /
serve as well — the langgraph dev server is shared across UI modes. So the
flag narrows or widens the agent API in every mode, and only `webui_host` is
actually WebUI-specific. Reported against cli/commands.py.

The documentation error hid a real gap: the PUBLIC BIND banner lived only in
deploy/server.py and deploy/webui.py, so a plain `EvoSci` session bound
0.0.0.0 with no runtime signal whatsoever — and `--help` is opt-in, so
fixing the text alone would not surface it. Added the same banner to the
shared CLI path, gated on `is_async_subagents_available()`: ensure_langgraph_dev
fails soft (async degrades to in-process delegation), and warning about a
bind that never happened would be worse than staying quiet.

READMEs (EN + zh-CN) get the same correction — the warning block sat inside
the Desktop WebUI section and read as WebUI-scoped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(deploy): strip the config-derived bind host, not just the CLI one

`deploy()` only stripped the `--host` branch. When the flag was omitted,
`getattr(config, "langgraph_dev_host", ...)` flowed unstripped into
`_is_port_occupied`, `is_langgraph_dev_running`, `start_langgraph_dev` and
the banner. `run_webui` already strips unconditionally; this aligns the two.

Reachable because `deploy()` reads through `getattr` and is routinely handed
duck-typed config objects (tests, embedders) that never run
`EvoScientistConfig.__post_init__`, which is what normally normalizes these
fields.

Worst case was not just a bad bind: `_is_loopback_host(" 127.0.0.1 ")` is
False, so a padded loopback value would print a false PUBLIC BIND warning
while binding a string socket.bind() rejects outright — a security banner
saying the opposite of the truth.

Three regression tests added, each verified to fail against the old code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* style: apply ruff format to the bind-host changes

The Lint workflow runs both `ruff check` and `ruff format --check`; I had
only been running the former locally, so five files landed unformatted and
failed CI. Whitespace and line-wrapping only — no semantic change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(security): keep the langgraph dev backend on loopback by default

The backend is an unauthenticated API whose agent can run shell commands,
and it is auto-started in every UI mode (tui/cli/webui/serve/deploy) — so a
0.0.0.0 default put it on the network for users who never asked. Restore
127.0.0.1 as the default and make 0.0.0.0 an explicit opt-in.

webui_host keeps its 0.0.0.0 default: the front-end serves the app shell
only and holds no credentials. run_webui already prints a remote-backend
hint when the front-end is exposed and the backend is not.

Help text and both READMEs are reframed around widening rather than
narrowing; the escape-hatch tests are inverted to assert the public-bind
opt-in survives into argv.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 16:15:49 +08:00
dependabot[bot] a9a57cd828 chore(deps): bump cryptography in the uv group across 1 directory (#406)
Bumps the uv group with 1 update in the / directory: [cryptography](https://github.com/pyca/cryptography).


Updates `cryptography` from 49.0.0 to 50.0.0
- [Changelog](https://github.com/pyca/cryptography/blob/main/CHANGELOG.rst)
- [Commits](https://github.com/pyca/cryptography/compare/49.0.0...50.0.0)

---
updated-dependencies:
- dependency-name: cryptography
  dependency-version: 50.0.0
  dependency-type: direct:production
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-05 15:30:02 +01:00
dependabot[bot] c7d93fb574 chore(deps): bump aiohttp in the uv group across 1 directory (#405)
Bumps the uv group with 1 update in the / directory: [aiohttp](https://github.com/aio-libs/aiohttp).


Updates `aiohttp` from 3.14.1 to 3.14.3
- [Changelog](https://github.com/aio-libs/aiohttp/blob/master/CHANGES.rst)
- [Commits](https://github.com/aio-libs/aiohttp/compare/v3.14.1...v3.14.3)

---
updated-dependencies:
- dependency-name: aiohttp
  dependency-version: 3.14.3
  dependency-type: direct:production
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-08-05 15:05:43 +01:00
houren Antony d97ee2b917 fix(middleware): normalize blank tool_call_id before provider request (#399)
* fix(middleware): normalize blank tool_call_id before provider request

Closes #345.

Some streaming providers (notably Kimi and Zhipu on streamed tool calls)
occasionally emit tool calls whose id is an empty string, whitespace, or
None. Strict providers reject the next turn with `invalid tool_call_id`
(HTTP 400, code 3), freezing the thread after the very first tool round.

The existing repair logic skipped blank ids via a truthy check
(`if tool_call_id := call.get("id")`), so an AIMessage carrying a blank
id was passed through unchanged while its paired ToolMessage was dropped
as an orphan -- the next model call then 400'd.

This adds a pre-pass (`_normalize_blank_tool_call_ids`) before the main
repair loop that:

* assigns each blank-id tool call on an AIMessage a fresh `_repair_<uuid>` id
* pairs subsequent blank-id ToolMessages in arrival order (FIFO) so existing
  exchanges stay paired
* lets unmatched blank calls fall through to the main loop, which synthesizes
  an interrupted-result ToolMessage with the fresh id
* drops `additional_kwargs["tool_calls"]` on touched messages so
  langchain-openai's serializer falls back to the now-valid parsed form
  instead of preferring the raw form (which still carries the blank id)
* pre-adds the fresh ids to `warned` so repair stays silent on subsequent
  model calls (the fresh ids are non-deterministic across calls)

Verified end-to-end: `_convert_message_to_dict` on the repaired history
puts no blank id on the wire and preserves AIMessage<->ToolMessage pairing.

* fix(middleware): scope blank-id FIFO per exchange; cover invalid_tool_calls

Addresses CodeRabbit review on PR #399.

(1) Critical -- pending_slots FIFO leak across exchanges
--------------------------------------------------------
The FIFO queue of fresh ids for blank-id calls was never closed at
non-ToolMessage boundaries, so a later exchange's blank ToolMessage
could be paired with a stale id from an earlier, already-interrupted
exchange. The main loop then dropped the real tool result as an orphan
and synthesized a fake interrupted result for the real call:

    AIMessage1(blank A) -- interrupted
    HumanMessage
    AIMessage2(blank B)
    ToolMessage(blank)  -> popped A_new from FIFO front, not B_new

Fix: close pending_slots at every AIMessage / HumanMessage /
SystemMessage boundary via _close_unclaimed, mirroring close_pending()
in the main loop. Unclaimed slots are merged into `warned` so the
main loop's synthesized interrupted-result for that id stays silent
across model calls.

(2) Major -- invalid_tool_calls with blank id were silently dropped
-------------------------------------------------------------------
`any_changed` was set only inside the `tool_calls` loop, so a message
with a blank id only in `invalid_tool_calls` never entered the
model_copy update path and the blank id survived untouched.

langchain-openai's serializer puts `tool_calls + invalid_tool_calls`
on the wire when either parsed list is non-empty (it does NOT skip
invalid calls), so a blank id on an invalid call reaches the provider
just as readily as one on a valid call -- verified by direct
inspection of `_convert_message_to_dict`.

Fix: track `invalid_changed` separately and include it in the
update-path condition; also push invalid fresh ids into pending_slots
(valid-first ordering ensures a real ToolMessage for a valid blank
call never accidentally claims an invalid call's id).

Tests
-----
* test_pending_slots_scoped_per_exchange_not_global_fifo: the exact
  cross-exchange leak scenario CodeRabbit described.
* test_normalizes_blank_id_in_invalid_tool_calls_only: the
  invalid-only case that previously slipped through.

All 26 tests in test_tool_history_repair_middleware.py pass; full
suite 3044 passed, 13 skipped (Windows-compatible subset).

* fix(middleware): deterministic repair ids; don't pair invalid calls with orphan results

Addresses din0s review on PR #399.

(1) IDs are now deterministic, not uuid4
----------------------------------------
Each blank id is rewritten as `_repair_{msg_idx}_{v|i}{call_idx}` so the same
blank call gets the same id on every model call. The middleware re-runs on
every request but can only rewrite the outgoing request, not the thread
state, so random uuids made the wire payload unstable and grew the `warned`
set unboundedly. Deterministic ids let the main loop's existing `warned`-set
dedup suppress the synthesized-result warning from the second call on --
no pre-add hack needed. The `warned` parameter is therefore dropped from
`_normalize_blank_tool_call_ids` (and the `_close_unclaimed` helper removed
in favor of plain `pending_slots.clear()`).

(2) invalid_tool_calls fresh ids are no longer pushed to pending_slots
----------------------------------------------------------------------
Invalid calls are never executed by LangGraph (args can't be parsed), so no
real ToolMessage can claim their slot. Pushing it let an orphan blank
ToolMessage from some other call mis-pair with the invalid call, surfacing
the orphan's content under the invalid call's name. Without the push the
orphan keeps its blank id and the main loop drops it, which is what we want.

Tests
-----
* Updated `test_blank_id_repair_does_not_spam_warnings` to expect the
  one-time-warning-then-silent pattern (matches
  `test_warning_deduplicates_across_calls`) and asserts the deterministic id.
* New `test_invalid_blank_id_not_pushed_to_pending_slots` reproduces the
  exact mis-pairing scenario (orphan blank result + invalid call) and
  asserts the orphan content never leaks into a repaired tool result.

27 tests in test_tool_history_repair_middleware.py pass; ruff clean.
2026-08-03 15:02:15 +01:00
Xi Zhang 2a17d14e33 chore: update version to v0.2.5 2026-08-01 17:33:12 +01:00
Xi Zhang d249e320bd feat: unified HITL approval for sync + async sub-agents (closes #387) (#396)
* feat: implement guard for dangerous commands and enhance HITL interrupt handling

* feat: enhance HITL approval mechanism and introduce session auto-approve decisions

* feat: add refuse_delete option to backend and enhance async delete guards

* feat: simplify delete method in CustomSandboxBackend and clarify adelete behavior

* feat: enhance allow-list behavior for command resolution and add related tests
2026-07-31 10:23:05 +01:00
Xi Zhang f81a8b086e feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395
- Introduced TodoListMiddleware to the middleware stack for better task management.
- Updated HITL interrupt configuration to include 'delete' operations requiring approval.
- Implemented error handling for delete operations in read-only and memory backends.
- Enhanced approval prompt formatting to display file paths for delete actions.
- Added tests to ensure delete operations are correctly blocked or prompted for approval.
- Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality.
2026-07-30 10:41:12 +01:00
nb213 4ddf7ebe52 Add Atlas Cloud LLM provider (#388)
* Add Atlas Cloud LLM provider

* Add Atlas Cloud onboarding support

* fix(validators): update atlascloud key validation to handle insufficient balance case

---------

Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-07-29 14:09:44 +01:00
Xi Zhang 562ce0eb83 fix(middleware): mount tool-history repair on sync subagents and harden raw tool-call vetting (#393)
* feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation

* fix(tests): add test for dropping non-list raw tool calls in repair_tool_history
2026-07-28 18:14:37 +01:00
nightcityblade a6a8e19dcc fix(llm): filter unnamed tool calls (#390)
* fix(llm): filter unnamed tool calls

* test(llm): cover tool call sanitization branches

* fix(llm): repair unnamed tool calls in middleware

---------

Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-28 17:37:32 +01:00
Thibault Jaigu 8cac50e2ef Add Requesty as an LLM provider (#346)
* Add Requesty as an LLM provider

* Address review: Requesty prompt caching, model ordering, key validation

- Declare Anthropic-style prompt caching for Requesty Claude models by
  default (mirroring the OpenRouter behavior), with an opt-out flag
  EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
  provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
  overrides native/OpenRouter models for names it shares with them
  (the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
  invalid/missing key (public catalog), so it cannot validate a key. Use a
  minimal authenticated /v1/chat/completions request instead (200 = valid,
  403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).

* Validate Requesty key against auth layer, not a specific model

The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-07-28 15:35:41 +01:00
jfilipiuk 10c032450e fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385)
* fix: propagate langgraph dev bind port into subprocess env for self-loop URL

* fix: keep parent env authoritative over workspace .env for mapped keys

* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins

* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS

* fix: merge .env via dotenv_values to close empty-value and RMW-race edges

* chore: align docstrings after .env-merge rework

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-28 12:26:22 +00:00
dinos 8b1451cdda refactor(runtime): centralize async bridges under an owned runtime (#376)
* feat(runtime): add application-scoped async runtime

* refactor(cli): use owned runtime for session stats

* refactor(onboard): use the owned async runtime

* docs(runtime): record async bridge ownership

* refactor(middleware): keep sync fallback synchronous

* refactor(mcp): load tools on an owned runtime

* refactor(cli): share owned runtime across entry points

* refactor(channels): make inbound sync bridge explicit

* refactor(stream): run Rich streaming on owned runtime

* chore(runtime): remove nest-asyncio dependency

* refactor(asyncio): require active loops in async code

* docs(runtime): document final event loop ownership

* fix(stream): cancel stalled owned streams

* fix(cli): recover cleanly from stream cancellation

* fix(runtime): drain executor work before shutdown

* fix(runtime): terminate cancelled shell process trees

* fix(models): let fallback bypass selector failures

* fix(cli): reset interrupt handling between turns

* docs: rm implementation spec

* fix(serve): cancel active turns during shutdown

* fix(runtime): protect settlement from waiter cancellation

* fix(backends): reject empty shell commands

* fix(runtime): terminate descendants after shell exit

* fix(mcp): keep standalone discovery off channel loop

* fix(cli): own and settle interactive prompt cancellation

* fix(serve): keep channel sends off runtime loop

* fix(stream): scope cancel context to iterator steps

* refactor(serve): require the owned async runtime

* fix(channels): keep interactive sends off runtime loop

* fix(selector): surface fallback without log spam

* test(runtime): normalize Windows shell marker

* fix(cli): serialize interactive session turns

* fix(shell): bound output drain after termination

* fix(ui): do not retry owned runtime failures

* fix(shell): allow signal-safe registry reentry

* fix(shell): avoid terminating reused process ids

* fix(channels): preserve streaming send order

* fix(cli): report runtime shutdown timeouts cleanly

* fix(mcp): guide async callers to async loader

* docs(runtime): clarify reserved async bridge APIs

* fix(runtime): bound code interpreter cleanup

* test(shell): use active Python for drain regression

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-27 14:17:57 +01:00
Xi Zhang ac58caab7b release: v0.2.4 (#389)
* fix: add support for new Anthropic models and enhance adaptive thinking tests

* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models

* fix: update version to v0.2.4 in badges, README, and project files

* fix: update Star History chart links in README and README.zh-CN

* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog

* fix: update wechat group image in assets
2026-07-26 14:58:48 +01:00
houren Antony fe70599d3d fix: default reasoning context for codex proxy Responses API (#380)
* fix: set langgraph and codex proxy runtime defaults

* fix: address runtime default review feedback

* fix: drop langgraph dev env defaults per maintainer review

langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment,
so the reported KeyError cannot come from this flow; the env defaults added
here were unnecessary. Scope the PR back to the codex proxy reasoning
context fix only.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-24 15:29:34 +00:00
jfilipiuk 1863f0730c fix: log missing async-subagent tools at DEBUG, not WARNING (#378)
* fix: log missing async-subagent tools at DEBUG, not WARNING

* fix: distinguish load_subagents callers via async_swap_pending flag

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-23 19:07:30 +00:00
jfilipiuk 172c8409d4 fix: scrub host path from skill_manager output and guard batch install (#377)
* fix: scrub host path from skill_manager output and guard batch install

* test: tighten install leak guards to catch host path in either tier

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-23 18:55:51 +00:00
Yougang Lyu cb4e127d8c Update README.zh-CN.md 2026-07-23 20:38:16 +02:00
Yougang Lyu 5bd3234973 Update README.md 2026-07-23 20:35:49 +02:00
Yougang Lyu 71e480d463 Update README.md 2026-07-23 20:35:11 +02:00
Sanjay Santhanam f802a49535 fix: repair interrupted tool call history (#366)
* fix: repair interrupted tool call history

Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths.

* fix: repair malformed tool calls and dedupe repair warnings

Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted
threads with syntactically invalid tool calls get synthesized error
results and are accepted by strict providers.

Preserve the originating tool call's name in the synthesized ToolMessage,
and deduplicate repair warnings per unique tool-call id via a warned set
owned by the middleware instance, since the middleware rewrites the
request but not thread state.

Document the middleware's scope versus deepagents' PatchToolCallsMiddleware
(orphan ToolMessage dropping and mid-run coverage).

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-22 15:43:58 +01:00
Xi Zhang e9857dde22 fix(deps): pin openrouter below 0.11 to avoid SSE stream regressions (#373)
* fix(openrouter): address SSE stream leak by closing response iterator

* refactor(openrouter): pass through SDK args in SSE leak patch

* test(openrouter): make SSE leak tests version-agnostic across SDK generations

* fix(openrouter): remove SSE stream leak patch and update dependencies
2026-07-22 16:03:57 +02:00
X-iZhang 1abfc8236d chore: update version to v0.2.3 2026-07-18 00:28:13 +01:00
Xi Zhang 042da63d54 feat(llm): add Kimi K3 support (#367)
* feat(context-window): add Kimi K3 model with 1M context window

* feat(openrouter): implement structured output for Kimi K3 and add 429 retry handling

* Refactor code structure for improved readability and maintainability
2026-07-18 00:12:08 +01:00
dinos 06a9511bdd fix(channels): telegram slash commands (#364) 2026-07-17 16:05:49 +01:00
jfilipiuk 584b9d24ac fix(tool-selector): cap chatter and streaming volume (#350)
* feat: add disable_streaming helper for tool-selector's internal model

* feat: apply disable_streaming to the tool-selector's model in the factory

* fix(tool-selector): hide selector model call from public event streams

* feat(tool-selector): log a WARNING when the selector's model returns a duplicate-tool_calls flood

* chore: log flood-detector errors, document parent-method drift risk, tighten tests

* refactor(tool-selector): switch to nostream tag via model-field wiring, drop subclass
2026-07-17 15:28:16 +01:00
dependabot[bot] 89bfb548b8 chore(deps): bump mcp in the uv group across 1 directory (#365)
---
updated-dependencies:
- dependency-name: mcp
  dependency-version: 1.28.1
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-16 22:24:53 +01:00
Xi Zhang e8399b7c94 fix(tui): window slash-command completions to terminal height (#362)
* feat(tui): implement completion popup rendering and windowing logic

* feat(tui): enhance completion popup with dynamic row budgeting and CSS adjustments

* Refactor picker widgets to use shared base class for improved code reuse

- Introduced `picker_base.py` to encapsulate common functionality for picker widgets.
- Updated `ModelPickerWidget`, `SkillBrowserWidget`, and `ThreadPickerWidget` to inherit from `PickerWidgetBase`.
- Implemented selection helpers (`first_selectable_index`, `move_selection`) in `picker_base.py` for consistent item navigation.
- Refactored rendering and selection logic in each widget to utilize the new base class methods.
- Added tests for picker functionality to ensure behavior remains consistent post-refactor.
2026-07-16 16:17:58 +01:00
Xi Zhang 0f709cff8b fix: cascade-cancel runs on thread deletion + startup orphan sweep (#358) (#359)
* feat: implement bulk cancellation of non-terminal runs before thread deletion

* test: enhance thread cancellation tests and add fake restore for orphaned runs sweep

* feat: enhance run cancellation logic to support status filtering during thread deletion

* feat: add langgraph-sdk dependency for enhanced functionality
2026-07-16 12:46:09 +01:00
dinos 05dfffbc73 fix(llm): use native langchain-deepseek SDK (#349) 2026-07-15 17:13:33 +01:00
dinos 01845f4311 refactor: route middleware display events through an injected event sink (#343)
* chore: add pytest-asyncio in auto mode

* test: migrate channel and stream tests to native async

Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.

* test: migrate command and model/middleware tests to native async

Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.

* test: migrate TUI, notifier, gateway, and session tests to native async

TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.

* test: replace direct asyncio.run() calls with native async tests

Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.

* test: drop undeclared anyio markers and delete run_async helper

The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).

* test: add autouse fixture for watcher cleanup

* refactor: remove redundant hasattr calls

* refactor: add typed middleware event sink and thread through assembly

Add MiddlewareEventSink protocol + NoOpSink in middleware/events.py
with a documented any-thread non-blocking contract (contract test uses a
deliberately-slow fake sink). Thread an optional `events` parameter
through create_cli_agent -> _get_default_middleware -> tool selector /
model fallback constructors; subagent stacks are always forced to
NoOpSink.

* refactor: inject a notifier port into async-watcher and background middleware

Add public pre_cancel_watcher() and enqueue_task_notification() to
cli/async_notifier.py and a small NotifierPort protocol
(middleware/notifier.py) that the module satisfies structurally.
AsyncWatcherMiddleware and BackgroundExecutionMiddleware now receive the
port by constructor injection at the composition root, deleting the lazy
'from ..cli import async_notifier' imports and the private
_watcher_by_thread / _enqueue pokes.

* refactor: invert tool-selection ownership onto a frontend event sink

The adaptive tool selector now reports on_tool_selection_started /
on_tool_selection / on_tool_selection_ended to the injected sink instead
of writing four process-global module variables. The frontend sink
(stream/sink.py FrontendEventSink) owns the selected/total/active state
with consume-once + dedup-vs-last-emitted semantics;
stream/tool_selection.py reads that sink object (a ToolSelectionView)
rather than reaching into tool_selector's globals.

Deleted: the 4 module globals, the cross-module mutations in
tool_selection.py, the track_stream_selection flag, the now-vestigial
_ToolSelectionTrackerMiddleware, reset_tool_selection_state_for_tests,
and the autouse conftest fixture. The sink is threaded from the two
interactive frontends through create_runtime_gateways ->
LocalGraphGateway (read side) and _load_agent -> create_cli_agent (write
side); subagent / headless stacks get NoOpSink.

* refactor: route model-fallback narration through the injected event sink

Delete the _ui_emit_fn / set_ui_emit module global and the
..stream.console import from model_fallback.py. The fallback middleware
now reports through its injected sink: the fallback transition via the
structured on_model_fallback (the frontend formats the '-> Falling back
to ...' line), and the surrounding narration (primary-failure header,
per-attempt outcome, exhaustion, non-fallbackable rejection) via
emit_fallback_notice, preserving the exact user-facing text. The TUI
binds its _append_system as the sink's fallback display where it used to
call set_ui_emit (cleared on exit); the Rich CLI's sink prints to the
console. _try_fallbacks / _guard_and_fallback take the sink.

* refactor: declare events on the GraphGateway protocol

Both gateway implementations now carry an explicit events attribute
(LangGraphServerGateway holds None — no frontend renders middleware
events across the HTTP boundary), so the four call sites use plain
attribute access instead of getattr probing an implicit contract.

* refactor: bind fallback display via the closure-scoped concrete sink

The App methods used gateway.events (typed as the read-side view) and
hasattr-probed for the concrete FrontendEventSink API. The enclosing
factory creates that sink two hundred lines up — close over it directly:
no probing, fully typed, and it becomes a constructor parameter
naturally when the App class is hoisted out of the factory.

* fix: end tool selection before fallback handler

* fix: keep fallback display errors non-fatal

* fix: preserve selector suppression for default streams

* fix: restore fallback notice console display

* refactor: consolidate fallback narration events

* refactor: clean middleware event sink plumbing

* fix: type gateway session events

* refactor: make all event protocols runtime-checkable

MiddlewareEventSink already carried @runtime_checkable (the stream
binding guard isinstance-checks it); ToolSelectionView and SessionEvents
now match, so mirroring that pattern against any of the three protocols
works instead of raising TypeError.

* fix(cli): close QuickJS workers after one-shot failures

* fix(cli): honor no-thinking in final output

* fix(channels): report failed startup accurately

* fix(channels): make Telegram cleanup idempotent

* fix(tui): skip command sync during exit

* fix(channels): preserve startup state during retries

* refactor(channels): share pending startup status

* refactor(cli): expose channel startup snapshot

* fix(tui): move channel startup off event loop

* test(channels): release retry gate on assertion failure

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-14 22:34:17 +00:00
dinos db1abce8d8 refactor: extract a shared HITL/ask_user interaction engine (#342)
* chore: add pytest-asyncio in auto mode

* test: migrate channel and stream tests to native async

Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.

* test: migrate command and model/middleware tests to native async

Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.

* test: migrate TUI, notifier, gateway, and session tests to native async

TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.

* test: replace direct asyncio.run() calls with native async tests

Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.

* test: drop undeclared anyio markers and delete run_async helper

The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).

* test: add autouse fixture for watcher cleanup

* refactor: remove redundant hasattr calls

* refactor: extract shared HITL/ask_user interaction grammar

Extract prompt/question formatting, the reply grammar (approval letters,
ask_user choice letters + the 'Other' sub-flow, stop-commands), the
ApprovalPolicy (config auto-approve rule + session registry +
session-key derivation), per-flow timeout constants, and the bilingual
feedback strings into channels/interaction.py. Both drivers now point at
the shared functions: this reverses cli/channel.py's imports of consumer
privates and closes the /stop drift at the parsing layer (serve-mode
ask_user now checks stop-commands before parsing an answer, matching the
CLI path).

* refactor: add interaction engine + registry; port InboundConsumer

Introduce InteractionIO (transport adapter Protocol),
PendingReplyRegistry (one asyncio-based reply router per process), and
the engine coroutines resolve_ask_user / resolve_approval in
channels/interaction.py. Port InboundConsumer onto them: a _ConsumerIO
adapter over bus.publish_outbound + the registry, one ApprovalPolicy
replacing the config/session auto-approve checks, and a single
reply-interception point (registry.try_resolve) replacing the parallel
ask_user/HITL pending dicts. _resolve_ask_user and the approval section
of _stream_with_hitl are now thin engine calls.

Behavior unification (serve mode): an unrecognized HITL reply now
declines with the shared 'Unrecognized reply' notice instead of
rejecting-and-refeeding as a fresh turn, and /stop mid-approval cancels
cleanly — both via the shared parser.

* refactor: port CLI channel bridge onto the interaction engine

Replace the ~250-line parallel bodies of channel_ask_user_prompt /
channel_hitl_prompt with thin bridges that run resolve_ask_user /
resolve_approval on the bus loop via
run_coroutine_threadsafe(...).result() (outer = engine per-flow timeout
+ slack, so the engine's own timeout fires first). The 15s send timeout
moves into the _BridgeIO adapter.

Delete the _pending_hitl / _hitl_lock / _hitl_auto_approve module
globals and the _register_hitl_wait / _try_set_hitl_reply /
_pop_hitl_reply helpers, absorbed by one bus-loop PendingReplyRegistry +
one ApprovalPolicy. The bus consumer feeds the registry via try_resolve
ahead of normal enqueue.

* refactor: restore serve-mode refeed for unrecognized HITL replies

Gate-review fix: the engine no longer decides transport policy for
unparseable approval replies. resolve_approval now returns an
ApprovalOutcome carrying unrecognized_reply (raw text) when parsing
fails, sending no feedback itself; recognized reject keeps the sharedi
rejection message.

Consumer driver (serve mode) restores the pre-engine semantics: an
unrecognized reply rejects the pending action, confirms with the
rejection message, and the text is re-dispatched as a NEW agent turn —
_stream_with_hitl returns the captured text and _handle_message starts
the refeed turn only after the current one has released the chat lock
(old fall-through ordering). CLI bridge keeps its old no-refeed path
byte-for-byte: 'Unrecognized reply. Action rejected.' and decline.

Tests: serve refeed pinned end-to-end (prompt → unrecognized text →
rejection feedback → text reaches the stream path as a new turn), CLI
no-refeed pinned (notice sent, nothing enqueued), engine test updated to
assert the outcome struct with no engine-side feedback.

* refactor: polish the interaction engine surface

- English feedback strings (Approved / Rejected / auto-approving)
- drop the consumer's backwards-compatible re-exports and both modules'
  private timeout aliases; callers use the canonical interaction names
- replace byte-for-byte prompt goldens with structural format tests and
  assert feedback via the shared constants instead of string literals
- strip audit/design shorthand (R1/R2/G3, stage numbers) from comments

* fix: propagate pending reply task cancellation

* fix: preserve reply context when refeeding HITL replies

* fix: honor HITL session grants without bus loop

* fix: bound bridge waits by send latency

* fix: handle empty ask_user replies explicitly

* chore: remove stale interaction helpers

* fix: harden interaction engine reply edge cases

Review follow-ups on the interaction engine:
- normalize ask_user choices before .get(): the tool args come from model
  JSON and only presence is validated, so plain-string choices must render
  and parse instead of crashing the turn
- treat only None as an approval timeout, so an empty/media-only reply
  flows through the unrecognized path and serve mode refeeds it with its
  preserved context
- intercept prompt replies before _get_thread_id so a consumed reply
  cannot create an orphan graph thread or touch the sender-session LRU
- close engine coroutines the bridge failed to schedule (no bus loop /
  scheduling error) to avoid never-awaited warnings
- clear pending-response and channel-request state in the bridge test
  fixture

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-14 14:59:41 +00:00
jfilipiuk 753c745405 fix: silence YAML-docstring noise from custom-app OpenAPI scan (#317) 2026-07-13 15:38:28 +01:00