452 Commits

Author SHA1 Message Date
X-iZhang b40b6f784d feat: clarify comment in NewCommand to align with explicit expert clearing 2026-08-07 17:07:01 +01:00
X-iZhang aff63ccd65 feat: enhance expert invitation handling with case-insensitive matching and session management 2026-08-07 17:07:01 +01:00
jfilipiuk ab1a6b0062 feat: adopt the AGENTS.md expert-skill contract (#404)
* feat: detect expert skills by AGENTS.md presence

* feat: let the orchestrator choose the expert dispatch tool

* refactor: drop per-skill expert dispatch classification

* chore: remove duplicate TestSkillManager test classes left by rebase

* fix: clear legacy actor fields when AGENTS.md declares the expert

* refactor: move the expert prompt into ActiveTeamMiddleware
2026-08-07 17:07:01 +01:00
jfilipiuk bd2464423a feat: agent-teams part D - async expert dispatch mechanism (#391)
* fix(deps): pin openrouter below 0.11 to avoid SSE stream regressions (#373)

* fix(openrouter): address SSE stream leak by closing response iterator

* refactor(openrouter): pass through SDK args in SSE leak patch

* test(openrouter): make SSE leak tests version-agnostic across SDK generations

* fix(openrouter): remove SSE stream leak patch and update dependencies

* fix: repair interrupted tool call history (#366)

* fix: repair interrupted tool call history

Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths.

* fix: repair malformed tool calls and dedupe repair warnings

Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted
threads with syntactically invalid tool calls get synthesized error
results and are accepted by strict providers.

Preserve the originating tool call's name in the synthesized ToolMessage,
and deduplicate repair warnings per unique tool-call id via a warned set
owned by the middleware instance, since the middleware rewrites the
request but not thread state.

Document the middleware's scope versus deepagents' PatchToolCallsMiddleware
(orphan ToolMessage dropping and mid-run coverage).

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Update README.md

* Update README.md

* Update README.zh-CN.md

* fix: scrub host path from skill_manager output and guard batch install (#377)

* fix: scrub host path from skill_manager output and guard batch install

* test: tighten install leak guards to catch host path in either tier

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: log missing async-subagent tools at DEBUG, not WARNING (#378)

* fix: log missing async-subagent tools at DEBUG, not WARNING

* fix: distinguish load_subagents callers via async_swap_pending flag

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix: default reasoning context for codex proxy Responses API (#380)

* fix: set langgraph and codex proxy runtime defaults

* fix: address runtime default review feedback

* fix: drop langgraph dev env defaults per maintainer review

langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment,
so the reported KeyError cannot come from this flow; the env defaults added
here were unnecessary. Scope the PR back to the codex proxy reasoning
context fix only.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* release: v0.2.4 (#389)

* fix: add support for new Anthropic models and enhance adaptive thinking tests

* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models

* fix: update version to v0.2.4 in badges, README, and project files

* fix: update Star History chart links in README and README.zh-CN

* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog

* fix: update wechat group image in assets

* refactor(runtime): centralize async bridges under an owned runtime (#376)

* feat(runtime): add application-scoped async runtime

* refactor(cli): use owned runtime for session stats

* refactor(onboard): use the owned async runtime

* docs(runtime): record async bridge ownership

* refactor(middleware): keep sync fallback synchronous

* refactor(mcp): load tools on an owned runtime

* refactor(cli): share owned runtime across entry points

* refactor(channels): make inbound sync bridge explicit

* refactor(stream): run Rich streaming on owned runtime

* chore(runtime): remove nest-asyncio dependency

* refactor(asyncio): require active loops in async code

* docs(runtime): document final event loop ownership

* fix(stream): cancel stalled owned streams

* fix(cli): recover cleanly from stream cancellation

* fix(runtime): drain executor work before shutdown

* fix(runtime): terminate cancelled shell process trees

* fix(models): let fallback bypass selector failures

* fix(cli): reset interrupt handling between turns

* docs: rm implementation spec

* fix(serve): cancel active turns during shutdown

* fix(runtime): protect settlement from waiter cancellation

* fix(backends): reject empty shell commands

* fix(runtime): terminate descendants after shell exit

* fix(mcp): keep standalone discovery off channel loop

* fix(cli): own and settle interactive prompt cancellation

* fix(serve): keep channel sends off runtime loop

* fix(stream): scope cancel context to iterator steps

* refactor(serve): require the owned async runtime

* fix(channels): keep interactive sends off runtime loop

* fix(selector): surface fallback without log spam

* test(runtime): normalize Windows shell marker

* fix(cli): serialize interactive session turns

* fix(shell): bound output drain after termination

* fix(ui): do not retry owned runtime failures

* fix(shell): allow signal-safe registry reentry

* fix(shell): avoid terminating reused process ids

* fix(channels): preserve streaming send order

* fix(cli): report runtime shutdown timeouts cleanly

* fix(mcp): guide async callers to async loader

* docs(runtime): clarify reserved async bridge APIs

* fix(runtime): bound code interpreter cleanup

* test(shell): use active Python for drain regression

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* feat: add payload-aware EvoAsyncSubAgentMiddleware

* feat: register expert_container_async graph for async expert dispatch

* feat: fold installed expert skills into async subagent registry

* fix: accept 'async' as valid default_dispatch value

* feat: dispatch-aware ActiveTeamMiddleware cue (task vs start_async_task)

* fix: drop future annotations in expert_async_subagent so ToolRuntime injects

* feat: surface output_path and skill_name to expert container as runtime cue

* fix: extend AsyncWatcher client cache with expert specs for completion nudge

* feat: teach main agent the async-expert return envelope shape

* feat: propagate cfg.model to expert-async runs.create via ClientCacheProxy

* chore: guard AsyncWatcher client-cache extension against upstream rename

* test: cover output_path runtime-context tail block and wrong-type guard

* docs: drop out-of-repo notes/ ref from expert_container_async module doc

* fix: warn on unrecognized default_dispatch frontmatter value

* fix: reject empty-body expert skills on async dispatch to match sync policy

* docs: explain why expert container includes general-purpose subagent

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385)

* fix: propagate langgraph dev bind port into subprocess env for self-loop URL

* fix: keep parent env authoritative over workspace .env for mapped keys

* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins

* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS

* fix: merge .env via dotenv_values to close empty-value and RMW-race edges

* chore: align docstrings after .env-merge rework

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* Add Requesty as an LLM provider (#346)

* Add Requesty as an LLM provider

* Address review: Requesty prompt caching, model ordering, key validation

- Declare Anthropic-style prompt caching for Requesty Claude models by
  default (mirroring the OpenRouter behavior), with an opt-out flag
  EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
  provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
  overrides native/OpenRouter models for names it shares with them
  (the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
  invalid/missing key (public catalog), so it cannot validate a key. Use a
  minimal authenticated /v1/chat/completions request instead (200 = valid,
  403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).

* Validate Requesty key against auth layer, not a specific model

The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix(llm): filter unnamed tool calls (#390)

* fix(llm): filter unnamed tool calls

* test(llm): cover tool call sanitization branches

* fix(llm): repair unnamed tool calls in middleware

---------

Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>

* fix(middleware): mount tool-history repair on sync subagents and harden raw tool-call vetting (#393)

* feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation

* fix(tests): add test for dropping non-list raw tool calls in repair_tool_history

* Add Atlas Cloud LLM provider (#388)

* Add Atlas Cloud LLM provider

* Add Atlas Cloud onboarding support

* fix(validators): update atlascloud key validation to handle insufficient balance case

---------

Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>

* fix: prepend EvoAsyncSubAgentMiddleware for prefix cache stability

* docs: clarify list_dispatchable_experts covers both dispatch shapes

* fix: guard async expert fold-in against reserved-name collisions

* fix: honest advertising surfaces for async expert dispatch

* fix: compose expert persona into base-stack system_message

* fix: drop payload from start_async_task, inject skill_name by construction

* feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395

- Introduced TodoListMiddleware to the middleware stack for better task management.
- Updated HITL interrupt configuration to include 'delete' operations requiring approval.
- Implemented error handling for delete operations in read-only and memory backends.
- Enhanced approval prompt formatting to display file paths for delete actions.
- Added tests to ensure delete operations are correctly blocked or prompted for approval.
- Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality.

* revert: drop skill_manager from sync expert-container tool_registry

* revert: drop skill_manager from async expert-container tools

* fix: drop removed ASYNC_TASK_SYSTEM_PROMPT import for deepagents 0.7.0

* fix: mock list_dispatchable_experts in single-cue test for CI

* chore: drop stale output_path from async container graph docstring

* fix: forward configurable_extra through owned-runtime and HITL re-invocations

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: Sanjay Santhanam <51058514+Sanjays2402@users.noreply.github.com>
Co-authored-by: Yougang Lyu <82445958+youganglyu@users.noreply.github.com>
Co-authored-by: houren Antony <2212222@mail.nankai.edu.cn>
Co-authored-by: dinos <dinospk1999@gmail.com>
Co-authored-by: Thibault Jaigu <84420566+Thibaultjaigu@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
Co-authored-by: nightcityblade <jackchen@haloailabs.com>
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: nb213 <binyangzhu000@gmail.com>
Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
2026-08-07 17:07:01 +01:00
jfilipiuk 3cda9894c7 feat: agent-teams part C - expert selection UX (depends on part B) (#371)
* feat: bias main-agent delegation toward configurable.active_teams

* feat: add /experts and /expert TUI commands for expert-skill summoning

* feat: align expert-selection wording with WebUI (invite/dismiss)

* chore: clear active_teams on /new, cleanup active_team.py comment

* fix: use local append_to_system_message in ActiveTeamMiddleware

* fix: prevent configurable_extra from overriding thread_id

* fix: cache expert-skill lookup for /expert completions

* fix: suppress /expert completions past first arg and on exact match

* fix: invalidate /expert completion cache on skill install/uninstall

* fix: refuse /expert invites for non-dispatchable expert skills

* fix: fire /expert cache invalidation on every install_skill / uninstall_skill path

* fix: propagate active_teams to Rich CLI and serve dispatch surfaces

* fix: keep invited experts across channel shutdown

* fix: match /expert completions case-insensitively
2026-08-07 17:07:01 +01:00
jfilipiuk 3a1dbf0a0f feat: agent-teams part B - expert-skill backend mechanism (#370)
* feat: add expert-skill schema and type filter to skill_manager

* feat: fold installed expert skills into main-agent subagent registry

* feat: add GET /api/teams listing expert skills for gallery

* chore: cache SKILL.md body on SkillInfo, cleanup expert-container comments

* fix: register skill_manager in expert-subagent tool_registry

* fix: catch UnicodeDecodeError in expert-skill body loader

* fix: guard expert subagent registration against name collisions

* fix: skip expert registration when SKILL.md body is empty

* fix: drop redundant str() guards on expert-skill frontmatter

* fix: harden SKILL.md parsing on expert-registration hot path
2026-08-07 17:07:01 +01:00
jfilipiuk b5b01d50c2 feat: agent-teams part A - TUI panel visibility + DELEGATION_STRATEGY routing (#369)
* feat: document sync / QuickJS panel / async dispatch modes in DELEGATION_STRATEGY

* feat: surface QuickJS panel dispatches in TUI

* fix: finalize running panel dispatches on turn cancellation

* fix: replace panel widget cancel path with public finalize API

* fix: drop dead isinstance guard on panel dispatch duration_ms

* fix: re-arm panel widget when a new dispatch arrives after finalize

* fix: stop panel timer when finalize_running finds zero running rows

* fix: register panel widget in cleanup dict before awaiting mount
2026-08-07 17:07:01 +01:00
Ziheng Zhang 33979e5371 fix(llm): support Volcengine Coding model aliases (#411)
* fix(llm): support Volcengine Coding model aliases

* refactor(llm): add Volcengine Coding provider

* style: format Volcengine Coding test
2026-08-07 14:30:23 +08:00
Xiaohui Yan 3c5cc831c0 Feat/configurable bind host (#402)
* feat: configurable bind host for WebUI and langgraph dev (refs #400)

WebUI mode was only reachable from the machine running it: the front-end
got no bind interface, and `start_langgraph_dev(...)` was called without a
host, so both servers stayed on loopback with no way to widen them.

Adds two config fields with deliberately different defaults:

  webui_host        = 0.0.0.0    front-end serves the app shell, no secrets
  langgraph_dev_host = 127.0.0.1  unauthenticated API, agent can run shell

The design hinges on separating bind address from client address. Only
bind() uses the configured interface; every consumer that *connects*
(health probes, occupancy checks, async sub-agent self-dispatch) goes
through the new `_probe_host`, which maps a wildcard bind back to
loopback and honors a pinned interface verbatim. `_can_bind_port` is the
one exception and binds the literal host, since it must replicate the
bind the server itself will attempt.

  - manager.py: `_probe_host`, `_is_loopback_host`, `_format_hostport`;
    host kwarg threaded through the probes and `start_langgraph_dev`,
    which now emits `--host` and propagates
    EVOSCIENTIST_LANGGRAPH_DEV_HOST to the subprocess
  - sdk.py: `langgraph_dev_url` tracks host as well as port;
    EvoScientist.py reuses it instead of an inline f-string
  - server.py: `--host` flag mirroring `--port`, plus a red PUBLIC BIND
    banner whenever the bind is not provably loopback
  - webui.py: forwards both hosts; the front-end is widened via HOSTNAME
    because @evoscientist/webui ships no --host flag — its bin launcher
    does `HOSTNAME: process.env.HOSTNAME || "127.0.0.1"`. The warning is
    gated on the backend host only, so the shipped front-end default
    doesn't print a banner on every launch

Verified end to end against a live server: requesting 0.0.0.0 yields a
socket listening on 0.0.0.0 with the health probe correctly resolved to
127.0.0.1, while the default still binds 127.0.0.1 only.

Note: webui_host defaulting to 0.0.0.0 is a behavior change — upgrading
users will find the front-end reachable from the LAN.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: default both bind hosts to 0.0.0.0, add --host and wizard host rendering (closes #400)

Completes the remaining items from #400.

  - `langgraph_dev_host` now defaults to 0.0.0.0, matching `webui_host`.
    Remote WebUI use needs both anyway (the UI reaches the backend from the
    browser, not server-side), so a loopback backend default just meant every
    remote user hit a silently failing UI. `_DEFAULT_HOST` and sdk's
    `DEFAULT_LANGGRAPH_DEV_HOST` follow, so there is one story about where
    these servers listen.

    SECURITY: this exposes an unauthenticated API whose agent can run shell
    commands. The red PUBLIC BIND banner consequently fires on every launch
    while exposed — kept deliberately, since the exposure is real and the
    escape hatch (`--host 127.0.0.1` / `config set langgraph_dev_host`) is
    only discoverable if we say so. READMEs now lead with the warning and
    document the SSH-tunnel alternative.

  - `EvoSci --host <ip>` on the WebUI launch path, driving both servers. In
    WebUI mode they are two halves of one surface; moving only one leaves the
    UI loading but unable to reach the agent. Blank values are dropped rather
    than written as an empty override that would beat the config file.

  - Onboarding wizard no longer prints hard-coded `http://127.0.0.1:{port}` /
    `http://localhost:{port}` (steps.py:160, :223) — both render the
    configured bind through `_base_url` / `_format_hostport`, so a pinned
    interface is reported honestly and a wildcard still shows loopback.

Verified against a live server: with no host argument at all, resolution
through EvoScientistConfig yields a socket listening on 0.0.0.0, a client URL
of http://127.0.0.1, and the warning gate returning True.

Still open and tracked separately: the front-end takes its backend URL from
browser input: `@evoscientist/webui` reads only HOSTNAME, PORT and
EVOSCIENTIST_LANGGRAPH_DEV_PORT, so advertising a backend URL needs a change
in that repo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: bump setup-uv v6 -> v9.0.0 to drop the deprecated node20 runtime

GitHub now warns that setup-uv@v6 targets Node.js 20 and is being forced
onto Node.js 24. v7.0.0 is the release that made that switch, so anything
>= v7 clears the warning; v9.0.0 is current.

Pinned to the full tag deliberately: setup-uv stopped publishing major and
minor tags in v8.0.0 as supply-chain hardening, so `@v9` and `@v8` return
404 and would fail the job outright. Releases are immutable from v8 on, so
the full tag is as tamper-proof as a SHA. Comment left in lint.yml because
"simplifying" this back to `@v9` is an easy and CI-breaking mistake.

actions/checkout@v5 is already node24 and needs no change.

Note: v9.0.0 flips the `prune-cache` default to false (upstream did this to
ease load on PyPI infrastructure). None of these workflows set it, so they
follow the new default and Actions cache usage may grow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(cli): correct --host help text and warn on public bind in non-WebUI modes

The --host help claimed "WebUI mode only", which is wrong in a way that
matters for security. `--host` writes `langgraph_dev_host` unconditionally,
and `_ensure_async_subagent_server` auto-starts that backend for tui / cli /
serve as well — the langgraph dev server is shared across UI modes. So the
flag narrows or widens the agent API in every mode, and only `webui_host` is
actually WebUI-specific. Reported against cli/commands.py.

The documentation error hid a real gap: the PUBLIC BIND banner lived only in
deploy/server.py and deploy/webui.py, so a plain `EvoSci` session bound
0.0.0.0 with no runtime signal whatsoever — and `--help` is opt-in, so
fixing the text alone would not surface it. Added the same banner to the
shared CLI path, gated on `is_async_subagents_available()`: ensure_langgraph_dev
fails soft (async degrades to in-process delegation), and warning about a
bind that never happened would be worse than staying quiet.

READMEs (EN + zh-CN) get the same correction — the warning block sat inside
the Desktop WebUI section and read as WebUI-scoped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(deploy): strip the config-derived bind host, not just the CLI one

`deploy()` only stripped the `--host` branch. When the flag was omitted,
`getattr(config, "langgraph_dev_host", ...)` flowed unstripped into
`_is_port_occupied`, `is_langgraph_dev_running`, `start_langgraph_dev` and
the banner. `run_webui` already strips unconditionally; this aligns the two.

Reachable because `deploy()` reads through `getattr` and is routinely handed
duck-typed config objects (tests, embedders) that never run
`EvoScientistConfig.__post_init__`, which is what normally normalizes these
fields.

Worst case was not just a bad bind: `_is_loopback_host(" 127.0.0.1 ")` is
False, so a padded loopback value would print a false PUBLIC BIND warning
while binding a string socket.bind() rejects outright — a security banner
saying the opposite of the truth.

Three regression tests added, each verified to fail against the old code.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* style: apply ruff format to the bind-host changes

The Lint workflow runs both `ruff check` and `ruff format --check`; I had
only been running the former locally, so five files landed unformatted and
failed CI. Whitespace and line-wrapping only — no semantic change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(security): keep the langgraph dev backend on loopback by default

The backend is an unauthenticated API whose agent can run shell commands,
and it is auto-started in every UI mode (tui/cli/webui/serve/deploy) — so a
0.0.0.0 default put it on the network for users who never asked. Restore
127.0.0.1 as the default and make 0.0.0.0 an explicit opt-in.

webui_host keeps its 0.0.0.0 default: the front-end serves the app shell
only and holds no credentials. run_webui already prints a remote-backend
hint when the front-end is exposed and the backend is not.

Help text and both READMEs are reframed around widening rather than
narrowing; the escape-hatch tests are inverted to assert the public-bind
opt-in survives into argv.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 16:15:49 +08:00
houren Antony d97ee2b917 fix(middleware): normalize blank tool_call_id before provider request (#399)
* fix(middleware): normalize blank tool_call_id before provider request

Closes #345.

Some streaming providers (notably Kimi and Zhipu on streamed tool calls)
occasionally emit tool calls whose id is an empty string, whitespace, or
None. Strict providers reject the next turn with `invalid tool_call_id`
(HTTP 400, code 3), freezing the thread after the very first tool round.

The existing repair logic skipped blank ids via a truthy check
(`if tool_call_id := call.get("id")`), so an AIMessage carrying a blank
id was passed through unchanged while its paired ToolMessage was dropped
as an orphan -- the next model call then 400'd.

This adds a pre-pass (`_normalize_blank_tool_call_ids`) before the main
repair loop that:

* assigns each blank-id tool call on an AIMessage a fresh `_repair_<uuid>` id
* pairs subsequent blank-id ToolMessages in arrival order (FIFO) so existing
  exchanges stay paired
* lets unmatched blank calls fall through to the main loop, which synthesizes
  an interrupted-result ToolMessage with the fresh id
* drops `additional_kwargs["tool_calls"]` on touched messages so
  langchain-openai's serializer falls back to the now-valid parsed form
  instead of preferring the raw form (which still carries the blank id)
* pre-adds the fresh ids to `warned` so repair stays silent on subsequent
  model calls (the fresh ids are non-deterministic across calls)

Verified end-to-end: `_convert_message_to_dict` on the repaired history
puts no blank id on the wire and preserves AIMessage<->ToolMessage pairing.

* fix(middleware): scope blank-id FIFO per exchange; cover invalid_tool_calls

Addresses CodeRabbit review on PR #399.

(1) Critical -- pending_slots FIFO leak across exchanges
--------------------------------------------------------
The FIFO queue of fresh ids for blank-id calls was never closed at
non-ToolMessage boundaries, so a later exchange's blank ToolMessage
could be paired with a stale id from an earlier, already-interrupted
exchange. The main loop then dropped the real tool result as an orphan
and synthesized a fake interrupted result for the real call:

    AIMessage1(blank A) -- interrupted
    HumanMessage
    AIMessage2(blank B)
    ToolMessage(blank)  -> popped A_new from FIFO front, not B_new

Fix: close pending_slots at every AIMessage / HumanMessage /
SystemMessage boundary via _close_unclaimed, mirroring close_pending()
in the main loop. Unclaimed slots are merged into `warned` so the
main loop's synthesized interrupted-result for that id stays silent
across model calls.

(2) Major -- invalid_tool_calls with blank id were silently dropped
-------------------------------------------------------------------
`any_changed` was set only inside the `tool_calls` loop, so a message
with a blank id only in `invalid_tool_calls` never entered the
model_copy update path and the blank id survived untouched.

langchain-openai's serializer puts `tool_calls + invalid_tool_calls`
on the wire when either parsed list is non-empty (it does NOT skip
invalid calls), so a blank id on an invalid call reaches the provider
just as readily as one on a valid call -- verified by direct
inspection of `_convert_message_to_dict`.

Fix: track `invalid_changed` separately and include it in the
update-path condition; also push invalid fresh ids into pending_slots
(valid-first ordering ensures a real ToolMessage for a valid blank
call never accidentally claims an invalid call's id).

Tests
-----
* test_pending_slots_scoped_per_exchange_not_global_fifo: the exact
  cross-exchange leak scenario CodeRabbit described.
* test_normalizes_blank_id_in_invalid_tool_calls_only: the
  invalid-only case that previously slipped through.

All 26 tests in test_tool_history_repair_middleware.py pass; full
suite 3044 passed, 13 skipped (Windows-compatible subset).

* fix(middleware): deterministic repair ids; don't pair invalid calls with orphan results

Addresses din0s review on PR #399.

(1) IDs are now deterministic, not uuid4
----------------------------------------
Each blank id is rewritten as `_repair_{msg_idx}_{v|i}{call_idx}` so the same
blank call gets the same id on every model call. The middleware re-runs on
every request but can only rewrite the outgoing request, not the thread
state, so random uuids made the wire payload unstable and grew the `warned`
set unboundedly. Deterministic ids let the main loop's existing `warned`-set
dedup suppress the synthesized-result warning from the second call on --
no pre-add hack needed. The `warned` parameter is therefore dropped from
`_normalize_blank_tool_call_ids` (and the `_close_unclaimed` helper removed
in favor of plain `pending_slots.clear()`).

(2) invalid_tool_calls fresh ids are no longer pushed to pending_slots
----------------------------------------------------------------------
Invalid calls are never executed by LangGraph (args can't be parsed), so no
real ToolMessage can claim their slot. Pushing it let an orphan blank
ToolMessage from some other call mis-pair with the invalid call, surfacing
the orphan's content under the invalid call's name. Without the push the
orphan keeps its blank id and the main loop drops it, which is what we want.

Tests
-----
* Updated `test_blank_id_repair_does_not_spam_warnings` to expect the
  one-time-warning-then-silent pattern (matches
  `test_warning_deduplicates_across_calls`) and asserts the deterministic id.
* New `test_invalid_blank_id_not_pushed_to_pending_slots` reproduces the
  exact mis-pairing scenario (orphan blank result + invalid call) and
  asserts the orphan content never leaks into a repaired tool result.

27 tests in test_tool_history_repair_middleware.py pass; ruff clean.
2026-08-03 15:02:15 +01:00
Xi Zhang d249e320bd feat: unified HITL approval for sync + async sub-agents (closes #387) (#396)
* feat: implement guard for dangerous commands and enhance HITL interrupt handling

* feat: enhance HITL approval mechanism and introduce session auto-approve decisions

* feat: add refuse_delete option to backend and enhance async delete guards

* feat: simplify delete method in CustomSandboxBackend and clarify adelete behavior

* feat: enhance allow-list behavior for command resolution and add related tests
2026-07-31 10:23:05 +01:00
Xi Zhang f81a8b086e feat(deps): upgrade deepagents to 0.7.0 with todos restore and delete gating- #395
- Introduced TodoListMiddleware to the middleware stack for better task management.
- Updated HITL interrupt configuration to include 'delete' operations requiring approval.
- Implemented error handling for delete operations in read-only and memory backends.
- Enhanced approval prompt formatting to display file paths for delete actions.
- Added tests to ensure delete operations are correctly blocked or prompted for approval.
- Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality.
2026-07-30 10:41:12 +01:00
nb213 4ddf7ebe52 Add Atlas Cloud LLM provider (#388)
* Add Atlas Cloud LLM provider

* Add Atlas Cloud onboarding support

* fix(validators): update atlascloud key validation to handle insufficient balance case

---------

Co-authored-by: binyangzhu000-sudo <224954946+binyangzhu000-sudo@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-07-29 14:09:44 +01:00
Xi Zhang 562ce0eb83 fix(middleware): mount tool-history repair on sync subagents and harden raw tool-call vetting (#393)
* feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation

* fix(tests): add test for dropping non-list raw tool calls in repair_tool_history
2026-07-28 18:14:37 +01:00
nightcityblade a6a8e19dcc fix(llm): filter unnamed tool calls (#390)
* fix(llm): filter unnamed tool calls

* test(llm): cover tool call sanitization branches

* fix(llm): repair unnamed tool calls in middleware

---------

Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-28 17:37:32 +01:00
Thibault Jaigu 8cac50e2ef Add Requesty as an LLM provider (#346)
* Add Requesty as an LLM provider

* Address review: Requesty prompt caching, model ordering, key validation

- Declare Anthropic-style prompt caching for Requesty Claude models by
  default (mirroring the OpenRouter behavior), with an opt-out flag
  EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
  provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
  overrides native/OpenRouter models for names it shares with them
  (the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
  invalid/missing key (public catalog), so it cannot validate a key. Use a
  minimal authenticated /v1/chat/completions request instead (200 = valid,
  403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).

* Validate Requesty key against auth layer, not a specific model

The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-07-28 15:35:41 +01:00
jfilipiuk 10c032450e fix: propagate langgraph dev bind port into subprocess env for self-loop URL (#385)
* fix: propagate langgraph dev bind port into subprocess env for self-loop URL

* fix: keep parent env authoritative over workspace .env for mapped keys

* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins

* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS

* fix: merge .env via dotenv_values to close empty-value and RMW-race edges

* chore: align docstrings after .env-merge rework

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-28 12:26:22 +00:00
dinos 8b1451cdda refactor(runtime): centralize async bridges under an owned runtime (#376)
* feat(runtime): add application-scoped async runtime

* refactor(cli): use owned runtime for session stats

* refactor(onboard): use the owned async runtime

* docs(runtime): record async bridge ownership

* refactor(middleware): keep sync fallback synchronous

* refactor(mcp): load tools on an owned runtime

* refactor(cli): share owned runtime across entry points

* refactor(channels): make inbound sync bridge explicit

* refactor(stream): run Rich streaming on owned runtime

* chore(runtime): remove nest-asyncio dependency

* refactor(asyncio): require active loops in async code

* docs(runtime): document final event loop ownership

* fix(stream): cancel stalled owned streams

* fix(cli): recover cleanly from stream cancellation

* fix(runtime): drain executor work before shutdown

* fix(runtime): terminate cancelled shell process trees

* fix(models): let fallback bypass selector failures

* fix(cli): reset interrupt handling between turns

* docs: rm implementation spec

* fix(serve): cancel active turns during shutdown

* fix(runtime): protect settlement from waiter cancellation

* fix(backends): reject empty shell commands

* fix(runtime): terminate descendants after shell exit

* fix(mcp): keep standalone discovery off channel loop

* fix(cli): own and settle interactive prompt cancellation

* fix(serve): keep channel sends off runtime loop

* fix(stream): scope cancel context to iterator steps

* refactor(serve): require the owned async runtime

* fix(channels): keep interactive sends off runtime loop

* fix(selector): surface fallback without log spam

* test(runtime): normalize Windows shell marker

* fix(cli): serialize interactive session turns

* fix(shell): bound output drain after termination

* fix(ui): do not retry owned runtime failures

* fix(shell): allow signal-safe registry reentry

* fix(shell): avoid terminating reused process ids

* fix(channels): preserve streaming send order

* fix(cli): report runtime shutdown timeouts cleanly

* fix(mcp): guide async callers to async loader

* docs(runtime): clarify reserved async bridge APIs

* fix(runtime): bound code interpreter cleanup

* test(shell): use active Python for drain regression

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-27 14:17:57 +01:00
Xi Zhang ac58caab7b release: v0.2.4 (#389)
* fix: add support for new Anthropic models and enhance adaptive thinking tests

* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models

* fix: update version to v0.2.4 in badges, README, and project files

* fix: update Star History chart links in README and README.zh-CN

* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog

* fix: update wechat group image in assets
2026-07-26 14:58:48 +01:00
houren Antony fe70599d3d fix: default reasoning context for codex proxy Responses API (#380)
* fix: set langgraph and codex proxy runtime defaults

* fix: address runtime default review feedback

* fix: drop langgraph dev env defaults per maintainer review

langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment,
so the reported KeyError cannot come from this flow; the env defaults added
here were unnecessary. Scope the PR back to the codex proxy reasoning
context fix only.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-24 15:29:34 +00:00
jfilipiuk 1863f0730c fix: log missing async-subagent tools at DEBUG, not WARNING (#378)
* fix: log missing async-subagent tools at DEBUG, not WARNING

* fix: distinguish load_subagents callers via async_swap_pending flag

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-23 19:07:30 +00:00
jfilipiuk 172c8409d4 fix: scrub host path from skill_manager output and guard batch install (#377)
* fix: scrub host path from skill_manager output and guard batch install

* test: tighten install leak guards to catch host path in either tier

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-23 18:55:51 +00:00
Sanjay Santhanam f802a49535 fix: repair interrupted tool call history (#366)
* fix: repair interrupted tool call history

Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths.

* fix: repair malformed tool calls and dedupe repair warnings

Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted
threads with syntactically invalid tool calls get synthesized error
results and are accepted by strict providers.

Preserve the originating tool call's name in the synthesized ToolMessage,
and deduplicate repair warnings per unique tool-call id via a warned set
owned by the middleware instance, since the middleware rewrites the
request but not thread state.

Document the middleware's scope versus deepagents' PatchToolCallsMiddleware
(orphan ToolMessage dropping and mid-run coverage).

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-22 15:43:58 +01:00
Xi Zhang e9857dde22 fix(deps): pin openrouter below 0.11 to avoid SSE stream regressions (#373)
* fix(openrouter): address SSE stream leak by closing response iterator

* refactor(openrouter): pass through SDK args in SSE leak patch

* test(openrouter): make SSE leak tests version-agnostic across SDK generations

* fix(openrouter): remove SSE stream leak patch and update dependencies
2026-07-22 16:03:57 +02:00
m4 3ce5614254 fix: harden tool-call protocol and fallback handling 2026-07-19 12:05:56 +08:00
Xi Zhang 042da63d54 feat(llm): add Kimi K3 support (#367)
* feat(context-window): add Kimi K3 model with 1M context window

* feat(openrouter): implement structured output for Kimi K3 and add 429 retry handling

* Refactor code structure for improved readability and maintainability
2026-07-18 00:12:08 +01:00
dinos 06a9511bdd fix(channels): telegram slash commands (#364) 2026-07-17 16:05:49 +01:00
jfilipiuk 584b9d24ac fix(tool-selector): cap chatter and streaming volume (#350)
* feat: add disable_streaming helper for tool-selector's internal model

* feat: apply disable_streaming to the tool-selector's model in the factory

* fix(tool-selector): hide selector model call from public event streams

* feat(tool-selector): log a WARNING when the selector's model returns a duplicate-tool_calls flood

* chore: log flood-detector errors, document parent-method drift risk, tighten tests

* refactor(tool-selector): switch to nostream tag via model-field wiring, drop subclass
2026-07-17 15:28:16 +01:00
Xi Zhang e8399b7c94 fix(tui): window slash-command completions to terminal height (#362)
* feat(tui): implement completion popup rendering and windowing logic

* feat(tui): enhance completion popup with dynamic row budgeting and CSS adjustments

* Refactor picker widgets to use shared base class for improved code reuse

- Introduced `picker_base.py` to encapsulate common functionality for picker widgets.
- Updated `ModelPickerWidget`, `SkillBrowserWidget`, and `ThreadPickerWidget` to inherit from `PickerWidgetBase`.
- Implemented selection helpers (`first_selectable_index`, `move_selection`) in `picker_base.py` for consistent item navigation.
- Refactored rendering and selection logic in each widget to utilize the new base class methods.
- Added tests for picker functionality to ensure behavior remains consistent post-refactor.
2026-07-16 16:17:58 +01:00
Xi Zhang 0f709cff8b fix: cascade-cancel runs on thread deletion + startup orphan sweep (#358) (#359)
* feat: implement bulk cancellation of non-terminal runs before thread deletion

* test: enhance thread cancellation tests and add fake restore for orphaned runs sweep

* feat: enhance run cancellation logic to support status filtering during thread deletion

* feat: add langgraph-sdk dependency for enhanced functionality
2026-07-16 12:46:09 +01:00
dinos 05dfffbc73 fix(llm): use native langchain-deepseek SDK (#349) 2026-07-15 17:13:33 +01:00
dinos 01845f4311 refactor: route middleware display events through an injected event sink (#343)
* chore: add pytest-asyncio in auto mode

* test: migrate channel and stream tests to native async

Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.

* test: migrate command and model/middleware tests to native async

Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.

* test: migrate TUI, notifier, gateway, and session tests to native async

TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.

* test: replace direct asyncio.run() calls with native async tests

Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.

* test: drop undeclared anyio markers and delete run_async helper

The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).

* test: add autouse fixture for watcher cleanup

* refactor: remove redundant hasattr calls

* refactor: add typed middleware event sink and thread through assembly

Add MiddlewareEventSink protocol + NoOpSink in middleware/events.py
with a documented any-thread non-blocking contract (contract test uses a
deliberately-slow fake sink). Thread an optional `events` parameter
through create_cli_agent -> _get_default_middleware -> tool selector /
model fallback constructors; subagent stacks are always forced to
NoOpSink.

* refactor: inject a notifier port into async-watcher and background middleware

Add public pre_cancel_watcher() and enqueue_task_notification() to
cli/async_notifier.py and a small NotifierPort protocol
(middleware/notifier.py) that the module satisfies structurally.
AsyncWatcherMiddleware and BackgroundExecutionMiddleware now receive the
port by constructor injection at the composition root, deleting the lazy
'from ..cli import async_notifier' imports and the private
_watcher_by_thread / _enqueue pokes.

* refactor: invert tool-selection ownership onto a frontend event sink

The adaptive tool selector now reports on_tool_selection_started /
on_tool_selection / on_tool_selection_ended to the injected sink instead
of writing four process-global module variables. The frontend sink
(stream/sink.py FrontendEventSink) owns the selected/total/active state
with consume-once + dedup-vs-last-emitted semantics;
stream/tool_selection.py reads that sink object (a ToolSelectionView)
rather than reaching into tool_selector's globals.

Deleted: the 4 module globals, the cross-module mutations in
tool_selection.py, the track_stream_selection flag, the now-vestigial
_ToolSelectionTrackerMiddleware, reset_tool_selection_state_for_tests,
and the autouse conftest fixture. The sink is threaded from the two
interactive frontends through create_runtime_gateways ->
LocalGraphGateway (read side) and _load_agent -> create_cli_agent (write
side); subagent / headless stacks get NoOpSink.

* refactor: route model-fallback narration through the injected event sink

Delete the _ui_emit_fn / set_ui_emit module global and the
..stream.console import from model_fallback.py. The fallback middleware
now reports through its injected sink: the fallback transition via the
structured on_model_fallback (the frontend formats the '-> Falling back
to ...' line), and the surrounding narration (primary-failure header,
per-attempt outcome, exhaustion, non-fallbackable rejection) via
emit_fallback_notice, preserving the exact user-facing text. The TUI
binds its _append_system as the sink's fallback display where it used to
call set_ui_emit (cleared on exit); the Rich CLI's sink prints to the
console. _try_fallbacks / _guard_and_fallback take the sink.

* refactor: declare events on the GraphGateway protocol

Both gateway implementations now carry an explicit events attribute
(LangGraphServerGateway holds None — no frontend renders middleware
events across the HTTP boundary), so the four call sites use plain
attribute access instead of getattr probing an implicit contract.

* refactor: bind fallback display via the closure-scoped concrete sink

The App methods used gateway.events (typed as the read-side view) and
hasattr-probed for the concrete FrontendEventSink API. The enclosing
factory creates that sink two hundred lines up — close over it directly:
no probing, fully typed, and it becomes a constructor parameter
naturally when the App class is hoisted out of the factory.

* fix: end tool selection before fallback handler

* fix: keep fallback display errors non-fatal

* fix: preserve selector suppression for default streams

* fix: restore fallback notice console display

* refactor: consolidate fallback narration events

* refactor: clean middleware event sink plumbing

* fix: type gateway session events

* refactor: make all event protocols runtime-checkable

MiddlewareEventSink already carried @runtime_checkable (the stream
binding guard isinstance-checks it); ToolSelectionView and SessionEvents
now match, so mirroring that pattern against any of the three protocols
works instead of raising TypeError.

* fix(cli): close QuickJS workers after one-shot failures

* fix(cli): honor no-thinking in final output

* fix(channels): report failed startup accurately

* fix(channels): make Telegram cleanup idempotent

* fix(tui): skip command sync during exit

* fix(channels): preserve startup state during retries

* refactor(channels): share pending startup status

* refactor(cli): expose channel startup snapshot

* fix(tui): move channel startup off event loop

* test(channels): release retry gate on assertion failure

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-14 22:34:17 +00:00
dinos db1abce8d8 refactor: extract a shared HITL/ask_user interaction engine (#342)
* chore: add pytest-asyncio in auto mode

* test: migrate channel and stream tests to native async

Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.

* test: migrate command and model/middleware tests to native async

Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.

* test: migrate TUI, notifier, gateway, and session tests to native async

TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.

* test: replace direct asyncio.run() calls with native async tests

Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.

* test: drop undeclared anyio markers and delete run_async helper

The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).

* test: add autouse fixture for watcher cleanup

* refactor: remove redundant hasattr calls

* refactor: extract shared HITL/ask_user interaction grammar

Extract prompt/question formatting, the reply grammar (approval letters,
ask_user choice letters + the 'Other' sub-flow, stop-commands), the
ApprovalPolicy (config auto-approve rule + session registry +
session-key derivation), per-flow timeout constants, and the bilingual
feedback strings into channels/interaction.py. Both drivers now point at
the shared functions: this reverses cli/channel.py's imports of consumer
privates and closes the /stop drift at the parsing layer (serve-mode
ask_user now checks stop-commands before parsing an answer, matching the
CLI path).

* refactor: add interaction engine + registry; port InboundConsumer

Introduce InteractionIO (transport adapter Protocol),
PendingReplyRegistry (one asyncio-based reply router per process), and
the engine coroutines resolve_ask_user / resolve_approval in
channels/interaction.py. Port InboundConsumer onto them: a _ConsumerIO
adapter over bus.publish_outbound + the registry, one ApprovalPolicy
replacing the config/session auto-approve checks, and a single
reply-interception point (registry.try_resolve) replacing the parallel
ask_user/HITL pending dicts. _resolve_ask_user and the approval section
of _stream_with_hitl are now thin engine calls.

Behavior unification (serve mode): an unrecognized HITL reply now
declines with the shared 'Unrecognized reply' notice instead of
rejecting-and-refeeding as a fresh turn, and /stop mid-approval cancels
cleanly — both via the shared parser.

* refactor: port CLI channel bridge onto the interaction engine

Replace the ~250-line parallel bodies of channel_ask_user_prompt /
channel_hitl_prompt with thin bridges that run resolve_ask_user /
resolve_approval on the bus loop via
run_coroutine_threadsafe(...).result() (outer = engine per-flow timeout
+ slack, so the engine's own timeout fires first). The 15s send timeout
moves into the _BridgeIO adapter.

Delete the _pending_hitl / _hitl_lock / _hitl_auto_approve module
globals and the _register_hitl_wait / _try_set_hitl_reply /
_pop_hitl_reply helpers, absorbed by one bus-loop PendingReplyRegistry +
one ApprovalPolicy. The bus consumer feeds the registry via try_resolve
ahead of normal enqueue.

* refactor: restore serve-mode refeed for unrecognized HITL replies

Gate-review fix: the engine no longer decides transport policy for
unparseable approval replies. resolve_approval now returns an
ApprovalOutcome carrying unrecognized_reply (raw text) when parsing
fails, sending no feedback itself; recognized reject keeps the sharedi
rejection message.

Consumer driver (serve mode) restores the pre-engine semantics: an
unrecognized reply rejects the pending action, confirms with the
rejection message, and the text is re-dispatched as a NEW agent turn —
_stream_with_hitl returns the captured text and _handle_message starts
the refeed turn only after the current one has released the chat lock
(old fall-through ordering). CLI bridge keeps its old no-refeed path
byte-for-byte: 'Unrecognized reply. Action rejected.' and decline.

Tests: serve refeed pinned end-to-end (prompt → unrecognized text →
rejection feedback → text reaches the stream path as a new turn), CLI
no-refeed pinned (notice sent, nothing enqueued), engine test updated to
assert the outcome struct with no engine-side feedback.

* refactor: polish the interaction engine surface

- English feedback strings (Approved / Rejected / auto-approving)
- drop the consumer's backwards-compatible re-exports and both modules'
  private timeout aliases; callers use the canonical interaction names
- replace byte-for-byte prompt goldens with structural format tests and
  assert feedback via the shared constants instead of string literals
- strip audit/design shorthand (R1/R2/G3, stage numbers) from comments

* fix: propagate pending reply task cancellation

* fix: preserve reply context when refeeding HITL replies

* fix: honor HITL session grants without bus loop

* fix: bound bridge waits by send latency

* fix: handle empty ask_user replies explicitly

* chore: remove stale interaction helpers

* fix: harden interaction engine reply edge cases

Review follow-ups on the interaction engine:
- normalize ask_user choices before .get(): the tool args come from model
  JSON and only presence is validated, so plain-string choices must render
  and parse instead of crashing the turn
- treat only None as an approval timeout, so an empty/media-only reply
  flows through the unrecognized path and serve mode refeeds it with its
  preserved context
- intercept prompt replies before _get_thread_id so a consumed reply
  cannot create an orphan graph thread or touch the sender-session LRU
- close engine coroutines the bridge failed to schedule (no bus loop /
  scheduling error) to avoid never-awaited warnings
- clear pending-response and channel-request state in the bridge test
  fixture

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-14 14:59:41 +00:00
m4 4fc74e7da7 EvoScientist Ai4Sci
Docker / build (push) Has been cancelled
Build / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
2026-07-14 22:07:14 +08:00
jfilipiuk 753c745405 fix: silence YAML-docstring noise from custom-app OpenAPI scan (#317) 2026-07-13 15:38:28 +01:00
jfilipiuk 88ac9f5ba1 fix: surface real exception class+message in SSE error events (#315)
* fix: surface real exception class+message in SSE error events

* fix: tighten SSE error patch scope and key redaction

* fix: redact base64-style secret suffixes fully

* style: remove notes/ reference from the dosctring

* fix: rebuild env cache on each error call

* fix: route BaseException through serde.default on SSE/webhook paths

* fix: distinguish routed providers by request URL host

* feat: normalize provider-SDK exceptions via ErrorNormalizationMiddleware

* refactor: drop json_dumpb dataclass-bypass wrappers, superseded by middleware

* fix: guard _extract_host against SDK properties that raise

* refactor: derive provider tag from ModelRequest.model, not the exception

* refactor: drop serde.default patch and exception-based inference; ProviderStreamError.model_dump handles the emit

* refactor: move envelope helpers from patches.py to errors.py

* feat: extend ErrorNormalizationMiddleware coverage to every model-call path

* chore: clean up review findings from middleware pivot

* fix: pass through all langgraph.errors

* fix: move langgraph.errors pass-through into _normalize

* fix: pass through ContextOverflowError in _normalize
2026-07-13 14:17:56 +01:00
jfilipiuk 952e68efe3 fix: scope quickjs snapshot to turn to keep checkpoints small (#316)
* fix: scope quickjs snapshot to turn to keep checkpoints small

* fix: strip _quickjs_snapshot_payload from state/history responses instead of dropping mode=thread

* fix: recurse strip into nested subgraph StateSnapshot

* fix: drop conditional-snapshot gate that leaked repl slots

* fix: assert LangGraph state-shape invariants at import time

* refactor: discover graphs to filter from langgraph.json

* test: assert copy() preserves subclass; iterate langgraph.json for subagent coverage

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 12:24:52 +00:00
Mani Saint-Victor 2b28c46caf fix(llm): make gpt-5.x usable through ccproxy Codex OAuth (#324)
* fix(llm): make gpt-5.x usable through ccproxy Codex OAuth

Two independent blockers made current OpenAI models fail when routed
through ccproxy's Codex OAuth endpoint:

1. ccproxy's default Codex model mappings rewrite any gpt-*/o1-*/o3-*/
   claude-* model to gpt-5.3-codex before forwarding, silently overriding
   the configured model and failing outright on accounts where
   gpt-5.3-codex is not served ("The 'gpt-5.3-codex' model is not
   supported when using Codex with a ChatGPT account").
   start_ccproxy() now generates a config with empty codex model
   mappings and passes it via 'ccproxy serve --config'.

2. ccproxy forwards the client's own User-Agent upstream and only
   gap-fills its Codex headers, so the backend gates current models on
   the client identity ("The '<model>' model requires a newer version
   of Codex"). get_chat_model() now sends Codex-CLI-shaped
   originator/version/User-Agent headers when the ccproxy Codex adapter
   is detected, overridable via EVOSCIENTIST_CODEX_CLIENT_VERSION.

Verified live: gpt-5.5 and gpt-5.4 complete successfully through
ccproxy Codex OAuth on a ChatGPT Plus account with both fixes; each
fails without them.

* fix(ccproxy): harden Codex client routing

* fix(llm): keep Codex client identity consistent

* docs: clarify Codex version floor

* style: ruff format models.py after merge

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-07-13 11:04:47 +00:00
Mani Saint-Victor da6ca38d53 fix(llm): respect reasoning_effort setting on native OpenAI path (#321)
* fix(llm): respect reasoning_effort setting on native OpenAI path

The native OpenAI provider path hardcoded reasoning effort to xhigh for
gpt-5.4/5.5/codex models, silently ignoring the user's reasoning_effort
config setting. The OpenRouter path already honors the
EVOSCIENTIST_REASONING_EFFORT env var that settings.py exports from that
setting; this applies the same lookup on the native path, falling back
to the previous defaults when unset.

Adds a regression test and isolates the existing xhigh test from the
env var.

* fix(llm): preserve model reasoning defaults

* fix(llm): preserve GPT-5.6 reasoning default

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 10:42:55 +00:00
Mani Saint-Victor 19888c2db6 fix(ccproxy): raise startup timeouts (auth check 10s→30s, serve health 30s→180s) (#328)
* fix(ccproxy): raise auth status check timeout to 30s

ccproxy's CLI initializes its full plugin system on every invocation;
a cold 'ccproxy auth status' takes ~10s wall time on Apple Silicon,
so the 10s subprocess timeout made OAuth startup fail intermittently
with 'Auth check timed out' even when credentials were valid.

* fix(ccproxy): raise serve health deadline to 120s

ccproxy boot includes plugin init plus Codex CLI detection; measured
~76s to first healthy response on an Apple Silicon Mac (ccproxy-api
0.2.9). The 30s deadline in start_ccproxy() killed the process before
it could come up, failing OAuth startup with 'ccproxy did not become
healthy within 30 seconds'.

* fix(ccproxy): widen serve health deadline to 180s

Full startup measured at ~111s on a second cold run (Apple Silicon,
ccproxy-api 0.2.9); 120s left too little headroom for boot variance.

* fix(ccproxy): centralize startup timeouts

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 10:36:05 +00:00
Zixin Dong f72f7b93d5 feat(llm): add OpenRouter app attribution headers (#339) (#344)
* feat(llm): add OpenRouter app attribution headers (#339)

Attach EvoScientist app-attribution at the shared model-init layer so all
OpenRouter calls are credited to the project. langchain-openrouter maps
app_url/app_title/app_categories -> HTTP-Referer / X-Title /
X-OpenRouter-Categories. Applied only for the openrouter provider, via
setdefault so explicit caller kwargs win. Configurable through new
openrouter_http_referer / openrouter_app_title / openrouter_app_categories
settings and their EVOSCIENTIST_OPENROUTER_* env vars.

Closes #339

* refactor(llm): centralize OpenRouter attribution defaults + cap categories

Address PR #344 review:
- Define the app-attribution default constants once in config/settings.py
  (the config fields and llm/models.py both use them) instead of duplicating
  the literals across the two modules.
- Reduce the default categories to creative-writing,personal-agent and cap the
  sent list to OpenRouter's 2-per-request limit, warning when a configured list
  exceeds it, so extras are dropped predictably (and surfaced) here rather than
  being silently truncated server-side.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 11:08:59 +01:00
dinos 49770949da fix(langgraph): prefer executable in venv over path (#341) 2026-07-11 11:59:27 +00:00
X-iZhang 9042068094 feat(models): add support for GPT-5.6 variants and update context windows for Grok models 2026-07-11 00:51:06 +01:00
dinos d2452c54d5 Refactor onboarding OAuth flow for auxiliary models (#337)
* refactor(onboard): shared flow for ccproxy providers

* feat(onboard): support oauth configuration for auxiliary models

* fix(onboard): reuse main model auth for same-provider auxiliary

* fix(onboard): reconcile oauth providers
2026-07-08 18:28:44 +00:00
dinos a7b9e175c1 fix(config): set config.yaml permissions to 0x600 (#336) 2026-07-08 19:25:01 +01:00
dinos be3dd272c3 test: deflake timing-dependent tests (#335)
* test: deflake timing-dependent tests

Inject a clock into channel dedup tests, replace fixed async sleeps with
events/explicit flushes, and avoid wall-clock waits in background tests.

* coderabbit nit
2026-07-07 08:25:35 +01:00
X-iZhang df54d8498c chore: update version to 0.2.1 2026-07-05 10:17:28 +01:00
Wiktor Cupiał 1d117ff277 feat: completion enchancements (#302)
* feat: completion enchancements

* fix: handle exception

* fix: duplicate view

* fix: remove deadcode

* fix tab
2026-07-05 05:14:42 +00:00
renaissancefieldlite f086d77756 Fix UTF-8 config reads on Windows (#318)
* Fix UTF-8 config reads on Windows

* test: cover utf8 production loaders

* fix: read and write settings as utf8

* Apply ruff formatting
2026-07-05 05:10:24 +00:00
kalisgd0h bd54eaa0a4 feat(cli): add --output-format stream-json for headless clients (#309)
* feat(cli): add --output-format stream-json for headless clients

Emit EvoScientist's native event stream as line-delimited JSON on stdout
in single-shot (-p) mode, with all human output redirected to stderr so
stdout stays pure JSONL. Intended as the integration surface for
programmatic clients (e.g. an agent runtime) that drive EvoSci headlessly.

- stream/json_sink.py: write_events_as_json + stream_json sink, plus
  redirect_console_to_stderr helper for stdout purity
- cli/interactive.py: cmd_run gains output_format; stream-json branch runs
  the sink instead of the Rich renderer
- cli/commands.py: --output-format option + validation (stream-json
  requires -p; value must be text|stream-json)
- docs/stream-json.md: event-schema contract + example transcript
- tests: json sink serialization, CLI dispatch, console redirect, validation

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(cli): honor explicit --no-auto-mode over config in stream-json

Address CodeRabbit review (discussion_r3514041123): the auto-mode override
block only wrote to cli_overrides when the resolved value was True, so an
explicit --no-auto-mode silently fell back to a config that enables
auto-mode -- breaking "explicit flags always win" and leaving stream-json
running unattended despite the warning. Write auto_mode=False when the flag
is explicitly False. Add regression tests that capture the overrides.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 04:53:52 +00:00