* refactor: rename AGENTS.md to EXPERT.md per EvoSkills convention
* feat: resolve newly installed experts on start_async_task miss
* chore: reword the /expert invite hint to state the dispatch boundary
* docs: note the resolve-on-miss caller in build_expert_async_subagent_specs
* fix: thread the construction cfg through resolve-on-miss and symmetric setdefault
* fix: guard agent_map iteration against concurrent resolve-on-miss writes
* fix: warn once per broken expert on repeated resolve-on-miss walks
* fix: scope the /expert invite hint to newly installed experts
* docs: note live uninstalls as a resolve-on-miss limitation
* chore: isolate the warn-once collision test key from the route-specs suite
* fix: _extract_retry_after returns None for non-retryable errors
* fix: update _extract_retry_after to handle generic transient errors with default retry delay
* fix: extend base _non_retryable_patterns in channel subclasses
* fix(channels): merge structured SDK error check into non-retryable step
* fix(channels): decouple status code and SDK error code extraction in retry logic
- Independently evaluate HTTP status codes and structured SDK error codes
- Fix misleading doc comments for Feishu and DingTalk patterns
- Remove redundant try-except AttributeError on getattr with default
- Expand test coverage for dual-signal matrix and header parsing
* refactor(channels): simplify status code and SDK error extraction via channel overrides
- Handle httpx and aiohttp exceptions in base Channel class
- Override _extract_status_code and _extract_sdk_error_code in SlackChannel and DiscordChannel
- Replace mock exception types in comprehensive test suite with real httpx and aiohttp errors
- Add dedicated Slack and Discord retry error extraction test suites
* fix(channels): clean up Slack and Discord error code extraction
- Remove defensive string checks and attribute guards in SlackChannel
- Directly access exc.response.status_code and exc.response.get('error') in SlackChannel
- Remove unnecessary _extract_sdk_error_code override in DiscordChannel
- Use real SlackApiError, SlackResponse, and discord.HTTPException in unit tests
* refactor: reorder retry logic to prioritize non-retryable checks, remove aiohttp dependency, and clean up exception handling in base and channel modules.
* test(channels): skip Slack/Discord retry tests when the SDK extra is absent
The retry-extraction tests build real SlackApiError / discord.HTTPException
objects, but slack-sdk and discord.py are optional extras that the dev
dependency group does not install. Under CI's `uv sync --dev` all nine
tests failed with ModuleNotFoundError raised from the channel override.
Gate both test classes with skipif(find_spec(...) is None) so the suite is
green without the extras and the tests still run wherever they are installed.
* refactor(channels): replace retry-delay lookup with _extract_retry_delay
_extract_retry_after still read the server-supplied delay by probing
exc.retry_after and exc.response.headers via getattr/hasattr, the last
remnant of the pattern the extractors moved away from. Replace both steps
with one overridable hook, _extract_retry_delay, implemented against the
real exception types:
- base: httpx.HTTPStatusError -> Retry-After header (httpx.Headers is
case-insensitive; HTTP-date form remains unsupported)
- SlackChannel: SlackApiError -> Retry-After, matched case-insensitively
because SlackResponse.headers is a plain dict whose casing depends on the
HTTP client (same approach as slack_sdk's RateLimitErrorRetryHandler)
- TelegramChannel: telegram.error.RetryAfter.retry_after (int, or timedelta
under PTB_TIMEDELTA)
- DiscordChannel: discord.RateLimited.retry_after, which the old duck-typed
getattr matched and would otherwise have been lost
Drop the isinstance(retry, bool) and val >= 0 guards; no SDK produces those.
Delete the test that asserted the duck-typed attribute; add real-object tests
for each override, guarded like the existing SDK-dependent classes.
* ci: install the all-channels extra so SDK-dependent channel tests run
The Slack, Discord, and Telegram retry tests build real SDK exception
objects and are skipped when the SDK is absent. CI only ran `uv sync --dev`,
so those tests never executed there. Install the existing all-channels
extra alongside the dev group; the skipif guards remain for lean local runs.
* fix(channels): honor HTTP-date Retry-After and tolerate malformed values
RFC 9110 allows Retry-After as either delay-seconds or an HTTP-date. The
httpx path treated a date as unparseable and fell back to the 1.0 s default,
so a 503 asking for a specific wait was retried too early. Add
Channel._parse_retry_after, which returns delay-seconds as-is and converts
an HTTP-date to the non-negative seconds until it (tz-less dates read as
UTC).
SlackChannel used a bare float() on the header. A non-numeric value raised
inside the retry predicate, which escapes retry_async and drops the chunk
instead of retrying. Route Slack through the same helper so a bad header
falls back to _rate_limit_delay.
Addresses CodeRabbit review comments on base.py:866 and slack/channel.py:229.
* fix(channels): treat HTTP 400 and 404 as non-retryable
Both are permanent for a given request, so retrying burns the attempt
budget for nothing. Add them to _non_retryable_status_codes alongside
401/403.
Deliberately not a 4xx range check: 408 and 425 are retryable by
definition and 429 is handled by the rate-limit path. A test pins 408 as
still retryable so the range shortcut is not reintroduced later.
Partially addresses CodeRabbit's outside-diff comment on base.py:749-750.
* fix(channels): guard Slack retry extractors against raw aiohttp responses
slack_sdk attaches the bare aiohttp.ClientResponse to SlackApiError when a
JSON-declared body fails to parse. That object has neither status_code nor
get(), so _extract_status_code raised AttributeError inside should_retry,
replacing the original error and skipping the remaining attempts. Narrow
both extractors to SlackResponse/AsyncSlackResponse so such errors fall
through to the message patterns and retry as before. Add a wire-level
regression test against a local aiohttp server.
---------
Co-authored-by: Dinos Papakostas <dinospk1999@gmail.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
OpenRouter keys app pages by HTTP-Referer; X-Title only renames that
page. A custom openrouter_app_title on the default referer therefore
renamed the shared EvoScientist app page for everyone. Force the default
title whenever the resolved referer is the default, silently, so usage
keeps being attributed to EvoScientist; a private fork still overrides
both together.
Scheduled tasks gain an optional `rubric`: an acceptance checklist graded
after each run by deepagents' RubricMiddleware (LLM-as-a-judge on the
auxiliary model) with one revision retry. No `rubric` key = no-op.
- cron/schedule.py: create_schedule/run_now carry the rubric in the run
input and cron metadata only when non-blank
- middleware/scheduler.py: schedule_task gains `rubric`; list marks graded rows
- commands/implementation/schedule.py: `/schedule add ... --rubric`,
`/schedule run` forwards the stored rubric, list gets a Rubric column
- subagents/_factory.py: RubricMiddleware mounted last on the scheduler
graph so a needs_revision jump skips the memory lifecycle until the
accepted run; grader is read-only (ls + read_file, eviction off),
bounded by a 12-call budget, and gets an explicit structured-output
strategy on OpenRouter (Gemini JSON mode, Anthropic tool calling);
warns at build for anthropic/claude-fable-5.1 via OpenRouter, which
grades under neither strategy today
- tests: 17 new cases; fix two pre-existing fixture leaks (callable
backend stub in test_hitl, import-under-patch in test_async_subagent_factory)
- Gateway internal identity: when a service token is configured, reject
wrong/missing tokens even from loopback (closes SSRF/local bypass).
- Terminal metering: classified AgentControlError propagates without
retry; exhausted retries raise BILLING_UNAVAILABLE instead of a
generic RuntimeError, keeping error attribution accurate.
* feat(llm): add GLM-5.3-Flash, Qwen3.8-Flash, and Tencent HY4 preview
GLM-5.3-Flash on Zhipu, Zhipu Coding Plan and OpenRouter; Qwen3.8-Flash on
DashScope, DashScope Coding Plan and OpenRouter; HY4 preview on OpenRouter.
All three get explicit 1M-class context-window entries so they are not
caught by the glm-5 family fallback (203K) or the 200K default.
* chore: update wechat_group image asset
* chore: update version to v0.2.9
uv.lock moves to deepagents 0.7.11 with its raised dependency floors
(langchain 1.3.18, langchain-core 1.6.1, langchain-anthropic 1.7.0,
langchain-google-genai 4.3.7, langsmith 0.11.2); langgraph stays 1.2.11.
* docs: use umbrella provider names in v0.2.9 changelog
* fix: bound cache-eviction race test by deadline to fit Windows CI timeout
* test: start cache-eviction race workers from a shared barrier and assert progress
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: defer eager observation-index build to first model call
* fix: add mtime-keyed cache to list_observation_documents to avoid re-parsing unchanged files
* fix: return a copy of the cached document list and strengthen the deletion test
* fix: copy cached document list on read and write to prevent caller mutations
* fix: split global and project cache to avoid duplicate parsing and cross-project invalidation
* fix: bump mtime explicitly in cache modification test for Windows NTFS resolution
* fix: bound project observation cache with LRU eviction
* fix: deduplicate path logic and strengthen cache typing
* fix: group cache tests under TestObservationCache with autouse fixture and fix f-string interpolation
* fix: reject non-positive observation cache cap in config validation
* fix: cache resolved observation docs and config cap to avoid repeated work
* fix: use st_mtime_ns and st_size in cache signature for NTFS reliability
* test: clear EVOSCIENTIST_MAX_CACHED_PROJECTS in test env cleanup fixtures
* test: cover per-file observation cache semantics
* fix: replace layered observation caches with per-file parse cache
* fix: serialize observation parse cache transactions
Standardize the stream delta on LangChain's own message dict (delta.message)
instead of a bespoke {text, tool_calls} shape, and raise when the stream ends
without the [DONE] sentinel so mid-stream truncation is no longer silent.
The legacy BaseChatModel.astream path does not forward run_manager to
_astream, so the per-call tracing run_id was unreachable and the stream
fell back to the conversation run_id, which the gateway's
verify_model_attempt rejected (401 RUN_ATTEMPT_NOT_ACCEPTED). Publish the
tracing run_id into the shared configurable dict from
on_chat_model_start and read it back in _attempt_id.
* docs: add subscription OAuth recipe (Claude + ChatGPT/Codex)
Standalone guide for running EvoScientist on Claude Pro/Max and ChatGPT
Plus/Pro subscriptions via ccproxy OAuth, previously only partially
covered inside the macOS deployment recipe.
Documents the two Codex-route pitfalls from #323 with their exact error
strings — ccproxy's default model mappings silently rewriting gpt-* to
gpt-5.3-codex, and the backend's client-identity gate ('requires a newer
version of Codex') — plus the manual ccproxy TOML fix for self-managed
instances, account-tier model availability, verification probes, and a
troubleshooting table. Adds the recipe to the docs index.
* docs: correct subscription OAuth guidance
* docs: remove stale Codex client version guidance
* docs: show the config section to add instead of a clobbering heredoc
* docs: correct reasoning_effort default behavior on the Codex route
* docs: drop private helper import from the Codex probe
* docs: warn that local ccproxy config files shadow the global one
---------
Co-authored-by: Dinos Papakostas <dinospk1999@gmail.com>
* Add Novita as an LLM provider
Registers Novita (novita.ai) as an OpenAI-routed provider, following the
same pattern as Requesty/Atlas Cloud/SiliconFlow: a base_url + API key env
var entry in _OPENAI_ROUTED_PROVIDERS, a handful of model registry entries
(DeepSeek/Qwen/GLM), onboarding wizard support (constants/steps/wizard/
helpers), a key validator using the auth-preflight sentinel pattern (Novita's
/v1/models endpoint returns the public catalog even for an invalid key, so
auth must be checked via a chat completion instead), and a host-to-provider
mapping entry for error attribution.
* Recommend Novita's current flagship models
The models listed for Novita were older ids that no longer reflect what
the platform leads with. Point the recommendations at the three current
flagships instead, each verified against api.novita.ai:
moonshotai/kimi-k3 1M context, native vision
zai-org/glm-5.2 1M context, long-horizon agentic work
deepseek/deepseek-v4-flash-0731 1M context, cheapest of the three
Context windows, output limits, input modalities and pricing were taken
from the live /openai/v1/models response rather than carried over.
* Keep branch CI workflow files unchanged (no workflow OAuth scope)
Co-authored-by: multica-agent <github@multica.ai>
* ci: restore workflow files to match main
---------
Co-authored-by: jax-novita <jax-novita@users.noreply.github.com>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Dinos Papakostas <dinospk1999@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Add a "BioNeMo Skills" entry to the onboard Step 7 checkbox so users can
opt into the 31 life-science skills (protein folding, docking, generative
chemistry, genomics, protein design) with one selection.
The pack installs through the existing GitHub shorthand path — no installer
change. Source points at the toolkit's plugin skills directory
(plugins/bionemo-agent-toolkit/skills), which is the flat aggregate of all
31 skills; installing from the repo root would only reach 14 because the
remaining skills sit three levels deep.
Closes#333
* fix: prevent session-emptying crash on resume commands
* docs: fix stale docstrings in goto=None crash tests and patch
- Correct checkpoint corruption claim: the error state replaces the
previous conversation state (messages: [], files: {}), it IS corrupted.
- Replace WebUI-specific language with UI-agnostic wording.
- Remove references to uncommitted local notes files.
- Add upstream issue reference (langchain-ai/langgraph#5656).
* fix: wrap _control_branch instead of reimplementing, use dataclasses.replace, add END routing and session preservation tests
* fix: rebind map_cmd on already-loaded consumers, rewrite crash-path tests to exercise __start__ via checkpoint deletion
Closes#392 (uncontroversial part).
WeChat (`_handle_message`) and Feishu (`_handle_event`) gated their
signature/decryption checks behind a condition the REQUEST controls:
- WeChat: `if encrypt and self._crypto:` -- a POST with no `<Encrypt>`
element took the false branch and reached `_safe_process_message`
without any verification, even when `encoding_aes_key` + `token` were
configured.
- Feishu: `if self.config.encrypt_key and "encrypt" in body:` -- a
plaintext body skipped decryption entirely and was processed directly.
Since the webhook port is the channel's only inbound boundary, an
attacker could POST forged plaintext and reach the agent, spoofing
`sender_id` / `FromUserName` (and, with an empty allowlist, passing the
sender gate).
Fix: when encryption is configured, an inbound POST MUST carry the
encrypted field (`<Encrypt>` / `encrypt`) -- otherwise it is rejected
with 403 and never reaches the agent. Plaintext mode (no encryption
configured) is unchanged, so existing plaintext deployments are not
affected. The remaining fail-closed question (what to do when
credentials are entirely unset) is left for the maintainers to decide
as the policy part of the issue.
Regression tests (9 new):
- WeChat: plaintext rejected / missing Encrypt rejected / bad signature
rejected / valid signature decrypts and processes / plaintext still
accepted when no crypto.
- Feishu: plaintext rejected / non-dict body rejected / encrypted body
decrypts and processes / plaintext still accepted when no encrypt_key.
93 tests in the two channel files pass; full suite 3045 passed, 13
skipped; ruff clean.
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix(mcp): give stdio subprocess a real stderr fd under redirected streams (#418)
On Windows the Textual TUI redirects sys.stderr to an in-memory capture
(textual.app._PrintCapture) whose fileno() returns -1. The MCP SDK forwards
that stderr to stdio server subprocesses via subprocess.Popen(stderr=...),
and Popen rejects the invalid handle with OSError: [Errno 9] Bad file
descriptor — so only stdio servers fail to load (HTTP/SSE are unaffected).
Wrap mcp.client.stdio.stdio_client so that, whenever the configured errlog
has no usable fileno, it falls back to sys.__stderr__ (or os.devnull in GUI
hosts). Idempotent, no-op when the SDK is absent, warns if the SDK renames
stdio_client. Adds 9 regression tests and a troubleshooting note.
* fix(mcp): validate live fd and close fallback errlog after stdio session
Address CodeRabbit review on #423:
- _stdio_errlog_is_usable now os.fstat()s the fd to reject closed streams
that still report their former positive fileno (prevents a deferred
[Errno 9] from subprocess.Popen).
- The stdio_client wrapper owns the devnull fallback it allocates and
closes it once the session exits, so repeated MCP reloads no longer leak
file descriptors. Caller-provided usable errlogs pass through untouched.
- Tests cover the closed-fd case, the fd-leak/closure invariant, and
confirm langchain-mcp-adapters binds the patched stdio_client.
* fix(mcp): rebind adapter stdio_client, forward errlog by kw, harden tests
Address CodeRabbit round-2 review on #423:
- The patch now also rebinds langchain_mcp_adapters.sessions.stdio_client,
which the adapter captures via a 'from' import at module load — so the
wrapped function reaches the adapter regardless of import order.
- errlog is forwarded to the SDK by keyword (original(server, *args,
errlog=errlog, **kwargs)) so a future SDK inserting a positional
parameter before errlog can't mis-bind the fallback.
- The fallback stream is now allocated inside the async context manager,
so it is closed on session exit even if the CM is constructed but never
entered (narrower fd-leak path).
- test_closed_fd_rejected now reaches the os.fstat branch (stale positive
fd stub) instead of the ValueError path; test_adapter_binds_patched_stdio_client
documents and asserts the import-order-independent rebind.
* fix(mcp): close fallback errlog when stdio_client construction fails
Address CodeRabbit round-3 review on #423: move the original(server, *args,
errlog=errlog, **kwargs) construction inside the try block so a failure
during subprocess/client setup still reaches the finally and closes the
wrapper-owned os.devnull stream. Added test_fallback_closed_when_construction_fails
covering the path.
* refactor(mcp): track fallback ownership via (stream, opened_by_us)
Address din0s review on #423:
- _safe_stdio_errlog() now returns (stream, opened_by_us); the wrapper closes
the fallback only when opened_by_us is True, instead of inferring ownership
from needs_fallback + an identity check against sys.__stderr__. Simpler and
less likely to regress.
- Removed dead try/finally in test_closed_fd_rejected.
- Added test_wrapped_stdio_client_swaps_explicit_bad_errlog covering the
'not _stdio_errlog_is_usable(errlog)' branch (explicit bad errlog, not the
default sentinel).
- Updated test_safe_errlog_returns_usable_stream for the tuple return.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat: add thread metadata index for improved performance in thread listing
- Implemented a new SQLite index on the `checkpoints` table to optimize thread listing queries by indexing relevant metadata fields.
- Updated the `list_threads` function to ensure the index is created if it does not exist.
- Added a test to verify the creation of the metadata index during thread listing.
feat: enhance workspace sidecar management with owner tracking
- Modified the workspace sidecar to include `owner_pids` to track the current process owners.
- Updated tests to validate the new owner tracking functionality and ensure proper behavior when managing workspace sidecars.
chore: introduce model registry for streamlined model management
- Created a new `registry.py` file to maintain a comprehensive model registry, including model names, IDs, providers, and routing tables.
- Added functions to retrieve models by provider and list available models, enhancing the modularity and maintainability of model management.
* feat: enhance workspace sidecar management and improve thread metadata indexing
* fix(tests): ensure sidecar correctly registers owner with original workspace and pid
* refactor: simplify workspace sidecar management by removing owner tracking
* feat(server): add commands to manage background langgraph dev server
- Introduced `server_app` for managing the langgraph dev server with commands to check status and stop the server.
- Enhanced workspace sidecar management to include configuration fingerprint for drift detection.
- Updated deployment functions to handle server configuration and state more effectively.
* feat(server): enhance server status command to display PID with stale record warning
* feat(langgraph_dev): exclusion-set config fingerprint, webui keepalive, unified stop guidance
* fix(cli): platform-specific manual-stop hint; document keepalive endpoint-change limitation