The runtime (llm/runtime.py, stream/stop.py) is a separate development line
whose stop adapter drives LangGraph internals. Upstream v0.3.0 bumps its
dependencies for currency, and two exact-version assertions in that adapter
turned the bump into a silent regression: stop ownership was refused, so
cancellation/continuation runs never reached a terminal state.
Resolved without touching the adapter's logic:
- stream/stop.py: claim ownership by *capability* instead of an exact version
string. The internals the adapter swaps (_graph_aiter / _pump_cond /
_exhausted / _aborting / _anext_task / _mux) and the SQLite saver's
connection lock are present and identical in langgraph 1.2.6 and 1.2.11, and
langgraph-checkpoint-sqlite 3.1.1 exposes the same barrier as 3.0.3. A new
patch release can no longer disable stop ownership by being newer; a release
that really drops the internals still fails closed with
CHECKPOINT_STOP_ADAPTER_UNSUPPORTED.
- EvoScientist.py: supply TodoListMiddleware only when deepagents' own default
chain lacks it. deepagents 0.7 dropped it (upstream adds one back); 0.6.x
still ships it, and a second instance collides by name in
langchain's create_agent.
- backends.py: fall back to a shape-compatible DeleteResult when deepagents
has no delete support, so upstream v0.3.0's delete refusals import and run on
either line.
- tests/test_backends.py: gate the delete-behaviour tests on the framework
actually providing backend deletion instead of asserting a specific stack.
Verified: 4216 passed / 33 skipped / 29 failed / 16 errors — every remaining
failure is pre-existing on the untouched pre-merge tree except two
(a google-stream cleanup-order assertion and one webui launcher test).
Post-merge validation fixes (upstream v0.3.0 + Ai4Sci fork):
- llm/patches.py: restore the two module-level patch calls the merge dropped
(_patch_openai_empty_sse_keepalive, _patch_deepagents_extracted_document_text)
and make _is_ccproxy_codex accept an explicit base_url/api_key so the
invocation plan can classify an endpoint without mutating the process env.
- llm/models.py: an explicit per-call plan now wins over
EVOSCIENTIST_USE_RESPONSES_API (env is only a default), an explicit caller
`reasoning` block survives an explicit use_responses_api=False, and the
third-party (openrouter) default effort stays the fork's fixed `medium`.
- EvoScientist.py: sub-agent stacks pass NO_OP_SINK as `events` instead of None.
- middleware/error_normalization.py: platform-generated diagnostics
(ModelOutputTruncatedError) keep their actionable text while provider SDK
errors still get the canned redacted message.
- pyproject.toml: hold google-genai 1.x (langchain-google-genai>=4.3.7,<4.4)
because llm/gemini_interactions.py drives the 1.x Interactions API; this is
also what deepagents 0.7.13 requires.
- config/settings.py: restore upstream's use_responses_api config field.
`reasoning_effort` stays deleted on purpose — Ai4Sci keeps reasoning an
invocation-plan parameter, never a deployment-env override.
- tests: align upstream tests that encode replaced behaviour (ccproxy
responses-api context, reasoning-effort-overrides-env, fingerprint coverage)
with the fork's contracts.
* refactor: rename AGENTS.md to EXPERT.md per EvoSkills convention
* feat: resolve newly installed experts on start_async_task miss
* chore: reword the /expert invite hint to state the dispatch boundary
* docs: note the resolve-on-miss caller in build_expert_async_subagent_specs
* fix: thread the construction cfg through resolve-on-miss and symmetric setdefault
* fix: guard agent_map iteration against concurrent resolve-on-miss writes
* fix: warn once per broken expert on repeated resolve-on-miss walks
* fix: scope the /expert invite hint to newly installed experts
* docs: note live uninstalls as a resolve-on-miss limitation
* chore: isolate the warn-once collision test key from the route-specs suite
* fix: _extract_retry_after returns None for non-retryable errors
* fix: update _extract_retry_after to handle generic transient errors with default retry delay
* fix: extend base _non_retryable_patterns in channel subclasses
* fix(channels): merge structured SDK error check into non-retryable step
* fix(channels): decouple status code and SDK error code extraction in retry logic
- Independently evaluate HTTP status codes and structured SDK error codes
- Fix misleading doc comments for Feishu and DingTalk patterns
- Remove redundant try-except AttributeError on getattr with default
- Expand test coverage for dual-signal matrix and header parsing
* refactor(channels): simplify status code and SDK error extraction via channel overrides
- Handle httpx and aiohttp exceptions in base Channel class
- Override _extract_status_code and _extract_sdk_error_code in SlackChannel and DiscordChannel
- Replace mock exception types in comprehensive test suite with real httpx and aiohttp errors
- Add dedicated Slack and Discord retry error extraction test suites
* fix(channels): clean up Slack and Discord error code extraction
- Remove defensive string checks and attribute guards in SlackChannel
- Directly access exc.response.status_code and exc.response.get('error') in SlackChannel
- Remove unnecessary _extract_sdk_error_code override in DiscordChannel
- Use real SlackApiError, SlackResponse, and discord.HTTPException in unit tests
* refactor: reorder retry logic to prioritize non-retryable checks, remove aiohttp dependency, and clean up exception handling in base and channel modules.
* test(channels): skip Slack/Discord retry tests when the SDK extra is absent
The retry-extraction tests build real SlackApiError / discord.HTTPException
objects, but slack-sdk and discord.py are optional extras that the dev
dependency group does not install. Under CI's `uv sync --dev` all nine
tests failed with ModuleNotFoundError raised from the channel override.
Gate both test classes with skipif(find_spec(...) is None) so the suite is
green without the extras and the tests still run wherever they are installed.
* refactor(channels): replace retry-delay lookup with _extract_retry_delay
_extract_retry_after still read the server-supplied delay by probing
exc.retry_after and exc.response.headers via getattr/hasattr, the last
remnant of the pattern the extractors moved away from. Replace both steps
with one overridable hook, _extract_retry_delay, implemented against the
real exception types:
- base: httpx.HTTPStatusError -> Retry-After header (httpx.Headers is
case-insensitive; HTTP-date form remains unsupported)
- SlackChannel: SlackApiError -> Retry-After, matched case-insensitively
because SlackResponse.headers is a plain dict whose casing depends on the
HTTP client (same approach as slack_sdk's RateLimitErrorRetryHandler)
- TelegramChannel: telegram.error.RetryAfter.retry_after (int, or timedelta
under PTB_TIMEDELTA)
- DiscordChannel: discord.RateLimited.retry_after, which the old duck-typed
getattr matched and would otherwise have been lost
Drop the isinstance(retry, bool) and val >= 0 guards; no SDK produces those.
Delete the test that asserted the duck-typed attribute; add real-object tests
for each override, guarded like the existing SDK-dependent classes.
* ci: install the all-channels extra so SDK-dependent channel tests run
The Slack, Discord, and Telegram retry tests build real SDK exception
objects and are skipped when the SDK is absent. CI only ran `uv sync --dev`,
so those tests never executed there. Install the existing all-channels
extra alongside the dev group; the skipif guards remain for lean local runs.
* fix(channels): honor HTTP-date Retry-After and tolerate malformed values
RFC 9110 allows Retry-After as either delay-seconds or an HTTP-date. The
httpx path treated a date as unparseable and fell back to the 1.0 s default,
so a 503 asking for a specific wait was retried too early. Add
Channel._parse_retry_after, which returns delay-seconds as-is and converts
an HTTP-date to the non-negative seconds until it (tz-less dates read as
UTC).
SlackChannel used a bare float() on the header. A non-numeric value raised
inside the retry predicate, which escapes retry_async and drops the chunk
instead of retrying. Route Slack through the same helper so a bad header
falls back to _rate_limit_delay.
Addresses CodeRabbit review comments on base.py:866 and slack/channel.py:229.
* fix(channels): treat HTTP 400 and 404 as non-retryable
Both are permanent for a given request, so retrying burns the attempt
budget for nothing. Add them to _non_retryable_status_codes alongside
401/403.
Deliberately not a 4xx range check: 408 and 425 are retryable by
definition and 429 is handled by the rate-limit path. A test pins 408 as
still retryable so the range shortcut is not reintroduced later.
Partially addresses CodeRabbit's outside-diff comment on base.py:749-750.
* fix(channels): guard Slack retry extractors against raw aiohttp responses
slack_sdk attaches the bare aiohttp.ClientResponse to SlackApiError when a
JSON-declared body fails to parse. That object has neither status_code nor
get(), so _extract_status_code raised AttributeError inside should_retry,
replacing the original error and skipping the remaining attempts. Narrow
both extractors to SlackResponse/AsyncSlackResponse so such errors fall
through to the message patterns and retry as before. Add a wire-level
regression test against a local aiohttp server.
---------
Co-authored-by: Dinos Papakostas <dinospk1999@gmail.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
OpenRouter keys app pages by HTTP-Referer; X-Title only renames that
page. A custom openrouter_app_title on the default referer therefore
renamed the shared EvoScientist app page for everyone. Force the default
title whenever the resolved referer is the default, silently, so usage
keeps being attributed to EvoScientist; a private fork still overrides
both together.
Scheduled tasks gain an optional `rubric`: an acceptance checklist graded
after each run by deepagents' RubricMiddleware (LLM-as-a-judge on the
auxiliary model) with one revision retry. No `rubric` key = no-op.
- cron/schedule.py: create_schedule/run_now carry the rubric in the run
input and cron metadata only when non-blank
- middleware/scheduler.py: schedule_task gains `rubric`; list marks graded rows
- commands/implementation/schedule.py: `/schedule add ... --rubric`,
`/schedule run` forwards the stored rubric, list gets a Rubric column
- subagents/_factory.py: RubricMiddleware mounted last on the scheduler
graph so a needs_revision jump skips the memory lifecycle until the
accepted run; grader is read-only (ls + read_file, eviction off),
bounded by a 12-call budget, and gets an explicit structured-output
strategy on OpenRouter (Gemini JSON mode, Anthropic tool calling);
warns at build for anthropic/claude-fable-5.1 via OpenRouter, which
grades under neither strategy today
- tests: 17 new cases; fix two pre-existing fixture leaks (callable
backend stub in test_hitl, import-under-patch in test_async_subagent_factory)
- Gateway internal identity: when a service token is configured, reject
wrong/missing tokens even from loopback (closes SSRF/local bypass).
- Terminal metering: classified AgentControlError propagates without
retry; exhausted retries raise BILLING_UNAVAILABLE instead of a
generic RuntimeError, keeping error attribution accurate.
* feat(llm): add GLM-5.3-Flash, Qwen3.8-Flash, and Tencent HY4 preview
GLM-5.3-Flash on Zhipu, Zhipu Coding Plan and OpenRouter; Qwen3.8-Flash on
DashScope, DashScope Coding Plan and OpenRouter; HY4 preview on OpenRouter.
All three get explicit 1M-class context-window entries so they are not
caught by the glm-5 family fallback (203K) or the 200K default.
* chore: update wechat_group image asset
* chore: update version to v0.2.9
uv.lock moves to deepagents 0.7.11 with its raised dependency floors
(langchain 1.3.18, langchain-core 1.6.1, langchain-anthropic 1.7.0,
langchain-google-genai 4.3.7, langsmith 0.11.2); langgraph stays 1.2.11.
* docs: use umbrella provider names in v0.2.9 changelog
* fix: bound cache-eviction race test by deadline to fit Windows CI timeout
* test: start cache-eviction race workers from a shared barrier and assert progress
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: defer eager observation-index build to first model call
* fix: add mtime-keyed cache to list_observation_documents to avoid re-parsing unchanged files
* fix: return a copy of the cached document list and strengthen the deletion test
* fix: copy cached document list on read and write to prevent caller mutations
* fix: split global and project cache to avoid duplicate parsing and cross-project invalidation
* fix: bump mtime explicitly in cache modification test for Windows NTFS resolution
* fix: bound project observation cache with LRU eviction
* fix: deduplicate path logic and strengthen cache typing
* fix: group cache tests under TestObservationCache with autouse fixture and fix f-string interpolation
* fix: reject non-positive observation cache cap in config validation
* fix: cache resolved observation docs and config cap to avoid repeated work
* fix: use st_mtime_ns and st_size in cache signature for NTFS reliability
* test: clear EVOSCIENTIST_MAX_CACHED_PROJECTS in test env cleanup fixtures
* test: cover per-file observation cache semantics
* fix: replace layered observation caches with per-file parse cache
* fix: serialize observation parse cache transactions
Standardize the stream delta on LangChain's own message dict (delta.message)
instead of a bespoke {text, tool_calls} shape, and raise when the stream ends
without the [DONE] sentinel so mid-stream truncation is no longer silent.
The legacy BaseChatModel.astream path does not forward run_manager to
_astream, so the per-call tracing run_id was unreachable and the stream
fell back to the conversation run_id, which the gateway's
verify_model_attempt rejected (401 RUN_ATTEMPT_NOT_ACCEPTED). Publish the
tracing run_id into the shared configurable dict from
on_chat_model_start and read it back in _attempt_id.
* docs: add subscription OAuth recipe (Claude + ChatGPT/Codex)
Standalone guide for running EvoScientist on Claude Pro/Max and ChatGPT
Plus/Pro subscriptions via ccproxy OAuth, previously only partially
covered inside the macOS deployment recipe.
Documents the two Codex-route pitfalls from #323 with their exact error
strings — ccproxy's default model mappings silently rewriting gpt-* to
gpt-5.3-codex, and the backend's client-identity gate ('requires a newer
version of Codex') — plus the manual ccproxy TOML fix for self-managed
instances, account-tier model availability, verification probes, and a
troubleshooting table. Adds the recipe to the docs index.
* docs: correct subscription OAuth guidance
* docs: remove stale Codex client version guidance
* docs: show the config section to add instead of a clobbering heredoc
* docs: correct reasoning_effort default behavior on the Codex route
* docs: drop private helper import from the Codex probe
* docs: warn that local ccproxy config files shadow the global one
---------
Co-authored-by: Dinos Papakostas <dinospk1999@gmail.com>
* Add Novita as an LLM provider
Registers Novita (novita.ai) as an OpenAI-routed provider, following the
same pattern as Requesty/Atlas Cloud/SiliconFlow: a base_url + API key env
var entry in _OPENAI_ROUTED_PROVIDERS, a handful of model registry entries
(DeepSeek/Qwen/GLM), onboarding wizard support (constants/steps/wizard/
helpers), a key validator using the auth-preflight sentinel pattern (Novita's
/v1/models endpoint returns the public catalog even for an invalid key, so
auth must be checked via a chat completion instead), and a host-to-provider
mapping entry for error attribution.
* Recommend Novita's current flagship models
The models listed for Novita were older ids that no longer reflect what
the platform leads with. Point the recommendations at the three current
flagships instead, each verified against api.novita.ai:
moonshotai/kimi-k3 1M context, native vision
zai-org/glm-5.2 1M context, long-horizon agentic work
deepseek/deepseek-v4-flash-0731 1M context, cheapest of the three
Context windows, output limits, input modalities and pricing were taken
from the live /openai/v1/models response rather than carried over.
* Keep branch CI workflow files unchanged (no workflow OAuth scope)
Co-authored-by: multica-agent <github@multica.ai>
* ci: restore workflow files to match main
---------
Co-authored-by: jax-novita <jax-novita@users.noreply.github.com>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Dinos Papakostas <dinospk1999@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Add a "BioNeMo Skills" entry to the onboard Step 7 checkbox so users can
opt into the 31 life-science skills (protein folding, docking, generative
chemistry, genomics, protein design) with one selection.
The pack installs through the existing GitHub shorthand path — no installer
change. Source points at the toolkit's plugin skills directory
(plugins/bionemo-agent-toolkit/skills), which is the flat aggregate of all
31 skills; installing from the repo root would only reach 14 because the
remaining skills sit three levels deep.
Closes#333
* fix: prevent session-emptying crash on resume commands
* docs: fix stale docstrings in goto=None crash tests and patch
- Correct checkpoint corruption claim: the error state replaces the
previous conversation state (messages: [], files: {}), it IS corrupted.
- Replace WebUI-specific language with UI-agnostic wording.
- Remove references to uncommitted local notes files.
- Add upstream issue reference (langchain-ai/langgraph#5656).
* fix: wrap _control_branch instead of reimplementing, use dataclasses.replace, add END routing and session preservation tests
* fix: rebind map_cmd on already-loaded consumers, rewrite crash-path tests to exercise __start__ via checkpoint deletion
Closes#392 (uncontroversial part).
WeChat (`_handle_message`) and Feishu (`_handle_event`) gated their
signature/decryption checks behind a condition the REQUEST controls:
- WeChat: `if encrypt and self._crypto:` -- a POST with no `<Encrypt>`
element took the false branch and reached `_safe_process_message`
without any verification, even when `encoding_aes_key` + `token` were
configured.
- Feishu: `if self.config.encrypt_key and "encrypt" in body:` -- a
plaintext body skipped decryption entirely and was processed directly.
Since the webhook port is the channel's only inbound boundary, an
attacker could POST forged plaintext and reach the agent, spoofing
`sender_id` / `FromUserName` (and, with an empty allowlist, passing the
sender gate).
Fix: when encryption is configured, an inbound POST MUST carry the
encrypted field (`<Encrypt>` / `encrypt`) -- otherwise it is rejected
with 403 and never reaches the agent. Plaintext mode (no encryption
configured) is unchanged, so existing plaintext deployments are not
affected. The remaining fail-closed question (what to do when
credentials are entirely unset) is left for the maintainers to decide
as the policy part of the issue.
Regression tests (9 new):
- WeChat: plaintext rejected / missing Encrypt rejected / bad signature
rejected / valid signature decrypts and processes / plaintext still
accepted when no crypto.
- Feishu: plaintext rejected / non-dict body rejected / encrypted body
decrypts and processes / plaintext still accepted when no encrypt_key.
93 tests in the two channel files pass; full suite 3045 passed, 13
skipped; ruff clean.
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>