Post-merge validation fixes (upstream v0.3.0 + Ai4Sci fork):
- llm/patches.py: restore the two module-level patch calls the merge dropped
(_patch_openai_empty_sse_keepalive, _patch_deepagents_extracted_document_text)
and make _is_ccproxy_codex accept an explicit base_url/api_key so the
invocation plan can classify an endpoint without mutating the process env.
- llm/models.py: an explicit per-call plan now wins over
EVOSCIENTIST_USE_RESPONSES_API (env is only a default), an explicit caller
`reasoning` block survives an explicit use_responses_api=False, and the
third-party (openrouter) default effort stays the fork's fixed `medium`.
- EvoScientist.py: sub-agent stacks pass NO_OP_SINK as `events` instead of None.
- middleware/error_normalization.py: platform-generated diagnostics
(ModelOutputTruncatedError) keep their actionable text while provider SDK
errors still get the canned redacted message.
- pyproject.toml: hold google-genai 1.x (langchain-google-genai>=4.3.7,<4.4)
because llm/gemini_interactions.py drives the 1.x Interactions API; this is
also what deepagents 0.7.13 requires.
- config/settings.py: restore upstream's use_responses_api config field.
`reasoning_effort` stays deleted on purpose — Ai4Sci keeps reasoning an
invocation-plan parameter, never a deployment-env override.
- tests: align upstream tests that encode replaced behaviour (ccproxy
responses-api context, reasoning-effort-overrides-env, fingerprint coverage)
with the fork's contracts.
OpenRouter keys app pages by HTTP-Referer; X-Title only renames that
page. A custom openrouter_app_title on the default referer therefore
renamed the shared EvoScientist app page for everyone. Force the default
title whenever the resolved referer is the default, silently, so usage
keeps being attributed to EvoScientist; a private fork still overrides
both together.
- Gateway internal identity: when a service token is configured, reject
wrong/missing tokens even from loopback (closes SSRF/local bypass).
- Terminal metering: classified AgentControlError propagates without
retry; exhausted retries raise BILLING_UNAVAILABLE instead of a
generic RuntimeError, keeping error attribution accurate.
* feat(llm): add GLM-5.3-Flash, Qwen3.8-Flash, and Tencent HY4 preview
GLM-5.3-Flash on Zhipu, Zhipu Coding Plan and OpenRouter; Qwen3.8-Flash on
DashScope, DashScope Coding Plan and OpenRouter; HY4 preview on OpenRouter.
All three get explicit 1M-class context-window entries so they are not
caught by the glm-5 family fallback (203K) or the 200K default.
* chore: update wechat_group image asset
* chore: update version to v0.2.9
uv.lock moves to deepagents 0.7.11 with its raised dependency floors
(langchain 1.3.18, langchain-core 1.6.1, langchain-anthropic 1.7.0,
langchain-google-genai 4.3.7, langsmith 0.11.2); langgraph stays 1.2.11.
* docs: use umbrella provider names in v0.2.9 changelog
Standardize the stream delta on LangChain's own message dict (delta.message)
instead of a bespoke {text, tool_calls} shape, and raise when the stream ends
without the [DONE] sentinel so mid-stream truncation is no longer silent.
The legacy BaseChatModel.astream path does not forward run_manager to
_astream, so the per-call tracing run_id was unreachable and the stream
fell back to the conversation run_id, which the gateway's
verify_model_attempt rejected (401 RUN_ATTEMPT_NOT_ACCEPTED). Publish the
tracing run_id into the shared configurable dict from
on_chat_model_start and read it back in _attempt_id.
* Add Novita as an LLM provider
Registers Novita (novita.ai) as an OpenAI-routed provider, following the
same pattern as Requesty/Atlas Cloud/SiliconFlow: a base_url + API key env
var entry in _OPENAI_ROUTED_PROVIDERS, a handful of model registry entries
(DeepSeek/Qwen/GLM), onboarding wizard support (constants/steps/wizard/
helpers), a key validator using the auth-preflight sentinel pattern (Novita's
/v1/models endpoint returns the public catalog even for an invalid key, so
auth must be checked via a chat completion instead), and a host-to-provider
mapping entry for error attribution.
* Recommend Novita's current flagship models
The models listed for Novita were older ids that no longer reflect what
the platform leads with. Point the recommendations at the three current
flagships instead, each verified against api.novita.ai:
moonshotai/kimi-k3 1M context, native vision
zai-org/glm-5.2 1M context, long-horizon agentic work
deepseek/deepseek-v4-flash-0731 1M context, cheapest of the three
Context windows, output limits, input modalities and pricing were taken
from the live /openai/v1/models response rather than carried over.
* Keep branch CI workflow files unchanged (no workflow OAuth scope)
Co-authored-by: multica-agent <github@multica.ai>
* ci: restore workflow files to match main
---------
Co-authored-by: jax-novita <jax-novita@users.noreply.github.com>
Co-authored-by: multica-agent <github@multica.ai>
Co-authored-by: Dinos Papakostas <dinospk1999@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: prevent session-emptying crash on resume commands
* docs: fix stale docstrings in goto=None crash tests and patch
- Correct checkpoint corruption claim: the error state replaces the
previous conversation state (messages: [], files: {}), it IS corrupted.
- Replace WebUI-specific language with UI-agnostic wording.
- Remove references to uncommitted local notes files.
- Add upstream issue reference (langchain-ai/langgraph#5656).
* fix: wrap _control_branch instead of reimplementing, use dataclasses.replace, add END routing and session preservation tests
* fix: rebind map_cmd on already-loaded consumers, rewrite crash-path tests to exercise __start__ via checkpoint deletion
* feat: add thread metadata index for improved performance in thread listing
- Implemented a new SQLite index on the `checkpoints` table to optimize thread listing queries by indexing relevant metadata fields.
- Updated the `list_threads` function to ensure the index is created if it does not exist.
- Added a test to verify the creation of the metadata index during thread listing.
feat: enhance workspace sidecar management with owner tracking
- Modified the workspace sidecar to include `owner_pids` to track the current process owners.
- Updated tests to validate the new owner tracking functionality and ensure proper behavior when managing workspace sidecars.
chore: introduce model registry for streamlined model management
- Created a new `registry.py` file to maintain a comprehensive model registry, including model names, IDs, providers, and routing tables.
- Added functions to retrieve models by provider and list available models, enhancing the modularity and maintainability of model management.
* feat: enhance workspace sidecar management and improve thread metadata indexing
* fix(tests): ensure sidecar correctly registers owner with original workspace and pid
* refactor: simplify workspace sidecar management by removing owner tracking
* feat(server): add commands to manage background langgraph dev server
- Introduced `server_app` for managing the langgraph dev server with commands to check status and stop the server.
- Enhanced workspace sidecar management to include configuration fingerprint for drift detection.
- Updated deployment functions to handle server configuration and state more effectively.
* feat(server): enhance server status command to display PID with stale record warning
* feat(langgraph_dev): exclusion-set config fingerprint, webui keepalive, unified stop guidance
* fix(cli): platform-specific manual-stop hint; document keepalive endpoint-change limitation
* Add Requesty as an LLM provider
* Address review: Requesty prompt caching, model ordering, key validation
- Declare Anthropic-style prompt caching for Requesty Claude models by
default (mirroring the OpenRouter behavior), with an opt-out flag
EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
overrides native/OpenRouter models for names it shares with them
(the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
invalid/missing key (public catalog), so it cannot validate a key. Use a
minimal authenticated /v1/chat/completions request instead (200 = valid,
403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).
* Validate Requesty key against auth layer, not a specific model
The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.
---------
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* fix: add support for new Anthropic models and enhance adaptive thinking tests
* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models
* fix: update version to v0.2.4 in badges, README, and project files
* fix: update Star History chart links in README and README.zh-CN
* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog
* fix: update wechat group image in assets
* fix: set langgraph and codex proxy runtime defaults
* fix: address runtime default review feedback
* fix: drop langgraph dev env defaults per maintainer review
langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment,
so the reported KeyError cannot come from this flow; the env defaults added
here were unnecessary. Scope the PR back to the codex proxy reasoning
context fix only.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat(context-window): add Kimi K3 model with 1M context window
* feat(openrouter): implement structured output for Kimi K3 and add 429 retry handling
* Refactor code structure for improved readability and maintainability
* fix: surface real exception class+message in SSE error events
* fix: tighten SSE error patch scope and key redaction
* fix: redact base64-style secret suffixes fully
* style: remove notes/ reference from the dosctring
* fix: rebuild env cache on each error call
* fix: route BaseException through serde.default on SSE/webhook paths
* fix: distinguish routed providers by request URL host
* feat: normalize provider-SDK exceptions via ErrorNormalizationMiddleware
* refactor: drop json_dumpb dataclass-bypass wrappers, superseded by middleware
* fix: guard _extract_host against SDK properties that raise
* refactor: derive provider tag from ModelRequest.model, not the exception
* refactor: drop serde.default patch and exception-based inference; ProviderStreamError.model_dump handles the emit
* refactor: move envelope helpers from patches.py to errors.py
* feat: extend ErrorNormalizationMiddleware coverage to every model-call path
* chore: clean up review findings from middleware pivot
* fix: pass through all langgraph.errors
* fix: move langgraph.errors pass-through into _normalize
* fix: pass through ContextOverflowError in _normalize
* fix(llm): make gpt-5.x usable through ccproxy Codex OAuth
Two independent blockers made current OpenAI models fail when routed
through ccproxy's Codex OAuth endpoint:
1. ccproxy's default Codex model mappings rewrite any gpt-*/o1-*/o3-*/
claude-* model to gpt-5.3-codex before forwarding, silently overriding
the configured model and failing outright on accounts where
gpt-5.3-codex is not served ("The 'gpt-5.3-codex' model is not
supported when using Codex with a ChatGPT account").
start_ccproxy() now generates a config with empty codex model
mappings and passes it via 'ccproxy serve --config'.
2. ccproxy forwards the client's own User-Agent upstream and only
gap-fills its Codex headers, so the backend gates current models on
the client identity ("The '<model>' model requires a newer version
of Codex"). get_chat_model() now sends Codex-CLI-shaped
originator/version/User-Agent headers when the ccproxy Codex adapter
is detected, overridable via EVOSCIENTIST_CODEX_CLIENT_VERSION.
Verified live: gpt-5.5 and gpt-5.4 complete successfully through
ccproxy Codex OAuth on a ChatGPT Plus account with both fixes; each
fails without them.
* fix(ccproxy): harden Codex client routing
* fix(llm): keep Codex client identity consistent
* docs: clarify Codex version floor
* style: ruff format models.py after merge
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* fix(llm): respect reasoning_effort setting on native OpenAI path
The native OpenAI provider path hardcoded reasoning effort to xhigh for
gpt-5.4/5.5/codex models, silently ignoring the user's reasoning_effort
config setting. The OpenRouter path already honors the
EVOSCIENTIST_REASONING_EFFORT env var that settings.py exports from that
setting; this applies the same lookup on the native path, falling back
to the previous defaults when unset.
Adds a regression test and isolates the existing xhigh test from the
env var.
* fix(llm): preserve model reasoning defaults
* fix(llm): preserve GPT-5.6 reasoning default
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat(llm): add OpenRouter app attribution headers (#339)
Attach EvoScientist app-attribution at the shared model-init layer so all
OpenRouter calls are credited to the project. langchain-openrouter maps
app_url/app_title/app_categories -> HTTP-Referer / X-Title /
X-OpenRouter-Categories. Applied only for the openrouter provider, via
setdefault so explicit caller kwargs win. Configurable through new
openrouter_http_referer / openrouter_app_title / openrouter_app_categories
settings and their EVOSCIENTIST_OPENROUTER_* env vars.
Closes#339
* refactor(llm): centralize OpenRouter attribution defaults + cap categories
Address PR #344 review:
- Define the app-attribution default constants once in config/settings.py
(the config fields and llm/models.py both use them) instead of duplicating
the literals across the two modules.
- Reduce the default categories to creative-writing,personal-agent and cap the
sent list to OpenRouter's 2-per-request limit, warning when a configured list
exceeds it, so extras are dropped predictably (and surfaced) here rather than
being silently truncated server-side.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat(middleware): reposition code interpreter middleware in the stack
* feat(models): add qwen3.7-plus model entry and update context window comment
* feat(models): add qwen3.7-max and qwen3.7-plus model entries for DashScope
* feat(auxiliary): implement auxiliary model support for background tasks and tool selection
- Added auxiliary model configuration to EvoScientistConfig.
- Introduced _ensure_auxiliary_chat_model function to manage auxiliary model instances.
- Updated onboarding steps to include auxiliary model selection.
- Modified middleware to route tool selection to the auxiliary model when applicable.
- Enhanced tests to cover auxiliary model functionality and configuration.
* feat(steps): update UI backend selection options and descriptions
* Refactor code structure for improved readability and maintainability
* feat(patches): implement OpenRouter response reasoning item stripping to prevent multi-turn errors
* feat: update version to v0.1.4 in badges, README, and pyproject.toml; adjust skill counts in steps.py
* feat(config): add auxiliary model and provider environment variables to test setup
* Enhance multimodal handling in LLM model
- Updated `_flatten_message_content` to preserve media blocks (images, files) while flattening text content.
- Introduced `_sanitize_messages` to manage media hoisting for tool messages, ensuring compatibility with OpenAI APIs.
- Modified `_patch_openai_compat_content` to accommodate new media handling logic, including retry mechanisms for media errors.
- Added comprehensive tests for media preservation, including various scenarios with images, files, and unsupported media types.
* fix: preserve order of text and media blocks in message flattening
* test: add tests for _strip_media_types to ensure position preservation and deduplication
* feat(llm): add dashscope-code provider for Alibaba Coding Plan keys
Alibaba Cloud Bailian "Coding Plan" subscription keys (sk-sp-*) route
through a separate endpoint (coding.dashscope.aliyuncs.com/v1) that the
standard `dashscope` provider can't reach. Add a sibling provider entry
matching the zhipu/zhipu-code and moonshot/kimi-coding precedents, with
its own validator (the coding endpoint returns 404 on /models, so probe
via chat.completions instead).
Closes#224
* fix(llm): keep dashscope as default provider for qwen3-coder shortcut
The MODELS dict is built from _MODEL_ENTRIES via a last-write-wins dict
comprehension. The initial commit listed dashscope-code AFTER dashscope,
which silently flipped the bare `get_chat_model("qwen3-coder")` shortcut
to the coding endpoint — breaking standard sk-* keys.
Reorder to match the zhipu-code / zhipu precedent: coding endpoint first,
general endpoint last so the general endpoint wins the collision and
remains the default for the shared "qwen3-coder" short name.
* feat(middleware): add ConfigurableModelMiddleware for dynamic model resolution
- Introduced ConfigurableModelMiddleware to resolve chat models from RunnableConfig.configurable on each call.
- Updated middleware initialization to include ConfigurableModelMiddleware.
- Enhanced context editing middleware tests to verify presence of ConfigurableModelMiddleware.
- Implemented tests for ConfigurableModelMiddleware to ensure correct model overriding and caching behavior.
- Added tests for deepagents model-passthrough patch to verify configuration injection in async tasks.
* feat(async-subagent): update middleware handling to prevent deadlocks in async sub-agents
* style: Refactor code formatting for improved readability in patches and test files
* refactor: streamline middleware construction and improve async handling in ConfigurableModelMiddleware
* fix: remove unused request parameter from _read_model_override function
* refactor: improve async handling in _ClientProxy and enhance logging in ConfigurableModelMiddleware
test: add behavior test to ensure AskUserMiddleware is excluded in async subagent mode