Feat/ai4sci tool protocol #1

Open
ouyangbo wants to merge 10 commits from feat/ai4sci-tool-protocol into Ai4Sci
Owner

Description

Type of change

  • Bug fix
  • New feature — link issue: #
  • Documentation / examples
  • Test improvement
  • Refactor (no behavior change)

Checklist

  • I have read the Contributing Guidelines
  • This targets core functionality used by the majority of users (niche features belong in EvoSkills)
  • I have added/updated tests where applicable
  • uv run ruff check . passes
  • uv run pytest passes
## Description <!-- What does this PR do? Link the related issue (e.g. "Closes #123"). --> ## Type of change <!-- Check the one that applies. --> - [ ] Bug fix - [ ] New feature — link issue: #<!-- issue number --> - [ ] Documentation / examples - [ ] Test improvement - [ ] Refactor (no behavior change) ## Checklist - [ ] I have read the [Contributing Guidelines](../CONTRIBUTING.md) - [ ] This targets **core functionality** used by the majority of users (niche features belong in [EvoSkills](https://github.com/EvoScientist/EvoSkills)) - [ ] I have added/updated tests where applicable - [ ] `uv run ruff check .` passes - [ ] `uv run pytest` passes
ouyangbo added 10 commits 2026-07-19 12:08:10 +08:00
* feat(llm): add OpenRouter app attribution headers (#339)

Attach EvoScientist app-attribution at the shared model-init layer so all
OpenRouter calls are credited to the project. langchain-openrouter maps
app_url/app_title/app_categories -> HTTP-Referer / X-Title /
X-OpenRouter-Categories. Applied only for the openrouter provider, via
setdefault so explicit caller kwargs win. Configurable through new
openrouter_http_referer / openrouter_app_title / openrouter_app_categories
settings and their EVOSCIENTIST_OPENROUTER_* env vars.

Closes #339

* refactor(llm): centralize OpenRouter attribution defaults + cap categories

Address PR #344 review:
- Define the app-attribution default constants once in config/settings.py
  (the config fields and llm/models.py both use them) instead of duplicating
  the literals across the two modules.
- Reduce the default categories to creative-writing,personal-agent and cap the
  sent list to OpenRouter's 2-per-request limit, warning when a configured list
  exceeds it, so extras are dropped predictably (and surfaced) here rather than
  being silently truncated server-side.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
get_effective_config() runs load_dotenv(find_dotenv(usecwd=True),
override=True), so any test that loads config injected the repo's real
.env into os.environ for the rest of the pytest process. An
empty-valued line like MINIMAX_BASE_URL= then made
os.environ.get(key, default) return '' instead of the default,
failing the MiniMax routing tests in full-suite runs while they
passed in isolation.

Generalizes the find_dotenv redirect that test_config.py's
temp_config_dir fixture already applied locally into a suite-wide
autouse fixture, pointing at a never-created path so tests writing
their own tmp_path/.env cannot collide with it. Adds a regression
test reproducing the leak.

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix(ccproxy): raise auth status check timeout to 30s

ccproxy's CLI initializes its full plugin system on every invocation;
a cold 'ccproxy auth status' takes ~10s wall time on Apple Silicon,
so the 10s subprocess timeout made OAuth startup fail intermittently
with 'Auth check timed out' even when credentials were valid.

* fix(ccproxy): raise serve health deadline to 120s

ccproxy boot includes plugin init plus Codex CLI detection; measured
~76s to first healthy response on an Apple Silicon Mac (ccproxy-api
0.2.9). The 30s deadline in start_ccproxy() killed the process before
it could come up, failing OAuth startup with 'ccproxy did not become
healthy within 30 seconds'.

* fix(ccproxy): widen serve health deadline to 180s

Full startup measured at ~111s on a second cold run (Apple Silicon,
ccproxy-api 0.2.9); 120s left too little headroom for boot variance.

* fix(ccproxy): centralize startup timeouts

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix(llm): respect reasoning_effort setting on native OpenAI path

The native OpenAI provider path hardcoded reasoning effort to xhigh for
gpt-5.4/5.5/codex models, silently ignoring the user's reasoning_effort
config setting. The OpenRouter path already honors the
EVOSCIENTIST_REASONING_EFFORT env var that settings.py exports from that
setting; this applies the same lookup on the native path, falling back
to the previous defaults when unset.

Adds a regression test and isolates the existing xhigh test from the
env var.

* fix(llm): preserve model reasoning defaults

* fix(llm): preserve GPT-5.6 reasoning default

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix(llm): make gpt-5.x usable through ccproxy Codex OAuth

Two independent blockers made current OpenAI models fail when routed
through ccproxy's Codex OAuth endpoint:

1. ccproxy's default Codex model mappings rewrite any gpt-*/o1-*/o3-*/
   claude-* model to gpt-5.3-codex before forwarding, silently overriding
   the configured model and failing outright on accounts where
   gpt-5.3-codex is not served ("The 'gpt-5.3-codex' model is not
   supported when using Codex with a ChatGPT account").
   start_ccproxy() now generates a config with empty codex model
   mappings and passes it via 'ccproxy serve --config'.

2. ccproxy forwards the client's own User-Agent upstream and only
   gap-fills its Codex headers, so the backend gates current models on
   the client identity ("The '<model>' model requires a newer version
   of Codex"). get_chat_model() now sends Codex-CLI-shaped
   originator/version/User-Agent headers when the ccproxy Codex adapter
   is detected, overridable via EVOSCIENTIST_CODEX_CLIENT_VERSION.

Verified live: gpt-5.5 and gpt-5.4 complete successfully through
ccproxy Codex OAuth on a ChatGPT Plus account with both fixes; each
fails without them.

* fix(ccproxy): harden Codex client routing

* fix(llm): keep Codex client identity consistent

* docs: clarify Codex version floor

* style: ruff format models.py after merge

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* fix: scope quickjs snapshot to turn to keep checkpoints small

* fix: strip _quickjs_snapshot_payload from state/history responses instead of dropping mode=thread

* fix: recurse strip into nested subgraph StateSnapshot

* fix: drop conditional-snapshot gate that leaked repl slots

* fix: assert LangGraph state-shape invariants at import time

* refactor: discover graphs to filter from langgraph.json

* test: assert copy() preserves subclass; iterate langgraph.json for subagent coverage

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: surface real exception class+message in SSE error events

* fix: tighten SSE error patch scope and key redaction

* fix: redact base64-style secret suffixes fully

* style: remove notes/ reference from the dosctring

* fix: rebuild env cache on each error call

* fix: route BaseException through serde.default on SSE/webhook paths

* fix: distinguish routed providers by request URL host

* feat: normalize provider-SDK exceptions via ErrorNormalizationMiddleware

* refactor: drop json_dumpb dataclass-bypass wrappers, superseded by middleware

* fix: guard _extract_host against SDK properties that raise

* refactor: derive provider tag from ModelRequest.model, not the exception

* refactor: drop serde.default patch and exception-based inference; ProviderStreamError.model_dump handles the emit

* refactor: move envelope helpers from patches.py to errors.py

* feat: extend ErrorNormalizationMiddleware coverage to every model-call path

* chore: clean up review findings from middleware pivot

* fix: pass through all langgraph.errors

* fix: move langgraph.errors pass-through into _normalize

* fix: pass through ContextOverflowError in _normalize
fix: harden tool-call protocol and fallback handling
Test / pytest (ubuntu-latest, 3.11) (pull_request) Has been cancelled
Build / build (pull_request) Has been cancelled
Docker / build (pull_request) Has been cancelled
Lint / ruff (pull_request) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (pull_request) Has been cancelled
Test / pytest (windows-latest, 3.11) (pull_request) Has been cancelled
Test / pytest (windows-latest, 3.12) (pull_request) Has been cancelled
3ce5614254
Some checks are pending
Test / pytest (ubuntu-latest, 3.11) (pull_request) Has been cancelled
Build / build (pull_request) Has been cancelled
Docker / build (pull_request) Has been cancelled
Lint / ruff (pull_request) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (pull_request) Has been cancelled
Test / pytest (windows-latest, 3.11) (pull_request) Has been cancelled
Test / pytest (windows-latest, 3.12) (pull_request) Has been cancelled
This pull request has changes conflicting with the target branch.
  • EvoScientist/EvoScientist.py
  • EvoScientist/llm/models.py
  • EvoScientist/llm/patches.py
  • tests/test_agent_factory_extensions.py
  • tests/test_llm.py
View command line instructions

Checkout

From your project repository, check out a new branch and test the changes.
git fetch -u origin feat/ai4sci-tool-protocol:feat/ai4sci-tool-protocol
git checkout feat/ai4sci-tool-protocol
Sign in to join this conversation.
No Reviewers
No Label
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: ouyangbo/EvoScientist#1