* fix(llm): make gpt-5.x usable through ccproxy Codex OAuth
Two independent blockers made current OpenAI models fail when routed
through ccproxy's Codex OAuth endpoint:
1. ccproxy's default Codex model mappings rewrite any gpt-*/o1-*/o3-*/
claude-* model to gpt-5.3-codex before forwarding, silently overriding
the configured model and failing outright on accounts where
gpt-5.3-codex is not served ("The 'gpt-5.3-codex' model is not
supported when using Codex with a ChatGPT account").
start_ccproxy() now generates a config with empty codex model
mappings and passes it via 'ccproxy serve --config'.
2. ccproxy forwards the client's own User-Agent upstream and only
gap-fills its Codex headers, so the backend gates current models on
the client identity ("The '<model>' model requires a newer version
of Codex"). get_chat_model() now sends Codex-CLI-shaped
originator/version/User-Agent headers when the ccproxy Codex adapter
is detected, overridable via EVOSCIENTIST_CODEX_CLIENT_VERSION.
Verified live: gpt-5.5 and gpt-5.4 complete successfully through
ccproxy Codex OAuth on a ChatGPT Plus account with both fixes; each
fails without them.
* fix(ccproxy): harden Codex client routing
* fix(llm): keep Codex client identity consistent
* docs: clarify Codex version floor
* style: ruff format models.py after merge
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* fix(llm): respect reasoning_effort setting on native OpenAI path
The native OpenAI provider path hardcoded reasoning effort to xhigh for
gpt-5.4/5.5/codex models, silently ignoring the user's reasoning_effort
config setting. The OpenRouter path already honors the
EVOSCIENTIST_REASONING_EFFORT env var that settings.py exports from that
setting; this applies the same lookup on the native path, falling back
to the previous defaults when unset.
Adds a regression test and isolates the existing xhigh test from the
env var.
* fix(llm): preserve model reasoning defaults
* fix(llm): preserve GPT-5.6 reasoning default
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat(llm): add OpenRouter app attribution headers (#339)
Attach EvoScientist app-attribution at the shared model-init layer so all
OpenRouter calls are credited to the project. langchain-openrouter maps
app_url/app_title/app_categories -> HTTP-Referer / X-Title /
X-OpenRouter-Categories. Applied only for the openrouter provider, via
setdefault so explicit caller kwargs win. Configurable through new
openrouter_http_referer / openrouter_app_title / openrouter_app_categories
settings and their EVOSCIENTIST_OPENROUTER_* env vars.
Closes#339
* refactor(llm): centralize OpenRouter attribution defaults + cap categories
Address PR #344 review:
- Define the app-attribution default constants once in config/settings.py
(the config fields and llm/models.py both use them) instead of duplicating
the literals across the two modules.
- Reduce the default categories to creative-writing,personal-agent and cap the
sent list to OpenRouter's 2-per-request limit, warning when a configured list
exceeds it, so extras are dropped predictably (and surfaced) here rather than
being silently truncated server-side.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* chore: add pytest-asyncio in auto mode
* test: migrate channel and stream tests to native async
Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.
* test: migrate command and model/middleware tests to native async
Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.
* test: migrate TUI, notifier, gateway, and session tests to native async
TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.
* test: replace direct asyncio.run() calls with native async tests
Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.
* test: drop undeclared anyio markers and delete run_async helper
The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).
* feat(middleware): reposition code interpreter middleware in the stack
* feat(models): add qwen3.7-plus model entry and update context window comment
* feat(models): add qwen3.7-max and qwen3.7-plus model entries for DashScope
* feat(auxiliary): implement auxiliary model support for background tasks and tool selection
- Added auxiliary model configuration to EvoScientistConfig.
- Introduced _ensure_auxiliary_chat_model function to manage auxiliary model instances.
- Updated onboarding steps to include auxiliary model selection.
- Modified middleware to route tool selection to the auxiliary model when applicable.
- Enhanced tests to cover auxiliary model functionality and configuration.
* feat(steps): update UI backend selection options and descriptions
* Refactor code structure for improved readability and maintainability
* feat(patches): implement OpenRouter response reasoning item stripping to prevent multi-turn errors
* feat: update version to v0.1.4 in badges, README, and pyproject.toml; adjust skill counts in steps.py
* feat(config): add auxiliary model and provider environment variables to test setup
* Enhance multimodal handling in LLM model
- Updated `_flatten_message_content` to preserve media blocks (images, files) while flattening text content.
- Introduced `_sanitize_messages` to manage media hoisting for tool messages, ensuring compatibility with OpenAI APIs.
- Modified `_patch_openai_compat_content` to accommodate new media handling logic, including retry mechanisms for media errors.
- Added comprehensive tests for media preservation, including various scenarios with images, files, and unsupported media types.
* fix: preserve order of text and media blocks in message flattening
* test: add tests for _strip_media_types to ensure position preservation and deduplication
* feat(llm): add dashscope-code provider for Alibaba Coding Plan keys
Alibaba Cloud Bailian "Coding Plan" subscription keys (sk-sp-*) route
through a separate endpoint (coding.dashscope.aliyuncs.com/v1) that the
standard `dashscope` provider can't reach. Add a sibling provider entry
matching the zhipu/zhipu-code and moonshot/kimi-coding precedents, with
its own validator (the coding endpoint returns 404 on /models, so probe
via chat.completions instead).
Closes#224
* fix(llm): keep dashscope as default provider for qwen3-coder shortcut
The MODELS dict is built from _MODEL_ENTRIES via a last-write-wins dict
comprehension. The initial commit listed dashscope-code AFTER dashscope,
which silently flipped the bare `get_chat_model("qwen3-coder")` shortcut
to the coding endpoint — breaking standard sk-* keys.
Reorder to match the zhipu-code / zhipu precedent: coding endpoint first,
general endpoint last so the general endpoint wins the collision and
remains the default for the shared "qwen3-coder" short name.
* Add status bar and compact summary widgets with context window resolution
- Implemented a shared status bar for CLI and TUI frontends, including helpers for managing session metrics and context windows.
- Created a `CompactSummaryWidget` for displaying manual summaries in a collapsible format.
- Introduced a `CompactingWidget` to indicate ongoing compacting processes.
- Added a base class `TimedStatusWidget` for widgets that require a timer.
- Developed context window resolution helpers to retrieve context window sizes from various model attributes.
- Enhanced tests for context window resolution and status bar functionalities, ensuring accurate behavior across different scenarios.
- Updated existing tests to cover new features and maintain code quality.
* refactor(Channel): simplify lambda function in _send_with_retry method
* feat: enhance context editing logic and improve error handling in StreamState
* refactor(Channel): streamline lambda function in _send_with_retry method
* feat: rename auto-approve option to auto-mode for unattended execution; update checkpoint queries to filter by agent name; improve compatibility validation logic
* feat: rename auto-approve option to auto-mode; update related logic and tests for improved unattended execution
* fix: correct formatting of console message for MCP server configuration status
* feat: add check for None summary_message in _apply_summarization_event to prevent errors
* feat: enhance _load_checkpoint_messages to validate message format and apply summarization event
* fix(ccproxy): update Responses API handling and patch system role conversion
* fix(ccproxy): streamline _agenerate method in system to developer patch
* fix(ccproxy): improve handling of None output in Codex compatibility patch
* fix(llm): patch _stream/_astream for OpenAI-compatible content flattening
_patch_openai_compat_content() only patched _generate/_agenerate but
EvoSci CLI uses streaming paths. This extends the content flattening
to _stream/_astream so strict OpenAI-compatible relays receive plain
string content during streaming calls.
Closes#142
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(test): use asyncio.run() instead of pytest-asyncio for CI compat
CI does not have pytest-asyncio installed, so async tests must use
asyncio.run() wrapper instead of @pytest.mark.asyncio decorator.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(test): use @pytest.mark.anyio for async tests (CI compat)
CI does not have pytest-asyncio. Use @pytest.mark.anyio consistent
with existing async tests in the project.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: add Moonshot and Kimi Coding Plan as LLM providers
Add two new providers for Moonshot AI:
- `moonshot`: OpenAI-compatible direct API (api.moonshot.cn/v1) with
kimi-k2.5, kimi-k2-thinking, moonshot-v1-auto/128k/32k/8k models
- `kimi-coding`: Anthropic-compatible Kimi Coding Plan endpoint
(api.kimi.com/coding/) with User-Agent header for compatibility
Both providers disable thinking to avoid multi-turn tool calling
errors caused by LangChain dropping reasoning_content from history.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* chore: update Moonshot thinking comment and add provider assertions
- Add clarifying comment for disabling thinking on all Moonshot models
- Add moonshot and kimi-coding assertions to test_entries_has_all_providers
* fix: exclude Moonshot and Kimi Coding from content patch
Tested and verified both APIs support standard list content format:
- Moonshot (OpenAI-compatible): supports list content, no patch needed
- Kimi Coding (Anthropic-compatible): supports list content, no patch needed
Only apply _patch_openai_compat_content to strict providers like DeepSeek.
* fix: set _original_provider in routed provider branches
Ensure _original_provider is set before provider is reassigned to
'openai' or 'anthropic', so the no-patch exclusion for Moonshot
and Kimi Coding works correctly.
* style: translate Moonshot comments to English
* style: translate comment to English to fix ruff lint error
* merge: resolve conflicts
* chore: revert uv.lock and translate Chinese comments to English
Revert unrelated uv.lock dependency changes and replace Chinese code
comments with English for codebase consistency per review feedback.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* style: fix ruff format for models.py
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: ypd <ypd@ypddeMac-mini.local>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: Xiaohui Yan <xhcloud@gmail.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* feat(llm): upgrade OpenAI reasoning effort from high to xhigh
The OpenAI Responses API supports "xhigh" as a reasoning effort level,
which provides deeper reasoning than "high". This is already used by
other CLI tools (e.g., OpenClaw) for OpenAI models.
Only affects the direct API key path; the ccproxy/OAuth path is
unchanged (reasoning is still skipped there).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(llm): limit xhigh reasoning to gpt-5.4+ and codex models
Only gpt-5.4 series and codex models support xhigh reasoning effort.
Older models (gpt-5, gpt-5.1, gpt-5.2, gpt-5.3) fall back to high.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Xi Zhang <zacharyzhang2022@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat: enable reasoning for OpenRouter via extra_body to prevent multi-turn errors
* feat: implement OpenRouter native reasoning support and patch langchain-openrouter bug
* feat: add OpenRouter reasoning effort configuration and update related tests
* feat: add langchain-openrouter dependency for enhanced reasoning support
* fix: correct spacing in reasoning effort choice label
* feat: implement patch for OpenRouter reasoning details to prevent Pydantic errors
* feat: add patches for OpenRouter reasoning and content handling utilities
* feat: prevent multiple patches of OpenRouter reasoning details by using a global flag
* feat: update OpenRouter reasoning patch to ensure single application with global flag
* feat: refine OpenAI responses API handling to apply only for OpenAI provider
* feat: Enhance TUI interaction by updating todo widget positioning and skipping empty tool call chunks
* feat: Update tool selector threshold and adjust logging level for selector failures
* feat: Temporarily disable timestamp toast in tool call widget for UX review
* feat: Re-enable timestamp toast in tool call widget on click
* feat: Upgrade ccproxy to version 0.2.7 and remove deprecated thinking tag handling
* feat: Enhance ccproxy compatibility and strip legacy thinking tags
* feat(config): use_responses_api (#98)
langchain-openai auto-switches to the Responses API when reasoning
params are set, which breaks OpenAI-compatible relays that only support
Chat Completions. This adds a user-facing config option to override
that behavior:
evosci config set use_responses_api false
# or EVOSCIENTIST_USE_RESPONSES_API=false
* fix: propagate use_responses_api from config file and add normalization tests
Address PR #105 review comments:
- apply_config_to_env() now sets EVOSCIENTIST_USE_RESPONSES_API so
config file values take effect (not just the env var directly)
- Add parametrized tests for case/whitespace normalization
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat: add DeepSeek as a recognized third-party provider
Register DeepSeek API (https://api.deepseek.com) with DEEPSEEK_API_KEY
env var and add model short names: deepseek-r1 → deepseek-reasoner,
deepseek-v3 → deepseek-chat.
* feat: add _flatten_message_content utility for list-to-string conversion
Extract text from content block lists while skipping thinking/reasoning
blocks. This handles the case where LangChain stores assistant messages
with content as a list of content blocks instead of a plain string.
* fix: flatten list content to strings for OpenAI-compatible providers
Add _patch_openai_compat_content() that wraps _generate/_agenerate to
sanitize message content before API calls. Apply it for all third-party
OpenAI-compat providers and native OpenAI proxies.
This fixes "invalid type: sequence, expected a string" errors from
strict APIs like DeepSeek that reject list-format content in assistant
messages during multi-turn conversations.
* feat: add DeepSeek API key validation and integrate into onboarding process
test: implement unit tests for content flattening utility in OpenAI-compatible providers
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* feat: update MiniMax integration to use Anthropic-compatible endpoint and enhance routing logic
* refactor: streamline OAuth install hint and update ccproxy health check timeout
Add MiniMax (api.minimax.io/v1) as a first-class third-party provider,
enabling direct API access without routing through NVIDIA/SiliconFlow/
OpenRouter intermediaries. Includes M2.5 and M2.5-highspeed models
with 204K context window.
Changes:
- Register "minimax" in _THIRD_PARTY_PROVIDERS with MINIMAX_API_KEY
- Add MiniMax-M2.5 and MiniMax-M2.5-highspeed model entries
- Add minimax_api_key to config, env mappings, and env export
- Add MiniMax to onboarding wizard with API key validation
- Update .env.example, README.md, README.zh-CN.md
- Add 9 unit tests and 3 integration tests (all passing)
Co-authored-by: PR Bot <pr-bot@minimaxi.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
- Added blank lines for better separation of test cases in multiple test files.
- Reformatted event handling in tests for clarity and consistency.
- Ensured consistent use of multi-line formatting for dictionary arguments in event handling.
- Improved assertions and test descriptions for better understanding.
- Updated test cases across various modules including test_stream_state, test_stream_utils, test_summarization, test_thread_selector, test_tool_error_handler, test_tui_widgets, test_ui_runtime, and test_wechat_channel.
Allow the Anthropic provider to accept a base_url override via the
ANTHROPIC_BASE_URL environment variable or anthropic_base_url config
field. This enables routing Anthropic API requests through local proxies
like ccproxy without losing provider-specific features (extended
thinking, adaptive effort) that would be dropped when using the generic
"custom" provider.
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
- Updated README.md to clarify tool allowlist supports glob wildcards.
- Enhanced docstrings in prompts.py, utils.py, and formatter.py for better understanding of function parameters and return values.
- Refactored test cases in test_llm.py and test_onboard.py to use consistent mocking style with unittest.mock.patch.
- Improved test coverage and clarity in test_skills_manager.py and test_stream_state.py by adding descriptive comments and organizing sections.
- Adjusted tool result formatting logic in formatter.py to streamline success checks.
- Ensured backward compatibility in various modules while enhancing functionality.