Commit Graph

21 Commits

Author SHA1 Message Date
m4 4c338ed914 fix(merge): resolve integration gaps found by running the v0.3.0 test suite
Post-merge validation fixes (upstream v0.3.0 + Ai4Sci fork):

- llm/patches.py: restore the two module-level patch calls the merge dropped
  (_patch_openai_empty_sse_keepalive, _patch_deepagents_extracted_document_text)
  and make _is_ccproxy_codex accept an explicit base_url/api_key so the
  invocation plan can classify an endpoint without mutating the process env.
- llm/models.py: an explicit per-call plan now wins over
  EVOSCIENTIST_USE_RESPONSES_API (env is only a default), an explicit caller
  `reasoning` block survives an explicit use_responses_api=False, and the
  third-party (openrouter) default effort stays the fork's fixed `medium`.
- EvoScientist.py: sub-agent stacks pass NO_OP_SINK as `events` instead of None.
- middleware/error_normalization.py: platform-generated diagnostics
  (ModelOutputTruncatedError) keep their actionable text while provider SDK
  errors still get the canned redacted message.
- pyproject.toml: hold google-genai 1.x (langchain-google-genai>=4.3.7,<4.4)
  because llm/gemini_interactions.py drives the 1.x Interactions API; this is
  also what deepagents 0.7.13 requires.
- config/settings.py: restore upstream's use_responses_api config field.
  `reasoning_effort` stays deleted on purpose — Ai4Sci keeps reasoning an
  invocation-plan parameter, never a deployment-env override.
- tests: align upstream tests that encode replaced behaviour (ccproxy
  responses-api context, reasoning-effort-overrides-env, fingerprint coverage)
  with the fork's contracts.
2026-09-13 16:57:03 +08:00
m4 470cf75722 merge: bring upstream v0.3.0 (72 commits) into Ai4Sci fork
Merged upstream/main (418abca, release v0.3.0) into our fork on a
dedicated branch. 21 conflicting files resolved; main worktree untouched.

Resolution policy and key decisions:
- Keep Ai4Sci runtime endpoints, durable dispatch, workspace scopes and
  the HITL/DynamicReview approval chain (approval path is product-critical).
- Adopt upstream model registry (llm/registry.py): our 136 model entries
  are a strict subset of upstream's 180, so dropping our inline table
  loses nothing and gains 44 new models.
- Adopt upstream native EvoChatDeepSeek; drop our obsolete
  _patch_deepseek_reasoning_passback monkey patch.
- Keep our six patches.py additions, ported onto upstream's new
  _OpenAICompatContent class: stable tool-call ids, tool-history
  sanitization, drop_reasoning_metadata, empty-SSE keepalive,
  extracted-document-text patch, _has_assistant_tool_protocol.
- Keep our skill-budget middleware path (skills=None) instead of passing
  skills through, to avoid double loading.
- Keep sanitized error labels (_safe_error_label) while adopting
  upstream's injected MiddlewareEventSink for fallback narration.
- Keep port 3076 and the LANGGRAPH_SERVER_URL override; adopt upstream's
  host/probe-host handling and CONFIG_DRIFT_SINCE_LAUNCH.
- Adopt upstream dependency stack: deepagents 0.7.6, langchain-quickjs
  0.3.7, langgraph-api 0.14; keep our extra deps (rfc8785, pillow,
  firecrawl-anydoc, nest-asyncio).
- Align call sites with upstream APIs: create_tool_selector_middleware
  now takes events= instead of track_stream_selection=.
2026-09-13 16:07:27 +08:00
jfilipiuk 0410b40f57 feat: inherit the caller's model for async sub-agent launch and update (#446)
* feat: inherit the caller's model for async sub-agent launch and update

* docs: tighten middleware related docstrings
2026-09-11 17:58:50 +01:00
m4 c683f6e739 feat: prepare EvoScientist 0.3.0
Docker / build (push) Has been cancelled
Build / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
Add bounded document ingestion, controlled web search, recoverable session support, subagent timeouts, and the native sandbox runtime contract. Unify package versioning and add release-focused regression coverage.
2026-09-03 06:55:56 +08:00
jfilipiuk b500ccc311 fix: prevent session-emptying crash on resume commands (#428)
* fix: prevent session-emptying crash on resume commands

* docs: fix stale docstrings in goto=None crash tests and patch

- Correct checkpoint corruption claim: the error state replaces the
  previous conversation state (messages: [], files: {}), it IS corrupted.
- Replace WebUI-specific language with UI-agnostic wording.
- Remove references to uncommitted local notes files.
- Add upstream issue reference (langchain-ai/langgraph#5656).

* fix: wrap _control_branch instead of reimplementing, use dataclasses.replace, add END routing and session preservation tests

* fix: rebind map_cmd on already-loaded consumers, rewrite crash-path tests to exercise __start__ via checkpoint deletion
2026-08-17 07:10:37 +00:00
m4 5a581c78a2 feat: add scoped model runtime configuration
Build / build (push) Has been cancelled
Docker / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
Introduce provider, model, and invocation contracts with encrypted configuration persistence. Add web runtime fencing, route fallback, recovery middleware, workspace scoping, and comprehensive tests.
2026-08-14 22:03:04 +08:00
Xi Zhang ac58caab7b release: v0.2.4 (#389)
* fix: add support for new Anthropic models and enhance adaptive thinking tests

* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models

* fix: update version to v0.2.4 in badges, README, and project files

* fix: update Star History chart links in README and README.zh-CN

* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog

* fix: update wechat group image in assets
2026-07-26 14:58:48 +01:00
m4 3ce5614254 fix: harden tool-call protocol and fallback handling 2026-07-19 12:05:56 +08:00
Xi Zhang 042da63d54 feat(llm): add Kimi K3 support (#367)
* feat(context-window): add Kimi K3 model with 1M context window

* feat(openrouter): implement structured output for Kimi K3 and add 429 retry handling

* Refactor code structure for improved readability and maintainability
2026-07-18 00:12:08 +01:00
dinos 05dfffbc73 fix(llm): use native langchain-deepseek SDK (#349) 2026-07-15 17:13:33 +01:00
m4 4fc74e7da7 EvoScientist Ai4Sci
Docker / build (push) Has been cancelled
Build / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
2026-07-14 22:07:14 +08:00
jfilipiuk 753c745405 fix: silence YAML-docstring noise from custom-app OpenAPI scan (#317) 2026-07-13 15:38:28 +01:00
Xi Zhang 63969b596d Release/v0.1.4 (#266)
* feat(middleware): reposition code interpreter middleware in the stack

* feat(models): add qwen3.7-plus model entry and update context window comment

* feat(models): add qwen3.7-max and qwen3.7-plus model entries for DashScope

* feat(auxiliary): implement auxiliary model support for background tasks and tool selection

- Added auxiliary model configuration to EvoScientistConfig.
- Introduced _ensure_auxiliary_chat_model function to manage auxiliary model instances.
- Updated onboarding steps to include auxiliary model selection.
- Modified middleware to route tool selection to the auxiliary model when applicable.
- Enhanced tests to cover auxiliary model functionality and configuration.

* feat(steps): update UI backend selection options and descriptions

* Refactor code structure for improved readability and maintainability

* feat(patches): implement OpenRouter response reasoning item stripping to prevent multi-turn errors

* feat: update version to v0.1.4 in badges, README, and pyproject.toml; adjust skill counts in steps.py

* feat(config): add auxiliary model and provider environment variables to test setup
2026-06-07 00:52:59 +01:00
Xi Zhang 9cffe9d457 Enhance multimodal handling in LLM model (#256)
* Enhance multimodal handling in LLM model

- Updated `_flatten_message_content` to preserve media blocks (images, files) while flattening text content.
- Introduced `_sanitize_messages` to manage media hoisting for tool messages, ensuring compatibility with OpenAI APIs.
- Modified `_patch_openai_compat_content` to accommodate new media handling logic, including retry mechanisms for media errors.
- Added comprehensive tests for media preservation, including various scenarios with images, files, and unsupported media types.

* fix: preserve order of text and media blocks in message flattening

* test: add tests for _strip_media_types to ensure position preservation and deduplication
2026-06-03 01:06:46 +01:00
Xi Zhang c407d2e20f Fix/async subagent model switch (#217)
* feat(middleware): add ConfigurableModelMiddleware for dynamic model resolution

- Introduced ConfigurableModelMiddleware to resolve chat models from RunnableConfig.configurable on each call.
- Updated middleware initialization to include ConfigurableModelMiddleware.
- Enhanced context editing middleware tests to verify presence of ConfigurableModelMiddleware.
- Implemented tests for ConfigurableModelMiddleware to ensure correct model overriding and caching behavior.
- Added tests for deepagents model-passthrough patch to verify configuration injection in async tasks.

* feat(async-subagent): update middleware handling to prevent deadlocks in async sub-agents

* style: Refactor code formatting for improved readability in patches and test files

* refactor: streamline middleware construction and improve async handling in ConfigurableModelMiddleware

* fix: remove unused request parameter from _read_model_override function

* refactor: improve async handling in _ClientProxy and enhance logging in ConfigurableModelMiddleware
test: add behavior test to ensure AskUserMiddleware is excluded in async subagent mode
2026-05-08 22:52:24 +01:00
Xi Zhang d7c0eec0e9 chore: update dependencies and remove unused OpenRouter patch (#211) 2026-05-05 18:51:00 +01:00
Xi Zhang 56cc2fef85 fix(deepseek): add empty-string fallback for reasoning_content in cross-provider scenarios (#192) 2026-04-27 19:19:50 +01:00
Xi Zhang 20c06d4897 feat(deepseek): implement reasoning_content passback for multi-turn s… (#190)
* feat(deepseek): implement reasoning_content passback for multi-turn scenarios

* fix(tests): ensure consistent import of EvoScientist.llm.patches in test cases

* fix(tests): streamline tool_calls formatting in TestPatchDeepseekReasoningPassback

* feat(patches): add reasoning_content capture and re-injection for DeepSeek assistant messages

* fix(deepseek): optimize reasoning_content extraction and assignment in passback

* fix(deepseek): refine reasoning_content handling in OpenAI capture patch
2026-04-26 14:53:08 +01:00
Xi Zhang d5b982c980 fix(ccproxy): update Responses API handling and patch system role con… (#149)
* fix(ccproxy): update Responses API handling and patch system role conversion

* fix(ccproxy): streamline _agenerate method in system to developer patch

* fix(ccproxy): improve handling of None output in Codex compatibility patch
2026-04-09 12:51:05 +02:00
Xi Zhang 4f11de23a2 fix(llm): patch _stream/_astream for OpenAI-compatible content flattening (#147)
* fix(llm): patch _stream/_astream for OpenAI-compatible content flattening

_patch_openai_compat_content() only patched _generate/_agenerate but
EvoSci CLI uses streaming paths. This extends the content flattening
to _stream/_astream so strict OpenAI-compatible relays receive plain
string content during streaming calls.

Closes #142

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): use asyncio.run() instead of pytest-asyncio for CI compat

CI does not have pytest-asyncio installed, so async tests must use
asyncio.run() wrapper instead of @pytest.mark.asyncio decorator.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): use @pytest.mark.anyio for async tests (CI compat)

CI does not have pytest-asyncio. Use @pytest.mark.anyio consistent
with existing async tests in the project.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 20:24:13 +01:00
Xi Zhang 68f3ab2962 feat: enable reasoning for OpenRouter via extra_body to prevent multi… (#124)
* feat: enable reasoning for OpenRouter via extra_body to prevent multi-turn errors

* feat: implement OpenRouter native reasoning support and patch langchain-openrouter bug

* feat: add OpenRouter reasoning effort configuration and update related tests

* feat: add langchain-openrouter dependency for enhanced reasoning support

* fix: correct spacing in reasoning effort choice label

* feat: implement patch for OpenRouter reasoning details to prevent Pydantic errors

* feat: add patches for OpenRouter reasoning and content handling utilities

* feat: prevent multiple patches of OpenRouter reasoning details by using a global flag

* feat: update OpenRouter reasoning patch to ensure single application with global flag

* feat: refine OpenAI responses API handling to apply only for OpenAI provider

* feat: Enhance TUI interaction by updating todo widget positioning and skipping empty tool call chunks

* feat: Update tool selector threshold and adjust logging level for selector failures

* feat: Temporarily disable timestamp toast in tool call widget for UX review

* feat: Re-enable timestamp toast in tool call widget on click
2026-04-03 11:04:34 +01:00