Commit Graph

253 Commits

Author SHA1 Message Date
m4 b68fdf6d57 feat: disable agent observation writes in web deploy
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-19 21:59:38 +08:00
m4 4445379f0a feat: resolve memory middleware dir per-call from configurable 2026-08-19 21:50:47 +08:00
m4 5041807a8e feat: resolve observation tool memory_dir per-call from configurable 2026-08-19 21:43:23 +08:00
m4 8376f56ab4 feat: native sandbox execution, dynamic review middleware, and workspace files
Adds native sandbox execution runtime, dynamic review middleware, and
workspace file handling, with supporting stream events, prompt, and scope
registry changes plus architecture docs.
2026-08-19 20:00:09 +08:00
m4 5a581c78a2 feat: add scoped model runtime configuration
Build / build (push) Has been cancelled
Docker / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
Introduce provider, model, and invocation contracts with encrypted configuration persistence. Add web runtime fencing, route fallback, recovery middleware, workspace scoping, and comprehensive tests.
2026-08-14 22:03:04 +08:00
m4 3ce5614254 fix: harden tool-call protocol and fallback handling 2026-07-19 12:05:56 +08:00
m4 4fc74e7da7 EvoScientist Ai4Sci
Docker / build (push) Has been cancelled
Build / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
2026-07-14 22:07:14 +08:00
jfilipiuk 753c745405 fix: silence YAML-docstring noise from custom-app OpenAPI scan (#317) 2026-07-13 15:38:28 +01:00
jfilipiuk 88ac9f5ba1 fix: surface real exception class+message in SSE error events (#315)
* fix: surface real exception class+message in SSE error events

* fix: tighten SSE error patch scope and key redaction

* fix: redact base64-style secret suffixes fully

* style: remove notes/ reference from the dosctring

* fix: rebuild env cache on each error call

* fix: route BaseException through serde.default on SSE/webhook paths

* fix: distinguish routed providers by request URL host

* feat: normalize provider-SDK exceptions via ErrorNormalizationMiddleware

* refactor: drop json_dumpb dataclass-bypass wrappers, superseded by middleware

* fix: guard _extract_host against SDK properties that raise

* refactor: derive provider tag from ModelRequest.model, not the exception

* refactor: drop serde.default patch and exception-based inference; ProviderStreamError.model_dump handles the emit

* refactor: move envelope helpers from patches.py to errors.py

* feat: extend ErrorNormalizationMiddleware coverage to every model-call path

* chore: clean up review findings from middleware pivot

* fix: pass through all langgraph.errors

* fix: move langgraph.errors pass-through into _normalize

* fix: pass through ContextOverflowError in _normalize
2026-07-13 14:17:56 +01:00
jfilipiuk 952e68efe3 fix: scope quickjs snapshot to turn to keep checkpoints small (#316)
* fix: scope quickjs snapshot to turn to keep checkpoints small

* fix: strip _quickjs_snapshot_payload from state/history responses instead of dropping mode=thread

* fix: recurse strip into nested subgraph StateSnapshot

* fix: drop conditional-snapshot gate that leaked repl slots

* fix: assert LangGraph state-shape invariants at import time

* refactor: discover graphs to filter from langgraph.json

* test: assert copy() preserves subclass; iterate langgraph.json for subagent coverage

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 12:24:52 +00:00
Mani Saint-Victor 2b28c46caf fix(llm): make gpt-5.x usable through ccproxy Codex OAuth (#324)
* fix(llm): make gpt-5.x usable through ccproxy Codex OAuth

Two independent blockers made current OpenAI models fail when routed
through ccproxy's Codex OAuth endpoint:

1. ccproxy's default Codex model mappings rewrite any gpt-*/o1-*/o3-*/
   claude-* model to gpt-5.3-codex before forwarding, silently overriding
   the configured model and failing outright on accounts where
   gpt-5.3-codex is not served ("The 'gpt-5.3-codex' model is not
   supported when using Codex with a ChatGPT account").
   start_ccproxy() now generates a config with empty codex model
   mappings and passes it via 'ccproxy serve --config'.

2. ccproxy forwards the client's own User-Agent upstream and only
   gap-fills its Codex headers, so the backend gates current models on
   the client identity ("The '<model>' model requires a newer version
   of Codex"). get_chat_model() now sends Codex-CLI-shaped
   originator/version/User-Agent headers when the ccproxy Codex adapter
   is detected, overridable via EVOSCIENTIST_CODEX_CLIENT_VERSION.

Verified live: gpt-5.5 and gpt-5.4 complete successfully through
ccproxy Codex OAuth on a ChatGPT Plus account with both fixes; each
fails without them.

* fix(ccproxy): harden Codex client routing

* fix(llm): keep Codex client identity consistent

* docs: clarify Codex version floor

* style: ruff format models.py after merge

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-07-13 11:04:47 +00:00
Mani Saint-Victor da6ca38d53 fix(llm): respect reasoning_effort setting on native OpenAI path (#321)
* fix(llm): respect reasoning_effort setting on native OpenAI path

The native OpenAI provider path hardcoded reasoning effort to xhigh for
gpt-5.4/5.5/codex models, silently ignoring the user's reasoning_effort
config setting. The OpenRouter path already honors the
EVOSCIENTIST_REASONING_EFFORT env var that settings.py exports from that
setting; this applies the same lookup on the native path, falling back
to the previous defaults when unset.

Adds a regression test and isolates the existing xhigh test from the
env var.

* fix(llm): preserve model reasoning defaults

* fix(llm): preserve GPT-5.6 reasoning default

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 10:42:55 +00:00
Mani Saint-Victor 19888c2db6 fix(ccproxy): raise startup timeouts (auth check 10s→30s, serve health 30s→180s) (#328)
* fix(ccproxy): raise auth status check timeout to 30s

ccproxy's CLI initializes its full plugin system on every invocation;
a cold 'ccproxy auth status' takes ~10s wall time on Apple Silicon,
so the 10s subprocess timeout made OAuth startup fail intermittently
with 'Auth check timed out' even when credentials were valid.

* fix(ccproxy): raise serve health deadline to 120s

ccproxy boot includes plugin init plus Codex CLI detection; measured
~76s to first healthy response on an Apple Silicon Mac (ccproxy-api
0.2.9). The 30s deadline in start_ccproxy() killed the process before
it could come up, failing OAuth startup with 'ccproxy did not become
healthy within 30 seconds'.

* fix(ccproxy): widen serve health deadline to 180s

Full startup measured at ~111s on a second cold run (Apple Silicon,
ccproxy-api 0.2.9); 120s left too little headroom for boot variance.

* fix(ccproxy): centralize startup timeouts

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 10:36:05 +00:00
Mani Saint-Victor f3e65a446f fix(tests): isolate the developer's real .env from the test suite (#329)
get_effective_config() runs load_dotenv(find_dotenv(usecwd=True),
override=True), so any test that loads config injected the repo's real
.env into os.environ for the rest of the pytest process. An
empty-valued line like MINIMAX_BASE_URL= then made
os.environ.get(key, default) return '' instead of the default,
failing the MiniMax routing tests in full-suite runs while they
passed in isolation.

Generalizes the find_dotenv redirect that test_config.py's
temp_config_dir fixture already applied locally into a suite-wide
autouse fixture, pointing at a never-created path so tests writing
their own tmp_path/.env cannot collide with it. Adds a regression
test reproducing the leak.

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 10:27:40 +00:00
Zixin Dong f72f7b93d5 feat(llm): add OpenRouter app attribution headers (#339) (#344)
* feat(llm): add OpenRouter app attribution headers (#339)

Attach EvoScientist app-attribution at the shared model-init layer so all
OpenRouter calls are credited to the project. langchain-openrouter maps
app_url/app_title/app_categories -> HTTP-Referer / X-Title /
X-OpenRouter-Categories. Applied only for the openrouter provider, via
setdefault so explicit caller kwargs win. Configurable through new
openrouter_http_referer / openrouter_app_title / openrouter_app_categories
settings and their EVOSCIENTIST_OPENROUTER_* env vars.

Closes #339

* refactor(llm): centralize OpenRouter attribution defaults + cap categories

Address PR #344 review:
- Define the app-attribution default constants once in config/settings.py
  (the config fields and llm/models.py both use them) instead of duplicating
  the literals across the two modules.
- Reduce the default categories to creative-writing,personal-agent and cap the
  sent list to OpenRouter's 2-per-request limit, warning when a configured list
  exceeds it, so extras are dropped predictably (and surfaced) here rather than
  being silently truncated server-side.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-07-13 11:08:59 +01:00
dinos 49770949da fix(langgraph): prefer executable in venv over path (#341) 2026-07-11 11:59:27 +00:00
dinos 690b903f85 test: standardize async tests on pytest-asyncio auto mode (#338)
* chore: add pytest-asyncio in auto mode

* test: migrate channel and stream tests to native async

Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.

* test: migrate command and model/middleware tests to native async

Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.

* test: migrate TUI, notifier, gateway, and session tests to native async

TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.

* test: replace direct asyncio.run() calls with native async tests

Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.

* test: drop undeclared anyio markers and delete run_async helper

The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).
2026-07-08 18:37:48 +00:00
dinos d2452c54d5 Refactor onboarding OAuth flow for auxiliary models (#337)
* refactor(onboard): shared flow for ccproxy providers

* feat(onboard): support oauth configuration for auxiliary models

* fix(onboard): reuse main model auth for same-provider auxiliary

* fix(onboard): reconcile oauth providers
2026-07-08 18:28:44 +00:00
dinos a7b9e175c1 fix(config): set config.yaml permissions to 0x600 (#336) 2026-07-08 19:25:01 +01:00
dinos be3dd272c3 test: deflake timing-dependent tests (#335)
* test: deflake timing-dependent tests

Inject a clock into channel dedup tests, replace fixed async sleeps with
events/explicit flushes, and avoid wall-clock waits in background tests.

* coderabbit nit
2026-07-07 08:25:35 +01:00
Wiktor Cupiał 1d117ff277 feat: completion enchancements (#302)
* feat: completion enchancements

* fix: handle exception

* fix: duplicate view

* fix: remove deadcode

* fix tab
2026-07-05 05:14:42 +00:00
renaissancefieldlite f086d77756 Fix UTF-8 config reads on Windows (#318)
* Fix UTF-8 config reads on Windows

* test: cover utf8 production loaders

* fix: read and write settings as utf8

* Apply ruff formatting
2026-07-05 05:10:24 +00:00
kalisgd0h bd54eaa0a4 feat(cli): add --output-format stream-json for headless clients (#309)
* feat(cli): add --output-format stream-json for headless clients

Emit EvoScientist's native event stream as line-delimited JSON on stdout
in single-shot (-p) mode, with all human output redirected to stderr so
stdout stays pure JSONL. Intended as the integration surface for
programmatic clients (e.g. an agent runtime) that drive EvoSci headlessly.

- stream/json_sink.py: write_events_as_json + stream_json sink, plus
  redirect_console_to_stderr helper for stdout purity
- cli/interactive.py: cmd_run gains output_format; stream-json branch runs
  the sink instead of the Rich renderer
- cli/commands.py: --output-format option + validation (stream-json
  requires -p; value must be text|stream-json)
- docs/stream-json.md: event-schema contract + example transcript
- tests: json sink serialization, CLI dispatch, console redirect, validation

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(cli): honor explicit --no-auto-mode over config in stream-json

Address CodeRabbit review (discussion_r3514041123): the auto-mode override
block only wrote to cli_overrides when the resolved value was True, so an
explicit --no-auto-mode silently fell back to a config that enables
auto-mode -- breaking "explicit flags always win" and leaving stream-json
running unattended despite the warning. Write auto_mode=False when the flag
is explicitly False. Add regression tests that capture the overrides.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 04:53:52 +00:00
dinos 2b244888ec feat(memory): autoskills (#319) 2026-07-03 11:16:33 +02:00
houren Antony 7568a6bc1f fix(tui): keep welcome banner at top after /new (#311)
* fix(tui): keep welcome banner at top after /new

PR #262 replaced scroll_end() with anchor() for free-scrolling.
When /new clears a long anchored conversation, the anchor kept
the viewport pinned to the (now empty) bottom, producing a
negative scroll_y and pushing the welcome banner out of view.

Reset the anchor and scroll to the top in clear_chat(), and
restore the follow/new-content flags so the fresh session starts
correctly.

Closes #301

* fix(tui): suppress anchor when chat content fits viewport

The previous fix for #301 only handled the /new path. din0s reported
that the banner still dropped to the bottom after a normal short turn
(user types 'hi', agent replies) — i.e. whenever the conversation
fit in the viewport. Root cause is in Textual's compositor
(textual._compositor): when a widget is anchored, scroll_y is
recomputed via set_reactive, which bypasses the validator. If the
anchored widget's content is shorter than the viewport, scroll_y
goes negative on the next layout pass and the welcome banner is
pushed below the visible region.

PR #262 made _stream_with_widgets re-engage the anchor at the end of
every turn via _anchor_chat, so the bug surfaced on any short reply
that fit in the viewport. Markdown re-renders, status-bar updates,
or any subsequent mount would then trip the compositor.

Fix in three places:

  * _anchor_chat: only engage the anchor when max_scroll_y > 0;
    otherwise release and scroll_home so the banner stays at the top.
  * streaming anchor loop: if content shrinks below the viewport
    mid-stream (e.g. loading widget removed), release the anchor
    instead of leaving _anchored=True for the compositor to trip on.
  * clear_chat: keep the unconditional reset (children are removed
    asynchronously so a max_scroll_y check would be stale) but
    document why.

Adds two regressions:
  * test_short_turn_keeps_banner_at_top_after_layout_refresh — the
    exact scenario din0s tested; fails with scroll_y=-10 on the
    previous code, passes with the fix.
  * test_long_turn_keeps_viewport_pinned_to_bottom — guards against
    regressing free-scrolling for overflowing conversations.

Manually verified: 'hi' -> reply (banner stays at top) -> /new
(banner at top) -> another turn (banner stays at top).

* test(tui): address review feedback on banner-position regressions

- extract `_release_anchor_and_pin_top` helper for the 3-line
  `anchor(False) + scroll_home(...)` pattern repeated in
  `clear_chat`, `_anchor_chat`, and the streaming loop
- replace `pytest.skip` in `_capture_app` with a hard `RuntimeError`
  so a broken capture never silently passes
- drop the redundant `load_agent` and `create_session_workspace`
  monkeypatches (the factory is given those as parameters, so the
  module-level symbols never run; added a comment explaining why)
- add a defensive `_FakeChannelRuntime` patch for symmetry with the
  other module-level fakes
- drop the local `_run` helper and use the `run_async` fixture from
  `conftest.py` (its teardown is better)

* test(tui): replace _FakeChannelRuntime with _auto_start_channel no-op

The _FakeChannelRuntime patch was ineffective because ChannelRuntime is
just a dataclass — the real channel manager still started via
_auto_start_channel, leaving pending tasks and non-hermetic test state.

Per review feedback, stub _auto_start_channel directly instead.
2026-06-30 19:15:08 +01:00
jfilipiuk 7214d4099d feat: expose model registry at GET /api/models (#308)
* feat: expose model registry at GET /api/models

* fix: include Ollama models in /api/models endpoint

* fix: honor env vars override in /api/models endpoint

* fix: offload get_effective_config to thread to satisfy blockbuster

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-26 22:48:37 +01:00
dinos f2f010a350 feat(memory): observation linking (#307)
* refactor(gateway): create module for launching async/bg agents

* refactor(memory): refactor worker launch around source context & output deltas

* refactor(gateway): generalize async/bg module

* refactor(memory): revamp worker launching

* feat(memory): add observation linking

* test(memory): remove redundant test branches

* fix(memory): make 'supersedes' relation directional

* fix(memory): don't create empty project observation dirs

* fix(memory): schedule direct observations for linking

* fix(cli): wait for observation linker before shutdown

* fix(memory): block arbitrary writes to /memories

* fix(linker): remove `linked_by` attribute from frontmatter

* refactor(linker): rename base relationship to `comlpements`

* fix(cli): bump worker wait to 2m

* feat(tools): catch malformed tool calls & retry

* feat(status): add linking result to statusbar

* fix(linker): don't launch linker when observations are disabled

* fix(memory): use posix paths

* fix(watcher): call abort hook on error status

* fix(watcher): delete thread on failed run creation

* fix(watcher): preserve url

* fix(observation): record session_id, drop unused fields

* fix(memory): reject unsupported worker source types

* refactor(backends): shared memory backend builder

* fix(scheduler): resolve linker inputs outside lock

* fix(memory): dont launch workers / record observations without thread_id

* feat(memory): include related observations in tool results

* fix(memory): skip malformed observation frontmatter

* revert(tools): drop tool error handling changes from this PR

* fix(memory): serialize observation link writes

* fix(memory): queue observations written by aborted workers

* fix(memory): track observation linker launch handoff

* fix(memory): resolve cross-project related observations

* fix(status): avoid recounting reason-only link updates

* fix(memory): avoid rereading file for content

* fix(linker): use neutral prose for bidirectional reasons

* test(memory): coverage for aborted/failed launches

* test(memory): cleanup & helpers

* feat(linker): add observations index hint
2026-06-26 22:20:52 +01:00
Xi Zhang 7ccfe68f3f feat: add scheduler functionality with cron-style task management (#306)
* feat: add scheduler functionality with cron-style task management

- Implemented a new scheduler subagent to automate recurring tasks using cron expressions.
- Enhanced the subagent factory to include the skill manager and auxiliary chat model for the scheduler.
- Created a YAML configuration for the scheduler with a detailed system prompt and toolset.
- Updated README files to include documentation on scheduled tasks and usage examples.
- Added tests for the scheduler, including command execution, scheduling tools, and middleware integration.
- Introduced new dependencies for timezone handling and ensured compatibility in the project configuration.

* fix(async-notifier): ensure fallback hint is used for unknown notification kinds

* feat: enhance scheduling functionality and improve system message handling
2026-06-25 17:31:23 +01:00
dinos b1dccf17ea fix(tool-selector): memory tools & state for main agent (#305)
* fix(tool-selector): always include memory tools

* fix(tool-selector): only track & show state for main agent

* docs: update docstring
2026-06-23 14:47:25 +01:00
X-iZhang c063a00c8f Add regression tests for code_interpreter middleware PTC allowlist
- Ensure 'task' is excluded from the default PTC allowlist to prevent ValueError in langchain-quickjs >=0.3.
- Verify that essential async dispatch tools remain in the allowlist.
- Confirm that the live quickjs filter accepts the default allowlist even with a 'task' tool present.
- Test the creation of the code_interpreter middleware to ensure it builds correctly.
2026-06-23 08:40:35 +01:00
dinos bd307f3a11 refactor: LangGraph gateway layer for UI-agnostic graph and thread access (#295)
* feat(gateway): graph gateway protocol

* refactor(cli): wire gateway in cli/tui

* refactor(gateway): centralize runtime gateway init

* chore(gateway): restrict RunRequest message type

* feat(gateway): add langgraph server gateway

* chore(cli): tighten serve runtime state typing

* refactor(cli): route async task state reads through graph gateway

* refactor(gateway): support graph targets in server gateway

* refactor(cli): route session commands through graph gateway

* refactor(cli): fold thread store under graph gateway

* refactor(gateway): route graph state access through gateway

* refactor(channels): wire graph gateway

* refactor(memory): preserve graph threads for cloning

* feat(gateway): add thread cloning

* fix(tui): pass effective workspace for thread creation

* chore(memory): add workspare dir to memory worker metadata

* fix(sessions): filter preloaded UUID registy entries by the current scope

* test(fakes): use https

* refactor(consumer): consolidate imports

* fix(stream): optional summarization event

* fix(gateway): resolve abbreviated thread IDs by search

* fix(gateway): page server thread listings

* fix(gateway): emit pending interrupt events

* style: fmt

* feat(gateway): persist workspace_dir & model in thread metadata

* fix(gateway): page server thread prefix resolution

* fix(gateway): expose server thread list metadata

* refactor: add back type def

* refactor: tighten types

* revert: add back worker thread deletion

The worker thread forking changes are out of scope for now, so to
maintain parity with the existing behavior we'll leave this intact.

* fix(gateway): apply compaction to server thread history

* refactor(stream): restore direct summary replay suppression

* fix(gateway): preserve compaction state and server stream output

* fix(gateway): close local stream generator on cancellation

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-22 13:54:14 +00:00
dinos 6eb467e70b feat(llm): make OpenRouter Anthropic prompt cache opt-out (#299)
* feat(llm): make OpenRouter Anthropic prompt cache opt-out

* fix(llm): restore truthy env flag helper
2026-06-18 15:09:22 +02:00
Ziheng Zhang 6593ac9b5b fix(cli): submit slash command on Enter when name prefixes another (#293) (#300)
Typing `/model` and pressing Enter did nothing in the TUI; the picker
only opened via `/model --save` or `/model <name>`. The completion popup
matched both `/model` and `/model-fallback` by prefix, so the
exact-match-hide guard (which required a single match) never fired. With
the popup still visible, the TUI's Enter handler completed the text
instead of submitting the command, so it never executed.

Treat the typed prefix as an exact match whenever it equals any matched
command name, not only when it is the sole match. This hides the popup
on a complete command name so Enter submits it, even when a longer
command shares the prefix.
2026-06-18 14:07:43 +01:00
dinos de3f588fbf fix: windows async MCP tool execution, restore graph state after interruptions (#290)
Co-authored-by: z00827015 <zhoulun1@huawei.com>
2026-06-16 21:03:26 +02:00
dinos f356de36a6 feat: memory retrieval (#281) 2026-06-16 09:14:29 +02:00
houren Antony 477481f9d9 feat(backends): use _platform_quote for Windows cmd.exe compatibility (#280)
* feat(backends): use _platform_quote for Windows cmd.exe compatibility

Resolves the 3 skipped E2E tests in test_backends.py that exercised
the /skills/... mount path. The path-rewriter was wrapping resolved
absolute paths via shlex.quote (POSIX single-quote style); cmd.exe
doesn't strip single quotes, so the literal ' characters ended up in
the subprocess argv and the python script failed to find its file.

Replace the 3 shlex.quote call sites in _resolve_virtual_mount_path
with _platform_quote, a thin platform dispatcher:
- POSIX: shlex.quote (unchanged)
- Windows: _cmd_quote uses cmd.exe-compatible double-quote wrapping
  and properly escapes embedded " and percent signs

Adds:
- backends.py: _is_windows, _cmd_quote, _platform_quote (~40 lines)
- test_backends.py: 6 TestPlatformQuote unit tests + _split_cmd
  cross-platform tokenizer helper to replace shlex.split in the 8
  sites that tokenize convert_virtual_paths_in_command results
  (POSIX shlex strips backslashes from bare Windows paths, which
  broke the 5 TestVirtualMountResolution assertions on Windows)

Removes:
- 3 @pytest.mark.skipif(sys.platform == "win32") markers on the
  E2E tests for /skills/... mount resolution

Refs #274.

* fix: escape % as %% in _cmd_quote instead of relying on double-quoting

cmd.exe expands %VAR% before processing quotes, so double-quoting
cannot neutralize percent signs.  Escape bare % as %% (the cmd.exe
idiom for a literal percent) before any other quoting logic.

Also updates _cmd_quote docstring and _resolve_virtual_mount_path
docstring to reflect the actual quoting strategy.

* style: fix ruff format (single → double quotes)

* fix: treat % as regular char in _cmd_quote, document limitation

%% escaping only collapses in .bat/.cmd files, not via cmd /c.
Since virtual-mount paths should never contain % in practice,
simpler to leave % alone and document the caveat.
2026-06-13 18:08:18 +01:00
houren Antony 76972449c7 feat(cli): multi-stage slash command completions with subcommand awareness (Phase 1 of #82) (#273)
* feat(cli): multi-stage slash command completions with subcommand awareness

Phase 1 of #82 — subcommand and argument awareness in completions.

- commands/base.py: add SubCommand dataclass and subcommands/category
  ClassVars to the Command ABC. Each SubCommand has name, description,
  and optional arguments.

- commands/manager.py: add get_subcommands() and list_subcommands()
  methods to expose subcommand metadata for completion rendering.

- commands/implementation/mcp.py: declare 6 subcommands (list, config,
  add, edit, remove, install).

- commands/implementation/model_fallback.py: declare 6 subcommands
  (list, add, remove, clear, save, help).

- commands/implementation/channel.py: declare 2 subcommands
  (status, stop).

- cli/tui_interactive.py: rewrite on_text_area_changed slash-completion
  branch. When the user types a command name + trailing space and the
  command has subcommands, show subcommand completions instead of hiding
  the popup. Filter subcommands by typed prefix in multi-token input.

- commands/implementation/general.py: /help now lists subcommands
  below each command that declares them.

Tests: 8 new tests covering SubCommand creation, CommandManager
subcommand lookup, and cross-command verification.
2293 passed baseline, no regressions.

* fix: subcommand completion preserves prefix + prompt_toolkit + tests

- _apply_selected_completion: preserve '/mcp ' prefix when completing
  subcommands via _comp_is_subcommand flag
- SlashCommandCompleter (Rich CLI): add subcommand completion support
- Fix trailing-space bug: rstrip prefix before top-level matching
- test_tui_widgets.py: update stub on_input_changed to match multi-stage
  logic; add 5 new subcommand tests
- test_command_manager.py: 8 tests for SubCommand + CommandManager

28 passed, 0 failed.

* style: ruff format tui_interactive.py + test_tui_widgets.py

* style: fix RUF012 ClassVar annotation on subcommands lists

* fix: sync test stub, add len>=3 guard, remove exact-match hide

- Sync test stub on_input_changed with real TUI code (remove exact-match
  hide for subcommands, add len(parts)>=3 guard)
- Update test_input_changed_exact_subcommand_hides -> shows_confirmation
- Add test_input_changed_three_parts_hides
- Remove unused category ClassVar (din0s: what is this for)

* refactor(commands): extract shared completion engine

Per din0s feedback: one shared completion engine (commands/_completion_engine.py)
that parses text + cursor once, returns structured CompletionCandidate objects
with replace_start/replace_end ranges.

- SlashCommandCompleter (Rich CLI): thin adapter, delegates to engine
- on_text_area_changed (TUI): thin adapter, delegates to engine
- _apply_selected_completion: uses candidate.replace_start/replace_end instead
  of _comp_is_subcommand flag
- Tests: engine tested directly (10 new tests), stub methods updated

29 passed, 0 failed.

* style: ruff format

* fix: preserve text after cursor when applying completion

CodeRabbit: replace_start only cuts from start to cursor,
dropping any suffix after the cursor. Use replace_start + replace_end
to correctly splice the replacement while preserving trailing text.

* fix(cli): repair slash-command completion (TUI crash, subcommand bugs, sort)

Apology + context: the previous push shipped a TUI-breaking change
(the new shared engine assumed ``event.text_area.cursor_position``
existed, but ``ChatTextArea`` / ``Changed`` don't expose it). User
caught the crash on ``/``; fixing that surfaced two more bugs in
the engine that din0s had already flagged. This commit addresses
all of them and drops a piece of dead stub code.

## Bug fixes

1. **TUI crash on ``/``** (``tui_interactive.py:2335``)
   ``event.cursor_position`` doesn't exist on the ``Changed`` event,
   and ``ChatTextArea`` (Textual ``TextArea`` subclass) doesn't expose
   ``cursor_position`` either. Pass ``len(event.text_area.text)``
   instead — the user types at the end of the input in practice.

2. **Subcommand trailing-space duplication** (``_completion_engine.py``)
   Typing ``/mcp a `` + Tab produced ``/mcp aadd``. The engine
   included the trailing space in ``replace_end``; the TUI apply
   unconditionally appended ``" "``, producing double-space output.
   Fix: ``replace_end`` excludes the trailing space; the TUI apply
   checks ``current[replace_end:].startswith(" ")`` and skips the
   separator when the suffix already has one.

3. **Subcommand exact-match confirmation noise** (``_completion_engine.py``)
   Typing ``/mcp list`` + Tab re-inserted ``list`` and the popup
   kept showing the same subcommand. Add a guard mirroring the
   top-level exact-match rule: when the only subcommand match is
   the prefix itself (no trailing space), return ``empty``.

4. **Alphabetical sort dropped in CLI** (``cli/interactive.py``)
   The new completer iterated ``result.candidates`` in registration
   order. Re-add ``sorted(result.candidates, key=lambda c: c.text)``.
   Same sort added to the TUI for consistency.

## Cleanup

- Drop the dead ``on_input_changed`` method from the ``_StubApp``
  test stub (0 call sites) plus the unused ``_slash_commands`` /
  ``_subcommands`` locals that fed it. This addresses din0s's
  comment about the stub duplicating real TUI logic — the inlined
  copy is no longer needed since the real completer now routes
  through the shared engine.

## Tests

- ``test_engine_exact_subcommand_shows_confirmation`` → renamed to
  ``test_engine_exact_subcommand_hides`` to match new behavior.
- New: ``test_engine_subcommand_trailing_space_excludes_space_from_range``
  and ``test_engine_subcommand_trailing_space_apply_does_not_double_space``.
- All 97 tests in ``test_tui_widgets.py`` pass.
- ``ruff check`` / ``ruff format`` clean.
- Local TUI smoke: ``/`` (no crash, top-level popup), ``/mcp ``
  (subcommand popup), ``/mcp a `` + Tab → ``/mcp add ``.

Refs the din0s review comments on PR #273. CLI path tests and the
``category`` ClassVar follow-up are deferred to a separate PR (the
former is a test-suite addition; the latter is already absent from
``base.py`` on the current branch).

* fix: address remaining review items (help duplication, CLI tests, stub sync, docstrings)

- mcp.py: auto-generate help text from subcommands ClassVar (#1)
- tests/test_cli_completion.py: add 9 CLI completer tests (#2c)
- test_tui_widgets.py: sync _apply_selected_completion stub with real code (#4)
- mcp.py + interactive.py: add docstrings to key functions (#8)

* fix: hide completions on exact subcommand match regardless of trailing space

Remove the
ot has_trailing_space guard from the exact-subcommand
check.  Previously /mcp list  (with trailing space) would still
return candidates, causing Tab to oscillate between adding and removing
the trailing whitespace.  Now the engine hides whenever the subcommand
is an exact match, same as the top-level rule.

Added test_engine_exact_subcommand_with_trailing_space_hides to cover
the scenario din0s flagged.

* refactor: use StrEnum for CompletionResult.kind

Replace plain str with CompletionKind(StrEnum) for type safety.
Backward-compatible with existing string comparisons.

* fix: normalize @file completion tuples to CompletionCandidate

complete_file_mention() returns list[tuple[str, str]] but the TUI
rendering/apply code expects objects with .text/.description.
Wrap tuples in CompletionCandidate to prevent AttributeError crash.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-13 17:06:23 +00:00
houren Antony 8f9dfa159e fix(backends): rewrite quoted virtual paths containing whitespace (#269)
* fix(backends): rewrite quoted virtual paths containing whitespace

The `convert_virtual_paths_in_command` regex
`(?<=\s)/[^\s;|&<>'"`]*` stopped at the first whitespace or quote,
so:

  - `python "/skills/my skill/main.py"` was left completely
    unchanged (the `(?<=\s)` lookbehind failed after the opening
    `"`), and the shell then broke the inner unquoted path at the
    embedded space.
  - `python /skills/my skill/main.py` was truncated to
    `python ./skills/my skill/main.py` (only `/skills/my` rewritten).

Replace the regex with `shlex.shlex(command, posix=True,
punctuation_chars=";|&<>")` so quoted regions stay whole, then
splice the rewrite back into the original command — extending the
splice span to include any matching quote chars around the path so
the fresh `shlex.quote` of the replacement isn't double-wrapped.

`_resolve_virtual_mount_path` now returns the unquoted path; the
caller owns shell-quoting, which avoids the previous
`shlex.quote` inside original `"…"` leaving literal `'` chars in
the argument value.

Unquoted paths with embedded whitespace remain a known limitation
(shlex has no way to know the user meant one path) — the
workaround of avoiding spaces in skill directory names still
applies, as flagged in the original issue.

Closes #237

* fix(backends): backslash-escaped paths, multi-path per token, subshell paths

- Fix backslash-escape handling: use unescape before rewriting
- Fix re.search→re.finditer: all /-paths in a token are rewritten
- Keep ( ) and backticks inside word tokens so paths spanning
  \ or wrapped in backticks are matched correctly
- Add _try_rewrite helper with URL detection and unescape logic
- Add 10 contract tests pinning the din0s review cases

* fix(backends): restore () and backtick as shell operators for validate_command

- Restore ( ) and backtick to the operator set in _shell_token_spans.
  Removing them caused a security regression: commands like (sudo ls)
  would not detect sudo as a blocked command because (sudo became one
  word token. With operators restored, validate_command correctly
  catches blocked commands inside subshells and command substitutions.

- Fix _value_span_to_raw_span: the 'quoted' flag from the tokenizer
  means the token *contains* a quoted segment (not necessarily starts
  with a quote). Replace raw[0] assumption with a forward scan for
  the first quote char, consuming unquoted prefix chars 1:1.

- Update test_system_path_with_shell_expansion:  paths are now
  partially rewritten because () are operators. Test updated to
  reflect this known limitation (security >  path rewriting).

* fix(test): cross-platform compatibility for pre-existing Windows failures

- python3 -> python in execute() calls (python is on PATH in any activated venv)
- sleep 10 -> _sleep_cmd(10) cross-platform helper
- str().endswith() -> Path().parts assertions (backslash-safe on Windows)
- shlex.quote exact-match assertions -> 'in' assertions (Windows quotes paths differently)
- mkdir -p E2E test -> preprocessor boundary test
- Skip 3 E2E tests on Windows: shlex.quote produces POSIX quoting incompatible with cmd.exe

141 passed, 3 skipped on Windows.

* fix: update docstring + strengthen shell-expansion test assertion

- Fix _value_span_to_raw_span docstring: no longer assumes raw[0] is
  the opening quote, scans forward for first quote char
- Strengthen test_system_path_with_shell_expansion: verify
  ./workspace/notes is rewritten, not just notes in result

* style: ruff format backends.py + test_backends.py

* refactor(backends): simplify quoted virtual path rewriting

Replace 500+ line shlex tokenizer with 12-line pre-process step. Match quoted args via regex, unescape, rewrite via _rewrite_quoted_path, substitute with shlex.quote. 133 passed, 3 skipped.

* fix: guard bare absolute paths from double-rewrite by post-process regex

On POSIX, shlex.quote returns bare paths (e.g. /tmp/memories/note.md).
The pre-process substitutes these into the command, then the post-process
regex re-matches and incorrectly rewrites them.

Fix: _guard_bare_absolute wraps bare /-paths in single quotes so the
post-process regex''s character class stops at the quote char.

* style: ruff format

* fix(backends): narrow pre-process to exclude system-prefixed paths

Only rewrite quoted paths that are NOT known system prefixes.

* fix: narrow quoted-path pre-process to virtual mounts only

Only rewrite quoted /... paths that resolve to actual virtual mounts (/skills/..., /memories/...) or workspace-prefixed system paths. Remove catch-all that incorrectly rewrote bare paths like echo /hi.

Addresses din0s review feedback on #269.

* docs: update docstring for narrower quoted-path rewrite scope
2026-06-13 17:56:30 +01:00
Xi Zhang 05a5e5f8a0 fix: session lost after evoscientist restart (supersedes #278) (#279)
* fix: session lost after evoscientist restart

* feat: Implement memory worker thread deletion on completion

- Added synchronous and asynchronous functions to delete memory worker threads after they finish execution, ensuring no residual checkpoints are left in the database.
- Enhanced `_watch_memory_worker_run_sync` and `_watch_memory_worker_run_async` to invoke deletion functions upon confirming worker completion.
- Introduced tests to verify that worker threads are deleted correctly upon successful completion and that failures in deletion do not affect the overall worker status.
- Updated session management to ensure that only relevant threads are restored from the database, preventing exposure of internal or unrelated workspace threads.
- Implemented a purge function to clean up leftover worker checkpoints during server startup.

* feat: Implement short thread ID display for CLI and session hints

* fix: ensure proper accounting and deletion order for memory worker threads

---------

Co-authored-by: z00827015 <zhoulun1@huawei.com>
2026-06-11 23:51:51 +01:00
X-iZhang 526c571b10 feat(tunnel): add Cloudflare tunnel support for EvoSci deploy and update documentation 2026-06-11 00:23:55 +01:00
Xi Zhang c02be519f6 feat(dangerous-mode): implement real-filesystem access with safety ch… (#276)
* feat(dangerous-mode): implement real-filesystem access with safety checks

- Introduced a 'dangerous mode' allowing the agent to operate on the real filesystem.
- Updated command validation to bypass path confinement while enforcing a blocklist for privileged commands.
- Added warnings and guidelines for users when operating in dangerous mode.
- Enhanced configuration to support dangerous mode and ensure it implies auto-approval.
- Updated tests to verify the behavior of commands and configurations in dangerous mode.

* feat(dangerous-mode): enhance logging and environment management for dangerous mode

* feat(dangerous-mode): improve handling of dangerous mode with environment flags and enhance test isolation
2026-06-10 18:19:13 +01:00
X-iZhang fc05b5eed2 fix(tests): ensure _check_npx is True in _patch_all_questionary to prevent extra prompts on slow runners 2026-06-10 16:11:49 +01:00
X-iZhang c438966b28 fix(tests): mock _ensure_npx in TestStepSkills to prevent integration issues on headless Windows CI 2026-06-10 15:59:10 +01:00
houren Antony d4f1fbd110 ci: add windows-latest to test matrix + fix 11 cross-platform test bugs (#271)
* ci: add windows-latest to test matrix + fix 11 cross-platform test bugs

The test workflow ran on ``ubuntu-latest`` only. Per the issue's
first bullet — the maintainer's explicit #1 priority — add
``windows-latest`` to the matrix so the manager and related
modules are exercised on Windows on every PR.

The matrix addition surfaces 18 pre-existing Windows-only test
failures. Without fixes the new leg would be 18+ reds from
day one and the matrix would just produce a wall of
``fail-fast`` noise. This PR fixes 11 of them; each fix is
a real (cross-platform) bug, not a Windows-specific hack —
most were already flagged by CodeRabbit on PR #236 but never
acted on. The remaining 4 failures need code refactors
(``os.killpg`` → ``psutil`` in ``background.py``,
``convert_virtual_paths_in_command`` Windows-aware quoting,
tilde expansion) that are documented as out-of-scope
follow-ups below.

## What changed

* ``.github/workflows/test.yml``
  - ``os: [ubuntu-latest, windows-latest]`` → 2 OS × 2 Python
    = 4 cells.
  - ``fail-fast: false`` so one bad cell doesn't cancel the
    rest while the Windows leg is being brought up. Removable
    in a future PR once the suite is fully green.

* ``tests/test_backends.py``
  - Hard-coded ``"python3"`` → ``{sys.executable}`` in 7
    test commands. Windows has no ``python3`` on PATH; using
    ``sys.executable`` is portable and matches what CodeRabbit
    flagged on PR #236.
  - Strict string comparisons → ``shlex.split`` round-trip in
    5 resolver tests. ``shlex.quote`` adds single quotes
    around backslash paths on Windows, which broke the
    direct ``==`` compare.
  - Cross-platform suffix checks in 2 path-resolution tests
    (``Path(resolved).parts[-2:]`` instead of
    ``str(resolved).endswith("src/main.py")``).
  - ``mkdir -p`` → ``sys.executable -c "import os;
    os.makedirs(...)"`` in the cwd-sanitization test.
  - ``skipif(sys.platform == "win32")`` on 3 e2e tests that
    hit the underlying ``shlex.quote`` + ``cmd.exe`` quoting
    bug (real, separate issue).

* ``tests/test_sessions.py``
  - ``test_uses_data_dir``: check ``.evoscientist`` in the
    long path form (via ``Path.resolve()``) rather than the
    short-path form ``get_db_path`` returns on Windows.

* ``tests/test_mcp_client.py``
  - ``endswith("python")`` → ``Path(result).stem.lower()`` so
    ``python.EXE`` matches on Windows.
  - ``endswith("npx")`` also accepts ``npx.cmd`` so the npm
    shim on Windows matches.

## Out of scope (follow-up issues to file)

* ``os.killpg`` doesn't exist on Windows
  (``EvoScientist/background.py:248``) — 3 background tests
  fail. Real fix is the same ``psutil`` walk pattern PR #200
  shipped in ``langgraph_dev/manager.py``.
* Tilde expansion in file mentions.
* Windows-aware shell quoting in
  ``convert_virtual_paths_in_command``.
* Path conventions (``~/.config/evoscientist/`` vs
  ``%APPDATA%\EvoScientist``) — needs design discussion +
  ``platformdirs`` migration.
* Cross-module audit of
  ``EvoScientist/tools/execute.py``,
  ``EvoScientist/ccproxy_manager.py``,
  ``EvoScientist/config/onboard.py``.

Closes #207 (step 1 only — CI matrix + the easy test
fixes; remaining bullets tracked separately).

* fix: cross-platform compatibility for Windows CI runners

- background.py: replace POSIX-only os.killpg/os.getpgid with
  cross-platform _kill_process_tree() helper. On Windows falls back
  to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX
  keeps existing os.killpg logic.

- test_backends.py: replace mkdir -p shell execution in
  test_literal_workspace_path_replaced with preprocessing-boundary
  assertion (patch LocalShellBackend.execute, capture command,
  assert workspace path was rewritten to ./). Avoids POSIX-only
  mkdir -p on Windows runners.

- test_file_mentions.py: monkeypatch USERPROFILE on Windows so
  ntpath.expanduser() resolves ~ to tmp_path even when HOME is
  unset on CI runners.

* fix(test): cross-platform sleep/true commands for Windows CI

Replace POSIX-only sleep/true with module-level helpers that use
ping -n / cmd /c on Windows. Also fix python3 -> sys.executable
in the non-timeout recovery test.

- test_background.py: 7 sleep/true fixes
- test_background_middleware.py: 6 sleep/true fixes
- test_backends.py: 4 sleep fixes + 1 python3 fix

2318 passed, 0 failed on Windows.

* fix(test): use shell-portable double quotes for python -c on Windows

cmd.exe does not treat single quotes as string delimiters, so
-c 'raise SystemExit(1)' was passed with literal quotes on Windows.
Switch to double quotes which work on both cmd.exe and POSIX sh.

* fix: use psutil for Windows process tree kill + avoid sys.executable under uv

- background.py: replace Popen.terminate()/kill() with psutil-based
  process tree walking on Windows. TerminateProcess does NOT cascade
  to grandchildren; psutil.Process.children(recursive=True) ensures
  the entire tree is signaled.

- test_backends.py: replace sys.executable with 'python' in sandbox
  execute() calls. Under uv, sys.executable is under the workspace
  and gets rewritten to ./ by prepare_sandbox_command, breaking
  Linux CI. The plain 'python' command resolves correctly in any
  activated venv.

* fix: broaden try/except in _kill_process_tree to cover proc.children()

If the process exits between Process(popen.pid) and children(recursive=True),
the children call raises an uncaught exception escaping stop(). Move it inside
the existing try/except block.

* fix: narrow exception to ProcessLookupError in POSIX _kill_process_tree

OSError is too broad — would silently swallow EPERM on SIGKILL, leaving
the process alive when we report it as stopped. Match original behavior
which only caught ProcessLookupError (process already gone).

* style: ruff format test_backends.py

* ci: trigger re-run for flaky prompt_toolkit test

* style: fix ruff check (import order + RUF005 unpacking)

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-10 15:43:18 +01:00
X-iZhang cf5e0dd3bd feat(models): add support for 'claude-fable-5' mode 2026-06-10 00:51:29 +01:00
houren Antony 2dc1e227eb fix(langgraph-dev): rotate langgraph_dev.log when it exceeds 50MB (#270)
* fix(langgraph-dev): rotate langgraph_dev.log when it exceeds 50MB

``_LOG_FILE`` (``~/.config/evoscientist/langgraph_dev.log``) was
opened in ``start_langgraph_dev`` with plain ``"ab"`` and never
rotated, so it grew unbounded over weeks/months of heavy use —
especially when chatty MCP servers spawned by langgraph dev
filled it, or when failure paths produced stack traces.

Implement the recommended option 1 from #209: filesize-based
rollover. When the active log exceeds 50MB on the next
``start_langgraph_dev`` invocation, rename it to
``langgraph_dev.log.1`` (overwriting any existing backup) via
``os.replace`` and start fresh. Single-backup policy keeps the
disk footprint bounded at roughly 2x threshold.

Rotation is best-effort: ``_rotate_log_if_needed`` logs and
swallows OSError so a permission error or racing rename can't
block langgraph dev from starting. The next ``start`` invocation
will try again — worst case the log grows for one more session.

Options 2 (timestamped per-session + 7-day sweep) and 3
(``RotatingFileHandler`` + pipe) are explicitly NOT done — option
1 is simplest, no async machinery, matches the issue's
recommendation.

Closes #209

* test(langgraph-dev): redirect _PID_DIR in rotate integration test

Address CodeRabbit review comment on #270: the
``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open``
test patched only ``_LOG_FILE`` to a tmp path, but
``start_langgraph_dev`` also calls ``_PID_DIR.mkdir(...)`` as part
of its prelude, which would create a real directory under
``~/.config/evoscientist/`` on a dev machine. Redirect
``_PID_DIR`` to ``tmp_path / "pids"`` too so the test stays
fully isolated. Add a final assertion that ``pid_dir.is_dir()``
holds, proving the function reached past the mkdir call.

* refactor(langgraph-dev): bundle runtime paths into LanggraphRuntimePaths

@din0s review follow-up on #270: the previous test isolation patched
only ``_LOG_FILE`` (and after a second round, ``_PID_DIR``), but
``start_langgraph_dev`` still touches 5 distinct on-disk paths. Patching
any subset of those still leaves the others pointing at the user's real
``~/.config/evoscientist/`` — exactly the case that produced the
"Port 6174 cannot be bound after waiting 60s" symptom on the
reviewer's machine.

Replace the five free-floating module-level constants
(``_PID_DIR`` / ``_PID_FILE`` / ``_LOG_FILE`` / ``_WORKSPACE_SIDECAR``
/ ``_FILE_LOCK_PATH``) with a single ``LanggraphRuntimePaths`` frozen
dataclass exposed as a module-level ``RUNTIME`` instance. Production
code accesses ``RUNTIME.pid_file`` etc.; tests can now substitute the
*whole* bundle in one assignment:

    monkeypatch.setattr(
        manager, "RUNTIME",
        manager.LanggraphRuntimePaths.for_directory(tmp_path / "runtime"),
    )

The classmethod ``for_directory(pid_dir)`` builds an isolated bundle
rooted at a single dir, so the test author doesn't spell out every
path field. Tests that only care about one field (e.g. pid_file
during the stale-process kill path) use
``dataclasses.replace(manager.RUNTIME, pid_file=X)`` — frozen
dataclass-friendly, no need to enumerate the other four fields.

The dataclass's docstring records the migration rationale (the old
five-name layout invited inconsistent patches).

External callers of the old constants updated:
- ``EvoScientist/deploy/server.py`` and ``webui.py`` now import
  ``RUNTIME`` and use ``RUNTIME.log_file`` for the on-screen log
  path hint. The other imports they had (``_DEFAULT_PORT``,
  ``_is_port_occupied``, ``_read_workspace_sidecar``) are still
  module-level functions/values, untouched.

Test updates:
- ``tests/test_langgraph_manager.py``: ``patch.object(manager, "_XXX",
  X)`` patterns now go through ``dataclasses.replace(manager.RUNTIME,
  xxx=X)``; the ``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open``
  test (from the previous #270 review iteration) uses
  ``for_directory`` for one-shot isolation.
- ``tests/test_langgraph_dev_workspace_sidecar.py``: each test now
  goes through a tiny ``_isolated_runtime(monkeypatch, tmp_path)``
  helper that calls ``for_directory``.
- ``tests/test_langgraph_dev_deploy_mode.py``: same ``for_directory``
  swap.

No production behavior change. All ``langgraph_dev``-side tests
(``test_langgraph_manager.py`` 26/26, ``test_langgraph_dev_workspace_sidecar.py``
14/14, ``test_langgraph_dev_deploy_mode.py`` 14/14, ``test_cli_deploy.py``
18/18 — which indirectly exercises deploy/server.py and deploy/webui.py
imports) pass. Full-project test count unchanged from baseline; the
remaining 22 Windows-only pre-existing failures (test_background
``os.killpg``, test_file_mentions tilde, mcp_client ``shutil.which``,
test_sessions 8.3 short path) are documented as out-of-scope for #207.

* style: apply ruff format to langgraph_dev test + module files

CI lint check on #270 failed:

  Run ruff format --check .
  Would reformat: EvoScientist/langgraph_dev/manager.py
  Would reformat: tests/test_langgraph_manager.py

Plus two test files touched by the prior consolidation commit that
``ruff format`` hadn't seen yet:

  tests/test_langgraph_dev_deploy_mode.py
  tests/test_langgraph_dev_workspace_sidecar.py

Just formatting. No logic change. All 75 refactor-related tests pass.

* fix(test): use for_directory for full path isolation + patch _can_bind_port to skip real socket ops

Two fixes for TestStartLanggraphDevRotatesLog:

1. Replace dataclasses.replace(manager.RUNTIME, ...) with
   LanggraphRuntimePaths.for_directory(pid_dir) so pid_file,
   workspace_sidecar, and lock_file are also temp-rooted
   (prevents leak to ~/.config/evoscientist/).

2. Monkeypatch _can_bind_port to always return True so the
   bind-poll loop in _wait_for_port_bindable passes immediately
   without touching real sockets (fixes 60s timeout on machines
   where port 6174 is already in use).

* fix: cross-platform compatibility for Windows CI runners

- background.py: replace POSIX-only os.killpg/os.getpgid with
  cross-platform _kill_process_tree() helper. On Windows falls back
  to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX
  keeps existing os.killpg logic.

- test_backends.py: replace mkdir -p shell execution in
  test_literal_workspace_path_replaced with preprocessing-boundary
  assertion (patch LocalShellBackend.execute, capture command,
  assert workspace path was rewritten to ./). Avoids POSIX-only
  mkdir -p on Windows runners.

- test_file_mentions.py: monkeypatch USERPROFILE on Windows so
  ntpath.expanduser() resolves ~ to tmp_path even when HOME is
  unset on CI runners.

* refactor(test): add runtime_paths fixture to isolate manager.RUNTIME

Adds a reusable fixture that monkeypatches manager.RUNTIME to a
temp-rooted LanggraphRuntimePaths.for_directory(). Tests that need
specific fields can still dataclasses.replace(runtime_paths, ...)
but the baseline is always temp-isolated, preventing leaks to
~/.config/evoscientist/.

Updated test_langgraph_dev_deploy_mode.py, test_langgraph_dev_workspace_sidecar.py,
and test_langgraph_manager.py to use the fixture, consolidating sequential
lock_file + pid_dir patches into single dataclasses.replace calls.

* Revert "fix: cross-platform compatibility for Windows CI runners"

This reverts commit eb025d24af32e195a982cd40f6d70dba885c4019.

* style: ruff format conftest.py

* fix: address review issues in log-rotation + runtime paths

- Use for_directory(tmp_path/pids) as base in ensure_langgraph_dev tests
  so pid_file/log_file are co-located with pid_dir, not split across paths
- Remove unused runtime_paths param from test_no_existing_file_is_noop
- Replace manager.RUNTIME with runtime_paths in two sidecar tests
- Use for_directory(DEFAULT_PID_DIR) instead of explicit construction
- Fix stale _LOG_FILE reference in TestRotateLogIfNeeded docstring

* style: ruff format test files

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-09 15:55:44 +01:00
dinos cd2baa9588 feat(models): add opt-in prompt caching support for anthropic via openrouter (#272)
* feat(models): add opt-in prompt caching support for anthropic via openrouter

* chore: don't coerce model_kwargs to dict
2026-06-09 15:13:49 +01:00
Ziheng Zhang 4b6a969df2 refactor(agent): make create_cli_agent(config=, chat_model=) pure (#267)
* refactor(agent): make create_cli_agent(config=, chat_model=) pure

Re-applies the #183 purity refactor on top of the observation-memory
lifecycle that landed in #259, integrating the two cleanly.

create_cli_agent gains a pure path: when both `config` and `chat_model`
are passed it builds the agent entirely from locals and writes none of
the cached module globals (`_config`, `_chat_model`, `_chat_model_key`,
`_EvoScientist_agent`). `/model` commits the switch via
`set_active_config` / `set_chat_model_instance` only after a successful
build, so a failed rebuild leaves the session on the original model
(replaces the old snapshot/restore rollback).

Supporting changes:
- Extract `set_active_config` (write-half of `_ensure_config`),
  `_apply_env_from_config`, `_build_chat_model`, and
  `set_chat_model_instance`.
- Thread `cfg` / `chat_model` through `_get_default_middleware`,
  `_build_base_kwargs`, `load_mcp_and_build_kwargs`,
  `_maybe_swap_async_subagents`, and `_inject_subagent_middleware` so the
  pure path never falls back to the global-writing `_ensure_config()` /
  `_ensure_chat_model()`.
- Integrate with #259's memory middleware: subagent context-editing
  middleware binds the threaded `chat_model`, and the configured system
  prompt / memory controls read the threaded `cfg` (new threading vs the
  original #183, required because #259 made these paths read config).
- Consolidate `cfg` resolution to one `cfg if cfg is not None else
  _ensure_config()` at the top of each kwargs builder, matching the
  pattern already used in the other config-aware helpers.

* fix(agent): keep pure tool selector off global cache

* fix(model): apply config switch in place to preserve reference integrity

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-06-08 18:38:13 +01:00
dinos 8bb1d6c0e3 refactor(stream): langgraph streaming v3 (#268)
* refactor(stream): langgraph streaming v3

* fix: address CR comments

* chore(stream): add success field to state

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-08 17:14:25 +01:00
Xi Zhang 63969b596d Release/v0.1.4 (#266)
* feat(middleware): reposition code interpreter middleware in the stack

* feat(models): add qwen3.7-plus model entry and update context window comment

* feat(models): add qwen3.7-max and qwen3.7-plus model entries for DashScope

* feat(auxiliary): implement auxiliary model support for background tasks and tool selection

- Added auxiliary model configuration to EvoScientistConfig.
- Introduced _ensure_auxiliary_chat_model function to manage auxiliary model instances.
- Updated onboarding steps to include auxiliary model selection.
- Modified middleware to route tool selection to the auxiliary model when applicable.
- Enhanced tests to cover auxiliary model functionality and configuration.

* feat(steps): update UI backend selection options and descriptions

* Refactor code structure for improved readability and maintainability

* feat(patches): implement OpenRouter response reasoning item stripping to prevent multi-turn errors

* feat: update version to v0.1.4 in badges, README, and pyproject.toml; adjust skill counts in steps.py

* feat(config): add auxiliary model and provider environment variables to test setup
2026-06-07 00:52:59 +01:00