- Introduced TodoListMiddleware to the middleware stack for better task management.
- Updated HITL interrupt configuration to include 'delete' operations requiring approval.
- Implemented error handling for delete operations in read-only and memory backends.
- Enhanced approval prompt formatting to display file paths for delete actions.
- Added tests to ensure delete operations are correctly blocked or prompted for approval.
- Updated dependencies to use deepagents 0.7.0 and langchain 1.5.3 for improved functionality.
* feat(middleware): add ToolHistoryRepairMiddleware and enhance tool call validation
* fix(tests): add test for dropping non-list raw tool calls in repair_tool_history
* Add Requesty as an LLM provider
* Address review: Requesty prompt caching, model ordering, key validation
- Declare Anthropic-style prompt caching for Requesty Claude models by
default (mirroring the OpenRouter behavior), with an opt-out flag
EVOSCIENTIST_REQUESTY_ANTHROPIC_PROMPT_CACHE. Requesty is an OpenAI-routed
provider, so the caching check now uses the original provider name.
- Move the Requesty model entries above OpenRouter so Requesty no longer
overrides native/OpenRouter models for names it shares with them
(the MODELS dict is last-entry-wins); drop the outdated gpt-4o-mini entry.
- Fix validate_requesty_key: Requesty's /v1/models returns 200 even for an
invalid/missing key (public catalog), so it cannot validate a key. Use a
minimal authenticated /v1/chat/completions request instead (200 = valid,
403 = invalid), verified against the live endpoint.
- Add tests for Requesty prompt caching (default on, opt-out, non-Anthropic skip).
* Validate Requesty key against auth layer, not a specific model
The onboarding validator probed /v1/chat/completions with a hardcoded
real model (openai/gpt-4o-mini), which tied key validation to that model
staying available upstream. The router resolves auth before the model, so
probe a deliberately nonexistent sentinel model (requesty/auth-preflight)
instead: a valid key yields 404 (model-not-found, auth passed), an invalid
key yields 401/403, and 429/5xx stay inconclusive so a transient outage
does not reject a good key. Add unit tests covering each case.
---------
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* fix: propagate langgraph dev bind port into subprocess env for self-loop URL
* fix: keep parent env authoritative over workspace .env for mapped keys
* fix: limit .env shadow-guard to EVOSCIENTIST_* keys so API keys keep .env-wins
* fix: snapshot EVOSCIENTIST_* env by prefix instead of filtering _ENV_MAPPINGS
* fix: merge .env via dotenv_values to close empty-value and RMW-race edges
* chore: align docstrings after .env-merge rework
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: add support for new Anthropic models and enhance adaptive thinking tests
* fix: implement patches for Anthropic protocol to handle foreign reasoning blocks and structured output for mandatory-thinking Kimi models
* fix: update version to v0.2.4 in badges, README, and project files
* fix: update Star History chart links in README and README.zh-CN
* fix: add support for Gemini 3.6 Flash and 3.5 Flash Lite models in model entries and update changelog
* fix: update wechat group image in assets
* fix: set langgraph and codex proxy runtime defaults
* fix: address runtime default review feedback
* fix: drop langgraph dev env defaults per maintainer review
langgraph dev patches DATABASE_URI/REDIS_URI itself via patch_environment,
so the reported KeyError cannot come from this flow; the env defaults added
here were unnecessary. Scope the PR back to the codex proxy reasoning
context fix only.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: log missing async-subagent tools at DEBUG, not WARNING
* fix: distinguish load_subagents callers via async_swap_pending flag
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: scrub host path from skill_manager output and guard batch install
* test: tighten install leak guards to catch host path in either tier
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: repair interrupted tool call history
Normalize incomplete tool exchanges before model calls so strict providers do not reject resumed sessions. Preserve completed exchanges and cover sync and async model paths.
* fix: repair malformed tool calls and dedupe repair warnings
Track AIMessage.invalid_tool_calls alongside tool_calls so interrupted
threads with syntactically invalid tool calls get synthesized error
results and are accepted by strict providers.
Preserve the originating tool call's name in the synthesized ToolMessage,
and deduplicate repair warnings per unique tool-call id via a warned set
owned by the middleware instance, since the middleware rewrites the
request but not thread state.
Document the middleware's scope versus deepagents' PatchToolCallsMiddleware
(orphan ToolMessage dropping and mid-run coverage).
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat(context-window): add Kimi K3 model with 1M context window
* feat(openrouter): implement structured output for Kimi K3 and add 429 retry handling
* Refactor code structure for improved readability and maintainability
* feat: add disable_streaming helper for tool-selector's internal model
* feat: apply disable_streaming to the tool-selector's model in the factory
* fix(tool-selector): hide selector model call from public event streams
* feat(tool-selector): log a WARNING when the selector's model returns a duplicate-tool_calls flood
* chore: log flood-detector errors, document parent-method drift risk, tighten tests
* refactor(tool-selector): switch to nostream tag via model-field wiring, drop subclass
* feat(tui): implement completion popup rendering and windowing logic
* feat(tui): enhance completion popup with dynamic row budgeting and CSS adjustments
* Refactor picker widgets to use shared base class for improved code reuse
- Introduced `picker_base.py` to encapsulate common functionality for picker widgets.
- Updated `ModelPickerWidget`, `SkillBrowserWidget`, and `ThreadPickerWidget` to inherit from `PickerWidgetBase`.
- Implemented selection helpers (`first_selectable_index`, `move_selection`) in `picker_base.py` for consistent item navigation.
- Refactored rendering and selection logic in each widget to utilize the new base class methods.
- Added tests for picker functionality to ensure behavior remains consistent post-refactor.
* feat: implement bulk cancellation of non-terminal runs before thread deletion
* test: enhance thread cancellation tests and add fake restore for orphaned runs sweep
* feat: enhance run cancellation logic to support status filtering during thread deletion
* feat: add langgraph-sdk dependency for enhanced functionality
* chore: add pytest-asyncio in auto mode
* test: migrate channel and stream tests to native async
Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.
* test: migrate command and model/middleware tests to native async
Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.
* test: migrate TUI, notifier, gateway, and session tests to native async
TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.
* test: replace direct asyncio.run() calls with native async tests
Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.
* test: drop undeclared anyio markers and delete run_async helper
The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).
* test: add autouse fixture for watcher cleanup
* refactor: remove redundant hasattr calls
* refactor: add typed middleware event sink and thread through assembly
Add MiddlewareEventSink protocol + NoOpSink in middleware/events.py
with a documented any-thread non-blocking contract (contract test uses a
deliberately-slow fake sink). Thread an optional `events` parameter
through create_cli_agent -> _get_default_middleware -> tool selector /
model fallback constructors; subagent stacks are always forced to
NoOpSink.
* refactor: inject a notifier port into async-watcher and background middleware
Add public pre_cancel_watcher() and enqueue_task_notification() to
cli/async_notifier.py and a small NotifierPort protocol
(middleware/notifier.py) that the module satisfies structurally.
AsyncWatcherMiddleware and BackgroundExecutionMiddleware now receive the
port by constructor injection at the composition root, deleting the lazy
'from ..cli import async_notifier' imports and the private
_watcher_by_thread / _enqueue pokes.
* refactor: invert tool-selection ownership onto a frontend event sink
The adaptive tool selector now reports on_tool_selection_started /
on_tool_selection / on_tool_selection_ended to the injected sink instead
of writing four process-global module variables. The frontend sink
(stream/sink.py FrontendEventSink) owns the selected/total/active state
with consume-once + dedup-vs-last-emitted semantics;
stream/tool_selection.py reads that sink object (a ToolSelectionView)
rather than reaching into tool_selector's globals.
Deleted: the 4 module globals, the cross-module mutations in
tool_selection.py, the track_stream_selection flag, the now-vestigial
_ToolSelectionTrackerMiddleware, reset_tool_selection_state_for_tests,
and the autouse conftest fixture. The sink is threaded from the two
interactive frontends through create_runtime_gateways ->
LocalGraphGateway (read side) and _load_agent -> create_cli_agent (write
side); subagent / headless stacks get NoOpSink.
* refactor: route model-fallback narration through the injected event sink
Delete the _ui_emit_fn / set_ui_emit module global and the
..stream.console import from model_fallback.py. The fallback middleware
now reports through its injected sink: the fallback transition via the
structured on_model_fallback (the frontend formats the '-> Falling back
to ...' line), and the surrounding narration (primary-failure header,
per-attempt outcome, exhaustion, non-fallbackable rejection) via
emit_fallback_notice, preserving the exact user-facing text. The TUI
binds its _append_system as the sink's fallback display where it used to
call set_ui_emit (cleared on exit); the Rich CLI's sink prints to the
console. _try_fallbacks / _guard_and_fallback take the sink.
* refactor: declare events on the GraphGateway protocol
Both gateway implementations now carry an explicit events attribute
(LangGraphServerGateway holds None — no frontend renders middleware
events across the HTTP boundary), so the four call sites use plain
attribute access instead of getattr probing an implicit contract.
* refactor: bind fallback display via the closure-scoped concrete sink
The App methods used gateway.events (typed as the read-side view) and
hasattr-probed for the concrete FrontendEventSink API. The enclosing
factory creates that sink two hundred lines up — close over it directly:
no probing, fully typed, and it becomes a constructor parameter
naturally when the App class is hoisted out of the factory.
* fix: end tool selection before fallback handler
* fix: keep fallback display errors non-fatal
* fix: preserve selector suppression for default streams
* fix: restore fallback notice console display
* refactor: consolidate fallback narration events
* refactor: clean middleware event sink plumbing
* fix: type gateway session events
* refactor: make all event protocols runtime-checkable
MiddlewareEventSink already carried @runtime_checkable (the stream
binding guard isinstance-checks it); ToolSelectionView and SessionEvents
now match, so mirroring that pattern against any of the three protocols
works instead of raising TypeError.
* fix(cli): close QuickJS workers after one-shot failures
* fix(cli): honor no-thinking in final output
* fix(channels): report failed startup accurately
* fix(channels): make Telegram cleanup idempotent
* fix(tui): skip command sync during exit
* fix(channels): preserve startup state during retries
* refactor(channels): share pending startup status
* refactor(cli): expose channel startup snapshot
* fix(tui): move channel startup off event loop
* test(channels): release retry gate on assertion failure
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* chore: add pytest-asyncio in auto mode
* test: migrate channel and stream tests to native async
Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.
* test: migrate command and model/middleware tests to native async
Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.
* test: migrate TUI, notifier, gateway, and session tests to native async
TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.
* test: replace direct asyncio.run() calls with native async tests
Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.
* test: drop undeclared anyio markers and delete run_async helper
The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).
* test: add autouse fixture for watcher cleanup
* refactor: remove redundant hasattr calls
* refactor: extract shared HITL/ask_user interaction grammar
Extract prompt/question formatting, the reply grammar (approval letters,
ask_user choice letters + the 'Other' sub-flow, stop-commands), the
ApprovalPolicy (config auto-approve rule + session registry +
session-key derivation), per-flow timeout constants, and the bilingual
feedback strings into channels/interaction.py. Both drivers now point at
the shared functions: this reverses cli/channel.py's imports of consumer
privates and closes the /stop drift at the parsing layer (serve-mode
ask_user now checks stop-commands before parsing an answer, matching the
CLI path).
* refactor: add interaction engine + registry; port InboundConsumer
Introduce InteractionIO (transport adapter Protocol),
PendingReplyRegistry (one asyncio-based reply router per process), and
the engine coroutines resolve_ask_user / resolve_approval in
channels/interaction.py. Port InboundConsumer onto them: a _ConsumerIO
adapter over bus.publish_outbound + the registry, one ApprovalPolicy
replacing the config/session auto-approve checks, and a single
reply-interception point (registry.try_resolve) replacing the parallel
ask_user/HITL pending dicts. _resolve_ask_user and the approval section
of _stream_with_hitl are now thin engine calls.
Behavior unification (serve mode): an unrecognized HITL reply now
declines with the shared 'Unrecognized reply' notice instead of
rejecting-and-refeeding as a fresh turn, and /stop mid-approval cancels
cleanly — both via the shared parser.
* refactor: port CLI channel bridge onto the interaction engine
Replace the ~250-line parallel bodies of channel_ask_user_prompt /
channel_hitl_prompt with thin bridges that run resolve_ask_user /
resolve_approval on the bus loop via
run_coroutine_threadsafe(...).result() (outer = engine per-flow timeout
+ slack, so the engine's own timeout fires first). The 15s send timeout
moves into the _BridgeIO adapter.
Delete the _pending_hitl / _hitl_lock / _hitl_auto_approve module
globals and the _register_hitl_wait / _try_set_hitl_reply /
_pop_hitl_reply helpers, absorbed by one bus-loop PendingReplyRegistry +
one ApprovalPolicy. The bus consumer feeds the registry via try_resolve
ahead of normal enqueue.
* refactor: restore serve-mode refeed for unrecognized HITL replies
Gate-review fix: the engine no longer decides transport policy for
unparseable approval replies. resolve_approval now returns an
ApprovalOutcome carrying unrecognized_reply (raw text) when parsing
fails, sending no feedback itself; recognized reject keeps the sharedi
rejection message.
Consumer driver (serve mode) restores the pre-engine semantics: an
unrecognized reply rejects the pending action, confirms with the
rejection message, and the text is re-dispatched as a NEW agent turn —
_stream_with_hitl returns the captured text and _handle_message starts
the refeed turn only after the current one has released the chat lock
(old fall-through ordering). CLI bridge keeps its old no-refeed path
byte-for-byte: 'Unrecognized reply. Action rejected.' and decline.
Tests: serve refeed pinned end-to-end (prompt → unrecognized text →
rejection feedback → text reaches the stream path as a new turn), CLI
no-refeed pinned (notice sent, nothing enqueued), engine test updated to
assert the outcome struct with no engine-side feedback.
* refactor: polish the interaction engine surface
- English feedback strings (Approved / Rejected / auto-approving)
- drop the consumer's backwards-compatible re-exports and both modules'
private timeout aliases; callers use the canonical interaction names
- replace byte-for-byte prompt goldens with structural format tests and
assert feedback via the shared constants instead of string literals
- strip audit/design shorthand (R1/R2/G3, stage numbers) from comments
* fix: propagate pending reply task cancellation
* fix: preserve reply context when refeeding HITL replies
* fix: honor HITL session grants without bus loop
* fix: bound bridge waits by send latency
* fix: handle empty ask_user replies explicitly
* chore: remove stale interaction helpers
* fix: harden interaction engine reply edge cases
Review follow-ups on the interaction engine:
- normalize ask_user choices before .get(): the tool args come from model
JSON and only presence is validated, so plain-string choices must render
and parse instead of crashing the turn
- treat only None as an approval timeout, so an empty/media-only reply
flows through the unrecognized path and serve mode refeeds it with its
preserved context
- intercept prompt replies before _get_thread_id so a consumed reply
cannot create an orphan graph thread or touch the sender-session LRU
- close engine coroutines the bridge failed to schedule (no bus loop /
scheduling error) to avoid never-awaited warnings
- clear pending-response and channel-request state in the bridge test
fixture
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix: surface real exception class+message in SSE error events
* fix: tighten SSE error patch scope and key redaction
* fix: redact base64-style secret suffixes fully
* style: remove notes/ reference from the dosctring
* fix: rebuild env cache on each error call
* fix: route BaseException through serde.default on SSE/webhook paths
* fix: distinguish routed providers by request URL host
* feat: normalize provider-SDK exceptions via ErrorNormalizationMiddleware
* refactor: drop json_dumpb dataclass-bypass wrappers, superseded by middleware
* fix: guard _extract_host against SDK properties that raise
* refactor: derive provider tag from ModelRequest.model, not the exception
* refactor: drop serde.default patch and exception-based inference; ProviderStreamError.model_dump handles the emit
* refactor: move envelope helpers from patches.py to errors.py
* feat: extend ErrorNormalizationMiddleware coverage to every model-call path
* chore: clean up review findings from middleware pivot
* fix: pass through all langgraph.errors
* fix: move langgraph.errors pass-through into _normalize
* fix: pass through ContextOverflowError in _normalize
* fix(llm): make gpt-5.x usable through ccproxy Codex OAuth
Two independent blockers made current OpenAI models fail when routed
through ccproxy's Codex OAuth endpoint:
1. ccproxy's default Codex model mappings rewrite any gpt-*/o1-*/o3-*/
claude-* model to gpt-5.3-codex before forwarding, silently overriding
the configured model and failing outright on accounts where
gpt-5.3-codex is not served ("The 'gpt-5.3-codex' model is not
supported when using Codex with a ChatGPT account").
start_ccproxy() now generates a config with empty codex model
mappings and passes it via 'ccproxy serve --config'.
2. ccproxy forwards the client's own User-Agent upstream and only
gap-fills its Codex headers, so the backend gates current models on
the client identity ("The '<model>' model requires a newer version
of Codex"). get_chat_model() now sends Codex-CLI-shaped
originator/version/User-Agent headers when the ccproxy Codex adapter
is detected, overridable via EVOSCIENTIST_CODEX_CLIENT_VERSION.
Verified live: gpt-5.5 and gpt-5.4 complete successfully through
ccproxy Codex OAuth on a ChatGPT Plus account with both fixes; each
fails without them.
* fix(ccproxy): harden Codex client routing
* fix(llm): keep Codex client identity consistent
* docs: clarify Codex version floor
* style: ruff format models.py after merge
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* fix(llm): respect reasoning_effort setting on native OpenAI path
The native OpenAI provider path hardcoded reasoning effort to xhigh for
gpt-5.4/5.5/codex models, silently ignoring the user's reasoning_effort
config setting. The OpenRouter path already honors the
EVOSCIENTIST_REASONING_EFFORT env var that settings.py exports from that
setting; this applies the same lookup on the native path, falling back
to the previous defaults when unset.
Adds a regression test and isolates the existing xhigh test from the
env var.
* fix(llm): preserve model reasoning defaults
* fix(llm): preserve GPT-5.6 reasoning default
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix(ccproxy): raise auth status check timeout to 30s
ccproxy's CLI initializes its full plugin system on every invocation;
a cold 'ccproxy auth status' takes ~10s wall time on Apple Silicon,
so the 10s subprocess timeout made OAuth startup fail intermittently
with 'Auth check timed out' even when credentials were valid.
* fix(ccproxy): raise serve health deadline to 120s
ccproxy boot includes plugin init plus Codex CLI detection; measured
~76s to first healthy response on an Apple Silicon Mac (ccproxy-api
0.2.9). The 30s deadline in start_ccproxy() killed the process before
it could come up, failing OAuth startup with 'ccproxy did not become
healthy within 30 seconds'.
* fix(ccproxy): widen serve health deadline to 180s
Full startup measured at ~111s on a second cold run (Apple Silicon,
ccproxy-api 0.2.9); 120s left too little headroom for boot variance.
* fix(ccproxy): centralize startup timeouts
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
get_effective_config() runs load_dotenv(find_dotenv(usecwd=True),
override=True), so any test that loads config injected the repo's real
.env into os.environ for the rest of the pytest process. An
empty-valued line like MINIMAX_BASE_URL= then made
os.environ.get(key, default) return '' instead of the default,
failing the MiniMax routing tests in full-suite runs while they
passed in isolation.
Generalizes the find_dotenv redirect that test_config.py's
temp_config_dir fixture already applied locally into a suite-wide
autouse fixture, pointing at a never-created path so tests writing
their own tmp_path/.env cannot collide with it. Adds a regression
test reproducing the leak.
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat(llm): add OpenRouter app attribution headers (#339)
Attach EvoScientist app-attribution at the shared model-init layer so all
OpenRouter calls are credited to the project. langchain-openrouter maps
app_url/app_title/app_categories -> HTTP-Referer / X-Title /
X-OpenRouter-Categories. Applied only for the openrouter provider, via
setdefault so explicit caller kwargs win. Configurable through new
openrouter_http_referer / openrouter_app_title / openrouter_app_categories
settings and their EVOSCIENTIST_OPENROUTER_* env vars.
Closes#339
* refactor(llm): centralize OpenRouter attribution defaults + cap categories
Address PR #344 review:
- Define the app-attribution default constants once in config/settings.py
(the config fields and llm/models.py both use them) instead of duplicating
the literals across the two modules.
- Reduce the default categories to creative-writing,personal-agent and cap the
sent list to OpenRouter's 2-per-request limit, warning when a configured list
exceeds it, so extras are dropped predictably (and surfaced) here rather than
being silently truncated server-side.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* chore: add pytest-asyncio in auto mode
* test: migrate channel and stream tests to native async
Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.
* test: migrate command and model/middleware tests to native async
Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.
* test: migrate TUI, notifier, gateway, and session tests to native async
TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.
* test: replace direct asyncio.run() calls with native async tests
Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.
* test: drop undeclared anyio markers and delete run_async helper
The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).
* refactor(onboard): shared flow for ccproxy providers
* feat(onboard): support oauth configuration for auxiliary models
* fix(onboard): reuse main model auth for same-provider auxiliary
* fix(onboard): reconcile oauth providers
* feat(cli): add --output-format stream-json for headless clients
Emit EvoScientist's native event stream as line-delimited JSON on stdout
in single-shot (-p) mode, with all human output redirected to stderr so
stdout stays pure JSONL. Intended as the integration surface for
programmatic clients (e.g. an agent runtime) that drive EvoSci headlessly.
- stream/json_sink.py: write_events_as_json + stream_json sink, plus
redirect_console_to_stderr helper for stdout purity
- cli/interactive.py: cmd_run gains output_format; stream-json branch runs
the sink instead of the Rich renderer
- cli/commands.py: --output-format option + validation (stream-json
requires -p; value must be text|stream-json)
- docs/stream-json.md: event-schema contract + example transcript
- tests: json sink serialization, CLI dispatch, console redirect, validation
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(cli): honor explicit --no-auto-mode over config in stream-json
Address CodeRabbit review (discussion_r3514041123): the auto-mode override
block only wrote to cli_overrides when the resolved value was True, so an
explicit --no-auto-mode silently fell back to a config that enables
auto-mode -- breaking "explicit flags always win" and leaving stream-json
running unattended despite the warning. Write auto_mode=False when the flag
is explicitly False. Add regression tests that capture the overrides.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(tui): keep welcome banner at top after /new
PR #262 replaced scroll_end() with anchor() for free-scrolling.
When /new clears a long anchored conversation, the anchor kept
the viewport pinned to the (now empty) bottom, producing a
negative scroll_y and pushing the welcome banner out of view.
Reset the anchor and scroll to the top in clear_chat(), and
restore the follow/new-content flags so the fresh session starts
correctly.
Closes#301
* fix(tui): suppress anchor when chat content fits viewport
The previous fix for #301 only handled the /new path. din0s reported
that the banner still dropped to the bottom after a normal short turn
(user types 'hi', agent replies) — i.e. whenever the conversation
fit in the viewport. Root cause is in Textual's compositor
(textual._compositor): when a widget is anchored, scroll_y is
recomputed via set_reactive, which bypasses the validator. If the
anchored widget's content is shorter than the viewport, scroll_y
goes negative on the next layout pass and the welcome banner is
pushed below the visible region.
PR #262 made _stream_with_widgets re-engage the anchor at the end of
every turn via _anchor_chat, so the bug surfaced on any short reply
that fit in the viewport. Markdown re-renders, status-bar updates,
or any subsequent mount would then trip the compositor.
Fix in three places:
* _anchor_chat: only engage the anchor when max_scroll_y > 0;
otherwise release and scroll_home so the banner stays at the top.
* streaming anchor loop: if content shrinks below the viewport
mid-stream (e.g. loading widget removed), release the anchor
instead of leaving _anchored=True for the compositor to trip on.
* clear_chat: keep the unconditional reset (children are removed
asynchronously so a max_scroll_y check would be stale) but
document why.
Adds two regressions:
* test_short_turn_keeps_banner_at_top_after_layout_refresh — the
exact scenario din0s tested; fails with scroll_y=-10 on the
previous code, passes with the fix.
* test_long_turn_keeps_viewport_pinned_to_bottom — guards against
regressing free-scrolling for overflowing conversations.
Manually verified: 'hi' -> reply (banner stays at top) -> /new
(banner at top) -> another turn (banner stays at top).
* test(tui): address review feedback on banner-position regressions
- extract `_release_anchor_and_pin_top` helper for the 3-line
`anchor(False) + scroll_home(...)` pattern repeated in
`clear_chat`, `_anchor_chat`, and the streaming loop
- replace `pytest.skip` in `_capture_app` with a hard `RuntimeError`
so a broken capture never silently passes
- drop the redundant `load_agent` and `create_session_workspace`
monkeypatches (the factory is given those as parameters, so the
module-level symbols never run; added a comment explaining why)
- add a defensive `_FakeChannelRuntime` patch for symmetry with the
other module-level fakes
- drop the local `_run` helper and use the `run_async` fixture from
`conftest.py` (its teardown is better)
* test(tui): replace _FakeChannelRuntime with _auto_start_channel no-op
The _FakeChannelRuntime patch was ineffective because ChannelRuntime is
just a dataclass — the real channel manager still started via
_auto_start_channel, leaving pending tasks and non-hermetic test state.
Per review feedback, stub _auto_start_channel directly instead.