- Gateway internal identity: when a service token is configured, reject
wrong/missing tokens even from loopback (closes SSRF/local bypass).
- Terminal metering: classified AgentControlError propagates without
retry; exhausted retries raise BILLING_UNAVAILABLE instead of a
generic RuntimeError, keeping error attribution accurate.
Standardize the stream delta on LangChain's own message dict (delta.message)
instead of a bespoke {text, tool_calls} shape, and raise when the stream ends
without the [DONE] sentinel so mid-stream truncation is no longer silent.
The legacy BaseChatModel.astream path does not forward run_manager to
_astream, so the per-call tracing run_id was unreachable and the stream
fell back to the conversation run_id, which the gateway's
verify_model_attempt rejected (401 RUN_ATTEMPT_NOT_ACCEPTED). Publish the
tracing run_id into the shared configurable dict from
on_chat_model_start and read it back in _attempt_id.
* fix: surface real exception class+message in SSE error events
* fix: tighten SSE error patch scope and key redaction
* fix: redact base64-style secret suffixes fully
* style: remove notes/ reference from the dosctring
* fix: rebuild env cache on each error call
* fix: route BaseException through serde.default on SSE/webhook paths
* fix: distinguish routed providers by request URL host
* feat: normalize provider-SDK exceptions via ErrorNormalizationMiddleware
* refactor: drop json_dumpb dataclass-bypass wrappers, superseded by middleware
* fix: guard _extract_host against SDK properties that raise
* refactor: derive provider tag from ModelRequest.model, not the exception
* refactor: drop serde.default patch and exception-based inference; ProviderStreamError.model_dump handles the emit
* refactor: move envelope helpers from patches.py to errors.py
* feat: extend ErrorNormalizationMiddleware coverage to every model-call path
* chore: clean up review findings from middleware pivot
* fix: pass through all langgraph.errors
* fix: move langgraph.errors pass-through into _normalize
* fix: pass through ContextOverflowError in _normalize
* fix(llm): make gpt-5.x usable through ccproxy Codex OAuth
Two independent blockers made current OpenAI models fail when routed
through ccproxy's Codex OAuth endpoint:
1. ccproxy's default Codex model mappings rewrite any gpt-*/o1-*/o3-*/
claude-* model to gpt-5.3-codex before forwarding, silently overriding
the configured model and failing outright on accounts where
gpt-5.3-codex is not served ("The 'gpt-5.3-codex' model is not
supported when using Codex with a ChatGPT account").
start_ccproxy() now generates a config with empty codex model
mappings and passes it via 'ccproxy serve --config'.
2. ccproxy forwards the client's own User-Agent upstream and only
gap-fills its Codex headers, so the backend gates current models on
the client identity ("The '<model>' model requires a newer version
of Codex"). get_chat_model() now sends Codex-CLI-shaped
originator/version/User-Agent headers when the ccproxy Codex adapter
is detected, overridable via EVOSCIENTIST_CODEX_CLIENT_VERSION.
Verified live: gpt-5.5 and gpt-5.4 complete successfully through
ccproxy Codex OAuth on a ChatGPT Plus account with both fixes; each
fails without them.
* fix(ccproxy): harden Codex client routing
* fix(llm): keep Codex client identity consistent
* docs: clarify Codex version floor
* style: ruff format models.py after merge
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
* fix(llm): respect reasoning_effort setting on native OpenAI path
The native OpenAI provider path hardcoded reasoning effort to xhigh for
gpt-5.4/5.5/codex models, silently ignoring the user's reasoning_effort
config setting. The OpenRouter path already honors the
EVOSCIENTIST_REASONING_EFFORT env var that settings.py exports from that
setting; this applies the same lookup on the native path, falling back
to the previous defaults when unset.
Adds a regression test and isolates the existing xhigh test from the
env var.
* fix(llm): preserve model reasoning defaults
* fix(llm): preserve GPT-5.6 reasoning default
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* fix(ccproxy): raise auth status check timeout to 30s
ccproxy's CLI initializes its full plugin system on every invocation;
a cold 'ccproxy auth status' takes ~10s wall time on Apple Silicon,
so the 10s subprocess timeout made OAuth startup fail intermittently
with 'Auth check timed out' even when credentials were valid.
* fix(ccproxy): raise serve health deadline to 120s
ccproxy boot includes plugin init plus Codex CLI detection; measured
~76s to first healthy response on an Apple Silicon Mac (ccproxy-api
0.2.9). The 30s deadline in start_ccproxy() killed the process before
it could come up, failing OAuth startup with 'ccproxy did not become
healthy within 30 seconds'.
* fix(ccproxy): widen serve health deadline to 180s
Full startup measured at ~111s on a second cold run (Apple Silicon,
ccproxy-api 0.2.9); 120s left too little headroom for boot variance.
* fix(ccproxy): centralize startup timeouts
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
get_effective_config() runs load_dotenv(find_dotenv(usecwd=True),
override=True), so any test that loads config injected the repo's real
.env into os.environ for the rest of the pytest process. An
empty-valued line like MINIMAX_BASE_URL= then made
os.environ.get(key, default) return '' instead of the default,
failing the MiniMax routing tests in full-suite runs while they
passed in isolation.
Generalizes the find_dotenv redirect that test_config.py's
temp_config_dir fixture already applied locally into a suite-wide
autouse fixture, pointing at a never-created path so tests writing
their own tmp_path/.env cannot collide with it. Adds a regression
test reproducing the leak.
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat(llm): add OpenRouter app attribution headers (#339)
Attach EvoScientist app-attribution at the shared model-init layer so all
OpenRouter calls are credited to the project. langchain-openrouter maps
app_url/app_title/app_categories -> HTTP-Referer / X-Title /
X-OpenRouter-Categories. Applied only for the openrouter provider, via
setdefault so explicit caller kwargs win. Configurable through new
openrouter_http_referer / openrouter_app_title / openrouter_app_categories
settings and their EVOSCIENTIST_OPENROUTER_* env vars.
Closes#339
* refactor(llm): centralize OpenRouter attribution defaults + cap categories
Address PR #344 review:
- Define the app-attribution default constants once in config/settings.py
(the config fields and llm/models.py both use them) instead of duplicating
the literals across the two modules.
- Reduce the default categories to creative-writing,personal-agent and cap the
sent list to OpenRouter's 2-per-request limit, warning when a configured list
exceeds it, so extras are dropped predictably (and surfaced) here rather than
being silently truncated server-side.
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* chore: add pytest-asyncio in auto mode
* test: migrate channel and stream tests to native async
Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.
* test: migrate command and model/middleware tests to native async
Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.
* test: migrate TUI, notifier, gateway, and session tests to native async
TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.
* test: replace direct asyncio.run() calls with native async tests
Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.
* test: drop undeclared anyio markers and delete run_async helper
The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).
* refactor(onboard): shared flow for ccproxy providers
* feat(onboard): support oauth configuration for auxiliary models
* fix(onboard): reuse main model auth for same-provider auxiliary
* fix(onboard): reconcile oauth providers
* feat(cli): add --output-format stream-json for headless clients
Emit EvoScientist's native event stream as line-delimited JSON on stdout
in single-shot (-p) mode, with all human output redirected to stderr so
stdout stays pure JSONL. Intended as the integration surface for
programmatic clients (e.g. an agent runtime) that drive EvoSci headlessly.
- stream/json_sink.py: write_events_as_json + stream_json sink, plus
redirect_console_to_stderr helper for stdout purity
- cli/interactive.py: cmd_run gains output_format; stream-json branch runs
the sink instead of the Rich renderer
- cli/commands.py: --output-format option + validation (stream-json
requires -p; value must be text|stream-json)
- docs/stream-json.md: event-schema contract + example transcript
- tests: json sink serialization, CLI dispatch, console redirect, validation
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(cli): honor explicit --no-auto-mode over config in stream-json
Address CodeRabbit review (discussion_r3514041123): the auto-mode override
block only wrote to cli_overrides when the resolved value was True, so an
explicit --no-auto-mode silently fell back to a config that enables
auto-mode -- breaking "explicit flags always win" and leaving stream-json
running unattended despite the warning. Write auto_mode=False when the flag
is explicitly False. Add regression tests that capture the overrides.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(tui): keep welcome banner at top after /new
PR #262 replaced scroll_end() with anchor() for free-scrolling.
When /new clears a long anchored conversation, the anchor kept
the viewport pinned to the (now empty) bottom, producing a
negative scroll_y and pushing the welcome banner out of view.
Reset the anchor and scroll to the top in clear_chat(), and
restore the follow/new-content flags so the fresh session starts
correctly.
Closes#301
* fix(tui): suppress anchor when chat content fits viewport
The previous fix for #301 only handled the /new path. din0s reported
that the banner still dropped to the bottom after a normal short turn
(user types 'hi', agent replies) — i.e. whenever the conversation
fit in the viewport. Root cause is in Textual's compositor
(textual._compositor): when a widget is anchored, scroll_y is
recomputed via set_reactive, which bypasses the validator. If the
anchored widget's content is shorter than the viewport, scroll_y
goes negative on the next layout pass and the welcome banner is
pushed below the visible region.
PR #262 made _stream_with_widgets re-engage the anchor at the end of
every turn via _anchor_chat, so the bug surfaced on any short reply
that fit in the viewport. Markdown re-renders, status-bar updates,
or any subsequent mount would then trip the compositor.
Fix in three places:
* _anchor_chat: only engage the anchor when max_scroll_y > 0;
otherwise release and scroll_home so the banner stays at the top.
* streaming anchor loop: if content shrinks below the viewport
mid-stream (e.g. loading widget removed), release the anchor
instead of leaving _anchored=True for the compositor to trip on.
* clear_chat: keep the unconditional reset (children are removed
asynchronously so a max_scroll_y check would be stale) but
document why.
Adds two regressions:
* test_short_turn_keeps_banner_at_top_after_layout_refresh — the
exact scenario din0s tested; fails with scroll_y=-10 on the
previous code, passes with the fix.
* test_long_turn_keeps_viewport_pinned_to_bottom — guards against
regressing free-scrolling for overflowing conversations.
Manually verified: 'hi' -> reply (banner stays at top) -> /new
(banner at top) -> another turn (banner stays at top).
* test(tui): address review feedback on banner-position regressions
- extract `_release_anchor_and_pin_top` helper for the 3-line
`anchor(False) + scroll_home(...)` pattern repeated in
`clear_chat`, `_anchor_chat`, and the streaming loop
- replace `pytest.skip` in `_capture_app` with a hard `RuntimeError`
so a broken capture never silently passes
- drop the redundant `load_agent` and `create_session_workspace`
monkeypatches (the factory is given those as parameters, so the
module-level symbols never run; added a comment explaining why)
- add a defensive `_FakeChannelRuntime` patch for symmetry with the
other module-level fakes
- drop the local `_run` helper and use the `run_async` fixture from
`conftest.py` (its teardown is better)
* test(tui): replace _FakeChannelRuntime with _auto_start_channel no-op
The _FakeChannelRuntime patch was ineffective because ChannelRuntime is
just a dataclass — the real channel manager still started via
_auto_start_channel, leaving pending tasks and non-hermetic test state.
Per review feedback, stub _auto_start_channel directly instead.
* feat: expose model registry at GET /api/models
* fix: include Ollama models in /api/models endpoint
* fix: honor env vars override in /api/models endpoint
* fix: offload get_effective_config to thread to satisfy blockbuster
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>