POST /api/model-registry/test (model_config:test) runs the section 9.4
flow: resolve_for_test, per-test credential resolution, build_chat_model
with both safe clients, one minimal chat call, and per-capability probes
(tools/structured_output/vision) whose failures only mark that capability
unverified. Results upsert the model_verifications five-tuple inside a
BEGIN IMMEDIATE transaction that re-checks the registry revision and
configuration hash, returning 409 MODEL_CONFIGURATION_CHANGED on any
concurrent change. resolve_for_test now also relaxes the passing-
verification gate, which the provider test itself produces.
effective_request_options reuses the redacted adapter.build_request
output; the OpenAPI contract and checked-in openapi.json are updated.
- add the missing rule-8 counterexample test: declared capabilities
exceeding the adapter protocol are rejected with
CAPABILITY_UNSUPPORTED_BY_ADAPTER (all eight section 9.2 checks now
have at least one negative test)
- raise CREDENTIAL_NOT_CONFIGURED explicitly in _check_enabled_model
when a required credential reference is null instead of relying on
resolve_parameters call ordering
The abort path was read-then-write with an unconditional UPDATE, so a bind
committing between the two calls was clobbered back to aborted, losing its
langgraph_run_id. Add a conditional store-level abort_run_snapshot
(prepared-only UPDATE, rowcount-checked) and re-read on a lost race, matching
the bind loop. Also pin the inherit selection_hash test to a hardcoded
SHA-256 literal instead of reimplementing the serialization in the test.
Resolver (8.1): validates provider/model/credential/capability/limits and
the 6.5 four-mode input budget, freezes ResolvedModelConfig; resolve_for_test
relaxes only the enabled-visibility check (9.4); compute_availability is the
single 4.3 six-state judgement (stale beats configured, selectable only when
enabled).
SnapshotService (8.2, shared by the Task 5 HTTP API and Task 7 local entry):
freezes both roles' full ResolvedModelConfig with adapter spec revision,
fixed reserves, capabilities, and credential revisions; selection-hash
idempotency with pre-resolution semantics; prepared(15min)/bound(+24h)/
expired/aborted lifecycle with atomic bind; binding-checked reads that
revalidate frozen spec revisions; per-call credential resolution against the
frozen revision with no in-process secret cache (5.2); public diagnostic
view limited to the 8.2 safe subset.
Store gains additive helpers (credential pointer lookup, verification
listing, active-triplet lookup, conditional bind, due-expiry sweep) and the
taxonomy gains SNAPSHOT_NOT_FOUND (404) for missing snapshots.
Review fixes for the Task 3 contract layer:
- build_chat_model now accepts http_async_client alongside http_client
(at least one required) and wires it into ChatOpenAI
(http_async_client), ChatAnthropic (seeded _async_client), and
ChatOllama (async_client_kwargs transport), closing the unsafe
default-async-client gap.
- ChatOllama safe transports move from the shared client_kwargs to
sync_client_kwargs/async_client_kwargs; langchain-ollama merges shared
kwargs into both clients, which poisoned the async client with a sync
transport and crashed ainvoke.
- Unsupported parameters now actually execute the contract-declared
normalizer (reject_non_auto) instead of a hardcoded raise, with a
fallback rejection if a normalizer would let a value through.
- build_chat_model rejects overlapping client_options/request_options
keys instead of silently overwriting.
Add the Task 3 parameter contract layer (design doc 6.1-6.4):
- adapters.py: versioned built-in contracts for the five phase-1
adapters plus the openai-compatible/glm-5.2 model-specific contract
(verbatim section 6.2 values); exact > longest glob > generic
matching with spec_revision pinning; resolve_parameters implementing
the section 6.1 inherit/omit semantics, contract validation with
stable error codes, and named normalizers (identity,
clamp_to_model_limit, omit_when_none, omit_when_auto,
reject_non_auto); Adapter.build_request as the single entry point
mapping ResolvedModelConfig to {client_options, request_options};
compute_effective_capabilities (protocol AND declared AND verified).
- factory.py: build_chat_model(resolved_config, http_client, *,
credential=None) with no **kwargs and no setdefault merging; injects
the safe HTTP client into ChatOpenAI/ChatAnthropic/ChatOllama, never
reads provider API-key environment variables, and strips the
OLLAMA_API_KEY authorization header for mode=none adapters.
- tests: per-adapter request-capturing fakes plus an httpx.MockTransport
outbound capture proving registry resolution matches the wire request.
EndpointPolicy validates provider base URLs (section 4.3): public https
endpoints with hostname and optional port pass; loopback, private,
link-local, multicast, unspecified, and cloud-metadata addresses are
denied unless the normalized URL exactly matches a registered
development_endpoints entry (no prefix or wildcard matching). URLs with
user info, fragments, or non-http(s) schemes are rejected with the new
stable 422 code ENDPOINT_NOT_ALLOWED.
SafeHttpTransport is the single network egress for adapters: a custom
httpcore NetworkBackend resolves DNS under control on every connect
(retries included), filters denied ranges, and connects directly to the
selected IP, while TLS SNI/certificate checks and the HTTP Host header
keep the original hostname. Redirects and env proxies are disabled;
every request origin re-passes URL-layer validation before any I/O.
* chore: add pytest-asyncio in auto mode
* test: migrate channel and stream tests to native async
Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.
* test: migrate command and model/middleware tests to native async
Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.
* test: migrate TUI, notifier, gateway, and session tests to native async
TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.
* test: replace direct asyncio.run() calls with native async tests
Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.
* test: drop undeclared anyio markers and delete run_async helper
The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).
* refactor(onboard): shared flow for ccproxy providers
* feat(onboard): support oauth configuration for auxiliary models
* fix(onboard): reuse main model auth for same-provider auxiliary
* fix(onboard): reconcile oauth providers
* feat(cli): add --output-format stream-json for headless clients
Emit EvoScientist's native event stream as line-delimited JSON on stdout
in single-shot (-p) mode, with all human output redirected to stderr so
stdout stays pure JSONL. Intended as the integration surface for
programmatic clients (e.g. an agent runtime) that drive EvoSci headlessly.
- stream/json_sink.py: write_events_as_json + stream_json sink, plus
redirect_console_to_stderr helper for stdout purity
- cli/interactive.py: cmd_run gains output_format; stream-json branch runs
the sink instead of the Rich renderer
- cli/commands.py: --output-format option + validation (stream-json
requires -p; value must be text|stream-json)
- docs/stream-json.md: event-schema contract + example transcript
- tests: json sink serialization, CLI dispatch, console redirect, validation
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(cli): honor explicit --no-auto-mode over config in stream-json
Address CodeRabbit review (discussion_r3514041123): the auto-mode override
block only wrote to cli_overrides when the resolved value was True, so an
explicit --no-auto-mode silently fell back to a config that enables
auto-mode -- breaking "explicit flags always win" and leaving stream-json
running unattended despite the warning. Write auto_mode=False when the flag
is explicitly False. Add regression tests that capture the overrides.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
* fix(tui): keep welcome banner at top after /new
PR #262 replaced scroll_end() with anchor() for free-scrolling.
When /new clears a long anchored conversation, the anchor kept
the viewport pinned to the (now empty) bottom, producing a
negative scroll_y and pushing the welcome banner out of view.
Reset the anchor and scroll to the top in clear_chat(), and
restore the follow/new-content flags so the fresh session starts
correctly.
Closes#301
* fix(tui): suppress anchor when chat content fits viewport
The previous fix for #301 only handled the /new path. din0s reported
that the banner still dropped to the bottom after a normal short turn
(user types 'hi', agent replies) — i.e. whenever the conversation
fit in the viewport. Root cause is in Textual's compositor
(textual._compositor): when a widget is anchored, scroll_y is
recomputed via set_reactive, which bypasses the validator. If the
anchored widget's content is shorter than the viewport, scroll_y
goes negative on the next layout pass and the welcome banner is
pushed below the visible region.
PR #262 made _stream_with_widgets re-engage the anchor at the end of
every turn via _anchor_chat, so the bug surfaced on any short reply
that fit in the viewport. Markdown re-renders, status-bar updates,
or any subsequent mount would then trip the compositor.
Fix in three places:
* _anchor_chat: only engage the anchor when max_scroll_y > 0;
otherwise release and scroll_home so the banner stays at the top.
* streaming anchor loop: if content shrinks below the viewport
mid-stream (e.g. loading widget removed), release the anchor
instead of leaving _anchored=True for the compositor to trip on.
* clear_chat: keep the unconditional reset (children are removed
asynchronously so a max_scroll_y check would be stale) but
document why.
Adds two regressions:
* test_short_turn_keeps_banner_at_top_after_layout_refresh — the
exact scenario din0s tested; fails with scroll_y=-10 on the
previous code, passes with the fix.
* test_long_turn_keeps_viewport_pinned_to_bottom — guards against
regressing free-scrolling for overflowing conversations.
Manually verified: 'hi' -> reply (banner stays at top) -> /new
(banner at top) -> another turn (banner stays at top).
* test(tui): address review feedback on banner-position regressions
- extract `_release_anchor_and_pin_top` helper for the 3-line
`anchor(False) + scroll_home(...)` pattern repeated in
`clear_chat`, `_anchor_chat`, and the streaming loop
- replace `pytest.skip` in `_capture_app` with a hard `RuntimeError`
so a broken capture never silently passes
- drop the redundant `load_agent` and `create_session_workspace`
monkeypatches (the factory is given those as parameters, so the
module-level symbols never run; added a comment explaining why)
- add a defensive `_FakeChannelRuntime` patch for symmetry with the
other module-level fakes
- drop the local `_run` helper and use the `run_async` fixture from
`conftest.py` (its teardown is better)
* test(tui): replace _FakeChannelRuntime with _auto_start_channel no-op
The _FakeChannelRuntime patch was ineffective because ChannelRuntime is
just a dataclass — the real channel manager still started via
_auto_start_channel, leaving pending tasks and non-hermetic test state.
Per review feedback, stub _auto_start_channel directly instead.
* feat: expose model registry at GET /api/models
* fix: include Ollama models in /api/models endpoint
* fix: honor env vars override in /api/models endpoint
* fix: offload get_effective_config to thread to satisfy blockbuster
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
* feat: add scheduler functionality with cron-style task management
- Implemented a new scheduler subagent to automate recurring tasks using cron expressions.
- Enhanced the subagent factory to include the skill manager and auxiliary chat model for the scheduler.
- Created a YAML configuration for the scheduler with a detailed system prompt and toolset.
- Updated README files to include documentation on scheduled tasks and usage examples.
- Added tests for the scheduler, including command execution, scheduling tools, and middleware integration.
- Introduced new dependencies for timezone handling and ensured compatibility in the project configuration.
* fix(async-notifier): ensure fallback hint is used for unknown notification kinds
* feat: enhance scheduling functionality and improve system message handling
- Ensure 'task' is excluded from the default PTC allowlist to prevent ValueError in langchain-quickjs >=0.3.
- Verify that essential async dispatch tools remain in the allowlist.
- Confirm that the live quickjs filter accepts the default allowlist even with a 'task' tool present.
- Test the creation of the code_interpreter middleware to ensure it builds correctly.
* feat(gateway): graph gateway protocol
* refactor(cli): wire gateway in cli/tui
* refactor(gateway): centralize runtime gateway init
* chore(gateway): restrict RunRequest message type
* feat(gateway): add langgraph server gateway
* chore(cli): tighten serve runtime state typing
* refactor(cli): route async task state reads through graph gateway
* refactor(gateway): support graph targets in server gateway
* refactor(cli): route session commands through graph gateway
* refactor(cli): fold thread store under graph gateway
* refactor(gateway): route graph state access through gateway
* refactor(channels): wire graph gateway
* refactor(memory): preserve graph threads for cloning
* feat(gateway): add thread cloning
* fix(tui): pass effective workspace for thread creation
* chore(memory): add workspare dir to memory worker metadata
* fix(sessions): filter preloaded UUID registy entries by the current scope
* test(fakes): use https
* refactor(consumer): consolidate imports
* fix(stream): optional summarization event
* fix(gateway): resolve abbreviated thread IDs by search
* fix(gateway): page server thread listings
* fix(gateway): emit pending interrupt events
* style: fmt
* feat(gateway): persist workspace_dir & model in thread metadata
* fix(gateway): page server thread prefix resolution
* fix(gateway): expose server thread list metadata
* refactor: add back type def
* refactor: tighten types
* revert: add back worker thread deletion
The worker thread forking changes are out of scope for now, so to
maintain parity with the existing behavior we'll leave this intact.
* fix(gateway): apply compaction to server thread history
* refactor(stream): restore direct summary replay suppression
* fix(gateway): preserve compaction state and server stream output
* fix(gateway): close local stream generator on cancellation
---------
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Typing `/model` and pressing Enter did nothing in the TUI; the picker
only opened via `/model --save` or `/model <name>`. The completion popup
matched both `/model` and `/model-fallback` by prefix, so the
exact-match-hide guard (which required a single match) never fired. With
the popup still visible, the TUI's Enter handler completed the text
instead of submitting the command, so it never executed.
Treat the typed prefix as an exact match whenever it equals any matched
command name, not only when it is the sole match. This hides the popup
on a complete command name so Enter submits it, even when a longer
command shares the prefix.