Commit Graph

602 Commits

Author SHA1 Message Date
m4 940db565b3 fix(model-registry): close review gaps in save-time validation
- add the missing rule-8 counterexample test: declared capabilities
  exceeding the adapter protocol are rejected with
  CAPABILITY_UNSUPPORTED_BY_ADAPTER (all eight section 9.2 checks now
  have at least one negative test)
- raise CREDENTIAL_NOT_CONFIGURED explicitly in _check_enabled_model
  when a required credential reference is null instead of relying on
  resolve_parameters call ordering
2026-07-21 09:38:37 +08:00
m4 dbb6b7abde feat(model-registry): add delegation-JWT auth and config/snapshot HTTP API
- BFF service token (constant-time, plaintext or SHA-256 hash) plus
  X-Evo-Actor delegation JWT verification (ES256/RS256, iss/aud, <=60s
  lifetime, required claims, thread binding) with atomic jti anti-replay
- Config API: GET/PUT /api/model-registry, credential rotation endpoint,
  GET /api/models selector; PUT runs the section 9.2 save-time checks
  inside the registry write transaction after credential writes
- Snapshot API: create/bind/delete routes delegating to SnapshotService
  with thread/deployment binding checks and 9.5 unified error payloads
- Platform security config loader (config.yaml fields), OpenAPI export
  (scripts/export_model_registry_schema.py -> model_registry/openapi.json)
- Mount new routes in langgraph_dev/http.py; retire the legacy
  GET /api/models and POST /api/runtime-snapshots handlers
- Declare PyJWT>=2.8 (previously transitive); extend the 9.5 error code
  table with the HTTP-layer codes (400/401/403/422/500)
2026-07-21 09:28:55 +08:00
m4 c8c46eab16 fix(model-registry): make snapshot abort atomic against concurrent bind
The abort path was read-then-write with an unconditional UPDATE, so a bind
committing between the two calls was clobbered back to aborted, losing its
langgraph_run_id. Add a conditional store-level abort_run_snapshot
(prepared-only UPDATE, rowcount-checked) and re-read on a lost race, matching
the bind loop. Also pin the inherit selection_hash test to a hardcoded
SHA-256 literal instead of reimplementing the serialization in the test.
2026-07-21 08:45:33 +08:00
m4 0cc995eb80 docs(model-registry): add task 4 resolver and snapshot service report 2026-07-21 08:34:50 +08:00
m4 b1233d42dc feat(model-registry): add ModelRegistryResolver and run snapshot service
Resolver (8.1): validates provider/model/credential/capability/limits and
the 6.5 four-mode input budget, freezes ResolvedModelConfig; resolve_for_test
relaxes only the enabled-visibility check (9.4); compute_availability is the
single 4.3 six-state judgement (stale beats configured, selectable only when
enabled).

SnapshotService (8.2, shared by the Task 5 HTTP API and Task 7 local entry):
freezes both roles' full ResolvedModelConfig with adapter spec revision,
fixed reserves, capabilities, and credential revisions; selection-hash
idempotency with pre-resolution semantics; prepared(15min)/bound(+24h)/
expired/aborted lifecycle with atomic bind; binding-checked reads that
revalidate frozen spec revisions; per-call credential resolution against the
frozen revision with no in-process secret cache (5.2); public diagnostic
view limited to the 8.2 safe subset.

Store gains additive helpers (credential pointer lookup, verification
listing, active-triplet lookup, conditional bind, due-expiry sweep) and the
taxonomy gains SNAPSHOT_NOT_FOUND (404) for missing snapshots.
2026-07-21 08:33:12 +08:00
m4 b2e28249fd fix(model-registry): inject async safe clients and split ollama transports
Review fixes for the Task 3 contract layer:

- build_chat_model now accepts http_async_client alongside http_client
  (at least one required) and wires it into ChatOpenAI
  (http_async_client), ChatAnthropic (seeded _async_client), and
  ChatOllama (async_client_kwargs transport), closing the unsafe
  default-async-client gap.
- ChatOllama safe transports move from the shared client_kwargs to
  sync_client_kwargs/async_client_kwargs; langchain-ollama merges shared
  kwargs into both clients, which poisoned the async client with a sync
  transport and crashed ainvoke.
- Unsupported parameters now actually execute the contract-declared
  normalizer (reject_non_auto) instead of a hardcoded raise, with a
  fallback rejection if a normalizer would let a value through.
- build_chat_model rejects overlapping client_options/request_options
  keys instead of silently overwriting.
2026-07-20 22:40:33 +08:00
m4 af4ae1aef5 feat(model-registry): add adapter parameter contracts and build_chat_model factory
Add the Task 3 parameter contract layer (design doc 6.1-6.4):

- adapters.py: versioned built-in contracts for the five phase-1
  adapters plus the openai-compatible/glm-5.2 model-specific contract
  (verbatim section 6.2 values); exact > longest glob > generic
  matching with spec_revision pinning; resolve_parameters implementing
  the section 6.1 inherit/omit semantics, contract validation with
  stable error codes, and named normalizers (identity,
  clamp_to_model_limit, omit_when_none, omit_when_auto,
  reject_non_auto); Adapter.build_request as the single entry point
  mapping ResolvedModelConfig to {client_options, request_options};
  compute_effective_capabilities (protocol AND declared AND verified).
- factory.py: build_chat_model(resolved_config, http_client, *,
  credential=None) with no **kwargs and no setdefault merging; injects
  the safe HTTP client into ChatOpenAI/ChatAnthropic/ChatOllama, never
  reads provider API-key environment variables, and strips the
  OLLAMA_API_KEY authorization header for mode=none adapters.
- tests: per-adapter request-capturing fakes plus an httpx.MockTransport
  outbound capture proving registry resolution matches the wire request.
2026-07-20 22:17:11 +08:00
m4 c46ae17084 feat(model-registry): add EndpointPolicy and SafeHttpTransport SSRF defenses
EndpointPolicy validates provider base URLs (section 4.3): public https
endpoints with hostname and optional port pass; loopback, private,
link-local, multicast, unspecified, and cloud-metadata addresses are
denied unless the normalized URL exactly matches a registered
development_endpoints entry (no prefix or wildcard matching). URLs with
user info, fragments, or non-http(s) schemes are rejected with the new
stable 422 code ENDPOINT_NOT_ALLOWED.

SafeHttpTransport is the single network egress for adapters: a custom
httpcore NetworkBackend resolves DNS under control on every connect
(retries included), filters denied ranges, and connects directly to the
selected IP, while TLS SNI/certificate checks and the HTTP Host header
keep the original hostname. Redirects and env proxies are disabled;
every request origin re-passes URL-layer validation before any I/O.
2026-07-20 21:25:15 +08:00
m4 c21fc0a272 feat(model-registry): add RegistryV4 schema, SQLite store, and unified error codes
New independent subpackage EvoScientist/model_registry implementing the
frozen unified model configuration design (v1.1.0, sections 4.2, 4.3,
5.1, 8.2, 9.5):

- schemas.py: single Pydantic v2 RegistryV4 schema (ModelRef, ProviderConfig,
  ModelConfig, AuthConfig) plus AdapterParameterSpec/AuthSpec/ParameterRule,
  ModelAvailability, ResolvedModelConfig (no secrets), CredentialStatus
- errors.py: all 23 stable error codes from the 9.5 table with HTTP status
  mapping and the unified {code, message, details, request_id} payload
- store.py: ModelRuntimeStore over model-runtime.sqlite3 (0700 dir, 0600
  file, WAL, foreign keys, busy_timeout) with BEGIN IMMEDIATE revision CAS,
  bootstrap->active atomic transition, immutable credential versions with
  masked status, model_verifications upsert, run_runtime_snapshots partial
  unique index, delegation_jtis, and shared-storage lock probe
- hashing.py: normalized SHA-256 configuration_hash

No existing module behavior changed. 89 new tests; full suite passes
(2922 passed, 10 skipped).
2026-07-20 20:43:28 +08:00
m4 8a0ab17936 chore: baseline WIP before unified model configuration implementation
Pre-existing uncommitted work (runtime snapshots, message budget middleware) preserved as baseline.
2026-07-20 20:15:38 +08:00
m4 38668c4ce5 feat: add workspace isolation and provider administration 2026-07-19 12:17:18 +08:00
m4 7a3fcc7c8e Merge remote-tracking branch 'upstream/main'
# Conflicts:
#	README.md
#	uv.lock
2026-07-13 09:46:07 +08:00
dinos 49770949da fix(langgraph): prefer executable in venv over path (#341) 2026-07-11 11:59:27 +00:00
X-iZhang 6f10406d5b chore: update version to v0.2.2
Docker / build (push) Has been cancelled
v0.2.2
2026-07-11 01:16:03 +01:00
X-iZhang 9042068094 feat(models): add support for GPT-5.6 variants and update context windows for Grok models 2026-07-11 00:51:06 +01:00
dependabot[bot] 81ff0519dc chore(deps): bump soupsieve in the uv group across 1 directory (#347)
Bumps the uv group with 1 update in the / directory: [soupsieve](https://github.com/facelessuser/soupsieve).


Updates `soupsieve` from 2.8.3 to 2.8.4
- [Release notes](https://github.com/facelessuser/soupsieve/releases)
- [Commits](https://github.com/facelessuser/soupsieve/compare/2.8.3...2.8.4)

---
updated-dependencies:
- dependency-name: soupsieve
  dependency-version: 2.8.4
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-10 22:47:18 +01:00
m4 e0acc6155e feat: improve WebUI run recovery
Build / build (push) Has been cancelled
Docker / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
2026-07-10 17:35:44 +08:00
dinos 690b903f85 test: standardize async tests on pytest-asyncio auto mode (#338)
* chore: add pytest-asyncio in auto mode

* test: migrate channel and stream tests to native async

Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.

* test: migrate command and model/middleware tests to native async

Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.

* test: migrate TUI, notifier, gateway, and session tests to native async

TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.

* test: replace direct asyncio.run() calls with native async tests

Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.

* test: drop undeclared anyio markers and delete run_async helper

The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).
2026-07-08 18:37:48 +00:00
dinos d2452c54d5 Refactor onboarding OAuth flow for auxiliary models (#337)
* refactor(onboard): shared flow for ccproxy providers

* feat(onboard): support oauth configuration for auxiliary models

* fix(onboard): reuse main model auth for same-provider auxiliary

* fix(onboard): reconcile oauth providers
2026-07-08 18:28:44 +00:00
dinos a7b9e175c1 fix(config): set config.yaml permissions to 0x600 (#336) 2026-07-08 19:25:01 +01:00
dinos be3dd272c3 test: deflake timing-dependent tests (#335)
* test: deflake timing-dependent tests

Inject a clock into channel dedup tests, replace fixed async sleeps with
events/explicit flushes, and avoid wall-clock waits in background tests.

* coderabbit nit
2026-07-07 08:25:35 +01:00
X-iZhang df54d8498c chore: update version to 0.2.1 2026-07-05 10:17:28 +01:00
Wiktor Cupiał 1d117ff277 feat: completion enchancements (#302)
* feat: completion enchancements

* fix: handle exception

* fix: duplicate view

* fix: remove deadcode

* fix tab
2026-07-05 05:14:42 +00:00
renaissancefieldlite f086d77756 Fix UTF-8 config reads on Windows (#318)
* Fix UTF-8 config reads on Windows

* test: cover utf8 production loaders

* fix: read and write settings as utf8

* Apply ruff formatting
2026-07-05 05:10:24 +00:00
jfilipiuk 5dacab7e4c fix: bump langchain-openrouter to 0.2.5 for upstream OpenRouter duplicate-terminal-chunk fix (#313) 2026-07-05 05:01:35 +00:00
kalisgd0h bd54eaa0a4 feat(cli): add --output-format stream-json for headless clients (#309)
* feat(cli): add --output-format stream-json for headless clients

Emit EvoScientist's native event stream as line-delimited JSON on stdout
in single-shot (-p) mode, with all human output redirected to stderr so
stdout stays pure JSONL. Intended as the integration surface for
programmatic clients (e.g. an agent runtime) that drive EvoSci headlessly.

- stream/json_sink.py: write_events_as_json + stream_json sink, plus
  redirect_console_to_stderr helper for stdout purity
- cli/interactive.py: cmd_run gains output_format; stream-json branch runs
  the sink instead of the Rich renderer
- cli/commands.py: --output-format option + validation (stream-json
  requires -p; value must be text|stream-json)
- docs/stream-json.md: event-schema contract + example transcript
- tests: json sink serialization, CLI dispatch, console redirect, validation

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(cli): honor explicit --no-auto-mode over config in stream-json

Address CodeRabbit review (discussion_r3514041123): the auto-mode override
block only wrote to cli_overrides when the resolved value was True, so an
explicit --no-auto-mode silently fell back to a config that enables
auto-mode -- breaking "explicit flags always win" and leaving stream-json
running unattended despite the warning. Write auto_mode=False when the flag
is explicitly False. Add regression tests that capture the overrides.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 04:53:52 +00:00
dinos 2b244888ec feat(memory): autoskills (#319) 2026-07-03 11:16:33 +02:00
houren Antony 7568a6bc1f fix(tui): keep welcome banner at top after /new (#311)
* fix(tui): keep welcome banner at top after /new

PR #262 replaced scroll_end() with anchor() for free-scrolling.
When /new clears a long anchored conversation, the anchor kept
the viewport pinned to the (now empty) bottom, producing a
negative scroll_y and pushing the welcome banner out of view.

Reset the anchor and scroll to the top in clear_chat(), and
restore the follow/new-content flags so the fresh session starts
correctly.

Closes #301

* fix(tui): suppress anchor when chat content fits viewport

The previous fix for #301 only handled the /new path. din0s reported
that the banner still dropped to the bottom after a normal short turn
(user types 'hi', agent replies) — i.e. whenever the conversation
fit in the viewport. Root cause is in Textual's compositor
(textual._compositor): when a widget is anchored, scroll_y is
recomputed via set_reactive, which bypasses the validator. If the
anchored widget's content is shorter than the viewport, scroll_y
goes negative on the next layout pass and the welcome banner is
pushed below the visible region.

PR #262 made _stream_with_widgets re-engage the anchor at the end of
every turn via _anchor_chat, so the bug surfaced on any short reply
that fit in the viewport. Markdown re-renders, status-bar updates,
or any subsequent mount would then trip the compositor.

Fix in three places:

  * _anchor_chat: only engage the anchor when max_scroll_y > 0;
    otherwise release and scroll_home so the banner stays at the top.
  * streaming anchor loop: if content shrinks below the viewport
    mid-stream (e.g. loading widget removed), release the anchor
    instead of leaving _anchored=True for the compositor to trip on.
  * clear_chat: keep the unconditional reset (children are removed
    asynchronously so a max_scroll_y check would be stale) but
    document why.

Adds two regressions:
  * test_short_turn_keeps_banner_at_top_after_layout_refresh — the
    exact scenario din0s tested; fails with scroll_y=-10 on the
    previous code, passes with the fix.
  * test_long_turn_keeps_viewport_pinned_to_bottom — guards against
    regressing free-scrolling for overflowing conversations.

Manually verified: 'hi' -> reply (banner stays at top) -> /new
(banner at top) -> another turn (banner stays at top).

* test(tui): address review feedback on banner-position regressions

- extract `_release_anchor_and_pin_top` helper for the 3-line
  `anchor(False) + scroll_home(...)` pattern repeated in
  `clear_chat`, `_anchor_chat`, and the streaming loop
- replace `pytest.skip` in `_capture_app` with a hard `RuntimeError`
  so a broken capture never silently passes
- drop the redundant `load_agent` and `create_session_workspace`
  monkeypatches (the factory is given those as parameters, so the
  module-level symbols never run; added a comment explaining why)
- add a defensive `_FakeChannelRuntime` patch for symmetry with the
  other module-level fakes
- drop the local `_run` helper and use the `run_async` fixture from
  `conftest.py` (its teardown is better)

* test(tui): replace _FakeChannelRuntime with _auto_start_channel no-op

The _FakeChannelRuntime patch was ineffective because ChannelRuntime is
just a dataclass — the real channel manager still started via
_auto_start_channel, leaving pending tasks and non-hermetic test state.

Per review feedback, stub _auto_start_channel directly instead.
2026-06-30 19:15:08 +01:00
X-iZhang 682922f690 docs: update changelog for version 0.2.0 release 2026-06-27 00:43:17 +01:00
dependabot[bot] 598b668131 chore(deps): bump langgraph-checkpoint (#310)
Bumps the uv group with 1 update in the / directory: [langgraph-checkpoint](https://github.com/langchain-ai/langgraph).


Updates `langgraph-checkpoint` from 4.1.0 to 4.1.1
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1)

---
updated-dependencies:
- dependency-name: langgraph-checkpoint
  dependency-version: 4.1.1
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-27 00:23:16 +01:00
X-iZhang 6546f4022f chore: update version to v0.2.0 2026-06-26 23:55:46 +01:00
jfilipiuk 7214d4099d feat: expose model registry at GET /api/models (#308)
* feat: expose model registry at GET /api/models

* fix: include Ollama models in /api/models endpoint

* fix: honor env vars override in /api/models endpoint

* fix: offload get_effective_config to thread to satisfy blockbuster

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-26 22:48:37 +01:00
dinos f2f010a350 feat(memory): observation linking (#307)
* refactor(gateway): create module for launching async/bg agents

* refactor(memory): refactor worker launch around source context & output deltas

* refactor(gateway): generalize async/bg module

* refactor(memory): revamp worker launching

* feat(memory): add observation linking

* test(memory): remove redundant test branches

* fix(memory): make 'supersedes' relation directional

* fix(memory): don't create empty project observation dirs

* fix(memory): schedule direct observations for linking

* fix(cli): wait for observation linker before shutdown

* fix(memory): block arbitrary writes to /memories

* fix(linker): remove `linked_by` attribute from frontmatter

* refactor(linker): rename base relationship to `comlpements`

* fix(cli): bump worker wait to 2m

* feat(tools): catch malformed tool calls & retry

* feat(status): add linking result to statusbar

* fix(linker): don't launch linker when observations are disabled

* fix(memory): use posix paths

* fix(watcher): call abort hook on error status

* fix(watcher): delete thread on failed run creation

* fix(watcher): preserve url

* fix(observation): record session_id, drop unused fields

* fix(memory): reject unsupported worker source types

* refactor(backends): shared memory backend builder

* fix(scheduler): resolve linker inputs outside lock

* fix(memory): dont launch workers / record observations without thread_id

* feat(memory): include related observations in tool results

* fix(memory): skip malformed observation frontmatter

* revert(tools): drop tool error handling changes from this PR

* fix(memory): serialize observation link writes

* fix(memory): queue observations written by aborted workers

* fix(memory): track observation linker launch handoff

* fix(memory): resolve cross-project related observations

* fix(status): avoid recounting reason-only link updates

* fix(memory): avoid rereading file for content

* fix(linker): use neutral prose for bidirectional reasons

* test(memory): coverage for aborted/failed launches

* test(memory): cleanup & helpers

* feat(linker): add observations index hint
2026-06-26 22:20:52 +01:00
Xi Zhang 7ccfe68f3f feat: add scheduler functionality with cron-style task management (#306)
* feat: add scheduler functionality with cron-style task management

- Implemented a new scheduler subagent to automate recurring tasks using cron expressions.
- Enhanced the subagent factory to include the skill manager and auxiliary chat model for the scheduler.
- Created a YAML configuration for the scheduler with a detailed system prompt and toolset.
- Updated README files to include documentation on scheduled tasks and usage examples.
- Added tests for the scheduler, including command execution, scheduling tools, and middleware integration.
- Introduced new dependencies for timezone handling and ensured compatibility in the project configuration.

* fix(async-notifier): ensure fallback hint is used for unknown notification kinds

* feat: enhance scheduling functionality and improve system message handling
2026-06-25 17:31:23 +01:00
dinos b1dccf17ea fix(tool-selector): memory tools & state for main agent (#305)
* fix(tool-selector): always include memory tools

* fix(tool-selector): only track & show state for main agent

* docs: update docstring
2026-06-23 14:47:25 +01:00
X-iZhang c063a00c8f Add regression tests for code_interpreter middleware PTC allowlist
- Ensure 'task' is excluded from the default PTC allowlist to prevent ValueError in langchain-quickjs >=0.3.
- Verify that essential async dispatch tools remain in the allowlist.
- Confirm that the live quickjs filter accepts the default allowlist even with a 'task' tool present.
- Test the creation of the code_interpreter middleware to ensure it builds correctly.
2026-06-23 08:40:35 +01:00
dependabot[bot] 9460b9ecd6 chore(deps): bump the uv group across 1 directory with 2 updates (#304)
Bumps the uv group with 2 updates in the / directory: [langsmith](https://github.com/langchain-ai/langsmith-sdk) and [pydantic-settings](https://github.com/pydantic/pydantic-settings).


Updates `langsmith` from 0.8.15 to 0.8.18
- [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases)
- [Commits](https://github.com/langchain-ai/langsmith-sdk/compare/v0.8.15...v0.8.18)

Updates `pydantic-settings` from 2.14.0 to 2.14.2
- [Release notes](https://github.com/pydantic/pydantic-settings/releases)
- [Commits](https://github.com/pydantic/pydantic-settings/compare/v2.14.0...v2.14.2)

---
updated-dependencies:
- dependency-name: langsmith
  dependency-version: 0.8.18
  dependency-type: indirect
  dependency-group: uv
- dependency-name: pydantic-settings
  dependency-version: 2.14.2
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-22 18:16:45 +01:00
X-iZhang 57461f671f chore: update version to v0.1.8 2026-06-22 18:10:21 +01:00
dinos bd307f3a11 refactor: LangGraph gateway layer for UI-agnostic graph and thread access (#295)
* feat(gateway): graph gateway protocol

* refactor(cli): wire gateway in cli/tui

* refactor(gateway): centralize runtime gateway init

* chore(gateway): restrict RunRequest message type

* feat(gateway): add langgraph server gateway

* chore(cli): tighten serve runtime state typing

* refactor(cli): route async task state reads through graph gateway

* refactor(gateway): support graph targets in server gateway

* refactor(cli): route session commands through graph gateway

* refactor(cli): fold thread store under graph gateway

* refactor(gateway): route graph state access through gateway

* refactor(channels): wire graph gateway

* refactor(memory): preserve graph threads for cloning

* feat(gateway): add thread cloning

* fix(tui): pass effective workspace for thread creation

* chore(memory): add workspare dir to memory worker metadata

* fix(sessions): filter preloaded UUID registy entries by the current scope

* test(fakes): use https

* refactor(consumer): consolidate imports

* fix(stream): optional summarization event

* fix(gateway): resolve abbreviated thread IDs by search

* fix(gateway): page server thread listings

* fix(gateway): emit pending interrupt events

* style: fmt

* feat(gateway): persist workspace_dir & model in thread metadata

* fix(gateway): page server thread prefix resolution

* fix(gateway): expose server thread list metadata

* refactor: add back type def

* refactor: tighten types

* revert: add back worker thread deletion

The worker thread forking changes are out of scope for now, so to
maintain parity with the existing behavior we'll leave this intact.

* fix(gateway): apply compaction to server thread history

* refactor(stream): restore direct summary replay suppression

* fix(gateway): preserve compaction state and server stream output

* fix(gateway): close local stream generator on cancellation

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-22 13:54:14 +00:00
dinos 6eb467e70b feat(llm): make OpenRouter Anthropic prompt cache opt-out (#299)
* feat(llm): make OpenRouter Anthropic prompt cache opt-out

* fix(llm): restore truthy env flag helper
2026-06-18 15:09:22 +02:00
Ziheng Zhang 6593ac9b5b fix(cli): submit slash command on Enter when name prefixes another (#293) (#300)
Typing `/model` and pressing Enter did nothing in the TUI; the picker
only opened via `/model --save` or `/model <name>`. The completion popup
matched both `/model` and `/model-fallback` by prefix, so the
exact-match-hide guard (which required a single match) never fired. With
the popup still visible, the TUI's Enter handler completed the text
instead of submitting the command, so it never executed.

Treat the typed prefix as an exact match whenever it equals any matched
command name, not only when it is the sole match. This hides the popup
on a complete command name so Enter submits it, even when a longer
command shares the prefix.
2026-06-18 14:07:43 +01:00
dinos 6894d45be8 chore: bump ruff in pre-commit (#298)
* chore(pre-commit): sync ruff version with uv

* chore(gitignore): don't ignore built-in skills dir

* style: unused var
2026-06-18 12:46:09 +01:00
X-iZhang 086425377d chore: update version to v0.1.7 2026-06-16 23:13:17 +01:00
dinos de3f588fbf fix: windows async MCP tool execution, restore graph state after interruptions (#290)
Co-authored-by: z00827015 <zhoulun1@huawei.com>
2026-06-16 21:03:26 +02:00
dependabot[bot] 5282a5028b chore(deps): bump the uv group across 1 directory with 4 updates (#291)
---
updated-dependencies:
- dependency-name: aiohttp
  dependency-version: 3.14.1
  dependency-type: direct:production
  dependency-group: uv
- dependency-name: cryptography
  dependency-version: 48.0.1
  dependency-type: direct:production
  dependency-group: uv
- dependency-name: python-multipart
  dependency-version: 0.0.31
  dependency-type: indirect
  dependency-group: uv
- dependency-name: starlette
  dependency-version: 1.3.1
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-16 11:21:10 +01:00
dependabot[bot] 42d8b9bf10 chore(deps): bump pyjwt in the uv group across 1 directory (#289)
Bumps the uv group with 1 update in the / directory: [pyjwt](https://github.com/jpadilla/pyjwt).


Updates `pyjwt` from 2.12.1 to 2.13.0
- [Release notes](https://github.com/jpadilla/pyjwt/releases)
- [Changelog](https://github.com/jpadilla/pyjwt/blob/master/CHANGELOG.rst)
- [Commits](https://github.com/jpadilla/pyjwt/compare/2.12.1...2.13.0)

---
updated-dependencies:
- dependency-name: pyjwt
  dependency-version: 2.13.0
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-16 09:56:31 +00:00
dinos f356de36a6 feat: memory retrieval (#281) 2026-06-16 09:14:29 +02:00
X-iZhang 8ea186441a chore(deps): bump deepagents 0.6.8→0.6.10 + langchain-quickjs →0.2.0 2026-06-13 18:43:51 +01:00
houren Antony 477481f9d9 feat(backends): use _platform_quote for Windows cmd.exe compatibility (#280)
* feat(backends): use _platform_quote for Windows cmd.exe compatibility

Resolves the 3 skipped E2E tests in test_backends.py that exercised
the /skills/... mount path. The path-rewriter was wrapping resolved
absolute paths via shlex.quote (POSIX single-quote style); cmd.exe
doesn't strip single quotes, so the literal ' characters ended up in
the subprocess argv and the python script failed to find its file.

Replace the 3 shlex.quote call sites in _resolve_virtual_mount_path
with _platform_quote, a thin platform dispatcher:
- POSIX: shlex.quote (unchanged)
- Windows: _cmd_quote uses cmd.exe-compatible double-quote wrapping
  and properly escapes embedded " and percent signs

Adds:
- backends.py: _is_windows, _cmd_quote, _platform_quote (~40 lines)
- test_backends.py: 6 TestPlatformQuote unit tests + _split_cmd
  cross-platform tokenizer helper to replace shlex.split in the 8
  sites that tokenize convert_virtual_paths_in_command results
  (POSIX shlex strips backslashes from bare Windows paths, which
  broke the 5 TestVirtualMountResolution assertions on Windows)

Removes:
- 3 @pytest.mark.skipif(sys.platform == "win32") markers on the
  E2E tests for /skills/... mount resolution

Refs #274.

* fix: escape % as %% in _cmd_quote instead of relying on double-quoting

cmd.exe expands %VAR% before processing quotes, so double-quoting
cannot neutralize percent signs.  Escape bare % as %% (the cmd.exe
idiom for a literal percent) before any other quoting logic.

Also updates _cmd_quote docstring and _resolve_virtual_mount_path
docstring to reflect the actual quoting strategy.

* style: fix ruff format (single → double quotes)

* fix: treat % as regular char in _cmd_quote, document limitation

%% escaping only collapses in .bat/.cmd files, not via cmd /c.
Since virtual-mount paths should never contain % in practice,
simpler to leave % alone and document the caveat.
2026-06-13 18:08:18 +01:00
houren Antony 76972449c7 feat(cli): multi-stage slash command completions with subcommand awareness (Phase 1 of #82) (#273)
* feat(cli): multi-stage slash command completions with subcommand awareness

Phase 1 of #82 — subcommand and argument awareness in completions.

- commands/base.py: add SubCommand dataclass and subcommands/category
  ClassVars to the Command ABC. Each SubCommand has name, description,
  and optional arguments.

- commands/manager.py: add get_subcommands() and list_subcommands()
  methods to expose subcommand metadata for completion rendering.

- commands/implementation/mcp.py: declare 6 subcommands (list, config,
  add, edit, remove, install).

- commands/implementation/model_fallback.py: declare 6 subcommands
  (list, add, remove, clear, save, help).

- commands/implementation/channel.py: declare 2 subcommands
  (status, stop).

- cli/tui_interactive.py: rewrite on_text_area_changed slash-completion
  branch. When the user types a command name + trailing space and the
  command has subcommands, show subcommand completions instead of hiding
  the popup. Filter subcommands by typed prefix in multi-token input.

- commands/implementation/general.py: /help now lists subcommands
  below each command that declares them.

Tests: 8 new tests covering SubCommand creation, CommandManager
subcommand lookup, and cross-command verification.
2293 passed baseline, no regressions.

* fix: subcommand completion preserves prefix + prompt_toolkit + tests

- _apply_selected_completion: preserve '/mcp ' prefix when completing
  subcommands via _comp_is_subcommand flag
- SlashCommandCompleter (Rich CLI): add subcommand completion support
- Fix trailing-space bug: rstrip prefix before top-level matching
- test_tui_widgets.py: update stub on_input_changed to match multi-stage
  logic; add 5 new subcommand tests
- test_command_manager.py: 8 tests for SubCommand + CommandManager

28 passed, 0 failed.

* style: ruff format tui_interactive.py + test_tui_widgets.py

* style: fix RUF012 ClassVar annotation on subcommands lists

* fix: sync test stub, add len>=3 guard, remove exact-match hide

- Sync test stub on_input_changed with real TUI code (remove exact-match
  hide for subcommands, add len(parts)>=3 guard)
- Update test_input_changed_exact_subcommand_hides -> shows_confirmation
- Add test_input_changed_three_parts_hides
- Remove unused category ClassVar (din0s: what is this for)

* refactor(commands): extract shared completion engine

Per din0s feedback: one shared completion engine (commands/_completion_engine.py)
that parses text + cursor once, returns structured CompletionCandidate objects
with replace_start/replace_end ranges.

- SlashCommandCompleter (Rich CLI): thin adapter, delegates to engine
- on_text_area_changed (TUI): thin adapter, delegates to engine
- _apply_selected_completion: uses candidate.replace_start/replace_end instead
  of _comp_is_subcommand flag
- Tests: engine tested directly (10 new tests), stub methods updated

29 passed, 0 failed.

* style: ruff format

* fix: preserve text after cursor when applying completion

CodeRabbit: replace_start only cuts from start to cursor,
dropping any suffix after the cursor. Use replace_start + replace_end
to correctly splice the replacement while preserving trailing text.

* fix(cli): repair slash-command completion (TUI crash, subcommand bugs, sort)

Apology + context: the previous push shipped a TUI-breaking change
(the new shared engine assumed ``event.text_area.cursor_position``
existed, but ``ChatTextArea`` / ``Changed`` don't expose it). User
caught the crash on ``/``; fixing that surfaced two more bugs in
the engine that din0s had already flagged. This commit addresses
all of them and drops a piece of dead stub code.

## Bug fixes

1. **TUI crash on ``/``** (``tui_interactive.py:2335``)
   ``event.cursor_position`` doesn't exist on the ``Changed`` event,
   and ``ChatTextArea`` (Textual ``TextArea`` subclass) doesn't expose
   ``cursor_position`` either. Pass ``len(event.text_area.text)``
   instead — the user types at the end of the input in practice.

2. **Subcommand trailing-space duplication** (``_completion_engine.py``)
   Typing ``/mcp a `` + Tab produced ``/mcp aadd``. The engine
   included the trailing space in ``replace_end``; the TUI apply
   unconditionally appended ``" "``, producing double-space output.
   Fix: ``replace_end`` excludes the trailing space; the TUI apply
   checks ``current[replace_end:].startswith(" ")`` and skips the
   separator when the suffix already has one.

3. **Subcommand exact-match confirmation noise** (``_completion_engine.py``)
   Typing ``/mcp list`` + Tab re-inserted ``list`` and the popup
   kept showing the same subcommand. Add a guard mirroring the
   top-level exact-match rule: when the only subcommand match is
   the prefix itself (no trailing space), return ``empty``.

4. **Alphabetical sort dropped in CLI** (``cli/interactive.py``)
   The new completer iterated ``result.candidates`` in registration
   order. Re-add ``sorted(result.candidates, key=lambda c: c.text)``.
   Same sort added to the TUI for consistency.

## Cleanup

- Drop the dead ``on_input_changed`` method from the ``_StubApp``
  test stub (0 call sites) plus the unused ``_slash_commands`` /
  ``_subcommands`` locals that fed it. This addresses din0s's
  comment about the stub duplicating real TUI logic — the inlined
  copy is no longer needed since the real completer now routes
  through the shared engine.

## Tests

- ``test_engine_exact_subcommand_shows_confirmation`` → renamed to
  ``test_engine_exact_subcommand_hides`` to match new behavior.
- New: ``test_engine_subcommand_trailing_space_excludes_space_from_range``
  and ``test_engine_subcommand_trailing_space_apply_does_not_double_space``.
- All 97 tests in ``test_tui_widgets.py`` pass.
- ``ruff check`` / ``ruff format`` clean.
- Local TUI smoke: ``/`` (no crash, top-level popup), ``/mcp ``
  (subcommand popup), ``/mcp a `` + Tab → ``/mcp add ``.

Refs the din0s review comments on PR #273. CLI path tests and the
``category`` ClassVar follow-up are deferred to a separate PR (the
former is a test-suite addition; the latter is already absent from
``base.py`` on the current branch).

* fix: address remaining review items (help duplication, CLI tests, stub sync, docstrings)

- mcp.py: auto-generate help text from subcommands ClassVar (#1)
- tests/test_cli_completion.py: add 9 CLI completer tests (#2c)
- test_tui_widgets.py: sync _apply_selected_completion stub with real code (#4)
- mcp.py + interactive.py: add docstrings to key functions (#8)

* fix: hide completions on exact subcommand match regardless of trailing space

Remove the
ot has_trailing_space guard from the exact-subcommand
check.  Previously /mcp list  (with trailing space) would still
return candidates, causing Tab to oscillate between adding and removing
the trailing whitespace.  Now the engine hides whenever the subcommand
is an exact match, same as the top-level rule.

Added test_engine_exact_subcommand_with_trailing_space_hides to cover
the scenario din0s flagged.

* refactor: use StrEnum for CompletionResult.kind

Replace plain str with CompletionKind(StrEnum) for type safety.
Backward-compatible with existing string comparisons.

* fix: normalize @file completion tuples to CompletionCandidate

complete_file_mention() returns list[tuple[str, str]] but the TUI
rendering/apply code expects objects with .text/.description.
Wrap tuples in CompletionCandidate to prevent AttributeError crash.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-13 17:06:23 +00:00