645 Commits

Author SHA1 Message Date
m4 11181c14d5 fix(release): upload assets via curl multipart, reuse existing release on resume 2026-08-13 11:09:51 +08:00
m4 96c380caa3 fix(release): idempotent bump/tag, push current branch, ignore state file
Docker / build (push) Has been cancelled
2026-08-13 11:03:15 +08:00
m4 59c77c65b1 feat(release): allow release branches via RELEASE_BRANCHES env (default main) 2026-08-13 10:35:55 +08:00
m4 ae461d1c2d chore: bump version to 0.2.3 2026-08-13 10:18:33 +08:00
m4 2a7cccd598 feat(model-registry): add admin config export endpoints with plaintext secrets 2026-08-12 20:07:59 +08:00
m4 aae8d0a379 feat: workspace file references, read-file-images middleware, image model enabled flag
In-progress work committed to unblock the config import/export plan:
- prompts: FILE_REFERENCES section for workspace-relative file citation
- backends: resolve quoted virtual absolute paths onto the sandbox workspace
- middleware: read_file_images middleware; message_budget extensions
- image_gen/model_registry: image model 'enabled' flag refactor
- memory/launch, gateway/background_runs, tools/image follow-ons
- scripts: dev_backend.sh, release.sh
- tests for the above
2026-08-12 19:43:35 +08:00
m4 f3ca381ab3 feat(release): local release script with unified version bump and Gitea publish
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-12 17:25:43 +08:00
m4 8f6a568646 feat(update): system update/status/rollback HTTP endpoints with system:write scope 2026-08-12 16:42:10 +08:00
m4 91e2a87be5 feat(update): rollback version list with BREAKING-DB truncation 2026-08-12 16:34:16 +08:00
m4 03771df508 feat(update): standalone updater runner (wait/install/restart/result) 2026-08-12 16:30:49 +08:00
m4 90e773b30b feat(update): plan builder, plan lock and updater spawn 2026-08-12 16:27:39 +08:00
m4 40b896bcb5 feat(update): deployment detection and install command builder 2026-08-12 16:24:17 +08:00
m4 174f03b92d feat(update): stage artifacts under config dir and reuse verified local files 2026-08-12 16:21:45 +08:00
m4 194402fc88 fix(usage): stop sharing EVOSCIENTIST_DEPLOYMENT_ID with scope partitioning
The usage identity exported the same variable the scope registry reads to
partition workspace scopes, so a backend started with the usage environment
(61d1b61b) could not see scopes provisioned under the workspace-derived id
(5a882492) and every scope lookup 404'd. Usage attribution now reads
EVOSCIENTIST_USAGE_DEPLOYMENT_ID; the scope side keeps the original variable.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-30 22:08:21 +08:00
m4 c6efdaa13f docs(webui): thinking-timer design for the chat message area
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-30 11:11:33 +08:00
m4 08aa0d0e05 feat(registry): mutually-exclusive sampling_override frozen into run snapshots 2026-07-30 10:08:26 +08:00
m4 8eff551bb6 docs(registry): implementation plan for mutually-exclusive sampling_override
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-30 09:48:34 +08:00
m4 862c1e9743 docs(registry): supersede flat temperature/top_p overrides with mutually-exclusive sampling_override
temperature and top_p cannot be set together; the flat-fields design
allowed both. Replaced by a discriminated union (default | temperature
| top_p) where overriding one omits the other from the request entirely.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-30 09:29:06 +08:00
m4 d7711484ef feat(registry): expose generation defaults on selectable models 2026-07-28 17:42:16 +08:00
m4 ba7d908276 feat(registry): freeze per-thread temperature/top_p overrides into run snapshots
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-28 17:23:26 +08:00
m4 e57ecd4588 docs(plan): per-thread temperature/top_p override implementation plan
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-28 16:59:37 +08:00
m4 261845830d docs(spec): per-thread temperature/top_p override design
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-28 16:17:16 +08:00
m4 8efd4ad0ab fix(image-gen): wrap corrupt-config GET in the 422 error envelope 2026-07-24 09:45:57 +08:00
m4 01e674e1bc feat(image-gen): add /api/image-generation config endpoints with masked keys 2026-07-24 09:36:29 +08:00
m4 7f26ecc19a fix(image-gen): sanitize config validation errors; widen expected_revision type
load_image_generation_settings now re-raises pydantic ValidationError as a
sanitized ImageGenError carrying only field locations and error types, so a
mis-indented config.yaml can never echo a literal API key into agent-visible
errors. Also widen save_registry's expected_revision annotation to
int | None to match the http_api caller (value remains ignored).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-24 07:51:36 +08:00
m4 395baab7d5 test(registry): add IMAGE_MODEL_NOT_CHAT_MODEL to error taxonomy mirror 2026-07-24 07:44:37 +08:00
m4 ccf4173990 feat(registry): reject image-only models in chat model saves 2026-07-24 07:36:34 +08:00
m4 a57c52c676 feat(image-gen): add image-artist skill for generation workflow 2026-07-24 07:26:23 +08:00
m4 384bc13a5b feat(image-gen): add generate_image/edit_image agent tools
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-23 22:36:06 +08:00
m4 5622b40cf3 feat(image-gen): add service layer with safe artifact saving 2026-07-23 22:10:14 +08:00
m4 802f71bd46 feat(image-gen): add Gemini (Imagen) image adapter 2026-07-23 21:53:40 +08:00
m4 cc9dfb1cc9 fix(image-gen): address review findings in OpenAI image adapter
Scope download Authorization header to the provider origin, translate
httpx errors in _download/edit into safe ImageGenError messages, and
prevent entry params from clobbering core payload keys; also harden
strip_data_uri against malformed values and drop dead code.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-23 21:44:57 +08:00
m4 12e4f34005 feat(image-gen): add OpenAI-compatible image adapter with responses fallback 2026-07-23 21:28:38 +08:00
m4 b0f9a8d785 feat(image-gen): add image_generation config section and model detection
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-23 21:03:50 +08:00
m4 15cc389b3d feat(registry): make expected_revision optional and ignored in save API
Regenerate the checked-in OpenAPI export to match the relaxed schema.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-23 20:55:52 +08:00
m4 3e67e64067 refactor(registry): make saves last-write-wins, ignore expected_revision
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-23 20:47:08 +08:00
m4 bb9bed82e1 feat(runtime)!: remove auxiliary model role, resolve all roles from run snapshot
ModelRole collapses to "primary": every role (main, tool selector, memory
agents, subagents, summarizer) resolves to the snapshot's frozen primary
model, per design 6.1/8.3 — users typically configure a single usable LLM,
so compile-time auxiliary bindings were bypassing run snapshots and
mis-attributing usage. Legacy auxiliary keys in stored snapshots, registry
JSON, and thread metadata are tolerated on read and dropped.

BREAKING CHANGE: ThreadModelSelection no longer carries an auxiliary ref;
snapshot selection_hash is computed over {primary, reasoning_effort} only;
ConfigurableModelMiddleware(role="auxiliary") is rejected.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-23 10:18:44 +08:00
m4 1a01fb5d74 docs(env): drop stale LLM-key and provider-admin-token entries from .env.example
Provider credentials now live exclusively in the Model Registry and the
x-evoscientist-admin-token / provider-admin-token mechanism was removed;
no code reads these variables anymore.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-21 21:14:31 +08:00
m4 087781556b fix(runtime): serve Config API in bootstrap and verify snapshot issuer by registered deployment set
Graph construction no longer raises on a bootstrap registry: build paths
bind a shared RegistryNotReadyChatModel placeholder that fails every call
with MODEL_REGISTRY_NOT_READY, so langgraph dev serves the Config API for
first-time configuration while run creation stays forbidden.

Run snapshot binding no longer compares configurable
'workspace_deployment_id' (the workspace-isolation scope id) against the
snapshot's issuing deployment — a mismatch that made every BFF run fail
with SNAPSHOT_NOT_FOUND. SnapshotService.get_for_run verifies thread_id
equality plus membership in the platform-registered deployment set
(local_deployment_id + webui_delegation_public_keys entries).

Blocking I/O moved off the event loop for langgraph dev's blockbuster:
Config API authentication (store mkdir/chmod, config.yaml read, jti
registration) and the message-budget snapshot read now run in threads,
with the immutable snapshot cached per run.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-21 20:21:32 +08:00
m4 a1bfbd92ca chore: untrack accidentally committed docx
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-21 18:10:40 +08:00
m4 421a664336 feat(runtime)!: complete legacy removal, local snapshot entries, and TTL cleanup
- Remove legacy provider profiles, admin-token auth, /model command,
  model picker widget, and config.yaml LLM fields (design doc section 10)
- Wire CLI/channels/cron and async sub-agents through the local snapshot
  entry; run creation rejects model config outside runtime_snapshot_id
- Add periodic run-snapshot TTL cleanup to the config service lifespan
- Isolate tests from the real config dir and activate the registry where
  run/model paths fail closed in bootstrap

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-21 18:10:23 +08:00
m4 57176b359a feat(runtime)!: switch middleware and agent factories to snapshot-driven models
Replace config.yaml-driven model selection with registry snapshot resolution
across the runtime chain:

- ConfigurableModelMiddleware reads configurable["runtime_snapshot_id"] only;
  model/model_provider overrides are rejected with MODEL_CONFIG_OUTSIDE_SNAPSHOT
- MessageBudgetMiddleware derives budgets from snapshot reserves
  (system/tools/attachments) and re-resolves the summarizer per snapshot
- Agent factory and subagent factory resolve models via SnapshotRuntime
  (auxiliary/tool_selector/scheduler -> defaults.auxiliary ?? defaults.primary)
- Remove ModelFallbackMiddleware, /model-fallback command, and fallback chain
- Add model_registry/runtime.py SnapshotRuntime glue layer

Legacy config.yaml LLM fields, /model command, and llm/models.py remain for
Task 7. Report: .superpowers/sdd/briefs/task-6-report.md
2026-07-21 12:54:35 +08:00
m4 b2660fc38c feat(model-registry): add provider test API and guarded verification recording
POST /api/model-registry/test (model_config:test) runs the section 9.4
flow: resolve_for_test, per-test credential resolution, build_chat_model
with both safe clients, one minimal chat call, and per-capability probes
(tools/structured_output/vision) whose failures only mark that capability
unverified. Results upsert the model_verifications five-tuple inside a
BEGIN IMMEDIATE transaction that re-checks the registry revision and
configuration hash, returning 409 MODEL_CONFIGURATION_CHANGED on any
concurrent change. resolve_for_test now also relaxes the passing-
verification gate, which the provider test itself produces.
effective_request_options reuses the redacted adapter.build_request
output; the OpenAPI contract and checked-in openapi.json are updated.
2026-07-21 11:39:43 +08:00
m4 940db565b3 fix(model-registry): close review gaps in save-time validation
- add the missing rule-8 counterexample test: declared capabilities
  exceeding the adapter protocol are rejected with
  CAPABILITY_UNSUPPORTED_BY_ADAPTER (all eight section 9.2 checks now
  have at least one negative test)
- raise CREDENTIAL_NOT_CONFIGURED explicitly in _check_enabled_model
  when a required credential reference is null instead of relying on
  resolve_parameters call ordering
2026-07-21 09:38:37 +08:00
m4 dbb6b7abde feat(model-registry): add delegation-JWT auth and config/snapshot HTTP API
- BFF service token (constant-time, plaintext or SHA-256 hash) plus
  X-Evo-Actor delegation JWT verification (ES256/RS256, iss/aud, <=60s
  lifetime, required claims, thread binding) with atomic jti anti-replay
- Config API: GET/PUT /api/model-registry, credential rotation endpoint,
  GET /api/models selector; PUT runs the section 9.2 save-time checks
  inside the registry write transaction after credential writes
- Snapshot API: create/bind/delete routes delegating to SnapshotService
  with thread/deployment binding checks and 9.5 unified error payloads
- Platform security config loader (config.yaml fields), OpenAPI export
  (scripts/export_model_registry_schema.py -> model_registry/openapi.json)
- Mount new routes in langgraph_dev/http.py; retire the legacy
  GET /api/models and POST /api/runtime-snapshots handlers
- Declare PyJWT>=2.8 (previously transitive); extend the 9.5 error code
  table with the HTTP-layer codes (400/401/403/422/500)
2026-07-21 09:28:55 +08:00
m4 c8c46eab16 fix(model-registry): make snapshot abort atomic against concurrent bind
The abort path was read-then-write with an unconditional UPDATE, so a bind
committing between the two calls was clobbered back to aborted, losing its
langgraph_run_id. Add a conditional store-level abort_run_snapshot
(prepared-only UPDATE, rowcount-checked) and re-read on a lost race, matching
the bind loop. Also pin the inherit selection_hash test to a hardcoded
SHA-256 literal instead of reimplementing the serialization in the test.
2026-07-21 08:45:33 +08:00
m4 0cc995eb80 docs(model-registry): add task 4 resolver and snapshot service report 2026-07-21 08:34:50 +08:00
m4 b1233d42dc feat(model-registry): add ModelRegistryResolver and run snapshot service
Resolver (8.1): validates provider/model/credential/capability/limits and
the 6.5 four-mode input budget, freezes ResolvedModelConfig; resolve_for_test
relaxes only the enabled-visibility check (9.4); compute_availability is the
single 4.3 six-state judgement (stale beats configured, selectable only when
enabled).

SnapshotService (8.2, shared by the Task 5 HTTP API and Task 7 local entry):
freezes both roles' full ResolvedModelConfig with adapter spec revision,
fixed reserves, capabilities, and credential revisions; selection-hash
idempotency with pre-resolution semantics; prepared(15min)/bound(+24h)/
expired/aborted lifecycle with atomic bind; binding-checked reads that
revalidate frozen spec revisions; per-call credential resolution against the
frozen revision with no in-process secret cache (5.2); public diagnostic
view limited to the 8.2 safe subset.

Store gains additive helpers (credential pointer lookup, verification
listing, active-triplet lookup, conditional bind, due-expiry sweep) and the
taxonomy gains SNAPSHOT_NOT_FOUND (404) for missing snapshots.
2026-07-21 08:33:12 +08:00
m4 b2e28249fd fix(model-registry): inject async safe clients and split ollama transports
Review fixes for the Task 3 contract layer:

- build_chat_model now accepts http_async_client alongside http_client
  (at least one required) and wires it into ChatOpenAI
  (http_async_client), ChatAnthropic (seeded _async_client), and
  ChatOllama (async_client_kwargs transport), closing the unsafe
  default-async-client gap.
- ChatOllama safe transports move from the shared client_kwargs to
  sync_client_kwargs/async_client_kwargs; langchain-ollama merges shared
  kwargs into both clients, which poisoned the async client with a sync
  transport and crashed ainvoke.
- Unsupported parameters now actually execute the contract-declared
  normalizer (reject_non_auto) instead of a hardcoded raise, with a
  fallback rejection if a normalizer would let a value through.
- build_chat_model rejects overlapping client_options/request_options
  keys instead of silently overwriting.
2026-07-20 22:40:33 +08:00
m4 af4ae1aef5 feat(model-registry): add adapter parameter contracts and build_chat_model factory
Add the Task 3 parameter contract layer (design doc 6.1-6.4):

- adapters.py: versioned built-in contracts for the five phase-1
  adapters plus the openai-compatible/glm-5.2 model-specific contract
  (verbatim section 6.2 values); exact > longest glob > generic
  matching with spec_revision pinning; resolve_parameters implementing
  the section 6.1 inherit/omit semantics, contract validation with
  stable error codes, and named normalizers (identity,
  clamp_to_model_limit, omit_when_none, omit_when_auto,
  reject_non_auto); Adapter.build_request as the single entry point
  mapping ResolvedModelConfig to {client_options, request_options};
  compute_effective_capabilities (protocol AND declared AND verified).
- factory.py: build_chat_model(resolved_config, http_client, *,
  credential=None) with no **kwargs and no setdefault merging; injects
  the safe HTTP client into ChatOpenAI/ChatAnthropic/ChatOllama, never
  reads provider API-key environment variables, and strips the
  OLLAMA_API_KEY authorization header for mode=none adapters.
- tests: per-adapter request-capturing fakes plus an httpx.MockTransport
  outbound capture proving registry resolution matches the wire request.
2026-07-20 22:17:11 +08:00
m4 c46ae17084 feat(model-registry): add EndpointPolicy and SafeHttpTransport SSRF defenses
EndpointPolicy validates provider base URLs (section 4.3): public https
endpoints with hostname and optional port pass; loopback, private,
link-local, multicast, unspecified, and cloud-metadata addresses are
denied unless the normalized URL exactly matches a registered
development_endpoints entry (no prefix or wildcard matching). URLs with
user info, fragments, or non-http(s) schemes are rejected with the new
stable 422 code ENDPOINT_NOT_ALLOWED.

SafeHttpTransport is the single network egress for adapters: a custom
httpcore NetworkBackend resolves DNS under control on every connect
(retries included), filters denied ranges, and connects directly to the
selected IP, while TLS SNI/certificate checks and the HTTP Host header
keep the original hostname. Redirects and env proxies are disabled;
every request origin re-passes URL-layer validation before any I/O.
2026-07-20 21:25:15 +08:00
m4 c21fc0a272 feat(model-registry): add RegistryV4 schema, SQLite store, and unified error codes
New independent subpackage EvoScientist/model_registry implementing the
frozen unified model configuration design (v1.1.0, sections 4.2, 4.3,
5.1, 8.2, 9.5):

- schemas.py: single Pydantic v2 RegistryV4 schema (ModelRef, ProviderConfig,
  ModelConfig, AuthConfig) plus AdapterParameterSpec/AuthSpec/ParameterRule,
  ModelAvailability, ResolvedModelConfig (no secrets), CredentialStatus
- errors.py: all 23 stable error codes from the 9.5 table with HTTP status
  mapping and the unified {code, message, details, request_id} payload
- store.py: ModelRuntimeStore over model-runtime.sqlite3 (0700 dir, 0600
  file, WAL, foreign keys, busy_timeout) with BEGIN IMMEDIATE revision CAS,
  bootstrap->active atomic transition, immutable credential versions with
  masked status, model_verifications upsert, run_runtime_snapshots partial
  unique index, delegation_jtis, and shared-storage lock probe
- hashing.py: normalized SHA-256 configuration_hash

No existing module behavior changed. 89 new tests; full suite passes
(2922 passed, 10 skipped).
2026-07-20 20:43:28 +08:00
m4 8a0ab17936 chore: baseline WIP before unified model configuration implementation
Pre-existing uncommitted work (runtime snapshots, message budget middleware) preserved as baseline.
2026-07-20 20:15:38 +08:00
m4 38668c4ce5 feat: add workspace isolation and provider administration 2026-07-19 12:17:18 +08:00
m4 7a3fcc7c8e Merge remote-tracking branch 'upstream/main'
# Conflicts:
#	README.md
#	uv.lock
2026-07-13 09:46:07 +08:00
dinos 49770949da fix(langgraph): prefer executable in venv over path (#341) 2026-07-11 11:59:27 +00:00
X-iZhang 6f10406d5b chore: update version to v0.2.2
Docker / build (push) Has been cancelled
2026-07-11 01:16:03 +01:00
X-iZhang 9042068094 feat(models): add support for GPT-5.6 variants and update context windows for Grok models 2026-07-11 00:51:06 +01:00
dependabot[bot] 81ff0519dc chore(deps): bump soupsieve in the uv group across 1 directory (#347)
Bumps the uv group with 1 update in the / directory: [soupsieve](https://github.com/facelessuser/soupsieve).


Updates `soupsieve` from 2.8.3 to 2.8.4
- [Release notes](https://github.com/facelessuser/soupsieve/releases)
- [Commits](https://github.com/facelessuser/soupsieve/compare/2.8.3...2.8.4)

---
updated-dependencies:
- dependency-name: soupsieve
  dependency-version: 2.8.4
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-07-10 22:47:18 +01:00
m4 e0acc6155e feat: improve WebUI run recovery
Build / build (push) Has been cancelled
Docker / build (push) Has been cancelled
Lint / ruff (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.11) (push) Has been cancelled
Test / pytest (ubuntu-latest, 3.12) (push) Has been cancelled
Test / pytest (windows-latest, 3.11) (push) Has been cancelled
Test / pytest (windows-latest, 3.12) (push) Has been cancelled
2026-07-10 17:35:44 +08:00
dinos 690b903f85 test: standardize async tests on pytest-asyncio auto mode (#338)
* chore: add pytest-asyncio in auto mode

* test: migrate channel and stream tests to native async

Convert run_async() wrapper tests to plain 'async def test_*' under
pytest-asyncio auto mode. collect_events() in stream_v3_fakes becomes a
coroutine awaited at every call site.

* test: migrate command and model/middleware tests to native async

Convert run_async() wrappers (import, alias, and fixture forms) to plain
'async def test_*'. Multi-call tests merge onto one loop as sequential
awaits; none asserted on loop identity.

* test: migrate TUI, notifier, gateway, and session tests to native async

TUI/notifier/gateway files convert run_async wrappers to plain async
tests. test_sessions.py's unittest.TestCase classes move to
unittest.IsolatedAsyncioTestCase (pytest-asyncio does not await async
methods on plain TestCase; converting blindly would have made ~70 tests
silently vacuous). Its setUpClass keeps a one-shot asyncio.run() since
IsolatedAsyncioTestCase has no async class-level hook. TestLoadingWidget
in test_tui_widgets.py drops its TestCase base for the same reason.

* test: replace direct asyncio.run() calls with native async tests

Convert tests that called asyncio.run() (directly or via a local _run
helper) to plain 'async def test_*'; delete the local helpers.

* test: drop undeclared anyio markers and delete run_async helper

The @pytest.mark.anyio tests relied on anyio being a transitive dep of
httpx; auto-mode pytest-asyncio collects them natively. run_async() and
its fixture are unreferenced after the migration, so remove them —
pytest-asyncio's per-test loop teardown covers the pending-task
cancellation the helper existed for (verified: full suite runs with no
'Event loop is closed' errors or destroyed-task warnings).
2026-07-08 18:37:48 +00:00
dinos d2452c54d5 Refactor onboarding OAuth flow for auxiliary models (#337)
* refactor(onboard): shared flow for ccproxy providers

* feat(onboard): support oauth configuration for auxiliary models

* fix(onboard): reuse main model auth for same-provider auxiliary

* fix(onboard): reconcile oauth providers
2026-07-08 18:28:44 +00:00
dinos a7b9e175c1 fix(config): set config.yaml permissions to 0x600 (#336) 2026-07-08 19:25:01 +01:00
dinos be3dd272c3 test: deflake timing-dependent tests (#335)
* test: deflake timing-dependent tests

Inject a clock into channel dedup tests, replace fixed async sleeps with
events/explicit flushes, and avoid wall-clock waits in background tests.

* coderabbit nit
2026-07-07 08:25:35 +01:00
X-iZhang df54d8498c chore: update version to 0.2.1 2026-07-05 10:17:28 +01:00
Wiktor Cupiał 1d117ff277 feat: completion enchancements (#302)
* feat: completion enchancements

* fix: handle exception

* fix: duplicate view

* fix: remove deadcode

* fix tab
2026-07-05 05:14:42 +00:00
renaissancefieldlite f086d77756 Fix UTF-8 config reads on Windows (#318)
* Fix UTF-8 config reads on Windows

* test: cover utf8 production loaders

* fix: read and write settings as utf8

* Apply ruff formatting
2026-07-05 05:10:24 +00:00
jfilipiuk 5dacab7e4c fix: bump langchain-openrouter to 0.2.5 for upstream OpenRouter duplicate-terminal-chunk fix (#313) 2026-07-05 05:01:35 +00:00
kalisgd0h bd54eaa0a4 feat(cli): add --output-format stream-json for headless clients (#309)
* feat(cli): add --output-format stream-json for headless clients

Emit EvoScientist's native event stream as line-delimited JSON on stdout
in single-shot (-p) mode, with all human output redirected to stderr so
stdout stays pure JSONL. Intended as the integration surface for
programmatic clients (e.g. an agent runtime) that drive EvoSci headlessly.

- stream/json_sink.py: write_events_as_json + stream_json sink, plus
  redirect_console_to_stderr helper for stdout purity
- cli/interactive.py: cmd_run gains output_format; stream-json branch runs
  the sink instead of the Rich renderer
- cli/commands.py: --output-format option + validation (stream-json
  requires -p; value must be text|stream-json)
- docs/stream-json.md: event-schema contract + example transcript
- tests: json sink serialization, CLI dispatch, console redirect, validation

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(cli): honor explicit --no-auto-mode over config in stream-json

Address CodeRabbit review (discussion_r3514041123): the auto-mode override
block only wrote to cli_overrides when the resolved value was True, so an
explicit --no-auto-mode silently fell back to a config that enables
auto-mode -- breaking "explicit flags always win" and leaving stream-json
running unattended despite the warning. Write auto_mode=False when the flag
is explicitly False. Add regression tests that capture the overrides.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-05 04:53:52 +00:00
dinos 2b244888ec feat(memory): autoskills (#319) 2026-07-03 11:16:33 +02:00
houren Antony 7568a6bc1f fix(tui): keep welcome banner at top after /new (#311)
* fix(tui): keep welcome banner at top after /new

PR #262 replaced scroll_end() with anchor() for free-scrolling.
When /new clears a long anchored conversation, the anchor kept
the viewport pinned to the (now empty) bottom, producing a
negative scroll_y and pushing the welcome banner out of view.

Reset the anchor and scroll to the top in clear_chat(), and
restore the follow/new-content flags so the fresh session starts
correctly.

Closes #301

* fix(tui): suppress anchor when chat content fits viewport

The previous fix for #301 only handled the /new path. din0s reported
that the banner still dropped to the bottom after a normal short turn
(user types 'hi', agent replies) — i.e. whenever the conversation
fit in the viewport. Root cause is in Textual's compositor
(textual._compositor): when a widget is anchored, scroll_y is
recomputed via set_reactive, which bypasses the validator. If the
anchored widget's content is shorter than the viewport, scroll_y
goes negative on the next layout pass and the welcome banner is
pushed below the visible region.

PR #262 made _stream_with_widgets re-engage the anchor at the end of
every turn via _anchor_chat, so the bug surfaced on any short reply
that fit in the viewport. Markdown re-renders, status-bar updates,
or any subsequent mount would then trip the compositor.

Fix in three places:

  * _anchor_chat: only engage the anchor when max_scroll_y > 0;
    otherwise release and scroll_home so the banner stays at the top.
  * streaming anchor loop: if content shrinks below the viewport
    mid-stream (e.g. loading widget removed), release the anchor
    instead of leaving _anchored=True for the compositor to trip on.
  * clear_chat: keep the unconditional reset (children are removed
    asynchronously so a max_scroll_y check would be stale) but
    document why.

Adds two regressions:
  * test_short_turn_keeps_banner_at_top_after_layout_refresh — the
    exact scenario din0s tested; fails with scroll_y=-10 on the
    previous code, passes with the fix.
  * test_long_turn_keeps_viewport_pinned_to_bottom — guards against
    regressing free-scrolling for overflowing conversations.

Manually verified: 'hi' -> reply (banner stays at top) -> /new
(banner at top) -> another turn (banner stays at top).

* test(tui): address review feedback on banner-position regressions

- extract `_release_anchor_and_pin_top` helper for the 3-line
  `anchor(False) + scroll_home(...)` pattern repeated in
  `clear_chat`, `_anchor_chat`, and the streaming loop
- replace `pytest.skip` in `_capture_app` with a hard `RuntimeError`
  so a broken capture never silently passes
- drop the redundant `load_agent` and `create_session_workspace`
  monkeypatches (the factory is given those as parameters, so the
  module-level symbols never run; added a comment explaining why)
- add a defensive `_FakeChannelRuntime` patch for symmetry with the
  other module-level fakes
- drop the local `_run` helper and use the `run_async` fixture from
  `conftest.py` (its teardown is better)

* test(tui): replace _FakeChannelRuntime with _auto_start_channel no-op

The _FakeChannelRuntime patch was ineffective because ChannelRuntime is
just a dataclass — the real channel manager still started via
_auto_start_channel, leaving pending tasks and non-hermetic test state.

Per review feedback, stub _auto_start_channel directly instead.
2026-06-30 19:15:08 +01:00
X-iZhang 682922f690 docs: update changelog for version 0.2.0 release 2026-06-27 00:43:17 +01:00
dependabot[bot] 598b668131 chore(deps): bump langgraph-checkpoint (#310)
Bumps the uv group with 1 update in the / directory: [langgraph-checkpoint](https://github.com/langchain-ai/langgraph).


Updates `langgraph-checkpoint` from 4.1.0 to 4.1.1
- [Release notes](https://github.com/langchain-ai/langgraph/releases)
- [Commits](https://github.com/langchain-ai/langgraph/compare/checkpoint==4.1.0...checkpoint==4.1.1)

---
updated-dependencies:
- dependency-name: langgraph-checkpoint
  dependency-version: 4.1.1
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-27 00:23:16 +01:00
X-iZhang 6546f4022f chore: update version to v0.2.0 2026-06-26 23:55:46 +01:00
jfilipiuk 7214d4099d feat: expose model registry at GET /api/models (#308)
* feat: expose model registry at GET /api/models

* fix: include Ollama models in /api/models endpoint

* fix: honor env vars override in /api/models endpoint

* fix: offload get_effective_config to thread to satisfy blockbuster

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-26 22:48:37 +01:00
dinos f2f010a350 feat(memory): observation linking (#307)
* refactor(gateway): create module for launching async/bg agents

* refactor(memory): refactor worker launch around source context & output deltas

* refactor(gateway): generalize async/bg module

* refactor(memory): revamp worker launching

* feat(memory): add observation linking

* test(memory): remove redundant test branches

* fix(memory): make 'supersedes' relation directional

* fix(memory): don't create empty project observation dirs

* fix(memory): schedule direct observations for linking

* fix(cli): wait for observation linker before shutdown

* fix(memory): block arbitrary writes to /memories

* fix(linker): remove `linked_by` attribute from frontmatter

* refactor(linker): rename base relationship to `comlpements`

* fix(cli): bump worker wait to 2m

* feat(tools): catch malformed tool calls & retry

* feat(status): add linking result to statusbar

* fix(linker): don't launch linker when observations are disabled

* fix(memory): use posix paths

* fix(watcher): call abort hook on error status

* fix(watcher): delete thread on failed run creation

* fix(watcher): preserve url

* fix(observation): record session_id, drop unused fields

* fix(memory): reject unsupported worker source types

* refactor(backends): shared memory backend builder

* fix(scheduler): resolve linker inputs outside lock

* fix(memory): dont launch workers / record observations without thread_id

* feat(memory): include related observations in tool results

* fix(memory): skip malformed observation frontmatter

* revert(tools): drop tool error handling changes from this PR

* fix(memory): serialize observation link writes

* fix(memory): queue observations written by aborted workers

* fix(memory): track observation linker launch handoff

* fix(memory): resolve cross-project related observations

* fix(status): avoid recounting reason-only link updates

* fix(memory): avoid rereading file for content

* fix(linker): use neutral prose for bidirectional reasons

* test(memory): coverage for aborted/failed launches

* test(memory): cleanup & helpers

* feat(linker): add observations index hint
2026-06-26 22:20:52 +01:00
Xi Zhang 7ccfe68f3f feat: add scheduler functionality with cron-style task management (#306)
* feat: add scheduler functionality with cron-style task management

- Implemented a new scheduler subagent to automate recurring tasks using cron expressions.
- Enhanced the subagent factory to include the skill manager and auxiliary chat model for the scheduler.
- Created a YAML configuration for the scheduler with a detailed system prompt and toolset.
- Updated README files to include documentation on scheduled tasks and usage examples.
- Added tests for the scheduler, including command execution, scheduling tools, and middleware integration.
- Introduced new dependencies for timezone handling and ensured compatibility in the project configuration.

* fix(async-notifier): ensure fallback hint is used for unknown notification kinds

* feat: enhance scheduling functionality and improve system message handling
2026-06-25 17:31:23 +01:00
dinos b1dccf17ea fix(tool-selector): memory tools & state for main agent (#305)
* fix(tool-selector): always include memory tools

* fix(tool-selector): only track & show state for main agent

* docs: update docstring
2026-06-23 14:47:25 +01:00
X-iZhang c063a00c8f Add regression tests for code_interpreter middleware PTC allowlist
- Ensure 'task' is excluded from the default PTC allowlist to prevent ValueError in langchain-quickjs >=0.3.
- Verify that essential async dispatch tools remain in the allowlist.
- Confirm that the live quickjs filter accepts the default allowlist even with a 'task' tool present.
- Test the creation of the code_interpreter middleware to ensure it builds correctly.
2026-06-23 08:40:35 +01:00
dependabot[bot] 9460b9ecd6 chore(deps): bump the uv group across 1 directory with 2 updates (#304)
Bumps the uv group with 2 updates in the / directory: [langsmith](https://github.com/langchain-ai/langsmith-sdk) and [pydantic-settings](https://github.com/pydantic/pydantic-settings).


Updates `langsmith` from 0.8.15 to 0.8.18
- [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases)
- [Commits](https://github.com/langchain-ai/langsmith-sdk/compare/v0.8.15...v0.8.18)

Updates `pydantic-settings` from 2.14.0 to 2.14.2
- [Release notes](https://github.com/pydantic/pydantic-settings/releases)
- [Commits](https://github.com/pydantic/pydantic-settings/compare/v2.14.0...v2.14.2)

---
updated-dependencies:
- dependency-name: langsmith
  dependency-version: 0.8.18
  dependency-type: indirect
  dependency-group: uv
- dependency-name: pydantic-settings
  dependency-version: 2.14.2
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-22 18:16:45 +01:00
X-iZhang 57461f671f chore: update version to v0.1.8 2026-06-22 18:10:21 +01:00
dinos bd307f3a11 refactor: LangGraph gateway layer for UI-agnostic graph and thread access (#295)
* feat(gateway): graph gateway protocol

* refactor(cli): wire gateway in cli/tui

* refactor(gateway): centralize runtime gateway init

* chore(gateway): restrict RunRequest message type

* feat(gateway): add langgraph server gateway

* chore(cli): tighten serve runtime state typing

* refactor(cli): route async task state reads through graph gateway

* refactor(gateway): support graph targets in server gateway

* refactor(cli): route session commands through graph gateway

* refactor(cli): fold thread store under graph gateway

* refactor(gateway): route graph state access through gateway

* refactor(channels): wire graph gateway

* refactor(memory): preserve graph threads for cloning

* feat(gateway): add thread cloning

* fix(tui): pass effective workspace for thread creation

* chore(memory): add workspare dir to memory worker metadata

* fix(sessions): filter preloaded UUID registy entries by the current scope

* test(fakes): use https

* refactor(consumer): consolidate imports

* fix(stream): optional summarization event

* fix(gateway): resolve abbreviated thread IDs by search

* fix(gateway): page server thread listings

* fix(gateway): emit pending interrupt events

* style: fmt

* feat(gateway): persist workspace_dir & model in thread metadata

* fix(gateway): page server thread prefix resolution

* fix(gateway): expose server thread list metadata

* refactor: add back type def

* refactor: tighten types

* revert: add back worker thread deletion

The worker thread forking changes are out of scope for now, so to
maintain parity with the existing behavior we'll leave this intact.

* fix(gateway): apply compaction to server thread history

* refactor(stream): restore direct summary replay suppression

* fix(gateway): preserve compaction state and server stream output

* fix(gateway): close local stream generator on cancellation

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-22 13:54:14 +00:00
dinos 6eb467e70b feat(llm): make OpenRouter Anthropic prompt cache opt-out (#299)
* feat(llm): make OpenRouter Anthropic prompt cache opt-out

* fix(llm): restore truthy env flag helper
2026-06-18 15:09:22 +02:00
Ziheng Zhang 6593ac9b5b fix(cli): submit slash command on Enter when name prefixes another (#293) (#300)
Typing `/model` and pressing Enter did nothing in the TUI; the picker
only opened via `/model --save` or `/model <name>`. The completion popup
matched both `/model` and `/model-fallback` by prefix, so the
exact-match-hide guard (which required a single match) never fired. With
the popup still visible, the TUI's Enter handler completed the text
instead of submitting the command, so it never executed.

Treat the typed prefix as an exact match whenever it equals any matched
command name, not only when it is the sole match. This hides the popup
on a complete command name so Enter submits it, even when a longer
command shares the prefix.
2026-06-18 14:07:43 +01:00
dinos 6894d45be8 chore: bump ruff in pre-commit (#298)
* chore(pre-commit): sync ruff version with uv

* chore(gitignore): don't ignore built-in skills dir

* style: unused var
2026-06-18 12:46:09 +01:00
X-iZhang 086425377d chore: update version to v0.1.7 2026-06-16 23:13:17 +01:00
dinos de3f588fbf fix: windows async MCP tool execution, restore graph state after interruptions (#290)
Co-authored-by: z00827015 <zhoulun1@huawei.com>
2026-06-16 21:03:26 +02:00
dependabot[bot] 5282a5028b chore(deps): bump the uv group across 1 directory with 4 updates (#291)
---
updated-dependencies:
- dependency-name: aiohttp
  dependency-version: 3.14.1
  dependency-type: direct:production
  dependency-group: uv
- dependency-name: cryptography
  dependency-version: 48.0.1
  dependency-type: direct:production
  dependency-group: uv
- dependency-name: python-multipart
  dependency-version: 0.0.31
  dependency-type: indirect
  dependency-group: uv
- dependency-name: starlette
  dependency-version: 1.3.1
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-16 11:21:10 +01:00
dependabot[bot] 42d8b9bf10 chore(deps): bump pyjwt in the uv group across 1 directory (#289)
Bumps the uv group with 1 update in the / directory: [pyjwt](https://github.com/jpadilla/pyjwt).


Updates `pyjwt` from 2.12.1 to 2.13.0
- [Release notes](https://github.com/jpadilla/pyjwt/releases)
- [Changelog](https://github.com/jpadilla/pyjwt/blob/master/CHANGELOG.rst)
- [Commits](https://github.com/jpadilla/pyjwt/compare/2.12.1...2.13.0)

---
updated-dependencies:
- dependency-name: pyjwt
  dependency-version: 2.13.0
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-16 09:56:31 +00:00
dinos f356de36a6 feat: memory retrieval (#281) 2026-06-16 09:14:29 +02:00
X-iZhang 8ea186441a chore(deps): bump deepagents 0.6.8→0.6.10 + langchain-quickjs →0.2.0 2026-06-13 18:43:51 +01:00
houren Antony 477481f9d9 feat(backends): use _platform_quote for Windows cmd.exe compatibility (#280)
* feat(backends): use _platform_quote for Windows cmd.exe compatibility

Resolves the 3 skipped E2E tests in test_backends.py that exercised
the /skills/... mount path. The path-rewriter was wrapping resolved
absolute paths via shlex.quote (POSIX single-quote style); cmd.exe
doesn't strip single quotes, so the literal ' characters ended up in
the subprocess argv and the python script failed to find its file.

Replace the 3 shlex.quote call sites in _resolve_virtual_mount_path
with _platform_quote, a thin platform dispatcher:
- POSIX: shlex.quote (unchanged)
- Windows: _cmd_quote uses cmd.exe-compatible double-quote wrapping
  and properly escapes embedded " and percent signs

Adds:
- backends.py: _is_windows, _cmd_quote, _platform_quote (~40 lines)
- test_backends.py: 6 TestPlatformQuote unit tests + _split_cmd
  cross-platform tokenizer helper to replace shlex.split in the 8
  sites that tokenize convert_virtual_paths_in_command results
  (POSIX shlex strips backslashes from bare Windows paths, which
  broke the 5 TestVirtualMountResolution assertions on Windows)

Removes:
- 3 @pytest.mark.skipif(sys.platform == "win32") markers on the
  E2E tests for /skills/... mount resolution

Refs #274.

* fix: escape % as %% in _cmd_quote instead of relying on double-quoting

cmd.exe expands %VAR% before processing quotes, so double-quoting
cannot neutralize percent signs.  Escape bare % as %% (the cmd.exe
idiom for a literal percent) before any other quoting logic.

Also updates _cmd_quote docstring and _resolve_virtual_mount_path
docstring to reflect the actual quoting strategy.

* style: fix ruff format (single → double quotes)

* fix: treat % as regular char in _cmd_quote, document limitation

%% escaping only collapses in .bat/.cmd files, not via cmd /c.
Since virtual-mount paths should never contain % in practice,
simpler to leave % alone and document the caveat.
2026-06-13 18:08:18 +01:00
houren Antony 76972449c7 feat(cli): multi-stage slash command completions with subcommand awareness (Phase 1 of #82) (#273)
* feat(cli): multi-stage slash command completions with subcommand awareness

Phase 1 of #82 — subcommand and argument awareness in completions.

- commands/base.py: add SubCommand dataclass and subcommands/category
  ClassVars to the Command ABC. Each SubCommand has name, description,
  and optional arguments.

- commands/manager.py: add get_subcommands() and list_subcommands()
  methods to expose subcommand metadata for completion rendering.

- commands/implementation/mcp.py: declare 6 subcommands (list, config,
  add, edit, remove, install).

- commands/implementation/model_fallback.py: declare 6 subcommands
  (list, add, remove, clear, save, help).

- commands/implementation/channel.py: declare 2 subcommands
  (status, stop).

- cli/tui_interactive.py: rewrite on_text_area_changed slash-completion
  branch. When the user types a command name + trailing space and the
  command has subcommands, show subcommand completions instead of hiding
  the popup. Filter subcommands by typed prefix in multi-token input.

- commands/implementation/general.py: /help now lists subcommands
  below each command that declares them.

Tests: 8 new tests covering SubCommand creation, CommandManager
subcommand lookup, and cross-command verification.
2293 passed baseline, no regressions.

* fix: subcommand completion preserves prefix + prompt_toolkit + tests

- _apply_selected_completion: preserve '/mcp ' prefix when completing
  subcommands via _comp_is_subcommand flag
- SlashCommandCompleter (Rich CLI): add subcommand completion support
- Fix trailing-space bug: rstrip prefix before top-level matching
- test_tui_widgets.py: update stub on_input_changed to match multi-stage
  logic; add 5 new subcommand tests
- test_command_manager.py: 8 tests for SubCommand + CommandManager

28 passed, 0 failed.

* style: ruff format tui_interactive.py + test_tui_widgets.py

* style: fix RUF012 ClassVar annotation on subcommands lists

* fix: sync test stub, add len>=3 guard, remove exact-match hide

- Sync test stub on_input_changed with real TUI code (remove exact-match
  hide for subcommands, add len(parts)>=3 guard)
- Update test_input_changed_exact_subcommand_hides -> shows_confirmation
- Add test_input_changed_three_parts_hides
- Remove unused category ClassVar (din0s: what is this for)

* refactor(commands): extract shared completion engine

Per din0s feedback: one shared completion engine (commands/_completion_engine.py)
that parses text + cursor once, returns structured CompletionCandidate objects
with replace_start/replace_end ranges.

- SlashCommandCompleter (Rich CLI): thin adapter, delegates to engine
- on_text_area_changed (TUI): thin adapter, delegates to engine
- _apply_selected_completion: uses candidate.replace_start/replace_end instead
  of _comp_is_subcommand flag
- Tests: engine tested directly (10 new tests), stub methods updated

29 passed, 0 failed.

* style: ruff format

* fix: preserve text after cursor when applying completion

CodeRabbit: replace_start only cuts from start to cursor,
dropping any suffix after the cursor. Use replace_start + replace_end
to correctly splice the replacement while preserving trailing text.

* fix(cli): repair slash-command completion (TUI crash, subcommand bugs, sort)

Apology + context: the previous push shipped a TUI-breaking change
(the new shared engine assumed ``event.text_area.cursor_position``
existed, but ``ChatTextArea`` / ``Changed`` don't expose it). User
caught the crash on ``/``; fixing that surfaced two more bugs in
the engine that din0s had already flagged. This commit addresses
all of them and drops a piece of dead stub code.

## Bug fixes

1. **TUI crash on ``/``** (``tui_interactive.py:2335``)
   ``event.cursor_position`` doesn't exist on the ``Changed`` event,
   and ``ChatTextArea`` (Textual ``TextArea`` subclass) doesn't expose
   ``cursor_position`` either. Pass ``len(event.text_area.text)``
   instead — the user types at the end of the input in practice.

2. **Subcommand trailing-space duplication** (``_completion_engine.py``)
   Typing ``/mcp a `` + Tab produced ``/mcp aadd``. The engine
   included the trailing space in ``replace_end``; the TUI apply
   unconditionally appended ``" "``, producing double-space output.
   Fix: ``replace_end`` excludes the trailing space; the TUI apply
   checks ``current[replace_end:].startswith(" ")`` and skips the
   separator when the suffix already has one.

3. **Subcommand exact-match confirmation noise** (``_completion_engine.py``)
   Typing ``/mcp list`` + Tab re-inserted ``list`` and the popup
   kept showing the same subcommand. Add a guard mirroring the
   top-level exact-match rule: when the only subcommand match is
   the prefix itself (no trailing space), return ``empty``.

4. **Alphabetical sort dropped in CLI** (``cli/interactive.py``)
   The new completer iterated ``result.candidates`` in registration
   order. Re-add ``sorted(result.candidates, key=lambda c: c.text)``.
   Same sort added to the TUI for consistency.

## Cleanup

- Drop the dead ``on_input_changed`` method from the ``_StubApp``
  test stub (0 call sites) plus the unused ``_slash_commands`` /
  ``_subcommands`` locals that fed it. This addresses din0s's
  comment about the stub duplicating real TUI logic — the inlined
  copy is no longer needed since the real completer now routes
  through the shared engine.

## Tests

- ``test_engine_exact_subcommand_shows_confirmation`` → renamed to
  ``test_engine_exact_subcommand_hides`` to match new behavior.
- New: ``test_engine_subcommand_trailing_space_excludes_space_from_range``
  and ``test_engine_subcommand_trailing_space_apply_does_not_double_space``.
- All 97 tests in ``test_tui_widgets.py`` pass.
- ``ruff check`` / ``ruff format`` clean.
- Local TUI smoke: ``/`` (no crash, top-level popup), ``/mcp ``
  (subcommand popup), ``/mcp a `` + Tab → ``/mcp add ``.

Refs the din0s review comments on PR #273. CLI path tests and the
``category`` ClassVar follow-up are deferred to a separate PR (the
former is a test-suite addition; the latter is already absent from
``base.py`` on the current branch).

* fix: address remaining review items (help duplication, CLI tests, stub sync, docstrings)

- mcp.py: auto-generate help text from subcommands ClassVar (#1)
- tests/test_cli_completion.py: add 9 CLI completer tests (#2c)
- test_tui_widgets.py: sync _apply_selected_completion stub with real code (#4)
- mcp.py + interactive.py: add docstrings to key functions (#8)

* fix: hide completions on exact subcommand match regardless of trailing space

Remove the
ot has_trailing_space guard from the exact-subcommand
check.  Previously /mcp list  (with trailing space) would still
return candidates, causing Tab to oscillate between adding and removing
the trailing whitespace.  Now the engine hides whenever the subcommand
is an exact match, same as the top-level rule.

Added test_engine_exact_subcommand_with_trailing_space_hides to cover
the scenario din0s flagged.

* refactor: use StrEnum for CompletionResult.kind

Replace plain str with CompletionKind(StrEnum) for type safety.
Backward-compatible with existing string comparisons.

* fix: normalize @file completion tuples to CompletionCandidate

complete_file_mention() returns list[tuple[str, str]] but the TUI
rendering/apply code expects objects with .text/.description.
Wrap tuples in CompletionCandidate to prevent AttributeError crash.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-13 17:06:23 +00:00
houren Antony 8f9dfa159e fix(backends): rewrite quoted virtual paths containing whitespace (#269)
* fix(backends): rewrite quoted virtual paths containing whitespace

The `convert_virtual_paths_in_command` regex
`(?<=\s)/[^\s;|&<>'"`]*` stopped at the first whitespace or quote,
so:

  - `python "/skills/my skill/main.py"` was left completely
    unchanged (the `(?<=\s)` lookbehind failed after the opening
    `"`), and the shell then broke the inner unquoted path at the
    embedded space.
  - `python /skills/my skill/main.py` was truncated to
    `python ./skills/my skill/main.py` (only `/skills/my` rewritten).

Replace the regex with `shlex.shlex(command, posix=True,
punctuation_chars=";|&<>")` so quoted regions stay whole, then
splice the rewrite back into the original command — extending the
splice span to include any matching quote chars around the path so
the fresh `shlex.quote` of the replacement isn't double-wrapped.

`_resolve_virtual_mount_path` now returns the unquoted path; the
caller owns shell-quoting, which avoids the previous
`shlex.quote` inside original `"…"` leaving literal `'` chars in
the argument value.

Unquoted paths with embedded whitespace remain a known limitation
(shlex has no way to know the user meant one path) — the
workaround of avoiding spaces in skill directory names still
applies, as flagged in the original issue.

Closes #237

* fix(backends): backslash-escaped paths, multi-path per token, subshell paths

- Fix backslash-escape handling: use unescape before rewriting
- Fix re.search→re.finditer: all /-paths in a token are rewritten
- Keep ( ) and backticks inside word tokens so paths spanning
  \ or wrapped in backticks are matched correctly
- Add _try_rewrite helper with URL detection and unescape logic
- Add 10 contract tests pinning the din0s review cases

* fix(backends): restore () and backtick as shell operators for validate_command

- Restore ( ) and backtick to the operator set in _shell_token_spans.
  Removing them caused a security regression: commands like (sudo ls)
  would not detect sudo as a blocked command because (sudo became one
  word token. With operators restored, validate_command correctly
  catches blocked commands inside subshells and command substitutions.

- Fix _value_span_to_raw_span: the 'quoted' flag from the tokenizer
  means the token *contains* a quoted segment (not necessarily starts
  with a quote). Replace raw[0] assumption with a forward scan for
  the first quote char, consuming unquoted prefix chars 1:1.

- Update test_system_path_with_shell_expansion:  paths are now
  partially rewritten because () are operators. Test updated to
  reflect this known limitation (security >  path rewriting).

* fix(test): cross-platform compatibility for pre-existing Windows failures

- python3 -> python in execute() calls (python is on PATH in any activated venv)
- sleep 10 -> _sleep_cmd(10) cross-platform helper
- str().endswith() -> Path().parts assertions (backslash-safe on Windows)
- shlex.quote exact-match assertions -> 'in' assertions (Windows quotes paths differently)
- mkdir -p E2E test -> preprocessor boundary test
- Skip 3 E2E tests on Windows: shlex.quote produces POSIX quoting incompatible with cmd.exe

141 passed, 3 skipped on Windows.

* fix: update docstring + strengthen shell-expansion test assertion

- Fix _value_span_to_raw_span docstring: no longer assumes raw[0] is
  the opening quote, scans forward for first quote char
- Strengthen test_system_path_with_shell_expansion: verify
  ./workspace/notes is rewritten, not just notes in result

* style: ruff format backends.py + test_backends.py

* refactor(backends): simplify quoted virtual path rewriting

Replace 500+ line shlex tokenizer with 12-line pre-process step. Match quoted args via regex, unescape, rewrite via _rewrite_quoted_path, substitute with shlex.quote. 133 passed, 3 skipped.

* fix: guard bare absolute paths from double-rewrite by post-process regex

On POSIX, shlex.quote returns bare paths (e.g. /tmp/memories/note.md).
The pre-process substitutes these into the command, then the post-process
regex re-matches and incorrectly rewrites them.

Fix: _guard_bare_absolute wraps bare /-paths in single quotes so the
post-process regex''s character class stops at the quote char.

* style: ruff format

* fix(backends): narrow pre-process to exclude system-prefixed paths

Only rewrite quoted paths that are NOT known system prefixes.

* fix: narrow quoted-path pre-process to virtual mounts only

Only rewrite quoted /... paths that resolve to actual virtual mounts (/skills/..., /memories/...) or workspace-prefixed system paths. Remove catch-all that incorrectly rewrote bare paths like echo /hi.

Addresses din0s review feedback on #269.

* docs: update docstring for narrower quoted-path rewrite scope
2026-06-13 17:56:30 +01:00
X-iZhang d82f0ed2a4 chore: update version to v0.1.6 2026-06-12 00:08:30 +01:00
Xi Zhang 05a5e5f8a0 fix: session lost after evoscientist restart (supersedes #278) (#279)
* fix: session lost after evoscientist restart

* feat: Implement memory worker thread deletion on completion

- Added synchronous and asynchronous functions to delete memory worker threads after they finish execution, ensuring no residual checkpoints are left in the database.
- Enhanced `_watch_memory_worker_run_sync` and `_watch_memory_worker_run_async` to invoke deletion functions upon confirming worker completion.
- Introduced tests to verify that worker threads are deleted correctly upon successful completion and that failures in deletion do not affect the overall worker status.
- Updated session management to ensure that only relevant threads are restored from the database, preventing exposure of internal or unrelated workspace threads.
- Implemented a purge function to clean up leftover worker checkpoints during server startup.

* feat: Implement short thread ID display for CLI and session hints

* fix: ensure proper accounting and deletion order for memory worker threads

---------

Co-authored-by: z00827015 <zhoulun1@huawei.com>
2026-06-11 23:51:51 +01:00
X-iZhang 526c571b10 feat(tunnel): add Cloudflare tunnel support for EvoSci deploy and update documentation 2026-06-11 00:23:55 +01:00
X-iZhang 49f23560fd chore: update version to v0.1.5 2026-06-10 22:37:52 +01:00
Xi Zhang c02be519f6 feat(dangerous-mode): implement real-filesystem access with safety ch… (#276)
* feat(dangerous-mode): implement real-filesystem access with safety checks

- Introduced a 'dangerous mode' allowing the agent to operate on the real filesystem.
- Updated command validation to bypass path confinement while enforcing a blocklist for privileged commands.
- Added warnings and guidelines for users when operating in dangerous mode.
- Enhanced configuration to support dangerous mode and ensure it implies auto-approval.
- Updated tests to verify the behavior of commands and configurations in dangerous mode.

* feat(dangerous-mode): enhance logging and environment management for dangerous mode

* feat(dangerous-mode): improve handling of dangerous mode with environment flags and enhance test isolation
2026-06-10 18:19:13 +01:00
X-iZhang fc05b5eed2 fix(tests): ensure _check_npx is True in _patch_all_questionary to prevent extra prompts on slow runners 2026-06-10 16:11:49 +01:00
X-iZhang c438966b28 fix(tests): mock _ensure_npx in TestStepSkills to prevent integration issues on headless Windows CI 2026-06-10 15:59:10 +01:00
houren Antony d4f1fbd110 ci: add windows-latest to test matrix + fix 11 cross-platform test bugs (#271)
* ci: add windows-latest to test matrix + fix 11 cross-platform test bugs

The test workflow ran on ``ubuntu-latest`` only. Per the issue's
first bullet — the maintainer's explicit #1 priority — add
``windows-latest`` to the matrix so the manager and related
modules are exercised on Windows on every PR.

The matrix addition surfaces 18 pre-existing Windows-only test
failures. Without fixes the new leg would be 18+ reds from
day one and the matrix would just produce a wall of
``fail-fast`` noise. This PR fixes 11 of them; each fix is
a real (cross-platform) bug, not a Windows-specific hack —
most were already flagged by CodeRabbit on PR #236 but never
acted on. The remaining 4 failures need code refactors
(``os.killpg`` → ``psutil`` in ``background.py``,
``convert_virtual_paths_in_command`` Windows-aware quoting,
tilde expansion) that are documented as out-of-scope
follow-ups below.

## What changed

* ``.github/workflows/test.yml``
  - ``os: [ubuntu-latest, windows-latest]`` → 2 OS × 2 Python
    = 4 cells.
  - ``fail-fast: false`` so one bad cell doesn't cancel the
    rest while the Windows leg is being brought up. Removable
    in a future PR once the suite is fully green.

* ``tests/test_backends.py``
  - Hard-coded ``"python3"`` → ``{sys.executable}`` in 7
    test commands. Windows has no ``python3`` on PATH; using
    ``sys.executable`` is portable and matches what CodeRabbit
    flagged on PR #236.
  - Strict string comparisons → ``shlex.split`` round-trip in
    5 resolver tests. ``shlex.quote`` adds single quotes
    around backslash paths on Windows, which broke the
    direct ``==`` compare.
  - Cross-platform suffix checks in 2 path-resolution tests
    (``Path(resolved).parts[-2:]`` instead of
    ``str(resolved).endswith("src/main.py")``).
  - ``mkdir -p`` → ``sys.executable -c "import os;
    os.makedirs(...)"`` in the cwd-sanitization test.
  - ``skipif(sys.platform == "win32")`` on 3 e2e tests that
    hit the underlying ``shlex.quote`` + ``cmd.exe`` quoting
    bug (real, separate issue).

* ``tests/test_sessions.py``
  - ``test_uses_data_dir``: check ``.evoscientist`` in the
    long path form (via ``Path.resolve()``) rather than the
    short-path form ``get_db_path`` returns on Windows.

* ``tests/test_mcp_client.py``
  - ``endswith("python")`` → ``Path(result).stem.lower()`` so
    ``python.EXE`` matches on Windows.
  - ``endswith("npx")`` also accepts ``npx.cmd`` so the npm
    shim on Windows matches.

## Out of scope (follow-up issues to file)

* ``os.killpg`` doesn't exist on Windows
  (``EvoScientist/background.py:248``) — 3 background tests
  fail. Real fix is the same ``psutil`` walk pattern PR #200
  shipped in ``langgraph_dev/manager.py``.
* Tilde expansion in file mentions.
* Windows-aware shell quoting in
  ``convert_virtual_paths_in_command``.
* Path conventions (``~/.config/evoscientist/`` vs
  ``%APPDATA%\EvoScientist``) — needs design discussion +
  ``platformdirs`` migration.
* Cross-module audit of
  ``EvoScientist/tools/execute.py``,
  ``EvoScientist/ccproxy_manager.py``,
  ``EvoScientist/config/onboard.py``.

Closes #207 (step 1 only — CI matrix + the easy test
fixes; remaining bullets tracked separately).

* fix: cross-platform compatibility for Windows CI runners

- background.py: replace POSIX-only os.killpg/os.getpgid with
  cross-platform _kill_process_tree() helper. On Windows falls back
  to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX
  keeps existing os.killpg logic.

- test_backends.py: replace mkdir -p shell execution in
  test_literal_workspace_path_replaced with preprocessing-boundary
  assertion (patch LocalShellBackend.execute, capture command,
  assert workspace path was rewritten to ./). Avoids POSIX-only
  mkdir -p on Windows runners.

- test_file_mentions.py: monkeypatch USERPROFILE on Windows so
  ntpath.expanduser() resolves ~ to tmp_path even when HOME is
  unset on CI runners.

* fix(test): cross-platform sleep/true commands for Windows CI

Replace POSIX-only sleep/true with module-level helpers that use
ping -n / cmd /c on Windows. Also fix python3 -> sys.executable
in the non-timeout recovery test.

- test_background.py: 7 sleep/true fixes
- test_background_middleware.py: 6 sleep/true fixes
- test_backends.py: 4 sleep fixes + 1 python3 fix

2318 passed, 0 failed on Windows.

* fix(test): use shell-portable double quotes for python -c on Windows

cmd.exe does not treat single quotes as string delimiters, so
-c 'raise SystemExit(1)' was passed with literal quotes on Windows.
Switch to double quotes which work on both cmd.exe and POSIX sh.

* fix: use psutil for Windows process tree kill + avoid sys.executable under uv

- background.py: replace Popen.terminate()/kill() with psutil-based
  process tree walking on Windows. TerminateProcess does NOT cascade
  to grandchildren; psutil.Process.children(recursive=True) ensures
  the entire tree is signaled.

- test_backends.py: replace sys.executable with 'python' in sandbox
  execute() calls. Under uv, sys.executable is under the workspace
  and gets rewritten to ./ by prepare_sandbox_command, breaking
  Linux CI. The plain 'python' command resolves correctly in any
  activated venv.

* fix: broaden try/except in _kill_process_tree to cover proc.children()

If the process exits between Process(popen.pid) and children(recursive=True),
the children call raises an uncaught exception escaping stop(). Move it inside
the existing try/except block.

* fix: narrow exception to ProcessLookupError in POSIX _kill_process_tree

OSError is too broad — would silently swallow EPERM on SIGKILL, leaving
the process alive when we report it as stopped. Match original behavior
which only caught ProcessLookupError (process already gone).

* style: ruff format test_backends.py

* ci: trigger re-run for flaky prompt_toolkit test

* style: fix ruff check (import order + RUF005 unpacking)

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-10 15:43:18 +01:00
X-iZhang cf5e0dd3bd feat(models): add support for 'claude-fable-5' mode 2026-06-10 00:51:29 +01:00
X-iZhang de3785346f feat(docs): update README to reflect Desktop WebUI changes and add demo video 2026-06-09 18:04:34 +01:00
houren Antony 2dc1e227eb fix(langgraph-dev): rotate langgraph_dev.log when it exceeds 50MB (#270)
* fix(langgraph-dev): rotate langgraph_dev.log when it exceeds 50MB

``_LOG_FILE`` (``~/.config/evoscientist/langgraph_dev.log``) was
opened in ``start_langgraph_dev`` with plain ``"ab"`` and never
rotated, so it grew unbounded over weeks/months of heavy use —
especially when chatty MCP servers spawned by langgraph dev
filled it, or when failure paths produced stack traces.

Implement the recommended option 1 from #209: filesize-based
rollover. When the active log exceeds 50MB on the next
``start_langgraph_dev`` invocation, rename it to
``langgraph_dev.log.1`` (overwriting any existing backup) via
``os.replace`` and start fresh. Single-backup policy keeps the
disk footprint bounded at roughly 2x threshold.

Rotation is best-effort: ``_rotate_log_if_needed`` logs and
swallows OSError so a permission error or racing rename can't
block langgraph dev from starting. The next ``start`` invocation
will try again — worst case the log grows for one more session.

Options 2 (timestamped per-session + 7-day sweep) and 3
(``RotatingFileHandler`` + pipe) are explicitly NOT done — option
1 is simplest, no async machinery, matches the issue's
recommendation.

Closes #209

* test(langgraph-dev): redirect _PID_DIR in rotate integration test

Address CodeRabbit review comment on #270: the
``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open``
test patched only ``_LOG_FILE`` to a tmp path, but
``start_langgraph_dev`` also calls ``_PID_DIR.mkdir(...)`` as part
of its prelude, which would create a real directory under
``~/.config/evoscientist/`` on a dev machine. Redirect
``_PID_DIR`` to ``tmp_path / "pids"`` too so the test stays
fully isolated. Add a final assertion that ``pid_dir.is_dir()``
holds, proving the function reached past the mkdir call.

* refactor(langgraph-dev): bundle runtime paths into LanggraphRuntimePaths

@din0s review follow-up on #270: the previous test isolation patched
only ``_LOG_FILE`` (and after a second round, ``_PID_DIR``), but
``start_langgraph_dev`` still touches 5 distinct on-disk paths. Patching
any subset of those still leaves the others pointing at the user's real
``~/.config/evoscientist/`` — exactly the case that produced the
"Port 6174 cannot be bound after waiting 60s" symptom on the
reviewer's machine.

Replace the five free-floating module-level constants
(``_PID_DIR`` / ``_PID_FILE`` / ``_LOG_FILE`` / ``_WORKSPACE_SIDECAR``
/ ``_FILE_LOCK_PATH``) with a single ``LanggraphRuntimePaths`` frozen
dataclass exposed as a module-level ``RUNTIME`` instance. Production
code accesses ``RUNTIME.pid_file`` etc.; tests can now substitute the
*whole* bundle in one assignment:

    monkeypatch.setattr(
        manager, "RUNTIME",
        manager.LanggraphRuntimePaths.for_directory(tmp_path / "runtime"),
    )

The classmethod ``for_directory(pid_dir)`` builds an isolated bundle
rooted at a single dir, so the test author doesn't spell out every
path field. Tests that only care about one field (e.g. pid_file
during the stale-process kill path) use
``dataclasses.replace(manager.RUNTIME, pid_file=X)`` — frozen
dataclass-friendly, no need to enumerate the other four fields.

The dataclass's docstring records the migration rationale (the old
five-name layout invited inconsistent patches).

External callers of the old constants updated:
- ``EvoScientist/deploy/server.py`` and ``webui.py`` now import
  ``RUNTIME`` and use ``RUNTIME.log_file`` for the on-screen log
  path hint. The other imports they had (``_DEFAULT_PORT``,
  ``_is_port_occupied``, ``_read_workspace_sidecar``) are still
  module-level functions/values, untouched.

Test updates:
- ``tests/test_langgraph_manager.py``: ``patch.object(manager, "_XXX",
  X)`` patterns now go through ``dataclasses.replace(manager.RUNTIME,
  xxx=X)``; the ``TestStartLanggraphDevRotatesLog::test_rotate_called_before_open``
  test (from the previous #270 review iteration) uses
  ``for_directory`` for one-shot isolation.
- ``tests/test_langgraph_dev_workspace_sidecar.py``: each test now
  goes through a tiny ``_isolated_runtime(monkeypatch, tmp_path)``
  helper that calls ``for_directory``.
- ``tests/test_langgraph_dev_deploy_mode.py``: same ``for_directory``
  swap.

No production behavior change. All ``langgraph_dev``-side tests
(``test_langgraph_manager.py`` 26/26, ``test_langgraph_dev_workspace_sidecar.py``
14/14, ``test_langgraph_dev_deploy_mode.py`` 14/14, ``test_cli_deploy.py``
18/18 — which indirectly exercises deploy/server.py and deploy/webui.py
imports) pass. Full-project test count unchanged from baseline; the
remaining 22 Windows-only pre-existing failures (test_background
``os.killpg``, test_file_mentions tilde, mcp_client ``shutil.which``,
test_sessions 8.3 short path) are documented as out-of-scope for #207.

* style: apply ruff format to langgraph_dev test + module files

CI lint check on #270 failed:

  Run ruff format --check .
  Would reformat: EvoScientist/langgraph_dev/manager.py
  Would reformat: tests/test_langgraph_manager.py

Plus two test files touched by the prior consolidation commit that
``ruff format`` hadn't seen yet:

  tests/test_langgraph_dev_deploy_mode.py
  tests/test_langgraph_dev_workspace_sidecar.py

Just formatting. No logic change. All 75 refactor-related tests pass.

* fix(test): use for_directory for full path isolation + patch _can_bind_port to skip real socket ops

Two fixes for TestStartLanggraphDevRotatesLog:

1. Replace dataclasses.replace(manager.RUNTIME, ...) with
   LanggraphRuntimePaths.for_directory(pid_dir) so pid_file,
   workspace_sidecar, and lock_file are also temp-rooted
   (prevents leak to ~/.config/evoscientist/).

2. Monkeypatch _can_bind_port to always return True so the
   bind-poll loop in _wait_for_port_bindable passes immediately
   without touching real sockets (fixes 60s timeout on machines
   where port 6174 is already in use).

* fix: cross-platform compatibility for Windows CI runners

- background.py: replace POSIX-only os.killpg/os.getpgid with
  cross-platform _kill_process_tree() helper. On Windows falls back
  to Popen.terminate()/Popen.kill() (TerminateProcess); on POSIX
  keeps existing os.killpg logic.

- test_backends.py: replace mkdir -p shell execution in
  test_literal_workspace_path_replaced with preprocessing-boundary
  assertion (patch LocalShellBackend.execute, capture command,
  assert workspace path was rewritten to ./). Avoids POSIX-only
  mkdir -p on Windows runners.

- test_file_mentions.py: monkeypatch USERPROFILE on Windows so
  ntpath.expanduser() resolves ~ to tmp_path even when HOME is
  unset on CI runners.

* refactor(test): add runtime_paths fixture to isolate manager.RUNTIME

Adds a reusable fixture that monkeypatches manager.RUNTIME to a
temp-rooted LanggraphRuntimePaths.for_directory(). Tests that need
specific fields can still dataclasses.replace(runtime_paths, ...)
but the baseline is always temp-isolated, preventing leaks to
~/.config/evoscientist/.

Updated test_langgraph_dev_deploy_mode.py, test_langgraph_dev_workspace_sidecar.py,
and test_langgraph_manager.py to use the fixture, consolidating sequential
lock_file + pid_dir patches into single dataclasses.replace calls.

* Revert "fix: cross-platform compatibility for Windows CI runners"

This reverts commit eb025d24af32e195a982cd40f6d70dba885c4019.

* style: ruff format conftest.py

* fix: address review issues in log-rotation + runtime paths

- Use for_directory(tmp_path/pids) as base in ensure_langgraph_dev tests
  so pid_file/log_file are co-located with pid_dir, not split across paths
- Remove unused runtime_paths param from test_no_existing_file_is_noop
- Replace manager.RUNTIME with runtime_paths in two sidecar tests
- Use for_directory(DEFAULT_PID_DIR) instead of explicit construction
- Fix stale _LOG_FILE reference in TestRotateLogIfNeeded docstring

* style: ruff format test files

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-09 15:55:44 +01:00
dinos cd2baa9588 feat(models): add opt-in prompt caching support for anthropic via openrouter (#272)
* feat(models): add opt-in prompt caching support for anthropic via openrouter

* chore: don't coerce model_kwargs to dict
2026-06-09 15:13:49 +01:00
Ziheng Zhang 4b6a969df2 refactor(agent): make create_cli_agent(config=, chat_model=) pure (#267)
* refactor(agent): make create_cli_agent(config=, chat_model=) pure

Re-applies the #183 purity refactor on top of the observation-memory
lifecycle that landed in #259, integrating the two cleanly.

create_cli_agent gains a pure path: when both `config` and `chat_model`
are passed it builds the agent entirely from locals and writes none of
the cached module globals (`_config`, `_chat_model`, `_chat_model_key`,
`_EvoScientist_agent`). `/model` commits the switch via
`set_active_config` / `set_chat_model_instance` only after a successful
build, so a failed rebuild leaves the session on the original model
(replaces the old snapshot/restore rollback).

Supporting changes:
- Extract `set_active_config` (write-half of `_ensure_config`),
  `_apply_env_from_config`, `_build_chat_model`, and
  `set_chat_model_instance`.
- Thread `cfg` / `chat_model` through `_get_default_middleware`,
  `_build_base_kwargs`, `load_mcp_and_build_kwargs`,
  `_maybe_swap_async_subagents`, and `_inject_subagent_middleware` so the
  pure path never falls back to the global-writing `_ensure_config()` /
  `_ensure_chat_model()`.
- Integrate with #259's memory middleware: subagent context-editing
  middleware binds the threaded `chat_model`, and the configured system
  prompt / memory controls read the threaded `cfg` (new threading vs the
  original #183, required because #259 made these paths read config).
- Consolidate `cfg` resolution to one `cfg if cfg is not None else
  _ensure_config()` at the top of each kwargs builder, matching the
  pattern already used in the other config-aware helpers.

* fix(agent): keep pure tool selector off global cache

* fix(model): apply config switch in place to preserve reference integrity

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-06-08 18:38:13 +01:00
dinos 8bb1d6c0e3 refactor(stream): langgraph streaming v3 (#268)
* refactor(stream): langgraph streaming v3

* fix: address CR comments

* chore(stream): add success field to state

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-08 17:14:25 +01:00
Wiktor Cupiał ac052bbb4c feat: free-scrolling (#262)
* feat: free-scrolling

* fix: anchor

* fix: textual private vars
2026-06-08 16:23:46 +02:00
Xi Zhang b4ffb35d71 feat(langgraph): add --no-reload option to start_langgraph_dev 2026-06-08 01:35:02 +01:00
Xi Zhang 63969b596d Release/v0.1.4 (#266)
* feat(middleware): reposition code interpreter middleware in the stack

* feat(models): add qwen3.7-plus model entry and update context window comment

* feat(models): add qwen3.7-max and qwen3.7-plus model entries for DashScope

* feat(auxiliary): implement auxiliary model support for background tasks and tool selection

- Added auxiliary model configuration to EvoScientistConfig.
- Introduced _ensure_auxiliary_chat_model function to manage auxiliary model instances.
- Updated onboarding steps to include auxiliary model selection.
- Modified middleware to route tool selection to the auxiliary model when applicable.
- Enhanced tests to cover auxiliary model functionality and configuration.

* feat(steps): update UI backend selection options and descriptions

* Refactor code structure for improved readability and maintainability

* feat(patches): implement OpenRouter response reasoning item stripping to prevent multi-turn errors

* feat: update version to v0.1.4 in badges, README, and pyproject.toml; adjust skill counts in steps.py

* feat(config): add auxiliary model and provider environment variables to test setup
2026-06-07 00:52:59 +01:00
Eliot Drizzle 3563c1d94f Update source for Scientific Skills in steps.py (#265) 2026-06-06 14:53:14 +01:00
dinos 92d95dee68 feat(memory): add observation memory lifecycle (#259)
* feat(memory): add observation memory lifecycle

Add file-backed observation memory with deterministic markdown records,
structured record_observation tooling, startup indexing, and
profile/observation prompt guidance.

Launch post-turn and post-subagent EvoMemory workers through LangGraph
dev so completed runs can update profile memory, save durable
observations, and write subagent execution summaries without blocking
the active agent.

Wire memory middleware into the main agent, subagents, async graphs, TUI
status reporting, worker activity accounting, and observation-aware
research prompts, with regression coverage for storage, lifecycle
scheduling, graph registration, status display, and stream reset
behavior.

* fix(cli): sync background agent server on resume

Resume flows now need to keep the LangGraph dev background server
aligned with the active workspace even when async subagents are
disabled. EvoMemory workers use that server too, so gating resume-time
sync on enable_async_subagents could leave workers pinned to the launch
workspace after resuming a thread from another workspace.

Run workspace sync unconditionally for Rich CLI and Textual resume
paths, while preserving WorkspaceMismatchError handling so failed sync
aborts the resume before mutating the active thread or workspace.

Propagate aborted resume callbacks through the command UI so
channel-issued /resume commands do not send false success or history
output. Channel slash dispatch now treats CommandManager-caught command
errors as command errors and skips completion hooks for those failed
commands.

Add regression coverage for disabled async subagents, callback aborts,
and channel command error reporting.

* fix(cli): prepare serve resume workspace before adopting

Load the resumed workspace agent and sync the background server as a
single pre-adoption step. Restore the previous active workspace if
preparation fails so serve mode keeps using the old session
consistently.

* fix(memory): untrack abandoned worker status watches

Stop treating watcher shutdown as confirmed worker completion. Terminal
worker statuses still count memory deltas, while poll failures or
watcher setup failures now remove the active run without crediting
partial outputs.

* fix(cli): report channel command failures accurately

Treat command_error as a None sentinel so empty error strings still
fail, and let TUI resumes continue only on non-mismatch
background-server sync failures while reporting degraded mode.

* fix(stream): clear memory counters for resume streams

Reset completed-memory counters for every new agent stream, including
Command-based HITL and resume streams, so saved-memory indicators do not
leak across turns.

* docs(tools): make observation recording guidance conditional

Clarify that agents should call record_observation only when the
observation tool is available, preserving the existing durability and
usefulness criteria.

* feat(config): add controls for profile and observation memory

Add config flags for profile memory, observation memory, observation
writer placement, and background memory workers.

Wire the controls through main agents, subagents, EvoMemory middleware,
and memory lifecycle workers so observation writes can be assigned to
the live agent, subagent worker, both, or neither. Keep turn memory
workers profile-only and make prompts reflect the available observation
read/write paths. Skip langgraph dev startup when neither async
subagents nor memory workers need the background server.

Add coverage for config parsing, prompt gating, middleware wiring, and
worker tool availability.

* test(cli): include memory defaults in serve config stubs

* fix(memory): offload async worker launch blocking calls

Run the langgraph-dev health check and memory-output snapshot in worker
threads from the async EvoMemory launcher so it does not block the event
loop.

* chore(memory): harden turn worker subagent guardrail

* chore(memory): refresh profile context per request

* fix(memory): offload async profile file reads

* fix(memory): offload async worker completion accounting
2026-06-05 15:11:20 +01:00
dependabot[bot] a27e5230c7 chore(deps): bump starlette in the uv group across 1 directory (#261)
Bumps the uv group with 1 update in the / directory: [starlette](https://github.com/Kludex/starlette).


Updates `starlette` from 1.0.0 to 1.0.1
- [Release notes](https://github.com/Kludex/starlette/releases)
- [Changelog](https://github.com/Kludex/starlette/blob/main/docs/release-notes.md)
- [Commits](https://github.com/Kludex/starlette/compare/1.0.0...1.0.1)

---
updated-dependencies:
- dependency-name: starlette
  dependency-version: 1.0.1
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-04 20:05:08 +01:00
dependabot[bot] 5c8975bc28 chore(deps): bump aiohttp in the uv group across 1 directory (#260)
---
updated-dependencies:
- dependency-name: aiohttp
  dependency-version: 3.14.0
  dependency-type: direct:production
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-06-04 11:06:05 +02:00
Zhen-Yi Zhou cedb3aa744 Fix ssh remote path handling in sandbox (#242)
* Fix ssh remote path handling in sandbox

* Address ssh remote command review feedback

* Format backend files with ruff

* Narrow SSH remote command handling

* Narrow SSH preprocessing to single-quoted remote args

* Address remaining SSH preprocessing review feedback

* Tighten SSH executable recognition

* Recognize only literal ssh wrapper

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-06-03 17:55:27 +01:00
X-iZhang 044a85ccd2 Update ResearchClawBench ranking details in README files 2026-06-03 17:35:10 +01:00
Wanghan Xu 0198e50e7f Add ResearchClawBench ranking news (#257) 2026-06-03 14:59:47 +01:00
X-iZhang faea53be52 chore: bump version to v0.1.3 2026-06-03 01:37:16 +01:00
Xi Zhang 9cffe9d457 Enhance multimodal handling in LLM model (#256)
* Enhance multimodal handling in LLM model

- Updated `_flatten_message_content` to preserve media blocks (images, files) while flattening text content.
- Introduced `_sanitize_messages` to manage media hoisting for tool messages, ensuring compatibility with OpenAI APIs.
- Modified `_patch_openai_compat_content` to accommodate new media handling logic, including retry mechanisms for media errors.
- Added comprehensive tests for media preservation, including various scenarios with images, files, and unsupported media types.

* fix: preserve order of text and media blocks in message flattening

* test: add tests for _strip_media_types to ensure position preservation and deduplication
2026-06-03 01:06:46 +01:00
dinos d348076f40 Add runtime context middleware (#255) 2026-06-02 18:47:26 +01:00
dinos 9285c6dad8 Migrate memory middleware to profile files (#253)
* feat(memory): migrate to profile memory files

* chore(stream): read profile headings from templates

* fix(display): keep assistant responses if response_text has started

* fix(memory): do not treat failed bootstraps as profile creation

* chore(memory): unlink blank legacy memory

* fix(memory): resolve project_id once

* fix(memory): preserve unreadable profile files

* chore(tui): render streamed narration inline with tool timeline

Update the TUI streaming timeline so assistant text emitted before or
between tool calls is rendered inline where it occurs, rather than being
kept as a single answer bubble above or below the tools.

If the model begins an assistant response and then emits another tool
call, the provisional response is converted into inline narration before
that tool. The final assistant message then renders only the remaining
response suffix, avoiding duplicate text in the completed transcript.

Stop/cancel handling now preserves any active inline narration, appends
the visible stopped marker only to the remaining displayed segment, and
still returns the full normalized stopped response for channel callers.

Completed tools continue to collapse while long runs are active, but
expand again when the turn reaches a final state so the completed
transcript shows the full tool timeline.

* fix(stream): preserve narration around tool timelines

Keep assistant narration attached to the tool call that follows it
instead of folding all streamed text into the final answer block.

Track narrated response segments in stream state, render them before
their corresponding regular or task tool entries, and keep final answers
limited to the response suffix that has not already been shown inline.
Preserve narration across normal completion, stop/error final frames,
sub-agent task calls, and collapsed live tool summaries.

Add regression coverage for pending tools, completed tools, sub-agent
task delegations, collapsed completed/running tool summaries, and final
stop frames.

* fix(tui): finalize inline narration transitions

* test(memory): use canonical project id helper
2026-06-02 18:21:14 +01:00
X-iZhang b56234317a fix: restrict textual version to avoid CJK input issues on iTerm2 2026-06-02 11:58:12 +01:00
X-iZhang 565d9647ac chore: update version to v0.1.2 in project files and badges 2026-06-02 00:32:09 +01:00
X-iZhang d53bfa35c5 feat: update MiniMax model entries and context window for M3 variant 2026-06-02 00:00:52 +01:00
Xi Zhang fbd1d709ca feat: add WebUI mode support with related configuration and onboarding (#252)
* feat: add WebUI mode support with related configuration and onboarding steps

* feat: enhance WebUI port configuration to prevent conflicts with backend port

* feat: add support for fresh interactive session detection in WebUI
2026-06-01 12:01:08 +01:00
Xi Zhang 3ce6523faf fix: update deepagents and langchain versions (#251)
* fix: update deepagents and langchain versions; enhance _reduce_messages_delta handling for None state

* fix: update langchain version constraint to >=1.3 in pyproject.toml and uv.lock
2026-05-31 22:12:13 +01:00
Xi Zhang a13904185d Feat/sandbox execute timeout (#243)
* feat: implement configurable sandbox execute timeout and enhance recovery instructions

* feat: add background process management tools and middleware for sandbox execution

* feat: enhance background process management with completion notifications and deduplication

* feat: enhance sandbox execution timeout validation and update related messages

* feat: enhance background process management with thread-specific completion notifications and HITL approval handling

* test: assert completion notification waits for process finish timestamp
2026-05-31 15:11:25 +01:00
Ziheng Zhang 2364e6b130 fix(cli): forward async-notifier replies back to originating channel (#244)
* fix(cli): forward async-notifier replies back to originating channel

When PR #214's auto-notifier fires a synthetic agent turn after a
channel-originated conversation, the synthesized response only rendered
to the local CLI/TUI — the channel user (iMessage etc.) saw nothing
and had to manually re-prompt to find out what happened.

Adds a per-thread channel-origin registry in cli/channel.py and wires
the three notifier paths (Rich CLI / TUI / serve) to publish the final
response back via bus.publish_outbound when the originating thread was
started by a channel turn. Publish is fire-and-forget (scheduled on the
bus loop + done-callback for failure logging) so the notifier turn
doesn't block on the asyncio / textual event loop.

The registry is cleared on /new and /resume rotation so stale entries
don't accumulate.

* fix(cli): address review feedback on channel-origin forwarding

Follow-up to the review on #244 (din0s, X-iZhang):

- Guard the /resume origin cleanup on a real thread change in Rich CLI
  and TUI (serve mode already did via thread_changed). Resuming the
  already-active thread no longer wipes its still-live origin, which
  would otherwise silently drop a later async-notifier forward — the
  exact gap this PR closes.
- Re-bind the now-current thread to its channel after a channel-issued
  /new or /resume slash command (which rotates the thread inside the
  dispatch), so notifier turns on the rotated thread still forward.
- Guard the publish done-callback against a cancelled future, whose
  .exception() raises CancelledError (rather than returning it) on
  bus-loop teardown, so the intended warning still logs.
- Mirror the normal reply path's manager.record_message(channel, "sent")
  for forwarded notifications so per-channel stats stay accurate.
- Print the closing "[channel: Replied to ...]" line in all three
  notifier paths (Rich CLI / TUI / serve) when a forward actually
  happened, so the forwarded block reads as terminated on screen.

Adds test_publish_records_sent_metric. ruff clean; notification-origin
suite (10) + related channel/CLI/serve suites (728) pass.

* fix(cli): store sender information separately from chat_id in channel origin

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-05-31 14:43:42 +01:00
X-iZhang 721a03c25b feat: update model version from claude-sonnet-4-5 to claude-sonnet-4-6 and related adjustments 2026-05-29 00:40:46 +01:00
Xi Zhang f75bfcda51 Add onboarding wizard with style and validation components (#241)
* Add onboarding wizard with style and validation components

- Introduced `style.py` for shared visual elements used in the onboarding wizard.
- Created `validators.py` for input validation, including integer and choice validators, and API key validation functions for various providers.
- Implemented `wizard.py` as the entry point for the onboarding process, managing user prompts and configuration steps.
- Added progress rendering and autosave functionality to enhance user experience during the onboarding process.

* feat(onboarding): enhance validation and configuration for onboarding wizard

- Added validation for UI backends, workspace modes, and providers in the onboarding command.
- Updated channel definitions to include secret field handling for sensitive tokens.
- Improved user prompts for required fields, ensuring sensitive data is masked.
- Introduced constants for valid providers, UI backends, and workspace modes to maintain consistency.
- Implemented tests to ensure alignment between constants and interactive choices in onboarding steps.

* feat(onboarding): improve WeChat account ID prompt and validation for newly enabled channels

* feat(onboarding): enhance WeChat backend credential prompts and validation

* feat(onboarding): refine WeChat backend credential prompts for wecom and wechatmp

* Refactor onboarding package for improved structure and clarity

- Simplified the onboarding package by removing unnecessary re-exports and consolidating public API to only include `run_onboard`.
- Updated `install_back_keys` to `install_navigation_keys` for clarity and consistency in the prompter module.
- Enhanced the `NonInteractivePrompter` class to support strict mode, allowing for better handling of non-interactive prompts.
- Adjusted the onboarding steps to utilize the new navigation keys installation method.
- Improved the `run_onboard` function to handle section implications based on user flags, enhancing the onboarding experience.
- Updated tests to reflect changes in imports and ensure compatibility with the new structure.

* feat(onboarding): enhance validation logic for non-interactive prompts

* refactor(onboarding): streamline onboarding module structure and enhance validation error handling

* refactor(onboarding): enhance config revert logic to preserve original file state

* refactor(onboarding): enhance tavily key validation and error handling in onboarding process
2026-05-28 12:42:49 +01:00
Xi Zhang b9ad694467 fix: resolve path correctly when workspace name appears in parent path 2026-05-23 12:37:21 +01:00
Ziheng Zhang d2283397a4 feat(feishu): scan-to-create QR onboarding + silence unsubscribed WS events (#239)
* feat(feishu): scan-to-create QR onboarding flow

Add a device-code flow against accounts.feishu.cn/oauth/v1/app/registration
that lets users scan a terminal QR code with Feishu / Lark mobile to
auto-create a PersonalAgent bot app with the required IM permissions
pre-attached. The poll endpoint returns app_id + app_secret, which the
onboarding wizard then writes into the channel config — no manual app
creation on open.feishu.cn required.

- channels/feishu/onboard.py: qr_register() public entry, init/begin/poll
  helpers, QR rendering via the soft qrcode dep, automatic feishu↔lark
  domain switch based on the scanning user's tenant_brand, and a
  best-effort bot probe to surface the bot name in the wizard
- channels/feishu/__init__.py: re-export qr_register (mirrors qq)
- config/onboard.py: offer "Scan QR code (recommended) / Enter manually"
  in the Feishu branch, ask for region (feishu vs lark), then call
  qr_register and populate feishu_app_id / feishu_app_secret /
  feishu_domain; add qrcode>=7.4 to the feishu pip extras

* fix(feishu): silently absorb unsubscribed WebSocket events

Feishu auto-subscribes PersonalAgent apps to many event types
(im.message.reaction.created_v1, message.read_v1, message.recalled_v1,
chat.member.*, ...) that EvoScientist doesn't register handlers for.
Without intervention, lark-oapi's dispatcher raises EventException
("processor not found, type: ..."), the WS client logs it at ERROR and
replies HTTP 500 on the frame, and Feishu marks the event as failed
and retries it.

The problem is amplified by _send_ack_reaction: every inbound message
triggers our own reaction, which Feishu echoes back as
reaction.created_v1, creating a continuous ERROR-log feedback loop and
pointless retries.

Wrap EventDispatcherHandler._do_without_validation after build() to
swallow "processor not found" EventExceptions (debug log + return None)
while letting all other errors propagate. Failure-safe: if lark-oapi's
internal API changes the wrapper degrades to the prior behavior rather
than breaking the channel.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-05-20 22:53:16 +08:00
dependabot[bot] b11932f91b chore(deps): bump idna in the uv group across 1 directory (#238)
Bumps the uv group with 1 update in the / directory: [idna](https://github.com/kjd/idna).


Updates `idna` from 3.13 to 3.15
- [Release notes](https://github.com/kjd/idna/releases)
- [Changelog](https://github.com/kjd/idna/blob/master/HISTORY.md)
- [Commits](https://github.com/kjd/idna/compare/v3.13...v3.15)

---
updated-dependencies:
- dependency-name: idna
  dependency-version: '3.15'
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-05-20 15:39:12 +01:00
Xi Zhang 7959495a13 feat(deploy): add EvoSci deploy subcommand (#228)
* feat(deploy): implement standalone LangGraph server and CLI command for deployment

* feat(deploy): enhance port validation and environment variable management for deployment

* Refactor langgraph dev deployment and introduce workspace sidecar protocol

- Updated the deployment mode handling in `server.py` to use a single environment variable `EVOSCIENTIST_DEPLOY_MODE` with values `full` and `stripped`.
- Enhanced the `manager.py` to implement a workspace fingerprint sidecar, allowing cross-process reuse of langgraph dev instances while ensuring workspace consistency.
- Introduced functions to write and read the workspace sidecar, with error handling for missing or corrupt data.
- Added tests for the workspace sidecar functionality, including validation of the JSON schema and ensuring proper error handling for workspace mismatches.
- Updated existing tests to reflect changes in deployment mode handling and added new tests for signal handling during shutdown.
- Ensured that cleanup routines remove the workspace sidecar alongside the PID file during shutdown.

* fix(langgraph): improve workspace sidecar checks for process ownership and stale handles
2026-05-20 11:06:35 +01:00
X-iZhang 9c7347eedb chore: update version to v0.1.1 in badges and pyproject.toml 2026-05-19 14:06:22 +01:00
Xi Zhang 331056cdc8 feat(middleware): upgrade deepagents 0.5.7 → 0.6.2 (#231)
* feat(middleware): add CodeInterpreterMiddleware with project-specific configuration

chore(config): increase checkpoint retention limit for runaway conversations

fix(tests): update database schema references from 'blob' to 'value'

chore(deps): update deepagents dependency to include quickjs support

* feat(deepagents): update to version 0.6.1 and add optional dependencies for quickjs

* feat(sessions): improve error handling for message deltas and update Overwrite type check

* Enhance PruningCheckpointer with DeltaChannel Awareness

- Introduced a new pruning strategy in `_prune_after_put` to preserve the `_DeltaSnapshot` chain during checkpoint pruning.
- Implemented methods to fetch recent checkpoint IDs and walk to snapshot ancestors, ensuring that necessary checkpoints are retained.
- Updated SQL queries to handle checkpoint and write deletions more efficiently.
- Added comprehensive tests for DeltaChannel-aware pruning, ensuring that the pruning logic correctly handles various checkpoint scenarios, including those with and without snapshot seeds.
- Refactored `_load_checkpoint_messages` to utilize the new saver interface, improving message reconstruction from checkpoints.

* feat(tests): add migration sweep test to preserve snapshot ancestor

* feat(sessions): enhance checkpoint retrieval to prevent transcript leakage in multi-agent scenarios

* feat(middleware): enhance CodeInterpreterMiddleware with configurable timeout and result character limit

feat(config): add CodeInterpreterMiddleware tuning parameters to EvoScientistConfig

feat(sessions): implement inline message delta reducer for improved message handling

* feat(dependencies): update deepagents version to 0.6.2 in pyproject.toml and uv.lock
2026-05-19 12:35:36 +01:00
Xi Zhang 385f9756c1 feat(backends): implement tier-aware virtual mount resolution for ski… (#236)
* feat(backends): implement tier-aware virtual mount resolution for skills and memories

* test: add end-to-end test for workspace tier shadowing global tier in CustomSandboxBackend

* feat(backends): enhance virtual mount resolution for skills and memories with tier paths and quoting

* fix(tests): update Python command in virtual mount resolution tests to use python3
2026-05-19 11:50:49 +01:00
Ziheng Zhang 7f1aa3b0f6 fix(cli): handle spaces in @file mentions (#234)
* fix(cli): handle spaces in @file mentions

The @file parser truncated at the first space, so dragging or pasting a
filename like `@PREPING_ Building Agent.pdf` only matched `@PREPING_`
and warned "file not found". Now supports `@"..."` / `@'...'` quoted
form for explicit paths, plus a greedy expansion fallback that walks
across whitespace until an existing file resolves (bounded by newlines,
the next `@`, and a 20-token cap). Autocomplete also returns quoted
mentions for any candidate containing a space.

* style: apply ruff format to file_mentions
2026-05-18 12:05:21 +01:00
X-iZhang 0d51406149 chore: update wechat_group image asset 2026-05-16 16:07:32 +01:00
Wiktor Cupiał a4c9c779c9 feat: status and elapsed time indicator (#218)
* feat: status and elapsed time indicator

* test: add tests for tui-status

* fix: move to enum+switch, change phase calculation

* feat: remove 'done' phase
2026-05-13 14:15:39 +01:00
dinos 4b0c91190a feat(llm): add dashscope-code provider for Alibaba Coding Plan keys (#225)
* feat(llm): add dashscope-code provider for Alibaba Coding Plan keys

Alibaba Cloud Bailian "Coding Plan" subscription keys (sk-sp-*) route
through a separate endpoint (coding.dashscope.aliyuncs.com/v1) that the
standard `dashscope` provider can't reach. Add a sibling provider entry
matching the zhipu/zhipu-code and moonshot/kimi-coding precedents, with
its own validator (the coding endpoint returns 404 on /models, so probe
via chat.completions instead).

Closes #224

* fix(llm): keep dashscope as default provider for qwen3-coder shortcut

The MODELS dict is built from _MODEL_ENTRIES via a last-write-wins dict
comprehension. The initial commit listed dashscope-code AFTER dashscope,
which silently flipped the bare `get_chat_model("qwen3-coder")` shortcut
to the coding endpoint — breaking standard sk-* keys.

Reorder to match the zhipu-code / zhipu precedent: coding endpoint first,
general endpoint last so the general endpoint wins the collision and
remains the default for the shared "qwen3-coder" short name.
2026-05-13 10:23:32 +01:00
dependabot[bot] 35ea2bfb52 chore(deps): bump urllib3 in the uv group across 1 directory (#222)
Bumps the uv group with 1 update in the / directory: [urllib3](https://github.com/urllib3/urllib3).


Updates `urllib3` from 2.6.3 to 2.7.0
- [Release notes](https://github.com/urllib3/urllib3/releases)
- [Changelog](https://github.com/urllib3/urllib3/blob/main/CHANGES.rst)
- [Commits](https://github.com/urllib3/urllib3/compare/2.6.3...2.7.0)

---
updated-dependencies:
- dependency-name: urllib3
  dependency-version: 2.7.0
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-12 08:53:08 +01:00
Ziheng Zhang 8fe774b056 Feat/qq interactive buttons (#220)
* feat(qq): add inline keyboard buttons for C2C HITL approval

QQ Bot supports inline buttons via `markdown + keyboard` payloads. Clicks
arrive as `interaction_create` events through the existing botpy
WebSocket gateway — no extra subscription needed beyond enabling the
`interaction` intent. Group-scope clicks are out of scope here (DM only).

Send path
- `_build_qq_keyboard(buttons)` mirrors the Feishu helper, mapping the
  generic `{text, value, type}` shape to QQ's `{render_data, action}`
  with action.type=1 (callback). One button per row for mobile clarity.
- `_send_chunk` extracts `metadata["buttons"]` and threads a `keyboard`
  payload into `_post_markdown_message` for C2C only.
- Markdown→plain fallback can't carry a keyboard, so when buttons were
  attached the fallback content gets a textual `Reply: 1=Approve, …`
  hint built from the button list. `_parse_approval_reply` accepts
  the same values typed manually, so the user is never stuck.

Receive path
- `on_interaction_create` is registered on the bot class.
- `_on_interaction` extracts `data.resolved.button_data`, builds an
  InboundMessage, runs it through inbound middleware (Dedup suppresses
  retry callbacks), and publishes directly to the bus — bypassing the
  per-sender debounce buffer so the click value isn't merged with any
  text typed in the same window.
- Always ACKs via `api.on_interaction_result(id, 0)` in `finally` so
  QQ doesn't show the button as "expired", even if middleware drops
  the click or something throws downstream.

`QQ.inline_buttons=True`; `_approval_prompt_metadata` now auto-attaches
the Approve/Reject/Approve-all button row for QQ HITL prompts.

* fix(qq): button-value coercion, ACK timing, HITL consumer wiring

Fixes 6 bugs found in the inline-keyboard commit and consolidates the
button helpers so the keyboard builder, plain-text fallback hint, and
interaction handler share one coercion path.

- Plain-text fallback no longer crashes on non-string `value` (e.g.
  `{"text": "OK", "value": 42}`).  Extracted `_normalize_button` is now
  the single place that resolves `(label, value)` and coerces non-strings.
- `metadata["button_value"]` is the coerced string instead of the raw
  payload, matching `content` and downstream string comparisons.
- `_on_interaction` ACKs first, before publishing to the bus, so the
  QQ button UI never shows "expired" if middleware is slow.
- Wire `_approval_prompt_metadata` + `_format_approval_prompt(with_buttons=)`
  into `InboundConsumer._stream_with_hitl` and `cli.channel.channel_hitl_prompt`
  so the QQ `inline_buttons=True` capability is actually used end-to-end
  (HITL prompts auto-attach Approve/Reject/Approve-all buttons when the
  channel advertises the capability).
- Trim contradictory `_QQ_DEFAULT_PERMISSION` comment.
- Fix `test_group_interaction_ignored` docstring (ACK runs first now,
  not in `finally` after a `return`).

Tests: `_normalize_button` covered indirectly via existing keyboard tests;
new regressions for non-string fallback hint, ACK-on-handler-throw, and
string-coerced `button_value` metadata.

* refactor(qq): slim button helpers and explicit has_buttons flag

Inline single-use _button_hint and the _QQ_BUTTON_STYLE/_QQ_DEFAULT_PERMISSION
constants in qq/channel.py; tighten _on_interaction (drop unreachable
"[button click]" sentinel and unused triggering_message_id metadata; collapse
"if resolved else" ternaries via `or ""`).

Replace the metadata round-trip ("buttons" in metadata) used to detect button
support in consumer.py and cli/channel.py with an explicit has_buttons bool
threaded through both the prompt formatter and metadata builder.

Apply ruff format to the previously unformatted blocks introduced earlier on
this branch so CI lint passes.

* feat(qq): send post-decision confirmation after HITL approval

Send a visible confirmation message ("✅ 已批准" / "❌ 已拒绝") right after
the user resolves a HITL approval — QQ Bot has no message-recall or edit API
for C2C, so a follow-up message is the only way to give the click/reply
strong feedback.

Bus consumer (consumer.py): only sends the confirmation when the user
actually responded (event was set), to avoid pretending the user approved
when the request really timed out and auto-approved.

CLI HITL prompt (cli/channel.py): mirrors the same set of confirmation
strings.  Timeout and unrecognized-reply paths keep their existing English
text since their semantics differ (auto-reject vs auto-approve, plus a
hint about the unparsed input).
2026-05-11 10:44:20 +01:00
dependabot[bot] ccb3083183 chore(deps): bump langchain-core in the uv group across 1 directory (#219)
Bumps the uv group with 1 update in the / directory: [langchain-core](https://github.com/langchain-ai/langchain).


Updates `langchain-core` from 1.3.2 to 1.3.3
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain-core==1.3.2...langchain-core==1.3.3)

---
updated-dependencies:
- dependency-name: langchain-core
  dependency-version: 1.3.3
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-09 16:50:25 +01:00
X-iZhang f00ee1b5cc chore: update version to v0.1.0 2026-05-08 22:57:30 +01:00
Xi Zhang c407d2e20f Fix/async subagent model switch (#217)
* feat(middleware): add ConfigurableModelMiddleware for dynamic model resolution

- Introduced ConfigurableModelMiddleware to resolve chat models from RunnableConfig.configurable on each call.
- Updated middleware initialization to include ConfigurableModelMiddleware.
- Enhanced context editing middleware tests to verify presence of ConfigurableModelMiddleware.
- Implemented tests for ConfigurableModelMiddleware to ensure correct model overriding and caching behavior.
- Added tests for deepagents model-passthrough patch to verify configuration injection in async tasks.

* feat(async-subagent): update middleware handling to prevent deadlocks in async sub-agents

* style: Refactor code formatting for improved readability in patches and test files

* refactor: streamline middleware construction and improve async handling in ConfigurableModelMiddleware

* fix: remove unused request parameter from _read_model_override function

* refactor: improve async handling in _ClientProxy and enhance logging in ConfigurableModelMiddleware
test: add behavior test to ensure AskUserMiddleware is excluded in async subagent mode
2026-05-08 22:52:24 +01:00
Ziheng Zhang 89b0ecdbf3 feat(qq): add QR-code scan-to-configure onboarding for QQ Bot (#213)
* feat(qq): add QR-code scan-to-configure onboarding for QQ Bot

Adds a `qr_register()` flow that drives q.qq.com's create_bind_task /
poll_bind_result APIs so the wizard can auto-fill `qq_app_id` and
`qq_app_secret` after the developer scans a QR code with a bound QQ
account, falling back to manual entry on failure or cancel.

- channels/qq/crypto.py: AES-256-GCM helpers for decrypting the bot's
  client_secret returned by poll_bind_result.
- channels/qq/onboard.py: portal API client + polling loop.
- channels/qq/__init__.py: re-export `qr_register`.
- config/onboard.py: QQ branch in `_step_channels` that offers
  "Scan QR code" vs "Enter manually", and skips the manual prompt
  loop when a scan succeeded.

* style(qq): fix ruff lint errors in onboard.py

Move `import os` to the top-level import block (E402), drop the legacy
`typing.Optional`/`typing.Tuple` imports (UP035), and use the PEP 585/604
builtin generics (`tuple[...]`, `X | None`) for the few annotations that
still referenced them (UP006/UP045). No behavior change.

* fix(qq): harden QR onboard error paths and declare scan deps

Address review feedback on PR #213:
- Declare cryptography>=41.0 and qrcode>=7.4 in [qq]/[all-channels]
  extras and in _CHANNEL_PIP_DEPS so the scan flow no longer fails
  with an opaque ImportError on a fresh `evoscientist[qq]` install.
- Polling loop logs each _poll_bind_result failure and aborts after
  5 consecutive errors instead of silently spinning until the 600s
  timeout, restoring the documented Raises: RuntimeError contract.
- Wrap decrypt_secret in try/except so failures honor the
  None-on-failure contract instead of letting exceptions escape.
- Preflight `import cryptography` in the scan branch and offer
  install or fall back to manual entry.

* style: ruff format collapse two over-wrapped log/console lines
2026-05-08 12:08:52 +01:00
Xi Zhang 4e04ac5b72 fix: Improve watcher logic to prevent false-positive notifications on… (#216)
* fix: Improve watcher logic to prevent false-positive notifications on clean stream exits

* fix: Update watcher logic to drop notifications on persistent runs.get failures

* fix: Refactor test for watcher persistent failure notification handling

* fix: Enhance watcher test to validate all notification queues are empty after reconnect budget exhaustion
2026-05-07 22:16:36 +01:00
Xi Zhang 80f1f4fa0f feat: Implement async sub-agent auto-notification system (#214)
* feat: Implement async sub-agent auto-notification system

- Added async notifier functionality to handle notifications for sub-agents reaching terminal states.
- Introduced `AsyncTaskNotification` dataclass for structured notification data.
- Implemented `watch_run_and_notify` to monitor agent runs and enqueue notifications.
- Created `spawn_watcher` to manage watcher tasks and ensure proper cancellation of previous watchers.
- Developed `consume_notifications` to process notifications, deduplicate them, and format messages for LLM.
- Added tests for notification handling, including draining, deduplication, and formatting.
- Patched deepagents to integrate the new watcher functionality into start and update tools.

* Enhance async notifier with per-thread notification routing and error handling

- Introduced `origin_cli_thread_id` to `AsyncTaskNotification` for routing notifications back to the originating CLI session.
- Implemented per-thread notification queues to handle notifications based on the originating thread.
- Updated `has_pending_notifications` and `drain_notifications` to respect thread-specific queues.
- Enhanced `watch_run_and_notify` to detect in-band error events from the SSE stream and handle clean exits.
- Modified tests to verify the new notification routing behavior and ensure proper handling of notifications across threads.
- Added a fixture to restore the async watcher patch state in tests to prevent state leakage.
- Updated deepagents patching to capture the main agent's CLI thread ID for notification routing.

* feat: Enhance async notifier with thread-specific watcher management and notification filtering

* test: Enhance notification draining logic for cleaner test setup

* refactor: Remove summary field from AsyncTaskNotification and update related tests

* feat: Enhance async notification handling with target thread ID support

* Refactor async notifier and middleware for improved task management

- Removed the no-op shutdown watcher loop from async_notifier.py as it is no longer needed.
- Updated watch_run_and_notify to clarify notification handling and race conditions.
- Cleaned up shutdown handling in commands.py, interactive.py, and tui_interactive.py by removing obsolete shutdown watcher calls.
- Deleted the deepagents async watcher patch from patches.py, transitioning to a new middleware approach.
- Introduced AsyncWatcherMiddleware to handle async task notifications directly during tool calls.
- Updated tests to validate the new middleware functionality and ensure proper watcher spawning and cancellation.
- Enhanced test coverage for async watcher middleware, including edge cases and error handling.

* feat(tests): add fixture to reset notifier state before each test
2026-05-07 16:03:02 +01:00
Ziheng Zhang 692dc491ac # feat(wechat): add personal-WeChat (iLink) backend with QR login (#212)
* feat(wechat): add personal-WeChat (iLink) backend with QR-code login

Adds a third WeChat backend alongside WeCom and Official Account:
``personal`` rides Tencent's iLink Bot long-poll gateway so a personal
WeChat account can act as a bot.  Credentials are obtained via QR-code
scan and persisted under ``DATA_DIR/wechat_personal/accounts/``.

- channels/wechat/personal.py: WeixinPersonalChannel + qr_login.
- channels/wechat/crypto.py: aes128_ecb_decrypt + parse_ilink_aes_key
  for the iLink CDN media protocol.
- channels/wechat/probe.py: validate_wechat_personal credential probe.
- channels/wechat/serve.py: --backend personal CLI + --qr-login flow.
- channels/wechat/__init__.py: factory dispatch on wechat_backend; pull
  in the new dependencies in the docstring.
- config/settings.py: wechat_personal_* fields.
- config/onboard.py: WeChat-backend picker + QR-scan flow in the wizard
  + personal-backend probe in _probe_channel.
- pyproject.toml / uv.lock: add qrcode + certifi to wechat & all-channels
  extras (aiohttp was already pulled in transitively).

* fix(wechat): address ruff failures and CodeRabbit review on personal-WeChat PR

- personal.py: drop unused imports (`field`, `PollingMixin`); replace
  `asyncio.TimeoutError` with builtin; hold references to background
  `asyncio.create_task` results so they aren't GC'd; wire `dm_policy`
  through `_process_message` (disabled/allowlist) so `wechat_personal_dm_policy`
  actually takes effect for DMs.
- onboard.py: import-check gate now validates the full WeChat dependency
  set (aiohttp, qrcode, Crypto, certifi) instead of only aiohttp; mask
  `WeCom Secret` and `MP App Secret` prompts via `questionary.password`;
  derive the QR-login hint path from `_account_dir()` instead of the
  hard-coded `~/.evoscientist/...`; stop copying the QR-login token into
  the main config (already persisted per-account on disk — copying broadens
  secret exposure and risks staleness).
- pyproject.toml: allow Chinese full-width punctuation in `allowed-confusables`
  for user-facing CN messages.

* style(wechat): apply ruff format

`ruff format --check` was failing CI on three files (one pre-existing in
`__init__.py` plus formatter-driven line-merges in the files touched by
the previous fix commit). Ran `ruff format` to bring them in line; both
`ruff check` and `ruff format --check` now pass.
2026-05-07 16:33:30 +02:00
Wiktor Cupiał 9e51ec6fdd feat(cmd): add /model-fallback command (#196)
* feat(cmd): add /model-fallback command

* fix: apply feedback

* fix: lock usage with _fallback_chain

* fix: apply feedback

* fix: apply feedback

* feat: add tests

* fix: tests

* Update EvoScientist/middleware/model_fallback.py

Co-authored-by: dinos <dinospk1999@gmail.com>

---------

Co-authored-by: dinos <dinospk1999@gmail.com>
2026-05-07 15:35:26 +02:00
Ziheng Zhang 22a65b640d refactor(channels): remove dead MessageBus dispatcher (#205)
* refactor(channels): remove dead MessageBus dispatcher

Outbound routing has two implementations: ``MessageBus.dispatch_outbound``
(subscriber-based) and ``ChannelManager._dispatch_outbound`` (registry
lookup).  Only the latter is ever started in production — the former
is reachable solely from tests, yet both consume from the same
``bus.outbound`` queue.  If anyone followed the bus's own API surface
they would silently steal messages from the real dispatcher.

Drop the unused machinery to leave a single, obvious outbound path:
- ``MessageBus.subscribe_outbound`` / ``dispatch_outbound`` / ``stop``
- ``_running`` flag and ``_outbound_subscribers`` map
- ``OutboundCallback`` type alias
- The lone ``bus.stop()`` call in ``cli/channel.py`` (was no-op)
- Four tests covering the removed code paths

* test(channels): drop empty MessageBus stubs after dispatcher removal

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-05-07 12:11:06 +01:00
Xi Zhang d7c0eec0e9 chore: update dependencies and remove unused OpenRouter patch (#211) 2026-05-05 18:51:00 +01:00
Xi Zhang f41584e10b Refactor sub-agent architecture and introduce async support (#200)
* Refactor sub-agent architecture and introduce async support

- Removed the legacy subagent.yaml file and replaced it with individual YAML files for each sub-agent in the subagents directory.
- Updated the load_subagents function to support both directory and single file layouts for loading sub-agent configurations.
- Added new langgraph_dev module for managing async sub-agent lifecycle and deployment.
- Created graphs for async sub-agents (writing-agent, data-analysis-agent) and updated langgraph.json for deployment.
- Introduced new sub-agent definitions for planner, research, debug, code, and writing agents with appropriate system prompts and configurations.
- Enhanced package data inclusion in pyproject.toml to accommodate new sub-agent YAML files.

* Refactor code for improved readability by consolidating conditional statements and formatting

* feat: enhance async sub-agent support with workspace synchronization and user feedback

- Added console status messages during async sub-agent server startup and workspace synchronization to improve user experience.
- Implemented a new WorkspaceSyncWidget for live feedback during workspace sync operations.
- Updated onboarding to reject occupied ports and ensure proper workspace handling for async sub-agents.
- Introduced locking mechanisms to manage concurrent access to langgraph dev processes and workspace states.

* feat: add async sub-agent configuration and server management functions

* feat: improve port occupation handling and log file management in start_langgraph_dev

* feat: enhance async sub-agent handling and introduce comprehensive tests

- Updated `_maybe_swap_async_subagents` to improve async sub-agent management, ensuring internal flags are stripped before handoff.
- Enhanced port management in `onboard.py` to allow reuse of occupied ports if already running by the same service.
- Introduced file locking in `manager.py` to prevent race conditions during concurrent CLI invocations.
- Added new tests for async sub-agent swapping and langgraph manager functionalities to ensure reliability and correctness.
- Updated dependencies in `pyproject.toml` to include `psutil` and `filelock`.

* fix(docs): clarify sub-agent configuration in README

* test(manager): isolate _PID_DIR + tighten reuse-path assertion

Addresses CodeRabbit review on tests/test_langgraph_manager.py:

- Patch _PID_DIR to tmp_path so the FileLock setup in
  ensure_langgraph_dev doesn't mkdir the user's real
  ~/.config/evoscientist/ dir as a test side-effect.
- Tighten "result is None or hasattr(result, 'poll')" to a strict
  "result is None" — the reuse path returns None unconditionally,
  so the OR clause was hiding potential regressions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(manager): clean up stale PID file when unrelated process reuses PID

* feat(tests): add validation tests for async flag in load_subagents

* fix(load_subagents): restrict to .yaml files and clarify configuration handling

* fix(load_subagents): improve error handling for non-dict specifications in YAML

* feat(onboard): add "LangGraph Port" step to onboarding process

* feat(langgraph): add concurrency configuration for langgraph dev workers

* feat(async-subagents): enhance MCP tool routing for async sub-agents

* fix(manager): update exception handling for connection errors and prevent zombie processes

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-05 12:45:50 +01:00
X-iZhang d5ffa00263 docs(README): enhance Docker section with installation details and usage instructions 2026-05-04 15:37:05 +01:00
X-iZhang a297445f92 fix(assets): update wechat_group image file 2026-05-04 10:51:33 +01:00
Xi Zhang ab3787f86a fix(markdown): ensure proper spacing for ATX headings in Markdown ren… (#201)
* fix(markdown): ensure proper spacing for ATX headings in Markdown rendering

* fix(markdown): improve docstrings and tests for heading spacing functionality
2026-05-02 00:59:30 +01:00
Xi Zhang 29e7fd383d fix/cli hitl (#202)
* feat(display): enhance approval prompt with questionary for better navigation

* feat(cli): add HumanInTheLoopMiddleware for user approval in main agent
2026-05-02 00:54:40 +01:00
dinos 0d8ac4f24b feat(docker): official image with all runtime deps pre-installed (#198)
* feat(docker): official image with all runtime deps pre-installed

Multi-stage build using uv for the EvoScientist core + all messaging-channel
extras, plus Node.js 24 LTS (for npx-based MCP servers) and uv (for runtime
Python MCP installs) in the runtime layer. Runs as non-root user evosci,
with workspace, app data, and config (XDG_CONFIG_HOME) all consolidated
under a single /home/evosci/.evoscientist volume so a single mount
persists everything across container restarts.

Includes a docker-compose.yml starter, a build/push GitHub Actions
workflow targeting ghcr.io with multi-arch (amd64/arm64) and PR-only
build verification, a .dockerignore, and a new Docker section in the
README documenting mounts, derivation recipes for the unbundled stt /
oauth / TinyTeX extras, and proxy/cert handling expectations.

* fix(docker): pin trixie base + drop redundant python image

Switch builder and runtime from `python:3.11-slim-bookworm` to a single
`ghcr.io/astral-sh/uv:python3.11-trixie-slim` base — trixie drops several
CRITICAL vulnerabilities that bookworm carries today, and reusing the uv
image for runtime eliminates the separate `COPY --from=…/uv` line.

* chore(docker): pin GitHub Actions to commit SHAs in workflow

Replace mutable major-version tags with full commit SHAs (with the
corresponding semver tag in a trailing comment) so a compromised /
retagged action release can't silently change what runs in the publish
pipeline.

* chore(deps): enable Dependabot version updates for Dockerfile pins

Adds a weekly `docker` ecosystem that watches the Dockerfile's `FROM` /
`COPY --from=` references — including the ARG-bound `BASE_IMAGE` and
`NODE_IMAGE` digests — and opens one grouped PR per cadence bumping
both the @sha256 digest and the trailing version comment. This keeps
the otherwise-frozen pins flowing with Debian point releases and
upstream patches.

* fix(docker): use nodejs alias stage so NODE_IMAGE ARG actually resolves

`COPY --from=${NODE_IMAGE}` left the dollar-curly literal at parse time
under buildkit 29.x — it expands ARGs in `FROM` but reads `--from=` as a
static stage/image name. Introduce a tiny `FROM ${NODE_IMAGE} AS nodejs`
alias and `COPY --from=nodejs …` against it, which preserves the
ARG-driven Dependabot updates without tripping the parser.

* fix(docker): harden venv ownership and PATH ordering

- Drop `--chown` on the `/opt/venv` COPY so the venv stays root-owned.
  The runtime user only needs read+execute (default Unix perms allow
  that); making it user-owned let the agent rewrite its own
  dependencies, which defeats the sandboxing premise. All persistent
  agent state already lives under /home/evosci/.evoscientist/.
- Reorder PATH so /opt/venv/bin precedes the user-writable
  UV_TOOL_BIN_DIR. Otherwise a stray binary dropped into the latter
  (e.g. via `uv tool install`) could shadow the canonical
  `evosci` / `python` / `pip` shipped with the image.

* docs: update README

* docs(docker): warn about non-root UID and `curl | sh` for derived images

- The image runs as `evosci` (UID 1000), so a host-side `./workspace`
  bind mount fails if the host user has a different UID — same gotcha
  that bites onboarding's `mcp.yaml` write. Add an !IMPORTANT block
  with the two practical fixes (`chown -R 1000:1000` once, or
  `--user "$(id -u):$(id -g)"` on each run).
- The TinyTeX derivation snippet pipes an unpinned remote installer
  into `sh`. Add a one-line pointer to fetching a pinned release
  tarball from `rstudio/tinytex-releases` for users who'd rather not
  trust the upstream script blindly. The official installer is kept
  as the default since that's what TinyTeX itself recommends.

* chore(docker): cancel in-flight workflow runs + flag iMessage as host-only

- Add `concurrency: cancel-in-progress: true` to the docker workflow
  so successive pushes on the same ref supersede the prior run rather
  than queueing in parallel — multi-arch buildx is the slowest job in
  CI, no point burning minutes on superseded builds.
- Spell out that the docker image installs the `all-chanels` extra and
  call out iMessage as a deliberate host-only exclusion: it requires
  the `imsg` CLI bridging to macOS's Messages.app, which no Linux
  container config can satisfy.
2026-05-01 13:24:14 +02:00
dinos 73928a2d78 fix(onboard): detect installed skill packs via install manifest (#199)
* fix(onboard): detect installed skill packs via install manifest

Onboarding's _step_skills only inspected USER_SKILLS_DIR and matched
recommended entries by directory-name hint, so a pack like
EvoScientist/EvoSkills@skills (which explodes into paper-writing/,
evo-memory/, etc. under GLOBAL_SKILLS_DIR) was never detected and kept
appearing as not-yet-installed.

skills_manager now writes a per-tier .installed.yaml mapping skill
directory name -> original install source on every install, removes the
entry on uninstall, and exposes installed_sources(). _step_skills checks
both tiers and treats a recommended source as installed when present in
any manifest -- so packs are recognized regardless of how their child
dirs are named.

* fix(onboard): write install manifest atomically

Stage to a sibling temp file, fsync, then os.replace into place. A crash
mid-write can no longer leave a half-written .installed.yaml behind,
which would otherwise wipe out pack detection until the next reinstall.

* style: fmt

* fix(onboard): catch decode errors when loading install manifest

read_text() can raise UnicodeDecodeError on a hand-edited or corrupt .installed.yaml; pin encoding="utf-8" and add UnicodeError to the except clause so the function honors its "returns {} on any error" contract.
2026-05-01 13:23:59 +02:00
X-iZhang 27a97c3097 Refactor code structure for improved readability and maintainability 2026-04-30 22:00:19 +01:00
dinos 5c829942d7 refactor(cli): replace channel module globals with ChannelRuntime (#197)
* refactor(cli): replace channel module globals with ChannelRuntime

Removes _cli_agent / _cli_thread_id from EvoScientist/cli/channel.py
and threads a ChannelRuntime via CommandContext.channel_runtime so
/model and /channel rebind without poking module-level state.

* fix(cli): address coderabbit review

- _auto_start_channel: bind ChannelRuntime only after
  _start_channels_bus_mode succeeds, so a startup failure no longer
  leaves a stale binding pointing at channels that never started.
- _sync_tui_command_completion (TUI) and the Rich CLI command-completion
  paths: rebind the runtime on thread rotation, not just agent swap, so
  /new and /resume keep ChannelRuntime in sync with the running thread
  (matches the serve-mode hook contract).
- test_hook_syncs_channel_runtime: pin ctx.thread_id explicitly so a
  bare MagicMock attribute can't silently mutate runtime.thread_id.
- New regression test covering the rebind-on-thread-rotation contract.
2026-04-30 18:06:22 +01:00
Xi Zhang 7cfec02416 Fix/sessions migration sweep race (#195)
* fix(sessions): ensure migration sweep runs before yielding checkpointer to prevent race conditions

* fix(sessions): enhance migration sweep with progress indication and ETA estimation
2026-04-29 02:06:25 +01:00
Xi Zhang 50719ef256 Implement PruningCheckpointer for efficient checkpoint management and… (#194)
* Implement PruningCheckpointer for efficient checkpoint management and add comprehensive tests

- Introduced `PruningCheckpointer` to manage checkpoint pruning after each `aput()`, ensuring only the latest checkpoints are retained based on a configurable limit.
- Added migration sweep functionality to clean up legacy checkpoints and prevent database bloat.
- Enhanced `get_checkpointer()` to utilize the new `PruningCheckpointer` and trigger migration sweeps when necessary.
- Developed a suite of integration tests for `PruningCheckpointer`, covering various scenarios including pruning behavior, concurrent writes, and retention policies.
- Implemented tests for migration sweep functionality, ensuring proper partitioning and user version management.
- Added diagnostic helper `db_stats` to provide insights into the database state, including thread and checkpoint counts.

* feat(sessions): enhance pruning logic to handle legacy DBs without writes table

* fix(tests): prevent atexit hook leakage in TestMigrationSweep

* feat(tests): enhance TestPruningCheckpointer to validate put+prune serialization

* feat(tests): refactor mock path implementation for get_db_path in test cases
2026-04-28 22:26:59 +02:00
Xi Zhang 56cc2fef85 fix(deepseek): add empty-string fallback for reasoning_content in cross-provider scenarios (#192) 2026-04-27 19:19:50 +01:00
Xi Zhang 52f8d3a3a5 Feat/llm context window patch table (#191)
* feat(context-window): add model context window patch table and apply function

* fix(tests): clean up formatting in context window tests

* fix(tests): update context window tests for Claude model exceptions
2026-04-27 18:28:57 +01:00
X-iZhang 26e3452ef6 chore: update version to v0.0.9 in badges, README, and project files 2026-04-26 15:16:12 +01:00
Xi Zhang 20c06d4897 feat(deepseek): implement reasoning_content passback for multi-turn s… (#190)
* feat(deepseek): implement reasoning_content passback for multi-turn scenarios

* fix(tests): ensure consistent import of EvoScientist.llm.patches in test cases

* fix(tests): streamline tool_calls formatting in TestPatchDeepseekReasoningPassback

* feat(patches): add reasoning_content capture and re-injection for DeepSeek assistant messages

* fix(deepseek): optimize reasoning_content extraction and assignment in passback

* fix(deepseek): refine reasoning_content handling in OpenAI capture patch
2026-04-26 14:53:08 +01:00
Xi Zhang e48bc1cb71 Refactor/system prompt structure (#189)
* feat: Refactor system prompt structure and enhance documentation for clarity

* refactor: Improve clarity and consistency in prompt documentation

* refactor(tests): Improve readability of first-person avoidance test assertion

* refactor: Remove redundant datetime imports and enhance prompt documentation
2026-04-26 11:08:08 +01:00
Ziheng Zhang da74c325d6 fix(channel): scope stop and restore resume history (#186)
* fix(channel): scope stop and restore resume history

* refactor(channel): simplify stop and resume patch

* Delete PR_MESSAGE.md

* fix(channel): address review feedback

* fix(channel): address remaining review bugs

* fix(channel): clean up stopped request handling

* fix(channel): preserve resolved replies and sync tui commands

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-25 16:12:49 +01:00
X-iZhang 27fee3256e fix(clipboard): remove unnecessary check for selection end in copy_selection_to_clipboard 2026-04-25 13:04:19 +01:00
Xi Zhang 558360b558 feat: Enhance ModelPickerWidget for Ollama integration (#187)
* feat: Enhance ModelPickerWidget for Ollama integration

- Implemented a sentinel row for "Custom Ollama model..." in ModelPickerWidget, allowing users to input arbitrary model names.
- Updated action handling in ModelPickerWidget to manage transitions between list and input modes.
- Added async model discovery for Ollama models, integrating with the /model command to fetch locally installed models.
- Created tests for Ollama model discovery and ModelPickerWidget behavior, ensuring proper functionality and user experience.
- Refactored validate_ollama_connection and discover_ollama_models for improved error handling and response management.

* fix: Simplify code by removing unnecessary line breaks in ModelPickerWidget and test cases

* fix: Restore globals on set_chat_model failure to prevent half-switched session

* fix: Improve error handling in ModelCommand by restoring globals on failure
2026-04-25 00:37:48 +01:00
Xi Zhang 49b03c36eb feat: add support for gpt-5.5 model in the LLM configuration and tests (#188) 2026-04-24 23:52:55 +01:00
Wiktor Cupiał 155c4eaa40 fix: text copy on remote sessions/legacy terminal emulators (#185)
* fix: text copy on remote sessions/legacy terminal emulators

* fix: display warning only once
2026-04-24 23:16:39 +02:00
Xi Zhang c134a16e19 Fix/channel slash rich cli (#184)
* feat(cli): implement slash command dispatch for channel messages

* fix(cli): make EvoSci serve exit on Ctrl+C and hot-swap /model (#181)

* fix(cli): streamline debug logging and formatting in channel command handling

* fix(cli): ensure proper handling of asyncio event loop in slash command processing

* fix(cli): add error handling for unexpected exceptions in slash command dispatch

* fix(cli): improve error messaging for slash command dispatch failures

* fix(tests): enhance test setup by restoring channel globals and simplifying assertions

* fix(cli): enhance slash command handling across UI surfaces and improve resume command warnings
2026-04-24 16:10:25 +01:00
Xi Zhang 3831198f19 fix(chat): resolve model switch lag by tracking model/provider key (#180)
* fix(chat): resolve model switch lag by tracking model/provider key for cache invalidation

* style: format set_chat_model function for improved readability

* fix(chat): improve model switching logic to prevent unnecessary cache rebuilds

* fix(model): ensure globals are restored on agent load failure to prevent model switch issues
2026-04-24 11:36:33 +01:00
Xi Zhang 92df3e1844 Refactor/cli command manager (#178)
* feat(cli): migrate command handling to CommandManager and enhance UI interactions

* feat(cli): add /clear and /help commands to enhance user experience

* feat(cli): implement /new and /resume commands with interactive session management

* Refactor MCP and Skills Command Handling

- Moved the interactive picker style to a centralized widget for consistency across MCP and Skills commands.
- Updated the MCP command to remove the old command dispatch logic, delegating to the new InstallMCPCommand.
- Enhanced the Skills command to utilize a new interactive picker for skill selection, improving user experience.
- Implemented cancellation handling in the picker to differentiate between user cancellations and empty selections.
- Added comprehensive tests for the new command structures and picker functionalities to ensure reliability.

* refactor(cli): streamline CommandManager dispatch and remove deprecated command set

* refactor(cli): enhance error handling and state management in ChannelCommand and RichCLICommandUI

* refactor(cli): update lifecycle callback terminology and improve async prompt handling in RichCLICommandUI

* refactor(cli): enhance SlashCommandCompleter to dynamically fetch workspace directory for autocompletion

* refactor(cli): unify quit handling in RichCLICommandUI with shared _stop helper

* refactor(cli): remove hardcoded slash commands and utilize command manager for dynamic completion

* refactor(cli): update MCP and skills command files for improved clarity and organization

* refactor(mcp_ui): remove unnecessary newline in _show_mcp_config function
2026-04-24 10:38:44 +01:00
X-iZhang f704e3d761 update 2026-04-23 14:33:02 +01:00
dinos 65828e9666 perf(cli): cut startup latency and defer MCP loading to the background (#171)
* perf(cli): cut startup time of `evosci --help` from ~2.2s to ~0.3s

Module-level imports were eagerly pulling in langchain.chat_models (with
the whole anthropic/openai/google stack), langgraph, textual, and
prompt_toolkit on every invocation — even for `--help` or `config list`.

Defer those with PEP 562 `__getattr__`, using `lazy_loader.attach` (SPEC-1,
the scientific-python standard) where it's a clean attach pattern:

- `EvoScientist/llm/__init__.py`: attach `models` lazily so importing
  `context_window` from this package no longer drags in langchain.
- `EvoScientist/stream/__init__.py`: attach display/events lazily; split
  the shared Rich `Console` singleton into a new lightweight
  `stream/console.py` so callers that only need `console` skip the
  `stream.events` → `langchain_core.messages` chain.
- `EvoScientist/cli/__init__.py`: hand-rolled `__getattr__` (reaches into
  `..stream.state`, which `lazy_loader` doesn't cover) so `commands` and
  `app` are the only eager loads.
- `EvoScientist/cli/commands.py`: move `cmd_interactive`/`cmd_run` to
  in-function imports so prompt_toolkit + textual only load when the
  interactive path actually runs.
- `EvoScientist/cli/_constants.py`: read `AGENT_NAME` on demand so
  `build_metadata` doesn't eagerly import `sessions` (langgraph/aiosqlite).

Adds `lazy-loader>=0.5` as a dependency.

* feat(cli): defer MCP tool loading with live per-server progress

The CLI was blocking ~5 s on MCP tool enumeration before the first
prompt appeared. Move the agent construction off the event loop and
surface per-server progress so the user can interact immediately and see
what's happening.

MCP client:
- Add an `on_progress` callback to `load_mcp_tools` / `aload_mcp_tools`
  / `_load_tools` emitting `start` / `success` / `error` events per
  server.
- Fan connection attempts out with `asyncio.gather` so latency no longer
  scales linearly with server count; cap simultaneous attempts at
  `_MAX_CONCURRENT_CONNECTIONS` (8) via a semaphore so a big stdio fleet
  doesn't spawn every subprocess at once.

Agent wiring:
- Plumb `on_mcp_progress` through `create_cli_agent` / `_load_agent` /
  `load_mcp_and_build_kwargs` so CLI and TUI can plug in collectors.

CLI (`cmd_interactive`):
- Run `_load_agent` in a background thread via `asyncio.to_thread`; the
  prompt and banner render immediately.
- `_await_agent_ready()` awaits the task before each agent-using site
  (first turn, channel messages, `/channel`, `/compact`). Raises if
  called without a prior `_start_agent_load` instead of silently
  reloading without the SQLite checkpointer.
- Pre-prime the progress dict from `load_mcp_config()` so the
  bottom-toolbar's `N/M` denominator is stable from the first render.
- Wrap `session.prompt_async` in `patch_stdout(raw=True)` so
  `console.print` from the worker-thread progress callback lands cleanly
  above the prompt as inline chat messages instead of stomping the
  prompt cursor.

TUI (`EvoTextualInteractiveApp`):
- Same background load + `_await_agent_ready()` gates on every
  `self._agent` read.
- New `MCPLoaderWidget` mounted at the top of `#input-shell` shows a
  header with `N/M` and one live row per server (spinner → ✓ / ✗ with
  tool count or error detail). On completion:
  - all-clean loads auto-dismiss ~2.5 s later;
  - cache hits (no events ever fired) dismiss immediately rather than
    flashing a misleading "0/N loaded";
  - failures keep the widget mounted so the user can read the errors.
  - `dismissed` property lets the app clear its ref so late events from
    slow servers become no-ops. The error branch of `_on_agent_loaded`
    also settles the widget so a load failure can't leave the spinner
    animating forever.
- Chat input is `disabled` while MCP resolves — no placeholder hack, no
"waiting…" system message.

Shared:
- Hoist braille spinner frames to `status_bar.SPINNER_FRAMES` and import
  them in the TUI widget so CLI and TUI animate in sync.

Tests:
- Extend `test_agent_mcp_cache` fakes to accept the new `on_progress`
  kwarg.
- New `TestLoadToolsProgressCallback` in `test_mcp_client` exercises the
  event sequence for success/failure/mixed fleets, verifies a buggy
  callback doesn't break the load, and asserts the semaphore caps
  in-flight connections.

* style: ruff

* chore: update uv.lock

* chore: uv.lock

* fix: coderabbit issues

* style: fmt

* fix: move _await_agent_ready inside try block

* fix(tui): auto-dismiss MCP loader widget on failure

The widget was designed to stay mounted on failure so the user could
read error detail, but since it's pinned above the input it never went
away in practice — just permanent banner clutter.

Auto-dismiss on failure too, with a longer grace (12s vs 2.5s) so the
error summary stays readable.

* fix: address second coderabbit pass

- Channel handlers (CLI + TUI): catch agent-load failures so the
  channel request doesn't hang; CLI moves `_await_agent_ready()`
  inside the existing try/except, TUI catches explicitly and calls
  `_set_channel_response` with the error.

- Stale background loads: `prev.cancel()` only stops the asyncio
  wrapper, not the thread running `_load_agent`. Added a generation
  token (`agent_load_id` / `self._agent_load_id`) and gated both
  progress and completion callbacks on it so a superseded load can't
  clobber the current session's state or UI.

- TUI prompt lifecycle: added `_agent_load_pending()` and gated the
  `_process_channel_message` / `_handle_command` finally blocks on it
  so `/new` or `/resume` invoked from a command keeps the prompt
  disabled until the fresh load settles.

- TUI readiness failures: `_run_turn` and `_handle_command` now
  catch exceptions from `_await_agent_ready()` and surface a
  "Agent failed to load: …" system message instead of letting the
  exception escape into Textual's traceback panel.

* refactor(cli): share background agent loader between CLI and TUI

The CLI and TUI were carrying near-identical copies of the same
background-load state machine: the `agent_task`, the `agent_load_id`
generation token, the gated progress/completion callbacks, and the
per-server progress dict. Every CodeRabbit finding on that lifecycle
had to be fixed in both files.

Extract it into `cli/_agent_loader.py`:

- `MCPProgressTracker` — owns the `server -> (state, detail)` dict;
  exposes `prime`, `record`, `snapshot`, `totals`.

- `BackgroundAgentLoader` — owns `agent`, the in-flight task, and the
  generation token. Exposes `start(**loader_kwargs)`, `await_ready()`,
  `is_pending`. Internally gates all progress/completion callbacks by
  generation so a superseded load can't clobber the current session.
  UI-specific rendering plugs in via `on_progress` / `on_success` /
  `on_failure` callbacks.

Both surfaces now just wire their UI hooks; the loader file holds no
Rich / prompt_toolkit / Textual dependencies. Net -345 lines from
`interactive.py` + `tui_interactive.py`; +20 unit tests pinning the
lifecycle (generation filtering, cache-hit short-circuit, failure
reset, progress ordering).

* refactor(cli): make _on_done the sole authority for agent state transitions

await_ready no longer sets self.agent — it just awaits the task and
reads what _on_done already wrote. Eliminates the dual-write overlap
(asyncio guarantees done-callbacks fire in registration order).

* fix(tui): let users type during MCP load, only block on send

Remove prompt-disabling during background agent load — the TUI now
matches the CLI approach where the input stays enabled and only gates
on await_ready() at submit time. The MCPLoaderWidget still provides
visual feedback that loading is in progress.

* fix(loader): preserve real load error on await_ready; dedup failure message

CodeRabbit flagged two issues with the new loader:

1. After a failed load, `_on_done` nulled `self._task`, so the next
   `await_ready()` hit the "before start()" branch and the CLI wrapper
   remapped it to a misleading "checkpointer not available" message —
   losing the real exception (bad MCP config, network, etc.).

   Keep `_task` set on failure so `await_ready` re-raises the real
   exception. Added `needs_restart` so TUI's auto-retry check stays a
   one-liner and doesn't need to reach into task internals.

2. TUI reported each load failure twice: once from
   `_on_agent_load_failure` (the done-callback) and once from each
   caller of `_await_agent_ready` (`_run_turn`,
   `_process_channel_message`, `_handle_command`) catching the re-raise.

   `_on_agent_load_failure` is now the sole local reporter; callers
   just handle control flow (return cleanly, set channel response to
   unblock remote).

* fix(cli): wire /model handler through the agent loader

The /model command from main (merged via f1f0d7c) still reached for
`state["agent"]` (CLI) and `self._agent` (TUI) — both removed by the
background-loader refactor. CLI raised KeyError on first invocation;
TUI raised AttributeError. Writes to the old fields also had no effect
because every other code path now reads from `agent_loader.agent`, so
the model switch would have silently failed.

Route everything through the loader: `await _await_agent_ready()` up
front so /model doesn't race with the initial background load, build
the `CommandContext` with the current agent, and sync `ctx.agent` back
into `agent_loader.agent` (plus channel globals) when the command
replaces it.

* fix(cli): isolate progress callback, capture awaited agent, gate by requires_agent

Three CodeRabbit findings on the loader + command dispatch path:

- Wrap ``_on_progress`` in try/except inside the loader's gated wrapper
  so a buggy UI adapter can't bubble into ``loader_fn`` and fail the
  whole background load. The MCP client already protects this, but
  defence-in-depth keeps the loader self-contained.

- In CLI channel + main-loop streaming, capture the agent returned by
  ``_await_agent_ready()`` and pass that into ``run_streaming`` rather
  than reading ``agent_loader.agent`` after a subsequent ``await``.
  A concurrent ``/new``/``/resume``/``/model`` could have swapped it.

- Add ``requires_agent: ClassVar[bool] = False`` to ``Command`` and
  mark ``/compact``, ``/model``, ``/channel`` as ``True``. TUI dispatch
  sites (channel and keyboard) now check ``cmd_manager.resolve(...)``
  and only wait for readiness when the command actually needs the
  agent. ``/mcp add``, ``/skills``, ``/new`` etc. no longer deadlock
  behind a failing MCP load they are meant to fix.

* fix(cli): guard sync-back, subcommand-aware gating, /model adopt-path

Three CodeRabbit findings on command dispatch:

- ``_handle_command`` unconditionally synced ``ctx.agent`` back into
  ``agent_loader``.  For non-agent commands ``ctx.agent`` is ``None``,
  so ``/threads`` / ``/mcp`` / ``/skills`` (etc.) could clobber a valid
  loaded agent — and rebind channel globals to ``None``.  Guard the
  sync on ``ctx.agent is not None``.

- ``/channel status`` and ``/channel stop`` don't touch ``ctx.agent``
  but the class-level ``requires_agent = True`` blocked them behind
  agent readiness.  Added ``Command.needs_agent(args)`` (defaults to
  ``requires_agent``) so ``/channel`` can override with subcommand
  awareness; kept the class flag for the common case.

- ``/model`` builds a new agent from scratch, it never reads the
  existing one — gating it on readiness meant a broken provider
  blocked the command that would fix it.  Flipped it to
  ``requires_agent = False`` and added ``BackgroundAgentLoader.adopt``
  so the UI can seat the replacement and supersede any in-flight
  load (the generation token keeps a late completion from clobbering
  the adopted agent).

Bonus cleanup: ``CommandManager.resolve`` now returns
``(command, args)`` so callers can invoke ``needs_agent`` without
re-implementing ``shlex`` parsing.

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-22 18:22:33 +02:00
Wiktor Cupiał 28f3e81b4c feat: add /model command for changing models inside TUI/CLI (#162)
* rebase main

* fix: apply pr comments

* fix: fix critical issue

* fix: ordering /model in cli mode

* feat: refactor /model command handling and add Rich CLI support

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-04-22 14:43:21 +01:00
Xi Zhang aa3dd00409 feat: add support for session resumption with --resume flag and enhan… (#170)
* feat: add support for session resumption with --resume flag and enhance thread ID resolution

* refactor(tests): streamline help output testing for --resume flag

* feat: enhance session resume functionality with improved thread ID resolution and SQL wildcard handling

* feat: improve error handling for resume hint retrieval in interactive modes

* refactor: streamline logging for print_resume_hint failure in interactive mode

* feat: implement deferred scrolling for Markdown-heavy content in interactive mode
2026-04-21 21:26:40 +01:00
dinos 05f54334ba fix(mcp): stdio env passthrough + durable package installs (#169)
* fix(mcp): forward proxy and CA bundle env vars to stdio subprocesses

The MCP SDK's stdio transport inherits only a minimal allowlist (HOME,
PATH, USER, …) from the parent, stripping http_proxy/https_proxy and
SSL_CERT_FILE/REQUESTS_CA_BUNDLE/etc. Behind a proxy or with a custom CA
bundle, stdio MCP servers silently hang on outbound requests while the
same server over HTTP transport works. Auto-forward the proxy and cert
vars when present; user-configured env still takes precedence.

* fix(mcp): use `uv tool install` so MCP packages survive uv sync

Source installs previously used `uv pip install --python $VENV <pkg>`,
which lands in the evosci venv but is not recorded in pyproject.toml or
uv.lock. A subsequent `uv sync` (typical after `git pull`) reconciles
the venv to the lockfile and removes the MCP package, forcing users to
re-run onboard.

Prefer `uv tool install <pkg>` for the non-uv-tool install path: the
binary symlink in ~/.local/bin survives uv sync and evosci upgrades,
and the MCP server gets its own isolated env (no dep conflicts).
Verify the expected CLI entry point resolves afterward; if not (package
has no console-script), fall through to the old uv-pip path so
command-less packages still work.

The uv-tool-env path (`uv tool install evoscientist --with <pkg>`) is
unchanged — it was already durable via uv's receipt.

* fix(mcp): gate standalone uv tool install on verify_command

Previously `install_pip_package` would route every install through
`uv tool install <pkg>` when `verify_command` was None, returning
success as long as the uv subprocess exited 0. Library callers
(`evoscientist[oauth]`, `lark-oapi`, etc.) expect the package to land
in the active venv so they can import it — a standalone uv tool env
is not importable, so the import fails at the next line.

Gate the `uv tool install <pkg>` branch on `verify_command` being
set: that signals the caller wants a durable CLI binary, which is
what `uv tool install` produces. Library callers omit it and go
straight to the pip-install-into-venv path.

Also: log info messages on every fall-through so stale-binary and
entry-point-missing failure modes are debuggable, and document the
--with → standalone recovery path.

* fix(mcp): resolve MCP binaries to `uv tool dir --bin`, not `.venv/bin`

Under `uv run`, the project venv's `bin/` comes first on PATH, so
`shutil.which("arxiv-mcp-server")` returns a stale `.venv/bin/` copy
left over from an earlier install instead of the fresh symlink that
`uv tool install` just placed in `~/.local/bin`. The venv copy gets
written to mcp.yaml and is then wiped by the next `uv sync` — exactly
the failure mode the durability fix was meant to prevent.

Query `uv tool dir --bin` directly and prefer binaries found there
over `shutil.which`. Same change to the post-install verify in
`install_pip_package` so a venv shadow can't falsely short-circuit
the fallback.

* refactor(mcp): split install_pip_package into install_library + install_cli_tool

`verify_command` was doing double duty: naming the CLI binary to check
*and* signaling "this is a CLI install, use the standalone `uv tool
install` path." Callers routed library installs through the CLI branch
any time they forgot to pass it, and the resulting standalone uv tool
env wasn't importable from the active venv.

Separate the two use cases into distinct functions, each with one
install strategy per environment shape. Shared logic lives in private
`_install_with_uv_tool_env` / `_install_via_pip` helpers.

- install_library(pkg): uv-tool-env --with → pip. Never uses standalone
  `uv tool install <pkg>` (not importable from active venv).
- install_cli_tool(pkg, *, verify_command): uv-tool-env --with →
  standalone `uv tool install` → pip. `verify_command` is now required.

Callers pick the right function at the call site: registry.py picks
based on whether `entry.command` is set; onboard.py call sites all
install libraries.
2026-04-21 16:59:06 +01:00
X-iZhang 1e4c011b7c fix(assets): update wechat_group image for improved clarity 2026-04-21 12:06:49 +01:00
Xi Zhang 06822f236c feat: enhance tool result handling with tool_call_id for concurrent execution 2026-04-19 23:15:20 +01:00
Ziheng Zhang bd501cce34 fix(channel/qq): deliver HITL approval prompts reliably (#166)
* fix(channel/qq): deliver HITL approval prompts reliably

QQ approval prompts were silently dropped when the markdown send hit
a QQ server-side error (e.g. template not configured, content audit)
because the fallback path only matched TypeError / specific string
patterns, and the plain-text retry reused the already-consumed
msg_seq which QQ then rejects as duplicate.

- Consume a fresh msg_seq for the plain-text fallback send
- Recognize QQ server error codes (304014/304023/304003/40034059)
  and CN fragments ("模版"/"审核") as markdown-fallback triggers
- Promote send failure logs from debug to warning/error with
  chat_id/msg_id/seq so real-world errors can be diagnosed
- Extend test_qq_channel with a server-error-code fallback case

* style(channel/qq): apply ruff formatter to approval-delivery fix

* Fix
2026-04-19 11:14:23 +01:00
Xi Zhang 58435dba52 Release/v0.0.8 (#167)
* chore(release): update version to v0.0.8 and dependencies in project files

* feat(models): add new model entries for Claude Opus 4-7 and update version handling

* Refactor code structure for improved readability and maintainability
2026-04-18 16:49:31 +01:00
Xi Zhang f4a3617646 refactor(paths): unify global data directory to ~/.evoscientist and u… (#164)
* refactor(paths): unify global data directory to ~/.evoscientist and update related paths

* refactor(paths): update legacy session migration to respect XDG_CONFIG_HOME

* refactor(tests): clear XDG_CONFIG_HOME in legacy session migration tests for deterministic behavior
2026-04-18 15:01:22 +01:00
Xi Zhang 7c6b6755f2 fix(docs): update survey literature and macOS deployment links for accuracy 2026-04-17 15:25:23 +01:00
Xi Zhang 4b2aaea49a fix(docs): update survey literature link to point to the correct GitHub path 2026-04-17 15:20:59 +01:00
Xi Zhang 1965dd8661 Add new asset images for survey literature examples
- Added model_selection.png to illustrate model selection process.
- Added prompt.png for visual representation of prompts used in surveys.
- Added skill_selection.png to depict skill selection criteria.
2026-04-17 15:18:06 +01:00
dependabot[bot] 03b4412c9a chore(deps): bump langchain-openai in the uv group across 1 directory (#163)
Bumps the uv group with 1 update in the / directory: [langchain-openai](https://github.com/langchain-ai/langchain).


Updates `langchain-openai` from 1.1.12 to 1.1.14
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain-openai==1.1.12...langchain-openai==1.1.14)

---
updated-dependencies:
- dependency-name: langchain-openai
  dependency-version: 1.1.14
  dependency-type: direct:production
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-17 10:27:38 +02:00
Xi Zhang 210e8864f6 feat(memory): migrate MEMORY.md to global path & enhance ask-user prompts (#161)
* feat(prompt): enhance user interaction with multiple-choice and free-text questions

* refactor(paths): rename MEMORY_DIR to MEMORIES_DIR for consistency

* style(tests): format code for better readability in test cases

* feat(prompt): add validation for 'other' option in user prompt

* feat(prompt): refactor validation logic for user prompts and add skip option

* feat(style): refactor to use shared _PICKER_STYLE from interactive module
2026-04-16 15:33:45 +01:00
dependabot[bot] 7f522cb4fe chore(deps): bump langsmith in the uv group across 1 directory (#160)
Bumps the uv group with 1 update in the / directory: [langsmith](https://github.com/langchain-ai/langsmith-sdk).


Updates `langsmith` from 0.7.30 to 0.7.31
- [Release notes](https://github.com/langchain-ai/langsmith-sdk/releases)
- [Commits](https://github.com/langchain-ai/langsmith-sdk/compare/v0.7.30...v0.7.31)

---
updated-dependencies:
- dependency-name: langsmith
  dependency-version: 0.7.31
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-16 10:38:20 +01:00
Xi Zhang 0b7c162d1b fix(skill-manager): update skill source filtering to include workspace and global tiers (#159) 2026-04-15 15:29:11 +01:00
Ziheng Zhang a2d2ddc5a2 fix(channel): avoid replaying thinking after resume (#154)
* fix(channel): avoid replaying thinking after resume

* fix(channel): relay fresh thinking after resume

* Fix

* Fix

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-15 14:05:47 +01:00
dinos 2961e5ee88 fix(minimax): correct API endpoint and add region selection (#158) 2026-04-15 14:30:26 +02:00
Xi Zhang c4237fecb8 feat(backends): rename MergedReadOnlyBackend to MergedSkillsBackend a… (#157)
* feat(backends): rename MergedReadOnlyBackend to MergedSkillsBackend and update documentation for clarity

refactor(paths): simplify ensure_dirs function to create only memory directory eagerly

fix(prompts): update skills availability description for accuracy

test(paths): adjust test to reflect skills directory creation on demand

* refactor(tests): format assertion for skills directory existence in ensure_dirs test
2026-04-15 01:52:31 +01:00
X-iZhang 64721f5967 update 2026-04-13 20:57:27 +01:00
Xi Zhang 65db3a4fcd Add status bar and compact summary widgets with context window resolu… (#152)
* Add status bar and compact summary widgets with context window resolution

- Implemented a shared status bar for CLI and TUI frontends, including helpers for managing session metrics and context windows.
- Created a `CompactSummaryWidget` for displaying manual summaries in a collapsible format.
- Introduced a `CompactingWidget` to indicate ongoing compacting processes.
- Added a base class `TimedStatusWidget` for widgets that require a timer.
- Developed context window resolution helpers to retrieve context window sizes from various model attributes.
- Enhanced tests for context window resolution and status bar functionalities, ensuring accurate behavior across different scenarios.
- Updated existing tests to cover new features and maintain code quality.

* refactor(Channel): simplify lambda function in _send_with_retry method

* feat: enhance context editing logic and improve error handling in StreamState

* refactor(Channel): streamline lambda function in _send_with_retry method

* feat: rename auto-approve option to auto-mode for unattended execution; update checkpoint queries to filter by agent name; improve compatibility validation logic

* feat: rename auto-approve option to auto-mode; update related logic and tests for improved unattended execution

* fix: correct formatting of console message for MCP server configuration status

* feat: add check for None summary_message in _apply_summarization_event to prevent errors

* feat: enhance _load_checkpoint_messages to validate message format and apply summarization event
2026-04-12 17:47:37 +01:00
Xi Zhang ff15f515cc Release/v0.0.7 (#151)
* chore(assets): update wechat_group image file

* Refactor code structure for improved readability and maintainability

* feat(backends): enhance MergedReadOnlyBackend with improved ls, grep, and glob methods

* fix(docs): update WeChat QR code image link in README files

* feat(skills): enhance skill management to support global and workspace tiers

* style: apply ruff format to skills_cmd and commands/implementation/skills

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(skills): improve uninstall_skill to prevent removal of built-in skills

* fix(docs): update skill installation documentation for clarity on global and user directories

* fix(skills): enhance uninstall_skill to validate skill directory before removal

* fix(skills): improve error handling in install_skill and uninstall_skill for directory creation and validation

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-10 20:04:05 +01:00
Xi Zhang d5b982c980 fix(ccproxy): update Responses API handling and patch system role con… (#149)
* fix(ccproxy): update Responses API handling and patch system role conversion

* fix(ccproxy): streamline _agenerate method in system to developer patch

* fix(ccproxy): improve handling of None output in Codex compatibility patch
2026-04-09 12:51:05 +02:00
Ziheng Zhang 3e493233cf feat(channels): simplify debug tracing and add serve debug mode (#143)
* feat(channels): simplify debug tracing and add serve debug mode

* Fix

* fix(channels): remove serve loop patch
2026-04-09 12:47:48 +02:00
dependabot[bot] 011dffdabc chore(deps): bump the uv group across 1 directory with 2 updates (#148)
Bumps the uv group with 2 updates in the / directory: [cryptography](https://github.com/pyca/cryptography) and [langchain-core](https://github.com/langchain-ai/langchain).


Updates `cryptography` from 46.0.6 to 46.0.7
- [Changelog](https://github.com/pyca/cryptography/blob/main/CHANGELOG.rst)
- [Commits](https://github.com/pyca/cryptography/compare/46.0.6...46.0.7)

Updates `langchain-core` from 1.2.25 to 1.2.28
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain-core==1.2.25...langchain-core==1.2.28)

---
updated-dependencies:
- dependency-name: cryptography
  dependency-version: 46.0.7
  dependency-type: indirect
  dependency-group: uv
- dependency-name: langchain-core
  dependency-version: 1.2.28
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-09 00:09:19 +01:00
Xi Zhang 4f11de23a2 fix(llm): patch _stream/_astream for OpenAI-compatible content flattening (#147)
* fix(llm): patch _stream/_astream for OpenAI-compatible content flattening

_patch_openai_compat_content() only patched _generate/_agenerate but
EvoSci CLI uses streaming paths. This extends the content flattening
to _stream/_astream so strict OpenAI-compatible relays receive plain
string content during streaming calls.

Closes #142

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): use asyncio.run() instead of pytest-asyncio for CI compat

CI does not have pytest-asyncio installed, so async tests must use
asyncio.run() wrapper instead of @pytest.mark.asyncio decorator.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): use @pytest.mark.anyio for async tests (CI compat)

CI does not have pytest-asyncio. Use @pytest.mark.anyio consistent
with existing async tests in the project.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 20:24:13 +01:00
Peidong Yang 066f36fa13 feat: add Moonshot and Kimi Coding Plan as LLM providers (#128)
* feat: add Moonshot and Kimi Coding Plan as LLM providers

Add two new providers for Moonshot AI:
- `moonshot`: OpenAI-compatible direct API (api.moonshot.cn/v1) with
  kimi-k2.5, kimi-k2-thinking, moonshot-v1-auto/128k/32k/8k models
- `kimi-coding`: Anthropic-compatible Kimi Coding Plan endpoint
  (api.kimi.com/coding/) with User-Agent header for compatibility

Both providers disable thinking to avoid multi-turn tool calling
errors caused by LangChain dropping reasoning_content from history.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: update Moonshot thinking comment and add provider assertions

- Add clarifying comment for disabling thinking on all Moonshot models
- Add moonshot and kimi-coding assertions to test_entries_has_all_providers

* fix: exclude Moonshot and Kimi Coding from content patch

Tested and verified both APIs support standard list content format:
- Moonshot (OpenAI-compatible): supports list content, no patch needed
- Kimi Coding (Anthropic-compatible): supports list content, no patch needed

Only apply _patch_openai_compat_content to strict providers like DeepSeek.

* fix: set _original_provider in routed provider branches

Ensure _original_provider is set before provider is reassigned to
'openai' or 'anthropic', so the no-patch exclusion for Moonshot
and Kimi Coding works correctly.

* style: translate Moonshot comments to English

* style: translate comment to English to fix ruff lint error

* merge: resolve conflicts

* chore: revert uv.lock and translate Chinese comments to English

Revert unrelated uv.lock dependency changes and replace Chinese code
comments with English for codebase consistency per review feedback.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: fix ruff format for models.py

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: ypd <ypd@ypddeMac-mini.local>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: Xiaohui Yan <xhcloud@gmail.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-04-08 19:53:49 +01:00
JackyFan 632159261c Fix: Support XDG_CONFIG_HOME for sessions.db on Windows with non-ASCII usernames (#102)
* fix: support XDG_CONFIG_HOME for sessions.db on Windows

Fixes SQLite database opening failure on Windows systems with non-ASCII
usernames (e.g., Chinese characters). The get_db_path() function now
supports the XDG_CONFIG_HOME environment variable, consistent with
get_config_dir() in settings.py.

Closes #101

* fix: auto-resolve Windows Unicode path for sqlite3 via 8.3 short path

Refactor get_db_path() to reuse get_config_dir() (XDG_CONFIG_HOME
support) and add _to_short_path() helper that converts the config
directory to its Windows 8.3 short form via GetShortPathNameW. This
automatically resolves sqlite3 failures on Windows systems with
non-ASCII usernames (e.g., Chinese characters) without requiring
manual environment variable configuration.

The short-path conversion is best-effort: it targets the directory
(which exists after mkdir) rather than the db file, and falls back
gracefully on non-Windows, non-NTFS, or when 8.3 naming is disabled.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-08 18:50:42 +01:00
Ziheng Zhang 3647ff1afa fix: preserve QQ markdown formatting and newline rendering (#144)
* fix: preserve QQ markdown formatting

* fix: remove stale qq trace fallback hook

* fix: narrow qq markdown fallback handling

* style: format qq channel with ruff
2026-04-08 13:52:45 +01:00
Xiaohui Yan e89b71aa60 feat(cli): add --debug flag for verbose logging in serve mode (#141)
* feat(cli): add --debug flag for verbose logging in serve mode

* feat(cli): add log_level config field with priority over env var

Replace dead `debug` parameter in `main()` with a proper `log_level`
config field in EvoScientistConfig. Enables `EvoSci config set log_level
debug` with priority: config file > EVOSCIENTIST_LOG_LEVEL env var.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-08 13:23:10 +08:00
Ziheng Zhang 729cab11be Fix late channel response delivery after timeout (#140)
* Fix late channel response delivery after timeout

* Fix
2026-04-07 15:48:12 +01:00
X-iZhang 080a5c06f3 feat: enhance OpenRouter support with additional reasoning handling and model entries 2026-04-03 17:14:19 +01:00
X-iZhang 117750eda7 chore: update wechat group image asset 2026-04-03 15:44:40 +01:00
Xi Zhang 028dbe79d3 Release/v0.0.6 (#138)
* chore: clean up empty code change sections in the changes log

* feat: add adaptive tools and context editing features to README
2026-04-03 15:20:45 +01:00
Allen d08ca535c6 fix:Telegram channel start failed #133 (#134)
* fix:Telegram channel start failed #133

* refactor: simplify bus thread and add nest_asyncio warning

- Remove unnecessary _run_as_task() wrapper; run_until_complete()
  already creates a Task internally via ensure_future()
- Add comment noting nest_asyncio.apply() is global and irreversible

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Xi Zhang <zacharyzhang2022@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-03 14:47:15 +01:00
Yinhan Lu 5aa8353613 feat(llm): upgrade OpenAI reasoning effort from high to xhigh (#136)
* feat(llm): upgrade OpenAI reasoning effort from high to xhigh

The OpenAI Responses API supports "xhigh" as a reasoning effort level,
which provides deeper reasoning than "high". This is already used by
other CLI tools (e.g., OpenClaw) for OpenAI models.

Only affects the direct API key path; the ccproxy/OAuth path is
unchanged (reasoning is still skipped there).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(llm): limit xhigh reasoning to gpt-5.4+ and codex models

Only gpt-5.4 series and codex models support xhigh reasoning effort.
Older models (gpt-5, gpt-5.1, gpt-5.2, gpt-5.3) fall back to high.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Xi Zhang <zacharyzhang2022@gmail.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-03 14:35:14 +01:00
Xiaohui Yan 91de173c78 feat: Add GLM-5.1 support for Zhipu providers (#137)
* feat: Add GLM-5.1 support for Zhipu providers

Add GLM-5.1 model entries to both zhipu-code (coding endpoint) and
zhipu (general endpoint) providers, following the existing pattern
for GLM models.

* feat: add glm-5v-turbo support for Zhipu providers

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: Xi Zhang <zacharyzhang2022@gmail.com>
2026-04-03 14:00:21 +01:00
dependabot[bot] 70629e254c chore(deps): bump aiohttp in the uv group across 1 directory (#132)
---
updated-dependencies:
- dependency-name: aiohttp
  dependency-version: 3.13.4
  dependency-type: direct:production
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-04-03 20:38:13 +08:00
Xi Zhang 68f3ab2962 feat: enable reasoning for OpenRouter via extra_body to prevent multi… (#124)
* feat: enable reasoning for OpenRouter via extra_body to prevent multi-turn errors

* feat: implement OpenRouter native reasoning support and patch langchain-openrouter bug

* feat: add OpenRouter reasoning effort configuration and update related tests

* feat: add langchain-openrouter dependency for enhanced reasoning support

* fix: correct spacing in reasoning effort choice label

* feat: implement patch for OpenRouter reasoning details to prevent Pydantic errors

* feat: add patches for OpenRouter reasoning and content handling utilities

* feat: prevent multiple patches of OpenRouter reasoning details by using a global flag

* feat: update OpenRouter reasoning patch to ensure single application with global flag

* feat: refine OpenAI responses API handling to apply only for OpenAI provider

* feat: Enhance TUI interaction by updating todo widget positioning and skipping empty tool call chunks

* feat: Update tool selector threshold and adjust logging level for selector failures

* feat: Temporarily disable timestamp toast in tool call widget for UX review

* feat: Re-enable timestamp toast in tool call widget on click
2026-04-03 11:04:34 +01:00
Xi Zhang 897b444d27 Feat/context management middleware (#127)
* feat: Add context management middleware for improved error handling and context editing

* feat: Implement LLMToolSelectorMiddleware for enhanced tool selection and tracking

* feat(tests): update test functions to include mock timestamp parameter

* refactor: simplify tool selection state storage and update comments in middleware

* feat: Enhance tool selection handling and suppress structured output for improved event streaming

* refactor: simplify patching in test_create_tool_selector functions

* feat: add model parameter to create_tool_selector_middleware for enhanced flexibility

* feat: enhance tool selection suppression with JSON buffering for improved accuracy
2026-04-02 22:39:20 +02:00
Xi Zhang d8720a0b35 feat: Upgrade ccproxy to version 0.2.7 and remove deprecated thinking… (#130)
* feat: Upgrade ccproxy to version 0.2.7 and remove deprecated thinking tag handling

* feat: Enhance ccproxy compatibility and strip legacy thinking tags
2026-04-02 15:09:20 +01:00
dinos b9e809aeb6 fix: use uv tool install --with for durable MCP server installs (#125)
* fix: use `uv tool install --with` for durable MCP server installs (#121)

When EvoScientist is installed via `uv tool install`, MCP server packages
added during onboarding were installed with `uv pip install`, which is
not tracked by uv. Running `uv tool upgrade evoscientist` would recreate
the venv from scratch and silently wipe the MCP server binaries.

Now `install_pip_package()` detects uv tool environments and uses
`uv tool install <tool> --with <package>`, which records the dependency
in uv-receipt.toml so it survives upgrades. Existing --with packages
are read from the receipt and preserved.

Falls back to the old `uv pip install` path if the durable method fails.

* style: fmt

* fix: preserve requirement specs and normalize dedup in uv tool installs

Address review feedback: _uv_tool_existing_requirements() now returns
a dict mapping bare names to full PEP 508 specs (preserving extras and
version constraints from uv-receipt.toml). Dedup check uses
_bare_package_name() to normalize the incoming package argument before
comparing against receipt entries.
2026-04-02 10:51:54 +02:00
dependabot[bot] 3c7e1e5b8f chore(deps): bump anthropic in the uv group across 1 directory (#129)
Bumps the uv group with 1 update in the / directory: [anthropic](https://github.com/anthropics/anthropic-sdk-python).


Updates `anthropic` from 0.86.0 to 0.87.0
- [Release notes](https://github.com/anthropics/anthropic-sdk-python/releases)
- [Changelog](https://github.com/anthropics/anthropic-sdk-python/blob/main/CHANGELOG.md)
- [Commits](https://github.com/anthropics/anthropic-sdk-python/compare/v0.86.0...v0.87.0)

---
updated-dependencies:
- dependency-name: anthropic
  dependency-version: 0.87.0
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-04-02 10:41:20 +02:00
Xi Zhang 6e5844a68f Fix/tui file mention display (#123)
* fix: enhance user message display in run_textual_interactive function

* fix: add binary file detection in file mention resolution
2026-03-31 16:52:46 +08:00
Xi Zhang dd6cbe3b45 docs: clarify uv upgrade command in README files (#120) 2026-03-30 09:58:04 +02:00
dependabot[bot] 5c9cd353b9 chore(deps): bump cryptography in the uv group across 1 directory (#117)
Bumps the uv group with 1 update in the / directory: [cryptography](https://github.com/pyca/cryptography).


Updates `cryptography` from 46.0.5 to 46.0.6
- [Changelog](https://github.com/pyca/cryptography/blob/main/CHANGELOG.rst)
- [Commits](https://github.com/pyca/cryptography/compare/46.0.5...46.0.6)

---
updated-dependencies:
- dependency-name: cryptography
  dependency-version: 46.0.6
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-03-29 13:58:49 +02:00
Yuyue Zhao 4505300c8d feat: Add **More Effort** code generation mode (#118)
* feat: Add **More Effort** code generation mode

* feat: Enhance code generation mode selection and update documentation

---------

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-03-28 18:17:47 +00:00
dependabot[bot] 29533cce7f chore(deps): bump langchain-core in the uv group across 1 directory (#114)
Bumps the uv group with 1 update in the / directory: [langchain-core](https://github.com/langchain-ai/langchain).


Updates `langchain-core` from 1.2.21 to 1.2.22
- [Release notes](https://github.com/langchain-ai/langchain/releases)
- [Commits](https://github.com/langchain-ai/langchain/compare/langchain-core==1.2.21...langchain-core==1.2.22)

---
updated-dependencies:
- dependency-name: langchain-core
  dependency-version: 1.2.22
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-03-28 13:23:59 +01:00
Xi Zhang db72fb507f fix: update wechat_group image asset (#116) 2026-03-28 13:19:27 +01:00
Xi Zhang 6e23e98cca fix: inject thread_id into LangGraph context for compact conversation (#112) 2026-03-27 20:46:12 +00:00
Yuyue Zhao 4a9562229f Go0day patch 1 (#113)
* Update news section with latest ranking information

* Update awards and rankings in README.zh-CN.md
2026-03-27 20:40:34 +00:00
Yuyue Zhao 10e0364163 Update awards section in README.md (#111) 2026-03-27 19:15:43 +00:00
Xi Zhang d3284ee881 v0.0.5 (#110)
* fix: update OpenRouter API key validation to use /auth/key endpoint and httpx

* Refactor code structure for improved readability and maintainability

* feat: enhance welcome banner to include file commands indication

* feat: update LaTeX setup prompt to use selection UI for better user experience

* feat: update version to v0.0.5 in badges and project configuration
2026-03-27 20:03:36 +01:00
dinos 6dc8a25579 feat: add use_responses_api config to force Chat Completions for OpenAI relays (#105)
* feat(config): use_responses_api (#98)

langchain-openai auto-switches to the Responses API when reasoning
params are set, which breaks OpenAI-compatible relays that only support
Chat Completions. This adds a user-facing config option to override
that behavior:

  evosci config set use_responses_api false
  # or EVOSCIENTIST_USE_RESPONSES_API=false

* fix: propagate use_responses_api from config file and add normalization tests

Address PR #105 review comments:
- apply_config_to_env() now sets EVOSCIENTIST_USE_RESPONSES_API so
  config file values take effect (not just the env var directly)
- Add parametrized tests for case/whitespace normalization

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-03-27 16:15:48 +00:00
MuXinCG a813f8afd9 fix(feishu): isolate SDK event loop to prevent cross-thread RuntimeError on Linux
lark_oapi.ws.client captures the main thread's event loop in a
module-level variable at import time. When the WebSocket SDK thread
calls loop.run_until_complete() on that shared loop, nest_asyncio's
global patches cause task-tracking conflicts on Linux/Python 3.12:
    - RuntimeError: Leaving task … does not match the current task
    - AttributeError: 'NoneType' object has no attribute 'select'

Replace the previous Handle._run monkey-patch (which only suppressed
symptoms) with a proper fix: create a fresh event loop in the SDK
thread and swap the module-level loop variable so the SDK operates
on a fully isolated loop with no cross-thread interaction.

Closes #97
2026-03-27 22:23:08 +08:00
Jan Piotrowski 5bc33adeb0 fix(tui): simplify behaviour of pasting clipboard to textbox
Closes #107
2026-03-27 14:39:34 +01:00
dependabot[bot] 5ae3d47945 chore(deps): bump requests in the uv group across 1 directory
Bumps the uv group with 1 update in the / directory: [requests](https://github.com/psf/requests).

Updates `requests` from 2.32.5 to 2.33.0
- [Release notes](https://github.com/psf/requests/releases)
- [Changelog](https://github.com/psf/requests/blob/main/HISTORY.md)
- [Commits](https://github.com/psf/requests/compare/v2.32.5...v2.33.0)

---
updated-dependencies:
- dependency-name: requests
  dependency-version: 2.33.0
  dependency-type: indirect
  dependency-group: uv
...

Signed-off-by: dependabot[bot] <support@github.com>
2026-03-27 14:31:30 +01:00
Xi Zhang fe6a2c4b83 fix: resolve #93 review comments and close #95 (#92 regression fix included) (#96)
* feat: enhance file mention parsing with deduplication and warning handling

* feat: optimize ancestor grouping by improving path comparison efficiency

* fix: update middleware injection to use extend for better readability

* Refactor code structure for improved readability and maintainability

* feat: update README files to include AstaBench ranking and adjust award image layout

* update

* fix: deduplicate file mentions and improve warning message formatting
2026-03-25 17:01:46 +00:00
Jan Piotrowski 259842243d feat: add context retry middleware when API returns 4xx errors cause by context exceeding limits (#92) 2026-03-25 15:41:06 +00:00
Xi Zhang fab5f85eee v0.0.4 (#93)
* feat(tui): enhance conversation history rendering and implement two-level thread hierarchy in picker

* feat(tui): improve conversation history display and enhance thread selection UI

* feat(file_mentions): implement @file mention parsing and completion for CLI and TUI

* feat(uv-tool): add compatibility checks and installation helpers for uv tool environments

* feat(dependencies): update package versions in uv.lock for compatibility and improvements

* feat(badges): update PyPI version to v0.0.4 in SVG assets and README files

* feat(tests): format code in TestUvToolCompat for improved readability
2026-03-24 18:13:42 +00:00
Xi Zhang cc266dd6b1 Feat/latex onboard setup and tui completion fix (#88)
* feat(onboard): add LaTeX setup steps and TinyTeX installation helpers

* fix(onboard): improve LaTeX status output to a single-line summary

* fix(assets): update wechat_group image
2026-03-23 10:51:42 +00:00
Ziheng Zhang 65cec64445 feat(feishu): add WebSocket long connection subscription mode (#87)
* feat(feishu): add WebSocket long connection subscription mode

Add WebSocket (长连接) mode as an alternative to webhook for Feishu
event subscription. This allows running without a public IP, port
forwarding, or tunnel — ideal for local dev and NAT/firewall setups.

- New `feishu_subscription_mode` config: "webhook" (default) or "websocket"
- WebSocket mode uses official `lark-oapi` SDK with thread-safe queue bridge
- Onboard wizard: mode selection, SDK install prompt for websocket
- CLI: `--mode webhook|websocket` for standalone serve
- `pip install evoscientist[feishu]` optional dependency
- 5 new tests covering config, SDK missing error, message bridge, cleanup
- Docs: subscription mode comparison table, prerequisites per mode

* Fix: Ruff

* Fix: small fix
2026-03-22 14:45:28 +00:00
Xi Zhang 3503142af6 feat(tui): UX polish — multi-line input, timestamps, update check & v0.0.3 (#85)
* feat: implement background update check and startup notifications

* feat: enhance user experience with timestamp notifications and UI polish

* feat: implement multi-line chat input with Enter-to-submit and modifier+Enter newline

* update

* update

* v0.0.3

* feat: improve code readability with consistent formatting in TUI and test files

* feat: update PyPI badge version to v0.0.3 in README files

* feat: add docstrings for test classes in test_update_check.py
2026-03-20 19:57:35 +00:00
Ziheng Zhang a7d3ef8087 fix: subagent summerize (#83)
* fix: subagent summerize

* fix: group subagent text by agent name for parallel fallback

The flat subagent_text_buffer list would interleave text from parallel
sub-agents into incoherent output. Replace with a dict grouped by
agent name so each sub-agent's text stays coherent, with [name]:
attribution when multiple agents contribute.

Add comprehensive tests for the new behavior (24 tests).

* fix: group subagent text by agent name for parallel fallback

The flat subagent_text_buffer list would interleave text from parallel
sub-agents into incoherent output. Replace with a dict grouped by
agent name so each sub-agent's text stays coherent, with [name]:
attribution when multiple agents contribute.

Add comprehensive tests for the new behavior (24 tests).

* fix: group subagent text by agent name for parallel fallback

The flat subagent_text_buffer list would interleave text from parallel
sub-agents into incoherent output. Replace with a dict grouped by
agent name so each sub-agent's text stays coherent, with [name]:
attribution when multiple agents contribute.

Add comprehensive tests for the new behavior (24 tests).

* chore: fix multiple agent

* test: add test

* fix linter

* remove redundant
2026-03-20 17:13:27 +00:00
Xi Zhang 372a6272b7 fix(onboard): Windows npx detection & rename /install-skills to /evoskills (#79)
* fix: improve npx detection on Windows using shutil.which

* fix: rename /install-skills command to /evoskills for consistency
2026-03-20 10:42:56 +00:00
Jiao Huifeng fdafebd47f feat: add STT voice transcription for all messaging channels (#28)
* feat: add STT voice transcription for all channels

Automatically transcribes audio/voice messages (Telegram, WeChat, Slack,
etc.) into text before the agent sees them. Enabled via config, off by default.

Changes:
- EvoScientist/stt.py: new STT engine using faster-whisper with lazy
  model loading and per-language model selection (zh/en/auto)
- EvoScientist/channels/base.py: hook in _enqueue_raw() to transcribe
  audio files and prepend transcript to message text; removes the raw
  [voice: ...] annotation after successful transcription so the agent
  does not attempt further audio processing
- EvoScientist/config/settings.py: stt_enabled (default False),
  stt_language (default "auto")
- pyproject.toml: optional [stt] dependency group (faster-whisper>=1.0)
- tests/test_stt.py: unit tests covering all backends and channel integration

Usage:
  pip install 'EvoScientist[stt]'
  EvoSci config set stt_enabled true
  EvoSci config set stt_language zh   # zh / en / auto

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: remove unused imports (ruff F401)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix: address PR #28 reviewer feedback

Changes per SemiGlassFace review (CHANGES_REQUESTED):

1. Cache config at channel __init__ — no longer calls load_config() on
   every incoming message; STT settings stored as instance attributes
   (_stt_enabled, _stt_language, _stt_model, _stt_device,
   _stt_compute_type) set once during Channel.__init__().

2. Replace deprecated asyncio.get_event_loop() with get_running_loop()
   to avoid DeprecationWarning on Python 3.12+.

3. Annotation removal now uses exact path matching instead of substring
   search — checks fp == a or a.endswith(f": {fp}]") so only the
   correct annotation is removed after transcription.

4. Expose stt_model, stt_device, stt_compute_type as config fields so
   users can override the HuggingFace model id, inference device, and
   quantisation without touching code. transcribe_file() forwards all
   three to the engine.

Also: _engines dict replaced with single _engine + _engine_key tuple
(model_id, device, compute_type) — reuses cached model unless settings
change, simpler than a dict.

Tests: 19 STT-specific tests all pass; total 1105 tests green, ruff clean.

* fix: resolve ruff lint errors (UP037, I001, PT006)

* style: apply ruff format

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-20 10:59:30 +01:00
Xi Zhang 210d71c590 feat(tui): double Ctrl+C quit confirmation & docs Examples/Recipes section (#78)
* feat: enhance quit handling with double Ctrl+C confirmation and cleanup logic

* feat: enhance TUI cancellation handling and improve user interruption messages

* feat: add Examples & Recipes section to documentation

* docs: remove guideline to follow the structure of existing recipes
2026-03-20 00:25:31 +00:00
Icy Fish 649f2ca121 fix: resolve TAB cursor disappearance and up/down double-handling (#58)
- Add priority binding for TAB to intercept before Textual's focus_next
- Remove duplicate up/down handling in on_key (now handled by priority
  bindings from PR #76)
- Update tests to use cmd_manager.list_commands() instead of removed
  _TUI_SLASH_COMMANDS

Closes #57

Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-19 23:28:34 +00:00
Ziheng Zhang c627aadd39 Fix/deepseek sequence content error (#77)
* feat: add DeepSeek as a recognized third-party provider

Register DeepSeek API (https://api.deepseek.com) with DEEPSEEK_API_KEY
env var and add model short names: deepseek-r1 → deepseek-reasoner,
deepseek-v3 → deepseek-chat.

* feat: add _flatten_message_content utility for list-to-string conversion

Extract text from content block lists while skipping thinking/reasoning
blocks. This handles the case where LangChain stores assistant messages
with content as a list of content blocks instead of a plain string.

* fix: flatten list content to strings for OpenAI-compatible providers

Add _patch_openai_compat_content() that wraps _generate/_agenerate to
sanitize message content before API calls. Apply it for all third-party
OpenAI-compat providers and native OpenAI proxies.

This fixes "invalid type: sequence, expected a string" errors from
strict APIs like DeepSeek that reject list-format content in assistant
messages during multi-turn conversations.

* feat: add DeepSeek API key validation and integrate into onboarding process
test: implement unit tests for content flattening utility in OpenAI-compatible providers

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-03-19 19:28:48 +00:00
Jiao Huifeng 28f0d15cda docs: add macOS 24/7 deployment guide for EvoSci + ccproxy + Telegram STT bot (#32)
Comprehensive step-by-step guide covering:
- ccproxy installation from source (patched fork required for Claude 4+)
- ccproxy OAuth login via `ccproxy auth login claude-api`
- EvoScientist install with telegram + stt extras using uv
- STT model pre-download to avoid first-message delay
- launchd plist setup for auto-start on login with KeepAlive
- Full troubleshooting section based on real deployment experience:
  - OAuth token not found (.credentials.json location)
  - Packages installed in wrong Python environment (conda vs .venv)
  - Shell glob eating brackets in pip install 'pkg[extra]'
  - Whisper hallucination / VAD filter
  - ffmpeg approval prompts → --auto-approve
  - heredoc variable expansion gotcha
  - External drive mount timing with launchd

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-03-19 18:45:34 +00:00
X-iZhang ac7fdcecf2 refactor: improve readability of command validation and path extraction logic 2026-03-19 18:38:57 +00:00
X-iZhang edd3823881 feat: implement absolute system path detection in command validation 2026-03-19 18:28:54 +00:00
Jan Piotrowski 8862531738 chore: add pre-commit config 2026-03-19 17:04:06 +01:00
Jan Piotrowski 2fdf961fee chore: fix ruff linting and async patterns in tests/ 2026-03-19 17:04:05 +01:00
Jan Piotrowski 81316ffb2d chore: resolve all ruff linting and static analysis errors in EvoScientist/
- Fix RUF006: Implement background task tracking in Discord, iMessage, WeChat, and TUI to prevent premature GC of fire-and-forget tasks.
- Fix B904: Add explicit exception chaining (raise ... from) across all exception handlers.
- Fix RUF012: Annotate mutable class attributes with ClassVar for command arguments and media maps.
- Fix B008: Refactor Typer commands in cli/commands.py to use Annotated for argument and option defaults.
- Fix B023/B018: Resolve late-binding issues in lambdas and remove useless expressions.
- Fix syntax errors in retry.py docstrings and models.py lambda parameter ordering.
2026-03-19 17:04:03 +01:00
Jan Piotrowski 4a3d6c0318 chore: add ruff lint rules and turn on formatting 2026-03-19 17:04:02 +01:00
X-iZhang 793b3f32af feat: add WeChat QR code to README files for community engagement 2026-03-19 15:32:55 +00:00
Allen 97d00808e1 fix: The input prompt box supports selection using the up and down arrow keys (#76)
* Add the ability to read text from the system clipboard

* The input prompt box supports selection using the up and down arrow keys.
2026-03-19 15:15:01 +00:00
Xi Zhang 47e5ef6219 Feat/minimax anthropic routing (#75)
* feat: update MiniMax integration to use Anthropic-compatible endpoint and enhance routing logic

* refactor: streamline OAuth install hint and update ccproxy health check timeout
2026-03-19 15:01:06 +00:00
Octopus a011dca693 feat: add MiniMax as direct LLM provider (#70)
Add MiniMax (api.minimax.io/v1) as a first-class third-party provider,
enabling direct API access without routing through NVIDIA/SiliconFlow/
OpenRouter intermediaries. Includes M2.5 and M2.5-highspeed models
with 204K context window.

Changes:
- Register "minimax" in _THIRD_PARTY_PROVIDERS with MINIMAX_API_KEY
- Add MiniMax-M2.5 and MiniMax-M2.5-highspeed model entries
- Add minimax_api_key to config, env mappings, and env export
- Add MiniMax to onboarding wizard with API key validation
- Update .env.example, README.md, README.zh-CN.md
- Add 9 unit tests and 3 integration tests (all passing)

Co-authored-by: PR Bot <pr-bot@minimaxi.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-03-19 13:02:40 +00:00
dinos ac5ea37d14 fix(mcp): eliminate duplicate config loading during startup (#69)
* fix(mcp): eliminate duplicate config loading during startup

load_mcp_config() was called twice on every startup — once in
_mcp_config_signature() to compute the cache key, and again inside
load_mcp_tools(). This caused warnings to appear twice.

Merge the two calls into _load_mcp_config_once() which returns both the
signature and the parsed config, then pass the config through to
load_mcp_tools() via a new optional parameter.

* test(mcp): fix existing cache tests and add coverage for single-load guarantee

- Update fake_load_mcp_tools to accept optional config kwarg
- Add test_load_mcp_config_called_once_per_cache_miss: verifies
  load_mcp_config is called exactly once per cache miss (the bug)
- Add test_cached_config_passed_to_load_mcp_tools: verifies the
  pre-loaded config dict is forwarded to load_mcp_tools

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-03-19 12:47:47 +00:00
dinos 7ebaa2cec0 feat: implement command to browse and install MCP servers (#65)
* refactor(mcp): extract MCP server registry from onboard into shared module

Move _RECOMMENDED_MCP_SERVERS, _install_pip_package, and
_pip_install_hint from config/onboard.py into mcp/registry.py as a
shared MCPServerEntry dataclass and registry functions. This enables
reuse by the new /install-mcp command and marketplace integration.

* feat(mcp): add /install-mcp command for browsing and installing MCP servers

Interactive browser for MCP servers (built-in registry + EvoSkills
marketplace), supporting three modes:
- /install-mcp         — interactive tag filter + checkbox selection
- /install-mcp <name>  — direct install by name or tag pre-filter
- /install-mcp file.yaml — import servers from arbitrary YAML file

Also available as /mcp install and EvoSci mcp install. Includes TUI
browser widget (MCPBrowserWidget) mirroring the skill browser UX.

* chore(tavily): conditionally pass tavily_search when TAVILY_API_KEY is set

* fix(tui): Enter key detection for /install-mcp

- Fix message handler names: Textual converts MCPBrowserWidget to
  mcpbrowser_widget (not mcp_browser_widget), so Confirmed/Cancelled
  messages were never received by the app
- Distinguish empty selection from cancel in result handling

* chore(widgets): stop auto-advancing cursor on Space toggle in browser widgets

* style: linter

* refactor(mcp): simplify MCP registry to marketplace-only

Remove built-in server list and arbitrary YAML import — all server
definitions now come from the EvoSkills marketplace (mcp/*.yaml).
Onboarding filters by the `onboarding` tag instead of a hardcoded list.

* refactor(mcp): consolidate /install-mcp into /mcp install

Remove standalone /install-mcp command — use /mcp install as the
single entry point. CLI adapter now delegates logic to the shared
InstallMCPCommand class, keeping only the questionary UI layer.

* chore(cli): rm reference to yaml import
2026-03-19 12:45:32 +00:00
Jan Piotrowski f9734c7e67 feat: extract commands from TUI to commands module
- adjust discord max message length to 2000 (the actual limit)
- /install-skill will fallback to searching EvoSkills repository if path not found to allow for installing skills from channels
- remove discord config tests (it asserted @dataclass behaviour)
2026-03-18 16:59:50 +01:00
z00827015 1c6f4a2078 Add the ability to read text from the system clipboard 2026-03-18 16:42:34 +01:00
akou 2f1622579f fix: improve ccproxy OAuth install hint (#55)
Co-authored-by: Dinos Papakostas <dinospk1999@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-18 09:57:06 +01:00
X-iZhang 16dc1121e4 docs: add tips for copying long outputs in CLI mode to README files 2026-03-18 01:00:18 +00:00
X-iZhang 9274df7405 chore(deps): bump deepagents 0.4.11, update langchain bounds, fix pyasn1 DoS (Dependabot #1) 2026-03-18 00:20:44 +00:00
X-iZhang dfe81ad795 docs: update installation instructions for latest version from GitHub in README files 2026-03-17 23:10:03 +00:00
X-iZhang 5d06d93ac5 docs: add instructions for installing the latest version from GitHub in README files 2026-03-17 23:08:25 +00:00
X-iZhang 7089d1c179 feat: implement _resolve_command function for command path resolution 2026-03-17 23:05:05 +00:00
X-iZhang 53699f9a4e fix: extend ccproxy health check timeout from 10 to 30 seconds 2026-03-17 22:51:02 +00:00
Wiktor Cupiał d5e50a690e fix: add config option for ccproxy port number (#52)
* fix: add config option for ccproxy port number

* feat: add user prompt for ccproxy port configuration and validation
fix: update is_ccproxy_running to use health check endpoint
test: enhance tests for ccproxy port handling and validation

* fix: streamline ccproxy installation process using _install_pip_package

* fix: auto-patch ccproxy adapter for correct OAuth beta header

---------

Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
Co-authored-by: X-iZhang <zacharyzhang2022@gmail.com>
2026-03-17 22:27:20 +00:00
dinos 97684eab7b fix: centralize dotenv loading inside get_effective_config (#45) (#50)
load_dotenv was called in three scattered places: commands.py (main and
serve entry points) and search.py (at module import time). This could
re-override env vars after config resolution.

I moved the single load_dotenv call into get_effective_config() so it
participates in the config priority chain, and removed it from all
other call sites.
2026-03-17 13:13:28 +00:00
dinos 2f855c07c7 feat: add /install-skills command with tag-based skill browsing (#48)
* feat: add tag support to skill metadata parser

- Add `tags` field to `SkillInfo` dataclass
- Extend `_parse_skill_md` to return `SkillInfo` directly (instead of
  dict), extracting tags from top-level `tags` or `metadata.tags`
fallback
- Accept `source` as keyword argument in `_parse_skill_md` to avoid
  post-hoc mutation
- Add `_normalize_tags` helper (handles list, comma-string, missing)
- Add `list_skills_by_tag()` for filtering installed skills by tag
- Add `get_all_tags()` returning tags sorted by count then
alphabetically
- Add `fetch_remote_skill_index()` with shallow-clone and 10-min cache
- Add 12 new tests covering tag parsing, filtering, and remote index

* feat: add /install-skills command and browse action to skill_manager tool

CLI:
- Add `/install-skills` slash command with interactive tag picker and
  skill checkbox (questionary-based, for CLI mode)
- Accepts optional tag argument for pre-filtering: `/install-skills
core`
- Update `/skills` listing to show tags per skill
- Register command in interactive.py dispatch

LangChain tool:
- Add `browse` action to `skill_manager` tool with optional `tag` filter
- Update `list` and `info` actions to include tags in output

* feat: add interactive skill browser widget for TUI

New widget (skill_browser.py):
- Two-phase keyboard-driven widget mounted inline in chat
- Phase 1: tag picker (arrow keys + Enter, Esc to cancel)
- Phase 2: skill checkbox (Space to toggle, Enter to install, Esc back)
- Width-aware description truncation with ellipsis
- Installed skills shown as non-toggleable with checkmark

TUI integration (tui_interactive.py):
- Register /install-skills command with async widget flow
- Echo executed commands in cyan before output
- Add SkillBrowserWidget keyboard delegation (up/down/esc)
- Refocus prompt input after any widget dismissal (also fixes
  pre-existing /resume and /delete focus bug)
- Show tags as bulleted newlines in /skills table
- Dynamic autocomplete padding based on longest item
2026-03-17 12:04:23 +00:00
MuXinCG d3e877c845 chore: add models 2026-03-17 10:42:38 +01:00
Xi Zhang 973edb6aa4 fix: update system prompt to include today's date and refactor related functions (#43) 2026-03-17 02:42:04 +00:00
X-iZhang 859a0aabbb fix: update asset URLs in README and README.zh-CN to use raw GitHub links 2026-03-17 01:04:40 +00:00
X-iZhang 1dac00c262 fix: correct typo in OAuth sign-in description in README.md 2026-03-16 20:26:32 +00:00
Xi Zhang e515f69bdc Fix/onboard oauth ux (#38)
* "fix(onboard): guide ccproxy install and auth in OAuth flow

  - Always show API Key / OAuth choice; prompt to install evoscientist[oauth]
    when ccproxy missing (mirrors iMessage imsg install UX)
  - Add _ccproxy_exe() helper: checks PATH then env bin dir (fixes conda envs
    where shutil.which may not find newly installed binaries)
  - Fix check_ccproxy_auth() false positive: ccproxy auth status exits 0 even
    when not authenticated; detect via output content + filter structlog noise
  - Silent install/login subprocesses; show browser URL as fallback
  - Reset anthropic/openai auth_mode to api_key when switching to non-Anthropic/
    OpenAI provider, preventing stale oauth config from triggering ccproxy error

  Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>"

* fix(onboard): update OAuth support details in README and zh-CN translation
2026-03-16 20:22:20 +00:00
Ziheng Zhang 4b648b3413 chore: add models (#30)
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-03-16 17:40:03 +00:00
Xi Zhang 6408ccb6db feat: add glm-5-turbo model entries for OpenRouter and Zhipu (#37) 2026-03-16 17:30:37 +00:00
dinos 0303c6d43d Add contribution templates and --version CLI flag (#36)
* feat(cli): add --version / -V flag

Uses importlib.metadata to read the version from the installed package.

* chore: improve bug report and feature request issue templates

- Bug report: replace web-app steps with CLI-oriented examples, add
  error output section, add Python version and LLM provider fields
- Feature request: add note directing niche features to EvoSkills,
  set default label

* chore: add documentation issue template and issue chooser config

- Add documentation template for reporting missing or unclear docs
- Add config.yml to disable blank issues and link to EvoSkills and
  Discord as contact options

* chore: add PR template and improve CONTRIBUTING.md

- Add PR template with type-of-change checkboxes, issue linking for
  new features, and CI checklist
- CONTRIBUTING.md: add development setup, PR workflow, and code style
  sections; fix wording; make Discord link clickable
2026-03-16 15:50:19 +00:00
Jan Piotrowski dd81a69585 chore: update CONTRIBUTING.md 2026-03-16 15:40:36 +01:00
Jan Piotrowski 55fb99f0f4 chore: update Issue templates 2026-03-16 15:30:50 +01:00
Jan Piotrowski c542e6403e chore(README): add community discord invite link 2026-03-16 13:52:19 +01:00
X-iZhang 79b572cdae Refactor code structure for improved readability and maintainability 2026-03-16 11:46:41 +00:00
X-iZhang 28ad1f0106 feat: update PyPI badge version to v0.0.2 in README files and SVG assets 2026-03-16 00:45:42 +00:00
X-iZhang 0fe30de92d Refactor code structure for improved readability and maintainability 2026-03-16 00:43:35 +00:00
X-iZhang ccaa3113d2 feat: update provider choices in onboarding wizard to include API/OAuth details 2026-03-16 00:37:25 +00:00
X-iZhang 62fe970476 feat: add OpenAI support with OAuth and environment configuration 2026-03-16 00:25:19 +00:00
X-iZhang ab709d4ec0 feat: update .gitignore to include environment variable files while excluding example 2026-03-15 21:35:03 +00:00
X-iZhang cc29229f35 feat: add community contributors section to README files 2026-03-15 21:31:28 +00:00
X-iZhang c5a4d559a2 Refactor test cases for improved readability and consistency
- Added blank lines for better separation of test cases in multiple test files.
- Reformatted event handling in tests for clarity and consistency.
- Ensured consistent use of multi-line formatting for dictionary arguments in event handling.
- Improved assertions and test descriptions for better understanding.
- Updated test cases across various modules including test_stream_state, test_stream_utils, test_summarization, test_thread_selector, test_tool_error_handler, test_tui_widgets, test_ui_runtime, and test_wechat_channel.
2026-03-15 21:12:51 +00:00
X-iZhang f27521d62b feat: enhance thinking configuration for Anthropic 4-5 models with proxy support 2026-03-15 17:11:00 +00:00
X-iZhang 1ce2d059ca feat: implement OAuth support for Anthropic using ccproxy and add related configuration options 2026-03-15 16:52:11 +00:00
Jiao Huifeng 741eb709b9 feat: support ANTHROPIC_BASE_URL for proxy-based deployments (e.g. ccproxy) (#25)
Allow the Anthropic provider to accept a base_url override via the
ANTHROPIC_BASE_URL environment variable or anthropic_base_url config
field. This enables routing Anthropic API requests through local proxies
like ccproxy without losing provider-specific features (extended
thinking, adaptive effort) that would be dropped when using the generic
"custom" provider.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Xi Zhang <106144707+X-iZhang@users.noreply.github.com>
2026-03-15 14:51:31 +00:00
X-iZhang 174c6044a7 update 2026-03-15 13:26:56 +00:00
X-iZhang b5a3f6de90 update 2026-03-15 10:37:13 +00:00
X-iZhang a792615340 update 2026-03-15 10:30:31 +00:00
X-iZhang 70a5f99c13 update 2026-03-15 10:23:37 +00:00
X-iZhang 7be3a90c09 update 2026-03-15 10:20:15 +00:00
X-iZhang d31ee86e79 update 2026-03-15 10:16:12 +00:00
X-iZhang fb158f2ce7 update 2026-03-15 10:08:41 +00:00
X-iZhang 21a875e026 feat: add /compact command to summarize conversation and free context 2026-03-14 20:09:38 +00:00
X-iZhang 46a1de9dd6 refactor: enhance summarization widget with streaming support and text accumulation 2026-03-14 19:41:27 +00:00
X-iZhang 1bcc38ea16 refactor: enhance pip installation process with improved command handling 2026-03-14 16:56:35 +00:00
Xi Zhang 27f8de5ad2 Refactor code structure for improved readability and maintainability 2026-03-14 16:18:33 +00:00
Ziheng Zhang bfecfe0afc chore: add timeout (#24) 2026-03-14 11:16:59 +00:00
X-iZhang 4b25816ae0 refactor: add optional development dependencies for testing and building 2026-03-14 10:51:01 +00:00
X-iZhang 438bcb82e1 update 2026-03-13 19:59:11 +00:00
X-iZhang 3f23ac3657 update 2026-03-13 19:50:48 +00:00
X-iZhang 4cbb19defc refactor: update recommended skills list by removing duplicates and adding new entries 2026-03-13 16:33:20 +00:00
X-iZhang 0e6fe8532d update 2026-03-13 15:42:57 +00:00
X-iZhang f4bdc85dfc update 2026-03-13 15:39:58 +00:00
X-iZhang a6186a67b8 update 2026-03-13 15:28:36 +00:00
X-iZhang fb92972ce5 update 2026-03-13 14:22:15 +00:00
X-iZhang 534a7cea06 update 2026-03-13 14:15:52 +00:00
X-iZhang 6f927daa89 update 2026-03-13 14:06:20 +00:00
X-iZhang 483ace6df0 update 2026-03-13 14:02:44 +00:00
X-iZhang a635699539 update 2026-03-13 14:00:16 +00:00
X-iZhang 01b94366e9 update 2026-03-13 00:36:00 +00:00
X-iZhang 6f97b1027a Revise README to include updated roadmap section and enhance navigation links 2026-03-13 00:32:16 +00:00
X-iZhang 04bd8c7841 Add framework image and enhance README with scientific workflow details 2026-03-13 00:28:59 +00:00
X-iZhang a93e58bccf Update README files to enhance demo section with video and table layout 2026-03-13 00:03:55 +00:00
X-iZhang 17f75d4a32 Add citation section to README files 2026-03-12 23:18:35 +00:00
X-iZhang ef80273480 Update onboarding skills label and add new NVIDIA models 2026-03-12 23:17:11 +00:00
X-iZhang 1ccdbbea90 update 2026-03-12 23:05:09 +00:00
X-iZhang 017f169015 update 2026-03-12 23:02:59 +00:00
X-iZhang 72b422d3a2 update 2026-03-12 23:00:39 +00:00
X-iZhang f7b5cc2a1c update 2026-03-12 22:59:12 +00:00
X-iZhang 8060ded803 update 2026-03-12 22:57:19 +00:00
X-iZhang 6034f07ea5 update 2026-03-12 22:50:42 +00:00
X-iZhang ee20adb021 Remove unused function _channel_ask_user_prompt_from_event from tui_interactive.py 2026-03-12 12:13:01 +00:00
Ziheng Zhang 438f9e56d5 Fix: channel can't ask user (#23)
* small fix

* small fix

* remove local file

* remove local file
2026-03-12 11:28:37 +00:00
X-iZhang 9bbde69f7f update 2026-03-12 03:12:06 +00:00
X-iZhang 1cbf6a8cd1 update 2026-03-12 02:52:30 +00:00
X-iZhang d226aabe01 update GitHub Actions workflows to use latest action versions 2026-03-12 02:47:51 +00:00
X-iZhang 5ab677c40d add demo video to README files 2026-03-12 02:39:51 +00:00
X-iZhang 1ce8837122 update 2026-03-12 01:36:29 +00:00
X-iZhang 774473eb30 update 2026-03-12 01:34:28 +00:00
X-iZhang cf225de644 update 2026-03-12 01:25:49 +00:00
X-iZhang 20c76cd7cf update 2026-03-12 01:21:03 +00:00
X-iZhang 5fb7094e6f update 2026-03-12 00:47:51 +00:00
X-iZhang fb5babbcc8 update 2026-03-12 00:46:33 +00:00
X-iZhang 9ad215b257 update 2026-03-12 00:44:13 +00:00
X-iZhang ad2b1c9f97 update 2026-03-12 00:43:13 +00:00
X-iZhang 8d7dfaafd0 update 2026-03-12 00:42:23 +00:00
X-iZhang 6eceb045ab update 2026-03-12 00:40:32 +00:00
X-iZhang 253abe8445 update 2026-03-12 00:22:08 +00:00
X-iZhang 0d4f736eae update 2026-03-12 00:15:27 +00:00
X-iZhang 1d91c028db update 2026-03-12 00:11:05 +00:00
X-iZhang 83845e2655 update 2026-03-11 23:28:18 +00:00
Xi Zhang 1db62de1bb Update LICENSE 2026-03-11 11:09:44 +00:00
X-iZhang 7784afa12b feat: add user prompt functionality to agent for clarification questions in README files 2026-03-11 00:34:37 +00:00
X-iZhang 28d8b334c7 fix: update contributor information and citation in README.zh-CN.md 2026-03-11 00:12:15 +00:00
X-iZhang 5ff000a9b2 feat: enhance timeout handling with recovery guidance and resource estimation in middleware 2026-03-11 00:03:18 +00:00
X-iZhang 1625caa194 fix: update dependencies to latest versions 2026-03-10 22:17:14 +00:00
X-iZhang c15c0dfdf0 feat: add ask_user middleware and interactive widget for user prompts
- Implemented `ask_user` middleware to facilitate agent-initiated questions during research workflows.
- Created `AskUserWidget` for interactive user prompts, supporting both text and multiple choice questions.
- Enhanced `StreamEventEmitter` to handle `ask_user` interrupts and updated event handling in `stream_agent_events`.
- Added state management for pending `ask_user` events in `StreamState`.
- Developed validation for question structures and parsing for responses.
- Introduced unit tests covering middleware functionality, event handling, and widget behavior.
2026-03-10 18:17:52 +00:00
Yougang Lyu bc7a0b753c Update README.md 2026-03-10 17:43:52 +01:00
Yougang Lyu 30dde5cf83 Update README.md 2026-03-10 17:43:16 +01:00
Yougang Lyu 9c56c73838 Update README.md 2026-03-10 17:42:20 +01:00
Yougang Lyu ae1da96fb6 Update README.md 2026-03-10 17:39:52 +01:00
Yougang Lyu b54acbde41 Update README.md 2026-03-10 11:51:19 +01:00
Yougang Lyu 4a06afd0f1 Update README.md 2026-03-10 10:24:39 +01:00
Yougang Lyu f41a9a6114 Update README.md 2026-03-10 10:23:33 +01:00
Yougang Lyu d3e7e22bb8 Update README.md 2026-03-10 08:57:52 +01:00
X-iZhang 311fb08676 feat(summarization): implement summarization feature with dedicated widget and event handling 2026-03-10 01:13:30 +00:00
X-iZhang 903ef4d158 feat(thread-picker): add ThreadPickerWidget for inline thread selection in TUI 2026-03-10 00:45:21 +00:00
X-iZhang aad431785c feat(diff): implement unified diff formatting for edit_file tool results 2026-03-09 22:18:54 +00:00
X-iZhang 4c53bb0480 fix(cli): update welcome slogans for clarity and engagement 2026-03-09 17:30:20 +00:00
X-iZhang 6817128a03 fix(models): remove deprecated gemini-3-pro model entry from registry 2026-03-09 16:46:09 +00:00
X-iZhang 18b1d38aa5 fix(onboard): update wechat dependencies and improve package display formatting
fix(cli): unify directory handling in print_banner and _build_welcome_banner functions
2026-03-09 15:01:26 +00:00
X-iZhang c4de31c772 fix(ui): map legacy UI backend values to current equivalents in normalization 2026-03-09 10:38:42 +00:00
Ziheng Zhang 68f432b39d Fix/onboard channel auto install (#22)
* small fix

* small fix

* small fix
2026-03-09 10:33:10 +00:00
Xiaohui Yan ce33de79af fix(onboard): map legacy 'textual' value to 'tui' for UI backend default (#21)
Add mapping for legacy 'textual' value to current 'tui' value when
selecting UI backend, preventing invalid default selection.
2026-03-09 10:19:47 +00:00
X-iZhang 3885197559 docs: update CONTRIBUTING.md with test count and headless mode instructions 2026-03-08 17:18:19 +00:00
X-iZhang f3bb70c74e feat(serve): implement headless serve mode with message processing and interrupt handling 2026-03-08 17:06:40 +00:00
X-iZhang 260ee95697 feat(ui): update UI backend references from 'rich' to 'tui' in onboarding and interactive components 2026-03-08 16:44:28 +00:00
X-iZhang a118d3ed23 refactor(onboard): simplify workspace step to return only selected mode
test(onboard): update tests for workspace step to reflect mode-only return
2026-03-08 16:16:40 +00:00
X-iZhang 07e57e110b feat(ui): update UI backend options to use 'cli' and 'tui' 2026-03-08 15:34:22 +00:00
X-iZhang 659fbc7f84 feat(interactive): implement history-based auto-suggest for TUI input 2026-03-08 15:07:17 +00:00
X-iZhang a91d6def40 fix(interactive): remove unused os import from interactive.py 2026-03-08 14:40:57 +00:00
X-iZhang ed376d1b5b fix(interactive): update history file path to use config directory 2026-03-08 14:38:11 +00:00
Yougang Lyu 085cef4463 Update README.md 2026-03-08 14:25:46 +01:00
X-iZhang 5ea9f3e938 feat(serve): add auto-approve option for tool executions
docs: update README to include HITL action approval details
tests: adjust mock for get_effective_config in CLI serve tests
2026-03-07 21:13:40 +00:00
X-iZhang 833fa0abfb Implement HITL (Human-in-the-Loop) approval mechanism
- Added HITL approval lifecycle to ToolCallWidget, including a new "rejected" status.
- Introduced configuration options for auto-approval and shell command allow list in EvoScientistConfig.
- Developed approval handling functions in display module to manage HITL interrupts and user decisions.
- Created ApprovalWidget for user interaction during approval prompts.
- Enhanced StreamEventEmitter to support interrupt events.
- Updated StreamState to track pending interrupts.
- Implemented tests for HITL functionality, including event structure, state handling, and approval logic.
2026-03-07 21:02:35 +00:00
X-iZhang bdd3e30284 feat(ollama): enable reasoning by default for Ollama models and update tests 2026-03-07 18:44:11 +00:00
X-iZhang 250c41c561 feat(token-usage): implement token usage tracking and display widget 2026-03-07 15:16:29 +00:00
X-iZhang cdb5507896 fix(dependencies): update deepagents version to 0.4.7 for compatibility 2026-03-07 12:49:07 +00:00
X-iZhang 6992e37ce5 feat(tui-interactive): update color scheme for improved visibility and aesthetics 2026-03-07 12:09:02 +00:00
X-iZhang 0c5d0049a4 feat(tui-interactive): implement throttled scrolling for long content and improve Markdown rendering efficiency 2026-03-07 01:28:44 +00:00
X-iZhang d2555ac10f feat(tui-interactive): implement message queueing and editing functionality 2026-03-07 01:09:52 +00:00
X-iZhang 0a823c6e0d feat(thinking-widget): enhance collapsible behavior and improve display logic 2026-03-07 00:10:53 +00:00
X-iZhang e7c06f8ddb feat(display): enhance streaming display with final frame rendering and improved thinking text handling 2026-03-06 23:50:04 +00:00
X-iZhang d69ba4ec8b feat(skill-creator): enhance HTML documentation and improve benchmark packaging validation 2026-03-06 23:29:37 +00:00
X-iZhang 5e1235627d feat(skill-creator): enhance documentation and validation scripts
- Updated SKILL.md for clarity on headless environments
- Adjusted runs_per_configuration logic in aggregate_benchmark.py
- Improved description placeholder in init_skill.py
- Added strict validation checks in quick_validate.py
- Ensured skill-creator root is included in sys.path for script imports
2026-03-06 23:12:37 +00:00
X-iZhang 0897f3a2e1 Update 2026-03-06 22:31:12 +00:00
X-iZhang 8c77ece5e5 docs(CONTRIBUTING): update test count and add details for skill-creator licensing 2026-03-06 21:35:49 +00:00
X-iZhang eeb0a21275 feat(license): add Apache License 2.0 to skill-creator directory
style(eval_review): update button colors and instructions for clarity

style(viewer): modify accent colors and update instructions for agent

docs(output-patterns): replace "Claude" with "agent" for consistency

docs(workflows): replace "Claude" with "agent" for consistency

fix(improve_description): update skill description context to refer to EvoScientist

fix(init_skill): update references to "Claude" to "agent" for consistency
2026-03-06 15:03:36 +00:00
X-iZhang 13858b860a feat(models): add zhipu and zhipu-code providers to models registry tests 2026-03-06 14:19:00 +00:00
Xiaohui Yan b14dceb143 feat(models): add glm-5, glm-4.7 from zhipu API (with code plan supoort) (#19)
* feat(models): add glm-5, glm-4.7 from zhipu API (with code plan supoort)

* fix: add zhipu and zhipu-code to valid providers in test
2026-03-06 08:57:33 +00:00
Yougang Lyu 5f2d963a97 Update README.md 2026-03-06 06:00:02 +01:00
X-iZhang c62c7d24e2 fix(run_eval): remove reasoning kwarg for OpenAI provider to allow default model behavior 2026-03-06 02:36:16 +00:00
X-iZhang 67a60b9ca1 feat(models): add gpt-5.4 model entry to the OpenAI provider 2026-03-06 02:08:55 +00:00
X-iZhang b05576cdbc feat: Add skill description improvement and evaluation scripts
- Implemented `improve_description.py` to enhance skill descriptions based on evaluation results using LLMs.
- Created `run_eval.py` to evaluate skill descriptions against a set of queries, determining trigger effectiveness.
- Developed `run_loop.py` to automate the evaluation and improvement process, tracking history and optimizing descriptions iteratively.
- Added utility functions in `utils.py` for parsing skill metadata from SKILL.md files.
- Enhanced `package_skill.py` to exclude unnecessary files and directories during skill packaging.
- Updated `quick_validate.py` to include compatibility checks in skill validation.
2026-03-06 00:38:07 +00:00
X-iZhang 313e48eb62 fix: update default model to claude-sonnet-4-6 and increase recursion limit to 1000 2026-03-04 20:23:02 +00:00
X-iZhang 3be3f5adf0 feat(subagent): implement visibility management for completed and running tools in SubAgentWidget 2026-03-04 19:26:08 +00:00
Jan Piotrowski bf4e297cbb fix(tui): Disable auto refresh during generation when using Rich
In some terminals there is heavy flicker when auto refresh was being used. No cause could be found
2026-03-04 16:21:28 +01:00
X-iZhang 5fa7c684bd feat(models): add support for gemini-3.1-flash-lite model in Google GenAI provider 2026-03-04 11:21:12 +00:00
X-iZhang c96d984869 feat(workflows): add cache-dependency-glob for improved dependency caching in build, lint, and test workflows 2026-03-03 01:20:06 +00:00
X-iZhang 32bbd1a1af feat(middleware): add structured output handling for OpenAI models in EvoMemoryMiddleware 2026-03-03 01:14:12 +00:00
X-iZhang a574492123 refactor(prompts, utils, subagent, skill_manager, think): update skill usage instructions and enhance reflection tool for better decision-making 2026-03-03 01:05:10 +00:00
X-iZhang 7e65bfe0b7 feat(models): expand support for additional providers and enhance auto-configuration logic 2026-03-01 00:09:44 +00:00
X-iZhang eaae62cf26 fix(docs): correct test file count and update pip installation instructions in CONTRIBUTING.md 2026-02-28 23:58:56 +00:00
X-iZhang 8ebc7dafc4 fix(docs): update website link in Chinese README to point to GitHub Pages 2026-02-28 17:27:05 +00:00
X-iZhang f286091607 fix(docs): update website link in README to point to GitHub Pages 2026-02-28 17:20:21 +00:00
X-iZhang 845fdd8710 feat(docs): enhance contribution guidelines for AI agents and update README files 2026-02-28 17:04:04 +00:00
X-iZhang 540f986c3c fix(docs): update contributor superscripts in README files 2026-02-28 16:18:35 +00:00
Yougang Lyu 78de46455f Update README.md 2026-02-28 11:45:30 +01:00
Yougang Lyu feb39bfbd9 Update README.md 2026-02-28 11:43:11 +01:00
X-iZhang 3251b8985b fix(deps): update deepagents dependency version to 0.4.4
refactor(backends): reorganize imports for clarity
2026-02-28 00:06:27 +00:00
Xi Zhang f2723ad674 Rename README.zh-cn.md to README.zh-CN.md 2026-02-27 23:25:56 +00:00
X-iZhang 2f0108d7f2 fix(docs): update installation instructions and improve formatting in README files 2026-02-27 23:24:24 +00:00
X-iZhang 409a1bb7d4 feat(docs): update roadmap section in English and Chinese README files 2026-02-27 22:12:11 +00:00
X-iZhang e793bc0797 fix(docs): correct links and update collaborator section in README files 2026-02-27 19:04:51 +00:00
X-iZhang 2f6cf824f2 refactor(config): remove max_concurrent and max_iterations from config and onboarding steps
feat(prompts): update system prompt to eliminate numeric limits for sub-agents and delegation rounds
test(tests): adjust tests to reflect changes in configuration and onboarding logic
2026-02-27 18:57:47 +00:00
X-iZhang 1130ddd8e9 fix(logging): reorder logging configuration for suppressing warnings 2026-02-27 18:41:37 +00:00
X-iZhang 3f42259441 feat(skills): enhance skill management with directory scanning and update skill labels 2026-02-27 18:37:57 +00:00
X-iZhang d81d41f756 feat(models): add new models for SiliconFlow and OpenRouter providers
test(models): update tests to validate new providers and registered models
2026-02-27 14:38:22 +00:00
X-iZhang c9e5d925b8 update 2026-02-27 13:04:35 +00:00
X-iZhang ebaaff6083 update readme 2026-02-27 13:00:37 +00:00
X-iZhang 2431d7e926 update 2026-02-27 02:21:53 +00:00
X-iZhang 08507f5d14 update 2026-02-27 02:18:11 +00:00
X-iZhang 200c5e0a12 feat(docs): add Chinese translation of README with visual elements and links 2026-02-27 02:17:08 +00:00
X-iZhang 29881d7392 docs(README): update language links, enhance project description, and revise team roles 2026-02-27 02:11:10 +00:00
X-iZhang 931fc8a88d feat(skills): implement batch installation of multiple skills and enhance install feedback 2026-02-25 23:38:08 +00:00
X-iZhang 35135ee79c feat(tests): add additional smoke tests for email, signal, and QQ channels
feat(tests): enhance relative time formatting tests for minutes, hours, days, and months
chore(docs): update contributing guidelines for testing commands
2026-02-25 10:05:29 +00:00
X-iZhang abf783d2d3 feat(cli): enhance agent loading with optional config parameter for improved initialization 2026-02-25 00:35:40 +00:00
X-iZhang 7a6488d51d feat(models): add new model entries for Anthropic and OpenAI, enhance reasoning and thinking parameters 2026-02-24 22:45:08 +00:00
X-iZhang ea00578cee feat(onboarding): add Ollama provider support with connection validation 2026-02-24 19:45:02 +00:00
Xi Zhang 5705720d0e Merge pull request #18 from EvoScientist/docs/uv
Use uv for dependency management
2026-02-24 12:19:32 +00:00
Dinos Papakostas 6faad4bb98 chore(workflows): update gh workflows to use uv 2026-02-24 20:06:19 +08:00
Dinos Papakostas fad6bde5a6 docs(README): add info about optional channel deps 2026-02-24 19:58:28 +08:00
Dinos Papakostas d39c82989d chore(pyproject): modernize pyproject - use dep groups for dev 2026-02-24 19:57:11 +08:00
Dinos Papakostas 5260b17c0b docs(README): add uv 2026-02-24 19:57:11 +08:00
X-iZhang 6c07e9dda7 update 2026-02-24 11:24:10 +00:00
X-iZhang f13d43a362 update 2026-02-24 11:19:10 +00:00
X-iZhang ef46411796 update 2026-02-24 03:32:58 +00:00
X-iZhang 317dfa1f63 update 2026-02-24 03:32:08 +00:00
X-iZhang 0eb6254083 update 2026-02-24 03:29:15 +00:00
X-iZhang bb7a4c12e9 update 2026-02-24 03:25:02 +00:00
X-iZhang 53f6761dec update 2026-02-24 03:16:31 +00:00
X-iZhang ae15c569c7 update 2026-02-22 19:54:46 +00:00
X-iZhang 9c6f2838ad update 2026-02-22 19:54:05 +00:00
X-iZhang 2da1d6944b update 2026-02-22 19:52:40 +00:00
X-iZhang 1ccff57acd update 2026-02-22 19:48:55 +00:00
X-iZhang 4b69903a94 update 2026-02-22 19:47:21 +00:00
X-iZhang 709f212289 update 2026-02-22 19:45:31 +00:00
X-iZhang b72db34a08 update readme 2026-02-22 19:38:59 +00:00
X-iZhang c9b4bf0367 Refactor code structure for improved readability and maintainability 2026-02-22 18:01:04 +00:00
X-iZhang 9b088aa0ed update 2026-02-22 17:58:03 +00:00
X-iZhang c88e9ccc08 update 2026-02-22 17:54:25 +00:00
X-iZhang 65e95ac4b3 update 2026-02-22 17:49:08 +00:00
X-iZhang a1c44cf95d update 2026-02-22 17:46:17 +00:00
X-iZhang 11b14d01a8 update 2026-02-22 17:42:19 +00:00
X-iZhang 11587250ae Refactor code structure for improved readability and maintainability 2026-02-22 17:39:57 +00:00
X-iZhang 5523a3c4a3 Refactor code structure for improved readability and maintainability 2026-02-22 17:37:09 +00:00
X-iZhang f3cf1dc896 Refactor code structure for improved readability and maintainability 2026-02-22 17:24:45 +00:00
X-iZhang c57e6132a5 Add tests for Rich markup escape safety in formatter
- Introduced a new test file `test_rich_escape.py` to validate the safety of Rich markup escaping in the ToolResultFormatter.
- Added tests to ensure that tool names and error messages containing brackets do not cause crashes during formatting.
2026-02-22 17:15:11 +00:00
X-iZhang 7dedf2b426 fix: improve cancellation handling in middleware and mixins for Python 3.12+ 2026-02-22 16:20:55 +00:00
Xi Zhang a2048e4127 Merge pull request #17 from EvoScientist/bug-fix
Bug fix
2026-02-22 14:32:07 +00:00
MuXinCG 663cc601d5 fix README 2026-02-21 18:42:55 +08:00
MuXinCG f211ecb7d7 bug fix 2026-02-21 17:18:23 +08:00
MuXinCG 258fbc4d34 some bug fix 2026-02-21 17:18:10 +08:00
X-iZhang 0e4d09f5d4 fix(tui): rename thread_id to conversation_tid for clarity in session management 2026-02-21 03:54:43 +00:00
X-iZhang bd167d6c85 feat(cli): add random welcome slogans to interactive and TUI modes 2026-02-21 03:17:29 +00:00
X-iZhang b5dbdc2a1e feat(widgets): enhance ToolCallWidget and SubAgentWidget for better tool management and status handling 2026-02-21 02:50:28 +00:00
X-iZhang d62c0cfb2e feat(clipboard): add clipboard utilities for copy-on-select functionality 2026-02-21 01:39:32 +00:00
X-iZhang 03f4134190 feat: add UserMessage widget and UI backend selection during onboarding
- Introduced UserMessage widget for displaying user input with a styled prompt.
- Updated onboarding steps to include UI backend selection (Rich CLI or Textual TUI).
- Modified EvoScientistConfig to store selected UI backend.
- Enhanced configuration handling to support UI backend environment variable.
- Updated README with new UI backend options and commands.
- Added tests for new UI backend functionality and UserMessage widget.
- Removed obsolete test files and ensured existing tests are updated accordingly.
2026-02-21 01:18:33 +00:00
X-iZhang 75ee7aeba9 feat(backends): enhance convert_virtual_paths_in_command to handle workspace paths 2026-02-20 23:07:28 +00:00
X-iZhang 6399855df7 fix(middleware): improve model binding in EvoMemoryMiddleware for reliable extraction 2026-02-20 22:11:21 +00:00
X-iZhang 4675d311fe fix(docs): update MCP integration guide link in README 2026-02-20 15:49:42 +00:00
X-iZhang ce4497fab5 feat(docs): update README for MCP integration and replace CLI help image 2026-02-20 15:47:39 +00:00
Xi Zhang abb237115f Merge pull request #16 from EvoScientist/channel-fix
fix(channels): critical async, cleanup, and performance bugs
2026-02-20 12:20:02 +00:00
MuXinCG 86a288b3a1 some small bug fix 2026-02-20 19:40:42 +08:00
MuXinCG 13d6852857 Merge remote-tracking branch 'origin/main' into channel-fix 2026-02-20 19:13:43 +08:00
X-iZhang e41e04a176 feat(models): add new Gemini model entries for Google GenAI 2026-02-19 22:11:16 +00:00
MuXinCG e49bac3a33 some small bug fix 2026-02-19 20:34:59 +08:00
MuXinCG 8a00cd3508 some small bug fix 2026-02-19 20:13:02 +08:00
X-iZhang 2dd331be02 feat(tests): refactor async coroutine handling with shared run_async function 2026-02-18 16:53:58 +00:00
Xi Zhang 145b31dc2c Merge pull request #15 from EvoScientist/feature/channel-email-qq-signal
Feature/channel email qq signal
2026-02-18 15:43:09 +00:00
MuXinCG db6e309fdd Add email qq and signal 2026-02-18 22:44:34 +08:00
MuXinCG 98e133564f Add email qq and signal 2026-02-18 21:28:50 +08:00
X-iZhang d09ae76a82 feat(assets): add new channel and MCP images 2026-02-18 01:48:35 +00:00
X-iZhang e2a5241a68 feat(cli): enhance channel thinking propagation and error handling in bus mode 2026-02-17 20:25:49 +00:00
Xi Zhang 8873b83083 Merge pull request #14 from EvoScientist/feat/tool-error-handler-subagents
feat(evosci): inject tool error handler middelware to subagents
2026-02-17 13:32:51 +00:00
Xi Zhang a786e181bf Merge pull request #13 from EvoScientist/feature/channel-dingtalk-feishu
Add Dingtalk Feishu
2026-02-17 13:31:40 +00:00
Dinos Papakostas 23c0a47455 feat(evosci): inject tool error handler middelware to subagents 2026-02-17 19:46:22 +08:00
MuXinCG 5fd8e1a8ad fix feishu bugs 2026-02-17 16:08:41 +08:00
MuXinCG bec3fc0b67 pass linter 2026-02-17 15:40:57 +08:00
MuXinCG b85faf7516 Add Dingtalk Feishu 2026-02-17 10:54:15 +08:00
X-iZhang 72e4353f54 feat(onboard): add header for channel settings in onboarding wizard 2026-02-16 14:26:09 +00:00
X-iZhang 65d53b4c24 fix(onboard): correct pip install command formatting for missing packages 2026-02-16 14:10:48 +00:00
X-iZhang 1446267360 feat(tests): mock import for telegram and discord in TestStepChannels 2026-02-16 13:52:01 +00:00
X-iZhang 283ae49a91 feat(onboard): enhance channel configuration with import checks and pip dependencies 2026-02-16 13:44:20 +00:00
X-iZhang b62e63ac74 feat(channel): improve error handling and logging in channel management
feat(interactive): refactor message sending to channel for better error handling
fix(display): remove unnecessary print statement in final results display
test(bus): update tests to match changes in _bus_inbound_consumer signature
2026-02-16 13:33:00 +00:00
Xi Zhang 1cc7c539be Merge pull request #12 from EvoScientist/feature/channel-slack-wechat
Feature/channel slack wechat
2026-02-16 11:58:33 +00:00
MuXinCG f39556afa0 pass pytest 2026-02-16 19:54:39 +08:00
X-iZhang 4fc7a035a2 feat(stream): track sent media paths to avoid duplicates in file handling 2026-02-16 11:51:42 +00:00
MuXinCG 2ee50e3c4f Add wait sing 2026-02-16 19:50:58 +08:00
Xi Zhang 1fc823b182 Merge pull request #11 from EvoScientist/feat/tool-error-handling
Fix tool execution errors crashing the agent loop
2026-02-16 11:22:53 +00:00
Dinos Papakostas 764e4a4cc8 feat(tools): add error handler middleware to prevent tool failures crashing the agent 2026-02-16 17:13:04 +08:00
MuXinCG 8ffc294ba1 add me 2026-02-16 16:56:59 +08:00
MuXinCG 12bc82226f add slack and wechat 2026-02-16 16:05:41 +08:00
X-iZhang 924ce3dd30 feat: implement channel message queue and response handling in bus mode 2026-02-16 03:01:19 +00:00
Xi Zhang 4daf290fee Merge pull request #10 from EvoScientist/feature/channel-telegram-discord
feat(channels): add unified channel framework with Telegram and Disco…
2026-02-16 01:13:24 +00:00
MuXinCG 8a2b6161f1 fix bugs 2026-02-16 02:15:28 +08:00
MuXinCG 616a90ef1c feat(channels): add unified channel framework with Telegram and Discord support
Introduce the new channel unification architecture:
- Core framework: base channel class, bus, consumer, channel manager,
  middleware, mixins, capabilities, formatter, retry, plugin system
- Telegram channel implementation with bot token validation
- Discord channel implementation with bot token validation
- Refactored iMessage to use new base channel architecture
- Updated CLI: /channel commands, serve mode, channel setup wizard
- Updated onboard wizard to support multi-channel selection
- Config settings for all channel types
- Stream events: media attachment support, done event content field
- Comprehensive test coverage for channels, bus, and manager
2026-02-16 02:09:25 +08:00
X-iZhang 3da590acaf refactor: rename imessage_send_thinking to channel_send_thinking for clarity 2026-02-15 17:31:02 +00:00
X-iZhang eb10135314 Revert "Merge pull request #8 from EvoScientist/feature/channel-unification"
This reverts commit 84da81256a, reversing
changes made to c2c8e27b46.
2026-02-15 16:32:44 +00:00
Xi Zhang 84da81256a Merge pull request #8 from EvoScientist/feature/channel-unification
Feature/channel unificationfeat: unified channel architecture with multi-platform support
2026-02-15 13:20:55 +00:00
MuXinCG 9b93cea12b fix linter bug 2026-02-15 21:07:47 +08:00
MuXinCG 91e0cfb7c1 pass linter and pytest 2026-02-15 20:57:21 +08:00
MuXinCG 358f63994b Merge remote-tracking branch 'origin/main' into feature/channel-unification
# Conflicts:
#	EvoScientist/EvoScientist.py
#	EvoScientist/channels/imessage/serve.py
#	EvoScientist/cli/commands.py
2026-02-15 20:37:12 +08:00
X-iZhang c2c8e27b46 feat(mcp): implement caching for MCP tools to optimize loading
feat(onboard): enhance API key validation logic to retain existing keys
test(mcp): add tests for MCP tool caching behavior
2026-02-15 12:11:10 +00:00
X-iZhang 3758ae1c1d feat(workspace): set default workspace root and ensure directories in CLI agent 2026-02-15 11:46:30 +00:00
X-iZhang 7b598dd6a6 feat(paths): implement set_workspace_root function and update related paths
refactor(cli): improve workspace handling in commands and agent modules
fix(interactive): update memory directory path resolution
docs(README): clarify workspace directory structure and environment variable usage
test(paths): add tests for set_workspace_root and ensure_dirs functions
test(skills_manager): update tests to use paths module for USER_SKILLS_DIR
2026-02-15 11:27:04 +00:00
MuXinCG 1ccc4158d9 Merge remote-tracking branch 'origin/main' into feature/channel-unification
# Conflicts:
#	EvoScientist/cli/commands.py
2026-02-15 18:10:24 +08:00
MuXinCG 01018c0d24 stable 2026-02-15 18:08:57 +08:00
MuXinCG 919aa15c5e stable 2026-02-15 18:07:41 +08:00
MuXinCG 68181fae49 reduce code complex 2026-02-15 18:01:53 +08:00
MuXinCG e672f71c12 Stable 2026-02-15 17:21:29 +08:00
MuXinCG 7a94951af9 Add me 2026-02-15 14:21:34 +08:00
MuXinCG 2f851cdd7b wechat fail 2026-02-15 14:01:41 +08:00
MuXinCG f4419dae80 wechat failed 2026-02-15 14:01:08 +08:00
MuXinCG 299054edbf test wechat 2026-02-15 13:35:06 +08:00
MuXinCG 123a5de179 test weichat 2026-02-15 13:33:59 +08:00
X-iZhang 8893906e94 update 2026-02-15 03:34:43 +00:00
X-iZhang 94184fe13b feat(workspace): defer directory creation to CLI and improve workspace handling 2026-02-15 01:10:48 +00:00
X-iZhang f80246c48f Revert "add sci-swarm skill"
This reverts commit d591fda338.
2026-02-14 20:31:16 +00:00
MuXinCG 61cf79e525 for merge 2026-02-15 02:11:38 +08:00
X-iZhang d591fda338 add sci-swarm skill 2026-02-13 20:53:53 +00:00
X-iZhang 8c67abf7ef feat(memory): add structured extraction schemas for user profiles, research preferences, and experiment conclusions 2026-02-13 19:34:07 +00:00
X-iZhang 7572fef50a feat(middleware): reorganize memory middleware structure and add memory management functionality 2026-02-13 18:54:44 +00:00
X-iZhang 418f2d58a7 Enhance documentation and refactor code for clarity and functionality
- Updated README.md to clarify tool allowlist supports glob wildcards.
- Enhanced docstrings in prompts.py, utils.py, and formatter.py for better understanding of function parameters and return values.
- Refactored test cases in test_llm.py and test_onboard.py to use consistent mocking style with unittest.mock.patch.
- Improved test coverage and clarity in test_skills_manager.py and test_stream_state.py by adding descriptive comments and organizing sections.
- Adjusted tool result formatting logic in formatter.py to streamline success checks.
- Ensured backward compatibility in various modules while enhancing functionality.
2026-02-13 18:54:33 +00:00
Xi Zhang b50299534f Merge pull request #7 from EvoScientist/feat/mcp-wildcards
Add wildcard support for MCP tool filtering
2026-02-13 17:32:31 +00:00
X-iZhang ccdf2157cb feat(middleware): remove skills middleware and update exports 2026-02-13 17:29:07 +00:00
Dinos Papakostas 97ac3e0401 feat(mcp): support wildcards in tool filtering 2026-02-14 00:06:08 +08:00
X-iZhang 2265fb70b7 feat(interactive): update thread limit to 0 for session commands and enhance session picker display 2026-02-13 00:29:38 +00:00
X-iZhang a7f1e166d7 feat: add support for custom API key and base URL in configuration 2026-02-13 00:12:47 +00:00
X-iZhang 8210192b03 feat: add OpenRouter API key support and enhance model handling 2026-02-13 00:03:36 +00:00
X-iZhang 349693f9d6 feat: add support for SiliconFlow API key and update model providers 2026-02-12 19:51:09 +00:00
X-iZhang 5e83f4caba feat(models): add support for new OpenAI models and enhance test coverage for providers 2026-02-12 18:30:32 +00:00
X-iZhang 5271931935 feat(tests): add tests for _merge_memory backslash safety and _step_mcp_servers npx handling 2026-02-12 18:30:27 +00:00
X-iZhang a61333d4c5 fix: correct formatting in docstrings and enhance checkbox rendering for installed items 2026-02-12 15:47:13 +00:00
Xi Zhang 8118d20dfd Merge pull request #6 from EvoScientist/feat/detect-mcp
feat(onboard): disable already configured mcp servers
2026-02-12 14:36:24 +00:00
Dinos Papakostas 29be24efe6 feat(onboard): disable already configured mcp servers 2026-02-12 17:09:18 +08:00
X-iZhang 56711aa321 feat: enhance session management to isolate agent data and prevent cross-agent visibility 2026-02-12 00:47:28 +00:00
X-iZhang 10266c660f feat: lazy load default agent and enhance CLI agent creation with optional checkpointer 2026-02-11 23:34:38 +00:00
X-iZhang f29b254d9a Implement session persistence and management features
- Introduced a new `sessions.py` module for handling session persistence using SQLite.
- Added CRUD operations for threads, including listing, checking existence, finding similar threads, and deleting threads.
- Enhanced the interactive CLI to support commands for managing sessions: `/current`, `/threads`, `/resume`, and `/delete`.
- Updated the `cmd_interactive` function to handle session metadata and improve user experience with session history rendering.
- Modified the streaming functions to include metadata for checkpoint persistence.
- Added unit tests for session management functionalities to ensure reliability and correctness.
2026-02-11 23:34:25 +00:00
X-iZhang 63517a2a04 feat: refactor iMessage channel handling and remove view_image tool
- Updated iMessage channel to read send_thinking preference from config.
- Modified _auto_start_channel to accept send_thinking parameter.
- Removed view_image tool and adjusted related functionality.
- Enhanced README to reflect changes in media handling.
- Updated tests to cover new behavior and removed tests for view_image.
2026-02-11 19:27:05 +00:00
X-iZhang 7cf55e2aa1 fix: trim whitespace from thinking text in display functions 2026-02-11 19:26:47 +00:00
X-iZhang b0644a2b53 feat(backends): enhance CustomSandboxBackend with LocalShellBackend and improved command validation
fix(utils): add skills support in load_subagents function

chore: update dependencies to latest versions

test: remove unused working_dir parameter in CustomSandboxBackend tests
2026-02-10 23:56:29 +00:00
X-iZhang 8404a8e013 fix: update allowed_senders handling and improve logging configuration 2026-02-10 22:20:56 +00:00
X-iZhang 100b74dd63 fix: remove unused import from agent module 2026-02-10 20:09:48 +00:00
X-iZhang e987ec1e61 Add interactive CLI mode and skill management commands
- Implemented interactive CLI in `interactive.py` for real-time user interaction with slash commands.
- Added MCP server management functionality in `mcp_ui.py`, including listing, adding, editing, and removing servers.
- Created skill management commands in `skills_cmd.py` for listing, installing, and uninstalling user skills.
- Enhanced user experience with rich text output and command completion features.
2026-02-10 19:52:15 +00:00
X-iZhang 27572dbd46 feat(onboard): add npx validation and installation prompt in onboarding process 2026-02-10 18:16:58 +00:00
X-iZhang 20db6be1d3 feat(onboard): standardize question mark style in onboarding steps 2026-02-10 16:15:54 +00:00
Xi Zhang 20751205e4 Merge pull request #5 from EvoScientist/feat/session-naming
Add `--name/-n` flag for naming run workspaces
2026-02-10 15:40:48 +00:00
Dinos Papakostas 7eb1d1bf4d docs: update README 2026-02-10 19:16:39 +08:00
Dinos Papakostas f1a7b957e7 chore(test): rm unused imports 2026-02-10 19:13:30 +08:00
Dinos Papakostas d1cc6096d2 test(cli): add tests for the deduplication logic 2026-02-10 19:09:49 +08:00
Dinos Papakostas e3ef17507b feat(cli): add --name flag 2026-02-10 19:07:00 +08:00
X-iZhang baf025da9a update 2026-02-10 02:00:49 +00:00
X-iZhang 3840a3dfea update 2026-02-09 20:11:16 +00:00
X-iZhang 250a23682c update 2026-02-09 02:54:41 +00:00
X-iZhang a32fe79bdd update 2026-02-08 20:52:37 +00:00
X-iZhang b923c73283 update 2026-02-08 20:28:57 +00:00
X-iZhang e2fbbf1b02 update 2026-02-08 18:50:04 +00:00
X-iZhang 61a3cb645f update 2026-02-08 16:46:38 +00:00
X-iZhang 7b15b3b8a0 update 2026-02-07 01:53:46 +00:00
X-iZhang 05b098ca63 update 2026-02-06 23:05:05 +00:00
X-iZhang d23acd9f01 update 2026-02-06 22:50:32 +00:00
X-iZhang 5518f4b1f9 update 2026-02-06 19:16:37 +00:00
Xi Zhang 64b2c56679 Merge pull request #4 from EvoScientist/feat/mcp
MCP Client Support
2026-02-06 19:14:15 +00:00
Dinos Papakostas f2d991b722 docs(README): add myself 2026-02-06 23:27:36 +08:00
Dinos Papakostas e07bcf2a91 test(mcp): add tests for MCP client 2026-02-06 23:17:45 +08:00
Dinos Papakostas 1898fdf33f feat(mcp): implement MCP client 2026-02-06 23:17:26 +08:00
Dinos Papakostas 53d38ea0c6 docs: add Gemini info to README 2026-02-06 21:33:44 +08:00
Dinos Papakostas b2f93bfda9 chore(gitignore): ignore uv.lock 2026-02-06 21:20:21 +08:00
X-iZhang f9c25de680 update 2026-02-06 01:51:44 +00:00
X-iZhang 5b7120a401 add view_image 2026-02-06 01:16:58 +00:00
X-iZhang 967670bfc2 update 2026-02-06 00:49:57 +00:00
X-iZhang 80baf85a2c update 2026-02-05 23:58:55 +00:00
X-iZhang 9341dc474b update 2026-02-05 23:14:11 +00:00
X-iZhang a507032236 update 2026-02-05 21:33:54 +00:00
X-iZhang a0397e449a fix 2026-02-05 20:56:27 +00:00
X-iZhang 6261cbd07c update 2026-02-05 20:48:58 +00:00
X-iZhang c3e473cc69 update imessge 2026-02-05 17:57:58 +00:00
Xi Zhang 3d54a5e3c5 Merge pull request #1 from EvoScientist/feature/imessage-channel
feat: Add iMessage channel integration with thinking and debounce support
2026-02-05 14:26:28 +00:00
MuXinCG ac485c30e8 fix some bugs 2026-02-05 20:25:42 +08:00
X-iZhang a267cabd6a feat(init): implement lazy imports for EvoScientist_agent and create_cli_agent 2026-02-05 11:15:22 +00:00
Xi Zhang 12f435b82a Merge pull request #3 from EvoScientist/feat/google-genai
Add Google GenAI (Gemini) Support
2026-02-05 11:08:39 +00:00
Xi Zhang 5eb7fb2a58 Merge pull request #2 from EvoScientist/fix/asyncio-event-loop
fix: Event loop closed error on alternating messages in interactive mode
2026-02-05 11:02:09 +00:00
Dinos Papakostas 10a63c785e test(llm): update tests 2026-02-05 18:57:56 +08:00
Dinos Papakostas 905dd5568a chore(deps): bump langchain-google-genai to a more recent version 2026-02-05 18:52:22 +08:00
Dinos Papakostas d3d1310986 feat(llm): support google genai models 2026-02-05 18:49:43 +08:00
Dinos Papakostas 6e2b37cc13 :Merge remote-tracking branch 'upstream' into fix/asyncio-event-loop 2026-02-05 18:19:01 +08:00
Dinos Papakostas cb7d3081a0 style: rm unused import 2026-02-05 18:18:19 +08:00
Dinos Papakostas fa304e5a55 test(asyncio): add event loop tests 2026-02-05 18:09:10 +08:00
Dinos Papakostas 55bdfef347 fix(streaming): don't close the event loop in interactive mode 2026-02-05 17:08:11 +08:00
MuXinCG 3d0805dd83 Revert "enable thinking mode for claude model"
This reverts commit c49be0956838a85406c18ebdaf3d721d18d69e40.
2026-02-05 16:21:34 +08:00
MuXinCG 67a561a1ee feat: add --thinking and --debounce options to /channel command 2026-02-05 16:18:52 +08:00
MuXinCG 3c29076e4c feat: add thinking message support for iMessage channel 2026-02-05 14:43:18 +08:00
MuXinCG 9f0ce14b94 feat: add message debounce for iMessage server 2026-02-05 10:12:37 +08:00
X-iZhang 97f104015b update 2026-02-04 23:53:24 +00:00
X-iZhang f46e328545 update 2026-02-04 23:45:09 +00:00
X-iZhang b677d8b08b update 2026-02-04 16:37:31 +00:00
X-iZhang c4698e118d update 2026-02-04 15:57:42 +00:00
MuXinCG 5810cdd7b1 enable thinking mode for claude model 2026-02-04 21:27:43 +08:00
MuXinCG b6f778e725 format telephone number and mail 2026-02-04 18:46:04 +08:00
MuXinCG 9f12a35b7f update 2026-02-04 15:48:21 +08:00
MuXinCG 0d789c33cd feat: add iMessage channel support with /channel command
- Add channels module with abstract Channel base class for extensiblemessaging integrations (iMessage, WeChat, etc.)
  - Implement iMessage channel using imsg CLI tool with RPC-based
    communication
  - Add /channel command to CLI for starting iMessage server
  - Support sender whitelist via --allow flag for security
2026-02-04 15:46:22 +08:00
X-iZhang 1aef7e450e update 2026-02-03 22:29:19 +00:00
X-iZhang e725fb08e2 add CI workflows 2026-02-03 22:24:32 +00:00
X-iZhang 7d4b8e6557 update 2026-02-03 17:23:23 +00:00
X-iZhang 51cf4a9acb update 2026-02-03 03:17:40 +00:00
X-iZhang 57aaa6309d update 2026-02-03 02:24:53 +00:00
X-iZhang e4b7a1eb62 add memory 2026-02-03 01:49:27 +00:00
X-iZhang 3eec63a447 update 2026-02-03 00:20:45 +00:00
X-iZhang f36fec0bf9 update 2026-02-02 23:22:18 +00:00
X-iZhang ecefcee1d0 update 2026-02-02 16:17:25 +00:00
X-iZhang 11dc675649 update 2026-02-02 02:34:14 +00:00
X-iZhang 80ee8c82cd update 2026-02-02 02:18:16 +00:00
X-iZhang 425e60e9f4 update 2026-02-01 23:34:40 +00:00
X-iZhang 3993029466 update cli 2026-02-01 20:52:08 +00:00
X-iZhang 068d6e135d add writing skills 2026-02-01 15:27:08 +00:00
X-iZhang 9fc25349ea update 2026-01-31 15:48:32 +00:00
X-iZhang 529527020b update 2026-01-29 23:01:24 +00:00
X-iZhang 5b0e57eb61 update 2026-01-29 22:45:59 +00:00
X-iZhang 19548016d9 EvoScientist Initial 2026-01-29 17:39:50 +00:00
X-iZhang eccbee0fa1 Update 2026-01-29 03:13:53 +00:00
Xi Zhang ad8f44af8d Create .gitignore for EvoScientist project
Add .gitignore file for Angular project dependencies and build artifacts.
2026-01-29 02:47:04 +00:00
Xi Zhang edeae89b2e Add MIT License to the project 2026-01-28 18:14:42 +00:00
X-iZhang f3126cba29 Initial 2026-01-27 13:34:55 +00:00
Xi Zhang b4036665c8 Add Main Branch section to README 2026-01-27 13:07:45 +00:00
507 changed files with 107874 additions and 45509 deletions
+65
View File
@@ -0,0 +1,65 @@
# VCS
.git
.gitignore
.gitattributes
# CI / project meta (image doesn't need these)
.github/
docs/
CONTRIBUTING.md
README.zh-CN.md
LICENSE
# Editor / tooling state
.vscode/
.idea/
.cursor/
.codex/
.claude/
.agents/
.cursorrules
.ruff_cache/
# Python build artifacts and caches
__pycache__/
*.py[cod]
*.egg-info/
*.egg
build/
dist/
.venv/
venv/
.pytest_cache/
.ruff_cache/
.coverage
# Tests aren't needed at runtime
tests/
# Notebooks
*.ipynb
.ipynb_checkpoints/
# Local runtime data (must never leak into the image)
.env
.env.*
!.env.example
runs/
workspace/
skills/
memory/
memories/
media/
conversation_history/
.deno_cache/
.langgraph_api/
large_tool_results/
*.log
botpy.log
# Docker outputs themselves
Dockerfile.*
docker-compose*.override.yml
# OS
.DS_Store
+38 -24
View File
@@ -1,27 +1,41 @@
# EvoScientist CLI environment variables
# The preferred configuration flow is `evosci onboard`, which writes
# ~/.evoscientist/config/settings.yaml. Environment variables can override it.
# EvoScientist — cp .env.example .env && fill in your keys
#
# LLM providers, models, and API keys are managed exclusively through the
# Model Registry (WebUI 大模型配置 / Config API); no provider credential is
# read from environment variables. See
# docs/unified-model-configuration-architecture.md.
# Optional application directories
# EVOSCIENTIST_HOME=~/.evoscientist
# EVOSCIENTIST_DATA_ROOT=~/.evoscientist/data
# Web search (optional)
TAVILY_API_KEY= # app.tavily.com
# Logging
EVOSCIENTIST_LOG_LEVEL=INFO
# EVOSCIENTIST_LOG_DIR=~/.evoscientist/data/logs
EVOSCIENTIST_LOG_RETENTION_DAYS=30
# WebUI conversation workspace policy. EVOSCIENTIST_WORKSPACE_DIR is the
# deployment root, not a per-conversation directory. In isolated modes each
# conversation is stored under <root>/.evoscientist/conversations/<scope-id>/.
#
# EVOSCIENTIST_WORKSPACE_ISOLATION accepts exactly:
# - legacy: all WebUI conversations share the deployment root. Compatibility
# rollback only; files are visible to every conversation using this deployment.
# - optional: default. New WebUI conversations receive isolated scope folders;
# missing Registry/token/scope fails the request instead of silently sharing.
# - required: isolated scopes plus strict runtime validation. It requires a
# completed cutover and a verified OCI executor; no legacy fallback exists.
#
# This is a deployment-startup security setting. Change it only during a
# maintenance window, restart backend and WebUI afterwards, and never use it to
# convert an existing conversation between shared and isolated directories.
EVOSCIENTIST_WORKSPACE_DIR=
EVOSCIENTIST_WORKSPACE_ISOLATION=optional
# Required mode supports only a single-host Registry topology in v1.
EVOSCIENTIST_SCOPE_REGISTRY_TOPOLOGY=single-host
# Required mode: use a pinned image digest, preserve single-host topology, and
# keep the Code Interpreter disabled unless its scoped implementation is enabled.
# Do not put EVOSCIENTIST_BACKEND_SERVICE_TOKEN here for a same-host `EvoSci
# deploy`: it is generated and passed privately at startup.
# EVOSCIENTIST_WORKSPACE_ISOLATION=required
# EVOSCIENTIST_STRICT_EXECUTOR=oci
# EVOSCIENTIST_STRICT_EXECUTOR_IMAGE=registry.example/evoscientist-runtime@sha256:replace-with-verified-digest
# EVOSCIENTIST_STRICT_CODE_INTERPRETER=disabled
# Optional PostgreSQL checkpoint storage. Without this value the CLI uses an
# in-memory checkpointer for the current process.
# EVOSCIENTIST_SESSION_DB_URL=postgresql://user:password@localhost:5432/evoscientist
# Model provider examples. Provider-specific settings can also be configured
# through `evosci onboard`.
# OPENAI_API_KEY=sk-...
# OPENAI_BASE_URL=https://api.openai.com/v1
# ANTHROPIC_API_KEY=sk-ant-...
# GOOGLE_API_KEY=...
# TAVILY_API_KEY=tvly-...
# Optional default model
# DEFAULT_MODEL=openai/gpt-5.4
# Conversation workspace isolation retention defaults (used by workspace_maintenance.py).
EVOSCIENTIST_DRAFT_WORKSPACE_TTL_HOURS=24
EVOSCIENTIST_WORKSPACE_TRASH_RETENTION_DAYS=7
Binary file not shown.

After

Width:  |  Height:  |  Size: 234 KiB

+1 -1
View File
@@ -5,5 +5,5 @@
<rect x="54" y="5" width="62" height="24" rx="6" fill="#1565c0"/>
<text x="85" y="22" text-anchor="middle"
font-family="Inter, -apple-system, system-ui, sans-serif"
font-size="13" font-weight="700" fill="#ffffff">v0.0.7</text>
font-size="13" font-weight="700" fill="#ffffff">v0.2.2</text>
</svg>

Before

Width:  |  Height:  |  Size: 555 B

After

Width:  |  Height:  |  Size: 555 B

+1 -1
View File
@@ -5,5 +5,5 @@
<rect x="54" y="5" width="62" height="24" rx="6" fill="#2563eb"/>
<text x="85" y="22" text-anchor="middle"
font-family="Inter, -apple-system, system-ui, sans-serif"
font-size="13" font-weight="700" fill="#ffffff">v0.0.7</text>
font-size="13" font-weight="700" fill="#ffffff">v0.2.2</text>
</svg>

Before

Width:  |  Height:  |  Size: 555 B

After

Width:  |  Height:  |  Size: 555 B

Binary file not shown.

Before

Width:  |  Height:  |  Size: 654 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 213 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 428 KiB

After

Width:  |  Height:  |  Size: 287 KiB

+20
View File
@@ -0,0 +1,20 @@
version: 2
updates:
# Base images in Dockerfile (BASE_IMAGE / NODE_IMAGE ARG defaults).
# Dependabot reads `FROM`, `COPY --from=`, and ARG-bound base refs, and
# bumps both the @sha256 digest and the trailing # vX.Y.Z comment.
- package-ecosystem: "docker"
directory: "/"
schedule:
interval: "weekly"
open-pull-requests-limit: 5
commit-message:
prefix: "chore(docker)"
labels:
- "dependencies"
- "docker"
# Single PR per cadence rather than one per image — keeps reviewer load low
# and lets us validate trixie/uv/node bumps as a coherent set.
groups:
base-images:
patterns: ["*"]
+67
View File
@@ -0,0 +1,67 @@
name: Docker
on:
push:
branches: ["main"]
tags: ["v*"]
pull_request:
paths:
- "Dockerfile"
- ".dockerignore"
- "pyproject.toml"
- "uv.lock"
- "EvoScientist/**"
- ".github/workflows/docker.yml"
workflow_dispatch:
concurrency:
group: docker-${{ github.ref }}
cancel-in-progress: true
env:
REGISTRY: ghcr.io
IMAGE_NAME: ${{ github.repository }}
jobs:
build:
runs-on: ubuntu-latest
timeout-minutes: 45
permissions:
contents: read
packages: write
steps:
- uses: actions/checkout@93cb6efe18208431cddfb8368fd83d5badbf9bfd # v5.0.1
- uses: docker/setup-qemu-action@c7c53464625b32c7a7e944ae62b3e17d2b600130 # v3.7.0
- uses: docker/setup-buildx-action@8d2750c68a42422c14e847fe6c8ac0403b4cbd6f # v3.12.0
- name: Log in to ${{ env.REGISTRY }}
if: github.event_name != 'pull_request'
uses: docker/login-action@c94ce9fb468520275223c153574b00df6fe4bcc9 # v3.7.0
with:
registry: ${{ env.REGISTRY }}
username: ${{ github.actor }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Extract image metadata
id: meta
uses: docker/metadata-action@c299e40c65443455700f0fdfc63efafe5b349051 # v5.10.0
with:
images: ${{ env.REGISTRY }}/${{ env.IMAGE_NAME }}
tags: |
type=ref,event=branch
type=ref,event=pr
type=semver,pattern={{version}}
type=semver,pattern={{major}}.{{minor}}
type=raw,value=latest,enable={{is_default_branch}}
- name: Build and push
uses: docker/build-push-action@10e90e3645eae34f1e60eeb005ba3a3d33f178e8 # v6.19.2
with:
context: .
platforms: linux/amd64,linux/arm64
push: ${{ github.event_name != 'pull_request' }}
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
cache-from: type=gha
cache-to: type=gha,mode=max
+33 -1
View File
@@ -7,13 +7,25 @@ on:
jobs:
pytest:
runs-on: ubuntu-latest
timeout-minutes: 15
# ``fail-fast: false`` so a single failing (os, python-version) cell
# doesn't cancel the rest of the matrix. Useful while the Windows
# leg is being brought up — we want to see all four cell results
# in one CI run instead of playing whack-a-mole one failure at a
# time. See #207.
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
python-version: ["3.11", "3.12"]
runs-on: ${{ matrix.os }}
steps:
- uses: actions/checkout@v5
- name: Check out shared usage fixtures
uses: actions/checkout@v5
with:
repository: EvoScientist/EvoScientist-WebUI
path: EvoScientist-WebUI
- uses: astral-sh/setup-uv@v6
with:
python-version: ${{ matrix.python-version }}
@@ -22,3 +34,23 @@ jobs:
run: uv sync --dev
- name: Run pytest
run: uv run pytest -v --timeout=30
env:
EVOSCIENTIST_USAGE_FIXTURES: ${{ github.workspace }}/EvoScientist-WebUI/docs/schemas/fixtures
usage-spool-benchmark:
timeout-minutes: 15
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]
runs-on: ${{ matrix.os }}
steps:
- uses: actions/checkout@v5
- uses: astral-sh/setup-uv@v6
with:
python-version: "3.11"
cache-dependency-glob: "**/pyproject.toml"
- name: Install dependencies
run: uv sync --dev
- name: Verify durable spool latency
run: uv run python scripts/benchmark_usage_spool.py
+7 -19
View File
@@ -9,13 +9,13 @@ dist/
build/
*.egg
*.pytest_cache/
.benchmarks/
.coverage
.ipynb_checkpoints/
# Environment
.env
.env.*
.env_*
!.env.example
.venv/
venv/
@@ -36,7 +36,9 @@ bridge/package-lock.json
.langgraph_api/
workspace/
skills/
!EvoScientist/skills/
memory/
!EvoScientist/memory/
media/
conversation_history/
.deno_cache/
@@ -46,22 +48,8 @@ conversation_history/
*meals/
botpy.log
large_tool_results/
runs/
# Docker runtime data
docker/data/
# Project-level config data
.data/
# Sensitive / credentials (never commit)
postgresql:*
_s3_backup/
# Local scratch / tooling data
.superpowers/
.test-home/
tmp/
# Root-level debug scratch scripts (proper tests live in tests/)
/test_*.py
/research_lookup_temp.py
# local runtime artifacts (scope tokens, control DBs)
.evoscientist/
.release-state.json
+2 -2
View File
@@ -1,10 +1,10 @@
repos:
- repo: https://github.com/astral-sh/ruff-pre-commit
# Ruff version.
rev: v0.9.9
rev: v0.15.17
hooks:
# Run the linter.
- id: ruff
- id: ruff-check
args: [ --fix ]
# Run the formatter.
- id: ruff-format
+132
View File
@@ -0,0 +1,132 @@
# Task 1 报告 — RegistryV4 schema + SQLite 存储层 + 统一错误码
## 实现摘要
按简报与设计文档 v1.1.0(4.2、4.3、5.1、8.2、9.5 节)在仓库内新建独立子包
`EvoScientist/model_registry/`,未改动任何现有文件(顶层 `EvoScientist/__init__.py`
采用惰性导出,无需修改)。
1. **`schemas.py` — 唯一 RegistryV4 schema(Pydantic v2)**
- `ProviderId`/`ModelKey`/`CredentialId`/`AdapterId` 共用锚定全匹配模式
`\A[a-z0-9][a-z0-9._-]{0,63}\z`(pydantic 的 pattern 约束默认是子串搜索,必须锚定)。
- `upstream_model_id` 仅限长 1–300、保留大小写;`ValidatedEndpoint` 限长 2048 且要求
http(s) scheme(EndpointPolicy 属后续任务)。
- `RegistryV4 { version: Literal[4], revision: PositiveInt, state, defaults, providers }`;
schema 级校验:Provider ID 唯一、Provider 内模型 key 唯一、defaults 引用必须存在、
`active` 状态下 primary 非空且引用已启用模型(auxiliary 同理)。
- `ProviderConfig.runtime`:`timeout_seconds` [10,600](缺省 120)、`max_retries` [0,5]
(缺省 2)、`default_temperature` [0,2]|null、`default_top_p` (0,1]|null、
`default_reasoning_effort` 缺省 `auto`。
- `ModelConfig.runtime`:`limit_mode` combined 要求 `context_window_tokens`、input_only
要求 `max_input_tokens`(model_validator);`min_effective_input_tokens` 缺省 4096、
下限 1024;三个 `fixed_*_reserve_tokens` 非负;temperature/top_p/reasoning_effort 与
`declared_capabilities` 按 4.3 定义。
- `AuthConfig`:`mode=none` 时 `credential_id` 必须为 null(model_validator)。
- 6.2/6.4/9.1 结构:`AdapterParameterSpec`(含 `connection` 可选块,对应 6.2 示例
`chat_model`/`model_field`/`base_url_field`)、`AuthSpec`、`ParameterRule`、
`ModelAvailability`(`VerificationInfo`)、`ResolvedModelConfig`
(`auth_ref={mode, credential_id?, credential_revision?}`,无任何 secret 字段)、
`CredentialStatus`、`CredentialWrite`(9.2 的 `operation: replace`)。
2. **`errors.py` — 统一错误码与载荷**
- 23 个稳定错误码常量 + `ERROR_HTTP_STATUS` 映射,覆盖 9.5 总表全部 17 组
(409×6、401×1、404×1、422×15)。
- `ErrorPayload {code, message, details:[{path, code}], request_id}` 与 9.5 结构一致;
`ModelRegistryError` 携带 code/`http_status`/`payload()`,未知 code 直接拒绝。
3. **`store.py` — `ModelRuntimeStore`**
- 数据库 `<config_dir>/model-runtime.sqlite3`(默认 `~/.config/evoscientist`,可注入);
目录 0700、文件 0600、WAL、外键、`busy_timeout=30000`。
- 手写 DDL(`CREATE TABLE IF NOT EXISTS`,无 alembic):`registry_state`(单行)、
`credential_pointers`、`credential_versions`((credential_id, revision) 主键)、
`model_verifications`(五元组主键,upsert 只留最近一次)、`run_runtime_snapshots`
(含部分唯一索引 `UNIQUE(deployment_id, thread_id, run_request_id)
WHERE status IN ('prepared','bound')`)、`delegation_jtis`。
- `load_registry()` 无行时返回 bootstrap/revision=1 空 RegistryV4;
`save_registry(expected_revision=, registry=, credential_writes=)` 在
`BEGIN IMMEDIATE` 事务内校验 revision(不符抛 `REGISTRY_REVISION_CONFLICT`)、
写入不可变凭据版本、registry revision+1;首次同时具备已启用模型+有效 primary+已配置
凭据(或 `mode=none`)时原子转为 `active`;任一失败整体回滚。每次保存强制执行
9.2 第 7 条(defaults 必须引用已启用模型,违反抛 `MODEL_DISABLED`)。
- 凭据:`write_credential_version`(递增 revision;重写相同当前密钥幂等返回原
revision)、`resolve_credential`(不存在/已销毁抛
`RUN_CREDENTIAL_REVISION_UNAVAILABLE`)、`retire_credential_version`、
`credential_status`(hint 末 4 位 `...abcd`,短于 4 字符的密钥 hint 为 null 绝不泄露;
绝不返回明文)。
- 快照:`insert_run_snapshot`/`set_run_snapshot_status`/`get_run_snapshot`
(状态机 prepared|bound|expired|aborted,部分唯一索引行为由测试覆盖)。
- `check_shared_storage()`:探测 `BEGIN IMMEDIATE` 写锁能力,失败抛
`SharedStorageError`(多节点不共享持久卷时启动失败)。
4. **`hashing.py` — `configuration_hash(provider, model)`**
- 覆盖 adapter、base_url、upstream_model_id、Provider 与 Model 全部运行参数(含声明
能力与限制),`json.dumps(sort_keys=True, separators=(",", ":"))` 规范化后 SHA-256。
## 文件清单
新增(无修改既有文件):
- `EvoScientist/model_registry/__init__.py`
- `EvoScientist/model_registry/schemas.py`
- `EvoScientist/model_registry/errors.py`
- `EvoScientist/model_registry/store.py`
- `EvoScientist/model_registry/hashing.py`
- `tests/test_model_registry_schemas.py`
- `tests/test_model_registry_store.py`
- `.superpowers/sdd/briefs/task-1-report.md`(本文件)
## 测试命令与输出
TDD 流程:先写两个测试文件并确认失败(`ModuleNotFoundError: No module named
'EvoScientist.model_registry'`),再实现。
```
$ .venv/bin/python -m pytest tests/test_model_registry_schemas.py tests/test_model_registry_store.py -x -q
........................................................................ [ 80%]
................. [100%]
89 passed in 0.22s
```
全量回归(无既有失败,无回归):
```
$ .venv/bin/python -m pytest tests/ -x -q
........sssss....................................... [100%]
2922 passed, 10 skipped, 1 warning in 73.00s
```
(warning 为 `test_langgraph_dev_http.py` 的 StarletteDeprecationWarning,既有、与本任务无关。)
lint 与格式:
```
$ .venv/bin/ruff check EvoScientist/model_registry tests/test_model_registry_schemas.py tests/test_model_registry_store.py
All checks passed!
$ .venv/bin/ruff format --check ... # 已格式化
```
## 自我审查发现(已处理)
1. **测试副作用污染真实配置目录**:初版 `test_default_config_dir` 未注入路径,运行时在
真实 `~/.config/evoscientist/` 创建了空的 `model-runtime.sqlite3`。已确认该库所有表
为空(确为测试副产物)后删除(含 -wal/-shm),并把测试改为 monkeypatch
`store.DEFAULT_CONFIG_DIR` 到 tmp_path,此后测试不再触碰真实 home。
2. **共享 `Field()` 实例**:初版 `_TEMPERATURE`/`_TOP_P` 在三个模型间复用同一 FieldInfo,
已改为各字段独立 `Field(...)`,规避 pydantic 共享元数据的潜在风险。
3. **docstring 混入中文**:`save_registry` 一处 docstring 误用中文,已改为英文以符合
仓库注释惯例。
4. **并发 CAS 断言**:`sorted()` 大小写排序导致误报,改为不区分大小写排序。
5. ruff 修复:`datetime.UTC` 别名、导入排序、`pytest.raises` 增加 `match=`。
## 遗留疑虑
1. **`connection` 块为可选**:6.2 正文的"至少包含"清单未列 `connection`,但 glm-5.2
示例契约包含它。schema 将其建模为可选字段,Task 2 落地五种 Adapter 契约时若确认
每个契约都有 connection,可考虑收紧为必填。
2. **激活就绪判定中的"满足 AuthSpec"**:4.3 要求激活时认证状态满足 AuthSpec,但
Adapter 契约属 Task 2。当前 store 仅做存储层判定(`mode=none` 或凭据已配置);
AuthSpec 级别校验(如 credential_kind 匹配)需在 Task 2/3 的 API 层补充。
3. **幂等语义解释**:简报称 `write_credential_version` 为"幂等 prepare",实现为"重写
与当前版本完全相同的密钥时返回现有 revision";不同密钥轮换仍产生新 revision。若
后续任务对幂等键有不同约定(如客户端提供 request id),需再对齐。
4. **`save_registry` 参数为 keyword-only**:与简报签名
`save_registry(expected_revision, registry, credential_writes=[])` 语义一致,仅调用
形式略异。
5. `MODEL_DISABLED` 用于 defaults 引用未启用模型的保存错误(9.2 第 7 条属 422 校验,
总表无更贴切码);引用不存在模型由 schema 层先行拒绝,store 内同名分支仅作防御。
+89
View File
@@ -0,0 +1,89 @@
# Task 4 报告 — ModelRegistryResolver + 运行快照服务
- 状态:DONE
- 分支:`feature/unified-model-config`
- Commit:`b1233d4` `feat(model-registry): add ModelRegistryResolver and run snapshot service`
- 规格来源:`docs/unified-model-configuration-architecture.md` v1.1.1,章节 4.3 / 5.2 / 6.1 / 6.4 / 6.5 / 8.1 / 8.2
## 交付物
### 新建 `EvoScientist/model_registry/resolver.py`
`ModelRegistryResolver(store, *, specs=None)`:
- `resolve(model_ref, role="primary", *, registry=None) -> ResolvedModelConfig`,校验顺序:
1. Provider 存在(否则 404 `MODEL_NOT_FOUND`)且 enabled(否则 422 `MODEL_NOT_AVAILABLE`);
2. 模型存在(`MODEL_NOT_FOUND`)且 enabled(否则 422 `MODEL_DISABLED`);
3. `find_adapter_spec` 匹配契约:未开放 Adapter 抛 `ADAPTER_NOT_SUPPORTED`,无匹配契约抛 `MODEL_NOT_AVAILABLE`(只可保持 configured);
4. `resolve_parameters` 重放保存期契约校验(AuthSpec、能力、参数范围/互斥);
5. 凭据:`credential_required` 时读 `credential_pointers.current_revision` 写入 `auth_ref`(指针缺失 → `CREDENTIAL_NOT_CONFIGURED`),**从不读 `secret_value`**;`mode=none` 冻结 `{mode: none}`;
6. 当前验证五元组(`configuration_hash` + 当前凭据 revision + 当前 `adapter_spec_revision`)必须存在且 `passed`,否则 `MODEL_NOT_AVAILABLE`;
7. `limits_status=confirmed`,否则 `MODEL_LIMITS_UNCONFIRMED`;
8. 预算(6.5):`combined` → `context_window - max_output`;`input_only` → `max_input`;四种工具/附件组合扣除固定预留后均须 ≥ `min_effective_input_tokens`,否则 `CONTEXT_BUDGET_UNSATISFIABLE`。冻结 `resolved_input_limit`、三个固定预留与 base `message_budget`(仅扣系统预留;逐次调用的预算由 MessageBudgetMiddleware 按 6.5 公式重算,Task 6 接线);
9. 能力按唯一规则 `protocol AND declared AND verified` 计算并冻结。
- `resolve_for_test(model_ref)`:仅放宽"模型必须 enabled",其余校验(含 Provider enabled、凭据、验证记录、限制、预算)全部执行(9.4)。
- `compute_availability(registry, verifications, *, credential_revisions=None) -> list[ModelAvailability]`:4.3 判定顺序 `unavailable → enabled → verified → verification_failed → verification_stale → configured`,首个命中生效;`verification_stale` 优先于 `configured`;`selectable` 仅 `enabled` 为 true。`verification` 报告当前记录的 passed/failed 或最新旧记录的 stale;`effective_capabilities` 仅在当前记录 passed 时非全 false。reason_code 取值:`PROVIDER_DISABLED` / `NO_ADAPTER_CONTRACT` / `MODEL_DISABLED` / 记录的错误码 / `VERIFICATION_STALE`。
- 6.1 角色映射实现为 `snapshots.config_for_role`(映射对象是快照而非 Registry,故放在快照模块):`primary → snapshot.primary`;`auxiliary`/`summary`/`tool_selector → snapshot.auxiliary ?? snapshot.primary`;未知角色 `ValueError`。Resolver 不猜测其他映射。
### 新建 `EvoScientist/model_registry/snapshots.py`
`SnapshotService(store, resolver)`(Task 5 HTTP API 与 Task 7 本地入口共用):
- `create(SnapshotCreateRequest)`:bootstrap → 422 `MODEL_REGISTRY_NOT_READY`;`primary=null`(inherit)解析 Registry `defaults.primary`,`auxiliary=null` 解析 `defaults.auxiliary`;冻结两个完整 `ResolvedModelConfig`(含 `adapter_spec_revision`、固定预留、能力、`auth_ref` 凭据版本)与 `registry_revision`、`model_selection_revision` 写入 `payload_json`,绝无 secret。`selection_hash` = 解析前 `{primary, auxiliary}`(inherit 以 null 参与)固定字段序 JSON 的 SHA-256;`model_selection_revision` 仅审计、不参与哈希。
- 幂等:同一三元组哈希相同返回原快照(`created=False`,200 语义),不同抛 409 `RUN_REQUEST_CONFLICT`;并发创建撞部分唯一索引时回退到同一幂等比较。`expired`/`aborted` 不占三元组,同三元组可重建(`created=True`,201)。
- `bind(snapshot_id, langgraph_run_id)`:`prepared→bound` 一次(条件 UPDATE 保证原子),并把 `expires_at` 延长到 +24h;相同 run id 重复 bind 幂等成功,不同值抛 `SNAPSHOT_ALREADY_BOUND`;终态抛 `SNAPSHOT_EXPIRED`;不存在抛 `SNAPSHOT_NOT_FOUND`。
- `abort(snapshot_id)`:仅 `prepared→aborted`;已 bound 抛 `SNAPSHOT_ALREADY_BOUND`;expired 抛 `SNAPSHOT_EXPIRED`;重复 abort 幂等成功。
- `get(snapshot_id, *, deployment_id, thread_id)`:绑定关系不匹配按 `SNAPSHOT_NOT_FOUND` 失败(不跨线程/部署泄露存在性);终态抛 `SNAPSHOT_EXPIRED`;读取时对两个冻结配置逐一重校验 `adapter_spec_revision` 仍存在(经 Resolver 的 spec 列表),已移除抛 `ADAPTER_NOT_SUPPORTED`,绝不静默替换。
- `cleanup_expired(now)`:到期 prepared(创建时 TTL **15 分钟**)与 bound(bind 时 **+24 小时**)置为终态 `expired`,返回迁移的 ID。
- `resolve_snapshot_credential(snapshot, role) -> str`:经 6.1 角色映射取冻结 `auth_ref`,按冻结 `credential_revision` 每次从凭据存储解析(**无进程内密钥缓存**);版本销毁抛 `RUN_CREDENTIAL_REVISION_UNAVAILABLE`;`mode=none` 返回 `""`。
- `public_snapshot_view(snapshot)`:仅 8.2 示例字段(`snapshot_id`、`registry_revision`、primary/auxiliary 的 `provider_id`/`model_key`/`adapter_spec_revision`/runtime 五项),无 base_url、无 secret。
### Task 1-3 文件的增补(均为纯新增,未改动任何既有对外行为)
- `errors.py`:新增 `SNAPSHOT_NOT_FOUND`(404)。9.5 承认"404 资源不存在"类别但总表只有 `MODEL_NOT_FOUND`(语义为 ModelRef);快照缺失需要独立稳定码,属对总表的增补,已同步更新 Task 1 的 `ALL_ERROR_CODES` 测试(仅加一行)。
- `store.py`:新增 `current_credential_revision`、`list_model_verifications`(可用性判定的输入)、`find_active_run_snapshot`(三元组幂等查询)、`bind_run_snapshot`(`WHERE status='prepared'` 条件绑定 + 延长 expires_at)、`expire_due_run_snapshots`;行→dict 映射提取为共享私有helper,既有方法签名不变。
- `__init__.py`:导出新符号。
- **未删除** `EvoScientist/llm/runtime_snapshots.py`(Task 6 切换)。
### Task 3 评审接线要求的遵守
本任务不构造任何 HTTP client / ChatModel:`resolve`/`resolve_for_test` 只产出 `ResolvedModelConfig`,`SnapshotService` 只做冻结与读取。因此 `build_chat_model` 双传 sync+async client、ollama `max_retries` 经 client builder `retries=` 执行这两条接线要求在本任务无适用点,也未被绕过;Task 5/7 构造 client 时仍须遵守(`build_safe_http_client(policy, retries=resolved.client_options.max_retries)` 双传)。
## 测试(TDD)
先写 `tests/test_resolver.py`(37 例)与 `tests/test_snapshots.py`(34 例)并确认红灯(模块不存在),再实现至全绿。覆盖简报全部要求:
- resolve 正/反例:完整冻结断言(含预算数值 1048576-32768-4096)、Provider disabled→`MODEL_NOT_AVAILABLE`、模型 disabled→`MODEL_DISABLED`、无匹配契约→不可用、验证缺失/失败/五元组不一致(配置哈希、凭据轮换)→不可用、`MODEL_LIMITS_UNCONFIRMED`、`CONTEXT_BUDGET_UNSATISFIABLE`、四种角色戳记、未开放 Adapter、`AUTH_MODE_UNSUPPORTED`、mode=none 无凭据解析、resolved config 不含 secret;
- resolve_for_test:放宽模型 enabled、其余校验不放宽;
- ModelAvailability 六态判定顺序、stale 优先 configured、凭据轮换致 stale、selectable 仅 enabled、全 Provider 全模型覆盖;
- 快照:inherit 解析 defaults(含 auxiliary default null 两条路径)、显式选择冻结双角色、bootstrap 422、selection_hash 解析前语义(显式等于默认仍不同哈希)、revision 不参与哈希、幂等 200/冲突 409、expired/aborted 后重建 201、payload 冻结凭据版本且无 secret、prepared TTL 15min 断言;
- bind 一次/同 id 幂等/异 id 409/expired/aborted/未知 404;abort 规则全集;get 绑定校验(跨线程/跨部署拒绝)、expired 拒绝、spec_revision 移除后读取 `ADAPTER_NOT_SUPPORTED`;cleanup 到期迁移;bind 后 +24h 断言;
- 凭据:冻结 revision 解析、轮换后旧版本仍可用、销毁后 `RUN_CREDENTIAL_REVISION_UNAVAILABLE`、新 store 实例(无进程缓存)仍可解析、辅助角色解析到 auxiliary 凭据、mode=none 返回空串;
- 6.1 角色映射(auxiliary 冻结/缺省两路径、未知角色)与 `public_snapshot_view` 形状(无 secret、无 base_url)。
## 验证结果
- `.venv/bin/python -m pytest tests/test_resolver.py tests/test_snapshots.py -x -q` → **71 passed**
- `.venv/bin/python -m pytest tests/ -x -q` → **3131 passed, 10 skipped**(无回归;期间修复一处:Task 1 错误码表测试因新增 `SNAPSHOT_NOT_FOUND` 需增补一行期望值)
- `ruff check` 与 `ruff format --check`(model_registry 全包 + 涉及测试)→ 全净
## 疑虑 / 后续注意
1. `SNAPSHOT_NOT_FOUND` 是对设计文档 9.5 错误码总表的增补(404 类别文档已承认,但总表未列快照缺失码);Task 5 HTTP 层应直接复用。
2. 终态(expired/aborted)快照的 bind/get 统一抛 `SNAPSHOT_EXPIRED`;文档只明文规定 expired 的情形,aborted 按同一终态语义处理。
3. 冻结的 `budget.message_budget` 取 base(仅扣系统预留);逐次调用的 has_tools/has_attachments 重算属 Task 6 的 MessageBudgetMiddleware。
4. `resolve` 支持 `registry=` 参数供 `create` 传入同一份 Registry,保证快照 `registry_revision` 与解析所用文档一致。
## 评审修复(2026-07-21,commit 见下)
1. **Important:`abort` read-then-write 竞态**。原实现先读后写且 `set_run_snapshot_status` 为无条件 UPDATE,并发 bind 在两次调用间提交时会把 bound 改写为 aborted 并丢失 `langgraph_run_id`。修复:store 层新增 `abort_run_snapshot`(`UPDATE ... SET status='aborted' WHERE snapshot_id=? AND status='prepared'`,按 rowcount 判定),`SnapshotService.abort` 改为与 `bind` 相同的读-条件写-失败重读循环;已 aborted 重复调用保持幂等成功。
2. **Minor:selection_hash 期望值自证**。`test_selection_hash_uses_pre_resolution_semantics` 原先用测试内重复实现的同一序列化逻辑计算期望值(两侧同变不红)。改为钉死离线算出的 SHA-256 字面值(`{"auxiliary":null,"primary":null}` → `697c0462...55abc`,常量 `INHERIT_SELECTION_HASH`),锁住对外契约;删除测试内的重复实现。
新增测试(先红后绿):
- `test_abort_run_snapshot_store_update_is_conditional`:store 层条件 UPDATE 的 rowcount 语义(prepared→True,重复/bound→False 且行不被改写)。
- `test_abort_losing_bind_race_keeps_bound_state`:monkeypatch `get_run_snapshot` 在 abort 读与写之间插入并发 bind,断言 abort 抛 `SNAPSHOT_ALREADY_BOUND` 且行保持 bound、`langgraph_run_id` 完好(旧实现此测试必红)。
验证:
- `.venv/bin/python -m pytest tests/test_snapshots.py -x -q` → **42 passed**
- `.venv/bin/python -m pytest tests/ -x -q` → **3133 passed, 10 skipped**(无回归)
- `ruff check` / `ruff format --check`(涉及文件)→ 全净
+1 -1
View File
@@ -69,7 +69,7 @@ EvoScientist is a multi-agent AI system for automated scientific experimentation
| Framework | [DeepAgents](https://github.com/langchain-ai/deepagents) + [LangChain](https://python.langchain.com/) + [LangGraph](https://langchain-ai.github.io/langgraph/) |
| Default model | `claude-sonnet-4-6` (Anthropic) |
| Tests | ~890 across 36 files, no API keys needed |
| Config file | `.data/.config/settings.yaml` (project root) |
| Config file | `~/.config/evoscientist/config.yaml` |
### Sub-Agents (defined in `EvoScientist/subagent.yaml`)
+75
View File
@@ -0,0 +1,75 @@
# syntax=docker/dockerfile:1.7
ARG BASE_IMAGE=ghcr.io/astral-sh/uv:python3.11-trixie-slim@sha256:7936cc6625ca04cafa6ecc3c2881ddfe90a747c55c74480cd4ac6ffad6a5af1e
ARG NODE_IMAGE=node:24-trixie-slim@sha256:735dd688da64d22ebd9dd374b3e7e5a874635668fd2a6ec20ca1f99264294086
FROM ${NODE_IMAGE} AS nodejs
# ---------- Builder ----------
FROM ${BASE_IMAGE} AS builder
ENV UV_COMPILE_BYTECODE=1 \
UV_LINK_MODE=copy \
UV_PYTHON_DOWNLOADS=never \
UV_PROJECT_ENVIRONMENT=/opt/venv
WORKDIR /src
COPY pyproject.toml uv.lock README.md ./
RUN --mount=type=cache,target=/root/.cache/uv \
uv sync --frozen --no-install-project --no-dev \
--extra all-channels
COPY EvoScientist ./EvoScientist
RUN --mount=type=cache,target=/root/.cache/uv \
uv sync --frozen --no-dev --no-editable \
--extra all-channels
# ---------- Runtime ----------
FROM ${BASE_IMAGE} AS runtime
RUN apt-get update \
&& apt-get install -y --no-install-recommends \
git \
ca-certificates \
tini \
curl \
&& rm -rf /var/lib/apt/lists/*
COPY --from=nodejs /usr/local/bin/node /usr/local/bin/node
COPY --from=nodejs /usr/local/lib/node_modules /usr/local/lib/node_modules
RUN ln -sf /usr/local/lib/node_modules/npm/bin/npm-cli.js /usr/local/bin/npm \
&& ln -sf /usr/local/lib/node_modules/npm/bin/npx-cli.js /usr/local/bin/npx
ARG UID=1000
ARG GID=1000
RUN groupadd --gid ${GID} evosci \
&& useradd --uid ${UID} --gid ${GID} --create-home --shell /bin/bash evosci
COPY --from=builder /opt/venv /opt/venv
ENV PATH="/opt/venv/bin:/home/evosci/.evoscientist/.local/bin:${PATH}" \
PYTHONUNBUFFERED=1 \
PYTHONDONTWRITEBYTECODE=1 \
EVOSCIENTIST_WORKSPACE_DIR=/workspace \
EVOSCIENTIST_DATA_DIR=/home/evosci/.evoscientist \
XDG_CONFIG_HOME=/home/evosci/.evoscientist/.config \
UV_TOOL_DIR=/home/evosci/.evoscientist/.local/share/uv/tools \
UV_TOOL_BIN_DIR=/home/evosci/.evoscientist/.local/bin
RUN mkdir -p /workspace \
/home/evosci/.evoscientist/.config/evoscientist \
/home/evosci/.evoscientist/.local/bin \
/home/evosci/.evoscientist/.local/share/uv/tools \
&& chown -R ${UID}:${GID} /workspace /home/evosci
USER evosci
WORKDIR /workspace
LABEL org.opencontainers.image.title="EvoScientist" \
org.opencontainers.image.description="EvoScientist agent with core + all-channels dependencies pre-installed." \
org.opencontainers.image.source="https://github.com/EvoScientist/EvoScientist" \
org.opencontainers.image.documentation="https://github.com/EvoScientist/EvoScientist#-docker" \
org.opencontainers.image.licenses="Apache-2.0"
ENTRYPOINT ["tini", "--", "evosci"]
File diff suppressed because it is too large Load Diff
+1 -12
View File
@@ -9,14 +9,13 @@ from __future__ import annotations
from importlib import import_module
from ._version import __version__
_EXPORTS: dict[str, tuple[str, str]] = {
# Agent graph (lazy to avoid expensive initialization at import time)
"EvoScientist_agent": (".EvoScientist", "EvoScientist_agent"),
"create_cli_agent": (".EvoScientist", "create_cli_agent"),
# Backends
"CustomSandboxBackend": (".backends", "CustomSandboxBackend"),
"MemoryFilesystemBackend": (".backends", "MemoryFilesystemBackend"),
"ReadOnlyFilesystemBackend": (".backends", "ReadOnlyFilesystemBackend"),
# Configuration
"EvoScientistConfig": (".config", "EvoScientistConfig"),
@@ -24,26 +23,16 @@ _EXPORTS: dict[str, tuple[str, str]] = {
"save_config": (".config", "save_config"),
"get_effective_config": (".config", "get_effective_config"),
"get_config_path": (".config", "get_config_path"),
# LLM
"get_chat_model": (".llm", "get_chat_model"),
"list_models": (".llm", "list_models"),
# Prompts
"get_system_prompt": (".prompts", "get_system_prompt"),
"RESEARCHER_INSTRUCTIONS": (".prompts", "RESEARCHER_INSTRUCTIONS"),
# Tools
"tavily_search": (".tools", "tavily_search"),
"web_search": (".tools", "web_search"),
"web_extract": (".tools", "web_extract"),
"web_crawl": (".tools", "web_crawl"),
"think_tool": (".tools", "think_tool"),
# Sessions
"get_checkpointer": (".sessions", "get_checkpointer"),
"generate_thread_id": (".sessions", "generate_thread_id"),
"list_threads": (".sessions", "list_threads"),
"delete_thread": (".sessions", "delete_thread"),
"get_storage_stats": (".sessions", "get_storage_stats"),
"get_aggregated_storage_stats": (".sessions", "get_aggregated_storage_stats"),
"list_all_session_db_paths": (".sessions", "list_all_session_db_paths"),
}
-1
View File
@@ -1 +0,0 @@
__version__ = "0.1.20"
+50
View File
@@ -0,0 +1,50 @@
"""Windows asyncio event-loop policy compatibility.
On Windows, MCP stdio servers are launched as subprocesses by the MCP SDK's
stdio transport, which uses ``anyio.open_process`` → ``asyncio`` async
subprocess support. ``asyncio``'s *Selector* event loop does **not** implement
async subprocess creation, so on a Selector loop the stdio transport falls back
to a synchronous ``subprocess.Popen`` inside an ``async`` function. Under
``langgraph dev`` (which enables ``blockbuster`` by default to police blocking
I/O) that synchronous call is flagged as a ``BlockingError`` — see issue #283.
The *Proactor* loop supports async subprocesses natively, so the fallback never
happens and ``blockbuster`` allows the (now genuinely async) spawn.
Windows has defaulted to ``WindowsProactorEventLoopPolicy`` since Python 3.8, so
this is normally a no-op. We set it explicitly anyway as a safeguard: a
dependency, IDE, or notebook host may have installed a Selector policy earlier
in the process, and the ``langgraph dev`` subprocess in particular runs code we
don't fully control. Calling this at each process entrypoint — **before any
event loop is created** — guarantees the MCP subprocess path stays async.
This must run at import/startup time, ahead of the first ``asyncio.run`` /
``new_event_loop`` call; once a loop exists, swapping the policy does not change
the already-running loop.
"""
from __future__ import annotations
import sys
def ensure_proactor_event_loop_policy() -> bool:
"""Install ``WindowsProactorEventLoopPolicy`` on Windows if needed.
Returns ``True`` if a Proactor policy is in effect afterwards (always
``False`` off Windows, where the concept doesn't apply). Safe and idempotent
to call multiple times; a no-op on non-Windows platforms.
"""
if sys.platform != "win32":
return False
import asyncio
proactor_policy = getattr(asyncio, "WindowsProactorEventLoopPolicy", None)
if proactor_policy is None: # pragma: no cover - non-Windows / stripped build
return False
current = asyncio.get_event_loop_policy()
if not isinstance(current, proactor_policy):
asyncio.set_event_loop_policy(proactor_policy())
return True
+876 -618
View File
File diff suppressed because it is too large Load Diff
+334
View File
@@ -0,0 +1,334 @@
"""Background OS-process execution for the sandbox.
A *process* here is a single detached OS process launched via ``run_in_background``
(distinct from an async sub-agent *task* and a future cron *schedule* — the word
"job" is intentionally never used).
The registry is **module-global (process-level)**: processes survive ``/new`` and
``/resume`` within the same CLI process, but are not persisted across a CLI restart.
The live ``Popen`` handle is held so ``poll()`` / ``returncode`` stay authoritative
(no PID-reuse risk).
Command validation and cwd resolution happen at the tool layer
(``middleware/background.py``); this module is the pure execution + tracking mechanism
and is safe to unit-test on its own. A future scheduler (cron) would reuse ``launch``.
"""
from __future__ import annotations
import logging
import os
import signal
import subprocess
import threading
import time
import uuid
from collections.abc import Callable
from dataclasses import dataclass, field
from datetime import UTC, datetime
from pathlib import Path
import psutil
logger = logging.getLogger(__name__)
_BG_DIRNAME = ".bg_processes"
_KILL_GRACE_SECONDS = 2.0
@dataclass
class BgProcess:
"""A tracked background OS process."""
process_id: str
name: str
command: str
popen: subprocess.Popen
pid: int
log_path: Path
started_at: str # ISO-8601 UTC (record/display)
started_ts: float # epoch seconds (elapsed computation)
origin_thread_id: str | None = None # CLI thread/session that launched it
returncode: int | None = None
finished_at: str | None = None
finished_ts: float | None = None # epoch at exit; freezes elapsed once done
stopped: bool = False # set by stop(); suppresses the completion notification
# epoch each thread last checked this process (status/list); keyed by thread_id
# so a check from one session can't dedup another session's completion ping.
last_checked_by_thread: dict[str | None, float] = field(default_factory=dict)
_PROCESSES: dict[str, BgProcess] = {}
_LOCK = threading.Lock()
def _now_iso() -> str:
return datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%SZ")
def _record_exit(proc: BgProcess) -> None:
"""Record terminal state on first observed exit. Caller MUST hold ``_LOCK``.
``finished_ts`` is set when the exit is first observed. The per-process daemon
watcher (:func:`_watch`) calls this right after ``popen.wait()`` returns, so in
practice ``finished_ts`` ≈ the real exit time. Calls from ``status`` / ``list_all`` /
``stop`` are a fallback for the brief window before the watcher runs.
"""
rc = proc.popen.poll()
if rc is not None and proc.returncode is None:
proc.returncode = rc
proc.finished_at = _now_iso()
proc.finished_ts = time.time()
def _elapsed(proc: BgProcess) -> int:
"""Seconds the process has run — frozen at first-observed exit once it has exited."""
end = proc.finished_ts if proc.finished_ts is not None else time.time()
return int(end - proc.started_ts)
def was_observed_done(process_id: str, origin_thread_id: str | None = None) -> bool:
"""True if ``origin_thread_id`` already saw this process's completion itself.
i.e. the process has exited AND was checked (``status``/``list_all``) from that thread
at or after it finished. Used to dedup the completion notification (routed to the
launching thread), so a check from a *different* session can't suppress it.
"""
with _LOCK:
proc = _PROCESSES.get(process_id)
if proc is None or proc.finished_ts is None:
return False
seen_ts = proc.last_checked_by_thread.get(origin_thread_id)
return seen_ts is not None and seen_ts >= proc.finished_ts
def _read_tail(log_path: Path, tail_bytes: int) -> str:
# Seek from the end so a huge log isn't fully read into memory on each status check.
try:
with log_path.open("rb") as f:
f.seek(0, os.SEEK_END)
size = f.tell()
if size == 0:
return "(no output yet)"
if size > tail_bytes:
f.seek(-tail_bytes, os.SEEK_END)
return "...(truncated)...\n" + f.read().decode("utf-8", "replace")
f.seek(0)
data = f.read()
except OSError:
return "(no output captured yet)"
return data.decode("utf-8", "replace")
def _watch(proc: BgProcess, on_exit: Callable[[BgProcess], None] | None) -> None:
"""Block until ``proc`` exits, record the exit promptly, then fire ``on_exit``.
Running in a daemon thread, ``popen.wait()`` lets us record ``finished_ts`` at (very
close to) the real exit time — fixing the observation-time inflation — and gives a
hook the CLI layer wires to a completion notification, without ``background.py``
importing the notifier (kept decoupled via the callback).
"""
try:
proc.popen.wait()
except Exception:
pass
with _LOCK:
_record_exit(proc)
if on_exit is not None:
try:
on_exit(proc)
except Exception:
logger.warning("background on_exit callback failed", exc_info=True)
def launch(
command: str,
cwd: str,
name: str | None = None,
*,
origin_thread_id: str | None = None,
on_exit: Callable[[BgProcess], None] | None = None,
) -> str:
"""Launch ``command`` detached in ``cwd``; return a short ``process_id``.
The command is run via ``shell=True`` with output redirected to a per-process log
file under ``<cwd>/.bg_processes/`` and ``start_new_session=True`` so the child is a
process-group leader (survives this call's return and can be killed as a group).
The caller is responsible for validating ``command`` first.
``origin_thread_id`` records the launching CLI session so ``list_all`` can scope to it.
``on_exit`` (optional) is called with the ``BgProcess`` from a daemon watcher thread
once the process exits — used by the CLI layer to emit a completion notification.
"""
process_id = uuid.uuid4().hex[:8]
log_dir = Path(cwd) / _BG_DIRNAME
log_dir.mkdir(parents=True, exist_ok=True)
log_path = log_dir / f"{process_id}.log"
log_file = open(log_path, "w")
try:
popen = subprocess.Popen(
command,
shell=True,
cwd=cwd,
stdout=log_file,
stderr=subprocess.STDOUT,
stdin=subprocess.DEVNULL,
start_new_session=True,
)
finally:
# The child inherited its own dup of the fd during spawn; the parent's copy
# is no longer needed (and must be closed so the pipe/file isn't held open).
log_file.close()
proc = BgProcess(
process_id=process_id,
name=name or command[:40],
command=command,
popen=popen,
pid=popen.pid,
log_path=log_path,
started_at=_now_iso(),
started_ts=time.time(),
origin_thread_id=origin_thread_id,
)
with _LOCK:
_PROCESSES[process_id] = proc
# Daemon watcher: records the precise exit time and fires on_exit when done.
threading.Thread(target=_watch, args=(proc, on_exit), daemon=True).start()
return process_id
def status(
process_id: str, *, thread_id: str | None = None, tail_bytes: int = 16_000
) -> str:
"""Return a human-readable status + recent output tail for ``process_id``."""
with _LOCK:
proc = _PROCESSES.get(process_id)
if proc is None:
return (
f"No such background process: {process_id!r}. "
"Use list_processes to see tracked processes."
)
_record_exit(proc)
proc.last_checked_by_thread[thread_id] = time.time() # this thread observed it
running = proc.returncode is None
elapsed = _elapsed(proc)
name, pid, command, returncode, log_path = (
proc.name,
proc.pid,
proc.command,
proc.returncode,
proc.log_path,
)
if running:
head = f"Process {process_id} (name={name!r}) RUNNING — {elapsed}s elapsed, pid {pid}."
else:
head = f"Process {process_id} (name={name!r}) EXITED code {returncode} after ~{elapsed}s."
tail = _read_tail(log_path, tail_bytes) # file IO outside the lock
return (
f"{head}\nCommand: {command}\n--- output (last {tail_bytes} bytes) ---\n{tail}"
)
def _kill_process_tree(popen: subprocess.Popen, *, forceful: bool) -> None:
"""Kill the process group/tree in a cross-platform way.
On POSIX ``start_new_session=True`` makes the child a process-group
leader; ``os.killpg`` terminates the entire group (shell + any
grandchildren). On Windows ``TerminateProcess`` (used by
``Popen.terminate()`` / ``Popen.kill()``) only kills the direct
child — it does *not* cascade to grandchildren. We use ``psutil``
to walk the process tree and signal every descendant.
"""
if os.name == "nt":
try:
proc = psutil.Process(popen.pid)
targets = [proc, *proc.children(recursive=True)]
except (psutil.NoSuchProcess, psutil.AccessDenied):
return
for p in targets:
try:
if forceful:
p.kill()
else:
p.terminate()
except (psutil.NoSuchProcess, psutil.AccessDenied):
pass
else:
sig = signal.SIGKILL if forceful else signal.SIGTERM
try:
os.killpg(os.getpgid(popen.pid), sig)
except ProcessLookupError:
pass
def stop(process_id: str) -> str:
"""Terminate ``process_id`` and its process group (SIGTERM, then SIGKILL)."""
with _LOCK:
proc = _PROCESSES.get(process_id)
if proc is None:
return f"No such background process: {process_id!r}."
if proc.popen.poll() is not None:
_record_exit(proc)
return f"Process {process_id} already finished (code {proc.returncode})."
# Mark as user-stopped so the watcher's on_exit suppresses the completion
# notification (the user already knows — no need to ping them).
proc.stopped = True
# The watcher's popen.wait() reaps without the lock, so a tiny PID-reuse race
# remains (getpgid on a recycled pid). On POSIX ProcessLookupError covers the
# common case; on Windows ``Popen.terminate()`` is a no-op on a dead handle
# so we poll after the call instead.
_kill_process_tree(proc.popen, forceful=False)
if proc.popen.poll() is not None:
_record_exit(proc)
return f"Process {process_id} is no longer running."
deadline = time.time() + _KILL_GRACE_SECONDS
while time.time() < deadline:
with _LOCK:
if proc.popen.poll() is not None:
_record_exit(proc)
break
time.sleep(0.1)
else:
with _LOCK:
if proc.popen.poll() is None:
_kill_process_tree(proc.popen, forceful=True)
_record_exit(proc)
with _LOCK:
_record_exit(proc)
name = proc.name
return f"Stopped background process {process_id} (name={name!r})."
def list_all(thread_id: str | None = None, *, include_all: bool = False) -> str:
"""List tracked background processes with live statuses.
Scoped to the launching session (``thread_id``) unless ``include_all`` is set.
"""
with _LOCK:
all_procs = list(_PROCESSES.values())
procs = (
all_procs
if include_all
else [p for p in all_procs if p.origin_thread_id == thread_id]
)
if not procs:
if all_procs and not include_all:
return (
"No background processes in this session "
f"({len(all_procs)} in other sessions — pass all_threads=True to see them)."
)
return "No background processes tracked."
lines = []
now = time.time()
for p in procs:
_record_exit(p)
p.last_checked_by_thread[thread_id] = now # this thread observed it
state = "RUNNING" if p.returncode is None else f"exited({p.returncode})"
lines.append(
f" {p.process_id} {state:12} {_elapsed(p)}s name={p.name!r}"
)
return f"{len(procs)} background process(es):\n" + "\n".join(lines)
+1 -1
View File
@@ -2,7 +2,7 @@
EvoScientist provides unified integration with 10 messaging platforms. This document covers the architecture overview, message processing pipeline, capability matrix, security model, deployment guides, and troubleshooting.
Configuration file: `~/.config/ai4scientist/settings.yaml` (or use environment variables with the `EVOSCIENTIST_` prefix).
Configuration file: `~/.config/evoscientist/config.yaml` (or use environment variables with the `EVOSCIENTIST_` prefix).
## Table of Contents
+2 -59
View File
@@ -2,23 +2,20 @@
Channels push messages to the inbound queue; the agent (or any consumer)
reads from inbound, processes, and pushes responses to the outbound queue.
A background dispatcher routes outbound messages to the correct channel
via subscriber callbacks.
``ChannelManager._dispatch_outbound`` routes outbound messages to the
correct channel by looking up its registered :class:`Channel` instance.
Deduplication is handled at the Channel level (single dedup point).
"""
import asyncio
import logging
from collections.abc import Awaitable, Callable
from ..debug import TraceMixin, debug_trace_enabled
from .events import InboundMessage, OutboundMessage
logger = logging.getLogger(__name__)
OutboundCallback = Callable[[OutboundMessage], Awaitable[None]]
class MessageBus(TraceMixin):
"""Async message bus that decouples chat channels from the agent core."""
@@ -28,8 +25,6 @@ class MessageBus(TraceMixin):
def __init__(self):
self.inbound: asyncio.Queue[InboundMessage] = asyncio.Queue(maxsize=5000)
self.outbound: asyncio.Queue[OutboundMessage] = asyncio.Queue(maxsize=5000)
self._outbound_subscribers: dict[str, list[OutboundCallback]] = {}
self._running = False
self._debug_trace = debug_trace_enabled()
self._trace_logger = logger
@@ -53,58 +48,6 @@ class MessageBus(TraceMixin):
"""Consume the next outbound message (blocks until available)."""
return await self.outbound.get()
# ── subscriber routing ──
def subscribe_outbound(
self,
channel: str,
callback: OutboundCallback,
) -> None:
"""Register a callback for outbound messages targeting *channel*."""
if channel not in self._outbound_subscribers:
self._outbound_subscribers[channel] = []
self._outbound_subscribers[channel].append(callback)
async def dispatch_outbound(self) -> None:
"""Route outbound messages to subscribed channels.
Run as a background task — loops until :meth:`stop` is called.
"""
self._running = True
while self._running:
try:
msg = await asyncio.wait_for(
self.outbound.get(),
timeout=1.0,
)
except TimeoutError:
continue
subscribers = self._outbound_subscribers.get(msg.channel, [])
if not subscribers:
self._trace_event(
"bus_dispatch_drop",
target_channel=msg.channel,
reason="no_subscriber",
chat_id=msg.chat_id,
)
logger.warning(f"No subscriber for channel: {msg.channel}")
continue
for callback in subscribers:
try:
await callback(msg)
except Exception as e:
self._trace_event(
"bus_dispatch_error",
target_channel=msg.channel,
chat_id=msg.chat_id,
error_type=type(e).__name__,
)
logger.error(f"Error dispatching to {msg.channel}: {e}")
def stop(self) -> None:
"""Stop the dispatcher loop."""
self._running = False
@property
def inbound_size(self) -> int:
return self.inbound.qsize()
+1
View File
@@ -160,6 +160,7 @@ QQ = ChannelCapabilities(
format_type="plain",
max_text_length=4096,
typing=False, # no typing API for QQ bots
inline_buttons=True, # markdown + keyboard payload (C2C only)
media_send=True,
media_receive=True,
voice=False, # qq-botpy does not expose voice as a distinct message type
+129 -90
View File
@@ -11,13 +11,12 @@ from __future__ import annotations
import asyncio
import logging
import time
import uuid
from collections import OrderedDict
from collections.abc import AsyncIterator, Callable
from dataclasses import dataclass
from typing import Any, TypeVar
from ..gateway import GraphGateway, GraphRunInput, GraphTarget, RunRequest
from .base import Channel
from .bus import MessageBus
from .bus.events import InboundMessage, OutboundMessage
@@ -119,7 +118,7 @@ def _should_auto_approve(action_requests: list[dict]) -> bool:
return True
try:
from ..config.settings import load_config
from ..config.settings import HITL_SHELL_TOOLS, load_config
cfg = load_config()
except Exception:
@@ -135,14 +134,10 @@ def _should_auto_approve(action_requests: list[dict]) -> bool:
)
for req in action_requests:
name = (
req.get("name", "") if isinstance(req, dict) else getattr(req, "name", "")
)
if name != "execute":
name = req.get("name", "")
if name not in HITL_SHELL_TOOLS:
continue
args = (
req.get("args", {}) if isinstance(req, dict) else getattr(req, "args", {})
)
args = req.get("args", {})
command = args.get("command", "") if isinstance(args, dict) else ""
cmd = command.strip()
if not any(cmd.startswith(prefix) for prefix in shell_allow_list):
@@ -150,16 +145,18 @@ def _should_auto_approve(action_requests: list[dict]) -> bool:
return True
def _format_approval_prompt(action_requests: list[dict]) -> str:
"""Format an approval prompt as a text message for channel users."""
def _format_approval_prompt(
action_requests: list[dict], *, with_buttons: bool = False
) -> str:
"""Format an approval prompt as a text message for channel users.
When *with_buttons* is True, the trailing "Reply: 1=Approve..."
instruction is dropped — the buttons replace the textual cue.
"""
lines = ["\u26a0\ufe0f Approval Required\n"]
for i, req in enumerate(action_requests, 1):
name = (
req.get("name", "") if isinstance(req, dict) else getattr(req, "name", "")
)
args = (
req.get("args", {}) if isinstance(req, dict) else getattr(req, "args", {})
)
name = req.get("name", "")
args = req.get("args", {})
if isinstance(args, dict):
command = args.get("command", args.get("path", ""))
else:
@@ -168,6 +165,7 @@ def _format_approval_prompt(action_requests: list[dict]) -> str:
lines.append(f" {i}. {name}: {command}")
else:
lines.append(f" {i}. {name}")
if not with_buttons:
lines.append("")
lines.append("Reply: 1=Approve, 2=Reject, 3=Approve all")
lines.append("(Auto-reject in 2 min if no reply)")
@@ -189,6 +187,25 @@ def _parse_approval_reply(text: str) -> str | None:
return None
def _approval_prompt_metadata(
base_metadata: dict | None, *, with_buttons: bool
) -> dict:
"""Outbound metadata for the HITL approval prompt.
When *with_buttons* is True, attaches Approve/Reject/Auto buttons whose
values match ``_parse_approval_reply`` so a click flows through the same
path as a typed ``"1"``/``"2"``/``"3"`` reply.
"""
metadata = dict(base_metadata or {})
if with_buttons:
metadata["buttons"] = [
{"text": "Approve", "value": "1", "type": "primary"},
{"text": "Reject", "value": "2", "type": "danger"},
{"text": "Approve all", "value": "3"},
]
return metadata
@dataclass
class _PendingInterrupt:
"""Stored state for a pending HITL interrupt awaiting channel user reply."""
@@ -217,9 +234,11 @@ class InboundConsumer:
manager:
The ChannelManager (used to look up channel instances).
agent:
The agent object (must support ``stream_agent_events``).
The local agent object used by local graph gateway targets.
thread_id:
Default thread ID for agent conversations.
graph_gateway:
Gateway used for thread creation and graph streaming.
send_thinking:
Whether to forward thinking messages to the channel.
on_message_received:
@@ -250,6 +269,7 @@ class InboundConsumer:
agent: Any,
thread_id: str,
*,
graph_gateway: GraphGateway,
send_thinking: bool = False,
on_message_received: Callable[[InboundMessage], None] | None = None,
on_streaming_event: Callable[[dict], None] | None = None,
@@ -263,6 +283,7 @@ class InboundConsumer:
self.manager = manager
self.agent = agent
self.thread_id = thread_id
self.graph_gateway = graph_gateway
self.send_thinking = send_thinking
self._on_message_received = on_message_received
self._on_streaming_event = on_streaming_event
@@ -296,7 +317,7 @@ class InboundConsumer:
# ask_user: pending reply per session_key
self._pending_ask_user_replies: dict[str, _PendingAskUserReply] = {}
def _get_thread_id(self, sender_id: str) -> str:
async def _get_thread_id(self, sender_id: str) -> str:
"""Get or create a thread ID for the given sender.
Uses LRU ordering: recently accessed senders are moved to the
@@ -312,7 +333,9 @@ class InboundConsumer:
if self.thread_id:
self._sessions[sender_id] = f"{self.thread_id}:{sender_id}"
else:
self._sessions[sender_id] = str(uuid.uuid4())
self._sessions[sender_id] = await self.graph_gateway.create_thread(
GraphTarget(local_graph=self.agent)
)
return self._sessions[sender_id]
def _get_channel(self, channel_name: str) -> Channel | None:
@@ -406,7 +429,7 @@ class InboundConsumer:
pass
channel = self._get_channel(msg.channel)
thread_id = self._get_thread_id(msg.sender_id)
thread_id = await self._get_thread_id(msg.sender_id)
session_key = msg.session_key # "channel:chat_id"
# Lazily create per-chat lock; evict stale locks when too many
@@ -449,15 +472,16 @@ class InboundConsumer:
session_key: str,
) -> None:
"""Stream agent events with HITL interrupt handling."""
from ..stream.events import stream_agent_events
from langgraph.types import Command
stream_input: Any = msg.content
_t0 = time.monotonic()
stream_input: GraphRunInput = msg.content
try:
if channel:
await channel.start_typing(msg.chat_id)
_last_sent_thinking: str | None = None
for _hitl_round in range(_MAX_HITL_ROUNDS):
final_content = ""
thinking_buffer: list[str] = []
@@ -466,14 +490,38 @@ class InboundConsumer:
thinking_sent = False
interrupt_data: dict | None = None
async def _flush_thinking_buffer(
buffer: list[str] = thinking_buffer,
) -> bool:
"""Send the current thinking buffer, dedup by content."""
nonlocal thinking_sent, _last_sent_thinking
if not channel or thinking_sent or not buffer:
return False
full_thinking = "".join(buffer).rstrip()
buffer.clear()
if not full_thinking or full_thinking == _last_sent_thinking:
return False
await channel.send_thinking_message(
msg.sender_id,
full_thinking,
msg.metadata,
)
thinking_sent = True
_last_sent_thinking = full_thinking
return True
async for event in _timeout_aiter(
stream_agent_events(
self.agent,
stream_input,
thread_id,
self.graph_gateway.stream_events(
RunRequest(
message=stream_input,
thread_id=thread_id,
media=msg.media or None
if isinstance(stream_input, str)
else None,
target=GraphTarget(local_graph=self.agent),
)
),
self._inference_timeout,
):
@@ -494,16 +542,7 @@ class InboundConsumer:
if event.get("name") == "write_todos" and not todo_sent:
todos = event.get("args", {}).get("todos", [])
if todos and channel:
if thinking_buffer and not thinking_sent:
full_thinking = "".join(thinking_buffer)
if full_thinking:
await channel.send_thinking_message(
msg.sender_id,
full_thinking,
msg.metadata,
)
thinking_sent = True
thinking_buffer.clear()
await _flush_thinking_buffer()
await channel.send_todo_message(
msg.sender_id,
_format_todo_list(todos),
@@ -514,14 +553,11 @@ class InboundConsumer:
elif event_type == "text":
final_content += event.get("content", "")
elif event_type == "progress":
# Internal agent planning — ignore for channel delivery.
# This text should NOT appear in the user-facing response.
pass
elif event_type == "subagent_text":
sa_name = event.get("subagent", "unknown")
instance_id = event.get("instance_id") or sa_name
instance_id = event.get("instance_id")
if not instance_id:
continue
if instance_id not in subagent_text_buffers:
subagent_text_buffers[instance_id] = (sa_name, [])
subagent_text_buffers[instance_id][1].append(
@@ -540,14 +576,7 @@ class InboundConsumer:
break # exit async for to handle ask_user
# Flush thinking
if thinking_buffer and not thinking_sent and channel:
full_thinking = "".join(thinking_buffer)
if full_thinking:
await channel.send_thinking_message(
msg.sender_id,
full_thinking,
msg.metadata,
)
await _flush_thinking_buffer()
# No interrupt — normal completion
if interrupt_data is None:
@@ -562,13 +591,6 @@ class InboundConsumer:
)
await self.bus.publish_outbound(outbound)
self._metrics.total_successes += 1
elapsed = time.monotonic() - _t0
logger.info(
"stream completed: %d chars, %.2fs, session=%s",
len(outbound.content),
elapsed,
session_key,
)
if self._on_message_sent:
try:
self._on_message_sent(outbound)
@@ -583,7 +605,6 @@ class InboundConsumer:
interrupt_data,
session_key,
)
from langgraph.types import Command # type: ignore[import-untyped]
stream_input = Command(resume=result)
continue
@@ -594,8 +615,6 @@ class InboundConsumer:
# Session auto-approve (user previously chose "Approve all")
if session_key in self._auto_approve_sessions:
from langgraph.types import Command # type: ignore[import-untyped]
stream_input = Command(
resume={"decisions": [{"type": "approve"} for _ in range(n)]}
)
@@ -603,21 +622,27 @@ class InboundConsumer:
# Config auto-approve (auto_approve, non-execute, allow_list)
if _should_auto_approve(action_reqs):
from langgraph.types import Command # type: ignore[import-untyped]
stream_input = Command(
resume={"decisions": [{"type": "approve"} for _ in range(n)]}
)
continue
# Needs user approval — send prompt to channel
prompt_text = _format_approval_prompt(action_reqs)
has_buttons = (
channel is not None and channel.capabilities.inline_buttons
)
prompt_text = _format_approval_prompt(
action_reqs, with_buttons=has_buttons
)
approval_metadata = _approval_prompt_metadata(
msg.metadata, with_buttons=has_buttons
)
await self.bus.publish_outbound(
OutboundMessage(
channel=msg.channel,
chat_id=msg.chat_id,
content=prompt_text,
metadata=msg.metadata,
metadata=approval_metadata,
)
)
@@ -629,35 +654,61 @@ class InboundConsumer:
)
self._pending_interrupts[session_key] = pending
timed_out = False
try:
await asyncio.wait_for(
pending.event.wait(),
timeout=_HITL_APPROVAL_TIMEOUT,
)
except TimeoutError:
# Auto-approve on timeout
pending.decision = "approve"
timed_out = True
finally:
# Unregister BEFORE any further await so a late reply can't flip
# the decision back to approve during the notification round-trip.
self._pending_interrupts.pop(session_key, None)
decision = pending.decision or "approve"
if decision == "reject":
if timed_out:
# Reject on timeout (fail-closed; matches cli/channel.py). Decision
# is a local constant, not pending.decision, so it can't be
# overwritten by a late reply after we unregistered above.
decision = "reject"
await self.bus.publish_outbound(
OutboundMessage(
channel=msg.channel,
chat_id=msg.chat_id,
content="Tool execution rejected.",
content="⏰ Approval timed out. Action rejected.",
metadata=msg.metadata,
)
)
else:
decision = pending.decision or "reject"
# Visible confirmation so the click/reply registers (QQ has no
# message recall API for C2C). Only fires when the user
# actually responded — silent on timeout to avoid claiming
# the user approved when they just walked away.
if pending.event.is_set():
feedback_text = {
"approve": "\u2705 已批准",
"auto": "\u2705 已批准(后续自动通过)",
"reject": "\u274c 已拒绝",
}.get(decision)
if feedback_text:
await self.bus.publish_outbound(
OutboundMessage(
channel=msg.channel,
chat_id=msg.chat_id,
content=feedback_text,
metadata=msg.metadata,
)
)
if decision == "reject":
return
if decision == "auto":
self._auto_approve_sessions.add(session_key)
from langgraph.types import Command # type: ignore[import-untyped]
stream_input = Command(
resume={"decisions": [{"type": "approve"} for _ in range(n)]}
)
@@ -665,14 +716,9 @@ class InboundConsumer:
except TimeoutError:
self._metrics.total_timeouts += 1
elapsed = time.monotonic() - _t0
logger.error(
"Inference timeout (%ds idle) for %s in %s, elapsed=%.2fs, response=%d chars",
self._inference_timeout,
msg.sender_id,
session_key,
elapsed,
len(final_content),
f"Inference timeout ({self._inference_timeout}s idle) "
f"for {msg.sender_id} in {session_key}"
)
await self.bus.publish_outbound(
OutboundMessage(
@@ -685,14 +731,7 @@ class InboundConsumer:
except Exception as e:
self._metrics.total_failures += 1
elapsed = time.monotonic() - _t0
logger.error(
"Agent error: %s | elapsed=%.2fs, response=%d chars, session=%s",
e,
elapsed,
len(final_content),
session_key,
)
logger.error(f"Agent error: {e}")
await self.bus.publish_outbound(
OutboundMessage(
channel=msg.channel,
+5 -2
View File
@@ -19,11 +19,15 @@ Examples:
import argparse
import logging
from ...logging_config import configure_logging_from_settings
from ..bus import MessageBus
from ..standalone import run_standalone
from .channel import DingTalkChannel, DingTalkConfig
logging.basicConfig(
level=logging.DEBUG,
format="%(asctime)s [%(levelname)s] %(name)s: %(message)s",
datefmt="%H:%M:%S",
)
logger = logging.getLogger(__name__)
@@ -68,7 +72,6 @@ def parse_args():
def main():
"""Entry point."""
configure_logging_from_settings(default_level=logging.INFO)
args = parse_args()
config = DingTalkConfig(
+5 -2
View File
@@ -19,11 +19,15 @@ Examples:
import argparse
import logging
from ...logging_config import configure_logging_from_settings
from ..bus import MessageBus
from ..standalone import run_standalone
from .channel import DiscordChannel, DiscordConfig
logging.basicConfig(
level=logging.DEBUG,
format="%(asctime)s [%(levelname)s] %(name)s: %(message)s",
datefmt="%H:%M:%S",
)
logger = logging.getLogger(__name__)
@@ -69,7 +73,6 @@ def parse_args():
def main():
"""Entry point."""
configure_logging_from_settings(default_level=logging.INFO)
args = parse_args()
config = DiscordConfig(
+5 -2
View File
@@ -19,11 +19,15 @@ Examples:
import argparse
import logging
from ...logging_config import configure_logging_from_settings
from ..bus import MessageBus
from ..standalone import run_standalone
from .channel import EmailChannel, EmailConfig
logging.basicConfig(
level=logging.DEBUG,
format="%(asctime)s [%(levelname)s] %(name)s: %(message)s",
datefmt="%H:%M:%S",
)
logger = logging.getLogger(__name__)
@@ -95,7 +99,6 @@ def parse_args():
def main():
"""Entry point."""
configure_logging_from_settings(default_level=logging.INFO)
args = parse_args()
config = EmailConfig(
+2 -1
View File
@@ -1,7 +1,8 @@
from ..channel_manager import _parse_csv, register_channel
from .channel import FeishuChannel, FeishuConfig
from .onboard import qr_register
__all__ = ["FeishuChannel", "FeishuConfig"]
__all__ = ["FeishuChannel", "FeishuConfig", "qr_register"]
def create_from_config(config) -> FeishuChannel:
+25
View File
@@ -376,6 +376,31 @@ class FeishuChannel(Channel, WebhookMixin, TokenMixin):
.build()
)
# Silently absorb events we don't have a handler for. Feishu auto-
# subscribes a PersonalAgent app to many event types (reactions,
# read receipts, recalls, member changes…) that EvoScientist doesn't
# care about. Without this wrapper, ``_do_without_validation``
# raises ``EventException("processor not found, type: ...")``,
# which lark-oapi's WS client (ws/client.py) catches and turns into
# an HTTP 500 reply on the WebSocket frame — Feishu then marks the
# event as failed and retries it. This is especially noisy because
# our own ``_send_ack_reaction`` triggers ``im.message.reaction.
# created_v1`` on every inbound message, causing a feedback loop.
from lark_oapi.core.exception import EventException
_original_dispatch = handler._do_without_validation
def _silent_dispatch(payload: bytes):
try:
return _original_dispatch(payload)
except EventException as exc:
if "processor not found" in str(exc):
logger.debug("Feishu: ignored unsubscribed event (%s)", exc)
return None
raise
handler._do_without_validation = _silent_dispatch
ws_client = lark.ws.Client(
self.config.app_id,
self.config.app_secret,
+361
View File
@@ -0,0 +1,361 @@
"""Feishu / Lark scan-to-create (QR code onboard) flow.
Drives the Feishu open-platform device-code flow at
``accounts.feishu.cn/oauth/v1/app/registration`` (and the Lark equivalent
at ``accounts.larksuite.com``). The user scans a terminal QR code with
Feishu / Lark mobile, the platform provisions a ``PersonalAgent``-archetype
bot application with the required IM permissions pre-attached, and the
poll endpoint returns ``client_id`` / ``client_secret`` — enough to fully
configure :class:`FeishuChannel`.
Domain auto-switches from ``feishu`` to ``lark`` if the poll response's
``user_info.tenant_brand`` reports a Lark tenant.
The HTTP shape mirrors RFC 8628 (OAuth Device Authorization Grant) with
a vendor-specific ``action`` form field selecting init / begin / poll.
Style follows :mod:`EvoScientist.channels.qq.onboard` — httpx, plain
``print`` for progress, and an optional ``qrcode`` dependency for ASCII
rendering.
"""
from __future__ import annotations
import logging
import time
from typing import Any
logger = logging.getLogger(__name__)
# ---------------------------------------------------------------------------
# Endpoints
# ---------------------------------------------------------------------------
_ACCOUNTS_URLS: dict[str, str] = {
"feishu": "https://accounts.feishu.cn",
"lark": "https://accounts.larksuite.com",
}
_OPEN_URLS: dict[str, str] = {
"feishu": "https://open.feishu.cn",
"lark": "https://open.larksuite.com",
}
_REGISTRATION_PATH = "/oauth/v1/app/registration"
_REQUEST_TIMEOUT_S = 10.0
_DEFAULT_POLL_INTERVAL_S = 5
_DEFAULT_EXPIRE_S = 600
def _accounts_base_url(domain: str) -> str:
return _ACCOUNTS_URLS.get(domain, _ACCOUNTS_URLS["feishu"])
def _open_base_url(domain: str) -> str:
return _OPEN_URLS.get(domain, _OPEN_URLS["feishu"])
# ---------------------------------------------------------------------------
# QR rendering
# ---------------------------------------------------------------------------
try:
import qrcode as _qrcode_mod
except (ImportError, TypeError):
_qrcode_mod = None # type: ignore[assignment]
def _render_qr(url: str) -> bool:
"""Render *url* as an ASCII QR in the terminal. Returns True on success."""
if _qrcode_mod is None:
return False
try:
qr = _qrcode_mod.QRCode(
error_correction=_qrcode_mod.constants.ERROR_CORRECT_M,
border=2,
)
qr.add_data(url)
qr.make(fit=True)
qr.print_ascii(invert=True)
return True
except Exception:
return False
# ---------------------------------------------------------------------------
# Registration HTTP
# ---------------------------------------------------------------------------
def _post_registration(base_url: str, body: dict[str, str]) -> dict:
"""POST form-encoded *body* to the registration endpoint.
The endpoint replies with JSON even on 4xx responses (``authorization_pending``
comes back as HTTP 400 with a parseable body), so we always read the body
and only fall back to raising if the bytes are missing or not JSON.
"""
import httpx
url = f"{base_url}{_REGISTRATION_PATH}"
headers = {"Content-Type": "application/x-www-form-urlencoded"}
with httpx.Client(timeout=_REQUEST_TIMEOUT_S, follow_redirects=True) as client:
resp = client.post(url, data=body, headers=headers)
# Don't raise_for_status — 4xx may still carry a usable JSON body.
try:
return resp.json()
except ValueError:
resp.raise_for_status() # re-raise underlying HTTP error
raise # pragma: no cover — raise_for_status already raised
def _init_registration(domain: str) -> None:
"""Probe the registration environment. Raises if client_secret auth is unavailable."""
res = _post_registration(_accounts_base_url(domain), {"action": "init"})
methods = res.get("supported_auth_methods") or []
if "client_secret" not in methods:
raise RuntimeError(
f"Feishu / Lark registration environment does not support "
f"client_secret auth (got: {methods})"
)
def _begin_registration(domain: str) -> dict:
"""Start the device-code flow.
Returns a dict with ``device_code``, ``qr_url``, ``user_code``,
``interval``, and ``expire_in``.
"""
res = _post_registration(
_accounts_base_url(domain),
{
"action": "begin",
"archetype": "PersonalAgent",
"auth_method": "client_secret",
"request_user_info": "open_id",
},
)
device_code = res.get("device_code")
if not device_code:
raise RuntimeError(
f"Feishu / Lark registration did not return a device_code: {res}"
)
qr_url = res.get("verification_uri_complete") or ""
sep = "&" if "?" in qr_url else "?"
qr_url = f"{qr_url}{sep}from=evoscientist&tp=evoscientist"
return {
"device_code": device_code,
"qr_url": qr_url,
"user_code": res.get("user_code", ""),
"interval": int(res.get("interval") or _DEFAULT_POLL_INTERVAL_S),
"expire_in": int(res.get("expire_in") or _DEFAULT_EXPIRE_S),
}
def _poll_registration(
*,
device_code: str,
interval: int,
expire_in: int,
domain: str,
) -> dict | None:
"""Poll until the user scans, or the device_code expires / is denied.
Auto-switches the polling domain to ``lark`` if the server reports
``user_info.tenant_brand == "lark"`` — the credentials only resolve
against the matching open-platform host.
Returns a dict with ``app_id``, ``app_secret``, ``domain``, ``open_id``
on success, or ``None`` on timeout / explicit denial.
"""
deadline = time.monotonic() + expire_in
current_domain = domain
domain_switched = False
poll_count = 0
while time.monotonic() < deadline:
try:
res = _post_registration(
_accounts_base_url(current_domain),
{
"action": "poll",
"device_code": device_code,
"tp": "ob_app",
},
)
except Exception as exc:
logger.debug("[Feishu onboard] poll request error: %s", exc)
time.sleep(interval)
continue
poll_count += 1
if poll_count == 1:
print(" Waiting for scan…", end="", flush=True)
elif poll_count % 6 == 0:
print(".", end="", flush=True)
# Domain auto-detection — the server may still return creds in
# this same poll, so we fall through rather than restarting.
user_info = res.get("user_info") or {}
if (
user_info.get("tenant_brand") == "lark"
and not domain_switched
and current_domain != "lark"
):
current_domain = "lark"
domain_switched = True
if res.get("client_id") and res.get("client_secret"):
print() # newline after the dots
return {
"app_id": res["client_id"],
"app_secret": res["client_secret"],
"domain": current_domain,
"open_id": user_info.get("open_id"),
}
error = res.get("error", "")
if error in {"access_denied", "expired_token"}:
print()
logger.warning("[Feishu onboard] Registration %s", error)
return None
# authorization_pending / slow_down / unknown — keep polling
time.sleep(interval)
print()
logger.warning("[Feishu onboard] Poll timed out after %ds", expire_in)
return None
# ---------------------------------------------------------------------------
# Bot probe (best-effort, uses tenant_access_token + /bot/v3/info)
# ---------------------------------------------------------------------------
def _probe_bot(app_id: str, app_secret: str, domain: str) -> dict | None:
"""Fetch bot name / bot_open_id via the open-platform REST API.
Best-effort: failures return ``None`` and the caller proceeds without
a friendly bot name. Uses raw HTTP so we don't require ``lark-oapi``
to be installed at onboard time (it's only needed for WebSocket mode).
"""
import httpx
base = _open_base_url(domain)
token_url = f"{base}/open-apis/auth/v3/tenant_access_token/internal"
info_url = f"{base}/open-apis/bot/v3/info"
try:
with httpx.Client(timeout=_REQUEST_TIMEOUT_S, follow_redirects=True) as client:
tok_resp = client.post(
token_url,
json={"app_id": app_id, "app_secret": app_secret},
)
tok_data = tok_resp.json()
if tok_data.get("code") != 0:
logger.debug("[Feishu onboard] token fetch failed: %s", tok_data)
return None
token = tok_data.get("tenant_access_token")
if not token:
return None
info_resp = client.get(
info_url,
headers={"Authorization": f"Bearer {token}"},
)
info_data = info_resp.json()
except Exception as exc:
logger.debug("[Feishu onboard] bot probe failed: %s", exc)
return None
if info_data.get("code") != 0:
return None
bot = info_data.get("bot") or info_data.get("data", {}).get("bot") or {}
return {
"bot_name": bot.get("app_name") or bot.get("bot_name"),
"bot_open_id": bot.get("open_id"),
}
# ---------------------------------------------------------------------------
# Public entry-point
# ---------------------------------------------------------------------------
def qr_register(
*,
initial_domain: str = "feishu",
timeout_seconds: int = 600,
) -> dict[str, Any] | None:
"""Run the Feishu / Lark scan-to-create QR registration flow.
Args:
initial_domain: ``"feishu"`` (default, mainland) or ``"lark"`` (overseas).
Auto-switches mid-flow if the scanning user is on the other tenant.
timeout_seconds: Wall-clock budget for the whole flow.
Returns on success::
{
"app_id": str,
"app_secret": str,
"domain": "feishu" | "lark",
"open_id": str | None,
"bot_name": str | None,
"bot_open_id": str | None,
}
Returns ``None`` on expected failures (network, denial, timeout).
"""
try:
return _qr_register_inner(
initial_domain=initial_domain,
timeout_seconds=timeout_seconds,
)
except Exception as exc:
logger.warning("[Feishu onboard] Registration failed: %s", exc)
return None
def _qr_register_inner(
*,
initial_domain: str,
timeout_seconds: int,
) -> dict[str, Any] | None:
print(" Connecting to Feishu / Lark…", end="", flush=True)
_init_registration(initial_domain)
begin = _begin_registration(initial_domain)
print(" done.")
print()
qr_url = begin["qr_url"]
if _render_qr(qr_url):
print(
f"\n Scan the QR code above with Feishu / Lark on your phone,\n"
f" or open this URL directly:\n {qr_url}"
)
else:
print(f" Open this URL in Feishu / Lark on your phone:\n\n {qr_url}\n")
print(
" Tip: pip install qrcode to display a scannable QR code here next time"
)
print()
result = _poll_registration(
device_code=begin["device_code"],
interval=begin["interval"],
expire_in=min(begin["expire_in"], timeout_seconds),
domain=initial_domain,
)
if not result:
return None
bot_info = _probe_bot(result["app_id"], result["app_secret"], result["domain"])
if bot_info:
result["bot_name"] = bot_info.get("bot_name")
result["bot_open_id"] = bot_info.get("bot_open_id")
else:
result["bot_name"] = None
result["bot_open_id"] = None
return result
+5 -2
View File
@@ -20,11 +20,15 @@ Examples:
import argparse
import logging
from ...logging_config import configure_logging_from_settings
from ..bus import MessageBus
from ..standalone import run_standalone
from .channel import FeishuChannel, FeishuConfig
logging.basicConfig(
level=logging.DEBUG,
format="%(asctime)s [%(levelname)s] %(name)s: %(message)s",
datefmt="%H:%M:%S",
)
logger = logging.getLogger(__name__)
@@ -92,7 +96,6 @@ def parse_args():
def main():
"""Entry point."""
configure_logging_from_settings(default_level=logging.INFO)
args = parse_args()
config = FeishuConfig(
-2
View File
@@ -19,7 +19,6 @@ Examples:
import argparse
import logging
from ...logging_config import configure_logging_from_settings
from ..bus import MessageBus
from ..standalone import run_standalone
from . import IMessageChannel, IMessageConfig
@@ -68,7 +67,6 @@ def parse_args():
def main():
"""Entry point."""
configure_logging_from_settings(default_level=logging.INFO)
args = parse_args()
config = IMessageConfig(
+10 -5
View File
@@ -75,11 +75,13 @@ class DedupCache:
max_size: int = _DEDUP_MAX,
trim_to: int = _DEDUP_TRIM,
ttl_seconds: float = _DEDUP_TTL,
clock: Callable[[], float] | None = None,
) -> None:
self._seen: OrderedDict[str, float] = OrderedDict()
self._max = max_size
self._trim = trim_to
self._ttl = ttl_seconds
self._clock = clock or time.monotonic
# ── public API ──────────────────────────────────────────────────
@@ -93,15 +95,16 @@ class DedupCache:
if not msg_id:
return False
self._prune()
now = self._clock()
self._prune(now)
if msg_id in self._seen:
# LRU: refresh position and timestamp
self._seen.move_to_end(msg_id)
self._seen[msg_id] = time.monotonic()
self._seen[msg_id] = now
return True
self._seen[msg_id] = time.monotonic()
self._seen[msg_id] = now
if len(self._seen) > self._max:
while len(self._seen) > self._trim:
self._seen.popitem(last=False)
@@ -118,9 +121,9 @@ class DedupCache:
# ── internal ────────────────────────────────────────────────────
def _prune(self) -> None:
def _prune(self, now: float | None = None) -> None:
"""Remove entries older than *ttl_seconds*."""
cutoff = time.monotonic() - self._ttl
cutoff = (self._clock() if now is None else now) - self._ttl
# OrderedDict is insertion-ordered; oldest entries are first.
while self._seen:
_key, ts = next(iter(self._seen.items()))
@@ -428,11 +431,13 @@ class DedupMiddleware(InboundMiddleware):
max_size: int = 1000,
trim_to: int = 500,
ttl_seconds: float = 3600.0,
clock: Callable[[], float] | None = None,
) -> None:
self._cache = DedupCache(
max_size=max_size,
trim_to=trim_to,
ttl_seconds=ttl_seconds,
clock=clock,
)
async def process_inbound(
+2 -1
View File
@@ -10,8 +10,9 @@ Usage in config:
from ..channel_manager import _parse_csv, register_channel
from .channel import QQChannel, QQConfig
from .onboard import qr_register
__all__ = ["QQChannel", "QQConfig"]
__all__ = ["QQChannel", "QQConfig", "qr_register"]
def create_from_config(config) -> QQChannel:
+211 -15
View File
@@ -26,6 +26,55 @@ except ImportError:
GroupMessage = None
# ── Inline keyboard (button) helpers ─────────────────────────────────
def _normalize_button(btn: dict) -> tuple[str, str] | None:
"""Return ``(label, value)`` for a button, or ``None`` if no label."""
label = (btn.get("text") or "").strip()
if not label:
return None
raw = btn.get("value")
return label, str(raw) if raw is not None else label
def _build_qq_keyboard(buttons: list[dict]) -> dict | None:
"""Build a QQ Bot keyboard payload (one button per row).
Render style: 1 = primary (blue), 0 = secondary (grey) — QQ has no danger.
``action.permission`` is required by the schema; ``type=2`` is harmless for
C2C (the click always comes from the DM peer). Returns ``None`` if no
button has a usable label.
"""
rows: list[dict] = []
for idx, btn in enumerate(buttons):
norm = _normalize_button(btn)
if norm is None:
continue
label, value = norm
style = 1 if btn.get("type") == "primary" else 0
rows.append(
{
"buttons": [
{
"id": btn.get("id") or f"btn_{idx}",
"render_data": {
"label": label,
"visited_label": label,
"style": style,
},
"action": {
"type": 1, # callback (server pushes interaction event)
"permission": {"type": 2},
"data": value,
},
}
]
}
)
return {"content": {"rows": rows}} if rows else None
@dataclass
class QQConfig(BaseChannelConfig):
app_id: str = ""
@@ -35,7 +84,11 @@ class QQConfig(BaseChannelConfig):
def _make_bot_class(channel: "QQChannel") -> "type[botpy.Client]":
"""Create a botpy Client subclass bound to the given channel."""
intents = botpy.Intents(public_messages=True, direct_message=True)
intents = botpy.Intents(
public_messages=True,
direct_message=True,
interaction=True, # button clicks → on_interaction_create
)
class _Bot(botpy.Client):
def __init__(self):
@@ -50,6 +103,9 @@ def _make_bot_class(channel: "QQChannel") -> "type[botpy.Client]":
async def on_group_at_message_create(self, message: "GroupMessage"):
await channel._on_msg(message, "group")
async def on_interaction_create(self, interaction):
await channel._on_interaction(interaction)
return _Bot
@@ -174,6 +230,76 @@ class QQChannel(Channel):
except Exception as e:
logger.error(f"Error handling QQ message: {e}")
async def _on_interaction(self, interaction) -> None:
"""Handle ``on_interaction_create`` (button click).
Surfaces the click as an :class:`InboundMessage` whose ``content`` is
the button's ``data`` verbatim — so a "1"/"approve"/… click flows
through ``_parse_approval_reply`` exactly like a typed reply.
The click runs through inbound middleware (Dedup suppresses QQ
retries) but is published directly to the bus so the per-sender
debounce buffer doesn't merge the click value with subsequent text.
Group-scope clicks are ignored (DM-only by design).
"""
# ACK first — QQ requires a response within ~5s or the button UI
# shows "expired". Code 0 just means "received"; downstream still
# decides the actual approval/rejection.
interaction_id = getattr(interaction, "id", "") or ""
if interaction_id and self._client:
try:
await self._client.api.on_interaction_result(interaction_id, 0)
except Exception as ack_exc:
logger.debug("QQ interaction ack failed: %s", ack_exc)
try:
user_openid = getattr(interaction, "user_openid", "") or ""
if not user_openid:
logger.debug("QQ interaction ignored (no user_openid; not C2C)")
return
resolved = getattr(getattr(interaction, "data", None), "resolved", None)
button_data = getattr(resolved, "button_data", "") or ""
button_id = getattr(resolved, "button_id", "") or ""
triggering_msg_id = getattr(resolved, "message_id", "") or ""
# QQ may serialize non-str values; coerce. Fall back to button id
# when no data — same path as a typed reply via _parse_approval_reply.
button_value = str(button_data) if button_data != "" else ""
text = button_value or button_id
# Stable id so DedupMiddleware suppresses any QQ retry callbacks.
message_id = (
f"{triggering_msg_id}:action:{interaction_id}"
if interaction_id
else f"qq_action:{datetime.now().timestamp()}"
)
raw = RawIncoming(
sender_id=user_openid,
chat_id=user_openid, # C2C: chat_id == user_openid
text=text,
timestamp=datetime.now(),
message_id=message_id,
metadata={
"chat_id": user_openid,
"msg_type": "c2c",
"event_id": triggering_msg_id,
"backend": "qq",
"button_click": True,
"button_id": button_id,
"button_value": button_value,
},
is_group=False,
was_mentioned=True,
)
inbound = await self._build_inbound_async(raw)
if inbound is not None and self._bus:
await self._bus.publish_inbound(inbound)
except Exception:
logger.exception("QQ interaction handler error")
# ── Send ──────────────────────────────────────────────────────
def _next_msg_seq(self, msg_id: str) -> int:
@@ -195,17 +321,81 @@ class QQChannel(Channel):
msg_type = (metadata or {}).get("msg_type", "c2c")
msg_id = (metadata or {}).get("event_id", "")
seq = self._next_msg_seq(msg_id)
# Inline keyboard is C2C-only here — group keyboards have stricter
# permission semantics and are out of scope for now.
buttons = (metadata or {}).get("buttons") if msg_type == "c2c" else None
keyboard = _build_qq_keyboard(buttons) if buttons else None
try:
await self._post_markdown_message(chat_id, raw_text, msg_type, msg_id, seq)
await self._post_markdown_message(
chat_id, raw_text, msg_type, msg_id, seq, keyboard=keyboard
)
return
except Exception as exc:
if not self._should_fallback_to_plain_text(exc):
logger.error(
"QQ markdown send failed with non-fallbackable error "
"(chat_id=%s, msg_id=%s, seq=%s): %r",
chat_id,
msg_id,
seq,
exc,
)
raise
self._record_markdown_fallback(chat_id, raw_text, exc)
logger.debug("QQ markdown send failed, falling back to plain text: %s", exc)
logger.warning(
"QQ markdown send failed, falling back to plain text "
"(chat_id=%s, msg_id=%s, seq=%s): %r",
chat_id,
msg_id,
seq,
exc,
)
# QQ may have already consumed `seq` server-side even on failure.
# Reusing it for the plain retry triggers "duplicate msg_seq", so
# always advance to a fresh seq before the fallback send.
fallback_seq = self._next_msg_seq(msg_id)
plain_text = self._plain_formatter.format(raw_text)
await self._post_plain_message(chat_id, plain_text, msg_type, msg_id, seq)
# Plain-text fallback can't carry a keyboard. Append `value=label`
# pairs so the user can still type "1"/"approve"/… instead of
# tapping (`_parse_approval_reply` accepts the same values).
if buttons:
pairs = []
for btn in buttons:
norm = _normalize_button(btn)
if norm is not None:
label, value = norm
pairs.append(f"{value}={label}")
if pairs:
plain_text = f"{plain_text}\n\nReply: {', '.join(pairs)}"
try:
await self._post_plain_message(
chat_id, plain_text, msg_type, msg_id, fallback_seq
)
except Exception as plain_exc:
logger.error(
"QQ plain fallback also failed (chat_id=%s, msg_id=%s, seq=%s): %r",
chat_id,
msg_id,
fallback_seq,
plain_exc,
)
raise
# QQ server-side error codes / fragments that indicate the markdown
# request itself is invalid (template not configured, format rejected,
# content audit, etc.). Seeing any of these means we should retry with
# plain text rather than re-raise.
_QQ_MARKDOWN_ERROR_MARKERS: ClassVar[tuple[str, ...]] = (
"304014", # markdown template not configured
"304003", # invalid markdown params
"40034059", # generic send message failed (often markdown-related)
"模板", # CN: template (standard form)
"模版", # CN: template (variant form)
"审核", # CN: audit
)
def _should_fallback_to_plain_text(self, exc: Exception) -> bool:
"""Return True only for markdown compatibility/validation failures."""
@@ -214,17 +404,20 @@ class QQChannel(Channel):
msg = str(exc).lower()
compatibility_tokens = ("unsupported", "unexpected", "unknown", "invalid")
return (
"unexpected keyword argument" in msg
or (
"markdown" in msg
and any(token in msg for token in compatibility_tokens)
)
or (
"msg_type" in msg
and any(token in msg for token in compatibility_tokens)
)
)
if "unexpected keyword argument" in msg:
return True
if "markdown" in msg and any(token in msg for token in compatibility_tokens):
return True
if "msg_type" in msg and any(token in msg for token in compatibility_tokens):
return True
# QQ-specific server error codes returned by qq-botpy as strings.
raw = str(exc)
for marker in self._QQ_MARKDOWN_ERROR_MARKERS:
if marker in raw or marker.lower() in msg:
return True
return False
def _record_markdown_fallback(
self,
@@ -254,6 +447,7 @@ class QQChannel(Channel):
msg_type: str,
msg_id: str,
seq: int,
keyboard: dict | None = None,
) -> None:
payload = {
"msg_type": 2,
@@ -261,6 +455,8 @@ class QQChannel(Channel):
"msg_id": msg_id,
"msg_seq": seq,
}
if keyboard is not None:
payload["keyboard"] = keyboard
if msg_type == "group":
await self._client.api.post_group_message(
group_openid=chat_id,
+49
View File
@@ -0,0 +1,49 @@
"""AES-256-GCM utilities for QQ Bot scan-to-configure credential decryption.
Ported from hermes-agent/gateway/platforms/qqbot/crypto.py — the q.qq.com
``create_bind_task`` / ``poll_bind_result`` flow uses AES-256-GCM to keep
the bot's *client_secret* off the wire in plaintext.
"""
from __future__ import annotations
import base64
import os
def generate_bind_key() -> str:
"""Generate a 256-bit random AES key, base64-encoded.
The key is sent to ``create_bind_task`` so the server can encrypt
the bot's *client_secret* before returning it. Only this client
holds the key, so the secret never travels in plaintext.
"""
return base64.b64encode(os.urandom(32)).decode()
def decrypt_secret(encrypted_base64: str, key_base64: str) -> str:
"""Decrypt a base64-encoded AES-256-GCM ciphertext.
Ciphertext layout (after base64-decoding)::
IV (12 bytes) ‖ ciphertext (N bytes) ‖ AuthTag (16 bytes)
Args:
encrypted_base64: The ``bot_encrypt_secret`` value returned by
``poll_bind_result``.
key_base64: The base64 AES key produced by :func:`generate_bind_key`.
Returns:
The decrypted *client_secret* as a UTF-8 string.
"""
from cryptography.hazmat.primitives.ciphers.aead import AESGCM
key = base64.b64decode(key_base64)
raw = base64.b64decode(encrypted_base64)
iv = raw[:12]
ciphertext_with_tag = raw[12:] # AESGCM expects ciphertext + tag concatenated
aesgcm = AESGCM(key)
plaintext = aesgcm.decrypt(iv, ciphertext_with_tag, None)
return plaintext.decode("utf-8")
+300
View File
@@ -0,0 +1,300 @@
"""QQ Bot scan-to-configure (QR code onboard) flow.
Ported from hermes-agent/gateway/platforms/qqbot/onboard.py.
Calls the ``q.qq.com`` ``create_bind_task`` / ``poll_bind_result`` APIs to
generate a QR code URL and poll for scan completion. On success the caller
receives the bot's *app_id*, *client_secret* (decrypted locally), and the
scanner's *user_openid* — enough to fully configure the QQ channel.
The bot must already be registered at https://q.qq.com — scanning binds
the QQ user (developer / admin) to the existing application; it does not
create a new one.
Reference: https://bot.q.qq.com/wiki/develop/api-v2/
"""
from __future__ import annotations
import logging
import os
import platform
import sys
import time
from enum import IntEnum
from urllib.parse import quote
from .crypto import decrypt_secret, generate_bind_key
logger = logging.getLogger(__name__)
# ---------------------------------------------------------------------------
# Endpoints / timing
# ---------------------------------------------------------------------------
# The portal domain is configurable for corporate proxies / sandbox routing.
PORTAL_HOST = os.getenv("QQ_PORTAL_HOST", "q.qq.com")
ONBOARD_CREATE_PATH = "/lite/create_bind_task"
ONBOARD_POLL_PATH = "/lite/poll_bind_result"
QR_URL_TEMPLATE = (
"https://q.qq.com/qqbot/openclaw/connect.html"
"?task_id={task_id}&_wv=2&source=evoscientist"
)
ONBOARD_API_TIMEOUT = 10.0
ONBOARD_POLL_INTERVAL = 2.0
_MAX_REFRESHES = 3
# ---------------------------------------------------------------------------
# Bind status
# ---------------------------------------------------------------------------
class BindStatus(IntEnum):
"""Status codes returned by ``poll_bind_result``."""
NONE = 0
PENDING = 1
COMPLETED = 2
EXPIRED = 3
# ---------------------------------------------------------------------------
# HTTP headers
# ---------------------------------------------------------------------------
def _get_evoscientist_version() -> str:
try:
from importlib.metadata import version
return version("evoscientist")
except Exception:
return "dev"
def _build_user_agent() -> str:
py_version = (
f"{sys.version_info.major}.{sys.version_info.minor}.{sys.version_info.micro}"
)
os_name = platform.system().lower()
return (
f"EvoScientistQQ/1.0.0 (Python/{py_version}; {os_name}; "
f"EvoScientist/{_get_evoscientist_version()})"
)
def _api_headers() -> dict[str, str]:
"""Standard HTTP headers for q.qq.com onboard API requests.
``q.qq.com`` requires ``Accept: application/json`` — without it,
the server returns a JavaScript anti-bot challenge page.
"""
return {
"Content-Type": "application/json",
"Accept": "application/json",
"User-Agent": _build_user_agent(),
}
# ---------------------------------------------------------------------------
# QR rendering
# ---------------------------------------------------------------------------
try:
import qrcode as _qrcode_mod
except (ImportError, TypeError):
_qrcode_mod = None # type: ignore[assignment]
def _render_qr(url: str) -> bool:
"""Render a QR code to the terminal. Returns True on success."""
if _qrcode_mod is None:
return False
try:
qr = _qrcode_mod.QRCode(
error_correction=_qrcode_mod.constants.ERROR_CORRECT_M,
border=2,
)
qr.add_data(url)
qr.make(fit=True)
qr.print_ascii(invert=True)
return True
except Exception:
return False
# ---------------------------------------------------------------------------
# HTTP helpers
# ---------------------------------------------------------------------------
def _create_bind_task(timeout: float = ONBOARD_API_TIMEOUT) -> tuple[str, str]:
"""Create a bind task and return *(task_id, aes_key_base64)*.
Raises:
RuntimeError: if the API returns a non-zero ``retcode``.
"""
import httpx
url = f"https://{PORTAL_HOST}{ONBOARD_CREATE_PATH}"
key = generate_bind_key()
with httpx.Client(timeout=timeout, follow_redirects=True) as client:
resp = client.post(url, json={"key": key}, headers=_api_headers())
resp.raise_for_status()
data = resp.json()
if data.get("retcode") != 0:
raise RuntimeError(data.get("msg", "create_bind_task failed"))
task_id = data.get("data", {}).get("task_id")
if not task_id:
raise RuntimeError("create_bind_task: missing task_id in response")
logger.debug("create_bind_task ok: task_id=%s", task_id)
return task_id, key
def _poll_bind_result(
task_id: str,
timeout: float = ONBOARD_API_TIMEOUT,
) -> tuple[BindStatus, str, str, str]:
"""Poll the bind result for *task_id*.
Returns:
``(status, bot_appid, bot_encrypt_secret, user_openid)``.
Raises:
RuntimeError: if the API returns a non-zero ``retcode``.
"""
import httpx
url = f"https://{PORTAL_HOST}{ONBOARD_POLL_PATH}"
with httpx.Client(timeout=timeout, follow_redirects=True) as client:
resp = client.post(url, json={"task_id": task_id}, headers=_api_headers())
resp.raise_for_status()
data = resp.json()
if data.get("retcode") != 0:
raise RuntimeError(data.get("msg", "poll_bind_result failed"))
d = data.get("data", {})
return (
BindStatus(d.get("status", 0)),
str(d.get("bot_appid", "")),
d.get("bot_encrypt_secret", ""),
d.get("user_openid", ""),
)
def build_connect_url(task_id: str) -> str:
"""Build the QR-code target URL for a given *task_id*."""
return QR_URL_TEMPLATE.format(task_id=quote(task_id))
# ---------------------------------------------------------------------------
# Public entry-point
# ---------------------------------------------------------------------------
def qr_register(timeout_seconds: int = 600) -> dict | None:
"""Run the QQ Bot scan-to-configure QR registration flow.
Handles create → display → poll → decrypt in one call. The QR
auto-refreshes up to ``_MAX_REFRESHES`` times if the user takes
too long to scan.
Args:
timeout_seconds: Total wall-clock budget across all refreshes.
Returns:
``{"app_id": ..., "client_secret": ..., "user_openid": ...}`` on
success, or ``None`` on failure / expiry / cancellation.
"""
deadline = time.monotonic() + timeout_seconds
for refresh_count in range(_MAX_REFRESHES + 1):
# ── Create bind task ──
try:
task_id, aes_key = _create_bind_task()
except Exception as exc:
logger.warning("[QQ onboard] Failed to create bind task: %s", exc)
return None
url = build_connect_url(task_id)
# ── Display QR code + URL ──
print()
if _render_qr(url):
print(f" Scan the QR code above, or open this URL on your phone:\n {url}")
else:
print(f" Open this URL in QQ on your phone:\n {url}")
print(" Tip: pip install qrcode to display a scannable QR code here")
print()
# ── Poll loop ──
consecutive_errors = 0
while time.monotonic() < deadline:
try:
status, app_id, encrypted_secret, user_openid = _poll_bind_result(
task_id
)
except Exception as exc:
consecutive_errors += 1
logger.warning(
"[QQ onboard] poll_bind_result failed (%d consecutive): %s",
consecutive_errors,
exc,
)
if consecutive_errors >= 5:
print(
"\n Repeated polling failures — aborting."
" See logs for details."
)
return None
time.sleep(ONBOARD_POLL_INTERVAL)
continue
consecutive_errors = 0
if status == BindStatus.COMPLETED:
try:
client_secret = decrypt_secret(encrypted_secret, aes_key)
except Exception as exc:
logger.warning("[QQ onboard] decrypt_secret failed: %s", exc)
return None
print()
print(f" QR scan complete! (App ID: {app_id})")
if user_openid:
print(f" Scanner's OpenID: {user_openid}")
return {
"app_id": app_id,
"client_secret": client_secret,
"user_openid": user_openid,
}
if status == BindStatus.EXPIRED:
if refresh_count >= _MAX_REFRESHES:
logger.warning(
"[QQ onboard] QR code expired %d times — giving up",
_MAX_REFRESHES,
)
return None
print(
f"\n QR code expired, refreshing... "
f"({refresh_count + 1}/{_MAX_REFRESHES})"
)
break # next outer iteration creates a new task
time.sleep(ONBOARD_POLL_INTERVAL)
else:
# deadline reached without completing
logger.warning("[QQ onboard] Poll timed out after %ds", timeout_seconds)
return None
return None
+5 -2
View File
@@ -19,11 +19,15 @@ Examples:
import argparse
import logging
from ...logging_config import configure_logging_from_settings
from ..bus import MessageBus
from ..standalone import run_standalone
from .channel import QQChannel, QQConfig
logging.basicConfig(
level=logging.DEBUG,
format="%(asctime)s [%(levelname)s] %(name)s: %(message)s",
datefmt="%H:%M:%S",
)
logger = logging.getLogger(__name__)
@@ -64,7 +68,6 @@ def parse_args():
def main():
"""Entry point."""
configure_logging_from_settings(default_level=logging.INFO)
args = parse_args()
config = QQConfig(
+5 -2
View File
@@ -19,11 +19,15 @@ Examples:
import argparse
import logging
from ...logging_config import configure_logging_from_settings
from ..bus import MessageBus
from ..standalone import run_standalone
from .channel import SignalChannel, SignalConfig
logging.basicConfig(
level=logging.DEBUG,
format="%(asctime)s [%(levelname)s] %(name)s: %(message)s",
datefmt="%H:%M:%S",
)
logger = logging.getLogger(__name__)
@@ -74,7 +78,6 @@ def parse_args():
def main():
"""Entry point."""
configure_logging_from_settings(default_level=logging.INFO)
args = parse_args()
config = SignalConfig(
+5 -2
View File
@@ -19,11 +19,15 @@ Examples:
import argparse
import logging
from ...logging_config import configure_logging_from_settings
from ..bus import MessageBus
from ..standalone import run_standalone
from .channel import SlackChannel, SlackConfig
logging.basicConfig(
level=logging.DEBUG,
format="%(asctime)s [%(levelname)s] %(name)s: %(message)s",
datefmt="%H:%M:%S",
)
logger = logging.getLogger(__name__)
@@ -74,7 +78,6 @@ def parse_args():
def main():
"""Entry point."""
configure_logging_from_settings(default_level=logging.INFO)
args = parse_args()
config = SlackConfig(
+3
View File
@@ -108,8 +108,10 @@ async def _async_main(
if use_agent:
logger.info("Loading EvoScientist agent...")
from ..EvoScientist import create_cli_agent
from ..gateway import create_runtime_gateways
agent = create_cli_agent()
runtime_gateways = create_runtime_gateways()
logger.info("Agent loaded")
consumer = InboundConsumer(
@@ -117,6 +119,7 @@ async def _async_main(
manager=manager,
agent=agent,
thread_id="",
graph_gateway=runtime_gateways.graph_gateway,
send_thinking=send_thinking,
)
manager.register_health_provider("consumer", lambda: consumer.metrics)
+5 -2
View File
@@ -19,11 +19,15 @@ Examples:
import argparse
import logging
from ...logging_config import configure_logging_from_settings
from ..bus import MessageBus
from ..standalone import run_standalone
from .channel import TelegramChannel, TelegramConfig
logging.basicConfig(
level=logging.DEBUG,
format="%(asctime)s [%(levelname)s] %(name)s: %(message)s",
datefmt="%H:%M:%S",
)
logger = logging.getLogger(__name__)
@@ -59,7 +63,6 @@ def parse_args():
def main():
"""Entry point."""
configure_logging_from_settings(default_level=logging.INFO)
args = parse_args()
config = TelegramConfig(
+48 -10
View File
@@ -5,13 +5,16 @@ Supports multiple WeChat backends:
— Most stable, pure HTTP, no third-party dependencies
- **wechatmp**: 微信公众号 (WeChat Official Account) via official API
— Pure HTTP webhook, suitable for public-facing bots
- **personal**: 个人微信 via Tencent's iLink Bot API
— Long-poll + AES-128-ECB CDN media protocol; QR-code login required.
Adapted from hermes-agent.
Both backends use httpx (already a core dependency) and receive messages
via HTTP webhook, send replies via REST API.
Backends 1+2 share the HTTP-webhook ``WeChatChannel``; backend 3 uses the
long-poll ``WeixinPersonalChannel``.
Usage in config:
channel_enabled = "wechat"
wechat_backend = "wecom" # or "wechatmp"
wechat_backend = "wecom" # or "wechatmp" or "personal"
# WeCom settings
wechat_wecom_corp_id = "..."
@@ -27,19 +30,54 @@ Usage in config:
wechat_mp_token = "..."
wechat_mp_encoding_aes_key = "..."
wechat_webhook_port = 9001
# OR: Personal WeChat (iLink Bot)
# First run `python -m EvoScientist.channels.wechat.serve --qr-login`
# to obtain an account_id + token via QR-code scan.
wechat_personal_account_id = "..."
wechat_personal_token = "..." # optional if persisted on disk
wechat_personal_dm_policy = "open" # open | allowlist
wechat_personal_group_policy = "disabled"
"""
from ..channel_manager import _parse_csv, register_channel
from .channel import WeChatChannel, WeChatMPConfig, WeComConfig
from .personal import WeixinPersonalChannel, WeixinPersonalConfig, qr_login
__all__ = ["WeChatChannel", "WeChatMPConfig", "WeComConfig"]
__all__ = [
"WeChatChannel",
"WeChatMPConfig",
"WeComConfig",
"WeixinPersonalChannel",
"WeixinPersonalConfig",
"qr_login",
]
def create_from_config(config) -> WeChatChannel:
backend = config.wechat_backend or "wecom"
allowed = _parse_csv(config.wechat_allowed_senders)
proxy = config.wechat_proxy or None
port = int(config.wechat_webhook_port or 9001)
def create_from_config(config):
"""Factory dispatched on ``config.wechat_backend``."""
backend = (getattr(config, "wechat_backend", "") or "wecom").lower()
allowed = _parse_csv(getattr(config, "wechat_allowed_senders", ""))
proxy = getattr(config, "wechat_proxy", "") or None
port = int(getattr(config, "wechat_webhook_port", 9001) or 9001)
if backend == "personal":
group_allowed = _parse_csv(getattr(config, "wechat_personal_group_allowed", ""))
cfg = WeixinPersonalConfig(
account_id=getattr(config, "wechat_personal_account_id", ""),
token=getattr(config, "wechat_personal_token", ""),
base_url=getattr(config, "wechat_personal_base_url", "")
or "https://ilinkai.weixin.qq.com",
cdn_base_url=getattr(config, "wechat_personal_cdn_base_url", "")
or "https://novac2c.cdn.weixin.qq.com/c2c",
dm_policy=getattr(config, "wechat_personal_dm_policy", "open") or "open",
group_policy=getattr(config, "wechat_personal_group_policy", "disabled")
or "disabled",
group_allowed_senders=group_allowed,
allowed_senders=allowed,
proxy=proxy,
)
return WeixinPersonalChannel(cfg)
if backend == "wechatmp":
mp_config = WeChatMPConfig(
@@ -52,7 +90,7 @@ def create_from_config(config) -> WeChatChannel:
proxy=proxy,
)
return WeChatChannel(mp_config, backend="wechatmp")
else:
wecom_config = WeComConfig(
corp_id=config.wechat_wecom_corp_id,
agent_id=config.wechat_wecom_agent_id,
+47
View File
@@ -83,6 +83,53 @@ def _aes_encrypt(key: bytes, iv: bytes, plaintext: bytes) -> bytes:
) from None
def aes128_ecb_decrypt(ciphertext: bytes, key: bytes) -> bytes:
"""AES-128-ECB decryption with PKCS#7 unpadding.
Used by the personal-WeChat (iLink) backend for CDN-encrypted media
payloads. Block size is 16; the WeChat CDN protocol pads with PKCS#7.
"""
if _HAS_PYCRYPTO:
cipher = AES.new(key, AES.MODE_ECB)
padded = cipher.decrypt(ciphertext)
else:
try:
import pyaes
decrypter = pyaes.Decrypter(pyaes.AESModeOfOperationECB(key))
padded = decrypter.feed(ciphertext)
padded += decrypter.feed()
except ImportError:
raise ImportError(
"WeChat CDN media decryption requires pycryptodome or pyaes. "
"Install with: pip install pycryptodome"
) from None
if not padded:
return padded
pad_len = padded[-1]
if 1 <= pad_len <= 16 and padded.endswith(bytes([pad_len]) * pad_len):
return padded[:-pad_len]
return padded
def parse_ilink_aes_key(aes_key_b64: str) -> bytes:
"""Parse the iLink CDN AES key.
iLink encodes the 16-byte key in two formats:
- direct base64 of 16 bytes
- base64 of a 32-char ASCII hex string (which decodes to 16 raw bytes)
"""
decoded = base64.b64decode(aes_key_b64)
if len(decoded) == 16:
return decoded
if len(decoded) == 32:
text = decoded.decode("ascii", errors="ignore")
if text and all(ch in "0123456789abcdefABCDEF" for ch in text):
return bytes.fromhex(text)
raise ValueError(f"unexpected aes_key format ({len(decoded)} decoded bytes)")
class WeChatCrypto:
"""Handles WeChat/WeCom message encryption and decryption.
File diff suppressed because it is too large Load Diff
+37
View File
@@ -70,3 +70,40 @@ async def validate_wechat_mp(
return False, f"Error ({data.get('errcode')}): {data.get('errmsg')}"
except Exception as e:
return False, f"Error: {e}"
async def validate_wechat_personal(
account_id: str,
token: str = "",
) -> tuple[bool, str]:
"""Validate that a personal WeChat (iLink) account has been logged in.
Personal-WeChat credentials are obtained via QR-code scan and persisted
on disk; there is no offline credential format the user can paste in.
This probe simply checks that:
- the account_id is set, and
- either *token* is supplied inline, or a saved-account file exists for
*account_id* under ``DATA_DIR/wechat_personal/accounts/``.
Online liveness is not checked because the iLink long-poll endpoint is
not designed for cheap probes.
"""
if not account_id:
return False, (
"account_id is required. Run "
"`python -m EvoScientist.channels.wechat.serve --qr-login` first."
)
if token:
return True, f"Personal WeChat account {account_id[:8]}… token provided"
from .personal import load_account
persisted = load_account(account_id)
if not persisted or not persisted.get("token"):
return False, (
f"No saved credentials for account_id={account_id[:8]}…. "
"Run `python -m EvoScientist.channels.wechat.serve --qr-login`."
)
return True, f"Personal WeChat account {account_id[:8]}… loaded from disk"
+81 -6
View File
@@ -20,6 +20,13 @@ Usage:
--token TOKEN \\
--aes-key AES_KEY
# Personal WeChat (个人微信 via iLink Bot)
# First, log in via QR scan to obtain credentials:
python -m EvoScientist.channels.wechat.serve --qr-login
# Then run with the saved account_id:
python -m EvoScientist.channels.wechat.serve \\
--backend personal --account-id <id>
Options:
--port PORT Webhook listen port (default: 9001)
--allow USER_ID Allowed sender (repeatable)
@@ -28,13 +35,24 @@ Options:
"""
import argparse
import asyncio
import logging
from ...logging_config import configure_logging_from_settings
from ..bus import MessageBus
from ..standalone import run_standalone
from .channel import WeChatChannel, WeChatMPConfig, WeComConfig
from .personal import (
WeixinPersonalChannel,
WeixinPersonalConfig,
load_account,
qr_login,
)
logging.basicConfig(
level=logging.DEBUG,
format="%(asctime)s [%(levelname)s] %(name)s: %(message)s",
datefmt="%H:%M:%S",
)
logger = logging.getLogger(__name__)
@@ -46,10 +64,15 @@ def parse_args():
)
parser.add_argument(
"--backend",
choices=["wecom", "wechatmp"],
choices=["wecom", "wechatmp", "personal"],
default="wecom",
help="WeChat backend type (default: wecom)",
)
parser.add_argument(
"--qr-login",
action="store_true",
help="Run interactive QR-code login for personal WeChat and exit",
)
parser.add_argument("--port", type=int, default=9001, help="Webhook port")
parser.add_argument(
"--allow",
@@ -85,6 +108,31 @@ def parse_args():
mp.add_argument("--app-id", default="", help="MP App ID")
mp.add_argument("--app-secret", default="", help="MP App Secret")
# Personal-WeChat settings
personal = parser.add_argument_group("Personal WeChat (iLink Bot)")
personal.add_argument(
"--account-id",
default="",
help="iLink account_id (obtained via --qr-login)",
)
personal.add_argument(
"--bot-token",
default="",
help="iLink bearer token; if omitted, loaded from disk via account-id",
)
personal.add_argument(
"--dm-policy",
choices=["open", "allowlist"],
default="open",
help="Direct-message policy (default: open)",
)
personal.add_argument(
"--group-policy",
choices=["open", "allowlist", "disabled"],
default="disabled",
help="Group-message policy (default: disabled — iLink rarely delivers)",
)
# Shared settings
parser.add_argument("--token", default="", help="Callback verification token")
parser.add_argument("--aes-key", default="", help="EncodingAESKey")
@@ -95,8 +143,14 @@ def parse_args():
def main():
"""Entry point."""
configure_logging_from_settings(default_level=logging.INFO)
args = parse_args()
if args.qr_login:
result = asyncio.run(qr_login())
if not result:
raise SystemExit(1)
return
allowed = set(args.allowed_senders) if args.allowed_senders else None
allowed_channels = set(args.allowed_channels) if args.allowed_channels else None
proxy = args.proxy or None
@@ -113,7 +167,8 @@ def main():
allowed_channels=allowed_channels,
proxy=proxy,
)
else:
channel = WeChatChannel(config, backend=args.backend)
elif args.backend == "wechatmp":
config = WeChatMPConfig(
app_id=args.app_id,
app_secret=args.app_secret,
@@ -124,11 +179,31 @@ def main():
allowed_channels=allowed_channels,
proxy=proxy,
)
channel = WeChatChannel(config, backend=args.backend)
else: # personal
token = args.bot_token
if not token and args.account_id:
persisted = load_account(args.account_id)
if persisted:
token = persisted.get("token", "")
if not args.account_id or not token:
raise SystemExit(
"Personal WeChat requires --account-id (and a saved token, "
"obtained via --qr-login)."
)
personal_config = WeixinPersonalConfig(
account_id=args.account_id,
token=token,
allowed_senders=allowed,
allowed_channels=allowed_channels,
dm_policy=args.dm_policy,
group_policy=args.group_policy,
proxy=proxy,
)
channel = WeixinPersonalChannel(personal_config)
send_thinking = args.thinking and args.agent
bus = MessageBus()
channel = WeChatChannel(config, backend=args.backend)
run_standalone(channel, bus, use_agent=args.agent, send_thinking=send_thinking)
+66 -26
View File
@@ -1,27 +1,67 @@
"""EvoScientist CLI package."""
"""EvoScientist CLI package.
# Backward-compat re-exports (tests import these from EvoScientist.cli)
from ..stream.state import ( # noqa: F401
StreamState,
SubAgentState,
_build_todo_stats,
_parse_todo_items,
)
Most re-exports are served lazily through ``__getattr__`` so that a bare
``import EvoScientist.cli`` only costs what ``main()`` actually needs. That
keeps ``evosci --help`` fast — the heavy chat-model/TUI/langgraph imports
only pay their cost when someone actually touches those names.
"""
from __future__ import annotations
from .. import deploy as _deploy_pkg # noqa: F401 — registers `deploy` @app.command
from . import commands # noqa: F401 — registers @app.command decorators
from ._app import app
from ._constants import WELCOME_SLOGANS # noqa: F401
from .agent import _deduplicate_run_name # noqa: F401
from .channel import _channels_is_running, _channels_stop # noqa: F401
# UI runtime re-exports (merged from former tui/ package)
from .tui_runtime import ( # noqa: F401
DEFAULT_UI_BACKEND,
SUPPORTED_UI_BACKENDS,
get_backend,
normalize_ui_backend,
resolve_ui_backend,
run_streaming,
)
__all__ = [
"DEFAULT_UI_BACKEND",
"SUPPORTED_UI_BACKENDS",
"WELCOME_SLOGANS",
"StreamState",
"SubAgentState",
"_build_todo_stats",
"_channels_is_running",
"_channels_stop",
"_deduplicate_run_name",
"_parse_todo_items",
"app",
"get_backend",
"main",
"normalize_ui_backend",
"resolve_ui_backend",
"run_streaming",
]
# Map attribute name -> (relative-module, attribute-in-module).
# Paths starting with ".." reach out of this package.
_LAZY_EXPORTS: dict[str, tuple[str, str]] = {
"StreamState": ("..stream.state", "StreamState"),
"SubAgentState": ("..stream.state", "SubAgentState"),
"_build_todo_stats": ("..stream.state", "_build_todo_stats"),
"_parse_todo_items": ("..stream.state", "_parse_todo_items"),
"WELCOME_SLOGANS": ("._constants", "WELCOME_SLOGANS"),
"_deduplicate_run_name": (".agent", "_deduplicate_run_name"),
"_channels_is_running": (".channel", "_channels_is_running"),
"_channels_stop": (".channel", "_channels_stop"),
"DEFAULT_UI_BACKEND": (".tui_runtime", "DEFAULT_UI_BACKEND"),
"SUPPORTED_UI_BACKENDS": (".tui_runtime", "SUPPORTED_UI_BACKENDS"),
"get_backend": (".tui_runtime", "get_backend"),
"normalize_ui_backend": (".tui_runtime", "normalize_ui_backend"),
"resolve_ui_backend": (".tui_runtime", "resolve_ui_backend"),
"run_streaming": (".tui_runtime", "run_streaming"),
}
def __getattr__(name: str):
target = _LAZY_EXPORTS.get(name)
if target is None:
raise AttributeError(f"module {__name__!r} has no attribute {name!r}")
from importlib import import_module
module_path, attr = target
module = import_module(module_path, package=__name__)
value = getattr(module, attr)
globals()[name] = value
return value
def main():
@@ -29,6 +69,12 @@ def main():
import os
import warnings
# Keep MCP stdio subprocess spawning async on Windows (see #283). Must run
# before any event loop is created, hence at the very top of the entrypoint.
from .._winloop import ensure_proactor_event_loop_policy
ensure_proactor_event_loop_policy()
warnings.filterwarnings("ignore", message=".*not known to support tools.*")
warnings.filterwarnings(
"ignore", message=".*type is unknown and inference may fail.*"
@@ -41,9 +87,3 @@ def main():
_log_level = os.environ.get("EVOSCIENTIST_LOG_LEVEL", "") or config.log_level
_configure_logging()
app()
def admin_main():
"""evo-admin CLI entry point — admin management commands."""
from ._app import admin_app
admin_app()
+199
View File
@@ -0,0 +1,199 @@
"""Background MCP/agent load lifecycle shared by CLI and TUI surfaces.
Holds no references to Rich, prompt_toolkit, or Textual — UI-specific
rendering and thread-hopping plug in via callbacks.
"""
from __future__ import annotations
import asyncio
import logging
from collections.abc import Callable
from typing import Any, Generic, TypeVar
_logger = logging.getLogger(__name__)
ProgressEvent = str # "start" | "success" | "error"
ProgressState = str # "pending" | "ok" | "error"
AgentT = TypeVar("AgentT")
ProgressCallback = Callable[[ProgressEvent, str, str], None]
SuccessCallback = Callable[[AgentT], None]
FailureCallback = Callable[[BaseException], None]
class MCPProgressTracker:
"""Per-server MCP load progress state.
Reads and writes are GIL-atomic but iteration must go through
:meth:`snapshot` — events can fire from a worker thread while the
main thread renders.
"""
__slots__ = ("progress",)
def __init__(self) -> None:
self.progress: dict[str, tuple[ProgressState, str]] = {}
def prime(self) -> None:
"""Seed a ``pending`` entry for every configured server.
Keeps the UI's "N / M" denominator stable from the first render.
"""
try:
from ..mcp import load_mcp_config
cfg = load_mcp_config() or {}
self.progress = dict.fromkeys(cfg, ("pending", ""))
except Exception:
self.progress = {}
def record(
self, event: ProgressEvent, server: str, detail: str
) -> ProgressState | None:
"""Apply an event and return the new state, or ``None`` if unknown."""
if event == "start":
self.progress.setdefault(server, ("pending", ""))
return "pending"
if event == "success":
self.progress[server] = ("ok", detail)
return "ok"
if event == "error":
self.progress[server] = ("error", detail)
return "error"
return None
def snapshot(self) -> list[tuple[ProgressState, str]]:
return list(self.progress.values())
def totals(self) -> tuple[int, int]:
"""``(done, total)`` — done excludes ``pending``."""
snap = self.snapshot()
total = len(snap)
done = sum(1 for state, _ in snap if state != "pending")
return done, total
class BackgroundAgentLoader(Generic[AgentT]):
"""Owns the background ``_load_agent`` task and its generation token.
Each :meth:`start` bumps an internal id; callbacks from a superseded
load (the old worker thread keeps running after cancel, since
``asyncio.to_thread`` can't preempt arbitrary Python code) compare
against it and drop silently.
``on_progress`` fires on the **worker thread**; UI callers hop
threads inside it if needed. ``on_success`` / ``on_failure`` fire
on the event loop when the task completes.
"""
def __init__(
self,
loader_fn: Callable[..., AgentT],
*,
on_progress: ProgressCallback | None = None,
on_success: SuccessCallback | None = None,
on_failure: FailureCallback | None = None,
) -> None:
self._loader_fn = loader_fn
self._on_progress = on_progress
self._on_success = on_success
self._on_failure = on_failure
self.agent: AgentT | None = None
self._task: asyncio.Task[AgentT] | None = None
self._load_id: int = 0
@property
def task(self) -> asyncio.Task[AgentT] | None:
return self._task
@property
def is_pending(self) -> bool:
return self.agent is None and self._task is not None and not self._task.done()
@property
def needs_restart(self) -> bool:
"""True when no load is in flight and no agent is ready.
Callers that want auto-retry behavior (e.g. TUI on the next
user send after a failure) check this before :meth:`start`.
"""
return self.agent is None and (self._task is None or self._task.done())
def start(self, **loader_kwargs: Any) -> None:
prev = self._task
if prev is not None and not prev.done():
prev.cancel()
self._load_id += 1
load_id = self._load_id
self.agent = None
def _gated_progress(event: str, server: str, detail: str) -> None:
if load_id != self._load_id:
return
if self._on_progress is None:
return
try:
self._on_progress(event, server, detail)
except Exception:
_logger.debug("MCP progress callback raised", exc_info=True)
self._task = asyncio.create_task(
asyncio.to_thread(
self._loader_fn,
on_mcp_progress=_gated_progress,
**loader_kwargs,
)
)
self._task.add_done_callback(lambda task, lid=load_id: self._on_done(task, lid))
def adopt(self, agent: AgentT) -> None:
"""Install an externally-built agent and supersede any in-flight load.
Used by any caller that constructs a replacement agent directly:
bumps the generation token so a late-arriving background load can't
clobber ``self.agent`` via the done-callback, cancels the in-flight
wrapper, and seats the new agent immediately.
"""
prev = self._task
if prev is not None and not prev.done():
prev.cancel()
self._load_id += 1
self._task = None
self.agent = agent
async def await_ready(self) -> AgentT:
"""Return the loaded agent; re-raises on load failure.
Idempotent. State transitions (setting ``self.agent``, calling
``on_success`` / ``on_failure``) are handled exclusively by
:meth:`_on_done`, which fires before this ``await`` resumes
(asyncio guarantees done-callbacks run in registration order).
"""
if self.agent is not None:
return self.agent
if self._task is None:
raise RuntimeError(
"BackgroundAgentLoader.await_ready called before start()"
)
await self._task
if self.agent is None:
raise RuntimeError("BackgroundAgentLoader completed without an agent")
return self.agent
def _on_done(self, task: asyncio.Task[AgentT], load_id: int) -> None:
if load_id != self._load_id:
return
if task.cancelled():
return
try:
self.agent = task.result()
except Exception as exc:
# Keep ``_task`` set so a later ``await_ready`` re-raises the
# real exception instead of the "before start()" sentinel.
self.agent = None
if self._on_failure is not None:
self._on_failure(exc)
return
if self._on_success is not None:
self._on_success(self.agent)
+15 -3
View File
@@ -52,6 +52,18 @@ app.add_typer(mcp_app, name="mcp")
channel_app = typer.Typer(help="Channel management commands")
app.add_typer(channel_app, name="channel")
# Admin subcommand group
admin_app = typer.Typer(help="Admin management commands")
app.add_typer(admin_app, name="admin")
# Sessions subcommand group — diagnostic tools for the LangGraph checkpoint DB
sessions_app = typer.Typer(
help="Inspect and manage the sessions DB (~/.evoscientist/sessions.db)",
invoke_without_command=True,
)
app.add_typer(sessions_app, name="sessions")
# Configure subcommand group — re-run a single onboarding section.
configure_app = typer.Typer(
help=(
"Re-run one onboarding section without going through the full wizard.\n"
"Example: EvoSci configure provider"
),
)
app.add_typer(configure_app, name="configure")
+15 -1
View File
@@ -2,8 +2,22 @@
from datetime import UTC, datetime
def _agent_name() -> str:
# Deferred import: ``sessions`` pulls in langgraph/aiosqlite (~300 ms)
# and is only needed when ``build_metadata`` is actually called.
from ..sessions import AGENT_NAME
return AGENT_NAME
# Dangerous-mode warning banner — shared by Rich CLI, Textual TUI, and serve so
# the wording never drifts. Label is rendered white-on-red, message in red.
DANGEROUS_BANNER_LABEL = "DANGEROUS MODE"
DANGEROUS_BANNER_MESSAGE = (
"Real-filesystem access • the agent can read/write/delete anywhere."
)
WELCOME_SLOGANS = [
"Ready for vibe research? What do you want cooking?",
"Science doesn't sleep. Neither do your sub-agents.",
@@ -34,7 +48,7 @@ LOGO_GRADIENT = ["#1a237e", "#1565c0", "#1e88e5", "#42a5f5", "#64b5f6", "#90caf9
def build_metadata(workspace_dir: str | None, model: str | None) -> dict:
"""Build metadata dict for LangGraph checkpoint persistence."""
return {
"agent_name": AGENT_NAME,
"agent_name": _agent_name(),
"updated_at": datetime.now(UTC).isoformat(),
"workspace_dir": workspace_dir or "",
"model": model or "",
+40 -3
View File
@@ -3,9 +3,13 @@
import os
from datetime import datetime
from pathlib import Path
from typing import TYPE_CHECKING
from ..paths import new_run_dir
if TYPE_CHECKING:
from langgraph.graph.state import CompiledStateGraph
def _shorten_path(path: str) -> str:
"""Shorten absolute path to relative path from current directory."""
@@ -58,18 +62,51 @@ def _create_session_workspace(name: str | None = None) -> str:
return workspace_dir
def _load_agent(workspace_dir: str | None = None, checkpointer=None, config=None):
def current_model_label() -> str:
"""Return the display label for the registry's primary default model.
The CLI no longer carries ``config.model``/``config.provider`` free
strings; the status bar and run metadata label the active model as
``provider_id/model_key`` from the Model Registry defaults. Returns
``"unconfigured"`` when the registry is still in bootstrap.
"""
try:
from ..model_registry.runtime import get_snapshot_runtime
primary, _ = get_snapshot_runtime().registry_default()
except Exception:
return "unconfigured"
return f"{primary.provider_id}/{primary.model_key}"
def _load_agent(
workspace_dir: str | None = None,
checkpointer=None,
config=None,
chat_model=None,
*,
on_mcp_progress=None,
) -> "CompiledStateGraph":
"""Load the CLI agent with optional persistent checkpointer.
Args:
workspace_dir: Optional per-session workspace directory.
checkpointer: Optional LangGraph checkpointer.
checkpointer: Optional LangGraph checkpointer (e.g. ``AsyncSqliteSaver``).
Falls back to ``InMemorySaver`` when ``None``.
config: Optional pre-loaded ``EvoScientistConfig``. Forwarded to
``create_cli_agent`` to avoid double config loading.
chat_model: Optional pre-built chat model. Forwarded to
``create_cli_agent``; combined with an explicit ``config`` it
selects the pure (no module-global write) build path.
on_mcp_progress: Optional per-server MCP progress callback.
Signature ``(event, server_name, detail) -> None``.
"""
from ..EvoScientist import create_cli_agent
return create_cli_agent(
workspace_dir=workspace_dir, checkpointer=checkpointer, config=config
workspace_dir=workspace_dir,
checkpointer=checkpointer,
config=config,
chat_model=chat_model,
on_mcp_progress=on_mcp_progress,
)
+595
View File
@@ -0,0 +1,595 @@
"""Async sub-agent auto-notification.
When a sub-agent on langgraph dev reaches a terminal state, a watcher coroutine
pushes a lightweight notification onto a thread-safe queue. The CLI loop drains
the queue, dedups against deepagents' async_tasks state, batches survivors,
and injects a synthetic user message that triggers one LLM turn.
"""
from __future__ import annotations
import asyncio
import json
import logging
import queue
import threading
from collections.abc import Awaitable, Callable
from dataclasses import dataclass
from datetime import UTC, datetime
from typing import TYPE_CHECKING, Final, TypeAlias, TypedDict
if TYPE_CHECKING:
from ..gateway import GraphGateway, GraphTarget
TERMINAL_STATUSES: Final = frozenset({"success", "error", "timeout", "interrupted"})
"""Aligned with langgraph_sdk.schema.RunStatus terminal values.
Cancel operations transition runs into ``interrupted`` (not ``cancelled``).
"""
# How many times the watcher will re-join the SSE stream when it closes
# cleanly but ``runs.get`` reports the run is still alive (typical cause:
# HTTP keep-alive timeout on long static periods). Bounded to prevent an
# unbounded loop if the server permanently misreports status.
_MAX_RECONNECT_ATTEMPTS: Final = 10
class AsyncTaskState(TypedDict, total=False):
status: str
last_checked_at: str
last_updated_at: str
AsyncTasksState: TypeAlias = dict[str, AsyncTaskState]
@dataclass(frozen=True)
class AsyncTaskNotification:
"""A completed-async-task signal pushed by a watcher."""
task_id: str
agent_name: str
status: str # one of TERMINAL_STATUSES
received_at: str # ISO-8601 UTC timestamp
prompt: str = "" # original task description sent to the sub-agent
kind: str = "agent" # "agent" (sub-agent) | "bg-process" (background shell)
# The CLI/main-agent thread_id under which the watcher was spawned. Used
# to route the notification back to the originating CLI session so a
# /new between launch and completion does not inject the synthetic
# message into an unrelated thread (where ``check_async_task`` cannot
# find the task_id). ``None`` means "unrouted" — the notification
# drains for any current_thread_id (back-compat for direct callers).
origin_cli_thread_id: str | None = None
# Per-thread routing: notifications with ``origin_cli_thread_id`` land in
# the matching sub-queue. Notifications without one go to ``_unrouted_queue``
# and drain regardless of current thread (back-compat for legacy callers
# and direct-put test paths).
_notifications_by_thread: dict[str, queue.Queue[AsyncTaskNotification]] = {}
_notifications_lock = threading.Lock()
_unrouted_queue: queue.Queue[AsyncTaskNotification] = queue.Queue()
# Public alias for the unrouted bucket — preserved so legacy tests and any
# external direct callers that did ``_notification_queue.put(...)`` keep
# working unchanged. New code should call ``_enqueue`` instead.
_notification_queue = _unrouted_queue
# Track active watcher tasks/futures for clean shutdown.
# dict[handle, origin_cli_thread_id] so the consumer's batching grace loop
# can filter for watchers tied to the current CLI thread (or unrouted)
# without being delayed by sibling-thread watchers.
_active_watchers: dict[object, str | None] = {}
# Map thread_id (sub-agent thread) → current watcher handle (supports
# replacement on update_async_task).
_watcher_by_thread: dict[str, asyncio.Task[None]] = {}
def _has_relevant_active_watchers(current_thread_id: str | None) -> bool:
"""Are there any in-flight watchers whose notifications would drain on
a ``consume_notifications`` call for ``current_thread_id``?
A watcher is relevant if its ``origin_cli_thread_id`` matches the
current CLI thread or is ``None`` (unrouted bucket drains for any
consumer). Sibling-thread watchers are ignored.
"""
if current_thread_id is None:
return bool(_active_watchers)
return any(
origin == current_thread_id or origin is None
for origin in _active_watchers.values()
)
logger = logging.getLogger(__name__)
def _enqueue(notification: AsyncTaskNotification) -> None:
"""Route a notification to its origin-thread queue or the unrouted bucket."""
tid = notification.origin_cli_thread_id
if not tid:
_unrouted_queue.put(notification)
return
with _notifications_lock:
q = _notifications_by_thread.get(tid)
if q is None:
q = queue.Queue()
_notifications_by_thread[tid] = q
q.put(notification)
def has_pending_notifications(current_thread_id: str | None = None) -> bool:
"""Cheap predicate for poller idle paths — true iff there's anything to consume.
If ``current_thread_id`` is given, only the matching thread queue and
the unrouted bucket count. With no argument, only the unrouted bucket
counts (legacy behavior).
"""
if not _unrouted_queue.empty():
return True
if current_thread_id is None:
return False
with _notifications_lock:
q = _notifications_by_thread.get(current_thread_id)
return q is not None and not q.empty()
def pending_thread_ids() -> set[str]:
"""Return the set of thread_ids with pending routed notifications."""
with _notifications_lock:
return {tid for tid, q in _notifications_by_thread.items() if not q.empty()}
async def read_async_tasks_from_gateway(
gateway: GraphGateway,
target: GraphTarget,
thread_id: str,
) -> AsyncTasksState:
"""Read async_tasks state through the active graph gateway."""
try:
values = await gateway.get_state_values(target, thread_id)
except Exception:
return {}
return values.get("async_tasks", {})
async def watch_run_and_notify(
client,
thread_id: str,
run_id: str,
agent_name: str,
prompt: str = "",
origin_cli_thread_id: str | None = None,
) -> None:
"""Subscribe to a run's event stream; enqueue notification when it terminates.
Status detection strategy (priority order):
1. **In-band ``event="error"`` SSE part** — authoritative error signal
from langgraph dev, no race against server-side state writeback.
2. **Server-side state via ``runs.get``** — invoked after the stream
closes (cleanly or with exception) to verify the run is actually
done. Required because SSE long-poll can close on HTTP keep-alive
timeout while the run is still running, which would otherwise be
misread as ``"success"`` (observed in production with long-running
literature search tasks under concurrency).
3. **Re-join loop** — if ``runs.get`` reports ``pending`` / ``running``,
the run is alive but we lost the stream; re-join up to
``_MAX_RECONNECT_ATTEMPTS`` times before giving up.
The previous implementation trusted clean stream exits as success
without any verification, which produced false-positive notifications
when SSE keep-alive timeouts closed the stream early.
Race-safety note: ``runs.get`` returning ``"error"`` immediately after
a clean stream close can be a transient state for an actually-successful
run (server hasn't finalized the writeback). We trust the absence of
in-band error event over a stale ``runs.get="error"`` — see the
``status == "error" and not saw_error_event`` branch below.
"""
for attempt in range(_MAX_RECONNECT_ATTEMPTS + 1):
stream_failed = False
saw_error_event = False
try:
async for chunk in client.runs.join_stream(
thread_id=thread_id, run_id=run_id, stream_mode="values"
):
ev = getattr(chunk, "event", None)
data = getattr(chunk, "data", None)
if ev == "error":
saw_error_event = True
logger.info(
"Watcher saw error event for task %s: %r", thread_id, data
)
except Exception:
stream_failed = True
logger.warning(
"Watcher stream failed for task %s", thread_id, exc_info=True
)
if saw_error_event:
status = "error"
break
# Verify with server before deciding the run is done — clean stream
# close does NOT guarantee terminal state.
try:
run = await client.runs.get(thread_id=thread_id, run_id=run_id)
raw = run.get("status", "")
except Exception:
# Cannot verify terminal state. Defaulting to "success" here would
# reintroduce the false-positive class this watcher exists to
# prevent (clean stream + transient runs.get failure → unverified
# success). Retry within the reconnect budget; on exhaustion drop
# the notification rather than guess.
if attempt >= _MAX_RECONNECT_ATTEMPTS:
logger.warning(
"Watcher runs.get failed for task %s after %d reconnects; "
"unable to verify terminal state, skipping notification",
thread_id,
_MAX_RECONNECT_ATTEMPTS,
exc_info=True,
)
return
logger.warning(
"Watcher runs.get failed for task %s; retrying after backoff "
"(attempt %d)",
thread_id,
attempt + 1,
exc_info=True,
)
await asyncio.sleep(min(0.25 * (attempt + 1), 2.0))
continue
if raw not in TERMINAL_STATUSES:
# Non-terminal status — includes the documented ``pending`` /
# ``running`` values AND any future / unknown status the SDK may
# introduce. Stream closed early but run is not done; re-join
# unless we've exhausted attempts. Treating unknown statuses as
# non-terminal is the safe default — better to retry once more
# than to enqueue a false-positive on an unrecognized state.
if attempt >= _MAX_RECONNECT_ATTEMPTS:
logger.warning(
"Watcher gave up on task %s after %d reconnects "
"(server still reports %r); skipping notification",
thread_id,
_MAX_RECONNECT_ATTEMPTS,
raw,
)
return
logger.info(
"Watcher SSE closed for task %s but run reports %r; "
"re-joining (attempt %d)",
thread_id,
raw,
attempt + 1,
)
continue
if raw == "error":
# Race-safe interpretation: no in-band error event → trust the
# absence over the server-side ``error`` (likely transient
# writeback state for a successful run). Stream-failure path
# is the one case where we DO trust ``error`` — the stream
# blowing up usually means something genuinely went wrong.
status = "error" if stream_failed else "success"
break
# success / timeout / interrupted — trust authoritative terminal status.
status = raw
break
else:
# Loop exhausted without a break — should be unreachable because the
# re-join branch returns explicitly when attempts are exhausted, but
# guard against future refactors.
return
notification = AsyncTaskNotification(
task_id=thread_id,
agent_name=agent_name,
status=status,
received_at=datetime.now(UTC).strftime("%Y-%m-%dT%H:%M:%SZ"),
prompt=prompt,
origin_cli_thread_id=origin_cli_thread_id,
)
_enqueue(notification)
logger.info(
"Enqueued async notification: task=%s agent=%s status=%s origin_thread=%s",
thread_id,
agent_name,
status,
origin_cli_thread_id or "<unrouted>",
)
def spawn_watcher(
client,
thread_id: str,
run_id: str,
agent_name: str,
prompt: str = "",
origin_cli_thread_id: str | None = None,
) -> asyncio.Task[None]:
"""Spawn a watcher on the caller's asyncio loop.
Replacement semantics support ``update_async_task`` which creates a new
run_id on the same thread_id — we want the new watcher to take over
without the old (now obsolete) watcher firing a stale notification.
Cancellation propagates ``CancelledError`` (a BaseException), which the
watcher's ``except Exception:`` does NOT catch — so ``_enqueue(...)``
never executes for the cancelled watcher (no stale notification).
``origin_cli_thread_id`` tags the resulting notification so the consumer
only injects it back into the originating CLI session.
Caller must already be in a running asyncio event loop. Serve mode's
ephemeral per-turn loop kills watchers spawned during a turn — that
limitation is tracked separately.
"""
old_task = _watcher_by_thread.get(thread_id)
if old_task is not None and not old_task.done():
old_task.cancel()
task = asyncio.create_task(
watch_run_and_notify(
client,
thread_id,
run_id,
agent_name,
prompt,
origin_cli_thread_id=origin_cli_thread_id,
)
)
_watcher_by_thread[thread_id] = task
_active_watchers[task] = origin_cli_thread_id
def _cleanup(t: asyncio.Task[None]) -> None:
_active_watchers.pop(t, None)
# Only remove if THIS task is still the registered one — could
# have been replaced by a newer spawn_watcher call already.
if _watcher_by_thread.get(thread_id) is t:
del _watcher_by_thread[thread_id]
task.add_done_callback(_cleanup)
return task
def _drain_one_queue(q: queue.Queue) -> list[AsyncTaskNotification]:
items: list[AsyncTaskNotification] = []
while True:
try:
items.append(q.get_nowait())
except queue.Empty:
return items
def drain_notifications(
current_thread_id: str | None = None,
) -> list[AsyncTaskNotification]:
"""Pull pending notifications off the queue (non-blocking).
With ``current_thread_id``: drains the matching per-thread queue plus
the unrouted bucket. Without it: drains EVERY queue (legacy behavior;
used by tests and diagnostics).
"""
if current_thread_id is None:
items: list[AsyncTaskNotification] = _drain_one_queue(_unrouted_queue)
with _notifications_lock:
queues = list(_notifications_by_thread.values())
for q in queues:
items.extend(_drain_one_queue(q))
return items
items = _drain_one_queue(_unrouted_queue)
with _notifications_lock:
q = _notifications_by_thread.get(current_thread_id)
if q is not None:
items.extend(_drain_one_queue(q))
return items
def dedup_notifications(
notifs: list[AsyncTaskNotification],
async_tasks: AsyncTasksState | None,
) -> list[AsyncTaskNotification]:
"""Filter notifications the agent has already 'seen' via prior check.
Logic: skip a notification if `async_tasks[task_id]` exists with a TERMINAL
status and `last_checked_at >= last_updated_at` (timestamps are ISO-8601
so lexicographic comparison is correct). Also skip if `last_checked_at`
is empty (brand-new task where agent hasn't checked yet).
"""
from .. import background # cli -> core import; lazy to avoid import-order issues
async_tasks = async_tasks or {}
survivors: list[AsyncTaskNotification] = []
for n in notifs:
if n.kind == "bg-process":
# Background process: skip if the launching session already inspected it
# after it finished (check_process / list_processes) — mirrors the task
# dedup below. Per-thread: another session's check doesn't suppress this.
if background.was_observed_done(n.task_id, n.origin_cli_thread_id):
logger.debug("Dedup: skipping shell notification for %s", n.task_id)
continue
survivors.append(n)
continue
task = async_tasks.get(n.task_id)
if (
task
and task.get("status") in TERMINAL_STATUSES
and task.get("last_checked_at", "") >= task.get("last_updated_at", "")
and task.get("last_checked_at", "") != ""
):
logger.debug(
"Dedup: skipping notification for already-checked task %s", n.task_id
)
continue
survivors.append(n)
return survivors
def _render_notification_group(
notifs: list[AsyncTaskNotification], title: str, label: str
) -> list[tuple[str, str]]:
"""Render one group of notifications inside a titled open-right frame.
Open-right compact frame; bottom matches the top's width:
╭── ✦ Agent Teams ✦ ────
✔ writing Task: ... success
╰─────────────────────────
"""
top_divider = "╭──" + title + "────" # 4 dashes on the right (2x of left)
bottom_divider = "╰" + "─" * (len(top_divider) - 1)
lines: list[tuple[str, str]] = [(top_divider, "dim")]
for n in notifs:
# `writing-agent` → `writing`.
name = n.agent_name.removesuffix("-agent")
if n.status == "success":
icon, color = "✔", "#e67e22" # carrot orange (CSS hex; Rich+Textual)
elif n.status == "error":
icon, color = "✗", "red"
else: # cancelled, timeout, interrupted
icon, color = "⚠", "yellow"
# Collapse newlines, truncate prompt/command preview to 60 chars.
prompt_preview = (n.prompt or "").replace("\n", " ").strip()
if len(prompt_preview) > 60:
prompt_preview = prompt_preview[:60] + "…"
if prompt_preview:
text = f" {icon} {name:18s} {label}: {prompt_preview} {n.status}"
else:
# Fallback: short task_id when no prompt is available
short_tid = (
f"{n.task_id[:8]}…{n.task_id[-4:]}"
if len(n.task_id) > 12
else n.task_id
)
text = f" {icon} {name:18s} ({short_tid}) {n.status}"
lines.append((text, color))
lines.append((bottom_divider, "dim"))
return lines
def format_notification_lines(
notifs: list[AsyncTaskNotification],
) -> list[tuple[str, str]]:
"""Render notifications as compact tool-result-style lines for screen display.
Async sub-agents and background processes get SEPARATE titled frames so a shell
background process is never mislabeled as an "Agent Team". Returns (text, rich_style)
tuples. The LLM still receives the full ``format_batch_message`` text; this is purely
the visual representation for the human operator.
"""
if not notifs:
return []
tasks = [n for n in notifs if n.kind == "agent"]
shell = [n for n in notifs if n.kind == "bg-process"]
unknown = [n for n in notifs if n.kind not in {"agent", "bg-process"}]
lines: list[tuple[str, str]] = []
if tasks:
lines += _render_notification_group(tasks, " ✦ Agent Teams ✦ ", "Task")
if shell:
lines += _render_notification_group(shell, " ✦ Background ✦ ", "Cmd")
if unknown:
# Fallback so a future kind is never silently dropped from the display.
lines += _render_notification_group(unknown, " ✦ Updates ✦ ", "Task")
return lines
def format_batch_message(notifs: list[AsyncTaskNotification]) -> str:
"""Compose the synthetic user message that wakes the supervisor.
Each task is rendered as a compact JSON object (one per line) so the LLM
can reliably parse agent name, status, and task_id without ambiguity.
``ensure_ascii=False`` lets non-ASCII agent names pass through unchanged.
Visual decoration lives in ``format_notification_lines``.
"""
if not notifs:
return ""
lines = ["[Async tasks update]"]
for n in notifs:
lines.append(
json.dumps(
{
"agent": n.agent_name,
"kind": n.kind,
"status": n.status,
"task_id": n.task_id,
},
ensure_ascii=False,
)
)
# bg-process is inspected with check_process; sub-agents with check_async_task.
hints: list[str] = []
if any(n.kind == "agent" for n in notifs):
hints.append("check_async_task (sub-agents)")
if any(n.kind == "bg-process" for n in notifs):
hints.append("check_process (background processes)")
# Fallback when a batch has only unrecognized kinds (hints empty).
hint_text = " or ".join(hints) if hints else "the appropriate status tool"
lines.append(
f"(Signal only — fetch full result via {hint_text} if relevant to "
"the current step, else acknowledge & continue.)"
)
return "\n".join(lines)
# Brief grace window after the last drain: catch one final burst of arrivals
NOTIFICATION_BATCH_GRACE_SECONDS = 0.3
# Max time we'll wait for in-flight watchers to settle before triggering the
# agent turn — bounds latency for long-running tasks while still batching
# co-completing ones.
NOTIFICATION_ACTIVE_WATCHER_WAIT_SECONDS = 3.0
async def consume_notifications(
run_message: Callable[[str, list[AsyncTaskNotification]], Awaitable[None]],
read_async_tasks_state: Callable[[], Awaitable[AsyncTasksState]],
current_thread_id: str | None = None,
) -> None:
"""Drain queue, dedup, batch, and inject as a synthetic user message.
Args:
run_message: async callable receiving (llm_text, notifs_list).
``llm_text`` is the full structured message for the LLM
(from ``format_batch_message``). ``notifs_list`` is the
survivors list so callers can render per-task visual lines
without re-parsing the text.
read_async_tasks_state: async callable returning current ``async_tasks``
from the agent's state for dedup.
current_thread_id: the active CLI thread id. When given, only
notifications whose ``origin_cli_thread_id`` matches (or that
were enqueued unrouted) are drained — notifications belonging
to other threads stay queued and naturally drain on the next
poller tick after the user ``/resume``s back into them. When
omitted (legacy callers / tests), every queue drains.
"""
notifs = drain_notifications(current_thread_id)
if not notifs:
return
# Adaptive grace: if other watchers tied to THIS thread (or unrouted) are
# still in flight, wait briefly for them to settle so co-completing tasks
# batch into a single agent turn. Sibling-thread watchers don't count —
# their notifications wouldn't drain on this tick anyway.
loop = asyncio.get_running_loop()
deadline = loop.time() + NOTIFICATION_ACTIVE_WATCHER_WAIT_SECONDS
while _has_relevant_active_watchers(current_thread_id) and loop.time() < deadline:
await asyncio.sleep(0.2)
notifs.extend(drain_notifications(current_thread_id))
# Final brief grace to catch arrivals enqueued just before this tick
await asyncio.sleep(NOTIFICATION_BATCH_GRACE_SECONDS)
notifs.extend(drain_notifications(current_thread_id))
try:
async_tasks = await read_async_tasks_state()
except Exception:
logger.warning("Failed to read async_tasks state for dedup", exc_info=True)
async_tasks = {}
survivors = dedup_notifications(notifs, async_tasks)
if not survivors:
logger.info(
"All %d notifications deduped (already known to agent)", len(notifs)
)
return
text = format_batch_message(survivors)
await run_message(text, survivors)
+587 -147
View File
@@ -9,20 +9,26 @@ enqueues a ``ChannelMessage`` on a thread-safe ``queue.Queue`` and waits
for the main thread to set a response via ``_set_channel_response()``.
"""
from __future__ import annotations
import asyncio
import logging
import queue
import threading
import time
import uuid
from collections.abc import Awaitable, Callable
from dataclasses import dataclass
from typing import Any
from typing import TYPE_CHECKING, Any
from rich.panel import Panel
from rich.table import Table
from rich.text import Text
from ..stream.display import console
from ..commands.base import ChannelRuntime
from ..stream.console import console
if TYPE_CHECKING:
from ..gateway import GraphGateway
_channel_logger = logging.getLogger(__name__)
@@ -59,6 +65,10 @@ _response_lock = threading.Lock()
_RESPONSE_TIMEOUT = 600.0
_LATE_RESPONSE_TIMEOUT = 86400.0
_LATE_RESPONSE_NOTICE = "Still working on it. I'll send the result when it's ready."
_channel_request_lock = threading.Lock()
_channel_requests: dict[str, dict[str, str]] = {}
_session_requests: dict[str, list[str]] = {}
_cancelled_channel_messages: set[str] = set()
def _enqueue_channel_message(msg: ChannelMessage) -> asyncio.Future[str]:
@@ -71,6 +81,7 @@ def _enqueue_channel_message(msg: ChannelMessage) -> asyncio.Future[str]:
"loop": loop,
"response": None,
}
_register_channel_request(msg)
_message_queue.put(msg)
return future
@@ -106,6 +117,336 @@ def _pop_channel_response(msg_id: str, *, cancel_pending: bool = False) -> str |
return slot["response"]
def _channel_session_key(channel_type: str, chat_id: str) -> str:
return f"{channel_type}:{chat_id}"
def _channel_message_session_key(msg: ChannelMessage) -> str:
return _channel_session_key(msg.channel_type, msg.chat_id)
def _channel_message_cancel_scope(msg: ChannelMessage) -> str:
return f"channel:{msg.channel_type}:{msg.chat_id}:{msg.msg_id}"
def _register_channel_request(msg: ChannelMessage) -> None:
"""Track a queued channel request so `/stop` can find it later."""
session_key = _channel_message_session_key(msg)
with _channel_request_lock:
_channel_requests[msg.msg_id] = {
"session_key": session_key,
"cancel_scope": _channel_message_cancel_scope(msg),
"state": "queued",
}
_session_requests.setdefault(session_key, []).append(msg.msg_id)
def _claim_channel_request(msg: ChannelMessage) -> bool:
"""Mark a queued request active. Returns False if it was cancelled first."""
with _channel_request_lock:
slot = _channel_requests.get(msg.msg_id)
if slot is None or msg.msg_id in _cancelled_channel_messages:
return False
slot["state"] = "active"
return True
def _claim_or_complete_channel_request(msg: ChannelMessage) -> bool:
"""Claim a request, or clean it up if `/stop` cancelled it while queued."""
if _claim_channel_request(msg):
return True
_complete_channel_request(msg.msg_id)
return False
def _channel_request_state(msg_id: str) -> str | None:
with _channel_request_lock:
slot = _channel_requests.get(msg_id)
return slot.get("state") if slot is not None else None
def _complete_channel_request(
msg_id: str,
*,
discard_cancel_scope: bool = True,
) -> None:
"""Forget a request once its waiter is resolved or cancelled."""
with _channel_request_lock:
slot = _channel_requests.pop(msg_id, None)
_cancelled_channel_messages.discard(msg_id)
if slot is not None:
request_ids = _session_requests.get(slot["session_key"])
if request_ids:
try:
request_ids.remove(msg_id)
except ValueError:
pass
if not request_ids:
_session_requests.pop(slot["session_key"], None)
if slot is not None and discard_cancel_scope:
from ..stream.display import discard_stream_cancel
discard_stream_cancel(slot["cancel_scope"])
def _cancel_channel_session(channel_type: str, chat_id: str) -> tuple[int, int]:
"""Cancel queued and active work for one channel chat session."""
session_key = _channel_session_key(channel_type, chat_id)
with _channel_request_lock:
request_ids: list[str] = []
cancelled_ids: list[str] = []
active_scopes: list[str] = []
with _response_lock:
for msg_id in tuple(_session_requests.get(session_key, ())):
request_slot = _channel_requests.get(msg_id)
if request_slot is None:
continue
response_slot = _pending_responses.get(msg_id)
response_resolved = False
if response_slot is not None:
future = response_slot["future"]
# Once a response is already resolved, leave the slot alone
# so the bus waiter can still publish it instead of falling
# back to "No response".
response_resolved = (
response_slot.get("response") is not None or future.done()
)
if not response_resolved:
request_ids.append(msg_id)
should_cancel = False
if response_slot is None:
should_cancel = request_slot.get("state") == "active"
else:
should_cancel = not response_resolved
if should_cancel:
cancelled_ids.append(msg_id)
if request_slot.get("state") == "active" and should_cancel:
active_scopes.append(request_slot["cancel_scope"])
_cancelled_channel_messages.update(cancelled_ids)
for msg_id in request_ids:
_pop_channel_response(msg_id, cancel_pending=True)
if active_scopes:
from ..stream.display import request_stream_cancel
for cancel_scope in active_scopes:
request_stream_cancel(cancel_scope)
return len(request_ids), len(active_scopes)
# ---------------------------------------------------------------------------
# Slash command dispatch for channel messages
# ---------------------------------------------------------------------------
# Shared by all three UI surfaces that accept inbound channel messages:
# Rich CLI (``cli/interactive.py::_process_channel_message``), Textual
# TUI (``cli/tui_interactive.py``'s channel handler), and headless
# serve (``cli/commands.py::_serve_process_message``). They all route
# ``/foo`` text through ``cmd_manager`` instead of feeding it to the
# LLM as a plain prompt.
async def dispatch_channel_slash_command(
msg: ChannelMessage,
*,
agent: Any,
thread_id: str,
workspace_dir: str | None,
checkpointer: Any,
append_system: Callable[[str, str], None],
graph_gateway: GraphGateway,
start_new_session_cb: Callable[[], Awaitable[None]] | None = None,
handle_session_resume_cb: Callable[..., Awaitable[None]] | None = None,
await_agent_ready: Callable[[], Awaitable[Any]] | None = None,
on_cmd_completed: Callable[..., Awaitable[None]] | None = None,
channel_runtime: ChannelRuntime | None = None,
) -> bool:
"""Dispatch a slash command from a channel message.
Returns True if the helper handled the message (successfully or with
an error) — the caller must then return without streaming anything
to the agent. Returns False for non-slash content or unresolved
slash commands, so the caller can fall through to the agent
streaming path (matches TUI behavior).
Parameters
----------
msg:
The inbound ``ChannelMessage`` to inspect.
agent:
Default agent handle for the ``CommandContext``. Commands that
do not need the agent use this value directly.
thread_id, workspace_dir, checkpointer:
Populate ``CommandContext``.
append_system:
``(text, style)`` callback for local CLI/TUI log output. Used
by ``ChannelCommandUI`` to surface system breadcrumbs and by
this helper to print the "Executed command from ..." line.
start_new_session_cb, handle_session_resume_cb:
Optional lifecycle callbacks forwarded to ``ChannelCommandUI``.
Headless serve passes ``None`` — ``/new`` and ``/resume`` degrade
gracefully via the default ``ChannelCommandUI`` messages.
graph_gateway:
Graph gateway forwarded to slash commands and channel resume-history
rendering.
await_agent_ready:
Optional async resolver that blocks until the background agent
load finishes. Called only when ``cmd.needs_agent(args)`` is
True. Headless serve passes ``None`` because the agent is
loaded up-front before the bus starts.
on_cmd_completed:
Optional ``async (ctx, original_agent, cmd) -> None`` callback
fired only after ``cmd_manager.execute`` returns True. The
``original_agent`` argument is the agent handle command execution
started against: ``agent_for_ctx`` after any ``await_agent_ready``
resolution, or the dispatcher's input agent when no resolver is
supplied. Callers can compare ``ctx.agent`` with
``original_agent`` to detect command-driven swaps. Used by Rich
CLI to (a) adopt an agent swap back into the
running session and (b) refresh the status snapshot for
commands that mutate session-level state (``/new``,
``/compact``) — mirrors the REPL dispatch at
``cli/interactive.py:1002-1030``. Headless serve passes
``None`` since it cannot hot-swap its polling-loop agent.
"""
if not msg.content.strip().startswith("/"):
return False
try:
return await _dispatch_channel_slash_impl(
msg,
agent=agent,
thread_id=thread_id,
workspace_dir=workspace_dir,
checkpointer=checkpointer,
append_system=append_system,
start_new_session_cb=start_new_session_cb,
handle_session_resume_cb=handle_session_resume_cb,
await_agent_ready=await_agent_ready,
on_cmd_completed=on_cmd_completed,
channel_runtime=channel_runtime,
graph_gateway=graph_gateway,
)
except Exception as exc:
# Last-ditch safety: any uncaught exception from inside the
# dispatch pipeline (lazy import failure, ChannelCommandUI
# construction, terminal I/O from ``append_system``, bus
# publish races, ...) must not take down the caller's polling
# loop — a crashed serve / dead channel queue task is worse
# than one failed command.
_channel_logger.exception(
"Unexpected slash dispatch failure for %s (msg=%s)",
msg.channel_type,
msg.msg_id,
)
try:
_set_channel_response(msg.msg_id, f"Command error: {exc}")
except Exception: # pragma: no cover — defensive
pass
# Return True so the caller treats the message as handled and
# does not fall through to the agent streaming path.
return True
async def _dispatch_channel_slash_impl(
msg: ChannelMessage,
*,
agent: Any,
thread_id: str,
workspace_dir: str | None,
checkpointer: Any,
append_system: Callable[[str, str], None],
graph_gateway: GraphGateway,
start_new_session_cb: Callable[[], Awaitable[None]] | None,
handle_session_resume_cb: Callable[..., Awaitable[None]] | None,
await_agent_ready: Callable[[], Awaitable[Any]] | None,
on_cmd_completed: Callable[..., Awaitable[None]] | None,
channel_runtime: ChannelRuntime | None,
) -> bool:
"""Inner body of ``dispatch_channel_slash_command``.
Split from the public wrapper so the wrapper can guard with a
top-level try/except without visually obscuring the main flow.
"""
# Lazy imports: avoid coupling the channel module to ``commands`` at
# import time (tui_interactive.py does the same).
from ..commands.base import CommandContext
from ..commands.channel_ui import ChannelCommandUI
from ..commands.manager import manager as cmd_manager
parsed = cmd_manager.resolve(msg.content)
if parsed is None:
# Unknown slash command — let the agent handle it (matches TUI).
return False
cmd, cmd_args = parsed
agent_for_ctx = agent
if cmd.needs_agent(cmd_args) and await_agent_ready is not None:
try:
agent_for_ctx = await await_agent_ready()
except Exception as exc:
_set_channel_response(msg.msg_id, f"Command error: {exc}")
return True
ui = ChannelCommandUI(
msg,
append_system_callback=append_system,
start_new_session_callback=start_new_session_cb,
handle_session_resume_callback=handle_session_resume_cb,
graph_gateway=graph_gateway,
)
ctx = CommandContext(
agent=agent_for_ctx,
thread_id=thread_id,
ui=ui,
workspace_dir=workspace_dir,
checkpointer=checkpointer,
channel_runtime=channel_runtime,
graph_gateway=graph_gateway,
)
try:
cmd_executed = await cmd_manager.execute(msg.content, ctx)
except Exception as exc:
_channel_logger.debug(f"Channel command error: {exc}", exc_info=True)
_set_channel_response(msg.msg_id, f"Command error: {exc}")
return True # must return — do NOT fall through to the agent
if cmd_executed:
if ctx.command_error is not None:
details = ctx.command_error or "(no details)"
_set_channel_response(msg.msg_id, f"Command error: {details}")
return True
if on_cmd_completed is not None:
try:
# Command output already flushed by ``cmd_manager.execute``
# via ``ctx.ui.flush()`` — the hook does internal state
# sync (agent adoption, status snapshot refresh) only,
# so swallowing its errors keeps the user-visible reply
# intact even if the sync path is broken.
await on_cmd_completed(ctx, agent_for_ctx, cmd)
except Exception as exc:
_channel_logger.debug(
f"Channel command post-exec callback error: {exc}",
exc_info=True,
)
append_system(
f"[{msg.channel_type}: Executed command from {msg.sender}]",
"dim",
)
_set_channel_response(msg.msg_id, f"Command executed: {msg.content}")
return True
# ``cmd_manager.execute`` returned False (empty / unparseable input).
# Fall through to the agent streaming path.
return False
# ---------------------------------------------------------------------------
# HITL approval intercept: bus thread ⇄ main CLI thread
# ---------------------------------------------------------------------------
@@ -121,6 +462,143 @@ _HITL_APPROVAL_TIMEOUT = 120.0 # seconds to wait for HITL approval reply
_ASK_USER_TIMEOUT = (
300.0 # seconds to wait for ask_user reply (longer for thinking time)
)
_STOP_COMMANDS = frozenset(("/stop", "/cancel"))
# ---------------------------------------------------------------------------
# Per-thread channel-origin registry
# ---------------------------------------------------------------------------
# When a channel-originated message starts an agent turn, the turn's
# thread_id is remembered against its (channel_type, chat_id, metadata).
# Later, when an async sub-agent notification fires a synthetic agent turn
# for that same thread_id, the notifier path pushes the synthesized final
# response back to the same chat — otherwise the follow-up would only render
# locally and the channel user would never see it. v1 forwards only the
# final response (no mid-turn thinking/todo/media).
@dataclass(frozen=True)
class _ChannelOrigin:
"""Channel destination remembered for a thread, for notifier push-back."""
channel_type: str
chat_id: str
sender: str
metadata: dict | None = None
_thread_channel_origins: dict[str, _ChannelOrigin] = {}
_thread_channel_origins_lock = threading.Lock()
def remember_channel_origin(thread_id: str | None, msg: ChannelMessage) -> None:
"""Record that ``thread_id`` is currently bound to ``msg``'s channel chat.
Called on entry to each channel-triggered agent turn (Rich CLI / TUI /
serve). The latest channel turn for a given thread wins — re-registering
is intentional, since the user can keep talking on the same thread from
the same channel and we always want the most recent metadata.
"""
if not thread_id:
return
with _thread_channel_origins_lock:
_thread_channel_origins[thread_id] = _ChannelOrigin(
channel_type=msg.channel_type,
chat_id=msg.chat_id,
sender=msg.sender,
metadata=dict(msg.metadata) if msg.metadata else None,
)
def get_channel_origin(thread_id: str | None) -> _ChannelOrigin | None:
"""Return the channel origin remembered for ``thread_id``, or ``None``."""
if not thread_id:
return None
with _thread_channel_origins_lock:
return _thread_channel_origins.get(thread_id)
def forget_channel_origin(thread_id: str | None) -> None:
"""Drop the registry entry for ``thread_id`` (e.g. on ``/new`` rotation)."""
if not thread_id:
return
with _thread_channel_origins_lock:
_thread_channel_origins.pop(thread_id, None)
def publish_to_channel_origin(thread_id: str | None, content: str) -> bool:
"""Schedule pushing ``content`` to the channel remembered for ``thread_id``.
Fire-and-forget: returns ``True`` iff a publish coroutine was scheduled
on the bus loop; returns ``False`` if no origin is registered, the bus
isn't running, ``content`` is empty/whitespace, or scheduling itself
fails. The publish runs asynchronously — failures inside the coroutine
are logged via a done-callback so callers (which are often on event
loops that must not block) don't pay any latency.
"""
from ..channels.bus.events import OutboundMessage
if not content or not content.strip():
return False
origin = get_channel_origin(thread_id)
if origin is None:
return False
loop = _bus_loop
manager = _manager
if loop is None or manager is None:
return False
bus = getattr(manager, "bus", None)
if bus is None:
return False
async def _publish_and_record() -> None:
await bus.publish_outbound(
OutboundMessage(
channel=origin.channel_type,
chat_id=origin.chat_id,
content=content,
metadata=origin.metadata or {},
)
)
# Mirror the normal channel-reply path, which records a "sent"
# message after a successful publish so per-channel stats stay
# accurate for forwarded notifications too.
manager.record_message(origin.channel_type, "sent")
try:
future = asyncio.run_coroutine_threadsafe(_publish_and_record(), loop)
except Exception as exc:
_channel_logger.warning(
"Async notification publish to %s:%s failed to schedule: %s",
origin.channel_type,
origin.chat_id,
exc,
)
return False
def _on_publish_done(fut) -> None:
"""Log any exception raised by the fire-and-forget publish coroutine."""
# A cancelled future raises CancelledError from .exception() rather
# than returning it (e.g. bus loop torn down mid-publish); treat that
# as a benign shutdown, not a failure to log.
if fut.cancelled():
return
exc = fut.exception()
if exc is not None:
_channel_logger.warning(
"Async notification publish to %s:%s failed: %s",
origin.channel_type,
origin.chat_id,
exc,
)
future.add_done_callback(_on_publish_done)
return True
def _is_stop_command(content: str | None) -> bool:
"""Whether incoming content is a stop/cancel slash command."""
return (content or "").strip().lower() in _STOP_COMMANDS
def _register_hitl_wait(channel_type: str, chat_id: str) -> threading.Event:
@@ -154,7 +632,7 @@ def _try_set_hitl_reply(channel_type: str, chat_id: str, content: str) -> bool:
def channel_ask_user_prompt(
ask_user_data: dict,
msg: "ChannelMessage | None" = None,
msg: ChannelMessage | None = None,
) -> dict:
"""Format ask_user questions and collect answers from a channel user.
@@ -186,7 +664,7 @@ def channel_ask_user_prompt(
channel=msg.channel_type,
chat_id=msg.chat_id,
content=content,
metadata=msg.metadata,
metadata=msg.metadata or {},
)
),
bus_loop,
@@ -243,6 +721,8 @@ def channel_ask_user_prompt(
return {"status": "cancelled"}
raw = reply_text.strip()
if _is_stop_command(raw):
return {"status": "cancelled"}
if raw.lower() == "cancel":
return {"status": "cancelled"}
@@ -260,6 +740,8 @@ def channel_ask_user_prompt(
if not replied or not other_text:
_send("\u23f0 Response timed out.")
return {"status": "cancelled"}
if _is_stop_command(other_text):
return {"status": "cancelled"}
if other_text.strip().lower() == "cancel":
return {"status": "cancelled"}
answers.append(other_text.strip())
@@ -279,7 +761,7 @@ def channel_ask_user_prompt(
def channel_hitl_prompt(
action_requests: list,
msg: "ChannelMessage",
msg: ChannelMessage,
) -> list[dict] | None:
"""Send HITL approval prompt to channel user and wait for reply.
@@ -290,6 +772,7 @@ def channel_hitl_prompt(
"""
from ..channels.bus.events import OutboundMessage
from ..channels.consumer import (
_approval_prompt_metadata,
_format_approval_prompt,
_parse_approval_reply,
)
@@ -304,7 +787,17 @@ def channel_hitl_prompt(
_channel_logger.debug("HITL: no bus_loop or bus_ref, rejecting")
return None
def _send(content: str) -> bool:
# Look up the channel instance so we can attach buttons when the channel
# supports `inline_buttons` (Feishu cards, QQ keyboards, …).
channel_obj = (
_manager.get_channel(msg.channel_type) if _manager is not None else None
)
has_buttons = channel_obj is not None and channel_obj.capabilities.inline_buttons
approval_metadata = _approval_prompt_metadata(
msg.metadata, with_buttons=has_buttons
)
def _send(content: str, *, metadata: dict | None = None) -> bool:
"""Send a message to the channel user. Returns True on success."""
try:
asyncio.run_coroutine_threadsafe(
@@ -313,7 +806,9 @@ def channel_hitl_prompt(
channel=msg.channel_type,
chat_id=msg.chat_id,
content=content,
metadata=msg.metadata,
metadata=metadata
if metadata is not None
else msg.metadata or {},
)
),
bus_loop,
@@ -324,8 +819,8 @@ def channel_hitl_prompt(
return False
# 1. Send approval prompt
prompt_text = _format_approval_prompt(action_requests)
if not _send(prompt_text):
prompt_text = _format_approval_prompt(action_requests, with_buttons=has_buttons)
if not _send(prompt_text, metadata=approval_metadata):
return None
# 2. Wait for channel user's reply
@@ -337,16 +832,24 @@ def channel_hitl_prompt(
_send("\u23f0 Approval timed out. Action rejected.")
return None
if _is_stop_command(reply_text):
# `/stop` already got its own immediate ack from the bus fast-path.
# Treat it as a pure cancel signal here so we don't send a second,
# contradictory "Unrecognized reply" message.
return None
# 3. Parse decision
decision = _parse_approval_reply(reply_text)
if decision == "auto":
_hitl_auto_approve.add(session_key)
_send("\u2705 已批准(后续自动通过)")
return [{"type": "approve"} for _ in action_requests]
if decision == "approve":
_send("\u2705 已批准")
return [{"type": "approve"} for _ in action_requests]
feedback = (
"Action rejected."
"\u274c 已拒绝"
if decision == "reject"
else "Unrecognized reply. Action rejected."
)
@@ -361,8 +864,6 @@ def channel_hitl_prompt(
_manager: Any | None = None # ChannelManager
_bus_loop: asyncio.AbstractEventLoop | None = None
_bus_thread: threading.Thread | None = None
_cli_agent: Any = None # shared agent reference (same as CLI)
_cli_thread_id: str | None = None # shared thread_id (same conversation)
def _channels_is_running(channel_type: str | None = None) -> bool:
@@ -380,9 +881,18 @@ def _channels_running_list() -> list[str]:
return _manager.running_channels() if _manager else []
def _channels_stop(channel_type: str | None = None) -> None:
"""Stop channel(s) and clean up module-level state."""
global _manager, _bus_loop, _bus_thread, _cli_agent, _cli_thread_id
def _channels_stop(
channel_type: str | None = None,
*,
runtime: ChannelRuntime | None = None,
) -> None:
"""Stop channel(s) and clean up module-level state.
``runtime`` is the ``ChannelRuntime`` whose binding should be
cleared once the channels are gone — the caller owns it (commands
keep a reference via ``ctx.channel_runtime``).
"""
global _manager, _bus_loop, _bus_thread
if channel_type is None:
# Stop everything
@@ -395,15 +905,13 @@ def _channels_stop(channel_type: str | None = None) -> None:
future.result(timeout=10)
except Exception as e:
_channel_logger.debug(f"Error stopping channels: {e}")
if _manager:
_manager.bus.stop()
if _bus_thread:
_bus_thread.join(timeout=5)
_manager = None
_bus_loop = None
_bus_thread = None
_cli_agent = None
_cli_thread_id = None
if runtime is not None:
runtime.clear()
return
# Stop a specific channel
@@ -417,9 +925,8 @@ def _channels_stop(channel_type: str | None = None) -> None:
except Exception as e:
_channel_logger.debug(f"Error removing channel {channel_type}: {e}")
if _manager and not _manager.running_channels():
_cli_agent = None
_cli_thread_id = None
if _manager and not _manager.running_channels() and runtime is not None:
runtime.clear()
def _start_channels_bus_mode(
@@ -533,6 +1040,20 @@ async def _bus_inbound_consumer(bus, manager) -> None:
except asyncio.CancelledError:
break
# /stop should preempt HITL interception so cancel works while
# waiting for approvals/questions. If a HITL wait is pending,
# still release it so the blocking prompt can unwind immediately.
if _is_stop_command(msg.content):
if _try_set_hitl_reply(msg.channel, msg.chat_id, msg.content):
_channel_logger.info(
f"[bus] stop request released HITL wait for "
f"{msg.channel}:{msg.chat_id}"
)
_task = asyncio.create_task(_handle_bus_message(bus, manager, msg))
_tasks.add(_task)
_task.add_done_callback(_tasks.discard)
continue
# Check if this message is a HITL approval reply
if _try_set_hitl_reply(msg.channel, msg.chat_id, msg.content):
_channel_logger.info(
@@ -561,6 +1082,37 @@ async def _handle_bus_message(bus, manager, msg) -> None:
)
manager.record_message(msg.channel, "received")
# Fast-path: /stop intercept. Handle on the bus task itself so we
# don't deadlock behind the main-thread stream we're trying to
# interrupt. No typing indicator, no queue entry.
if _is_stop_command(msg.content):
cancelled_count, active_count = _cancel_channel_session(
msg.channel, msg.chat_id
)
try:
await bus.publish_outbound(
OutboundMessage(
channel=msg.channel,
chat_id=msg.chat_id,
content="Stopped.",
reply_to=msg.message_id or None,
metadata=msg.metadata,
)
)
manager.record_message(msg.channel, "sent")
except Exception as e:
_channel_logger.error(f"[bus] /stop ack send error: {e}")
else:
if cancelled_count or active_count:
_channel_logger.info(
"[bus] /stop cancelled %d request(s) (%d active) for %s:%s",
cancelled_count,
active_count,
msg.channel,
msg.chat_id,
)
return
channel = manager.get_channel(msg.channel)
typing_active = False
if channel:
@@ -630,6 +1182,8 @@ async def _handle_bus_message(bus, manager, msg) -> None:
f"for {cm.msg_id}"
)
_pop_channel_response(cm.msg_id, cancel_pending=True)
if _channel_request_state(cm.msg_id) != "active":
_complete_channel_request(cm.msg_id)
return
response = _pop_channel_response(cm.msg_id) or "No response"
@@ -645,6 +1199,8 @@ async def _handle_bus_message(bus, manager, msg) -> None:
manager.record_message(msg.channel, "sent")
except asyncio.CancelledError:
_pop_channel_response(cm.msg_id, cancel_pending=True)
if _channel_request_state(cm.msg_id) != "active":
_complete_channel_request(cm.msg_id)
raise
except Exception as e:
_channel_logger.error(f"[bus] Outbound error: {e}")
@@ -682,131 +1238,13 @@ def _print_channel_panel(channels: list[tuple[str, bool, str]]) -> None:
console.print()
def _cmd_channel(
args: str,
agent: Any,
thread_id: str,
*,
send_thinking: bool | None = None,
) -> None:
"""Start a channel in background using bus mode.
Usage:
/channel [telegram|discord|imessage] -- start channel (default from config)
/channel status -- show current channel status
/channel stop -- stop running channel
"""
global _cli_agent, _cli_thread_id
from ..config import load_config
app_config = load_config()
channel_type = args.strip().lower() if args and args.strip() else ""
if channel_type == "status":
running = _channels_running_list()
if running and _manager:
detailed = _manager.get_detailed_status()
table = Table(title="Channel Status", show_header=True, expand=False)
table.add_column("Channel", style="cyan")
table.add_column("Status")
table.add_column("Uptime", style="dim")
table.add_column("Rx", justify="right")
table.add_column("Tx", justify="right")
for ch_name in running:
info = detailed.get(ch_name, {})
secs = info.get("uptime_seconds", 0)
mins, s = divmod(int(secs), 60)
hours, mins = divmod(mins, 60)
uptime = f"{hours}h{mins:02d}m" if hours else f"{mins}m{s:02d}s"
rx = str(info.get("received", 0))
tx = str(info.get("sent", 0))
table.add_row(ch_name, "[green]running[/green]", uptime, rx, tx)
console.print(table)
console.print()
else:
console.print("[dim]No channel running[/dim]\n")
return
if not channel_type:
channel_type = app_config.channel_enabled
if not channel_type:
console.print("[yellow]No channel configured.[/yellow]")
console.print(
"[dim]Run[/dim] evosci onboard [dim]or specify:[/dim] /channel telegram\n"
)
return
requested = [t.strip() for t in channel_type.split(",") if t.strip()]
if _channels_is_running():
running = _channels_running_list()
results: list[tuple[str, bool, str]] = []
for ct in requested:
if ct in running:
results.append((ct, True, "already running"))
else:
try:
_add_channel_to_running_bus(
ct,
app_config,
send_thinking=send_thinking,
)
results.append((ct, True, "connected (bus)"))
except Exception as e:
results.append((ct, False, str(e)))
_print_channel_panel(results)
return
_cli_agent = agent
_cli_thread_id = thread_id
# Override channel_enabled for this invocation
original = app_config.channel_enabled
app_config.channel_enabled = channel_type
try:
_start_channels_bus_mode(
app_config,
agent,
thread_id,
send_thinking=send_thinking,
)
results = [(ct, True, "connected (bus)") for ct in requested]
except Exception as e:
results = [(ct, False, str(e)) for ct in requested]
finally:
app_config.channel_enabled = original
_print_channel_panel(results)
def _cmd_channel_stop(channel_type: str | None = None) -> None:
"""Stop background channel(s).
Args:
channel_type: Specific channel to stop, or None to stop all.
"""
if not _channels_is_running():
console.print("[dim]No channel running[/dim]\n")
return
if channel_type:
if not _channels_is_running(channel_type):
console.print(f"[dim]{channel_type} is not running[/dim]\n")
return
_channels_stop(channel_type)
console.print(f"[dim]{channel_type} stopped[/dim]\n")
else:
running = _channels_running_list()
_channels_stop()
console.print(f"[dim]{', '.join(running)} stopped[/dim]\n")
def _auto_start_channel(
agent: Any,
thread_id: str,
config,
*,
send_thinking: bool | None = None,
runtime: ChannelRuntime | None = None,
) -> None:
"""Start channels automatically from config (bus mode).
@@ -814,21 +1252,23 @@ def _auto_start_channel(
agent: Compiled agent graph.
thread_id: Current thread ID.
config: EvoScientistConfig with channel settings.
runtime: Caller-owned ``ChannelRuntime`` to bind so commands
running over the channels can swap the agent later. ``None``
is accepted for callers that don't yet pass one.
"""
global _cli_agent, _cli_thread_id
if not config.channel_enabled:
return
_cli_agent = agent
_cli_thread_id = thread_id
_start_channels_bus_mode(
config,
agent,
thread_id,
send_thinking=send_thinking,
)
# Bind only after startup succeeds; a failure above must not leave
# a stale runtime binding pointing at channels that never started.
if runtime is not None:
runtime.bind(agent, thread_id)
types = [t.strip() for t in config.channel_enabled.split(",") if t.strip()]
results = [(ct, True, "connected (bus)") for ct in types]
_print_channel_panel(results)
+44 -15
View File
@@ -21,7 +21,7 @@ import os
import pathlib
import subprocess
import sys
from typing import TYPE_CHECKING
from typing import TYPE_CHECKING, Any
if TYPE_CHECKING:
from textual.app import App
@@ -29,6 +29,15 @@ if TYPE_CHECKING:
logger = logging.getLogger(__name__)
_PREVIEW_MAX = 40
_pyperclip_notify_shown = False
def _is_remote_session() -> bool:
"""Return True when running over SSH without a local display."""
if os.environ.get("SSH_CLIENT") or os.environ.get("SSH_CONNECTION"):
if not os.environ.get("DISPLAY") and not os.environ.get("WAYLAND_DISPLAY"):
return True
return False
# ── Platform-native clipboard read ────────────────────────────────
@@ -132,8 +141,6 @@ def copy_selection_to_clipboard(app: App) -> None:
if not hasattr(widget, "text_selection") or not widget.text_selection:
continue
selection = widget.text_selection
if selection.end is None:
continue
try:
result = widget.get_selection(selection)
except (AttributeError, TypeError, ValueError, IndexError) as exc:
@@ -154,37 +161,59 @@ def copy_selection_to_clipboard(app: App) -> None:
combined = "\n".join(selected_texts)
# Try methods in priority order
copy_methods = [app.copy_to_clipboard]
# Build method list: (fn, reliable) — reliable means we *know* the text
# reached the system clipboard (e.g. pyperclip). OSC 52 / Textual write
# to the terminal and succeed even when the terminal silently ignores the
# sequence (PuTTY, older terminals).
copy_methods: list[tuple[Any, bool]] = [
(app.copy_to_clipboard, False),
]
try:
import pyperclip
copy_methods.insert(0, pyperclip.copy)
copy_methods.insert(0, (pyperclip.copy, True))
except ImportError:
pass
global _pyperclip_notify_shown
if not _pyperclip_notify_shown:
_pyperclip_notify_shown = True
app.notify(
'Failed to import "pyperclip", text copying might not work.',
severity="information",
timeout=3,
)
copy_methods.append(_copy_osc52)
copy_methods.append((_copy_osc52, False))
for fn in copy_methods:
remote = _is_remote_session()
for fn, reliable in copy_methods:
try:
fn(combined)
except (OSError, RuntimeError, TypeError) as exc:
logger.debug(
"Clipboard method %s failed: %s", getattr(fn, "__name__", repr(fn)), exc
)
continue
if reliable or not remote:
app.notify(
f'"{_shorten(selected_texts)}" copied',
severity="information",
timeout=2,
markup=False,
)
except (OSError, RuntimeError, TypeError) as exc:
logger.debug(
"Clipboard method %s failed: %s", getattr(fn, "__name__", repr(fn)), exc
)
continue
else:
# OSC 52 over SSH — may be silently ignored (e.g. Windows Terminal, PuTTY)
app.notify(
"Copied text - if paste fails, use Shift+mouse-select for native copy",
severity="information",
timeout=3,
)
return
app.notify(
"Failed to copy — no clipboard method available",
"Copy failed — use Shift+mouse-select for native terminal copy",
severity="warning",
timeout=3,
)
+1202 -313
View File
File diff suppressed because it is too large Load Diff
+150 -24
View File
@@ -19,16 +19,35 @@ from pathlib import Path
_PATH_CHARS = r"A-Za-z0-9._~/\\:-"
FILE_MENTION_PATTERN = re.compile(r"@(?P<path>(?:\\.|[" + _PATH_CHARS + r"])+)")
FILE_MENTION_PATTERN = re.compile(
r"@(?:"
r'"(?P<dquoted>[^"\n]+)"'
r"|'(?P<squoted>[^'\n]+)'"
r"|(?P<bare>(?:\\.|[" + _PATH_CHARS + r"])+)"
r")"
)
"""Matches ``@path/to/file`` in user input.
Escaped spaces (``@my\\\\ folder/file``) are supported. Bare ``@`` with no
path characters is not matched (uses ``+`` not ``*``).
Three forms are supported, in priority order:
1. ``@"path with spaces.pdf"`` — explicit double-quoted path
2. ``@'path with spaces.pdf'`` — explicit single-quoted path
3. ``@bare/path`` — backslash-escaped spaces (``@my\\\\ folder/file``) work;
raw unescaped spaces are handled via greedy expansion in
:func:`parse_file_mentions`.
Bare ``@`` with no path characters is not matched (uses ``+`` not ``*``).
"""
_EMAIL_PREFIX = re.compile(r"[a-zA-Z0-9._%+-]$")
"""If the character immediately before ``@`` matches this, it's an email address."""
# Hard cap on tokens consumed during greedy expansion across whitespace.
_GREEDY_MAX_TOKENS = 20
# Trailing punctuation stripped before checking if a greedy candidate exists.
_GREEDY_TRAIL_PUNCT = ",;:!?)]}>"
# Files larger than this are referenced by path only (not embedded inline).
_MAX_EMBED_BYTES = 256 * 1024 # 256 KB
@@ -193,6 +212,75 @@ def _read_file(path: Path) -> str:
return f"\n### {path.name}\nPath: `{path}`\n```\n{content}\n```"
def _resolve_path(raw: str, cwd: Path) -> Path | None:
"""Resolve *raw* to an existing file path, or ``None``.
Honors backslash-escaped spaces and ``~`` expansion. Returns ``None``
when the path does not exist, is not a regular file, or raises
``OSError``/``RuntimeError`` during resolution.
"""
clean = raw.replace("\\ ", " ")
try:
p = Path(clean).expanduser()
if not p.is_absolute():
p = cwd / p
resolved = p.resolve()
except (OSError, RuntimeError):
return None
if resolved.is_file():
return resolved
return None
def _greedy_extend(
text: str,
raw: str,
match_end: int,
cwd: Path,
) -> tuple[str, Path, int] | None:
"""Try to extend *raw* across whitespace until the path resolves.
Walks the text after *match_end*, capped at the next newline or the
start of another ``@`` mention. Tries the longest plausible suffix
first and shrinks one token at a time, stripping trailing punctuation
that is unlikely to be part of a filename.
Returns ``(extended_raw, resolved_file, new_end_pos)`` on success,
else ``None``.
"""
rest = text[match_end:]
# Hard boundaries that should never be crossed.
boundary = len(rest)
nl = rest.find("\n")
if nl >= 0:
boundary = nl
next_at = re.search(r"\s@", rest[:boundary])
if next_at:
boundary = next_at.start()
region = rest[:boundary]
if not region or not region[0].isspace():
return None
tokens = list(re.finditer(r"\S+", region))
if not tokens:
return None
# Try longest-first so we prefer the most specific match.
for i in range(min(len(tokens), _GREEDY_MAX_TOKENS), 0, -1):
end = tokens[i - 1].end()
suffix = region[:end].rstrip(_GREEDY_TRAIL_PUNCT)
if not suffix:
continue
candidate = raw + suffix
resolved = _resolve_path(candidate, cwd)
if resolved is not None:
return candidate, resolved, match_end + len(suffix)
return None
def parse_file_mentions(
text: str,
cwd: Path | None = None,
@@ -218,23 +306,39 @@ def parse_file_mentions(
files: list[Path] = []
warnings: list[str] = []
seen: set[Path] = set()
# finditer would normally re-scan from each match's end, but greedy
# expansion can consume bytes past that point. Track a manual cursor
# and skip matches that start before it.
cursor = 0
for match in FILE_MENTION_PATTERN.finditer(text):
if match.start() < cursor:
continue
# Skip email addresses — character immediately before @ is alphanumeric
before = text[: match.start()]
if before and _EMAIL_PREFIX.search(before):
continue
raw = match.group("path")
clean = raw.replace("\\ ", " ")
dquoted = match.group("dquoted")
squoted = match.group("squoted")
bare = match.group("bare")
quoted_raw = dquoted if dquoted is not None else squoted
raw = quoted_raw if quoted_raw is not None else bare
is_quoted = quoted_raw is not None
try:
p = Path(clean).expanduser()
if not p.is_absolute():
p = cwd / p
resolved = p.resolve()
if not resolved.exists() or not resolved.is_file():
resolved = _resolve_path(raw, cwd)
end_pos = match.end()
if resolved is None and not is_quoted:
extended = _greedy_extend(text, raw, match.end(), cwd)
if extended is not None:
raw, resolved, end_pos = extended
cursor = end_pos
if resolved is None:
warnings.append(f"@file not found: {raw}")
continue
# Deduplicate: skip paths already seen in this message.
if resolved in seen:
continue
@@ -250,8 +354,6 @@ def parse_file_mentions(
f"@{raw} is outside the workspace "
f"({workspace_root}) — embedding may expose sensitive files"
)
except (OSError, RuntimeError) as exc:
warnings.append(f"invalid @file path {raw!r}: {exc}")
return files, warnings
@@ -302,6 +404,13 @@ def _type_hint(rel_path: str) -> str:
return suffix or "file"
def _format_mention(rel_path: str) -> str:
"""Render *rel_path* as an ``@`` mention, quoting if it contains spaces."""
if " " in rel_path:
return f'@"{rel_path}"'
return f"@{rel_path}"
def complete_file_mention(
text: str,
workspace_dir: str | None = None,
@@ -320,13 +429,22 @@ def complete_file_mention(
List of ``(completion_string, type_hint)`` tuples, e.g.
``[("@results/v2.json", "json"), ("@README.md", "md")]``.
Directories have a trailing ``/`` and type hint ``"dir"``.
Paths containing spaces are returned in double-quoted form,
e.g. ``@"my docs/file.pdf"``.
"""
# Find the last @token
match = re.search(r"@([^\s]*)$", text)
# Find the last @token. Allow whitespace inside a quoted partial so
# completion keeps working as the user types ``@"PRE`` → ``@"PREPING_ B``.
quoted_match = re.search(r'@"([^"\n]*)$', text)
if quoted_match:
partial = quoted_match.group(1)
quoted = True
else:
match = re.search(r"@([^\s\"']*)$", text)
if not match:
return []
partial = match.group(1).replace("\\ ", " ")
quoted = False
base_str = workspace_dir or str(Path.cwd())
base = Path(base_str)
@@ -344,10 +462,10 @@ def complete_file_mention(
rel = entry.relative_to(base)
suffix = "/" if entry.is_dir() else ""
candidates_raw.append(rel.as_posix() + suffix)
except OSError:
except (OSError, ValueError):
return []
return [
(f"@{r}", "dir" if r.endswith("/") else _type_hint(r))
(_format_mention(r), "dir" if r.endswith("/") else _type_hint(r))
for r in candidates_raw[:10]
]
@@ -363,16 +481,24 @@ def complete_file_mention(
except OSError:
pass
combined = all_files + dir_candidates
# Determine query: if partial has a slash, search within that subtree
if "/" in partial:
# Filter candidates to those starting with the directory prefix
# Search within the subtree for the given directory prefix
combined = all_files + dir_candidates
dir_prefix = partial.rsplit("/", 1)[0] + "/"
file_query = partial.rsplit("/", 1)[1]
subtree = [c for c in combined if c.startswith(dir_prefix)]
results = _fuzzy_search(file_query, subtree)
else:
results = _fuzzy_search(partial, combined)
# Depth 1 only: top-level files and directories
top_files = [f for f in all_files if "/" not in f]
results = _fuzzy_search(partial, top_files + dir_candidates)
return [(f"@{r}", "dir" if r.endswith("/") else _type_hint(r)) for r in results]
if quoted:
# User opened a quoted mention — close it for them.
return [
(f'@"{r}"', "dir" if r.endswith("/") else _type_hint(r)) for r in results
]
return [
(_format_mention(r), "dir" if r.endswith("/") else _type_hint(r))
for r in results
]
+1 -1
View File
@@ -1,7 +1,7 @@
"""History-based auto-suggest for Textual TUI Input widget.
Reads prompt_toolkit FileHistory format so Rich CLI and TUI share the same
history file at ~/.config/ai4scientist/history.
history file at ~/.evoscientist/history.
"""
from __future__ import annotations
File diff suppressed because it is too large Load Diff
+4 -15
View File
@@ -9,7 +9,6 @@ from __future__ import annotations
from collections import Counter
import questionary
from prompt_toolkit.styles import Style as PtStyle
from questionary import Choice
from ..mcp.registry import (
@@ -21,18 +20,8 @@ from ..mcp.registry import (
install_mcp_server,
install_mcp_servers,
)
from ..stream.display import console
_PICKER_STYLE = PtStyle.from_dict(
{
"questionmark": "#888888",
"question": "",
"pointer": "bold",
"highlighted": "bold",
"text": "#888888",
"answer": "bold",
}
)
from ..stream.console import console
from .widgets.thread_selector import PICKER_STYLE
_INSTALLED_INDICATOR = ("fg:#4caf50", "\u2713 ")
@@ -57,7 +46,7 @@ def _checkbox_ask(choices, message: str, **kwargs):
return questionary.checkbox(
message,
choices=choices,
style=_PICKER_STYLE,
style=PICKER_STYLE,
qmark="\u276f",
**kwargs,
).ask()
@@ -100,7 +89,7 @@ def _browse_and_select(
selected_tag = questionary.select(
"Filter by tag:",
choices=tag_choices,
style=_PICKER_STYLE,
style=PICKER_STYLE,
qmark="\u276f",
).ask()
+2 -128
View File
@@ -1,10 +1,10 @@
"""MCP server display, operations, and /mcp slash-command dispatcher."""
"""Shared UI helpers for MCP server display and operations (used by the Typer `mcp` commands)."""
from typing import Any
from rich.table import Table
from ..stream.display import console
from ..stream.console import console
def _mcp_list_servers() -> None:
@@ -181,129 +181,3 @@ def _show_mcp_config(name: str = "", *, show_blank_line: bool = True) -> str:
if show_blank_line:
console.print()
return "ok"
def _cmd_mcp_add(args_str: str) -> None:
"""Handle ``/mcp add ...``."""
import shlex
from ..mcp import parse_mcp_add_args
if not args_str.strip():
console.print("[bold]Usage:[/bold] /mcp add <name> <command-or-url> [args...]")
console.print()
console.print(
"[dim]Transport is auto-detected: URLs \u2192 http, commands \u2192 stdio[/dim]"
)
console.print()
console.print("[bold]Examples:[/bold]")
console.print(
" /mcp add sequential-thinking npx -y @modelcontextprotocol/server-sequential-thinking"
)
console.print(" /mcp add docs-langchain https://docs.langchain.com/mcp")
console.print(
" /mcp add my-sse http://localhost:9090/sse --transport sse --expose-to research-agent"
)
console.print()
console.print("[dim]Options:[/dim]")
console.print(" --transport T Transport type (default: auto-detect)")
console.print(
" --tools t1,t2 Tool allowlist (supports wildcards: *_exa, read_*)"
)
console.print(" --expose-to a1,a2 Target agents (default: main)")
console.print(" --header Key:Value HTTP header (repeatable)")
console.print(" --env KEY=VALUE Env var for stdio (repeatable)")
console.print(
" --env-ref KEY Env var as runtime ${KEY} reference (repeatable)"
)
console.print()
return
try:
tokens = shlex.split(args_str)
kwargs = parse_mcp_add_args(tokens)
_mcp_add_server_from_kwargs(kwargs, show_reload_hint=True)
except ValueError as exc:
console.print(f"[red]{exc}[/red]")
console.print()
def _cmd_mcp_edit(args_str: str) -> None:
"""Handle ``/mcp edit <name> --field value ...``."""
import shlex
from ..mcp import parse_mcp_edit_args
if not args_str.strip():
console.print("[bold]Usage:[/bold] /mcp edit <name> --<field> <value> ...")
console.print()
console.print(
"[dim]Fields:[/dim] --transport, --command, --url, --args, --tools, --expose-to, --header, --env"
)
console.print(
"[dim]Use[/dim] --tools none [dim]or[/dim] --expose-to none [dim]to clear a field.[/dim]"
)
console.print()
console.print("[bold]Examples:[/bold]")
console.print(" /mcp edit filesystem --expose-to main,code-agent")
console.print(" /mcp edit filesystem --tools read_file,write_file")
console.print(" /mcp edit my-api --url http://new-host:8080/mcp")
console.print(" /mcp edit my-api --tools none")
console.print()
return
try:
tokens = shlex.split(args_str)
name, fields = parse_mcp_edit_args(tokens)
_mcp_edit_server_fields(name, fields, show_reload_hint=True)
except ValueError as exc:
console.print(f"[red]{exc}[/red]")
console.print()
def _cmd_mcp_remove(name: str) -> None:
"""Handle ``/mcp remove <name>``."""
_mcp_remove_server(name, show_reload_hint=True)
console.print()
def _cmd_mcp_config(name: str) -> None:
"""Handle ``/mcp config [name]``."""
_show_mcp_config(name, show_blank_line=True)
def _cmd_mcp(args: str) -> None:
"""Dispatch ``/mcp`` subcommands."""
args = args.strip()
if not args:
_mcp_list_servers()
return
parts = args.split(maxsplit=1)
subcmd = parts[0].lower()
subargs = parts[1] if len(parts) > 1 else ""
if subcmd == "list":
_mcp_list_servers()
elif subcmd == "add":
_cmd_mcp_add(subargs)
elif subcmd == "edit":
_cmd_mcp_edit(subargs)
elif subcmd == "remove":
_cmd_mcp_remove(subargs)
elif subcmd == "config":
_cmd_mcp_config(subargs)
elif subcmd == "install":
from .mcp_install_cmd import _cmd_install_mcp
_cmd_install_mcp(subargs)
else:
console.print("[bold]MCP commands:[/bold]")
console.print(" /mcp List configured servers")
console.print(" /mcp list List configured servers")
console.print(" /mcp config Show detailed server config")
console.print(" /mcp add ... Add a server")
console.print(" /mcp edit ... Edit an existing server")
console.print(" /mcp remove ... Remove a server")
console.print(" /mcp install ... Browse and install servers")
console.print()
+21
View File
@@ -0,0 +1,21 @@
"""Helper for printing the session-exit Goodbye message and resume hint."""
from __future__ import annotations
from rich.console import Console
from rich.markup import escape
def print_resume_hint(
thread_id: str | None,
console: Console | None = None,
) -> None:
"""Print ``Goodbye!`` and, when available, a resume hint for *thread_id*."""
out = console or Console()
out.print("[dim]Goodbye![/dim]")
if thread_id:
from ..sessions import short_thread_id
out.print()
out.print("[dim]Resume this session with:[/dim]")
out.print(f"[cyan]EvoSci --resume {escape(short_thread_id(thread_id))}[/cyan]")
+221
View File
@@ -0,0 +1,221 @@
"""CommandUI Protocol adapter for the Rich CLI surface.
Lifecycle methods (``request_quit``, ``force_quit``, ``clear_chat``,
``start_new_session``, ``handle_session_resume``, ``update_status_after_compact``)
are callback-driven: when their corresponding ``on_*`` constructor kwarg
is ``None``, the method is a silent no-op, mirroring
``ChannelCommandUI``'s fallback pattern. Callers that need a specific
side-effect (REPL quit flag flip, status-bar refresh, …) wire the
callback at construction time; non-interactive surfaces (tests,
alternate REPLs) can leave callbacks unset without crashing.
``wait_for_*`` methods return ``None`` on cancel / fallback and are
always safe to ``await``.
"""
from __future__ import annotations
from collections.abc import Awaitable, Callable
from typing import Any
from rich.console import Console
from ..commands.base import CommandUI
class RichCLICommandUI(CommandUI):
"""CommandUI implementation that prints to a Rich ``Console``.
Commands that affect CLI-closure state (session lifecycle, exit flag,
status-bar snapshot) go through optional callbacks wired by the REPL.
This mirrors ``ChannelCommandUI``'s injection pattern and keeps
``interactive.py``'s ``state`` dict as the single source of truth.
"""
def __init__(
self,
console: Console,
*,
on_request_quit: Callable[[], None] | None = None,
on_force_quit: Callable[[], None] | None = None,
on_clear_chat: Callable[[], None] | None = None,
on_status_after_compact: Callable[[int], None] | None = None,
on_start_new_session: Callable[[], Awaitable[None]] | None = None,
on_handle_session_resume: (
Callable[[str, str | None], Awaitable[None]] | None
) = None,
) -> None:
self.console = console
self._on_request_quit = on_request_quit
self._on_force_quit = on_force_quit
self._on_clear_chat = on_clear_chat
self._on_status_after_compact = on_status_after_compact
self._on_start_new_session = on_start_new_session
self._on_handle_session_resume = on_handle_session_resume
# Bound ``console.status(...)`` context manager used by
# /compact's start/stop indicator pair.
self._compact_status_ctx: Any = None
# ── Core I/O ─────────────────────────────────────────────
@property
def supports_interactive(self) -> bool:
return True
def append_system(self, text: str, style: str = "dim") -> None:
self.console.print(text, style=style)
def mount_renderable(self, renderable: Any) -> None:
self.console.print(renderable)
async def flush(self) -> None:
# Rich console flushes synchronously; nothing to await.
return
# ── Interactive pickers ────────────────────────────────
async def wait_for_thread_pick(
self, threads: list[dict], current_thread: str, title: str
) -> str | None:
"""Interactive workspace-grouped thread picker using ``questionary``.
Ported from the pre-migration ``_cmd_resume`` implementation.
Returns the selected ``thread_id`` string, or ``None`` on cancel.
Callers (``ResumeCommand``/``DeleteCommand``) pre-check for
empty thread lists before invoking this method.
"""
import questionary # type: ignore[import-untyped]
from prompt_toolkit.layout.dimension import ( # type: ignore[import-untyped]
Dimension,
)
from questionary.prompts.common import ( # type: ignore[import-untyped]
InquirerControl,
)
from ..sessions import _format_relative_time
from .widgets.thread_selector import PICKER_STYLE, _build_items
choices: list[Any] = []
for item in _build_items(threads):
if item["type"] == "header":
choices.append(questionary.Separator(f"── \U0001f4c2 {item['label']}"))
elif item["type"] == "subheader":
choices.append(questionary.Separator(f" {item['label']}"))
else:
t = item["thread"]
tid = t["thread_id"]
preview = t.get("preview", "") or ""
msgs = t.get("message_count", 0)
model = t.get("model", "") or ""
when = _format_relative_time(t.get("updated_at"))
indent = " " if item.get("indented") else " "
marker = " *" if tid == current_thread else ""
parts = [f"{indent}{tid}{marker}"]
if preview:
parts.append(preview[:40] + "…" if len(preview) > 40 else preview)
parts.append(f"({msgs} msgs)")
if model:
parts.append(model)
if when:
parts.append(when)
label = " ".join(parts)
choices.append(questionary.Choice(title=label, value=tid))
prompt = questionary.select(title, choices=choices, style=PICKER_STYLE)
# Limit visible list to 10 rows with scrolling. Touches
# questionary/prompt-toolkit private internals so guard against
# library-shape changes — picker stays functional at default
# height even if the cap fails.
try:
for window in prompt.application.layout.find_all_windows():
if isinstance(window.content, InquirerControl):
window.height = Dimension(max=10)
break
except Exception:
pass
# ``ask_async`` (questionary >= 2.0.1) avoids blocking the
# asyncio event loop while the user interacts with the picker.
return await prompt.ask_async()
# ── Lifecycle callbacks ───────────────────────────────
def clear_chat(self) -> None:
if self._on_clear_chat is not None:
self._on_clear_chat()
else:
self.console.clear()
def request_quit(self) -> None:
if self._on_request_quit is not None:
self._on_request_quit()
def force_quit(self) -> None:
if self._on_force_quit is not None:
self._on_force_quit()
async def start_new_session(self) -> None:
if self._on_start_new_session is not None:
await self._on_start_new_session()
async def handle_session_resume(
self, thread_id: str, workspace_dir: str | None = None
) -> None:
if self._on_handle_session_resume is not None:
await self._on_handle_session_resume(thread_id, workspace_dir)
# /compact indicator pair — duck-typed by ``CompactCommand`` via
# ``getattr``, not declared on the ``CommandUI`` Protocol.
def start_compacting_indicator(self) -> None:
# Idempotent: close any lingering context before starting a new
# one so a double-call (e.g. two overlapping /compact attempts
# via the message queue) can't leak a Rich Live handle.
if self._compact_status_ctx is not None:
try:
self._compact_status_ctx.__exit__(None, None, None)
except Exception:
pass
self._compact_status_ctx = None
status = self.console.status("[cyan]Compacting conversation...[/cyan]")
status.__enter__()
self._compact_status_ctx = status
def stop_compacting_indicator(self) -> None:
ctx = self._compact_status_ctx
self._compact_status_ctx = None
if ctx is not None:
try:
ctx.__exit__(None, None, None)
except Exception:
pass
def update_status_after_compact(self, input_tokens: int) -> None:
if self._on_status_after_compact is not None:
self._on_status_after_compact(input_tokens)
# ── Skill / MCP browse (delegated to worker threads) ──
async def wait_for_skill_browse(
self, index: list[dict], installed_names: set[str], pre_filter_tag: str
) -> list[str] | None:
"""Delegate to the extracted questionary picker on a worker
thread — questionary blocks the event loop so the call must
not happen on the main asyncio thread."""
import asyncio
from .skills_cmd import _pick_skills_interactive
return await asyncio.to_thread(
_pick_skills_interactive, index, installed_names, pre_filter_tag
)
async def wait_for_mcp_browse(
self, servers: list, installed_names: set[str], pre_filter_tag: str
) -> list | None:
"""Delegate to the MCP browse picker on a worker thread."""
import asyncio
from .mcp_install_cmd import _browse_and_select
return await asyncio.to_thread(
_browse_and_select, servers, installed_names, pre_filter_tag
)
+30 -224
View File
@@ -1,171 +1,30 @@
"""Slash commands for skill management: /skills, /install-skill, /uninstall-skill, /evoskills."""
"""Shared UI helpers for skill-management commands (picker used by /evoskills)."""
from pathlib import Path
from ..stream.display import console
from .agent import _shorten_path
from ..stream.console import console
def _cmd_list_skills() -> None:
"""List all available skills (workspace, global, and built-in)."""
from ..paths import GLOBAL_SKILLS_DIR, USER_SKILLS_DIR
from ..tools.skills_manager import list_skills
def _pick_skills_interactive(
index: list[dict],
installed_names: set[str],
pre_filter_tag: str,
) -> list[str] | None:
"""Interactive questionary picker for EvoSkills browse.
skills = list_skills(include_system=True)
Two-phase picker:
1. tag filter — ``questionary.select`` (skipped if ``pre_filter_tag``)
2. multi-select — ``questionary.checkbox`` with installed items disabled
if not skills:
console.print("[dim]No skills available.[/dim]")
console.print("[dim]Install with:[/dim] /install-skill <path-or-url>")
console.print(
f"[dim]Global skills:[/dim] [cyan]{_shorten_path(str(GLOBAL_SKILLS_DIR))}[/cyan]"
)
console.print()
return
workspace_skills = [s for s in skills if s.source == "workspace"]
global_skills = [s for s in skills if s.source == "global"]
builtin_skills = [s for s in skills if s.source == "builtin"]
sections = [
("Workspace Skills", workspace_skills, "green"),
("Global Skills", global_skills, "cyan"),
("Built-in Skills", builtin_skills, "blue"),
]
printed = False
for title, group, color in sections:
if not group:
continue
if printed:
console.print()
console.print(f"[bold]{title}[/bold] ({len(group)}):")
for skill in group:
tags_str = f" [dim]({', '.join(skill.tags)})[/dim]" if skill.tags else ""
console.print(
f" [{color}]{skill.name}[/{color}] - {skill.description}{tags_str}"
)
printed = True
console.print(
f"\n[dim]Global skills:[/dim] [cyan]{_shorten_path(str(GLOBAL_SKILLS_DIR))}[/cyan]"
)
console.print(
f"[dim]Workspace skills:[/dim] [green]{_shorten_path(str(USER_SKILLS_DIR))}[/green]"
)
console.print()
def _cmd_install_skill(args: str) -> None:
"""Install a skill from local path or GitHub URL.
By default, installs to the global skills directory (~/.config/ai4scientist/skills/).
Append --local to install to the current workspace instead.
Usage: /install-skill <path-or-url> [--local]
"""
from ..paths import GLOBAL_SKILLS_DIR, USER_SKILLS_DIR
from ..tools.skills_manager import install_skill
# Parse --local flag out of the args string
local = "--local" in args.split()
source = args.replace("--local", "").strip()
if not source:
console.print("[red]Usage:[/red] /install-skill <path-or-url> [--local]")
console.print("[dim]Examples:[/dim]")
console.print(" /install-skill ./my-skill")
console.print(
" /install-skill https://github.com/user/repo/tree/main/skill-name"
)
console.print(" /install-skill user/repo@skill-name")
console.print(
" /install-skill ./my-skill --local [dim](workspace only)[/dim]"
)
console.print()
return
dest_label = (
f"[cyan]{_shorten_path(str(USER_SKILLS_DIR))}[/cyan] [dim](workspace)[/dim]"
if local
else f"[cyan]{_shorten_path(str(GLOBAL_SKILLS_DIR))}[/cyan] [dim](global)[/dim]"
)
console.print(f"[dim]Installing skill from:[/dim] {source}")
console.print(f"[dim]Destination:[/dim] {dest_label}")
result = install_skill(source, global_install=not local)
if result.get("batch"):
# Batch install — multiple skills
for item in result.get("installed", []):
console.print(f"[green]Installed:[/green] {item['name']}")
console.print(
f" [dim]Description:[/dim] {item.get('description', '(none)')}"
)
console.print(
f" [dim]Path:[/dim] [cyan]{_shorten_path(item['path'])}[/cyan]"
)
for item in result.get("failed", []):
console.print(f"[red]Failed:[/red] {item['name']} — {item['error']}")
installed_count = len(result.get("installed", []))
if installed_count:
console.print(f"\n[green]{installed_count} skill(s) installed.[/green]")
console.print("[dim]Reload with /new to apply.[/dim]")
elif result["success"]:
console.print(f"[green]Installed:[/green] {result['name']}")
console.print(f"[dim]Description:[/dim] {result.get('description', '(none)')}")
console.print(f"[dim]Path:[/dim] [cyan]{_shorten_path(result['path'])}[/cyan]")
console.print()
console.print("[dim]Reload with /new to apply.[/dim]")
else:
console.print(f"[red]Failed:[/red] {result['error']}")
console.print()
def _cmd_uninstall_skill(name: str) -> None:
"""Uninstall a user-installed skill."""
from ..tools.skills_manager import uninstall_skill
if not name:
console.print("[red]Usage:[/red] /uninstall-skill <skill-name>")
console.print("[dim]Use /skills to see installed skills.[/dim]")
console.print()
return
result = uninstall_skill(name)
if result["success"]:
console.print(f"[green]Uninstalled:[/green] {name}")
console.print("[dim]Reload with /new to apply.[/dim]")
else:
console.print(f"[red]Failed:[/red] {result['error']}")
console.print()
def _cmd_install_skills(args: str = "") -> None:
"""Browse and install skills from the EvoSkills repository.
Args:
args: Optional tag name to pre-filter (e.g. "core").
Returns:
list of ``install_source`` strings selected by the user,
``None`` if the user cancelled at either phase, or
``[]`` if nothing was selectable / all-installed in the filter.
"""
from collections import Counter
import questionary
from prompt_toolkit.styles import Style as PtStyle
from questionary import Choice
from ..paths import GLOBAL_SKILLS_DIR, USER_SKILLS_DIR
from ..tools.skills_manager import fetch_remote_skill_index, install_skill
_PICKER_STYLE = PtStyle.from_dict(
{
"questionmark": "#888888",
"question": "",
"pointer": "bold",
"highlighted": "bold",
"text": "#888888",
"answer": "bold",
}
)
from .widgets.thread_selector import PICKER_STYLE
# Installed-item indicator style for disabled checkbox choices.
_INSTALLED_INDICATOR = ("fg:#4caf50", "✓ ")
@@ -190,46 +49,24 @@ def _cmd_install_skills(args: str = "") -> None:
return questionary.checkbox(
message,
choices=choices,
style=_PICKER_STYLE,
style=PICKER_STYLE,
qmark="❯",
**kwargs,
).ask()
finally:
InquirerControl._get_choice_tokens = original
# Step 1: Fetch remote index
console.print("[dim]Fetching skill index...[/dim]")
try:
index = fetch_remote_skill_index()
except Exception as e:
console.print(f"[red]Failed to fetch skill index: {e}[/red]")
console.print(
"[dim]Try installing directly: /install-skill EvoScientist/EvoSkills@skills[/dim]"
)
console.print()
return
pre_filter_tag = (pre_filter_tag or "").strip().lower()
if not index:
console.print("[yellow]No skills found in the repository.[/yellow]")
console.print()
return
# Detect already-installed skills (both global and workspace tiers)
installed_names: set[str] = set()
for skills_dir in (Path(GLOBAL_SKILLS_DIR), Path(USER_SKILLS_DIR)):
if skills_dir.exists():
installed_names.update(e.name for e in skills_dir.iterdir() if e.is_dir())
pre_filter_tag = args.strip().lower() if args else ""
# Step 2: Tag filter (skip if pre-filtered via args)
# Phase 1: tag filter (skip if pre-filtered via args)
if pre_filter_tag:
filtered = [
s for s in index if pre_filter_tag in [t.lower() for t in s.get("tags", [])]
]
if not filtered:
console.print(f"[yellow]No skills found with tag: {args.strip()}[/yellow]")
# Show available tags
console.print(
f"[yellow]No skills found with tag: {pre_filter_tag}[/yellow]"
)
tag_counter: Counter[str] = Counter()
for s in index:
for t in s.get("tags", []):
@@ -238,10 +75,8 @@ def _cmd_install_skills(args: str = "") -> None:
sorted_tags = sorted(tag_counter.items(), key=lambda x: (-x[1], x[0]))
tags_str = ", ".join(f"{tag} ({count})" for tag, count in sorted_tags)
console.print(f"[dim]Available tags: {tags_str}[/dim]")
console.print()
return
return []
else:
# Build tag choices for interactive picker
tag_counter = Counter()
for s in index:
for t in s.get("tags", []):
@@ -255,13 +90,12 @@ def _cmd_install_skills(args: str = "") -> None:
selected_tag = questionary.select(
"Filter by tag:",
choices=tag_choices,
style=_PICKER_STYLE,
style=PICKER_STYLE,
qmark="❯",
).ask()
if selected_tag is None:
console.print()
return
return None
if selected_tag == "__all__":
filtered = index
@@ -272,14 +106,12 @@ def _cmd_install_skills(args: str = "") -> None:
if selected_tag in [t.lower() for t in s.get("tags", [])]
]
# Step 3: Skill selection checkbox
all_installed = all(s["name"] in installed_names for s in filtered)
if all_installed:
# Phase 2: skill selection checkbox
if all(s["name"] in installed_names for s in filtered):
console.print(
"[green]All skills in this category are already installed.[/green]"
)
console.print()
return
return []
choices = []
for s in filtered:
@@ -305,31 +137,5 @@ def _cmd_install_skills(args: str = "") -> None:
selected = _checkbox_ask(choices, "Select skills to install:")
if selected is None:
console.print()
return
if not selected:
console.print("[dim]No skills selected.[/dim]")
console.print()
return
# Step 4: Install selected skills (default: global)
installed_count = 0
for source in selected:
result = install_skill(source, global_install=True)
if result.get("batch"):
for item in result.get("installed", []):
console.print(f"[green]Installed:[/green] {item['name']}")
installed_count += 1
for item in result.get("failed", []):
console.print(f"[red]Failed:[/red] {item['name']} — {item['error']}")
elif result.get("success"):
console.print(f"[green]Installed:[/green] {result['name']}")
installed_count += 1
else:
console.print(f"[red]Failed:[/red] {result.get('error', 'unknown')}")
if installed_count:
console.print(f"\n[green]{installed_count} skill(s) installed.[/green]")
console.print("[dim]Reload with /new to apply.[/dim]")
console.print()
return None
return list(selected)
+110 -13
View File
@@ -4,7 +4,7 @@ from __future__ import annotations
from dataclasses import dataclass, replace
from datetime import datetime
from typing import Any
from typing import TYPE_CHECKING, Any
from langchain_core.messages import AIMessage, HumanMessage
from langchain_core.messages.utils import count_tokens_approximately
@@ -13,7 +13,15 @@ from ..llm.context_window import (
DEFAULT_CONTEXT_WINDOW_FALLBACK,
resolve_context_window,
)
from ..sessions import get_thread_messages
from ..memory.worker_activity import (
MemoryWorkerStatusSnapshot,
ObservationLinkerStatusSnapshot,
memory_worker_status,
observation_linker_status,
)
if TYPE_CHECKING:
from ..gateway import GraphGateway
_FALLBACK_CONTEXT_WINDOW = DEFAULT_CONTEXT_WINDOW_FALLBACK
STATUS_BAR_BG = "#171a20"
@@ -26,6 +34,11 @@ STATUS_BAD = "#d08c61"
STATUS_CRITICAL = "#d86f6f"
STATUS_HINT_IDLE = "#8b9bb0"
STATUS_HINT_BUSY = "#f0c36a"
STATUS_HINT_WRITING = "#7eb8e0"
# Braille spinner frames used by the CLI bottom toolbar and TUI status bar
# to animate the "Loading MCP tools" indicator.
SPINNER_FRAMES = "\u280b\u2819\u2839\u2838\u283c\u2834\u2826\u2827\u2807\u280f"
@dataclass(slots=True)
@@ -86,15 +99,15 @@ def format_token_count_compact(value: int) -> str:
"""Format large token counts into a compact human-readable form."""
abs_value = abs(int(value))
if abs_value >= 1_000_000:
num = value / 1_000_000
num = float(value) / 1_000_000
suffix = "M"
elif abs_value >= 1_000:
num = value / 1_000
num = float(value) / 1_000
suffix = "K"
else:
return str(value)
if num.is_integer():
if num == int(num):
return f"{int(num)}{suffix}"
return f"{num:.1f}{suffix}"
@@ -155,15 +168,10 @@ def trim_status_text(text: str, max_width: int) -> str:
if max_width <= ellipsis_width:
return ellipsis[:max_width]
try:
from prompt_toolkit.utils import get_cwidth
except Exception:
get_cwidth = None
out: list[str] = []
width = 0
for ch in text:
ch_width = get_cwidth(ch) if get_cwidth else len(ch)
ch_width = _display_width(ch)
if width + ch_width + ellipsis_width > max_width:
break
out.append(ch)
@@ -171,13 +179,94 @@ def trim_status_text(text: str, max_width: int) -> str:
return "".join(out).rstrip() + ellipsis
def get_memory_worker_status() -> MemoryWorkerStatusSnapshot | None:
"""Read completed EvoMemory save counts without making rendering fail."""
try:
return memory_worker_status()
except Exception:
return None
def get_observation_linker_status() -> ObservationLinkerStatusSnapshot | None:
"""Read active observation-linker status without making rendering fail."""
try:
return observation_linker_status()
except Exception:
return None
def _plural(count: int, singular: str, plural: str | None = None) -> str:
word = singular if count == 1 else (plural or f"{singular}s")
return f"{count} {word}"
def _memory_activity_label(
*,
worker_status: MemoryWorkerStatusSnapshot | None,
linker_status: ObservationLinkerStatusSnapshot | None,
) -> str:
parts: list[str] = []
if worker_status is not None and worker_status.is_running:
parts.append("🧠")
if linker_status is not None and linker_status.is_running:
parts.append("🔗")
saved: list[str] = []
if worker_status is not None:
if worker_status.profile_updates:
saved.append(_plural(worker_status.profile_updates, "profile edit"))
if worker_status.observations_recorded:
saved.append(_plural(worker_status.observations_recorded, "observation"))
if saved:
parts.append(f"Saved {', '.join(saved)}")
if linker_status is not None and linker_status.relations_linked:
parts.append(
f"Created {_plural(linker_status.relations_linked, 'memory link')}"
)
return " ".join(parts)
def _append_memory_indicator(
frags: list[tuple[str, str]],
*,
worker_status: MemoryWorkerStatusSnapshot | None,
linker_status: ObservationLinkerStatusSnapshot | None,
width: int,
) -> None:
if worker_status is None and linker_status is None:
return
label = _memory_activity_label(
worker_status=worker_status,
linker_status=linker_status,
)
if not label:
return
tail: list[tuple[str, str]] = []
if frags and frags[-1] == ("class:status-bar", " "):
tail.append(frags.pop())
separator = " │ " if width >= 76 else " · "
frags.extend(
[
("class:status-bar-dim", separator),
("class:status-bar-warn", label),
]
)
frags.extend(tail)
def build_status_fragments(
snapshot: SessionStatusSnapshot,
started_at: datetime,
width: int,
) -> list[tuple[str, str]]:
"""Build prompt_toolkit formatted-text fragments for the status bar."""
duration_label = format_duration_compact(started_at)
now = datetime.now()
duration_label = format_duration_compact(started_at, now=now)
percent = snapshot.context_percent
percent_label = f"{percent}%"
if width < 52:
@@ -215,6 +304,13 @@ def build_status_fragments(
("class:status-bar", " "),
]
_append_memory_indicator(
frags,
worker_status=get_memory_worker_status(),
linker_status=get_observation_linker_status(),
width=width,
)
total_width = sum(_display_width(text) for _, text in frags)
if total_width > width:
plain_text = "".join(text for _, text in frags)
@@ -345,11 +441,12 @@ async def build_session_status_snapshot(
model_name: str | None = None,
model_obj: Any | None = None,
pending_user_text: str | None = None,
graph_gateway: GraphGateway,
) -> SessionStatusSnapshot:
"""Count current thread context and return a display snapshot."""
resolved_name = _resolve_model_name(model_name, model_obj)
window = _resolve_context_window(model_obj)
messages = list(await get_thread_messages(thread_id))
messages = list(await graph_gateway.get_thread_messages(thread_id))
pending = (pending_user_text or "").strip()
if pending:
+7
View File
@@ -6,6 +6,7 @@ from collections.abc import Callable
from dataclasses import dataclass
from typing import Any, Protocol
from ..gateway import GraphGateway
from ..stream.display import _run_streaming
@@ -30,6 +31,8 @@ class StreamingTUIBackend(Protocol):
metadata: dict | None = None,
hitl_prompt_fn: Callable[[list], list[dict] | None] | None = None,
ask_user_prompt_fn: Callable[[dict], dict] | None = None,
cancel_scope: str | None = None,
gateway: GraphGateway,
) -> str:
"""Run streaming and return final response text."""
@@ -56,6 +59,8 @@ class RichStreamingBackend:
metadata: dict | None = None,
hitl_prompt_fn: Callable[[list], list[dict] | None] | None = None,
ask_user_prompt_fn: Callable[[dict], dict] | None = None,
cancel_scope: str | None = None,
gateway: GraphGateway,
) -> str:
return _run_streaming(
agent=agent,
@@ -71,4 +76,6 @@ class RichStreamingBackend:
metadata=metadata,
hitl_prompt_fn=hitl_prompt_fn,
ask_user_prompt_fn=ask_user_prompt_fn,
cancel_scope=cancel_scope,
gateway=gateway,
)
File diff suppressed because it is too large Load Diff
+13 -2
View File
@@ -5,11 +5,16 @@ from __future__ import annotations
from collections.abc import Callable
from typing import Any
from ..stream.display import console
from ..gateway import GraphGateway
from ..stream.console import console
from .tui_backends import RichStreamingBackend, StreamingTUIBackend
DEFAULT_UI_BACKEND = "cli"
SUPPORTED_UI_BACKENDS = ("cli", "tui")
# "webui" launches the browser front-end instead of an in-terminal UI; it is
# intercepted earlier (cli/commands.py:_main_callback) and never reaches the
# streaming backends, but is listed here so normalize/resolve preserve it
# rather than falling back to "cli".
SUPPORTED_UI_BACKENDS = ("cli", "tui", "webui")
_LEGACY_BACKEND_MAP = {"textual": "tui", "rich": "cli"}
@@ -74,6 +79,8 @@ def run_streaming(
metadata: dict | None = None,
hitl_prompt_fn: Callable[[list], list[dict] | None] | None = None,
ask_user_prompt_fn: Callable[[dict], dict] | None = None,
cancel_scope: str | None = None,
gateway: GraphGateway,
) -> str:
"""Run streaming with the selected backend."""
backend = get_backend(ui_backend, warn_fallback=True)
@@ -92,6 +99,8 @@ def run_streaming(
metadata=metadata,
hitl_prompt_fn=hitl_prompt_fn,
ask_user_prompt_fn=ask_user_prompt_fn,
cancel_scope=cancel_scope,
gateway=gateway,
)
except RuntimeError:
requested = normalize_ui_backend(ui_backend)
@@ -113,5 +122,7 @@ def run_streaming(
metadata=metadata,
hitl_prompt_fn=hitl_prompt_fn,
ask_user_prompt_fn=ask_user_prompt_fn,
cancel_scope=cancel_scope,
gateway=gateway,
)
raise
+2
View File
@@ -6,6 +6,7 @@ from .assistant_message import AssistantMessage
from .compact_summary_widget import CompactSummaryWidget
from .compacting_widget import CompactingWidget
from .loading_widget import LoadingWidget
from .mcp_loader_widget import MCPLoaderWidget
from .subagent_widget import SubAgentWidget
from .summarization_widget import SummarizationWidget
from .system_message import SystemMessage
@@ -23,6 +24,7 @@ __all__ = [
"CompactSummaryWidget",
"CompactingWidget",
"LoadingWidget",
"MCPLoaderWidget",
"SubAgentWidget",
"SummarizationWidget",
"SystemMessage",
+3 -15
View File
@@ -105,11 +105,7 @@ class ApprovalWidget(Widget):
self._option_widgets = []
count = len(self._action_requests)
if count == 1:
name = (
self._action_requests[0].get("name", "")
if isinstance(self._action_requests[0], dict)
else getattr(self._action_requests[0], "name", "")
)
name = self._action_requests[0].get("name", "")
title = f">>> {name} Requires Approval <<<"
else:
title = f">>> {count} Tool Calls Require Approval <<<"
@@ -117,16 +113,8 @@ class ApprovalWidget(Widget):
# Show each action request as a compact line
for req in self._action_requests:
name = (
req.get("name", "")
if isinstance(req, dict)
else getattr(req, "name", "")
)
args = (
req.get("args", {})
if isinstance(req, dict)
else getattr(req, "args", {})
)
name = req.get("name", "")
args = req.get("args", {})
if isinstance(args, dict):
command = args.get("command", args.get("path", ""))
else:
+10 -4
View File
@@ -5,6 +5,7 @@ from __future__ import annotations
from textual.containers import Vertical
from textual.widgets import Markdown
from ...stream.display import _fix_markdown_heading_spacing
from .timestamp_mixin import TimestampClickMixin
@@ -39,8 +40,11 @@ class AssistantMessage(TimestampClickMixin, Vertical):
yield Markdown("")
def on_mount(self) -> None:
"""Render ``initial_content`` once the widget enters the DOM."""
if self._content:
self.query_one(Markdown).update(self._content)
self.query_one(Markdown).update(
_fix_markdown_heading_spacing(self._content)
)
async def append_content(self, text: str) -> None:
"""Append text and schedule a debounced Markdown re-render."""
@@ -50,12 +54,14 @@ class AssistantMessage(TimestampClickMixin, Vertical):
self.set_timer(0.1, self._flush_markdown)
def _flush_markdown(self) -> None:
"""Flush accumulated content to the Markdown widget."""
"""Flush accumulated content to the Markdown widget on a display copy."""
self._flush_pending = False
self.query_one(Markdown).update(self._content)
self.query_one(Markdown).update(_fix_markdown_heading_spacing(self._content))
async def stop_stream(self) -> None:
"""Finalize the stream — ensure final content is rendered."""
self._flush_pending = False
if self._content:
self.query_one(Markdown).update(self._content)
self.query_one(Markdown).update(
_fix_markdown_heading_spacing(self._content)
)
@@ -0,0 +1,191 @@
"""Live-updating widget that shows per-server MCP load progress.
Mounted above the chat input while MCP tools are being fetched in the
background. Re-renders on a 100 ms tick so the spinner animates and the
per-server states transition smoothly from pending → ok/error.
When the load finishes:
- All-success runs auto-dismiss after a short grace period so the chat
area isn't permanently crowded.
- Failures stick around longer so the user has time to read the error
detail, then auto-dismiss — otherwise the widget pins itself above
the input forever.
"""
from __future__ import annotations
import time
from rich.text import Text
from textual.widgets import Static
from ..status_bar import SPINNER_FRAMES
_DIM = "#7c8594"
_STRONG = "#e5e7eb"
_GOOD = "#5fcf8b"
_WARN = "#d7b45a"
_BAD = "#d86f6f"
# How long to wait after an all-success load before auto-dismissing.
_AUTO_DISMISS_SECONDS = 2.5
# Longer grace on failure so the user has time to read error detail.
_AUTO_DISMISS_ON_ERROR_SECONDS = 12.0
class MCPLoaderWidget(Static):
"""Shows a header line + one line per MCP server with its live status."""
DEFAULT_CSS = """
MCPLoaderWidget {
height: auto;
padding: 0 1;
margin: 0 0 1 0;
}
"""
TICK_SECONDS = 0.1
def __init__(self, servers: list[str]) -> None:
# server_name -> (state, detail); state ∈ {"pending","ok","error"}.
self._progress: dict[str, tuple[str, str]] = dict.fromkeys(
servers, ("pending", "")
)
self._frame = 0
self._tick_handle = None
self._finished = False
self._dismissed = False
self._auto_dismiss_at: float | None = None
# Seed with real content so Textual can measure us before the
# first tick; ``self.update()`` during ``__init__`` is unsafe
# (widget isn't attached yet), but we can pass the renderable
# straight into ``Static.__init__``.
super().__init__(self._build_renderable())
def on_mount(self) -> None:
self._tick_handle = self.set_interval(self.TICK_SECONDS, self._tick)
def on_unmount(self) -> None:
if self._tick_handle is not None:
self._tick_handle.stop()
self._tick_handle = None
# ── Public API ───────────────────────────────────────────────────
@property
def dismissed(self) -> bool:
"""Whether the widget has already removed itself from the DOM."""
return self._dismissed
def update_server(self, name: str, state: str, detail: str = "") -> None:
"""Record a progress event for one server and re-render."""
if self._dismissed or state not in ("pending", "ok", "error"):
return
# First-time-seen servers (e.g., ones missing from the initial
# prime set because the config file changed mid-load) just get
# appended — order stays stable for already-known entries.
self._progress[name] = (state, detail)
self._refresh_content()
def mark_finished(self) -> None:
"""Call once the background load task resolves (success or error).
If nothing ever progressed past ``pending``, the load was served
from cache (no events emitted) — drop the widget immediately
instead of flashing a misleading "0/N loaded" header.
Otherwise schedule an auto-dismiss: short on full success so the
chat area isn't cluttered, longer on failure so the user has
time to read the error detail before it goes away.
"""
if self._finished:
return
self._finished = True
progressed = any(state != "pending" for state, _ in self._progress.values())
if not progressed:
self._dismiss()
return
has_errors = any(state == "error" for state, _ in self._progress.values())
delay = _AUTO_DISMISS_ON_ERROR_SECONDS if has_errors else _AUTO_DISMISS_SECONDS
self._auto_dismiss_at = time.monotonic() + delay
self._refresh_content()
# ── Internal ─────────────────────────────────────────────────────
def _tick(self) -> None:
self._frame = (self._frame + 1) % len(SPINNER_FRAMES)
if (
self._auto_dismiss_at is not None
and time.monotonic() >= self._auto_dismiss_at
):
self._auto_dismiss_at = None
self._dismiss()
return
if not self._finished:
self._refresh_content()
def _dismiss(self) -> None:
"""Stop the tick timer and detach from the DOM.
Sets :attr:`dismissed` so the app can clear its widget reference
and late progress events become no-ops.
"""
if self._dismissed:
return
self._dismissed = True
if self._tick_handle is not None:
self._tick_handle.stop()
self._tick_handle = None
# Fire-and-forget remove — nothing awaits us.
self.remove()
def _build_renderable(self) -> Text:
spinner = SPINNER_FRAMES[self._frame]
pending = sum(1 for state, _ in self._progress.values() if state == "pending")
total = len(self._progress)
done = total - pending
header = Text()
if self._finished:
errors = sum(1 for state, _ in self._progress.values() if state == "error")
if errors:
header.append("✗ MCP ", style=f"{_BAD} bold")
header.append(
f"{done - errors}/{total} loaded, {errors} failed",
style=_STRONG,
)
else:
header.append("✓ MCP ", style=f"{_GOOD} bold")
header.append(f"{done}/{total} servers loaded", style=_STRONG)
else:
header.append(f"{spinner} ", style=f"{_WARN} bold")
header.append("Loading MCP tools ", style=_STRONG)
header.append(f"{done}/{total}", style=_DIM)
lines: list[Text] = [header]
for name, (state, detail) in self._progress.items():
line = Text(" ")
if state == "pending":
line.append(f"{spinner} ", style=_WARN)
line.append(name, style=_DIM)
elif state == "ok":
line.append("✓ ", style=_GOOD)
line.append(name, style=_STRONG)
if detail:
line.append(f" {detail} tools", style=_DIM)
else: # error
line.append("✗ ", style=_BAD)
line.append(name, style=_STRONG)
if detail:
summary = detail if len(detail) <= 80 else detail[:77] + "…"
line.append(f" {summary}", style=_BAD)
lines.append(line)
return Text("\n").join(lines)
def _refresh_content(self) -> None:
# NB: don't name this ``_render`` — that shadows Textual's internal
# ``Widget._render`` which must return a ``Visual``. Silently
# breaking that contract triggers ``'NoneType' object has no
# attribute 'get_height'`` during layout.
self.update(self._build_renderable())
+4 -3
View File
@@ -22,7 +22,7 @@ class SubAgentWidget(Vertical):
┌ ▶ Cooking with research-agent — Search literature ─┐
│ ✓ 8 completed │
│ ● web_search query="LLM attention" │
│ ● tavily_search query="LLM attention" │
│ ✓ 3 results │
└─────────────────────────────────────────────────────┘
@@ -208,10 +208,11 @@ class SubAgentWidget(Vertical):
widget.set_success(content)
else:
widget.set_error(content)
# Move from running to completed
# Move from running to completed (dedup guards against repeat
# deliveries of the same tool result inflating the collapse summary).
if matched_key and matched_key in self._running_ids:
self._running_ids.remove(matched_key)
if matched_key:
if matched_key and matched_key not in self._completed_ids:
self._completed_ids.append(matched_key)
self._update_visibility()
+19 -2
View File
@@ -18,6 +18,7 @@ from __future__ import annotations
from typing import TYPE_CHECKING, Any, ClassVar
from prompt_toolkit.styles import Style as PtStyle # type: ignore[import-untyped]
from rich.text import Text
from textual.binding import Binding, BindingType
from textual.containers import Container
@@ -30,6 +31,22 @@ if TYPE_CHECKING:
from textual.app import ComposeResult
# Style for questionary pickers used by the Rich CLI ``/resume`` and
# ``/delete`` interactive selectors. Matches the slash-completion menu's
# visual language: gray (#888888) for non-selected, bold for selected,
# no background changes.
PICKER_STYLE = PtStyle.from_dict(
{
"questionmark": "#888888",
"question": "",
"pointer": "bold",
"highlighted": "bold",
"text": "#888888",
"answer": "bold",
}
)
# ---------------------------------------------------------------------------
# Path helpers
# ---------------------------------------------------------------------------
@@ -185,9 +202,9 @@ def build_row_text(
indented: bool = False,
) -> Text:
"""Thread row. *indented* adds extra leading space for L2-grouped rows."""
from ...sessions import _format_relative_time
from ...sessions import _format_relative_time, short_thread_id
tid = thread["thread_id"]
tid = short_thread_id(thread["thread_id"])
preview = thread.get("preview", "") or ""
msgs = thread.get("message_count", 0)
model = thread.get("model", "") or ""
+9 -1
View File
@@ -16,7 +16,12 @@ class UsageWidget(Static):
}
"""
def __init__(self, input_tokens: int, output_tokens: int) -> None:
def __init__(
self,
input_tokens: int,
output_tokens: int,
elapsed: str | None = None,
) -> None:
stats = Text(justify="right")
stats.append("[", style="dim italic")
stats.append("Usage: ", style="dim italic")
@@ -24,5 +29,8 @@ class UsageWidget(Static):
stats.append(" in · ", style="dim italic")
stats.append(f"{output_tokens:,}", style="green italic")
stats.append(" out", style="dim italic")
if elapsed:
stats.append(" · ", style="dim italic")
stats.append(f"Elapsed: {elapsed}", style="dim italic")
stats.append("]", style="dim italic")
super().__init__(stats)
@@ -0,0 +1,39 @@
"""Transient widget shown while a /resume restarts the langgraph dev subprocess.
Mirrors ``CompactingWidget`` — a timer-backed status line that ticks elapsed
seconds so the user has live feedback during the up-to-60s langgraph dev
workspace sync (subprocess stop + restart so deployed sub-agents see the
resumed thread's workspace).
"""
from __future__ import annotations
from .timed_status_widget import TimedStatusWidget
class WorkspaceSyncWidget(TimedStatusWidget):
"""Timer-backed status line for an in-progress workspace sync."""
DEFAULT_CSS = """
WorkspaceSyncWidget {
height: auto;
color: #94a3b8;
padding: 0 0;
margin: 0 0 1 0;
}
"""
def __init__(self) -> None:
super().__init__()
def _refresh_display(self) -> None:
self.update(
f"Syncing async sub-agent server to resumed workspace... "
f"({self.elapsed_seconds}s)"
)
async def cleanup(self) -> None:
"""Stop timer and remove from DOM."""
self._stop_timer()
if self.is_mounted:
await self.remove()
+2 -1
View File
@@ -1,7 +1,7 @@
from __future__ import annotations
from . import implementation
from .base import Argument, Command, CommandContext, CommandUI
from .base import Argument, Command, CommandContext, CommandUI, SubCommand
from .channel_ui import ChannelCommandUI
from .manager import CommandManager, manager
@@ -12,6 +12,7 @@ __all__ = [
"CommandContext",
"CommandManager",
"CommandUI",
"SubCommand",
"implementation",
"manager",
]
+177
View File
@@ -0,0 +1,177 @@
from __future__ import annotations
from dataclasses import dataclass
from enum import StrEnum
class CompletionKind(StrEnum):
"""Discriminator for the kind of completion result."""
COMMANDS = "commands"
SUBCOMMANDS = "subcommands"
EMPTY = "empty"
_CATEGORY_ORDER = ["Session", "Skills", "MCP", "Channels", "General"]
@dataclass(frozen=True)
class CompletionCandidate:
"""A single completion suggestion with its replacement range."""
text: str
description: str
replace_start: int
replace_end: int
category: str = ""
@dataclass(frozen=True)
class CompletionResult:
"""The result of parsing a slash command input for completions."""
kind: CompletionKind
candidates: list[CompletionCandidate]
def compute_completions(text: str, cursor_pos: int) -> CompletionResult:
"""Parse *text* up to *cursor_pos* and return completion candidates.
This is the shared engine used by both the Rich CLI
(``SlashCommandCompleter``) and the TUI (``on_text_area_changed``).
Both thin adapters only need to translate the returned candidates
into their respective render/apply primitives.
"""
from .manager import manager as cmd_manager
before = text[:cursor_pos]
if not before.startswith("/"):
return CompletionResult(CompletionKind.EMPTY, [])
parts = before.split()
if not parts:
return CompletionResult(CompletionKind.EMPTY, [])
cmd_name = parts[0].lower()
has_trailing_space = before.endswith(" ")
# --- Top-level command completion ---
if len(parts) == 1:
prefix = before.lower().rstrip()
# Match commands by canonical name AND aliases
by_cat: dict[str, list[tuple[str, str]]] = {}
for cmd in cmd_manager.get_all_commands():
all_names = [cmd.name.lower()] + [
a.lower() if a.startswith("/") else f"/{a.lower()}" for a in cmd.alias
]
if any(n.startswith(prefix) for n in all_names):
by_cat.setdefault(cmd.category, []).append((cmd.name, cmd.description))
# Whether the typed prefix resolves to an exact command/alias
exact_cmd = cmd_manager.get_command(prefix)
if exact_cmd and not has_trailing_space:
return CompletionResult(CompletionKind.EMPTY, [])
if exact_cmd and has_trailing_space:
completions = exact_cmd.get_completions([""])
if completions:
insert_pos = len(before)
return CompletionResult(
CompletionKind.SUBCOMMANDS,
[
CompletionCandidate(
text=name,
description=desc,
replace_start=insert_pos,
replace_end=insert_pos,
)
for name, desc in completions
],
)
return CompletionResult(CompletionKind.EMPTY, [])
all_matches = [v for vs in by_cat.values() for v in vs]
if not all_matches:
return CompletionResult(CompletionKind.EMPTY, [])
# Build candidates ordered by category
candidates: list[CompletionCandidate] = []
for cat in _CATEGORY_ORDER:
for cmd_text, desc in by_cat.get(cat, []):
candidates.append(
CompletionCandidate(
text=cmd_text,
description=desc,
replace_start=0,
replace_end=len(before),
category=cat,
)
)
for cat, items in by_cat.items():
if cat not in _CATEGORY_ORDER:
for cmd_text, desc in items:
candidates.append(
CompletionCandidate(
text=cmd_text,
description=desc,
replace_start=0,
replace_end=len(before),
category=cat,
)
)
return CompletionResult(CompletionKind.COMMANDS, candidates)
# --- Subcommand / argument completion (len(parts) >= 2) ---
cmd = cmd_manager.get_command(cmd_name)
if cmd is None:
return CompletionResult(CompletionKind.EMPTY, [])
# Delegate to Command.get_completions for all depths
tokens = parts[1:]
if has_trailing_space:
tokens.append("")
completions = cmd.get_completions(tokens)
if not completions:
return CompletionResult(CompletionKind.EMPTY, [])
# Compute replacement range.
if tokens[-1]:
# User is typing a partial — replace it
sub_start = before.rfind(tokens[-1])
if sub_start < 0:
sub_start = len(before)
replace_end = len(before)
elif len(tokens) >= 2 and tokens[-2]:
# Trailing space after a token. Check if the previous token is
# a known subcommand name — if so, the completion is for the
# NEXT argument (insert at cursor). If not, the completions
# refine the partial (replace it).
prev = tokens[-2]
is_known_sub = any(sc.name == prev for sc in cmd.subcommands)
if not is_known_sub:
sub_start = before.rfind(prev)
if sub_start < 0:
sub_start = len(before)
else:
sub_start = len(before)
replace_end = len(before)
else:
sub_start = len(before)
replace_end = len(before)
return CompletionResult(
CompletionKind.SUBCOMMANDS,
[
CompletionCandidate(
text=name,
description=desc,
replace_start=sub_start,
replace_end=replace_end,
)
for name, desc in completions
],
)
+81 -3
View File
@@ -1,8 +1,11 @@
from __future__ import annotations
from abc import ABC, abstractmethod
from dataclasses import dataclass
from typing import Any, ClassVar, Protocol, runtime_checkable
from dataclasses import dataclass, field
from typing import TYPE_CHECKING, Any, ClassVar, Protocol, runtime_checkable
if TYPE_CHECKING:
from ..gateway import GraphGateway
@dataclass
@@ -15,6 +18,15 @@ class Argument:
required: bool = True
@dataclass
class SubCommand:
"""A subcommand of a parent slash command."""
name: str
description: str
arguments: list[Argument] = field(default_factory=list)
@runtime_checkable
class CommandUI(Protocol):
"""Protocol for UI operations that commands can perform."""
@@ -38,13 +50,29 @@ class CommandUI(Protocol):
def clear_chat(self) -> None: ...
def request_quit(self) -> None: ...
def force_quit(self) -> None: ...
def start_new_session(self) -> None: ...
async def start_new_session(self) -> None: ...
async def handle_session_resume(
self, thread_id: str, workspace_dir: str | None = None
) -> None: ...
async def flush(self) -> None: ...
@dataclass
class ChannelRuntime:
"""Mutable handle to the agent + thread bound to running channels."""
agent: Any = None
thread_id: str | None = None
def bind(self, agent: Any, thread_id: str) -> None:
self.agent = agent
self.thread_id = thread_id
def clear(self) -> None:
self.agent = None
self.thread_id = None
@dataclass
class CommandContext:
"""Context passed to commands during execution."""
@@ -55,6 +83,9 @@ class CommandContext:
workspace_dir: str | None = None
checkpointer: Any = None
config: Any = None
channel_runtime: ChannelRuntime | None = None
graph_gateway: GraphGateway | None = None
command_error: str | None = None
# Real LLM input token count from last usage_metadata (includes system
# prompt + tool schemas). Used by /compact for accurate display.
input_tokens_hint: int | None = None
@@ -67,6 +98,53 @@ class Command(ABC):
alias: ClassVar[list[str]] = []
description: str
arguments: ClassVar[list[Argument]] = []
category: ClassVar[str] = "General"
subcommands: ClassVar[list[SubCommand]] = []
# When False, callers may dispatch this command without waiting for
# the background agent load to finish — important so recovery
# commands like ``/mcp add`` can run even when the MCP load is
# failing and ``_await_agent_ready`` would hang.
requires_agent: ClassVar[bool] = False
def needs_agent(self, args: list[str]) -> bool:
"""Whether this specific invocation needs the agent.
Default returns :attr:`requires_agent`. Override when a command
has a mix of agent-using and agent-free subcommands (e.g.
``/channel start`` vs ``/channel status``).
"""
return self.requires_agent
def get_completions(self, tokens: list[str]) -> list[tuple[str, str]]:
"""Return completions for args typed after the command name.
Default walks :attr:`subcommands` for the first positional token
only. Override for deeper levels (e.g. server names, thread IDs).
"""
if not self.subcommands:
return []
if len(tokens) <= 1:
prefix = tokens[0].lower() if tokens else ""
matches = [
(sc.name, sc.description)
for sc in self.subcommands
if sc.name.startswith(prefix)
]
# Exact match — subcommand already complete, hide popup
if len(matches) == 1 and matches[0][0] == prefix:
return []
return matches
# partial + trailing space: /mcp a → still show "add"
if len(tokens) == 2 and tokens[1] == "":
prefix = tokens[0].lower()
if any(sc.name == prefix for sc in self.subcommands):
return []
return [
(sc.name, sc.description)
for sc in self.subcommands
if sc.name.startswith(prefix)
]
return []
@abstractmethod
async def execute(self, ctx: CommandContext, args: list[str]) -> None:
+99 -7
View File
@@ -1,14 +1,23 @@
from __future__ import annotations
import asyncio
from typing import Any
import logging
from collections.abc import Awaitable, Callable
from typing import TYPE_CHECKING, Any
from .base import CommandUI
if TYPE_CHECKING:
from ..gateway import GraphGateway
_logger = logging.getLogger(__name__)
class ChannelCommandUI(CommandUI):
"""CommandUI implementation for messaging channels with output buffering."""
_TEXT_CHUNK_LIMIT = 3500
@property
def supports_interactive(self) -> bool:
return False
@@ -16,24 +25,67 @@ class ChannelCommandUI(CommandUI):
def __init__(
self,
channel_msg: Any,
*,
graph_gateway: GraphGateway,
append_system_callback: Any = None,
start_new_session_callback: Any = None,
start_new_session_callback: Callable[[], Awaitable[None]] | None = None,
handle_session_resume_callback: Any = None,
):
self.msg = channel_msg
self.append_system_callback = append_system_callback
self.start_new_session_callback = start_new_session_callback
self.handle_session_resume_callback = handle_session_resume_callback
self.graph_gateway = graph_gateway
self._system_buffer: list[str] = []
def append_system(self, text: str, style: str = "dim") -> None:
if self.append_system_callback:
def _queue_system(
self,
text: str,
style: str = "dim",
*,
mirror_local: bool = True,
) -> None:
if mirror_local and self.append_system_callback:
self.append_system_callback(text, style)
# Buffer the text for grouped delivery to the channel
# We ignore style for grouping but keep it for individual lines if needed
self._system_buffer.append(text)
def append_system(self, text: str, style: str = "dim") -> None:
self._queue_system(text, style)
@staticmethod
def _extract_message_text(message: Any) -> str:
content = getattr(message, "content", "") or ""
if isinstance(content, list):
parts = [
block.get("text", "")
for block in content
if isinstance(block, dict) and block.get("type") == "text"
]
content = " ".join(parts) if parts else ""
return str(content).strip()
async def _send_text_chunks(self, text: str, *, mirror_local: bool = True) -> None:
"""Flush long plain-text payloads in channel-safe chunks."""
text = (text or "").strip()
if not text:
return
pending = text
while pending:
chunk = pending[: self._TEXT_CHUNK_LIMIT]
if len(pending) > self._TEXT_CHUNK_LIMIT:
split_at = chunk.rfind("\n")
if split_at > 0:
chunk = chunk[:split_at]
chunk = chunk.rstrip()
if not chunk:
chunk = pending[: self._TEXT_CHUNK_LIMIT]
self._queue_system(chunk, mirror_local=mirror_local)
await self.flush()
pending = pending[len(chunk) :].lstrip("\n")
async def flush(self) -> None:
"""Send all buffered system messages as a single grouped message."""
if not self._system_buffer:
@@ -129,9 +181,9 @@ class ChannelCommandUI(CommandUI):
def force_quit(self) -> None:
self.request_quit()
def start_new_session(self) -> None:
async def start_new_session(self) -> None:
if self.start_new_session_callback:
self.start_new_session_callback()
await self.start_new_session_callback()
else:
self.append_system(
"New session requested. Please restart the channel link or use /new if supported."
@@ -140,5 +192,45 @@ class ChannelCommandUI(CommandUI):
async def handle_session_resume(
self, thread_id: str, workspace_dir: str | None = None
) -> None:
mirror_local = self.handle_session_resume_callback is None
if self.handle_session_resume_callback:
await self.handle_session_resume_callback(thread_id, workspace_dir)
lines = [f"Resumed session: {thread_id}"]
try:
messages = await self.graph_gateway.get_thread_messages(thread_id)
except Exception as exc:
_logger.exception(
"Failed to load saved history for resumed thread %s",
thread_id,
)
lines.append(f"(history unavailable: {exc})")
await self._send_text_chunks("\n".join(lines), mirror_local=mirror_local)
return
display = [m for m in messages if getattr(m, "type", None) in ("human", "ai")]
if not display:
if messages:
lines.append("No displayable messages in this session.")
else:
lines.append("No saved messages in this session.")
await self._send_text_chunks("\n".join(lines), mirror_local=mirror_local)
return
HISTORY_WINDOW = 10
if len(display) > HISTORY_WINDOW:
display = display[-HISTORY_WINDOW:]
lines.append(f"Conversation history (last {HISTORY_WINDOW} messages):")
else:
lines.append("Conversation history:")
for message in display:
text = self._extract_message_text(message)
if not text:
continue
if getattr(message, "type", None) == "human":
lines.append(f"User: {text}")
else:
lines.append(f"EvoScientist: {text}")
await self._send_text_chunks("\n".join(lines), mirror_local=mirror_local)
@@ -1,5 +1,21 @@
from __future__ import annotations
from . import channel, general, mcp, session, skills
from . import (
autoskills,
channel,
general,
mcp,
schedule,
session,
skills,
)
__all__ = ["channel", "general", "mcp", "session", "skills"]
__all__ = [
"autoskills",
"channel",
"general",
"mcp",
"schedule",
"session",
"skills",
]
@@ -0,0 +1,435 @@
from __future__ import annotations
import asyncio
from enum import Enum
from typing import ClassVar
from rich.table import Table
from ..base import Command, CommandContext, SubCommand
from ..manager import manager
AUTOSKILLS_COMMAND = "/autoskills"
_PROPOSAL_STATUSES = {
"review": "pending",
"approved": "approved",
"rejected": "rejected",
}
class AutoSkillsCommand(Command):
"""Manage EvoMemory AutoSkills proposals."""
name = AUTOSKILLS_COMMAND
alias: ClassVar[list[str]] = ["/skills-review"]
description = "Review EvoMemory autoskill proposals"
subcommands: ClassVar[list[SubCommand]] = [
SubCommand("status", "Show AutoSkills config and proposals for review"),
SubCommand("help", "Show AutoSkills command examples"),
SubCommand("list", "List autoskill proposals, optionally filtered by status"),
SubCommand("review", "Review autoskill proposals awaiting a decision"),
SubCommand("approve", "Approve an autoskill proposal by id"),
SubCommand("reject", "Reject an autoskill proposal by id"),
SubCommand("run", "Run AutoSkills once now"),
SubCommand("on", "Enable periodic AutoSkills"),
SubCommand("off", "Disable periodic AutoSkills"),
SubCommand("mode", "Set review or auto approval mode"),
SubCommand("cadence", "Set nightly, weekly, or monthly cadence"),
SubCommand("time", "Set local run time as HH:MM"),
]
async def execute(self, ctx: CommandContext, args: list[str]) -> None:
sub = args[0].lower() if args else "help"
rest = args[1:]
if sub in {"help", "-h", "--help", "?"}:
self._show_help(ctx)
elif sub in {"status", "show"}:
await self._status(ctx)
elif sub in {"list", "ls", "proposals"}:
await self._list_command(ctx, rest)
elif sub == "review":
await self._list(ctx, status="pending")
elif sub in {"approve", "accept"}:
await self._approve(ctx, self._first_arg(rest))
elif sub in {"reject", "deny", "decline"}:
await self._reject(ctx, self._first_arg(rest))
elif sub in {"run", "now"}:
await self._run(ctx)
elif sub in {"on", "enable"}:
await self._set_config(ctx, "memory_skill_synthesis_enabled", "true")
elif sub in {"off", "disable"}:
await self._set_config(ctx, "memory_skill_synthesis_enabled", "false")
elif sub == "mode":
await self._set_config(
ctx,
"memory_skill_synthesis_mode",
self._first_arg(rest),
)
elif sub in {"auto", "automatic"}:
await self._set_config(ctx, "memory_skill_synthesis_mode", "auto")
elif sub == "manual":
await self._set_config(ctx, "memory_skill_synthesis_mode", "review")
elif sub == "cadence":
await self._set_config(
ctx,
"memory_skill_synthesis_cadence",
self._first_arg(rest),
)
elif sub in {"nightly", "weekly", "monthly"}:
await self._set_config(ctx, "memory_skill_synthesis_cadence", sub)
elif sub == "time":
await self._set_config(
ctx,
"memory_skill_synthesis_time",
self._first_arg(rest),
)
else:
self._show_help(ctx, prefix=f"Unknown AutoSkills command: {sub}")
async def _status(self, ctx: CommandContext) -> None:
from ... import paths
from ...config import get_effective_config
from ...memory.autoskills.proposals import list_skill_proposals
from ...memory.autoskills.schedule import alist_autoskill_schedules
cfg = get_effective_config()
workspace_dir = self._workspace_dir(ctx)
pending = list_skill_proposals(
paths.MEMORIES_DIR,
status="pending",
workspace_dir=workspace_dir,
)
ctx.ui.append_system(
(
"AutoSkills: "
f"{'on' if cfg.memory_skill_synthesis_enabled else 'off'} | "
f"mode={cfg.memory_skill_synthesis_mode.value} | "
f"cadence={cfg.memory_skill_synthesis_cadence.value} | "
f"time={cfg.memory_skill_synthesis_time}"
),
style="dim",
)
ctx.ui.append_system(
f"AutoSkill proposal(s) ready for review: {len(pending)}",
style="yellow" if pending else "dim",
)
if pending:
ctx.ui.append_system(
(
f"Next: {AUTOSKILLS_COMMAND} review, then "
f"{AUTOSKILLS_COMMAND} approve <id> or "
f"{AUTOSKILLS_COMMAND} reject <id>."
),
style="dim",
)
elif cfg.memory_skill_synthesis_enabled:
ctx.ui.append_system(
f"Next: {AUTOSKILLS_COMMAND} run to search now, or "
f"{AUTOSKILLS_COMMAND} help for commands.",
style="dim",
)
else:
ctx.ui.append_system(
f"Next: {AUTOSKILLS_COMMAND} run to search once, or "
f"{AUTOSKILLS_COMMAND} on to enable scheduled runs.",
style="dim",
)
if cfg.memory_skill_synthesis_enabled:
try:
rows = await alist_autoskill_schedules(cfg, limit=1)
except Exception:
rows = []
if rows:
ctx.ui.append_system(
f"Background schedule id: {str(rows[0].get('cron_id', ''))[:8]}",
style="dim",
)
async def _list_command(self, ctx: CommandContext, args: list[str]) -> None:
if not args or args[0].lower() == "all":
await self._list(ctx)
return
status = _PROPOSAL_STATUSES.get(args[0].lower())
if status is None:
ctx.ui.append_system(
f"Usage: {AUTOSKILLS_COMMAND} list [review|approved|rejected|all]",
style="yellow",
)
return
await self._list(ctx, status=status)
async def _list(self, ctx: CommandContext, *, status: str | None = None) -> None:
from ... import paths
from ...memory.autoskills.proposals import list_skill_proposals
workspace_dir = self._workspace_dir(ctx)
proposals = list_skill_proposals(
paths.MEMORIES_DIR,
status=status,
workspace_dir=workspace_dir,
)
if not proposals:
if status:
label = self._status_label(status)
ctx.ui.append_system(
f"No autoskill proposals {label}.",
style="dim",
)
else:
ctx.ui.append_system("No autoskill proposals.", style="dim")
return
title = "EvoMemory AutoSkill Proposals"
if status:
title = (
f"EvoMemory AutoSkill Proposals {self._status_label(status).title()}"
)
table = Table(title=title, show_header=True)
table.add_column("ID", style="cyan")
table.add_column("Action", style="magenta")
table.add_column("AutoSkill", style="green")
table.add_column("Status", style="yellow")
table.add_column("Observations", justify="right")
table.add_column("Description", style="dim")
for proposal in proposals:
table.add_row(
proposal.proposal_id,
proposal.operation,
proposal.skill_name,
proposal.status,
str(len(proposal.source_observation_ids)),
proposal.description,
)
ctx.ui.mount_renderable(table)
ctx.ui.append_system(
f"Use {AUTOSKILLS_COMMAND} approve <id> or "
f"{AUTOSKILLS_COMMAND} reject <id>.",
style="dim",
)
async def _approve(self, ctx: CommandContext, proposal_id: str | None) -> None:
from ... import paths
from ...memory.autoskills.proposals import approve_skill_proposal
if not proposal_id:
ctx.ui.append_system(
f"Usage: {AUTOSKILLS_COMMAND} approve <id>",
style="yellow",
)
ctx.ui.append_system(
f"Run {AUTOSKILLS_COMMAND} review to copy a proposal ID.",
style="dim",
)
return
workspace_dir = self._workspace_dir(ctx)
result = await asyncio.to_thread(
approve_skill_proposal,
paths.MEMORIES_DIR,
proposal_id,
workspace_dir=workspace_dir,
)
if result.get("approved"):
verb = "Updated" if result.get("operation") == "update" else "Approved"
ctx.ui.append_system(
f"{verb} autoskill: {result['skill_name']} ({result['path']})",
style="green",
)
ctx.ui.append_system(
"Reload with /new to apply the new skill.", style="dim"
)
else:
ctx.ui.append_system(f"Approval failed: {result.get('error')}", style="red")
async def _reject(self, ctx: CommandContext, proposal_id: str | None) -> None:
from ... import paths
from ...memory.autoskills.proposals import reject_skill_proposal
if not proposal_id:
ctx.ui.append_system(
f"Usage: {AUTOSKILLS_COMMAND} reject <id>",
style="yellow",
)
ctx.ui.append_system(
f"Run {AUTOSKILLS_COMMAND} review to copy a proposal ID.",
style="dim",
)
return
workspace_dir = self._workspace_dir(ctx)
result = await asyncio.to_thread(
reject_skill_proposal,
paths.MEMORIES_DIR,
proposal_id,
workspace_dir=workspace_dir,
)
if result.get("rejected"):
ctx.ui.append_system(
f"Rejected proposal: {result['proposal_id']}",
style="green",
)
else:
ctx.ui.append_system(f"Reject failed: {result.get('error')}", style="red")
async def _run(self, ctx: CommandContext) -> None:
from ...config import get_effective_config
from ...memory.autoskills.schedule import arun_autoskill_now
workspace_dir = self._workspace_dir(ctx)
try:
result = await arun_autoskill_now(
get_effective_config(),
workspace_dir=workspace_dir,
)
except Exception as exc:
ctx.ui.append_system(f"Failed to start AutoSkills: {exc}", style="red")
return
ctx.ui.append_system(
f"Started AutoSkills run {result['run_id']}.",
style="green",
)
async def _set_config(
self,
ctx: CommandContext,
key: str,
value: str | None,
) -> None:
from ...config import get_effective_config, set_config_value
from ...memory.autoskills.schedule import reconcile_autoskill_schedule
workspace_dir = self._workspace_dir(ctx)
if not value:
cfg = get_effective_config()
current = self._display_value(getattr(cfg, key))
ctx.ui.append_system(
f"Current {self._config_label(key)}: {current}",
style="dim",
)
ctx.ui.append_system(
f"Usage: {self._config_usage(key)}",
style="yellow",
)
return
if not await asyncio.to_thread(set_config_value, key, value):
valid = self._config_values(key)
suffix = f" Valid values: {valid}." if valid else ""
ctx.ui.append_system(
f"Invalid value for {self._config_label(key)}: {value}.{suffix}",
style="red",
)
return
cfg = get_effective_config()
if ctx.config is not None and hasattr(ctx.config, key):
setattr(ctx.config, key, getattr(cfg, key))
await asyncio.to_thread(
reconcile_autoskill_schedule,
cfg,
workspace_dir=workspace_dir,
)
ctx.ui.append_system(
f"Updated {self._config_label(key)} = {self._display_value(getattr(cfg, key))}",
style="green",
)
@staticmethod
def _config_label(key: str) -> str:
labels = {
"memory_skill_synthesis_enabled": "AutoSkills",
"memory_skill_synthesis_mode": "AutoSkills mode",
"memory_skill_synthesis_cadence": "AutoSkills cadence",
"memory_skill_synthesis_time": "AutoSkills time",
}
return labels.get(key, key)
@staticmethod
def _display_value(value: object) -> object:
return getattr(value, "value", value)
@staticmethod
def _enum_values(enum_type: type[Enum], *, separator: str = ", ") -> str:
return separator.join(str(member.value) for member in enum_type)
@classmethod
def _config_usage(cls, key: str) -> str:
from ...config import MemorySkillSynthesisCadence, MemorySkillSynthesisMode
if key == "memory_skill_synthesis_mode":
values = cls._enum_values(MemorySkillSynthesisMode, separator="|")
return f"{AUTOSKILLS_COMMAND} mode {values}"
if key == "memory_skill_synthesis_cadence":
values = cls._enum_values(MemorySkillSynthesisCadence, separator="|")
return f"{AUTOSKILLS_COMMAND} cadence {values}"
if key == "memory_skill_synthesis_time":
return f"{AUTOSKILLS_COMMAND} time HH:MM"
return f"{AUTOSKILLS_COMMAND} <value>"
@classmethod
def _config_values(cls, key: str) -> str | None:
from ...config import MemorySkillSynthesisCadence, MemorySkillSynthesisMode
if key == "memory_skill_synthesis_mode":
return cls._enum_values(MemorySkillSynthesisMode)
if key == "memory_skill_synthesis_cadence":
return cls._enum_values(MemorySkillSynthesisCadence)
if key == "memory_skill_synthesis_time":
return "24-hour local time, for example 03:00"
return None
@staticmethod
def _status_label(status: str) -> str:
if status == "pending":
return "ready for review"
return status
@staticmethod
def _show_help(ctx: CommandContext, *, prefix: str | None = None) -> None:
if prefix:
ctx.ui.append_system(prefix, style="yellow")
ctx.ui.append_system(
(
f"Usage: {AUTOSKILLS_COMMAND} "
"[status|review|approve|reject|run|on|off|mode|cadence|time]"
),
style="bold",
)
table = Table(title="AutoSkills Commands", show_header=True)
table.add_column("Command", style="cyan")
table.add_column("Use when", style="dim")
rows = [
(AUTOSKILLS_COMMAND, "Show this command reference"),
(f"{AUTOSKILLS_COMMAND} status", "Show config and the next useful action"),
(f"{AUTOSKILLS_COMMAND} review", "Review proposals waiting for a decision"),
(f"{AUTOSKILLS_COMMAND} approve <id>", "Install a reviewed autoskill"),
(f"{AUTOSKILLS_COMMAND} reject <id>", "Dismiss a reviewed proposal"),
(f"{AUTOSKILLS_COMMAND} run", "Start a one-off background autoskill run"),
(f"{AUTOSKILLS_COMMAND} on|off", "Enable or disable scheduled runs"),
(f"{AUTOSKILLS_COMMAND} auto|manual", "Switch approval behavior"),
(
f"{AUTOSKILLS_COMMAND} nightly|weekly|monthly",
"Set the built-in schedule cadence",
),
(f"{AUTOSKILLS_COMMAND} time 03:00", "Set the local schedule time"),
(
f"{AUTOSKILLS_COMMAND} list [status]",
"List all proposals or filter by review, approved, or rejected",
),
]
for command, description in rows:
table.add_row(command, description)
ctx.ui.mount_renderable(table)
ctx.ui.append_system(
"Aliases: /skills-review, ls, proposals, accept, deny, enable, disable, now.",
style="dim",
)
@staticmethod
def _workspace_dir(ctx: CommandContext) -> str:
from ... import paths
return str(ctx.workspace_dir or paths.WORKSPACE_ROOT)
@staticmethod
def _first_arg(args: list[str]) -> str | None:
return args[0] if args else None
manager.register(AutoSkillsCommand())
@@ -1,10 +1,12 @@
from __future__ import annotations
from typing import ClassVar
from rich.panel import Panel
from rich.table import Table
from rich.text import Text
from ..base import Command, CommandContext
from ..base import Command, CommandContext, SubCommand
from ..manager import manager
@@ -13,6 +15,27 @@ class ChannelCommand(Command):
name = "/channel"
description = "Configure messaging channels"
category = "Channels"
subcommands: ClassVar[list[SubCommand]] = [
SubCommand("status", "Show channel status"),
SubCommand("stop", "Stop running channels"),
SubCommand("telegram", "Start Telegram channel"),
SubCommand("discord", "Start Discord channel"),
SubCommand("slack", "Start Slack channel"),
SubCommand("feishu", "Start Feishu channel"),
SubCommand("dingtalk", "Start DingTalk channel"),
SubCommand("wechat", "Start WeChat channel"),
SubCommand("email", "Start Email channel"),
SubCommand("imessage", "Start iMessage channel"),
]
def needs_agent(self, args: list[str]) -> bool:
# ``status`` and ``stop`` are introspection / teardown; they
# must work even when the agent load is still in flight or has
# failed. Only start/add flows feed ``ctx.agent`` into
# ``_start_channels_bus_mode``.
subcmd = args[0].lower() if args else ""
return subcmd not in {"status", "stop"}
async def execute(self, ctx: CommandContext, args: list[str]) -> None:
import EvoScientist.cli.channel as _ch_mod
@@ -63,10 +86,10 @@ class ChannelCommand(Command):
ctx.ui.append_system("No channels are running.", style="dim")
else:
if target:
_channels_stop(target)
_channels_stop(target, runtime=ctx.channel_runtime)
ctx.ui.append_system(f"Channel '{target}' stopped.", style="green")
else:
_channels_stop()
_channels_stop(runtime=ctx.channel_runtime)
ctx.ui.append_system("All channels stopped.", style="green")
return
@@ -90,6 +113,11 @@ class ChannelCommand(Command):
ctx.ui.append_system(f"Adding channel(s): {', '.join(requested)}...")
from ...cli.channel import _add_channel_to_running_bus
# Bind the runtime up-front so partial-success states (one
# channel attached, next one raises) still leave the bus
# observing the latest agent/thread refs.
if ctx.channel_runtime is not None:
ctx.channel_runtime.bind(ctx.agent, ctx.thread_id)
try:
for ct in requested:
_add_channel_to_running_bus(ct, config, send_thinking=send_thinking)
@@ -123,6 +151,8 @@ class ChannelCommand(Command):
ctx.thread_id,
send_thinking=send_thinking,
)
if ctx.channel_runtime is not None:
ctx.channel_runtime.bind(ctx.agent, ctx.thread_id)
# Show status panel
if _ch_mod._manager:
@@ -29,6 +29,9 @@ class HelpCommand(Command):
if cmd.alias:
desc += f" (aliases: {', '.join(cmd.alias)})"
help_text.append(f"{desc}\n", style="dim")
if cmd.subcommands:
names = ", ".join(sc.name for sc in cmd.subcommands)
help_text.append(f" subcommands: {names}\n", style="dim italic")
ctx.ui.mount_renderable(help_text)
@@ -49,17 +52,13 @@ class CurrentCommand(Command):
f"Workspace: {_shorten_path(ctx.workspace_dir)}",
style="dim",
)
memory_path = paths.MEMORY_DIR
memory_path = paths.MEMORIES_DIR
if memory_path:
from ...cli.agent import _shorten_path
ctx.ui.append_system(
f"Memory dir: {_shorten_path(str(memory_path))}", style="dim"
)
# How to determine UI type here?
# Maybe ctx.ui has a name or we pass it in ctx.
# For now, let's keep it simple.
ctx.ui.append_system("UI: auto", style="dim")
# Register commands
+51 -16
View File
@@ -1,8 +1,10 @@
from __future__ import annotations
from typing import ClassVar
from rich.table import Table
from ..base import Command, CommandContext
from ..base import Command, CommandContext, SubCommand
from ..manager import manager
@@ -11,8 +13,46 @@ class MCPCommand(Command):
name = "/mcp"
description = "Manage MCP servers"
category = "MCP"
subcommands: ClassVar[list[SubCommand]] = [
SubCommand("list", "List configured MCP servers"),
SubCommand("config", "Show server configuration details"),
SubCommand("add", "Add a new MCP server"),
SubCommand("edit", "Edit an MCP server configuration"),
SubCommand("remove", "Remove an MCP server"),
SubCommand("install", "Browse and install MCP servers"),
]
_server_names_cache: list[str] | None = None
def _get_server_names(self) -> list[str]:
if self._server_names_cache is None:
try:
from ...mcp import load_mcp_config
self._server_names_cache = list(load_mcp_config().keys())
except Exception:
return []
return self._server_names_cache
def _invalidate_server_cache(self) -> None:
self._server_names_cache = None
def get_completions(self, tokens: list[str]) -> list[tuple[str, str]]:
if len(tokens) <= 1:
return super().get_completions(tokens)
subcmd = tokens[0].lower()
if subcmd in ("config", "remove", "edit") and len(tokens) == 2:
prefix = tokens[1].lower()
return [
(name, "")
for name in self._get_server_names()
if name.lower().startswith(prefix)
]
return super().get_completions(tokens)
async def execute(self, ctx: CommandContext, args: list[str]) -> None:
"""Dispatch to the appropriate MCP subcommand."""
if not args or args[0] == "list":
await self._mcp_list(ctx)
return
@@ -24,35 +64,26 @@ class MCPCommand(Command):
await self._mcp_config(ctx, subargs[0] if subargs else "")
elif subcmd == "add":
await self._mcp_add(ctx, subargs)
self._invalidate_server_cache()
elif subcmd == "edit":
await self._mcp_edit(ctx, subargs)
self._invalidate_server_cache()
elif subcmd == "remove":
await self._mcp_remove(ctx, subargs[0] if subargs else "")
self._invalidate_server_cache()
elif subcmd == "install":
from .mcp_install import InstallMCPCommand
await InstallMCPCommand().execute(ctx, subargs)
else:
ctx.ui.append_system("MCP commands:", style="bold")
for sub in self.subcommands:
ctx.ui.append_system(
" /mcp List configured servers", style="dim"
)
ctx.ui.append_system(
" /mcp list List configured servers", style="dim"
)
ctx.ui.append_system(
" /mcp config Show detailed server config", style="dim"
)
ctx.ui.append_system(" /mcp add ... Add a server", style="dim")
ctx.ui.append_system(
" /mcp edit ... Edit an existing server", style="dim"
)
ctx.ui.append_system(" /mcp remove ... Remove a server", style="dim")
ctx.ui.append_system(
" /mcp install ... Browse and install servers", style="dim"
f" /mcp {sub.name:<12} {sub.description}", style="dim"
)
async def _mcp_list(self, ctx: CommandContext) -> None:
"""Display a table of all configured MCP servers."""
from ...mcp import load_mcp_config
from ...mcp.client import USER_MCP_CONFIG
@@ -85,6 +116,7 @@ class MCPCommand(Command):
ctx.ui.append_system(f"Config file: {USER_MCP_CONFIG}", style="dim")
async def _mcp_config(self, ctx: CommandContext, name: str) -> None:
"""Show detailed configuration for one or all MCP servers."""
from ...mcp import load_mcp_config
from ...mcp.client import USER_MCP_CONFIG
@@ -133,6 +165,7 @@ class MCPCommand(Command):
ctx.ui.append_system(f"Config file: {USER_MCP_CONFIG}", style="dim")
async def _mcp_add(self, ctx: CommandContext, tokens: list[str]) -> None:
"""Add a new MCP server from parsed arguments."""
from ...mcp import add_mcp_server, parse_mcp_add_args
if not tokens:
@@ -153,6 +186,7 @@ class MCPCommand(Command):
ctx.ui.append_system(f"Error: {exc}", style="red")
async def _mcp_edit(self, ctx: CommandContext, tokens: list[str]) -> None:
"""Edit fields of an existing MCP server configuration."""
from ...mcp import edit_mcp_server, parse_mcp_edit_args
if not tokens:
@@ -170,6 +204,7 @@ class MCPCommand(Command):
ctx.ui.append_system(f"Error: {exc}", style="red")
async def _mcp_remove(self, ctx: CommandContext, name: str) -> None:
"""Remove an MCP server by name."""
from ...mcp import remove_mcp_server
if not name:
@@ -10,6 +10,7 @@ class InstallMCPCommand(Command):
name = "/install-mcp"
description = "Browse and install MCP servers"
category = "MCP"
arguments: ClassVar[list[Argument]] = [
Argument(
name="source",
@@ -0,0 +1,234 @@
from __future__ import annotations
import asyncio
import re
from typing import ClassVar
from rich.table import Table
from ..base import Command, CommandContext, SubCommand
from ..manager import manager
class ScheduleCommand(Command):
"""Manage scheduled (cron) tasks."""
name = "/schedule"
description = "Manage scheduled (cron) tasks"
subcommands: ClassVar[list[SubCommand]] = [
SubCommand("add", 'Add: /schedule add <m h dom mon dow> "<prompt>"'),
SubCommand("list", "List scheduled tasks"),
SubCommand("remove", "Remove a schedule by id"),
SubCommand("run", "Run a schedule's prompt once now (test)"),
SubCommand("pause", "Disable a schedule by id"),
SubCommand("resume", "Enable a schedule by id"),
]
async def execute(self, ctx: CommandContext, args: list[str]) -> None:
"""Dispatch to the appropriate /schedule subcommand."""
cfg = getattr(ctx, "config", None)
if cfg is not None and not getattr(cfg, "enable_scheduler", True):
ctx.ui.append_system(
"Scheduled tasks are disabled (`enable_scheduler` is off).",
style="yellow",
)
return
from ...cron import schedule as crons
# Cron SDK calls are sync HTTP; offload to a thread so backend latency
# can never freeze the interactive event loop.
if not await asyncio.to_thread(crons.is_available):
ctx.ui.append_system(
"Scheduler unavailable: the langgraph dev backend is not running.",
style="yellow",
)
return
if not args or args[0].lower() == "list":
await self._list(ctx, crons)
return
sub = args[0].lower()
rest = args[1:]
if sub == "add":
await self._add(ctx, crons, rest)
elif sub == "remove":
await self._remove(ctx, crons, rest[0] if rest else "")
elif sub == "run":
await self._run(ctx, crons, rest[0] if rest else "")
elif sub in ("pause", "resume"):
await self._set_enabled(
ctx, crons, rest[0] if rest else "", sub == "resume"
)
else:
ctx.ui.append_system("Schedule commands:", style="bold")
for s in self.subcommands:
ctx.ui.append_system(
f" /schedule {s.name:<8} {s.description}", style="dim"
)
async def _add(self, ctx: CommandContext, crons, rest: list[str]) -> None:
# Cron may arrive as 5 separate tokens (unquoted) or 1 token (shlex-quoted).
# Split on any whitespace so extra spaces don't break detection; the
# backend rejects genuinely malformed expressions.
if rest and len(rest[0].split()) == 5: # quoted 5-field cron
schedule, prompt_tokens = " ".join(rest[0].split()), rest[1:]
elif len(rest) >= 5: # 5 separate cron fields
schedule, prompt_tokens = " ".join(rest[:5]), rest[5:]
else:
ctx.ui.append_system(
'Usage: /schedule add "<m h dom mon dow>" "<prompt>"', style="yellow"
)
return
prompt = " ".join(prompt_tokens).strip().strip('"').strip("'")
if not prompt:
ctx.ui.append_system("A task prompt is required.", style="yellow")
return
# B3: strip unsafe chars; keep only alphanumerics + hyphens (kebab-case).
raw = prompt[:48].lower()
name = re.sub(r"[^a-z0-9]+", "-", raw).strip("-")[:32] or "task"
try:
rec = await asyncio.to_thread(
crons.create_schedule, name=name, schedule=schedule, prompt=prompt
)
except Exception as exc:
ctx.ui.append_system(f"Error: {exc}", style="red")
return
ctx.ui.append_system(
f"Scheduled '{name}' [{schedule}] — id {rec.get('cron_id')}. "
"Runs unattended in the background.",
style="green",
)
async def _list(self, ctx: CommandContext, crons) -> None:
# B1: guard SDK call — backend may die after the is_available() check.
try:
rows = await asyncio.to_thread(crons.list_schedules)
except Exception as exc:
ctx.ui.append_system(f"Error: {exc}", style="red")
return
if not rows:
ctx.ui.append_system(
"No scheduled tasks. Add one: /schedule add ...", style="dim"
)
return
table = Table(title="Scheduled Tasks", show_header=True)
table.add_column("ID", style="cyan")
table.add_column("Name", style="magenta")
table.add_column("Schedule", style="green")
table.add_column("Enabled", style="yellow")
table.add_column("Next run (UTC)", style="white")
for r in rows:
meta = r.get("metadata") or {}
table.add_row(
str(r.get("cron_id", ""))[:8],
str(meta.get("name", "")),
str(r.get("schedule", "")),
"yes" if r.get("enabled", True) else "no",
str(r.get("next_run_date", "")),
)
ctx.ui.mount_renderable(table)
_AMBIGUOUS = object() # B2: sentinel returned when multiple crons match a prefix
_BACKEND_ERROR = object() # sentinel returned when list_schedules() raises
async def _resolve(self, crons, prefix: str):
"""Return the unique matching record, _AMBIGUOUS if >1 match, _BACKEND_ERROR on error, or None."""
# B1: guard SDK call — backend may die after is_available() check.
try:
all_rows = await asyncio.to_thread(crons.list_schedules)
except Exception as exc:
# Store the exception text so _resolve_or_report can surface it.
self._last_backend_exc = exc
return self._BACKEND_ERROR
# B2: collect ALL matches; ambiguous prefix → sentinel so callers can warn.
matches = [r for r in all_rows if str(r.get("cron_id", "")).startswith(prefix)]
if len(matches) > 1:
return self._AMBIGUOUS
return matches[0] if matches else None
async def _resolve_or_report(self, ctx: CommandContext, crons, prefix: str):
"""Resolve prefix → record, emit UI error on ambiguity/miss/error, return None on failure."""
match = await self._resolve(crons, prefix)
if match is self._BACKEND_ERROR:
exc = getattr(self, "_last_backend_exc", None)
ctx.ui.append_system(
f"Error: scheduler backend unavailable ({exc})",
style="red",
)
return None
if match is self._AMBIGUOUS:
ctx.ui.append_system(
f"Multiple schedules match '{prefix}' — use a longer id.",
style="yellow",
)
return None
if not match:
ctx.ui.append_system(f"No schedule matching {prefix}.", style="yellow")
return None
return match
async def _remove(self, ctx: CommandContext, crons, prefix: str) -> None:
if not prefix:
ctx.ui.append_system("Usage: /schedule remove <id>", style="yellow")
return
match = await self._resolve_or_report(ctx, crons, prefix)
if match is None:
return
cron_id = str(match.get("cron_id", ""))
try:
await asyncio.to_thread(crons.delete_schedule, cron_id)
except Exception as exc:
ctx.ui.append_system(f"Error: {exc}", style="red")
return
ctx.ui.append_system(f"Removed schedule {cron_id}.", style="green")
async def _run(self, ctx: CommandContext, crons, prefix: str) -> None:
if not prefix:
ctx.ui.append_system("Usage: /schedule run <id>", style="yellow")
return
match = await self._resolve_or_report(ctx, crons, prefix)
if match is None:
return
prompt = (match.get("metadata") or {}).get("prompt", "")
if not str(prompt).strip():
ctx.ui.append_system(
f"Schedule {prefix} has no stored prompt — cannot run it.",
style="yellow",
)
return
try:
rec = await asyncio.to_thread(crons.run_now, prompt)
except Exception as exc:
ctx.ui.append_system(f"Error: {exc}", style="red")
return
# Don't promise a location; the task's own prompt decides where output goes.
ctx.ui.append_system(
f"Fired schedule {prefix} once now (run {rec.get('run_id')}). "
"Any output goes wherever the task's instruction specifies.",
style="green",
)
async def _set_enabled(
self, ctx: CommandContext, crons, prefix: str, enabled: bool
) -> None:
if not prefix:
ctx.ui.append_system("Usage: /schedule pause|resume <id>", style="yellow")
return
match = await self._resolve_or_report(ctx, crons, prefix)
if match is None:
return
cron_id = str(match.get("cron_id", ""))
try:
await asyncio.to_thread(crons.set_enabled, cron_id, enabled)
except Exception as exc:
ctx.ui.append_system(f"Error: {exc}", style="red")
return
ctx.ui.append_system(
f"{'Resumed' if enabled else 'Paused'} schedule {cron_id}.", style="green"
)
# Register schedule command
manager.register(ScheduleCommand())
+45 -40
View File
@@ -5,15 +5,24 @@ from typing import ClassVar
from rich.table import Table
from ...gateway import GraphGateway, GraphTarget
from ..base import Argument, Command, CommandContext
from ..manager import manager
def _graph_gateway(ctx: CommandContext) -> GraphGateway:
if ctx.graph_gateway is None:
raise RuntimeError("Session commands require a graph_gateway")
return ctx.graph_gateway
class CompactCommand(Command):
"""Compact conversation to free context."""
name = "/compact"
description = "Compact conversation to free context"
requires_agent = True
category = "Session"
async def execute(self, ctx: CommandContext, args: list[str]) -> None:
from ...cli.commands import (
@@ -35,8 +44,12 @@ class CompactCommand(Command):
try:
result = await compact_conversation(
agent=ctx.agent,
graph_gateway=_graph_gateway(ctx),
thread_id=ctx.thread_id,
target=GraphTarget(
local_graph=ctx.agent,
workspace_dir=ctx.workspace_dir,
),
input_tokens_hint=ctx.input_tokens_hint,
)
finally:
@@ -70,11 +83,13 @@ class ThreadsCommand(Command):
name = "/threads"
description = "List recent sessions"
category = "Session"
async def execute(self, ctx: CommandContext, args: list[str]) -> None:
from ...sessions import _format_relative_time, list_threads
from ...sessions import _format_relative_time, short_thread_id
threads = await list_threads(
gateway = _graph_gateway(ctx)
threads = await gateway.list_threads(
limit=0,
include_message_count=True,
include_preview=True,
@@ -102,7 +117,7 @@ class ThreadsCommand(Command):
marker = " *" if thread_id_value == ctx.thread_id else ""
row = [
f"{thread_id_value}{marker}",
f"{short_thread_id(thread_id_value)}{marker}",
thread.get("preview", "") or "",
str(thread.get("message_count", 0)),
]
@@ -112,6 +127,11 @@ class ThreadsCommand(Command):
table.add_row(*row)
ctx.ui.mount_renderable(table)
if not is_channel:
ctx.ui.append_system(
" /resume to continue a session "
"/delete <id> to remove /new to start fresh",
)
class ResumeCommand(Command):
@@ -119,6 +139,7 @@ class ResumeCommand(Command):
name = "/resume"
description = "Resume a previous session"
category = "Session"
arguments: ClassVar[list[Argument]] = [
Argument(
name="thread_id",
@@ -129,14 +150,10 @@ class ResumeCommand(Command):
]
async def execute(self, ctx: CommandContext, args: list[str]) -> None:
from ...sessions import (
get_thread_metadata,
list_threads,
)
gateway = _graph_gateway(ctx)
arg = args[0] if args else ""
if not arg:
threads = await list_threads(
threads = await gateway.list_threads(
limit=0,
include_message_count=True,
include_preview=True,
@@ -160,7 +177,7 @@ class ResumeCommand(Command):
if not resolved:
return
metadata = await get_thread_metadata(resolved)
metadata = await gateway.get_thread_metadata(resolved)
restored_workspace = (metadata or {}).get("workspace_dir", "")
if restored_workspace:
ctx.workspace_dir = restored_workspace
@@ -172,21 +189,16 @@ class ResumeCommand(Command):
await ctx.ui.handle_session_resume(resolved, restored_workspace)
async def _resolve_thread_id(self, prefix: str, ctx: CommandContext) -> str | None:
from ...sessions import find_similar_threads, thread_exists
resolution = await _graph_gateway(ctx).resolve_thread(prefix)
if resolution.thread_id:
return resolution.thread_id
if await thread_exists(prefix):
return prefix
similar = await find_similar_threads(prefix)
if len(similar) == 1:
return similar[0]
if len(similar) > 1:
if resolution.matches:
ctx.ui.append_system(
f"Ambiguous thread ID '{prefix}'. Use a longer prefix.",
style="yellow",
)
for thread in similar:
for thread in resolution.matches:
ctx.ui.append_system(f" - {thread}", style="dim")
return None
@@ -199,9 +211,10 @@ class NewCommand(Command):
name = "/new"
description = "Start a new session"
category = "Session"
async def execute(self, ctx: CommandContext, args: list[str]) -> None:
ctx.ui.start_new_session()
await ctx.ui.start_new_session()
class ClearCommand(Command):
@@ -209,6 +222,7 @@ class ClearCommand(Command):
name = "/clear"
description = "Clear chat history"
category = "Session"
async def execute(self, ctx: CommandContext, args: list[str]) -> None:
ctx.ui.clear_chat()
@@ -219,6 +233,7 @@ class DeleteCommand(Command):
name = "/delete"
description = "Delete a saved session"
category = "Session"
arguments: ClassVar[list[Argument]] = [
Argument(
name="thread_id",
@@ -229,16 +244,10 @@ class DeleteCommand(Command):
]
async def execute(self, ctx: CommandContext, args: list[str]) -> None:
from ...sessions import (
delete_thread,
find_similar_threads,
list_threads,
thread_exists,
)
gateway = _graph_gateway(ctx)
arg = args[0] if args else ""
if not arg:
threads = await list_threads(
threads = await gateway.list_threads(
limit=0,
include_message_count=True,
include_preview=True,
@@ -258,22 +267,17 @@ class DeleteCommand(Command):
arg = selected
# Resolve thread_id
resolved = None
if await thread_exists(arg):
resolved = arg
else:
similar = await find_similar_threads(arg)
if len(similar) == 1:
resolved = similar[0]
elif len(similar) > 1:
resolution = await gateway.resolve_thread(arg)
if resolution.matches:
ctx.ui.append_system(
f"Ambiguous thread ID '{arg}'. Use a longer prefix.",
style="yellow",
)
for thread in similar:
for thread in resolution.matches:
ctx.ui.append_system(f" - {thread}", style="dim")
return
resolved = resolution.thread_id
if not resolved:
ctx.ui.append_system(f"Session '{arg}' not found.", style="red")
return
@@ -285,7 +289,7 @@ class DeleteCommand(Command):
)
return
deleted = await delete_thread(resolved)
deleted = await gateway.delete_thread(resolved)
if deleted:
ctx.ui.append_system(f"Deleted session {resolved}.", style="green")
else:
@@ -298,6 +302,7 @@ class ExitCommand(Command):
name = "/exit"
alias: ClassVar[list[str]] = ["/quit", "/q"]
description = "Quit EvoScientist"
category = "Session"
async def execute(self, ctx: CommandContext, args: list[str]) -> None:
ctx.ui.force_quit()
+12 -1
View File
@@ -13,6 +13,7 @@ class SkillsCommand(Command):
name = "/skills"
description = "List installed skills"
category = "Skills"
async def execute(self, ctx: CommandContext, args: list[str]) -> None:
from ...cli.agent import _shorten_path
@@ -65,6 +66,7 @@ class InstallSkill(Command):
name: ClassVar[str] = "/install-skill"
description: ClassVar[str] = "Add a skill from path or GitHub"
category: ClassVar[str] = "Skills"
arguments: ClassVar[list[Argument]] = [
Argument(
name="source",
@@ -142,6 +144,7 @@ class InstallSkills(Command):
description: ClassVar[str] = (
"Browse and install EvoSkills (optional: /evoskills <tag>)"
)
category: ClassVar[str] = "Skills"
arguments: ClassVar[list[Argument]] = [
Argument(
name="tag", type=str, description="Tag to filter skills by", required=False
@@ -213,11 +216,18 @@ class InstallSkills(Command):
pre_filter_tag=tag,
)
if not selected_sources:
# ``None`` means user cancelled (Esc / Ctrl-C). An empty list means
# the picker handled a "nothing to do" state (all-installed / no
# tag matches) and already printed its own specific message; the
# outer layer should stay silent rather than claim a cancel.
if selected_sources is None:
if not is_channel:
ctx.ui.append_system("Browse cancelled.", style="dim")
return
if not selected_sources:
return
# Install selected skills
installed_count = 0
for source in selected_sources:
@@ -248,6 +258,7 @@ class UninstallSkill(Command):
name: ClassVar[str] = "/uninstall-skill"
description: ClassVar[str] = "Remove an installed skill"
category: ClassVar[str] = "Skills"
arguments: ClassVar[list[Argument]] = [
Argument(
name="name",
+35 -1
View File
@@ -3,7 +3,7 @@ from __future__ import annotations
import logging
import shlex
from .base import Command, CommandContext
from .base import Command, CommandContext, SubCommand
_logger = logging.getLogger(__name__)
@@ -27,6 +27,27 @@ class CommandManager:
"""Lookup a command by name."""
return self._commands.get(name.lower())
def resolve(self, command_str: str) -> tuple[Command, list[str]] | None:
"""Return ``(command, args)`` for the dispatch of ``command_str``.
Uses the same parsing as :meth:`execute` so callers can inspect
metadata (e.g. call :meth:`Command.needs_agent`) without
re-implementing ``shlex`` quirks.
"""
command_str = command_str.strip()
if not command_str:
return None
try:
parts = shlex.split(command_str)
except ValueError:
parts = command_str.split()
if not parts:
return None
cmd = self.get_command(parts[0])
if cmd is None:
return None
return cmd, parts[1:]
def list_commands(self) -> list[tuple[str, str]]:
"""List all registered command names and descriptions."""
seen = set()
@@ -37,6 +58,17 @@ class CommandManager:
seen.add(cmd)
return results
def get_subcommands(self, command_name: str) -> list[SubCommand]:
"""Return subcommands declared by *command_name*, or empty list."""
cmd = self.get_command(command_name)
if cmd is None:
return []
return cmd.subcommands
def list_subcommands(self, command_name: str) -> list[tuple[str, str]]:
"""Return ``(name, description)`` pairs for completion rendering."""
return [(sc.name, sc.description) for sc in self.get_subcommands(command_name)]
def get_all_commands(self) -> list[Command]:
"""Return all registered command instances."""
seen = set()
@@ -72,12 +104,14 @@ class CommandManager:
if not cmd:
return False
ctx.command_error = None
try:
await cmd.execute(ctx, args)
await ctx.ui.flush()
return True
except Exception as e:
_logger.exception(f"Error executing command {cmd_name}: {e}")
ctx.command_error = str(e)
ctx.ui.append_system(f"Error executing {cmd_name}: {e}", style="red")
await ctx.ui.flush()
return True
+14
View File
@@ -10,11 +10,18 @@ The onboard module is loaded lazily because it pulls in heavy dependencies
from .settings import (
EvoScientistConfig,
MemoryControls,
MemoryObservationTarget,
MemoryObservationWriter,
MemorySkillSynthesisCadence,
MemorySkillSynthesisMode,
apply_config_to_env,
get_config_dir,
get_config_path,
get_config_value,
get_default_workspace_dir,
get_effective_config,
is_config_applied_env,
list_config,
load_config,
reset_config,
@@ -24,12 +31,19 @@ from .settings import (
__all__ = [
"EvoScientistConfig",
"MemoryControls",
"MemoryObservationTarget",
"MemoryObservationWriter",
"MemorySkillSynthesisCadence",
"MemorySkillSynthesisMode",
"apply_config_to_env",
# settings
"get_config_dir",
"get_config_path",
"get_config_value",
"get_default_workspace_dir",
"get_effective_config",
"is_config_applied_env",
"list_config",
"load_config",
"reset_config",
-316
View File
@@ -1,316 +0,0 @@
"""Configuration helpers for dedicated image-generation models.
Image generation models are service/tool models, not chat models. Keeping
them in a separate config section prevents image-only models such as
``gpt-image-2`` from being offered in the normal chat model selector.
"""
from __future__ import annotations
import os
import tempfile
from typing import Any
import yaml
from pydantic import BaseModel, Field, field_validator
IMAGE_GENERATION_SECTION = "image_generation"
DEFAULT_IMAGE_GENERATION_TIMEOUT_SECONDS = 120.0
IMAGE_GENERATION_USAGE_NOTES = [
"Normal chat model lists filter out image-only models such as gpt-image-* and dall-e-*.",
"If an image-only model is accidentally saved in LLM settings, it is moved into image_generation.",
"If the frontend sends an image-only model as the chat model, the backend returns IMAGE_MODEL_NOT_CHAT_MODEL.",
]
_IMAGE_MODEL_PREFIXES = (
"gpt-image",
"chatgpt-image",
"dall-e",
"dalle",
)
_IMAGE_MODEL_MARKERS = (
"wanx",
"seedream",
)
class ImageModelEntry(BaseModel):
"""A single dedicated image-generation model."""
id: str = Field(..., description="Model ID sent to the image API")
name: str = Field("", description="Display name or short alias")
provider: str = Field("openai-compatible", description="Provider label")
api_key: str = Field("", description="API key or ${ENV_VAR} reference")
base_url: str = Field("", description="API base URL or ${ENV_VAR} reference")
supports_generation: bool = True
supports_edit: bool = True
default_size: str = "1024x1024"
default_quality: str = "auto"
params: dict[str, Any] = Field(default_factory=dict)
@field_validator("id")
@classmethod
def id_not_empty(cls, value: str) -> str:
value = value.strip()
if not value:
raise ValueError("image model id must not be empty")
return value
def resolved_api_key(self) -> str:
return _resolve_env_ref(self.api_key)
def resolved_base_url(self) -> str:
return _resolve_env_ref(self.base_url)
def display_name(self) -> str:
return self.name or self.id
class ImageGenerationSettings(BaseModel):
"""Dedicated image-generation model settings."""
default_model: str = ""
timeout_seconds: float = Field(
DEFAULT_IMAGE_GENERATION_TIMEOUT_SECONDS,
description="Per-request timeout for image provider HTTP calls",
)
models: list[ImageModelEntry] = Field(default_factory=list)
@field_validator("timeout_seconds")
@classmethod
def timeout_must_be_positive(cls, value: float) -> float:
if value <= 0:
raise ValueError("image generation timeout_seconds must be greater than 0")
return value
def is_image_generation_model(model_ref: str | None) -> bool:
"""Return True when a model ID is known to be image-generation only."""
if not model_ref:
return False
value = str(model_ref).strip().lower()
if "/" in value:
value = value.rsplit("/", 1)[1]
return value.startswith(_IMAGE_MODEL_PREFIXES) or any(
marker in value for marker in _IMAGE_MODEL_MARKERS
)
def load_image_generation_settings(
*,
include_legacy: bool = True,
resolve_env: bool = True,
) -> ImageGenerationSettings:
"""Load dedicated image-generation settings with legacy fallback."""
raw = _load_settings_yaml()
legacy_present = _legacy_config_present(raw)
section = raw.get(IMAGE_GENERATION_SECTION)
if isinstance(section, dict):
source = _resolve_nested_env(section) if resolve_env else section
settings = ImageGenerationSettings.model_validate(source)
else:
settings = ImageGenerationSettings()
legacy = _legacy_model_entry(raw)
if not settings.models:
if include_legacy or legacy_present:
settings.models = [legacy]
elif include_legacy and _legacy_env_overrides_present():
_merge_entry(settings, legacy)
if not settings.default_model:
if settings.models:
settings.default_model = settings.models[0].id
elif include_legacy or legacy_present:
settings.default_model = legacy.id
return settings
def save_image_generation_settings(settings: ImageGenerationSettings) -> ImageGenerationSettings:
"""Save image-generation settings into ``settings.yaml``."""
validated = ImageGenerationSettings.model_validate(settings.model_dump(mode="python"))
raw = _load_settings_yaml()
raw[IMAGE_GENERATION_SECTION] = validated.model_dump(mode="python")
_atomic_write_settings_yaml(raw)
return validated
def merge_image_generation_models(
entries: list[ImageModelEntry],
*,
default_model: str | None = None,
) -> ImageGenerationSettings:
"""Merge image model entries into the dedicated image model list."""
raw = _load_settings_yaml()
settings = load_image_generation_settings(include_legacy=_legacy_config_present(raw))
for entry in entries:
_merge_entry(settings, entry)
if default_model:
settings.default_model = default_model
elif entries and not settings.default_model:
settings.default_model = entries[0].id
return save_image_generation_settings(settings)
def resolve_image_generation_model(model_ref: str | None = None) -> ImageModelEntry:
"""Resolve a model ID/name to a configured image model entry."""
settings = load_image_generation_settings()
requested = (model_ref or settings.default_model or "").strip()
if not requested and settings.models:
requested = settings.models[0].id
for entry in settings.models:
if requested in {entry.id, entry.name}:
return _with_legacy_fallbacks(entry)
if requested:
legacy = _legacy_model_entry(_load_settings_yaml())
if requested == legacy.id:
return legacy
raise ValueError(f"Unknown image generation model: {requested}")
raise ValueError("No image generation model configured")
def list_image_generation_models(*, include_sensitive: bool = False) -> dict[str, Any]:
"""Return image-generation model settings for API/tool display."""
settings = load_image_generation_settings()
models = []
for entry in settings.models:
item = entry.model_dump(mode="python")
item["name"] = entry.display_name()
if include_sensitive:
item["api_key"] = entry.resolved_api_key()
item["base_url"] = entry.resolved_base_url()
else:
item["api_key"] = _mask_secret(entry.resolved_api_key())
item["base_url"] = entry.resolved_base_url()
models.append(item)
return {
"default_model": settings.default_model,
"timeout_seconds": settings.timeout_seconds,
"models": models,
"usage_notes": IMAGE_GENERATION_USAGE_NOTES,
}
def _merge_entry(settings: ImageGenerationSettings, entry: ImageModelEntry) -> None:
for idx, existing in enumerate(settings.models):
if existing.id == entry.id:
data = existing.model_dump(mode="python")
update = entry.model_dump(mode="python")
for key, value in update.items():
if value not in ("", None, {}, []):
data[key] = value
settings.models[idx] = ImageModelEntry.model_validate(data)
return
settings.models.append(entry)
def _with_legacy_fallbacks(entry: ImageModelEntry) -> ImageModelEntry:
legacy = _legacy_model_entry(_load_settings_yaml())
data = entry.model_dump(mode="python")
if not data.get("api_key") or _is_masked_secret(str(data.get("api_key") or "")):
data["api_key"] = legacy.api_key
if not data.get("base_url"):
data["base_url"] = legacy.base_url
return ImageModelEntry.model_validate(data)
def _legacy_env_overrides_present() -> bool:
return any(os.environ.get(key) for key in ("IMAGE_GEN_MODEL", "IMAGE_GEN_API_KEY", "IMAGE_GEN_BASE_URL"))
def _legacy_config_present(raw: dict[str, Any]) -> bool:
return _legacy_env_overrides_present() or any(
str(raw.get(key) or "").strip()
for key in ("image_gen_model", "image_gen_api_key", "image_gen_base_url")
)
def _legacy_model_entry(raw: dict[str, Any]) -> ImageModelEntry:
model = (
os.environ.get("IMAGE_GEN_MODEL")
or str(raw.get("image_gen_model") or "").strip()
or "dall-e-3"
)
api_key = (
os.environ.get("IMAGE_GEN_API_KEY")
or str(raw.get("image_gen_api_key") or "").strip()
or os.environ.get("OPENAI_API_KEY", "")
)
base_url = (
os.environ.get("IMAGE_GEN_BASE_URL")
or str(raw.get("image_gen_base_url") or "").strip()
or os.environ.get("OPENAI_BASE_URL", "")
)
return ImageModelEntry(
id=model,
name=model,
api_key=api_key,
base_url=base_url,
provider="openai-compatible",
)
def _load_settings_yaml() -> dict[str, Any]:
from EvoScientist.config.settings import get_config_path
path = get_config_path()
if not path.exists():
return {}
try:
with path.open(encoding="utf-8") as fh:
data = yaml.safe_load(fh) or {}
except Exception:
return {}
return data if isinstance(data, dict) else {}
def _atomic_write_settings_yaml(data: dict[str, Any]) -> None:
from EvoScientist.config.settings import get_config_path
path = get_config_path()
path.parent.mkdir(parents=True, exist_ok=True)
fd, tmp_path = tempfile.mkstemp(
dir=str(path.parent),
prefix=".settings_",
suffix=".yaml.tmp",
text=True,
)
try:
with os.fdopen(fd, "w", encoding="utf-8") as fh:
yaml.safe_dump(data, fh, default_flow_style=False, sort_keys=False)
fh.flush()
os.fsync(fh.fileno())
os.replace(tmp_path, path)
finally:
if os.path.exists(tmp_path):
os.unlink(tmp_path)
def _resolve_nested_env(value: Any) -> Any:
if isinstance(value, dict):
return {k: _resolve_nested_env(v) for k, v in value.items()}
if isinstance(value, list):
return [_resolve_nested_env(v) for v in value]
if isinstance(value, str):
return _resolve_env_ref(value)
return value
def _resolve_env_ref(value: str) -> str:
if isinstance(value, str) and value.startswith("${") and value.endswith("}"):
return os.environ.get(value[2:-1], "")
return value or ""
def _mask_secret(value: str) -> str:
if not value:
return ""
if len(value) <= 8:
return "***"
return f"{value[:3]}...{value[-4:]}"
def _is_masked_secret(value: str) -> bool:
return bool(value and (value == "***" or value == "********" or "..." in value))
+132
View File
@@ -0,0 +1,132 @@
"""Startup detection of pre-Registry legacy model configuration artifacts.
Design doc section 10 step 4: this project is in development and keeps no
historical compatibility. When the config service or the CLI finds legacy
artifacts — an old ``providers.yaml``, an old
``run-runtime-snapshots.sqlite3``, or LLM fields left behind in
``config.yaml`` — it must refuse to start and log an explicit reset guide
instead of partially reading them.
"""
from __future__ import annotations
import logging
from pathlib import Path
import yaml
from .settings import get_config_dir
logger = logging.getLogger(__name__)
#: LLM configuration keys that ``config.yaml`` must no longer carry
#: (section 10 step 2). Platform configuration (workspace, MCP, ports,
#: storage, scheduling, security) stays; everything model/provider/credential
#: related lives in the Model Registry (``model-runtime.sqlite3``) only.
LEGACY_CONFIG_YAML_KEYS = frozenset(
{
"provider",
"model",
"model_catalog",
"model_fallbacks",
"auxiliary_provider",
"auxiliary_model",
"anthropic_api_key",
"anthropic_base_url",
"anthropic_auth_mode",
"openai_api_key",
"openai_auth_mode",
"nvidia_api_key",
"google_api_key",
"minimax_api_key",
"minimax_base_url",
"siliconflow_api_key",
"openrouter_api_key",
"deepseek_api_key",
"zhipu_api_key",
"volcengine_api_key",
"dashscope_api_key",
"moonshot_api_key",
"kimi_api_key",
"custom_openai_api_key",
"custom_openai_base_url",
"custom_anthropic_api_key",
"custom_anthropic_base_url",
"ollama_base_url",
"use_responses_api",
"openrouter_anthropic_prompt_cache",
}
)
_LEGACY_PROVIDERS_FILE = "providers.yaml"
_LEGACY_SNAPSHOTS_DB = "run-runtime-snapshots.sqlite3"
_RESET_GUIDANCE = """\
EvoScientist no longer reads legacy model configuration (design doc §10).
To reset the development environment:
1. Delete {config_dir}/providers.yaml (Provider Profiles are superseded
by the Model Registry).
2. Delete {config_dir}/run-runtime-snapshots.sqlite3 (the old snapshot
store; run snapshots now live in model-runtime.sqlite3).
3. Remove the leftover LLM fields listed above from {config_path} —
platform fields (workspace, MCP, ports, scheduling, security) stay.
4. Configure providers/models through the Model Registry (WebUI
configuration page or the model-registry API), run the provider test,
then enable the models you need.
Startup is refused — no legacy artifact is read, even partially.\
"""
class LegacyArtifactsError(RuntimeError):
"""Raised at startup when pre-Registry configuration artifacts remain."""
def find_legacy_artifacts(config_dir: Path | None = None) -> list[str]:
"""Return human-readable descriptions of every legacy artifact found."""
config_dir = config_dir if config_dir is not None else get_config_dir()
found: list[str] = []
providers_yaml = config_dir / _LEGACY_PROVIDERS_FILE
if providers_yaml.exists():
found.append(f"legacy Provider Profiles file: {providers_yaml}")
snapshots_db = config_dir / _LEGACY_SNAPSHOTS_DB
if snapshots_db.exists():
found.append(f"legacy run snapshot database: {snapshots_db}")
config_path = config_dir / "config.yaml"
if config_path.exists():
try:
with open(config_path, encoding="utf-8") as handle:
data = yaml.safe_load(handle) or {}
except yaml.YAMLError:
data = {}
if isinstance(data, dict):
leftover = sorted(LEGACY_CONFIG_YAML_KEYS & data.keys())
if leftover:
found.append(
f"leftover LLM fields in {config_path}: {', '.join(leftover)}"
)
return found
def assert_no_legacy_artifacts(config_dir: Path | None = None) -> None:
"""Refuse startup when any legacy model configuration artifact remains.
Logs the findings plus the explicit reset guide (section 10 step 4) and
raises :class:`LegacyArtifactsError`. Nothing is read partially: the
caller must not catch-and-continue.
"""
found = find_legacy_artifacts(config_dir)
if not found:
return
config_dir = config_dir if config_dir is not None else get_config_dir()
guidance = _RESET_GUIDANCE.format(
config_dir=config_dir,
config_path=config_dir / "config.yaml",
)
message = "Legacy model configuration artifacts detected:\n" + "\n".join(
f" - {item}" for item in found
)
logger.error("%s\n%s", message, guidance)
raise LegacyArtifactsError(f"{message}\n\n{guidance}")
-947
View File
@@ -1,947 +0,0 @@
"""Structured LLM provider/model configuration.
This module defines the YAML-driven configuration for providers, models,
and their parameters. It supports:
- Multiple providers with api_key, base_url, and protocol
- Multiple models per provider with alias, capabilities, and params
- Three-level parameter inheritance: defaults ← model ← runtime
- Environment variable references in api_key fields
"""
from __future__ import annotations
import logging
import os
import random
import re
import tempfile
import threading
import time
from pathlib import Path
from typing import Any, Literal
logger = logging.getLogger(__name__)
import yaml
from pydantic import BaseModel, Field, field_validator
# ─── Environment variable reference pattern ──────────────────────────────
_ENV_REF_RE = re.compile(r"^\$\{(\w+)\}$")
_FLAT_LLM_CONFIG_KEYS = {
"provider_routes",
"provider",
"model",
"reasoning_effort",
"anthropic_api_key",
"anthropic_base_url",
"openai_api_key",
"nvidia_api_key",
"google_api_key",
"minimax_api_key",
"siliconflow_api_key",
"openrouter_api_key",
"deepseek_api_key",
"zhipu_api_key",
"volcengine_api_key",
"dashscope_api_key",
"moonshot_api_key",
"kimi_api_key",
"custom_openai_api_key",
"custom_openai_base_url",
"custom_anthropic_api_key",
"custom_anthropic_base_url",
"ollama_base_url",
}
def _resolve_env_ref(value: str) -> str:
"""Resolve ${ENV_VAR} references in string values.
If the value matches the pattern ${ENV_VAR}, return the environment
variable's value. Otherwise return the string as-is.
"""
m = _ENV_REF_RE.match(value.strip())
if m:
return os.environ.get(m.group(1), "")
return value
# ─── Weighted round-robin load balancer ───────────────────────────────────
# Per-provider counter for round-robin position
_round_robin_counters: dict[str, int] = {}
# ─── Endpoint call statistics ──────────────────────────────────────────
class EndpointStats:
"""Thread-safe in-memory endpoint call and token usage statistics.
Tracks per-endpoint:
- Call counts (from resolve_model)
- Token usage (input_tokens, output_tokens)
- Last call timestamp
- Per-model call distribution
Stats are kept in memory. Call ``snapshot()`` for a point-in-time
copy, ``reset()`` to clear counters, or ``summary()`` for a
human-readable report.
"""
def __init__(self) -> None:
self._lock = threading.Lock()
# key = (provider_name, endpoint_name)
# value = dict with keys: calls, last_call_ts, models, input_tokens, output_tokens
self._stats: dict[tuple[str, str], dict[str, Any]] = {}
# Track last endpoint selected per model for token attribution
# key = model_id, value = (provider, endpoint_name)
self._last_endpoint: dict[str, tuple[str, str]] = {}
# Track most recently used endpoint (for fallback when model_id is empty)
self._last_recorded: tuple[str, str] | None = None
# Stack of pending endpoint attributions for sequential matching.
# Each call to record() pushes, each record_tokens_for_model() pops.
# This ensures tokens are attributed to the correct endpoint in tool-call loops
# where resolve_model() is called multiple times before usage_metadata arrives.
self._attribution_stack: list[tuple[str, str]] = []
def _ensure_entry(self, key: tuple[str, str]) -> dict[str, Any]:
"""Get or create a stats entry for the given key."""
return self._stats.setdefault(key, {
"calls": 0,
"last_call_ts": 0.0,
"models": {},
"input_tokens": 0,
"output_tokens": 0,
})
def record(self, provider: str, endpoint: str, model_id: str) -> None:
"""Record one endpoint selection (called from resolve_model)."""
with self._lock:
entry = self._ensure_entry((provider, endpoint))
entry["calls"] += 1
entry["last_call_ts"] = time.time()
entry["models"][model_id] = entry["models"].get(model_id, 0) + 1
# Remember last endpoint for this model (for token attribution)
self._last_endpoint[model_id] = (provider, endpoint)
# Track most recently used endpoint overall
self._last_recorded = (provider, endpoint)
# Push to attribution stack for sequential token matching
self._attribution_stack.append((provider, endpoint))
def record_tokens(
self,
provider: str,
endpoint: str,
input_tokens: int = 0,
output_tokens: int = 0,
) -> None:
"""Record token usage for an endpoint.
Can be called after an API response is received to accumulate
token counts. The endpoint entry is created if it doesn't exist
(e.g. for single-endpoint providers that aren't tracked by
``record()``).
"""
if not input_tokens and not output_tokens:
return
with self._lock:
entry = self._ensure_entry((provider, endpoint))
entry["input_tokens"] += input_tokens
entry["output_tokens"] += output_tokens
def record_tokens_for_model(
self,
model_id: str,
input_tokens: int = 0,
output_tokens: int = 0,
) -> None:
"""Record token usage, auto-routing to the correct endpoint.
Resolution order:
1. Attribution stack — pops the oldest pending endpoint (FIFO matching
with record() calls, handles tool-call loops correctly).
2. ``_last_endpoint[model_id]`` — direct model→endpoint mapping.
3. Scan ``_stats`` for any endpoint that has this model registered.
4. ``_last_recorded`` — fallback for empty model_id.
5. ``("unknown", "unknown")`` — last resort (debug level).
"""
if not input_tokens and not output_tokens:
return
with self._lock:
provider, endpoint = None, None
# Strategy 1: Pop from attribution stack (matches record() calls in order)
if self._attribution_stack:
provider, endpoint = self._attribution_stack.pop(0)
# Strategy 2: Direct lookup by model_id
if provider is None and model_id:
p, e = self._last_endpoint.get(model_id, (None, None))
if p is not None:
provider, endpoint = p, e
# Strategy 3: Scan _stats for this model
if provider is None and model_id:
for (p, e), data in self._stats.items():
if model_id in data.get("models", {}):
provider, endpoint = p, e
self._last_endpoint[model_id] = (p, e)
break
# Strategy 4: Use most recently recorded endpoint
if provider is None and self._last_recorded:
provider, endpoint = self._last_recorded
# Strategy 5: Unknown — no resolve_model() was called beforehand
if provider is None:
# When model_id is empty and all state is empty, this is a
# known race condition (TOCTOU between events.py guard and
# this method's lock). Silently skip — no useful attribution
# is possible and logging it just creates noise.
if not model_id and not self._last_recorded:
return
provider, endpoint = "unknown", "unknown"
# Common when LLM calls bypass resolve_model() (e.g. LangChain
# internal bindings, sub-agents). Debug level to avoid log spam.
logger.debug(
"EndpointStats: model %r not found, stack=%d _last_endpoint=%s "
"_last_recorded=%s",
model_id,
len(self._attribution_stack),
list(self._last_endpoint.keys()),
self._last_recorded,
)
entry = self._ensure_entry((provider, endpoint))
entry["input_tokens"] += input_tokens
entry["output_tokens"] += output_tokens
def snapshot(self) -> dict[str, Any]:
"""Return a deep copy of current stats."""
with self._lock:
import copy
return copy.deepcopy(self._stats)
async def restore_today_from_db(self) -> int:
"""Restore today's endpoint stats from endpoint_usage_daily.
Called once at gateway startup so that the in-memory EndpointStats
reflects today's accumulated usage after a process restart.
Returns the number of rows restored.
"""
try:
from EvoScientist.runtime_integrations import current_date, get_app_connection
db = await get_app_connection()
today = current_date()
rows = await db.execute_fetchall(
"""
SELECT provider, endpoint, model, calls,
input_tokens, output_tokens
FROM endpoint_usage_daily
WHERE date = $1
""",
(today,),
)
with self._lock:
for r in rows:
key = (r["provider"], r["endpoint"])
entry = self._ensure_entry(key)
entry["calls"] += r["calls"]
entry["input_tokens"] += r["input_tokens"]
entry["output_tokens"] += r["output_tokens"]
if r["model"]:
entry["models"][r["model"]] = (
entry["models"].get(r["model"], 0) + r["calls"]
)
self._last_endpoint[r["model"]] = key
restored = len(rows)
if restored:
logger.info(
"EndpointStats: restored %d endpoint entries from DB for today",
restored,
)
return restored
except Exception:
logger.warning("EndpointStats: failed to restore from DB", exc_info=True)
return 0
def reset(self) -> None:
"""Clear all counters."""
with self._lock:
self._stats.clear()
@staticmethod
def _fmt_tokens(n: int) -> str:
"""Format token count compactly."""
if n >= 1_000_000:
return f"{n / 1_000_000:.1f}M"
if n >= 1_000:
return f"{n / 1_000:.1f}K"
return str(n)
def summary(self) -> str:
"""Human-readable summary for CLI / logging."""
snap = self.snapshot()
if not snap:
return "No endpoint calls recorded."
total_calls = sum(d["calls"] for d in snap.values())
total_input = sum(d["input_tokens"] for d in snap.values())
total_output = sum(d["output_tokens"] for d in snap.values())
if total_calls == 0:
return "No endpoint calls recorded."
lines: list[str] = []
for (provider, endpoint), data in sorted(snap.items()):
calls = data["calls"]
pct = calls / total_calls * 100
elapsed = time.time() - data["last_call_ts"] if data["last_call_ts"] else 0
if elapsed < 60:
ago = f"{elapsed:.0f}s ago"
elif elapsed < 3600:
ago = f"{elapsed / 60:.0f}m ago"
else:
ago = f"{elapsed / 3600:.1f}h ago"
bar = "█" * int(pct / 5) + "░" * (20 - int(pct / 5))
inp = self._fmt_tokens(data["input_tokens"])
out = self._fmt_tokens(data["output_tokens"])
models = ", ".join(f"{m}({c})" for m, c in sorted(data["models"].items()))
token_str = ""
if data["input_tokens"] or data["output_tokens"]:
token_str = f" in={inp} out={out}"
lines.append(
f" {provider}/{endpoint} {bar} {calls} ({pct:.0f}%) last {ago}{token_str}\n"
f" models: {models}"
)
header = f"Endpoint usage — {total_calls} calls"
if total_input or total_output:
header = (
f"Endpoint usage — {total_calls} calls, "
f"{self._fmt_tokens(total_input)} in / {self._fmt_tokens(total_output)} out tokens"
)
return header + ":\n" + "\n".join(lines)
# Global singleton
_endpoint_stats = EndpointStats()
def get_endpoint_stats() -> EndpointStats:
"""Return the global endpoint call statistics instance."""
return _endpoint_stats
def _select_endpoint(
endpoints: list,
provider_name: str,
preferred_name: str = "",
) -> tuple:
"""Select an endpoint using weighted round-robin or explicit pinning.
Args:
endpoints: List of EndpointConfig objects with valid credentials.
provider_name: Provider name for round-robin counter.
preferred_name: If set, try to find an endpoint with this name first.
Returns:
The selected EndpointConfig.
"""
if not endpoints:
raise ValueError(f"No available endpoints for provider '{provider_name}'")
# 1. If model pins to a specific endpoint, find it
if preferred_name:
for ep in endpoints:
if ep.name == preferred_name and ep.resolved_api_key():
return ep
# 2. Filter to endpoints with valid credentials
available = [ep for ep in endpoints if ep.resolved_api_key()]
if not available:
raise ValueError(f"No endpoints with valid credentials for provider '{provider_name}'")
if len(available) == 1:
return available[0]
# 3. Weighted round-robin selection
total_weight = sum(ep.weight for ep in available)
if total_weight <= 0:
return available[0]
key = provider_name
pos = _round_robin_counters.get(key, 0)
_round_robin_counters[key] = (pos + 1) % total_weight
# Walk through endpoints by cumulative weight
cumulative = 0
for ep in available:
cumulative += ep.weight
if pos < cumulative:
return ep
return available[-1]
# ═══════════════════════════════════════════════════════════════════════════
# Pydantic models for the new structured config
# ═══════════════════════════════════════════════════════════════════════════
class EndpointConfig(BaseModel):
"""A single API endpoint within a provider — its own api_key, base_url, and optional weights."""
name: str = Field("", description="Endpoint name for reference (e.g. 'coding', 'general')")
api_key: str = Field("", description="API key (literal or ${ENV_VAR} reference)")
base_url: str = Field("", description="API base URL override (empty = provider default)")
weight: int = Field(1, description="Load-balancing weight (higher = more traffic)")
extra_body: dict[str, Any] = Field(
default_factory=dict,
description="Extra JSON body fields for requests through this endpoint",
)
default_headers: dict[str, str] = Field(
default_factory=dict,
description="Custom HTTP headers for requests through this endpoint",
)
params: dict[str, Any] = Field(
default_factory=dict,
description="Endpoint-level model/client parameter overrides",
)
def resolved_api_key(self) -> str:
"""Return the API key with environment variable references resolved."""
return _resolve_env_ref(self.api_key) if self.api_key else ""
def resolved_base_url(self) -> str:
"""Return the base URL with environment variable references resolved."""
return _resolve_env_ref(self.base_url) if self.base_url else ""
class AccessConfig(BaseModel):
"""Plan/role access constraints for providers and models.
Empty lists mean unrestricted. Gateway treats the admin role as allowed.
"""
allowed_plans: list[str] = Field(default_factory=list, description="Allowed subscription plans")
allowed_roles: list[str] = Field(default_factory=list, description="Allowed user roles")
class ModelEntry(BaseModel):
"""A single model definition within a provider."""
id: str = Field(..., description="Full model ID sent to the API")
alias: str = Field("", description="Short alias for easy reference")
tier: str = Field("", description="Billing tier controlled by the server config")
currency: str = Field("CNY", description="Settlement currency for model pricing")
max_tokens: int = Field(4096, description="Maximum output tokens")
supports_vision: bool = Field(False, description="Whether the model supports image inputs")
supports_reasoning: bool = Field(False, description="Whether the model supports reasoning/thinking")
endpoint: str = Field("", description="Preferred endpoint name (empty = auto-select via load balancing)")
params: dict[str, Any] = Field(
default_factory=dict,
description="Model-level parameter overrides (temperature, thinking, reasoning, etc.)",
)
pricing: dict[str, float | str | None] = Field(
default_factory=dict,
description="Optional billing price: input_per_million, output_per_million, cached_input_per_million.",
)
permission_mode: Literal["inherit", "custom"] = Field(
"inherit",
description="Model access mode: inherit provider permissions or use model-level access.",
)
access: AccessConfig = Field(
default_factory=AccessConfig,
description="Model-level access constraints used when permission_mode is custom.",
)
@field_validator("id")
@classmethod
def id_not_empty(cls, v: str) -> str:
if not v.strip():
raise ValueError("model id must not be empty")
return v.strip()
class ProviderConfig(BaseModel):
"""Configuration for a single LLM provider.
Supports two modes:
1. Single endpoint (legacy): set api_key / base_url directly at provider level.
2. Multi-endpoint (new): define an 'endpoints' list, each with its own
api_key, base_url, weight. Models can pin to a specific endpoint or
let the system auto-select via weighted round-robin for load balancing.
"""
api_key: str = Field("", description="API key (literal or ${ENV_VAR} reference)")
base_url: str = Field("", description="API base URL override (empty = provider default)")
protocol: Literal["openai", "anthropic", "google-genai", "ollama", "openrouter"] = Field(
"openai",
description="LLM protocol to use for this provider",
)
extra_body: dict[str, Any] = Field(
default_factory=dict,
description="Extra JSON body fields for all requests to this provider",
)
default_headers: dict[str, str] = Field(
default_factory=dict,
description="Custom HTTP headers for all requests to this provider",
)
params: dict[str, Any] = Field(
default_factory=dict,
description="Provider-level model/client parameter defaults",
)
access: AccessConfig = Field(
default_factory=AccessConfig,
description="Provider-level access constraints.",
)
models: list[ModelEntry] = Field(
default_factory=list,
description="Models available under this provider",
)
endpoints: list[EndpointConfig] = Field(
default_factory=list,
description="Multiple API endpoints for load balancing (overrides top-level api_key/base_url when set)",
)
def resolved_api_key(self) -> str:
"""Return the API key with environment variable references resolved.
For multi-endpoint providers, returns the first available resolved key.
"""
# Multi-endpoint mode: return first non-empty key
if self.endpoints:
for ep in self.endpoints:
key = ep.resolved_api_key()
if key:
return key
return ""
# Single-endpoint mode (legacy)
return _resolve_env_ref(self.api_key) if self.api_key else ""
def get_resolved_endpoints(self) -> list[EndpointConfig]:
"""Return only endpoints that have valid credentials."""
return [ep for ep in self.endpoints if ep.resolved_api_key() or (not ep.api_key and self.resolved_api_key())]
def has_credentials(self) -> bool:
"""Check if this provider has any usable credentials."""
if self.endpoints:
return any(ep.resolved_api_key() for ep in self.endpoints)
return bool(self.resolved_api_key())
class ModelDefaults(BaseModel):
"""Global default parameters for all models."""
temperature: float | None = Field(None, description="Sampling temperature")
max_tokens: int = Field(4096, description="Default max output tokens")
reasoning_effort: str | None = Field("high", description="Reasoning effort level (low/medium/high/xhigh)")
stream_usage: bool = Field(True, description="Whether to enable streaming token usage stats")
class StructuredConfig(BaseModel):
"""Top-level structured configuration for EvoScientist LLM settings."""
default_model: str = Field("", description="Default model ID or alias")
providers: dict[str, ProviderConfig] = Field(
default_factory=dict,
description="Provider definitions keyed by name",
)
model_defaults: ModelDefaults = Field(
default_factory=ModelDefaults,
description="Global default model parameters",
)
# ═══════════════════════════════════════════════════════════════════════════
# Config loading with layered discovery
# ═══════════════════════════════════════════════════════════════════════════
def _get_global_settings_path() -> Path:
"""Get global settings.yaml path via get_config_path()."""
from EvoScientist.config.settings import get_config_path
return get_config_path()
def _load_yaml_file(path: Path) -> dict:
"""Load a YAML file, returning empty dict on failure."""
if not path.is_file():
return {}
try:
with open(path) as f:
data = yaml.safe_load(f)
return data if isinstance(data, dict) else {}
except Exception:
return {}
def _detect_new_config_format(data: dict) -> bool:
"""Detect whether a YAML dict uses the new structured format.
The new format is identified by the presence of a 'providers' top-level key
that is a dict (not a flat string value).
"""
providers = data.get("providers")
return isinstance(providers, dict) and len(providers) > 0
def _deep_merge(base: dict, override: dict) -> dict:
"""Deep merge override into base. Override values take precedence."""
result = base.copy()
for key, value in override.items():
if key in result and isinstance(result[key], dict) and isinstance(value, dict):
result[key] = _deep_merge(result[key], value)
elif key in result and isinstance(result[key], list) and isinstance(value, list):
# For lists (like models), override completely replaces
result[key] = value
else:
result[key] = value
return result
def load_structured_config(
cli_overrides: dict[str, Any] | None = None,
) -> StructuredConfig:
"""Load structured config from ``settings.yaml`` (single source of truth).
Returns code defaults if the file doesn't exist or is invalid.
Args:
cli_overrides: Optional CLI argument overrides.
Returns:
StructuredConfig instance.
"""
settings_path = _get_global_settings_path()
settings_data = _load_yaml_file(settings_path)
config = StructuredConfig()
if _detect_new_config_format(settings_data):
config = StructuredConfig(**settings_data)
_remove_image_generation_models(config)
# Apply CLI overrides
if cli_overrides:
if "default_model" in cli_overrides and cli_overrides["default_model"]:
config.default_model = cli_overrides["default_model"]
return config
def _remove_image_generation_models(config: StructuredConfig) -> list[tuple[str, ProviderConfig, ModelEntry]]:
"""Remove image-only models from chat LLM providers and return them."""
from EvoScientist.config.image_models import is_image_generation_model
removed: list[tuple[str, ProviderConfig, ModelEntry]] = []
for prov_name, provider in config.providers.items():
chat_models: list[ModelEntry] = []
for model in provider.models:
if is_image_generation_model(model.id) or is_image_generation_model(model.alias):
removed.append((prov_name, provider, model))
else:
chat_models.append(model)
provider.models = chat_models
if config.default_model and is_image_generation_model(config.default_model):
config.default_model = ""
return removed
def move_image_models_to_dedicated_config(config: StructuredConfig) -> StructuredConfig:
"""Move image-only model entries out of chat LLM config."""
removed = _remove_image_generation_models(config)
if not removed:
return config
from EvoScientist.config.image_models import ImageModelEntry, merge_image_generation_models
entries: list[ImageModelEntry] = []
for prov_name, provider, model in removed:
api_key = provider.api_key
base_url = provider.base_url
if model.endpoint:
for endpoint in provider.endpoints:
if endpoint.name == model.endpoint:
api_key = endpoint.api_key or api_key
base_url = endpoint.base_url or base_url
break
entries.append(
ImageModelEntry(
id=model.id,
name=model.alias or model.id,
provider=prov_name,
api_key=api_key,
base_url=base_url,
supports_generation=True,
supports_edit=True,
default_size=str(model.params.get("size") or "1024x1024"),
default_quality=str(model.params.get("quality") or "auto"),
params=model.params,
)
)
merge_image_generation_models(entries, default_model=entries[0].id if entries else None)
return config
def save_structured_config(config: StructuredConfig | dict[str, Any]) -> StructuredConfig:
"""Validate and atomically persist the global structured config."""
validated = config if isinstance(config, StructuredConfig) else StructuredConfig.model_validate(config)
validated = move_image_models_to_dedicated_config(validated)
path = _get_global_settings_path()
path.parent.mkdir(parents=True, exist_ok=True)
existing = _load_yaml_file(path)
output = existing if isinstance(existing, dict) else {}
for key in _FLAT_LLM_CONFIG_KEYS:
output.pop(key, None)
output.update(validated.model_dump(mode="python"))
yaml_text = yaml.safe_dump(
output,
allow_unicode=False,
default_flow_style=False,
sort_keys=False,
)
fd, tmp_path = tempfile.mkstemp(
dir=str(path.parent),
prefix=".llm_config_",
suffix=".yaml.tmp",
text=True,
)
try:
with os.fdopen(fd, "w", encoding="utf-8") as f:
f.write(yaml_text)
f.flush()
os.fsync(f.fileno())
os.replace(tmp_path, path)
finally:
if os.path.exists(tmp_path):
os.unlink(tmp_path)
return validated
# ═══════════════════════════════════════════════════════════════════════════
# Model registry: lookup and resolution
# ═══════════════════════════════════════════════════════════════════════════
class ResolvedModel(BaseModel):
"""Fully resolved model information ready for get_chat_model()."""
provider_name: str = Field(..., description="Provider name (e.g. 'anthropic', 'deepseek')")
model_id: str = Field(..., description="Full model ID for the API call")
protocol: str = Field(..., description="Protocol to use (openai/anthropic/google-genai/ollama)")
api_key: str = Field("", description="Resolved API key")
base_url: str = Field("", description="Resolved base URL")
endpoint_name: str = Field("", description="Name of the selected endpoint (empty for single-endpoint providers)")
params: dict[str, Any] = Field(default_factory=dict, description="Merged parameters")
supports_vision: bool = False
supports_reasoning: bool = False
max_tokens: int = 4096
def _build_alias_index(config: StructuredConfig) -> dict[str, tuple[str, str]]:
"""Build alias → (provider_name, model_id) index.
Also indexes model IDs directly so both alias and full ID can be looked up.
"""
index: dict[str, tuple[str, str]] = {}
for prov_name, prov in config.providers.items():
for model in prov.models:
# Index by alias (if set)
if model.alias:
index[model.alias] = (prov_name, model.id)
# Index by full model ID
index[model.id] = (prov_name, model.id)
return index
def resolve_model(
model_ref: str,
config: StructuredConfig | None = None,
runtime_params: dict[str, Any] | None = None,
) -> ResolvedModel:
"""Resolve a model reference to a fully configured ResolvedModel.
Args:
model_ref: Model ID, alias, or provider-prefixed name (e.g. "anthropic/claude-sonnet-4-6").
config: StructuredConfig to use (loaded automatically if None).
runtime_params: Additional runtime parameter overrides.
Returns:
ResolvedModel with all fields populated.
Raises:
ValueError: If the model cannot be resolved.
"""
if config is None:
config = load_structured_config()
runtime_params = runtime_params or {}
# Handle provider-prefixed references: "anthropic/claude-sonnet-4-6"
explicit_provider = None
if "/" in model_ref:
parts = model_ref.split("/", 1)
explicit_provider = parts[0]
model_ref = parts[1]
# Try alias/ID lookup from structured configuration.
alias_index = _build_alias_index(config)
if model_ref in alias_index:
prov_name, model_id = alias_index[model_ref]
if explicit_provider and prov_name != explicit_provider:
prov_name = explicit_provider
model_id = model_ref
elif explicit_provider:
# Not in registry but user specified provider — use as-is
prov_name = explicit_provider
model_id = model_ref
else:
raise ValueError(
f"Model '{model_ref}' is not declared in structured LLM config. "
"Add it under providers.*.models or call it as 'provider/model'."
)
# Get provider config
provider = config.providers.get(prov_name)
if not provider:
raise ValueError(
f"Provider '{prov_name}' is not declared in structured LLM config."
)
# Find the model entry (if registered)
model_entry = None
for m in provider.models:
if m.id == model_id or m.alias == model_ref:
model_entry = m
break
# Three-level parameter merge: defaults ← model ← runtime
params: dict[str, Any] = {}
defaults = config.model_defaults
# Level 1: global defaults
if defaults.temperature is not None:
params["temperature"] = defaults.temperature
params["max_tokens"] = defaults.max_tokens
params["stream_usage"] = defaults.stream_usage
# Level 2: provider-level params
params.update(provider.params)
# Level 3: model-level params
if model_entry:
params.update(model_entry.params)
params["max_tokens"] = model_entry.max_tokens
# Override defaults with model-level values
if "temperature" not in model_entry.params and defaults.temperature is not None:
params["temperature"] = defaults.temperature
# ── Resolve credentials from provider or endpoint ────────────
resolved_api_key = provider.resolved_api_key()
resolved_base_url = _resolve_env_ref(provider.base_url) if provider.base_url else ""
resolved_extra_body = dict(provider.extra_body) if provider.extra_body else {}
resolved_headers = dict(provider.default_headers) if provider.default_headers else {}
selected_endpoint_name = ""
if provider.endpoints:
# Multi-endpoint mode: select endpoint via load balancing
preferred_ep = model_entry.endpoint if model_entry else ""
available_endpoints = [ep for ep in provider.endpoints if ep.resolved_api_key()]
if available_endpoints:
selected = _select_endpoint(available_endpoints, prov_name, preferred_ep)
resolved_api_key = selected.resolved_api_key()
resolved_base_url = selected.resolved_base_url() or resolved_base_url
selected_endpoint_name = selected.name
# Merge endpoint-level extra_body and headers (provider-level first, endpoint overrides)
if selected.extra_body:
resolved_extra_body.update(selected.extra_body)
if selected.default_headers:
resolved_headers.update(selected.default_headers)
if selected.params:
params.update(selected.params)
# Record endpoint call statistics
_endpoint_stats.record(prov_name, selected.name or "default", model_id)
logger.debug(
"endpoint selected: provider=%s endpoint=%s model=%s",
prov_name, selected.name or "default", model_id,
)
else:
# Single-endpoint provider — still record for token attribution
# so that usage_stats events carry correct endpoint info
_endpoint_stats.record(prov_name, "default", model_id)
# Level 4: runtime params
params.update(runtime_params)
# Store merged extra_body/headers in params for downstream consumption
if resolved_extra_body:
params["_extra_body"] = resolved_extra_body
if resolved_headers:
params["_default_headers"] = resolved_headers
return ResolvedModel(
provider_name=prov_name,
model_id=model_id,
protocol=provider.protocol,
api_key=resolved_api_key,
base_url=resolved_base_url,
endpoint_name=selected_endpoint_name,
params=params,
supports_vision=model_entry.supports_vision if model_entry else False,
supports_reasoning=model_entry.supports_reasoning if model_entry else False,
max_tokens=model_entry.max_tokens if model_entry else defaults.max_tokens,
)
def get_default_model(config: StructuredConfig | None = None) -> str:
"""Get the default model reference from config."""
if config is None:
config = load_structured_config()
return config.default_model
def list_available_models(config: StructuredConfig | None = None) -> list[dict[str, Any]]:
"""List all available models from configured providers.
Returns a list of dicts with keys: id, alias, provider, protocol,
max_tokens, supports_vision, supports_reasoning.
"""
if config is None:
config = load_structured_config()
models = []
for prov_name, provider in config.providers.items():
# Only include models from providers that have credentials
# (either via top-level api_key or via at least one endpoint)
has_creds = provider.has_credentials() or (
prov_name == "ollama" and bool(provider.base_url)
)
if not has_creds:
continue
for model in provider.models:
models.append({
"id": model.id,
"alias": model.alias,
"provider": prov_name,
"protocol": provider.protocol,
"max_tokens": model.max_tokens,
"supports_vision": model.supports_vision,
"supports_reasoning": model.supports_reasoning,
"provider_access": provider.access.model_dump(mode="python"),
"permission_mode": model.permission_mode,
"model_access": model.access.model_dump(mode="python"),
})
return models
File diff suppressed because it is too large Load Diff
+33
View File
@@ -0,0 +1,33 @@
"""Onboarding package.
The wizard's only package-level public entry point is :func:`run_onboard`.
Everything else lives in submodules — import directly from them:
- :mod:`EvoScientist.config.onboard.wizard` — orchestrator, ``run_onboard``,
``STEPS``, ``render_progress``
- :mod:`EvoScientist.config.onboard.steps` — per-step functions
- :mod:`EvoScientist.config.onboard.channels` — channel selection + setup
- :mod:`EvoScientist.config.onboard.helpers` — API-key prompt,
npx/node, LaTeX, iMessage helpers
- :mod:`EvoScientist.config.onboard.style` — Rich styles + ``_checkbox_ask``
- :mod:`EvoScientist.config.onboard.validators` — input validators
- :mod:`EvoScientist.config.onboard.prompter` — ``NonInteractivePrompter``
(CLI-answer container) + ``select_navigation_active`` for keyboard nav
- :mod:`EvoScientist.config.onboard.constants` — canonical valid-value sets
This module used to re-export every symbol from every submodule for
backward compat during the initial refactor; those re-exports have since
been removed to keep the public surface narrow. New code should always
import from the submodule that owns the symbol; test code should use
``patch("EvoScientist.config.onboard.<submodule>.<name>")`` paths.
"""
from __future__ import annotations
# Sole package-level public entry. ``EvoScientist.config`` re-exports this
# (via lazy ``__getattr__``) so ``from EvoScientist.config import
# run_onboard`` keeps working — that import path is used by the CLI and is
# the only documented external API.
from .wizard import run_onboard
__all__ = ["run_onboard"]
+958
View File
@@ -0,0 +1,958 @@
"""Channel selection + per-channel configuration.
`_step_channels` is the big one — over 700 lines that walk the user through
selecting which messaging channels to enable and collecting credentials for
each.
"""
from __future__ import annotations
import questionary
from questionary import Choice
from ..settings import EvoScientistConfig
from .helpers import (
_setup_imessage,
)
from .style import (
QMARK,
WIZARD_STYLE,
console,
)
def _step_channels(config: EvoScientistConfig) -> dict[str, object]:
"""Step: Select channels to enable on startup.
Presents a multi-select list of supported channels.
For each selected channel, prompts for required credentials
and validates them via the channel's probe function.
Args:
config: Current configuration.
Returns:
Dict mapping config field names to their new values.
Empty dict when the user skips or selects nothing.
"""
# Currently enabled channels
_currently_enabled = {
t.strip()
for t in (getattr(config, "channel_enabled", "") or "").split(",")
if t.strip()
}
# Legacy iMessage compat
if (
getattr(config, "imessage_enabled", False)
and "imessage" not in _currently_enabled
):
_currently_enabled.add("imessage")
# Direct pip packages for each channel extra. Used to install the
# exact dependency without requiring the evoscientist package itself
# to be resolvable on PyPI (e.g. editable / dev installs).
_CHANNEL_PIP_DEPS: dict[str, list[str]] = {
"telegram": ["python-telegram-bot>=21.0"],
"discord": ["discord.py>=2.3"],
"slack": ["slack-sdk>=3.27", "aiohttp>=3.9"],
"feishu": ["aiohttp>=3.9", "qrcode>=7.4"],
"dingtalk": ["aiohttp>=3.9"],
"wechat": [
"pycryptodome>=3.20",
"aiohttp>=3.9",
"qrcode>=7.4",
"certifi>=2024.0",
],
"qq": ["qq-botpy>=1.0", "cryptography>=41.0", "qrcode>=7.4"],
}
# Channel definitions:
# (value, display_name, required_fields, import_check, pip_extra)
# required_fields entries are (field_name, prompt_label, is_secret).
# ``is_secret=True`` triggers a password prompt (no echo, no default echo)
# so bot tokens / OAuth secrets / IMAP+SMTP passwords don't leak into
# terminal scrollback, screen recordings, or support sessions.
_CHANNELS = [
(
"telegram",
"Telegram",
[("telegram_bot_token", "Bot token (from @BotFather)", True)],
"telegram",
"telegram",
),
(
"discord",
"Discord",
[("discord_bot_token", "Bot token", True)],
"discord",
"discord",
),
(
"slack",
"Slack",
[
("slack_bot_token", "Bot token (xoxb-...)", True),
("slack_app_token", "App token for Socket Mode (xapp-...)", True),
],
"slack_sdk",
"slack",
),
(
"feishu",
"Feishu",
[
("feishu_app_id", "App ID", False),
("feishu_app_secret", "App Secret", True),
],
"aiohttp",
"feishu",
),
(
"dingtalk",
"DingTalk",
[
("dingtalk_client_id", "Client ID (AppKey)", False),
("dingtalk_client_secret", "Client Secret (AppSecret)", True),
],
"aiohttp",
"dingtalk",
),
(
"wechat",
"WeChat",
[], # backend-specific fields prompted in the wechat branch below
("aiohttp", "qrcode", "Crypto", "certifi"),
"wechat",
),
(
"email",
"Email",
[
("email_imap_host", "IMAP host", False),
("email_imap_username", "IMAP username", False),
("email_imap_password", "IMAP password", True),
("email_smtp_host", "SMTP host", False),
("email_smtp_username", "SMTP username", False),
("email_smtp_password", "SMTP password", True),
("email_from_address", "From address", False),
],
None,
None,
),
(
"qq",
"QQ",
[
("qq_app_id", "App ID", False),
("qq_app_secret", "App Secret", True),
],
"botpy",
"qq",
),
(
"signal",
"Signal",
[("signal_phone_number", "Phone number (E.164)", False)],
None,
None,
),
("imessage", "iMessage", [], None, None), # handled via _setup_imessage()
]
choices = [
Choice(
title=display,
value=value,
checked=value in _currently_enabled,
)
for value, display, *_ in _CHANNELS
]
selected = questionary.checkbox(
"Select channels to enable (Space to toggle, Enter to confirm):",
choices=choices,
style=WIZARD_STYLE,
qmark=QMARK,
).ask()
if selected is None:
raise KeyboardInterrupt()
updates: dict[str, object] = {}
if not selected:
updates["channel_enabled"] = ""
updates["imessage_enabled"] = False
return updates
from ...mcp.registry import install_library, pip_install_hint
# Build a lookup for channel definitions
_ch_lookup = {
v: (v, d, fields, imp, extra) for v, d, fields, imp, extra in _CHANNELS
}
enabled_channels: list[str] = []
for ch_name in selected:
_, display, required_fields, import_check, pip_extra = _ch_lookup[ch_name]
console.print(f"\n [bold cyan]── {display} ──[/bold cyan]")
# Check pip dependency before proceeding
if import_check:
_required_imports: tuple[str, ...] = (
(import_check,)
if isinstance(import_check, str)
else tuple(import_check)
)
_pkg_ready = False
try:
for _module_name in _required_imports:
__import__(_module_name)
_pkg_ready = True
except ImportError:
console.print(" [yellow]✗ Required package not installed.[/yellow]")
# Determine packages to install
_pip_pkgs = _CHANNEL_PIP_DEPS.get(pip_extra, []) if pip_extra else []
_pkg_display = (
" ".join(f'"{p}"' for p in _pip_pkgs)
if _pip_pkgs
else f'"evoscientist[{pip_extra}]"'
)
install_now = questionary.confirm(
f"Install {_pkg_display} now?",
default=True,
style=WIZARD_STYLE,
qmark=f" {QMARK}",
).ask()
if install_now is None:
raise KeyboardInterrupt() from None
if install_now:
console.print(f" [dim]Installing {_pkg_display}...[/dim]")
if _pip_pkgs:
_ok = all(install_library(p) for p in _pip_pkgs)
else:
_ok = install_library(f"evoscientist[{pip_extra}]")
if _ok:
# Verify the imports actually work now
try:
for _module_name in _required_imports:
__import__(_module_name)
console.print(" [green]✓ Installed successfully.[/green]")
_pkg_ready = True
except ImportError:
console.print(
" [red]✗ Package installed but import failed.[/red]"
)
console.print(
" [dim]Try restarting and running:[/dim] evosci channel setup"
)
else:
console.print(" [red]✗ Installation failed.[/red]")
console.print(
f" [dim]Run manually:[/dim] {pip_install_hint()} {_pkg_display}"
)
if not _pkg_ready:
# Previously-enabled channels are silently dropped from
# ``channel_enabled`` if we just ``continue`` — warn.
if ch_name in _currently_enabled:
console.print(
f" [bold yellow]⚠ {display} will be DISABLED[/bold yellow]"
" [dim](dependency missing — re-run after install)[/dim]"
)
else:
console.print(
f" [dim]Skipping {display} — dependency not installed.[/dim]"
)
continue
# Special handling for iMessage
if ch_name == "imessage":
ready = _setup_imessage()
if not ready:
console.print()
enable_anyway = questionary.confirm(
"Enable iMessage anyway? (will try to connect on startup)",
default=False,
style=WIZARD_STYLE,
qmark=f" {QMARK}",
).ask()
if enable_anyway is None:
raise KeyboardInterrupt()
if not enable_anyway:
continue
# Allowed senders
senders = questionary.text(
"Allowed senders (comma-separated, empty = all):",
default=getattr(config, "imessage_allowed_senders", ""),
style=WIZARD_STYLE,
qmark=f" {QMARK}",
).ask()
if senders is None:
raise KeyboardInterrupt()
updates["imessage_enabled"] = True
updates["imessage_allowed_senders"] = senders.strip()
enabled_channels.append("imessage")
continue
# QQ: offer scan-to-configure before falling back to manual entry.
# The bot must already exist at q.qq.com — scanning binds the
# developer's QQ account to it and returns app_id + client_secret.
_qq_scanned = False
_feishu_scanned = False
if ch_name == "qq":
scan_choices = [
Choice(
title="Scan QR code (recommended — auto-fill App ID & Secret)",
value="scan",
),
Choice(title="Enter App ID and Secret manually", value="manual"),
]
scan_choice = questionary.select(
"Configure QQ Bot:",
choices=scan_choices,
default="scan",
style=WIZARD_STYLE,
qmark=f" {QMARK}",
use_indicator=True,
).ask()
if scan_choice is None:
raise KeyboardInterrupt()
if scan_choice == "scan":
# Preflight: AES-GCM decryption needs `cryptography`.
# `qrcode` is a soft dep — onboard.py degrades to URL-only display.
try:
import cryptography # noqa: F401
except ImportError:
console.print(
' [yellow]✗ QR scan requires "cryptography".[/yellow]'
)
install_now = questionary.confirm(
'Install "cryptography" now?',
default=True,
style=WIZARD_STYLE,
qmark=f" {QMARK}",
).ask()
if install_now is None:
raise KeyboardInterrupt() from None
if install_now and install_library("cryptography>=41.0"):
console.print(" [green]✓ Installed cryptography.[/green]")
else:
console.print(
" [yellow]⚠ Falling back to manual entry.[/yellow]"
)
scan_choice = "manual"
if scan_choice == "scan":
from ...channels.qq.onboard import qr_register
console.print(
" [dim]Make sure the bot is registered at"
" https://q.qq.com first — scanning binds an"
" existing app, it does not create one.[/dim]"
)
try:
creds = qr_register()
except Exception as exc:
console.print(f" [red]✗ Scan failed: {exc}[/red]")
creds = None
if creds:
updates["qq_app_id"] = creds["app_id"]
updates["qq_app_secret"] = creds["client_secret"]
console.print(
f" [green]✓ Bound QQ Bot (App ID: {creds['app_id']})[/green]"
)
_qq_scanned = True
else:
console.print(
" [yellow]⚠ Scan did not complete — falling"
" back to manual entry.[/yellow]"
)
# Feishu: offer scan-to-create before falling back to manual entry.
# Unlike QQ, this provisions a brand-new PersonalAgent app with the
# required IM permissions attached, then returns app_id + app_secret.
if ch_name == "feishu":
scan_choices = [
Choice(
title="Scan QR code (recommended — auto-create app, fill App ID & Secret)",
value="scan",
),
Choice(title="Enter App ID and Secret manually", value="manual"),
]
scan_choice = questionary.select(
"Configure Feishu / Lark:",
choices=scan_choices,
default="scan",
style=WIZARD_STYLE,
qmark=f" {QMARK}",
use_indicator=True,
).ask()
if scan_choice is None:
raise KeyboardInterrupt()
if scan_choice == "scan":
# `qrcode` is the only soft dep needed — onboard prints the URL
# if it's missing, but the UX is much worse, so offer to install.
try:
import qrcode # noqa: F401
except ImportError:
console.print(
' [yellow]✗ QR scan looks best with "qrcode".[/yellow]'
)
install_now = questionary.confirm(
'Install "qrcode" now?',
default=True,
style=WIZARD_STYLE,
qmark=f" {QMARK}",
).ask()
if install_now is None:
raise KeyboardInterrupt() from None
if install_now and install_library("qrcode>=7.4"):
console.print(" [green]✓ Installed qrcode.[/green]")
else:
console.print(
" [yellow]⚠ Falling back to manual entry.[/yellow]"
)
scan_choice = "manual"
if scan_choice == "scan":
# Region selection — accounts.feishu.cn vs accounts.larksuite.com.
# The poll endpoint auto-switches if the scanning user is on the
# other tenant, so this is just a starting hint.
region_choices = [
Choice(title="Feishu (飞书, mainland China)", value="feishu"),
Choice(title="Lark (overseas)", value="lark"),
]
region = questionary.select(
"Region:",
choices=region_choices,
default="feishu",
style=WIZARD_STYLE,
qmark=f" {QMARK}",
use_indicator=True,
).ask()
if region is None:
raise KeyboardInterrupt()
if scan_choice == "scan":
from ...channels.feishu.onboard import qr_register
console.print(
" [dim]A QR code will be printed below — open Feishu or"
" Lark on your phone and scan it. The platform will"
" auto-create a bot app with IM permissions and return"
" the credentials here.[/dim]"
)
try:
creds = qr_register(initial_domain=region)
except Exception as exc:
console.print(f" [red]✗ Scan failed: {exc}[/red]")
creds = None
if creds:
updates["feishu_app_id"] = creds["app_id"]
updates["feishu_app_secret"] = creds["app_secret"]
# Sync open-platform domain to the resolved region
updates["feishu_domain"] = (
"https://open.larksuite.com"
if creds.get("domain") == "lark"
else "https://open.feishu.cn"
)
bot_name = creds.get("bot_name")
if bot_name:
console.print(
f' [green]✓ Bound Feishu bot "{bot_name}"'
f" (App ID: {creds['app_id']})[/green]"
)
else:
console.print(
f" [green]✓ Bound Feishu app"
f" (App ID: {creds['app_id']})[/green]"
)
_feishu_scanned = True
else:
console.print(
" [yellow]⚠ Scan did not complete — falling"
" back to manual entry.[/yellow]"
)
# WeChat: pick backend (wecom / wechatmp / personal), then prompt
# backend-specific fields. Personal-WeChat has no static credentials —
# we offer an interactive QR-scan that obtains and persists them.
if ch_name == "wechat":
backend_choices = [
Choice(
title="WeCom (企业微信应用) — most stable, official API",
value="wecom",
),
Choice(
title="Official Account (微信公众号) — public-facing bots",
value="wechatmp",
),
Choice(
title="Personal WeChat (个人微信, iLink) — QR-code scan login",
value="personal",
),
]
wechat_backend = questionary.select(
"WeChat backend:",
choices=backend_choices,
default=getattr(config, "wechat_backend", "") or "wecom",
style=WIZARD_STYLE,
qmark=f" {QMARK}",
use_indicator=True,
).ask()
if wechat_backend is None:
raise KeyboardInterrupt()
updates["wechat_backend"] = wechat_backend
# Both WeCom and WeChat MP need the same non-empty-required
# treatment as the generic required_fields loop below — newly
# enabling either with blank credentials would leave the channel
# half-configured and only fail at first message.
wechat_newly_enabled = "wechat" not in _currently_enabled
wechat_fields_for_backend: list[tuple[str, str, bool]] = []
if wechat_backend == "wecom":
wechat_fields_for_backend = [
("wechat_wecom_corp_id", "WeCom Corp ID", False),
("wechat_wecom_agent_id", "WeCom Agent ID", False),
("wechat_wecom_secret", "WeCom Secret", True),
]
elif wechat_backend == "wechatmp":
wechat_fields_for_backend = [
("wechat_mp_app_id", "Official Account App ID", False),
("wechat_mp_app_secret", "Official Account App Secret", True),
]
if wechat_backend in ("wecom", "wechatmp"):
for field_name, prompt_label, is_secret in wechat_fields_for_backend:
current = getattr(config, field_name, "")
while True:
if is_secret:
masked_hint = (
f" (current: ***{current[-4:]})" if current else ""
)
value = questionary.password(
f"{prompt_label}{masked_hint}:",
style=WIZARD_STYLE,
qmark=f" {QMARK}",
).ask()
else:
value = questionary.text(
f"{prompt_label}:",
default=current,
style=WIZARD_STYLE,
qmark=f" {QMARK}",
).ask()
if value is None:
raise KeyboardInterrupt()
value = value.strip()
if not value and current:
break # keep existing
if not value and wechat_newly_enabled:
console.print(
f" [yellow]{prompt_label} is required to "
"enable WeChat. Press Ctrl+C to cancel.[/yellow]"
)
continue
updates[field_name] = value
break
elif wechat_backend == "personal":
personal_choices = [
Choice(
title="Scan QR code now (recommended — login to a personal WeChat account)",
value="scan",
),
Choice(
title="I already have an account_id — enter it manually",
value="manual",
),
]
personal_choice = questionary.select(
"Personal WeChat login:",
choices=personal_choices,
default="scan",
style=WIZARD_STYLE,
qmark=f" {QMARK}",
use_indicator=True,
).ask()
if personal_choice is None:
raise KeyboardInterrupt()
if personal_choice == "scan":
from ...channels.wechat.personal import _account_dir as _wp_dir
_accounts_path = _wp_dir()
console.print(
" [dim]A QR code will be printed below — open WeChat on"
" your phone and scan it. The session token is saved"
f" to {_accounts_path}.[/dim]"
)
try:
import asyncio
from ...channels.wechat.personal import qr_login
creds = asyncio.run(qr_login())
except Exception as exc:
console.print(f" [red]✗ Scan failed: {exc}[/red]")
creds = None
if creds:
updates["wechat_personal_account_id"] = creds["account_id"]
# Token is persisted on disk by qr_login(); the channel
# reads it from the per-account store at runtime, so we
# intentionally do NOT copy it into the main config here
# (avoids stale duplicates and broader secret exposure).
console.print(
f" [green]✓ Logged in (account_id: "
f"{creds['account_id'][:12]}…)[/green]"
)
else:
console.print(
" [yellow]⚠ QR login did not complete — falling"
" back to manual entry.[/yellow]"
)
personal_choice = "manual"
if personal_choice == "manual":
current_id = getattr(config, "wechat_personal_account_id", "")
wechat_newly_enabled = "wechat" not in _currently_enabled
while True:
account_id = questionary.text(
"iLink account_id (from a previous --qr-login run):",
default=current_id,
style=WIZARD_STYLE,
qmark=f" {QMARK}",
).ask()
if account_id is None:
raise KeyboardInterrupt()
account_id = account_id.strip()
if not account_id and current_id:
break # keep existing
if not account_id and wechat_newly_enabled:
console.print(
" [yellow]account_id is required to enable "
"WeChat Personal. Press Ctrl+C to cancel.[/yellow]"
)
continue
updates["wechat_personal_account_id"] = account_id
break
# Prompt for required fields. Secret fields use ``questionary.password``
# so the entered value (and the existing one shown as a hint) are
# never echoed to the terminal — see _CHANNELS docstring above.
if not _qq_scanned and not _feishu_scanned:
# "Newly enabled" = this channel wasn't in the user's prior
# ``channel_enabled`` list. Required fields with no existing
# value must be non-empty for newly enabled channels — saving
# blanks leaves the channel half-configured and only surfaces
# the problem on the first message.
newly_enabled = ch_name not in _currently_enabled
for field_name, prompt_label, is_secret in required_fields:
current = getattr(config, field_name, "")
while True:
if is_secret:
masked_hint = (
f" (current: ***{current[-4:]})" if current else ""
)
value = questionary.password(
f"{prompt_label}{masked_hint}:",
style=WIZARD_STYLE,
qmark=f" {QMARK}",
).ask()
else:
value = questionary.text(
f"{prompt_label}:",
default=current,
style=WIZARD_STYLE,
qmark=f" {QMARK}",
).ask()
if value is None:
raise KeyboardInterrupt()
value = value.strip()
# Empty input + existing value → keep existing (this is
# the "re-run wizard, no change to this field" path).
if not value and current:
break
# Empty input + newly enabling channel → not OK; the
# channel would be enabled with broken creds. Re-prompt.
if not value and newly_enabled:
console.print(
f" [yellow]{prompt_label} is required to enable "
f"{display}. Press Ctrl+C to cancel instead.[/yellow]"
)
continue
# Empty input + previously enabled but never set
# (unlikely, but tolerate) → still allow blank-through
# so the user isn't blocked re-running configure later.
updates[field_name] = value
break
# Feishu: subscription mode + optional fields
if ch_name == "feishu":
mode_choices = [
Choice(
title="Webhook (requires public IP / port forwarding)",
value="webhook",
),
Choice(
title="WebSocket long connection (no public IP needed)",
value="websocket",
),
]
sub_mode = questionary.select(
"Subscription mode:",
choices=mode_choices,
default="webhook",
style=WIZARD_STYLE,
qmark=f" {QMARK}",
use_indicator=True,
).ask()
if sub_mode is None:
raise KeyboardInterrupt()
updates["feishu_subscription_mode"] = sub_mode
if sub_mode == "websocket":
# WebSocket mode needs lark-oapi SDK
try:
__import__("lark_oapi")
except ImportError:
console.print(
' [yellow]✗ WebSocket mode requires "lark-oapi".[/yellow]'
)
install_sdk = questionary.confirm(
'Install "lark-oapi>=1.4.0" now?',
default=True,
style=WIZARD_STYLE,
qmark=f" {QMARK}",
).ask()
if install_sdk is None:
raise KeyboardInterrupt() from None
if install_sdk:
console.print(' [dim]Installing "lark-oapi"...[/dim]')
if install_library("lark-oapi>=1.4.0"):
console.print(" [green]✓ Installed successfully.[/green]")
else:
console.print(" [red]✗ Installation failed.[/red]")
console.print(
f" [dim]Run manually:[/dim] {pip_install_hint()} "
'"lark-oapi>=1.4.0"'
)
else:
# Webhook mode: prompt optional verification/encryption fields.
# Both are credentials — use password() so they don't echo.
console.print(
" [dim]The following fields are optional"
" (press Enter to skip):[/dim]"
)
for field_name, prompt_label in [
("feishu_verification_token", "Verification Token (optional)"),
("feishu_encrypt_key", "Encrypt Key (optional)"),
]:
current = getattr(config, field_name, "")
masked_hint = f" (current: ***{current[-4:]})" if current else ""
value = questionary.password(
f"{prompt_label}{masked_hint}:",
style=WIZARD_STYLE,
qmark=f" {QMARK}",
).ask()
if value is None:
raise KeyboardInterrupt()
value = value.strip()
if not value and current:
# Keep existing value when user just presses Enter.
continue
updates[field_name] = value
# Allowed senders (common for all channels)
senders_field = f"{ch_name}_allowed_senders"
if hasattr(config, senders_field):
senders = questionary.text(
"Allowed senders (comma-separated, empty = all):",
default=getattr(config, senders_field, ""),
style=WIZARD_STYLE,
qmark=f" {QMARK}",
).ask()
if senders is None:
raise KeyboardInterrupt()
updates[senders_field] = senders.strip()
# Probe validation
_probe_channel(ch_name, config, updates)
enabled_channels.append(ch_name)
updates["channel_enabled"] = ",".join(enabled_channels)
# Keep legacy field in sync
updates["imessage_enabled"] = "imessage" in enabled_channels
# --- Common prompt: send thinking (shown when any channel is enabled) ---
if enabled_channels:
console.print("\n [bold cyan]── Channel Settings ──[/bold cyan]")
thinking_choices = [
Choice(title="On (forward model reasoning)", value=True),
Choice(title="Off (only send final responses)", value=False),
]
send_thinking = questionary.select(
"Send thinking panel in channel?",
choices=thinking_choices,
default=config.channel_send_thinking,
style=WIZARD_STYLE,
qmark=f" {QMARK}",
use_indicator=True,
).ask()
if send_thinking is None:
raise KeyboardInterrupt()
updates["channel_send_thinking"] = send_thinking
return updates
def _probe_channel(
ch_name: str,
config: EvoScientistConfig,
updates: dict[str, object],
) -> None:
"""Run the probe for a channel type and print the result.
Non-fatal: prints a warning on failure but does not prevent enabling.
"""
import asyncio
def _val(key: str, fallback: str = "") -> str:
"""Get a value from updates first, then config, then fallback."""
if key in updates:
return str(updates[key])
return str(getattr(config, key, fallback))
console.print(" [dim]Validating credentials...[/dim]")
async def _run() -> tuple[bool, str]:
if ch_name == "telegram":
from ...channels.telegram.probe import validate_telegram_token
return await validate_telegram_token(
_val("telegram_bot_token"),
_val("telegram_proxy") or None,
)
elif ch_name == "discord":
from ...channels.discord.probe import validate_discord_token
return await validate_discord_token(
_val("discord_bot_token"),
_val("discord_proxy") or None,
)
elif ch_name == "slack":
from ...channels.slack.probe import validate_slack_tokens
return await validate_slack_tokens(
_val("slack_bot_token"),
_val("slack_app_token") or None,
_val("slack_proxy") or None,
)
elif ch_name == "wechat":
backend = _val("wechat_backend", "wecom")
if backend == "wechatmp":
from ...channels.wechat.probe import validate_wechat_mp
return await validate_wechat_mp(
_val("wechat_mp_app_id"),
_val("wechat_mp_app_secret"),
_val("wechat_proxy") or None,
)
elif backend == "personal":
from ...channels.wechat.probe import validate_wechat_personal
return await validate_wechat_personal(
_val("wechat_personal_account_id"),
_val("wechat_personal_token"),
)
else:
from ...channels.wechat.probe import validate_wecom
return await validate_wecom(
_val("wechat_wecom_corp_id"),
_val("wechat_wecom_secret"),
_val("wechat_proxy") or None,
)
elif ch_name == "feishu":
from ...channels.feishu.probe import validate_feishu_credentials
return await validate_feishu_credentials(
_val("feishu_app_id"),
_val("feishu_app_secret"),
_val("feishu_domain", "https://open.feishu.cn"),
)
elif ch_name == "dingtalk":
from ...channels.dingtalk.probe import validate_dingtalk
return await validate_dingtalk(
_val("dingtalk_client_id"),
_val("dingtalk_client_secret"),
_val("dingtalk_proxy") or None,
)
elif ch_name == "email":
from ...channels.email.probe import validate_email_imap
return await validate_email_imap(
_val("email_imap_host"),
int(_val("email_imap_port", "993")),
_val("email_imap_username"),
_val("email_imap_password"),
_val("email_imap_use_ssl", "True").lower() not in ("false", "0", "no"),
)
elif ch_name == "qq":
from ...channels.qq.probe import validate_qq
return await validate_qq(
_val("qq_app_id"),
_val("qq_app_secret"),
)
elif ch_name == "signal":
from ...channels.signal.probe import validate_signal
return await validate_signal(
_val("signal_phone_number"),
_val("signal_cli_path", "signal-cli"),
int(_val("signal_rpc_port", "7583")),
)
else:
return True, "No probe available"
try:
try:
loop = asyncio.get_event_loop()
if loop.is_running():
import nest_asyncio # type: ignore[import-untyped]
nest_asyncio.apply()
except RuntimeError:
loop = asyncio.new_event_loop()
asyncio.set_event_loop(loop)
ok, detail = loop.run_until_complete(_run())
if ok:
console.print(f" [green]✓ {detail}[/green]")
else:
console.print(f" [yellow]⚠ {detail}[/yellow]")
console.print(
" [dim]Channel will still be enabled — check credentials later.[/dim]"
)
except Exception as e:
console.print(f" [yellow]⚠ Could not validate: {e}[/yellow]")
console.print(
" [dim]Channel will still be enabled — check credentials later.[/dim]"
)
# =============================================================================
# Progress Rendering (for tests and potential future use)
# =============================================================================

Some files were not shown because too many files have changed in this diff Show More