Commit Graph

640 Commits

Author SHA1 Message Date
m4 aae8d0a379 feat: workspace file references, read-file-images middleware, image model enabled flag
In-progress work committed to unblock the config import/export plan:
- prompts: FILE_REFERENCES section for workspace-relative file citation
- backends: resolve quoted virtual absolute paths onto the sandbox workspace
- middleware: read_file_images middleware; message_budget extensions
- image_gen/model_registry: image model 'enabled' flag refactor
- memory/launch, gateway/background_runs, tools/image follow-ons
- scripts: dev_backend.sh, release.sh
- tests for the above
2026-08-12 19:43:35 +08:00
m4 f3ca381ab3 feat(release): local release script with unified version bump and Gitea publish
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-12 17:25:43 +08:00
m4 8f6a568646 feat(update): system update/status/rollback HTTP endpoints with system:write scope 2026-08-12 16:42:10 +08:00
m4 91e2a87be5 feat(update): rollback version list with BREAKING-DB truncation 2026-08-12 16:34:16 +08:00
m4 03771df508 feat(update): standalone updater runner (wait/install/restart/result) 2026-08-12 16:30:49 +08:00
m4 90e773b30b feat(update): plan builder, plan lock and updater spawn 2026-08-12 16:27:39 +08:00
m4 40b896bcb5 feat(update): deployment detection and install command builder 2026-08-12 16:24:17 +08:00
m4 174f03b92d feat(update): stage artifacts under config dir and reuse verified local files 2026-08-12 16:21:45 +08:00
m4 194402fc88 fix(usage): stop sharing EVOSCIENTIST_DEPLOYMENT_ID with scope partitioning
The usage identity exported the same variable the scope registry reads to
partition workspace scopes, so a backend started with the usage environment
(61d1b61b) could not see scopes provisioned under the workspace-derived id
(5a882492) and every scope lookup 404'd. Usage attribution now reads
EVOSCIENTIST_USAGE_DEPLOYMENT_ID; the scope side keeps the original variable.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-30 22:08:21 +08:00
m4 c6efdaa13f docs(webui): thinking-timer design for the chat message area
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-30 11:11:33 +08:00
m4 08aa0d0e05 feat(registry): mutually-exclusive sampling_override frozen into run snapshots 2026-07-30 10:08:26 +08:00
m4 8eff551bb6 docs(registry): implementation plan for mutually-exclusive sampling_override
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-30 09:48:34 +08:00
m4 862c1e9743 docs(registry): supersede flat temperature/top_p overrides with mutually-exclusive sampling_override
temperature and top_p cannot be set together; the flat-fields design
allowed both. Replaced by a discriminated union (default | temperature
| top_p) where overriding one omits the other from the request entirely.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-30 09:29:06 +08:00
m4 d7711484ef feat(registry): expose generation defaults on selectable models 2026-07-28 17:42:16 +08:00
m4 ba7d908276 feat(registry): freeze per-thread temperature/top_p overrides into run snapshots
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-28 17:23:26 +08:00
m4 e57ecd4588 docs(plan): per-thread temperature/top_p override implementation plan
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-28 16:59:37 +08:00
m4 261845830d docs(spec): per-thread temperature/top_p override design
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-28 16:17:16 +08:00
m4 8efd4ad0ab fix(image-gen): wrap corrupt-config GET in the 422 error envelope 2026-07-24 09:45:57 +08:00
m4 01e674e1bc feat(image-gen): add /api/image-generation config endpoints with masked keys 2026-07-24 09:36:29 +08:00
m4 7f26ecc19a fix(image-gen): sanitize config validation errors; widen expected_revision type
load_image_generation_settings now re-raises pydantic ValidationError as a
sanitized ImageGenError carrying only field locations and error types, so a
mis-indented config.yaml can never echo a literal API key into agent-visible
errors. Also widen save_registry's expected_revision annotation to
int | None to match the http_api caller (value remains ignored).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-24 07:51:36 +08:00
m4 395baab7d5 test(registry): add IMAGE_MODEL_NOT_CHAT_MODEL to error taxonomy mirror 2026-07-24 07:44:37 +08:00
m4 ccf4173990 feat(registry): reject image-only models in chat model saves 2026-07-24 07:36:34 +08:00
m4 a57c52c676 feat(image-gen): add image-artist skill for generation workflow 2026-07-24 07:26:23 +08:00
m4 384bc13a5b feat(image-gen): add generate_image/edit_image agent tools
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-23 22:36:06 +08:00
m4 5622b40cf3 feat(image-gen): add service layer with safe artifact saving 2026-07-23 22:10:14 +08:00
m4 802f71bd46 feat(image-gen): add Gemini (Imagen) image adapter 2026-07-23 21:53:40 +08:00
m4 cc9dfb1cc9 fix(image-gen): address review findings in OpenAI image adapter
Scope download Authorization header to the provider origin, translate
httpx errors in _download/edit into safe ImageGenError messages, and
prevent entry params from clobbering core payload keys; also harden
strip_data_uri against malformed values and drop dead code.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-23 21:44:57 +08:00
m4 12e4f34005 feat(image-gen): add OpenAI-compatible image adapter with responses fallback 2026-07-23 21:28:38 +08:00
m4 b0f9a8d785 feat(image-gen): add image_generation config section and model detection
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-23 21:03:50 +08:00
m4 15cc389b3d feat(registry): make expected_revision optional and ignored in save API
Regenerate the checked-in OpenAPI export to match the relaxed schema.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-23 20:55:52 +08:00
m4 3e67e64067 refactor(registry): make saves last-write-wins, ignore expected_revision
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-23 20:47:08 +08:00
m4 bb9bed82e1 feat(runtime)!: remove auxiliary model role, resolve all roles from run snapshot
ModelRole collapses to "primary": every role (main, tool selector, memory
agents, subagents, summarizer) resolves to the snapshot's frozen primary
model, per design 6.1/8.3 — users typically configure a single usable LLM,
so compile-time auxiliary bindings were bypassing run snapshots and
mis-attributing usage. Legacy auxiliary keys in stored snapshots, registry
JSON, and thread metadata are tolerated on read and dropped.

BREAKING CHANGE: ThreadModelSelection no longer carries an auxiliary ref;
snapshot selection_hash is computed over {primary, reasoning_effort} only;
ConfigurableModelMiddleware(role="auxiliary") is rejected.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-23 10:18:44 +08:00
m4 1a01fb5d74 docs(env): drop stale LLM-key and provider-admin-token entries from .env.example
Provider credentials now live exclusively in the Model Registry and the
x-evoscientist-admin-token / provider-admin-token mechanism was removed;
no code reads these variables anymore.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-21 21:14:31 +08:00
m4 087781556b fix(runtime): serve Config API in bootstrap and verify snapshot issuer by registered deployment set
Graph construction no longer raises on a bootstrap registry: build paths
bind a shared RegistryNotReadyChatModel placeholder that fails every call
with MODEL_REGISTRY_NOT_READY, so langgraph dev serves the Config API for
first-time configuration while run creation stays forbidden.

Run snapshot binding no longer compares configurable
'workspace_deployment_id' (the workspace-isolation scope id) against the
snapshot's issuing deployment — a mismatch that made every BFF run fail
with SNAPSHOT_NOT_FOUND. SnapshotService.get_for_run verifies thread_id
equality plus membership in the platform-registered deployment set
(local_deployment_id + webui_delegation_public_keys entries).

Blocking I/O moved off the event loop for langgraph dev's blockbuster:
Config API authentication (store mkdir/chmod, config.yaml read, jti
registration) and the message-budget snapshot read now run in threads,
with the immutable snapshot cached per run.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-21 20:21:32 +08:00
m4 a1bfbd92ca chore: untrack accidentally committed docx
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-21 18:10:40 +08:00
m4 421a664336 feat(runtime)!: complete legacy removal, local snapshot entries, and TTL cleanup
- Remove legacy provider profiles, admin-token auth, /model command,
  model picker widget, and config.yaml LLM fields (design doc section 10)
- Wire CLI/channels/cron and async sub-agents through the local snapshot
  entry; run creation rejects model config outside runtime_snapshot_id
- Add periodic run-snapshot TTL cleanup to the config service lifespan
- Isolate tests from the real config dir and activate the registry where
  run/model paths fail closed in bootstrap

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-21 18:10:23 +08:00
m4 57176b359a feat(runtime)!: switch middleware and agent factories to snapshot-driven models
Replace config.yaml-driven model selection with registry snapshot resolution
across the runtime chain:

- ConfigurableModelMiddleware reads configurable["runtime_snapshot_id"] only;
  model/model_provider overrides are rejected with MODEL_CONFIG_OUTSIDE_SNAPSHOT
- MessageBudgetMiddleware derives budgets from snapshot reserves
  (system/tools/attachments) and re-resolves the summarizer per snapshot
- Agent factory and subagent factory resolve models via SnapshotRuntime
  (auxiliary/tool_selector/scheduler -> defaults.auxiliary ?? defaults.primary)
- Remove ModelFallbackMiddleware, /model-fallback command, and fallback chain
- Add model_registry/runtime.py SnapshotRuntime glue layer

Legacy config.yaml LLM fields, /model command, and llm/models.py remain for
Task 7. Report: .superpowers/sdd/briefs/task-6-report.md
2026-07-21 12:54:35 +08:00
m4 b2660fc38c feat(model-registry): add provider test API and guarded verification recording
POST /api/model-registry/test (model_config:test) runs the section 9.4
flow: resolve_for_test, per-test credential resolution, build_chat_model
with both safe clients, one minimal chat call, and per-capability probes
(tools/structured_output/vision) whose failures only mark that capability
unverified. Results upsert the model_verifications five-tuple inside a
BEGIN IMMEDIATE transaction that re-checks the registry revision and
configuration hash, returning 409 MODEL_CONFIGURATION_CHANGED on any
concurrent change. resolve_for_test now also relaxes the passing-
verification gate, which the provider test itself produces.
effective_request_options reuses the redacted adapter.build_request
output; the OpenAPI contract and checked-in openapi.json are updated.
2026-07-21 11:39:43 +08:00
m4 940db565b3 fix(model-registry): close review gaps in save-time validation
- add the missing rule-8 counterexample test: declared capabilities
  exceeding the adapter protocol are rejected with
  CAPABILITY_UNSUPPORTED_BY_ADAPTER (all eight section 9.2 checks now
  have at least one negative test)
- raise CREDENTIAL_NOT_CONFIGURED explicitly in _check_enabled_model
  when a required credential reference is null instead of relying on
  resolve_parameters call ordering
2026-07-21 09:38:37 +08:00
m4 dbb6b7abde feat(model-registry): add delegation-JWT auth and config/snapshot HTTP API
- BFF service token (constant-time, plaintext or SHA-256 hash) plus
  X-Evo-Actor delegation JWT verification (ES256/RS256, iss/aud, <=60s
  lifetime, required claims, thread binding) with atomic jti anti-replay
- Config API: GET/PUT /api/model-registry, credential rotation endpoint,
  GET /api/models selector; PUT runs the section 9.2 save-time checks
  inside the registry write transaction after credential writes
- Snapshot API: create/bind/delete routes delegating to SnapshotService
  with thread/deployment binding checks and 9.5 unified error payloads
- Platform security config loader (config.yaml fields), OpenAPI export
  (scripts/export_model_registry_schema.py -> model_registry/openapi.json)
- Mount new routes in langgraph_dev/http.py; retire the legacy
  GET /api/models and POST /api/runtime-snapshots handlers
- Declare PyJWT>=2.8 (previously transitive); extend the 9.5 error code
  table with the HTTP-layer codes (400/401/403/422/500)
2026-07-21 09:28:55 +08:00
m4 c8c46eab16 fix(model-registry): make snapshot abort atomic against concurrent bind
The abort path was read-then-write with an unconditional UPDATE, so a bind
committing between the two calls was clobbered back to aborted, losing its
langgraph_run_id. Add a conditional store-level abort_run_snapshot
(prepared-only UPDATE, rowcount-checked) and re-read on a lost race, matching
the bind loop. Also pin the inherit selection_hash test to a hardcoded
SHA-256 literal instead of reimplementing the serialization in the test.
2026-07-21 08:45:33 +08:00
m4 0cc995eb80 docs(model-registry): add task 4 resolver and snapshot service report 2026-07-21 08:34:50 +08:00
m4 b1233d42dc feat(model-registry): add ModelRegistryResolver and run snapshot service
Resolver (8.1): validates provider/model/credential/capability/limits and
the 6.5 four-mode input budget, freezes ResolvedModelConfig; resolve_for_test
relaxes only the enabled-visibility check (9.4); compute_availability is the
single 4.3 six-state judgement (stale beats configured, selectable only when
enabled).

SnapshotService (8.2, shared by the Task 5 HTTP API and Task 7 local entry):
freezes both roles' full ResolvedModelConfig with adapter spec revision,
fixed reserves, capabilities, and credential revisions; selection-hash
idempotency with pre-resolution semantics; prepared(15min)/bound(+24h)/
expired/aborted lifecycle with atomic bind; binding-checked reads that
revalidate frozen spec revisions; per-call credential resolution against the
frozen revision with no in-process secret cache (5.2); public diagnostic
view limited to the 8.2 safe subset.

Store gains additive helpers (credential pointer lookup, verification
listing, active-triplet lookup, conditional bind, due-expiry sweep) and the
taxonomy gains SNAPSHOT_NOT_FOUND (404) for missing snapshots.
2026-07-21 08:33:12 +08:00
m4 b2e28249fd fix(model-registry): inject async safe clients and split ollama transports
Review fixes for the Task 3 contract layer:

- build_chat_model now accepts http_async_client alongside http_client
  (at least one required) and wires it into ChatOpenAI
  (http_async_client), ChatAnthropic (seeded _async_client), and
  ChatOllama (async_client_kwargs transport), closing the unsafe
  default-async-client gap.
- ChatOllama safe transports move from the shared client_kwargs to
  sync_client_kwargs/async_client_kwargs; langchain-ollama merges shared
  kwargs into both clients, which poisoned the async client with a sync
  transport and crashed ainvoke.
- Unsupported parameters now actually execute the contract-declared
  normalizer (reject_non_auto) instead of a hardcoded raise, with a
  fallback rejection if a normalizer would let a value through.
- build_chat_model rejects overlapping client_options/request_options
  keys instead of silently overwriting.
2026-07-20 22:40:33 +08:00
m4 af4ae1aef5 feat(model-registry): add adapter parameter contracts and build_chat_model factory
Add the Task 3 parameter contract layer (design doc 6.1-6.4):

- adapters.py: versioned built-in contracts for the five phase-1
  adapters plus the openai-compatible/glm-5.2 model-specific contract
  (verbatim section 6.2 values); exact > longest glob > generic
  matching with spec_revision pinning; resolve_parameters implementing
  the section 6.1 inherit/omit semantics, contract validation with
  stable error codes, and named normalizers (identity,
  clamp_to_model_limit, omit_when_none, omit_when_auto,
  reject_non_auto); Adapter.build_request as the single entry point
  mapping ResolvedModelConfig to {client_options, request_options};
  compute_effective_capabilities (protocol AND declared AND verified).
- factory.py: build_chat_model(resolved_config, http_client, *,
  credential=None) with no **kwargs and no setdefault merging; injects
  the safe HTTP client into ChatOpenAI/ChatAnthropic/ChatOllama, never
  reads provider API-key environment variables, and strips the
  OLLAMA_API_KEY authorization header for mode=none adapters.
- tests: per-adapter request-capturing fakes plus an httpx.MockTransport
  outbound capture proving registry resolution matches the wire request.
2026-07-20 22:17:11 +08:00
m4 c46ae17084 feat(model-registry): add EndpointPolicy and SafeHttpTransport SSRF defenses
EndpointPolicy validates provider base URLs (section 4.3): public https
endpoints with hostname and optional port pass; loopback, private,
link-local, multicast, unspecified, and cloud-metadata addresses are
denied unless the normalized URL exactly matches a registered
development_endpoints entry (no prefix or wildcard matching). URLs with
user info, fragments, or non-http(s) schemes are rejected with the new
stable 422 code ENDPOINT_NOT_ALLOWED.

SafeHttpTransport is the single network egress for adapters: a custom
httpcore NetworkBackend resolves DNS under control on every connect
(retries included), filters denied ranges, and connects directly to the
selected IP, while TLS SNI/certificate checks and the HTTP Host header
keep the original hostname. Redirects and env proxies are disabled;
every request origin re-passes URL-layer validation before any I/O.
2026-07-20 21:25:15 +08:00
m4 c21fc0a272 feat(model-registry): add RegistryV4 schema, SQLite store, and unified error codes
New independent subpackage EvoScientist/model_registry implementing the
frozen unified model configuration design (v1.1.0, sections 4.2, 4.3,
5.1, 8.2, 9.5):

- schemas.py: single Pydantic v2 RegistryV4 schema (ModelRef, ProviderConfig,
  ModelConfig, AuthConfig) plus AdapterParameterSpec/AuthSpec/ParameterRule,
  ModelAvailability, ResolvedModelConfig (no secrets), CredentialStatus
- errors.py: all 23 stable error codes from the 9.5 table with HTTP status
  mapping and the unified {code, message, details, request_id} payload
- store.py: ModelRuntimeStore over model-runtime.sqlite3 (0700 dir, 0600
  file, WAL, foreign keys, busy_timeout) with BEGIN IMMEDIATE revision CAS,
  bootstrap->active atomic transition, immutable credential versions with
  masked status, model_verifications upsert, run_runtime_snapshots partial
  unique index, delegation_jtis, and shared-storage lock probe
- hashing.py: normalized SHA-256 configuration_hash

No existing module behavior changed. 89 new tests; full suite passes
(2922 passed, 10 skipped).
2026-07-20 20:43:28 +08:00
m4 8a0ab17936 chore: baseline WIP before unified model configuration implementation
Pre-existing uncommitted work (runtime snapshots, message budget middleware) preserved as baseline.
2026-07-20 20:15:38 +08:00
m4 38668c4ce5 feat: add workspace isolation and provider administration 2026-07-19 12:17:18 +08:00
m4 7a3fcc7c8e Merge remote-tracking branch 'upstream/main'
# Conflicts:
#	README.md
#	uv.lock
2026-07-13 09:46:07 +08:00