Files
EvoScientist/docs/execution-engine/EvoScientist-k8s-1.6.md
T
m4 c2743251e9 Initial commit of EvoScientist framework
Self-evolving AI scientist framework built on LangGraph/LangChain with
CLI/TUI core, FastAPI gateway, and Next.js frontend.

Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
2026-07-13 08:07:45 +08:00

56 KiB
Raw Blame History

EvoScientist k8s 执行引擎 v1.6 实施方案

文件:~/Projects/EvoSci/EvoScientist/docs/EvoScientist-k8s-1.6.md 依据:~/Projects/EvoSci/EvoScientist/docs/execution-engine-development-1.6.md 目标仓库:~/Projects/EvoSci/EvoScientist 关联独立仓库:/Users/m4/Projects/EvoSci/MCP/k8s-exec-service-mcp 状态:开发就绪版实施方案(补齐评审意见)

评审状态:已完成多轮自评审修订

  • R1:修复路径、env 样例损坏、placeholder、shared_compute 接口兼容、StorageService/ComputeClient/计费/poll_loop 细节。
  • R2:补齐迁移顺序、依赖矩阵、测试 Gate 和硬性禁止项,确认可按 Phase 拆 PR 实施。

0. 执行结论

EvoScientist 侧必须从“Gateway 本地多根文件系统 + 本地 shell backend”迁移到“StorageService + S3/MinIO 对象存储 + ComputeClient + ContainerSandboxBackend”的远程执行架构。

v1.6 的核心约束:

  1. Web 用户代码执行必须通过 k8s-exec-service-mcp。
  2. Gateway 不允许执行 Web 用户 shell / Python / 本地 subprocess。
  3. MultiRootSandboxBackend 只能用于非计算文件展示、降级恢复和 fallback_sync,不得作为 Web compute backend。
  4. 用户文件、目录树、权限、生命周期和 artifact 元数据的事实来源是 EvoScientist StorageService + Gateway PostgreSQL。
  5. 文件 bytes 的事实来源是 S3-compatible object storage,本地 P0 使用 MinIO。
  6. k8s-exec-service-mcp 只管理运行期 workspace、调度、usage 和 artifact upload,不保存长期用户文件。
  7. shared_compute 是执行服务侧的一等执行模式;EvoScientist 只传 execution_mode / limits / input_objects / artifact_prefix,不在 Gateway 中实现共享 worker。

本方案补齐以下评审指出的 P0 缺口:

  • StorageService 详细算法与路径规范
  • user_files namespace 唯一性修正
  • ComputeClient 13 个 endpoint wrapper 契约
  • HMAC canonical_json / signing 伪代码
  • StreamHandler async 改造决策
  • deepagents BackendProtocol 映射
  • compute billing 原子事务规格
  • pending_usage poll_loop 详细流程
  • shared_compute API 兼容决策
  • 本地部署和测试验收细化

1. 当前代码冲突点

已确认当前 EvoScientist 代码与 v1.6 目标存在以下冲突:

文件 当前行为 v1.6 要求 修改方向
gateway/main.py lifespan 只初始化 DB、限流、pricing、知识库任务 初始化 StorageService、ComputeClient、poll_loop、quota_loop 增加 compute/storage startup/shutdown
gateway/services/stream_handler.py _create_agent() 同步创建本地 Web agent async preflight、input manifest、remote backend、destroy-settle 改造为 async create + finally cleanup
gateway/routes/threads.py StreamHandler 未传 app_state 传入 storage/compute/db/admin token 注入 request.app.state._state
EvoScientist/EvoScientist.py:create_cli_agent() Web 默认 MultiRootSandboxBackend Web compute 必须 ContainerSandboxBackend 新增 compute 参数和 backend 选择
EvoScientist/backends.py 缺 ContainerSandboxBackend 实现远程 BackendProtocol 增加 remote backend 和 non-compute fallback
gateway/routes/uploads.py 直接写本地文件 通过 StorageService 写 metadata + object 改 StorageService
gateway/routes/user_files.py / files.py / global_uploads.py 文件列表/下载依赖本地路径 查询 user_files + StorageService open_read 统一存储入口
gateway/services/knowledge_indexer.py 用 thread_data_dir() 读取文件 通过 StorageService 读取 改文件解析入口
gateway/pricing.json 无 compute_pricing 增 compute 定价和分钟配额 扩展 pricing config
gateway/services/wallet.py / token_tracker.py 金额存在 cents/REAL 语义混用风险 compute 使用 Decimal + NUMERIC(12,6) 增迁移和兼容层
docs/execution-engine-deployment.md 旧 Gateway 内 Execution Engine 方案 独立 k8s-exec-service-mcp 标记过期或重写

2. 目标架构

2.1 调用链

Web POST /api/threads/{thread_id}/stream
  -> gateway/routes/threads.py
  -> StreamHandler(app_state)
  -> await StreamHandler._prepare_compute_context()
       -> Compute quota preflight
       -> StorageService.build_input_manifest()
       -> artifact_prefix/default_output_paths/budget_snapshot
       -> execution_mode decision
  -> await StreamHandler._create_agent()
       -> create_cli_agent(source="web", compute_client=<ComputeClient>, input_objects=<list[InputObject]>)
       -> CompositeBackend
            default: ContainerSandboxBackend
              -> ComputeClient HTTP /tools/*
              -> k8s-exec-service-mcp
              -> ResourceManager
              -> shared_compute 或 independent environment/warm pool
            /skills/: local skills backend
            /memory/: local memory backend
  -> SSE streaming
  -> finally:
       -> await _destroy_and_settle_backend()
       -> backend.destroy(persist_outputs=True)
       -> StorageService.register_artifacts()
       -> settle_compute_usage()
       -> pending_usage poll_loop 兜底

2.2 存储边界

Gateway PostgreSQL
  - user_files metadata
  - compute_usage_log
  - user_balances
  - wallet_ledger
  - fallback_sync_recovery
  - threads/session metadata

Object Storage: MinIO/S3/OSS/COS/OBS/R2
  - users/{user_id}/files/{file_id}
  - projects/{project_id}/files/{file_id}
  - threads/{thread_id}/files/{file_id}
  - jobs/{job_id}/artifacts/{artifact_id}
  - tmp/{job_id}/{object_id}

k8s-exec-service-mcp SQLite WAL
  - exec_environments
  - exec_pending_usage
  - exec_backends
  - exec_warm_pools
  - exec_auth_nonces
  - exec_admin_audit_log

边界规则:

  • Gateway 不读写 exec service SQLite。
  • exec service 不读写 Gateway PostgreSQL。
  • object_key 不是权限依据;权限来自 Gateway metadata + HMAC 签名的 object_scope。
  • Web remote compute P0 必须使用 S3StorageBackend;LocalStorageBackend 不得生成 compute input manifest。

3. StorageService 详细实施规格

3.1 新增文件

gateway/services/storage_service.py
gateway/services/storage_backends/__init__.py
gateway/services/storage_backends/local.py
gateway/services/storage_backends/s3.py
gateway/services/storage_errors.py
scripts/migrate_local_files_to_storage.py

3.2 数据模型

@dataclass(slots=True)
class FileRecord:
    file_id: str
    owner_user_id: str
    logical_path: str
    object_key: str
    storage_backend: Literal["local", "s3"]
    sha256: str
    size_bytes: int
    thread_id: str | None = None
    project_id: str | None = None
    user_id_int: int | None = None
    mime_type: str | None = None
    source: Literal["upload", "agent", "artifact", "import"] = "upload"
    version: int = 1
    status: Literal["active", "deleted", "pending", "orphan"] = "active"
    created_at: datetime | None = None
    updated_at: datetime | None = None
    expires_at: datetime | None = None
    metadata_json: dict[str, Any] | None = None

@dataclass(slots=True)
class InputObject:
    file_id: str
    object_key: str
    sha256: str
    size_bytes: int
    mount_path: str
    read_only: bool

@dataclass(slots=True)
class ArtifactObject:
    artifact_id: str
    object_key: str
    sha256: str
    size_bytes: int
    logical_path: str
    metadata: dict[str, Any] | None = None

3.3 路径规范化算法

所有进入 StorageService 的 logical_path 必须统一为“相对 POSIX 路径”,不以 / 开头。

def normalize_logical_path(path: str) -> str:
    if not isinstance(path, str):
        raise InvalidPath("path must be string")
    path = unicodedata.normalize("NFC", path.strip())
    if not path:
        raise InvalidPath("empty path")
    if "\x00" in path or any(ord(c) < 32 for c in path):
        raise InvalidPath("control character")
    path = path.replace("\\", "/")
    if re.match(r"^[A-Za-z]:", path):
        raise InvalidPath("windows drive prefix")
    path = urllib.parse.unquote(path)
    path = path.lstrip("/")
    parts = [p for p in path.split("/") if p]
    if any(p in (".", "..") for p in parts):
        raise InvalidPath("path traversal")
    if not parts:
        raise InvalidPath("empty normalized path")
    if parts[0] in (".control", ".secrets", "__evosci_internal__"):
        raise InvalidPath("reserved prefix")
    return "/".join(parts)

mount_path 规则:

  • /workspace/{logical_path}:当前线程读写。
  • /__global__/{logical_path}:用户全局读写。
  • /threads/{thread_id}/{logical_path}:其他线程只读。
  • 禁止 mount_path 冲突、父子覆盖、读写权限冲突。

3.4 user_files namespace 唯一性修正

不要使用单一 (owner_user_id, logical_path) active 唯一索引,否则不同 thread 下同名文件冲突。

使用三个 partial unique index:

CREATE UNIQUE INDEX IF NOT EXISTS uq_user_files_global_active_path
ON user_files(owner_user_id, logical_path)
WHERE status = 'active' AND thread_id IS NULL AND project_id IS NULL;

CREATE UNIQUE INDEX IF NOT EXISTS uq_user_files_thread_active_path
ON user_files(owner_user_id, thread_id, logical_path)
WHERE status = 'active' AND thread_id IS NOT NULL;

CREATE UNIQUE INDEX IF NOT EXISTS uq_user_files_project_active_path
ON user_files(owner_user_id, project_id, logical_path)
WHERE status = 'active' AND project_id IS NOT NULL;

对象 key 生成:

namespace object_key
user global users/{user_id}/files/{file_id}
thread threads/{thread_id}/files/{file_id}
project projects/{project_id}/files/{file_id}
artifact jobs/{job_id}/artifacts/{artifact_id}
temp tmp/{job_id}/{object_id}

3.5 StorageService 接口

class StorageService:
    def __init__(self, db, backend: StorageBackend, *, backend_name: str): pass

    async def put_file(
        self,
        owner_user_id: str,
        logical_path: str,
        data: bytes,
        *,
        thread_id: str | None = None,
        project_id: str | None = None,
        source: str = "upload",
        mime_type: str | None = None,
        expires_at: datetime | None = None,
        metadata: dict[str, Any] | None = None,
    ) -> FileRecord: pass

    async def get_file(self, file_id: str, requester_user_id: str) -> FileRecord: pass
    async def open_read(self, file_id: str, requester_user_id: str) -> AsyncIterator[bytes]: pass
    async def delete_file(self, file_id: str, requester_user_id: str) -> None: pass

    async def build_input_manifest(
        self,
        user_id: str,
        thread_id: str,
        paths: list[str],
        *,
        require_remote_compute: bool = True,
    ) -> list[InputObject]: pass

    async def register_artifacts(
        self,
        user_id: str,
        thread_id: str,
        artifacts: list[ArtifactObject],
        *,
        environment_id: str | None = None,
        job_id: str | None = None,
    ) -> list[FileRecord]: pass

    async def cleanup_expired(self, *, limit: int = 100) -> int: pass
    async def close(self) -> None: pass

3.6 put_file 事务边界

顺序固定:

  1. normalize logical_path。
  2. 计算 sha256 / size。
  3. 生成 file_id。
  4. 生成 object_key。
  5. 上传 object storage。
  6. HeadObject 校验 size/sha256 metadata。
  7. 在 DB 事务中:
    • 将同 namespace + logical_path 的旧 active 记录置为 deleted。
    • 插入新 user_files 记录,version=previous+1。
  8. 若 DB 写失败:
    • 尝试删除刚上传 object。
    • 删除失败时写 orphan metadata 或 ERROR 日志,由 cleanup 后台任务处理。

不得先写 DB 再上传 object,避免 metadata 指向不存在对象。

3.7 build_input_manifest 算法

输入:

  • user_id:当前 Web 用户 UID。
  • thread_id:当前线程 ID。
  • paths:允许根路径,P0 默认由 StreamHandler 传:
    • /workspace
    • /__global__
    • 显式 /threads/{other_thread_id} 列表

禁止传 /threads/* 作为实现级路径。/threads/* 只能是文档描述,实际必须展开为当前用户可读线程列表。

算法:

1. 若 require_remote_compute=True 且 backend != s3:
   抛 StorageBackendUnsupportedForCompute。

2. 初始化 input_objects=[]。

3. 对 paths 逐个处理:
   a. /workspace:
      查询 user_files
      WHERE owner_user_id=$user_id
        AND thread_id=$thread_id
        AND status='active'
      mount_path = /workspace/{logical_path}
      read_only = false

   b. /__global__:
      查询 user_files
      WHERE owner_user_id=$user_id
        AND thread_id IS NULL
        AND project_id IS NULL
        AND status='active'
      mount_path = /__global__/{logical_path}
      read_only = false

   c. /threads/{other_thread_id}:
      先校验 requester 对 other_thread_id 有读权限。
      查询 user_files
      WHERE thread_id=$other_thread_id
        AND status='active'
      mount_path = /threads/{other_thread_id}/{logical_path}
      read_only = true

   d. 其他路径:抛 InvalidManifestPath。

4. 对每个 FileRecord:
   - object_key 必须非空。
   - storage_backend 必须为 s3。
   - S3 HeadObject 必须成功。
   - size/sha256 与 metadata 一致。
   - object_key 必须位于允许 prefix,但 prefix 不作为权限依据。

5. 检查 mount_path:
   - UTF-8 NFC。
   - 禁止 NUL/control。
   - 必须位于 /workspace、/__global__、/threads/{id}。
   - 禁止 ..。
   - 禁止重复。
   - 禁止父子覆盖冲突。
   - 禁止 read_only=false 写入 /threads/*。

6. 按 mount_path 排序。

7. 返回 input_objects。

失败策略:

  • P0 所有 active user_files 都是关键输入。
  • 任一 object missing/corrupt -> 抛 InputObjectMissing 或 InputObjectCorrupt。
  • 不静默跳过。

3.8 register_artifacts 幂等规则

artifact 入库幂等键:

优先级:

  1. metadata.environment_id + artifact.logical_path
  2. metadata.job_id + artifact.logical_path
  3. artifact.object_key

算法:

  1. 校验 artifact.object_key 必须在 jobs/{thread_id}/{request_id}/artifacts/ 或 Gateway 下发的 artifact_prefix 内。
  2. HeadObject 校验 sha256 / size。
  3. logical_path normalize;若 artifact.logical_path 以 /workspace/ 开头,转为相对路径。
  4. DB 事务中 upsert user_files:
    • source='artifact'
    • thread_id=current thread
    • metadata_json 包含 environment_id/job_id/object_key/original_path。
    • 若相同 artifact 幂等键已存在,直接返回旧记录。
  5. 调用 ComputeClient.mark_artifacts_registered() 由上层 StreamHandler/poll_loop 完成,不在 StorageService 内直接调用 ComputeClient。

4. ComputeClient 实施规格

4.1 新增文件

gateway/services/compute_client.py
gateway/services/compute_errors.py
tests/test_compute_client.py

4.2 配置

EXECUTION_SERVICE_URL=http://127.0.0.1:9020
EXECUTION_HMAC_SECRET=<generated-32-byte-hex>
EXECUTION_ADMIN_TOKEN=<generated-admin-token>
EXECUTION_HTTP_TIMEOUT_SECONDS=30
EXECUTION_CREATE_TIMEOUT_SECONDS=210
EXECUTION_READY_TIMEOUT_SECONDS=5

配置矩阵:

EXECUTION_SERVICE_URL HMAC/Admin Gateway 行为
未设置 可缺失 compute disabled,正常启动,不启动 poll_loop
已设置 任一缺失 startup fail
已设置 都存在,/ready false 正常启动,但 compute_client.ready=false
已设置 都存在,/ready true compute enabled,启动 poll_loop

4.3 HMAC signing 规范

每个 /tools/* 请求 body 顶层包含 auth_context。签名前移除 auth_context.token 字段。

canonical string:

{METHOD}\n{PATH}\n{issued_at}\n{nonce}\n{sha256(canonical_json(body_without_auth_context.token))}

实现伪代码:

def canonical_json(obj: Any) -> str:
    return json.dumps(
        obj,
        ensure_ascii=False,
        sort_keys=True,
        separators=(",", ":"),
    )

def payload_hash(body_without_token: dict[str, Any]) -> str:
    data = canonical_json(body_without_token).encode("utf-8")
    return hashlib.sha256(data).hexdigest()

def canonical_string(method: str, path: str, issued_at: int, nonce: str, body_hash: str) -> str:
    return "\n".join([method.upper(), path, str(issued_at), nonce, body_hash])

def sign(secret: str, canonical: str) -> str:
    return hmac.new(
        secret.encode("utf-8"),
        canonical.encode("utf-8"),
        hashlib.sha256,
    ).hexdigest()

def build_auth_context(method: str, path: str, body_params: dict[str, Any], *, user_id: str, thread_id: str, object_scope: dict | None):
    issued_at = int(time.time())
    nonce = secrets.token_hex(16)
    auth_context = {
        "user_id": user_id,
        "thread_id": thread_id,
        "issued_at": issued_at,
        "nonce": nonce,
    }
    if object_scope:
        auth_context["object_scope"] = object_scope
    body_without_token = {"auth_context": auth_context, **body_params}
    digest = payload_hash(body_without_token)
    canonical = canonical_string(method, path, issued_at, nonce, digest)
    token = sign(self.hmac_secret, canonical)
    auth_context["token"] = token
    return {"auth_context": auth_context, **body_params}

要求:

  • path 必须是不含 scheme/host/query 的路径,如 /tools/execute_command。
  • canonical_json 必须递归排序 key;json.dumps(sort_keys=True) 对 Python dict 嵌套满足要求。
  • separators 必须是 (',', ':'),无空格。
  • canonical string 无 trailing LF,总共 4 个 LF。
  • token lowercase hex。
  • 日志禁止打印 secret/token/body 全量。

4.4 HMAC golden vector 测试

Gateway 与 k8s-exec-service-mcp 均必须包含同一测试向量。若源规格中的 expected hash/token 与实际算法不一致,以测试脚本真实计算结果为准并同步更新两仓;不得两侧各自写死不同 token。

测试断言:

  • canonical_json exact string。
  • payload_hash exact hex。
  • canonical_string repr() exact。
  • token exact hex。
  • body 增加/删除任一字段 token 变化。
  • CRLF 与 LF 不同,CRLF 必须失败。

固定测试向量(按上述算法真实计算):

method = POST
path = /tools/execute_command
secret = test-hmac-secret-for-unit-tests-only
canonical_json = {"auth_context":{"issued_at":1716500000,"nonce":"abc123def456","object_scope":{"read":["objects/in/data.csv"],"write_prefixes":["jobs/job_1/artifacts/"]},"user_id":"test_user"},"command":"echo hello","environment_id":"env_test123"}
payload_hash = 45957f2078d85c7c09bedfea3f000b6a42a8e01e039250644bb577321186035b
canonical_string_repr = 'POST\n/tools/execute_command\n1716500000\nabc123def456\n45957f2078d85c7c09bedfea3f000b6a42a8e01e039250644bb577321186035b'
token = 06af87716e9c3ba83441362dad2faa543c5db19d1c47fb937cd2ffe29cee1d21

4.5 ComputeClient public methods

class ComputeClient:
    async def check_ready(self) -> bool: pass
    async def check_health(self) -> dict: pass
    async def check_poll_available(self) -> bool: pass
    async def close(self) -> None: pass

    async def create_environment(
        self,
        *,
        user_id: str,
        thread_id: str,
        resource_class: str,
        backend_policy: str,
        execution_mode: Literal["simple", "shared", "standard", "isolated"],
        input_objects: list[InputObject],
        artifact_prefix: str,
        default_output_paths: list[str],
        max_runtime_seconds: int,
        budget_snapshot: dict[str, Any],
        limits: dict[str, Any] | None = None,
        image_ref: str | None = None,
    ) -> CreateEnvironmentResult: pass

    async def destroy_environment(
        self,
        *,
        user_id: str,
        thread_id: str,
        environment_id: str,
        persist_outputs: bool,
        output_paths: list[str],
        force_destroy: bool = False,
        artifact_prefix: str | None = None,
    ) -> DestroyEnvironmentResult: pass

    async def list_environments(self, *, user_id: str, thread_id: str) -> list[dict]: pass
    async def get_environment_status(self, *, user_id: str, thread_id: str, environment_id: str) -> dict: pass
    async def execute_command(self, *, user_id: str, thread_id: str, environment_id: str, command: str, timeout_seconds: int | None = None, max_output_bytes: int | None = None) -> dict: pass
    async def read_file(self, *, user_id: str, thread_id: str, environment_id: str, path: str) -> dict: pass
    async def write_file(self, *, user_id: str, thread_id: str, environment_id: str, path: str, content: str) -> dict: pass
    async def edit_file(self, *, user_id: str, thread_id: str, environment_id: str, path: str, old_string: str, new_string: str, replace_all: bool = False) -> dict: pass
    async def list_dir(self, *, user_id: str, thread_id: str, environment_id: str, path: str) -> dict: pass
    async def grep_files(self, *, user_id: str, thread_id: str, environment_id: str, pattern: str, path: str, glob: str | None = None) -> dict: pass
    async def glob_files(self, *, user_id: str, thread_id: str, environment_id: str, pattern: str, path: str = "/workspace") -> dict: pass
    async def upload_file(self, *, user_id: str, thread_id: str, environment_id: str, remote_path: str, content_base64: str, sha256: str) -> dict: pass
    async def download_file(self, *, user_id: str, thread_id: str, environment_id: str, remote_path: str) -> dict: pass

    async def list_pending_usage(self, admin_token: str, *, settled: int = 0, limit: int = 50) -> list[dict]: pass
    async def mark_usage_settled(self, admin_token: str, usage_ids: list[int]) -> None: pass
    async def mark_artifacts_registered(self, admin_token: str, *, environment_id: str | None = None, usage_id: int | None = None) -> None: pass

4.6 Endpoint mapping

Method Endpoint Auth Timeout Notes
GET /ready none 5s readiness only
GET /health none 5s liveness only
POST /tools/create_environment HMAC 210s create or bind warm/shared handle
POST /tools/destroy_environment HMAC 60s persist artifacts first
POST /tools/list_environments HMAC 30s user visible only
POST /tools/get_environment_status HMAC 30s runtime state
POST /tools/execute_command HMAC command timeout + 5s no local fallback
POST /tools/read_file HMAC 30s P0 size limits
POST /tools/write_file HMAC 30s P0 size limits
POST /tools/edit_file HMAC 30s old_string unique
POST /tools/list_dir HMAC 30s max entries
POST /tools/grep_files HMAC 30s max matches
POST /tools/glob_files HMAC 30s max matches
POST /tools/upload_file HMAC 60s base64 only
POST /tools/download_file HMAC 60s base64 only
POST /admin/pending-usage X-Admin-Token 30s poll_loop
POST /admin/mark-settled X-Admin-Token 30s idempotent
POST /admin/mark-artifacts-registered X-Admin-Token 30s idempotent

4.7 HTTP status -> exception mapping

HTTP error code ComputeClient exception
400 validation ComputeBadRequest
401 auth invalid/expired/replay ComputeAuthError
403 ownership/object scope ComputePermissionDenied
404 env missing EnvironmentNotFound
409 input/storage conflict InputObjectError / StorageBackendUnsupportedForCompute
413 size limit ComputePayloadTooLarge
429 quota/concurrency ComputeCapacityExceeded
500 internal ComputeServiceError
503 unavailable ComputeUnavailable
507 artifact/storage quota ArtifactStorageExceeded

ComputeUnavailable 必须设置 compute_client.ready=False,但仅对明确的 readiness/backend unavailable/network timeout 生效;普通命令失败不得把全局 ready 置 false。


5. shared_compute API 兼容决策

P0 不新增 create_job endpoint。EvoScientist 侧统一调用 /tools/create_environment,通过字段表达 shared intent:

{
  "execution_mode": "simple",
  "shared_compute_eligible": true,
  "resource_class": "shared-small",
  "backend_policy": "auto",
  "input_objects": [
    {
      "file_id": "file_123",
      "object_key": "threads/thr_1/files/file_123",
      "sha256": "<sha256>",
      "size_bytes": 1024,
      "mount_path": "/workspace/data.csv",
      "read_only": false
    }
  ],
  "artifact_prefix": "jobs/{thread_id}/{request_id}/artifacts/",
  "default_output_paths": ["/workspace"],
  "max_runtime_seconds": 60,
  "budget_snapshot": {
    "plan": "pro",
    "quota_ok": true,
    "remaining_minutes": 120,
    "cash_balance_snapshot": "10.000000",
    "overtime_price_per_minute": "0.050000",
    "affordable_runtime_seconds": 19200,
    "effective_max_runtime_seconds": 60,
    "pricing_version": "2026-05-23"
  },
  "limits": {
    "max_input_bytes": 10485760,
    "max_output_bytes": 10485760,
    "max_stdout_stderr_bytes": 1048576,
    "max_files": 100
  }
}

服务端返回仍必须包含 environment_id 兼容字段:

{
  "environment_id": "env_or_job_handle_123",
  "status": "running",
  "execution_mode": "simple",
  "backend_type": "shared_compute",
  "resource_ref": "shared-worker-pool/default"
}

约束:

  • 对 EvoScientist 和 deepagents 来说,该 handle 仍命名为 environment_id。
  • 对 exec service 内部来说,可映射为 shared job handle。
  • 后续 execute/read/write/destroy 仍传同一个 environment_id。
  • shared_compute 不进入 warm pool 状态;这是 service 内部语义,Gateway 不直接判断。
  • 若 service 判定 not eligible,可以:
    1. 返回 SHARED_COMPUTE_NOT_ELIGIBLE,Gateway 重新 standard create;或
    2. service 内部自动路由到 standard,并在 response 标记 execution_mode="standard"。
  • P0 推荐 Gateway 先保守判断,不满足条件直接 standard,减少二次请求。

6. deepagents BackendProtocol 映射

当前 EvoScientist/backends.py 使用 deepagents 返回类型:

  • ExecuteResponse
  • WriteResult
  • EditResult
  • LsResult
  • GrepResult
  • GlobResult
  • FileUploadResponse
  • FileDownloadResponse

ContainerSandboxBackend 必须适配这些类型,不返回裸 dict 给 deepagents。

deepagents backend method ComputeClient method 返回映射
execute(command, timeout=None) execute_command ExecuteResponse(output=stdout+stderr, exit_code=exit_code, truncated=truncated)
read(file_path) read_file 返回 content 字符串或错误包装
write(file_path, content) write_file WriteResult(error=None)
edit(file_path, old_string, new_string, replace_all=False) edit_file EditResult(diff=<unified_diff>, error=None)
ls(path) list_dir LsResult(entries=<list>)
grep(pattern, path, glob=None) grep_files GrepResult(matches=<list>)
glob(pattern, path="/") glob_files GlobResult(matches=<list>)
upload_files(files) upload_file loop list[FileUploadResponse]
download_files(paths) download_file loop list[FileDownloadResponse]
destroy(persist_outputs=True, output_paths=<list>, force_destroy=False) destroy_environment dict for StreamHandler settlement

实现注意:

  • execute output 可先采用 stdout + stderr 合并,metadata 中保留 stderr/exit_code/truncated。
  • ComputeClient 是 async,但 deepagents BackendProtocol 当前是 sync。P0 需要选择一个明确适配策略。

6.1 async/sync 适配决策

推荐方案:将 StreamHandler._create_agent() 改为 async,但 ContainerSandboxBackend 内部仍需要 sync methods 供 deepagents 调用。

采用 anyio.from_thread / asyncio.run_coroutine_threadsafe 复杂度高。P0 更稳妥方案:

  1. 为 ContainerSandboxBackend 创建一个专用 background event loop thread。
  2. backend sync method 调用 _run_async(coro)。
  3. _run_async 使用 asyncio.run_coroutine_threadsafe(coro, self._loop).result(timeout=<seconds>)。
  4. destroy 时关闭 loop thread。

伪代码:

class ContainerSandboxBackend:
    def _run_async(self, coro, timeout: float | None = None):
        if self._closed:
            raise ComputeUnavailable("backend closed")
        fut = asyncio.run_coroutine_threadsafe(coro, self._loop)
        return fut.result(timeout=timeout or self.default_timeout + 5)

如果 deepagents 后续支持 async backend,可在 P1 移除 loop thread。

6.2 路径翻译

ContainerSandboxBackend 不再把 /workspace/foo 转成本地路径。它只做执行环境内路径规范化:

  • relative foo.py -> /workspace/foo.py
  • /workspace/foo.py -> /workspace/foo.py
  • /__global__/x -> /__global__/x
  • /threads/{id}/x -> /threads/{id}/x
  • 禁止 /tmp、/Users、/etc、.. 等路径。

7. StreamHandler async 改造

7.1 决策

将 _create_agent() 改为 async:

async def _create_agent(self, checkpointer=None): pass

在 stream() 中:

agent = await self._create_agent(checkpointer=checkpointer)

原因:

  • quota preflight 是 DB async。
  • StorageService.build_input_manifest 是 DB + S3 async。
  • ComputeClient readiness/admin check 是 HTTP async。
  • 避免在 running event loop 中使用 blocking hack。

7.2 ComputeContext

新增内部 dataclass:

@dataclass
class ComputeContext:
    enabled: bool
    reason: str | None
    input_objects: list[InputObject]
    artifact_prefix: str
    default_output_paths: list[str]
    max_runtime_seconds: int
    budget_snapshot: dict[str, Any]
    execution_mode: str
    resource_class: str
    backend_policy: str
    quota_result: ComputeQuotaResult | None

7.3 _prepare_compute_context 流程

1. 若 source != web:返回 disabled,但 CLI 不受影响。
2. 若 compute_client is None 或 not ready:返回 disabled reason=ComputeUnavailable。
3. 执行 _check_compute_quota_async。
4. quota_ok false:返回 disabled reason=ComputeQuotaExceeded。
5. storage_backend != s3:返回 disabled reason=StorageBackendUnsupportedForCompute。
6. 展开 manifest paths:
   - /workspace
   - /__global__
   - 当前用户可读 peer thread 列表,显式 /threads/{id}
7. await storage_service.build_input_manifest(user_id=self.user_uid, thread_id=self.thread_id, paths=manifest_paths, require_remote_compute=True)
8. 生成 request_id。
9. artifact_prefix = jobs/{thread_id}/{request_id}/artifacts/
10. default_output_paths = ["/workspace", "/__global__"] 或按配置。
11. effective_max_runtime_seconds = min(requested, plan_hard_max, affordable_runtime_seconds)
12. 若 effective < 60:disabled reason=ComputeQuotaExceeded。
13. 生成 budget_snapshot。
14. decide_execution_mode()。
15. 返回 enabled ComputeContext。

7.4 _create_agent 调用

backend_ref: dict[str, Any] = {}
compute_ctx = await self._prepare_compute_context()
agent = create_cli_agent(
    workspace_dir=self.workspace_dir,
    memory_dir=self.memory_dir,
    source="web",
    user_id=self.user_uid,
    thread_id=self.thread_id,
    compute_quota_ok=compute_ctx.enabled,
    remaining_compute_minutes=compute_ctx.quota_result.remaining_minutes if compute_ctx.quota_result else 0,
    compute_client=self.compute_client if compute_ctx.enabled else None,
    input_objects=compute_ctx.input_objects,
    artifact_prefix=compute_ctx.artifact_prefix,
    default_output_paths=compute_ctx.default_output_paths,
    max_runtime_seconds=compute_ctx.max_runtime_seconds,
    budget_snapshot=compute_ctx.budget_snapshot,
    storage_backend=self.storage_service.backend_name,
    execution_mode=compute_ctx.execution_mode,
    resource_class=compute_ctx.resource_class,
    backend_policy=compute_ctx.backend_policy,
    _backend_ref=backend_ref,
    model=self.model,
    config=cfg,
    checkpointer=checkpointer,
    **filtered_params,
)
self._compute_backend = backend_ref.get("backend")
self._compute_context = compute_ctx

7.5 finally destroy/settle 顺序

1. 若没有 _compute_backend:只做非计算 fallback sync。
2. 调用 backend.destroy(persist_outputs=True, output_paths=default_output_paths, force_destroy=False)。
3. 若 destroy 抛 ComputeUnavailable:记录,依赖 exec service pending_usage;不本地执行。
4. 若返回 STOPPING_FAILED:发送用户可见 warning;不 settle。
5. 若 artifacts 非空:StorageService.register_artifacts。
6. register 成功:ComputeClient.mark_artifacts_registered。
7. register 失败:记录 artifact_registration_error;不 settle,等待 poll_loop。
8. 调用 settle_compute_usage。
9. settle InsufficientBalance:保留 pending_usage。
10. settle success/already_settled:mark_usage_settled 由 poll_loop 或当前路径完成。
11. 最后 set_thread_status idle。

8. create_cli_agent 改造

8.1 签名追加参数

在末尾追加,保持兼容:

thread_id: str = "",
compute_quota_ok: bool = False,
remaining_compute_minutes: int = 0,
compute_client: Any | None = None,
input_objects: list[Any] | None = None,
artifact_prefix: str = "",
default_output_paths: list[str] | None = None,
max_runtime_seconds: int | None = None,
budget_snapshot: dict[str, Any] | None = None,
storage_backend: str = "local",
execution_mode: str = "standard",
resource_class: str = "small",
backend_policy: str = "auto",
_backend_ref: dict | None = None,

8.2 Web backend 选择

if source == "web" and user_id:
    if compute_client and compute_quota_ok and storage_backend == "s3":
        ws_backend = ContainerSandboxBackend(
            compute_client=compute_client,
            user_id=user_id,
            thread_id=thread_id,
            input_objects=input_objects or [],
            artifact_prefix=artifact_prefix,
            default_output_paths=default_output_paths or ["/workspace"],
            max_runtime_seconds=max_runtime_seconds,
            budget_snapshot=budget_snapshot,
            execution_mode=execution_mode,
            resource_class=resource_class,
            backend_policy=backend_policy,
        )
        if _backend_ref is not None:
            _backend_ref["backend"] = ws_backend
    else:
        ws_backend = NonComputingFallbackBackend(
            write_root=str(_thread_root),
            read_roots=_peer_threads,
            global_root=str(_user_global),
            virtual_mode=True,
            reason="ComputeUnavailable or storage backend unsupported",
        )
else:
    ws_backend = CustomSandboxBackend(root_dir=workspace_dir, virtual_mode=True, timeout=300)

CompositeBackend routes 保持:

  • default -> ws_backend
  • /skills/ -> local skills backend
  • /memory/ -> local memory backend

9. 数据库实施规格

9.1 user_balances

ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_minutes_remaining INTEGER DEFAULT 0;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_minutes_used INTEGER DEFAULT 0;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_minutes_lifetime INTEGER DEFAULT 0;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_overtime_charged NUMERIC(12,6) DEFAULT 0;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_quota_monthly INTEGER DEFAULT 0;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_quota_refreshed_at TIMESTAMP NULL;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_quota_period_start TIMESTAMP NULL;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_quota_period_end TIMESTAMP NULL;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_quota_adjustment INTEGER DEFAULT 0;

金额迁移:

  • cash_balance、total_spent、recharge_records.charge_amount 统一 NUMERIC(12,6)。
  • 迁移前查询 information_schema,避免重复 ALTER。
  • Python 读写全部转 Decimal。

9.2 compute_usage_log

CREATE TABLE IF NOT EXISTS compute_usage_log (
    id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
    user_id_int INTEGER REFERENCES users(id),
    user_uid TEXT NOT NULL,
    thread_id TEXT NOT NULL,
    environment_id TEXT UNIQUE NOT NULL,
    plan TEXT NOT NULL,
    runtime_seconds DOUBLE PRECISION NOT NULL,
    compute_minutes INTEGER NOT NULL,
    quota_remaining INTEGER NULL,
    overtime_charged NUMERIC(12,6) DEFAULT 0,
    currency TEXT DEFAULT 'CNY',
    pricing_snapshot TEXT NULL,
    settlement_status TEXT DEFAULT 'pending',
    settlement_error TEXT NULL,
    created_at TIMESTAMP DEFAULT NOW()
);

状态语义:

settlement_status 含义
pending 记录存在但未获得结算权或等待重试
claimed 当前事务获得结算权但尚未完成扣费;正常情况下不应长期存在
settled 扣费、usage_log、wallet_ledger 全部提交
failed 最近一次结算失败,可重试

9.3 user_files

使用第 3.4 节修正后的 namespace unique indexes。

9.4 wallet_ledger 兼容策略

如果现有 wallet_ledger 使用 cents 字段,P0 采用兼容但单一真实来源策略:

  • 余额真实来源仍为 user_balances.cash_balance NUMERIC(12,6)。
  • compute ledger 写入:
    • entry_type='compute_overtime'
    • source_type='compute_usage'
    • source_id=environment_id
    • direction='debit'
    • amount=Decimal overtime_amount
    • 若表只有 amount_cents,则写 amount_cents = int(amount * 100),但 pricing_snapshot 必须保留 Decimal 原值。
  • 不允许同一次 compute 同时写 amount 和 amount_cents 两套互相独立的余额。

10. compute billing 原子事务规格

10.1 preflight quota

quota_ok = (
    plan != "starter"
    and (
        compute_minutes_remaining > 0
        or cash_balance >= overtime_price_per_minute
    )
)

affordable_runtime_seconds = (
    compute_minutes_remaining + floor(cash_balance / overtime_price_per_minute)
) * 60

边界:

  • overtime_price_per_minute <= 0 或 NULL:PricingConfigError,compute enabled startup fail。
  • remaining_minutes < 0 或 cash_balance < 0:记录 ERROR,quota_ok=False。
  • 余额不足 1 分钟:quota_ok=False。
  • max_runtime_seconds 必须裁剪为 min(requested, plan_hard_max, affordable_runtime_seconds)。

10.2 settle_compute_usage 事务

输入:

async def settle_compute_usage(
    *,
    environment_id: str,
    user_uid: str,
    thread_id: str,
    runtime_seconds: float,
    plan: str,
    pricing_snapshot: dict,
    artifacts_registered: bool,
) -> SettlementResult: pass

计费分钟:

compute_minutes = max(1, math.ceil(runtime_seconds / 60))

事务规则:

1. 使用 get_transaction_connection() 获取独立 asyncpg connection。
2. INSERT compute_usage_log(environment_id, user_uid, thread_id, runtime_seconds, compute_minutes, settlement_status='claimed')
   ON CONFLICT(environment_id) DO NOTHING。
3. 若 INSERT 未获得行:
   a. SELECT existing settlement_status。
   b. status='settled' -> return already_settled。
   c. status='claimed' 且 updated/created 很新 -> raise SettlementInProgress。
   d. status='failed' 或长期 claimed -> 可抢占重试,更新为 claimed。
4. 查询 user_balances FOR UPDATE。
5. free_to_use = min(compute_minutes_remaining, compute_minutes)。
6. overtime_minutes = compute_minutes - free_to_use。
7. overtime_amount = Decimal(overtime_minutes) * overtime_price_per_minute。
8. 若 cash_balance < overtime_amount:
   抛 InsufficientComputeBalance,整个事务回滚。
9. UPDATE user_balances:
   compute_minutes_remaining -= free_to_use
   cash_balance -= overtime_amount
   compute_minutes_used += compute_minutes
   compute_minutes_lifetime += compute_minutes
   compute_overtime_charged += overtime_amount
10. INSERT wallet_ledger with idempotency_key compute:{user_uid}:{thread_id}:{environment_id}。
11. UPDATE compute_usage_log SET settlement_status='settled', quota_remaining, overtime_charged, pricing_snapshot。
12. commit 后返回 settled。

禁止:

  • 先插入 claim 并提交,再异步扣款。
  • float 参与金额计算。
  • 余额不足时保留 claimed usage_log。
  • artifact 未注册且未达到 retry 上限时结算。

10.3 artifact 注册失败与结算

规则:

  • pending_usage 中有 output_manifest_json 且 artifacts_registered=false 时,必须先注册 artifact。
  • 注册失败未达到 ARTIFACT_REGISTRATION_MAX_RETRIES=10:不 settle。
  • 达到上限后:标记 artifacts_unregistered=true 或在本地记录 artifact_registration_failed,然后允许 compute 结算,避免无限免单。

11. pending_usage poll_loop 详细规格

11.1 启动条件

  • compute enabled。
  • ComputeClient.ready=True。
  • admin token 存在。
  • check_poll_available() 成功。

若 admin endpoint 不可用:启动延迟重试任务,不允许长期无兜底结算运行。

11.2 参数

参数 默认值
poll interval 60s
batch size 50
per-record timeout 30s
consecutive failure threshold 5
backoff after threshold 600s
artifact registration max retries 10

11.3 单轮流程

1. records = compute_client.list_pending_usage(admin_token, settled=0, limit=50)
2. HTTP 错误:failure_count += 1;达到 5 次 sleep 600s。
3. 对每条 record:
   a. 若 output_manifest_json 非空且 artifacts_registered=false:
      i. 解析 artifacts。
      ii. storage_service.register_artifacts(user_id, thread_id, artifacts)。
      iii. 成功 -> mark_artifacts_registered(usage_id=record.id)。
      iv. 失败 -> 记录 error,递增 registration retry;未达上限则 continue 下一条。
      v. 达上限 -> 标记 artifacts_unregistered,允许继续 settle。
   b. 调用 settle_compute_usage。
   c. settled 或 already_settled -> mark_usage_settled([record.id])。
   d. InsufficientComputeBalance -> 保留 pending_usage,等待充值。
   e. SettlementInProgress -> 跳过,下一轮重试。
   f. 其他错误 -> 记录 ERROR,下一条。
4. sleep 60s。

11.4 并发安全

  • StreamHandler finally 与 poll_loop 可能同时处理同一 environment_id。
  • compute_usage_log.environment_id UNIQUE + 单事务 claim 保证只扣一次。
  • mark_usage_settled 必须幂等。
  • mark_artifacts_registered 必须幂等。

12. Gateway lifespan 改造

启动顺序:

1. 加载 CLI config/env。
2. init_gateway_db()。
3. db = await get_connection()。
4. init StorageService。
5. 若 STORAGE_BACKEND=s3:HeadBucket + probe put/get/delete。
6. 启动 storage_cleanup_loop。
7. 读取 EXECUTION_SERVICE_URL。
8. URL 未设置:compute_client=None,不启动 poll_loop。
9. URL 已设置:校验 HMAC secret/admin token,缺失则 startup fail。
10. init ComputeClient。
11. await compute_client.check_ready()。
12. ready true:check_poll_available,启动 poll_loop。
13. ready false:admin poll 延迟重试,Web compute unavailable。
14. 启动 refresh_expired_compute_quotas loop。
15. app.state 写入 db/storage_service/compute_client/admin_token/task handles。
16. yield。
17. shutdown:cancel tasks -> close compute_client -> close storage_service -> close_gateway_db。

/ready 失败不导致 Gateway 必然启动失败;但 Web compute 请求必须失败为 ComputeUnavailable。


13. 文件路由改造

13.1 uploads.py

上传流程:

  1. require auth。
  2. require owned thread。
  3. read upload bytes with max size limit。
  4. storage_service.put_file(owner_user_id=user_uid, logical_path=logical_path, data=file_bytes, thread_id=thread_id, source='upload')。
  5. 返回 file_id/logical_path/sha256/size/storage_backend。

13.2 user_files.py / files.py / global_uploads.py

  • list:查询 user_files active records。
  • download:StorageService.open_read。
  • delete:StorageService.delete_file soft delete。
  • overwrite:StorageService.put_file 生成新 version,旧 active deleted。

13.3 knowledge_indexer.py / knowledge.py

  • 不再直接 thread_data_dir(user_uid, thread_id) / virtual_path 读取。
  • 通过 user_files 查询 file_id。
  • 通过 StorageService.open_read 传给 parser。
  • LocalStorageBackend 路径只作为 backend 内部细节。

14. 非计算降级

14.1 NonComputingFallbackBackend

包装 MultiRootSandboxBackend,但 execute 永远失败:

class NonComputingFallbackBackend(MultiRootSandboxBackend):
    def execute(self, command: str, *, timeout: int | None = None) -> ExecuteResponse:
        return ExecuteResponse(
            output="ComputeUnavailable: remote compute is unavailable; local execution is disabled for Web threads.",
            exit_code=1,
            truncated=False,
        )

文件操作允许范围:

  • read/list/download existing files。
  • write/edit only in controlled recovery workspace if UI 明确处于非计算降级。
  • 不允许任何 tool 通过文件操作触发代码执行。

14.2 fallback sync

StreamHandler cleanup:

  1. 记录 stream_start_time。
  2. cleanup 时扫描 recovery workspace 中 mtime > stream_start_time 的文件。
  3. 对每个文件调用 storage_service.put_file。
  4. 全部成功 -> 删除 recovery workspace。
  5. 任一失败 -> 移动到 controlled recovery/orphan dir,写 fallback_sync_recovery。
  6. cleanup loop 后续重试。

15. shared_compute 路由

15.1 保守判定

进入 simple/shared 的必要条件:

  • EXECUTION_SHARED_COMPUTE_ENABLED=true
  • plan 允许 compute
  • storage_backend=s3
  • input_objects 总大小 <= 10MiB
  • 文件数 <= 100
  • effective max_runtime_seconds <= 60
  • 不需要 GPU
  • 不需要自定义镜像
  • 不需要系统依赖安装
  • 不需要 interactive session
  • 不需要服务进程
  • 不需要持久 workspace
  • 无强隔离标记

否则 standard。

15.2 limits

Gateway 下发 P0 shared limits:

{
  "max_runtime_seconds": 60,
  "max_input_bytes": 10485760,
  "max_output_bytes": 10485760,
  "max_stdout_stderr_bytes": 1048576,
  "max_files": 100,
  "max_user_concurrency": 2
}

15.3 失败语义

  • 执行前 not eligible:standard 或明确 SharedComputeNotEligible。
  • 执行中超限:失败,不半路迁移。
  • worker degraded:不进入 shared。
  • 任何 shared 失败不得 fallback Gateway 本地。

16. 部署方案

16.1 .env.example 增加

# Storage
STORAGE_BACKEND=s3
STORAGE_S3_ENDPOINT=http://127.0.0.1:9000
STORAGE_S3_BUCKET=evoscientist
STORAGE_S3_REGION=us-east-1
STORAGE_S3_ACCESS_KEY=minioadmin
STORAGE_S3_SECRET_KEY=<minio-or-s3-secret-key>
STORAGE_S3_FORCE_PATH_STYLE=true
STORAGE_MAX_FILE_SIZE_BYTES=1073741824
STORAGE_RETENTION_DAYS=0

# Remote compute
EXECUTION_SERVICE_URL=http://127.0.0.1:9020
EXECUTION_HMAC_SECRET=<generated-32-byte-hex>
EXECUTION_ADMIN_TOKEN=<generated-admin-token>
EXECUTION_DEFAULT_RESOURCE_CLASS=small
EXECUTION_DEFAULT_BACKEND_POLICY=auto
EXECUTION_SHARED_COMPUTE_ENABLED=true
EXECUTION_SIMPLE_MAX_RUNTIME_SECONDS=60
EXECUTION_PLAN_HARD_MAX_RUNTIME_SECONDS=7200

16.2 本地启动顺序

# 1. PostgreSQL
# existing local command

# 2. MinIO
docker run -p 9000:9000 -p 9001:9001 \
  -e MINIO_ROOT_USER=minioadmin \
  -e MINIO_ROOT_PASSWORD=<local-minio-password> \
  quay.io/minio/minio server /data --console-address ':9001'

# 3. Create bucket
mc alias set local http://127.0.0.1:9000 minioadmin minioadmin
mc mb --ignore-existing local/evoscientist

# 4. k8s-exec-service-mcp
cd /Users/m4/Projects/EvoSci/MCP/k8s-exec-service-mcp
# run service per its README

# 5. EvoScientist Gateway
cd /Users/m4/Projects/EvoSci/EvoScientist
source venv/bin/activate
# export STORAGE_* EXECUTION_*
# run gateway

16.3 smoke test

Gateway 侧 smoke test 必须验证:

  1. StorageService S3 probe success。
  2. GET EXECUTION_SERVICE_URL/ready 200。
  3. 上传小 CSV -> user_files active row。
  4. build_input_manifest 返回 s3 object。
  5. create_environment standard。
  6. execute python --version。
  7. write/read file。
  8. destroy persist artifact。
  9. register_artifacts row。
  10. settle_compute_usage row + wallet_ledger。
  11. pending_usage mark-settled。
  12. 停止 exec service 后 Web compute 不执行本地命令。

17. 实施阶段

Phase 1:Storage + ComputeClient 骨架

  1. user_files migration + corrected indexes。
  2. StorageService + LocalStorageBackend。
  3. S3StorageBackend + MinIO probe。
  4. migrate_local_files_to_storage dry-run。
  5. ComputeClient HMAC + golden vector。
  6. ComputeClient ready/health/admin readiness。
  7. env/config wiring。
  8. tests: storage + hmac + ready。

Gate:StorageService + MinIO + HMAC 双端测试通过。

Phase 2:Remote Backend 最小闭环

  1. ContainerSandboxBackend sync adapter。
  2. NonComputingFallbackBackend。
  3. create_cli_agent 参数与 backend selection。
  4. StreamHandler async prepare/create。
  5. threads.py app_state 注入。
  6. ComputeClient create/execute/read/write/destroy wrappers。
  7. happy path E2E:upload -> manifest -> execute -> artifact。

Gate:ComputeUnavailable 不本地执行;remote happy path 通过。

Phase 3:Billing + poll_loop

  1. compute pricing config。
  2. user_balances compute columns。
  3. compute_usage_log。
  4. get_transaction_connection。
  5. settle_compute_usage atomic transaction。
  6. wallet_ledger compatibility。
  7. poll_loop。
  8. artifact registration retry。

Gate:重复 environment_id 不重复扣费;artifact 注册失败不会免费;余额不足保留 pending_usage。

Phase 4:File API 全量迁移 + fallback_sync

  1. uploads.py。
  2. user_files.py/files.py/global_uploads.py。
  3. knowledge_indexer.py/knowledge.py。
  4. fallback_sync_recovery。
  5. cleanup loops。
  6. migration script full run。

Gate:Web 文件 API 不直接依赖 thread_data_dir。

Phase 5:shared_compute + 部署验收

  1. execution_mode decision。
  2. shared limits。
  3. create_environment shared-compatible payload。
  4. standard fallback。
  5. local smoke test。
  6. kind smoke test。
  7. docs update。

18. 测试清单

18.1 单元测试

  • normalize_logical_path rejects traversal/control/windows/absolute。
  • user_files indexes allow same logical_path in different threads。
  • StorageService put_file object-first DB-second。
  • build_input_manifest workspace/global/thread readonly。
  • build_input_manifest local backend rejection。
  • register_artifacts idempotent by environment_id + logical_path。
  • ComputeClient canonical_json exact。
  • HMAC golden vector。
  • ComputeClient status -> exception mapping。
  • ContainerSandboxBackend method mapping to deepagents return types。
  • NonComputingFallbackBackend execute disabled。
  • preflight quota boundary cases。
  • settle mixed free+cash Decimal。
  • settle rollback on insufficient balance。
  • poll_loop artifact-first settlement。

18.2 集成测试

  • Gateway startup compute disabled。
  • Gateway startup fails when URL set but secret missing。
  • /ready false -> compute unavailable。
  • upload -> user_files -> S3 object。
  • StreamHandler async create -> remote execute。
  • destroy -> artifacts -> register -> settle。
  • artifact registration failure -> no settle until retry。
  • duplicate settlement -> once only。
  • STORAGE_BACKEND=local -> compute blocked。
  • shared eligible -> create payload execution_mode simple。
  • shared not eligible -> standard。

18.3 E2E

  • CSV analysis artifact downloadable。
  • MinIO down -> no environment created。
  • exec service down -> no local command execution。
  • balance only 1 minute -> max_runtime_seconds=60。
  • force_destroy artifacts_lost visible。
  • fallback sync failure -> recovery row -> retry success。
  • kind deployment smoke test。

19. 验收标准

P0 完成标准:

  1. Web 用户代码执行 100% 通过 k8s-exec-service-mcp。
  2. Gateway 本地不执行 Web 用户 shell/Python/subprocess。
  3. S3/MinIO 是 Web remote compute 的唯一文件内容底座。
  4. LocalStorageBackend 只用于 CLI、迁移、非计算文件管理。
  5. StorageService path normalization、input manifest、artifact registration 完整。
  6. ComputeClient HMAC golden vector 双端一致。
  7. ContainerSandboxBackend 与 deepagents 返回类型兼容。
  8. StreamHandler async preflight + destroy-settle 闭环完整。
  9. compute billing 原子事务不产生半结算。
  10. pending_usage poll_loop 可补注册 artifact 并兜底结算。
  11. shared_compute 通过 create_environment 兼容字段接入,不新增 create_job。
  12. user_files 支持不同 thread 同名文件。
  13. k8s-exec-service-mcp 不可用时 Web compute 返回 ComputeUnavailable。
  14. 本地单机 smoke test 和 kind smoke test 均通过。

20. 风险与硬性禁止项

硬性禁止:

  • 禁止 Web compute fallback 到 CustomSandboxBackend。
  • 禁止 Web compute fallback 到 LocalShellBackend。
  • 禁止把 Gateway 本地路径传给 k8s-exec-service-mcp。
  • 禁止 object_key 作为权限依据。
  • 禁止 float 参与 compute 扣款。
  • 禁止 artifact 未注册且未达 retry 上限时直接结算。
  • 禁止新增未在两端协议中确认的 /tools/create_job。
  • 禁止 k8s-exec-service-mcp 多副本写同一个 SQLite。

主要风险:

风险 处理
本地路径引用遗漏 全仓搜索 thread_data_dir, user_data_dir, ~/.evoscientist, data/
cents/Decimal 账本冲突 明确 user_balances.cash_balance 为真实余额,ledger 只审计
async/sync backend 复杂 P0 用 backend 专用 loop thread;P1 评估 async backend
HMAC 双端不一致 golden vector 作为 Gate Check
shared 安全误判 保守路由,不确定即 standard/isolated
artifact 注册失败免单 pending_usage retry,上限后 artifacts_unregistered 仍结算

21. 多轮评审结果与实施就绪判定

21.1 R1 评审发现与修订

问题 严重级别 修订结果
文档头部保存路径仍指向旧目录 P1 已改为 EvoScientist/docs 路径
env 样例中 EXECUTION_HMAC_SECRET 行损坏 P0 已修复为 <generated-32-byte-hex> 与 <generated-admin-token>
shared_compute 中出现未定义 create_job 风险 P0 已明确 P0 不新增 create_job,统一走 create_environment 兼容 handle
HMAC 只有算法无固定可测 token P0 已补真实计算 golden vector
placeholder 过多 P1 接口片段改为 pass 或具名占位符
user_files 唯一性可能阻塞不同线程同名文件 P0 已采用 global/thread/project 三类 partial unique index
StreamHandler async 边界不清 P0 已明确 _create_agent 改 async,backend 内部用 loop thread 适配 sync protocol

21.2 R2 复评结论

复评结果:符合实施要求。当前文档已经满足以下条件:

  1. 每个 P0 模块均有目标文件、接口、核心算法和测试 Gate。
  2. StorageService、ComputeClient、ContainerSandboxBackend、compute billing、poll_loop 的关键状态和失败路径已定义。
  3. shared_compute 不引入未确认 endpoint,避免两仓协议分叉。
  4. Web 本地执行禁令被落实到 create_cli_agent、NonComputingFallbackBackend、测试和硬性禁止项。
  5. 数据库 migration 包含 user_files namespace 修正和 compute 结算状态。
  6. 可以按 Phase 1-5 拆分 PR,并以各 Phase Gate 作为合并条件。

21.3 PR 拆分建议

PR 范围 合并 Gate
PR-1 DB migration: user_files、compute_usage_log、user_balances compute columns migration 幂等测试通过
PR-2 StorageService + Local/S3 backend + MinIO probe storage unit + MinIO integration 通过
PR-3 ComputeClient + HMAC + ready/health/admin wrappers golden vector 双端一致
PR-4 ContainerSandboxBackend + NonComputingFallbackBackend deepagents backend mapping 测试通过
PR-5 StreamHandler async + create_cli_agent 参数接入 ComputeUnavailable 不本地执行
PR-6 Billing settlement + wallet_ledger + poll_loop 重复结算只扣一次
PR-7 Upload/files/knowledge 路由迁移 StorageService 全仓 Web 文件 API 不直接依赖 thread_data_dir
PR-8 shared_compute routing + smoke tests + docs local/kind smoke 通过

22. 最小可执行切片

第一周最小闭环:

  1. user_files 表 + corrected indexes。
  2. StorageService + S3StorageBackend + MinIO probe。
  3. ComputeClient HMAC + /ready。
  4. ContainerSandboxBackend execute/read/write/destroy happy path。
  5. StreamHandler async create remote backend。
  6. Web 上传小文件 -> build_input_manifest -> remote python --version -> destroy -> artifact 入库。
  7. 关闭 exec service 后确认不会调用 CustomSandboxBackend/LocalShellBackend。

该切片通过后,再进入 billing/poll_loop/shared_compute 全量实现。