Self-evolving AI scientist framework built on LangGraph/LangChain with CLI/TUI core, FastAPI gateway, and Next.js frontend. Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
56 KiB
EvoScientist k8s 执行引擎 v1.6 实施方案
文件:
~/Projects/EvoSci/EvoScientist/docs/EvoScientist-k8s-1.6.md依据:~/Projects/EvoSci/EvoScientist/docs/execution-engine-development-1.6.md目标仓库:~/Projects/EvoSci/EvoScientist关联独立仓库:/Users/m4/Projects/EvoSci/MCP/k8s-exec-service-mcp状态:开发就绪版实施方案(补齐评审意见)
评审状态:已完成多轮自评审修订
- R1:修复路径、env 样例损坏、placeholder、shared_compute 接口兼容、StorageService/ComputeClient/计费/poll_loop 细节。
- R2:补齐迁移顺序、依赖矩阵、测试 Gate 和硬性禁止项,确认可按 Phase 拆 PR 实施。
0. 执行结论
EvoScientist 侧必须从“Gateway 本地多根文件系统 + 本地 shell backend”迁移到“StorageService + S3/MinIO 对象存储 + ComputeClient + ContainerSandboxBackend”的远程执行架构。
v1.6 的核心约束:
- Web 用户代码执行必须通过
k8s-exec-service-mcp。 - Gateway 不允许执行 Web 用户 shell / Python / 本地 subprocess。
MultiRootSandboxBackend只能用于非计算文件展示、降级恢复和 fallback_sync,不得作为 Web compute backend。- 用户文件、目录树、权限、生命周期和 artifact 元数据的事实来源是 EvoScientist
StorageService + Gateway PostgreSQL。 - 文件 bytes 的事实来源是 S3-compatible object storage,本地 P0 使用 MinIO。
k8s-exec-service-mcp只管理运行期 workspace、调度、usage 和 artifact upload,不保存长期用户文件。shared_compute是执行服务侧的一等执行模式;EvoScientist 只传 execution_mode / limits / input_objects / artifact_prefix,不在 Gateway 中实现共享 worker。
本方案补齐以下评审指出的 P0 缺口:
- StorageService 详细算法与路径规范
- user_files namespace 唯一性修正
- ComputeClient 13 个 endpoint wrapper 契约
- HMAC canonical_json / signing 伪代码
- StreamHandler async 改造决策
- deepagents BackendProtocol 映射
- compute billing 原子事务规格
- pending_usage poll_loop 详细流程
- shared_compute API 兼容决策
- 本地部署和测试验收细化
1. 当前代码冲突点
已确认当前 EvoScientist 代码与 v1.6 目标存在以下冲突:
| 文件 | 当前行为 | v1.6 要求 | 修改方向 |
|---|---|---|---|
gateway/main.py |
lifespan 只初始化 DB、限流、pricing、知识库任务 | 初始化 StorageService、ComputeClient、poll_loop、quota_loop | 增加 compute/storage startup/shutdown |
gateway/services/stream_handler.py |
_create_agent() 同步创建本地 Web agent |
async preflight、input manifest、remote backend、destroy-settle | 改造为 async create + finally cleanup |
gateway/routes/threads.py |
StreamHandler 未传 app_state | 传入 storage/compute/db/admin token | 注入 request.app.state._state |
EvoScientist/EvoScientist.py:create_cli_agent() |
Web 默认 MultiRootSandboxBackend |
Web compute 必须 ContainerSandboxBackend |
新增 compute 参数和 backend 选择 |
EvoScientist/backends.py |
缺 ContainerSandboxBackend |
实现远程 BackendProtocol | 增加 remote backend 和 non-compute fallback |
gateway/routes/uploads.py |
直接写本地文件 | 通过 StorageService 写 metadata + object | 改 StorageService |
gateway/routes/user_files.py / files.py / global_uploads.py |
文件列表/下载依赖本地路径 | 查询 user_files + StorageService open_read | 统一存储入口 |
gateway/services/knowledge_indexer.py |
用 thread_data_dir() 读取文件 |
通过 StorageService 读取 | 改文件解析入口 |
gateway/pricing.json |
无 compute_pricing | 增 compute 定价和分钟配额 | 扩展 pricing config |
gateway/services/wallet.py / token_tracker.py |
金额存在 cents/REAL 语义混用风险 | compute 使用 Decimal + NUMERIC(12,6) | 增迁移和兼容层 |
docs/execution-engine-deployment.md |
旧 Gateway 内 Execution Engine 方案 | 独立 k8s-exec-service-mcp | 标记过期或重写 |
2. 目标架构
2.1 调用链
Web POST /api/threads/{thread_id}/stream
-> gateway/routes/threads.py
-> StreamHandler(app_state)
-> await StreamHandler._prepare_compute_context()
-> Compute quota preflight
-> StorageService.build_input_manifest()
-> artifact_prefix/default_output_paths/budget_snapshot
-> execution_mode decision
-> await StreamHandler._create_agent()
-> create_cli_agent(source="web", compute_client=<ComputeClient>, input_objects=<list[InputObject]>)
-> CompositeBackend
default: ContainerSandboxBackend
-> ComputeClient HTTP /tools/*
-> k8s-exec-service-mcp
-> ResourceManager
-> shared_compute 或 independent environment/warm pool
/skills/: local skills backend
/memory/: local memory backend
-> SSE streaming
-> finally:
-> await _destroy_and_settle_backend()
-> backend.destroy(persist_outputs=True)
-> StorageService.register_artifacts()
-> settle_compute_usage()
-> pending_usage poll_loop 兜底
2.2 存储边界
Gateway PostgreSQL
- user_files metadata
- compute_usage_log
- user_balances
- wallet_ledger
- fallback_sync_recovery
- threads/session metadata
Object Storage: MinIO/S3/OSS/COS/OBS/R2
- users/{user_id}/files/{file_id}
- projects/{project_id}/files/{file_id}
- threads/{thread_id}/files/{file_id}
- jobs/{job_id}/artifacts/{artifact_id}
- tmp/{job_id}/{object_id}
k8s-exec-service-mcp SQLite WAL
- exec_environments
- exec_pending_usage
- exec_backends
- exec_warm_pools
- exec_auth_nonces
- exec_admin_audit_log
边界规则:
- Gateway 不读写 exec service SQLite。
- exec service 不读写 Gateway PostgreSQL。
- object_key 不是权限依据;权限来自 Gateway metadata + HMAC 签名的 object_scope。
- Web remote compute P0 必须使用 S3StorageBackend;LocalStorageBackend 不得生成 compute input manifest。
3. StorageService 详细实施规格
3.1 新增文件
gateway/services/storage_service.py
gateway/services/storage_backends/__init__.py
gateway/services/storage_backends/local.py
gateway/services/storage_backends/s3.py
gateway/services/storage_errors.py
scripts/migrate_local_files_to_storage.py
3.2 数据模型
@dataclass(slots=True)
class FileRecord:
file_id: str
owner_user_id: str
logical_path: str
object_key: str
storage_backend: Literal["local", "s3"]
sha256: str
size_bytes: int
thread_id: str | None = None
project_id: str | None = None
user_id_int: int | None = None
mime_type: str | None = None
source: Literal["upload", "agent", "artifact", "import"] = "upload"
version: int = 1
status: Literal["active", "deleted", "pending", "orphan"] = "active"
created_at: datetime | None = None
updated_at: datetime | None = None
expires_at: datetime | None = None
metadata_json: dict[str, Any] | None = None
@dataclass(slots=True)
class InputObject:
file_id: str
object_key: str
sha256: str
size_bytes: int
mount_path: str
read_only: bool
@dataclass(slots=True)
class ArtifactObject:
artifact_id: str
object_key: str
sha256: str
size_bytes: int
logical_path: str
metadata: dict[str, Any] | None = None
3.3 路径规范化算法
所有进入 StorageService 的 logical_path 必须统一为“相对 POSIX 路径”,不以 / 开头。
def normalize_logical_path(path: str) -> str:
if not isinstance(path, str):
raise InvalidPath("path must be string")
path = unicodedata.normalize("NFC", path.strip())
if not path:
raise InvalidPath("empty path")
if "\x00" in path or any(ord(c) < 32 for c in path):
raise InvalidPath("control character")
path = path.replace("\\", "/")
if re.match(r"^[A-Za-z]:", path):
raise InvalidPath("windows drive prefix")
path = urllib.parse.unquote(path)
path = path.lstrip("/")
parts = [p for p in path.split("/") if p]
if any(p in (".", "..") for p in parts):
raise InvalidPath("path traversal")
if not parts:
raise InvalidPath("empty normalized path")
if parts[0] in (".control", ".secrets", "__evosci_internal__"):
raise InvalidPath("reserved prefix")
return "/".join(parts)
mount_path 规则:
/workspace/{logical_path}:当前线程读写。/__global__/{logical_path}:用户全局读写。/threads/{thread_id}/{logical_path}:其他线程只读。- 禁止 mount_path 冲突、父子覆盖、读写权限冲突。
3.4 user_files namespace 唯一性修正
不要使用单一 (owner_user_id, logical_path) active 唯一索引,否则不同 thread 下同名文件冲突。
使用三个 partial unique index:
CREATE UNIQUE INDEX IF NOT EXISTS uq_user_files_global_active_path
ON user_files(owner_user_id, logical_path)
WHERE status = 'active' AND thread_id IS NULL AND project_id IS NULL;
CREATE UNIQUE INDEX IF NOT EXISTS uq_user_files_thread_active_path
ON user_files(owner_user_id, thread_id, logical_path)
WHERE status = 'active' AND thread_id IS NOT NULL;
CREATE UNIQUE INDEX IF NOT EXISTS uq_user_files_project_active_path
ON user_files(owner_user_id, project_id, logical_path)
WHERE status = 'active' AND project_id IS NOT NULL;
对象 key 生成:
| namespace | object_key |
|---|---|
| user global | users/{user_id}/files/{file_id} |
| thread | threads/{thread_id}/files/{file_id} |
| project | projects/{project_id}/files/{file_id} |
| artifact | jobs/{job_id}/artifacts/{artifact_id} |
| temp | tmp/{job_id}/{object_id} |
3.5 StorageService 接口
class StorageService:
def __init__(self, db, backend: StorageBackend, *, backend_name: str): pass
async def put_file(
self,
owner_user_id: str,
logical_path: str,
data: bytes,
*,
thread_id: str | None = None,
project_id: str | None = None,
source: str = "upload",
mime_type: str | None = None,
expires_at: datetime | None = None,
metadata: dict[str, Any] | None = None,
) -> FileRecord: pass
async def get_file(self, file_id: str, requester_user_id: str) -> FileRecord: pass
async def open_read(self, file_id: str, requester_user_id: str) -> AsyncIterator[bytes]: pass
async def delete_file(self, file_id: str, requester_user_id: str) -> None: pass
async def build_input_manifest(
self,
user_id: str,
thread_id: str,
paths: list[str],
*,
require_remote_compute: bool = True,
) -> list[InputObject]: pass
async def register_artifacts(
self,
user_id: str,
thread_id: str,
artifacts: list[ArtifactObject],
*,
environment_id: str | None = None,
job_id: str | None = None,
) -> list[FileRecord]: pass
async def cleanup_expired(self, *, limit: int = 100) -> int: pass
async def close(self) -> None: pass
3.6 put_file 事务边界
顺序固定:
- normalize logical_path。
- 计算 sha256 / size。
- 生成 file_id。
- 生成 object_key。
- 上传 object storage。
- HeadObject 校验 size/sha256 metadata。
- 在 DB 事务中:
- 将同 namespace + logical_path 的旧 active 记录置为 deleted。
- 插入新 user_files 记录,version=previous+1。
- 若 DB 写失败:
- 尝试删除刚上传 object。
- 删除失败时写 orphan metadata 或 ERROR 日志,由 cleanup 后台任务处理。
不得先写 DB 再上传 object,避免 metadata 指向不存在对象。
3.7 build_input_manifest 算法
输入:
- user_id:当前 Web 用户 UID。
- thread_id:当前线程 ID。
- paths:允许根路径,P0 默认由 StreamHandler 传:
/workspace/__global__- 显式
/threads/{other_thread_id}列表
禁止传 /threads/* 作为实现级路径。/threads/* 只能是文档描述,实际必须展开为当前用户可读线程列表。
算法:
1. 若 require_remote_compute=True 且 backend != s3:
抛 StorageBackendUnsupportedForCompute。
2. 初始化 input_objects=[]。
3. 对 paths 逐个处理:
a. /workspace:
查询 user_files
WHERE owner_user_id=$user_id
AND thread_id=$thread_id
AND status='active'
mount_path = /workspace/{logical_path}
read_only = false
b. /__global__:
查询 user_files
WHERE owner_user_id=$user_id
AND thread_id IS NULL
AND project_id IS NULL
AND status='active'
mount_path = /__global__/{logical_path}
read_only = false
c. /threads/{other_thread_id}:
先校验 requester 对 other_thread_id 有读权限。
查询 user_files
WHERE thread_id=$other_thread_id
AND status='active'
mount_path = /threads/{other_thread_id}/{logical_path}
read_only = true
d. 其他路径:抛 InvalidManifestPath。
4. 对每个 FileRecord:
- object_key 必须非空。
- storage_backend 必须为 s3。
- S3 HeadObject 必须成功。
- size/sha256 与 metadata 一致。
- object_key 必须位于允许 prefix,但 prefix 不作为权限依据。
5. 检查 mount_path:
- UTF-8 NFC。
- 禁止 NUL/control。
- 必须位于 /workspace、/__global__、/threads/{id}。
- 禁止 ..。
- 禁止重复。
- 禁止父子覆盖冲突。
- 禁止 read_only=false 写入 /threads/*。
6. 按 mount_path 排序。
7. 返回 input_objects。
失败策略:
- P0 所有 active user_files 都是关键输入。
- 任一 object missing/corrupt -> 抛
InputObjectMissing或InputObjectCorrupt。 - 不静默跳过。
3.8 register_artifacts 幂等规则
artifact 入库幂等键:
优先级:
metadata.environment_id + artifact.logical_pathmetadata.job_id + artifact.logical_pathartifact.object_key
算法:
- 校验 artifact.object_key 必须在
jobs/{thread_id}/{request_id}/artifacts/或 Gateway 下发的 artifact_prefix 内。 - HeadObject 校验 sha256 / size。
- logical_path normalize;若 artifact.logical_path 以
/workspace/开头,转为相对路径。 - DB 事务中 upsert user_files:
- source='artifact'
- thread_id=current thread
- metadata_json 包含 environment_id/job_id/object_key/original_path。
- 若相同 artifact 幂等键已存在,直接返回旧记录。
- 调用 ComputeClient.mark_artifacts_registered() 由上层 StreamHandler/poll_loop 完成,不在 StorageService 内直接调用 ComputeClient。
4. ComputeClient 实施规格
4.1 新增文件
gateway/services/compute_client.py
gateway/services/compute_errors.py
tests/test_compute_client.py
4.2 配置
EXECUTION_SERVICE_URL=http://127.0.0.1:9020
EXECUTION_HMAC_SECRET=<generated-32-byte-hex>
EXECUTION_ADMIN_TOKEN=<generated-admin-token>
EXECUTION_HTTP_TIMEOUT_SECONDS=30
EXECUTION_CREATE_TIMEOUT_SECONDS=210
EXECUTION_READY_TIMEOUT_SECONDS=5
配置矩阵:
| EXECUTION_SERVICE_URL | HMAC/Admin | Gateway 行为 |
|---|---|---|
| 未设置 | 可缺失 | compute disabled,正常启动,不启动 poll_loop |
| 已设置 | 任一缺失 | startup fail |
| 已设置 | 都存在,/ready false | 正常启动,但 compute_client.ready=false |
| 已设置 | 都存在,/ready true | compute enabled,启动 poll_loop |
4.3 HMAC signing 规范
每个 /tools/* 请求 body 顶层包含 auth_context。签名前移除 auth_context.token 字段。
canonical string:
{METHOD}\n{PATH}\n{issued_at}\n{nonce}\n{sha256(canonical_json(body_without_auth_context.token))}
实现伪代码:
def canonical_json(obj: Any) -> str:
return json.dumps(
obj,
ensure_ascii=False,
sort_keys=True,
separators=(",", ":"),
)
def payload_hash(body_without_token: dict[str, Any]) -> str:
data = canonical_json(body_without_token).encode("utf-8")
return hashlib.sha256(data).hexdigest()
def canonical_string(method: str, path: str, issued_at: int, nonce: str, body_hash: str) -> str:
return "\n".join([method.upper(), path, str(issued_at), nonce, body_hash])
def sign(secret: str, canonical: str) -> str:
return hmac.new(
secret.encode("utf-8"),
canonical.encode("utf-8"),
hashlib.sha256,
).hexdigest()
def build_auth_context(method: str, path: str, body_params: dict[str, Any], *, user_id: str, thread_id: str, object_scope: dict | None):
issued_at = int(time.time())
nonce = secrets.token_hex(16)
auth_context = {
"user_id": user_id,
"thread_id": thread_id,
"issued_at": issued_at,
"nonce": nonce,
}
if object_scope:
auth_context["object_scope"] = object_scope
body_without_token = {"auth_context": auth_context, **body_params}
digest = payload_hash(body_without_token)
canonical = canonical_string(method, path, issued_at, nonce, digest)
token = sign(self.hmac_secret, canonical)
auth_context["token"] = token
return {"auth_context": auth_context, **body_params}
要求:
path必须是不含 scheme/host/query 的路径,如/tools/execute_command。- canonical_json 必须递归排序 key;
json.dumps(sort_keys=True)对 Python dict 嵌套满足要求。 - separators 必须是
(',', ':'),无空格。 - canonical string 无 trailing LF,总共 4 个 LF。
- token lowercase hex。
- 日志禁止打印 secret/token/body 全量。
4.4 HMAC golden vector 测试
Gateway 与 k8s-exec-service-mcp 均必须包含同一测试向量。若源规格中的 expected hash/token 与实际算法不一致,以测试脚本真实计算结果为准并同步更新两仓;不得两侧各自写死不同 token。
测试断言:
- canonical_json exact string。
- payload_hash exact hex。
- canonical_string
repr()exact。 - token exact hex。
- body 增加/删除任一字段 token 变化。
- CRLF 与 LF 不同,CRLF 必须失败。
固定测试向量(按上述算法真实计算):
method = POST
path = /tools/execute_command
secret = test-hmac-secret-for-unit-tests-only
canonical_json = {"auth_context":{"issued_at":1716500000,"nonce":"abc123def456","object_scope":{"read":["objects/in/data.csv"],"write_prefixes":["jobs/job_1/artifacts/"]},"user_id":"test_user"},"command":"echo hello","environment_id":"env_test123"}
payload_hash = 45957f2078d85c7c09bedfea3f000b6a42a8e01e039250644bb577321186035b
canonical_string_repr = 'POST\n/tools/execute_command\n1716500000\nabc123def456\n45957f2078d85c7c09bedfea3f000b6a42a8e01e039250644bb577321186035b'
token = 06af87716e9c3ba83441362dad2faa543c5db19d1c47fb937cd2ffe29cee1d21
4.5 ComputeClient public methods
class ComputeClient:
async def check_ready(self) -> bool: pass
async def check_health(self) -> dict: pass
async def check_poll_available(self) -> bool: pass
async def close(self) -> None: pass
async def create_environment(
self,
*,
user_id: str,
thread_id: str,
resource_class: str,
backend_policy: str,
execution_mode: Literal["simple", "shared", "standard", "isolated"],
input_objects: list[InputObject],
artifact_prefix: str,
default_output_paths: list[str],
max_runtime_seconds: int,
budget_snapshot: dict[str, Any],
limits: dict[str, Any] | None = None,
image_ref: str | None = None,
) -> CreateEnvironmentResult: pass
async def destroy_environment(
self,
*,
user_id: str,
thread_id: str,
environment_id: str,
persist_outputs: bool,
output_paths: list[str],
force_destroy: bool = False,
artifact_prefix: str | None = None,
) -> DestroyEnvironmentResult: pass
async def list_environments(self, *, user_id: str, thread_id: str) -> list[dict]: pass
async def get_environment_status(self, *, user_id: str, thread_id: str, environment_id: str) -> dict: pass
async def execute_command(self, *, user_id: str, thread_id: str, environment_id: str, command: str, timeout_seconds: int | None = None, max_output_bytes: int | None = None) -> dict: pass
async def read_file(self, *, user_id: str, thread_id: str, environment_id: str, path: str) -> dict: pass
async def write_file(self, *, user_id: str, thread_id: str, environment_id: str, path: str, content: str) -> dict: pass
async def edit_file(self, *, user_id: str, thread_id: str, environment_id: str, path: str, old_string: str, new_string: str, replace_all: bool = False) -> dict: pass
async def list_dir(self, *, user_id: str, thread_id: str, environment_id: str, path: str) -> dict: pass
async def grep_files(self, *, user_id: str, thread_id: str, environment_id: str, pattern: str, path: str, glob: str | None = None) -> dict: pass
async def glob_files(self, *, user_id: str, thread_id: str, environment_id: str, pattern: str, path: str = "/workspace") -> dict: pass
async def upload_file(self, *, user_id: str, thread_id: str, environment_id: str, remote_path: str, content_base64: str, sha256: str) -> dict: pass
async def download_file(self, *, user_id: str, thread_id: str, environment_id: str, remote_path: str) -> dict: pass
async def list_pending_usage(self, admin_token: str, *, settled: int = 0, limit: int = 50) -> list[dict]: pass
async def mark_usage_settled(self, admin_token: str, usage_ids: list[int]) -> None: pass
async def mark_artifacts_registered(self, admin_token: str, *, environment_id: str | None = None, usage_id: int | None = None) -> None: pass
4.6 Endpoint mapping
| Method | Endpoint | Auth | Timeout | Notes |
|---|---|---|---|---|
| GET | /ready |
none | 5s | readiness only |
| GET | /health |
none | 5s | liveness only |
| POST | /tools/create_environment |
HMAC | 210s | create or bind warm/shared handle |
| POST | /tools/destroy_environment |
HMAC | 60s | persist artifacts first |
| POST | /tools/list_environments |
HMAC | 30s | user visible only |
| POST | /tools/get_environment_status |
HMAC | 30s | runtime state |
| POST | /tools/execute_command |
HMAC | command timeout + 5s | no local fallback |
| POST | /tools/read_file |
HMAC | 30s | P0 size limits |
| POST | /tools/write_file |
HMAC | 30s | P0 size limits |
| POST | /tools/edit_file |
HMAC | 30s | old_string unique |
| POST | /tools/list_dir |
HMAC | 30s | max entries |
| POST | /tools/grep_files |
HMAC | 30s | max matches |
| POST | /tools/glob_files |
HMAC | 30s | max matches |
| POST | /tools/upload_file |
HMAC | 60s | base64 only |
| POST | /tools/download_file |
HMAC | 60s | base64 only |
| POST | /admin/pending-usage |
X-Admin-Token | 30s | poll_loop |
| POST | /admin/mark-settled |
X-Admin-Token | 30s | idempotent |
| POST | /admin/mark-artifacts-registered |
X-Admin-Token | 30s | idempotent |
4.7 HTTP status -> exception mapping
| HTTP | error code | ComputeClient exception |
|---|---|---|
| 400 | validation | ComputeBadRequest |
| 401 | auth invalid/expired/replay | ComputeAuthError |
| 403 | ownership/object scope | ComputePermissionDenied |
| 404 | env missing | EnvironmentNotFound |
| 409 | input/storage conflict | InputObjectError / StorageBackendUnsupportedForCompute |
| 413 | size limit | ComputePayloadTooLarge |
| 429 | quota/concurrency | ComputeCapacityExceeded |
| 500 | internal | ComputeServiceError |
| 503 | unavailable | ComputeUnavailable |
| 507 | artifact/storage quota | ArtifactStorageExceeded |
ComputeUnavailable 必须设置 compute_client.ready=False,但仅对明确的 readiness/backend unavailable/network timeout 生效;普通命令失败不得把全局 ready 置 false。
5. shared_compute API 兼容决策
P0 不新增 create_job endpoint。EvoScientist 侧统一调用 /tools/create_environment,通过字段表达 shared intent:
{
"execution_mode": "simple",
"shared_compute_eligible": true,
"resource_class": "shared-small",
"backend_policy": "auto",
"input_objects": [
{
"file_id": "file_123",
"object_key": "threads/thr_1/files/file_123",
"sha256": "<sha256>",
"size_bytes": 1024,
"mount_path": "/workspace/data.csv",
"read_only": false
}
],
"artifact_prefix": "jobs/{thread_id}/{request_id}/artifacts/",
"default_output_paths": ["/workspace"],
"max_runtime_seconds": 60,
"budget_snapshot": {
"plan": "pro",
"quota_ok": true,
"remaining_minutes": 120,
"cash_balance_snapshot": "10.000000",
"overtime_price_per_minute": "0.050000",
"affordable_runtime_seconds": 19200,
"effective_max_runtime_seconds": 60,
"pricing_version": "2026-05-23"
},
"limits": {
"max_input_bytes": 10485760,
"max_output_bytes": 10485760,
"max_stdout_stderr_bytes": 1048576,
"max_files": 100
}
}
服务端返回仍必须包含 environment_id 兼容字段:
{
"environment_id": "env_or_job_handle_123",
"status": "running",
"execution_mode": "simple",
"backend_type": "shared_compute",
"resource_ref": "shared-worker-pool/default"
}
约束:
- 对 EvoScientist 和 deepagents 来说,该 handle 仍命名为 environment_id。
- 对 exec service 内部来说,可映射为 shared job handle。
- 后续 execute/read/write/destroy 仍传同一个 environment_id。
- shared_compute 不进入 warm pool 状态;这是 service 内部语义,Gateway 不直接判断。
- 若 service 判定 not eligible,可以:
- 返回
SHARED_COMPUTE_NOT_ELIGIBLE,Gateway 重新 standard create;或 - service 内部自动路由到 standard,并在 response 标记
execution_mode="standard"。
- 返回
- P0 推荐 Gateway 先保守判断,不满足条件直接 standard,减少二次请求。
6. deepagents BackendProtocol 映射
当前 EvoScientist/backends.py 使用 deepagents 返回类型:
ExecuteResponseWriteResultEditResultLsResultGrepResultGlobResultFileUploadResponseFileDownloadResponse
ContainerSandboxBackend 必须适配这些类型,不返回裸 dict 给 deepagents。
| deepagents backend method | ComputeClient method | 返回映射 |
|---|---|---|
execute(command, timeout=None) |
execute_command |
ExecuteResponse(output=stdout+stderr, exit_code=exit_code, truncated=truncated) |
read(file_path) |
read_file |
返回 content 字符串或错误包装 |
write(file_path, content) |
write_file |
WriteResult(error=None) |
edit(file_path, old_string, new_string, replace_all=False) |
edit_file |
EditResult(diff=<unified_diff>, error=None) |
ls(path) |
list_dir |
LsResult(entries=<list>) |
grep(pattern, path, glob=None) |
grep_files |
GrepResult(matches=<list>) |
glob(pattern, path="/") |
glob_files |
GlobResult(matches=<list>) |
upload_files(files) |
upload_file loop |
list[FileUploadResponse] |
download_files(paths) |
download_file loop |
list[FileDownloadResponse] |
destroy(persist_outputs=True, output_paths=<list>, force_destroy=False) |
destroy_environment |
dict for StreamHandler settlement |
实现注意:
- execute output 可先采用
stdout + stderr合并,metadata 中保留 stderr/exit_code/truncated。 - ComputeClient 是 async,但 deepagents BackendProtocol 当前是 sync。P0 需要选择一个明确适配策略。
6.1 async/sync 适配决策
推荐方案:将 StreamHandler._create_agent() 改为 async,但 ContainerSandboxBackend 内部仍需要 sync methods 供 deepagents 调用。
采用 anyio.from_thread / asyncio.run_coroutine_threadsafe 复杂度高。P0 更稳妥方案:
- 为
ContainerSandboxBackend创建一个专用 background event loop thread。 - backend sync method 调用
_run_async(coro)。 _run_async使用asyncio.run_coroutine_threadsafe(coro, self._loop).result(timeout=<seconds>)。- destroy 时关闭 loop thread。
伪代码:
class ContainerSandboxBackend:
def _run_async(self, coro, timeout: float | None = None):
if self._closed:
raise ComputeUnavailable("backend closed")
fut = asyncio.run_coroutine_threadsafe(coro, self._loop)
return fut.result(timeout=timeout or self.default_timeout + 5)
如果 deepagents 后续支持 async backend,可在 P1 移除 loop thread。
6.2 路径翻译
ContainerSandboxBackend 不再把 /workspace/foo 转成本地路径。它只做执行环境内路径规范化:
- relative
foo.py->/workspace/foo.py /workspace/foo.py->/workspace/foo.py/__global__/x->/__global__/x/threads/{id}/x->/threads/{id}/x- 禁止
/tmp、/Users、/etc、..等路径。
7. StreamHandler async 改造
7.1 决策
将 _create_agent() 改为 async:
async def _create_agent(self, checkpointer=None): pass
在 stream() 中:
agent = await self._create_agent(checkpointer=checkpointer)
原因:
- quota preflight 是 DB async。
- StorageService.build_input_manifest 是 DB + S3 async。
- ComputeClient readiness/admin check 是 HTTP async。
- 避免在 running event loop 中使用 blocking hack。
7.2 ComputeContext
新增内部 dataclass:
@dataclass
class ComputeContext:
enabled: bool
reason: str | None
input_objects: list[InputObject]
artifact_prefix: str
default_output_paths: list[str]
max_runtime_seconds: int
budget_snapshot: dict[str, Any]
execution_mode: str
resource_class: str
backend_policy: str
quota_result: ComputeQuotaResult | None
7.3 _prepare_compute_context 流程
1. 若 source != web:返回 disabled,但 CLI 不受影响。
2. 若 compute_client is None 或 not ready:返回 disabled reason=ComputeUnavailable。
3. 执行 _check_compute_quota_async。
4. quota_ok false:返回 disabled reason=ComputeQuotaExceeded。
5. storage_backend != s3:返回 disabled reason=StorageBackendUnsupportedForCompute。
6. 展开 manifest paths:
- /workspace
- /__global__
- 当前用户可读 peer thread 列表,显式 /threads/{id}
7. await storage_service.build_input_manifest(user_id=self.user_uid, thread_id=self.thread_id, paths=manifest_paths, require_remote_compute=True)
8. 生成 request_id。
9. artifact_prefix = jobs/{thread_id}/{request_id}/artifacts/
10. default_output_paths = ["/workspace", "/__global__"] 或按配置。
11. effective_max_runtime_seconds = min(requested, plan_hard_max, affordable_runtime_seconds)
12. 若 effective < 60:disabled reason=ComputeQuotaExceeded。
13. 生成 budget_snapshot。
14. decide_execution_mode()。
15. 返回 enabled ComputeContext。
7.4 _create_agent 调用
backend_ref: dict[str, Any] = {}
compute_ctx = await self._prepare_compute_context()
agent = create_cli_agent(
workspace_dir=self.workspace_dir,
memory_dir=self.memory_dir,
source="web",
user_id=self.user_uid,
thread_id=self.thread_id,
compute_quota_ok=compute_ctx.enabled,
remaining_compute_minutes=compute_ctx.quota_result.remaining_minutes if compute_ctx.quota_result else 0,
compute_client=self.compute_client if compute_ctx.enabled else None,
input_objects=compute_ctx.input_objects,
artifact_prefix=compute_ctx.artifact_prefix,
default_output_paths=compute_ctx.default_output_paths,
max_runtime_seconds=compute_ctx.max_runtime_seconds,
budget_snapshot=compute_ctx.budget_snapshot,
storage_backend=self.storage_service.backend_name,
execution_mode=compute_ctx.execution_mode,
resource_class=compute_ctx.resource_class,
backend_policy=compute_ctx.backend_policy,
_backend_ref=backend_ref,
model=self.model,
config=cfg,
checkpointer=checkpointer,
**filtered_params,
)
self._compute_backend = backend_ref.get("backend")
self._compute_context = compute_ctx
7.5 finally destroy/settle 顺序
1. 若没有 _compute_backend:只做非计算 fallback sync。
2. 调用 backend.destroy(persist_outputs=True, output_paths=default_output_paths, force_destroy=False)。
3. 若 destroy 抛 ComputeUnavailable:记录,依赖 exec service pending_usage;不本地执行。
4. 若返回 STOPPING_FAILED:发送用户可见 warning;不 settle。
5. 若 artifacts 非空:StorageService.register_artifacts。
6. register 成功:ComputeClient.mark_artifacts_registered。
7. register 失败:记录 artifact_registration_error;不 settle,等待 poll_loop。
8. 调用 settle_compute_usage。
9. settle InsufficientBalance:保留 pending_usage。
10. settle success/already_settled:mark_usage_settled 由 poll_loop 或当前路径完成。
11. 最后 set_thread_status idle。
8. create_cli_agent 改造
8.1 签名追加参数
在末尾追加,保持兼容:
thread_id: str = "",
compute_quota_ok: bool = False,
remaining_compute_minutes: int = 0,
compute_client: Any | None = None,
input_objects: list[Any] | None = None,
artifact_prefix: str = "",
default_output_paths: list[str] | None = None,
max_runtime_seconds: int | None = None,
budget_snapshot: dict[str, Any] | None = None,
storage_backend: str = "local",
execution_mode: str = "standard",
resource_class: str = "small",
backend_policy: str = "auto",
_backend_ref: dict | None = None,
8.2 Web backend 选择
if source == "web" and user_id:
if compute_client and compute_quota_ok and storage_backend == "s3":
ws_backend = ContainerSandboxBackend(
compute_client=compute_client,
user_id=user_id,
thread_id=thread_id,
input_objects=input_objects or [],
artifact_prefix=artifact_prefix,
default_output_paths=default_output_paths or ["/workspace"],
max_runtime_seconds=max_runtime_seconds,
budget_snapshot=budget_snapshot,
execution_mode=execution_mode,
resource_class=resource_class,
backend_policy=backend_policy,
)
if _backend_ref is not None:
_backend_ref["backend"] = ws_backend
else:
ws_backend = NonComputingFallbackBackend(
write_root=str(_thread_root),
read_roots=_peer_threads,
global_root=str(_user_global),
virtual_mode=True,
reason="ComputeUnavailable or storage backend unsupported",
)
else:
ws_backend = CustomSandboxBackend(root_dir=workspace_dir, virtual_mode=True, timeout=300)
CompositeBackend routes 保持:
- default -> ws_backend
/skills/-> local skills backend/memory/-> local memory backend
9. 数据库实施规格
9.1 user_balances
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_minutes_remaining INTEGER DEFAULT 0;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_minutes_used INTEGER DEFAULT 0;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_minutes_lifetime INTEGER DEFAULT 0;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_overtime_charged NUMERIC(12,6) DEFAULT 0;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_quota_monthly INTEGER DEFAULT 0;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_quota_refreshed_at TIMESTAMP NULL;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_quota_period_start TIMESTAMP NULL;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_quota_period_end TIMESTAMP NULL;
ALTER TABLE user_balances ADD COLUMN IF NOT EXISTS compute_quota_adjustment INTEGER DEFAULT 0;
金额迁移:
cash_balance、total_spent、recharge_records.charge_amount统一 NUMERIC(12,6)。- 迁移前查询 information_schema,避免重复 ALTER。
- Python 读写全部转 Decimal。
9.2 compute_usage_log
CREATE TABLE IF NOT EXISTS compute_usage_log (
id BIGINT GENERATED BY DEFAULT AS IDENTITY PRIMARY KEY,
user_id_int INTEGER REFERENCES users(id),
user_uid TEXT NOT NULL,
thread_id TEXT NOT NULL,
environment_id TEXT UNIQUE NOT NULL,
plan TEXT NOT NULL,
runtime_seconds DOUBLE PRECISION NOT NULL,
compute_minutes INTEGER NOT NULL,
quota_remaining INTEGER NULL,
overtime_charged NUMERIC(12,6) DEFAULT 0,
currency TEXT DEFAULT 'CNY',
pricing_snapshot TEXT NULL,
settlement_status TEXT DEFAULT 'pending',
settlement_error TEXT NULL,
created_at TIMESTAMP DEFAULT NOW()
);
状态语义:
| settlement_status | 含义 |
|---|---|
| pending | 记录存在但未获得结算权或等待重试 |
| claimed | 当前事务获得结算权但尚未完成扣费;正常情况下不应长期存在 |
| settled | 扣费、usage_log、wallet_ledger 全部提交 |
| failed | 最近一次结算失败,可重试 |
9.3 user_files
使用第 3.4 节修正后的 namespace unique indexes。
9.4 wallet_ledger 兼容策略
如果现有 wallet_ledger 使用 cents 字段,P0 采用兼容但单一真实来源策略:
- 余额真实来源仍为
user_balances.cash_balance NUMERIC(12,6)。 - compute ledger 写入:
entry_type='compute_overtime'source_type='compute_usage'source_id=environment_iddirection='debit'amount=Decimal overtime_amount- 若表只有 amount_cents,则写
amount_cents = int(amount * 100),但 pricing_snapshot 必须保留 Decimal 原值。
- 不允许同一次 compute 同时写 amount 和 amount_cents 两套互相独立的余额。
10. compute billing 原子事务规格
10.1 preflight quota
quota_ok = (
plan != "starter"
and (
compute_minutes_remaining > 0
or cash_balance >= overtime_price_per_minute
)
)
affordable_runtime_seconds = (
compute_minutes_remaining + floor(cash_balance / overtime_price_per_minute)
) * 60
边界:
- overtime_price_per_minute <= 0 或 NULL:PricingConfigError,compute enabled startup fail。
- remaining_minutes < 0 或 cash_balance < 0:记录 ERROR,quota_ok=False。
- 余额不足 1 分钟:quota_ok=False。
- max_runtime_seconds 必须裁剪为
min(requested, plan_hard_max, affordable_runtime_seconds)。
10.2 settle_compute_usage 事务
输入:
async def settle_compute_usage(
*,
environment_id: str,
user_uid: str,
thread_id: str,
runtime_seconds: float,
plan: str,
pricing_snapshot: dict,
artifacts_registered: bool,
) -> SettlementResult: pass
计费分钟:
compute_minutes = max(1, math.ceil(runtime_seconds / 60))
事务规则:
1. 使用 get_transaction_connection() 获取独立 asyncpg connection。
2. INSERT compute_usage_log(environment_id, user_uid, thread_id, runtime_seconds, compute_minutes, settlement_status='claimed')
ON CONFLICT(environment_id) DO NOTHING。
3. 若 INSERT 未获得行:
a. SELECT existing settlement_status。
b. status='settled' -> return already_settled。
c. status='claimed' 且 updated/created 很新 -> raise SettlementInProgress。
d. status='failed' 或长期 claimed -> 可抢占重试,更新为 claimed。
4. 查询 user_balances FOR UPDATE。
5. free_to_use = min(compute_minutes_remaining, compute_minutes)。
6. overtime_minutes = compute_minutes - free_to_use。
7. overtime_amount = Decimal(overtime_minutes) * overtime_price_per_minute。
8. 若 cash_balance < overtime_amount:
抛 InsufficientComputeBalance,整个事务回滚。
9. UPDATE user_balances:
compute_minutes_remaining -= free_to_use
cash_balance -= overtime_amount
compute_minutes_used += compute_minutes
compute_minutes_lifetime += compute_minutes
compute_overtime_charged += overtime_amount
10. INSERT wallet_ledger with idempotency_key compute:{user_uid}:{thread_id}:{environment_id}。
11. UPDATE compute_usage_log SET settlement_status='settled', quota_remaining, overtime_charged, pricing_snapshot。
12. commit 后返回 settled。
禁止:
- 先插入 claim 并提交,再异步扣款。
- float 参与金额计算。
- 余额不足时保留 claimed usage_log。
- artifact 未注册且未达到 retry 上限时结算。
10.3 artifact 注册失败与结算
规则:
- pending_usage 中有 output_manifest_json 且 artifacts_registered=false 时,必须先注册 artifact。
- 注册失败未达到
ARTIFACT_REGISTRATION_MAX_RETRIES=10:不 settle。 - 达到上限后:标记
artifacts_unregistered=true或在本地记录 artifact_registration_failed,然后允许 compute 结算,避免无限免单。
11. pending_usage poll_loop 详细规格
11.1 启动条件
- compute enabled。
- ComputeClient.ready=True。
- admin token 存在。
check_poll_available()成功。
若 admin endpoint 不可用:启动延迟重试任务,不允许长期无兜底结算运行。
11.2 参数
| 参数 | 默认值 |
|---|---|
| poll interval | 60s |
| batch size | 50 |
| per-record timeout | 30s |
| consecutive failure threshold | 5 |
| backoff after threshold | 600s |
| artifact registration max retries | 10 |
11.3 单轮流程
1. records = compute_client.list_pending_usage(admin_token, settled=0, limit=50)
2. HTTP 错误:failure_count += 1;达到 5 次 sleep 600s。
3. 对每条 record:
a. 若 output_manifest_json 非空且 artifacts_registered=false:
i. 解析 artifacts。
ii. storage_service.register_artifacts(user_id, thread_id, artifacts)。
iii. 成功 -> mark_artifacts_registered(usage_id=record.id)。
iv. 失败 -> 记录 error,递增 registration retry;未达上限则 continue 下一条。
v. 达上限 -> 标记 artifacts_unregistered,允许继续 settle。
b. 调用 settle_compute_usage。
c. settled 或 already_settled -> mark_usage_settled([record.id])。
d. InsufficientComputeBalance -> 保留 pending_usage,等待充值。
e. SettlementInProgress -> 跳过,下一轮重试。
f. 其他错误 -> 记录 ERROR,下一条。
4. sleep 60s。
11.4 并发安全
- StreamHandler finally 与 poll_loop 可能同时处理同一 environment_id。
compute_usage_log.environment_id UNIQUE+ 单事务 claim 保证只扣一次。mark_usage_settled必须幂等。mark_artifacts_registered必须幂等。
12. Gateway lifespan 改造
启动顺序:
1. 加载 CLI config/env。
2. init_gateway_db()。
3. db = await get_connection()。
4. init StorageService。
5. 若 STORAGE_BACKEND=s3:HeadBucket + probe put/get/delete。
6. 启动 storage_cleanup_loop。
7. 读取 EXECUTION_SERVICE_URL。
8. URL 未设置:compute_client=None,不启动 poll_loop。
9. URL 已设置:校验 HMAC secret/admin token,缺失则 startup fail。
10. init ComputeClient。
11. await compute_client.check_ready()。
12. ready true:check_poll_available,启动 poll_loop。
13. ready false:admin poll 延迟重试,Web compute unavailable。
14. 启动 refresh_expired_compute_quotas loop。
15. app.state 写入 db/storage_service/compute_client/admin_token/task handles。
16. yield。
17. shutdown:cancel tasks -> close compute_client -> close storage_service -> close_gateway_db。
/ready 失败不导致 Gateway 必然启动失败;但 Web compute 请求必须失败为 ComputeUnavailable。
13. 文件路由改造
13.1 uploads.py
上传流程:
- require auth。
- require owned thread。
- read upload bytes with max size limit。
- storage_service.put_file(owner_user_id=user_uid, logical_path=logical_path, data=file_bytes, thread_id=thread_id, source='upload')。
- 返回 file_id/logical_path/sha256/size/storage_backend。
13.2 user_files.py / files.py / global_uploads.py
- list:查询 user_files active records。
- download:StorageService.open_read。
- delete:StorageService.delete_file soft delete。
- overwrite:StorageService.put_file 生成新 version,旧 active deleted。
13.3 knowledge_indexer.py / knowledge.py
- 不再直接
thread_data_dir(user_uid, thread_id) / virtual_path读取。 - 通过 user_files 查询 file_id。
- 通过 StorageService.open_read 传给 parser。
- LocalStorageBackend 路径只作为 backend 内部细节。
14. 非计算降级
14.1 NonComputingFallbackBackend
包装 MultiRootSandboxBackend,但 execute 永远失败:
class NonComputingFallbackBackend(MultiRootSandboxBackend):
def execute(self, command: str, *, timeout: int | None = None) -> ExecuteResponse:
return ExecuteResponse(
output="ComputeUnavailable: remote compute is unavailable; local execution is disabled for Web threads.",
exit_code=1,
truncated=False,
)
文件操作允许范围:
- read/list/download existing files。
- write/edit only in controlled recovery workspace if UI 明确处于非计算降级。
- 不允许任何 tool 通过文件操作触发代码执行。
14.2 fallback sync
StreamHandler cleanup:
- 记录 stream_start_time。
- cleanup 时扫描 recovery workspace 中 mtime > stream_start_time 的文件。
- 对每个文件调用 storage_service.put_file。
- 全部成功 -> 删除 recovery workspace。
- 任一失败 -> 移动到 controlled recovery/orphan dir,写 fallback_sync_recovery。
- cleanup loop 后续重试。
15. shared_compute 路由
15.1 保守判定
进入 simple/shared 的必要条件:
EXECUTION_SHARED_COMPUTE_ENABLED=true- plan 允许 compute
- storage_backend=s3
- input_objects 总大小 <= 10MiB
- 文件数 <= 100
- effective max_runtime_seconds <= 60
- 不需要 GPU
- 不需要自定义镜像
- 不需要系统依赖安装
- 不需要 interactive session
- 不需要服务进程
- 不需要持久 workspace
- 无强隔离标记
否则 standard。
15.2 limits
Gateway 下发 P0 shared limits:
{
"max_runtime_seconds": 60,
"max_input_bytes": 10485760,
"max_output_bytes": 10485760,
"max_stdout_stderr_bytes": 1048576,
"max_files": 100,
"max_user_concurrency": 2
}
15.3 失败语义
- 执行前 not eligible:standard 或明确 SharedComputeNotEligible。
- 执行中超限:失败,不半路迁移。
- worker degraded:不进入 shared。
- 任何 shared 失败不得 fallback Gateway 本地。
16. 部署方案
16.1 .env.example 增加
# Storage
STORAGE_BACKEND=s3
STORAGE_S3_ENDPOINT=http://127.0.0.1:9000
STORAGE_S3_BUCKET=evoscientist
STORAGE_S3_REGION=us-east-1
STORAGE_S3_ACCESS_KEY=minioadmin
STORAGE_S3_SECRET_KEY=<minio-or-s3-secret-key>
STORAGE_S3_FORCE_PATH_STYLE=true
STORAGE_MAX_FILE_SIZE_BYTES=1073741824
STORAGE_RETENTION_DAYS=0
# Remote compute
EXECUTION_SERVICE_URL=http://127.0.0.1:9020
EXECUTION_HMAC_SECRET=<generated-32-byte-hex>
EXECUTION_ADMIN_TOKEN=<generated-admin-token>
EXECUTION_DEFAULT_RESOURCE_CLASS=small
EXECUTION_DEFAULT_BACKEND_POLICY=auto
EXECUTION_SHARED_COMPUTE_ENABLED=true
EXECUTION_SIMPLE_MAX_RUNTIME_SECONDS=60
EXECUTION_PLAN_HARD_MAX_RUNTIME_SECONDS=7200
16.2 本地启动顺序
# 1. PostgreSQL
# existing local command
# 2. MinIO
docker run -p 9000:9000 -p 9001:9001 \
-e MINIO_ROOT_USER=minioadmin \
-e MINIO_ROOT_PASSWORD=<local-minio-password> \
quay.io/minio/minio server /data --console-address ':9001'
# 3. Create bucket
mc alias set local http://127.0.0.1:9000 minioadmin minioadmin
mc mb --ignore-existing local/evoscientist
# 4. k8s-exec-service-mcp
cd /Users/m4/Projects/EvoSci/MCP/k8s-exec-service-mcp
# run service per its README
# 5. EvoScientist Gateway
cd /Users/m4/Projects/EvoSci/EvoScientist
source venv/bin/activate
# export STORAGE_* EXECUTION_*
# run gateway
16.3 smoke test
Gateway 侧 smoke test 必须验证:
- StorageService S3 probe success。
GET EXECUTION_SERVICE_URL/ready200。- 上传小 CSV -> user_files active row。
- build_input_manifest 返回 s3 object。
- create_environment standard。
- execute
python --version。 - write/read file。
- destroy persist artifact。
- register_artifacts row。
- settle_compute_usage row + wallet_ledger。
- pending_usage mark-settled。
- 停止 exec service 后 Web compute 不执行本地命令。
17. 实施阶段
Phase 1:Storage + ComputeClient 骨架
- user_files migration + corrected indexes。
- StorageService + LocalStorageBackend。
- S3StorageBackend + MinIO probe。
- migrate_local_files_to_storage dry-run。
- ComputeClient HMAC + golden vector。
- ComputeClient ready/health/admin readiness。
- env/config wiring。
- tests: storage + hmac + ready。
Gate:StorageService + MinIO + HMAC 双端测试通过。
Phase 2:Remote Backend 最小闭环
- ContainerSandboxBackend sync adapter。
- NonComputingFallbackBackend。
- create_cli_agent 参数与 backend selection。
- StreamHandler async prepare/create。
- threads.py app_state 注入。
- ComputeClient create/execute/read/write/destroy wrappers。
- happy path E2E:upload -> manifest -> execute -> artifact。
Gate:ComputeUnavailable 不本地执行;remote happy path 通过。
Phase 3:Billing + poll_loop
- compute pricing config。
- user_balances compute columns。
- compute_usage_log。
- get_transaction_connection。
- settle_compute_usage atomic transaction。
- wallet_ledger compatibility。
- poll_loop。
- artifact registration retry。
Gate:重复 environment_id 不重复扣费;artifact 注册失败不会免费;余额不足保留 pending_usage。
Phase 4:File API 全量迁移 + fallback_sync
- uploads.py。
- user_files.py/files.py/global_uploads.py。
- knowledge_indexer.py/knowledge.py。
- fallback_sync_recovery。
- cleanup loops。
- migration script full run。
Gate:Web 文件 API 不直接依赖 thread_data_dir。
Phase 5:shared_compute + 部署验收
- execution_mode decision。
- shared limits。
- create_environment shared-compatible payload。
- standard fallback。
- local smoke test。
- kind smoke test。
- docs update。
18. 测试清单
18.1 单元测试
- normalize_logical_path rejects traversal/control/windows/absolute。
- user_files indexes allow same logical_path in different threads。
- StorageService put_file object-first DB-second。
- build_input_manifest workspace/global/thread readonly。
- build_input_manifest local backend rejection。
- register_artifacts idempotent by environment_id + logical_path。
- ComputeClient canonical_json exact。
- HMAC golden vector。
- ComputeClient status -> exception mapping。
- ContainerSandboxBackend method mapping to deepagents return types。
- NonComputingFallbackBackend execute disabled。
- preflight quota boundary cases。
- settle mixed free+cash Decimal。
- settle rollback on insufficient balance。
- poll_loop artifact-first settlement。
18.2 集成测试
- Gateway startup compute disabled。
- Gateway startup fails when URL set but secret missing。
- /ready false -> compute unavailable。
- upload -> user_files -> S3 object。
- StreamHandler async create -> remote execute。
- destroy -> artifacts -> register -> settle。
- artifact registration failure -> no settle until retry。
- duplicate settlement -> once only。
- STORAGE_BACKEND=local -> compute blocked。
- shared eligible -> create payload execution_mode simple。
- shared not eligible -> standard。
18.3 E2E
- CSV analysis artifact downloadable。
- MinIO down -> no environment created。
- exec service down -> no local command execution。
- balance only 1 minute -> max_runtime_seconds=60。
- force_destroy artifacts_lost visible。
- fallback sync failure -> recovery row -> retry success。
- kind deployment smoke test。
19. 验收标准
P0 完成标准:
- Web 用户代码执行 100% 通过 k8s-exec-service-mcp。
- Gateway 本地不执行 Web 用户 shell/Python/subprocess。
- S3/MinIO 是 Web remote compute 的唯一文件内容底座。
- LocalStorageBackend 只用于 CLI、迁移、非计算文件管理。
- StorageService path normalization、input manifest、artifact registration 完整。
- ComputeClient HMAC golden vector 双端一致。
- ContainerSandboxBackend 与 deepagents 返回类型兼容。
- StreamHandler async preflight + destroy-settle 闭环完整。
- compute billing 原子事务不产生半结算。
- pending_usage poll_loop 可补注册 artifact 并兜底结算。
- shared_compute 通过 create_environment 兼容字段接入,不新增 create_job。
user_files支持不同 thread 同名文件。- k8s-exec-service-mcp 不可用时 Web compute 返回 ComputeUnavailable。
- 本地单机 smoke test 和 kind smoke test 均通过。
20. 风险与硬性禁止项
硬性禁止:
- 禁止 Web compute fallback 到
CustomSandboxBackend。 - 禁止 Web compute fallback 到
LocalShellBackend。 - 禁止把 Gateway 本地路径传给 k8s-exec-service-mcp。
- 禁止 object_key 作为权限依据。
- 禁止 float 参与 compute 扣款。
- 禁止 artifact 未注册且未达 retry 上限时直接结算。
- 禁止新增未在两端协议中确认的
/tools/create_job。 - 禁止 k8s-exec-service-mcp 多副本写同一个 SQLite。
主要风险:
| 风险 | 处理 |
|---|---|
| 本地路径引用遗漏 | 全仓搜索 thread_data_dir, user_data_dir, ~/.evoscientist, data/ |
| cents/Decimal 账本冲突 | 明确 user_balances.cash_balance 为真实余额,ledger 只审计 |
| async/sync backend 复杂 | P0 用 backend 专用 loop thread;P1 评估 async backend |
| HMAC 双端不一致 | golden vector 作为 Gate Check |
| shared 安全误判 | 保守路由,不确定即 standard/isolated |
| artifact 注册失败免单 | pending_usage retry,上限后 artifacts_unregistered 仍结算 |
21. 多轮评审结果与实施就绪判定
21.1 R1 评审发现与修订
| 问题 | 严重级别 | 修订结果 |
|---|---|---|
| 文档头部保存路径仍指向旧目录 | P1 | 已改为 EvoScientist/docs 路径 |
env 样例中 EXECUTION_HMAC_SECRET 行损坏 |
P0 | 已修复为 <generated-32-byte-hex> 与 <generated-admin-token> |
shared_compute 中出现未定义 create_job 风险 |
P0 | 已明确 P0 不新增 create_job,统一走 create_environment 兼容 handle |
| HMAC 只有算法无固定可测 token | P0 | 已补真实计算 golden vector |
| placeholder 过多 | P1 | 接口片段改为 pass 或具名占位符 |
| user_files 唯一性可能阻塞不同线程同名文件 | P0 | 已采用 global/thread/project 三类 partial unique index |
| StreamHandler async 边界不清 | P0 | 已明确 _create_agent 改 async,backend 内部用 loop thread 适配 sync protocol |
21.2 R2 复评结论
复评结果:符合实施要求。当前文档已经满足以下条件:
- 每个 P0 模块均有目标文件、接口、核心算法和测试 Gate。
- StorageService、ComputeClient、ContainerSandboxBackend、compute billing、poll_loop 的关键状态和失败路径已定义。
- shared_compute 不引入未确认 endpoint,避免两仓协议分叉。
- Web 本地执行禁令被落实到 create_cli_agent、NonComputingFallbackBackend、测试和硬性禁止项。
- 数据库 migration 包含 user_files namespace 修正和 compute 结算状态。
- 可以按 Phase 1-5 拆分 PR,并以各 Phase Gate 作为合并条件。
21.3 PR 拆分建议
| PR | 范围 | 合并 Gate |
|---|---|---|
| PR-1 | DB migration: user_files、compute_usage_log、user_balances compute columns | migration 幂等测试通过 |
| PR-2 | StorageService + Local/S3 backend + MinIO probe | storage unit + MinIO integration 通过 |
| PR-3 | ComputeClient + HMAC + ready/health/admin wrappers | golden vector 双端一致 |
| PR-4 | ContainerSandboxBackend + NonComputingFallbackBackend | deepagents backend mapping 测试通过 |
| PR-5 | StreamHandler async + create_cli_agent 参数接入 | ComputeUnavailable 不本地执行 |
| PR-6 | Billing settlement + wallet_ledger + poll_loop | 重复结算只扣一次 |
| PR-7 | Upload/files/knowledge 路由迁移 StorageService | 全仓 Web 文件 API 不直接依赖 thread_data_dir |
| PR-8 | shared_compute routing + smoke tests + docs | local/kind smoke 通过 |
22. 最小可执行切片
第一周最小闭环:
- user_files 表 + corrected indexes。
- StorageService + S3StorageBackend + MinIO probe。
- ComputeClient HMAC + /ready。
- ContainerSandboxBackend execute/read/write/destroy happy path。
- StreamHandler async create remote backend。
- Web 上传小文件 -> build_input_manifest -> remote
python --version-> destroy -> artifact 入库。 - 关闭 exec service 后确认不会调用 CustomSandboxBackend/LocalShellBackend。
该切片通过后,再进入 billing/poll_loop/shared_compute 全量实现。