Files
EvoScientist/docs/execution-engine/execution-engine-deployment.md
T
m4 c2743251e9 Initial commit of EvoScientist framework
Self-evolving AI scientist framework built on LangGraph/LangChain with
CLI/TUI core, FastAPI gateway, and Next.js frontend.

Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
2026-07-13 08:07:45 +08:00

620 lines
16 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# EvoScientist 执行引擎 — 部署运维文档
> 基于 SWE-ReX 架构模式(Docker 本地 + Modal 云端双模式)
## 1. 概述
EvoScientist 执行引擎为实验代码提供安全隔离的运行环境。核心思路:将代码执行从 Agent 主进程中剥离,放入独立的 Docker 容器或 Modal 云端 Sandbox 中运行,通过 HTTP API 交互。
### 1.1 架构总览
```
┌──────────────────┐
│ Gateway/API │
│ (FastAPI 8065) │
└────────┬─────────┘
│
┌────────▼─────────┐
│ Execution Engine │
│ (调度核心) │
└──┬────────────┬───┘
│ │
┌────────────▼──┐ ┌────▼────────────┐
│ DockerProvider│ │ ModalProvider │
│ (本地容器池) │ │ (云端GPU弹性) │
└───────┬───────┘ └──────┬──────────┘
│ │
┌────────▼──────┐ ┌────────▼──────┐
│ Container ×N │ │ Modal Sandbox │
│ evosci-exec │ │ (GPU 可用) │
│ (FastAPI) │ │ (按需启停) │
└───────┬───────┘ └────────┬──────┘
│ │
┌────▼────────────────────▼────┐
│ 共享存储 (NFS/S3/本地) │
│ ~/.evoscientist/data/ │
│ {user_id}/{thread_id}/ │
└──────────────────────────────┘
```
### 1.2 核心组件
| 组件 | 职责 | 运行位置 |
|------|------|---------|
| Execution Engine | 任务调度、路由、容器生命周期管理 | Gateway 进程内 |
| evosci-exec-server | 容器内 FastAPI 服务,接收并执行命令 | 每个执行容器内 |
| DockerProvider | 管理本地 Docker 容器池 | Gateway 进程内 |
| ModalProvider | 管理 Modal 云端 Sandbox | Gateway 进程内 |
| 执行容器镜像 | 预装科学计算库的隔离环境 | Docker Hub / 本地构建 |
### 1.3 数据流
```
1. 用户发消息 → Gateway
2. Agent 调用 execute 工具
3. ExecutionEngine.submit(task)
4. 路由判断:
- LIGHT 任务 → DockerProvider → 启动/复用容器 → HTTP /execute
- HEAVY 任务 → ModalProvider → 创建 Sandbox → HTTP /execute
5. 容器内 evosci-exec-server 执行命令
6. 结果原路返回 → Agent → 用户
```
---
## 2. 环境要求
### 2.1 硬件要求
| 场景 | CPU | 内存 | 磁盘 | GPU |
|------|-----|------|------|-----|
| 开发/测试 | 4 核 | 8 GB | 50 GB | 无 |
| 生产(轻量) | 8 核 | 32 GB | 200 GB | 无 |
| 生产(GPU) | 16 核 | 64 GB | 500 GB | 1× A100/T4 |
### 2.2 软件要求
| 依赖 | 版本 | 用途 | 必需 |
|------|------|------|------|
| Docker Engine | ≥ 24.0 | 容器运行时 | 是 |
| Docker Compose | ≥ 2.20 | 服务编排 | 是 |
| Python | ≥ 3.11 | 后端运行时 | 是 |
| PostgreSQL | ≥ 15 | 主数据库(已有) | 是 |
| Redis | ≥ 7.0 | 任务队列(P3 阶段) | P3 |
| NVIDIA Container Toolkit | ≥ 1.14 | GPU 容器支持 | GPU 场景 |
| Modal CLI | ≥ 0.6 | 云端执行 | Modal 场景 |
### 2.3 网络要求
```
端口规划:
8065 — Gateway API (已有)
8880 — evosci-exec-server (容器内部, 不对外暴露)
5432 — PostgreSQL (已有)
6379 — Redis (P3 阶段)
Docker 网络:
evosci-net — 执行容器与 Gateway 的通信网络
容器内网络可设为 none (无外网) 或 bridge (有外网)
```
---
## 3. 部署步骤
### 3.1 构建执行容器镜像
```bash
cd ~/Projects/EvoSci/EvoScientist
# 构建轻量版 (CPU, 无 GPU)
docker build -t evosci-exec:latest \
-f docker/evosci-exec/Dockerfile .
# 构建 GPU 版 (需要 NVIDIA Container Toolkit)
docker build -t evosci-exec-gpu:latest \
-f docker/evosci-exec/Dockerfile.gpu .
# 验证镜像
docker run --rm evosci-exec:latest python -c "import numpy; print('OK')"
```
**Dockerfile** — 轻量版:
```dockerfile
# docker/evosci-exec/Dockerfile
FROM python:3.11-slim AS base
RUN apt-get update && apt-get install -y --no-install-recommends \
gcc g++ git curl wget ca-certificates \
&& rm -rf /var/lib/apt/lists/*
RUN pip install --no-cache-dir \
numpy pandas matplotlib scipy scikit-learn \
httpx fastapi uvicorn pydantic
COPY execution/server.py /opt/evosci/evosci-exec-server.py
WORKDIR /workspace
EXPOSE 8880
HEALTHCHECK --interval=10s --timeout=3s --retries=3 \
CMD curl -f http://localhost:8880/is_alive || exit 1
CMD ["python", "/opt/evosci/evosci-exec-server.py", "--port", "8880"]
```
**Dockerfile** — GPU 版:
```dockerfile
# docker/evosci-exec/Dockerfile.gpu
FROM nvidia/cuda:12.4.0-runtime-ubuntu22.04 AS base
RUN apt-get update && apt-get install -y --no-install-recommends \
python3 python3-pip python3-dev gcc g++ git curl wget \
&& rm -rf /var/lib/apt/lists/*
RUN pip3 install --no-cache-dir \
numpy pandas matplotlib scipy scikit-learn \
torch --index-url https://download.pytorch.org/whl/cu124 \
httpx fastapi uvicorn pydantic
COPY execution/server.py /opt/evosci/evosci-exec-server.py
WORKDIR /workspace
EXPOSE 8880
CMD ["python3", "/opt/evosci/evosci-exec-server.py", "--port", "8880"]
```
### 3.2 配置执行引擎
在 `settings.yaml` 中新增 `execution` 段:
```yaml
# settings.yaml
execution:
# 全局开关
enabled: true
# 默认执行提供者: docker | modal | auto
# auto: 轻量任务走 Docker, GPU 任务走 Modal
provider: auto
docker:
# 镜像名称
image: evosci-exec:latest
gpu_image: evosci-exec-gpu:latest
# 容器池限制
max_containers: 20 # 最大并发容器数
idle_timeout: 1800 # 空闲容器回收 (秒)
startup_timeout: 60 # 容器启动超时 (秒)
# 默认资源
default_cpu: 2 # 每容器 CPU 核数
default_memory: 4g # 每容器内存
max_cpu: 8 # 单容器上限
max_memory: 16g
# 网络
network_mode: none # none (隔离) | bridge (可联网)
# 清理
remove_on_stop: true # 容器停止后自动删除
volume_cleanup: true # 清理容器卷
modal:
# Modal App 名称
app_name: evoscientist
# 镜像
base_image: python:3.11-slim
# 资源
default_timeout: 3600 # Sandbox 默认超时 (秒)
deployment_timeout: 7200 # 部署最大存活 (秒)
default_cpu: 4
default_memory: "16GiB"
gpu_type: A100 # A100 | T4 | L4
max_gpu: 4
# 沙箱额外参数
sandbox_kwargs: {}
# 用户配额
quota:
# 默认用户 (普通套餐)
default:
max_concurrent: 2 # 最大并发容器
max_cpu: 4
max_memory: "8g"
max_gpu: 0
daily_budget_minutes: 100 # 每日执行总时长上限
max_task_timeout: 1800 # 单任务超时上限 (秒)
# 高级用户 (premium 套餐)
premium:
max_concurrent: 5
max_cpu: 8
max_memory: "32g"
max_gpu: 2
daily_budget_minutes: 500
max_task_timeout: 7200
# 任务调度
scheduler:
queue_type: memory # memory (初版) | redis (P3)
priority_levels: 3 # 优先级级数
retry_on_failure: true
max_retries: 2
```
### 3.3 Docker Compose 更新
```yaml
# docker/docker-compose.yml — 完整版
services:
backend:
image: ${BACKEND_IMAGE:-ghcr.io/jakeyang886/evoscientist-backend:latest}
build:
context: ../
dockerfile: docker/Dockerfile
args:
APP_VERSION: "${APP_VERSION:-0.0.0}"
ports:
- "8065:8065"
env_file:
- ../.env
environment:
- CORS_ORIGINS=${CORS_ORIGINS:-http://localhost:3065}
- BASE_URL=${BASE_URL:-http://localhost:3065}
- GATEWAY_HOST=0.0.0.0
- GATEWAY_PORT=8065
- EVOSCIENTIST_HOME=/app/.data
volumes:
- ../.data:/app/.data
# Docker-in-Docker: 让 backend 能管理执行容器
- /var/run/docker.sock:/var/run/docker.sock
# 镜像构建缓存
- evosci-images:/var/lib/docker
depends_on:
redis:
condition: service_healthy
restart: unless-stopped
# 资源限制 (backend 本身)
deploy:
resources:
limits:
cpus: "4"
memory: 8G
frontend:
image: ${FRONTEND_IMAGE:-ghcr.io/jakeyang886/evoscientist-frontend:latest}
build:
context: ../
dockerfile: docker/Dockerfile.frontend
args:
INTERNAL_GATEWAY_URL: http://backend:8065
NEXT_PUBLIC_GATEWAY_URL: ""
ports:
- "3065:3065"
environment:
- PORT=3065
- INTERNAL_GATEWAY_URL=http://backend:8065
depends_on:
backend:
condition: service_healthy
restart: unless-stopped
redis:
image: redis:7-alpine
ports:
- "6379:6379"
volumes:
- redis-data:/data
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 5s
timeout: 3s
retries: 5
restart: unless-stopped
volumes:
evosci-images:
redis-data:
```
### 3.4 环境变量
在 `.env` 中新增:
```bash
# .env
# === 执行引擎 ===
EVOSCIENTIST_EXECUTION_ENABLED=true
EVOSCIENTIST_EXECUTION_PROVIDER=auto
# Docker
EVOSCIENTIST_DOCKER_IMAGE=evosci-exec:latest
EVOSCIENTIST_DOCKER_MAX_CONTAINERS=20
# Modal (云端 GPU)
MODAL_TOKEN_ID=ak-xxx # Modal 凭证
MODAL_TOKEN_SECRET=as-xxx
# 用户配额
EVOSCIENTIST_DEFAULT_MAX_CONCURRENT=2
EVOSCIENTIST_PREMIUM_MAX_CONCURRENT=5
```
### 3.5 验证部署
```bash
# 1. 构建镜像
docker build -t evosci-exec:latest -f docker/evosci-exec/Dockerfile .
# 2. 启动所有服务
docker compose -f docker/docker-compose.yml up -d
# 3. 检查服务状态
curl http://localhost:8065/health
# 4. 测试执行引擎
# 通过 API 创建一个执行任务
curl -X POST http://localhost:8065/api/execution/submit \
-H "Authorization: Bearer <token>" \
-H "Content-Type: application/json" \
-d '{
"command": "python -c \"print(1+1)\"",
"timeout": 30
}'
# 5. 查看执行容器
docker ps --filter "name=evosci-"
# 6. 查看容器日志
docker logs evosci-<container-id>
```
---
## 4. 运维操作
### 4.1 容器池管理
```bash
# 查看所有执行容器
docker ps --filter "name=evosci-" --format "table {{.Names}}\t{{.Status}}\t{{.CreatedAt}}"
# 清理所有空闲容器
docker stop $(docker ps -q --filter "name=evosci-session-") 2>/dev/null
docker rm $(docker ps -aq --filter "name=evosci-session-") 2>/dev/null
# 清理退出的容器
docker container prune --filter "label=evosci-exec"
# 查看资源使用
docker stats --no-stream --format "table {{.Name}}\t{{.CPUPerc}}\t{{.MemUsage}}" \
$(docker ps -q --filter "name=evosci-")
```
### 4.2 镜像管理
```bash
# 构建新版本镜像
docker build -t evosci-exec:v1.1.0 -f docker/evosci-exec/Dockerfile .
# 更新运行中的镜像 (滚动更新)
# 1. 更新 settings.yaml 中的 image 字段
# 2. 重启 Gateway
docker compose -f docker/docker-compose.yml restart backend
# 清理旧镜像
docker image prune --filter "label=evosci-exec" --filter "until=168h"
```
### 4.3 GPU 支持 (本地)
```bash
# 检查 NVIDIA Container Toolkit
nvidia-ctk --version
# 测试 GPU 容器
docker run --rm --gpus all evosci-exec-gpu:latest \
python -c "import torch; print(torch.cuda.is_available())"
# 在 settings.yaml 中设置
# docker.gpu_image: evosci-exec-gpu:latest
# docker.default_gpu: 1 # 默认分配 GPU 数
```
### 4.4 Modal 云端配置
```bash
# 安装 Modal CLI
pip install modal
# 登录
modal profile create
# 验证连接
modal app list
# 查看 Sandbox 使用情况
modal sandbox list
# 设置凭证 (在 .env 中)
MODAL_TOKEN_ID=ak-xxx
MODAL_TOKEN_SECRET=as-xxx
```
### 4.5 监控指标
```bash
# 通过 Gateway API 获取执行引擎状态
curl http://localhost:8065/api/execution/status
# 返回示例:
# {
# "docker": {
# "running_containers": 5,
# "idle_containers": 2,
# "total_cpu_used": "8/32 cores",
# "total_memory_used": "20/64 GB"
# },
# "modal": {
# "active_sandboxes": 1,
# "gpu_hours_today": 2.5
# },
# "queue": {
# "pending_tasks": 3,
# "running_tasks": 5,
# "completed_today": 127
# }
# }
```
---
## 5. 故障排查
### 5.1 常见问题
| 问题 | 原因 | 解决 |
|------|------|------|
| 容器启动超时 | 镜像未拉取 / 磁盘不足 | `docker pull evosci-exec:latest`; 清理磁盘 |
| `docker.sock` 权限不足 | 后端容器无 Docker 访问权限 | 检查 socket 挂载和用户组 |
| 容器内命令超时 | 任务复杂度超预期 | 增大 `timeout` 参数; 使用 Modal |
| 内存 OOM | 容器内存不足 | 增大 `default_memory`; 检查内存泄漏 |
| Modal 连接失败 | 凭证过期 / 网络不通 | 刷新 `MODAL_TOKEN_*`; 检查防火墙 |
| 容器池满 | 并发用户超过 `max_containers` | 增大限制; 检查空闲容器是否被回收 |
### 5.2 日志查看
```bash
# Gateway 日志
docker compose -f docker/docker-compose.yml logs -f backend
# 特定执行容器日志
docker logs -f evosci-session-<user-id>-<thread-id>
# evosci-exec-server 日志 (容器内)
docker exec evosci-xxx cat /var/log/evosci-exec.log
```
### 5.3 紧急操作
```bash
# 强制停止所有执行容器
docker stop $(docker ps -q --filter "label=evosci-exec")
# 清理所有执行容器和镜像
docker rm -f $(docker ps -aq --filter "name=evosci-")
docker rmi $(docker images -q --filter "label=evosci-exec")
# 禁用执行引擎 (回退到旧模式)
# 在 settings.yaml 中设置:
# execution.enabled: false
```
---
## 6. 安全注意事项
### 6.1 容器隔离
- **网络隔离**: 生产环境建议 `network_mode: none`,用户代码无法访问外网
- **文件系统**: 只挂载用户自身的工作区目录,其他路径只读
- **资源限制**: 通过 cgroup 限制 CPU/内存/进程数
- **用户隔离**: 不同用户的容器之间完全隔离
### 6.2 命令安全
- 容器内 evosci-exec-server 维护命令黑名单(sudo、rm -rf / 等)
- 禁止访问容器元数据服务(169.254.169.254)
- 禁止安装内核模块
- 网络隔离模式下禁止 curl/wget 外部资源
### 6.3 Docker Socket 安全
Docker-in-Docker 方案(挂载 docker.sock)存在提权风险。生产环境建议:
1. **方案 A**: 使用 Docker Socket Proxy(推荐)
```yaml
# docker-compose.yml
docker-proxy:
image: tecnativa/docker-socket-proxy
environment:
CONTAINERS: 1
IMAGES: 1
NETWORKS: 0
VOLUMES: 0
POST: 1
volumes:
- /var/run/docker.sock:/var/run/docker.sock:ro
```
2. **方案 B**: 使用独立执行节点(推荐大规模部署)
- 执行引擎通过 TCP 连接远程 Docker daemon
- Gateway 与执行节点物理隔离
---
## 7. 扩展与升级
### 7.1 增加执行节点
当单机容器池不够时,可横向扩展:
```
Gateway ──► Docker Node 1 (192.168.1.10:2376)
──► Docker Node 2 (192.168.1.11:2376)
──► Modal Cloud (GPU 任务)
```
在 settings.yaml 中配置:
```yaml
execution:
docker:
nodes:
- host: tcp://192.168.1.10:2376
max_containers: 20
labels: [cpu]
- host: tcp://192.168.1.11:2376
max_containers: 10
labels: [gpu]
gpu_available: true
```
### 7.2 升级到 Kubernetes
当规模进一步增长(50+ 并发用户),可切换到 K8s:
```yaml
execution:
provider: kubernetes
kubernetes:
namespace: evosci-exec
image: evosci-exec:latest
namespace_quota:
max_pods: 100
gpu_limit: 10
```
### 7.3 迁移路径
```
阶段 1: Docker 单机 (当前)
↓
阶段 2: Docker 多节点
↓
阶段 3: Kubernetes 集群
↓
阶段 4: 混合云 (本地 K8s + Modal GPU)
```