c2743251e9
Self-evolving AI scientist framework built on LangGraph/LangChain with CLI/TUI core, FastAPI gateway, and Next.js frontend. Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
16 KiB
16 KiB
EvoScientist 执行引擎 — 部署运维文档
基于 SWE-ReX 架构模式(Docker 本地 + Modal 云端双模式)
1. 概述
EvoScientist 执行引擎为实验代码提供安全隔离的运行环境。核心思路:将代码执行从 Agent 主进程中剥离,放入独立的 Docker 容器或 Modal 云端 Sandbox 中运行,通过 HTTP API 交互。
1.1 架构总览
┌──────────────────┐
│ Gateway/API │
│ (FastAPI 8065) │
└────────┬─────────┘
│
┌────────▼─────────┐
│ Execution Engine │
│ (调度核心) │
└──┬────────────┬───┘
│ │
┌────────────▼──┐ ┌────▼────────────┐
│ DockerProvider│ │ ModalProvider │
│ (本地容器池) │ │ (云端GPU弹性) │
└───────┬───────┘ └──────┬──────────┘
│ │
┌────────▼──────┐ ┌────────▼──────┐
│ Container ×N │ │ Modal Sandbox │
│ evosci-exec │ │ (GPU 可用) │
│ (FastAPI) │ │ (按需启停) │
└───────┬───────┘ └────────┬──────┘
│ │
┌────▼────────────────────▼────┐
│ 共享存储 (NFS/S3/本地) │
│ ~/.evoscientist/data/ │
│ {user_id}/{thread_id}/ │
└──────────────────────────────┘
1.2 核心组件
| 组件 | 职责 | 运行位置 |
|---|---|---|
| Execution Engine | 任务调度、路由、容器生命周期管理 | Gateway 进程内 |
| evosci-exec-server | 容器内 FastAPI 服务,接收并执行命令 | 每个执行容器内 |
| DockerProvider | 管理本地 Docker 容器池 | Gateway 进程内 |
| ModalProvider | 管理 Modal 云端 Sandbox | Gateway 进程内 |
| 执行容器镜像 | 预装科学计算库的隔离环境 | Docker Hub / 本地构建 |
1.3 数据流
1. 用户发消息 → Gateway
2. Agent 调用 execute 工具
3. ExecutionEngine.submit(task)
4. 路由判断:
- LIGHT 任务 → DockerProvider → 启动/复用容器 → HTTP /execute
- HEAVY 任务 → ModalProvider → 创建 Sandbox → HTTP /execute
5. 容器内 evosci-exec-server 执行命令
6. 结果原路返回 → Agent → 用户
2. 环境要求
2.1 硬件要求
| 场景 | CPU | 内存 | 磁盘 | GPU |
|---|---|---|---|---|
| 开发/测试 | 4 核 | 8 GB | 50 GB | 无 |
| 生产(轻量) | 8 核 | 32 GB | 200 GB | 无 |
| 生产(GPU) | 16 核 | 64 GB | 500 GB | 1× A100/T4 |
2.2 软件要求
| 依赖 | 版本 | 用途 | 必需 |
|---|---|---|---|
| Docker Engine | ≥ 24.0 | 容器运行时 | 是 |
| Docker Compose | ≥ 2.20 | 服务编排 | 是 |
| Python | ≥ 3.11 | 后端运行时 | 是 |
| PostgreSQL | ≥ 15 | 主数据库(已有) | 是 |
| Redis | ≥ 7.0 | 任务队列(P3 阶段) | P3 |
| NVIDIA Container Toolkit | ≥ 1.14 | GPU 容器支持 | GPU 场景 |
| Modal CLI | ≥ 0.6 | 云端执行 | Modal 场景 |
2.3 网络要求
端口规划:
8065 — Gateway API (已有)
8880 — evosci-exec-server (容器内部, 不对外暴露)
5432 — PostgreSQL (已有)
6379 — Redis (P3 阶段)
Docker 网络:
evosci-net — 执行容器与 Gateway 的通信网络
容器内网络可设为 none (无外网) 或 bridge (有外网)
3. 部署步骤
3.1 构建执行容器镜像
cd ~/Projects/EvoSci/EvoScientist
# 构建轻量版 (CPU, 无 GPU)
docker build -t evosci-exec:latest \
-f docker/evosci-exec/Dockerfile .
# 构建 GPU 版 (需要 NVIDIA Container Toolkit)
docker build -t evosci-exec-gpu:latest \
-f docker/evosci-exec/Dockerfile.gpu .
# 验证镜像
docker run --rm evosci-exec:latest python -c "import numpy; print('OK')"
Dockerfile — 轻量版:
# docker/evosci-exec/Dockerfile
FROM python:3.11-slim AS base
RUN apt-get update && apt-get install -y --no-install-recommends \
gcc g++ git curl wget ca-certificates \
&& rm -rf /var/lib/apt/lists/*
RUN pip install --no-cache-dir \
numpy pandas matplotlib scipy scikit-learn \
httpx fastapi uvicorn pydantic
COPY execution/server.py /opt/evosci/evosci-exec-server.py
WORKDIR /workspace
EXPOSE 8880
HEALTHCHECK --interval=10s --timeout=3s --retries=3 \
CMD curl -f http://localhost:8880/is_alive || exit 1
CMD ["python", "/opt/evosci/evosci-exec-server.py", "--port", "8880"]
Dockerfile — GPU 版:
# docker/evosci-exec/Dockerfile.gpu
FROM nvidia/cuda:12.4.0-runtime-ubuntu22.04 AS base
RUN apt-get update && apt-get install -y --no-install-recommends \
python3 python3-pip python3-dev gcc g++ git curl wget \
&& rm -rf /var/lib/apt/lists/*
RUN pip3 install --no-cache-dir \
numpy pandas matplotlib scipy scikit-learn \
torch --index-url https://download.pytorch.org/whl/cu124 \
httpx fastapi uvicorn pydantic
COPY execution/server.py /opt/evosci/evosci-exec-server.py
WORKDIR /workspace
EXPOSE 8880
CMD ["python3", "/opt/evosci/evosci-exec-server.py", "--port", "8880"]
3.2 配置执行引擎
在 settings.yaml 中新增 execution 段:
# settings.yaml
execution:
# 全局开关
enabled: true
# 默认执行提供者: docker | modal | auto
# auto: 轻量任务走 Docker, GPU 任务走 Modal
provider: auto
docker:
# 镜像名称
image: evosci-exec:latest
gpu_image: evosci-exec-gpu:latest
# 容器池限制
max_containers: 20 # 最大并发容器数
idle_timeout: 1800 # 空闲容器回收 (秒)
startup_timeout: 60 # 容器启动超时 (秒)
# 默认资源
default_cpu: 2 # 每容器 CPU 核数
default_memory: 4g # 每容器内存
max_cpu: 8 # 单容器上限
max_memory: 16g
# 网络
network_mode: none # none (隔离) | bridge (可联网)
# 清理
remove_on_stop: true # 容器停止后自动删除
volume_cleanup: true # 清理容器卷
modal:
# Modal App 名称
app_name: evoscientist
# 镜像
base_image: python:3.11-slim
# 资源
default_timeout: 3600 # Sandbox 默认超时 (秒)
deployment_timeout: 7200 # 部署最大存活 (秒)
default_cpu: 4
default_memory: "16GiB"
gpu_type: A100 # A100 | T4 | L4
max_gpu: 4
# 沙箱额外参数
sandbox_kwargs: {}
# 用户配额
quota:
# 默认用户 (普通套餐)
default:
max_concurrent: 2 # 最大并发容器
max_cpu: 4
max_memory: "8g"
max_gpu: 0
daily_budget_minutes: 100 # 每日执行总时长上限
max_task_timeout: 1800 # 单任务超时上限 (秒)
# 高级用户 (premium 套餐)
premium:
max_concurrent: 5
max_cpu: 8
max_memory: "32g"
max_gpu: 2
daily_budget_minutes: 500
max_task_timeout: 7200
# 任务调度
scheduler:
queue_type: memory # memory (初版) | redis (P3)
priority_levels: 3 # 优先级级数
retry_on_failure: true
max_retries: 2
3.3 Docker Compose 更新
# docker/docker-compose.yml — 完整版
services:
backend:
image: ${BACKEND_IMAGE:-ghcr.io/jakeyang886/evoscientist-backend:latest}
build:
context: ../
dockerfile: docker/Dockerfile
args:
APP_VERSION: "${APP_VERSION:-0.0.0}"
ports:
- "8065:8065"
env_file:
- ../.env
environment:
- CORS_ORIGINS=${CORS_ORIGINS:-http://localhost:3065}
- BASE_URL=${BASE_URL:-http://localhost:3065}
- GATEWAY_HOST=0.0.0.0
- GATEWAY_PORT=8065
- EVOSCIENTIST_HOME=/app/.data
volumes:
- ../.data:/app/.data
# Docker-in-Docker: 让 backend 能管理执行容器
- /var/run/docker.sock:/var/run/docker.sock
# 镜像构建缓存
- evosci-images:/var/lib/docker
depends_on:
redis:
condition: service_healthy
restart: unless-stopped
# 资源限制 (backend 本身)
deploy:
resources:
limits:
cpus: "4"
memory: 8G
frontend:
image: ${FRONTEND_IMAGE:-ghcr.io/jakeyang886/evoscientist-frontend:latest}
build:
context: ../
dockerfile: docker/Dockerfile.frontend
args:
INTERNAL_GATEWAY_URL: http://backend:8065
NEXT_PUBLIC_GATEWAY_URL: ""
ports:
- "3065:3065"
environment:
- PORT=3065
- INTERNAL_GATEWAY_URL=http://backend:8065
depends_on:
backend:
condition: service_healthy
restart: unless-stopped
redis:
image: redis:7-alpine
ports:
- "6379:6379"
volumes:
- redis-data:/data
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 5s
timeout: 3s
retries: 5
restart: unless-stopped
volumes:
evosci-images:
redis-data:
3.4 环境变量
在 .env 中新增:
# .env
# === 执行引擎 ===
EVOSCIENTIST_EXECUTION_ENABLED=true
EVOSCIENTIST_EXECUTION_PROVIDER=auto
# Docker
EVOSCIENTIST_DOCKER_IMAGE=evosci-exec:latest
EVOSCIENTIST_DOCKER_MAX_CONTAINERS=20
# Modal (云端 GPU)
MODAL_TOKEN_ID=ak-xxx # Modal 凭证
MODAL_TOKEN_SECRET=as-xxx
# 用户配额
EVOSCIENTIST_DEFAULT_MAX_CONCURRENT=2
EVOSCIENTIST_PREMIUM_MAX_CONCURRENT=5
3.5 验证部署
# 1. 构建镜像
docker build -t evosci-exec:latest -f docker/evosci-exec/Dockerfile .
# 2. 启动所有服务
docker compose -f docker/docker-compose.yml up -d
# 3. 检查服务状态
curl http://localhost:8065/health
# 4. 测试执行引擎
# 通过 API 创建一个执行任务
curl -X POST http://localhost:8065/api/execution/submit \
-H "Authorization: Bearer <token>" \
-H "Content-Type: application/json" \
-d '{
"command": "python -c \"print(1+1)\"",
"timeout": 30
}'
# 5. 查看执行容器
docker ps --filter "name=evosci-"
# 6. 查看容器日志
docker logs evosci-<container-id>
4. 运维操作
4.1 容器池管理
# 查看所有执行容器
docker ps --filter "name=evosci-" --format "table {{.Names}}\t{{.Status}}\t{{.CreatedAt}}"
# 清理所有空闲容器
docker stop $(docker ps -q --filter "name=evosci-session-") 2>/dev/null
docker rm $(docker ps -aq --filter "name=evosci-session-") 2>/dev/null
# 清理退出的容器
docker container prune --filter "label=evosci-exec"
# 查看资源使用
docker stats --no-stream --format "table {{.Name}}\t{{.CPUPerc}}\t{{.MemUsage}}" \
$(docker ps -q --filter "name=evosci-")
4.2 镜像管理
# 构建新版本镜像
docker build -t evosci-exec:v1.1.0 -f docker/evosci-exec/Dockerfile .
# 更新运行中的镜像 (滚动更新)
# 1. 更新 settings.yaml 中的 image 字段
# 2. 重启 Gateway
docker compose -f docker/docker-compose.yml restart backend
# 清理旧镜像
docker image prune --filter "label=evosci-exec" --filter "until=168h"
4.3 GPU 支持 (本地)
# 检查 NVIDIA Container Toolkit
nvidia-ctk --version
# 测试 GPU 容器
docker run --rm --gpus all evosci-exec-gpu:latest \
python -c "import torch; print(torch.cuda.is_available())"
# 在 settings.yaml 中设置
# docker.gpu_image: evosci-exec-gpu:latest
# docker.default_gpu: 1 # 默认分配 GPU 数
4.4 Modal 云端配置
# 安装 Modal CLI
pip install modal
# 登录
modal profile create
# 验证连接
modal app list
# 查看 Sandbox 使用情况
modal sandbox list
# 设置凭证 (在 .env 中)
MODAL_TOKEN_ID=ak-xxx
MODAL_TOKEN_SECRET=as-xxx
4.5 监控指标
# 通过 Gateway API 获取执行引擎状态
curl http://localhost:8065/api/execution/status
# 返回示例:
# {
# "docker": {
# "running_containers": 5,
# "idle_containers": 2,
# "total_cpu_used": "8/32 cores",
# "total_memory_used": "20/64 GB"
# },
# "modal": {
# "active_sandboxes": 1,
# "gpu_hours_today": 2.5
# },
# "queue": {
# "pending_tasks": 3,
# "running_tasks": 5,
# "completed_today": 127
# }
# }
5. 故障排查
5.1 常见问题
| 问题 | 原因 | 解决 |
|---|---|---|
| 容器启动超时 | 镜像未拉取 / 磁盘不足 | docker pull evosci-exec:latest; 清理磁盘 |
docker.sock 权限不足 |
后端容器无 Docker 访问权限 | 检查 socket 挂载和用户组 |
| 容器内命令超时 | 任务复杂度超预期 | 增大 timeout 参数; 使用 Modal |
| 内存 OOM | 容器内存不足 | 增大 default_memory; 检查内存泄漏 |
| Modal 连接失败 | 凭证过期 / 网络不通 | 刷新 MODAL_TOKEN_*; 检查防火墙 |
| 容器池满 | 并发用户超过 max_containers |
增大限制; 检查空闲容器是否被回收 |
5.2 日志查看
# Gateway 日志
docker compose -f docker/docker-compose.yml logs -f backend
# 特定执行容器日志
docker logs -f evosci-session-<user-id>-<thread-id>
# evosci-exec-server 日志 (容器内)
docker exec evosci-xxx cat /var/log/evosci-exec.log
5.3 紧急操作
# 强制停止所有执行容器
docker stop $(docker ps -q --filter "label=evosci-exec")
# 清理所有执行容器和镜像
docker rm -f $(docker ps -aq --filter "name=evosci-")
docker rmi $(docker images -q --filter "label=evosci-exec")
# 禁用执行引擎 (回退到旧模式)
# 在 settings.yaml 中设置:
# execution.enabled: false
6. 安全注意事项
6.1 容器隔离
- 网络隔离: 生产环境建议
network_mode: none,用户代码无法访问外网 - 文件系统: 只挂载用户自身的工作区目录,其他路径只读
- 资源限制: 通过 cgroup 限制 CPU/内存/进程数
- 用户隔离: 不同用户的容器之间完全隔离
6.2 命令安全
- 容器内 evosci-exec-server 维护命令黑名单(sudo、rm -rf / 等)
- 禁止访问容器元数据服务(169.254.169.254)
- 禁止安装内核模块
- 网络隔离模式下禁止 curl/wget 外部资源
6.3 Docker Socket 安全
Docker-in-Docker 方案(挂载 docker.sock)存在提权风险。生产环境建议:
-
方案 A: 使用 Docker Socket Proxy(推荐)
# docker-compose.yml docker-proxy: image: tecnativa/docker-socket-proxy environment: CONTAINERS: 1 IMAGES: 1 NETWORKS: 0 VOLUMES: 0 POST: 1 volumes: - /var/run/docker.sock:/var/run/docker.sock:ro -
方案 B: 使用独立执行节点(推荐大规模部署)
- 执行引擎通过 TCP 连接远程 Docker daemon
- Gateway 与执行节点物理隔离
7. 扩展与升级
7.1 增加执行节点
当单机容器池不够时,可横向扩展:
Gateway ──► Docker Node 1 (192.168.1.10:2376)
──► Docker Node 2 (192.168.1.11:2376)
──► Modal Cloud (GPU 任务)
在 settings.yaml 中配置:
execution:
docker:
nodes:
- host: tcp://192.168.1.10:2376
max_containers: 20
labels: [cpu]
- host: tcp://192.168.1.11:2376
max_containers: 10
labels: [gpu]
gpu_available: true
7.2 升级到 Kubernetes
当规模进一步增长(50+ 并发用户),可切换到 K8s:
execution:
provider: kubernetes
kubernetes:
namespace: evosci-exec
image: evosci-exec:latest
namespace_quota:
max_pods: 100
gpu_limit: 10
7.3 迁移路径
阶段 1: Docker 单机 (当前)
↓
阶段 2: Docker 多节点
↓
阶段 3: Kubernetes 集群
↓
阶段 4: 混合云 (本地 K8s + Modal GPU)