# EvoScientist 执行引擎 — 部署运维文档 > 基于 SWE-ReX 架构模式(Docker 本地 + Modal 云端双模式) ## 1. 概述 EvoScientist 执行引擎为实验代码提供安全隔离的运行环境。核心思路:将代码执行从 Agent 主进程中剥离,放入独立的 Docker 容器或 Modal 云端 Sandbox 中运行,通过 HTTP API 交互。 ### 1.1 架构总览 ``` ┌──────────────────┐ │ Gateway/API │ │ (FastAPI 8065) │ └────────┬─────────┘ │ ┌────────▼─────────┐ │ Execution Engine │ │ (调度核心) │ └──┬────────────┬───┘ │ │ ┌────────────▼──┐ ┌────▼────────────┐ │ DockerProvider│ │ ModalProvider │ │ (本地容器池) │ │ (云端GPU弹性) │ └───────┬───────┘ └──────┬──────────┘ │ │ ┌────────▼──────┐ ┌────────▼──────┐ │ Container ×N │ │ Modal Sandbox │ │ evosci-exec │ │ (GPU 可用) │ │ (FastAPI) │ │ (按需启停) │ └───────┬───────┘ └────────┬──────┘ │ │ ┌────▼────────────────────▼────┐ │ 共享存储 (NFS/S3/本地) │ │ ~/.evoscientist/data/ │ │ {user_id}/{thread_id}/ │ └──────────────────────────────┘ ``` ### 1.2 核心组件 | 组件 | 职责 | 运行位置 | |------|------|---------| | Execution Engine | 任务调度、路由、容器生命周期管理 | Gateway 进程内 | | evosci-exec-server | 容器内 FastAPI 服务,接收并执行命令 | 每个执行容器内 | | DockerProvider | 管理本地 Docker 容器池 | Gateway 进程内 | | ModalProvider | 管理 Modal 云端 Sandbox | Gateway 进程内 | | 执行容器镜像 | 预装科学计算库的隔离环境 | Docker Hub / 本地构建 | ### 1.3 数据流 ``` 1. 用户发消息 → Gateway 2. Agent 调用 execute 工具 3. ExecutionEngine.submit(task) 4. 路由判断: - LIGHT 任务 → DockerProvider → 启动/复用容器 → HTTP /execute - HEAVY 任务 → ModalProvider → 创建 Sandbox → HTTP /execute 5. 容器内 evosci-exec-server 执行命令 6. 结果原路返回 → Agent → 用户 ``` --- ## 2. 环境要求 ### 2.1 硬件要求 | 场景 | CPU | 内存 | 磁盘 | GPU | |------|-----|------|------|-----| | 开发/测试 | 4 核 | 8 GB | 50 GB | 无 | | 生产(轻量) | 8 核 | 32 GB | 200 GB | 无 | | 生产(GPU) | 16 核 | 64 GB | 500 GB | 1× A100/T4 | ### 2.2 软件要求 | 依赖 | 版本 | 用途 | 必需 | |------|------|------|------| | Docker Engine | ≥ 24.0 | 容器运行时 | 是 | | Docker Compose | ≥ 2.20 | 服务编排 | 是 | | Python | ≥ 3.11 | 后端运行时 | 是 | | PostgreSQL | ≥ 15 | 主数据库(已有) | 是 | | Redis | ≥ 7.0 | 任务队列(P3 阶段) | P3 | | NVIDIA Container Toolkit | ≥ 1.14 | GPU 容器支持 | GPU 场景 | | Modal CLI | ≥ 0.6 | 云端执行 | Modal 场景 | ### 2.3 网络要求 ``` 端口规划: 8065 — Gateway API (已有) 8880 — evosci-exec-server (容器内部, 不对外暴露) 5432 — PostgreSQL (已有) 6379 — Redis (P3 阶段) Docker 网络: evosci-net — 执行容器与 Gateway 的通信网络 容器内网络可设为 none (无外网) 或 bridge (有外网) ``` --- ## 3. 部署步骤 ### 3.1 构建执行容器镜像 ```bash cd ~/Projects/EvoSci/EvoScientist # 构建轻量版 (CPU, 无 GPU) docker build -t evosci-exec:latest \ -f docker/evosci-exec/Dockerfile . # 构建 GPU 版 (需要 NVIDIA Container Toolkit) docker build -t evosci-exec-gpu:latest \ -f docker/evosci-exec/Dockerfile.gpu . # 验证镜像 docker run --rm evosci-exec:latest python -c "import numpy; print('OK')" ``` **Dockerfile** — 轻量版: ```dockerfile # docker/evosci-exec/Dockerfile FROM python:3.11-slim AS base RUN apt-get update && apt-get install -y --no-install-recommends \ gcc g++ git curl wget ca-certificates \ && rm -rf /var/lib/apt/lists/* RUN pip install --no-cache-dir \ numpy pandas matplotlib scipy scikit-learn \ httpx fastapi uvicorn pydantic COPY execution/server.py /opt/evosci/evosci-exec-server.py WORKDIR /workspace EXPOSE 8880 HEALTHCHECK --interval=10s --timeout=3s --retries=3 \ CMD curl -f http://localhost:8880/is_alive || exit 1 CMD ["python", "/opt/evosci/evosci-exec-server.py", "--port", "8880"] ``` **Dockerfile** — GPU 版: ```dockerfile # docker/evosci-exec/Dockerfile.gpu FROM nvidia/cuda:12.4.0-runtime-ubuntu22.04 AS base RUN apt-get update && apt-get install -y --no-install-recommends \ python3 python3-pip python3-dev gcc g++ git curl wget \ && rm -rf /var/lib/apt/lists/* RUN pip3 install --no-cache-dir \ numpy pandas matplotlib scipy scikit-learn \ torch --index-url https://download.pytorch.org/whl/cu124 \ httpx fastapi uvicorn pydantic COPY execution/server.py /opt/evosci/evosci-exec-server.py WORKDIR /workspace EXPOSE 8880 CMD ["python3", "/opt/evosci/evosci-exec-server.py", "--port", "8880"] ``` ### 3.2 配置执行引擎 在 `settings.yaml` 中新增 `execution` 段: ```yaml # settings.yaml execution: # 全局开关 enabled: true # 默认执行提供者: docker | modal | auto # auto: 轻量任务走 Docker, GPU 任务走 Modal provider: auto docker: # 镜像名称 image: evosci-exec:latest gpu_image: evosci-exec-gpu:latest # 容器池限制 max_containers: 20 # 最大并发容器数 idle_timeout: 1800 # 空闲容器回收 (秒) startup_timeout: 60 # 容器启动超时 (秒) # 默认资源 default_cpu: 2 # 每容器 CPU 核数 default_memory: 4g # 每容器内存 max_cpu: 8 # 单容器上限 max_memory: 16g # 网络 network_mode: none # none (隔离) | bridge (可联网) # 清理 remove_on_stop: true # 容器停止后自动删除 volume_cleanup: true # 清理容器卷 modal: # Modal App 名称 app_name: evoscientist # 镜像 base_image: python:3.11-slim # 资源 default_timeout: 3600 # Sandbox 默认超时 (秒) deployment_timeout: 7200 # 部署最大存活 (秒) default_cpu: 4 default_memory: "16GiB" gpu_type: A100 # A100 | T4 | L4 max_gpu: 4 # 沙箱额外参数 sandbox_kwargs: {} # 用户配额 quota: # 默认用户 (普通套餐) default: max_concurrent: 2 # 最大并发容器 max_cpu: 4 max_memory: "8g" max_gpu: 0 daily_budget_minutes: 100 # 每日执行总时长上限 max_task_timeout: 1800 # 单任务超时上限 (秒) # 高级用户 (premium 套餐) premium: max_concurrent: 5 max_cpu: 8 max_memory: "32g" max_gpu: 2 daily_budget_minutes: 500 max_task_timeout: 7200 # 任务调度 scheduler: queue_type: memory # memory (初版) | redis (P3) priority_levels: 3 # 优先级级数 retry_on_failure: true max_retries: 2 ``` ### 3.3 Docker Compose 更新 ```yaml # docker/docker-compose.yml — 完整版 services: backend: image: ${BACKEND_IMAGE:-ghcr.io/jakeyang886/evoscientist-backend:latest} build: context: ../ dockerfile: docker/Dockerfile args: APP_VERSION: "${APP_VERSION:-0.0.0}" ports: - "8065:8065" env_file: - ../.env environment: - CORS_ORIGINS=${CORS_ORIGINS:-http://localhost:3065} - BASE_URL=${BASE_URL:-http://localhost:3065} - GATEWAY_HOST=0.0.0.0 - GATEWAY_PORT=8065 - EVOSCIENTIST_HOME=/app/.data volumes: - ../.data:/app/.data # Docker-in-Docker: 让 backend 能管理执行容器 - /var/run/docker.sock:/var/run/docker.sock # 镜像构建缓存 - evosci-images:/var/lib/docker depends_on: redis: condition: service_healthy restart: unless-stopped # 资源限制 (backend 本身) deploy: resources: limits: cpus: "4" memory: 8G frontend: image: ${FRONTEND_IMAGE:-ghcr.io/jakeyang886/evoscientist-frontend:latest} build: context: ../ dockerfile: docker/Dockerfile.frontend args: INTERNAL_GATEWAY_URL: http://backend:8065 NEXT_PUBLIC_GATEWAY_URL: "" ports: - "3065:3065" environment: - PORT=3065 - INTERNAL_GATEWAY_URL=http://backend:8065 depends_on: backend: condition: service_healthy restart: unless-stopped redis: image: redis:7-alpine ports: - "6379:6379" volumes: - redis-data:/data healthcheck: test: ["CMD", "redis-cli", "ping"] interval: 5s timeout: 3s retries: 5 restart: unless-stopped volumes: evosci-images: redis-data: ``` ### 3.4 环境变量 在 `.env` 中新增: ```bash # .env # === 执行引擎 === EVOSCIENTIST_EXECUTION_ENABLED=true EVOSCIENTIST_EXECUTION_PROVIDER=auto # Docker EVOSCIENTIST_DOCKER_IMAGE=evosci-exec:latest EVOSCIENTIST_DOCKER_MAX_CONTAINERS=20 # Modal (云端 GPU) MODAL_TOKEN_ID=ak-xxx # Modal 凭证 MODAL_TOKEN_SECRET=as-xxx # 用户配额 EVOSCIENTIST_DEFAULT_MAX_CONCURRENT=2 EVOSCIENTIST_PREMIUM_MAX_CONCURRENT=5 ``` ### 3.5 验证部署 ```bash # 1. 构建镜像 docker build -t evosci-exec:latest -f docker/evosci-exec/Dockerfile . # 2. 启动所有服务 docker compose -f docker/docker-compose.yml up -d # 3. 检查服务状态 curl http://localhost:8065/health # 4. 测试执行引擎 # 通过 API 创建一个执行任务 curl -X POST http://localhost:8065/api/execution/submit \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "command": "python -c \"print(1+1)\"", "timeout": 30 }' # 5. 查看执行容器 docker ps --filter "name=evosci-" # 6. 查看容器日志 docker logs evosci- ``` --- ## 4. 运维操作 ### 4.1 容器池管理 ```bash # 查看所有执行容器 docker ps --filter "name=evosci-" --format "table {{.Names}}\t{{.Status}}\t{{.CreatedAt}}" # 清理所有空闲容器 docker stop $(docker ps -q --filter "name=evosci-session-") 2>/dev/null docker rm $(docker ps -aq --filter "name=evosci-session-") 2>/dev/null # 清理退出的容器 docker container prune --filter "label=evosci-exec" # 查看资源使用 docker stats --no-stream --format "table {{.Name}}\t{{.CPUPerc}}\t{{.MemUsage}}" \ $(docker ps -q --filter "name=evosci-") ``` ### 4.2 镜像管理 ```bash # 构建新版本镜像 docker build -t evosci-exec:v1.1.0 -f docker/evosci-exec/Dockerfile . # 更新运行中的镜像 (滚动更新) # 1. 更新 settings.yaml 中的 image 字段 # 2. 重启 Gateway docker compose -f docker/docker-compose.yml restart backend # 清理旧镜像 docker image prune --filter "label=evosci-exec" --filter "until=168h" ``` ### 4.3 GPU 支持 (本地) ```bash # 检查 NVIDIA Container Toolkit nvidia-ctk --version # 测试 GPU 容器 docker run --rm --gpus all evosci-exec-gpu:latest \ python -c "import torch; print(torch.cuda.is_available())" # 在 settings.yaml 中设置 # docker.gpu_image: evosci-exec-gpu:latest # docker.default_gpu: 1 # 默认分配 GPU 数 ``` ### 4.4 Modal 云端配置 ```bash # 安装 Modal CLI pip install modal # 登录 modal profile create # 验证连接 modal app list # 查看 Sandbox 使用情况 modal sandbox list # 设置凭证 (在 .env 中) MODAL_TOKEN_ID=ak-xxx MODAL_TOKEN_SECRET=as-xxx ``` ### 4.5 监控指标 ```bash # 通过 Gateway API 获取执行引擎状态 curl http://localhost:8065/api/execution/status # 返回示例: # { # "docker": { # "running_containers": 5, # "idle_containers": 2, # "total_cpu_used": "8/32 cores", # "total_memory_used": "20/64 GB" # }, # "modal": { # "active_sandboxes": 1, # "gpu_hours_today": 2.5 # }, # "queue": { # "pending_tasks": 3, # "running_tasks": 5, # "completed_today": 127 # } # } ``` --- ## 5. 故障排查 ### 5.1 常见问题 | 问题 | 原因 | 解决 | |------|------|------| | 容器启动超时 | 镜像未拉取 / 磁盘不足 | `docker pull evosci-exec:latest`; 清理磁盘 | | `docker.sock` 权限不足 | 后端容器无 Docker 访问权限 | 检查 socket 挂载和用户组 | | 容器内命令超时 | 任务复杂度超预期 | 增大 `timeout` 参数; 使用 Modal | | 内存 OOM | 容器内存不足 | 增大 `default_memory`; 检查内存泄漏 | | Modal 连接失败 | 凭证过期 / 网络不通 | 刷新 `MODAL_TOKEN_*`; 检查防火墙 | | 容器池满 | 并发用户超过 `max_containers` | 增大限制; 检查空闲容器是否被回收 | ### 5.2 日志查看 ```bash # Gateway 日志 docker compose -f docker/docker-compose.yml logs -f backend # 特定执行容器日志 docker logs -f evosci-session-- # evosci-exec-server 日志 (容器内) docker exec evosci-xxx cat /var/log/evosci-exec.log ``` ### 5.3 紧急操作 ```bash # 强制停止所有执行容器 docker stop $(docker ps -q --filter "label=evosci-exec") # 清理所有执行容器和镜像 docker rm -f $(docker ps -aq --filter "name=evosci-") docker rmi $(docker images -q --filter "label=evosci-exec") # 禁用执行引擎 (回退到旧模式) # 在 settings.yaml 中设置: # execution.enabled: false ``` --- ## 6. 安全注意事项 ### 6.1 容器隔离 - **网络隔离**: 生产环境建议 `network_mode: none`,用户代码无法访问外网 - **文件系统**: 只挂载用户自身的工作区目录,其他路径只读 - **资源限制**: 通过 cgroup 限制 CPU/内存/进程数 - **用户隔离**: 不同用户的容器之间完全隔离 ### 6.2 命令安全 - 容器内 evosci-exec-server 维护命令黑名单(sudo、rm -rf / 等) - 禁止访问容器元数据服务(169.254.169.254) - 禁止安装内核模块 - 网络隔离模式下禁止 curl/wget 外部资源 ### 6.3 Docker Socket 安全 Docker-in-Docker 方案(挂载 docker.sock)存在提权风险。生产环境建议: 1. **方案 A**: 使用 Docker Socket Proxy(推荐) ```yaml # docker-compose.yml docker-proxy: image: tecnativa/docker-socket-proxy environment: CONTAINERS: 1 IMAGES: 1 NETWORKS: 0 VOLUMES: 0 POST: 1 volumes: - /var/run/docker.sock:/var/run/docker.sock:ro ``` 2. **方案 B**: 使用独立执行节点(推荐大规模部署) - 执行引擎通过 TCP 连接远程 Docker daemon - Gateway 与执行节点物理隔离 --- ## 7. 扩展与升级 ### 7.1 增加执行节点 当单机容器池不够时,可横向扩展: ``` Gateway ──► Docker Node 1 (192.168.1.10:2376) ──► Docker Node 2 (192.168.1.11:2376) ──► Modal Cloud (GPU 任务) ``` 在 settings.yaml 中配置: ```yaml execution: docker: nodes: - host: tcp://192.168.1.10:2376 max_containers: 20 labels: [cpu] - host: tcp://192.168.1.11:2376 max_containers: 10 labels: [gpu] gpu_available: true ``` ### 7.2 升级到 Kubernetes 当规模进一步增长(50+ 并发用户),可切换到 K8s: ```yaml execution: provider: kubernetes kubernetes: namespace: evosci-exec image: evosci-exec:latest namespace_quota: max_pods: 100 gpu_limit: 10 ``` ### 7.3 迁移路径 ``` 阶段 1: Docker 单机 (当前) ↓ 阶段 2: Docker 多节点 ↓ 阶段 3: Kubernetes 集群 ↓ 阶段 4: 混合云 (本地 K8s + Modal GPU) ```