Files
EvoScientist/docs/image-skill.md
T
m4 c2743251e9 Initial commit of EvoScientist framework
Self-evolving AI scientist framework built on LangGraph/LangChain with
CLI/TUI core, FastAPI gateway, and Next.js frontend.

Co-Authored-By: Claude Opus 4 <noreply@anthropic.com>
2026-07-13 08:07:45 +08:00

16 KiB

EvoScientist Image Skill Design

Goal

Add image generation and image editing to EvoScientist through a dedicated image-artist skill plus backend tools that call the configured image model.

Supported user workflows:

  • Text-to-image: generate covers, posters, illustrations, icons, diagrams, and visual assets from a prompt.
  • Image editing: restyle, retouch, remove or replace backgrounds, generate variations, and perform image-to-image edits from uploaded or artifact images.

The image model is configured in settings.yaml:

image_gen_base_url: https://hunnuapi.top
image_gen_model: chatgpt-image-2
image_gen_api_key: <server-side secret>

The API key must remain server-side only and must not be committed to the repository, copied into skill files, shown in frontend code, or returned in tool responses.

Architecture

User request
  -> main agent detects image task
  -> image-artist skill provides workflow and prompt rules
  -> generate_image / edit_image tools execute the request
  -> Gateway ImageGenerationService calls the image model
  -> generated bytes are saved under artifacts/
  -> artifact registration exposes the image to the frontend
  -> ArtifactsPanel previews/downloads the image

Responsibilities:

  • image-artist skill: task classification, prompt construction, workflow selection, output naming guidance, quality checks, and retry strategy.
  • Agent tools: expose generate_image and edit_image to the agent.
  • Gateway service: load config, call the OpenAI-compatible image endpoint, validate paths, decode image results, write artifacts, and return metadata.
  • Frontend: reuse artifact image preview, with optional quality-of-life actions.

Current Repository Fit

Existing capabilities that should be reused:

  • settings.yaml already supports image_gen_api_key, image_gen_base_url, and image_gen_model.
  • Admin service settings already expose image generation fields and mask the API key in read responses.
  • ArtifactsPanel already previews common image formats.
  • File routes already serve thread files and image blobs.
  • Stream cleanup already registers or mirrors files under artifacts/ for FileGateway/local workspace flows.
  • Path Semantics already defines artifacts/... as the user-visible output path.

Implementation should reuse these pieces instead of creating a parallel storage or preview path.

Skill Design

Create an internal skill:

EvoScientist/skills/image-artist/
├── SKILL.md
├── references/
│   ├── prompt-patterns.md
│   ├── edit-workflows.md
│   └── quality-checks.md
└── assets/

SKILL.md frontmatter:

---
name: image-artist
description: Use when the user asks to draw, generate, create, edit, retouch, restyle, remove background from, or create variations of images, covers, posters, illustrations, icons, diagrams, or visual assets.
---

Core workflow:

  1. Classify the task as generation, editing, variation, background removal, style transfer, poster/cover design, icon generation, or diagram asset creation.
  2. Ask a clarification only when a critical constraint is missing, such as target size, visual style, target use, transparency, or mandatory text.
  3. Use generate_image for new images.
  4. Use edit_image when an input image is available.
  5. Save outputs under artifacts/....
  6. Verify the artifact exists before reporting completion.
  7. Report only user-visible logical paths such as artifacts/cover.png.

The skill should keep SKILL.md short. Put reusable prompt templates and variant-specific guidance in reference files.

Agent Tools

Add a tool module:

EvoScientist/tools/image_generation.py

Expose two required tools.

generate_image(
    prompt: str,
    size: str = "1024x1024",
    quality: str = "auto",
    background: str = "auto",
    output_path: str | None = None,
    n: int = 1,
)
edit_image(
    image_path: str,
    prompt: str,
    mask_path: str | None = None,
    size: str = "1024x1024",
    quality: str = "auto",
    output_path: str | None = None,
)

Optional follow-up tool:

describe_image(image_path: str)

Tool success result:

{
  "ok": true,
  "path": "artifacts/generated_cover.png",
  "mime_type": "image/png",
  "model": "chatgpt-image-2",
  "size": "1024x1024"
}

Tool failure result:

{
  "ok": false,
  "error": "IMAGE_GEN_API_KEY is missing"
}

Tool rules:

  • Never return the API key.
  • Never accept absolute output paths.
  • Normalize output paths to artifacts/....
  • Reject path traversal with ...
  • Use PNG by default.
  • Verify the returned artifact path exists before reporting success.

Tool registration:

  • Import the tool functions from EvoScientist/tools/image_generation.py.
  • Add them to the main agent base tool list.
  • Make them available to the main agent whenever image generation config is present.
  • Do not add image tools to planner-agent; planner-agent should only plan.
  • image-artist is a skill, not a tool. It gives the agent the workflow for deciding when and how to call generate_image or edit_image.

Gateway Service

Add:

gateway/models/image_generation.py
gateway/services/image_generation.py
gateway/routes/image_generation.py

Suggested routes:

POST /api/image-generation/generate
POST /api/image-generation/edit

Route requirements:

  • Routes must require the same authenticated user context as chat/file routes.
  • thread_id must belong to the authenticated user.
  • The service must derive user_uid from the authenticated user/session, not from an arbitrary request body field.
  • For web threads, resolve the workspace using the same helper used by file routes/FileGateway integration.
  • For CLI/local use, the tool may call the service layer directly with the active workspace instead of going through HTTP.

The service should:

  • Load image_gen_api_key, image_gen_base_url, and image_gen_model from settings.yaml via existing config loading.
  • Validate the current user and thread.
  • Resolve input image paths only inside the current thread workspace.
  • Restrict edit inputs to uploads/... and artifacts/....
  • Restrict outputs to artifacts/....
  • Call the configured OpenAI-compatible endpoint.
  • Decode returned base64 images or download returned URLs.
  • Write image bytes to the thread workspace.
  • Return artifact metadata.
  • Register or make the artifact discoverable using the same artifact pipeline as other generated files.

Generation request:

{
  "thread_id": "thread_id",
  "prompt": "生成一张小学数字化转型主题封面图,科技感,明亮,适合课题申报书",
  "size": "1024x1024",
  "output_path": "artifacts/topic_cover.png"
}

Generation response:

{
  "path": "artifacts/topic_cover.png",
  "mime_type": "image/png",
  "model": "chatgpt-image-2",
  "size": "1024x1024"
}

OpenAI-Compatible API Strategy

The configured provider uses:

base_url = https://hunnuapi.top
model = chatgpt-image-2

Prefer image endpoints first:

POST {base_url}/v1/images/generations
POST {base_url}/v1/images/edits

Generation payload shape:

{
  "model": "chatgpt-image-2",
  "prompt": "...",
  "size": "1024x1024",
  "quality": "auto",
  "n": 1,
  "response_format": "b64_json"
}

Edit payload shape:

multipart/form-data
  model=chatgpt-image-2
  prompt=...
  image=@input.png
  mask=@mask.png       # optional
  size=1024x1024
  quality=auto
  response_format=b64_json

If the provider exposes image generation through the Responses API, support a fallback path:

POST {base_url}/v1/responses
model: chatgpt-image-2
tools: [{"type": "image_generation"}]

Response parsing should support:

  • b64_json
  • url
  • Responses API image output blocks containing base64 data

Endpoint URL construction should avoid double /v1 when the configured base URL already includes it.

Provider compatibility checks:

  • On startup or first use, fail with a clear error if image_gen_api_key, image_gen_base_url, or image_gen_model is missing.
  • If /v1/images/generations returns "not found" or "unsupported model", retry the Responses API fallback only when the response is compatible with that interpretation.
  • If both endpoints fail, return a concise user-facing error without exposing request headers, API key, or raw provider payload.

Artifact Rules

Default paths:

artifacts/generated_YYYYMMDD_HHMMSS.png
artifacts/edited_YYYYMMDD_HHMMSS.png

User-provided paths are allowed only when they are relative paths under artifacts/.

Reject:

/tmp/out.png
/workspace/artifacts/out.png
../out.png
uploads/out.png

Accept:

artifacts/out.png
artifacts/covers/topic_cover.png

If the user omits an extension, append .png.

Artifact registration details:

  • In FileGateway mode, write to the mounted thread workspace under artifacts/... and call the existing artifact/file registration path if the file needs to appear before stream finalization.
  • In local/NFS mode, write to the thread workspace under artifacts/...; stream finalization can register it, but the tool should still return a path that is immediately readable by file preview routes.
  • Do not create duplicate root-level files next to artifacts/.
  • Store only the generated image bytes and minimal metadata. Do not store raw API responses unless needed for debugging and scrubbed of secrets.

Frontend Integration

The current ArtifactsPanel already previews common image types:

  • png
  • jpg/jpeg
  • gif
  • webp
  • bmp
  • svg
  • ico

MVP frontend work can reuse this panel.

Recommended enhancements:

  • Show image thumbnails in the artifact list.
  • Add image actions in the preview panel:
    • Download
    • Copy path
    • Continue editing
    • Generate variation
  • Add optional prompt shortcuts near the chat input:
    • Generate image
    • Edit selected image
    • Create variation

Admin Configuration

The services admin page already has image generation fields:

image_gen_api_key
image_gen_base_url
image_gen_model

Required behavior:

  • Save values into settings.yaml.
  • Mask API keys in read responses.
  • Reload gateway config after updates.
  • Provide an optional test connection button.

Suggested test route:

POST /api/admin/services/image-gen/test

The test should not expose the key and should avoid generating expensive output unless explicitly requested.

Admin reload behavior:

  • Updating image generation fields should persist to settings.yaml.
  • Gateway config should reload without a full server restart when possible.
  • Agent processes should see updated environment/config values on the next tool invocation.

Security

Required controls:

  • Keep the API key server-side only.
  • Never log or return the full API key.
  • Reject absolute paths and path traversal.
  • Allow edit_image to read only current-thread uploads/... and artifacts/....
  • Allow writes only to current-thread artifacts/....
  • Limit input file types to png, jpg, jpeg, and webp.
  • Enforce an input image size limit, for example 20 MB or the provider-specific limit.
  • Limit n, size, and quality according to user plan.
  • Do not persist provider raw responses if they contain sensitive metadata.
  • Sanitize provider errors before returning them to the user.
  • Redact Authorization headers and API keys from logs.
  • Avoid sending uploaded images to the provider unless the user explicitly asks for an image edit or variation.

Billing And Limits

Track image usage separately from chat token usage:

service = image_generation
model = chatgpt-image-2
units = image_count

Suggested limits:

  • Default n = 1.
  • Single request maximum n = 4.
  • Free or starter users have a daily quota.
  • High quality or large size consumes more quota.
  • Charge only after the provider returns a successful image.
  • Do not charge failed validation or provider failures.

Quota enforcement point:

  • Validate quota before calling the provider.
  • Record usage only after a successful provider response is decoded and written to artifacts/....
  • Record the model, operation (generate or edit), image count, and requested size.

Prompt Strategy

Text-to-image prompt structure:

Subject:
Style:
Composition:
Lighting:
Color palette:
Use case:
Avoid:
Output:

Example:

一张用于课题申报书封面的插画。主题是小学教育数字化转型与学生问题意识培养。
画面包含智慧教室、学生提问、数字化学习终端、柔和科技光效。
风格为现代教育科技插画,明亮、专业、干净,适合正式申报材料。
不要出现乱码文字,不要出现品牌 logo,不要过度卡通。

Image-edit prompt structure:

Keep:
Change:
Style:
Constraints:
Output:

Example:

保留原图主体人物和构图,将整体风格改为现代科技教育海报。
增强蓝白色科技感光效,背景加入智慧课堂元素。
不要改变人物面部特征,不要添加无关文字,不要出现品牌 logo。

Chinese text warning:

  • Image models may render Chinese text unreliably.
  • For exact Chinese text, prefer generating a text-free image first and adding text later through deterministic image composition or frontend editing.

Tests

Backend tests:

  • Load image config from settings.yaml.
  • Build endpoint URLs correctly when base URL has or lacks /v1.
  • Mock generation endpoint returning base64 and verify PNG artifact write.
  • Mock edit endpoint returning base64 and verify PNG artifact write.
  • Mock URL response and verify image download/write.
  • Mock Responses API image output and verify parsing.
  • Reject absolute output paths.
  • Reject path traversal.
  • Reject edit input outside the thread workspace.
  • Reject unsupported image extensions.
  • Reject requests for a thread not owned by the authenticated user.
  • Return clear errors for missing API key or model.
  • Do not expose the API key in errors or logs.
  • Correctly parse base64 and URL responses.
  • Record usage only on successful image generation.

Tool tests:

  • generate_image schema is valid.
  • edit_image schema is valid.
  • Tools normalize output to artifacts/....
  • Tools return stable user-visible paths without leading slash.
  • Tools are included in the main agent tool list when configured.
  • Tools do not appear in planner-agent.
  • Tool errors are concise and do not expose provider secrets.

Skill tests:

  • Trigger terms include draw, generate image, edit image, retouch, remove background, restyle, variation, cover, poster, illustration, icon, and diagram.
  • Skill instructions do not expose API keys.
  • Skill instructions do not use /workspace/... as a user-visible output path.

Frontend tests:

  • Image artifacts preview correctly.
  • Download works.
  • Copy path works.
  • Masked API key renders in admin settings.
  • Test connection failure displays a clear error.

End-to-end smoke tests:

  • Text prompt -> generate_image -> artifacts/*.png -> preview route returns image bytes.
  • Uploaded image -> edit_image -> artifacts/*_edited.png -> preview route returns image bytes.
  • A generated path in the final answer is shown as artifacts/..., never as a /workspace/... path.

Implementation Plan

Phase 1: MVP

  1. Create image-artist skill.
  2. Add ImageGenerationService.
  3. Add generate_image and edit_image tools.
  4. Call chatgpt-image-2 through the configured OpenAI-compatible endpoint.
  5. Save PNG outputs to artifacts/....
  6. Reuse the current artifact preview panel.
  7. Add backend and tool tests.

Phase 2: User Experience

  1. Add artifact thumbnails.
  2. Add Continue editing and Generate variation actions.
  3. Add prompt shortcuts in the chat input.
  4. Add admin test connection.

Phase 3: Governance And Extension

  1. Add image generation billing and daily quotas.
  2. Add per-plan limits for image count, size, and quality.
  3. Add optional provider fallback.
  4. Add local ComfyUI or Stable Diffusion provider support.
  5. Add image version history.

Example User Flows

Text-to-image:

User: 帮我画一张“小学数字化转型与学生问题意识培养”的课题封面图。
Agent: Uses image-artist, calls generate_image, saves artifacts/topic_cover.png.

Image editing:

User: 把我上传的图片改成更正式的申报书封面风格。
Agent: Uses image-artist, calls edit_image on uploads/source.png, saves artifacts/source_formal_cover.png.

Variation:

User: 基于这张图再生成 3 个不同风格版本。
Agent: Uses image-artist, calls generate/edit variation workflow, saves artifacts/variation_*.png.