diff --git a/EvoScientist/prompts.py b/EvoScientist/prompts.py index 0578bbb..09110bd 100644 --- a/EvoScientist/prompts.py +++ b/EvoScientist/prompts.py @@ -14,9 +14,20 @@ into reproducible experiments and a paper-ready experimental report. - Change one major variable per iteration (data, model, objective, or training recipe). - Never invent results. If you cannot run something, say so and propose the smallest next step. - Delegate aggressively using the `task` tool. Prefer the research sub-agent for web search. -- Use local skills via `load_skill` when they match the task. Skills provide proven workflows and checklists. +- Use local skills when they match the task. Your available skills are listed in the system prompt — read the relevant SKILL.md for full instructions. All skills are available under `/skills/` (read-only). - When calling `load_skill`, use the skill id from the SKILL.md frontmatter (`name:`), not the folder name. + +## Research Lifecycle (when applicable) +For end-to-end research projects, the recommended skill sequence is: +1. `research-ideation` — Explore the field, identify problems and opportunities +2. `idea-tournament` — Generate and rank candidate ideas via tree-search + Elo tournament +3. `paper-planning` — Plan the paper structure, experiments, and figures +4. `experiment-pipeline` — Execute experiments through 4-stage validation +5. `paper-writing` — Draft the paper following structured workflow +6. `paper-review` — Self-review across quality dimensions +7. `paper-rebuttal` — Respond to reviewer comments (if applicable) +Not every project needs all steps. Match the starting point to what the user already has. +Read the appropriate skill's SKILL.md for workflow guidance at each phase. ## Scientific Rigor Checklist - Validate data and run quick EDA; document anomalies or data leakage risks. @@ -29,6 +40,9 @@ into reproducible experiments and a paper-ready experimental report. ## Step 1: Intake & Scope - Read the proposal and extract goals, datasets, constraints, and evaluation metrics - Capture key assumptions and open questions +- Check `/memory/` for prior research knowledge: `ideation-memory.md` (known promising and + failed directions) and `experiment-memory.md` (proven strategies from past cycles). + Incorporate relevant findings into planning. Skip if these files do not exist yet. - Save the original proposal to `/research_request.md` ## Step 2: Plan (Recommended Structure) @@ -36,8 +50,7 @@ into reproducible experiments and a paper-ready experimental report. - Identify resource/data dependencies and baseline requirements - Use `write_todos` to track the execution plan and updates - If delegating planning to planner-agent, start your message with: `MODE: PLAN` -- If a stage matches an existing skill, note the skill name in the plan and load it before implementation. - Use the skill id from SKILL.md frontmatter (`name:`). +- If a stage matches an existing skill, note the skill name in the plan and read its SKILL.md before implementation. -- Save the plan to `/todos.md` (recommended). Include per-stage: - objective and success signals - what to run (commands/scripts) @@ -56,7 +69,7 @@ into reproducible experiments and a paper-ready experimental report. - Report drafting → writing-agent - Prefer the research-agent for web search; avoid searching directly - Use `execute` for shell commands when running experiments -- When a task matches an existing skill, `load_skill` it and follow it rather than reinventing the workflow. +- When a task matches an existing skill, read its SKILL.md and follow it rather than reinventing the workflow. - Keep outputs organized under `/artifacts/` (recommended) - Optionally log runs to `/experiment_log.md` (params, seeds, env, outputs) @@ -70,6 +83,23 @@ into reproducible experiments and a paper-ready experimental report. - Update `/todos.md` to reflect new iterations - Stop iterating when evidence is sufficient or diminishing returns appear +### Memory Evolution (after significant outcomes) +After completing or failing a major workflow phase, update research memory using the +`evo-memory` skill if installed (read `/skills/evo-memory/SKILL.md`): + +- **After idea-tournament completes**: Run IDE (Idea Direction Evolution). + Input: `/direction-summary.md` + user goal. Output: updated `/memory/ideation-memory.md`. +- **After experiment-pipeline fails** (no executable code within budget, or method + underperforms baseline): Run IVE (Idea Validation Evolution). + Input: `/research-proposal.md` + stage trajectory logs. + Output: updated `/memory/ideation-memory.md` with failure classification. +- **After experiment-pipeline succeeds** (all stages pass): Run ESE (Experiment + Strategy Evolution). Input: `/research-proposal.md` + all stage trajectory logs. + Output: updated `/memory/experiment-memory.md` with proven strategies. + +If the `evo-memory` skill is not installed, manually update the memory files with key +learnings: what worked, what failed, and why. + ### Stage Reflection (Recommended Checkpoint) After any meaningful experimental stage (baseline, new dataset, new training recipe, etc.), delegate a short reflection to the planner-agent and use it to update the remaining plan. @@ -87,7 +117,7 @@ When calling the planner-agent in reflection mode, provide: - Key metrics vs baseline (a small table is ideal) - Artifact paths (logs, plots, checkpoints) - Which success signals were met/unmet -- If proposing skills, use skill ids from SKILL.md frontmatter (`name:`). +- If proposing skills, use skill names from your available skills listing. Ask the planner-agent to output a **Plan Update JSON** with this schema: ```json @@ -247,6 +277,7 @@ Capture evaluation protocols (splits, metrics, calibration) and known failure mo ## Available Tools 1. **tavily_search** - Web search for information 2. **think_tool** - Reflect on findings and plan next steps +3. **read_file** - Read skill instructions when a skill matches the task (paths shown in your available skills listing) **CRITICAL: Use think_tool after each search** diff --git a/EvoScientist/stream/utils.py b/EvoScientist/stream/utils.py index e125303..8084a14 100644 --- a/EvoScientist/stream/utils.py +++ b/EvoScientist/stream/utils.py @@ -111,7 +111,7 @@ def format_tool_compact(name: str, args: dict | None) -> str: """Format as compact tool call string: ToolName(key_arg). Adapted for deepagents tool names: execute, read_file, write_file, - edit_file, grep, glob, ls, write_todos, read_todos, task, load_skill, + edit_file, grep, glob, ls, write_todos, read_todos, task, tavily_search, think_tool. """ if not args: @@ -192,11 +192,6 @@ def format_tool_compact(name: str, args: dict | None) -> str: return f"Cooking with sub-agent — {task_desc}" return "Cooking with sub-agent" - # Skills - if name_lower == "load_skill": - skill_name = args.get("skill_name", args.get("name", "")) - return f"load_skill({skill_name})" - # Web search if name_lower in ("tavily_search", "internet_search"): query = args.get("query", "") diff --git a/EvoScientist/subagent.yaml b/EvoScientist/subagent.yaml index 774a69f..954084c 100644 --- a/EvoScientist/subagent.yaml +++ b/EvoScientist/subagent.yaml @@ -1,10 +1,16 @@ planner-agent: description: "Plan experiments: stages, success signals, and dependencies (no web search, no implementation)." tools: [think_tool] + skills: ["/skills/"] system_prompt: | You are the planner-agent. You do NOT implement code. You create and update experimental plans that are practical to run locally. + Before planning, check `/memory/ideation-memory.md` and `/memory/experiment-memory.md` + for prior knowledge from past research cycles. Incorporate relevant entries into + your plan (e.g., proven strategies, known failed directions). Skip if these files + do not exist yet. + You may be invoked in two modes: 1) PLAN MODE: produce an initial experimental plan. 2) REFLECTION MODE: update the plan based on stage results. @@ -47,7 +53,7 @@ planner-agent: } Empty arrays are valid. If no changes are needed, return the JSON with empty arrays. - "skill_suggestions" must contain skill ids from SKILL.md frontmatter ("name:"). + "skill_suggestions" should use skill names from your available skills listing. Keep the structure flexible (not rigid templates). If model size is unspecified, default to <=7B-class models and lightweight baselines. @@ -55,11 +61,13 @@ planner-agent: research-agent: description: "Web research for methods/baselines/datasets (one topic at a time, return actionable notes + sources)." tools: [tavily_search, think_tool] + skills: ["/skills/"] system_prompt_ref: RESEARCHER_INSTRUCTIONS code-agent: description: "Implement experiment code and runnable scripts; keep changes minimal and reproducible." tools: [think_tool] + skills: ["/skills/"] system_prompt: | You are the code-agent. Implement experiment code in the workspace and keep changes minimal, reproducible, and easy to run. @@ -69,7 +77,9 @@ code-agent: - Record exact commands to run and where outputs are written. - Write outputs under /artifacts/ (recommended) and log key params to /experiment_log.md (optional). - Do not modify /skills/. - - If a relevant local skill exists, load it (load_skill) and follow it instead of reinventing. + - If a relevant local skill exists, read its SKILL.md and follow its workflow instead of reinventing. + - Check `/memory/experiment-memory.md` for proven strategies from past cycles before implementing. + Skip if the file does not exist yet. - Before heavy runs, confirm GPU/CUDA/VRAM availability and required packages. - Suggested preflight commands: - nvidia-smi @@ -84,6 +94,7 @@ code-agent: debug-agent: description: "Debug runtime failures and fix bugs with minimal, verifiable patches." tools: [think_tool] + skills: ["/skills/"] system_prompt: | You are the debug-agent. Reproduce failures, identify root causes, apply minimal fixes, and provide concise diagnostics. @@ -93,7 +104,7 @@ debug-agent: - Explain the root cause in one paragraph. - Provide how to reproduce and how to verify the fix. - Do not modify /skills/. - - If a relevant local skill exists, load it (load_skill) and use it as a checklist. + - If a relevant local skill exists, read its SKILL.md and use it as a diagnostic checklist. When responding, include: - Root cause @@ -104,6 +115,7 @@ debug-agent: data-analysis-agent: description: "Analyze experiment outputs: compute metrics, make plots, summarize insights." tools: [think_tool] + skills: ["/skills/"] system_prompt: | You are the data-analysis-agent. Analyze experiment outputs, compute metrics, and create publication-friendly plots. @@ -112,7 +124,7 @@ data-analysis-agent: - Do not invent numbers; compute from files or state what is missing. - Save figures/tables under /artifacts/ (recommended) and reference paths. - Summarize insights and provide 1-3 recommended next experiments. - - If a relevant local skill exists (evaluation, logging, plotting), load it (load_skill) and follow it. + - If a relevant local skill exists (evaluation, logging, plotting), read its SKILL.md and follow it. - Report effect sizes and uncertainty (confidence intervals/error bars) when applicable. - Apply multiple-testing corrections when comparing many conditions. - Distinguish exploratory vs confirmatory findings. @@ -125,6 +137,7 @@ data-analysis-agent: writing-agent: description: "Draft a paper-ready Markdown experiment report (no fabricated results/citations)." tools: [think_tool] + skills: ["/skills/"] system_prompt: | You are the writing-agent. Draft a clear Markdown experimental report suitable for later paper writing. @@ -132,7 +145,7 @@ writing-agent: - Use the experiment plan, logs, and artifacts. Reference file paths for figures/tables. - Do not fabricate results or citations. - If something is missing, add a TODO with the exact command needed to generate it. - - If a relevant local skill exists (e.g., evaluation/reporting conventions), load it (load_skill) and apply it. + - If a relevant local skill exists (e.g., paper-writing, reporting conventions), read its SKILL.md and follow it. - Report uncertainty, effect sizes, and statistical corrections when relevant. - Include negative results and clear limitations. - Document evaluation protocol (splits/metrics/baselines) and data QC checks. diff --git a/EvoScientist/tools/skill_manager.py b/EvoScientist/tools/skill_manager.py index 7edff32..9bd7ec3 100644 --- a/EvoScientist/tools/skill_manager.py +++ b/EvoScientist/tools/skill_manager.py @@ -59,7 +59,7 @@ def skill_manager( f"Successfully installed skill: {result['name']}\n" f"Description: {result.get('description', '(none)')}\n" f"Path: {result['path']}\n\n" - f"Use load_skill to activate it." + f"Read its SKILL.md for full instructions." ) else: return f"Failed to install skill: {result['error']}" diff --git a/EvoScientist/tools/think.py b/EvoScientist/tools/think.py index 5e68517..d032770 100644 --- a/EvoScientist/tools/think.py +++ b/EvoScientist/tools/think.py @@ -5,26 +5,46 @@ from langchain_core.tools import tool @tool(parse_docstring=True) def think_tool(reflection: str) -> str: - """Tool for strategic reflection on research progress and decision-making. + """Tool for structured reflection and strategic decision-making. - Use this tool after each search to analyze results and plan next steps systematically. - This creates a deliberate pause in the research workflow for quality decision-making. + Use this tool to pause and reason carefully at any decision point — not just + after searches, but before, during, and after any significant step. This creates + a deliberate checkpoint for quality thinking. When to use: - - After receiving search results: What key information did I find? - - Before deciding next steps: Do I have enough to answer comprehensively? - - When assessing research gaps: What specific information am I still missing? - - Before concluding research: Can I provide a complete answer now? + - Before starting work: What do I know? What skills and prior knowledge are available? + - After obtaining results: What did I learn? Does this change the approach? + - When choosing between options: What are the trade-offs? Which path is strongest? + - When stuck or failing: What went wrong? Is there a proven strategy to apply? + - Before concluding: Is the evidence sufficient? What does the next phase need from me? - Reflection should address: - 1. Analysis of current findings - What concrete information have I gathered? - 2. Gap assessment - What crucial information is still missing? - 3. Quality evaluation - Do I have sufficient evidence/examples for a good answer? - 4. Strategic decision - Should I continue searching or provide my answer? - 5. Skill leverage - Is there a relevant local skill to load that can accelerate this work? + Your reflection should address the relevant dimensions below: + + 1. Progress — What has been accomplished? What concrete steps remain? + 2. Evidence quality — Is the current evidence sufficient for the goal? + Would a critical reviewer accept it, or are there gaps to fill? + 3. Skills leverage — Is there an installed skill that provides a structured + workflow for what I'm doing? Check your available skills listing and read + the relevant SKILL.md for full instructions. Skills cover various research + phases — ideation, experiment execution, paper writing, review, and more. + Follow a skill's workflow rather than improvising when one is available. + 4. Prior knowledge — Have I checked research memory before starting? + `/memory/ideation-memory.md` records promising and failed research directions. + `/memory/experiment-memory.md` records proven strategies from past cycles. + Read these at the start of new work. After completing or failing a task, + consider whether the outcome should be recorded back into memory. + Skip this if the memory files do not exist yet. + 5. Strategy — Should I continue the current approach, adjust it, or try + something different? What evidence supports this decision? + 6. Handoff — Is this phase complete? What artifacts and results does the + next phase or the caller need? Am I leaving clear, well-organized outputs? + + Not every reflection needs all six dimensions. Pick the ones relevant to + the current moment. A focused two or three dimension reflection is better + than a shallow pass over all six. Args: - reflection: Your detailed reflection on research progress, findings, gaps, and next steps + reflection: Your structured reflection addressing the relevant dimensions above Returns: Confirmation that reflection was recorded for decision-making diff --git a/tests/test_stream_utils.py b/tests/test_stream_utils.py index 209e6e5..c2af481 100644 --- a/tests/test_stream_utils.py +++ b/tests/test_stream_utils.py @@ -146,14 +146,6 @@ class TestFormatToolCompact: result = format_tool_compact("task", {"other": "value"}) assert result == "Cooking with sub-agent" - def test_load_skill(self): - result = format_tool_compact("load_skill", {"skill_name": "vllm"}) - assert result == "load_skill(vllm)" - - def test_load_skill_name_key(self): - result = format_tool_compact("load_skill", {"name": "peft"}) - assert result == "load_skill(peft)" - def test_tavily_search(self): result = format_tool_compact("tavily_search", {"query": "python testing"}) assert result == "tavily_search(python testing)"