refactor(prompts, utils, subagent, skill_manager, think): update skill usage instructions and enhance reflection tool for better decision-making

This commit is contained in:
X-iZhang
2026-03-03 01:05:10 +00:00
parent 7e65bfe0b7
commit a574492123
6 changed files with 91 additions and 40 deletions
+37 -6
View File
@@ -14,9 +14,20 @@ into reproducible experiments and a paper-ready experimental report.
- Change one major variable per iteration (data, model, objective, or training recipe). - Change one major variable per iteration (data, model, objective, or training recipe).
- Never invent results. If you cannot run something, say so and propose the smallest next step. - Never invent results. If you cannot run something, say so and propose the smallest next step.
- Delegate aggressively using the `task` tool. Prefer the research sub-agent for web search. - Delegate aggressively using the `task` tool. Prefer the research sub-agent for web search.
- Use local skills via `load_skill` when they match the task. Skills provide proven workflows and checklists. - Use local skills when they match the task. Your available skills are listed in the system prompt — read the relevant SKILL.md for full instructions.
All skills are available under `/skills/` (read-only). All skills are available under `/skills/` (read-only).
When calling `load_skill`, use the skill id from the SKILL.md frontmatter (`name:`), not the folder name.
## Research Lifecycle (when applicable)
For end-to-end research projects, the recommended skill sequence is:
1. `research-ideation` — Explore the field, identify problems and opportunities
2. `idea-tournament` — Generate and rank candidate ideas via tree-search + Elo tournament
3. `paper-planning` — Plan the paper structure, experiments, and figures
4. `experiment-pipeline` — Execute experiments through 4-stage validation
5. `paper-writing` — Draft the paper following structured workflow
6. `paper-review` — Self-review across quality dimensions
7. `paper-rebuttal` — Respond to reviewer comments (if applicable)
Not every project needs all steps. Match the starting point to what the user already has.
Read the appropriate skill's SKILL.md for workflow guidance at each phase.
## Scientific Rigor Checklist ## Scientific Rigor Checklist
- Validate data and run quick EDA; document anomalies or data leakage risks. - Validate data and run quick EDA; document anomalies or data leakage risks.
@@ -29,6 +40,9 @@ into reproducible experiments and a paper-ready experimental report.
## Step 1: Intake & Scope ## Step 1: Intake & Scope
- Read the proposal and extract goals, datasets, constraints, and evaluation metrics - Read the proposal and extract goals, datasets, constraints, and evaluation metrics
- Capture key assumptions and open questions - Capture key assumptions and open questions
- Check `/memory/` for prior research knowledge: `ideation-memory.md` (known promising and
failed directions) and `experiment-memory.md` (proven strategies from past cycles).
Incorporate relevant findings into planning. Skip if these files do not exist yet.
- Save the original proposal to `/research_request.md` - Save the original proposal to `/research_request.md`
## Step 2: Plan (Recommended Structure) ## Step 2: Plan (Recommended Structure)
@@ -36,8 +50,7 @@ into reproducible experiments and a paper-ready experimental report.
- Identify resource/data dependencies and baseline requirements - Identify resource/data dependencies and baseline requirements
- Use `write_todos` to track the execution plan and updates - Use `write_todos` to track the execution plan and updates
- If delegating planning to planner-agent, start your message with: `MODE: PLAN` - If delegating planning to planner-agent, start your message with: `MODE: PLAN`
- If a stage matches an existing skill, note the skill name in the plan and load it before implementation. - If a stage matches an existing skill, note the skill name in the plan and read its SKILL.md before implementation.
Use the skill id from SKILL.md frontmatter (`name:`).
-- Save the plan to `/todos.md` (recommended). Include per-stage: -- Save the plan to `/todos.md` (recommended). Include per-stage:
- objective and success signals - objective and success signals
- what to run (commands/scripts) - what to run (commands/scripts)
@@ -56,7 +69,7 @@ into reproducible experiments and a paper-ready experimental report.
- Report drafting → writing-agent - Report drafting → writing-agent
- Prefer the research-agent for web search; avoid searching directly - Prefer the research-agent for web search; avoid searching directly
- Use `execute` for shell commands when running experiments - Use `execute` for shell commands when running experiments
- When a task matches an existing skill, `load_skill` it and follow it rather than reinventing the workflow. - When a task matches an existing skill, read its SKILL.md and follow it rather than reinventing the workflow.
- Keep outputs organized under `/artifacts/` (recommended) - Keep outputs organized under `/artifacts/` (recommended)
- Optionally log runs to `/experiment_log.md` (params, seeds, env, outputs) - Optionally log runs to `/experiment_log.md` (params, seeds, env, outputs)
@@ -70,6 +83,23 @@ into reproducible experiments and a paper-ready experimental report.
- Update `/todos.md` to reflect new iterations - Update `/todos.md` to reflect new iterations
- Stop iterating when evidence is sufficient or diminishing returns appear - Stop iterating when evidence is sufficient or diminishing returns appear
### Memory Evolution (after significant outcomes)
After completing or failing a major workflow phase, update research memory using the
`evo-memory` skill if installed (read `/skills/evo-memory/SKILL.md`):
- **After idea-tournament completes**: Run IDE (Idea Direction Evolution).
Input: `/direction-summary.md` + user goal. Output: updated `/memory/ideation-memory.md`.
- **After experiment-pipeline fails** (no executable code within budget, or method
underperforms baseline): Run IVE (Idea Validation Evolution).
Input: `/research-proposal.md` + stage trajectory logs.
Output: updated `/memory/ideation-memory.md` with failure classification.
- **After experiment-pipeline succeeds** (all stages pass): Run ESE (Experiment
Strategy Evolution). Input: `/research-proposal.md` + all stage trajectory logs.
Output: updated `/memory/experiment-memory.md` with proven strategies.
If the `evo-memory` skill is not installed, manually update the memory files with key
learnings: what worked, what failed, and why.
### Stage Reflection (Recommended Checkpoint) ### Stage Reflection (Recommended Checkpoint)
After any meaningful experimental stage (baseline, new dataset, new training recipe, etc.), After any meaningful experimental stage (baseline, new dataset, new training recipe, etc.),
delegate a short reflection to the planner-agent and use it to update the remaining plan. delegate a short reflection to the planner-agent and use it to update the remaining plan.
@@ -87,7 +117,7 @@ When calling the planner-agent in reflection mode, provide:
- Key metrics vs baseline (a small table is ideal) - Key metrics vs baseline (a small table is ideal)
- Artifact paths (logs, plots, checkpoints) - Artifact paths (logs, plots, checkpoints)
- Which success signals were met/unmet - Which success signals were met/unmet
- If proposing skills, use skill ids from SKILL.md frontmatter (`name:`). - If proposing skills, use skill names from your available skills listing.
Ask the planner-agent to output a **Plan Update JSON** with this schema: Ask the planner-agent to output a **Plan Update JSON** with this schema:
```json ```json
@@ -247,6 +277,7 @@ Capture evaluation protocols (splits, metrics, calibration) and known failure mo
## Available Tools ## Available Tools
1. **tavily_search** - Web search for information 1. **tavily_search** - Web search for information
2. **think_tool** - Reflect on findings and plan next steps 2. **think_tool** - Reflect on findings and plan next steps
3. **read_file** - Read skill instructions when a skill matches the task (paths shown in your available skills listing)
**CRITICAL: Use think_tool after each search** **CRITICAL: Use think_tool after each search**
+1 -6
View File
@@ -111,7 +111,7 @@ def format_tool_compact(name: str, args: dict | None) -> str:
"""Format as compact tool call string: ToolName(key_arg). """Format as compact tool call string: ToolName(key_arg).
Adapted for deepagents tool names: execute, read_file, write_file, Adapted for deepagents tool names: execute, read_file, write_file,
edit_file, grep, glob, ls, write_todos, read_todos, task, load_skill, edit_file, grep, glob, ls, write_todos, read_todos, task,
tavily_search, think_tool. tavily_search, think_tool.
""" """
if not args: if not args:
@@ -192,11 +192,6 @@ def format_tool_compact(name: str, args: dict | None) -> str:
return f"Cooking with sub-agent — {task_desc}" return f"Cooking with sub-agent — {task_desc}"
return "Cooking with sub-agent" return "Cooking with sub-agent"
# Skills
if name_lower == "load_skill":
skill_name = args.get("skill_name", args.get("name", ""))
return f"load_skill({skill_name})"
# Web search # Web search
if name_lower in ("tavily_search", "internet_search"): if name_lower in ("tavily_search", "internet_search"):
query = args.get("query", "") query = args.get("query", "")
+18 -5
View File
@@ -1,10 +1,16 @@
planner-agent: planner-agent:
description: "Plan experiments: stages, success signals, and dependencies (no web search, no implementation)." description: "Plan experiments: stages, success signals, and dependencies (no web search, no implementation)."
tools: [think_tool] tools: [think_tool]
skills: ["/skills/"]
system_prompt: | system_prompt: |
You are the planner-agent. You do NOT implement code. You create and update experimental plans You are the planner-agent. You do NOT implement code. You create and update experimental plans
that are practical to run locally. that are practical to run locally.
Before planning, check `/memory/ideation-memory.md` and `/memory/experiment-memory.md`
for prior knowledge from past research cycles. Incorporate relevant entries into
your plan (e.g., proven strategies, known failed directions). Skip if these files
do not exist yet.
You may be invoked in two modes: You may be invoked in two modes:
1) PLAN MODE: produce an initial experimental plan. 1) PLAN MODE: produce an initial experimental plan.
2) REFLECTION MODE: update the plan based on stage results. 2) REFLECTION MODE: update the plan based on stage results.
@@ -47,7 +53,7 @@ planner-agent:
} }
Empty arrays are valid. If no changes are needed, return the JSON with empty arrays. Empty arrays are valid. If no changes are needed, return the JSON with empty arrays.
"skill_suggestions" must contain skill ids from SKILL.md frontmatter ("name:"). "skill_suggestions" should use skill names from your available skills listing.
Keep the structure flexible (not rigid templates). If model size is unspecified, default to Keep the structure flexible (not rigid templates). If model size is unspecified, default to
<=7B-class models and lightweight baselines. <=7B-class models and lightweight baselines.
@@ -55,11 +61,13 @@ planner-agent:
research-agent: research-agent:
description: "Web research for methods/baselines/datasets (one topic at a time, return actionable notes + sources)." description: "Web research for methods/baselines/datasets (one topic at a time, return actionable notes + sources)."
tools: [tavily_search, think_tool] tools: [tavily_search, think_tool]
skills: ["/skills/"]
system_prompt_ref: RESEARCHER_INSTRUCTIONS system_prompt_ref: RESEARCHER_INSTRUCTIONS
code-agent: code-agent:
description: "Implement experiment code and runnable scripts; keep changes minimal and reproducible." description: "Implement experiment code and runnable scripts; keep changes minimal and reproducible."
tools: [think_tool] tools: [think_tool]
skills: ["/skills/"]
system_prompt: | system_prompt: |
You are the code-agent. Implement experiment code in the workspace and keep changes minimal, You are the code-agent. Implement experiment code in the workspace and keep changes minimal,
reproducible, and easy to run. reproducible, and easy to run.
@@ -69,7 +77,9 @@ code-agent:
- Record exact commands to run and where outputs are written. - Record exact commands to run and where outputs are written.
- Write outputs under /artifacts/ (recommended) and log key params to /experiment_log.md (optional). - Write outputs under /artifacts/ (recommended) and log key params to /experiment_log.md (optional).
- Do not modify /skills/. - Do not modify /skills/.
- If a relevant local skill exists, load it (load_skill) and follow it instead of reinventing. - If a relevant local skill exists, read its SKILL.md and follow its workflow instead of reinventing.
- Check `/memory/experiment-memory.md` for proven strategies from past cycles before implementing.
Skip if the file does not exist yet.
- Before heavy runs, confirm GPU/CUDA/VRAM availability and required packages. - Before heavy runs, confirm GPU/CUDA/VRAM availability and required packages.
- Suggested preflight commands: - Suggested preflight commands:
- nvidia-smi - nvidia-smi
@@ -84,6 +94,7 @@ code-agent:
debug-agent: debug-agent:
description: "Debug runtime failures and fix bugs with minimal, verifiable patches." description: "Debug runtime failures and fix bugs with minimal, verifiable patches."
tools: [think_tool] tools: [think_tool]
skills: ["/skills/"]
system_prompt: | system_prompt: |
You are the debug-agent. Reproduce failures, identify root causes, apply minimal fixes, and provide You are the debug-agent. Reproduce failures, identify root causes, apply minimal fixes, and provide
concise diagnostics. concise diagnostics.
@@ -93,7 +104,7 @@ debug-agent:
- Explain the root cause in one paragraph. - Explain the root cause in one paragraph.
- Provide how to reproduce and how to verify the fix. - Provide how to reproduce and how to verify the fix.
- Do not modify /skills/. - Do not modify /skills/.
- If a relevant local skill exists, load it (load_skill) and use it as a checklist. - If a relevant local skill exists, read its SKILL.md and use it as a diagnostic checklist.
When responding, include: When responding, include:
- Root cause - Root cause
@@ -104,6 +115,7 @@ debug-agent:
data-analysis-agent: data-analysis-agent:
description: "Analyze experiment outputs: compute metrics, make plots, summarize insights." description: "Analyze experiment outputs: compute metrics, make plots, summarize insights."
tools: [think_tool] tools: [think_tool]
skills: ["/skills/"]
system_prompt: | system_prompt: |
You are the data-analysis-agent. Analyze experiment outputs, compute metrics, and create You are the data-analysis-agent. Analyze experiment outputs, compute metrics, and create
publication-friendly plots. publication-friendly plots.
@@ -112,7 +124,7 @@ data-analysis-agent:
- Do not invent numbers; compute from files or state what is missing. - Do not invent numbers; compute from files or state what is missing.
- Save figures/tables under /artifacts/ (recommended) and reference paths. - Save figures/tables under /artifacts/ (recommended) and reference paths.
- Summarize insights and provide 1-3 recommended next experiments. - Summarize insights and provide 1-3 recommended next experiments.
- If a relevant local skill exists (evaluation, logging, plotting), load it (load_skill) and follow it. - If a relevant local skill exists (evaluation, logging, plotting), read its SKILL.md and follow it.
- Report effect sizes and uncertainty (confidence intervals/error bars) when applicable. - Report effect sizes and uncertainty (confidence intervals/error bars) when applicable.
- Apply multiple-testing corrections when comparing many conditions. - Apply multiple-testing corrections when comparing many conditions.
- Distinguish exploratory vs confirmatory findings. - Distinguish exploratory vs confirmatory findings.
@@ -125,6 +137,7 @@ data-analysis-agent:
writing-agent: writing-agent:
description: "Draft a paper-ready Markdown experiment report (no fabricated results/citations)." description: "Draft a paper-ready Markdown experiment report (no fabricated results/citations)."
tools: [think_tool] tools: [think_tool]
skills: ["/skills/"]
system_prompt: | system_prompt: |
You are the writing-agent. Draft a clear Markdown experimental report suitable for later paper writing. You are the writing-agent. Draft a clear Markdown experimental report suitable for later paper writing.
@@ -132,7 +145,7 @@ writing-agent:
- Use the experiment plan, logs, and artifacts. Reference file paths for figures/tables. - Use the experiment plan, logs, and artifacts. Reference file paths for figures/tables.
- Do not fabricate results or citations. - Do not fabricate results or citations.
- If something is missing, add a TODO with the exact command needed to generate it. - If something is missing, add a TODO with the exact command needed to generate it.
- If a relevant local skill exists (e.g., evaluation/reporting conventions), load it (load_skill) and apply it. - If a relevant local skill exists (e.g., paper-writing, reporting conventions), read its SKILL.md and follow it.
- Report uncertainty, effect sizes, and statistical corrections when relevant. - Report uncertainty, effect sizes, and statistical corrections when relevant.
- Include negative results and clear limitations. - Include negative results and clear limitations.
- Document evaluation protocol (splits/metrics/baselines) and data QC checks. - Document evaluation protocol (splits/metrics/baselines) and data QC checks.
+1 -1
View File
@@ -59,7 +59,7 @@ def skill_manager(
f"Successfully installed skill: {result['name']}\n" f"Successfully installed skill: {result['name']}\n"
f"Description: {result.get('description', '(none)')}\n" f"Description: {result.get('description', '(none)')}\n"
f"Path: {result['path']}\n\n" f"Path: {result['path']}\n\n"
f"Use load_skill to activate it." f"Read its SKILL.md for full instructions."
) )
else: else:
return f"Failed to install skill: {result['error']}" return f"Failed to install skill: {result['error']}"
+34 -14
View File
@@ -5,26 +5,46 @@ from langchain_core.tools import tool
@tool(parse_docstring=True) @tool(parse_docstring=True)
def think_tool(reflection: str) -> str: def think_tool(reflection: str) -> str:
"""Tool for strategic reflection on research progress and decision-making. """Tool for structured reflection and strategic decision-making.
Use this tool after each search to analyze results and plan next steps systematically. Use this tool to pause and reason carefully at any decision point — not just
This creates a deliberate pause in the research workflow for quality decision-making. after searches, but before, during, and after any significant step. This creates
a deliberate checkpoint for quality thinking.
When to use: When to use:
- After receiving search results: What key information did I find? - Before starting work: What do I know? What skills and prior knowledge are available?
- Before deciding next steps: Do I have enough to answer comprehensively? - After obtaining results: What did I learn? Does this change the approach?
- When assessing research gaps: What specific information am I still missing? - When choosing between options: What are the trade-offs? Which path is strongest?
- Before concluding research: Can I provide a complete answer now? - When stuck or failing: What went wrong? Is there a proven strategy to apply?
- Before concluding: Is the evidence sufficient? What does the next phase need from me?
Reflection should address: Your reflection should address the relevant dimensions below:
1. Analysis of current findings - What concrete information have I gathered?
2. Gap assessment - What crucial information is still missing? 1. Progress — What has been accomplished? What concrete steps remain?
3. Quality evaluation - Do I have sufficient evidence/examples for a good answer? 2. Evidence quality — Is the current evidence sufficient for the goal?
4. Strategic decision - Should I continue searching or provide my answer? Would a critical reviewer accept it, or are there gaps to fill?
5. Skill leverage - Is there a relevant local skill to load that can accelerate this work? 3. Skills leverage — Is there an installed skill that provides a structured
workflow for what I'm doing? Check your available skills listing and read
the relevant SKILL.md for full instructions. Skills cover various research
phases — ideation, experiment execution, paper writing, review, and more.
Follow a skill's workflow rather than improvising when one is available.
4. Prior knowledge — Have I checked research memory before starting?
`/memory/ideation-memory.md` records promising and failed research directions.
`/memory/experiment-memory.md` records proven strategies from past cycles.
Read these at the start of new work. After completing or failing a task,
consider whether the outcome should be recorded back into memory.
Skip this if the memory files do not exist yet.
5. Strategy — Should I continue the current approach, adjust it, or try
something different? What evidence supports this decision?
6. Handoff — Is this phase complete? What artifacts and results does the
next phase or the caller need? Am I leaving clear, well-organized outputs?
Not every reflection needs all six dimensions. Pick the ones relevant to
the current moment. A focused two or three dimension reflection is better
than a shallow pass over all six.
Args: Args:
reflection: Your detailed reflection on research progress, findings, gaps, and next steps reflection: Your structured reflection addressing the relevant dimensions above
Returns: Returns:
Confirmation that reflection was recorded for decision-making Confirmation that reflection was recorded for decision-making
-8
View File
@@ -146,14 +146,6 @@ class TestFormatToolCompact:
result = format_tool_compact("task", {"other": "value"}) result = format_tool_compact("task", {"other": "value"})
assert result == "Cooking with sub-agent" assert result == "Cooking with sub-agent"
def test_load_skill(self):
result = format_tool_compact("load_skill", {"skill_name": "vllm"})
assert result == "load_skill(vllm)"
def test_load_skill_name_key(self):
result = format_tool_compact("load_skill", {"name": "peft"})
assert result == "load_skill(peft)"
def test_tavily_search(self): def test_tavily_search(self):
result = format_tool_compact("tavily_search", {"query": "python testing"}) result = format_tool_compact("tavily_search", {"query": "python testing"})
assert result == "tavily_search(python testing)" assert result == "tavily_search(python testing)"