refactor(prompts, utils, subagent, skill_manager, think): update skill usage instructions and enhance reflection tool for better decision-making
This commit is contained in:
+37
-6
@@ -14,9 +14,20 @@ into reproducible experiments and a paper-ready experimental report.
|
||||
- Change one major variable per iteration (data, model, objective, or training recipe).
|
||||
- Never invent results. If you cannot run something, say so and propose the smallest next step.
|
||||
- Delegate aggressively using the `task` tool. Prefer the research sub-agent for web search.
|
||||
- Use local skills via `load_skill` when they match the task. Skills provide proven workflows and checklists.
|
||||
- Use local skills when they match the task. Your available skills are listed in the system prompt — read the relevant SKILL.md for full instructions.
|
||||
All skills are available under `/skills/` (read-only).
|
||||
When calling `load_skill`, use the skill id from the SKILL.md frontmatter (`name:`), not the folder name.
|
||||
|
||||
## Research Lifecycle (when applicable)
|
||||
For end-to-end research projects, the recommended skill sequence is:
|
||||
1. `research-ideation` — Explore the field, identify problems and opportunities
|
||||
2. `idea-tournament` — Generate and rank candidate ideas via tree-search + Elo tournament
|
||||
3. `paper-planning` — Plan the paper structure, experiments, and figures
|
||||
4. `experiment-pipeline` — Execute experiments through 4-stage validation
|
||||
5. `paper-writing` — Draft the paper following structured workflow
|
||||
6. `paper-review` — Self-review across quality dimensions
|
||||
7. `paper-rebuttal` — Respond to reviewer comments (if applicable)
|
||||
Not every project needs all steps. Match the starting point to what the user already has.
|
||||
Read the appropriate skill's SKILL.md for workflow guidance at each phase.
|
||||
|
||||
## Scientific Rigor Checklist
|
||||
- Validate data and run quick EDA; document anomalies or data leakage risks.
|
||||
@@ -29,6 +40,9 @@ into reproducible experiments and a paper-ready experimental report.
|
||||
## Step 1: Intake & Scope
|
||||
- Read the proposal and extract goals, datasets, constraints, and evaluation metrics
|
||||
- Capture key assumptions and open questions
|
||||
- Check `/memory/` for prior research knowledge: `ideation-memory.md` (known promising and
|
||||
failed directions) and `experiment-memory.md` (proven strategies from past cycles).
|
||||
Incorporate relevant findings into planning. Skip if these files do not exist yet.
|
||||
- Save the original proposal to `/research_request.md`
|
||||
|
||||
## Step 2: Plan (Recommended Structure)
|
||||
@@ -36,8 +50,7 @@ into reproducible experiments and a paper-ready experimental report.
|
||||
- Identify resource/data dependencies and baseline requirements
|
||||
- Use `write_todos` to track the execution plan and updates
|
||||
- If delegating planning to planner-agent, start your message with: `MODE: PLAN`
|
||||
- If a stage matches an existing skill, note the skill name in the plan and load it before implementation.
|
||||
Use the skill id from SKILL.md frontmatter (`name:`).
|
||||
- If a stage matches an existing skill, note the skill name in the plan and read its SKILL.md before implementation.
|
||||
-- Save the plan to `/todos.md` (recommended). Include per-stage:
|
||||
- objective and success signals
|
||||
- what to run (commands/scripts)
|
||||
@@ -56,7 +69,7 @@ into reproducible experiments and a paper-ready experimental report.
|
||||
- Report drafting → writing-agent
|
||||
- Prefer the research-agent for web search; avoid searching directly
|
||||
- Use `execute` for shell commands when running experiments
|
||||
- When a task matches an existing skill, `load_skill` it and follow it rather than reinventing the workflow.
|
||||
- When a task matches an existing skill, read its SKILL.md and follow it rather than reinventing the workflow.
|
||||
- Keep outputs organized under `/artifacts/` (recommended)
|
||||
- Optionally log runs to `/experiment_log.md` (params, seeds, env, outputs)
|
||||
|
||||
@@ -70,6 +83,23 @@ into reproducible experiments and a paper-ready experimental report.
|
||||
- Update `/todos.md` to reflect new iterations
|
||||
- Stop iterating when evidence is sufficient or diminishing returns appear
|
||||
|
||||
### Memory Evolution (after significant outcomes)
|
||||
After completing or failing a major workflow phase, update research memory using the
|
||||
`evo-memory` skill if installed (read `/skills/evo-memory/SKILL.md`):
|
||||
|
||||
- **After idea-tournament completes**: Run IDE (Idea Direction Evolution).
|
||||
Input: `/direction-summary.md` + user goal. Output: updated `/memory/ideation-memory.md`.
|
||||
- **After experiment-pipeline fails** (no executable code within budget, or method
|
||||
underperforms baseline): Run IVE (Idea Validation Evolution).
|
||||
Input: `/research-proposal.md` + stage trajectory logs.
|
||||
Output: updated `/memory/ideation-memory.md` with failure classification.
|
||||
- **After experiment-pipeline succeeds** (all stages pass): Run ESE (Experiment
|
||||
Strategy Evolution). Input: `/research-proposal.md` + all stage trajectory logs.
|
||||
Output: updated `/memory/experiment-memory.md` with proven strategies.
|
||||
|
||||
If the `evo-memory` skill is not installed, manually update the memory files with key
|
||||
learnings: what worked, what failed, and why.
|
||||
|
||||
### Stage Reflection (Recommended Checkpoint)
|
||||
After any meaningful experimental stage (baseline, new dataset, new training recipe, etc.),
|
||||
delegate a short reflection to the planner-agent and use it to update the remaining plan.
|
||||
@@ -87,7 +117,7 @@ When calling the planner-agent in reflection mode, provide:
|
||||
- Key metrics vs baseline (a small table is ideal)
|
||||
- Artifact paths (logs, plots, checkpoints)
|
||||
- Which success signals were met/unmet
|
||||
- If proposing skills, use skill ids from SKILL.md frontmatter (`name:`).
|
||||
- If proposing skills, use skill names from your available skills listing.
|
||||
|
||||
Ask the planner-agent to output a **Plan Update JSON** with this schema:
|
||||
```json
|
||||
@@ -247,6 +277,7 @@ Capture evaluation protocols (splits, metrics, calibration) and known failure mo
|
||||
## Available Tools
|
||||
1. **tavily_search** - Web search for information
|
||||
2. **think_tool** - Reflect on findings and plan next steps
|
||||
3. **read_file** - Read skill instructions when a skill matches the task (paths shown in your available skills listing)
|
||||
|
||||
**CRITICAL: Use think_tool after each search**
|
||||
|
||||
|
||||
@@ -111,7 +111,7 @@ def format_tool_compact(name: str, args: dict | None) -> str:
|
||||
"""Format as compact tool call string: ToolName(key_arg).
|
||||
|
||||
Adapted for deepagents tool names: execute, read_file, write_file,
|
||||
edit_file, grep, glob, ls, write_todos, read_todos, task, load_skill,
|
||||
edit_file, grep, glob, ls, write_todos, read_todos, task,
|
||||
tavily_search, think_tool.
|
||||
"""
|
||||
if not args:
|
||||
@@ -192,11 +192,6 @@ def format_tool_compact(name: str, args: dict | None) -> str:
|
||||
return f"Cooking with sub-agent — {task_desc}"
|
||||
return "Cooking with sub-agent"
|
||||
|
||||
# Skills
|
||||
if name_lower == "load_skill":
|
||||
skill_name = args.get("skill_name", args.get("name", ""))
|
||||
return f"load_skill({skill_name})"
|
||||
|
||||
# Web search
|
||||
if name_lower in ("tavily_search", "internet_search"):
|
||||
query = args.get("query", "")
|
||||
|
||||
@@ -1,10 +1,16 @@
|
||||
planner-agent:
|
||||
description: "Plan experiments: stages, success signals, and dependencies (no web search, no implementation)."
|
||||
tools: [think_tool]
|
||||
skills: ["/skills/"]
|
||||
system_prompt: |
|
||||
You are the planner-agent. You do NOT implement code. You create and update experimental plans
|
||||
that are practical to run locally.
|
||||
|
||||
Before planning, check `/memory/ideation-memory.md` and `/memory/experiment-memory.md`
|
||||
for prior knowledge from past research cycles. Incorporate relevant entries into
|
||||
your plan (e.g., proven strategies, known failed directions). Skip if these files
|
||||
do not exist yet.
|
||||
|
||||
You may be invoked in two modes:
|
||||
1) PLAN MODE: produce an initial experimental plan.
|
||||
2) REFLECTION MODE: update the plan based on stage results.
|
||||
@@ -47,7 +53,7 @@ planner-agent:
|
||||
}
|
||||
|
||||
Empty arrays are valid. If no changes are needed, return the JSON with empty arrays.
|
||||
"skill_suggestions" must contain skill ids from SKILL.md frontmatter ("name:").
|
||||
"skill_suggestions" should use skill names from your available skills listing.
|
||||
|
||||
Keep the structure flexible (not rigid templates). If model size is unspecified, default to
|
||||
<=7B-class models and lightweight baselines.
|
||||
@@ -55,11 +61,13 @@ planner-agent:
|
||||
research-agent:
|
||||
description: "Web research for methods/baselines/datasets (one topic at a time, return actionable notes + sources)."
|
||||
tools: [tavily_search, think_tool]
|
||||
skills: ["/skills/"]
|
||||
system_prompt_ref: RESEARCHER_INSTRUCTIONS
|
||||
|
||||
code-agent:
|
||||
description: "Implement experiment code and runnable scripts; keep changes minimal and reproducible."
|
||||
tools: [think_tool]
|
||||
skills: ["/skills/"]
|
||||
system_prompt: |
|
||||
You are the code-agent. Implement experiment code in the workspace and keep changes minimal,
|
||||
reproducible, and easy to run.
|
||||
@@ -69,7 +77,9 @@ code-agent:
|
||||
- Record exact commands to run and where outputs are written.
|
||||
- Write outputs under /artifacts/ (recommended) and log key params to /experiment_log.md (optional).
|
||||
- Do not modify /skills/.
|
||||
- If a relevant local skill exists, load it (load_skill) and follow it instead of reinventing.
|
||||
- If a relevant local skill exists, read its SKILL.md and follow its workflow instead of reinventing.
|
||||
- Check `/memory/experiment-memory.md` for proven strategies from past cycles before implementing.
|
||||
Skip if the file does not exist yet.
|
||||
- Before heavy runs, confirm GPU/CUDA/VRAM availability and required packages.
|
||||
- Suggested preflight commands:
|
||||
- nvidia-smi
|
||||
@@ -84,6 +94,7 @@ code-agent:
|
||||
debug-agent:
|
||||
description: "Debug runtime failures and fix bugs with minimal, verifiable patches."
|
||||
tools: [think_tool]
|
||||
skills: ["/skills/"]
|
||||
system_prompt: |
|
||||
You are the debug-agent. Reproduce failures, identify root causes, apply minimal fixes, and provide
|
||||
concise diagnostics.
|
||||
@@ -93,7 +104,7 @@ debug-agent:
|
||||
- Explain the root cause in one paragraph.
|
||||
- Provide how to reproduce and how to verify the fix.
|
||||
- Do not modify /skills/.
|
||||
- If a relevant local skill exists, load it (load_skill) and use it as a checklist.
|
||||
- If a relevant local skill exists, read its SKILL.md and use it as a diagnostic checklist.
|
||||
|
||||
When responding, include:
|
||||
- Root cause
|
||||
@@ -104,6 +115,7 @@ debug-agent:
|
||||
data-analysis-agent:
|
||||
description: "Analyze experiment outputs: compute metrics, make plots, summarize insights."
|
||||
tools: [think_tool]
|
||||
skills: ["/skills/"]
|
||||
system_prompt: |
|
||||
You are the data-analysis-agent. Analyze experiment outputs, compute metrics, and create
|
||||
publication-friendly plots.
|
||||
@@ -112,7 +124,7 @@ data-analysis-agent:
|
||||
- Do not invent numbers; compute from files or state what is missing.
|
||||
- Save figures/tables under /artifacts/ (recommended) and reference paths.
|
||||
- Summarize insights and provide 1-3 recommended next experiments.
|
||||
- If a relevant local skill exists (evaluation, logging, plotting), load it (load_skill) and follow it.
|
||||
- If a relevant local skill exists (evaluation, logging, plotting), read its SKILL.md and follow it.
|
||||
- Report effect sizes and uncertainty (confidence intervals/error bars) when applicable.
|
||||
- Apply multiple-testing corrections when comparing many conditions.
|
||||
- Distinguish exploratory vs confirmatory findings.
|
||||
@@ -125,6 +137,7 @@ data-analysis-agent:
|
||||
writing-agent:
|
||||
description: "Draft a paper-ready Markdown experiment report (no fabricated results/citations)."
|
||||
tools: [think_tool]
|
||||
skills: ["/skills/"]
|
||||
system_prompt: |
|
||||
You are the writing-agent. Draft a clear Markdown experimental report suitable for later paper writing.
|
||||
|
||||
@@ -132,7 +145,7 @@ writing-agent:
|
||||
- Use the experiment plan, logs, and artifacts. Reference file paths for figures/tables.
|
||||
- Do not fabricate results or citations.
|
||||
- If something is missing, add a TODO with the exact command needed to generate it.
|
||||
- If a relevant local skill exists (e.g., evaluation/reporting conventions), load it (load_skill) and apply it.
|
||||
- If a relevant local skill exists (e.g., paper-writing, reporting conventions), read its SKILL.md and follow it.
|
||||
- Report uncertainty, effect sizes, and statistical corrections when relevant.
|
||||
- Include negative results and clear limitations.
|
||||
- Document evaluation protocol (splits/metrics/baselines) and data QC checks.
|
||||
|
||||
@@ -59,7 +59,7 @@ def skill_manager(
|
||||
f"Successfully installed skill: {result['name']}\n"
|
||||
f"Description: {result.get('description', '(none)')}\n"
|
||||
f"Path: {result['path']}\n\n"
|
||||
f"Use load_skill to activate it."
|
||||
f"Read its SKILL.md for full instructions."
|
||||
)
|
||||
else:
|
||||
return f"Failed to install skill: {result['error']}"
|
||||
|
||||
+34
-14
@@ -5,26 +5,46 @@ from langchain_core.tools import tool
|
||||
|
||||
@tool(parse_docstring=True)
|
||||
def think_tool(reflection: str) -> str:
|
||||
"""Tool for strategic reflection on research progress and decision-making.
|
||||
"""Tool for structured reflection and strategic decision-making.
|
||||
|
||||
Use this tool after each search to analyze results and plan next steps systematically.
|
||||
This creates a deliberate pause in the research workflow for quality decision-making.
|
||||
Use this tool to pause and reason carefully at any decision point — not just
|
||||
after searches, but before, during, and after any significant step. This creates
|
||||
a deliberate checkpoint for quality thinking.
|
||||
|
||||
When to use:
|
||||
- After receiving search results: What key information did I find?
|
||||
- Before deciding next steps: Do I have enough to answer comprehensively?
|
||||
- When assessing research gaps: What specific information am I still missing?
|
||||
- Before concluding research: Can I provide a complete answer now?
|
||||
- Before starting work: What do I know? What skills and prior knowledge are available?
|
||||
- After obtaining results: What did I learn? Does this change the approach?
|
||||
- When choosing between options: What are the trade-offs? Which path is strongest?
|
||||
- When stuck or failing: What went wrong? Is there a proven strategy to apply?
|
||||
- Before concluding: Is the evidence sufficient? What does the next phase need from me?
|
||||
|
||||
Reflection should address:
|
||||
1. Analysis of current findings - What concrete information have I gathered?
|
||||
2. Gap assessment - What crucial information is still missing?
|
||||
3. Quality evaluation - Do I have sufficient evidence/examples for a good answer?
|
||||
4. Strategic decision - Should I continue searching or provide my answer?
|
||||
5. Skill leverage - Is there a relevant local skill to load that can accelerate this work?
|
||||
Your reflection should address the relevant dimensions below:
|
||||
|
||||
1. Progress — What has been accomplished? What concrete steps remain?
|
||||
2. Evidence quality — Is the current evidence sufficient for the goal?
|
||||
Would a critical reviewer accept it, or are there gaps to fill?
|
||||
3. Skills leverage — Is there an installed skill that provides a structured
|
||||
workflow for what I'm doing? Check your available skills listing and read
|
||||
the relevant SKILL.md for full instructions. Skills cover various research
|
||||
phases — ideation, experiment execution, paper writing, review, and more.
|
||||
Follow a skill's workflow rather than improvising when one is available.
|
||||
4. Prior knowledge — Have I checked research memory before starting?
|
||||
`/memory/ideation-memory.md` records promising and failed research directions.
|
||||
`/memory/experiment-memory.md` records proven strategies from past cycles.
|
||||
Read these at the start of new work. After completing or failing a task,
|
||||
consider whether the outcome should be recorded back into memory.
|
||||
Skip this if the memory files do not exist yet.
|
||||
5. Strategy — Should I continue the current approach, adjust it, or try
|
||||
something different? What evidence supports this decision?
|
||||
6. Handoff — Is this phase complete? What artifacts and results does the
|
||||
next phase or the caller need? Am I leaving clear, well-organized outputs?
|
||||
|
||||
Not every reflection needs all six dimensions. Pick the ones relevant to
|
||||
the current moment. A focused two or three dimension reflection is better
|
||||
than a shallow pass over all six.
|
||||
|
||||
Args:
|
||||
reflection: Your detailed reflection on research progress, findings, gaps, and next steps
|
||||
reflection: Your structured reflection addressing the relevant dimensions above
|
||||
|
||||
Returns:
|
||||
Confirmation that reflection was recorded for decision-making
|
||||
|
||||
@@ -146,14 +146,6 @@ class TestFormatToolCompact:
|
||||
result = format_tool_compact("task", {"other": "value"})
|
||||
assert result == "Cooking with sub-agent"
|
||||
|
||||
def test_load_skill(self):
|
||||
result = format_tool_compact("load_skill", {"skill_name": "vllm"})
|
||||
assert result == "load_skill(vllm)"
|
||||
|
||||
def test_load_skill_name_key(self):
|
||||
result = format_tool_compact("load_skill", {"name": "peft"})
|
||||
assert result == "load_skill(peft)"
|
||||
|
||||
def test_tavily_search(self):
|
||||
result = format_tool_compact("tavily_search", {"query": "python testing"})
|
||||
assert result == "tavily_search(python testing)"
|
||||
|
||||
Reference in New Issue
Block a user