- Added HITL approval lifecycle to ToolCallWidget, including a new "rejected" status.
- Introduced configuration options for auto-approval and shell command allow list in EvoScientistConfig.
- Developed approval handling functions in display module to manage HITL interrupts and user decisions.
- Created ApprovalWidget for user interaction during approval prompts.
- Enhanced StreamEventEmitter to support interrupt events.
- Updated StreamState to track pending interrupts.
- Implemented tests for HITL functionality, including event structure, state handling, and approval logic.
- Updated SKILL.md for clarity on headless environments
- Adjusted runs_per_configuration logic in aggregate_benchmark.py
- Improved description placeholder in init_skill.py
- Added strict validation checks in quick_validate.py
- Ensured skill-creator root is included in sys.path for script imports
style(eval_review): update button colors and instructions for clarity
style(viewer): modify accent colors and update instructions for agent
docs(output-patterns): replace "Claude" with "agent" for consistency
docs(workflows): replace "Claude" with "agent" for consistency
fix(improve_description): update skill description context to refer to EvoScientist
fix(init_skill): update references to "Claude" to "agent" for consistency
- Implemented `improve_description.py` to enhance skill descriptions based on evaluation results using LLMs.
- Created `run_eval.py` to evaluate skill descriptions against a set of queries, determining trigger effectiveness.
- Developed `run_loop.py` to automate the evaluation and improvement process, tracking history and optimizing descriptions iteratively.
- Added utility functions in `utils.py` for parsing skill metadata from SKILL.md files.
- Enhanced `package_skill.py` to exclude unnecessary files and directories during skill packaging.
- Updated `quick_validate.py` to include compatibility checks in skill validation.
feat(prompts): update system prompt to eliminate numeric limits for sub-agents and delegation rounds
test(tests): adjust tests to reflect changes in configuration and onboarding logic