
running-skills-edd-cycle
by taisukeoe
This is a foundation to create effective Agentic AI Skills
SKILL.md
name: running-skills-edd-cycle description: Guides evaluation-driven development (EDD) process for agent skills. Use when setting up skill testing workflows, creating skill evaluation scenarios, or establishing Claude A/B feedback loops for skill validation. Provides development methodology, not content guidance. license: Apache-2.0 allowed-tools: "Skill(creating-effective-skills) Skill(improving-skills) Skill(reviewing-skills) Skill(evaluating-skills-with-models)" metadata: author: Softgraphy GK version: "0.2.0"
Running Skills EDD Cycle
Run evaluation-driven development cycle for agent skills.
Workflow
Step 1: Build Evaluations First
Create evaluations BEFORE writing documentation. This ensures skills solve real problems.
- Run Claude on representative tasks WITHOUT the skill
- Document specific failures or missing context
- Create 3+ evaluation scenarios that test these gaps
Evaluation scenarios are saved to tests/scenarios.md as the final step of /creating-effective-skills workflow.
Step 2: Establish Baseline
Measure Claude's performance WITHOUT the skill:
- Run each evaluation scenario
- Record: success/failure, missing context, wrong approaches
- This becomes comparison baseline
Step 3: Write Minimal Instructions
Create just enough content to address the gaps:
- Start with core workflow only
- Add detail only when tests fail
- Avoid over-explaining
REQUIRED: Use the Skill tool to invoke creating-effective-skills before writing any skill content. This ensures proper naming, description format, and structure from the start.
Step 4: Evaluate with Multiple Models
Note: This step requires Claude Code CLI. Skip if using Claude.ai.
REQUIRED: Use the Skill tool to invoke evaluating-skills-with-models with the skill path.
This will:
- Auto-load scenarios from
tests/scenarios.md - Execute with sub-agents across models (sonnet, opus, haiku)
- Evaluate against expected behaviors
- Determine recommended model (least capable with full compatibility)
After evaluation: Document recommended model in skill's metadata.
REQUIRED: Use the Skill tool to invoke improving-skills when observations reveal issues.
Step 5: Final Review
Before considering the skill complete:
REQUIRED: Use the Skill tool to invoke reviewing-skills to verify compliance with best practices.
- Address all compliance issues identified
- Re-run evaluations after fixes
- Repeat until skill passes review
Step 6: User Validation Guide
After all reviews pass, output instructions for user to validate in a fresh session:
## Test Your Skill
Run this command in a new terminal to test with a fresh Claude session:
claude --model {recommended_model} "{evaluation_query}"
After testing, paste the output file or result back to this session for final confirmation.
Replace:
{recommended_model}: Model determined in Step 4 (e.g.,sonnet){evaluation_query}: A representative query from your evaluations
Quick Reference
Cycle
Identify gaps -> Create evaluations -> Baseline -> Write minimal -> Model eval (sub-agents) -> Review -> User validation
What Observations Indicate
| Observation | Indicates |
|---|---|
| Unexpected file reading order | Structure not intuitive |
| Missed references | Links need to be explicit |
| Repeated reads of same file | Move content to SKILL.md |
| Never accessed file | Unnecessary or poorly signaled |
Score
Total Score
Based on repository quality metrics
SKILL.mdファイルが含まれている
ライセンスが設定されている
100文字以上の説明がある
GitHub Stars 100以上
3ヶ月以内に更新がある
10回以上フォークされている
オープンIssueが50未満
プログラミング言語が設定されている
1つ以上のタグが設定されている
Reviews
Reviews coming soon