What problem does it solve? It is hard to know whether a new or changed agent skill actually improves agent behaviour, since convincing instructions and self-reported compliance are not evidence. This Skill provides a structured evaluation workflow that produces real evidence about triggering accuracy and task outcomes. ## Core Features & Use Cases - Triggering Evaluation: Test a skill description against realistic prompts that should and should not trigger it, recording expected versus actual results. - Blinded Candidate Comparison: Run the new skill, old skill, or a no-skill baseline on identical tasks, hide candidate identities, and judge artifacts against a shared rubric. - Rubric and Decision Guidance: Write short observable rubrics and reach a clear decision to keep, revise, remove, or gather more evidence. - Use Case: After rewriting a technical-writing skill, run both versions on three real documentation tasks, blind the outputs, score them against a rubric covering factual accuracy and clarity, and decide whether the change ships. ## Quick Start Use the skill-evaluation skill to compare my new skill version against the old one on three realistic tasks with a shared rubric and blinded review.