eval

Define eval criteria and track pass@k metrics for AI development.

24|5|Updated Feb 8, 2026
One-click install
npx skills add https://github.com/Luohaothu/everything-codex --skill eval-luohaothu
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: eval
Source: https://github.com/Luohaothu/everything-codex/tree/main/skills/eval
Command: npx skills add https://github.com/Luohaothu/everything-codex --skill eval-luohaothu

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Evals before coding establish expectations and enable continuous quality checks for AI systems.

Core Features & Use Cases

  • Define capability and regression evals to guide development.
  • Track pass@k metrics and generate reports to monitor regressions.
  • Use in AI product workflows to ship reliable capabilities.

Quick Start

Run an evaluation plan by defining evals and executing checks with the /eval commands.

Frequently Asked Questions about eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define eval criteria before coding to guide AI development?▼

To define eval criteria before coding, you establish formal evaluation expectations and execute checks using standardized commands. This eval-driven approach enables continuous quality assurance across feature development and model updates.

What is eval-driven development and how does it ensure AI quality?▼

Eval-driven development is a process where you define capability and regression evaluations before writing code. It ensures AI quality by establishing expectations early and tracking pass@k metrics to monitor for regressions across release cycles.

Does this evaluation framework support regression testing for model updates?▼

Yes, the evaluation framework supports regression testing for model updates. It tracks pass@k metrics and generates reports to monitor regressions, ensuring reliable capabilities are shipped across release cycles.

Can I use multiple grader types for AI capability evaluations?▼

Yes, you can use multiple grader types for AI capability evaluations. The framework supports various grader types and standardized workflows to assess and ensure the quality of AI-driven development.

What is the best way to store and manage evaluation data for AI systems?▼

The best way to store and manage evaluation data is using a formal eval framework with standardized workflows. It provides eval storage and tracks pass@k metrics to maintain continuous quality checks for AI systems.