eval-harness

Benchmark pixl-crew skills, agents, and prompts with configurable evaluation runs.

2|Updated Mar 16, 2026
One-click install
npx skills add https://github.com/hamzaPixl/pixl-ai --skill eval-harness-hamzapixl
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: eval-harness
Source: https://github.com/hamzaPixl/pixl-ai/tree/main/packages/crew/skills/eval-harness
Command: npx skills add https://github.com/hamzaPixl/pixl-ai --skill eval-harness-hamzapixl

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This evaluation harness provides a structured, repeatable framework to assess pixl-crew skills, agents, and prompts, enabling objective quality metrics and regression tracking.

Core Features & Use Cases

  • Automated, configurable evaluations (capability and regression) across versions.
  • Generation of structured eval reports and dashboards to surface regressions and performance trends.
  • Test-case discovery and rubric-driven scoring to guide ongoing skill/agent/prompt improvement.

Quick Start

Run pixl with the eval-harness target to start a capability or regression evaluation.

Frequently Asked Questions about eval-harness

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate agent and prompt quality for automated workflows?▼

You benchmark agent and prompt quality by running an evaluation harness that applies structured criteria and rubric-driven scoring to generate per-test reports tracking regressions over time.

What is regression testing for AI prompts and how does it work?▼

Regression testing for AI prompts runs configurable evaluations across versions to validate capability, generating per-test reports that track performance trends and surface regressions over time using structured criteria.

How do I benchmark skills across multiple versions to track performance?▼

Benchmark skills across versions by executing a regression evaluation with a configurable run count, enforcing structured criteria to generate per-test reports that track performance trends over time.

Can I export evaluation reports for skills and prompts?▼

Yes, the evaluation harness supports results export, generating structured per-test reports that surface regressions and performance trends from your configurable capability and regression test runs.

Do I need to define test cases manually to measure prompt quality?▼

No, the evaluation harness features automated test-case discovery, using rubric-driven scoring and structured criteria to measure prompt quality and guide ongoing skill, agent, and prompt improvement.

When should I use a repeatable evaluation harness for testing agents?▼

Use a repeatable evaluation harness when you need objective quality metrics and regression tracking for agents and prompts, enabling capability and regression testing to surface regressions over time.