eval-harness

Evaluate agent accuracy, efficiency, alignment, and quality across sessions.

7|1|Updated Feb 5, 2026
One-click install
npx skills add https://github.com/besync-labs/antigravity-ai-kit --skill eval-harness-besync-labs
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: eval-harness
Source: https://github.com/besync-labs/antigravity-ai-kit/tree/main/.agent/skills/eval-harness
Command: npx skills add https://github.com/besync-labs/antigravity-ai-kit --skill eval-harness-besync-labs

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Evaluation of agent performance across accuracy, efficiency, alignment, and quality to guide improvements.

Core Features & Use Cases

  • Evaluation Dimensions: Accuracy, Efficiency, Alignment, and Quality with concrete questions to assess behavior.
  • Evaluation Metrics: Target thresholds (e.g., first-time success rate >80%, iterations <3, test coverage >80%, build 100%).
  • Report Format: Generates structured evaluation reports and learnings for future sessions.
  • Integration: Run at session end for learning and continuous improvement.

Quick Start

Run the evaluation harness at the end of a session to measure agent performance on the selected task.

Frequently Asked Questions about eval-harness

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I measure AI agent performance across different tasks?▼

Measuring AI agent performance involves applying an evaluation harness to quantify accuracy, efficiency, alignment, and quality across standardized testing scenarios, enabling repeatable metrics for cross-task comparisons and continuous improvement.

What metrics should I use for agent evaluation and benchmarking?▼

Agent evaluation metrics should include target thresholds like first-time success rate over 80%, iterations under 3, test coverage over 80%, and successful builds. These standardized benchmarks quantify performance across accuracy, efficiency, alignment, and quality dimensions.

How do I generate standardized reports for AI testing sessions?▼

Generate standardized reports for AI testing by running an evaluation harness at the end of a session. This produces structured evaluation reports capturing performance learnings and metrics to guide future improvements and continuous quality assurance.

Can I use an evaluation harness to compare agent performance across sessions?▼

Yes, you can use an evaluation harness to compare agent performance across sessions. It standardizes testing scenarios with defined dimensions and target metrics, producing repeatable results that enable direct cross-task comparisons and track continuous improvement over time.

When should I run agent performance evaluation in my workflow?▼

Run agent performance evaluation at the end of a session. Integrating the evaluation harness at this stage captures structured learnings and generates reports for continuous improvement without disrupting active task execution.