harness-eval

Evaluate software engineering tasks with a 15-task benchmark and scoring rubric.

13|6|Updated Apr 14, 2026
One-click install
npx skills add https://github.com/baekenough/second-brain --skill harness-eval
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: harness-eval
Source: https://github.com/baekenough/second-brain/tree/main/.claude/skills/harness-eval
Command: npx skills add https://github.com/baekenough/second-brain --skill harness-eval

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Structured SE task evaluation uses a 15-task benchmark to provide objective scores for agent performance.

Core Features & Use Cases

  • 15-task benchmark suite with a standardized scoring rubric.
  • Presets (all / quick) for rapid benchmarking and comparison.
  • Integrates with the evaluator-optimizer workflow to produce per-task scores and aggregate grade.

Quick Start

Run harness-eval with all benchmarks to generate the full results report.

Frequently Asked Questions about harness-eval

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
What is a structured software engineering benchmark and how does it evaluate tasks?▼

A structured software engineering benchmark evaluates agent performance using a standardized 15-task harness suite. It applies a standardized rubric across API design, data modeling, authentication flow, logging, and configuration to generate objective per-task scores and an aggregate grade.

How do I run a software engineering benchmark evaluation on all tasks?▼

To run a software engineering benchmark evaluation, use the all preset to execute the full 15-task harness suite. This generates per-task scores, an aggregate grade, and exports the formatted results to the predefined .claude/outputs path.

Can I use a quick preset for rapid software engineering benchmarking and comparison?▼

Yes, you can use the quick preset for rapid software engineering benchmarking. It runs a faster subset of the 15-task suite to produce objective scores quickly, enabling rapid comparison of agent performance across standard tasks.

Does the benchmark evaluation integrate with an evaluator-optimizer workflow?▼

Yes, the benchmark evaluation integrates directly with the evaluator-optimizer workflow. It formats per-task scores and aggregate grades into outputs optimized for the evaluator-optimizer pipeline to consume and process.

What software engineering tasks are covered by the benchmark scoring rubric?▼

The benchmark scoring rubric covers 15 tasks including API design, data modeling, authentication flow, logging, and configuration. It generates objective per-task scores and an aggregate grade for each.

Where are the software engineering benchmark results exported?▼

Software engineering benchmark results are exported to the predefined .claude/outputs path. The evaluation outputs are formatted specifically for consumption by the evaluator-optimizer pipeline.