hamel-husain

Build production-grade eval sets with human-validated LLM-as-judge workflows.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/voidborne-d/master-skill --skill hamel-husain
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: hamel-husain
Source: https://github.com/voidborne-d/master-skill/tree/main/prototypes/monetize-agents-master/output/sub-skills/hamel-husain
Command: npx skills add https://github.com/voidborne-d/master-skill --skill hamel-husain

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Identify and close AI agent reliability gaps by building production-grade eval sets before making prompts or model changes.

Core Features & Use Cases

  • Evals-first workflow: manual trace review, human-validated rubrics, and LLM-as-judge alignment.
  • Role-play and identity guidance: define role-specific agent behavior and governance.
  • Production tracing: instrumentation and measurement to support continuous improvement.

Quick Start

Identify your current agent's evals gap and align on building an evals-first engagement plan.

Frequently Asked Questions about hamel-husain

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I build production-grade LLM eval sets before changing prompts or models?▼

Production-grade LLM eval sets require a domain expert to manually annotate 50 traces, implement an LLM-as-judge with human validation, and establish instrumentation for continuous trace collection and evaluation.

What is an evals-first workflow for AI agents?▼

An evals-first workflow prioritizes manual trace review, human-validated rubrics, and LLM-as-judge alignment to identify and close AI agent reliability gaps before making any prompt or model modifications.

How do I validate an LLM-as-judge against human annotations?▼

Validating an LLM-as-judge involves having a domain expert manually annotate traces to create a baseline, then aligning the automated judge scoring against these human-validated rubrics to ensure measurement accuracy.

How many manually annotated traces do I need for reliable LLM evaluation?▼

Reliable LLM evaluation requires a domain expert to manually annotate 50 production traces to build a human-validated baseline for establishing accurate LLM-as-judge rubrics.

Does setting up AI agent tracing require production instrumentation?▼

Yes, production tracing requires instrumentation to measure agent behavior, collect traces continuously, and support ongoing evaluation to guide system design and model improvements across different domains.

When should I not use an evals-first approach for AI projects?▼

An evals-first approach is not suitable for AI projects lacking production traces or when a domain expert is unavailable to manually annotate 50 traces for human-validated rubric alignment.