selection-science-auditor

Audit hiring assessments and evaluation designs for validity, reliability, and fairness.

1|Updated Aug 3, 2026
One-click install
npx skills add https://github.com/getyak/talent-signal --skill selection-science-auditor-getyak
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: selection-science-auditor
Source: https://github.com/getyak/talent-signal/tree/main/.agents/skills/selection-science-auditor
Command: npx skills add https://github.com/getyak/talent-signal --skill selection-science-auditor-getyak

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve? Hiring tools and AI products often make unsupported claims about candidate quality, fit, or potential, and product evaluations often rely on self-ratifying LLM judges or averaged metrics that hide catastrophic errors. This Skill applies industrial-organizational selection science to challenge those claims and enforce rigorous evaluation design. ## Core Features & Use Cases - Selection Boundary Audit: Classifies claims about candidates, vetoes unsupported scoring of personality, culture fit, potential, or acceptance likelihood derived from conversations, and blocks protected-trait inference and candidate ranking by sentiment or responsiveness. - Evaluation Design Audit: Reviews rubrics, graders, gold sets, and experiments used to assess the product itself, requiring deterministic checks, calibrated LLM judges with order-swap and repeat-stability testing, and human gold-set comparison. - Structured Verdicts: Scores nine dimensions on a 0-4 rubric (construct, job analysis, reliability, validity, fairness, uncertainty, eval-set quality, grader integrity, outcome inference) and returns pass, pass_with_changes, fail, or abstain. - Use Case: Before shipping a feature that scores candidate responsiveness, run this audit to confirm the construct is defined, the criterion is job-relevant, subgroup effects are measured, and no prohibited candidate ranking slips into the product. ## Quick Start Use the selection-science-auditor to review this candidate scoring feature and tell me whether its validity and fairness evidence holds up.

Frequently Asked Questions about selection-science-auditor

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I audit an AI hiring assessment for validity?▼

Define the target job and observable criterion first, then check whether the predictor is tied to actual job requirements, uses structured behaviorally anchored scoring, and has local validation evidence. Published average validity from other populations does not transfer to an unvalidated implementation.

How do I evaluate an LLM judge for product evaluation?▼

Bind the judge to atomic rubric criteria with evidence citations, blind it where practical, and run order-swap and repeat-stability checks. Compare its scores against a human gold set, and never let the same unconstrained model generate and ratify its own answer.

What candidate assessments are prohibited in recruiting chat products?▼

Prohibited items include candidate quality, personality, culture fit, potential, or acceptance scores derived from chats, protected-trait inference or proxies, and ranking candidates by responsiveness, sentiment, or communication style. Only explicit facts like deadlines, constraints, and commitments are permitted.

When should an evaluation audit return abstain instead of fail?▼

Return abstain when the target population, criterion, sample, implementation, or results are unavailable, so no evidence-based judgment is possible. Fail applies when there is prohibited candidate scoring, an undefined consequential construct, or a gate hidden by an average.

Why should critical errors not be averaged into overall evaluation scores?▼

Averaging lets a rare catastrophic failure, such as a wrong candidate identity or unauthorized external action, be hidden inside a pleasant mean. Critical identity, privacy, evidence, and action-write failures must be treated as gates, while usability dimensions may be scored separately.