What problem does it solve? Hiring tools and AI products often make unsupported claims about candidate quality, fit, or potential, and product evaluations often rely on self-ratifying LLM judges or averaged metrics that hide catastrophic errors. This Skill applies industrial-organizational selection science to challenge those claims and enforce rigorous evaluation design. ## Core Features & Use Cases - Selection Boundary Audit: Classifies claims about candidates, vetoes unsupported scoring of personality, culture fit, potential, or acceptance likelihood derived from conversations, and blocks protected-trait inference and candidate ranking by sentiment or responsiveness. - Evaluation Design Audit: Reviews rubrics, graders, gold sets, and experiments used to assess the product itself, requiring deterministic checks, calibrated LLM judges with order-swap and repeat-stability testing, and human gold-set comparison. - Structured Verdicts: Scores nine dimensions on a 0-4 rubric (construct, job analysis, reliability, validity, fairness, uncertainty, eval-set quality, grader integrity, outcome inference) and returns pass, pass_with_changes, fail, or abstain. - Use Case: Before shipping a feature that scores candidate responsiveness, run this audit to confirm the construct is defined, the criterion is job-relevant, subgroup effects are measured, and no prohibited candidate ranking slips into the product. ## Quick Start Use the selection-science-auditor to review this candidate scoring feature and tell me whether its validity and fairness evidence holds up.