What problem does it solve? Static benchmark scores and output-only review are losing authority as evidence of AI research quality, and this Skill provides a framework for evaluating AI-produced results through process evidence, contamination checks, and an explicit verification hierarchy. ## Core Features & Use Cases - Verification Hierarchy Tagging: Classify every claim by its verification level, from formal verifiers down to self-assessment, and require at least one claim per project to be checked above self-assessment. - Process-Based Auditing: Require trace logs and code for load-bearing numbers, raising fabrication detection from 55% (paper-only review) to 82%. - Benchmark Contamination Checks: Treat benchmark numbers as contaminated until provenance, leakage paths, and training-data timing are verified. - Use Case: When reviewing an AI-generated research report claiming SOTA results, use this Skill to tag each claim's verification level, re-derive key numbers from logs, check benchmark provenance, and name an independent verifier before accepting the conclusions. ## Quick Start Evaluate this AI-generated research report by tagging each claim's verification level and checking the benchmark provenance and trace logs.