What problem does it solve? When studies conflict or a single exciting trial dominates the conversation, decision-makers lack a disciplined way to say how much trust a whole body of evidence deserves. This Skill applies the published GRADE method (Guyatt et al., 2008; GRADE Handbook, 2013) to rate certainty per outcome — high, moderate, low, or very low — with every downgrade or upgrade tied to a named reason. ## Core Features & Use Cases - Per-outcome GRADE rating: Frames a PICO, starts randomized bodies at High and observational bodies at Low, then applies the five rate-down factors (risk of bias, inconsistency, indirectness, imprecision, publication bias) and three rate-up factors (large effect, dose-response, opposing residual confounding) with explicit reasons. - Deterministic companion tool: scripts/grade.py (standard library only) performs the rating arithmetic, clamps to the four-level scale, prints summary-of-findings rows, and warns when rating up meets rating down or randomized evidence. - Calibrated communication: Produces GRADE-language certainty statements, bans words like "proven" and "no evidence", and names the surveillance trigger — the new evidence that would change the grade. - Use Case: A health researcher has five RCTs of CBT-I versus sleep-hygiene education with conflicting effect sizes. The Skill grades the body per outcome, producing a summary-of-findings row (e.g., ⊕⊕◯◯ Low, −1 inconsistency, −1 imprecision) and a calibrated statement for the recommendation. ## Quick Start Ask the assistant to GRADE the certainty of the evidence for your question, providing the assembled study set with per-study risk-of-bias ratings and any pooled effect estimates.