result-reliability-checker

Audits biomedical study results for design integrity, statistical risk, validation strength, and claim overreach.

Updated Sep 15, 2026
One-click install
npx skills add https://github.com/mrsonord2240/openscience-specialists --skill result-reliability-checker-mrsonord2240
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: result-reliability-checker
Source: https://github.com/mrsonord2240/openscience-specialists/tree/main/specialists/critical-appraisal-specialist/versions/1.0.0/package/skills/result-reliability-checker
Command: npx skills add https://github.com/mrsonord2240/openscience-specialists --skill result-reliability-checker-mrsonord2240

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve? Readers of biomedical papers often cannot tell whether headline findings are trustworthy, fragile, or overstated. This Skill audits the full evidence chain behind a study's results so you can decide whether to cite, use cautiously, or treat the findings as exploratory only. ## Core Features & Use Cases - Evidence Chain Audit: Reconstructs the data, methods, and validation behind each main claim before judging reliability. - Design, Bias, and Statistics Checks: Evaluates design fit, confounding, leakage risk, sample size adequacy, multiple testing, and overfitting. - Validation Chain Grading: Distinguishes internal validation, external validation, orthogonal validation, replication, and prospective support. - Use Case: Given a machine-learning prognosis paper with impressive AUROC values, the Skill checks for leakage, overfitting, and weak validation, then assigns a per-claim reliability judgment instead of accepting the metrics at face value. ## Quick Start Ask the AI to audit whether the results in this attached paper are reliable enough to cite, including bias, statistics, and validation checks.

Frequently Asked Questions about result-reliability-checker

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I check if a research paper's results are reliable?▼

Provide the paper, abstract, or methods and results text and request a reliability audit. The Skill reconstructs the evidence chain behind each claim, checks design, bias, statistics, and validation, then assigns a per-claim reliability judgment from high to low.

How to assess overfitting and leakage risk in machine-learning medical papers?▼

The audit checks training and test separation, cross-validation discipline, tuning isolation, sample size versus model complexity, and temporal or information leakage. Suspiciously high metrics relative to sample size and design trigger explicit caution.

What is the difference between internal and external validation in study appraisal?▼

Internal validation uses splits or resampling within one cohort and only supports internal consistency. External validation tests the finding in a distinct cohort, dataset, or center, providing stronger but still context-dependent evidence of generalizability.

Can this skill give patient-specific clinical advice?▼

No. Patient-specific clinical decision support is explicitly out of scope, and the Skill responds with a redirect. It only audits whether reported research results are reliable enough to treat as evidence.

Does a high AUROC or significant p-value mean study results are trustworthy?▼

No. The Skill's hard rules state that statistical significance never equals reliability and high AUROC or accuracy does not equal robustness. Calibration, validation strength, bias control, and claim discipline must all support the conclusion.

What happens when a paper lacks enough methodological detail to audit?▼

Missing methods, statistics, or validation steps are labeled as unresolved rather than filled in. The Skill never invents study features, references, PMIDs, or DOIs, and states clearly when judgments rest only on user-provided text.