l5-antifragile_big-data-cherry-picking

Explains how Big Data enables spurious correlations and publication cherry-picking.

Updated Jun 29, 2026
One-click install
npx skills add https://github.com/curation-labs/taleb-mind --skill l5-antifragile-big-data-cherry-picking-curation-labs
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: l5-antifragile_big-data-cherry-picking
Source: https://github.com/curation-labs/taleb-mind/tree/main/skills/l5-antifragile_big-data-cherry-picking
Command: npx skills add https://github.com/curation-labs/taleb-mind --skill l5-antifragile-big-data-cherry-picking-curation-labs

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve? It helps analysts, researchers, and decision-makers recognize why large datasets produce spurious correlations and why published Big Data findings are often noise dressed up as signal. ## Core Features & Use Cases - Spurious Correlation Analysis: Explains the combinatorial growth of false patterns as dataset variables increase. - Incentive Critique: Frames publication bias as a free-option trade where researchers keep upside and bury failed hypotheses. - Use Case: When evaluating a data-driven study claiming a statistically significant finding, apply this reasoning to question how many hypotheses were tested and whether skepticism scaled with data volume. ## Quick Start Ask the AI to apply the Big Data cherry-picking critique to evaluate a research claim or data-driven finding you are reviewing.

Frequently Asked Questions about l5-antifragile_big-data-cherry-picking

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
Why does more data lead to more spurious correlations?▼

As the number of variables in a dataset grows, the number of possible correlations grows combinatorially. With enough variables, some patterns will appear statistically significant purely by chance, so the false discovery rate increases rather than decreases with more data.

What is the cherry-picking problem in Big Data research?▼

Researchers can mine enormous datasets, test thousands of hypotheses, and publish only the ones that appear to work. The failed hypotheses are never reported, so a large proportion of published findings are noise presented as signal.

How do publication incentives create false discoveries?▼

Researchers collect the upside of publication such as citations and tenure while bearing no cost for false positives. Since failed hypotheses and replications are rarely published, the incentive structure guarantees a high proportion of false discoveries.

When should I be skeptical of a data-driven finding?▼

Be skeptical when a finding comes from mining a large dataset without disclosure of how many hypotheses were tested. More data requires a proportional increase in skepticism, not less, because the potential for self-deception grows with data volume.

What are the limitations of this critique of Big Data?▼

This is a conceptual and epistemological argument, not a statistical tool. It does not compute false discovery rates or test specific datasets; it provides a reasoning framework for evaluating research claims and incentive structures.