Anti-Benchmark

Identify and challenge benchmark assumptions to improve evaluation rigor.

393|34|Updated Feb 10, 2026
One-click install
npx skills add https://github.com/Pthahnix/De-Anthropocentric-Research-Engine --skill anti-benchmark
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: Anti-Benchmark
Source: https://github.com/Pthahnix/De-Anthropocentric-Research-Engine/tree/main/skills/sop/anti-benchmark
Command: npx skills add https://github.com/Pthahnix/De-Anthropocentric-Research-Engine --skill anti-benchmark

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Identify and challenge benchmark assumptions to improve evaluation rigor.

Core Features & Use Cases

  • Ability to critique foundational benchmark claims and surface gaps
  • Works across domains where benchmarks govern evaluation results
  • Facilitates generating alternative benchmarking perspectives for robust testing

Quick Start

Provide a benchmark and its context to receive a structured critique and proposed alternatives.

Frequently Asked Questions about Anti-Benchmark

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I challenge benchmark assumptions to improve evaluation rigor?▼

To challenge benchmark assumptions, provide a benchmark and its context to receive a structured critique containing challenged assumptions, variants, and proposed alternatives for robust testing.

What is benchmark assumption critique and when do I need it?▼

Benchmark assumption critique is the process of stress-testing foundational claims in scientific and engineering benchmarks. You need it to surface gaps and improve evaluation rigor across domains where benchmarks govern results.

How do I generate alternative benchmarking perspectives for robust testing?▼

Generate alternative benchmarking perspectives by invoking a critique workflow on your benchmark and context, which returns an AntiBenchmarkResult containing challenged assumptions, variants, and proposed alternatives.

Does benchmark critique work across different scientific and engineering domains?▼

Yes, benchmark critique works across scientific and engineering domains where foundational claims must be stress-tested. It systematically surfaces gaps to improve evaluation rigor regardless of the specific benchmarking domain.

What do I need to provide to start a benchmark assumption critique?▼

You need to provide a benchmark and its context to the critique workflow. The system processes these inputs and returns a structured result with challenged assumptions, variants, and proposed alternatives.

What is the best way to surface gaps in foundational benchmark claims?▼

The best way to surface gaps in foundational claims is systematic assumption critique. By challenging benchmark assumptions, you receive a structured AntiBenchmarkResult exposing variants and proposed alternatives for more robust evaluation.