evaluation-running

Automate running and back-testing AI safety evaluations across model generations.

Updated Feb 24, 2026
One-click install
npx skills add https://github.com/DouwMarx/evaluating-evaluations --skill evaluation-running-douwmarx
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: evaluation-running
Source: https://github.com/DouwMarx/evaluating-evaluations/tree/main/evaluation-running
Command: npx skills add https://github.com/DouwMarx/evaluating-evaluations --skill evaluation-running-douwmarx

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) and references (resource) components.

What problem does it solve?

The evaluation-running skill addresses the challenge of running and back-testing AI safety evaluations across different model generations. It automates the process of validating instruments, executing evals, and analyzing results, enabling efficient and thorough safety assessments.

Core Features & Use Cases

  • Evaluation Automation: Automate the execution of AI safety evaluations across multiple models and generations.
  • Validation of Instruments: Validate the instruments used in evaluations to ensure accuracy and reliability.
  • Analysis of Results: Provide detailed analysis of evaluation results, including statistical tests and trend lines.
  • Use Case: Use this skill to validate an AI safety evaluation instrument across various model generations and analyze the results to identify any safety issues.

Quick Start

Run the evaluation-running skill with the following command: /evaluation-running run-eval --evaluation-id [evaluation_id] --model-panel [model_panel] --budget [budget]

Frequently Asked Questions about evaluation-running

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate AI safety evaluations across multiple model generations?▼

You can automate AI safety evaluations across model generations by running a command that executes evals against a specified model panel and budget, validating instruments, and analyzing results with statistical tests.

What statistical analysis is included when back-testing AI safety evaluations?▼

Back-testing AI safety evaluations includes statistical tests and trend analysis to validate instruments and identify safety issues across different model generations.

How do I validate AI safety evaluation instruments before running them?▼

Validating AI safety evaluation instruments is automated by the skill, which checks accuracy and reliability across the specified model panel before executing the full evaluation.

What do I need to set up before running automated AI safety evaluations?▼

You need access to multiple model generations and evaluation instruments, plus an evaluation ID, model panel, and budget to execute the automated evaluation run.

Can I run AI safety evaluations with a limited compute budget?▼

Yes, you can specify a budget parameter when running evaluations to control resource consumption while still automating execution and statistical analysis across model generations.

Why use automated evaluation running instead of manual AI safety testing?▼

Automating AI safety evaluation running enables efficient back-testing across model generations, validating instruments and generating trend analysis that manual testing cannot easily scale to.