llmbench

Run a prompt across four models in parallel and render a 2x2 comparison video.

1|Updated Feb 14, 2026
One-click install
npx skills add https://github.com/az9713/my-agent-skills --skill llmbench
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: llmbench
Source: https://github.com/az9713/my-agent-skills/tree/main/.claude/skills/llmbench
Command: npx skills add https://github.com/az9713/my-agent-skills --skill llmbench

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This skill eliminates the guesswork in comparing large language models by automating a full benchmark pipeline that runs a given prompt across four models in parallel and produces a shared, visual comparison video.

Core Features & Use Cases

  • Parallel model execution across four model endpoints using OpenCode to speed up benchmarking.
  • Remotion-based 2x2 video grid that visually juxtaposes outputs for easy assessment.
  • Use cases include prompt design evaluation, model capability comparison, and end-to-end workflow testing across model families.

Quick Start

Run the llmbench pipeline with a prompt to test four models in parallel and render the comparison video.

Frequently Asked Questions about llmbench

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark multiple LLMs in parallel with a single prompt?▼

You can benchmark multiple LLMs in parallel by running a single prompt across four model endpoints simultaneously using OpenCode, which automates the execution to speed up the comparison process.

Can I render a side-by-side video comparison of LLM outputs?▼

Yes, you can render a side-by-side video comparison of LLM outputs using a Remotion-based 2x2 video grid that visually juxtaposes the results for easy assessment.

What is the best way to compare model capabilities across different families?▼

The best way to compare model capabilities across different families is using an automated benchmark pipeline that executes your prompt across four models in parallel and renders a visual comparison video.

Do I need Remotion to visualize prompt design evaluation results?▼

Yes, you need Remotion to visualize prompt design evaluation results because it renders the 2x2 video grid that juxtaposes the parallel outputs for easy assessment.

Does the benchmark workflow support end-to-end orchestration for testing model families?▼

Yes, the benchmark workflow supports end-to-end orchestration for testing model families by handling parallel task execution, Remotion video rendering, and Bash-based workflow automation.

Why use a 2x2 video grid for LLM evaluation instead of text logs?▼

You use a 2x2 video grid for LLM evaluation because it visually juxtaposes outputs in a shared format, eliminating the guesswork in comparing model capabilities and prompt design results.