search-benchmark

Benchmark Switchboard's search tool across opus, sonnet, and haiku models.

15|7|Updated Feb 26, 2026
One-click install
npx skills add https://github.com/daltoniam/switchboard --skill search-benchmark
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: search-benchmark
Source: https://github.com/daltoniam/switchboard/tree/main/.agents/skills/search-benchmark
Command: npx skills add https://github.com/daltoniam/switchboard --skill search-benchmark

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Benchmark Switchboard's search tool across models to measure cross-model search quality for tool discovery.

Core Features & Use Cases

  • Cross-model benchmark across opus, sonnet, and haiku in parallel
  • Generates a comparison table and identifies optimization opportunities
  • Use cases: performance assessment of search suggestions, model-tier differences, and when evaluating Phase 2 tag impact

Quick Start

Run the search-benchmark skill to dispatch identical scenarios to opus, sonnet, and haiku in parallel and generate a comparison report

Frequently Asked Questions about search-benchmark

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark cross-model search quality for tool discovery?▼

You can benchmark cross-model search quality by dispatching identical scenarios to opus, sonnet, and haiku in parallel. This measures search suggestions equally across model tiers and produces a structured comparison report.

What is cross-model search benchmarking used for?▼

Cross-model search benchmarking is used to evaluate search tool quality for tool discovery. It identifies performance differences across model tiers and assesses Phase 2 tag impact to reveal optimization opportunities.

Can I compare search results across opus, sonnet, and haiku simultaneously?▼

Yes, you can compare search results across opus, sonnet, and haiku simultaneously. The benchmark dispatches identical scenarios to all three models in parallel to ensure equal treatment and generate a comparison table.

How do I generate a comparison report for multimodel search results?▼

You generate a comparison report for multimodel search results by running a benchmark that applies synthetic and live-model scenarios across opus, sonnet, and haiku. It collects the results and outputs a structured comparison table.

Does the search benchmark support synthetic and live-model scenarios?▼

Yes, the search benchmark supports both synthetic and live-model scenarios. It applies these scenarios equally across opus, sonnet, and haiku to compare results and measure cross-model search quality.

What are the limitations of multimodel search benchmarking?▼

A limitation of multimodel search benchmarking is that it requires equal treatment across opus, sonnet, and haiku to compare results accurately. It focuses strictly on search tool quality and Phase 2 tag impact rather than general model performance.