ERQA

Evaluates multimodal QA models on spatial reasoning and embodied visual grounding benchmarks.

Updated May 7, 2026
One-click install
npx skills add https://github.com/EurecaMoment/BenchClaw --skill erqa
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: ERQA
Source: https://github.com/EurecaMoment/BenchClaw/tree/main/BenchClaw/benchmarkDatasetCards/ERQA
Command: npx skills add https://github.com/EurecaMoment/BenchClaw --skill erqa

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

This Skill provides a reusable benchmark-dataset card for ERQA, enabling models to evaluate spatial reasoning, world-knowledge understanding, and embodied visual grounding.

Core Features & Use Cases

  • Multimodal Question-Answering: Evaluates models on answering questions based on both text and images.
  • Spatial Reasoning: Tests models' ability to reason about spatial relations and positions.
  • World Knowledge: Assesses models on combining real-world semantics with visual evidence.
  • Embodied Reasoning: Evaluates models on movement, viewpoint changes, and visibility.
  • Use Case: For a model to assess its capabilities in multi-image reasoning and spatial relations, it can use the ERQA dataset as a benchmark.

Quick Start

Run the ERQA Skill to evaluate your model's performance on spatial reasoning and multimodal question-answering tasks.

Frequently Asked Questions about ERQA

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate spatial reasoning in multimodal question-answering models?▼

To evaluate spatial reasoning in multimodal question-answering models, you can use the ERQA benchmark dataset to assess capabilities in multi-image reasoning, spatial relations, and embodied visual grounding.

What benchmarks test embodied reasoning and world-knowledge understanding?▼

The ERQA benchmark dataset tests embodied reasoning and world-knowledge understanding by evaluating models on movement, viewpoint changes, visibility, and combining real-world semantics with visual evidence.

Can I use this dataset to assess multi-image reasoning capabilities?▼

Yes, you can use the ERQA dataset to assess multi-image reasoning capabilities. It specifically provides benchmark tasks for evaluating models on spatial relations and embodied visual grounding across multiple images.

Does multimodal question-answering require world knowledge for visual grounding?▼

Multimodal question-answering requires world knowledge for visual grounding. The ERQA dataset evaluates how effectively models combine real-world semantics with visual evidence to answer complex spatial questions.

What is the best way to benchmark embodied visual grounding tasks?▼

The best way to benchmark embodied visual grounding tasks is using the ERQA dataset, which provides structured evaluations for viewpoint changes, movement, and spatial relations in multimodal question-answering scenarios.

When do I need a benchmark for spatial relations and world knowledge?▼

You need a benchmark for spatial relations and world knowledge when assessing a model's ability to reason about spatial positions and combine real-world semantics with visual evidence in multimodal tasks.