researcher-evaluation

Evaluate GenAI agents with G-Eval methodology and structured performance reports.

1|Updated Jan 23, 2026
One-click install
npx skills add https://github.com/VibeTechnologies/VibeTeam --skill researcher-evaluation
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: researcher-evaluation
Source: https://github.com/VibeTechnologies/VibeTeam/tree/main/.opencode/skills/researcher-evaluation
Command: npx skills add https://github.com/VibeTechnologies/VibeTeam --skill researcher-evaluation

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Evaluates GenAI agents using the G-Eval methodology to produce structured performance reports.

Core Features & Use Cases

  • Provides a standardized Required Output Format for evaluating multiple frameworks (AutoGen, CrewAI, OpenHands)
  • Includes a comprehensive methodology (G-Eval) that scores accuracy, reasoning, and actionability across tasks
  • Offers a CLI workflow to run benchmarks and compare frameworks, with clear evaluation dimensions and a summary of results

Quick Start

Run the evaluation workflow on a sample task using the G-Eval methodology to compare agent outputs.

Frequently Asked Questions about researcher-evaluation

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I evaluate GenAI agent performance across different frameworks?▼

You can evaluate GenAI agents by applying the G-Eval methodology to score accuracy, reasoning, and actionability. This Skill benchmarks agent performance across frameworks like AutoGen, CrewAI, and OpenHands to generate structured reports.

What is the G-Eval methodology for benchmarking GenAI agents?▼

The G-Eval methodology is a standardized evaluation process that scores GenAI agents on accuracy, reasoning, and feedback. It uses a defined scoring scale and evaluation dimensions to produce structured performance reports.

Can I use this to benchmark AutoGen against CrewAI?▼

Yes, you can benchmark AutoGen against CrewAI. The Skill provides a CLI workflow to run benchmarks and compare agent outputs across multiple frameworks using a standardized required output format.

How do I run a GenAI agent benchmark using a CLI workflow?▼

You can run a GenAI agent benchmark by using the provided CLI workflow to execute the G-Eval methodology on sample tasks. This process compares framework outputs and summarizes results based on defined evaluation dimensions.

What evaluation dimensions are scored when assessing GenAI agents?▼

The GenAI agent evaluation scores dimensions including accuracy, reasoning, and actionability. These metrics are calculated using the G-Eval methodology to provide a comprehensive summary of agent performance.

Do I need any external dependencies to generate GenAI agent evaluation reports?▼

No external dependencies are required to generate GenAI agent evaluation reports. The Skill operates independently to apply the G-Eval methodology and output structured performance summaries.