evaluation-methodology

Automate plugin and skill quality measurement with static analysis, judge scoring, and Monte Carlo simulations.

1|Updated Apr 14, 2026
One-click install
npx skills add https://github.com/Sumeet138/qwen-code-agents --skill evaluation-methodology-sumeet138
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: evaluation-methodology
Source: https://github.com/Sumeet138/qwen-code-agents/tree/main/plugins/plugin-eval/skills/evaluation-methodology
Command: npx skills add https://github.com/Sumeet138/qwen-code-agents --skill evaluation-methodology-sumeet138

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

PluginEval quality methodology provides a structured, objective framework to measure plugin and skill quality using defined evaluation layers, rubrics, and scoring formulas, enabling consistent comparisons and trusted badges.

Core Features & Use Cases

  • Layered evaluation with static analysis, a judge-based scoring process, and Monte Carlo simulations
  • Anchored rubrics across dimensions and badge calibration
  • Guidance for improving triggering accuracy, orchestration fitness, and output quality
  • Use cases: evaluating skills before marketplace publishing; auditing governance; calibrating scores for partner communication

Quick Start

Run plugin-eval score ./path/to/skill to view the static and judge assessment results.

Frequently Asked Questions about evaluation-methodology

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I measure plugin quality using structured rubrics?▼

Measure plugin quality by applying layered evaluation rubrics that automate static analysis, judge-based scoring, and Monte Carlo simulations to produce objective quality metrics and calibrated badges.

What is a judge-based scoring process for skill evaluation?▼

A judge-based scoring process evaluates skills by applying anchored rubrics across multiple dimensions to generate comparative Elo rankings and compute standardized quality badges.

How do I calibrate quality score thresholds for marketplace curation?▼

Calibrate quality score thresholds by interpreting static and judge assessment results, adjusting anchored rubrics, and configuring badge computation to align with marketplace curation standards.

Can I use Monte Carlo simulations to evaluate plugin triggering accuracy?▼

Yes, Monte Carlo simulations evaluate plugin triggering accuracy and orchestration fitness by running probabilistic assessments against anchored rubrics to generate quality metrics.

Does plugin evaluation methodology require external dependencies to run?▼

No, the plugin evaluation methodology operates with zero external dependencies, running static analysis and badge calibration autonomously to produce self-contained quality score reports.

What is the best way to explain quality badges to stakeholders?▼

Explain quality badges to stakeholders by interpreting computed Elo rankings and score analysis outputs, translating calibrated threshold metrics into governance and partner communication summaries.