evaluate-cortex-agent

Evaluate Cortex Agents in Snowflake and review results in Snowsight.

Updated Mar 7, 2026
One-click install
npx skills add https://github.com/randoneering/nix-flake-mirror --skill evaluate-cortex-agent
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: evaluate-cortex-agent
Source: https://github.com/randoneering/nix-flake-mirror/tree/main/home/programs/opencode/skills/snowflake/agent_optimization/evaluate-cortex-agent
Command: npx skills add https://github.com/randoneering/nix-flake-mirror --skill evaluate-cortex-agent

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

This Skill provides a structured, repeatable workflow to evaluate Cortex Agents using Snowflake’s native Agent Evaluations, enabling objective benchmarking and comparison of agent performance across configurations.

Core Features & Use Cases

  • Define evaluation datasets for Cortex Agents and track metrics such as correctness, tool_selection_accuracy, tool_execution_accuracy, and logical_consistency.
  • Automate setup of evaluation runs in Snowflake and generate Snowsight reports.
  • Support scenario-based comparisons to measure improvements after prompts, tool changes, or configuration updates.

Quick Start

Configure the target agent, select metrics, build or choose a dataset, run the evaluation, and review results in Snowsight.

Frequently Asked Questions about evaluate-cortex-agent

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I benchmark a Cortex Agent in Snowflake?▼

To benchmark a Cortex Agent in Snowflake, you select metrics like correctness and logical_consistency, supply an evaluation dataset, run the evaluation, and review the performance results in Snowsight.

What metrics are available for Cortex Agent evaluation?▼

Cortex Agent evaluation metrics include correctness, tool_selection_accuracy, tool_execution_accuracy, and logical_consistency to objectively measure and compare agent performance across configurations.

Can I compare Cortex Agent performance after prompt or tool changes?▼

Yes, you can compare Cortex Agent performance after prompt or tool changes by running scenario-based evaluations to measure improvements and benchmark configurations against previous results in Snowsight.

Do I need to prepare a dataset to evaluate a Cortex Agent?▼

Yes, you need to prepare or supply a dataset to evaluate a Cortex Agent, which is necessary to measure metrics like correctness and tool_selection_accuracy during the evaluation run.

What is the best way to track logical consistency in AI agents?▼

The best way to track logical consistency in AI agents is using Snowflake native Agent Evaluations to benchmark performance, which requires selecting the logical_consistency metric and running an evaluation with a prepared dataset.