observe-production

Evaluate deployed service health using SLOs, error rates, latency, throughput, and alerts.

5|Updated Jul 25, 2025
One-click install
npx skills add https://github.com/tomzx/agents --skill observe-production
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: observe-production
Source: https://github.com/tomzx/agents/tree/main/skills/observe-production
Command: npx skills add https://github.com/tomzx/agents --skill observe-production

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Checks the health of a deployed service or feature by reviewing SLOs, error rates, latency, throughput, and recent alerts. Produces a health report suitable for maintenance reviews, post-deploy verification, or incident triage.

Core Features & Use Cases

  • Monitor SLOs/SLIs and alert status across services.
  • Assess error rates, latency percentiles, throughput, and recent alerts for incident triage.
  • Produce a structured health report for maintenance reviews, post-deploy verification, or incident response.

Quick Start

Check the health of the deployed service and generate a concise production health report.

Frequently Asked Questions about observe-production

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I check the health of a deployed service after a release?▼

To check deployed service health, evaluate SLOs, error rates, latency, throughput, and recent alerts. This Skill aggregates observability data to produce a structured health report suitable for post-deploy verification and incident triage.

What metrics do I need to assess production service health for incident triage?▼

Assessing production service health requires reviewing SLOs, error rates, latency percentiles, throughput, and recent alerts. Aggregating these metrics helps generate a structured health report for incident triage and routine maintenance reviews.

Can I use this to monitor SLO and SLI status across individual services?▼

Yes, you can monitor SLO and SLI status across individual services or features. It evaluates error rates, latency, throughput, and alert status to generate a structured health report for maintenance reviews and incident response.

How do I generate a structured health report from observability data?▼

Generate a structured health report by aggregating observability data from monitoring tools to evaluate SLOs, error rates, latency, and throughput. This produces a concise report suitable for maintenance reviews and troubleshooting.

What is the best way to verify SLO compliance and detect production issues fast?▼

The best way to verify SLO compliance and detect production issues is by evaluating service health through error rates, latency, throughput, and alerts. This approach produces a structured health report for fast incident triage and troubleshooting.