sre-engineer

Define SLOs, error budgets, monitoring, and automation scripts for production systems.

1|Updated Jan 19, 2026
One-click install
npx skills add https://github.com/camelranchentertainment/Booking-Platform --skill sre-engineer-camelranchentertainment
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/camelranchentertainment/Booking-Platform/tree/main/.claude/skills/sre-engineer
Command: npx skills add https://github.com/camelranchentertainment/Booking-Platform --skill sre-engineer-camelranchentertainment

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Defines how to establish reliability for production systems by setting SLOs/SLIs, creating error budgets, designing incident response, modeling capacity, and delivering monitoring configurations and automation scripts.

Core Features & Use Cases

  • Define quantitative SLOs/SLIs and error budgets with burn-rate monitoring.
  • Create runbooks, blameless postmortems, and incident response procedures; establish golden signals monitoring.
  • Produce automation scripts and capacity planning models for scalable reliability.

Quick Start

Configure SLOs, error budgets, incident response procedures, and monitoring setups for a production service.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define SLOs and error budgets for production services?▼

Define SLOs and error budgets by establishing quantitative SLIs and configuring burn-rate monitoring to track reliability thresholds for production services.

What is the best way to set up incident management and blameless postmortems?▼

Set up incident management by creating runbooks and incident response procedures, then conduct blameless postmortems to analyze failures without assigning individual fault.

How do I configure golden signals monitoring for scalable services?▼

Configure golden signals monitoring by tracking latency, traffic, errors, and saturation metrics to establish measurable SLIs across scalable service infrastructure.

Can I automate runbooks and capacity planning for site reliability engineering?▼

Yes, you can automate runbooks and generate capacity planning models to reduce operational toil and deliver quantitative reliability automation scripts for scalable services.

When do I need chaos engineering for site reliability engineering?▼

Apply chaos engineering when you need to proactively test production system resilience by running controlled experiments that validate incident response procedures and reliability thresholds.