What problem does it solve? Teams shipping to production often lack defined reliability targets, actionable alerting, tested disaster recovery, and runbooks an on-call engineer can actually follow at 3 AM. This Skill turns existing infrastructure and monitoring configs into a complete site reliability engineering practice. ## Core Features & Use Cases - Production Readiness Review: Audits Kubernetes manifests, Terraform configs, and application settings against a checklist covering health probes, graceful shutdown, timeouts, retries, and resource limits, producing findings with severity and remediation steps. - SLO & Error Budget Definition: Generates SLI definitions, SLO targets, error budget policies, and multi-window burn-rate alert rules in Prometheus format. - Chaos Engineering & Game Days: Authors Chaos Mesh scenarios (pod failure, network partition, dependency failure, resource pressure) with steady-state hypotheses and a game-day playbook with abort criteria. - Incident Management & Runbooks: Produces severity classifications, on-call rotations, escalation policies, communication templates, war-room procedures, and per-service runbooks with exact commands and decision trees. - Capacity Planning: Builds load models at 1x/10x/100x scale, validated HPA configurations, cost projections, and bottleneck analysis. - Use Case: After deploying a Kubernetes-based API, run this Skill to get SLO dashboards, burn-rate alerts, four chaos scenarios, and runbooks for high error rate, latency, OOM, and dependency failures. ## Quick Start Ask the agent to run the SRE skill against your infrastructure directory to produce a production readiness review, SLO definitions, and on-call runbooks.