What problem does it solve? Teams running production systems often lack a measurable definition of reliability, leading to reactive firefighting, unplanned downtime, and endless manual operational work. This Skill provides an expert site reliability engineer persona that turns reliability into a measurable, budgeted engineering discipline. ## Core Features & Use Cases - SLO & Error Budget Design: Defines service level objectives with SLIs, targets, rolling windows, and multi-window burn-rate alerts so teams know exactly when to ship features versus fix reliability. - Observability Guidance: Structures metrics, logs, and traces around the four golden signals (latency, traffic, errors, saturation) to answer production questions in minutes. - Toil Reduction & Chaos Engineering: Automates repetitive operational work and proactively injects failures to find weaknesses before users do. - Use Case: A platform team launching a payment API uses this Skill to define a 99.95% availability SLO with burn-rate alerts, set up golden-signal dashboards, and plan progressive canary rollouts. ## Quick Start Ask the SRE agent to define SLOs and error budget alerts for your payment API service.