What problem does it solve? Teams often lack measurable reliability targets, drown in noisy alerts, and burn engineering time on repetitive operational work. This Skill turns vague reliability goals into concrete SLIs/SLOs, error-budget policies, and toil-reduction plans grounded in what the system can actually deliver. ## Core Features & Use Cases - SLI/SLO Design: Select measurable indicators (latency, error rate, saturation) and set targets that reflect real user-facing experience rather than easy-to-measure low-signal metrics. - Error-Budget & Alert Policy: Define burn-rate thresholds, enforcement triggers, and alert-quality audits so every alert maps to an actionable response with a clear owner. - Toil Reduction & Resilience Architecture: Identify manual operational work worth automating and apply patterns like circuit breakers, bulkheads, retries with backoff, and graceful degradation. - Use Case: Before a major launch, ask for a reliability review of your checkout service to get SLO recommendations, capacity headroom analysis, and a prioritized list of resilience gaps. ## Quick Start Ask the agent to help define SLOs and an error-budget policy for your service based on its current architecture and traffic profile.