What problem does it solve? Production systems fail silently without proper telemetry, and teams struggle to detect, diagnose, and resolve incidents before users are impacted. This Skill provides expert guidance for building comprehensive observability systems covering metrics, logs, traces, alerting, and reliability targets. ## Core Features & Use Cases - Monitoring & Metrics Design: Architect Prometheus, Grafana, DataDog, or CloudWatch monitoring with recording rules, dashboards, and high-cardinality handling. - Distributed Tracing & Logging: Implement OpenTelemetry instrumentation, Jaeger tracing, and ELK/Loki log aggregation for root cause analysis. - SLI/SLO & Incident Response: Define service level objectives, error budgets, alert routing with PagerDuty, and blameless postmortem workflows. - Use Case: A platform team running 50 microservices needs to meet a 99.9% availability target. Use this Skill to design the SLI/SLO framework, deploy OpenTelemetry tracing, build Grafana dashboards, and configure noise-reduced alerting with escalation policies. ## Quick Start Ask the assistant to design a monitoring and alerting strategy with SLOs for your microservices architecture.