What problem does it solve? Cloud operations teams struggle with noisy alarms, missing metrics, and unclear thresholds across AWS services. This Skill provides structured troubleshooting workflows for CloudWatch alarms, metric gaps, alert fatigue, and SLI/SLO definitions so engineers can resolve observability issues systematically instead of guessing. ## Core Features & Use Cases - Alarm Investigation Decision Trees: Step-by-step flows for ALARM, INSUFFICIENT_DATA, and flapping alarm states, including evaluation periods, datapoints-to-alarm, and missing-data treatment. - Per-Service Metric & Threshold Reference: Recommended warning/critical thresholds for EC2, RDS, Lambda, ALB, ECS, DynamoDB, ElastiCache, API Gateway, SQS, and SNS with ready-to-run AWS CLI commands. - CloudWatch Best Practices: Metric math expressions (error rate, cache hit ratio), composite alarms, cross-account dashboards, CloudWatch Agent configuration, and EMF for Lambda. - Use Case: An on-call engineer sees a flapping CPU alarm. The Skill guides them to widen evaluation periods, switch to an anomaly detection band, and build a composite alarm combining CPU and error rate to cut noise. ## Quick Start Ask the agent to investigate why a specific CloudWatch alarm is stuck in INSUFFICIENT_DATA and recommend the right threshold and evaluation period settings.