What problem does it solve? Production systems fail silently when teams lack a defined monitoring strategy: alerts fire without actionability, baselines are guessed rather than observed, and capacity issues surface only after they become incidents. This Skill structures how teams decide what to observe, what healthy looks like, and which signals deserve alerts. ## Core Features & Use Cases - Monitoring Strategy Definition: Establishes observation coverage across health, performance, and capacity dimensions for every critical user path. - Signal and Threshold Design: Grounds thresholds in observed baselines with recorded rationale instead of guesses. - Capacity Forecasting: Projects trends from observed data to feed procurement and scaling decisions before capacity becomes an incident. - Alert-Worthiness Judgement: Ensures every alert is actionable and routes everything else to dashboards, reducing alert fatigue. - Use Case: When onboarding a new service, use this Skill to define its signal catalogue, set baseline-derived thresholds, and produce a capacity forecast that infrastructure teams can act on. ## Quick Start Ask the AI to define a monitoring strategy with signals, thresholds, and a capacity forecast for your service using its architecture and incident history.