What problem does it solve? Responding to critical production incidents without a structured process leads to missed context, wasted queries on expired telemetry, and premature conclusions. This Skill enforces a disciplined incident response workflow that verifies data availability, validates findings against the incident description, and keeps customer impact front and center. ## Core Features & Use Cases - Pre-flight Incident Checklist: Confirms incident ID, reporter, description, and a complete timestamp (date, time, timezone) before any investigation begins, with data-retention guidance based on incident age. - Initial Assessment Workflow: Identifies affected components via health checks, validates findings against the incident description, and determines customer impact. - Health Check Procedures: Component-level and global cross-system health analysis using dashboards and knowledge base documents. - Incident Status Updates: Generates delta reports comparing current health against previous assessments to show whether the situation is improving or worsening. - Use Case: An on-call engineer receives incident CI-326 reporting checkout failures at 2026-03-05 14:30 UTC. The Skill walks them through the checklist, runs component health checks for the affected services, confirms customer impact, and produces a summary for stakeholders. ## Quick Start Investigate critical incident CI-326 that started at 2026-03-05 14:30 UTC affecting the checkout service and give me an initial assessment.