incident-response

Triage production incidents and provide mitigation steps and post-mortem analysis.

Updated Aug 23, 2026
One-click install
npx skills add https://github.com/4asaanAI/Claude-patches --skill incident-response-4asaanai
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: incident-response
Source: https://github.com/4asaanAI/Claude-patches/tree/main/framework-foundry/Claude%20Plugins/layaa-ai/skills/incident-response
Command: npx skills add https://github.com/4asaanAI/Claude-patches --skill incident-response-4asaanai

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Provides a structured process to triage, diagnose, mitigate, communicate, and perform post-mortems for production incidents and outages, reducing downtime and coordinating clear, timely responses.

Core Features & Use Cases

  • Triage & Prioritization: Rapidly assess impact, scope, and severity to prioritize actions and stakeholders.
  • Diagnosis & Mitigation Guidance: Guide investigations, suggest immediate mitigations or rollbacks, and outline follow-up fixes.
  • Communication & Post-mortem: Draft incident communications for stakeholders and produce a post-mortem with root cause analysis and action items.
  • Use Case: An on-call engineer receives alerts about increased error rates; use this Skill to triage impact, propose immediate mitigations, coordinate on-call actions, and produce a post-mortem.

Quick Start

Triage the production outage for service X by summarizing impact, listing immediate mitigation steps, hypothesizing root causes, and drafting a stakeholder update.

Frequently Asked Questions about incident-response

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I triage a production outage for a microservice?▼

To triage a production outage, assess the incident's impact, scope, and severity to prioritize mitigation actions. This structured approach guides on-call engineers through investigations and proposes immediate mitigations or rollbacks to restore service availability.

What is the best way to write a post-mortem after an incident?▼

Writing a post-mortem involves performing a root cause analysis and outlining specific remediation steps and action items. It provides a structured process to document the incident timeline, diagnose underlying causes, and prevent future outages.

Can I use this incident response process for cloud-based degraded performance?▼

Yes, this incident response process is explicitly applicable to cloud-based and microservice environments. It handles various service availability issues including degraded performance, data integrity problems, and emergency rollbacks.

How do I draft stakeholder communications during a service outage?▼

Drafting stakeholder communications during a service outage uses provided communication templates to deliver clear, timely updates. It ensures coordinated messaging regarding impact assessment, mitigation progress, and resolution expectations.

What immediate mitigation steps should I take for increased error rates?▼

For increased error rates, immediate mitigation steps include hypothesizing root causes and executing emergency rollbacks or configuration changes. The process guides investigations to propose quick fixes that stabilize service availability.