operations-incident

Triage and mitigate production incidents using a structured playbook.

Updated Jan 5, 2025
One-click install
npx skills add https://github.com/pkuppens/pkuppens --skill operations-incident
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: operations-incident
Source: https://github.com/pkuppens/pkuppens/tree/main/skills/operations/operations-incident
Command: npx skills add https://github.com/pkuppens/pkuppens --skill operations-incident

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Incidents in production can cause downtime and confusion; this Skill provides a repeatable, fast-response workflow to triage, mitigate, and document incidents, reducing response time and improving incident reports.

Core Features & Use Cases

  • Triage steps, mitigation actions, status communications, and post-incident reporting.
  • Use cases include degraded services, alert fires, and root-cause investigations with structured timelines.

Quick Start

Follow the incident playbook to triage, mitigate, and compose a post-incident timeline.

Frequently Asked Questions about operations-incident

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I triage a production incident when an alert fires?▼

To triage a production incident when an alert fires, follow a structured playbook that identifies the issue across services and deployment environments, then applies mitigation actions to reduce downtime.

What is the best way to document a post-incident timeline for root-cause analysis?▼

The best way to document a post-incident timeline for root-cause analysis is to enforce a structured post-incident reporting workflow that captures verification steps and mitigation actions.

Can I use this incident response workflow across different deployment environments and data stores?▼

Yes, you can use this incident response workflow across different deployment environments and data stores because the playbook applies to outages and alert fires spanning multiple services.

What steps should I follow to mitigate degraded services during an outage?▼

To mitigate degraded services during an outage, follow the incident playbook to identify the root-cause, apply mitigation actions, execute verification steps, and compose a post-incident timeline.

Why do I need a structured playbook for on-call incident response?▼

You need a structured playbook for on-call incident response because production incidents cause downtime and confusion, and a repeatable fast-response workflow reduces response time and improves reports.

Does this incident triage process include status communications and post-incident reporting?▼

Yes, this incident triage process includes status communications and post-incident reporting as core features to manage degraded services, alert fires, and root-cause investigations.