agency-incident-response-commander

Coordinate production incident response and post-mortem facilitation for distributed services.

Updated Apr 11, 2026
One-click install
npx skills add https://github.com/omeraltn/ice_cream_website_testing --skill agency-incident-response-commander-omeraltn
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: agency-incident-response-commander
Source: https://github.com/omeraltn/ice_cream_website_testing/tree/main/.antigravity/agency-incident-response-commander
Command: npx skills add https://github.com/omeraltn/ice_cream_website_testing --skill agency-incident-response-commander-omeraltn

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

This Skill helps engineering teams turn chaotic production incidents into structured, timely responses by providing severity classification, role-based coordination, runbooks, communication templates, and post-mortem processes so incidents are resolved faster and organizational learning is captured.

Core Features & Use Cases

  • Structured Incident Command: Assign IC, comms lead, technical lead, and scribe with clear timeboxed decision steps and escalation triggers.
  • Runbooks & Remediation Playbooks: Templates for detection, diagnosis, rollback, restart, scaling, and verification to reduce MTTR.
  • Post-Mortem Facilitation: Blameless post-mortem templates, 5 Whys, action item tracking, and lessons learned to prevent repeats.
  • SLO/SLI & On-Call Design: SLO definitions, burn rate policies, and on-call rotation designs to guide when to page and when to pause feature work.
  • Use Case: Lead a SEV1 outage for a checkout API: declare severity, coordinate rollback or mitigation, communicate to stakeholders, verify SLIs, and produce a post-mortem with tracked actions.

Quick Start

Ask the agent to act as Incident Response Commander for a SEV2 outage on the checkout-api: assign roles, run diagnostics using runbook steps, propose immediate mitigations, and produce a timestamped timeline plus a post-mortem action list.

Frequently Asked Questions about agency-incident-response-commander

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate incident response for a production outage in distributed services?▼

Incident response coordination assigns IC, comms lead, technical lead, and scribe roles with timeboxed decision steps and escalation triggers to resolve production outages faster. You get structured severity classification, runbook-driven diagnosis, and stakeholder communication templates.

What is a blameless post-mortem and how does it prevent repeat incidents?▼

A blameless post-mortem uses 5 Whys analysis and action item tracking to capture organizational learning without assigning fault. It produces templates and tracked remediation actions that prevent similar distributed service incidents from recurring.

How do I classify incident severity for a checkout API outage?▼

Incident severity classification uses severity matrices to evaluate impact on SLOs and user-facing functionality for distributed services. You apply SEV classification levels to trigger appropriate escalation, stakeholder communications, and runbook-driven mitigation steps.

Can I design on-call rotations and paging policies based on SLO burn rates?▼

On-call rotation designs integrate SLO definitions and burn rate policies to guide paging thresholds and feature work pauses. You get schedule designs and SLO/SLI frameworks that determine when to page engineers during distributed service degradation.

How do I run game-day exercises for incident response preparedness?▼

Game-day exercises simulate production incidents using runbook templates and severity matrices to test distributed service response workflows. You practice role-based coordination, runbook-driven diagnosis, and post-mortem facilitation before real outages occur.

What is the best way to reduce MTTR during a SEV1 incident?▼

Reducing MTTR during a SEV1 incident involves applying runbook templates for detection, diagnosis, rollback, restart, and scaling alongside structured incident command. You verify SLIs after mitigation and produce a timestamped timeline with tracked post-mortem actions.