sre-engineer

Define SLOs, error budgets, and incident response procedures for production systems.

Updated Jan 9, 2026
One-click install
npx skills add https://github.com/dieu-donnee/luxtrax --skill sre-engineer-dieu-donnee
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/dieu-donnee/luxtrax/tree/main/.agent/skills/sre-engineer
Command: npx skills add https://github.com/dieu-donnee/luxtrax --skill sre-engineer-dieu-donnee

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

The SRE Engineer skill guides teams in designing and operationalizing reliability practices, including SLO/SLI definitions, error budget policies, incident response playbooks, capacity planning, and monitoring configurations.

Core Features & Use Cases

  • Define and align SLOs/SLIs to user impact across services.
  • Create and manage error budget policies, burn rate alerts, and runbooks.
  • Design incident response procedures, postmortems, capacity planning models, and monitoring setups.

Quick Start

Provide a concise SLO plan and initial monitoring outline for your production service.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I define and align SLOs and SLIs to measure user impact across services?▼

To define SLOs and SLIs, you establish quantifiable reliability targets mapped directly to user impact across distributed services. This Skill generates SLO plans and initial monitoring outlines, translating service-level indicators into actionable error budget policies.

What is an error budget policy and how does it guide release decisions?▼

An error budget policy dictates the permissible failure rate within an SLO framework. This Skill designs burn-rate alerts and error budget management rules to balance feature velocity against reliability requirements for production systems.

How do I set up incident response procedures and blameless postmortems for distributed services?▼

Setting up incident response procedures involves creating runbooks and postmortem templates for production incidents. This Skill designs blameless postmortem workflows and incident management playbooks tailored for enterprise-grade software teams.

Can I use this for capacity planning and monitoring configurations in enterprise environments?▼

Yes, this Skill targets enterprise-grade software teams by generating capacity planning models and monitoring configurations. It enforces end-to-end reliability requirements including automation templates and burn-rate alerts for distributed services.

What is the best way to start operationalizing reliability practices for a new production service?▼

The best way to start operationalizing reliability practices is to provide a concise SLO plan and initial monitoring outline. This Skill guides teams in building incident response playbooks and aligning SLI measurements to user impact from day one.

SRE Engineer: does it support generating runbooks and automation templates for incident management?▼

Yes, the SRE Engineer Skill supports generating runbooks and automation templates for incident management. It enforces end-to-end reliability requirements by designing incident response procedures and blameless postmortems for distributed services.