sre-engineer

Defines and tracks SLOs, configures monitoring, and automates operational tasks.

Updated May 31, 2026
One-click install
npx skills add https://github.com/fanguyun/SkillManager --skill sre-engineer-fanguyun
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/fanguyun/SkillManager/tree/main/sre-engineer
Command: npx skills add https://github.com/fanguyun/SkillManager --skill sre-engineer-fanguyun

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires prometheus, kubernetes, python, go, terraform, and includes scripts (resource) and references (resource) components.

What problem does it solve?

This Skill unit provides a comprehensive set of tools and best practices for Site Reliability Engineers (SREs) to manage production systems more effectively, reduce toil, and maintain high service reliability.

Core Features & Use Cases

  • Service Level Objective (SLO) Management: Define and track SLOs for availability, latency, and error budgets.
  • Error Budget Policies: Calculate and manage error budgets to plan for and respond to system failures.
  • Monitoring and Alerting: Implement golden signals monitoring and configure alerting based on SLOs.
  • Automation: Automate repetitive tasks and toil reduction with scripts and tools.
  • Chaos Engineering: Design and execute chaos experiments to test system resilience and recovery.
  • Incident Response: Develop incident response procedures, runbooks, and postmortems for improved MTTR.

Quick Start

Load the sre-engineer skill to start defining SLOs, creating error budget policies, and automating incident response procedures.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
Can I configure Prometheus monitoring and alerting based on SLOs?▼

Yes, you can configure Prometheus monitoring and alerting based on SLOs. It implements golden signals monitoring to track system health and triggers alerts when your service levels are breached.

How do I design and execute chaos engineering experiments to test system resilience?▼

The best way to reduce operational toil is by automating repetitive tasks with Python and Go scripts. This approach minimizes manual intervention and streamlines incident response procedures.

Do I need to know Terraform and Kubernetes before using this SRE approach?▼

You design and execute chaos engineering experiments to actively test system resilience and recovery. This validates your infrastructure's ability to withstand failures under controlled conditions.

How do I create incident response runbooks to improve MTTR?▼

Yes, you need familiarity with Kubernetes, Terraform, Prometheus, and SRE best practices. This foundational knowledge is required to effectively manage production system reliability and define infrastructure policies.

How do I create incident response runbooks to improve MTTR?▼

You develop incident response procedures, runbooks, and postmortems to improve MTTR. This structured documentation guides responders through standardized recovery steps during production system failures.