sre-engineer

Identify reliability risks and propose minimal verifiable fixes with rollback plans.

22|2|Updated Mar 24, 2026
One-click install
npx skills add https://github.com/jshsakura/awesome-opencode-skills --skill sre-engineer-jshsakura
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: sre-engineer
Source: https://github.com/jshsakura/awesome-opencode-skills/tree/main/skills/sre-engineer
Command: npx skills add https://github.com/jshsakura/awesome-opencode-skills --skill sre-engineer-jshsakura

SYSTEM DOCUMENTATION & REQUIREMENTS

What problem does it solve?

Own site reliability engineering work as production-safety and operability engineering, not checklist completion.

Core Features & Use Cases

  • Focus on establishing robust SLO-aligned reliability improvements, incident runbooks, and safe rollback strategies.
  • Provide measurable telemetry and guardrails to prevent uncontrolled blast radius.
  • Coordinate across control plane, data plane, and dependencies for end-to-end resilience.

Quick Start

Propose the smallest safe reliability improvement with rollback options for the current incident.

Frequently Asked Questions about sre-engineer

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I propose the smallest safe reliability improvement during an incident?▼

To propose safe reliability improvements, identify the current risk and limit analysis to control plane, data plane, and dependency edges, providing evidence-based rationale with rollback options validated against production telemetry.

What is the best way to establish SLO-aligned incident runbooks?▼

Establishing SLO-aligned incident runbooks involves coordinating across control plane, data plane, and dependencies to ensure end-to-end resilience, while providing measurable telemetry and guardrails to prevent uncontrolled blast radius.

How does risk analysis work for site reliability engineering?▼

Risk analysis for site reliability engineering works by identifying current reliability risks and proposing the smallest, verifiable fix that preserves security boundaries, ensuring recommendations reference measurable indicators.

Can I validate SLO recommendations against production telemetry?▼

Yes, you can validate SLO recommendations against production telemetry, ensuring that proposed reliability improvements include rollback or degrade-path plans and evidence-based rationale for safe deployment.

When do I need rollback strategies for dependency edges?▼

You need rollback strategies for dependency edges when coordinating end-to-end resilience, ensuring that any proposed safe changes include verifiable rollback options to prevent uncontrolled blast radius.

Why focus on small changes for incident management rather than complete overhauls?▼

Focusing on small changes for incident management ensures verifiable fixes that preserve security boundaries, limiting blast radius and providing measurable telemetry guardrails instead of risking uncontrolled disruptions during reliability improvements.