astra-sre

Orchestrate health scans, incident triage, guided repair, and learning loops for multi-node infrastructure.

1|Updated Jun 18, 2026
One-click install
npx skills add https://github.com/alrcatraz/astra-aiagent-infra --skill astra-sre
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: astra-sre
Source: https://github.com/alrcatraz/astra-aiagent-infra/tree/main/skills/devops/astra-sre
Command: npx skills add https://github.com/alrcatraz/astra-aiagent-infra --skill astra-sre

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill requires astra-hub, astra-sre-fix-e2ee, astra-sre-restart-service, astra-sre-fix-gfw, astra-sre-fix-mcp, astra-sre-fix-vps-recovery, server-restart-recovery, server-health-audit, infrastructure-device-inventory, service-inventory, crash-marker-pattern, full-e2ee-recovery-after-server-rebuild, pre-upgrade-server-backup, and includes scripts (resource) and references (resource) components.

What problem does it solve?

The astra-sre skill provides a unified SRE (Site Reliability Engineering) coordination layer for multi-node infrastructure, addressing the challenges of health scanning, incident triage, guided repair, and learning across all managed devices.

Core Features & Use Cases

  • Health Scanning: Conducts comprehensive health scans across all devices and services.
  • Incident Triage: Assesses the severity and impact of incidents, routing known faults to appropriate sub-skills.
  • Guided Repair: Offers a step-by-step repair plan based on diagnosis, with verification probes and rollback mechanisms.
  • Learning Loop: Continuously learns from incidents and suggests creating new sub-skills for recurring issues.
  • Use Case: When an incident occurs, astra-sre can automatically diagnose the problem, propose a repair plan, and guide the user through the process to ensure a reliable infrastructure.

Quick Start

To begin, activate astra-sre to initiate a health scan of your infrastructure.

Frequently Asked Questions about astra-sre

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I coordinate incident triage and health scanning for multi-node infrastructure?▼

Infrastructure health scanning and incident triage are coordinated by assessing incident severity, routing known faults to sub-skills, and conducting comprehensive health checks across all managed devices and services.

What is the best way to automate guided repair and fault diagnostics for reliability engineering?▼

Guided repair for reliability engineering is automated by offering a step-by-step repair plan based on fault diagnostics, complete with verification probes and rollback mechanisms to ensure stable multi-node infrastructure.

How does a site reliability engineering learning loop handle recurring infrastructure incidents?▼

A site reliability engineering learning loop handles recurring incidents by continuously learning from fault diagnostics and suggesting the creation of new sub-skills for known issues to prevent future infrastructure failures.

Do I need pre-upgrade server backups and device inventory data to use SRE incident management?▼

Yes, SRE incident management requires pre-upgrade server backups, device inventory data, and service inventory integration to accurately assess fault severity and route incidents to appropriate repair sub-skills.

Can I use automated SRE fixes for full e2ee recovery after a server rebuild?▼

Yes, automated SRE fixes can handle full e2ee recovery after a server rebuild by orchestrating guided repair plans and utilizing specialized sub-skills like e2ee recovery and server restart mechanisms.

What are the limitations of using a unified SRE coordination layer for multi-node infrastructure management?▼

The limitations of unified SRE coordination include its dependency on multiple sub-skills for specific fixes like gfw or mcp issues, meaning it routes faults rather than independently resolving all multi-node infrastructure anomalies.