holmesgpt-skill

Investigate Kubernetes and cloud-native infrastructure issues using live observability data.

7|1|Updated Jul 12, 2026
One-click install
npx skills add https://github.com/julianobarbosa/claude-code-skills --skill holmesgpt-skill
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: holmesgpt-skill
Source: https://github.com/julianobarbosa/claude-code-skills/tree/main/skills/holmesgpt-skill
Command: npx skills add https://github.com/julianobarbosa/claude-code-skills --skill holmesgpt-skill

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

HolmesGPT provides AI-powered troubleshooting across cloud-native ecosystems, integrating with observability data to perform root-cause analysis while operating with read-only RBAC.

Core Features & Use Cases

  • Root Cause Analysis: Investigates alerts across Kubernetes, Prometheus, and incident systems.
  • Multi-Source Integrations: 30+ toolsets for Kubernetes, Grafana, Loki, Tempo, etc.
  • Alert Integration: Integrates with AlertManager, PagerDuty, OpsGenie, Jira, Slack.
  • Use Case: Automatically investigate a paged incident and summarize root causes.

Quick Start

Install via Helm and connect providers; start an investigation with holmesgpt.

Frequently Asked Questions about holmesgpt-skill

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I investigate Kubernetes alerts automatically with AI?▼

AI-powered troubleshooting investigates Kubernetes alerts by analyzing live observability data from Prometheus, AlertManager, and PagerDuty to identify root causes and suggest remediation steps without manual log analysis.

Can I use AI troubleshooting with Prometheus and Grafana together?▼

Yes, multi-source integrations support simultaneous connections to Kubernetes, Prometheus, Grafana, Loki, Tempo, and DataDog, enabling root-cause analysis across your entire observability stack in a single investigation.

Does cloud-native troubleshooting work with read-only RBAC permissions?▼

Root-cause analysis operates under read-only RBAC constraints, allowing secure investigations without elevated cluster access or write permissions to your infrastructure.

How do I set up AI incident investigation with PagerDuty and Slack?▼

Connect via Helm or CLI installation, configure your AI provider (Anthropic, OpenAI, Azure, AWS Bedrock, Google Gemini, or Vertex AI), and integrate alert systems; investigations automatically trigger from PagerDuty and post summaries to Slack.

What observability platforms does cloud-native root-cause analysis support?▼

30+ toolsets integrate with Kubernetes, Prometheus, Grafana, Loki, Tempo, DataDog, AlertManager, OpsGenie, and Jira, supporting alert-driven investigations across heterogeneous cloud-native environments.

Can I customize troubleshooting runbooks and toolsets for my infrastructure?▼

YAML configuration allows custom toolset definitions and runbook templates, enabling tailored investigation workflows that align with your team's procedures and incident response patterns.