mx-diagease

Automate GPU hardware diagnostics and health monitoring for AI deployments.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/dongg622/china-ai-chip-skill --skill mx-diagease
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: mx-diagease
Source: https://github.com/dongg622/china-ai-chip-skill/tree/main/MetaX/mx-diagease
Command: npx skills add https://github.com/dongg622/china-ai-chip-skill --skill mx-diagease

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes scripts (resource) components.

What problem does it solve?

The Skill simplifies the process of diagnosing and monitoring AI GPU devices, saving time and reducing errors during hardware troubleshooting.

Core Features & Use Cases

  • Device Diagnostics: Perform comprehensive PCIe, memory, and power diagnostics on GPU hardware.
  • Real-Time Monitoring: Continuously track GPU temperature, power, and link status during operation.
  • Use Case: An FAE needs to quickly identify hardware issues such as PCIe bottlenecks or thermal faults on AI clusters using automated scripts.
  • Automated Testing: Run stress tests and validate device stability under load.

Quick Start

Launch the diagnostic tool with sudo to list connected GPU devices and perform health checks directly from the command line.

Frequently Asked Questions about mx-diagease

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I automate GPU diagnostics and health monitoring for AI hardware?▼

Automate GPU diagnostics by running script-based health checks to detect PCIe, memory, and power issues on AI hardware. This approach minimizes manual effort and provides comprehensive device status validation for maintenance teams.

What is the best way to troubleshoot PCIe bottlenecks and thermal faults on AI clusters?▼

Troubleshoot PCIe bottlenecks and thermal faults on AI clusters by executing automated stress tests and real-time monitoring scripts. This validates device stability under load and tracks temperature, power, and link status continuously.

Can I run stress tests to validate GPU stability under load?▼

Yes, you can run stress tests to validate GPU stability under load. The automated scripts test device stability during operation, helping hardware engineers quickly identify performance degradation and hardware faults.

Do I need sudo permissions to perform GPU health checks?▼

Yes, you need sudo permissions to perform GPU health checks. Launching the diagnostic tool with sudo allows the automated scripts to list connected GPU devices and directly access hardware status for comprehensive checks.

How does real-time GPU monitoring work for AI deployments?▼

Real-time GPU monitoring works by continuously tracking GPU temperature, power, and link status during operation. This mechanism ensures that hardware engineers can detect anomalies and prevent failures in AI deployments.