slime-rl-training

Guide RL-based post-training of LLMs with Megatron-LM and SGLang.

Updated Aug 27, 2026
One-click install
npx skills add https://github.com/helix4u/hermes-agent --skill slime-rl-training-helix4u
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: slime-rl-training
Source: https://github.com/helix4u/hermes-agent/tree/main/skills/mlops/training/slime
Command: npx skills add https://github.com/helix4u/hermes-agent --skill slime-rl-training-helix4u

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

slime provides guidance for RL-based post-training of LLMs using the slime framework, helping researchers and engineers structure experiments and scale RL training workflows.

Core Features & Use Cases

  • Megatron-LM integration with SGLang-based rollout for scalable RL post-training.
  • Data-generation workflow templates and buffer management for RL tasks.
  • End-to-end guidance for configuring experiments, monitoring progress, and debugging RL pipelines.

Quick Start

Launch slime training with your model checkpoint and prompt data to begin RL post-training.

Frequently Asked Questions about slime-rl-training

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I configure RL post-training for LLMs using Megatron-LM and SGLang?▼

RL post-training for LLMs using Megatron-LM and SGLang is configured by launching the slime framework with your model checkpoint and prompt data. This provides end-to-end guidance for structuring scalable training experiments across GPUs.

Do I need a specific environment to run slime RL training workflows?▼

Yes, running slime RL training workflows requires a Megatron-LM environment, the slime framework, and access to model checkpoints and prompt data. This setup is necessary for researchers and engineers performing scalable post-training.

What is the best way to manage data generation and buffers for RL tasks?▼

The best way to manage data generation and buffers for RL tasks is using the slime framework's custom data-generation workflow templates. These templates help structure data pipelines and buffer management for reinforcement learning.

How does SGLang-based rollout work with Megatron-LM for scalable training?▼

SGLang-based rollout integrates with Megatron-LM to enable scalable RL post-training across GPUs. The slime framework provides the necessary templates to configure experiments, monitor progress, and debug these distributed pipelines.

Why use slime for reinforcement learning post-training instead of other tools?▼

Use slime for reinforcement learning post-training to get end-to-end guidance on configuring experiments, monitoring progress, and debugging pipelines. It specifically integrates Megatron-LM and SGLang for scalable custom data generation workflows.

Can I debug RL pipelines while monitoring training progress?▼

Yes, you can debug RL pipelines while monitoring training progress using the slime framework. It provides end-to-end guidance for configuring experiments, tracking metrics, and troubleshooting Megatron-LM and SGLang deployments.