slime-rl-training

Automate end-to-end RL post-training workflows for GLM models with Megatron-LM and SGLang.

Updated Apr 3, 2026
One-click install
npx skills add https://github.com/handsomelong922/my-codex-skills --skill slime-rl-training-handsomelong922
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: slime-rl-training
Source: https://github.com/handsomelong922/my-codex-skills/tree/main/skills/slime
Command: npx skills add https://github.com/handsomelong922/my-codex-skills --skill slime-rl-training-handsomelong922

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

slime-rl-training provides a complete guide for post-training LLMs with RL, integrating Megatron-LM and SGLang to scale data generation, rollout, and evaluation.

Core Features & Use Cases

  • Megatron-LM based training with SGLang rollout
  • High-throughput data generation and rollout orchestration
  • Broad model support (GLM, Qwen3, Llama3, DeepSeek)
  • End-to-end RL post-training workflows with evaluation and monitoring

Quick Start

Install slime-rl-training, configure your Megatron-LM + SGLang environment, and run the standard GRPO workflow to start RL-based post-training for your LLM.

Frequently Asked Questions about slime-rl-training

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I scale reinforcement learning post-training for large language models using Megatron-LM and SGLang?▼

Reinforcement learning post-training for large language models is scaled by integrating Megatron-LM for training and SGLang for high-throughput rollouts. This workflow automates data generation, rollout orchestration, and evaluation for models like GLM, Qwen3, and Llama3.

What prerequisites do I need to run GRPO workflows for LLM post-training?▼

GRPO workflows for LLM post-training require a configured Megatron-LM and SGLang environment. You also need access to training data and model configuration to orchestrate the training, rollout, and evaluation steps effectively.

How does high-throughput data generation work during LLM rollout orchestration?▼

High-throughput data generation during LLM rollout orchestration works by coordinating SGLang with Megatron-LM. This integration automates custom data generation and rollouts, enabling efficient end-to-end RL post-training workflows in research and production settings.

Does this RL training workflow support models outside the GLM family?▼

Yes, this RL training workflow supports models outside the GLM family. It provides broad model support including GLM, Qwen3, Llama3, and DeepSeek, applying Megatron-LM based training and SGLang rollout across these architectures.

What is the best way to automate end-to-end RL post-training workflows with evaluation?▼

The best way to automate end-to-end RL post-training workflows is using slime to orchestrate training, rollout, and evaluation. This applies custom data generation and Megatron-LM plus SGLang integration for large language models.

Can I use SGLang for rollout orchestration in production RL training settings?▼

Yes, you can use SGLang for rollout orchestration in production RL training settings. It integrates with Megatron-LM to support high-throughput data generation and end-to-end evaluation for large language model post-training.