slime-rl-training

Implement RL post-training for GLM models with slime, Megatron-LM, and SGLang.

Updated Apr 26, 2026
One-click install
npx skills add https://github.com/dawsonblock/HERMY --skill slime-rl-training-dawsonblock
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: slime-rl-training
Source: https://github.com/dawsonblock/HERMY/tree/main/hermes-agent-2026.4.23/optional-skills/mlops/slime
Command: npx skills add https://github.com/dawsonblock/HERMY --skill slime-rl-training-dawsonblock

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

Provides guided, structured methodologies for reinforcement-learning-based post-training of large language models using slime, enabling researchers to implement end-to-end RL pipelines with Megatron-LM and SGLang.

Core Features & Use Cases

  • Framework integration: Seamlessly connects slime with Megatron-LM and SGLang for RL post-training workflows.
  • Custom data generation and rollout management: Supports data generation, rollout strategies, and evaluation hooks for RL experiments.
  • Use Case: Researchers can prototype RL-based post-training for GLM-family models with custom reward functions and multi-turn interactions.

Quick Start

Install slime and run the included quick-start example to begin RL post-training with Megatron-LM and SGLang.

Frequently Asked Questions about slime-rl-training

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I set up RL post-training for LLMs using slime and Megatron-LM?▼

To set up RL post-training for LLMs using slime and Megatron-LM, install slime and run the included quick-start example to begin configuring distributed training and rollout optimization workflows.

Can I use custom reward functions for on-policy and off-policy training with slime?▼

Yes, slime supports custom reward functions and multi-turn interactions for both on-policy and off-policy training, allowing researchers to implement tailored rollout-based optimization for GLM-family models.

Does slime integrate with SGLang for rollout data handling?▼

Slime integrates with SGLang to manage rollout data handling and routing, providing structured data generation workflows and evaluation hooks for RL post-training experiments.

What is the best way to configure Megatron-LM for GLM-family model training?▼

The best way to configure Megatron-LM for GLM-family model training is using slime's guided methodologies, which provide structured configuration for rollout-based optimization and custom data generation.

Are there limitations when using slime for RL post-training in research environments?▼

Slime is designed for research environments and specifically targets GLM-family models, meaning its structured RL post-training methodologies are optimized for experimental rollout-based optimization rather than production-scale deployment.