slime-rl-training

Coordinate RL-based post-training for LLMs using slime's Megatron-LM and SGLang integration.

Updated Apr 12, 2026
One-click install
npx skills add https://github.com/thisismynewfmail-ui/Monika-agent --skill slime-rl-training-thisismynewfmail-ui
Or copy as Structured Prompt for Agent▼
Please help me install this Agent Skill.
Skill: slime-rl-training
Source: https://github.com/thisismynewfmail-ui/Monika-agent/tree/main/optional-skills/mlops/slime
Command: npx skills add https://github.com/thisismynewfmail-ui/Monika-agent --skill slime-rl-training-thisismynewfmail-ui

SYSTEM DOCUMENTATION & REQUIREMENTS

💡 This Skill includes references (resource) components.

What problem does it solve?

RL-based post-training for large language models requires coordinating model training, rollout generation, and reward-based optimization. This skill guides users through building scalable RL workflows with slime (Megatron-LM + SGLang) to streamline post-training pipelines.

Core Features & Use Cases

  • End-to-end RL post-training setup for GLM-like models
  • Flexible data generation and rollout orchestration with Megatron-LM and SGLang
  • Tools for monitoring, evaluation, and iterative improvement in production RL tasks

Quick Start

Train a RL-based post-training workflow for your GLM models using slime to coordinate Megatron-LM and SGLang.

Frequently Asked Questions about slime-rl-training

High-intent search queries and answers about installing and using this skill.

FAQPage Schema
How do I scale RL post-training for LLMs using Megatron-LM?▼

Scale RL post-training for LLMs by using slime to coordinate Megatron-LM training and SGLang rollout generation, enabling modular configuration and data buffering for production-scale reward-based optimization.

What is the best way to set up end-to-end RL workflows for GLM models?▼

Set up end-to-end RL workflows for GLM models by configuring slime's Megatron-LM and SGLang integration to orchestrate rollout generation, data buffering, and reward-based optimization.

Can I use slime for custom data generation workflows in reinforcement learning?▼

Yes, slime supports custom data generation workflows in reinforcement learning by providing flexible rollout orchestration and data buffering to coordinate generation and training pipelines.

Does slime RL post-training work with architectures beyond GLM-family models?▼

Slime RL post-training applies to GLM-family models and related architectures, coordinating Megatron-LM scale training and SGLang integration to support compatible reinforcement learning workflows.

What components are needed to configure modular RL training with slime?▼

Configuring modular RL training with slime requires setting up Megatron-LM and SGLang integration, utilizing data buffering and optional reference components to enable end-to-end post-training workflows.