Python Apache-2.0

slime

slime is an LLM post-training framework for RL Scaling.

T

THUDM

Dernière activité 29 sept. 2026
THUDM/slime

8,6 k

étoiles

1,3 k

forks

523

issues ouvertes

Ce README est souvent en anglais.

slime

中文版

Documentation CI Ask DeepWiki

slime is an LLM post-training framework for RL scaling, providing two core capabilities:

  1. High-Performance Training: Supports efficient training in various modes by connecting Megatron with SGLang;
  2. Flexible Data Generation: Enables arbitrary training data generation workflows through custom data generation interfaces and server-based engines.

slime's design goal is to make these two capabilities reinforce each other without turning the system into a heavy stack of disconnected trainers, rollout services, and agent frameworks. Megatron training, SGLang rollout, custom data generation, reward computation, verifier feedback, and environment interaction all flow through the same training / rollout / Data Buffer path.

This makes slime one of the most battle-tested open RL post-training frameworks: small enough to understand and extend, but validated through complete training loops behind SOTA-level model releases.

Why This Design Matters

  • Production experience: slime is the RL framework behind GLM-5.3-Flash, GLM-5.3, GLM-5.2, GLM-5.1, GLM-5, GLM-4.7, GLM-4.6, and GLM-4.5.
  • Direct engine integration: Megatron arguments are available directly; SGLang arguments use the --sglang- prefix.
  • Customizable data generation: generation functions, reward functions, verifiers, and environments connect through documented interfaces.
  • Correctness and reliability: separate rollout-only and train-only debugging, reproducibility controls, recovery, and CPU/GPU tests support long-running experiments.

Supported Models

Alongside the GLM family, slime supports Qwen (Qwen3.6, Qwen3.5, Qwen3-Next, Qwen3 MoE, Qwen3, Qwen2.5), DeepSeek (V3, V3.1, R1), and Llama 3. See the model configurations in scripts/models and the training examples.

Engine Configuration and Deployment

Use Megatron arguments directly for parallelism, optimizers, checkpoints, and model configuration. Prefix installed SGLang arguments with --sglang-; for example, --mem-fraction-static becomes --sglang-mem-fraction-static.

For more involved deployments:

  • SGLang Config: YAML configuration for server groups, multiple models, and per-group overrides.
  • PD Disaggregation: separate prefill and decode resources.
  • Delta Weight Sync: send changed weight bytes over shared storage.
  • External Rollout Engines: connect serving processes managed outside the training job, including disk-based updates across different GPU fleets.

Correctness, Stability, and CI

CPU tests cover core behavior and customization contracts. GPU tests exercise dense and MoE training, rollout deployment, checkpointing, precision, fully asynchronous rollout, distillation, and debug replay. See CI for the test matrix and how to run it.

Engineering guides: Debugging, Reproducibility, Fault Tolerance, Tracing, and Profiling.

Blogs

Table of Contents

Architecture Overview

arch

Module Descriptions:

  • training (Megatron): Responsible for the main training process, reads data from the Data Buffer, and synchronizes parameters to the rollout module after training.
  • rollout (SGLang + router): Generates new data (including rewards/verifier outputs) and stores it in the Data Buffer. Custom generate functions can wrap this with multi-turn loops, tool calls, environment/sandbox interaction, and verifier-based reward.
  • data buffer: A bridge module that manages prompt initialization, custom data, and rollout generation methods (including agentic workflows that produce samples through the same interface).

The default payload transport is Ray object-store. Selecting --rollout-data-transport straw uses straw for persistent prompt tasks, rollout continuations and training batches on shared JuiceFS storage. Ray carries control messages and references; generation and training processes read and write packed payloads directly. See the straw usage and recovery guide.

Quick Start

Follow the Quick Start Guide for environment setup, data preparation, and your first training run.

We also provide examples for some use cases not covered in the quick start guide; please check examples.

Agentic RL examples

These agentic RL examples use the standard rollout and data buffer interfaces:

  • examples/multi_agent: Multi-agent generation via --custom-generate-function-path inside the standard rollout loop.
  • examples/search-r1: Search/RAG-style multi-turn generation via --custom-generate-function-path.
  • examples/fully_async: Fully-async rollout, useful for long-tail agentic generation where some samples take much longer than others.
  • examples/coding_agent_rl: End-to-end SWE coding-agent RL with sandboxed tool use, test-based rewards, and token-correct trajectory segments via --custom-generate-function-path.

See the Customization Guide for which interface to use for a given agentic workflow.

Ecosystem Built on slime

These independent projects build on slime for model post-training, agentic RL, domain applications, and rollout research.

Dressage

Dressage — Alibaba Accio’s agentic RL framework, with Paddock, Sandbox, and Proxy layers for interaction, execution placement, and token-level trajectory capture. It supports black-box agents and configurable sandboxes.

Miles

Miles — RadixArk’s large-model post-training framework, extending slime with SGLang integration, deployment and operations tooling, LoRA, TITO, and low-precision training.

vime

vime — A post-training framework maintained by the vLLM project. It retains slime’s Megatron training and data generation design while using vLLM and vllm-router for rollout.

Relax

Relax — RedAI Infra’s multimodal agentic RL framework, using Ray Serve, TransferQueue, and asynchronous checkpoint synchronization to separate training, rollout, and teacher/reference computation.

OpenClaw-RL

OpenClaw-RL — Personalized OpenClaw training from conversation feedback, using GRPO or on-policy distillation alongside ongoing API serving.

P1

P1 — Physics reasoning models trained with multi-stage RL, adaptive task difficulty, and training stabilization.

RLVE

RLVE — RL across 400 procedurally generated, verifiable environments, with task difficulty adapted to the current policy.

TritonForge

TritonForge — GPU kernel generation using SFT followed by RL with multi-turn compilation feedback.

APRIL

APRIL — Rollout throughput optimization through additional in-flight requests and active management of partial generations.

qqr

qqr — Open-ended agent training with ArenaRL tournament ranking and MCP tool environments.

ART

ART — An SDK for training production agents on AWS Bedrock AgentCore Runtime. It reuses agent workflows, captures trajectories at the model gateway, and offers slime as a training backend.

Arguments Walkthrough

Arguments in slime are divided into three categories:

  1. Megatron arguments: slime reads Megatron arguments directly. You can configure Megatron by passing arguments like --tensor-model-parallel-size 2.
  2. SGLang arguments: All arguments for the installed SGLang are supported through pass-through. These arguments must be prefixed with --sglang-. For example, --mem-fraction-static should be passed as --sglang-mem-fraction-static.
  3. slime-specific arguments: Please refer to: slime/utils/arguments.py

For complete usage instructions, please refer to the Usage Documentation.

Code Reading Path

Start from the training loop and follow the calls only as deep as needed:

train.py: train
├─ slime/ray/placement_group.py       Ray resource and worker initialization
├─ slime/ray/rollout.py              RolloutManager.generate: rollout orchestration
│  └─ slime/rollout/sglang_rollout.py  Sample generation and reward computation
└─ slime/ray/actor_group.py          RayTrainGroup.async_train: training dispatch
   └─ slime/backends/megatron_utils/actor.py
      ├─ model.py                    Megatron model execution
      └─ loss.py                     RL losses and advantages

On a first pass, treat slime/utils/arguments.py as the configuration entry point. The deployment details in slime/backends/sglang_utils/ and the weight-sync implementations under slime/backends/megatron_utils/update_weight/ can also wait until you need to change those areas.

Developer Guide

apt install pre-commit -y
pre-commit install

# run pre-commit to ensure code style consistency
pre-commit run --all-files --show-diff-on-failure --color=always

FAQ & Acknowledgements

  • For frequently asked questions, please see the Q&A
  • Special thanks to the following projects & communities: SGLang, Megatron‑LM, mbridge, OpenRLHF, veRL, Pai-Megatron-Patch and others.
  • To quote slime, please use:
@misc{slime_github,
  author       = {Zilin Zhu and Chengxing Xie and Xin Lv and slime Contributors},
  title        = {slime: An LLM post-training framework for RL Scaling},
  year         = {2025},
  howpublished = {\url{https://github.com/THUDM/slime}},
  note         = {GitHub repository. Corresponding author: Xin Lv},
  urldate      = {2025-06-19}
}

Projets similaires

An LLM post-training framework with vLLM for RL Scaling

Python
Vvllm-project
478 étoiles98

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

Python
Rradixark
3 k étoiles526

A project implementing various agentic RL based on the Slime post-training framework

Python
LLMIS-ORG
539 étoiles39