# THUDM/slime: an SGLang-native RL post-training framework for Megatron users

> slime connects Megatron training with SGLang rollout so that RL data generation, reward computation and weight synchronization run through one path. It is a narrow, deliberate bet on one inference backend, and that is the point.

**THUDM/slime** — slime is an LLM post-training framework for RL Scaling.

- Repository: https://github.com/THUDM/slime
- Website: https://thudm.github.io/slime
- Stars: 8,571 · Forks: 1,284
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/thudm-slime

## The gap slime is built to close

If you have tried to do reinforcement learning on a large language model at any real scale, you have probably assembled the same three pieces by hand: a training engine, a rollout engine, and a growing pile of glue that moves weights and samples between them. The glue is where the bugs live. A weight synchronization that is off by one step, a buffer that drops the last partial batch, a reward function that sees a slightly different prompt than the policy did. None of these crash. They just make your reward curve look plausible and wrong.

slime is an attempt to make that glue a first-class part of the framework rather than something each team writes for itself. The README describes it as an LLM post-training framework for RL scaling with two capabilities: high-performance training by connecting Megatron with SGLang, and flexible data generation through custom data generation interfaces and server-based engines. The audience is specific. This is for teams that already have Megatron in their stack and are willing to run SGLang as their only rollout backend. It is not a general-purpose RL library, and it does not pretend to be one.

## Megatron arguments pass through, SGLang arguments keep their names

The architectural decision that shapes everything else is pass-through. slime does not wrap Megatron in an abstraction; per the README, it reads Megatron arguments directly, so Megatron-side parallelism, optimizer, checkpointing and model options stay available without wrapper code. On the serving side, every argument supported by the installed SGLang can be used by adding a `--sglang-` prefix. The README gives the concrete example of passing `--mem-fraction-static` as `--sglang-mem-fraction-static`.

That prefix convention is the whole integration story in miniature. It means slime does not need to track SGLang's argument surface. When SGLang adds a flag, you get it the day you upgrade SGLang, not the day slime ships a wrapper for it. The cost is that slime cannot validate those arguments for you, and a typo in a prefixed flag is your problem, not the framework's.

The dataflow itself is described as one path: Megatron training, SGLang rollout, custom data generation, reward computation, verifier feedback and environment interaction all flow through the same training, rollout and Data Buffer path. For multi-turn and agentic workloads there is PD disaggregation, documented at docs/en/advanced/pd-disaggregation.md, which separates prefill and decode resource needs. For training and inference disaggregation there is delta weight sync, documented at docs/en/advanced/delta-weight-sync.md. The repository also includes an external rollout engine path, docs/en/advanced/external-rollout-engines.md, where the SGLang serving side can run in an independent environment, and with disk transport can even run on different GPU models or vendors, using full-checkpoint update from disk or delta update over a shared filesystem.

## Installing slime and running a first training job

The repository ships a setup.py that declares `python_requires=">=3.10"` and pulls its dependencies from requirements.txt, which is where the real weight sits: ray[default], sglang-router>=0.3.0, transformers, datasets, wandb, and a set of extras including e2b, mcp[cli] and openai-agents for agentic data generation. There is also a build_conda.sh at the top level and a docker/ directory, which is the path most people will take given the GPU and CUDA requirements.

A source install looks like this. Run it from the repository root, and expect the install to pull the full requirements list rather than a trimmed runtime subset.

```bash
pip install -r requirements.txt
pip install -e .
```

The package version in setup.py is 0.3.2, matching the most recent release listed for the project. Python 3.10, 3.11 and 3.12 are the declared classifiers, and the environment classifier is NVIDIA CUDA, so there is no CPU-only story here.

For a first real run, the examples/ directory is the entry point. It contains self-contained setups such as examples/geo3k_vlm/, examples/search-r1/, examples/retool/, examples/tau-bench/ and examples/fully_async/. The top-level train.py and train_async.py are the two training entry scripts. The pattern across the examples is a launch script that invokes one of those entry points with a model path, a data path and the Megatron and `--sglang-` arguments for your topology. Read examples/README.md before adapting anything, because the examples encode the expected argument shape and the repository does not document a generic default configuration outside them.

If you want to debug without training, the README mentions separate rollout-only and train-only debugging paths, with details in docs/en/developer_guide/debug.md. That split is worth knowing about before your first multi-node run.

## One rollout backend is a decision, not an oversight

slime chooses SGLang and only SGLang. The README is explicit about why: multi-backend frameworks often have to abstract over the common subset of several inference engines, which can hide the strongest features of each backend. By committing to one, slime can use SGLang-specific serving, routing, caching, disaggregation and weight-sync behavior directly.

That reasoning is sound and the consequence is real. If your organization has standardized on vLLM for serving, slime is the wrong tool, and no amount of configuration will fix it. You would be adopting a second inference stack for training and maintaining two sets of operational knowledge. The external rollout engines path softens this somewhat, since the serving side can run in an independent environment and, with disk transport, on different GPU models or vendors. But the engine is still SGLang.

The second limitation is scale. This is a framework for large-scale RL on Megatron models. The production validation section lists Qwen3.6 through Qwen2.5, DeepSeek V3 and V3.1 and R1, and Llama 3, alongside the GLM family. If you are fine-tuning a 1B model on a single node, the Megatron plus SGLang plus Ray stack is more machinery than the problem needs, and a simpler single-process RL setup will get you to an answer faster. slime's own framing supports this: it says it focuses deeply on the Megatron plus SGLang path used for large-scale RL.

## How slime differs from a general RL training library

The obvious comparison point is a general-purpose RL post-training library that treats the inference backend as a pluggable component. The difference is architectural rather than feature-level. A pluggable design gives you backend choice and a stable internal API, at the cost of an abstraction layer that has to be maintained against every backend's changes. slime removes the abstraction layer and the backend choice together.

That trade shows up in upgrade behavior. With a pluggable library, a new SGLang optimization reaches you when the library adds support for it. With slime, the README's claim is that upstream engine improvements remain accessible as the engines evolve, because the arguments pass through. The flip side is that slime inherits breakage from upstream too. A Megatron or SGLang change that alters argument semantics lands in your training run without an intermediary to absorb it.

The second difference is where the project puts its engineering effort. The README lists CI, debugging, reproducibility, fault tolerance, trace viewing and profiling as first-class concerns, with docs under docs/en/developer_guide/ and docs/en/advanced/. The stated reason is that RL bugs are often silent. Whether the documentation delivers on that is something you should check against your own failure modes, but the intent is clearly to compete on correctness rather than on the number of supported algorithms.

## Maintenance, releases and the Apache-2.0 terms

The repository is not archived and the last push was on 2026-09-03, with v0.3.2 released on 2026-08-28, v0.3.1 on 2026-08-06 and v0.3.0 on 2026-05-31. That is a steady cadence across the last several months, and the version in setup.py tracks the latest release.

The practical upgrade cost is not in slime itself. It is in the two engines underneath it. Because arguments pass through, a Megatron or SGLang version bump is effectively a slime upgrade for your configuration, and the failure mode is an argument that no longer means what it did. Pin both engines, and test a version bump with the rollout-only and train-only debug paths before you run it on a full training job. The repository's CI documentation at docs/en/developer_guide/ci.md describes what the project itself tests across dense and MoE models, checkpointing, numerical precision, async rollout and PPO-style workflows; that tells you what is covered upstream, not what is covered in your fork.

On licensing, slime is Apache-2.0, which is permissive and includes an explicit patent grant. The dependencies are the part to look at: requirements.txt pulls in ray[default], sglang-router, transformers, datasets and wandb, each with its own licence, and the agentic extras bring in e2b, mcp[cli] and openai-agents. Those licences are not all the same shape as Apache-2.0, and if you redistribute a container image, you are redistributing all of them. That is a question for your own legal review, not something the repository answers.

## Conclusion

Adopt slime if you already run Megatron and SGLang and your problem is the RL loop between them, not the engines themselves. Skip it if you need a multi-backend rollout layer, or if your serving stack is vLLM-based and you have no intention of moving. Before committing, verify the SGLang version your cluster runs against the pass-through arguments you plan to use, confirm your GPU topology supports the PD disaggregation or delta weight sync path you want, and read docs/en/advanced/reproducibility.md, because RL bugs in this class of system are silent.

## FAQ

### What is THUDM/slime exactly?

It is an LLM post-training framework for RL scaling that connects Megatron training with SGLang rollout. Its README describes two core capabilities: high-performance training across modes, and flexible data generation through custom interfaces and server-based engines.

### How to use THUDM/slime for a first training run?

The repository provides train.py and train_async.py as the training entry scripts, with self-contained setups under examples/ such as examples/geo3k_vlm/ and examples/search-r1/. Read examples/README.md first, since the examples carry the expected argument shape.

### How do I install THUDM/slime?

Install from source with pip install -r requirements.txt followed by pip install -e . from the repository root. The package requires Python 3.10 or newer and declares an NVIDIA CUDA environment classifier, so a GPU machine is required.

### Does THUDM/slime support inference backends other than SGLang?

No. The README states that slime chooses SGLang as its single rollout backend deliberately, so that it can use SGLang-specific serving, routing, caching, disaggregation and weight-sync behavior directly rather than abstracting over several engines.

### How do I pass an SGLang argument through THUDM/slime?

Add the --sglang- prefix to the argument. The README gives --mem-fraction-static as the example, which becomes --sglang-mem-fraction-static. Megatron arguments are read directly without a prefix.

## Sources

- [License: Apache-2.0](https://github.com/THUDM/slime/blob/main/LICENSE)
- [Project website](https://thudm.github.io/slime)
- [README](https://github.com/THUDM/slime/blob/main/README.md)
- [Releases](https://github.com/THUDM/slime/releases)
- [THUDM/slime on GitHub](https://github.com/THUDM/slime)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/thudm-slime
