Library / SDK
NVIDIA-NeMo/RL avatar
NVIDIA-NeMo/RL

NeMo RL: NVIDIA's Post-Training Library for GRPO, DPO and PPO at Scale

Project brief: Scalable toolkit for efficient model reinforcement. [04/06/2026] New Model Support Added support for Qwen3.5 dense and MoE models (LLM and VLM) for GRPO training.

2,042 stars576 forksPythonApache-2.0

At a glance

What is it?
NeMo RL is a Python post-training library that runs from one GPU to thousands, with recipes for GRPO, DPO, PPO, SFT and on-policy distillation. It is documented around NVIDIA hardware and is not a drop-in replacement for a single-GPU fine-tuning script.
Who is it for?
Adopt NeMo RL if you are already on NVIDIA GPUs, need GRPO or DPO at multi-node scale, and can work from the YAML recipes in examples/configs. Do not adopt it for a single-GPU LoRA experiment on consumer hardware, or if you need CPU-only training: the dependency list pins torch==2.11.0 and pulls in ray, and the recipes assume multi-GPU layouts such as 1n8g.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What NeMo RL actually solves, and who it is for

Post-training a language model with reinforcement learning means running at least two large workloads at once: a policy being optimized and a generation engine producing rollouts. Getting those two to agree on weights, share a cluster, and stay numerically stable is the part that eats engineering weeks. NeMo RL is NVIDIA's answer to that problem, packaged as a Python library with Hydra configs and example entry points.

The scope is stated in pyproject.toml as "Post-Training Library for Models Ranging from 1 GPU to 1000s, and from Tiny to >100B Parameters". The algorithms visible in the repository are GRPO, DPO, PPO, SFT, reward modeling, VLM variants of GRPO, MPO and SFT, and on-policy distillation. There is also a separate example for cross-tokenizer off-policy distillation.

The intended user is not someone fine-tuning a 7B model on a single card. It is a team that has a cluster, a model checkpoint, and a reward signal, and needs a training loop that survives multi-node execution. The README's news entries are a fair signal of that audience: Nemotron-3-Ultra, Nemotron-3-Super and Nemotron-3-Nano were post-trained with this library, and NVIDIA publishes reproducible recipes for each.

The mechanism: policy trainer, generation backend, and weight transfer

The architecture implied by the repository layout separates three concerns. The nemo_rl/ package holds the library. examples/ holds runnable entry points, one per algorithm: run_grpo.py, run_dpo.py, run_ppo.py, run_sft.py, run_rm.py, run_eval.py, run_vlm_grpo.py, run_distillation.py. Configuration lives in examples/configs/recipes/, and the filenames encode the parallelism plan.

That naming convention is the most informative thing in the repository. A recipe called grpo-qwen3.5-9b-1n8g-megatron.yaml describes GRPO on a 9B Qwen3.5 model, one node with eight GPUs, using the Megatron Core backend. A second recipe, grpo-qwen3.5-35ba3b-2n8g-megatron-ep16tp2cp2.yaml, is a 35B-total, 3B-active MoE model across two nodes and eight GPUs, with expert parallelism 16, tensor parallelism 2 and context parallelism 2. You can read the resource plan off the filename before opening the file.

Two training backends appear across the release notes: DTensor and Megatron Core. LoRA support was added to both, for SFT, GRPO and DPO. Generation backends include vLLM and, as of v0.6.0, SGLang. The same release added speculative decoding, the Muon optimizer, YaRN long-context training, chunked cross entropy loss, and top-p/top-k sampling for training.

Weight transfer between the trainer and the generator is treated as a first-class engineering problem. A discussion linked from the README is titled "Journey of Optimizing Weight Transfer in Large MoE Models by 10x", which tells you where the bottlenecks were found.

Installing NeMo RL and running a first GRPO job

The README does not carry a quickstart block; installation instructions live in the documentation at docs.nvidia.com/nemo/rl. Two distribution paths are confirmed in the repository. The first is the NGC container, published since v0.4.0 and available for both linux/amd64 and linux/arm64 as of v0.5.0. The second is a source install from the repository, which uses pyproject.toml with setuptools and a uv.lock file at the top level.

The Python requirement is narrow: requires-python is ">=3.13.14,<3.14". Torch is pinned exactly at 2.11.0. Ray, Weights and Biases, datasets, Hydra, OmegaConf and Triton are all direct dependencies. The container tag follows the release version:

bash
docker pull nvcr.io/nvidia/nemo-rl:v0.5.0

A source install from a clone of the default branch runs through uv, since uv.lock is committed. The repository pins the interpreter in .python-version:

bash
uv sync

Once installed, the entry points are plain Python scripts under examples/. A GRPO run takes a recipe YAML. The README points to grpo-qwen3.5-9b-1n8g-megatron.yaml as the recipe for a 9B Qwen3.5 model on one node of eight GPUs, and the repository ships examples/run_grpo.py as the corresponding entry point. Expect the process to initialize Ray, load the policy, start the generation backend, and begin writing metrics. The dependency list includes wandb and tensorboard, so both logging paths are available; the release notes for v0.5.0 and v0.4.0 also link Google Colab notebooks with release run metrics. The README does not document a CPU-only or single-GPU path for these recipes, and the 1n8g naming suggests eight accelerators is the floor for the published configurations.

Where NeMo RL is the wrong choice

The narrowest constraint is the Python pin. ">=3.13.14,<3.14" means 3.14 is excluded, and so is anything from 3.12 downward. If your cluster image ships Python 3.11 or your internal tooling is stuck on 3.12, you are changing that before you train anything.

The second constraint is hardware. Every recipe in the repository is named with an n8g suffix, and the performance benchmarks referenced in the v0.6.0 notes are labeled GB200 BF16. There is no documented path for a workstation with one or two consumer GPUs. If that is your environment, a library built around Ray and a distributed generation backend adds setup cost without adding capability.

The third is scope. NeMo RL does post-training: SFT, preference optimization, RL, distillation. It is not a pretraining framework and the README does not present it as one. If you need to train a model from scratch, the recipes here start from an existing checkpoint.

Finally, the documentation is not uniform across features. Some capabilities have a dedicated guide in docs/guides, such as the DAPO guide. Others are announced in the news list with a recipe filename and nothing else. The README does not document rollback or checkpoint-compatibility guarantees between releases, which matters if you plan to resume a long run across a version bump.

NeMo-RL vs VeRL and the Megatron Core question

The comparison people search for is NeMo-RL vs VeRL. Both are RL post-training libraries with a trainer and a generation engine, and both target multi-GPU clusters. The difference visible in this repository is the backend story. NeMo RL offers two training backends, DTensor and Megatron Core, and the Megatron Core path is what the large MoE recipes use, with explicit expert, tensor and context parallelism in the config filename. That is a bet on NVIDIA's own parallelism stack rather than on PyTorch-native sharding alone.

The second difference is the recipe library. NeMo RL ships named, versioned YAML files per model and per cluster shape, and NVIDIA publishes the post-training recipes for its own Nemotron releases. If you are training a Nemotron or a recent Qwen or GLM checkpoint, there is likely a starting config. If you are training something unusual, you are writing the parallelism plan yourself.

A third option in the same space is NeMo-Aligner, which also appears in the related searches. This repository does not describe Aligner's internals, so the honest framing is that NeMo RL is the current NVIDIA post-training library and Aligner is a separate, earlier project; check Aligner's own documentation before assuming feature parity in either direction.

Maintenance, releases and the Apache-2.0 licence

The repository is not archived and the last push was on 2026-07-29, which is the same day as the v0.7.0 release. Releases have landed roughly quarterly: v0.5.0 on 2026-01-31, v0.6.0 on 2026-04-30, v0.7.0 on 2026-07-29. Each release notes a substantial feature set rather than a patch list, which means upgrade cost is real. v0.6.0 added an SGLang backend, Muon, speculative decoding, YaRN long-context training and chunked cross entropy loss. v0.7.0 added PPO, MOPD, cross-tokenizer support, router replay and CISPO, plus model support for Qwen3-Omni, Nemotron Nano v3 Omni, Gemma 4 and GLM 5.1.

That cadence is good for capability and awkward for reproducibility. A recipe that runs on v0.6.0 may need edits on v0.7.0 if a config key moved. The README does not publish a deprecation policy or a config migration guide, so pinning your container tag is the practical way to keep a run reproducible. The NGC container tags map to release versions, which makes pinning straightforward.

The licence is Apache-2.0, declared both in the repository LICENSE file and in pyproject.toml as "Apache 2.0". That is a permissive licence with an explicit patent grant, which matters for a library that depends on NVIDIA's own parallelism stack. It does not, by itself, settle the licensing of the model weights you train with it. Those carry their own terms from their publishers, and this repository says nothing about them. This is not legal advice; if you plan to ship a model trained with NeMo RL, check the base model's licence separately.

Editorial conclusion

Adopt NeMo RL if you are already on NVIDIA GPUs, need GRPO or DPO at multi-node scale, and can work from the YAML recipes in examples/configs. Do not adopt it for a single-GPU LoRA experiment on consumer hardware, or if you need CPU-only training: the dependency list pins torch==2.11.0 and pulls in ray, and the recipes assume multi-GPU layouts such as 1n8g. Before committing, check the recipe closest to your model size, confirm which backend it uses (Megatron Core or DTensor), and verify the pinned Python range of >=3.13.14,<3.14 against your cluster image.

Frequently asked questions

What is NeMo RL?

NeMo RL is NVIDIA's open source post-training library for language and vision-language models, covering SFT, GRPO, DPO, PPO, reward modeling and on-policy distillation. pyproject.toml describes it as a post-training library for models ranging from 1 GPU to 1000s and from tiny to over 100B parameters.

Is NVIDIA NeMo free to use?

The NeMo RL repository is licensed under Apache-2.0, declared in both the LICENSE file and pyproject.toml. That covers the library code, not the model weights you train with it, which carry their own licences from their publishers.

What is NVIDIA NeMo used for?

NeMo RL is used for post-training models rather than pretraining them. The repository ships entry points for GRPO, DPO, PPO, SFT, reward modeling, VLM training and on-policy distillation, and NVIDIA states that Nemotron-3-Ultra, Nemotron-3-Super and Nemotron-3-Nano were post-trained with it.

What is NVIDIA NeMo Gym?

The release notes list NeMo-Gym plus NeMo-RL support as a v0.5.0 feature, and an examples/nemo_gym/ directory exists in the repository with reproducible recipes for Nemotron-3.5-lightning. The README does not describe NeMo-Gym's own architecture, so its internals have to be read from its own documentation.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvidia-nemo-rl.svg)](https://hysenlabs.com/projects/nvidia-nemo-rl)