Open-source project
NVlabs/QeRL avatar
NVlabs/QeRL

QeRL: Quantization-Enhanced Reinforcement Learning for 32B LLMs on a Single GPU

[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.

522 stars52 forksPythonApache-2.0

At a glance

What is it?
QeRL is an NVIDIA research framework that combines NVFP4 quantization with Low-Rank Adaptation (LoRA) to make reinforcement learning training feasible for 32-billion-parameter language models on a single H100 80 GB GPU. It also demonstrates that quantization noise itself improves RL exploration, enabling faster reward growth than standard 16-bit LoRA.
Who is it for?
ML researchers training 32B models with RL on constrained GPU budgets have a clear path with QeRL: clone the repository, set up the quantization environment to produce an NVFP4 model, then train with the provided DAPO scripts. The hardware requirement is firm: NVFP4 is only available on GPUs like the H100, RTX 5090, and B100, and the README states 64 GB of RAM is needed alongside the GPU.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Why RL Training at 32B Scale Is Hard

Reinforcement learning for language models requires generating many full-length rollouts (outputs) from the model at each training step and then using those rollouts to compute a reward signal. For a 32B-parameter model in 16-bit precision, just holding the model weights in GPU memory consumes roughly 64 GB, leaving almost nothing for the KV cache, the optimizer state, or the rollout buffer on an 80 GB H100.

The standard approach is QLoRA: freeze the model weights in 4-bit NF4, and train low-rank adapter weights in higher precision. This cuts the memory footprint significantly, but NF4 quantization during RL adds noise to the policy in a way that QeRL's paper argues is actually suboptimal for exploration.

QeRL proposes a different path. Use NVFP4, NVIDIA's hardware-accelerated 4-bit float format, for both the weight representation and the forward pass. The NVFP4 inference is faster than 16-bit inference on supported hardware because the H100's tensor cores handle NVFP4 natively. The rollout phase, which dominates wall-clock time in RL training, therefore runs faster than in any 16-bit approach.

How NVFP4 Quantization Reduces Memory and Speeds Rollout

NVFP4 represents each weight in 4 bits using NVIDIA's FP4 format, which is hardware-supported on H100, RTX 5090, and B100 GPUs with triton. At NVFP4, a 32B model's weights occupy roughly 16 GB rather than 64 GB, leaving substantial headroom for the LoRA adapter parameters and the rollout buffer.

The README and paper claim over 1.5x speedup in the rollout phase compared to 16-bit training. The speedup is larger for longer reasoning traces because rollout time scales with sequence length; a longer trace multiplies the benefit of faster per-token inference.

The memory reduction also makes it practical to run a 32B model on a single H100 80 GB GPU. The README states this is the first framework to enable RL training of a 32B LLM on a single H100 80 GB GPU while delivering overall training speedups. The paper reports benchmarks on Qwen2.5-7B-Instruct and Qwen2.5-32B-Instruct.

Adaptive Quantization Noise and Its Effect on RL Exploration

QeRL's second finding is that quantization noise is not merely a trade-off to accept but an advantage to control. In reinforcement learning, exploration refers to the policy's ability to try novel outputs rather than always picking the highest-probability token. High initial entropy means more exploration; low entropy means the model commits to familiar patterns early.

NVFP4 quantization introduces stochastic noise into each forward pass that increases the model's initial policy entropy, compared to a clean 16-bit model. The README states that quantization noise increases policy entropy, enhancing exploration, and enabling the discovery of better strategies during RL.

To control this effect, QeRL introduces the Adaptive Quantization Noise (AQN) mechanism. AQN adjusts the noise level dynamically during training: more noise early in training (to promote exploration) and less noise later (to consolidate the learned policy). The combination of NVFP4 memory savings, faster rollout, and AQN-guided exploration is what the paper reports as achieving faster reward growth and higher final accuracy than 16-bit LoRA and QLoRA.

Setting Up Two Conda Environments

QeRL requires two separate conda environments because the quantization tool uses Python 3.12 while the training code uses Python 3.10. Start with the main training environment:

bash
git clone https://github.com/NVlabs/QeRL
cd QeRL
conda create -n qerl python=3.10 -y
conda activate qerl
conda install nvidia/label/cuda-12.4.1::cuda
conda install -c nvidia/label/cuda-12.4.1 cudatoolkit
sh setup_env.sh

Then set up the quantization environment in the llm-compressor subdirectory:

bash
cd llm-compressor
conda create -n llmcompressor python=3.12 -y
conda activate llmcompressor

pip install -e .
pip install nvidia-ml-py

The quantization script converts a Hugging Face model to NVFP4 format. For the four Qwen2.5 sizes used in the paper:

bash
python quantize_nvfp4.py --model Qwen/Qwen2.5-3B-Instruct
python quantize_nvfp4.py --model Qwen/Qwen2.5-7B-Instruct
python quantize_nvfp4.py --model Qwen/Qwen2.5-14B-Instruct
python quantize_nvfp4.py --model Qwen/Qwen2.5-32B-Instruct

The quantized model is then used in the training step.

Running RL Training with QeRL

After quantization, switch back to the qerl environment and run training. For QeRL on a single GPU with DAPO:

bash
bash training/dapo_qwen2.5-7b_nvfp4_single_gpu.sh

For a baseline comparison using vanilla 16-bit LoRA:

bash
bash training/dapo_qwen2.5-7b_bf16_single_gpu.sh

The README gives two tuning hints. First, set --vllm-gpu-memory-utilization to a lower value than the default because the NVFP4 model uses less memory; leaving the default may over-allocate for the rollout buffer. Second, tune --perdevice_train_batch_size and --gradient_accumulation_steps to match available memory and the desired effective batch size.

The README truncates the full text for the second hint, but the pattern is consistent with standard GPU-limited training: reduce batch size if memory is tight, use gradient accumulation to keep the effective batch size constant.

Hardware Requirements and Production Constraints

The README states the tested setup: an Nvidia GPU supporting NVFP4 weight format with triton, such as the RTX 5090, H100, or B100; Linux operating system; and 64 GB RAM. Other hardware setups could also work but have not been tested according to the README.

NVFP4 support is hardware-gated. The format is available on NVIDIA Hopper (H100) and newer architectures. Teams using V100, A100, or older cards cannot run QeRL without hardware changes.

The framework depends on a specific CUDA version (12.4.1 as installed in the conda commands) and a specific version of vllm (0.8.5.post1 in the Makefile). These pinned versions reflect the hardware and software environment at the time of the paper's publication.

The project was last pushed on 2026-03-30. There are no GitHub releases. Teams integrating QeRL into a production pipeline should pin to a specific commit hash.

How QeRL Compares to QLoRA

QLoRA (Quantized LoRA) is a widely used parameter-efficient fine-tuning method that quantizes model weights to NF4 (4-bit NormalFloat) for memory savings while keeping LoRA adapter training in higher precision. QLoRA was designed primarily for supervised fine-tuning, not reinforcement learning.

QeRL uses NVFP4 instead of NF4. NVFP4 is a hardware-accelerated format specific to NVIDIA Hopper GPUs that enables faster inference during rollout, not just memory savings. QeRL's AQN mechanism is also absent from QLoRA: QLoRA treats quantization noise as an undesirable artifact to minimize, while QeRL treats it as an adjustable parameter for controlling exploration.

The paper reports that QeRL achieves faster reward growth and higher final accuracy than QLoRA on mathematical benchmarks. On GSM8K, QeRL achieves 90.8% with the 7B model. On MATH 500, it achieves 77.4% with the 7B model. These results are specific to the experimental setup described in the paper and the Qwen2.5 model family.

Editorial conclusion

ML researchers training 32B models with RL on constrained GPU budgets have a clear path with QeRL: clone the repository, set up the quantization environment to produce an NVFP4 model, then train with the provided DAPO scripts. The hardware requirement is firm: NVFP4 is only available on GPUs like the H100, RTX 5090, and B100, and the README states 64 GB of RAM is needed alongside the GPU. Teams using older GPU generations that do not support NVFP4 cannot use QeRL without hardware changes. The last push was on 2026-03-30. The Apache-2.0 license permits commercial use and modification, and the setup.py copyright header explicitly credits the HuggingFace Team as the origin of the build tooling.

Frequently asked questions

What GPU hardware does QeRL require?

The README specifies an Nvidia GPU that supports NVFP4 weight format with triton. The tested examples are the RTX 5090, H100, and B100. The H100 80 GB is the reference GPU for the 32B model experiments. Older architectures like V100 and A100 do not support NVFP4.

Why does QeRL use two conda environments?

The quantization tool in the llm-compressor subdirectory requires Python 3.12, while the main RL training code runs on Python 3.10. The README gives separate conda create commands for each environment: qerl on Python 3.10 for training, and llmcompressor on Python 3.12 for converting models to NVFP4.

What is the Adaptive Quantization Noise mechanism in QeRL?

AQN dynamically adjusts the quantization noise level during RL training. Higher noise early in training increases policy entropy and promotes exploration; lower noise later helps the policy converge. The README states that quantization noise increases policy entropy and enables the discovery of better strategies during RL.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. NVlabs/QeRL on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvlabs-qerl.svg)](https://hysenlabs.com/projects/nvlabs-qerl)