Open-source project
huggingface/open-r1 avatar
huggingface/open-r1

Open R1: Hugging Face's Open Reproduction of the DeepSeek-R1 Training Pipeline

Fully open reproduction of DeepSeek-R1

26,477 stars2,452 forksPythonApache-2.0

At a glance

What is it?
Open R1 is an Apache-2.0 Python toolkit that reproduces the DeepSeek-R1 reasoning model pipeline, from distillation to reinforcement learning. It provides SFT and GRPO training scripts, published datasets including Mixture-of-Thoughts, and requires CUDA 12.4 on multi-GPU hardware.
Who is it for?
Open R1 is the right starting point for ML researchers and engineers who want to reproduce or extend the DeepSeek-R1 reasoning pipeline without access to proprietary training infrastructure. Anyone planning to run training should verify their hardware matches the documented configuration of 8xH100 GPUs at 80GB each, since the README states the scripts are configured for that specific setup and require batch size or gradient accumulation tuning for different hardware.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Open R1 Reproduces and Who It Is For

Open R1 is a community effort to rebuild the training pipeline behind DeepSeek-R1, a reasoning-focused language model. The README states the goal plainly: "build the missing pieces of the R1 pipeline such that everybody can reproduce and build on top of it." The target audience is ML researchers and engineers who want to train or fine-tune reasoning-capable language models using the methods described in the DeepSeek-R1 technical report, without relying on proprietary infrastructure.

The repository is explicitly declared a work in progress. Step 1, which involves reproducing the R1 distillation models by training on high-quality reasoning traces, was completed in May 2025 with the release of the OpenR1-Distill-7B model and the Mixture-of-Thoughts dataset. Steps 2 and 3, which involve replicating the pure reinforcement learning pipeline and demonstrating multi-stage training from a base model, are not yet finished.

The project is licensed under Apache-2.0 and hosted by Hugging Face. Unlike a finished product release, the repository is structured around iterative milestones documented in the news section of the README, meaning the scope of what is ready to run changes as the team completes each stage.

The Three-Stage Pipeline Architecture

The repository follows the DeepSeek-R1 technical report as its guide. The three stages are: first, replicate R1-Distill models by distilling a high-quality reasoning corpus from DeepSeek-R1; second, replicate the pure RL pipeline that was used to create R1-Zero; third, demonstrate going from a base model to an RL-tuned model via multi-stage training.

The codebase for stage 1 consists of three main scripts under src/open_r1/: grpo.py trains a model with Group Relative Policy Optimization on a given dataset, sft.py performs supervised fine-tuning, and generate.py generates synthetic data from a model using the Distilabel library. A Makefile provides easy-to-run commands for each step.

The README notes the training commands are configured for a node of 8 H100 GPUs at 80GB each, and recommends scaling the per-device batch size or gradient accumulation steps proportionally when changing the number of GPUs.

Installing Open R1 on a CUDA Cluster

The README requires CUDA 12.4 and warns that segmentation faults may occur if the version does not match. Installation uses the uv package manager. The Makefile's install target runs the full sequence:

bash
uv venv openr1 --python 3.11 && source openr1/bin/activate && uv pip install --upgrade pip

Next, install vLLM and FlashAttention. The README specifies an exact vLLM version because the binaries are compiled for a specific PyTorch version:

bash
uv pip install vllm==0.8.5.post1
uv pip install setuptools && uv pip install flash-attn --no-build-isolation

Then install the project's dependencies with the dev mode:

bash
GIT_LFS_SKIP_SMUDGE=1 uv pip install -e ".[dev]"

After installation, authenticate with the Hugging Face Hub and Weights and Biases:

bash
huggingface-cli login
wandb login

The README also requires Git LFS for loading and pushing models and datasets to the Hugging Face Hub. For Hugging Face cluster users, the README recommends adding `export UV_LINK_MODE=copy` to .bashrc to suppress cache warnings.

Running SFT and GRPO Training

Training supports two parallelism modes: DDP and DeepSpeed with ZeRO-2 and ZeRO-3. To run SFT on the Mixture-of-Thoughts dataset, the README gives this command:

bash
accelerate launch --config_file=recipes/accelerate_configs/zero3.yaml src/open_r1/sft.py \
    --model_name_or_path open-r1/Qwen2.5-Math-7B-RoPE-300k \
    --dataset_name open-r1/Mixture-of-Thoughts \
    --dataset_config all \
    --eos_token '<|im_end|>' \
    --learning_rate 4.0e-5 \
    --num_train_epochs 5 \
    --max_seq_length 32768 \
    --per_device_train_batch_size 2 \
    --gradient_checkpointing \
    --bf16 \
    --use_liger_kernel \
    --output_dir data/OpenR1-Distill-7B

By default, the script pushes each trained model to the user's Hugging Face Hub account. The scripts also accept YAML config files as an alternative to command-line arguments. Evaluation uses lighteval via the Makefile's evaluate target, with support for both data-parallel and tensor-parallel inference through vLLM.

Published Datasets: Mixture-of-Thoughts, OpenR1-Math-220k and CodeForces-CoTs

Open R1 has produced three datasets published on the Hugging Face Hub:

Mixture-of-Thoughts contains 350,000 verified reasoning traces distilled from DeepSeek-R1, spanning mathematics, coding and science tasks. The README describes it as designed to teach language models step-by-step reasoning. It was released in May 2025 alongside the OpenR1-Distill-7B model as the completion of step 1.

OpenR1-Math-220k contains 220,000 traces distilled from R1 on a version of NuminaMath. The README notes that models trained on this dataset match the performance of DeepSeek's own distilled models.

CodeForces-CoTs contains 10,000 competitive programming problems and 100,000 solutions distilled from R1. A 7B Qwen model trained on this dataset can outperform Claude 3.7 Sonnet on the IOI24 benchmark, according to the README; a 32B model can outperform R1 itself on the same benchmark.

All three datasets are public on the Hugging Face Hub and are the primary output of the step 1 work.

Limitations: Hardware Requirements and an Unfinished Pipeline

The hardware requirement is concrete and significant. The README states the training commands are configured for a node of 8 H100 GPUs with 80GB each, and notes that different hardware requires tuning the batch size or gradient accumulation steps. Running the full training pipeline without that hardware class means either accepting longer training times with reduced batch sizes or adapting the configuration manually. The SFT command in the README specifies per_device_train_batch_size 2 with gradient_checkpointing enabled, which is already conservative for H100 memory; smaller GPUs will need further adjustment.

The pipeline is explicitly unfinished. The README says "This repo is a work in progress, let's build it together." Steps 2 and 3, covering the pure RL pipeline and multi-stage training, are not complete. Researchers who want a finished, end-to-end reproducible DeepSeek-R1 clone will find that the current repository only covers distillation.

The README also notes that CUDA version mismatches cause segmentation faults, which is a harder failure mode than a clear error message. Verifying nvcc --version before installation is a required step. Additionally, the repository depends on vLLM pinned to version 0.8.5.post1 and PyTorch v2.6.0, meaning upgrading either independently of the other is likely to break the build.

License and Development Status

Open R1 is licensed under Apache-2.0, which permits commercial use, modification and redistribution with attribution. This is a permissive license, and the datasets published on the Hugging Face Hub each carry their own separate license terms that should be reviewed independently before use in a product. The setup.py file includes the Apache-2.0 header from the HuggingFace transformers project, indicating shared tooling and a consistent license baseline across the Hugging Face ecosystem.

The last push to the huggingface/open-r1 repository was on 2026-04-02. The repository is not archived. The repository has no GitHub releases; version tracking follows the commit log and the PyPI package if one is published. The README's news section stops at May 2025, and the last recorded repository push was April 2026, suggesting the initial phase of active development has slowed. The Makefile's test target runs pytest against the tests/ directory excluding tests/slow/, and a separate slow_test target runs the slower integration tests under tests/slow/, providing a path to verify the environment is correctly set up before attempting full training runs.

Editorial conclusion

Open R1 is the right starting point for ML researchers and engineers who want to reproduce or extend the DeepSeek-R1 reasoning pipeline without access to proprietary training infrastructure. Anyone planning to run training should verify their hardware matches the documented configuration of 8xH100 GPUs at 80GB each, since the README states the scripts are configured for that specific setup and require batch size or gradient accumulation tuning for different hardware. The pipeline is explicitly a work in progress; step 1 (distillation) is complete, but steps 2 and 3 (pure RL and multi-stage training) are not finished.

Frequently asked questions

What is Open R1?

Open R1 is an Apache-2.0 Python toolkit from Hugging Face that aims to fully reproduce the DeepSeek-R1 reasoning model training pipeline. It provides scripts for supervised fine-tuning and GRPO reinforcement learning, along with published datasets like Mixture-of-Thoughts.

What hardware does Open R1 require to run training?

The README states the training commands are configured for a node of 8 H100 GPUs at 80GB each. Different hardware requires adjusting the per-device batch size or gradient accumulation steps to keep the global batch size constant. CUDA 12.4 is also required; mismatches cause segmentation faults.

What datasets does Open R1 provide?

Open R1 has published three datasets on the Hugging Face Hub: Mixture-of-Thoughts (350k reasoning traces covering math, coding and science), OpenR1-Math-220k (220k traces on NuminaMath), and CodeForces-CoTs (10k competitive programming problems with 100k solutions distilled from R1).

Official sources

  1. huggingface/open-r1 on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/huggingface-open-r1.svg)](https://hysenlabs.com/projects/huggingface-open-r1)