Open-source project
allenai/open-instruct avatar
allenai/open-instruct

allenai/open-instruct: a post-training stack for SFT, DPO and RLVR

AllenAI's post-training codebase

3,881 stars592 forksPythonApache-2.0

At a glance

What is it?
Open Instruct is AllenAI's codebase for instruction tuning and post-training pretrained language models, the same stack behind the Tulu 3 recipes. It is built for people with multi-GPU machines, not for laptops.
Who is it for?
Adopt Open Instruct if you already have a multi-GPU Linux box and want to reproduce or extend the Tulu 3 style pipeline, from SFT through DPO to RLVR with verifiable rewards, and you are comfortable reading scripts rather than a polished CLI. Skip it if you only need to fine-tune a small model on one GPU, or if you want a maintained evaluation suite: the README states the in-repo evaluations are unmaintained and points to OLMES instead.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Open Instruct is for, and who it is aimed at

Post-training a pretrained language model is not one job. It is at least three: supervised fine-tuning on instruction data, preference optimization against chosen and rejected pairs, and reinforcement learning against a reward signal. Each stage has its own data format, its own loss, and its own failure modes. Open Instruct exists to put those stages in one repository with a unified data format, so that a model can move from base checkpoint to SFT to DPO to RLVR without re-engineering the training loop each time.

The README frames the project as an open effort on instruction-tuning and post-training popular pretrained language models on publicly available datasets. The audience is implied by what ships alongside the code: checkpoints for Llama 3.1 8B and 70B, OLMo-2 7B and 13B, at the SFT, DPO, final RLVR and reward-model stages. If you want to know what the pipeline produces before you run it, those artifacts are the reference points.

The project is not a library you import into an existing trainer. It is a collection of training entry points, configs and scripts, plus supporting pieces like a decontamination directory and a human evaluation directory. The intended user is someone who will read those scripts and adapt them, not someone looking for a one-line API.

How the training stages and data flow fit together

The pipeline is staged. A base model is fine-tuned on instruction data to produce an SFT checkpoint. That checkpoint is then optimized against preference pairs with DPO, or trained with reinforcement learning using verifiable rewards, which is the RLVR stage named in the README. The final Tulu 3 models in the table are the RLVR stage; the SFT and DPO checkpoints are published separately, which tells you the stages are meant to be inspected and reused independently.

A reward model is trained alongside. The table lists 8B and 7B reward models, and notes that the 70B and 13B rows reuse the smaller reward model rather than training a larger one. That is a design decision worth noticing: reward modeling capacity is not scaled with policy size in the published configuration.

The mechanics live in the dependency list. Ray and vLLM are both present, which points at a distributed training loop where generation for rollout or evaluation is served by an inference engine rather than the training process. DeepSpeed and accelerate handle the training side. PEFT, bitsandbytes and LoRA-related tooling cover parameter-efficient runs. The repository also carries mason.py at the top level, a launcher script, and test_mason.py next to it.

Evaluation is explicitly out of scope now. The README says the codebase supports some evaluations natively but that these are now unmaintained, and directs users to OLMES, the tool used for Tulu 3. Treat the eval scripts as historical.

Setting up Open Instruct and checking the checkout

The project pins Python tightly: pyproject.toml declares requires-python as 3.12.*, so 3.11 or 3.13 will not satisfy it. Dependencies are managed with uv, and requirements.txt is generated from the lockfile, as its header comment shows. The Makefile runs everything through uv as well, for both formatting and type checking.

If you would rather not manage CUDA, toolchain and driver versions yourself, the repository ships a Dockerfile that builds on NVIDIA base images. The build argument selects the CUDA line:

dockerfile
ARG CUDA_VERSION=12
FROM nvidia/cuda:12.8.1-devel-ubuntu22.04 AS cuda12
FROM nvidia/cuda:13.0.3-devel-ubuntu22.04 AS cuda13
FROM cuda${CUDA_VERSION}

After setup, the code-quality targets are the quickest way to confirm the checkout is intact:

bash
make style-check
make quality-check

The Makefile exports PYTHONPATH as open_instruct so that scripts import the local checkout rather than an installed copy, and it runs ruff format, ruff check and ty check over open_instruct and mason.py. Passing those targets means the environment resolves, not that a training run will start. For the actual training entry points, the README and repository layout point at the configs and scripts directories; the documentation does not spell out a single canonical first command, so expect to read a config before launching anything.

Platform limits: where Open Instruct will not run

The dependency markers are unusually explicit, and they define the supported surface. vLLM is marked for platforms other than Darwin. flash-attn carries the same Darwin exclusion plus a platform_machine != 'aarch64' condition, and bitsandbytes is also excluded on Darwin. On macOS you get a partial install at best, and the generation path that the training loop depends on is not there.

The aarch64 exclusion on flash-attn is the sharper edge. Linux on ARM, which includes many cloud instances and Apple-silicon-adjacent environments, is excluded from that package even though the system is Linux. The Dockerfile targets x86_64 Debian packages throughout, including the Mellanox drivers and the Google Cloud CLI repository, so the container route assumes x86_64 as well.

There is also a version-availability constraint. The pinned transformers requirement is >=5.4.0 and torch is >=2.10.0. Those are recent major lines, and the rest of the stack (vLLM, deepspeed, peft, liger-kernel) has to be compatible with them simultaneously. If your existing training environment is pinned to an older transformers, Open Instruct will not slot into it without an upgrade you may not control.

Finally, the README states that native evaluations are unmaintained. If your workflow depends on running evals inside the same repository as training, you are adopting a codebase whose own maintainers moved that responsibility elsewhere.

How Open Instruct differs from Hugging Face TRL and Axolotl

The closest general-purpose alternative is TRL from Hugging Face, which also covers SFT, DPO and reward modeling, and which is designed as a library of trainers you import and compose. Open Instruct is the opposite shape: a research codebase with scripts and configs, where you adopt the whole pipeline or copy the parts you need. TRL integrates with the wider Hugging Face ecosystem and its Trainer abstractions; Open Instruct carries its own launcher and its own assumptions about Ray and vLLM being present.

Axolotl is the other common comparison point, a YAML-driven fine-tuning tool that tries to make configuration declarative and to hide the training loop. Open Instruct does not hide the loop. The trade-off is real in both directions: Axolotl and TRL will get a single-GPU LoRA run going faster, while Open Instruct gives you the exact configuration AllenAI used to produce published Tulu 3 checkpoints at 8B and 70B scale. If your goal is to match a known result rather than to get any result, that difference matters. If your goal is a quick adapter on a 7B model, it does not.

Maintenance, licensing and what an upgrade costs

The repository is not archived and the last push was on 2026-09-23. Releases are tagged: v0.1.0 in January 2026, v0.2.0 in March 2026, v0.3.0 in June 2026. That is a steady cadence rather than a frozen snapshot, and the CHANGELOG.md at the top level is where the project records what moved between them.

The licence is Apache-2.0, which permits commercial use and modification with the usual notice and patent terms. That covers the code in this repository. It does not automatically cover the model checkpoints linked from the README, which live on Hugging Face under their own terms, and it does not cover the datasets the recipes are trained on. If you intend to ship a derivative model, the dataset licences are a separate question from the codebase licence. This is not legal advice; read the terms attached to each artifact you use.

Upgrade cost is dominated by the dependency pins, not by the Python code. A jump between releases can move torch, transformers, vLLM and deepspeed together, and each of those has its own CUDA and driver requirements. The Dockerfile is the mitigation: rebuilding the image against the CUDA line you have is cheaper than reconciling a hand-built environment. Budget for a full environment rebuild, not a package bump, when moving between tagged versions.

Editorial conclusion

Adopt Open Instruct if you already have a multi-GPU Linux box and want to reproduce or extend the Tulu 3 style pipeline, from SFT through DPO to RLVR with verifiable rewards, and you are comfortable reading scripts rather than a polished CLI. Skip it if you only need to fine-tune a small model on one GPU, or if you want a maintained evaluation suite: the README states the in-repo evaluations are unmaintained and points to OLMES instead. Before you commit, check pyproject.toml for the requires-python pin of 3.12 and the Darwin and aarch64 markers on vllm, flash-attn and bitsandbytes, then confirm your CUDA driver matches the 12.8.1 or 13.0.3 base image the Dockerfile builds from.

Frequently asked questions

What Python version does allenai/open-instruct require?

The pyproject.toml pins requires-python to 3.12.*, so the project expects exactly the 3.12 series. Earlier or later minor versions will not satisfy the constraint as declared.

Can I run allenai/open-instruct on macOS or on an ARM machine?

The dependency markers exclude vLLM, flash-attn and bitsandbytes on Darwin, and flash-attn is also excluded when platform_machine is aarch64. That leaves the generation path the training loop relies on unavailable on those platforms.

Which training stages does allenai/open-instruct cover?

The README describes finetuning with instruction datasets, DPO and preference finetuning, and reinforcement learning with verifiable rewards, referred to as RLVR. Published checkpoints exist for the SFT, DPO, RLVR and reward-model stages.

Does allenai/open-instruct include evaluation scripts?

It contains evaluation code, but the README states that the native evaluations are now unmaintained and recommends OLMES instead, which is what was used for Tulu 3.

What licence does allenai/open-instruct use?

The repository is licensed under Apache-2.0. That covers the code; the model checkpoints and datasets referenced by the README carry their own separate terms.

Official sources

  1. allenai/open-instruct on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/allenai-open-instruct.svg)](https://hysenlabs.com/projects/allenai-open-instruct)