huggingface/alignment-handbook: post-training recipes for SFT, DPO and ORPO
Robust recipes to align language models with human and AI preferences
At a glance
- What is it?
- The Alignment Handbook packages the training pipeline behind Zephyr and SmolLM as YAML recipes plus scripts, covering continued pretraining, supervised fine-tuning, preference alignment and evaluation. The trade-off is a pinned, hardware-specific stack that assumes you already have GPUs and a Hugging Face account.
- Who is it for?
- Adopt the Alignment Handbook if you already have multi-GPU hardware and want a documented path from continued pretraining through SFT to DPO or ORPO, using the Zephyr and SmolLM recipes as templates for your own YAML. Skip it if you only need a single fine-tune on one GPU or you cannot match the pinned PyTorch 2.6.0 and flash-attn 2.7.4.post1 environment, because the README treats that version as a reproducibility requirement.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap the Alignment Handbook fills
The README opens with a short history: a year before it was written, chatbots were out of fashion, then ChatGPT and the Llama releases pushed the ML community into building its own assistants. The ecosystem that followed, in the project's own framing, mostly taught models to follow instructions through supervised fine-tuning. Preference data was the missing half. The InstructGPT and Llama2 papers are cited as evidence that adding human or AI preferences on top of SFT produces gains in helpfulness and safety, and the README states plainly that public resources on how to train these models, what data to collect and what metrics to measure were scarce.
That is the audience: engineers and researchers who already know the transformer stack and want the post-training pipeline spelled out rather than reverse-engineered from a paper. The repository is not a library you import into a larger application. It is a set of training scripts, YAML configuration files and dataset-formatting instructions that you copy, edit and run. If you are looking for a one-line fine-tuning API, this is the wrong shape of project.
Four training stages and how the scripts are organised
The repository layout is deliberately flat. The README describes the project as simple by design, consisting of a scripts directory for training and evaluation and a recipes directory for reproducing named models. Four steps are listed: continued pretraining, supervised fine-tuning for chat, preference alignment with DPO, and ORPO, which combines SFT and preference alignment in a single stage. Each script supports distributed training of full model weights with DeepSpeed ZeRO-3, or LoRA and QLoRA for parameter-efficient fine-tuning. That split matters more than it first appears. ZeRO-3 shards optimizer state, gradients and parameters across devices, so it lets you train models that do not fit on one GPU, at the cost of heavy inter-node communication. LoRA and QLoRA keep the base weights frozen or quantised and train small adapters, which fits on far less hardware but changes what you can claim about the resulting checkpoint.
Recipes are YAML files, and the README is explicit that each one contains all parameters associated with a single training run. A gpt2-nl recipe is included specifically to demonstrate language or domain adaptation: continue pretraining on a different language, then SFT and DPO the result. The techniques listed in the contents section go beyond the four scripts: reward modeling, rejection sampling, DPO as an alternative to PPO, and ORPO. The release history shows the recipes being used on real models, including Zephyr 141B (A35B) with ORPO on Mixtral 8x22B, StarChat2 15B, Zephyr 7B Gemma with RLAIF, and a Constitutional AI recipe.
Installing the handbook with uv and PyTorch 2.6.0
The README gives a linear installation sequence. First create a Python 3.11 virtual environment with uv and upgrade pip inside it:
uv venv handbook --python 3.11 && source handbook/bin/activate && uv pip install --upgrade pipIf uv is not installed, the README points to the UV installation guide rather than describing the process. Next comes PyTorch, and the README flags the version as important for reproducibility. The command it gives targets CUDA 12.6 wheels:
uv pip install torch==2.6.0 --index-url https://download.pytorch.org/whl/cu126Because the correct build depends on your hardware, the README also points to the PyTorch installation page instead of claiming one command works everywhere. The remaining dependencies come from the repository itself:
uv pip install .Flash Attention 2 is a separate install with build isolation disabled, which is the usual workaround for a package that compiles against the local CUDA toolchain:
uv pip install "flash-attn==2.7.4.post1" --no-build-isolationFinally, log in to the Hugging Face Hub and install Git LFS so you can push models:
huggingface-cli login
sudo apt-get install git-lfsA first real use is not a single command. The README recommends replicating Zephyr-7b-beta by following the recipe instructions in recipes/zephyr-7b-beta/README.md, and points to scripts/README.md for the dataset formatting rules if you want to train on your own data. Expect the first run to be a configuration exercise: pick a recipe, point it at your dataset, and decide between the ZeRO-3 and LoRA paths before launching.
What the pinned stack costs you
The installation section is unusually strict for a research repository, and the strictness is the point. PyTorch v2.6.0 is pinned with an explicit note that the precise version matters for reproducibility, and flash-attn is pinned to 2.7.4.post1. That is good practice for reproducing published checkpoints, and a real constraint if your cluster ships a different CUDA version or a vendor-patched PyTorch. You are not told how much of the pipeline still works on other versions, because the README does not make that claim.
The deeper limitation is hardware. Distributed training with DeepSpeed ZeRO-3 and full weight updates is not a laptop workflow, and the recipes span models from small ones up to Mixtral 8x22B in the Zephyr 141B release. Copying a 141B recipe onto a single GPU will fail, and the README does not provide a table mapping recipe to minimum hardware. The parameter-efficient path is the escape hatch, but LoRA and QLoRA are not drop-in equivalents for full fine-tuning, and the handbook does not present them as such. There is also no rollback story: the README does not document how to revert a training run or recover a checkpoint, so version control of your YAML and your Hub repository is on you. Finally, the FAQ-style questions people actually search for, about types of alignment or principles of an alignment framework, are not answered by this repository. It is a training codebase, not a textbook.
How it compares with TRL alone or a hosted fine-tuning service
The closest comparison is TRL, the Hugging Face library that provides the trainers themselves. The handbook's dependencies include TRL, and the recipes are the layer above it: the part that decides which trainer, which hyperparameters and which data format. If you use TRL directly you write the training loop and the configuration yourself and get full control over every knob. If you use the handbook you inherit a configuration that has already produced a named model, and you inherit its assumptions about sequence length, batch size and hardware along with it. Neither is strictly better. The handbook saves you the first week of guessing and costs you flexibility when your setup diverges from the recipe.
A hosted fine-tuning service sits at the opposite end. You upload a dataset, choose a base model, and never see a YAML file, a DeepSpeed config or a flash-attn build. The trade-off is that you cannot continue pretraining on a new language, you cannot run ORPO as a distinct stage, and you cannot inspect the exact parameters that produced the result. The handbook's value is precisely that all of that is visible in the repository. Its cost is that all of it is your problem when something breaks. The gpt2-nl recipe is a good illustration of where the handbook wins: a three-stage language adaptation pipeline is not something a hosted service exposes, and it is not something TRL gives you as a single artifact either.
Maintenance, licence and upgrade surface
The repository is not archived, and the last push was on 2026-05-26. That is a recent commit history but not a continuous release cadence, and the news section shows the pattern: dated releases tied to model launches, from the Zephyr-7b-beta training code in November 2023 through the SmolLM2-Instruct recipe in November 2024 to the SmolLM3-3B post-training recipe in July 2025. The handbook tracks the models Hugging Face is publishing rather than following a fixed schedule, so plan for upgrades to arrive when a new recipe lands, not on a version cadence.
Upgrade cost concentrates in the pinned dependencies. Because setup.py pins fast-moving packages such as transformers to exact versions, moving to a newer transformers or PyTorch means editing the dependency list and re-validating the recipe rather than just bumping a number. The Makefile exposes the quality gates you would run before sending a change: black at line length 119, isort, and flake8 with a max line length of 119 over src, tests and scripts. There is also a release helper invoked as python src/alignment/release.py with --patch and --post_release variants, which suggests the project manages its own version bumps through that script.
The licence is Apache-2.0, declared in the LICENSE file and reproduced in the header of setup.py. That is a permissive licence, but it covers the handbook code, not the models or datasets you train with it. The README links to a collection of models and datasets and to a technical report, and those carry their own terms. Nothing here is legal advice; check the licence of each base model and dataset before you redistribute a checkpoint.
Editorial conclusion
Adopt the Alignment Handbook if you already have multi-GPU hardware and want a documented path from continued pretraining through SFT to DPO or ORPO, using the Zephyr and SmolLM recipes as templates for your own YAML. Skip it if you only need a single fine-tune on one GPU or you cannot match the pinned PyTorch 2.6.0 and flash-attn 2.7.4.post1 environment, because the README treats that version as a reproducibility requirement. Before committing, verify that your dataset matches the formatting instructions in scripts/README.md, that the recipe you copy targets the model size you can actually fit, and that the DeepSpeed ZeRO-3 path in the script you pick supports your cluster's interconnect.
Frequently asked questions
What are the four types of alignment in the alignment-handbook?
The README lists four training steps: continued pretraining, supervised fine-tuning for chat, preference alignment with DPO, and supervised fine-tuning with preference alignment using ORPO. Each has a corresponding script, and each supports DeepSpeed ZeRO-3 or LoRA and QLoRA.
What does instructional alignment mean in the alignment-handbook?
The repository does not use that term. Its closest equivalent is supervised fine-tuning, described in the README as teaching language models to follow instructions, along with guidance on collecting and curating the training dataset.
What are the key principles of alignment in the alignment-handbook?
The handbook does not present a principles list. It offers training recipes: continued pretraining, supervised fine-tuning, reward modeling, rejection sampling, DPO and ORPO, each backed by a script and a YAML recipe rather than a set of stated principles.
What is the purpose of the Alignment Framework in the alignment-handbook?
The README does not describe an Alignment Framework by that name. Its stated purpose is to fill the gap in public resources on how to train preference-aligned models, what data to collect and what metrics to measure, by providing recipes that span the whole pipeline.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/huggingface-alignment-handbook)