Hugging Face TRL: post-training foundation models with SFT, DPO and GRPO
Train transformer language models with reinforcement learning. Built on top of the Transformers ecosystem, TRL supports a variety of model architectures and modalities, and can be scaled-up across various hardware setups.
At a glance
- What is it?
- TRL is a Python library that wraps the Transformers trainer with post-training methods such as SFT, DPO, GRPO and KTO, plus a CLI for running them without writing code. It fits teams whose data is already in the Hub format and who want the standard recipes; it is the wrong tool when you need a custom training loop or a non-Transformers model.
- Who is it for?
- Adopt TRL if your post-training data is already in a Hub dataset and you want SFT, DPO, GRPO or KTO without writing a training loop, especially if you need DeepSpeed, FSDP or PEFT scaling. Do not adopt it if you need a non-Transformers model, a fully custom optimizer step, or a stable API surface, since pyproject.toml still declares Development Status 2 - Pre-Alpha.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What problem TRL solves, and who it is for
Pretraining a language model is not the hard part for most teams. The hard part is what comes after: turning a base checkpoint into something that follows instructions, prefers one answer over another, or solves a math problem by reasoning. Each of those goals has its own algorithm and its own data format, and each algorithm has its own paper, loss function and set of hyperparameters.
TRL collects those algorithms behind trainer classes that share the Transformers trainer interface. The README describes the library as "a comprehensive library to post-train foundation models" and lists SFTTrainer, GRPOTrainer, DPOTrainer, KTOTrainer and RewardTrainer among the available trainers. The audience is developers and research groups who already work in the Hugging Face ecosystem and want the published recipe rather than a reimplementation of it.
The design assumption is worth stating plainly: TRL assumes your data lives on the Hub or in a datasets.Dataset. Every quick-start example in the README loads a dataset with load_dataset before constructing a trainer.
The trainer classes and what each algorithm assumes about your data
The trainers are not interchangeable, because the underlying algorithms consume different shapes of supervision. SFTTrainer takes plain text or prompt-completion pairs and does next-token training. DPOTrainer expects paired preferences, which is why the README example loads trl-lib/ultrafeedback_binarized, a dataset of chosen and rejected responses. KTOTrainer instead takes binary desirable/undesirable labels, and the README example loads trl-lib/kto-mix-14k. RewardTrainer trains a reward model on the same paired-preference data that DPO consumes.
GRPOTrainer is the one that changes the shape of the work rather than just the loss. It implements Group Relative Policy Optimization, which the README describes as "more memory-efficient than PPO" and credits as the method used to train Deepseek AI's R1. Instead of a trained reward model, it takes reward functions. The README example passes accuracy_reward directly as reward_funcs, and a note adds that reasoning models should use reasoning_accuracy_reward() instead. That means the reward is code you write and control, not a checkpoint you train first.
Each trainer is described in the README as "a light wrapper around the Transformers trainer" and is said to natively support DDP, DeepSpeed ZeRO and FSDP. That wrapper framing is accurate and also the main constraint: you inherit the Transformers training loop, including its TrainingArguments surface, rather than defining your own.
Installing TRL and running a first SFT job
The README gives three installation paths. The standard one is pip:
pip install trlIf you need features that have landed on main but are not in a release yet, the README offers a source install:
pip install git+https://github.com/huggingface/trl.gitAnd if you want the example scripts that ship with the project, clone the repository:
git clone https://github.com/huggingface/trl.gitBefore installing, check the version floors in pyproject.toml. TRL requires Python 3.10 or newer and pins accelerate>=1.4.0, datasets>=4.7.0 and transformers>=4.56.2. A stale transformers in your environment is the most likely cause of an import error.
The smallest working example from the README is the SFTTrainer, which needs only a model name and a dataset:
from trl import SFTTrainer
from datasets import load_dataset
dataset = load_dataset("trl-lib/Capybara", split="train")
trainer = SFTTrainer(
model="Qwen/Qwen2.5-0.5B",
train_dataset=dataset,
)
trainer.train()Running that script downloads Qwen/Qwen2.5-0.5B and the Capybara dataset, then starts a supervised fine-tuning run. The model is small enough that the point is to confirm the pipeline works, not to produce a useful checkpoint.
If you would rather not write Python at all, the project ships a CLI entry point, declared in pyproject.toml as trl = "trl.cli:main". The README's SFT invocation is:
trl sft --model_name_or_path Qwen/Qwen2.5-0.5B \
--dataset_name trl-lib/Capybara \
--output_dir Qwen2.5-0.5B-SFTThe same pattern covers DPO and KTO, with trl dpo and trl kto and their own --model_name_or_path, --dataset_name and --output_dir flags. The README points to the CLI documentation section and to --help for the rest of the options.
Where TRL stops being the right tool
The most concrete limitation is in the packaging metadata. pyproject.toml declares "Development Status :: 2 - Pre-Alpha" as a classifier, even though the project is at v1.12.0 with a stable DistillationTrainer announced in the README. That mismatch is not cosmetic. A repository that also ships a MIGRATION.md at the top level is telling you that APIs move between releases and that upgrading is a task with its own document.
The wrapper design is the second constraint. Because each trainer wraps the Transformers trainer, anything the base trainer cannot express is awkward here. If you need a custom optimizer schedule, a non-standard loss composition, or gradient manipulation between backward and step, you are fighting the abstraction rather than using it. TRL is a library of implemented algorithms, not a framework for inventing new ones.
The third constraint is the ecosystem boundary. TRL is built on Transformers, and the README's scaling story runs through Accelerate, DeepSpeed, PEFT and Unsloth. A model that does not load through Transformers is out of scope, and so is a training stack built on a different tensor framework.
Finally, reward functions in GRPOTrainer are code you supply. The README's examples use accuracy_reward and reasoning_accuracy_reward, which are verifiable for math and reasoning tasks. For open-ended objectives there is no built-in reward, and a poorly specified reward function will be optimized against exactly as written.
TRL compared with Unsloth and verl
The two comparisons people search for most are Unsloth and verl, and they differ from TRL in different directions.
Unsloth is a kernel-level optimization project. The TRL README treats it as an integration rather than a rival: it lists Unsloth among the things TRL integrates for "accelerating training using optimized kernels". So the relationship is closer to layering than substitution. If your goal is the fastest possible single-GPU LoRA run on a supported model, Unsloth's kernels are the reason to reach for it. If your goal is the algorithm itself, SFT, DPO, GRPO or KTO as published, TRL is the layer that implements it and Unsloth is an accelerator underneath.
verl sits on the other side. It is a reinforcement learning training framework for language models, and the practical difference is scope: TRL covers supervised fine-tuning, preference optimization and reward modeling in the same package as GRPO, while a dedicated RL framework assumes reinforcement learning is the whole job. That matters if you need large-scale RL rollout infrastructure, and it matters less if half your work is SFT on a single node.
The honest summary is that TRL's advantage is breadth behind one interface, and its cost is that no single trainer is as specialized as a tool built for one method.
Maintenance, upgrades and the Apache-2.0 licence
The repository is not archived, and the last push was on 2026-08-26, which is recent enough that the project is under active development. The release cadence supports that: v1.10.0 on 2026-08-13, v1.11.0 and v1.12.0 both on 2026-08-26. Two releases in one day suggests a fast merge-to-release loop, and it also means release notes are worth reading before you pin a version.
The upgrade cost is real and mostly self-inflicted by the project's own pace. MIGRATION.md exists at the top level for a reason, and the Pre-Alpha classifier means the maintainers are not promising API stability. Pin trl to an exact version in your lockfile, and read MIGRATION.md before bumping it. The dependency floors move too: pyproject.toml requires transformers>=4.56.2 and datasets>=4.7.0, and the deepspeed extra carries an explicit transformers!=5.1.0 exclusion with a comment pointing at an upstream issue. If you use DeepSpeed, that pin is not optional.
On licensing, TRL is Apache-2.0, with the licence text in LICENSE and declared in pyproject.toml via license = "Apache-2.0" and license-files = ["LICENSE"]. That is a permissive licence, but it covers TRL's code only. The models you fine-tune and the datasets you load carry their own licences, and TRL does nothing to reconcile them. Checking the terms on Qwen/Qwen2.5-0.5B and trl-lib/Capybara is your responsibility, not the library's. This is a description of the licence field, not legal advice.
Editorial conclusion
Adopt TRL if your post-training data is already in a Hub dataset and you want SFT, DPO, GRPO or KTO without writing a training loop, especially if you need DeepSpeed, FSDP or PEFT scaling. Do not adopt it if you need a non-Transformers model, a fully custom optimizer step, or a stable API surface, since pyproject.toml still declares Development Status 2 - Pre-Alpha. Before committing, check that transformers>=4.56.2 and datasets>=4.7.0 resolve in your environment, and read MIGRATION.md for the breaking changes between releases.
Frequently asked questions
What is Hugging Face TRL?
TRL is a Python library for post-training foundation models, described in its README as a comprehensive library for that purpose. It provides trainers for supervised fine-tuning, preference optimization and reinforcement learning methods including SFT, DPO, GRPO and KTO, built on top of the Transformers ecosystem.
What is TRL AI?
TRL is the Transformers Reinforcement Learning library maintained under the huggingface organization. Its README describes it as a comprehensive library to post-train foundation models, offering trainers such as SFTTrainer, GRPOTrainer, DPOTrainer and KTOTrainer.
What does Hugging Face Transformers do?
Transformers is the library TRL is built on. The README states that TRL is built on top of the Transformers ecosystem, and pyproject.toml requires transformers>=4.56.2 as a dependency.
What is Hugging Face used for?
In the context of TRL, Hugging Face is where the models and datasets used in the README's examples live. The quick-start examples load datasets such as trl-lib/Capybara and models such as Qwen/Qwen2.5-0.5B from the Hub.
What is huggingface trl?
huggingface/trl is the GitHub repository for the TRL library. It is written in Python, licensed Apache-2.0, and its README points to the documentation at hf.co/docs/trl.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/huggingface-trl)
Community notes