TextRL is a config dataclass and four trainers over HuggingFace TRL, and PPO is the one it removed
Implementation of ChatGPT RLHF (Reinforcement Learning with Human Feedback) on any generation model in huggingface's transformer (blommz-176B/bloom/gpt/bart/T5/MetaICL)
At a glance
- What is it?
- A thin layer that puts one dataclass in front of the upstream trainers, collapses eleven preference losses into a single loss-type argument, and takes rewards as ordinary callables. The design is small on purpose. The parts that will surprise you are the removals: PPO and five neighbouring algorithms are gone because upstream dropped them, and asking for one raises an error with a migration hint rather than a deprecation warning.
- Who is it for?
- Use it if you want TRL's trainers without hand-assembling configuration dicts, since the value is the collapse rather than the algorithms. Two checks before you start.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 163 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One config dataclass, four trainer families, rewards as callables
The stated ambition is ergonomic rather than novel. Configuration is a single dataclass, each algorithm family gets one trainer class, rewards are plain functions, and parameter-efficient fine-tuning, distributed launching, and fast rollout are all first-class options rather than integrations you assemble yourself:
pip install textrl # core
pip install 'textrl[quant]' # + bitsandbytes (QLoRA)
pip install 'textrl[vllm]' # + vLLM rollout
pip install 'textrl[quant,vllm,rewards]' # kitchen sinkEverything else, including the dev extra with pytest, a parallel test runner, a linter and a type checker, is optional. The layer sits on top of HuggingFace TRL rather than reimplementing it, which means your results depend on upstream defaults you do not see.
PPO is in the keywords and not in the library
The algorithm table has four rows and then a paragraph about what is missing. Online covers GRPO, RLOO and REINFORCE++. Pairwise preference covers eleven named losses including DPO, IPO, Hinge, several variants of APO and BCO, NCA, and two hybrids, all through the same trainer with a loss-type argument. Binary preference is KTO. Reward model training has its own trainer. Then the removals: PPO, OnlineDPO, ORPO, CPO, SimPO, and the binary form of BCO are not supported, because they were removed upstream in a later TRL release rather than by choice here. Asking for one raises an error carrying a migration hint. Two details make that list more than a changelog entry. The package keywords still include the abbreviation for the algorithm that is no longer available, and the project description still advertises support for a 176-billion-parameter model family that the body of the documentation never mentions.
Eleven preference losses collapse into one argument
The preference trainer is the part of the API worth copying. Instead of one class per objective, the family takes a loss-type parameter, so moving from DPO to a hinge loss to a null-anchored variant is a string change rather than a class swap. That matters most for research code, where the objective changes far more often than the data pipeline does. The four data shapes are equally small: prompts alone for the online algorithms, prompt with chosen and rejected for pairwise work, prompt with completion and a boolean label for binary feedback, and chosen with rejected for reward-model training. Datasets are built with three helpers, from a list, from a JSON lines file, or from a hub dataset, and you can pass any dataset object directly if you already have one. Nothing in the layer owns your preprocessing.
With adapters set, the reference model is free
The model loader returns a triple, the policy, the tokenizer, and a reference model or none, and the third element has a rule worth internalising. When a parameter-efficient configuration is supplied, the reference is none, because the reference forward pass is produced by disabling the adapters rather than by holding a second copy of the weights in memory. The loader also takes the adapter rank, its alpha, and a target-module selector that accepts a catch-all for all linear layers, a four-bit quantisation mode described as normal-float QLoRA, a dtype, and an attention implementation flag. There is a separate switch to skip loading a reference at all, and the documentation tells you to use it for the online algorithms to save memory. That advice only applies when you are not using adapters, since with adapters there is no second model to save.
Distributed training adds no scaffolding, and deepspeed arrives by smuggling
Multi-node and multi-GPU are handed to the accelerator library entirely. The project's own words are that TextRL adds no scaffolding of its own, and the launch is one command pointing at the package's training module with a config file. That is a defensible boundary and it also means every distributed problem is an upstream problem. The interesting detail is how a distributed strategy gets through. The config dataclass carries a distributed field holding a strategy name and a zero stage, and that field is not a first-class parameter of the trainer at all: it is forwarded to the underlying library through a generic extras dictionary. So the same mechanism carries the vLLM flags, and nothing in the type signature stops you putting a misspelled key in either place. Anything relying on it is untyped by construction.
Fast rollout exists for one algorithm only
The vLLM integration is the standard one: two extras in the config, one to enable it and one to set how much GPU memory the rollout engine may take, or a helper that builds the same dictionary for you. The limitation is the heading, which says GRPO only. The online family has three members, and the two others have no fast rollout path in this layer, so a run using them pays full generation cost through the training loop instead. That is a real throughput difference on the algorithm families where generation dominates, and it is easy to miss because the vLLM extra installs cleanly for every algorithm and only fails to apply to two of the three. If you are benchmarking the online family, the fastest path exists for one member of it.
A deprecated console script still ships on the install path
The command line has four entry points. One trains from a YAML config, one merges a parameter-efficient adapter into a standalone checkpoint, and one evaluates by rolling out against a dataset and a reward function specified as a module and function name, reporting reward statistics without training. The fourth is marked as a deprecated alias for the merge command. It is still registered as a console script, so it is installed on every machine alongside the three that work, and calling it presumably warns rather than failing. The config file format is the same object as the dataclass expressed in YAML, including a nested model name, a dataset specified by hub identifier with a split expression, and a reward given in module-colon-function notation. That last convention means a typo in a config file is a runtime import error rather than a schema failure.
The licence, the classifiers and the dependency floor all disagree
Three small inconsistencies, each of which will cost somebody time. The repository metadata reports the licence as MIT while the packaging configuration declares Apache and carries the matching classifier, so the licence a distributor reads and the licence the project page shows are different. The classifier set also declares the development status as production and stable, on a release whose headline change is the removal of a legacy API. And the dependency floor for the upstream library is set twelve releases below the version in which the removals happened, so a fresh install can legitimately resolve to a combination where the algorithms the README says are gone still exist upstream. The build also uses a legacy setup-time dependency declaration for a version helper even though the version is hardcoded, and there is no pyproject file, so the package builds through the older path only.
Editorial conclusion
Use it if you want TRL's trainers without hand-assembling configuration dicts, since the value is the collapse rather than the algorithms. Two checks before you start. Pick your algorithm from the supported list rather than from the package keywords, which still advertise an algorithm the library deliberately refuses. And note that the declared floor for the upstream dependency is far below the version where the removals happened, so pip can hand you a combination the README describes differently than the one you get.
Frequently asked questions
What is TextRL?
A thin, opinionated layer on top of HuggingFace TRL that makes reinforcement learning for text generation easier to use: one dataclass for configuration, one trainer class per algorithm family, rewards as ordinary callables, and first-class support for PEFT, accelerate, and vLLM. It is a layer, not a reimplementation, so defaults come from upstream.
Which algorithms does TextRL support?
Four families. Online covers GRPO, RLOO, and REINFORCE++. Pairwise preference covers eleven losses through one trainer and a loss-type argument, including DPO, IPO, Hinge, APO, BCO-pair, NCA-pair, AOT, DiscoPOP, SPPO-hard, and EXO-pair. Binary preference is KTO, and reward-model training has its own trainer. PPO, OnlineDPO, ORPO, CPO, SimPO, and binary BCO are not supported.
What happened to the old TextRL API?
Version 1.0 removed the legacy reinforcement-learning environment and gym-style API, including the environment class, the actor class, and the training entry point that took evaluation arguments. Requesting one of the removed algorithms now raises an error with a migration hint, and a migration document in the docs directory covers the change.
How do I write a reward function for TextRL?
As a plain callable taking prompts, completions, and any extra dataset columns, returning a list of floats. Decorating it with the reward helper coerces it into a protocol object, a base class exists for stateful rewards such as a loaded classifier, several rewards can be composed with weights, and a classifier wrapper takes any HuggingFace pipeline plus a target label.
What extras does TextRL provide?
Four. The quant extra adds bits-and-bytes for QLoRA, the vllm extra adds a fast rollout engine and works with GRPO only, the rewards extra adds the evaluation library with two scoring packages, and the dev extra adds pytest, a parallel test runner, a linter, and a type checker.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/voidful-textrl)