torchtune: PyTorch's post-training library, now wound down
PyTorch native post-training library
At a glance
- What is it?
- torchtune is a PyTorch-native library of hackable recipes for SFT, DPO, PPO, GRPO, knowledge distillation and QAT. Development wound down in 2025, so the practical question is whether its recipe set still fits your stack.
- Who is it for?
- torchtune fits teams already inside PyTorch who want readable recipe code and YAML configs for LoRA SFT, DPO or QAT, and who can accept a library whose development wound down in 2025. It is the wrong pick if you need a supported dependency with a release cadence, if your method is full-weight DPO or PPO with LoRA, or if you want a GUI.
- Can I use it commercially?
- Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 21 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What torchtune solves, and for whom
Post-training an LLM usually means stitching together a trainer, a tokenizer, a checkpoint format and a distributed launcher from different projects. torchtune's answer is to keep all of that inside PyTorch. The README describes it as a PyTorch library for authoring, post-training and experimenting with LLMs, with hackable training recipes, simple PyTorch implementations of models like Llama, Gemma, Mistral, Phi and Qwen, and YAML configs for training, evaluation, quantization or inference.
The audience is narrow and specific. If you already write PyTorch and want to read the training loop rather than configure someone else's trainer, the recipes are the point. Each recipe is a Python file under recipes/, and each config is a YAML file under recipes/configs/, so the thing you modify is ordinary Python and ordinary YAML. If you would rather never open a training loop, this is not aimed at you.
The support matrix is where the library is most honest. SFT with full weights or LoRA/QLoRA runs on one device, multiple devices and multiple nodes. Knowledge distillation supports LoRA/QLoRA on one or many devices but not multiple nodes, and full-weight KD is marked unsupported everywhere. DPO supports full weights only on more than one device, and LoRA/QLoRA on one or many devices, but not multiple nodes. PPO supports full weights on a single device only. GRPO shows full-weight multi-device and multi-node as under construction, with LoRA/QLoRA unsupported. QAT supports full weights on one or many devices and LoRA/QLoRA on more than one device, neither across nodes. Read that table before you plan anything.
How the recipe and config split actually works
There are two layers. The torchtune/ package holds the model builders, tokenizers, datasets and training utilities. The recipes/ directory holds runnable entry points such as lora_finetune_single_device, knowledge_distillation_distributed, lora_dpo_single_device and qat_distributed. A recipe is invoked through the tune command-line tool, which is declared in pyproject.toml as the console script tune = "torchtune._cli.tune:main".
Configuration is separate from code. A recipe takes a --config argument naming a YAML file, and the configs live in per-model directories. The README gives the pattern directly: tune run lora_finetune_single_device --config llama3_2/3B_lora_single_device. The same pattern appears for the other methods, for example tune run knowledge_distillation_distributed --config qwen2/1.5B_to_0.5B_KD_lora_distributed and tune run qat_distributed --config llama3_1/8B_qat_lora.
Because the config name encodes the model family and size, the practical workflow is to list what exists before writing anything. The README states that tune ls lora_finetune_single_device prints the full list of available configs, and the same command works for other recipes, such as tune ls full_dpo_distributed. That is the fastest way to find out whether your model and method combination already has a starting point.
Dependencies are declared in pyproject.toml and are ordinary PyTorch-adjacent packages: torchdata, datasets, huggingface_hub with the hf_transfer extra, safetensors, kagglehub, sentencepiece, tiktoken, blobfile>=2, tokenizers, numpy, tqdm, omegaconf, psutil and Pillow>=9.4.0. Two pins stand out. pyarrow is capped below 21.0.0 to avoid a breaking change, and the dev extra pins urllib3<2.0.0. The async_rl extra pulls torchrl and tensordict straight from git commits rather than a release, which the file itself flags as temporary.
Installing torchtune and running a first LoRA finetune
The README points to its installation section for setup and does not reproduce the commands in the text available here, so treat the package metadata as the reliable source for constraints. Python must be 3.9 or newer, and the package is published as torchtune on PyPI with the console script tune. A GPU environment is implied by the recipe table, since every recipe is listed per device count.
Start by confirming the interpreter and installing the library.
python --version
pip install torchtunepython --version should report 3.9 or later. After the install, the tune entry point should be on your PATH, since pyproject.toml declares it under [project.scripts].
Next, find a config that matches your model and hardware instead of guessing at one. This prints every config registered for the single-device LoRA recipe.
tune ls lora_finetune_single_deviceThe output is a list of config names in the form model_family/size_variant. Pick one from that list, then run the recipe with the same name you saw.
tune run lora_finetune_single_device --config llama3_2/3B_lora_single_deviceThat command is the README's own example. The recipe reads the YAML config, builds the model and tokenizer, loads the dataset, and runs the LoRA training loop. To change hyperparameters, edit the YAML file under recipes/configs/ rather than passing flags, which is the design the library is built around.
If you need a different method, the shape stays the same: swap the recipe name and the config. For distributed DPO the README's example is tune run lora_dpo_single_device --config llama3_1/8B_dpo_single_device, and for QAT it is tune run qat_distributed --config llama3_1/8B_qat_lora.
The maintenance question is not a footnote
The README opens with a warning that torchtune is no longer actively maintained, stating that development wound down in 2025 and linking to an issue titled The future of torchtune. That banner is the single most important fact about the project, and it changes how you should read everything else on this page.
The repository is not archived, and the last push was on 2026-09-09, so the code is still there and the branch still moves. That does not make it actively developed. The last tagged release listed is v0.6.1 from 2025-04-07, preceded by v0.6.0 on 2025-03-24 and v0.5.0 on 2024-12-20. The feature announcements in the README stop in May 2025 with Qwen3 support, following Llama4 in April 2025 and multi-node training in February 2025.
For a library you plan to depend on for a year, that timeline matters more than any individual feature. A wound-down project will not receive fixes for a new PyTorch release that changes an internal API, and the pinned pyarrow<21.0.0 will not be revisited. The recipes remain readable and forkable, which is the mitigation the library's own design offers: because a recipe is a Python file you can copy into your repository, the exit path is shorter than it would be with a framework that hides its training loop.
Where the support matrix rules torchtune out
The gaps in the README's own tables are the clearest limitations. Full-weight knowledge distillation is marked unsupported on one device, multiple devices and multiple nodes, so KD here means LoRA or QLoRA only. That is a real constraint if your goal is to compress a large teacher into a small student by moving all the weights.
DPO has a similar shape. Full-weight DPO is unsupported on a single device, and neither full-weight nor LoRA DPO is listed for multiple nodes. PPO is narrower still: full weights on one device, and LoRA/QLoRA unsupported everywhere. GRPO shows full-weight multi-device and multi-node as under construction, with LoRA/QLoRA unsupported. If your plan is parameter-efficient RLHF across several nodes, the table says no.
Multi-node is not uniform either. SFT carries the checkmarks across one device, more than one device and more than one node for both full and LoRA/QLoRA. Every other method drops at least one cell in that column. That asymmetry is worth internalizing before you size a cluster.
There is also a dependency risk that has nothing to do with training quality. The async_rl extra installs torchrl and tensordict from specific git commits rather than published releases, and pyproject.toml notes that this will be updated to a stable release. A wound-down project is unlikely to complete that update, so anything depending on async_rl is pinned to those commits.
torchtune compared with Unsloth and Axolotl
The most common comparison is with Unsloth, and the difference is architectural rather than a matter of speed. Unsloth is built around optimized kernels for a narrower set of model families, so you get a small surface area with heavy optimization underneath. torchtune keeps the training loop in plain PyTorch and exposes it as a recipe file you can read and edit. If you want to change how the loss is computed or how the dataloader batches, torchtune's model is to open the recipe. If you want the fastest path to a finetuned model without touching the loop, that is the opposite trade.
Axolotl sits in a different place again. It is a configuration-driven trainer where the YAML is the interface and the training code is not meant to be edited. torchtune also uses YAML, which makes the two look similar from a distance, but the configs here point at recipes you are expected to modify. The practical test: if your team's instinct on hitting a limitation is to fork the trainer, torchtune's layout supports that. If the instinct is to file an issue and wait, a configuration-first tool is a better fit.
The Hugging Face ecosystem is the third reference point, and it is less a competitor than an adjacency. torchtune depends on datasets, huggingface_hub[hf_transfer] and safetensors, and the README states that models are pulled from the Hugging Face Hub or Kaggle Hub. So the comparison is not torchtune versus Hugging Face but whether you want the training loop expressed in PyTorch primitives or through Transformers' Trainer abstraction.
Licence, upgrade cost and what a fork inherits
torchtune is BSD-3-Clause, declared in the LICENSE file at the repository root and referenced from pyproject.toml through license = {file = "LICENSE"}. That is a permissive licence, which matters for the fork path: if you copy a recipe into your own repository after development stops, the licence does not force you to publish your changes. This is a description of the licence text, not legal advice, and your own counsel should review anything you redistribute.
The upgrade cost is the part that is easy to underestimate. Because recipes are Python files and configs are YAML, an upgrade is not a dependency bump you can absorb silently. A new PyTorch release that changes an internal API can break a recipe, and with development wound down there is no upstream fix to pull. The realistic maintenance plan is to pin your torch version alongside torchtune, and to treat any recipe you rely on as code you now own.
That also means the version you pick matters more than usual. v0.6.1 from 2025-04-07 is the last release listed, and the features documented in the README, including Llama4 and Qwen3 support, are the ceiling of what you get. There is no roadmap to wait for.
Editorial conclusion
torchtune fits teams already inside PyTorch who want readable recipe code and YAML configs for LoRA SFT, DPO or QAT, and who can accept a library whose development wound down in 2025. It is the wrong pick if you need a supported dependency with a release cadence, if your method is full-weight DPO or PPO with LoRA, or if you want a GUI. Before adopting, check the support matrix in the README against your exact method, device count and model, then confirm the pinned pyarrow<21.0.0 and Python >=3.9 constraints resolve in your environment.
Frequently asked questions
How do I install torchtune?
The README points to its installation section, and the package metadata requires Python 3.9 or newer and publishes the package as torchtune with the console script tune. Installing it puts the tune command on your PATH, which is what you use to list configs and run recipes.
How does torchtune compare with Axolotl?
Both use YAML configuration, but torchtune's configs point at recipe files under recipes/ that you are expected to read and edit, while a configuration-first trainer treats the YAML as the interface. If your instinct on hitting a limit is to fork the training loop, torchtune's layout suits that better.
What are the alternatives to torchtune?
The README's model is to keep the training loop in plain PyTorch, which contrasts with Unsloth's optimized kernels for a narrower set of model families and with configuration-driven trainers that treat YAML as the interface. torchtune also depends on Hugging Face packages such as datasets and safetensors rather than replacing them.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/meta-pytorch-torchtune)