Library / SDK
huggingface/finetrainers avatar
huggingface/finetrainers

finetrainers: one training entry point for diffusion and video models

Scalable and memory-optimized training of diffusion models

1,359 stars142 forksPythonApache-2.0

At a glance

What is it?
A Hugging Face library that replaces per-model finetuning scripts with a standardized model specification and two trainers, SFT and Control. Useful for video diffusion work on limited GPUs, with a support matrix that is candid about where the numbers are not known yet.
Who is it for?
finetrainers is at its best for video diffusion fine-tuning on a single consumer GPU, where a LoRA run on LTX-Video fits in 5 GB and a full finetune of CogVideoX-5b fits in 53, and where writing your own training loop would mean reimplementing attention backends and precomputation.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 19 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

From one script per model to a standardized specification

The idea behind finetrainers is stated plainly in the README: it is a work-in-progress library to support accessible training of diffusion models and various commonly used training algorithms. The structural change arrived on 2025-03-03, when the project shipped a complete refactor for multi-backend distributed training, better precomputation handling for large datasets, FSDP support, and what the README calls a model specification format that is externally usable for training custom models.

That last part is the real contribution. Instead of every supported model carrying its own training script with its own argument names, the model is described in a standard form and the trainers consume that description. Two trainers exist today: an SFT Trainer and a Control Trainer. Anything you add to the specification is then available to both, which is why supporting Wan image-conditioning on a text-to-video model or CogView4 control conditioning was a feature entry rather than a new script.

The trade is that you accept a layer of indirection. When training misbehaves, the debugging path runs through the specification and the trainer rather than through a script you can read top to bottom.

Installing against the diffusers main branch on purpose

The quickstart is short and contains an instruction worth heeding. Requirements are installed from the file, and diffusers is installed from source rather than from a release:

bash
pip install -r requirements.txt
pip install git+https://github.com/huggingface/diffusers

The README explains why: requirements.txt pins `diffusers>=0.32.1`, but the recommendation is the `main` branch for the latest features and bugfixes. Requirements also include torch at 2.5.1 or higher, torchvision, torchao, accelerate, bitsandbytes, peft, datasets at 3.3.2 or higher, transformers at 4.45.2 or higher, decord for video decoding, wandb, kornia, sentencepiece and imageio-ffmpeg. That list tells you the intended workload: video diffusion, with frame decoding and dataset handling treated as first-class.

There is a hard version note as well. The README says PyTorch 2.5.1 or above is recommended, and that earlier versions can produce completely black videos or out-of-memory errors and are not tested. Reproducible runs are pointed at docs/environment.md.

bash
git fetch --all --tags
git checkout tags/v0.2.0

That second block is the escape hatch. Because `main` is also the development branch, the README is explicit that stable support should come from the release tags, and it sends you to a separate release branch for the instructions that match v0.2.0.

Reading the support matrix as a VRAM budget

The support matrix is the most useful table in the README, and it carries a warning: the numbers were obtained from the release branch, and `main` is unstable at the moment and may use higher memory. In other words, the table is a floor from a known-good point, not a promise about the branch you just cloned.

For text-to-video, LTX-Video needs 5 GB for LoRA and 21 GB for full finetuning. HunyuanVideo needs 32 GB for LoRA, and full finetuning is listed as OOM. CogVideoX-5b needs 18 GB for LoRA and 53 GB for full finetuning. Wan is listed as TODO for both. CogView4 appears in the matrix with its own row.

Reading across the rows gives a clear instruction. LoRA is the only path most people have a card for, and the spread between 5 GB and 32 GB means model choice matters more than any flag in docs/args.md. Full-rank finetuning is a multi-GPU or workstation activity, and HunyuanVideo is explicitly out of reach that way.

Two caveats sit under the table. The asterisk and caret footnotes are not spelled out in the visible text, so the exact conditions behind each figure are in the model docs rather than the table. And a row reading TODO means the number is unknown, not small.

Distributed backends, attention providers, and FP8

The feature list is where this library separates from a single training script. DDP, FSDP-2, HSDP and context parallelism are all supported, so the same trainer covers one GPU and many. Alongside that sit memory-efficient single-GPU training, LoRA and full-rank finetuning, conditional Control training, and memory-efficient precomputation with or without on-the-fly precomputation for large scale datasets.

Attention is treated as a pluggable choice. Four providers are listed: `flash`, `flex`, `sage` and `xformers`, each documented under docs/models/attention.md. A separate news entry on 2025-04-25 records support for different attention providers, and 2025-04-08 added `torch.compile`. For anyone who has lost a day to attention memory on a long video sequence, having the provider be an argument rather than a patch is the useful part.

On precision, the feature list claims fake FP8 training with QAT upcoming, and a January 2025 entry describes naive FP8 weight-casting training that the README says allows training HunyuanVideo in under 24 GB up to specific resolutions. Note the tension with the support matrix: 32 GB for HunyuanVideo LoRA in one place, under 24 GB in another. The matrix numbers are flagged as coming from the release branch, which is the likeliest explanation, but the README does not reconcile the two.

Dataset handling is automatic where it can be. The library detects commonly used dataset formats, supports combined image and video datasets, chains multiple local or remote datasets, and does multi-resolution bucketing. That removes the usual week of writing a preprocessing script.

What moved from a personal repo, and the packaging leftovers

This project changed homes, and the metadata shows it. The repository now sits under the Hugging Face organisation, but setup.py still declares its url as the a-r-r-o-w namespace, and the README points at release branch URLs on that same older namespace. The author field names Aryan V S. If you are following documentation, expect to see both paths.

The license needs a note too. The repository's declared license is Apache-2.0 and the LICENSE file at the root is the Apache text, while the trove classifiers in setup.py still declare Apache-2.0 as the license argument but list the OSI-approved MIT License among the classifiers. Those two statements disagree. Apache-2.0 is the one the repository metadata and the LICENSE file agree on.

Smaller inconsistencies exist too. setup.py pins the dev extras to pytest 8.3.2 and ruff 0.1.5, while requirements.txt pins ruff at 0.9.10. The Makefile exposes `make quality` and `make style` targets that run ruff check and ruff format over the finetrainers package, tests, examples and train.py, excluding examples/_legacy. pyproject.toml sets a line length of 119, ignores E501 entirely, and configures double quotes with space indentation, so the code style is enforced rather than aspirational.

Development status is not a guess here. The last push was on 2026-09-17 and the repository is not archived, while the newest release tag is v0.2.0 from 2025-04-25. The news list stops in April 2025 even though the branch keeps moving, which is why the README frames tags as the stable surface and main as where features land.

Editorial conclusion

finetrainers is at its best for video diffusion fine-tuning on a single consumer GPU, where a LoRA run on LTX-Video fits in 5 GB and a full finetune of CogVideoX-5b fits in 53, and where writing your own training loop would mean reimplementing attention backends and precomputation. It suits less well if you need a stable tagged release for production training, since the latest tag is v0.2.0 from April 2025 while the README directs you to the main branch for features, and it does not cover image-only diffusion beyond what Control conditioning adds. Start from one of the reproducible example scripts under examples/training/sft, match the environment pinned in docs/environment.md, and check the support matrix row for your model before budgeting VRAM.

Frequently asked questions

How much VRAM does finetrainers need for LoRA training?

The README support matrix lists 5 GB for LTX-Video, 18 GB for CogVideoX-5b and 32 GB for HunyuanVideo on text-to-video LoRA. The table notes these figures come from the release branch, and that the main branch is unstable and may use higher memory.

Does finetrainers support full-rank finetuning as well as LoRA?

Both. The matrix lists full finetuning figures of 21 GB for LTX-Video and 53 GB for CogVideoX-5b, and lists HunyuanVideo full finetuning as OOM. The feature list also names conditional Control training as a third mode.

Can I train a custom model with finetrainers?

Yes, through the model specification format. The 2025-03-03 refactor entry describes it as externally usable for training custom models, so a new model is described rather than wrapped in its own script. The two trainers, SFT and Control, both consume that specification.

Which version of PyTorch does finetrainers expect?

The README recommends PyTorch 2.5.1 or above and states that earlier versions can lead to completely black videos, out-of-memory errors or other issues, and are not tested. For reproducible runs the README points to docs/environment.md.

What license is finetrainers released under?

The repository metadata and the LICENSE file both state Apache-2.0. The trove classifiers in setup.py are inconsistent, listing an OSI-approved MIT License alongside the Apache-2.0 license argument, so the LICENSE file is the more reliable statement.

Official sources

  1. huggingface/finetrainers on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/huggingface-finetrainers.svg)](https://hysenlabs.com/projects/huggingface-finetrainers)