Library / SDK
lucidrains/denoising-diffusion-pytorch avatar
lucidrains/denoising-diffusion-pytorch

denoising-diffusion-pytorch: a minimal DDPM implementation you can train on your own images

Implementation of Denoising Diffusion Probabilistic Model in Pytorch

10,700 stars1,280 forksPythonMIT

At a glance

What is it?
lucidrains/denoising-diffusion-pytorch packages a DDPM in about the smallest API that still trains: a Unet, a GaussianDiffusion wrapper, and a Trainer that reads a folder of images. It is a research and teaching codebase, not a production image generator, and the README makes that boundary fairly clear.
Who is it for?
Adopt it if you want to read, modify and train a diffusion model on your own image folder without pulling in a full training framework, or if you need a 1D variant for sequences. Do not adopt it if you need a production inference service, a hosted checkpoint, or a maintained evaluation pipeline for non-image data.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What denoising-diffusion-pytorch actually is, and who it is for

This is an implementation of the Denoising Diffusion Probabilistic Model paper (Ho, Jain, Abbeel, NeurIPS 2020) in PyTorch, written by Phil Wang. The README describes it as "a new approach to generative modeling" that uses denoising score matching to estimate the gradient of the data distribution, followed by Langevin sampling. It is not a product. It is a library of the pieces you need to build one.

The intended reader is someone who wants to train a diffusion model on their own images and is comfortable reading Python. The README's own example is four lines of setup and one line of loss. That is the whole pitch: the abstraction boundary is thin enough that you can see the noise schedule, the timestep sampling and the sampling loop without digging through a framework.

It is the wrong tool for anyone who wants a pretrained model they can call over HTTP. There is no checkpoint in the repository, no inference server, and no hosted demo. The README points to the official TensorFlow implementation as inspiration and to a Flax port and a Hugging Face annotated walkthrough as adjacent material, but the package itself ships code, not weights.

How the Unet and GaussianDiffusion objects fit together

The data flow is two objects. Unet is the noise predictor: you give it dim, dim_mults and optionally flash_attn, and it returns a network shaped for the image size you will declare later. GaussianDiffusion wraps that network, holds image_size and timesteps, and turns a batch of normalized images into a scalar loss.

Training is a single call. The README's example passes a tensor of shape (8, 3, 128, 128) with values normalized from 0 to 1 into diffusion(training_images), gets back a loss, and calls backward on it. Sampling is a separate call, diffusion.sample(batch_size = 4), which returns a tensor of shape (4, 3, 128, 128).

The sampling_timesteps parameter is where the speed trade-off lives. The README notes that setting it (250 in the example, against 1000 training timesteps) uses DDIM for faster inference, citing the DDIM paper. So the same object supports the slower ancestral sampler and the faster deterministic one, and you choose by setting a constructor argument rather than swapping classes.

The 1D path is a parallel set of classes: Unet1D, GaussianDiffusion1D, Trainer1D and Dataset1D. GaussianDiffusion1D takes seq_length and an objective argument, which the README shows as 'pred_v'. That is a different parameterization from the image path, and it matters if you are porting image recipes to sequence data.

Installing denoising-diffusion-pytorch and running a first training job

The README gives one install command. It pulls the dependencies declared in pyproject.toml, which include torch>=2.0, accelerate, einops, ema-pytorch>=0.4.2, pytorch-fid and scipy. The Python requirement is >=3.8. Note that torch is a hard dependency, so the install will resolve a CUDA or CPU build according to your environment.

bash
$ pip install denoising_diffusion_pytorch

The fastest way to see the loss path work is the README's minimal example. It builds a Unet with dim 64 and dim_mults (1, 2, 4, 8), enables flash attention, wraps it in GaussianDiffusion at image_size 128 and timesteps 1000, and computes a loss on random data. This does not train anything useful; it confirms the shapes line up on your machine.

python
import torch
from denoising_diffusion_pytorch import Unet, GaussianDiffusion

model = Unet(dim = 64, dim_mults = (1, 2, 4, 8), flash_attn = True)
diffusion = GaussianDiffusion(model, image_size = 128, timesteps = 1000)

training_images = torch.rand(8, 3, 128, 128)
loss = diffusion(training_images)
loss.backward()

For a real run, the Trainer class takes a folder path instead of a tensor. The README states that samples and model checkpoints are logged to ./results periodically, so the directory is created relative to wherever you launch the script.

python
from denoising_diffusion_pytorch import Unet, GaussianDiffusion, Trainer

model = Unet(dim = 64, dim_mults = (1, 2, 4, 8), flash_attn = True)
diffusion = GaussianDiffusion(model, image_size = 128, timesteps = 1000, sampling_timesteps = 250)

trainer = Trainer(
    diffusion,
    'path/to/your/images',
    train_batch_size = 32,
    train_lr = 8e-5,
    train_num_steps = 700000,
    gradient_accumulate_every = 2,
    ema_decay = 0.995,
    amp = True,
    calculate_fid = True
)

trainer.train()

The default train_num_steps of 700000 is not a suggestion for a quick experiment. Lower it deliberately for a first pass, and set calculate_fid to False if you do not want the FID computation running during training.

Multi-GPU training goes through Accelerate, not through the library

The Trainer is equipped with Hugging Face Accelerator, and the README describes multi-GPU as a two-step process using their CLI. At the project root, where the training script lives, you run accelerate config and answer its prompts, then launch the script through accelerate launch. There is no distributed flag on Trainer itself.

python
$ accelerate config
$ accelerate launch train.py

This is a reasonable division of labor: the library does not reimplement process launching, and you inherit whatever Accelerate supports. The cost is that the README does not document how batch size, gradient_accumulate_every and the number of processes interact. If you scale from one GPU to four, the effective batch size changes unless you adjust train_batch_size yourself. The README is silent on this, so treat it as something to work out from the Accelerate documentation rather than from this repository.

Where denoising-diffusion-pytorch stops short

The most explicit limitation is in the 1D section. The README states that Trainer1D does not evaluate the generated samples in any way, since the type of data is not known. It suggests adding a suitable metric to the training loop yourself after an editable install with pip install -e . That is an honest admission, and it means sequence work has no built-in quality signal at all.

The image path is better served, but only through FID. calculate_fid = True pulls in pytorch-fid, and FID on small datasets is a noisy number. The README does not discuss dataset size requirements, so if you are training on a few hundred images, expect the metric to be of limited use.

There is also a design note in the README that cuts against the project's own premise. An update reads: "Turns out none of the technicalities really matters at all," pointing to the Cold Diffusion paper and Muse. That is the author telling you the noise schedule and the exact forward process are less load-bearing than the original paper implies. It is unusual to see a library undercut its own subject matter, and it is worth reading before you spend a week tuning timesteps.

Finally, the package metadata classifies the project as 'Development Status :: 4 - Beta'. The last release listed is 2.2.6 on 2026-02-11, and the last push to main was on 2026-09-01. That is recent, but the README does not document a deprecation policy, a versioning contract, or a migration guide between minor versions. Pin your version.

How this differs from the Hugging Face Diffusers approach

Diffusers is the obvious alternative, and the difference is architectural rather than qualitative. Diffusers is a model zoo and pipeline framework: you load a pretrained checkpoint, pick a scheduler, and call a pipeline object. The scheduler is a swappable component, and the library's value is in the breadth of pretrained models and the standardization of the pipeline interface.

denoising-diffusion-pytorch inverts that. There is no checkpoint registry and no pipeline abstraction. You construct a Unet, you construct a GaussianDiffusion, and you train. The sampling method is selected by setting sampling_timesteps rather than by instantiating a different scheduler class. If your goal is to fine-tune an existing model or to run inference on a published checkpoint, Diffusers is the shorter path. If your goal is to understand or modify the training objective itself, this repository puts fewer layers between you and the loss.

The README also links a Flax implementation by YiYi Xu and an annotated walkthrough from Hugging Face research staff. Those are reading material, not drop-in replacements, but they are the natural next stops if the PyTorch code alone is not enough to follow the math.

Licence and the real cost of upgrading

The project is MIT licensed, declared both in the repository LICENSE file and in pyproject.toml under license = "MIT". The build backend is hatchling, and the version is read dynamically from denoising_diffusion_pytorch/version.py via tool.hatch.version. That is a permissive licence with no copyleft obligation, and it is compatible with commercial use. This is not legal advice; check the LICENSE file for the exact terms.

The upgrade cost is not in the licence, it is in the dependency surface. The declared dependencies include torch>=2.0, accelerate, einops, ema-pytorch>=0.4.2, pytorch-fid, scipy, numpy, pillow, torchvision and tqdm. A torch major-version bump will touch more of your environment than any change in this library. The project's own version moved from 2.2.4 to 2.2.5 to 2.2.6 across 2025-07-30, 2025-08-04 and 2026-02-11, which is a patch-level cadence with no documented breaking-change policy. Pin denoising_diffusion_pytorch in your requirements and upgrade deliberately.

The training compute is the larger line item. The README's default train_num_steps is 700000 with a batch size of 32, which is not a number you arrive at by accident. Budget for it before you start.

Editorial conclusion

Adopt it if you want to read, modify and train a diffusion model on your own image folder without pulling in a full training framework, or if you need a 1D variant for sequences. Do not adopt it if you need a production inference service, a hosted checkpoint, or a maintained evaluation pipeline for non-image data. Before committing, verify the torch>=2.0 and accelerate requirements in pyproject.toml against your environment, check whether calculate_fid is worth the pytorch-fid dependency for your dataset size, and read the Trainer class to confirm what it logs to ./results. The MIT license is the least of your concerns here; the training compute is.

Frequently asked questions

What are denoising diffusion models?

The README describes them as a generative modeling approach that uses denoising score matching to estimate the gradient of the data distribution, followed by Langevin sampling to sample from the true distribution. It cites the Denoising Diffusion Probabilistic Model paper as the reference.

How does DDPM work?

In this implementation the mechanism is a Unet that predicts noise and a GaussianDiffusion object that holds the image size and timestep count and produces a training loss. Sampling is a separate call, and setting sampling_timesteps switches to DDIM for faster inference.

Can you explain the diffusion model in a simple way?

The README's own framing is that the model learns to denoise, with score matching estimating the data distribution's gradient and Langevin sampling drawing from it. The minimal example is five lines: build a Unet, wrap it, pass normalized images, take the loss, call backward.

What is the difference between DDPMs and GANs?

The README states that diffusion is a new approach to generative modeling that may have the potential to rival GANs, and links to an external post on score matching with Langevin sampling. It does not give a technical comparison beyond that framing.

Official sources

  1. Issues
  2. License: MIT
  3. lucidrains/denoising-diffusion-pytorch on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lucidrains-denoising-diffusion-pytorch.svg)](https://hysenlabs.com/projects/lucidrains-denoising-diffusion-pytorch)