# dLLM: a training, inference and evaluation stack for diffusion language models

> dLLM wraps LLaDA, Dream, BERT-Chat and Edit Flows recipes behind a transformers Trainer and lm-evaluation-harness, so you can finetune a masked diffusion model without writing your own sampler. It is research tooling, not a serving framework.

**ZHZisZZ/dllm** — dLLM: Simple Diffusion Language Modeling

- Repository: https://github.com/ZHZisZZ/dllm
- Website: https://arxiv.org/pdf/2602.22661
- Stars: 2,700 · Forks: 281
- Language: Python
- License: Apache-2.0
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/zhziszz-dllm

## What dLLM is for, and who it is not for

Diffusion language models generate text by iteratively denoising a sequence rather than emitting one token at a time. That changes almost everything downstream of the model: the sampler, the loss, the way you score a checkpoint, the way you cache during inference. dLLM exists to keep those pieces in one repository. The README describes it as a library that unifies the training and evaluation of diffusion language models, and the examples directory backs that up: examples/llada, examples/dream, examples/a2d, examples/bert, examples/editflow, examples/fastdllm, examples/rl, plus inference-only directories for LLaDA2.0 and LLaDA2.1.

The intended reader is someone who already knows what a masked diffusion model is and wants a working recipe for one. If you are looking for an OpenAI-compatible endpoint to put behind a chat product, this is the wrong repository. There is no server, no batching scheduler, no quantized runtime in the core package. The README points at Fast-dLLM for accelerated inference and evaluation, and the optional dependency group lists vllm and flash-attn, but the project itself is organized around training runs and benchmark scores, not request throughput.

## The three layers: Trainer, harness, and per-model recipes

The architecture is deliberately shallow. dllm/ holds the shared code, examples/ holds one directory per model family, and lm-evaluation-harness sits at the repository root as a submodule.

The training layer builds on the transformers Trainer. The README states that dLLM provides scalable training pipelines based on Trainer with support for LoRA, DeepSpeed and FSDP. That is a real constraint as much as a feature: your training loop inherits Trainer's callbacks, its checkpoint format, its argument parsing, and its assumptions about what a forward pass returns. A diffusion objective does not naturally fit the next-token prediction contract, so the per-model pipeline code under dllm/pipelines/ is where the adaptation lives. The pyproject.toml black configuration excludes dllm/pipelines/llada/models, dllm/pipelines/llada2/models, dllm/pipelines/llada21/models, dllm/pipelines/dream/models and dllm/pipelines/a2d/models from formatting, which tells you those directories are treated as vendored or lightly adapted model code rather than as first-class library modules.

The evaluation layer wraps lm-evaluation-harness. The README says the pipelines abstract away inference details so that customization is simple. In practice that means a diffusion sampler has to be expressed as something the harness can call, which is the part most people underestimate when they try to add a new benchmark.

The recipe layer is where the project's actual knowledge sits. examples/llada covers pretraining, finetuning and evaluating LLaDA and LLaDA-MoE. examples/dream does the same for Dream. examples/a2d finetunes any autoregressive model into a masked or block diffusion generator. examples/bert turns a BERT encoder into a lightweight chatbot. examples/editflow is described as an educational reference for extending existing DLLMs with insertion, deletion and substitution operations. examples/rl covers GRPO training for LLaDA and Tiny-A2D across GSM8K, MATH, Countdown, Sudoku and Code.

## Installing dLLM and running a first finetune

The README's setup section starts with a conda environment on Python 3.10 and installs a CUDA 12.4 toolchain, then PyTorch. The README notes that other PyTorch and CUDA versions should also work, so the pinned versions below are the documented path rather than a hard requirement.

```bash
conda create -n dllm python=3.10 -y
conda activate dllm
conda install cuda=12.4 -c nvidia
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 \
    --index-url https://download.pytorch.org/whl/cu124
```

The package itself is declared in pyproject.toml with name dllm and version 0.1.0, requiring Python 3.10 or newer. Its dependencies are pinned tightly: transformers==4.57.0, accelerate==1.11.0, deepspeed==0.18.0, peft==0.17.1, datasets==4.2.0 and sentencepiece==0.2.0, alongside torchmetrics, tyro, wandb, omegaconf, tqdm, matplotlib, pytest and rich. There is an optional extra group containing bitsandbytes==0.48.1, vllm==0.8.5.post1 and flash-attn==2.8.3.

Those pins are the first thing to check before you invest time. transformers==4.57.0 is an exact pin, not a floor, so installing dLLM into an environment that already holds a different transformers release will either downgrade it or fail to resolve. The same applies to peft and deepspeed.

After installation, the top-level README does not give a single canonical training command. It defers to the per-example READMEs. The llada example directory is the place to look for pretraining, finetuning and evaluation instructions, and the a2d directory is the place to look if your starting point is an existing autoregressive checkpoint such as Qwen, LLaMA or GPT-2. Read the README inside that directory before running anything, because the flags live there and not at the repository root.

One structural detail matters at install time: lm-evaluation-harness appears in the repository listing and the black exclude list treats it as a directory to leave alone. If you clone without submodules, the evaluation pipelines will not have their dependency present.

## Where dLLM gets in your way

The dependency pins are aggressive. An exact transformers==4.57.0 pin means dLLM competes with every other library in your environment that wants a different transformers version. For a research repository that is a defensible choice, since model code under dllm/pipelines/ is adapted per release and a minor bump can change attention implementations. For a team trying to add diffusion training to an existing stack, it is a real coordination cost.

The training abstraction is Trainer-shaped. Anything that does not fit the Trainer contract, such as a custom sampling schedule tied to a curriculum, has to be expressed through the pipeline code rather than through a clean extension point. The formatting exclusions in pyproject.toml suggest those files are not maintained to the same standard as the rest of the package.

Evaluation is only as good as the harness integration. The README claims the evaluation pipelines abstract away inference details, and that is the right goal, but a diffusion sampler is stateful in ways a standard autoregressive generate call is not. If you add a benchmark that assumes a single forward pass per example, expect to write adapter code.

The README does not document rollback, versioning policy, or how breaking changes to the pinned dependencies will be handled. There are no retrieved releases, so there is no changelog to consult. The repository was last pushed on 2026-07-17, which is two months before the date of writing; that is recent enough that the code is not stale, but there is no release cadence to plan upgrades against.

Finally, the README's own framing is worth taking literally. The commented-out note in the source describes the repository as primarily for educational purposes and explicitly disclaims 100 percent exact reproduction of official models. That note is not rendered in the README as published, but the disclaimer is consistent with the example structure, where recipes are demonstrations of a training algorithm rather than drop-in replacements for vendor checkpoints.

## dLLM against vendor-specific diffusion repositories

The natural alternative is a model-specific repository. If all you want is to run LLaDA, you can use the LLaDA authors' own code; if all you want is Dream, use the Dream authors' code. Those repositories give you the exact checkpoint the paper evaluated and the exact sampler the paper used, which is what you want for reproduction.

dLLM takes the opposite approach. It puts LLaDA, LLaDA-MoE, Dream, BERT-Chat, Edit Flows and the a2d conversion recipes behind shared training and evaluation plumbing. The difference shows up the moment you want to compare two of them. In a vendor repository you would write your own evaluation loop and your own data loading to get a comparable number. Here, the README states that evaluation runs through lm-evaluation-harness, so two model families evaluated through the same harness produce numbers that are at least produced the same way.

The cost of that uniformity is fidelity. A shared Trainer-based pipeline cannot reproduce every quirk of every paper's training script. If your goal is to match a published number to the decimal, the vendor repository is the better starting point. If your goal is to finetune a diffusion model on your own data and know whether it got better, dLLM's structure saves you from rebuilding the evaluation side.

A second comparison point is the a2d and bert examples. Converting an existing autoregressive model or a BERT encoder into a diffusion model is not something a vendor repository for LLaDA or Dream offers, because those projects start from their own architectures. If that conversion is your actual task, dLLM is one of the few places the README points to a recipe for it.

## Licence and upgrade cost

dLLM is Apache-2.0. That is a permissive licence with an explicit patent grant, and it is compatible with commercial use, modification and redistribution provided you keep the notices and state changes. It does not give legal advice and it does not settle the licence of anything you load into it.

That last point is the one to think about. The library is Apache-2.0, but the checkpoints it loads are not covered by that licence. LLaDA, Dream, Tiny-A2D, BERT-Chat and the LLaDA2.x models live on the dllm-hub Hugging Face organization, and each carries its own terms. The same applies to any base model you convert through examples/a2d, such as Qwen, LLaMA or GPT-2. Your obligations come from those model licences, not from dLLM's.

Upgrade cost is dominated by the dependency pins. Because transformers, peft, accelerate and deepspeed are pinned to exact versions, moving dLLM forward means moving all of them together, and the model code under dllm/pipelines/ is excluded from formatting precisely because it is adapted from upstream releases. Budget for a re-test of your training run after any bump. There is no retrieved release history, so you cannot judge how often those bumps happen; the last push to the default branch was on 2026-07-17.

## Conclusion

Adopt dLLM if you are researching masked or block diffusion training and want the LLaDA, Dream, BERT-Chat and Edit Flows recipes in one place with the evaluation harness already wired in. Do not adopt it as a production text-generation server; the README points at Fast-dLLM for accelerated inference and the repository ships no serving layer. Before committing, check that the pinned transformers==4.57.0 and torch==2.6.0 versions resolve in your environment, confirm the lm-evaluation-harness submodule is checked out, and read the notes in the specific examples/ directory you plan to use, since the top-level README defers to those files for training and inference instructions.

## FAQ

### What is a dLLM?

In this repository the name refers to the dLLM library, which the README describes as a library that unifies the training and evaluation of diffusion language models. It is not a model itself; it is the training, inference and evaluation code around models such as LLaDA and Dream.

### What is a diffusion language model?

The README groups the supported algorithms under masked diffusion, block diffusion and edit flows, and the examples cover models such as LLaDA, Dream, BERT-Chat and Edit Flows. In dLLM these are trained through a transformers Trainer-based pipeline and evaluated through lm-evaluation-harness.

### Can you give me an example of a diffusion model?

The README names LLaDA and Dream as open-weight diffusion models with minimal training, inference and evaluation recipes in examples/llada and examples/dream. It also points to Tiny-A2D, a collection of small 0.5B and 0.6B diffusion models adapted from autoregressive models.

## Sources

- [Issues](https://github.com/ZHZisZZ/dllm/issues)
- [License: Apache-2.0](https://github.com/ZHZisZZ/dllm/blob/main/LICENSE)
- [Project website](https://arxiv.org/pdf/2602.22661)
- [README](https://github.com/ZHZisZZ/dllm/blob/main/README.md)
- [ZHZisZZ/dllm on GitHub](https://github.com/ZHZisZZ/dllm)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zhziszz-dllm
