# LLaDA: an 8B masked diffusion language model you can load with transformers

> LLaDA is the official PyTorch implementation of Large Language Diffusion with mAsking, an 8B text model trained from scratch as a diffusion model rather than an autoregressive one. It ships inference, chat and evaluation code, but not the training framework.

**ML-GSAI/LLaDA** — Official PyTorch implementation for "Large Language Diffusion Models"

- Repository: https://github.com/ML-GSAI/LLaDA
- Stars: 3,994 · Forks: 290
- Language: Python
- License: not declared
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/ml-gsai-llada

## What LLaDA replaces, and for whom

Almost every open text model you can download generates left to right, one token per forward pass. LLaDA does not. It is a masked diffusion model: the model is trained to recover text from a randomly masked sequence, and generation works by starting from a fully masked answer and filling it in over several steps. The README frames the goal narrowly: "Our goal is to explore a theoretically complete language modeling approach, masked diffusion models." The repository exists so other people can run that model, evaluate it, and fine-tune on top of it.

The audience is therefore narrower than the download count suggests. If you want a drop-in chat model for a product, LLaDA is a research artifact with an inference path, not a serving stack. If you want to study non-autoregressive generation, compare it against a transformer of similar size, or fine-tune a diffusion language model on your own data, the repository gives you the pieces: generation code, likelihood code, a chat script, and evaluation harnesses wired to lm-evaluation-harness and OpenCompass. The README also points to LLaDA-V for vision-language work, LLaDA 1.5 for preference alignment, LLaDA-MoE for a mixture-of-experts variant, and iLLaDA for an improved generation-efficiency line.

## How masked diffusion generation actually works

The mechanism visible in the repository is conditional generation through iterative unmasking. You give the model a prompt, and the answer region begins as mask tokens. At each step the model predicts tokens for the masked positions, a portion of those predictions is committed, and the rest stay masked for the next round. The number of steps is a runtime parameter of the generate() function in generate.py, not a property of the weights, which is why the same checkpoint can produce different quality and latency trade-offs.

That is the architectural difference from a causal decoder. There is no KV cache to grow token by token, and no single pass that produces the whole reply; the compute is spread across refinement passes over the full sequence. The README contrasts this with BERT directly: LLaDA uses a masking ratio that varies randomly between 0 and 1, while BERT uses a fixed ratio, and the README states that LLaDA's training objective is an upper bound on the negative log-likelihood of the model distribution. That claim is what the authors say makes it a generative model capable of in-context learning and instruction following rather than an encoder trained for representation. Whether you accept the theory, the practical consequence is that generation cost scales with the number of unmasking steps you choose.

The repository also exposes get_log_likelihood() in get_log_likelihood.py for conditional likelihood evaluation, which is the natural companion to a diffusion objective: you can score a completion without generating it.

## Installing LLaDA and running a first generation

The README pins one dependency explicitly: install transformers==4.38.2 before loading the model. The checkpoints live on Hugging Face as GSAI-ML/LLaDA-8B-Base and GSAI-ML/LLaDA-8B-Instruct, and the loading example in the README uses trust_remote_code=True, which means the repository's own modeling code runs rather than a stock transformers class.

```bash
pip install transformers==4.38.2
```

The README's loading snippet, reproduced as given:

```python
from transformers import AutoModel, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained('GSAI-ML/LLaDA-8B-Base', trust_remote_code=True)
model = AutoModel.from_pretrained('GSAI-ML/LLaDA-8B-Base', trust_remote_code=True, torch_dtype=torch.bfloat16)
```

Note that the snippet references torch.bfloat16 without importing torch; add the import and move the model to a GPU before generating. The README does not spell out the VRAM requirement for an 8B model in bfloat16, so treat that as something to check against your own hardware.

For a first real use, the repository's own entry points are the fastest route. Running the chat script gives multi-round conversation against the instruct checkpoint:

```bash
python chat.py
```

For a browser interface, install Gradio and run the demo app:

```bash
pip install gradio
python app.py
```

If you would rather call the pieces yourself, generate.py provides generate() for conditional generation and get_log_likelihood.py provides get_log_likelihood() for scoring. The README describes these as the two supported operations but does not document their argument lists, so read the source before wiring them into anything.

## The training framework is deliberately absent

This is the limitation that matters most, and the README states it plainly: "We will not provide the training framework and data as most open-source LLMs do." You get inference, chat, and evaluation. You do not get the pipeline that produced the weights, nor the data.

What you get instead is GUIDELINES.md, which the README says contains guidelines for pre-training and supervised fine-tuning, plus a pointer to SMDM, a separate repository with a similar training process and an open-sourced framework. The README's own claim is that adapting an autoregressive training codebase takes "just a few lines of code." That may be true for the loss function, but a training run also needs the masking schedule, the data pipeline, and the evaluation loop, none of which ship here. If your plan is to reproduce the model, budget for reading the paper and GUIDELINES.md carefully and for borrowing heavily from SMDM.

A second boundary is the TODO list. vLLM integration is listed as unchecked. There is no served inference path in this repository, so throughput-oriented deployments have nothing to build on here yet. Batch inference support was added, according to the news entries, alongside evaluation code for the base, instruct and 1.5 models, but batching is not the same as a production server.

## Evaluation tooling and the iLLaDA drop-in

Evaluation is where this repository is more generous than most research releases. EVAL.md holds the instructions, and the top level contains eval_llada.py, eval_illada.py, eval_reverse.py and shell wrappers for lm-evaluation-harness and OpenCompass, with an opencompass/ directory alongside them. The news entries date the lm-evaluation-harness support to 2025.05.04 and the batch inference plus full evaluation code to 2025.10.27.

The iLLaDA line is worth understanding before you pick a checkpoint. According to the news entry, iLLaDA improves on LLaDA in benchmark performance and generation efficiency, and by setting mask_id=5 it can reuse the existing LLaDA inference code. That single configuration key is the whole migration path: point the loader at an iLLaDA checkpoint, set mask_id to 5, and the same generate() and chat scripts apply. If you are starting fresh, the README's own framing suggests iLLaDA is the newer option, while LLaDA-8B remains the checkpoint with the most evaluation code written against it.

A practical caution on the numbers: the README compares LLaDA to LLaMA3 8B and reports that LLaDA-MoE-7B-A1B-Instruct is comparable to Qwen2.5-3B-Instruct. Those are the authors' benchmark claims, run with the harnesses in this repository. Reproduce them with EVAL.md before you rely on them for a decision.

## Where an autoregressive model is still the better answer

The honest alternative is not another diffusion model; it is a standard causal decoder of similar size, such as LLaMA3 8B, which the README itself uses as the comparison point. The difference in approach is fundamental. An autoregressive model commits to each token as it emits it, so the output is stable and cacheable, and the ecosystem around serving, quantization, and speculative decoding assumes that shape. LLaDA revises the whole answer across refinement steps, which buys a different generation dynamic and costs you the tooling built for left-to-right decoding.

If your workload is high-throughput chat behind an API, an autoregressive model with a mature serving engine will beat this repository on operational maturity today, because LLaDA has no vLLM path yet. If your workload needs exact token-level likelihoods for scoring, both families can do it, and LLaDA's get_log_likelihood() is the entry point here. The case for LLaDA is research and experimentation: studying masked diffusion at 8B scale, fine-tuning on top of it, or testing whether iterative refinement helps a task where left-to-right generation struggles.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-07-15, which is recent. The news entries show a steady cadence across the LLaDA, LLaDA-V, LLaDA 1.5, LLaDA-MoE and iLLaDA lines, and the TODO list still carries the unchecked vLLM item, so the project is moving but not finished.

Upgrade cost is concentrated in the pinned dependency. transformers==4.38.2 is old relative to current releases, and the loading path uses trust_remote_code=True, so the model code comes from the checkpoint repository rather than from your installed transformers. That means a transformers upgrade is not automatically a LLaDA upgrade, and a checkpoint-side change can alter behavior without a version bump you control. Pin both the library and the checkpoint revision if you care about reproducibility.

On licensing, no licence identifier is given for this repository, and the README does not state one. The checkpoints live on separate Hugging Face repositories with their own terms. Before commercial use, read the LICENSE file in the repository root and the model card on Hugging Face; this is a factual gap in the documentation, not a legal conclusion.

## Conclusion

Adopt LLaDA if you want to study or build on a masked diffusion language model at 8B scale and you are comfortable with transformers==4.38.2 and trust_remote_code=True. Do not adopt it if you need a supported training pipeline, a vLLM serving path, or a model whose generation loop you can leave untouched: the README states it will not provide the training framework and data, and vLLM integration is still an open TODO. Before committing, check the licence file in the repository, since the README does not state one, and run python chat.py once to see how many refinement steps your hardware can afford per reply.

## FAQ

### What does LLaDA stand for?

The README expands it as Large Language Diffusion with mAsking, and the repository is the official PyTorch implementation for the paper "Large Language Diffusion Models."

### What is diffusion language in the context of LLaDA?

In LLaDA, generation starts from a masked sequence and fills it in over several refinement steps, with a masking ratio that varies randomly between 0 and 1 during training. The README contrasts this with BERT, which uses a fixed masking ratio.

### How do I install LLaDA and run it?

Install transformers==4.38.2, load GSAI-ML/LLaDA-8B-Base or GSAI-ML/LLaDA-8B-Instruct with trust_remote_code=True, then run python chat.py for conversation or python app.py for the Gradio demo.

### Does the LLaDA repository include training code?

No. The README states the training framework and data will not be provided, and points to GUIDELINES.md plus the separate SMDM repository for a similar open-sourced training process.

### Can I use iLLaDA with the existing LLaDA inference code?

Yes. According to the news entry, setting mask_id=5 lets iLLaDA directly reuse the existing LLaDA inference code.

### Does LLaDA support vLLM?

Not yet. Integrating the vLLM inference engine is listed as an unchecked item in the repository's TODO list.

## Sources

- [Issues](https://github.com/ML-GSAI/LLaDA/issues)
- [ML-GSAI/LLaDA on GitHub](https://github.com/ML-GSAI/LLaDA)
- [README](https://github.com/ML-GSAI/LLaDA/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ml-gsai-llada
