Model or dataset
ML-GSAI/LLaDA avatar
ML-GSAI/LLaDA

LLaDA: an 8B diffusion language model with no training code

Official PyTorch implementation for "Large Language Diffusion Models"

3,994 stars290 forksPythonLicense varies

At a glance

What is it?
The official PyTorch implementation for masked diffusion language models, published as flat scripts with a transformers pin from 2024 and no packaging manifest. Useful for inference and evaluation; not for reproducing the model.
Who is it for?
LLaDA is worth your time if you want to see what a masked diffusion language model actually does at inference, or you need a second opinion on an evaluation suite, since the repository carries lm-evaluation-harness integration plus separate OpenCompass and likelihood scripts. It is not a research starting point, because the training framework and data are deliberately withheld and only a guidelines document is offered.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 83 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 6, 2026, and from our analysis. They are not legal advice.

Editorial analysis

LLaDA is Large Language Diffusion with mAsking

The name is a mnemonic for the mechanism. L comes from Large, La from Language, D from Diffusion, and the final A from mAsking, with the letter highlighted in the README so you can see where each piece comes from. That last letter is the whole design: instead of appending tokens one at a time, the model starts from a sequence that is mostly mask tokens and denoises it, the same shape as a masked image generator applied to text. The project describes itself as a diffusion model at an 8B scale, trained entirely from scratch, and states that it rivals LLaMA3 8B in performance, which is the authors' claim rather than an independent result. The paper is on arXiv at 2502.09992 and the weights are on Hugging Face as LLaDA-8B-Base and LLaDA-8B-Instruct.

transformers==4.38.2 is the load-bearing pin

Loading a model here is two objects and one very old dependency. The instruction is to install transformers==4.38.2 first, then load the base checkpoint through the standard auto classes.

python
from transformers import AutoModel, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained('GSAI-ML/LLaDA-8B-Base', trust_remote_code=True)
model = AutoModel.from_pretrained('GSAI-ML/LLaDA-8B-Base', trust_remote_code=True, torch_dtype=torch.bfloat16)

The exact pin matters more than it looks. An exact version from 2024 means this model definition predates the transformers releases that added most later architectures, so a current transformers will not necessarily have what it needs, and conversely a newer one may break the remote code. In a shared environment, put this in its own virtualenv. bfloat16 is the stated dtype, which is a reasonable minimum for an 8B model and rules out running it on hardware without bf16 support.

trust_remote_code means the checkpoint ships Python that runs on your machine

Both load calls pass trust_remote_code=True, and that flag is the security decision in this repository. It tells transformers to fetch the modelling code out of the Hugging Face repository and execute it locally rather than using a built-in architecture. For a diffusion language model that is not in transformers, there is not much alternative, but it means the trust boundary is the model repository rather than your pinned library version. Anyone who can publish to that repository can change what runs when you call from_pretrained. Read the modelling file before the weights if the checkpoint is not one you have evaluated, and keep bfloat16 and the exact pin in mind when you decide whether the environment is one you want to touch. The same reasoning applies to the Gradio demo and to chat.py, since both load the model the same way.

Sampling is slower than autoregressive, and the project explains why

The FAQ is unusually candid here, listing three reasons rather than hedging. LLaDA samples with a fixed context length, so the work is done for the whole window instead of for the tokens emitted so far. It cannot yet use techniques like KV-Cache, which is the mechanism that makes autoregressive decoding cheap on repeated context. And it reaches optimal performance when the number of sampling steps equals the response length, which means cutting steps to go faster directly costs quality. That last point is the one to sit with: there is no separate fast mode, only a quality dial. On top of this, batch inference support only arrived on 27 October 2025, so the repository was single-item oriented for most of its life, which is the other reason throughput is not a selling point.

The vLLM integration is still an unchecked item

The TODO list has two entries and only one is finished. Releasing the evaluation code for iLLaDA is marked done; integrating the vLLM inference engine is not. That single unchecked box tells you more about the project's serving story than any benchmark number. vLLM is the usual route to putting a large model behind a throughput-oriented server, and its kernels are built around autoregressive decoding with a paged key-value cache, both of which the FAQ already names as unavailable here. So the honest position is that LLaDA has no high-throughput serving path in this repository today, and anyone running it in anger is doing batch inference on a research stack. The iLLaDA release, dated 24 June 2026, improved on both benchmark performance and generation efficiency, so that is where to watch.

The training framework and the data are deliberately absent

This is the section most people skim and it is the one that decides whether a project is useful to you. The README states plainly that the training framework and data will not be provided, as is the case with most open-source LLMs. What is offered instead is GUIDELINES.md for pre-training and supervised fine-tuning, with the claim that adapting an existing autoregressive training codebase takes a few lines of changes. That claim is worth believing in principle, because the architecture is a Transformer, but it is untested by you until you do it. For a working example, the SMDM repository follows the same training process and has open-sourced its training framework, which makes it the more practical starting point for fine-tuning work. The FAQ points at both.

Evaluation arrives through three separate harnesses

Rather than one entry point, the repository carries three evaluation paths, and knowing which is which saves time. Evaluation code based on EleutherAI's lm-evaluation-harness was added on 4 May 2025 for the base model, and batch inference support plus the full evaluation code arrived on 27 October 2025 for the base model, the instruct model and LLaDA 1.5. Alongside that are shell scripts for OpenCompass, a separate OpenCompass configuration directory, a conditional likelihood script in get_log_likelihood.py and its eval_reverse.py counterpart, and an eval_instruct directory. The naming is the map: eval_llada_lm_eval.sh is the harness, eval_llada_opencompass.sh is OpenCompass, and EVAL.md is the document that says which to run. Two entry points for the same models producing different numbers is a normal state of affairs, so treat any single figure with care.

The repository is a flat directory of scripts with no manifest

There is no package directory, no setup.py, no pyproject.toml, no requirements.txt and no licence file at the top level, and the licence field is unset. What there is: README.md, EVAL.md, GUIDELINES.md, then the scripts themselves, app.py for the Gradio demo, chat.py for multi-turn conversation with the instruct model, generate.py holding generate(), get_log_likelihood.py holding get_log_likelihood(), eval_llada.py, eval_illada.py, eval_reverse.py, the three shell drivers, and directories for data, images, visualization and the OpenCompass config. In practice that means clone the repository, read the dependency list out of the README by hand, and install transformers, then whatever the script you want imports. For a research reference this is fine and keeps the surface small. For anything you intend to depend on, pin the commit and keep your own environment description, because there is nothing here that captures it for you.

Editorial conclusion

LLaDA is worth your time if you want to see what a masked diffusion language model actually does at inference, or you need a second opinion on an evaluation suite, since the repository carries lm-evaluation-harness integration plus separate OpenCompass and likelihood scripts. It is not a research starting point, because the training framework and data are deliberately withheld and only a guidelines document is offered. Three things to verify before you commit. The transformers pin is 4.38.2, which is old enough to conflict with other work in the same environment, so use a separate virtualenv. Loading uses trust_remote_code, which executes Python from the model repository on your machine, so read what you are trusting. And the vLLM integration is still an unchecked item on the project's own list, which tells you where the serving story stands.

Frequently asked questions

What does LLaDA stand for?

Large Language Diffusion with mAsking, with the letters of the acronym highlighted in the README so the mechanism shows through the name. It is an 8B diffusion language model trained entirely from scratch, published as LLaDA-8B-Base and LLaDA-8B-Instruct.

How is LLaDA different from BERT?

LLaDA uses a masking ratio that varies randomly between 0 and 1, while BERT uses a fixed ratio. The project says that makes LLaDA's training objective an upper bound on the negative log-likelihood of the model distribution, which is what makes it a generative model rather than an encoder.

How is LLaDA different from GPT?

The architecture is the same: LLaDA adopts the Transformer, like GPT. The difference is in probabilistic modelling, where GPT predicts the next token autoregressively and LLaDA uses a diffusion process over masked tokens.

Why is LLaDA slower than an autoregressive model?

The project gives three reasons: it samples with a fixed context length, it cannot yet use techniques like KV-Cache, and it performs optimally when the number of sampling steps equals the response length, so reducing steps reduces performance.

Does the LLaDA repository include the training code?

No. The training framework and data are deliberately not provided, as with most open-source LLMs. GUIDELINES.md covers pre-training and supervised fine-tuning, and the SMDM repository follows the same training process with its framework open-sourced.

How do I run the LLaDA chat demo?

Install transformers==4.38.2 first. Then python chat.py gives a multi-round conversation with LLaDA-8B-Instruct, while the Gradio demo needs pip install gradio followed by python app.py. Both load the checkpoint with trust_remote_code=True.

Official sources

  1. Issues
  2. ML-GSAI/LLaDA on GitHub
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ml-gsai-llada.svg)](https://hysenlabs.com/projects/ml-gsai-llada)