Fast-dLLM: KV Cache and Parallel Decoding for Diffusion LLMs
Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"
At a glance
- What is it?
- NVlabs' Fast-dLLM repository collects four acceleration techniques for diffusion language models, from a training-free v1 path to a block-diffusion v2 with a released 7B checkpoint. The v1 route is the one you can adopt without retraining anything.
- Who is it for?
- Adopt v1 if you already run Dream or LLaDA and want faster sampling without touching weights, and adopt v2 only if you can fine-tune and are willing to take the released Fast_dLLM_v2_7B checkpoint as your starting point. Skip the repository entirely if you need vLLM serving today, since the TODO list still shows vLLM support as pending.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 123 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Fast-dLLM is for, and who should care
Diffusion language models generate text by iteratively denoising a whole token block rather than emitting tokens left to right. That makes them parallel in principle, but the naive sampling loop recomputes attention over the full sequence at every denoising step, and it commits to tokens one position at a time. Fast-dLLM exists to attack both of those costs. The repository is a family rather than a single tool: v1 covers text-only training-free acceleration, v2 covers block diffusion with fine-tuning, fast_dvlm covers vision plus text, and fast_ddrive covers vision, text and driving actions.
The audience is narrow and technical. You need to be running a diffusion LLM such as Dream or LLaDA, or willing to fine-tune Qwen2.5 into a block-diffusion model. If you are serving a standard autoregressive model, nothing here applies to you. The v1 README frames the technique as training-free, which is the property that matters most for adoption: you point the existing evaluation scripts at a pretrained checkpoint and change sampling flags, rather than producing a new model.
How v1 gets speed without retraining
Two mechanisms sit at the centre of v1, and the paper title names them: KV cache and parallel decoding. A diffusion LM's denoising loop revisits the same prefix many times, so caching the key and value projections of tokens whose values have already been fixed avoids recomputing them on later steps. The README exposes this as a boolean flag, use_cache=true, inside the model_args string of the Dream evaluation command, which tells you the cache is a runtime switch rather than an architectural change.
Parallel decoding is the second half. Instead of decoding one position per step, a confidence-threshold strategy decodes every position whose model confidence clears a threshold. In the Dream example the threshold is 0.9 and the algorithm is named confidence_threshold. Raise the threshold and fewer positions qualify per step, so you trade steps for accuracy; lower it and you commit to more tokens that the model is less sure about. The repository also mentions a factor-based parallel strategy added to v1/llada/eval_gsm8k.sh in July 2025, which is a second way to decide how many tokens move forward at once.
The v2 branch changes the premise. Rather than accelerating an existing model, it fine-tunes Qwen2.5 into a block-diffusion model with hierarchical caching, and the released Fast_dLLM_v2_7B checkpoint is the product of that training. The README's comparison table lists block diffusion plus hierarchical caching as v2's key techniques, which means v2's speed is partly bought with a training run and a different weight set.
Installing v1 and running a first LLaDA chat
The Quick Start section keeps installation minimal: change into v1 and install its requirements file. There is no package published to PyPI for v1, so this is a clone-and-install workflow.
cd v1
pip install -r requirements.txtWith dependencies in place, the README's first real use is an interactive LLaDA chat. The three flags control generation length, the number of denoising steps, and the block size the sampler works over.
python llada/chat.py --gen_length 128 --steps 128 --block_size 32Expect a prompt loop in the terminal. The block_size of 32 with gen_length 128 means the sequence is handled in four blocks, and steps 128 is the total denoising budget. If you want to see the acceleration flags in action rather than chat, the Dream evaluation command is the better entry point because it passes use_cache and threshold explicitly:
accelerate launch dream/eval.py --model dream \
--model_args pretrained=Dream-org/Dream-v0-Base-7B,max_new_tokens=256,diffusion_steps=8,add_bos_token=true,alg=confidence_threshold,threshold=0.9,use_cache=true \
--tasks gsm8k --num_fewshot 5 --batch_size 1Note that diffusion_steps is 8 here while max_new_tokens is 256. That ratio is the whole point of parallel decoding: fewer denoising steps than tokens produced. The command runs the gsm8k task with 5 few-shot examples and a batch size of 1.
Installing v2 and its Gradio demo
v2 installs differently. It is an editable install from inside the v2 directory, which reflects that the directory contains a training framework (LMFlow, under src/) rather than a loose script collection.
cd v2
pip install -e .The README's first use for v2 is a Gradio web demo started with a single command:
python app.pyGradio prints a local URL when it starts, and that is where the chat interface lives. Evaluation is a shell script rather than a direct Python invocation:
bash eval_script.shThe repository layout also lists run_chatbot.py alongside app.py, configs/ for DeepSpeed configurations, and train_scripts/ for fine-tuning. The presence of DeepSpeed configs is a fair signal about the hardware expectation: v2 training is not a laptop exercise, and the released 7B checkpoint exists precisely so that most users can skip that step.
Where Fast-dLLM stops being the right tool
The clearest limitation is stated in the project's own TODO list: vLLM support is still marked as pending. If your deployment plan assumes a vLLM server with paged attention and continuous batching, Fast-dLLM does not give you that today, and no release notes in the repository announce it. You would be running the provided scripts, not a serving stack.
The second constraint is version coupling. v1's acceleration is implemented against specific backbones, Dream and LLaDA, and the flags live inside those evaluation scripts. A diffusion model from a different lineage is not covered by the code paths described in the README. The third is that v2 is not training-free at all. If you cannot run fine-tuning, or do not want to adopt the Fast_dLLM_v2_7B weights, the v2 branch is out of reach and only v1 is relevant to you.
There is also a maturity asymmetry worth naming. The repository spans four research lines with separate READMEs, separate requirements files and separate entry points. That is good for reproducing papers and awkward for anyone hoping for one coherent interface.
How Fast-dLLM differs from an autoregressive serving stack
The natural comparison is not another diffusion accelerator but the autoregressive pipeline people already run. With an AR model on vLLM or SGLang you get a mature server, but you inherit strictly sequential decoding: token N cannot be produced before token N-1. Fast-dLLM's parallel decoding breaks that dependency by letting several positions advance in the same denoising step when confidence allows. The trade is that you give up the serving ecosystem and accept a confidence threshold as a new tuning knob that has no equivalent in AR inference.
The README itself makes this comparison concrete for the driving variant, describing Fast-dDrive as reaching over 200 TPS on a single H100 and up to 12x over the AR baseline with SGLang. That framing, AR baseline versus diffusion acceleration, is the axis the project thinks in. If your workload is latency-sensitive generation on a diffusion backbone, the comparison is meaningful. If your workload is already served well by an AR model, the comparison is irrelevant because the underlying model class is different.
Licence, maintenance and upgrade cost
The repository is licensed Apache-2.0, which permits commercial use and modification provided the licence and notices are preserved. That is the licence text in the repository root; it says nothing about the pretrained checkpoints hosted on Hugging Face, which carry their own terms you would need to check separately. Nothing here is legal advice.
The last push to the default branch was on 2026-05-30, and the repository is not archived. The news entries show a steady cadence through 2026: v1 and v2 accepted at ICLR 2026 in January, Fast-dVLM in April, Fast-dDrive in May. Upgrades are not automatic. v1 and v2 have independent requirements.txt files, so bumping one does not bump the other, and v2's editable install means a pull can change the installed package in place. Pinning your dependency versions before a pull is the cheap insurance.
Editorial conclusion
Adopt v1 if you already run Dream or LLaDA and want faster sampling without touching weights, and adopt v2 only if you can fine-tune and are willing to take the released Fast_dLLM_v2_7B checkpoint as your starting point. Skip the repository entirely if you need vLLM serving today, since the TODO list still shows vLLM support as pending. Verify first that your installed transformers and accelerate versions match v1/requirements.txt, and that the Dream-org/Dream-v0-Base-7B checkpoint identifier in the evaluation command resolves on the Hub.
Frequently asked questions
What is a dLLM?
A dLLM is a diffusion-based large language model. Instead of generating tokens left to right, it denoises a block of tokens over several steps, which is what allows techniques like parallel decoding and KV caching to be applied to it.
Does Fast-dLLM v1 require any training?
No. The README describes v1 as training-free inference acceleration, and the Quick Start shows it being applied to existing Dream and LLaDA checkpoints by passing flags such as use_cache=true and threshold=0.9 to the evaluation script.
What is the difference between Fast-dLLM v1 and v2?
v1 is training-free acceleration of an existing diffusion LLM using KV cache and parallel decoding. v2 uses block diffusion with hierarchical caching and requires fine-tuning, with a released Fast_dLLM_v2_7B checkpoint on Hugging Face.
Can Fast-dLLM be served with vLLM?
Not yet according to the repository. The TODO list in the README still shows vLLM support as pending, so the documented entry points are the provided Python scripts and shell scripts rather than a vLLM server.
Which models does Fast-dLLM support?
The comparison table lists Dream and LLaDA as the v1 backbones, Qwen2.5 for v2, and Qwen2.5-VL for Fast-dVLM and Fast-dDrive. The acceleration code is written against those specific model families.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvlabs-fast-dllm)