z-lab/dflash: a draft model that guesses whole blocks of tokens at once
DFlash: Block Diffusion for Flash Speculative Decoding
At a glance
- What is it?
- DFlash is a block diffusion draft model for speculative decoding, shipped as a Python CLI with Transformers, MLX and OpenAI-compatible server backends. It is a beta package whose usefulness depends entirely on which checkpoint you pair it with.
- Who is it for?
- Adopt dflash if you already serve one of the listed model families and want to test parallel drafting against your current setup, starting with the pip install and a single dflash generate call. Do not adopt it if your model is not in the checkpoint collections, if you need a stable API rather than a 0.1.0 beta, or if you cannot run a supported SGLang or vLLM build.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 44 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem DFlash targets: drafting tokens one at a time is slow
Speculative decoding works by having a small draft model propose tokens that a larger target model then verifies in parallel. The bottleneck is the drafting step itself. A conventional autoregressive draft model produces one token per forward pass, so the draft becomes the serial part of an otherwise parallel pipeline.
DFlash is a block diffusion model built specifically as that draft. The README describes it as "a lightweight block diffusion model designed for speculative decoding" that enables "efficient and high-quality parallel drafting". The distinction matters: an autoregressive draft and a block diffusion draft have different failure modes and different latency profiles, and the project exists to explore the second.
The audience is narrow and technical. You need a target model from the supported list, a machine that can hold both the target and the draft, and enough familiarity with speculative decoding to interpret acceptance rates. This is not a drop-in library for application developers who just want faster chat completions.
How block diffusion drafting fits between target and draft
The architecture is split across two generations in the README. DFlash 2 is presented first, with its own blog post and model collection, and DFlash follows with a paper link, a blog and a separate collection. Both are draft-side artefacts: you supply a target model and a draft checkpoint, and the CLI orchestrates them.
The repository layout is thin. The top level holds .github/, .gitignore, LICENSE, README.md, dflash/ and pyproject.toml, and the package entry point is dflash.cli:main. There is no separate inference engine inside the repository. For local use the optional local extra pulls in torch, transformers and jinja2 on Linux, or mlx, mlx-lm and huggingface-hub on Apple Silicon. For serving, the README points elsewhere entirely: you install a supported SGLang or vLLM build, launch its OpenAI-compatible server with DFlash, and pass its --base-url to the dflash CLI.
That split is the design. The project owns the draft checkpoints and a thin client, and delegates the heavy serving path to projects that already have optimised kernels. It also means the quality of your experience is largely determined by code you did not install from this repository.
Installing dflash and running a first generation
The package installs from PyPI. The base install is enough for talking to an OpenAI-compatible server; the local extra adds the dependencies for running inference on your own machine, which the README says means MLX on Apple Silicon and Transformers on Linux.
pip install dflash
pip install "dflash[local]" # local inferenceThe CLI exposes a generate subcommand with one subcommand per backend. The Transformers example in the README uses Muse-Glimmer-30B as the target and the matching DFlash 2 draft checkpoint, with a reasoning strength setting:
dflash generate transformers \
--model meta-models/Muse-Glimmer-30B \
--draft z-lab/Muse-Glimmer-30B-DFlash2 \
--reasoning high --temperature 1 --top-p 0.95 --top-k 64 \
"How many positive whole-number divisors does 196 have?"On Apple Silicon the MLX backend takes a different reasoning flag, reasoning_effort, and separate quantisation controls for the draft. The README warns that quantized targets or drafts should use block_size <= 5 because MLX's quantized matmul kernel becomes less efficient at larger verify widths:
dflash generate mlx \
--model mlx-community/Qwen3.8-27B-4bit \
--draft z-lab/Qwen3.8-27B-DFlash2 \
--draft-bits 4 --block-size 5 --reasoning xhigh \
"How many positive whole-number divisors does 196 have?"If you already run SGLang or vLLM, skip the local extra. Launch the server separately and point the CLI at it, which is the shortest path to a working call:
dflash generate openai \
--base-url http://127.0.0.1:8000 --model Qwen/Qwen3.8-27B \
"How many positive whole-number divisors does 196 have?"The output is a completion printed to the terminal. Nothing in the README describes a JSON output mode or a streaming flag for the CLI.
Benchmarking is a first-class command, and its defaults matter
The same CLI carries a benchmark subcommand with the same three backend choices. The datasets are named in the README: gsm8k, math500, humaneval, mbpp and mt-bench, downloaded and cached by Hugging Face Datasets. That caching is a practical detail: the first run of a benchmark pulls data over the network, so an air-gapped machine needs the cache populated in advance.
The OpenAI-compatible path exposes concurrency and prompt count directly, which is how you would compare a DFlash draft against a stock server configuration:
dflash benchmark openai \
--base-url http://127.0.0.1:8000 --model Qwen/Qwen3.8-27B \
--dataset gsm8k --num-prompts 128 --concurrency 1 --reasoning xhigh \
--temperature 1 --top-p 0.95 --top-k 20The local backends use --max-samples instead of --num-prompts, and the MLX path repeats the block size and draft bit width. Note that the README gives sampling parameters per example rather than documenting defaults, so a benchmark you run without those flags is not necessarily the benchmark the project ran. If you intend to compare numbers, pin temperature, top-p, top-k and reasoning level explicitly and keep them identical across configurations.
Where DFlash is the wrong tool
The supported model list is the hard boundary. DFlash 2 covers Muse-Glimmer-30B and Qwen3.8-27B. DFlash covers Qwen3.6, Qwen3.5, Qwen3, Gemma 4, MiniMax M2.5 and M2.7, Kimi K2.5 through K2.7-Code, GPT-OSS, Llama-3.1-8B, GLM 5.1 and Alpamayo. If your production model is not on that list, there is no draft checkpoint to pair with it, and speculative decoding without a matched draft is not a configuration you can assemble from this repository.
The README is explicit that the Transformers and MLX backends are for "their explicitly listed model families", and that other checkpoints can only be benchmarked through an OpenAI-compatible SGLang or vLLM server. That is a real constraint, not a footnote. It means the local backends are narrower than the serving path.
The second limitation is maturity. The package version is 0.1.0, the single release is v0.1.0 dated 2026-08-18, and pyproject.toml classifies the project as "Development Status :: 4 - Beta". The dependency pins are exact rather than ranged, including transformers==5.15.0 and torch==2.13.0 on Linux, so dflash will not float forward with your existing environment without a resolver conflict. Expect to isolate it.
Third, the serving path depends on forks and pull requests. The README links to specific SGLang and vLLM pull requests and to a signed oMLX disk image from a z-lab fork for Apple Silicon. Those are not upstream releases in the ordinary sense, and the README does not describe what happens when they diverge from mainline.
DFlash against multi-token prediction
The natural comparison is multi-token prediction, which trains the target model itself with extra heads that emit several future tokens in one pass. MTP removes the separate draft model entirely and folds the extra prediction into the model you are already serving.
DFlash takes the opposite route. It keeps a distinct draft model, but replaces the autoregressive draft with a block diffusion one, so the draft proposes a block of tokens in parallel rather than a sequence of single-token steps. The practical consequences differ. MTP requires a target checkpoint that was trained with those heads, which usually means adopting a specific model release. DFlash requires a target from its list plus a matching draft checkpoint, which means holding two models in memory and managing two sets of quantisation settings, as the MLX example's --draft-bits flag shows.
The advantage of the DFlash approach is that the draft is a separate artefact you can swap, quantise or retrain without touching the target. The cost is operational: two models, two sets of weights, and a draft quality that determines whether the parallel proposal is accepted often enough to pay for itself. The README does not publish acceptance-rate figures, so the only way to judge that trade on your workload is to run dflash benchmark against your own server.
Licence, maintenance and what an upgrade costs
DFlash is MIT licensed, declared both as license = "MIT" in pyproject.toml and via license-files = ["LICENSE"]. MIT is permissive and imposes no copyleft obligation on your own code. What it does not cover is the model weights. The checkpoints live in Hugging Face collections under z-lab, and their terms are separate from the repository licence. If you plan to ship a product built on a DFlash draft, read the model card for the specific checkpoint rather than assuming the MIT grant extends to it. This is a description of the licence files, not legal advice.
The repository is not archived and the last push was on 2026-08-18, which is the same timestamp as the v0.1.0 release. That is a single-release project with a short history, so there is no upgrade cadence to plan around yet. The upgrade cost that is visible today comes from the pins: torch, transformers, mlx and mlx-lm are all pinned to exact versions in the local extra, and the serving integrations point at specific pull requests. Moving any one of those forward is a coordinated change, not a version bump. Budget for a virtual environment per project rather than a shared one.
The README also points to a feedback form for requesting new model support, which suggests the supported list is expected to grow through requests rather than through a documented contribution path.
Editorial conclusion
Adopt dflash if you already serve one of the listed model families and want to test parallel drafting against your current setup, starting with the pip install and a single dflash generate call. Do not adopt it if your model is not in the checkpoint collections, if you need a stable API rather than a 0.1.0 beta, or if you cannot run a supported SGLang or vLLM build. Verify three things first: that a DFlash draft checkpoint exists for your exact target model, that the backend you intend to use lists that family in the README, and that your block size and quantisation combination matches the MLX constraint of block_size <= 5 for quantized targets.
Frequently asked questions
What is DFlash?
DFlash is a lightweight block diffusion model used as a draft model for speculative decoding, which the README describes as enabling efficient and high-quality parallel drafting. It ships as a Python package with a dflash command-line tool.
What is DFlash 2?
DFlash 2 is the newer generation listed in the README, with its own blog post and Hugging Face collection, and available checkpoints for Muse-Glimmer-30B and Qwen3.8-27B. The Transformers backend supports DFlash 2 for Muse-Glimmer-30B and the MLX backend supports it for Qwen3.8-27B.
How do I use dflash?
Install it with pip install dflash, or pip install "dflash[local]" for local inference, then run dflash generate with the transformers, mlx or openai backend. The openai backend takes a --base-url pointing at an SGLang or vLLM server you launch separately.
Is dflash lossless?
The README does not make a losslessness claim, so it cannot be confirmed from the available material. What it does state is that DFlash is a draft model for speculative decoding, where a target model verifies the drafted tokens.
What is DFlash speculative decoding?
It is speculative decoding where the draft model is a block diffusion model rather than an autoregressive one, so the draft proposes tokens in parallel. The README frames this as high-quality parallel drafting, with the target model still performing verification.
What is DFlash in llm terms?
In LLM serving, DFlash is the draft side of a speculative decoding pair. You supply a target model from the supported list plus a matching DFlash checkpoint, and run it through Transformers, MLX or an OpenAI-compatible SGLang or vLLM server.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/z-lab-dflash)