mlx-dspark: Lossless Speculative Decoding for Apple Silicon
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
At a glance
- What is it?
- mlx-dspark ports DeepSeek's DSpark and z-lab's DFlash speculative decoding to Apple Silicon via MLX, reaching up to 4.3x faster token generation with a lossless guarantee: the target model verifies every draft token, so outputs are bit-for-bit identical to plain decoding.
- Who is it for?
- mlx-dspark is the right choice for any developer running Gemma-4, Qwen3, LFM2.5, Nemotron, or the other supported targets on an Apple Silicon Mac and wanting faster responses without changing model output. Before adopting it, verify that your specific target and quantization appear in the auto-resolve registry: models outside that set still benefit from drafter-free speculation but will not reach the headline ratios.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The Problem: Slow Autoregressive Decoding on a Mac
Standard LLM decoding generates one token per forward pass through the full target model. On Apple Silicon, where the GPU and CPU share unified memory and the memory bandwidth is high relative to a discrete GPU, the bottleneck is not raw compute but the ratio of memory reads to useful work per pass. Each pass reads the entire model weight just to produce one token.
Speculative decoding breaks that bottleneck. A small drafter model proposes a batch of candidate tokens in one step; the larger target model verifies the whole batch in a single forward pass, accepting correct tokens and stopping at the first mismatch. When the drafter is accurate, one verify pass produces several tokens instead of one, multiplying throughput without changing the output distribution.
mlx-dspark is for Mac-based developers, researchers, and coding-agent users who run inference locally and find that plain mlx-lm decoding is too slow for their workflow. It is not a general-purpose inference server and it does not run on non-Apple hardware.
DSpark and DFlash: Two Drafter Architectures
The library implements two EAGLE-family drafter architectures under a single verify loop.
DSpark is semi-autoregressive. It originates from DeepSeek's DeepSpec codebase, where it was used to accelerate DeepSeek-V4. In mlx-dspark, DSpark runs against consumer-sized targets that have published DSpark checkpoint pairs, not against DeepSeek-V4 itself. The README is explicit on this point: the targets are models such as Gemma-4, Qwen3, and Nemotron, with published drafters, so this brings the drafter method to a Mac rather than running V4 inference.
DFlash uses block diffusion. The DFlash 2 drafter for Qwen3.8-27B, published by incoai, is the project's current best-measured pair: 4.22x math, 3.80x code, and 2.99x chat on an M4 Pro at 8-bit quantization.
Both drafters run inside the same verify pass. The `--mode auto` flag (the default) resolves the best available drafter for a given model and quantization; `--mode dspark` forces the DSpark path when you want to compare them directly. `--drafter` accepts any DeepSpec-native checkpoint for targets outside the auto-resolve registry.
Installing mlx-dspark and Running a Benchmark
The package installs from PyPI. Python 3.10 or newer is required, and the hardware must be Apple Silicon running macOS.
pip install mlx-dsparkThe install pulls mlx (0.32.0 or newer), mlx-lm (0.31.3 or newer), mlx-vlm (0.6.12 or newer), numpy, and huggingface-hub. The mlx-vlm floor is the first release that includes the Muse Glimmer module; earlier versions cannot load that target at all.
Once installed, the `mlx-dspark` CLI is available. The benchmark subcommand measures your specific machine's verify and drafter cost curves and prints per-content speedup ratios:
mlx-dspark benchmark --trials 3The README documents three benchmark contents: chat, code, and math. The printed medians show which content type benefits most from speculation on your hardware. That profile varies: copy-heavy code editing can push further than the table suggests, with the README citing 4.5x on Gemma-12B for code-refactoring workloads.
For production or tool use, mlx-dspark serves an OpenAI-compatible API endpoint, so any client that speaks the OpenAI chat completions format, including LM Studio and Claude Code, can point at it without modification. A native Mac app ships alongside the CLI with a chat interface, a model manager, live speculative-decoding telemetry, and a per-machine roofline view that shows where the loaded model sits relative to the measured memory bandwidth ceiling.
Reading the Benchmark Table: Where the Speedups Come From
The README publishes warm medians measured on an M4 Pro. The numbers vary by model, quantization, and content type. Several patterns are worth understanding before picking a model.
First, speedup is non-monotone in quantization. The README gives a full sweep for Ornith-1.0-9B: 4-bit achieves 1.38x, 8-bit achieves 2.17x, and bf16 drops to 1.54x on the code benchmark. The explanation is that MLX's unquantized matrix multiply has a roughly 2x cost cliff at the verify width used by speculation. Choosing bf16 to preserve precision comes at the cost of both ratio and absolute speed.
Second, the auto-tune system adjusts draft cap to your actual machine. The README explains that mlx-dspark measures your Mac's verify-to-drafter cost curves once on first use (around five seconds, cached per model, quant, and MLX version) and derives the optimal maximum draft length from them. The `--max-draft auto` flag additionally adapts the cap per round during generation. An M1 Pro therefore gets different cap values than the M4 Pro rows in the published table.
Third, code-editing workloads run faster than chat workloads on the same model. The README attributes this to match-scaled lookup drafts, where the model is re-emitting or refactoring content already in the context, producing longer accepted draft sequences.
Limitations: Platform Constraint and Model Coverage
mlx-dspark is macOS-only. The pyproject.toml classifier reads `Operating System :: MacOS`, and the underlying MLX framework is Apple Silicon-specific. Linux and Windows users have no supported path.
The auto-resolve registry covers the pairs that the project has measured and vouches for. As of the README, that includes Gemma-4 12B, Qwen3-8B, Qwen3-14B, Qwen3.6-27B, Qwen3.8-27B, Qwen3-4B, LFM2.5-1.2B, LFM2.5-2.6B, Muse-Glimmer-30B, Ornith-1.0-9B, Nemotron-3.5-Lightning-30B-A3B, and Ternary-Bonsai-27B. Any model outside that set can still be run with `--drafter` pointing at a compatible checkpoint, but the auto-resolve convenience does not apply and the user is responsible for verifying drafter-target compatibility.
MoE (mixture-of-experts) targets show lower ratios. The README table shows Qwen3.6-35B-A3B at 1.67x math and 1.05x chat, and Nemotron at 1.34x math. The README does not explain why MoE targets benefit less, but the pattern is consistent across those entries.
Finally, the published speedup numbers are short-context figures. The README notes that at agent-sized prompts of 16k to 32k tokens, the numbers differ and points to a separate long-context section dated 2026-09-27 for that data.
DSpark vs DFlash: What the Comparison Shows
Both drafters are available for Qwen3.8-27B, making it the only target in the registry with a head-to-head comparison in the README. At 8-bit, DFlash 2 (the `incoai/Qwen3.8-27B-DFlash2` checkpoint) is the default and the faster pair: 4.22x math versus what the README describes as lower ratios for the DSpark heads on the same target. The README notes that DSpark heads for Qwen3.8-27B remain available with `--mode dspark` for users who want the comparison.
The broader implication is that drafter architecture is model-specific. DFlash 2's block-diffusion approach happens to suit Qwen3.8-27B better than the semi-autoregressive DSpark approach, but the project does not generalize this to other targets. The auto mode resolves the registered best for each model, so the user does not need to choose unless benchmarking.
Alternative and Maintenance
The direct alternative is plain mlx-lm decoding, which mlx-dspark builds on. mlx-lm handles inference on Apple Silicon without speculative decoding: a simpler dependency footprint, broader model compatibility, and no drafter-target pairing required, but single-token-per-pass throughput. For users who need speculative decoding and are not on Apple Silicon, llama.cpp supports draft models on macOS, Linux, and Windows, but it uses a different inference stack and does not run the MLX graph.
The last push to the repository was on 2026-09-10, and the most recent release, v0.20.1, landed on 2026-09-27. The project publishes regular minor releases. The MIT license allows commercial use, modification, and redistribution without restriction. The models accelerated by mlx-dspark carry their own terms: Qwen Image has a research and evaluation license that requires a separate agreement for commercial use, and YuE2 and ternary Bonsai weights carry noncommercial model terms, as the README notes in context.
Editorial conclusion
mlx-dspark is the right choice for any developer running Gemma-4, Qwen3, LFM2.5, Nemotron, or the other supported targets on an Apple Silicon Mac and wanting faster responses without changing model output. Before adopting it, verify that your specific target and quantization appear in the auto-resolve registry: models outside that set still benefit from drafter-free speculation but will not reach the headline ratios. The bf16 quantization quirk in MLX means 8-bit is usually the better pick for both speed and ratio, which matters if you are choosing a download. The MIT license imposes no restrictions on commercial use of the library itself, though the models it accelerates carry their own terms.
Frequently asked questions
Is mlx-dspark output identical to plain mlx-lm decoding?
Yes. The lossless guarantee comes from the verify step: after the drafter proposes tokens, the target model runs a forward pass to check them all. Only tokens that match the target's distribution are accepted. The README states the output is identical to normal decoding.
Which models work with mlx-dspark's automatic drafter resolution?
The auto-resolve registry covers Gemma-4 12B, Qwen3-4B, Qwen3-8B, Qwen3-14B, Qwen3.6-27B, Qwen3.8-27B, LFM2.5-1.2B, LFM2.5-2.6B, Muse-Glimmer-30B, Ornith-1.0-9B, Nemotron-3.5-Lightning-30B-A3B, and Ternary-Bonsai-27B. Other models can still use speculation via the --drafter flag with a compatible checkpoint.
Does mlx-dspark run on Intel Macs or Linux?
No. The library targets Apple Silicon only. The pyproject.toml classifier specifies macOS, and MLX itself requires Apple Silicon hardware. There is no supported Linux or Windows path.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/arahim3-mlx-dspark)