DeepSpec: A Full-Stack Codebase for Training and Evaluating Speculative Decoding Draft Models
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
At a glance
- What is it?
- deepseek-ai/DeepSpec is an MIT-licensed Python repository that covers data preparation, draft model training, and evaluation for speculative decoding algorithms. It includes three implemented algorithms (DSpark, DFlash, Eagle3), pre-released Hugging Face checkpoints, and an evaluation suite across nine benchmark datasets.
- Who is it for?
- DeepSpec is the right tool for researchers who want to train speculative decoding draft models for a specific target LLM or study the DSpark, DFlash, and Eagle3 algorithms on a working codebase. It is the wrong choice for teams that only need inference-time speculative decoding: those teams should use an inference engine that provides the feature out of the box rather than training a new draft model.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 82 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
Speculative decoding and what DeepSpec trains
Speculative decoding is an inference acceleration technique for large language models. A smaller draft model generates a sequence of candidate tokens quickly. The larger target model then verifies those candidates in a single forward pass. When the target model accepts the draft tokens, the result is that multiple tokens are produced per target model step rather than one, reducing inference latency without changing the output distribution.
The draft model is the variable that determines how effective this approach is. A draft model that produces candidates the target model frequently rejects provides little speedup. Training a high-acceptance-rate draft model that is fast enough to be worth running is a non-trivial task, and it requires dataset preparation, custom training procedures, and evaluation tooling specific to speculative decoding.
DeepSpec provides all three components. The README describes it as a full-stack codebase covering data preparation utilities, draft model implementations, training code, and evaluation scripts. It targets the workflow that a researcher needs when developing a new draft model for a specific target LLM, not the deployment workflow of someone running inference with an existing model.
The three-stage pipeline: data, training, evaluation
The README organizes the workflow into three stages that run in sequence, with each stage's output feeding the next.
The first stage is data preparation: downloading prompts, regenerating target model answers using an inference engine to serve the target model, and building the target cache. The README references scripts/data/README.md for the step-by-step pipeline and includes a storage warning: the target cache can be very large, with approximately 38 TB for the default Qwen/Qwen3-4B setting. This is a hard infrastructure requirement that must be in place before training starts.
The second stage is training. The training script launches one worker per visible GPU and reads its configuration from a file in the config/ directory. The third stage is evaluation. The eval script runs the trained draft checkpoint against a fixed set of benchmark datasets and measures speculative decoding acceptance. The benchmark suite covers nine tasks: gsm8k, math500, aime25, humaneval, mbpp, livecodebench, mt-bench, alpaca, and arena-hard-v2.
Installing DeepSpec and running training
The requirements are pinned to specific versions. Install them with:
python -m pip install -r requirements.txtThe requirements.txt lists: torch==2.9.1, transformers==5.10.2, triton==3.5.1, numpy==2.4.4, tqdm==4.67.3, tensorboard==2.20.0, matplotlib==3.10.9, sentencepiece==0.2.1, safetensors==0.7.0, and prettytable==3.17.0 for the core dependencies. Data preparation additionally requires datasets==4.8.5 and openai==2.6.1 for regenerating answers from a served target model.
Once the target cache is built, run training with:
bash scripts/train/train.shThis launches train.py and spawns one worker per visible GPU. The configuration is selected by pointing `config_path` at a file in the config/ directory. The README gives `config/dspark/dspark_qwen3_4b.py` as an example. Checkpoints are written to `~/checkpoints/<project_name>/<exp_name>/step_*`. The `--opts` flag supports overriding individual config fields without editing the file.
To evaluate a trained checkpoint:
bash scripts/eval/eval.shThe eval script requires two parameters: `target_name_or_path` specifying the target model the draft was trained against, and `draft_name_or_path` pointing to the checkpoint directory or a Hugging Face repository ID.
The three algorithms: DSpark, DFlash, and Eagle3
DeepSpec implements three draft model approaches, each represented as a configuration set in the config/ directory and a corresponding Hugging Face checkpoint family.
DSpark is the algorithm described in the paper at arXiv:2607.05147, titled "DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation." The published checkpoints are available for Qwen/Qwen3-4B, Qwen/Qwen3-8B, Qwen/Qwen3-14B, and google/gemma-4-12B-it target models.
DFlash (arXiv:2602.06036) is the second algorithm, with its own checkpoint family covering the same target model set. Eagle3 (arXiv:2503.01840) is the third. The Eagle3 implementation in DeepSpec is adapted from SpecForge (Apache-2.0), with the README listing the specific portions borrowed: modeling, loss, optimizer, attention, and evaluation code. Adapted files carry in-file attribution comments, and a full notice is recorded in the NOTICE file at the repository root.
All three algorithms are trained on open-perfectblend data generated by the corresponding target model in non-thinking mode. The README includes an explicit note about comparison validity: when citing these results in a new paper, align the setup with the training settings in this repository; otherwise, the comparison is not meaningful. For domain-specific use cases, the README recommends fine-tuning the draft model further, especially when the target model runs in thinking mode.
Hardware requirements: 8 GPUs and the 38 TB cache constraint
The default configurations and training scripts assume a single node with eight GPUs. The README states this directly and provides a workaround for smaller setups: reduce CUDA_VISIBLE_DEVICES to the number of available GPUs. There is no documented multi-node training path.
The more constraining requirement is storage for the target cache. The README states that the cache can be very large, with approximately 38 TB for the default Qwen/Qwen3-4B setting. This is not a theoretical maximum but the documented default expectation. Teams that do not have access to 38 TB of fast storage cannot run the default data preparation pipeline without modification.
The pinned dependency versions are also a constraint. torch==2.9.1 and triton==3.5.1 require a compatible CUDA version. The requirements.txt does not specify the required CUDA version explicitly, but the comment at the top of the file notes that the default PyTorch wheel may not be appropriate for all environments and suggests installing the correct CUDA build separately.
Released checkpoints and when to use them directly
DeepSpec publishes trained checkpoints on Hugging Face for all three algorithms across the four supported target model sizes. The Hugging Face repository IDs follow the pattern `deepseek-ai/<algorithm>_<target_model>_<config>`. For example, the DSpark checkpoint for Qwen/Qwen3-4B is at `deepseek-ai/dspark_qwen3_4b_block7`.
These checkpoints are the ones used to produce Table 1 in the DSpark paper, and the README notes that they are the direct output of the training configurations in the config/ directory. A researcher who wants to reproduce the paper's results can use these checkpoints directly with the eval script without training from scratch.
Using a published checkpoint avoids the 38 TB data preparation step entirely, which makes it practical to run evaluations on hardware that cannot store the full target cache. The limitation is that these checkpoints were trained on non-thinking-mode outputs from the listed target models. Teams that need speculative decoding for a different target model, a domain-specific dataset, or a target running in thinking mode must train from scratch using the provided pipeline.
vLLM as the inference-only alternative
vLLM is a widely deployed inference engine for large language models that includes speculative decoding as a built-in feature. The fundamental difference from DeepSpec is the scope: vLLM provides speculative decoding at inference time using existing draft and target model pairs; DeepSpec trains the draft model itself.
A team that wants faster inference and is happy using a published draft model checkpoint should use vLLM's speculative decoding configuration rather than DeepSpec. The DeepSpec-published Hugging Face checkpoints are compatible with vLLM's speculative decoding feature, which means the workflow is: train or download a draft model checkpoint using DeepSpec, then deploy the target and draft models together in vLLM for production inference.
DeepSpec is the right tool only when the goal is training or benchmarking draft models. If the goal is faster inference from a known model, deploying an existing checkpoint in an inference engine eliminates the training and data preparation steps entirely.
Editorial conclusion
DeepSpec is the right tool for researchers who want to train speculative decoding draft models for a specific target LLM or study the DSpark, DFlash, and Eagle3 algorithms on a working codebase. It is the wrong choice for teams that only need inference-time speculative decoding: those teams should use an inference engine that provides the feature out of the box rather than training a new draft model. The 38 TB cache requirement for the default Qwen/Qwen3-4B configuration is a concrete infrastructure constraint to verify before starting. If the goal is to compare against the paper results, align the training setup with the configurations in this repository: the README explicitly warns that comparisons using different setups are not meaningful.
Frequently asked questions
What is DeepSpec?
DeepSpec is a Python codebase from deepseek-ai for training and evaluating draft models used in speculative decoding. It covers data preparation, training with three implemented algorithms (DSpark, DFlash, Eagle3), and evaluation across nine benchmark datasets including gsm8k, humaneval, and mt-bench.
What hardware does DeepSpec require for training?
The default configurations assume a single node with eight GPUs. For fewer GPUs, the README instructs reducing CUDA_VISIBLE_DEVICES. Data preparation for the default Qwen/Qwen3-4B setting requires approximately 38 TB of storage for the target cache.
Can I use DeepSpec checkpoints without running the full training pipeline?
Yes. DeepSpec publishes trained checkpoints on Hugging Face for DSpark, DFlash, and Eagle3 across multiple target model sizes. These checkpoints can be passed directly to the eval script using the draft_name_or_path parameter, bypassing the data preparation and training stages.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/deepseek-ai-deepspec)
Community notes