# Speculators: training speculative decoding draft models for vLLM

> Speculators is a Red Hat library that trains draft models for speculative decoding and packages them in a Hugging Face-compatible format for deployment in vLLM. The training pipeline is well defined; the deployment path still depends on vLLM support for each algorithm.

**vllm-project/speculators** — A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

- Repository: https://github.com/vllm-project/speculators
- Website: https://docs.vllm.ai/projects/speculators
- Stars: 859 · Forks: 238
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/vllm-project-speculators

## The problem Speculators solves: draft models that never reach production

Speculative decoding is a lossless technique. A smaller draft model proposes several tokens, the base model verifies them in a single forward pass, and every accepted token is guaranteed to match what the base model would have produced on its own. The speedup comes from verification being cheaper than generation, not from approximating the output. The hard part is not the idea. It is producing a draft model that an inference server will actually load.

Research repositories each train draft models their own way, with their own checkpoint layout and their own assumptions about the verifier. Getting one of those checkpoints into a serving engine means writing conversion code, matching hidden state extraction to what the trainer expects, and hoping the engine's kernel supports the architecture. Speculators targets that gap directly. It is a library for training draft models that deploy into vLLM, and it standardizes the checkpoint format so a trained model is not trapped in the repository that produced it.

The audience is narrow and specific. You need a vLLM deployment, a base model you cannot or will not swap, and latency headroom worth recovering. If you serve through a different engine, or you are still choosing a base model, this library adds a training stage without a payoff.

## Inside the pipeline: hidden states, a trainer, and a Hugging Face format

The README describes two stages. First, offline training data generation uses vLLM to produce hidden states. Samples are saved to disk rather than held in memory, which is what makes training on pre-extracted states possible without loading the full verifier. Second, draft model training runs end to end on that data, and the repository layout reflects the split: examples/train/, examples/evaluate/ and examples/convert/ sit alongside src/ and tests/.

The hs_connectors plugin package handles the online case. It is a separate workspace member, declared in pyproject.toml under tool.uv.workspace and resolved as a local source, with pluggable backends for moving hidden states between vLLM and the trainer across nodes. The file-based backend uses a shared filesystem. The Mooncake backend uses a distributed store for environments without shared storage. That is the design decision worth noticing: the project treats hidden state transport as a swappable component instead of baking in one assumption about cluster topology.

Algorithm support is where the library has grown fastest. DFlash uses anchored-block drafting with auxiliary hidden states from multiple verifier layers. DSpark extends it with a Markov head that conditions each draft position on the previous token in the block, plus a confidence head that predicts per-position acceptance probability, and DSpark checkpoints can warm-start from existing DFlash checkpoints. P-EAGLE extends EAGLE-3 with parallel multi-token prediction via Conditional-On-Distribution sampling, predicting multiple tokens in one forward pass instead of sequentially. MTP finetuning targets the native multi-token prediction heads of models like Qwen3-Next, which the README puts at roughly 100M to 400M parameters.

DFlash and DSpark use sliding window attention on all draft layers by default. The --sliding-window flag sets the window size and --full-attention-indices opts specific layers into full attention. The stated trade-off is KV cache allocation for long-context sequences against per-position acceptance rates, and the README claims sliding window can improve acceptance compared to full attention. Treat that as a claim from the project, not a measured result.

## Installing Speculators from PyPI and running a first evaluation

The package is published on PyPI as speculators, requires Python 3.10 or newer, and the badge in the README lists support through 3.13. The repository is built with setuptools and uses a uv workspace for hs_connectors, so a plain pip install gets the base library while the connector plugin is resolved at build time based on build type.

```bash
pip install speculators
```

After that, the fastest way to see real behaviour is the evaluation example directory rather than a training run. Training a draft model is a GPU-hours commitment; evaluation tells you whether a checkpoint is worth deploying. The repository keeps examples/evaluate/ next to examples/train/ and examples/convert/, and the convert path is the one to use if you are bringing a checkpoint in from an external research repository.

The README does not reproduce the full flag list for the example scripts, so read the script or the documentation site at docs.vllm.ai/projects/speculators before assuming a flag exists. This is a library with example scripts, not a CLI with a stable surface across every algorithm.

For the training path, the CLI options mentioned in the README include --sliding-window and --full-attention-indices for DFlash and DSpark, plus block size and max anchor options for DFlash. Those flags belong to the training entry points in examples/train/. If you are fine-tuning an MTP head, the README's point is that the head is small enough to train on pre-extracted hidden states without loading the full verifier, which changes the memory profile of the job substantially compared to training a multi-layer draft model from scratch.

## Where Speculators will cost you time

The deployment half of the promise is not entirely in this repository's hands. The README notes that DFlash models trained through Speculators can run in vLLM as of vLLM PR #38300. That phrasing is a signal: algorithm support in vLLM is a separate change, tracked separately, and it lands on its own schedule. A training algorithm appearing in Speculators does not mean the corresponding checkpoint loads in your vLLM build. If you are on an older vLLM release, check the engine side before you start a training run.

Training cost is the second constraint. The library generates hidden states offline using vLLM, which means a full pass over your data with the verifier model before training even begins. That is the price of the design, and it is why the MTP case is attractive: a head in the 100M to 400M parameter range can be trained on pre-extracted states without loading the verifier, while a conventional multi-layer draft model cannot.

Finally, this is the wrong tool if your bottleneck is not decoding latency. Speculative decoding adds a second model to the serving path and consumes KV cache for the draft. If you are memory-bound at long context, the sliding window defaults exist precisely because that budget matters, and you should understand the window size before assuming the default suits your sequence lengths. The README does not document a rollback procedure for a deployed speculator, so plan for that on the serving side.

## Speculators against hand-rolled EAGLE pipelines

The obvious alternative is the research repository for whichever algorithm you want. EAGLE-3 and its relatives have public training code, and if you only ever intend to train one algorithm for one base model, that code will work. The difference is in what surrounds it. A research repository typically assumes you will adapt the data pipeline, and it does not ship a conversion path from other formats or a checkpoint layout that a serving engine reads without changes.

Speculators inverts that. The training code is one part; the standardized Hugging Face-compatible format for defining speculative models and the tools to convert from external research repositories are the other. That matters when you have more than one algorithm in play, which is exactly the situation the README describes, with DFlash, DSpark, P-EAGLE and MTP finetuning all supported under one interface. The cost of that breadth is that each algorithm's options are exposed through the same example scripts, and the documentation site, not the README, is where the per-algorithm detail lives.

A second alternative is to skip draft models entirely and buy latency with hardware or quantization. That is a legitimate answer if your quality budget allows it. Speculative decoding is the option that preserves output exactly, which is the reason to pay its complexity cost.

## Maintenance, licensing and the upgrade path

The repository is not archived, and the last push was on 2026-09-10. Releases are frequent: v0.8.0 on 2026-09-03, v0.7.0.1 on 2026-08-13, and v0.6.0.1 on 2026-08-12. Version numbers are derived from git tags by setuptools-git-versioning, and setup.py distinguishes release, nightly, alpha and dev build types, so nightly builds carry a different version shape than tagged releases. If you pin Speculators in a production image, pin the tagged release, not a nightly.

The licence is Apache-2.0, declared in pyproject.toml and shipped as a LICENSE file. That is a permissive licence with an explicit patent grant, which is generally the easy case for commercial deployment, but the trained checkpoints you produce are a separate question from the library licence, and the terms attached to the base model you train against are not addressed here. Read the base model's licence before you ship a speculator built on it.

Upgrade cost is dominated by the vLLM pairing rather than by Speculators itself. Because checkpoint compatibility depends on engine-side support, an upgrade means checking both sides. The repository's own quality gate is visible in the Makefile: ruff check, ruff format --check, mdformat on markdown files, mypy with --check-untyped-defs, and a provenance sync check. That tells you the project enforces typing and formatting on contributions, which is a reasonable proxy for how much churn an upgrade will bring, though it says nothing about API stability between minor versions.

## Conclusion

Adopt Speculators if you already serve models on vLLM and have the GPU budget to train or fine-tune a draft model, especially for the algorithms the README lists as deployable. Do not adopt it if you expect a pretrained speculator to drop into your stack without training, or if your inference engine is not vLLM. Before committing, verify that your target algorithm has landed in vLLM, and check the sliding window attention defaults for DFlash and DSpark against your context length.

## FAQ

### What is Speculators in the vLLM project?

It is a library for training speculative decoding draft models that deploy into vLLM. It provides an end-to-end framework for training draft models with reusable formats, and it is published on PyPI as speculators under the Apache-2.0 licence.

### What Python versions does Speculators support?

The README badge lists Python 3.10 through 3.13, and pyproject.toml sets requires-python to >=3.10. The package builds with setuptools and uses a uv workspace for the hs_connectors plugin.

### How do I install Speculators?

It is published on PyPI, so pip install speculators installs the base library. The hs_connectors plugin is a separate workspace member resolved at build time based on build type.

### Does a Speculators-trained model run in vLLM?

The README states that DFlash models trained through Speculators can run in vLLM as of vLLM PR #38300. Support is tracked on the vLLM side per algorithm, so confirm your engine version before starting a training run.

### What are the sliding window attention flags in Speculators?

DFlash and DSpark speculators use sliding window attention on all draft layers by default. The --sliding-window flag sets the window size, and --full-attention-indices opts specific layers into full attention.

## Sources

- [License: Apache-2.0](https://github.com/vllm-project/speculators/blob/main/LICENSE)
- [Project website](https://docs.vllm.ai/projects/speculators)
- [README](https://github.com/vllm-project/speculators/blob/main/README.md)
- [Releases](https://github.com/vllm-project/speculators/releases)
- [vllm-project/speculators on GitHub](https://github.com/vllm-project/speculators)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vllm-project-speculators
