SpecForge: Training EAGLE3 and DFlash Draft Models for SGLang
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
At a glance
- What is it?
- SpecForge is the SGLang team's training framework for speculative decoding draft models, with typed configs for online and offline topologies. It is opinionated about serving and thin on installation detail.
- Who is it for?
- Adopt SpecForge if you already serve models with SGLang and want draft checkpoints that load without porting work, and if you can read the example configs under examples/configs to learn the topology options. Do not adopt it if you serve with vLLM or TensorRT-LLM, or if you need a documented installation walkthrough: the README points to docs.sglang.io/SpecForge and shows a single train command, and pyproject.toml pins sglang==0.5.18 alongside torch==2.13.0.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What SpecForge is for, and who it is not for
Speculative decoding pairs a small draft model with a large target model so the draft proposes tokens and the target verifies them in a batch. The hard part is not the serving loop, it is producing a draft model whose proposals the target actually accepts. SpecForge is the SGLang team's framework for training those draft models and moving them into SGLang serving. The README states the project exists because most open-source speculative decoding projects are "not well-maintained or not directly compatible with SGLang", and it names three goals: runnable out of the box, direct compatibility with SGLang, and one runtime for online disaggregated training plus colocated and disaggregated offline training.
The intended user is an engineer who already runs SGLang and wants an acceptance-rate gain without writing a conversion layer between a third-party trainer and the serving stack. The README does not frame SpecForge as a general research toolkit. If you serve with a different engine, the compatibility argument does not apply to you, and the value proposition shrinks to the training code alone.
The project is an "ecosystem project developed by the SGLang team" according to the README, licensed MIT, written in Python, and the last push to the default branch was on 2026-09-09.
One typed entry point, many methods
The design decision that shapes everything else is that there are no method-specific Python training entry points. The README says so directly, and the command surface is a single CLI verb:
specforge train --config examples/configs/online/disaggregated/external/qwen3-8b-eagle3-disaggregated.yamlThe path under examples/configs encodes three things: feature mode, topology, and who owns the online services. In an external recipe, SpecForge supervises the producer and consumer on one trainer node while the user or scheduler owns Mooncake and SGLang. Recipes under managed-local additionally start those services on the local host. That split is the most interesting part of the project. It means the config file, not a Python flag, is the unit of configuration, and unsupported combinations are rejected during config validation or run assembly rather than silently falling back to an older trainer. Rejecting early is the right call for a training job that may burn hours before failing.
The supported methods are EAGLE3, P-EAGLE, EAGLE3.1, DFlash, DFlash2, Domino, and DSpark. Each row in the README table links an example config, and several list an optimization: LK loss for EAGLE3, D-PACE for DFlash and DFlash2. DSpark is described as confidence-scheduled semi-autoregressive generation. Note that only EAGLE3 and DFlash have offline colocated examples listed; DFlash2 appears with an online managed-local example only.
Installing SpecForge and running a first training job
The README does not carry an installation section; it points to the documentation at docs.sglang.io/SpecForge for getting started. What the repository does give is packaging metadata in pyproject.toml, which declares the console script and the dependency set. The project requires Python 3.10 or newer and pins sglang==0.5.18, torch==2.13.0, and transformers==5.12.1, so the environment is narrow by design. Installing from a checkout would look like this:
python -m pip install -e .That installs the specforge package and registers the specforge command. Optional extras are declared for data (openai), dev (pre-commit), fa (flash-attn, ninja, packaging), and liger (liger-kernel). A ROCm requirements file, requirements-rocm.txt, sits at the repository root.
Once installed, the first real run is a config-driven train invocation. The README's own example uses an external online disaggregated recipe for Qwen3-8B with EAGLE3:
specforge train --config examples/configs/online/disaggregated/external/qwen3-8b-eagle3-disaggregated.yamlBecause the recipe is external, you own Mooncake and SGLang in this setup; SpecForge only supervises the producer and consumer on the trainer node. If you would rather have those services started on the local host, the managed-local recipes under examples/configs do that. A simpler starting point is the offline colocated EAGLE3 example, examples/configs/offline/colocated/qwen3-8b-eagle3-offline.yaml, which avoids the online service split entirely. The README does not document expected log output or a success marker, so treat the first run as an environment check rather than a benchmark.
Where SpecForge gets awkward
The dependency pinning is the first real constraint. torch==2.13.0, transformers==5.12.1, and sglang==0.5.18 are exact pins, not ranges. If your serving cluster runs a different SGLang build, you are either changing the pin or running SpecForge in a separate environment from the one that serves the model. The README does not discuss that scenario.
The second limitation is documentation coverage. The README is a launch page: it links a training guide and a disaggregated guide under docs/sections/basic_usage/, and it links a SpecBundle collection and performance dashboard. It does not document rollback, checkpoint compatibility with older SGLang versions, or what happens when the target model has no matching example config. For a project whose selling point is "no additional porting effort", the missing piece is a written statement of which SGLang versions can load a checkpoint trained today.
The third is scope. SpecForge trains draft models; it does not serve them, and it does not evaluate acceptance rate for you. The README claims SpecBundle models give "up to 4x speedup for inference" together with SGLang, but that claim is attached to the SpecBundle checkpoints, not to whatever you train yourself. There is no promise in the README about the acceptance rate of a model you produce.
SpecForge against a general fine-tuning stack
The obvious alternative is to train a draft model with a general-purpose fine-tuning stack: Hugging Face transformers with TRL or plain PyTorch training loops, then convert the result for serving. The difference in approach is not the optimizer, it is the data flow. SpecForge is built around generating and consuming target-model hidden features and logits, which is why the online recipes need a producer and a consumer and why external recipes leave Mooncake and SGLang to the user. A generic fine-tuning stack has no concept of a producer process feeding training features; you would write that pipeline yourself.
The trade-off runs the other way too. A generic stack supports any architecture and any serving engine, and it does not pin sglang==0.5.18. SpecForge supports seven named methods and one serving target. The README's own framing is that this narrowness is the point: existing speculative decoding projects were not directly compatible with SGLang, so the team built one that is. If your serving engine is not SGLang, that argument disappears and the pinned dependency becomes pure cost.
Licence and maintenance cost
SpecForge is MIT licensed, and the LICENSE file sits at the repository root. MIT is permissive: it allows commercial use and modification, and it requires that the copyright notice and permission notice be preserved. The README does not discuss the licences of the model checkpoints you train or of the SpecBundle collection, and it does not discuss the licences of dependencies. Whether a draft model trained on a given target model inherits constraints from that target model's licence is a question the repository does not answer, and it is worth asking before shipping a checkpoint.
On maintenance, the last push to the default branch was on 2026-09-09, and the repository is not archived. Upgrade cost is dominated by the exact pins: moving to a newer SGLang means editing pyproject.toml and re-validating that your configs still assemble, since the README states unsupported combinations are rejected at config validation or run assembly. There are no retrieved releases, so there is no changelog to read between versions.
What to check before you commit a run
Three things are worth confirming first. Does an example config exist for your target model and method combination? The README table lists Qwen3-8B, Qwen3-4B, Qwen3-30B-A3B, and Qwen3.6-27B across its examples, and the docs are the place to check for anything else. Second, which topology do you actually want? Offline colocated is the smallest moving part; online external means you own Mooncake and SGLang. Third, can your serving environment accept the pinned sglang==0.5.18, or will you run training and serving in separate environments?
The README is explicit that online target parallelism belongs to SGLang while deployment.trainer owns trainer DP and offline EAGLE3 USP process groups. That sentence is the clearest statement of the boundary between the two systems, and it is the sentence to re-read when a run fails to assemble.
Editorial conclusion
Adopt SpecForge if you already serve models with SGLang and want draft checkpoints that load without porting work, and if you can read the example configs under examples/configs to learn the topology options. Do not adopt it if you serve with vLLM or TensorRT-LLM, or if you need a documented installation walkthrough: the README points to docs.sglang.io/SpecForge and shows a single train command, and pyproject.toml pins sglang==0.5.18 alongside torch==2.13.0. Verify that your target model has a matching example config before you commit a training run.
Frequently asked questions
What is SpecForge?
SpecForge is a framework from the SGLang team for training speculative decoding models and porting them to the SGLang serving framework. It supports methods including EAGLE3, P-EAGLE, DFlash, DFlash2, Domino, and DSpark through a single typed training entry point.
How do I install SpecForge?
The README does not include installation steps and instead points to the documentation at docs.sglang.io/SpecForge. The repository's pyproject.toml declares a specforge console script and requires Python 3.10 or newer.
How do I train a model with SpecForge?
You run the specforge train command with a config path, for example specforge train --config examples/configs/online/disaggregated/external/qwen3-8b-eagle3-disaggregated.yaml. The path under examples/configs identifies the feature mode, topology, and online service ownership, and there are no method-specific Python training entry points.
Which speculative decoding methods does SpecForge support?
The README lists EAGLE3, P-EAGLE, EAGLE3.1, DFlash, DFlash2, Domino, and DSpark, each with at least one example config. EAGLE3 lists LK loss as an optimization and DFlash and DFlash2 list D-PACE.
Does SpecForge serve the models it trains?
No. SpecForge trains speculative decoding models for use with the SGLang serving framework, and in external online recipes the user or scheduler owns Mooncake and SGLang while SpecForge supervises the producer and consumer on the trainer node.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/sgl-project-specforge)