Lucebox: a speculative inference server tuned per chip, not per framework
LLM speculative inference server for heterogeneous hardware & consumer GPUs
At a glance
- What is it?
- Lucebox is a C++ LLM inference server from Luce-Org that pairs speculative prefill and decoding with hand-written kernels for specific consumer GPUs and APUs. The measured results are tied to named model and hardware pairs, and that specificity is both the pitch and the constraint.
- Who is it for?
- Adopt Lucebox if you own one of the tested machines, specifically an R9700, Strix Halo, RX 7900 XT or XTX, RTX 3090 or RTX 5090, and you are willing to run the exact model and drafter pair the project publishes for that target. Skip it if you need a portable server that runs the same way across many GPUs, or if your workload is prompt-heavy at long context, where the published prefill numbers apply only to the Laguna XS 2.1 33B configuration.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem Lucebox solves: one chip, one tuned engine
Most local inference stacks try to run acceptably everywhere. Lucebox takes the opposite position. The project description in pyproject.toml states the intent directly: "open LLM inference, rewritten by hand for one specific chip at a time." That sentence is the whole design thesis, and it explains why the README is organised as a table of measured setups rather than a list of features.
The audience is narrow and identifiable. You have a consumer GPU or an APU, you want to serve a model locally, and you are prepared to accept a build that is compiled for your exact device. The tested machines table names RDNA4 gfx1201 (Radeon AI PRO R9700), RDNA3.5 gfx1151 (Ryzen AI MAX+ 395, Strix Halo), RDNA3 gfx1100 (RX 7900 XT and XTX), Ampere sm_86 (RTX 3090) and Blackwell sm_120 (RTX 5090). If your card is not on that list, the project has no measured configuration for you, and you are on your own.
The payoff for that constraint is that the optimisations are named and scoped. DFlash2, DSpark, PFlash, KVFlash, Luce Spark and the Megakernel are not general features; each is attached to a model and a device in the README tables. That is unusual, and it is the reason the numbers are worth reading even if you never install anything.
How speculative prefill and decoding actually run here
Lucebox is a speculative inference server, which means a small drafter model proposes tokens and a larger target model verifies them in a batch. The README separates this into two phases. PFlash and KVFlash operate on prefill, the stage that processes the prompt. DFlash, DFlash2, DSpark and Luce Spark operate on decode, the token-by-token generation stage. A configuration therefore needs two artefacts: the target weights and a matching drafter.
The README makes that pairing explicit. For Qwen 3.8 27B UD-IQ4_XS on an R9700, the DFlash2 source checkpoint is published on Hugging Face and the README notes it is "converted to Q8_0." For Laguna XS 2.1 33B, the DFlash drafter is already published as a GGUF quant. For DeepSeek V4 Flash on Strix Halo, the target is a ROCmFPX MIX Strix GGUF and the drafter is a DSpark Q4RMFP4 file. Conversion is sometimes your job, sometimes not.
Around the speculation, the server implements paged attention and continuous batching, which the README lists as serving multiple clients concurrently. The measured setups there are Qwen 3.8 27B with DFlash2 on R9700 at 300.9 tok/s total across 5 clients, and DeepSeek V4 Flash autoregressive on Strix Halo at 48.4 tok/s output-window across 4 clients. Heterogeneous execution is the other architectural piece: the README reports DeepSeek V4 on an R9700 plus Strix Halo combination at 86 tok/s decode and 788 tok/s prefill at 2K. That implies the engine can split work across a discrete GPU and an APU, though the README does not document how the split is decided.
The Megakernel is a different kind of optimisation. It targets Qwen 3.5 0.8B BF16 on RTX 3090 and is reported at 413 tok/s and 1.87 tok/J, with the results file living under optimizations/megakernel/RESULTS.md. The energy-per-token figure is the interesting one, because it suggests the work was measured with power draw in mind rather than throughput alone.
Installing Lucebox and running the Qwen 3.8 R9700 quick start
The repository ships a Makefile that describes itself as the "single entry point for the common dev/CI ops on lucebox-hub," and it shells out to uv and docker buildx bake. The Makefile comment warns that these are "pre-release software" targets and that they assume bash, GNU coreutils, a working docker buildx and uv on PATH. Start by listing what is available.
make helpThat target prints every documented target with its one-line description, extracted from the Makefile itself. Next, narrow the CUDA architecture list. The Makefile states that restricting the list to your local GPU cuts build time by 5 to 6 times, and the default covers sm_75 through sm_120.
make build DFLASH_CUDA_ARCHES=120That builds the local image tagged lucebox-hub:cuda12 through docker buildx bake. The Dockerfile confirms the prebuilt CUDA path is CUDA 12.8.1 on Ubuntu 22.04 with DFLASH_CUDA_ARCHES defaulting to "75;80;86;89;90;120". It also notes that each architecture adds roughly 50 to 200 MB of fat-binary kernel code and 3 to 5 minutes of nvcc time per translation unit, which is the reason to trim the list.
Models are bind-mounted from a directory the Makefile defaults to $(HOME)/models, and the serve target runs the local image in the foreground with a gemma-4-26b default. The README points at a Qwen 3.8 R9700 quick start under the run-the-server heading, so expect the exact model path and flags to come from that section rather than from the Makefile. For AMD targets, the README is explicit that HIP builds "should target the device's exact gfx architecture," and the tested machines table gives the values: gfx1201 for R9700, gfx1151 for Strix Halo, gfx1100 for RX 7900 XT and XTX. There is a separate Dockerfile.rocm, so the CUDA Dockerfile is not the path for those cards.
One build note comes from pyproject.toml and matters if you touch the Megakernel. Its CUDAExtension links against torch's C++ libraries, and uv's default isolated build environment resolves torch from PyPI instead of the cu128 index, producing an ABI-incompatible .so with an undefined torch symbol on import. The file's remedy is a two-pass install starting with uv sync, and the comment states isolation must be skipped.
Where Lucebox is the wrong tool
The first limitation is coverage. The engine is not tied to one reference card, but it is tied to a finite list of them. The Dockerfile comment says Thor and GB10 prebuilt-image coverage is "intentionally omitted," and pre-Turing architectures sm_60, sm_61, sm_70 and sm_72 are excluded because dflash's BF16 and WMMA paths have no fallback below sm_75. If you are running a P100 or a V100, this is not a build-flag problem; the kernels assume instructions your card does not have.
The second limitation is the model matrix. The supported models and drafters table is short and specific, and several entries require conversion work. Qwen 3.8 27B needs the DFlash2 source converted to Q8_0 before it is usable. If your model is not in that table, you have no published drafter, and speculative decoding without a matching drafter is not what this server is for.
The third is the nature of the published results. The README's optimisation table lists a single measured setup per row. PFlash plus KVFlash at 6.1 times prefill, 411 s to 67.3 s, is Laguna XS 2.1 33B at 256K on RTX 3090. The 208.1 tok/s average for DFlash2 is Qwen 3.8 27B on one R9700. None of these numbers generalise to a different model on the same card, and the README does not present them as if they do. Treat every figure as a configuration result, not a property of the engine.
The fourth is operational. There are no retrieved releases, so installation means building from source or from the local image you bake yourself. The Makefile calls the targets pre-release and favours simplicity over portability. There is no documented rollback path in the README, and no upgrade procedure. For a production serving deployment, that is a real gap.
Lucebox compared with llama.cpp
llama.cpp is the obvious reference point, and the README invites the comparison directly: one entry reports 6.4 times speedup versus Lucebox autoregressive and 3.8 times versus llama.cpp with the same drafter. That second number is the informative one, because it holds the drafter constant and isolates the engine.
The difference in approach is where the optimisation lives. llama.cpp optimises across a broad range of backends and quantisations so that one build runs on many devices. Lucebox compiles kernels for one architecture at a time and pairs each with a named model and drafter. The Dockerfile's architecture list and the per-arch build cost are the direct consequence: a fat binary covering six architectures takes longer to build, and the Makefile tells you to cut the list.
The trade shows up in portability. With llama.cpp, you copy a binary or pull an image and run whatever GGUF you have. With Lucebox, you need a supported device, a published drafter for your model, and often a conversion step. In exchange, the README reports decode speedups over llama.cpp on the same drafter for the configurations it measured. If you move to a model outside the supported table, that advantage has no published evidence behind it, and llama.cpp's broader coverage becomes the more useful property.
Licence and the cost of keeping up
Lucebox is Apache-2.0, declared in both the LICENSE file at the repository root and the license field in pyproject.toml. That is a permissive licence, and it is the same licence llama.cpp uses, so mixing the two in a product does not create an obvious licensing conflict. This is not legal advice; if you are redistributing a built image or a modified kernel, read the licence text and the third-party dependency licences yourself, particularly the torch and CUDA components the Dockerfile pulls in.
The upgrade cost is the part to weigh before adopting. The project last received a push on 2026-09-10, so it is current. But there are no retrieved releases, and the optimisations are per-model and per-device. A new model version, a new ROCm or CUDA release, or a driver change can invalidate a tuned configuration, and the README shows no compatibility matrix beyond the runtime column in the tested machines table (ROCm 7.2 for RDNA4 and RDNA3.5, ROCm 6 or later for RDNA3, CUDA 12 or later for Ampere and Blackwell). Budget for rebuilding the image when you change any of those, and for re-verifying the drafter pairing rather than assuming it still holds.
The Python side is lighter. The workspace requires Python 3.12 through 3.12.x and declares dependencies on lucebox-dflash and pflash, with an optional megakernel extra and a dev extra of pytest, mypy and ruff. The lint configuration in pyproject.toml is explicitly staged: only harness and scripts are included, and the comment notes that server-internal and optimisation Python carries pre-existing style debt. That is a fair signal about where contributor attention has gone.
Editorial conclusion
Adopt Lucebox if you own one of the tested machines, specifically an R9700, Strix Halo, RX 7900 XT or XTX, RTX 3090 or RTX 5090, and you are willing to run the exact model and drafter pair the project publishes for that target. Skip it if you need a portable server that runs the same way across many GPUs, or if your workload is prompt-heavy at long context, where the published prefill numbers apply only to the Laguna XS 2.1 33B configuration. Before committing, verify three things: that your GPU appears in the tested machines table, that a published GGUF and drafter exist for your model, and that your ROCm or CUDA version matches the runtime column, since the README lists ROCm 7.2 for RDNA4 and RDNA3.5 and CUDA 12 or later for Ampere and Blackwell.
Frequently asked questions
What is Lucebox?
Lucebox is a C++ LLM inference server for speculative prefill and decoding on consumer GPUs and APUs. Its own project metadata describes it as open LLM inference rewritten by hand for one specific chip at a time, and each optimisation is paired with a named model and hardware target.
Does Lucebox work on AMD GPUs?
Yes, for the AMD devices in the tested machines table: RDNA4 gfx1201 (Radeon AI PRO R9700), RDNA3.5 gfx1151 (Ryzen AI MAX+ 395, Strix Halo) and RDNA3 gfx1100 (RX 7900 XT and XTX). The README states HIP builds should target the device's exact gfx architecture, and there is a separate Dockerfile.rocm.
What is Lucebox DFlash?
DFlash is one of the decode-phase speculative optimisations in the supported models and drafters table, with a DFlash Q4 drafter published for Laguna XS 2.1 33B and DFlash2 used for Qwen 3.8 27B on the R9700. The README reports 1.7 times speedup at 256K for the Laguna configuration.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/luce-org-lucebox)