Self-hosted service
eugr/spark-vllm-docker avatar
eugr/spark-vllm-docker

spark-vllm-docker: packaging vLLM for DGX Spark boxes

Docker configuration for running VLLM on dual DGX Sparks

2,326 stars407 forksPythonMIT

At a glance

What is it?
A community Docker configuration that builds vLLM from source for NVIDIA's GB10 desktop machines, wires multi-node NCCL networking through an env file, and keeps its own patched wheels for the hardware quirks.
Who is it for?
This repository is worth your time if you own DGX Spark hardware and want a serving stack that the community has already compiled and regression tested for GB10, because the alternative is running vLLM's own build yourself against a CUDA image on hardware NVIDIA treats as a side case. It is close to useless without that hardware: the gencode flag, the B12X image tag and the prebuilt wheel releases all name DGX Spark specifically.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 8 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the Dockerfile actually pins

The whole project rests on one Dockerfile, and its build arguments tell you what hardware the thing was written for. The base image is `nvidia/cuda:13.0.2-devel-ubuntu24.04`, and the NCCL gencode flag targets a single architecture:

dockerfile
ARG CUDA_IMAGE=nvidia/cuda:13.0.2-devel-ubuntu24.04
ARG NCCL_NVCC_GENCODE="-gencode=arch=compute_121,code=sm_121"
ARG TORCH_VERSION=2.13.0
ARG TORCHVISION_VERSION=0.28.0
ARG TORCHAUDIO_VERSION=2.11.0
ARG CUTLASS_DSL_VERSION=4.7.0

`sm_121` is the compute capability of GB10, the chip in DGX Spark, which tells you immediately that this is not a portable base image. PyTorch is pinned by argument rather than by a lockfile, and the same pattern applies to TorchVision, TorchAudio and CUTLASS DSL.

Build parallelism is a first-class argument, `BUILD_JOBS` defaults to 16, and it is exported four separate ways because four tools need telling: `MAX_JOBS`, `CMAKE_BUILD_PARALLEL_LEVEL`, `NINJAFLAGS` and `MAKEFLAGS`. The comment above it says why, limiting build parallelism to reduce out-of-memory situations. This is compiling vLLM, and on a desktop-sized machine with unified memory, that is the difference between finishing and not. There is also an empty `FROM scratch AS vllm_source` stage, which is the hook for pointing the build at a checkout you already have via `--build-context vllm_source=/path/to/checkout`.

Recipes are the interface, not a Dockerfile you write

The fastest path in the README is one command, and it is worth understanding what it hides. On a single Spark, you check out the repository on the head node and hand a YAML recipe to `run-recipe.sh`:

bash
git clone https://github.com/eugr/spark-vllm-docker.git
cd spark-vllm-docker
./run-recipe.sh recipes/qwen3.8-flash-next-nvfp4-solo.yaml --solo --setup

`--solo` means one node, `--setup` means do the preparation as well as the launch. The dual-node version swaps in a different recipe and drops the flag:

bash
git clone https://github.com/eugr/spark-vllm-docker.git
cd spark-vllm-docker
./run-recipe.sh recipes/deepseek-v4-flash-vision-exp.yaml --setup

So the design decision here is that a serving setup is a data file in `recipes/`, and `run-recipe.sh` plus its Python companion `run-recipe.py` interpret it. A team running several models ends up maintaining YAML rather than remembering a wall of `vllm serve` flags, which is the right way round. The dual-node path also assumes you have already enabled passwordless SSH between the machines and read `docs/NETWORKING.md`, including its three-node mesh instructions.

The manual path separates image, model and server

The longer quick start splits what the recipe does into three visible stages, which is the better shape if you want to understand what is happening. First the image, then the weights, then the server. Preparation is two commands: the build script plus one `hf-download.sh` call per model, main and draft.

bash
./build-and-copy.sh
./hf-download.sh nvidia/Qwen3.8-27B-NVFP4
./hf-download.sh z-lab/Qwen3.8-27B-DFlash2

The second download is a draft model for speculative decoding, which is why the launch configuration later carries a `--speculative-config` flag naming it. `build-and-copy.sh` also distributes the image to other nodes using `COPY_HOSTS`, so the same script covers both the local build and the cluster push.

The launch is `launch-cluster.sh` wrapping a normal `vllm serve`, which is the useful thing to notice: nothing about the serving API changes, and every flag you already know still works.

bash
./launch-cluster.sh --solo -t vllm-node \
  exec vllm serve nvidia/Qwen3.8-27B-NVFP4 \
    --kv-cache-dtype fp8 \
    --gpu-memory-utilization 0.7 \
    --max-model-len 262144 \
    --max-num-seqs 8 \
    --max-num-batched-tokens 16384 \
    --enable-chunked-prefill \
    --async-scheduling \
    --enable-prefix-caching \
    --load-format instanttensor

`-t vllm-node` selects the regular image tag; the Flash Next and DeepSeek Vision examples use `vllm-node-b12x` instead. The full example in the README also passes `--host 0.0.0.0`, `--port 8000`, `--trust-remote-code`, the speculative config for DFlash2, `--reasoning-parser qwen3`, `--tool-call-parser qwen3_xml` and `--enable-auto-tool-choice`. A 262144 token context with `--gpu-memory-utilization 0.7` is the kind of combination that only makes sense when memory is shared with the CPU, which is again the GB10 constraint.

The env file is the multi-node contract

Cluster configuration lives in a `.env` file, and `.env.example` in the repository root is the template. The comment at the top tells you to copy it and customise. The variables fall into groups, and the grouping tells you what each one controls.

bash
CLUSTER_NODES="192.168.177.11,192.168.177.12"
ETH_IF="enp1s0f1np1"
IB_IF="rocep1s0f1,roceP2p1s0f1"
LOCAL_IP="192.168.177.11"
MASTER_PORT="29501"
CONTAINER_NAME="vllm_node"

`CLUSTER_NODES` is a comma-separated list where the first entry is the head node, and `MASTER_PORT` is the cluster coordination port. `ETH_IF` and `IB_IF` are optional and auto-detected when left unset, which is convenient but worth setting explicitly in production, because picking the wrong RoCE interface produces a cluster that looks healthy and runs at Ethernet speed. `LOCAL_IP` exists for solo mode and for overriding detection.

The `CONTAINER_` prefix is a convention with a rule attached: any variable starting with it, except `CONTAINER_NAME`, becomes a `-e` flag on the container. `CONTAINER_NCCL_DEBUG` set to `INFO` reaches the container as `NCCL_DEBUG=INFO`, and `CONTAINER_HF_TOKEN` is how the Hugging Face token gets in. The template demonstrates the rule with an inline example, which is the sort of detail that saves an afternoon.

The project patches vLLM rather than wrapping it

The `mods/` directory and two patch files, `fastsafetensors.patch` and `fastsafetensors_mxfp4.patch`, are the reason this is more than a Dockerfile. Three specific fixes are described in the README, and they are all platform workarounds rather than feature additions.

The first is a CUDA memory reporting bug under WSL on integrated NVIDIA GPUs: vLLM was replacing CUDA's reported free memory with guest RAM availability, and the fix keeps the reported figure. The second trims unused glibc CPU heap pages after startup garbage collection in API servers and workers, which complements the existing CUDA allocator cleanup before KV cache sizing. It runs after warmup and is included in the exported wheels, and platforms without `malloc_trim` skip it. The third is a runtime default, `VLLM_WSL2_ENABLE_PIN_MEMORY=1`, which you can turn off by passing `-e VLLM_WSL2_ENABLE_PIN_MEMORY=0`.

Autotuning is the opposite choice: `B12X_AUTOTUNE=0` is the default on every GPU architecture, and passing `-e B12X_AUTOTUNE=1` enables it. That is a maintainer making the safe choice the default and asking you to opt in. There is a second Dockerfile, `Dockerfile.mxfp4`, for the MXFP4 model format, and the launches use `--load-format instanttensor` for faster weight loading, so the loading path is treated as part of the configuration surface rather than left to the default.

Prebuilt wheels and the staging trap

Building vLLM from source on a desktop machine takes hours, so the release tags matter. On 2026-09-28 two were published: `prebuilt-vllm-current` carrying wheels built from `0.30.1rc1.dev254+gccfd1cea7.d20260928`, and `prebuilt-flashinfer-current` carrying FlashInfer `0.7.0-85bfa5e4-d20260928`. Both names end in `-current` and both descriptions say DGX Spark only, which is an admission that they are rolling pointers rather than a version history you can pin.

The build script exposes that choice directly. `--use-wheels` builds the runner from precompiled vLLM and FlashInfer wheels and never falls back to compiling anything: if a wheel cannot be downloaded or found locally, the command stops with an error. That is a good default behaviour, since a silent fallback would turn a five-minute build into a multi-hour one without telling you. When you do need source, the flags are `--rebuild-vllm`, `--vllm-repo`, `--vllm-ref` and `--vllm-source-dir` for vLLM, and `--rebuild-flashinfer`, `--flashinfer-ref` and `--apply-flashinfer-pr` for FlashInfer.

One tag in that list is a trap worth knowing about: `staging-current-1787746223` is named "Staging Wheels (Do Not Use)" and describes itself as a temporary staging environment for uploading wheels. Anything matching that pattern should be ignored.

What this will not do for you

The scope is narrow and the README does not pretend otherwise. It is not affiliated with NVIDIA, describes itself as a community effort, and builds a container for a machine that most readers do not own. On any other GPU the gencode flag is wrong and the images have never been claimed to work, so the honest comparison against alternatives is not with another wrapper but with building vLLM yourself: vLLM's own Docker images and install instructions are the general path, and they will run on hardware this repository ignores. `Scitrera/DGX Spark-vllm` also turns up as another Spark-specific image if you want a second reference point.

Within its own scope the README is unusually candid. By default `build-and-copy.sh` pulls `eugr/spark-vllm:latest`, a nightly image, and the maintainers explain that nightly images are tested on several models in both cluster and solo configuration before `latest` moves. Then comes the sentence that matters: vLLM is a rapidly developing platform and some things may break. `--exp-b12x` is explicitly experimental and pulls a separately tested `eugr/spark-vllm-b12x:latest`, which is the flag to avoid in production.

What you get in exchange is a tested starting point rather than a maintained service: nightly verification across models, documented interconnect setup, and hardware fixes you would otherwise have to rediscover on GB10.

Editorial conclusion

This repository is worth your time if you own DGX Spark hardware and want a serving stack that the community has already compiled and regression tested for GB10, because the alternative is running vLLM's own build yourself against a CUDA image on hardware NVIDIA treats as a side case. It is close to useless without that hardware: the gencode flag, the B12X image tag and the prebuilt wheel releases all name DGX Spark specifically. Two things to check before you commit a service to it. First, the project offers no stability promise for the upstream project it wraps, and the README says outright that things may break as vLLM moves. Second, image provenance is a nightly DockerHub tag rather than a versioned release, so what you run depends on when you ran `build-and-copy.sh`. Pin a recipe, read `docs/NETWORKING.md` for the interconnect, and keep the manual `docker run` path in mind as the fallback.

Frequently asked questions

Where can I find spark vLLM docker recipes?

In the repository's `recipes/` directory, as YAML files that `run-recipe.sh` reads. The README's quick start points at `recipes/qwen3.8-flash-next-nvfp4-solo.yaml` for a single Spark and `recipes/deepseek-v4-flash-vision-exp.yaml` for a dual-node cluster, and also references `recipes/qwen3.8-27b-nvfp4-dflash2.yaml` for the Qwen3.8-27B example.

How to run vLLM on nvidia DGX Spark?

Clone the repository on the head node and run `./run-recipe.sh` with a recipe and the `--solo` flag for one machine, adding `--setup` to build the image and fetch the model. For a cluster, drop `--solo` and connect the machines with passwordless SSH first, as described in `docs/NETWORKING.md`. The manual alternative is `build-and-copy.sh`, then `hf-download.sh` per model, then `launch-cluster.sh` wrapping `vllm serve`.

Can DGX Spark be used for inference?

Yes, and that is what this repository is for. It builds and launches vLLM on DGX Spark from a single node up to multi-node clusters over Ray or vLLM's native PyTorch distributed mode, with InfiniBand and RDMA support through NCCL. The launch flags in the README, such as a 262144 token context with `--gpu-memory-utilization 0.7`, are tuned for the GB10's shared memory.

What is the best LLM model for DGX Spark?

The repository does not rank models, it ships tested recipes for them: Qwen3.8-27B-NVFP4, Qwen3.8 Flash Next in NVFP4, and a DeepSeek V4 Flash Vision configuration. The README says nightly images are built and tested on several models in both cluster and solo mode before `latest` is advanced, and that the pipeline's model selection will keep expanding, so the README's own answer is to use a recipe rather than to trust a leaderboard.

Official sources

  1. eugr/spark-vllm-docker on GitHub
  2. Issues
  3. License: MIT
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/eugr-spark-vllm-docker.svg)](https://hysenlabs.com/projects/eugr-spark-vllm-docker)