# Two DGX Sparks serve one Qwen checkpoint, and the docs say so

> MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks is a set of shell scripts that bring up vLLM across two DGX Spark nodes with tensor parallel 2, expert parallelism and 3 speculative tokens. Its own README admits that one of its headline measurements came from a container running an older config value than the file it ships.

**MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks** — Qwen3.8-Flash-Next-NVFP4-vLLM-Dual-DGX-Spark

- Repository: https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks
- Website: https://x.com/MiaAI_lab
- Stars: 396 · Forks: 47
- Language: HTML
- License: AGPL-3.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/miaai-lab-qwen3-8-flash-next-dual-dgx-sparks

## Both nodes hold the whole checkpoint by default

The hardware assumption is narrow and stated as a requirement list: two DGX Spark nodes with GB10, 128 GB of unified memory and sm_121, connected over ConnectX RoCE or InfiniBand, with passwordless SSH between them and Docker on both. Anything missing from that list is something you have to solve before the scripts are useful.

Disk is the requirement people miss. The prerequisite is roughly 126 GiB free on each node, because by default both nodes keep their own copy of the checkpoint in `~/.cache/huggingface`. `start.sh` rsyncs the worker copy from the head once and skips the transfer if the worker already has it. NFS weight sharing is the documented alternative, which removes the worker copy entirely at the cost of the head cache being read over the network on every load.

So the default shape is two full copies and a one-time sync, and the alternative is one copy over NFS. Which one is cheaper depends on whether the link is fast enough to stream weights, which is a decision the flags push onto you: `--nfs` to share, `--no-nfs` to force rsync even when `NFS_SHARE=true` sits in `.env`.

## Every lane launches a container called vllm-fn on the same port

The repository is a pile of shell scripts rather than an application, and the collision between them is stated plainly: `./start.sh` and `./start-fp8.sh` both launch containers named `vllm-fn` on the same port, so you stop first with `./stop.sh` before switching checkpoints.

The scripts at the root are `download.sh`, `start.sh`, `start-fp8.sh`, `start-tp1.sh`, `start-v030.sh` and `stop.sh`, alongside `check-weights.sh` and `verify-weights.py`. Each start script is a different lane over the same two nodes: the stock NVFP4 launch, the official FP8 checkpoint, a single node TP1 path with its own `tp1/` directory, and a vLLM 0.30 path. They share the container name and the port, which is why the workflow is stop, start, stop, start rather than two services on two ports.

The flags on those scripts divide into two groups. Distribution flags are `--no-download`, `--no-launch`, `--launch`, `--nfs` and `--no-nfs`, which together let you fetch weights without serving, distribute without serving, or serve without touching the network. `ABLIT=1` is the odd one out because it is an environment or `.env` flag rather than a command line flag, and it selects a gated checkpoint.

## Two of the seven steps patch files out of the container image

The What Happens list is seven steps, and steps four and five are the reason this repository exists at all.

The PLE patch extracts `ple_layer.py` from the image and writes a patched version to `files/ple_layer_patched.py`, with no image rebuild and a bind mount at runtime. The MXFP8 patch does the same with `modelopt.py`, producing `files/modelopt_patched.py`, and its purpose is specific: it routes the MXFP8 shapes that FlashInfer's `mm_mxfp8` kernel cannot run to the BF16 emulation kernel. Both are also bind-mounted, so the container image stays stock and the fixes live in the repository as ordinary files you can read and diff.

Step six is a preflight. It refuses to launch if another process holds a GPU on either node, and `REQUIRE_IDLE_GPU=false` overrides that. The same step prepares the `FP8_DENSE` and `QSA_PROFILE` bind-mounts. Step seven is the launch itself: the worker at rank 1 starts first, then the head at rank 0 serves on `:8888`, with the head's rendered script kept as `.last_head_launch.sh` for inspection.

## Three checkpoints, and the shard counts are not comparable

Which weights you serve is a single variable in `.env`, and the three documented builds mix different precisions in different places:

```
MODEL_ID="nvidia/Qwen3.8-Flash-Next-NVFP4"    # 133G / 11 shards: BF16 dense + NVFP4 experts + FP8 PLE
MODEL_ID="local-inference-lab/Qwen3.8-Flash-Next-NVFP4"  # 99G / 37 shards: MXFP8 attention + NVFP4 PLE
MODEL_ID="RadixArk/Qwen3.8-Flash-Next-NVFP4"             # 126G / 206 shards: BF16 attention + FP8 PLE
```

The size and shard counts are the part to read carefully. Eleven shards for the nvidia build, thirty seven for local-inference-lab, and two hundred and six for RadixArk, at 133G, 99G and 126G respectively. So a larger download is not a coarser checkpoint here, and neither is a smaller one. Whichever you pick, `./download.sh` pulls that `MODEL_ID` into the head cache, and step two distributes it.

The network settings sit directly above the model setting in the same file. `HEAD_IP` and `WORKER_IP` are literal addresses to edit, `WORKER_USER` is blank by default meaning the same user as the head, `IFACE` names the inter-node ethernet interface and the comment warns that ConnectX ports on DGX Spark often carry np0 or np1 suffixes, and `IB_HCA` is set to `=rocep1s0f0` with a single exact match device and a warning never to list a port cabled to another cluster. `IB_GID_INDEX` is 3, with the note that if 3 does not come up, try 5.

## The FP8 lane has a shorter context because the weights are bigger

The optional FP8 lane serves `Qwen/Qwen3.8-Flash-Next-FP8` instead of the NVFP4 checkpoint, keeping the same two-node launch and the same `.env` for cluster IPs, tensor parallel, expert parallelism, image and ports. The API name changes to `qwen3.8-flash-next-fp8`.

The cost is context. Context is native 262K with YaRN off, and the stated reason is arithmetic: FP8 weights leave too little KV cache for 1M on this kit. Available KV cache is around 500k tokens there, against 3.65M for the nvidia NVFP4 checkpoint with `KV_CACHE_DTYPE=fp8`. That is a difference of more than seven times, so the two lanes are not interchangeable for long context work.

Two naming traps sit in this section. The lane is not `FP8_DENSE=true`, which is the hybrid of NVFP4 experts with FP8 dense projections, and it is not fp8 KV on stock NVFP4. And `ABLIT=1` is ignored when `./start-fp8.sh` sets `OVERRIDE_MODEL_ID`, so a gated checkpoint request silently does nothing on this lane.

## Breakable CUDA graphs are switched off on purpose

The vLLM 0.30 lane is the most measured part of the repository, and the reason it exists is a stability problem rather than a speed one. vLLM 0.30 turns `VLLM_USE_BREAKABLE_CUDAGRAPH` on for this model, and on GB10 that made decode vary by 10 to 18 percent between identical runs and lose 13 to 19 percent at two or more concurrent streams. The lane sets it to 0, and `V030_BREAKABLE_CUDAGRAPH=1` restores the upstream default if you want it back.

The speed claims are given as measurements against the day-0 lane: prose decode of 60.0 against 56.3 tokens per second at one stream, code decode of 441 against 422 at eight streams, and multi-turn time to first token of 0.28 to 0.40 seconds against 0.63, which the README calls twice as fast. A first repeat miss, logged as issue number 62, is fixed.

KV cache defaults to BF16 on this lane at 1.52M tokens with GPU memory utilisation at 0.80. Setting `OVERRIDE_KV_CACHE_DTYPE=fp8` reaches 2.64M tokens through a backport of vllm#55557 in `files/patch_qsa_fp8_kv_v030.py`, at about 5 percent slower decode, and the file carries an instruction to delete it once the image reaches vLLM 0.31 or later. Not supported here: `FP8_DENSE`, `QSA_PROFILE` and the determinism knobs.

## The README's own numbers came from an older config

The provenance note near the top is the most useful paragraph in the file. Every number in the KV cache budget and the default runtime section was read from the running server with `docker logs vllm-fn` and `docker inspect vllm-fn`, so they are measurements rather than estimates. Then it says which measurement: that container runs `GPU_MEMORY_UTILIZATION=0.835`, which is the value `.env.sample` shipped until 26 September 2026, and the file now says 0.80.

So the documented KV numbers were taken under a higher memory utilisation setting than the one you get after cloning, and the CHANGELOG is where the change is recorded. Anyone comparing their own run against the README is comparing against a configuration the repository no longer ships.

A second note starts the same way and stops mid-sentence. On GB10 the page cache shares one unified memory pool with the model, so a checkpoint left resident from a download, an rsync or a previous launch competes with the weights for memory. That is enough to explain why the rsync-then-launch order matters, and why the text does not get to the rest of its argument.

## Conclusion

Fit for someone who has exactly two GB10 nodes, a ConnectX link between them, and a checkpoint whose PLE and MXFP8 layers do not run unpatched on the stock image, since the patches are extracted from the image at runtime rather than baked into a rebuild. A poor fit for a single machine or a growing cluster, because the scripts assume one head and one worker, every lane launches a container with the same name on the same port, and the worker is assumed to have no internet. Before running anything, edit .env for your IPs, interface and HCA, confirm the IB_GID_INDEX works on your ConnectX card, and read the CHANGELOG, since the shipped GPU_MEMORY_UTILIZATION no longer matches the number the README's own measurements were taken under.

## FAQ

### What is a DGX Spark used for?

In this repository a DGX Spark is one half of a two-node vLLM deployment. The prerequisites call for 2 DGX Spark nodes with GB10, 128 GB unified memory and sm_121, connected via ConnectX RoCE or InfiniBand, with Docker on both and passwordless SSH between them, used to serve Qwen3.8-Flash-Next-NVFP4 with tensor parallel 2, expert parallelism and 3 speculative tokens.

### How many DGX sparks can be connected together?

This setup assumes exactly two. One node is the head at rank 0 and serves on port :8888, the other is the worker at rank 1, which starts first, and .env names HEAD_IP and WORKER_IP separately. Nothing in the repository describes a cluster larger than a head and a worker.

### Can DeepSeek V4 Flash run on DGX Spark?

The repository does not address that model. What it serves is Qwen3.8-Flash-Next in three builds selected through MODEL_ID in .env: nvidia NVFP4 at 133G and 11 shards, local-inference-lab MXFP8 attention with NVFP4 PLE at 99G and 37 shards, and RadixArk BF16 attention with FP8 PLE at 126G and 206 shards.

## Sources

- [Issues](https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks/issues)
- [License: AGPL-3.0](https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks/blob/main/LICENSE)
- [MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks on GitHub](https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks)
- [Project website](https://x.com/MiaAI_lab)
- [README](https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/miaai-lab-qwen3-8-flash-next-dual-dgx-sparks
