# Qwen3.8 27B on DGX Spark: three engines and no release tag

> MiaAI-Lab's repository serves Qwen3.8-27B under SGLang on a 128 GB DGX Spark with three swap-in speculative decoding modes, and every launch flag is pinned to a number measured on that box rather than copied from the cookbook. The DFlash2 path depends on an upstream development image held in place by a digest, because no released tag contains the commit.

**MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark** — Qwen3.8 27B on SGLang for DGX Spark

- Repository: https://github.com/MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark
- Website: https://x.com/MiaAI_lab
- Stars: 464 · Forks: 52
- Language: Shell
- License: MIT
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/miaai-lab-qwen3-8-27b-sglang-dgx-spark

## The three engines disagree, and which one wins depends on the probe

There are three swap-in serving modes and the README refuses to declare a winner, which is the most useful thing about it. On code, DSpark and DFlash2 are both faster than MTP. Against MTP, DSpark gives the essay back and DFlash2 does not, so DFlash2 trades long-form generation for code throughput. On everyday chat the same streamed probe can come out a DFlash2 win once tokens are counted the right way, which is a reminder that a chat benchmark measured on a streaming probe and a code benchmark measured on completion can point in opposite directions. The quick start comments carry the numbers: DSpark at roughly 51.5 on code, 23 on default chat and 18 on a long essay; MTP at 34.5, 21 and 24; DFlash2 at 50.9 on code, 25.4 on essay and 29 to 67 on chat under streaming. A bf16 base for DFlash2 is present but explicitly unbenchched on this box. The stated policy is that all measured numbers, ranges, counting notes and caveats live in one section, with nothing repeated elsewhere in the document, which is a discipline worth copying because a tuning README that repeats a number in three places is a tuning README you cannot update. The reason this repository can publish numbers at all is that it owns the hardware. A single DGX Spark with a fixed core topology, a fixed memory configuration and one model means the same measurement can be repeated after an upgrade to the serving stack, which is what turns a benchmark into a regression test. Everything the scripts pin downstream is derived from that, so the numbers and the flags are two halves of one artefact rather than a table someone pasted in next to a script someone else tuned.

## A digest pin stands in for a release that does not exist

The DFlash2 path needs an image that has never been released, and the repository is explicit about why. `start-dflash.sh` runs the official multi-arch `lmsysorg/sglang:dev-qwen38-27b-dflash2`, the tag the SGLang cookbook maps to DGX Spark, pinned by its index digest beginning `sha256:616a3e97`, corresponding to upstream build `5f55db35e` on a branch named `dflash2-pin-1cf2b8c-nccl` dated 22 August 2026. It is pulled from Docker Hub on first run at about 14 GB compressed, with no git step and no local build. The reason no released tag is used is stated directly: no released tag ships DFlash2 yet, because version 0.5.18 was tagged afterwards but does not contain the commit. So what you are running is a CI build of an upstream development branch from the same publisher as the other image, frozen by digest rather than by version. That image carries DFlash2 itself and the quantized head selector, which is what makes both NVFP4 exports work including the packed FP4 head. The earlier approach, a self-built image with its own patch directory, was retired on 5 September 2026.

## YaRN and the DSpark draft configuration cannot both be on

Long context is available but it costs you the DSpark path, and the sample config says so in a warning rather than letting you discover it. YaRN rope scaling is required for any context above the native 262,144 tokens: off by default, only valid at or below that, and turned on for 500K, 786K or 1M. Two behaviours are easy to trip over. YaRN switches on implicitly at exactly 1,000,000 even if you set it to zero, so an explicit 0 is not a guarantee at that length. And the scaling factor is derived for you as the target length rounded against the native length, which gives 2.0 at 524,288, 3.0 at 786,432 and 4.0 at 1,000,000, where 2.0 and 4.0 are the model card's validated values and intermediate factors work but sit outside the tested points. The consequence is the warning: YaRN is not compatible with the DSpark speculative draft alternative on this SGLang build, so if you switch the speculative algorithm to DSpark you should keep YaRN off and context at the native length. DFlash2 shares the same draft configuration leak.

## Three layers of configuration, and a restart to apply any of it

Configuration is a plain `VAR=value` file read by the start scripts, and the precedence is written out rather than discovered. Highest first come shell environment variables already exported when you run the start script, then values in `./.env`, and last the inline defaults inside the start script itself. So exporting a variable in your shell beats the file, and the file beats the built-in default. The sample ships the native configuration, with YaRN off, a context length of 262,144 and ten concurrent requests, which means a fresh clone serves a 262K context with rope scaling off and ten concurrent requests after a single copy command. The detail that saves an afternoon is the next one: changes only take effect on the next launch, so editing the file while a server is up does nothing until you stop and start again, and `./stop.sh` stops whichever engine happens to be running. The whole loop is short enough to fit on screen:

```bash
# 1. Copy the sample config once (creates ./.env if you don't have one)
cp .env.sample .env

# 2. Start the server
./start-dspark.sh

# 3. Use it
curl http://127.0.0.1:8888/v1/models

# 4. Stop it
./stop.sh
```

Two comments in the script header are worth carrying with you. One notes that DSpark and DFlash2 both cannot use YaRN or a context above 262,144 on this build, the DFlash2 case being the same draft configuration leak. The other notes that a DFlash2 run against a bf16 base is unbenchched on this box, so those numbers are not a measurement you can lean on. The checkpoint selector is the other variable worth reading: four options, the NVFP4 repository with a dense BF16 head as the default, the packed FP4 variant, and full FP8 and BF16 alternatives.

## Ten big cores, and the scheduler is kept off the little ones

The hardware requirement is specific: an NVIDIA DGX Spark, GB10, aarch64, SM121, with 128 GB of unified memory, plus Docker with the NVIDIA Container Toolkit working well enough for a GPU passthrough run. The tuning choices are where the box's asymmetry shows. The launch flags are pinned to GB10's ten 3.9 GHz Cortex-X5 cores with an explicit cpuset, so the scheduler and tokenizer never land on the 2.8 GHz A725 efficiency cores, and the measured effect on decode is two to seven percent. Everything else follows the same pattern of overriding the cookbook with a local measurement. The GDN state runs in bf16 because the cookbook's float32 was three percent worse. Memory fraction is 0.90. Chunk size is 8192. DSpark gets seven blocks of eight draft tokens. Torch compile and decode graphs are on. The KV cache is FP8 in the `fp8_e4m3` format for roughly a halving of KV memory, and it reuses the calibration scales from the NVFP4 checkpoint rather than needing its own calibration pass. The state pool itself is sized from the concurrency setting at four state slots per concurrent request.

## The checkpoint arrives inside the container, about 24 GB

There is no download step in the documentation because there is no download step in the design. The container pulls the checkpoint into `./.cache/huggingface` on first start, roughly 24 GB for the default NVFP4 repository with the dense BF16 head, and about 1.7 GB smaller on disk for the packed FP4 twin. The size breakdown is given so you can tell whether the pull is behaving: the cookbook cites around 16.5 GB for the NVFP4 language model weights alone before the MTP head, and the dense head adds about 1.7 GB on disk or 3.2 GB at runtime. A Hugging Face token defined in `~/.bashrc` is picked up automatically and buys higher rate limits, which is the difference between a fast first boot and a throttled one. The images are listed separately from the checkpoint because there are two of them: the model-specific build used for MTP and DSpark, and the separate digest-pinned development image used for DFlash2. The only CLI prerequisites are `docker` and `curl`.

## Idempotent scripts, a privileged container, and a poll instead of a health check

The start scripts share one shape, and the details are the ones that matter in practice. They launch with `docker run -d` on the host network with a 32 GB shared memory setting and running privileged, stream logs into `.sglang.log`, record the container ID in `.sglang.pid`, and then poll the models endpoint until the server answers. Polling a real endpoint rather than trusting the process is why a cold start on a fresh box does not produce a race where you query a server that has not loaded weights. All of them are idempotent: a container that is already running produces a message and an exit rather than a second one, and a stopped container is removed before a new one is created. Monitoring is on by default with Prometheus metrics and cache reporting enabled. The MTP path uses EAGLE speculative decoding with its step, top-k and draft settings pinned, and the DSpark path swaps the algorithm while keeping the rest. One engine-specific detail is called out rather than left implicit: the specification verification window is a separate engine-side buffer, not part of the state pool, which is verified in the build's cache configurator. That distinction is the kind of thing that silently halves your concurrency if you get it wrong, since sizing the pool as if the verify window lived inside it leaves you short at exactly the moment you raise the request count. The rest of the repository is small and tells you where the measurements live: a `bench/` directory, a `docs/` directory, a changelog and a licence, alongside the three start scripts and the stop script. Nothing here builds an image any more, which is the clearest signal of how the DFlash2 story was resolved: the patch directory and its builder are gone from the tree and the answer is a digest.

## Conclusion

Use this if you have an actual DGX Spark and want Qwen3.8-27B serving without a week of flag archaeology, because the value is not the scripts but the measurements behind them: the cpuset, the bf16 GDN state, the memory fraction, the chunk size and the draft token count were all tuned on the box rather than inherited from a cookbook written for different hardware. Two things to check before you commit. Which mode matches your workload, because the three disagree depending on whether you are generating code, chatting or writing a long essay, and the code and essay probes rank differently. And whether you need more than 262K context, because YaRN and the DSpark draft configuration are mutually exclusive on this build, so long context and DSpark are not both available to you. Also treat the DFlash2 image as what it is: a digest-pinned CI build of an upstream development branch, frozen rather than released.

## FAQ

### What does this repository give me for a DGX Spark?

Ready-to-run scripts that serve Qwen3.8-27B under SGLang in Docker on an NVIDIA DGX Spark, or GB10, with 128 GB of unified memory. The checkpoint is pulled into ./.cache/huggingface on first start at roughly 24 GB, and there is no separate download step.

### Which of the three serving modes should I use?

It depends on the workload, and the repository declines to pick one. On code, DSpark and DFlash2 both beat MTP. Against MTP, DSpark holds up better on long essays while DFlash2 does not. Everyday chat comes out a DFlash2 win on the streamed probe once tokens are counted correctly.

### How do I get more than 262K context on DGX Spark?

Set YaRN to 1 and raise the context length, which is derived into a scaling factor against the native 262,144. The model card validates 2.0 at 524,288 and 4.0 at 1,000,000, intermediate factors work but sit outside the tested points, and YaRN turns on implicitly at exactly 1,000,000.

### Why does the DFlash2 image use a development tag?

Because no released tag ships DFlash2. Version 0.5.18 was tagged afterwards but does not contain the commit, so the image is the official multi-arch dev tag pinned by index digest, a CI build of an upstream development branch frozen by digest rather than by version number.

### Can I use YaRN together with DSpark speculative decoding?

Not on this SGLang build. The sample config carries a warning that YaRN is not compatible with the DSpark speculative draft alternative and says to keep rope scaling off with context at the native 262,144 if you switch the speculative algorithm to DSpark. DFlash2 has the same draft configuration issue.

## Sources

- [Issues](https://github.com/MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark/issues)
- [License: MIT](https://github.com/MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark/blob/main/LICENSE)
- [MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark on GitHub](https://github.com/MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark)
- [Project website](https://x.com/MiaAI_lab)
- [README](https://github.com/MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/miaai-lab-qwen3-8-27b-sglang-dgx-spark
