DeepSeek V4 Flash on one DGX Spark: a launcher pinned to 384k context
DeepSeek v4 Flash EXL3 on one DGX Spark
At a glance
- What is it?
- A single node Docker recipe for serving the EXL3 build of DeepSeek V4 Flash 0731 on one DGX Spark, with DSpark speculative decoding and a compressed KV cache. The headline decode number is 44 to 47 tok/s, and the interesting part is how the KV pool is whatever memory the weights leave behind.
- Who is it for?
- This is worth reading if you own a DGX Spark and want one box serving a very long context, and pointless without one: the image is aarch64 only, the first boot pulls about 107 GB of weights, and the configuration assumes you are willing to give up concurrent sequences and disable EarlyOOM. Before running it, read your own boot's KV pool number rather than the 439,622 quoted here, because that figure moves between boots, and budget ten minutes for a full length prefill.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 38 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Tensor parallel 1 on one box, where the official build wants two
The pitch is arithmetic before it is anything else. This launcher serves the 0xSero/deepseek-v4-flash-0731-spark build at 3.0 bpw EXL3 on a single DGX Spark with tensor parallel 1, where the official FP4 build is described as needing TP2 across two Sparks. The serving stack is sparkinfer, formerly called b12x, shipped as a self-contained Docker image tuned for speed and KV cache headroom on one device. Speed comes from DSpark speculative decoding with a K64 draft model, selected with MODE=dspark, running a fixed K5 configuration. Two further defaults matter for concurrency: the nvfp4_ds_mla compressed KV cache with the B12X_MLA_SPARSE attention and MoE backend, and CUDA graph capture tuned to the batch sizes [6,12,24] so concurrent decode stays on captured graphs instead of dropping into eager execution. A decode starvation guard for long prefills is enabled natively.
The KV pool is the memory left over after the weights land
The headline figure, a 439,622 token KV pool, is not a configured constant. It is derived: the utilization budget minus weights, minus the profiled activation peak, minus non-torch overhead. That is why the number moves between boots on the same host, where the values observed were 337,841 on a two sequence boot, then 402,334, 430,909 and 440,461 on one sequence boots, then 439,622 cold on the 384k boot. Two effects dominate. First, the hybrid cache split: the model keeps 128 token sliding window layers alongside full depth global layers, and the split of KV bytes shifts with MAX_NUM_SEQS, so a single sequence boot routes more of the budget into the global cache that defines the reported pool. Second, cold versus warm JIT, worth about 0.6 GiB or roughly 28k tokens, because a boot compiling fresh kernels has less memory free when KV sizing happens. The worst boot observed still cleared MAX_MODEL_LEN with at least 1.14x headroom, and a genuinely bad boot trips the boot-time KV check and stops cleanly under restart: on-failure:1.
Two needle tests at 320k and 370k, both with zero preemptions
The evidence offered for the deep context configuration is two runs on 2026-08-21, each on a fresh boot at about 96% of that boot's MAX_MODEL_LEN. The method is built to defeat prefix caching: a single large user prompt of random word filler, generated so that cached_tokens is 0 and every token needs fresh KV, a secret passphrase planted near token 20, and the recall question at the very end, with temperature 0, thinking disabled through chat_template_kwargs and max_completion_tokens 128. Run 1 used MAX_MODEL_LEN 334,000, a 402,334 token pool and a 320,037 token prompt, 1.33 MB, about 80% of the pool, and returned the needle XQ-7741-BLUE exactly with finish_reason stop and 0 preemptions in 517 seconds. Run 2 used the now default 384,000, a 439,622 token pool and a 370,104 token prompt, 1.54 MB, about 84% of the pool, and returned ZK-9931-AMBER exactly, again stop and 0 preemptions, in 594 seconds. After both, /health returned 200 and KV usage was back to 0%.
Prefill decays from 1,024 tok/s to a few hundred past 300k
The 44 to 47 tok/s decode figure is the steady state, and the prefill side is where the time goes. Throughput starts near 1,024 tok/s at the beginning of a request and falls to roughly 350 to 614 tok/s once more than 300k tokens have accumulated, which puts a full length 384k prefill at about ten minutes end to end. The two stress runs are consistent with that shape: effective prefill of about 630 tok/s and about 625 tok/s, producing 517 and 594 second totals. The reason the long runs are trustworthy at all is a bug fixed upstream in the patched kernel directory, image-patch/sparkinfer/. Before the fix the NVFP4 dual-cache prefill path returned NaN on any prompt of 7 tokens or more, so a 320k token test would have been impossible rather than slow. What you give up for the depth is concurrency: MAX_NUM_SEQS is pinned to 1 in the deep context defaults, so this configuration is built for one long conversation rather than several short concurrent ones.
Four env values define the deep context default, changed 2026-08-21
start.sh now serves what the file calls the deep-context NVFP4 config, and the switch happened on 2026-08-21. Four values carry it. KV_RECORD=stock432 selects native 432 byte records. GPU_MEMORY_UTILIZATION=0.94 claims 94% of the 128 GiB unified memory. MAX_MODEL_LEN=384000 sets the context ceiling. MAX_NUM_SEQS=1 accepts a single sequence, and with DSpark K5 healthy the resulting pool is the 439,622 tokens quoted above. The reason for the change is stated plainly: the NVFP4 dual-cache prefill bugs behind the previous either/or configuration are fixed, and the write-up of that debugging is kept local rather than committed, so the repository gives you the configuration without the postmortem. Anyone auditing the numbers has to take the table at face value, since the detailed account of what went wrong is not in the tree.
EarlyOOM has to be disabled or it kills a healthy server
One host level change is mandatory rather than optional. EarlyOOM, if the host has it, has to be turned off:
sudo systemctl disable --now earlyoomThe reasoning is specific: the server intentionally holds about 94% of the 128 GiB unified memory, so a user space OOM killer cannot distinguish a healthy server from a leak and may kill it mid-serve. Beyond that, the hardware floor is one DGX Spark with the GB10 GPU at SM121 and at least 128 GiB of unified memory, with GPU passthrough into Docker through the NVIDIA Container Toolkit. The operating system requirement is Linux aarch64 on DGX OS, and the runtime image is aarch64 only, which rules out running it on an x86 workstation no matter how much memory the host has. Software is Docker Engine with Compose v2, curl, and roughly 110 GiB of free local disk for the weights.
start.sh writes the compose file, then launches it
There is no compose file in the repository. start.sh generates compose.yml on the first run and rewrites it on every launch, so the file you would want to read is generated output rather than something to edit or commit. Launching is one script:
./start.sh
./start.sh --no-waitThe first boot is deliberately slow and does five things in order: pulls the image, downloads about 107 GB of weights locally into ./hf-hub, coalesces TP4 to TP1 losslessly, builds the K64 draft model, and captures CUDA graphs. The container is marked healthy only once the OpenAI-compatible endpoint actually responds, so a healthy status is a real signal rather than a container being up. The download is entirely local with no remote host involved, and no HuggingFace login is needed because the repository and image are public; an optional HF_TOKEN can be set in start.sh or through the environment for rate limits or a private repository. An opt-in LAN mode reuses one copy of the weights across machines over an SSHFS share, with sshfs and fuse3 the only extra requirements.
Kernel fixes arrive as read-only bind mounts over image-patch
The repository is small at the top level: start.sh, stop.sh, download.sh, and three directories, files/, image-patch/ and the usual metadata. What matters is image-patch/, which holds two upstream kernel backports applied as read-only bind-mounts over the image rather than baked into a fork, and it is the directory that carries the NVFP4 dual-cache prefill fix described above. The backend names in the configuration map onto the same code: nvfp4_ds_mla for the compressed cache layout and B12X_MLA_SPARSE for the attention and MoE path, with the b12x name kept only as the former name of the stack. The last push to the repository was on 2026-08-24, it is not archived, and it publishes no GitHub releases, so there is no tag to pin and no changelog beyond the README. The README also carries a ko-fi link beside the lab's account, which is worth noting when you are deciding how much of this configuration you intend to reproduce yourself.
Editorial conclusion
This is worth reading if you own a DGX Spark and want one box serving a very long context, and pointless without one: the image is aarch64 only, the first boot pulls about 107 GB of weights, and the configuration assumes you are willing to give up concurrent sequences and disable EarlyOOM. Before running it, read your own boot's KV pool number rather than the 439,622 quoted here, because that figure moves between boots, and budget ten minutes for a full length prefill. Verify EarlyOOM is off first, since a healthy server holding 94% of unified memory looks exactly like a leak to a user space OOM killer.
Frequently asked questions
Can I run DeepSeek V4 on DGX Spark?
On one DGX Spark, yes, with this launcher. It serves the DeepSeek V4 Flash 0731 build in EXL3 at 3.0 bpw with tensor parallel 1, where the official FP4 build is described as needing TP2 across two Sparks. Reported decode is 44 to 47 tok/s at 384k context.
What operating system does the DGX Spark run?
Linux aarch64 on DGX OS is the requirement, and the runtime image is aarch64 only. GPU passthrough into Docker comes from the NVIDIA Container Toolkit, and the hardware floor is one Spark with the GB10 GPU at SM121 and at least 128 GiB of unified memory.
Is DGX Spark using unified memory?
Yes, 128 GiB of it, shared between CPU and GPU on the GB10 at SM121. That shared budget is why GPU_MEMORY_UTILIZATION is set to 0.94 and why the server holds about 94% of it, which is the stated reason EarlyOOM has to be disabled on the host.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/miaai-lab-deepseek-v4-flash-one-dgx-spark)