# ds4-on-spark serves DeepSeek-V4-Flash from one DGX Spark

> Entrpi/ds4-on-spark is an installer for a batched fork of antirez/ds4 that runs DeepSeek-V4-Flash entirely on a single DGX Spark, with continuous batching, DSpark speculative decode and a memory governor that reports a byte shortfall instead of crashing. The fork is pinned at v0.6.5 and the model is an 81 GiB asymmetric quant.

**Entrpi/ds4-on-spark** — Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full continuous batch support

- Repository: https://github.com/Entrpi/ds4-on-spark
- Stars: 405 · Forks: 30
- Language: Shell
- License: MIT
- Published: 2026-09-20 · Updated: 2026-09-20 · Language: en
- Canonical page: https://hysenlabs.com/projects/entrpi-ds4-on-spark

## One command verifies, builds, downloads, smoke tests, serves

The whole install is a single piped command on a DGX Spark that already has CUDA 13:

```bash
curl -sSL https://raw.githubusercontent.com/entrpi/ds4-on-spark/main/install.sh | bash -s -- --start
```

Seven things happen inside it, and they are listed rather than implied. It verifies the host first: aarch64, GB10 with SM121, CUDA 13 and at least 120 GiB free on disk. It clones the `Entrpi/ds4` fork at the pinned release tag, currently `v0.6.5`, into `~/code/ds4` or `$DS4_SRC_DIR`. It builds `ds4`, `ds4-server` and `ds4-bench` with `CUDA_ARCH=sm_121`, which the README puts at about eight seconds.

Then the weights come down, the smoke test runs, the launcher is installed to `~/.local/bin`, and the server starts on port 8000 with a context of 524288, which is 512K. The smoke test is deliberately trivial: it asks for the capital of France and asserts that Paris appears in the output, so a broken build fails before anything is left running.

You can look before you leap, since the same script answers `--help` and prints without doing any of it.

## The model is an 81 GiB asymmetric quant with no MTP head

The engine and the model are separate downloads and the README is precise about both. The engine is the `Entrpi/ds4` fork pinned at `v0.6.5`, described as a major feature fork of `antirez/ds4` (DwarfStar 4) that builds the upstream engine into a batched multi-request server for Blackwell CUDA. Upstream stays the architectural foundation and the model recipe, and this repository's job is pinning and packaging the fork.

The model is the DeepSeek-V4-Flash-0731 GGUF from `antirez/deepseek-v4-gguf`, about 81 GiB, and it is quantised asymmetrically by projection type: IQ2_XXS on the routed gate and up projections, Q2_K on the routed down, Q8_0 on everything else that is dense, and F16 for the compressor and indexer. The comment next to those weights is worth internalising: FP8 inside ds4 is a runtime KV cache format, not a stored weight format.

A second, smaller download is the DSpark drafter at about 6.5 GiB from a different repository on the Hub. The 0731 checkpoint has no MTP head, which is why no MTP file is fetched and why DSpark is the only speculation available on this model.

## Context stopped being a memory decision in v0.6

The launcher exists because the context argument stopped meaning what it used to. `-c` used to be a way to reserve memory, and in v0.6 unused context is demand-mapped instead, so a deep setting costs almost nothing until a request actually fills it. That is why the default launch context is 512K rather than something smaller, and why raising it is not the trade it looks like.

The claim is qualified rather than asserted. The model's full 1M window, set with `-c 1048576`, is described as qualified to the last token, with the deepest proof being a 1,029,340-token prompt where a needle placed at 99.9 percent depth was retrieved exactly. That is a specific test rather than a general assurance, and it is the kind of number worth asking to see the receipt for.

The memory governor goes further still. It is measured to hold up to 3 million tokens of active context resident and warm at once on a single Spark, and 2.26 million with every shipped default intact. The gap between those two numbers is the honest part: the headline figure assumes configuration you have to ask for.

## The memory governor reports a byte shortfall instead of crashing

The v0.6 line is described as memory truth, and the mechanism is that the engine keeps a single account of its memory and makes every decision from it. Concretely, requests are charged what they will actually use, and every floor and margin in the plan is derived from a measurement rather than being a constant.

Two behaviours follow from that. Idle memory returns to the pool after a couple of minutes of quiet rather than staying pinned. And when something genuinely does not fit, the result is a clear refusal that says how many bytes were missing and whether retrying will help, never a crash. For a long ingestion job that distinction is the whole point, since a refusal with a byte count is actionable and an out of memory abort is not.

The engine also audits itself while serving. An idle reconciliation line checks the box's raw memory drop against the ledger and logs the residual, which is a self check on the accounting rather than a claim about it.

## A request pays for its prompt plus its decode budget

Two numbers decide how much context a configuration can hold, and the README explains both.

The first is what one request costs: its prompt plus its decode budget. The decode budget is `max_tokens`, and the assumption made when a client omits it is 32768. That default is worth knowing about, because a client that sends a short question without a limit is charged against a 32K decode allowance in the accounting.

The second is how many requests can hold context at once, which is expressed as the bank count. Since requests are charged what they use rather than what they might, a large bank count with small requests and a small one with large requests behave differently, and the two knobs interact with the demand-mapped unused context described earlier.

Against upstream the fork claims 2.4 to 3.3 times the prefill throughput and 1.33 to 1.47 times the decode speed across the full 2k to 128k context range, plus continuous batching and prefix caching with disk-persisted KV banks that survive restarts. The README says a delta table itemises exactly what changed and how each claim was measured.

## DSpark is the only speculation on this checkpoint

Speculation has three switches and only one of them is available here. The full stack is the default, described as lossless, quench-governed and armed at every depth, with the break-even guard calibrated to the 0731 model identity rather than left at a generic threshold. `--no-dspark` serves plain continuous decode instead and skips the drafter download entirely, which is the option to pick when you want the 6.5 GiB back.

`--with-mtp` is the odd one. MTP-2 speculation, worth a modest 1.08 times, applies only to the legacy generation's base and MTP files, so it does nothing for the 0731 checkpoint this installer downloads. On 0731 the fallback ladder is DSpark and then plain decode, and there is no third rung.

The fork's agent-facing hardening sits alongside this. Client `reasoning_effort` fields compat-map to the prefix-free level, and deep tool conversations carry a protocol reminder by default, which the README calls the measured fix for agent tasks dying to prose turns at depth. Two optional levers can render scaffold-echoed reasoning out of the prompt or resample a prose turn once.

## curl to bash, and a --force flag that skips the check

The install path is a script fetched over HTTPS and piped into bash, which is the normal shape for this kind of one-liner and also the part to think about. The README gives you the option to print it instead: the same command with `--help` shows the interface without touching the machine.

The host check is the guard rail, and `--force` removes it. Without it the installer refuses unless it finds aarch64, GB10 with SM121, CUDA 13 and at least 120 GiB free disk. That disk figure is about the weights rather than about runtime memory: the 81 GiB checkpoint and the 6.5 GiB drafter have to land somewhere before the server starts.

Other overrides exist for machines that differ from the reference one. `--cuda-arch sm_120` builds for RTX PRO 6000 and 5090 class Blackwell, and datacenter B200 and B300 are `sm_100` and described as untested. `--no-download` reuses an existing GGUF, `--src-dir` and `--gguf-dir` move the two directories from their defaults of `~/code/ds4` and `~/gguf`, and `--ctx` and `--port` set the launch values.

## One batch serves four API shapes, foreground only

Connecting an agent is the reason for the batch. `ds4-server` speaks four APIs from one continuous batch: OpenAI chat completions, OpenAI completions, the OpenAI Responses surface that Codex uses, and Anthropic Messages, the surface Claude Code expects. One model, one pool of KV banks, four request shapes.

The three launcher examples show the shape of the arguments:

```bash
ds4-serve                              # full stack, ctx 524288, 127.0.0.1:8000
ds4-serve -c 1048576                   # the model's full 1M window
ds4-serve -c 32768 --host 0.0.0.0      # smaller context, reachable from LAN
```

Anything you pass goes straight to `ds4-server` and overrides the defaults, with `ds4-server --help` listing the rest including `--cors`. The default bind is loopback, so the third line is the one that changes who can reach the model.

One operational detail is stated plainly: the process runs in the foreground, and supervising it with nohup, systemd or tmux is left to you. The repository tree matches the parts, with `bin/` for the launcher, `bench/` for benchmarking, `codex/` for the Codex side, `docs/`, `scripts/` and a `LATEST` file at the root.

## Conclusion

Fit for someone with a GB10 who wants DeepSeek-V4-Flash served locally against the OpenAI and Anthropic API shapes without hand building an engine, since the installer verifies the host, builds three binaries, pulls the weights and smoke tests the result before handing over a launcher. Not fit if you need more than one node, a model other than the 0731 checkpoint, or a stack without CUDA 13 on Blackwell silicon. Before running it, read the disk requirement rather than the memory requirement, because the roughly 81 GiB checkpoint plus a 6.5 GiB drafter is what the 120 GiB free disk check is protecting, and decide whether to accept a curl to bash install or to read install.sh first.

## FAQ

### Can DeepSeek-V4 be run locally?

Yes, on one DGX Spark. The installer starts ds4-server on port 8000 with a 512k context, serving the DeepSeek-V4-Flash-0731 checkpoint from a local 81 GiB asymmetric GGUF entirely on-device, on a GB10 with 128 GB of unified memory of which about 119 GiB is usable. The engine is Entrpi/ds4, a batched fork of antirez/ds4.

### What are the capabilities of DeepSeek V4 flash?

The 0731 checkpoint used here is an asymmetric quant: IQ2_XXS on the routed gate and up projections, Q2_K on the routed down, Q8_0 on everything else dense, and F16 for the compressor and indexer. It has no MTP head, so DSpark is the only speculation available, and the model is 1M native with a 512k default launch context.

### How much disk and hardware does ds4-on-spark need?

The installer verifies aarch64, a GB10 with SM121, CUDA 13 and at least 120 GiB of free disk before doing anything, then clones the pinned fork into ~/code/ds4, builds three binaries in about eight seconds, and downloads roughly 81 GiB of model plus a 6.5 GiB drafter into ~/gguf. --cuda-arch sm_120 builds for RTX PRO 6000 and 5090 class Blackwell, and --force skips the host check.

## Sources

- [Entrpi/ds4-on-spark on GitHub](https://github.com/Entrpi/ds4-on-spark)
- [Issues](https://github.com/Entrpi/ds4-on-spark/issues)
- [License: MIT](https://github.com/Entrpi/ds4-on-spark/blob/main/LICENSE)
- [README](https://github.com/Entrpi/ds4-on-spark/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/entrpi-ds4-on-spark
