Model or dataset
Entrpi/ds4-on-spark avatar
Entrpi/ds4-on-spark

Entrpi/ds4-on-spark: a one-command DeepSeek-V4-Flash server for the DGX Spark

Entrpi/ds4, a Blackwell CUDA perf fork of antirez/ds4 on NVIDIA DGX Spark: one-command install, ~3x upstream prefill, ~1.5x decode, DSpark, and full continuous batch support

405 stars30 forksShellMIT

At a glance

What is it?
A packaging repo that pins a Blackwell CUDA fork of antirez/ds4 and installs it on GB10 hardware with a single curl command. The trade-off is a narrow hardware target and a build that reaches outside the repository for the engine and the weights.
Who is it for?
Adopt it if you already own a DGX Spark with CUDA 13 and want DeepSeek-V4-Flash answering on port 8000 without assembling a build yourself. Skip it on datacenter Blackwell, on non-NVIDIA accelerators, or if you need an upstream project you can track commit by commit.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 35 days ago.
What is it written in?
Mainly Shell, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What ds4-on-spark packages, and who it is for

Entrpi/ds4-on-spark is not an inference engine. It is a packaging and pinning repository. The engine itself is Entrpi/ds4, described in the README as a major-feature fork of antirez/ds4 (DwarfStar 4) that turns the upstream engine into a batched multi-request server for Blackwell CUDA. This repository clones that fork at a release tag, builds it for the local GPU architecture, downloads the weights, runs a smoke test, and installs a launcher.

The audience is narrow and specific. You need an NVIDIA DGX Spark with its GB10 chip at SM121, CUDA 13 installed, and at least 120 GiB of free disk according to the installer's host check. The README also says RTX PRO 6000 and 5090-class sm_120 cards build through an override flag, and that datacenter B200/B300 (sm_100) is untested. If none of those describe your machine, the repository is not aimed at you.

The problem it removes is assembly work. Getting a MoE model of this size running means building a CUDA engine for the right architecture, sourcing a quantized GGUF plus a speculative drafter, and wiring a server that speaks the API your client already uses. The README frames the result as one command that ends with a server listening on :8000.

How the stack is put together

Three pieces move independently. The engine is cloned from Entrpi/ds4 at a pinned tag, currently v0.6.5, and built native for sm_121 by the installer. The model is an ~81 GiB asymmetric quant of DeepSeek-V4-Flash-0731 from antirez/deepseek-v4-gguf, with IQ2_XXS on the routed gate and up projections, Q2_K on the routed down projection, Q8_0 on everything dense, and F16 for the compressor and indexer. A separate ~6.5 GiB DSpark drafter comes from bleysg/DeepSeek-V4-Flash-DSpark-drafter-GGUF.

The README is explicit that the 0731 checkpoint has no MTP head, so DSpark is the only speculation path on this model. The fork's claimed gains over upstream are 2.4 to 3.3 times prefill throughput and 1.33 to 1.47 times decode speed across a 2k to 128k context range, plus continuous batching and prefix caching with KV banks persisted to disk so they survive restarts.

The memory story is the part worth reading closely. Since v0.6 the engine keeps a single ledger of its memory and charges each request for its prompt plus its decode budget, assuming 32768 tokens when the client omits max_tokens. Unused context is demand-mapped rather than reserved, which is why the default launch context of 512k does not cost 512k worth of memory up front. The README says the fork has been measured at 2.26 million tokens of active context with shipped defaults intact, and up to 3 million with tuning. When a request genuinely does not fit, the documented behaviour is a refusal stating how many bytes were missing and whether a retry will help.

Installing on a DGX Spark and serving a first request

The installer is fetched and piped to bash with a flag that starts the server when it finishes. The README gives this exact form:

bash
curl -sSL https://raw.githubusercontent.com/entrpi/ds4-on-spark/main/install.sh | bash -s -- --start

Expect the script to check the host, clone the fork into ~/code/ds4 (or $DS4_SRC_DIR), build ds4, ds4-server and ds4-bench with CUDA_ARCH=sm_121, then download roughly 81 GiB of weights plus the 6.5 GiB drafter into ~/gguf (or $DS4_GGUF_DIR). It finishes by running a "capital of France" smoke test that asserts "Paris" appears in the output, installing ds4-serve into ~/.local/bin, and starting ds4-server on :8000 with a 524288-token context.

If you want to see what the script does before running it, there is a help path:

bash
curl -sSL https://raw.githubusercontent.com/entrpi/ds4-on-spark/main/install.sh | bash -s -- --help

After installation, ds4-serve is the launcher. The README's examples show the defaults being overridden by passing flags straight through to ds4-server:

bash
ds4-serve                              # full stack, ctx 524288, 127.0.0.1:8000
ds4-serve -c 1048576                   # the model's full 1M window
ds4-serve -c 32768 --host 0.0.0.0      # smaller context, reachable from LAN

The process runs in the foreground, so supervision is your choice: nohup, systemd or tmux. On the client side, ds4-server is documented as speaking four APIs from the same continuous batch: OpenAI chat completions and completions, OpenAI Responses (what Codex uses), and Anthropic Messages (what Claude clients use).

Where the packaging bites back

The first constraint is hardware. The installer's host check wants aarch64, GB10/SM121, CUDA 13 and 120 GiB free. The README states that datacenter B200/B300 at sm_100 is untested, so a Blackwell datacenter card is not a supported target even though it is the same generation. --force skips the host check, but skipping a check is not the same as passing it.

The second constraint is provenance. The installer pulls code from a raw GitHub URL, a fork pinned at a tag, and two Hugging Face repositories. Nothing in the README describes signature verification or checksum pinning for the script or the weights. You are trusting four external sources at install time, and the 81 GiB download is not something you want to repeat casually.

The third is the fork relationship itself. Entrpi/ds4 carries the features; this repository pins it. That means upstream fixes reach you when the pin moves, not when antirez commits. If your team needs to follow upstream development directly, or to patch the engine, you are one layer removed from the code you would be editing.

Finally, the README does not document rollback. There is no described procedure for returning to a previous pinned release or for cleaning up a partially completed install, and no uninstall instructions appear in the README.

ds4-on-spark against building antirez/ds4 yourself

The real alternative is going to the source: clone antirez/ds4, build it for your architecture, fetch the GGUF from antirez/deepseek-v4-gguf, and run the upstream server. That path gives you the architectural foundation the README credits for the model recipe, without a packaging layer in between.

The difference is in what you get and what you do. Upstream, per the README's own comparison, does not carry continuous batching, disk-persisted prefix caching, DSpark speculative decode, or the memory governor that lets a 512k default context be cheap. Those are fork features. Building upstream yourself means accepting the lower prefill and decode figures the README quotes as the baseline, in exchange for a codebase you can track commit by commit.

A second alternative, for readers who do not own a Spark, is simply not to run this locally. The README's framing is entirely on-device inference on a 128 GB unified-memory box. If your hardware is a workstation with a 24 GB card, neither this repository nor the fork it packages is a route to serving a model whose quantized weights alone are around 81 GiB.

Licence, upgrade cost and what the pin implies

The repository is MIT licensed. That covers this packaging layer. The engine it clones and the weights it downloads are separate artifacts with their own licensing, and the README does not state their terms, so check each source before deploying anything commercially. Nothing in this article is legal advice.

Upgrade cost is dominated by the pin. Moving to a newer fork release means the installer clones a different tag, which in practice means another build (the README quotes roughly 8 seconds for the compile) and possibly another weight download if the model recipe changed. The README notes that --no-download lets you reuse existing GGUF files, which is the lever to pull when only the engine moved.

There is also a compatibility note worth flagging. The README says --with-mtp applies only to the legacy generation's base plus MTP files, and that on 0731 the fallback ladder is DSpark then plain decode. So the flags you pass are coupled to which model generation you have on disk, not just to your hardware. Changing one without the other produces a configuration the README does not cover.

Editorial conclusion

Adopt it if you already own a DGX Spark with CUDA 13 and want DeepSeek-V4-Flash answering on port 8000 without assembling a build yourself. Skip it on datacenter Blackwell, on non-NVIDIA accelerators, or if you need an upstream project you can track commit by commit. Before you run the installer, confirm the free disk space, decide whether the 81 GiB download is acceptable, and read the fork's CHANGELOG at the pinned v0.6.5 tag to see which upstream commits sit underneath.

Frequently asked questions

Can DeepSeek-V4 be run locally with ds4-on-spark?

Yes, that is the entire purpose of the repository. It installs a pinned fork of the ds4 engine and the DeepSeek-V4-Flash-0731 GGUF on a DGX Spark, then starts a server on port 8000. The weights are roughly 81 GiB plus a 6.5 GiB drafter, and the installer requires at least 120 GiB of free disk.

What are the capabilities of DeepSeek-V4-Flash as served by ds4-on-spark?

The README describes a 1M-native context window, with a default launch context of 524288 tokens and the full window reachable through -c 1048576. The server speaks OpenAI chat completions and completions, OpenAI Responses, and Anthropic Messages from one continuous batch. DSpark speculative decode is armed by default on the 0731 checkpoint.

What hardware does ds4-on-spark require?

The installer checks for aarch64, GB10/SM121, CUDA 13 and at least 120 GiB of free disk, which points at an NVIDIA DGX Spark. The README says RTX PRO 6000 and 5090-class sm_120 cards build with --cuda-arch sm_120, and that datacenter B200/B300 at sm_100 is untested.

How do I serve without speculative decoding in ds4-on-spark?

Pass --no-dspark to the installer, which serves plain continuous decode instead and skips the drafter download, or pass --no-dspark or --no-spec to ds4-serve. The README notes that on the 0731 checkpoint the fallback ladder is DSpark then plain, and that --with-mtp applies only to the legacy generation's base plus MTP files.

Does ds4-on-spark persist its KV cache across restarts?

The README states that the fork supports prefix caching with disk-persisted KV banks that survive restarts. That is a fork feature listed among the additions over the upstream engine, alongside continuous batching and the memory governor.

Official sources

  1. Entrpi/ds4-on-spark on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/entrpi-ds4-on-spark.svg)](https://hysenlabs.com/projects/entrpi-ds4-on-spark)