Open-source project
antirez/ds4 avatar
antirez/ds4

ds4 (DwarfStar): A Purpose-Built Inference Engine for DeepSeek V4 Flash on Consumer Hardware

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm.

22,719 stars2,182 forksCMIT

At a glance

What is it?
ds4, also called DwarfStar, is a C-language MIT-licensed inference engine that runs DeepSeek V4 Flash and a small set of related large language models on Apple Silicon Metal, NVIDIA CUDA, and AMD ROCm hardware. It is deliberately narrow in scope: it uses its own GGUF format rather than accepting arbitrary GGUF files, and it targets engineers with 96 GB Macs, DGX Spark systems, or Strix Halo workstations.
Who is it for?
Engineers with 96 GB Apple Silicon Macs, a DGX Spark, or Strix Halo hardware who specifically want to run DeepSeek V4 Flash or the supported GLM models locally will find DwarfStar the most targeted tool for that purpose. Anyone expecting to load arbitrary GGUF models from the broader llama.cpp ecosystem will hit a hard boundary: DwarfStar uses its own GGUF files and does not accept general model files from other sources.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 10 days ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What DwarfStar Solves and Who It Is For

DwarfStar exists to answer one question: how do you run a capable large language model on hardware that an engineer can actually own? The README frames the project's purpose as serving that goal, identifying MacBook Pros with 96 GB RAM, NVIDIA DGX Spark workstations, and AMD Strix Halo systems (such as the Framework Desktop) as the primary hardware targets.

The project name in the repository is ds4, and the binary is also called ds4. The internal project name DwarfStar appears in the README. The author is Salvatore Sanfilippo, also known as antirez, the creator of Redis.

The software is beta quality. The README says directly: 'Consider it beta quality. Before each release, a big QA run is executed, however instabilities and regressions are definitely possible.' There are no GitHub releases in the repository; the main branch receives updates directly. The last push was on 2026-09-20.

The supported hardware list has three tiers. Metal on Apple Silicon is the primary development target. NVIDIA CUDA support targets the DGX Spark as the main goal, with additional support for multi-GPU configurations including Ada Lovelace cards such as the L40S. ROCm support targets AMD Strix Halo systems. The README notes that DwarfStar can run DeepSeek V4 Flash on Ada Lovelace when other inference backends require newer GPU architectures.

Supported Models: DeepSeek V4, GLM 5.x, and Qwen3

The model list is small and intentional. DwarfStar supports DeepSeek V4 Flash (including an experimental vision variant), DeepSeek V4.1 Flash (Metal and text inference on CUDA), GLM 5.2, GLM 5.3, GLM 5.3 Flash, DeepSeek V4 PRO, and Qwen3.8 Flash Next (Metal and CUDA).

The README is explicit about why the list is short: 'Model support is intentionally opportunistic. The project follows the best open weights for useful local machine sizes, especially 128 GB laptops and 256/512 GB workstations. A model may be removed when a better replacement arrives.' This means the model roster can shrink as well as grow, and users relying on a specific model should monitor the repository for changes.

Three technical properties justify this model selection, as the README explains: capable open-weight models now fit on high-end personal machines; DeepSeek V4 Flash and PRO along with GLM 5.2 tolerate aggressive routed-expert quantization; compressed KV caches combined with fast local SSDs make long contexts practical on consumer hardware.

The project requires its own GGUF files rather than accepting arbitrary GGUF files. The README states: 'you need to use the GGUF files the project produces, that are part of the project itself.' This is a hard constraint that separates DwarfStar from general GGUF runners.

Building DwarfStar and Running a First Inference

Start by cloning the repository:

sh
git clone https://github.com/antirez/ds4.git
cd ds4

The build command depends on hardware. For Metal on Apple Silicon:

sh
make

For the NVIDIA DGX Spark:

sh
make cuda-spark

For Strix Halo (Framework Desktop and similar AMD ROCm systems):

sh
make strix-halo

For other CUDA configurations, including multi-GPU and Ada Lovelace cards:

sh
make cuda-generic

Each build target has a corresponding platform guide in the docs/ directory: docs/METAL.md, docs/DGX_SPARK.md, docs/STRIX_HALO.md, and docs/CUDA_MULTI_GPU.md cover hardware-specific prerequisites, memory sizing, and setup details.

With the binary built, download the DeepSeek V4 Flash Q2 quantization as a starting point:

sh
./download_model.sh ds4f-q2

The script saves the model to the gguf/ directory and can be interrupted and resumed. Once downloaded, the four primary binaries cover inference modes:

sh
./ds4
./ds4 -p "Explain Redis streams in one paragraph."
./ds4-agent
./ds4-server --ctx 32768

The default model is ds4flash.gguf, a symlink updated by main-model downloads. Pass `-m FILE` to specify a different model file. The HTTP server mode at `./ds4-server` exposes an API endpoint. The agent mode at `./ds4-agent` runs a coding agent. The context size for the server can be set with `--ctx`.

SSD Streaming, Multi-GPU, and RDMA Parallelism

For Macs with less than 96 GB of RAM, the README documents SSD streaming as a path to running models that exceed available memory. Fast local NVMe SSDs make this practical according to the documentation; the expected tradeoff is lower generation speed compared to running entirely in RAM.

For multi-GPU setups, DwarfStar supports Ada Lovelace cards including L40S. The README reports achieving approximately 126 tokens per second aggregate generation with 16 sessions across eight L40S cards. Ada Lovelace support is notable because the README says other inference implementations for some of these newer models require newer GPU architectures, while DwarfStar's implementation covers Ada Lovelace.

Two distributed inference modes are documented. Tensor parallelism works across two 128 GB Macs connected with RDMA, and can run 4-bit DeepSeek Flash or GLM 5.3 Flash at that configuration. Pipeline parallelism connects multiple systems to sum their RAM, enabling larger model quants that would not fit on any single machine. The GLM 5.2 full (non-Flash) quants require this approach on 128 GB single systems.

The README also mentions an 'engram' feature visible in the repository (ds4_engram.c and ds4_engram.h) but does not describe it in the README text, leaving its purpose to be discovered through the code or the AGENT.md and CONTRIBUTING.md files.

DwarfStar vs. llama.cpp: Narrow vs. General

llama.cpp is the most direct alternative. The README acknowledges this relationship in detail: 'ds4.c does not link against GGML, but it exists thanks to the path opened by the llama.cpp project.' The GGUF quantization formats, GGUF ecosystem, and kernels in DwarfStar are derived from llama.cpp's work, and some source-level pieces are retained or adapted under MIT.

The functional difference is scope. llama.cpp is a general GGUF runner that supports hundreds of model architectures; any model with a compatible GGUF conversion is usable. DwarfStar deliberately supports only a small set of models and uses its own GGUF files rather than the broader ecosystem's files. This narrows compatibility but allows deeper per-model optimization: DwarfStar can implement model-specific kernel paths, quantization strategies, and caching designs that a general runner would have to generalize away.

For the specific use case of running DeepSeek V4 Flash on a DGX Spark or a 96 GB Mac, DwarfStar may offer optimizations that llama.cpp does not yet implement. For running other models, llama.cpp is the only option: DwarfStar will not load them. The choice depends entirely on whether your target model appears in DwarfStar's supported list.

Maintenance, MIT License, and AI-Assisted Development

The last push to the repository was on 2026-09-20. There are no tagged releases; the README describes the project as beta quality with QA runs before major updates. The pace of visible file changes, including model-specific CUDA header files (ds4_deepseek41_cuda.cuh, ds4_qwen4_cuda.cuh, ds4_glm53_vision_gpu.cuh), suggests ongoing development across model architectures.

The README makes an unusually direct disclosure about development methodology: 'This software is developed with strong assistance from AI coding agents and with humans leading the ideas, testing, and debugging. We say this openly because it shaped how the project was built. If you are not happy with AI-developed code, this software is not for you.' This transparency is rare in open-source projects and sets expectations about code provenance for auditors or contributors.

The MIT license permits modification, redistribution, and commercial use. The license file includes the GGML authors' copyright notice because certain quantization formats and kernel logic are retained or adapted from llama.cpp. The CONTRIBUTING.md and EVAL_DATA.md files are present in the repository, indicating contribution and evaluation processes are documented. The project author recommends that users with coding agent access use agents to modify and extend the project for specific hardware setups rather than waiting for official support.

Editorial conclusion

Engineers with 96 GB Apple Silicon Macs, a DGX Spark, or Strix Halo hardware who specifically want to run DeepSeek V4 Flash or the supported GLM models locally will find DwarfStar the most targeted tool for that purpose. Anyone expecting to load arbitrary GGUF models from the broader llama.cpp ecosystem will hit a hard boundary: DwarfStar uses its own GGUF files and does not accept general model files from other sources. Before downloading, check docs/MODELS.md in the repository to confirm that the specific quant or variant of DeepSeek V4 Flash you need is available and compatible with your hardware configuration.

Frequently asked questions

Can I run DeepSeek V4 Flash locally with ds4?

Yes. DwarfStar is designed specifically for running DeepSeek V4 Flash locally on Apple Silicon Macs with 96 GB or more of RAM, NVIDIA DGX Spark hardware, and AMD ROCm systems. Macs with less RAM can use SSD streaming to run the model, though at reduced speed. The download_model.sh script fetches the required GGUF files directly.

What hardware does ds4 require to run DeepSeek V4 Flash?

The README names Macs with 96 GB or more as the Metal target, the NVIDIA DGX Spark as the primary CUDA target, and AMD Strix Halo systems (such as the Framework Desktop) for ROCm. Macs with less RAM can use SSD streaming, and multi-GPU CUDA setups including Ada Lovelace cards are supported. Multi-Mac RDMA tensor parallelism is also documented for larger model quants.

Does ds4 work with GGUF files from other sources?

No. DwarfStar uses its own GGUF files produced by the project; it is not a general GGUF runner. The README states this directly: you must use the GGUF files the project generates, not arbitrary GGUF files from the broader llama.cpp ecosystem. The download_model.sh script fetches the project's own model files.

Official sources

  1. Official README
  2. Project repository
Community notes

Community notes