# distributed-llama: tensor-parallel LLM inference across home machines

> distributed-llama splits one quantized model across a root node and 2^n - 1 workers over Ethernet. It is a C++ tensor-parallel runtime for CPU and Vulkan, not a general serving stack, and its node-count and quantization rules are strict.

**b4rtaz/distributed-llama** — Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.

- Repository: https://github.com/b4rtaz/distributed-llama
- Stars: 3,097 · Forks: 253
- Language: C++
- License: MIT
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/b4rtaz-distributed-llama

## What problem distributed-llama solves, and for whom

A 70B model quantized to q40 is 40 GB of weights. Most home machines do not have 40 GB of usable RAM in one box, but several of them together often do. distributed-llama exists to turn that pile of hardware into one inference cluster. The README frames it plainly: connect home devices into a cluster, and more devices mean faster inference. The intended user is someone with two to eight spare machines on the same network, not a datacenter operator. The project targets Linux, macOS and Windows, and it is optimized for ARM and x86_64 AVX2 CPUs, with experimental Vulkan support for GPUs. The repository is C++ with a Python launcher, MIT licensed, and the last push was on 2026-07-05.

## Root node, worker nodes and how a token moves through the cluster

The README's architecture diagram shows a switch or router with one root device and several worker devices, each worker listening on port 9999. The root node loads the model and weights and forwards them to workers, and it also synchronizes the state of the neural network. The root is itself a worker: it processes its own slice of the network. Worker nodes hold no model configuration of their own; they receive their slice and compute on it. The README describes the parallelism as tensor parallelism with high-speed synchronization over Ethernet, so the split is across tensors rather than across layers, which means every token involves an all-reduce-like exchange between nodes. That is the design's cost: the network is in the critical path of every generated token. The README notes that RAM usage of the network is split across all nodes, and that the root node requires a bit more RAM than workers. The two operating modes are inference and chat, plus a worker mode and a separate dllama-api server.

## Installing distributed-llama and running a first prompt

The README's fastest path is the single-command launcher. It requires Python 3 and a C++ compiler, and the command downloads both the model and the tokenizer. This example pulls the 1.7 GB Llama 3.2 1B Instruct Q40 model, which is the smallest listed and the least painful first run.

```bash
python launch.py llama3_2_1b_instruct_q40
```

After the download finishes, the launcher builds the C++ binary through the Makefile and starts the root node. The repository's Makefile compiles with g++, -std=c++11, -O3, and adds -march=native -mtune=native unless TERMUX_VERSION is defined. Vulkan support is opt-in: defining DLLAMA_VULKAN adds -DDLLAMA_VULKAN, links -lvulkan, and compiles the shader sources under src/nn/vulkan through glslc.

To add workers, the README documents a separate worker command. Each worker binds a host and port, and the root receives the worker addresses as a space-separated list. The examples directory contains n-workers.sh, which is where the multi-node wiring lives.

```bash
dllama worker --host 10.0.0.2 --port 9999 --nthreads 4
dllama chat --model dllama_model_meta-llama-3-8b_q40.m --tokenizer dllama_tokenizer_llama3.t --buffer-float-type q80 --workers 10.0.0.2:9999
```

The README also lists dllama inference for a benchmark run with --prompt and --steps, and dllama-api for an HTTP server. The examples folder includes chat-api-client.js, which is the reference client for that API.

## The 2^n node rule and the quantization pairing are hard constraints

The known limitations section is unusually blunt, and it is the part most likely to disqualify a deployment. You can run distributed-llama only on 1, 2, 4, 8 and so on nodes. A three-node cluster is not supported. The maximum number of nodes is capped by the number of KV heads in the model, which the README links to issue #70. Quantization is also paired: a q40 model must use q80 for buffer-float-type, and an f32 model must use f32. There is no documented q40-with-f32 or q80-with-f32 combination. The README does not document rollback or downgrade behaviour between releases, and it does not publish a compatibility matrix for models outside the launch.py table. The node-count rule is the one that bites first, because it is a property of the topology rather than of the model.

## Where distributed-llama is the wrong tool

If your model already fits on one machine, this project adds network synchronization to every token and buys you nothing. Single-node inference with llama.cpp or a similar runtime will be simpler and will not fail when a worker drops off the network. If you need an arbitrary number of nodes, or you want to scale to ten or twenty machines, the 2^n rule and the KV-head ceiling make that impossible. If your primary interface must be an OpenAI-compatible HTTP endpoint with batching and multi-user scheduling, the README describes dllama-api as an API server but does not document batching, request queues or multi-tenant behaviour; treat it as an interface to a single cluster, not a serving platform. And if your interconnect is Wi-Fi, the tensor-parallel design will spend its time waiting on synchronization. The README's own framing is Ethernet.

## distributed-llama compared with llama.cpp RPC and Exo

The closest comparison is llama.cpp's RPC backend, which people search for as llama.cpp RPC or llama.cpp distributed inference. Both split a model across machines, but the split is different. llama.cpp's RPC approach offloads whole layers or tensors to remote backends and keeps the orchestration in the main process; distributed-llama instead runs a dedicated root node that owns the model and the synchronization, with workers that hold no configuration. That makes distributed-llama's worker setup simpler (a host, a port, a thread count) but makes the root node a single point of coordination. Exo, the other name that comes up in searches, is a Python-based cluster runtime with a different deployment model and a broader device-discovery story. The real difference to weigh is not raw speed but operational shape: distributed-llama gives you a small C++ binary and an explicit root, while the alternatives give you a larger runtime with more automatic behaviour. The README does not publish a benchmark against either, so any speed comparison you read elsewhere is not sourced from this repository.

## Maintenance, upgrades and the MIT licence

The repository is not archived, and the last push was on 2026-07-05. The most recent release listed is v0.16.5 on 2026-02-02, following v0.16.4 on 2026-01-17 and v0.16.3 on 2025-10-26. The release cadence in that window is uneven: two releases two weeks apart, then a gap of roughly three months to the next. The project has absorbed a fundamental codebase refactor (v0.12.0, February 2025) and added Vulkan support (v0.13.0, March 2025) inside the same period, so upgrading across those boundaries is not a drop-in operation; the README does not describe a migration path. Upgrade cost is mostly operational: you rebuild the C++ binary, and because the Makefile defaults to -march=native -mtune=native, the binary is tied to the build machine's CPU unless you set TERMUX_VERSION or override CXXFLAGS. The licence is MIT, which permits commercial use and modification with the copyright notice retained. That is a permissive licence, but it says nothing about the licences of the model weights you load, which are governed separately by whoever published the model.

## Hardware that actually makes sense for this

The README links a discussion titled Llama 3.3 70B on 4 x Mac Mini M4 Pro 24GB RAM, and there is a dedicated Raspberry Pi guide in docs. Those two endpoints describe the realistic range. On the low end, a handful of Raspberry Pis can run the smaller Qwen 3 and Llama 3.2 models, where the 0.9 GB Qwen 3 0.6B Q40 is the smallest entry in the launch table. On the high end, four machines with 24 GB each can host a 70B q40 model whose weights total 40 GB, which is the point of the exercise. The launch.py table is the practical shopping list: it lists each model with its size and its exact command, from 0.9 GB up to the 238 GB Llama 3.1 405B Instruct Q40. If your target model is not in that table, the README points to a separate Hugging Face conversion guide, and that conversion step is where most of the friction will be.

## Conclusion

Adopt distributed-llama if you have 2, 4 or 8 machines with enough combined RAM for a model that will not fit on one box, and you accept q40 weights with q80 buffer-float-type. Do not adopt it if you need a single-node server, an arbitrary node count, or an OpenAI-style HTTP API as your primary interface. Before committing, verify that your node count is a power of two and does not exceed the model's KV head count, that your target model is in the launch.py list or convertible through the Hugging Face conversion doc, and that your interconnect is Ethernet fast enough that synchronization does not dominate.

## FAQ

### What is distributed-llama?

It is a C++ runtime that connects home devices into a cluster to accelerate LLM inference, using tensor parallelism and synchronization over Ethernet. The README describes a root node that loads and forwards the model plus worker nodes that each process a slice of the neural network.

### What does distributed AI mean?

In this project's terms, it means splitting one neural network across several machines so the RAM usage of the network is divided among them, with the root node synchronizing the state of the network. Each node computes its own slice rather than running an independent model.

### How many nodes can I use with distributed-llama?

Only 1, 2, 4 and other powers of two, according to the known limitations. The maximum number of nodes is also bounded by the number of KV heads in the model, which the README links to issue #70.

### Which quantizations does distributed-llama support?

The README lists two pairings: a q40 model with q80 buffer-float-type, and an f32 model with f32 buffer-float-type. Other combinations are not documented.

### Can I run distributed-llama on a Raspberry Pi?

Yes. The README links a dedicated guide, How to Run on Raspberry Pi, alongside the Linux, macOS and Windows guide and the GPU guide. The project is optimized for ARM and x86_64 AVX2 CPUs.

## Sources

- [b4rtaz/distributed-llama on GitHub](https://github.com/b4rtaz/distributed-llama)
- [Issues](https://github.com/b4rtaz/distributed-llama/issues)
- [License: MIT](https://github.com/b4rtaz/distributed-llama/blob/main/LICENSE)
- [README](https://github.com/b4rtaz/distributed-llama/blob/main/README.md)
- [Releases](https://github.com/b4rtaz/distributed-llama/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/b4rtaz-distributed-llama
