# cake: A Rust-Based Distributed AI Inference Server for Heterogeneous Devices

> cake is a multimodal AI inference server written in Rust that runs on a single machine or shards a model's transformer blocks across a heterogeneous cluster of iOS, Android, macOS, Linux, and Windows devices. It targets developers who want to run large models that do not fit on any single device by pooling the memory and compute of multiple machines they already own.

**evilsocket/cake** — Distributed inference for mobile, desktop and server.

- Repository: https://github.com/evilsocket/cake
- Stars: 3,126 · Forks: 208
- Language: Rust
- License: NOASSERTION
- Published: 2026-09-24 · Updated: 2026-09-24 · Language: en
- Canonical page: https://hysenlabs.com/projects/evilsocket-cake

## What cake Does and Who It Is For

Running a large language model typically requires more GPU memory than a single consumer device holds. cake's approach is to shard the transformer blocks of a model across multiple devices, each handling a portion of the forward pass. The master node holds the model weights, assigns layers to workers based on their available VRAM or compute, and streams only the required weight shards to each worker. Workers cache received data locally so subsequent runs do not retransfer weights.

The README describes the motivation as leveraging "planned obsolescence": old phones, laptops, and desktops that are no longer useful for their original purpose may still contribute compute to a distributed inference cluster.

The use cases are: running inference on a single machine (the simplest path), building a small home cluster of heterogeneous devices for models that exceed a single device's memory, or using a Docker Compose cluster of CUDA workers for a team setup.

cake supports three modalities. Text generation covers 15 model families with architecture auto-detected from HuggingFace `config.json` files. Image generation supports Stable Diffusion and FLUX (6 image model variants). Voice synthesis uses VibeVoice TTS with voice cloning support. The README lists 2 TTS models.

This is not a production inference service. The README states: "This is experimental code that's being actively developed and changed very quickly."

## Building cake for Different Backends

cake is written in Rust and must be compiled from source. The README gives separate build commands for each hardware backend:

```sh
cargo build --release --features cuda        # Linux (NVIDIA)
cargo build --release --features metal       # macOS (Apple Silicon GPU)
cargo build --release --features accelerate  # macOS (Apple Silicon CPU, F32 models)
cargo build --release --features vulkan      # Linux (AMD/Intel/Steam Deck)
cargo build --release                        # CPU only (portable)
```

The correct feature flag depends on the hardware. CUDA targets NVIDIA GPUs on Linux. Metal targets Apple Silicon GPUs on macOS. Vulkan covers AMD and Intel GPUs on Linux and also the Steam Deck. The `accelerate` flag uses Apple's Accelerate framework for CPU-based inference on Apple Silicon when Metal is not needed.

The workspace is organized into three crates: `cake-core` (the inference engine), `cake-cli` (the command-line interface), and `cake-mobile` (Kotlin Multiplatform bindings for iOS and Android). The mobile crate uses UniFFI to generate Kotlin bindings and produces a static library (`libcake_mobile.a`) for iOS and a shared library for Android.

## Downloading Models and Running Single-Node Inference

Models are downloaded from HuggingFace using the `cake pull` command:

```sh
cake pull evilsocket/Qwen3-0.6B             # text model (600M params)
cake pull evilsocket/flux1-dev               # image model (FLUX.1-dev FP8)
cake pull evilsocket/VibeVoice-1.5B          # voice synthesis model
```

Models are stored in the standard HuggingFace cache directory (`~/.cache/huggingface/hub/`) and shared with other tools that use the same cache. `cake list` shows locally available models; `cake rm` deletes a cached model.

For single-node usage, architecture is auto-detected from the model's `config.json`:

```sh
cake run evilsocket/Qwen3-0.6B "Explain quantum computing in simple terms"
```

For interactive chat, use the TUI client:

```sh
cake chat Qwen/Qwen3-0.6B
```

To start an OpenAI-compatible REST API server with a built-in web UI:

```sh
cake serve evilsocket/Qwen3-0.6B
```

For image generation:

```sh
cake run evilsocket/flux1-dev --model-type image-model --image-model-arch flux1 \
  "a cyberpunk cityscape at night"
```

For voice synthesis with voice cloning from a reference audio file:

```sh
cake run evilsocket/VibeVoice-1.5B --model-type audio-model \
  --voice-prompt voice.wav "Hello world"
```

## Distributed Inference: Clustering Across Multiple Machines

The distributed mode uses a `--cluster-key` shared secret to connect workers to the master. Workers run on any machine; they do not need the model data because the master streams the required weight shards to them.

Start workers on any machines in the cluster:

```sh
cake run --cluster-key mysecret --name gpu-server-1    # machine A
cake run --cluster-key mysecret --name macbook          # machine B
```

Run inference from the master (which has the model):

```sh
cake run evilsocket/Qwen3-0.6B "Hello" --cluster-key mysecret
```

Or start an API server as the master:

```sh
cake serve evilsocket/Qwen3-0.6B --cluster-key mysecret
```

The master discovers workers via mDNS on the local network ("zero-config mDNS clustering" per the README). No manual IP addresses or port configuration are needed for same-network discovery. The README also mentions support for manual topology files for scenarios where mDNS does not work.

Weight transfer between nodes uses zstd compression and CRC32 checksums for integrity verification. Workers cache received shards locally, so only the first run pays the transfer cost for each shard.

The Docker Compose path provides an example cluster for Linux/NVIDIA. It requires the NVIDIA Container Toolkit and a topology file specifying which layers map to which worker containers. The `docker-compose.yml` shows a master plus two worker services, all on a `cake-net` bridge network.

## Limitations and Comparison with llama.cpp

cake is experimental and changes quickly. The README says this explicitly. Model support is currently 15 text model families, 6 image model variants, and 2 TTS models. Models outside these families are not supported. The auto-detection reads `config.json` from the HuggingFace checkpoint; if a model's architecture is not implemented in cake's codebase, it will fail to load.

The Docker image only supports Linux with NVIDIA GPUs. macOS users with Apple Silicon must build and run natively, as the README's docker-compose.yml comments state: "Docker on macOS cannot access Metal GPUs."

Distributed inference adds network latency to every forward pass. For a fast single GPU that can already run the model, distributing across slower devices on a home network will be slower, not faster. The benefit is enabling models that would not fit in any single device's memory.

llama.cpp is the most common alternative for running LLMs on consumer hardware without a full GPU. It supports CPU inference across many platforms and has a large community. The key difference is that llama.cpp is a single-node CPU/GPU inference engine while cake is a distributed multi-device inference framework. cake also handles image generation and TTS in addition to text, while llama.cpp focuses on text models. For a developer who wants to run a 7B parameter text model on a single MacBook, llama.cpp is more mature and has broader model support.

## The FAIR License and Commercial Use

cake is released under the FAIR License (Free for Attribution and Individual Rights) v1.0.0. This is not an MIT, Apache, or GPL license. The README documents three tiers:

Non-commercial use (personal, educational, research, non-profit) is freely permitted.

Commercial use (SaaS, paid applications, any monetization) requires visible attribution to the project and its author. The license text contains the specific attribution requirement.

Business use (any use by or on behalf of a business entity) requires a signed commercial agreement with the author. The README lists the contact email as `evilsocket@gmail.com`.

This means that deploying cake as part of a product, offering it as a service, or using it within a company's internal tooling all require contacting the author. Developers who need an inference server under a permissive license compatible with commercial use should evaluate alternatives. To check the licenses of cake's own dependencies, the README instructs: install `cargo-license` with `cargo install cargo-license` and run `cargo license`.

## Repository Structure and Mobile Support

The repository workspace contains `cake-core/` (the inference engine), `cake-cli/` (command-line tools), `cake-mobile/` (Rust bindings for mobile), and `cake-mobile-app/` (the Kotlin Multiplatform application for iOS and Android). The `autoresearch/` directory and `CLAUDE.md` at the root suggest active development tooling.

The Makefile contains targets for syncing to remote build machines (`sync_bahamut`, `sync_blade`) and for building the mobile library for iOS and Android. iOS requires `cargo build --release --target=aarch64-apple-ios` with the `metal` feature; the build produces `libcake_mobile.a` for Xcode's cinterop linker. Android requires `cargo ndk` and generates Kotlin bindings via UniFFI.

The `docs/` directory contains the full usage guide, model list, clustering documentation, Docker setup, image generation guide, and voice generation guide. These are separate Markdown files referenced by links in the README.

## Conclusion

cake is the right tool if you have multiple personal devices and want to run a model that is too large for any single one of them, or if you want to experiment with distributed inference using consumer hardware. The FAIR License terms are specific: non-commercial and personal use are freely permitted; commercial use requires attribution; any business use requires a signed agreement with the author. Before adopting it, verify whether the model family you need is among the 15 supported text model families and check the clustering documentation for the network topology requirements. The last push was on 2026-04-24.

## FAQ

### What is cake AI?

cake is an open-source multimodal AI inference server written in Rust. It can run text generation, image generation (Stable Diffusion and FLUX), and voice synthesis (VibeVoice TTS) either on a single machine or sharded across a cluster of iOS, Android, macOS, Linux, and Windows devices using the FAIR License.

### Is cake AI free to use?

Non-commercial use including personal, educational, research, and non-profit use is freely permitted under the FAIR License. Commercial use requires visible attribution. Any business use requires a signed agreement with the author; contact `evilsocket@gmail.com` for those inquiries.

### How do you build and run cake AI?

Clone the repository and build with the appropriate Cargo feature flag for your hardware: `cargo build --release --features metal` for Apple Silicon, `cargo build --release --features cuda` for NVIDIA on Linux, or `cargo build --release` for CPU-only. Then run `cake pull <model>` to download a model and `cake run <model> "prompt"` to generate output.

## Sources

- [evilsocket/cake on GitHub](https://github.com/evilsocket/cake)
- [Issues](https://github.com/evilsocket/cake/issues)
- [README](https://github.com/evilsocket/cake/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/evilsocket-cake
