# exo shards a model across your Macs and stops short of Windows

> exo is a Python and Rust cluster runtime that pools several Apple Silicon machines into one inference host, with an OpenAI-compatible API and a dashboard on a single localhost port. Every install path is a source build, and the cluster layer is patched to a personal fork of zenoh.

**exo-explore/exo** — Run frontier AI locally.

- Repository: https://github.com/exo-explore/exo
- Stars: 47,683 · Forks: 3,542
- Language: Python
- License: Apache-2.0
- Published: 2026-08-17 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/exo-explore-exo

## Two Mac Studios only become one computer with Thunderbolt 5 RDMA

exo's premise is that the machines you already own can hold a model that does not fit on one of them. Devices running exo find each other with no configuration, and the runtime splits a model across whatever has joined. The transport is the part with the fine print. The feature list claims day-0 support for RDMA over Thunderbolt 5 and a 99 percent reduction in latency between devices, and the quick start ends with a pointer to the RDMA section to enable that feature on macOS 26.2 or later. Below that OS version the reduction does not apply, and the sharding still works over whatever link exo falls back to.

The two numbers worth holding on to are the tensor parallelism figures: up to 1.8x speedup on 2 devices and 3.2x on 4 devices. Both are given as feature-list claims with no methodology in the repository, and both assume the fast path above. Read them as the ceiling for a well-cabled cluster, not as a promise for the two laptops on your desk.

## The cluster layer is a zenoh fork pinned to one branch

Cargo.toml is where the trust boundary of the project is. The workspace holds two crates, rust/exo_rs and rust/networking, on edition 2024, and its networking stack is zenoh pinned to the exact version 1.9.0 alongside netwatcher for interface watching. Then the manifest replaces crates.io entirely for the zenoh family:

```toml
[patch.crates-io]
zenoh = { git = "https://github.com/evanev7/zenoh.git", branch = "exo" }
zenoh-buffers = { git = "https://github.com/evanev7/zenoh.git", branch = "exo" }
zenoh-codec = { git = "https://github.com/evanev7/zenoh.git", branch = "exo" }
zenoh-collections = { git = "https://github.com/evanev7/zenoh.git", branch = "exo" }
zenoh-config = { git = "https://github.com/evanev7/zenoh.git", branch = "exo" }
zenoh-core = { git = "https://github.com/evanev7/zenoh.git", branch = "exo" }
zenoh-crypto = { git = "https://github.com/evanev7/zenoh.git", branch = "exo"
```

Seven crates now come from a fork under one contributor's account, and pidfile-rs is pulled from another personal repository. If that branch is rebased, force-pushed or deleted, a source build stops resolving and there is no released fallback to move to, because the patch section overrides what crates.io would have served. That is the first thing to check on any machine that will not build.

## Automatic discovery means cluster membership is a property of the network

The discovery story is one sentence in the quick start: devices running exo discover each other, without manual configuration. Nothing in that sentence mentions a peer list, a token or an allowlist, and the README does not document a way to restrict which machines may join. Combined with the zenoh-based networking layer, the practical reading is that exo is a peer on whatever network it finds, which is convenient for a desk full of Macs and wrong for a shared office segment.

The same paragraph carries the thing you will use first. Each device provides an API and a dashboard for its cluster, and both live at http://localhost:52415. The API is compatible with the OpenAI Chat Completions API, the Claude Messages API, the OpenAI Responses API and the Ollama API, so an existing client can point at that address unchanged. The README documents no authentication setting and no TLS option for that port, so treat the bind address as a decision you have to verify yourself before the port sits on anything but loopback.

## The macOS path needs a pinned macmon fork because the Homebrew build crashes

Install exo from source on macOS and the first surprise is that Xcode is a prerequisite, because MLX compiles against the Metal toolchain. Past brew, uv, node and a nightly Rust toolchain, the hardware monitor is where the instructions get specific: install the pinned fork revision used by this repo rather than the Homebrew package, because Homebrew macmon 0.6.1 still crashes on Apple M5.

```bash
cargo install --git https://github.com/vladkens/macmon \
  --rev a1cd06b6cc0d5e61db24fd8832e74cd992097a7d \
  macmon \
  --force
```

The shortcut is Nix, and it costs you a system config edit. Adding trusted-users and experimental-features set to nix-command flakes in /etc/nix/nix.conf, then restarting the daemon with `sudo launchctl kickstart -k system/org.nixos.nix-daemon`, buys you one command:

```bash
nix run .#exo
```

The rest of the macOS path is four steps, and the dashboard is built from source in the second one rather than shipped as a binary:

```bash
# Clone exo
git clone https://github.com/exo-explore/exo

# Build dashboard
cd exo/dashboard && npm install && npm run build && cd ..

# Install Python dependencies, including the MLX backend
uv sync --extra mlx

# Run exo
uv run exo
```

After that the dashboard and API answer on http://localhost:52415/. On Linux the same four commands work with a different extra for your hardware, and macmon is not needed there at all; the Linux path asks for node 18 or higher while the macOS path names no node version.

## Python 3.13 only, with anyio and MLX pinned to exact versions

The manifest header is short and decides who can install it:

```toml
[project]
name = "exo"
version = "0.3.70"
description = "Exo"
readme = "README.md"
requires-python = "==3.13.*"
```

That specifier accepts any 3.13 patch and nothing else, so a 3.12 or 3.14 interpreter fails at sync rather than at runtime. Several dependencies are exact pins for the same reason: anyio==4.11.0, mlx==0.32.0, mflux==0.17.5, and mlx-cpu==0.31.2 on the CPU-only Linux path against mlx==0.32.0 everywhere else. Picking up a bugfix in any of them means editing pyproject.toml, and the two MLX lines mean the CPU-only path resolves a different MLX build than the Apple path.

The rest of the dependency list is telling about what exo is. exo-rs is a hard dependency, so a Python install also compiles Rust. huggingface-hub is there for custom models, tiktoken is called out as required for the Kimi K2 tokenizer, openai-harmony and transformers cover the model formats, and hypercorn and fastapi serve the API. Even the simplest install is a compiler toolchain, not a wheel.

## The published benchmarks are links to somebody else's blog post

The Benchmarks section is three collapsed entries and no tables. Qwen3-235B at 8-bit, DeepSeek v3.1 671B at 8-bit and Kimi K2 Thinking at native 4-bit, all on 4 x M3 Ultra Mac Studio with tensor parallel RDMA, and all three credit the same source: a Jeff Geerling post on 15 TB VRAM on Mac Studio and RDMA over Thunderbolt 5. The dashboard image caption repeats the same hardware with the same two models. Nothing in the repository publishes tokens per second, memory headroom, or first-token latency for its own runs.

The consequence for an evaluator is that the project's only quantitative claims are the 1.8x and 3.2x speedup figures and the 99 percent latency reduction, none of which come with a measurement description. The hardware in the benchmark titles is specific enough to act on, 4 x 512GB M3 Ultra, but two machines wired differently will not reproduce those numbers, and the README gives no guidance on the cabling or the split that produced them.

## Building the app shells out to Xcode, so macOS is the only packaged target

The repository root carries an Xcode project under app/, a PyInstaller spec under packaging/, a Rust workspace and a Svelte dashboard under dashboard/. The justfile wires them together, and the app recipe is the honest summary of the build: rust-rebuild, then a clean sync, then the PyInstaller pass, then xcodebuild. That last step is a plain shell line you can run yourself:

```bash
env -u LD xcodebuild build -project app/EXO/EXO.xcodeproj -scheme EXO -configuration Debug -derivedDataPath app/EXO/build
```

Building it on anything other than a Mac is not possible, which settles the Windows question the README never answers. The same file shows the day-to-day commands a contributor uses: `uv run pytest src` for tests, `uv run ruff check --fix` for lint, `uv run basedpyright --project pyproject.toml` for types, and `uv sync --all-packages --extra mlx` to reset the environment. Note that a PyInstaller binary and a fork-patched Rust dependency both arrive in the same artifact, so a team distributing it internally is shipping code it did not build from upstream sources.

## A v1.0.71 tag sits next to a 0.3.70 manifest, five months behind the last push

The release history and the manifest disagree about which version of exo exists. GitHub tags run 1.0.69, 1.0.70 and 1.0.71, with 1.0.71 published on 2026-04-23, while pyproject.toml still reads version 0.3.70. The last push to main was on 2026-09-29, which is roughly five months after that tag, and the project is not archived.

So there is no unambiguous version to pin. A tag tells you when a release was cut, the manifest tells you what the source you just cloned calls itself, and neither is connected to a published package index: every documented install builds from a clone. If you run exo in a pipeline, record the commit rather than a version string, and re-check the Cargo patch section at the same time, since that branch can move independently of any tag.

## Conclusion

exo suits a team that already owns several Apple Silicon machines and Thunderbolt 5 cabling and accepts building from source on macOS 26.2 or later. It does not suit a single laptop, a Windows shop, or anyone wanting a package index install, because every documented path compiles Rust and MLX. Verify first that the zenoh fork branch named exo still resolves, since the cluster layer is patched to a personal repository rather than to crates.io.

## FAQ

### What is exo used for?

exo connects multiple devices into an AI cluster so a model larger than one machine can run across them, and each device serves an API and dashboard at http://localhost:52415. It shards the model based on a realtime view of device resources and the latency and bandwidth of each link.

### What is exo for AI?

The inference backend is MLX, with MLX distributed handling communication between devices, and the API is compatible with the OpenAI Chat Completions, Claude Messages, OpenAI Responses and Ollama APIs. Custom models load from the HuggingFace hub.

### What is the software EXO?

It is a Python and Rust cluster runtime whose entry point is the exo script mapping to exo.main:main, with the Rust side exposed through the exo-rs bindings and pyinstaller used to package the binary. The project is licensed Apache-2.0 and maintained by exo labs.

## Sources

- [Official README](https://github.com/exo-explore/exo#readme)
- [Project repository](https://github.com/exo-explore/exo)
- [Release notes](https://github.com/exo-explore/exo/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/exo-explore-exo
