exo: running frontier models across a cluster of your own machines
Run frontier AI locally.
At a glance
- What is it?
- exo turns several devices into one inference cluster, using MLX as the backend and a Rust networking layer for discovery. It is Apple-silicon first, and the README's own prerequisites are the best test of whether it fits you.
- Who is it for?
- exo is for people with several Apple silicon machines, or a Linux box with an NVIDIA GPU, who want to serve a model too large for one device and are willing to build from source. It is not for anyone who wants a single pip install, a Windows host, or a stable API surface: pyproject.toml pins requires-python to 3.13.*, the README gives no Windows instructions, and the Rust bindings need a nightly toolchain.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What exo is solving, and who has that problem
A frontier model at 8-bit does not fit in one desktop. The README's own examples make the scale concrete: DeepSeek v3.1 671B at 8-bit and Kimi-K2-Thinking at 4-bit, shown running across four 512GB M3 Ultra Mac Studio machines. The stated goal is to connect all your devices into an AI cluster, so that models larger than any single device can hold still run, and so that adding devices makes inference faster rather than just bigger.
The audience is narrower than the tagline suggests. This is for someone who already owns more than one capable machine, or who has a Linux host with an NVIDIA card, and who is willing to assemble a build toolchain rather than install a binary. The README lists Xcode, Homebrew, uv, Node, a nightly Rust toolchain and a pinned fork of macmon as prerequisites for macOS. That list is the real audience definition. If any line of it makes you close the tab, exo is not aimed at you yet.
How the cluster actually forms: Zenoh discovery and MLX sharding
Two layers do the work. The Python side, versioned 0.3.70 in pyproject.toml while the releases are tagged v1.0.71, holds the model logic and the HTTP surface. The Rust workspace in rust/exo_rs and rust/networking holds the transport, built on Zenoh. Cargo.toml pins zenoh to 1.9.0 and patches crates.io with a fork on the branch named exo, which tells you the project needed behaviour upstream did not provide. Peer discovery is therefore not a bespoke protocol; it rides on Zenoh's networking, which is why the README can claim devices find each other without manual configuration.
On top of that sits placement. The README describes topology-aware auto parallel: exo builds a realtime view of device resources and the latency and bandwidth of each link, then decides how to split the model. Tensor parallelism is the supported sharding mode, and the README states up to 1.8x speedup on two devices and 3.2x on four. Those are the project's numbers, not independent measurements. The interesting design choice is that the topology graph is recomputed rather than configured, which is what makes hot-adding a device plausible, and also what makes behaviour harder to predict when a link degrades.
Installing exo from source on macOS and Linux
The README gives two paths. On macOS with Nix already installed, one command runs it. This is the shortest route and the one the project puts first.
nix run .#exoThe README notes that accepting the Cachix binary cache avoids building the Xcode Metal toolchain yourself, and shows the two lines to add to /etc/nix/nix.conf: trusted-users and experimental-features = nix-command flakes. Without that, expect a long compile.
The manual path clones the repository, builds the dashboard with Node, syncs Python dependencies with uv, and starts the server.
# Clone exo
git clone https://github.com/exo-explore/exo
# Build dashboard
cd exo/dashboard && npm install && npm run build && cd ..
# Install Python dependencies, including the MLX backend
uv sync --extra mlx
# Run exo
uv run exoThe README states this starts the dashboard and API at http://localhost:52415/. Each device exposes its own API and dashboard, so you open that URL on whichever machine you are sitting at.
On Linux the prerequisites shrink: uv, Node 18 or higher, and a nightly Rust toolchain. macmon is explicitly macOS-only and not required. The backend extra differs by hardware, and the README names mlx-cuda13, with mlx-cuda12 and a CPU variant also defined in pyproject.toml.
uv sync --extra mlx-cuda13
uv run exoIf you already run Ollama or an OpenAI-compatible client, the README states exo is compatible with the OpenAI Chat Completions API, the Claude Messages API, the OpenAI Responses API and the Ollama API, so the first real use is pointing an existing client at the local endpoint rather than learning a new one.
The macmon pin and other sharp edges in the toolchain
One prerequisite deserves its own paragraph because it is the clearest sign of how young the stack is. The README instructs you to install macmon from a specific git revision rather than from Homebrew, with the reason stated plainly: Homebrew macmon 0.6.1 still crashes on Apple M5. The command uses cargo install with --git and --rev a1cd06b6cc0d5e61db24fd8832e74cd992097a7d, plus --force.
If you skip that and install the Homebrew package, hardware monitoring is the part that breaks, not necessarily inference. Still, it means the documented setup is pinned to a commit rather than a release, and pinned dependencies age. The same pattern appears in Cargo.toml, where Zenoh is patched to a personal fork branch. Neither is unusual for a project at this stage, and both are things a reviewer should weigh: you are tracking someone else's branch, and a force-push or a deleted branch breaks your build with no upstream fallback.
The Python requirement is equally strict. pyproject.toml sets requires-python to ==3.13.*, an exact minor pin, not a floor. If your environment is on 3.12 or 3.14, uv will resolve a different interpreter or refuse, and you will be debugging that before you debug the cluster.
Where exo is the wrong tool
Three cases stand out. The first is a single machine. If everything fits on one device, the cluster layer adds discovery, a Rust build and a pinned toolchain for no gain. A plain MLX or llama.cpp setup is less to maintain.
The second is Windows. The README documents macOS and Linux only, and PLATFORMS.md exists in the repository precisely because platform support is a matrix rather than a given. Anyone arriving from a search for exo on Windows will not find install steps here.
The third is production serving. The README promises API compatibility with four different interfaces, which is a broad surface to keep in step, and the version numbers do not line up: releases are tagged v1.0.71 while pyproject.toml declares 0.3.70. Nothing in the README describes a stability guarantee, a deprecation policy or a rollback path. Treat exo as a way to run models you own on hardware you own, not as infrastructure you promise someone else.
How exo differs from llama.cpp and Ollama
The obvious comparison is llama.cpp, and the difference is architectural. llama.cpp is a single-process inference engine with its own quantisation formats and a broad CPU and GPU backend list. It scales across GPUs within one host, and it targets hardware exo does not mention. exo assumes a network of separate machines and puts effort into the part llama.cpp leaves to you: discovering peers, measuring the links between them, and choosing a split. Its backend is MLX, which ties the fast path to Apple silicon in a way llama.cpp is not tied to any vendor.
Ollama is the closer comparison for a newcomer, because exo speaks Ollama's API. Ollama is a model runner with a curated library and a one-command install; it does not attempt to pool memory across hosts. exo is the opposite trade: much heavier setup, in exchange for a model too large for any single device and a speedup that grows with device count. The README's tensor parallelism figures, up to 3.2x on four devices, are the claim that justifies the extra work.
Maintenance, licence and what upgrading costs you
The last push to the default branch was on 2026-04-23, the same day as the v1.0.71 release, and the two prior releases landed on 2026-04-17 and 2026-03-27. The repository is not archived. That is a steady recent cadence, but it is also a project that ships often, which matters for upgrade cost.
The dependency pins are the real maintenance burden. Zenoh is patched to a fork branch in Cargo.toml, macmon is installed from a commit hash, and the Python requirement is an exact 3.13.*. Upgrading means re-resolving a workspace that spans uv, Cargo and npm, which is what the justfile is for: just sync runs uv sync --all-packages --extra mlx, and just rust-rebuild regenerates the Python stubs with cargo run --bin stub_gen before reinstalling exo_rs. If you change the Rust side, that second command is not optional.
The licence is Apache-2.0, per both the LICENSE file and the badge in the README. That is a permissive licence with an explicit patent grant, which is generally the easy case for internal and commercial use. It is not legal advice, and the bundled dependencies carry their own licences, so check those if you redistribute a packaged build.
Editorial conclusion
exo is for people with several Apple silicon machines, or a Linux box with an NVIDIA GPU, who want to serve a model too large for one device and are willing to build from source. It is not for anyone who wants a single pip install, a Windows host, or a stable API surface: pyproject.toml pins requires-python to 3.13.*, the README gives no Windows instructions, and the Rust bindings need a nightly toolchain. Before committing, check PLATFORMS.md for your exact hardware and extra, confirm the macmon fork revision installs on your machine, and verify that your API client works against the OpenAI-compatible endpoint on port 52415 rather than assuming full parity.
Frequently asked questions
What is exo for AI?
exo connects multiple devices into one AI cluster so you can run models larger than any single device can hold, and it uses MLX as the inference backend with MLX distributed for communication. The README also states it is compatible with the OpenAI Chat Completions, Claude Messages, OpenAI Responses and Ollama APIs.
What is the software EXO?
It is a Python project, versioned 0.3.70 in pyproject.toml, that runs a dashboard and API on each device at http://localhost:52415/. Devices running it discover each other automatically, with no manual configuration, according to the README.
What is Exo used for?
Running frontier models locally across a cluster of machines. The README shows DeepSeek v3.1 671B at 8-bit and Kimi-K2-Thinking at 4-bit on four 512GB M3 Ultra Mac Studio systems, and states tensor parallelism gives up to 1.8x speedup on two devices and 3.2x on four.
Is EXO AI used in healthcare?
The README and repository files describe exo as a local inference cluster and do not mention healthcare or any regulated deployment. Apache-2.0 permits such use, but the project documents no compliance features.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/exo-explore-exo)
Community notes