Mesh-LLM: pool GPUs across machines behind one OpenAI-compatible endpoint
Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat.
At a glance
- What is it?
- Mesh-LLM is a Rust project that turns several machines into one inference pool and exposes it at http://localhost:9337/v1. It installs from a shell script, runs a public or private mesh, and can split a dense model across nodes. The MoA gateway is marked experimental.
- Who is it for?
- Adopt Mesh-LLM if you have two or more boxes with GPUs or unified memory and you want one OpenAI-compatible endpoint at http://localhost:9337/v1 without writing a scheduler yourself; the local-model-only mode is also a clean way to serve a single GGUF file with no mesh networking. Do not adopt it if you need a stable, versioned API contract today: the MoA gateway is labelled experimental, and the mesh is still at v0.76.0-rc9, a release candidate.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Mesh-LLM actually pools, and who needs that
The README frames the problem in one line: Mesh LLM pools GPUs and memory across machines and exposes the result as one OpenAI-compatible API at http://localhost:9337/v1. That is a narrower claim than "distributed training" and a broader one than "remote inference proxy".
The intended user is someone who has more than one machine with a usable accelerator and wants a single endpoint that clients already understand. Because the surface is the OpenAI chat completions shape, existing tooling that points at a base URL keeps working; the README shows a plain curl against /v1/chat/completions with a model id.
It is also aimed at agent users. The table of workflows lists mesh-llm goose, mesh-llm opencode, mesh-llm claude and mesh-llm pi as first-class commands, which tells you the project expects coding agents to be the main consumer rather than a chat UI alone. If you have one workstation and one model that fits on it, this is more moving parts than you need.
Routing, QUIC transport and the Skippy split path
The mesh decides placement in a fixed order. Single-machine fit comes first: if one node can host the full model, it serves locally with no stage traffic. Only after that does mesh routing apply, where every node exposes the same /v1 API and requests are routed by the model field to a peer that can serve that model.
Peer traffic runs over QUIC and is described as end-to-end encrypted, covering inference requests, responses and split-model activations. Iroh relays forward encrypted packets without reading their payload, so a relay is a bandwidth hop, not a trust boundary. That distinction matters when you are deciding whether to join a public mesh.
For models too large for one box, the Skippy path loads a dense model as package-backed layer stages. The coordinator plans contiguous layer ranges, starts downstream stages first, waits for readiness, then publishes the stage-0 route. Layer packages contain model-package.json plus GGUF fragments, so a peer fetches only the pieces for its assigned stage instead of the whole model. The ordering is the interesting part: publishing stage 0 last means the entry point does not accept traffic before the later stages can answer.
Discovery splits by intent. Published meshes advertise through Nostr discovery; private meshes stay invite-token based. The repository layout backs this up: crates/skippy-coordinator, crates/skippy-topology, crates/skippy-scheduler and crates/mesh-llm-routing are separate crates, so planning, scheduling and routing are distinct components rather than one module.
Installing Mesh-LLM and serving a first model
The README gives a one-line install for Unix-like systems. It fetches install.sh from the main branch and pipes it to bash.
curl -fsSL https://raw.githubusercontent.com/Mesh-LLM/mesh-llm/main/install.sh | bashOn Windows the equivalent is a PowerShell one-liner, and the README notes that Windows commands use mesh-llm.exe.
irm https://raw.githubusercontent.com/Mesh-LLM/mesh-llm/main/install.ps1 | iexApple Silicon users have a Homebrew formula: brew install Mesh-LLM/tap/mesh-llm. The README points to the Mesh-LLM/mesh-packaging repository for versioned formulas, Ubuntu and Arch packages, checksums, SBOMs and OCI images, and to the platform install guides for the supported package matrix. Those are the places to check before assuming your distribution is covered.
After installing, run setup once.
mesh-llm setupThen the shortest path to a working endpoint is the auto mode, which the README says chooses a backend flavor, downloads a suitable model if needed, joins the best discovered public mesh, starts the local API on port 9337 and the web console on port 3131.
mesh-llm serve --autoTo confirm what is being served, query the models endpoint; the README pipes it through jq to list ids.
curl -s http://localhost:9337/v1/models | jq '.data[].id'If you would rather not join anything public, serve a single named model as a private mesh instead.
mesh-llm serve --model Qwen3-8B-Q4_K_MFor server deployments the README suggests adding --headless, which hides the web UI while keeping the management API on the --console port.
mesh-llm serve --auto --headlessUninstalling is staged. mesh-llm uninstall --dry-run previews the cleanup and mesh-llm uninstall --yes performs it; the README states that ~/.mesh-llm configuration and identity data are preserved unless you pass --purge-config. That default is the right one for anyone experimenting, since it means a reinstall keeps your node identity.
Local-model-only mode and its hard failure rule
The most predictable configuration is the one that opts out of the mesh entirely. With --local-model-only, Mesh-LLM starts the OpenAI frontend and one local Skippy model runtime and nothing else: no QUIC, no discovery, no peer maintenance, no split planning, no plugins, no release lookup, no web console, no management API.
mesh-llm serve \
--local-model-only \
--model /models/model.gguf \
--port 9337The README states that startup fails if the complete model does not fit within detected local capacity or --max-vram, and that it never falls back to distributed serving. That is a deliberate constraint, and it is worth reading as a feature: a deployment script that expects a local model will fail loudly rather than silently start routing inference to strangers. The same section imposes a path rule. For --local-model-only, the values of --model, --gguf and --mmproj must be absolute paths and must not be symlinks. Container images and Nix-style stores that symlink model files into place will trip this, so resolve paths before launch. Binding beyond loopback requires --listen-all, which the README advises adding only when the endpoint must be reachable off-host.
The MoA gateway is a preview, and the version number says so
Sending "model": "mesh" fans a request out to every model available in the mesh in parallel, arbitrates the responses with deterministic logic, and returns one OpenAI-compatible reply. The arbiter runs in code rather than as another model call, and escalates to a reducer LLM only on genuine conflict. Tool calls flow through the full pipeline.
The README is unusually direct about the status: the MoA gateway is new, and behavior, routing heuristics, error shapes and tuning knobs may change between versions. It tells you to treat model: "mesh" as a preview and to use a specific model id when you need stable semantics. Take that at face value. Fan-out to every model in a mesh also means latency is set by the slowest responder and cost is paid on every node, so it is a poor default for high-volume traffic even once it stabilizes.
The release history reinforces the same reading. The three most recent releases listed are v0.76.0-rc9, v0.76.0-rc8 and v0.76.0-rc7, all release candidates. The last push to the repository was on 2026-09-09 and the repository is not archived, so work is ongoing, but a project shipping release candidates is not promising a frozen API. Pin a version if you deploy it.
Where Mesh-LLM is the wrong tool, and what to compare it against
Mesh-LLM is the wrong tool when the model already fits on one machine and you have no interest in sharing capacity. Every mesh feature is overhead in that case, and the local-model-only path exists precisely because the project knows the difference.
It is also a poor fit when you need a documented rollback procedure. The README documents uninstall with --dry-run and --yes, and it documents that config and identity survive unless --purge-config is passed, but it does not document rolling back to a previous version after an upgrade. If your deployment needs a described downgrade path, that gap is on you to solve.
The closest comparison in the documentation is not another mesh product but the direct path inside this project. A conventional single-node OpenAI-compatible server, such as the widely used llama.cpp server, keeps one process and one model file with no peer transport, no discovery and no coordinator. Mesh-LLM's difference is the placement decision: it will run a model locally when it fits, route to a peer when another node can serve the model, and split contiguous layer ranges across stages when the model is too large for any single box. If your problem is "this model does not fit on my hardware", that is the feature you are buying. If your problem is "I want a small local endpoint", the single-process server is simpler and has fewer failure modes.
The other boundary is the public mesh itself. Joining with --auto means accepting whatever the discovery layer finds. The README notes that published meshes advertise through Nostr discovery while private meshes stay invite-token based, so an invite token is the mechanism for controlling membership. If you cannot describe who is in your mesh, do not publish one.
Licence, packaging and what a version bump costs
Mesh-LLM is Apache-2.0. That is a permissive licence with an explicit patent grant and a requirement to preserve notices when redistributing. For most internal deployments the practical effect is that you can run modified builds without publishing your changes. This is not legal advice; if you redistribute binaries or embed the project in a product, read the licence text in the repository and the notices your packaging path generates.
Upgrade cost is the harder question. The README documents install and uninstall, and it points at the separate mesh-packaging repository for versioned formulas, Ubuntu and Arch packages, checksums, SBOMs and OCI images. That packaging repository is where you should look for a pinned artifact rather than tracking main. The install script itself fetches from the main branch, which means a fresh install is not reproducible by default.
The workspace is large. Cargo.toml lists dozens of member crates spanning the CLI, config, identity, routing, the OpenAI frontend, the Skippy runtime and coordinator, plugin host and API, FFI and Node.js bindings, plus a test harness. Building from source with the Justfile target is a real compile, not a small one. If you are upgrading a fleet, expect the version skew between nodes to matter: the README describes the mesh plane as designed for mixed-version compatibility, which suggests the project treats heterogeneous node versions as a normal state rather than an error.
Editorial conclusion
Adopt Mesh-LLM if you have two or more boxes with GPUs or unified memory and you want one OpenAI-compatible endpoint at http://localhost:9337/v1 without writing a scheduler yourself; the local-model-only mode is also a clean way to serve a single GGUF file with no mesh networking. Do not adopt it if you need a stable, versioned API contract today: the MoA gateway is labelled experimental, and the mesh is still at v0.76.0-rc9, a release candidate. Before committing, verify three things in your own environment: that the model you want fits detected local capacity (startup fails rather than falling back), that your paths satisfy the absolute-path, no-symlink rule for --model, --gguf and --mmproj, and that your Windows or WSL2 setup matches the troubleshooting guide rather than the Linux quick start.
Frequently asked questions
What is Mesh-LLM?
It is a Rust project that pools GPUs and memory across machines and exposes the result as one OpenAI-compatible API at http://localhost:9337/v1. The mesh decides whether a model runs locally, routes to a peer, or uses Skippy stage splits for models too large for one box.
What is an LLM mesh?
In Mesh-LLM's terms it is a set of nodes that each expose the same /v1 API, with requests routed by the model field to a peer that can serve that model. Peer traffic runs over QUIC, and published meshes advertise through Nostr discovery while private ones use invite tokens.
What is mesh AI used for?
The README targets serving models that do not fit on one machine, and driving coding agents: it ships mesh-llm goose, mesh-llm opencode, mesh-llm claude and mesh-llm pi commands. It can also fan a prompt to every model in the mesh via the experimental model: "mesh" gateway.
What is a mesh model?
In this project, a mesh model is the placement outcome rather than a model format: a model that a node can serve locally, that another node in the mesh can serve, or that is loaded as package-backed layer stages through Skippy when it is too large for a single box.
What is an example of a mesh in Mesh-LLM?
The README's workflow table gives concrete examples: mesh-llm serve --auto joins the best discovered public mesh, mesh-llm serve --model Qwen3-8B-Q4_K_M starts a private mesh, and mesh-llm serve --model Qwen3-8B-Q4_K_M --publish publishes your own.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mesh-llm-mesh-llm)