# Parallax: a peer to peer serving mesh whose pins, tags and quick start disagree with each other

> GradientHQ/parallax shards models across mixed personal devices over a Lattica peer layer, with SGLang, vLLM and MLX backends. The packaging carries three exact pins, two mutually exclusive GPU routes that both install MLX, and a build backend that never reads its own setuptools table.

**GradientHQ/parallax** — Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere

- Repository: https://github.com/GradientHQ/parallax
- Stars: 1,395 · Forks: 148
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/gradienthq-parallax

## The quick install serves a checkpoint the model table never mentions

The quick install is four shell lines: clone the repo, run the bundled script, activate a virtual environment, and serve one checkpoint.

```sh
git clone https://github.com/GradientHQ/parallax.git
cd parallax
./install.sh
source .venv/bin/activate
parallax serve -m Qwen/Qwen3.5-0.8B
```

That last line is the entire visible command surface of the project, and its model does not appear in the supported model table. The Qwen row there points at Qwen3.6-35B-A3B, an open weight mixture of experts model, and Qwen3.5-0.8B is named nowhere else on the page. Nothing says whether the 0.8B dense checkpoint is a supported target, a placeholder for smoke testing, or a leftover from an earlier revision of the guide. The same row adds a second loose end: its description says Qwen3-Next and larger Qwen3 MoE variants are represented, while the HuggingFace column holds a single link. Which Qwen weights the serving layer accepts is decided in code under src/, not on the front page, and the only other route into the project is the docs link, which points at an install page and a quick start page without repeating the model name.

## Two GPU backends, two unrelated version lines, no switch to flip

The About section names two GPU backends at once: SGLang and vLLM. pyproject turns them into mutually exclusive extras that pin entirely different version lines.

```toml
gpu = [
  "sglang[all]==0.5.12",
  "kernels<0.15",
  "accelerate",
  "mlx-lm==0.31.3",
  "mlx[cpu]==0.31.2",
]

vllm = [
  "vllm==0.14.0",
  "mlx-lm==0.31.3",
  "mlx[cpu]==0.31.2",
]
```

SGLang sits at 0.5.12 while vLLM sits at 0.14.0, and neither is a floor that lets a resolver take a later release. Choosing a backend is therefore an edit to a dependency list rather than a runtime switch, and no page says which route is the tested one, whether the two can coexist in one environment, or what the serving layer does when both are importable. The About section simply lists them side by side underneath a bullet about dynamic request scheduling and routing. Note also that torch is declared only in the mac extra, at torch==2.8.0, while the GPU extras leave it to arrive transitively through sglang[all] or vllm, which puts the accelerator runtime version in the hands of those two projects rather than of this one.

## Both GPU routes install the Mac runtime, and disagree on how

The Mac extra and the GPU extras ask for the same Apple packages in two different ways.

```toml
mac = [
  "nanobind==2.12.0",
  "torch==2.8.0",
  "mlx-lm==0.31.3",
  "mlx==0.31.2",
]
```

The Mac route requests mlx==0.31.2, while the gpu and vllm routes request mlx[cpu]==0.31.2, the same version with the cpu extra spelled out. So a Linux host with accelerators installs Apple MLX wheels built for CPU either way, and all three routes drag in mlx-lm==0.31.3 regardless of whether a Mac is anywhere in the cluster. That may be deliberate, since the About section credits paged KV cache management and continuous batching to the Mac backend, but nothing states whether the Apple code paths stay inactive on non Darwin hosts or are merely unused there. A second gap sits next to it: cmake>=3.27 and ninja>=1.11 live in the dev extra, not in mac, gpu or vllm. nanobind and the MLX packages ship as native wheels, so a source build of any pinned component needs a toolchain the runtime extras never declare.

## Open floors on the model stack, exact pins on the protocol layer

The base dependency list mixes three policies in twenty entries. Most are open floors: msgpack>=1.0.7, safetensors>=0.5.1, transformers>=4.57.1, huggingface-hub, modelscope, jinja2>=3.1.0, numpy>=1.26, pyzmq>=25.0, psutil>=5.9.5, requests, httpx[socks]>=0.26.0, aiohttp, uvicorn, uvloop, fastapi, pydantic, orjson. Three are exact pins: protobuf==6.31.1, dijkstar==2.6.0 and lattica==1.0.21. One, kernels<0.15 in the gpu extra, is a ceiling with no floor, the only dependency written that way.

The three exact pins all sit on the protocol side. lattica==1.0.21 is the peer layer the About section credits with the whole mesh design, protobuf==6.31.1 covers wire messages, and dijkstar==2.6.0 is named after its authors. A new release of any of those cannot be taken without editing pyproject, while the model stack underneath floats on whatever satisfies its floors, so a peer layer upgrade and a transformers upgrade are two very different kinds of work here. Python is bounded in the same spirit: requires-python is >=3.11,<3.14, excluding 3.14 by declaration. The top level of the repository holds no changelog file, so nothing in the tree records why those numbers were chosen.

## A poetry-core backend on a project configured for setuptools

Parallax declares a poetry-core build backend and then configures itself like a setuptools project.

```toml
[build-system]
requires = ["setuptools>=68", "wheel", "poetry-core"]
build-backend = "poetry.core.masonry.api"
```

The project table lists its packages in PEP 621 style with a from key pointing into src, and the same file also carries a setuptools discovery table pointing at the same directory, which the declared backend never reads. setuptools>=68 and wheel are installed into the build environment for a backend that ignores them. The part with teeth is the four names shipped out of src: parallax, scheduling, parallax_utils and parallax_extensions. Three are namespaced by the project, while scheduling lands in site unspaced as a bare generic name, so any other distribution using that name in the same environment collides on import. parallax_extensions is packaged but never mentioned in the README: no extension registry, entry point group or loading order is described, so what lives in that package and how a third party would attach to it has to be read from source.

## Three tags, a fourth version in the News block, and seven months past the last tag

The tag list and the front page tell different release stories. Three tags are listed: v0.1.0 on 2025-11-11, v0.1.1 on 2025-11-26 and v0.1.2 on 2025-12-02. The version field in pyproject reads 0.1.2, matching the newest tag exactly, so there is no drift between the two. The last push on the default branch is 2026-07-01, about seven months after that tag, which means everything on main since December 2025 is unreleased.

The News block has not tracked any of it. Its newest line is the 2026/02 entry announcing OpenClaw integration, and there is nothing after it even though commits continued into July. The oldest line is stranger: it announces that version 0.0.1 shipped in 2025/10, and no such tag appears among the three listed, which start at 0.1.0 in November. So a reader looking for a version to pin finds three tags, a front page referring to a fourth version number that is not in the list, and a gap between the newest tag and the newest commit. The repository is not archived, and none of these dates appear in the page as anything other than news and tag dates.

## Scheduling claims with no numbers, and a benchmark extra with no entry point

The About section lists five core features and not one of them carries a figure.

```
- Host local LLM on personal devices
- Cross-platform support
- Pipeline parallel model sharding
- Paged KV cache management & continuous batching for Mac
- Dynamic request scheduling and routing for high performance
```

Paged KV cache management, continuous batching and dynamic routing are exactly the claims that need throughput, latency and memory numbers to be checkable, and the page offers none, for any model or any hardware mix. A benchmark extra does exist, carrying transformers, tqdm, datasets and pillow, and the top level holds both scripts/ and tests/, but no documented command runs a benchmark and no path points at one. The pillow dependency inside a benchmark extra raises a question the page leaves open: it points to image inputs somewhere in a harness, while the model table describes text and multimodal reasoning rather than a vision pipeline. Anyone who wants to know whether the scheduling claims hold has to open scripts/, find the entry point, and reconstruct the invocation from scratch.

## A mesh engine is described in three sentences and joined in zero

The About section calls Parallax a fully decentralized inference engine developed by Gradient, and credits P2P communication to Lattica. The base dependencies hint at what carries that traffic: pyzmq>=25.0 for messaging, httpx[socks]>=0.26.0 for requests through proxies, aiohttp and requests as clients, uvicorn and uvloop with fastapi on the serving side. A SOCKS capable HTTP client in a mesh engine base list is a hint that peers reach each other through something other than a local network.

What the visible pages never state is how a node finds another node, which ports it opens, whether peers must be mutually reachable, or what happens when a node disappears in the middle of a request. The User Guide is three links, to install, to quick start and to working with OpenClaw, plus a contributing guide, so those answers sit under docs/ rather than on the front page. The front page is thin elsewhere too: a heading reading Trusted by Partners has an empty container beneath it, the two visible badges both point at the issues tracker, and the Product Hunt anchor wraps nothing while repeating utm_source in its query string.

## Conclusion

Parallax is worth reading as architecture rather than as a deployment you can copy. Anyone planning to run it should first check whether Qwen3.5-0.8B from the quick install is still a valid target, whether lattica==1.0.21 and protobuf==6.31.1 match the peer layer their nodes need, and whether the MLX CPU wheels that both GPU extras pull in are acceptable on their host. For anything resembling production serving, remember that the newest tag is from 2025-12-02 while the last push on the default branch is 2026-07-01, so reading the default branch means reading unreleased code.

## FAQ

### What does GradientHQ/parallax actually do?

It is a distributed model serving framework. The README describes a fully decentralized inference engine that shards a model across nodes with different configurations and physical locations, using pipeline parallelism for model sharding and a Lattica peer layer for communication.

### Which models does GradientHQ/parallax serve?

The supported table names DeepSeek-V3.2, DeepSeek-R1, MiniMax-M3, MiniMax-M2.7, GLM-5.2, GLM-5.1, GLM-4.7, Kimi-K2-Thinking, Kimi-K2-Instruct-0905, Qwen3.6-35B-A3B, gpt-oss-120b, gpt-oss-safeguard-120b and Step-3.5-Flash. The quick install line serves Qwen/Qwen3.5-0.8B, which is not among them.

### Does GradientHQ/parallax need NVIDIA GPUs, or will it run on a Mac?

Both shapes are declared. The GPU extras install SGLang or vLLM, and the mac extra installs MLX LM, which the About section credits with paged KV cache management and continuous batching. The mac extra is the only one that pins torch, at torch==2.8.0.

### How do I install GradientHQ/parallax?

Clone the repository, run ./install.sh, activate .venv/bin/activate, then call the serve command, for example parallax serve -m Qwen/Qwen3.5-0.8B. The console script is declared in pyproject.toml as parallax = parallax.cli:main, and Python must be 3.11 through 3.13.

### Is GradientHQ/parallax still cutting releases?

The newest listed tag is v0.1.2 from 2025-12-02 and pyproject carries the matching version 0.1.2, while the last push on the default branch is 2026-07-01. The News block's most recent entry is the 2026/02 OpenClaw integration.

## Sources

- [GradientHQ/parallax on GitHub](https://github.com/GradientHQ/parallax)
- [Issues](https://github.com/GradientHQ/parallax/issues)
- [License: Apache-2.0](https://github.com/GradientHQ/parallax/blob/main/LICENSE)
- [README](https://github.com/GradientHQ/parallax/blob/main/README.md)
- [Releases](https://github.com/GradientHQ/parallax/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/gradienthq-parallax
