# vLLM Metal: running vLLM on Apple Silicon through MLX

> vLLM Metal is a community hardware plugin that lets the vLLM server, scheduler and paged block manager drive MLX model layers on Apple Silicon. It is a strong fit for Mac-based serving; it is not a drop-in replacement for CUDA vLLM.

**vllm-project/vllm-metal** — Community maintained hardware plugin for vLLM on Apple Silicon

- Repository: https://github.com/vllm-project/vllm-metal
- Website: https://docs.vllm.ai/projects/vllm-metal
- Stars: 1,787 · Forks: 271
- Language: Python
- License: Apache-2.0
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/vllm-project-vllm-metal

## The problem vLLM Metal solves on Apple Silicon

Upstream vLLM is built around CUDA and, to a lesser degree, other server-class accelerators. An Apple Silicon Mac has neither. The usual workaround is to run an MLX-only inference server and give up the vLLM request path entirely: the continuous batching scheduler, the paged block manager, the OpenAI-compatible API surface. vLLM Metal exists to keep those pieces and swap out only the compute layer.

The README describes it as a plugin that enables vLLM to run on Apple Silicon Macs using MLX as the primary compute backend, and it states that the plugin unifies MLX and PyTorch under a single lowering path. The audience is narrow and specific: developers who already know the vLLM CLI and want it to work on a Mac, and researchers who want to serve a quantized MLX checkpoint through vLLM's server rather than through mlx_lm's own tooling. If you have never used vLLM and only want a local chat server, this plugin adds a layer you do not need.

## How the plugin splits work between vLLM and mlx_lm

The architecture section of the README is unusually explicit about the division of labour. Upstream vLLM supplies the API server, the scheduler and the paged block manager. mlx_lm supplies the token-wise model layers. vllm-metal itself owns what the README calls the request-aware attention path: the paged varlen kernel, M5 NAX prefill, and speculative decoding.

That split explains most of the project's behaviour. Because the scheduler and block manager come from upstream, queueing, batching and KV-cache paging follow vLLM semantics rather than MLX semantics. Because the model layers come from mlx_lm, the set of runnable checkpoints tracks what mlx_lm and the MLX community publish. The genuinely new code is the attention path, and it is also the part most tightly coupled to the hardware: the README notes that as of 2026/08 the plugin uses M5 NAX tensor units to accelerate MHA, GQA and MQA prefill.

The repository layout matches that story. There is a Python package directory (vllm_metal/), a Rust extension declared in Cargo.toml with crate-type cdylib and a pyo3 dependency, and a build backend of maturin in pyproject.toml. So a wheel build compiles Rust and links against MLX, which is where the fragility described below comes from.

## Installing vLLM Metal and serving a first model

The README gives a single install command, a shell script fetched over curl and piped to bash. It creates a virtual environment at ~/.venv-vllm-metal by default and installs the plugin, vLLM core and the related libraries into it.

```bash
curl -fsSL https://raw.githubusercontent.com/vllm-project/vllm-metal/main/install.sh | bash
```

After it finishes, activating that environment puts the vllm CLI on your PATH. The README says you can then access vLLM right away, and points at the official vLLM CLI guide for usage, so the command surface is upstream vLLM's rather than something the plugin invents.

```bash
source ~/.venv-vllm-metal/bin/activate
vllm --help
```

Two requirements gate the install before you start. The README lists macOS 15 (Sequoia) or later on Apple Silicon, and native arm64 Python 3.12, with Rosetta and x86_64 Python explicitly unsupported. pyproject.toml is slightly wider on the Python side, declaring requires-python >=3.12,<3.14, so 3.13 is inside the package metadata even though the README names 3.12. Treat the README as the safer target.

For the model itself, the repository points to docs/supported_models.md as the full matrix rather than listing checkpoints in the README. One concrete example it does give is mlx-community/Qwen3.8-27B-8bit, described as a 27B hybrid SDPA + GDN linear model served on a single Apple Silicon Mac.

## The MLX pin is the real operational constraint

The dependency block in pyproject.toml is the most informative part of the repository. mlx is pinned with ==0.32.1, not a range, and the accompanying comment explains why: the prebuilt _paged_ops.so links MLX private headers and subclasses mlx::core::Primitive, and libmlx.dylib carries no SONAME version, so a wheel built against one MLX is only ABI-safe against that exact MLX. The comment ends with an instruction to rebuild and re-release the wheel on every MLX bump.

That is an honest description of a real failure mode. It means you cannot casually upgrade MLX in the environment, and it means the project carries a maintenance obligation proportional to MLX's release cadence. It also means a source build is not a casual operation: you need maturin, a Rust toolchain, and MLX headers that match.

The second dependency is stranger. mlx-lm is pulled not from a release but from a Git SHA dated 2026-08-25, with a comment stating that the revision requires MLX >=0.32.1 and includes cache APIs not available in the latest release (0.31.3). Depending on a commit rather than a tag is a deliberate choice here, and it is the kind of thing that makes reproducible installs harder for anyone building their own lockfile. The README does not document a rollback procedure, and it does not document an uninstall path for the install script.

## Where vLLM Metal is the wrong tool

The plugin is macOS-only by construction. pyproject.toml marks the MLX dependencies with platform_system == 'Darwin' and platform_machine == 'arm64', and the README rules out Rosetta and x86_64 Python. If your deployment target is Linux, or a CI runner without Apple Silicon, nothing here applies.

There is also a maturity signal worth reading literally. The package classifiers include Development Status :: 3 - Alpha. The README's own news items quantify progress against the project's own earlier version, not against other engines: it states that v0.2.0 brought a unified paged varlen Metal kernel as the default attention backend, with 83x TTFT and 3.6x throughput compared to v0.1.0. Those are self-reported comparisons across the project's own history, and they say nothing about how it compares to a CUDA host or to mlx_lm's own server.

Finally, the supported model set is not the whole Hugging Face catalogue. The README defers to docs/supported_models.md, which is the file to read before assuming your checkpoint works. A model that runs under mlx_lm is not automatically a model that runs under this plugin, because the attention path is the plugin's own code.

## vLLM Metal compared with mlx_lm and vLLM-MLX

The natural alternative is mlx_lm, the MLX library the plugin itself depends on for token-wise model layers. The difference is architectural rather than numerical. Running mlx_lm directly gives you model loading and generation in Python without vLLM's scheduler, paged block manager or OpenAI-compatible server. vLLM Metal keeps those upstream components and replaces only the compute and attention layer. If you want a script that generates text, mlx_lm is less machinery. If you want a server that batches concurrent requests with paged KV cache, the plugin is the one that provides it.

A second comparison point is vLLM-MLX, which people search for alongside this project. The repository does not describe vLLM-MLX's internals, so the only defensible statement is about this project's own positioning: vLLM Metal is described as a community maintained hardware plugin inside the vllm-project organisation, with a documented split between upstream vLLM components and its own attention path. Anyone choosing between the two should compare their supported-model lists and their MLX version policies directly, because both are MLX-backed and both will inherit MLX's release cadence.

## Maintenance, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-09-14. Releases are frequent: v0.28.0 on 2026-09-01, v0.29.0 on 2026-09-11, and a development build v0.29.0.dev20260914040516 on 2026-09-14. Version numbers track upstream vLLM's 0.29 line rather than the plugin's own history, which is consistent with the plugin model. Note that pyproject.toml declares version 0.29.0 while Cargo.toml declares 0.3.0 for the Rust crate; they version independently.

Licensing is Apache-2.0 across the repository: the LICENSE file, the pyproject.toml license field, the Cargo.toml license field, and the SPDX header at the top of pyproject.toml all agree. That is permissive and compatible with commercial use, but it is worth noting that the plugin depends on mlx-lm at a specific Git commit, so your own distribution obligations depend on what that commit contains. That is a question for your legal team, not something this article can settle.

The upgrade cost is the part to plan for. Because mlx is pinned exactly and the prebuilt shared object is ABI-bound to it, upgrading MLX means waiting for a matching vllm-metal wheel or rebuilding from source with maturin and a Rust toolchain. Budget for that rather than treating MLX upgrades as routine.

## Conclusion

Adopt vLLM Metal if you serve models on an Apple Silicon Mac and want the vLLM request API, scheduler and paged block manager rather than a separate MLX-only server. Skip it if you need x86_64 or Rosetta Python, a Linux or CUDA host, or a stable interface, since the project classifies itself as Alpha and pins mlx to an exact version. Before committing, check docs/supported_models.md for your model and confirm your Python is native arm64 3.12.

## FAQ

### How do I install vLLM Metal?

The README gives a single install command that pipes install.sh to bash, which creates an environment at ~/.venv-vllm-metal by default and installs the plugin, vLLM core and related libraries. Activating that environment puts the vllm CLI on your PATH.

### What is vLLM Metal?

It is a community maintained hardware plugin that enables vLLM to run on Apple Silicon Macs using MLX as the primary compute backend, unifying MLX and PyTorch under a single lowering path. Upstream vLLM supplies the API server, scheduler and paged block manager, while the plugin owns the attention path.

### How do I use vLLM Metal?

After installing and activating ~/.venv-vllm-metal, the vllm CLI becomes available and the README says you can access vLLM right away. It directs you to the official vLLM CLI guide for command usage.

## Sources

- [License: Apache-2.0](https://github.com/vllm-project/vllm-metal/blob/main/LICENSE)
- [Project website](https://docs.vllm.ai/projects/vllm-metal)
- [README](https://github.com/vllm-project/vllm-metal/blob/main/README.md)
- [Releases](https://github.com/vllm-project/vllm-metal/releases)
- [vllm-project/vllm-metal on GitHub](https://github.com/vllm-project/vllm-metal)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vllm-project-vllm-metal
