# candle-vllm: a Rust inference server for local LLMs with an OpenAI compatible API

> candle-vllm serves local models from Rust, with CUDA and Metal backends, paged attention and a built-in Web UI. It is a good fit when you want a single binary instead of a Python stack, and a poor fit when your model is not on its supported list.

**EricLBuehler/candle-vllm** — Efficent platform for inference and serving local LLMs including an OpenAI compatible API server.

- Repository: https://github.com/EricLBuehler/candle-vllm
- Stars: 728 · Forks: 91
- Language: Rust
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ericlbuehler-candle-vllm

## What candle-vllm solves, and for whom

Serving a local model usually means a Python process, a virtual environment, and a dependency tree that changes under you. candle-vllm takes the other route: it is a Rust crate and binary, published as candle-vllm version 0.9.1 in Cargo.toml, that loads a model and exposes an OpenAI compatible HTTP API on port 2000 by default. The README describes it as an "Efficient, easy-to-use platform for inference and serving local LLMs including an OpenAI compatible API server."

The audience is narrow and specific. You are comfortable building Rust, or you are willing to use the prebuilt install script or Docker image. You have an NVIDIA GPU on Linux or an Apple Silicon machine, because those are the two backends the README lists: CUDA on Linux and Metal on macOS, from the same codebase and the same API. You want an endpoint that existing OpenAI client libraries can talk to without a translation layer. If any of those do not hold, the project costs more than it returns.

The README's feature list is where the design intent shows. PagedAttention for KV cache management, continuous batching for requests arriving over time, chunked prefilling with a default chunk size of 8K, prefix caching, CUDA Graphs, and multi-GPU plus multi-node inference over TCP. That is the vocabulary of a serving engine, not of a script that loads a checkpoint.

## The serving path: paged KV cache, continuous batching, and TurboQuant

Requests enter through an axum HTTP server, which is the web framework listed first in Cargo.toml. From there the work is the standard serving loop. Continuous batching means the engine does not wait for one request to finish before starting the next; it batches decoding steps across requests that are in flight at the same time. Chunked prefilling breaks a long prompt into chunks, 8K by default, so a single enormous prompt does not block the batch.

The memory side is where candle-vllm makes its own choices. PagedAttention manages the KV cache in blocks rather than one contiguous buffer per sequence, and prefix caching reuses the KV state of a shared prefix across requests, which matters when many callers send the same system prompt. On top of that sits TurboQuant, which the README describes as 2 to 4 bit KV cache compression that "extends context up to 4.7× with minimal quality loss." The supported modes are turbo8, turbo4 and turbo3, with native flash attention kernels. Treat the 4.7x figure as the project's own claim; the README does not give the evaluation setup behind it.

A second layer of compression applies to weights rather than the cache. In-situ quantization with --isq q4k converts a model on the fly, and the README also lists GPTQ, AWQ, Marlin, MXFP4, NVFP4, block-wise FP8 and GGUF (including split shards) as input formats. That is an unusually wide format surface for a project this size, and it is the main reason the performance table has so many rows.

The trait-based architecture is the extension point. The README describes it as enabling "rapid implementation of new model pipelines," which is accurate but also a warning: a model family that is not already implemented requires Rust work against those traits, not a config change.

## Installing candle-vllm and serving a first model

The README gives three install paths. The fastest is the one-line script, which fetches a DEB package or a binary. The script is hosted on the project's GitHub Pages site, so it is a download rather than a build.

```bash
curl -sSL https://ericlbuehler.github.io/candle-vllm/install.sh | bash
```

If you would rather build, the CUDA path installs the binary with the cuda, nccl, flashinfer and cutlass features. The README notes that flashinfer and cutlass should be removed for sm_70 and sm_75 architectures, which is the first sign that the feature set is architecture-dependent.

```bash
git clone git@github.com:EricLBuehler/candle-vllm.git
cd candle-vllm
cargo install --features cuda,nccl,flashinfer,cutlass --path .
```

On macOS the equivalent build uses the metal feature only, and the README shows the same command shape with a different feature list.

```bash
cargo install --features metal --path .
```

The Docker route goes through the build script in the repository root. The README shows the feature string as the first argument and notes that a custom SM version and CUDA version can be passed as the second and third arguments, with sm_90 and 13.0.0 as the example.

```bash
./build_docker.sh "cuda,nccl,flashinfer,cutlass"
```

Once the binary exists, serving a model from Hugging Face is one command. The model ID is passed with --m, and --ui-server adds the built-in ChatGPT-style Web UI. The README states that the UI server listens on the API port minus one, so with the default API port of 2000 the UI is on 1999.

```bash
candle-vllm --m Qwen/Qwen3.6-27B-FP8 --ui-server
```

For local weights the same flag takes a directory of safetensors, a single GGUF file, or a directory of GGUF files that the loader auto-detects. A GGUF file needs --f when the directory holds more than one quantization.

```bash
candle-vllm --m unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF --f Qwen3-30B-A3B-Instruct-2507-Q4_K_M.gguf --ui-server
```

Multi-GPU selection uses --d with a comma-separated device list, as in --d 0,1,2,3,4,5,6,7 for an eight-way run. The repository also ships Python examples under examples/, including chat.py and benchmark.py, which are the natural place to start if you want to hit the endpoint from an existing OpenAI client rather than from the UI.

## Where candle-vllm stops being the right tool

The hardest constraint is model coverage. The README's performance table names LLAMA, Mistral, Phi3 and Phi4, Qwen2 and Qwen3 dense and MoE, Qwen3-Next, Qwen3.5 through 3.8, Yi, StableLM, Gemma-2 through Gemma-4, DeepSeek V2, V3, V3.2 and R1, QwQ-32B, GLM4, GLM4.7 Flash, GLM-5.2, Llama4 and MiniMax-M2.5 and M2.7. Several rows read TBD rather than a number, which means the row is listed but the project does not publish a decode figure for it. If your model is not in that table, there is no documented path short of implementing a pipeline in Rust.

Hardware is the second boundary. There is no CPU backend in the README's cross-platform row, which names CUDA on Linux and Metal on macOS. The build script's supported architecture set is 7.0, 7.5, 8.0, 8.6, 8.9 and 9.0, so older cards are out, and the README's own note about dropping flashinfer and cutlass for sm_70 and sm_75 means the fastest kernels are not available on those parts. The Dockerfile defaults CUDA_COMPUTE_CAP to 80, so a build for a different architecture needs that argument set explicitly.

The performance table is also easy to misread. It is single-request decode speed at 4k input and 1k output on an 80GB Hopper part, with tensor parallelism noted per row. It says nothing about throughput under concurrent load, and it says nothing about time to first token. If your workload is many short requests, those numbers do not predict your result. The DeepSeek row is the clearest warning: roughly 20 tokens per second for a 671B AWQ model across eight GPUs with offloading is a usable demonstration, not a production figure.

Finally, the README does not document rollback, upgrade or downgrade procedures for a running deployment. There is a version field in Cargo.toml and dated releases, but no migration notes in the README.

## candle-vllm compared with vLLM and llama.cpp

The obvious comparison is with vLLM, and the names invite it. Both serve an OpenAI compatible API, both use PagedAttention, and both do continuous batching. The difference is the runtime. vLLM is a Python project built on PyTorch, with a large model zoo and a plugin ecosystem that follows from that. candle-vllm is Rust on top of the Candle tensor library, and Cargo.toml pins candle-core and candle-nn to a specific git revision of a fork rather than a crates.io release. That pinning is a deliberate trade: the project controls its tensor layer, and you inherit the consequences at build time.

The second comparison is with llama.cpp. That project also targets local inference and GGUF, and candle-vllm reads GGUF too, including split shards. The difference is the serving model. llama.cpp centers on a single-process runner with a server wrapper, while candle-vllm is built around batched serving from the start: continuous batching, chunked prefilling, prefix caching, and multi-node coordination over TCP. If you want one user talking to one model on a laptop, llama.cpp's shape fits better. If you want an endpoint that several clients hit at once, candle-vllm's design addresses that case directly.

Neither comparison is settled by the README. What the README does establish is that candle-vllm is not trying to be a general inference framework. It is a serving engine with a fixed, documented model list and two GPU backends.

## Licence and the cost of keeping up

The repository is MIT licensed, and the LICENSE file sits at the top level. MIT is permissive: you can use, modify and redistribute the code, including in commercial products, provided the copyright notice and permission notice travel with it. That is the general shape of the licence, not legal advice for your situation.

The practical licence question is not candle-vllm's own terms but the models you serve with it. The README's examples pull weights from Hugging Face under IDs such as Qwen, zai-org and unsloth, and candle-vllm's MIT licence says nothing about those weights. Check each model's own licence separately.

Upgrade cost is the part worth budgeting for. The last push to master was on 2026-09-07, and the release history shows v0.9.1 on 2026-08-20, v0.9.0 on 2026-07-29 and v0.8.9 on 2026-07-22. Three releases in about a month is a fast cadence, and the dependency list explains why that cadence has a price: candle-core, candle-nn, attention-rs and range-checked are all git dependencies pinned by revision, and the Dockerfile exposes CUDA_VERSION, CUDA_FLAVOR and CUDA_COMPUTE_CAP as build arguments. A version bump can therefore mean a new pinned revision, a new CUDA base image, or a different feature set for your GPU architecture. Plan to rebuild and re-verify rather than to hot-swap a binary, and read the release notes for each version you move across.

## Conclusion

Adopt candle-vllm if you want a Rust binary serving an OpenAI compatible endpoint on CUDA or Metal and your model family appears in the README's performance table. Do not adopt it if you need a model outside that list, since the pipeline is trait-based and a new architecture means writing Rust. Before committing, verify the build features for your GPU architecture: the README says to remove flashinfer and cutlass for sm_70 and sm_75, and the Dockerfile defaults CUDA_COMPUTE_CAP to 80.

## FAQ

### What is candle-vllm?

It is a Rust inference and serving platform for local LLMs, built on the Candle tensor library. It exposes an OpenAI compatible API server on port 2000 by default and can also launch a built-in ChatGPT-style Web UI with the --ui-server flag.

### How do I install candle-vllm?

The README gives three options: a one-line install script that fetches a DEB package or binary, a cargo install from a cloned repository with features such as cuda,nccl,flashinfer,cutlass or metal, and a Docker build through the repository's build_docker.sh script.

### Does candle-vllm run on macOS?

Yes. The README lists CUDA on Linux and Metal on macOS as the two supported backends, built from the same codebase with the same API. The macOS build uses the metal feature.

### Which model formats does candle-vllm support?

The README lists safetensors directories, GGUF files including split shards, and quantized formats: GPTQ, AWQ, Marlin, MXFP4, NVFP4, block-wise FP8, and in-situ quantization via --isq. FP8 and FP4 models are documented separately in the usage section.

### What is the difference between candle-vllm and vLLM?

Both serve an OpenAI compatible API with PagedAttention and continuous batching. candle-vllm is written in Rust on Candle, with candle-core and candle-nn pinned to a specific git revision, while vLLM is a Python project built on PyTorch with a larger model zoo.

## Sources

- [EricLBuehler/candle-vllm on GitHub](https://github.com/EricLBuehler/candle-vllm)
- [Issues](https://github.com/EricLBuehler/candle-vllm/issues)
- [License: MIT](https://github.com/EricLBuehler/candle-vllm/blob/master/LICENSE)
- [README](https://github.com/EricLBuehler/candle-vllm/blob/master/README.md)
- [Releases](https://github.com/EricLBuehler/candle-vllm/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ericlbuehler-candle-vllm
