mistral.rs: a Rust inference engine with OpenAI and Anthropic compatible serving
Fast, flexible LLM inference
At a glance
- What is it?
- mistral.rs is a Rust LLM inference engine that loads Hugging Face and GGUF models automatically and serves them behind OpenAI and Anthropic compatible endpoints. Its own benchmarks show it ahead of llama.cpp on quantized prefill and behind vLLM on BF16 prefill for larger models.
- Who is it for?
- Adopt mistral.rs if you want a single Rust or Python process that loads GGUF and Hugging Face checkpoints, serves OpenAI and Anthropic compatible endpoints, and gives you a CLI plus SDKs to script against. Do not adopt it if your workload is BF16 serving of large MoE models on datacenter GPUs, where the project's own v0.8.2 numbers put vLLM well ahead, or if you need a documented rollback path, which the README does not describe.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What mistral.rs solves, and who it is aimed at
Running a local model usually means choosing between a C++ runtime with a GGUF file and a Python serving stack with a Hugging Face checkpoint. mistral.rs targets both in one engine. The README describes automatic model loading: architecture, weight format and chat template are detected for supported Hugging Face models and GGUF files, with flags available when you want to select explicitly. That detection step is the product's centre of gravity. You point the CLI at a repository or a local file and it works out what it is looking at.
The audience is engineers who want to self-host and script against the result. There is a Rust crate on crates.io and a Python SDK, so the same engine is reachable from a compiled service or a notebook. The workspace layout confirms this: mistralrs-core, mistralrs-pyo3, mistralrs-cli, mistralrs-server-core and a set of narrower crates for quantization, paged attention, vision, audio, MCP and sandboxed code execution. If you only wanted a chat window, this is more machinery than you need.
The project is not archived and the last push was on 2026-09-08. The most recent release listed is v0.9.3 on 2026-09-07, with v0.9.2 and v0.9.1 in the weeks before. The workspace version in Cargo.toml reads 0.9.3, matching the tag.
How model loading, quantization and serving fit together
The README describes two load paths. For GGUF, you either pass a local file with -f or select a published artifact with --quant; tokenizer, configuration and multimodal projector files are discovered when the available metadata identifies them unambiguously. That last clause matters. Discovery depends on metadata being unambiguous, which is a weaker guarantee than a lookup table, and the README does not enumerate the cases where ambiguity causes a failure.
For other Hugging Face repositories, --quant uses a prebuilt UQFF when one is available and otherwise applies ISQ. UQFF is the project's own quantized format, and the v0.8.2 benchmark tables compare UQFF q8 against llama.cpp's GGUF Q8_0 directly. ISQ is the fallback path, so the quantization you get depends on whether someone has already published a UQFF artifact for that model.
Serving is one process. The README states that mistralrs serve exposes OpenAI-compatible /v1 endpoints and Anthropic-compatible /v1/messages and /v1/messages/count_tokens endpoints from the same process. It also exposes a /metrics endpoint in Prometheus format, recording per-request counts and latency labeled by method, route and status, and a built-in web UI at /ui that can be disabled with --no-ui. The Dockerfile confirms the default port: it sets EXPOSE 1234 as the default port of mistralrs serve, and sets HF_HOME to /data so downloaded models persist in a mounted volume.
Around the HTTP surface sits an agentic runtime: web search, local Python code execution, shell execution, Skills bundles uploaded to /v1/skills, file inputs through /v1/files, and session management. The mistralrs-mcp, mistralrs-code-exec and mistralrs-sandbox crates in the workspace are the implementation of that runtime, which means tool execution is not bolted on from outside.
Installing mistral.rs and serving a first model
The README gives a one-line installer for Linux and macOS. It downloads a self-contained prebuilt binary for your platform (Metal on Apple Silicon; per-GPU CUDA or CPU on Linux; CPU on Windows) and falls back to a source build if none matches.
curl -fsSL https://mistralrs.dev/install.sh | shOn Windows the README gives the PowerShell equivalent.
irm https://mistralrs.dev/install.ps1 | iexAfter installation the binary is mistralrs. The Dockerfile sets its entrypoint to mistralrs with CMD ["--help"], so running the image without arguments prints usage. That is the cheapest way to confirm which subcommands your build actually has before committing to a model download.
The README's own quick-start route for GGUF is to pass a local file with -f, or to name a published artifact with --quant. The documentation also describes a mistralrs tune subcommand that recommends quantization and device mapping from the model config and your detected hardware. Run that before your first real load: the recommendation is derived from your detected hardware, so it is the one piece of guidance that is specific to your machine rather than to the model. If you are scripting against the server instead of the CLI, the Python SDK is documented under a separate getting-started guide on docs.mistralrs.dev.
One practical note from the repository rather than the README: the Makefile has docs-regen and docs-check targets that regenerate the CLI reference, the OpenAPI document and the supported-models list, and then verify the committed copies match. The generated reference is therefore tied to the code, and the docs-check target is what keeps it that way.
Where the benchmarks put mistral.rs, and where they do not
The project publishes its own v0.8.2 CUDA numbers with commands, model revisions and host metadata in releases/v0.8.2/report.md. Read them closely, because they do not say one thing.
Against llama.cpp on Q8, mistral.rs is ahead across the table. On Gemma 4 E4B prefill it reports 7395.7 tokens per second on GB10 against 3973.7, and 27705.6 against 11992.4 on B200. Decode margins are much thinner: 44.1 against 40.5 on GB10, and for Gemma 4 26B-A4B on GB10, 46.8 against 46.4. That is a near tie, and it is the quantized path where the two projects are most directly comparable.
Against vLLM in BF16 the picture splits by model size. For Gemma 4 E4B, mistral.rs reports 5838.9 prefill on GB10 against 5812.9, and 43547.8 against 39431.2 on B200. For Gemma 4 26B-A4B, vLLM is far ahead on prefill: 3878.6 against 592.2 on GB10, 28532.8 against 3467.3 on B200, and 26295.9 against 2766.0 on H100 SXM. Decode for that model is closer, and on B200 vLLM leads there too, 220.2 against 159.6.
So the honest reading is that mistral.rs is competitive on quantized inference and on smaller models, and that BF16 serving of a larger mixture-of-experts model on datacenter GPUs is not where its numbers are strongest. These are the maintainers' own measurements, published with their methodology, and they are the only performance figures I can point to. They are not a reason to skip your own measurement on your own hardware.
Limitations the README does not resolve
The documentation is uneven, and the gaps are not random. The README describes loading, serving, quantization and the agentic runtime in reasonable detail and links out to docs.mistralrs.dev for each. What it does not describe is rollback. There is no documented procedure for reverting to a previous binary or pinning a version, which matters because the installer fetches whatever the current release is. If you need reproducible deployments, the Dockerfile is a better starting point than the install script, but note what it says about itself: the prebuilt binary is staged into dist/ by the release workflow and copied in, and the image is not compiled here. Building the image therefore depends on artifacts produced by that workflow rather than on a self-contained build.
Model support is also a moving target. The README advertises automatic detection for supported models, and the repository has a docs-check target that verifies the supported-models list matches the code, but supported is doing real work in that sentence. A model that is not in the generated list falls back to explicit flags at best. Check the supported models reference before planning around a specific checkpoint.
The agentic features widen the surface further. Shell execution, local Python execution and file inputs through /v1/files are all described as built in. The README does not describe the isolation model in the sections shown here, and the existence of a mistralrs-sandbox crate tells you isolation is treated as a separate concern rather than assumed. If you expose these endpoints beyond a trusted network, that is the part to read before anything else.
Finally, the hardware story is uneven by design. macOS gets Metal, Linux gets per-GPU CUDA or CPU, Windows gets CPU. There is no Windows GPU path in the installer description, and no ROCm mention anywhere in the repository's own description of supported hardware.
mistral.rs against llama.cpp and vLLM
The two comparisons people search for are mistral.rs versus llama.cpp and mistral.rs versus vLLM, and the projects differ in more than speed.
llama.cpp is the closest comparison on the quantized path. Both consume GGUF, and the v0.8.2 tables compare them on the same models and hardware. The architectural difference is where quantization comes from: llama.cpp centres on the GGUF format and its own quantized kernels, while mistral.rs treats GGUF as one input among several and prefers its own UQFF artifacts when they exist, falling back to ISQ on the fly when they do not. If your workflow is already a directory of GGUF files, llama.cpp asks nothing new of you. mistral.rs asks you to accept that the quantized artifact may be produced at load time instead of downloaded.
The vLLM comparison is about scope. vLLM is a Python serving stack built around high-throughput GPU inference, and the BF16 numbers above show what that focus buys on a larger MoE model. mistral.rs is a Rust engine with a CLI, a Rust crate, a Python SDK, and a built-in agentic runtime with tool execution and a web UI. If you want a Python-native serving layer and nothing else, vLLM is the narrower tool. If you want one binary that can also run tools and expose both OpenAI and Anthropic compatible endpoints, that combination is not what vLLM is for.
The searchers asking about mistral.rs versus candle should note that the relationship is not a comparison. Cargo.toml pins candle-core, candle-nn, candle-flash-attn-v3 and candle-metal-kernels to a specific git revision of huggingface/candle. mistral.rs is built on candle, not alongside it.
Licence and the cost of keeping up
The workspace declares license = "MIT" and the repository carries a LICENSE file, so the code is permissively licensed. That covers the engine. It does not cover the model weights you load with it, which carry their own terms from wherever you fetched them, and it does not cover the prebuilt binaries distributed by the install script. If your legal position depends on the exact terms of a redistributed binary rather than the source, that is a question for your own counsel, not something the repository answers.
The upgrade cost is visible in the release cadence. v0.9.1, v0.9.2 and v0.9.3 landed within roughly three weeks of each other, and the README's Latest section lists changes to model support, GGUF discovery, Skills, file inputs and a new block-diffusion generation path. Features arrive quickly and the supported-models list is regenerated from the code. Pin a version, and expect the list of models your pinned build handles to be a snapshot. The workspace also declares rust-version = "1.94", which is the floor for anyone building from source rather than using the prebuilt binary.
Editorial conclusion
Adopt mistral.rs if you want a single Rust or Python process that loads GGUF and Hugging Face checkpoints, serves OpenAI and Anthropic compatible endpoints, and gives you a CLI plus SDKs to script against. Do not adopt it if your workload is BF16 serving of large MoE models on datacenter GPUs, where the project's own v0.8.2 numbers put vLLM well ahead, or if you need a documented rollback path, which the README does not describe. Verify first: run mistralrs tune against your actual model and hardware, and check the supported models reference before assuming your architecture is covered.
Frequently asked questions
What is mistral.rs?
It is a Rust LLM inference engine that loads supported Hugging Face models and GGUF files, detects architecture, weight format and chat template automatically, and serves them through OpenAI-compatible /v1 endpoints and Anthropic-compatible Messages endpoints from one process. It ships a CLI, a Rust crate and a Python SDK.
How does mistral.rs performance compare with llama.cpp?
In the project's own v0.8.2 CUDA tables, mistral.rs UQFF q8 leads llama.cpp GGUF Q8_0 on prefill by a wide margin and on decode by a narrow one, with Gemma 4 26B-A4B decode on GB10 reported at 46.8 against 46.4 tokens per second. Those are the maintainers' measurements, published with commands and host metadata in releases/v0.8.2/report.md.
How is mistral.rs related to candle?
mistral.rs is built on candle rather than being an alternative to it. Cargo.toml pins candle-core, candle-nn, candle-flash-attn-v3 and candle-metal-kernels to a specific git revision of huggingface/candle.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ericlbuehler-mistral-rs)