# Rapid-MLX: Apple Silicon LLM Server Built for Tool-Calling Reliability

> Rapid-MLX is an Apache 2.0 open-source LLM inference server for Apple Silicon Macs, built on Apple's MLX framework, with 27 tool-call parser modules, radix prefix caching saved to disk across restarts, and drop-in OpenAI and Anthropic API compatibility. It ships as both a Python package and a signed macOS desktop app.

**raullenchai/Rapid-MLX** — The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.

- Repository: https://github.com/raullenchai/Rapid-MLX
- Website: https://pypi.org/project/rapid-mlx
- Stars: 3,871 · Forks: 424
- Language: Python
- License: NOASSERTION
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/raullenchai-rapid-mlx

## What Rapid-MLX Solves and Who Uses It

Running large language models locally on Apple Silicon is possible with several tools, but tool calling, the mechanism that lets AI agents invoke external functions, has historically been unreliable in local inference servers. Different models use different token formats to delimit tool calls, and a server that handles one format fails silently on another. Rapid-MLX was built to address this. The README describes it as 'focused on reliable tool calling for coding agents.' It supports an OpenAI-compatible API at /v1/chat/completions and /v1/responses, plus an Anthropic Messages-compatible API at /v1/messages, making it a drop-in replacement for clients that already target either provider. The README lists tested compatibility with coding agents including Claude Code, Aider, OpenHands, Codex, and others. Five Tier-1 agents are run end-to-end on real model weights before each release.

## The 27-Parser Tool-Calling Architecture

Each model family uses a different format to encode tool calls in its output token stream. A parser that works for Qwen's format may fail on Llama's, and a model that was not explicitly tested may use a format that matches no existing parser. Rapid-MLX addresses this with 27 parser modules under rapid_mlx/tool_parsers, including an auto-detect fallback that tries to identify the format at inference time. The README links to this directory directly in its comparison table. The auto-detect fallback means the server degrades gracefully when it encounters an untested format rather than producing an unparsed string that breaks the calling agent. This is the main functional difference between Rapid-MLX and alternatives like mlx-lm's built-in server, which supports tool calls for models whose tokenizer explicitly declares support, but has no auto-detect path for undeclared formats.

## Installing Rapid-MLX and Starting the Server

Rapid-MLX is available as a PyPI package, a Homebrew formula (formulae.brew.sh/formula/rapid-mlx), and a macOS desktop app. The pyproject.toml records the package name as rapid-mlx, requires Python 3.10 or later, and classifies the package as macOS-only. The install section of the README was in the CLI and server section, which covers all three paths. A signed macOS desktop app is available at rapidmlx.com/desktop and from the GitHub releases page tagged with rapid-mac-v; it handles model management, vision, files, voice, and image generation in a single window without the command line. Once the CLI is installed, the server command is straightforward. The pyproject.toml comments describe the minimum invocation as:

```bash
rapid-mlx serve <text-model>
```

The Makefile provides a developer test suite. Running `make smoke` executes lint, a CLI-to-config fidelity audit, and the unit test suite in sequence. Running `make stress` exercises 8 concurrent scenarios against a live server. These targets are for contributors, not end users, but they show what the project verifies before each release.

## Radix Prefix Cache and Disk Persistence

Prompt caching is one of the largest practical speed gains in repeated LLM calls. When a coding agent sends the same system prompt and context on every turn, a prefix cache means only the new tokens need to be computed. Rapid-MLX uses a radix prefix cache in memory. Beyond that, it saves the cache state to disk on shutdown and restores it at startup, so the cache is warm across server restarts. It also supports state snapshots for hybrid models. The comparison table in the README shows that Ollama reuses only the previous request's KV cache, with no cross-restart persistence. mlx-lm's server holds the prefix cache in memory only. The disk-backed approach in Rapid-MLX is particularly relevant for coding agent workflows where the development session may span multiple server restarts but the underlying project context rarely changes.

## Measured Speed and Where Rapid-MLX Is Not Faster

The README states a measured 3.0x aggregate decode throughput advantage over Ollama at 8 concurrent streams using Qwen3.6-35B-A3B on a 32 GB M2 Pro Mac mini (82.9 vs 27.2 tok/s). The description field in the repository claims 4.2x faster than Ollama, but the README itself gives 3.0x as the verified aggregate decode number and links to the method and raw data. The README is explicit about where the advantage does not apply: single-stream decode was about 1.5x; whole-batch throughput including prefill was 1.6x; dense 12B models were no faster; and llama.cpp-family engines prefilled cold prompts faster. The 3.0x figure applies specifically to a mixture-of-experts model at 8 concurrent streams. Teams with a single-user local setup will see a smaller gain.

## macOS Only: No Windows or Linux Desktop

Rapid-MLX depends on Apple's MLX framework, which runs exclusively on Apple Silicon hardware. The pyproject.toml classifies the package as 'Operating System :: MacOS' and requires Python 3.10 or later. The README notes that Windows and Linux desktop builds are not available yet, and the Homebrew formula also targets macOS. This is not a deployment-time limitation that can be worked around with Docker: Docker on macOS does not expose the Metal GPU required by MLX. Teams running inference on Linux servers, cloud GPUs, or Windows workstations cannot use Rapid-MLX and should look at vLLM or llama.cpp-based servers instead. The CLI and Python library also require macOS; the restriction is not limited to the desktop app.

## Rapid-MLX vs Ollama

Ollama is the most widely used local LLM tool and runs on macOS, Linux, and Windows. It uses the GGML/GGUF engine by default, with an optional MLX engine for safetensors models on macOS. Rapid-MLX uses MLX exclusively. The practical differences, per the comparison table in the README, are: Ollama defaults to one parallel request per model (configurable via OLLAMA_NUM_PARALLEL), while Rapid-MLX uses continuous batching throughout. Ollama reuses only the previous request's KV cache, with no cross-restart persistence, while Rapid-MLX saves its radix prefix cache to disk. Ollama supports tool calling through a defined format, while Rapid-MLX has 27 parser modules and an auto-detect fallback. If you need cross-platform support or GGUF model compatibility, Ollama is the practical choice. If you are on Apple Silicon, use MoE models with many concurrent agent calls, and need reliable tool calling, Rapid-MLX is the stronger option.

## Conclusion

Rapid-MLX is the most capable option for running LLM inference on Apple Silicon when tool calling is the primary requirement and you are willing to accept a macOS-only constraint. The 27-parser tool-call architecture and the disk-persisted prefix cache address real friction points in local agent workflows. Teams on Linux, Windows, or non-M-series hardware have no supported path. Before adopting it as a drop-in replacement, check the tested compatibility matrix at rapidmlx.com/docs/matrix.html for exact API coverage for the specific client you use.

## FAQ

### What is Rapid-MLX?

Rapid-MLX is an open-source LLM inference server for Apple Silicon Macs, built on Apple's MLX framework. It exposes OpenAI-compatible and Anthropic-compatible endpoints and ships 27 tool-call parser modules to handle diverse model output formats reliably.

### What are the differences between MLX and Rapid-MLX?

MLX is Apple's open-source machine learning framework; mlx-lm is a separate package built on it that provides basic LLM serving. Rapid-MLX is built on MLX and mlx-lm but adds 27 tool-call parsers, a disk-persisted radix prefix cache, Anthropic Messages API compatibility, and continuous batching as a full server implementation focused on coding agent workloads.

### How does Rapid-MLX compare to llama.cpp?

llama.cpp targets GGUF quantized models and runs on Linux, Windows, and macOS with CPU and GPU backends. Rapid-MLX targets Apple Silicon via MLX and does not support GGUF models. The README notes that llama.cpp-family engines prefilled cold prompts faster in their benchmark, while Rapid-MLX had higher aggregate decode throughput at 8 concurrent streams for MoE models.

## Sources

- [Issues](https://github.com/raullenchai/Rapid-MLX/issues)
- [Project website](https://pypi.org/project/rapid-mlx)
- [raullenchai/Rapid-MLX on GitHub](https://github.com/raullenchai/Rapid-MLX)
- [README](https://github.com/raullenchai/Rapid-MLX/blob/main/README.md)
- [Releases](https://github.com/raullenchai/Rapid-MLX/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/raullenchai-rapid-mlx
