# mlx-serve: a Zig inference server for Apple Silicon that speaks OpenAI, Anthropic and Ollama APIs

> mlx-serve runs MLX and GGUF models on Apple Silicon behind OpenAI, Anthropic and Ollama compatible endpoints, with a menu-bar app called MLX Core. It is a good fit if your clients already speak one of those wires; it is not a fit if you are not on a recent macOS release.

**ddalcu/mlx-serve** — Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.

- Repository: https://github.com/ddalcu/mlx-serve
- Website: http://mlxserve.com/
- Stars: 1,682 · Forks: 173
- Language: Zig
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ddalcu-mlx-serve

## The gap mlx-serve is trying to fill on a Mac

Local inference on Apple Silicon has been split between two habits. One is the Python route, where mlx-lm and friends give you native MLX execution but no server, no CLI, and no drop-in API for tools that expect one. The other is the GGUF route, where llama.cpp and Ollama give you the API surface but treat MLX as an afterthought. mlx-serve takes the position that a Mac user should not have to choose. The README describes it as a native Zig server that runs MLX-format models and every GGUF on HuggingFace, and exposes OpenAI-compatible and Anthropic-compatible HTTP APIs from the same process. The audience is narrow and specific: people on Apple Silicon who want to point Claude Code, the OpenAI SDK, Continue, Cursor or Open WebUI at a local endpoint without a Python runtime in the loop. The README states the project requires no Python, no cloud and no Electron, and the repository layout backs that up, with a build.zig at the top level and a src/ directory rather than a pyproject.toml. The README also claims decode speed above LM Studio on identical MLX weights, with a table entry of +26 percent geomean. That number comes from the project's own comparison and its own benchmarks.md, so treat it as a claim to reproduce on your hardware, not as an independent measurement.

## How the server, the app and the API surfaces fit together

The architecture visible in the repository is a single server binary plus a macOS app that bundles it. The server owns model loading and inference and listens on http://localhost:11234. The README is explicit that MLX Core, the menu-bar app, runs the same binary as the CLI on the same port, so an app running in your menu bar is enough for an external client to connect. That is a cleaner arrangement than shipping two engines, and it means the app is a front end rather than a separate product. On top of that one port, the server presents several protocol surfaces. The OpenAI-compatible surface covers the usual clients. The Anthropic Messages API surface is what makes Claude Code work against it. The README also says the server speaks the Ollama API at /api/chat, /api/generate, /api/tags, /api/embed and /api/pull, which is a deliberate compatibility play: Ollama-connected tools such as Raycast, Obsidian, Enchanted and Open WebUI keep working if you repoint them. Model loading is on demand by name, so a serve process can hold many pulled models and load the one a request names. The repository also lists containers/ and a Brewfile, which suggests the build is meant to be reproducible from a clean machine rather than assembled by hand.

## Installing mlx-serve with Homebrew and running a first model

The README gives Homebrew as the shortest install path, with a separate cask for the app and a formula for the CLI and server. The tap is added from the repository URL, then you choose one of the two packages. The cask is what the README recommends, because it brings the GUI, but the formula is enough if you only want a server.

```bash
brew tap ddalcu/mlx-serve https://github.com/ddalcu/mlx-serve
brew install --cask mlx-core   # the app (recommended)
brew install mlx-serve         # CLI + server only, no GUI
```

After that, the CLI follows an Ollama-shaped workflow. The run subcommand downloads a model, serves it and drops you into a terminal chat. The pull subcommand downloads without serving, and the README notes the download is resumable and comes straight from Hugging Face. The list subcommand shows what is already on disk.

```bash
mlx-serve run gemma4        # downloads Gemma 4 E4B (4-bit), serves it, chats right in your terminal
mlx-serve pull qwen3.6:27b  # just download (resumable, straight from Hugging Face)
mlx-serve list              # what's on disk
mlx-serve serve             # serve everything you've pulled — models load on demand by name
```

The README states that short names, org/repo HuggingFace ids and name:tag all resolve, so qwen3.6:27b and a full repository path are both acceptable. For scripts and headless Macs, the README points to direct --model and --model-dir invocations and says every server flag is documented in docs/cli.md. Once a serve process is up, the practical first test is to point an existing client at http://localhost:11234 rather than writing a new one. If you are on the app path, the README says Claude Code and any OpenAI or Anthropic client can connect while the app is running, on the same port.

Building from source is a different route and has its own prerequisites. The README requires Xcode 26.2 or newer with the Metal Toolchain component, and gives a check for it: if xcrun -sdk macosx metal --version fails, run xcodebuild -downloadComponent MetalToolchain. The clone uses --recurse-submodules, then Brewfile installs cmake and webp, then ./app/build.sh produces the app and server with an ad-hoc signature.

```bash
git clone --recurse-submodules https://github.com/ddalcu/mlx-serve && cd mlx-serve
brew bundle install --file=Brewfile   # cmake + webp
./app/build.sh                        # app + server, ad-hoc signed
```

The README states that Zig, mlx and llama.cpp are pinned and fetched or built by the script, and that server-only builds are covered in docs/building.md.

## The macOS 26.2 floor and other things that will stop you

The hardest constraint is stated plainly in the README: macOS 26.2 or newer on Apple Silicon. That is not a soft recommendation. It rules out Intel Macs entirely, and it rules out any Mac that cannot move to that OS version. If your fleet is on an older release, nothing else in this article matters. The second constraint is the platform itself. There is no Linux or Windows server here, and the topics list Apple Silicon and macOS as the target. The third is the build toolchain if you go from source: Xcode 26.2 or newer plus the Metal Toolchain component, which the README treats as a separate download. A machine that cannot install that component cannot build the project. Beyond those, the honest limitation is that this is a young project with a fast release cadence. Three releases landed in the weeks before the last push on 2026-09-10, at v26.9.2, v26.9.1 and v26.8.11, and the release notes describe per-model settings, chat providers, terminals in the sidebar and a faster Flash Next path. That cadence is good for features and bad for stability expectations: flags, defaults and app behaviour can move between minor versions. There is also a needs_additional_testing.md file at the top level of the repository, which is a candid signal about where coverage is thin. Finally, the README's headline speed claim is a project-authored comparison. If you are switching specifically for the +26 percent figure, reproduce it on your own weights and your own prompts before you commit a workflow to it.

## Where mlx-serve is the wrong tool, and what to use instead

If you need to serve models to a team from a Linux box or a machine with an NVIDIA GPU, mlx-serve is simply not the answer, because it targets Apple Silicon and macOS only. For that case, vLLM or llama.cpp on Linux is the conventional choice. The more interesting comparison is with Ollama, because the two overlap heavily on the API surface and on the CLI habit. Ollama is cross-platform, has a large library of prebuilt GGUF models, and its API is the native one that mlx-serve emulates. mlx-serve's difference in approach is that it runs MLX natively and embeds llama.cpp for GGUF in the same process, so a single server covers both model formats rather than treating MLX as a special case. The README's comparison table also claims continuous batching, KV-cache quantization with 4-bit and 8-bit plus TurboQuant, speculative decoding with PLD, drafter and native MTP, and an Anthropic Messages surface, and marks Ollama as partial or absent on several of those rows. Read that table as the project's own positioning. The practical question is whether you need the MLX path and the Anthropic path at all. If you only run GGUF models and only speak the Ollama API, staying on Ollama costs you nothing and avoids the macOS 26.2 floor. If you want native MLX execution and Claude Code against a local endpoint, the overlap argument disappears and mlx-serve is doing something Ollama does not.

## Licence, release cadence and what an upgrade actually costs

The repository metadata reports the licence as NOASSERTION, while the README carries an MIT badge and a comparison table row that lists the licence as MIT. The repository also contains both LICENSE and LICENSE-APACHE-2.0 files plus a NOTICE file, which is the pattern of a dual or mixed licensing arrangement rather than a single plain MIT grant. That is worth reading directly rather than inferring from a badge. This is not legal advice; if the distinction matters for your use, read LICENSE and NOTICE and get your own answer. On upgrades, the release history shows a monthly-to-weekly rhythm: v26.8.11 on 2026-08-29, v26.9.1 on 2026-09-03, v26.9.2 on 2026-09-09, with the last push on 2026-09-10. Homebrew makes the mechanics cheap, since brew upgrade on the cask or the formula pulls the new build, and the CLI and app share a binary so they move together. The cost is not the upgrade command, it is the surface area that changes. Release notes mention per-model settings and chat providers arriving in a single minor bump, which means configuration you set in one version may need revisiting in the next. If you pin a version for a production workflow, pin it deliberately and read the CHANGELOG.md before moving. The README does not document a rollback procedure for the app or the server, so plan your own.

## Conclusion

Adopt mlx-serve if you are on Apple Silicon with macOS 26.2 or newer and your clients already speak OpenAI, Anthropic or Ollama HTTP. Skip it if you need Linux, a non-Apple GPU, or an OS older than the stated floor. Verify first that the server starts and answers on http://localhost:11234 before you point Claude Code or an OpenAI SDK at it, and check the LICENSE and NOTICE files for the exact terms, since the repository metadata reports the licence as NOASSERTION even though the README badge says MIT.

## FAQ

### What is mlx-serve?

It is a native Zig inference server for Apple Silicon that runs MLX-format models and GGUF models, and exposes OpenAI-compatible, Anthropic-compatible and Ollama-compatible HTTP APIs on http://localhost:11234. It ships alongside MLX Core, a macOS menu-bar app that bundles the same server binary.

### What is MLX good for in mlx-serve?

In mlx-serve, MLX is the native Apple execution path that runs MLX-format weights, as opposed to the embedded llama.cpp path used for GGUF models. The README positions the MLX path as the reason the server can claim faster decode than LM Studio on identical MLX weights.

### What is the difference between GGUF and MLX in mlx-serve?

mlx-serve runs both. MLX-format models go through the native Apple path, while GGUF models go through an embedded llama.cpp, so the README describes the server as covering every GGUF on HuggingFace as well as MLX weights. The choice affects which engine executes the model, not which API your client uses.

### How do I use mlx-serve?

The README gives two routes. Install the MLX Core app and use its chat, agent mode and model browser, or install the CLI and run mlx-serve run gemma4 to download a model, serve it and chat in the terminal. Either way the server answers on http://localhost:11234.

## Sources

- [ddalcu/mlx-serve on GitHub](https://github.com/ddalcu/mlx-serve)
- [Issues](https://github.com/ddalcu/mlx-serve/issues)
- [Project website](http://mlxserve.com/)
- [README](https://github.com/ddalcu/mlx-serve/blob/main/README.md)
- [Releases](https://github.com/ddalcu/mlx-serve/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ddalcu-mlx-serve
