# SwiftLM: an MLX Swift inference server with an OpenAI-compatible API

> SwiftLM compiles MLX model serving into a single Swift binary for Apple Silicon, adds SSD streaming for oversized MoE models, and speaks the OpenAI chat completions API. The README's own benchmarks show where TurboQuant helps and where MTP stops paying off.

**SharpAI/SwiftLM** — ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app.

- Repository: https://github.com/SharpAI/SwiftLM
- Stars: 776 · Forks: 54
- Language: Swift
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/sharpai-swiftlm

## What SwiftLM solves, and for whom

Serving an MLX model usually means a Python process: mlx-lm or a similar stack, a runtime, and the memory behaviour that comes with it. SwiftLM's README frames the project as the alternative. It is a native Swift inference server that serves MLX models behind what it calls a strict OpenAI-compatible API, with no Python runtime and no GIL, compiled to a single binary.

The audience is narrow and specific. You need Apple Silicon, because the project targets Metal and MLX rather than CUDA. You probably want an OpenAI-shaped endpoint so existing client code keeps working, and you may be running a Mixture-of-Experts model large enough that loading every expert into unified memory is uncomfortable. The README's own example of that case is gemma-4-26b-a4b-it-4bit, a 26B MoE with roughly 4B active parameters per token at 4-bit precision.

The project also ships SwiftBuddy, an iOS app, alongside the server. That matters for scope: this is not only a headless daemon. The repository layout has a top-level SwiftBuddy/ directory and a build_swiftbuddy.sh script next to the server's build.sh, so the mobile target is a first-class part of the tree rather than a demo folder.

## How the Swift binary, Metal kernels and MLX submodules fit together

The architecture visible in the repository is a Swift package wrapping MLX. Package.swift and Package.resolved sit at the top level, Sources/ holds the Swift code, and two git submodules, mlx-swift and mlx-swift-lm, carry the MLX bindings and the language-model layer. A Packages/ directory holds additional Swift package dependencies.

The build is not a plain swift build. The README states that build.sh initialises submodules, installs cmake through Homebrew if it is missing, compiles mlx.metallib from the Metal kernel sources, and then builds the SwiftLM binary in release mode. That ordering explains why the release archive is described as self-contained: mlx.metallib is bundled next to the binary, so a downloaded tarball does not need the kernel compilation step.

Model weights are handled by MLX rather than by a separate downloader. The README says models download automatically if they are not cached, and both examples use mlx-community repository names, so the model identifier is an MLX community path such as mlx-community/Qwen2.5-3B-Instruct-4bit.

The interesting layer is the optional one. The --stream-experts flag is documented as the way to run oversized MoE models by streaming expert layers directly from NVMe SSD instead of relying on macOS virtual memory swapping. That is a deliberate trade: the README's benchmark table puts SSD streaming at roughly 10.8 tok/s at 512 context against 77.5 tok/s for the full-RAM MoE configuration, and it notes SSD Stream enables long-context inference on 24 GB Macs at around 22 to 27 GB of RAM. You are buying the ability to run the model at all, and paying for it in tokens per second.

## Installing SwiftLM and serving your first model

There are two paths. The fastest is the pre-built release tarball, which the README describes as self-contained because mlx.metallib is bundled alongside the binary. Download it from the Releases page, unpack it, and start the server against an MLX community model on port 5413:

```bash
tar -xzf SwiftLM-<version>-macos-arm64.tar.gz
./SwiftLM --model mlx-community/Qwen2.5-3B-Instruct-4bit --port 5413
```

If the model is not already cached, the README says it downloads automatically, so the first start will take longer than later ones.

The source path goes through build.sh, which the README says handles submodules, cmake, Metal kernel compilation and the release build:

```bash
git clone --recursive https://github.com/SharpAI/SwiftLM
cd SwiftLM
./build.sh
```

After that, the binary lives under .build/release/. The README's example starts it with a larger MoE model:

```bash
.build/release/SwiftLM \
  --model mlx-community/gemma-4-26b-a4b-it-4bit \
  --port 5413
```

For models that do not fit comfortably in RAM, the README says to add --stream-experts, which streams expert layers from NVMe SSD and bypasses macOS virtual memory swapping. Expect the throughput drop documented in the benchmark table before you commit to that flag for interactive use.

Because the API is described as OpenAI-compatible, the practical first test is to point an existing OpenAI client at http://localhost:5413. The README does not include a curl example for the chat completions route, so the exact request shape is something to confirm against the running server rather than assume.

## TurboQuant and MTP: what the benchmark tables actually say

The README publishes two benchmark tables for gemma-4-26b-a4b-it on an M5 Pro with 64 GB, one at 4-bit and one at 8-bit, and the two tell different stories about the same features.

TurboQuant, the KV cache compression path, is the consistent winner. At 4-bit and 100K context the README reports 66.9 tok/s for Vanilla + TurboQuant against 27.5 tok/s for vanilla, and OS RAM at 40K context dropping from 48.7 GB to 18.2 GB. At 8-bit and 100K it reports 48.3 tok/s against 14.9 tok/s vanilla. The mechanism is straightforward: compress the KV cache and long-context memory pressure falls, which removes the bandwidth bottleneck that dominates at depth.

MTP speculative decoding is conditional. The README states that at 4-bit the model is compute-bound on MoE expert dispatch, batch verification scales linearly with token count, and MTP provides no net throughput gain. At 8-bit the same table reports +20% at 40K and +51% at 100K, because the heavier weights make the model bandwidth-bound and the KV reads in the three-token verification batch amortise across all three queries. The README's own summary is that MTP alone is free at 4-bit in the sense of adding no memory, not in the sense of adding speed.

The combination is the part most likely to surprise. The README states plainly that TQ + MTP undercuts TQ alone at 4-bit, and the 8-bit table shows the same pattern: Vanilla + MTP + TurboQuant lands at 23.3 tok/s at 100K against 48.3 tok/s for TurboQuant alone. The explanation given is that once the KV cache is tiny, MTP adds verification overhead without removing a bottleneck that still exists. If you take one configuration decision from this project, take that one.

## Where SwiftLM is the wrong tool

The first boundary is hardware. This is Apple Silicon only. If your inference fleet is NVIDIA, nothing here applies, and the MLX dependency is not something you can swap out.

The second is the SSD streaming path, which is easy to misread as a general memory fix. The README's 4-bit table shows SSD + TurboQuant collapsing to 2.5 tok/s at 40K context and 1.6 tok/s at 100K, against 10.4 and 9.0 tok/s for SSD Stream alone. Streaming experts from NVMe is a way to fit a model that would otherwise not run on a 24 GB machine. It is not a way to make that model pleasant to use, and combining it with TurboQuant appears to make things worse rather than better.

The third is operational maturity. Releases are frequent and tagged with build numbers (b711, b710, b709 within about a week), and the README does not document a rollback procedure, a version compatibility matrix for model formats, or a migration path between builds. Nothing in the README suggests a stable long-term support line. If you need pinned, audited serving infrastructure with a documented downgrade path, the README gives you nothing to plan around.

Finally, the documentation is uneven. There is a reproducible profiling command for the benchmarks, but no curl example for the API, no description of which OpenAI endpoints beyond chat completions are implemented, and no concurrency or batching guidance. The word strict in the README's description of the API compatibility is a claim, not a specification.

## How it differs from mlx-lm and llama.cpp

The closest comparison is mlx-lm, the Python MLX serving stack. Both target Apple Silicon and both run MLX models. The difference is the runtime: mlx-lm runs in Python, and SwiftLM's stated reason for existing is to avoid that, along with the GIL and the memory copies the README associates with it. The practical consequence is deployment shape. SwiftLM ships as a single binary with the Metal library bundled, which is a different distribution problem from managing a Python environment, and it has no Python-side extension point. If you want to script model behaviour in Python, mlx-lm is the more natural home; if you want a self-contained process, that is SwiftLM's argument.

llama.cpp is the other reference point, and the difference is the model format and the kernel path. llama.cpp works from GGUF and supports CPU and multiple GPU backends across platforms. SwiftLM works from MLX community model repositories and targets Metal on Apple Silicon specifically. That means SwiftLM cannot serve a GGUF file you already have, and it cannot run on a Linux box with an NVIDIA card, but it also means the quantisation and kernel work is aimed at one architecture instead of many.

The features without a direct equivalent in either are worth naming. SSD streaming of expert layers for oversized MoE models, and the MTP speculative decoding path, are specific to this project's design. TurboQuant has conceptual cousins in other KV cache compression schemes, but the README's numbers are for this implementation and this hardware.

## Licence and the cost of tracking frequent builds

SwiftLM is MIT licensed, and the LICENSE file sits at the top level of the repository. MIT is permissive, so the usual obligations apply: keep the copyright notice and the licence text with redistributed copies. That covers SwiftLM's own code. It does not automatically cover the mlx-swift and mlx-swift-lm submodules, the Packages/ dependencies, or the model weights you download from mlx-community, each of which carries its own licence that you need to check separately. This is a description of the licence file, not legal advice.

The upgrade cost is the more practical concern. The release cadence visible in the repository is roughly daily to weekly, with build-numbered tags rather than semantic versions. The README does not describe a changelog, a deprecation policy, or how flags and model compatibility behave across builds. That means each upgrade is a fresh validation: rebuild or re-download, start the server on your model, and confirm the flags you depend on (--model, --port, --stream-experts) still behave as before. The README does point to a reproducible profiling script, scripts/profiling/profile_runner.py, which is the one concrete tool available for checking whether a new build changed throughput on your own hardware.

## Conclusion

SwiftLM is worth trying if you run Apple Silicon hardware, want an OpenAI-compatible endpoint without a Python runtime, and are willing to check model-specific behaviour yourself. It is the wrong pick if you need CUDA, multi-GPU serving, or a documented rollback path between the frequent b-tagged builds, because the README does not describe one. Before adopting it, run the server on your target Mac with your own model and context length, and read the benchmark table for the quantisation you actually plan to serve, since the README shows MTP helping at 8-bit and doing nothing at 4-bit.

## FAQ

### How do I install SwiftLM on a Mac?

Download the pre-built tarball from the Releases page and unpack it, since the README says the archive is self-contained with mlx.metallib bundled next to the binary. Alternatively, clone the repository with --recursive and run ./build.sh, which initialises submodules, installs cmake via Homebrew if needed, compiles mlx.metallib, and builds the release binary.

### Does SwiftLM need Python to run?

No. The README describes it as a native Swift inference server with no Python runtime and no GIL, compiled to a single binary. Models are downloaded automatically by MLX if they are not already cached.

### What does the --stream-experts flag do in SwiftLM?

The README says to add --stream-experts when running oversized MoE models, to bypass macOS virtual memory swapping and stream expert layers directly from NVMe SSD. The benchmark table shows this enables long-context inference on 24 GB Macs at around 22 to 27 GB of RAM, but at roughly 9 to 11 tok/s rather than the 27 to 77 tok/s range of the full-RAM configurations.

### Should I combine TurboQuant and MTP in SwiftLM?

The README's benchmark tables say no. At 4-bit it states that TQ + MTP undercuts TQ alone, and at 8-bit and 100K context the table shows 23.3 tok/s for the combined configuration against 48.3 tok/s for TurboQuant alone. TurboQuant is the configuration the README highlights as the headline result.

### Which models can SwiftLM serve?

MLX models, referenced by their mlx-community repository names. Both README examples use that namespace: mlx-community/Qwen2.5-3B-Instruct-4bit and mlx-community/gemma-4-26b-a4b-it-4bit. The README does not document serving GGUF files or non-MLX formats.

## Sources

- [Issues](https://github.com/SharpAI/SwiftLM/issues)
- [License: MIT](https://github.com/SharpAI/SwiftLM/blob/main/LICENSE)
- [README](https://github.com/SharpAI/SwiftLM/blob/main/README.md)
- [Releases](https://github.com/SharpAI/SwiftLM/releases)
- [SharpAI/SwiftLM on GitHub](https://github.com/SharpAI/SwiftLM)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sharpai-swiftlm
