Model or dataset
jundot/omlx avatar
jundot/omlx

oMLX: an Apple Silicon LLM server with continuous batching and tiered KV caching

LLM inference server with continuous batching & SSD caching for Apple Silicon, managed from the macOS menu bar.

22,367 stars1,944 forksPythonApache-2.0

At a glance

What is it?
oMLX is a Python LLM inference server for Apple Silicon Macs that keeps KV cache in a hot memory tier and a cold SSD tier, exposes an OpenAI-compatible endpoint on port 8000, and is managed from the macOS menu bar. It is honest about its own limits: custom kernels need full Xcode, and some model families fall back to much slower generic paths.
Who is it for?
Adopt oMLX if you run local models on an M-series Mac and want an OpenAI-compatible endpoint on port 8000 with a menu bar and an admin dashboard, and if you are willing to install full Xcode or use the DMG when you need the custom kernels. Do not adopt it if you are on Intel, Linux, or an older macOS, or if you need a stable API surface: pyproject.toml marks it Development Status 3, Alpha.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The choice oMLX is built to avoid

The README opens with the author's own framing: every LLM server he tried made him choose between convenience and control. The stated goal is to pin everyday models in memory, auto-swap heavier ones on demand, set context limits, and manage all of it from a menu bar. That is a narrow but real problem. On a Mac, a local inference server is usually one of two things: a Python process you launch from a shell and forget, or a desktop app that hides everything. oMLX tries to be both, with a Homebrew service, a DMG app, and a CLI shim at ~/.omlx/bin/omlx that lets terminal commands and Apple Shortcuts drive the app-managed server.

The target user is a developer running local models on Apple Silicon for real coding work, the README names Claude Code specifically. It is not aimed at teams serving many users on Linux GPUs. The dependencies make that explicit: mlx==0.32.2, macOS 15.0 (Sequoia) or later, Python 3.11 to 3.13, and Apple Silicon M1 through M5. Intel Macs are out. So is anything that is not a Mac.

Continuous batching, tiered KV cache, and what persists

Two mechanisms carry the design. The first is continuous batching, which the README describes as shared with the VLM path: vision-language models run on the same batching and cache stack as text LLMs, not a separate code path. The second is the tiered KV cache. The README states that oMLX persists KV cache across a hot in-memory tier and a cold SSD tier, and that even when context changes mid-conversation, all past context stays cached and reusable across requests. That last clause is the interesting one. A prefix cache normally invalidates when the prefix changes; the claim here is that past context survives a context change and stays reusable.

Model discovery is directory-based. The server scans subdirectories under the model directory and picks up LLMs, VLMs, embeddings, and rerankers automatically. That means the model directory is the configuration surface: what you put in it is what the server serves. There is an admin dashboard at /admin for monitoring, model management, chat, benchmark, and per-model settings, and a chat UI at /admin/chat. The dashboard ships in English, Korean, Japanese, Chinese, French, Russian, Spanish, and Brazilian Portuguese, and the README notes that all CDN dependencies are vendored so the dashboard works offline. For an air-gapped Mac that detail matters more than the language list.

Installing oMLX on macOS and serving a first model

There are three install paths. The DMG is the shortest: download it from Releases, drag to Applications, and the app handles its own updates. It also ships the custom kernels precompiled, which avoids the build problem described below. Homebrew is the path for a background service.

bash
brew tap jundot/omlx https://github.com/jundot/omlx
brew install jundot/omlx/omlx

# Upgrade to the latest version
brew update && brew upgrade omlx

# Run as a background service (auto-restarts on crash)
omlx start

The service runs omlx serve with zero-config defaults: ~/.omlx/models and port 8000. omlx start, omlx stop, and omlx restart are described as the portable lifecycle commands; on a Homebrew install they delegate to brew services. To change the model directory or port, either set environment variables such as OMLX_MODEL_DIR and OMLX_PORT, or run the serve command once with a flag so the setting is written to ~/.omlx/settings.json.

bash
# Foreground server attached to this terminal
omlx serve --model-dir ~/models

With the server up, any OpenAI-compatible client points at http://localhost:8000/v1. If you installed from source instead, the README gives pip install -e . for the core and pip install -e ".[mcp]" when you want Model Context Protocol support. Logs land in two places: the service log under $(brew --prefix)/var/log/omlx.log and the structured server log at ~/.omlx/logs/server.log. When something fails to start, the second file is the one to read.

The custom kernel trap, and how to check whether you fell into it

This is the sharpest limitation in the documentation, and the README states it plainly. A plain pip install -e . does not build the native custom kernels. The affected model families then silently fall back to much slower generic paths. The README's own figure for GLM-5.2 is that the fused DSA prefill is roughly 30x faster with the kernels, measured at 845 versus about 29 tok/s on an M3 Ultra, and the fallback also uses more memory. Silent is the operative word: nothing in the default install tells you which path you are on.

Building the kernels requires the Metal toolchain, which Command Line Tools alone do not provide. The README quotes the failure: xcrun: error: unable to find utility "metal". The fixes are full Xcode, the official DMG which ships the kernels precompiled, or a Homebrew HEAD build that also needs full Xcode.

bash
brew install jundot/omlx/omlx --HEAD --with-custom-kernel

After any source install, verify rather than assume:

bash
python -c "from omlx.custom_kernels import native_kernel_status; print(native_kernel_status())"

The ABI coupling explains why this is fragile. pyproject.toml pins nanobind==2.15.0 specifically to match the ABI version MLX 0.32.2 was built with, and comments that a mismatched nanobind isolates the mlx NB_DOMAIN so that custom kernel extensions reject every mlx.core.array at the type caster. The bundled kernel binaries are ABI-coupled to that exact mlx version, so bumping the pin requires rebuilding them. If you install oMLX from source and later upgrade MLX independently, expect this to break.

Multi-Mac inference is experimental, and the README says so

Source builds can split one downloaded language model across Macs with unequal memory using MLX pipeline ranks over Ring or Thunderbolt RDMA/JACCL. The Cluster dashboard handles read-only peer discovery, strict SSH and runtime verification, byte-aware unequal shard planning, measured compute and link rebalancing, headroom-aware execution tuning, activation, and a live shard and performance map on both machines. Three profiles (interactive, balanced, throughput) expose coalesced batching, prompt-cache affinity, rotating-KV limits, Ring connection tuning, and a capability-gated experimental token-only output path.

The word experimental is the project's own. The README points to docs/distributed-cluster.md for setup, security boundaries, current limitations, and a physical-hardware validation checklist. Note what that implies: the feature is validated against physical hardware by the user, not by a test suite you can inspect. If you need two Macs to behave like one server today, treat this as something to evaluate on your own hardware rather than something to deploy.

Where oMLX is the wrong tool

The platform constraint is absolute. macOS 15.0+, Apple Silicon, Python 3.11 to 3.13. A Linux box with an NVIDIA card cannot run this at all, and neither can an Intel Mac. If your serving target is a shared cluster, oMLX is not in the conversation.

The second case is stability. pyproject.toml classifies the project as Development Status 3, Alpha, and the release history shows a run of release candidates (v0.6.3rc2, v0.6.3rc3) before v0.6.3. An alpha project with ABI-pinned native extensions is not a good fit for a service you cannot easily rebuild. The nanobind and mlx pins are exact, not ranges, which means a dependency resolver conflict with another package that wants a different MLX version has no clean resolution.

The third case is anyone who wants the model directory to be managed elsewhere. Discovery is by scanning subdirectories, so a layout that scatters weights across several roots, or pulls them at request time, does not map onto how oMLX finds models.

How it differs from llama.cpp and Ollama

The closest comparison is llama.cpp's server, and the difference is the cache design rather than the model formats. llama.cpp server offers prompt caching and, in recent versions, slot-based prompt reuse, but the cache lives in process memory and dies with the process. oMLX's stated design is a two-tier cache where the cold tier is SSD, so cached context outlives the process and is reused across requests. That is the claim to test on your own workload, because it is the one thing the README asserts that a generic OpenAI-compatible server does not do.

Ollama is the other reference point, and the split is control. Ollama optimizes for pulling a model and running it with one command, and it hides the serving parameters. oMLX exposes per-model settings in the admin dashboard, context limits, and lifecycle commands, and it assumes you already have weights in a directory. It also assumes a Mac. The trade is convenience for the ability to pin some models in memory and swap others on demand, which is exactly the choice the author says he was trying to stop making.

Licence, upgrade cost, and what to verify first

The licence is Apache-2.0, declared both in the repository and in pyproject.toml's license field, and the classifier is OSI Approved :: Apache Software License. Apache-2.0 is permissive and includes a patent grant; if you redistribute oMLX or a modified build, the usual obligations around notices and attribution apply. That is a description of the licence text, not legal advice, and the bundled kernel binaries and the MLX dependency carry their own terms that you should read separately.

Upgrade cost splits by install path. The DMG app has in-app auto-update, so the cost is a click plus a restart. Homebrew is brew update && brew upgrade omlx. Source installs are the expensive case: any MLX version bump invalidates the prebuilt kernels, and rebuilding them needs full Xcode and OMLX_WITH_CUSTOM_KERNEL=1. The macOS app also installs the ~/.omlx/bin/omlx shim, so a terminal workflow and a menu bar app can coexist over the same server.

The maintenance signal is the last push on 2026-08-27, the same day as v0.6.3. Before adopting, verify three things: the output of native_kernel_status() on your machine, whether your model family is one of the ones that silently falls back, and whether the settings you need are persisted in ~/.omlx/settings.json after a single omlx serve --model-dir run.

Editorial conclusion

Adopt oMLX if you run local models on an M-series Mac and want an OpenAI-compatible endpoint on port 8000 with a menu bar and an admin dashboard, and if you are willing to install full Xcode or use the DMG when you need the custom kernels. Do not adopt it if you are on Intel, Linux, or an older macOS, or if you need a stable API surface: pyproject.toml marks it Development Status 3, Alpha. Before committing, run the native kernel status check, confirm your model family is not silently on the generic path, and read docs/distributed-cluster.md if you plan to split a model across two Macs.

Frequently asked questions

How does oMLX work?

It runs an OpenAI-compatible inference server on Apple Silicon using MLX, with continuous batching and a KV cache split across a hot in-memory tier and a cold SSD tier. Model directories are scanned automatically for LLMs, VLMs, embeddings, and rerankers, and the server is managed from the macOS menu bar or from the omlx CLI.

How to set up oMLX?

Download the DMG from Releases and drag it to Applications, or run brew tap jundot/omlx https://github.com/jundot/omlx followed by brew install jundot/omlx/omlx, then omlx start. The service defaults to ~/.omlx/models and port 8000, and the README also documents a source install with pip install -e .

How to get an oMLX API key?

The README does not document an API key mechanism. It states that any OpenAI-compatible client can connect to http://localhost:8000/v1, and the server runs locally on the same Mac, so there is no key issuance step described.

What is omlx?

oMLX is an LLM inference server optimized for Apple Silicon, written in Python and licensed Apache-2.0. The README describes it as continuous batching and tiered KV caching managed directly from the macOS menu bar.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jundot-omlx.svg)](https://hysenlabs.com/projects/jundot-omlx)