Model or dataset
jjang-ai/vmlx avatar
jjang-ai/vmlx

vMLX: a self-hosted MLX inference server for Apple Silicon

vMLX - JANGTQ Uber Compressed MLX Models - L2 Disk Cache (survives restart) + L1 Paged (super fast ttft) + Hybrid SSM Scheduler + Cont Batching + etc!

879 stars93 forksPythonApache-2.0

At a glance

What is it?
vMLX packages an OpenAI, Anthropic and Ollama compatible HTTP API around MLX models, with a two-tier prompt cache and hybrid SSM scheduling. It is macOS only, published as a beta, and the README's headline quality claim rests on JANG-quantized checkpoints rather than on stock MLX weights.
Who is it for?
Adopt vMLX if you already run MLX models on an Apple Silicon Mac and want an OpenAI-compatible endpoint with a persistent prompt cache, and you accept the macOS-only constraint and the Beta classifier in pyproject.toml. Do not adopt it if you need Linux, CUDA, or a stable API surface.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What vMLX replaces on a Mac

vLLM and TensorRT-LLM assume CUDA. On Apple Silicon the practical serving layer is MLX, and mlx-lm ships a generation loop rather than a server. vMLX fills that gap: it is a self-hosted inference server for LLMs, VLMs and image generation on Apple Silicon, exposing an OpenAI, Anthropic and Ollama compatible HTTP API. The README states that no third-party API keys are required, which is the whole point of the project. It is aimed at developers who want a local endpoint their existing SDK code can point at by changing a base URL.

The scope is wider than text. The model support table lists text LLMs (Qwen, Llama, Mistral/Mixtral, Gemma, Phi-4, DeepSeek, GLM, MiniMax, Nemotron, Kimi), vision LLMs (Qwen-VL, Pixtral, InternVL, LLaVA), MoE models, hybrid SSM architectures such as Nemotron-H, Jamba and GatedDeltaNet, image generation via mflux (Flux Schnell/Dev, Z-Image Turbo), image editing via Qwen Image Edit, embeddings, reranking, Kokoro TTS and Whisper STT. That breadth is the project's main differentiator against a plain mlx-lm script, and also its main risk surface, since each family can carry its own cache and scheduling rules.

Two cache tiers, one scheduler, and why the split matters

The feature table describes a layered caching design. L1 is a paged KV cache: block-based caching with content-addressable deduplication, so identical prompt prefixes map to the same blocks instead of being stored twice. L2 is a disk cache that persists prompt caches to SSD and, per the feature description, survives server restarts. A separate block disk cache pairs per-block persistence with the paged KV cache. KV cache quantization to q4 or q8 is listed as a further 2-4x memory saving on top.

The practical consequence is that time-to-first-token separates into two cases. A cold prompt pays full prefill. A prompt whose prefix is already in L1 or in the on-disk L2 skips most of that work. The restart survival of L2 is the interesting part: it means the cache is worth something across process lifetimes, not just within one session, which is unusual for a local server.

Scheduling is handled by continuous batching for concurrent requests, with a Hybrid SSM Scheduler that the README calls out for Mamba and GatedDeltaNet layers handled alongside attention. That is a real architectural problem rather than a marketing line: recurrent state does not behave like a KV block, so batching and paging have to treat those layers differently. The README also mentions native MTP artifact detection and family-specific cache policy gates, described as keeping speculative and cache settings explicit and model-safe. Read that as an admission that these settings are not universally safe to enable, and that the server will refuse some combinations rather than guess.

Installing vMLX from PyPI and serving a first model

The README publishes the package on PyPI as vmlx and gives three install routes. The recommended one uses uv. Running these commands installs the tool and starts a server on port 8000 with an OpenAI and Anthropic compatible API.

bash
brew install uv
uv tool install vmlx
vmlx serve mlx-community/Qwen3-8B-4bit

The README notes that on macOS 14 and later a bare pip install fails with "externally-managed-environment", so pipx or a virtual environment are the alternatives it offers. Once the server is up, existing client code works by pointing at the local base URL. This is the OpenAI SDK example the README gives, with a placeholder API key because the server does not require one.

python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
response = client.chat.completions.create(
    model="local",
    messages=[{"role": "user", "content": "Hello!"}],
    stream=True,
)
for chunk in response:
    print(chunk.choices[0].delta.content or "", end="", flush=True)

Anthropic's SDK works the same way against the same base URL, and the README also gives a curl form posting to /v1/chat/completions with streaming enabled. The model field is "local" in these examples regardless of which checkpoint was loaded. For a first real use, pick a model from mlx-community so you are testing the server rather than a quantization scheme; the README points at that organization as the source of thousands of ready models.

The JANG quantization claim is not a claim about stock MLX

The README opens with a table comparing JANG_2L at 2 bits against MLX 4-bit, 3-bit and 2-bit on MiniMax M2.5, reporting 74% MMLU for JANG_2L at 89 GB against 26.5% for MLX 4-bit at 120 GB. Those numbers are attributed to jangq.ai with models hosted under the JANGQ-AI organization on Hugging Face.

Read carefully, that is a claim about a specific quantization recipe, not about vMLX as a server. The mechanism described is adaptive mixed-precision, keeping critical layers at higher precision while the rest go to 2 bits. If you serve a stock mlx-community checkpoint, this table tells you nothing about what you will get. The comparison is also a single model family and a single benchmark, and the README does not describe the evaluation harness. Treat the table as a pointer to the JANG model collection, not as a reason to expect unusual accuracy from the server itself.

Where vMLX is the wrong tool

The platform constraint is absolute. pyproject.toml declares the operating system classifier as MacOS and requires Python 3.11 through 3.14. The dependency list pins mlx>=0.32.2 on Darwin and mlx>=0.29.0 elsewhere, with a comment explaining that the newer floor exists because MLX 0.32.2 adds M5/G17 small-row QMM and quantized-MoE kernels plus fused head-dim-256 prefill attention, and that those Metal-only optimizations cannot run on non-Darwin platforms. If your deployment target is Linux or CUDA, vMLX is not a candidate; the dependency pins say so directly.

The second constraint is maturity. pyproject.toml carries the classifier Development Status :: 4 - Beta, and the JIT compilation feature is marked experimental in the feature table. A beta server that manages persistent disk caches and per-family cache policy gates has more state than a plain generation loop, and the README does not document rollback, cache invalidation, or what happens when a cached block no longer matches the loaded model. If you need a frozen API surface with a support contract, this is the wrong layer to build on.

The third constraint is that the headline speed features are conditional. Speculative decoding needs a draft model and is described as a 20-90% speedup, a range wide enough to be meaningless without knowing the workload. Prompt lookup decoding needs no draft model and is described as best for structured or repetitive output such as code, JSON or schemas, enabled with --enable-pld. Neither helps a workload of short, unique, non-repetitive prompts.

Alternatives and the actual difference

The nearest alternative is mlx-lm itself. It provides the model loading and generation primitives, and vMLX depends on it, pinning mlx-lm>=0.31.3 as the release-bundle floor for current JANG families. The difference is what sits on top: mlx-lm does not ship an HTTP server with OpenAI, Anthropic and Ollama compatible routes, continuous batching, a paged KV cache, or a disk cache that survives restarts. If you only need to run one prompt at a time in a Python process, mlx-lm is the smaller dependency and vMLX adds a server you will not use.

The other alternative is a native macOS application rather than a server. The README points readers looking for a native Swift macOS app or Swift inference engine to osaurus.ai, and separately offers MLX Studio, a macOS app from the same author with a chat UI, model management, image generation and developer tools, distributed as a DMG. The difference is integration versus interface: MLX Studio is for someone who wants a window and no terminal, while vMLX is for someone who wants a base URL their own code can call. Choosing vMLX means you own the process, the port and the cache directory.

A third option worth naming is simply running the model remotely behind a hosted API. vMLX's stated value is that it is self-hosted with no third-party API keys, which matters when prompts cannot leave the machine. If that constraint does not apply to you, the operational cost of a local server is hard to justify.

Licence, maintenance and upgrade cost

vMLX is Apache-2.0, stated in both the LICENSE file and the license field in pyproject.toml. That is a permissive licence with an explicit patent grant, and it does not impose copyleft obligations on your own code. It says nothing about the models you load: checkpoints from mlx-community, JANGQ-AI and elsewhere carry their own licences, and the Apache-2.0 grant on the server does not extend to them. That is a factual boundary, not legal advice.

The repository was last pushed on 2026-08-29 and is not archived. The release history shows v1.6.45 on 2026-08-29, v1.6.44 on 2026-08-28 and v1.6.43 on 2026-08-27, so three releases landed within three days at the end of that window. pyproject.toml declares version 1.6.58, which is ahead of the newest release listed, so the packaging metadata and the tagged releases are not in lockstep. Pin the version you install rather than tracking the latest tag.

Upgrade cost concentrates in the dependency floors. The mlx-lm pin at 0.31.3 is annotated as the release-bundle floor for current JANG families, citing ArraysCache.lengths/advance, SequenceStateMachine, LRUPromptCache and a native Gemma 4 text MoE, and the comment notes that vMLX patches the baseline model registry with bailing_hybrid for Ling. That means vMLX reaches into mlx-lm internals. An mlx-lm release that renames or restructures those symbols can break the server in ways the version constraint will not catch, because the constraint is a floor, not a ceiling.

Editorial conclusion

Adopt vMLX if you already run MLX models on an Apple Silicon Mac and want an OpenAI-compatible endpoint with a persistent prompt cache, and you accept the macOS-only constraint and the Beta classifier in pyproject.toml. Do not adopt it if you need Linux, CUDA, or a stable API surface. Before committing, verify two things yourself: whether the model family you intend to serve is listed as gated for speculative or cache settings, and whether your chosen checkpoint is a JANG bundle or a stock mlx-community one, because the README's quantization comparison is drawn from JANG models.

Frequently asked questions

Can I run vMLX on Linux or Windows?

No. pyproject.toml declares the operating system classifier as MacOS, and the dependency comment states that the MLX 0.32.2 Metal-only optimizations cannot run on non-Darwin platforms. The non-Darwin mlx floor exists only so the package resolves elsewhere, not because vMLX is supported there.

How do I install vMLX and start the server?

The README recommends uv: brew install uv, then uv tool install vmlx, then vmlx serve mlx-community/Qwen3-8B-4bit. pipx and a Python virtual environment are given as alternatives, and the README notes that a bare pip install fails on macOS 14 and later with "externally-managed-environment".

Which models can vMLX serve?

The README says it runs any MLX model and you point it at a HuggingFace repo or a local path. The support table covers text LLMs, vision LLMs, MoE and hybrid SSM architectures, image generation and editing via mflux, embeddings, reranking, Kokoro TTS and Whisper STT.

Does the vMLX prompt cache survive a server restart?

The feature table describes the L2 disk cache as persisting prompt caches to SSD and surviving server restarts, with a separate block disk cache paired with the paged KV cache. The README does not document cache invalidation or what happens when a persisted block no longer matches the loaded model.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/jjang-ai-vmlx.svg)](https://hysenlabs.com/projects/jjang-ai-vmlx)