MLX-VLM: running vision language models locally on Apple Silicon
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
At a glance
- What is it?
- MLX-VLM is a Python package for inference and fine-tuning of vision language models on a Mac using MLX. It ships a CLI, a FastAPI server and a Gradio chat UI, and its supported model list is unusually broad.
- Who is it for?
- Adopt MLX-VLM if you are on an Apple Silicon Mac and want one Python package that covers CLI inference, a FastAPI server and fine-tuning across many vision architectures. Skip it if your deployment target is a CUDA host or a Linux server, because the project's own CUDA and CPU extras are marked beta in the package classifier and the primary path is MLX on Apple Silicon.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What MLX-VLM solves, and who it is actually for
Running a vision language model normally means a GPU host, a container image and a serving stack. MLX-VLM removes that requirement for one class of hardware: Apple Silicon Macs. It is a Python package that performs inference and fine-tuning of VLMs, and of Omni Models that add audio and video, using MLX as the array framework. The pyproject.toml declares requires-python >= 3.10 and a Beta development status classifier, so the intended audience is developers rather than production platform teams.
The practical target is someone with an M-series machine who wants to feed an image to a model and get text back without uploading the image anywhere. The README's own example uses mlx-community/Qwen2-VL-2B-Instruct-4bit, a small quantized checkpoint, which sets the tone: this is local, modest-hardware work, not a datacenter replacement. If your images cannot leave the laptop, that constraint alone justifies looking at it.
How the package is put together
The repository layout tells most of the story. mlx_vlm/ holds the library and per-model implementations; each architecture gets its own directory under mlx_vlm/models, and several of them carry a dedicated README covering prompt formats and best practices. The README's model-specific documentation table lists entries such as DeepSeek-OCR, PaddleOCR-VL, Phi-4 Multimodal, MiniCPM-o, Moondream2 and Moondream3, LLaVA-OneVision, Gemma 4 and Granite Vision 3.2. That table is the real index of what the project supports; the top-level README only summarizes.
Six console entry points are declared in pyproject.toml: mlx_vlm.generate, mlx_vlm.convert, mlx_vlm.server, mlx_vlm.chat, mlx_vlm.chat_ui and mlx_vlm.moe_offload. So the same installed package gives you a one-shot CLI, a weight converter, an HTTP server and two chat front ends. Dependencies in requirements.txt include mlx>=0.32.2, transformers>=5.14.0, fastapi, starlette, uvicorn, websockets, opencv-python, llguidance and mlx-audio. That dependency set explains the feature surface: FastAPI and websockets for serving, llguidance for structured outputs, mlx-audio for the Omni models.
Optional extras are declared separately: ui pulls gradio>=5.19.0, train pulls datasets>=2.19.1, realtime pulls sounddevice>=0.5.0, and there are cuda and cpu extras pointing at mlx-cuda and mlx-cpu. The base install does not include Gradio, which is why the README gives a second install command for the chat UI.
Installing MLX-VLM and generating your first caption
The README states the easiest path is pip. One command installs the package and pulls the MLX, transformers and FastAPI dependencies from requirements.txt.
pip install -U mlx-vlmIf you want the Gradio chat interface, install the ui extra. The README warns to quote the package name so shells that expand square brackets, such as zsh, do not treat [ui] as a glob pattern.
pip install -U 'mlx-vlm[ui]'For a first real use, the README shows the generate entry point with a small quantized checkpoint. Text generation takes a prompt; the image path uses the same entry point with an image flag, which the truncated README example begins to show but does not complete.
mlx_vlm.generate --model mlx-community/Qwen2-VL-2B-Instruct-4bit --max-tokens 100 --prompt "Hello, how are you?"The first run downloads the checkpoint from Hugging Face, so expect a wait and a cache directory that grows. After that, output appears in the terminal. To serve instead of generating once, pyproject.toml declares mlx_vlm.server, and the README documents continuous batching, automatic prefix caching and KV cache quantization as server-side options. Those three are worth reading before you size a deployment, because they change memory behaviour rather than output quality.
Where MLX-VLM runs into trouble
The first limitation is hardware. Everything in the README's usage section assumes MLX, which is Apple's framework. The cuda and cpu extras exist, and the README even has an Activation Quantization (CUDA) section, but the package classifier still reads Development Status :: 4 - Beta, and the primary documented workflow is a Mac. Treat non-Apple paths as secondary until you have verified them yourself.
The second is model coverage versus model quality. The model-specific documentation table is long, and length is not the same as uniformity. Each architecture has its own prompt format and its own caveats, which is exactly why those per-model READMEs exist. A generic chat template that works for one checkpoint can silently produce worse output on another. When results look wrong, check the model's own README before blaming the runtime.
The third is the README itself. It is dense and partly truncated in places: the CLI image-generation example is cut off mid-flag, and features such as speculative decoding, DFlash, Gemma 4 MTP, EAGLE-3, TurboQuant KV cache and distributed inference are listed as headings with detail that this overview cannot verify. Do not plan around a feature because it has a heading. Open the section and read it, or open the corresponding file under mlx_vlm/.
MLX-VLM versus llama.cpp and Ollama
The comparison people search for most is against llama.cpp and Ollama, and the difference is architectural rather than a matter of speed. llama.cpp is a C/C++ inference engine with its own GGUF weight format and broad hardware support; Ollama wraps that engine in a model manager and a local API. Both are designed to run on many platforms, including Linux and CUDA machines.
MLX-VLM is a Python package built on MLX, which targets Apple Silicon's unified memory. It does not use GGUF. Its conversion path is mlx_vlm.convert, which the README's agent-skills table describes as converting and quantizing Hugging Face models to MLX, with bits and group size, quant modes, RTN and AWQ, and mixed recipes. That is a different pipeline from GGUF conversion, and it means weights prepared for llama.cpp will not load here.
The trade-off is straightforward. Choose llama.cpp or Ollama when you need one runtime across Linux, Windows and macOS, or when a GGUF build of your model already exists. Choose MLX-VLM when you are on a Mac, you want the model in its Hugging Face form, and you want Python-level access to the sampling, the cache and the fine-tuning loop rather than an opaque server binary.
Maintenance, licence and upgrade cost
The repository is not archived, and the last push was on 2026-09-09. Releases are frequent: v0.7.0 on 2026-09-07, v0.7.0rc0 on 2026-08-31, v0.6.17 on 2026-08-26. That cadence is the upgrade cost. A package that adds model architectures this quickly will move its dependency floor with it, and requirements.txt already pins mlx>=0.32.2 and transformers>=5.14.0. If you pin mlx-vlm in a lockfile, expect to revisit that pin when you want a newly added architecture.
The licence is MIT, declared both in the repository LICENSE file and in pyproject.toml as license = {text = "MIT"}. MIT is permissive, so the package itself imposes few obligations. The models are a separate question: checkpoints published on Hugging Face carry their own licences, and the README does not claim to normalize them. A permissive runtime licence does not grant you rights to the weights you download with it. That is a factual boundary, not legal advice; read the model card before commercial use.
The repository also ships an agent-skills bundle under skills/, with entries for cli-inference, server-inference, convert-quantize, add-new-model, benchmarking, contributing, hf-cache-models and reproducible-github-issues. The README gives an install path for Claude Code, Codex CLI and Gemini CLI, plus a validation script at skills/scripts/validate_skills.py. If your team uses coding agents, that bundle is the cheapest way to keep them from inventing flags.
Editorial conclusion
Adopt MLX-VLM if you are on an Apple Silicon Mac and want one Python package that covers CLI inference, a FastAPI server and fine-tuning across many vision architectures. Skip it if your deployment target is a CUDA host or a Linux server, because the project's own CUDA and CPU extras are marked beta in the package classifier and the primary path is MLX on Apple Silicon. Before committing, run the pip install, generate one image caption with mlx_vlm.generate, and confirm the specific checkpoint you need appears in the repository's model-specific documentation table; that table, not the README prose, is where per-model prompt formats and caveats live.
Frequently asked questions
What is the difference between mlx-vlm and mlx-lm?
MLX-VLM handles vision language models and Omni models with audio and video support, while mlx-lm is the text-only counterpart in the same MLX family. MLX-VLM's own description scopes it to inference and fine-tuning of VLMs on a Mac. If your input is only text, mlx-lm is the narrower tool.
How do I install MLX-VLM?
The README states the easiest way is pip install -U mlx-vlm. The Gradio chat UI needs a separate extra, pip install -U 'mlx-vlm[ui]', and the README advises quoting the package name so zsh does not expand the brackets as a glob.
How do I use MLX-VLM?
The documented entry points are the CLI, a Gradio chat UI, a Python script interface and a FastAPI server. The README's first example runs mlx_vlm.generate with a model, a max-tokens value and a prompt. pyproject.toml also declares mlx_vlm.convert, mlx_vlm.chat, mlx_vlm.chat_ui, mlx_vlm.server and mlx_vlm.moe_offload as console scripts.
What is MLX-VLM?
It is a Python package for inference and fine-tuning of Vision Language Models and Omni Models on a Mac using MLX. It is distributed on PyPI as mlx-vlm under the MIT licence and requires Python 3.10 or later.
How does MLX-VLM compare with Ollama?
Ollama wraps the llama.cpp engine and a model manager, and targets multiple platforms. MLX-VLM is built on MLX, which is Apple's framework, and converts Hugging Face models to MLX with mlx_vlm.convert rather than using GGUF weights. The README's usage and server sections assume a Mac.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/blaizzy-mlx-vlm)