# mlx-swift-lm: A Swift Package for Running LLMs and VLMs on Apple Silicon

> A Swift package from the ML-Explore team that lets developers add large language model and vision language model inference to Apple platform applications. It abstracts model loading, tokenization, and generation behind a common API, supports LoRA and full fine-tuning, and provides a FoundationModels bridge so MLX-loaded models can be used with Apple's standard LanguageModelSession interface.

**ml-explore/mlx-swift-lm** — LLMs and VLMs with MLX Swift

- Repository: https://github.com/ml-explore/mlx-swift-lm
- Stars: 822 · Forks: 412
- Language: Swift
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ml-explore-mlx-swift-lm

## What mlx-swift-lm Solves and Who Should Use It

Running a large language model on an Apple device requires solving three problems that do not exist when calling a cloud API: loading the model weights into memory efficiently, tokenizing input without a Python runtime, and generating output tokens within the memory and thermal constraints of a mobile or desktop chip. The mlx-swift-lm package handles all three. It provides a Swift-native implementation of LLM and VLM inference using the MLX framework, which is Apple's array computation library optimized for the unified memory architecture of Apple Silicon. The package is aimed at iOS, macOS, and visionOS developers who want to ship on-device AI features, machine learning researchers who prototype models in Swift, and application developers who want to use Hugging Face-hosted models without setting up a Python environment. Because the package is a Swift Package Manager dependency, integrating it into an existing Xcode project requires adding a dependency declaration, not installing additional runtimes.

## Package Structure: Six Libraries with Distinct Responsibilities

The package contains six libraries. MLXLLMCommon defines the shared protocol layer for model loading and generation that both LLM and VLM libraries implement. MLXLLM implements large language model architectures. MLXVLM implements vision language model architectures. MLXEmbedders contains encoder and embedding model implementations. MLXGuidedGeneration adds grammar-constrained generation: the library accepts a JSON Schema or EBNF grammar and constrains the model's output to tokens that conform to that grammar, applicable to any MLX model. MLXFoundationModels bridges any MLX model to Apple's FoundationModels framework, exposing it as a LanguageModelSession-compatible model. The FoundationModels integration requires the macOS/iOS/visionOS 27.0 SDK. The README links to MLX Swift Examples as a reference for applications and command-line tools built on these libraries.

## Installing and Adding mlx-swift-lm to a Swift Project

The package is distributed through Swift Package Manager. Add it to Package.swift with a version requirement:

```swift
.package(url: "https://github.com/ml-explore/mlx-swift-lm", .upToNextMajor(from: "3.31.3")),
```

Then add the specific library targets your application needs. The full dependency block for a project that uses the Hugging Face integration looks like this:

```swift
dependencies: [
    .package(url: "https://github.com/ml-explore/mlx-swift-lm", .upToNextMajor(from: "3.31.3")),
    .package(url: "https://github.com/huggingface/swift-huggingface", from: "0.9.0"),
    .package(url: "https://github.com/huggingface/swift-transformers", from: "1.3.0"),
],
```

The Hugging Face packages provide the tokenizer and model downloader integration. The swift-transformers package handles tokenization and the swift-huggingface package handles model file downloading from the Hugging Face Hub. With these in place, the MLXHuggingFace macro resolves a model configuration from the registry and downloads the weights automatically at runtime on first launch. Subsequent launches use the cached weights from disk. The package supports three integration styles for tokenizers and downloaders, trading convenience for control: a fully managed macro approach, a semi-managed approach with explicit downloader configuration, and a fully manual approach for local-only or custom weight sources. The documentation at swiftpackageindex.com covers all three in detail, including the use of custom downloaders for models not hosted on Hugging Face.

## Loading a Model and Running a Chat Session

With the MLXHuggingFace integration in place, loading a model and running a chat conversation requires a few lines of Swift:

```swift
import MLXLLM
import MLXLMCommon
import MLXHuggingFace
import HuggingFace
import Tokenizers

let model = try await #huggingFaceLoadModelContainer(
    configuration: LLMRegistry.gemma3_1B_qat_4bit
)

let session = ChatSession(model)
print(try await session.respond(to: "What are two things to see in San Francisco?"))
print(try await session.respond(to: "How about a great place to eat?"))
```

The ChatSession type maintains conversation history across calls automatically, so the second question receives the context of the first answer without any additional code. The #huggingFaceLoadModelContainer macro accepts a model configuration from the LLMRegistry, which includes quantized variants of widely-used open models such as Gemma. Quantized models use less memory than full-precision versions, which matters on devices with constrained unified memory budgets. The model files download from the Hugging Face Hub on first use and are cached locally; subsequent application launches load the model from disk without re-downloading. For applications that cannot access the internet or must use custom model weights, the manual integration approach described in the documentation replaces the macro with an explicit downloader and tokenizer configuration.

## FoundationModels Bridge and Grammar-Constrained Generation

MLXFoundationModels allows an MLX model to be used anywhere Apple's FoundationModels framework is used. The bridge builds an MLXLanguageModel, passes it to a LanguageModelSession, and generation goes through the standard FoundationModels API. This means an application that already uses LanguageModelSession for Apple's on-device model can substitute an MLX-loaded Hugging Face model with minimal code changes, subject to the macOS/iOS/visionOS 27.0 SDK requirement. MLXGuidedGeneration adds grammar-constrained output to this path: by combining the two libraries and requesting a @Generable type, the response is constrained to a JSON schema at token-generation time rather than via post-processing. The README example shows a Recommendation struct annotated with @Generable receiving structured location data from the model, with the constraint enforced by the grammar rather than by prompt engineering alone. Other capabilities available through the FoundationModels bridge include vision input, tool calling, and reasoning mode.

## Breaking Changes in 3.x and the Upgrade Path

The 3.x release introduced breaking changes relative to the 2.x line. The README notes that the main branch is a new major version and that breaking changes were introduced to decouple the package from specific tokenizer and downloader packages. Existing integrations built against 2.x will not compile against 3.x without changes. The upgrading documentation at swiftpackageindex.com/ml-explore/mlx-swift-lm/main/documentation/mlxlmcommon/upgrade covers the specific changes required. The most recent GitHub releases listed in the repository are 3.31.4 from 2026-06-30, 3.31.3 from 2026-04-15, and 2.31.3 from 2026-04-01. The package uses swift-format version 603.0.0 pinned in CI for consistent code formatting; a different version of swift-format may produce different output when formatting the package source.

## What mlx-swift-lm Does Not Do and Alternatives

The package does not provide a server or HTTP API layer; it runs inference in-process within the application. It does not include model training infrastructure beyond LoRA and full fine-tuning, and it does not provide model conversion utilities for formats other than those already supported. The package requires Apple Silicon; it does not run on Intel Macs or non-Apple hardware. The FoundationModels bridge requires macOS/iOS/visionOS 27.0 or later, which means applications targeting earlier OS versions cannot use the bridge and must use the MLXLLMCommon API directly. The closest alternative for on-device inference in Swift is llama.cpp, which provides C++ inference with Swift bindings. llama.cpp runs on a broader range of hardware including Intel Macs and Linux, but does not provide the MLX performance optimizations for Apple Silicon unified memory or the FoundationModels framework integration. The package is MIT-licensed with no restrictions on commercial use.

## Conclusion

mlx-swift-lm is the right package for Swift developers who want on-device LLM and VLM inference on Apple Silicon without a server dependency. The FoundationModels bridge makes MLX-loaded models a drop-in replacement for Apple's on-device model in applications that already use LanguageModelSession, but it requires the macOS/iOS/visionOS 27.0 SDK. For applications that target older OS versions, the MLXLLMCommon API is still available. The 3.x line introduced breaking changes from 2.x; check the upgrading documentation before updating an existing integration. The library is MIT-licensed and the last push was on 2026-09-22.

## FAQ

### How to install mlx lm in a Swift project?

Add the Swift Package Manager dependency to Package.swift using the GitHub URL https://github.com/ml-explore/mlx-swift-lm with .upToNextMajor(from: "3.31.3"). Then add the specific library targets you need, such as MLXLLM or MLXVLM, to your target's dependencies.

### Can MLX run on iPhone with mlx-swift-lm?

The package supports iOS. The FoundationModels bridge requires iOS 27.0 or later. The base MLXLLMCommon and MLXLLM APIs are available on earlier iOS versions that support the MLX Swift framework. Model size and device memory are the practical limits for on-device inference on iPhone.

### What breaking changes came with mlx-swift-lm 3.x?

The 3.x line decoupled the package from specific tokenizer and downloader packages, requiring code changes for any integration built against 2.x. The upgrading documentation at swiftpackageindex.com covers the specific changes. The main branch targets 3.x and is not backwards-compatible with 2.x.

## Sources

- [Issues](https://github.com/ml-explore/mlx-swift-lm/issues)
- [License: MIT](https://github.com/ml-explore/mlx-swift-lm/blob/main/LICENSE)
- [ml-explore/mlx-swift-lm on GitHub](https://github.com/ml-explore/mlx-swift-lm)
- [README](https://github.com/ml-explore/mlx-swift-lm/blob/main/README.md)
- [Releases](https://github.com/ml-explore/mlx-swift-lm/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ml-explore-mlx-swift-lm
