# Swama: MLX-based LLM inference on macOS with a native Swift stack

> Swama is a pure Swift runtime for local LLM and VLM inference on Apple Silicon, built on MLX and shipped as a CLI plus a menu bar app. The interesting part is the OpenAI-compatible server; the constraint is that everything depends on macOS 15 and Apple Silicon.

**Trans-N-ai/swama** — High-performance MLX-based LLM inference engine for macOS with native Swift implementation

- Repository: https://github.com/Trans-N-ai/swama
- Stars: 591 · Forks: 31
- Language: Swift
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/trans-n-ai-swama

## What Swama solves, and who it is actually for

Running a local model on a Mac usually means picking between a Python stack that drags in its own environment and a GUI app that hides the serving layer. Swama takes a third position: a Swift runtime built directly on Apple's MLX framework, packaged as three artifacts. SwamaKit is the core framework library holding the business logic. Swama CLI is the command line tool for model management and inference. Swama.app is the macOS menu bar application with a graphical interface and background services.

The audience is narrow and specific. The README lists macOS 15.0 or later, Apple Silicon (M1/M2/M3/M4), Xcode 16.0+ for compilation, and Swift 6.2+. That is a developer or power user on a recent Mac who wants an OpenAI-shaped HTTP surface in front of local weights without running a Python server. If your deployment target is a Linux box with an NVIDIA card, nothing here applies to you.

## The MLX and Swift architecture, and how a request moves

Swama is not a wrapper around llama.cpp or a Python process. It is a pure Swift runtime on top of MLX, Apple's array framework, which is why the system requirements pin Swift 6.2 and macOS 15. The README describes the split into SwamaKit, the CLI, and the app, and the repository layout reflects it: swama/ holds the Swift package, swama-macos/ holds the Xcode project for the app, with Tests/ and Tools/ alongside.

The serving surface is the part worth paying attention to. Swama exposes OpenAI-compatible endpoints: /v1/chat/completions, /v1/embeddings, /v1/audio/transcriptions, and /v1/audio/speech, with tool calling supported on the chat endpoint. The README marks the speech endpoint as experimental. That means an existing client that already talks to the OpenAI API can be pointed at a local Swama instance, and the request flow becomes: client sends an OpenAI-shaped JSON body, Swama maps it onto an MLX model in memory, and streaming responses come back as server-sent events. Embeddings and transcription are served from the same process, so semantic search or RAG work does not need a second service.

Model acquisition is handled inside the runtime rather than by the user. The README states that models are downloaded automatically on first use and cached for future use, with HuggingFace Hub as the source. There is no separate pull step required before a run, though a list command exists for inspecting what is already cached.

## Installing Swama and running a first model

Homebrew is the recommended path. The README gives a single formula name, and after it completes the swama binary is on your PATH.

```bash
brew install swama
```

If you prefer the graphical route, the README points to the Releases page for a Swama.dmg, which you mount, drag Swama.app into Applications, and launch. It notes that macOS may show a security warning on first launch, with the workaround being the Open Anyway button under System Settings, or right-clicking the app and choosing Open. The CLI is installed from inside the app: open Swama from the menu bar and click "Install Command Line Tool..." to add the swama command to your PATH.

The quickest real use is a single run command with a model alias. Aliases are short names that map to full HuggingFace model identifiers, and the README says the model is downloaded automatically if it is not already cached.

```bash
swama run qwen3 "Hello, AI"
swama run llama3.2 "Tell me a joke"
swama list
```

The first command resolves qwen3 to mlx-community/Qwen3-8B-4bit, a 4.3 GB download, and prints the model's reply. The second uses llama3.2, which maps to mlx-community/Llama-3.2-3B-Instruct-4bit at 1.7 GB. The third lists what is already downloaded. Expect the first invocation of any alias to spend its time on the download rather than on generation.

Vision and audio follow the same pattern. Passing an image with the -i flag routes to a VLM alias, and the README gives this example:

```bash
swama run gemma3 "What's in this image?" -i /path/to/image.jpg
```

For building from source, the README documents cloning the repository, running swift build -c release inside the swama directory, renaming the binary to swama-bin, and building the macOS app separately with xcodebuild against Swama.xcodeproj. That path is labelled advanced, and it is the only one that requires Xcode 16.0+.

## Where Swama is the wrong tool

The platform lock is absolute. macOS 15.0 or later and Apple Silicon are hard requirements, so a Linux CI runner, a Windows workstation, or an Intel Mac cannot run this at all. If your inference has to live next to a CUDA GPU, Swama is not a candidate regardless of how well the Swift layer is written.

The model catalogue is curated, not open. The README publishes a fixed alias table mapping short names to specific mlx-community and lmstudio-community repositories, with sizes attached. You can also pass a full model name directly, as the README shows with mlx-community/Llama-3.2-1B-Instruct-4bit, but the aliases are what the quick start leans on. A model that has no MLX conversion does not become available just because Swama exists.

Memory is the other wall. The alias table is honest about scale: qwen3-235b is listed at 123.2 GB and qwen3.5-397b-a17b at roughly 220 GB. Those entries exist in the table, but they are not runnable on a typical laptop. The README does not document a fallback, a quantized variant for those specific aliases, or a CPU offload path, so treat the large aliases as entries for machines that can hold them. There is also an experimental label on /v1/audio/speech, which is the README's own signal that this endpoint should not be the foundation of anything you depend on.

One more gap: the README does not document rollback, version pinning for cached models, or what happens when a cached model is updated upstream. If reproducibility of a specific weight revision matters to you, that is unaddressed in the documentation.

## Swama against llama.cpp and Ollama

The obvious comparison is Ollama, and the difference is the runtime underneath. Ollama wraps llama.cpp and runs its own model registry and daemon; it is portable across macOS, Linux and Windows, and its GGUF-based models are widely available. Swama instead sits on MLX, Apple's own framework, and is written in Swift. The practical consequence is that Swama is not portable at all, while Ollama is. The benefit Swama claims in exchange is integration with Apple Silicon through MLX, plus a native menu bar app and a Swift package you can link into your own macOS code through SwamaKit.

Against llama.cpp used directly, the difference is the serving layer. llama.cpp gives you a binary and a server, but you assemble the model management, the API compatibility and the app around it. Swama ships an OpenAI-compatible endpoint set including embeddings and audio transcription in the same process, and an alias system that downloads on first use. If you already have a llama.cpp pipeline tuned and scripted, Swama's value proposition is mostly the API surface and the Swift integration, not the inference core.

There is a second axis: language. Swama is a Swift framework first. If you are building a macOS app and want inference as a library rather than as a subprocess, SwamaKit is the piece that has no direct equivalent in either alternative.

## Licence, releases and what an upgrade costs

Swama is MIT licensed, which permits commercial and closed-source use, modification and redistribution provided the copyright notice and permission notice are preserved. That is the standard permissive position and it is compatible with shipping SwamaKit inside a proprietary macOS application. This is a description of the licence text, not legal advice; if you are redistributing the binary or the framework, have your own counsel confirm the notice requirements for your distribution format.

The release cadence visible in the repository is three releases over roughly three months: v2.2.0 on 2026-06-08, v2.3.0 on 2026-08-26, and v2.4.0 on 2026-09-09. The last push to the default branch was on 2026-09-09, the same day as v2.4.0. A CHANGELOG.md sits at the repository root, which is where upgrade notes would live, and a VERSION file is also present.

Upgrade cost is dominated by the model cache rather than by the binary. Homebrew handles the CLI; the app is replaced by dragging a new build into Applications. The expensive part is that a model alias can point at a different upstream repository or revision between releases, and the README does not describe a pinning mechanism or a migration step for an existing cache. In practice, budget disk space and download time for a re-fetch after a version bump, and check CHANGELOG.md before upgrading rather than after.

## Conclusion

Adopt Swama if you are on an Apple Silicon Mac running macOS 15 or later and you want a local model server that speaks the OpenAI chat, embeddings and audio endpoints, with a menu bar app for model management. Do not adopt it if you need Linux, CUDA, or a server you can run on non-Apple hardware; the repository targets Apple Silicon exclusively. Before committing, verify three things on your own machine: that swama run qwen3 completes an auto-download and prints a reply, that your client works against the /v1/chat/completions endpoint including tool calling, and that the audio transcription endpoint behaves as you need, since the README labels the speech endpoint experimental. The README does not document rollback or downgrade steps for a model cache that has already been populated.

## FAQ

### What is Swama and what makes it different from other local LLM runners?

Swama is a machine learning runtime written in pure Swift for macOS, built on Apple's MLX framework, for local LLM and VLM inference. The distinguishing parts are the native Swift implementation, the OpenAI-compatible endpoints including embeddings and audio transcription, and the split into SwamaKit, a CLI and a menu bar app.

### How do I install Swama on macOS?

The README recommends Homebrew with brew install swama. Alternatively you can download Swama.dmg from the Releases page, drag Swama.app into Applications, and then use the menu bar app's "Install Command Line Tool..." option to put the swama command on your PATH.

### What are the system requirements for Swama?

The README lists macOS 15.0 or later, Apple Silicon (M1/M2/M3/M4), Xcode 16.0+ for compilation, and Swift 6.2+. The Xcode requirement applies to building from source rather than to the Homebrew install.

### Does Swama need a separate step to download models?

No. The README states that models are downloaded automatically on first use and cached for future use, so swama run qwen3 will fetch the model if it is not already present. The swama list command shows which models are already downloaded.

## Sources

- [Issues](https://github.com/Trans-N-ai/swama/issues)
- [License: MIT](https://github.com/Trans-N-ai/swama/blob/main/LICENSE)
- [README](https://github.com/Trans-N-ai/swama/blob/main/README.md)
- [Releases](https://github.com/Trans-N-ai/swama/releases)
- [Trans-N-ai/swama on GitHub](https://github.com/Trans-N-ai/swama)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/trans-n-ai-swama
