# LLM.swift: Running GGUF Models Locally on Apple Platforms

> LLM.swift wraps llama.cpp in a small Swift API for macOS, iOS, watchOS, tvOS and visionOS apps. It is easy to adopt, but the README leaves memory limits, threading and model licensing to the developer.

**eastriverlee/LLM.swift** — LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.

- Repository: https://github.com/eastriverlee/LLM.swift
- Stars: 878 · Forks: 126
- Language: Swift
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/eastriverlee-llm-swift

## What LLM.swift solves, and who it is actually for

Running a language model inside a Swift app normally means linking llama.cpp, writing the tokenizer plumbing, and deciding when to keep context between turns. LLM.swift packages that work behind a class you subclass. The README describes it as "a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS." The audience is Apple-platform developers who want inference on the device rather than a call to a hosted endpoint.

The interesting part is the platform list. watchOS and tvOS are unusual targets for local inference, so the library is not just a thin macOS wrapper. The topics list includes gguf, ios, llm, llm-inference, macos, swift, tvos, visionos and watchos, which matches the README's claim.

It is not a chat application. There is no server, no HTTP endpoint, no prompt-management UI. You get a model handle, a preprocess step, a completion call, and a streaming output property that your own SwiftUI view can observe.

## How inference works: gguf files, embedded templates and incremental context

The unit of work is a gguf file. You either bundle one in your app or let HuggingFaceModel download one at runtime. The LLM class loads it, and getCompletion returns generated text.

Chat templates changed in v3. The README states that by default LLM.swift renders conversations using the template embedded in the gguf file, executed by llama.cpp's Jinja engine, so the template: argument is optional when the model ships one. Passing a Template explicitly overrides the embedded one, which the README recommends when a model's gguf metadata is broken or missing.

Context handling is also incremental. A note in the README says conversation context is maintained between turns and only new tokens are evaluated, rather than re-feeding the whole history each turn. That matters on mobile, where re-evaluating a long conversation would dominate latency.

Reasoning models get their own surface. The README says bot.thinking separates thinking output automatically for models that support it, with no marker configuration. The SwiftUI example renders bot.thinking in gray above bot.output.

The @Generatable macro takes a different route to reliability. Annotating a struct or enum generates a JSON schema that constrains the model's output, and respond(to:as:) returns a value. The README calls this "100% reliable" type-safe structured output. Treat that as the project's own claim, not an independently measured result.

## Installing LLM.swift and getting a first completion

The README does not print a package installation command. The repository layout shows a Package.swift at the top level, so the package is consumed through Swift Package Manager in the usual way: add the dependency in Xcode or in your own Package.swift, pointing at the repository. The README's own examples all start from an import, so the first real step after adding the package is the import line.

Once the package is linked, the minimal path is to bundle a gguf file and construct an LLM. The README gives this example verbatim:

```swift
let bot = LLM(from: Bundle.main.url(forResource: "gemma-3-4b-it-q4_0", withExtension: "gguf")!, template: .gemma)
let question = bot.preprocess("What's the meaning of life?", [])
let answer = await bot.getCompletion(from: question)
print(answer)
```

The force unwrap on the bundle URL is the README's own choice. In shipping code you would guard that lookup, because a missing resource crashes here rather than returning an error.

If you would rather not ship weights inside the app, the Hugging Face path downloads the gguf and reports progress. The README shows the model identifier and quantization explicitly:

```swift
let systemPrompt = "You are a sentient AI with emotions."
let bot = await LLM(from: HuggingFaceModel("unsloth/Qwen3-0.6B-GGUF", .Q4_K_M, template: .chatML(systemPrompt)))!
```

For a streaming UI, the README's SwiftUI example subclasses LLM, calls await bot.respond(to: input) inside a Task, and calls bot.stop() to cancel. The view reads bot.output and bot.thinking directly, which is why the class is observable.

## The maxTokenCount trade-off the README warns about

The most useful piece of guidance in the README is a tip about maxTokenCount, the parameter passed when initializing LLM. It says the value should be tuned because of the memory and computation it requires, that lowering it helps speed on mobile devices, and that setting it too low, "to a point where two turns cannot even fit", causes quality to decrease because context is cut off.

That is a real constraint stated plainly. Context length is not free: it is allocated up front and it costs both memory and time per token. On a phone or a watch, the ceiling is the device, not the model. The README does not give a table of recommended values per device, so the number is something you determine by measuring your own app on your own hardware.

The README does point elsewhere for scale. It notes that mistral 7B based models worked on an iPad Air 5th gen at Q5_K_M and an iPhone 12 mini at Q2_K, and says that generally, for mobile devices, 3B or larger parameter models are recommended, linking to a llama.cpp benchmark discussion for details. Those are the author's reported test devices, not a published benchmark from this repository.

## Where LLM.swift is the wrong tool

The README documents no rollback, no migration notes and no deprecation policy. The release history shows three releases inside roughly twelve hours on 2026-07-18 and 2026-07-19, which is a fast-moving surface. If your app pins a version and you plan to upgrade on your own schedule, budget time to read the diff rather than assuming the API is frozen.

There is also no documented memory budget. The README tells you to tune maxTokenCount and warns that too low a value degrades quality, but it does not say how much memory a given model and context length will consume on a given device. You will find that out by running the app.

If your problem is serving many concurrent users, this is the wrong shape entirely. There is no batching, no request queue, no server mode. Local single-user inference is the design.

Model licensing is another boundary. The library is MIT, but the weights are not covered by that licence, and the README does not discuss model terms. Downloading from Hugging Face at runtime also means the app depends on that host being reachable at first run.

## LLM.swift against llama.cpp bindings and MLX Swift

The closest comparison is using llama.cpp directly, or through a thin Swift binding. LLM.swift is built on llama.cpp, so the inference engine is the same family. The difference is what sits on top: an LLM class you subclass, a preprocess and getCompletion pair, an observable output string for SwiftUI, a Hugging Face downloader with a progress callback, and the @Generatable macro for schema-constrained output. With raw llama.cpp you would write that layer yourself, and you would also control exactly which llama.cpp revision you build against.

MLX Swift takes a different route. It targets Apple silicon through Apple's MLX array framework rather than gguf and llama.cpp. That means a different model format and a different conversion path, and it is not the format the gguf ecosystem publishes. If your existing pipeline already produces gguf files, LLM.swift stays closer to that pipeline; if you are starting fresh on Apple silicon and want the MLX toolchain, the two do not share artifacts.

A third option is calling a hosted API. That removes device memory limits and model download size, at the cost of network dependency, per-request billing and sending user data off device. LLM.swift exists precisely for the cases where that trade is unacceptable.

## Licence, maintenance and what an upgrade costs

The repository is MIT licensed, and the LICENSE file is at the top level. That covers the library source. It does not cover the gguf weights you bundle or download, which carry their own terms from whoever published them. Nothing here is legal advice; check the model card for the model you actually ship.

Maintenance is visible but not documented as a policy. The last push was on 2026-07-19, and v3.0.3 carries that same timestamp. The repository is not archived. There is a CONTRIBUTING.md and an update.sh at the top level, and the presence of Tests/ and LLM.xctestplan suggests the package is exercised by a test plan rather than only by the example app.

The upgrade cost is mostly the model-facing API. The v3 line moved chat templates to the gguf-embedded default and made HuggingFaceModel's template parameter optional, so code written against the older explicit-template style still works but is no longer the documented default. The README does not publish a migration guide, so an upgrade is a read-the-source exercise. Pinning to a tag and testing your own prompts after each bump is the practical approach, because template handling directly changes model output.

## Conclusion

Adopt LLM.swift if you are building a native Apple app and want a local GGUF model behind a small Swift API, and you can bundle or download the weights yourself. Do not adopt it if you need a server-side serving stack, a stable API surface across major versions, or a documented memory budget for older devices. Before writing UI code, verify three things from the repository itself: that your deployment targets satisfy Package.swift, that your chosen model fits the device after quantization, and that the model's own licence permits your distribution model.

## FAQ

### What is LLM.swift?

It is a Swift library for interacting with large language models locally on macOS, iOS, watchOS, tvOS and visionOS, built around gguf model files. The README describes it as simple and readable, with an LLM class you subclass and a getCompletion call for generation.

### How do I install LLM.swift in a Swift project?

The README does not print an installation command, but the repository has a top-level Package.swift, so it is added as a Swift Package Manager dependency. After linking it, the README's examples begin with import LLM.

### Can LLM.swift load a model from Hugging Face instead of bundling one?

Yes. The README shows initializing LLM with a HuggingFaceModel value that names the repository, a quantization such as .Q4_K_M, and an optional template, and it also shows an initializer that reports download progress through a callback.

### Does LLM.swift need a chat template passed in?

No. The README states that by default conversations are rendered with the chat template embedded in the gguf file, executed by llama.cpp's Jinja engine, so the template argument is optional. Passing a Template explicitly overrides the embedded one when a model's metadata is broken or missing.

### What does maxTokenCount affect in LLM.swift?

The README's tip says the value affects memory and computation, that lowering it improves speed on mobile devices, and that setting it too low so two turns cannot fit causes quality to decrease because context is cut off.

## Sources

- [eastriverlee/LLM.swift on GitHub](https://github.com/eastriverlee/LLM.swift)
- [Issues](https://github.com/eastriverlee/LLM.swift/issues)
- [License: MIT](https://github.com/eastriverlee/LLM.swift/blob/main/LICENSE)
- [README](https://github.com/eastriverlee/LLM.swift/blob/main/README.md)
- [Releases](https://github.com/eastriverlee/LLM.swift/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/eastriverlee-llm-swift
