Llama for macOS: a menu bar front end for llama.cpp
A cosy home for your LLMs.
At a glance
- What is it?
- Llama is a 4 MB macOS menu bar app that starts a local llama.cpp server on port 9931, recommends GGUF models that fit your Mac, and shares the Hugging Face cache with other tools. Here is how it installs, how its config layering works, and where it stops being the right choice.
- Who is it for?
- Adopt Llama if you want a local OpenAI-compatible endpoint on a Mac without hand-writing llama serve flags, and you accept that the app owns the generated models.ini. Skip it if you need a headless Linux server, Docker deployment, or a fully hand-maintained llama.cpp configuration.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 10 days ago.
- What is it written in?
- Mainly Swift, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Llama solves on a Mac, and for whom
Running a local model on macOS usually means three separate chores: getting a llama.cpp binary that matches your architecture, finding a GGUF quantization that fits your unified memory, and remembering the right llama serve flags for that model. Llama collapses those into a menu bar item. The README describes it as "a macOS menu bar app for running local LLMs," and the feature list is explicit that it is built on llama.cpp from the same GGML org, developed alongside it.
The intended user is a Mac owner who wants a chat UI and an API endpoint without managing a server process by hand. The app is 4 MB and native, so it is not an Electron shell around a Python stack. If you already run llama.cpp, the README says models you have installed through it appear in the app automatically, which means Llama is positioned as a front end rather than a separate model ecosystem. If you have never installed llama.cpp, the app installs a prebuilt binary for your Mac instead, so the dependency is satisfied either way.
The second audience is tooling that speaks the OpenAI HTTP shape. Coding agents, chat UIs and editors are named in the README as things you connect to the local server, which is why the API surface matters more here than the chat window.
The server on port 9931 and how models load
Starting Llama starts a local server at http://localhost:9931/v1. That path suffix is the giveaway: the app exposes an OpenAI-compatible route set, and the README points at the llama.cpp server documentation for the complete API reference rather than restating it. Two endpoints are shown in the README: a model listing and a chat completion.
Model lifecycle is the part worth understanding before you adopt it. Models load when requested and unload when idle, so memory is released when nothing is asking for inference. That is a deliberate trade against latency: the first request after an idle period pays the load cost, and a workflow that alternates between two models will keep paying it. The README does not document a keep-alive setting or an idle timeout value, so if predictable first-token latency matters to you, treat this as an open question and check the app's settings yourself.
Storage is the other design decision. Models live in the Hugging Face cache, shared with llama.cpp and other tools, rather than in an app-private directory. Nothing is duplicated, and deleting the app does not strand tens of gigabytes of weights in an unexpected place. The cost is that the app's view of "installed models" is really a view of that shared cache.
Install and first request
The README gives a Homebrew cask as the primary install path, with a Releases download as the fallback. The cask name contains a hyphen, so copy it exactly.
brew install --cask llama-appAfter launching the app, the README says a local server is running at http://localhost:9931/v1. Confirm that and see which models are visible to the app with the listing endpoint:
curl http://localhost:9931/v1/modelsThe response is a model list in the OpenAI shape. If it is empty, no GGUF weights have been found in the Hugging Face cache yet, and the app's built-in recommendation list is the intended way to fill it: the README describes a list of models your Mac can run, installable in one click.
A first completion uses the model identifier exactly as the listing reports it. The README's own example sends a single user message:
curl http://localhost:9931/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "ggml-org/gpt-oss-20b-GGUF:MXFP4",
"messages": [{"role": "user", "content": "Hello"}]
}'If you would rather not touch curl, the built-in WebUI chats with any model the app can see, and the README treats that as the default first experience.
Two config files, and which one you are allowed to edit
This is the part of Llama that most deserves a careful read, because the app writes one file and reads another. It generates models.ini on every launch, deriving settings from your Mac's hardware. Editing that file is pointless: it is regenerated. Your changes go in ~/.config/llama/models.user.ini, which the app only reads and merges into the generated config. A commented template is written to that path on first launch.
Section headers must match the app's own format, org/repo:QUANT, and you list only the keys you want to change. Keys are llama serve options without the leading dashes, which is why the README's example uses temp, ctx-size and cache-type-k rather than their dashed spellings:
[ggml-org/gemma-4-E4B-it-GGUF:Q8_0]
temp = 0.7
ctx-size = 32768
cache-type-k = q4_1A section for a model the app did not discover is passed through as-is. That is how you point at weights from another repo, or attach a draft model for speculative decoding, by setting model, spec-type and spec-draft-model yourself.
The precedence rule is stated plainly: anything you set wins over the app's own value, including settings the app derived from your memory size. That is a real footgun. If you set ctx-size to a value your Mac cannot hold alongside the weights, the app does not protect you from it. The one guardrail is error handling: if a key cannot be applied, the app ignores the user file entirely and the menu names the option at fault. Note the blast radius. One bad key discards every override in the file, not just the offending line.
Reaching the server from other devices, and the risk in that
By default the server is reachable only from the Mac it runs on. The README documents two network modes behind an "Allow network access" setting, and they are not equivalent.
The Tailscale option binds the server to your Tailscale address. Other devices reach it from anywhere, Tailscale handles authentication and encryption, and the server stays invisible on whatever local network you happen to be on. The README calls this the option to use on a laptop, and it is shown only when Tailscale is installed and signed in. The second option, "This network", binds all interfaces at 0.0.0.0 so anything on your current network can reach it. The README's own warning is the important sentence here: the server has no password, so it should be used only on a network you trust, and never together with agent mode on a network you do not own.
There is an escape hatch for pinning to an address the app does not offer, set by hand through a defaults write against app.llama.Llama exposeToNetwork. The app leaves that value alone and falls back to localhost whenever the address is not present on any interface. If you need an API key, the README's route is custom server arguments, appended to the llama serve command after the app's own flags, with the caveat that they only override where the server honors the later occurrence of a flag. They take effect on the next server start, not immediately.
Where Llama is the wrong tool
The most obvious mismatch is platform. This is a macOS app distributed as a Homebrew cask and built with Xcode, and the repository's top level is an Xcode project plus a Swift source directory. There is no documented Linux or Windows path, no container image, and no headless mode described in the README. If your inference host is a Linux box in a rack, Llama is not a candidate, however good the menu bar experience is.
The second limitation is the generated config. Because the app regenerates models.ini on every launch and merges your file on top, you are always working in a two-layer system. A setting you deliberately chose can be shadowed by nothing, but a setting the app derives from hardware can change when your hardware or the app's heuristics change, and the generated file is not the place to record why a value was chosen. Teams that need a reviewable, version-controlled server configuration will find that the app's model of configuration is the opposite of theirs.
The third is the unauthenticated default. The README is direct that the server has no password. For a single-user laptop that is fine. For a shared machine or an office network, the only documented mitigations are Tailscale binding or passing an API key through extraServerArgs, and the latter is described as experimental and dependent on flag ordering.
Finally, the README does not document rollback, downgrade, or what happens to your models.user.ini across app updates. If you pin a config that matters, that silence is worth resolving before you depend on it.
How it compares to Ollama and to raw llama.cpp
The nearest alternative in daily use is Ollama, which also runs local models behind an HTTP API on macOS. The difference in approach is where the models live and who owns the configuration. Ollama maintains its own model store and pulls models through its own registry and manifest format. Llama keeps weights in the Hugging Face cache and installs GGUF files from Hugging Face, which is why the README can say models you installed via llama.cpp show up automatically. If you already have a Hugging Face cache full of GGUFs, Llama reuses it and Ollama does not, at least not without duplication.
The other alternative is running llama.cpp's server yourself. That gives you the full flag surface directly, a config file you fully control, and no app layer regenerating anything. What you give up is the recommendation list, the one-click install, the menu bar controls, and the automatic per-machine settings. Llama is essentially that manual setup with the tuning decisions made for you and a small Swift shell around them. If you enjoy writing llama serve invocations, you are not the target user. If you want the endpoint to exist without thinking about it, you are.
Editorial conclusion
Adopt Llama if you want a local OpenAI-compatible endpoint on a Mac without hand-writing llama serve flags, and you accept that the app owns the generated models.ini. Skip it if you need a headless Linux server, Docker deployment, or a fully hand-maintained llama.cpp configuration. Before trusting it, install it with brew install --cask llama-app, run curl http://localhost:9931/v1/models, and check which llama.cpp binary the app picked up, because that choice determines which serve options your models.user.ini keys can actually reach.
Frequently asked questions
Does Llama for macOS have a GUI?
Yes. It is a macOS menu bar app, and the README says you can chat with any model in the built-in WebUI. You can also skip the GUI entirely and use the local API on port 9931.
How do I install Llama on macOS?
The README gives a Homebrew cask, brew install --cask llama-app, with a download from the GitHub Releases page as the alternative. After launching the app it runs a local server at http://localhost:9931/v1.
Can I use llama.cpp on macOS with Llama?
Yes, and the README says Llama uses your existing llama.cpp install if you have one, otherwise it installs a prebuilt binary for your Mac. Models already installed through llama.cpp appear in the app automatically because both share the Hugging Face cache.
Can I use Ollama on macOS instead of Llama?
Ollama runs on macOS, but it is not covered in this repository's documentation, so nothing here describes how it behaves. The relevant difference for Llama is that it stores models in the Hugging Face cache shared with llama.cpp, while Ollama keeps its own model store.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ggml-org-llama-macos)