llama-swap: one OpenAI-shaped endpoint in front of every local model server
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
At a glance
- What is it?
- A Go binary that starts, stops and routes between llama.cpp, vLLM, stable-diffusion.cpp and ComfyUI on demand, so a client that only knows the OpenAI API can reach any of them without holding every model in memory at once.
- Who is it for?
- llama-swap solves a problem that only appears once you run more than one local model: the GPU can hold one of them at a time, but every client you use insists on talking the OpenAI protocol to a fixed address. Its configuration file, a documented JSON schema and a large endpoint surface make that mismatch disappear at the cost of one more process in the path.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem is memory, not protocol
Running a local model server is easy. Running several at once is not, because a single GPU holds a limited amount of model weights and the inference servers are built assuming they own it. The usual workarounds are all bad: run everything and let the driver thrash, or restart a server by hand before each prompt.
llama-swap sits between your client and those servers. It exposes one address that speaks the OpenAI and Anthropic API shapes, looks at which model the request names, and starts that model's server if it is not already running. When the request moves to a different model, the previous server is stopped. The client sees a stable endpoint and never learns that a process was launched on its behalf.
The README lists the servers it can front as llama.cpp and its forks, vLLM, stable-diffusion.cpp, audio.cpp and ComfyUI, with the framing that it is future-proof against upgrading any of them. That is the right way to think about the design: llama-swap owns routing and lifecycle, the model servers own inference, and the seam between them is the OpenAI-compatible HTTP surface.
Every endpoint it will proxy, listed explicitly
The feature list is unusually specific about routes, which makes it easy to check compatibility before you commit. The OpenAI-shaped surface covers `v1/completions`, `v1/chat/completions`, `v1/responses`, `v1/embeddings`, `v1/models`, `v1/audio/speech`, `v1/audio/transcriptions`, `v1/audio/voices`, `v1/images/generations` and `v1/images/edits`. The Anthropic shape is narrower, covering `v1/messages` and `v1/messages/count_tokens`.
Then come the server-native routes. llama-server's own surface adds `v1/rerank`, `v1/reranking`, `/rerank`, `/infill` for code infilling, `/completion`, `/models` and `/props`. stable-diffusion.cpp's SDAPI routes are proxied too, including `/sdapi/v1/txt2img`, `/sdapi/v1/img2img` and `/sdapi/v1/loras`, where the `model` field in the request body selects which LoRA set to load. audio.cpp gets `/audioapi/v1/tasks/run`, and ComfyUI gets `/comfyui/`.
Several of those routes carry an explicit caveat in the documentation, which is the kind of honesty that makes a route list usable. `/props` requires a `?model=` query parameter and ignores the autoload parameter. `/sdapi/v1/loras` needs `model` in the body to fetch the right LoRAs. Knowing that up front saves a debugging session.
Control and observability routes beyond proxying
Past the model routes, llama-swap adds its own API, and this is where it stops being a thin proxy. `/upstream/:model_id` bypasses the router and talks to a specific upstream directly, which is what you want when you suspect the router is the problem. `/running` lists what is currently loaded, and `POST /api/models/unload` or `POST /api/models/unload/:model_id` shut things down on demand.
Profiles let you change model ID routing at runtime. `GET /api/profiles` lists configured profiles and the active selection, and `PUT /api/profiles/active` switches them, so a client that hardcodes `local-model` can be pointed at a different physical model without touching the client.
Log streaming is unusually thorough. `/logs` returns buffered plain text and redirects to the UI if you send `Accept: text/html`. `/logs/stream` holds a connection open for live output, buffered history first unless you add `?no-history`. There are separate streams for proxy logs only, upstream process logs only, and one named model including IDs containing slashes like `author/model`. `/health` returns OK and `/metrics` exposes system and GPU metrics for Prometheus.
There is also API key support to restrict access, and a `/ui` web interface with a playground, token metrics, request inspection, manual load and unload, and live log streaming.
Configuration: models, ttl, filters and the swap matrix
Configuration is one YAML file, and the repository ships both an example and a schema: `config.example.yaml` and `config-schema.json`. The README also links to the docs site for the full reference.
Four options do most of the work. A `ttl` unloads a model after a timeout, which is the simple answer to GPU memory. Profiles switch model ID routing at runtime. Filters modify requests before they reach the upstream, using `stripParams`, `setParams` and `setParamsByID` to remove or force inference parameters per model or per request. `hooks` preload models at startup.
The interesting one is the swap matrix. The README describes it as a custom DSL that lets you run concurrent models, referenced through issue 643. That is a different arrangement from sequential swapping: instead of one model at a time, you declare which combinations may coexist, which is what you need when a small embedding model has to stay resident alongside a large generation model. Docker and Podman support comes through `cmd` and `cmdStop` together, so a container image can be started and stopped as a unit.
For anyone scripting model lifecycle outside the tool, the JSON schema is the piece to keep. It is a machine-readable contract you can validate a config against, which is rare and useful in a project that has been releasing weekly since at least v253 in September 2026.
Zero dependencies in the README, thirty modules in go.mod
The README says llama-swap has zero dependencies and no external dependencies, twice, and that it is built in Go for performance and simplicity. The literal claim is about deployment: you run one binary with one config file and you do not need a service mesh or a container runtime to get going. Read that way, it is accurate.
Read as a statement about the build, it is not what `go.mod` shows. The module requires Gin for HTTP, Bubble Tea, Bubbles and Lip Gloss for the terminal interface, gjson and sjson for JSON handling, gojq for queries, CBOR, Goose for migrations, gopsutil for metrics, modernc.org/sqlite, and a pinned pre-release of tailscale.com that the v255 release added for the tailcat peer connectivity feature.
That last detail is worth flagging as a practical matter rather than a criticism. Depending on a version string like `v1.103.0-pre.0.20260904030409-31d8badb3bfb` means a build can change behaviour when the upstream dependency moves, and it is the kind of pin that gets reviewed when a transitive vulnerability appears. The build is reproducible from the committed `go.mod`, but it is not a dependency-free program.
test:
go test -short -count=1 ./internal/...The Makefile keeps that distinction honest in practice. The default `all` target builds macOS, Linux and a simple responder binary, and every platform target depends on a `ui` target that runs an npm build first, so the embedded web interface is compiled from TypeScript before the Go binary is linked.
Container images and which tag to pull
Two image families are built nightly, and the README is candid that one of them is legacy. The unified image builds llama-server, ik-llama-server, stable-diffusion.cpp, whisper.cpp, audio.cpp and llama-swap all from source, and the project recommends it. The legacy image is llama.cpp's own container with llama-swap copied in, so it carries only what that base image ships and has no image generation, speech or ik-llama-server.
The unified family has three tags and the choice is driven by your GPU:
$ docker pull ghcr.io/mostlygeek/llama-swap:unified-cuda13`unified-cuda13` covers NVIDIA Ampere through Blackwell, lists amd64 and arm64, and is a multi-arch tag so a `docker pull` resolves correctly on an aarch64 host such as a DGX Spark. `unified-cuda` is amd64 only and covers Pascal through Ada with CUDA 12, for cards where CUDA 13 is no longer offered. `unified-vulkan` is the AMD and other Vulkan-capable path.
Installation is otherwise broad: Docker, Homebrew on macOS and Linux, MacPorts, WinGet, release binaries and source builds. What is not documented on the repository page is how you get model weights, and that is correctly outside the tool's remit: it manages servers, not files. Point the models volume at wherever you keep weights and the swap behaviour does not change.
Editorial conclusion
llama-swap solves a problem that only appears once you run more than one local model: the GPU can hold one of them at a time, but every client you use insists on talking the OpenAI protocol to a fixed address. Its configuration file, a documented JSON schema and a large endpoint surface make that mismatch disappear at the cost of one more process in the path. Read `config.example.yaml` against `config-schema.json` before you write your own file, then decide whether the swap matrix or simple TTL unloading matches how you actually move between models.
Frequently asked questions
What does llama swap do?
It acts as a proxy that starts and stops local model servers on demand while presenting one stable OpenAI or Anthropic compatible endpoint. A request naming a model causes that model's server to be launched if it is not already running, and moving to a different model stops the previous one.
Does llama swap work with Vllm?
Yes. The README lists vLLM among the servers llama-swap can front, alongside llama.cpp and its forks, stable-diffusion.cpp, audio.cpp and ComfyUI, on the basis that the tool works with any OpenAI and Anthropic API compatible server.
What are the fees associated with LlamaSwap?
llama-swap is an MIT licensed project with no fees for the software itself. Its cost is operational: you need the hardware to hold the model weights, and the container images are published on GitHub's own registry. Some routes also assume an upstream server such as llama.cpp or vLLM is already available to launch.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/mostlygeek-llama-swap)