Olla routes LLM requests across self-hosted inference nodes, with a proxy engine you pick at startup
High-performance lightweight proxy and load balancer for LLM infrastructure. Intelligent routing, automatic failover and unified model discovery across local and remote inference backends.
At a glance
- What is it?
- Olla is a Go proxy and load balancer that sits in front of self-hosted inference servers such as Ollama, LM Studio, vLLM, vLLM-MLX, llama.cpp, and Docker Model Runner, and is meant to be used alongside an existing gateway or orchestration platform rather than instead of one. Its distinguishing details are operational: two proxy engines, KV-cache-aware sticky sessions, and a configuration file whose path decides whether your server binds to loopback.
- Who is it for?
- Olla fits a team already running several local inference backends that keep disagreeing about which one is healthy and which model each one actually has. It does not replace an API gateway or an orchestration platform; the project's own framing is that it makes existing infrastructure reliable, sitting beside LiteLLM or GPUStack rather than in place of them.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Two proxy engines, one for maintainability and one for speed
The engine choice is explicit and it is the first thing to settle. Sherpa is the simple, maintainable option; Olla is the high performance one, carrying the advanced features such as circuit breakers and connection pooling. Everything else in the project, routing, failover, model discovery, is the same either way.
The positioning is deliberately narrow. Olla is described as working alongside API gateways like LiteLLM or orchestration platforms like GPUStack, with the goal of making your existing LLM infrastructure reliable through intelligent routing and failover. So the deployment shape is a proxy in front of your inference nodes, not a replacement for your gateway layer, and adopting it does not mean removing what you already have.
Distribution is a single CLI application with a single config file, and the release is a Go binary built with goreleaser, published as `v0.0.29` most recently, with the container image on GitHub Container Registry as `ghcr.io/thushan/olla:latest`.
The project is Apache 2.0 licensed and its module targets Go 1.24.0. Larger deployments are pointed elsewhere on purpose: large GPU deployments, enterprise, and data centre use are directed to TensorFoundry FoundryOS, and an inference control plane is directed to Alloy.
Sticky sessions exist because the KV cache is what you are paying for
Routing is priority-based, with automatic failover and connection retry, but the feature that changes request distribution most is sticky sessions. They are described as KV-cache-aware affinity routing that pins a multi-turn conversation to the same backend.
The reason is that prefix caching lives in the backend's KV cache. If turn two of a conversation lands on a different node, that node has to recompute the prompt prefix instead of reading it from cache, which turns a cheap continuation into a full prefill. Load balancing alone will do this constantly on a multi-turn workload, because each request looks independent to a naive router.
Failure handling is separate from affinity. Intelligent retry handles connection failures with immediate transparent endpoint failover, so a node that dies is skipped rather than retried against. Health monitoring runs continuous endpoint health checks with circuit breakers and automatic recovery, and self-healing refreshes the model discovery catalogue when an endpoint comes back.
Those three mechanisms are separable by design. A circuit breaker stops sending requests to a node that is failing, the retry path moves a single request elsewhere, and the catalogue refresh means a recovered node is correctly described again rather than silently missing its models.
Passthrough versus translation is decided per backend, and the versions differ
Olla speaks the Anthropic Messages API, and the integration table records for each backend whether `/olla/anthropic/` requests are forwarded untouched or translated from Anthropic format to OpenAI format on the fly. That single column is the compatibility fact most people need.
The version markers are uneven, and instructively so. Ollama is marked passthrough from v0.14.0+, LM Studio from v0.4.1+, vLLM from v0.11.1+, vLLM-MLX as plain passthrough, and llama.cpp from `b4847+`, which is a build identifier rather than a release number.
Two conclusions follow. The first is that a backend on an old version is not passthrough-capable, so Claude Code and other Anthropic-shaped clients will be translated instead, and translation is a different code path with different behaviour. The second is that upgrading a backend can change which path Olla takes for your traffic, so a backend upgrade is also a routing change to test.
The supported set is broader than the rows shown here, with LMDeploy, SGLang, Lemonade, and omlx all linked from the project, and the claim is that the list is extensible for endpoints not natively supported.
The container copies config/config.yaml because the loader searches that path first
There is a configuration trap in the container build, and it is documented in a comment inside the Dockerfile rather than in prose.
The image copies `build/docker-config.yaml` to `config/config.yaml`, not the repository root `config.yaml`. The reason given is that the loader searches `config/config.yaml` first in `internal/config/config.go`, so in a repo-root build context the bare-metal config would win and bind the server to loopback.
That is a real deployment difference, not a cosmetic one: a container that binds to loopback is unreachable from anywhere but the host, and the symptom appears as a dashboard that works locally and refuses connections from the network. Two `config` locations in one repository, one loader precedence rule, and a comment that exists because somebody got it wrong.
The same Dockerfile also explains why files are copied individually rather than with `COPY . .`: the image is built from two different contexts, the repository root for a local build and goreleaser's synthesised extra_files context, and naming the files keeps both images identical instead of depending on `.dockerignore` to filter one of them. It runs on Alpine as an unprivileged `olla` user, exposes 40114, and its healthcheck is a `wget` spider against the internal health endpoint.
The dashboard is embedded in the binary and loopback-only by default
Since v0.0.29 the binary embeds a read-only dashboard at `/internal/ui/`, showing fleet health, per-endpoint latency, and the discovered model inventory. There is no extra port and no separate web server; it is served from the same process.
wget --no-verbose --tries=1 --spider http://localhost:40114/internal/health || exit 1That line is the container healthcheck, and it points at `/internal/health`, the endpoint the compose healthcheck also uses every 30 seconds with a 10 second timeout, 3 retries, and a 40 second start period.
The binding is the part to notice. The dashboard is on by default and loopback-only, which means browsing the herd means being on the machine or tunnelling to it. That is the right default for something that reveals which models you have and how fast each node is answering, and it means anyone wanting it reachable has to make that decision deliberately.
The visual side has a light and a dark theme, and the page states that nothing leaves your network.
Filtering is glob patterns plus an expression engine over profiles
Advanced filtering is documented as profile and model filtering with glob patterns, and the repository layout backs that up: a `config/profiles/` directory is copied into the container, alongside a separate `config/models.yaml`, and both files are copied by name in the Dockerfile.
The Go module explains the mechanics. `expr-lang/expr` is a direct dependency, which is the expression language used to evaluate filtering rules, and `tidwall/gjson` is there to pull values out of JSON responses without unmarshalling everything. Caching for model catalogues comes from `jellydator/ttlcache`, so discovery results expire and are refetched rather than cached forever.
Model unification is the other half. Per-provider unification plus OpenAI-compatible cross-provider routing means a client can ask for a model by name and be routed to whichever endpoint has it, with a unified catalogue per provider. The output carries response headers and statistics, which is what makes per-request tracking possible at the proxy layer.
Backend authentication is configured per endpoint, covering local backends that need an API key, a bearer token, or basic auth, so a fleet with mixed auth requirements can sit behind one proxy.
The whole proxy is expected to fit in under 50 megabytes of RAM
The performance claims are specific enough to check. Endpoint selection is sub-millisecond using lock-free atomic statistics, and the process is described as lightweight, running on less than 50 megabytes of RAM. Design is streaming-first, with timeouts optimised for long inference rather than for short HTTP requests.
The dependency list supports the shape. `puzpuzpuz/xsync` provides concurrent maps, `golang.org/x/sync` handles the waitgroup style coordination, and `lumberjack` does log rotation so a long-running proxy does not fill a disk with logs. `go-isatty` is there to decide whether to colour terminal output, and `pterm` provides the progress and status rendering, so the CLI presentation is part of the binary rather than a shell script wrapper.
Production features are present rather than aspirational: rate limiting, request size limits, optional CORS for browser clients, configurable proxy and server timeouts, and graceful shutdown.
The container runs the same single binary as the CLI, `ENTRYPOINT ["olla"]`, and the compose file mounts a local config read-only, sets `OLLA_CONFIG_FILE` and `OLLA_LOG_LEVEL=info`, and joins a bridge network so an Ollama container can be attached alongside it.
Editorial conclusion
Olla fits a team already running several local inference backends that keep disagreeing about which one is healthy and which model each one actually has. It does not replace an API gateway or an orchestration platform; the project's own framing is that it makes existing infrastructure reliable, sitting beside LiteLLM or GPUStack rather than in place of them. Before adopting it, decide which proxy engine you actually need, check that each backend's Anthropic passthrough version is new enough for your clients, and read which config file wins, because copying the wrong one silently binds the server to loopback.
Frequently asked questions
What is Olla and what does it actually do?
Olla is a Go proxy and load balancer for LLM infrastructure that routes requests across self-hosted inference nodes with priority-based routing, automatic failover, sticky sessions, and unified model catalogues per provider. It is designed to sit alongside a gateway like LiteLLM or an orchestrator like GPUStack to make existing infrastructure more reliable.
Which inference backends does Olla support?
The integration table covers Ollama, LM Studio, vLLM, vLLM-MLX, llama.cpp, and Docker Model Runner, with LMDeploy, SGLang, Lemonade, and omlx also linked from the project. For Anthropic API requests each backend is either passed through natively or translated from Anthropic format to OpenAI format.
What is the difference between the Sherpa and Olla proxy engines?
You choose between them at startup: Sherpa is the option for simplicity and maintainability, while Olla is the high performance engine that carries advanced features such as circuit breakers and connection pooling. Routing, failover, and model discovery work the same way in both.
How do I run Olla in Docker?
The compose file uses ghcr.io/thushan/olla:latest with port 40114 mapped to the host, a config file mounted read-only at /config/config.yaml, OLLA_CONFIG_FILE pointing at it, and OLLA_LOG_LEVEL=info. The healthcheck calls http://localhost:40114/internal/health every 30 seconds with a 40 second start period.
Where is the Olla dashboard and can I expose it?
Since v0.0.29 a read-only dashboard is embedded in the binary at /internal/ui/, served from the same process with no extra port, and it is on by default but bound to loopback only. It shows fleet health, per-endpoint latency, and the discovered model inventory.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/thushan-olla)