Model or dataset
yym68686/uni-api avatar
yym68686/uni-api

uni-api is a config file in front of a dozen model providers

This is a project that unifies the management of LLM APIs. It can call multiple backend services through a unified API interface, convert them to the OpenAI format uniformly, and support load balancing. Currently supported backend services include: OpenAI, Anthropic, DeepBricks, OpenRouter, Gemini, Vertex, etc.

1,266 stars159 forksPythonApache-2.0

At a glance

What is it?
An API gateway with no web interface: you write a YAML file naming providers, base URLs and keys, and it exposes OpenAI-shaped endpoints that fan out, retry, cool down and load balance across them. The published container is now a Rust binary while the release tags still track a Python version line, and that split shows up in the build files.
Who is it for?
uni-api fits a single user with several provider keys who wants OpenAI-shaped endpoints without operating a panel, and the channel cooling plus automatic retry are the parts that make a multi-provider setup survive a provider outage. Two things to check before you point real traffic at it.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

No interface, and a configuration file you must supply

The design premise is stated as a contrast with a larger alternative. One API gateway is described as too complex for personal use, carrying commercial features an individual does not need, and uni-api is offered for people who do not want a web interface and would rather support more models. So the entire control surface is a YAML file, and the project pitches itself as beginner-friendly with a detailed configuration guide.

There is no way to start without that file. Two routes are offered and they are mutually exclusive in intent. You either set a CONFIG_URL environment variable holding the configuration file's address, which is downloaded automatically when uni-api starts, or you mount a file named api.yaml into the container. The filename is fixed.

The structure is small. A providers list gives each entry a provider name, which is required but can be any string rather than an enforced identifier, a base_url, and an API key. An api_keys list holds the keys clients present to uni-api itself, which is the credential the gateway issues rather than one of the upstream credentials. Within a provider you can configure multiple API keys, and each provider can serve multiple models, with load balancing applied across them.

That layering is what separates uni-api from a proxy pointed at one upstream: the downstream keys are yours, and the upstream keys are yours, and they are configured in different places.

Models are discovered, or listed, or renamed

Model configuration has two modes. If you omit the model list, uni-api fetches everything available using the base URL and key through the provider's /v1/models endpoint, so adding a provider does not require enumerating its catalogue. If you do supply the list, it is used instead, with each entry giving a usable model name.

The list also carries a renaming syntax, written as a target name and a source name separated by a colon, for example mapping a dated model identifier such as claude-sonnet-4-5-20250929 onto the shorter name claude-sonnet-4-5. That matters more than it first appears in a gateway, because the name your client sends is the name you have to keep stable in your own configuration while the upstream moves to a new snapshot. A rename entry lets the client keep asking for the short name.

Two supporting controls sit alongside. Fine-grained model timeout settings allow a different timeout duration per model, which is the right answer to a fast model on a slow channel. Fine-grained permission control uses wildcards to restrict which models a given api_keys channel may reach, so one downstream key can be scoped to a cheap model while another is not.

Rate limiting is configured as a maximum number of requests per interval, given as an integer and a period: 2/min, 5/hour, 10/day, 10/month, 10/year are the forms given, with a default of 60 per minute. The intervals extend past anything most gateways offer, which suggests the design assumes someone sharing a key with a collaborator rather than running a service.

Three load balancing modes, and two of them are off by default

Load balancing is listed as three separate mechanisms rather than one, and the defaults are the interesting part. Channel-level weighted load balancing distributes requests according to per-channel weights; it is not enabled by default and requires configuring channel weights. Channel-level sequential load balancing is likewise off by default and requires setting SCHEDULING_ALGORITHM to round_robin. The third is automatic API key-level round-robin across multiple API keys inside a single channel, and no configuration is described for it.

That last asymmetry is the one to notice. If you add a second API key to a provider expecting it to be shared, it participates in key-level rotation without further setup, while splitting traffic across two providers needs a config change you will not find by reading the example file. Two of the three modes also describe the intent rather than the mechanism, with weighted and sequential described as improving an immersive translation experience in the second case.

Running the container is two commands, one per startup method:

bash
docker run --user root -p 8001:8000 --name uni-api -dit \
-v ./api.yaml:/home/api.yaml \
yym68686/uni-api:latest
bash
docker run --user root -p 8001:8000 --name uni-api -dit \
-e CONFIG_URL=http://file_url/api.yaml \
yym68686/uni-api:latest

Both pass --user root, and both publish 8001 on the host against 8000 in the container. Nothing in either command restricts access to the published port, which matters for a gateway whose whole job is holding credentials.

Cooling a channel is what makes retry safe

Two mechanisms handle a failing upstream, and they are separate. Automatic retry means that when an API channel's response fails, uni-api moves on to the next channel. Channel cooling means the failing channel is excluded and cooled for a period, with requests to it stopped, and then automatically restored when the cooling period ends, until it fails again and is cooled once more.

Retry alone is a bad default in front of a rate-limited provider: if a channel is failing because the key is exhausted or the account is throttled, retrying immediately just burns the quota and extends the failure. Cooling converts a dead channel into a temporarily excluded one, and the restore-on-expiry behaviour means a provider that recovers is picked up again without restarting the container.

The same idea appears at a finer grain in the Codex integration. When you add a provider with engine: codex, you can list several account credentials, each given as a comma-separated pair of account_id and refresh_token, and uni-api mints and refreshes the access token itself. When an account runs out of quota, that token is cooled down and the next one is used, with a default cooldown of six hours that you can change through api_key_quota_cooldown_period. That is per-credential cooling rather than per-channel, and six hours is a long time to be locked out of a single account by default.

There is also a moderation step. OpenAI moderation can review user messages and return an error when content is flagged, described as a way to reduce the risk of the backend API being banned by providers. That protects the gateway's own upstream accounts rather than the user.

Eight OpenAI-shaped endpoints behind one port

The gateway exposes the standard OpenAI surface rather than only chat completions. The listed paths are /v1/chat/completions, /v1/responses, /v1/images/generations, /v1/embeddings, /v1/audio/transcriptions, /v1/audio/speech, /v1/moderations and /v1/models.

/v1/responses is the newer of these and has its own integration path. The Codex section describes pointing a Codex CLI or OpenAI Responses API client at uni-api, supplying a uni-api api_keys value as the client credential, and adding a provider with engine: codex to handle the account credentials. The documentation states that this engine setting forces a particular store or storage behaviour automatically, though that sentence is cut off in the text available here.

Beyond the OpenAI shape, the gateway converts several non-OpenAI providers into it. Anthropic, Gemini, Vertex AI, Azure, AWS, xai, Cohere, Groq, Cloudflare and 0-0.pro are all supported, with Vertex able to serve both Claude and Gemini through one provider. Native tool use function calls and native image recognition are supported for OpenAI, Anthropic, Gemini, Vertex, Azure, AWS and xai, meaning those are translated rather than approximated through the chat interface. OpenAI Dalle-3 image generation is called out separately.

The 0-0.pro entry is worth noticing: it is both a supported backend and the project's own homepage, so the vendor you may be routing through is the vendor publishing this.

The example configuration contains credential-shaped strings

The minimal api.yaml in the documentation is short, a providers list with a name, a base_url and an API key, plus an api_keys list for the downstream credential. The advanced version adds the optional model list and the rename syntax described earlier.

What is worth flagging is the shape of the example values. Both blocks carry keys written in the conventional prefixed format for two different providers, and they are presented as things to copy. Even if they are placeholders, a configuration example that looks exactly like a working credential is a bad default, because that file is the one most likely to be edited in place rather than rewritten. The upstream keys belong in a file that is never committed, and the downstream api_keys entry is the one a client will actually hold.

So treat the example as a shape to copy and not as a set of values to keep. Replace every credential before the first run, and keep the file out of version control, since it holds both sides of the proxy.

The key names are deliberately plain, provider and base_url and model inside a providers entry, and api_keys at the top level, rather than the prefixed environment variable style some gateways use. Since the YAML file is the documented interface, that is what you configure against, and it also means there is no environment variable escape hatch for the provider list described in the documentation. The two documented environment variables are CONFIG_URL and the port mapping.

The container is Rust now, and versioned separately

The most consequential detail is not in the README. The repository is recorded as primarily Python, with release tags in the v1.7.x line at 1.7.276, 1.7.275 and 1.7.274, all three published within about fifteen minutes on 2026-09-12. But Cargo.toml declares a package named uni-api-native at version 0.1.91 with rust-version 1.84, producing a binary called uni-api-front from src/main.rs. There is a .python-version file and a .cargo directory in the tree, so both toolchains are represented in the repository while the published artefact is the Rust one.

The Dockerfile confirms the direction. It is a four-stage cargo-chef build pinned by digest at tag 0.1.71-rust-1.84-bullseye, copies the manifest and lockfile first so dependency compilation is cached separately from source, and the final stage is debian:bookworm-slim with ca-certificates and the compiled uni-api-front binary. There is no Python interpreter in the runtime image. So a v1.7.276 tag and a 0.1.91 crate version describe two different things, and knowing which one you are running matters when you compare behaviour.

The dependency pinning is unusually disciplined and has one exception. A comment explains that exact versions keep the production Rust 1.84 toolchain reproducible and that the resolver config selects only transitive releases whose declared MSRV fits. Twenty-four of the twenty-five direct dependencies use the exact equals form, axum at 0.8.4 and tokio at 1.47.1 among them. The exception is memchr, written as a floating caret range rather than an exact pin, which contradicts the stated rule in the one place you would not notice.

The compose file configures both methods at once

docker-compose.yml in the repository root sets a CONFIG_URL environment variable and mounts ./api.yaml into the container at the same time:

yaml
services:
  uni-api:
    container_name: uni-api
    image: yym68686/uni-api:latest
    environment:
      - CONFIG_URL=http://file_url/api.yaml
    ports:
      - 8001:8000
    volumes:
      - ./api.yaml:/home/api.yaml
      - ./uniapi_db:/home/data

The two startup methods described in the documentation are alternatives, so running this file as shipped means the gateway will try to download the configuration from a placeholder URL at startup, and your mounted api.yaml may never be read. Edit one out before using it. The placeholder is not a subtle value: it is the literal string file_url.

The compose file is more informative than the Dockerfile about runtime state, because it mounts a named volume at /home/data, which is the working directory the Dockerfile sets and therefore where the gateway keeps its database. The Cargo manifest backs that up: rusqlite is a dependency with the bundled feature, and tokio-postgres is present alongside it, so both an embedded SQLite store and a Postgres path exist in the binary. Which one is in use is decided by configuration rather than by the image.

The release profile is tuned for a small, fast binary, with fat link time optimisation, a single codegen unit, opt-level 3, panic set to abort and symbols stripped. The runtime environment also carries glibc malloc tuning, MALLOC_ARENA_MAX at 2 with 131072 thresholds for mmap and trim, which is the usual response to allocator behaviour under many concurrent connections.

Editorial conclusion

uni-api fits a single user with several provider keys who wants OpenAI-shaped endpoints without operating a panel, and the channel cooling plus automatic retry are the parts that make a multi-provider setup survive a provider outage. Two things to check before you point real traffic at it. The container runs as root and the shipped compose file configures both startup methods at once, with a placeholder CONFIG_URL that will be fetched before your mounted api.yaml is read, so fix that before starting. And note the runtime is now a Rust binary with no Python in the image, so if you are following the Python release tags, check which artefact you actually pulled.

Frequently asked questions

How do I start uni-api?

A configuration file is required, named api.yaml, and there are two ways to supply it: mount it into the container at /home/api.yaml, or set a CONFIG_URL environment variable holding its address so it is downloaded at startup. The image is then run with docker, publishing 8001 against the container's 8000.

Does uni-api need a web interface to configure providers?

No. There is no front end, and providers are configured entirely in the YAML file with a provider name, a base URL and an API key. Keys your clients present are configured separately in the api_keys list.

Which load balancing modes does uni-api support and are they on by default?

Three. Channel-level weighted balancing and channel-level sequential round_robin are both off by default and need channel weights or SCHEDULING_ALGORITHM set to round_robin. Automatic API key-level round-robin across several keys inside one channel needs no extra configuration.

What happens when a uni-api provider channel fails?

uni-api retries on the next channel, and cools the failing channel down, excluding it for a period and then restoring it automatically when the cooling period ends. Per-credential cooldown also applies to Codex accounts, defaulting to six hours and configurable with api_key_quota_cooldown_period.

Is uni-api written in Python or Rust?

The repository is recorded as primarily Python and the release tags run in the v1.7.x line, but the published container is a Rust binary built from a package named uni-api-native at 0.1.91, and the runtime image contains no Python interpreter.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. yym68686/uni-api on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/yym68686-uni-api.svg)](https://hysenlabs.com/projects/yym68686-uni-api)