# GPT-Load: a self-hosted AI gateway for multi-channel, multi-credential setups

> GPT-Load puts API-key channels and subscription accounts behind one base URL, with scheduling, failover and usage accounting in a bundled UI. The 2.0 line is still at release candidate, and it cannot read 1.x data.

**tbphp/gpt-load** — Self-hosted AI gateway for multi-channel, multi-credential setups — API keys and subscription accounts, scheduling, failover, request logs and usage. 自托管 AI 网关：多渠道多凭据统一接入，含密钥与订阅账号、调度容错、日志与用量。

- Repository: https://github.com/tbphp/gpt-load
- Website: https://www.gpt-load.com
- Stars: 7,009 · Forks: 778
- Language: Go
- License: MIT
- Published: 2026-09-09 · Updated: 2026-09-09 · Language: en
- Canonical page: https://hysenlabs.com/projects/tbphp-gpt-load

## The credential sprawl GPT-Load is built to absorb

The project targets a specific kind of mess. An application has grown to talk to several upstreams: an OpenAI-compatible endpoint, an Anthropic-style endpoint, a cloud platform, a relay, plus one or more subscription accounts such as Codex or Claude that authenticate through OAuth rather than a static key. Each of those has its own base URL, its own credential, its own rate limits and its own way of failing. The client code ends up holding all of it.

GPT-Load inverts that. The README states that your application needs one base URL and one AccessKey, while providers, accounts, credentials, models and routing policy live in the management UI. The unit of configuration is the AccessKey, which is scoped to groups and to client protocols, so a given application only sees the models and routes it was granted. The README also documents a read-only home for AccessKey sign-in, where the holder sees only its own groups, models, requests, usage and cost allowance. That is a meaningfully different posture from handing out the upstream keys themselves.

The intended user is an operator, not an end user. Someone who runs a small platform, an internal tool, or a team's shared LLM access and is tired of rotating keys by hand. If you have exactly one provider and one key, the gateway is overhead with no payoff.

## Channels, groups and AccessKeys: how the routing actually layers

The configuration model is three levels, and the README's initial setup names them in order. A channel represents an upstream service and holds one or more API keys; for subscription channels you complete an OAuth flow or import credentials. A group picks a channel and then defines the available models and the runtime policy. An AccessKey binds groups to the client protocols it may use.

Scheduling and failure handling sit between the group and the channel. The README lists multi-credential scheduling, configurable weights, retries, cooldown, blacklisting and session affinity as the mechanisms that reduce the impact of overloaded or failing credentials. The important word is affinity: without it, a multi-turn conversation could be routed to a different credential each turn, which breaks prompt caching and can break stateful upstream behaviour. The README does not spell out how affinity keys are derived, so that is something to read in the source rather than assume.

Protocol handling is native rather than translated. The scope table lists OpenAI Chat Completions at POST /v1/chat/completions, OpenAI Responses under /v1/responses and its resource paths, OpenAI Images, OpenAI Embeddings at POST /v1/embeddings, and Rerank at POST /v1/rerank. The README also states that clients keep their OpenAI, Anthropic or Gemini native interfaces. That means the gateway is not normalising everything into one dialect; it is preserving the caller's dialect and routing by it.

The dependency list in go.mod backs up the architecture claims. Gin for HTTP, GORM with drivers for SQLite, MySQL and PostgreSQL, gjson and sjson for surgical JSON edits on request and response bodies, brotli and klauspost/compress for compression, and gorilla/websocket for streaming paths. There is also an embedded component pulled from a separate module path, github.com/router-for-me/CLIProxyAPI/v7/gptload-embedded, which suggests the subscription-account handling is not written from scratch inside this repository. That is a real architectural dependency, and it is worth knowing before you fork.

## Installing GPT-Load with Docker Compose and making a first request

The README requires Docker and Docker Compose and gives a four-command quick start. Clone shallow, copy the environment template, bring the stack up.

```bash
git clone --depth 1 https://github.com/tbphp/gpt-load.git
cd gpt-load

cp .env.example .env
docker compose up -d
```

The compose file pulls ghcr.io/tbphp/gpt-load:2 and publishes the service on ${PORT:-3001}, bound to ${BIND_ADDRESS:-${HOST:-127.0.0.1}}. With the default .env, HOST is 127.0.0.1, so the port is reachable only from the host itself. The compose file also sets HOST: 0.0.0.0 inside the container, which is the correct split: loopback on the host, all interfaces inside the container network.

Confirm the service is healthy before touching the UI. The README gives this check:

```bash
curl --fail http://127.0.0.1:3001/health
```

A successful call returns without error. The compose healthcheck uses the same endpoint via wget, with a 40 second start period and 3 retries at 30 second intervals, so a container that reports healthy has already answered /health at least once.

On first start the service generates a management key. The README tells you to read it and store it safely:

```bash
docker compose exec gpt-load sh -c 'cat /app/data/auth.key'
```

Open http://127.0.0.1:3001 and sign in with that value. Alternatively, set AUTH_KEY explicitly in .env before starting, which is the better choice if you are provisioning from a secrets manager rather than reading a file out of a container.

From there the README describes three configuration steps: add a channel with one or more API keys (or complete OAuth for a subscription channel), create a group that selects the channel and configures models and runtime policy, then create an AccessKey scoped to the groups and client protocols it may use. The generated AccessKey is what you hand to your application, replacing the upstream credentials.

## The OAuth callback ports and the one-instance-per-host limit

This is the constraint most likely to bite during installation, and the README calls it out directly. The Codex, Claude and Antigravity OAuth clients use fixed callback ports. The compose file publishes three of them: 1455, 54545 and 51121. Because the ports are fixed by the upstream clients rather than chosen by GPT-Load, only one default Compose instance can run on a host at a time.

That rules out a few otherwise reasonable setups. You cannot run a staging and a production copy side by side on the same machine if both need to complete an OAuth flow. You cannot trivially run two projects that both embed GPT-Load on one host. The workaround is not documented in the README; the compose file exposes OAUTH_CALLBACK_BIND_ADDRESS, which lets you change the bind address, but the host-side port numbers are still 1455, 54545 and 51121, so changing the address does not remove the conflict.

The compose file also ties the publishing address to HOST. The README notes that setting HOST=0.0.0.0 also publishes those callback ports on all host interfaces. That is a security-relevant side effect of a setting most people change for unrelated reasons, and it deserves more than a parenthetical. If you set HOST=0.0.0.0 to reach the console from another machine, you have also exposed three OAuth callback listeners.

The second documented friction point is remote browser flows. When working over SSH or from a remote browser, the browser's localhost may not reach GPT-Load, so the README instructs you to paste the full callback URL into the authorization dialog to finish the flow. That works, but it means the OAuth step is not a click-through on a headless server.

## Storage, key material and what happens if you lose encryption.key

State lives in a database. The .env.example comment states that an empty DATABASE_DSN uses ${DATA_DIR}/gpt-load.db, which is a SQLite file, and that otherwise you set a SQLite, MySQL or PostgreSQL URL. In the container, DATA_DIR is /app/data and the compose file mounts the named volume gpt-load-data there, so the default deployment persists across container recreation.

The security model is file-based by default and the .env.example is explicit about the consequence: empty AUTH_KEY and ENCRYPTION_KEY values cause generated auth.key and encryption.key files to be written into DATA_DIR, and the file states that you should back up encryption.key with the database because losing it makes credentials unreadable. That is a hard boundary, not a caution. The upstream API keys and subscription credentials stored in the database are encrypted with that key material. A database dump alone is not a restorable backup.

This is also the argument for setting both values explicitly rather than letting them be generated. If AUTH_KEY and ENCRYPTION_KEY come from your environment, the recovery story is the same as every other secret you manage. If they come from files inside the data volume, your backup procedure has to include those files, and the README's quick start does not walk through that.

On the schema side, the project supports three backends through GORM, which is more choice than most self-hosted gateways offer. SQLite is the default and is fine for a single instance. The README does not document what happens if you run multiple GPT-Load replicas against one PostgreSQL database, so horizontal scaling is unverified territory rather than a supported mode.

## GPT-Load against a plain reverse proxy or LiteLLM

The obvious question is why not put nginx or Caddy in front of your providers and call it done. A reverse proxy can hold multiple upstreams and do round-robin, and it costs nothing to learn. What it cannot do is understand credentials as a pool. It has no concept of a key that is cooling down after a 429, no blacklisting after repeated failures, no session affinity keyed to a conversation, and no usage accounting per AccessKey. It also will not complete an OAuth flow on your behalf. If your problem is purely "route this path to this host", a reverse proxy is the right tool and GPT-Load is not.

The closer comparison is LiteLLM, which the README does not describe, so the honest difference is one of scope rather than feature parity. GPT-Load's stated design keeps client protocols native and routes by them, with the scope table listing OpenAI Chat Completions, Responses, Images, Embeddings and Rerank as separate entries. It is a Go binary with an embedded web UI, a single compose file and a choice of three SQL databases. The trade-off you accept is that the project is younger and currently shipping release candidates: the newest release is v2.0.0-rc.11, published 2026-09-09, with rc.10 and rc.9 in the two days before that. Three release candidates in three days is a fast iteration cadence, which is good for fixes and bad for anyone who wants to pin a version and forget about it.

There is also a structural alternative worth naming: keep the credentials in your application and write the retry and cooldown logic yourself. That is entirely reasonable for two providers. It stops being reasonable once you need per-application usage attribution, which is the feature GPT-Load's AccessKey model provides and hand-rolled retry code usually does not.

## Maintenance cost, licence and the 1.x migration wall

The repository is not archived and the last push was on 2026-09-09, so the project is being worked on. The release history supports that: v2.0.0-rc.11 landed on 2026-09-09, with rc.10 on 2026-09-08 and rc.9 on 2026-09-07. The practical consequence is that you should expect to move versions, and that pinning to a specific rc tag is the sane default rather than tracking the moving 2 tag that the compose file references.

The licence is MIT, and the repository carries a LICENSES/ directory and a THIRD_PARTY_NOTICES.md alongside the LICENSE file. MIT is permissive and imposes no copyleft obligation on your own code. It does not, however, resolve the terms of the upstream services you route through the gateway, and it does not cover the separate module the project embeds for subscription-account handling. If your organisation reviews third-party dependencies, the go.mod require block is the place to start, and the embedded CLIProxyAPI module is the entry most likely to need a second look. This is a description of what the repository contains, not legal advice.

The upgrade cost is dominated by one fact the README puts in a warning box before the quick start: if you are using 1.x, 2.0 cannot open, import or migrate 1.x data in place. There is no migration path in the documentation. For an existing 1.x deployment that means a parallel install and manual re-entry of channels, groups and AccessKeys, or staying on 1.x. For a new deployment it means nothing, which is why the 2.0 release candidate line is a reasonable starting point today and a poor upgrade target for an existing install.

## Conclusion

Adopt GPT-Load if you already juggle several upstream providers or subscription accounts and want one AccessKey per application, with cooldown and blacklisting handled by the gateway instead of your client code. Do not adopt it as a drop-in upgrade from 1.x: the README states 2.0 cannot open, import or migrate 1.x data in place, and the current release is v2.0.0-rc.11 from 2026-09-09. Before you commit, verify three things on your own host: that the fixed OAuth callback ports 1455, 54545 and 51121 are free, that you have a backup of encryption.key alongside the database, and that your chosen DATABASE_DSN points at a path the container can actually reach.

## FAQ

### What does GPT stand for?

The name is not an acronym in the README; GPT-Load is the project name, and it describes itself as a self-hosted AI gateway for multi-channel, multi-credential setups. No expansion is documented.

### Why won't GPT load?

The README does not document a troubleshooting path for a failed start. What it does give is a health endpoint: curl --fail http://127.0.0.1:3001/health, which the compose healthcheck also uses with a 40 second start period and 3 retries.

### Why is GPT loading so slow?

The README does not describe performance characteristics or latency. It does list retries, cooldown and blacklisting as scheduling mechanisms, all of which can add delay when credentials are failing, but no timings are given.

### What is the GPT file format?

GPT-Load does not define a file format of its own. Configuration is read from .env, and state is stored in a database: an empty DATABASE_DSN uses ${DATA_DIR}/gpt-load.db, otherwise a SQLite, MySQL or PostgreSQL URL.

### chatgpt loading slowly

The README does not discuss ChatGPT latency. For GPT-Load itself, the documented scheduling mechanisms are multi-credential scheduling, configurable weights, retries, cooldown, blacklisting and session affinity, and none of them carry published timings.

## Sources

- [License: MIT](https://github.com/tbphp/gpt-load/blob/main/LICENSE)
- [Project website](https://www.gpt-load.com)
- [README](https://github.com/tbphp/gpt-load/blob/main/README.md)
- [Releases](https://github.com/tbphp/gpt-load/releases)
- [tbphp/gpt-load on GitHub](https://github.com/tbphp/gpt-load)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/tbphp-gpt-load
