# CliRelay: a self-hosted gateway that puts Claude Code, Codex and Gemini CLI behind one endpoint

> CliRelay is a Go-based fork of CLIProxyAPI that turns existing AI CLI subscriptions and API keys into a single OpenAI, Claude, Gemini or Codex compatible endpoint, with a multi-tenant web console, request logging and spend quotas. It is built for teams that need to operate that traffic for more than one person.

**kittors/CliRelay** — Self-hosted AI gateway for coding CLIs — one OpenAI/Claude/Gemini/Codex-compatible endpoint, with a multi-tenant web console, request logs, and spend quotas.

- Repository: https://github.com/kittors/CliRelay
- Stars: 1,026 · Forks: 125
- Language: Go
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kittors-clirelay

## The problem CliRelay solves for teams running several coding CLIs

A single developer using Claude Code or the OpenAI Codex CLI configures one credential and moves on. That stops working when five people on a team each hold a subscription, each client speaks a slightly different protocol, and nobody can say which key burned the monthly budget. CliRelay takes those CLI subscriptions, OAuth credentials, API keys and compatible upstream services and exposes them as one managed API layer on port 8317. The README describes it as a unified proxy server for AI CLI tools, and the diagram in the repository shows coding tools on the left, CliRelay in the middle, and Gemini, OpenAI/Codex, Anthropic Claude, Qwen, iFlow, Antigravity, xAI, Vertex, Bedrock, OpenCode, ClinePass, Ollama and Amp on the right.

The audience is explicit in the README: the project is "built to be operated by more than one person." Tenants, users, roles and a resource.action permission model such as governance.tenants, models.write and providers.test decide which pages and actions an account can reach, and security-sensitive changes land in an audit log. Portal accounts let an end user hold several API keys under one identity and read their own usage without an administrator in the loop. If you are one person with one key, none of that machinery pays for itself.

## How the proxy, routing groups and failover actually fit together

CliRelay is a Go service (module github.com/router-for-me/CLIProxyAPI/v6, go 1.26.0) that fronts upstream providers and speaks the OpenAI Chat Completions protocol on its own surface. The README states that any upstream that speaks that protocol works, which is why the provider list can be long without a bespoke adapter for each one. Incoming requests are matched to channels: round-robin or fill-first scheduling spreads load across multiple API keys for the same provider, and channels can be bound into groups so an API key is restricted to the groups it is allowed to use. Custom path namespaces let a team or workload get its own URL prefix on the same process.

Failover is the part worth understanding before you rely on it. The README says CliRelay "automatically switches to backup channels when quotas are exhausted or errors occur." That is channel-level failover, not request-level retry semantics, and the README does not document how many attempts a single request gets or which error classes trigger the switch. Treat the grouping configuration as the real control surface: if two channels for the same provider sit in one group, you have redundancy; if they do not, you do not.

The runtime data stack is PostgreSQL 15+ and Redis 7+, with Ent ORM generating the schema. PostgreSQL is the source of truth; Redis holds cache, locks, limits, queues and rebuildable state. That split matters operationally. Losing Redis degrades limits and caching, while losing PostgreSQL loses the logs and quota state that make the gateway worth running.

## Installing CliRelay with Docker Compose and making a first request

The repository ships a Dockerfile, a docker-compose.yml and an install.sh, and the compose file pulls ghcr.io/kittors/clirelay:latest by default. The compose stack defines two services: clirelay-init, which runs the clirelay-init-env command to generate the deployment .env, and cli-proxy-api, which is the gateway itself. The Dockerfile builds the management panel from the separate codeProxy repository at build time, pinned through FRONTEND_REPOSITORY and FRONTEND_REF, so a local build compiles the current panel rather than relying on a checked-out frontend directory.

Start the stack from the repository root:

```bash
docker compose up -d
```

The compose file maps 8317:8317 for the gateway and binds several other ports to 127.0.0.1 only. The README describes one endpoint at http://localhost:8317 fronting every configured provider. The container sets AUTH_PATH to /CLIProxyAPI/auths by default; that is where credentials live, and it is the directory to mount if you want auth state to survive a container rebuild.

Configuration is YAML. The repository root contains config.example.yaml, and the container reads a generated .env through its entrypoint, which sources /clirelay-deploy/.env before exec-ing docker-entrypoint.sh. The compose file exposes CLIRELAY_LOCALE, CLIRELAY_UPDATE_CHANNEL, CLIRELAY_LANG, CLIRELAY_LANGUAGE and LC_ALL as environment variables, with CLIRELAY_LOCALE defaulting to zh. The README does not spell out a copy step for the example config, so check the docs site at help.router-for.me for the config file name the binary expects before you edit it.

Once the service is up, point a client at the gateway. The README gives http://localhost:8317 as the base URL, and any OpenAI-compatible client can be aimed at it. The README does not publish a full request example, so verify the exact path your client uses against the docs site before assuming a route is live. A request that reaches CliRelay should appear in the request log with timestamp, model, token counts (in, out, reasoning, cache), latency, status and source channel.

## What gets logged, and why the storage choices are a trade-off

Every API request is logged to PostgreSQL with model, token counts, latency, status and source channel, and the README states that full request and response message content is captured in compressed PostgreSQL storage. Content and metadata have separate retention settings. That is a genuinely useful capability for debugging a failing channel or attributing spend, and it is also the part of CliRelay most likely to surprise you in production: prompt bodies are stored, compressed, in the same database that serves quota enforcement.

The README describes pre-computed analytics dashboards (daily trends, model distribution, hourly heatmaps, per-key statistics), a health score engine that produces a 0 to 100 value from success rate, latency, active channels and error patterns, and live system stats (CPU, memory, goroutines, network I/O, DB size) streamed over WebSocket. All of that is downstream of PostgreSQL. If the database grows faster than your retention policy removes content, query performance and disk usage are your problem, not the gateway's.

Quotas are enforced per key: max token and request limits, with period resets for a chosen window. The README does not specify what happens to in-flight requests when a quota is crossed mid-stream, so test that boundary yourself rather than assuming a clean cutoff.

## Where CliRelay is the wrong tool

CliRelay is a fork, and the README says so directly: it is a heavily enhanced fork of CLIProxyAPI, rebuilt with a management layer, web control panel hosting and a terminal TUI. Forks inherit the upstream's protocol assumptions and diverge over time. If you need a thin proxy that mirrors upstream releases closely, tracking the parent project is the simpler path; if you need tenants, audit logs and quotas, CliRelay is the one that has them.

The heavier constraint is the dependency footprint. PostgreSQL 15+ and Redis 7+ are mandatory for the runtime data stack, not optional. A developer who wants to route a laptop's CLI traffic through a local process now runs two stateful services, a container with several bound ports, and a database whose retention settings need a decision. That is a poor fit for a single-user setup, for an air-gapped machine where you cannot pull ghcr.io images, or for anyone unwilling to store prompt bodies. The README does not document rollback or downgrade steps between releases, and it does not state what happens when Redis is unavailable mid-request, so a deployment without a tested recovery path is a deployment you have not finished.

## CliRelay compared with running CLIProxyAPI directly

The honest comparison is with CLIProxyAPI, the project CliRelay forks. Upstream gives you the proxy engine: one endpoint in front of multiple providers, with the protocol translation that lets OpenAI, Claude, Gemini and Codex style clients share a backend. It does not give you tenants, roles, portal accounts, per-key quotas, audit logs or the /manage web hosting that CliRelay adds, and the README frames those as the reason the fork exists.

So the difference is not the proxying, it is the operating layer. If you are the only operator and you can read a config file, CLIProxyAPI's smaller surface is easier to reason about. If more than one person needs an account, a quota that resets on a schedule, and a log that answers "who spent this," you would be rebuilding CliRelay's governance features on top of the upstream engine. The cost of choosing CliRelay is the PostgreSQL and Redis dependency and a management panel built from a second repository at image build time; the cost of choosing upstream is building the tenancy yourself.

## Maintenance, licence and upgrade expectations

CliRelay is MIT licensed, and the LICENSE file sits at the repository root. MIT is permissive: you can run it commercially, modify it and redistribute it, provided the copyright notice and permission notice travel with copies. That is a statement about the licence text, not legal advice for your situation.

The project is not archived. The last push to main was on 2026-09-09, and the most recent releases are v0.5.1 (a Codex native web search fix) on 2026-09-08, v0.5.0 (channel scheduling and protocol fixes) on 2026-09-07, and v0.4.34 (support for the GPT-6 Astra series and long-context deep reasoning) on 2026-09-05. Releases are frequent and the changelog is in Chinese, which is worth knowing if your team reads release notes in English only.

Upgrade cost comes from three places. The management panel is cloned from the codeProxy repository at image build time and pinned by FRONTEND_REF and FRONTEND_COMMIT, so a rebuild can pull panel changes you did not review. The compose file sets CLIRELAY_UPDATE_CHANNEL to main by default and the repository contains an updater module, so an online update path exists. And the Ent-generated schema means database migrations ride along with releases. The README does not document rollback, so pin CLI_PROXY_IMAGE to a specific tag rather than accepting latest if you need to go back.

## Conclusion

Adopt CliRelay if you already run Claude Code, Codex or Gemini CLI accounts for several people and want one endpoint, per-key quotas and an audit trail instead of passing credentials around. Do not adopt it if you are a single developer with one API key; a direct client configuration is less work than PostgreSQL, Redis and a Compose stack. Before committing, verify that your upstream providers are on the supported list, confirm the Go 1.26 toolchain and PostgreSQL 15+ and Redis 7+ requirements fit your infrastructure, and check the CHANGELOG for the migration path between releases, since the README does not document rollback.

## FAQ

### What is CliRelay?

CliRelay is a self-hosted AI gateway, described in its README as a unified proxy server for AI CLI tools. It puts Claude Code, Gemini CLI, OpenAI Codex and other coding tools behind one endpoint on port 8317, and adds a multi-tenant web console, request logs and spend quotas.

### How do I install CliRelay?

The repository ships a Dockerfile and a docker-compose.yml, and the compose file pulls ghcr.io/kittors/clirelay:latest by default. Running docker compose up -d from the repository root starts the clirelay-init service, which generates the deployment .env, and the cli-proxy-api service, which is the gateway.

### What database and cache does CliRelay require?

The README states the runtime data stack is PostgreSQL 15+ and Redis 7+ with Ent ORM. PostgreSQL is the source of truth for runtime data, while Redis is used for cache, locks, limits, queues and rebuildable state.

### Is CliRelay a fork of another project?

Yes. The README describes CliRelay as a heavily enhanced fork of CLIProxyAPI, rebuilt with a production-grade management layer, web control panel hosting and a terminal TUI for day-2 operations.

### What licence does CliRelay use?

CliRelay is MIT licensed, and the LICENSE file is at the repository root. The licence permits commercial use, modification and redistribution as long as the copyright and permission notices are kept with copies.

## Sources

- [Issues](https://github.com/kittors/CliRelay/issues)
- [kittors/CliRelay on GitHub](https://github.com/kittors/CliRelay)
- [License: MIT](https://github.com/kittors/CliRelay/blob/main/LICENSE)
- [README](https://github.com/kittors/CliRelay/blob/main/README.md)
- [Releases](https://github.com/kittors/CliRelay/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kittors-clirelay
