# The AI gateway that converts four protocols and admits the mapping is lossy

> QuantumNous/new-api is a self-hosted Go gateway that puts one API in front of many model providers, with channels, routing, quotas, and a web console behind it. The interesting parts are the ones the project states plainly: conversion between protocols does not map exactly, the default request timeout is no limit, and the install paths bind differently.

**QuantumNous/new-api** — GitHub describes it as A unified AI model hub for aggregation & distribution. It supports cross-converting various LLMs into OpenAI-compatible, Claude-compatible, or Gemini-compatible formats. A centralized gateway for personal and enterprise model management. 🍥. The repository metadata lists Go as its primary language. The metadata lists the AGPL-3.0 license. This article stays within the project description and details documented in the GitHub repository README.

- Repository: https://github.com/QuantumNous/new-api
- Website: https://www.newapi.ai
- Stars: 48,931 · Forks: 11,748
- Language: Go
- License: AGPL-3.0
- Published: 2026-08-13 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/quantumnous-new-api

## RelayKit converts between four text protocols, and the mapping is lossy

One sentence does more work than the whole capability table: available features depend on the channel, the upstream model, and the conversion path, and protocol-specific tools and fields may not map exactly. RelayKit is the component that converts requests, responses, and streaming output between the four text protocols, OpenAI Chat and Responses, Anthropic Messages, and Gemini. The endpoint table shows the three shapes a client can speak, and the Gemini routes look nothing like the others, `POST /v1beta/models/{model}:generateContent` and its streaming sibling, rather than a chat completions path. Because the gateway normalizes these, a client written against one shape keeps working against an upstream built for another, and that is the entire point. The cost is that the normalized surface is only the intersection. A tool definition, a reasoning field, or a multimodal input that exists on the protocol you send may have no representation on the protocol you reach, and you find that out at runtime, per model, not at configuration time.

## The one-line quick start binds to localhost while the compose file binds to everything

Two install paths ship in the same repository, and they differ in the way that matters most. The quick start command publishes on 127.0.0.1:3000, so the console and the API are reachable from that machine only:

```bash
mkdir -p data
docker run --name new-api -d --restart unless-stopped \
  -p 127.0.0.1:3000:3000 \
  -e TZ=Asia/Shanghai \
  -v "$(pwd)/data:/data" \
  calciumion/new-api:latest
```

Open http://localhost:3000 afterwards and the setup wizard creates the administrator account, while the data directory keeps the SQLite database across container replacements. The compose service publishes 3000:3000 instead, which binds every interface the host has, so the console, the wizard, and the API are all on the network before you have configured anything. That path also runs a bigger stack: PostgreSQL and Redis, with an optional separate log database, rather than the SQLite file the single container uses. Nothing forces you to read the port mapping, and that is the line most people copy without looking.

## The compose file ships working default passwords and a line-numbered MySQL switch

The compose file opens with a warning to change all default passwords before deploying to production, and then supplies working defaults: the PostgreSQL connection string is root with the password 123456, and the Redis connection string carries the same one. The console creates its administrator through a setup wizard on first run, so the account is not the weak point here, but those two credentials are, and they sit in the file as literals rather than being generated. The MySQL instructions are worth reading twice. They are four steps pinned to line numbers: comment out the postgres service and its SQL_DSN on line 15, uncomment the mysql service and its SQL_DSN on line 16, uncomment mysql in depends_on on line 28, and uncomment mysql_data in the volumes section on line 64. Documentation anchored to line numbers stops being correct the first time somebody reformats the file, and following it after that means guessing which line moved.

## RELAY_TIMEOUT defaults to zero, so nothing bounds a hung upstream

The timeout defaults deserve reading before you point this at a production provider. RELAY_TIMEOUT is 0, which the example environment file defines as no limit, and it governs the request as a whole. RELAY_IDLE_CONN_TIMEOUT defaults to 90 seconds for the HTTP client idle connection. RELAY_RESPONSE_HEADER_TIMEOUT defaults to 1800 seconds and constrains only the wait for response headers, with an explicit note that streaming after the headers arrive is unaffected, and a second note that a non-streaming request usually has to wait for the upstream to finish generating before headers come back, so the value needs margin. STREAMING_TIMEOUT defaults to 300 seconds, and the stated remedy for an empty completion is to raise it. Together the defaults mean a stalled upstream can hold a request open for as long as it likes, and raising the header timeout changes nothing for a stream that already started. Every one of these ships commented out in .env.example, so they are decisions, not defaults you inherited.

## Task plugins bring their own routes and their own refund rules

Asynchronous work is not a fixed set of endpoints. Task plugins are JavaScript, the routes are `POST /v1/tasks/{pluginKey}` and `GET /v1/tasks/{taskId}`, plus whatever routes each plugin declares for itself. That last clause is the one to hold onto: both the surface you must build and the surface you must support come from the plugin rather than the gateway, so an image or video capability is exactly as good as the plugin somebody wrote for it. Failure handling is unusually deliberate. TASK_TIMEOUT_MINUTES defaults to 1440, a hard timeout measured from submission time, after which an unfinished task is marked failed and refunded. TASK_POLL_MAX_FAILURES defaults to 20 consecutive polling failures, counting upstream 429, 5xx, 401, and 403 responses, network errors, and responses the poller cannot recognize; reaching it marks the task failed and refunds it, and a single successful poll resets the counter to zero. Set the timeout to 0 and the hard stop is disabled.

## Log retention is off by default, and the log database is a separate choice

Usage and audit logging can go somewhere other than the main database. LOG_SQL_DSN is optional and can point at a separate PostgreSQL database for logs, and the compose file offers a ClickHouse connection string in the same slot, with the note that using it means uncommenting the clickhouse service and adding it to depends_on. Retention is then entirely your decision, and the default keeps nothing: LOG_SQL_CLICKHOUSE_TTL_DAYS unset or 0 disables automatic deletion, while a value such as 30 keeps 30 days. The main database has its own tuning in the same file, with SQL_MAX_IDLE_CONNS at 100, SQL_MAX_OPEN_CONNS at 1000, SQL_MAX_LIFETIME at 60 seconds, and SQL_SLOW_THRESHOLD_MS at 200, where 0 disables slow query logging and a value outside the 0 to 3600000 range falls back to 200. ERROR_LOG_ENABLED turns on error logging. The gap is that nothing in the configuration warns you when a log table starts growing.

## Two image registries, one module path, and three Go versions

The names drift in three directions, and it helps to know which is which. The Go module is github.com/QuantumNous/new-api. The container images are published as calciumion/new-api, which is also where the license badge in the README points, at Calcium-Ion rather than QuantumNous, and a Docker Hub link and an AtomGit mirror both appear in the badge row. Then there are three Go versions: go.mod declares go 1.25.1, the build stage uses golang:1.26.1-alpine, and a Heroku build comment still reads goVersion go1.18. None of that blocks a build, since the toolchain image is newer than the module requires, but it does mean the go.mod line is not what tells you the minimum. The build has one wrinkle worth knowing before you try to reproduce it: relaykit is a local submodule reached through a replace directive, and its go.mod has to be copied into the build before go mod download runs, or the main module graph will not resolve.

## Every recent tag is a release candidate while the quick start pulls latest

The three newest tags are v1.0.0-rc.38, v1.0.0-rc.39, and v1.0.0-rc.40, released within three days of each other, and their titles name the changes: Responses WebSocket and Request Policies, then Task Plugins and Quota Settings. So the 1.0 line is still in release candidate form while the quick start command pulls calciumion/new-api:latest, a floating tag that will follow those candidates on its own. The project states plainly that its README describes the current source tree and that you should check the release notes for the version you deploy, which is the right way to read something moving at that pace. Two related details follow from the same file set. The README ships in five languages, English, Simplified Chinese, Traditional Chinese, French, and Japanese, while the web console is listed as available in seven, adding Russian and Vietnamese, so documentation and interface do not cover the same set.

## Conclusion

Choose new-api if you need one endpoint in front of several model providers and you want per-team keys, quotas, and a playground to manage them, and do not choose it if you need every protocol feature to survive the round trip or a request timeout that fires by default. Before deploying, change the database passwords the compose file ships with, decide whether the service should be reachable off localhost, set RELAY_RESPONSE_HEADER_TIMEOUT and STREAMING_TIMEOUT to values your upstreams can meet, turn on log retention, and pin an image tag rather than tracking latest, since the 1.0 line is still at release candidates.

## FAQ

### What is new-api in QuantumNous/new-api used for?

It is a self-hosted AI gateway written in Go. You connect upstream model services such as OpenAI, Anthropic, Google Gemini, Azure OpenAI, AWS Bedrock, Vertex AI, DeepSeek, and Qwen, and it presents one API to your own clients while you manage routing, access, usage, and costs in one place.

### How do I use new-api for the first time?

Create a data directory, run the calciumion/new-api image with that directory mounted, then open port 3000 on localhost and complete the setup wizard to create the administrator account. After that, add a channel with an upstream key and a group, set model pricing and quota, and create an API key in the console.

### Does new-api serve Anthropic and Gemini clients, or only OpenAI ones?

It serves all three shapes: POST /v1/chat/completions and POST /v1/responses for OpenAI, POST /v1/messages for Anthropic Messages, and the generateContent and streamGenerateContent routes for Gemini. Conversion between them runs through RelayKit, and the project warns that protocol-specific tools and fields may not map exactly.

### What database does new-api use by default?

The single-container quick start uses SQLite kept in the mounted data directory. The compose file uses PostgreSQL with Redis, offers a MySQL alternative that takes four commented lines to enable, and can send logs to a separate PostgreSQL or ClickHouse database where retention stays off until you set a TTL.

## Sources

- [Official documentation](https://www.newapi.ai)
- [Official README](https://github.com/QuantumNous/new-api#readme)
- [Project repository](https://github.com/QuantumNous/new-api)
- [Release notes](https://github.com/QuantumNous/new-api/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/quantumnous-new-api
