# MAX API is a Go model gateway that holds the governance boundary between your apps, agents, and upstream model vendors

> A unified model gateway and control plane, in Go, that sits between applications and upstream AI vendors and keeps permissions, cost, evidence, and step-up verification in one place. Excellent for multi-vendor routing and recoverable billing; the autonomy layer is still a stated blueprint, and SQLite is only for local use.

**MAX-API-Next/MAX-API** — MAX API Next 社区是由来自科研机构和高校的 AGI 爱好者组织发起、维护和运营的研究驱动型技术社区，聚焦 AI Models 与 Agents 治理方向，致力于建设面向 AGI 应用时代的开放基础设施。社区关注多模型与多平台接入、国产模型持续适配、AgentOps 工程实践、成本审计、权限与安全边界、私有化部署和长期运营优化，目标是把模型服务、Agent 应用、用户组织和上游平台之间的共性治理问题沉淀为稳定、可复用、可持续演进的工程能力。

- Repository: https://github.com/MAX-API-Next/MAX-API
- Website: https://max-api.com
- Stars: 406 · Forks: 24
- Language: Go
- License: AGPL-3.0
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/max-api-next-max-api

## Docker pins MAX API by digest and binds it to loopback

The quick-start path assumes nothing but Docker and defaults to a local SQLite database. What makes it interesting is that the image is pinned by content digest rather than a floating tag, so the label `latest` is decorative and the exact build you tested is the one that runs:

```bash
MAX_API_IMAGE=cscitechtop/max-api:latest@sha256:006d5d86887a261baab4d71ec3797d429e3771a4836e5899734aee0e7f66f2ab

docker pull "$MAX_API_IMAGE"

docker run --name max-api -d --restart always -p 127.0.0.1:3000:3000 -e TZ=Asia/Shanghai -v ./data:/data "$MAX_API_IMAGE"
```

Two choices in that one-liner carry real intent. The port mapping binds to `127.0.0.1:3000`, not `0.0.0.0:3000`, so a fresh container is loopback-only and someone has to deliberately expose it. The volume mounts `./data` so the SQLite file and any local state survive a container replacement. Once it is up you visit `http://localhost:3000` and do exactly three things: create or confirm the admin account, add an upstream channel with a legitimately authorized key, then mint an access token and point your application's Base URL at MAX API. Until that third step the gateway has nothing to serve, so the empty first-run state is by design.

## SQLite for a first run, MySQL 8.4 or PostgreSQL 14 once real traffic arrives

MAX API draws a hard line between local and production databases, and it is one of the clearer boundaries in the README. SQLite is for local experience, development, and small-scale tests. Production should run MySQL that is still inside its vendor security support window (8.4 LTS is the suggestion) or PostgreSQL 14 or newer, alongside Redis, HTTPS, and a backup and recovery plan. The compatibility floor is deliberately lower than the recommendation: the code still supports MySQL 5.7.8 and PostgreSQL 9.6, but the project says plainly not to run those in production. So the database is a place where "it starts" and "it is supported" are different bars.

The compose file shows what a fuller deployment looks like, and it ships with placeholder secrets you must change:

```yaml
- SQL_DSN=postgresql://root:123456@postgres:5432/max-api
- REDIS_CONN_STRING=redis://:123456@redis:6379
```

Both default passwords are `123456`, flagged in the file as needing a change before production. One more operational switch deserves attention: `QUOTA_DATA_AGGREGATE_MIGRATION_ENABLED` is off by default and is meant to be turned on by hand, on a single node, during a low-traffic window, only when you need to consolidate an old `quota_data` table, merge duplicate aggregate groups, and build a unique index on the aggregate key. That is a maintenance migration you run deliberately, not something the gateway does for you at boot.

## Billing pre-deducts then settles, and every step is idempotent

Recoverable billing is the mechanism the project leans on hardest, and it is built around a specific failure: an abnormal retry, a long-running task, or a crash window causing a duplicate charge, a wrong refund, or an untraceable result. The flow pre-deducts from the quota when a request starts, records the charge idempotently, then performs a final settlement once the real usage is known, refunding on failure. For asynchronous tasks the system polls and keeps an explicit `pending` or `manual` state so nothing is silently lost.

The stated principle behind it is One Billing Truth: production billing, quotas, and settlement keep a single source of truth, and no second ledger is built for an agent or a plugin. That matters when agents generate requests at machine speed, because a naive per-request charge without idempotency is exactly where double-billing creeps in. Cost accounting is also broken out per user, token, model, channel, and group, so you can attribute spend rather than receiving one blended invoice. In a gateway whose whole pitch is that vendors supply the models, agent frameworks orchestrate the business, and MAX API holds the access and governance boundary, this pre-deduct-then-settle loop is the load-bearing piece. Get it wrong and every other governance feature sits on top of incorrect money.

## Model switching happens in channel config, not in your code

The pitch is a comparison table, and the interesting column is the one about switching models. Wiring an application straight to several vendors means maintaining separate SDKs, protocols, auth, and error formats, and switching a model means editing code, keys, and deployment config. Routing through MAX API moves all of that into channels: weights, priority, groups, model mapping, failure retry, and cross-vendor switching. Availability stops being each application's problem and becomes a central setting where you configure retry and failover once.

Protocol coverage is where a gateway earns its keep, and the supported surface spans OpenAI Compatible, Responses, Claude Messages, Gemini, Realtime, plus multimodal and asynchronous task interfaces. There is also a compatibility layer for reasoning and tool context: Reasoning Effort, tool definitions, Tool Call, correlation between a tool call and its response, and multi-turn context transformation, all intended to reduce the loss of reasoning information and tool-call semantics when you move between vendors. A few environment knobs tie into the upstream layer directly, such as `GEMINI_VISION_MAX_IMAGE_NUM` for how many images a Gemini request may carry and `COHERE_SAFETY_SETTING` for Cohere safety level. So the abstraction is not a lowest-common-denominator proxy; it is a translation layer, which is also why the upstream extension points (protocol adaptation, path and header overrides, model discovery, task-status mapping) matter to anyone running a nonstandard endpoint.

## Sensitive actions demand a scope-bound re-check and revoke old sessions

Security and organization governance is where MAX API gets specific rather than aspirational. High-risk operations use scope-bound step-up verification: Passkey, 2FA, Telegram, API token, and session revocation each require a fresh check tied to that particular action's scope, so completing a login does not silently authorize a later sensitive change. Old sessions are revoked through a `session_generation` counter, which gives you a cheap way to invalidate every session issued before a given moment. The Go dependencies back this up concretely, with go-webauthn for Passkey and an OTP library for two-factor codes.

Above that sits a stated principle, Governance before Autonomy: identity, permissions, budgets, approval, audit, and rollback must exist before autonomous capability is granted. The fully autonomous layer is described as a long-term blueprint rather than a shipped feature, listing Policy, Budget, Approval, Shadow, Canary, Rollback, and an isolated coding workspace as the intended pieces. Read that as a roadmap, not a toggle you flip today. What you can rely on now is the enforced half: per-user and per-token budgets and quotas, admin permissions, audit trails, and rate limits on sensitive operations, backed by the principle that evidence comes before action. Diagnostics, suggestions, and any automatic move are supposed to rest on traceable logs, errors, retries, and settlement evidence rather than on inference from a prompt.

## The image bakes a React build and a Go binary into a slim Debian layer

One `docker pull` hides a three-stage build, and reading the Dockerfile tells you what the project actually is. A Bun stage compiles the React frontend under `web/`, stamping the version into the bundle from the repository's `VERSION` file. A Go stage on Alpine downloads modules, builds with `CGO_ENABLED=0` and `GOEXPERIMENT=greenteagc`, and injects the same version into the binary at link time. The final image is a slim Debian layer that copies in only the compiled `max-api` binary, the built frontend, and the license files, then exposes port 3000 and runs the binary from `/data`. Shipping media codecs, a web framework, Redis, JWT, Stripe, Pyroscope, and an i18n stack, yet arriving as a small static binary with no C runtime, is a deliberate footprint choice.

The release gate matches that discipline. Go tests, frontend Bun tests, TypeScript type checking, JSON wrapping rules, and synchronized test images together form what has to pass before a release ships. The module targets a recent Go toolchain (go 1.25.1 declared in go.mod) and carries both MySQL and PostgreSQL drivers plus SQLite, which is consistent with the cross-database promise. None of this is a reason to trust the billing logic blindly, but it does mean the maintainers treat the build and the test matrix as part of the product rather than an afterthought.

## The current releases are 2.0.0 SmartOps pre-release tags

Recent tagged releases are all pre-releases of the same line: v2.0.0-smartops.pre1, then pre2, then pre3, each roughly a week to two apart. The README links to the latest stable release notes rather than to these tags, and its production advice is to pin a confirmed stable tag, back up the database, and have a rollback plan ready before upgrading. That is the right posture for a gateway holding billing state and access tokens, and it means the SmartOps center described in the capabilities table (active alerts, channel and model performance, system info, and billing settlement reconciliation evidence) is arriving through a pre-release train rather than a settled stable line.

Two more boundaries are worth naming for anyone reading further. First, the gateway expects you to add only upstream channels you are legally authorized to use, and if you offer generative AI services to the public you are responsible for upstream authorization, filing and permits, content safety, real-name verification, log retention, tax, payment, and user agreements in your jurisdiction; log audit and content retention are meant to be enabled only where there is a legal basis, clear notice, permission isolation, and data security. Second, private deployment is supported across SQLite, MySQL, PostgreSQL, Redis, multi-node, and a separate log database, but that support is a statement about what the code can talk to, not a recipe for operating it. The README here is also cut off partway through its use-cases section, so the scenario list beyond team and organization model access is not visible in this copy.

## Conclusion

MAX API fits a team that has outgrown hardcoding one vendor per application and now needs per-token budgets, auditable billing, and one place to watch channel performance. It does not fit a single-key hobby project, because the operational floor, a supported MySQL or PostgreSQL plus Redis, HTTPS, backups, is real work. Verify the version tag before you deploy: the recent releases are 2.0.0 SmartOps pre-release tags, and the README tells you to pin a confirmed stable tag and back up the database before upgrading. Start on the SQLite Docker path to confirm routing and settlement behave, then move to PostgreSQL 14 or MySQL 8.4 LTS before real traffic, not after.

## FAQ

### What does MAX API actually do, in one sentence?

MAX API is a unified model gateway, governance control plane, and operations entry point that sits between applications, agents, and upstream model services. It unifies multi-vendor access and holds the permission, cost, and security boundary.

### Which database should I run MAX API on in production?

Not SQLite, which is for local use, development, and small tests. Production should use MySQL still in its security support window (8.4 LTS suggested) or PostgreSQL 14+, with Redis, HTTPS, and backups configured. The compatibility floor is MySQL 5.7.8 and PostgreSQL 9.6, but neither is recommended for production.

### How does MAX API avoid charging twice when a request retries?

Billing pre-deducts at request start, records the charge idempotently, then does a final settlement once real usage is known, with refunds on failure and explicit pending or manual states for async tasks. The gateway keeps one billing truth rather than a separate ledger per agent.

### How do I run MAX API locally to try it?

Pull the Docker image, which is pinned by digest, and run it with the port bound to 127.0.0.1:3000 and a ./data volume, using SQLite by default. Then open http://localhost:3000, create the admin account, add an authorized upstream channel and key, and mint an access token for your app.

### What license is MAX API released under?

AGPL-3.0. The Dockerfile also copies the LICENSE and THIRD-PARTY-LICENSES.md files into the image's /licenses directory.

## Sources

- [License: AGPL-3.0](https://github.com/MAX-API-Next/MAX-API/blob/main/LICENSE)
- [MAX-API-Next/MAX-API on GitHub](https://github.com/MAX-API-Next/MAX-API)
- [Project website](https://max-api.com)
- [README](https://github.com/MAX-API-Next/MAX-API/blob/main/README.md)
- [Releases](https://github.com/MAX-API-Next/MAX-API/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/max-api-next-max-api
