# CoAI (coaidev/coai): a self-hosted LLM gateway with built-in billing and admin

> CoAI is a Go backend plus React frontend that puts a multi-provider LLM gateway, a channel routing layer and a credit or subscription billing system behind one admin panel. It is aimed at operators who want to resell or meter model access rather than at someone who just wants a nicer chat window.

**coaidev/coai** — 🚀 Next Gen Multi-tenant AI One-Stop Solution. Builtin Admin & Billing System. Enterprise-Grade Unified LLM Gateway Support for 200+ Models And 35+ Providers, Load Balacing w/ Priority-base Routing, Cost Management, Chat Share, Cloud Sync, Credit/Subscription Billing, All File Parsing, Web Search, Built-in Model Cache.

- Repository: https://github.com/coaidev/coai
- Website: https://coai.dev
- Stars: 9,307 · Forks: 1,226
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/coaidev-coai

## The problem CoAI solves is billing and routing, not chatting

Most open source chat frontends assume you have one API key and one user. CoAI assumes the opposite. The README describes it as a multi-tenant solution with a built-in admin and billing system, and the feature list backs that up: subscription billing and elastic billing, per-request or per-token metering, gift codes and redemption codes, user grouping, and a model market that the site owner configures. The intended reader is someone operating a site where other people consume model capacity, which means every request has to be attributed to an account, priced, and routed to a provider that is currently healthy.

That framing explains several choices that look odd in a chat app. Channel priority and weight exist because you rarely have exactly one upstream for a given model. Redemption codes exist because prepaid credit is how small resellers actually collect money. The README even states the project positions itself above the combination of a chat frontend and a standalone API distribution tool, which is a claim about scope rather than about quality. Whether the integration is worth it depends on whether you need all three layers at once.

## How the Go backend, channel layer and React app fit together

The repository layout makes the architecture readable without running anything. `main.go` sits at the root, with `adapter/`, `channel/`, `connection/`, `manager/`, `middleware/`, `migration/` and `auth/` as sibling directories. `app/` holds the frontend, and the Dockerfile confirms this: a Go 1.20 Alpine stage builds a static binary named `chat`, a Node 18 stage runs `pnpm install` and `pnpm run build` inside `app/`, and the final Alpine image copies `/backend/chat` to `/chat` and the frontend build to `/app/dist`. One binary serves both.

Persistence is MySQL plus Redis, per both `go.mod` and the Compose file. `go.mod` pulls `github.com/go-sql-driver/mysql`, `github.com/go-redis/redis/v8` and `github.com/mattn/go-sqlite3`, so SQLite is a build dependency even if the shipped Compose stack uses MySQL. Token accounting has a real dependency behind it: `github.com/pkoukk/tiktoken-go` is in the require block, which is consistent with the README's per-token billing option. File parsing is not hand-rolled either: `github.com/lukasjarosch/go-docx` appears in `go.mod` alongside the README's claim of PDF, Docx, Pptx and Excel support, and `github.com/chai2010/webp` handles image conversion.

The channel layer is where the routing claims live. The README lists multi-channel management, priority for call order, weight for load balancing among equal-priority channels, automatic retry on failure, model redirection and built-in upstream hiding. Those are configuration concepts rather than algorithms you can inspect from the README, and the README does not document the retry policy, the health check interval, or what happens to in-flight requests when a channel is disabled. That is a real gap if you plan to tune failover.

## Installing CoAI with Docker Compose and reaching the admin panel

The repository ships a `docker-compose.yaml` that brings up MySQL, Redis and the application together. The application image is `programzmh/chatnio`, and the container listens on 8094 internally, mapped to 8000 on the host. Start it from the repository root:

```bash
docker-compose up -d
```

The Compose file creates three named volumes for the app: `./config`, `./logs` and `./storage`. The database credentials are set through environment variables, with `MYSQL_HOST`, `MYSQL_USER`, `MYSQL_PASSWORD`, `MYSQL_DB` and the `REDIS_*` family passed into the container. The defaults in the file are development values (`MYSQL_ROOT_PASSWORD: root`, `MYSQL_PASSWORD: chatnio123456!`), so change them before the stack is reachable from outside your machine.

The Dockerfile declares `VOLUME ["/config", "/logs", "/storage"]` and `EXPOSE 8094`, and the Compose file maps that to the host with `"8000:8094"`. The first thing to configure is not a model but a channel: the README describes channel settings, model market and price settings as the core admin surface, and it also mentions one-click synchronization with an upstream site for those settings. The Dockerfile copies `config.example.yaml` to `/config.example.yaml` in the image, so configuration lives outside the container and survives restarts. The README does not document the full set of keys in that file, so read it directly from the image or the repository before assuming a setting exists.

## Model caching and where it can surprise you

CoAI's cache is keyed on a hash of the request parameters. The README states that if the same request parameter hash has been seen before, the cached result is returned directly, and that a cache hit is not billed. It also says caching is opt-in per model, with configurable cache duration and a configurable number of cached results.

That design has a sharp edge. A cache keyed on parameters alone cannot distinguish between two tenants asking the same question, so a hit on one account's request can serve another account's response. For a gateway that bills per token, this is exactly the mechanism that reduces cost, and it is also the mechanism that makes cache scope a security decision rather than a performance one. The README does not state whether the cache is partitioned per user, per channel or globally. If you enable caching, that is the first thing to determine from the code or from a controlled test, because the answer changes what the feature is safe for.

## Web search depends on a separate SearXNG instance

The README describes full model internet search built on the SearXNG open source engine, listing Google, Bing, DuckDuckGo, Yahoo, Wikipedia, Arxiv and Qwant as supported engines, with safe search mode, content truncation, image proxy and a test-search-availability function. SearXNG is a separate project with its own deployment. CoAI does not embed it.

That means the search feature is only as available as the SearXNG instance you point it at, and the README does not document how that endpoint is configured or what happens when it is unreachable. If search is part of your product, treat the SearXNG deployment as a second service with its own uptime, not as a checkbox inside CoAI. The same pattern applies to the file parsing feature, which the README points at a separate project, CoAI.Dev Blob Service, for cloud image storage across S3, R2 and MinIO. Two of the headline features are integrations with other repositories, and each one adds an operational surface.

## What CoAI is not good at

CoAI is the wrong tool if you want a single-user chat interface. You would be running MySQL, Redis, a Go binary and a React build to serve one person, and the features that justify that stack (billing, channel priority, redemption codes) would sit unused. A client that stores conversations locally is a better fit for that case.

The second limitation is documentation depth. The README is a feature list, and feature lists read well but do not specify behavior. It does not document rollback, migration behavior between the major versions, the retry policy, the cache scope, or the full configuration schema. The release history shows why that matters: v3.10 shipped in March 2024, v3.11.1 in January 2025, and v4.0.0 in October 2025. A jump across a major version with no documented upgrade path is a real risk if you have production data in MySQL, and the `migration/` directory in the repository is the only place that would tell you what the schema changes actually are.

The third limitation is the commercial split. The README lists a Pro version with TTS and STT, a plugin marketplace, a RAG knowledge base, more payment providers, more authentication methods including SMS and OAuth login, and channel health detection with automatic switching on failure. Several of those overlap with the reliability features you would otherwise expect in the open source build, notably channel health detection. If automatic failover is central to your setup, confirm which parts of it exist in the Apache-2.0 code before you plan around them.

## Compared with running a chat frontend and a gateway separately

The conventional alternative is two deployments: a chat frontend such as Next Web for the interface, and a separate API distribution or proxy layer for keys, channels and accounting. The README explicitly frames CoAI as the union of those two categories. The practical difference is where the user identity lives. In a two-service setup, the frontend authenticates the user and the gateway authenticates the frontend, so billing attribution depends on the frontend passing a stable identifier through. In CoAI, the same Go process handles authentication, conversation storage and channel dispatch, so the account that sends a message is the account that gets billed, with no translation layer in between.

That integration is also the cost. You cannot swap out the chat UI without touching the gateway, and you cannot replace the gateway without losing the billing ledger that the UI writes to. A two-service setup lets you change one half. If your team already runs a gateway and only needs a frontend, adding CoAI means running a second gateway alongside the first, which is worse than either option alone.

## Licence, upgrade cost and what to check before you commit

CoAI is licensed under Apache-2.0, and the README describes that as business-friendly for commercial secondary development and distribution, with the caveat that the licence terms still apply and the project must not be used for illegal purposes. Apache-2.0 includes an explicit patent grant and requires that you preserve notices and state significant changes. If you modify and redistribute the code, that obligation is yours; this is a description of the licence text, not legal advice.

Upgrade cost is the part the repository does not answer. There are three tagged releases, and the gap between v3.11.1 and v4.0.0 is roughly nine months. The Compose file pins `mysql:latest` and `redis:latest` rather than fixed versions, which means a fresh deployment can pull a database major version that the `migration/` code was never run against. Pinning those two images is the cheapest risk reduction available. The repository's last push was on 2026-03-12, so check the commit log after that date before you decide how much of your own work to invest in a fork.

## Conclusion

Adopt CoAI if you are running a site that sells or meters model access and you need channel priority, retries and a billing ledger in one deployment. Do not adopt it if you only want a personal chat client, because the gateway, database and admin surface are overhead you will never use. Before committing, verify two things against your own deployment: that your MySQL and Redis versions behave with the Compose file's `mysql:latest` and `redis:latest` tags, and that `config.example.yaml` exposes the channel and billing keys you plan to depend on. The last push to the repository was on 2026-03-12, so check the commit history after that date before you build a long-lived fork on it.

## FAQ

### What is CoAI and who is it for?

CoAI is a self-hosted LLM gateway with a built-in admin panel and billing system, covering channel routing, model market configuration, subscription and elastic billing, and conversation sharing. It is aimed at operators running a site where multiple users consume model capacity, not at individual chat users.

### How do I install CoAI?

The repository ships a docker-compose.yaml that starts MySQL, Redis and the application image programzmh/chatnio together. Running docker-compose up -d from the repository root brings the stack up, and the site is reachable on host port 8000, which maps to port 8094 in the container.

### Does CoAI support OpenAI-compatible API calls?

The README states that CoAI supports calling various large models in OpenAI API standard format, and that a single deployment can serve both business and consumer use through that interface. The README does not list every endpoint, so check the adapter directory in the repository for the supported request shapes.

## Sources

- [coaidev/coai on GitHub](https://github.com/coaidev/coai)
- [License: Apache-2.0](https://github.com/coaidev/coai/blob/main/LICENSE)
- [Project website](https://coai.dev)
- [README](https://github.com/coaidev/coai/blob/main/README.md)
- [Releases](https://github.com/coaidev/coai/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/coaidev-coai
