# Inside caozhiyuan/copilot-api: one local port serving three AI protocols

> copilot-api is a MIT licensed TypeScript gateway that serves OpenAI Chat Completions, OpenAI Responses and Anthropic Messages from a single localhost endpoint, routing each request to GitHub Copilot, a built-in Codex provider or a third-party API. Its most interesting engineering is not the routing table but the guard rails around the data directory and the model tier mapping that lets a Claude Code session run against a non-Anthropic model.

**caozhiyuan/copilot-api** — GitHub Copilot, OpenAI Codex, OpenCode Go, and third-party AI provider gateway with OpenAI and Anthropic API compatibility.

- Repository: https://github.com/caozhiyuan/copilot-api
- Website: https://caozhiyuan.github.io/copilot-api/
- Stars: 1,067 · Forks: 239
- Language: TypeScript
- License: MIT
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/caozhiyuan-copilot-api

## One process, three protocol surfaces, port 4141

The premise is a single local process that speaks three wire formats at once. Start it and it listens on `http://localhost:4141`, exposing an OpenAI Chat Completions endpoint, an OpenAI Responses endpoint and an Anthropic Messages endpoint. Any client that speaks one of those three shapes can point at the same address, and the gateway handles the rest.

```sh
npx @jeffreycao/copilot-api@latest start
```

The quickest way to confirm the process is alive is to ask it what it can serve:

```sh
curl http://localhost:4141/v1/models
```

There is a second entry point for credentials rather than for the server:

```sh
npx @jeffreycao/copilot-api@latest auth login
```

The runtime floor is stated twice in the documentation, once as a runtime requirement and once as a storage requirement: Node 22.13.0 or newer, or Bun 1.2.x or newer. Token usage storage is the part that needs the newer runtime, so a machine on an older Node can run the gateway and still fail to keep usage records. That distinction is easy to miss and worth noting before anyone builds reporting on top of it.

The published package exposes a single binary named `copilot-api` pointing at the built entry script, and ships two directories rather than one: the compiled output and a `pages` directory. A top level `plugin` directory and a `.claude-plugin` directory sit alongside it, which is how the same gateway can be presented as a plugin rather than only as a process you start yourself.

## GitHub Copilot is optional, and that changes the design

Most gateways in this space assume one upstream. This one treats the upstream as a list, and the list has a default member that can be removed. GitHub Copilot is the first entry, followed by a built-in provider named `codex` and a set of third-party providers including Kimi, DeepSeek, DashScope, OpenRouter, OpenCode Go, or a custom provider of your own defining. The documentation is explicit that with at least one provider enabled the server starts without a GitHub token at all, in what it calls provider-only mode.

That single design choice cascades through the rest of the project. If no token is required to boot, then a fresh install is a matter of starting a process, and the only mandatory configuration belongs to whichever provider the user actually wants. It also means the gateway is not a Copilot client with extras bolted on, which is the framing a reader is most likely to bring to the name.

Protocol support is per model rather than per provider, and the documentation says so twice, once in the compatibility notes and once in the provider section. Chat Completions needs a native endpoint. Responses and Messages can go through supported adapters. The built-in `codex` provider uses Responses natively, and third-party providers declare one of `anthropic`, `openai-compatible` or `openai-responses`, with the ability to override per model. Per model overrides matter because provider catalogues change faster than protocol implementations do, and a single provider level declaration would strand models whenever an upstream adds a new surface.

## The Claude Code tier mapping, in two ways

Coding agents are treated as first class clients, and Claude Code gets the most elaborate path because of a mismatch: that client asks for models by family name rather than by model identifier. The gateway answers with a launcher that resolves the ambiguity at start time.

```sh
npx @jeffreycao/copilot-api@latest start --claude-code
```

With that flag the interactive model picker is skipped. The gateway looks at the models it can currently serve and maps each Claude Code size tier to the newest available model in that family, opus to the newest Opus, sonnet to the newest Sonnet, haiku to the newest Haiku. A tier with no matching model is left out of the result rather than filled with a substitute. The generated command is copied to the clipboard and it sets the three tier variables so the next terminal starts with the right mapping already in place.

The alternative is to write the mapping yourself, and the documentation shows what that looks like. The project level settings file carries the base URL, a placeholder auth token, an explicit model, one variable per tier, and a set of switches that turn off other provider paths and tune the auto-compaction window:

```json
{
  "env": {
    "ANTHROPIC_BASE_URL": "http://localhost:4141",
    "ANTHROPIC_AUTH_TOKEN": "dummy",
    "ANTHROPIC_MODEL": "gpt-5.6-sol[1m]",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "gpt-5.6-sol[1m]",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "gpt-5.6-sol[1m]",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "gpt-5.6-luna[1m]",
    "CLAUDE_CODE_AUTO_COMPACT_WINDOW": "272000",
    "CLAUDE_CODE_USE_VERTEX": "0",
    "CLAUDE_CODE_USE_BEDROCK": "0"
  }
}
```

Two things stand out in that file. The auth token is a placeholder, which tells you the gateway is not expecting the client to hold a real secret. And the compaction window is set far above a default, which is a signal about what the upstream tolerates rather than about the client. Both are choices a reader copying this file should re-evaluate for their own context, since a very large context window assumption is exactly the kind of setting that breaks when the model behind the alias changes.

## Streaming is SSE on the way in, WebSocket on the way out

Client-facing traffic is uniform. All three protocols stream over server-sent events, so a client does not have to learn a second mechanism to get tokens back.

Upstream traffic is not uniform, and this is the part of the design that is genuinely per provider. For Copilot, Responses requests pick WebSocket or HTTP from the endpoints each model advertises, so the choice follows the catalogue rather than a global setting. For the built-in `codex` provider, streamed Responses traffic uses WebSocket by default, and there is a named flag, `useResponsesApiWebSocket`, that turns that back to HTTP. The existence of that flag is a useful signal: it means the default is a bet that can be lost, and the project leaves the escape hatch visible rather than burying it.

The client matrix makes the trade-off explicit. Claude Code is served over Anthropic Messages, natively or through the adapter. OpenCode is listed as native on Chat Completions and Responses and as reaching Messages through a third-party SDK. Codex is served over Responses. Generic OpenAI clients get Chat Completions, generic Anthropic clients get Messages. The recommended column picks one protocol per client, and reading it as a recommendation rather than a limitation is the faster interpretation, because the whole point of a translating gateway is that the client keeps its native shape and the translation happens underneath.

## An Electron app, three platform packages, and a default branch called dev

Beyond the command line there is a desktop application in the `desktop` directory, built as a separate bundle with its own build config, its own lint and typecheck scripts, and its own hook in the pre-commit pipeline. It covers GitHub Copilot sign-in, Codex OAuth, API key configuration for the third-party providers, one-click start and stop of the local server, and a window that shows the endpoint, the auth header, the available models, usage and logs together.

Packaging is per platform rather than universal: a Windows x64 executable, a macOS Apple Silicon disk image, and a Linux x64 AppImage, published in the project's GitHub releases. The omission of an Intel macOS build is the kind of detail that saves an afternoon, since an Apple Silicon image will not install cleanly on older hardware.

One piece of repository metadata is worth pausing on. The default branch is `dev`, not `main`, which means a plain clone lands on work in progress and the README on the default branch is the current one. The release cadence supports that reading: three tagged versions in the last week, with 2.6.0 and 2.6.1 on the same day and 2.6.2 a few days later. The desktop directory is listed in the repository root but is not part of the published package files, so the npm install path and the desktop path are maintained together while shipping separately.

## The data directory guard in the Compose file

The most careful code in this repository is not in the gateway. It is a shell preamble in the Compose file, and it exists because the gateway stores its state on a mounted host directory that the user names through an environment variable. A mistyped value pointing at a filesystem root is a realistic way to break a machine, so the file refuses several classes of value before the gateway ever starts.

The first pass normalises the path so that equivalent spellings collapse to one form: repeated leading slashes are reduced, trailing slashes are stripped, and a trailing dot segment is removed. The comment explains why, namely that a literal blocklist can otherwise be bypassed by writing the same directory a different way. The second pass then checks the normalised value against a fixed set of system directories, and the third closes the gap that a string check cannot see. A path can look safe and still resolve to a system root through a parent segment or a symbolic link, so the script inspects the mounted directory itself and walks a list of marker names, refusing the mount if it finds any of them inside. The failure messages are specific about which condition tripped, and the final repair step touches only the gateway's own paths rather than rewriting whatever else the directory happens to contain.

The container image is built on the same principle at a different layer. It runs as a non-root user, creates its data directory with owner and mode set explicitly, and exposes a volume, which means the permissions are established by the image rather than inherited from whatever created the volume on the host.

## A Bun build with two bundles, a hook pipeline and a dead code check

The build runs on Bun rather than Node, in two passes. The main bundle is produced with tsdown, and the desktop application has its own configuration file and its own output, which is how one repository ships a library and a GUI without either one bloating the other.

```
FROM oven/bun:1.4.2-alpine AS runner
WORKDIR /app

ENV NODE_ENV=production \
    NODE_USE_SYSTEM_CA=1 \
    COPILOT_API_HOME=/data
```

The runtime image sets the production flag, enables the system certificate store, and points the gateway's home at the data volume so that state and credentials live in the mounted directory rather than inside the container's own filesystem. The build stage installs from a frozen lockfile, and the runner stage repeats the install with production only and scripts disabled. A health check polls the local port on a fixed interval with a start period long enough to survive a cold boot, and the entry point is a shell script rather than a binary so that the image can decide how to start depending on the mounted state.

The surrounding tooling is worth a mention because it explains how a project shipping this fast stays coherent. A typecheck covers both the main package and the desktop directory, as does linting. A dead code check is wired into the scripts, which is the kind of tool that matters in a codebase translating between three protocols, since unreachable branches accumulate there faster than anywhere else. Publishing runs through a version bumper in a single command, and a pre-commit hook runs the linter on every staged file with a separate configuration for the desktop tree. None of this is visible in the feature list, and all of it is the reason a project with three protocol adapters and an Electron app can release three tagged versions in a week without losing its structure.

## Conclusion

copilot-api is worth your time if you want a single local endpoint that several coding agents can share, and if you are willing to treat provider credentials as the project's most sensitive asset. Before pointing anything at it, read the compose file's data directory guard and understand that it exists because a mistyped mount can otherwise reach system directories, and confirm which of the three protocols your client actually speaks before assuming the adapter path will behave like the native one. Teams that need audit logs, per-user quotas or a shared hosted deployment are outside what this project offers, since it is built to listen on localhost and be started and stopped by the person using it.

## FAQ

### What port does copilot-api listen on and how do I check it is running?

It listens on http://localhost:4141 by default. Run the start command and then request the models endpoint with curl to confirm the gateway is up.

### Do I need a GitHub Copilot token to run copilot-api?

No. With at least one provider enabled the server starts in provider-only mode without a GitHub token, routing to the built-in codex provider or to a configured third-party provider instead.

### Which API formats can copilot-api serve at once?

OpenAI Chat Completions at /v1/chat/completions, the OpenAI Responses API at /v1/responses and Anthropic Messages at /v1/messages, all from the same local endpoint with server-sent event streaming.

### What runtime does copilot-api need?

Node 22.13.0 or newer, or Bun 1.2.x or newer. Token usage storage specifically requires the newer Node or Bun, and the container image is built on Bun 1.4.2 Alpine.

### How do I point Claude Code at copilot-api without picking models by hand?

Start the gateway with the --claude-code flag. It maps each Claude Code size tier to the newest available model in that family, omits tiers with no match, and copies a command that sets the three default tier variables to your clipboard.

### Is there a desktop version of copilot-api and which platforms does it cover?

There is an Electron desktop app covering Copilot sign-in, Codex OAuth, API key configuration, token usage and logs. Packages are published for Windows x64, macOS Apple Silicon and Linux x64 in the GitHub releases.

## Sources

- [caozhiyuan/copilot-api on GitHub](https://github.com/caozhiyuan/copilot-api)
- [License: MIT](https://github.com/caozhiyuan/copilot-api/blob/dev/LICENSE)
- [Project website](https://caozhiyuan.github.io/copilot-api/)
- [README](https://github.com/caozhiyuan/copilot-api/blob/dev/README.md)
- [Releases](https://github.com/caozhiyuan/copilot-api/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/caozhiyuan-copilot-api
