# claude-code-proxy: pointing the Claude Code CLI at any OpenAI-compatible endpoint

> A small FastAPI service that accepts Anthropic's /v1/messages requests, rewrites them as OpenAI chat completions, maps three Claude model tiers onto whatever you configure, and streams the answer back.

**fuergaosi233/claude-code-proxy** — Claude Code to OpenAI API Proxy

- Repository: https://github.com/fuergaosi233/claude-code-proxy
- Stars: 2,792 · Forks: 416
- Language: Python
- License: MIT
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/fuergaosi233-claude-code-proxy

## A request translator, not a new client

The description is two words, Claude Code to OpenAI API Proxy, and the mechanism is a protocol translation. Claude Code speaks Anthropic's Messages API. Most model providers, including Ollama, Azure OpenAI and a long tail of compatible services, speak OpenAI's chat completions shape. This project sits between the two and rewrites the request on the way through, so the CLI does not need to know which provider is answering.

The surface area is one endpoint, `/v1/messages`, described as complete Claude API compatibility. The feature list adds streaming over SSE, function calling with proper tool-use conversion, base64 image input, automatic injection of custom HTTP headers, and error handling with logging.

The audience is narrow and the project says so by staying small. `pyproject.toml` declares version 1.0.0 with a Beta development status classifier, requires Python 3.9 or newer, and lists five runtime dependencies: FastAPI with its standard extras, uvicorn, pydantic, python-dotenv and the OpenAI Python SDK. That is the entire dependency surface for a translation layer, which is a good sign about how little of the request path is custom.

One detail in the packaging metadata is worth flagging because it can mislead. The project URLs in `pyproject.toml` point at a different GitHub owner than the repository they live in, so the links in the package metadata do not resolve back to this project. The console script name, `claude-code-proxy`, is wired to `src.main:main`, so an installed package exposes a single entry point.

## Three environment variables standing in for Claude's model tiers

The cleverest part of the design is how little configuration Claude Code's model selection requires. Claude Code asks for models by tier, opus, sonnet or haiku, and the proxy matches on the substring in the requested model name. Requests containing haiku go to `SMALL_MODEL`, sonnet to `MIDDLE_MODEL`, and opus to `BIG_MODEL`.

That gives you one switch per tier. With `BIG_MODEL` set to a large model and `SMALL_MODEL` set to a small fast one, the CLI's internal routing between heavy and light requests keeps working while both land on the same provider. The defaults are `gpt-4o` for both the big and middle tiers and `gpt-4o-mini` for the small one.

There is a small inconsistency in the documentation worth knowing about, because it is exactly the kind of thing that wastes an afternoon. The environment variable list describes `MIDDLE_MODEL` as the model for Claude opus requests, while the mapping table says sonnet requests go to `MIDDLE_MODEL` and opus requests go to `BIG_MODEL`. The table is the one consistent with the variable's name and with the two-tier behaviour the feature is built around, so treat the prose line as a typo.

Two other variables shape behaviour. `MAX_TOKENS_LIMIT` caps tokens, defaulting to 4096, and `REQUEST_TIMEOUT` sets the upstream timeout in seconds, defaulting to 90. `LOG_LEVEL` defaults to `WARNING`, which is a sensible default for something that sits between a CLI and a paid API.

## Client authentication is optional, and that is a real choice

`OPENAI_API_KEY` is the only required variable, and it is the key for the upstream provider. Everything else has a default. The upstream base URL defaults to `https://api.openai.com/v1`, the server binds `0.0.0.0` on port `8082`.

Client-side validation is handled by `ANTHROPIC_API_KEY`, and the README is explicit about both branches of its behaviour. If you set it, clients must present that exact key to reach the proxy. If you leave it unset, any API key is accepted.

That second branch is what makes the quick start work, because Claude Code insists on having some key value set:

```bash
# If ANTHROPIC_API_KEY is not set in the proxy:
ANTHROPIC_BASE_URL=http://localhost:8082 ANTHROPIC_API_KEY="any-value" claude

# If ANTHROPIC_API_KEY is set in the proxy:
ANTHROPIC_BASE_URL=http://localhost:8082 ANTHROPIC_API_KEY="exact-matching-key" claude
```

The one-line change from any-value to the matching key is the whole security model. Note the interaction with the default bind address: the server listens on `0.0.0.0` unless you change it, so leaving `ANTHROPIC_API_KEY` unset on anything other than a loopback machine exposes an open relay to your provider credentials. Setting `HOST` to `127.0.0.1` or supplying the key are the two ways to close that.

Configuration is loaded from a `.env` file in the project root through python-dotenv, so you can copy `.env.example` and edit it instead of exporting variables in your shell.

## Custom headers as an environment variable convention

Some providers authenticate or fingerprint requests in ways the OpenAI SDK does not model, so the proxy lets you inject arbitrary headers through environment variables with a `CUSTOM_HEADER_` prefix. The name transformation is mechanical: `CUSTOM_HEADER_ACCEPT` becomes the `ACCEPT` header, `CUSTOM_HEADER_X_API_KEY` becomes `X-API-KEY`, and `CUSTOM_HEADER_AUTHORIZATION` becomes `AUTHORIZATION`.

The README groups the supported header types into four categories: content type for `ACCEPT` and `CONTENT-TYPE`, authentication for `AUTHORIZATION` and `X-API-KEY`, client identification for `USER-AGENT`, `X-CLIENT-ID` and `X-CLIENT-VERSION`, and tracking for `X-REQUEST-ID`, `X-TRACE-ID` and `X-SESSION-ID`. The commented examples in `.env.example` show the shape, including an `AUTHORIZATION` value set to a bearer token, which is how you would add a second credential on top of the OpenAI key.

This is a small feature with more reach than it first appears. Because the mapping is prefix-based rather than an allowlist, adding a header a provider invents later needs no code change, just a new variable.

Running it is equally short. The README gives three routes: `python start_proxy.py` directly, `uv run claude-code-proxy` for the installed console script, or `docker compose up -d`. The compose file builds from the repository `Dockerfile`, publishes 8082, and mounts your `.env` into the container at `/app/.env`. The Dockerfile itself is thin: it starts from Astral's uv image on bookworm-slim, copies the project in, runs `uv sync --locked` so the lockfile is asserted as current, and launches `start_proxy.py`.

## Provider examples, and where the project is thin

The README gives four worked configurations. OpenAI is the default with `gpt-4o` for the big and middle tiers and `gpt-4o-mini` for small. Azure OpenAI points `OPENAI_BASE_URL` at a deployment path under your resource and uses the older `gpt-4` and `gpt-35-turbo` names. Local models use Ollama at `http://localhost:11434/v1` with `llama3.1:70b` for the large tiers and `llama3.1:8b` for small, and the README notes that `OPENAI_API_KEY` is still required but can be a dummy value. Anything else that speaks the OpenAI shape works by setting its base URL.

Development setup follows the same pattern as the runtime one. `uv sync` installs, `uv run black src/` and `uv run isort src/` format, `uv run mypy src/` type checks, and the mypy configuration in `pyproject.toml` is strict: `warn_return_any`, `warn_unused_configs` and `disallow_untyped_defs` are all on, pinned to Python 3.9. There is a `tests/` directory and a root-level `test_cancellation.py`, and the README points at `python src/test_claude_to_openai.py` for a functional check.

A few things signal where this project is thin, and they are worth knowing before you depend on it. There are no published releases. The repository tree carries `BINARY_PACKAGING.md` and `pyinstaller` in the dev dependencies, so a packaged single-file build is planned or in progress rather than shipped. `CLAUDE.md` in the root suggests the repository is developed with AI coding assistance, and `QUICKSTART.md` duplicates some of the README's ground.

The last push is dated 2026-03-12, with 2,792 stars, 416 forks and 23 open issues. For a translation layer whose correctness depends entirely on upstream model names and endpoint shapes, that date is the thing to check before you commit to it: nothing in the design adapts to a provider renaming a model, and the mapping is substring-based on the tier word, so a provider change is where breakage would show up first.

## Conclusion

This proxy is at its best when you want one interface, the Claude Code CLI, pointed at several different backends, including a local Ollama model, without rewriting your workflow each time. The translation layer is the whole product: `/v1/messages` in, an OpenAI chat completion out, with haiku, sonnet and opus requests routed to three separate environment variables, so a large tool-heavy request and a quick classification can hit different models. Two things to weigh first. The client API key is optional, and if you leave `ANTHROPIC_API_KEY` unset the proxy accepts any key, which is fine on loopback and wrong on a shared host. And with the last push dated 2026-03-12, check that the model names you configure still exist upstream, since nothing here adapts to a provider renaming its models.

## FAQ

### Can Claude Code use a proxy?

Yes, by pointing ANTHROPIC_BASE_URL at a server that speaks Anthropic's Messages API. Claude Code sends its requests there instead of to Anthropic, and this proxy accepts them at /v1/messages, translates each one into an OpenAI chat completion, and streams the response back as SSE.

### How do I map Claude's opus, sonnet and haiku models onto other providers?

The proxy matches on the tier word in the requested model name. Names containing haiku go to SMALL_MODEL, sonnet to MIDDLE_MODEL, and opus to BIG_MODEL, with defaults of gpt-4o-mini for small and gpt-4o for the other two. One caveat: the README's variable list describes MIDDLE_MODEL as the opus target, which contradicts its own mapping table and the variable name.

### Does claude-code-proxy need an Anthropic API key?

Not for the upstream call, which needs your provider's key in OPENAI_API_KEY. The optional ANTHROPIC_API_KEY controls client validation: set it and clients must present that exact value, leave it unset and any key is accepted. Since the server binds 0.0.0.0 on port 8082 by default, leaving it unset on a shared machine exposes an open relay.

### Can I use a local model such as Ollama with Claude Code?

Yes. Point OPENAI_BASE_URL at the Ollama endpoint at http://localhost:11434/v1 and set the tier variables to your model names, for example llama3.1:70b for the big and middle tiers and llama3.1:8b for small. OPENAI_API_KEY is still required by the configuration even though Ollama ignores it, so the README suggests a dummy value.

## Sources

- [fuergaosi233/claude-code-proxy on GitHub](https://github.com/fuergaosi233/claude-code-proxy)
- [Issues](https://github.com/fuergaosi233/claude-code-proxy/issues)
- [License: MIT](https://github.com/fuergaosi233/claude-code-proxy/blob/main/LICENSE)
- [README](https://github.com/fuergaosi233/claude-code-proxy/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/fuergaosi233-claude-code-proxy
