claude-code-proxy: Redirect Claude Code to OpenAI or Gemini Backends
Run Claude Code on OpenAI models
At a glance
- What is it?
- claude-code-proxy is a self-hosted Python proxy that intercepts Anthropic API calls from Claude Code and routes them to OpenAI, Gemini, or Anthropic models via LiteLLM. It trades complete feature parity for provider flexibility, and works best for local or small-team setups.
- Who is it for?
- Engineers who want Claude Code's terminal workflow but prefer OpenAI or Google billing, or who must route through Vertex AI for compliance reasons, will find claude-code-proxy workable for local and small-team use. Teams that need production-grade reliability, versioned API contracts, or authenticated multi-tenant access should look elsewhere.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 101 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Claude Code Clients Against Non-Anthropic Backends
Claude Code is an Anthropic terminal client that expects the Anthropic Messages API on the other end. Engineers who want to use that interface but pay OpenAI or Google billing face a compatibility gap: the Anthropic API format and the OpenAI API format are structurally different, and the Claude Code client cannot speak directly to GPT or Gemini endpoints.
claude-code-proxy closes that gap. It runs a local server that presents an Anthropic-compatible endpoint on port 8082. The Claude Code CLI sends its requests there, and the proxy translates and forwards them to whichever backend is configured. The round trip is transparent to Claude Code: the client does not know or care that the model on the other end is not Anthropic.
The primary audience is developers who already use Claude Code for its terminal workflow but want to route requests to a different provider, for cost management, for API access reasons, or because their organization holds a Google Cloud Vertex AI contract. A secondary use case is running the proxy in transparent mode: with PREFERRED_PROVIDER set to anthropic, it passes calls through to Anthropic unchanged, which lets you insert a local middleware layer without modifying any client configuration.
LiteLLM as the Translation Layer
The proxy is a FastAPI application. When a request arrives on port 8082, the server reads the model name from the Anthropic-format request body, applies the configured alias mapping, prepends the appropriate provider prefix, and passes the translated call to LiteLLM. LiteLLM is a Python library that normalizes calls to dozens of provider APIs behind a single interface, handling per-provider authentication and wire format.
The runtime dependencies declared in pyproject.toml are fastapi, uvicorn, httpx, pydantic, litellm, python-dotenv, google-auth, and google-cloud-aiplatform. The google-auth and google-cloud-aiplatform packages load at startup regardless of the configured provider. A Docker image built from the repository ships those Google Cloud libraries even for deployments that only use OpenAI.
One architectural boundary matters here: the proxy translates model selection and request format. It is not a semantic validator. Whether LiteLLM's Anthropic compatibility layer correctly handles a given Claude API behavior, such as multi-turn tool use with interleaved content blocks or streaming with partial JSON tool inputs, depends entirely on LiteLLM's own implementation and on the upstream provider's support. The proxy does not check compatibility between what Claude Code sends and what the chosen backend accepts.
Installing with uv or Docker
The README documents two installation paths.
The from-source path uses uv, a Python package manager. Clone the repository and enter the directory:
git clone https://github.com/1rgs/claude-code-proxy.git
cd claude-code-proxyCopy the example environment file and edit it to supply API keys and model preferences:
cp .env.example .envStart the server. The uv run command reads pyproject.toml and resolves dependencies automatically:
uv run uvicorn server:app --host 0.0.0.0 --port 8082 --reloadThe --reload flag restarts the server on file changes; the README describes it as optional and intended for development.
The Docker path uses a published image. The preferred method is Docker Compose:
services:
proxy:
image: ghcr.io/1rgs/claude-code-proxy:latest
restart: unless-stopped
env_file: .env
ports:
- 8082:8082Or run it directly as a container:
docker run -d --env-file .env -p 8082:8082 ghcr.io/1rgs/claude-code-proxy:latestWith the server running on port 8082, connect Claude Code by setting ANTHROPIC_BASE_URL before the claude command:
ANTHROPIC_BASE_URL=http://localhost:8082 claudeFrom that point, Claude Code sends its API requests to the proxy, which forwards them to the configured backend.
Mapping haiku and sonnet to Real Model Names
Claude Code refers to models by family alias rather than exact model identifiers. The proxy intercepts two aliases: haiku for the smaller model and sonnet for the larger. Three environment variables control what those aliases resolve to: PREFERRED_PROVIDER, BIG_MODEL, and SMALL_MODEL.
When PREFERRED_PROVIDER is openai (the default), haiku maps to SMALL_MODEL with an openai/ prefix, which defaults to gpt-4.1-mini, and sonnet maps to BIG_MODEL with an openai/ prefix, which defaults to gpt-4.1. When PREFERRED_PROVIDER is google, the proxy checks whether the configured model name appears in its internal GEMINI_MODELS list. If the model is on that list, it receives a gemini/ prefix; if not, the mapping falls back to the OpenAI configuration. Setting PREFERRED_PROVIDER to anthropic disables remapping entirely and passes haiku and sonnet through to Anthropic with an anthropic/ prefix, ignoring BIG_MODEL and SMALL_MODEL.
The README gives this as Example 2a for a Gemini configuration:
GEMINI_API_KEY="your-google-key"
OPENAI_API_KEY="your-openai-key" # Needed for fallback
PREFERRED_PROVIDER="google"
# BIG_MODEL="gemini-2.5-pro" # Optional, it's the default for Google pref
# SMALL_MODEL="gemini-2.5-flash" # Optional, it's the default for Google prefA practical consequence of the prefix logic: if you set BIG_MODEL to a Gemini model name not on the server's GEMINI_MODELS list, the proxy will not add the gemini/ prefix, and the LiteLLM call will fail. The README documents gemini-2.5-pro and gemini-2.5-flash as the two recognized Gemini model names. Any other Gemini model identifier requires verifying that it appears in the GEMINI_MODELS list inside server.py before use.
Vertex AI and Application Default Credentials
For Google Cloud environments where a static API key is not practical, the proxy supports Application Default Credentials (ADC). Setting USE_VERTEX_AUTH=true tells the server to use the ADC credential chain instead of a GEMINI_API_KEY. This is the standard authorization mechanism for workloads running inside Google Cloud.
With ADC enabled, two additional variables are required. VERTEX_PROJECT must hold the Google Cloud Project ID, and VERTEX_LOCATION must hold the region. The README provides this configuration as Example 2b:
OPENAI_API_KEY="your-openai-key" # Needed for fallback
PREFERRED_PROVIDER="google"
VERTEX_PROJECT="your-gcp-project-id"
VERTEX_LOCATION="us-central1"
USE_VERTEX_AUTH=true
# BIG_MODEL="gemini-2.5-pro" # Optional, it's the default for Google pref
# SMALL_MODEL="gemini-2.5-flash" # Optional, it's the default for Google prefThe README notes that OPENAI_API_KEY must still be present even when PREFERRED_PROVIDER is google, because the proxy falls back to the OpenAI mapping when a Gemini model name is not recognized. If the key is absent and a fallback occurs, the request fails. No documentation in the repository covers what happens if ADC token refresh fails mid-session or if Vertex AI quota is exceeded.
What the Proxy Does Not Address
The repository has no GitHub releases and no changelog. The last push was on 2026-06-23. Without a versioned release contract, a dependency update to LiteLLM could change translation behavior silently, and no update policy is stated in the repository.
The proxy has no authentication layer of its own. Any client that can reach port 8082 can issue requests that consume the API keys stored in .env. The Docker Compose example binds port 8082 to 0.0.0.0, which exposes it on all network interfaces. For anything beyond local development, restricting network access or placing an authenticating reverse proxy in front is necessary; the repository does not document how to do this.
Feature compatibility is not formally documented. LiteLLM's Anthropic compatibility covers model selection and basic text generation, but whether it correctly replicates Anthropic-specific behaviors, such as vision inputs, streaming with interleaved content blocks, or complex multi-turn tool-use patterns, depends on LiteLLM's own coverage for each backend provider. The proxy has no mechanism to detect or report when a Claude API feature is silently dropped or mishandled.
The repository does not state a license. The absence of a license declaration means the default copyright law applies, and incorporating the code into commercial products or redistributing modified versions may require explicit permission.
LiteLLM Proxy Server as the Direct Comparison
LiteLLM ships its own proxy server that routes to the same set of providers but exposes the OpenAI Chat Completions format rather than the Anthropic Messages format. Clients that speak OpenAI, which includes most third-party tooling and many AI frameworks, connect to LiteLLM Proxy directly without any adapter. Clients that speak Anthropic, such as Claude Code and the Anthropic Python SDK, cannot connect to LiteLLM Proxy without an intermediate translation layer.
claude-code-proxy occupies that narrower slot: it exists specifically to give Anthropic-format clients access to non-Anthropic backends. The trade-off is a much smaller codebase with fewer operational controls. LiteLLM Proxy ships with a built-in virtual key system, per-key budget limits, a YAML configuration format for defining routing rules across multiple models, and a web interface for monitoring usage. claude-code-proxy has none of those features.
For teams that already operate LiteLLM Proxy for other services, routing claude-code-proxy through the existing LiteLLM Proxy instance rather than directly to the upstream provider may be more maintainable. A single LiteLLM instance handles provider authentication and rate limiting, and claude-code-proxy acts only as the Anthropic-to-OpenAI format adapter in front of it.
Editorial conclusion
Engineers who want Claude Code's terminal workflow but prefer OpenAI or Google billing, or who must route through Vertex AI for compliance reasons, will find claude-code-proxy workable for local and small-team use. Teams that need production-grade reliability, versioned API contracts, or authenticated multi-tenant access should look elsewhere. Before committing to it, verify that the LiteLLM backend correctly handles the specific Claude API features your code uses, particularly tool-call streaming and multimodal content, since the proxy does not validate semantic compatibility and the repository documents no test coverage for those paths.
Frequently asked questions
What is claude-code-proxy?
claude-code-proxy is a Python FastAPI server that accepts Anthropic-format API requests from clients like Claude Code and routes them to OpenAI, Gemini, or Anthropic models via LiteLLM. It lets you use Claude Code's interface while directing compute to a different provider.
How do you use claude-code-proxy?
Run the proxy server locally with uv or Docker, then start Claude Code with the ANTHROPIC_BASE_URL environment variable pointing to your local proxy. The command is ANTHROPIC_BASE_URL=http://localhost:8082 claude.
How do you install claude-code-proxy?
Clone the repository, copy .env.example to .env and fill in your API keys, then run uv run uvicorn server:app --host 0.0.0.0 --port 8082. Docker Compose is also supported using the published image ghcr.io/1rgs/claude-code-proxy:latest.
How do you configure claude-code-proxy?
Set PREFERRED_PROVIDER in .env to openai, google, or anthropic to choose the backend. Set BIG_MODEL and SMALL_MODEL to name the exact models that Claude Code's sonnet and haiku aliases should resolve to.
What is an alternative to claude-code-proxy?
LiteLLM Proxy Server routes to the same set of providers but exposes the OpenAI Chat Completions API format instead of the Anthropic Messages format, so it serves OpenAI-compatible clients rather than Anthropic clients like Claude Code.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/1rgs-claude-code-proxy)