my-free-code stops falling back the moment output is committed
Open-source multi-provider AI gateway for Claude Code and other coding agents, with model routing, streaming, tools, reasoning, fallbacks, and local model support
At a glance
- What is it?
- my-free-code is an independent, MIT-licensed gateway that puts one local address in front of roughly forty model providers, speaks both the Anthropic Messages protocol and an OpenAI Responses-compatible one, routes by Claude tier name, and walks an ordered fallback list until the first byte of a streamed response is committed, at which point it refuses to switch rather than duplicate a turn.
- Who is it for?
- my-free-code fits someone running a coding agent locally who wants to move between a free provider, a paid one and a model on their own machine without reconfiguring the client, and who values a stable model name over knowing which upstream answered. It does not fit a team that needs hardened internet exposure, since the project states it is intended for local use and lists the loopback binding and non-trivial token as requirements.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One local address, two wire protocols
The gateway exists so a client does not have to know which provider answered. It listens on a loopback address, on a port set in the environment, and exposes two protocol surfaces.
One is the Anthropic Messages protocol, with its token counting endpoint alongside it. The other is an OpenAI Responses-compatible surface. Both converge on the same router, which is the point: a client written against either protocol gets the same routing and the same fallback behaviour.
There are also the operational endpoints a self-hosted service needs: model discovery, a health check, a local admin interface, and an authenticated admin API for scripted inspection.
The identity rule is the detail that makes the routing invisible. The public model identity stays as the gateway model even when a request is routed to another upstream provider. So a client asks for one name and gets a different engine behind it, without the client learning that happened.
The project is explicit about its provenance in one line: it is an independent implementation and is not affiliated with the company whose protocol it speaks. That matters for anyone reading the source and wondering what they are looking at.
The provider list is honest about adapters
The catalog names roughly forty providers, from the large hosted platforms through regional and specialist gateways to four local runtimes. The list includes the major cloud model hosts, several aggregator gateways, a couple of coding-specific plans, and, at the end, a local inference server, a local studio application and a local runtime.
What makes this section credible is the sentence that follows the list. Provider entries are not fake claims of universal support: providers with unusual authentication or protocols require a dedicated adapter, while the common OpenAI-compatible providers use the shared transport.
So the catalog is not forty hand-written integrations. It is one generic adapter for the many providers that speak a common shape, plus a smaller set of adapters for the ones that do not. The directory layout in the repository matches that split, with a general adapter file, a specialised one, a catalog and a runtime.
The environment template matches it too: it has a separate key variable per provider, most of them empty, plus three base URLs for local servers. A deployment is configured by filling in the ones you use rather than by declaring which of forty you want.
That is also the practical maintenance story. Adding a provider that speaks the common shape is configuration; adding one with its own authentication is code.
Fallback stops the moment output is committed
The fallback mechanism is documented with a worked example, and the rule at its end is the important part.
You configure a primary model per tier and an ordered fallback list. A request goes to the primary. If the provider fails before any output has been produced, it moves to the first fallback, and then to the next on the same condition. The environment template carries the concurrency and rate-window controls as well, with defaults for maximum concurrent requests per provider, a rate limit and the window it applies over, and an overall HTTP timeout.
The condition is doing all the work. Once a streaming response has committed output, the gateway does not silently switch providers and duplicate the turn.
That is the correct behaviour for a streaming protocol, and it is easy to get wrong. A client that has already rendered half a sentence cannot un-render it, so switching mid-stream would either duplicate the prefix or produce text stitched from two models. The project chooses the failure mode of an incomplete response over the failure mode of a incoherent one.
Provider health backoff sits alongside this as a listed capability, so repeated failures on one provider do not have to be paid for on every request.
Tier names stay, engines change
Routing is expressed in the client's own vocabulary, which is what makes the gateway transparent.
Four tier variables exist, one each for the four Claude tiers, and a general model variable. The project also lists gateway identifiers for the no-thinking case, so a client can request a model with reasoning off without needing a second name.
The example configuration is worth reading as a template.
MODEL=open_router/openrouter/free
MODEL_SONNET=deepseek/deepseek-chat
MODEL_HAIKU=groq/llama-3.3-70b-versatile
MODEL_OPUS=nvidia_nim/meta/llama-3.3-70b-instruct
FALLBACK_MODELS=deepseek/deepseek-chat,ollama/llama3.1A general model points at a free tier on a gateway. The mid tier points at one hosted provider, the small tier at another with a specific model, the large tier at a third, and the fallback list names two more, ending with a local model.
Read as a whole, that example says something about the intended use: this is a configuration where the quality tiers are not tied to the engines behind them, and where the last fallback is a model running on the same machine. Nothing about the client's request changes when the engine does.
The design goal underneath all of this is the stable public identity, and the tier variables are where that goal is implemented.
Reasoning is normalised before any provider sees it
Thinking and reasoning settings get their own layer, and the reason is a separation of concerns.
The gateway accepts reasoning intent in the style the Anthropic protocol uses, and keeps it separate from provider-specific request translation. In practice that means three normalised modes, on, off and auto, plus an optional effort level of low, medium or high.
Provider adapters then map that normalised policy onto whatever fields their upstream actually documents. Some providers have an effort parameter, some have a boolean, some have neither, and the adapter is where that difference is absorbed.
Doing this in the gateway rather than in each adapter has two benefits. A client can express intent once in the protocol it already speaks, and adding a provider does not require inventing a new reasoning vocabulary.
The same separation appears elsewhere in the architecture, which is described as deliberately keeping wire protocols apart from routing and provider code, with HTTP handling, routing and execution, provider runtime and client adapters treated as separate concerns.
Launchers set two variables and step aside
A launcher adapter does very little, and the readme says exactly what: it prepares the local proxy environment and delegates arguments to the installed client.
There are adapters for nine clients, covering the one the project is named for plus several other coding agents. Each is reachable as a subcommand of a single launcher entry point, and each expects the client to already exist on the machine, which the readme states as a precondition.
For the Anthropic-protocol client the environment is three values: the base URL pointing at the local gateway, an auth token, and a flag enabling model discovery through the gateway.
The discovery flag is the one worth understanding. Without it a client may not ask the gateway what models it has, and will instead use whatever it believes the provider offers, which defeats the whole premise. With it, the client learns the gateway's catalogue.
So the launcher abstraction is not a wrapper that manages a client's lifecycle. It is a few environment variables and an argument pass-through, which means it can add a new client by learning its three variable names rather than by integrating with it.
Local runtimes and a localhost-only posture
Local models are a first-class configuration rather than a workaround, with three examples given.
One sets a base URL and a model name for a local server on its default port. A second does the same for a local studio application on another default port. A third does it for a local runtime server. In each case the model name is namespaced the same way hosted models are, so a fallback can cross from a paid provider to a local model without any other change.
The security section is short because the intended deployment is narrow: this is for local use. The guidance is to keep the host bound to loopback, set a non-trivial proxy auth token, never commit the environment file, do not expose the admin endpoints to the internet, and rely on provider credentials remaining in configuration and never being forwarded to another provider.
That last point is the one with teeth in a gateway. A router that holds credentials for forty providers is exactly the kind of process you would want to exfiltrate, so the credential rule is the project's own acknowledgement of that.
One practical note: the example environment file ships the proxy token set to a placeholder value, so the first thing to change is that token rather than the port.
The test suite covers routing, protocol conversion, authentication, reasoning, the model catalogue and the streaming primitives, and the package itself depends on four libraries: a web framework, an application server, an HTTP client and a dotenv loader.
Editorial conclusion
my-free-code fits someone running a coding agent locally who wants to move between a free provider, a paid one and a model on their own machine without reconfiguring the client, and who values a stable model name over knowing which upstream answered. It does not fit a team that needs hardened internet exposure, since the project states it is intended for local use and lists the loopback binding and non-trivial token as requirements. Before you run it, change the proxy auth token, because the example environment file ships a placeholder, and decide your fallback order in advance, since the chain only works while nothing has been streamed yet.
Frequently asked questions
What is my-free-code?
An MIT-licensed multi-provider gateway for coding agents. It serves the Anthropic Messages protocol and an OpenAI Responses-compatible surface from one local address, routes by Claude tier name, walks an ordered fallback list when a provider fails before output, and ships launcher adapters for nine coding clients.
Is my-free-code affiliated with Anthropic?
No. The readme states that it is an independent implementation and is not affiliated with Anthropic. It speaks the Anthropic Messages protocol so that existing clients work unchanged, but the implementation is not from that company.
Does my-free-code support local models?
Yes. The configuration examples set a base URL and a namespaced model name for a local server, a local studio application and a local runtime, and because they speak the same shape as hosted providers a fallback list can end on a model running on the same machine.
What happens if a provider fails in the middle of a request?
Fallback applies only while nothing has been committed. Once a streaming response has produced output, the gateway does not switch providers, because doing so would duplicate the turn for a client that has already rendered part of the response.
How does my-free-code handle reasoning or thinking settings?
It accepts reasoning intent in the Anthropic style and keeps it separate from provider-specific translation, normalising it to a mode of on, off or auto with an optional effort of low, medium or high. Each provider adapter maps that policy onto whatever fields its upstream documents.
Is my-free-code safe to expose to a network?
It is intended for local use. The guidance is to keep the host bound to loopback, set a non-trivial proxy auth token, never commit the environment file, keep the admin endpoints off the public internet, and note that provider credentials stay in configuration and are never sent to another provider.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hkqr-my-free-code)