LiteLLM routes on the model string, and its default branch is an internal one
Self-hosted AI gateway with a Rust core and Python SDK that calls 100+ LLM providers in OpenAI format, adding cost tracking, guardrails, and load balancing.
At a glance
- What is it?
- LiteLLM puts one OpenAI-format interface in front of more than a hundred model providers, usable as a Python library or as a proxy server holding your keys. The routing mechanism is a prefixed model name, and the operational decisions sit in the virtual keys and the release channel.
- Who is it for?
- Adopt LiteLLM when more than one provider is in play and the cost of rewriting client code per provider is real, especially if you need per-team keys and spend attribution. Do not adopt it as a library when you call exactly one provider, because the gateway is an extra hop and a credential store you now own.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The model string is the router
Everything in LiteLLM follows from one design decision, which is that the provider is part of the model name rather than a property of the client object. A call carries a string such as openai/gpt-4o or anthropic/claude-sonnet-4-20250514, and that prefix selects the backend, the credential and the request translation. The reason the project exists is stated plainly: every model comes with different SDKs, different auth patterns, different request formats and different error types, and the library collapses that into one interface using the OpenAI format. Swapping a provider is then a string edit rather than a rewrite, which is the drop-in compatibility claim made in the project description. The same mechanism has a cost, since a malformed prefix fails at call time and the set of valid prefixes is a moving target as providers are added.
Two ways in, and they are different products
There are two installations and picking the wrong one costs you a deployment. As a library for direct integration, the command is short:
uv add litellmAs the AI Gateway, which the project also calls the Proxy Server and describes as a production-ready service for a team or organisation, the install carries an extra:
uv tool install 'litellm[proxy]'
litellm --model gpt-4oThe bracketed proxy extra is what pulls in the server, and the second line starts it against one model. Both forms expose the same call surface, so the choice is about whether the credential handling lives in your process or in a service somebody connects to. Deploy buttons are offered for Render, Railway, AWS CloudShell and Google CloudShell, which tells you the project expects the gateway to be a long-lived deployment rather than a library initialised per request.
The proxy answers on port 4000 and accepts an unmodified OpenAI client
The point of running the gateway rather than the library is that the code on the other side does not change. A plain OpenAI client is pointed at the proxy's base URL and handed a placeholder key:
import openai
client = openai.OpenAI(api_key="anything", base_url="http://0.0.0.0:4000")
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Hello!"}],
)The literal string anything is the tell. The proxy is not validating an OpenAI credential, it is doing its own authentication, which is why the gateway side of the project lists virtual keys, spend tracking, guardrails, load balancing and an admin dashboard as things you get out of the box. Note also that the model name here is a bare gpt-4o, not the prefixed form the library uses, because the prefix is resolved by the proxy rather than the caller. The endpoint surface is wider than chat, covering /chat/completions, /responses, /embeddings, /images, /audio, /batches, /rerank, /a2a and /messages.
Virtual keys turn the gateway into a service in your path
The operational difference between the SDK and the proxy is not performance, it is that the proxy holds state. With the library your provider keys are environment variables in your own process, set as OPENAI_API_KEY or ANTHROPIC_API_KEY in the example given. With the gateway the real provider credentials sit in the proxy, callers present virtual keys, and the proxy records what each one spent. That is what makes per-team attribution and budget enforcement possible at all, and it is the reason teams adopt this shape rather than the library. The same change makes the proxy a dependency: it is now the component that can see every prompt your organisation sends, and the component whose outage takes every model call with it. Both facts follow from the feature list rather than from anything surprising in the design.
The default branch is litellm_internal_staging
Version control here is unusual in a way that affects how you install. The default branch is not main, not master and not a release branch, it is litellm_internal_staging, which is a name that describes the project's own workflow rather than a public release line. The recent releases make the cadence visible: v1.102.2, then v1.103.1, then v1.104.0-rc.2, all three dated 2026-09-30, and the last push to the repository was on 2026-09-25. A release candidate appearing in the same day's stream as two final versions tells you the project ships continuously and tags often. For a reader the practical rule is that a floating version specifier will move under you, and that a major-minor-plus-patch scheme at the hundredth minor is where dependency churn accumulates, so record the exact version that passed your tests.
The 8ms P95 figure is the project's own measurement
The performance claim appears as a single line, 8ms P95 latency at 1k RPS, with a link to a benchmarks page in the documentation rather than to a method or a machine specification. It is worth attributing precisely that way, because a latency number without the provider, the payload and the hardware behind it is not something a reader can transfer to their own traffic. What the project does state as its mechanism is a Rust core with a Python SDK, so the claim is plausibly about the gateway's own overhead rather than about model response time, which is the only reading under which the figure is meaningful. The related description also names the providers covered, including Bedrock, Azure, OpenAI, Anthropic, VertexAI, vLLM and Nvidia NIM. If overhead is what you are deciding on, measure it on your own payload.
MCP is marked experimental and A2A wants a2a-sdk>=1.1.0
Two newer surfaces sit alongside the core, and they are at different levels of maturity in how the project presents them. The MCP bridge is reached through a function literally named experimental_mcp_client, which loads tools from a connected server session in OpenAI format so they can be used with any LiteLLM model. An experimental prefix in a public API is a promise that the signature may change, and code written against it will need revisiting. The A2A side is the more structured of the two, with a dedicated A2AClient, a documented set of supported agent providers including LangGraph, Vertex AI Agent Engine, Azure AI Foundry, Bedrock AgentCore and Pydantic AI, and a stated requirement of a2a-sdk>=1.1.0. Agent addressing goes through the proxy under a path of the form /a2a/agent-name, authenticated with a master key or a virtual key.
An enterprise tier and a hosted proxy sit next to the gateway
The project is presented in three forms, and only one of them is the thing most readers are evaluating. There is the LiteLLM Proxy Server, described as the AI Gateway and open source, there is a hosted proxy at an address the README links as a separate product, and there is an enterprise tier with its own page. The header navigation lists all three side by side, which is a reasonable way to run a company but leaves the licensing picture for the open source gateway less settled than the marketing suggests: the repository facts record no standard licence identifier at all, so a reader planning to modify and redistribute the gateway should read the licence file in the repository rather than assume one. The README also names a single OSS adopter, Netflix, in an otherwise empty table.
Editorial conclusion
Adopt LiteLLM when more than one provider is in play and the cost of rewriting client code per provider is real, especially if you need per-team keys and spend attribution. Do not adopt it as a library when you call exactly one provider, because the gateway is an extra hop and a credential store you now own. Verify first which release channel you are on, since three versions shipped on the same day including a release candidate, and the default branch is named litellm_internal_staging rather than a release branch, so pin an exact version before this reaches production.
Frequently asked questions
Does LiteLLM cost money?
The proxy server is presented as an open source AI Gateway, but the project also offers a hosted proxy and an enterprise tier as separate products with their own pages. Current pricing is not stated in the repository.
What are the benefits of LiteLLM?
One interface in the OpenAI format across 100+ providers, so changing provider is a change to the model string rather than to your client code. Running it as the gateway adds virtual keys, spend tracking, guardrails, load balancing and an admin dashboard.
how to install litellm
As a library, uv add litellm. As the AI Gateway, uv tool install 'litellm[proxy]' followed by litellm --model gpt-4o to start the server, which then answers on port 4000.
how to use litellm proxy
Start it with the proxy extra installed, then point any OpenAI client at its base URL with a placeholder key such as api_key="anything". The proxy does its own authentication, which is where virtual keys come in.
Which companies are currently using LiteLLM?
The OSS adopters table in the README names one company, Netflix, and no others. Beyond that single entry the repository does not list adopters.
Official sources
Where this project is recommended
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/berriai-litellm)