OrcaRouter Lite: a self-hosted LLM router with model="auto" and a hosted fallback
Self-hosted LLM router with a managed safety net. OpenAI-compatible. BYOK. Single-workspace. Streaming. For more advanced routing choose hosted OrcaRouter
At a glance
- What is it?
- OrcaRouter Lite is the MIT-licensed, single-workspace edition of OrcaRouter: a FastAPI server that speaks OpenAI, Anthropic and Gemini wire formats, routes with model="auto", and can fall back to the hosted service. Here is what it does, how to run it, and where it stops being the right tool.
- Who is it for?
- Adopt OrcaRouter Lite if you want a self-hosted OpenAI-compatible endpoint with a dashboard, no Postgres or Redis requirement, and a hosted account you can switch on as a fallback provider. Do not adopt it if you need multi-tenant key management, a documented upgrade path, or routing rules more expressive than cheapest-capable; the repository has no migration guide and the README never describes rollback.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem OrcaRouter Lite actually solves
Most teams end up with provider keys scattered across services, a cost-optimization branch somewhere in application code, and no single place to see which model answered a request. OrcaRouter Lite puts a server in front of those providers. You point any OpenAI SDK at http://localhost:8000/v1, keep your own keys, and get a dashboard for providers, routing, analytics and keys. The README frames the audience directly: developers who want to run it on a laptop or ship it inside a product, and who do not want to manage keys for every model in the long tail. The single-workspace constraint is the dividing line between this and the hosted product. There is one workspace, so there is no built-in story for handing each of your own customers an isolated key and budget. If you are building a multi-tenant gateway, this is the wrong edition, and the README says so by pointing at hosted OrcaRouter for advanced routing.
model="auto" and the routing pipeline behind the three protocols
The headline feature is a model string. Send model="auto" and the router selects the cheapest model among your configured providers that satisfies the request's capability requirements, which the README lists as tools, vision and JSON mode. A request carrying an image routes to the cheapest vision-capable model your keys cover. The resolved choice comes back in the x-orca-resolved-model response header, which is the part that matters operationally: without it, auto-routing would be a black box in your logs. Lite accepts three inbound protocols (OpenAI, Anthropic and Gemini) and translates them at the edge into one internal pipeline, so model="auto", the cross-provider prompt cache, routing strategies and analytics behave the same regardless of which SDK you used. The README states the cache is shared across protocols, which is a real design decision rather than a marketing line: a prompt sent via the Anthropic endpoint can hit a cache entry created by an OpenAI-shaped request. Streaming follows the standard SSE framing with data: lines and a terminal [DONE] sentinel, so existing streaming clients need no adapter. The routing logic is described only in terms of cheapest-capable; the README does not document weighting, latency preference or per-request overrides beyond naming the model explicitly.
Installing OrcaRouter Lite and making a first routed call
The README gives a 60-second quickstart for the self-hosted path. Clone the repository, copy the environment template, add at least one provider key, and start the compose stack:
git clone https://github.com/Continuum-AI-Corp/OrcaRouter-Lite.git
cd OrcaRouter-Lite
cp .env.example .env
# add at least one: OPENAI_API_KEY=sk-... (or ORCAROUTER_API_KEY=...)
docker compose upThe compose file builds the image, publishes port 8000, loads .env, and mounts a named volume at /data with DATABASE_URL set to sqlite+aiosqlite:////data/orca.db. On startup the logs print a generated key in the form sk-orca-abc123.... That key is what your clients authenticate with. The dashboard is served at http://localhost:8000/ on this path.
Then call it with any OpenAI SDK. This is the Python example from the README, with model="auto" so the router picks the model:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="sk-orca-abc123...",
)
r = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Hello!"}],
)
print(r.choices[0].message.content)If you would rather test the endpoint without an SDK, the README's curl form is a single POST to /v1/chat/completions with a Bearer token and a JSON body carrying model and messages. The response is a normal chat completion, and the x-orca-resolved-model header tells you which provider model was billed. If you have no local provider keys at all, the second path is to register at www.orcarouter.ai, copy an sk-orca-* key, and set it as ORCAROUTER_API_KEY so the hosted service becomes one more provider in the chain.
Where the Lite edition runs out of road
The first limitation is in the name. Single-workspace means one set of keys and one set of logs. The README's own comparison table lists BYOK and a built-in dashboard but nothing about per-tenant isolation, quotas or team-level access control, and the pyproject description repeats the single-workspace framing. If your product needs to hand a customer a scoped key with its own budget, this edition does not offer it, and the README directs that audience to hosted OrcaRouter instead.
The second is the upgrade story. The repository contains RELEASING.md and alembic is a declared dependency, which suggests schema migrations exist, but the README does not document rollback, downgrade or what happens to the SQLite database when the schema changes. There is exactly one release, v0.1.0 from 2026-05-11, and pyproject classifies the project as Development Status :: 4 - Beta. Anyone running this against a production traffic path should treat an upgrade as an untested operation until they have walked it themselves.
The third is key persistence. The .env.example warns that CREDENTIAL_ENCRYPTION_KEY and API_KEY_PEPPER are auto-generated on first run if empty, and that once generated they should be set explicitly so DB-stored provider keys remain decryptable across restarts. That is a real failure mode: a container that is recreated without those values set can leave encrypted provider keys unreadable. The README states the requirement; it does not describe a recovery procedure.
OrcaRouter Lite against LiteLLM, OpenRouter and Ollama
The README's own table is the clearest statement of positioning, and it is worth reading literally. LiteLLM is a library, not a server, and the table marks it as lacking model="auto" and a built-in dashboard. If your application is Python and you are happy to own the routing code, LiteLLM's approach embeds the abstraction in your process; OrcaRouter Lite's approach puts a network boundary and a database in front of it, which buys you the dashboard and the RequestLog rows but adds a service to operate. OpenRouter is hosted and closed-source, with no BYOK and no self-hosted option, so the difference is not features but control: you cannot run OpenRouter on a laptop or inside a private network. Ollama is local-only and the table shows it as not multi-provider, so it does not compete on the routing problem at all. The distinctive claim in the README is the combination of a self-hosted server with a managed fallback, which none of the three offer in the same shape. That combination is also the lock-in: the fallback path is a hosted OrcaRouter account, so the safety net is theirs.
Maintenance, licence and what an upgrade costs you
The project is MIT licensed, and the Dockerfile labels the image org.opencontainers.image.licenses="MIT". MIT imposes no copyleft obligation on your own code, but it also means no warranty and no support commitment from Continuum AI Corp. The dependencies are ordinary permissive Python packages (FastAPI, SQLAlchemy, LiteLLM, Pydantic), with asyncpg and redis[hiredis] as opt-in extras rather than defaults, which keeps the base install small.
On maintenance: the repository is not archived, and the last push was on 2026-09-04, so the code is recent. That is not the same as a stable release cadence. The only tagged release is v0.1.0 from 2026-05-11, and the README does not describe a deprecation policy, a support window or a compatibility guarantee between versions. Alembic is present, so schema changes are expected to be versioned, but the README does not tell you how to apply them or reverse them. The practical upgrade cost is therefore your own testing time plus a database backup, and the .env.example's note about CREDENTIAL_ENCRYPTION_KEY is the one piece of configuration you should pin before any upgrade, because losing it makes stored provider keys undecryptable.
Editorial conclusion
Adopt OrcaRouter Lite if you want a self-hosted OpenAI-compatible endpoint with a dashboard, no Postgres or Redis requirement, and a hosted account you can switch on as a fallback provider. Do not adopt it if you need multi-tenant key management, a documented upgrade path, or routing rules more expressive than cheapest-capable; the repository has no migration guide and the README never describes rollback. Before committing, verify three things on your own machine: that the sk-orca-* key printed at startup survives a container restart once CREDENTIAL_ENCRYPTION_KEY and API_KEY_PEPPER are set explicitly, that the x-orca-resolved-model header reports a model your keys actually cover, and that your Anthropic or Gemini client works against the native endpoints. The last push was on 2026-09-04 and the only release is v0.1.0 from 2026-05-11, so treat the project as a beta with a thin release history rather than a settled dependency.
Frequently asked questions
How do I install OrcaRouter Lite?
Clone the repository, copy .env.example to .env, add at least one provider key such as OPENAI_API_KEY, then run docker compose up. The README states the API is then available at http://localhost:8000/v1 with a generated sk-orca-* key printed in the logs.
What does model="auto" do in OrcaRouter Lite?
It routes the request to the cheapest model among your configured providers that meets the request's capability requirements, which the README lists as tools, vision and JSON mode. The model that was actually used is returned in the x-orca-resolved-model response header.
Does OrcaRouter Lite need Postgres or Redis?
No. The default database is SQLite via sqlite+aiosqlite, and the README's comparison table lists no Postgres or Redis requirement. Postgres and Redis are optional extras in pyproject.toml, and the .env.example notes Redis also enables cross-worker cache invalidation.
Can I use OrcaRouter Lite without any provider keys of my own?
Yes, by setting ORCAROUTER_API_KEY to a hosted sk-orca-* key from www.orcarouter.ai. The README describes this as try-before-you-buy, with hosted acting as one more provider in the routing chain.
Does OrcaRouter Lite work with the Anthropic or Gemini SDKs?
The README states that Lite accepts OpenAI, Anthropic and Gemini inbound protocols, translating them into the same internal pipeline. For Claude Code you set ANTHROPIC_BASE_URL to http://localhost:8000 with no /v1 suffix, and the google-genai SDK connects through HttpOptions with the same base URL.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/continuum-ai-corp-orcarouter-lite)