# OrcaRouter Lite runs on SQLite with no Redis, and its compose file publishes the port that holds your provider keys

> Continuum-AI-Corp/OrcaRouter-Lite is a self-hosted OpenAI-compatible LLM router with one release tag five months behind the branch, three inbound wire protocols, and an optional managed upstream that the same vendor sells, while its compose file binds the API to all interfaces and keeps provider keys in a named volume encrypted with a key you generate yourself.

**Continuum-AI-Corp/OrcaRouter-Lite** — Self-hosted LLM router with a managed safety net. OpenAI-compatible. BYOK. Single-workspace. Streaming. For more advanced routing choose hosted OrcaRouter

- Repository: https://github.com/Continuum-AI-Corp/OrcaRouter-Lite
- Website: https://www.orcarouter.ai
- Stars: 1,717 · Forks: 269
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/continuum-ai-corp-orcarouter-lite

## The compose file publishes the port that also holds your provider keys

The default deployment is one service, one volume, and a published port:

```yaml
services:
  api:
    build: .
    container_name: orcarouter-lite
    ports:
      - "8000:8000"
    env_file:
      - .env
    volumes:
      - lite-data:/data
    environment:
      - DATABASE_URL=sqlite+aiosqlite:////data/orca.db

volumes:
  lite-data:
```

The server host is set separately in the environment example, and it is not left at the loopback address:

```bash
HOST=0.0.0.0
PORT=8000
LOG_LEVEL=info
```

Taken together those two files mean the API listens on every interface and Docker forwards the host port, so a default install on a laptop joined to a network is reachable by anything on that network. Nothing in the example file restricts the binding to localhost, and the dashboard with its providers, routing, analytics and keys pages sits on the same port. If you run this on shared infrastructure, that is the first line to change.

The storage side is simpler than the comparison table implies. SQLite is the default and lives in the named volume, Postgres is an opt-in extra, and the table row that says no Postgres and no Redis required is describing this compose file accurately.

## Provider keys are encrypted at rest, and rotating the key breaks them

The environment example is unusually honest about failure modes. Keys entered through the dashboard are encrypted at rest, and the encryption key is yours to generate, once:

```bash
# Generate once and keep stable: `openssl rand -hex 32`.
# Rotating CREDENTIAL_ENCRYPTION_KEY without a migrate step
# makes dashboard-stored provider keys undecryptable (chat
# 503s; the providers page shows "Can't decrypt").
# CREDENTIAL_ENCRYPTION_KEY=    # 64 hex chars (32 bytes)
# During rotation only: the prior key. Startup re-encrypts
# stored provider keys, then you can unset this.
# CREDENTIAL_ENCRYPTION_PREVIOUS_KEY=
# API_KEY_PEPPER=               # 32+ char string
```

So the rotation path exists, in the form of a previous-key variable that startup reads and then you unset, but skipping it has a documented symptom rather than a silent one: chat requests return 503 and the providers page reports that it cannot decrypt. A separate pepper variable of at least 32 characters sits alongside the encryption key.

Keys can also come from the environment instead of the dashboard, and the file settles the precedence question: environment variables are described as convenient for ephemeral environments, and UI-set keys take precedence when both are configured. Six provider variables are listed as examples, covering OpenAI, Anthropic, Google, Groq, Together and Fireworks.

## One tag at 0.1.0, five months of commits, and a beta classifier

There is a single release, v0.1.0, dated 2026-05-11. The manifest carries the same version:

```toml
name = "orcarouter-lite"
version = "0.1.0"
requires-python = ">=3.11"
license = "MIT"
```

The default branch is `main`, the repository is not archived, and its most recent push is dated 2026-09-30. So the only tag describes a state of the code from roughly five months ago, and nothing on the page says the intervening work is unreleased or that the tag is meant as a stable point. The classifier is honest about the maturity instead: Development Status 4 - Beta.

Runtime expectations are pinned from both ends. The floor is Python 3.11, and the classifier list names 3.11, 3.12 and 3.13 with no 3.10 entry. The container images are built on Python 3.12, so the shipped runtime sits in the middle of that range rather than at the floor. The dependency set is about fourteen packages, led by FastAPI and uvicorn for the server, litellm for provider translation, SQLAlchemy with aiosqlite and alembic for storage, and cryptography with passlib for the credential handling described above.

## The container never installs the project, so its console script is not in the image

The build is a two stage Docker build, and the interesting part is what it deliberately leaves out. The builder stage copies only `pyproject.toml`, reads the dependency list out of it with a one line `tomllib` script, and installs those packages. The comment explains both halves of the choice: the project itself is not installed in the builder, because the runtime stage puts the source on `PYTHONPATH` instead, and that way the dependency layer is invalidated only when the dependency list changes rather than on every code or README edit.

```dockerfile
FROM python:3.12-slim AS builder
WORKDIR /app
COPY pyproject.toml .
COPY app/ app/
```

The manifest declares a console entry point, `orcarouter-lite = "app.cli:main"`, which a normal `pip install` would give you as a command. Putting the source on `PYTHONPATH` rather than installing the package means that entry point is never generated in the image, so the documented route into the container is the HTTP server and the dashboard, not a CLI.

The builder also installs `gcc`, `libffi-dev` and `libssl-dev` for source-built wheels, and the comment is candid that this is currently unnecessary: every listed dependency ships wheels for linux amd64 and arm64. The roughly 80 megabytes land in a layer that is discarded before the runtime stage, and the stated reason is to avoid a silent break the next time a native dependency is added.

## Three inbound wire formats, and the base URL is not the same for all of them

The router accepts OpenAI, Anthropic and Gemini wire formats and translates them at the edge into one internal pipeline, so `model="auto"`, the cross-provider prompt cache, routing strategies and the dashboard behave the same whichever protocol you speak. For OpenAI SDK clients the base URL carries the version suffix:

```bash
curl http://localhost:8000/v1/chat/completions \
  -H "Authorization: Bearer sk-orca-abc123..." \
  -H "Content-Type: application/json" \
  -d '{"model":"auto","messages":[{"role":"user","content":"Hello!"}]}'
```

For Claude Code it does not:

```bash
export ANTHROPIC_BASE_URL=http://localhost:8000
export ANTHROPIC_API_KEY=sk-orca-...
claude
```

The comment above it states the reason, no `/v1` suffix in the base URL, and the Gemini example passes `http://localhost:8000` the same way through the SDK's HTTP options. If you are moving an existing client over from a provider that uses a versioned path, that difference is the thing to get wrong, and it is easy to end up with a 404 that looks like a routing failure.

Streaming is plain OpenAI-compatible server-sent events with the standard data framing and a terminal `[DONE]` sentinel, so an SDK that already streams from OpenAI needs no changes.

## model="auto" picks the cheapest capable model and tells you which one

The headline feature is a model name that is not a model. Sending `model="auto"` makes the router choose the cheapest model among the providers you have configured keys for that satisfies the request's capability requirements, which it lists as tools, vision and JSON mode.

```python
client.chat.completions.create(
    model="auto",
    messages=[{"role": "user", "content": [
        {"type": "text", "text": "What's in this image?"},
        {"type": "image_url", "image_url": {"url": "data:..."}},
    ]}],
)
# → routes to the cheapest VISION-capable model your keys cover
```

Two details make it usable in production rather than a demo. The model that was actually chosen comes back in an `x-orca-resolved-model` response header, so it can be logged or displayed without parsing the body. And the cheapest criterion is scoped by capability, so an image request cannot land on a text-only model just because it is cheaper.

The overhead the page claims to remove is in your code rather than in the network: no manual routing rules, no rate limit handling, no conditional cost branches in the caller. The header is the only way to know what you paid for, which makes logging it a choice you have to make.

## Hosted-as-upstream turns a BYOK router into a router with someone else's bill

The self-hosted edition and the managed one are the same product at different points on the routing chain, and the page is direct about the three reasons to wire them together. Set `ORCAROUTER_API_KEY` and the hosted service becomes one more provider in your chain. The stated use cases are trying before you buy, because no local keys are needed to start; local logging, because hosted handles the routing while Lite stores RequestLog rows for the dashboard; and failover, where hosted catches the requests your local providers cannot.

```bash
# .env
ORCAROUTER_API_KEY=sk-orca-hosted-abc...
```

That last use case is also the one to read twice. The header line describes `model="auto"` absorbing a provider outage in real time with no code change, and a managed upstream is how that is achieved when your own keys run out. The cost is that the fallback is billed per token on a vendor account, and the requests that go through it carry your prompts.

The comparison table scores the project against LiteLLM as a library, OpenRouter as closed source hosted, and Ollama as local only, and the row that distinguishes it is hosted as fallback. Nobody else in that table can offer it, which is also the reason the boundary between the open source edition and the hosted one is worth deciding deliberately rather than by accident.

## Conclusion

OrcaRouter Lite suits people who want one OpenAI-shaped endpoint in front of several providers, want the cheapest capable model chosen automatically, and would rather hold the keys themselves. Before you run it, settle the terms of the managed side first, because the project's own comparison table advertises hosted as a fallback, which means some requests will leave your machine and be billed per token on a vendor account, and that is a decision about data rather than about routing. Then check two operational facts. The compose file publishes port 8000 while the server binds all interfaces, so put it behind something before it sits on a shared network. And generate the encryption key once and keep it, because rotating it without the documented migration makes stored provider keys undecryptable and chat requests fail with 503s.

## FAQ

### What does OrcaRouter Lite do with model="auto"?

It routes to the cheapest model among the providers you have keys for that meets the request's capability requirements, which are tools, vision and JSON mode. The model it resolved to is returned in the x-orca-resolved-model response header so callers can log it.

### Does OrcaRouter Lite need Postgres or Redis to run?

No. The compose file sets the database to SQLite inside a named volume, and both Postgres and Redis are optional extras in the manifest. Without Redis, the environment example notes that sibling workers keep serving stale configuration until the cache expires.

### How does OrcaRouter Lite store provider keys?

Dashboard-entered keys are encrypted at rest with a CREDENTIAL_ENCRYPTION_KEY you generate, 64 hex characters from openssl rand -hex 32. The environment example warns that rotating it without the migration step makes stored keys undecryptable, which shows up as chat 503s.

### Can I use Claude Code or the Gemini SDK against OrcaRouter Lite?

Yes, it accepts Anthropic and Gemini wire formats alongside OpenAI and translates them into the same pipeline. Claude Code is pointed at the base URL without a version suffix, while OpenAI SDK clients use the /v1 path.

### How do I run OrcaRouter Lite on my own machine?

Clone the repository, copy .env.example to .env, add at least one provider key or an ORCAROUTER_API_KEY, and run docker compose up. The base URL is http://localhost:8000/v1 and the API key is the sk-orca-* value printed in the startup logs.

## Sources

- [Continuum-AI-Corp/OrcaRouter-Lite on GitHub](https://github.com/Continuum-AI-Corp/OrcaRouter-Lite)
- [License: MIT](https://github.com/Continuum-AI-Corp/OrcaRouter-Lite/blob/main/LICENSE)
- [Project website](https://www.orcarouter.ai)
- [README](https://github.com/Continuum-AI-Corp/OrcaRouter-Lite/blob/main/README.md)
- [Releases](https://github.com/Continuum-AI-Corp/OrcaRouter-Lite/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/continuum-ai-corp-orcarouter-lite
