Model or dataset
mozilla-ai/otari avatar
mozilla-ai/otari

mozilla-ai/otari: a self-hosted OpenAI-compatible LLM gateway

Open-source, OpenAI-compatible LLM gateway you run yourself. One endpoint for 40+ providers, with virtual keys, budgets, and usage tracking.

494 stars58 forksPythonApache-2.0

At a glance

What is it?
Otari puts one OpenAI- and Anthropic-compatible endpoint in front of more than 40 model providers, with virtual keys, budget enforcement and usage records. It is aimed at teams that want provider credentials to stay behind their own gateway, and it needs PostgreSQL or SQLite plus a master key before it does anything.
Who is it for?
Adopt Otari if you already run your own infrastructure and the problem you have is credential sprawl and untracked spend across several model providers, since the gateway's job is to authenticate, resolve credentials, check budgets before dispatch and record usage afterwards. Do not adopt it if you want a managed endpoint with no database to operate, or if you only ever call one provider from one service, where a direct SDK call is less machinery.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The credential and budget problem Otari was built for

The README describes the position Otari takes in a stack in one sentence: it sits between your applications and model providers, authenticates requests, resolves provider credentials, enforces budgets before dispatch, and records usage afterwards. That ordering is the whole argument. Most teams that call three or four providers end up with provider keys copied into each service, no shared view of who spent what, and no way to revoke access for one person without rotating a key that several services share.

Otari's answer is a revocable API key with user, workspace and model scope, issued by the gateway rather than by the provider. Applications authenticate against Otari with that key. The provider credential stays behind the gateway, so a leaked application key does not hand over the OpenAI account. The same layer is where budget checks happen, and the README is explicit that the check runs before spend while the usage record is written after settlement. Those are two different moments, and the gap between them is where overspend can still occur if a request is expensive and the budget is nearly exhausted.

Who this is for: platform or infrastructure teams running their own services, who need per-user or per-workspace accounting and want one endpoint that both OpenAI and Anthropic clients can talk to. Who it is not for: a single developer calling one model from one script. The gateway adds a database, a master key and a config file to a problem that a direct SDK call already solves.

How requests flow through the gateway

Provider calls go through any-llm, per the README, and that dependency is visible in pyproject.toml as any-llm-sdk[all]>=1.27.1. Otari is the layer that decides which provider credential to use and whether the caller is allowed to spend, and any-llm is the layer that actually speaks each provider's protocol. This is why the model string in a request carries a provider prefix: the quickstart sends "openai:gpt-4o-mini", so routing information travels in the model field rather than in a separate header.

The core completion routes are POST /v1/chat/completions, POST /v1/messages and POST /v1/responses. The first two are the OpenAI and Anthropic shapes; the third is the Responses API shape. A running server publishes Swagger UI at /docs and OpenAPI at /openapi.json, so the full surface is inspectable without reading source.

Routing is not just a lookup table. The README lists local routing policies for failover, weighting and learned selection, which means the gateway can hold more than one candidate for a request and pick between them. Budget enforcement sits in front of that dispatch step, and usage recording sits behind it. The pyproject.toml comments also show a deliberate separation for guardrails: the gateway imports only the guardrail catalog (GuardrailName and get_parameter_schema) and runs guardrails against an operator-supplied any-guardrail sidecar rather than constructing them in-process. That keeps a model backend out of the gateway's own dependency tree, at the cost of one more service to run if you want guardrails at all.

Running otari serve in Docker and sending a first request

The README's quickstart runs an ephemeral standalone gateway. This container uses SQLite inside the container and is deleted when it stops, so it is a way to see the thing work, not a deployment. The master key and the provider key are passed as environment variables, and OTARI_CONFIG_YAML carries inline configuration.

bash
docker run --rm -p 8000:8000 \
  -e OTARI_MASTER_KEY=SET_A_MASTER_KEY \
  -e OPENAI_API_KEY=YOUR_OPENAI_KEY \
  -e OTARI_CONFIG_YAML='default_pricing: true' \
  mzdotai/otari:latest \
  otari serve

On the first empty database, Otari creates an API key and prints it once. The README shows the log line and the shape of the key.

text
No API keys found. Created bootstrap key for first run. Save this key now:
gw-...

That key is the credential your applications use. Send a request with it, and note that the model field carries the provider prefix.

bash
curl http://localhost:8000/v1/chat/completions \
  -H "Authorization: Bearer gw-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai:gpt-4o-mini",
    "messages": [{"role": "user", "content": "Say hello."}]
  }'

Existing OpenAI clients work by setting base_url to http://localhost:8000/v1. For anything you intend to keep, the README points at the Compose setup instead, which runs Otari with PostgreSQL and reads a config.yml you copy from config.example.yml. The dashboard is served at http://localhost:8000/, and to store provider keys through the dashboard you set OTARI_SECRET_KEY to a Fernet key generated by otari gen-secret-key.

Where the gateway gets in your way

The first constraint is the database. Standalone mode uses local storage, and the README's development section tells you to change database_url to sqlite+aiosqlite:///./otari.db if you want to work without PostgreSQL. That is convenient for development and a poor fit for anything with concurrent writers. The Compose path exists because PostgreSQL is the intended production store.

The second is the bootstrap key. On a first empty database the gateway prints a key once, and the README's phrasing is a warning: save this key now. There is no documented recovery path for a bootstrap key you failed to copy, and the README does not document rollback or key re-issuance for that case. Treat the first run as a step that needs someone watching the logs.

The third is the budget gap already noted. A budget check happens before dispatch and the usage record is written after settlement, so a request that is in flight when the budget is nearly exhausted can push spend past the limit. If your requirement is a hard cap that no single request can breach, this design does not promise that.

The fourth is scope. Otari is a gateway, not a model host and not an evaluation harness. If your problem is picking the best model for a task, or running prompts against a dataset, this is the wrong layer. It also assumes you can supply provider credentials at all; it does not give you access to models you have no account for.

Otari against a direct SDK integration and against hosted routing

The obvious alternative to a self-hosted gateway is no gateway: each service imports a provider SDK and holds its own key. The difference is where state lives. With direct SDK calls, budget limits and usage records have to be built per service, and revocation means rotating a key that every service shares. With Otari, those live in one database behind one endpoint, at the cost of running that endpoint, its database and its migrations. For a two-service system this is a bad trade. For a platform with several teams and several providers, the accounting is the product.

The other alternative is a hosted router: send traffic to someone else's endpoint and let them hold the provider keys. That removes the database and the Compose file from your plate. It also means your prompts and your provider credentials pass through a third party, and your availability is tied to theirs. Otari's own runtime modes table makes the middle ground explicit: hybrid mode runs a data-plane gateway that resolves credentials locally and reports usage to otari.ai, selected when OTARI_AI_TOKEN is set and OTARI_MODE is unset. So the project itself supports the split between where inference credentials live and where the control plane lives, rather than forcing one answer.

One more distinction worth naming: Otari's routing policies operate locally, per the README, which means failover and weighting decisions do not depend on a remote service being reachable. That is the design bet. It is a good one if you distrust a network hop in the request path, and a burden if you would rather not own the routing configuration.

Maintenance, licensing and the cost of upgrading

The repository is not archived, and the last push was on 2026-09-10. Releases v0.5.0 and v0.5.1 both landed on 2026-09-08, with v0.4.0 before them on 2026-07-24. The version cadence is irregular: two releases in one day, then roughly six weeks before the next minor. A project in the 0.5 range makes no compatibility promise, and the pyproject.toml comments show the maintainers pinning dependencies defensively, for example any-guardrail>=0.7.7,<0.8.0 with an explanation that a 0.x minor is not trusted to keep the three symbols they import. Expect to read changelogs rather than assume upgrades are drop-in.

Upgrade cost has three parts. Schema changes come through alembic, which is a dependency and has its own alembic.ini and alembic/ directory in the repository, so database migrations are part of the upgrade path. The dashboard is a separate pnpm project under web/, built during the Docker image build rather than committed, which means a source build needs Node and pnpm in addition to Python. And the runtime itself requires Python 3.13 or newer, per both pyproject.toml and the README badge.

On licensing: the project is Apache-2.0 and ships a LICENSE file. That permits commercial use and modification, and it includes a patent grant. It is not legal advice, and if you redistribute Otari or a modified version you should read the licence text and any notice requirements yourself. Note one packaging detail that matters if you build from source: the distribution and import package are named gateway, not otari, because the otari name on PyPI belongs to the Otari client SDK. The CLI, environment variables and documentation all say otari; only the internal import path says gateway.

Editorial conclusion

Adopt Otari if you already run your own infrastructure and the problem you have is credential sprawl and untracked spend across several model providers, since the gateway's job is to authenticate, resolve credentials, check budgets before dispatch and record usage afterwards. Do not adopt it if you want a managed endpoint with no database to operate, or if you only ever call one provider from one service, where a direct SDK call is less machinery. Before committing, verify two things in your own environment: that the provider you depend on is reachable through the any-llm path Otari dispatches on, and that your config.yml carries the master key, provider credentials and pricing the README asks for, because an empty database only prints a bootstrap key once and that key is the only way in.

Frequently asked questions

What is mozilla-ai/otari?

It is an OpenAI-compatible LLM gateway that you run yourself, described in the README as a way to route one endpoint to more than 40 providers, issue virtual keys, enforce budgets and track usage. It sits between your applications and model providers, and provider calls go through any-llm.

How do I install and start Otari?

The README's quickstart runs an ephemeral standalone gateway with docker run, passing OTARI_MASTER_KEY, a provider key and OTARI_CONFIG_YAML, with the command otari serve. For persistent data it points at the Compose setup, which runs Otari with PostgreSQL from a config.yml copied from config.example.yml.

Where does Otari store provider credentials?

Provider credentials stay behind the gateway, per the README, and applications authenticate with revocable API keys scoped to a user, workspace and model. To store provider keys through the dashboard you set OTARI_SECRET_KEY to a Fernet key generated by otari gen-secret-key.

Does Otari need PostgreSQL?

The Compose setup runs Otari with PostgreSQL, and the README's development section says that for local development without PostgreSQL you can change database_url to sqlite+aiosqlite:///./otari.db. The quickstart container uses SQLite inside the container and is deleted when it stops.

Which API routes does Otari expose?

The README lists POST /v1/chat/completions, POST /v1/messages and POST /v1/responses as the core completion routes, and says standalone also serves the broader OpenAI-compatible and management APIs. A running server publishes Swagger UI at /docs and OpenAPI at /openapi.json.

Official sources

  1. License: Apache-2.0
  2. mozilla-ai/otari on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mozilla-ai-otari.svg)](https://hysenlabs.com/projects/mozilla-ai-otari)