Model or dataset
tashfeenahmed/freellmapi avatar
tashfeenahmed/freellmapi

FreeLLMAPI: One /v1 Endpoint Across 34 Free LLM Providers

7.4 billion tokens per month. 34 free LLM providers. 635 free model endpoints. All behind one /v1 endpoint, plus any custom OpenAI-compatible endpoint. Smart routing, automatic failover, encrypted keys. Personal experimentation only.

26,877 stars3,669 forksTypeScriptMIT

At a glance

What is it?
FreeLLMAPI stacks the free tiers of 34 providers behind a single OpenAI-compatible endpoint, with a router that fails over on rate limits and tracks per-key usage. It is a personal-experimentation tool, not shared infrastructure.
Who is it for?
Adopt FreeLLMAPI if you are an individual developer who wants free-tier capacity for side projects and local agents, and you accept that the free catalog snapshot lags the live feed by 30 days. Do not adopt it for multi-tenant or internet-facing production: the README calls it single-user and the Docker port binds to 127.0.0.1 by default.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 13 days ago.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem FreeLLMAPI solves, and who it is for

Every major lab ships a free tier. Individually each one is small: a few million tokens a month, a few thousand requests a day. The README's framing is that stacking them by hand is the painful part, because that means thirty-four SDKs, thirty-four rate-limit schemes, and thirty-four places a request can fail. FreeLLMAPI's claim is arithmetic: stacked, those tiers add up to roughly 7.4 billion tokens per month across 474 model families and 635 provider endpoints.

The target user is an individual. The README says so directly in the disclaimer: personal experimentation only. The Docker compose file repeats the point in a comment, noting that FreeLLMAPI is single-user and must not be exposed to the internet. If you are building a product for other people, this is the wrong shape of tool, and the project itself says so before you find out the hard way.

How the router, failover and key storage actually work

The architecture is a local server that speaks the OpenAI wire format on /v1 and translates outward. You add provider keys once; the router picks a model per request, and when a provider returns a rate-limit response it falls over to the next one. Per-key usage is tracked so you stay under each free-tier cap rather than discovering the ceiling through errors.

Keys are stored encrypted. The .env.example documents the precedence chain for the encryption key: the ENCRYPTION_KEY environment variable first, then a key file named .encryption-key written next to the SQLite database with 0600 permissions, then a legacy key found in the old settings table and migrated to the file on first boot, then a freshly generated key. In production the variable is required; outside production it is optional. That is a sensible degradation path, but it also means a production deployment with no ENCRYPTION_KEY set is a misconfiguration the app will not silently paper over.

The catalog itself is not static. The router pulls a signed model catalog from freellmapi.co on its own, so new free models, quota changes and compatibility fixes arrive without a git pull. The README is explicit about the trade-off: free installs get the monthly snapshot, so a model reaches them 30 days after it joins the live feed, while the paid tier at $19/yr gets it the same day. That is a deliberate delay, and it is the main functional difference between paying and not paying.

Installing FreeLLMAPI with Docker and sending a first request

The README points at three install paths: a Docker Compose file, a desktop app for macOS and Windows, and a Google Play build. Docker is the one with the most detail in the repository, so start there.

The compose file publishes the container on port 3001, bound to 127.0.0.1 by default. It also sets host.docker.internal to host-gateway, which is what makes the documented PROXY_URL=http://host.docker.internal:7890 setting resolve on plain Linux Docker, not just Docker Desktop.

yaml
services:
  freellmapi:
    image: ghcr.io/tashfeenahmed/freellmapi:latest
    env_file:
      - .env
    environment:
      NODE_ENV: production
      PORT: 3001
    ports:
      - "${HOST_BIND:-127.0.0.1}:${PORT:-3001}:3001"
    volumes:
      - freellmapi-data:/app/server/data

Copy .env.example to .env before the first run. The file asks for a 64-character hex encryption key, and gives the command to generate one.

bash
node -e "console.log(require('crypto').randomBytes(32).toString('hex'))"

Paste the output into ENCRYPTION_KEY. Then bring the stack up and open the dashboard on the same machine.

bash
docker compose up -d

A healthcheck polls http://127.0.0.1:3001/api/ping every 30 seconds, so docker compose ps will show the container as unhealthy rather than running-but-dead if the server fails to start. The first account is created through the dashboard. If the server is reachable from other devices, that first account also needs a one-time setup code printed in the server logs at startup while no account exists, which stops a stranger from claiming a freshly exposed install.

After you add provider keys, you point any OpenAI client library at the local server. The README's claim is that routing is transparent from that point on. What you should see is your existing client code working unchanged, with the dashboard showing which provider handled each request and how much of each free-tier cap is consumed.

If you prefer not to use Docker, the repository is an npm workspace monorepo. The root package.json requires Node >=20.18.0 and <25.0.0 and npm >=10.0.0, and exposes npm run dev, which starts the server and client together through concurrently.

Where FreeLLMAPI breaks down

The 30-day catalog lag on free installs is the first real limitation, and it is structural rather than a bug. Free-tier providers retire models and change quotas without notice. If a model you depend on is retired on a Tuesday, a free install keeps routing to it until the next monthly snapshot lands. The paid tier exists precisely to close that window, which makes the delay a pricing lever as much as an engineering decision.

Second, the deployment model is deliberately narrow. The compose file binds to 127.0.0.1, the README calls the proxy single-user, and the .env.example warns that opening HOST_BIND=0.0.0.0 should only happen on a trusted network because the proxy is guarded only by the unified API key. There is no documented multi-user story, no per-tenant isolation, and no rate-limit scheme beyond the default 120 requests per minute per client IP mentioned in the example env file.

Third, the free tiers themselves are the ceiling. Aggregating 34 providers gives you breadth, not depth: you are still subject to each provider's own quota, and the router's failover only helps when another provider has headroom. A workload that needs consistent latency or a specific model version will spend its time being routed somewhere else.

Finally, the encryption key handling has a sharp edge. Outside production, an unset ENCRYPTION_KEY causes a key to be generated and written to .encryption-key next to the SQLite database. Lose that file and the stored provider keys are unreadable. The precedence chain is documented, but the consequence of losing the key material is not spelled out in the README.

FreeLLMAPI versus OpenRouter

OpenRouter is the obvious comparison, and it appears in the related searches for this project. The difference is where the credentials live. OpenRouter is a hosted gateway: you hold one OpenRouter account and one key, and OpenRouter holds the relationships with upstream providers. FreeLLMAPI runs on your machine, and you bring your own keys for each of the 34 providers.

That changes the failure modes. With a hosted gateway, quota changes and model retirements are handled upstream and you find out through their status. With FreeLLMAPI, you are the one holding 34 sets of credentials, and the catalog feed is how the project keeps you current. It also changes the cost model: OpenRouter takes its cut in the routing layer, while FreeLLMAPI's paid tier is a flat $19/yr for a same-day catalog rather than a per-token markup.

The custom provider slot is the other distinction. FreeLLMAPI lets you point chat, embedding, image or audio models at any OpenAI-compatible endpoint, naming llama.cpp, LM Studio, vLLM, a local Ollama, or a remote gateway. That makes it usable as a single front door for both free hosted tiers and your own local inference, which a hosted gateway cannot offer.

Licence, upgrades and the cost of keeping it running

The repository is MIT licensed, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is a permissive licence, but it says nothing about the terms of the 34 upstream providers. Each free tier has its own acceptable-use policy, and aggregating them behind one endpoint does not merge those terms. Whether your use of a given free tier complies with that provider's policy is a separate question from the MIT grant, and the README's personal-experimentation framing suggests the author is aware of the gap.

Upgrade cost depends on install method. Docker users pull a new image; the compose file does not pin a digest, so :latest moves. The desktop builds are distributed through GitHub releases, and the current release line is v0.9.x, with v0.9.8 pushed on 2026-09-07. Source installs are an npm workspace monorepo, and the server ships database migrations with npm run db:migration:up, db:migration:down, db:migration:fresh and db:migration:status, plus a dedicated test:migrations script. Migration tooling that includes a down path is a good sign for upgrade safety, though the README does not document a rollback procedure for a failed upgrade.

Editorial conclusion

Adopt FreeLLMAPI if you are an individual developer who wants free-tier capacity for side projects and local agents, and you accept that the free catalog snapshot lags the live feed by 30 days. Do not adopt it for multi-tenant or internet-facing production: the README calls it single-user and the Docker port binds to 127.0.0.1 by default. Before committing, verify that your Node version satisfies the engines field, that ENCRYPTION_KEY is set in production, and that the providers you actually need appear in the catalog your install receives.

Frequently asked questions

What is FreeLLMAPI?

It is a local server that aggregates the free tiers of 34 LLM providers, plus any custom OpenAI-compatible endpoint, behind a single OpenAI-compatible /v1 API. A router picks a model per request, fails over when a provider is rate-limited, and tracks per-key usage against free-tier caps.

How do I install FreeLLMAPI?

The README documents Docker Compose, a desktop app for macOS and Windows, and a Google Play build. The Docker path uses the ghcr.io/tashfeenahmed/freellmapi image on port 3001, bound to 127.0.0.1 by default, with an .env file copied from .env.example.

Which LLM API is free?

FreeLLMAPI does not make any single provider free; it aggregates the free tiers that providers already offer. The README puts the combined capacity at roughly 7.4 billion tokens per month across 474 model families and 635 endpoints, with each provider's own quota still applying.

Do free API keys exist for FreeLLMAPI?

You bring your own keys. FreeLLMAPI stores them encrypted and routes across whichever providers you have added, but it does not issue keys on a provider's behalf. The README's disclaimer limits the project to personal experimentation.

Can I use FreeLLMAPI with VS Code or Claude Code?

The README lists compatible CLIs and coding agents as a supported use case, and the repository ships an opencode.json at the top level. Because the endpoint is OpenAI-compatible, any client that accepts a custom base URL can point at the local server.

How does FreeLLMAPI differ from OmniRoute?

The repository does not describe OmniRoute, so a direct comparison is not possible from the available material. What FreeLLMAPI documents about itself is a self-hosted, single-user proxy with encrypted key storage and a signed catalog feed, versus a hosted gateway model.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. tashfeenahmed/freellmapi on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tashfeenahmed-freellmapi.svg)](https://hysenlabs.com/projects/tashfeenahmed-freellmapi)