# free-coding-models: a CLI that pings 228 free coding models before you pick one

> The tool probes free and free-limited LLM endpoints in parallel, ranks them by a latency and stability score, then writes the winner into your coding assistant's config. Here is how it installs, how the scoring works, and where it falls short.

**vava-nessa/free-coding-models** — Find, benchmark and install in CLI 170+ FREE coding LLM models across 15+ providers in real time

- Repository: https://github.com/vava-nessa/free-coding-models
- Website: https://freecodingmodels.vercel.app
- Stars: 2,791 · Forks: 297
- Language: HTML
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/vava-nessa-free-coding-models

## What free-coding-models actually decides for you

The catalog of free and free-limited coding models is large enough that picking one by reputation is guesswork. The README puts the current count at 24 providers and 228 live models, generated from sources.js. The project's argument is that average latency is a poor selector: a model that occasionally spikes to six seconds is worse than one with a slightly higher mean and no spikes. So the tool ranks candidates on a live Stability Score from 0 to 100, described as combining p95 latency, jitter, spike rate and uptime.

That score, not a leaderboard position, is the thing it sells. The audience is developers who already hold one or more free API keys and want to know which endpoint is healthy at this moment, then hand that endpoint to a coding assistant without editing config by hand. The README lists support for OpenCode CLI, Desktop and WebUI, OpenClaw, Crush, Goose, Aider, Kilo CLI, Qwen Code, OpenHands, Amp, Hermes, Continue, Cline, Xcode, Pi, ZCode, ForgeCode, Copilot, jcode and Caveman Code. If your tool is not on that list, the launch step has nowhere to write.

## Parallel probes, a 24h cache, and a score built from p95 and jitter

On first run the tool prompts for provider keys, then pings every model in parallel and lights rows green as responses arrive. A persistent 24-hour probe cache is shared across the TUI, the web dashboard, the agent extensions and the router, so switching surfaces does not restart the measurement. Health probes consume provider quota, which the README states plainly. The mitigation is automatic: a provider that returns 429 is paused, failing models back off exponentially, and a footer chip shows while a provider rests. Leaving a key empty still allows anonymous liveness probes for most providers.

Beyond latency, the project layers in AI Speed Test benchmarks that run real completions and report AI Latency plus TPS, a Smart Recommend wizard with three questions, live quota read from response headers, real-world telemetry scores, and models.dev enrichment with drift detection. The tier scale is anchored to SWE-bench Verified, from S+ at 70 percent and above down to C. Provider metadata lives in sources.js, and docs/providers.md is generated from it by node scripts/generate-provider-table.mjs, which is how the counts stay in sync rather than being edited by hand.

## Installing free-coding-models and launching your first model

The package needs Node.js 18 or later, has no native build step, and the README says it never needs sudo. Install it globally and read the flags before committing to a provider.

```bash
npm install -g free-coding-models
free-coding-models --help   # prints every flag
```

The help output is the authoritative list of flags, so treat it as the reference rather than copying options from a blog post. Next, obtain one free key. The README names Groq at console.groq.com/keys, Cerebras at cloud.cerebras.ai, described as the lowest latency in the catalog, and NVIDIA NIM at build.nvidia.com, described as the biggest no-credit-card quota. One key is enough to start; more can be added later with P inside the app.

Launching with no arguments opens the TUI and prompts for keys on first run, where Enter skips a provider. Rows light up green as models respond. From there, arrow keys navigate and Enter writes the selected model into your tool's config and launches it. You can pre-target a tool and filter from the command line instead.

```bash
free-coding-models --goose --tier S      # Goose, pre-filtered to S-tier only
free-coding-models --crush --origin groq # Crush, Groq models only
free-coding-models --fiable              # print the single most reliable model and exit
```

The --fiable form is the useful one for scripting: it prints the most reliable model and exits, so it can feed a shell variable rather than an interactive session. If you prefer a browser, free-coding-models web opens the web dashboard. For a single endpoint with auto-failover, free-coding-models --daemon-bg starts the Smart Model Router.

## Running the dashboard in Docker on port 19280

The repository ships a Dockerfile and a docker-compose.yml for the web dashboard. The image runs as a non-root user, listens on 19280, and sets FCM_HOST to 0.0.0.0 and FREE_CODING_MODELS_TELEMETRY to 0. A healthcheck polls http://127.0.0.1:19280/health every 30 seconds.

```yaml
services:
  fcm:
    image: ${FCM_IMAGE:-ghcr.io/vava-nessa/free-coding-models:latest}
    container_name: fcm
    restart: unless-stopped
    ports:
      - "19280:19280"
    environment:
      FREE_CODING_MODELS_TELEMETRY: "0"
      FCM_HOST: "0.0.0.0"
```

The compose file passes provider keys through from the host environment, including NVIDIA_API_KEY, GROQ_API_KEY, CEREBRAS_API_KEY, OPENROUTER_API_KEY, DASHSCOPE_API_KEY, OLLAMA_API_KEY and others. Its own comment points at src/core/config.js and its ENV_VARS map as the source of truth, and warns that a new provider must be added there and in the compose file. FCM_ALLOWED_ORIGINS controls which origins may reach the web UI from outside the container; leaving it empty is the conservative default. If you expose the dashboard beyond localhost, that variable is the one to set deliberately.

## Where free-coding-models is the wrong tool

Two constraints deserve attention before you build a workflow on it. First, probing is not free in the quota sense. The README states that health probes consume provider quota, and while automatic pausing and exponential backoff limit the damage, a provider with a small free allowance will see that allowance spent on measurement rather than on completions. If your free tier is measured in a handful of requests per day, a tool whose core loop is pinging every model is a poor fit.

Second, the value proposition depends on your coding tool being on the supported list. The one-key launch writes into a known config format. For an editor or agent that is not listed, you are back to copying a base URL and model identifier by hand, and the tool's main convenience disappears. There is also a structural dependency worth naming: the catalog is generated from sources.js, so a provider that changes its free tier, deprecates a model or starts requiring a credit card will not be reflected until that file is updated and a release ships. The version cadence is frequent, with v0.5.91 published on 2026-09-10 following v0.5.90 the day before, but that cadence is a property of the project, not a guarantee about any individual provider's terms.

## How it compares with OpenRouter's own model list

OpenRouter appears in the provider table with 19 models and its own free tier, and it publishes a browsable model list with pricing and context information. The difference is what each one measures. OpenRouter's list tells you what exists and what it costs; it does not ping endpoints from your machine and report p95 latency, jitter and spike rate for your network path. free-coding-models is a measurement layer that happens to include OpenRouter as one of 24 sources.

The trade-off runs the other way too. OpenRouter is a single account, a single key and a single billing relationship, and its availability is a vendor commitment. free-coding-models aggregates providers whose only shared property is that they currently offer something free, so a model disappearing from your ranked list is normal operation rather than an incident. If you want one endpoint with a support contact, the aggregator wins. If you want to know which of several free endpoints is answering quickly right now, the measurement tool wins, and the Smart Model Router with auto-failover is the project's own attempt to get both.

## Maintenance, licence and upgrade cost

The repository is not archived and the last push was on 2026-09-10, with releases v0.5.89 on 2026-09-07, v0.5.90 on 2026-09-09 and v0.5.91 on 2026-09-10. That is a fast release rhythm, and it carries a matching upgrade cost: provider metadata, tiers and benchmark data move together, and the generated provider table is regenerated rather than hand-edited, so local modifications to docs/providers.md will be overwritten. The npm package ships bin/, src/, web/, sources.js, the two patch-openclaw scripts and changelog/, and deliberately excludes web/src and web/public from the published files.

On licensing, package.json declares MIT, while the repository metadata reports NOASSERTION. Those two signals disagree, and the LICENSE file at the repository root is the document that governs. Read it before redistributing the package or bundling it into a product; this is a description of what the files say, not legal advice. The package also declares a prepack script, which matters if you install from a git checkout rather than from the registry.

## Conclusion

Adopt free-coding-models if you already juggle several free provider keys and want one ranked view plus a one-key write into your coding tool's config. Skip it if you need a single vendor's uptime commitment, or if you cannot tolerate probe traffic counting against provider quotas. Verify first that your target tool is on the supported list and that your provider's rate limits survive repeated health probes.

## FAQ

### Which is the best free model for coding?

free-coding-models does not publish a single permanent winner. It ranks models by a live Stability Score from 0 to 100 that combines p95 latency, jitter, spike rate and uptime, and the ranking changes as endpoints respond. The README notes that average latency alone is misleading, so the top of the list is a current measurement rather than a fixed recommendation.

### Which AI model is totally free?

The project tracks free and free-limited models across 24 providers and 228 live models, generated from sources.js. The README does not claim any of them are unconditionally free; it describes free-tier limits in docs/providers.md and notes that health probes consume provider quota, which is why it auto-pauses a provider on a 429 response.

### What is the cheapest model for coding?

The tool is aimed at free and free-limited endpoints rather than at price comparison, so it does not rank models by cost. It reports live latency, a Stability Score, tier ratings anchored to SWE-bench Verified and quota read from response headers. For pricing across paid models, the provider's own console is the place to look.

## Sources

- [Issues](https://github.com/vava-nessa/free-coding-models/issues)
- [Project website](https://freecodingmodels.vercel.app)
- [README](https://github.com/vava-nessa/free-coding-models/blob/main/README.md)
- [Releases](https://github.com/vava-nessa/free-coding-models/releases)
- [vava-nessa/free-coding-models on GitHub](https://github.com/vava-nessa/free-coding-models)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/vava-nessa-free-coding-models
