# llmfit: matching local LLM models to your actual RAM, CPU and GPU

> llmfit is a Rust terminal tool that detects your hardware and ranks a catalog of LLM models by memory fit, estimated speed, quality and context. It estimates rather than measures, and the README is upfront about that.

**AlexsJones/llmfit** — llmfit checks local hardware against model requirements and recommends models and providers that fit available memory and compute.

- Repository: https://github.com/AlexsJones/llmfit
- Stars: 37,157 · Forks: 2,368
- Language: Rust
- License: MIT
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/alexsjones-llmfit

## The question llmfit answers before you download a 40 GB checkpoint

The failure mode this targets is mundane and expensive: you pick a model from a leaderboard, start the download, and discover it either will not load or loads so slowly that it is unusable. llmfit inverts the order. It reads your RAM, CPU, GPU and VRAM, then scores every model in its catalog across four dimensions the README names as memory fit, estimated speed, quality and context, and ranks them.

The intended user is someone running models locally on hardware they already own, not someone provisioning a cluster. That shows in the platform surface: Scoop for Windows, Homebrew and MacPorts for macOS and Linux, a curl install script, a uv tool install, and a container image. Multi-GPU setups and MoE architectures are called out as supported, which matters because a mixture-of-experts model's resident memory is not its total parameter count. The README explicitly notes that llm-checker, its own listed alternative, treats every model as dense and therefore overstates memory for models like Mixtral or DeepSeek-V3.

What it is not is a serving stack. It recommends and it measures; it does not host. The sibling projects named in the README, llmserve and llama-panel, cover the serving side.

## Hardware detection, a four-axis score, and where the numbers come from

The pipeline is detect, then score, then rank. Detection produces RAM, CPU, GPU/VRAM and backend. Scoring runs each catalog entry through four dimensions. Ranking is what the TUI displays at the top of the screen and what the CLI prints as a table.

The interesting part is the speed estimate. According to the README, it comes from a memory-bandwidth model grounded in runtime sampling and real community measurements, and every estimate ships its inputs, so llmfit info shows what a number assumes and how to verify it on the machine in front of you. That last property is the design decision worth noting: rather than presenting a single opaque tok/s figure, the tool exposes its assumptions. A bandwidth-derived estimate is only as good as the bandwidth figure it was seeded with, and on a machine with unusual memory or a partially supported GPU that assumption is where the error lives.

The newer benchmark-and-share feature closes the loop. You download a model, serve it, measure real tok/s on your hardware, and submit the result as a pull request from inside the TUI. The README states that every run is saved locally first, that your own measurements replace estimates in the fit table, and that merged submissions ship in the next release so anyone on identical hardware sees measured numbers rather than estimates. The contribution path is a PR, not a third-party account or the gh CLI.

The workspace splits into llmfit-core, llmfit-tui and llmfit-desktop, with the first two as default members. A Dockerfile builds web assets with Node 20 before the Rust stages, which is worth knowing if you build the image yourself: the container is not a pure Rust build.

## Installing llmfit on Windows, macOS, Ubuntu and Docker

On Windows the README gives Scoop as the install path. If Scoop is not present, the README points at the Scoop installation guide rather than offering a manual binary drop.

```bash
scoop install llmfit
```

On macOS and Linux, Homebrew is described as the recommended prebuilt route. There are two formulae, and the difference is not cosmetic: the tap installs a prebuilt binary, while the homebrew-core formula builds from source on macOS versions that have no bottle available.

```bash
brew install AlexsJones/llmfit/llmfit
```

Ubuntu users can take the same Homebrew route, or the quick install script, which downloads the latest release binary from GitHub and places it in /usr/local/bin, falling back to ~/.local/bin when there is no sudo. The script takes a --local flag to force the no-sudo location.

```bash
curl -fsSL https://llmfit.axjns.dev/install.sh | sh -s -- --local
```

If you prefer a Python-managed tool, uv installs and updates it, and uvx runs it without installing. The README also says it can be installed as a normal Python package with pip or uv.

```bash
uv tool install -U llmfit
```

Docker is the odd one out: the image defaults to printing JSON from the recommend command rather than opening the TUI. To get the interactive interface you pass the global --tui flag.

```bash
docker run ghcr.io/alexsjones/llmfit
```

That default is deliberate and useful in pipelines, since the JSON can be filtered with jq. The README's own example pipes recommend --use-case coding into jq to list model names. To launch the TUI in a container instead, the README shows passing --tui with an interactive terminal.

```bash
docker run --rm -it ghcr.io/alexsjones/llmfit --tui
```

For a first real use, run the binary with no arguments. The README says that opens the TUI with your detected specs at the top and every model scored. If you would rather stay in the shell, the CLI subcommands are fit for a ranked table, recommend --json for machine consumption, info for one model including its estimate basis and verify commands, bench to measure against a running provider, and doctor for a hardware detection report intended for bug reports. The doctor output is the first thing to check when a recommendation looks wrong, because a misdetected GPU or backend invalidates everything downstream.

## Where llmfit gets it wrong, and when it is the wrong tool

The catalog is a snapshot. The Makefile exposes make update-models, which fetches model data from HuggingFace, and make update-docker-models, which scrapes the Docker Model Runner catalog from Docker Hub. Both are maintainer-facing scripts that regenerate embedded data, not something a user runs against an installed binary. So a model released after your build is simply absent until you upgrade, and the README does not document a way to refresh the catalog in place.

Estimates are estimates. The README is unusually direct about this: speed comes from a memory-bandwidth model, and the benchmark feature exists precisely because estimates are not measurements. If your decision hinges on whether a model hits a specific tokens-per-second threshold on your exact machine, llmfit's default output cannot settle it. Run llmfit bench against a running provider, or treat the estimate as a filter and measure the survivors.

Hardware detection is the other soft spot. Anything the detector misreads propagates: a GPU that is not recognized, an unusual backend, or a container without device passthrough. The Docker image is a good illustration. Running it without GPU access means it sees the container's view of the machine, not the host's, and the README does not document a GPU passthrough invocation for the container. If you are evaluating a machine you are not sitting at, the detection report is only as trustworthy as the environment it ran in.

Finally, this is not a serving tool. If your question is how to run a model, llmfit answers which one, not how. The README points at llmserve and llama-panel for that, and at the runtime providers documentation for Ollama, llama.cpp, MLX, Docker Model Runner and LM Studio integration.

## llmfit compared with llm-checker: estimate first or run first

The README names one alternative directly: llm-checker, a Node.js CLI with Ollama integration. The difference in approach is the whole argument. llm-checker pulls and benchmarks models by actually running them through Ollama on your hardware, which the README describes as a more hands-on approach. llmfit estimates from specs and offers benchmarking as an opt-in step, with the results contributed back to a shared catalog.

That produces opposite cost profiles. llm-checker spends disk and time before it answers; you cannot get a verdict on a 70B model without downloading it. llmfit answers immediately from detection plus catalog data, which is what makes it usable across several machines in one sitting, but the answer is only as good as the bandwidth model behind it. The README also flags a concrete correctness gap in the alternative: llm-checker does not support MoE architectures and treats all models as dense, so its memory estimates for Mixtral or DeepSeek-V3 reflect total parameter count rather than the active subset.

If you already run Ollama and want ground truth on one model, llm-checker's approach is the more direct route. If you are triaging a catalog against unfamiliar hardware, or you want a JSON list your automation can consume, llmfit's estimate-first design is the better fit. The two are not exclusive: llmfit's bench subcommand measures against your running provider, so the estimate-first path can end in a real number.

## Licence, release cadence and what an upgrade actually costs

llmfit is MIT licensed, and the workspace Cargo.toml declares the same licence for the workspace package. MIT is permissive, so redistribution and modification are permitted subject to the licence text; that is a description of the terms, not legal advice, and anyone embedding it in a commercial product should read the LICENSE file rather than this summary.

The release history is dense. Three releases landed between 2026-08-17 and 2026-08-28, and the last push to main was on 2026-08-28. The repository is not archived. That cadence is a real cost signal: if you pin a version, expect the catalog to age quickly, and expect the upgrade path to be the install command you already used. Homebrew, Scoop, uv and the curl script all fetch the latest release, so the friction is low, but there is no documented in-place catalog refresh for an installed binary.

One upgrade detail is worth flagging for anyone building from source or from the Dockerfile. The Dockerfile comments state that rustc 1.95 or newer is required because sysinfo 0.39.x raised its minimum supported Rust version, and that both build and runtime stages are pinned to Debian bookworm to avoid a GLIBC mismatch that makes the binary fail to start with a GLIBC_2.39 not found error. The same comments record that a previous release failed because a BUILDPLATFORM declaration pulled the amd64 toolchain onto an arm64 runner. None of this affects installing a prebuilt binary, but it does mean the container build is more fragile than the CLI's install story suggests.

Windows binaries are signed via SignPath.io with a certificate from the SignPath Foundation, and the README says signing happens automatically in the release pipeline.

## Conclusion

Adopt llmfit if you are picking local models across more than one machine and want a single ranked answer plus a JSON output your scripts can consume. Skip it if you need measured throughput before you commit: the catalog numbers are estimates until you run llmfit bench, and the benchmarking guide describes contributing those results back. Before trusting any recommendation, run llmfit doctor and confirm the detected RAM, GPU and backend match the box you are actually on, then run llmfit info on the one model you care about and read the estimate basis it prints.

## FAQ

### How to install llmfit?

The README lists Scoop on Windows, Homebrew or MacPorts on macOS and Linux, a curl install script that places the binary in /usr/local/bin or ~/.local/bin, uv tool install -U llmfit, or a container image. Building from source is a cargo build --release in a cloned repository, with the binary at target/release/llmfit.

### How do I uninstall llmfit?

The README does not document an uninstall procedure. The removal step follows from whichever install route you used: Scoop, Homebrew, MacPorts, uv and the curl script each own their own files, and the curl script's binary lands in /usr/local/bin or ~/.local/bin.

### How can llmfit help me find LLM models that run on my hardware?

It detects your RAM, CPU, GPU/VRAM and backend, then scores every catalog model across memory fit, estimated speed, quality and context, and ranks them. The TUI shows the ranked list against your detected specs, and llmfit fit prints the same as a table.

### How to use llmfit?

Running llmfit with no arguments opens the interactive TUI. For scripts there are subcommands: fit for a ranked table, recommend --json for machine-readable output, info for a single model with its estimate basis, bench to measure against a running provider, and doctor for a hardware detection report.

### Is llmfit safe?

The repository is MIT licensed and the Windows release binaries are digitally signed via SignPath.io with a certificate from the SignPath Foundation. The curl install script downloads a release binary from GitHub, so anyone using that route should be comfortable piping a remote script into a shell.

### How to install llmfit on Windows?

The README gives a single command, scoop install llmfit, and points at the Scoop installation guide if Scoop is not already present. Windows release binaries are Authenticode signed through SignPath.io.

## Sources

- [Official README](https://github.com/AlexsJones/llmfit#readme)
- [Project repository](https://github.com/AlexsJones/llmfit)
- [Release notes](https://github.com/AlexsJones/llmfit/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/alexsjones-llmfit
