# whichllm: Hardware-Aware Local LLM Ranking from Real Benchmarks

> whichllm is a Python CLI that auto-detects your GPU and RAM, then ranks local LLMs from HuggingFace by combined benchmark score and estimated inference speed rather than by model size alone. It runs without a project setup and also simulates any hardware configuration before you buy.

**Andyyyy64/whichllm** — Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.

- Repository: https://github.com/Andyyyy64/whichllm
- Stars: 6,654 · Forks: 366
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/andyyyy64-whichllm

## The Problem whichllm Solves

When choosing a local LLM to run on a specific GPU, most tools answer the wrong question. They tell you which models fit in VRAM, not which of those models is actually the best. The README states directly: fitting a model into your VRAM is the easy part; the hard part is knowing which of the models that fit is the best.

whichllm is a command-line tool for engineers and researchers who run local models and want a ranked list that weighs benchmark quality alongside hardware constraints. It targets users of NVIDIA, AMD, Intel, and Apple Silicon hardware, and handles CPU-only setups. Version 0.5.19 is the current release, available on PyPI and via Homebrew.

## How the Ranking Works

The ranking algorithm draws from several benchmark sources: LiveBench, Artificial Analysis, Aider, multimodal/vision evaluations, Chatbot Arena ELO, and the Open LLM Leaderboard. The README describes the ranking as evidence-graded: every score is tagged with one of five confidence levels (direct, variant, base, interpolated, or self-reported) and discounted by that confidence. Self-reported scores from model uploaders and scores borrowed across family boundaries (a small fine-tune claiming its much larger base model's score) are actively rejected.

The ranking is also recency-aware. Stale leaderboard entries are demoted along a model's lineage, so a 2024 model cannot outrank a current-generation one by an outdated score. The date of the benchmark data is printed beneath every ranking so an outdated recommendation is visible rather than silently trusted.

The README gives a concrete example using an RTX 4090: the 32B model fits the card, but whichllm ranks the 27B model first because it scores higher on real benchmarks and represents a newer generation. A size-only tool would surface the larger model; whichllm does not.

## Installing and Running whichllm

The recommended one-off invocation requires no install:

```bash
uvx whichllm@latest
```

To simulate a specific GPU before buying hardware:

```bash
uvx whichllm@latest --gpu "RTX 4090"
```

For repeated use, install as a persistent tool:

```bash
uv tool install whichllm
uv tool upgrade whichllm
```

Alternative install paths are available:

```bash
brew install andyyyy64/whichllm/whichllm
pip install whichllm
```

The package requires Python 3.11 or higher and is distributed as `whichllm` on PyPI at version 0.5.19. Dependencies include typer, rich, httpx, psutil, dbgpu, and nvidia-ml-py.

## Common Workflows and Output Flags

After install, running `whichllm` without arguments outputs a ranked list for the detected hardware. The output shows model name, parameter count, quantization format, benchmark score, and estimated tokens per second. Speed is color-coded by practical usability: under 4 t/s is red, 4-10 is yellow, 10-30 is green, and 30-plus is bright green.

For safer LM Studio-style recommendations that exclude partial-offload candidates and slow models:

```bash
uvx whichllm@latest --gpu-only --speed usable --vram-headroom 1GB
```

Other useful invocations from the README:

```bash
whichllm --vram 8 --ram-bandwidth 68
whichllm --gpu "2x RTX 4090"
whichllm --speed fast
whichllm --markdown
whichllm upgrade "RTX 4090" "RTX 5090" "H100"
```

`--markdown` outputs a GitHub-flavored Markdown table suitable for pasting into a Slack or GitHub discussion. `upgrade` compares the current hardware against candidate GPUs. `--json` enables pipeline-friendly output for use with tools like jq. The `whichllm run` subcommand downloads and starts a chat session with the top-ranked model directly.

## Architecture-Aware VRAM and Speed Estimation

The README describes the VRAM estimation formula as: weights plus GQA KV cache plus activations plus overhead. This is more precise than a simple weights-only estimate and accounts for the memory the runtime needs beyond the raw model parameters.

Speed estimation is bandwidth-bound with per-quantization efficiency factors, per-backend factors (such as whether the runtime is optimized for the GPU type), and a MoE active-versus-total split. For MoE models, speed is ranked on active parameters while quality is ranked on total parameters, since the two numbers serve different roles. Unified-memory architectures (Apple Silicon) and discrete PCIe partial-offload configurations are modeled separately.

The --vram-headroom flag adds a safety margin on top of the estimate, which is useful when LM Studio or another runtime reports a model is slightly too large despite whichllm's estimate showing it fits. Increasing the headroom value by 0.5 GB increments is the documented approach for resolving edge fits.

## Limitations

whichllm does not run any model; it only recommends one. Users still need a separate runtime such as Ollama or llama.cpp to actually load and run the selected model. The `whichllm run` shortcut creates an isolated environment via `uv`, but this is a convenience wrapper rather than a production inference server.

The benchmark data is fetched live from HuggingFace at runtime and falls back to curated frozen data when the API is unavailable or rate-limited. This means that on a rate-limited network, recommendations may be based on the frozen fallback data rather than the latest published scores. The printed benchmark date makes this visible.

Hardware auto-detection covers NVIDIA, AMD, Intel, and Apple Silicon, but the accuracy of iGPU and unified-memory detection may require manual correction using `--vram` and `--ram-bandwidth` overrides. The README specifically notes this use case.

## Comparison with LM Studio

LM Studio is a desktop GUI for downloading, managing, and running local LLMs. The README explicitly uses LM Studio as a reference point for the `--gpu-only --speed usable` flag set, describing this combination as a "more comfortable LM Studio-style recommendation."

The architectural difference is that LM Studio filters models by what fits in VRAM and presents them visually for the user to choose. whichllm ranks the models that fit by their benchmark quality and estimated speed, then prints the best one. LM Studio also runs models and provides an OpenAI-compatible server; whichllm only selects a model.

For users who prefer a graphical interface and want to also run the model from the same tool, LM Studio is the more complete choice. For users who want a scriptable, benchmark-ranked recommendation as part of a CLI workflow or hardware-planning pipeline, whichllm covers that gap.

## Maintenance and MIT License

The latest release is v0.5.19, published on 2026-09-19. The last code push was on 2026-09-14. The repository is not archived and shows active release activity. The project is released under the MIT license, which permits use, modification, and redistribution without restriction beyond attribution.

The project is classified as Development Status 4 (Beta) in its PyPI metadata. The pyproject.toml declares required-version for the uv build tool, pinning uv to 0.11.33, which means certain build workflows may need that specific version of uv to reproduce the package exactly.

## Conclusion

whichllm is the right tool for engineers who run local LLMs and want a principled ranking that goes beyond parameter count or VRAM fit. It is not the right fit for users who want a GUI, for teams using cloud APIs only, or for workflows that require running the model itself rather than just selecting one. Before relying on the ranking for a new hardware configuration, run it once to check the benchmark data date printed in the output, since the live HuggingFace data and benchmark scores that underpin the recommendations are fetched at runtime and may differ from what is shown in static documentation.

## FAQ

### Which local LLM is best for a 24GB VRAM GPU?

Run `uvx whichllm@latest --gpu "RTX 3090"` or `--gpu "RTX 4090"` (both are 24 GB cards) to get a ranked list. The README's example snapshot for an RTX 4090 shows Qwen3.6-27B at Q5_K_M as the top pick with a score of 92.8 and approximately 27 tokens per second, ahead of the larger Qwen3-32B.

### Which LLM is best for running on CPU-only hardware?

Run `whichllm` on a CPU-only machine and the tool auto-detects the absence of a discrete GPU, ranking models by what is feasible with RAM bandwidth alone. The README's snapshot shows a MoE model at Q4_K_M as the CPU-only top pick, with an estimated 6 tokens per second.

### How to use whichllm?

Run `uvx whichllm@latest` for a one-off recommendation against your detected hardware, or install with `uv tool install whichllm` for repeated use. Add `--gpu "RTX 4090"` to simulate a specific card, `--gpu-only` to exclude partial-offload candidates, and `--speed usable` to hide models that would be too slow to be practical.

### How to install whichllm?

The quickest path is `uvx whichllm@latest`, which runs without a permanent install. For persistent use: `uv tool install whichllm`, `brew install andyyyy64/whichllm/whichllm`, or `pip install whichllm`. Python 3.11 or higher is required.

## Sources

- [Andyyyy64/whichllm on GitHub](https://github.com/Andyyyy64/whichllm)
- [Issues](https://github.com/Andyyyy64/whichllm/issues)
- [License: MIT](https://github.com/Andyyyy64/whichllm/blob/main/LICENSE)
- [README](https://github.com/Andyyyy64/whichllm/blob/main/README.md)
- [Releases](https://github.com/Andyyyy64/whichllm/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/andyyyy64-whichllm
