Model or dataset
signerless/llm-checker avatar
signerless/llm-checker

llm-checker: pick the right local LLM for your exact hardware

Advanced CLI tool that scans your hardware and tells you exactly which LLM or sLLM models you can run locally, with full Ollama integration.

3,000 stars202 forksJavaScriptNOASSERTION

At a glance

What is it?
llm-checker is a Node.js CLI that scores Ollama-compatible models against your detected GPU, VRAM and RAM, then hands back a ranked list and the commands to pull them. The scoring is deterministic and the catalog is a packaged snapshot, which is both its selling point and its main constraint.
Who is it for?
Adopt llm-checker if you run Ollama on one machine and keep guessing which quantized model will actually fit in VRAM; the hw-detect plus recommend flow answers that in two commands. Skip it if you need a benchmark of real tokens per second on your workload, or if you are not on Ollama at all, because the catalog and the run path are built around that runtime.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 25 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem llm-checker targets: model choice on fixed hardware

The README frames the problem plainly: thousands of model variants and quantization levels exist, and picking one requires understanding memory bandwidth and VRAM limits. That is the gap this tool fills. It is aimed at a single-machine user, someone with one laptop or one workstation who wants to know what will run before downloading tens of gigabytes. The command surface reflects that: hw-detect, recommend, sync, and ai-run. The README's own two-minute flow puts hw-detect first and recommend second, which tells you the intended order of operations. It is not a serving framework, not an evaluation harness, and not a multi-node scheduler. If you already know your hardware budget and you just want to pull a model, this is a convenience layer rather than a necessity.

How the scoring and the packaged registry actually work

The mechanism is a deterministic scoring pass, not a model call. The README describes scoring across four dimensions it names Quality, Speed, Fit and Context, weighted by use case, with a --category flag selecting the weighting. Fit is where the hardware detection feeds in: the tool reads Apple Silicon, NVIDIA CUDA, AMD ROCm, Intel Arc or plain CPU, and estimates memory with what the README calls a bytes-per-parameter formula validated against real Ollama sizes. Those estimates are then matched against a packaged catalog. According to the README the package ships a synced Ollama SQLite catalog of 200+ models plus a multi-source registry of roughly 33,700 exact artifacts from Hugging Face, Ollama and GPT4All, with per-source commands and runtime targeting. Two details matter here. First, the catalog is a snapshot inside the package, so without running sync your recommendations reflect the version you installed, not the current Ollama library. Second, the scoring is described as deterministic, which means the same hardware and the same catalog produce the same ranking. That is reproducible and it is also only as good as the memory formula behind it. The README notes the formula was validated against real Ollama sizes, but it does not publish the validation results.

Installing llm-checker and running a first recommendation

Installation is a global npm install. The README lists Node.js 18+ as the requirement, with 18, 20 and 22 as tested release lines, and notes Ollama must be installed separately for actually running models. The package declares sql.js as an optional dependency; if your package manager skips optional dependencies, the README says database commands will report sql.js missing and you should reinstall with optional dependencies enabled.

bash
npm install -g llm-checker

If you would rather not install globally, the README gives npx as an alternative and uses hw-detect as the example. Either way, the first command to run is hardware detection, because every later score depends on it.

bash
npx llm-checker hw-detect

You should see your detected accelerator and memory reported back. From there, ask for a category. The README uses coding as its example category, and the output is a ranked set of recommendations with pull and run commands.

bash
llm-checker recommend --category coding

When you want the catalog to reflect the current Ollama library rather than the snapshot shipped in the package, the README points to sync.

bash
llm-checker sync

The end-to-end path is ai-run, which selects a model for the category and launches it, reporting tokens per second next to the output. The README's example prompt is a hello world in Python.

bash
llm-checker ai-run --category coding --prompt "Write a hello world in Python"

One behaviour worth knowing before you run any of this: recommendations and auto-selection exclude models labelled uncensored, abliterated or heretic by default. The README states that opting in requires --include-uncensored, and that ai-run needs the same flag before it will select or launch one of those installed models. If you are troubleshooting an install that reports sql.js missing, the documented fix is a reinstall with the optional flag.

bash
npm install -g llm-checker --include=optional

Where llm-checker gets it wrong or does not apply

The tokens-per-second figure from ai-run is a live measurement of one run, not a benchmark, and the README does not present it as anything more. Treat it as a sanity check, not a capacity plan. The deeper limitation is the memory estimation itself. A bytes-per-parameter formula is a model of reality; it does not account for the KV cache growth at long context, for concurrent requests, or for whatever else is holding VRAM on your desktop. The README says the formula was validated against real Ollama sizes, which is a claim about typical model files, not about your session. Expect the boundary cases, very long contexts and multi-model setups, to be where a recommendation that scored well still fails to load. There is a second, harder boundary: this is an Ollama-shaped tool. The registry includes Hugging Face and GPT4All artifacts with per-source commands, but the run path and the catalog sync are built around Ollama. If your inference stack is llama.cpp directly, vLLM, or LM Studio, the recommendation half may still be useful and the execution half is not. Finally, the README documents no rollback procedure for sync, so if a refresh pulls in a catalog you dislike, the documented path back is reinstalling a pinned version rather than undoing the sync.

llm-checker compared with llmfit

The README addresses this comparison directly rather than leaving it to the reader. It positions llm-checker as hardware-aware model selection for local inference, producing ranked recommendations, compatibility scores and pull or run commands. It describes llmfit as broader LLM workflow support and model-fit evaluation from a different angle, with a different optimization workflow and selection heuristics. The practical difference the README draws is intent: if the question is what should I run on this exact machine right now, use llm-checker; if the goal is experimentation across custom pipelines, the README says the two can be complementary. Read that as an honest split rather than a dismissal. llm-checker is narrow on purpose, and the narrowness is what makes the output a single ranked list instead of a configuration surface. If your work is pipeline plumbing rather than machine-specific model choice, the comparison table is telling you that this is not the tool for that job.

Licence, distribution and the cost of staying current

The licence field on the repository is NOASSERTION, and the README badge points at NPDL-1.0 while the LICENSE file is what actually governs. That mismatch is worth resolving yourself before you vendor anything from this project into a product. The README does clarify one dependency question: the structural verification behind verify, ai-run --verify, structural policy validation and the MCP verify_model tool comes from ModelVet, created by Tetsuo AI, and llm-checker ships that WebAssembly integration under ModelVet's MIT license. There is a THIRD_PARTY_NOTICES file in the repository root, which is where the remaining attributions live. On distribution, the README is explicit that npm is the recommended channel for the newest release, that GitHub Releases carries the history, and that the scoped @pavelevich/llm-checker GitHub Packages mirror is legacy and may lag. It even gives the fix for a stale scoped install: uninstall the scoped package, install llm-checker@latest, rehash the shell, and check --version. The upgrade cost is the catalog. Because a snapshot ships inside the package, npm updates and sync refreshes are two separate maintenance actions, and the last push to the repository was on 2026-09-06.

Editorial conclusion

Adopt llm-checker if you run Ollama on one machine and keep guessing which quantized model will actually fit in VRAM; the hw-detect plus recommend flow answers that in two commands. Skip it if you need a benchmark of real tokens per second on your workload, or if you are not on Ollama at all, because the catalog and the run path are built around that runtime. Before trusting a recommendation, run llm-checker hw-detect and check that the reported VRAM and accelerator match what your OS reports, since every score downstream is computed from that single detection pass.

Frequently asked questions

How do I install llm-checker?

Install it globally with npm install -g llm-checker, or run it without installing via npx llm-checker hw-detect. Node.js 18 or newer is required, and Ollama must be installed separately if you want to run models.

How do I use llm-checker?

The README's two-minute flow is hw-detect, then recommend --category coding, then sync when you want current Ollama references, then ai-run with a prompt. The ai-run command selects a model and reports tokens per second alongside the output.

What is llm-checker?

It is a CLI that analyzes your hardware and recommends which local LLM or sLLM models will run on it, using a deterministic four-dimension score over a packaged multi-source registry and the Ollama catalog. It integrates with Ollama to pull and run the models it recommends.

How does llm-checker differ from llmfit?

The README describes llm-checker as hardware-aware model selection for local inference, returning ranked recommendations and pull or run commands, while it characterizes llmfit as broader workflow support and model-fit evaluation with different selection heuristics. It suggests using llm-checker when the question is what to run on this machine right now.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. signerless/llm-checker on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/signerless-llm-checker.svg)](https://hysenlabs.com/projects/signerless-llm-checker)