# Ollama Grid Search: Comparing Local LLMs on macOS, Windows and Linux

> A Tauri desktop app that fans one prompt across several Ollama models and parameter values, then lets you read the responses side by side. It is a comparison harness, not a benchmark suite.

**dezoito/ollama-grid-search** — A multi-platform desktop application to evaluate and compare LLM models, written in Rust and React.

- Repository: https://github.com/dezoito/ollama-grid-search
- Stars: 954 · Forks: 57
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/dezoito-ollama-grid-search

## The comparison problem Ollama Grid Search was built around

Picking a model for a local task is usually done by hand. You pull a few tags, run ollama run against each, paste the same prompt into a terminal, and try to remember which answer you liked. Once the prompt has a system message, a temperature and a repeat penalty attached, the comparison stops being reproducible. You are no longer comparing models, you are comparing whatever settings you happened to type that evening.

Ollama Grid Search turns that into a matrix. The README describes the goal as automating "the process of selecting the best models, prompts, or inference parameters for a given use-case, allowing you to iterate over their combinations and to visually inspect the results." The prompt is submitted once for each parameter value, for each selected model, so a run of three models against two temperature values produces six responses in one view.

The audience is narrow and specific. You need Ollama running, either on localhost or on a remote server, and you need a desktop machine to run the app on. The README states the project assumes Ollama is installed and serving endpoints. If your models live behind a hosted API, this tool has nothing to point at.

The README is also honest about the name. It notes that grid search in machine learning usually means tuning training hyperparameters such as batch_size, learning_rate or number_of_epochs, and that the concept here is only similar. What is being swept is inference configuration, not training. That distinction matters when you search for the term and land on scikit-learn documentation instead.

## How the app is put together and where inference actually happens

The repository is a Tauri application. The front end is React 18 with Vite, TypeScript, Tailwind, Radix UI primitives, Jotai for state and TanStack Query for data fetching. The native shell lives in src-tauri, which is where the Rust side and the packaging configuration sit. Inference is not performed by the app itself. The app talks to an Ollama server over its HTTP API, fetches the list of available models, and issues the generation calls.

That architecture explains most of the behaviour you will notice. Because model discovery is a live call, the model picker reflects whatever the configured server currently has pulled, and the README notes that model selection can be filtered by name. Because generation is also a live call, the app can offer limited concurrency or synchronous inference calls, which the README frames as a way to prevent spamming servers. A remote Ollama box shared with other people is exactly the case where you want the synchronous option.

Results are not just text. The README lists optional output of inference parameters and response metadata covering inference time, tokens and tokens per second. Individual inference calls can be refetched, which is useful when one response in a grid comes back empty or truncated and you want to know whether the model or the transport was at fault.

Experiments are persisted rather than discarded. You can list them, inspect them in readable views, download them as JSON, and re-run a past experiment while cloning or modifying the parameters that were used. There is a configurable inference timeout, and settings hold custom default parameters and system prompts. A prompt archive supports saving and managing prompts, with autocomplete triggered by typing "/" in the inputs, a pattern the README says was inspired by Open WebUI.

## Installing from the releases page and running a first grid

There is no package manager install for the app itself. The README points at the releases page, which carries builds for macOS, Windows and Linux. Download the artifact for your platform, install it, and make sure Ollama is already serving before you launch.

On Apple Silicon there is one extra step, because the build is ad-hoc signed rather than notarized with a paid Apple Developer certificate. The README gives this sequence: download the aarch64 .dmg, drag the app to Applications, then right-click or Control+click the app and choose Open, and confirm Open in the security dialog. After that first launch, double-clicking works normally.

Once the app is open, the first thing to check is the server URL. If Ollama is on the same machine, the default localhost endpoint applies. If you run llmman, which the README describes as a local model runner that serves the Ollama API as a drop-in replacement, open Settings and set the Ollama Server URL to:

```bash
http://localhost:17434
```

Models pulled through llmman then appear in the model list and can be selected like any others. This is the only non-default endpoint the README documents, so treat it as the worked example rather than a general recipe.

With a server reachable, the first real experiment is small on purpose. Select two or three models, write one prompt, and add a single parameter with two values, for instance a temperature pair. The README's own illustration is a simple prompt tested on three different models. Run it, then open the experiment list and download the JSON. Reading the exported file is the fastest way to confirm which model, prompt and parameter values actually produced each row.

If you prefer to build from source rather than use a release artifact, the repository is a Vite project driven by Bun, and the development notes under docs/ carry setup instructions, sequence diagrams and workflow charts. The scripts in package.json are the standard ones:

```bash
bun install
bun run tauri dev
```

The README directs contributors with feature ideas to open an issue for discussion before writing a pull request, which is worth respecting if you plan to change behaviour rather than fix a typo.

## Where the grid approach breaks down

The tool compares outputs. It does not score them. The README lists grading results and filtering by grade under future features, which means that as of the current documentation there is no built-in way to mark a response as good and have the app sort by that judgement. For a handful of models and a short prompt, reading six answers yourself is fine. For a grid of ten models against five parameter values, you have fifty responses and no ranking mechanism inside the app. You will end up exporting JSON and doing the comparison elsewhere.

Cost is real even though no API bill arrives. Every cell in the grid is a full generation on your own hardware. The README offers limited concurrency or synchronous calls precisely because a large grid can hammer a server, and the configurable inference timeout exists because some combinations will take a long time. On a shared remote Ollama instance, a wide grid is a good way to make yourself unpopular.

The scope is Ollama's API. Anything that does not expose that interface is out of reach, and the llmman entry exists only because that runner implements the same API. If your evaluation involves a hosted model, a different inference server, or a mix of local and remote providers, this is the wrong tool and no setting will change that.

There is also a naming trap. Searching for grid search will surface training-time hyperparameter tuning material, and the README itself flags that the term is used loosely here. Nothing in this app touches batch size, learning rate or epoch counts. It sweeps inference parameters and prompts against already-trained models.

## Ollama Grid Search compared with promptfoo and lm-evaluation-harness

The closest comparison in spirit is promptfoo, which also runs one prompt across multiple models and parameter sets. The difference is where the work happens and what you get back. promptfoo is a CLI and config-file tool that targets many providers, including hosted APIs, and its central abstraction is an assertion: you declare what a passing output looks like, and the run reports pass and fail per case. Ollama Grid Search is a desktop GUI whose central abstraction is the visual grid, with results inspected by eye and exported as JSON. If your evaluation criteria can be written as code, promptfoo gives you automation and a pass rate. If your criteria are taste, and you want to see the answers next to each other while you decide, the grid view is the more direct instrument.

lm-evaluation-harness sits at a different level entirely. It runs standardized academic benchmarks with fixed datasets and published scoring, which is what you want when the question is how a model performs in general. Ollama Grid Search answers a narrower and more practical question: how does this model handle my prompt, with my system message, at this temperature. The two are not substitutes. A model can score well on a public benchmark and still produce the wrong shape of answer for your specific task, which is the gap the grid view is designed to expose.

Against both, the distinguishing feature here is that it is a native desktop application. There is no config file to write and no test suite to maintain. That is a genuine convenience for a one-off comparison and a genuine limitation for anything you want to run on a schedule.

## Maintenance, licence and the upgrade path

The repository is not archived, and the last push was on 2026-09-07. Releases have been spaced out: v0.9.0 in December 2024, v0.9.1 in April 2025, and v0.9.2 in November 2025. The project version in package.json matches the latest release tag at 0.9.2, so the tagged artifacts and the source tree are in step. The CHANGELOG.md at the repository root is where release-by-release detail lives; the README does not document a rollback procedure, so if an upgrade misbehaves, downgrading means reinstalling an earlier artifact from the releases page.

Upgrade cost for a user is close to zero. The app is a self-contained desktop binary, there is no server component to migrate, and experiments are exported as JSON files that you control. The one thing to watch across versions is the format of those exported experiment files, since re-running past experiments is a documented feature and depends on stored parameters remaining readable.

Building from source is a heavier commitment. The stack spans Rust and the Tauri toolchain on one side and a Bun and Vite front end on the other, so a contributor needs both installed. The docs/DEVELOPMENT.md notes are the entry point rather than the README.

The licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. That is a permissive arrangement with few obligations, but it is not legal advice and the LICENSE file at the repository root is the text that governs. One practical consequence of MIT combined with ad-hoc macOS signing is that nothing stops you from building and signing your own binary if the Gatekeeper prompt is unacceptable in a managed environment.

## Conclusion

Adopt it if you already run Ollama and your question is which of your installed models handles a prompt best under a given temperature or system prompt. Skip it if you need scored benchmarks, hosted API models, or anything that runs without a local Ollama endpoint. Before trusting a run, confirm the Ollama Server URL in Settings points at the right host, and re-read the exported JSON to check that the parameter values you typed are the ones that were actually sent.

## FAQ

### How do I do a grid search with Ollama Grid Search?

Select the models you want, write a prompt, and add parameter values to sweep; the README states the prompt is submitted once for each parameter value, for each selected model. The results appear as a set of responses you can inspect side by side and later export as JSON.

### How do I find the Ollama base URL for Ollama Grid Search?

The README does not document a discovery procedure. It states the app assumes Ollama is serving endpoints on localhost or a remote server, and that the server URL is set in Settings. The one non-default value the README gives is http://localhost:17434 for llmman.

### What does the grid search method mean in Ollama Grid Search?

The README notes that grid search normally refers to sweeping training hyperparameters such as batch_size, learning_rate or number_of_epochs, and that the concept here is only similar. In this app the sweep covers models, prompts and inference parameters, and the prompt is submitted once per parameter value per model.

### What is the difference between grid search and random search in Ollama Grid Search?

The README does not discuss random search, so the app's documentation offers no comparison. It does describe its own sweep as iterating over combinations of models, prompts and inference parameters, with results inspected and exported rather than sampled at random.

## Sources

- [dezoito/ollama-grid-search on GitHub](https://github.com/dezoito/ollama-grid-search)
- [Issues](https://github.com/dezoito/ollama-grid-search/issues)
- [License: MIT](https://github.com/dezoito/ollama-grid-search/blob/main/LICENSE)
- [README](https://github.com/dezoito/ollama-grid-search/blob/main/README.md)
- [Releases](https://github.com/dezoito/ollama-grid-search/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/dezoito-ollama-grid-search
