# Gollama: A Terminal Interface for Managing Local Ollama Models

> Gollama is a Go-based TUI for macOS and Linux that replaces individual Ollama CLI calls with a single interactive screen for sorting, deleting, editing Modelfiles, and estimating vRAM. The author has stated publicly that maintenance is slowing.

**sammcj/gollama** — Go manage your Ollama models

- Repository: https://github.com/sammcj/gollama
- Website: https://smcleod.net
- Stars: 1,838 · Forks: 113
- Language: Go
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/sammcj-gollama

## What Gollama Replaces in an Ollama Workflow

Ollama ships its own command-line verbs for every management operation: list models, remove a model, copy or rename one, push to a registry, or show model details. Each call handles one model at a time. If you maintain a growing library of local models and want to compare quantisation levels across several of them, identify the largest files for removal, or delete a set of outdated variants in one pass, the single-verb approach requires many individual commands and manual bookkeeping between them.

Gollama replaces that pattern with a TUI that presents all installed models on one screen. The tool targets macOS and Linux developers who use Ollama for local inference and need to keep their model library manageable. It does not run inference itself; it delegates all model serving to Ollama. Its scope is the management layer: seeing the full model set at once, sorting and selecting within it, and getting vRAM estimates without leaving the terminal.

## Architecture: Bubbletea over the Ollama HTTP API

The application is built on the Charmbracelet stack. The go.mod lists charmbracelet/bubbles, bubbletea, and lipgloss as core dependencies. Bubbletea follows an Elm-inspired model-update-view pattern for terminal UI construction. Model data is fetched from Ollama's HTTP API, defaulting to http://localhost:11434. The -h or --host flag at startup points the tool at a different host or port.

Configuration is managed by Viper with fsnotify watching the config file for live changes. Logging uses zerolog, with the level overridable via --log-level debug or --log-level error at the command line. The top-level repository layout separates responsibilities clearly: app_model.go holds the core application state, item_delegate.go controls list row rendering, and operations.go contains the model manipulation functions that call Ollama's API. The vramestimator/ subdirectory is an independent package for the vRAM calculation logic.

## Installing Gollama and Running It for the First Time

The README recommends the Go install path, which fetches the latest tagged release and places the binary in $GOPATH/bin:

```bash
go install github.com/sammcj/gollama/v2@latest
```

If the shell reports that gollama is not found after installation, the Go bin directory is not in your PATH. The README gives the fix for zsh:

```shell
echo 'export PATH=$PATH:$HOME/go/bin' >> ~/.zshrc
source ~/.zshrc
```

A curl-based installer is also documented for systems where Go is not installed:

```shell
curl -sL https://raw.githubusercontent.com/sammcj/gollama/refs/heads/main/scripts/install.sh | bash
```

The README notes this path is harder to update when new versions ship. Once installed, launch the TUI with no arguments:

```sh
gollama
```

The full model list appears. Press Space to select a model, D to delete the selected set, i to inspect details, and c to copy or rename. For a non-interactive listing that exits immediately, the -l flag works without opening the TUI:

```shell
gollama -l
```

## The vRAM Estimator and Command-Line Search

Beyond the TUI, Gollama ships a vRAM estimation tool usable directly from the command line. The --vram flag accepts both Ollama model identifiers in name:tag format and HuggingFace repository IDs in author/name format. The README gives examples such as llama3.1:8b-instruct-q6_K and qwen2:14b-q4_0 as accepted Ollama-style identifiers.

The --fits parameter takes available GPU memory in gigabytes and calculates how much context fits within that budget. The --vram-to-nth or --context flag sets the maximum context length to analyse. The --quant flag overrides the quantisation level for what-if calculations, which is useful when comparing two quantisation variants of the same model before downloading either.

For search operations without opening the TUI, the -s flag filters models by name. The OR operator uses a pipe character inside single quotes, and the AND operator uses an ampersand:

```shell
gollama -s my-model
gollama -s 'my-model|my-other-model'
gollama -s 'my-model&instruct'
```

To edit the Modelfile for a specific model from the command line:

```shell
gollama -e my-model
```

## Platform Limits, LM Studio Removal, and What the TUI Cannot Do

Gollama runs on macOS and Linux only. The README describes it explicitly as a macOS and Linux tool, and there is no documented Windows build or installation path. Users on Windows cannot use it.

The application requires Ollama to be running. Without a reachable Ollama API, the TUI opens but cannot fetch models. The most common user-reported symptom is an error fetching models message, which in most cases means Ollama is not running or the host flag points to the wrong address. For remote or non-default Ollama instances, the -H flag is a shortcut for connecting to http://localhost:11434, while --host accepts any custom URL. The --ollama-dir flag lets you point at a custom Ollama models directory if your setup departs from the default location.

LM Studio integration was removed in v2.0.1. The README records that maintaining the link between Ollama and LM Studio across upstream changes in both applications consumed too much of the author's time for a feature they rarely used themselves. The feature will not return. Users who depended on syncing models between Ollama and LM Studio must manage that transfer through LM Studio's own interface.

The tool provides no web-based or graphical interface. A terminal renderer is the only option.

## Gollama vs. Ollama's Built-in Commands

Ollama's own CLI covers the same functional ground: ollama list, ollama rm, ollama cp, ollama push, and ollama show. For a single, well-known operation on a specific model, typing the Ollama command directly is faster than opening a TUI. Gollama's advantage is in multi-model workflows: selecting a group of models for deletion, sorting by size to identify large candidates for removal, and comparing quantisation levels across a set, all from a single screen without rerunning commands.

The vRAM estimator is not present in Ollama's built-in CLI. That functionality belongs to Gollama alone.

For users who want a browser-based interface for local Ollama management, Open WebUI provides a web panel that covers model management alongside a chat interface. Open WebUI is aimed at interactive chat as the primary use case. Gollama is aimed at developers who prefer terminal workflows and need to work with many models at once. The two do not compete on the same use case.

## Maintenance Status and License

The repository is not archived. The last push to the main branch was on 2026-07-20, and release v2.0.6 was tagged on 2026-09-28. The author published a note in the README dated 2025-12-02 stating that maintenance is slowing. The reason given is that they have moved to llama.cpp with llama-swap and to LM Studio for local model serving, and no longer use Ollama regularly.

This matters for anyone planning to rely on Gollama long term. Bugs may go unpatched and new Ollama API changes may not be reflected promptly. The MIT license places no restrictions on forking or modification, so adopters who need a fix have the option to apply it themselves.

Upgrading is low-effort for users who installed via go install: rerunning the same command fetches the latest release without any additional steps.

## Conclusion

Developers on macOS or Linux who run multiple Ollama models and want to batch-delete, compare sizes, or estimate vRAM without issuing individual CLI commands will find Gollama useful. Windows users, users who relied on the LM Studio integration (removed in v2.0.1), and users who want a browser-based panel have no supported path with this tool. Before adopting it, confirm your Go bin directory is in PATH and that an Ollama instance is reachable. The author's own note in the README, dated 2025-12-02, states that maintenance is slowing because they have moved primarily to llama.cpp with llama-swap.

## FAQ

### How do you use Gollama?

Run gollama with no arguments to open the TUI. Press Space to select models, D to delete the selected set, and i to inspect a model's details. For non-interactive use, gollama -l lists all models and exits, and gollama -s <term> filters models by name.

### Does Gollama run on Windows?

No. The README describes Gollama as a macOS and Linux tool. There is no documented Windows build or installation path.

### How does Gollama estimate vRAM usage for a model?

The --vram flag accepts an Ollama model identifier or a HuggingFace repository ID. The estimator calculates approximate VRAM requirements for a given context length. The --fits parameter specifies available GPU memory in GB, and --quant overrides the quantisation level for what-if calculations.

## Sources

- [License: MIT](https://github.com/sammcj/gollama/blob/main/LICENSE)
- [Project website](https://smcleod.net)
- [README](https://github.com/sammcj/gollama/blob/main/README.md)
- [Releases](https://github.com/sammcj/gollama/releases)
- [sammcj/gollama on GitHub](https://github.com/sammcj/gollama)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sammcj-gollama
