# llama.vim: local FIM completion inside Vim, wired to a llama.cpp server

> llama.vim is a Vim plugin that requests fill-in-the-middle completions and instruction-based edits from a llama.cpp server you run yourself. It is small, MIT-licensed, and useless without that server.

**ggml-org/llama.vim** — Vim plugin for LLM-assisted code/text completion

- Repository: https://github.com/ggml-org/llama.vim
- Stars: 2,178 · Forks: 125
- Language: Vim Script
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ggml-org-llama-vim

## What llama.vim actually replaces

The plugin targets one workflow: you are typing in Insert mode, you pause or move the cursor, and a grey suggestion appears ahead of the cursor. Press Tab and it is inserted. Press Shift+Tab and only the first line is inserted. That is the whole loop, and it is deliberately narrower than what a hosted assistant does. There is no chat pane, no file-level refactor command, no agent that edits multiple buffers.

The second mode is instruction-based editing, bound by default to <leader>lli. You give an instruction and the plugin asks the server to rewrite the region, with <leader>llr to rerun, <leader>llc to continue, Tab to accept and Esc to cancel. So the plugin covers two things people actually did with Copilot-style tools in Vim: accept a completion, and ask for a local edit.

Who it is for: developers who already have a machine with a GPU, who are comfortable running a server process alongside their editor, and who care that the code never leaves their network. The README's own framing is "Local LLM-assisted text completion," and the llama.cpp dependency is not optional.

## The client-server split and the context ring buffer

llama.vim does not contain a model. It is an HTTP client. The README states the plugin requires a llama.cpp server instance running at g:llama_config.endpoint_fim and/or g:llama_config.endpoint_inst. Completion requests go to the FIM endpoint; instruction-based edits go to the instruction endpoint. You can point them at the same server or two different ones.

The interesting part is how context is assembled. The plugin keeps a ring buffer of chunks pulled from open files, edited files and yanked text, and sends a selected slice of that ring alongside the text around the cursor. The README links this to a llama.cpp change and describes it as supporting "very large contexts even on low-end hardware via smart context reuse." The mechanism matters because a naive client would resend the entire context every keystroke; here the server can reuse the prompt prefix it already has cached, so only newly computed tokens cost time.

The README's own example makes the accounting visible. On an M1 Pro with Qwen2.5-Coder 1.5B Q8_0, the status line showed 15186 tokens of context against a 32768 maximum, 30 chunks in the ring out of 64, 1 chunk evicted, 260 newly computed prompt tokens and 24 generated tokens, taking 1245 ms after typing a single letter. That is the trade-off in one line: most of the context was reused, a small part was new, and the latency is still measured in over a second.

## Installing llama.vim and getting a first completion

The plugin ships no server, so the install is two halves. First the plugin manager entry. With vim-plug:

```vim
Plug 'ggml-org/llama.vim'
```

With lazy.nvim, the README gives a bare spec:

```lua
{
    'ggml-org/llama.vim',
}
```

Second, llama.cpp. On macOS the README says `brew install llama.cpp`; on Windows, `winget install llama.cpp`. On other systems it points at the llama.cpp releases page or a source build. Then start the server with a FIM-capable model sized to your VRAM. The README's recommended settings are:

```bash
llama-server --fim-qwen-7b-default
```

That flag is the one for more than 16GB VRAM. The README lists --fim-qwen-30b-default above 64GB, --fim-qwen-3b-default below 16GB, and --fim-qwen-1.5b-default below 8GB. After that, open a file in Vim, enter Insert mode and start typing. The README's screenshots show the suggestion in orange and a green stats line with token counts and milliseconds. If nothing appears, the endpoint default in autoload/llama.vim is the first thing to check against the port your server actually bound to.

## Profiles: pointing one Vim at several machines

Because the plugin is a client, the server can live elsewhere. g:llama_config.profiles maps names to base URLs, and g:llama_config.profile selects one:

```vim
let g:llama_config.profiles = {
    \ 'spark':   'http://192.168.0.66:8080',
    \ 'gmktec':  'http://192.168.0.65:8080',
    \ }
let g:llama_config.profile = 'spark'
```

:LlamaProfile lists the profiles, :LlamaProfile spark switches both FIM and instruction requests, and profile names support command-line completion. The selection is persisted, and :LlamaProfileReset restores whatever .vimrc configured. This is the feature that makes a desktop-plus-home-server setup practical: the laptop runs Vim, the GPU box runs llama-server, and switching between two GPU boxes is one command.

The README is explicit about the limit here. Profiles only specify hosts, not models; models come from model_fim and model_inst. The suggested workaround is consistent naming across servers or server-side aliases, so that model_fim resolves to the same alias name on every host. That is a reasonable design, but it does mean a profile switch can silently pair your FIM requests with a model you did not intend if the aliases diverge.

## Where llama.vim is the wrong tool

The hard dependency is the first limitation. If you cannot run llama.cpp somewhere reachable, the plugin does nothing. There is no bundled model, no fallback to a hosted API, and no offline mode beyond your own server.

Model choice is constrained too. The README states the plugin requires FIM-compatible models and links a Hugging Face collection for them. A general chat model that was not trained for fill-in-the-middle will not produce useful inline completions here, no matter how good it is at conversation.

Latency is the second constraint, and the README's own numbers set expectations rather than hide them. The M1 Pro example shows 1245 ms for a 24-token suggestion on a 1.5B model. On a small model and consumer hardware, that is a suggestion you wait for, not one that keeps pace with typing. The README also shows an M2 Ultra with a 7B model working across a large codebase, but the smaller-hardware case is the one most readers will hit.

Finally, this is a Vim plugin written in Vim Script. If your team standardises on VS Code or a JetBrains IDE, none of this transfers. The plugin's value is tied to the editor, not to the server.

## How it differs from a hosted assistant and from raw llama.cpp

The obvious comparison is GitHub Copilot. The difference is not the feature list, which is similar at the completion level, but where the request goes. Copilot sends your buffer to a remote service and returns a suggestion; llama.vim sends the same kind of request to an endpoint you control, and the model is whatever you loaded. That changes cost, latency and data handling at once. It also changes quality: a 1.5B local model and a large hosted model are not equivalent, and the README's own example numbers show the local model being asked to do the job in about a second.

The second comparison is llama.cpp's own server. llama-server already exposes a completion endpoint, and you could curl it from a Vim function. What llama.vim adds is the editor-side machinery: trigger on cursor movement in Insert mode, the keymaps for accepting a word, a line or the full suggestion, the instruction-edit loop with rerun and continue, and the ring buffer that decides which chunks of your other files get sent. That context assembly is the part that is tedious to rebuild by hand, and it is the reason to use the plugin rather than a shell script.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-08, nine days before this writing. The only release listed is v0.1.0, dated 2026-08-24. So the project is moving, but it is early: one tagged release, and a README that points readers to :help llama_config and autoload/llama.vim for the full option list rather than documenting every key in the README itself. Expect configuration to be discovered by reading the source.

The licence is MIT, which permits commercial use and modification with the copyright notice retained. That matters less for the plugin than for the model you load behind it: model weights carry their own licences, and the README links a Hugging Face collection without discussing terms. Check the model licence separately; the plugin's MIT badge says nothing about it.

Upgrade cost is mostly on the llama.cpp side. The plugin talks to the server over HTTP, and the README ties its context reuse to a specific llama.cpp change. If you pin an old llama.cpp build for stability, you are also pinning whatever behaviour that change depends on. There is no documented rollback procedure in the README for a plugin version that misbehaves against a given server build, so keeping the previous plugin revision in your plugin manager's lockfile is the practical safeguard.

## Conclusion

Adopt llama.vim if you already run llama.cpp locally or on a LAN box and want FIM completions without sending code to a hosted service. Skip it if you want a plugin that works out of the box with no server, or if you need a non-Vim editor. Before committing, confirm which FIM-capable model your GPU can hold, start llama-server with the matching --fim-qwen-*-default flag, and check that g:llama_config.endpoint_fim points at that port.

## FAQ

### What is Vim and why is it used?

The README does not cover this. It assumes you are already in Vim, and everything it documents (Insert mode, Tab, Shift+Tab, <leader> keymaps) is written for someone who edits there daily.

### How is LLaMA different from GPT?

The README does not compare model families. It only states that the plugin requires FIM-compatible models and links a Hugging Face collection of them, and that the server behind it is llama.cpp.

### What does Vim stand for?

The README says nothing about this. It describes llama.vim as a Vim plugin for local LLM-assisted text completion and leaves the editor's history out of scope.

### What does LLaMA stand for?

The README does not expand the acronym. It uses LLaMA only in the project name and points to llama.cpp as the required server, with no note on what the letters mean.

## Sources

- [ggml-org/llama.vim on GitHub](https://github.com/ggml-org/llama.vim)
- [Issues](https://github.com/ggml-org/llama.vim/issues)
- [License: MIT](https://github.com/ggml-org/llama.vim/blob/master/LICENSE)
- [README](https://github.com/ggml-org/llama.vim/blob/master/README.md)
- [Releases](https://github.com/ggml-org/llama.vim/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ggml-org-llama-vim
