Model or dataset
huggingface/llm.nvim avatar
huggingface/llm.nvim

llm.nvim: Hugging Face Code Completion and LLM Integration for Neovim

LLM powered development for Neovim

1,189 stars58 forksLuaApache-2.0

At a glance

What is it?
llm.nvim is a Neovim plugin from Hugging Face that adds ghost-text code completion to the editor through an llm-ls language server backend. It supports the Hugging Face Inference API, Ollama, OpenAI-compatible endpoints, and Text Generation Inference, with the model choice open to any HTTP endpoint that matches the expected request format.
Who is it for?
llm.nvim is the right choice for Neovim users who want code completion that is not tied to GitHub Copilot's model selection and subscription, or who want to run the completion model locally via Ollama. The ghost-text experience is close to Copilot's inline suggestion model.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 13 days ago.
What is it written in?
Mainly Lua, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What llm.nvim Does in the Editor

llm.nvim adds one primary capability to Neovim: ghost-text code completion, which displays suggested code inline at the cursor position as you type, similar to how GitHub Copilot's inline completion works. The suggestion appears as faded text after the cursor. Pressing the configured key accepts it; continuing to type dismisses it.

The plugin is influenced by copilot.vim and tabnine-nvim, as the README states directly. The distinction from those tools is the model source: llm.nvim sends completion requests to any HTTP endpoint you configure, rather than routing all requests through a single vendor's service. This openness is the practical reason to choose llm.nvim: you can point it at a Hugging Face model through the Inference API, at a locally running Ollama instance, at a Text Generation Inference server you host yourself, or at any server that exposes an OpenAI-compatible completions endpoint.

The plugin was formerly named hfcc.nvim and has been renamed to llm.nvim. The backend language server is llm-ls, which handles the actual HTTP communication and prompt construction. The plugin installs llm-ls automatically on first load by downloading the binary from the llm-ls release page.

Backend Support: From Hugging Face to Ollama

The plugin supports four backends, each with a different configuration path. The huggingface backend calls the Hugging Face Inference API and requires an API token. The token can be passed in plugin opts, set as LLM_NVIM_HF_API_TOKEN environment variable, stored at $HF_HOME/token, or read from the token set by huggingface-cli login. The model is set with the model option or the LLM_NVIM_MODEL environment variable.

The ollama backend connects to a locally running Ollama server. A minimal configuration for CodeLlama looks like this:

lua
{
  model = "codellama:7b",
  url = "http://localhost:11434", -- llm-ls uses "/api/generate"
  request_body = {
    options = {
      temperature = 0.2,
      top_p = 0.95,
    }
  }
}

The openai backend handles any server that exposes an OpenAI-compatible completions endpoint, including llama-cpp-python and similar local inference servers. The tgi backend connects to Text Generation Inference servers. The URL construction in all four backends follows the same pattern: llm-ls appends the correct path based on the backend setting, so you provide the base URL and let the server handle routing.

The LLM_NVIM_URL environment variable overrides the backend URL at runtime, which is useful for switching between a local development server and a remote one without changing the Neovim configuration file.

The tgi backend connects to a Text Generation Inference server. TGI is a production inference server from Hugging Face that supports many open models with optimized throughput. It requires more setup than Ollama but handles concurrent requests better and supports larger models with tensor parallelism. An example configuration for TGI points to a local server running StarCoder at http://localhost:8080.

The backend setting also controls how the prompt is formatted. The huggingface backend builds a fill-in-the-middle prompt with the model's specific FIM tokens inserted between the context before the cursor and the context after it. Getting this format right matters because a model that does not recognize its own FIM tokens will treat the surrounding code as regular text rather than as cursor context, producing suggestions that repeat existing code rather than completing the gap.

Installing llm.nvim and llm-ls

llm.nvim is a standard Neovim plugin and installs with any Neovim plugin manager. The VS Code Marketplace page is at the URL listed in the README for reference. The README describes it as installing like any other vscode extension, but for Neovim the installation follows the standard plugin manager approach.

By default, llm-ls is downloaded automatically from its GitHub release page on first load and stored in:

lua
vim.api.nvim_call_function("stdpath", { "data" }) .. "/llm_nvim/bin"

The lsp.version setting controls which llm-ls version is downloaded. If you use mason.nvim, you can install llm-ls through Mason instead:

vim
:MasonInstall llm-ls

After a Mason install, set the lsp.bin_path to the Mason binary path:

lua
{
  lsp = {
    bin_path = vim.api.nvim_call_function("stdpath", { "data" }) .. "/mason/bin/llm-ls",
  },
}

The lsp.cmd_env option sets environment variables for the llm-ls process specifically, which is useful when the API token should not be in the global shell environment. The llm-ls binary can also start in TCP mode with the --port flag for debugging.

Configuring Models and the Tokenizer

The default model is bigcode/starcoder with a context window of 8192 tokens. The plugin ships with a preconfigured block for StarCoder that includes the fill-in-the-middle tokens:

lua
{
  tokens_to_clear = { "<|endoftext|>" },
  fim = {
    enabled = true,
    prefix = "<fim_prefix>",
    middle = "<fim_middle>",
    suffix = "<fim_suffix>",
  },
  model = "bigcode/starcoder",
  context_window = 8192,
  tokenizer = {
    repository = "bigcode/starcoder",
  }
}

The fim (fill-in-the-middle) configuration tells the plugin how to format the prompt so the model can complete from the cursor position using both the text before and after the cursor. Each model has its own FIM token format; CodeLlama uses different prefix, middle, and suffix tokens with specific spacing requirements, as noted in the README.

The tokenizer configuration determines how the prompt is sized to fit the context window. The three options are: no tokenizer (llm-ls counts characters instead), a local tokenizer.json file, or a repository on the Hugging Face Hub from which llm-ls downloads the tokenizer.json automatically. The repository path is the Hugging Face model identifier.

Keybindings and Suggestion Behavior

The plugin sets two keybindings by default. Cmd+Shift+L triggers inline suggestion manually, which corresponds to the editor.action.inlineSuggest.trigger command. The llm.enableAutoSuggest setting controls whether suggestions appear automatically as you type.

The llm.documentFilter setting restricts completions to specific files or directories. To enable completions only in a particular project directory, set the pattern to match that path. To enable completions only in Python and Rust files, set the pattern to **/*.{py,rs}. This prevents the plugin from making requests for file types where code completion is not useful.

The request delay is configurable through llm.requestDelay with a default of 150 milliseconds. This debounces the trigger so the plugin does not fire on every keystroke. Increasing the delay reduces API calls at the cost of slightly less responsive suggestions.

Limitations and Cases Where llm.nvim Is the Wrong Tool

llm.nvim provides code completion only. It has no chat panel, no code explanation feature, no inline editing, and no context-aware refactoring. Engineers who need a broader AI assistant experience in Neovim will find the plugin scope too narrow. There is no multi-file context: the prompt is built from the current file around the cursor position, and other open buffers are not included.

The Hugging Face Inference API has rate limits on the free tier. The README notes that subscribing to the PRO plan avoids rate limiting, but this adds a recurring cost. Using Ollama locally avoids API costs entirely at the expense of GPU memory and the latency of local inference. On machines without a GPU, local inference is slow enough to make auto-suggest impractical; manual triggering with Cmd+Shift+L is a better pattern in that case.

Model quality for fill-in-the-middle completion varies significantly by model. The default StarCoder model is trained specifically for code completion with FIM tokens and produces better inline suggestions than a general-purpose chat model would. Switching to a chat model that does not support FIM tokens will produce worse suggestions even if the model is otherwise capable.

GitHub Copilot for Neovim is the direct alternative. It provides inline code completion with a similar ghost-text interface, plus a chat panel. The tradeoff is that Copilot uses a fixed model with no option to point the plugin at a different endpoint. llm.nvim is better when you want model flexibility or local inference; Copilot is better when you want a more complete AI coding experience in one plugin.

The last push to the repository was on 2026-09-17, eleven days before this article was written. The license is Apache-2.0.

Editorial conclusion

llm.nvim is the right choice for Neovim users who want code completion that is not tied to GitHub Copilot's model selection and subscription, or who want to run the completion model locally via Ollama. The ghost-text experience is close to Copilot's inline suggestion model. The primary limitation is that the plugin supports only code completion: there is no chat interface, no code explanation, and no in-editor assistant beyond inline suggestions. Engineers who need a broader AI coding experience in Neovim should look at GitHub Copilot for Neovim, which provides a chat panel alongside completion. Verify the llm-ls binary version matches your platform before relying on the plugin in a production workflow; the README notes that platform support depends on the llm-ls release page.

Frequently asked questions

What is llm.nvim used for?

llm.nvim adds ghost-text code completion to the Neovim editor. It sends your current file context to a language model backend, such as the Hugging Face Inference API or a local Ollama server, and displays the model's suggestion inline at the cursor. Accepting the suggestion inserts it into the file.

Can llm.nvim use a local model instead of the Hugging Face API?

Yes. The ollama backend routes requests to a locally running Ollama server, which runs the model on your own hardware with no API key required. You can also use any server that exposes an OpenAI-compatible completions endpoint through the openai backend setting.

How is llm.nvim different from GitHub Copilot for Neovim?

llm.nvim supports configurable backends including local Ollama servers, Hugging Face models, and any OpenAI-compatible endpoint, while Copilot uses a fixed model. Copilot provides both inline completion and a chat panel; llm.nvim provides only inline completion. Choose llm.nvim for model flexibility and local inference; choose Copilot for a broader in-editor AI experience.

Official sources

  1. huggingface/llm.nvim on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/huggingface-llm-nvim.svg)](https://hysenlabs.com/projects/huggingface-llm-nvim)