# llm-vscode: Hugging Face Code Completion Extension for Visual Studio Code

> llm-vscode is a VS Code extension from Hugging Face that adds ghost-text code completion backed by the Hugging Face Inference API, Ollama, Text Generation Inference, or any OpenAI-compatible endpoint. It is the VS Code sibling of the same team's Neovim and IntelliJ extensions, using llm-ls as the shared backend.

**huggingface/llm-vscode** — LLM powered development for VSCode

- Repository: https://github.com/huggingface/llm-vscode
- Stars: 1,315 · Forks: 142
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/huggingface-llm-vscode

## What llm-vscode Does and Who It Is For

llm-vscode adds one capability to Visual Studio Code: ghost-text code completion, identical in feel to GitHub Copilot's inline suggestions. The plugin displays a suggested completion at the cursor position as you type. The key difference from Copilot is the model source: llm-vscode sends requests to the endpoint you configure, which can be the Hugging Face Inference API, a local Ollama server, a Text Generation Inference server, or any server that exposes an OpenAI-compatible completions API.

The plugin was formerly named huggingface-vscode. The Hugging Face team also maintains a Neovim version (llm.nvim), a Jupyter extension (jupytercoder), and an IntelliJ extension (llm-intellij) that share the same llm-ls backend. If you work across editors, all four use compatible configuration patterns.

The default model is bigcode/starcoder, routed through the Hugging Face Inference API. Using the free tier of the Inference API will eventually hit rate limits; the README recommends the PRO plan for uninterrupted use. Engineers who want no API dependency can configure Ollama as the backend instead.

## Installing the Extension

Install llm-vscode from the VS Code Marketplace like any other extension. Search for llm-vscode under the HuggingFace publisher, or install it directly from the marketplace page. The display name is llm-vscode and the publisher identifier is HuggingFace.

After installing, supply your Hugging Face API token if you are using the Inference API backend:

1. Open the VS Code command palette with Cmd/Ctrl+Shift+P
2. Type Llm: Login

If you have previously run huggingface-cli login on your system, the extension reads the token from disk automatically. The token stored by the CLI is compatible with the extension's authentication. For the Ollama backend, no token is required; you configure the URL instead.

The bundled llm-ls binary is included with the extension and does not need a separate download. If you prefer to use the Mason-managed binary or a custom build, set llm.lsp.binaryPath in the extension settings to point to that binary. The extension requires VS Code 1.82.0 or later.

## Configuring the Backend

The extension supports four backends: huggingface, ollama, openai, and tgi. The default is huggingface with the starcoder model. Switching to a local Ollama server requires setting the backend to ollama and providing the local server URL. Switching to TGI requires the TGI server URL.

The request body format varies by backend. For the huggingface backend, the extension builds a fill-in-the-middle prompt using the model's FIM tokens, which are model-specific. The README shows what the request looks like internally:

```js
const inputs = `{start token}import numpy as np\nimport scipy as sp\n{end token}def hello_world():\n    print("Hello world"){middle token}`
const data = { inputs, ...configuration.requestBody };
```

URL construction follows a consistent rule across all backends: llm-ls appends the correct path for each backend type (for example, /v1/completions for the openai backend) unless the base URL already ends with that path. Setting llm.disableUrlPathCompletion to true disables this automatic path appending when you want full control over the URL.

The llm.requestBody setting in the extension configuration passes additional parameters to the model, such as temperature and top_p. These map directly to the request body the backend expects.

## Tokenizer and Context Window Management

One practical detail that distinguishes llm-vscode from simpler completion plugins is automatic prompt sizing. The llm-ls backend uses Hugging Face's tokenizers library to count the tokens in the prompt and truncate it to fit within the model's context window. This prevents silent failures from sending prompts that are too long.

The tokenizer configuration has four options. Setting llm.tokenizer to null disables tokenization and uses character counting instead:

```json
{
  "llm.tokenizer": null
}
```

Pointing to a Hugging Face repository downloads the tokenizer.json automatically:

```json
{
  "llm.tokenizer": {
    "repository": "myusername/myrepo",
    "api_token": null
  }
}
```

When api_token is null, the extension uses the token set via Llm: Login. The tokenizer can also be loaded from a local file path or from an HTTP URL. The starcoder default configuration uses the repository option pointing to bigcode/starcoder, which is the recommended approach for any model hosted on the Hugging Face Hub.

Using the correct tokenizer matters when the code file near the cursor is large. Without accurate tokenization, the prompt can silently exceed the context window and the model receives a truncated or malformed prompt.

## Code Attribution: Checking Generated Code Against The Stack

llm-vscode includes a code attribution feature that checks whether generated code appears in The Stack, the training dataset for StarCoder. Press Cmd+Shift+A after accepting a suggestion to run the check. The result is a first-pass attribution check, not a definitive legal analysis.

The check queries stack.dataportraits.org using a Bloom filter approach. Sequences of at least 50 characters are checked for matches. The README notes that false positives are possible because the method checks for n-gram matches rather than exact file matches, and sufficient surrounding context is required for accurate results. The dedicated Stack search tool at hf.co/spaces/bigcode/search provides a more complete second-pass check for cases where the attribution result is important.

This feature is relevant mainly for teams using StarCoder or StarCoder-based models. Other models from different training datasets do not have attribution coverage through the same endpoint.

## Practical Limitations

llm-vscode is a completion-only plugin. There is no chat sidebar, no code explanation command, no refactoring assistant, and no context from open files beyond the current cursor position and its surrounding code. VS Code's built-in Copilot integration, available since VS Code 1.82.0, provides inline completion plus a chat interface in one extension. For engineers who need both, Copilot is the more complete option.

The Inference API rate limits on the free tier are a real constraint for active development sessions. The README's recommendation to subscribe to the PRO plan is the documented mitigation. Using Ollama avoids rate limits entirely but requires hardware capable of running the model locally and introduces local inference latency.

The suggestion trigger by default fires automatically as you type, controlled by llm.enableAutoSuggest. On slow connections or with large models, the auto-suggest behavior can introduce noticeable latency. Setting llm.enableAutoSuggest to false and using the Cmd+Shift+L manual trigger gives more control over when the model is called.

The extension supports VS Code 1.82.0 or later and targets both the Intel and Apple Silicon builds of VS Code on macOS. The last push to the repository was on 2026-09-23, five days before this article was written, and the current version is 0.2.2. The license is Apache-2.0.

## Conclusion

llm-vscode is the practical choice for VS Code users who want LLM code completion without being locked into GitHub Copilot's model and pricing. Pointing it at a local Ollama server removes the API cost entirely. The code attribution feature is a useful first-pass tool for checking whether generated code appears in The Stack. The plugin scope is limited to completion: there is no chat interface, no inline editing commands, and no context-aware refactoring beyond what the model generates. GitHub Copilot for VS Code remains a better option when those additional capabilities matter. The last push was on 2026-09-23, and the extension is available on the VS Code Marketplace under the HuggingFace publisher.

## FAQ

### How do I use a local LLM with VS Code instead of the Hugging Face API?

Set the backend to ollama in the extension configuration and provide your local Ollama server URL, typically http://localhost:11434. Ollama must be running with a model downloaded before the extension will return completions. No API token is required for the Ollama backend.

### What is the difference between llm-vscode and GitHub Copilot?

llm-vscode supports configurable backends including local Ollama servers and any Hugging Face model, while Copilot uses a fixed Microsoft-hosted model. Copilot includes a chat panel and additional AI editing commands; llm-vscode provides only inline completion. The code attribution feature in llm-vscode is specific to models trained on The Stack dataset.

### Does VS Code have a built-in AI?

VS Code 1.82.0 and later include a built-in GitHub Copilot integration. llm-vscode is a separate extension that provides similar inline completion through configurable model endpoints rather than through Copilot's fixed model.

## Sources

- [huggingface/llm-vscode on GitHub](https://github.com/huggingface/llm-vscode)
- [Issues](https://github.com/huggingface/llm-vscode/issues)
- [License: Apache-2.0](https://github.com/huggingface/llm-vscode/blob/master/LICENSE)
- [README](https://github.com/huggingface/llm-vscode/blob/master/README.md)
- [Releases](https://github.com/huggingface/llm-vscode/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/huggingface-llm-vscode
