llama.vscode: local FIM completions and an agent inside VS Code
VS Code extension for LLM-assisted code/text completion
At a glance
- What is it?
- The ggml-org extension wires VS Code to a local llama.cpp server for fill-in-the-middle completion, chat and an agent view. It is a thin client, and its quality is the quality of the model you point it at.
- Who is it for?
- Adopt llama.vscode if you already run llama.cpp or are willing to, you want completion that never leaves the machine, and you accept that model choice decides output quality. Skip it if you want a hosted assistant with no local runtime, or if you cannot give a FIM-capable model enough VRAM to be useful.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly TypeScript, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap llama.vscode fills for VS Code users with local models
Hosted completion is convenient, but it sends your buffer to someone else's server and it stops working when the network does. llama.vscode takes the opposite position. It is a VS Code extension that talks to a llama.cpp server running on your own machine, and the README describes it as "Local LLM-assisted text completion, chat with AI and agentic coding extension for VS Code". The target reader is someone who already has a GPU or a reasonably fast CPU and would rather pay in VRAM than in API calls.
The extension is written in TypeScript and published by ggml-org, the same organisation behind llama.cpp and the llama.vim plugin. The README states the initial implementation was done by Ivaylo Gardev using llama.vim as a reference, which explains the shape of the feature list: inline suggestions, a Tab to accept, a word-at-a-time accept, and a manual trigger. This is a completion tool first. Chat and the agent view were added on top of that base rather than the other way round.
One consequence is worth stating plainly. The extension does not ship a model. Everything you get out of it depends on which GGUF you load and how much memory you can give it.
How the completion loop actually works
The extension is a client. It sends a fill-in-the-middle request to a llama.cpp server and renders the returned text as a ghost suggestion in the editor. FIM means the model is asked to produce the missing middle between a prefix and a suffix, which is what makes inline completion feel like continuation rather than chat.
Context assembly is where the design gets interesting. The README lists configuration for the scope of context around the cursor, plus something it calls "Ring context with chunks from open and edited files and yanked text". In other words, the prompt is not just the current file. Open buffers, recently edited regions and clipboard content can be pulled in. That is a real benefit for cross-file work and a real risk for privacy hygiene inside a shared repo, since unrelated buffers can end up in the prompt.
The README also points at smart context reuse, linking a llama.cpp pull request, and claims support for "very large contexts even on low-end hardware". Read that as a statement about the server-side KV cache, not about the extension being clever on its own. The extension's job is to keep the prefix stable so the server can reuse what it already computed. Change the prefix constantly and the reuse advantage shrinks.
Generation is bounded by a max text generation time setting, which the README lists as a feature. That is the right knob for an inline tool: a suggestion that arrives after you have already typed the line is worthless, so a time budget matters more than a token budget.
Installing llama.vscode and getting a first suggestion
Install the extension from the VS Code marketplace under the name llama-vscode, published by ggml-org. The README notes it is also available on Open VSX. Note the engine requirement in package.json: VS Code ^1.109.0. On an older build the extension will not activate.
The README states that llama.cpp setup is now automatic on starting llama-vscode, pulling from the official llama.cpp shell script or from brew on Mac and Linux or Winget on Windows. The manual route is still documented. On macOS:
brew install llama.cppOn Windows:
winget install llama.cppOn Linux the README points at the latest binaries from the llama.cpp releases and says to add the bin folder to the path. There is no Linux package command given, so plan on a manual download.
Once llama.cpp is present, open the llama-vscode menu from the status bar or with Ctrl+Shift+M and choose "Install/Upgrade llama.cpp", then "Select/start env...". The env concept groups models; selecting an env selects all the models in it. The README lists predefined envs for different use cases, including completion only and chat plus completion.
The server itself is started with llama serve. The README gives recommended commands by VRAM, for example:
llama serve --fim-qwen-7b-defaultThat one is listed for machines with more than 16GB VRAM. For less than 8GB the README suggests the 1.5B variant. For CPU-only hardware it gives explicit settings, including a port and cache reuse:
llama serve \
-hf ggml-org/Qwen2.5-Coder-1.5B-Q8_0-GGUF \
--port 8012 -ub 512 -b 512 --ctx-size 0 --cache-reuse 256The README warns that CPU-only quality will be significantly lower. Models pulled with -hf land in ~/Library/Caches/llama.cpp/ on macOS, ~/.cache/llama.cpp on Linux, and LOCALAPPDATA on Windows.
With the server running and an env selected, typing in a supported file should produce a grey suggestion. Tab accepts it, Shift+Tab accepts only the first line, Ctrl/Cmd+Right Arrow accepts the next word, and Ctrl+L toggles the suggestion manually.
The agent view, MCP tools and the Telegram bridge
Beyond completion, the extension ships a Llama Agent with its own UI in the Explorer view, opened with Ctrl+Shift+A or from the menu. The README says it works with local models and calls gpt-oss 20B the best choice for now, and that it can also work with external models such as those from OpenRouter.
The agent has nine internal tools according to the README, including custom_tool, which returns the content of a file or a web page, and custom_eval_tool, which lets you write your own tool in JavaScript as a function with an input and a string return value. It supports MCP, using tools from MCP servers installed and started inside VS Code. You can attach the current selection to the context and configure a maximum number of loops.
Two details stand out as worth thinking about before you enable any of this. First, a maximum loop count is the only guard mentioned, so an agent that keeps calling tools is bounded by that number rather than by a cost ceiling. Second, the README documents a Telegram bot that gives access to the llama-vscode agents from a phone, and deep links of the form vscode://ggml-org.llama-vscode?view=agent&prompt=Hello that open VS Code and prefill a prompt. Both are convenient and both widen the surface: a link or a chat message becomes a way to drive an agent that can read files.
The README also mentions dynamic tool loading when Kimi K3 is used from moonshot.ai, which is a provider-specific behaviour rather than something the extension implements generally.
Where llama.vscode is the wrong choice
The extension requires FIM-compatible models; the README points at a Hugging Face collection for these. Point it at a chat-only model and inline completion is not what you will get. That single constraint rules out a large part of the model ecosystem people already have downloaded.
It also requires llama.cpp to be installed and running. The README makes this easier than it used to be, but it is still a second process, a second thing to update, and a second thing that can fail. If your team standardises on a hosted assistant, adding a local runtime per developer is friction you may not want.
Then there is the hardware floor. The README's own CPU-only section says quality will be significantly lower, and the VRAM tiers start at a 1.5B model for under 8GB. A 1.5B completion model is usable for boilerplate and noticeably weak on anything that requires reasoning about a codebase. If you have no GPU and no appetite for small-model output, this is not the tool that will change your mind.
Finally, the context features cut both ways. Ring context pulls from open and edited files and yanked text. On a machine where you keep credentials, customer data or an unrelated proprietary repo open in other tabs, that content can enter the prompt. The README does not document a per-file exclusion mechanism, so the practical control is what you keep open in the window.
How it differs from Continue and Cline
The closest comparisons people search for are Continue and Cline, and the difference is architectural rather than cosmetic.
Continue is a general assistant layer for VS Code and JetBrains. It is built around configurable providers and a broader set of assistant interactions, with local models as one option among many. llama.vscode is narrower and more opinionated: it assumes llama.cpp, it assumes FIM, and it optimises the inline completion path specifically, including the context reuse behaviour it inherits from the server. If you want one assistant that can point at OpenAI, Anthropic or a local server depending on the task, Continue is the more natural fit.
Cline is agent-first. Its centre of gravity is multi-step task execution with tool use, and completion is not the point. llama.vscode has an agent, but the README presents completion as the primary feature and the agent as an addition, with gpt-oss 20B named as the current best local choice for it. If your daily need is an agent that plans and edits across files, Cline is aimed at that problem more directly. If your daily need is a suggestion that appears as you type and costs nothing per keystroke, llama.vscode is aimed at yours.
There is also llama.vim from the same organisation, for Vim and Neovim users. Same idea, different editor.
Maintenance, licence and what upgrading costs you
The repository is not archived, and the last push was on 2026-09-05, which is recent. Releases are frequent and the version numbers are small: v0.0.65 on 2026-09-05, v0.0.64 on 2026-09-01, v0.0.63 on 2026-08-19. A 0.0.x line with releases a few weeks apart tells you the project is moving and that the API surface is not frozen. Expect to update the extension and llama.cpp together rather than independently.
The licence is MIT, which is permissive and places few obligations on how you redistribute or modify it. That is a statement about the licence text, not legal advice; if you plan to ship the extension inside a commercial product, read the LICENSE file and talk to someone qualified.
The real upgrade cost is not the extension, it is the model. Each llama.cpp release can change how models are served, and the README's recommended commands are tied to specific model names such as --fim-qwen-7b-default. When those presets move, your working setup can stop matching the documentation. Pin the llama.cpp version you have validated, and treat an upgrade as a change to test rather than a background update.
Editorial conclusion
Adopt llama.vscode if you already run llama.cpp or are willing to, you want completion that never leaves the machine, and you accept that model choice decides output quality. Skip it if you want a hosted assistant with no local runtime, or if you cannot give a FIM-capable model enough VRAM to be useful. Before committing, verify that llama.cpp is on your PATH, that the env you pick actually starts, and that your chosen model is FIM-compatible.
Frequently asked questions
How do I set up llama.vscode in VS Code?
Install the llama-vscode extension from the VS Code marketplace or Open VSX, then open the llama-vscode menu from the status bar or with Ctrl+Shift+M and choose Install/Upgrade llama.cpp. After that, select an env from the menu with Select/start env. The README states llama.cpp setup is now automatic on starting the extension, with manual brew or winget commands left for reference.
How do I use llama in VS Code for code completion?
Start a llama.cpp server with a FIM-compatible model, select an env in llama-vscode, then type in the editor. Suggestions appear automatically, Tab accepts one, Shift+Tab accepts the first line, Ctrl/Cmd+Right Arrow accepts the next word, and Ctrl+L toggles the suggestion manually.
Can I run llama.cpp on Windows 11 with llama.vscode?
Yes. The README gives winget install llama.cpp for Windows and lists Windows Package Manager as a prerequisite, and it states the extension can install llama.cpp automatically on Mac and Windows. Models downloaded with the -hf flag are stored in LOCALAPPDATA on Windows.
What language is VS Code coded in?
The README does not answer this. The repository's package.json shows llama-vscode is written in TypeScript, but the question is about VS Code itself, which llama.vscode is an extension for rather than a part of.
Does VS Code have a built-in AI?
The README does not answer this. It documents llama.vscode as an extension that provides local LLM-assisted completion, chat and an agent, and notes it can be installed from the VS Code extension marketplace or Open VSX.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ggml-org-llama-vscode)