llm-ls: a Rust LSP server that puts LLM completions behind any editor
LSP server leveraging LLMs for code completion (and more?)
At a glance
- What is it?
- llm-ls is a Language Server Protocol server that talks to Hugging Face Inference API, text-generation-inference, ollama and OpenAI-compatible endpoints so IDE extensions can stay thin. It is a work in progress, and the README says so.
- Who is it for?
- Adopt llm-ls if you already run an OpenAI-compatible endpoint, ollama, or Hugging Face text-generation-inference and you want the same backend to serve llm.nvim, llm-vscode and llm-intellij without writing HTTP and tokenization code three times. Do not adopt it if you need a stable, documented protocol: the README labels the project a work in progress, and the newest release listed is 0.5.3 from 2024-05-24, so the client extensions are the real integration contract.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly Rust, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem llm-ls solves: one LLM backend, many editors
Every editor extension that wants inline completions ends up reimplementing the same three things: build a prompt from the current file, cut it to the model's context window, and call an HTTP endpoint. llm-ls moves that work into a separate process that speaks the Language Server Protocol, so the extension only has to forward editor state and render the result. The README states the goal directly: llm-ls exists "to provide a common platform for IDE extensions to be build on", and it "takes care of the heavy lifting with regards to interacting with LLMs so that extension code can be as lightweight as possible".
The audience is therefore narrow and specific. It is not an end user tool. It is infrastructure for people writing or maintaining editor plugins, and for users of the three extensions the README marks as compatible: llm.nvim, llm-vscode and llm-intellij. A fourth, jupytercoder, is listed unchecked, meaning the README does not claim it works. If you are not using one of those clients and you are not writing one, running llm-ls by hand gives you an LSP server with no UI.
How the prompt, the context window and the AST parser fit together
The mechanism has three stages, and each one is a place where behaviour can surprise you.
First, prompt construction. The server uses the current file as context and can operate in fill-in-the-middle mode or not, depending on configuration. Fill-in-the-middle means the prompt carries both the text before the cursor and the text after it, which is what makes mid-file completion possible rather than only appending at the end.
Second, window management. The README says llm-ls "makes sure that you are within the context window of the model by tokenizing the prompt". This is the detail that separates it from a naive wrapper: the server counts tokens rather than characters, so the truncation point depends on the model's tokenizer, and switching models changes what survives truncation.
Third, completion shape. llm-ls "parses the AST of the code to determine if completions should be multi line, single line or empty". Empty is a real outcome: the server can decide that no completion should be offered at all, rather than sending a request and letting the model produce something unhelpful. That decision is driven by syntax, not by the model.
The roadmap in the README lists what is not there yet, and the list is honest: multi-file workspace context, a suffix_percent setting to control the prefix-to-suffix token ratio, a context window fill percent or a switch from context_window to max_tokens, and filtering of bad suggestions such as repetitive output or output identical to the text below the cursor. Until suggestion filtering lands, a model that repeats itself will repeat itself in your editor.
Backends: Hugging Face Inference API, text-generation-inference, ollama, OpenAI-compatible servers
llm-ls does not ship a model. It is a client, and the README names four backend families: Hugging Face's Inference API, Hugging Face's text-generation-inference, ollama, and OpenAI-compatible APIs, with the python llama.cpp server bindings given as an example of the last category.
That last category is the widest door. Any server exposing an OpenAI-compatible route can sit behind llm-ls, which is why the self-hosted topic on the repository makes sense: with ollama or a local llama.cpp server, no prompt text leaves the machine. The README is explicit about the outbound side too. It notes that llm-ls "does not export any data anywhere (other than setting a user agent when querying the model API)", and that telemetry is written to a log file at ~/.cache/llm_ls/llm-ls.log when the log level is set to info.
Read that sentence carefully. Telemetry collection and telemetry export are different things. The server gathers request and completion information locally, which the README says "can enable retraining", and it stays on disk. If you set the log level to info, that file will contain prompt and completion text. On a shared machine, that is a file worth knowing about.
Installing llm-ls and getting a first completion
The README does not contain an install section, so there is no documented set of commands to reproduce here. What the repository layout does show is a Cargo workspace: Cargo.toml declares members ["xtask/", "crates/*"], and the top level holds crates/, xtask/, Cargo.lock and Cargo.toml. That means building from source is a cargo build in the workspace root, and xtask exists for distribution tasks, since Cargo.toml mentions that miniz_oxide's opt-level is raised "to speed up `cargo xtask dist`".
In practice most users will not build it. They install the client extension, and the extension obtains the server binary. The three extensions the README marks compatible are llm.nvim, llm-vscode and llm-intellij.
What you configure is not the server binary directly but the client. In llm.nvim, for example, the plugin needs to know which backend to talk to. The README does not document the configuration keys, so treat the extension's own README as the source of truth for the exact field names. The shape of the decision is this: point the client at a Hugging Face Inference API endpoint, a text-generation-inference deployment, an ollama instance, or an OpenAI-compatible base URL, and supply whatever token that backend needs.
A minimal local setup uses ollama, since it runs on the same machine and needs no API key. The README lists ollama as a supported backend but does not give a port or an environment variable, so the address is whatever your ollama installation listens on. Start ollama, pull a code model, and point the extension at it:
ollama serve
ollama pull <model>What you should see is the extension registering with llm-ls and, on the first keystroke that triggers a completion, a request in the ollama server log. If nothing appears, the problem is almost always the client-to-server handshake rather than the model.
To confirm the server is running at all, check the log path the README gives. With the log level at info, the file is created here:
ls ~/.cache/llm_ls/llm-ls.logIf that file exists and grows while you type, llm-ls is receiving LSP traffic. If it does not exist, the client never connected.
Where llm-ls is the wrong tool
The README opens with an important callout: "This is currently a work in progress, expect things to be broken!" That is not marketing hedging, and it should govern your decision. There is no stable protocol version documented, no compatibility guarantee, and no rollback procedure described in the README. The release list shows 0.5.3 on 2024-05-24, preceded by 0.5.2 and 0.5.1 in February 2024, so the tagged release cadence is not a signal you can plan a migration around.
The repository itself is not archived, and the last push was on 2026-05-26, so commits continue. But a recent push is not the same as a stable interface, and the README's own roadmap lists unfinished work rather than completed guarantees.
Concretely, llm-ls is the wrong choice in three situations. If you want a completion experience that never produces repeated or duplicated text, the suggestion filter is still on the roadmap, so nothing in the server removes a bad suggestion before it reaches your editor. If you work in a large codebase where completions need context from other files, the README lists multi-file workspace context as a roadmap item, not a feature, so the prompt is built from the current file only. And if you need a documented, versioned API to build a product on, the README does not describe one, which means your extension is coupled to implementation details rather than a contract.
There is also a language dependency hidden in the AST step. The server parses the AST to decide between multi-line, single-line and empty completions. The README does not enumerate which languages that parser handles, so for a language outside its coverage the completion-shape decision is the part to test before trusting it.
How llm-ls differs from GitHub Copilot and from editor-native assistants
The closest alternative in daily use is GitHub Copilot, and the difference is architectural rather than cosmetic. Copilot is a hosted service with a fixed set of clients: the model, the prompt construction and the serving stack are all on GitHub's side, and you choose a subscription rather than a model. llm-ls inverts that. It is a local process that you point at a backend you choose, which can be ollama or a llama.cpp OpenAI-compatible server running on your own hardware.
That inversion buys two things. First, model choice: switching from a Hugging Face Inference API model to a local ollama model is a client configuration change, not a vendor decision. Second, data path: with a local backend, prompt text does not leave the machine, and the README confirms the server itself exports nothing beyond a user agent on model API calls.
What it costs is everything Copilot hides. You run the backend, you manage the model, you accept that the README calls the project a work in progress, and you accept that a recent release is 0.5.3 from May 2024. Copilot is a product with a support contract; llm-ls is a component in a stack you assemble. For a team already running text-generation-inference, that trade is reasonable. For someone who wants completions to simply work after installing an extension, it is not.
Licence, build cost and what upgrading actually involves
llm-ls is licensed under Apache-2.0, declared in both the repository's licence field and the Cargo.toml workspace package. Apache-2.0 is a permissive licence with an explicit patent grant and a requirement to preserve notices; it does not impose copyleft on your extension code. This is a description of the licence text, not legal advice, and if you redistribute a modified server you should read the notice and attribution clauses yourself.
Upgrade cost is dominated by the fact that llm-ls is a server consumed by three client extensions. The version that matters is the one your extension expects, not the one on the repository's release page. Because the README does not document a protocol version or a compatibility matrix, pinning is the practical answer: pin the client extension version, and if you build the server from source, pin the commit.
Building from source has its own cost, and Cargo.toml hints at it. The workspace sets debug = 0 in the release profile and comments that disabling debug info speeds up builds, and it raises miniz_oxide to opt-level 3 specifically to speed up cargo xtask dist. Those are the adjustments of a project whose release build was slow enough to need tuning. Expect a Rust toolchain build rather than a quick script.
On the running side, the log file at ~/.cache/llm_ls/llm-ls.log is the operational artifact. With the log level at info it accumulates request and completion data. That is the file to check when something breaks, and the file to rotate or delete when it grows.
Editorial conclusion
Adopt llm-ls if you already run an OpenAI-compatible endpoint, ollama, or Hugging Face text-generation-inference and you want the same backend to serve llm.nvim, llm-vscode and llm-intellij without writing HTTP and tokenization code three times. Do not adopt it if you need a stable, documented protocol: the README labels the project a work in progress, and the newest release listed is 0.5.3 from 2024-05-24, so the client extensions are the real integration contract. Before wiring it into a daily driver, verify three things: which release tag your editor extension pins, whether your backend speaks the OpenAI-compatible route or needs the Hugging Face Inference API route, and whether the AST-driven completion mode matches your language, because the tokenizer and the AST parser decide what the server sends.
Frequently asked questions
What is llm-ls?
It is a Language Server Protocol server written in Rust that uses LLMs to produce code completions. It is meant as a shared platform for IDE extensions, so llm.nvim, llm-vscode and llm-intellij can stay thin and let the server handle prompt building, tokenization and backend calls.
Which backends can llm-ls talk to?
The README lists Hugging Face's Inference API, Hugging Face's text-generation-inference, ollama, and OpenAI-compatible APIs such as the python llama.cpp server bindings. With ollama or a local OpenAI-compatible server, the model runs on your own machine.
Does llm-ls send my code anywhere?
The README states that llm-ls does not export any data anywhere other than setting a user agent when querying the model API. Telemetry about requests and completions is stored locally in ~/.cache/llm_ls/llm-ls.log when the log level is set to info.
Which editor extensions work with llm-ls?
The README marks llm.nvim, llm-vscode and llm-intellij as compatible. jupytercoder is listed unchecked, so the README does not claim support for it.
Is llm-ls ready for production use?
The README carries a callout saying it is currently a work in progress and to expect things to be broken. Several capabilities, including multi-file workspace context and filtering of bad suggestions, are listed as roadmap items rather than finished features.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/huggingface-llm-ls)