Model or dataset
microsoft/LLMLingua avatar
microsoft/LLMLingua

LLMLingua: prompt compression for long-context LLM calls

[EMNLP'23, ACL'24] To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.

6,708 stars431 forksPythonMIT

At a glance

What is it?
Microsoft's LLMLingua series compresses prompts before they reach a black-box LLM, trading a small local model pass for fewer billed tokens. It is a good fit for RAG pipelines and long meeting or code transcripts, and a poor fit when every character of the input must survive intact.
Who is it for?
Adopt LLMLingua if you pay per token for a black-box model and your inputs are long, redundant, or retrieved in bulk: RAG contexts, meeting transcripts, chain-of-thought traces. Do not adopt it if the downstream task is sensitive to exact wording, if you need a guarantee that a specific string survives compression, or if you are already running a small model locally under no token budget.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 19 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem LLMLingua targets: prompt length is a bill and a ceiling

The README frames the motivation with three questions: hitting the token limit when asking ChatGPT to summarize a lengthy text, losing earlier instructions after a long context, and paying high GPT-3.5/4 API costs for experiments. Those are the same failure in different clothes. A prompt is charged by length, and a long prompt is also harder for the model to attend to evenly, which is the effect LongLLMLingua calls "lost in the middle".

The intended user is an application developer who sends long inputs to a hosted model and cannot change that model. LLMLingua sits before the API call. It does not fine-tune the target model, does not require access to its weights, and does not change the answer format the caller expects. The README describes the output as compression of the prompt and KV-Cache, with the claim of up to 20x compression and minimal performance loss. That number is a headline from the papers, not a guarantee for any particular prompt, and the repository does not publish a per-task table of what survives at 20x.

How the compression actually works, and why there are three methods

LLMLingua uses a compact, well-trained language model such as GPT2-small or LLaMA-7B to identify and remove non-essential tokens in the prompt. The mechanism is token-level selection driven by a small model's probability estimates, not summarization. Nothing is rewritten; tokens are dropped. That distinction matters when you debug a bad output, because the compressed prompt is a subsequence of the original rather than a paraphrase.

LongLLMLingua is the variant aimed at long-context scenarios. The README says it mitigates the lost-in-the-middle issue and improves RAG performance by up to 21.4% using only a quarter of the tokens, and the repository ships examples/Retrieval.ipynb and examples/RAG.ipynb for that path. LLMLingua-2 is a different design: a BERT-level encoder trained by data distillation from GPT-4 for token classification, described as task-agnostic and 3x to 6x faster than LLMLingua, with better handling of out-of-domain data. If you only read one paragraph of the README, read that one, because it decides which class you import. There is also SecurityLingua, a guardrail model that uses security-aware compression to expose jailbreak intent, with the README claiming 100x lower token cost than comparable guardrail approaches. The release history supports the split: v0.2.1 added LLMLingua-2, and v0.2.0 added customized compression specs.

Installing LLMLingua and running a first compression

The package is published as llmlingua and installs from PyPI. The setup.py lists transformers>=4.26.0, accelerate, torch, tiktoken, nltk and numpy as install requirements, so expect a Torch download on a clean machine. The Makefile shows the maintainers' own path, building a wheel with setup.py bdist_wheel.

bash
pip install llmlingua

The README's quick-start pattern constructs a PromptCompressor with a model name and then calls compress_prompt on a string. The repository's own examples under ./examples, including LLMLingua2.ipynb, RAG.ipynb, OnlineMeeting.ipynb, CoT.ipynb and Code.ipynb, show the exact call shape; open one of those before you copy anything, since the repository documents more than one method and the model name differs between them.

bash
make install

The Makefile's install target builds a wheel from setup.py and then installs the resulting distribution. The Makefile's test target runs pytest across llmlingua and tests, which is the check to run after a dependency bump.

The return value of a compression call is a dictionary; the README's examples read the compressed prompt out of it rather than treating the call as returning a plain string. The compression rate is the main dial: a lower rate means more tokens removed. The v0.2.0 release added customized compression specs, which is the mechanism to use when a plain global rate is too blunt and you want different treatment for different parts of the prompt. For RAG, the repository points at examples/RAG.ipynb and examples/RAGLlamaIndex.ipynb; for meetings, examples/OnlineMeeting.ipynb.

Where LLMLingua is the wrong tool

Token dropping is lossy by construction, and the repository does not document a rollback or a guarantee that any given span survives. If your prompt contains a contract clause, a JSON schema, a function signature, or an identifier that the downstream model must reproduce exactly, a compressor that removes tokens on probability grounds is a risk you cannot audit from the outside. The README does not document a protected-span or must-keep API.

The second limit is the local cost. LLMLingua needs a model loaded in your process, which means GPU memory or CPU time depending on what you pass as model_name. The README's cost-saving argument is about API tokens, not about your own compute. If your prompts are short, the compression pass can cost more than it saves, and if you already run a small local model you may be paying twice for the same hardware.

The third limit is evaluation. The 20x figure and the 21.4% RAG improvement come from papers on their own benchmarks. The README does not ship a harness that tells you what compression rate your task tolerates. You have to build that yourself, and until you do, any rate you pick is a guess.

LLMLingua compared with retrieval-side trimming and with summarization

The nearest alternative is not another compressor but a different place to cut: retrieve fewer or shorter chunks. A retriever that returns three passages instead of ten reduces tokens without touching the text, so nothing inside a surviving passage is distorted. The trade-off is recall: fewer chunks means more chance the answer was in the chunk you dropped, and you cannot recover it at prompt time. LLMLingua keeps more of the source material and thins it token by token, which tends to preserve coverage while risking local damage.

The other alternative is asking a model to summarize or rewrite the context before the main call. That produces fluent, readable text and handles redundancy well, but it needs a generation call, it can introduce facts, and the output is no longer traceable to the source. LLMLingua's output is a subsequence of the input, which makes provenance checks possible in principle. The repository also lists integrations with LangChain and LlamaIndex, so if you already use either framework you can compare the compressor against their existing node post-processors without rebuilding your pipeline.

Maintenance, licence and upgrade cost

The licence is MIT, and setup.py declares it as such. MIT is permissive: you can use the code commercially and modify it, provided the copyright notice and permission notice are retained. That is a statement about the licence text, not legal advice for your product.

The maintenance picture needs care. The repository is not archived, and the last push was on 2026-04-08, which is recent. The releases are older: v0.2.2 landed on 2024-04-09, v0.2.1 on 2024-03-20, and v0.2.0 on 2024-03-13. So the code has moved since the last tagged release, and the changelog in the releases does not describe that work. If you pin a version, pin v0.2.2 and read the commit history for anything after it. setup.py classifies the project as Development Status 3 - Alpha, which is the maintainers' own label and worth taking at face value when you decide how much of your pipeline to hang on it.

Upgrade cost is dominated by the model, not the library. The install requirements pin transformers>=4.26.0 with no upper bound, so a major transformers release can change behaviour under you. The Makefile's test target runs pytest across llmlingua and tests, which is the check to run after any dependency bump.

Editorial conclusion

Adopt LLMLingua if you pay per token for a black-box model and your inputs are long, redundant, or retrieved in bulk: RAG contexts, meeting transcripts, chain-of-thought traces. Do not adopt it if the downstream task is sensitive to exact wording, if you need a guarantee that a specific string survives compression, or if you are already running a small model locally under no token budget. Before committing, run your own evaluation set through the compression spec you intend to ship, check which of the three methods (LLMLingua, LongLLMLingua, LLMLingua-2) your task tolerates, and confirm the rate_limits and compression_rate settings that the README documents for your model. The repository's last push was on 2026-04-08, and the newest release listed is v0.2.2 from 2024-04-09, so treat release notes as the source of truth for what changed since.

Frequently asked questions

What is LLMLingua?

It is a prompt compression library from Microsoft that uses a small language model to remove non-essential tokens before a prompt is sent to a larger LLM. The README describes it as achieving up to 20x compression with minimal performance loss, and the repository also covers LongLLMLingua for long-context and RAG use.

What is prompt compression?

In this project, it means dropping tokens from a prompt so the downstream model receives fewer of them. LLMLingua does this with a compact model that scores tokens and removes the ones judged non-essential, rather than rewriting the text.

How to use LLMLingua?

Install the llmlingua package, construct a PromptCompressor with a model name, and call compress_prompt on your text with a rate. The return value is a dictionary containing the compressed prompt, and the README's examples show the exact keys to read.

What is LLMLingua 2?

LLMLingua-2 is a separate method in the same repository, built on a BERT-level encoder trained by data distillation from GPT-4 for token classification. The README describes it as task-agnostic, 3x to 6x faster than LLMLingua, and better on out-of-domain data; it was added in release v0.2.1.

What does LLM mean in a chat?

In this project's context, LLM stands for large language model, the system that receives the compressed prompt. The README discusses LLMs such as ChatGPT and GPT-4 as the models whose prompt length and prompt-based pricing LLMLingua is designed to address.

Official sources

  1. License: MIT
  2. microsoft/LLMLingua on GitHub
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/microsoft-llmlingua.svg)](https://hysenlabs.com/projects/microsoft-llmlingua)
Community notes

Community notes