Model or dataset
vllm-project/llm-compressor avatar
vllm-project/llm-compressor

llm-compressor: quantizing Hugging Face models for vLLM deployment

Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM

3,827 stars674 forksPythonApache-2.0

At a glance

What is it?
LLM Compressor is a Python library from the vLLM project that applies quantization and pruning algorithms to Transformers models and writes them out in the compressed-tensors format. It is aimed at teams who serve models with vLLM and need smaller weights, not at people who just want a smaller file on disk.
Who is it for?
Adopt llm-compressor if you already serve with vLLM and want weights, activations or KV cache in a lower precision without leaving the Hugging Face ecosystem. Skip it if your serving stack is something else, since the saved compressed-tensors checkpoints are built for vLLM, or if you have no calibration data and no GPU time to spare, because most of the recipes here are data-driven.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap llm-compressor fills between a trained model and a servable one

A model that trains fine is usually too large to serve at the precision it was trained in. Quantization tooling exists, but most of it is either tied to one algorithm or produces checkpoints that a serving engine will not load without conversion work. llm-compressor sits in that gap. It is a Python library that applies compression algorithms to Transformers models and saves the result in the compressed-tensors format, which the README describes as compatible with vLLM.

The intended user is an engineer who has a Hugging Face checkpoint and a vLLM deployment, and who wants the two connected. The README lists weight, activation, KV cache and attention quantization as targets, plus pruning modifiers, and it points to published quantized checkpoints such as RedHatAI/Qwen3.8-27B-INT4 and RedHatAI/Kimi-K3-NVFP4 as examples of the output. If you serve with something other than vLLM, the format choice matters more than the algorithms, and that is the first thing to check before investing time.

How a recipe, a calibration pass and a modifier pipeline produce a compressed checkpoint

The library is organized around recipes and modifiers. A recipe names the algorithm and the precision you want, and modifiers are the units that implement the transforms. The repository layout reflects this: src/llmcompressor/modifiers/pruning/reap is the REAP expert pruning modifier, and the README says that implementation can be used as a template for other expert pruning algorithms. So the modifier interface is the extension point, not a fixed list of supported methods.

Data flow is one-directional. You load a model through the Transformers integration, run a calibration pass over sample data, let the modifiers collect statistics or adjust weights, then save. The README states that models are saved in the compressed-tensors format. For very large models the README mentions DDP and disk offloading, and the examples directory has dedicated folders for disk_offloading and big_models_with_sequential_onloading, which tells you the project treats memory pressure during compression as a first-class problem rather than an afterthought.

The example tree also shows how wide the algorithm surface is. Separate directories exist for awq, autoround, imatrix, quantization_w8a8_fp8, quantization_w8a8_int8, quantization_w4a16, quantization_w4a4_fp4, quantization_w4a4_mxfp4, quantization_w4a8_fp8, quantization_kv_cache, quantization_attention and quantizing_moe. That is a lot of ground, and it means the practical question is rarely whether a method exists, but which one your architecture and your serving engine actually support.

Installing llm-compressor from PyPI and running a first quantization

The README links to PyPI under the name llmcompressor, so installation is a pip install. Because the library applies GPU-heavy transforms, do this in an environment where the CUDA stack matches your PyTorch build; the repository has a pytest-xpu.ini and a test-xpu target in the Makefile, which indicates an Intel XPU path exists alongside CUDA, but the README does not spell out the extra requirements for it.

bash
pip install llmcompressor

After that, the fastest way to see real behaviour is to run one of the scripts in examples/ rather than writing a recipe from scratch. The README points at examples/quantization_w8a8_fp8/nemotron_3_5_lightning_example.py for an FP8 checkpoint and examples/quantization_w4a16/qwen3_8_gptq_awq_example.py for an INT4 one, so a first run looks like this.

bash
python examples/quantization_w8a8_fp8/nemotron_3_5_lightning_example.py

Expect the script to download the base model, run a calibration pass, and write a compressed checkpoint. The exact directory depends on the script's output argument, which the README does not document; read the file before running it so you know where the weights land and how much disk they will take. For larger models, the README's examples for GLM-5.2 describe DDP plus disk offloading used to quantize a model whose full precision footprint is 1.6T of VRAM, so plan for offload space when your model does not fit in memory.

If you build from source rather than PyPI, the Makefile defines the expected workflow: make style runs ruff format and ruff check, make test runs pytest with per-directory ignores driven by the TARGETS variable, and make build produces a wheel through setup.py. Note that setup.py derives the version from setuptools_scm and reads a BUILD_TYPE environment variable whose allowed values are release, nightly and dev, with nightly builds published as pre-releases.

Where llm-compressor is the wrong tool

The clearest limitation is the output format. Checkpoints are saved in compressed-tensors for vLLM, and the README does not describe a path for producing GGUF, ONNX or another runtime's format. If you serve with a different engine, you are either converting afterwards or choosing a different library.

Second, most of the listed algorithms are data-driven. The REAP modifier, for instance, is described in the README as computing a saliency metric from calibration forward pass data to decide which experts to remove. That means you need a calibration set that resembles your traffic. A generic sample can produce a checkpoint that looks fine in file size and behaves badly on your domain, and there is no way to know that from the library alone; you have to evaluate.

Third, the resource floor is real. Disk offloading and DDP exist because compressing large models does not fit on one device, but they are still bounded by the hardware you have. The GLM-5.2 example in the README describes a full precision footprint of 1.6T of VRAM reduced by more than 70 percent after NVFP4 and FP8 quantization with DDP and disk offloading in under two hours. That is a statement about a specific published run on specific hardware, not a guarantee for your model or your machine.

Finally, the README is oriented toward what is new and what has been published. It does not document rollback, nor does it describe how to recover the uncompressed model if a recipe produces a checkpoint you do not like. Keeping your own copy of the source weights is the obvious mitigation, but the library does not manage that for you.

llm-compressor against a general-purpose quantization toolkit

The natural alternative is a broad quantization toolkit such as AutoGPTQ or AutoAWQ, or the quantization utilities bundled with a serving framework. The difference is scope and target. Those tools generally center on one family of algorithms, typically GPTQ or AWQ, and produce checkpoints for whichever runtimes their community has integrated.

llm-compressor covers a wider set of precisions in one library, including W8A8 in int8 and fp8, W4A16, W4A4 in fp4 and mxfp4, and KV cache and attention quantization, and it treats vLLM as the deployment target rather than one option among many. The cost of that focus is portability. A GPTQ checkpoint produced by a general toolkit has a long tail of community tooling around it; a compressed-tensors checkpoint produced here assumes vLLM on the other end.

There is a second difference in how the work is packaged. llm-compressor publishes full example scripts per model family, including MoE-specific ones in examples/quantizing_moe, and the README cites concrete quantized checkpoints for GLM, Qwen, Nemotron, Kimi and Hy3. That is a recipe-driven workflow. General toolkits more often expect you to assemble the pipeline yourself. If you want a supported path for a specific model, the recipe approach saves time. If you want maximum freedom in the output format, it constrains you.

Maintenance, licensing and the cost of staying current

The repository is not archived and the last push was on 2026-09-10, so the codebase is being changed. Recent releases include 0.13.0 on 2026-08-11, 0.12.0.1 on 2026-07-31 and 0.10.0.3 on 2026-07-28. The gap between 0.10.0.3 and 0.12.0.1 is a few days, which suggests patch releases land frequently rather than on a slow cadence.

The upgrade cost is the part to budget for. The README's What's New section is dominated by new quantized checkpoints and new example scripts, which means new model architectures arrive as new examples rather than as changes to a stable core. Pinning a version and reading the release notes before moving is the realistic approach, because a change in a modifier's behaviour can alter the checkpoint you produce without changing your own code.

Licensing is Apache-2.0, with a NOTICE file at the repository root alongside LICENSE. Apache-2.0 is permissive and includes a patent grant, but this article is not legal advice. Two things are worth checking with whoever handles licensing on your side: the terms of the base models you compress, which are separate from this library's licence, and whether the NOTICE file imposes attribution requirements on redistribution of a modified build.

Editorial conclusion

Adopt llm-compressor if you already serve with vLLM and want weights, activations or KV cache in a lower precision without leaving the Hugging Face ecosystem. Skip it if your serving stack is something else, since the saved compressed-tensors checkpoints are built for vLLM, or if you have no calibration data and no GPU time to spare, because most of the recipes here are data-driven. Before committing, verify three things on your own model: that the precision you picked is supported for that architecture, that the saved checkpoint loads in your vLLM version, and that your evaluation scores hold after compression.

Frequently asked questions

What is llm-compressor?

It is a Python library from the vLLM project for applying compression algorithms to Transformers models. It covers weight, activation, KV cache and attention quantization plus pruning, and saves models in the compressed-tensors format for deployment with vLLM.

How do I install llm-compressor?

It is published on PyPI as llmcompressor, so the documented route is a pip install. The README does not list extra requirements beyond that, so match your PyTorch and CUDA build to your machine before running any example.

How do you compress an AI model with llm-compressor?

You pick a recipe and run a calibration pass over sample data, then save. The README points to ready-made scripts in examples/, such as the FP8 Nemotron 3.5 Lightning example and the INT4 Qwen3.8 example, which is the fastest way to see the full flow.

What is compression in AI?

In this project's terms it means reducing the precision of weights, activations, KV cache or attention, and sometimes removing experts outright, so the model takes less memory and runs on fewer devices with vLLM.

Official sources

  1. License: Apache-2.0
  2. Project website
  3. README
  4. Releases
  5. vllm-project/llm-compressor on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/vllm-project-llm-compressor.svg)](https://hysenlabs.com/projects/vllm-project-llm-compressor)