Model or dataset
ModelTC/LightCompress avatar
ModelTC/LightCompress

LightCompress: a Python toolkit for quantizing LLMs, VLMs and video diffusion models

[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.

749 stars78 forksPythonApache-2.0

At a glance

What is it?
LightCompress, formerly LLMC, compresses large language models, vision-language models and video generative models with AWQ, GPTQ, SmoothQuant and token reduction. It is aimed at engineers who already have a GPU box and a model checkpoint, not at people who want a hosted service.
Who is it for?
Adopt LightCompress if you already run inference on vLLM, SGLang or lightx2v and want to produce quantized checkpoints yourself instead of downloading someone else's. Do not adopt it if you need a supported service, a stable API surface, or compression for a model family outside the ones the documentation lists.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 139 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 22, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What LightCompress does that a plain quantization script does not

Most quantization code you find in the wild is a single script for a single model family. LightCompress, formerly LLMC, is the opposite: a repository that collects many compression algorithms behind a shared configuration layer and a shared export path. The README describes it as an "off-the-shell tool designed for compressing aigc models(LLM, VLM, Diffusion ...)", and the release history backs that up. The v1.4.0 release in February 2025 was still branded LLMC; v1.5.0 in November 2025 carries the LightCompress name.

The audience is narrow and specific. You need a GPU, a model checkpoint you are allowed to modify, and a reason to own the quantization step rather than consume a pre-quantized model from a hub. Typical cases: you fine-tuned a model and want to ship it at INT4, you need FP8 weights for a specific serving stack, or you are researching how a compression algorithm behaves on your own data. The repository also carries benchmark and evaluation tooling, and the topics list mentions vLLM, SGLang-adjacent deployment and lm-evaluation-harness, so the intended workflow runs from a dense checkpoint through compression to a serving backend.

The scope claim is broad: LLMs, VLMs, and video generative models. That breadth is the selling point and also the source of the main risk, because a toolkit that claims to cover Mixtral, DeepSeek-V3, Qwen2VL and Wan2.1 cannot be equally deep on all of them.

How the compression pipeline is organized

The repository layout tells you most of the architecture. There is an llmc/ package directory, a configs/ directory, a scripts/ directory, an examples/ directory with an examples/backend/ subdirectory, and a vendored lm-evaluation-harness alongside a .gitmodules file that suggests further submodules. Compression is driven by configuration files rather than by long command lines, which is why configs/ and scripts/ exist as separate top-level entries.

The algorithms named in the README and topics split into two families. The first is weight and activation quantization: integer and floating-point quantization, AWQ, GPTQ, SmoothQuant, QuaRot, and static per-tensor activation quantization. The second is token reduction for vision-language models, described as covering token merging, token pruning and token reduction, with the August 2025 announcement citing over 20 algorithms across token reduction and quantization.

The output side matters as much as the compression side. The README states that LightCompress exports "real quantized(INT4, INT8)" and FP8 (E4M3, E5M2) models for vLLM, SGLang, AutoAWQ and MLC-LLM, and that the Wan2.1 video models export INT8/FP8 weights compatible with lightx2v. That export path is the reason the toolkit exists as a pipeline rather than a set of scripts: the checkpoint it produces is meant to be loaded by an inference engine without a conversion step in between, which the README calls out for DeepSeek FP8 weights specifically.

Installing LightCompress and running a first quantization

The README does not give a pip install line. It points at two Docker images, one on Docker Hub and one on Alibaba Cloud for users in mainland China, and it recommends Python 3.11 for local development because that matches the project's Docker images and CI configuration.

bash
# docker hub: https://hub.docker.com/r/llmcompression/llmc
docker pull llmcompression/llmc:pure-latest

# aliyun docker: registry.cn-hangzhou.aliyuncs.com/yongyang/llmcompression:[tag]
docker pull registry.cn-hangzhou.aliyuncs.com/yongyang/llmcompression:pure-latest

After the pull completes you have an image tagged llmcompression/llmc:pure-latest. The repository also ships a Dockerfile and a Dockerfile_cu124 at the top level, and that Dockerfile builds from nvidia/cuda:12.1.0-cudnn8-devel-ubuntu22.04, installs python3.11, then installs FlashAttention and fast-hadamard-transform from local source directories before installing requirements/runtime.txt. The fast-hadamard-transform dependency is a hint about which algorithms expect Hadamard rotations.

For a source install, requirements.txt is a one-line file that pulls in requirements/runtime.txt, so the dependency list lives under requirements/. The README does not document the exact pip invocation, so the Docker route is the one the project actually supports.

The first real use is a config-driven run. The repository keeps its configs in configs/ and its entry scripts in scripts/, and the documentation site at llmc-en.readthedocs.io is where the README sends readers for algorithm-specific instructions, including a Best Practice page. Because the README does not inline a complete config for any single model, the honest first step is to pick a config from configs/ that matches your model family, read it, and run it through the project's script rather than inventing one. The README does not document rollback or a dry-run mode, so run against a copy of the checkpoint.

Where LightCompress is the wrong tool

The documentation is organized around specific supported model families, not around arbitrary checkpoints. If your model is not one of the ones the docs cover, you are adapting the toolkit, not using it. The README names DeepSeek-V2/V2.5, DeepSeek-V3, DeepSeek-R1, DeepSeek-R1-zero, Qwen2VL, Llama3.2, Llama-3.1-405B, Mixtral, InternLM2 and Wan2.1. A custom architecture with unusual attention or a non-standard MoE layout is out of scope until someone writes the integration.

The hardware floor is real. The README states that AWQ and RTN quantization for the 671B DeepSeek models can run on a single 80GB GPU. That is the floor for the largest supported models, not a general requirement, but it tells you the toolkit is not aimed at laptop or single-consumer-card work for frontier-scale MoE models. Smaller models will fit on smaller cards, and the README does not publish a general memory table, so you have to reason from the model size and the algorithm.

There is also a maintenance-shape caveat. The last push to the default branch was on 2026-05-14, roughly four months before this writing, and the most recent release is v1.5.0 from 2025-11-19. That is a research-lab cadence tied to paper deadlines, not a vendor cadence. If you need a deprecated-config policy or a versioned API guarantee, the repository does not offer one. The README documents no rollback path for a failed compression run, so keep the original checkpoint.

LightCompress compared with llm-compressor and AutoAWQ

The nearest alternative in spirit is llm-compressor, which also targets the vLLM serving path and also works from recipes. The difference is scope. llm-compressor concentrates on LLM quantization recipes for the vLLM ecosystem. LightCompress covers LLMs but extends the same config-driven approach to vision-language token reduction and to video generative models, which is a different problem: token reduction changes what the model sees at inference time, not just the numeric precision of its weights. If your work is text-only LLM quantization and you want the tightest integration with vLLM's own tooling, llm-compressor is the more focused choice. If you need one pipeline that also handles Qwen2VL token pruning or Wan2.1 video quantization, LightCompress is the one that claims that coverage.

AutoAWQ is a different kind of alternative: it is an implementation of one algorithm family rather than a collection. The README lists AutoAWQ as one of the export targets LightCompress can write to, which makes the relationship partly complementary. You can use LightCompress to produce AWQ-format weights and AutoAWQ to load them. Choosing AutoAWQ directly makes sense when AWQ is the only method you will ever use and you want the smallest possible dependency surface.

The third comparison point is a pre-quantized checkpoint from a model hub. The README links to INT4 and INT8 Llama-3.1-405B weights that were quantized with this toolkit, which means for that specific model you can skip the compression step entirely. Running the toolkit yourself only pays off when your model or your calibration data differs from what is already published.

Licence, upgrade cost and what the repository does not promise

LightCompress is Apache-2.0, and the repository carries the standard LICENSE file at the top level. Apache-2.0 is permissive and includes a patent grant, which matters for a toolkit that implements published quantization methods. It does not resolve the licence of the models you compress, of the calibration datasets you feed in, or of the inference backends you export to. Those are separate questions and the repository does not address them. Nothing here is legal advice.

The upgrade cost is the more practical concern. The project renamed itself from LLMC to LightCompress between v1.4.0 and v1.5.0, and the README keeps a notice about the rename at the top. Documentation URLs still use the llmc- prefix (llmc-en.readthedocs.io and llmc-zhcn.readthedocs.io), and the Docker Hub image is still llmcompression/llmc. Expect the old name to persist in paths and tags for a while, and expect config files written against LLMC-era releases to need review when you move to v1.5.0.

The release cadence is uneven: v1.3.0 in October 2024, v1.4.0 in February 2025, v1.5.0 in November 2025. Feature announcements in the README do not always line up with releases, since VLM support and Wan2.1 quantization were announced in 2025 outside the tagged versions. If you pin to a release, you may be pinning to something older than the features the README describes.

Editorial conclusion

Adopt LightCompress if you already run inference on vLLM, SGLang or lightx2v and want to produce quantized checkpoints yourself instead of downloading someone else's. Do not adopt it if you need a supported service, a stable API surface, or compression for a model family outside the ones the documentation lists. Before committing, verify three things: that your target model appears in the docs, that your GPU memory matches the algorithm you picked (the repository states AWQ and RTN for DeepSeek-V3 at 671B fit on a single 80GB GPU), and that the exported checkpoint loads in your chosen backend. The last push was on 2026-05-14, so the code is recent but the project is a research toolkit, not a product with an SLA.

Frequently asked questions

What is LightCompress used for?

It compresses large models, including LLMs, VLMs and video generative models, using quantization and token reduction algorithms, and exports the result as real quantized INT4, INT8 or FP8 weights for inference backends such as vLLM, SGLang, AutoAWQ, MLC-LLM and lightx2v.

How do I install LightCompress?

The README gives Docker as the supported route, with llmcompression/llmc:pure-latest on Docker Hub and registry.cn-hangzhou.aliyuncs.com/yongyang/llmcompression:pure-latest on Alibaba Cloud for users in mainland China. It recommends Python 3.11 for local development and installation, but does not give a pip install command.

Is LightCompress the same project as LLMC?

Yes. The README states that the repository was formerly known as LLMC and has been renamed to LightCompress. The v1.4.0 release is still labelled LLMC while v1.5.0 uses the LightCompress name, and the documentation URLs and Docker image name still use the old llmc spelling.

Which models does LightCompress support for quantization?

The README names DeepSeek-V2/V2.5, DeepSeek-V3, DeepSeek-R1, DeepSeek-R1-zero, Qwen2VL, Llama3.2, Llama-3.1-405B, Mixtral, InternLM2 and the Wan2.1 video generation series. It also states that AWQ and RTN quantization for the 671B DeepSeek models can run on a single 80GB GPU.

Does LightCompress support vision-language models?

Yes. An August 2025 announcement in the README states that the project open-sourced a compression solution for vision-language models covering over 20 algorithms across token reduction and quantization, with details in the token reduction section of the documentation.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. ModelTC/LightCompress on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/modeltc-lightcompress.svg)](https://hysenlabs.com/projects/modeltc-lightcompress)