Library / SDK
nunchux-ai/nunchaku avatar
nunchux-ai/nunchaku

Nunchaku: 4-bit Diffusion Inference Under SVDQuant

[ICLR2025 Spotlight] SVDQuant: Absorbing Outliers by Low-Rank Components for 4-Bit Diffusion Models

3,957 stars279 forksPythonApache-2.0

At a glance

What is it?
Nunchaku is a 4-bit inference engine for diffusion models, built on the SVDQuant method from the ICLR 2025 spotlight paper. It targets engineers who want FLUX and Qwen-Image generation to fit in less VRAM without retraining the base model.
Who is it for?
Use Nunchaku if you run FLUX or Qwen-Image generation on a single GPU and VRAM, not raw speed, is the binding constraint; the asynchronous offloading path is documented to bring Qwen-Image down to as little as 3 GiB. Do not adopt it if your models are outside the supported families, or if you need a pure-Python install with no CUDA toolchain.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 26 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What the 4-bit diffusion engine actually replaces

Diffusion transformers are memory-hungry at inference. The weights are large, and the activations during a multi-step denoising loop add more on top. Nunchaku's answer is to keep the model in 4-bit form and run the denoising loop against those weights, rather than loading full precision and relying on offloading to swap tensors in and out.

The project is the reference implementation of SVDQuant, described in the README as a method that absorbs outliers using low-rank components. That framing matters. Naive 4-bit quantization of a diffusion transformer tends to fail because a small number of activation channels carry extreme values; the low-rank branch exists to absorb those outliers so the 4-bit main path stays usable. The quantization library itself lives in a separate repository, DeepCompressor, which the README points to.

The audience is narrow and specific. If you are serving FLUX.1 or Qwen-Image variants on a workstation GPU and you keep hitting an out-of-memory wall, this is aimed at you. If you are training or fine-tuning diffusion models, it is not: this is an inference engine, and the README describes it as such.

The architecture: a CUDA extension with a Python front end

The repository is not a pure-Python package. The build backend is setuptools with a custom build extension, and setup.py imports CUDAExtension and BuildExtension from torch.utils.cpp_extension. That means installing from source compiles CUDA kernels, and it means the install step depends on a working nvcc.

setup.py probes nvcc for its version and derives which SM targets to build. The source reads the nvcc release string, then compares it against thresholds: sm120 support requires nvcc 12.8 or later, and sm121 requires 13.0 or later. There is also an install mode environment variable, NUNCHAKU_INSTALL_MODE, which defaults to FAST. In FAST mode the build enumerates the local CUDA devices and targets their capabilities rather than compiling for every architecture. That is a sensible default for a workstation install and a poor one for building a wheel you intend to ship to machines with different GPUs.

The Python side is layered. The README notes that a Python backend became available, with FLUX models under nunchaku/models/transformers/ and a modular 4-bit linear layer at nunchaku/models/linear.py. The repository also carries a src/ directory alongside the nunchaku/ package, which is consistent with a compiled core wrapped by Python modules. The examples directory is the practical entry point: it holds per-model scripts such as examples/flux.1-dev.py, examples/flux.1-kontext-dev.py, and examples/v1/qwen-image.py.

Installing Nunchaku and running a first FLUX generation

The README does not spell out a single canonical install command in the excerpt available here; it links to the documentation site at nunchaku.tech and to tutorial videos in English and Chinese, and the project publishes nightly wheels under version tags such as v1.3.0dev20260306. Treat the documentation site as the source of truth for the exact command for your platform.

What the repository does pin down is the Python and dependency floor. pyproject.toml requires Python 3.10 or newer, and lists torch>=2.7 in the build requirements. Runtime dependencies include diffusers>=0.36, transformers>=4.54, peft>=0.17, accelerate>=1.9, and huggingface-hub>=0.34. A source install therefore goes through setuptools and compiles the CUDA extension, so nvcc must be on PATH and the CUDA toolkit version must satisfy the thresholds in setup.py. If the build fails at the nvcc detection step, the error comes from that check, not from Python packaging.

Once installed, the example scripts are the fastest way to a real run. The FLUX.1-dev example is the natural starting point, and the repository lists it at examples/flux.1-dev.py. Expect the script to download the corresponding 4-bit checkpoint from Hugging Face or ModelScope and write generated images to disk. The README also points to a ComfyUI integration in a separate repository, nunchux-ai/ComfyUI-nunchaku, for users who would rather work through a node graph than a script.

Where Nunchaku is the wrong tool

The hard boundary is model coverage. Nunchaku is not a general quantization pass you can point at any diffusion checkpoint. The README lists support for FLUX.1-dev and its canny, depth, fill, and Kontext variants, FLUX.1-Krea-dev, Qwen-Image, Qwen-Image-Edit and Qwen-Image-Edit-2509, and Z-Image-Turbo. Each supported model has a matching 4-bit release on Hugging Face or ModelScope. If your model is not on that list, there is nothing here for you, regardless of how much VRAM you would save.

The second boundary is the build. A CUDA toolchain is not optional for a source install. On a machine without nvcc, or with a CUDA version below the thresholds setup.py checks, the install stops. The project publishes nightly wheels, which sidesteps compilation, but a wheel is tied to a Python version and a CUDA combination; if yours is not published, you are back to compiling.

The third is scope. This is inference only. There is no training path, and the README directs anyone working on the underlying quantization method to DeepCompressor rather than this repository. Teams that need to quantize a custom fine-tune themselves are outside the intended workflow.

How Nunchaku differs from bitsandbytes and torchao

The obvious comparison is with general-purpose quantization libraries such as bitsandbytes and torchao. Those take a broad approach: they expose quantization as a configuration applied to a model, and they support a wide range of architectures because they do not assume anything specific about the model family.

The difference in Nunchaku is that the low-rank branch is part of the method, not an option. SVDQuant's premise is that a 4-bit main path alone loses too much on diffusion transformers because of outlier channels, so a low-rank correction absorbs them. That is a heavier design than a plain weight-only quantization pass, and it is why the project ships per-model 4-bit checkpoints rather than a runtime flag.

The trade is coverage for fit. bitsandbytes and torchao will accept a model Nunchaku has never seen. Nunchaku will, for the models it does support, produce a checkpoint that is already quantized and ready to load, with example scripts that run it. If your model is supported, the second path requires less experimentation. If it is not, the first path is the only one available.

Maintenance cadence, licence, and upgrade cost

The repository is not archived, and the last push was on 2026-09-06. Recent releases are nightlies under the v1.3.0dev series, the most recent tagged v1.3.0dev20260306 on 2026-03-06. The gap between the last release tag and the last push is worth noting: release tags are not the only signal of activity, and the README's news section carries entries dated later than the most recent tag, including a v1.2.0 announcement on 2026-01-12.

The project is licensed Apache-2.0, with the licence file at LICENCE.txt in the repository root. That is a permissive licence, and it is the same licence family used by much of the surrounding ecosystem. It does not, by itself, settle the question of what licence the model weights carry: the 4-bit checkpoints are distributed separately on Hugging Face and ModelScope, and their terms are set there, not by this repository. Check the model card for the checkpoint you intend to use.

Upgrade cost is dominated by the compiled extension. Because setup.py builds against the local CUDA toolkit and, in FAST mode, the local GPU capabilities, moving to a new machine or a new CUDA version means a rebuild. Nightly wheels reduce that cost when a matching wheel exists. The version string is dynamic, so the installed package version reflects the build rather than a fixed value in pyproject.toml.

What to check before you commit to a deployment

Start with the model list. Confirm that the exact checkpoint you plan to serve has a released 4-bit counterpart under the nunchaku-ai organisation on Hugging Face or nunchaku-tech on ModelScope. A supported architecture is not the same as a supported checkpoint.

Then check the toolchain. Run nvcc --version and compare it against the thresholds in setup.py: 12.8 for sm120 targets and 13.0 for sm121. If your CUDA is older, a source build will fail at the detection step before it compiles anything.

Finally, decide between a wheel and a source build. The nightly releases are tagged with dates, so the wheel you install is tied to a specific build. If your Python version is 3.10 through 3.14 and a wheel exists for your CUDA combination, that is the shorter path. If not, plan for the compile time and for the NUNCHAKU_INSTALL_MODE setting, which defaults to FAST and therefore targets only the GPUs present on the build machine.

Editorial conclusion

Use Nunchaku if you run FLUX or Qwen-Image generation on a single GPU and VRAM, not raw speed, is the binding constraint; the asynchronous offloading path is documented to bring Qwen-Image down to as little as 3 GiB. Do not adopt it if your models are outside the supported families, or if you need a pure-Python install with no CUDA toolchain. Before committing, check the CUDA and nvcc versions against the build requirements, confirm the exact model checkpoint you plan to run has a released 4-bit counterpart, and verify that the wheel for your Python and CUDA combination exists rather than assuming a source build will succeed.

Frequently asked questions

How do I install Nunchaku for ComfyUI?

The README links to a separate repository, nunchux-ai/ComfyUI-nunchaku, for the ComfyUI integration. The core package is installed from this repository, and the ComfyUI nodes are distributed separately.

How do I install the Nunchaku wheel?

The project publishes nightly wheels under version tags such as v1.3.0dev20260306. A wheel avoids the CUDA compile step, but it is tied to a Python version and a CUDA combination, so check that yours is covered before relying on it.

What is Nunchaku ComfyUI?

It is the ComfyUI integration for the Nunchaku 4-bit inference engine, hosted in a separate repository at nunchux-ai/ComfyUI-nunchaku. The README lists ComfyUI as one of the supported front ends alongside the Python example scripts.

How do I use Nunchaku with ComfyUI?

The README points to the separate ComfyUI-nunchaku repository for node-based use. The README also notes that v1.2.0 added LoRA support with native ComfyUI nodes and compatibility with ComfyUI 0.7.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. nunchux-ai/nunchaku on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nunchux-ai-nunchaku.svg)](https://hysenlabs.com/projects/nunchux-ai-nunchaku)