ComfyUI-nunchaku: 4-bit Diffusion Models in ComfyUI
ComfyUI Plugin of Nunchaku
At a glance
- What is it?
- ComfyUI-nunchaku is the ComfyUI plugin for the Nunchaku inference engine, which runs 4-bit models quantized with SVDQuant. It targets FLUX, Qwen-Image and Z-Image workflows that no longer fit comfortably in VRAM.
- Who is it for?
- Adopt ComfyUI-nunchaku if your ComfyUI graphs are FLUX, Qwen-Image or Z-Image based and VRAM, not model quality, is what stops you. Skip it if you are not prepared to keep a separately built Nunchaku wheel in step with a plugin that has already moved through v1.0.0, v1.1.0, v1.2.0 and v1.2.1.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 27 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ComfyUI-nunchaku solves, and for whom
Diffusion checkpoints are large, and the usual answer to running them is to buy more VRAM or accept slower offloading. ComfyUI-nunchaku takes a different route: it is the ComfyUI plugin for Nunchaku, described in the README as an efficient inference engine for 4-bit neural networks quantized with SVDQuant. The quantization itself is not done here. The README points at the separate Nunchaku repository for the engine and DeepCompressor for the quantization library, and the pre-quantized weights live on Hugging Face and ModelScope under the nunchaku-ai and nunchaku-tech organizations.
The audience is therefore narrow and specific. You are already running ComfyUI, your graphs are built around FLUX, Qwen-Image or Z-Image, and the thing blocking you is memory rather than model choice. The release notes give the shape of that: the v1.0.0 announcement states that Qwen-Image supports asynchronous offloading, cutting Transformer VRAM usage to as little as 3 GiB with no performance loss. That is the claim to design around, and it is a claim from the project, not an independent measurement.
If you only run SD 1.5 or SDXL checkpoints, this plugin adds a dependency chain for no benefit. The nodes are written for the Nunchaku engine and the quantized model families the project has published.
How the plugin, the wheel and the quantized weights fit together
Three pieces have to line up. The first is the Nunchaku engine itself, distributed as a wheel that must match your Python version, PyTorch build and CUDA. The second is this repository, the ComfyUI-side plugin that registers the loader and sampler nodes and patches the model. The third is a pre-quantized checkpoint in SVDQuant 4-bit form.
The repository layout reflects that split. model_configs/ and model_patcher/ handle per-architecture configuration and the patching of the loaded model, model_base/ and models/ hold the model-side code, nodes/ holds the ComfyUI nodes, and wrappers/ sits around the engine calls. The README also mentions a NunchakuWheelInstaller node, added in v0.3.1, which exists precisely because getting the right wheel is the part users get wrong.
Because quantization happens offline, there is no on-the-fly conversion path documented here. You either download a published 4-bit checkpoint or you produce one with DeepCompressor. That is the central design trade-off: the plugin is fast because the hard numerical work was done before you installed anything, and it is limited because its supported model list is whatever the project has quantized and published.
Installing ComfyUI-nunchaku and running a first FLUX graph
The plugin is a normal ComfyUI custom node package. The README does not spell out a clone command, but the repository is laid out for ComfyUI's custom_nodes directory, and the official documentation at nunchaku.tech/docs/ComfyUI-nunchaku/ is where the project says to look for guides. The package metadata requires Python >=3.10,<3.14, so check that first.
pip install -r requirements.txtThat installs the runtime dependencies listed in requirements.txt: diffusers, transformers, sentencepiece, protobuf, huggingface_hub, tomli, peft, accelerate, insightface, opencv-python, facexlib, onnxruntime and timm. After a ComfyUI restart the Nunchaku nodes should appear in the node menu. If they do not, the missing piece is almost always the engine wheel rather than the plugin: the README notes that from v0.3.2 you can install or update the Nunchaku wheel from inside ComfyUI using the install_wheel.json workflow, and v0.3.1 introduced the NunchakuWheelInstaller node for the same job. Load that workflow in ComfyUI and run it, then restart.
The final step is a model. The example_workflows/ directory holds ready-made graphs, including nunchaku-flux.1-dev.json, nunchaku-qwen-image.json and nunchaku-z-image-turbo.json. Open the one matching the checkpoint you downloaded and let the loader node point at the file. The README's own instruction for the Z-Image-Turbo release is to download the 4-bit model from Hugging Face or ModelScope and try it with the matching workflow in example_workflows/.
Where ComfyUI-nunchaku breaks or is the wrong tool
The most common failure is "import failed", which is what the related searches show people typing. The cause is structural: the plugin is a thin layer over a compiled engine, and ComfyUI's Python environment has to match the wheel's. A wheel built for a different torch or CUDA combination will not import, and no amount of reinstalling the plugin fixes that. The NunchakuWheelInstaller node exists to reduce the guesswork, not to remove the constraint.
The second limitation is coverage. Support arrives model family by model family, and the news entries read as a sequence of additions: FLUX.1-Kontext-dev in v0.3.3, Qwen-Image in v1.0.0, Qwen-Image-Edit and its Lightning variants in September 2025, Z-Image-Turbo in v1.1.0. If your model is not on that list, the plugin has nothing to load. The README's own note on Qwen-Image in v1.0.0 says LoRA support is coming soon, which is a reminder that even within a supported family, features land incrementally.
The third is that this is the wrong tool when you want to quantize something yourself, quickly, inside ComfyUI. Nothing in the repository suggests an in-app quantization path. And if your existing workflow depends on a custom node that expects full-precision weights, swapping in a 4-bit checkpoint changes the tensors that node receives.
How it compares with the usual low-VRAM alternatives
The obvious alternative is the GGUF quantization family of ComfyUI nodes. Both approaches give you a smaller checkpoint, but the difference in method matters. GGUF-style quantization generally compresses weights to a lower-bit format and relies on dequantization kernels at runtime. Nunchaku's SVDQuant, as described in the linked paper and the project's own framing, is a quantization method that also handles the outlier problem, and the engine is written specifically to execute those 4-bit networks rather than to dequantize into a standard kernel.
The second alternative is the model's own low-precision format, notably FP8, which ComfyUI already supports for several architectures. The README makes a direct comparison point here: the upgraded 4-bit T5 encoder released with v0.3.0 is described as matching FP8 T5 in quality. That is the project's claim about its text encoder, not a general statement about all 4-bit weights, but it tells you where the project believes the quality line sits.
The third alternative is simply offloading, which ComfyUI does natively and which costs nothing to install. If offloading already gets you acceptable speed, the extra wheel and the pre-quantized checkpoint download are overhead you do not need.
Maintenance, licensing and what an upgrade actually costs
The repository is not archived and the last push was on 2026-09-06. Releases have been frequent: v1.1.0 on 2025-12-27, v1.2.0 on 2026-01-12, v1.2.1 on 2026-01-26. That cadence is the maintenance story, and it also sets the cost. Because the plugin, the engine wheel and the quantized checkpoints are versioned separately, an upgrade is not a single action. You update the plugin, then confirm the wheel still imports, then confirm your downloaded checkpoints are still the ones the current nodes expect.
On licensing, the repository ships a LICENCE.txt and pyproject.toml declares the project under that file, with the repository metadata listing Apache-2.0. That covers this plugin. It does not automatically cover the quantized checkpoints, which are hosted separately on Hugging Face and ModelScope and may carry their own terms tied to the upstream model. Check the licence on each checkpoint you download. Nothing here is legal advice.
The dependency list is worth reading before you commit. requirements.txt pulls in diffusers, transformers, peft, accelerate, insightface, opencv-python, facexlib, onnxruntime and timm. insightface, opencv-python and onnxruntime are there for the face-related nodes rather than for the diffusion path, and they are the kind of dependencies that can conflict with an existing ComfyUI environment.
Editorial conclusion
Adopt ComfyUI-nunchaku if your ComfyUI graphs are FLUX, Qwen-Image or Z-Image based and VRAM, not model quality, is what stops you. Skip it if you are not prepared to keep a separately built Nunchaku wheel in step with a plugin that has already moved through v1.0.0, v1.1.0, v1.2.0 and v1.2.1. Before committing, check your Python version against the >=3.10,<3.14 range in pyproject.toml, confirm a pre-quantized checkpoint exists for the exact model you intend to run, and open example_workflows/ to see whether a matching graph is already there.
Frequently asked questions
How to install ComfyUI-nunchaku?
Install the dependencies from requirements.txt inside your ComfyUI environment and restart ComfyUI. The Nunchaku engine wheel is installed separately, and from v0.3.2 you can do that from inside ComfyUI with the install_wheel.json workflow or the NunchakuWheelInstaller node.
What is ComfyUI-nunchaku?
It is the ComfyUI plugin for Nunchaku, an inference engine for 4-bit neural networks quantized with SVDQuant. It adds the ComfyUI nodes needed to load and run those quantized checkpoints.
Which models does ComfyUI-nunchaku support?
The news entries document support added over time for FLUX.1-Kontext-dev, Qwen-Image, Qwen-Image-Edit and its Lightning variants, and Tongyi-MAI/Z-Image-Turbo. Each requires a pre-quantized checkpoint published by the project on Hugging Face or ModelScope.
Why does ComfyUI-nunchaku fail to import?
The plugin depends on a separately built Nunchaku wheel that must match your Python, PyTorch and CUDA versions. A mismatch there prevents the nodes from loading, which is why the project ships a NunchakuWheelInstaller node and an install_wheel.json workflow.
Can ComfyUI-nunchaku quantize a model for me?
The README does not describe an in-app quantization path. Quantization is handled by DeepCompressor, and this plugin consumes checkpoints that have already been quantized and published.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nunchux-ai-comfyui-nunchaku)