# Intel Neural Compressor: low-bit quantization for PyTorch, TensorFlow and JAX models

> Intel Neural Compressor packages post-training quantization, quantization-aware training and weight-only LLM quantization behind a prepare/convert API, with the deepest testing on Intel hardware. Here is what the repository documents, where it stops, and what to check before you commit a model to it.

**intel/neural-compressor** — SOTA low-bit LLM quantization (INT8/FP8/MXFP8/INT4/MXFP4/NVFP4) & sparsity; leading model compression techniques on PyTorch, TensorFlow, and ONNX Runtime

- Repository: https://github.com/intel/neural-compressor
- Website: https://intel.github.io/neural-compressor/
- Stars: 2,709 · Forks: 324
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/intel-neural-compressor

## What Intel Neural Compressor compresses, and for whom

The problem is the gap between a trained checkpoint and a checkpoint that fits the memory and latency budget of the machine you actually deploy on. Intel Neural Compressor is a Python library that closes that gap with a menu of compression techniques: static quantization, dynamic quantization, SmoothQuant, weight-only quantization, quantization-aware training and mixed precision. The README also lists auto-tuning, pruning, knowledge distillation and sparsity among the topics the repository covers.

The intended user is an engineer who already has a working model in PyTorch, TensorFlow or JAX and wants lower precision without rewriting the model. The library targets large language models and vision-language models specifically, naming LLaMA, Qwen, DeepSeek, Flux and FramePack in the README, and it handles low-precision data types from INT8 down to INT4, FP8, MXFP8, MXFP4 and NVFP4.

The hardware assumption is the part worth reading twice. The README says Intel Gaudi AI accelerators, Core Ultra processors, Xeon Scalable processors, Xeon CPU Max Series and Data Center GPU Flex and Max Series get extensive testing. AMD CPU, ARM CPU and NVidia GPU support is described as limited testing. That sentence should shape your evaluation more than any feature list.

## The prepare/convert mechanism behind the quantization API

The API is deliberately small. For PyTorch you import a config object and two functions, prepare and convert, from neural_compressor.torch.quantization. You build a config, hand the model and config to prepare, run calibration data through the prepared model, then call convert to get the final quantized model. The README's FP8 example does exactly this with an FP8Config(fp8_config="E4M3") and a single dummy forward pass on a ResNet-18.

That two-stage split is the whole design idea. prepare inserts observers or conversion hooks; the calibration pass collects the statistics those observers need; convert freezes them into the quantized graph. It means the quality of your result depends on the data you feed during the middle step, and the README's example explicitly labels its calibration input as dummy. A dummy tensor is fine for showing that the code path runs. It is not a calibration set.

For LLMs there is a second entry point. The load function takes a model name or path, a format such as "huggingface", a device such as "hpu", and a torch_dtype. The README notes that on first load the library converts an auto-gptq checkpoint to HPU format and writes hpu_model.safetensors into the local cache directory for subsequent loads, so the first load is slow and later loads are not. The documentation points to AutoRound as the integration behind the advanced low-bit LLM and VLM paths, and the docs directory contains separate pages for FP8, MXFP8/MXFP4, NVFP4 and AutoRound workflows.

## Installing Intel Neural Compressor and running a first quantization

Installation is split by backend so you do not pull TensorFlow into a PyTorch environment. The README gives three pip targets: neural-compressor-pt for PyTorch, neural-compressor-tf for TensorFlow and neural-compressor-jax for JAX. Each is described as framework extension API plus that framework's dependency.

```bash
pip install neural-compressor-pt
```

After that, the framework itself is your responsibility. For PyTorch on CPU or Intel GPU the README points you at the matching intel_extension_for_pytorch build; for other platforms it points at the standard PyTorch install page. On Gaudi, the README recommends starting from the Habana container image rather than installing on bare metal.

```bash
docker run -it --runtime=habana -e HABANA_VISIBLE_DEVICES=all -e OMPI_MCA_btl_vader_single_copy_mechanism=none --cap-add=sys_nice --net=host --ipc=host vault.habana.ai/gaudi-docker/1.24.0/ubuntu24.04/habanalabs/pytorch-installer-2.10.0:latest
```

Inside that environment, the first real use is an FP8 quantization of a small vision model. Note the device string: the example moves tensors to "hpu", not "cuda".

```python
from neural_compressor.torch.quantization import FP8Config, prepare, convert
import torch
import torchvision.models as models

model = models.resnet18()
qconfig = FP8Config(fp8_config="E4M3")
model = prepare(model, qconfig)
model(torch.randn(1, 3, 224, 224).to("hpu"))
model = convert(model)
output = model(torch.randn(1, 3, 224, 224).to("hpu")).to("cpu")
print(output.shape)
```

The expected result is a printed tensor shape equal to the batch and class dimensions of the input, confirming the converted model still runs. The README warns separately that from Habana software 1.21.0 onward PT_HPU_LAZY_MODE=0 is the default, and that most low-precision functions such as convert_from_uint4 do not support that setting. It recommends PT_HPU_LAZY_MODE=1 for compatibility.

```bash
export PT_HPU_LAZY_MODE=1
```

If you skip that export, expect failures in low-precision paths rather than a graceful fallback.

## Where Intel Neural Compressor is the wrong tool

The hardware asymmetry is the first limitation, and it is stated by the project rather than inferred. If your production target is an NVIDIA GPU, you are on the limited-testing side of the README's own split, and nothing in the repository description promises that the FP8 or MXFP4 paths behave the same there as on Gaudi. Choosing this library for a CUDA-only fleet means betting on a code path the documentation does not claim to validate.

The second limitation is version coupling. On Gaudi, the README says there is a version mapping between Intel Neural Compressor and the Gaudi software stack and directs you to a table to keep them matched. That is a real operational constraint: upgrading one side without the other is not a supported combination, and it turns a routine dependency bump into a coordinated change.

The third is the experimental label. FP8 for Keras/JAX, FP8 KV cache and attention static quantization, NVFP4 and MXFP8/MXFP4 are all marked experimental in the What's New list. Experimental in this repository means the API and the numerics may still move between releases, so pinning a version is the safe posture for anything in that group.

Finally, calibration quality is on you. The library provides the mechanism to collect activation statistics, not the data to collect them from. A dummy tensor, as in the README example, proves the pipeline runs and tells you nothing about accuracy.

## How it differs from LLM Compressor and framework-native quantization

LLM Compressor appears in the related searches alongside this project, and the comparison is instructive even though the two overlap. LLM Compressor is built around compressed-tensors and the vLLM serving stack, so its natural output is a checkpoint that a vLLM server loads. Intel Neural Compressor is built around a prepare/calibrate/convert API that produces a quantized model you keep using in your own Python process, with format conversion handled at load time, as the auto-gptq to HPU conversion shows.

The practical difference is where the quantization happens relative to serving. If your deployment is a vLLM endpoint, a serving-oriented tool fits the shape of the problem better. If your deployment is a Python service or a batch job that calls model(...) directly, the prepare/convert shape is the shorter path, and the framework-native quantization in PyTorch or TensorFlow is shorter still if you only need dynamic quantization and do not care about Intel-specific kernels.

What Intel Neural Compressor adds over framework-native quantization is the breadth of the low-precision menu and the Intel hardware kernels behind it. What it costs is a dependency on Intel's extension packages and, on Gaudi, on a matched software stack.

## Maintenance, releases and the Apache-2.0 licence

The repository is not archived. The last push was on 2026-09-10, and the most recent release is v3.9 from 2026-07-02, preceded by v3.8 on 2026-06-02 and v3.7 on 2025-12-25. The gap between v3.7 and v3.8 is roughly five months, while v3.8 to v3.9 is one month, so the cadence is uneven rather than metronomic. Plan upgrades around releases, not around a monthly rhythm.

The upgrade cost is concentrated in the version mapping. Because the library tracks intel-extension-for-pytorch and the Gaudi software stack, a version bump can require moving the framework underneath it at the same time. The pyproject.toml shows Black, isort, Ruff and codespell configured with a 120-character line length and a Python 3.10 target, so if you vendor or patch the source you inherit those formatting expectations.

The licence is Apache-2.0, which permits commercial use and modification and requires that you preserve notices and state changes. The repository also ships a third-party-programs.txt file, which is the place to look for bundled dependencies with different terms. That is a pointer, not legal advice; your counsel should read the file before you redistribute anything.

## Conclusion

Adopt it if you are quantizing PyTorch or JAX models for Intel Xeon, Core Ultra, Gaudi or Data Center GPU hardware, and you want one API for static, dynamic, weight-only and FP8 paths. Do not adopt it if your deployment target is an NVIDIA GPU and you need a guarantee of performance parity, because the README states that NVidia GPU support is tested only in a limited way. Before writing it into a pipeline, verify three things yourself: that the pinned intel-extension-for-pytorch or Gaudi software version matches the version mapping table, that PT_HPU_LAZY_MODE=1 is set on Gaudi, and that your accuracy tolerance survives the calibration sample you choose, since the README's own example uses a dummy calibration input.

## FAQ

### How can I compress a neural network with Intel Neural Compressor?

Install the backend package such as neural-compressor-pt, build a quantization config, call prepare on the model, run calibration data through it, then call convert to get the quantized model. The README's FP8 example uses FP8Config, prepare and convert on a ResNet-18.

### What is compression in AI, in the context of Intel Neural Compressor?

The library treats compression as a set of model compression techniques applied to a trained checkpoint: static and dynamic quantization, SmoothQuant, weight-only quantization, quantization-aware training and mixed precision. The goal is a model that fits the memory and latency budget of the target hardware.

### What is a compressor plugin for Intel Neural Compressor?

The README does not describe a plugin mechanism. What it does describe is per-framework install targets (neural-compressor-pt, neural-compressor-tf, neural-compressor-jax) and an integration with AutoRound for advanced low-bit LLM quantization.

## Sources

- [intel/neural-compressor on GitHub](https://github.com/intel/neural-compressor)
- [License: Apache-2.0](https://github.com/intel/neural-compressor/blob/main/LICENSE)
- [Project website](https://intel.github.io/neural-compressor/)
- [README](https://github.com/intel/neural-compressor/blob/main/README.md)
- [Releases](https://github.com/intel/neural-compressor/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/intel-neural-compressor
