Model or dataset
openvinotoolkit/nncf avatar
openvinotoolkit/nncf

NNCF: Post-Training and Training-Time Compression for OpenVINO, PyTorch and ONNX Models

Neural Network Compression Framework for enhanced OpenVINO™ inference

1,204 stars305 forksPythonApache-2.0

At a glance

What is it?
The Neural Network Compression Framework is Intel's Python package for quantizing, pruning and sparsifying models before they run under OpenVINO. It is most useful when you already have a model and a small calibration set, and least useful when you need a compression scheme the framework has not implemented for your backend.
Who is it for?
Adopt NNCF when your target runtime is OpenVINO and you want 8-bit quantization driven by a calibration set of roughly 300 samples, or when you need quantization-aware training and pruning on PyTorch. Skip it if you need activation sparsity on ONNX or OpenVINO, since the feature table marks that combination unsupported, or if you are working outside the PyTorch, TorchFX, ONNX and OpenVINO model families.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What NNCF compresses, and for whom

NNCF is a compression layer that sits between a trained model and an inference runtime. The README describes it as a suite of post-training and training-time algorithms for optimizing inference of neural networks in OpenVINO with a minimal accuracy drop. It accepts models from PyTorch, TorchFX, ONNX and OpenVINO, which matters because it means you do not have to retrain from scratch in a new framework to shrink a model.

The audience is fairly narrow. You are a Python engineer who already has a working model, a validation split, and a deployment target that cares about latency, memory or model size. If you are still choosing an architecture, NNCF has nothing to offer you. If you are serving a model that is already small enough, the calibration step is pure overhead.

The package is versioned as v3.3.0, released on 2026-08-05, with the previous two releases landing on 2026-06-01 and 2026-04-08. The last push to the develop branch was on 2026-09-09, so the project is being worked on. The repository is not archived.

The compression surface: what is supported on which backend

The feature table in the README is the most useful page in the repository, because it separates what is production-supported from what is experimental. Post-training quantization is supported on OpenVINO, PyTorch and ONNX, and marked experimental on TorchFX. Weights compression follows the same pattern. Activation sparsity is the odd one out: experimental on PyTorch, and explicitly not supported on OpenVINO, TorchFX or ONNX.

Training-time algorithms are PyTorch-only. Quantization aware training, weight-only quantization aware training with LoRA and NLS, and pruning are all listed under a single PyTorch column. There is no ONNX or OpenVINO column for that table, which is a clear statement that if you need to fine-tune through the compression, you are doing it in PyTorch.

The README also notes that OpenVINO is the preferred backend for post-training quantization. That phrasing is worth reading literally: PyTorch and ONNX work, but the OpenVINO path is the one the project points you at first.

Beyond the algorithm tables, the framework offers automatic, configurable model graph transformation, a common interface across compression methods, GPU-accelerated layers for fine-tuning compressed models, and distributed training support. There is a git patch for huggingface-transformers that demonstrates integrating NNCF into a custom training pipeline, and an export path from PyTorch compressed models to ONNX checkpoints, plus compressed models to SavedModel or Frozen Graph format for use with the OpenVINO toolkit.

Installing NNCF and running a first quantization

The README's installation section points at the project's documentation rather than listing pip commands inline, so the canonical install instruction is the docs site at docs.openvino.ai/nncf. The package itself is published as nncf on PyPI, the build backend is setuptools with a minimum of setuptools>=77.0, and the required Python version is 3.10 or newer. Linux, Windows and MacOS are listed as supported operating systems.

The dependency list is worth reading before you install, because it is heavier than a typical quantization helper. It includes networkx, ninja, numpy, openvino-telemetry, packaging, psutil, pydot, rich, safetensors, scikit-learn, scipy and tabulate. Plotting support is a separate optional extra called plots, which pulls in kaleido, matplotlib, pandas, pillow and plotly-express. If you want the compression-ratio and accuracy plots, install that extra explicitly.

The README gives a complete OpenVINO example. It assumes an already-compressed or uncompressed model on disk, a folder-style image dataset, and a transform function that pulls the tensor out of each batch item:

python
import nncf
import openvino as ov
import torch
from torchvision import datasets, transforms

model = ov.Core().read_model("/model_path")
val_dataset = datasets.ImageFolder("/path", transform=transforms.Compose([transforms.ToTensor()]))
dataset_loader = torch.utils.data.DataLoader(val_dataset, batch_size=1)

def transform_fn(data_item):
    images, _ = data_item
    return images

calibration_dataset = nncf.Dataset(dataset_loader, transform_fn)
quantized_model = nncf.quantize(model, calibration_dataset)

The same three-step shape applies to PyTorch, where the model comes from torchvision instead of the OpenVINO core. You instantiate the model, wrap the validation split into an nncf.Dataset with a transform function, and call nncf.quantize. What you should see is a quantized model object of the same family as the input, ready for export or further evaluation.

The README states that PTQ needs only your model and a small calibration dataset of roughly 300 samples. That is the selling point, and it is also the constraint: the calibration set has to be representative, and NNCF does not tell you whether yours is.

Where NNCF stops being the right tool

The clearest limitation is the support matrix itself. If your deployment pipeline is ONNX-based and you need activation sparsity, the README says that combination is not supported. There is no workaround documented in the README, so you would be looking at a different tool or a different backend.

TorchFX is the second soft spot. Post-training quantization and weights compression are both marked experimental there, which is a different label from supported, and the README does not spell out what experimental means in terms of API stability or accuracy guarantees. Anyone building a production pipeline on the TorchFX path is doing so without a stability promise in the README.

The third constraint is the accuracy trade-off itself. NNCF is built around the idea of a minimal accuracy drop, and the README acknowledges the failure mode directly: if post-training quantization does not meet quality requirements, the suggested path is to fine-tune the quantized PyTorch model with quantization-aware training. That is a real escalation. You go from a small calibration run to a full training loop, and the training-time algorithms are PyTorch-only, so an ONNX or OpenVINO model that quantizes badly has no fine-tuning escape hatch inside this framework.

A fourth point is scale. The framework is a Python package with a dependency tree that includes scikit-learn, scipy and pydot. It is not a lightweight utility you drop into a constrained build image without thinking about the install footprint.

NNCF versus running the model uncompressed

The honest alternative is not another compression library. It is skipping compression and serving the original model.

The difference in approach is straightforward. NNCF transforms the model graph itself, inserting quantization operations and rewriting weights so that the runtime executes a smaller graph. The calibration dataset is used to collect activation statistics that determine the quantization ranges. Serving uncompressed means the runtime executes exactly the graph you trained, at full precision, with no calibration step and no graph rewriting.

That alternative wins in specific cases. If your model already fits the latency and memory budget, the calibration run, the accuracy evaluation and the extra dependency surface are cost without benefit. If your validation data is not representative of production inputs, calibration statistics can be misleading, and an uncompressed model will not have that failure mode. If your team cannot evaluate accuracy regressions reliably, introducing a lossy transform is a risk you cannot measure.

NNCF wins when the model is the bottleneck. The README frames the whole project around inference optimization, and the supported backends are exactly the ones where a compressed graph pays off. The trade is explicit: you spend a calibration dataset and an accuracy evaluation to get a smaller model.

Maintenance, releases and licence

The release cadence visible in the repository is roughly every two months: v3.1.0 on 2026-04-08, v3.2.0 on 2026-06-01, v3.3.0 on 2026-08-05. The develop branch received a push on 2026-09-09. That is a project with a regular release rhythm, and the version is declared dynamically from custom_version.version rather than hardcoded in pyproject.toml.

Upgrade cost is dominated by the dependency constraints, not by NNCF's own API. The pins are tight in places: networkx is bounded at <=3.6.1, numpy at <2.5.0, ninja at <1.14, pydot at <=4.0.1, and pandas, when you install the plots extra, at <2.4. A major bump in any of those can force a coordinated upgrade across your environment. The Python floor is 3.10, which is a real constraint if you maintain older environments.

The licence is Apache-2.0, declared both in the LICENSE file and in pyproject.toml under the SPDX identifier Apache-2.0. Apache-2.0 is a permissive licence with an explicit patent grant and requires that you preserve notices and state changes. It does not impose copyleft obligations on your own code. That is a description of the licence text, not legal advice; if you are redistributing a modified NNCF inside a product, have your own counsel read the NOTICE and patent clauses.

Editorial conclusion

Adopt NNCF when your target runtime is OpenVINO and you want 8-bit quantization driven by a calibration set of roughly 300 samples, or when you need quantization-aware training and pruning on PyTorch. Skip it if you need activation sparsity on ONNX or OpenVINO, since the feature table marks that combination unsupported, or if you are working outside the PyTorch, TorchFX, ONNX and OpenVINO model families. Before committing, check that your model's operators are handled by the backend you plan to use, and confirm which optional extras your pipeline needs, because the plots extra is not installed by default.

Frequently asked questions

What is OpenVINO used for, and how does NNCF relate to it?

The README describes NNCF as a framework for optimizing inference of neural networks in OpenVINO with a minimal accuracy drop, and OpenVINO is the preferred backend for running post-training quantization. NNCF is the compression step; OpenVINO is the runtime the compressed model targets.

Is OpenVINO only for Intel hardware?

The repository does not answer this. The README lists Linux, Windows and MacOS as supported operating systems and names OpenVINO as a backend, but it makes no statement about which hardware OpenVINO runs on.

How do I install NNCF?

The README's installation section points to the documentation at docs.openvino.ai/nncf rather than listing commands inline. The package is named nncf, requires Python 3.10 or newer, and is built with setuptools>=77.0.

How much calibration data does NNCF post-training quantization need?

The README states that PTQ needs only your model and a small calibration dataset of roughly 300 samples. The calibration set is wrapped in an nncf.Dataset with a transform function before being passed to nncf.quantize.

What should I do if NNCF quantization hurts model accuracy?

The README notes that if post-training quantization does not meet quality requirements, you can fine-tune the quantized PyTorch model, and points to a quantization-aware training example for a PyTorch ResNet18. Training-time algorithms are listed as PyTorch-only.

Which compression algorithms does NNCF support on ONNX?

The feature table lists post-training quantization and weights compression as supported on ONNX. Activation sparsity is listed as not supported on ONNX, and the training-time algorithms appear only under the PyTorch column.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. openvinotoolkit/nncf on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/openvinotoolkit-nncf.svg)](https://hysenlabs.com/projects/openvinotoolkit-nncf)