# Optimum Quanto: eager-mode PyTorch quantization that skips the compiler

> A quantization backend for Hugging Face Optimum that trades speed for reach, working in eager mode on any device including MPS and XPU. The README now labels the project maintenance-only and points new work at bitsandbytes or torchao.

**huggingface/optimum-quanto** — A pytorch quantization backend for optimum

- Repository: https://github.com/huggingface/optimum-quanto
- Stars: 1,053 · Forks: 92
- Language: Python
- License: Apache-2.0
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/huggingface-optimum-quanto

## Maintenance mode is stated before anything else

The very first thing in the README is a blockquote, and it is worth reading before any feature list: this project is currently in maintenance mode. Pull requests are accepted only for minor bug fixes, documentation improvements and other maintenance tasks. Major new features or breaking changes are unlikely to be merged. For production-ready quantization features or active development, the README names bitsandbytes and torchAO as alternatives.

That framing changes how everything below should be read. The last push was on 2026-08-25, so the repository is not archived and is not stale, but recent activity is maintenance activity rather than feature work. The features described are complete and working. They are not going to grow.

The library is a PyTorch quantization backend for Optimum, Hugging Face's optimization hub, authored by David Corvoysier and maintained by the Hugging Face Special Ops team. It is licensed Apache-2.0, requires Python 3.9 or newer and, as of v0.2.7, PyTorch 2.6 or newer.

## Eager mode is the decision that defines the library

Optimum Quanto is built around a choice most quantization libraries make differently. Every feature is available in eager mode, which the README explains works with non-traceable models. Nothing requires a graph rewrite, a custom operator registration pass, or a compiler pass.

The consequences are concrete. Quantized models can be placed on any device, including CUDA and MPS, and the release history shows the work to widen that reach. HIP support and a Marlin int4 kernel landed in v0.2.6, while v0.2.7 enabled XPU testing and restricted CUDA extension compilation to Linux so that installation stops failing on macOS and Windows.

The price is speed. There is no graph-level fusion, no kernel autotuning, and no kernel coverage for every mixed-precision combination on every device. The README is direct about this under features yet to be implemented: dynamic activations smoothing, kernels for all mixed matrix multiplications on all devices, and compatibility with the torch compiler, which it calls dynamo.

Accelerated paths do exist where they were written. The README lists accelerated matrix multiplications on CUDA devices for int8-int8, fp16-int4, bf16-int8 and bf16-int4, plus int2, int4, int8 and float8 weights with int8 and float8 activations.

## Installing and quantizing a language model in four lines

Installation is a single pip command.

```sh
pip install optimum-quanto
```

For a Hugging Face causal language model the high-level path is short. You load the model, ask a quantized wrapper class to quantize it with a weight config, and optionally exclude a module:

```python
from transformers import AutoModelForCausalLM
from optimum.quanto import QuantizedModelForCausalLM, qint4

model = AutoModelForCausalLM.from_pretrained('meta-llama/Meta-Llama-3-8B')
qmodel = QuantizedModelForCausalLM.quantize(model, weights=qint4, exclude='lm_head')
```

Saving and reloading both go through the familiar Hugging Face methods, `save_pretrained` and `from_pretrained`, which is what makes the result usable as an ordinary model directory rather than a special artifact. Serialization is compatible with PyTorch's `weight_only` loading and with safetensors.

One detail in that snippet carries real consequences. The quantized weights come out frozen. If you intend to train them, the README says to call `optimum.quanto.quantize` directly instead of going through the wrapper class.

## Quantize, calibrate, tune, freeze as an ordered sequence

The low-level API is where the design shows. It is presented as five ordered steps, and the default behavior is the trap: weights are quantized dynamically by default, so an explicit freeze call is required.

Step one converts a float model into a dynamically quantized one. Passing both a weight and an activation config changes inference to quantize weights dynamically.

```python
from optimum.quanto import quantize, qint8

quantize(model, weights=qint8, activations=qint8)
```

Step two is calibration, optional when activations are not quantized. A `Calibration` context manager records activation ranges while representative samples pass through the model, and it activates activation quantization in the quantized modules on exit.

```python
from optimum.quanto import Calibration

with Calibration(momentum=0.9):
    model(samples)
```

Step three is optional quantization-aware tuning. If quality degrades too far, run a few epochs and recover float-level performance, since the forward pass returns a dequantized output. Step four freezes, replacing float weights with integer weights. Step five serializes, and the README notes you must store the quantization map alongside the weights to reload them.

The automatic work happens underneath all of it. The README lists automatic insertion of quantization and dequantization stubs, quantized functional operations and quantized modules, a path from float to dynamic to static quantization, and a standardized model specification format for models that live outside the Hugging Face Hub.

## Diffusers submodels, not only whole pipelines

The library reaches past text models. Quantized classes exist for diffusers submodules, and the README example quantizes the `transformer` inside a PixArt pipeline rather than the pipeline as a whole, then saves it and swaps it back into a freshly loaded pipeline at inference time.

```python
from diffusers import PixArtTransformer2DModel
from optimum.quanto import QuantizedPixArtTransformer2DModel, qfloat8

model = PixArtTransformer2DModel.from_pretrained("PixArt-alpha/PixArt-Sigma-XL-2-1024-MS", subfolder="transformer")
qmodel = QuantizedPixArtTransformer2DModel.quantize(model, weights=qfloat8)
qmodel.save_pretrained("./pixart-sigma-fp8")
```

That pattern generalizes to any pipeline with a heavy internal model, which covers most interesting cases in text-to-image and video generation. The tree has an `examples/` directory, the pyproject examples extra pulls torchvision, transformers, diffusers, datasets, accelerate, sentencepiece and scipy, and v0.2.5 added both a Whisper speech recognition example and a ViT classification example, so audio and vision models are covered too.

Measurements live in a separate `bench/` directory holding detailed per-use-case results, and the README sends you there rather than asking you to trust a summary.

## What the performance claims actually rest on

The README summarizes performance in three claims. Accuracy for models compiled with int8 or float8 weights and float8 activations is described as very close to full precision. Latency for quantized inference is comparable to full precision when only weights are quantized and an optimized kernel exists. Device memory is approximately divided by the ratio of float bits to integer bits.

Two of those three are conditional and the conditions matter. Latency parity requires an optimized kernel for the specific combination, and memory scaling assumes you are not fighting fragmentation or activation memory. The accuracy claim is scoped to two specific configurations rather than to int2 or int4.

The evidence offered is a worked example for meta-llama/Meta-Llama-3.1-8B, presented as WikiText perplexity and latency charts, with the README stating that the paragraph is only an example and directing readers to the bench directory for detailed results per use case. The project publishes its own numbers and cites no outside measurements, so treat the charts as a starting point rather than a settled comparison.

Repository hygiene is ordinary. The Makefile exposes `check`, `style` and `test` targets running ruff and pytest over the optimum, tests, bench and examples directories. pyproject.toml uses ruff with a 119-character line length, ignores E501, and pins the build to setuptools with setuptools_scm for versioning. v0.2.6 also records the switch from black to ruff.

## Conclusion

Optimum Quanto earns a place when a model will not cooperate with torch.compile, when you need the same quantization code to run on a Mac, on XPU and on CUDA, or when you want one API covering int2 through float8 without learning a kernel library. Skip it for a production service: the README states the project is in maintenance mode, accepts pull requests only for minor fixes and documentation, and directs you to bitsandbytes or torchao for production-ready quantization or active development. Install the package, quantize one model with the low-level API, and compare perplexity and latency against the bench directory before committing, since the published figures come from the project's own benchmarks rather than independent measurement.

## FAQ

### Is Optimum Quanto still actively developed?

The README states the project is in maintenance mode: pull requests are accepted only for minor bug fixes and documentation improvements, and major features or breaking changes are unlikely to be merged. It points to bitsandbytes and torchao for production-ready quantization or active development. The last push was on 2026-08-25.

### Which quantization formats does Optimum Quanto support?

The README lists int2, int4, int8 and float8 weights, plus int8 and float8 activations. Accelerated matrix multiplications on CUDA cover int8-int8, fp16-int4, bf16-int8 and bf16-int4, and v0.2.5 added float8 e4f3mnuz support.

### Why are my quantized weights not updating during training?

They are frozen on purpose. The README notes that weights quantized through the wrapper class are frozen, and that you must call optimum.quanto.quantize directly if you want to keep them unfrozen and train them.

### Does Optimum Quanto work without torch.compile?

Yes. All features are available in eager mode, which the README says works with non-traceable models, and compatibility with the torch compiler is listed as not yet implemented. That is the main trade: reach across devices instead of graph-level speed.

### How do I reload a model quantized with Optimum Quanto?

For a Hugging Face model, save it with save_pretrained and load it back with QuantizedModelForCausalLM.from_pretrained. For the low-level API, save the state dict with safetensors and also store the quantization map, since both are needed to reload the weights.

## Sources

- [huggingface/optimum-quanto on GitHub](https://github.com/huggingface/optimum-quanto)
- [Issues](https://github.com/huggingface/optimum-quanto/issues)
- [License: Apache-2.0](https://github.com/huggingface/optimum-quanto/blob/main/LICENSE)
- [README](https://github.com/huggingface/optimum-quanto/blob/main/README.md)
- [Releases](https://github.com/huggingface/optimum-quanto/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/huggingface-optimum-quanto
