What is Model quantization?
Model quantization is the process of storing a model's weights and sometimes its activations in lower-precision numbers, such as 8-bit or 4-bit integers instead of 16-bit floats. The aim is to cut memory use and speed up inference, usually at some cost to accuracy.
How quantization works
A neural network is a long list of numbers. Weights are the learned parameters, activations are the intermediate values produced at run time, and in a KV cache the model stores past attention keys and values. Training normally keeps these in 16-bit or 32-bit floating point. Quantization replaces that representation with a smaller one, most often 8-bit or 4-bit integers. The mapping is a scale and a zero point: a real value x becomes round(x / s) + z, and the integer is turned back into an approximate float at use time. The scale is chosen per tensor, per channel or per group of weights, and the choice of granularity decides how much error the rounding introduces. A coarse per-tensor scale is cheap but blunt; finer groups cost more metadata and more compute.
Quantization is not one technique. Post-training quantization (PTQ) takes a trained checkpoint and calibrates the scales on a small sample of data, with no gradient updates. Quantization-aware training (QAT) inserts fake quantization into the forward pass during training so the model learns to tolerate rounding. Weight-only quantization stores the weights in low precision but runs the matrix multiply in a higher precision, which saves memory without needing special integer kernels. Weight and activation quantization goes further and runs the arithmetic itself in low precision, which is where the largest speedups live but also where accuracy is hardest to hold.
There are two distinct families in open-source work. Integer quantization is the classic path, for example 8-bit or 4-bit weights with an integer scale. Microscaling formats such as MXFP8 and NVFP4 group values into blocks that share an exponent, which keeps a wider dynamic range than a single global scale. Vector quantization is different again: it replaces a continuous vector with the index of the nearest entry in a learned codebook, which is the mechanism behind VQ-VAE style models rather than a way to shrink an existing LLM.
When you need quantization, and when you do not
The main reason to quantize is memory. A model's weight footprint is roughly parameter count times bytes per parameter, so moving from 16-bit to 8-bit halves it and moving to 4-bit quarters it. That is what makes a model fit on a single accelerator or a phone. The second reason is bandwidth: decoding a token requires reading the weights, so fewer bytes per weight can mean faster generation when the workload is memory bound. The third is cost, since a smaller model fits on cheaper hardware and may need fewer devices.
Quantization is less useful when the workload is compute bound rather than memory bound, for example a large prefill with a long prompt on a fast GPU. It is also a poor fit when the model is already small enough, when accuracy on a narrow task is the only thing that matters, or when the serving stack has no low-precision kernels for your hardware. Weight-only quantization saves memory but does not speed up the matrix multiply by itself. Activation quantization can hurt accuracy more than weight quantization because activations have outliers, which is why techniques that keep a few channels in high precision exist. Quantizing the KV cache is a separate decision from quantizing the weights, and it targets long-context memory rather than the model file.
Pitfalls and limits
The first pitfall is measuring the wrong thing. Perplexity is a common proxy, but a small perplexity change can still break structured output, tool calling or long reasoning chains. The second is calibration. PTQ scales are chosen from a calibration set, and a set that does not resemble production traffic can produce scales that clip real activations. The third is granularity and kernel support: a scheme that looks good on paper may have no fast path on your accelerator, so it runs slower than the 16-bit baseline.
There are also practical traps around tooling. Quantized checkpoints are often tied to the serving framework that produced them, and moving between formats can require a full re-quantization. Some toolkits document their limitations clearly and others do not; where a project's documentation is silent, that silence is itself a risk to check before committing a model. Mixed precision is a partial answer: keeping the first and last layers, or the attention output projection, in higher precision often recovers much of the lost accuracy at a small memory cost. Finally, quantization does not fix a weak model. It trades precision for size, and if the starting checkpoint is poor, the quantized version will be worse.
How quantization shows up in open-source projects
bitsandbytes-foundation/bitsandbytes brings 8-bit optimizers, LLM.int8() inference and QLoRA 4-bit training to PyTorch, and its analysis notes support for CPU, NVIDIA, AMD, Intel and Apple hardware. It is a library for quantizing inside PyTorch rather than a conversion pipeline. pytorch/ao (TorchAO) is PyTorch native quantization for training and inference, spanning float8 and MXFP8 training, int4 inference, QAT and sparsity; it installs with pip, but its analysis says choosing the right workflow for your hardware is the real work. NVIDIA/Model-Optimizer (ModelOpt) turns a Hugging Face, PyTorch or ONNX checkpoint into a compressed one that TensorRT-LLM, TensorRT, vLLM or SGLang can serve; the analysis warns that the useful path runs through NVIDIA's stack and that the docs are uneven outside the LLM examples.
On the vendor side, qualcomm/aimet provides post-training techniques such as AdaRound and SeqMSE plus QAT for trained PyTorch and ONNX models, with clean PyPI installs and a CMake and CUDA source build. intel/neural-compressor packages PTQ, QAT and weight-only LLM quantization behind a prepare/convert API, with the deepest testing on Intel hardware. intel/auto-round is an Apache-2.0 toolkit that uses sign-gradient descent to hold accuracy at 2 to 4 bits, exports to AutoRound, AutoAWQ, AutoGPTQ and GGUF, and plugs into Transformers, vLLM and SGLang.
For deployment, cactus-compute/cactus is a C++ engine for mobiles, wearables, smart home and robots, and its analysis highlights the CQ2 quantization tier as a headline trade-off that collapses reasoning scores. ModelTC/LightX2V wraps diffusion video models behind one inference stack and, per its release notes dated February 27, 2026, added FP8 and NVFP4 quantization for autoregressive video generation models. 0xSero/turboquant implements a KV cache compression scheme as Triton kernels plus a vLLM patch, and its analysis notes that the README retracts several headline numbers and narrows the real gain to the full-attention layers. lucidrains/vector-quantize-pytorch is a different kind of project: it turns continuous feature vectors into discrete codebook indices with residual, grouped and scalar variants, and its analysis describes it as a building block for VQ-VAE style models rather than an end-to-end trainer.
Choosing a starting point
Start by deciding what you are trying to save. If the goal is fitting a larger model into the same memory, weight-only 4-bit or 8-bit PTQ is usually the first thing to try. If the goal is throughput on a specific accelerator, check whether your serving stack has low-precision kernels before you convert anything, because a format without a fast path can be slower than the baseline. If accuracy is tight, consider QAT or a mixed-precision scheme that keeps sensitive layers in higher precision. If the problem is long-context memory rather than the model file, look at KV cache quantization separately.
Read each project's own limitations section before you commit a checkpoint. Some of the projects above document where their methods stop working, and some do not. The ones that do are easier to trust with a production model. Maintenance status is also worth checking: an archived repository or one whose last push is many months old may still work, but it will not track changes in the frameworks it depends on.
In practice
Quantization is a memory and bandwidth trade, not a free speedup. Decide first whether you are constrained by model size, by KV cache size or by kernel throughput, then pick the precision and the granularity to match. Read the limitations section of whichever project you choose, and verify accuracy on your own task rather than on a perplexity number. A reasonable next step is to read the bitsandbytes, TorchAO and ModelOpt documentation side by side, since they cover three different points in the same pipeline: in-framework quantization, native PyTorch training and inference, and conversion for a serving stack.