AIMET: Quantization and Compression for PyTorch and ONNX Models
AIMET is a library that provides advanced quantization and compression techniques for trained neural network models.
At a glance
- What is it?
- AIMET is Qualcomm's toolkit for quantizing trained PyTorch and ONNX models, with post-training techniques such as AdaRound and SeqMSE plus QAT support. The PyPI packages install cleanly; the source build is a CMake and CUDA affair.
- Who is it for?
- Adopt AIMET if you have a trained PyTorch or ONNX model and a target where INT8 inference matters, and start with the PyPI packages rather than the source build. Do not adopt it expecting ONNX support across the board: OmniQuant is marked PyTorch-only in the feature table, and the repository's build path assumes CMake, CUDA architectures and a Docker environment.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What AIMET solves, and who ends up using it
A trained model that runs in float32 is often too large and too slow for the device it needs to run on. AIMET is a library for quantizing and compressing those trained models, and its stated goal is to reduce compute load and memory footprint so the model can be deployed on edge devices such as mobile phones or laptops. The README gives the arithmetic behind that motivation: 8-bit precision models have a 4x smaller footprint than 32-bit precision models, and models run 5x-15x faster on the Qualcomm Hexagon DSP than on the Qualcomm Kyro CPU.
The audience is narrower than "anyone with a model". You need a trained model in PyTorch or ONNX, and you need to care about integer inference. AIMET works on models you already have; it is not a training framework, and it does not design architectures. The README describes it as designed to automate optimization of neural networks, avoiding manual tweaking, and it exposes APIs that can be called directly from a PyTorch pipeline. If your deployment target is a server with abundant memory and float throughput is not the bottleneck, the techniques here solve a problem you do not have.
How quantization actually proceeds: calibration, then correction
The mechanism is a pipeline of techniques applied to a trained model, not a single conversion call. The README's feature table is the clearest map of it. Calibration computes quantization parameters and is available for both ONNX and PyTorch. On top of that base sit techniques that repair the accuracy lost during quantization: AdaRound rounds quantized weights, SeqMSE optimizes encodings for each layer, Cross Layer Equalization rescales weights to reduce range imbalance, BatchNorm Folding folds batchnorm to bridge the gap between simulation and on-target behaviour, and BatchNorm re-estimation recomputes batchnorm statistics. AdaScale learns per-weight scales through blockwise optimization, and SpinQuant reduces activation outliers via Hadamard rotations.
Two entries in that table are not symmetric across frameworks. OmniQuant, which optimizes quantized weights, is marked PyTorch only. Everything else in the PTQ table carries both checkmarks. That asymmetry is the first thing to check against your own model, because it determines which techniques are even reachable from your pipeline.
For Quantization Aware Training, AIMET supports QAT through aimet-torch. The README recommends a workflow that combines QAT with the advanced PTQ techniques, and illustrates it with a diagram rather than a step list. Compression is a separate track: Spatial SVD splits a large layer into two smaller ones through tensor decomposition, Channel Pruning removes redundant input channels and reconstructs layer weights, and per-layer compression-ratio selection picks how much to compress each layer automatically. Visualization covers weight ranges, so you can inspect whether a model is a candidate for Cross Layer Equalization and see the effect after applying it.
Installing AIMET from PyPI and running a first quantization pass
The README states that aimet-onnx and aimet-torch are available on PyPI, and points to the Quick Start page for the latest package. The package names are the ones to install:
pip install aimet-torch
pip install aimet-onnxPick the one that matches your model. aimet-torch is the package the README names for Quantization Aware Training; aimet-onnx covers the ONNX side of the PTQ table. The project requires Python 3.10 or newer according to pyproject.toml, and the wheel is built for the stable ABI with py-api cp310, so a single wheel is intended to work across Python 3.10 and later.
Once installed, the first real step is calibration, which the table describes as computing quantization parameters. The repository ships worked code under Examples/, split into Examples/onnx/ and Examples/torch/, and Examples/README.md is the entry point for those. Read that file before writing your own script, because the examples encode the API sequence the documentation expects.
If you need to build from source instead, the README directs you to the Docker-based build instructions rather than a bare pip install. The build system is scikit-build-core driving CMake, with a minimum CMake version of 3.26, and pyproject.toml sets CUDA architectures to 70;75;80 with CMAKE_BUILD_TYPE RelWithDebInfo. That is a heavier environment than the PyPI path and is worth avoiding unless you need the develop branch.
Where AIMET stops helping
The framework split is the most immediate limitation. If your model is ONNX and the technique you want is OmniQuant, the table says no. The README does not document a workaround, and it does not explain why OmniQuant is PyTorch-only. Plan around the checkmarks rather than assuming parity.
The second limitation is that quantization is not free accuracy. The README is explicit that maintaining model accuracy when quantizing is often challenging, which is the reason the toolkit exists at all. Calibration alone computes parameters; the techniques that recover accuracy, AdaRound and SeqMSE among them, are additional steps with their own cost. Nothing in the README estimates how long those steps take on a given model, so budget for experimentation rather than treating this as a one-shot conversion.
The third is the build path. The PyPI packages are the documented quick start, but the source route assumes Docker, CMake 3.26 or newer and CUDA architectures. If your environment cannot provide those, the source build is closed to you. The README does not document rollback or recovery if a quantization run degrades a model past usefulness, so keep the original checkpoint.
AIMET against general-purpose quantization in the training framework
PyTorch ships its own quantization APIs, and that is the real alternative for most teams already inside the PyTorch ecosystem. The difference is in what the two are trying to do. Framework-native quantization gives you the conversion mechanics and a stable API surface tied to the framework release cycle. AIMET layers research-derived techniques on top of a trained model: AdaRound for weight rounding, SeqMSE for per-layer encoding search, Cross Layer Equalization for range imbalance, SpinQuant for activation outliers, and data-free quantization, which the README names as a technique that provides state-of-the-art INT8 results on several popular models.
That extra layer is the trade. You get a larger menu of accuracy-recovery methods and a compression track (Spatial SVD, Channel Pruning) that framework-native quantization does not offer. You take on a dependency whose install path and feature matrix are tied to Qualcomm's release cadence, currently visible as 2.37.0, 2.38.0 and 2.39.0 across August and September 2026, and whose most advanced ONNX coverage is incomplete. If calibration plus basic conversion gets you the accuracy you need, the framework's own path is fewer moving parts. Reach for AIMET when calibration alone is not enough.
Maintenance cadence, licence and the cost of upgrading
The repository is not archived, and the last push was on 2026-09-09. Releases are frequent: 2.37.0 on 2026-08-12, 2.38.0 on 2026-08-24 and 2.39.0 on 2026-09-08. That cadence cuts both ways. Fixes and new techniques arrive quickly, but so does churn, and the README's links point at a releases/latest path, meaning documentation moves with the release rather than staying pinned to the version you installed.
Upgrade cost is dominated by the build, not the API. The PyPI packages are the low-friction path. The source build pulls in scikit-build-core, pybind11, Cython 3.0 or newer, CMake 3.26 and CUDA architecture flags from pyproject.toml, so a source-based workflow means re-validating that toolchain whenever you move versions.
The licence field deserves care. pyproject.toml declares BSD-3-Clause, while the repository metadata reports NOASSERTION, and the root LICENSE file is the authoritative text. The repository also carries a .pre-commit-license-header.txt and a repolint.json, which suggests header checks are enforced on contributions. Read LICENSE before you redistribute anything built with AIMET; nothing here is legal advice.
Editorial conclusion
Adopt AIMET if you have a trained PyTorch or ONNX model and a target where INT8 inference matters, and start with the PyPI packages rather than the source build. Do not adopt it expecting ONNX support across the board: OmniQuant is marked PyTorch-only in the feature table, and the repository's build path assumes CMake, CUDA architectures and a Docker environment. Before committing, verify which techniques your framework actually supports, check that the licence terms in LICENSE match your distribution plan, and confirm the accuracy you get from a calibration-only run before investing in AdaRound or QAT.
Frequently asked questions
What is aimet-torch and how does it differ from aimet-onnx?
They are the two PyPI packages the README names for working with the toolkit, split by model framework. aimet-torch is the package the README specifies for Quantization Aware Training, while aimet-onnx covers the ONNX side of the post-training table.
How do I install AIMET?
The README states that aimet-onnx and aimet-torch are available on PyPI and points to the Quick Start page for the latest package. Building from source instead follows the Docker-based build instructions, and pyproject.toml requires Python 3.10 or newer.
Which quantization techniques does AIMET support for ONNX models?
Calibration, AdaRound, SeqMSE, BatchNorm Folding, Cross Layer Equalization, BatchNorm re-estimation, AdaScale and SpinQuant are all marked as available for both ONNX and PyTorch. OmniQuant is marked PyTorch only.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/qualcomm-aimet)
Community notes