Brevitas: quantization-aware training in PyTorch, from AMD Research
Brevitas: neural network quantization in PyTorch
At a glance
- What is it?
- Brevitas is a PyTorch library for post-training quantization and quantization-aware training, maintained as a research project by AMD Research rather than an official Xilinx product. It is built for engineers who need quantized layers they can drop into an existing torchvision model.
- Who is it for?
- Adopt Brevitas if you already train in PyTorch and want quantized layers you can swap into an existing model, and you accept that it is a research project rather than a supported Xilinx product. Do not adopt it if you need a vendor-backed support contract or an ONNX export path that the README does not describe.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
What Brevitas solves, and who it is actually for
Quantization shrinks the numeric precision of a neural network's tensors, and doing that well is not a single function call. Brevitas exists to make quantization a first-class part of a PyTorch model rather than a post-hoc conversion step. The README frames it as a library for neural network quantization with support for both post-training quantization (PTQ) and quantization-aware training (QAT).
The audience is narrow and identifiable. If you have a torchvision model and you want to measure what happens when you quantize it, the project ships an example user flow under brevitas_examples.imagenet_classification.ptq that quantizes an input torchvision model under different configurations, for example bit-width and granularity of scale. That is a research-and-evaluation workflow, not a deployment pipeline.
The README is explicit about status: Brevitas is a research project and not an official Xilinx product. That sentence should shape every adoption decision. It means the API can move between releases, and it means nobody owes you a support contract. The project is published from an AMD Research address and authored by a named list of researchers, which is consistent with a research artifact that happens to be pip-installable.
QuantConv2d and friends: how the quantization is actually wired in
The mechanism is layer substitution. Brevitas offers quantized implementations of common PyTorch layers under brevitas.nn, including QuantConv1d, QuantConv2d, QuantConvTranspose1d, QuantConvTranspose2d, QuantMultiheadAttention, QuantRNN and QuantLSTM. You build or convert a model so those layers replace their torch equivalents.
The granularity is the interesting part. For each of these layers, quantization of different tensors (inputs, weights, bias, outputs and so on) can be individually tuned according to a wide range of quantization settings. So a single convolution can have one setting for its weights and a different one for its activations, rather than a single global precision for the whole graph.
That design is what separates QAT from a conversion tool. Because the quantized layers are ordinary nn.Module subclasses living inside your model, gradients flow through them during training, which is what makes quantization-aware training possible at all. The cost is that you are editing model code, not pointing a converter at a checkpoint. The repository layout reflects this: the library lives under src/, with tests/ and a noxfile.py alongside it, and the project description in setup.py is simply quantization-aware training in PyTorch.
Installing Brevitas and quantizing a first layer
The README gives one installation command for the latest release from PyPI:
pip install brevitasBefore you run it, check the stated requirements. Python 3.10 or newer is required. PyTorch must be at least 1.13 and at most 2.13, and the README notes that more recent versions would be untested. Windows, Linux and macOS are all listed as supported, and GPU training-time acceleration is described as optional but recommended.
Once installed, the first real use is importing a quantized layer from brevitas.nn and placing it in a model. The README names the available layers but does not print a worked constructor example, so the exact keyword arguments for a given quantizer are a documentation question rather than something to guess at. What you should see is a module whose class names carry the Quant prefix, so it is obvious at a glance which parts of the graph are quantized. For a fuller PTQ walkthrough, the README points at brevitas_examples.imagenet_classification.ptq, which takes a torchvision model and applies different quantization configurations to it. The README does not document an expected accuracy number for that flow, so treat any figure you generate yourself as your own result, not a project guarantee.
The PyTorch version ceiling is the constraint to plan around
The most concrete limitation is stated in the requirements: PyTorch must be at least 1.13 and at most 2.13, with more recent versions described as untested. If your team upgrades PyTorch on a schedule, Brevitas is a dependency that can hold that upgrade back until the maintainers widen the range. That is a real cost, and it is not a hypothetical one for a library that hooks into layer internals.
The second limitation is the research status. A research project can rename a quantizer, change a default, or drop a layer between releases. The release history shows a steady cadence, with 0.13.0, 0.13.1 and 0.13.2 all published within roughly two months of each other, and the last push to the repository was on 2026-09-10. Frequent releases are not the same as a stability promise, and nothing in the README offers one.
Where it is the wrong tool: if your goal is to take a finished checkpoint and emit a deployable quantized artifact without touching training code, Brevitas is aimed elsewhere. It is a library you adopt inside a PyTorch model, and its example flow is an evaluation script over torchvision models. The README does not document rollback, a conversion-only mode, or a supported export path in the text available, so do not assume one exists.
How Brevitas differs from a generic PTQ toolkit
The clearest contrast is with PyTorch's own quantization facilities, which the README implicitly positions against by describing Brevitas as supporting both PTQ and QAT. The difference is in approach rather than in feature lists. PyTorch's built-in path is oriented around converting an existing eager or FX graph, with quantization settings applied through observer configuration and a conversion step.
Brevitas inverts that. You place quantized layers into the model yourself, and each tensor within a layer can be tuned separately. That makes the quantizer part of the model definition, which is what allows quantization-aware training to run as ordinary training. The trade-off is that porting a model means editing its architecture, and the settings live in constructor arguments rather than in a central conversion configuration.
There is also a hardware angle visible in the repository topics, which include fpga and hardware-acceleration. Brevitas comes from Xilinx and AMD Research, and the quantized layer set reads like something intended to line up with fixed-point accelerator datapaths. The README does not document a hardware deployment flow, so treat that as context for why the library exists rather than as a documented feature.
Licence, maintenance and the cost of upgrading
The licence is reported as NOASSERTION at the repository level, but setup.py carries an SPDX identifier of BSD-3-Clause and a copyright line for Advanced Micro Devices. Those two signals disagree, and the discrepancy matters if your organisation has an allowlist of approved licences. Check the LICENSE file in the repository root directly rather than relying on either the metadata or this article; nothing here is legal advice.
On maintenance, the evidence is concrete. The repository is not archived, and the last push was on 2026-09-10, which is recent. Releases are frequent: 0.13.0 on 2026-07-14, 0.13.1 on 2026-08-25 and 0.13.2 on 2026-08-28. The project also carries a Zenodo DOI and asks to be cited, which is a research-project habit rather than a product one.
The upgrade cost has two parts. First, the PyTorch ceiling means a Brevitas upgrade and a PyTorch upgrade are coupled decisions. Second, because quantization settings live in layer constructors, a change to a quantizer's arguments becomes a code change across every model that uses it. Pin the version in your requirements file and read the release notes for each bump; the README links them per release.
Editorial conclusion
Adopt Brevitas if you already train in PyTorch and want quantized layers you can swap into an existing model, and you accept that it is a research project rather than a supported Xilinx product. Do not adopt it if you need a vendor-backed support contract or an ONNX export path that the README does not describe. Before committing, verify that your PyTorch version sits inside the documented range of 1.13 to 2.13, that Python 3.10 or newer is available, and that the quantized layers you need are present under brevitas.nn.
Frequently asked questions
What does Brevitas do?
Brevitas is a PyTorch library for neural network quantization, supporting both post-training quantization and quantization-aware training. It provides quantized versions of common PyTorch layers under brevitas.nn, such as QuantConv2d and QuantLinear.
What is Brevitas?
It is a research project from AMD Research and Xilinx that implements quantized layers for PyTorch, installable from PyPI. The README states plainly that it is a research project and not an official Xilinx product.
What are the benefits of using Brevitas?
Each tensor in a quantized layer, including inputs, weights, bias and outputs, can be tuned individually according to a wide range of quantization settings. It supports both PTQ and QAT, so quantization can be part of training rather than only a conversion step.
What are some examples of Brevitas usage?
The README points to an example user flow under brevitas_examples.imagenet_classification.ptq that quantizes an input torchvision model under different quantization configurations, such as bit-width and granularity of scale.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/xilinx-brevitas)
Community notes