Model or dataset
lucidrains/vector-quantize-pytorch avatar
lucidrains/vector-quantize-pytorch

vector-quantize-pytorch: A Vector Quantization Library for PyTorch Models

Vector (and Scalar) Quantization, in Pytorch

4,009 stars340 forksPythonMIT

At a glance

What is it?
A PyTorch package that turns continuous feature vectors into discrete codebook indices, with residual, grouped and scalar variants. It is a building block for VQ-VAE style models, not an end-to-end trainer.
Who is it for?
Adopt vector-quantize-pytorch if you are building a VQ-VAE, RQ-VAE, or audio codec model in PyTorch and want the quantization layer, the commitment loss, and the codebook update rules already implemented. Do not adopt it if you need a trained model, a data pipeline, or a training loop: the package ships the layer and examples, not an end-to-end system.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What vector-quantize-pytorch actually solves

Neural encoders output continuous vectors. Many generative models need a discrete token instead, because a language model, an autoregressive prior, or a lookup table of embeddings can only consume symbols. Vector quantization is the bridge: each encoder output is replaced by the nearest entry in a learned codebook, and the model trains against that replacement.

The package exists because writing that layer correctly is fiddly. The codebook has to be updated, the nearest-neighbour search has to be batched, the gradient has to pass through a non-differentiable argmin, and the encoder has to be pushed toward the codebook it is being quantized against. The README describes the library as originally transcribed from Deepmind's tensorflow implementation, made conveniently into a package, and it uses exponential moving averages to update the dictionary. That is the whole scope: the quantization layer, not the model around it.

It is for people who already have an encoder and a decoder and need the middle. If you are training a VQ-VAE, an RQ-VAE, or an audio codec in PyTorch, the layer is the part you would otherwise write yourself. If you want a pretrained image or audio model, this is not that.

The VectorQuantize forward pass and what it returns

The core class is VectorQuantize. It takes the feature dimension, the codebook size, the exponential moving average decay, and the commitment weight. The forward pass accepts a tensor shaped (batch, sequence, dim) and returns three values: the quantized tensor with the same shape as the input, integer indices of shape (batch, sequence), and a scalar commitment loss.

The indices are the discrete codes. The quantized tensor is what the decoder receives. The commitment loss is added to the reconstruction loss during training; the README gives its default weight as 1.0 through the commitment_weight argument. The decay argument controls how fast the dictionary moves, and the README states that a lower value means the dictionary will change faster.

The important structural point is that the layer is stateless with respect to your model. It does not own an optimizer, a scheduler, or a training step. You call it inside your own forward pass and add the returned loss to yours. That is why the package can support so many variants without becoming a framework.

Residual, grouped and scalar variants in one package

A single codebook with 1024 entries gives you 10 bits per vector. To get finer reconstruction you can quantize the residual repeatedly, which is what ResidualVQ does. It takes num_quantizers and returns indices of shape (batch, sequence, num_quantizers), so the second stage quantizes what the first stage got wrong. Passing return_all_codes = True returns every intermediate code as well.

GroupedResidualVQ splits the feature dimension into groups and runs residual quantization inside each group. The README cites a paper reporting results equivalent to Encodec while using far fewer codebooks, and the returned index tensor gains a leading group axis. The repository also contains example files for scalar and finite scalar quantization, including examples/autoencoder_fsq.py and examples/autoencoder_lfq.py, so the scalar path is present in the tree even though the README excerpt focuses on the vector classes.

These are not interchangeable. ResidualVQ multiplies your index sequence length by num_quantizers, which changes what your prior model has to predict. GroupedResidualVQ keeps the sequence shorter per group but adds a group axis. Pick based on what your decoder and your prior can consume, not on which name sounds more advanced.

Installing vector-quantize-pytorch and running a first quantization

The README gives one install command. It pulls the package from PyPI along with its runtime dependencies, which pyproject.toml lists as einops, einx, torch, and torch-einops-utils, with torch pinned at 2.4 or newer and Python at 3.9 or newer.

bash
$ pip install vector-quantize-pytorch

After that, the smallest useful program is the README's own example. It builds a quantizer over 256-dimensional vectors with a 512-entry codebook and runs a batch of 1024 such vectors through it.

python
import torch
from vector_quantize_pytorch import VectorQuantize

vq = VectorQuantize(
    dim = 256,
    codebook_size = 512,     # codebook size
    decay = 0.8,             # the exponential moving average decay, lower means the dictionary will change faster
    commitment_weight = 1.   # the weight on the commitment loss
)

x = torch.randn(1, 1024, 256)
quantized, indices, commit_loss = vq(x) # (1, 1024, 256), (1, 1024), (1)

The shapes in the trailing comment are the thing to check first. If your encoder emits (batch, channels, length) rather than (batch, length, channels), you need a permute before the call and the inverse after it. If the indices come back as (1, 1024) and your decoder expects a different layout, fix that at the boundary rather than reshaping the codebook.

For a first real use, start with kmeans_init = True. The README attributes the idea to the SoundStream paper, which proposes initializing the codebook with kmeans centroids of the first batch. The flag is available on both VectorQuantize and ResidualVQ, and kmeans_iters controls the number of kmeans iterations used to compute those centroids.

python
from vector_quantize_pytorch import ResidualVQ

residual_vq = ResidualVQ(
    dim = 256,
    codebook_size = 256,
    num_quantizers = 4,
    kmeans_init = True,   # set to True
    kmeans_iters = 10     # number of kmeans iterations to calculate the centroids for the codebook on init
)

x = torch.randn(1, 1024, 256)
quantized, indices, commit_loss = residual_vq(x)

Note that kmeans_init consumes the first batch to compute centroids, so the first forward pass is not representative of steady-state behaviour. Warm up on real data rather than on random noise if you care about the initialization.

Gradient estimators: straight-through, rotation trick and DiVeQ

The argmin in vector quantization has no useful gradient, so the layer has to choose how to pass one backward. The traditional answer is the straight-through estimator: the gradient flows around the VQ layer rather than through it. The README describes this as the default behaviour and gives rotation_trick = True as the alternative.

The rotation trick, from a 2024 paper cited in the README, transforms the gradient through the VQ layer so the relative angle and magnitude between the input vector and the quantized output are encoded into the gradient. The third option is directional_reparam, which the README attributes to the DiVeQ paper. It models quantization as adding a simulated quantization error to the input, with the direction aligned to the nearest codeword and the magnitude equal to the actual quantization error. That makes the quantized output differentiable with respect to both the input and the selected codeword, so the codebook can be learned by gradients without auxiliary losses.

python
from vector_quantize_pytorch import VectorQuantize

vq_layer = VectorQuantize(
    dim = 256,
    codebook_size = 256,
    rotation_trick = True,   # Set to False to use the STE gradient estimator or True to use the rotation trick.
)

The directional variant adds a variance parameter, given as directional_reparam_variance = 5e-3 in the README example, and the README notes that a small value aligns the simulated error with the nearest codeword. This is the least settled part of the library: three estimators with different papers behind them, no benchmark in the README comparing them, and no guidance on which to pick for a given architecture. Treat the choice as a hyperparameter you have to test, not a switch with a documented best setting.

Dead codebook entries and the techniques aimed at them

The README is direct about the main failure mode: dead codebook entries, which it calls a common problem when using vector quantizers. A codebook entry that never wins the nearest-neighbour search receives no updates and stays where it is, which wastes capacity and shrinks the effective codebook. Nothing in the layer prevents this, and the README excerpt does not show an exported metric for codebook utilization, so you have to count unused indices yourself during training.

The repository collects several published answers. Lower codebook dimension, from the Improved VQGAN paper, keeps the codebook in a lower dimension: encoder values are projected down before quantization and projected back up afterward. Stochastic sampling replaces the always-nearest-match rule with sampling, controlled by sample_codebook_temp, where the README states that a temperature of 0 is equivalent to non-stochastic behaviour. Shared codebooks across all quantizers, from the RQ-VAE paper, are enabled with shared_codebook = True and reduce the total parameter count.

These interact. A shared codebook with stochastic sampling and a low temperature is a very different training regime from a per-quantizer codebook with hard nearest-neighbour lookup, and the README does not quantify the difference. The honest summary is that the library hands you the published remedies and leaves the diagnosis to you.

Where vector-quantize-pytorch is the wrong tool

This is a layer library, not a model zoo and not a training framework. There is no pretrained checkpoint, no dataset loader, and no training loop in the package. The examples directory contains full autoencoder scripts, such as examples/autoencoder.py, but those are examples you read and adapt, not a supported CLI. If you wanted to quantize an existing model's activations without writing training code, this package will not do it for you.

There is also no inference-time API for encoding a new sample with a saved codebook. The README documents the module's forward pass, not a serialization format for the codebook plus encoder. You would save the module state dict yourself and reconstruct the same architecture to load it.

The package is also not a general-purpose compression tool. Vector quantization here is a differentiable training-time operation tied to a reconstruction loss. If you want to compress a fixed dataset of embeddings offline, a k-means or product quantization implementation from a nearest-neighbour library is the more direct route, because it does not carry a commitment loss or a gradient estimator.

One more boundary: pyproject.toml declares the development status as Beta. That is the project's own classifier, and it matches the surface area, which changes often across releases.

Alternatives and the difference in approach

The closest alternative in the PyTorch ecosystem is the residual quantization layer inside the Encodec and SoundStream style implementations shipped with audio codec projects. Those bundle the quantizer with a convolutional encoder, a decoder, and a training script. The difference is scope: vector-quantize-pytorch gives you a configurable layer that drops into any architecture, while a codec repository gives you a working pipeline whose quantizer is not designed to be swapped out. If you are building an audio codec and want the published architecture, the codec repository saves you the assembly. If you are building something the codec repository does not anticipate, the layer is the better starting point.

For non-neural compression of embeddings, FAISS is the standard reference. It implements product quantization and its variants as an index for similarity search, with no training loss and no straight-through estimator. The difference is fundamental: FAISS quantizes vectors to answer nearest-neighbour queries quickly, while this package quantizes vectors so that gradients can flow back into an encoder. Choosing between them is a question of whether you are training a model or indexing a dataset.

Within the package itself, the alternative is writing the layer yourself. That is a few hundred lines for the basic case, and it is what many teams do before discovering that the residual, grouped, and initialization variants each need their own care. The library's value is in having those variants behind consistent argument names.

Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-02, roughly two weeks before this writing, so the codebase is being touched. The recent release list shows 1.27.21 on 2026-02-12, with 1.27.20 and 1.27.19 in the weeks before that, while pyproject.toml carries version 1.31.4. That gap between the release tags and the declared version is worth noting if you pin by version number rather than by commit.

Upgrade cost is real here. The package has no stability guarantee beyond its Beta classifier, argument names have accumulated across many papers, and new quantization techniques arrive as new keyword arguments rather than as a separate module. If you pin a version, read the diff before moving, because a change to the default gradient estimator or the index layout would silently alter your model's behaviour rather than raise an error.

The licence is MIT, declared in pyproject.toml and shipped as LICENSE. That is permissive and imposes no copyleft obligation on your own code. It does not follow that the techniques are unencumbered: the README cites papers by Deepmind, OpenAI, and others, and the package is described as originally transcribed from Deepmind's tensorflow implementation. Patent questions around specific quantization methods are separate from the software licence, and this is not legal advice.

Editorial conclusion

Adopt vector-quantize-pytorch if you are building a VQ-VAE, RQ-VAE, or audio codec model in PyTorch and want the quantization layer, the commitment loss, and the codebook update rules already implemented. Do not adopt it if you need a trained model, a data pipeline, or a training loop: the package ships the layer and examples, not an end-to-end system. Before committing, verify three things against your own tensors: the shape of the indices your downstream decoder expects, whether rotation_trick or directional_reparam improves reconstruction in your setup, and how many codebook entries stay unused after a few thousand steps, since dead entries are the documented failure mode and no metric for them is exported by the layer.

Frequently asked questions

What does it mean to quantize a vector?

It means replacing a continuous vector with the nearest entry from a fixed set, called the codebook. In this library the replacement is returned as an integer index per position, and the codebook is updated with exponential moving averages by default.

What is vector quantization and how does it work in vector-quantize-pytorch?

An encoder output is matched to the closest codebook entry, and the matched entry is passed to the decoder in place of the original vector. The VectorQuantize class returns the quantized tensor, the indices, and a commitment loss that is added to the reconstruction loss during training.

What is the difference between product quantization and vector quantization?

The README does not cover product quantization, so this package offers no direct comparison. What it does document is residual and grouped residual quantization, where multiple quantizers or multiple feature groups are used instead of a single codebook.

What does VQ-VAE stand for and what is its purpose?

The README refers to VQ-VAE models and to VQ-VAE-2 as a use of vector quantization by Deepmind for high quality image generation, and notes that VQ-VAEs are traditionally trained with the straight-through estimator. The package supplies the quantization layer such models are built around.

Official sources

  1. Issues
  2. License: MIT
  3. lucidrains/vector-quantize-pytorch on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lucidrains-vector-quantize-pytorch.svg)](https://hysenlabs.com/projects/lucidrains-vector-quantize-pytorch)