Model or dataset
bitsandbytes-foundation/bitsandbytes avatar
bitsandbytes-foundation/bitsandbytes

bitsandbytes: GPU-Accelerated Inference and Training at Lower Bit Depths

Accessible large language models via k-bit quantization for PyTorch.

8,484 stars929 forksPythonMIT

At a glance

What is it?
bitsandbytes reduces memory consumption for large language models via 8-bit and 4-bit quantization. LLM.int8() halves inference memory without performance loss; QLoRA enables training on consumer GPUs. Integrates with PyTorch and Hugging Face libraries.
Who is it for?
Use bitsandbytes if you need to run large models on consumer GPUs without custom quantization code. Do not use it if you need maximum throughput or if your hardware falls outside the supported platforms.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 22 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

Why quantization matters for LLM deployment

Large language models contain billions or trillions of parameters, each typically stored as 32-bit floating-point values. A 70 billion parameter model as float32 requires roughly 280 GB of memory. Most consumer GPUs have 8 to 48 GB of VRAM, making inference impossible without partitioning across multiple GPUs or loading weights from disk repeatedly.

Quantization reduces memory by storing weights at lower precision. An 8-bit integer requires one-quarter the memory of float32. A 4-bit value requires one-eighth. bitsandbytes enables inference with half the required memory and training with several memory-saving techniques without compromising performance.

How LLM.int8() quantizes weights and activations

LLM.int8() uses vector-wise quantization to handle the distribution of values in neural networks. Most weights in a trained model cluster near zero; a small number of outliers can be orders of magnitude larger. Quantizing outliers aggressively causes quality loss.

LLM.int8() splits the computation: weights and activations that fit well in 8-bit are quantized, while outliers are handled separately with 16-bit matrix multiplication.

The implementation provides `bitsandbytes.nn.Linear8bitLt`, a drop-in replacement for PyTorch's Linear layer that performs quantization internally. You can swap Linear layers without changing training loops or model architecture.

QLoRA for memory-efficient fine-tuning

Training a model requires not only storing weights but also gradients and optimizer state. With Adam optimizer, a 70 billion parameter model needs roughly 1.4 TB of memory for training. QLoRA reduces this by quantizing the base model to 4-bit and freezing it, then adding a small set of trainable Low-Rank Adaptation (LoRA) weights.

LoRA inserts learnable matrices with far fewer parameters than the original weights. A 4-bit quantized base model plus LoRA weights can fit on a 24 GB GPU where a full 8-bit training would fail.

bitsandbytes provides `bitsandbytes.nn.Linear4bit` for this approach. The library packages QLoRA with Hugging Face PEFT, which handles LoRA weight creation and merging automatically.

8-bit optimizers for full-model training

When you do want to train a full model but memory is still tight, bitsandbytes provides 8-bit optimizers. They use block-wise quantization to reduce the size of optimizer states, which otherwise can be as large as the model itself.

These are available through `bitsandbytes.optim` module and can replace the standard PyTorch optimizers. The trade-off is compute: quantizing and dequantizing during each training step adds overhead, so 8-bit optimizers run slower than float32 optimizers but use far less memory.

Hardware support and platform requirements

bitsandbytes requires Python 3.10 or later and PyTorch 2.4 or later. The library aims for wide backwards compatibility but recommends using the latest PyTorch version. On Linux with x86-64 CPUs, bitsandbytes supports NVIDIA GPUs with SM60 or later (minimum; SM75 recommended), AMD GPUs with CDNA or RDNA architectures, Intel GPUs including Arc A-Series and Data Center GPU Max, and Intel Gaudi2 or Gaudi3. For ARM systems (aarch64), CPU mode works along with NVIDIA SM75+ GPUs. CPU-only mode also works on x86-64 with AVX2 support or better.

On Windows 11 and Server 2022, support covers NVIDIA SM60+, AMD CDNA and RDNA, Intel Arc, and CPU mode with AVX2. On macOS 14 and later, Apple M1 and later support CPU and Metal acceleration, though Metal optimization for 8-bit optimizers is still in planning.

CPU and Apple Metal runs are supported but may lack optimization. The README notes that support matrices evolve; users on older stable releases should check the version-specific tag in the repository for their hardware's support status.

Integration with Hugging Face ecosystem

bitsandbytes is not a standalone replacement for PyTorch but a library that extends it. The official documentation lists integration with Transformers for quantized model loading, Diffusers for image model quantization, and PEFT for LoRA fine-tuning.

You do not directly instantiate quantized models; instead, you load a Hugging Face Transformers model and pass quantization arguments to the loader. The library handles the weight conversion. This integration is the primary use case and the most tested configuration. The library is listed as a dependency in projects like vLLM and integrated deeply into the Hugging Face ecosystem, making it the standard choice for quantization in that community.

When quantization breaks things

Not all models or tasks work with quantization. Some architectures may not have been tested with 8-bit or 4-bit weights. The library does not provide guidance on which models are known to work or fail.

Quantization introduces numerical rounding. For some use cases like machine translation or code generation, this rounding is negligible; for others like sensitive classification tasks, it may matter. The README does not document when quality loss becomes unacceptable.

Build complexity is another constraint. bitsandbytes includes C++ and CUDA code that must be compiled during installation. The setup.py documentation notes that CMake is required; if compilation fails, you must debug the toolchain yourself. For Intel Gaudi accelerators, QLoRA 4-bit quantization is only partially supported and 8-bit optimizers are not supported. On Apple Metal (macOS with M-series chips), 8-bit optimizer support is still in planned status while 4-bit and 8-bit inference work but may lack performance optimizations.

Editorial conclusion

Use bitsandbytes if you need to run large models on consumer GPUs without custom quantization code. Do not use it if you need maximum throughput or if your hardware falls outside the supported platforms. Before adopting it, verify that your GPU and PyTorch version are on the support matrix: the README documents support for NVIDIA SM60+, AMD CDNA and RDNA, Intel Arc, and Apple M1+ Metal.

Frequently asked questions

What is bitsandbytes?

bitsandbytes is a PyTorch library that enables lower-precision quantization for large language models. It provides 8-bit quantization for inference, 4-bit quantization for training, and 8-bit optimizers for full-model training with reduced memory consumption.

How do I install bitsandbytes with GPU support?

Install via pip: `pip install bitsandbytes`. The package includes precompiled wheels for NVIDIA CUDA and AMD ROCm. For Intel GPUs or custom builds, refer to the official documentation at huggingface.co/docs/bitsandbytes.

What is bitsandbytes quantization?

bitsandbytes quantization reduces memory by storing weights at lower bit depths: 8-bit for inference without accuracy loss, and 4-bit for training. Most weights are quantized while outliers are handled separately to preserve quality.

How do I use bitsandbytes for model training?

For fine-tuning, use QLoRA: quantize the model to 4-bit with `bitsandbytes.nn.Linear4bit`, add trainable LoRA weights through the PEFT library, and train. For full-model training, use 8-bit optimizers from `bitsandbytes.optim`.

How does bitsandbytes compare to AWQ?

bitsandbytes focuses on reducing memory for inference and training via 8-bit and 4-bit quantization. AWQ is an alternative quantization method. Both aim to reduce model size, but bitsandbytes integrates with PyTorch and Hugging Face libraries.

Official sources

  1. bitsandbytes-foundation/bitsandbytes on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/bitsandbytes-foundation-bitsandbytes.svg)](https://hysenlabs.com/projects/bitsandbytes-foundation-bitsandbytes)
Community notes

Community notes