kyegomez/BitNet: PyTorch Implementation of 1-Bit Transformer Quantization
Implementation of "BitNet: Scaling 1-bit Transformers for Large Language Models" in pytorch
At a glance
- What is it?
- BitNet is a PyTorch library that implements the linear layer quantization method from the paper 'BitNet: Scaling 1-bit Transformers for Large Language Models.' It provides drop-in replacements for nn.Linear that binarize weights, and includes a full BitNetTransformer, BitMGQA attention and BitFeedForward. Models must be trained from scratch or fine-tuned with BitLinear; swapping layers in an already-trained model does not work.
- Who is it for?
- kyegomez/BitNet is the right starting point for researchers and engineers who want to experiment with 1-bit quantized transformers in PyTorch without building the binarization logic from scratch. The critical constraint is that BitLinear is not a post-training quantization method: the README is explicit that a model needs to be trained from scratch or fine-tuned with BitLinear, and that swapping the layers in an already-trained model will not work.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What BitNet Does and the Problem It Targets
Standard transformer models use 16-bit or 32-bit floating point weights. The BitNet paper proposes replacing the linear projection layers with 1-bit quantized layers, where each weight is either +1 or -1. The motivation is that 1-bit weights reduce memory bandwidth and can replace matrix multiplications with additions, which are cheaper in hardware.
This library implements that substitution as a drop-in replacement for `nn.Linear` in PyTorch. The replacement is called `BitLinear`. The forward pass for BitLinear follows the sequence: tensor, then layer normalization, then binarization, then absolute max quantization, then dequantization. Once the library is installed, swapping an existing `nn.Linear` for `BitLinear` requires changing one constructor call.
The README adds a key constraint directly: 'A model obviously needs to be finetuned from scratch to use BitLinear, just changing the linear methods in an already trained model isn't going to work.' This means the library is for training new quantized models or fine-tuning existing ones with binarized weights from the start, not for post-training quantization of a checkpoint.
Installation and Basic Usage
Install from PyPI:
pip3 install bitnetThe minimal `BitLinear` usage:
import torch
from bitnet import BitLinear
x = torch.randn(10, 1000, 512)
layer = BitLinear(512, 400)
y = layer(x)For a complete transformer:
import torch
from bitnet import BitNetTransformer
x = torch.randint(0, 20000, (1, 1024))
bitnet = BitNetTransformer(
num_tokens=20000,
dim=1024,
depth=6,
heads=8,
ff_mult=4,
)
logits = bitnet(x)The `BitNetTransformer` includes multi-head attention and `BitFeedForward` layers with residual connections. The README notes it can handle not only text but also images and potentially video or audio processing.
Available Modules: BitLinear Variants, Attention and Feed-Forward
The library ships several components. `BitLinear` is the base 1-bit layer from the original paper. `BitLinearNew` is an updated variant with a groups parameter. `BitMGQA` is a multi-grouped query attention layer that uses `BitLinear` for projections; the README notes that Multi-Grouped Query Attention is recognized for fast decoding and long context handling.
`BitFeedForward` implements the feed-forward block from the diagram with BitLinear and GELU activation: Linear then GELU then Linear. It accepts parameters for hidden dimension, number of layers, swish activation, post-activation layer normalization and dropout:
ff = BitFeedForward(512, 512, 4, swish=True, post_act_ln=True, dropout=0.1)A drop-in replacement function `replace_linears_in_pytorch_model` traverses an existing PyTorch model and replaces all `nn.Linear` layers with `BitLinear`. A Hugging Face variant `replace_linears_in_hf` does the same for Hugging Face Transformers models. Both functions exist for convenience but the README's warning still applies: the model must then be trained or fine-tuned from scratch with the replaced layers.
A CUDA kernel (`gemm_lowbit_ext`) is included for optimized low-bit matrix multiplication. Building it requires running `python setup.py build_ext --inplace` before use.
The 1.58-Bit Extension and Its Current State
The NEWS section references a second paper, 'The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits,' which extends the binarization to ternary weights (+1, 0, -1) rather than binary. The library aims to implement this as `BitLinearNew`.
However, the README explicitly states: 'The new BitLinear 1.5 is still in progress. There are still some bugs like with the dequantization algorithm and we still need to replace the multiplication with elementwise addition.' This is a direct caution from the repository maintainer. Users who need the 1.58-bit variant should verify the current state of `BitLinearNew` before relying on it.
The same NEWS section invites contributors to join the Agora discord to help implement the 1.58-bit variant. This signals that the library is under active community development rather than being a mature, production-ready implementation.
Hugging Face and Generic Model Integration
The library includes a path for replacing linear layers in a Hugging Face Transformers model. The `replace_linears_in_hf` function takes a loaded model object and traverses it:
from transformers import AutoModelForSequenceClassification
from bitnet import replace_linears_in_hf
model = AutoModelForSequenceClassification.from_pretrained("bert-base-uncased")
replace_linears_in_hf(model)After replacement, the model must be trained or fine-tuned. For inference on a previously trained BitNet model, `BitNetInference` handles loading a checkpoint and generating text:
from bitnet import BitNetInference
bitnet = BitNetInference()
bitnet.load_model("../model_checkpoint.pth")
output_str = bitnet.generate("The dog jumped over the ", 512)The `examples/` directory includes scripts for bit attention, bit feed-forward, the BitLinearNew variant, Mamba integration, mixture of experts, Hugging Face usage, a kernel test, and a one-bit Vision Transformer. These cover the main usage patterns beyond the basic README examples.
Limitations and When to Look Elsewhere
The most important limitation is that BitLinear is a training-time quantization method, not a post-training quantization method. If you have a trained model and want to reduce its memory footprint or inference cost, this library does not apply directly. Post-training quantization libraries such as GPTQ, AWQ or BitsAndBytes quantize existing checkpoints and can be applied without retraining.
The CUDA kernel requires building a C extension (`python setup.py build_ext --inplace`) and a CUDA installation. The standard pip install does not include compiled CUDA kernels.
The library depends on `torch`, `einops` and `zetascale`. The `zetascale==2.1.6` pin in `requirements.txt` may create conflicts with environments that use different zetascale versions. The `pyproject.toml` pins Python to 3.6 in the classifiers, though the actual `requires-python` field is set to `^3.10`, so Python 3.10 or newer is required.
Microsoft published its own bitnet.cpp repository, which provides an official CPU inference implementation for 1-bit language models. That project focuses on inference efficiency on CPUs without GPU requirements, which is a different goal from this library's training-focused approach.
Maintenance Status and Licence
The library is at version 0.2.5 according to `pyproject.toml`. The last push to the repository was on 2026-09-21, seven days before today's date. The repository is not archived. There are no GitHub releases; version tracking is through PyPI.
The project is MIT licensed and built by Kye Gomez under the Agora organization. The README includes an appreciation section naming specific contributors: Dimitry and Nullonix for code review and revision, and Vyom for providing a 4080 GPU for training experiments.
The build system uses Poetry. The lint dependencies in `pyproject.toml` include ruff, mypy-protobuf and black. The `setup.py` handles the optional CUDA extension build and includes platform detection for darwin, linux and win32. The `data/` directory at the repository root and the `train.py` script suggest that training examples or small datasets are provided for validating the BitNetTransformer, though the README does not document what data is included.
The `examples/` directory covers a range of use cases: `bit_attention.py`, `bit_ffn.py`, `bit_linear_new.py`, `bit_mamba.py`, `bit_moe_example.py`, `huggingface_example.py`, `kernel_test.py`, `one_bit_vit.py` and `transformer_example.py`. These demonstrate integrations beyond the basic text transformer, including a Mamba state-space model variant, a mixture-of-experts setup, and a Vision Transformer, indicating that the binarization approach is being tested across different architecture families.
Editorial conclusion
kyegomez/BitNet is the right starting point for researchers and engineers who want to experiment with 1-bit quantized transformers in PyTorch without building the binarization logic from scratch. The critical constraint is that BitLinear is not a post-training quantization method: the README is explicit that a model needs to be trained from scratch or fine-tuned with BitLinear, and that swapping the layers in an already-trained model will not work. The BitLinear 1.5 implementation (for the 1.58-bit variant from the second paper) is listed in the NEWS section as still in progress with known bugs. Before building production pipelines on this library, verify that the specific module you need (BitLinear, BitLinearNew, BitMGQA or the full transformer) is in a stable state in the current version.
Frequently asked questions
What is BitNet?
BitNet is a PyTorch library that implements 1-bit linear layers for transformer models, based on the BitNet paper. It provides drop-in replacements for nn.Linear that binarize weights. Models using BitLinear must be trained from scratch or fine-tuned with the quantized layers; swapping layers in an already-trained model does not work.
What is BitNet b1.58?
BitNet b1.58 refers to a follow-up paper titled 'The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits,' which uses ternary weights (+1, 0, -1) instead of binary. The kyegomez/BitNet library is working on implementing this as BitLinearNew, but the README notes the implementation still has bugs in the dequantization algorithm.
How do I install BitNet?
Run `pip3 install bitnet`. The optional CUDA kernel for optimized low-bit matrix multiplication requires a separate build step: `python setup.py build_ext --inplace`. The library requires Python 3.10 or newer.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/kyegomez-bitnet)