# kyegomez/BitNet: PyTorch Implementation of 1-Bit Transformer Quantization

> BitNet is a PyTorch library that implements the linear layer quantization method from the paper 'BitNet: Scaling 1-bit Transformers for Large Language Models.' It provides drop-in replacements for nn.Linear that binarize weights, and includes a full BitNetTransformer, BitMGQA attention and BitFeedForward. Models must be trained from scratch or fine-tuned with BitLinear; swapping layers in an already-trained model does not work.

**kyegomez/BitNet** — Implementation of "BitNet: Scaling 1-bit Transformers for Large Language Models" in pytorch

- Repository: https://github.com/kyegomez/BitNet
- Website: https://discord.gg/qUtxnK2NMf
- Stars: 1,946 · Forks: 174
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kyegomez-bitnet

## What BitNet Does and the Problem It Targets

Standard transformer models use 16-bit or 32-bit floating point weights. The BitNet paper proposes replacing the linear projection layers with 1-bit quantized layers, where each weight is either +1 or -1. The motivation is that 1-bit weights reduce memory bandwidth and can replace matrix multiplications with additions, which are cheaper in hardware.

This library implements that substitution as a drop-in replacement for `nn.Linear` in PyTorch. The replacement is called `BitLinear`. The forward pass for BitLinear follows the sequence: tensor, then layer normalization, then binarization, then absolute max quantization, then dequantization. Once the library is installed, swapping an existing `nn.Linear` for `BitLinear` requires changing one constructor call.

The README adds a key constraint directly: 'A model obviously needs to be finetuned from scratch to use BitLinear, just changing the linear methods in an already trained model isn't going to work.' This means the library is for training new quantized models or fine-tuning existing ones with binarized weights from the start, not for post-training quantization of a checkpoint.

## Installation and Basic Usage

Install from PyPI:

```bash
pip3 install bitnet
```

The minimal `BitLinear` usage:

```python
import torch
from bitnet import BitLinear

x = torch.randn(10, 1000, 512)
layer = BitLinear(512, 400)
y = layer(x)
```

For a complete transformer:

```python
import torch
from bitnet import BitNetTransformer

x = torch.randint(0, 20000, (1, 1024))
bitnet = BitNetTransformer(
    num_tokens=20000,
    dim=1024,
    depth=6,
    heads=8,
    ff_mult=4,
)
logits = bitnet(x)
```

The `BitNetTransformer` includes multi-head attention and `BitFeedForward` layers with residual connections. The README notes it can handle not only text but also images and potentially video or audio processing.

## Available Modules: BitLinear Variants, Attention and Feed-Forward

The library ships several components. `BitLinear` is the base 1-bit layer from the original paper. `BitLinearNew` is an updated variant with a groups parameter. `BitMGQA` is a multi-grouped query attention layer that uses `BitLinear` for projections; the README notes that Multi-Grouped Query Attention is recognized for fast decoding and long context handling.

`BitFeedForward` implements the feed-forward block from the diagram with BitLinear and GELU activation: Linear then GELU then Linear. It accepts parameters for hidden dimension, number of layers, swish activation, post-activation layer normalization and dropout:

```python
ff = BitFeedForward(512, 512, 4, swish=True, post_act_ln=True, dropout=0.1)
```

A drop-in replacement function `replace_linears_in_pytorch_model` traverses an existing PyTorch model and replaces all `nn.Linear` layers with `BitLinear`. A Hugging Face variant `replace_linears_in_hf` does the same for Hugging Face Transformers models. Both functions exist for convenience but the README's warning still applies: the model must then be trained or fine-tuned from scratch with the replaced layers.

A CUDA kernel (`gemm_lowbit_ext`) is included for optimized low-bit matrix multiplication. Building it requires running `python setup.py build_ext --inplace` before use.

## The 1.58-Bit Extension and Its Current State

The NEWS section references a second paper, 'The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits,' which extends the binarization to ternary weights (+1, 0, -1) rather than binary. The library aims to implement this as `BitLinearNew`.

However, the README explicitly states: 'The new BitLinear 1.5 is still in progress. There are still some bugs like with the dequantization algorithm and we still need to replace the multiplication with elementwise addition.' This is a direct caution from the repository maintainer. Users who need the 1.58-bit variant should verify the current state of `BitLinearNew` before relying on it.

The same NEWS section invites contributors to join the Agora discord to help implement the 1.58-bit variant. This signals that the library is under active community development rather than being a mature, production-ready implementation.

## Hugging Face and Generic Model Integration

The library includes a path for replacing linear layers in a Hugging Face Transformers model. The `replace_linears_in_hf` function takes a loaded model object and traverses it:

```python
from transformers import AutoModelForSequenceClassification
from bitnet import replace_linears_in_hf

model = AutoModelForSequenceClassification.from_pretrained("bert-base-uncased")
replace_linears_in_hf(model)
```

After replacement, the model must be trained or fine-tuned. For inference on a previously trained BitNet model, `BitNetInference` handles loading a checkpoint and generating text:

```python
from bitnet import BitNetInference

bitnet = BitNetInference()
bitnet.load_model("../model_checkpoint.pth")
output_str = bitnet.generate("The dog jumped over the ", 512)
```

The `examples/` directory includes scripts for bit attention, bit feed-forward, the BitLinearNew variant, Mamba integration, mixture of experts, Hugging Face usage, a kernel test, and a one-bit Vision Transformer. These cover the main usage patterns beyond the basic README examples.

## Limitations and When to Look Elsewhere

The most important limitation is that BitLinear is a training-time quantization method, not a post-training quantization method. If you have a trained model and want to reduce its memory footprint or inference cost, this library does not apply directly. Post-training quantization libraries such as GPTQ, AWQ or BitsAndBytes quantize existing checkpoints and can be applied without retraining.

The CUDA kernel requires building a C extension (`python setup.py build_ext --inplace`) and a CUDA installation. The standard pip install does not include compiled CUDA kernels.

The library depends on `torch`, `einops` and `zetascale`. The `zetascale==2.1.6` pin in `requirements.txt` may create conflicts with environments that use different zetascale versions. The `pyproject.toml` pins Python to 3.6 in the classifiers, though the actual `requires-python` field is set to `^3.10`, so Python 3.10 or newer is required.

Microsoft published its own bitnet.cpp repository, which provides an official CPU inference implementation for 1-bit language models. That project focuses on inference efficiency on CPUs without GPU requirements, which is a different goal from this library's training-focused approach.

## Maintenance Status and Licence

The library is at version 0.2.5 according to `pyproject.toml`. The last push to the repository was on 2026-09-21, seven days before today's date. The repository is not archived. There are no GitHub releases; version tracking is through PyPI.

The project is MIT licensed and built by Kye Gomez under the Agora organization. The README includes an appreciation section naming specific contributors: Dimitry and Nullonix for code review and revision, and Vyom for providing a 4080 GPU for training experiments.

The build system uses Poetry. The lint dependencies in `pyproject.toml` include ruff, mypy-protobuf and black. The `setup.py` handles the optional CUDA extension build and includes platform detection for darwin, linux and win32. The `data/` directory at the repository root and the `train.py` script suggest that training examples or small datasets are provided for validating the BitNetTransformer, though the README does not document what data is included.

The `examples/` directory covers a range of use cases: `bit_attention.py`, `bit_ffn.py`, `bit_linear_new.py`, `bit_mamba.py`, `bit_moe_example.py`, `huggingface_example.py`, `kernel_test.py`, `one_bit_vit.py` and `transformer_example.py`. These demonstrate integrations beyond the basic text transformer, including a Mamba state-space model variant, a mixture-of-experts setup, and a Vision Transformer, indicating that the binarization approach is being tested across different architecture families.

## Conclusion

kyegomez/BitNet is the right starting point for researchers and engineers who want to experiment with 1-bit quantized transformers in PyTorch without building the binarization logic from scratch. The critical constraint is that BitLinear is not a post-training quantization method: the README is explicit that a model needs to be trained from scratch or fine-tuned with BitLinear, and that swapping the layers in an already-trained model will not work. The BitLinear 1.5 implementation (for the 1.58-bit variant from the second paper) is listed in the NEWS section as still in progress with known bugs. Before building production pipelines on this library, verify that the specific module you need (BitLinear, BitLinearNew, BitMGQA or the full transformer) is in a stable state in the current version.

## FAQ

### What is BitNet?

BitNet is a PyTorch library that implements 1-bit linear layers for transformer models, based on the BitNet paper. It provides drop-in replacements for nn.Linear that binarize weights. Models using BitLinear must be trained from scratch or fine-tuned with the quantized layers; swapping layers in an already-trained model does not work.

### What is BitNet b1.58?

BitNet b1.58 refers to a follow-up paper titled 'The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits,' which uses ternary weights (+1, 0, -1) instead of binary. The kyegomez/BitNet library is working on implementing this as BitLinearNew, but the README notes the implementation still has bugs in the dequantization algorithm.

### How do I install BitNet?

Run `pip3 install bitnet`. The optional CUDA kernel for optimized low-bit matrix multiplication requires a separate build step: `python setup.py build_ext --inplace`. The library requires Python 3.10 or newer.

## Sources

- [Issues](https://github.com/kyegomez/BitNet/issues)
- [kyegomez/BitNet on GitHub](https://github.com/kyegomez/BitNet)
- [License: MIT](https://github.com/kyegomez/BitNet/blob/main/LICENSE)
- [Project website](https://discord.gg/qUtxnK2NMf)
- [README](https://github.com/kyegomez/BitNet/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kyegomez-bitnet
