Framework
microsoft/BitNet avatar
microsoft/BitNet

bitnet.cpp announced version 1.0 in its own news feed and has published no releases

GitHub describes it as Official inference framework for 1-bit LLMs. The repository metadata lists C++ as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.

40,352 stars3,739 forksC++MIT

At a glance

What is it?
bitnet.cpp is Microsoft's inference framework for 1-bit LLMs, with optimized CPU and GPU kernels and NPU support described as coming next. The repository is candid about its models and its speedup ranges and quiet about distribution: there are no GitHub releases, the last commit is 2026-07-27, every Python dependency is inherited from llama.cpp, and the headline 100B capability has no published model behind it.
Who is it for?
Evaluate bitnet.cpp if you are running a BitNet b1.58 model on x86 or ARM silicon today and care about the energy numbers, which are stated per architecture and per model, and if you are willing to build from source. Do not plan around a version pin, because the repository has no GitHub releases and the last commit is 2026-07-27, so there is no artifact to pin.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 65 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

There is a news post about version 1.0 and no releases to go with it

The repository has no GitHub releases. That is the whole distribution story, and it sits in tension with the news feed at the top of the file, which contains a line reading 10/17/2024, bitnet.cpp 1.0 released.

The word released is doing several jobs in that feed. It introduces a version in 2024, an official GPU inference kernel in May 2025, a CPU inference optimization write-up in January 2026, an I2_S conversion guide in July 2026, and two embedding models in July 2026. A version, a kernel, a document and a pair of model weights are all announced the same way.

Now the dates. The most recent entry is 07/23/2026, the VibeASR.cpp release. The last commit to the repository is 2026-07-27, four days later, and nothing has been pushed since.

The consequence is that there is no tag to pin, no release note to diff, and no changelog outside the news feed itself. If you are tracking this project for a deployment, your only version handle is a commit hash, and the word released in the documentation does not tell you whether a given item is downloadable.

Every Python requirement is inherited from llama.cpp, and the file forbids adding to it

The top-level requirements file is five lines of includes and nothing else:

code
-r 3rdparty/llama.cpp/requirements/requirements-convert_legacy_llama.txt
-r 3rdparty/llama.cpp/requirements/requirements-convert_hf_to_gguf.txt
-r 3rdparty/llama.cpp/requirements/requirements-convert_hf_to_gguf_update.txt
-r 3rdparty/llama.cpp/requirements/requirements-convert_llama_ggml_to_gguf.txt
-r 3rdparty/llama.cpp/requirements/requirements-convert_lora_to_gguf.txt

The header above them says these requirements include all dependencies for all top-level python scripts for llama.cpp, that you should avoid adding packages there directly, and that package versions must stay compatible across all top-level python scripts.

Every include is a conversion script. Legacy llama conversion, Hugging Face to GGUF, the update variant, GGML to GGUF, LoRA to GGUF. None of them is an inference dependency, and the Python surface they serve is not this project's.

The consequence is that installing the Python environment for a repo whose top level holds run_inference.py, run_inference_server.py and setup_env.py means installing a model conversion toolchain. And the one file you would naturally edit to add a missing package carries an instruction not to edit it, so a contributor has to work out on their own where a new dependency is supposed to live.

3rdparty is a submodule, so the requirement includes start out pointing at nothing

The tree has a .gitmodules file, and a 3rdparty/ directory. The requirements file then reaches into 3rdparty/llama.cpp/requirements/ for all five of its includes.

So a clone that has not initialised submodules has an empty 3rdparty, and every one of those -r lines resolves to a path that does not exist. The failure is not a missing package and not a version conflict. It is a file-not-found on the include line itself, which points at your submodule state rather than at anything about BitNet.

The README does not name a submodule step. The build path it points to is an anchor in this same file, and the GPU path is a separate README of its own, and neither of them mentions that the repository has a vendored dependency that has to be fetched separately.

The consequence is a first-run experience that looks like a broken requirements file. The fix is one command, but the error message will not tell you that, and a reader who has not seen a .gitmodules file in a project before will reasonably conclude the paths in requirements.txt are simply wrong.

The 100B headline has no published model behind it

The Overview makes the largest claim in the file: bitnet.cpp can run a 100B BitNet b1.58 model on a single CPU, achieving speeds comparable to human reading at 5 to 7 tokens per second, which the file presents as significantly enhancing the potential for running LLMs on local devices.

Now count the models that actually ship. The Model Releases section has three: BitNet-b1.58-2B-4T with 2.4B parameters trained on 4 trillion tokens, BitNet-embedding-0.6B, and BitNet-embedding-270M.

So the demonstrated capability and the downloadable artifact are about two orders of magnitude apart. The 100B figure is attributed to the technical report rather than to anything in this repository, and the models here are 2.4B, 0.6B and 270M.

The consequence is worth being blunt about, because the two facts get quoted together. The single-CPU claim is about a model you cannot download from this project, and the models you can download are small enough that a CPU or a modest accelerator handles them without the argument being necessary. Reading the 100B line as a description of what the released weights will do on your machine is a mistake the file invites.

The speedup numbers split by architecture, and the flagship model is the small end

The Overview gives two sets of numbers and they are not the same. On ARM CPUs, speedups run 1.37x to 5.07x with energy consumption reduced by 55.4% to 70.0%. On x86 CPUs, speedups run 2.37x to 6.17x with energy reductions between 71.9% and 82.2%.

Both are ranges, not per-model figures, and the file adds that larger models experience greater performance gains. That single sentence is the one to hold on to, because the largest language model published here is 2.4B parameters.

So the architecture spread is roughly a factor of two at both ends, and the direction of improvement across model size points the other way from the flagship model. The 2.4B model is quoted at up to 6.17x on x86, which is the top of the x86 range.

The consequence is that any number you repeat from this file needs three qualifiers to be honest: which architecture, which end of the range, and which model. A reader who quotes 6.17x without them is quoting the best case on the smaller model on one processor family. The energy percentages have the same structure and deserve the same treatment.

One-bit covers three different quantization schemes under one name

The project name says 1-bit. The material underneath it describes at least three distinct schemes, and the news feed moves between them without flagging the change.

The original line is ternary weights, which is what b1.58 means and what the flagship model section calls a ternary, 1.58-bit, language model. Then there is a4.8, 4-bit activations for 1-bit LLMs, introduced in November 2024 as a further efficiency gain. And the newest work is I2_S quantization, which the embedding bullets pair with lossless inference with 2 bits per weight, and which VibeASR.cpp also uses.

The CPU inference optimization from January 2026 adds a third axis again, since the file credits configurable tiling and embedding quantization support alongside parallel kernels.

The consequence is that one-bit in this project is a family name rather than a specification. The headline numbers in the Overview are attached to b1.58 specifically, while the most recent releases are built on I2_S. Comparing a claim about bitnet.cpp's one-bit-ness against a different system requires knowing which of the three you are reasoning about, and the file will not tell you from the project name.

NPU support is still in the future tense two months after the last commit

The Overview states that the framework supports inference on CPU and GPU, with NPU support coming next. That parenthetical is the only mention of NPUs in the file.

The GPU path is not hypothetical. An official GPU inference kernel shipped in May 2025 and has its own README in the gpu/ directory, and the flagship model's release notes list GPU support as an available capability alongside the CPU speedups and the conversational mode.

So two of the three accelerator classes are live, one has been live for more than a year, and one is described in the future tense. The last commit is 2026-07-27 and the newest news entry is 07/23/2026, so the parenthetical has not been revised since the most recent work landed.

The consequence is that an NPU deployment is a research question, not a configuration question. There is no flag to set, no build option named, and no list of supported devices, because there is no NPU support to configure yet. If your target hardware is an NPU rather than a discrete GPU or a CPU, the honest reading of this file is that you are waiting on a future release.

Editorial conclusion

Evaluate bitnet.cpp if you are running a BitNet b1.58 model on x86 or ARM silicon today and care about the energy numbers, which are stated per architecture and per model, and if you are willing to build from source. Do not plan around a version pin, because the repository has no GitHub releases and the last commit is 2026-07-27, so there is no artifact to pin. Before you start, settle three things: that 3rdparty is a submodule you must initialise, since all five requirement includes resolve into it and will otherwise fail on missing files; that the Python dependency list is llama.cpp's conversion toolchain rather than anything inference-related, and the file forbids adding to it directly; and which of the three quantization schemes you actually mean, because one-bit in this project covers ternary weights, 4-bit activations, and I2_S with 2 bits per weight. If you need an NPU path, it is still described as coming next.

Frequently asked questions

What is BITNET used for?

bitnet.cpp is the official inference framework for 1-bit LLMs such as BitNet b1.58. It offers a suite of optimized kernels for fast and lossless inference of 1.58-bit models on CPU and GPU, with NPU support described as coming next. The published models are one 2.4B language model and two embedding models.

What is BITNET b1 58?

It is the ternary, 1.58-bit, model line introduced by the paper The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits. The official model here is BitNet-b1.58-2B-4T, with 2.4B parameters trained on 4 trillion tokens, and it is quoted at up to 6.17x speedup on x86 CPUs against full-precision models.

how to install bitnet cpp

The README gives no install command. It offers an online demo to try immediately, or building and running it yourself, with a build-from-source anchor in the file and a separate GPU README. The top level holds setup_env.py, run_inference.py and run_inference_server.py, and requirements.txt resolves entirely into a 3rdparty llama.cpp submodule.

how to use bitnet cpp

You build it and run one of the top-level Python scripts, run_inference.py for inference or run_inference_server.py for the server, after setup_env.py. The CPU kernels are described in src/README.md, including parallel implementations with configurable tiling and embedding quantization support, and the GPU kernels have their own guide at gpu/README.md.

what is bitnet model

Three are published: BitNet-b1.58-2B-4T, a 2.4B-parameter ternary language model trained on 4 trillion tokens that supports conversational mode, plus BitNet-embedding-0.6B and BitNet-embedding-270M, embedding models quoted at 1.42x to 2.28x and 1.32x to 1.74x over F16 on prefill with 8 threads on x86.

Official sources

  1. Official README
  2. Project repository