Model or dataset
OpenBMB/BMInf avatar
OpenBMB/BMInf

OpenBMB/BMInf: running 10B-parameter language models on a 6GB GTX 1060

Efficient Inference for Big Models

586 stars67 forksPythonApache-2.0

At a glance

What is it?
BMInf is a Python inference package that wraps transformer models with quantized linear layers and an offload-aware block list so large pretrained language models fit on modest GPUs. This article covers the wrapper mechanism, the pip install path, the hardware floor, and where the approach stops being the right tool.
Who is it for?
Adopt BMInf if you have an NVIDIA GPU with compute capability 6.1 or higher, at least 16GB of system memory, and a transformer checkpoint you control well enough to wrap or to edit layer by layer. Do not adopt it if you need a model hub, a serving stack, or a supported path to non-NVIDIA hardware, because the README documents none of those.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 85 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem BMInf solves: a 10B model on a 6GB card

Serving a ten-billion-parameter language model normally means a datacenter GPU or a multi-GPU setup. BMInf takes the opposite position. The README states that it supports running models with more than 10 billion parameters on a single NVIDIA GTX 1060 in its minimum configuration, with 16GB of system memory and a PCI-E 3.0 x16 slot. The recommended configuration is 24GB of memory and a Tesla V100 16GB. The package targets Python developers who already have a transformer checkpoint and a consumer or workstation GPU, and who want inference rather than a hosted endpoint.

The scope widened over time. The 2.0.0 release note says BMInf can now be applied to any transformer-based model, where earlier versions were tied to specific CPM checkpoints. That matters for adoption: the package is not a model distribution, it is a set of replacement layers you apply to a model you already have. If your work is Chinese-language generation with CPM-2 or EVA, the repository's references point at the papers behind those models, and the examples directory contains directories for glm-130B and huggingface. If your work is a different transformer, the wrapper is the entry point to check.

One caveat sits in the README itself: it notes that examples of CPM-1/2 and EVA will be published soon, and that the BMInf-1 README lives in the old_docs directory. So the documentation is split across versions, and a reader following an older tutorial may land on the 1.x interface.

How the wrapper and TransformerBlockList mechanism work

BMInf does not reimplement the model. It substitutes components. The README's quick start shows the intended flow: build the model on CPU, load the state dict, then call bminf.wrapper(model) inside a torch.cuda.device context. The wrapper walks the model and swaps in BMInf's own layers, which is why the README insists the state dict be loaded before wrapping.

When the wrapper does not fit, the README gives the manual route. torch.nn.ModuleList becomes bminf.TransformerBlockList, which takes the list of blocks and a CUDA device index. torch.nn.Linear becomes bminf.QuantizedLinear, constructed from an existing linear layer. QuantizedLinear is where the memory saving comes from; the transformer block list is where execution is organized. Both replacements are structural, so a model whose attention or feed-forward code deviates from the standard pattern may need edits rather than a single wrapper call.

The dependency list confirms the design. requirements.txt names torch, cpm_kernels>=1.0.9 and typing_extensions. cpm_kernels is the piece that carries the low-level CUDA work, which is why the README requires CUDA 10.1 or newer and compute capability 6.1 or higher. The README also records that 1.0.0 removed a cupy dependency and added PyTorch backpropagation support, so gradients are possible, though the project describes itself as an inference package and the paper title mentions tuning as well.

Installing BMInf and wrapping a model for the first time

The README gives two installation paths. The pip path is a single command, and the source path downloads the package and runs setup.py. The Dockerfile in the repository shows what a container build looks like: an nvidia/cuda:10.1-runtime base, python3 and pip3 installed from apt, and bminf installed from a PyPI mirror at a version passed as a build argument.

Start with the pip install:

bash
pip install bminf

That pulls torch, cpm_kernels>=1.0.9 and typing_extensions through requirements.txt. The README states CUDA 10.1 or newer is required and that dependencies install automatically.

Then apply the wrapper. The README's example, in full:

python
import bminf

# initialize your model on CPU
model = MyModel()

# load state_dict before using wrapper
model.load_state_dict(model_checkpoint)

# apply wrapper
with torch.cuda.device(CUDA_DEVICE_INDEX):
    model = bminf.wrapper(model)

Two things to watch. The model must exist on CPU first, and load_state_dict must run before the wrapper, because the wrapper replaces layers and the checkpoint keys belong to the original modules. The CUDA_DEVICE_INDEX constant is a placeholder in the README, not a package symbol; substitute your own device index.

If the wrapper misbehaves, drop to the manual replacements the README documents:

python
module_list = bminf.TransformerBlockList([
	# ...
], [CUDA_DEVICE_INDEX])

linear = bminf.QuantizedLinear(torch.nn.Linear(...))

The block list takes the original blocks and a list containing the device index. QuantizedLinear takes an already-constructed linear layer. After either route, the README points at benchmark/cpm2/encoder.py and benchmark/cpm2/decoder.py for measuring speed on your own machine, which is the only number that reflects your hardware.

Where BMInf stops being the right tool

The hardware floor is the first boundary. Compute capability 6.1 or higher rules out older cards, and the README sends readers to an external CUDA table to check their own GPU rather than listing supported models. If your card falls below that line, nothing in the package helps.

Decoder throughput is the second boundary. The README's performance table reports BMInf decoder speeds of 4.4 tokens per second on a GTX 1060, 12 on a 1080Ti, 19 on a 2080Ti, 20 on a V100 and 26 on an A100. Encoder speeds are far higher, from 718 to 4365 tokens per second on the same range. Those are the project's own figures for CPM2. The gap between encoder and decoder tells you what the package is good at: classification, embedding and other encoder workloads are comfortable, while open-ended generation on a 1060 runs at roughly four tokens a second. For a long response that is a visible wait, and no amount of configuration changes it.

The third boundary is integration. BMInf is a library you import into your own process. The README documents no HTTP server, no batching server, no model download command and no quantization configuration file. If your requirement is a service that several applications call, you are building that layer yourself on top of the wrapped model. And because wrapping is structural, a model that does not decompose into standard transformer blocks and linear layers will need manual surgery, with the README's two replacement classes as the only documented tools.

BMInf compared with plain PyTorch inference

The most direct alternative is running the same checkpoint in plain PyTorch on the same GPU, and the README provides the comparison itself. On a Tesla V100, PyTorch decodes at 3 tokens per second against BMInf's 20; on an A100, 7 against 26. The README also notes that even where GPU memory is sufficient for large model inference, such as on a V100 or A100, BMInf still has a significant performance improvement over the existing PyTorch implementation.

The difference in approach is what produces that gap. Plain PyTorch keeps the model's layers as written and relies on the framework's own kernels and memory management. BMInf replaces the linear layers with quantized equivalents and groups transformer blocks into its own list, backed by cpm_kernels. That is a trade: you gain lower memory use and higher throughput, and you give up the ability to treat the model as an ordinary nn.Module tree without checking what the wrapper did to it. Debugging a wrapped model means reasoning about QuantizedLinear rather than nn.Linear.

A second alternative is simply using a smaller model that fits your GPU without any of this. BMInf's value proposition is the parameter count, not convenience. If a model that fits in memory unmodified already meets your quality bar, the wrapper adds a dependency and a layer of indirection for no benefit. The project's own framing, efficient inference for big models, tells you which side of that line it is aimed at.

Maintenance, versioning and the Apache 2.0 licence

The repository is not archived, and the last push was on 2026-07-07. The most recent tagged release is 2.0.1 from 2023-01-24, following 2.0.0 in July 2022 and 1.0.2 in January 2022. So commits have continued after the last release, but the release cadence has been slow, and anyone pinning a version should expect to sit on 2.0.1 for a while.

Version handling has a quirk worth knowing. setup.py resolves the version in a fixed order: an explicit CONFIG value, then the BM_VERSION environment variable, then CI_COMMIT_TAG, then CI_COMMIT_SHA, then __version__ in bminf/version.py, and finally the literal string "test". It also writes the resolved value back into bminf/version.py during the build. Installing from source outside a CI environment with none of those variables set can therefore produce a package stamped "test". Install from pip if the version string matters to you, or set BM_VERSION explicitly.

The licence is Apache 2.0, per the README and the LICENSE file. That is a permissive licence, but it covers the BMInf code, not the model weights you load into it. CPM-2, EVA, GLM-130B and any Hugging Face checkpoint carry their own terms, and those terms are what govern your use of the model. Check the model licence separately; the package licence tells you nothing about it. This is a description of the files, not legal advice.

Editorial conclusion

Adopt BMInf if you have an NVIDIA GPU with compute capability 6.1 or higher, at least 16GB of system memory, and a transformer checkpoint you control well enough to wrap or to edit layer by layer. Do not adopt it if you need a model hub, a serving stack, or a supported path to non-NVIDIA hardware, because the README documents none of those. Before committing, check your GPU against the CUDA compute capability table the README links, confirm your CUDA version is 10.1 or newer, and run benchmark/cpm2/decoder.py on your own card, since the published decoder numbers span 4.4 to 26 tokens per second across GPUs and that range decides whether interactive use is realistic.

Frequently asked questions

What GPU do I need to run OpenBMB/BMInf?

The README lists an NVIDIA GeForce GTX 1060 6GB as the minimum GPU and a Tesla V100 16GB as recommended, with 16GB and 24GB of system memory respectively. GPUs must have compute capability 6.1 or higher, and the README links to an external CUDA table for checking a specific card.

How do I install BMInf?

The README gives two paths: pip install bminf, or downloading the package and running python setup.py install. Dependencies including torch and cpm_kernels>=1.0.9 are installed automatically, and CUDA 10.1 or newer is required.

Can BMInf run models other than CPM?

The 2.0.0 release note states that BMInf can now be applied to any transformer-based model. The README shows bminf.wrapper as the general entry point, with bminf.TransformerBlockList and bminf.QuantizedLinear as manual replacements when the wrapper does not fit.

Why is BMInf generation slow on a GTX 1060?

The README's performance table reports 4.4 decoder tokens per second on a GTX 1060, against 718 encoder tokens per second on the same card. Decoder speed rises to 26 tokens per second on an A100, so generation throughput is bounded by the GPU rather than by configuration.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. OpenBMB/BMInf on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/openbmb-bminf.svg)](https://hysenlabs.com/projects/openbmb-bminf)