Model or dataset
ridgerchu/matmulfreellm avatar
ridgerchu/matmulfreellm

MatMul-Free LM: running a ternary-weight language model through Hugging Face Transformers

Implementation for MatMul-free LM.

3,092 stars203 forksPythonApache-2.0

At a glance

What is it?
The matmulfreellm repository packages the HGRNBit architecture as a Transformers-compatible model, so you can load a 370M to 2.7B checkpoint and generate text without a single dense matrix multiplication. It is research code with a CUDA build path, not a drop-in replacement for a standard decoder stack.
Who is it for?
Adopt it if you are reproducing the scalable MatMul-free language modeling paper, studying ternary-weight training, or evaluating HGRNBit as a linear-attention alternative inside an existing Transformers pipeline. Do not adopt it if you need a production inference server, CPU-only deployment, or a model whose weights are ordinary fp16 matrices.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 25 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What MatMul-Free LM removes, and who that matters to

Every transformer block spends most of its arithmetic on matrix multiplication: the query, key, value and output projections, plus the two or three matrices in the feed-forward network. MatMul-Free LM replaces those dense products with ternary weights and element-wise operations, so the heavy work becomes addition and gating rather than multiply-accumulate. The repository README describes the project as "a language model architecture that eliminates the need for Matrix Multiplication (MatMul) operations" and ships an implementation compatible with the Hugging Face Transformers library.

The audience is narrow but real. Researchers who want to reproduce the scaling comparison between Transformer++ and HGRNBit at 370M, 1.3B and 2.7B parameters are the primary users. Engineers probing whether ternary weights can cut energy or memory traffic on hardware that favors accumulation are the second group. If you are looking for a model to put behind an API tomorrow, this is the wrong repository, and the README does not present it as anything else.

HGRNBit: ternary projections, gated linear attention and the FusedBitLinear block

The architecture is visible in the model dump the README prints. Each HGRNBitBlock holds an attention sublayer and an MLP sublayer. The attention module contains four projections, i_proj, f_proj, g_proj and o_proj, all of them FusedBitLinear instances with bias disabled. The MLP holds gate_proj and down_proj, also FusedBitLinear. The names tell you where the multiplication would normally live, and the Bit suffix tells you what replaced it: the weights are ternarized and the linear layer fuses a normalization step, each FusedBitLinear carrying its own RMSNorm with eps=1e-08.

Two details stand out. First, the attention is gated linear attention rather than softmax attention, which is why the block has separate input, forget and gate projections instead of a single query-key-value bundle. Second, the MLP is asymmetric: gate_proj maps 2048 to 11264, but down_proj maps 5632 back to 2048. That mismatch is not a typo in the dump, it is the shape produced by gating a subset of channels, and anyone porting the block by hand should copy the shapes from the config rather than assume symmetry.

The README's scaling-law section claims the projection for this model "exhibits a steeper descent compared to Transformer++, suggesting our architecture is more efficient in leveraging additional compute to improve performance." That is the paper's own reading of its curves. Treat it as a hypothesis to reproduce, not a settled result.

Installing matmulfreellm from git and generating your first tokens

The README lists three requirements before installation: PyTorch 2.0 or newer, Triton 2.2 or newer, and einops. The install command pulls directly from the repository rather than PyPI, which means you are building the package from the default branch:

bash
pip install -U git+https://github.com/ridgerchu/matmulfreellm

setup.py reveals a wrinkle worth knowing before you run that. The package name is mmfreelm, and CUDA compilation is controlled by environment variables that read FLA_FORCE_BUILD, FLA_SKIP_CUDA_BUILD and FLA_FORCE_CXX11_ABI. FLA_SKIP_CUDA_BUILD defaults to "TRUE", so a plain install skips the CUDA build and expects prebuilt wheels or a later compile step. If nvcc is missing, setup.py only warns rather than failing, which means a broken kernel build can surface later as a runtime error.

Once installed, you can instantiate the model from its config without downloading weights, which is the fastest way to confirm the package imported correctly:

python
from mmfreelm.models import HGRNBitConfig
from transformers import AutoModel
config = HGRNBitConfig()
AutoModel.from_config(config)

For a real generation run, the README points at generate.py and shows the pattern: set TOKENIZERS_PARALLELISM to "false", load a tokenizer and a causal LM from a checkpoint name, move the model to CUDA in half precision, tokenize a prompt, and call generate with max_length=32, do_sample=True, top_p=0.4 and temperature=0.6. The checkpoint name is left blank in the example, so you fill it with one of the three models from the MatMul-free LM collection on Hugging Face, for instance the 370M entry. Expect a decode of a single string from the batch.

The CUDA and Triton dependency is the real deployment boundary

Nothing in the README promises CPU inference, and the generation example calls .cuda().half() unconditionally. The kernels underneath come from flash-linear-attention, which the README credits as the source this repository was adapted from, and that lineage means Triton is not optional. Triton 2.2 or newer is a stated requirement, and Triton targets NVIDIA GPUs. If your target is an Apple laptop, a CPU-only CI runner, or an inference host without a CUDA toolchain, the project as documented will not run there.

The second constraint is precision. The example casts the model to half, and the ternary weights interact with that cast in ways the README does not discuss. There is no documented guidance on numerical stability at long context, no statement about what happens when max_length grows well past the 32 tokens in the example, and no rollback or fallback path if a Triton kernel fails to compile. Those are gaps in the documentation, not necessarily defects, but they are the gaps you will hit first.

Finally, the pre-trained zoo is small. Three checkpoints, 370M, 1.3B and 2.7B, trained on 15B, 100B and 100B tokens respectively. There is no instruction-tuned variant, no chat template, and no tokenizer training documentation in the README. For a base language model that is fine; for anything that needs to follow instructions, you would be doing that work yourself.

How this differs from flash-linear-attention and from a standard Llama-style stack

The closest reference point is flash-linear-attention, the project this repository was adapted from. That library is a broad collection of linear attention kernels and layers, designed so you can swap one attention variant for another inside a training script. matmulfreellm takes one of those directions, HGRNBit, and turns it into a full Hugging Face model class with configs, a published checkpoint collection, and a reproducible release tied to a Nature Computational Science manuscript. If you want to compare many linear attention variants, flash-linear-attention is the better starting point. If you want the specific MatMul-free model from the paper, with weights you can download, this repository is the one that ships them.

The other comparison is against an ordinary decoder such as Llama. A Llama block multiplies fp16 or bf16 matrices and relies on softmax attention; HGRNBit ternarizes the projection weights and uses gated linear attention, which changes both the memory profile and the training dynamics. The practical difference is that you cannot take a Llama fine-tuning recipe and apply it unchanged, because the layer names, the projection shapes and the numerical behavior all differ. The upside, if the paper's scaling claim holds, is better loss per unit of compute at the sizes tested. The downside is an ecosystem that is roughly one repository wide.

Version 0.1.0, the Apache-2.0 licence and what upgrading costs

Release v0.1.0 is the only release listed, published on 2026-09-06. The README describes it as "the archival software release associated with the Nature Computational Science manuscript", with runtime code based on commit f24cfe5 and citation, version and licence metadata added on top. That framing matters for upgrade planning: this is a frozen artifact for citation, not a package with a deprecation policy. There is no changelog in the repository listing, so the only way to know what changed between the archival commit and the default branch is to diff them yourself.

The licence is Apache-2.0, which permits commercial use, modification and redistribution provided you keep the notices and state changes. The README asks that you cite both the Zenodo software release, DOI 10.5281/zenodo.22501850, and the arXiv paper 2406.02528. Citation is a request, not a licence condition, but the CITATION.cff file exists specifically so that tools can pick up the metadata. If you ship a product built on these weights, check the licence terms on the Hugging Face checkpoints separately, since a model card can carry its own terms distinct from the code repository. That is a question for your own review, not something this article can settle.

Upgrade cost is low in the sense that there is almost nothing to upgrade, and high in the sense that the Triton and PyTorch floors will move under you. Pinning PyTorch and Triton versions in your environment is the only reliable way to keep the kernels compiling as those projects release.

What the repository does not tell you

Several things a careful adopter would want are absent. There is no documented training script in the README, only the model definition, the generation example and the checkpoint table, so reproducing the 100B-token runs from scratch is not something the README walks you through. There is no benchmark table for throughput or memory, only the scaling-law curve image. There is no stated minimum GPU, no VRAM figure for the 2.7B checkpoint, and no note on whether the half-precision cast in the example is required or merely conventional. There is also no rollback guidance if a kernel build fails, and setup.py's decision to warn rather than error when nvcc is missing means the failure can be silent at install time.

None of this is unusual for a paper release. It does mean that the honest way to evaluate matmulfreellm is to treat it as a reference implementation: read the model code in mmfreelm/, check the shapes against the config dump, and run the generation example on one checkpoint before planning anything larger around it.

Editorial conclusion

Adopt it if you are reproducing the scalable MatMul-free language modeling paper, studying ternary-weight training, or evaluating HGRNBit as a linear-attention alternative inside an existing Transformers pipeline. Do not adopt it if you need a production inference server, CPU-only deployment, or a model whose weights are ordinary fp16 matrices. Before committing, verify that your PyTorch and Triton versions satisfy the stated floors, that the Triton kernels compile on your GPU, and that you can load one of the three published checkpoints from the MatMul-free LM collection on Hugging Face.

Frequently asked questions

What does MatMul stand for and what does it do?

MatMul is short for matrix multiplication, the dense multiply-accumulate operation that dominates transformer projections. MatMul-Free LM removes those operations from its layers by using ternary weights and element-wise gating instead, which the README describes as eliminating the need for Matrix Multiplication operations.

Is llm free to use?

The matmulfreellm code is released under the Apache License 2.0, which permits commercial use and modification with notice retention. The README separately asks that you cite the Zenodo release and the associated paper, and the Hugging Face checkpoints may carry their own terms.

What is the difference between the np.dot() and np.matmul() functions in NumPy?

The repository does not cover NumPy, so this question cannot be answered from the available material. matmulfreellm is a PyTorch and Triton project that avoids matrix multiplication at the model layer rather than comparing NumPy APIs.

Is AI just matrix multiplication?

MatMul-Free LM is a direct counterexample at the architecture level: it replaces the dense matrix products in its attention and MLP projections with ternary weights and element-wise operations. The README frames this as an architecture that eliminates the need for Matrix Multiplication operations, so matrix multiplication is a dominant but not mandatory primitive.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. README
  4. Releases
  5. ridgerchu/matmulfreellm on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ridgerchu-matmulfreellm.svg)](https://hysenlabs.com/projects/ridgerchu-matmulfreellm)