Library / SDK
libxsmm/libxsmm avatar
libxsmm/libxsmm

LIBXSMM: JIT code generation for small matrix multiplications

Library for specialized dense and sparse matrix operations, and deep learning primitives.

978 stars205 forksCBSD-3-Clause

At a glance

What is it?
LIBXSMM is a C library that generates specialized kernels at runtime for small dense and sparse matrix operations, and it now positions itself as the reference implementation of Tensor Processing Primitives. It is a good fit when your matrices are small and your kernel shapes repeat; it is the wrong tool when you want one large GEMM call.
Who is it for?
Adopt LIBXSMM when your workload is many small GEMMs or elementwise primitives with repeated shapes, and when you can accept a build step plus a runtime dispatch call per kernel shape. Do not adopt it if you need one large single GEMM, if you cannot ship a JIT that emits code at runtime, or if you are on a target the README does not list.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem LIBXSMM targets: many small GEMMs, not one big one

A conventional BLAS GEMM is tuned for large matrices. When m, n and k are small (single digits to low hundreds), the call overhead and the generic loop structure dominate, and the arithmetic units sit idle. LIBXSMM is built for that regime. The README describes it as a high performance library for small dense and sparse linear algebra operations including GEMM and elementwise primitives often seen in deep learning applications.

The intended user is someone writing a C or Fortran kernel where the same small shape is dispatched over and over: a convolution lowered to matrix multiplication, a fully connected layer, a small stencil update. The README also states that LIBXSMM serves as the reference implementation of Tensor Processing Primitives, a programming abstraction for deep learning and HPC workloads. That framing matters, because it tells you the project is not trying to be a drop-in replacement for a vendor BLAS. It is a layer you write against directly.

How JIT specialization works, and what dispatch actually does

The mechanism is runtime code generation. Rather than compiling a library of every shape in advance, LIBXSMM takes a shape descriptor, generates machine code for that exact shape, and hands back a function pointer. The README says code generation is mainly based on Just-In-Time code specialization for compiler-independent performance, covering matrix multiplications, transpose/copy, sparse functionality and tensor primitives.

The data flow in the Hello LIBXSMM example is explicit. You build a shape with libxsmm_create_gemm_shape, pass it to libxsmm_dispatch_gemm along with flags and a prefetch setting, and receive a libxsmm_gemmfunction. You then fill a libxsmm_gemm_param struct with pointers to a, b and c, and call the returned function. The same pattern repeats for elementwise kernels: libxsmm_dispatch_meltw_unary, libxsmm_dispatch_meltw_binary and libxsmm_dispatch_meltw_ternary each return a typed function pointer, and you invoke it with a matching param struct.

The README claims this supports build once and deploy everywhere: no special target flags are needed to exploit the available performance. That is the real argument for JIT here. A binary compiled with generic flags still gets ISA-specific kernels at runtime. The trade-off is that dispatch is not free, which is why the root Makefile exposes a CACHE option described as a thread-local cache of recently dispatched kernels, with values 0 to disable, 1 to enable, or a small power-of-two number. If you call dispatch inside a hot loop instead of caching the function pointer yourself, you are fighting the design.

Building LIBXSMM and running the Hello example

The README gives the build and run sequence directly. Build the shared library from the repository root. The STATIC=0 setting selects a shared library rather than a static one.

bash
cd /path/to/libxsmm
make STATIC=0

The build writes into lib/ under the repository root, and headers live under include/. Save the Hello program as hello.c, then compile it against that include directory and link -lxsmm plus -lm.

bash
gcc -I/path/to/libxsmm/include hello.c -L/path/to/libxsmm/lib -lxsmm -lm -o hello

Run it with the library path pointed at the build output. LIBXSMM_VERBOSE=2 is the README's own setting for the example, and it makes the library report what it generated.

bash
LD_LIBRARY_PATH=/path/to/libxsmm/lib LIBXSMM_VERBOSE=2 ./hello

The program itself dispatches four kernels for m=13, n=5, k=7 in FP32: a GEMM producing C = A * B, a unary ReLU, a binary bias add, and a ternary multiply-add. All matrices are column-major. The dispatch calls look like this.

c
const libxsmm_gemm_shape gshape = libxsmm_create_gemm_shape(m, n, k, m, k, m,
  LIBXSMM_DATATYPE_F32, LIBXSMM_DATATYPE_F32, LIBXSMM_DATATYPE_F32, LIBXSMM_DATATYPE_F32);
libxsmm_gemmfunction gemm = libxsmm_dispatch_gemm(gshape,
  LIBXSMM_GEMM_FLAG_NONE, LIBXSMM_GEMM_PREFETCH_NONE);

After that you zero a libxsmm_gemm_param, assign gp.a.primary, gp.b.primary and gp.c.primary, and call gemm(&gp). The sample returns 0 and prints nothing on its own; the observable output is what LIBXSMM_VERBOSE emits. The README points to samples/xgemm and the other sample directories for more complete drivers, and notes that equivalent C and Fortran versions exist under samples/hello.

Where LIBXSMM is the wrong tool

The first limitation is scope. LIBXSMM is aimed at small operations. If your workload is a handful of large GEMMs, a tuned vendor BLAS is the better instrument, and adding a JIT dispatch layer buys you nothing.

The second is the runtime code generation itself. JIT means the library emits and executes machine code while your process runs. That is a deployment constraint, not a footnote: environments that forbid writable-executable memory, or that need a fully auditable binary image, will not accept it. The README does not document a fallback path for such environments.

The third is target coverage. The README lists Intel Architecture with SSE, AVX, AVX2, AVX-512 (with VNNI and Bfloat16) and AMX; AArch64 with NEON, SVE and SME; RISC-V with RVV; and PowerPC 64-bit little-endian (POWER10 with VSX and MMA). Anything outside that list is undocumented territory, and the README does not describe what happens on an unsupported target beyond the generic build-once claim.

The fourth is API surface. Version 2.0 moved application-specific pieces out: the README states that deep-learning operators and convolution drivers now live in LIBXSMM-DNN, the PyTorch integration in TPP PyTorch Extension, and the spectral-element reproducers in their own repositories. If you used those from the 1.x tree, upgrading means adding companion dependencies rather than just bumping a version. The README does not describe a migration path for that split.

How LIBXSMM differs from BLAS and from oneDNN

The clearest comparison is with a vendor BLAS such as OpenBLAS or Intel MKL. Those ship a fixed set of precompiled kernels and select among them at call time. LIBXSMM generates the kernel for the shape you ask for. For large matrices the two converge, because a well-tuned blocked kernel is already close to optimal. For small matrices the precompiled approach pays dispatch and setup costs on every call, which is exactly the gap LIBXSMM was built to close. The cost of LIBXSMM's approach is compile time on first dispatch for a new shape, which is why the kernel cache exists.

The closer comparison is oneDNN, which also targets deep learning primitives and also generates or selects optimized implementations. The difference in approach is the abstraction level. oneDNN exposes operators such as convolution, pooling and normalization as first-class primitives, with layout and memory format handled inside the library. LIBXSMM 2.0 exposes the lower layer: GEMM, BRGEMM and elementwise primitives, from which the README says higher-level operators such as convolutions, fully-connected layers, normalization and pooling are composed. If you want an operator library, oneDNN is the shorter path. If you want to write the composition yourself and control the kernels underneath, LIBXSMM gives you that. The README's own answer to the operator question is to point at LIBXSMM-DNN rather than to claim the core library does it.

Maintenance cost, licensing and what a 2.0 upgrade involves

The repository is not archived, and the last push was on 2026-09-03. Releases 2.0.0 and 2.1.0 landed on 2026-06-30 and 2026-07-25 respectively, after a long gap from 1.17 on 2021-12-03. That gap is worth noting if you are planning an upgrade: the 1.x to 2.x jump is not a routine patch, and the README frames 2.0 as a repositioning of the project around Tensor Processing Primitives, with several application-specific components moved to companion repositories.

The build system is a plain Makefile at the repository root, with CMakeLists.txt and fpm.toml also present for CMake and Fortran package manager users. The root Makefile defines install subdirectories such as PINCDIR, POUTDIR and PPKGDIR, so packaging follows the usual PREFIX conventions. If you vendor LIBXSMM into a larger build, note that the Makefile exposes MNK for generating M,N,K combinations at build time, which is how the static code generation path works alongside JIT.

On licensing: the repository carries the BSD 3-Clause License, and the README links LICENSE.md. The Makefile installs that file into LICFDIR under the documentation share directory. BSD 3-Clause is permissive and does not carry the patent grant language of Apache 2.0, which is a consideration if patent exposure matters to your organization. This is a description of the licence text's general shape, not legal advice; have counsel read LICENSE.md for your own case.

Editorial conclusion

Adopt LIBXSMM when your workload is many small GEMMs or elementwise primitives with repeated shapes, and when you can accept a build step plus a runtime dispatch call per kernel shape. Do not adopt it if you need one large single GEMM, if you cannot ship a JIT that emits code at runtime, or if you are on a target the README does not list. Before committing, verify three things on your own machine: that your CPU falls under one of the documented targets (Intel SSE/AVX/AVX2/AVX-512/AMX, AArch64 NEON/SVE/SME, RISC-V RVV, POWER10), that the datatype you need is in the supported list, and that your deployment can link the shared library and set LD_LIBRARY_PATH. The samples/hello program is the cheapest way to confirm all three at once.

Frequently asked questions

How do I multiply 3x3 matrices in C++ with LIBXSMM?

You do not write the loop yourself. You create a shape with libxsmm_create_gemm_shape for your m, n and k, dispatch it with libxsmm_dispatch_gemm to get a function pointer, fill a libxsmm_gemm_param with pointers to a, b and c, and call that pointer. The README's Hello example uses m=13, n=5, k=7 in FP32 with column-major layout, and the same pattern applies to any small shape.

What does NXM mean in matrix shapes for LIBXSMM?

The README does not define an NXM notation. LIBXSMM's shape APIs take m, n and k separately, and the root Makefile uses MNK to describe generating M,N,K combinations at build time. The README's example is the safest reference for argument order.

How do I perform sparse matrix multiplication with LIBXSMM?

The README states that LIBXSMM covers sparse functionality and sparse matrix operations alongside dense GEMM, and lists samples/xgemm_sparse and samples/xgemm_sparse_Ainregs among the sample directories. It does not give a sparse code example in the README itself, so those sample directories are where the concrete API usage lives.

What is GEMM matrix multiplication in LIBXSMM terms?

GEMM is the general matrix multiply that LIBXSMM dispatches as one of its Tensor Processing Primitives. In the Hello example it computes C = A * B for FP32 column-major matrices, and the README lists supported GEMM datatypes ranging from FP64 and FP32 down to int1 and various low-precision combinations.

Official sources

  1. libxsmm/libxsmm on GitHub
  2. License: BSD-3-Clause
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/libxsmm-libxsmm.svg)](https://hysenlabs.com/projects/libxsmm-libxsmm)