# ARM Compute Library: Low-Level Arm CPU and GPU Kernels for CV and ML

> The Compute Library is a C++ collection of over 100 machine learning and computer vision functions tuned for Arm Cortex-A, Neoverse and Mali GPUs, released under MIT and updated through v53.3.0 on 2026-09-09. It is a kernel library, not a framework, and that distinction decides whether you want it.

**ARM-software/ComputeLibrary** — The Compute Library is a set of computer vision and machine learning functions optimised for both Arm CPUs and GPUs using SIMD technologies.

- Repository: https://github.com/ARM-software/ComputeLibrary
- Stars: 3,194 · Forks: 820
- Language: C++
- License: not declared
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/arm-software-computelibrary

## What the Compute Library actually is, and who it is for

The Compute Library is a set of low-level machine learning and computer vision functions optimised for Arm Cortex-A, Neoverse and Mali GPU architectures. The README describes it as a collection of functions, and that wording is accurate: this is a kernel library, not an inference framework. There is no model file format, no graph optimiser that decides which device runs which operator, and no Python binding described in the README.

The audience follows from that. You are writing C++ and you want direct control over which primitive runs where. Typical cases are embedded or edge inference on Cortex-A and Cortex-R, server-side Arm work on Neoverse, and GPU compute on Mali through OpenCL. If your job is to load a trained model and get predictions with minimal code, this library gives you the pieces but not the assembly.

The README claims superior performance to other open source alternatives and immediate support for new Arm technologies such as SVE2. Those are the project's own claims. What is verifiable from the repository is the breadth: over 100 functions, multiple convolution algorithms, and datatype support spanning FP32, FP16, INT8, UINT8 and BFLOAT16.

## How the CPU and GPU kernels are organised

The repository separates public headers from implementation: arm_compute/ holds the interfaces, src/ the kernels, and include/ the headers consumed by examples and tests. The examples directory is the clearest map of what the library can do end to end. It contains graph-level samples for AlexNet, GoogLeNet, Inception v3 and v4, MobileNet, MobileNet v2, ResNet50, ResNeXt50, SqueezeNet, ShuffleNet and others, plus a DeepSpeech v0.4.1 sample and an EDSR super-resolution sample.

Two execution models coexist. The single-kernel path uses the low-level functions directly, which is what examples/cl_sgemm.cpp and examples/cl_cache.cpp exercise. The graph path builds a network of nodes and lets the library manage the sequence, which is what the graph_*.cpp examples do. The graph API is a convenience layer over the same kernels, not a different engine.

On the CPU side, optimisation targets Neon on Cortex-A and Cortex-X1, the Neoverse family, and Cortex-R with Armv8-R AArch64. On the GPU side, Mali-G and Mali-T are supported through OpenCL. The README lists several convolution strategies: GeMM, Winograd, FFT, Direct and indirect-GeMM. Choosing among them is a real decision with accuracy and memory consequences, and the library does not make it for you.

There are tuning mechanisms rather than fixed defaults. An OpenCL tuner and GeMM-optimised heuristics adjust behaviour per device and per workload. Kernel fusion and fast-math enablement are listed as optimisation techniques, which means results can differ from a strict reference implementation depending on how you configure the build and runtime.

## Building the Compute Library and running a first graph

The README points to pre-built binaries on the releases page and to the documentation for the build guide. Bazel and CMake builds are flagged as experimental CPU-only builds, so the supported path for GPU work is the SCons build described in the how-to-build documentation. The repository root carries SConstruct and SConscript, which is the entry point for that build, and the build is invoked with scons from that directory.

```bash
scons
```

The README does not reproduce the full option set, so flag names such as the backend selection and example compilation belong to the how-to-build page rather than this article. Read that page before running the command, because the defaults do not necessarily enable the GPU backend or the examples.

With examples enabled, the build produces the graph samples listed in examples/, and the README's linked tutorial on running AlexNet on a Raspberry Pi with Compute Library is the worked end-to-end walkthrough the project itself points to. Start there rather than with a custom graph.

The README notes that pre-built binaries are compiled with a long list of warning and hardening flags, including -Wall, -Wextra, -fstack-protector-strong and -pedantic. If you build from source, matching that flag set is on you.

One naming change is worth planning around. The README states that libarm_compute-static.a will be renamed to libarm_compute.a in distributed pre-built binaries, effective from the first release in or after January 2026. Any build script that hardcodes the old static library name needs updating against a release that carries the new name.

## Where the Compute Library is the wrong tool

The most common mismatch is expecting a framework. There is no documented model importer, no ONNX or TFLite front end in the README, and no automatic operator scheduling across CPU and GPU. You assemble the graph yourself or call kernels yourself. Teams used to writing Python and getting a working pipeline in an afternoon will find the setup cost here measured in build configuration and C++ integration, not in lines of model code.

Portability is narrower than the function count suggests. Supported systems include Android, Linux, macOS, Tizen and bare metal, with OpenBSD, QNX and FreeBSD marked experimental. The x86 entry in the supported architectures list exists, but the optimisation story in the README is entirely about Arm CPUs and Mali GPUs. If your deployment target is x86, you are outside the design centre of this project and should compare against libraries built for that platform.

GPU support is OpenCL-specific. There is no Vulkan, Metal or CUDA path documented. A team on Apple silicon GPUs or on non-Mali mobile GPUs gets the CPU kernels and not the acceleration they may have assumed.

The convolution algorithm choice is another trap. Winograd and FFT reduce arithmetic at the cost of numerical behaviour that differs from direct convolution. The README lists the algorithms but does not tell you which to pick for a given tolerance. That decision requires your own accuracy testing.

## How it differs from Arm NN and from oneDNN

Arm NN is the natural comparison and the two are often confused. Arm NN is a framework layer: it takes a model, builds a graph, and dispatches work to backends, with the Compute Library serving as one of those backends. The Compute Library itself sits below that line. If you want a model-level API, Arm NN is the layer that provides it; if you want the kernels and the ability to place them yourself, the Compute Library is what you call directly. Choosing between them is really choosing how much of the graph machinery you want to own.

oneDNN occupies a similar position for other architectures: a library of optimised primitives for deep learning, with its own set of supported hardware. The difference in approach is the target. oneDNN's optimisation work is aimed at Intel hardware; the Compute Library's is aimed at Arm Cortex, Neoverse and Mali. If your fleet is Arm, the primitives here are written for your instruction sets, including SVE2 support that the README calls out explicitly.

KleidiAI appears in the README's third-party section as an Apache 2.0 library used by this project, and KleidiCV appears in related searches alongside it. Both belong to the same Arm software family. The README states only that KleidiAI code is included, so treat the relationship as a dependency, not as an alternative.

For DSP-oriented embedded work, CMSIS-DSP is a different kind of answer: signal processing primitives for Cortex-M class devices rather than neural network and computer vision functions for Cortex-A and Neoverse. The overlap is smaller than the topic lists suggest.

## Licence, third-party code and what maintenance costs

The README states the software is provided under the MIT license, and contributions are accepted under the same license. The repository also carries a LICENSES/ directory and a REUSE.toml, which is the machine-readable form of that licensing information. Note that the GitHub repository metadata does not report a licence, so the README and the LICENSES directory are the sources to trust here.

MIT is permissive, but this project is not solely MIT. The README lists bundled code with other terms: the OpenCL header library under Apache 2.0, the half library and libnpy under MIT, stb under MIT or public domain, KleidiAI under Apache 2.0, and GoogleTest and Benchmark used by KleidiAI under BSD-3-Clause and Apache 2.0 respectively. The original licence text is included in those source files. If your organisation runs licence scanning, expect a mixed report and check the Apache 2.0 components specifically. This is a description of what the repository says, not legal advice.

The release cadence is visible: v53.1.0 on 2026-05-20, v53.2.0 on 2026-07-01, and v53.3.0 on 2026-09-09. The last push to the default branch was on 2026-09-09. That means roughly every six to eight weeks, which sets the upgrade cost. Each release can change kernel behaviour, tuning heuristics, or artefact names, as the static library rename shows. Pinning a version and testing before moving is cheaper than tracking main.

Contributions require a Developer Certificate of Origin sign-off on every commit, with a real name. The README gives the required trailer form for that sign-off:

```
Signed-off-by: John Doe <john.doe@example.org>
```

Technical discussion happens on the acl-dev mailing list hosted by Linaro. If you plan to patch the library for your hardware, that is the channel and the process.

## Conclusion

Adopt it if you are writing C++ inference or CV code that must run on Arm Cortex-A, Neoverse or Mali hardware and you want the kernels rather than a framework on top of them. Do not adopt it if you expect a model loader, an operator graph with automatic placement, or a Python-first workflow; those live in other projects. Before committing, check the build guide for your target, confirm whether you need the pre-built binaries or a source build, and read the errata page for the release you pick.

## FAQ

### What libraries are used for computer vision?

The Compute Library is one option: it provides over 100 computer vision and machine learning functions optimised for Arm CPUs and Mali GPUs, with graph-level examples for networks such as AlexNet, ResNet50 and MobileNet. It is a C++ kernel library, so it is used alongside your own application code rather than as a standalone tool.

### What is Arm Cortex used for in the context of the Compute Library?

The README lists the Arm Cortex-A processor family using Neon technology, the Cortex-R family with Armv8-R AArch64, and the Cortex-X1 as supported CPU targets. The kernels are micro-architecture optimised for these cores, which is why the library is tied to Arm hardware rather than being generic.

### What are libraries in computing, and how does the Compute Library fit that description?

A library is a collection of reusable functions that an application calls rather than a program you run on its own. The Compute Library fits this exactly: it exposes low-level machine learning and computer vision functions through headers in arm_compute/ and include/, and you link it into your own C++ program.

## Sources

- [ARM-software/ComputeLibrary on GitHub](https://github.com/ARM-software/ComputeLibrary)
- [Issues](https://github.com/ARM-software/ComputeLibrary/issues)
- [README](https://github.com/ARM-software/ComputeLibrary/blob/main/README.md)
- [Releases](https://github.com/ARM-software/ComputeLibrary/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/arm-software-computelibrary
