CMSIS-DSP: the performance contract is a list of compiler flags, two of them forbidden
CMSIS-DSP embedded compute library for Cortex-M and Cortex-A
At a glance
- What is it?
- Arm's compute kernels for Cortex-M and Cortex-A, covering filters, transforms, statistics and classical machine learning across seven datatypes, where the documented build options include two you must not use because the library's speed depends on the compiler recognising a memcpy idiom, and where enabling the A-profile vector extension silently breaks the optional C++ layer.
- Who is it for?
- Adopt CMSIS-DSP if you are on Arm hardware and need filters, transforms or fixed-point arithmetic close to the metal, because the kernel coverage across seven datatypes and two core families is the reason it exists, and the documentation tells you exactly which compiler options make it fast.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 20 days ago.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.
Editorial analysis
DSP is in the name for legacy reasons
The project says so in the second line of its about section, and it is the key to reading the rest of the repository. CMSIS-DSP is a compute library for embedded systems, and the digital signal processing in the name is historical. What the kernel list actually contains is broader: basic mathematics in real, complex, quaternion and linear algebra forms plus fast math functions, filtering, transforms including fast Fourier transforms, MFCC and the discrete cosine transform, statistics, and classical machine learning including support vector machines and distance functions for clustering. Quaternions and support vector machines in a library named for signal processing tells you the scope has moved. The other dimension is data type, and there are seven of them: three floating point widths including half precision, and three fixed-point widths plus a seven-bit one. That is a lot of instantiations of every kernel, and it is the reason the library is large and the reason a fixed-point variant exists for a part with no floating point unit at all. The architecture story is the third axis. Kernels are provided for both core families, and most functions have a vectorised variant used when the Helium extension is available on the microcontroller profile or the Neon extension on the application profile. So the same algorithm exists in several places along three axes, and choosing the right instantiation is part of using the library rather than an afterthought.
Two flags are required, two are forbidden, and the reason is a memcpy idiom
The build section is unusual in stating a performance contract explicitly, and the forbidden list is the interesting half. The required options are straightforward: aggressive optimisation and fast math, strongly advised when the vector extension is in use, and a note that one compiler family is currently not producing good results on the microcontroller vector extension so a different compiler is recommended. Then there are the options to avoid. One is a flag that stops the compiler treating standard functions as builtins. The other is a freestanding flag, and the reason given is that it enables the first one. The explanation of why is the most concrete sentence in this repository. The library does type punning: it reads a 32-bit word from memory as a pair of 16-bit fixed-point values or a quadruple of 7-bit values, and it does that manipulation through memory copy functions. Compilers are expected to fold those calls away when the length is small, four bytes in the example given. With builtins disabled that folding does not happen, and the stated consequence is a very bad impact on performance. So this is not a style preference, it is a load-bearing idiom, and a build system that adds a no-builtin flag for safety, which some do, will make this library several times slower while every test still passes. A third option is needed by some compilers to declare that unaligned access is in use. The remaining advice is about memory rather than flags: if the data and the constant tables can live in a particular fast region of memory, put them there, and if the core has a cache, enable it.
The vector extension is automatic on one core family and manual on the other, and that breaks the C++ layer
Here is the asymmetry that will cost you an afternoon if nobody warned you. When building with the microcontroller vector extension, support is detected automatically. When building with the application profile's vector extension, it is not, and you must define the option yourself for the C compilation, or turn it on through the CMake option:
-O3 -ffast-math
-fno-builtin
-ffreestandingThe second and third lines in that block are the two options the project tells you not to use, and they are there to be recognisable rather than to be copied. The real finding is further down, in a paragraph that is set in bold. The C++ layer is currently unsupported in builds with that vector macro enabled. On the application profile, enabling the macro to get the vectorised C kernels also affects the C++ layer, whose Neon implementation is incomplete, and the result can be compilation errors with no automatic scalar fallback. Applications that use only the C API are unaffected and can use the vector extension normally. So the contract is precise: on the application profile you can have the vectorised C kernels or the C++ layer, not both. On the microcontroller profile you can have both, because the vector extension there is detected rather than enabled by a macro. A project that adopted the C++ layer and then turned on the vector extension for the C kernels would find its build broken, and the natural reaction, which is to turn the macro off, gives back a large amount of performance on the C path. Decide which side of that trade you are on before you build.
The C++ layer is header-only, optional, and new in the current release
The C++ layer, called the C++ API, exists to do something the C kernels cannot: combine existing kernels into a new algorithm while keeping performance. The mechanism it names is loop fusion, which merges several operations into a single loop so that intermediate results never become temporary arrays and the data is not passed over repeatedly. That is a real and non-obvious win in a memory-bound embedded workload, and it is the kind of thing you would otherwise write by hand. What makes the design easy to adopt is stated carefully and is worth repeating. The layer is entirely optional and header-only, so it requires no separate library build. Applications using only the C API are unaffected, no code or build changes are needed, and the extension adds no code size, memory use or runtime overhead for code that does not use it. That last claim is the important one for a library shipping to constrained devices: a header-only optional layer costs nothing if unused. The limitation is timing. The headers live in their own directory and are included in the CMSIS pack starting with version 1.18.0, which is the release at the top of the current list and the one published on the same day as the last commit. So if you are on an earlier release, the C++ layer is not in your pack yet, and you would be vendoring the headers yourself. There is also a loop-unrolling option exposed through the build system, with a note that it maps to a C-level option and may or may not be needed depending on the compiler, which is a fair description of a tuning knob.
The autodiff extension is experimental and describes itself narrowly
There is a third component, built on top of the C++ layer, and the way it is described tells you how much weight to put on it. It is an automatic differentiation extension, and it is marked experimental in bold. The project is careful about the claim, stating that it is not a new machine learning framework, which is a specific and welcome thing for a firmware library to say. Its stated focus is on-device fine-tuning, using a subset of the existing kernels. So the ambition is bounded: differentiate through operations the library already has, on the device, without pulling in a training framework. That is a reasonable and useful goal, because on-device fine-tuning otherwise means carrying a framework that will not fit in the memory you have. The example directory listing includes a dedicated autodiff example alongside the C++ examples, the build-tool examples and the architecture examples, which tells you how to see it work. The caution is the one the project makes itself. Experimental, on a layer that is new in the current release, on an architecture that has a known vector-extension gap. If you are evaluating, treat it as a demonstration of a direction rather than a component to plan around, and read the introduction page linked from the readme before you model anything on it.
The Python wrapper compiles C on the machine you install it on
The Python package is presented as the easy way to design an algorithm and then move it to a C implementation, with an API close to the C API, NumPy compatibility, fixed-point support, and a claim that it works in Google Colab. The setup file tells you what that last part involves. It builds a native extension module, not a pure Python package. It reads the version out of a Python file with a regular expression, adds the C include directories to the extension, and then inspects the machine architecture and the operating system to decide the compiler flags. On an Apple silicon or other 64-bit Arm machine it sets the vectorisation macro and the architecture flags, on an Intel Mac it does not, and on Windows it compiles with unaligned access support disabled. The build requirements are declared in the project configuration and include a compiler driver and a build system generator as build requirements, and the packaging library has a different minimum depending on the Python version, three separate floors for three version ranges. So installing the Python package compiles this C library on your machine with whatever compiler you have, using the flags that file computes. That is a genuine surprise for something described as a convenience wrapper, and it has two consequences. The performance advice from the previous sections applies to the wrapper, and your notebook in a hosted environment is running a build made there rather than one Arm produced. And the fixed-point and half-precision paths, the ones that matter for embedded work, are the ones most likely to expose the compiler differences the library is sensitive to.
Two package manager manifests, five build systems, two vendored kernel libraries
The build section of the readme lists five ways to build this library, and the spread is itself a signal about who uses it. There is the IDE and pack route for the microcontroller ecosystem, a plain make route, a CMake route, an installable CMake package, a Zephyr route, and a catch-all for any other build system. On top of that there are two port manifests for the C++ package manager, one for the default build and one whose name marks it as the vector-extension configuration. Two manifests for one project means the two architecture targets are not the same build, which is consistent with everything else in the repository: the vector extension is detected on one core family and must be enabled on the other, and the C++ layer is unavailable in one of those configurations. The top-level listing also shows two vendored library directories alongside the library's own source and headers, one named for Arm's Ne10 and one named for the compute library, so some of the optimised kernels arrive from those rather than from this tree. That is a normal arrangement for performance libraries and it also means your build includes code you did not write and are not reviewing. One more file is worth naming. There is a pack descriptor and a script that generates the pack, which is how a library with this many instantiations gets distributed into the IDE-based toolchain, and it means the repository's primary distribution channel is not a package manager at all.
Editorial conclusion
Adopt CMSIS-DSP if you are on Arm hardware and need filters, transforms or fixed-point arithmetic close to the metal, because the kernel coverage across seven datatypes and two core families is the reason it exists, and the documentation tells you exactly which compiler options make it fast. Do not adopt it expecting the optional C++ layer to work everywhere, because the project states that the C++ API is unsupported in builds with the A-profile vector macro enabled and has no automatic scalar fallback, so a Neon build on Cortex-A gives you the C API and a compile error for the C++ API. Two things to check before you build. Read the forbidden options list as seriously as the required one, because the library relies on the compiler folding a small memory copy, and disabling builtins makes the performance collapse rather than merely degrade. And if you install the Python package, remember you are compiling C on the machine you install it on, which means the same flags apply to your laptop and the wrapper inherits whatever compiler happens to be installed there.
Frequently asked questions
What does CMSIS-DSP provide and for which processors?
Optimized compute kernels for Cortex-M and Cortex-A, covering basic mathematics including quaternion and linear algebra, filtering, transforms such as FFT, MFCC and DCT, statistics, and classical machine learning such as support vector machines. Kernels are provided for f64, f32, f16, q31, q15 and q7, with vectorised variants when Helium or Neon is available.
Which compiler options does CMSIS-DSP need for good performance?
Optimisation level 3 and fast math are required, and strongly advised with the Helium vector extension, where the project recommends the Arm compiler because GCC is not currently performing well. You should avoid the no-builtin and freestanding options, because the library relies on the compiler folding small memory copies and disabling builtins removes that optimisation.
Does the CMSIS-DSP C++ API work with Neon enabled?
No. The readme states the C++ API is currently unsupported in builds with the vector macro enabled on the application profile, that its Neon implementation is incomplete and can cause compilation errors, and that there is no automatic scalar fallback. Applications using only the C API are unaffected.
What is DSP++ in CMSIS-DSP?
A higher-level C++ API for combining existing kernels into new algorithms while keeping performance, using loop fusion to merge operations and avoid temporary arrays. It is entirely optional and header-only, needs no separate library build, adds no overhead when unused, and its headers are included in the CMSIS pack starting with version 1.18.0.
Does the CMSIS-DSP Python package need a compiler?
Yes. The setup file builds a native extension, inspects the machine architecture and platform to choose compiler flags, disables unaligned access support on Windows, and the build requirements include a compiler driver, a build system generator and a minimum version of the packaging library that varies by Python version.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/arm-software-cmsis-dsp)