SIMDe: Run SSE, AVX and NEON Intrinsics on Hardware That Does Not Have Them
Implementations of SIMD instruction sets for systems which don't natively support them.
At a glance
- What is it?
- SIMDe is a header-only C library that implements x86 and ARM SIMD intrinsics in portable code, so a codebase can be ported to a new architecture before anyone rewrites the hot loops. The trade-off is that portability, not speed, is the default.
- Who is it for?
- Adopt SIMDe when you need an x86 or ARM SIMD codebase to build and run on a different architecture today, and you accept that the portable path is a correctness bridge rather than a performance one. Do not adopt it expecting the emulated path to match native throughput, and do not adopt it as a general vector math library: it exposes instruction sets, not algorithms.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly C, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The porting problem SIMDe was written to remove
Code written against SSE, AVX or NEON intrinsics is bound to one instruction set. The intrinsics are not a portable API; they are named after the instructions they compile to, and a compiler for a different architecture has no reason to know them. The usual response is to rewrite the vectorized sections per architecture, or to guard them behind preprocessor branches that multiply with every target added. SIMDe takes a third route: it supplies headers that define the same intrinsic names, implemented in portable C, so the existing source compiles unchanged on a machine that has no such instructions. The intended user is an engineer with a working x86 or ARM codebase who needs it to build and run somewhere else, and who wants that port finished before deciding which loops deserve hand-written native code.
How the headers decide between native and emulated paths
SIMDe is header-only, so there is no library to link and no runtime to initialize. Each header for an instruction set defines the intrinsics for that set, and each definition is selected at compile time. Where the target hardware and compiler support the instruction, SIMDe forwards to the native intrinsic, which is why the README states there is no performance penalty when the hardware supports the native implementation. Where they do not, SIMDe falls back to a portable implementation. That fallback is not a single technique. The README lists several: intrinsics from other instruction sets (using NEON to implement SSE, for example), compiler vector extensions and built-ins such as __builtin_shufflevector and __builtin_convertvector, and auto-vectorization hints via OpenMP 4 SIMD, Cilk Plus, GCC loop-specific pragmas and clang loop hint directives. The design goal stated in the README is to expose the entire functionality of the underlying instruction set rather than the lowest common denominator, which is what separates it from abstraction layers that only cover the intersection of several ISAs. The practical consequence of header-only plus compile-time selection is that the same source file can hold SSE and NEON code side by side, and the build decides which one is real.
Trying SIMDe without installing anything, then adding it to a build
The README points to an online demo on Compiler Explorer using an amalgamated SIMDe header, which is the fastest way to see whether a given intrinsic behaves as expected before touching a build system. For local use, the repository ships an amalgamate.py script at the top level that produces a single combined header from the simde/ directory, which is the same artifact used by the online demo. A minimal compile looks like this, with the include path pointing at the directory containing the generated header:
What the portable path costs you
The README is explicit that the accelerated implementations are a work in progress: the current focus is on writing complete portable implementations, and a large number of functions already have accelerated implementations. That ordering matters. It means the project's guarantee is correctness on any target, not speed on every target. A function that has no accelerated path still compiles and still produces the right result, but the result comes from portable code that the compiler may or may not turn into good vector instructions. The README also states that you will eventually want to test on the actual hardware you are targeting, which is the honest boundary of the approach: running NEON code on an x86 machine without an emulator is useful for development, but it is not a substitute for the target. A second limitation is coverage. The README lists complete implementations for NEON, MMX, SSE, SSE2, SSE3, SSSE3, SSE4.1, CRC32, AVX, AVX2, F16C, FMA, GFNI, XOP, SVML, AVX512VPOPCNTDQ, AVX512_BITALG, AVX512_VBMI and AVX512_VNNI. SSE4.2 is not in that list. If your code calls an intrinsic outside the complete set, the header may still exist, but the README's own status section does not claim completeness for it, and that is the first thing to check before a port.
SIMDe against sse2neon and against portable vector libraries
The closest comparison is sse2neon, which appears among the search phrases people use around this project. sse2neon translates one direction: SSE intrinsics to NEON, for running x86 code on ARM. SIMDe is not a translator with a single source and target. It defines intrinsics for many instruction sets and chooses an implementation per target at compile time, so the same headers can carry SSE, AVX and NEON code, and the native path is used where the hardware has it. That breadth is the difference in approach, and it is also why SIMDe's headers are larger and why its per-function acceleration is uneven: a project with one source ISA and one target ISA can specialize more aggressively than one that must cover both directions across many ISAs. On the other side, libraries like simdjson are not alternatives at all. They are applications built on SIMD for a specific problem, whereas SIMDe is infrastructure that makes other code buildable elsewhere. Choosing between them is not a comparison; it is a question of whether you are writing the vectorized code or consuming someone else's.
Build integration, licence and the cost of upgrading
The repository carries CMakeLists.txt and meson.build at the top level, so both CMake and Meson are supported entry points for building the tests and the amalgamation. The CI configuration is unusually wide: .appveyor.yml, .azure-pipelines.yml, .circleci/, .travis.yml and a .drone.star.disabled file are all present, alongside a docker/ directory and a .packit.yml with a .packit/ directory for downstream packaging. That spread is a signal about the maintenance surface rather than a quality claim. Because SIMDe is header-only, upgrading means replacing headers and recompiling; there is no ABI to match and no shared object to keep in sync. The real upgrade cost is behavioural: a new release can change which path a function takes on your compiler, and the accelerated paths are the ones most likely to change. The release history shows v0.8.4-rc1 in February 2025, v0.8.4-rc2 in September 2025 and v0.8.4-rc3 in February 2026, so the 0.8.4 line has been in release-candidate state across multiple cycles. The repository is not archived and the last push was on 2026-09-23. On licensing, the repository is labelled MIT and ships a COPYING file; SIMDe is a header-only library, so its code is compiled into your binary rather than linked, and the exact terms in COPYING are what you should read before distributing. This is not legal advice.
Editorial conclusion
Adopt SIMDe when you need an x86 or ARM SIMD codebase to build and run on a different architecture today, and you accept that the portable path is a correctness bridge rather than a performance one. Do not adopt it expecting the emulated path to match native throughput, and do not adopt it as a general vector math library: it exposes instruction sets, not algorithms. Before committing, verify that every intrinsic your code uses appears in the README's list of complete implementations, check that your compiler is one of the supported ones for the accelerated paths, and confirm the exact licence text in COPYING rather than trusting the MIT label on the repository.
Frequently asked questions
What is SIMDe used for?
It provides portable implementations of SIMD intrinsics so code written for one instruction set, such as SSE or AVX, compiles and runs on hardware that does not support it, for example NEON code on ARM or x86 code on ARM. The README describes the goal as making ports easier and letting different instruction sets coexist in the same implementation.
How do you pronounce SIMD?
The README does not say how the name is pronounced; it only expands the subject as SIMD intrinsics. The project name SIMDe is written as one word in the repository and its documentation.
What does SIMD mean?
The README links the term to the SIMD article on Wikipedia and describes the library as providing portable implementations of SIMD intrinsics, the functions that map to single-instruction-multiple-data operations on a processor.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/simd-everywhere-simde)