# xsimd: portable C++ SIMD wrappers for SSE, AVX, AVX512, NEON and SVE

> xsimd gives library authors one C++17 API for batch arithmetic across x86, ARM, WebAssembly, PowerPC, RISC-V, LoongArch and IBM Z. The design is sound, the install paths are short, and the real cost is in how you choose the instruction set.

**xtensor-stack/xsimd** — C++ wrappers for SIMD intrinsics and parallelized, optimized mathematical functions (SSE, AVX, AVX512, NEON, SVE, WebAssembly, VSX, RISC-V))

- Repository: https://github.com/xtensor-stack/xsimd
- Website: https://xsimd.readthedocs.io/
- Stars: 2,758 · Forks: 308
- Language: C++
- License: BSD-3-Clause
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/xtensor-stack-xsimd

## The problem xsimd solves for library authors

SIMD instructions perform one operation on a batch of values, and the README states plainly that they "differ between microprocessor vendors and compilers." That is the whole motivation. If you write intrinsics directly, you write them once per architecture, and every new target is a new code path. xsimd's pitch is narrower than "make my code fast": it is a unified means for using these features for library authors. The audience is people shipping a numerical library, not people tuning one binary for one machine. The README lists the adopters, and the list tells you what kind of code this is: Mozilla Firefox, Apache Arrow, Pandas, KDE Krita, Meta Velox, Milvus, Pythran. Those are projects that need one source tree to build well on x86 laptops, ARM servers and WebAssembly. If that is not your situation, the abstraction is overhead you did not ask for.

## How the batch type maps onto real instructions

The core type is xsimd::batch, parameterized by element type and by an instruction set tag. In the README example, xs::batch<double, xs::avx2> holds four doubles and supports the same arithmetic operators as a single double: (a + b) / 2 compiles to vector add and vector divide. The tag is not decoration. It selects which intrinsic family the wrapper expands to, and the README is explicit that you must enable the matching extension in the build (the -mavx flag on gcc and clang, /arch:AVX on MSVC). That is the central trade-off of the design: you get portability of source, not portability of binary. The instruction set is fixed at compile time by the type you name, so one translation unit is one target. The README also notes that version 8 was a complete rewrite with "slight differences" from 7.x and that a migration guide "will be available soon," which is a real gap if you are upgrading an existing codebase rather than starting fresh.

## Installing xsimd and running a first batch

Three install routes are documented. The conda-forge package is the shortest path on a machine that already has mamba or conda:

```bash
mamba install -c conda-forge xsimd
```

Spack users get a package too, with a separate load step:

```bash
spack install xsimd
spack load xsimd
```

From sources, the README gives a plain CMake invocation with an install prefix, followed by make install. Note that the README writes the CMake call with a trailing dot, so run it from the directory that holds CMakeLists.txt:

```bash
cmake -D CMAKE_INSTALL_PREFIX=your_install_prefix .
make install
```

With the headers in place, the first real use is the mean of two batches. This is the README's own example, with the AVX2 tag:

```cpp
#include <iostream>
#include "xsimd/xsimd.hpp"

namespace xs = xsimd;

int main(int argc, char* argv[])
{
    xs::batch<double, xs::avx2> a = {1.5, 2.5, 3.5, 4.5};
    xs::batch<double, xs::avx2> b = {2.5, 3.5, 4.5, 5.5};
    auto mean = (a + b) / 2;
    std::cout << mean << std::endl;
    return 0;
}
```

Compile it with the extension enabled, for example -mavx on gcc or clang, or /arch:AVX on MSVC. The README states the program prints (2.0, 3.0, 4.0, 5.0). If you build the same file without the matching flag, the tag and the compiler target disagree, and the README does not describe what happens in that case.

## Fixed instruction set tags versus runtime dispatch

The README example hardcodes xs::avx2, which is the easiest thing to demonstrate and the wrong thing to ship to a heterogeneous fleet. A binary built for AVX2 will not run on a machine without it, and the README does not document a dispatch layer in the text available here. The topics list includes vectorization and the CI matrix includes cross-compilation workflows for RISC-V and SVE, so the project clearly cares about many targets, but caring about targets and handing you a runtime dispatcher are different things. Treat this as the main thing to verify before you commit: if your product ships one binary to unknown CPUs, confirm how xsimd expects you to select a batch type at runtime, and do not assume the compile-time tag is enough. The related search phrase for runtime dispatch exists because this is the question people hit first.

## Where xsimd is the wrong tool

If your loop is a simple elementwise pass over a contiguous array, a modern compiler will often vectorize it without a library, and adding xsimd buys you a dependency and a build flag for nothing. xsimd pays off when you need the accelerated mathematical functions operating on batches, or when you need the same source to compile against NEON, SVE, WASM, VSX, RISC-V and LSX/LASX rather than one of them. The second limitation is the dependency table: xtl is optional, and the README states it is required only if you want vectorization for xtl::xcomplex. If you do not use that type, you can skip it. The third is the migration story. The README says 8.x is a complete rewrite, and the migration guide is promised rather than present, so a 7.x codebase faces reading the version 7 and version 8 examples side by side.

## How xsimd differs from SIMDe and from xtensor

SIMDe, which appears in the related searches, takes the opposite approach: it implements the intrinsics themselves so that code written against one vendor's headers compiles on another. You keep writing _mm256_add_ps and SIMDe translates it. xsimd asks you to stop writing intrinsics and use batch types instead. The difference matters at the boundary: with SIMDe your source is still platform-specific intrinsics, with xsimd your source is portable but the instruction set is a template parameter. xtensor is a different comparison again, and it is not a competitor. The README lists xsimd under the same xtensor-stack organization, and the adoption note says "Beyond Xtensor, Xsimd has been adopted by...", which positions xtensor as the multidimensional container layer that consumes xsimd for its expression evaluation. If you want arrays and lazy expressions, you want xtensor. If you want the batch primitive underneath, you want xsimd.

## Maintenance, licence and what an upgrade costs

The repository is not archived, and the last push was on 2026-09-22, six days before this writing, so there is no reason to describe it as dormant. The licence is BSD-3-Clause, which is permissive and, unlike a copyleft licence, does not require you to publish your own source; the LICENSE file is at the repository root if you need the exact wording, and this is not legal advice. The upgrade cost is where I would push back on the project's own communication. The README calls 8.x a complete rewrite and says a migration guide "will be available soon." For a library whose adopters include Firefox and Arrow, that sentence has been true for a while, and it means the practical migration path is reading two sets of examples. The CI matrix is broad (android, cross-rvv, cross-sve, cross, cxx-no-exceptions, cxx-versions, emscripten, linux, macos, windows), so the tested surface is wide, but the tested compilers are a floor, not a suggestion: MSVC 2022 and above, g++ 10 and above, clang 16 and above.

## Conclusion

Adopt xsimd if you are writing a C++17 library that needs one batch API across x86, ARM and WebAssembly, and you are willing to pick the instruction set per translation unit. Do not adopt it if you only care about one architecture and your compiler already auto-vectorizes your loops, or if you need dynamic dispatch that the documentation does not spell out. Before committing, verify three things: that your compiler meets the CI-tested minimums (MSVC 2022, g++ 10, clang 16), that the instruction set you name in the batch type matches the flag you pass to the build, and whether you need the optional xtl dependency for xtl::xcomplex vectorization.

## FAQ

### Is SIMD worth it?

The README frames SIMD as a way to "significantly accelerate code execution" by performing a single operation on a batch of values, and xsimd exists to make that available across vendors and compilers rather than in one vendor's intrinsics.

### When should you use SIMD?

The README positions xsimd for library authors who need batch arithmetic and accelerated mathematical functions while targeting more than one instruction set, listing adopters such as Apache Arrow, Pandas and Mozilla Firefox.

### How does SIMD work?

The README states that SIMD instructions perform a single operation on a batch of values at once. In xsimd that batch is the xsimd::batch type, which takes an element type and an instruction set tag and supports the same arithmetic operators as a scalar.

### What are the disadvantages of SIMD?

The README notes that SIMD instructions differ between microprocessor vendors and compilers, which is the portability problem xsimd wraps; it also requires a C++17 compiler and, for an explicit instruction set tag, the matching build flag such as -mavx.

## Sources

- [Issues](https://github.com/xtensor-stack/xsimd/issues)
- [License: BSD-3-Clause](https://github.com/xtensor-stack/xsimd/blob/master/LICENSE)
- [Project website](https://xsimd.readthedocs.io/)
- [README](https://github.com/xtensor-stack/xsimd/blob/master/README.md)
- [xtensor-stack/xsimd on GitHub](https://github.com/xtensor-stack/xsimd)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/xtensor-stack-xsimd
