# google/highway: Portable SIMD in C++ with Runtime Dispatch

> Highway is a C++17 library that wraps SIMD intrinsics behind one API across seven architectures, with dynamic dispatch to pick the best instruction set at runtime. It is aimed at engineers who want predictable vectorization instead of hoping autovectorization fires.

**google/highway** — Performance-portable, length-agnostic SIMD with runtime dispatch

- Repository: https://github.com/google/highway
- Stars: 5,909 · Forks: 474
- Language: C++
- License: NOASSERTION
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/google-highway

## What Highway solves, and who it is actually for

Writing SIMD by hand means writing the same kernel several times: SSE4.2, AVX2, AVX-512, NEON, SVE, WASM. Each has its own intrinsics, its own header, its own quirks. The README frames Highway as a C++ library providing portable SIMD/vector intrinsics, so the same application code can target various instruction sets, including those with scalable vectors whose size is unknown at compile time.

The audience is narrower than "C++ developers". It is engineers who have a data-parallel kernel that matters, who need it to run on more than one instruction set, and who distrust autovectorization. The README makes that distrust explicit: Highway is described as more predictable and robust to code changes and compiler updates than autovectorization. That is the pitch. You get intrinsics-level control with one source, and you stop re-auditing generated assembly after every compiler upgrade.

The README also lists where the library is already used, and the spread is informative: browsers (Chromium, Firefox), image codecs (JPEG XL, Jpegli, libaom, OpenHTJ2K), image processing (libvips), machine learning (gemma.cpp, TensorFlow, NumPy), cryptography, information retrieval and vector search. If your domain is on that list, you are in well-trodden territory. If it is not, the README says new use-cases may require additional ops and invites an issue, which is a polite way of saying the operation set is not infinite.

## One API, seven architectures, and the dispatch switch that decides everything

The architecture is straightforward once you see the dispatch mechanism. You write your kernel against Highway's types and functions. You compile that same source for several targets. At runtime, one line decides how it is reached.

The README states that application code is the same in both deployment modes except for swapping HWY_STATIC_DISPATCH with HWY_DYNAMIC_DISPATCH plus one line of code. Static dispatch targets a single instruction set with no runtime overhead; dynamic dispatch picks the best available instruction set at runtime, which is what lets one binary run on a heterogeneous cloud or across client devices.

That is the whole design in one sentence, and it explains the trade-off. Dynamic dispatch buys portability and costs you a dispatch step plus the build complexity of compiling the kernel multiple times. Static dispatch buys you zero overhead and costs you the ability to move the binary to a machine with a different instruction set.

The README points at two Compiler Explorer demos that illustrate the split: one for multiple targets with dynamic dispatch, described as more complicated but flexible and using the best available SIMD, and one for a single target using -m flags, described as simpler but requiring and only using the instruction set enabled by compiler flags. Both links are in the README; reading them side by side is the fastest way to understand what you are choosing between.

Highway requires C++17 (language features, not necessarily the library) and supports four families of compilers. The README names seven supported architectures and says that support for other platforms should be raised as an issue.

## Building Highway and running a first dynamic-dispatch kernel

The repository carries build definitions for CMake (CMakeLists.txt, CMakeLists.txt.in, cmake/), Bazel (BUILD, MODULE.bazel, WORKSPACE), and Meson (meson.build, meson_options.txt). pkg-config templates are present at the top level: libhwy.pc.in, libhwy-contrib.pc.in and libhwy-test.pc.in. The README itself does not spell out an install sequence, so the concrete starting point it does give is the example set under hwy/examples, whose own README covers them.

The README groups those examples into three tiers. Basics: sum_array_advanced.cc (4x unrolling and remainder handling), dot_product_unroll.cc (the same plus reduction), and matrix_transpose_scatter_gather (compares Scatter and Gather). Infrastructure: benchmark.cc (dot product with remainder-free loops and benchmarking), profiler_example (the built-in profiler for timing annotated zones), and skeleton* (a complete module with runtime dispatch). Challenges: masks_and_logic.cc and ctf_aes.cc.

For a first real use, the skeleton example is the one that matters, because it demonstrates the full dispatch structure rather than a single kernel. The README describes it as a complete example of a module with runtime dispatch, which is exactly the shape you will copy into your own tree.

The repository also ships a test runner at the top level. Running it is the shortest way to confirm your toolchain builds the targets Highway expects:

```bash
./run_tests.sh
```

What you should see: the example binaries compile for more than one target and the dispatch path resolves at runtime. If your kernel builds only under a static target, you have not yet exercised the part of Highway that justifies the dependency.

## Where Highway is the wrong tool

The README is unusually candid about this, and it is worth taking at face value. It says the biggest gains are unlocked by designing algorithms and data structures for scalable vectors, and names the techniques: batching, structure-of-array layouts, and aligned or padded allocations.

Read that as a constraint, not marketing. If your data lives in array-of-structures form and you cannot change it, Highway gives you Gather, MaskedLoad and FixedTag to soften the blow, but the README presents those as tools for legacy data structures rather than a route to the same speedup. You will get something. You will not get what a restructured layout would give you.

There is a second boundary. The README says new use-cases may require additional ops and that the maintainers are happy to add them where it makes sense, with the parenthetical that they avoid performance cliffs on some architectures. That parenthetical is doing real work: an operation that maps cleanly to AVX-512 may map badly to a narrower target, and the library's stated policy is to decline additions that create that asymmetry. If your kernel depends on an op that is not there yet, the answer may be no, and the reason may be a target you do not care about.

The third boundary is the one people hit last. A single-instruction-set build with -m flags is simpler, per the README's own description of its Compiler Explorer example. If you only ever deploy to one machine class and you are willing to pin compiler flags, dynamic dispatch is overhead you are choosing to pay for portability you will not use.

## Highway against raw intrinsics and autovectorization

The real alternative is not another library. It is writing intrinsics directly, or writing plain C++ and letting the compiler vectorize.

Raw intrinsics give you exactly the instruction you asked for, with no dispatch layer and no build matrix, at the cost of one kernel per architecture and a rewrite every time you add a target. Highway's answer is that you write the kernel once and the dispatch layer handles the rest. The difference in approach is where the per-architecture knowledge lives: in your source, or in the library.

Autovectorization is the opposite trade. You change nothing, and the compiler decides. Highway's README argues the resulting code is more predictable and robust to code changes and compiler updates than autovectorization, which is the strongest claim in the document. Whether that holds for your code is something you can only settle by compiling both versions and comparing, and the README's own Compiler Explorer links exist precisely so you can do that without setting up a project.

There is a middle option worth naming because the README names it: the single-target -m flag path. You get Highway's portable API and a simpler build, and you give up runtime selection. For a lot of internal services this is the correct answer, and it is under-discussed relative to dynamic dispatch.

The README also links an external Evaluation of C++ SIMD Libraries, which it quotes as finding that Highway excelled with a strong performance across multiple SIMD extensions. That is a third-party comparison, not a project benchmark, and it is the right place to look if you want numbers rather than architecture.

## Maintenance, licensing and what upgrades cost you

The repository is not archived and the last push was on 2026-09-22. Releases are not frequent: 1.4.0 landed on 2026-04-23, 1.3.0 on 2025-08-14, and 1.2.0 on 2024-05-31. That cadence suggests a library that is stable rather than fast-moving, and it means an upgrade decision comes up roughly annually rather than monthly.

The upgrade cost is concentrated in one place: your dispatch setup and your build files. A Highway version bump can change the build definitions (CMake, Bazel via MODULE.bazel and WORKSPACE, Meson, plus the pkg-config templates), and if your kernel is compiled once per target, every target is rebuilt and retested. That is the price of the multi-target model, and it is paid at upgrade time, not at runtime.

Licensing has changed and the README states it plainly: previously licensed under Apache 2, now dual-licensed as Apache 2 / BSD-3. The repository's LICENSE file is the authoritative text, and the machine-readable license field on the repository is NOASSERTION, which means the hosting platform has not classified it automatically. If your organization has a license allowlist, check LICENSE directly rather than relying on the repository metadata, and route the dual-license question through whoever handles that for you. This is a description of what the README says, not legal advice.

The README also links a contributor guide (CONTRIBUTING) and invites issues for new platforms and new operations, so the path for adding an architecture is documented in the repository rather than being a fork-only exercise.

## Conclusion

Adopt Highway if you ship one binary across heterogeneous x86, Arm or WASM targets and want the same vector code to run on all of them, or if you need an instruction set that autovectorization will not reliably produce. Do not adopt it if your hot loop is already well served by compiler flags and you are unwilling to restructure data into batches and structure-of-array layouts, because the README is explicit that the largest gains come from redesigning algorithms around scalable vectors. Before committing, verify which of the seven architectures your CI actually exercises, confirm whether you want HWY_DYNAMIC_DISPATCH or the zero-overhead HWY_STATIC_DISPATCH path, and check that the operations your kernel needs exist rather than assuming the set is complete.

## FAQ

### What is google/highway?

It is a C++ library that provides portable SIMD/vector intrinsics, so the same application code can target multiple instruction sets, including those with scalable vectors whose size is unknown at compile time. It requires C++17 language features and supports seven architectures and four compiler families.

### How do I install google/highway?

The repository ships build definitions for CMake, Bazel and Meson, along with pkg-config templates, and the README does not give a single install sequence of its own. The practical starting point it points to is the example set under hwy/examples, whose README covers how to build and run them.

### What is the difference between HWY_STATIC_DISPATCH and HWY_DYNAMIC_DISPATCH in google/highway?

Static dispatch targets a single instruction set with no runtime overhead, while dynamic dispatch chooses the best available instruction set at runtime. The README states the application code is the same in both cases except for swapping the two macros plus one line of code.

### Which platforms does google/highway support?

The README states Highway supports seven architectures and four families of compilers, and that applications can run on heterogeneous clouds or client devices by choosing the best available instruction set at runtime. It adds that support for other platforms should be raised as an issue.

### How is google/highway licensed?

The README states the project was previously licensed under Apache 2 and is now dual-licensed as Apache 2 / BSD-3. The repository's LICENSE file carries the authoritative text, and the repository metadata shows the license as NOASSERTION.

## Sources

- [google/highway on GitHub](https://github.com/google/highway)
- [Issues](https://github.com/google/highway/issues)
- [README](https://github.com/google/highway/blob/master/README.md)
- [Releases](https://github.com/google/highway/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/google-highway
