# ArrayFire: a GPU tensor library for C++ that hides the CUDA and OpenCL plumbing

> ArrayFire is a BSD-3-Clause tensor library that runs the same C++ code on CUDA, oneAPI, OpenCL or the CPU. It suits teams that want GPU speed without writing kernels, and it stops being the right tool once you need custom device code.

**arrayfire/arrayfire** — ArrayFire: a general purpose GPU library.

- Repository: https://github.com/arrayfire/arrayfire
- Website: https://arrayfire.com
- Stars: 4,904 · Forks: 555
- Language: C++
- License: BSD-3-Clause
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/arrayfire-arrayfire

## The portability problem ArrayFire solves for C++ numerics

Writing fast numeric C++ usually means picking a device API first. A CUDA version of your code does not run on an AMD card, an OpenCL version behaves differently across vendors, and a CPU fallback is a third implementation you now have to keep in sync. ArrayFire's pitch is that you write against one abstraction and the library picks the device underneath. The README describes it as a general-purpose tensor library that "simplifies the software development process for the parallel architectures found in CPUs, GPUs, and other hardware acceleration devices."

The audience is narrower than that sentence suggests. This is a library for people who already write C++ (or use one of the language bindings) and who have array-shaped work: linear algebra, image processing, signal processing, statistics, computer vision. The README lists those function areas explicitly, along with machine learning and vector algorithms. If your problem is a loop with pointer chasing and branches, ArrayFire will not help you, because there is no array operation to map it onto.

## How af::array and the JIT turn C++ into device kernels

The central object is `af::array`. The README states that developers write code operating on ArrayFire arrays, "which, in turn, are automatically translated into near-optimal kernels that execute on the computational device." That translation is the whole product. You are not calling cuBLAS or clBLAS by hand; you are building an expression graph and the backend lowers it.

Two consequences follow from that design. First, chained element-wise operations can be fused, so a sequence of small expressions does not necessarily cost a kernel launch each. Second, and more important for debugging, the timing of work is not obvious from reading the source: an expression may not execute where you wrote it. Anyone coming from hand-written CUDA has to unlearn the habit of reasoning about execution order line by line.

The repository layout reflects the scope. `src/` holds the core, `include/` the public headers, `test/` the test suite, and `extern/` the bundled third-party dependencies. `CMakeModules/` and `CMakePresets.json` at the top level are the build system, and `conanfile.py` plus `vcpkg.json` exist for package-manager builds. The README also points to an `examples/` directory, which is split by domain: `lin_algebra/`, `image_processing/`, `machine_learning/`, `computer_vision/`, `pde/`, `financial/`, `graphics/`, `helloworld/` and `unified/`.

Backends are the other half of the architecture. The README claims cross-platform support for CUDA, oneAPI, OpenCL and native CPU on Windows, Mac and Linux. Two further backends, Metal for macOS and Apple silicon, and HIP for AMD ROCm, live on `experimental/metal` and `experimental/hip`. The README is direct about their status: they are community contributed, they build from source only, and they are not covered by the release test matrix. Pull requests for them go to their branch, not to `master`. Treat those as research branches, not as supported targets.

## Installing ArrayFire and running the Game of Life example

The README does not carry install commands. It says that instructions to install ArrayFire or to build it from source can be found on the wiki, and the homepage is arrayfire.com. So the first step is the wiki, not a package command copied from this article.

Once the library is available, the README's own first example is Conway's Game of Life, which is a compact demonstration of the array style. It builds a 3x3 convolution kernel on the host, seeds a random state, and loops.

```cpp
static const float h_kernel[] = { 1, 1, 1, 1, 0, 1, 1, 1, 1 };
static const array kernel(3, 3, h_kernel, afHost);

array state = (randu(128, 128, f32) > 0.5).as(f32); // Init state
Window myWindow(256, 256);
while(!myWindow.close()) {
    array nHood = convolve(state, kernel); // Obtain neighbors
    array C0 = (nHood == 2);  // Generate conditions for life
    array C1 = (nHood == 3);
    state = state * C0 + C1;  // Update state
    myWindow.image(state);    // Display
}
```

The `afHost` flag on the kernel constructor is the detail worth noticing: it tells ArrayFire the data is already on the host, so the library moves it to the device rather than assuming it is there. The boolean comparisons produce arrays, and multiplying `state` by `C0` and adding `C1` is how the survival rule is expressed without a branch. The README links the complete source separately.

If you would rather not write C++, the README lists official and community language APIs. C++, Python and Rust are shown without a dagger mark; Julia and Nim carry a dagger that the README defines as community maintained. A second group is labelled "In-Progress Wrappers": .NET, Fortran, Go, Java, Lua, NodeJS, R and Ruby. That distinction matters more than the badge layout suggests. An in-progress wrapper is not a supported surface, and the README gives no completion timeline for any of them.

## Where ArrayFire is the wrong tool

The abstraction is the limitation. ArrayFire gives you hundreds of functions and no documented escape hatch for custom device code in the README. If your kernel does something the library does not expose, you cannot write it inside the ArrayFire expression language; you either restructure the problem into operations that exist, or you drop to the vendor API and lose the portability that made ArrayFire worth adopting. For research on new operators, that is a real ceiling.

The experimental backends carry a second, sharper caveat. Metal and HIP build from source only and sit outside the release test matrix. A team that needs Apple silicon acceleration or AMD ROCm today is taking on a branch that is explicitly not part of releases, and the README offers no statement about when either would be folded in.

Release cadence is a third consideration. The repository shows v3.10.0 in September 2025, following v3.9.0 in August 2023 and v3.8.3 in February 2023. The gap between 3.8.3 and 3.9.0 was roughly six months; the gap before 3.10.0 was about two years. The last push to the repository was on 2026-09-12, so work is happening, but the release history suggests you should not plan around a fast major-version treadmill. Pin a version and read the release notes before moving.

Finally, the README does not document rollback, downgrade paths or backend selection at runtime in the text available here. If your deployment needs to switch backends in production, confirm that mechanism exists in the documentation before you design around it.

## ArrayFire against CUDA and PyTorch

The comparison people actually search for is ArrayFire versus CUDA. They are not competitors in the same layer. CUDA is a programming model: you write kernels, manage streams, and own the memory movement. ArrayFire is a library that sits above that and generates the kernels for you. Choosing CUDA means maximum control and a single-vendor lock-in. Choosing ArrayFire means less control and code that can also target oneAPI, OpenCL and the CPU. The honest framing is that ArrayFire removes the kernel-writing step, and if you wanted that step, you have picked the wrong tool.

Against PyTorch the difference is ecosystem and language, not arithmetic. PyTorch is a Python-first framework with autograd and a large model zoo; ArrayFire is a C++ tensor library with bindings, aimed at embedding acceleration into existing native applications. If your product ships a C++ binary and cannot take a Python runtime dependency, PyTorch's strengths are largely out of reach and ArrayFire's C++ core is the point. If your work is training neural networks in notebooks, the reverse is true, and ArrayFire's machine learning functions are a small part of a broad numeric library rather than a training framework.

For teams searching for an ArrayFire alternative, the realistic substitutes are the vendor libraries directly (cuBLAS, cuFFT and friends) if you only ever target NVIDIA, or a framework binding if you can accept a heavier runtime. Neither gives you the same single-source, multi-backend C++ story, which is the specific thing ArrayFire sells.

## Licence, packaging and the cost of staying current

ArrayFire is BSD-3-Clause, which the README describes as "commercially friendly open-source licensing." That is a permissive licence in the usual sense, and it is a meaningful difference from copyleft numeric stacks for anyone shipping a closed-source product. Two practical notes. The repository has both a `LICENSE` file and a `LICENSES/` directory, and the README asks that anyone redistributing ArrayFire follow the terms established in the license. The bundled dependencies under `extern/` may carry their own terms, so read the directory rather than assuming one licence covers the whole build. This is not legal advice; if redistribution is central to your product, have someone qualified read it.

Upgrade cost is mostly a build-system question. The repository ships `CMakeLists.txt`, `CMakeModules/`, `CMakePresets.json`, `conanfile.py` and `vcpkg.json`, so there are three plausible integration routes: CMake presets, Conan, or vcpkg. Which one is maintained best is not stated in the README. If you depend on a specific backend, the upgrade risk concentrates there rather than in the API, which has been stable enough to carry the same `af::array` style across the 3.8 to 3.10 line.

Support has a commercial side. The README lists enterprise support from ArrayFire, along with consulting, support and training links, and notes that development is funded by AccelerEyes LLC and third parties. Community channels are a Slack invite and a Google Groups forum. For a library used in production, the existence of a paid support path is worth knowing before you standardise on it.

## Conclusion

Adopt ArrayFire if you have C++ numerics that map onto array operations and you want to reach CUDA, oneAPI, OpenCL or the CPU without maintaining four code paths. Do not adopt it if your workload needs custom kernels, because the library gives you no place to put them. Before committing, verify which backends your target machines actually have, check whether the language wrapper you want is listed as community maintained or in progress, and read the LICENSE and LICENSES/ directory yourself, since redistribution terms are your call and not something this article can settle.

## FAQ

### What is array computing in the context of ArrayFire?

It means expressing computation as operations on whole arrays rather than element-by-element loops. ArrayFire's README describes the library as a general-purpose tensor library and exposes an `af::array` object that holds data on the accelerator, so your code issues array operations and the library translates them into device kernels.

### What are array functions in ArrayFire?

The README lists the function areas as array handling, computer vision, image processing, linear algebra, machine learning, standard math, signal processing, statistics and vector algorithms, described as hundreds of accelerated tensor computing functions. The full list is in the function reference linked from the README.

### How does ArrayFire compare with CUDA?

CUDA is the programming model you write kernels in; ArrayFire is a library that generates the kernels for you from array expressions. The trade is control for portability, since ArrayFire's README claims the same code can target CUDA, oneAPI, OpenCL and native CPU.

### How does ArrayFire compare with PyTorch?

ArrayFire is a C++ tensor library with official and community language bindings, while PyTorch is a Python-first framework. The README positions ArrayFire around tensor computing functions across linear algebra, image and signal processing, not around a training workflow, so the choice usually comes down to whether your application is native C++ or Python.

### What are alternatives to ArrayFire?

If you only target NVIDIA hardware, the vendor libraries directly (cuBLAS, cuFFT and similar) remove the abstraction layer. If you can accept a heavier runtime, a framework binding is another route. Neither offers ArrayFire's single-source path across CUDA, oneAPI, OpenCL and CPU.

## Sources

- [arrayfire/arrayfire on GitHub](https://github.com/arrayfire/arrayfire)
- [License: BSD-3-Clause](https://github.com/arrayfire/arrayfire/blob/master/LICENSE)
- [Project website](https://arrayfire.com)
- [README](https://github.com/arrayfire/arrayfire/blob/master/README.md)
- [Releases](https://github.com/arrayfire/arrayfire/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/arrayfire-arrayfire
