Library / SDK
AnswerDotAI/gpu.cpp avatar
AnswerDotAI/gpu.cpp

gpu.cpp: WebGPU compute from a header-only C++20 interface

A lightweight library for portable low-level GPU computation using WebGPU.

3,986 stars199 forksC++Apache-2.0

At a glance

What is it?
gpu.cpp wraps Dawn so C++ and Python code can dispatch WGSL or SPIR-V kernels across Metal, Vulkan and the browser. It is small, explicit about pipeline layouts, and still alpha.
Who is it for?
Adopt gpu.cpp if you want WebGPU compute from C++ or Python without writing Dawn resource management yourself, and you accept an alpha-status dependency pinned to a Dawn revision. Do not adopt it if you need CUDA-specific features, arbitrary Vulkan SPIR-V passthrough, or a stable ABI.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 40 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap gpu.cpp fills between raw WebGPU and a full framework

Writing GPU compute against Dawn directly means owning the instance, adapter, device and queue, wiring descriptor structs for every pipeline and bind group, and deciding what happens when an asynchronous operation fails. gpu.cpp takes that ownership into a Context object and removes the repetitive descriptor code, while leaving shader resource declarations and the execution model explicit. That is the design line the README draws: the library handles native resources, asynchronous work and errors, but shaders still say what they bind and how they run.

The audience is narrow on purpose. This is for C++20 developers who want portable compute across Metal on macOS, Vulkan on Linux, and Emdawnwebgpu in browsers, and who are willing to write WGSL or a WebGPU-profile SPIR-V module. It is not a tensor framework and it does not ship autograd, model loading or scheduling. If your problem is matrix multiplication kernels you want to dispatch yourself, the abstraction level fits. If your problem is training a transformer, this is a layer you would build on, not a layer that replaces one.

Context, Tensor, Kernel: how a dispatch actually flows

The API is four ownership types. Context owns the instance, native adapter, device, queue and error state. Tensor owns a WebGPU buffer plus its shape and numeric type. WGSL and SPIRV own shader source and entry-point metadata. Kernel owns a reusable compute pipeline, a bind group, and an optional uniform parameter buffer.

The flow is: create a context, upload host data into tensors, build a kernel from shader source plus explicit bindings, dispatch, wait, then read back. dispatchKernel() and toCPU() return gpu::Future. wait() pumps Dawn events on native targets or suspends through JSPI in a browser, and it propagates asynchronous failures rather than swallowing them. toGPU() pushes new data into an existing tensor or into a kernel's parameter buffer, which is the path for reusing a pipeline across iterations.

The bindings are explicit for a reason the README states plainly: WebGPU pipeline layouts must agree with the shader declarations. read() and readWrite() are not decoration, they are the contract, and Dawn validates that contract when createKernel() builds the pipeline. If you get a binding wrong you find out at pipeline creation, not at dispatch. All WebGPU handles go through Dawn's generated wgpu C++ RAII facade, so gpu.cpp does not keep a parallel resource pool or release C handles by hand.

Building gpu.cpp and running a first kernel

The README lists CMake, Ninja, Python 3 and a C++20 compiler as prerequisites. Linux also needs a Vulkan driver, and Mesa's Vulkan driver is described as sufficient for development and CI. The first setup step builds the pinned Dawn revision and stages its headers, library and spirv-as under third_party/dawn/.

bash
./test --rebuild-dawn  # first setup, or after changing the pinned Dawn revision

After that, ./test configures, builds and runs the native GPU stories, and the hello-world example can be run directly once built.

bash
./test
./build/hello_gpu

The README's example program doubles a vector of int32 values. The shader is WGSL with a templated workgroup size, and the host side creates tensors, builds a kernel with read() and readWrite() bindings, dispatches, waits, and downloads the result into a std::vector.

cpp
#include "gpu.hpp"

using namespace gpu;

int main() {
  auto context = createContext();
  std::vector<int32_t> input{1, 2, 3, 4};
  std::vector<int32_t> output(input.size());
  auto gpuInput = createTensor(context, {input.size()}, ki32, input);
  auto gpuOutput = createTensor(context, {output.size()}, ki32);
  auto kernel = createKernel(context, WGSL{twice, 4},
                             Bindings{read(gpuInput), readWrite(gpuOutput)});
  auto dispatched = dispatchKernel(context, kernel);
  wait(context, dispatched);
  auto downloaded = toCPU(context, gpuOutput, output);
  wait(context, downloaded);
}

To consume it from an existing CMake project, add the repository as a subdirectory and link the interface target.

cmake
add_subdirectory(path/to/gpu.cpp)
target_link_libraries(my_program PRIVATE gpucpp)

There is also a Python path. The README says to install the self-contained macOS ARM64 wheel from PyPI, and the optional pybind11 module is built by default when gpu.cpp is the top-level CMake project.

bash
pip install gpu-cpp

The Python binding accepts C-contiguous NumPy arrays and preserves tensor shape and float16, float32 or int32 dtype information. Futures own their context and work both in blocking scripts and under native asyncio or in a notebook. Set GPUCPP_BUILD_PYTHON=OFF when embedding gpu.cpp in a CMake project that does not need the module. The pyproject.toml classifiers mark the project as Development Status :: 3 - Alpha, which is worth taking literally.

Where gpu.cpp stops: adapters, SPIR-V scope and browser limits

Native tests must have access to a real Metal or Vulkan adapter. gpu.cpp disables Dawn's Null backend, so an inaccessible GPU fails clearly instead of producing plausible-looking no-op executions. That is the right default for correctness, and it is also a hard constraint: headless CI machines without a GPU driver cannot run the native stories, and a container that silently lacks a Vulkan driver will fail rather than fall back.

SPIR-V support is narrower than the name suggests. It is an opt-in Dawn instance feature, enabled with createContext({.enableSPIRV = true}), and the input must satisfy Dawn and Tint's WebGPU SPIR-V profile. The README states that the current pinned Dawn accepts SPIR-V 1.3, with tests/write42.spvasm as the executable contract fixture, and calls this a WebGPU-facing compiler contract rather than arbitrary Vulkan SPIR-V passthrough. If you have a Vulkan shader that uses features outside that profile, gpu.cpp will not accept it. It is also native-only: the README says SPIR-V is intentionally not part of the browser build.

Browser support comes with its own set of conditions. It requires Chrome with WebGPU and JSPI, plus a sibling ../emsdk clone, and the build moves that clone to the exact emsdk revision pinned by Dawn. Embind functions that call createContext(), wait(), or browser-facing helpers containing them must use emscripten::async() so JSPI can suspend the Wasm stack. That is a real constraint on how you structure the boundary between C++ and JavaScript, not a footnote.

gpu.cpp versus writing Dawn directly

The honest alternative is Dawn itself. gpu.cpp is a thin layer over it, so the difference is not capability but ergonomics and control. With Dawn directly you manage wgpu::Instance, adapter request, device request, queue submission, buffer mapping and error scopes yourself, and you get to decide how resources are pooled, how work is scheduled across queues, and how errors surface. You also inherit the full surface area of the API, including anything gpu.cpp has not wrapped.

With gpu.cpp you get one Context, RAII handles, Future-based dispatch and readback, and a Kernel type that caches a pipeline and bind group for reuse. The trade is that you accept gpu.cpp's opinions about how those things are shaped. If you need a resource pool with a custom lifetime policy, or you want to interleave gpu.cpp dispatches with your own raw wgpu calls in a way the library does not anticipate, you are working against the abstraction rather than with it. The README's note that gpu.cpp does not maintain parallel resource pools is a design statement: there is exactly one owner, and it is the Context.

The other realistic comparison is to CUDA or HIP, and there the difference is not ergonomics but reach. CUDA gives you a mature toolchain and vendor-specific features, and it runs on one vendor's hardware. gpu.cpp targets the WebGPU abstraction and runs on Metal, Vulkan and the browser from one code path. If your deployment is a single NVIDIA machine and you need cooperative groups or vendor intrinsics, CUDA is the correct answer and gpu.cpp is the wrong tool.

Maintenance cost, the pinned Dawn revision, and the licence

The repository is not archived, and the last push was on 2026-08-21. Release v0.2.0 is dated 2026-08-21, and the previous release, 0.1.0, is dated 2024-08-13. That gap is worth noting when you plan an upgrade: the version number moved by one minor step across roughly two years, and the pyproject.toml still carries Development Status :: 3 - Alpha.

The upgrade cost concentrates in one place. tools/build_dawn.sh checks out exact revisions, builds a monolithic shared Dawn library, and stages its headers, library and spirv-as under third_party/dawn/. Changing the pinned Dawn revision means rerunning ./test --rebuild-dawn, and the README recommends that after any change to the pin. Because SPIR-V acceptance is tied to the pinned Dawn, a Dawn bump can change which SPIR-V modules compile. DEV.md is the file the README points to for the dependency layout and update process, and CHANGELOG.md carries release notes.

Licensing is Apache-2.0, declared in both the repository LICENSE and the pyproject.toml license field. That is permissive and includes an explicit patent grant, which matters for a graphics stack. Dawn itself is a separate project with its own terms, and gpu.cpp stages a build of it under third_party/dawn/, so your distribution has to account for both. That is a question for your own counsel, not something the README resolves.

Editorial conclusion

Adopt gpu.cpp if you want WebGPU compute from C++ or Python without writing Dawn resource management yourself, and you accept an alpha-status dependency pinned to a Dawn revision. Do not adopt it if you need CUDA-specific features, arbitrary Vulkan SPIR-V passthrough, or a stable ABI. Before committing, verify that your target machine exposes a real Metal or Vulkan adapter (the Null backend is disabled by design), confirm the pinned Dawn revision in tools/build_dawn.sh matches what your project can build, and check whether the macOS ARM64 wheel covers your platform or whether you must build from source.

Frequently asked questions

Does gpu.cpp run on macOS and Linux, or only in the browser?

All three. The README states that gpu.cpp uses Metal on macOS, Vulkan on Linux, and Emdawnwebgpu in browsers, from the same gpucpp CMake target. Browser builds require Chrome with WebGPU and JSPI plus a sibling emsdk clone.

Do I need to write WGSL, or can I use SPIR-V with gpu.cpp?

Both are supported, but SPIR-V is opt-in through createContext({.enableSPIRV = true}) and must satisfy Dawn and Tint's WebGPU SPIR-V profile. The README says the current pinned Dawn accepts SPIR-V 1.3 and describes this as a WebGPU-facing compiler contract, not arbitrary Vulkan SPIR-V passthrough. SPIR-V is native-only.

How do I install the Python package for gpu.cpp?

The README says to install the self-contained macOS ARM64 wheel from PyPI with pip install gpu-cpp. The optional pybind11 module is also built by default when gpu.cpp is the top-level CMake project, and GPUCPP_BUILD_PYTHON=OFF disables it.

Why does a gpu.cpp test fail instead of falling back to a software device?

gpu.cpp disables Dawn's Null backend deliberately, so an inaccessible GPU fails clearly rather than producing plausible-looking no-op executions. Native tests therefore need access to a real Metal or Vulkan adapter.

Official sources

  1. AnswerDotAI/gpu.cpp on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/answerdotai-gpu-cpp.svg)](https://hysenlabs.com/projects/answerdotai-gpu-cpp)