Library / SDK
KomputeProject/kompute avatar
KomputeProject/kompute

Kompute: a Vulkan-based GPU compute framework for C++ and Python

General purpose GPU compute framework built on Vulkan to support 1000s of cross vendor graphics cards (AMD, Qualcomm, NVIDIA & friends). Blazing fast, mobile-enabled, asynchronous and optimized for advanced GPU data processing usecases. Backed by the Linux Foundation.

2,573 stars199 forksC++Apache-2.0

At a glance

What is it?
Kompute wraps Vulkan compute shaders in a manager, tensor and sequence API so the same C++ or Python code runs on AMD, Qualcomm and NVIDIA GPUs. It is backed by the LF AI & Data Foundation, and its last push was on 2026-08-15.
Who is it for?
Adopt Kompute if you already ship Vulkan or need one compute path across AMD, Qualcomm and NVIDIA hardware, and if you can build the Python module through CMake 3.15 or newer. Do not adopt it if you want a turnkey CUDA replacement, a stable API across releases, or GPU debugging that does not require the Vulkan validation layers.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 47 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Kompute solves, and who is expected to use it

Vulkan exposes compute through queues, descriptor sets, command buffers and explicit memory types. That is portable across vendors, and it is also a large amount of boilerplate before a single shader runs. Kompute sits on top of that layer. The README describes it as a general purpose GPU compute framework for cross vendor graphics cards (AMD, Qualcomm, NVIDIA and others), with a Python module for scripting and a C++ SDK for optimisation.

The intended audience is narrow and identifiable. You are writing C++ that already links Vulkan, or you are writing Python and want GPU dispatch without leaving the language. The repository lists projects that use it, including GPT4ALL for on-edge language models and vkJAX, a JAX interpreter for Vulkan. Those are both cases where the vendor-neutral constraint matters more than peak throughput on one vendor.

If your workload is already comfortable in CUDA or a vendor library, Kompute is not competing for that job. Its value is the case where you cannot assume an NVIDIA device: Android via the NDK, Qualcomm GPUs, or desktop machines with AMD cards. The README also markets a bring-your-own-Vulkan design, meaning it can be introduced into an existing Vulkan application rather than replacing its device setup.

Manager, tensors, sequences: the mechanism behind a Kompute dispatch

The C++ example in the README shows the whole data flow in one function. A kp::Manager is constructed with default settings, meaning device 0, the first queue, and no extensions. Tensors are then created through that manager: mgr.tensor({2., 2., 2.}) for a default float tensor, and mgr.tensorT<uint32_t>(...) for an explicit type, with uint32, int32, double, float and bool listed as supported. Each tensor is a kp::Memory, and the vector of shared pointers passed into mgr.algorithm is what the shader binds against.

The algorithm object combines the shader, a workgroup size, spec constants and push constants. In the example the workgroup is {3, 1, 1}, spec constants are {2}, and two push constant vectors are prepared so the same algorithm can be dispatched twice with different values. That is the core abstraction: shader plus binding list plus dispatch geometry.

Execution goes through a sequence. The example records kp::OpSyncDevice, then kp::OpAlgoDispatch, then calls eval(); it then records a second dispatch with overridden push constants and evaluates only that. Asynchronous work uses the same object: evalAsync<kp::OpSyncLocal>(params) starts the transfer back to host memory, evalAwait() blocks until it finishes. The README's stated outputs are {4, 8, 12} and {10, 10, 10}, which is how the two push constant sets differ.

Memory ownership is explicit rather than hidden. The README links to a memory management page and says relationships for GPU and host memory ownership are explicit. In practice that means you decide when data lives on the device and when it comes back, and the manager releases both CPU and GPU resources when it goes out of scope. This is more work than a framework that copies on demand, and it is also why the same API can be used for workloads that would otherwise thrash the bus.

Building Kompute from source and running a first dispatch

The repository ships a setup.py that drives CMake rather than compiling a Python extension directly. Its CMakeBuild class checks for CMake and raises an error if the version is below 3.15, then passes flags including -DKOMPUTE_OPT_BUILD_PYTHON=ON, -DKOMPUTE_OPT_LOG_LEVEL=Off and -DKOMPUTE_OPT_USE_SPDLOG=Off to the build. So the Python module is a build of the C++ SDK, which is worth knowing before you file a bug against the Python layer.

A source install therefore looks like this:

bash
pip install .

That runs setup.py in the repository root and requires CMake 3.15 or newer to be on PATH. The README does not document a published wheel, so expect a compile rather than a download.

Once built, the Python interface follows the same manager and tensor shape as C++. The README points to examples/python_naive_matmul as a worked case, and the C++ listing above is the reference for the object graph. The README gives that example in C++, not Python, so treat any Python names as something to check against the module you built. The C++ side is the part with a full listing, and its first two steps are the manager and the tensors:

bash
make build_linux

The Makefile defines that target, and the Dockerfile builds on nvidia/vulkan:1.1.121, installs g++, copies the repository to /workspace and runs the same make target. If you want to avoid host toolchain setup, that Dockerfile is the shortest path to a working build.

For a first real run, start from examples/array_multiplication or examples/python_naive_matmul rather than writing a shader from scratch. The README notes that shaders can be written as raw strings or compiled to SPIRV and C++ headers with the Kompute tools, and the Makefile references glslangValidator for that step.

Where Kompute is the wrong tool

The most concrete limitation is visible in the Makefile. It carries a FILTER_TESTS variable excluding TestAsyncOperations.TestManagerParallelExecution, TestSequence.SequenceTimestamps and TestPushConstants.TestConstantsDouble, with a comment that these do not work with swiftshader but can be run directly with Vulkan. SwiftShader is the software Vulkan implementation, and it is what you get in CI containers and on machines without a GPU. So a meaningful part of the async, timestamp and double push-constant behaviour is not exercised in a software-only environment. If your CI has no GPU, you are not testing those paths.

Versioning is the second problem. The README badge says 0.7.0 while the most recent release listed is v0.9.0 from 2024-01-20, and the previous release, v0.8.1, dates to 2022-04-13. That is a long gap between tagged releases. The master branch has moved since, with the last push on 2026-08-15, so if you build from source you are tracking master, not a release. Anything you depend on should be pinned to a commit rather than a tag if you need recent behaviour.

Finally, consider what you are taking on. Vulkan compute on a mobile device means driver behaviour varies by vendor and by driver version, and Kompute cannot abstract that away. The framework removes boilerplate, not the underlying variability. If your workload fits a single vendor and you have no portability requirement, the Vulkan layer is cost without benefit.

Kompute compared with CUDA and with raw Vulkan

Against CUDA the difference is not performance, it is the device set. CUDA runs on NVIDIA hardware only. Kompute targets any Vulkan-capable GPU, which the README frames as cross vendor support for AMD, Qualcomm, NVIDIA and others, plus mobile through the Android NDK. If your deployment includes Android phones, Qualcomm-based devices or AMD desktops, CUDA is not an option at all and the comparison ends there. If everything you ship has an NVIDIA GPU, CUDA gives you a mature toolchain, profilers and libraries that Kompute does not attempt to match.

Against raw Vulkan the difference is the amount of code between you and a dispatch. Raw Vulkan requires you to create the instance, pick a physical device, create a logical device and queues, allocate and bind memory, build descriptor set layouts and pipelines, record command buffers and manage synchronisation. Kompute's Manager, tensors and sequences collapse that into a handful of calls, and the README's bring-your-own-Vulkan design means you can pass in an existing device rather than letting Kompute create one. That is the honest trade: you give up direct control of the Vulkan objects in exchange for not writing them. For teams whose expertise is in the shader and the data layout rather than in Vulkan setup, that is a good trade. For teams already fluent in Vulkan with a tuned pipeline, the abstraction is a layer to debug through.

Maintenance, upgrade cost and licence

Kompute is not archived, and the last push was on 2026-08-15. The release history is uneven: v0.8.0 in 2021, v0.8.1 in 2022, then v0.9.0 in January 2024, with master continuing afterwards. A project with that cadence is best treated as source-first. Pin a commit, read CHANGELOG.md before moving, and do not assume a fix from master is in a tagged release.

The repository carries the artefacts of a governed project: GOVERNANCE.md, CODE_OF_CONDUCT.md, CONTRIBUTING.md, SECURITY.md and a CII Best Practices badge, and the README states it is hosted by the LF AI & Data Foundation. That matters for procurement more than for engineering, but it does mean the project has a documented process for reporting security issues and for changes.

Upgrade cost is dominated by the build, not the API. Because the Python module is compiled through setup.py and CMake, a Python version bump or a CMake change can break an install even when the Kompute API is unchanged. The setup.py flags disable the Vulkan version check and spdlog, which reduces surface area but also means the build trusts your Vulkan headers. Budget for a rebuild in CI whenever the toolchain moves.

The licence is Apache-2.0, per the LICENSE file and the badge in the README. Apache-2.0 includes an explicit patent grant and requires you to preserve notices and state changes. It is permissive, so it can be used in closed products, but if you modify Kompute itself, the notice obligations apply. This is a description of the licence, not legal advice; check with your own counsel for anything beyond that.

Editorial conclusion

Adopt Kompute if you already ship Vulkan or need one compute path across AMD, Qualcomm and NVIDIA hardware, and if you can build the Python module through CMake 3.15 or newer. Do not adopt it if you want a turnkey CUDA replacement, a stable API across releases, or GPU debugging that does not require the Vulkan validation layers. Verify first that your target device exposes the queue families and memory types your workload needs, and check the v0.9.0 release notes against the master branch before pinning a version.

Frequently asked questions

What is a GPU compute, and how does Kompute relate to it?

GPU compute means running general-purpose work on the graphics processor rather than drawing pixels. Kompute is a framework that dispatches that work through Vulkan compute shaders, exposing it as a Manager, tensors and sequences in C++ and Python.

How do I install Kompute?

The repository provides setup.py, which drives a CMake build and requires CMake 3.15 or newer. Running pip install . from the repository root builds the Python module from the C++ SDK; the Dockerfile offers an alternative container build using nvidia/vulkan:1.1.121.

Can Kompute run on AMD and Qualcomm GPUs, or only NVIDIA?

The README describes Kompute as a cross vendor framework for AMD, Qualcomm, NVIDIA and others, and it uses Vulkan rather than a vendor-specific API. The Dockerfile is based on an NVIDIA Vulkan image, but that is a build convenience, not a device restriction.

What is the difference between Kompute and CUDA?

CUDA runs on NVIDIA hardware only, while Kompute targets any Vulkan-capable GPU and includes mobile support through the Android NDK. The README positions Kompute for cross vendor and mobile use cases, so the choice is about which devices you must support rather than about raw speed.

Official sources

  1. KomputeProject/kompute on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/komputeproject-kompute.svg)](https://hysenlabs.com/projects/komputeproject-kompute)