Open-source project
NVlabs/cuda-oxide avatar
NVlabs/cuda-oxide

cuda-oxide: compiling Rust SIMT kernels to CUDA PTX with a custom rustc backend

cuda-oxide is a Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign language bindings, just Rust.

3,632 stars285 forksRustApache-2.0

At a glance

What is it?
cuda-oxide is an alpha-stage rustc codegen backend from NVlabs that turns #[kernel] functions in ordinary Rust into CUDA PTX, with a single-source build driven by cargo oxide. It is a real toolchain with real setup costs, and the README is explicit that APIs will break.
Who is it for?
Adopt cuda-oxide if you already write Rust and want to keep host and device code in one crate, and you can accept nightly Rust, a CUDA 13.x driver and a project the README itself labels alpha. Do not adopt it if you need a stable kernel API, Windows support, or a replacement for hand-tuned CUDA C on a shipping product.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What cuda-oxide solves for Rust developers writing GPU kernels

Writing a GPU kernel in Rust today usually means one of two compromises: a domain-specific language that is not Rust, or a foreign function layer where the kernel body lives in another language and Rust only orchestrates. cuda-oxide takes a third route. It is a custom rustc backend that compiles #[kernel] functions written in standard Rust into CUDA PTX. The README describes the goal as letting CUDA SIMT kernels be written natively in pure Rust, with no DSLs and no foreign language bindings.

The intended audience is narrower than "Rust programmers". You need to be comfortable with a nightly toolchain, with a cargo subcommand that wraps the codegen backend, and with GPU concepts like thread indexing, shared memory and launch configuration. The README's own status paragraph is blunt: the project is in an early stage (alpha) and under active development, and you should expect bugs, incomplete features, and API breakage. That sentence should shape how you read everything else on this page.

The compilation pipeline: Rust MIR to Pliron IR to LLVM IR to PTX

The mechanism is a rustc codegen backend, not a transpiler that rewrites source text. The README states the pipeline as Rust to Rust MIR to Pliron IR to LLVM IR to PTX. Pliron is described as an MLIR-like IR framework written in Rust, and the workspace pins it to upstream struct-layout, GEP no-wrap, volatile, global-constant, inline-asm and generic LLVM attribute semantics.

The workspace layout shows how that is split up. crates/mir-importer, crates/dialect-mir, crates/mir-lower and crates/mir-transforms handle the MIR side. crates/dialect-ptx, crates/ptx-parse and crates/ptx-schedule handle the PTX side. crates/cuda-oxide-codegen and crates/llvm-export sit in between. The actual rustc backend, crates/rustc-codegen-cuda, is deliberately not a workspace member: the Cargo.toml comment says it requires special rustc nightly features and a different build process, and directs you to build examples with cargo oxide run.

On the host side, the split is unusual and worth knowing before you file an issue. cuda-bindings, cuda-core and cuda-async are shared with cutile-rs and published from NVlabs/cutile-rs, and the cuda-oxide SIMT surface lives under their simt modules. The local copies are gone. A bug in DeviceBuffer or in the launch runtime is therefore not necessarily a cuda-oxide bug.

Installing cuda-oxide and running your first kernel

The requirements section is specific. You need cargo-oxide, the cargo subcommand that drives the build pipeline. You need Rust nightly with the rust-src, rustc-dev and llvm-tools components, pinned in rust-toolchain.toml. You need CUDA Toolkit 13.0 or later including the cuRAND headers (libcurand-dev on Ubuntu), and the README notes that the shared cuda-bindings crate loads libcuda at run time and needs a CUDA 13.x driver (R580+). You also need Clang plus libclang dev headers, named as clang-21 and libclang-common-21-dev, because bindgen needs them when building the host cuda-bindings crate. The tested platform is Linux, specifically Ubuntu 24.04.

Inside the cuda-oxide repo, cargo oxide works out of the box through a workspace alias. For your own projects you install it separately. The README's setup section covers this under an Install heading for cargo-oxide; the truncated text does not show the exact install command, so check that section of the README before assuming a crate name.

Once it is available, the repository's own examples are the fastest way in. This builds and runs the host_closure example:

bash
cargo oxide run host_closure

If you want to see what the compiler produced rather than just that it ran, inspect the generated PTX for the vecadd example:

bash
cargo oxide inspect vecadd

And to watch each stage of the pipeline in order, Rust MIR through dialect-mir and mem2reg, then the LLVM dialect, LLVM IR and finally PTX:

bash
cargo oxide pipeline vecadd

For a first kernel of your own, the README's quick start shows the shape. A #[cuda_module] module holds a #[kernel] function, the host side calls kernels::load(&ctx) to get a typed module, and launching goes through a generated method such as module.map::<f32, _>(...). The README is explicit that LaunchConfig is intentionally raw data and that using it to launch a kernel is unsafe, because the caller must prove the dimensions and resources match the kernel. Kernels annotated with #[launch_contract(...)] instead get a checked PreparedLaunch through the generated safe method. That distinction is the part most likely to bite a first-time user.

Where cuda-oxide gets in your way: alpha status, platform and driver constraints

The first limitation is stated by the project itself. The README says the project is alpha and that API breakage is expected. Three releases exist, v0.1.0 in May 2026, v0.2.0 in June 2026 described as the first community release, and v0.2.1 in June 2026. A codegen backend that is still moving will change the surface your kernels compile against, and the README gives no compatibility policy for that.

The second is the environment. Linux is the tested platform, on Ubuntu 24.04. The README does not document Windows or macOS support, and it does not document a container or WSL path, even though the repository carries a .devcontainer directory and a flake.nix. If your team develops on Windows, the README is silent on whether that works.

The third is toolchain coupling. You are on nightly Rust with rust-src, rustc-dev and llvm-tools, plus clang and libclang headers, plus CUDA 13.0 or later, plus an R580 or newer driver. That is a lot of moving parts to keep aligned, and a nightly bump can interact with the pinned Pliron dependency. The workspace also notes that the rustc backend crate needs special nightly features and a different build process, which means the backend is not built the way the rest of the workspace is.

Finally, the unsafe boundary is real. The README frames LaunchConfig as raw data whose correctness is the caller's responsibility. If you want the compiler to check your launch geometry, you have to opt in with #[launch_contract(...)]. That is a reasonable design, but it means the default quick-start path puts a proof obligation on you.

cuda-oxide compared with writing CUDA C or using another Rust GPU route

The closest comparison is plain CUDA C or CUDA C++ compiled by nvcc. That path is mature, has years of tooling behind it, and does not ask you to adopt a nightly compiler or a young codegen backend. What it does not give you is a single Rust crate where host and device code share types, generics and closures. The README's map example is the clearest illustration of the difference: the kernel is generic over T and over a closure type F, and rustc monomorphizes the closure to a concrete type, with the captured factor scalarized and passed as a kernel parameter. Reproducing that in CUDA C means writing the specialization by hand.

The other comparison, which the search data reflects, is Rust versus cuda-oxide and CUDA versus cuda-oxide. Those are not the same axis. CUDA is the platform; cuda-oxide is one way to produce PTX for it. Rust without cuda-oxide means either a DSL or FFI, which is exactly what this project is trying to remove.

The honest trade-off: cuda-oxide buys you type safety, generics and one build command, and it charges you an alpha API, a nightly toolchain and a Linux-only, CUDA 13.x-only environment. For a research prototype or an internal tool where the Rust ergonomics pay off, that is a defensible trade. For a product that ships kernels to customers on a schedule, the README's own warning about API breakage is the deciding fact.

Debugging, sanitizing and inspecting kernels with cargo oxide

The subcommand list is the part of the project that feels most finished, and it covers the loop you actually work in. Beyond run, inspect and pipeline, there is sanitize, which the README shows running CUDA correctness checks through a named tool:

bash
cargo oxide sanitize vecadd --tool memcheck

For interactive debugging there is a cuda-gdb wrapper with a TUI flag:

bash
cargo oxide debug vecadd --tui

Two more subcommands round out the set. cargo oxide test runs Cargo tests through the cuda-oxide backend, and cargo oxide emit-ltoir compiles a crate's device code to a binary LTOIR artifact in one step. There is also cargo oxide clean to remove project-local build outputs and generated artifacts, and cargo oxide update to refresh the cached codegen backend.

That last one matters more than it looks. Because the backend is cached, a stale backend is a plausible source of confusing failures after a toolchain change, and cargo oxide update is the documented remedy. The README states that cargo oxide --help lists every subcommand, which is the right place to check when you need a flag this page does not cover.

Licence, maintenance and what upgrading costs you

cuda-oxide is Apache-2.0, and the workspace manifest sets that licence for the workspace packages. The repository carries a THIRD_PARTY_NOTICES file and a dependency-licenses.csv, which suggests the dependency tree is audited rather than assumed. That matters here because the pipeline pulls in Pliron, LLVM, the CUDA toolkit and bindgen, each with its own terms.

Apache-2.0 includes an express patent grant and requires you to preserve notices and state changes. It does not require you to open your own kernels. None of that is legal advice; if you are shipping commercially, have your own counsel read the notices files rather than this paragraph.

The maintenance picture is mixed and the repository facts support only part of it. The repository is not archived, and the last push was on 2026-09-23. Three releases landed between May and June 2026, and no release has appeared since. The README describes the project as under active development, and the push date is consistent with that, but the release cadence is not something the dates alone let me characterise.

The upgrade cost is concrete. Because the backend is a rustc codegen backend tied to a pinned nightly in rust-toolchain.toml, and because Pliron is pinned to specific upstream semantics, moving your Rust nightly forward is not a routine bump. The README's alpha warning means you should expect to read release notes before every upgrade rather than treat them as optional.

Editorial conclusion

Adopt cuda-oxide if you already write Rust and want to keep host and device code in one crate, and you can accept nightly Rust, a CUDA 13.x driver and a project the README itself labels alpha. Do not adopt it if you need a stable kernel API, Windows support, or a replacement for hand-tuned CUDA C on a shipping product. Before committing, verify three things on your own machine: that rust-toolchain.toml pins a nightly your toolchain can install, that your driver meets the CUDA 13.x / R580+ requirement the README states, and that cargo oxide pipeline vecadd prints the full Rust MIR to Pliron IR to LLVM IR to PTX chain on your GPU.

Frequently asked questions

What is cuda-oxide?

It is a custom rustc backend from NVlabs that compiles GPU kernels written in pure Rust into CUDA PTX. The README describes it as letting you write CUDA SIMT kernels natively in Rust, with no DSLs and no foreign language bindings, built with a single cargo oxide build command.

Can Rust be used with CUDA through cuda-oxide?

Yes. cuda-oxide compiles #[kernel] functions in a #[cuda_module] to PTX, and the README's quick start shows a generic map kernel launched from Rust host code with a captured closure passed as a kernel parameter. Host and device code live in the same file and are built together.

What happens when you run a CUDA kernel launched from cuda-oxide?

The host side loads the generated module with kernels::load(&ctx) and launches through a generated typed method such as module.map::<f32, _>(...), passing a LaunchConfig and the buffers. The README states that LaunchConfig is intentionally raw data, so launching with it is unsafe and the caller must prove the dimensions and resources match the kernel.

Official sources

  1. License: Apache-2.0
  2. NVlabs/cuda-oxide on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvlabs-cuda-oxide.svg)](https://hysenlabs.com/projects/nvlabs-cuda-oxide)