Library / SDK
Zaneham/Booth avatar
Zaneham/Booth

Zaneham/Booth: a C99 CUDA, HIP and Triton compiler that also targets CPUs

Open-source CUDA, Triton and HIP compiler targeting multiple GPU and CPU architectures.

1,750 stars93 forksCApache-2.0

At a glance

What is it?
Booth takes CUDA C, HIP and Triton source and emits AMD RDNA binaries, NVIDIA PTX, Tenstorrent Metalium C++ or native RV32IM, plus plain x86-64 that runs without a GPU. It is a small C codebase with a single-binary release and a narrow set of documented frontends.
Who is it for?
Booth fits engineers who already have CUDA C, HIP or Triton kernels and want a second path to AMD RDNA, NVIDIA PTX, Tenstorrent Metalium C++ or plain x86-64, especially if they want to run a Triton kernel on a machine with no GPU. Skip it if you need a documented rollback story, a stable CLI name, or a vendor support contract: the binary is kath, the release is a tarball, and the README points at docs/features.md for what compiles today.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 16 days ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Booth compiles, and who that is actually for

Booth accepts CUDA C, HIP and Triton source, the same files you would otherwise hand to nvcc, ROCm or Triton's JIT, and lowers them to AMD RDNA 2/3/4 binaries, NVIDIA PTX, Tenstorrent Metalium C++ or native RV32IM. It also emits plain x86-64, which is the part the README itself flags as unusual: a Triton kernel, matmul included, can be run on a machine that has never seen a GPU, with no LLVM in the path. The README states the author has not come across anyone else doing Triton this way and invites correction.

Two further frontends read the output of another compiler rather than the source directly. Fortran do concurrent kernels arrive through LFortran, which emits CUDA source that Booth compiles the rest of the way, and the results are checked against SLATEC values in CI. OCaml kernels are written as ordinary functions, type-checked by ocamlc, and Booth reads the resulting .cmt file. The OCaml path is the more interesting of the two because type errors surface before Booth sees anything: writing an int where a 32-bit device integer belongs, or reading a block-shared array as if it were global, fails at ocamlc. The README gives an Asian option pricer written this way as an example that runs on an RTX 4060 Ti and agrees with its closed-form reference.

The audience is narrow and specific. If you already write CUDA and want to see what your kernel looks like as an AMD binary, or you write Triton and want a CPU fallback for CI, Booth is aimed at you. It is not a drop-in replacement for nvcc in a build system that assumes nvcc's flag surface.

How the compiler is put together: frontends, IR, backends

The repository layout shows the pipeline split into directories: src/fe for frontends, src/ir for the intermediate representation, src/tdf, then one directory per target (src/amdgpu, src/tensix, src/nvidia, src/metal, src/intel, src/triton, src/cpu), with src/build and src/exec alongside runtime/include. The Makefile's include list matches that structure, so the backends are separable compilation units rather than one monolithic code generator.

The Makefile also reveals a deliberate split between the compiler proper and the host-side launchers. Its comment states that the launchers dlopen a vendor driver and link into trunner only, never into kath. nv_rt is described as portable, using LoadLibraryA on Windows and dlopen elsewhere, while bc_runtime is Linux-only and depends on dlfcn.h and libhsa. That means the compiler binary itself does not carry a vendor runtime dependency, and the runtime pieces are loaded at execution time instead.

The README credits specific academic work behind the backend: dominators from Cooper, Harvey and Kennedy, SSA spilling from Braun and Hack, and divergence analysis from Sampaio, Souza, Collange and Pereira. It also states the SSA register allocator exists because of a conversation with Fernando Magno Quintão Pereira at UFMG. That is a useful signal about the design lineage, and it tells you the register allocation and divergence handling are the parts the author considers load-bearing.

Installing Booth from the release tarball and compiling a first kernel

The README's preferred route is the release archive rather than a build. It states the archive comes with no dependencies and does not require running make. The commands below are the ones the README gives; the binary inside is named kath, not booth, because there is already a booth in the Linux HA stack, so both may end up on your PATH.

bash
tar xzf booth-*-linux-x86_64.tar.gz
cd booth-*-linux-x86_64
./kath --version

After that you should see a version string from the unpacked binary. The README says builds exist for Linux, macOS and Windows.

The first real use is compiling a CUDA kernel to an AMD GPU binary. The README gives this exact invocation, writing a .hsaco file:

bash
./kath --amdgpu-bin kernel.cu -o kernel.hsaco

The output is an AMD GPU binary you would load through a host-side launcher. The full flag set and every backend are in docs/usage.md, which the README points to as the command reference. The examples directory contains runnable starting points, including examples/launch_saxpy.c, examples/cpu_launch_vadd.c and examples/cpu_launch_matmul.c for the CPU path, and examples/koyeb_tensix_launch.cpp for Tenstorrent.

If you would rather build from source, the README gives one command and one requirement:

bash
make

A C99 compiler is the whole dependency list. The Makefile confirms this: it sets -std=c99 with gcc as the default CC, and adds -lm. LFortran and OCaml are needed only if you want the Fortran and OCaml frontends respectively, and neither is required to build Booth.

The kath binary, the missing rollback story, and other rough edges

The naming is the first thing that will trip up a scripted install. The project is Booth, the repository is Zaneham/Booth, and the binary is kath. The README explains the choice and acknowledges the collision with the existing booth in the Linux HA stack. Any wrapper, container image or PATH assumption you write has to account for that.

The second limitation is that the README does not document rollback. There is no stated procedure for reverting a release, no version pinning guidance, and no compatibility promise between the v5.01 and v0.5.x release lines. The release list itself is confusing: v0.5.2, then v5.01, then v0.5.0 (BarraCUDA 0.5). Anyone deciding whether to upgrade needs to read CHANGELOG.md, which the README describes as a running log of what has changed, because the version numbers alone do not tell you which line is current.

The third is coverage. The README is explicit that docs/features.md records what compiles today and what does not yet. That document, not the README, is where you find out whether your kernel is supported on your target. The README's claim about Triton-to-CPU is framed as a personal observation rather than a compatibility statement, and the author invites anyone who has seen it elsewhere to say so.

Finally, the Makefile carries a portability caveat in a comment: Apple's clang rejects GCC-only warning flags such as -Wstack-usage and -Wredundant-decls as a hard error under -Werror, so on clang those flags and -Werror itself are dropped until the tree is proven clang-clean. That is a maintainer's own note that the strict build is not yet uniform across compilers.

Booth against nvcc, ROCm and Triton's JIT

The comparison that matters is not Booth versus another hobby compiler. It is Booth versus the vendor toolchain you already use. nvcc and ROCm are single-vendor by construction: nvcc targets NVIDIA, ROCm targets AMD, and moving a kernel between them means rewriting or going through a portability layer. Booth takes the same CUDA C, HIP or Triton input and offers several targets from one frontend, including PTX for NVIDIA and RDNA binaries for AMD. The practical difference is that you keep one source file and choose the backend at the command line.

The second difference is the CPU path. Triton's JIT compiles for a GPU and expects one to be present. Booth's README describes compiling a Triton kernel, matmul included, straight to native x86-64 with no LLVM and running it on a machine with no GPU. That is a genuinely different capability rather than a faster or slower version of the same thing, and it changes what CI can do: a kernel test can run on an ordinary build machine.

The trade-off is maturity and support. nvcc and ROCm ship with vendor documentation, release cadences and bug trackers tied to hardware generations. Booth is a C99 codebase maintained by one person, with a documented feature-status page that admits gaps, a licence that says do whatever you want, and an email address for bug reports. If your kernel is on the critical path for a product, the vendor toolchain is the lower-risk choice; Booth is the one that tells you whether your kernel is portable at all.

Licence, maintenance and what an upgrade costs

Booth is Apache-2.0. The README's own summary is that you can do whatever you want, and it adds that the author would like to hear about production use. Apache-2.0 includes an explicit patent grant and requires attribution and notice retention, which matters if you vendor the source into a product. This is a description of the licence text, not legal advice; if you are shipping Booth inside something commercial, have counsel read the NOTICE and attribution requirements against your distribution model.

The repository is not archived, and the last push was on 2026-08-08, the same date as the v0.5.2 release. That is recent enough that the project is being worked on rather than parked, but the release numbering across v0.5.2, v5.01 and v0.5.0 suggests the version scheme has moved around, so treat any upgrade as a change to read about rather than a patch to apply blindly.

Upgrade cost has three components. The tarball install is cheap because there is nothing to build, but it is also the case where you have the least visibility into what changed, so CHANGELOG.md is the only signal. Building from source costs one make invocation with a C99 compiler, and the strict warning set in the Makefile means a build on a compiler the tree has not been proven clean against may need the same flag relaxation the Makefile already applies to clang. The third component is the launchers: because they dlopen the vendor driver and link into trunner rather than kath, a driver-side change can affect execution without the compiler binary changing at all.

Editorial conclusion

Booth fits engineers who already have CUDA C, HIP or Triton kernels and want a second path to AMD RDNA, NVIDIA PTX, Tenstorrent Metalium C++ or plain x86-64, especially if they want to run a Triton kernel on a machine with no GPU. Skip it if you need a documented rollback story, a stable CLI name, or a vendor support contract: the binary is kath, the release is a tarball, and the README points at docs/features.md for what compiles today. Before committing, check docs/features.md against your kernel, confirm the target in docs/hardware.md, and run the tarball's ./kath --version to see which build you actually have.

Frequently asked questions

How do I install Zaneham/Booth?

The README points at the latest release archive, which it says comes with no dependencies and does not require make. You unpack it and run ./kath --version inside the extracted directory. Alternatively, running make builds it from source with a C99 compiler.

What is Booth written in, and what does it need to build?

The primary language is C, and the Makefile sets -std=c99 with gcc as the default compiler. The README states you need a C99 compiler and nothing else. LFortran and OCaml are only required if you want the Fortran and OCaml frontends.

Why is the Booth binary called kath?

The README says the binary is kath after Kathleen Booth, and that it is not named booth because there is already a booth in the Linux HA stack, so you may end up with both on your PATH.

Which GPU architectures does Booth target?

The README lists AMD RDNA 2/3/4 binaries, NVIDIA PTX, Tenstorrent Metalium C++ and native RV32IM, plus plain x86-64 that runs without a GPU. The repository also has backend directories for metal and intel, though the README does not describe those targets.

What licence does Zaneham/Booth use?

The repository is Apache-2.0. The README summarises it as do whatever you want and asks to hear about production use.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zaneham-booth.svg)](https://hysenlabs.com/projects/zaneham-booth)