gpu-burn: holding a GPU at full load to find out whether it holds
Multi-GPU CUDA stress test
At a glance
- What is it?
- A small CUDA stress test whose real design decision is that the compute kernel is compiled at build time into a fat binary, so the running program needs no CUDA toolchain. What it does not tell you is which temperature killed your card.
- Who is it for?
- gpu-burn answers one question well: does this card hold a stable result under sustained load, across every device in the machine at once. Its whole structure serves that goal, from the multi-GPU default to the build-time kernel compilation and the timed exit that makes it scriptable.
- Can I use it commercially?
- Yes. BSD-2-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 5 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Two files, one of which is compiled before you run it
The tree is small enough to read in full: `compare.cu`, `gpu_burn-drv.cpp`, a `Makefile`, a `Dockerfile`, a man page at `gpu-burn.8`, a `.clang-format`, a `LICENSE`, and a `win/` directory. That split into exactly two source files is the design.
`compare.cu` is the CUDA source, and `gpu_burn-drv.cpp` is the host driver. The Makefile turns the first into `compare.fatbin` and compiles the second into `gpu_burn`. The `%.fatbin: %.cu` rule invokes `nvcc` with `-fatbin` to produce the device code ahead of time.
The consequence is that nvcc is a build-time dependency and not a runtime one. You can run `gpu_burn` on a machine that has the CUDA runtime and driver libraries but no compiler, which is the situation on most managed GPU instances. The link line in the Makefile confirms what the running program needs: `libcuda`, `libcublas`, and `libcudart`, plus rpath entries pointing back at the CUDA `lib64` and `lib` directories.
This is a deliberate trade in the other direction too. Because the kernel is baked at build time, it targets one compute capability unless you build a fat binary deliberately. Which is exactly what the build knobs are for.
Compute capability is the build-time decision
The default compute capability is 7.5, matching what the README says NVIDIA's compute capability table specifies, and the Makefile expresses it as `COMPUTE ?= 75`. Passing `COMPUTE` on the make line sets a single virtual architecture through the default `-arch` flag.
For anything more precise you blank the variable and drive the build entirely from `-gencode` flags. The README's example targets two architectures, which produces a single fat binary that will run on either:
make COMPUTE= NVCCFLAGS='-gencode=arch=compute_86,code=sm_86 -gencode=arch=compute_90,code=sm_90'The same mechanism reaches newer targets that plain `-arch` cannot express. The README calls out architecture-conditional features with the `sm_90a` suffix and family-specific ones with `compute_100f`, both of which are Hopper and Blackwell era requirements that predate the flag's default behaviour. If you are building for a card that needs one of those, this is the supported route and it requires knowing the syntax yourself.
The pattern holds for the other toolchain variables. `CFLAGS` and `LDFLAGS` are appended to rather than replacing the defaults, so `make CFLAGS=-Wall` adds a warning flag without dropping `-O3`, `-std=c++11`, and the include path. `LDFLAGS` behaves the same way. `NVCCFLAGS` follows the append rule too, which is what makes the `-gencode` example above compose with the default include path instead of replacing it.
CUDAPATH: the README and the Makefile disagree
The README says `CUDAPATH` points to a non standard install or a specific version of the CUDA toolkit, and gives `/usr/local/cuda-<version>` as the example, with the default stated as `/usr/local/cuda`. The Makefile does something else, and it does it before your override ever applies.
ifneq ("$(wildcard /usr/bin/nvcc", "")
CUDAPATH ?= /usr
else ifneq ("$(wildcard /usr/local/cuda/bin/nvcc", "")
CUDAPATH ?= /usr/local/cuda
endifTwo details matter here. The probe order puts a system-wide `nvcc` at `/usr/bin` ahead of the conventional `/usr/local/cuda`, so on a distribution that ships nvcc in the default PATH the effective default is `/usr` rather than `/usr/local/cuda`. And the assignment uses `?=`, so an explicit `CUDAPATH` on the command line or in the environment still wins over both branches. The README's statement of the default is only true when no system nvcc exists.
This is the kind of divergence worth knowing about before you debug a build failure. If your headers and libraries appear to be coming from the wrong toolkit, the Makefile's probe order is the reason, and passing `CUDAPATH` explicitly is the fix. The sibling variable `CCPATH` selects the host compiler, defaulting to `/usr/bin`, and the same wildcard-free approach applies.
One more detection detail the README never mentions: the Makefile greps `/proc/device-tree/model` for the word Jetson and compiles with `-DIS_JETSON=true` or `-DIS_JETSON=false` accordingly. Jetson support is therefore automatic on a real device tree and silent everywhere else, with no flag to set by hand.
Reading the usage output
The `gpu-burn.8` man page exists in the tree, though the README's own summary of options is the more useful reference because it fits on a screen. The program takes an optional duration in seconds as its only positional argument, plus a short set of flags.
GPU Burn
Usage: gpu_burn [OPTIONS] [TIME]
-m X Use X MB of memory
-m N% Use N% of the available GPU memory
-d Use doubles
-tc Try to use Tensor cores (if available)
-l List all GPUs in the system
-i N Execute only on GPU N
-h Show this help message
Example:
gpu_burn -d 3600Four of these change what you are testing rather than how long you test for. `-d` switches to double precision, which on consumer cards with poor FP64 throughput will make the reported rate drop by orders of magnitude, and a low number there is a property of the card rather than a fault. `-tc` asks for Tensor cores where they exist. The `-m` flag in both its absolute megabyte and percentage forms is the one that makes this usable in a shared environment, since a full-occupancy burn on a GPU with other tenants on it will disturb them.
`-l` and `-i N` are the operational pair. Listing the devices first, then selecting one, is the workflow for bisecting a multi-GPU machine where only some cards are unstable. Without `-i`, the default is to run on everything at once, which is the right default for a burn-in on new hardware and the wrong one for finding the culprit.
The positional TIME argument is what makes the tool scriptable. The Dockerfile's default command is `60`, so the container image runs for one minute and exits, which is a sensible smoke test default and a good thing to know before you run the image expecting a longer run.
The Docker path is the shortest route to a run
Four commands take you from nothing to a loaded GPU, and the first one is a plain clone.
git clone https://github.com/wilicc/gpu-burn
cd gpu-burn
docker build -t gpu-burn .
docker run --rm --gpus all gpu-burn`--gpus all` is the part people get wrong on first contact, and it is the reason the README leads with Docker rather than with `make`. Runtime container GPU access requires the NVIDIA container toolkit, and without it the container starts and immediately fails to find a device.
The `Dockerfile` is a two-stage build that mirrors the Makefile's own separation. The builder stage is `nvidia/cuda:${CUDA_VERSION}-devel-${IMAGE_DISTRO}` and runs `make COMPUTE=${COMPUTE}`, the devel tag being what supplies nvcc. The final stage is the matching `-runtime` image and copies in only `gpu_burn` and `compare.fatbin`. The entrypoint is `./gpu_burn` and the default command is `60`.
All three build arguments have defaults in the Dockerfile, and they are the same defaults as the Makefile: `CUDA_VERSION=11.8.0`, `IMAGE_DISTRO=ubi8`, and `COMPUTE=75`. Override them at build time as the README shows:
docker build --build-arg CUDA_VERSION=13.0.0 --build-arg COMPUTE=75 --build-arg IMAGE_DISTRO=ubi8 -t gpu-burn .The Makefile also has an `image` target that forwards the same three variables and honours an `IMAGE_NAME` override, so `make IMAGE_NAME=myregistry.private.com/gpu-burn image` is the route when you want the image tagged for a private registry without retyping the build command.
The ubi8 default base is worth noting if you deploy this, since it means the runtime layer is a Red Hat UBI image rather than a bare CUDA runtime, which is a deliberate choice for enterprise hosts and an unfamiliar one on a workstation.
Packaging, platforms, and what the project does not measure
There are no GitHub releases at all. The README's answer to how you get a binary is Repology, the packaging database, which is where distributions that ship gpu-burn get their builds. Combined with a `win/` directory in the tree and a man page for Unix, that describes the project's packaging posture accurately: a source repository with downstream packagers, not a project that ships artifacts.
The name is worth pausing on too, because the README links to the author's own write-up of the tool and the word `burn` refers to the thermal stress being applied. The repository records 59 open issues and the last push was on 2026-05-31, so the tree is current even though no tag has ever been cut.
What the output does not include is the number that most people running a burn actually want. gpu_burn reports a computation rate, and it prints a line when a device reports an error, which is what catches a dying card. It does not report temperature, power draw, or clock throttle reasons, so a card that passes while sitting at its thermal limit is indistinguishable in this output from one with plenty of headroom. If you are chasing a machine that crashes under load rather than verifying a new one, this tool will tell you that something is wrong without telling you what.
There is also no assertion mode. The program computes and reports; it does not exit non-zero on a bad result, so a CI job checking GPU health has to parse the numbers rather than trust the exit status.
Editorial conclusion
gpu-burn answers one question well: does this card hold a stable result under sustained load, across every device in the machine at once. Its whole structure serves that goal, from the multi-GPU default to the build-time kernel compilation and the timed exit that makes it scriptable. What it deliberately leaves to you is thermal context, since the output is a computation rate and not a temperature reading, which means a run that looks fine at ninety degrees and a run that looks fine at sixty are not the same result. Pair it with whatever temperature monitoring your distribution already installs, keep the run shorter than you think you need, and treat the presence of a `win/` directory as promising rather than as evidence, since the README documents only the Linux path.
Frequently asked questions
What does gpu-burn measure?
It drives CUDA GPUs at sustained full load and reports the computation rate achieved, exiting after the duration you give it. A device reporting an error is printed, which is how an unstable card shows up. It does not report temperature, power draw, or clock throttle reasons.
How do I build gpu-burn?
Run `make` in the repository. The default compute capability is 7.5; pass `COMPUTE=` with a value to change it, or blank it and supply your own `-gencode` flags through `NVCCFLAGS` to build a fat binary for several architectures. `make clean` removes the built artifacts.
How do I run gpu-burn in Docker?
Build the image with `docker build -t gpu-burn .` and run it with `docker run --rm --gpus all gpu-burn`. Passing the GPUs through requires the NVIDIA container toolkit on the host. The image's default command is `60`, so it runs for one minute unless you override the duration.
Can I run gpu-burn on only one GPU in a multi-GPU machine?
Yes. Use `-l` to list all GPUs in the system, then `-i N` to execute only on GPU N. Leaving both off runs on every device at once, which is the right default for burn-in and the wrong one for isolating which card is unstable.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/wilicc-gpu-burn)