# TornadoVM: the CUDA C tax, itemised

> This project's README makes its case in a single comment block listing what you still have to write on the host in CUDA C: allocate and copy per buffer, set grid dimensions, launch, synchronise, copy back, free, then build a binary per GPU and rewrite it for the other vendors. Everything else in the framework is a way of not writing that list.

**beehive-lab/TornadoVM** — Write Java. Run on GPUs. Fast.

- Repository: https://github.com/beehive-lab/TornadoVM
- Website: https://www.tornadovm.org
- Stars: 1,514 · Forks: 145
- Language: Java
- License: Apache-2.0
- Published: 2026-09-30 · Updated: 2026-09-30 · Language: en
- Canonical page: https://hysenlabs.com/projects/beehive-lab-tornadovm

## Bytecode to PTX to cubin, at run time

The compilation path is the whole idea, and the README states it as a chain: Java bytecode goes through Graal IR, becomes CUDA PTX, goes through NVRTC, and lands as a native cubin. On the other backends the same bytecode becomes OpenCL C or Apple's Metal Shading Language.

Every step of that happens at run time, and the kernel is specialised to the data sizes and the GPU you actually have. That word, specialised, is the part worth pausing on. Ahead-of-time GPU compilation forces you to commit to a shape before you know it, and most kernels are shape-sensitive: a tile size that suits a thousand rows is wrong for a million. Compiling when the shapes are known lets the JIT specialise the arithmetic for the sizes it is about to process, and it removes the need to ship per-GPU binaries at all, which is what the README means by no native toolchain in your application.

The kernel itself is Java with a CUDA-shaped indexing model, so the mental translation from an existing kernel is close to mechanical:

```java
void mxv(KernelContext ctx,
         FloatArray m,
         FloatArray v,
         FloatArray out,
         int rows, int cols) {
  int i = ctx.globalIdx;
  if (i < rows) {
    float sum = 0f;
    for (int j=0; j<cols; j++)
      sum += m.get(i*cols+j)*v.get(j);
    out.set(i, sum);
  }
}
```

A context object replaces the built-in variables, and array access goes through a typed wrapper rather than pointer arithmetic. The array type is an off-heap buffer, flat and row-major, which is a necessary consequence of the model: the data has to live in device memory, so it is allocated in a way the runtime can copy rather than on the Java heap.

There is a lower-control alternative for people who do not want that context object, and the README is explicit that it exists: annotate the loop instead and the framework infers the thread mapping. Both styles compose in the same graph, so a kernel that needs a barrier can have one and a kernel that does not need to think about it does not.

## What CUDA C hides in a comment

The README puts the same matrix-vector kernel in Java and in CUDA C side by side, and the kernels are the same length. The difference is a comment underneath the C version, and that comment is the argument.

It lists what you still write on the host: allocation and a copy for every buffer, the grid and block dimensions, the launch, a synchronisation, a copy back, a free. Then two more lines that matter more than all of that: an nvcc build plus per-GPU binaries, and a rewrite for non-NVIDIA GPUs.

The plumbing is tedious but it is ordinary tedium. The last two lines are the ones that scale badly, because they multiply. A per-GPU binary is one artefact per architecture per toolkit version per compute capability, and it is a release engineering problem rather than a coding one. A rewrite for the other vendors is not a port, it is a second implementation of the kernel with its own bug surface, and it is the reason most CUDA codebases support exactly one vendor and call it portability.

So the honest summary of what this framework buys you is not less code. It is one kernel body, one graph, and no per-vendor rebuild. The number of lines you write on the host goes down because the runtime owns the transfers, but the number of places your kernel logic exists goes to one, and that is the difference that survives contact with a second GPU.

There is a cost on the other side of the ledger, and it is the one people forget: your kernel is no longer a language the vendor's toolchain sees directly. Whatever the JIT cannot model, you do not get. That is a real constraint on exotic code, and it is why the framework's recent work has been about giving you the escape hatches rather than about hiding the boundary.

## Transfer modes are a scheduling decision

The host side is short, and almost all of it is a declaration about data movement.

```java
FloatArray m   = new FloatArray(rows * cols);
FloatArray v   = new FloatArray(cols);
FloatArray out = new FloatArray(rows);

KernelContext ctx  = new KernelContext();
WorkerGrid worker  = new WorkerGrid1D(rows);
GridScheduler grid = new GridScheduler("compute.mxv", worker);

TaskGraph tg = new TaskGraph("compute")
    .transferToDevice(DataTransferMode.FIRST_EXECUTION, m, v)
    .task("mxv", Kernels::mxv, ctx, m, v, out, rows, cols)
    .transferToHost(DataTransferMode.EVERY_EXECUTION, out);
```

Look at the two transfer modes. One buffer is declared to move on the first execution only, the other on every execution. That single annotation tells the runtime an input does not change between iterations, which means it can upload it once and keep it resident across a loop. For a matrix that is larger than the transfer path, that is the difference between paying for the copy once and paying for it a thousand times, and there is no way for a runtime to know which you meant.

So this is a scheduling interface rather than a convenience wrapper. The graph is a description of what depends on what and what may be reused, and the runtime compiles that description into device work.

Then the execution plan, which is a value rather than a call:

```java
try (TornadoExecutionPlan plan = new TornadoExecutionPlan(tg.snapshot())) {
    plan.withGridScheduler(grid).execute();
}
```

Two things are happening. The graph is snapshotted, which means the thing you execute is an immutable description that could be inspected, logged or compared before anything runs. And the plan is closed by a try-with-resources block, which means whatever device resources the plan holds are released deterministically rather than at the mercy of a garbage collector.

The grid scheduler is attached at execution time rather than baked into the graph, which is the same separation of concerns: what to compute, in the task graph; how to map it onto the hardware, at execution.

## One stream, shared buffers, no host round trip

The NVIDIA work in this project is not really about generating kernels any more. It is about making a generated kernel and a vendor library call be the same kind of thing.

The stated property is that generated kernels and native library calls share framework-managed device buffers on one CUDA stream, so a compiled kernel can feed a library call and consume its output with no extra copies and no host synchronisation. That is a much stronger claim than being able to call a library, and it is the whole point.

Consider what it takes to call cuBLAS from a host-managed system. You allocate the output, you copy the input across, you call the library, you copy the result back, you synchronise, and if a JIT kernel runs next you copy again. Every boundary is a copy and a synchronisation point. The alternative is one stream with resident buffers, where the library call is just another node in a schedule and the data never moves.

The library surface is broad. The cuBLAS and cuBLASLt tasks cover single and strided-batched GEMM plus a set of mixed-precision and tensor-core entry points, with plan caching, and fused epilogues where one library task replaces a GEMM followed by a separate activation kernel. The cuFFT tasks cover the complex, real-to-complex and inverse transforms in one and two dimensions with plan caching that is safe to capture into a graph. The cuDNN tasks go through the graph API, including fused attention.

Then there is CUDA graph capture, and it applies to the whole thing:

The execution plan can record kernels, library calls and transfers into a captured graph and replay the whole schedule with a single launch. That is the mechanism by which the transfer modes you declared earlier stop costing anything per iteration, and it is why the plan-caching in the library tasks is described as capture-safe rather than merely cached.

For a pipeline that alternates your own kernels with a vendor library, the practical question is how many synchronisation points remain. If the answer is zero, the pipeline is a schedule. If it is one per library call, it is a loop with a lot of asynchronous parts.

## Tensor Cores without a binding

The most technically aggressive line in the README is that Tensor Core matrix-multiply-accumulate instructions are exposed from pure Java, with a caveat that is worth reading exactly as written: not a binding, real CUDA generated from Java.

The distinction matters. A binding is a native function that takes your arguments and returns a result, which means you are calling into compiled code through a foreign function interface and you have accepted that boundary: an argument marshalling cost per call, a synchronisation per call, and a set of types you cannot mix with your own buffers. What is described here is instead access to the instructions through the same kernel context you use for thread IDs, with load, execute and store operations for the operand fragments, in half precision producing single-precision output and in eight-bit producing 32-bit output, and with the shared memory staging described as swizzled.

That last word is the argument. Shared memory banks are organised so that consecutive accesses from different threads in a warp hit different banks, and the natural layout of a matrix fragment is contiguous in rows, which means consecutive threads collide. The fix is to swizzle the layout so the access pattern is permuted, and getting a swizzle right is genuinely fiddly work that you do once per kernel shape. Generating it is exactly the case for a compiler rather than a wrapper.

The value proposition is that this closes the last gap in the porting argument. If the only way to get tensor cores was a native library call, then a Java codebase would be fast for its own kernels and slow for its matrix multiplies, and the vendor's tuned implementation would become a bottleneck you could not avoid without dropping out of Java. With the instructions reachable from Java, the same graph can mix them.

The framework also offers the simpler path for the many cases where none of this is needed, and the README is clear that both styles compose in one task graph.

## A Makefile that fails early, three separate times

The build system is worth reading on its own, because its comments are about the person running it rather than about the code.

There are three distinct ways a backend argument can go wrong, and each one gets an explicit error at the top of the Makefile instead of a confusing failure inside the build tool. The supported set is defined once:

```make
SUPPORTED_BACKENDS := opencl cuda metal
UNSUPPORTED_BACKENDS := $(filter-out $(SUPPORTED_BACKENDS),$(BACKEND_LIST))
```

And the comment above it explains why the check exists in two places at all: the installer performs the same check, and without the Makefile's version an unsupported name would be handed to the compile script, become a non-existent Maven profile, and fail much later and far less clearly.

The second trap is subtler and the comment shows the author thinking it through. A variable set to an empty value on the command line is not unset, so a defaulted assignment does not restore the default, and passing an empty backend value would build nothing and only fail deep inside the build. So there is an explicit check for an empty list, with the supported names in the error message.

The third is the JDK profile, and it is derived from the environment rather than hardcoded, with a long comment explaining why: the compile script itself checks that the environment's JDK matches the requested profile, so a fixed default can only ever fail once the environment moves to a newer JDK, and it fails confusingly because the goal that fails is not the one the user named. Deriving it removes the failure instead of documenting it.

There is also a separate makefile for Windows, an argument file example, an IDE run-configuration directory, and a Java compiler arguments file committed at the root with a timestamp in its name, which is a build artefact that should not be there and is a small reminder that a repository this large accumulates what nobody meant to add.

## Twenty modules, three licences, and a fuzzer

The module list is the map of the project's scope, and it is worth reading as a table of contents.

There are modules for each vendor library integration separately, which tells you the integrations are first-class rather than bolted on. There is an annotation module for the inferred-thread-mapping style, an API module that is the one published to Maven Central, an assembly module, a runtime module, a drivers module, and modules for matrices, benchmarks, examples and unit tests.

Two entries in that list are unusual and both are good signs. A fuzzing module means somebody is fuzzing a compiler that turns user-written Java into device code, which is the correct thing to do given that the input is a program and the output runs on a coprocessor. And a benchmarks module at the top level rather than in a separate repository means performance work is part of the product, not a side project.

Licensing is the other thing the tree tells you. There are three licence files, which is a tri-licensed project: the permissive pair you would expect from a foundation project, plus a copyleft variant, so you can take it under whichever terms suit your distribution.

Distribution runs through two documented paths. The build and run instructions are separate documents, and there is a dedicated guide for the hybrid API, which is the surface where a compiled kernel and a native library call meet. An SDKMAN badge suggests the install is a version-manager entry rather than a download, and the API artifact is on Maven Central, which means the parts of the framework you compile against are ordinary dependencies.

One detail to check before you start. The supported JDK list is given as 21, 25, 26 and 27, which is not every release above 21. If you are on 22, 23 or 24 you are outside what the README claims, and that is worth resolving before you spend an afternoon on a build failure.

## Where the portability claim stops

The honest boundary is the vendor library integration.

The framework's own code generation is genuinely multi-vendor. Your kernels compile to CUDA, OpenCL C or Metal, which means the same Java kernel runs on NVIDIA hardware, on AMD and Intel hardware, on integrated GPUs and on multi-core CPUs, and on Apple Silicon through the native Metal backend added in this release. That claim is structural, not a promise.

The library tasks are not. cuBLAS, cuFFT and cuDNN are NVIDIA products, and a task that calls one is a task that only runs on an NVIDIA device. So a pipeline built this way is portable in its kernels and vendor-specific in its libraries, and a port to AMD hardware means replacing those nodes. That is a reasonable trade, because a vendor's tuned library will usually beat what a portable stack provides, and it is worth being explicit that the library integration does not make the program portable. It makes it fast on one vendor.

The build system reflects the same shape. Three backends are supported, the default is OpenCL, and the topics list includes Level Zero alongside OpenCL and CUDA, which suggests interest in the vendor-neutral interface but does not by itself promise a fourth backend.

The other limit is the one from the previous section: the JIT has to be able to model your kernel. For a well-shaped array kernel it can, and the tensor-core intrinsics show how far the boundary has been pushed. For anything whose performance depends on a pattern the compiler cannot see, you are back to reasoning about what it emits.

And the release cadence suggests a project moving quickly. Two major versions eight days apart, a third release carrying a JDK version in its tag name, and a changelog linked from the README. That is fine to follow and awkward to depend on, so pick a version and read the changelog before you upgrade rather than after.

## Conclusion

Adopt TornadoVM if your kernels already live in Java and you need them on more than one GPU vendor, since the per-vendor binary and the second rewrite are the costs that multiply. Check your JDK against the supported list before anything else, because the README names 21, 25, 26 and 27 rather than every release above 21. Read the Makefile's backend validation before your first build, since the default is OpenCL and an unsupported name otherwise fails deep inside Maven. And treat the cuBLAS, cuFFT and cuDNN tasks as NVIDIA-only: they do not make your code portable to AMD or Intel.

## FAQ

### What is TornadoVM and which GPUs does it support?

TornadoVM is a GPU programming framework for Java that JIT-compiles bytecode at run time into CUDA, OpenCL C or Apple Metal, so the same kernel runs on NVIDIA, AMD, Intel and integrated GPUs as well as multi-core CPUs. It works with JDK 21, 25, 26 and 27 according to the README.

### How does TornadoVM differ from writing a CUDA kernel in C?

The kernel logic is comparable in length, but CUDA C also requires host-side allocation and copies per buffer, grid and block dimensions, launch, synchronisation, copy back and free, plus an nvcc build producing per-GPU binaries and a rewrite for non-NVIDIA vendors. TornadoVM manages the transfers through a task graph and compiles one Java kernel to whichever backend is present.

### What do the data transfer modes in a TornadoVM TaskGraph do?

They declare when a buffer moves. A transfer declared for the first execution only means the runtime can upload it once and keep it resident, so a repeated graph does not pay for the copy each iteration, while a transfer on every execution means the buffer is refreshed. That annotation is the scheduling decision the runtime cannot infer.

### Can Java code call cuBLAS, cuFFT and cuDNN through TornadoVM?

Yes, as library tasks in the same task graph as your compiled kernels, sharing framework-managed device buffers on one CUDA stream so no extra copies or host synchronisation are needed. Coverage includes single and batched GEMM with fused epilogues, one and two dimensional transforms, and fused attention through the cuDNN graph API. These tasks are NVIDIA-specific.

### Does TornadoVM expose Tensor Core instructions to Java?

Yes. The matrix-multiply-accumulate instruction is reachable through the kernel context, with load, execute and store operations, half precision producing single-precision output and eight-bit producing 32-bit output, with swizzled shared-memory staging. The README stresses this is generated CUDA rather than a foreign function binding.

### How do I choose a backend when building TornadoVM from source?

Pass it as a comma-separated list, for example an OpenCL-only or OpenCL-plus-CUDA build. The Makefile defines the supported set as opencl, cuda and metal with OpenCL as the default, and it validates the list up front so an unsupported name produces an immediate error rather than a missing Maven profile failing much later.

## Sources

- [beehive-lab/TornadoVM on GitHub](https://github.com/beehive-lab/TornadoVM)
- [License: Apache-2.0](https://github.com/beehive-lab/TornadoVM/blob/master/LICENSE)
- [Project website](https://www.tornadovm.org)
- [README](https://github.com/beehive-lab/TornadoVM/blob/master/README.md)
- [Releases](https://github.com/beehive-lab/TornadoVM/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/beehive-lab-tornadovm
