Framework
NVlabs/tiny-cuda-nn avatar
NVlabs/tiny-cuda-nn

Tiny CUDA Neural Networks: a C++/CUDA framework for fully fused MLPs and hash encodings

Lightning fast C++/CUDA neural network framework

4,538 stars581 forksC++NOASSERTION

At a glance

What is it?
NVlabs/tiny-cuda-nn is a small C++/CUDA library for training and querying small neural networks on the GPU, with a multiresolution hash encoding and an optional JIT fusion path added in v2.0. It is built for people who want the network inside their own CUDA application, not for people who want a general deep learning stack.
Who is it for?
Adopt tiny-cuda-nn if your problem is a small network queried at high frequency from your own CUDA code, and you are willing to build against CUDA yourself or accept the pip wheel path for the PyTorch bindings. Do not adopt it as a general training framework: it has no data loader, no checkpointing story in the README, and no CPU fallback, so anything that needs portable training runs belongs elsewhere.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What tiny-cuda-nn is for, and who it is not for

Most neural network libraries assume the network is the whole program. You hand them a dataset, they own the training loop, and you get a model file back. tiny-cuda-nn inverts that. It is a library you call from inside your own CUDA application, and the README describes it as "a small, self-contained framework for training and querying neural networks" whose headline components are a "fully fused" multi-layer perceptron and a multiresolution hash encoding. The intended user is someone writing a renderer, a simulation, or an interactive tool who needs a tiny network evaluated millions of times per frame and does not want a Python runtime in the loop.

The performance chart in the README compares fully fused networks against TensorFlow v2.5.0 with XLA, measured on 64 and 128 neuron wide MLPs on an RTX 3090, generated by benchmarks/bench_ours.cu and benchmarks/bench_tensorflow.py using data/config_oneblob.json. That is the claim to evaluate: not accuracy on a benchmark suite, but throughput on small dense networks.

The wrong user is anyone who wants a general deep learning framework. There is no dataset abstraction, no distributed training, no CPU execution path described, and no model zoo. If your network has millions of parameters per layer or your training data lives in a data warehouse, this is the wrong tool and the README does not pretend otherwise.

The two mechanisms: a fully fused MLP and a multiresolution hash encoding

The interesting design decision is that the network is not a sequence of cuBLAS calls. The README points to a diagram of the "fully fused" MLP and a technical paper behind it, and the JIT fusion section explains the underlying machinery: "JIT fusion works by converting a given tiny-cuda-nn model to a CUDA device function and then compiling it into a kernel using CUDA's runtime compilation (RTC) feature." In other words, the model is not just a set of weights but a piece of generated CUDA source, compiled at runtime and launched as one kernel. That is why the speedup exists and also why the failure modes are unusual: JIT compilation can fail, and the README states that if it does, "a warning will be emitted and JIT compilation mode automatically turned off" and your code falls back to the tiny-cuda-nn 1.X code path.

The second mechanism is the input encoding. The multiresolution hash encoding, also backed by a paper, is configured through a JSON object in the C++ API. The README's example uses otype "HashGrid" with n_levels 16, n_features_per_level 2, log2_hashmap_size 19, base_resolution 16 and per_level_scale 2.0. These parameters decide how a low-dimensional input coordinate is expanded into a higher-dimensional feature vector before it reaches the MLP. The README also says the framework supports "various other input encodings, losses, and optimizers", so HashGrid is one option among several rather than the only path.

The data flow through the API is explicit. You build a config, call create_from_config, and get back a struct with a trainer, a network, and the encoding. Training happens by calling model.trainer->training_step(inputs, targets, &loss) in a loop you write. Inference happens by calling model.network->inference(inputs, outputs). The library does not call you back.

Installing tiny-cuda-nn and running the 2D image sample

The README does not spell out a full installation procedure in the text available here, so the practical starting point is the repository layout. There is a CMakeLists.txt at the top level, a dependencies/ directory, a bindings/ directory for the PyTorch integration, and a samples/ directory containing mlp_learning_an_image.cu and mlp_learning_an_image_pytorch.py. If you want the C++ path, you configure and build with CMake against your CUDA toolkit; if you want the Python path, the bindings directory is what produces the tinycudann module.

The README does give one concrete command for the bundled sample application, which learns an image function (x,y) -> (R,G,B). It is run from the repository root after building:

sh
./build/mlp_learning_an_image data/images/albert.jpg data/config_hash.json

The README says this produces an image every couple of training steps, and the truncated text ends mid-sentence at "Each 1000 steps", so the exact output cadence beyond that is not documented in the README. Expect a build directory containing the compiled sample, and expect the process to write images as it trains.

On the Python side, the README shows the binding surface directly. The model is constructed as a NetworkWithInputEncoding and JIT fusion is toggled as a property:

python
import tinycudann as tcnn

model = tcnn.NetworkWithInputEncoding(...) # Or any other tcnn model
model.jit_fusion = tcnn.supports_jit_fusion() # Enable JIT if the system supports it

Note the comment in the README: enabling JIT fusion through the PyTorch bindings gives a lower speed-up than the C++ path, "particularly during training", because the JIT compiler cannot see the whole compute graph.

JIT fusion is not free, and the README says so

The most useful part of the documentation is the list of cases where the new feature hurts. JIT fusion arrived with v2.0 and is described as optional and "almost always" recommended, with a stated performance boost of 1.5x to 2.5x depending on model and GPU, and larger speedups on newer GPUs. But the same section states that JIT fusion can slow down training when your model has very large hash grids (roughly 20 million or more parameters), when MLPs have layer sizes larger than 128 neurons, or when your GPU is an RTX 3000 series or earlier. It adds that inference is rarely slowed too, and recommends enabling JIT fusion separately for training and inference to measure which is faster.

That is a real constraint, not a footnote. It means the correct configuration is hardware and model dependent, and the only way to know is to measure on your own setup. The README explicitly invites users to open an issue if a slowdown appears in a situation outside that list, which tells you the boundary is empirical rather than fully characterised.

The second limitation is structural. JIT fusion generates CUDA source and compiles it at runtime, so it depends on the CUDA runtime compilation path being available and working. The README's fallback behaviour is graceful, a warning and a return to the 1.X code path, but it also means a deployment can silently lose its performance characteristic without failing. If your application depends on the fused path for latency, you need to detect that fallback rather than assume it.

The third limitation is the batch size rule stated in the C++ example: batch_size must be a multiple of tcnn::BATCH_SIZE_GRANULARITY. There is no padding helper shown in the README, so the caller owns that arithmetic.

How tiny-cuda-nn differs from a general PyTorch model

The obvious alternative is to write the same small MLP in PyTorch and call it from your application. The difference is where the network lives at runtime. A PyTorch module is a Python object backed by kernels that the framework selects; tiny-cuda-nn generates a device function for your specific model and, with JIT fusion enabled, compiles it into a kernel alongside your own code. The README's manual JIT fusion example shows the shape of that integration: you write your kernel as a string, prepend the model's device function via model->generate_device_function("model_fun"), and hand the result to tcnn::CudaRtcKernel.

That is a different contract. In PyTorch you get a large ecosystem, autograd, checkpointing, and portability across backends. With tiny-cuda-nn you get a narrow API, a hard CUDA dependency, and the ability to fuse the network into a kernel that also does your ray marching or whatever else the application needs. The README cites Instant NGP as the extreme case, achieving a 5x speedup by fusing the entire NeRF ray marcher into a single kernel.

The constraint that comes with that: the fused kernel requires all 32 threads of the warp to be active at the model call, as the comment in the manual JIT example states, and the launch thread count must be a multiple of 32. Those are not suggestions you can ignore. If your kernel has divergent warp behaviour around the network call, the fused path is not available to you.

For pure Python users who never intend to write CUDA, the PyTorch bindings are the pragmatic option, and the README is honest that the speedup there is smaller. If you are already happy with PyTorch performance, the bindings add a dependency without obviously changing your outcome.

Maintenance, versioning and licence

The repository is not archived, and the last push was on 2026-04-21. The release history is uneven: v1.5 in April 2022, v1.6 in December 2022, then a gap to v2.0 in July 2025. That gap matters if you are planning around the 1.X API, because JIT fusion and the rtc_kernel.h header only exist from v2.0 onward. Pinning to a 1.X tag means you do not get the fused path at all, and the README's fallback language ("Your code will still run using the tiny-cuda-nn 1.X code path") implies the 1.X path is still maintained inside the 2.0 tree rather than as a separate branch.

Upgrade cost from 1.X to 2.0 looks low for existing code, since JIT fusion is opt-in through set_jit_fusion and the create_from_config API is unchanged in the examples. The cost is in verification, not in code changes: you have to measure whether the fused path helps or hurts on your hardware and model size, per the README's own caveats.

The licence is the item to check before shipping. The repository metadata reports the licence as NOASSERTION, which means GitHub could not map LICENSE.txt to a recognised SPDX identifier. The file exists at the top level, but the README does not state its terms. Read LICENSE.txt directly, and if your product is distributed in binary form, confirm what attribution or redistribution terms apply before you build it into a shipped artifact. This is a factual gap in the metadata, not a legal opinion.

Editorial conclusion

Adopt tiny-cuda-nn if your problem is a small network queried at high frequency from your own CUDA code, and you are willing to build against CUDA yourself or accept the pip wheel path for the PyTorch bindings. Do not adopt it as a general training framework: it has no data loader, no checkpointing story in the README, and no CPU fallback, so anything that needs portable training runs belongs elsewhere. Before committing, verify three things on your own hardware: that your GPU supports JIT fusion and whether it actually speeds up your model, since the README states that JIT fusion can slow down training on RTX 3000 series or earlier GPUs and on very large hash grids or MLPs wider than 128 neurons; that your batch size is a multiple of tcnn::BATCH_SIZE_GRANULARITY, which the API comment requires; and that the LICENSE.txt terms suit your distribution model, because the repository metadata reports the licence as NOASSERTION rather than a recognised SPDX identifier.

Frequently asked questions

How do you install tiny-cuda-nn?

The repository provides a top-level CMakeLists.txt, a dependencies/ directory and a bindings/ directory for the PyTorch integration. The README does not lay out a step-by-step install procedure, so the build path is CMake against your CUDA toolkit, and the Python path comes from the bindings. The one command the README does give is for the bundled sample: ./build/mlp_learning_an_image data/images/albert.jpg data/config_hash.json.

What is tiny-cuda-nn?

It is a small, self-contained C++/CUDA framework for training and querying neural networks, from NVlabs. Its main components are a fully fused multi-layer perceptron and a multiresolution hash encoding, plus support for other encodings, losses and optimizers.

What is the difference between cuDNN and CUDA?

The README for tiny-cuda-nn does not describe cuDNN or explain how it relates to CUDA, so this question cannot be answered from the available sources. What the README does establish is that tiny-cuda-nn depends on CUDA, including CUDA runtime compilation for its optional JIT fusion feature.

Is CUDA written in C or C++?

The README for tiny-cuda-nn does not address the language history of CUDA itself. It does show that tiny-cuda-nn exposes a C++/CUDA API, with headers such as tiny-cuda-nn/common.h and tiny-cuda-nn/rtc_kernel.h, and that models are configured through JSON objects in C++.

Official sources

  1. Issues
  2. NVlabs/tiny-cuda-nn on GitHub
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/nvlabs-tiny-cuda-nn.svg)](https://hysenlabs.com/projects/nvlabs-tiny-cuda-nn)