# NVIDIA Warp: JIT-compiled Python kernels for GPU simulation

> Warp turns annotated Python functions into CPU or CUDA kernels and ships physics, robotics and geometry primitives on top. It fits differentiable simulation pipelines; it is a poor fit if you have no NVIDIA GPU and no CUDA toolchain.

**NVIDIA/warp** — A Python framework for GPU-accelerated simulation, robotics, and machine learning.

- Repository: https://github.com/NVIDIA/warp
- Website: https://nvidia.github.io/warp/stable/
- Stars: 7,148 · Forks: 628
- Language: Python
- License: Apache-2.0
- Published: 2026-08-08 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-warp

## The gap Warp fills between NumPy loops and hand-written CUDA

A Python simulation loop over a million particles is slow because each element is visited in interpreter bytecode. The usual escape is to rewrite the inner loop in CUDA C++, which means a second build system, a second language, and a copy of the data across the host-device boundary every time the model changes. Warp's answer is to keep the loop in Python and compile it. You decorate a function with wp.kernel, and Warp JIT compiles it into kernel code that runs on the CPU or on a CUDA device, according to the README. The target audience is simulation and robotics engineers who already write Python and already have NVIDIA hardware: the pyproject.toml classifier lists Environment :: GPU :: NVIDIA CUDA and Development Status :: 5 - Production/Stable, and the package requires Python 3.10 or newer. The second audience is machine-learning researchers who need a differentiable physics step inside a training graph. Warp states that its kernels are differentiable and can be used as part of machine-learning pipelines with PyTorch, JAX and Paddle. That combination, a physics primitive library plus a compiler plus autodiff, is the reason to look at Warp rather than writing raw CUDA.

## How the JIT path works: wp.kernel, wp.array and wp.tid

The data flow is visible in the README's gravity example. Host data lives in wp.array objects, which are typed: the example builds positions and velocities as wp.array with dtype=wp.vec3 from a NumPy array. Inside a kernel, wp.tid() returns the index of the current thread, which is the same idea as threadIdx in CUDA but expressed without block or grid arithmetic in user code. The kernel body reads pos[i], computes a softened inverse-square acceleration, and writes back to vel[i] and pos[i]. Warp's compiler turns that into device code; the launch is implicit in how the kernel is invoked. The primitives layer sits on top: the README describes a set of primitives for physics simulation, robotics and geometry processing, and the warp/examples directory is organised into core (dem, fluid, graph capture, marching cubes, mesh, nvdb, raycast, raymarch, sample mesh, sph, torch) and other subdirectories, with tile-based GPU programming represented among the examples. Differentiation is the second axis: because kernels are differentiable, a simulation step can be embedded in a training loop and gradients can flow back through the solver rather than through a hand-derived adjoint. That is the architectural claim worth testing on your own model, since the README does not quantify the cost of the backward pass.

## Installing warp-lang and running a first kernel

Warp publishes wheels on PyPI under the name warp-lang for Windows x86-64, Linux x86-64 and AArch64, and macOS Apple Silicon. The README states that the Windows x86-64 and Linux wheels support CPU execution and CUDA acceleration, and that CUDA acceleration requires a supported NVIDIA GPU and driver. The macOS wheels support CPU execution but not Metal acceleration, which is the single most important platform constraint in the project. The plain install is one command:

```bash
pip install warp-lang
```

If you intend to run the bundled examples or work with USD files, the README gives a second form that pulls the example dependencies:

```bash
pip install warp-lang[examples]
```

On Linux aarch64 systems such as NVIDIA DGX Spark, the README notes that the [examples] extra installs usd-exchange instead of usd-core automatically, because usd-core wheels are not published for that platform. Once installed, the smallest real program is the README's gravity step. It simulates one million particles under mutual gravitational attraction in about twenty lines:

```python
import warp as wp
import numpy as np

num_particles = 1_000_000
dt = 0.01

@wp.kernel
def gravity_step(pos: wp.array[wp.vec3], vel: wp.array[wp.vec3]):
    i = wp.tid()
    position = pos[i]
    dist_sq = wp.length_sq(position) + 0.01  # softened distance
    acc = -1000.0 / dist_sq * wp.normalize(position)  # gravitational pull toward origin
    vel[i] = vel[i] + acc * dt
    pos[i] = pos[i] + vel[i] * dt
```

The host side allocates the arrays from a seeded NumPy generator, so the run is reproducible:

```python
rng = np.random.default_rng(42)
positions = wp.array(rng.normal(size=(num_particles, 3)), dtype=wp.vec3)
velocities = wp.array(rng.normal(size=(num_particles, 3)), dtype=wp.vec3)
```

After that you call gravity_step on the two arrays for as many steps as you need. What you should see is motion toward the origin that accelerates as particles approach it; the 0.01 term in dist_sq keeps the force finite at the centre. The README does not print or plot the result, so the visible output is whatever you choose to read back from the arrays. To see a finished simulation rather than a bare kernel, the README points at the example modules, which are run from the command line as python -m warp.examples.<example_subdir>.<example>, for instance the SPH fluid or marching cubes demos. A handful of examples require a CUDA-capable device, and the README says those are marked at the top of the script. Some examples write USD files containing time-sampled animations into the current working directory, which can be opened in UsdView, Blender or any USD-compatible viewer. There is also python -m warp.examples.browse, which opens the directory where the example sources live.

## Where Warp stops being the right tool

The macOS limitation is not a footnote. Apple Silicon wheels run on the CPU only, so a Mac-based developer gets the programming model and the primitives but none of the acceleration that motivates the framework. If your team standardises on MacBooks and your workload needs GPU throughput, Warp will not deliver it on that hardware, and the README says so plainly rather than promising Metal support. The second boundary is the CUDA toolchain itself. CUDA acceleration requires a supported NVIDIA GPU and driver, and the README defers the exact driver requirements to the Installation Guide rather than listing them inline, so the compatibility check is something you do before you commit, not after. The third boundary is problem shape. Warp is built around kernels over typed arrays, which suits particle systems, mesh operations and grid solvers. A workload that is really a dense matrix multiplication, a convolution stack, or a transformer block is already served by PyTorch or JAX kernels that have been tuned for those shapes; reimplementing them as Warp kernels is work with no obvious payoff. The fourth is the build path. Wheels cover the common platforms, but the repository contains build_lib.py, build_llvm.py, CMakeLists.txt and a deps directory, which tells you that source builds and custom LLVM toolchains exist as a supported path. That is an option, not a requirement, but it is the kind of option that becomes a maintenance obligation if you take it.

## Warp against Taichi, Triton and a plain PyTorch tensor loop

The nearest comparison people search for is Taichi, and the difference is in what ships alongside the compiler. Taichi is a general-purpose kernel language with its own data layout and its own autodiff; Warp pairs a comparable JIT model with a domain library for physics, robotics and geometry, so you get primitives for the simulation rather than only the means to write one. If your problem is a fluid or a rigid-body scene, Warp's example set (sph, fluid, dem, mesh) is the head start; if your problem is a custom stencil with unusual indexing, the two are closer and the choice comes down to which compiler handles your access patterns. Against Triton, the split is different again. Triton targets tensor programs, block-level operations and the kind of kernels that sit inside a neural network, with a programming model built around tiles; Warp's examples include tile-based GPU programming, so the two overlap there, but Warp's centre of gravity is spatial simulation with typed vectors and matrices rather than fused attention or matmul kernels. Against a plain PyTorch loop, the difference is granularity. PyTorch gives you vectorised tensor ops and autograd over them; Warp gives you per-element kernels with an explicit thread index, which is the right level when each element needs branching, neighbour lookups or a small local solve. If your update is expressible as a handful of tensor operations, PyTorch is less code and fewer concepts. The README also notes that Warp kernels work with PyTorch, JAX and Paddle, so this is not strictly either-or: the realistic pattern is a Warp solver step inside a PyTorch training loop.

## Licence, releases and what an upgrade actually costs

Warp is Apache-2.0, declared in pyproject.toml as license = { text = "Apache-2.0" }, in the README badge and in LICENSE.md. That is a permissive licence with an explicit patent grant, which matters for a simulation framework that may end up inside a commercial product. It is not legal advice, and the usual caveat applies: read LICENSE.md and the licences directory, which the repository keeps alongside the source, before you ship. On cadence, the repository is not archived and the last push was on 2026-08-03, the same day as the v1.16.0 release; v1.15.0 landed on 2026-07-07. The project also publishes a separate LLVM SDK artefact, tagged llvm-sdk-22.1.8-warp.1 on 2026-08-07, which reflects the compiler side of the stack rather than the Python package. For upgrade planning, two things in the repository layout are worth reading before you pin a version: CHANGELOG.md at the root and the changelog directory, which together are where API changes are recorded. Because kernels are compiled from Python source at import or first use, a version bump can change compilation behaviour without changing your code, so a pinned warp-lang version plus a CI run of your own kernels is the practical safeguard. The VERSION.md file and the asv directory (airspeed velocity configuration and benchmarks) indicate that performance regressions are tracked upstream, which is useful context when a version bump makes your kernel slower rather than wrong.

## Conclusion

Adopt Warp if your simulation or robotics workload is already CUDA-bound and you want the solver loop expressed in Python with autodiff hooks into PyTorch, JAX or Paddle. Do not adopt it if you need GPU acceleration on macOS, if your team cannot manage CUDA drivers, or if a pure PyTorch tensor formulation already meets your throughput target. Before committing, check the Installation Guide for the CUDA driver requirements and the supported GPU list, confirm which examples run on CPU only, and read the changelog for the 1.15 to 1.16 API changes that touch the code you plan to write.

## FAQ

### How do I install NVIDIA Warp?

The README's recommended route is pip install warp-lang from PyPI, which requires Python 3.10 or newer. Add the [examples] extra, pip install warp-lang[examples], if you want to run the bundled examples or use USD-related features.

### How do I use NVIDIA Warp?

You decorate a Python function with wp.kernel, index it with wp.tid(), and pass wp.array objects typed with dtypes such as wp.vec3. Warp JIT compiles the kernel to CPU or CUDA code, and the README's gravity example shows the whole pattern in about twenty lines.

### What is NVIDIA Warp?

It is a Python framework for GPU-accelerated simulation, robotics and machine learning. It takes regular Python functions, JIT compiles them to kernel code for CPU or GPU, and ships primitives for physics simulation, robotics and geometry processing.

### Is NVIDIA Warp open source?

Yes. The repository is licensed Apache-2.0, as declared in pyproject.toml, the README badge and LICENSE.md, and the source is published on GitHub under NVIDIA/warp.

### Is NVIDIA Warp the same as a CUDA warp?

No. A CUDA warp is a group of threads, which is a hardware concept; NVIDIA Warp is the Python framework described here, and its kernels use wp.tid() for the thread index rather than exposing warp-level primitives directly.

### How does NVIDIA Warp compare with CUDA?

Warp is a Python layer that JIT compiles kernels to CPU or CUDA code, so you write the loop in Python instead of CUDA C++. The README states that CUDA acceleration still requires a supported NVIDIA GPU and driver.

## Sources

- [Official documentation](https://nvidia.github.io/warp/stable/)
- [Official README](https://github.com/NVIDIA/warp#readme)
- [Project repository](https://github.com/NVIDIA/warp)
- [Release notes](https://github.com/NVIDIA/warp/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-warp
