NVSHMEM: GPU-initiated one-sided communication for CUDA clusters
NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmers to perform one-sided communication from within CUDA kernels and on CUDA streams.
At a glance
- What is it?
- NVSHMEM gives CUDA kernels and streams a partitioned global address space across GPUs, so communication starts inside the kernel rather than through a host-side library. It is for teams already running multi-GPU jobs on data center hardware.
- Who is it for?
- Adopt NVSHMEM if your kernels need one-sided puts, gets, atomics or signals issued from device code across two or more data center GPUs, and you already have a process launcher such as Open MPI, Hydra or Slurm in place. Do not adopt it for single-GPU work, for host-only coordination, or on hardware outside the release notes' supported platform list, because the prerequisites are two NVIDIA data center GPUs and a Linux system.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 7 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The host round trip NVSHMEM removes
The problem NVSHMEM addresses is the cost of moving data between GPUs when the decision to move it lives on the CPU. In a conventional setup a kernel finishes, control returns to the host, a communication library is called, and a second kernel picks up the result. NVSHMEM instead exposes a partitioned global address space across NVIDIA GPUs and lets threads inside a CUDA kernel, or work submitted on a CUDA stream, issue one-sided transfers, atomics, signaling and synchronization directly. The README describes the goal as reducing multi-process communication and coordination overheads by allowing programmers to perform one-sided communication from within CUDA kernels and on CUDA streams.
The audience is narrow and specific. You need a Linux system, an NVIDIA driver and a CUDA Toolkit supported by the current release, two NVIDIA data center GPUs, and a compatible process launcher such as Open MPI, Hydra or Slurm. That list rules out laptops, single-GPU workstations and any workflow where the communication pattern is decided on the host. If your job is one GPU with host-driven copies, NVSHMEM adds a build dependency and a launch requirement without changing anything you care about.
Symmetric memory, PEs and three interface layers
The architecture rests on symmetric memory: each processing element, or PE, allocates buffers at corresponding addresses, and remote access works because every PE knows the layout. That is what makes one-sided operations possible without a matching receive on the far side.
NVSHMEM exposes three interface layers. The host interface is called from CPU threads. The device interface is called from within CUDA kernels. The stream interface is ordered against CUDA stream work. The README states that CPU or GPU threads can initiate communication, which is the practical difference from a library that only accepts host-side calls.
Because the layers are separable, the build reflects them. Installed packages provide imported CMake targets named nvshmem::nvshmem_host and nvshmem::nvshmem_device. Applications that use only host-side or stream-ordered APIs can link only the host target, which keeps the device library out of a build that does not need it. Bootstrap is a separate concern: the perftest commands set NVSHMEM_BOOTSTRAP or NVSHMEM_BOOTSTRAP_PMI so that initialization matches whichever launcher is present. The repository layout shows the split clearly, with perftest/, examples/, src/, nvshmem4py/ and separate symbol files for host, transport and bootstrap.
Installing NVSHMEM and running the put bandwidth test
The README points at the NVSHMEM product page for packages and at the installation guide for the procedure. After installing, set NVSHMEM_PREFIX to the installation directory and put its lib directory on the loader path:
export NVSHMEM_PREFIX=/path/to/nvshmem
export LD_LIBRARY_PATH="$NVSHMEM_PREFIX/lib${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}"The verification step launches two PEs, one per GPU. With Open MPI the README gives this command, which runs the put bandwidth perftest over message sizes from 8 bytes to 1M:
NVSHMEM_BOOTSTRAP=MPI \
mpirun -np 2 \
"$NVSHMEM_PREFIX/bin/perftest/device/pt-to-pt/shmem_put_bw" \
--min_size 8 --max_size 1M --use_smem 0A successful run reports both GPUs, prints put bandwidth for a range of message sizes, and exits with status zero. If initialization fails, the README says to rerun with NVSHMEM_DEBUG=INFO and consult the launching guide. Hydra and Slurm variants are documented with NVSHMEM_BOOTSTRAP_PMI=PMI and NVSHMEM_BOOTSTRAP_PMI=PMI-2 respectively.
To build from source instead, the README gives a standard CMake sequence that clones the repository, sets NVSHMEM_PREFIX to a local install directory, configures, builds in parallel and installs:
git clone https://github.com/NVIDIA/nvshmem.git
cd nvshmem
export NVSHMEM_PREFIX="$PWD/install"
cmake -S . -B build \
-DNVSHMEM_PREFIX="$NVSHMEM_PREFIX" \
-DCMAKE_INSTALL_PREFIX="$NVSHMEM_PREFIX"
cmake --build build --parallel
cmake --install buildThe README warns that the development branch may be newer than the latest published release and its documentation, and that a release-aligned build should check out the corresponding branch or tag first. That warning matters because the default branch is devel.
For a first application, link against the imported targets:
find_package(NVSHMEM REQUIRED CONFIG)
target_link_libraries(my_target PRIVATE
nvshmem::nvshmem_host
nvshmem::nvshmem_device
)The README names examples/put-block.cu as a complete CUDA program and points Python users at the NVSHMEM4Py quick start in nvshmem4py/README.md.
Where NVSHMEM is the wrong tool
The prerequisites are the first limitation, and they are hard. Two NVIDIA data center GPUs and a supported driver and CUDA Toolkit are not negotiable; the README lists them as prerequisites rather than recommendations. A single-GPU application cannot use the interface at all in any meaningful way, because there is no remote PE to address.
Launch is the second constraint. NVSHMEM programs need a compatible process launcher, and the bootstrap mechanism differs between them: MPI, PMI-1 and PMI-2 are separate code paths selected by environment variables. A cluster whose scheduler integration is unusual, or whose MPI build does not expose the expected PMI interface, will fail at initialization rather than at compile time. The README directs you to NVSHMEM_DEBUG=INFO and the launching guide when that happens, which is a reasonable debugging path but not a fix.
The third limitation is documentation scope in the repository itself. The README is a pointer document: it links to the installation guide, launching guide, release notes, environment variable reference, best practices, API reference and Using NVSHMEM guide on docs.nvidia.com. It does not restate supported platform versions or known issues. Anyone deciding whether a specific driver and CUDA combination is supported has to open the release notes, and the README does not document rollback or downgrade procedures if a new release misbehaves on an existing cluster.
NVSHMEM against NCCL, MPI and OpenSHMEM
The comparison people reach for is NVSHMEM versus NCCL, and the difference is the programming model rather than the transport. NCCL is a collective library: you call it, typically from the host, and it executes an all-reduce, all-gather or similar operation across ranks. NVSHMEM exposes a partitioned global address space and one-sided operations, so a kernel can put data into a peer's symmetric buffer, signal, and continue, without a collective call and without the peer participating in a matching operation. If your workload is a dense collective such as an all-reduce in a training loop, NCCL is the direct fit and NVSHMEM is a detour. If your workload needs irregular point-to-point traffic issued from inside device code, collectives are the wrong primitive.
Against MPI the distinction is similar but sharper. MPI is a message-passing interface where the host process drives communication and the receiver typically matches the sender. NVSHMEM is OpenSHMEM-based: one-sided, symmetric-memory, and callable from device code. The repository ships examples for both worlds, including mpi-based-init.cu and shmem-based-init.cu, which shows that MPI is expected to remain present for initialization even when NVSHMEM handles the data movement.
Against OpenSHMEM itself, NVSHMEM is the GPU extension of that model. OpenSHMEM targets CPU processes in a partitioned global address space; NVSHMEM keeps the interface shape and adds CUDA kernel and CUDA stream entry points. That lineage is why the API names and environment variable conventions look familiar to OpenSHMEM users.
Release cadence, licence and the cost of staying current
The release history is dense. v3.7.0-0 was published on 2026-06-11, v3.7.1-0 on 2026-06-30, and v3.7.2-0 on 2026-07-17, roughly one release every two to three weeks across that window. The default branch is devel and the last push was on 2026-08-27. A cadence that fast means the upgrade cost is real: each release carries its own supported-platform matrix in the release notes, and the README explicitly warns that the development branch may be ahead of the latest published release and its documentation. Pinning to a tag rather than tracking devel is the documented way to keep code and documentation aligned.
On licensing, NVSHMEM is Apache-2.0, and the repository carries both License.txt and a NOTICE file. Apache-2.0 permits commercial use and modification and includes a patent grant, but it also imposes notice and attribution obligations when redistributing, and the NOTICE file exists for that purpose. The README does not describe any additional terms for the prebuilt packages distributed from the product page, so a team that ships a product containing NVSHMEM should read the package terms separately rather than assuming they match the repository licence. This is a description of what the files say, not legal advice.
Contributions go through pull requests and require a Developer Certificate of Origin sign-off with git commit -s, per the contributing guide referenced in the README.
What to check before committing to NVSHMEM
Start with the release notes, not the README. The README states that supported platforms, compatibility and known issues live there, and the prerequisites are defined by the current release rather than by the project in general. Confirm your driver and CUDA Toolkit appear in that matrix before anything else.
Then confirm your launcher. The three documented paths are Open MPI with NVSHMEM_BOOTSTRAP=MPI, Hydra with NVSHMEM_BOOTSTRAP_PMI=PMI, and Slurm with NVSHMEM_BOOTSTRAP_PMI=PMI-2. If your site uses something else, the README offers no fourth option, and the launching guide is the place to look.
Finally, read examples/put-block.cu and the Using NVSHMEM guide before writing your own kernel. The README points at that example as a complete CUDA program, and the guide covers the programming model, application lifecycle, compilation and launch requirements. Committing to NVSHMEM without understanding the application lifecycle is the most common way to end up debugging initialization instead of your algorithm.
Editorial conclusion
Adopt NVSHMEM if your kernels need one-sided puts, gets, atomics or signals issued from device code across two or more data center GPUs, and you already have a process launcher such as Open MPI, Hydra or Slurm in place. Do not adopt it for single-GPU work, for host-only coordination, or on hardware outside the release notes' supported platform list, because the prerequisites are two NVIDIA data center GPUs and a Linux system. Verify first that your driver and CUDA Toolkit match the current release notes, then run the shmem_put_bw perftest with two PEs and confirm it exits with status zero before writing any application code.
Frequently asked questions
What is NVSHMEM?
It is an OpenSHMEM-based parallel programming interface from NVIDIA that provides a partitioned global address space across NVIDIA GPUs. Applications use symmetric memory for one-sided transfers, atomics, signaling, synchronization and collective operations, with host, CUDA kernel and CUDA stream interfaces.
How to install NVSHMEM?
Download a package from the NVSHMEM product page and follow the installation guide, or build from source with a CMake configure, build and install sequence after cloning the repository. Either way, set NVSHMEM_PREFIX to the installation directory and add its lib directory to LD_LIBRARY_PATH.
How to use NVSHMEM?
Link your target against the imported CMake targets nvshmem::nvshmem_host and nvshmem::nvshmem_device, then launch your program with a compatible process launcher such as Open MPI, Hydra or Slurm. The README points to examples/put-block.cu as a complete CUDA program and to the Using NVSHMEM guide for the programming model and launch requirements.
Is NVSHMEM open source?
Yes. The repository is licensed under the Apache License 2.0 and carries a NOTICE file alongside License.txt.
What is the difference between NVSHMEM and MPI?
MPI is a message-passing interface where host processes drive communication and receivers typically match senders. NVSHMEM is OpenSHMEM-based, one-sided and callable from CUDA kernels and streams, and the repository ships both mpi-based-init.cu and shmem-based-init.cu, suggesting MPI is often kept for initialization while NVSHMEM handles data movement.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-nvshmem)