NVIDIA NCCL: Building and Installing the Multi-GPU Collective Library
Optimized primitives for collective multi-GPU communication
At a glance
- What is it?
- NCCL is NVIDIA's stand-alone library of collective communication routines for GPUs, covering all-reduce, all-gather, reduce, broadcast and reduce-scatter over PCIe, NVLink, NVSwitch, InfiniBand Verbs or TCP sockets. This article covers how it is built and packaged, where the boundaries of the design sit, and what to check before you depend on it.
- Who is it for?
- Adopt NCCL when your workload is a collective over GPUs that share PCIe, NVLink or NVSwitch, or that talk over InfiniBand Verbs or TCP, and when you can either consume the official build from developer.nvidia.com or reproduce the make targets on your own toolchain. Do not adopt it as a general CPU message-passing layer, and do not assume the repository ships a runnable test suite: tests live at github.com/nvidia/nccl-tests.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 13 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What NCCL solves, and who is actually supposed to use it
Training and inference jobs that span more than one GPU need the same handful of operations over and over: sum gradients across ranks, gather shards, scatter results back. Writing those by hand means writing topology-aware transport code for PCIe, NVLink, NVSwitch, InfiniBand Verbs and plain TCP sockets, and then keeping it correct as hardware changes. NCCL is the library that already contains those routines. The README describes it as a stand-alone library of standard communication routines for GPUs implementing all-reduce, all-gather, reduce, broadcast and reduce-scatter, plus any send/receive based pattern.
The intended audience is narrow in a useful way. The README states that NCCL supports an arbitrary number of GPUs in a single node or across multiple nodes, and that it can be used in single-process or multi-process applications, MPI being the example given. So the reader is someone writing C or C++ CUDA code, or a framework author wiring a collective into a training loop, who is willing to reason about rank counts and transport. It is not a user-facing tool. There is no CLI to learn and no config file to tune; the API is the product.
One thing the README does not do is explain when a single GPU or a single process makes the library pointless. That judgement is left to the reader, and it is the first decision to get right.
How the collectives get from your call to the wire
The repository layout tells you more about the architecture than the README does. The top level holds src/, plugins/, bindings/, pkg/, makefiles/, cmake/ and contrib/. The build is driven by a thin root Makefile that dispatches into those directories: src.% recurses into src, pkg.% into pkg, nccl4py.% into bindings/nccl4py, and ir.% into bindings/ir. That split matters because it means the core library, the packaging layer, the Python binding and the IR emission path are separate build products with separate targets, not one monolithic compile.
The transport story is spelled out in the README rather than inferred: NCCL is optimized for high bandwidth on platforms using PCIe, NVLink and NVSwitch, and for networking over InfiniBand Verbs or TCP/IP sockets. Those are the layers the library selects among. The plugins/ directory exists at the top level, which is consistent with a design where transport and other behaviours can be extended outside the core tree, though the README does not document the plugin interface.
The IR path is the least visible part. The root Makefile defines EMIT_LLVM_IR and NCCL_EMIT_LTO_IR, both defaulting to 0, and gates llvm_ir and ltoir behind them. Setting either to a non-zero value adds a goal to the default target. The README says nothing about what these artifacts are for, so treat that as an internal build concern rather than a documented feature.
Installing NVIDIA NCCL: official builds versus make targets
The README is explicit that the official and tested builds are downloaded from developer.nvidia.com/nccl, and that you can skip the build steps entirely if you use those. That is the shortest path and the one most users should take. The rest of this section covers the source route, which is what the repository documents.
To build the library from a checkout, change into the directory and run the build target. The README gives this exact form:
$ cd nccl
$ make -j src.buildIf CUDA is not in the default /usr/local/cuda path, the CUDA location is passed on the command line. The README shows the variable name and the placeholder:
$ make src.build CUDA_HOME=<path to cuda install>The output lands in build/ unless BUILDDIR is set, which the README states directly and the Makefile confirms with its BUILDDIR ?= $(abspath ./build) default. By default NCCL compiles for all supported architectures. To cut compile time and binary size, the README suggests redefining NVCC_GENCODE, which lives in makefiles/common.mk, and gives a single-architecture example:
$ make -j src.build NVCC_GENCODE="-gencode=arch=compute_90,code=sm_90"Installing on the system means building a package and installing it as root. Debian and Ubuntu users install the packaging tools first, then build the deb. The README lists the tools by name:
$ sudo apt install build-essential devscripts debhelper fakeroot
$ make pkg.debian.build
$ ls build/pkg/deb/RedHat and CentOS follow the same shape with rpm-build and rpmdevtools, and an OS-agnostic tarball is available through make pkg.txz.build, which writes to build/pkg/txz/. There is also a Python wheel target, make pkg.python_wheel.build, which the README notes also builds the .txz archive as an intermediate; it requires uv, installed via the curl command the README gives. The Makefile shows that pkg.debian.prep and pkg.txz.prep depend on the lic target, so the licence file is copied into the build directory before packaging.
For a first real use, the README points elsewhere: tests are maintained separately at github.com/nvidia/nccl-tests. The documented sequence clones that repository, builds it, and runs the all-reduce benchmark with a start size, end size, step factor and GPU count:
$ git clone https://github.com/NVIDIA/nccl-tests.git
$ cd nccl-tests
$ make
$ ./build/all_reduce_perf -b 8 -e 256M -f 2 -g <ngpus>Replace the GPU count placeholder with the number of devices you actually have. If the binary runs and reports bandwidth for each size, your NCCL install is reachable from the test harness. If it fails at initialization, the problem is almost always the library path or the visible device set, not the collective itself.
Where NCCL is the wrong tool
The clearest limitation is that NCCL is GPU-specific. Every routine in the README's list is described as a communication routine for GPUs. If your workload is CPU-to-CPU across nodes, NCCL adds a GPU dependency and a CUDA requirement to solve a problem it was not designed for. That is not a defect, it is the boundary.
The second limitation is the test story. The README states plainly that tests are maintained separately at github.com/nvidia/nccl-tests. Nothing in the repository itself gives you a first-party suite to run after a build. That means a source build is verified by an external project, and any mismatch between your NCCL version and the nccl-tests revision is yours to diagnose. For a library that sits under distributed training, that is a real gap, and it is worth knowing before you plan a CI pipeline around it.
The third is packaging weight. Installing on the system requires creating a package and installing it as root, and the Debian path pulls in build-essential, devscripts, debhelper and fakeroot. If you only need a shared library inside a container image, that full packaging route is more machinery than the task requires, and the tarball target is the lighter option.
Finally, the README does not document rollback, version pinning or upgrade procedures. The release list shows versioned tags such as v2.31.2-1, but nothing in the README describes how to move between them safely. Plan for that yourself.
NCCL versus MPI, and what the Python binding changes
The comparison people reach for is MPI, and the difference is not subtle. MPI is a general message-passing standard that spans CPUs and accelerators and gives you the full vocabulary of point-to-point and collective operations plus process management. NCCL implements a fixed set of collective primitives optimized for GPU bandwidth across PCIe, NVLink, NVSwitch, InfiniBand Verbs and TCP, and the README frames MPI as a possible host for a multi-process NCCL application rather than as a competitor. In practice the two are often layered: MPI launches and coordinates the ranks, NCCL moves the tensors. If your job is CPU-heavy with occasional GPU work, MPI alone is the simpler answer. If your job is gradient exchange across eight GPUs on one node, MPI's collectives will not touch the NVLink bandwidth NCCL is built for.
The bindings/ directory and the nccl4py release tags (nccl4py-v0.5.0 and nccl4py-v0.4.1 are both in the recent release list) indicate a Python-facing path exists alongside the C++ core. The README documents the wheel build target but does not describe the Python API, so anyone planning to use NCCL from Python should treat the binding as a separate surface to evaluate rather than assuming the C++ documentation transfers. The same caution applies to the IR targets gated behind EMIT_LLVM_IR and NCCL_EMIT_LTO_IR: they exist in the build system, and the README does not explain them.
Licence, maintenance and the cost of staying current
The repository's licence field is NOASSERTION, and the source files carry SPDX headers. The root Makefile states SPDX-License-Identifier: Apache-2.0 and points to LICENSE.txt for more license information, and the README's copyright line reads 2015-2020 NVIDIA CORPORATION. Those two signals do not fully agree on scope, and the README does not resolve it. If you are redistributing NCCL inside a product, read LICENSE.txt and ThirdPartyNotices.txt directly and get your own legal review; nothing here substitutes for that.
On maintenance, the repository is not archived and the last push was on 2026-09-10, which is recent. Releases are versioned and frequent enough to be worth tracking: v2.31.2-1, nccl4py-v0.4.1 and nccl4py-v0.5.0 all appear within roughly a month of each other in the release list. The practical upgrade cost is the mismatch surface. A new NCCL build has to line up with your CUDA toolkit, your GPU architecture set (the NVCC_GENCODE variable controls which architectures are compiled in), your framework's expected library version, and your nccl-tests revision. Building for all supported architectures avoids the architecture mismatch but lengthens compile time and grows the binary, which is exactly the trade-off the README flags when it suggests narrowing NVCC_GENCODE.
The cheapest maintenance posture is to consume the official builds from developer.nvidia.com/nccl and pin the version your framework expects. The source route is for people who need a custom architecture set, a specific package format, or a wheel built against their own Python environment.
Editorial conclusion
Adopt NCCL when your workload is a collective over GPUs that share PCIe, NVLink or NVSwitch, or that talk over InfiniBand Verbs or TCP, and when you can either consume the official build from developer.nvidia.com or reproduce the make targets on your own toolchain. Do not adopt it as a general CPU message-passing layer, and do not assume the repository ships a runnable test suite: tests live at github.com/nvidia/nccl-tests. Before you commit, verify three things on your own hardware: that the make target for your package format produces a package under build/pkg/, that all_reduce_perf from nccl-tests runs at the GPU count you plan to use, and that your CUDA toolkit version matches the build path you chose. The repository was last pushed on 2026-09-10 and is not archived, but the README does not document rollback or upgrade procedures, so the upgrade path is something you will have to establish yourself.
Frequently asked questions
What is NVIDIA NCCL?
It is a stand-alone C++ library of standard communication routines for GPUs, implementing all-reduce, all-gather, reduce, broadcast and reduce-scatter plus any send/receive based pattern. It is optimized for high bandwidth over PCIe, NVLink, NVSwitch, InfiniBand Verbs and TCP/IP sockets.
What does the acronym NCCL stand for?
The README only gives the pronunciation, "Nickel", and does not spell out the acronym. The repository description calls it optimized primitives for collective multi-GPU communication.
Does NVIDIA NCCL use NVLink?
Yes. The README states that NCCL has been optimized to achieve high bandwidth on platforms using PCIe, NVLink and NVSwitch, as well as networking using InfiniBand Verbs or TCP/IP sockets.
How do I install NVIDIA NCCL?
The README says the official and tested builds can be downloaded from developer.nvidia.com/nccl, and that you can skip the build steps if you use those. From source, you run make -j src.build and then build a package with a target such as make pkg.debian.build or make pkg.txz.build.
Is NVIDIA NCCL open source?
The source code is in a public repository, and the root Makefile carries an SPDX-License-Identifier of Apache-2.0 pointing to LICENSE.txt. The repository's licence field itself reads NOASSERTION, so check LICENSE.txt and ThirdPartyNotices.txt for the terms that apply to you.
How does NVIDIA NCCL compare to MPI?
MPI is a general message-passing standard, while NCCL is a GPU collective library, and the README presents multi-process NCCL applications as something that can run under MPI rather than as a replacement for it. In practice MPI often handles rank launch and coordination while NCCL moves the tensors.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-nccl)