# NVIDIA/cccl: Thrust, CUB and libcudacxx in a single header-only repository

> CCCL unifies NVIDIA's three CUDA C++ libraries into one header-only repository that ships with the CUDA Toolkit. It is a build-system decision before it is an API decision.

**NVIDIA/cccl** — CUDA Core Compute Libraries

- Repository: https://github.com/NVIDIA/cccl
- Website: https://nvidia.github.io/cccl/
- Stars: 2,523 · Forks: 507
- Language: C++
- License: NOASSERTION
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-cccl

## What CCCL replaces, and why the merge matters more than the APIs

Thrust, CUB and libcudacxx existed as three separate repositories with overlapping goals. The README describes the unification as the first step toward filling the role the standard C++ library fills for standard C++: general-purpose tools so CUDA developers can focus on the problem rather than the plumbing. The repository now holds thrust/, cub/, libcudacxx/ and cudax/ as top-level directories, alongside examples/, benchmarks/, test/ and a ci/ tree.

The practical effect is versioning. Previously, a project that needed CUB's block primitives and Thrust's algorithms tracked two release cadences and hoped they agreed. Now one tag covers all three, and the releases listed for the repository (v3.4.2 and v3.4.1, plus a separate python-1.2.0 line) are the unit of upgrade. The audience is narrow and specific: people writing CUDA C++ kernels who want either a high-level algorithm layer, a low-level block-scoped primitive layer, or both in the same translation unit.

## Three layers, one include path: how the libraries divide the work

The split is by altitude, not by domain. Thrust is the C++ parallel algorithms library, with configurable backends that allow using CUDA, TBB or OpenMP, which is what gives it performance portability between GPUs and multicore CPUs. CUB sits lower and is CUDA-specific, aimed at device-wide algorithms plus cooperative algorithms such as block-wide reduction and warp-wide scan, so kernel authors get building blocks rather than a whole algorithm. libcudacxx implements the C++ standard library for both host and device code and adds CUDA-specific abstractions for synchronization primitives, cache control and atomics.

The README's example shows how the layers compose. A hand-written kernel uses cub::BlockReduce for the per-block sum, then cuda::atomic_ref from libcudacxx with cuda::memory_order_relaxed to accumulate across blocks, and the same reduction is then run through thrust::reduce. The two results are compared and asserted equal. That structure is the clearest statement of intent in the repository: the high-level algorithm and the hand-written kernel are expected to agree, so you can prototype with Thrust and drop to CUB when the profile demands it.

## Installing CCCL and building a first reduction

Everything in CCCL is header-only, so there is nothing to link. The README states that the easiest route is the CUDA Toolkit, which includes the CCCL headers; when you compile with nvcc it adds them to the include path automatically, so a plain include works with no extra configuration. If you compile with another compiler, you point the include search path at the toolkit's headers, for example /usr/local/cuda/include.

```cpp
#include <thrust/device_vector.h>
#include <cub/cub.cuh>
#include <cuda/std/atomic>
```

Those three includes are the ones the README gives for the toolkit route. If you compile with nvcc, nothing else is needed. If you use a different compiler, add the toolkit include directory to your build system's search path first.

The README also documents a GitHub route for users who want to stay on the cutting edge of CCCL development. The compatibility rule it states is asymmetric: using a newer version of CCCL with an older CUDA Toolkit is supported, but not the other way around. The repository ships a CMakeLists.txt, CMakePresets.json and a cmake/ directory, and top-level examples/ with subdirectories such as examples/basic/ and examples/thrust_flexible_device_system/.

The smallest real use is the reduction from the README's example. A device_vector is filled with ones, a custom kernel reduces it with cub::BlockReduce and cuda::atomic_ref, and thrust::reduce produces the same value through the high-level path. The README points to a live Godbolt link for the full program, so you can read the complete source before compiling anything locally.

## The header-only model has a versioning trap

Header-only removes link steps and ABI questions, but it moves the risk into the compiler and toolkit pairing. Because the headers come from the toolkit by default, the CCCL version you compile against is decided by whichever CUDA Toolkit is installed, not by your project's manifest. A team that upgrades its toolkit silently upgrades Thrust, CUB and libcudacxx at the same time.

The README's compatibility statement is the constraint to internalize: newer CCCL with an older toolkit is supported, the reverse is not. That means pinning a GitHub checkout is the safe direction and pinning an old CCCL against a new toolkit is not. The README does not document a rollback procedure for a CCCL upgrade, and it does not describe how to detect a version mismatch at configure time. If your build matrix spans several toolkit versions, that gap is where you will spend your time.

There is a second boundary worth stating plainly. CCCL is CUDA C++. A host-only CPU program that wants parallel algorithms has no reason to include these headers, and a project targeting a non-NVIDIA accelerator is outside the scope of what this repository provides.

## Choosing between CCCL and a portability layer

The honest alternative is a portability framework rather than another CUDA library. Thrust's configurable backends already let the same algorithm run on CUDA, TBB or OpenMP, which covers CPU fallback within one codebase. If the requirement is running the same kernel source on non-NVIDIA hardware, that is a different kind of tool, and CCCL does not attempt it: CUB is described as CUDA-specific and tuned for GPU architectures, and libcudacxx is the CUDA C++ standard library.

The distinction is where the abstraction sits. A portability layer abstracts the device and the runtime. CCCL abstracts the algorithm and the primitive while assuming CUDA underneath. If your constraint is vendor neutrality, CCCL is the wrong layer to reach for. If your constraint is getting the most out of NVIDIA hardware without hand-writing every scan and reduce, the three-library split gives you a documented path from thrust::reduce down to cub::BlockReduce without leaving the same include path.

## Maintenance, licensing and the cost of staying current

The repository is not archived, and the last push was on 2026-09-28. Releases are frequent enough that the listed set includes v3.4.2 and v3.4.1 on the same day, plus a separate python-1.2.0 line, so the C++ libraries and the Python packaging move on different schedules. The top-level tree carries .clang-format, .clang-tidy, .clangd, .pre-commit-config.yaml and a pyproject.toml configuring ruff, mypy and codespell, which tells you the project enforces its own style on contributions. That matters to you only if you vendor headers and patch them.

The licence field reports NOASSERTION, and the repository has a LICENSE file at the top level. Because GitHub could not classify it, read that file yourself before shipping, particularly if you redistribute headers inside a product. Nothing here is legal advice, and the licence implications of vendoring headers differ from those of linking a binary, so treat the LICENSE file as the source of truth rather than the repository metadata.

The upgrade cost is the toolkit coupling described above. If you take CCCL from the toolkit, upgrades arrive with CUDA Toolkit upgrades and you inherit whatever else that brings. If you take it from GitHub, you own the pin, and the README's compatibility rule tells you which direction is safe.

## Conclusion

Adopt CCCL if you already compile CUDA C++ and want Thrust's parallel algorithms, CUB's block and warp primitives, and libcudacxx's host-and-device standard library from one include path. Do not adopt it if you are not writing CUDA C++ at all: there is nothing here for a host-only CPU program, and OpenCL or a portable runtime is a different decision. Before you build anything, confirm which CCCL version your CUDA Toolkit carries, because the README states that a newer CCCL with an older toolkit is supported while the reverse is not, and that asymmetry decides whether you can pin a GitHub checkout or must stay on the toolkit's headers.

## FAQ

### Is OpenCL better than CUDA?

The README does not compare CCCL to OpenCL. It describes CCCL as CUDA C++, with CUB described as CUDA-specific and libcudacxx as the CUDA C++ standard library, so the repository does not position itself as a cross-vendor alternative.

### Why would I need CUDA?

The README frames the goal as giving CUDA C++ developers building blocks that make it easier to write safe and efficient code, with Thrust offering parallel algorithms, CUB offering block-wide and warp-wide primitives, and libcudacxx offering host-and-device standard library facilities.

### Is OpenCL still used?

The README says nothing about OpenCL. It mentions CUDA, TBB and OpenMP as the backends Thrust's configurable execution policies can target.

### Which is better for Premiere Pro, OpenCL or CUDA?

The README does not discuss Premiere Pro or any video editing application. CCCL is a set of CUDA C++ libraries for writing accelerated code, not a runtime setting for creative software.

## Sources

- [Issues](https://github.com/NVIDIA/cccl/issues)
- [NVIDIA/cccl on GitHub](https://github.com/NVIDIA/cccl)
- [Project website](https://nvidia.github.io/cccl/)
- [README](https://github.com/NVIDIA/cccl/blob/main/README.md)
- [Releases](https://github.com/NVIDIA/cccl/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-cccl
