# NVIDIA/stdexec: the C++26 sender/receiver reference implementation

> stdexec is a header-only reference implementation of std::execution (P2300), the C++26 model for asynchronous and parallel programming. It is aimed at C++ engineers who want lazy, composable sender pipelines today, on CPUs, io_uring, or NVIDIA GPUs.

**NVIDIA/stdexec** — `std::execution`, the standard C++ framework for asynchronous and parallel programming.

- Repository: https://github.com/NVIDIA/stdexec
- Website: https://nvidia.github.io/stdexec/
- Stars: 2,443 · Forks: 274
- Language: C++
- License: Apache-2.0
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/nvidia-stdexec

## The problem stdexec addresses for C++ teams

Asynchronous C++ has historically been split across incompatible idioms: std::thread and condition variables, std::future and std::async, callbacks, and third-party executors. Each has a different composition story, and mixing them produces code where cancellation, error propagation and lifetime are handled ad hoc. stdexec is a reference implementation of std::execution, the model tracked by P2300 for C++26, and it targets exactly that gap: a single vocabulary for asynchronous and parallel work that composes the same way whether it runs on a thread pool, a Linux io_uring context, or a GPU. The README describes the library as letting you express asynchronous work as composable, lazy sender pipelines with structured concurrency guarantees. The audience is C++ library authors and application engineers on C++20 or later who are willing to work against an evolving standard. It is not a runtime, not a scheduler service, and not a replacement for a thread pool you already operate.

## How sender pipelines actually run

The mechanism is lazy composition. A sender describes work but does not start it. Algorithms such as then, let_value, when_all, bulk, split and transfer build a pipeline graph, and nothing executes until a consumer connects to it. The README example builds three squares with when_all over ex::on(sched, ex::just(i) | ex::then(fun)) and only then calls ex::sync_wait, which launches the work and blocks for the result. That separation is what makes the model testable and movable between execution contexts: the scheduler is a parameter of the pipeline, not a property of the algorithm. The repository also ships structured concurrency primitives such as async_scope, task, finally, when_any and repeat_n, which the README lists alongside the core algorithms. GPU support is layered on the same abstraction through <nvexec/...> schedulers: nvexec::stream_scheduler in <nvexec/stream_context.cuh> targets device 0, and nvexec::multi_gpu_stream_scheduler in <nvexec/multi_gpu_context.cuh> spans all visible devices. The same pipeline shape is reused; only the scheduler changes. Coroutine interop goes both ways, with senders awaitable and awaitables usable as senders, according to the README. Extensions that are not yet in the standard live under <exec/...>.

## Installing stdexec and running a first pipeline

The library is header-only with no external dependencies, so the lightest integration is a CMake target. The README recommends CPM, which fetches the repository and configures it from your CMakeLists.txt. The snippet below is the quick start the README gives, with the CPM include line it references.

```cmake
cmake_minimum_required(VERSION 3.25.0)
project(stdexec_example LANGUAGES CXX)

include(CPM.cmake)

CPMAddPackage(
  NAME stdexec
  GITHUB_REPOSITORY NVIDIA/stdexec
  GIT_TAG main
)

add_executable(example example.cpp)
target_link_libraries(example PRIVATE STDEXEC::stdexec)
```

Linking STDEXEC::stdexec matters: the README notes that using the CMake target is recommended because it sets the required compile flags. If you prefer not to use CMake, clone the repository and add the include directory directly, since header-only means the include path is enough. The README also documents Conan through the provided conanfile.py, and the NVIDIA HPC SDK path, where stdexec is bundled with nvc++ starting with NVHPC SDK 22.11 and is enabled with --experimental-stdpar, plus -stdpar=gpu for GPU features.

```bash
git clone https://github.com/NVIDIA/stdexec.git
```

With the target linked, the README's own example is the shortest real program. It schedules three squares on the parallel scheduler and prints their results.

```c++
#include <stdexec/execution.hpp>
#include <cstdio>

namespace ex = stdexec;

int main() {
    auto sched = ex::get_parallel_scheduler();
    auto fun   = [](int i) { return i * i; };

    auto work = ex::when_all(ex::on(sched, ex::just(0) | ex::then(fun)),
                             ex::on(sched, ex::just(1) | ex::then(fun)),
                             ex::on(sched, ex::just(2) | ex::then(fun)));

    auto [i, j, k] = ex::sync_wait(std::move(work)).value();
    std::printf("%d %d %d\n", i, j, k); // prints "0 1 4"
}
```

Compile it with -std=c++20 or later, as the README requires. The expected output is 0 1 4. For GPU work the README points to the nvexec examples and requires the nvc++ compiler at version 25.9 or newer. A larger set of runnable programs sits in examples/, including hello_world.cpp, hello_coro.cpp, scope.cpp, io_uring.cpp, sudoku.cpp and the server_theme/ directory, which the README describes as server-style patterns using let_value, split, bulk and transfer.

## Where stdexec is the wrong choice

The README carries an explicit warning: stdexec is experimental, tracks an evolving standard, and APIs may change without notice, with NVIDIA stating it does not guarantee fitness for any particular purpose. That is not boilerplate to skim past. If you are building a long-lived binary that must compile unchanged for years, a moving API surface is a real cost, and the repository publishes no release history that the README documents, so pinning to a stable version is not something the README helps you plan. Compiler support is another hard boundary: GCC 12, Clang 16, MSVC 14.43, Xcode 16, and nvc++ 25.9 for GPU. Teams on older toolchains are simply excluded. The README also states that stdexec does not yet support NVIDIA's nvcc compiler, so a CUDA codebase built around nvcc cannot adopt the GPU schedulers without moving to nvc++. Finally, the model itself has a learning cost. Sender pipelines are not callbacks with nicer syntax; readers unfamiliar with the vocabulary of senders, receivers and completion signatures will not be productive in an afternoon, and the README points to external documentation rather than teaching the concepts inline.

## stdexec compared with libunifex and std::execution::par

The closest relative is libunifex, the earlier Facebook implementation of the same sender/receiver ideas. Both provide lazy pipelines and scheduler abstraction, and both predate the C++26 wording. The practical difference is alignment: stdexec is described in its README as a reference implementation of std::execution and P2300, so its names and semantics track the committee's direction, while libunifex explored a related design that fed into that work. If your goal is to write code that survives into the standard, stdexec is the one whose API is explicitly tied to the standard being written. The other comparison people reach for is std::execution::par, the parallel execution policy passed to algorithms like std::for_each. That is a different layer entirely: a policy tells an existing algorithm to parallelize its loop, with no composition, no cancellation model and no scheduler abstraction. stdexec's bulk algorithm covers the parallel-loop case inside a sender pipeline, but the reason to use stdexec is everything around it: when_all for fan-out, let_value for dependent continuations, transfer to move work between contexts, and async_scope for structured lifetimes. If you only need to parallelize a loop over a container, the policy is far less machinery.

## Maintenance, licensing and upgrade cost

The repository is not archived, and the last push was on 2026-09-28, so work is ongoing. That does not change the experimental status the README states, and it does not guarantee that any specific API you adopt today will look the same after the next standard meeting. Budget for periodic recompilation against a newer checkout rather than assuming a frozen interface. The licence is Apache-2.0, and the README's badge text specifies Apache 2.0 with LLVM-exception; the LICENSE.txt file at the repository root is the authoritative text. The LLVM exception matters to some organisations because it relaxes certain conditions when combining with code under other licences, but whether that fits your situation is a question for your own legal review, not something this article can settle. Practical upgrade cost is dominated by compiler support rather than the library: because it is header-only, an upgrade is a new include path or a new CPM tag, but a new tag can require a newer compiler than the one you have, and the GPU path additionally requires a matching nvc++ release.

## Conclusion

Adopt stdexec if you are writing C++20 or later and want to learn or prototype the sender/receiver model before it lands in your standard library, or if you need NVIDIA GPU schedulers from nvc++. Do not adopt it as a stable dependency for a long-lived production service: the README states it is experimental and that APIs may change without notice. Verify first that your compiler meets the minimum (GCC 12, Clang 16, MSVC 14.43, Xcode 16, nvc++ 25.9 for GPU), that you can build with -std=c++20 or later, and that nvcc is not in your path for this code, since the README says stdexec does not yet support it.

## FAQ

### What is stdexec from NVIDIA?

It is a reference implementation of std::execution, the C++26 model for asynchronous and parallel programming tracked by P2300. The README describes it as header-only with no external dependencies, offering composable sender pipelines that run on threads, thread pools, GPUs or a custom execution context.

### Which compilers does stdexec support?

The README lists GCC 12, Clang 16, MSVC 14.43, Xcode (Apple Clang) 16, and nvc++ 25.9 for GPU support, with -std=c++20 or later required. It also notes that stdexec does not yet support NVIDIA's nvcc compiler.

### How do I add stdexec to a CMake project?

The README recommends CPM: call CPMAddPackage with NAME stdexec and GITHUB_REPOSITORY NVIDIA/stdexec, then link STDEXEC::stdexec to your target. It notes that using the CMake target is recommended because it sets the required compile flags.

### How does stdexec run work on a GPU?

GPU schedulers ship in the <nvexec/...> headers and require the nvc++ compiler. nvexec::stream_scheduler in <nvexec/stream_context.cuh> targets device 0, while nvexec::multi_gpu_stream_scheduler in <nvexec/multi_gpu_context.cuh> spans all visible devices.

### Is stdexec production ready?

The README states that stdexec is experimental and tracks an evolving standard, that APIs may change without notice, and that NVIDIA does not guarantee fitness for any particular purpose. Treat it as a way to work with the std::execution model now rather than as a frozen dependency.

## Sources

- [Issues](https://github.com/NVIDIA/stdexec/issues)
- [License: Apache-2.0](https://github.com/NVIDIA/stdexec/blob/main/LICENSE)
- [NVIDIA/stdexec on GitHub](https://github.com/NVIDIA/stdexec)
- [Project website](https://nvidia.github.io/stdexec/)
- [README](https://github.com/NVIDIA/stdexec/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/nvidia-stdexec
