Library / SDK
ai-dynamo/nixl avatar
ai-dynamo/nixl

The wheel ships two CUDA backends and chooses at run time

NVIDIA Inference Xfer Library (NIXL)

1,263 stars451 forksC++NOASSERTION

At a glance

What is it?
NIXL is NVIDIA's transfer layer for disaggregated inference, and its abstraction is memory rather than messages: a plug-in per place a buffer can live, from GPU memory and host memory to POSIX files, object stores and parallel filesystems. The packaging decision worth understanding is that one pip install gives you both the CUDA 12 and CUDA 13 backends and the correct one is selected at run time from the CUDA version PyTorch reports.
Who is it for?
Adopt NIXL if you are building a disaggregated inference stack and the bottleneck is moving key-value cache and activations between devices rather than the compute itself, since a plug-in per memory type is the abstraction you would otherwise write yourself. Do not adopt it for a single-GPU service, where a process-local copy is cheaper than a network transfer, and do not expect it to run on macOS or Windows, since the project supports Linux only.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 10 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 20, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The unit of transfer is a buffer, not a message

The one-sentence description is precise and it is the whole abstraction. NIXL accelerates point-to-point communication in AI inference frameworks such as NVIDIA Dynamo, and it does that by providing an abstraction over various types of memory, such as CPU and GPU, and storage, such as file, block and object store, through a modular plug-in architecture. Note what is absent from that sentence. There is no request, no response, no remote procedure, no session, no serialisation format. The library moves buffers and describes where they are, which means the thing being transferred is a contiguous piece of data that already has meaning to both ends, and the meaning is somebody else's problem. That is the right level for an inference stack, where the expensive thing is not the message but the memory traffic, and where a framework usually ends up writing its own ad-hoc transfer layer for exactly this reason.

One plug-in per place a buffer can live

The plug-in list is the best single description of the infrastructure the library expects to meet, and the ROCm section names almost all of it. UCX is the primary transport. POSIX is local files. OBJ is object storage, and AZURE_BLOB is a specific object store. HF3FS is a parallel filesystem, MOONCAKE is a storage library, GUSLI is another storage component, and UCCL is a collective communication library. GDS and its multi-threaded variant are GPU data storage paths, GPUNETIO is a network device path, and LIBFABRIC is a fabric library. Each of those is a different way for a buffer to be addressed and a different set of costs to move it. The build system exposes them as a comma-separated list with two mutually exclusive switches, one to include plugins and one to exclude them, and the README is explicit that the two cannot be combined. There is also a separate option for building a set of them statically, which is what you want when the plugin has to be present in a container image rather than discovered at load time.

pip install nixl gives you two backends and a run-time choice

The pre-built distribution section is short and the mechanism is unusual. The Python API and libraries, including UCX, are on PyPI, and a single command installs them:

bash
pip install nixl

What that command gives you is both the CUDA 12 and the CUDA 13 backends in one artefact, and the correct one is selected automatically at run time based on the CUDA version reported by PyTorch. That is a deliberate design choice with a clear trade. A wheel that carries two backends works on both a 12 host and a 13 host without anyone having to know which CUDA is on the machine before installing, and the cost is a larger download and a wheel that is not as small as it could be. For a container image that gets rebuilt across a fleet, that is the right side of the trade. The build system also has a wheel variant option that overrides the detection, and the named example is a ROCm variant, where passing the variant option yields a wheel with a different name so a ROCm build and a CUDA build can live side by side in an index.

Vendor-neutral, and the limits are written down

Most projects claiming vendor neutrality do not say where it stops. This one does, in a paragraph that is worth reading twice. NIXL itself builds vendor-neutrally, and CPU-side hardware detection, named as a count of AMD GPUs, discovers AMD cards by PCI vendor identifier whether or not a ROCm toolchain is present, which means detection is not conditional on the build environment. GPU-side ROCm and HIP build support is narrower, being available for the benchmark tool and for the UCX plug-in unit tests rather than for the library as a whole. And then the plug-in matrix is given explicitly: UCX is the primary transport for AMD GPU memory and needs UCX itself built with ROCm support; POSIX, OBJ, AZURE_BLOB, HF3FS, MOONCAKE, GUSLI and UCCL are vendor-neutral and build unchanged; and the CUDA-specific plug-ins, the GPU data storage paths, the device path and the fabric library, are skipped automatically on a host with no CUDA. For a project whose main risk is being quietly CUDA-only, publishing that table is the strongest thing in the README.

Linux only, stated in one line

The supported platforms section is the shortest section in the README and it settles the most questions. NIXL is supported in a Linux environment only, tested on Ubuntu 22.04 and 24.04 and on Fedora, and macOS and Windows are not currently supported, with the instruction to use a Linux host or a container or virtual machine. For a 2026 C++20 library that is an unusual position, and it follows from what the thing is built on: the transport is UCX, the fast paths involve GPU data storage, and the whole point is a multi-GPU serving stack. The practical consequence for a reader is simple and worth stating before anyone falls in love with the abstraction: if you develop on a Mac, every experiment happens inside a virtual machine, and the feedback loop for a data-movement library is the part of the work that suffers most. The build prerequisites reinforce the point, with a C++20 compiler, so GCC 11 or newer or Clang 14 or newer, plus the ordinary build tools and a Python set including meson, ninja, pybind11 and tomlkit.

The fast path needs two projects built from source

If you build NIXL from source rather than taking the wheel, the cost profile changes, and the README is upfront about it. NIXL was tested with UCX version 1.23.x, and the documentation includes the full configure line for building that version from its repository, with shared libraries enabled, static ones disabled, documentation disabled, optimisations on, automatic CPU management on, development headers on, and CUDA, verbs, device memory and GDRCopy pointed at your installations. Then it is built, installed with a stripped install, and registered with the dynamic linker cache. GDRCopy is called out separately and honestly: it is available on GitHub and necessary for maximum performance, but UCX and NIXL will work without it. That is the correct way to phrase an optional accelerator, because it tells you the difference between a working installation and a fast one. The library's own build is then four commands, a meson setup, a change into the build directory, ninja, and ninja install, with a debug variant and a set of documented options for docs, UCX path, headers, plug-in selection and CUDA paths.

ETCD is optional, and only for distributed coordination

One optional dependency appears in the prerequisites and it is worth understanding why it is optional. NIXL can use ETCD for metadata distribution and coordination between nodes in distributed environments, which is a cluster problem rather than a transfer problem. The README gives both routes for the server, a package install of the server and client, or a container image on the standard ETCD port, and then a separate build of the ETCD C++ API from source, which needs the gRPC and C++ REST development packages and a CMake build. That is a real amount of machinery, and it is only needed if you are running more than one node. The scoping is the interesting part, because it tells you the library's own model: the data transfer is a local, in-process concern with a plug-in per memory type, and only the coordination of metadata is a distributed-system concern, delegated to something that already solves it. If you are running a single node, none of that is yours.

Three attribution files and a licence with parts

The licensing deserves a close read, because the signals disagree in an informative way. Every source file in the repository carries an SPDX header identifying NVIDIA copyright and naming the Apache 2.0 licence, and the README badge says Apache 2.0. The Python packaging metadata says something different and more specific, with the licence field set to MIT and Apache-2.0 together, and a licence-files list that includes the repository licence file and a second file under a licenses directory with a proprietary name. The repository's own licence metadata records a custom licence that could not be classified. Put together, the reasonable reading is that most of the code is under familiar open terms, that the distribution as a whole includes at least one component that is not, and that anyone redistributing it has to work out which is which rather than assuming the badge. The repository makes that work possible with three attribution files, one each for the C++, Python and Rust trees, which is an unusual amount of provenance bookkeeping and suggests the language bindings did not all arrive from the same place. There is also a security policy, a code of conduct, a codeowners file and a Rust workspace alongside the meson build.

Editorial conclusion

Adopt NIXL if you are building a disaggregated inference stack and the bottleneck is moving key-value cache and activations between devices rather than the compute itself, since a plug-in per memory type is the abstraction you would otherwise write yourself. Do not adopt it for a single-GPU service, where a process-local copy is cheaper than a network transfer, and do not expect it to run on macOS or Windows, since the project supports Linux only. Verify four things: that you are on a supported distribution, since testing covers Ubuntu 22.04 and 24.04 and Fedora, that your UCX version is in the range the project tested, which is 1.23.x, whether you need the extra performance GDRCopy provides, since the README says the fast path wants it and things work without it, and which licence covers which file, because the packaging declares MIT and Apache-2.0 together and lists a proprietary licence file among them, so read the licence files before you redistribute. The last release is v1.4.1 from 2026-09-01, the packaging metadata on main already says 1.5.0, and the last push was 2026-09-19.

Frequently asked questions

What is NIXL?

NVIDIA Inference Xfer Library, a library for accelerating point-to-point communication in AI inference frameworks such as NVIDIA Dynamo. It abstracts over memory types such as CPU and GPU and over storage such as file, block and object store, through a modular plug-in architecture.

Which platforms does NIXL support?

Linux only. The README says it is tested on Ubuntu 22.04 and 24.04 and on Fedora, and that macOS and Windows are not currently supported, so you should use a Linux host, or a container or virtual machine. A C++20 compiler is required, meaning GCC 11 or newer or Clang 14 or newer.

How does the NIXL Python wheel choose a CUDA backend?

The wheel installs both the CUDA 12 and the CUDA 13 backends, and the correct one is selected automatically at run time based on the CUDA version reported by PyTorch. A build option overrides that detection, and the named example is a ROCm variant which produces a wheel with a different name.

Which transfer backends can NIXL use?

The plug-in list includes UCX, POSIX, OBJ, AZURE_BLOB, HF3FS, MOONCAKE, GUSLI, UCCL, GDS, GDS_MT, GPUNETIO and LIBFABRIC. On a host without CUDA the vendor-neutral ones build unchanged and the CUDA-specific ones are skipped automatically, with UCX as the primary transport for AMD GPU memory.

Do I need ETCD to run NIXL?

No, it is optional and only for metadata distribution and coordination between nodes in distributed environments. Single-node use does not need it. If you do want it, the README covers both a packaged server and a container, plus a build of the ETCD C++ API from source.

What licence is NIXL released under?

The source files carry SPDX headers naming Apache 2.0, but the Python packaging declares the licence as MIT and Apache-2.0 together and lists a proprietary licence file among its licence files, so read the licence files before redistributing. The repository also keeps separate attribution files for the C++, Python and Rust trees. The last release is v1.4.1 from 2026-09-01.

Official sources

  1. ai-dynamo/nixl on GitHub
  2. Issues
  3. README
  4. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/ai-dynamo-nixl.svg)](https://hysenlabs.com/projects/ai-dynamo-nixl)