UCCL: a drop-in NCCL/RCCL replacement with its own network transport
UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)
At a glance
- What is it?
- UCCL is a C++ GPU communication library covering collectives, P2P transfers and expert parallelism. It keeps the NCCL API but replaces the transport underneath, and ships as a Python wheel that loads as a net plugin.
- Who is it for?
- Adopt UCCL if you run multi-node PyTorch training or inference on H100, T4, MI300X or EFA-backed instances and you want to swap the collective transport without touching model code, since the README describes it as a drop-in replacement for NCCL/RCCL. Do not adopt it if you need a stable ABI across many Python versions, since pyproject.toml requires Python 3.12 or newer, or if you need TPU or Trainium support, which the roadmap still lists as unchecked.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What UCCL replaces, and for whom
Most multi-GPU training stacks sit on NCCL or RCCL, and most teams treat that layer as fixed. UCCL's premise is that the collective library is worth re-architecting while the API stays the same. The README calls UCCL-collective a drop-in replacement for NCCL/RCCL that requires no changes to application code, which is the whole adoption argument: you keep your training loop and change an environment variable.
The library covers three workloads rather than one. UCCL-collective handles the standard collectives. UCCL-P2P provides NIXL-style initiator-target transfer APIs, aimed at KV cache movement and reinforcement learning weight transfer. UCCL-EP targets expert parallelism, letting DeepEP run on heterogeneous hardware. The audience is infrastructure engineers running multi-node training or inference who have hit a wall with the default transport, not application developers who want a new programming model. If you have never looked at NCCL_NET_PLUGIN, this project is not aimed at you.
Packet spraying: the transport change that matters
The design claim in the README is specific. Existing transports under NCCL, kernel TCP and RDMA, stream large volumes over one or a few network paths, which makes them prone to congestion in datacenter networks. UCCL-collective instead implements packet spraying in software so traffic spreads across the available paths. The README lists four consequences: packet spraying with 256 paths, congestion control that can be latency-based or receiver-driven, loss recovery by selective repeat, and usability in public clouds with legacy NICs and Ethernet.
That last point is the practical one. A transport that assumes modern RDMA hardware excludes a lot of deployments, including cloud instances where you get whatever NIC the provider offers. The README's own performance examples reflect this spread: on six HGX servers with 8x400G CX-7 RoCE NICs and 8xH100 GPUs it reports up to 2.5x over NCCL for AllReduce, and on two AWS g4dn.8xlarge instances with 1x50G ENA NICs and 1xT4 GPUs it reports up to 3.7x. Those are the project's figures, not independent measurements, and the second case is the interesting one because it is a modest cloud instance rather than a flagship cluster.
The trade-off is that software packet spraying and custom congestion control move work that hardware or a mature kernel path would otherwise absorb. The README does not document the CPU cost of that decision, nor how the 256-path spraying behaves when the fabric is small.
Building UCCL and loading it as a net plugin
The README's quick start is a build script rather than a plain pip install, because the build depends on your accelerator stack. The script detects the Python version of the current environment; you can pass a specific one such as 3.10.
git clone https://github.com/uccl-project/uccl.git && cd uccl
# Eg, bash build.sh cu12 ep --install
bash build.sh [cu12|cu13|roc7|roc6|therock] [all|ccl_rdma|ccl_efa|p2p|ep] \
[py_version] [rocm_index_url] --installThe first positional argument picks the platform. By default cu12 targets CUDA 12.8 and roc7 targets ROCm 7.1; cu13 and roc6 target CUDA 13.0 and ROCm 6.4. The second picks the component to build. If you are building for ROCm through TheRock, the default index url is https://rocm.prereleases.amd.com/whl/gfx94X-dcgpu and the README warns it may not be what you want.
Once built, you do not call UCCL from your training code. You export a plugin path and let NCCL pick it up. For NCCL over IB/RoCE on x86 or GH200 ARM hosts:
NCCL_NET_PLUGIN=`python -c "import uccl; print(uccl.nccl_plugin_path())"`The README gives the same pattern with uccl.rccl_plugin_path() for RCCL over IB/RoCE on x86 hosts. For AWS EFA NICs, which the README limits to p4d and p4de, two variables are needed:
LD_PRELOAD=`python -c "import uccl; print(uccl.efa_nccl_path())"`
NCCL_NET_PLUGIN=`python -c "import uccl; print(uccl.efa_plugin_path())"`After that the README says you can run your PyTorch applications as usual. The repository ships examples/ddp_train.py and examples/ddp_run.sh if you want a DDP job to point the variables at. The thing to check before trusting any of this is that the printed path exists and is non-empty; the README does not describe what happens when the plugin path resolves to nothing.
Python 3.12, wheel variants and what that costs you
pyproject.toml sets requires-python to ">=3.12", and the single runtime dependency is intervaltree. That is a narrow surface, but the version floor is real. setup.py explains the reason: nanobind's stable ABI requires Python 3.12 or newer. On 3.12 and above the build emits a cp312-abi3 wheel, one wheel for all 3.12+ interpreters. On older Pythons the stub extension still exists to force a platform-specific wheel tag, but without the limited-API flag, producing a version-specific cpXY-cpXY wheel.
The packaging is also split by backend. setup.py describes a single uccl package with variants distinguished by PEP 440 local version identifiers, for example uccl-0.1.0+cu13 or uccl-0.1.0+cu12.efa. The default cu12 build carries no local version and is the one published to PyPI; every other variant is distributed through GitHub Releases. So pip install uccl gets you CUDA 12 and nothing else. If you need cu13, EFA, or a ROCm build, you are downloading a wheel file from a release page and installing it by filename, which is a different and more manual upgrade path than the PyPI one.
For TheRock builds the README adds an extra step: pass the index url to pip and add the [rocm] extra to the wheel, for example pip install --extra-index-url https://rocm.prereleases.amd.com/whl/gfx94X-dcgpu wheelhouse-therock/uccl-0.0.1.post4-py3-none-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl[rocm]. Note the version in that example, 0.0.1.post4, does not match the 0.1.1 in pyproject.toml; the README example lags the release it documents.
Where UCCL is the wrong tool
The roadmap is the honest limit. TPU and Trainium support is listed as an unchecked item under expert-parallel communication, so if your accelerators are not AMD or Nvidia GPUs, the project does not claim to run there. The same roadmap shows SM-efficient communication kernels and fine-grained compute-communication overlapping as in progress, and device kernels in vendor-agnostic Triton as unchecked. Collectives on consumer GPUs such as 4090, 5090 and GB10 are also marked in progress, which means the README's headline comparisons do not describe those cards yet.
The second limit is operational. Loading UCCL means exporting an environment variable that changes the network transport for every collective in the process. A bug in the plugin surfaces as a hung or corrupted collective inside your training job, and the README does not document a rollback procedure, a health check, or a way to fall back to the stock transport for a single job. Compare that to a library you call explicitly, where you can wrap one transfer and leave everything else alone. The plugin model buys you zero code changes and charges you a much harder failure mode to isolate.
Finally, the performance claims are the project's own, from its technical report and README, on specific hardware. Nothing in the repository tells you what the gap looks like on your fabric.
UCCL against NCCL, and against NIXL-style transfer libraries
The comparison the project invites is UCCL versus NCCL. The difference is not the API, which is deliberately identical, but the layer underneath. NCCL's network transports use kernel TCP and RDMA and, per the README, push large volumes over one or a few paths. UCCL replaces that with a software transport doing packet spraying across up to 256 paths, plus its own congestion control and selective-repeat loss recovery. So the two are interchangeable at the call site and not interchangeable in behaviour: congestion handling, path selection and loss recovery all move from the existing stack into UCCL's code.
UCCL-P2P is a different kind of comparison. It offers NIXL-style initiator-target transfer APIs, so it competes with point-to-point transfer libraries rather than with collectives, and the README states it is designed for next-gen 800Gbps NICs with multi-threaded transfer engines. That is a narrower target than UCCL-collective's. UCCL-EP is narrower still: it exists so DeepEP can run on AMD and Nvidia GPUs and on RDMA NICs including AWS EFA and Broadcom, at what the README calls IBGDA-level performance. If you are not running DeepEP, that component is not relevant to you. Treat the three as separate decisions, because they have separate hardware assumptions.
Maintenance, licence and upgrade cost
The repository is not archived. The last push was on 2026-05-10, which is the same date as the v0.1.1 release; the previous release, v0.0.1.post6, was on 2026-03-14. Two releases in roughly two months, then nothing pushed since May. That is a young project with a burst of activity rather than a long track record, and the README's TheRock example still referencing 0.0.1.post4 is a small sign of documentation drift.
Upgrading is not uniform across backends. The PyPI path is a normal pip upgrade of the cu12 build. The GitHub Releases path means tracking wheel filenames that encode the backend in a local version identifier, and the ROCm path additionally requires the right index url and the [rocm] extra. Budget for that if you run anything other than the default CUDA 12 build.
The licence is Apache-2.0, declared both in the repository LICENSE file and in pyproject.toml under the license field. Apache-2.0 is permissive and includes an explicit patent grant, which matters for a library that plugs into vendor networking stacks, but it also carries attribution and notice requirements. Whether those obligations affect your distribution is a question for your own legal review, not something the repository answers.
Editorial conclusion
Adopt UCCL if you run multi-node PyTorch training or inference on H100, T4, MI300X or EFA-backed instances and you want to swap the collective transport without touching model code, since the README describes it as a drop-in replacement for NCCL/RCCL. Do not adopt it if you need a stable ABI across many Python versions, since pyproject.toml requires Python 3.12 or newer, or if you need TPU or Trainium support, which the roadmap still lists as unchecked. Verify first that your exact NIC and GPU combination appears in the build matrix, then confirm the plugin path resolves with python -c "import uccl; print(uccl.nccl_plugin_path())" before you point NCCL_NET_PLUGIN at it.
Frequently asked questions
What is UCCL?
UCCL is an efficient communication library for GPUs covering collectives, P2P transfers such as KV cache and RL weight transfer, and expert parallelism. It is written in C++ with Python bindings and is licensed Apache-2.0.
How do I install UCCL?
The README's quick start clones the repository and runs build.sh with a platform argument such as cu12 or roc7 and a component argument such as ep, plus --install. The default cu12 build is also published to PyPI as the uccl package; other variants are distributed through GitHub Releases.
Does UCCL work as a drop-in replacement for NCCL?
The README describes UCCL-collective as a drop-in replacement for NCCL/RCCL that requires no changes to application code. In practice you set NCCL_NET_PLUGIN to the path returned by uccl.nccl_plugin_path() and run your PyTorch job unchanged.
Which Python versions does UCCL support?
pyproject.toml sets requires-python to >=3.12. setup.py explains that nanobind's stable ABI needs Python 3.12 or newer, which is why 3.12+ gets a single cp312-abi3 wheel while older Pythons would get version-specific wheels.
Does UCCL support AMD GPUs?
Yes. The build script accepts roc7 and roc6 targets for ROCm, and the roadmap marks AMD GPU support as done for KV cache transfer and for expert-parallel communication. The README also states UCCL has been adopted as part of the AMD TheRock ecosystem.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/uccl-project-uccl)