Library / SDK
openucx/ucc avatar
openucx/ucc

UCC: a collective API that sits above UCX and below MPI

Unified Collective Communication Library

314 stars131 forksCBSD-3-Clause

At a glance

What is it?
Unified Collective Communication is a BSD-3-Clause C library from the OpenUCX project that exposes one collective API across CPU, GPU and DPU transports. It is for HPC and AI runtime authors who need collectives without writing transport-specific code.
Who is it for?
Adopt UCC if you are building a runtime or framework that needs one collective API across InfiniBand, RoCE, shared memory, CUDA or NCCL backends, and if you are willing to build UCX first. Do not adopt it if you only need collectives inside a single MPI job and your MPI implementation already performs well, because adding UCC means another library in the path.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 13 days ago.
What is it written in?
Mainly C, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap UCC fills between a transport and a programming model

MPI has collectives. NCCL has collectives. Frameworks such as PyTorch have their own. Each of those is tied to a runtime, a transport or a vendor. UCC's premise is that the collective operation itself is a separate concern from the programming model that calls it and from the transport that moves the bytes.

The README frames this as an API and library for "current and emerging programming models and runtimes". The design goals list nonblocking collective operations that cover a variety of programming models, a flexible resource allocation model, a relaxed ordering model, a flexible synchronous model, and repetitive collective operations that are initialized once and invoked multiple times. Hardware collectives, meaning in-network reduction engines, are described as a first-class citizen rather than an optimization bolted on later.

The audience is therefore not an application developer writing a solver. It is the person building the layer underneath: an MPI implementation, an OpenSHMEM implementation, a deep learning framework, or an in-house runtime that has to run the same allreduce on a CPU cluster, a GPU node and a DPU-equipped fabric without three code paths.

The UCC component architecture and what actually moves data

The repository ships a component diagram at docs/images/ucc_components.png, and the supported transport list is the clearest statement of the data flow. UCC does not implement network transports itself. It reaches the fabric through UCX/UCP, which the README lists as covering InfiniBand, RoCE, Cray Gemini and Aries, and shared memory. On top of that it lists SHARP, CUDA, NCCL, RCCL and MLX5 as supported transports.

That layering is the whole design. A collective call enters UCC, UCC selects a collective algorithm and a transport, and the transport does the work. SHARP is the interesting entry: it names Mellanox's in-network reduction, which is what the design goal about hardware collectives being first-class refers to. A reduction can complete inside the switch rather than at every rank.

The dependency on UCX is not optional in the way the GPU dependencies are. The README states plainly that UCC uses utilities provided by UCX's UCS component. CUDA and HIP are marked optional and only matter if you want GPU collectives. So the minimum viable UCC is UCX plus a C toolchain, and everything else is a transport you either have or do not.

Building UCC from source and running a first MPI job

There is no package manager line in the README. The documented path is a source build with autotools. The developer's build runs autogen.sh, then configure with a prefix and a pointer to your UCX installation, then make.

bash
./autogen.sh
./configure --prefix=<ucc-install-path> --with-ucx=<ucx-install-path>
make

The two placeholders are the part people get wrong. --prefix is where UCC itself lands. --with-ucx must point at an existing UCX install, not a source tree, because UCC links against UCX's UCS utilities.

If you do not already have UCX, the README gives the full sequence. It clones UCX, builds and installs it, then clones UCC and configures against the UCX prefix.

bash
git clone https://github.com/openucx/ucx
cd ucx
./autogen.sh; ./configure --prefix=<ucx-install-path>; make -j install
bash
git clone https://github.com/openucx/ucc
cd ucc
./autogen.sh; ./configure --prefix=<ucc-install-path> --with-ucx=<ucx-install-path>; make -j install

Documentation is a separate build. Configuring with --with-docs-only and running make docs generates the API reference through Doxygen, which the README lists as a required package for that purpose.

The first real use in the README is not a standalone UCC program. It is Open MPI, built with both UCX and UCC, then run with UCC's collective component enabled through MCA parameters.

bash
mpirun -np 2 --mca coll_ucc_enable 1 --mca coll_ucc_priority 100 ./my_mpi_app

Setting coll_ucc_priority to 100 pushes UCC ahead of the other collective components in Open MPI's selection order. If the run completes and the collectives behave, UCC is in the path. To confirm it rather than assume it, you would need Open MPI's component verbosity output, which the README does not document here. OpenSHMEM has the parallel form, with scoll_ucc_enable and scoll_ucc_priority instead of coll_ucc_enable and coll_ucc_priority.

Where UCC is the wrong layer to reach for

The most common mismatch is treating UCC as a drop-in replacement for NCCL in a PyTorch training script. UCC is a C library with a C API. It is not a Python package, and the README shows no pip install, no Python binding and no framework integration code. If your requirement is a one-line change in a training loop, UCC is not that.

The second constraint is the build. UCC is configured and built from source with autotools, and it needs a UCX installation that exposes UCS. That means no binary distribution path is described in the README, and no version compatibility matrix is given either. The README does not state which UCX versions are supported against which UCC release, so a mismatch between the two is a build-time or link-time problem you discover yourself.

The third is that GPU support is conditional at compile time. CUDA collectives require CUDA 11.0 or above according to the README, and HIP support requires a ROCm/HIP installation. A UCC built without those flags will not give you GPU collectives, and nothing in the README suggests a runtime fallback that would tell you so.

Finally, the Open MPI integration is opt-in through MCA parameters. Enabling coll_ucc_enable does not guarantee UCC handles every collective. Component selection in Open MPI depends on priority and on whether a component can service a given operation, and the README documents the two parameters but not the selection rules.

How UCC differs from NCCL and from MPI's own collectives

NCCL is the obvious comparison because UCC lists it as a supported transport. The difference in approach is directional. NCCL is a collective library built for NVIDIA GPUs, with the GPU as the assumed execution target and the network as something it drives. UCC treats NCCL as one backend among several, alongside CUDA, RCCL, SHARP and UCX/UCP. A UCC user on an AMD node and a UCC user on an NVIDIA node can call the same collective API while the transport underneath differs.

The comparison with MPI's built-in collectives is subtler. UCC is not competing with MPI as a programming model. It is competing with the collective implementations inside MPI. The README's own integration path makes this explicit: Open MPI is configured with --with-ucc and then selects UCC through coll_ucc_enable. UCC replaces the algorithm, not the API the application sees.

That is a real architectural difference from both. NCCL asks you to adopt its API and its hardware assumptions. MPI collectives ask you to accept whatever algorithm your MPI build chose. UCC asks you to accept a build dependency on UCX in exchange for choosing the algorithm and transport per collective, including in-network reduction through SHARP where the fabric supports it.

Maintenance, licensing and upgrade cost

The last push to the default branch was on 2026-09-08, the same day v1.9.0 was tagged. v1.8.0 came on 2026-06-03, and a v1.9.0-rc1 preceded the release on 2026-08-12. The repository is not archived. That cadence, roughly a quarterly minor release with a release candidate ahead of it, is what the release list shows.

Upgrade cost is dominated by the UCX relationship rather than by UCC's own API. Because UCC links against UCX's UCS utilities and the README does not publish a compatibility matrix, moving to a new UCC release means checking the UCX version you build against. A cluster where UCX is pinned by a vendor stack may not be able to follow UCC's release cadence.

On licensing: UCC is BSD-style licensed per the README, and the LICENSE file is BSD-3-Clause. That is permissive and generally compatible with shipping in a larger product, but the README also notes that contributors must comply with the UCF Consortium's membership and export-compliance policies. If you plan to contribute patches rather than only consume the library, read CONTRIBUTING.md and those policies before your first pull request. This is not legal advice; the LICENSE file is the authoritative text.

Editorial conclusion

Adopt UCC if you are building a runtime or framework that needs one collective API across InfiniBand, RoCE, shared memory, CUDA or NCCL backends, and if you are willing to build UCX first. Do not adopt it if you only need collectives inside a single MPI job and your MPI implementation already performs well, because adding UCC means another library in the path. Before committing, verify that your UCX build exposes the UCS utilities UCC needs, that your GPU stack matches the CUDA or HIP version you intend to compile against, and which of the supported transports is actually available on your fabric.

Frequently asked questions

What is UCC (Unified Collective Communication) used for?

It provides a collective communication operations API and library for programming models and runtimes, covering HPC, AI/ML and I/O workloads. It is used by MPI and OpenSHMEM implementations and by runtimes that need collectives across CPU, GPU and DPU transports.

How do I install and build UCC?

The README documents a source build: run ./autogen.sh, then ./configure --prefix=<ucc-install-path> --with-ucx=<ucx-install-path>, then make. UCX must already be installed because UCC uses utilities from UCX's UCS component. CUDA and HIP are optional and only needed for GPU collectives.

Does UCC work with Open MPI, and how do I enable it?

Yes. Open MPI is configured with --with-ucc=<ucc-install-path> alongside --with-ucx, and UCC's collective component is then enabled at run time with the MCA parameters coll_ucc_enable set to 1 and coll_ucc_priority set to 100. OpenSHMEM uses the corresponding scoll_ucc_enable and scoll_ucc_priority parameters.

Which transports and hardware does UCC support?

The README lists UCX/UCP (covering InfiniBand, RoCE, Cray Gemini and Aries, and shared memory), SHARP, CUDA, NCCL, RCCL and MLX5. Hardware collectives are described as a first-class citizen in the design goals, which is what the SHARP transport relates to.

Official sources

  1. License: BSD-3-Clause
  2. openucx/ucc on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/openucx-ucc.svg)](https://hysenlabs.com/projects/openucx-ucc)