Library / SDK
deepseek-ai/DeepEP avatar
deepseek-ai/DeepEP

DeepEP V2: the ElasticBuffer rewrite, the NCCL Gin backend, and what it costs you in buffer size

Project brief: DeepEP: an efficient expert-parallel communication library. Example use in model training or inference prefilling V2 unifies the dispatch and combine APIs into a single ElasticBuffer interface.

10,223 stars1,456 forksCudaMIT

At a glance

What is it?
DeepEP is DeepSeek's expert-parallel all-to-all library for MoE dispatch and combine. V2 replaces NVSHMEM with NCCL Gin and merges the high-throughput and low-latency APIs into one ElasticBuffer, at the price of a larger memory footprint.
Who is it for?
Adopt DeepEP V2 if you run MoE training or inference prefilling on SM90 or SM100 GPUs with NVLink inside the node and RDMA between nodes, and you can afford the larger intermediate buffers the V2 notes call out. Do not adopt it if you are on non-NVIDIA accelerators, if you still rely on the 0 SM RDMA low-latency EP path from V1, or if you need a stable, documented pipeline-, context- or Engram-parallel API, since the README labels those experimental.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Cuda, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What DeepEP solves, and which teams it is actually for

Mixture-of-experts models split their experts across GPUs. Every token has to reach the experts that were selected for it, and the results have to come back. That is two all-to-all exchanges per MoE layer, and at scale they dominate the step time. DeepEP exists to make those two exchanges fast: the README describes it as providing high-throughput and low-latency all-to-all GPU kernels for MoE dispatch and combine, with low-precision support including FP8.

The audience is narrow and specific. You need Hopper (SM90) GPUs or another architecture with SM90 PTX ISA support, NVLink for intranode communication and an RDMA network for internode communication. The requirements list PyTorch 2.10 and above, NCCL 2.30.4 and above, and CUDA 12.3 and above for SM90. If your cluster is not NVIDIA, or your interconnect is plain Ethernet without RDMA, this library is not aiming at you. The README also mentions experimental primitives for pipeline parallelism, context parallelism and remote memory access under the name Engram, all designed for zero or minimal SM occupation. Those are side features; the core product is expert parallelism.

One detail matters for anyone evaluating the install cost: all kernels are compiled at runtime through a Just-In-Time module, so the README states no CUDA compilation is required during installation. That changes what a deployment looks like. You are not shipping prebuilt kernels for a fixed set of architectures; the first run pays the compile cost, and the cache directory is configurable through environment variables such as EP_JIT_CACHE_DIR, which appears in setup.py's list of persistent environment names.

V2 architecture: one ElasticBuffer, NCCL Gin instead of NVSHMEM

The V2 release is described in the README as a complete refactoring of expert parallelism, and the two structural changes are worth separating.

First, the backend. V1 used NVSHMEM. V2 switched to what the README calls the more lightweight NCCL Gin backend, which it describes as header-only and able to reuse existing NCCL communicators. Reusing existing communicators is the part that matters operationally: you are not standing up a second communication library alongside the one your training framework already initialized. NVSHMEM has not disappeared from the dependency list, though. The README says DeepEP still depends on NVSHMEM to provide support for legacy methods, and points to docs/nvshmem.md for installation. So a V2 install is not a smaller dependency graph than V1 in every respect, even if the hot path no longer runs through NVSHMEM.

Second, the API. In V2 the high-throughput and low-latency paths are unified under a single ElasticBuffer interface with a new GEMM layout. The README's example constructs the buffer from MoE settings directly: number of max tokens per rank, hidden dimension, top-k, expert count, and a use_fp8_dispatch flag. SM and QP counts are calculated analytically rather than auto-tuned, which the release notes frame as removing a tuning step. Whether that is a win depends on your situation. Analytical calculation is predictable and reproducible, but it also means the library decides the resource split for you, and the README does not describe an override for the SM count beyond noting it is set at buffer creation.

The scale claim is up to EP2048, with both hybrid and direct modes still supported. For V3-like legacy training, the README states SM usage drops from 24 to 4 to 6 while maintaining equivalent or better performance. Treat those numbers as the project's own figures from its stated test configuration, not as a guarantee for your workload.

Installing DeepEP V2 and running the elastic EP test

The README recommends installing NCCL through pip so DeepEP can locate it inside the Python environment. The command pins a minimum version:

bash
pip install "nvidia-nccl-cu13>=2.30.4" --no-deps

The --no-deps flag is deliberate; it keeps pip from pulling in a dependency set that might conflict with the CUDA stack you already have. After that, the README says DeepEP also depends on NVSHMEM for legacy methods and directs you to docs/nvshmem.md rather than giving an inline command.

For a development checkout, the README gives a build step followed by a symbolic link whose name you may need to adjust for your platform:

bash
python setup.py build
ln -s build/lib.linux-x86_64-cpython-38/deep_ep_cpp.cpython-38-x86_64-linux-gnu.so

The path in that link encodes both the platform tag and the CPython version, which is why the README tells you to modify the SO name. For a normal installation the README gives a single command:

bash
python setup.py install

Then import deep_ep in your Python project. The first real use is the elastic EP test. The README warns explicitly that you may need to modify the init_dist function in tests/utils/envs.py according to your own cluster settings before launching across multiple nodes:

bash
python tests/elastic/test_ep.py
python tests/elastic/test_agrs.py
python tests/elastic/test_engram.py
python tests/elastic/test_pp.py

Run test_ep.py first. It exercises the dispatch and combine path on your actual topology, and a failure there usually means a version or interconnect problem rather than a bug in your model code.

The limitations the README admits, and the ones it does not

The notes section is unusually direct. Buffer size consumption is larger than V1. That is the headline cost of the rewrite, and it is not quantified anywhere in the README. For a large MoE model where weights, optimizer state and activations already compete for HBM, an unquantified increase in intermediate buffer size is a real planning risk. You will not know whether it fits until you allocate it on your own hardware.

The second admitted regression: 0 SM RDMA low-latency EP is no longer supported. V1 users who built around that property cannot simply upgrade. If your design assumed communication kernels that occupy no SM slots on the RDMA path, V2 removes that option, and the README offers no migration note for it.

Third, Engram, PP and CP are labeled experimental. The news section advertises 0 SM Engram with RDMA, 0 SM PP with RDMA and 0 SM CP with Copy Engine, but the notes classify all three as experimental features. Advertising zero-SM operation and then calling the feature experimental in the same document is a tension worth taking at face value: do not build a production pipeline on them yet.

There is also a platform constraint that is easy to skim past. The requirements say Hopper (SM90) GPUs or other architectures with SM90 PTX ISA support. The performance table includes SM100 rows, so newer hardware is clearly exercised, but the stated requirement is framed around SM90 PTX ISA support. Verify your target architecture against that phrasing rather than against the benchmark table.

DeepEP compared with a general collective: what all-to-all in NCCL does differently

The obvious alternative is to skip DeepEP and use NCCL's own all-to-all directly, since V2 already runs on an NCCL backend. The difference is what each one knows about the data being moved.

A generic all-to-all moves a fixed, dense tensor between ranks. DeepEP's dispatch and combine are MoE-aware: the buffer is created with num_max_tokens_per_rank, hidden, num_topk and num_experts, and the dispatch call takes topk_idx and topk_weights. That means the library knows which tokens go to which experts and can lay out the exchange accordingly, including an FP8 dispatch path selected by use_fp8_dispatch. A generic collective has no concept of expert routing, so you would be responsible for the permutation, the layout and the precision conversion yourself, and you would lose the analytical SM and QP calculation that V2 performs at buffer creation.

The trade-off runs the other way too. A generic NCCL collective is a stable, widely deployed interface with behavior you can reason about from public documentation. DeepEP's kernels are compiled at runtime, its API is versioned with breaking changes between V1 and V2, and its buffer sizing is not documented in numbers. If your MoE is small, or your all-to-all is not the bottleneck, the generic path is less machinery to own. DeepEP earns its place when the exchange is the bottleneck and you are willing to accept the platform constraints and the buffer cost.

Version status, licence and the upgrade question

The latest release is v1.2.1, dated 2025-09-16, and the last push to the default branch was on 2025-09-16. That is roughly a year before today, so the repository is not under active development in any meaningful sense right now. It is not archived, and the README lists a set of still on-going features: elastic GPU and CPU buffers mapping a contiguous virtual address space over a hybrid of GPU and CPU physical memory, reduced intermediate buffer sizes via EP replay for load imbalance, and all-gather updates plus reduce-scatter implementations for DP and TP. Those are stated intentions. Nothing in the README indicates when or whether they shipped.

The practical consequence for an upgrade decision is that V2 is the current line and V1 is documented only in docs/legacy.md. Moving from V1 to V2 is not a version bump; it is an API migration, since the high-throughput and low-latency APIs are now merged into ElasticBuffer and the backend changed. Budget for it as a port, not a dependency update, and check the 0 SM RDMA low-latency EP removal against your existing design before you start.

The licence is MIT, which is permissive and places few conditions on redistribution and modification. That is a statement about the licence identifier in the repository, not legal advice; if you are embedding DeepEP in a distributed product, read the LICENSE file and the licences of the third-party and NVSHMEM dependencies yourself. The third-party directory and the .gitmodules file in the repository root indicate submodules are in play, so a source build pulls in code under its own terms.

Editorial conclusion

Adopt DeepEP V2 if you run MoE training or inference prefilling on SM90 or SM100 GPUs with NVLink inside the node and RDMA between nodes, and you can afford the larger intermediate buffers the V2 notes call out. Do not adopt it if you are on non-NVIDIA accelerators, if you still rely on the 0 SM RDMA low-latency EP path from V1, or if you need a stable, documented pipeline-, context- or Engram-parallel API, since the README labels those experimental. Before committing, verify three things on your own cluster: that your NCCL is at 2.30.4 or above with the Gin backend available, that PyTorch is at 2.10 or above, and that the buffer allocation for your EP domain fits alongside your weights and optimizer state. Then run tests/elastic/test_ep.py with init_dist adapted to your cluster, because the README says that function has to be modified before the tests will launch across multiple nodes.

Frequently asked questions

What is DeepEP?

DeepEP (DeepEveryParallel) is a communication library for modern machine learning training and inference, focused on expert parallelism. It provides high-throughput and low-latency all-to-all GPU kernels for MoE dispatch and combine, with low-precision support including FP8.

What is DeepEP used for?

It is used for the two all-to-all exchanges in a mixture-of-experts layer, dispatch and combine, in model training or inference prefilling. The README also lists experimental primitives for pipeline parallelism, context parallelism and remote memory access through Engram.

How do I install DeepEP?

The README recommends installing NCCL with pip install "nvidia-nccl-cu13>=2.30.4" --no-deps, then running python setup.py install. NVSHMEM is still required for legacy methods and is covered in docs/nvshmem.md rather than in the quick start.

What changed in DeepEP V2?

V2 is a complete refactoring of expert parallelism. It switched from the NVSHMEM backend to the NCCL Gin backend, unified the high-throughput and low-latency APIs into a single ElasticBuffer interface with a new GEMM layout, and calculates SM and QP counts analytically instead of auto-tuning.

What are the hardware requirements for DeepEP?

Hopper (SM90) GPUs or other architectures with SM90 PTX ISA support, Python 3.8 and above, CUDA 12.3 and above for SM90, PyTorch 2.10 and above, NCCL 2.30.4 and above, NVLink for intranode communication and an RDMA network for internode communication.

Does DeepEP V2 still support the 0 SM RDMA low-latency EP path?

No. The README notes state that 0 SM RDMA low-latency EP is no longer supported in V2. The 0 SM features that remain are listed for Engram, PP and CP, and those are described as experimental.

Official sources

  1. Official README
  2. Project repository
  3. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/deepseek-ai-deepep.svg)](https://hysenlabs.com/projects/deepseek-ai-deepep)
Community notes

Community notes