Library / SDK
MoonshotAI/MoonEP avatar
MoonshotAI/MoonEP

MoonEP: Balanced Expert Parallelism for MoE Training via Dynamic Redundant Experts

MoonEP: A Perfectly Balanced Expert Parallelism Library via Dynamic Redundant Experts

1,156 stars134 forksPythonMIT

At a glance

What is it?
MoonEP is a Python library from MoonshotAI that keeps token loads perfectly balanced across ranks during Mixture-of-Experts training by planning and prefetching redundant experts at runtime. It eliminates the routing-skew degradation and out-of-memory failures that affect naive expert parallelism implementations.
Who is it for?
MoonEP is the right tool when you are training a Mixture-of-Experts model with expert parallelism on NVIDIA GPUs and routing imbalance is causing variable iteration times or out-of-memory failures. The static activation shapes eliminate memory fragmentation, and the planning kernel adds negligible overhead compared to the communication cost savings at high imbalance.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 10 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The Expert Parallelism Imbalance Problem MoonEP Solves

In Mixture-of-Experts models, each token is routed to a small number of experts, typically the top-k by router score. Under expert parallelism, each GPU rank holds a subset of experts. When routing is balanced, each rank receives roughly the same number of tokens and computation is evenly distributed. When routing is skewed, some ranks receive far more tokens than others, and the overall iteration time is set by the hottest rank.

The README uses a specific metric to describe imbalance: maxvio, defined as the maximum ratio of tokens routed to any single expert divided by the expected number of tokens per expert under perfect balance, minus one. A maxvio of zero means perfectly balanced; any positive value means some expert is receiving more tokens than expected.

Under naive expert parallelism, both communication time and computation time grow with the hottest rank's load. The README describes the end-to-end effect: iteration time climbs as maxvio grows, and the ever-changing activation tensor shapes fragment GPU memory until training eventually runs out of memory at high imbalance levels. MoonEP addresses both problems by guaranteeing that every rank processes exactly S times K tokens per layer, where S is the input tokens per rank and K is the routed top-k per token, regardless of routing skew.

How Redundant Expert Planning Achieves Perfect Balance

MoonEP achieves perfect balance by planning and prefetching redundant copies of experts before computation begins. The planner runs a near-optimal GPU kernel that, given the current router outputs, determines which experts need redundant copies on which ranks. It then prefetches those expert weights to the ranks that need them before the expert computation layer executes.

The README describes the prefetch mechanism: each rank has a fixed number of prefetch slots equal to the number of local experts. The planner moves experts from at most one remote home group to each destination rank, so these slots cover every remote expert segment. The prefetch pool is process-global and shared by all layers, which means the extra physical memory cost is one set of expert weights per projection across all layers, not per layer.

Zero-copy dispatch is the second key mechanism. Tokens are written directly to their final expert-grouped positions on remote ranks during communication. No copy from a communication buffer to a user buffer is needed. The README states that this copy normally dominates the communication epilogue, and eliminating it makes MoonEP's raw communication time consistently faster than DeepEP v2 even before accounting for imbalance effects.

The backward pass mirrors the forward. The gradient of the dispatch operation is a combine operation that sums each token's K dispatched gradient copies back to the token-major layout. Gradients for prefetched experts are reduced back to their home ranks using a dedicated reduce buffer.

Installing and Building the CUDA Extension

MoonEP requires compiling a CUDA extension before it can be used. The setup.py file in the repository builds a pybind11 module named moonep._C from the source in csrc/bindings.cu.

The recommended installation is an editable install:

bash
pip install -e .

For a quick in-place build without installing:

bash
python setup.py build_ext --inplace

The setup.py specifies one runtime dependency: nvidia-cutlass-dsl version 4.6.2. This is the NVIDIA CUTLASS DSL, which the CUDA kernels use for high-performance matrix operations. Mismatching this version is likely to cause build failures or incorrect behavior.

The build uses nvcc with optimization flags including -O3, --use_fast_math, and -std=c++20. The CUDA extension lands at moonep/_C.<abi-tag>.so after a successful build. After building, import works as follows:

python
from moonep import Buffer

The Buffer API: Dispatch, Combine, Prefetch, and Reduce

The central object is Buffer, which is initialized with the problem dimensions:

python
from moonep import Buffer

buffer = Buffer(S=4096, H=7168, K=8, E=256, num_ep_ranks=8,
                num_sms=32, token_padding=128)

Here S is tokens per rank, H is hidden size, K is routed top-k per token, E is total routed experts in the expert parallelism group, and num_ep_ranks is the number of ranks in the EP group. The num_sms parameter defaults to 32 and controls how many streaming multiprocessors the planning kernel uses. The token_padding parameter adds padding to each VM group's token slot count.

The four main operations are dispatch for the forward pass, combine for gathering results, prefetch_weight for loading expert weights before computation, and reduce_grad for accumulating gradients in the backward pass. All four accept an async_finish parameter: passing True runs the operation on the communication stream and returns a CUDA event, enabling overlap with computation.

The dispatch operation returns a plan object alongside the dispatched tensors. That plan must be saved and passed to all subsequent prefetch, combine, and backward operations in the same layer forward-backward cycle.

Performance Comparison with DeepEP v2

The README describes two benchmark categories: raw communication time and end-to-end training, both run on H20 hardware with EP=8 and sweeping the maxvio parameter.

For communication time, the README states that MoonEP's communication time stays almost flat as maxvio grows, while DeepEP v2's latency is set by the hottest rank and degrades steadily. The MoonEP bars in the benchmark charts include the planning and weight-prefetch kernels that DeepEP does not need, and the README states that even with this overhead included, total dispatch time is on par with DeepEP v2 at zero imbalance and pulls ahead as imbalance grows. Combine is described as significantly faster at every imbalance level.

For end-to-end training, the README states that DeepEP degrades with imbalance because hottest ranks receive more tokens, causing iteration time to climb and eventually OOM at high imbalance due to changing activation shapes. MoonEP's iteration time stays flat at every imbalance level because every rank always computes exactly S times K tokens per layer. The static memory shapes mean no fragmentation.

The benchmark script is available at benchmarks/bench_vs_deepep.py in the repository.

Limitations and Hardware Requirements

MoonEP currently supports NVIDIA GPUs. Zhenwu PPU support is listed in the README as under review and coming soon. No other hardware is mentioned. Teams using AMD GPUs or other accelerators cannot use this library.

The planning kernel adds overhead that DeepEP does not have. The README explicitly notes this and states that the overhead is negligible compared to the communication cost at the scales benchmarked, but teams working with small EP groups or very short sequences may see a different trade-off than the benchmarks show.

The API is designed for integration at the MoE layer level inside a training framework. It is not a drop-in replacement for a complete training framework. The README documents how to integrate it at the weight buffer and gradient buffer level, which requires the framework to expose expert weights in the specific layout MoonEP expects.

The package version is 0.0.1 as listed in setup.py, and there are no GitHub releases. The library is early-stage research code rather than a production-hardened package.

Maintenance and License

The repository is published under the MIT license and maintained under the MoonshotAI organization on GitHub. The last push was on 2026-09-20, and the repository is not archived. No GitHub releases are present.

The MIT license permits unrestricted use including in commercial training infrastructure. The dependency nvidia-cutlass-dsl has its own license from NVIDIA, which teams building on MoonEP should review separately from the MoonEP MIT license.

The tests directory in the repository provides test coverage for the core operations. The benchmarks directory holds the comparison script against DeepEP v2 and additional benchmark tooling. Teams who want to verify their own setup before relying on MoonEP in a training run should use those benchmarks with their own hardware and sequence lengths to confirm the behavior matches expectations.

Editorial conclusion

MoonEP is the right tool when you are training a Mixture-of-Experts model with expert parallelism on NVIDIA GPUs and routing imbalance is causing variable iteration times or out-of-memory failures. The static activation shapes eliminate memory fragmentation, and the planning kernel adds negligible overhead compared to the communication cost savings at high imbalance. Teams training on hardware other than NVIDIA GPUs cannot use it: Zhenwu PPU support is listed as under review. The package requires nvidia-cutlass-dsl version 4.6.2 as a dependency, and the build requires a CUDA extension to compile via pip install -e . before the library is importable.

Frequently asked questions

What is the difference between MoonEP and DeepEP for expert parallelism?

MoonEP guarantees perfect token balance across ranks by planning and prefetching redundant experts before each layer, which keeps iteration time flat as routing imbalance grows and prevents out-of-memory failures from variable activation shapes. DeepEP v2 does not redistribute tokens, so its performance degrades when routing is skewed and it OOMs at high imbalance according to the benchmark results in the README.

Does MoonEP work with AMD or other non-NVIDIA GPUs?

The README lists NVIDIA GPU as the supported device. Zhenwu PPU support is described as under review. No AMD GPU or other hardware support is documented.

How much extra GPU memory does MoonEP require for redundant expert storage?

Each rank allocates a prefetch pool of expert weights equal to one set of local expert weights per projection (gate, up, down). The README states this pool is process-global and shared across all layers, so the total extra cost is one set of expert weights per projection in total, not per layer.

Official sources

  1. Issues
  2. License: MIT
  3. MoonshotAI/MoonEP on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/moonshotai-moonep.svg)](https://hysenlabs.com/projects/moonshotai-moonep)