Mixture-of-Kittens: a deterministic MoE megakernel for NVL72 racks
Mixture-of-experts (MoE) training megakernel for NVL72s
At a glance
- What is it?
- Cursor's MoK fuses MoE compute and inter-GPU communication into one kernel for Blackwell NVL72 systems, and it only runs there. Here is what the repository actually documents, how to install it, and where it stops.
- Who is it for?
- Adopt MoK only if you train mixture-of-experts models on GB200 or GB300 NVL72 hardware and can tune five kernel hyperparameters per workload; on any other GPU it will not build or run, since the Makefile targets SM100 or SM103 only.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Mixture-of-Kittens solves, and for whom
A mixture-of-experts layer is two workloads glued together: a local expert computation and an all-to-all exchange that moves tokens to the ranks that own their experts. In a conventional implementation those two phases alternate, so the network sits idle while the SMs compute and the SMs sit idle while the network moves tokens. MoK's premise is that on an NVL72 rack the interconnect is fast enough that the two should be overlapped inside a single persistent kernel rather than scheduled as separate CUDA work.
The README states that MoK fuses all MoE computation and communication into one kernel, overlaps them at configurable granularity, and fully eliminates CPU-GPU synchronization. It covers forward and backward passes in BF16 and MXFP8, and the README says it powers production training of Composer at Cursor. The intended reader is an infrastructure engineer who already runs expert parallelism across many GPUs and is willing to write kernel-level glue. This is not a drop-in replacement for a Hugging Face MoE block; it is a lower layer that assumes you own the data layout, the symmetric buffers, and the launch schedule.
The two API layers and the schedule object
The repository exposes two levels. The ops layer in mok/ops.py is the low-level surface: you call the CUDA kernels directly and manage your own data, including symmetric tensors, coordinating kernel calls yourself. The functional layer in mok/functional.py is the higher-level surface: it maintains scratch memory and coordinates kernel launches, and consists of workspace creation functions, a schedule function, and forward and backward functions. The README says the functional layer is the choice for production training and recommends it unless you have specific needs.
The data flow in the functional layer is narrow. You call schedule(...) once to build the dispatch and combine schedule, then pass that same schedule object to forward(...) and backward(...). That schedule is what makes the design deterministic: routing decisions are materialized before the kernels run, so the same inputs produce the same dispatch pattern instead of being resolved on the fly inside the kernel. The cost is that the schedule must be sized for the worst case, which is what schedule_capacity_multiplier controls.
Two objects have to exist before any of that. A MoKConfig dataclass carries five hyperparameters, and a workspace holds the PyTorch symmetric memory buffers plus metadata and scratchpads. Symmetric buffers are identically sized allocations across all GPUs in an expert-parallel group, and the README notes they are expensive to allocate, which is why it recommends one workspace per model reused across layers. get_workspace(...) caches workspaces with identical properties such as EP group, device and model shapes; create_workspace(...) does not cache and leaves lifetime management to you.
Installing MoK on a Blackwell NVL72 node
The requirements are unusually narrow, so check them before anything else. You need NVIDIA Blackwell SM100 or SM103 GPUs such as GB200 NVL72 or GB300 NVL72, Python 3.12 or later, PyTorch 2.10 or later, and CUDA toolkit 13.0 or later. setup.py enforces several of these at build time: it raises if the platform is not Linux, if GNU Make is missing, if Python is older than 3.12, if Python.h is absent, or if PyTorch falls outside >=2.10.0,<3.
The installed PyTorch build must target CUDA 13.0 or later, and its CUDA version must match the major and minor version of the system toolkit. To get a matching wheel:
python -m pip install "torch==2.10.0+cu130" --index-url https://download.pytorch.org/whl/cu130MoK builds for SM103 by default. From the repository root, the README gives two equivalent install paths:
pip install . --no-build-isolationpython setup.py installIf your hardware is SM100 rather than SM103, override the architecture through the environment variable. Note that the Makefile's own ARCH default is SM103, and the variable the README documents for the pip path is MOK_ARCH:
MOK_ARCH=SM100 pip install . --no-build-isolationAfter the build finishes, confirm the extension imports and reports a version:
python -c "import mok; print(mok.__version__)"For iterating on the CUDA sources, the README recommends an editable install followed by make for fast rebuilds. The Makefile compiles csrc/bindings.cu against the ThunderKittens headers under third_party/ThunderKittens and writes the extension to mok/_C with the interpreter's extension suffix:
pip install -e . --no-build-isolation
makeThe README points multi-GPU tests at torchrun with pytest, with the process count set to your GPU count:
torchrun --standalone --nproc-per-node=<num-gpus> -m pytest -s <test-path>The README does not document rollback, an uninstall procedure, or a CPU-only fallback path.
Tuning the five hyperparameters before production
MoK exposes exactly five hyperparameters that affect MoE execution performance, and the README is direct that optimal values depend heavily on the workload, so they should be swept before production use. fwd_num_comm_sms and bwd_num_comm_sms set the number of communication SMs in each direction, with a recommended range of 4 to 52. minibatch_size sets the granularity of compute-communication overlap and the README calls it an important parameter that must be tuned properly, recommending 2048 to 16384. macrobatch_size is the token ring buffer size; a large value such as 131072 means the ring buffer is used only once, and the guidance is to maximize it to fill available GPU memory.
schedule_capacity_multiplier is the one with a training-schedule implication. It defaults to 0.5 and should be set to the worst-case fraction of tokens routed to a single rank. Setting it to 1 assumes the absolute worst case of all tokens going to one rank, but adds kernel scheduling overhead, and because of expert padding the actual worst case sits slightly above 1.0. The README's suggested pattern is a higher value during the first training steps when expert imbalance is bad, reduced to around 0.5 later. It also notes that lowering this value does not save GPU memory meaningfully, since the schedule table is at most a few megabytes. That is a useful correction to the intuition that a smaller capacity multiplier buys headroom elsewhere.
For MXFP8 mode, the workflow is asymmetric: pass activations as-is in BF16 while prequantizing the weights to MXFP8. The README explains the choice, saying the kernels could quantize weights internally but that prequantizing leaves better opportunities for things like FSDP, so quantization is kept separate and exposed as mxfp8_quantize(...) at the ops layer. If your training stack already shards weights with FSDP, that separation is the reason the integration works at all; if you expected the kernel to handle quantization, you will be writing that step yourself.
Where MoK is the wrong tool
The hardware requirement is the first wall. MoK targets SM100 and SM103 only, and the Makefile selects between them with -gencode arch=compute_103a,code=sm_103a for the SM103 path. On Hopper, Ada, or any non-Blackwell part there is no supported build target, and setup.py refuses non-Linux platforms outright. A team training MoE models on H100 clusters cannot use this project at all, regardless of how well the design would map.
The second wall is the integration surface. MoK is a kernel library, not a training framework. It gives you schedule, forward and backward, and expects you to bring symmetric-memory tensors, an expert-parallel layout, and your own routing. The README's example assumes inputs like topk_experts as an int64 tensor of shape [num_local_tokens, topk] and router_weights as float32, so the router itself is outside the library. If you want a complete MoE training loop with checkpointing, optimizer state, and data loading, this repository does not provide one.
The third issue is the benchmark framing. The README reports up to 2.37x for MXFP8 forward and 1.78x for MXFP8 backward against the fastest baseline, with the methodology and full results in the linked blog post rather than in the repository. Those are standalone MoE layer comparisons, not end-to-end training throughput, and the README describes end-to-end evaluation separately on an internal production stack across multiple NVL72 racks. Treat the layer-level multipliers as an upper bound on what you might see, not a prediction for your training run.
How MoK differs from a general MoE kernel library
The obvious comparison is with libraries that provide MoE primitives for a range of GPUs, such as DeepSpeed-MoE or Megatron-LM's expert-parallel layers. Those projects prioritize portability and integration with a full training stack: they ship the router, the auxiliary losses, the checkpoint format, and kernels that run on several generations of hardware. MoK makes the opposite set of choices. It supports one hardware family, exposes five tuning knobs, and assumes the surrounding training system already exists. The difference in approach is that MoK treats the compute-communication overlap as the whole problem and solves it with a hand-written megakernel, while general libraries treat the MoE layer as one component among many and accept less overlap in exchange for running anywhere.
A second comparison point is the schedule model. Frameworks that dispatch tokens dynamically inside the kernel avoid a precomputed schedule table but pay for it with data-dependent control flow. MoK's schedule(...) call materializes dispatch and combine decisions up front, which the README frames as fully deterministic behavior. Determinism is worth real money when a training run diverges and you need to reproduce a step, but it also means the schedule must be sized for the worst case through schedule_capacity_multiplier, and the README notes that expert padding pushes the true worst case slightly above 1.0. You are trading dynamic flexibility for a predictable, replayable execution plan.
Licence, maintenance and upgrade cost
MoK is licensed under Apache-2.0, and pyproject.toml declares license = "Apache-2.0" with license-files = ["LICENSE"]. That is a permissive licence, but note that the build pulls in a submodule: third_party/ThunderKittens appears in .gitmodules and the Makefile includes headers from ${THUNDERKITTENS_ROOT}/include. If you redistribute a built wheel, check the ThunderKittens licence separately, since Apache-2.0 on this repository does not automatically cover a vendored dependency. Nothing here is legal advice; read both licences if redistribution matters to you.
The pinned dependency range is torch>=2.10.0,<3 in both pyproject.toml and setup.py, so a PyTorch 3.x upgrade will require a code change, not just a version bump. The CUDA constraint is stricter in practice: the PyTorch build's CUDA version must match the major and minor version of the system toolkit, so upgrading the driver stack usually means reinstalling PyTorch and rebuilding the extension. The build is compiled, not pure Python, so every CUDA or PyTorch change triggers a recompile through make or a fresh pip install.
On maintenance, the last push to the default branch was on 2026-08-14, and v0.1.0 was released on 2026-08-03. The repository is not archived. The README does not state a support policy, a compatibility matrix beyond the CUDA 13.0 baseline, or a deprecation plan, so plan for rebuilds on your own schedule rather than relying on a published upgrade path.
Editorial conclusion
Adopt MoK only if you train mixture-of-experts models on GB200 or GB300 NVL72 hardware and can tune five kernel hyperparameters per workload; on any other GPU it will not build or run, since the Makefile targets SM100 or SM103 only. Before committing, verify that your PyTorch build and system CUDA toolkit agree at the major and minor version, and measure your own minibatch_size and schedule_capacity_multiplier sweep, because the README states that optimal values depend heavily on the workload.
Frequently asked questions
What GPU do I need to run Mixture-of-Kittens?
The README requires NVIDIA Blackwell SM100 or SM103 GPUs, such as GB200 NVL72 or GB300 NVL72, and setup.py rejects non-Linux platforms. MoK builds for SM103 by default, and you set MOK_ARCH=SM100 to build for SM100 instead.
What PyTorch and CUDA versions does Mixture-of-Kittens require?
Python 3.12 or later, PyTorch 2.10 or later, and CUDA toolkit 13.0 or later. The README states that the installed PyTorch build must target CUDA 13.0 or later and that its CUDA version must match the major and minor version of the system CUDA toolkit.
How do I install Mixture-of-Kittens from the repository?
From the repository root, the README gives pip install . --no-build-isolation or python setup.py install, both with --no-build-isolation. You can then verify the install with python -c "import mok; print(mok.__version__)".
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/cursor-mixture-of-kittens)