Library / SDK
mirage-project/mirage avatar
mirage-project/mirage

Mirage Persistent Kernel: compiling an LLM into a single megakernel

Mirage Persistent Kernel: Compiling LLMs into a MegaKernel. Quickstart Mirage allows you to compile LLMs from the Hugging Face model zoo into a megakernel using just a few dozen lines of Python-mainly to define the kernel s inputs and outputs.

2,514 stars257 forksCudaApache-2.0

At a glance

What is it?
MPK is a compiler and runtime that fuses multi-GPU LLM inference into one kernel launch. The repository documents the Python API for defining layers and tensors, but leaves the supported hardware and error handling mostly unstated.
Who is it for?
MPK is worth a look if you already run multi-GPU LLM serving on Hopper or Blackwell class hardware, are comfortable reading CUDA and CMake, and want to see whether removing per-kernel launch overhead helps your latency. It is not a drop-in replacement for vLLM or TensorRT-LLM: it is a research compiler with a source build, a hard z3-solver pin, and an API that expects you to describe the computation graph yourself.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly Cuda, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The latency problem MPK attacks, and who feels it

Standard LLM inference on a GPU is a sequence of kernel launches: one for each normalization, matrix multiply, attention step and collective. Each launch has fixed overhead, and the gaps between launches leave SMs idle. MPK's premise, as the README states, is to transform inference into a single megakernel, a fused kernel that performs all necessary computation and communication within one launch. The README claims this reduces latency by 1.2x to 6.7x. That number comes from the project, not from an independent run.

The audience is narrow. You need multiple GPUs, a CUDA toolchain, and enough familiarity with thread block geometry to reason about grid_dim and block_dim. The README suggests keeping the total number of thread blocks a multiple of the worker count to avoid outliers, which assumes you know what a worker is in this system. Teams serving a single model on one GPU with a mature runtime like vLLM will not get much from this. Researchers and engineers studying kernel fusion, or serving latency-sensitive workloads where launch overhead is a measurable share of the budget, are the intended users.

How MPK turns a model into one kernel launch

MPK is a Python library that builds a computation graph and compiles it. You instantiate a PersistentKernel with a world size, an MPI rank, and three scheduler and worker counts. The README states these must match the number of physical SMs, with the relation num_workers + (num_local_schedulers + num_remote_schedulers) / 4. That is a hardware constraint expressed as arithmetic, and getting it wrong is on you.

Tensors enter the graph in two ways. attach_input wraps an existing torch.Tensor under a name that the generated CUDA refers to. new_tensor allocates a fresh one with dims, dtype and an io_category of either cuda_tensor or nvshmem_tensor. The README notes nvshmem_tensor is required for remote GPU access, for example during all-reduce, so the memory category is not cosmetic: it decides whether a tensor is reachable across ranks.

The graph itself is a chain of fused layer calls. The README's example is rmsnorm_linear_layer, which fuses an RMSNorm and a Linear layer into one node, taking weight_norm and weight_linear separately. Each call carries grid_dim and block_dim, so the task granularity is exposed at the Python level rather than decided by a scheduler. Two meta tensors are mandatory: step, an integer array tracking the decoding step that MPK increments after each iteration, and tokens, shaped [num_requests, seq_length], holding prompts and generated tokens. Compilation is mpk.compile(), execution is mpk().

Installing MPK from source and running the Qwen3 demo

The README gives source installation as the fastest path, with pre-built wheels still listed as work in progress. The build is recursive because the repository carries submodules under deps/.

bash
git clone --recursive --branch mpk https://www.github.com/mirage-project/mirage
cd mirage
pip install -e . -v
export MIRAGE_HOME=$(pwd)

After that, the README points to demo/qwen3/demo.py as the reference for compiling Qwen3-8B. The first run uses native Triton and FlashInfer kernels as a baseline; the second compiles and executes the megakernel. The third turns on profiling, which the README says visualizes the execution timeline of each task.

bash
python demo/qwen3/demo.py
python demo/qwen3/demo.py --use-mirage
python demo/qwen3/demo.py --use-mirage --profiling

The dependencies are pinned tighter than most projects. requirements.txt fixes z3-solver==4.16.0.0, transformers==4.57.1 and accelerate==1.8.0, and pyproject.toml repeats the z3 pin in the build requirements with a comment explaining why: the Cython extension is linked against that z3's libz3 SONAME, and an unpinned build dependency resolves to a newer library that fails to load at import. If you have a different z3 in your environment, expect an import error rather than a helpful message.

Where MPK is the wrong tool

The README does not document rollback, error recovery, or what happens when a megakernel fails partway through a decoding step. That absence matters. A single fused launch means a fault or a hang takes down the whole step, and there is no per-kernel boundary at which to retry. The documentation is silent on this, so treat it as an open question rather than a solved one.

The scheduler arithmetic is another sharp edge. Because num_workers plus a quarter of the schedulers must equal the SM count, the same Python code is not portable across GPU models without editing those numbers. The README does not publish a table of valid combinations per architecture. There are demo directories for Hopper and Blackwell, which suggests those are the tested targets, but the README does not state a minimum compute capability.

The build itself is a filter. It needs nvcc, CMake 3.24 or newer, and the exact z3 pin. The project also lists torch>=2.4 and cuda-python, and pulls tg4perfetto from a git URL, so the install reaches outside PyPI. If your environment forbids git dependencies in the build, this is a blocker before you write a line of model code.

Mirage versus a serving runtime like vLLM

vLLM and MPK both aim at low-latency LLM inference, but they sit at different layers. vLLM is a serving system: it manages request scheduling, paged KV cache, batching and an OpenAI-compatible HTTP API, and it dispatches to existing kernels. MPK is a compiler: it takes a description of the computation graph and emits a fused kernel, leaving request handling to you. The README's example constructs the kernel with a fixed num_requests dimension in the tokens tensor, which is a static shape decision a serving system would normally make dynamically.

The difference in effort is real. With vLLM you point it at a model and send HTTP requests. With MPK you write the layer chain, choose grid and block dimensions, allocate each intermediate tensor and pick its io_category, then compile. The README frames this as a few dozen lines of Python, which is accurate for a demo but understates the work of mapping a full model architecture, including attention variants and MoE routing, onto the available fused layer primitives. The repository has demos for llama3, qwen3, deepseek_v3, glm4_moe and gpt_oss, so the primitives cover more than one architecture, but the README does not document a general fallback for a layer MPK does not have.

Maintenance, licensing and the cost of upgrading

The last push to the default branch was on 2026-06-05, which is the same timestamp as the osdi26-ae-v1 release, an artifact release tied to the OSDI 2026 paper. The repository is not archived. Between that artifact release and the earlier v0.2.4 in March 2025 there is a gap of more than a year in tagged releases, so the version numbers you see on PyPI-style installs and the code on the mpk branch are not moving together. Pin to a commit rather than a tag if you need reproducibility.

Licensing is Apache-2.0, stated in both the README and the setup.py header. That is permissive, but it covers the Mirage code, not the dependencies. The build pulls z3-solver, torch, transformers, accelerate, graphviz, cuda-python, fastapi, uvicorn and a git-sourced tg4perfetto package, each under its own terms. CUDA itself is not open source. If you plan to redistribute a binary built from this, check the dependency licences separately; this article is not legal advice.

Upgrade cost is dominated by the z3 pin. Any change to the pinned version requires rebuilding the Cython extension against a matching libz3, and the pyproject comment makes clear the build and runtime environments must agree. That is a coordinated change across your build image and your runtime image, not a version bump in a requirements file.

Editorial conclusion

MPK is worth a look if you already run multi-GPU LLM serving on Hopper or Blackwell class hardware, are comfortable reading CUDA and CMake, and want to see whether removing per-kernel launch overhead helps your latency. It is not a drop-in replacement for vLLM or TensorRT-LLM: it is a research compiler with a source build, a hard z3-solver pin, and an API that expects you to describe the computation graph yourself. Before committing, verify that your GPU generation is covered by the demo directory, that your CUDA toolkit matches what setup.py expects, and that the two meta tensors (step and tokens) fit the request shape you plan to serve.

Frequently asked questions

How do I install Mirage Persistent Kernel?

The README gives a source install: clone the repository recursively on the mpk branch, run pip install -e . -v, then export MIRAGE_HOME to the checkout directory. Pre-built wheels are described as work in progress, so the source path is the documented one.

What does Mirage Persistent Kernel compile?

It compiles LLM inference into a single megakernel, a fused GPU kernel that performs all necessary computation and communication within one kernel launch. The README points to demo/qwen3/demo.py as the example that compiles Qwen3-8B.

How do I run the Mirage Persistent Kernel demo?

The README lists three commands: python demo/qwen3/demo.py for the native Triton and FlashInfer baseline, the same command with --use-mirage to compile and execute the megakernel, and adding --profiling to visualize the execution timeline of each task.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/mirage-project-mirage.svg)](https://hysenlabs.com/projects/mirage-project-mirage)