Library / SDK
microsoft/MInference avatar
microsoft/MInference

microsoft/MInference: dynamic sparse attention for long-context pre-filling

[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.

1,229 stars83 forksPythonMIT

At a glance

What is it?
MInference is a Microsoft research library that classifies each attention head into a sparse pattern offline, then builds the sparse index online during pre-filling. It targets engineers serving 128K to 1M token prompts on A100-class hardware, and its kernel now ships inside SGLang and vLLM.
Who is it for?
Adopt MInference when your cost is concentrated in pre-filling long prompts on A100-class GPUs and you can accept a one-time offline head-pattern search per model. Skip it if your workload is short-prompt or decode-bound, or if you cannot run the CUDA and Triton build path.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 20 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The pre-filling bottleneck MInference attacks

Long-context inference has two phases with different cost profiles. Decoding generates one token at a time and is memory-bandwidth bound. Pre-filling ingests the whole prompt at once and is compute bound, because every new token attends to every earlier token. At 1M tokens that quadratic term dominates the wall clock, and it is the phase MInference targets. The README states the goal plainly: process 1M context 10x faster in a single A100 using long-context models such as LLaMA-3-8B-1M and GLM-4-1M.

The audience is narrow and specific. You are serving long prompts (128K and up), you own the GPU, and you can afford a one-time offline analysis pass over the model before serving traffic. If you are calling a hosted API, this library is not in your path unless the provider already integrated it, which the news section says happened for Qwen2.5-1M online services. If your prompts are a few thousand tokens, the sparse pattern never becomes the bottleneck and the added machinery buys nothing.

How the head classification and online sparse index work

MInference rests on an empirical claim from the paper: attention in long-context LLMs is dynamically sparse, and the sparsity is not uniform across heads. Different heads settle into different static shapes, so a single global sparsity mask wastes compute on heads that need dense attention and misses savings on heads that do not.

The pipeline is two-stage. Offline, the library determines which sparse pattern each head belongs to, drawing from a set the README lists: MInference, xAttention, FlexPrefill, A-shape, Tri-shape, MInference w/ static, Dilated and Strided. Online, during pre-filling, it approximates the sparse index for the current input and computes attention with custom kernels chosen per head. That split matters: the expensive search happens once per model, and the per-request work is index approximation plus a sparse kernel call rather than a dense matmul over the full KV cache.

The repository layout reflects this. minference/ holds the Python package, csrc/ holds the CUDA sources, and setup.py builds a wheel through torch.utils.cpp_extension with CUDA_HOME, BuildExtension and CUDAExtension. The Makefile excludes minference/ops and minference/modules/minference_forward.py from the style checks, which tells you the kernel-facing code is not held to the same formatting rules as the rest. Build environment variables MINFERENCE_FORCE_BUILD and SKIP_CUDA_BUILD control whether you compile locally or reuse a prebuilt wheel from the GitHub release, with the wheel URL templated on the release tag.

Installing minference and running a first long-context pass

The README gives a single install command. The requirements list is Torch, Triton, optionally FlashAttention-2, and Transformers >= 4.46.0. Note that setup.py declares transformers>=4.37.0 in INSTALL_REQUIRES while the README asks for 4.46.0 or newer; follow the README, since the newer API surface is what the integration code expects.

bash
pip install minference

Once installed, the package exposes a configuration class you can query before writing any serving code. This is the cheapest way to confirm which attention and KV-cache methods your installed version actually carries, rather than trusting the README list.

python
from minference import MInferenceConfig
supported_attn_types = MInferenceConfig.get_available_attn_types()
supported_kv_types = MInferenceConfig.get_available_kv_types()

The repository ships runnable entry points under examples/. The README does not reproduce their contents, so read the files directly: examples/run_hf.py for a plain Hugging Face generation loop, examples/run_hf_streaming.py with a companion examples/run_hf_streaming.sh for the streaming path, and examples/run_vllm.py for the vLLM route. The shell script is the fastest way to see the expected arguments without reconstructing them from the Python source.

If you would rather not build the kernels yourself, the news section states that SGLang and vLLM merged the MInference sparse attention kernel, and that installing SGLang gives you the speedups without touching this repository. The README quotes up to 1.64x at 64K, 2.4x at 96K, 2.9x at 128K, 5.2x at 256K, 8x at 512K and 15x at 1M in that integration, and notes SGLang also adapted it for FlashAttention-3. Those numbers come from the project's own announcement, not from an independent run.

Where MInference is the wrong tool

The sparsity is approximate, and the library says so in its own description. On tasks where the answer depends on a small number of tokens buried in a very long prompt, an approximate index that misses those positions produces a wrong answer rather than a slow one. The README claims maintained accuracy, but accuracy is task-dependent, and the project's own SCBench work exists precisely because long-context methods behave differently across retrieval, global-information and multi-task settings. Treat the accuracy claim as a hypothesis to test on your data, not a property of the library.

The hardware constraint is equally real. The build path goes through CUDA and Triton, and setup.py imports CUDA_HOME and CUDAExtension at module scope, so a CPU-only or non-CUDA environment cannot complete the standard install. The 10x figure is stated for an A100; the README does not publish equivalent numbers for other accelerators.

The offline head-pattern search is a per-model cost. Swap the checkpoint and you owe that analysis again. The README does not document how long the search takes, whether it can be cached across machines, or how to roll back to dense attention if the sparse pattern degrades quality on your workload. That last gap is the one I would press on before a production rollout.

Finally, MInference accelerates pre-filling. If your traffic is dominated by decoding many short outputs, the pre-fill is a small fraction of the bill and the integration complexity is not repaid.

How MInference differs from xAttention, FlexPrefill and RetrievalAttention

The README lists several long-context methods under the same heading, which makes them look interchangeable. They are not, and the differences are structural.

xAttention, per the linked paper title, scores blocks by antidiagonal sums to decide which to keep. That is a block-selection rule applied at attention time. MInference instead precomputes a per-head pattern offline and approximates the index online, so its decision is model-specific and made once rather than per block per layer.

FlexPrefill is described as flexible, and the name signals the trade-off: more per-request adaptivity, more runtime decision work. MInference pushes that work into the offline stage and keeps the online path thin, which is why it can be baked into a serving kernel like the SGLang and vLLM integrations.

RetrievalAttention takes a different route entirely. The README describes it as KV cache offloading that accelerates long-context inference via vector retrieval, so it moves KV data out of GPU memory and fetches it back by similarity. That changes the memory profile, not just the attention compute. If your problem is that the KV cache does not fit, RetrievalAttention addresses that; MInference does not, and the two are complementary rather than competing.

MInference also ships SCBench, a KV cache-centric evaluation harness covering 12 tasks across two shared context modes and four capability categories. Whatever method you pick, that harness is the tool the project itself uses to compare them.

Licence, maintenance cadence and the cost of upgrading

MInference is MIT licensed, with the copyright header in setup.py and the Makefile reading "Copyright (c) 2024-2025 Microsoft" and "Licensed under The MIT License". MIT is permissive: you can use, modify and redistribute it, including in closed products, provided you keep the copyright and permission notice. That is a summary of the licence text, not legal advice; read LICENSE and your own counsel's view before shipping.

The repository is not archived, and the last push was on 2026-09-10. The most recent tagged release is v0.1.6 on 2025-06-17, which added SCBench. Before that, v0.1.5.post1 on 2024-08-13 supported LLaMA-3-70B, multi-GPU, and fixed a kernel and a sqrt(dk) issue; v0.1.5 landed on 2024-07-24. The gap between the latest tag and the latest commit means the main branch carries work that is not in a release, so pinning a tag and pinning a commit give you different code.

Upgrade cost concentrates in two places. The Python surface is small (MInferenceConfig plus the example scripts), so API churn is unlikely to be your problem. The kernels are. csrc/ and the CUDAExtension build mean a Torch or CUDA toolkit bump can force a rebuild, and MINFERENCE_FORCE_BUILD exists for exactly that case. If you consume MInference through SGLang or vLLM instead, your upgrade cadence is theirs, not this repository's, which is simpler but puts you behind on new patterns.

Editorial conclusion

Adopt MInference when your cost is concentrated in pre-filling long prompts on A100-class GPUs and you can accept a one-time offline head-pattern search per model. Skip it if your workload is short-prompt or decode-bound, or if you cannot run the CUDA and Triton build path. Before committing, verify two things: that your target model is in the supported list (the README names LLaMA-3-8B-1M, GLM-4-1M and meta-llama/Meta-Llama-3.1-8B-Instruct), and that the sparse pattern search for your checkpoint reproduces the accuracy you need on your own long-context task, since the README does not document a rollback path if it does not.

Frequently asked questions

What Python packages does MInference require?

The README lists Torch, Triton, optionally FlashAttention-2, and Transformers >= 4.46.0. The setup.py INSTALL_REQUIRES list is slightly looser, naming transformers>=4.37.0, torch, triton and einops.

How do I install MInference?

The README gives one command: pip install minference. Because setup.py builds a CUDA extension through torch.utils.cpp_extension, the install expects a CUDA toolchain unless you use the prebuilt wheel path controlled by MINFERENCE_FORCE_BUILD and SKIP_CUDA_BUILD.

Can I use MInference through SGLang or vLLM instead of installing it directly?

Yes. The news section states that SGLang and vLLM merged the MInference sparse attention kernel, that installing SGLang is enough to use it, and that SGLang also adapted it for FlashAttention-3.

Which long-context models does MInference support?

The README names LLaMA-3-8B-1M, GLM-4-1M and meta-llama/Meta-Llama-3.1-8B-Instruct, and release v0.1.5.post1 added LLaMA-3-70B. The installed package exposes MInferenceConfig.get_available_attn_types() and get_available_kv_types() to list what your version carries.

Is MInference the same as MMInference?

No. MMInference is a separate multi-modality work from the same group that applies modality-aware permutation sparse attention to long-context VLMs during pre-filling, and it has its own paper and repository link in the README.

Official sources

  1. License: MIT
  2. microsoft/MInference on GitHub
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/microsoft-minference.svg)](https://hysenlabs.com/projects/microsoft-minference)