Library / SDK
flashinfer-ai/flashinfer avatar
flashinfer-ai/flashinfer

FlashInfer: a kernel library for LLM serving, installed with pip

FlashInfer: Kernel Library for LLM Serving. It provides unified APIs for attention, GEMM, and MoE operations with multiple backend implementations including FlashAttention-2/3, cuDNN, CUTLASS, and TensorRT-LLM.

6,506 stars1,501 forksPythonApache-2.0

At a glance

What is it?
FlashInfer packages attention, GEMM, MoE and sampling kernels behind one Python API and picks a backend per GPU. The pip install is one line; the architecture coverage table is where the caveats start.
Who is it for?
Adopt FlashInfer if you are serving LLMs on NVIDIA GPUs from Turing through Blackwell and want attention, GEMM, MoE and sampling behind one API rather than hand-written CUDA. Do not adopt it if your hardware falls outside the supported compute capabilities, or if you need a CPU or non-NVIDIA path, because none is described.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 4 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem FlashInfer solves for serving engineers

Serving a large language model is not one kernel problem. A request touches paged attention over a KV cache, a matrix multiply, sometimes a mixture-of-experts router, and a sampler. Each of those has a different best implementation depending on the GPU generation and the batch shape, and each of those implementations has its own calling convention. FlashInfer's answer is to put them behind unified APIs for attention, GEMM and MoE operations, with multiple backend implementations including FlashAttention-2/3, cuDNN, CUTLASS and TensorRT-LLM. The README states that the library automatically selects the best backend for your hardware and workload.

The audience is narrow and specific: people building or maintaining an inference server on NVIDIA hardware who would otherwise wire several kernel libraries together by hand. The README's feature list reads like a serving stack inventory rather than a research toolkit. Paged and ragged KV-cache, decode, prefill and append phases, cascade attention for shared prefixes, fused MoE with DeepSeek-V3 and Llama-4 routing, sorting-free Top-K, Top-P and Min-P sampling. If your work is training, or a single-sequence notebook experiment, most of that list is irrelevant to you.

How the backends, JIT compilation and cubin packages fit together

The packaging is the part most readers underestimate. There are three distributions. flashinfer-python is the core package that, per the README, compiles or downloads kernels on first use. flashinfer-cubin holds pre-compiled kernel binaries for all supported GPU architectures. flashinfer-jit-cache holds a pre-built kernel cache for specific CUDA versions. The repository layout matches this: flashinfer-cubin/ and flashinfer-jit-cache/ are top-level directories alongside csrc/, include/ and the Python package in flashinfer/.

That split explains the first-run behaviour. A bare flashinfer-python install defers compilation to the moment you call a kernel, which costs time and requires a working toolchain. Installing the cubin and jit-cache packages moves that work to install time, which the README frames as faster initialization and offline usage. The jit-cache build is architecture-specific: the documented source build sets FLASHINFER_CUDA_ARCH_LIST to a space-separated list such as "7.5 8.0 8.9 9.0a 10.0a 10.3a 10.7a 11.0a 12.0f". Build for fewer architectures and the wheel is smaller; build for the wrong ones and the kernels you need are not in it.

Backend selection is not a single switch. The README lists FlashAttention-2/3, cuDNN, CUTLASS and TensorRT-LLM as implementations behind the unified APIs, and says selection is automatic. The README does not document a public API for forcing a specific backend per call, so treat the choice as the library's to make unless you find otherwise in the documentation.

Installing FlashInfer and running a first decode

The quickstart is a single pip command. The README gives it as-is, and the package name is flashinfer-python, not flashinfer.

bash
pip install flashinfer-python

With only that installed, kernels compile or download on first use. To move that cost to install time, the README shows two follow-up commands that pull the pre-compiled binaries and the pre-built cache.

bash
pip install flashinfer-python
flashinfer install-cubin-wheel
flashinfer install-jit-cache-wheel

On Blackwell (SM100 and later) the README points to the CUDA 13 extra for the CuTe DSL kernels. Note the extra name: cu13.

bash
pip install flashinfer-python[cu13]

The README's verification step is a CLI command, which prints the configuration the library resolved. Run it before writing any serving code, because a missing cubin or a mismatched CUDA version shows up here rather than at the first kernel call.

bash
flashinfer show-config

The README's basic usage example is a single decode attention call. The tensor shapes carry the contract: q is [num_qo_heads, head_dim], while k and v are [kv_len, num_kv_heads, head_dim].

python
import torch
import flashinfer

q = torch.randn(32, 128, device="cuda", dtype=torch.float16)
k = torch.randn(2048, 32, 128, device="cuda", dtype=torch.float16)
v = torch.randn(2048, 32, 128, device="cuda", dtype=torch.float16)

output = flashinfer.single_decode_with_kv_cache(q, k, v)

If you build from source instead, the README clones with --recursive, which matters because 3rdparty/ and 3rdparty_patches/ are top-level directories. The editable install path has a documented failure mode: with --no-build-isolation, pip does not install build dependencies, and the README notes that FlashInfer requires setuptools>=77. The symptom it names is an AttributeError about prepare_metadata_for_build_editable, and the fix given is upgrading pip and setuptools first.

Where FlashInfer stops: architecture gaps and install-time traps

The GPU support table is the honest part of the README and also the biggest constraint. Turing is SM 7.5, Ampere SM 8.0 and 8.6, Ada SM 8.9, Hopper SM 9.0, and Blackwell spans SM 10.0, 10.3, 11.0, 12.0 and 12.1. Directly under that table the README warns that not all features are supported across all compute capabilities. That sentence is doing a lot of work. FP4 GEMM is described as being for Blackwell GPUs, and the README's own news item dates Blackwell support to v0.4.0. If you are on a T4 and your plan depends on FP4 or on the CuTe DSL path, the architecture table is where that plan dies.

The dependency notes in requirements.txt are the second trap, and they are unusually candid. The file explains that nccl.ep moved out of nccl4py into nccl-extensions, and that depending on nccl4py alone therefore resolves to a wheel with no nccl.ep at all. It also explains why nvidia-nccl-cu13 is deliberately not a base dependency: torch's cu13 wheels pin it exactly, so adding a floor makes the resolver evict torch, and on aarch64 it backtracks to a CPU-only torch wheel. The >=2.30.7 B200 floor is instead enforced at runtime in flashinfer/moe_ep/core/validation/common.py. Read that as a warning: this package lives close to the edge of the CUDA dependency graph, and a resolver that reorders your install can produce a broken environment that looks fine until you hit the EP path.

The README does not document rollback, a downgrade procedure, or a way to pin a backend per call. If you need reproducible kernel selection across driver upgrades, that is not described here.

FlashInfer against Triton and against FlashAttention alone

The comparison people actually search for is FlashInfer versus Triton. The difference is one of scope and of who writes the kernel. Triton is a language and compiler: you write the kernel, you own the tiling and the memory movement, and you get portability across GPU generations for free. FlashInfer is the opposite trade. You call flashinfer.single_decode_with_kv_cache and the library owns the kernel, the backend choice and the architecture-specific tuning. The README describes it as a library and kernel generator, so the generator half exists, but the documented entry point is the unified API, not a Triton-like authoring surface. If your bottleneck is a kernel nobody has written yet, Triton is the tool for that job and FlashInfer is not.

Against FlashAttention on its own, the difference is breadth. FlashAttention-2 and FlashAttention-3 are listed inside FlashInfer as backend implementations, so the attention kernel is not the dividing line. What FlashInfer adds is everything around it: GEMM with per-tensor and groupwise FP8 scaling, grouped GEMM for LoRA and multi-expert routing, fused MoE with DeepSeek-V3 and Llama-4 routing, sorting-free sampling, RoPE, RMSNorm and fused SiLU and GELU gating. A team that only needs attention and already has FlashAttention wired in is not obviously better off adding a second kernel library. A team assembling a full serving loop is.

Maintenance, licensing and what an upgrade actually costs

The repository is not archived, and the last push was on 2026-08-29, the same day as the v0.6.18 release. Nightly builds are published on a near-daily cadence, with nightly-v0.6.18-20260819 and nightly-v0.6.18-20260818 both present. The README documents the nightly install path explicitly, using --pre against the project's own index, followed by a plain install for dependencies and the two CLI commands with --nightly.

The upgrade cost is not the pip line. It is the kernel cache. Because flashinfer-jit-cache is built for a specific CUDA version and a specific FLASHINFER_CUDA_ARCH_LIST, moving to a new CUDA release or adding a GPU generation to your fleet means rebuilding or reinstalling that wheel, and the README's source-build instructions show the architecture list being set by hand. The cubin package is broader (all supported architectures) but still tied to the release you installed. Plan for the fact that a version bump can invalidate pre-compiled work.

Licensing is Apache-2.0, declared in pyproject.toml and shipped as LICENSE with license-files listing LICENSE and LICENSE*.txt. That is a permissive licence with an explicit patent grant, and it carries the usual requirement to preserve notices. The repository also has a NOTICE file and a licenses/ directory, which is where bundled third-party terms live. Read NOTICE and licenses/ before redistributing a wheel, since the library links against and vendors components from several vendors. This is a description of the files present, not legal advice.

Editorial conclusion

Adopt FlashInfer if you are serving LLMs on NVIDIA GPUs from Turing through Blackwell and want attention, GEMM, MoE and sampling behind one API rather than hand-written CUDA. Do not adopt it if your hardware falls outside the supported compute capabilities, or if you need a CPU or non-NVIDIA path, because none is described. Before committing, run flashinfer show-config on the target machine and read the README note that not all features are supported across all compute capabilities, then match the features you actually call against that table.

Frequently asked questions

What does FlashInfer do?

It is a library and kernel generator for inference that provides unified APIs for attention, GEMM and MoE operations, with backend implementations including FlashAttention-2/3, cuDNN, CUTLASS and TensorRT-LLM. It also covers sampling, RoPE, normalization and activations.

How do I install FlashInfer?

The README's quickstart is pip install flashinfer-python. For faster initialization and offline usage it then shows flashinfer install-cubin-wheel and flashinfer install-jit-cache-wheel, and for Blackwell CuTe DSL kernels it uses the extra pip install flashinfer-python[cu13].

Does vLLM use FlashInfer?

The README does not state that vLLM depends on FlashInfer. It lists TensorRT-LLM among the backend implementations and points readers to the project's own documentation and discussion forum, so the vLLM relationship is not something this material confirms.

What is the difference between FlashInfer and Triton?

Triton is a kernel authoring language, while FlashInfer presents unified APIs and selects a backend for your hardware and workload. The README describes FlashInfer as a library and kernel generator, so the documented entry point is calling its kernels rather than writing them.

How do I verify a FlashInfer installation?

The README's verification step is the command flashinfer show-config, which reports the resolved configuration. Running it before writing serving code surfaces a missing cubin or a CUDA mismatch earlier than the first kernel call would.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/flashinfer-ai-flashinfer.svg)](https://hysenlabs.com/projects/flashinfer-ai-flashinfer)