flashinfer-ai
★ 6,038flashinfer
FlashInfer: Kernel Library for LLM Serving. It provides unified APIs for attention, GEMM, and MoE operations with multiple backend implementations including FlashAttention-2/3, cuDNN, CUTLASS, and TensorRT-LLM.
HYSEN LABS DIRECTORY
Verified repositories, classifications and analysis from the flashinfer-ai organization.
FlashInfer: Kernel Library for LLM Serving. It provides unified APIs for attention, GEMM, and MoE operations with multiple backend implementations including FlashAttention-2/3, cuDNN, CUTLASS, and TensorRT-LLM.