HPC-Ops: Tencent's Production LLM Inference Kernel Library for NVIDIA GPUs
High Performance LLM Inference Operator Library
At a glance
- What is it?
- HPC-Ops is a C++ and CUDA operator library from Tencent's Hunyuan AI Infra team that provides production-optimized kernels for the hot paths in LLM inference: attention, MoE, GEMM, sampling, normalization, and communication-compute fusion. It exposes a Python API and is benchmarked against vLLM, SGLang, FlashInfer, NCCL, cuBLAS, and TensorRT-LLM.
- Who is it for?
- HPC-Ops is the right choice for an inference team running LLM serving at scale on NVIDIA H20 or H100 hardware that wants to drop in optimized kernels rather than relying on generic library implementations. The operators target production serving latency and throughput, not offline batch processing.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 23 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What HPC-Ops Is and Who Operates It
LLM inference at scale runs the same computational hot paths repeatedly: prefill and decode attention over growing KV caches, sparse expert routing in MoE models, GEMM for each transformer layer, sampling from the output distribution, and normalization followed by AllReduce in tensor-parallel setups. Each of these can be implemented generically or heavily optimized for the specific GPU architecture and batch shape in use.
HPC-Ops takes the optimization path. The README describes it as a production-grade operator library developed by Tencent's Hunyuan AI Infra team and used in large-scale production inference at Tencent. The target hardware is NVIDIA H20 GPUs, with build support for SM90, SM100, and SM103 architectures. The library targets inference engineers integrating kernels into serving frameworks like vLLM or SGLang, not end users of a complete serving system.
Building HPC-Ops for a Target GPU Architecture
The build system uses Python's setuptools wrapping CMake. Each GPU architecture gets its own build directory and its own compiled module. The build process detects the local GPU architecture automatically if SM_ARCH is not set.
For a default local build:
python3 setup.py buildOr using the Makefile:
makeThe Makefile documents the architecture targeting approach: `make sm90` builds a wheel for SM90 only. `SM_ARCH=90,100 make` builds two modules, one for sm90 and one for sm100. A release wheel targeting all supported architectures uses:
make wheelThe three supported values for SM_ARCH are 90, 100, and 103. The same list appears in CMakeLists.txt, setup.py, and the Makefile; adding a new architecture requires updating all three files plus a corresponding HPC_ARCH_EVAL block in src/utils/utils.h. At import time, the Python package selects the module matching the local GPU.
The Kernel Catalog: Attention, MoE, and Sampling
The June 2026 updates documented in the README describe six operator categories.
Dynamic Decode Attention addresses tail latency in mixed-length decode batches. Static split-k scheduling cannot adapt to variable KV-cache lengths within a batch. HPC-Ops introduces a dynamic task scheduling path that splits requests into uniform KV tiles, assigns tiles before each decode step with a greedy bin-packing strategy, and balances work across CTAs. Attention kernels consume the generated task map; a combine kernel merges split-k results.
Sparse Attention provides FP8 block-sparse prefill attention for long-context workloads. A precomputed block mask lets the kernel skip irrelevant KV tiles entirely. Per-tile FP8 scaling preserves numerical quality across the sparse pattern.
The Fused Sampler collapses the full decode sampling pipeline (repetition penalty, temperature scaling, top-k, top-p, softmax, random sampling, and penalty-mask update) into two CUDA kernels. A lighter temperature-only fast path dispatches automatically when applicable. The penalty mask update stays on GPU rather than requiring a CPU roundtrip.
Route GEMM and Mixed-Precision Compute
MoE router GEMM and state-compress GEMM need BF16 activations but precision-sensitive FP32 weights. Modern Tensor Cores do not provide native FP32 throughput, so naive FP32 GEMM falls back to CUDA cores, while using BF16 or TF32 directly introduces accuracy degradation.
HPC-Ops solves this by decomposing each FP32 weight into a high BF16 component and a low BF16 residual component with a fixed scale of 1/256, then computing the result as a fused linear combination of two BF16 Tensor Core GEMMs. Both GEMMs run inside one kernel, sharing input movement and keeping intermediate accumulators in registers. The final result is written once. The README states this approach preserves FP32-level accuracy with significantly higher throughput than a CUDA-core FP32 fallback.
The Fused MoE path targets low-latency MoE inference by fusing routing, Gate-Up GEMM, activation quantization, Down GEMM, and top-k weighted reduction into one pipelined execution. Routing and index preprocessing reduce global atomic pressure with shared-memory counting.
Communication-Compute Fusion with NVLink
Tensor-parallel inference runs AllReduce, residual add, and RMSNorm as separate stages, creating extra kernel launches and repeated HBM reads and writes around an already communication-heavy path. HPC-Ops fuses AllReduce, Residual Add, and RMSNorm into NVLink-native kernels.
Two modes cover different working shapes. The high-throughput mode uses CUDA multicast for large-token prefill-like shapes. The low-latency mode uses a Lamport P2P two-kernel design with PDL overlap for small-token decode shapes. Both modes use a two-shot communication schedule to reduce overhead while keeping normalization fused into the collective path.
The README frames this as addressing the specific bottleneck where tensor-parallel inference creates a logical operation (normalize the reduced activation plus residual) that is split across multiple sequential kernels by default. The fused kernel eliminates the intermediate HBM traffic.
License, Integration Requirements, and Limitations
The repository lists its license as NOASSERTION, which means the license file does not use a standard SPDX identifier. The actual LICENSE.txt file is present in the repository but its terms are not reproduced in the README. Any organization considering commercial use of HPC-Ops should read LICENSE.txt and seek legal review before deployment.
HPC-Ops provides kernels, not a complete inference server. Integrating it into a serving workflow requires wrapping the Python API into an existing framework such as vLLM or SGLang, since HPC-Ops does not include request routing, batching, KV cache management, or tokenization.
The library targets NVIDIA H20 GPU architecture (SM90) as its primary benchmark target, with SM100 and SM103 listed as supported. AMD GPUs and other accelerators are not covered. There are no GitHub releases, so users work from the main branch. The repository does not document a stable Python API version; the interface may change between commits.
vLLM as a Framework-Level Alternative
vLLM is a widely used open-source LLM inference and serving engine. It provides a complete serving stack: an HTTP server, continuous batching, PagedAttention for KV cache management, multi-GPU tensor parallelism, and support for many model architectures. vLLM uses its own attention kernels and can also integrate external kernel libraries.
The comparison is not really between HPC-Ops and vLLM at the same level: they operate at different layers. vLLM is the serving framework; HPC-Ops is a kernel library that targets the same attention, MoE, and sampling hot paths that vLLM implements. The README explicitly mentions benchmarking against vLLM.
For a team that needs a complete, deployable serving system quickly, vLLM is the starting point and HPC-Ops may be an optimization layer later. For a team that already has a serving framework and wants to replace specific kernels with higher-performance implementations optimized for H20 hardware and production workloads, HPC-Ops is the appropriate tool.
Editorial conclusion
HPC-Ops is the right choice for an inference team running LLM serving at scale on NVIDIA H20 or H100 hardware that wants to drop in optimized kernels rather than relying on generic library implementations. The operators target production serving latency and throughput, not offline batch processing. The project is not a complete inference framework, so it requires integration work with a serving system like vLLM or SGLang. The license file is not a standard SPDX identifier; legal review is required before using HPC-Ops in a commercial product. The last push was on September 7, 2026.
Frequently asked questions
How do I build HPC-Ops for a specific NVIDIA GPU architecture?
Set SM_ARCH to the target compute capability (90, 100, or 103) before running make or python3 setup.py build. The Makefile also provides per-architecture shortcuts: make sm90 builds a wheel for SM90 only. Without SM_ARCH set, the build detects the local GPU architecture automatically.
What is the difference between HPC-Ops and vLLM?
HPC-Ops is a CUDA kernel library that provides optimized implementations of attention, MoE, GEMM, and sampling operators. vLLM is a complete LLM serving framework with request routing, batching, and a serving API. HPC-Ops targets teams who want to replace specific kernels in a framework; vLLM is a full serving system. The README benchmarks HPC-Ops against vLLM.
Does HPC-Ops support AMD or other non-NVIDIA hardware?
The README states HPC-Ops targets NVIDIA H20 GPUs, and the build system lists three supported SM architectures: 90, 100, and 103. AMD GPUs and other accelerator types are not mentioned in the repository.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tencent-hpc-ops)