DeepSelect: CUDA TopK Kernels for DeepSeek Sparse Attention and Sampling
DeepSelect: TopK kernels for DeepSeek Sparse Attention (DSA) and Samplers
At a glance
- What is it?
- DeepSelect is an open-source CUDA kernel library that provides a high-performance TopK implementation used in DeepSeek's sparse attention mechanism and its sampler. It achieves two to twenty times the throughput of PyTorch's built-in torch.topk on supported workloads, with the constraint that topk must not exceed 4096.
- Who is it for?
- DeepSelect is the right tool for inference teams running DeepSeek models or building systems that need fast GPU-side TopK on bfloat16 or float32 inputs with topk values of 4096 or lower. It is not a general-purpose TopK replacement: inputs with topk above 4096 are unsupported, the input tensor must satisfy a stride alignment requirement, and NaN-checking is always on and will abort by default.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 19 days ago.
- What is it written in?
- Mainly Cuda, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What DeepSelect Solves and Who Needs It
TopK is a fundamental operation in large-language-model inference: it appears in sparse attention to select the most relevant keys and in the sampler to pick the top candidate tokens. PyTorch ships torch.topk as a general CPU and GPU implementation. It is correct across all input shapes and dtypes, but it is not optimized for the specific input distributions that appear in production LLM inference.
DeepSelect targets exactly those distributions. It is the TopK kernel used internally in DeepSeek Sparse Attention (DSA), which powers the DeepSeek V3.2, V4, and V4.1 models. The README reports two to twenty times the effective memory bandwidth of torch.topk on the supported workloads. The metric used is effective memory bandwidth rather than floating-point throughput, because TopK does no floating-point arithmetic: it reads the input once and writes the top-k indices and optionally the top-k values.
The library is aimed at inference engineers building or optimizing pipelines that run DeepSeek models, or at anyone whose workload matches the two documented scenarios.
Two Supported Scenarios and Their Constraints
DeepSelect defines two distinct workload scenarios with different dtype and shape requirements.
The Lightning Indexer scenario uses bfloat16 input. It supports any batch size and any vocabulary size. The topk value must be 4096 or less; larger values are not supported. The README recommends disabling sorted_index unless the output must be ordered, and setting return_value to False when values are not needed, since both flags reduce throughput.
The Sampling scenario uses float32 input. It supports any batch size, and the vocabulary size should be around 128,000 tokens. The topk constraint is the same: 4096 or less.
Beyond the topk limit, the input tensor has a stride requirement. The row stride of the input must be a multiple of the value returned by deep_select.get_stride_requirement()[0] bytes, and the last dimension must be contiguous. For unaligned inputs, padding is necessary before calling the kernel. Output tensors allocated by the call also have a stride alignment, so they may be non-contiguous.
Installing DeepSelect from Source
The library has no published PyPI release and must be built from source. It requires CUDA and a compatible C++ toolchain. The installation sequence is:
git clone https://github.com/deepseek-ai/DeepSelect.git
cd DeepSelect
git submodule update --init --recursive
pip install -v .The -v flag passes verbose output so you can see the CUDA compilation steps, which can take several minutes. The setup.py shows that the build generates a large number of CUDA kernel instantiation files under csrc/cuda_kernels/, covering combinations of dtype, thread count, occupancy, and block size. These are generated via scripts/generate_instantiations.py.
After installation, the benchmark in tests/test.py reports performance ratios against torch.topk:
python3 tests/test.py --perf-onlyThis runs the benchmark in performance-only mode and prints the effective memory bandwidth ratio on the current hardware.
Using the API: Variable-Length Row Support
The core entry point is deep_select.topk. For the standard case, you pass the input tensor, the target topk, and optional flags for sorting and output dtype.
For workloads where rows have different lengths, the end parameter sets a per-row upper bound. Rows shorter than topk are padded using the fill values you supply:
batch_size, vocab_size = 2, 129280
x = torch.randn(batch_size, vocab_size, dtype=torch.float32, device="cuda")
end = torch.tensor([129280, 100000], dtype=torch.int32, device="cuda")
values, indices = deep_select.topk(x, 1000, end=end, sorted=True,
indices_type=torch.int64)The note in the README points out that 129280 is a multiple of 256, which satisfies the stride alignment requirement for float32. For cases where alignment is not naturally satisfied, the input must be padded before the call.
The full function signature is documented in deep_select/interface.py. An optional output_idx parameter lets you write indices into a pre-allocated buffer, which must meet the same stride requirement as an output tensor.
NaN Handling and Abort Behavior
DeepSelect always checks for NaN values. The default behavior when a NaN is found is to invoke trap() and abort the process. This is controlled by the abort_when_nan_found parameter, which defaults to True.
There is one exception: rows whose length is less than or equal to topk are never NaN-checked. The README states this explicitly. The implication is that if you are using the end parameter to specify short rows and some rows are shorter than topk, those rows bypass the NaN check entirely.
For inference systems that receive external inputs and need NaN detection to raise a Python exception rather than abort the process, the behavior of abort_when_nan_found=False should be verified against the interface documentation in deep_select/interface.py. The README does not describe what the library does when NaN is found and abort is disabled.
Output Allocation, Pre-Allocated Buffers, and the Deep Dive
By default, deep_select.topk allocates both the values output and the indices output internally. The allocated tensors satisfy the stride alignment requirement for output tensors, which means they may be non-contiguous. Code that assumes contiguous outputs may need adjustment.
For inference systems that need to control memory allocation, the output_idx parameter accepts a pre-allocated tensor for the indices output. That buffer must satisfy the same stride alignment as an internally allocated output tensor. The full parameter list is documented in deep_select/interface.py.
On 2026-09-10, the project released a detailed technical analysis of the algorithm and implementation. The analysis is available in English at docs/DeepSelect-deep-dive.md and in Chinese at docs/DeepSelect-deep-dive.zh.md. Teams evaluating the library for a production deployment should read this document before integrating, since it explains the kernel selection strategy and the reasoning behind the two workload scenarios.
Kernel Instantiation Design and Build Time
The build generates a large number of CUDA kernel instantiation files. The setup.py shows that csrc/cuda_kernels/v3/instantiations/ contains files named by their parameters: value dtype, output index dtype, sorted-by-index flag, sorted-by-value flag, return-value flag, max topk, thread count, occupancy, block sizes, and TMA version. Each combination gets its own .cu file.
This instantiation strategy compiles all the kernel variants ahead of time so the runtime can select the optimal kernel for a given input without JIT compilation overhead. The cost is a long build time during pip install and a large compiled artifact. The -v flag on the install command shows the compilation progress.
The scripts/generate_instantiations.py script produces these files, which means the instantiation coverage can be extended or customized by regenerating them with different parameters if needed.
DeepSelect Compared to torch.topk
PyTorch's torch.topk is the natural baseline. It ships with PyTorch, works on CPU and GPU, handles any dtype, any batch size, and any topk value without constraints on stride alignment or maximum k. It is correct and well-tested across a wide range of inputs.
DeepSelect narrows that generality to gain throughput. It targets two specific dtype-and-shape profiles, imposes a topk ceiling of 4096, requires stride-aligned inputs, and is a source-only build. The performance advantage the README describes applies within those profiles: two to twenty times the effective memory bandwidth against torch.topk on the same input. Outside those profiles, torch.topk is the safer choice.
For teams running DeepSeek sparse attention or sampling at scale, where the workloads naturally fall within the supported scenarios, the reported throughput improvement is significant. For teams with general TopK needs across varied shapes and dtypes, torch.topk avoids the constraints.
Editorial conclusion
DeepSelect is the right tool for inference teams running DeepSeek models or building systems that need fast GPU-side TopK on bfloat16 or float32 inputs with topk values of 4096 or lower. It is not a general-purpose TopK replacement: inputs with topk above 4096 are unsupported, the input tensor must satisfy a stride alignment requirement, and NaN-checking is always on and will abort by default. Teams whose workloads fall outside the two documented scenarios should benchmark against torch.topk before switching. The last push was on 2026-09-10.
Frequently asked questions
Does DeepSelect work with topk values larger than 4096?
No. The README states that topk must be 4096 or less in both the Lightning Indexer and Sampling scenarios. Larger values are not supported by the current implementation.
What happens if the input tensor is not stride-aligned?
The row stride of the input must be a multiple of the value returned by deep_select.get_stride_requirement()[0] bytes, and the last dimension must be contiguous. The README states that padding is necessary for unaligned inputs before calling the kernel.
Which DeepSeek models use DeepSelect internally?
According to the README, DeepSelect is the TopK kernel used in DeepSeek Sparse Attention (DSA), which is used in DeepSeek V3.2, DeepSeek V4, and DeepSeek V4.1 models.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/deepseek-ai-deepselect)