# KernelBench pins Python to one minor version, keeps an archived requirements file that disagrees with it, and marks its own metric script as missing

> KernelBench is a benchmark and evaluation environment for asking language models to write GPU kernels that are both correct and faster than the PyTorch operator they replace, across four levels of difficulty. The metric and the level taxonomy are precise and well argued. The packaging is where the project leaks: a manifest that pins one Python minor version while claiming to be the single source of truth for versioning, a second dependency file marked archived that contradicts it, and a metric section carrying an unfinished note.

**ScalingIntelligence/KernelBench** — KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs)

- Repository: https://github.com/ScalingIntelligence/KernelBench
- Website: https://scalingintelligence.stanford.edu/blogs/kernelbench/
- Stars: 1,274 · Forks: 196
- Language: Jupyter Notebook
- License: NOASSERTION
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/scalingintelligence-kernelbench

## The manifest calls itself the version source of truth and pins one Python minor

Two decisions in the packaging manifest deserve to be read together. The first is a comment placed directly above the project table saying this should be the single source of truth for versioning. The second is the value immediately under the package name, a development version of the next minor release. So the declared single source of truth contains a hand-written development version, which means it will drift from whatever the code actually is, and the comment is aspirational rather than descriptive. The manifest also pins the interpreter to an exact minor series rather than a floor, requiring one specific Python version and excluding every later one outright. That is an unusual choice for a 2026 project whose framework dependency asks for a recent stable release, and it is not explained anywhere in the visible documentation. The versioning story outside the manifest is loose by comparison. There are no tagged releases. The documentation refers to versions as branches, with an original release branch and a later one, and says the stable version lives on the default branch. The last push to that branch is dated 2026-03-24, roughly six months before this snapshot, so the code on the branch has not moved in half a year.

## The archived requirements file contradicts the manifest and installs the GPU stack unconditionally

There are two dependency manifests and the older one says so itself. Its first line marks it archived, explains that the project is transitioning to the other file and a different package manager workflow, and says it is provided as a backup for now. Read side by side, they disagree in four ways. The dataset library has a floor of a much later major release in the archived file than in the manifest. The array library is unpinned in the manifest and pinned to an exact version in the archived file. The framework library is unpinned in the manifest and given a floor in the archived file. Most consequentially, the GPU packages, which the manifest puts behind an optional extra for local GPU evaluation, are unconditional requirements in the archived file. So the file marked as a backup is the one that pulls in a compiler toolkit, two alternative kernel languages, a CUDA array library, and a profiler whether or not you have a GPU. The documentation does point at the newer workflow first and mentions the older one as a fallback, but the fallback is not the same install, and a contributor who follows the older advice on a laptop gets a very different environment.

## The metric section carries an unfinished note about its own measurement script

The benchmark's headline metric is defined twice, identically, in two adjacent subsections. A task counts toward the metric only if the generated kernel is both correct and faster than the PyTorch reference by a factor above a threshold, and the speedup is the ratio of the reference operator's wall-clock time to the generated kernel's time. Three concrete instantiations are given: a threshold of one means correct and faster than the baseline, a threshold of two means correct and at least twice as fast, and a threshold of zero is just the correctness rate with the performance condition removed. The design is the interesting part, since folding correctness and speed into one number means a model cannot buy a score by being fast and wrong. The problem is stated immediately after the definition. An HTML comment in the text says the measurement script still needs to be provided. The section that claims to compute the overall benchmark performance names a script, and the evaluation guidelines document is labelled as work in progress. So the metric is specified well enough to reimplement and the tooling to produce it is admittedly incomplete, which is worth knowing before you plan a leaderboard.

## Three levels have problem counts and the fourth does not

The problem set is organised as four levels of increasing scope, and the counts are given for the first three. The first level is single-kernel operators, one hundred problems, described as the building blocks of neural networks such as convolutions, matrix multiplies, and layer normalisation. The second is simple fusion patterns, also one hundred, where a fused kernel should beat separated ones, with the examples given as convolution plus bias plus activation, and matrix multiply plus scale plus activation. The third is whole model architectures, fifty problems, with four named architectures. The fourth level is described in one line as optimising whole model architectures drawn from a model hub, and it carries no count at all. So the benchmark has a bounded first three levels and an open-ended fourth, which is a reasonable design for a set that grows with the model ecosystem, but it also means the headline score depends on which subset of level four you include. Anyone comparing two systems needs to state the level three and level four denominators alongside the score, and the documentation does not supply defaults.

## Correctness is sampled and speed is timed, and the GPU architecture is an argument you supply

The evaluation has two checks and both are counts you control. Correctness is verified against the reference operator a configurable number of times on randomised inputs, so a kernel has to be right on inputs it has never seen rather than on one fixture. Performance is measured by timing the reference and the generated kernel a configurable number of trials and taking the ratio. Randomised inputs and repeated trials are the right choices for a kernel benchmark, because GPU timing on a single run is dominated by launch overhead and clock behaviour. Two configuration details limit reproducibility, and both are the user's problem. The first is the GPU architecture argument, which the documentation flags as something you may need to adjust to match your hardware, since a kernel compiled for one architecture will not run on another and a run on the wrong target either fails or silently differs. The second is precision, which is an argument with a named float format as an example value, and the sentence explaining which precision the reported results use stops mid-word. Neither default is stated in the manifest, so two people running the same command on different hardware can produce two different numbers and neither will be obviously wrong.

## There are two evaluation paths, a local one and a serverless one, chosen by a flag

You can evaluate on your own GPU or on rented GPUs, and the project supports both with slightly different plumbing. A convenience script takes one candidate source and one reference source and reports correctness and speedup, with the destination chosen by an argument that has a local value and a serverless value. The serverless path uses a third-party GPU platform, and setting it up means creating an account and running a token command, after which a separate script named for the serverless path handles generation and evaluation together. Everything is invoked through the modern package manager, which selects the right environment for you, and a single-problem run takes the dataset source, the level, the problem id, the server type, and the model name as arguments:
```bash
# for example, run level 2 problem 40 from huggingface and use google gemini 2.5 flash for generation

uv run python scripts/generate_and_eval_single_sample.py dataset_src=huggingface level=2 problem_id=40 server_type=google model_name=gemini/gemini-2.5-flash

# dataset_src could be "local" or "huggingface"
# add .verbose_logging for more visibility
```
That argument style is a configuration system rather than a flag parser, which suits sweeps, and the dataset can be read from a local copy or fetched from the model hub. There is also a tutorial notebook in the notebooks directory and a hosted copy of it for people with no GPU at all.

## Seven provider keys, three kernel languages, and an instruction to clone the repository instead

Three things frame how much this project expects of you. First, credentials. An example environment file lists keys for seven hosted model providers, including three that are placeholders shaped like real keys, plus a separate variable for a locally served model with three named server implementations in the comment. Model calls go through a multi-provider library, with the archived dependency file noting that this library handles cloud providers while a second client handles local ones. Second, the GPU side. The optional extra adds a compiler toolkit, two alternative kernel authoring languages, a CUDA array library, and a profiler, and the AMD path needs a recent ROCm release, with the documented procedure being to remove the framework package and add a build from an AMD-specific index, and a recommendation to use a container because of the stated complexity. Third, and most unusual, the documentation tells you not to treat this as a library. It says the repository provides core functionality and convenience scripts, that it is not intended to provide complex scaffolds that solve the task, and that you should clone and modify it or consume it as a submodule. It also says the project continues to be extended to other kernel languages and to AMD hardware, while the last push is dated 2026-03-24.

## Conclusion

KernelBench is worth using if you are evaluating models on kernel generation, because the level taxonomy, the two-axis correctness and speed check, and the threshold metric are all specified rather than implied, and the scripts take the arguments that matter in plain sight. Treat it as a research artefact rather than a maintained tool. The repository has no tagged releases at all, versioning is by branch, the manifest carries a development version, and the last push to the default branch is dated 2026-03-24. Before you build results on it, check four things. Which Python you are on, since the manifest requires one minor version exactly and a newer interpreter will simply refuse to install. Which dependency file you are reading, because the archived one lists a lower floor for the dataset library, pins the array library, and treats the GPU packages as unconditional. Which precision and GPU architecture you set, because both are arguments you supply and both change the numbers. And which script produces the headline metric, because the section that defines it still carries a note saying the measurement script is not yet provided.

## FAQ

### What is KernelBench?

It is a benchmark and evaluation environment that asks language models to generate correct and efficient CUDA or other language kernels for PyTorch programs on a target GPU. It was published as an ICML 2025 paper with an arXiv entry, a blog post, and a dataset on the model hub, and it is licensed under a file at the repository root while its metadata records no licence value.

### How does KernelBench score a generated kernel?

On two axes at once. Correctness is checked against the reference PyTorch operator a configurable number of times on randomised inputs, and performance is a wall-clock ratio between the reference and the generated kernel. Both conditions must hold for a task to count, which stops a model from scoring by being fast and wrong.

### What do the four KernelBench levels contain?

Single-kernel operators with one hundred problems, simple fusion patterns with one hundred, whole model architectures with fifty, and a fourth level covering whole architectures drawn from a model hub that carries no problem count at all. The first three levels are bounded and the fourth is open-ended, so a score needs its denominators stated.

### Which Python version does KernelBench need?

The manifest pins the interpreter to one exact minor version and excludes later ones, and the archived requirements file documents a conda environment on the same version. The framework dependency asks for a recent stable release, so the pin is narrower than the rest of the stack would suggest and is not explained in the visible documentation.

### Can I run KernelBench without a GPU?

For setup, yes. The base install is documented as working without a local GPU, with GPU packages behind an optional extra. For running and profiling kernels, no, unless you set up an account on a serverless GPU platform, run its token command, and use the separate script for that path, or work through the provided tutorial notebook on a hosted notebook service.

### How is KernelBench versioned?

Loosely. There are no tagged releases, the documentation refers to versions as branches, the manifest carries a development version under a comment saying it should be the single source of truth for versioning, and the dataset on the model hub is updated to an earlier release than the code. The last push to the default branch is dated 2026-03-24.

## Sources

- [Issues](https://github.com/ScalingIntelligence/KernelBench/issues)
- [Project website](https://scalingintelligence.stanford.edu/blogs/kernelbench/)
- [README](https://github.com/ScalingIntelligence/KernelBench/blob/main/README.md)
- [ScalingIntelligence/KernelBench on GitHub](https://github.com/ScalingIntelligence/KernelBench)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/scalingintelligence-kernelbench
