Kernel Design Agents: NVIDIA's Agent Workflow for High-Performance CUDA Kernels
Kernel Design Agents (KDA) is a agent-centric workflow to write high-performance CUDA Kernels.
At a glance
- What is it?
- KDA is an early-stage research workflow from NVIDIA Labs that coordinates coding agents through prompt templates and specialist skills to research, implement, and verify CUDA kernel performance. It targets GPU engineers who want to use agent-driven iteration rather than writing every optimisation step by hand.
- Who is it for?
- KDA suits GPU engineers who already understand CUDA and can write a precise task contract: objective, constraints, validation command, and promotion criteria. Engineers who cannot yet formulate those constraints will not get useful output from the agent.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What KDA Addresses and Who It Targets
Writing a high-performance CUDA kernel involves profiling, analysing memory access patterns, choosing tile sizes, and iterating on micro-benchmarks. That cycle is demanding even for experienced GPU engineers. Kernel Design Agents (KDA) addresses this by treating the optimisation loop as an agent task: a coding agent, supplied with structured prompt templates and two specialist skills, handles the research, drafts an implementation, verifies correctness, and records candidate results.
KDA comes from NVIDIA Labs (NVlabs) and the MIT HAN Lab. It targets ML researchers and GPU engineers who already understand CUDA and are willing to coordinate work through an agent session rather than writing every line by hand. The repository describes itself as an early research prototype, which sets expectations directly. There is no importable library, no GUI, and no cloud service. What you get is a documented process, a set of prompt templates, and the skill files a compatible coding agent can use.
The Architecture: Prompt Templates, Skills, and CLAUDE.md
KDA structures agent work through three layers. First, CLAUDE.md at the repository root gives the agent its standing instructions: what context to maintain, what files to record, and what constitutes a valid candidate. Second, `prompts/basic-flow.md` is a generic starter prompt that you fill in with task-specific details before handing it to the agent. Third, two skills extend what the agent can do on demand: KernelWiki provides reference material for NVIDIA GPU architecture and CUDA programming patterns, and ncu-report-skill lets the agent read and interpret Nsight Compute profiling reports directly.
The workflow itself is documented in `docs/agent-flow.md` and follows eight steps. You create a separate implementation workspace for the target task, define the task contract (objective, constraints, validation command, and promotion criteria), start an agent session in that workspace, supply the filled-in starter prompt, have the agent write a short plan draft to `docs/draft.md`, convert the draft into an executable plan, implement in small iterations with verification after each meaningful change, and record candidates together with profiling evidence and promotion decisions.
This separation of the reference repository from the implementation workspace is deliberate. KDA does not mix the workflow machinery with the artifact being built. The repository recommends a workspace layout including `docs/`, `runs/`, `outputs/`, `profile/`, `benchmark.csv`, and `candidates.jsonl`. The README notes that the exact files can change by domain, but the agent should always record enough context for another engineer to understand what was tried, what passed validation, and why the final candidate was selected.
Installing KDA and Running Your First Agent Session
Installation requires Git, a compatible Claude Code setup, and NVIDIA hardware for profiling. The repository includes two Git submodules, so the clone command must use --recurse-submodules:
git clone --recurse-submodules https://github.com/mit-han-lab/kernel-design-agents.git
cd kernel-design-agentsAfter cloning, link both skills into the Claude Code skills directory so the agent can load them:
mkdir -p ~/.claude/skills
ln -s "$(pwd)/skills/ncu-report-skill" ~/.claude/skills/ncu-report-skill
ln -s "$(pwd)/skills/KernelWiki" ~/.claude/skills/KernelWikiThe ncu-report-skill can also be cloned independently if you prefer not to use the pinned submodule:
mkdir -p ~/.claude/skills && cd ~/.claude/skills
git clone https://github.com/mit-han-lab/ncu-report-skill.gitThe README cautions that a direct upstream checkout of KernelWiki may contain artifact snapshots governed by additional terms that are omitted from this repository's distribution, so the pinned submodule version is the recommended path.
To install the Humanize planning plugin, open Claude Code and run:
/plugin marketplace add PolyArch/humanize
/plugin install humanize@PolyArchWith skills linked and the plugin installed, create a fresh directory for your kernel task, define the task contract in prose, and pass `prompts/basic-flow.md` filled in with your specifics to a new agent session. The agent writes a plan draft to `docs/draft.md` in the workspace, and you or the Humanize plugin then convert that draft into an executable plan before implementation begins.
The Community Kernel Wishlist
KDA includes a community mechanism for collecting kernel optimisation requests. Submissions go in as pull requests against the `wishlist` branch rather than issues. The PR must include a reproducible definition, representative workloads, and your best-known baseline implementation. Request files go in `requests/<github-username>-<kernel-name>/README.md`, and the wishlist PR template provides the short description format.
Once a request is merged, the submission is recorded and optimisation progress links remain on the original PR. The README states directly that the wishlist currently supports NVIDIA B200 and B300 GPUs only. Teams targeting other architectures cannot submit through this mechanism, though running the KDA workflow independently on other hardware is not restricted.
You can browse open wish requests and upvote them via the GitHub PR list, or open a discussion issue if you need help preparing the submission files. The project website at nvlabs.github.io/kda lists further documentation; its source lives on the `pages` branch.
Limitations and When KDA Is the Wrong Choice
KDA describes itself as an early research prototype with no GitHub releases. That matters for anyone evaluating it for a shipping pipeline.
The wishlist mechanism is scoped to NVIDIA B200 and B300 hardware. Teams targeting older architectures cannot contribute to or benefit from the wishlist-based optimisation track.
The workflow is not automated end to end. You must prepare the task contract yourself: writing a clear objective, deciding what the validation command is, and determining the promotion criteria. An engineer who cannot already describe what a well-optimised kernel should look like will not get useful output from the agent. KDA amplifies an existing understanding of CUDA; it does not replace one.
License complexity adds real overhead. The repository carries three separate license regimes: CC BY 4.0 for first-party documentation and prompts, Apache 2.0 for first-party source code, and MIT plus upstream terms for the submodules. The KernelWiki submodule contains restricted CuTe DSL artifacts whose terms are not reproduced here. Anyone redistributing or embedding parts of KDA must read THIRD_PARTY_NOTICES.md carefully before proceeding.
KDA vs. Writing Kernels with Triton
The most common alternative to agent-assisted CUDA development is hand-coding kernels either in CUDA C++ directly or through Triton. Triton, developed by OpenAI, is a Python-based domain-specific language that compiles to GPU code and manages tiling and memory layout automatically. It targets engineers who want higher abstraction and faster iteration without working in raw CUDA.
KDA operates at a different level. It does not change the kernel language or the GPU programming model. Instead, it structures how an agent researches, implements, verifies, and records CUDA work. A team using KDA could target CUDA C++, PTX, or any kernel language the task contract specifies. Triton reduces the decision space through abstraction; KDA keeps the full CUDA decision space and adds agent coordination on top.
The practical consequence is that KDA is the more demanding choice. It requires CUDA knowledge, NVIDIA hardware, and familiarity with agent workflows. Triton is more accessible to Python engineers who want GPU performance without deep CUDA experience. The two are not direct substitutes; they sit at different layers of the GPU programming stack.
License Structure and Current Maintenance Activity
The first-party portions of KDA carry a split license. Documentation, prompts, and skills-style content are under the Creative Commons Attribution 4.0 International License (CC BY 4.0). First-party source code is under the Apache License 2.0. Both licenses permit modification and redistribution with attribution.
The submodules have distinct terms. The ncu-report-skill is MIT licensed. The KernelWiki is MIT for original material, with a note that embedded artifacts retain their upstream terms and that restricted CuTe DSL artifacts are omitted from this distribution. Using the pinned submodule version rather than a direct upstream checkout avoids those restricted artifacts. The full verbatim license texts are in `third_party_licenses/` and `THIRD_PARTY_NOTICES.md`.
The repository is under active development; the last push was on 2026-09-27. The README describes it as an early research prototype and explicitly welcomes community feedback and contributions. The contribution process follows the Developer Certificate of Origin model, documented in CONTRIBUTING.md. There are no formal releases as of the time this article was written.
Editorial conclusion
KDA suits GPU engineers who already understand CUDA and can write a precise task contract: objective, constraints, validation command, and promotion criteria. Engineers who cannot yet formulate those constraints will not get useful output from the agent. Teams targeting hardware other than NVIDIA B200 or B300 should verify that the wishlist mechanism applies to their device before submitting requests. Anyone redistributing parts of the repository must read THIRD_PARTY_NOTICES.md, because the KernelWiki submodule carries restricted CuTe DSL artifacts under terms separate from the Apache 2.0 and CC BY 4.0 first-party licenses. The repository had its last push on 2026-09-27 and is under active development as a research prototype.
Frequently asked questions
What NVIDIA hardware does KDA require?
The community wishlist currently supports NVIDIA B200 and B300 GPUs only. The workflow itself does not restrict other hardware in its documentation, but running and profiling CUDA kernels requires an NVIDIA GPU capable of executing the code being optimised.
How is the KDA repository organised?
The repository holds prompt templates in `prompts/`, agent-facing documentation in `docs/`, skill submodules in `skills/`, and repository-facing agent instructions in `CLAUDE.md`. Implementation work is meant to happen in a separate workspace directory, not inside the KDA repository itself.
Can I contribute a kernel request for a GPU other than B200 or B300?
The community wishlist currently accepts requests targeting NVIDIA B200 and B300 GPUs only. The README does not document a path for contributing requests for other GPU architectures.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvlabs-kda)