Model or dataset
datawhalechina/llm-algo-leetcode avatar
datawhalechina/llm-algo-leetcode

llm-algo-leetcode: a runnable notebook curriculum for LLM algorithms and GPU systems

LLM algorithm practice lab with theory, solutions, and test cases.《大模型算法与系统教程》面向大模型入门到进阶的算法实战教程,覆盖原理讲解、答案解析、测试用例与 CUDA/Triton 实战。

620 stars124 forksJupyter NotebookNOASSERTION

At a glance

What is it?
The datawhalechina project packages LLM algorithm practice as Jupyter notebooks with theory, answer cells and test cases, then walks from PyTorch down to Triton and CUDA. Part 04 is still under construction, and the licence files need reading before reuse.
Who is it for?
Adopt it if you already write PyTorch and want a structured path into Triton kernels, memory accounting and profiling, and you are willing to work in Chinese-language notebooks. Do not adopt it if you need a finished CUDA C++ course, an English-language reference, or a library to depend on in production; Part 04 is marked as in progress and Part 05 is reserved.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 3 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The gap llm-algo-leetcode fills between algorithm tutorials and kernel work

Most LLM learning material stops at a working PyTorch model. The step after that, where you ask why an attention block costs what it costs and how to rewrite it as a GPU kernel, usually has no graded exercises attached. This repository is an attempt to supply those exercises. The README describes it as "Runnable notebooks for LLM algorithms and systems" and organises content into five numbered parts, from prerequisites through PyTorch algorithms, Triton kernels, CUDA and system optimisation, with a fifth part reserved for CUDA Rust.

The intended audience is stated plainly: people learning LLM algorithms who want to understand Transformer training, fine-tuning, inference and compression through PyTorch notebooks, and people who want to go further into memory, profiling, communication and GPU optimisation. A third group is named as project practitioners who want reproducible comparisons of throughput, latency, memory, quality and cost using benchmarks and real GPU experiments. That third group is the interesting one, because it implies the material expects you to measure rather than accept a claim.

The structure is not strictly linear. The README says you do not have to read from part 00 in order, and it recommends starting at Part 02 to build practical feel, then going back to Parts 00 and 01 as needed. That is a reasonable choice for anyone who already writes Python and PyTorch, and an awkward one for a complete beginner, who will find the recommended entry point assumes tensor and autograd fluency.

How the notebook structure and topic routes work together

Two organising systems run in parallel. The first is the numbered parts. Part 00 covers prerequisites in five groups of four lessons each: Python basics and data representation, PyTorch tensors and autograd, model construction, training intuition, and debugging with performance awareness. Part 01 covers hardware, maths and systems in five groups and 33 lessons, including numeric foundations, single-GPU memory access, multi-GPU communication, heterogeneous scheduling and operator programming, and compiler optimisation. Part 02 is the PyTorch algorithm core, ten groups spanning basic operators, model architecture, the training and fine-tuning loop, preference optimisation, backpropagation and memory, inference optimisation, advanced decoding, compression and quantisation, distributed parallelism, and a project section. Part 03 is Triton kernel development, 15 lessons across basics, a transition section, attention optimisation, inference optimisation and projects. Part 04, CUDA C++ and system optimisation, is listed as in progress with 16 lessons planned.

The second system is topic_discussion, which cuts across parts. Each topic has an intro page and a casebook page. The main routes are supervised fine-tuning, inference optimisation, memory optimisation, and operator and compiler optimisation, the last of which the README marks as in progress. Supporting topics cover quantisation and compression, communication and parallelism, profiling, and post-training and alignment. Each topic entry names the parts it draws from, so the fine-tuning route spans Parts 01 and 02 while the memory route spans Parts 00 through 02.

A learner who wants one specific skill, say reading a profiler trace to decide whether a kernel is memory-bound, can enter through the profiling casebook rather than reading three parts in sequence. The cost is that casebooks assume the vocabulary introduced in the parts they cite. The repository also carries a team_study directory for group study records and a benchmarks directory, though the README does not describe what the benchmark files contain.

Setting up llm-algo-leetcode and running your first notebook

The repository ships two environment files, environment.yml and environment-gpu.yml, plus a requirements.txt whose entire content is the two lines shown below. The split between a base environment and a GPU-specific one matters because Parts 03 and 04 assume a CUDA device.

text
-r requirements/base.txt
-r requirements/dev.txt

The README does not give conda or pip commands, so there is no install command to quote from the repository. What it does give is a starting point: it recommends beginning at Part 02, opening the intro page for that part and picking the first group, 2.1 basic operators. Each lesson is a notebook with a problem area, an answer area and a basic verification step, which is what the README means by "Notebook-first". Expect to edit cells rather than read them.

The repository root also contains verify.py, and the README does not document its command-line interface, so check the file itself for how it is invoked before pointing it at a directory. If you are working through the Triton material, confirm your GPU is visible to PyTorch first, because a notebook that cannot allocate a device will fail before it reaches the kernel logic.

Where the material is thin: Part 04, language, and verification scope

The clearest limitation is completeness. Part 04, CUDA C++ and system optimisation, is marked as in progress in the README's asset table, and the operator and compiler optimisation topic route carries the same marker. Part 05, CUDA Rust, is reserved with no content. Anyone whose goal is CUDA C++ should treat the current repository as preparation rather than a finished course, and should check the 04_CUDA_and_System_Optimization directory for what actually exists before planning around it.

The second constraint is language. The README is bilingual and the notebooks are written for a Chinese-speaking audience, with the English version presented as a translation of the same page. If your team reads English only, you are relying on translated material for the parts where precision matters most, such as memory layout and scheduling terminology.

The third is verification scope. The repository includes verify.py and a benchmarks directory, and the README describes test cases and basic verification as a feature. It does not state which notebooks are covered by the automated checks, so a passing run of verify.py is not evidence that every lesson has been executed. Treat verification as per-notebook unless you confirm otherwise.

Finally, the licence. The repository carries both LICENSE and LICENSE-CODE, and the project metadata reports the licence as NOASSERTION, meaning no standard identifier was detected. The README does not summarise the terms. If you plan to reuse notebook code in a product or a paid course, read both files, and if the split between text and code licensing is unclear, that is a question for your own legal review rather than something the repository answers.

llm-algo-leetcode compared with a general algorithm practice site

The name invites a comparison with LeetCode-style practice, and the difference in approach is worth stating. A general algorithm site gives you a problem statement, a judge and a pass or fail result. The unit of work is a function with defined inputs and outputs, and the environment is fixed by the platform.

This project keeps the problem-and-answer shape but changes the unit of work to a notebook cell and the judge to your own hardware. The exercises are about model computation, memory behaviour and kernel implementation rather than data structures. A question might ask you to implement an attention variant, then measure it. The answer area shows one implementation, and the verification step checks that it runs and produces the expected result, not that it beats a reference runtime on a hidden test set.

That makes it weaker as a competitive practice ground and stronger as a guided lab. There is no leaderboard and no hidden test suite to game. The trade-off is that you cannot get a quick correctness signal without a working environment, and for the Triton and CUDA material that means a GPU. A learner on a laptop without an NVIDIA device is limited to Parts 00 through 02 and the parts of the topic routes that do not require kernel execution. The repository does not document a CPU fallback for the kernel sections.

Maintenance, upgrade cost and licence questions

The last push to the default branch was on 2026-09-03, and the repository is not archived. That is recent enough that the material is being touched, and the README's own status column says Parts 00 through 03 are complete and under continuous optimisation while Part 04 is under construction. Continuous optimisation is also a warning: lesson content can change under you, so if you fork the notebooks for a course, pin the commit you built on.

Upgrade cost is dominated by the GPU toolchain rather than by Python. environment-gpu.yml pins whatever CUDA and PyTorch combination the maintainers tested against. Moving to a newer driver or a newer GPU generation means re-resolving that file, and the Triton and CUDA lessons are the ones most likely to break, because kernel APIs and compilation behaviour change between releases. The requirements.txt indirection through requirements/base.txt and requirements/dev.txt means a dependency bump is a two-file edit, which is a small convenience.

On licensing, the presence of separate LICENSE and LICENSE-CODE files suggests the prose and the code are licensed differently, which is a common arrangement for tutorial repositories. The metadata does not resolve to a standard identifier, and the README does not restate the terms. Nothing here should be read as legal advice; the practical step is to open both files and confirm whether commercial reuse of the notebook code is permitted before you build on it.

Editorial conclusion

Adopt it if you already write PyTorch and want a structured path into Triton kernels, memory accounting and profiling, and you are willing to work in Chinese-language notebooks. Do not adopt it if you need a finished CUDA C++ course, an English-language reference, or a library to depend on in production; Part 04 is marked as in progress and Part 05 is reserved. Verify three things first: which environment file matches your CUDA driver, whether verify.py covers the notebooks you intend to use, and what LICENSE and LICENSE-CODE actually permit, since the repository is classified NOASSERTION and the README does not restate the terms.

Frequently asked questions

Does llm-algo-leetcode require a GPU?

Parts 03 and 04 cover Triton and CUDA kernel development and the repository ships a separate environment-gpu.yml, which indicates a GPU is expected for that material. Parts 00 through 02 are PyTorch-focused, and the README does not document a CPU fallback for the kernel sections.

Which part of llm-algo-leetcode should a beginner start with?

The README recommends starting from Part 02 to build practical algorithm feel, then going back to Parts 00 and 01 as needed. It explicitly says you do not have to read from Part 00 in order.

Is llm-algo-leetcode finished?

Parts 00 through 03 are listed as complete and under continuous optimisation. Part 04, CUDA C++ and system optimisation, is marked as in progress, and Part 05, CUDA Rust, is reserved with no content.

What licence does llm-algo-leetcode use?

The repository contains both LICENSE and LICENSE-CODE files, and the project metadata reports the licence as NOASSERTION with no standard identifier detected. The README does not summarise the terms, so the files themselves are the source.

Official sources

  1. datawhalechina/llm-algo-leetcode on GitHub
  2. Issues
  3. Project website
  4. README
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/datawhalechina-llm-algo-leetcode.svg)](https://hysenlabs.com/projects/datawhalechina-llm-algo-leetcode)
Community notes

Community notes