Model or dataset
harleyszhang/llm_note avatar
harleyszhang/llm_note

harleyszhang/llm_note: a Chinese-language LLM inference and CUDA study repository

LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.

891 stars90 forksPythonLicense varies

At a glance

What is it?
This repository collects notes on transformer structure, quantization, FlashAttention, Triton kernels and vLLM source reading, plus a paid course on building an inference framework. The notes are free; the course is not.
Who is it for?
Adopt it as a reading list if you can read Chinese and want a curated path through FlashAttention, quantization and Triton kernel work. Skip it if you need runnable code, an installable package or English documentation, because the repository is a set of Markdown notes with no setup instructions and no stated licence.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 43 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What harleyszhang/llm_note actually contains

This is a notes repository, not a library. The README describes it as LLM notes covering model inference, high performance computing programming, transformer model structure, and vLLM framework code analysis. The top level of the repository holds five numbered directories plus an images folder: 1-hpc_note, 2-transformer_model, 3-llm_infer_optimize, 4-llm_parallel, 5-framework. The README's own table of contents uses a different numbering scheme (transformer model, LLM quantization inference, LLM inference optimization, high performance computing, framework analysis), so the directory names and the README headings do not line up one to one. Anyone cloning the repository should expect to navigate by directory listing rather than by the README map.

The intended reader is an engineer preparing for or working in GPU inference roles. The README states the course behind the notes is oriented toward interviews and toward high performance computing and inference framework positions, and it includes a section of interview questions grouped by topic. That framing runs through the whole repository: the notes read as study material for a specific job market rather than as reference documentation for a shipping system.

The primary language is Python, which reflects the subject matter (PyTorch, Triton, model code) rather than the repository's own content, since what is stored here is prose and images. There are no releases and no homepage. The last push was on 2026-08-19, so the repository is not archived and has been touched within the past two months.

The reading path through transformers, quantization and FlashAttention

The transformer section collects paper walkthroughs and code explanations: the transformer paper, a transformer implementation, llama1-3 model structure, ViT, GPT1-3, sinusoidal positional encoding, and MLA structure with code and optimization notes. The quantization section covers SmoothQuant and AWQ, each with a paper reading and a source-code analysis. Note that the README's link for the AWQ paper reading points at the SmoothQuant file path, so one of the two AWQ entries is mislinked in the index.

The inference optimization section is the densest. It splits into three parts: overall performance analysis (translations of LLM inference survey papers and a summary of inference serving framework features), algorithm-level optimization (online-softmax, FlashAttention 1, 2 and 3 individually plus a combined summary, prompt cache, and a note on vLLM's CUDA graph usage), and parallel acceleration (tensor parallelism). The high performance computing section then goes down to the metal: six Triton kernel development notes, a set of NVIDIA GPU architecture and multi-GPU topology notes, a Roofline model explanation, and CUDA notes covering the programming model, memory organization, execution model, kernel launch configuration and thread indexing, optimization strategies, and streams.

The framework section reads source: TGI, vLLM's inference flow, a vLLM optimization overview, and LightLLM. Taken together the sequence is deliberate. You start with the model architecture, learn why inference is memory-bound through Roofline, then see how FlashAttention and paged attention attack that, and finally read how serving frameworks assemble the pieces. That ordering is the repository's main value, and it is not something a scattered set of blog posts gives you.

Installing and using harleyszhang/llm_note: there is nothing to install

The README gives no installation steps, no package name, no environment variables and no runnable entry point, because the repository is documentation. The only way to use it is to clone it and read the Markdown files, and the images referenced by those files live in the images directory at the repository root.

bash
git clone https://github.com/HarleysZhang/llm_note.git
cd llm_note
ls

The listing should show the five numbered directories, README.md and images. After that, read the files directly. For example, to follow the Triton kernel notes, which the README lists under 4.1:

bash
ls 4-hpc_basic

The README names files such as trito内核开发基础0.md through trito内核开发基础5.md (the README spells Triton as trito in these filenames). Because there is no build step, no test suite and no sample output committed alongside the notes, the only verification available is reading the file and checking whether its claims match the papers it cites. Treat each note as a secondary source and keep the original paper open next to it.

The one thing in the repository that is not free is the course. The README states the price is 499 and describes a project in which you build an inference framework on Triton plus PyTorch, supporting FlashAttention V1, V2 and V3, GQA, paged attention, fused kernels such as KV linear layer fusion, and Qwen3, Qwen2.5, Llama3 and Llava1.5 models. The README claims a speedup of up to 4x over the transformers library on Llama3 1B and 3B. That number is the README's own claim about the course project, and nothing in the repository lets a reader reproduce it.

Where the notes stop and the gaps begin

The most obvious limitation is language. Every note is in Chinese. For an engineer who does not read Chinese, this repository is unusable regardless of its technical depth, and there is no translation layer.

The second is that notes are not code. The README's course description promises a framework with clean architecture and detailed comments, but the repository's own top-level entries are directories of Markdown and an images folder. A reader looking for a working Triton GEMM kernel to drop into a project will not find one here; the notes explain, they do not execute. The README does point outward to code, listing kernl, unsloth, Liger-Kernel and CUDA-Kernels-Learn-Notes as external resources, which is an honest signal that the runnable implementations live elsewhere.

The third is licence. No licence is stated for this repository. Without one, the default position is that the author retains rights, and readers who want to reuse the text or images in their own material have no granted permission to point to. The README's own reference list credits external sources, but that does not settle the status of the notes themselves.

Finally, the README's resource recommendations are opinionated in a way that is useful but also personal. It recommends the GPU Architecture and Programming PDF as a first document, the CUDA Tutorial site for systematic learning, and learn-cuda for advanced asynchronous material, while saying outright that the Chinese translation of the CUDA C Programming Guide is poor and that it does not recommend it. That kind of judgement is rare and worth having, but it is one practitioner's view, not a survey.

How this differs from vLLM, LightLLM or TGI documentation

The natural comparison is not another notes repository but the projects the notes describe. vLLM, LightLLM and TGI each ship their own documentation, and that documentation is written to get you running: install commands, configuration keys, API endpoints, deployment guides. Their source is also available, and the vLLM codebase is the authoritative description of what vLLM does.

The difference in approach is that this repository reads those projects rather than operating them. The README lists a vLLM inference flow analysis and a vLLM optimization overview, and a LightLLM inference overview. What you get is a guided walk through the code path, in Chinese, with the intent of explaining why the design is what it is. What you do not get is versioning. Framework documentation is tied to a release and changes with it; notes about vLLM's inference flow have no stated version anchor, so a reader cannot tell which vLLM release the walkthrough describes. If you are debugging a specific vLLM deployment, go to vLLM's own documentation and source. If you are trying to understand the paged attention idea well enough to reason about it, the note is the faster entry point.

The same relationship holds for the CUDA and Triton material. NVIDIA's CUDA C++ Programming Guide is the complete reference, and the README says so, calling it an encyclopedia to consult for gaps and detail rather than a first read. The notes here are the opposite artifact: a first read that points at the encyclopedia.

Maintenance, licence and what to check before depending on it

The repository is not archived and the last push was on 2026-08-19. There are no releases, so there is no version number to pin and no changelog to read. Updates arrive as commits to Markdown files, which means the practical way to see what changed is the commit history rather than a release page.

The README states that the course content will continue to be updated and optimized, which is a statement about the course, not a commitment about the notes. Nothing in the repository states an update cadence for the notes themselves.

On licence: none is declared. That matters if you plan to quote the notes, reuse the diagrams, or republish translations. Without an explicit licence you have no stated permission, and the README's practice of citing external references does not change that. If reuse matters to you, the repository provides no mechanism to resolve it, and this is a question for the author rather than something a reader can settle.

The README also carries a Star History chart image and a QR code image for contacting the course author. Both are part of the repository's promotional framing, and neither tells you anything about the technical content. Judge the notes by opening them.

Editorial conclusion

Adopt it as a reading list if you can read Chinese and want a curated path through FlashAttention, quantization and Triton kernel work. Skip it if you need runnable code, an installable package or English documentation, because the repository is a set of Markdown notes with no setup instructions and no stated licence. Before relying on any file, open it and check whether it is complete, since the README links to several documents whose content the README itself does not describe.

Frequently asked questions

Is harleyszhang/llm_note free?

The notes themselves are in a public GitHub repository with no stated price. The README separately describes a paid course on building an inference framework, which it states costs 499, so the free part and the paid part are distinct.

What topics does harleyszhang/llm_note cover?

The README lists transformer model structure, LLM quantization with SmoothQuant and AWQ, inference optimization including FlashAttention and prompt cache, high performance computing with Triton and CUDA notes, and framework analysis of TGI, vLLM and LightLLM.

Do I need to install anything to use harleyszhang/llm_note?

No. The repository contains Markdown notes and an images folder, and the README gives no installation steps, package name or runnable entry point. Cloning the repository and reading the files is the whole workflow.

Is harleyszhang/llm_note available in English?

The notes are written in Chinese, and the README gives no indication of an English version. Readers who do not read Chinese will not be able to use the material.

Official sources

  1. harleyszhang/llm_note on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/harleyszhang-llm-note.svg)](https://hysenlabs.com/projects/harleyszhang-llm-note)