LLM-RL-Visualized: 100+ Original Diagrams for LLM and RL Algorithms
🌟100+ 原创 LLM / RL 原理图📚,《大模型算法》作者巨献!💥(100+ LLM/RL Algorithm Maps )
At a glance
- What is it?
- A diagram-first reference by the author of 《大模型算法》, covering LLM, VLM, SFT, DPO, RLHF and policy optimization. It is a reading aid, not a runnable library, and the licence file does not name a standard licence.
- Who is it for?
- Adopt it if you already know the vocabulary and want a browsable map of how PPO, GRPO, DPO, RLHF and the rest connect, or if you teach these topics and need SVG figures you can zoom and select text in. Do not adopt it expecting runnable code: the repository is images and Markdown, with no package to install and no release history.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 20 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Why a diagram repository exists at all
Papers and course notes on reinforcement learning for language models assume you can hold several moving parts in your head at once. In the RLHF section alone, the README lists diagrams for the two-stage RLHF pipeline, the reward model structure, the four cooperating models inside PPO, their shared backbone, the KL distance computation, and the PPO-based RLHF schematic. Each one is a small system. Read as text, they blur together.
The repository is an attempt to fix that by making the picture the primary artifact. The README describes it as 100+ original architecture diagrams covering LLM and VLM principles, training algorithms (RL, RLHF, GRPO, DPO, SFT, CoT distillation), and optimization techniques such as RAG. It is aimed at people who already read papers and want a second, visual pass over the same material. The author's book, 《大模型算法:强化学习、微调与对齐》, is the companion text, and the README points to it for more detailed interpretation of the diagrams.
That framing matters for adoption. This is closer to a set of lecture slides than to software. If you are looking for a library you import, you are in the wrong repository.
What is actually in the repository tree
The top level holds README.md, LICENSE, two index files (AI-Roadmap(AI知识架构).md and LLM-VLM-index (汇总).md), a src directory, two image directories, and two PDFs: 强化学习算法图谱 (rl-algo-map).pdf and 策略梯度(Policy Gradient)-强化学习(PPO&GRPO等)之根基.pdf. The README also links to src/README_EN.md for an English version, so the diagram set is mirrored rather than translated in place.
The README's table of contents is the real index, and it runs to eleven parts and roughly 125 anchors. Part 2 covers LLM structure, decoding, input and output layers, VLM and MLLM structure, and the training flow. Part 3 is SFT: LoRA in two diagrams, Prefix-Tuning, token ID mapping, cross-entropy loss, instruction data sources, and packing. Part 4 is DPO, including a comparison of RLHF and DPO training architectures and the effect of the beta parameter. Parts 6 through 8 are the reinforcement learning core: MDP versus Markov chains, Q and V relationships, Monte Carlo versus TD, DQN, policy gradient, GAE, TRPO, PPO-Clip, PPO pseudocode, PPO versus GRPO, then RLHF and RLAIF with reward models, rejection sampling, constitutional AI and OpenAI RBR.
The breadth is the selling point and also the risk. A single author maintaining 125 diagrams across eleven topics will inevitably have uneven depth, and the README does not flag which diagrams are mature and which are sketches.
Getting the diagrams: no install, just clone or download
There is no package, no build step and no CLI. The README does not document an installation procedure because there is nothing to install. You get the figures by cloning the repository or downloading individual files from it. The README states that clicking an image opens a high-resolution version, and that the .svg files in the repository directory are vector graphics you can scale without limit and whose text you can select.
A clone is the most direct route if you want the whole set locally:
https://github.com/changyeyu/LLM-RL-Visualized.git
cd LLM-RL-Visualized
ls images_chinese images_englishAfter that you should see the two mirrored image directories and the PDFs at the top level. If you only want one figure, open it in the GitHub file view and use the download button; the README explicitly says the .svg files are the high-quality option.
For a first real use, treat the table of contents as a syllabus. Suppose you are trying to understand why GRPO differs from PPO. Open README.md, jump to the anchor named 第7部分, and find the entry PPO 与 GRPO. Read the two diagrams side by side, then go to 第8部分 for the four-model RLHF diagrams that show what PPO is actually updating. The README gives no recommended reading order, so the sequence is yours to choose.
Where the repository stops being useful
The most concrete limitation is that nothing here executes. There is no training code, no evaluation harness, no configuration files, and no releases. The README describes the repository as 长期勘误、追加, meaning it is corrected and extended over time, but a diagram set cannot validate a claim the way a runnable example can. If you want to see GRPO converge on a toy task, this repository will not show you that.
The second limitation is the licence. The repository reports NOASSERTION, and the README does not state terms for reuse. That is a real constraint for anyone who wants to put one of these figures into slides, a course, or a product document. NOASSERTION is what GitHub reports when it cannot match the LICENSE file to a known licence, so the file may say something specific, but you have to open it and read it yourself. I cannot tell you what it permits.
The third is language. The primary README is in Chinese, with an English version at src/README_EN.md. The diagram labels in images_chinese and images_english are presumably split along the same line, but the README does not state that every diagram exists in both languages. If you need English labels, check the specific figure before you plan around it.
Finally, scope. The title says LLM and RL, and the content leans heavily toward reinforcement learning for language models. If your interest is classical RL for control, the DQN and DDPG entries are there, but they are a minority of the set.
How it compares with a runnable RL library
The obvious alternative is a framework you can actually run, such as Stable-Baselines3 or a post-training stack like TRL. The difference is not quality, it is category. Stable-Baselines3 gives you implementations of PPO, DDPG and related algorithms with training loops, environments and logs; you learn by changing a hyperparameter and watching the curve move. LLM-RL-Visualized gives you the conceptual layer above that: what the advantage function is doing, why the clip range exists, how the four models in RLHF relate.
A second alternative is a written course or textbook, and here the comparison is closer. A textbook can explain a derivation step by step in a way a single diagram cannot. The repository's answer is the companion book, which the README links for deeper interpretation. So the honest description is that this repository is a supplement to a text, not a replacement for one.
If you are choosing between them, pick based on what you are missing. If you can derive PPO on paper but cannot picture where the reward model sits relative to the policy, the diagrams address your gap. If you cannot derive PPO on paper, a diagram will give you vocabulary without understanding.
Maintenance, upgrade cost and licence questions
The repository is not archived, and the last push was on 2026-08-29. There are no releases, so there is no version to pin and no changelog to read. Upgrading means pulling the latest commit and re-reading whatever changed, which for a diagram set is a visual diff rather than a dependency bump. In practice the cost is low: clone once, and pull when you want the additions.
The cost that is not low is citation. The README includes a 参考文献 section and a BibTeX section, so the project does expect academic use. If you cite it, use the BibTeX the README provides rather than reconstructing one from the repository metadata.
On licence, the only thing I can state is what the repository reports: NOASSERTION. That is not a licence name and not a permission. Whether you can redistribute the SVGs, embed them in paid course material, or modify them is a question the LICENSE file answers, not this article. Check it before you build anything on top of the figures, and if the file is ambiguous, treat that as a reason to ask the author directly through the repository.
Editorial conclusion
Adopt it if you already know the vocabulary and want a browsable map of how PPO, GRPO, DPO, RLHF and the rest connect, or if you teach these topics and need SVG figures you can zoom and select text in. Do not adopt it expecting runnable code: the repository is images and Markdown, with no package to install and no release history. Before you rely on it, open the LICENSE file, because the repository reports NOASSERTION rather than a named licence, and check the two PDF maps and the images_chinese and images_english directories to confirm the figures you need are present.
Frequently asked questions
What is RL for LLM?
The repository treats reinforcement learning for language models as a distinct topic with its own section, covering how a language model is modeled as an RL problem, the two-stage RLHF training flow, reward model structure and the PPO-based RLHF schematic. It also includes the policy optimization algorithms those pipelines depend on, such as GAE, TRPO, PPO-Clip and GRPO.
Can you give me a simple example of reinforcement learning?
The repository does not present a single worked example; it presents diagrams. The closest entries are the DQN application example and the DQN model diagrams in the reinforcement learning basics section, alongside the running-trajectory and return-calculation diagrams that show how an episode is structured.
Is reinforcement learning part of AI?
The README places reinforcement learning inside its coverage of machine learning, with a diagram comparing the three major machine learning paradigms and another tracing the development history of reinforcement learning. It does not make a broader claim about how the field should be classified.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/changyeyu-llm-rl-visualized)