# Modern LLM Notebook: building tokenizers, attention and RLHF in 26 PyTorch notebooks

> A bilingual, notebook-first course from WalkingLabs that rebuilds LLM components in PyTorch instead of wrapping a library. Useful for engineers who want the mechanism, not a production framework.

**walkinglabs/modern-llm-notebook** — A hands-on course for building modern LLMs from scratch in PyTorch, with 26 runnable Jupyter Notebooks covering tokenizers, attention, MoE, RLHF, inference, evaluation, and distillation.

- Repository: https://github.com/walkinglabs/modern-llm-notebook
- Website: https://walkinglabs.github.io/modern-llm-notebook/
- Stars: 210 · Forks: 41
- Language: Jupyter Notebook
- License: NOASSERTION
- Published: 2026-08-27 · Updated: 2026-08-27 · Language: en
- Canonical page: https://hysenlabs.com/projects/walkinglabs-modern-llm-notebook

## The gap Modern LLM Notebook targets: engineers who call models but cannot rebuild them

Most people entering LLM work start at the API layer. They call a tokenizer, load a checkpoint, run a generate loop, and never see the tensor shapes in between. The README states the project's aim plainly: build the core components yourself, from Tokenizer and Transformer to training, inference, alignment and production. The audience is named in the same document: software engineers who know Python and want to move into LLM engineering, ML practitioners who use model libraries but want to know what happens underneath, students preparing to read papers, and self-learners who prefer runnable examples over derivations. The stated background is light. Basic Python, arrays, functions, classes and simple matrix operations. Calculus, probability and PyTorch are described as helpful but not required on day one, and no prior knowledge of Tokenizer, Embedding, Self-Attention or Transformer internals is assumed. That is a deliberate positioning choice. This is not a course for someone who already tunes distributed training runs. It is for the engineer who can read a forward pass in a paper and wants to have written one.

## How the notebooks are built: intuition, hand calculation, implementation, experiment

Every notebook in the course follows one loop, quoted in the README as `intuition -> hand calculation -> implementation -> experiment`. The ordering is the mechanism. A topic opens with the problem it solves, then works a small numeric example by hand, then turns that arithmetic into PyTorch, then runs a controlled comparison and prints or plots the result. Six design principles back this up: motivation before mechanics, intuition before notation, hand calculation before abstraction, readable implementations over black boxes, experiments explain behavior, and one concept at a time. The practical consequence is that components stay explicit. You are meant to see the attention scores, the mask, the softmax and the weighted sum as separate steps rather than as one library call. The stated goal is not to reproduce a production framework line by line but to build a durable mental model of what each component does, why it exists and how numbers flow through it. The repository layout matches the claim: a `notebooks/` directory holds the Chinese source edition, `notebooks-en/` the English mirror, and the README notes that the Chinese course is the source edition while the English one is updated alongside it. There is also an `llm_train/` directory, a `karpathy_models.py` file at the root, an `external/` directory, and a `web/` React interface driven by the root `package.json` scripts.

## Installing Modern LLM Notebook and running your first tokenizer notebook

The repository ships a `requirements.txt` at the root, so the local path is a virtual environment plus pip. The pins are explicit and worth reading before you install, because they set the floor for what the notebooks can call.

```bash
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
```

After that install, the declared dependencies include `torch>=2.0.0`, `transformers>=4.30.0`, `datasets>=2.12.0`, `accelerate>=0.20.0`, `tiktoken>=0.7.0`, `modelscope>=1.9.0`, `jupyter>=1.0.0` and `ipykernel>=6.0.0`. Note that `tiktoken` appears twice in the file, at `>=0.7.0` and `>=0.5.0`. pip resolves that without complaint, but it is a sign the file has been edited over time rather than regenerated.

The README's fastest route skips local setup entirely. The header links a Colab badge that opens the first notebook directly:

```text
https://colab.research.google.com/github/walkinglabs/modern-llm-notebook/blob/main/notebooks-en/part1-foundation/01-tokenizer-basics.ipynb
```

That path tells you the naming convention: `notebooks-en/part1-foundation/01-tokenizer-basics.ipynb`. Part 1 is the foundation track, notebook 01 is the tokenizer. If you prefer to run locally, start Jupyter from the repository root and open the same file, and you should see the character, word and BPE tokenizer implementations the README lists under what you will build. The project also publishes a rendered version at `https://walkinglabs.github.io/modern-llm-notebook/`, which is the right place to skim before committing to a local install.

For the web interface, the root `package.json` defines `dev`, `build` and `preview` scripts that delegate to a `web/` subproject, with a `postinstall` hook that installs the web dependencies. That is a separate concern from the notebooks themselves.

## What the 26 notebooks actually cover, from BPE to speculative decoding

The curriculum is staged. Text to tokens covers character, word and BPE tokenizers. Tokens to vectors covers token embeddings and position encodings. The Transformer core rebuilds self-attention, multi-head attention, Transformer blocks and a Mini-GPT. A training system stage covers cross-entropy, batching, gradient flow and scaling-law experiments. Later stages move into MoE, LoRA, RLHF, decoding, KV Cache, long-context techniques, VLM components, evaluation and distillation. The README's learning outcomes list is specific enough to check against: trace data flow from raw text to tokens, hidden states, logits and generated text; implement a compact GPT-style model; connect cross-entropy, gradients, batching and data quality to training behavior; explain RoPE, RMSNorm, SwiGLU, GQA, MLA and MoE; compare LoRA, reward modeling, PPO and DPO; reason about latency, memory, KV Cache and speculative decoding. The August 2026 note says Part 3, notebooks 20 through 26, was fully rebuilt in the Part 1 house style, with three self-checking homework problems per notebook. The quantization notebook is described as covering FP8 and FP4 formats with a grid experiment, GGUF and K-quant details, and an end-to-end walkthrough producing GPTQ/FP8 via llm-compressor, AWQ via AutoAWQ, and GGUF via llama.cpp with imatrix, then serving each. Notebook 23 is described as a runnable speculative-sampling loop with measured acceptance and speedup. Notebook 24 covers batching, paging and prefix-caching simulators plus vLLM and SGLang deployment workflows. Notebook 25 walks an eval pipeline with example items from MMLU, C-Eval, CMMLU, GSM8K and HumanEval, a tooling map across lm-evaluation-harness, OpenCompass and EvalScope, confidence intervals, and a lab that registers a custom Chinese benchmark into lm-eval via YAML and scores GPT-2 against Qwen2.5-0.5B. That is a wide surface for a course, and it is the part most likely to age as tooling moves.

## Where the notebook format costs you: drift, hardware and the missing rollback story

Notebooks rot faster than libraries. The README states the Chinese course is the source edition and the English mirror is being updated alongside it, which means the two can diverge and the English reader may hit a cell that lags the Chinese original. There is a file named `NOTEBOOK_RUN_REPORT_20260630.md` at the repository root, which suggests the maintainers do track execution status, but the README does not explain what that report covers or how often it is regenerated. Treat it as a signal, not a guarantee.

The dependency floor is another constraint. `torch>=2.0.0` and `transformers>=4.30.0` are lower bounds, not pins, so a fresh install today can pull versions newer than the ones the notebooks were written against. Notebooks that print tensor shapes or loss values are sensitive to that. If a cell produces a shape you did not expect, check the installed version before assuming the notebook is wrong.

Hardware is the third limit. The course is designed to run small, and the README's framing supports that, but the Part 3 notebooks touch quantization pipelines, vLLM and SGLang deployment, and a benchmark lab that scores Qwen2.5-0.5B. Those are heavier than a tokenizer cell. The README does not publish per-notebook hardware requirements, and it does not document rollback or version-pinning guidance for the Part 3 tooling. If you need a fixed, reproducible environment for a team, this repository does not describe one.

Finally, the scope boundary matters. The README says the goal is not to reproduce a production framework line by line. If you want a training framework, a serving stack or tuned production hyperparameters, this is the wrong tool. It teaches the mechanism, and the mechanism is deliberately smaller than what you would deploy.

## Modern LLM Notebook versus nanoGPT and similar minimal implementations

The closest comparison is Andrej Karpathy's nanoGPT, and the repository itself gestures at that lineage: there is a `karpathy_models.py` file at the root and an `external/` directory. The difference is in what each optimizes for. A minimal GPT implementation like nanoGPT is a single training script you read top to bottom and then modify. It is compact by design, and its value is that you can hold the whole thing in your head. Modern LLM Notebook spreads the same material across 26 notebooks and adds stages that a single training script does not cover: tokenizers as a standalone topic, MoE, LoRA, RLHF with reward modeling and PPO and DPO, KV Cache, long context, VLMs, quantization, speculative decoding, evaluation harnesses and distillation. The trade-off runs both ways. You get breadth and a staged learning loop, and you lose the single-file coherence that makes a minimal implementation easy to reason about as a whole. A reader who wants to understand one GPT forward pass and stop will find nanoGPT faster. A reader who wants to reach speculative decoding and lm-eval benchmark registration without assembling a reading list will find the notebook sequence more direct. The course also ships a bilingual edition, which a single-author script does not attempt.

## Conclusion

Use it if you write Python and want to trace text through tokens, hidden states, logits and sampling without a framework hiding the steps. Do not adopt it as a training framework, a serving stack or a source of production hyperparameters; the README frames the goal as a mental model, not a line-by-line reproduction of production code. Before starting, open the requirements.txt pins and one Part 3 notebook, since the README states the English mirror is being updated alongside the Chinese source edition and a file named NOTEBOOK_RUN_REPORT_20260630.md sits at the repository root.

## FAQ

### What is Modern LLM Notebook?

It is an open, hands-on course for rebuilding the essential machinery of large language models in PyTorch, delivered as 26 runnable Jupyter Notebooks. The README describes it as notebook-first and from-scratch, covering tokenizers, attention, training, MoE, LoRA, RLHF, inference, evaluation and distillation.

### Does Modern LLM Notebook need a GPU to run?

The README does not publish per-notebook hardware requirements, so this cannot be answered from the documentation. The course is built around small, independently runnable notebooks, and the README links a Colab badge for the first one, which suggests the foundation track runs without local hardware.

### How do I install Modern LLM Notebook?

The repository ships a requirements.txt at the root, so the local path is creating a virtual environment and running pip install -r requirements.txt. The README also links a Colab badge that opens notebooks-en/part1-foundation/01-tokenizer-basics.ipynb directly, which skips local setup.

### Is Modern LLM Notebook available in Chinese?

Yes. The README states that the Chinese course is the source edition and the English mirror is being updated alongside it, with separate README.md and README-CN.md files and separate notebooks/ and notebooks-en/ directories in the repository.

### What background do I need before starting Modern LLM Notebook?

The README asks for comfortable basic Python and familiarity with arrays, functions, classes and simple matrix operations. Basic calculus, probability and PyTorch are described as helpful but not required on day one, and no prior knowledge of Tokenizer, Embedding, Self-Attention or Transformer internals is assumed.

## Sources

- [Official documentation](https://walkinglabs.github.io/modern-llm-notebook/)
- [Official README](https://github.com/walkinglabs/modern-llm-notebook#readme)
- [Project repository](https://github.com/walkinglabs/modern-llm-notebook)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/walkinglabs-modern-llm-notebook
