Model or dataset
raiyanyahya/how-to-train-your-gpt avatar
raiyanyahya/how-to-train-your-gpt

How to Train Your GPT: A 12-Chapter Notebook Course That Builds a LLaMA-Style Model Line by Line

Build a modern LLM from scratch. Every line commented. Explained like we are five.

3,369 stars419 forksJupyter NotebookMIT

At a glance

What is it?
raiyanyahya/how-to-train-your-gpt is a Jupyter Notebook textbook that walks a Python developer through a BPE tokenizer, RoPE, multi-head attention, a 151M parameter decoder and a training loop. It is a learning resource, not a framework, and the README says so.
Who is it for?
Adopt this if you already write Python and want to understand attention, RoPE and RMSNorm by typing them yourself rather than calling an API. Do not adopt it if you need a maintained training framework, a distributed or multi-GPU recipe, or a model you intend to ship: the README labels the project learning only, and the repository ships no releases.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 31 days ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap this course fills between calling an API and reading a paper

Two kinds of material dominate LLM learning. One is the three-line fine-tuning script where the model is a black box. The other is the original attention paper and its descendants, full of notation that assumes a graduate degree. This repository positions itself between them. The README's own comparison table puts it as "5-year-old analogies → full working code" against "model = GPT().fit(data)" on one side and dense papers on the other.

The stated audience is a Python developer who knows functions, classes and pip install, and has no ML background. The README says no calculus, no linear algebra and no PyTorch experience are required, and that those are taught as the chapters go. That is a specific claim, and it sets the difficulty of the first four chapters: tokenization, embeddings and positional encoding are explained before any matrix multiplication matters.

The author is explicit about motive. The README says the project was made with the goal of learning something not completely understood, specifically the attention part, and that AI was used to understand and verify key concepts. That framing matters when you decide how to read it: this is one person's study notes cleaned up into a course, not a specification written by the architecture's authors.

What the twelve chapters actually build, component by component

The course is sequential and the README asks you to start at chapter 0 and read in order. Chapters 0 through 4 cover the overview, environment setup, BPE tokenization, embeddings and RoPE. Chapter 5 is flagged as the core: Q, K, V, the scaling factor, the causal mask and an eight-step walkthrough. Chapter 6 assembles the transformer block from RMSNorm, SwiGLU and residual connections, and discusses pre-norm versus post-norm. Chapter 7 produces the complete model at 151M parameters with weight tying. Chapters 8 and 9 cover training and inference. Chapter 10 is a single runnable main.py and chapter 11 is a glossary with an architecture provenance table.

The README gives a line budget per component, which is the most useful sizing signal in the repository. The BPE tokenizer is about 60 lines, embeddings about 30, RoPE about 70, multi-head attention about 120, the transformer block about 50, the full model about 200, the training pipeline about 250 and the inference engine about 80. The README totals roughly 860 lines of core model code against about 2,600 lines of explanation and diagrams. If you are deciding whether to commit, that ratio is the honest answer: expect to read three lines of prose for every line of code.

Alongside the chapters the README lists 28 standalone topic explainers covering RoPE, attention, RMSNorm, SwiGLU, KV cache, AdamW and mixed precision, plus two narrative walkthroughs that trace one sentence through the whole model. The repository tree confirms the layout: chapters/, notebooks/, fine-tuning/, main.py and requirements.txt at the top level, with a directory named "explanations and examples WIP".

Installing it and running the full script for the first time

There is no package to install. The README's quick start clones the repository, and the dependency list lives in requirements.txt, which pins torch>=2.0 and adds tiktoken, datasets, numpy, matplotlib and jupyter. Note the first line: it adds the PyTorch CPU wheel index, so a plain install pulls the CPU build unless you change that line.

bash
git clone https://github.com/raiyanyahya/how-to-train-your-gpt.git
cd how-to-train-your-gpt
pip install -r requirements.txt

After that, chapter 1 covers virtual environments, GPU versus CPU and PyTorch basics, so the README does not treat installation as a single command. The repository also carries a Colab badge, and the README's quick start links to notebooks/colab_train.ipynb, which is the path of least resistance if you do not have a local GPU.

Chapter 10 is described as a runnable main.py containing everything in one file, and main.py sits at the repository root. That is the file to execute once you have read far enough to follow it.

bash
python main.py

The README does not document expected output, runtime, checkpoint paths or command line flags for main.py, so treat the first run as an experiment on your own hardware rather than a reproducible benchmark. If you want the guided version instead, open the notebooks in order and run the cells as you read.

Where the course stops: scale, data and reproducibility

The model is 151M parameters. That is a teaching size, and the README never pretends otherwise. It is roughly three orders of magnitude below the models whose architecture it copies, and the techniques that only matter at scale (tensor parallelism, pipeline parallelism, sharded optimizers, activation checkpointing across many devices) are not part of the chapter list. Mixed precision and gradient accumulation are covered in chapter 8, which is the extent of the efficiency material.

The data story is equally thin. The requirements include the datasets library, but the README does not name a corpus, a preprocessing step or a token budget, and the repository tree shows no data directory. You will supply your own text and decide your own tokenizer vocabulary, which the README's BPE chapter teaches but does not script for you.

Reproducibility is the third gap. The repository has no releases, so there is no version to pin. If you are working through it over weeks, a change to main.py or a chapter lands directly on master and you have no changelog to diff against. That is normal for a personal teaching repository, and it is a real cost if you plan to cite it or reuse its code in something graded.

Finally, the README labels the project purpose as learning only. Nothing in the repository suggests a path to a production checkpoint, and the fine-tuning/ directory is not described in the README at all, so its contents and intent are undocumented.

How this differs from nanoGPT and from a fine-tuning workflow

The obvious comparison is Andrej Karpathy's nanoGPT, which is the reference implementation most people reach for when they want a small GPT they can read in an afternoon. The difference is in what is annotated and which architecture you end up with. nanoGPT is a compact, working GPT-2 style training script where the code is the explanation. This project inverts that: the README claims 100% of the code is commented, and the line budget shows explanation outweighing code by roughly three to one. It also targets a newer architecture than GPT-2, using RoPE, RMSNorm, SwiGLU and pre-norm, which the README attributes to LLaMA, Mistral, Qwen, Gemma and PaLM.

If you already know PyTorch and want a training script you can modify, nanoGPT is the shorter path. If you do not know what a residual stream is, the chapter sequence here is the gentler one, because it spends chapters 2 through 4 on tokenization, embeddings and positional encoding before attention appears.

The other comparison is fine-tuning. Fine-tuning takes a pretrained checkpoint and adjusts it on your data. This course starts from random weights and teaches the pretraining objective, cross-entropy over next-token prediction, with AdamW and cosine warmup. Both are useful and they answer different questions. If your goal is a model that behaves well on your domain next week, fine-tuning is the right tool and this repository is not. If your goal is to stop treating attention as magic, the order is reversed.

Licence, maintenance and what an upgrade costs you

The repository is MIT licensed, with the LICENSE file at the top level. For a course you are reading and adapting, that is permissive: you can copy the tokenizer or the attention block into your own project and keep your own licence, provided you preserve the copyright notice and the licence text. That is the general shape of MIT, not advice about your situation. If you are pasting the model code into a commercial product, have someone qualified read the LICENSE file rather than this paragraph.

Maintenance is the part to weigh. The last push was on 2026-08-30, and the repository is not archived. There are no releases, so there is no version boundary between what you read today and what a future reader will see. The dependency set is small (torch, tiktoken, datasets, numpy, matplotlib, jupyter) and the torch pin is a floor rather than a ceiling, which means a future PyTorch release can change behaviour under a reader without any commit in this repository.

Upgrading is therefore not a task with a procedure. There is nothing to upgrade. What you re-verify is your environment: after any torch bump, re-run main.py and confirm the training loop still steps and the loss still falls. The README does not document a rollback path for dependency changes, and there is no changelog to consult.

Editorial conclusion

Adopt this if you already write Python and want to understand attention, RoPE and RMSNorm by typing them yourself rather than calling an API. Do not adopt it if you need a maintained training framework, a distributed or multi-GPU recipe, or a model you intend to ship: the README labels the project learning only, and the repository ships no releases. Before you start, open chapter 1 and confirm your PyTorch install matches the requirements.txt pin, then run main.py end to end on CPU to see how long one epoch actually takes on your machine.

Frequently asked questions

How can I train my own GPT with how-to-train-your-gpt?

Clone the repository, install the dependencies from requirements.txt, and work through the chapters in order from chapter 0. Chapter 10 provides a runnable main.py that contains the whole pipeline in one file, and the README also links a Colab notebook for running it without a local GPU.

Can I develop my own GPT using this course?

The course builds a 151M parameter decoder-only model from scratch, including the tokenizer, attention, transformer block, training loop and inference engine, so you do develop one. The README labels the purpose as learning only, which means the result is a teaching model rather than something intended for production.

Can I train my own AI with how-to-train-your-gpt?

You train a small language model, not a general AI system. The README describes a 12-chapter, 7,500+ line course that ends with a runnable main.py, and it states that only basic Python is required as a prerequisite.

Official sources

  1. Issues
  2. License: MIT
  3. raiyanyahya/how-to-train-your-gpt on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/raiyanyahya-how-to-train-your-gpt.svg)](https://hysenlabs.com/projects/raiyanyahya-how-to-train-your-gpt)