train-llm-from-scratch: a pure-PyTorch LLM pipeline from raw text to GRPO
A straightforward method for training your LLM, from downloading data to generating text.
At a glance
- What is it?
- FareedKhan-dev's repository walks a single small Transformer through pretraining, SFT, reward modelling, PPO, DPO and GRPO without trl, peft or transformers. It is a teaching codebase, and its GPU table is the real gate.
- Who is it for?
- Adopt it if you want to read and run the whole pipeline yourself, from tokenizer to GRPO, on one GPU, and you accept that the 13M parameter model is the verified path and the billion parameter claim is a config change you have to validate on your own hardware. Skip it if you need a maintained framework with releases, versioned APIs and support, because the repository has no releases and the last push was on 2026-08-17.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 44 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What train-llm-from-scratch actually is, and who it is written for
Most "train your own LLM" repositories hand you a wrapper around Hugging Face Trainer and call it a course. This one does the opposite. The README states the whole pipeline is hand written in plain PyTorch, with no trl, no peft and no transformers, and the path it lays out runs from raw text through tokens, a Transformer, next-token loss, a base model, then SFT, a reward model, PPO or DPO, GRPO, evaluation and chat. The author frames it as one idea repeated: turn text into numbers, predict the next token, then keep changing the data and the loss until the model does what you want.
The audience is stated in three parts. Students read top to bottom because each code block follows a plain explanation and most blocks are followed by expected output. Developers get commands and file paths to copy. Researchers are pointed at the post-training half, which the README describes as SFT, a Bradley-Terry reward model, PPO with GAE, DPO/ORPO/KTO and GRPO, all on the same small Transformer and trained on public datasets. That split matters, because the repository is two products in one tree: a readable pretraining tutorial and a from-scratch RLHF implementation. The second is the part you cannot easily find elsewhere without pulling in a framework.
The mechanism: one Transformer, five training loops
The architecture is deliberately small and built bottom-up. The README's code structure section walks through an MLP, single-head attention, multi-head attention, a Transformer block, then the full Transformer, and the pretraining step trains that object on next-token loss. That is the base model. Everything after it reuses the same weights and changes only the objective.
SFT comes first: supervised fine-tuning on instruction data. Then a reward model, described as Bradley-Terry, which scores completions. Then the repository offers several preference-optimisation routes on top of that reward signal: PPO with GAE, the DPO/ORPO/KTO family, and GRPO/RLVR. The README's pipeline diagram puts PPO and DPO side by side as alternatives branching off the reward model, with GRPO after them, which is the honest way to draw it. You are not meant to run every algorithm in sequence; you pick a preference objective and compare.
The colour convention in the diagrams is the closest thing to a legend for the data flow: green for raw data, teal for tokenized data stored on disk, blue for a plain processing step, yellow for the model or a training step, orange for reinforcement learning and reward parts, red for loss, grey for a saved checkpoint, purple for final output or evaluation. The tokenized-data-on-disk step is the practical hinge. Pretraining reads tokenized shards rather than raw text, so the data preparation step is a separate, repeatable stage rather than something buried inside the training loop.
Installing train-llm-from-scratch and generating your first tokens
The README's setup section is three commands. The editable install is the important one: it puts config, src, data_loader and ui on your import path, so you no longer need to set PYTHONPATH by hand. Run these from a shell with Python 3.9 or newer.
git clone https://github.com/FareedKhan-dev/train-llm-from-scratch.git
cd train-llm-from-scratch
pip install -e .The base install pulls torch, numpy, h5py, tqdm, tiktoken, zstandard and requests. Two optional extras cover the rest. The train extra adds datasets and wandb for downloading data and logging, and the ui extra adds streamlit, pandas and altair for the control panel.
pip install -e ".[train]"
pip install -e ".[ui]"One caveat the packaging files make explicit: the project install does not pin a CUDA build of torch. A separate requirements.txt adds the cu118 index and torchvision and torchaudio, and the pyproject comment says to install the right torch wheel for your machine first if needed. So on a CUDA box, install torch yourself before the editable install, or you may end up with a CPU wheel and a training script that will not use your GPU.
The README's own starting point is the 13 million parameter model, and it publishes a sample of that model's output so you know what a working small run looks like: text that is grammatical in patches, drifts across topics, and invents place names. Treat that as the expected ceiling for the smallest config, not as a failure.
GPU memory is the real constraint, and the flags that move it
The prerequisites section is blunt: you need a GPU to train. A free Colab or Kaggle T4 is enough for the 13 million parameter model but will not fit a billion parameter model. The README's table gives practical ceilings per card, and the numbers are worth reading as a budget rather than a compatibility list. A 40 GB A100 tops out around 6B to 8B parameters for training. A 24 GB RTX 4090 and RTX 3090 land around 4B and 3.5B to 4B respectively. A 16 GB V100, RTX 4080 or Tesla T4 sits near 2B, 2B and 1.5B to 2B. An 8 GB RTX 4060 is listed at about 1B. The RTX 5090 row is the one to read carefully: it says 13M verified, larger configs TBD. That is the author being straight with you, and it should temper any plan built on a single GPU and the word billion in the README's opening.
For configs that do not fit, the pretraining script has opt-in flags: --amp, --grad-checkpointing and --grad-accum. The README says these bring memory down a lot. They are opt-in rather than defaults, which is the right call for a teaching repo, because it keeps the first run simple and makes the memory trade-off a deliberate step. Note what the repository does not document: there is no stated quality cost for enabling these flags, and the README does not give a table of memory savings per flag. You will have to measure that yourself on your own card.
Where this repository is the wrong tool
Two failure modes are worth naming before you invest a weekend.
The first is scale. The verified path is a 13 million parameter model. The README says you can train a billion or million parameter model on a single GPU, and the GPU table supports a multi-billion ceiling on large cards, but the published output sample is from the 13M model and the RTX 5090 row is marked TBD for larger configs. If your goal is a model that answers questions usefully, this repository is a way to learn the mechanics, not a way to produce the artifact. The tokenizer and data pipeline will happily feed a larger config; whether that config trains to a useful loss on your hardware is an open question the README does not close.
The second is support. There are no releases, and the last push was on 2026-08-17. The API surface is the file layout itself: config, src, data_loader, ui. If you build on top of these modules, an upstream refactor of src/models or src/post_training can break you, and there is no version tag to pin against. A project with a stable interface and a support contract is a better fit when the model is a means to a product rather than the thing you want to understand.
The README also does not document rollback or checkpoint resumption behaviour, so plan your runs around the saved checkpoints the diagram marks in grey rather than assuming you can restart mid-run.
How it differs from the framework route
The obvious alternative is the stack this repository refuses to use: transformers plus trl plus peft, optionally with accelerate for multi-GPU. The difference is not quality, it is where the abstraction sits. In the framework route you configure a trainer, pass a dataset and a model identifier, and the library owns the training loop, the attention implementation, the loss and the distributed strategy. You get multi-GPU, mixed precision, checkpointing and a long list of supported architectures for free, and you spend your time on data and evaluation.
Here you own the loop. Multi-head attention, the Transformer block, the Bradley-Terry reward model and the PPO objective are all in the tree, which means you can print any tensor, change any loss term and see exactly what happens to the output. That is why the README can promise students that every block comes after an explanation of what it does. The cost is that nothing is optimised for you: no fused kernels, no sharded training, no distributed launcher in the described setup, and no upstream bug fixes arriving through a package manager. Choose the framework when you want a model. Choose this repository when you want to know how the model got there.
A middle option exists and the repository supports it: use the framework stack for the base model and come here only for the post-training algorithms, since SFT, DPO and GRPO are the parts most people have never implemented themselves.
Licence, upgrade cost and what the repository commits to
The licence is MIT, declared both in the README badge and in pyproject.toml as license = { text = "MIT" }. For practical purposes that means you can read, modify and redistribute the code, including in commercial work, provided you keep the licence notice. The repository ships no model weights and no dataset, so the licence covers the code and the documentation, not any artifact you train with it. Whether your training data or your resulting model carries its own obligations is a separate question this repository cannot answer for you, and it is not a legal opinion to rely on.
The upgrade cost is low in dependency terms and high in interface terms. Dependencies are unpinned apart from the Python floor of 3.9, so a fresh pip install -e . will pull current torch, tiktoken and h5py, and a breaking change in any of them lands on you. The package list in pyproject.toml is explicit, covering config, data_loader, src, src.models, src.post_training, src.post_training.rewards and ui, and the JSON configs ship as package data. That is a small, legible surface. What you cannot do is pin the project itself: with no releases, your only version reference is a commit hash. Record the hash you installed from if you plan to compare runs across weeks. Optional extras are additive, so a missing wandb or streamlit install degrades the UI and logging rather than the training path.
Editorial conclusion
Adopt it if you want to read and run the whole pipeline yourself, from tokenizer to GRPO, on one GPU, and you accept that the 13M parameter model is the verified path and the billion parameter claim is a config change you have to validate on your own hardware. Skip it if you need a maintained framework with releases, versioned APIs and support, because the repository has no releases and the last push was on 2026-08-17. Before you commit, run pip install -e . and then python src/train.py --help on your machine to confirm the flags the README lists actually exist in your checkout.
Frequently asked questions
Can I train my own AI LLM with train-llm-from-scratch?
Yes, within the limits the README sets. You need a GPU, and the verified path is a 13 million parameter model that fits on a free Colab or Kaggle T4. Larger configs are supported by the scripts but the README marks larger configs on some cards as TBD.
How can I build an LLM from scratch using train-llm-from-scratch?
Clone the repository, run pip install -e . so config, src, data_loader and ui are on your import path, then follow the README's steps in order: prepare the data, build the Transformer, pretrain the base model, generate text, then run the post-training steps. The README states every algorithm is hand written in plain PyTorch with no trl, peft or transformers.
Can I train an LLM locally with train-llm-from-scratch?
The README says you need a GPU to train and gives practical parameter ceilings per card, from about 1B on an 8 GB RTX 4060 up to 6B to 8B on a 40 GB A100. The 13 million parameter model is the configuration the repository verifies.
How do you train an AI LLM end to end in train-llm-from-scratch?
The README lays out the path as raw text, tokens, a Transformer, next-token loss, a base model, then SFT, a reward model, PPO or DPO, GRPO, and finally evaluation and chat. The same small Transformer is reused across the post-training algorithms.
How to train an LLM from scratch with train-llm-from-scratch?
The README's setup is git clone followed by pip install -e ., with optional extras for datasets and wandb or for the Streamlit control panel. From there the steps run from data preparation through pretraining, text generation and the post-training algorithms.
Is there a how to train an LLM from scratch book for train-llm-from-scratch?
The repository is not a book; it is a README plus a documentation site at fareedkhan-dev.github.io/train-llm-from-scratch and a sft_rlhf_guide.ipynb notebook in the repository root. The README says students can read it top to bottom, with each code block preceded by an explanation.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/fareedkhan-dev-train-llm-from-scratch)