nanoGPT in practice: two 300-line files, one A100, and a deprecation notice
GitHub describes it as The simplest, fastest repository for training/finetuning medium-sized GPTs.. The repository metadata lists Python as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.
At a glance
- What is it?
- karpathy's nanoGPT is a deliberately tiny training repository built from two 300-line Python files. Its own banner now marks it deprecated in favor of nanochat, so it is best read as reference code you study and edit rather than a project under current work.
- Who is it for?
- nanoGPT is worth reading and adapting if you want every line of a GPT training loop in front of you, and it is a fine sandbox for teaching. Do not pick it expecting ongoing work, pinned releases, or a cheap path to GPT-2 scale.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Activity is slowing. The repository last received commits 10 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
nanoGPT's own banner redirects you to a different project
The first thing the top banner says is that nanoGPT has a new and improved cousin called nanochat, and that most people arriving here meant to reach for nanochat instead. It goes on to call nanoGPT very old and deprecated, left up for posterity. That framing matches the repository's own history: the last push was on 2025-11-12. So when the material describes the code reproducing GPT-2 124M on OpenWebText in about four days, read that as a historical snapshot, not a promise of new work. For a reader the practical consequence is to treat nanoGPT as reference code you read, copy, and edit in place, and to look at nanochat if you want the project its author is currently pointing people toward. Nothing in the banner tells you to expect bug fixes, new backends, or releases here.
Two 300-line files carry the whole training loop
nanoGPT is a rewrite of minGPT that prioritizes teeth over education, meaning working code over textbook exposition. The design is deliberately plain: train.py is about a 300-line boilerplate training loop, and model.py is about a 300-line GPT model definition that can optionally load the GPT-2 weights from OpenAI. That smallness is the whole point. Because the pipeline is readable end to end, you can hack it to your needs, train a model from scratch, or finetune a pretrained checkpoint, with the largest available starting point mentioned being the GPT-2 1.3B model from OpenAI. The repository around those two files is small too: configurator.py, sample.py, bench.py, a data directory, a config directory, and notebooks for scaling laws and transformer sizing. There is little abstraction to lean on when your needs outgrow the script, which is exactly the point at which you graduate to something larger.
The Shakespeare run fits a three minute budget on one A100
The quick start trains a character-level GPT on the works of Shakespeare. You turn the raw text into one large stream of integers first:
python data/shakespeare_char/prepare.pyThat creates a train.bin and a val.bin in the data directory. If you have a GPU, train with the provided config:
python train.py config/train_shakespeare_char.pyThat config trains a GPT with a context size of up to 256 characters, 384 feature channels, and a 6-layer Transformer with 6 heads in each layer. On one A100 it takes about 3 minutes and the best validation loss is 1.4697. Checkpoints are written to the out_dir, out-shakespeare-char, and you sample with:
python sample.py --out_dir=out-shakespeare-charThe output is loose Shakespeare pastiche. What this path cannot do is produce a genuinely good model from scratch: the documentation says better results come from finetuning a pretrained GPT-2 model on the same dataset instead. The quick start proves the pipeline runs, not that it converges on anything.
The CPU recipe disables compile and shrinks every dimension
Without a GPU, the walkthrough dials everything down for a MacBook. Here is the full command:
python train.py config/train_shakespeare_char.py --device=cpu --compile=False --eval_iters=20 --log_interval=1 --block_size=64 --batch_size=12 --n_layer=4 --n_head=4 --n_embd=128 --max_iters=2000 --lr_decay_iters=2000 --dropout=0.0Every flag here earns its place. On CPU you set --device=cpu and turn off PyTorch 2.0 compile with --compile=False. --eval_iters=20 gives a faster but noisier estimate, down from 200. The context size drops to 64 characters instead of 256, and the batch size to 12 per iteration instead of 64. The Transformer shrinks to 4 layers, 4 heads, and 128 embedding size, and iterations drop to 2000 with the learning rate decayed to match via --lr_decay_iters. Because the network is small, regularization eases off with --dropout=0.0. What this recipe cannot do is match the GPU result: the documentation says it still runs in about 3 minutes but reaches a loss of only 1.88 and therefore worse samples. Sample it with:
python sample.py --out_dir=out-shakespeare-char --device=cpuFor a reader, the CPU path is a way to feel the pipeline, not a path to a meaningful model.
Apple Silicon needs the mps flag turned on by hand
On Apple Silicon MacBooks, with a recent PyTorch version, you have to add --device=mps, short for Metal Performance Shaders, so that PyTorch uses the on-chip GPU. The documentation says this can significantly accelerate training, on the order of 2-3X, and let you use larger networks, and it points to Issue 28 for more. The catch is that it is not automatic. If you copy the plain CPU command onto an M-series Mac, you leave the on-chip GPU idle and quietly run the slow CPU path instead, which is why the mps note is folded in as a correction near the end rather than set as the default. A reader who wants the fast local route has to remember the flag and has to be on a recent PyTorch for it to pay off. Without both, the cheap-computer suggestion becomes the slowest one without ever raising an error.
Reproducing GPT-2 124M needs an eight A100 node for four days
The serious path is reproducing GPT-2. You first tokenize the dataset, OpenWebText, an open reproduction of OpenAI's private WebText:
python data/openwebtext/prepare.pyThis downloads and tokenizes the dataset and creates a train.bin and val.bin holding the GPT2 BPE token ids in one sequence, stored as raw uint16 bytes. To reproduce GPT-2 124M you need at least an 8X A100 40GB node:
torchrun --standalone --nproc_per_node=8 train.py config/train_gpt2.pyThis runs for about 4 days using PyTorch Distributed Data Parallel and goes down to a loss of about 2.85, against a GPT-2 model evaluated on OpenWebText at roughly 3.11. What nanoGPT cannot do is make that scale cheap: the four day, eight GPU requirement is the floor, not a suggestion. A reader on a single consumer card is not reproducing GPT-2 here, whatever they call the run. The raw uint16 storage also ties token streams to that data width, so a different dataset needs its own prepare script rather than a flag.
One pip line and no releases means you own the version pins
Installation is a single command:
pip install torch numpy transformers datasets tiktoken wandb tqdmEach dependency earns its place. pytorch and numpy are the base runtime. transformers is there to load GPT-2 checkpoints. datasets is only needed if you want to download and preprocess OpenWebText. tiktoken provides OpenAI's fast BPE code. wandb is for optional logging and tqdm is for progress bars. Two consequences follow for a reader. There is no version pin in that command, and the project ships no GitHub releases, so the only fixed reference point is the master branch, whose last push was 2025-11-12. And the top-level entries list no lock file or pinned requirements file to fall back on. If you want to reproduce a run later, you have to record your own versions, because the project does not hand you a tested set to pin against.
Editorial conclusion
nanoGPT is worth reading and adapting if you want every line of a GPT training loop in front of you, and it is a fine sandbox for teaching. Do not pick it expecting ongoing work, pinned releases, or a cheap path to GPT-2 scale. Before you rely on it, check that your PyTorch version matches what the code expects, record your own dependency versions because none are pinned, and look at nanochat if you want the author's current direction.
Frequently asked questions
what is nanogpt
nanoGPT is an MIT-licensed Python repository described as the simplest, fastest repository for training and finetuning medium-sized GPTs. It is a rewrite of minGPT that prioritizes teeth over education, with train.py and model.py each about 300 lines.
how to use nanogpt
Run the Shakespeare data prepare script, train with a config file, then sample from the output directory. The quick start trains a character-level GPT on the works of Shakespeare in about 3 minutes on one A100 GPU.
how to set up nanogpt
The install is one pip line for torch, numpy, transformers, datasets, tiktoken, wandb, and tqdm. Then run the relevant prepare script under data/ to build train.bin and val.bin before launching train.py.
Is NanoGPT free?
This repository is MIT-licensed training code, not a hosted product. Its top banner now marks it very old and deprecated and points to a newer project called nanochat.