# Modded-NanoGPT: the NanoGPT speedrun that trains a 124M model in under 75 seconds

> Modded-NanoGPT is a competitive benchmark repo that searches for the fastest way to train a 124M-parameter language model to 3.28 FineWeb validation loss on 8 H100 GPUs. The current record is under 75 seconds, and every technique that got it there is documented in the training script.

**KellerJordan/modded-nanogpt** — NanoGPT (124M) in 90 seconds

- Repository: https://github.com/KellerJordan/modded-nanogpt
- Stars: 5,883 · Forks: 897
- Language: Python
- License: MIT
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/kellerjordan-modded-nanogpt

## What the NanoGPT speedrun actually measures

The repository hosts the NanoGPT speedrun, a collaborative and competitive search for the fastest algorithm that trains a language model to 3.28 cross-entropy loss on the FineWeb validation set using 8 NVIDIA H100 GPUs. The target comes from Andrej Karpathy's GPT-2 replication in llm.c, which reached that loss after 45 minutes. The speedrun code descends from llm.c's PyTorch trainer, which itself descends from NanoGPT, which explains the repository name.

The audience is narrow and specific. This is for researchers and engineers who want to test training optimizations under a fixed, reproducible target. It is not a library you import into a product. The README states that the current algorithm attains the target in under 75 seconds on 8xH100 and under 400M tokens, compared with 45 minutes and 10B tokens for the llm.c GPT-2 replication. There is also an optimization track in records/track_3_optimization that minimizes steps subject to fixed architecture, data and batch size with an unlimited wallclock budget. The main track and the optimization track measure different things, and mixing them up is a common source of confusion.

## The techniques stacked into train_gpt.py

The improvement comes from a long list of changes the README enumerates. On the architecture side: rotary embeddings, QK-Norm, ReLU², zero-initialized projections in a muP-like scheme, skip connections from the embedding to every block and from block 3 to 6, extra embeddings mixed into attention values, a Smear module for one-token lookback, bigram hash embedding on a quarter of model_dim with a sign trick, MUDD skip connections, learnable XSA, paired head attention, and multi-token prediction. Training-side changes include the Muon optimizer, FP8 for the head and the MLP forward pass, asymmetric rescale and softcap logits, gradient accumulation for two steps on the embedding and lm_head, a batch size schedule, a max sequence length schedule, cautious weight decay tied to the learning rate, untied embedding and lm_head at two thirds of training, and a prefix token prediction auxiliary loss. Flash Attention 3 with a long-short sliding window pattern, window size warmup with YaRN, and aligning batch starts with EoS round out the list.

That is a lot of interacting changes, and the repository does not present them as independent. The README lists them as the accumulated result of many contributors, and the world record history section tracks the progression. If you want to isolate the effect of one technique, this repository is not organized for that. It is organized to find the fastest combination.

## Installing Modded-NanoGPT and running the current record

The README gives a four-command path to run the current record. The clone and install steps are standard. The data script downloads only the first 900M training tokens to save time, and run.sh launches training.

```bash
git clone https://github.com/KellerJordan/modded-nanogpt.git && cd modded-nanogpt
pip install -r requirements.txt
python data/cached_fineweb10B.py 9
./run.sh
```

If ./run.sh fails with torchrun: command not found, the README says to add torchrun to your path. The README also warns that torch.compile will add around 7 minutes of latency the first time you run the code, so a cold start is not representative of the steady-state speed.

For systems where CUDA or NCCL versions are not compatible, the README recommends Docker as an alternative that standardizes CUDA, NCCL, CUDNN and Python versions. An NVIDIA driver must already be installed on the host. The Dockerfile builds from nvidia/cuda:12.6.2-cudnn-devel-ubuntu24.04 with Python 3.12.7 and installs the requirements plus a nightly PyTorch build from the cu126 index.

```bash
git clone https://github.com/KellerJordan/modded-nanogpt.git && cd modded-nanogpt
sudo docker build -t modded-nanogpt .
sudo docker run -it --rm --gpus all -v $(pwd):/modded-nanogpt modded-nanogpt python data/cached_fineweb10B.py 8
sudo docker run -it --rm --gpus all -v $(pwd):/modded-nanogpt modded-nanogpt sh run.sh
```

The Docker path downloads 8 units of data rather than 9, which is a detail worth noticing if you compare timings between the two routes. For an interactive container, the README gives a bash variant of the same run command. Official records are timed on 8 NVIDIA H100 GPUs from PrimeIntellect, which has sponsored recent validation runs.

## Where Modded-NanoGPT breaks down

The first constraint is hardware. The target and the record are defined on 8 NVIDIA H100 GPUs. Running on fewer GPUs, on A100s, or on consumer cards changes the comparison entirely, and the README does not document expected results for those configurations. If you do not have access to that hardware, you cannot reproduce the record, and the repository offers no fallback path.

The second constraint is that this is a benchmark, not a training framework. There is no configuration system, no checkpoint management documented in the README, and no stated rollback procedure. The README does not document rollback. The training script is a single file, train_gpt.py, with a medium variant alongside it. If your goal is to train a model you intend to deploy, you are looking at the wrong artifact. This repository exists to answer one question: how fast can this specific target be reached?

The third constraint is timing methodology. The README notes that torch.compile adds around 7 minutes of first-run latency, and that official records are timed on PrimeIntellect hardware. If you time a run on your own cluster without accounting for compilation and without matching CUDA and NCCL versions, your number will not be comparable to the leaderboard. The Docker path exists precisely to reduce that variance, but it does not eliminate differences in the underlying hardware.

## How this differs from NanoGPT and llm.c

NanoGPT, the project this one descends from, is a clean, readable implementation meant for learning and for reproducing GPT-2 training. It prioritizes clarity. Modded-NanoGPT inverts that priority: the training script is the accumulated output of dozens of contributors stacking micro-optimizations, and readability is not the goal. If you want to understand how transformer training works, read NanoGPT. If you want to see how far the training loop can be pushed, read this repository.

llm.c is the other reference point. Andrej Karpathy's GPT-2 replication in llm.c reached 3.28 validation loss after 45 minutes and 10B tokens. Modded-NanoGPT reaches the same loss in under 75 seconds and under 400M tokens. The difference is not one trick. It is the full list of architecture, optimizer, precision and systems changes applied together. A reader who wants a single change to adopt should look at the Muon optimizer, which has its own writeup and repository linked from the README, rather than trying to port the whole stack.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-18, four days before this article. That is recent enough that the code is moving, and the README's contributor list grows with each new record. The practical consequence is that the record changes. A technique that is in train_gpt.py today may be superseded, and the world record history section exists to track that progression. If you fork the repository to run your own experiments, expect to rebase against a moving target.

The licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are included. That is a permissive licence, but it applies to the code, not to the FineWeb dataset or to any model weights you train. The README does not address dataset licensing, so check the FineWeb dataset page before using it in a commercial context. This is not legal advice.

The upgrade cost is mostly dependency churn. requirements.txt pins torch==2.10 and typing-extensions==4.15.0, and the Dockerfile installs a nightly PyTorch build from the cu126 index on top of the pinned requirements. That combination means the Docker image and the pip path can diverge in PyTorch version. If you need reproducibility, the Dockerfile is the more controlled route because it fixes CUDA, NCCL and Python 3.12.7 together. The README does not document a version pinning policy beyond this, so treat the container as the reference environment.

## Conclusion

Adopt Modded-NanoGPT if you have 8 H100 GPUs and want to study or contribute to the fastest known training algorithm for a 124M model. Do not adopt it if you lack that hardware, need a production training pipeline, or want a general-purpose framework. Before running, verify that your CUDA and NCCL versions are compatible, since the README recommends Docker precisely for that case, and check whether torch.compile's roughly 7-minute first-run latency fits your timing budget.

## FAQ

### What hardware does Modded-NanoGPT require?

The speedrun target is defined on 8 NVIDIA H100 GPUs, and official records are timed on that hardware from PrimeIntellect. The README does not document expected results for other GPU configurations.

### How do I install and run Modded-NanoGPT?

Clone the repository, run pip install -r requirements.txt, download data with python data/cached_fineweb10B.py 9, then execute ./run.sh. The README notes that torch.compile adds around 7 minutes of latency on the first run.

### What is the difference between the main track and the optimization track in Modded-NanoGPT?

The main track minimizes wallclock time to reach 3.28 FineWeb validation loss on 8 H100s. The optimization track in records/track_3_optimization minimizes steps subject to fixed architecture, data and batch size with an unlimited wallclock budget.

## Sources

- [Issues](https://github.com/KellerJordan/modded-nanogpt/issues)
- [KellerJordan/modded-nanogpt on GitHub](https://github.com/KellerJordan/modded-nanogpt)
- [License: MIT](https://github.com/KellerJordan/modded-nanogpt/blob/master/LICENSE)
- [README](https://github.com/KellerJordan/modded-nanogpt/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kellerjordan-modded-nanogpt
