Open-source project
karpathy/nanochat avatar
karpathy/nanochat

nanochat derives every hyperparameter from --depth, and the recipe that won the speedrun is a moving commit

GitHub describes it as The best ChatGPT that $100 can buy.. The repository metadata lists Python as its primary language. The metadata lists the MIT license. This article stays within the project description and details documented in the GitHub repository README.

58,350 stars8,176 forksPythonMIT

At a glance

What is it?
A single-node LLM training harness that covers tokenization through inference and computes all but one hyperparameter for you, with a wall-clock leaderboard where GPT-2 grade capability went from 168 hours to 1.65, and a cost story that rests on a $3 per GPU hour assumption.
Who is it for?
Adopt it if your goal is to read, change and re-run a full LLM training pipeline rather than to ship a model, and if you can rent an 8XH100 node for the reference run or accept an eight times longer single-GPU version of the same experiment. Do not adopt it for a production assistant: a speedrun model is a 4e19 FLOPs model that hallucinates physics.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 22 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

One dial, --depth, and everything else is computed for you

nanochat bills itself as the simplest experimental harness for training LLMs: a single GPU node, minimal and hackable code, covering tokenization, pretraining, finetuning, evaluation and inference in one place. The design decision that follows from that is that you get one knob. `--depth` is the number of layers in the GPT transformer, GPT-2 capability sits at about depth 26, and the width of the transformer, the number of heads, the learning rate adjustments, the training horizons and the weight decays are all calculated automatically in what the file calls an optimal way.

The payoff is that a whole miniseries of compute-optimal models can be trained by changing a single number, which is exactly what runs/miniseries.sh is for. The cost is that the hyperparameters you might most want to study are the ones you cannot set. Anyone investigating a specific learning rate or decay has to go around the dial rather than turn it, and the repository does not document an escape hatch for overriding a computed value.

The project names the modded-nanogpt repository as the thing it took the speedrun leaderboard idea from, and the difference in approach shows in what gets optimised: a single wall-clock number on one fixed hardware shape, with everything else derived, rather than a general recipe you tune by hand.

The leaderboard is a wall-clock race, and the GPT-2 row took 168 hours

Time to GPT-2 is the wall-clock time needed to beat the GPT-2 (1.6B) CORE metric of 0.256525 on an 8XH100 node, and the table makes the scale of the claim obvious. Row zero is OpenAI's original 2019 GPT-2 checkpoint at 168 hours. Row one is a depth 24 baseline, slightly overtrained, at 3.04 hours on Jan 29 2026. From there: 2.91 hours for depth 26 slightly undertrained with fp8, 2.76 after bumping the total batch size to 1M tokens, 2.02 after changing the dataset to NVIDIA ClimbMix on Mar 4 2026 by two contributors, 1.80 for autoresearch round 1, and 1.65 for round 2 on Mar 14 2026.

Two details make the table usable rather than decorative. Every row records the commit that produced it, so any entry can be checked out and reproduced. And runs/speedrun.sh is stated to always reflect the current reference way to train a GPT-2 grade model and talk to it, which means the script and the leaderboard move together.

That second detail has a consequence for you. The 1.65 hour run is a property of commit a825e63, not of the repository as you clone it. master today runs a different recipe, and the file says the main focus of development is tuning the pretraining stage, which is where the compute goes. dev/LEADERBOARD.md is where the file says to look for how to interpret the table and how to contribute a row.

The cost claim is $3 per GPU hour, and that number belongs to your cloud provider

The headline figure comes with its arithmetic shown. GPT-2 cost about $43,000 to train in 2019, an 8XH100 node is about $24 per hour at roughly $3 per GPU hour, so two hours is about $48, and on a spot instance the total can land closer to $15. Those are provider prices, not measurements from this repository, and they are the load-bearing assumption in the whole argument.

If your provider prices that node shape differently, the number scales from $24 per hour. If spot capacity is unavailable when you want it, the spot figure stops applying. Nothing in the file documents a fallback for a provider that does not offer 8xH100 at all, and the scripts are described as designed for that node.

What you get for the money is worth being blunt about. A speedrun model is a 4e19 FLOPs capability model, described as a bit like talking to a kindergartener, and the example conversation in the file shows it inventing a physical explanation for why the sky is blue, complete with particles that bend light. The repository is a training pipeline you can take apart, not a chat product, and the dialogue with it is there to show the loop works end to end.

uv sync with one extra, then speedrun.sh does the rest

Dependencies come from uv, and you pick exactly one of two extras:

bash
uv sync --extra gpu    # Use for CUDA (A100/H100/etc.)
uv sync --extra cpu    # (or) Use for CPU-only / MPS
source .venv/bin/activate

For the research extras, add the dev group with `uv sync --extra gpu --group dev`, which brings in pytest, matplotlib, ipykernel and transformers. The two extras are declared mutually exclusive in pyproject.toml, so uv refuses to install both into one environment:

toml
[tool.uv]
default-groups = []
conflicts = [
    [
        { extra = "cpu" },
        { extra = "gpu" },
    ],
]

The routing behind that choice is two explicit indexes, pytorch-cpu and pytorch-cu128, and the comment in pyproject.toml says the target is CUDA 12.8 or CPU, with torch pinned to exactly 2.9.1. On hardware with a different CUDA version, neither configured index is the one you want, and fixing that means editing pyproject.toml.

With the environment active, the whole run is two commands, and the file suggests a screen session because it takes about an hour and a half:

bash
bash runs/speedrun.sh
python -m scripts.chat_cli

Under 80GB of VRAM, --device-batch-size is the knob you turn

The reference configuration assumes memory the reader may not have. If your GPUs have less than 80GB you will have to tune some hyperparameters or you will run out of VRAM, and the file points you at `--device-batch-size` in the scripts, to be reduced from the default of 32 down to 16, 8, 4, 2 or even 1, with the note that below that you need to know more than the file tells you.

The single-GPU path is a documented fallback rather than a broken one: omit `torchrun` and all the code runs on one GPU, automatically switching to gradient accumulation, producing roughly identical results, and taking eight times as long. The Ampere 8xA100 node also works, a bit slower.

Worth being clear about what that means for the numbers. A single-GPU run that matches the results takes eight times the wall clock, so it is the same experiment on a different scale, and it invalidates every time-to-GPT-2 figure in the leaderboard, which is defined on an 8xH100 node. The file also notes that most of the code is fairly vanilla PyTorch and should run on xpu or mps, while adding that not all of those code paths have been exercised and there might be sharp edges. Treat the CPU and MPS routes as untested configurations, not as supported ones.

The research path is a 12-layer run with evaluation switched off

For quick experimentation, with pretraining runs of about five minutes, the file recommends a 12-layer model, GPT-1 sized, launched like this:

bash
OMP_NUM_THREADS=1 torchrun --standalone --nproc_per_node=8 -m scripts.base_train -- \
    --depth=12 \
    --run="d12" \
    --model-tag="d12" \
    --core-metric-every=999999 \
    --sample-every=-1 \
    --save-every=-1

Those last three flags are the interesting part. `--core-metric-every=999999` pushes the CORE evaluation far out of the way, `--sample-every=-1` and `--save-every=-1` turn off periodic sampling and checkpointing, so a short run spends its time on the training step instead of on measurement and disk.

Researchers are pointed at two other scripts, runs/scaling_laws.sh and runs/miniseries.sh, with a Jan 7 miniseries v1 discussion thread linked for background. The test layout is conventional: pytest with testpaths set to tests, a slow marker you can deselect with -m "not slow", and a uv.lock committed at the root so the Python environment is reproducible.

The limit of this path is scale. A 12-layer GPT-1 sized run is for touching the code and seeing the loop move, not for producing anything you would evaluate as a language model, and the leaderboard metric it is aimed at is a pretraining speed number on hardware you are borrowing.

MIT with no releases, so you pin a commit or you pin nothing

The licence is MIT, with a LICENSE file at the root, and pyproject.toml names the project at version 0.1.0. There are no GitHub releases, so there is no tagged version to install, and the default branch is master with the last push on 2026-09-07.

For a repository whose central artefact is a leaderboard of wall-clock times, that is workable but worth being deliberate about. Each row is tied to a commit hash, and master is the only branch, so the reproducible unit is a commit rather than a version. If you need a fixed training recipe for your own records, write the commit down, because runs/speedrun.sh will not be the same script next month.

The tree also shows what kind of project this is: .claude/, .gitignore, .python-version, dev/, nanochat/, runs/, scripts/, tasks/, tests/ and uv.lock. The committed lock file is the part that helps most, since torch is pinned to exactly 2.9.1 and the two PyTorch indexes are explicit, so a fresh sync gives you the same packages. It also means a new torch release will not arrive on its own, and bumping it is an edit to pyproject.toml rather than a version bump you inherit.

Editorial conclusion

Adopt it if your goal is to read, change and re-run a full LLM training pipeline rather than to ship a model, and if you can rent an 8XH100 node for the reference run or accept an eight times longer single-GPU version of the same experiment. Do not adopt it for a production assistant: a speedrun model is a 4e19 FLOPs model that hallucinates physics. Before you start, check that your CUDA version matches the pytorch-cu128 index in pyproject.toml, that your GPUs have the 80GB the default --device-batch-size of 32 assumes, and that the spot price you have matches the roughly $24 per hour the cost claim is built on.

Frequently asked questions

How do I use nanochat?

Install with uv by picking one extra, gpu for CUDA hardware or cpu for CPU-only and MPS, then activate the environment with source .venv/bin/activate. The whole pipeline is in runs/speedrun.sh, which the file says is designed for an 8XH100 GPU node and takes about an hour and a half, after which you chat with python -m scripts.chat_cli.

What is nanochat?

It describes itself as the simplest experimental harness for training LLMs, running on a single GPU node with minimal, hackable code, and covering tokenization, pretraining, finetuning, evaluation and inference. The full run to chat loop is contained in one script, runs/speedrun.sh.

What is karpathy nanochat?

A Python training harness published under karpathy/nanochat, MIT licensed, with version 0.1.0 in pyproject.toml and no GitHub releases. The default branch is master, and the last push was on 2026-09-07.

How big is nanochat?

GPT-2 capability corresponds to a depth of about 26 layers, and a speedrun model is a 4e19 FLOPs model, described as a bit like talking to a kindergartener. For quick experiments the file suggests a 12-layer, GPT-1 sized model, and the leaderboard measures whether a model beats the GPT-2 (1.6B) CORE score of 0.256525.

Official sources

  1. Official README
  2. Project repository