Nanotron: 3D-Parallel LLM Pretraining Without the Megatron Boilerplate
Minimalistic large language model 3D-parallelism training
At a glance
- What is it?
- Hugging Face's Nanotron is a minimalistic library for pretraining transformer models with data, tensor and pipeline parallelism. It trades Megatron-LM's breadth for explicit, debuggable APIs and a small config surface.
- Who is it for?
- Adopt Nanotron if you are pretraining a transformer from scratch on a cluster you control and you want tensor and pipeline parallelism you can step through in a debugger. Do not adopt it if you need a finished, packaged training product with built-in evaluation and rollout tooling; the README does not document either.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 6 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 24, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Nanotron Actually Solves
Pretraining a transformer at scale means splitting one model across many GPUs along three axes at once: data parallelism, tensor parallelism and pipeline parallelism. Most teams that attempt this end up either adopting Megatron-LM, which carries a large amount of framework code, or writing their own sharding layer, which is where most of the schedule bugs live. Nanotron sits between those options. The README describes it as a library for pretraining transformer models that provides a simple and flexible API to pretrain models on custom datasets, and the repository is small enough that the parallelism logic is readable rather than buried.
The intended user is an engineer who already knows what a pipeline bubble is and wants to control it, not a team looking for a managed training service. The repository layout supports that reading: run_train.py and run_generate.py sit at the top level, with training configurations living in examples/ as YAML files that map onto dataclasses. Nanotron's own feature list includes 3D parallelism (DP+TP+PP), expert parallelism for MoEs, AFAB and 1F1B schedules for pipeline parallelism, and ZeRO-1. Those are the primitives, and the library exposes them rather than hiding them.
How the Parallelism and Config Layer Fit Together
The mechanism is a config-driven launch. A YAML file describes the model, the parallelism degrees, the optimizer and the data source; run_train.py loads it, builds the process groups, and hands the model to a training loop that applies the chosen pipeline schedule. The README states that the project provides explicit APIs for tensor and pipeline parallelism which enables easy debugging, and that is the design claim worth weighing. Explicit APIs mean you can attach a debugger to a specific rank and inspect the tensor shard, instead of tracing a fused kernel. The cost is that you write more of the wiring yourself.
Two scheduling options are named in the feature list: AFAB and 1F1B. The choice affects memory pressure versus pipeline bubble size, and Nanotron leaves that decision to the config rather than picking for you. ZeRO-1 shards optimizer state, and FP32 gradient accumulation is listed as a separate feature, which matters because accumulating in reduced precision is a common source of silent divergence in long runs. Parameter tying and sharding are supported, as is custom module checkpointing for large models, which is the path you take when a single checkpoint file would exceed what one rank can write.
The repository also ships slurm_launcher.py for multi-node jobs and a docs/multi-node-training.md guide, so the intended deployment is a Slurm cluster rather than a single workstation.
Installing Nanotron and Training a Tiny Llama
The README's installation path uses uv rather than plain pip. First create the virtual environment and upgrade pip inside it:
uv venv nanotron --python 3.11 && source nanotron/bin/activate && uv pip install --upgrade pipNext install PyTorch from the CUDA 12.4 wheel index. Note that pyproject.toml declares requires-python as ~=3.10, so the 3.11 environment above is within range but 3.12 is not:
uv pip install torch --index-url https://download.pytorch.org/whl/cu124Then install the core package in editable mode, followed by the extras needed for the example scripts and fused kernels:
uv pip install -e .
uv pip install datasets transformers datatrove[io] numba wandb
uv pip install ninja triton "flash-attn>=2.5.0" --no-build-isolationThe README notes that flash-attn is pinned below 2.7.0 in pyproject.toml's fast-modeling extra, so a newer release may not resolve. Log in to Hugging Face and Weights and Biases before the first run:
huggingface-cli login
wandb loginFor the first real run, the README gives a single-node command that trains a tiny Llama on 8 x H100s in about 10 minutes:
CUDA_DEVICE_MAX_CONNECTIONS=1 torchrun --nproc_per_node=8 run_train.py --config-file examples/config_tiny_llama.yamlThe checkpoints land in the checkpoints directory named in the config file. To sample from the result, run generation against the checkpoint with tensor parallelism set to 1:
torchrun --nproc_per_node=1 run_generate.py --ckpt-path checkpoints/{checkpoint_number}/ --tp 1 --pp 1If you want to generate the config programmatically rather than edit YAML, the README points at examples/config_tiny_llama.py. For debugging, it documents a VSCode launch.json that runs torchrun with NANOTRON_BENCHMARK available as an environment flag and WANDB_MODE set to disabled.
Where Nanotron Is the Wrong Tool
The most concrete limitation is hardware. The quick start assumes 8 x H100s and a CUDA 12.4 PyTorch build. There is no documented CPU path, no documented single-consumer-GPU path, and no documented Apple Silicon path. If your only machine is a laptop with one GPU, the tiny Llama example is not runnable as written, and nothing in the README suggests a reduced configuration for that case.
A second limit is scope. The feature list covers training mechanics, not the surrounding workflow. The README documents run_evals.py only as a filename in the repository tree; it does not describe what evaluations it runs or how to configure them. There is no documented data cleaning pipeline inside Nanotron itself; the datatrove example delegates that to an external library. And the README does not document rollback for a partially written checkpoint, which matters when custom module checkpointing is writing shards across ranks and a job dies mid-write.
The release history is worth reading carefully. The most recent release listed is v0.4 from 2024-03-04, while pyproject.toml still declares version 0.4, yet the README's benchmark section refers to nanotron v0.5. So the package metadata, the release tags and the documentation disagree about the current version. The last push to the default branch was on 2026-09-23, so the repository is being touched, but the release cadence does not match the documentation's claims.
Nanotron Versus Megatron-LM
The comparison people search for is nanotron vs megatron, and the difference is one of surface area rather than capability. Both implement tensor, pipeline and data parallelism for transformer pretraining. Megatron-LM is a broader framework with a longer history and more built-in components; Nanotron's README frames the project around simplicity, stating that it provides a simple and flexible API to pretrain models on custom datasets.
In practice that means Nanotron asks you to bring your own data pipeline, your own evaluation harness and your own experiment tracking wiring, while giving you explicit tensor and pipeline parallel APIs you can read. If your team's bottleneck is understanding why a pipeline schedule stalls, the smaller codebase is the argument. If your bottleneck is that you need a component that Megatron-LM already ships, Nanotron will not have it, and you will be writing it. The benchmark data referenced in the README lives in a separate Hugging Face dataset, ultrascale-playbook-data, alongside a companion guide, the Ultrascale Playbook, so the tuning knowledge is documented outside the repository itself.
Maintenance, Licensing and Upgrade Cost
The last push to the default branch was on 2026-09-23, which is recent. The repository is not archived. However, the release tags stop at v0.4, dated 2024-03-04, and pyproject.toml pins version to 0.4 while the README's benchmark section discusses v0.5. Anyone tracking releases rather than commits should treat the version number as unreliable and pin to a specific commit hash instead.
Upgrade cost is shaped by the dependency pins. numpy is constrained below 2, and flash-attn is constrained to >=2.5.0,<2.7.0 in the fast-modeling extra. Both constraints will bite when you try to combine Nanotron with a newer stack. The datatrove dependency in the nanosets extra is pulled from a git URL rather than a released version, so that extra is not reproducible from PyPI alone.
Nanotron is licensed Apache-2.0, which permits commercial use and modification, and the LICENSE file is at the repository root. That is a permissive licence, and it differs from some pretraining frameworks that carry custom or copyleft terms. This is a factual note about the licence identifier, not legal advice; if you are redistributing a modified version or bundling it into a product, have your own counsel review the notice requirements.
Editorial conclusion
Adopt Nanotron if you are pretraining a transformer from scratch on a cluster you control and you want tensor and pipeline parallelism you can step through in a debugger. Do not adopt it if you need a finished, packaged training product with built-in evaluation and rollout tooling; the README does not document either. Before committing, verify that your GPU generation matches the CUDA index URL used in the install steps, and that your checkpoint size fits the custom module checkpointing path, because the README does not document rollback behaviour for a partially written checkpoint.
Frequently asked questions
How does Nanotron compare to Megatron for LLM pretraining?
Both implement tensor, pipeline and data parallelism. Nanotron's README frames it around simplicity and a flexible API for custom datasets, and it exposes explicit tensor and pipeline parallel APIs for debugging, while Megatron-LM ships a broader set of built-in components that you would otherwise write yourself.
How do I install Nanotron?
The README creates a virtual environment with uv, installs PyTorch from the CUDA 12.4 wheel index, then runs uv pip install -e . from the repository root. Example scripts need additional packages including datasets, transformers, datatrove[io], numba and wandb.
What hardware does the Nanotron quick start require?
The README's tiny Llama command targets a single node of 8 x H100s and uses torchrun with --nproc_per_node=8. No CPU, single-GPU or Apple Silicon path is documented.
What licence does Nanotron use?
The repository is licensed Apache-2.0, with the LICENSE file at the repository root. That permits commercial use and modification subject to the notice requirements in the licence text.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/huggingface-nanotron)