# nnue-pytorch: training Stockfish's NNUE evaluation networks

> nnue-pytorch is the PyTorch trainer behind Stockfish's NNUE evaluation. It ships three Dockerfiles, a data loader that has to be compiled, and an optional loop that plays the checkpoints against each other.

**official-stockfish/nnue-pytorch** — Stockfish NNUE (Chess evaluation) trainer in Pytorch

- Repository: https://github.com/official-stockfish/nnue-pytorch
- Stars: 499 · Forks: 153
- Language: Python
- License: GPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/official-stockfish-nnue-pytorch

## What nnue-pytorch is for, and who ends up using it

Stockfish's playing strength comes from an evaluation function implemented as a small neural network, and nnue-pytorch is the repository that trains that network in PyTorch. The README describes it plainly as a "Stockfish NNUE (Chess evaluation) trainer in Pytorch", and the top-level layout backs that up: train.py, a trainer/ package, a model/ package, a data_loader/ directory, and a serialize.py that converts trained checkpoints into the format the engine consumes.

The audience is narrow and technical. This is not a library you import into a chess app. It is a training pipeline for people who want to produce or modify the evaluation weights that ship inside a Stockfish binary. That means engine developers, people running their own net-training experiments, and contributors to the Stockfish project itself. If your goal is to analyze games or build a chess product, you want the engine, not the trainer.

The README points readers at two wiki pages for the actual training procedure, one labelled the "hard way" for train.py and one labelled the "easier way" for easy_train.py. That split is a useful signal: the repository expects two kinds of users, someone who wants to drive the training loop themselves and someone who wants a wrapper that handles the common path.

## The training loop, the compiled data loader and the serialization step

Three components carry the work. The data loader reads training positions; the trainer in trainer/ runs the optimization loop; serialize.py turns checkpoints into .nnue files. The README credits Sopel for "the amazing fast sparse data loader", which is the performance-sensitive part: NNUE input features are sparse, and a generic PyTorch DataLoader would waste most of its time on padding.

That loader is not pure Python. The repository ships compile_data_loader.sh and a Windows counterpart, compile_data_loader.bat, at the top level, and setup_script.sh is run after installing requirements. So the install is a two-stage affair: Python dependencies first, then a compilation step for the loader. If that compilation fails, training does not start.

The third stage is what makes this project specific to Stockfish rather than a generic chess model. Checkpoints are intermediate artifacts; serialize.py converts them into the .nnue format the engine loads. run_games.py automates the evaluation of that conversion. According to the README it "Automatically converts all .ckpt found under run96 to .nnue and runs games to find the best net", plays those games with c-chess-cli, and ranks the nets with ordo. It runs in a loop and monitors the directory for new checkpoints, so it can run alongside training on idle cores.

That design has a consequence worth naming: the quality signal is not a loss curve, it is game results against other nets. A checkpoint that looks good in TensorBoard can still lose matches.

## Installing nnue-pytorch with Docker and running a first training job

The README's default path is Docker, because it "eliminates the need for local Python environment setup and C++ compilation". The repository provides three Dockerfiles (Dockerfile.NVIDIA, Dockerfile.AMD and Dockerfile.CPU) plus run_docker.sh, which builds and starts the container. The README states the container includes CUDA 12.x or ROCm 6.4.3 and all required dependencies, and that your local CUDA or ROCm toolkit version does not matter. It also warns that building the container takes time and roughly 30 to 60 GB of disk space.

Start the container with the provided script:

```bash
./run_docker.sh
```

You are prompted for the target GPU vendor (or CPU only, for testing) and for the path to your data directory, which gets mounted into the container. The README notes the script also supports non-interactive workflows when all necessary arguments are passed on the command line. Once inside, training commands run directly.

For Apple Silicon the README takes a different route, since native MPS acceleration does not work with Docker. It recommends conda or micromamba:

```bash
conda create -n nnue_pytorch -c pytorch -c conda-forge \
    python=3.12 \
    pytorch \
    torchvision \
    torchaudio \
    compilers \
    llvm-openmp \
    jpeg \
    libjpeg-turbo \
    cmake \
    make
conda activate nnue_pytorch
```

Then install the Python requirements and run the setup script, which is where the data loader gets compiled:

```bash
pip install --no-cache-dir -r requirements.txt
./setup_script.sh
```

The requirements file pins python-chess==0.31.4, numpy<2.0 and ruff==0.16.0, and pulls in torchmetrics, schedulefree, numba, tensorboard and tyro. Those pins matter more than usual here: numpy<2.0 is an explicit ceiling, and python-chess is held at an old release.

Progress is watched through TensorBoard. The README gives the command and the address:

```bash
tensorboard --logdir=logs
```

After that, http://localhost:6006/ serves the dashboard. For the training commands themselves the README defers to the wiki, so the first real run means opening the train.py or easy_train.py page rather than copying a command from the README.

## Where nnue-pytorch gets in your way

The first limitation is the README itself. It does not document the training commands, the expected data format, or how to obtain training data. It links to wiki pages for the training procedure and marks the logging and game-running sections with "TODO: Move to wiki". A reader arriving at the repository cannot go from clone to trained net using only README.md.

The second is the resource profile. A 30 to 60 GB container build is not a casual experiment, and it needs a driver stack that matches the vendor-specific Dockerfile. CPU-only mode exists but the README frames it as "for testing purposes", which tells you it is not the intended training path.

The third is platform friction. Apple Silicon users get a separate Conda recipe precisely because Docker cannot give them MPS acceleration. That is a maintenance cost: two documented install paths, each with its own failure modes.

The fourth is the dependency pinning. numpy<2.0 and python-chess==0.31.4 will collide with a modern environment you already have. Use the container or the Conda environment rather than trying to install into an existing interpreter.

Finally, this is the wrong tool if you want a stable API. It is a training pipeline tied to one engine's network format, and serialize.py exists to produce files that Stockfish loads, not a general model export.

## nnue-pytorch compared with a standalone NNUE trainer

The obvious alternative class is a self-contained trainer such as the ones distributed as single binaries, where the whole pipeline including the data loader is compiled once and the training data format is fixed by that tool. The difference in approach is where the compilation boundary sits. nnue-pytorch keeps the training loop in Python and compiles only the data loader through compile_data_loader.sh, which is why the README pushes Docker: the container hides the C++ step behind an image build. A single-binary trainer has no Python environment at all, but it also gives you no PyTorch graph to modify.

That matters if your reason for training is research rather than producing a net. If you want to change the loss function, the feature set or the optimizer, the PyTorch loop in trainer/ is the point. The README's acknowledgements mention connormcmonigle and seer-nnue for "loss function advice", and dkappe for suggesting the ranger optimizer, which is a hint that these are the kinds of changes people actually make here.

If you only want to reproduce the current Stockfish network, the extra flexibility buys you nothing and the Docker build cost is pure overhead.

## Maintenance, licensing and the cost of staying current

The repository is not archived, and the last push was on 2026-07-26, so there is recent activity. There are no releases in the repository's recent history, which fits a project consumed from a branch rather than from versioned artifacts: you track master and rebuild.

That changes the upgrade calculus. Because there are no tagged releases, pinning means pinning a commit yourself. The requirements file pins individual Python packages, but nothing pins the repository's own state, so two people following the README a month apart may build different trainers.

The licence is GPL-3.0, and LICENSE sits at the top level. For anyone planning to distribute a modified trainer or a derivative of its code, that is a copyleft licence, and the practical consequence is that derivative distributions carry the same terms. This is not legal advice; if the licence affects a product you intend to ship, read LICENSE and get proper counsel. Note also that the trained .nnue weights are a separate question from the trainer's source code, and the README does not address it.

Upgrade cost is dominated by the container rebuild and the dependency pins rather than by API churn, since there is no API to speak of. Budget disk and build time, not migration work.

## Conclusion

Adopt nnue-pytorch if you already have a machine with a supported GPU, tens of gigabytes of free disk for the container image and a Stockfish-format training dataset, and you want to reproduce or modify the evaluation network rather than just use it. Do not adopt it if you want a general chess engine to embed, or if you expect a pip install and a quick fine-tune on a laptop CPU: the README routes Apple Silicon users to a Conda environment and notes that native MPS acceleration does not work with Docker. Before committing, verify three things: that your data directory is large enough to be worth mounting, that run_docker.sh finishes building (the README warns it takes time and roughly 30 to 60 GB of disk), and that compile_data_loader.sh produces a working loader, since training depends on it.

## FAQ

### What does NNUE stand for in nnue-pytorch?

The README does not expand the acronym. It describes the project only as a "Stockfish NNUE (Chess evaluation) trainer in Pytorch", so the expansion is not stated there.

### Does Stockfish use NNUE, and is nnue-pytorch how it is trained?

The repository is the official Stockfish organisation's NNUE trainer, and serialize.py converts trained checkpoints into .nnue files for the engine. The README does not describe the engine's internals beyond that.

### Is nnue-pytorch written in C++?

No, the primary language is Python, and the top-level files are train.py, config.py and serialize.py. There is one compiled component: the data loader, built through compile_data_loader.sh or compile_data_loader.bat.

### How does Stockfish work with the networks nnue-pytorch produces?

The README does not explain the engine's internals. What it shows is the boundary: run_games.py converts .ckpt files to .nnue and plays them with c-chess-cli, ranking the nets with ordo.

## Sources

- [Issues](https://github.com/official-stockfish/nnue-pytorch/issues)
- [License: GPL-3.0](https://github.com/official-stockfish/nnue-pytorch/blob/master/LICENSE)
- [official-stockfish/nnue-pytorch on GitHub](https://github.com/official-stockfish/nnue-pytorch)
- [README](https://github.com/official-stockfish/nnue-pytorch/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/official-stockfish-nnue-pytorch
