# byol-pytorch: wrapping any PyTorch image network for BYOL self-supervised pretraining

> lucidrains/byol-pytorch is a small wrapper that turns an existing image model into a BYOL learner without contrastive negative pairs. It is a research-grade library, not a training pipeline, and the README leaves checkpointing and rollback to you.

**lucidrains/byol-pytorch** — Usable Implementation of "Bootstrap Your Own Latent" self-supervised learning, from Deepmind, in Pytorch

- Repository: https://github.com/lucidrains/byol-pytorch
- Stars: 1,900 · Forks: 247
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/lucidrains-byol-pytorch

## What byol-pytorch actually removes from your training loop

BYOL, from Deepmind, learns representations without negative pairs. The README calls it an "astoundingly simple method" that surpasses SimCLR without contrastive learning and without designating negative pairs. That matters because contrastive methods need large batches to supply enough negatives, and large batches are the expensive part of self-supervised training. byol-pytorch packages that idea as a wrapper: you hand it an image-based neural network, and it adds the projection head, the prediction head, the target encoder and the loss. The README frames the audience as anyone with a residual network, a discriminator or a policy network who wants to "immediately start benefitting from unlabelled image data." If your labelled set is small and your unlabelled set is large, that is the case this library is built for. If you already have a working contrastive pipeline and enough negatives, the wrapper adds little.

## The two-encoder mechanism and the moving-average update you must call

The wrapper holds an online encoder and a target encoder. The online branch produces a projection, then a prediction from that projection; the target branch produces a projection only. The loss pushes the prediction toward the target projection. The target encoder is not trained by gradients. It follows the online encoder by an exponential moving average, and the README's example calls learner.update_moving_average() after every optimizer step. That call is not optional in the default configuration. Skip it and the target encoder never moves, which breaks the learning signal the method depends on. The moving_average_decay keyword controls the decay factor, and the README says it is already set to what the paper recommends. The README also links later work that replaced batch norm with group norm plus weight standardization, so the batch-statistics question is not settled by this repository; the wrapper simply exposes the original design.

## Installing byol-pytorch and running a first training step

The package installs from PyPI. The README gives a single command:

```bash
$ pip install byol-pytorch
```

After that, the usage example wraps a torchvision ResNet-50. You specify the image size and the name of the hidden layer whose output becomes the latent representation. The README uses the string 'avgpool' for that.

```python
import torch
from byol_pytorch import BYOL
from torchvision import models

resnet = models.resnet50(pretrained=True)

learner = BYOL(
    resnet,
    image_size = 256,
    hidden_layer = 'avgpool'
)
```

The README then builds an Adam optimizer at lr=3e-4 and runs a loop where each iteration samples a batch, computes the loss, and calls learner.update_moving_average(). What you should see is a scalar loss returned by learner(images); the README does not state expected loss values, so do not read anything into the first few hundred steps. When training finishes, the README saves the improved network with torch.save(resnet.state_dict(), './improved-net.pt'). Note that this saves the wrapped network, not the wrapper, so the projection and prediction heads are not in that file.

## Turning off momentum: the SimSiam variant and what changes

The README documents a second mode. A paper from Kaiming He suggests BYOL does not need the target encoder to be an exponential moving average of the online encoder, and byol-pytorch exposes that as use_momentum = False. In that configuration you no longer call update_moving_average at all, and the README's example omits it. This is a real simplification of the training loop, and it removes the decay hyperparameter from your concerns. The trade-off is that you are no longer running BYOL as described in the original paper; you are running SimSiam under a BYOL-shaped API. If your goal is to reproduce the paper's numbers, keep momentum on. If your goal is a lighter loop and you are willing to treat the target branch as a stop-gradient copy, the flag is there.

## Augmentations: the library's defaults and where they stop being enough

By default the wrapper uses the augmentations from the SimCLR paper, which the README notes are also used in the BYOL paper. You can replace them by passing augment_fn, and the README shows a kornia sequence. There is a second slot, augment_fn2, and the README explains why it exists: the paper appears to ensure that one of the augmentations has a higher Gaussian blur probability than the other. The README's example pairs a plain horizontal flip with a flip plus GaussianBlur2d((3, 3), (1.5, 1.5)). This asymmetry is easy to miss. If you pass a single augment_fn and expect the library to build the asymmetric pair for you, the README does not say that it does. The practical consequence is that the augmentation pipeline is your responsibility to get right, and the library will not warn you if both views are too similar.

## Distributed training through BYOLTrainer and Huggingface Accelerate

The repository adds a BYOLTrainer that takes your own Dataset and handles the loop. You configure Accelerate first, then launch the script through its CLI:

```bash
$ accelerate config
$ accelerate launch ./train.py
```

The trainer constructor in the README takes the network, the dataset, image_size, hidden_layer, learning_rate, num_train_steps, batch_size and checkpoint_every. The README states that with checkpoint_every = 1000 the improved model is saved periodically to a ./checkpoints folder. That is the extent of the checkpoint story in the README. There is no documented resume flag, no documented way to continue from a saved checkpoint, and no documented rollback if a run diverges. For a short experiment that is fine. For a multi-day run on a large unlabelled set, treat checkpoint handling as something you will have to build around the trainer rather than something it gives you.

## Where this wrapper is the wrong tool

Two limits stand out. First, the wrapper is image-based by design. The README describes wrapping image networks, and the augmentation defaults are image augmentations. If your data is text, audio or tabular, the pieces you would need to replace are most of the value the library provides. Second, the README points segmentation work elsewhere: it recommends a separate repository, pixel-level-contrastive-learning, for downstream tasks that involve segmentation. That is an explicit statement that this package is not the right starting point for pixel-level learning. A third limit is the packaging status. The pyproject.toml carries the classifier "Development Status :: 4 - Beta", and the pinned dependencies are torch>=1.6 and torchvision>=0.8 with a requires-python of >= 3.6. The last push to the repository was on 2026-04-27, and the most recent release listed is 0.8.2 from 2024-07-15, while pyproject.toml declares version 0.9.1. That gap between the declared version and the latest release is worth checking before you pin a version in a lockfile.

## Compared with DINO and other self-supervised wrappers

The related searches pair BYOL with DINO, and the difference in approach is worth stating plainly. BYOL trains two views of the same image through an online encoder and a momentum-updated target encoder, and the loss is a similarity between a prediction and a target projection. There are no negative pairs and no teacher network trained on a separate objective. DINO-style methods also use a momentum teacher, but they add a centering and sharpening step on the teacher's output to avoid collapse, which introduces temperature and centering parameters that BYOL does not have. byol-pytorch exposes neither of those, which is why its configuration surface is small: projection_size, projection_hidden_size, moving_average_decay and the augmentation functions. If you want the smaller configuration surface and you are training on images, this package is the shorter path. If you need the collapse-avoidance machinery or you are working outside images, look at a different implementation rather than bending this one.

## Licence and the cost of keeping up

The repository is MIT licensed, and pyproject.toml declares license = { text = "MIT" } with the matching OSI classifier. MIT is permissive, so wrapping it into a commercial training pipeline is not the constraint here; the constraint is that you inherit the dependency graph, which includes accelerate, beartype, einops, torch and torchvision. Upgrades are the real cost. The API is small enough that a breaking change would be visible quickly, but the release history shows long gaps between published versions, so a fix you need may exist on the default branch before it exists on PyPI. If you depend on a specific behaviour, pin the version and read the source in byol_pytorch/ rather than assuming the README describes the current code. This is not legal advice; check the licence text in LICENSE for the terms that apply to you.

## Conclusion

Adopt byol-pytorch if you already have a PyTorch image model and unlabelled images, and you want the BYOL objective without writing the projection head, predictor and moving-average update yourself. Do not adopt it if you need a maintained training pipeline with checkpoint resume, dataset plumbing and experiment tracking; the repository is a wrapper, and the README documents neither rollback nor a resume path. Before committing, verify three things: that your hidden_layer name matches the module you intend to pool, that your augmentation pipeline supplies the two differently blurred views the paper assumes, and that your PyTorch and torchvision versions satisfy the torch>=1.6 and torchvision>=0.8 pins in pyproject.toml.

## FAQ

### What does BYOL stand for in byol-pytorch?

BYOL stands for Bootstrap Your Own Latent, the self-supervised learning method from Deepmind that byol-pytorch implements. The README describes it as a method that surpasses SimCLR without contrastive learning and without designating negative pairs.

### What is BYOL (Bootstrap Your Own Latent) in byol-pytorch?

It is a self-supervised approach where an online encoder and a momentum-updated target encoder are trained on two views of the same image, and the loss pushes the online branch's prediction toward the target branch's projection. byol-pytorch wraps an existing image network so you can apply that objective to unlabelled images.

### How do I install byol-pytorch?

The README gives one command, pip install byol-pytorch, which installs the package from PyPI. The project also declares dependencies on accelerate, beartype, einops, torch>=1.6 and torchvision>=0.8, so those come along with it.

### Does byol-pytorch need negative pairs or a contrastive loss?

No. The README states that BYOL surpasses SimCLR without contrastive learning and without having to designate negative pairs, which is the reason the wrapper does not ask you for a batch of negatives.

### Can I use byol-pytorch without the momentum target encoder?

Yes. Setting use_momentum = False switches to the SimSiam variant described in the README, and in that mode you no longer call update_moving_average. The README notes this follows a paper from Kaiming He suggesting the exponential moving average target is not required.

## Sources

- [Issues](https://github.com/lucidrains/byol-pytorch/issues)
- [License: MIT](https://github.com/lucidrains/byol-pytorch/blob/master/LICENSE)
- [lucidrains/byol-pytorch on GitHub](https://github.com/lucidrains/byol-pytorch)
- [README](https://github.com/lucidrains/byol-pytorch/blob/master/README.md)
- [Releases](https://github.com/lucidrains/byol-pytorch/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lucidrains-byol-pytorch
