LightlySSL: a PyTorch framework for self-supervised image pretraining
A python library for self-supervised learning on images.
At a glance
- What is it?
- LightlySSL exposes the building blocks of contrastive and distillation methods as PyTorch modules, with PyTorch Lightning handling distributed training. It fits teams that want to assemble their own SSL pipeline rather than call a pretrained model.
- Who is it for?
- Adopt LightlySSL if you have a PyTorch codebase, your own augmentation pipeline, and a reason to pretrain rather than fine-tune an existing checkpoint. Skip it if you want a single command that returns a pretrained encoder, since the README points those users at LightlyTrain and the commercial offering instead.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What LightlySSL solves, and who it is actually for
Self-supervised pretraining on images is not hard to describe and it is tedious to implement. Each method in the literature has its own loss, its own projection head, its own treatment of the momentum encoder or the queue, and its own augmentation recipe. Reimplementing SimCLR or SwaV from a paper means reproducing details that papers often leave implicit, and a mistake in the temperature or the queue update shows up only as a worse linear probe weeks later.
LightlySSL packages those pieces. The README describes it as a "computer vision framework for self-supervised learning" and lists a modular design that "exposes low-level building blocks such as loss functions and model heads" while staying in what the project calls a PyTorch-like style. That is the real pitch: you keep ownership of the training loop and the data pipeline, and you import the parts that are easy to get wrong.
The audience is narrower than the README's tone suggests. This is for engineers who already have a PyTorch project, a dataset loader, and an opinion about augmentations. It is not for someone who wants a pretrained checkpoint in an afternoon. The README itself redirects that reader: it points to LightlyTrain for SSL and distillation pretraining "in just a few lines of code", and to a commercial version with Docker support and pretraining models for embedding, classification, detection, and segmentation tasks.
The method zoo and what each entry costs you
The supported model table is the most informative page in the README, because it maps methods to years and papers rather than to difficulty. MoCo (2019), SimCLR (2020), BYOL (2020), SwaV (2020), DenseCL (2020) and SimSiam (2020) all appear, each with links to docs, a PyTorch Colab notebook, and a PyTorch Lightning Colab notebook. The repository also carries LeJEPA support, announced in the README's news list on May 28, 2026.
The list matters because these methods are not interchangeable in cost. SimCLR needs large batches to produce enough negatives, which is why the distributed examples exist at all. BYOL removes negatives and therefore removes the batch-size pressure, at the price of a target network and its momentum schedule. SwaV adds a queue and a multi-crop strategy. DenseCL targets dense prediction rather than a single image-level embedding. Choosing one is a decision about your GPU budget before it is a decision about accuracy.
The examples directory mirrors this: examples/pytorch/, examples/pytorch_lightning/, and examples/pytorch_lightning_distributed/ each hold runnable scripts, and examples/notebooks/ holds the Colab versions. The distributed folder is the honest signal that the single-GPU path is not the only intended one.
Installing LightlySSL and running a first pretraining step
The package is on PyPI as lightly, so the install is a normal pip install. The pyproject.toml declares dependencies on torch, torchvision, pytorch_lightning, hydra-core, numpy and tqdm, which means pip will pull a full PyTorch stack if you do not already have one.
pip install lightlyAfter that, the documented entry point for the modular API is the lightly.models and lightly.loss namespaces. The README's feature list describes loss functions and model heads as the exposed low-level pieces, and the per-method docs pages (for example the SimCLR page linked from the model table) carry the full training scripts. Those scripts import the building blocks the same way the per-method Colab notebooks do, wiring a backbone, a projection head and a contrastive loss together inside a training loop you control.
If you would rather not write that loop, the PyTorch Lightning examples in examples/pytorch_lightning/ wrap the same components in a LightningModule, and the distributed variants in examples/pytorch_lightning_distributed/ show the multi-GPU configuration. The repository keeps a PyTorch and a PyTorch Lightning example for every model in the table, so you can read both versions of the same method side by side.
For contributors, the Makefile is the other entry point. It runs everything through uv and states that no manual virtual environment activation is needed. The install-dev target installs the package for local development, and the Makefile pins a dependency cutoff with EXCLUDE_NEWER_DATE so that CI does not shift when a new dependency release lands.
Where the framework stops and you have to start
The modularity is the selling point and also the source of the friction. Because the framework exposes losses and heads rather than a complete training pipeline, everything around them is yours: the dataset, the two-view augmentation pipeline that contrastive methods depend on, the optimizer and schedule, checkpointing, and the linear-probe evaluation that tells you whether pretraining worked. The README does not document a built-in evaluation protocol or a rollback path for a pretraining run, so those are on you as well.
The Python floor is 3.8, stated in pyproject.toml with the note that it is the lowest version the project supports and tests. If your environment is pinned to 3.7 or older, this package is not installable. The dependency set is also not light: pytorch_lightning is a hard dependency even if you intend to use the plain PyTorch path, and hydra-core comes along with it.
There is a second boundary worth naming. If your goal is a production embedding model rather than a research pipeline, the README routes you elsewhere twice: to LightlyTrain for few-line SSL and distillation pretraining, and to the commercial version for Docker support and pretrained models across embedding, classification, detection and segmentation. LightlySSL is the layer underneath those, and treating it as a drop-in replacement for them will disappoint.
LightlySSL against timm and plain torchvision training
The closest alternative for many teams is not another SSL framework but timm plus a hand-written pretraining script. timm gives you backbones, augmentation helpers and a large collection of pretrained weights, and it lets you fine-tune a supervised checkpoint on your own labels. The difference in approach is what the checkpoint is trained on. timm's weights come from supervised ImageNet training; LightlySSL produces an encoder from your unlabeled data using a contrastive or distillation objective, with no labels at any point.
That distinction decides the use case. If you have a few thousand labeled images and a domain close to natural photographs, a supervised timm checkpoint will usually be the shorter path, and the LightlySSL README does not claim otherwise. If you have a large unlabeled corpus from a domain that ImageNet does not represent (medical, industrial, satellite), pretraining on that corpus is the reason to be here. The other practical difference is that timm hands you a model to fine-tune, while LightlySSL hands you a loss and expects you to build the run.
The project also lists a lightly optional extra that pulls in matplotlib, minimal, timm and video, so the two are not mutually exclusive: timm backbones can sit under a LightlySSL loss. The README's feature list explicitly supports custom backbone models for self-supervised pre-training.
Maintenance, releases and what the MIT licence covers
The repository is not archived, and the last push was on 2026-09-08. The release cadence is visible in the tags: v1.5.24 on 2026-05-28, v1.5.25 on 2026-06-17, v1.5.26 on 2026-07-27. That is a steady patch rhythm rather than a burst, and the version numbers stay in the 1.5.x line.
The licence is MIT, declared in LICENSE.txt and in the pyproject.toml classifier ("License :: OSI Approved :: MIT License"). For adopters that means the usual MIT terms: permissive reuse with the copyright notice retained. It says nothing about the commercial product the README mentions, which is a separate offering with its own terms and is not covered by this repository's licence. Nothing here is legal advice; read LICENSE.txt and, if you are shipping a product, have counsel read it too.
Upgrade cost is bounded by the dependency surface. Because pytorch_lightning and torch are direct dependencies, a major Lightning or PyTorch release is the event most likely to force a version bump on your side. The Makefile's EXCLUDE_NEWER_DATE pin exists precisely because dependency drift breaks CI, which tells you the maintainers expect that pressure. The project also ships a uv.lock and a uv-based Makefile, so reproducing a known-good environment is a single command rather than a manual pin list.
Editorial conclusion
Adopt LightlySSL if you have a PyTorch codebase, your own augmentation pipeline, and a reason to pretrain rather than fine-tune an existing checkpoint. Skip it if you want a single command that returns a pretrained encoder, since the README points those users at LightlyTrain and the commercial offering instead. Before committing, verify which method matches your compute budget and confirm that the loss and head you pick are exposed as modules you can wire into your own training loop.
Frequently asked questions
How do I install LightlySSL for self-supervised learning?
The package is published on PyPI as lightly, so pip install lightly is the install path. It requires Python 3.8 or newer and pulls in torch, torchvision and pytorch_lightning as dependencies.
What Python versions does lightly-ai/lightly support?
pyproject.toml sets requires-python to >=3.8 and lists classifiers for 3.8 through 3.12, with a comment that 3.8 is the lowest version the project supports and tests. The Makefile uses the same range for its minimum and maximum test versions.
Which self-supervised learning methods does lightly-ai/lightly implement?
The README's model table lists MoCo, SimCLR, BYOL, SwaV, DenseCL and SimSiam, each with a paper link and PyTorch and PyTorch Lightning Colab notebooks. The news section also notes LeJEPA support as of May 28, 2026.
Does lightly-ai/lightly include pretrained models I can download?
The README does not describe pretrained checkpoints in the open source package. It points readers to a commercial version with pretraining models for embedding, classification, detection and segmentation tasks, and to the separate LightlyTrain project for few-line pretraining.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lightly-ai-lightly)