Library / SDK
Lightning-AI/pytorch-lightning avatar
Lightning-AI/pytorch-lightning

PyTorch Lightning: the LightningModule and Trainer, and when plain PyTorch is enough

Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.

31,362 stars3,801 forksPythonApache-2.0

At a glance

What is it?
PyTorch Lightning wraps your PyTorch training loop in a LightningModule and a Trainer so the same code runs on one device or many. Here is what that buys you, what it costs, and how Fabric differs.
Who is it for?
Adopt PyTorch Lightning if you are maintaining your own training loop across multiple GPUs or nodes and want the loop written once. Skip it if you are on a single device with a short script, or if you need to modify step-level control flow that the Trainer owns; Lightning Fabric exists for that case.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 9 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 26, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The repetitive engineering code PyTorch Lightning removes

A plain PyTorch training script mixes two things that change at different rates: the model and the optimization logic, and the plumbing around them (device placement, gradient accumulation, mixed precision, distributed process groups, checkpoint files). The README states the second category is "error-prone and often reimplemented for every project," which is the honest version of the pitch. PyTorch Lightning's answer is to move the plumbing into a Trainer object and leave you a class with a fixed set of methods to fill in.

The audience is narrow but deep. If you train on one GPU with a short script, the abstraction costs you more than it saves. If you already have a multi-node launcher, gradient accumulation, and a checkpointing convention that three people understand, Lightning is aimed squarely at you. The README frames the target as scaling "from CPU to multi-node GPUs without changing your core code," and the project description claims the same code runs on 1 or 10,000+ GPUs. Those are claims about the intended design, not measured results, and the README does not publish scaling numbers to support them.

The README also points to LitServe for serving, which is a separate package. Lightning itself is a training framework; do not expect it to give you an inference server.

LightningModule, Trainer, and where the loop actually lives

The central object is the LightningModule, described in the README as an nn.Module subclass that "defines a full system" such as an LLM, a diffusion model, an autoencoder, or an image classifier. You keep your layers as ordinary PyTorch modules and add the training hooks. The README's toy example builds an encoder with nn.Sequential(nn.Linear(28 * 28, 128), nn.ReLU(), nn.Linear(128, 3)) inside a LitAutoEncoder class. The truncated snippet does not show the training_step, configure_optimizers or dataloader methods, but the class shape makes the contract clear: you describe what a step computes, and the Trainer decides when to call it, how to move tensors to devices, and when to save.

The second package is Lightning Fabric, which the README presents as the expert-control option. Where the Trainer owns the loop, Fabric gives you the primitives and leaves the loop to you. The README's framing is that Lightning "gives you granular control over how much abstraction you want to add over PyTorch," and that PyTorch experts can opt into the expert-level control path. In practice that means a project can start on the Trainer and drop to Fabric for a component that needs bespoke control flow, without leaving the ecosystem.

The analogy the README offers is that "if PyTorch is Javascript, PyTorch Lightning is ReactJS or NextJS." It is a useful way to set expectations: you are not replacing PyTorch, you are adopting a structure on top of it, and the underlying modules, tensors and optimizers remain PyTorch objects.

Installing PyTorch Lightning and running the first module

The README gives a single install command for the umbrella package, which pulls in both PyTorch Lightning and Fabric. Run it in the environment where your PyTorch is already installed:

bash
pip install lightning

After that, `import lightning as L` is the import the README's example uses. The project also documents optional dependency installs and a conda path, which are worth knowing about if you are behind a pip proxy or manage environments with conda:

bash
pip install lightning['extra']
conda install lightning -c conda-forge

The README's quick start example needs torchvision in addition to Lightning, and it is run as a normal Python file:

bash
pip install torchvision
python main.py

The example file begins with the imports and the module definition. The README shows the class opening, not the full file, so the first thing to check after install is whether your environment resolves `lightning` and `torchvision` to compatible versions. If you want the bleeding-edge build instead of the released one, the README documents installing from the master branch zip, and explicitly labels it "no guarantees." For a first real use, the README points to a set of hosted examples covering image classification, image segmentation and object detection, each described as a finetune of a specific architecture such as ResNet-34, ResNet-50 or Faster R-CNN. Starting from one of those is a shorter path than writing a module from scratch.

What you give up when the Trainer owns the loop

The trade-off is control over the order of operations. Once the Trainer drives the loop, code that needs to interleave custom logic between backward and optimizer step, or that changes behavior per batch based on state the Trainer does not know about, has to be expressed through hooks the framework provides. If your research depends on a nonstandard inner loop, the abstraction becomes a fight rather than a convenience. Fabric is the documented escape hatch for exactly this, and choosing it means writing the loop yourself.

The second limitation is version coupling. PyTorch Lightning tracks PyTorch closely, and a training stack that pins an old PyTorch for hardware or driver reasons can find itself outside the supported range. The repository's own pyproject.toml targets Python 3.9 for its linting configuration, which tells you the project's floor but not its ceiling.

The third is that the README is a marketing surface as much as a manual. It does not document rollback behavior, checkpoint compatibility across major versions, or what happens to a run when a node fails mid-training. The repository has a docs directory and a SECURITY.md, and those are the places to look for the details the README omits. Anyone evaluating this for production training should read the version upgrade notes before pinning, because the release cadence is fast: 2.6.1, 2.6.4 and 2.6.5 all shipped between January and May 2026.

PyTorch Lightning versus plain PyTorch, and versus Fabric

The comparison people search for is PyTorch Lightning against PyTorch, and the honest answer is that it is not a replacement. The README says Lightning is "just organized PyTorch," and the underlying model is still an nn.Module. The difference is who writes the training loop. In plain PyTorch you write the epoch loop, the device moves, the gradient scaling and the checkpoint save. In Lightning you write the step logic and the Trainer does the rest. For a single-GPU experiment the plain version is shorter and has fewer moving parts. For a codebase that has to run on one GPU for debugging and eight for a real run, the Lightning version is the one that does not fork.

The more interesting comparison is inside the project. The Trainer and Fabric solve the same problem at different levels. The Trainer is the opinionated path: define a module, hand it to the Trainer, get distributed training, precision and checkpointing configured through arguments. Fabric is the unopinionated path: you keep your loop and call into Fabric for the parts that need to know about devices and processes. The README positions Fabric for "expert control," and the repository keeps separate example directories for the two, examples/pytorch/ and examples/fabric/, with separate runner scripts. That split is a reasonable signal that the two are maintained as distinct entry points rather than one being a thin alias for the other.

Keras is a different axis entirely. Lightning stays inside the PyTorch object model, so your nn.Module, optimizer and scheduler remain usable outside Lightning. A framework that owns the model definition does not give you that.

Maintenance, releases and the Apache-2.0 licence

The repository is not archived, and the last push was on 2026-05-27, the same date as the 2.6.5 release. That is recent enough to call the project maintained on the evidence available, but the release history shows the shape of the cost: 2.6.1 in late January 2026, then 2.6.4 and 2.6.5 in May, with patch releases arriving close together. Pinning a version and reading its upgrade notes is cheaper than tracking the branch. The README documents a bleeding-edge install from the master zip and labels it as having no guarantees, which is the project telling you directly that the branch is not the supported path.

The licence is Apache-2.0, and the repository carries a LICENSE file at the top level. For most teams that means you can use, modify and redistribute the library, including in commercial products, provided you keep the licence and notice files. Apache-2.0 also includes an explicit patent grant, which matters more to some legal departments than the permission to copy. This is not legal advice, and if your organization has a policy on copyleft or on contributor licence agreements, the LICENSE file and the repository's contribution setup are the primary sources to check rather than a summary.

The Makefile shows the development setup path, which is useful if you intend to patch the library rather than only consume it. It uninstalls any pre-installed lightning packages, installs the requirements files for both the pytorch and fabric trees plus typing and test requirements, and installs the project in editable mode with the "all" extra before running pre-commit install. That is a heavier environment than a normal user needs, and it is the right one only if you are contributing.

Editorial conclusion

Adopt PyTorch Lightning if you are maintaining your own training loop across multiple GPUs or nodes and want the loop written once. Skip it if you are on a single device with a short script, or if you need to modify step-level control flow that the Trainer owns; Lightning Fabric exists for that case. Before committing, verify on your own hardware that the Trainer's strategy, precision and checkpoint settings match what your current loop does, and read the version notes for the release you pin.

Frequently asked questions

What is PyTorch Lightning used for?

It is a deep learning framework for pretraining and finetuning models, built on top of PyTorch. It organizes training code into a LightningModule and a Trainer so that backpropagation, mixed precision, multi-GPU and distributed training do not have to be reimplemented per project.

How do I install PyTorch Lightning?

The README gives a single command, pip install lightning, which installs both the PyTorch Lightning and Fabric packages. It also documents pip install lightning['extra'] for optional dependencies and conda install lightning -c conda-forge.

What are the key differences between PyTorch Lightning and Keras?

The README does not compare the two. What it does state is that Lightning is organized PyTorch, so your models remain nn.Module subclasses and your optimizers stay PyTorch objects. The README does compare the abstraction level to React or NextJS on top of JavaScript.

What is PyTorch Lightning Fabric?

The README describes Fabric as the expert-control package, the second of Lightning's two core packages. Where the Trainer owns the training loop, Fabric leaves the loop to you and provides the pieces needed to scale it, so you choose how much abstraction to add over PyTorch.

What is a PyTorch Lightning module?

A LightningModule is an nn.Module subclass that defines a full system, such as an LLM, a diffusion model, an autoencoder or an image classifier. You keep your ordinary PyTorch layers inside it and add the training logic the Trainer calls.

What does PyTorch Lightning do?

It automates the training infrastructure around a PyTorch model: backpropagation, mixed precision, multi-GPU and distributed training, and checkpointing, while leaving the model logic to you. The README describes it as organized PyTorch, with the Trainer driving the loop and a LightningModule defining the system.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/lightning-ai-pytorch-lightning.svg)](https://hysenlabs.com/projects/lightning-ai-pytorch-lightning)