Lightning Thunder: a PyTorch compiler you can actually read the decisions of
PyTorch compiler that accelerates training and inference. Get built-in optimizations for performance, memory, parallelism, and easily write your own.
At a glance
- What is it?
- Thunder traces a PyTorch program into a Python-level intermediate representation, applies composable transforms over it, and hands each op to an executor such as nvFuser or cuDNN. The trade is inspectability in exchange for a 40 percent headline that belongs to the README, not to your model.
- Who is it for?
- Thunder is the rare PyTorch optimization project where the failed paths are visible rather than hidden behind a runtime log. You can look at the trace, see which transform fired, swap the executor for one you trust, and get numerical parity checked against eager PyTorch in a single `torch.testing.assert_close` call.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 3 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 8, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the three-layer design in the README actually says
The README describes Thunder through three components, and they are the right place to start because they explain everything else in the repository. The first is a simple, Pythonic intermediate representation that captures the entire computation. The second is a system of transforms that operates on the IR, the model and the weights at the same time. The third is an extensible dispatch mechanism pointing at fusers and optimized kernel libraries.
That third layer is the unusual part. Instead of a single hardcoded execution engine, Thunder lets each operation in the trace be handed to a pluggable executor. The README lists custom Triton kernels as one of the things you can do with that, which means a kernel you write by hand and a kernel from cuDNN are routed through the same path rather than requiring you to give up the rest of the pipeline to use yours.
The bullet list above that section is a feature manifest rather than a claim list, and it is worth reading as one. Quantization, kernel fusion, FP4, FP6 and FP8 precision, distributed tensor, pipeline and data parallelism, CUDA Graphs and readiness for NVIDIA Blackwell hardware all appear there, alongside training recipes, inference recipes and the note that it works for LLMs and non-LLMs alike. The last item is the interesting one: compose all of the above.
Composition is what distinguishes this from a list of flags. The README frames Thunder as a way to move from unoptimized to optimized by bundling optimizations into recipes that can be ported across model families, so the same set of transforms applies to a BERT forward pass and to a Llama training step.
The install matrix is the first thing you will hit
Installation from PyPI is one line, and then the README immediately tells you to update torch, which is a hint that this is not a drop-in that works with whatever you already have:
pip install lightning-thunder
pip install -U torch torchvision
pip install nvfuser-cu128-torch28 nvidia-cudnn-frontend # if NVIDIA GPU is presentThe nvFuser package name encodes two things: the CUDA version and the torch version. So the second and third arguments are `cu128` and `torch28`. The same line appears four times in the README, once for the current pairing and once each for torch 2.7 with CUDA 12.8, torch 2.6 with CUDA 12.6 and torch 2.5 with CUDA 12.4, with the matching `nvfuser-cu127-torch27`, `nvfuser-cu126-torch26` and `nvfuser-cu124-torch25` package names and matching torchvision versions. NVIDIA's cuDNN frontend is the other piece, and both are conditional on having an NVIDIA GPU present.
Three more paths sit in a collapsed advanced section. Optional executors, specifically Float8 support, come from installing `transformer_engine[pytorch]`, and the README warns in a comment that this compiles from source. Bleeding edge installs go straight from the `main` branch with pip and a git URL. Development installs clone the repository and run `pip install -e .`.
That branch install is arguably the more interesting option for reading the code, and it is the same code path the project itself uses for CI. The Makefile installs from `requirements.txt` and `requirements/test.txt` before running anything, and the base requirements file is a one-line `-r requirements/base.txt` redirect, so dependency resolution for tests goes through the same `requirements/` directory that setup.py reads at build time.
A hello world that checks its own numerics
The README's hello world is a three-layer linear model, which is a deliberate choice: it has two matmuls and a ReLU, so any trace is short enough to print and any numerical error is easy to attribute.
import thunder
import torch
thunder_model = thunder.compile(model)
x = torch.randn(64, 2048)
y = thunder_model(x)
torch.testing.assert_close(y, model(x))The final line is the part that makes this a compiler example rather than a demo. `torch.testing.assert_close` compares the compiled output against the eager output, so numerical parity is part of the first thing you run rather than something you add later. For a source-to-source compiler that rewrites operations into fused kernels, that check is the contract.
The two worked examples that follow are both realistic. The training example uses LitGPT with Llama 3.2 1B in bfloat16 on CUDA, runs a forward and backward pass, and requires `pip install --no-deps 'litgpt[all]'` so that LitGPT does not drag torch back to a different version. The inference example uses Hugging Face Transformers with `bert-large-uncased` in bfloat16, calls `requires_grad_(False)` and `eval()`, and the README notes that version 4.50.2 or above is recommended.
Both examples share the same shape: define the model in the framework you already use, call `thunder.compile` on it, run it. That is the whole adoption story. If your model loads through `from_pretrained` or a library constructor, Thunder never needs to know about it.
What the packaging setup reveals about release engineering
The packaging is unusual in a way that tells you how releases are produced. `pyproject.toml` declares `dynamic = ["version", "dependencies", "readme"]`, so the version, the dependency list and the long description are all filled in at build time by `setup.py` rather than written into the manifest. The dependencies come from the `requirements/` directory: setup.py walks each requirements file, skips comments, and drops any line containing `@` or `://`, which is how local paths and direct URLs get filtered out before they reach the published metadata. Every requirement is then re-serialized through `packaging.requirements.Requirement`, so the published dependency strings are normalized rather than copied.
The version itself is read out of `thunder/__about__.py` by scanning for the line that starts with `__version__`. There is an environment switch, `CONVERT_VERSION2NIGHTLY`, which turns that into a dated nightly version instead, and the readme badges show a pre-commit service running on every push. The `requirements/` directory is also where the optional dependency groups live: the `[project.optional-dependencies]` block in the manifest is present but entirely commented out, with `sdpa`, `apex` and `triton>=2.1.0` listed as commented examples and a note that the block gets populated from the requirements files during the build.
The version bounds are worth reading closely. `requires-python` is `>=3.10, <3.15`, while the classifiers stop at Python 3.13, and one classifier carries an inline comment saying to add 3.13 if supported and remove the `<3.14` guard when that happens. So the upper bound was set ahead of the tested range. The development status classifier says `3 - Alpha`, which is the single most useful piece of metadata in the file.
Testing, docs and the CI directories in the tree
The tree shows a repository with real engineering around it: `.azure/` and `.github/` for two CI systems, `.lightning/` for Lightning-specific configuration, a `.codecov.yml`, a `.pre-commit-config.yaml` and a `.git-blame-ignore-revs`. That last file only exists in projects that have run an automated formatter over their entire history, which tells you the formatting was fixed in one sweep rather than argued about commit by commit.
Tests run through the Makefile, and the default target is the interesting one:
python -m coverage run --source thunder -m pytest thunder tests -v
python -m coverage reportCoverage is measured over the `thunder` package while pytest collects from both the package and a `tests` directory, and the target depends on `clean`, which removes the mypy and pytest caches along with the docs build directory and any leftover checkpoint files named `_ckpt_*`.
The docs target is more unusual. It installs awscli, creates a `dist/` directory and runs `aws s3 sync --no-sign-request` against a public S3 bucket to fetch the `lai-sphinx-theme`, then installs it from that local directory with a `-f` find-links flag. After that it installs the package in editable mode together with `requirements/docs.txt` pointed at the CPU-only torch wheel index, and finally builds with `python -m sphinx -b html -W --keep-going`. The `-W` flag turns warnings into errors and `--keep-going` continues past the first one, which is the combination you want in a docs build that must not silently drop a page.
The public documentation URL in the manifest is `lightning-thunder.rtfd.io`, the README badge points at the `readthedocs.io` host, and the in-README navigation links to a `lightning.ai/docs/thunder/` path. Three different documentation addresses for one project, which is normal for a project hosted under a company domain but worth knowing before you file a docs issue.
Reading the README's performance claim carefully
The README opens with a grid of claims that includes running PyTorch 40 percent faster, and the project description calls it a compiler that focuses on making it simple to optimize models for training and inference. Both statements are marketing copy rather than a measurement you can check, and the honest way to read them is as the ceiling of what the project aims at rather than as a number attached to your workload.
The README gives you the tools to find out for yourself, and that is the more durable part of the pitch. The listed capabilities include profiling deep learning programs, mapping individual operations to kernels and inspecting programs interactively, and programmatically replacing sequences of operations with optimized ones to see the effect on performance. A compiler you can inspect mid-flight answers different questions than one that prints a speedup number, and it is the inspection path that matters when a fused kernel turns out to be slower than the unfused pair.
The remaining capabilities sit in the same category. Acquiring full computation graphs without graph breaks by extending the interpreter, modifying programs to take advantage of new kernel libraries on specific hardware, writing models for a single GPU and transforming them to run distributed, and iterating on mixed precision and quantization strategies to search for combinations that barely affect quality. Each of those is a statement about what you can do inside the compiler rather than about a fixed result.
The repository has no benchmark directory in its tree, only `examples/coverage/` and `examples/quickstart/`, and the three most recent releases are 0.2.4 in June 2025, 0.2.5 in September 2025 and 0.2.6 in October 2025. With the last push dated 2026-09-15 and the project still unarchived, the gap between the newest tag and the newest commit is a fair signal that `main` moves faster than releases do.
Editorial conclusion
Thunder is the rare PyTorch optimization project where the failed paths are visible rather than hidden behind a runtime log. You can look at the trace, see which transform fired, swap the executor for one you trust, and get numerical parity checked against eager PyTorch in a single `torch.testing.assert_close` call. If you want to understand why a fused kernel is faster, or you want to hand-write a Triton kernel and have it slot into the same dispatch path as cuDNN, that is the project to read. The counters to weigh against are the alpha classifier in `pyproject.toml`, the 40 percent speedup claim that belongs to the README banner rather than to any single model, and a gap of nearly a year between the 0.2.6 release in October 2025 and the last push in September 2026, which suggests the main line of work is landing on `main` faster than it is being tagged.
Frequently asked questions
What is Lightning Thunder used for?
It is a source-to-source compiler for PyTorch that traces a program into a Python-level intermediate representation, applies transforms to that IR, and dispatches each operation to an executor such as nvFuser or cuDNN. The stated use is moving models from unoptimized to optimized for training and inference, including quantization, mixed precision, kernel fusion and distributed execution.
How do you install Lightning Thunder?
Run `pip install lightning-thunder`, then update torch and torchvision and install the nvFuser build matching your CUDA and torch versions along with `nvidia-cudnn-frontend`. The README documents separate pairings for torch 2.8 with CUDA 12.8, torch 2.7 with CUDA 12.8, torch 2.6 with CUDA 12.6 and torch 2.5 with CUDA 12.4.
What is thunder.compile and how do I use it?
You pass a model or function to `thunder.compile` and then call the returned object like the original. The README's hello world compiles a small `nn.Sequential`, runs a random input tensor through it, and finishes with `torch.testing.assert_close` to confirm the compiled output matches eager PyTorch.
Can I write my own kernels and use them with Lightning Thunder?
Custom Triton kernels are listed among the things Thunder supports, because dispatch to fusers and kernel libraries is an extensible mechanism rather than a fixed execution engine. That lets a kernel you write by hand sit in the same dispatch path as the ones from cuDNN or nvFuser.
What licence and Python versions does Lightning Thunder support?
The project is Apache-2.0, with the licence file in the repository root and the identifier repeated in `pyproject.toml`. The manifest requires Python 3.10 or newer and below 3.15, with classifiers listed for 3.10 through 3.13, and marks the development status as alpha.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lightning-ai-lightning-thunder)