PyTorch: tape-based autograd, GPU tensors, and the setup.py exit plan
PyTorch is a Python package for GPU-accelerated tensor computation and deep neural networks built on a tape-based autograd system, extensible via NumPy and SciPy.
At a glance
- What is it?
- PyTorch is a Python tensor and deep learning library whose defining choice is a tape-based autograd system rather than a static graph. This review covers what it does, how to install it, where it stops being the right tool, and how the build system is changing.
- Who is it for?
- Adopt PyTorch if you need GPU tensor computation or a deep learning framework where the network structure can change between iterations, and if your team is comfortable with Python tooling. Do not adopt it if you need a static graph compiled ahead of time, or if you are building a project that cannot absorb a large binary dependency.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
DEEP OPEN-SOURCE ANALYSIS
Two features, and the second one is the reason people stay
The README describes PyTorch as a Python package with two high-level features: tensor computation with GPU acceleration, and deep neural networks built on a tape-based autograd system. The first is a NumPy replacement that can move arrays onto a GPU. The second is the part that shapes how you write code.
The audience is narrow but deep. If you are doing scientific computing on matrices and want GPU acceleration without leaving Python, the tensor half is enough. If you are training or fine-tuning neural networks and you want to inspect intermediate values while the network runs, the autograd half is the reason you are here. The README frames the library as usable either as a NumPy replacement or as a research platform, and those are genuinely different levels of commitment.
The README also lists the components that make up the library: torch for tensors, torch.autograd for differentiation, torch.jit (TorchScript) for producing serializable models, torch.nn for layers, torch.multiprocessing for sharing tensors across processes, and torch.utils for DataLoader and similar helpers. That list is worth reading before you start, because it tells you which import you need for which job.
Tape-based autograd versus the static graph
The mechanism the README emphasizes is reverse-mode automatic differentiation, described as using and replaying a tape recorder. Operations are recorded as they execute, and the backward pass replays that recording. The README contrasts this with frameworks it names directly (TensorFlow, Theano, Caffe, CNTK) that it says take a static view of the world, where the network structure is built once and reused.
The practical difference is what happens when the structure changes. With a static graph, changing how the network behaves means rebuilding it. With a tape, the README states you can change the network's behaviour arbitrarily with zero lag or overhead. That claim is the project's own framing, and it is the design decision everything else follows from. Dynamic control flow inside a forward pass, per-example branching, loops whose length depends on the input: these are ordinary Python in PyTorch and awkward in a static-graph framework.
The README also notes that tape-based autograd is not unique to PyTorch and credits prior work including autograd and Chainer. The claim being made is about the quality of the implementation, not the invention of the technique. That is a more honest framing than most project READMEs offer, and it is worth taking at face value: you are choosing an implementation, not a paradigm nobody else has.
The cost of the tape is that you cannot see the whole graph before execution. Compilation, ahead-of-time optimization and deployment to environments without a Python runtime are the areas where that shows up, and the torch.jit component exists to address part of it.
Installing PyTorch and running a first training loop
The README's Installation section covers three routes: binaries, building from source, and a Docker image. For most readers the binary route is the one that matters. The download page at pytorch.org is where the project points for the platform-specific install command, because the correct pip command depends on whether you want CPU-only, a CUDA build, or a ROCm build. The README does not print a single universal pip line, and that is deliberate rather than an omission.
The Dockerfile in the repository shows how the project's own images are assembled. It defaults to ubuntu:24.04, installs build-essential, cmake, git, python3-dev and related packages, then installs Python dependencies from requirements.txt. For the CUDA path it takes build arguments:
ARG BASE_IMAGE=ubuntu:24.04
ARG CUDA_PATH=cu121
ARG INSTALL_CHANNEL=whl/nightlyThe Dockerfile also carries a note that torchaudio does not publish wheels against the CUDA 13.2 index yet, so it is skipped when CUDA_PATH=cu132. That is a concrete example of the kind of version-matching problem you will hit: the three packages (torch, torchvision, torchaudio) do not always move together across CUDA builds.
Once installed, the README points to the beginner tutorial at pytorch.org/tutorials/beginner/basics/intro.html for learning the basics. The core loop you write has three parts: define a module with torch.nn, compute a loss, call backward on it, and step an optimizer. The autograd tape is what makes the backward call work without you writing derivatives by hand.
The setup.py deprecation is the most concrete migration cost
If you have build scripts that call python setup.py install, they are on a clock. The repository's setup.py contains an explicit deprecation schedule. PyTorch 2.14 and 2.15 forward install and develop to pip, while other commands fail with instructions. PyTorch 2.16 and 2.17 make all commands fail with instructions. PyTorch 2.18 removes the file.
The replacement commands are named in the file itself. For an install, it is spin install, or pip install --no-build-isolation -v . For an editable install, spin develop, or pip install --no-build-isolation -v -e . For a wheel, python -m build --wheel --no-isolation, and for an sdist, python -m build --sdist. Build customization through environment variables such as DEBUG=1, USE_CUDA=0 and MAX_JOBS continues to work with the replacement commands, according to the file.
The reason this matters beyond a renamed command is what sits underneath. pyproject.toml declares scikit-build-core as the build backend, and the comment in setup.py states that pip and python -m build never run setup.py at all through the PEP 517 interface. If your tooling assumed setuptools, the build path has already changed even if your command still works today.
pyproject.toml also shows how the parallelism knob is wired. MAX_JOBS is described as PyTorch's umbrella parallelism knob, aliased to CMAKE_BUILD_PARALLEL_LEVEL, and a user-set CMAKE_BUILD_PARALLEL_LEVEL always wins. If you have been setting MAX_JOBS and wondering why a build ignores it, that precedence rule is the answer.
Where PyTorch is the wrong choice
The tape is a runtime structure. If your deployment target is an embedded device, a mobile runtime, or a serving stack where you want a frozen graph with no Python interpreter present, the dynamic model works against you. The README's own framing points at this: the features it lists are research flexibility and GPU tensor math, not deployment. TorchScript exists in the component list as a compilation stack for serializable and optimizable models, but the README does not present it as a complete answer to every deployment scenario, and the component table is the only description it gives.
Build time is the second constraint. Building from source requires the prerequisites listed in the README, including CUDA support, ROCm support or Intel GPU support depending on your hardware, plus the dependencies in requirements.txt. The Dockerfile installs cmake, ccache, git and a long list of system packages before any Python work begins. For a team that wants a small dependency and a fast CI, a pure-Python numerical library will be a better fit, and the README itself suggests NumPy, SciPy and Cython as packages you can keep using alongside PyTorch rather than replacing.
There is also a version-matching burden that the Dockerfile makes visible. CUDA builds, ROCm builds and the companion packages (torchvision, torchaudio) have to line up. The file's own comment about torchaudio and the cu132 index is a small symptom of a recurring class of problem.
TensorFlow as the other side of the design argument
The README names TensorFlow among the frameworks with a static view of the world, where you build a network and reuse the same structure. That is the honest comparison, and it is a design difference rather than a feature checklist difference.
In a static-graph framework, you declare the computation first and then execute it. The graph can be inspected, optimized and serialized before it ever runs. That is an advantage for production pipelines where you want the whole computation visible to a compiler, and for deployment targets that cannot host a Python interpreter. The price is that the structure is fixed: changing behaviour means rebuilding.
In PyTorch, execution and graph construction happen together. You get a stack trace that points to where your code was defined, which the README calls out explicitly as a goal, and you can drop into a debugger mid-forward-pass. The price is that the graph does not exist until you run it, so ahead-of-time optimization has less to work with.
The README's own list of influences (torch-autograd, autograd, Chainer) is a reminder that this is a settled argument in the field, not an open one. Both approaches have production users. The question is which cost you would rather pay.
Maintenance cadence, licensing and what the release notes show
The repository is not archived. The last push was on 2026-07-08, which is the same timestamp as the v2.13.0 release. Before that, v2.12.1 landed on 2026-06-18 as a bug fix release, and v2.12.0 on 2026-05-13. The pattern visible in those three entries is a minor release roughly every month or two with patch releases in between, which is a faster cadence than many libraries and a real operational cost if you pin versions.
The README points to hud.pytorch.org for trunk health and CI signals, which is where you would look before building from a specific commit rather than a tagged release.
On licensing: the repository contains a LICENSE file and a NOTICE file at the top level, and the README's table of contents has a License section. The licence identifier was not available in the files reviewed here, so the file itself is the thing to read. If you are redistributing PyTorch inside a product, or shipping a container image built from the repository's Dockerfile, read LICENSE and NOTICE directly rather than relying on a summary. This is not legal advice, and the terms of the dependencies bundled into a build are a separate question from the terms of PyTorch itself.
Editorial conclusion
Adopt PyTorch if you need GPU tensor computation or a deep learning framework where the network structure can change between iterations, and if your team is comfortable with Python tooling. Do not adopt it if you need a static graph compiled ahead of time, or if you are building a project that cannot absorb a large binary dependency. Before committing, verify which CUDA build your hardware needs, check the release notes for the version you plan to pin, and read the deprecation schedule in setup.py if your build scripts call it directly.
Frequently asked questions
Is PyTorch a Python library?
Yes. The README describes PyTorch as a Python package, and its component table lists Python modules including torch, torch.autograd, torch.nn and torch.utils. It is built to be deeply integrated into Python rather than being a binding into a separate C++ framework, according to the README.
What is PyTorch versus TensorFlow?
The README frames the difference as static versus dynamic graph construction. It describes TensorFlow and several other frameworks as taking a static view where the network is built once and reused, while PyTorch records operations as they execute using reverse-mode autodifferentiation, which the README calls a tape-based approach.
How do I install PyTorch?
The README's Installation section lists binaries, building from source, and a Docker image. The binary route depends on your platform and whether you want a CUDA or ROCm build, which is why the README points to pytorch.org for the specific command rather than printing one universal pip line.
How do I use PyTorch with CUDA?
The README lists CUDA support under the prerequisites for building from source, and the repository Dockerfile takes a CUDA_PATH build argument defaulting to cu121. The README notes that torchaudio does not publish wheels against the CUDA 13.2 index yet, so that combination is skipped in the project's own Dockerfile.
How do I use PyTorch to train a model?
The README points to the beginner tutorial at pytorch.org/tutorials/beginner/basics/intro.html for learning the basics. The library components it lists for this are torch.nn for layers, torch.autograd for the backward pass, and torch.utils for the DataLoader.
Is ChatGPT written in PyTorch?
The repository files reviewed here do not state what ChatGPT is implemented in, so this cannot be answered from them. The README lists the components of PyTorch and its influences but does not mention ChatGPT or any specific commercial model.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/pytorch-pytorch)
Community notes