Model or dataset
pytorch/vision avatar
pytorch/vision

TorchVision: Datasets, Transforms and Pre-trained Models for PyTorch

Datasets, Transforms and Models specific to Computer Vision

17,916 stars7,260 forksPythonBSD-3-Clause

At a glance

What is it?
TorchVision is the official companion library to PyTorch for computer vision. It bundles dataset loaders, image transforms and pre-trained model definitions, but its version table and dataset licensing terms are the parts most teams underestimate.
Who is it for?
TorchVision fits teams already committed to PyTorch who want dataset loaders, transform pipelines and pre-trained backbones without assembling them from separate repositories. It is the wrong choice if you need a framework-independent preprocessing layer, if you cannot pin torch and torchvision together, or if you plan to ship a pre-trained model whose training dataset license you have not checked.
Can I use it commercially?
Yes. BSD-3-Clause is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 12 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What TorchVision actually provides and who it is for

TorchVision is not a framework. The README describes it as a package of "popular datasets, model architectures, and common image transformations for computer vision", and that three-part scope is the whole product. Dataset classes wrap public corpora and handle download and preparation. Transform classes convert and augment images before they reach a model. Model definitions give you pre-trained weights for common architectures without writing the network yourself.

The intended user is someone already inside the PyTorch ecosystem. If your training loop is written in PyTorch, TorchVision slots in without an adapter layer. If you are working in TensorFlow, JAX or a non-Python stack, there is no reason to add it. The library's value comes from being native to PyTorch tensors and the torch data pipeline, and that value disappears the moment you leave that pipeline.

A second audience is less obvious: people who need a reference implementation. Model definitions and transform semantics are documented on the PyTorch website, and the repository also carries a references/ directory and an examples/ directory with C++ and Python examples. Those are useful when you want to check how a published architecture was translated into code rather than use it directly.

How the pieces fit: datasets, transforms, models and image backends

The data flow is deliberately simple. A dataset object yields samples, a transform pipeline converts those samples into tensors of the shape a model expects, and the model consumes them. TorchVision supplies each stage but does not force a particular composition on you.

The image backend is the part that surprises people. The README lists supported backends as torch tensors and PIL images, with Pillow as the standard implementation and Pillow-SIMD described as a drop-in replacement with SIMD. That means the transform layer has a real dependency on how images are represented before they become tensors, and the backend you pick affects what operations are available and how they behave.

The native side is more involved than a pure Python package. The repository contains a torchvision/csrc directory and a CMakeLists.txt at the top level, and setup.py reads environment variables including FORCE_CUDA, FORCE_MPS, DEBUG, TORCHVISION_USE_PNG, TORCHVISION_USE_JPEG, TORCHVISION_USE_WEBP and TORCHVISION_USE_NVJPEG. Those switches control which image codecs and accelerators get compiled in. The build script also defines TORCHVISION_INCLUDE and TORCHVISION_LIBRARY for pointing at headers and libraries, and TORCHVISION_PACKAGE_NAME for renaming the built package. In other words, a source build is a compiled extension build, not a copy of Python files.

Installing TorchVision and running a first transform pipeline

The README does not give a pip command. It points to the official PyTorch installation instructions at pytorch.org/get-started/locally/ for stable versions of torch and torchvision, and to CONTRIBUTING.md for a source build. So the first step is not a command from this repository at all; it is the version table.

The table maps torch releases to torchvision releases and to supported Python versions. For example, torch 2.13 maps to torchvision 0.28, and torch 2.12 maps to torchvision 0.27. Nightly torch pairs with nightly torchvision. Every row from torch 2.10 upward lists Python >=3.10 and <=3.14. Older rows narrow that range; torch 1.13 with torchvision 0.14 supports Python >=3.7.2 and <=3.10. Installing a mismatched pair is the most common way to get an import error that looks unrelated to versions.

The README's installation section gives the destination rather than a command, so the source build path from CONTRIBUTING.md is the one the repository documents in detail. The build reads its configuration from environment variables before it compiles the extension.

bash
python setup.py install

Running setup.py in the repository root triggers the build described above. Before it compiles anything, the script prints a "Torchvision build configuration" block listing the values it resolved for FORCE_CUDA, FORCE_MPS, DEBUG, USE_PNG, USE_JPEG, USE_WEBP, USE_NVJPEG, NVCC_FLAGS, TORCHVISION_INCLUDE and TORCHVISION_LIBRARY. Read that block first: it tells you whether CUDA sources will be built at all, since BUILD_CUDA_SOURCES depends on torch.cuda.is_available() and CUDA_HOME, or on FORCE_CUDA being set to 1.

Once a matching torch and torchvision pair is installed, the transforms module is documented at pytorch.org/vision/stable/transforms.html, which is where the repository points for transform behaviour. The repository's own examples live under examples/python and examples/cpp, and the README does not reproduce a transform snippet, so the API documentation is the place to confirm class names and argument order before writing a pipeline.

What you should expect from that documentation: a pipeline that resizes an image, crops it, and converts it to a tensor, with the conversion step placed after the geometric operations. If your tensor shape comes out wrong, check the order of operations before checking anything else.

The version table is a hard constraint, not a suggestion

TorchVision ships compiled extensions, and those extensions are built against a specific torch ABI. The README's compatibility table is therefore a compatibility requirement rather than documentation of a convention. The v0.29.0 release is titled "TorchVision 0.29: ABI stability!", which signals that ABI behaviour is treated as a release-level concern by the maintainers.

The practical consequence is that you cannot freely upgrade one of the two packages. Upgrading torch without upgrading torchvision, or the reverse, risks a binary incompatibility. This also affects environments where torch is pinned by another component, such as a CUDA-specific wheel or a vendor distribution. If that pin does not correspond to a torchvision release in the table, you have to either move the pin or build from source.

Python version support is part of the same constraint. The current rows cap at Python 3.14, and older torchvision releases cap lower. A team on an older Python interpreter may find that the torchvision release matching their torch release does not support their interpreter at all.

Datasets and pre-trained weights carry their own licenses

The BSD-3-Clause license on the repository covers the code. It does not cover the data or the weights, and the README says so twice, in a disclaimer on datasets and a separate section on pre-trained model licenses.

The dataset disclaimer states that TorchVision is a utility library that downloads and prepares public datasets, that it does not host or distribute them, and that it does not vouch for their quality or fairness or claim that you have license to use them. Determining whether you have permission under the dataset's own license is placed on you. The pre-trained model section makes the same point for weights, noting that they may carry licenses derived from the training dataset, and gives a concrete example: SWAG models are released under CC-BY-NC 4.0.

That example is worth pausing on. CC-BY-NC 4.0 is a non-commercial license. A team that pulls a model by name and ships it in a commercial product has a licensing problem that has nothing to do with the BSD-3-Clause code license they reviewed. The same reasoning applies per dataset: the library makes download convenient, and convenience is not permission.

When TorchVision is the wrong tool

The clearest case is a project that does not use PyTorch. TorchVision's transforms operate on PIL images and torch tensors, and its models are torch modules. If inference runs in a different runtime, you would be importing a large dependency to use a subset of it, and you would still need to reimplement the preprocessing to match what the pre-trained weights expect.

A second case is a preprocessing pipeline that must be identical across training and serving in a non-Python environment. TorchVision gives you the transforms in Python. Reproducing resize and crop semantics exactly in another language is your problem, and small differences in interpolation or rounding change model outputs.

A third case is a team that cannot pin torch. If your deployment environment resolves torch versions dynamically, or if multiple services share one environment with different torch requirements, TorchVision's ABI coupling turns every upgrade into a coordinated change.

Finally, if you only need one specific pre-trained architecture and nothing else, pulling TorchVision for its dataset loaders and transform library may be more surface area than the task requires. The library is broad by design, and breadth is a cost when you use one corner of it.

Alternatives and how their approach differs

The most direct alternative for preprocessing is to use Pillow, or Pillow-SIMD, directly and write your own transform functions. TorchVision already depends on that layer, and the README lists Pillow and Pillow-SIMD as supported image backends. Writing transforms yourself gives you full control over interpolation and ordering, at the cost of reimplementing what the library already tests. The trade-off is control versus maintenance.

On the model side, the alternative is to define architectures yourself or take them from a model-specific repository. TorchVision's advantage is that model definitions ship with the same versioning and the same ABI guarantees as the rest of the package. A model pulled from an individual research repository has no such coupling, which means it can break independently of your torch upgrade.

For dataset handling, the alternative is downloading and parsing each corpus yourself. TorchVision's dataset classes standardize that work, but as the disclaimer makes clear, they do not standardize the licensing. A team with strict data provenance requirements may prefer to control the download and preparation step directly, even if it means more code.

Maintenance, upgrade cost and license implications

The repository is not archived, and the last push was on 2026-09-09. Releases are frequent: v0.29.0 on 2026-09-02, v0.28.0 on 2026-07-08 and v0.27.1 on 2026-06-18. A cadence that tight is good for fixes and bad for anyone who wants to stay on an old version, because the compatibility table grows and old rows move into the collapsed "older versions" section.

The upgrade cost is the ABI coupling described earlier. Practically, an upgrade means moving torch, torchvision and your Python interpreter together, then re-validating any custom compiled code against the new pair. The release title for v0.29.0 suggests ABI stability is an active concern for the maintainers, which may reduce that cost over time, but the table remains the authority for what pairs are supported.

On licensing, the code is BSD-3-Clause, which is permissive. The datasets and the pre-trained weights are separate questions, and the README explicitly declines to answer them for you. The SWAG example shows that at least one family of weights is non-commercial. This is a description of what the documentation states, not legal advice; a team shipping a product should have its own review of the specific dataset and weight licenses it uses.

Editorial conclusion

TorchVision fits teams already committed to PyTorch who want dataset loaders, transform pipelines and pre-trained backbones without assembling them from separate repositories. It is the wrong choice if you need a framework-independent preprocessing layer, if you cannot pin torch and torchvision together, or if you plan to ship a pre-trained model whose training dataset license you have not checked. Before adopting it, verify three things: that your torch version maps to the torchvision version in the README table, that your Python version falls inside the stated range, and that the license of the specific dataset or pre-trained weight you intend to use permits your deployment.

Frequently asked questions

What is TorchVision used for?

It provides datasets, image transforms and pre-trained model architectures for computer vision in PyTorch. The README describes it as a package of popular datasets, model architectures and common image transformations.

How do I install TorchVision?

The README does not give a pip command; it refers to the official PyTorch installation instructions for stable torch and torchvision, and to CONTRIBUTING.md for a source build. Check the compatibility table first, since torch 2.13 maps to torchvision 0.28 and torch 2.12 maps to torchvision 0.27.

Which Python versions does TorchVision support?

The README lists Python >=3.10 and <=3.14 for torch 2.10 and newer, including nightly. Older pairings narrow that range, such as torch 1.13 with torchvision 0.14 supporting Python >=3.7.2 and <=3.10.

Can I use TorchVision's pre-trained models commercially?

The README states that pre-trained models may carry licenses derived from their training dataset, and that determining permission is your responsibility. It gives SWAG models as an example, released under CC-BY-NC 4.0.

Does TorchVision host the datasets it downloads?

No. The README's dataset disclaimer states that it is a utility library that downloads and prepares public datasets, and that it does not host or distribute them or claim you have license to use them.

Official sources

  1. License: BSD-3-Clause
  2. Project website
  3. pytorch/vision on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pytorch-vision.svg)](https://hysenlabs.com/projects/pytorch-vision)