Library / SDK
facebookresearch/pytorchvideo avatar
facebookresearch/pytorchvideo

PyTorchVideo: a video understanding library that ships models, datasets and transforms in one package

A deep learning library for video understanding research.

3,566 stars427 forksPythonApache-2.0

At a glance

What is it?
PyTorchVideo bundles video models, dataset loaders and video-specific transforms behind a plain pip install. It is useful for research code that needs a pretrained X3D or SlowFast checkpoint, and awkward for anything that expects long-term release cadence.
Who is it for?
Adopt PyTorchVideo if you want a pretrained video backbone such as X3D or SlowFast loaded through a TorchHub entry point and you are comfortable pinning to the 0.1.3 release from 2021. Do not adopt it if you need a library with a recent release cadence or documented upgrade path between versions; the repository has been pushed to as recently as 2026-05-05, but the last tagged release is still 0.1.3.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 147 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The gap PyTorchVideo fills between raw PyTorch and a full video pipeline

PyTorch itself has no opinion about video. A clip is a tensor with a time axis, and every research group ends up rewriting the same three pieces: a model that consumes that time axis, a dataset class that decodes frames without loading an entire file into memory, and transforms that crop and sample consistently between training and evaluation. PyTorchVideo packages those three pieces together. The README describes it as a library that provides "reusable, modular and efficient components needed to accelerate the video understanding research", and the top-level layout backs that up: pytorchvideo/ holds the library, tests/ holds its test suite, and tutorials/ holds runnable notebooks.

The audience is narrow on purpose. This is research infrastructure, not an application framework. If you are training or fine-tuning an action recognition model and you want a known-good SlowFast or X3D implementation to start from, the library removes a week of plumbing. If you are building a product that ingests user video and needs a supported, versioned dependency with a release every few months, the project's own history is a warning sign rather than a feature.

How the model zoo, transforms and data loaders fit together

The library splits into three layers that meet at the tensor boundary. Models live under pytorchvideo/models, with the vision transformer implementation in vision_transformers.py noted in the README's August 2021 update. Datasets and their loaders live separately, and transforms sit between them, converting decoded frames into the normalized clip tensor a model expects.

The glue is hubconf.py at the repository root. That file is what makes torch.hub.load work, and it is the reason you can pull a pretrained checkpoint without cloning the repository. The model zoo documentation lists the available pretrained models and their associated benchmarks, so the practical workflow is: pick a name from the model zoo page, load it through hubconf, and feed it clips shaped the way the corresponding transform produces them.

The dependency list in setup.py is short and worth reading before you install. It pins fvcore, av, parameterized, iopath and networkx. The av package is the PyAV binding for FFmpeg, which is where video decoding actually happens. That means the decode path is not pure Python and not pure PyTorch; it inherits FFmpeg's codec support and its platform packaging problems. On a machine where PyAV wheels are unavailable, the install fails before any PyTorchVideo code runs.

Installing PyTorchVideo and loading a pretrained model

The README gives one install command, inside a conda environment with Python 3.7 or newer:

bash
pip install pytorchvideo

For anything beyond that, the README points at INSTALL.md rather than repeating the steps. That file is the place to look when the plain pip install does not resolve, because the dependency chain includes PyAV and a PyTorch build that must match your CUDA setup.

The package metadata in setup.py declares python_requires=">=3.7" and the install_requires list shown above. If you are pinning dependencies in a lockfile, those five names are the ones to carry over.

Loading a model goes through TorchHub, which is why hubconf.py exists at the repository root. The README does not print a load call, so the shape below follows the standard torch.hub.load signature; substitute the model name you actually want from the model zoo listing rather than assuming a name resolves:

python
import torch

model = torch.hub.load("facebookresearch/pytorchvideo", "x3d_m", pretrained=True)

What you should see is a model object returned by the hub entry point. If the name is wrong, the failure comes from hubconf rather than from PyTorch, and it will list the entry points that do exist. That error message is the fastest way to discover the real model names on your installed version.

A pretrained checkpoint is not a deployment story

The README's headline demonstration is an X3D model running on a Samsung Galaxy S10 phone, accelerated by PyTorchVideo, processing one second of video in roughly 130 ms. That is a genuine result and it is also the most misread line in the README. It describes a model that was exported and optimized for mobile inference. It does not describe what happens when you call torch.hub.load on a server and run the model in eager mode. Those are different pipelines, and the library's contribution to the mobile case is the model architecture, not a runtime you get for free.

The second limitation is release cadence. The most recent tagged release is 0.1.3, published on 2021-09-10, preceded by 0.1.0 on 2021-04-13. The repository's last push was on 2026-05-05, so work has continued, but none of it has been cut into a version. If your dependency policy requires a release with a changelog, you are choosing between an old tag and tracking main. The README's Development section states that the maintainers "will be actively maintaining this library", and the commit history is consistent with that; the release history is not.

Third, the library assumes clips. It is built around fixed-length sampled segments, which is exactly right for action recognition benchmarks and awkward for streaming or variable-length input where you need frame-by-frame state. If your problem is online detection over an unbounded stream, the transform and loader abstractions fight you rather than help.

PyTorchVideo against TorchVision and a hand-rolled pipeline

The closest comparison is TorchVision, which most people already have installed. TorchVision ships image models, image transforms and image datasets. It has video dataset helpers and a couple of video classification models, but it does not ship the video-specific transforms or the pretrained video model zoo that PyTorchVideo does. The practical difference is that with TorchVision alone you write the temporal sampling and clip normalization yourself; with PyTorchVideo those arrive as library code you can call, and the models that consume them are pretrained and benchmarked in the same repository.

The other alternative is writing the pipeline directly against PyTorch and PyAV. That is what PyTorchVideo is, minus the shared interface. You gain full control over decode behavior and clip construction, which matters when your data does not look like a benchmark dataset. You lose the pretrained checkpoints and the ability to hand a colleague a model name instead of a repository path. For a team already comfortable with FFmpeg and PyTorch, the library's value is mostly the checkpoints, and if you only need one architecture it may be simpler to copy that model definition than to adopt the whole package.

Licence, maintenance cost and what an upgrade actually involves

PyTorchVideo is released under Apache-2.0, stated in both the README and the LICENSE file, and setup.py declares license="Apache 2.0". That is a permissive licence with a patent grant, and it is the same licence family as PyTorch itself, so combining the two in one project raises no obvious friction. This is a description of what the repository says, not legal advice; if you are redistributing a modified copy or shipping it in a product, have counsel read the actual LICENSE text.

The maintenance cost is the more interesting question. Because the last release is 0.1.3 from 2021 and the repository has been pushed to as recently as 2026-05-05, an upgrade is not a version bump. There is no published 0.2.0 changelog to read, so moving forward means diffing main against the tag yourself and running the test suite in tests/ to see what breaks. The upside is that the dependency list is stable and short, which limits how much can rot underneath you. The downside is that if a model name or a transform signature changes on main, nothing in the release metadata tells you.

If you pin to 0.1.3, budget for the possibility that a future PyTorch release breaks it and no patch version will arrive. That is the real cost of adopting a research library whose release train stopped.

Editorial conclusion

Adopt PyTorchVideo if you want a pretrained video backbone such as X3D or SlowFast loaded through a TorchHub entry point and you are comfortable pinning to the 0.1.3 release from 2021. Do not adopt it if you need a library with a recent release cadence or documented upgrade path between versions; the repository has been pushed to as recently as 2026-05-05, but the last tagged release is still 0.1.3. Before committing, verify that the installed torch version satisfies the pinned dependencies in setup.py and that the specific model name you intend to use actually resolves in hubconf.py.

Frequently asked questions

What exactly is PyTorchVideo used for?

It is a deep learning library for video understanding research, providing video models, video datasets and video-specific transforms built on PyTorch. The README describes it as offering reusable, modular and efficient components to accelerate video understanding research.

Is PyTorchVideo built on PyTorch rather than being a separate framework?

Yes. The README states that PyTorchVideo is developed using PyTorch and is built on it, which is what makes the rest of the PyTorch ecosystem usable alongside it. The library is installed as the pytorchvideo package.

Does PyTorchVideo work with TensorFlow?

Nothing in the repository suggests it does. PyTorchVideo is built on PyTorch, and the README's key features list describes it as based on PyTorch with access to PyTorch ecosystem components.

Official sources

  1. facebookresearch/pytorchvideo on GitHub
  2. License: Apache-2.0
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/facebookresearch-pytorchvideo.svg)](https://hysenlabs.com/projects/facebookresearch-pytorchvideo)