flownet2-pytorch: NVIDIA's PyTorch Port of FlowNet 2.0 Optical Flow
Pytorch implementation of FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks
At a glance
- What is it?
- flownet2-pytorch is NVIDIA's implementation of the FlowNet 2.0 optical flow estimation network in PyTorch, covering six network architectures and providing pre-trained weights converted from the original Caffe models. Its PyTorch 0.4.1 requirement and CUDA 9.0 Docker base image make it a research reference rather than a drop-in component for modern training pipelines.
- Who is it for?
- flownet2-pytorch is suited for researchers who need to reproduce or study FlowNet 2.0 results in a PyTorch environment, particularly those working with the MPI-Sintel benchmark. Engineers building production optical flow pipelines on modern PyTorch 2.x should evaluate whether they need a fork with updated dependencies, as the repository's stated requirement is PyTorch 0.4.1.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Optical Flow Estimation Is and Why FlowNet 2.0 Matters
Optical flow estimation is the problem of computing per-pixel velocity fields between consecutive video frames, representing how each region of an image has moved. FlowNet 2.0 is a convolutional neural network approach published at CVPR 2017, building on the original FlowNet by stacking multiple sub-networks and training them on progressively harder datasets to improve accuracy. The paper is by Ilg, Mayer, Saikia, Keuper, Dosovitskiy, and Brox.
flownet2-pytorch is the PyTorch port of that paper by Fitsum Reda, Robert Pottorff, Jon Barker, and Bryan Catanzaro at NVIDIA. It provides training and inference code, pre-trained model weights converted from the original Caffe implementation, dataset loaders for the standard evaluation benchmarks, and custom CUDA kernel implementations of the specialized layers the FlowNet architecture requires. The primary use case is computer vision research: training new optical flow models on standard datasets or running inference to reproduce the paper's results. Parts of the code were derived from ClementPinard/FlowNetPytorch, as acknowledged in the README.
Six Network Architectures and Their CUDA Requirements
The repository implements six FlowNet 2.0 architecture variants: FlowNet2S, FlowNet2C, FlowNet2CS, FlowNet2CSS, FlowNet2SD, and FlowNet2. A batch normalization version of each network is also available. The differences between variants relate to their internal structure: FlowNet2C uses a correlation layer between the two input image feature maps, while FlowNet2S uses a simpler stacked approach. The stacked variants (FlowNet2CS, FlowNet2CSS) chain sub-networks sequentially, with each network refining the flow estimate from the previous one. The FlowNet2SD sub-network is trained specifically on the ChairsSDHom dataset, which is designed for small displacements.
FlowNet2 and the FlowNet2C variants rely on two custom layers: Resample2d and Correlation. The README states that PyTorch implementations of these layers with CUDA kernels are available in the ./networks directory. There is a specific constraint here: the README notes that half-precision kernels are not currently available for the Resample2d and Correlation layers. Inference using fp16 (half-precision) is supported for architectures that do not use those custom layers, but not for FlowNet2 or the FlowNet2C family. This means a researcher wanting to run a full FlowNet2 or FlowNet2C variant with reduced memory footprint cannot simply enable fp16 and expect correct results from those custom layers.
Installation and the PyTorch Version Constraint
The README's installation sequence is:
git clone https://github.com/NVIDIA/flownet2-pytorch.git
cd flownet2-pytorch
bash install.shThe install.sh script compiles the custom CUDA kernels for Resample2d and Correlation. The README lists these Python requirements: numpy, PyTorch, scipy, scikit-image, tensorboardX, colorama, tqdm, and setproctitle. The stated PyTorch version is == 0.4.1. An older branch (python36-PyTorch0.4) is mentioned for PyTorch 0.4.0 and earlier.
PyTorch 0.4.1 was released in 2018. Current PyTorch versions are in the 2.x series. The README does not document compatibility with any PyTorch version beyond 0.4.1. The Dockerfile in the repository targets nvidia/cuda:9.0-cudnn7-devel-ubuntu16.04, which is CUDA 9.0 on Ubuntu 16.04, both of which reached end of life years ago. This means running the repository with its official Docker environment requires hardware and drivers that predate CUDA 10, 11, and 12. Researchers building on this code today would need to update the CUDA and PyTorch dependencies and recompile the custom kernels for a modern toolchain.
Pre-trained Weights and the Separate License
The README lists seven pre-trained model files converted from the original Caffe models, hosted on Google Drive. Their sizes range from 148 MB (FlowNet2-S) to 620 MB (FlowNet2). The full list is FlowNet2 (620 MB), FlowNet2-C (149 MB), FlowNet2-CS (297 MB), FlowNet2-CSS (445 MB), FlowNet2-CSS-ft-sd (445 MB), FlowNet2-S (148 MB), and FlowNet2-SD (173 MB). The CSS-ft-sd variant is fine-tuned on a sintel distractor dataset, indicated by the ft-sd suffix.
The README states that using these pre-trained weights requires adhering to the license agreements linked from a separate Google Drive file. The code repository itself carries a NOASSERTION license identifier, meaning GitHub's automated detection could not resolve it to a known SPDX expression. Researchers should read both the code LICENSE file and the pre-trained weight license before using either in any publication or product. The two licenses may carry different terms. A download script download_caffe_models.sh is present in the repository root, which the README implies is used to fetch the pre-trained files, though the exact invocation is not shown in the README text.
Training and Inference with Dataset Loaders
The README provides the main.py entry point for both training and inference, with --help as the documented path to see all available arguments:
python main.py --helpFor inference on the MPI-Sintel Clean dataset:
python main.py --inference --model FlowNet2 --save_flow --inference_dataset MpiSintelClean \
--inference_dataset_root /path/to/mpi-sintel/clean/dataset \
--resume /path/to/checkpointsFor training on MPI-Sintel Final with the L1 loss and Adam optimizer:
python main.py --batch_size 8 --model FlowNet2 --loss=L1Loss --optimizer=Adam --optimizer_lr=1e-4 \
--training_dataset MpiSintelFinal --training_dataset_root /path/to/mpi-sintel/final/dataset \
--validation_dataset MpiSintelClean --validation_dataset_root /path/to/mpi-sintel/clean/datasetThe README also shows a second training example using MultiScale loss on FlowNet2C with FlyingChairs as the training dataset and MpiSintelClean as validation, with --loss_numScales=5 and --loss_startScale=4, and a crop size of 384 by 512. This illustrates how the same main.py entry point handles different loss functions, crop sizes, and dataset pairings through command-line flags. Multiple GPU training is supported. The dataset loaders in datasets.py cover FlyingChairs, FlyingThings, ChairsSDHom, and ImagesFromFolder. L1 and L2 losses with multi-scale support are implemented in losses.py. The run_a_pair.py script is available for running inference on a single image pair without setting up a full dataset directory structure.
Comparing with RAFT and the Repository Status
RAFT (Recurrent All-Pairs Field Transforms) is a more recent open-source optical flow method, published in 2020, that achieves significantly better accuracy on the Sintel and KITTI benchmarks than FlowNet 2.0. RAFT is available in official PyTorch repositories and does not require custom CUDA kernels for its core method. Researchers beginning a new optical flow project today would typically start from RAFT or a derivative, not FlowNet 2.0. FlowNet 2.0 remains relevant for reproducibility work on benchmarks that include it as a baseline, or for studying the stacked-network design choices that influenced subsequent work.
NVIDIA also developed PWC-Net (code in both Caffe and PyTorch), which the README links to as related optical flow work. PWC-Net uses pyramid, warping, and cost volume operations and was a step between FlowNet 2.0 and RAFT in the progression of the field.
The last push to the flownet2-pytorch master branch was recorded on 2026-03-30. The repository is not archived.
Editorial conclusion
flownet2-pytorch is suited for researchers who need to reproduce or study FlowNet 2.0 results in a PyTorch environment, particularly those working with the MPI-Sintel benchmark. Engineers building production optical flow pipelines on modern PyTorch 2.x should evaluate whether they need a fork with updated dependencies, as the repository's stated requirement is PyTorch 0.4.1. Before using the pre-trained weights, read the separate license agreement linked from the README, since the code repository's license identifier is NOASSERTION and the model weights carry their own terms.
Frequently asked questions
Does flownet2-pytorch work with modern PyTorch versions?
The README specifies PyTorch 0.4.1 as the required version. It does not document compatibility with any later PyTorch version. The Dockerfile uses CUDA 9.0 and Ubuntu 16.04, both of which are outdated. Using the repository with current PyTorch 2.x would require updating dependencies and recompiling the custom CUDA kernels.
What datasets does flownet2-pytorch support for training?
The dataset loaders in datasets.py cover FlyingChairs, FlyingThings, ChairsSDHom, and ImagesFromFolder. The README shows training and inference examples using the MPI-Sintel Clean and Final datasets, and states that the same commands can be adapted for other datasets.
Are the flownet2-pytorch pre-trained weights under the same license as the code?
No. The README states that using the pre-trained weights requires adhering to a separate license agreement linked from the README. The code repository itself carries a NOASSERTION license identifier that GitHub could not resolve to a known SPDX expression. Both licenses should be reviewed before use.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nvidia-flownet2-pytorch)