# CoTracker3: Transformer-Based Joint Point Tracking for Video from Meta AI

> CoTracker3 is a research model from Meta AI and the University of Oxford VGG group that tracks any pixel or set of pixels across every frame of a video. Its third version achieves state-of-the-art tracking accuracy with a lightweight architecture trained using pseudo-labeling on real videos, requiring roughly 1000 times less data than previous top-performing models to reach comparable results.

**facebookresearch/co-tracker** — CoTracker is a model for tracking any point (pixel) on a video.

- Repository: https://github.com/facebookresearch/co-tracker
- Website: https://co-tracker.github.io/
- Stars: 5,131 · Forks: 392
- Language: Jupyter Notebook
- License: NOASSERTION
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/facebookresearch-co-tracker

## What CoTracker3 does and what problem it addresses

Point tracking in video is the task of following specific pixels or regions across frames. Traditional optical flow methods compute a dense motion field for every pixel between consecutive frames but do not maintain point identity across long sequences. CoTracker addresses the limitation of frame-to-frame tracking by computing tracks jointly across multiple frames simultaneously, which improves accuracy when objects temporarily leave the frame, are occluded, or deform.

CoTracker3 can track any pixel in a video, a quasi-dense set of pixels selected together, or a grid of points sampled automatically on any video frame. The model is a transformer architecture that processes temporal windows of frames and propagates tracking information across the window. The README describes it as bringing to tracking some of the benefits of optical flow, meaning it inherits the dense motion understanding of flow while adding temporal consistency across longer sequences.

The primary audience is computer vision researchers and practitioners working on problems that require knowing where specific points move over time: video editing, motion analysis, 3D reconstruction from video, sports analytics, and similar applications. The README notes that a related project, VGGSfM, uses CoTracker's point tracking approach to recover camera poses and 3D structure from image sequences.

## Loading CoTracker3 from PyTorch Hub

The fastest way to use CoTracker3 is through PyTorch Hub, which downloads the pretrained model weights automatically. The offline mode processes the entire video at once:

```python
import torch
# Download the video
url = 'https://github.com/facebookresearch/co-tracker/raw/refs/heads/main/assets/apple.mp4'

import imageio.v3 as iio
frames = iio.imread(url, plugin="FFMPEG")  # plugin="pyav"

device = 'cuda'
grid_size = 10
video = torch.tensor(frames).permute(0, 3, 1, 2)[None].float().to(device)  # B T C H W

# Run Offline CoTracker:
cotracker = torch.hub.load("facebookresearch/co-tracker", "cotracker3_offline").to(device)
pred_tracks, pred_visibility = cotracker(video, grid_size=grid_size) # B T N 2,  B T N 1
```

The input tensor shape is B T C H W: batch size, number of frames, channels, height, and width. The output `pred_tracks` has shape B T N 2 representing the x,y coordinates of each tracked point at each frame. The output `pred_visibility` has shape B T N 1 indicating whether each point is visible at each frame.

The `imageio` package with the FFMPEG plugin is a prerequisite for loading video files. Install it with `pip install imageio[ffmpeg]` before running the offline mode example.

## Online mode for long or streaming videos

The online mode processes video in chunks rather than all at once, making it more memory-efficient for long videos. The API differs from offline mode in that it is stateful:

```python
cotracker = torch.hub.load("facebookresearch/co-tracker", "cotracker3_online").to(device)

# Run Online CoTracker, the same model with a different API:
# Initialize online processing
cotracker(video_chunk=video, is_first_step=True, grid_size=grid_size)

# Process the video
for ind in range(0, video.shape[1] - cotracker.step, cotracker.step):
    pred_tracks, pred_visibility = cotracker(
        video_chunk=video[:, ind : ind + cotracker.step * 2]
    )  # B T N 2,  B T N 1
```

The README cautions that the above example requires the full video length to be known in advance for the loop bounds. For video from an unknown-length live stream, the `online_demo.py` script in the repository demonstrates the pattern for tracking from an online stream.

Online and offline modes use the same underlying model weights. The difference is in the inference API: online mode processes overlapping temporal windows and merges them, while offline mode processes the full sequence at once. Offline mode typically produces slightly more accurate results on short videos where the full context is available from the start.

## Installing CoTracker3 locally for evaluation and training

For running local demos, evaluating on benchmarks, or training the model, the repository must be cloned and installed. The README assumes PyTorch and TorchVision are already installed with CUDA support:

```bash
git clone https://github.com/facebookresearch/co-tracker
cd co-tracker
pip install -e .
pip install matplotlib flow_vis tqdm tenso
```

The `setup.py` lists `cotracker` as version 3.0 with no required install dependencies, relying on the environment having PyTorch available separately. After installation, offline and online demos can be run directly:

```bash
python demo.py --grid_size 10
```

and for the online path:

```bash
python online_demo.py
```

The training scripts at the repository root, including `train_on_kubric.py` and `train_on_real_data.py`, require the corresponding datasets. The Kubric dataset used for CoTracker3, containing 6000 high-resolution sequences of 512x512 pixels at 120 frames each, is available on Hugging Face as `facebook/CoTracker3_Kubric`.

For visualization, the `Visualizer` class saves tracked points overlaid on the original video.

## Architecture and the pseudo-labeling training approach

CoTracker3's title is "Simpler and Better Point Tracking by Pseudo-Labelling Real Videos." The key architectural insight described in the paper is that training on real video with pseudo-labels, generated by running a teacher model on unlabeled video and using its predictions as supervision, achieves state-of-the-art performance with roughly 1000 times less labeled data than previous approaches.

Previous top-performing point tracking models were trained primarily on synthetic data because real video with reliable ground-truth point annotations is expensive to generate. CoTracker3 sidesteps that requirement by generating pseudo-labels on real videos, which gives the model exposure to the distribution of real-world motion, texture, and occlusion patterns that synthetic datasets do not fully capture.

The model processes groups of points jointly rather than tracking each point independently. The transformer attention mechanism allows each tracked point to use information from neighboring points in the group when resolving ambiguous positions. This joint processing is what the original CoTracker paper described in its title as "better to track together."

CoTracker2, the previous version, could track up to 265x265 points jointly with a more memory-efficient implementation. CoTracker3 is described as having a lightweight architecture, with the specific parameter count and memory footprint detailed in the ArXiv papers linked from the README.

## Limitations and maintenance status

CoTracker3 requires a GPU for practical use. The README states this recommendation explicitly. CPU inference is technically possible for small tasks but is not a supported path for production use.

The model tracks points in videos where it has been given coordinates to track or a grid to sample. It does not automatically identify which points in a scene are interesting to track. If you need to track specific objects rather than arbitrary pixels, you must first detect and localize those objects with a separate model and then pass the relevant coordinates to CoTracker.

Code readability is also a constraint. The retainer-based tracking approach relies on being able to see what was tracked. For very fast motion, severe occlusion, or scenes where the tracked point disappears for many frames, the visibility output will reflect the uncertainty.

The last push to this repository was on 2026-03-03, which is more than six months before 2026-09-28. The repository is not archived. For research and experimentation, the existing model checkpoints and code remain usable regardless of recent activity. Building a production dependency on the repository while it may be inactive carries maintenance risk.

PIPS (Persistent Independent Particles) is a related approach to long-range point tracking. It takes a different path by tracking each point independently without joint inference, which makes it simpler to analyze but less accurate in cases where joint reasoning over nearby points would help.

## Conclusion

CoTracker3 is the right tool for computer vision researchers and practitioners who need accurate point tracking in video and are comfortable working in a PyTorch environment with GPU access. It is not suitable for real-time embedded deployment: a GPU is strongly recommended and the README explicitly states so. Before building a pipeline around it, verify that your video source can be loaded into the expected tensor format (B T C H W) and that your target environment supports CUDA. The last push to this repository was on 2026-03-03, which is more than six months before 2026-09-28. Active maintenance may have resumed on a private branch or a separate project, but check the repository status before depending on it for ongoing work.

## FAQ

### What is the difference between CoTracker3's offline and online modes?

Offline mode processes the entire video at once and produces tracks for all frames in one pass. Online mode processes the video in overlapping temporal chunks, making it more memory-efficient for long videos. Both modes use the same model weights. The README notes that offline mode is slightly more accurate on short videos where full context is available from the start.

### Does CoTracker3 require a GPU to run?

A GPU is strongly recommended, as the README states explicitly. CPU inference is technically possible for small tasks but the README does not present it as a supported deployment path. The quick-start examples load the model with `.to(device)` where device is set to `cuda`.

### What input format does CoTracker3 accept for video?

CoTracker3 expects a PyTorch tensor with shape B T C H W: batch size, number of frames, channels, height, and width. The quick-start example uses imageio with the FFMPEG plugin to load a video file into frames and converts them to the required tensor format with `torch.tensor(frames).permute(0, 3, 1, 2)[None].float()`.

## Sources

- [facebookresearch/co-tracker on GitHub](https://github.com/facebookresearch/co-tracker)
- [Issues](https://github.com/facebookresearch/co-tracker/issues)
- [Project website](https://co-tracker.github.io/)
- [README](https://github.com/facebookresearch/co-tracker/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/facebookresearch-co-tracker
