# tapnet: Google DeepMind's Tracking Any Point Models and Benchmarks

> tapnet is the official Google DeepMind repository for the Tracking Any Point (TAP) family of video models, including TAPIR, BootsTAPIR, TAPNext, and TAPNext++, along with the TAP-Vid and TAPVid-3D evaluation benchmarks and pre-trained weights in JAX and PyTorch.

**google-deepmind/tapnet** — Tracking Any Point (TAP)

- Repository: https://github.com/google-deepmind/tapnet
- Website: https://deepmind-tapir.github.io/blogpost.html
- Stars: 1,998 · Forks: 193
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/google-deepmind-tapnet

## What the Tracking Any Point Task Involves

Tracking Any Point addresses a specific video understanding task: given a query point in one frame of a video, locate that same physical point in every other frame, even when the point leaves and re-enters view. This differs from optical flow, which estimates dense motion between adjacent frames, and from object tracking, which follows bounding boxes around detected objects. TAP queries a single pixel coordinate and maintains its identity throughout the clip.

The application domains described in the README include robotics manipulation, where TAPIR point tracks drive imitation learning from demonstration videos in the RoboTAP extension. Evaluation of generative video models is another use case: the TRAJAN trajectory autoencoder can measure whether generated videos produce motion consistent with real-world physics.

The repository centralises the output of several related research projects: TAP-Vid, TAPIR, BootsTAP, RoboTAP, TAPVid-3D, TAPNext, TRAJAN, and TAPNext++. Each has its own paper and project page, and the repository provides the code, benchmarks, and checkpoints that accompany them.

## The TAPIR Algorithm and Its Successors

TAPIR is a two-stage algorithm. The first stage, matching, independently locates a candidate match for the query point in every other frame of the video. The second stage, refinement, updates both the trajectory estimate and the query features based on local spatial correlations around each candidate. The README states that TAPIR surpasses prior methods by a significant margin on the TAP-Vid benchmark.

BootsTAPNEXT (BootsTAPIR) improves on TAPIR using unlabelled real-world video. The training objective enforces consistency across spatial transformations, corruptions, and different choices of query points, rather than requiring ground-truth point annotations. The result, BootsTAPIR, shares TAPIR's architecture but outperforms it on TAP-Vid.

TAPNext reformulates the entire tracking problem as next-token prediction, propagating information about query point locations through a neural network one frame at a time. This allows online tracking rather than requiring the full clip upfront. TAPNext++, the most recent checkpoint in the repository, adds fine-tuning on 1024-frame synthetic sequences and achieves 40x longer stable tracking than the base TAPNext, along with the ability to track through occlusions and re-detect points that disappear and return.

## Running the Models: Colabs and Local Setup

The README identifies colab demos as the simplest path to running the models. Six colabs are available: TAPNext++ for long-term tracking with occlusion and re-detection, BootsTAPNext for the most capable per-frame online model, BootsTAPNext in PyTorch, standard TAPIR and BootsTAPIR for whole-video offline tracking, online causal TAPIR for real-time GPU tracking, and a rainbow visualisation colab that applies foreground/background segmentation and camera-motion correction before displaying point trajectories as colour trails. Each colab accepts an uploaded video for interactive experimentation.

For local execution, the README instructs cloning the repository and running a real-time demo on personal hardware. The base dependencies in requirements.txt include jax, jaxline, dm-haiku, optax, chex, mediapy, opencv-python, einshape, and einops. For training rather than inference, the pyproject.toml defines a train extra group that adds TensorFlow, TensorFlow Datasets, TensorFlow Graphics, Kubric, and RecurrentGemma as additional dependencies. A torch extra group adds torch and torchvision for the PyTorch-backed models.

The README does not specify minimum GPU memory requirements. Model sizes across the different checkpoints range from tens of megabytes to multi-gigabyte weights depending on the model variant. The colabs/ directory in the repository contains all six demo notebooks, each targeting a specific model or visualisation style.

## The TAP-Vid and TAPVid-3D Benchmarks

TAP-Vid is an evaluation benchmark for point tracking models. It provides ground-truth point trajectories on both real-world videos and synthetic ones, enabling standardised comparison between tracking methods. TAPIR's original results were reported on TAP-Vid, establishing it as the reference benchmark for the task.

TAPVid-3D extends the benchmark to three-dimensional point tracking. The README states it contains over one million computed ground-truth trajectories on more than 4,000 real-world videos, with a new set of metrics for the 3D tracking task. Sample evaluation code and generation scripts are included in the repository.

RoboTAP is a third evaluation dataset focused on robotics manipulation videos. It provides ground-truth points annotated on real manipulation sequences and includes clustering code that uses point tracks to segment manipulation tasks. The README positions RoboTAP as both a benchmark and a system that demonstrates how point tracking can make robot imitation learning from video demonstrations more sample-efficient, by identifying and reusing motion primitives across multiple demonstration clips.

## TRAJAN: The Trajectory Autoencoder

TRAJAN is a distinct model in the repository that takes a different approach to video understanding. Rather than tracking individual points forward through a clip, it encodes a set of support trajectories into a compact embedding and uses that embedding to reconstruct trajectories for a held-out set of query points.

The embedding space learned by TRAJAN can be used for three purposes described in the README: comparing distributions of videos by comparing their trajectory embeddings, comparing motion across different videos independently of object appearance, and evaluating the realism and consistency of videos produced by generative video models. The last use case positions TRAJAN as a quality metric for video generation, providing a way to check whether a generated video's motion patterns match those of real-world reference videos.

TRAJAN has its own colab demo in the repository. It is separate from the TAPIR and TAPNext tracking models and is designed for embedding and comparison tasks rather than frame-by-frame tracking.

## Limitations and What the Repository Does Not Cover

The repository contains no GitHub releases. Checkpoints are distributed separately for TAP-Net, TAPIR, and BootsTAPIR in both JAX and PyTorch. The README does not give precise checkpoint URLs inline; the checkpoints section links to separate instructions. Training TAPIR or TAP-Net requires the Kubric synthetic data generation pipeline, which is maintained in a separate Google Research repository and adds a non-trivial setup step before training can begin.

The primary interface is JAX, and the Haiku deep learning library is used throughout the model code. PyTorch implementations of TAPNext and BootsTAPNext are described as containing the exact architecture and weights as the JAX models, but the PyTorch path is a port rather than the original implementation. Developers working in pure TensorFlow have no direct inference path; TensorFlow appears only as a training-time dependency in the pyproject.toml optional group.

The six colab notebooks and the README provide the main user-facing documentation. Architecture and training details require reading the individual project papers linked from the README for each model variant. The TRAJAN and TAPVid-3D tools each have their own project pages with separate documentation. The last push to the repository was on 2026-09-15, indicating the project is under active maintenance. Teams integrating multiple model variants should track each checkpoint's update history separately, as not all variants are updated in each repository push.

## Conclusion

tapnet is the right starting point for researchers and engineers who need to track specific points through video, whether for robotics, video analysis, or generative model evaluation. The colab demos make the models accessible without any local setup. Running the real-time demo on your own hardware requires cloning the repository and installing the JAX or PyTorch dependencies, which are non-trivial. The TAPVid-3D benchmark and TRAJAN trajectory autoencoder are specialised tools that require reading the individual project pages before committing to an integration. Verify that your target platform supports JAX or PyTorch and has sufficient GPU memory for the checkpoint size you intend to use.

## FAQ

### What is tapnet used for?

Google DeepMind's tapnet repository provides models and benchmarks for the Tracking Any Point task: given a query point in one video frame, the models locate that point in every other frame, including through occlusions. Applications described in the README include robotics imitation learning, video motion analysis, and evaluating the realism of generated videos.

### How does BootsTAPIR differ from TAPIR?

BootsTAPIR is trained using unlabelled real-world video with a consistency objective, rather than relying solely on ground-truth point annotations. The README states that BootsTAPIR is architecturally similar to TAPIR but substantially outperforms it on the TAP-Vid benchmark.

### Do I need a GPU to run the tapnet models?

The README does not specify a GPU requirement for the colab demos, which run on Google Colab's provided hardware. For the local real-time demo, the README describes running on your own hardware without specifying minimum GPU memory. Model sizes in the repository range from small checkpoints to multi-gigabyte weights, so hardware requirements vary by model.

## Sources

- [google-deepmind/tapnet on GitHub](https://github.com/google-deepmind/tapnet)
- [Issues](https://github.com/google-deepmind/tapnet/issues)
- [License: Apache-2.0](https://github.com/google-deepmind/tapnet/blob/main/LICENSE)
- [Project website](https://deepmind-tapir.github.io/blogpost.html)
- [README](https://github.com/google-deepmind/tapnet/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/google-deepmind-tapnet
