Model or dataset
facebookresearch/vjepa2 avatar
facebookresearch/vjepa2

V-JEPA 2: What the facebookresearch/vjepa2 Repository Actually Ships

PyTorch code and models for VJEPA2 self-supervised learning from video.

4,692 stars584 forksPythonMIT

At a glance

What is it?
The official PyTorch codebase for V-JEPA 2, V-JEPA 2-AC and V-JEPA 2.1 is a research training stack, not a video understanding API. Here is what installs cleanly, what the checkpoints are for, and where the repository stops short.
Who is it for?
Adopt vjepa2 if you need to reproduce masked latent feature prediction training or fine-tune a frozen video encoder on your own labels; the ViT-L/16 checkpoint and the configs/train/vitl16 directory are the shortest path in. Do not adopt it if you want a packaged video question answering endpoint, since the repository exposes training and evaluation scaffolding rather than a serving layer.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 27, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What vjepa2 Solves That a Supervised Video Classifier Does Not

Labeled video is expensive. The vjepa2 repository exists to train video encoders without those labels, using internet-scale footage and a masked latent feature prediction objective. The encoder sees part of a clip, a predictor guesses the latent representation of the missing part, and the loss is computed in representation space rather than on pixels. That design choice is the whole point: predicting raw pixels rewards low-level texture, while predicting latents lets the model spend capacity on motion and physical structure. The README states that V-JEPA 2 attains state-of-the-art performance on motion understanding and human action anticipation, and the benchmark table lists EK100 at 39.7 percent against a previous best of 27.6 percent for PlausiVL, with SSv2 probing at 77.3 percent. Those are the numbers the project reports for itself, not independent reproductions. The audience is narrow and identifiable: research engineers who want a frozen or fine-tuned video backbone, and robotics groups interested in the action-conditioned variant. It is not aimed at application developers who want to drop a video into a function and get a caption out.

The Encoder, Predictor and World Model Split Inside the Codebase

V-JEPA 2 pre-training is a two-network arrangement. An encoder turns visible tokens into latent features, and a predictor is trained to produce the latents of masked tokens from the visible context. Both are trained jointly through self-supervision, and the encoder is what you keep. V-JEPA 2-AC changes the conditioning: it is post-trained from V-JEPA 2 using a small amount of robot trajectory interaction data, and the README describes it as a latent action-conditioned world model that solves manipulation tasks without environment-specific data collection, task-specific training or calibration. Planning happens against image goals on a Franka arm with a single monocular RGB camera. V-JEPA 2.1, released 2026-03-16, revises the recipe rather than the architecture family. The README names four changes: a Dense Predictive Loss where both visible and masked tokens contribute to the loss, Deep Self-Supervision that applies the loss at multiple intermediate encoder representations, multi-modal tokenizers for images and videos, and scaling of model and data. The repository layout reflects this: configs/train for the original runs, configs/train_2_1 for the new recipe, evals for downstream evaluation, and src for the library itself.

Installing vjepa2 and Running a First Forward Pass

setup.py declares the package name vjepa2, version 0.0.2, and python_requires of at least 3.11, with dependencies read straight from requirements.txt. That file pulls torch>=2, torchvision, decord, webdataset, timm, transformers, peft, submitit, wandb, tensorboard, fire, python-box and a long tail of scientific Python. Two of those deserve attention before you start: decord is a video reader with compiled components, and webdataset shapes how training data is streamed. Neither is optional in a training run. A source install from the repository root looks like this.

Model Sizes, Resolutions and Which Checkpoint to Pick

The models table splits into two families. V-JEPA 2 offers ViT-L/16 at 300M parameters and resolution 256, ViT-H/16 at 600M and 256, ViT-g/16 at 1B and 256, and a ViT-g/16 variant at 1B and resolution 384. V-JEPA 2.1 starts with a ViT-B/16 at 80M parameters and resolution 384, checkpoint vjepa2_1_vitb_dist_vitG_384.pt, configs under configs/train_2_1/vitb16, and continues into a ViT-L/16 entry. The README excerpt cuts off mid-row in that second table, so the full V-JEPA 2.1 lineup is not something this article can state. The practical reading: resolution is a separate axis from parameter count, and the 384-resolution models are not simply bigger versions of the 256 ones. If you are memory-constrained, the 80M ViT-B/16 is the only small option listed, and it is a V-JEPA 2.1 model, so you inherit the dense-feature recipe along with it. If you need the configuration that produced a specific published number, match the checkpoint to its config directory rather than mixing a 2.1 checkpoint with a configs/train entry.

Where vjepa2 Is the Wrong Tool

There is no serving layer here. The repository has no inference server, no REST endpoint, and no batched video ingestion pipeline; the top-level entries are training and evaluation oriented (configs, evals, src, app, notebooks, tests). If your requirement is a hosted video QA API, this is the wrong dependency and you will spend your time building the wrapper the project never intended to ship. The dependency list is also a real constraint. decord, webdataset and opencv-python all carry native build steps, and a training environment that satisfies them is heavier than a pure-Python install. The README gives no rollback instructions, no supported-platform matrix, and no guidance on what to do when a checkpoint download fails, so operational recovery is on you. Finally, the V-JEPA 2.1 paper link in the README points at arxiv.org/abs/TODO, which means the recipe is described in the repository text and configs but not yet backed by a citable paper. If your review process requires a published methods section, that gap matters. The V-JEPA 2 paper at arxiv.org/abs/2506.09985 is the citable one, and it covers the earlier model, not 2.1.

V-JEPA 2 Against a Contrastive Video Encoder

The natural comparison is a contrastive video model in the VideoMAE or InternVideo2 lineage, and the difference is in the objective. Contrastive and reconstruction methods learn by pulling augmented views together or by rebuilding pixels. V-JEPA 2 predicts masked latents, which the README frames as learning physical world structure rather than appearance. The benchmark table makes the practical consequence visible: V-JEPA 2 reports 90.2 percent on Diving48 probing against 86.4 percent for InternVideo2-1B, and 77.3 percent on SSv2 probing against 69.7 percent. On video question answering the margins narrow sharply, with MVP at 44.5 percent against 39.9 percent for InternVL-2.5 and TempCompass at 76.9 percent against 75.3 percent for Tarsier 2. That pattern is worth reading carefully. The advantage concentrates in motion and temporal tasks, and almost disappears where a language model is doing the heavy lifting. If your downstream task is captioning or open-ended QA, a multimodal model with a language head is the more direct answer. V-JEPA 2-AC is a different kind of alternative too: against Octo and Cosmos on Franka manipulation, the README table shows VJEPA 2-AC at 60 percent on cup grasping and 80 percent on cup pick-and-place, where Octo reports 10 percent on both and Cosmos reports 0 percent on cup grasping.

Licence, Maintenance and the Cost of Upgrading to 2.1

The repository is MIT licensed, and setup.py carries the Meta Platforms copyright header pointing at the LICENSE file. MIT is permissive, so redistribution and commercial use are not restricted by the licence text itself, but nothing here is legal advice: check the licence file and the terms attached to each checkpoint download separately, because a permissive code licence does not automatically settle model weights. On maintenance, the last push to the default branch was 2026-03-23, and the repository is not archived. The README announces V-JEPA 2.1 on 2026-03-16, so the most recent activity lines up with that release. There are no retrieved releases, which means versioning runs through the repository rather than tagged artifacts; setup.py pins the package at 0.0.2. Upgrading from V-JEPA 2 to 2.1 is not a drop-in swap. Configs move from configs/train to configs/train_2_1, the recipe changes at the loss level with dense prediction and deep self-supervision, and the smallest model in the 2.1 table is 80M where the original family starts at 300M. Budget for a re-evaluation rather than a version bump. A CHANGELOG.md exists at the top level, so that file is the place to check what moved between the two recipes.

Editorial conclusion

Adopt vjepa2 if you need to reproduce masked latent feature prediction training or fine-tune a frozen video encoder on your own labels; the ViT-L/16 checkpoint and the configs/train/vitl16 directory are the shortest path in. Do not adopt it if you want a packaged video question answering endpoint, since the repository exposes training and evaluation scaffolding rather than a serving layer. Before committing, verify that your Python is at least 3.11 as setup.py requires, that the decord and webdataset dependencies build on your platform, and that the checkpoint you intend to use matches the resolution listed in the models table.

Frequently asked questions

What is V-JEPA 2?

It is a self-supervised approach to training video encoders on internet-scale video, described in the README as attaining state-of-the-art performance on motion understanding and human action anticipation. The facebookresearch/vjepa2 repository is the official PyTorch codebase for it, along with V-JEPA 2-AC and V-JEPA 2.1.

What is the latest JEPA model in the vjepa2 repository?

The README announces V-JEPA 2.1 on 2026-03-16, describing it as a new family trained with a recipe for high quality and temporally consistent dense features. Its changes are a Dense Predictive Loss, Deep Self-Supervision at multiple intermediate representations, multi-modal tokenizers, and model and data scaling.

What is the JEPA model and how does it work?

In this repository the encoder and predictor are pre-trained by self-supervised learning from video using a masked latent feature prediction objective, so the model predicts representations of masked content rather than raw pixels. V-JEPA 2-AC then post-trains that model with a small amount of robot trajectory data into a latent action-conditioned world model.

Official sources

  1. facebookresearch/vjepa2 on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/facebookresearch-vjepa2.svg)](https://hysenlabs.com/projects/facebookresearch-vjepa2)