# VGGT-Omega: A CVPR 2026 Oral Feed-Forward Model for Camera and Depth Reconstruction from Multiple Images

> VGGT-Omega is a feed-forward neural network model from the Visual Geometry Group at Oxford and Meta AI that predicts camera poses (extrinsics and intrinsics) and depth maps from an unordered set of input images in a single forward pass. It was presented as an Oral at CVPR 2026 and is designed for computer vision researchers and practitioners who need to recover 3D scene structure without iterative optimization.

**facebookresearch/vggt-omega** — [CVPR 2026 Oral] VGGT Omega

- Repository: https://github.com/facebookresearch/vggt-omega
- Stars: 4,601 · Forks: 360
- Language: Python
- License: NOASSERTION
- Published: 2026-09-17 · Updated: 2026-09-17 · Language: en
- Canonical page: https://hysenlabs.com/projects/facebookresearch-vggt-omega

## What VGGT-Omega Is and What It Predicts

VGGT-Omega is a 1-billion-parameter feed-forward model that takes an unordered set of images as input and predicts camera extrinsics, intrinsics, and depth maps in a single forward pass. The model does not require an initial pose estimate or iterative refinement: the predictions emerge directly from the forward computation.

The work is from the Visual Geometry Group at the University of Oxford (affiliation 1 in the README) and Meta AI (affiliation 2). The arXiv preprint is at 2605.15195. The README lists the authors as Jianyuan Wang, Minghao Chen, Shangzhan Zhang, Nikita Karaev, Johannes Schonberger, Patrick Labatut, Piotr Bojanowski, David Novotny, Andrea Vedaldi, and Christian Rupprecht.

The model was accepted as an Oral presentation at CVPR 2026, which is the top computer vision conference. The project page is at vggt-omega.github.io and a public interactive demo is available on Hugging Face Spaces at huggingface.co/spaces/facebook/vggt-omega.

## Three Checkpoints: In-the-Wild, Benchmarking, and Text Alignment

The README lists three pretrained model checkpoints, each suited to a different use case.

VGGT-Omega-1B-512 uses 512 as its image resolution, was published in May 2026, and is described as 'Recommended for in-the-wild applications, no benchmarking.' This is the checkpoint to use when running the model on real-world images outside a controlled evaluation setting.

VGGT-Omega-1B-416-Reproduction uses 416 as its image resolution, was published in September 2026, and is described as 'Retrained for benchmarking, as detailed in reproduction.md.' The README explains this checkpoint was released to address a potential concern with the original checkpoint and should serve as the reference for future comparisons on reported benchmarks.

VGGT-Omega-1B-256-Text-Alignment uses 256 as its image resolution and adds text alignment. It is initialized with VGGTOmega(enable_alignment=True) and exposes a predictions['text_alignment_embedding'] field that is not present in the other two checkpoints.

Access to all checkpoints requires a request on Hugging Face (at huggingface.co/facebook/VGGT-Omega). The README states access requests are reviewed by an automated process and the authors are not involved in approving or rejecting individual applications.

## Installation and Running Inference

VGGT-Omega requires Python 3.10 or later and PyTorch 2.6 or later. To install:

```bash
git clone git@github.com:facebookresearch/vggt-omega.git
cd vggt-omega
pip install -r requirements.txt
pip install -e .
```

The requirements.txt specifies torch>=2.6, torchvision>=0.21, numpy<2, Pillow, einops, safetensors, and opencv-python. Once installed and with a checkpoint downloaded, inference looks like:

```python
import torch
from vggt_omega.models import VGGTOmega
from vggt_omega.utils.load_fn import load_and_preprocess_images

checkpoint_path = "path/to/vggt_omega_1b_512.pt"
image_names = ["path/to/imageA.png", "path/to/imageB.png", "path/to/imageC.png"]

model = VGGTOmega().to("cuda").eval()
model.load_state_dict(torch.load(checkpoint_path, map_location="cpu"))

images = load_and_preprocess_images(image_names, image_resolution=512).to("cuda")

with torch.inference_mode():
    predictions = model(images)
```

The predictions dictionary contains the camera extrinsics and intrinsics, plus depth information. For the text-aligned checkpoint, pass VGGTOmega(enable_alignment=True) and use image_resolution=256.

The load_and_preprocess_images function accepts a mode parameter. The default mode='balanced' produces 624x416 inputs for roughly 3:2 landscape images at resolution=512. Setting mode='max_size' resizes the longest side to 512 instead, giving approximately 512x336 inputs and using less GPU memory.

## The Gradio Demo for Local Interactive Exploration

VGGT-Omega includes a local Gradio-based interactive demo. It requires additional dependencies installed via a separate requirements file:

```bash
pip install -r requirements_demo.txt
```

The demo requirements include gradio>=5.17.1, viser>=0.2.23, tqdm, scipy, trimesh, matplotlib, requests, and onnxruntime. To launch the demo:

```bash
python demo_gradio.py \
  --checkpoint checkpoints/VGGT-Omega-1B-512/model.pt \
  --image-resolution 512
```

The demo accepts either uploaded images or a video as input. It runs camera and depth inference, then visualizes the depth-unprojected point cloud and predicted cameras as a GLB scene. The GLB output can be downloaded and opened in any 3D viewer.

The repository also includes four video examples under examples/: desert_road.mp4, forest_road.mp4, lake_speedboat.mp4, and snow_lift.mp4. These can be used to test the model locally before using your own footage.

## GPU Memory Scaling from 1 to 500 Input Frames

The README provides a GPU memory benchmark for VGGT-Omega-1B-512 on a single NVIDIA A100 GPU with 624x416 input images. The measurement covers the full inference program including model weights and inference buffers.

At 1 input frame, peak memory is 6.02 GB. At 10 frames: 6.67 GB. At 25 frames: 7.80 GB. At 50 frames: 9.66 GB. At 100 frames: 13.37 GB. At 200 frames: 20.82 GB. At 300 frames: 28.26 GB. At 400 frames: 35.71 GB. At 500 frames: 43.15 GB.

This scaling is important for planning hardware requirements. A GPU with 24 GB of VRAM can handle approximately 200 input frames at the default resolution setting. Using mode='max_size' instead of mode='balanced' reduces input resolution and therefore reduces memory consumption at a given frame count.

The A100 benchmark in the README represents an upper-tier research GPU. Practitioners running on consumer GPUs with 16 or 24 GB will be constrained to shorter sequences at this resolution.

## Training Code and Dataset Preparation

The training code and training infrastructure were released progressively in September 2026. On September 8, 2026, the training code was released alongside a reproduction checkpoint. On September 9, 2026, guides and tools for preparing training data were added, covering dataset collection, conversion, cleaning, and agent-assisted visual review. A supervised geometric filtering pipeline was also added.

On September 10, 2026, the sequence lists for eight datasets used in training were released. The reannotated UCo3D data used to train VGGT-Omega is being uploaded to Hugging Face at huggingface.co/datasets/facebook/uco3d/tree/main/vggt_omega_anno.

The training optional dependencies are listed in pyproject.toml: torch>=2.6, hydra-core>=1.3, omegaconf>=2.3, fvcore, tensorboard, and wcmatch. The curation tools require catboost, joblib, lightgbm, scikit-learn>=1.4, scipy, and xgboost. Export to COLMAP format requires pycolmap>=3.10.0.

The reproduction.md file in the repository documents the steps to reproduce the benchmarking checkpoint, which is relevant for researchers who need to verify claims against the reported numbers.

## Limitations: Access-Gated Weights, GPU Requirements, and a Non-Standard License

The primary operational limitation is that the model weights require an access request on Hugging Face. The public demo runs without a downloaded checkpoint, but local inference does not. The automated review process means you cannot start using the model immediately; you must wait for approval.

The memory requirements rule out inference on laptops or systems without a discrete GPU with substantial VRAM. The minimum useful configuration is a GPU with at least 6 GB of free memory for single-image inference, and practical multi-image use requires 8 to 15 GB or more.

The license is marked as NOASSERTION in the GitHub metadata. The README directs users to the LICENSE file in the repository for the actual terms. This is not a standard MIT or Apache license, and anyone planning to incorporate the model into a commercial product or redistribute it must read the LICENSE file carefully before doing so.

The repository has no GitHub releases. The model is at version 0.0.1 per pyproject.toml, and updates are pushed directly to the main branch.

## Alternative: DUSt3R

DUSt3R is another feed-forward multi-view 3D reconstruction model, from Naver Labs Europe and released at ECCV 2024. Like VGGT-Omega, it takes unordered images as input and predicts 3D structure in a single forward pass without iterative optimization.

The difference in approach is in what each model directly predicts. DUSt3R predicts point maps (3D point positions for each pixel in a reference frame), from which camera poses are recovered as a post-processing step. VGGT-Omega directly predicts camera extrinsics, intrinsics, and depth maps as structured outputs from the model forward pass. This makes VGGT-Omega's camera predictions more direct, while DUSt3R's point map output is more flexible for tasks that do not require explicit camera parameters.

Both models target researchers and practitioners in multi-view geometry and structure-from-motion. The choice between them depends on whether your downstream task requires explicit camera parameters (where VGGT-Omega is more direct) or dense point cloud reconstruction (where DUSt3R's output format may be more convenient).

## Conclusion

VGGT-Omega is the appropriate tool for researchers and practitioners who need feed-forward camera pose and depth estimation from multiple unordered images without running an iterative optimization pipeline. The three checkpoints serve different purposes: VGGT-Omega-1B-512 for in-the-wild use, the 1B-416-Reproduction checkpoint for benchmarking, and the 256-text-aligned checkpoint for text-conditioned tasks. Before using it, request checkpoint access on Hugging Face (reviewed by an automated process), and verify that your hardware meets the minimum GPU memory requirements: 6.02 GB for a single image, scaling to 43.15 GB for 500 frames on an A100. The license is non-standard; check the LICENSE file before distributing any derived work.

## FAQ

### How do I download the VGGT-Omega checkpoints?

Checkpoint access requires a request at huggingface.co/facebook/VGGT-Omega. The README states that access is reviewed by an automated process and the authors are not involved. Once approved, the three checkpoints are available for download. The Hugging Face demo at huggingface.co/spaces/facebook/vggt-omega is publicly accessible without a checkpoint request.

### What GPU memory does VGGT-Omega require?

The README benchmarks peak GPU memory on an A100 with 624x416 inputs: 6.02 GB for 1 frame, 9.66 GB for 50 frames, 13.37 GB for 100 frames, and 43.15 GB for 500 frames. Using mode='max_size' instead of the default mode='balanced' reduces resolution and lowers memory usage.

### What is the difference between VGGT-Omega and the original VGGT?

VGGT-Omega is a successor to the original VGGT model, with a separate paper (arXiv 2605.15195) and separate checkpoints. The README does not document the architectural differences between the two; for a technical comparison, the arXiv paper is the primary reference.

## Sources

- [facebookresearch/vggt-omega on GitHub](https://github.com/facebookresearch/vggt-omega)
- [Issues](https://github.com/facebookresearch/vggt-omega/issues)
- [README](https://github.com/facebookresearch/vggt-omega/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/facebookresearch-vggt-omega
