VGGT: feed-forward 3D reconstruction from one image or hundreds
[CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
At a glance
- What is it?
- VGGT is a Python model from Meta AI and Oxford's Visual Geometry Group that regresses cameras, depth, point maps and tracks in a single forward pass. It is a research release, not a pipeline, and the licence splits the original checkpoint from the commercial one.
- Who is it for?
- Adopt VGGT if you need camera poses and dense geometry from a small unordered image set and you can accept a research checkpoint plus a separate licence application for commercial work. Do not adopt it if you need a maintained product with a stable API, a documented rollback path, or a permissively licensed model out of the box, since the original checkpoint stays non-commercial.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 134 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What VGGT replaces in a 3D reconstruction pipeline
Classical structure-from-motion is a sequence of stages. You detect keypoints, match them across views, estimate an initial camera pair, triangulate, then run bundle adjustment to make the whole thing consistent. Each stage has its own failure mode, and a bad match at the front contaminates everything behind it. VGGT collapses that chain into one feed-forward network. The README describes it as a network that "directly infers all key 3D attributes of a scene, including extrinsic and intrinsic camera parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views, within seconds." The audience is researchers and engineers who need geometry from images without tuning a matching pipeline: multi-view reconstruction, novel view synthesis as an input stage, and robotics or vision-language work that needs camera poses as a side product. The single-image case is the interesting one, because there is no correspondence to match at all.
One aggregator, several prediction heads
The architecture in the code is a shared trunk with task-specific heads. model.aggregator(images) returns aggregated_tokens_list and ps_idx. Those tokens are then consumed by model.camera_head, model.depth_head, and the point and track heads. The camera head emits a pose encoding that pose_encoding_to_extri_intri converts into extrinsic and intrinsic matrices following the OpenCV convention, meaning camera-from-world. Depth maps come with a confidence map, and unproject_depth_map_to_point_map turns depth into a point map using the predicted intrinsics. That data flow matters for two reasons. First, the heads share one representation, so the tasks are not independent estimates bolted together. Second, you can call the aggregator once and run only the heads you need, which is what the detailed usage example in the README does. The README does not document the token dimensionality or how ps_idx should be used downstream, so anyone writing a custom head is reading the source rather than the docs.
Installing VGGT and running the first prediction
The README gives a clone-and-pip path. Dependencies listed there are torch, torchvision, numpy, Pillow and huggingface_hub; requirements.txt pins torch==2.3.1, torchvision==0.18.1 and numpy==1.26.1. The numpy pin is below 2 in pyproject.toml as well, so a global numpy 2 install will conflict.
git clone [email protected]:facebookresearch/vggt.git
cd vggt
pip install -r requirements.txtThe README also notes you can install VGGT as a package, pointing to docs/package.md for the details. The package name is vggt and requires-python is ">= 3.10".
A first run loads weights from Hugging Face and predicts on a list of image paths. The dtype choice in the README is conditional: bfloat16 on Ampere (compute capability 8.0) and above, float16 otherwise.
import torch
from vggt.models.vggt import VGGT
from vggt.utils.load_fn import load_and_preprocess_images
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if torch.cuda.get_device_capability()[0] >= 8 else torch.float16
model = VGGT.from_pretrained("facebook/VGGT-1B").to(device)
image_names = ["path/to/imageA.png", "path/to/imageB.png", "path/to/imageC.png"]
images = load_and_preprocess_images(image_names).to(device)
with torch.no_grad():
with torch.cuda.amp.autocast(dtype=dtype):
predictions = model(images)The README warns that the weight download happens on first run and may take a while. If it stalls, the documented fallback is to download model.pt manually and load it with torch.hub.load_state_dict_from_url against the Hugging Face resolve URL. For a scene export, demo_colmap.py saves predictions in COLMAP format with optional bundle adjustment; the README states those files can be used directly with gsplat or other NeRF and Gaussian splatting libraries.
What VGGT does not do, and where it breaks
VGGT predicts. It does not refine. There is no documented iterative optimisation inside the model, and the README does not describe a rollback or recovery path when a prediction is wrong, because there is no state to roll back to: you get one forward pass and a tensor. If your scene needs survey-grade accuracy, the feed-forward estimate is a starting point that you would still hand to bundle adjustment, which is why the COLMAP export script exposes that as an option rather than a default. Memory is the other hard constraint. The May 15, 2026 update states that an implementation issue was "keeping redundant intermediate tensors in memory" and that with the same GPU memory budget VGGT can now run on roughly 2-3x more input frames. That is a fix to a real ceiling: frame count was previously bounded by memory in a way the model itself did not explain. Anyone running an older checkout inherits the old ceiling. Finally, the README does not document supported input resolutions, video handling beyond the examples/videos/ directory, or behaviour on degenerate image sets, so those are empirical questions for your own data.
VGGT versus COLMAP-style incremental SfM
The alternative most teams already have is COLMAP: feature extraction, matching, incremental reconstruction, bundle adjustment. The difference in approach is not accuracy, it is where the work happens. COLMAP optimises an explicit geometric model against measured correspondences, so its errors are traceable to specific matches and its output is a refined camera model. VGGT learns a mapping from images to geometry, so it produces poses and dense point maps in one pass without matching, and its errors are not decomposable into a stage you can inspect. That makes VGGT faster to stand up and easier to run on sparse or weakly textured sets where matching struggles, and it makes COLMAP the better choice when you need the refinement loop and the intermediate outputs. The two are not exclusive: demo_colmap.py exists precisely because the VGGT output is useful as COLMAP-format input to downstream tooling.
Licence: two checkpoints, two rules
This is the part most likely to block adoption, and the README is explicit about it. The July 29, 2025 update says the licence was changed to permit commercial use excluding military applications, and that all code in the repository is under a commercial-use-friendly licence. But the checkpoint is separate: only VGGT-1B-Commercial is licensed for commercial usage, while the original checkpoint remains non-commercial. The new checkpoint requires completing an application form, processed automatically by a system the README compares to LLaMA's approval workflow. The README states the new checkpoint delivers similar performance to the original and asks users to submit an issue if they see a significant discrepancy. So the code licence and the weights licence do not move together, and the repository's LICENSE.txt is the file to read rather than the README summary. This is a description of what the project states, not legal advice; if commercial deployment is the plan, the checkpoint you may load is a decision to make before writing code.
Maintenance, training and upgrade cost
The last push to the default branch was on 2026-05-19, four months before this writing, and the repository is not archived. The update log is dense through 2025 and into 2026: training code arrived on July 6, 2025 in the training folder with an example for finetuning on a custom dataset, Co3D evaluation code lives on an evaluation branch, and the May 2026 updates cover both the memory fix and the release of VGGT-Omega as the next step. There are no retrieved releases, so there is no tagged version to pin against; pyproject.toml still reads version = "0.0.1". That combination, active commits plus no release tags, means upgrading is a git pull and a re-read of the update list rather than a dependency bump. Budget for it: the memory fix alone changes how many frames fit on a given GPU, and the licence change moved a checkpoint. Anyone pinning to a commit from before May 2026 is running the older memory behaviour and the older licence terms.
Editorial conclusion
Adopt VGGT if you need camera poses and dense geometry from a small unordered image set and you can accept a research checkpoint plus a separate licence application for commercial work. Do not adopt it if you need a maintained product with a stable API, a documented rollback path, or a permissively licensed model out of the box, since the original checkpoint stays non-commercial. Before committing, verify three things: which checkpoint your use case is allowed to load, whether your GPU supports bfloat16 (Ampere or newer), and whether the COLMAP export from demo_colmap.py feeds the downstream library you already use.
Frequently asked questions
What is VGGT and what does it stand for?
VGGT stands for Visual Geometry Grounded Transformer. It is a feed-forward neural network from Meta AI and the University of Oxford's Visual Geometry Group that infers camera parameters, point maps, depth maps and 3D point tracks from one or many views of a scene.
Which VGGT checkpoint can I use commercially?
Only VGGT-1B-Commercial is licensed for commercial usage; the original checkpoint remains non-commercial. Access to the commercial checkpoint requires completing an application form, and the README says the repository code is under a commercial-use-friendly licence excluding military applications.
How many images can VGGT process at once?
The README does not give a fixed number. It states that a May 15, 2026 fix stopped redundant intermediate tensors being kept in memory, so with the same GPU memory budget VGGT can now run on roughly 2-3x more input frames than before.
Does VGGT need a GPU?
The README's example selects CUDA when it is available and falls back to CPU otherwise, and it chooses bfloat16 only on Ampere GPUs (compute capability 8.0+) and float16 below that. No CPU performance figures are given.
Can VGGT output be used with Gaussian splatting tools?
Yes, according to the README. The demo_colmap.py script saves predictions in COLMAP format with optional bundle adjustment, and the README states those files can be used directly with gsplat or other NeRF and Gaussian splatting libraries.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/facebookresearch-vggt)