Library / SDK
cvg/vidmap avatar
cvg/vidmap

VidMap asks you to build COLMAP twice, and PyTorch is not in its dependency list

VidMap: Exploiting Temporal Structure for Video-Based Structure-from-Motion (ECCV 2026)

464 stars36 forksPythonApache-2.0

At a glance

What is it?
An offline structure-from-motion system for video, built as an extension to COLMAP's global mapper with Hydra-configured hyperparameters and Hydra-generated config dumps. The pipeline design is coherent and the configuration story is good. The installation is the hard part: a from-source COLMAP build, a second CMake build inside the wheel, three git-pinned dependencies, and a first run that downloads roughly ten gigabytes.
Who is it for?
Use VidMap if you already run COLMAP from source, have an NVIDIA GPU on Linux x86-64, and want a video SfM front end that hands mapper inputs to a pipeline you can tune with typed Hydra configs. Do not plan around the pip install alone, because PyTorch and xFormers are absent from the dependency list and the first frontend run pulls about 10 GB with no pre-staging flag.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

COLMAP gets built twice, by hand and then by the wheel

The setup instructions ask for a from-source COLMAP 4.2 build with PyCOLMAP bindings, pointing at the COLMAP build documentation and noting that GPU mapper acceleration needs Ceres 2.3 or newer with CUDA and cuDSS.

Then the package's own build system does the same thing again. The manifest uses scikit-build-core with a CMake source directory pointing into the repository:

toml
[tool.scikit-build]
cmake.version = ">=3.28"
cmake.source-dir = "extensions/colmap"
cmake.build-type = "Release"
cmake.define.BUILD_TESTING = "OFF"
build.tool-args = ["-j8"]

So there is a COLMAP extension compiled by the Python build, with pybind11 generating the bindings, sitting alongside the COLMAP you were told to build and install first. A C++ extension and a separately installed COLMAP are not the same artifact, and both are needed.

The build also hardcodes its parallelism. build.tool-args sets -j8 for every machine, with no manifest-level way to change it. On a four-core box that oversubscribes; on a large build host it leaves cores idle. Anyone who hits a slow or memory-hungry compile will not find the knob in pyproject.toml.

BUILD_TESTING is switched off, and the repository has no test directory at the top level, so the C++ side of this project ships with its own tests disabled and no Python test suite beside it.

The clone needs submodules, which is why the one-liner carries the recursive flag:

bash
git clone --recursive https://github.com/cvg/vidmap.git && cd vidmap

There is a .gitmodules and a third_party/ directory, and a plain clone leaves both empty.

PyTorch, TorchVision and xFormers are absent from the dependency list

The installation text says to install PyTorch, xFormers and VidMap, and then shows a single command:

bash
pip install -e .

That command does not install PyTorch or xFormers. The dependency list is twenty-four entries and none of them is torch, torchvision or xformers. What is there is numpy 1.26.4, scipy 1.15.3, pillow 11.3.0, opencv-python 4.11.0.86, h5py, pydantic, hydra-core, omegaconf, pyyaml, tqdm, natsort, evo, matplotlib, huggingface-hub, safetensors, accelerate, einops, addict, requests, and tomli for Python below 3.11.

So the two heaviest and most version-sensitive dependencies in the whole stack, the ones that decide which GPU kernel set you get and whether compiled models load at all, are left to you. That is a defensible choice for a research system where the CUDA build is user-specific, but the sentence groups them with the install command, which reads as though they are covered.

The tested versions are stated precisely, and they are the real contract. Linux x86-64, Python 3.10, COLMAP and PyCOLMAP 4.2, PyTorch 2.14.0, TorchVision 0.29.0 and xFormers 0.0.35, on NVIDIA GPUs. There is no macOS, no Windows, no CPU-only path and no AMD path in that sentence, and requires-python is set to 3.10 or newer.

One version constraint is stated separately and easy to miss: fast loading of cached compiled models requires PyTorch 2.10 or newer. The tested 2.14.0 clears it, but someone on an older torch silently loses the cache benefit rather than getting an error.

The first frontend run downloads about 10 GB with no way to pre-stage it

The size of the first run is stated in one sentence: it automatically downloads approximately 9 GB of model checkpoints and caches approximately 1.2 GB of compiled models.

Two properties of that sentence matter operationally. The download is automatic, so there is no documented flag to pre-fetch, verify or pin a checkpoint set before a run, and no environment variable for an offline model directory. On a machine without 10 GB of headroom or on a link where 10 GB is a bad idea, the first `python -m vidmap.frontend` is where you find out.

The second is that two different things are happening. Roughly 9 GB is fetched from somewhere remote, and roughly 1.2 GB is compiled locally from it. The compiled layer is what the PyTorch 2.10 requirement is about, and it is the part you would want to keep across runs rather than rebuild.

Which models are involved is visible from the packaging rather than from the instructions. The wheel excludes depth_anything_3/app, depth_anything_3/bench, depth_anything_3/cli.py, depth_anything_3/services and romav2/benchmarks, and the wheel license files name third_party/Depth-Anything-3, third_party/MegaLoc and third_party/RoMaV2. So Depth Anything 3, MegaLoc and RoMa v2 are vendored in the repository, and those are the depth and localisation models the frontend pulls weights for.

Since VidMap is described as an offline system, note what offline means here. It describes the reconstruction, not the model supply chain. The first run needs the network.

An offline reconstruction system whose browser viewer needs the internet

The opening line calls VidMap an offline Structure-from-Motion system for video. The visualization section then states that the browser viewer requires internet access.

Both statements are about different things, and the tension is worth naming because it is the kind of thing that surfaces on an air-gapped machine. The reconstruction runs locally against your own checkpoints. The HTML viewer, opened in a Chromium-based browser, wants the network.

The viewer is a file in the package, vidmap/visualization/vidmap-viewer.html, and you point it at a run folder containing `rec/` and `mapper_inputs/`. There is also a path that avoids the folder-selection step by embedding the reconstruction:

bash
python -m vidmap.visualization.html --run-dir "$OUTPUT_DIR"

That writes vidmap-viewer-embedded.html into the output directory so it opens directly. The documentation does not say whether the embedded variant drops the network requirement, and the browser requirement is stated only for the general viewer.

The Rerun path is the local alternative. Installing rerun-sdk and passing --save-playback-trace to the mapping command writes a trace to the output directory, and two commands turn it into recordings:

bash
python -m vidmap.visualization.rerun.playback "$OUTPUT_DIR" \
  --view isometric \
  --output "$VISUALS/solver-playback.rrd"

python -m vidmap.visualization.rerun.flythrough "$OUTPUT_DIR" \
  --reconstruction final \
  --view follow \
  --output "$VISUALS/flythrough.rrd"

Local images, frame timing and missing ground truth are detected automatically, so a dataset without ground-truth poses still plays back.

The depth visualisation has a caching requirement worth knowing. Depth lift replaces the sparse highlight with an RGB-coloured full depth prediction, which needs the predictions retained during reconstruction via --cache-depth-maps. Miss that flag on the first run and the backfill is not free: you re-run the frontend with --mapper-inputs and --cache-depth-maps, which the documentation says avoids repeating tracking and verification but is still a second frontend pass.

The sdist is a hand-maintained allowlist and the wheel carries three model licenses

The distribution configuration is unusually explicit, in both directions.

The source distribution starts by excluding everything:

toml
sdist.exclude = ["**"]
sdist.include = [
  "/pyproject.toml",
  "/README.md",
  "/extensions/colmap/src/**/*.cc",
  "/extensions/colmap/src/**/*.h",
  "/extensions/colmap/python/vidmap_native/*.py",
  "/extensions/colmap/CMakeLists.txt",

An allowlist. Every file in the sdist has to be added by hand, including the source of any dependency that is compiled rather than shipped as a wheel. A new module added to the package and not added here disappears from the sdist without any error, which is a quiet failure mode for anyone installing from source rather than from the git clone.

The wheel is cut the other way, by exclusion, and it declares which third-party licenses travel with it:

toml
wheel.license-files = [
  "third_party/Depth-Anything-3/LICENSE",
  "third_party/MegaLoc/LICENSE",
  "third_party/RoMaV2/LICENSE",
]

So the project is Apache-2.0 and redistributes code from three model projects under their own terms, which the packaging surface acknowledges by shipping their license files. Anyone deploying this needs to read all four, not just the top-level one.

The version is dynamic, taken from the build rather than written into the manifest, which is normal for scikit-build-core but means the version is only knowable by building. Combined with the release situation below, that means there is no published version string to cite.

Both release tags are assets, published 32 seconds apart

The release list for this repository contains two entries and neither is a version of the code.

They are named presentation-assets-v1 and dataset-assets-v2, and their titles are VidMap presentation assets v1 and VidMap dataset assets v2. Both were published on 2026-07-30, one at 02:37:36 and the other at 02:38:08. Thirty-two seconds apart, and neither carrying a code version.

So cvg/vidmap has no tagged release of the software. The paper, the video, the poster and the slides are all hosted somewhere else entirely: the poster and slide links resolve into a different repository, Zador-Pataki/VidMap-assets, under an eccv-2026 tag, and the arXiv identifier is 2607.27194. The assets repository is where the ECCV material is versioned; this repository gets asset tags that point at nothing a user installs.

That leaves the git clone as the only acquisition path, which is consistent with the build story. A clone with submodules, a from-source COLMAP, a CMake build of extensions/colmap, and hand-installed PyTorch and xFormers, with the version determined by whatever commit you cloned.

The last push was 2026-09-30, so the branch is active. What is missing is not activity but a version marker, and that matters for a research artifact other people build on: there is no tag to record in a paper's reproduction section or in a container image label.

Authors are Zador Pataki, Paul-Edouard Sarlin and Marc Pollefeys, and the venue is ECCV 2026.

Three git dependencies, two of them cvg forks rather than upstream

Twenty-one of the dependencies are pinned to exact released versions with the double-equals form. Three are git URLs, and each one is pinned to a full commit hash:

toml
poselib @ git+https://github.com/PoseLib/PoseLib.git@48b0fdd032534f1d76ffb5bc1536b337d5517c50
lightglue @ git+https://github.com/cvg/LightGlue.git@c2f8561576f0160b2cdaa3e4634b6b61d6cf6cb8
geocalib @ git+https://github.com/cvg/GeoCalib.git@982f625d822fbb2b50c65d123fd1362a5c01ec2c

The hashes are the good practice here. A build is reproducible against those three, which is more than a branch name would give you, and it means a broken upstream main cannot break an existing checkout.

What is worth noticing is the ownership. PoseLib is referenced at its upstream location. LightGlue and GeoCalib are referenced under the cvg organisation rather than upstream, so two of the three are the group's own forks of widely used components. That is a normal pattern for research code and it means the project can patch them, but it also means the version you get is not the version the upstream projects publish, and diffing against upstream is your job.

One small version detail: tomli is conditional, marked for Python below 3.11, so a 3.11 or 3.12 environment drops it and pulls the stdlib tomllib instead.

The rest of the surface is plain: repository has .gitattributes, .github/, .gitignore, .gitmodules, LICENSE, README.md, assets/, extensions/, pyproject.toml, third_party/ and vidmap/. No changelog, no contributing guide and no security policy at the root, which is thin for an Apache-2.0 project that ships C++ extensions.

Speed up VidMap and Benchmarks are promised in the quickstart and missing from the page

The quick-start list at the top of the page names six sections. Four are present in the readable text.

Setup, Execution, Visualization and Configuration all appear. The last two entries do not: Speed up VidMap, framed as finding settings that make extraction and mapping faster, and Benchmarks, framed as preparing datasets and evaluating reconstruction accuracy.

The missing benchmarks section is the one that would matter most to an adopter. This is an ECCV paper repository, the description names the venue, and the visible page contains no accuracy table, no dataset list and no comparison against a baseline. A structure-from-motion system is judged on pose error and map quality, so the absence of any of those numbers means the only performance claims available here are qualitative.

The absent speed section points at a real tension in the design. The first frontend run downloads about 9 GB of checkpoints and compiles about 1.2 GB of models, which is slow by construction, and the quickstart promises settings that change that. Those settings live in the frontend configs, which the configuration section does describe as controlling tracks, keyframes, depth and calibration, grouped by pipeline stage.

The configuration layer itself is the strongest part of what is visible. Hydra and OmegaConf provide typed defaults, and the defaults are printable rather than something you dig out of YAML:

python
from vidmap.configuration import FrontendConfig, MappingConfig, summarize_cfg

print(summarize_cfg(FrontendConfig()))
print(summarize_cfg(MappingConfig()))

Overrides are per-file and inherit through a defaults list, with the example living under a frontend directory named uncalib, which implies the uncalibrated-camera path is a first-class configuration rather than a special case.

Editorial conclusion

Use VidMap if you already run COLMAP from source, have an NVIDIA GPU on Linux x86-64, and want a video SfM front end that hands mapper inputs to a pipeline you can tune with typed Hydra configs. Do not plan around the pip install alone, because PyTorch and xFormers are absent from the dependency list and the first frontend run pulls about 10 GB with no pre-staging flag. Verify first that a plain build of extensions/colmap against Ceres 2.3 with CUDA and cuDSS completes, since that CMake step is not something pip can recover from.

Frequently asked questions

What does VidMap estimate from a video?

It is an offline structure-from-motion system for video that combines temporal tracks, loop closures, metric depth and global optimization to estimate camera poses, camera intrinsics and a sparse 3D map. Both mapping commands write the final COLMAP model under the output directory in a rec subfolder.

What are the system requirements for running VidMap?

It was last tested on Linux x86-64 with Python 3.10, COLMAP and PyCOLMAP 4.2, PyTorch 2.14.0, TorchVision 0.29.0 and xFormers 0.0.35 on NVIDIA GPUs. COLMAP 4.2 and its PyCOLMAP bindings must be built from source, with Ceres 2.3 or newer plus CUDA and cuDSS for GPU mapper acceleration. requires-python is 3.10 or newer.

Why does installing VidMap not install PyTorch?

PyTorch, TorchVision and xFormers are not in the dependency list, which is twenty-four pinned packages starting at numpy 1.26.4. The setup text names them alongside the install command but leaves them to be installed separately. Fast loading of cached compiled models additionally requires PyTorch 2.10 or newer.

How much data does the first VidMap run download?

The first frontend run automatically downloads approximately 9 GB of model checkpoints and caches approximately 1.2 GB of compiled models. The vendored model projects named in the wheel configuration are Depth-Anything-3, MegaLoc and RoMaV2.

Does VidMap have any code releases?

No. The two release tags are named presentation-assets-v1 and dataset-assets-v2 and were published 32 seconds apart on 2026-07-30, carrying paper and dataset assets rather than a code version. The poster and slides are hosted in a separate repository, Zador-Pataki/VidMap-assets. The package version is dynamic and derived at build time.

Official sources

  1. cvg/vidmap on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cvg-vidmap.svg)](https://hysenlabs.com/projects/cvg-vidmap)