DVLT loops one attention block over depth indices, and its config chapter stops mid-sentence
Official implementation of Déjà View: Looping Transformers for Multi-View 3D Reconstruction
At a glance
- What is it?
- A recurrent transformer that turns an unordered set of images into per-pixel rays, depth, confidence, and camera poses, shipped with a released checkpoint and wrappers for five feed-forward baselines. What the repository pins down hard is its dependency stack; what it leaves open is where your own data paths go.
- Who is it for?
- DVLT fits a research group that already has ScanNet++ data, a CUDA machine, and patience for exact-pinned dependencies, and the offline demo path is the cheapest way to see whether a looped block earns its keep on your own footage.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The quickstart snippet ends one line into the load call
The only inference example in the repository is a Python block. It imports `DVLT` along with `load_sequence` and `preprocess_images`, sets `checkpoint_path = "nvidia/dvlt"` with a comment noting that a local directory, an HTTPS URL, or a Hugging Face Hub repo id all work, wraps the run in `Accelerator(mixed_precision="bf16")`, and builds `DVLT(img_size=504)`. The block then stops mid-call on `model.load_pretrained(checkpoint_path, strict`. The argument after `strict` is never shown, the loader's return value is never bound to a name, and the per-pixel rays, depth, confidence, and poses the overview promises are never read out. The same gap hides the compute dial: the overview calls the number of refinement steps `K` an inference-time knob, and the one snippet that would set it is the snippet that ends early. The input side is settled, though. A comment on `load_sequence` states it accepts a directory, a single video, or an explicit list of files.
import torch
from accelerate import Accelerator
from dvlt.model.dvlt.model import DVLT
from dvlt.util.preprocess import load_sequence, preprocess_images
checkpoint_path = "nvidia/dvlt" # local dir, HTTPS URL, or HF Hub repo id
# load_sequence accepts a directory, a single video, or an explicit list of files.
input_path = "path/to/scene_dir"
# Or: input_path = "path/to/clip.mp4"
# Or: from glob import glob; input_path = sorted(glob("path/to/scene_dir/*.png"))
accelerator = Accelerator(mixed_precision="bf16")
model = DVLT(img_size=504)
model.load_pretrained(checkpoint_path, strictData paths open in a sentence that never finishes
The configuration chapter walks the Hydra layer first, then opens a subsection titled User configuration (data paths) and begins it with `Per-user settings (most importantly,`. The sentence has no ending. No file name, no directory, no example follows it, so the place where a user points the dataset loaders at their own images is not written down in the visible chapter. The packaging file ends at a similar boundary: the readable portion of pyproject.toml stops under a `# Using setup` comment, right after the optional-dependency groups, so whatever that comment introduced stays unstated. What the visible text does fix is the shape of the config system. Top-level experiment configs live in `src/dvlt/config/experiments/`, a resume run points Hydra at `outputs/<run>` through `--config-dir`, and step-level changes ride on the command line as dotted overrides.
# Resume
python -m dvlt.scripts.train \
--config-dir=outputs/<run> \
--config-name=config.yaml \
trainer.resume_from_checkpoint=latestVersion 0.0.1 and an Alpha classifier next to an Oral badge
pyproject.toml declares `version = "0.0.1"` and lists `Development Status :: 3 - Alpha`, while the header of the README carries a NeurIPS 2026 Oral badge and three outbound links: a research project page, an arXiv entry, and the Hugging Face repo `nvidia/dvlt`. There are no GitHub releases, which leaves the packaging metadata as the only version signal in the repository, and it has never left 0.0.1 even though the tree now holds a Stage-2 fine-tune recipe, an inference alias for a released checkpoint, and wrappers for five baselines. The release-status checklist is blunter than the version field. Inference code, model weights, evaluation code with its dataset preprocess and loaders, and training code are all ticked, and under training code only the ScanNet++ loader is ticked, with other training dataset loaders left open. The default branch was last pushed on 2026-09-25, and the repository is not archived.
The license field says NOASSERTION, the classifier says Apache
The license field in the repository's own metadata comes back as NOASSERTION. Inside the tree, pyproject.toml sets `license = { file = "LICENSE" }` and the classifier block includes `License :: OSI Approved :: Apache Software License`. The top level also carries a `LICENSE` file, a `LICENSES/` directory, and a separate `THIRD_PARTY_LICENSES.md`, so the project itself acknowledges more than one licensing document without saying, in the visible text, which one covers what. Three signals, three readings: an automated check that trusts the metadata field sees an unknown license, a human reading the classifier sees an Apache license, and a reader of the tree sees a bundle whose internal split is unstated. Settle that before shipping anything derived from the code, and read the directory and the third-party file rather than treating the classifier line as the answer.
Every dependency is pinned with an equals sign, and the dev tools ride along
The core dependency list pins every entry exactly, from `torch==2.5.1` and `torchvision==0.20.1` down to `numpy==2.4.4`, `open3d==0.19.0`, and `safetensors==0.5.3`, and the optional groups are pinned the same way. Three details sit oddly against that strictness. The install recipe opens a conda environment with `python=3.12`, while the packaging metadata asks for `requires-python = ">= 3.11"` and lists only `Programming Language :: Python :: 3.11` among its classifiers. `pip install -e .[all]` resolves to a group defined as `dvlt[demos,dev]`, so a first install also brings `black==25.1.0`, `ruff==0.12.1`, `pytest==8.4.1`, and `pre-commit==4.2.0` into the same environment as the model. And the comment on the visualization block explains why `matplotlib`, `plotly`, and `rerun-sdk` sit in core dependencies at all: a callbacks package imports them transitively, so an inference-only user installs a plotting stack whether or not anything gets rendered.
conda create -n dvlt python=3.12 && conda activate dvlt
conda install pytorch=2.5.1 torchvision pytorch-cuda=12.4 -c pytorch -c nvidia -c conda-forge
pip install -e .[all]Two benchmark scopes, one results table, no baseline rows
Evaluation comes in two sizes. `benchmark_lite` covers DTU, ETH3D, and 7Scenes, described as the datasets that skip heavy preprocessing, and the full `benchmark` adds ScanNet++ and NuScenes, each with its own preprocessing document under `src/dvlt/scripts/preprocess/`. Reference numbers appear for the full scope only, across four columns: pose AUC@3, pose AUC@30, depth inlier@3%, and depth AbsRel. DTU is the strongest row at 0.8319 pose AUC@3 and 0.0093 AbsRel, 7Scenes the weakest on pose at 0.1393, and NuScenes the weakest on depth at 0.5853 inlier with 0.0673 AbsRel. The table holds no baseline column and no parameter counts, while the overview claims the model matches or outperforms substantially larger feed-forward baselines at a fraction of their parameters. The wrappers for those baselines ship in the tree, so the comparison is runnable, but the numbers that would settle the claim are not printed.
python -m dvlt.scripts.test --config-name dvlt data=benchmark
# multi-GPU: accelerate launch --num-processes <N> -m dvlt.scripts.test --config-name dvlt data=benchmark
python -m dvlt.scripts.test --config-name dvlt data=benchmark_lite
# multi-GPU: accelerate launch --num-processes <N> -m dvlt.scripts.test --config-name dvlt data=benchmark_liteFive baseline wrappers, and a demo that needs each upstream package
The evaluation layer wraps VGGT, VGGT-Omega, Depth-Anything-3, MapAnything, and Pi3, and the config table expands some of those into more names than the prose lists: `da3-{base,large,giant}` and `pi3x` sit alongside `vggt`, `vggt_omega`, and `mapanything`. Each wrapper imports its upstream package, and those packages are installed separately, which is why `docs/INSTALL.md` gets linked from four separate places. The demo shows the same dependency from the browser side. Its dropdown switches between DVLT and the wrappers, and each baseline needs its upstream package present. `python -m dvlt.scripts.gradio_demo` serves on http://localhost:7860 with DVLT preselected. Add `--offline` and Gradio is skipped entirely, with a `.glb` and a `.rrd` written per sequence and model into `demo_outputs/<sequence_name>/`, `--input` repeatable across sequences, and `--models` taking a comma-separated list of registry names or `all`.
# Run two models on two sequences (one image dir, one video)
python -m dvlt.scripts.gradio_demo --offline \
--input /path/to/scene_dir \
--input /path/to/clip.mp4 \
--models dvlt
# Run every registered model on one sequence
python -m dvlt.scripts.gradio_demo --offline --input /path/to/scene_dir --models allThe released checkpoint has a conv depth head, the Stage-1 recipe does not
The config table separates three names that are easy to conflate. `dvlt-large` is the Stage-1 recipe: large model, full training schedule, linear depth head. `dvlt-large-depthconv-stage2` is the fine-tune of the depth-conv head, and it is the entry that matches the released checkpoint and the model default `depth_head_type="conv"`. `dvlt` is an inference-only alias for that Stage-2 checkpoint. Training from the Stage-1 recipe and then pointing inference at the alias therefore means the head has to take a second pass to land on the shipped configuration. The ablations branch off the same parent. `dvlt-large-ablation` toggles decoupled blocks, no-`s_out`, and no-depthscale through overrides, and `dvlt-large-ablation-decoupled` switches looping off with `recurrence_mode=none`, giving a distinct block per step at a fixed 16 steps. The shared-block path reuses one block instead, which is what turns the same count into an inference-time dial.
Editorial conclusion
DVLT fits a research group that already has ScanNet++ data, a CUDA machine, and patience for exact-pinned dependencies, and the offline demo path is the cheapest way to see whether a looped block earns its keep on your own footage. Settle three things before building on it: which document under LICENSES/ covers the code you intend to reuse, whether the s_out or depth-scaling ablations shift your accuracy target enough to matter at your step budget, and where the per-user data path configuration actually lives. Read the published benchmark rows as DVLT-only figures, not as evidence about the five baselines shipped alongside them.
Frequently asked questions
What does nv-tlabs/dvlt predict from a set of images?
Per-pixel rays, depth, confidence, and camera poses. It is a recurrent transformer that loops a shared block of frame/global attention with discrete depth indexing, and it takes an unordered set of images as input.
Can I run nv-tlabs/dvlt inference without training it?
Yes. The `dvlt` config is an inference-only alias for the released Stage-2 checkpoint, and the weights sit on the Hugging Face Hub as `nvidia/dvlt`, which the quickstart block accepts directly as a repo id. There are no GitHub releases to install from.
Which training datasets does nv-tlabs/dvlt support today?
Only ScanNet++ has a loader. The release checklist ticks the ScanNet++ training dataset loader and leaves other training dataset loaders unticked, so training on a different dataset means writing a loader for it.
Do the baseline wrappers in nv-tlabs/dvlt work out of the box?
No. Each wrapper imports its upstream package, and VGGT, VGGT-Omega, Depth-Anything-3, MapAnything, and Pi3 are installed separately, with the steps in `docs/INSTALL.md`. A baseline you have not installed is simply absent from the comparison.
Which Python and torch versions does nv-tlabs/dvlt expect?
The install recipe creates a conda environment with `python=3.12` and installs `pytorch=2.5.1` with `pytorch-cuda=12.4`, while the packaging metadata declares `requires-python = ">= 3.11"` and pins `torch==2.5.1` and `torchvision==0.20.1` with exact `==` pins throughout.
Where do I point nv-tlabs/dvlt at my own images?
The configuration chapter opens a User configuration (data paths) subsection and stops at `Per-user settings (most importantly,`, so the file that holds those paths is not named in the visible text. `load_sequence` still takes a directory, a single video, or an explicit list of files, which is enough to drive the demo.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/nv-tlabs-dvlt)