Open-source project
zju3dv/InfiniSplat avatar
zju3dv/InfiniSplat

InfiniSplat turns one photo into a Gaussian splat, with or without a depth sensor

[SIGGRAPH Asia 2026][TOG 2026] InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis

346 stars14 forksPythonNOASSERTION

At a glance

What is it?
Inference code from a SIGGRAPH Asia 2026 journal track paper, released as a single command-line module with two input modes: plain RGB or RGB paired with depth. The setup work sits outside pip, since PyTorch has to be matched to your CUDA build first.
Who is it for?
InfiniSplat suits researchers and engineers who already have a CUDA-capable machine and want a splat representation from a handful of images rather than a captured scene. It is not a turnkey tool for a workstation with no GPU, because the PyTorch build is a separate manual step and the checkpoint arrives from Hugging Face on first run.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 20 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Two input modes, one demo module

The released inference code is a single Python module, src.demo.infer_batch_images, and the difference between the two supported modes is one flag. Without it, the module takes an RGB image and reconstructs a 3D Gaussian representation from it. With it, the same module takes an RGB image plus a depth map and uses the sensor reading as guidance. The bundled example calls that flag --mode lidar, and the depth examples live in examples/data/lidar_demo while the plain images sit in examples/data/rgb_demo.

A single file in RGB mode:

bash
python -m src.demo.infer_batch_images --input examples/data/rgb_demo/pexels-masi.jpg

The same command pointed at a directory sweeps every example image in it:

bash
python -m src.demo.infer_batch_images --input examples/data/rgb_demo

So the batch path is not a separate script or a config toggle. It is the same entry point with a directory where a filename used to be, which keeps the depth branch consistent with the RGB one.

Depth mode pairs files by matching filename stems

The depth branch adds a constraint that the RGB branch does not have. An RGB image and its depth map have to sit in the same directory and carry matching filename stems, so the module can decide which depth belongs to which photograph. Nothing in the invocation names a depth file separately; the pairing is inferred from the names.

A single pair:

bash
python -m src.demo.infer_batch_images \
  --mode lidar \
  --input examples/data/lidar_demo/eth3d_kicker.png

The bundled pairs use the directory form of the same flag:

bash
python -m src.demo.infer_batch_images \
  --mode lidar \
  --input examples/data/lidar_demo

Because of the stem matching, dropping a depth map into a folder of RGB images under a different name does not produce an error, it produces a run that quietly treats files as unpaired. Renaming files to agree before the run is the part of this mode that has to be done by hand.

PyTorch is installed first, outside the requirements file

The dependency file opens with a comment instead of a torch line, and that is the first thing to get right. A CUDA-matched PyTorch build is installed separately, before the file itself, because the correct build depends on the driver and GPU on the machine. The comment gives the shape of the command and an example for CUDA 12.8:

bash
pip install torch torchvision xformers --index-url https://download.pytorch.org/whl/cu128

The index URL is the part that matters. Without it pip resolves torch from the default index, which does not know which CUDA build belongs on your hardware.

Everything after that comment is ordinary: hydra-core, omegaconf and dacite for configuration, numpy, scipy, pandas, h5py, einops, jaxtyping and opencv-python-headless for the data path, then huggingface-hub, timm and xformers for the backbone and the checkpoint fetch. Environment setup and checkpoint download are handled in INSTALL.md rather than in this file.

Hydra and dacite sit behind the flags

Three of the declared dependencies explain how the command line behaves. hydra-core and omegaconf bring a structured configuration system in, and dacite validates typed configuration objects, which means a malformed setting fails at load time rather than somewhere inside the reconstruction. The repository root carries a config directory for that purpose.

The consequence is that the two flags shown in the quick examples are the small surface. Camera parameters, output control and the remaining optional arguments are described in docs/inference.md, and the root also holds a scripts directory alongside src, so anything beyond the documented flags is a configuration change or a script rather than a third mode.

The geometry stack in the same file shows what the reconstruction leans on: scikit-learn, Pillow, imageio with ffmpeg, matplotlib, plyfile and torchmetrics. plyfile in particular is the format writer, which is consistent with the output the demo hands back as PLY.

numpy is held below 2.0 and two packages are pinned exactly

Three version constraints in the file deserve attention because they decide whether the environment resolves cleanly. numpy is capped below 2.0, which is an unusual ceiling for a new project and means an existing environment with numpy 2 already installed will be changed. gradio is pinned to exactly 6.20.0 and nodejs-wheel to exactly 22.20.0, both written with the full three-part version.

Those two pins belong to the local demo and the standalone HTML viewer export. A Node runtime arriving as a Python wheel is what lets the demo produce a self-contained HTML file, and holding both versions exactly means the export has one known-good combination rather than whatever the resolver finds that week.

The remaining packages use lower bounds: einops at 0.4.1 or newer, and open-ended entries for the rest. Nothing in the file caps hydra-core, torchmetrics or timm, so a fresh install pulls the current release of each.

The Gradio demo fetches a checkpoint unless one already exists

The browser interface is a single command and a fixed local port:

bash
python demo.py

Opening http://127.0.0.1:7860 gets RGB reconstruction, staged PLY and standalone HTML downloads, and an interactive Gaussian viewer. The checkpoint handling is the part worth knowing before the first run. The demo reuses checkpoints/infinisplat_rgb.ckpt when that file is present, and otherwise downloads the released RGB checkpoint from the PLUS-WAVE/InfiniSplat repository on Hugging Face. Setting INFINISPLAT_CHECKPOINT points it at a checkpoint somewhere else, which is the route for anyone holding a fine-tuned or converted weights file.

A hosted copy of the same interface is published as a Hugging Face Space, so the interface can be looked at before any local environment exists. The project page sits at zju3dv.github.io/InfiniSplat, and the paper is on arXiv as 2608.02437.

The licence asks companies to register before anything else

InfiniSplat ships under the Project Registration License v1.0, and it draws a line between individual and organisational use rather than between commercial and non-commercial. Personal, academic, educational and evaluation use is free with no registration, and research results may be published without registering first.

Organisational use is also free, but it requires prior registration, and the licence is explicit that this covers non-commercial projects, internal research and development, testing, deployment and production. Completing an accurate registration grants permission automatically, with no approval step and no fee. Each materially distinct project is registered once, which means one organisation working on several projects files several registrations.

The registration form is a Google Form, and the repository root carries a Word template for the same purpose. For a lab or a company, that paperwork is the first task in using this code, not an afterthought.

Four upstream projects carry the reconstruction

The acknowledgments name the work InfiniSplat is built on, and each one maps onto a part of the pipeline. DINOv3 supplies the visual backbone, Depth Pro from Apple is the depth model behind the depth-guided path, InfiniDepth from the same lab covers the depth side of their own earlier work, and gsplat provides the Gaussian rasterisation.

That list also sets expectations about upgrade behaviour. Three of the four are active upstream projects that move on their own schedules, and none of them is pinned to a version in the requirements file. A base image model gaining a new default config or a rasteriser changing an interface is the kind of change that would show up here as a run that fails rather than as a version conflict at install time.

The project itself is a research release with no tagged version. The last commit on the main branch is dated 2026-09-11, and no GitHub release has been published, so the checkpoint and the code are matched by date rather than by a version string.

Editorial conclusion

InfiniSplat suits researchers and engineers who already have a CUDA-capable machine and want a splat representation from a handful of images rather than a captured scene. It is not a turnkey tool for a workstation with no GPU, because the PyTorch build is a separate manual step and the checkpoint arrives from Hugging Face on first run. Read the licence before any company trial: personal and academic use costs nothing, while internal research, testing, deployment and production all require registration first, which is a paperwork step rather than an approval queue. Camera parameters and output control live in docs/inference.md, so check that file before assuming the two documented flags are all there is.

Frequently asked questions

What does InfiniSplat reconstruct from a single image?

It reconstructs a 3D Gaussian Splat representation from one RGB image. Passing --mode lidar adds a depth map as guidance, taking an RGB image plus depth as input for the same output.

Do I have to install PyTorch before InfiniSplat?

Yes. The requirements file opens with a comment saying a CUDA-matched PyTorch build is installed separately, and gives CUDA 12.8 as the example index URL.

How do I run the depth-guided mode in InfiniSplat?

Use --mode lidar with --input. The RGB image and its depth map must be in the same directory with matching filename stems, so the demo runs on the bundled examples in examples/data/lidar_demo.

Where does the InfiniSplat web demo get its checkpoint?

The demo reuses checkpoints/infinisplat_rgb.ckpt when it exists, and otherwise downloads the released RGB checkpoint from the PLUS-WAVE/InfiniSplat repository on Hugging Face. INFINISPLAT_CHECKPOINT overrides the path.

Can a company use InfiniSplat without registering?

No. Under the Project Registration License v1.0, organisational use including non-commercial work, internal R&D, testing, deployment and production requires prior registration. A complete registration grants permission automatically, with no approval and no fee.

What packages does InfiniSplat pin to an exact version?

gradio is pinned at 6.20.0 and nodejs-wheel at 22.20.0 for the local demo and the standalone HTML export, and numpy is held below 2.0. Everything else uses a lower bound.

Official sources

  1. Issues
  2. Project website
  3. README
  4. zju3dv/InfiniSplat on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/zju3dv-infinisplat.svg)](https://hysenlabs.com/projects/zju3dv-infinisplat)