# Pi3: 3D reconstruction from images with no reference view

> A feed-forward network that predicts camera poses and point maps from an unordered set of images, permutation-equivariant by design, plus a PyTorch checkpoint you can run from the command line. Geometry without a chosen keyframe.

**yyfz/Pi3** — [ICLR 2026] π^3: Permutation-Equivariant Visual Geometry Learning

- Repository: https://github.com/yyfz/Pi3
- Website: https://yyfz.github.io/pi3/
- Stars: 2,197 · Forks: 172
- Language: Python
- License: BSD-3-Clause
- Published: 2026-10-08 · Updated: 2026-10-08 · Language: en
- Canonical page: https://hysenlabs.com/projects/yyfz-pi3

## Dropping the reference frame that most pipelines depend on

Most multi-view geometry methods pick one image as the reference and express everything else relative to it. The README's argument against that convention is practical rather than theoretical: a designated reference frame introduces instability, and if you pick a bad one the reconstruction suffers.

Pi3 removes the choice. It is a feed-forward network that takes an unordered set of images and predicts directly from them, which the README describes as permutation-equivariant. Two consequences follow. Reordering the input images does not change what the model means by them, and the outputs are expressed in forms that do not depend on a scale reference: affine-invariant camera poses and scale-invariant local point maps.

The paper claims an emergent property rather than an engineered one: a dense, structured latent representation of the camera pose manifold, learned from a bias-free design without added priors or a special training scheme. If that holds, it is a cleaner account of geometry than a hand-built reference graph, and the claimed results span camera pose estimation, monocular and video depth estimation, and dense point map estimation.

Two supporting facts are worth holding onto. The project description carries an ICLR 2026 tag while the arXiv identifier in the README is 2507.13347, which places the preprint in mid 2025, and the README itself does not mention a conference venue anywhere in its text. The paper and the project page are the places to confirm publication status; the repository metadata is not.

The repository is BSD-3-Clause licensed with a LICENSE file at the root, and carries no GitHub releases, so there is no versioned checkpoint trail to consult. Version information lives in the packaging file and on the model hosting platform.

## Pi3X replaces the output head and adds conditioning

The December 2025 update replaced the original model with Pi3X, and the README describes the change in four parts.

The output head became convolutional, which the README credits with removing grid-like artifacts and producing smoother point clouds. That is a rendering-quality fix rather than a modelling change, and if your complaint about the original model was banding in the point cloud, this is the part that addresses it.

Input conditioning became optional. Pi3X can accept camera poses, intrinsics and depth as additional inputs, so if you already have partial priors from another stage of a pipeline you can inject them rather than making the model guess. This is the feature that most changes what you can build, because it turns a standalone estimator into a component that accepts help.

Confidence scoring also changed. The original approximated a binary mask; Pi3X predicts continuous quality levels, which the README describes as significantly more reliable for filtering noise. For anyone post-processing point clouds, a usable per-point confidence is the difference between keeping a cloud and hand-editing it.

Finally there is approximate metric scale. The original predictions were purely scale-invariant, meaning you recovered shape but not size; Pi3X adds an approximate metric scale reconstruction. The word approximate is doing real work there, and the README says so.

One naming trap. The scripts are example.py for the original model and example_mm.py for Pi3X, and the README marks the latter as recommended while leaving the former working. The bare name pi3 refers to both.

## Running inference on a folder of images or a video

Getting started is a clone and a dependency install:

```bash
git clone https://github.com/yyfz/Pi3.git
cd Pi3
pip install -r requirements.txt
```

Then one script. The README shows the default example video and the two variants side by side:

```bash
# Run with the default example video
# python example.py    # Inference with Pi3 (Original)
python example_mm.py   # [New] Inference with Pi3X (Recommended)

# Run on your own data (image folder or .mp4 file)
# python example.py --data_path <path/to/data>     # Pi3
python example_mm.py --data_path <path/to/data>    # Pi3X
```

The defaults are visible in the argument list: the data path falls back to examples/skating.mp4 and the output to examples/result.ply, so running the script with no arguments produces a result rather than an error. The frame sampling interval has two different defaults, one for an image directory and a coarser one for video, which tells you the video path is subsampling rather than processing every frame.

Checkpoints download automatically from the model hosting platform. The README covers the slow case: fetch model.safetensors manually and point the script at it with the checkpoint argument.

The output is a ply point cloud, which is the right format to open in a mesh viewer and the wrong format to feed to anything expecting a mesh. There is no meshing step in the repository.

Sample data is included rather than merely referenced. The tree carries video files for skating, parkour and skiing, plus directories for a room, a house, a valley, a long walking sequence, and a set of gradio examples, so you can check the shape of the output before supplying your own capture.

## Multimodal conditioning needs its own data format

The conditioning path is documented separately from the basic run because the input format is not self-describing. The README's comparison runs the same scene twice, once with priors and once without:

``` bash
# 1. Inference WITH conditioning (poses, intrinsics, etc.)
python example_mm.py --data_path examples/room/rgb --conditions_path examples/room/condition.npz --save_path examples/room_with_conditions.ply

# 2. Inference WITHOUT conditioning (image only)
python example_mm.py --data_path examples/room/rgb --save_path examples/room_no_conditions.ply
```

Two things fall out of those paths. The priors arrive as a NumPy npz file, not as a text file or a directory, which means you have to write whatever produces that archive yourself, and the README does not specify its internal structure, only pointing at the script for formatting details. And the two runs write to different output paths so you can compare the point clouds, which is the intended way to judge whether the conditioning is helping or hurting for your data.

The room example in the repository has an rgb subdirectory and a condition.npz file, so you have a working example to inspect rather than a specification to guess at. That matters more than it sounds: the sample conditions file is the only concrete description of the format available.

Everything about the conditioning path is Pi3X only. The original model's script does not accept it, which is the practical reason to settle the Pi3 versus Pi3X question before writing any data plumbing.

## Training and evaluation live on branches you will not get by cloning

The default branch holds inference, and the updates section in the README is explicit that the other two halves live elsewhere. Training code is on a branch named training, updated on September 3, 2025. Evaluation code is on a branch named evaluation, released on July 29, 2025.

Neither is in the tree of the default branch. What you get on the default branch is example.py, example_mm.py, another script named example_vo.py, a gradio demo script, a capacity benchmark script, two benchmark result files for Pi3 and Pi3X, the pi3 package directory, and a scripts directory. The benchmark result files being checked in is a small good sign: it means the numbers can be compared across a checkpoint change without rerunning anything.

So the reproduction path is three clones rather than one. That is a deliberate choice common to research code, and it has a real cost: training and evaluation are versioned separately from the inference code, so a paper result and the current default branch are not guaranteed to correspond. If you are reproducing a number from the paper, clone the branch that produced it.

The dependencies are pinned in a requirements.txt that fixes torch at 2.5.1, torchvision at 0.20.1 and numpy at 1.26.4, plus pillow, opencv-python, plyfile, huggingface_hub and safetensors. A separate requirements_demo.txt exists for the gradio demo, which keeps the interactive dependencies out of the inference path.

The pins are old relative to the current release cadence of PyTorch, and that is deliberate pinning rather than carelessness. It means a newer GPU stack may not be installable without editing the file, and it is the first thing to change if your environment refuses the install.

## A packaging file that declares one dependency and installs nine

Here is an inconsistency worth knowing before you install anything. The pyproject.toml in the repository declares the project name pi3 at version 0.1, requires Python 3.10 or newer, and lists exactly one dependency, safetensors. The requirements.txt next to it lists nine packages including torch and torchvision.

So a pip install of the package gives you a package that cannot run anything, because the only thing it brings is the safe tensor reader. The instructions in the README are right: install from requirements.txt after cloning. But the packaging metadata is misleading to anyone who finds the project through a package index or tries to depend on it from another project, and the README does not call this out.

The setuptools configuration also disables automatic package discovery and names the pi3 package explicitly, with package data included by glob. That is deliberate: without it, the repository's example scripts and data directories would be picked up as packages during a build. It is a sign the author hit that problem.

So the honest summary of the packaging is that this is research code published for use, not a library published for dependency resolution. Inference works from a clone. Depending on pi3 from another Python project would require declaring torch and the rest yourself, and would give you a project still numbered 0.1.

For the interactive route there is a Hugging Face space and a gradio demo script in the tree, which is the fastest way to see what the model produces before committing GPU time to a local run.

## Conclusion

Pi3 is worth trying if you are building something that needs depth, camera pose or point maps from a handful of images without training a model of your own, and the equivariance is the reason to try it over a fixed-reference pipeline. The practical surface is small: two example scripts, a checkpoint download, and a ply file at the end. Two things should shape your expectations. The original model and Pi3X are different things with different entry points, and the README recommends Pi3X while keeping the original script working, so decide which one you want before you start. And the packaging is uneven, with pyproject.toml declaring only safetensors as a dependency while requirements.txt pins torch and torchvision separately, so install from requirements.txt rather than installing the package and assuming the environment came with it. If you need training or evaluation code, note that the README points at separate branches for both rather than the default branch you would clone.

## FAQ

### What is the difference between Pi3 and Pi3X?

Pi3X is the updated version of Pi3, released on December 2025. It replaces the original output head with a convolutional head for smoother point clouds, adds optional conditioning on camera poses, intrinsics and depth, predicts continuous confidence levels instead of a binary mask, and supports approximate metric scale reconstruction. The original remains available through example.py, while Pi3X is recommended and runs through example_mm.py.

### How do I run inference on my own images with Pi3?

Clone the repository, run pip install -r requirements.txt, then pass your image folder or video file with the data path argument, for example python example_mm.py --data_path <path/to/data>. The default input is an example video and the default output is a ply point cloud. If the automatic checkpoint download is slow, download model.safetensors manually and pass its path with the checkpoint argument.

### Is Pi3's output metric scale or only relative scale?

The original model predicts scale-invariant local point maps, so it recovers shape without absolute size. Pi3X adds metric scale reconstruction, described in the README as approximate. If your downstream work needs real units, plan to treat that scale as an estimate to be refined rather than a measurement.

### Where do I find the training and evaluation code for Pi3?

Not on the default branch. The README's updates section points to a branch named training for the training code and a separate branch named evaluation for the evaluation code. The default branch carries inference scripts, a gradio demo, a capacity benchmark and two checked-in benchmark result files, so reproducing a paper number means cloning the branch that produced it.

## Sources

- [Issues](https://github.com/yyfz/Pi3/issues)
- [License: BSD-3-Clause](https://github.com/yyfz/Pi3/blob/main/LICENSE)
- [Project website](https://yyfz.github.io/pi3/)
- [README](https://github.com/yyfz/Pi3/blob/main/README.md)
- [yyfz/Pi3 on GitHub](https://github.com/yyfz/Pi3)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/yyfz-pi3
