CLI tool
apple-aiml-research/ml-depth-pro avatar
apple-aiml-research/ml-depth-pro

ml-depth-pro: Apple's reference implementation for metric depth from one photo

Depth Pro: Sharp Monocular Metric Depth in Less Than a Second.

5,739 stars432 forksPythonNOASSERTION

At a glance

What is it?
A retrained Depth Pro checkpoint that predicts absolute-scale depth and focal length from a single image, packaged as a small Python library with a one-image command line entry point.
Who is it for?
Depth Pro is a good fit when a downstream consumer needs depth in metres rather than a relative ordering of pixels, because that is the property most monocular models give up. The reference implementation is small enough to read, with the network and preprocessing under `src/depth_pro`, a single shell script for weights, and a CLI that takes one image path.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 26 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Depth Pro predicts and why metric scale is the hard part

Most monocular depth models output relative depth: nearer pixels get larger numbers, and the absolute scale is arbitrary. Depth Pro is built around removing that ambiguity. The README describes a foundation model for zero-shot metric monocular depth estimation, and the predictions are metric, with absolute scale, produced without relying on metadata such as camera intrinsics. That is a research claim rather than an implementation detail, and it is the claim that separates this project from the relative-depth family.

The model returns two things at once. The depth map is in metres, so a predicted 2.0 means two metres from the camera, and the focal length in pixels comes out of the same forward pass rather than from EXIF metadata. The README attributes this to state-of-the-art focal length estimation from a single image, which is what lets the network compensate for the fact that a monocular image does not otherwise tell you how far the camera was from the scene.

Sharpness is the other emphasis. The model synthesizes high-resolution depth maps with fine boundary detail, and the stated speed is a 2.25-megapixel depth map in 0.3 seconds on a standard GPU. The technical contributions listed are an efficient multi-scale vision transformer for dense prediction, a training protocol that mixes real and synthetic datasets, dedicated boundary metrics, and the focal length estimation. Four contributions rather than one, which is a sign the sharpness and the metric accuracy were treated as separate problems.

Installing the package and pulling the pretrained checkpoints

The README recommends a virtual environment and gives a miniconda route as the example. The install is an editable install of the repository itself, and the Python version is pinned to 3.9 by the project's own tooling:

bash
conda create -n depth-pro -y python=3.9
conda activate depth-pro

pip install -e .

`pyproject.toml` shows what that pulls in. The dependency list is short and unsurprising for a PyTorch project: `torch`, `torchvision`, `timm` for the vision transformer backbone, `numpy<2`, `pillow_heif`, and `matplotlib`. Two of those are worth a second look. The `numpy<2` ceiling is a real constraint rather than caution on paper, and `pillow_heif` is there so the loader can read HEIC files, which matters on a machine where most photos come off an iPhone.

Weights arrive through a shell script rather than the package manager, so the install and the checkpoint fetch are two separate steps:

bash
source get_pretrained_models.sh   # Files will be downloaded to `checkpoints` directory.

The script is at the repository root and writes into a `checkpoints` directory. Configuration in the same file is worth noticing for anyone planning to contribute: ruff runs with a line length of 100 and selects the `E`, `F`, `D` and `I` rules, pyright is configured for Python 3.9, and pytest is pointed at a `tests` path. The repository tree lists `src/`, `data/`, the script, `pyproject.toml`, `LICENSE`, `README.md` and `ACKNOWLEDGEMENTS.md`, with no `tests` directory visible among them.

Running one image from the command line or from Python

For a single image the helper script is the shortest path, and the README keeps it deliberately short:

bash
# Run prediction on a single image:
depth-pro-run -i ./data/example.jpg
# Run `depth-pro-run -h` for available options.

`depth-pro-run` is declared in `pyproject.toml` as `depth_pro.cli:run_main`, so it is a console script rather than something in the scripts directory. The example image at `data/example.jpg` is in the repository, which means the first run needs no input of your own.

The Python path is the one you want if the output goes into a pipeline. The sequence is: build model and transform together, switch to evaluation mode, load and preprocess, infer, then read the dictionary:

python
from PIL import Image
import depth_pro

# Load model and preprocessing transform
model, transform = depth_pro.create_model_and_transforms()
model.eval()

# Load and preprocess an image.
image, _, f_px = depth_pro.load_rgb(image_path)
image = transform(image)

# Run inference.
prediction = model.infer(image, f_px=f_px)
depth = prediction["depth"]  # Depth in [m].
focallength_px = prediction["focallength_px"]  # Focal length in pixels.

Two details in that snippet carry weight. `load_rgb` returns a three-tuple and the middle value is discarded, which is the signature of a function written to serve several model families rather than one. And the focal length in pixels is passed straight back into `infer` as `f_px`, so if you already know the focal length from your own calibration you can supply it instead of relying on the estimate.

The paper's numbers belong to a different set of weights

This is the single most important caveat in the repository and it appears early, in its own paragraph. The README states that the model here is a reference implementation which has been re-trained, and that its performance is close to the model reported in the paper but does not match it exactly. So the headline figures in the paper are not claims about the checkpoint you download.

The paper itself is Depth Pro: Sharp Monocular Metric Depth in Less Than a Second, by Aleksei Bochkovskii, Amael Delaunoy, Hugo Germain, Marcel Santos, Yichao Zhou, Stephan R. Richter and Vladlen Koltun, on arXiv as 2410.02073. The README gives an International Conference on Learning Representations bibtex entry with the year 2025 and asks that the paper be cited if the work is useful. It also says to check the paper for the complete list of references and datasets, which is a clear signal that the training data description lives outside this repository.

On the practical side of maintenance, the repository is not archived and the last push was on 2026-09-11, with 5,723 stars, 433 forks and 79 open issues. There are no published releases through the repository's releases listing, so the version to expect is the one in `pyproject.toml`, which reads `0.1`. Both the sample code and the model weights are released under the terms in the `LICENSE` file, and the acknowledgements file credits the open source work the codebase is built on.

Boundary metrics for the cases where average depth is not enough

Standard depth error metrics average over an image, which quietly hides exactly the failures this project claims to fix. If a hairline, a chair leg or a window frame gets smeared into the background, the mean error barely moves while the object is unusable for anything that cares about shape. Depth Pro ships metrics aimed at that case, in `eval/boundary_metrics.py`:

python
# for a depth-based dataset
boundary_f1 = SI_boundary_F1(predicted_depth, target_depth)

# for a mask-based dataset (image matting / segmentation)
boundary_recall = SI_boundary_Recall(predicted_depth, target_mask)

The two functions split by what kind of ground truth you have, which is the more useful thing about them. With a depth-based dataset such as a laser scan or a rendered scene, you get an F1 score computed against the target depth. With a mask-based dataset from image matting or segmentation, where you know where the object is but not how far away it is, the recall variant works instead.

That distinction matters if you are deciding whether this model fits your data. Plenty of projects have segmentation masks and no depth ground truth, and for those the boundary recall path is the only one of the two you can evaluate at all. Boundary accuracy is listed among the paper's contributions, so this file is part of the argument rather than an afterthought added later.

What the repository leaves to the paper and the reader

The README is short, which suits a reference implementation and leaves several questions open. There is no documented batch path. The CLI takes a single image with `-i` and points you at `-h` for the rest of the options, and the Python example processes one image, so anything larger is your loop to write.

Hardware is unstated beyond the single mention of a standard GPU. The 0.3 second figure for a 2.25-megapixel map is not tied to a specific card, and nothing in the repository says what CPU inference costs, whether a CUDA device is selected automatically, or what memory the checkpoint needs. The inference snippet calls `model.eval()` and stops there, which is the correct minimum but leaves device placement to you.

The evaluation story is also thinner than the model story. `eval/boundary_metrics.py` computes boundary scores, but the repository does not ship a script that runs a model over a benchmark, a table of numbers, or instructions for reproducing the paper's comparisons. Licensing is handled cleanly, with code and weights both under `LICENSE`, but the repository metadata carries no asserted license identifier, so anyone deploying this should read that file rather than infer terms from the repository page.

Editorial conclusion

Depth Pro is a good fit when a downstream consumer needs depth in metres rather than a relative ordering of pixels, because that is the property most monocular models give up. The reference implementation is small enough to read, with the network and preprocessing under `src/depth_pro`, a single shell script for weights, and a CLI that takes one image path. What it does not settle is how the retrained weights compare to the paper on your data, how to batch images, or what hardware the 0.3 second figure assumes, since none of that is written down here. Start with `depth-pro-run -i ./data/example.jpg`, then read `eval/boundary_metrics.py` if your application cares about edges rather than averages.

Frequently asked questions

How can I convert an image to a depth image?

With this repository, install the package, fetch the checkpoints with `get_pretrained_models.sh`, then point the helper script at an image: `depth-pro-run -i ./data/example.jpg`. From Python, call `depth_pro.create_model_and_transforms()`, load the image with `depth_pro.load_rgb`, and read `prediction["depth"]`, which is in metres.

What does Depth Pro return besides a depth map?

The same forward pass also estimates focal length in pixels, returned as `prediction["focallength_px"]`. Because the model predicts focal length from the image itself, it does not need camera intrinsics as input. If you already know the focal length, you can pass it into `infer` as `f_px`.

How do I get the pretrained model weights?

Run `source get_pretrained_models.sh` from the repository root after installing the package. The script downloads the checkpoints into a `checkpoints` directory. The package install and the weights are separate steps, so the second one is easy to forget.

Official sources

  1. apple-aiml-research/ml-depth-pro on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/apple-aiml-research-ml-depth-pro.svg)](https://hysenlabs.com/projects/apple-aiml-research-ml-depth-pro)