Open-source project
prs-eth/Marigold avatar
prs-eth/Marigold

Marigold: Diffusion Models Repurposed for Monocular Depth Estimation

[CVPR 2024 - Oral, Best Paper Award Candidate] Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation

3,236 stars212 forksPythonApache-2.0

At a glance

What is it?
Marigold adapts pretrained latent diffusion models such as Stable Diffusion into dense prediction models for depth, surface normals and intrinsic decomposition. It is research code from the CVPR 2024 paper, and its cost profile and dependency set matter as much as its accuracy.
Who is it for?
Adopt Marigold if you need zero-shot monocular depth, surface normals or intrinsic decomposition and you can run PyTorch 2.4.1 on a GPU, or if you want to study how a latent diffusion model is adapted to a dense regression task. Do not adopt it if you need CPU inference, a stable versioned API, or a maintained product with support commitments; the repository is research code whose last push was on 2026-09-06.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 28 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Marigold actually solves, and for whom

Monocular depth estimation means recovering a depth map from a single RGB image, with no stereo pair and no LiDAR. The README frames Marigold's core principle as leveraging the visual knowledge stored in modern generative image models. Instead of training a depth network from scratch on a large labelled dataset, the project starts from a pretrained latent diffusion model such as Stable Diffusion and fine-tunes it with synthetic data, so that the denoising process produces depth rather than an image. The README states that this requires minimal modification of the pretrained architecture, trains with small synthetic datasets on a single GPU over a few days, and demonstrates state-of-the-art zero-shot generalization. That last property is the reason to care: a model that transfers to unseen data without task-specific fine-tuning is useful when you have images but no ground-truth depth for them. The audience is therefore narrow and specific. It is researchers and engineers who need dense geometric predictions on arbitrary photographs, and who are willing to run a diffusion sampling loop to get them. It is not a library for a mobile app that needs a depth map in a few milliseconds.

How a generative model becomes a depth predictor

The mechanism follows the standard latent diffusion pipeline, with the output reinterpreted. In text-to-image diffusion, the denoising network is conditioned on a text embedding and predicts noise that is removed over a schedule of timesteps, ending in an image latent that a VAE decoder turns into pixels. Marigold keeps that schedule and that network, and changes what the latent represents: the target is a depth map rather than a photograph, and the conditioning signal is the input image. The fine-tuning protocol, described in the CVPR 2024 paper, teaches the model to denoise toward depth, and the README notes that the v1.1 depth checkpoint was trained with updated noise scheduler settings (zero-SNR and trailing timestamps) plus augmentations. Because the output is produced by iterative denoising, a single prediction is a sample from a distribution rather than a deterministic answer. The repository and the paper describe aggregating multiple samples to obtain a final depth map, which is why inference cost scales with the number of denoising steps and the number of samples you average. The same recipe is extended in the follow-up work to surface normals and to intrinsic decomposition, where the README lists checkpoints predicting Albedo, diffuse Shading and non-diffuse Residual, and a second set predicting Albedo, Roughness and Metallicity.

Installing Marigold and running a first inference

The README offers several ways to interact with the project: hosted demos on Hugging Face Spaces, the pipelines integrated into the diffusers library, and running the demo locally. The local route requires a GPU. The repository root carries three requirement files, requirements.txt, requirements+.txt and requirements++.txt, which suggests a base install plus optional extras, but the README does not document what the plus variants add, so treat that as unverified. The base file pins the core stack:

bash
pip install -r requirements.txt

That installs accelerate, diffusers, matplotlib, scipy, torch 2.4.1, torchvision 0.19.1 and transformers. The exact pins matter: torch==2.4.1 and torchvision==0.19.1 are fixed versions, so installing into an environment that already has a different PyTorch build will either fail resolution or force a downgrade. The README states that Marigold pipelines were merged into diffusers core starting with v0.28.0, and points to the diffusers usage page for the pipeline API. The repository's own usage section describes running the demo locally rather than printing the command, so check the script directory for the entry point before assuming a flag. A no-install alternative exists: the Colab notebook linked from the README, and the Hugging Face Spaces demos for depth, normals and image intrinsics.

Where Marigold is the wrong tool

The first constraint is hardware. The README states that running the demo locally requires a GPU, and the diffusion sampling loop is the reason: every prediction is produced by iterating a denoising network, and aggregating multiple samples multiplies that cost. For a pipeline that needs depth at video frame rates, or on a CPU-only machine, this is a mismatch, and no amount of tuning the checkpoint changes the shape of the computation. The second constraint is output semantics. Diffusion output is stochastic, and the depth values are relative rather than metric unless you apply your own scale and shift alignment against a reference; the README does not present Marigold as a metric depth sensor replacement. If your application needs absolute distances in metres, you must supply the alignment step yourself. The third constraint is packaging. This is research code with no retrieved releases, a version-pinned dependency set, and a repository whose last push was on 2026-09-06. There is no documented deprecation policy, no rollback guidance in the README, and no compatibility statement for future PyTorch versions. Teams that need a frozen, supported artifact should look elsewhere or vendor a specific commit and accept the maintenance burden.

Marigold against Depth Anything and MiDaS

The obvious alternatives are discriminative depth models such as Depth Anything and MiDaS. The difference is architectural, not just numerical. A discriminative model is trained with a regression or ranking loss to map an image to a depth map in one forward pass; inference is a single network evaluation, it is deterministic, and it runs comfortably on modest hardware. Marigold instead inherits a generative prior. Its forward pass is a multi-step denoising process, its output is a sample, and its claimed advantage is zero-shot generalization that comes from the visual knowledge in Stable Diffusion rather than from depth labels. That trade is the whole story: you pay in latency and hardware for a different generalization profile and for the ability to produce an ensemble of plausible depth maps instead of one answer. The same generative framing is what lets the project extend to surface normals and intrinsic decomposition with the same recipe, which a single-purpose depth regressor does not do. If you only need a depth map quickly and deterministically, a discriminative model is the simpler choice; if you need the generative prior or the multi-modality, Marigold is the one built for that.

Maintenance, dependencies and licence split

The repository is not archived, and the last push was on 2026-09-06, which is recent enough that the code is being touched, but the project publishes no retrieved releases, so there is no version tag to pin against and no changelog to read before upgrading. Upgrades therefore mean tracking the main branch, and the pinned requirements mean a PyTorch bump is a project-level decision rather than a routine one. The dependency surface is small but heavy: accelerate, diffusers, transformers, torch, torchvision, scipy and matplotlib. Two licence files sit at the root. LICENSE.txt is the Apache License, Version 2.0, which per the news section was adopted on 2023-12-19. LICENSE-MODEL.txt is separate, and the existence of a distinct model licence is the detail that matters for anyone shipping a product: the code and the weights are not necessarily under the same terms. The README does not explain the scope of LICENSE-MODEL.txt, so read both files and, if the distinction affects your distribution, take your own advice on it rather than treating this as legal guidance.

Editorial conclusion

Adopt Marigold if you need zero-shot monocular depth, surface normals or intrinsic decomposition and you can run PyTorch 2.4.1 on a GPU, or if you want to study how a latent diffusion model is adapted to a dense regression task. Do not adopt it if you need CPU inference, a stable versioned API, or a maintained product with support commitments; the repository is research code whose last push was on 2026-09-06. Before committing, check that the pinned torch==2.4.1 and torchvision==0.19.1 versions fit your existing environment, and read LICENSE-MODEL.txt separately from the Apache-2.0 code licence, since the two files cover different things.

Frequently asked questions

What is Marigold used for?

Marigold is a family of conditional generative models for dense image analysis. The README lists monocular depth estimation, surface normal prediction and intrinsic decomposition, all built by adapting pretrained latent diffusion models such as Stable Diffusion.

Does Marigold need a GPU to run?

The README states that running the demo locally requires a GPU. The Hugging Face Spaces demos and the Colab notebook are the alternatives that avoid a local GPU setup.

How do I install Marigold?

The repository provides requirements.txt, requirements+.txt and requirements++.txt at the root. Installing requirements.txt pulls accelerate, diffusers, matplotlib, scipy, torch 2.4.1, torchvision 0.19.1 and transformers. The README does not document what the plus variants add.

Can I use Marigold without cloning the repository?

Yes. The README states that Marigold pipelines were merged into the diffusers core starting with v0.28.0, and links a diffusers usage page. There are also hosted demos on Hugging Face Spaces for depth, normals and image intrinsics.

Which pretrained model does Marigold build on?

The README describes Marigold as derived from Stable Diffusion and fine-tuned with synthetic data. The repository links checkpoints on Hugging Face for depth, normals and the two intrinsic decomposition variants.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. Project website
  4. prs-eth/Marigold on GitHub
  5. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/prs-eth-marigold.svg)](https://hysenlabs.com/projects/prs-eth-marigold)