Open-source project
huawei-bayerlab/marigold-v2 avatar
huawei-bayerlab/marigold-v2

Marigold V2: single-step dense prediction from a frozen Qwen-Image-Edit backbone

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

903 stars74 forksPythonApache-2.0

At a glance

What is it?
Marigold V2 fine-tunes LoRA adapters on a frozen diffusion transformer to produce depth, normals and albedo in one step. The repository ships inference, evaluation and training code, but the install path assumes Linux, a CUDA GPU and a long first load.
Who is it for?
Adopt Marigold V2 if you need affine-invariant depth or normals on Linux with a CUDA GPU and at least 17 GB of free memory at 1024 squared, and you are willing to wait through the one-time 4-bit quantization of the DiT. Do not adopt it if you need metric depth straight out of the model, if you are on CPU or macOS, or if you want a pip-installable library rather than a cloned repository with a conda setup script.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 6 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Marigold V2 predicts and who the repository is aimed at

Marigold V2 is a family of models plus a fine-tuning protocol that turns a pretrained diffusion transformer into a single-step dense predictor. The README lists the supported outputs: depth, see-through depth, surface normals, albedo, and what it calls other dense modalities. The framing is explicit that the target user is not a large lab. Fine-tuning is described as taking less than a week on a single consumer GPU, which the README says puts it within reach of individual practitioners and small labs.

The depth checkpoints are all affine-invariant, meaning the output is correct up to an unknown scale and shift per image. That single sentence in the README is the most consequential thing on the page for anyone deciding whether to adopt it. If your downstream pipeline needs depth in metres, affine-invariant log depth is not that, and the README only mentions metric depth completion as an application the models unlock, not as the default output.

The released checkpoints also differ in parameterization, and the README is careful about the direction of each. Log and linear depth increase with distance; disparity decreases. Mixing them up in a visualization step produces an inverted image that looks plausible at a glance, which is a real source of confusion for a first-time user.

The frozen backbone, the LoRA adapters and why the text encoder never loads

Every checkpoint shares the same frozen Qwen-Image-Edit-2509 backbone. What is trained and stored is a `trainables.safetensors` file containing LoRA adapters and, where the recipe calls for it, the VAE decoder. That is the whole architecture in one line: the diffusion transformer stays put, and a small set of weights is fitted on top of it.

The savings show up in two places. First, training: the README describes a two-stage depth recipe where Stage 1 uses latent MSE plus L1, gradient and iREPA losses over 160k steps, and Stage 2 adds SinkLoss and fine-tunes the VAE decoder. Second, inference: precomputed Qwen text-prompt embeddings under `Marigold-V2/qwen_text_embeddings/` replace the text encoder, so the 7B text encoder is never loaded. That is a concrete memory and startup win, not a vague efficiency claim.

The cost is that the DiT is quantized to 4 bit when it is first loaded, and the README says this takes a few minutes. Every subsequent run in the same environment avoids that, but the first one does not. Plan for it rather than assuming the process has hung.

One gap worth flagging: the checkpoint table lists `depth/Uniform-layered` with the config column reading "not included". If you want that specific see-through linear depth variant, the repository does not ship the training config for it, so you cannot reproduce its training from this repo alone.

Installing Marigold V2 and running the example images

The README states the requirements plainly: Linux, Python 3.10, and a CUDA GPU. There is no pip path documented and no macOS or Windows instructions. Inference at 1024 squared needs about 17 GB of GPU memory; 2048 squared needs about 29 GB.

The quick start clones the repository and runs a conda setup script that creates an environment named `marigold-v2` with CUDA 12.8 wheels. The script takes an argument to change the CUDA wheel set, with `cu126` given as the example.

bash
git clone https://github.com/huawei-bayerlab/marigold-v2.git
cd marigold-v2
bash setup/setup_env.sh            # conda env "marigold-v2", CUDA 12.8 wheels; pass cu126 etc. to change
conda activate marigold-v2

Next, download the assets. The `--skip-datasets` flag keeps the download to the Qwen-Image-Edit-2509 backbone and the Marigold V2 checkpoints, which is what you want for inference only. Everything the repository downloads goes to `assets/`, or to `$DEPTH_ASSETS_DIR` if that variable is set.

bash
python scripts/download_assets.py --skip-datasets   # Qwen-Image-Edit-2509 + all Marigold V2 checkpoints

Then run inference over the bundled example images:

bash
python scripts/infer.py --image_dir assets/examples --output_dir output/examples

According to the README, depth predictions land in `output/examples/images/predictions_npy/` as float32 `.npy` files, and PNG visualizations land in `visualizations/depth_spectral/`. The `.npy` files are the ones to feed into a pipeline; the spectral PNGs are for looking at. If your output directory stays empty after a run, the first thing to check is whether the download step actually completed, since inference has nothing to load without the checkpoint.

Picking between the log, uniform and disparity depth checkpoints

The default checkpoint is `depth/Log-stage2`, described as the paper model: affine-invariant log depth, trained through Stage 1 and then Stage 2 with SinkLoss and a fine-tuned VAE decoder. If you have no reason to choose otherwise, that is the one the README points at.

The rest of the table is a mix of ablations and genuinely different capabilities, and the two are easy to confuse. `depth/Log-stage1` is the initialization for Stage 2 and for layered fine-tuning, trained with latent MSE, L1, gradient and iREPA losses over 160k steps. `depth/Uniform-base` and `depth/Disparity-base` are parameterization ablations trained with the Stage 1 recipe for 30k steps, so they exist to answer a research question rather than to be the best predictor.

The layered variants are the other category. `depth/Log-layered` is `Log-stage1` fine-tuned with SinkLoss on layer 8 of LayeredDepth-Syn, and the README says it predicts geometry behind glass. That is see-through depth, and it is a different task from ordinary monocular depth, not a better version of it. If your scene has no transparent surfaces, the layered checkpoints are the wrong pick.

The non-depth checkpoints follow the same pattern: `normals` outputs camera-space unit normals trained with an angular loss plus iREPA and SinkLoss, and `albedo` outputs linear RGB albedo in the range 0 to 1. Both were trained for 30k steps with a fine-tuned VAE decoder. Their configs live under `marigoldv2/experiments/20260728_qwen_normals/` and `marigoldv2/experiments/20260803_qwen_albedo/`, while the depth configs sit in `marigoldv2/experiments/20260316_qwen_depth/`.

Where Marigold V2 is the wrong tool

The affine-invariant output is the hard limitation. Depth values come back up to an unknown scale and shift per image. Two images from the same scene will not be on a common scale without an alignment step you write yourself, and nothing in the README suggests the repository provides one. The README does mention metric depth completion as an application the models unlock, but it does not present it as a documented workflow with a script.

The environment requirements rule out a large set of users outright. Linux, Python 3.10, CUDA. No CPU fallback is documented, and no macOS path. A 17 GB floor at 1024 squared is beyond most laptop GPUs, and 29 GB at 2048 squared exceeds a 24 GB card. The 4-bit quantization of the DiT reduces the footprint but does not make the model small.

The first load is slow in a way that is easy to misread as a failure. A few minutes of quantization before any prediction appears is normal per the README, but a job runner with a short timeout will kill it.

Finally, this is a research repository with a paper attached, not a maintained library. There are no retrieved releases, the version in `pyproject.toml` is `2.0.0`, and the dependency list pins exact versions across roughly two dozen packages including `diffusers==0.38.0`, `transformers==5.4.0` and `bitsandbytes==0.49.2`. Those pins will conflict with an existing environment, which is presumably why the setup script builds a fresh conda env rather than installing into yours.

How Marigold V2 differs from Marigold V1 and from depth-anything style models

The comparison the README invites is with Marigold V1, and it is specific. V1 used a linear depth parameterization; V2's default is affine-invariant log depth. The repository keeps `depth/Uniform-base` as an explicit parameterization ablation, described as Marigold V1 style, which means you can compare the two parameterizations within the same codebase rather than across two projects. That is a more useful comparison than most release notes offer.

The other axis is the backbone. V1 built on a Stable Diffusion style model; V2 uses Qwen-Image-Edit-2509 and replaces the text encoder with precomputed embeddings. The consequence is a 7B text encoder that never has to be loaded at inference or training time.

The comparison the README does not make is against discriminative depth models, the depth-anything family being the obvious reference point. The difference in approach is structural: those models regress depth directly in a single forward pass with a purpose-built network, while Marigold V2 repurposes a generative diffusion transformer and distills it down to a single step with LoRA adapters. The README's claims for the diffusion route are about detail fidelity, specifically sharp edges, fur and hair-thin details. If your scenes are indoor rooms with flat walls, that advantage is worth less than the memory cost of loading a quantized DiT.

Licence, maintenance and the cost of tracking this repository

The repository is Apache-2.0, with a `LICENSE` file and a separate `NOTICE` file at the top level. The `pyproject.toml` declares the licence by file reference rather than as an SPDX string. Apache-2.0 permits commercial use and modification, and it includes a patent grant, but it also carries notice and attribution obligations. The `NOTICE` file exists for that reason and should travel with any redistribution. This is a description of the licence text, not legal advice; if you are shipping a product, have someone read the actual terms.

The weights are a separate question from the code. The checkpoints live at `huawei-bayerlab/marigold-v2-0` on Hugging Face, and the frozen backbone is Qwen-Image-Edit-2509, which is a third-party model with its own terms. The repository's Apache-2.0 licence covers the code in this repo; it does not automatically relicense the backbone or the hosted weights.

Maintenance looks current rather than dormant. The last push was on 2026-09-13, four days before this writing, and the repository is not archived. The news entries track a real sequence: initial release on 2026-09-08, ModelScope mirroring on 2026-09-13, and a note that the work is to appear in ACM Transactions on Graphics 45(6) at SIGGRAPH Asia 2026. There are no retrieved releases, so upgrades mean pulling from `main`.

That is the upgrade cost. Exact pins in `pyproject.toml` mean a `git pull` can move `diffusers`, `transformers` or `bitsandbytes` under you, and the setup script's CUDA wheel argument is the escape hatch when a wheel set stops matching your driver. The `AGENTS.md` and `.pre-commit-config.yaml` at the top level suggest the repository expects contributions, and `pyproject.toml` configures ruff with a line length of 88 against Python 3.10.

Editorial conclusion

Adopt Marigold V2 if you need affine-invariant depth or normals on Linux with a CUDA GPU and at least 17 GB of free memory at 1024 squared, and you are willing to wait through the one-time 4-bit quantization of the DiT. Do not adopt it if you need metric depth straight out of the model, if you are on CPU or macOS, or if you want a pip-installable library rather than a cloned repository with a conda setup script. Verify three things before committing: that your GPU has the memory the README quotes for your target resolution, that the checkpoint you pick matches the depth parameterization you actually want (log, linear or disparity, since they are not interchangeable), and that your use of the Qwen-Image-Edit-2509 backbone and the released LoRA adapters fits the Apache-2.0 terms you are accepting.

Frequently asked questions

How much GPU memory does Marigold V2 need for inference?

The README states that inference at 1024 squared needs about 17 GB of GPU memory and 2048 squared about 29 GB. A CUDA GPU on Linux is required, and no CPU or macOS path is documented.

Does Marigold V2 output metric depth?

The depth checkpoints are affine-invariant, so values are correct only up to an unknown scale and shift per image. The README lists metric depth completion as an application the models unlock, but it does not document it as a default output.

Why does the first Marigold V2 run take several minutes before producing output?

The README says the DiT is quantized to 4 bit when it is first loaded, which takes a few minutes. Later runs in the same environment reuse that quantized model.

Official sources

  1. huawei-bayerlab/marigold-v2 on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/huawei-bayerlab-marigold-v2.svg)](https://hysenlabs.com/projects/huawei-bayerlab-marigold-v2)