Marigold V2
Revisiting Diffusion Transformers for Monocular Depth Estimation
ACM Transactions on Graphics (SIGGRAPH Asia 2026)
Igor Pavlovic1,2,,†, Thiemo Wandel2,, Anton Obukhov2,§
Luca Bartolomei3, Andrey Davydov2, Fabio Tosi3, Matteo Poggi3, Sabine Süsstrunk1, Dengxin Dai2
1EPFL · 2HUAWEI Bayer Lab · 3University of Bologna · *Equal contribution · †Internship · §Project lead
Marigold V2 is a family of models and a cost-effective fine-tuning protocol that repurposes a pretrained diffusion transformer into single-step dense predictors: depth, see-through depth, surface normals, albedo, and other dense modalities. Fine-tuning takes less than a week on a single consumer GPU, within reach of individual practitioners and small labs, and the results are state of the art, faithfully reproducing sharp edges, fur, and hair-thin details. The same models also unlock applications such as metric depth completion.

This repository contains inference, evaluation, and training code, plus the configs of every released checkpoint.
- News
- Quick start
- Checkpoints
- Inference
- Evaluation
- Training
- Extending to a new task
- Repository layout
- Checklist
- Troubleshooting
- Contributing
- Citation
News
2026-12: To appear in ACM Transactions on Graphics 45(6) and to be presented at SIGGRAPH Asia 2026.
2026-09-13: Mirrored on ModelScope:
2026-09-08: Initial release: inference, evaluation, and training code, the released checkpoints, and the demo.
Quick start
Requirements: Linux, Python 3.10, a CUDA GPU. Inference at 1024² needs about 17 GB of GPU memory and 2048² about 29 GB; the DiT is quantized to 4 bit when it is first loaded, which takes a few minutes.
git clone https://github.com/huawei-bayerlab/marigold-v2.git
cd marigold-v2
bash setup/setup_env.sh # conda env "marigold-v2", CUDA 12.8 wheels; pass cu126 etc. to change
conda activate marigold-v2
python scripts/download_assets.py --skip-datasets # Qwen-Image-Edit-2509 + all Marigold V2 checkpoints
python scripts/infer.py --image_dir assets/examples --output_dir output/examples
Depth predictions land in output/examples/images/predictions_npy/ as float32
.npy files and in visualizations/depth_spectral/ as PNGs. Everything the
repository downloads goes to assets/ (or $DEPTH_ASSETS_DIR).
Checkpoints
All checkpoints share the frozen Qwen-Image-Edit-2509 backbone and are stored as
trainables.safetensors (LoRA adapters and, where trained, the VAE decoder) in
huawei-bayerlab/marigold-v2-0. The
depth checkpoints differ in how depth is parameterized before it enters the VAE
and in the training stage.
| Checkpoint | Output | Training | Config |
|---|---|---|---|
depth/Log-stage2 (default) |
affine-invariant log depth | Stage 1 → Stage 2 (SinkLoss, VAE decoder fine-tuned). The paper model. | training_relative_log_depth_config_stage2.yaml |
depth/Log-stage1 |
affine-invariant log depth | Stage 1 only: latent MSE + L1 + gradient + DINOv3 spatial perceptual loss, 160k steps. Initialization for Stage 2 and layered fine-tuning. | training_relative_log_depth_config.yaml |
depth/Log-layered |
see-through log depth | Log-stage1 fine-tuned with SinkLoss on layer 8 of LayeredDepth-Syn: predicts geometry behind glass. |
training_relative_log_depth_layered_config.yaml |
depth/Uniform-base |
affine-invariant linear depth (Marigold V1 style) | Stage 1 recipe, 30k steps. Parameterization ablation. | training_relative_config.yaml |
depth/Disparity-base |
affine-invariant inverse depth | Stage 1 recipe with VAE decoder fine-tuning, 30k steps. Parameterization ablation. | training_relative_disp_config.yaml |
depth/Disparity-layered |
see-through inverse depth | Stage 1 recipe on layer 8 of LayeredDepth-Syn. | training_relative_disp_layered_config.yaml |
depth/Uniform-layered |
see-through linear depth | LayeredDepth-Syn variant of Uniform-base. |
not included |
normals |
camera-space unit normals | angular loss + DINOv3 spatial perceptual loss + SinkLoss, VAE decoder fine-tuned, 30k steps | training_normals.yaml |
albedo |
linear RGB albedo in [0, 1] | L1 + DINOv3 spatial perceptual loss, VAE decoder fine-tuned, 30k steps | training_albedo.yaml |
Depth configs live in marigoldv2/experiments/20260316_qwen_depth/, the normals
and albedo configs in marigoldv2/experiments/20260728_qwen_normals/ and
marigoldv2/experiments/20260803_qwen_albedo/. Log and linear depth increase
with distance, disparity decreases; the values are affine-invariant, i.e. up to
an unknown scale and shift per image.
The precomputed Qwen text-prompt embeddings under
Marigold-V2/qwen_text_embeddings/ replace the text encoder at inference and
training time, so the 7B text encoder is never loaded.
Inference
python scripts/infer.py --modality depth --image_dir /path/to/images --output_dir output/depth
python scripts/infer.py --modality normals --image_dir /path/to/images --output_dir output/normals
python scripts/infer.py --modality albedo --image_dir /path/to/images --output_dir output/albedo
| Flag | Meaning |
|---|---|
--modality |
depth (default), normals, or albedo; selects the default checkpoint and prompt embedding |
--checkpoint |
checkpoint directory or Hugging Face repo[/subfolder], e.g. assets/checkpoints/Marigold-V2/depth/Log-layered |
--width, --height |
run at a fixed resolution instead of the native one (rounded up to a multiple of 16) |
--seed |
seed for the VAE encoder sampling (default 2025) |
Outputs mirror the input folder structure under /images/:
predictions_npy/*.npy (depth: [H, W]; normals and albedo: [3, H, W]) and
visualizations//*.png. Predictions are resized back to the input
resolution.
Evaluation
Download the depth and surface-normal benchmarks separately using the instructions
below; scripts/download_assets.py does not download these evaluation datasets.
Run these commands from the repository root. They use assets/ by default, or
$DEPTH_ASSETS_DIR when set, matching the evaluation launchers.
For depth estimation, download the Marigold evaluation datasets and extract each tar archive into its containing directory:
(
set -e
mkdir -p "${DEPTH_ASSETS_DIR:-assets}/datasets/marigold_depth_eval"
cd "${DEPTH_ASSETS_DIR:-assets}/datasets/marigold_depth_eval"
wget -r -np -nH --cut-dirs=4 -R "index.html*" -P . https://share.phys.ethz.ch/~pf/bingkedata/marigold/evaluation_dataset/
find . -type f -name '*.tar' -print0 | while IFS= read -r -d '' archive; do
tar -xf "$archive" -C "$(dirname "$archive")"
done
)
For depth completion, download and preprocess the iBims-1, NYUv2, KITTI DC, and DDAD datasets:
bash scripts/marigold_dc/download_and_preprocess_depth_completion.sh
For normal estimation, create the destination directory and download the official Marigold evaluation archive:
NORMALS_ROOT="${DEPTH_ASSETS_DIR:-assets}/datasets/marigold_normals_eval"
mkdir -p "$NORMALS_ROOT"
(
set -e
cd "$NORMALS_ROOT"
wget -O evaluation_dataset.zip \
https://share.phys.ethz.ch/~pf/bingkedata/marigold/marigold_normals/evaluation_dataset.zip
unzip -n evaluation_dataset.zip
)
Download the preprocessed Sintel benchmark separately into the same directory, then extract it:
(
set -e
cd "$NORMALS_ROOT"
wget -O sintel.zip \
https://share.phys.ethz.ch/~pf/bingkedata/marigold/marigold_normals/sintel.zip
unzip -n sintel.zip
)
Both archives should extract directly under marigold_normals_eval/; do not
add another enclosing directory. The dataset directories must sit directly
under the benchmark root as shown below (downloaded archives can remain):
assets/ # or $DEPTH_ASSETS_DIR
└── datasets/
├── marigold_depth_eval/
│ ├── diode/
│ ├── eth3d/
│ ├── kitti/
│ ├── nyuv2/
│ └── scannet/
└── marigold_normals_eval/
├── ibims/ibims/
├── nyuv2/test/
├── scannet/
└── sintel/
Each launcher below writes predictions and metrics under output/eval_runs/ and
prints where. Add MAX_SAMPLES as the second argument to any launcher for a
quick partial run.
Zero-shot depth on NYUv2, KITTI, ETH3D, ScanNet, and DIODE with the Pixel-Perfect Depth protocol (RANSAC alignment in log space, native resolution; ETH3D is upsampled to 2048×1360 before scoring):
bash evaluation/depth/run_infer_and_eval.sh # Log-stage2, all datasets
bash evaluation/depth/run_infer_and_eval.sh "" 1 evaluation/config/inference_depth.yaml kitti,eth3d
bash evaluation/depth/run_eval_from_preds.sh output/eval_runs/depth_/predictions # rescore existing predictions
| NYUv2 | KITTI | ETH3D | ScanNet | DIODE | |
|---|---|---|---|---|---|
| AbsRel ↓ / δ1 ↑ | 3.6 / 98.0 | 5.4 / 97.4 | 2.8 / 99.2 | 3.7 / 97.9 | 5.2 / 97.1 |
Soft Edge Error on the Hypersim test split at native resolution. Needs the Hypersim training data:
python scripts/download_assets.py
bash evaluation/depth_see/run_hypersim_origres_edge_eval.sh # SEE_1,3,5,7; paper: 0.352 / 0.333 / 0.320 for k = 3, 5, 7
See-through depth on LayeredDepth-Syn validation layer 8.
python scripts/download_assets.py --include-layereddepth-syn
bash evaluation/depth_see_through/run_layereddepth_l8_eval.sh # Log-stage2 and Log-layered
| Checkpoint | AbsRel ↓ | δ1 ↑ |
|---|---|---|
| Log-stage2 | 13.66 | 83.96 |
| Log-layered | 8.17 | 92.65 |
Surface normals on NYUv2, ScanNet, iBims-1, and Sintel:
bash evaluation/normals/run_qwen_normals_infer_and_eval.sh # mean angular error / % within 11.25°
| NYUv2 | ScanNet | iBims-1 | Sintel | |
|---|---|---|---|---|
| mean err ↓ / 11.25° ↑ | 16.6 / 61.2 | 14.1 / 67.4 | 15.9 / 70.9 | 28.7 / 27.6 |
The Hypersim normals edge metric (SAEE) needs the Hypersim normals dataset:
python scripts/download_assets.py
bash scripts/hypersim_normals/download_and_preprocess_hypersim_normals.sh
bash evaluation/normals_saee/run_hypersim_normals_origres_edge_eval.sh
Albedo on the Hypersim IID test split needs the Hypersim albedo dataset; expected PSNR 20.78, SSIM 0.811, LPIPS 0.195:
python scripts/download_assets.py
bash scripts/hypersim_albedo/download_and_preprocess_hypersim_albedo.sh
bash evaluation/albedo/run_qwen_albedo_infer.sh assets/checkpoints/Marigold-V2/albedo output/eval_runs/albedo
bash evaluation/albedo/run_qwen_albedo_eval.sh output/eval_runs/albedo/hypersim_test_albedo_qwen_native output/eval_runs/albedo/metrics
Depth completion on iBims-1, NYUv2, KITTI-DC, and DDAD. The guide in evaluation/depth_completion/REPRODUCING_DEPTH_COMPLETION.md covers the data, the protocol, and the per-sample reference results. Cells are grouped by the GPU memory they need: 8 of the 11 fit a 32 GB card, the other 3 need 48 to 80 GB.
GPUS=0,1,2,3 bash evaluation/depth_completion/reproduce_table.sh subset 32gb # check the setup against the reference results, ~45 min on 4 GPUs
bash evaluation/depth_completion/reproduce_table.sh full 32gb # the 8 cells that fit a 32 GB card
bash evaluation/depth_completion/reproduce_table.sh full 80gb # the remaining 3 cells
| MAE (m) ↓ | iBims-1 | NYUv2 | KITTI-DC | DDAD |
|---|---|---|---|---|
| Marigold-V2-DC (LoRA) | 0.042 | 0.045 | 0.349 | 1.549 |
| + high-res inference | 0.034 | 0.044 | 0.340 | 1.465 |
| + tiled local adaptation | 0.030 | 0.044 | 0.318 | 1.226 |
Training
All released models were trained on a single 32 GB GPU with batch size 1. Stage 1 of the depth model (160k steps) takes about five days, every other config (30k steps) about one day.
Data. Depth trains on Hypersim and Virtual KITTI 2 as repackaged for
Marigold V1. The DINOv3 spatial perceptual loss needs the DINOv3 checkpoint,
which is gated on Hugging Face (accept the
terms on the model page,
then hf auth login):
python scripts/download_assets.py --include-dinov3 # checkpoints, DINOv3, and training data
python scripts/download_assets.py --include-layereddepth-syn --skip-checkpoints # optional; downloads and prepares LayeredDepth-Syn
bash scripts/hypersim_normals/download_and_preprocess_hypersim_normals.sh # optional, ~1.2 TB, for normals
bash scripts/hypersim_albedo/download_and_preprocess_hypersim_albedo.sh # optional, ~500 GB, for albedo
The two Hypersim scripts run the Marigold V1.1 preprocessors scene by scene and
resume after interruption; see their READMEs under scripts/.
Reproducing the depth model.
# Stage 1: DINOv3 spatial perceptual + pixel losses on Hypersim + vKITTI (160k steps)
python marigoldv2/script/train/train.py \
--config marigoldv2/experiments/20260316_qwen_depth/training_relative_log_depth_config.yaml \
--output_dir output/train_runs --no_wandb
# Stage 2: SinkLoss with VAE decoder fine-tuning, initialized from Stage 1 (30k steps)
python marigoldv2/script/train/train.py \
--config marigoldv2/experiments/20260316_qwen_depth/training_relative_log_depth_config_stage2.yaml \
--output_dir output/train_runs --no_wandb
The Stage 2 config initializes from assets/checkpoints/Marigold-V2/depth/Log-stage1;
point paths.env_paths.checkpoint at your own Stage 1 run to chain them. Any
other config in the table above trains the same way.
A run writes to //: config.yaml, a code snapshot,
TensorBoard logs, validation visualizations, and checkpoint/checkpoint-{best,latest,}/.
Each checkpoint folder holds trainables.safetensors, which is exactly the
format of the released checkpoints, so it can be passed to scripts/infer.py --checkpoint or to any evaluation launcher directly. Resume an interrupted run
with --resume_run /checkpoint/checkpoint-latest. Weights & Biases logging
is on unless --no_wandb is given; add --add_datetime_prefix to keep several
runs of one config apart.
Multi-GPU training uses accelerate launch --multi_gpu marigoldv2/script/train/train.py ...;
each GPU keeps the configured per-process batch size, increasing the effective
global batch size. optimization.gpu_scaling divides total and periodic step
counts by the number of processes, keeping the number of seen samples approximately
constant.
Extending to a new task
The training framework is a small registry of building blocks composed by YAML. A training config declares:
base_config: dataset definition files frommarigoldv2/config/datasets/. A dataset is amanifest_graph(how to list samples), atransformlist (how to load and augment one sample), and, for validation sets,validation_steps.register_modules: Python modules whose@register(...)classes the config may reference.network_components: loaders that put models into the registry, here the Qwen VAE and the quantized DiT with LoRA.network_graph: an ordered list of steps that read from and write to thebatchdict, e.g. encode RGB → one DiT step → decode → pick channels.loss_graph: losses that read prediction and target keys from the batch.optimization: schedule, quantization, and LoRA settings.
The surface-normal experiment is the smallest complete example of a new task,
in marigoldv2/experiments/20260728_qwen_normals/: data.py adds a manifest
loader and normal-aware transforms, network_graph.py adds one output step
that turns the decoded RGB into unit normals, loss.py adds the angular loss,
validation.py adds metrics and visualizations, and training_normals.yaml
wires them together with the shared depth blocks. The albedo experiment follows
the same pattern.
To add a task: copy a dataset YAML and adapt its manifest loader and
transforms; write the output step, loss, and validation steps in a new
experiment folder; list the new modules under register_modules; then run a
short training with --output_dir output/debug. For inference on plain image
folders, add an output adapter to marigoldv2/validation/folder_steps.py and a
MODALITIES entry in scripts/infer.py.
Repository layout
setup/setup_env.sh creates the conda environment and installs the package
scripts/ infer.py, download_assets.py, Marigold-DC and Hypersim dataset builders
marigoldv2/ training framework: core registry, datasets, losses, trainer, validation
experiments/ one folder per released model family with its configs and task-specific modules
config/datasets/ dataset definitions shared by the training configs
script/train/train.py training entry point
evaluation/ benchmark launchers (depth, depth_completion, depth_see, depth_see_through, normals, normals_saee, albedo), their configs, data splits
evaluation/src/ Marigold V1 benchmark datasets and metrics
assets/ downloads (git-ignored) plus tracked example images
Checklist
- [x] Depth completion code
- [x] See-through evaluation
- [ ] Diffusers integration
- [x] ComfyUI plugin
Troubleshooting
| Problem | Solution |
|---|---|
CUDA out of memory |
Inference needs about 17 GB at 1024² and 29 GB at 2048². Reduce --width and --height; both must stay multiples of 16. |
| The same image gives slightly different predictions | The VAE encoder samples its latent. Pass --seed to make runs reproducible. |
| A config cannot find a dataset or a checkpoint | Point DEPTH_ASSETS_DIR at your assets folder. Download checkpoints and training data with python scripts/download_assets.py; for depth, depth completion, and normals benchmarks, follow Evaluation. |
Contributing
Bug reports are welcome. Before sending a pull request, please open an issue to discuss the change with the maintainers.
AGENTS.md documents the repository layout, the registry the configs are built
on, and the conventions to keep.
Citation
@article{pavlovic2026marigoldv2,
author = {Pavlovic, Igor and Wandel, Thiemo and Obukhov, Anton and Bartolomei, Luca and Davydov, Andrey and Tosi, Fabio and Poggi, Matteo and S{\"u}sstrunk, Sabine and Dai, Dengxin},
title = {Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation},
year = {2026},
issue_date = {December 2026},
publisher = {Association for Computing Machinery},
volume = {45},
number = {6},
url = {https://doi.org/10.1145/3842528},
doi = {10.1145/3842528},
journal = {ACM Trans. Graph.},
month = dec,
articleno = {204},
numpages = {14}
}
Acknowledgements
Marigold V2 builds on Marigold V1, whose evaluation code and dataset preprocessing are reused here, and on Qwen-Image-Edit-2509. The depth evaluation protocol follows Pixel-Perfect Depth, the Sintel normals protocol follows Lotus-2, and see-through depth uses LayeredDepth-Syn.
License
Code and models are released under the Apache License, Version 2.0 (see LICENSE and NOTICE). Qwen-Image-Edit-2509 and the datasets keep their own licenses.