Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation Marigold V2, a diffusion transformer-based model for monocular depth estimation, revisits the approach to improve generalization to out-of-distribution inputs, addressing a persistent challenge in the field with applications in scene reconstruction, computational photography, and robotics. The model builds on prior work to enhance robustness across diverse real-world scenarios. Monocular depth estimation is a ubiquitous yet highly ill-posed computer vision task, with downstream applications in scene reconstruction, computational photography, and robotics, among others. Despite the field's maturity, recent models still struggle to generalize to out-of-distribution inputs an