Pith. sign in

REVIEW 11 cited by

Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.09358 v1 pith:NEJ2W45L submitted 2025-05-14 cs.CV cs.LG

Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis

classification cs.CV cs.LG
keywords modelsimagediffusiondatasetslatentlearningmarigoldpretrained
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The success of deep learning in computer vision over the past decade has hinged on large labeled datasets and strong pretrained models. In data-scarce settings, the quality of these pretrained models becomes crucial for effective transfer learning. Image classification and self-supervised learning have traditionally been the primary methods for pretraining CNNs and transformer-based architectures. Recently, the rise of text-to-image generative models, particularly those using denoising diffusion in a latent space, has introduced a new class of foundational models trained on massive, captioned image datasets. These models' ability to generate realistic images of unseen content suggests they possess a deep understanding of the visual world. In this work, we present Marigold, a family of conditional generative models and a fine-tuning protocol that extracts the knowledge from pretrained latent diffusion models like Stable Diffusion and adapts them for dense image analysis tasks, including monocular depth estimation, surface normals prediction, and intrinsic decomposition. Marigold requires minimal modification of the pre-trained latent diffusion model's architecture, trains with small synthetic datasets on a single GPU over a few days, and demonstrates state-of-the-art zero-shot generalization. Project page: https://marigoldcomputervision.github.io

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MUSE: Unlocking Timestep as Native Task Steering for One-Step Dense Prediction

    cs.CV 2026-06 unverdicted novelty 7.0

    MUSE shows that the native timestep embedding in diffusion models acts as a parameter-free steering signal for multi-task monocular depth and normal estimation via manifold decoupling in latent space.

  2. Monocular Depth Estimation via Neural Network with Learnable Algebraic Group and Ring Structures

    cs.CV 2026-04 unverdicted novelty 7.0

    LAGRNet embeds learnable algebraic group, ring, and sheaf structures into a neural network to improve accuracy and generalization in monocular depth estimation.

  3. CDPR: Cross-modal Diffusion with Polarization for Reliable Monocular Depth Estimation

    cs.CV 2026-04 unverdicted novelty 7.0

    CDPR integrates polarization priors into a diffusion-based monocular depth estimator via shared latent space and adaptive gating, outperforming RGB-only methods in challenging scenes.

  4. UniDAC: Universal Metric Depth Estimation for Any Camera

    cs.CV 2026-03 unverdicted novelty 7.0

    UniDAC achieves universal metric depth estimation across camera types by decoupling relative depth prediction from spatially varying scale estimation using a depth-guided module and distortion-aware positional embedding.

  5. A Unified and Controllable Framework for Layered Image Generation with Visual Effects

    cs.CV 2026-01 unverdicted novelty 7.0

    LASAGNA produces layered images with integrated visual effects in a single pass, enabling drift-free edits via alpha compositing while releasing a 48K dataset and a 242-sample benchmark.

  6. StereoSpace: Depth-Free Synthesis of Stereo Geometry via End-to-End Diffusion in a Canonical Space

    cs.CV 2025-12 unverdicted novelty 7.0

    A viewpoint-conditioned diffusion model generates stereo image pairs from monocular input in a canonical rectified space without using depth or explicit warping.

  7. In Depth We Trust: Reliable Monocular Depth Supervision for Gaussian Splatting

    cs.CV 2026-04 unverdicted novelty 6.0

    A selective regularization framework lets scale-ambiguous monocular depth priors improve Gaussian Splatting geometry and rendering by isolating and supervising only ill-posed regions.

  8. UniSER: A Foundation Model for Unified Soft Effects Removal

    cs.CV 2025-11 unverdicted novelty 6.0

    UniSER is a unified diffusion transformer foundation model that removes diverse soft image degradations by training on a large curated dataset of semi-transparent occlusions with fine-grained controls.

  9. Towards Consistent Video Geometry Estimation

    cs.CV 2026-05 unverdicted novelty 5.0

    ViGeo is a feed-forward transformer for video geometry that introduces dynamic chunking attention and a completion-based data refinement framework to achieve SOTA on depth, normals, and point map estimation.

  10. DealMaTe: Multi-Dimensional Material Transfer via Diffusion Transformer

    cs.GR 2026-05 unverdicted novelty 5.0

    DealMaTe proposes a simplified diffusion framework for material transfer that injects multi-dimensional 3D conditions via Multi-Dim 3D Shader LoRA and Shader Causal Mutual Attention with KV caching.

  11. LTM: Large-scale Terrain Model for Wildfire-prone Landscapes

    cs.CV 2026-07 reject novelty 4.0

    A ray-tracing pipeline aligns ground-level image pixels to outdated DEM rasters for real-time 3D terrain reconstruction in wildfire zones, validated primarily through a custom simulator.