Pith. sign in

REVIEW 14 cited by

What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06090 v4 pith:2L4IJG76 submitted 2024-03-10 cs.CV

classification cs.CV
keywords perceptiontasksdiffusionfine-tuningdensevisualdatamodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Extensive pre-training with large data is indispensable for downstream geometry and semantic visual perception tasks. Thanks to large-scale text-to-image (T2I) pretraining, recent works show promising results by simply fine-tuning T2I diffusion models for dense perception tasks. However, several crucial design decisions in this process still lack comprehensive justification, encompassing the necessity of the multi-step stochastic diffusion mechanism, training strategy, inference ensemble strategy, and fine-tuning data quality. In this work, we conduct a thorough investigation into critical factors that affect transfer efficiency and performance when using diffusion priors. Our key findings are: 1) High-quality fine-tuning data is paramount for both semantic and geometry perception tasks. 2) The stochastic nature of diffusion models has a slightly negative impact on deterministic visual perception tasks. 3) Apart from fine-tuning the diffusion model with only latent space supervision, task-specific image-level supervision is beneficial to enhance fine-grained details. These observations culminate in the development of GenPercept, an effective deterministic one-step fine-tuning paradigm tailed for dense visual perception tasks. Different from the previous multi-step methods, our paradigm has a much faster inference speed, and can be seamlessly integrated with customized perception decoders and loss functions for image-level supervision, which is critical to improving the fine-grained details of predictions. Comprehensive experiments on diverse dense visual perceptual tasks, including monocular depth estimation, surface normal estimation, image segmentation, and matting, are performed to demonstrate the remarkable adaptability and effectiveness of our proposed method.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Unified Video Dense Prediction from Disjoint Data

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A single video backbone predicts eight dense scene tasks from separate single-task datasets via latent distillation from diffusion-based specialists, with no co-annotated data or pseudo-labels.

  2. MUSE: Unlocking Timestep as Native Task Steering for One-Step Dense Prediction

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    MUSE shows that the native timestep embedding in diffusion models acts as a parameter-free steering signal for multi-task monocular depth and normal estimation via manifold decoupling in latent space.

  3. DepthMaster: Unified Monocular Depth Estimation for Perspective and Panoramic Images

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    DepthMaster unifies metric monocular depth estimation for perspective and panoramic images by patching panoramas into perspective views, adding a consistency loss and virtual cameras, and training mostly on perspectiv...

  4. Monocular Depth Estimation via Neural Network with Learnable Algebraic Group and Ring Structures

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    LAGRNet embeds learnable algebraic group, ring, and sheaf structures into a neural network to improve accuracy and generalization in monocular depth estimation.

  5. CDPR: Cross-modal Diffusion with Polarization for Reliable Monocular Depth Estimation

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    CDPR integrates polarization priors into a diffusion-based monocular depth estimator via shared latent space and adaptive gating, outperforming RGB-only methods in challenging scenes.

  6. Video Generation Models are General-Purpose Vision Learners

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A video-diffusion backbone fine-tuned as a single-step multi-task perceiver matches or beats specialists on depth, normals, pose and segmentation, with high data efficiency and sim-to-real transfer.

  7. UniGP: Taming Diffusion Transformer for Prior-Preserved Unified Generation and Perception

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    UniGP unifies controllable generation and dense prediction in an MMDiT-based diffusion model through simple joint training that preserves backbone priors.

  8. UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    UniVidX unifies diverse video generation tasks into one conditional diffusion model using stochastic condition masking, decoupled gated LoRAs, and cross-modal self-attention.

  9. LuxDiT: Lighting Estimation with Video Diffusion Transformer

    cs.GR 2025-09 conditional novelty 6.0 of 10

    A video diffusion transformer fine-tuned on synthetic and real data predicts HDR environment maps from images/videos, cutting peak light-direction error by roughly 45% on sunny outdoor scenes versus DiffusionLight.

  10. Ouroboros: Single-step Diffusion Models for Cycle-consistent Forward and Inverse Rendering

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Ouroboros uses two single-step diffusion models with cycle consistency for forward and inverse rendering, extending intrinsic decomposition to indoor/outdoor scenes with faster inference than multi-step methods.

  11. SDMatte: Grafting Diffusion Models for Interactive Matting

    cs.CV 2025-08 conditional novelty 6.0 of 10

    SDMatte adapts Stable Diffusion to interactive matting via visual-prompt cross-attention, opacity/coordinate embeddings, and masked self-attention, reporting SOTA results on multiple benchmarks.

  12. Depth Anything V2

    cs.CV 2024-06 unverdicted novelty 6.0 of 10

    Depth Anything V2 delivers finer, more robust monocular depth predictions by replacing real labeled images with synthetic data, scaling the teacher model, and using large-scale pseudo-labeled real images for student training.

  13. FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A coarse-to-fine diffusion SIRR method gates ControlNet residuals with a VLM intensity prior times a multi-scale high-frequency prior, then refines geometry and detail in image space.

  14. DepthMaster: Taming Diffusion Models for Monocular Depth Estimation

    cs.CV 2025-01 unverdicted novelty 5.0 of 10

    DepthMaster proposes a single-step diffusion model with Feature Alignment and Fourier Enhancement modules in a two-stage training process to improve generalization and detail preservation in monocular depth estimation...

Pith tools