Pith. sign in

REVIEW 13 cited by

What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06090 v4 pith:2L4IJG76 submitted 2024-03-10 cs.CV

classification cs.CV
keywords perceptiontasksdiffusionfine-tuningdensevisualdatamodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Extensive pre-training with large data is indispensable for downstream geometry and semantic visual perception tasks. Thanks to large-scale text-to-image (T2I) pretraining, recent works show promising results by simply fine-tuning T2I diffusion models for dense perception tasks. However, several crucial design decisions in this process still lack comprehensive justification, encompassing the necessity of the multi-step stochastic diffusion mechanism, training strategy, inference ensemble strategy, and fine-tuning data quality. In this work, we conduct a thorough investigation into critical factors that affect transfer efficiency and performance when using diffusion priors. Our key findings are: 1) High-quality fine-tuning data is paramount for both semantic and geometry perception tasks. 2) The stochastic nature of diffusion models has a slightly negative impact on deterministic visual perception tasks. 3) Apart from fine-tuning the diffusion model with only latent space supervision, task-specific image-level supervision is beneficial to enhance fine-grained details. These observations culminate in the development of GenPercept, an effective deterministic one-step fine-tuning paradigm tailed for dense visual perception tasks. Different from the previous multi-step methods, our paradigm has a much faster inference speed, and can be seamlessly integrated with customized perception decoders and loss functions for image-level supervision, which is critical to improving the fine-grained details of predictions. Comprehensive experiments on diverse dense visual perceptual tasks, including monocular depth estimation, surface normal estimation, image segmentation, and matting, are performed to demonstrate the remarkable adaptability and effectiveness of our proposed method.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unified Video Dense Prediction from Disjoint Data

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A single video backbone predicts eight dense scene tasks from separate single-task datasets via latent distillation from diffusion-based specialists, with no co-annotated data or pseudo-labels.

  2. Video Generation Models are General-Purpose Vision Learners

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A video-diffusion backbone fine-tuned as a single-step multi-task perceiver matches or beats specialists on depth, normals, pose and segmentation, with high data efficiency and sim-to-real transfer.

  3. LuxDiT: Lighting Estimation with Video Diffusion Transformer

    cs.GR 2025-09 conditional novelty 6.0 of 10

    A video diffusion transformer fine-tuned on synthetic and real data predicts HDR environment maps from images/videos, cutting peak light-direction error by roughly 45% on sunny outdoor scenes versus DiffusionLight.

  4. SDMatte: Grafting Diffusion Models for Interactive Matting

    cs.CV 2025-08 conditional novelty 6.0 of 10

    SDMatte adapts Stable Diffusion to interactive matting via visual-prompt cross-attention, opacity/coordinate embeddings, and masked self-attention, reporting SOTA results on multiple benchmarks.

  5. BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    BenchDepth evaluates eight depth foundation models by their performance on five downstream tasks, finding Depth Anything V2's relative version to be the most practically useful.

  6. UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Fine-tuning a pretrained video diffusion transformer to predict geometry in one shared global frame produces consistent, camera-free surface normals and coordinates across entire video clips.

  7. DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A frozen conditional diffusion model can be inverted via gradient-based discrete optimization, plus a learned layout prior, to perform object detection and faster classification without training a discriminative head.

  8. FiffDepth: Feed-forward Transformation of Diffusion-Based Generators for Detailed Depth Estimation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    FiffDepth transforms a pre-trained diffusion image generator into a feed-forward monocular depth estimator that combines generative detail with DINOv2-based robustness.

  9. Video Depth without Video Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A single-image latent diffusion model extended with cross-frame attention and global scale-shift alignment produces state-of-the-art video depth without a video diffusion model.

  10. FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A coarse-to-fine diffusion SIRR method gates ControlNet residuals with a VLM intensity prior times a multi-scale high-frequency prior, then refines geometry and detail in image space.

  11. DidSee: Diffusion-Based Depth Completion for Material-Agnostic Robotic Perception and Manipulation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    DidSee is a diffusion-based depth completion model that combines a zero terminal-SNR noise scheduler, single-step training, and a semantic segmentation enhancer to achieve state-of-the-art results on non-Lambertian objects.

  12. Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Marigold adapts Stable Diffusion via a simple latent-space fine-tuning recipe to competitive zero-shot depth, normals, and intrinsic decomposition, using only small synthetic training sets.

  13. Multi-agent Embodied AI: Advances and Future Directions

    cs.AI 2025-05 conditional novelty 3.0 of 10

    A survey that maps multi-agent embodied AI methods and benchmarks across control, learning, and generative-model categories, and lists open challenges.

Pith tools