REVIEW 13 cited by
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Extensive pre-training with large data is indispensable for downstream geometry and semantic visual perception tasks. Thanks to large-scale text-to-image (T2I) pretraining, recent works show promising results by simply fine-tuning T2I diffusion models for dense perception tasks. However, several crucial design decisions in this process still lack comprehensive justification, encompassing the necessity of the multi-step stochastic diffusion mechanism, training strategy, inference ensemble strategy, and fine-tuning data quality. In this work, we conduct a thorough investigation into critical factors that affect transfer efficiency and performance when using diffusion priors. Our key findings are: 1) High-quality fine-tuning data is paramount for both semantic and geometry perception tasks. 2) The stochastic nature of diffusion models has a slightly negative impact on deterministic visual perception tasks. 3) Apart from fine-tuning the diffusion model with only latent space supervision, task-specific image-level supervision is beneficial to enhance fine-grained details. These observations culminate in the development of GenPercept, an effective deterministic one-step fine-tuning paradigm tailed for dense visual perception tasks. Different from the previous multi-step methods, our paradigm has a much faster inference speed, and can be seamlessly integrated with customized perception decoders and loss functions for image-level supervision, which is critical to improving the fine-grained details of predictions. Comprehensive experiments on diverse dense visual perceptual tasks, including monocular depth estimation, surface normal estimation, image segmentation, and matting, are performed to demonstrate the remarkable adaptability and effectiveness of our proposed method.
Forward citations
Cited by 13 Pith papers
-
Unified Video Dense Prediction from Disjoint Data
A single video backbone predicts eight dense scene tasks from separate single-task datasets via latent distillation from diffusion-based specialists, with no co-annotated data or pseudo-labels.
-
Video Generation Models are General-Purpose Vision Learners
A video-diffusion backbone fine-tuned as a single-step multi-task perceiver matches or beats specialists on depth, normals, pose and segmentation, with high data efficiency and sim-to-real transfer.
-
LuxDiT: Lighting Estimation with Video Diffusion Transformer
A video diffusion transformer fine-tuned on synthetic and real data predicts HDR environment maps from images/videos, cutting peak light-direction error by roughly 45% on sunny outdoor scenes versus DiffusionLight.
-
SDMatte: Grafting Diffusion Models for Interactive Matting
SDMatte adapts Stable Diffusion to interactive matting via visual-prompt cross-attention, opacity/coordinate embeddings, and masked self-attention, reporting SOTA results on multiple benchmarks.
-
BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?
BenchDepth evaluates eight depth foundation models by their performance on five downstream tasks, finding Depth Anything V2's relative version to be the most practically useful.
-
UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation
Fine-tuning a pretrained video diffusion transformer to predict geometry in one shared global frame produces consistent, camera-free surface normals and coordinates across entire video clips.
-
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
A frozen conditional diffusion model can be inverted via gradient-based discrete optimization, plus a learned layout prior, to perform object detection and faster classification without training a discriminative head.
-
FiffDepth: Feed-forward Transformation of Diffusion-Based Generators for Detailed Depth Estimation
FiffDepth transforms a pre-trained diffusion image generator into a feed-forward monocular depth estimator that combines generative detail with DINOv2-based robustness.
-
Video Depth without Video Models
A single-image latent diffusion model extended with cross-frame attention and global scale-shift alignment produces state-of-the-art video depth without a video diffusion model.
-
FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal
A coarse-to-fine diffusion SIRR method gates ControlNet residuals with a VLM intensity prior times a multi-scale high-frequency prior, then refines geometry and detail in image space.
-
DidSee: Diffusion-Based Depth Completion for Material-Agnostic Robotic Perception and Manipulation
DidSee is a diffusion-based depth completion model that combines a zero terminal-SNR noise scheduler, single-step training, and a semantic segmentation enhancer to achieve state-of-the-art results on non-Lambertian objects.
-
Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis
Marigold adapts Stable Diffusion via a simple latent-space fine-tuning recipe to competitive zero-shot depth, normals, and intrinsic decomposition, using only small synthetic training sets.
-
Multi-agent Embodied AI: Advances and Future Directions
A survey that maps multi-agent embodied AI methods and benchmarks across control, learning, and generative-model categories, and lists open challenges.
Discussion (0). Continue with ORCID to comment.