REVIEW 10 cited by
Generative Models: What Do They Know? Do They Know Things? Let's Find Out!
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Generative models excel at mimicking real scenes, suggesting they might inherently encode important intrinsic scene properties. In this paper, we aim to explore the following key questions: (1) What intrinsic knowledge do generative models like GANs, Autoregressive models, and Diffusion models encode? (2) Can we establish a general framework to recover intrinsic representations from these models, regardless of their architecture or model type? (3) How minimal can the required learnable parameters and labeled data be to successfully recover this knowledge? (4) Is there a direct link between the quality of a generative model and the accuracy of the recovered scene intrinsics? Our findings indicate that a small Low-Rank Adaptators (LoRA) can recover intrinsic images-depth, normals, albedo and shading-across different generators (Autoregressive, GANs and Diffusion) while using the same decoder head that generates the image. As LoRA is lightweight, we introduce very few learnable parameters (as few as 0.04% of Stable Diffusion model weights for a rank of 2), and we find that as few as 250 labeled images are enough to generate intrinsic images with these LoRA modules. Finally, we also show a positive correlation between the generative model's quality and the accuracy of the recovered intrinsics through control experiments.
Forward citations
Cited by 10 Pith papers
-
Orchid: Image Latent Diffusion for Joint Appearance and Geometry Generation
Orchid jointly generates color, depth, and surface normals in one latent diffusion model, and also supports image-conditioned depth-normal prediction and joint inpainting.
-
Relighting as a Probe of Visual Priors via Augmented Latent Intrinsics
Semantic encoders can harm relighting, and ALI—fusing dense visual features with latent intrinsics—improves relighting on glossy and specular materials.
-
DNF-Intrinsic: Deterministic Noise-Free Diffusion for Indoor Inverse Rendering
A diffusion model fine-tuned with flow matching maps a single RGB indoor image directly to five scene properties, beating prior inverse rendering methods on InteriorVerse and real-world benchmarks.
-
Nested Diffusion Models Using Hierarchical Latent Priors
Nested diffusion models that generate images by progressively synthesizing hierarchical semantic latents from a frozen pretrained encoder improve image quality over single-level baselines at modest extra cost.
-
LumiNet: Latent Intrinsics Meets Diffusion Models for Indoor Scene Relighting
LumiNet transfers lighting between indoor scenes from images alone by conditioning a diffusion model on latent intrinsics from the source and a lighting code from the target.
-
Multiview Equivariance Improves 3D Correspondence Understanding with Minimal Feature Finetuning
Finetuning ViT features with SmoothAP on Objaverse multiview correspondences improves 3D correspondence tasks, with meaningful gains even from a single object and a single iteration.
-
DepthDark: Robust Monocular Depth Estimation for Low-Light Environments
DepthDark obtains state-of-the-art low-light depth estimates by jointly introducing synthetic nighttime data generation and an efficient fine-tuning strategy for a pretrained depth foundation model.
-
Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment
A reparameterization recipe that lets pre-trained Stable Diffusion checkpoints be finetuned as flow matching models, giving faster convergence and better performance under parameter-efficient constraints.
-
SAIL: Self-supervised Albedo Estimation from Real Images with a Latent Diffusion Model
SAIL learns a latent-space decomposition into albedo and lighting components via a relighting surrogate objective, producing consistent albedo estimates from unlabeled real images.
-
Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis
Marigold adapts Stable Diffusion via a simple latent-space fine-tuning recipe to competitive zero-shot depth, normals, and intrinsic decomposition, using only small synthetic training sets.
Discussion (0). Continue with ORCID to comment.