Pith. sign in

REVIEW 5 cited by

Generative Models: What Do They Know? Do They Know Things? Let's Find Out!

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17137 v3 pith:AODJQC2O submitted 2023-11-28 cs.CV cs.AIcs.GRcs.LG

classification cs.CVcs.AIcs.GRcs.LG
keywords modelsgenerativeintrinsicmodeldiffusionlorarecoverthey
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative models excel at mimicking real scenes, suggesting they might inherently encode important intrinsic scene properties. In this paper, we aim to explore the following key questions: (1) What intrinsic knowledge do generative models like GANs, Autoregressive models, and Diffusion models encode? (2) Can we establish a general framework to recover intrinsic representations from these models, regardless of their architecture or model type? (3) How minimal can the required learnable parameters and labeled data be to successfully recover this knowledge? (4) Is there a direct link between the quality of a generative model and the accuracy of the recovered scene intrinsics? Our findings indicate that a small Low-Rank Adaptators (LoRA) can recover intrinsic images-depth, normals, albedo and shading-across different generators (Autoregressive, GANs and Diffusion) while using the same decoder head that generates the image. As LoRA is lightweight, we introduce very few learnable parameters (as few as 0.04% of Stable Diffusion model weights for a rank of 2), and we find that as few as 250 labeled images are enough to generate intrinsic images with these LoRA modules. Finally, we also show a positive correlation between the generative model's quality and the accuracy of the recovered intrinsics through control experiments.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Relighting as a Probe of Visual Priors via Augmented Latent Intrinsics

    cs.CV 2026-02 conditional novelty 6.0 of 10

    Semantic encoders can harm relighting, and ALI—fusing dense visual features with latent intrinsics—improves relighting on glossy and specular materials.

  2. DNF-Intrinsic: Deterministic Noise-Free Diffusion for Indoor Inverse Rendering

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A diffusion model fine-tuned with flow matching maps a single RGB indoor image directly to five scene properties, beating prior inverse rendering methods on InteriorVerse and real-world benchmarks.

  3. DepthDark: Robust Monocular Depth Estimation for Low-Light Environments

    cs.CV 2025-07 conditional novelty 5.0 of 10

    DepthDark obtains state-of-the-art low-light depth estimates by jointly introducing synthetic nighttime data generation and an efficient fine-tuning strategy for a pretrained depth foundation model.

  4. Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A reparameterization recipe that lets pre-trained Stable Diffusion checkpoints be finetuned as flow matching models, giving faster convergence and better performance under parameter-efficient constraints.

  5. SAIL: Self-supervised Albedo Estimation from Real Images with a Latent Diffusion Model

    cs.CV 2025-05 reject novelty 5.0 of 10

    SAIL learns a latent-space decomposition into albedo and lighting components via a relighting surrogate objective, producing consistent albedo estimates from unlabeled real images.

Pith tools