Pith. sign in

REVIEW 3 cited by

Generative Models as a Data Source for Multiview Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.05258 v3 pith:FKHRTJ3J submitted 2021-06-09 cs.CV

classification cs.CV
keywords datagenerativelearningmodelsgeneratorlatentrepresentationrepresentations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative models are now capable of producing highly realistic images that look nearly indistinguishable from the data on which they are trained. This raises the question: if we have good enough generative models, do we still need datasets? We investigate this question in the setting of learning general-purpose visual representations from a black-box generative model rather than directly from data. Given an off-the-shelf image generator without any access to its training data, we train representations from the samples output by this generator. We compare several representation learning methods that can be applied to this setting, using the latent space of the generator to generate multiple "views" of the same semantic content. We show that for contrastive methods, this multiview data can naturally be used to identify positive pairs (nearby in latent space) and negative pairs (far apart in latent space). We find that the resulting representations rival or even outperform those learned directly from real data, but that good performance requires care in the sampling strategy applied and the training method. Generative models can be viewed as a compressed and organized copy of a dataset, and we envision a future where more and more "model zoos" proliferate while datasets become increasingly unwieldy, missing, or private. This paper suggests several techniques for dealing with visual representation learning in such a future. Code is available on our project page https://ali-design.github.io/GenRep/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dataset Augmentation by Mixing Visual Concepts

    cs.CV 2024-12 conditional novelty 5.0 of 10

    MVC fine-tunes Stable Diffusion with mixed CLIP caption embeddings to produce in-domain synthetic images, improving classifier accuracy on several benchmarks.

  2. SGIA: Enhancing Fine-Grained Visual Classification with Sequence Generative Image Augmentation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A sequence generated from a single image with a latent diffusion model improves fine-grained classification accuracy slightly over strong baselines.

  3. $\textit{Revelio}$: Interpreting and leveraging semantic information in diffusion models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Diffusion model features, when decoded with k-sparse autoencoders, reveal interpretable visual concepts, and a lightweight classifier on the best layer (up_ft1 at t=25) beats prior diffusion-based classifiers on fine-...

Pith tools