REVIEW 2 cited by
Diffusion Models as Masked Autoencoders
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
There has been a longstanding belief that generation can facilitate a true understanding of visual data. In line with this, we revisit generatively pre-training visual representations in light of recent interest in denoising diffusion models. While directly pre-training with diffusion models does not produce strong representations, we condition diffusion models on masked input and formulate diffusion models as masked autoencoders (DiffMAE). Our approach is capable of (i) serving as a strong initialization for downstream recognition tasks, (ii) conducting high-quality image inpainting, and (iii) being effortlessly extended to video where it produces state-of-the-art classification accuracy. We further perform a comprehensive study on the pros and cons of design choices and build connections between diffusion models and masked autoencoders.
Forward citations
Cited by 2 Pith papers
-
Improving Joint Embedding Predictive Architecture with Diffusion Noise
Injecting EDM-style noise into masked-token position embeddings and adding two auxiliary losses improves I-JEPA's linear-probing accuracy by about 1.5 points on ImageNet-1K.
-
The model is the message: Lightweight convolutional autoencoders applied to noisy imaging data for planetary science and astrobiology
A simple convolutional autoencoder reconstructs planetary images with up to 99% pixel loss, and the author argues its latent space could be a more efficient data product than raw imagery.
Discussion (0). Sign in to comment.