Pith. sign in

REVIEW 2 cited by

Diffusion Models as Masked Autoencoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.03283 v1 pith:I2ML3I7O submitted 2023-04-06 cs.CV

classification cs.CV
keywords diffusionmodelsmaskedautoencoderspre-trainingrepresentationsstrongvisual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

There has been a longstanding belief that generation can facilitate a true understanding of visual data. In line with this, we revisit generatively pre-training visual representations in light of recent interest in denoising diffusion models. While directly pre-training with diffusion models does not produce strong representations, we condition diffusion models on masked input and formulate diffusion models as masked autoencoders (DiffMAE). Our approach is capable of (i) serving as a strong initialization for downstream recognition tasks, (ii) conducting high-quality image inpainting, and (iii) being effortlessly extended to video where it produces state-of-the-art classification accuracy. We further perform a comprehensive study on the pros and cons of design choices and build connections between diffusion models and masked autoencoders.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Improving Joint Embedding Predictive Architecture with Diffusion Noise

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Injecting EDM-style noise into masked-token position embeddings and adding two auxiliary losses improves I-JEPA's linear-probing accuracy by about 1.5 points on ImageNet-1K.

  2. The model is the message: Lightweight convolutional autoencoders applied to noisy imaging data for planetary science and astrobiology

    astro-ph.EP 2025-07 conditional novelty 4.0 of 10

    A simple convolutional autoencoder reconstructs planetary images with up to 99% pixel loss, and the author argues its latent space could be a more efficient data product than raw imagery.

Pith tools