Pith. sign in

REVIEW 2 cited by

A Survey on Masked Autoencoder for Self-supervised Learning in Vision and Beyond

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.00173 v1 pith:K44GRPE2 submitted 2022-07-30 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords maskedvisionautoencoderautoencoderslearningbertbeyondgenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Masked autoencoders are scalable vision learners, as the title of MAE \cite{he2022masked}, which suggests that self-supervised learning (SSL) in vision might undertake a similar trajectory as in NLP. Specifically, generative pretext tasks with the masked prediction (e.g., BERT) have become a de facto standard SSL practice in NLP. By contrast, early attempts at generative methods in vision have been buried by their discriminative counterparts (like contrastive learning); however, the success of mask image modeling has revived the masking autoencoder (often termed denoising autoencoder in the past). As a milestone to bridge the gap with BERT in NLP, masked autoencoder has attracted unprecedented attention for SSL in vision and beyond. This work conducts a comprehensive survey of masked autoencoders to shed insight on a promising direction of SSL. As the first to review SSL with masked autoencoders, this work focuses on its application in vision by discussing its historical developments, recent progress, and implications for diverse applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GViT: Representing Images as Gaussians for Visual Recognition

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Images encoded as a few hundred learnable 2D Gaussians, steered by classifier gradients, support a ViT that reaches 76.9% top-1 on ImageNet-1k, close to patch-based ViTs.

  2. MM-Prompt: Cross-Modal Prompt Tuning for Continual Visual Question Answering

    cs.CV 2025-05 conditional novelty 5.0 of 10

    MM-Prompt couples the visual and language prompt paths in continual VQA, and reports higher average accuracy and lower forgetting than existing prompt-based methods.

Pith tools