Pith. sign in

REVIEW

Attention-Guided Masked Autoencoders For Learning Image Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15172 v1 pith:DQ4S5R62 submitted 2024-02-23 cs.CV cs.LG

classification cs.CVcs.LG
keywords representationsattention-guidedautoencodersemphasisestablishedfunctionimagelearn
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Masked autoencoders (MAEs) have established themselves as a powerful method for unsupervised pre-training for computer vision tasks. While vanilla MAEs put equal emphasis on reconstructing the individual parts of the image, we propose to inform the reconstruction process through an attention-guided loss function. By leveraging advances in unsupervised object discovery, we obtain an attention map of the scene which we employ in the loss function to put increased emphasis on reconstructing relevant objects, thus effectively incentivizing the model to learn more object-focused representations without compromising the established masking strategy. Our evaluations show that our pre-trained models learn better latent representations than the vanilla MAE, demonstrated by improved linear probing and k-NN classification results on several benchmarks while at the same time making ViTs more robust against varying backgrounds.

Discussion (0). Continue with ORCID to comment.

Pith tools