A masked image modeling method that predicts latent cluster assignments from an EMA teacher reaches 83.8% ImageNet accuracy and 32.1 ADE20K mIoU, outperforming prior MIM methods and approaching DINOv2.
data2vec: A general framework for self-supervised learning in speech, vision and language
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Cluster and Predict Latent Patches for Improved Masked Image Modeling
A masked image modeling method that predicts latent cluster assignments from an EMA teacher reaches 83.8% ImageNet accuracy and 32.1 ADE20K mIoU, outperforming prior MIM methods and approaching DINOv2.