Pith. sign in

REVIEW 3 cited by

MILAN: Masked Image Pretraining on Language Assisted Representation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.06049 v3 pith:N6QRLZKV submitted 2022-08-11 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords imagemaskedmilanpretraininglanguageprevioussemanticaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-attention based transformer models have been dominating many computer vision tasks in the past few years. Their superb model qualities heavily depend on the excessively large labeled image datasets. In order to reduce the reliance on large labeled datasets, reconstruction based masked autoencoders are gaining popularity, which learn high quality transferable representations from unlabeled images. For the same purpose, recent weakly supervised image pretraining methods explore language supervision from text captions accompanying the images. In this work, we propose masked image pretraining on language assisted representation, dubbed as MILAN. Instead of predicting raw pixels or low level features, our pretraining objective is to reconstruct the image features with substantial semantic signals that are obtained using caption supervision. Moreover, to accommodate our reconstruction target, we propose a more effective prompting decoder architecture and a semantic aware mask sampling mechanism, which further advance the transfer performance of the pretrained model. Experimental results demonstrate that MILAN delivers higher accuracy than the previous works. When the masked autoencoder is pretrained and finetuned on ImageNet-1K dataset with an input resolution of 224x224, MILAN achieves a top-1 accuracy of 85.4% on ViT-Base, surpassing previous state-of-the-arts by 1%. In the downstream semantic segmentation task, MILAN achieves 52.7 mIoU using ViT-Base on ADE20K dataset, outperforming previous masked pretraining results by 4 points.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Guided Masked Autoencoder

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A Masked Autoencoder that masks the object cluster found by its own early patch-clustering signal learns better representations than random masking, with no external labels or models.

  2. M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision

    cs.CV 2025-07 conditional novelty 6.0 of 10

    M-SpecGene is a Siamese masked-autoencoder foundation model for RGB-thermal vision, trained on the RGBT550K dataset with a GMM-CMSS progressive masking strategy, and evaluated on four downstream tasks.

  3. Symmetry Understanding of 3D Shapes via Chirality Disentanglement

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    Abstract-level claim: decorate 3D shape vertices with chirality features drawn from 2D foundation models via Diff3F, enabling left-right disentanglement; the supplied full text is a different paper, so the claim is un...

Pith tools