Pith. sign in

REVIEW 1 cited by

MaskCLIP: Masked Self-Distillation Advances Contrastive Language-Image Pretraining

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.12262 v2 pith:GMMOVXDB submitted 2022-08-25 cs.CV

classification cs.CV
keywords maskedself-distillationcontrastivemaskcliprepresentationlocalpretrainingbenefits
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents a simple yet effective framework MaskCLIP, which incorporates a newly proposed masked self-distillation into contrastive language-image pretraining. The core idea of masked self-distillation is to distill representation from a full image to the representation predicted from a masked image. Such incorporation enjoys two vital benefits. First, masked self-distillation targets local patch representation learning, which is complementary to vision-language contrastive focusing on text-related representation. Second, masked self-distillation is also consistent with vision-language contrastive from the perspective of training objective as both utilize the visual encoder for feature aligning, and thus is able to learn local semantics getting indirect supervision from the language. We provide specially designed experiments with a comprehensive analysis to validate the two benefits. Symmetrically, we also introduce the local semantic supervision into the text branch, which further improves the pretraining performance. With extensive experiments, we show that MaskCLIP, when applied to various challenging downstream tasks, achieves superior results in linear probing, finetuning, and zero-shot performance with the guidance of the language encoder. Code will be release at \url{https://github.com/LightDXY/MaskCLIP}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. sDREAMER: Self-distilled Mixture-of-Modality-Experts Transformer for Automatic Sleep Staging

    cs.LG 2025-01 conditional novelty 5.0 of 10

    A mixture-of-modality-experts transformer with self-distillation reports improved mouse sleep staging and enables single-channel inference after multi-channel training.

Pith tools