REVIEW 3 cited by
SILC: Improving Vision Language Pretraining with Self-Distillation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Image-Text pretraining on web-scale image caption datasets has become the default recipe for open vocabulary classification and retrieval models thanks to the success of CLIP and its variants. Several works have also used CLIP features for dense prediction tasks and have shown the emergence of open-set abilities. However, the contrastive objective used by these models only focuses on image-text alignment and does not incentivise image feature learning for dense prediction tasks. In this work, we introduce SILC, a novel framework for vision language pretraining. SILC improves image-text contrastive learning with the simple addition of local-to-global correspondence learning by self-distillation. We show that distilling local image features from an exponential moving average (EMA) teacher model significantly improves model performance on dense predictions tasks like detection and segmentation, while also providing improvements on image-level tasks such as classification and retrieval. SILC models sets a new state of the art for zero-shot classification, few shot classification, image and text retrieval, zero-shot segmentation, and open vocabulary segmentation. We further show that SILC features greatly benefit open vocabulary detection, captioning and visual question answering.
Forward citations
Cited by 3 Pith papers
-
Visual Pre-Training on Unlabeled Images using Reinforcement Learning
Casting image-crop consistency as temporal-difference value learning improves visual representations on unlabeled web, scene, and video data.
-
CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections
NoLA combines LLM class descriptions, DINO feature alignment, and visual prompt tuning to improve CLIP zero-shot classification without labels, averaging 3.6% over LaFTer on 11 datasets.
-
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
UniMed-CLIP, trained on 5.3M open-source medical image-text pairs with LLM-generated captions, reports strong zero-shot gains but is undermined by evaluation datasets that overlap with its pretraining data.
Discussion (0). Continue with ORCID to comment.