Pith. sign in

REVIEW 1 cited by

Misalign, Contrast then Distill: Rethinking Misalignments in Language-Image Pretraining

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.12661 v1 pith:6PNAOWC7 submitted 2023-12-19 cs.CV

classification cs.CV
keywords misalignmentstrainingcontrasttextadditionalapproachaugmentationdistill
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contrastive Language-Image Pretraining has emerged as a prominent approach for training vision and text encoders with uncurated image-text pairs from the web. To enhance data-efficiency, recent efforts have introduced additional supervision terms that involve random-augmented views of the image. However, since the image augmentation process is unaware of its text counterpart, this procedure could cause various degrees of image-text misalignments during training. Prior methods either disregarded this discrepancy or introduced external models to mitigate the impact of misalignments during training. In contrast, we propose a novel metric learning approach that capitalizes on these misalignments as an additional training source, which we term "Misalign, Contrast then Distill (MCD)". Unlike previous methods that treat augmented images and their text counterparts as simple positive pairs, MCD predicts the continuous scales of misalignment caused by the augmentation. Our extensive experimental results show that our proposed MCD achieves state-of-the-art transferability in multiple classification and retrieval downstream datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion Models

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Masking the image-feature dimensions most correlated with the style reference's content text reduces content leakage and improves text fidelity in text-to-image style transfer diffusion models.

Pith tools