Pith. sign in

REVIEW 1 cited by

Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.01397 v2 pith:PVLHFZZ4 submitted 2022-05-03 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords robustnesstrainingdistributiongainslanguagelanguage-imagecausesclip
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contrastively trained language-image models such as CLIP, ALIGN, and BASIC have demonstrated unprecedented robustness to multiple challenging natural distribution shifts. Since these language-image models differ from previous training approaches in several ways, an important question is what causes the large robustness gains. We answer this question via a systematic experimental investigation. Concretely, we study five different possible causes for the robustness gains: (i) the training set size, (ii) the training distribution, (iii) language supervision at training time, (iv) language supervision at test time, and (v) the contrastive loss function. Our experiments show that the more diverse training distribution is the main cause for the robustness gains, with the other factors contributing little to no robustness. Beyond our experimental results, we also introduce ImageNet-Captions, a version of ImageNet with original text annotations from Flickr, to enable further controlled experiments of language-image training.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A K-5-only pretraining corpus and 5B model show that language model capabilities track the knowledge boundary of the training data, and standard post-training methods do not cross it.

Pith tools