REVIEW 1 cited by
Is a Caption Worth a Thousand Images? A Controlled Study for Representation Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The development of CLIP [Radford et al., 2021] has sparked a debate on whether language supervision can result in vision models with more transferable representations than traditional image-only methods. Our work studies this question through a carefully controlled comparison of two approaches in terms of their ability to learn representations that generalize to downstream classification tasks. We find that when the pre-training dataset meets certain criteria -- it is sufficiently large and contains descriptive captions with low variability -- image-only methods do not match CLIP's transfer performance, even when they are trained with more image data. However, contrary to what one might expect, there are practical settings in which these criteria are not met, wherein added supervision through captions is actually detrimental. Motivated by our findings, we devise simple prescriptions to enable CLIP to better leverage the language information present in existing pre-training datasets.
Forward citations
Cited by 1 Pith paper
-
ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos
ACE fine-tunes video-language models with stochastically sampled action synonyms and shadow negatives, improving zero-shot classification of unseen procedural actions by up to 16 percent harmonic mean on ATA, IKEA, and GTEA.
Discussion (0). Continue with ORCID to comment.