Pith. sign in

REVIEW 1 cited by

Is a Caption Worth a Thousand Images? A Controlled Study for Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.07635 v1 pith:V7B474SP submitted 2022-07-15 cs.CV cs.LGstat.ML

classification cs.CVcs.LGstat.ML
keywords clipcaptionscontrolledcriteriaimage-onlylanguagemethodspre-training
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The development of CLIP [Radford et al., 2021] has sparked a debate on whether language supervision can result in vision models with more transferable representations than traditional image-only methods. Our work studies this question through a carefully controlled comparison of two approaches in terms of their ability to learn representations that generalize to downstream classification tasks. We find that when the pre-training dataset meets certain criteria -- it is sufficiently large and contains descriptive captions with low variability -- image-only methods do not match CLIP's transfer performance, even when they are trained with more image data. However, contrary to what one might expect, there are practical settings in which these criteria are not met, wherein added supervision through captions is actually detrimental. Motivated by our findings, we devise simple prescriptions to enable CLIP to better leverage the language information present in existing pre-training datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ACE: Action Concept Enhancement of Video-Language Models in Procedural Videos

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ACE fine-tunes video-language models with stochastically sampled action synonyms and shadow negatives, improving zero-shot classification of unseen procedural actions by up to 16 percent harmonic mean on ATA, IKEA, and GTEA.

Pith tools