Pith. sign in

REVIEW 1 cited by

The Role of Pre-training Data in Transfer Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.13602 v2 pith:XMQV7L2C submitted 2023-02-27 cs.CV cs.LG

classification cs.CVcs.LG
keywords pre-trainingdatatransfercontrastivefindfine-tuninglearningrole
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The transfer learning paradigm of model pre-training and subsequent fine-tuning produces high-accuracy models. While most studies recommend scaling the pre-training size to benefit most from transfer learning, a question remains: what data and method should be used for pre-training? We investigate the impact of pre-training data distribution on the few-shot and full fine-tuning performance using 3 pre-training methods (supervised, contrastive language-image and image-image), 7 pre-training datasets, and 9 downstream datasets. Through extensive controlled experiments, we find that the choice of the pre-training data source is essential for the few-shot transfer, but its role decreases as more data is made available for fine-tuning. Additionally, we explore the role of data curation and examine the trade-offs between label noise and the size of the pre-training dataset. We find that using 2000X more pre-training data from LAION can match the performance of supervised ImageNet pre-training. Furthermore, we investigate the effect of pre-training methods, comparing language-image contrastive vs. image-image contrastive, and find that the latter leads to better downstream accuracy

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Does the Spatial Distribution of Pre-training Data Affect Geospatial Foundation Models?

    cs.LG 2025-01 conditional novelty 6.0 of 10

    Balanced, globally representative pre-training data generally outperforms region-specific sampling for two geospatial foundation models in few-shot downstream tasks, and the advantage shrinks as finetuning data grows.

Pith tools