REVIEW 2 cited by
The Devil is in the Details: A Deep Dive into the Rabbit Hole of Data Filtering
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The quality of pre-training data plays a critical role in the performance of foundation models. Popular foundation models often design their own recipe for data filtering, which makes it hard to analyze and compare different data filtering approaches. DataComp is a new benchmark dedicated to evaluating different methods for data filtering. This paper describes our learning and solution when participating in the DataComp challenge. Our filtering strategy includes three stages: single-modality filtering, cross-modality filtering, and data distribution alignment. We integrate existing methods and propose new solutions, such as computing CLIP score on horizontally flipped images to mitigate the interference of scene text, using vision and language models to retrieve training samples for target downstream tasks, rebalancing the data distribution to improve the efficiency of allocating the computational budget, etc. We slice and dice our design choices, provide in-depth analysis, and discuss open questions. Our approach outperforms the best method from the DataComp paper by over 4% on the average performance of 38 tasks and by over 2% on ImageNet.
Forward citations
Cited by 2 Pith papers
-
Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning
The authors argue that many-to-many cross-modal correspondences, termed 'multiplicity', are inevitable and require rethinking multimodal learning, training, evaluation, and dataset construction.
-
Quality over Quantity: Boosting Data Efficiency Through Ensembled Multimodal Data Curation
EcoDatum filters web image-text data by ensembling eight unimodal and multimodal quality scorers with weak-supervision weighting, reporting a DataComp small-scale average score of 0.182.
Discussion (0). Continue with ORCID to comment.