Pith. sign in

REVIEW 2 cited by

The Devil is in the Details: A Deep Dive into the Rabbit Hole of Data Filtering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.15954 v1 pith:2SOEKMKL submitted 2023-09-27 cs.CV cs.LG

classification cs.CVcs.LG
keywords datafilteringdatacompmodelsdesigndifferentdistributionfoundation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The quality of pre-training data plays a critical role in the performance of foundation models. Popular foundation models often design their own recipe for data filtering, which makes it hard to analyze and compare different data filtering approaches. DataComp is a new benchmark dedicated to evaluating different methods for data filtering. This paper describes our learning and solution when participating in the DataComp challenge. Our filtering strategy includes three stages: single-modality filtering, cross-modality filtering, and data distribution alignment. We integrate existing methods and propose new solutions, such as computing CLIP score on horizontally flipped images to mitigate the interference of scene text, using vision and language models to retrieve training samples for target downstream tasks, rebalancing the data distribution to improve the efficiency of allocating the computational budget, etc. We slice and dice our design choices, provide in-depth analysis, and discuss open questions. Our approach outperforms the best method from the DataComp paper by over 4% on the average performance of 38 tasks and by over 2% on ImageNet.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    The authors argue that many-to-many cross-modal correspondences, termed 'multiplicity', are inevitable and require rethinking multimodal learning, training, evaluation, and dataset construction.

  2. Quality over Quantity: Boosting Data Efficiency Through Ensembled Multimodal Data Curation

    cs.LG 2025-02 conditional novelty 5.0 of 10

    EcoDatum filters web image-text data by ensembling eight unimodal and multimodal quality scorers with weak-supervision weighting, reporting a DataComp small-scale average score of 0.182.

Pith tools