Pith. sign in

REVIEW 4 cited by

Finetuned Multimodal Language Models Are High-Quality Image-Text Data Filters

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.02677 v1 pith:ULKWBEBK submitted 2024-03-05 cs.CV cs.CL

classification cs.CVcs.CL
keywords datamodelsclipscorefiltersimage-textmlmsdesignfilter
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose a novel framework for filtering image-text data by leveraging fine-tuned Multimodal Language Models (MLMs). Our approach outperforms predominant filtering methods (e.g., CLIPScore) via integrating the recent advances in MLMs. We design four distinct yet complementary metrics to holistically measure the quality of image-text data. A new pipeline is established to construct high-quality instruction data for fine-tuning MLMs as data filters. Comparing with CLIPScore, our MLM filters produce more precise and comprehensive scores that directly improve the quality of filtered data and boost the performance of pre-trained models. We achieve significant improvements over CLIPScore on popular foundation models (i.e., CLIP and BLIP2) and various downstream tasks. Our MLM filter can generalize to different models and tasks, and be used as a drop-in replacement for CLIPScore. An additional ablation study is provided to verify our design choices for the MLM filter.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Synthetic Visual Genome

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A GPT-4V/GPT-4o pipeline for completing and refining scene graph annotations yields a dense synthetic dataset that, after instruction tuning, gives a 3B model strong relationship understanding and grounding results.

  2. Active Data Curation Effectively Distills Large-Scale Multimodal Models

    cs.CV 2024-11 conditional novelty 7.0 of 10

    Selecting training data by a reference model's loss acts as an implicit distillation, and combining it with explicit distillation yields more FLOP-efficient vision-language models that beat prior SoTA on 27 benchmarks.

  3. SCAN: Bootstrapping Contrastive Pre-training for Data Efficiency

    cs.CV 2024-11 conditional novelty 6.0 of 10

    SCAN dynamically prunes and regrows training data during contrastive pre-training, matching full-data accuracy within about 1% on average while using 30-35% less data.

  4. Object detection in adverse weather conditions for autonomous vehicles using Instruct Pix2Pix

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Training on images with artificially added weather using InstructPix2Pix improves Faster R-CNN under adverse conditions but not YOLOv10 in real-world tests.

Pith tools