Pith. sign in

REVIEW 2 cited by

Curriculum Learning with Quality-Driven Data Selection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.00102 v2 pith:CQ5UV3XA submitted 2024-06-27 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords dataqualityselectioncapabilitiesmllmsspacecurriculumdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The impressive multimodal capabilities demonstrated by OpenAI's GPT-4 have generated significant interest in the development of Multimodal Large Language Models (MLLMs). Visual instruction tuning of MLLMs with machine-generated instruction-following data has shown to enhance zero-shot capabilities across various tasks. However, there has been limited exploration into controlling the quality of the instruction data.Current methodologies for data selection in MLLMs often rely on single, unreliable scores or use downstream tasks for selection, which is time-consuming and can lead to potential overfitting on the chosen evaluation datasets. To mitigate these limitations, we propose a novel data selection methodology that utilizes image-text correlation and model perplexity to evaluate and select data of varying quality. This approach leverages the distinct distribution of these two attributes, mapping data quality into a two-dimensional space that allows for the selection of data based on their location within this distribution. By utilizing this space, we can analyze the impact of task type settings, used as prompts, on data quality. Additionally, this space can be used to construct multi-stage subsets of varying quality to facilitate curriculum learning. Our research includes comprehensive experiments conducted on various datasets. The results emphasize substantial enhancements in five commonly assessed capabilities compared to using the complete dataset. Our codes, data, and models are publicly available at: https://anonymous.4open.science/r/EHIT-31B4

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples

    cs.CL 2025-02 conditional novelty 6.0 of 10

    Preference examples vary in difficulty; overly difficult examples degrade DPO alignment, and filtering them out improves AlpacaEval 2 win rates by 9-16 percentage points.

  2. DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Decoupling inter-class ratio search from intra-class convex allocation yields VLM data recipes that beat stacking and transfer from small proxies to larger scales.

Pith tools