Pith. sign in

REVIEW 1 cited by

MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.13835 v1 pith:IND5BG7B submitted 2025-04-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords dataspacedatasetssemanticdiversityinformationmethodtextbf
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data quality and diversity are key to the construction of effective instruction-tuning datasets. % With the increasing availability of open-source instruction-tuning datasets, it is advantageous to automatically select high-quality and diverse subsets from a vast amount of data. % Existing methods typically prioritize instance quality and use heuristic rules to maintain diversity. % However, this absence of a comprehensive view of the entire collection often leads to suboptimal results. % Moreover, heuristic rules generally focus on distance or clustering within the embedding space, which fails to accurately capture the intent of complex instructions in the semantic space. % To bridge this gap, we propose a unified method for quantifying the information content of datasets. This method models the semantic space by constructing a label graph and quantifies diversity based on the distribution of information within the graph. % Based on such a measurement, we further introduce an efficient sampling method that selects data samples iteratively to \textbf{M}aximize the \textbf{I}nformation \textbf{G}ain (MIG) in semantic space. % Experiments on various datasets and base models demonstrate that MIG consistently outperforms state-of-the-art methods. % Notably, the model fine-tuned with 5\% Tulu3 data sampled by MIG achieves comparable performance to the official SFT model trained on the full dataset, with improvements of +5.73\% on AlpacaEval and +6.89\% on Wildbench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GemMaroc: Unlocking Darija Proficiency in LLMs with Minimal Data

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Fine-tuning Gemma 3-4B and 27B on about 50,000 mixed Darija and English instructions produces a 27B model that matches Atlas-Chat on DarijaMMLU and exceeds it on DarijaHellaSwag, using 48 GPU-hours.

Pith tools