Pith. sign in

REVIEW 3 cited by

Data Management For Training Large Language Models: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.01700 v3 pith:OMZTFOFH submitted 2023-12-04 cs.CL cs.AI

classification cs.CLcs.AI
keywords datamanagementtrainingllmssurveycurrentefficientfine-tuning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data plays a fundamental role in training Large Language Models (LLMs). Efficient data management, particularly in formulating a well-suited training dataset, is significant for enhancing model performance and improving training efficiency during pretraining and supervised fine-tuning stages. Despite the considerable importance of data management, the underlying mechanism of current prominent practices are still unknown. Consequently, the exploration of data management has attracted more and more attention among the research community. This survey aims to provide a comprehensive overview of current research in data management within both the pretraining and supervised fine-tuning stages of LLMs, covering various aspects of data management strategy design. Looking into the future, we extrapolate existing challenges and outline promising directions for development in this field. Therefore, this survey serves as a guiding resource for practitioners aspiring to construct powerful LLMs through efficient data management practices. The collection of the latest papers is available at https://github.com/ZigeW/data_management_LLM.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Unified Framework for In-Context Learning with Causal and Masked Language Models

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Masked and causal pretraining yield same-order k-shot excess-risk bounds under Wasserstein regularity, and a Masked Pair Encoder matches GPT-2-style ICL on synthetic function classes.

  2. On the Effect of Instruction Tuning Loss on Generalization

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Weighted Instruction Tuning, with low-to-moderate prompt weight and moderate-to-high response weight, beats the standard response-only instruction tuning loss in most of the 75 (model, dataset, benchmark) settings tested.

  3. ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    ClusterUCB uses gradient clustering plus a modified UCB bandit to match full-budget gradient influence data selection at a 20% computing budget.

Pith tools