Pith. sign in

REVIEW 10 cited by

A Survey on Data Synthesis and Augmentation for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.12896 v1 pith:K6PAPKSP submitted 2024-10-16 cs.CL

classification cs.CL
keywords datagenerationllmsaugmentationfuturehigh-qualitylanguagelarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The success of Large Language Models (LLMs) is inherently linked to the availability of vast, diverse, and high-quality data for training and evaluation. However, the growth rate of high-quality data is significantly outpaced by the expansion of training datasets, leading to a looming data exhaustion crisis. This underscores the urgent need to enhance data efficiency and explore new data sources. In this context, synthetic data has emerged as a promising solution. Currently, data generation primarily consists of two major approaches: data augmentation and synthesis. This paper comprehensively reviews and summarizes data generation techniques throughout the lifecycle of LLMs, including data preparation, pre-training, fine-tuning, instruction-tuning, preference alignment, and applications. Furthermore, We discuss the current constraints faced by these methods and investigate potential pathways for future development and research. Our aspiration is to equip researchers with a clear understanding of these methodologies, enabling them to swiftly identify appropriate data generation strategies in the construction of LLMs, while providing valuable insights for future exploration.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DocMEdit: Towards Document-Level Model Editing

    cs.CL 2025-05 conditional novelty 7.0 of 10

    DocMEdit, a dataset of nearly 38,000 Wikipedia article updates, shows that existing model editing methods achieve low accuracy and cause large side effects on document-level editing tasks.

  2. A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A new taxonomy characterizes 48 post-training AI adaptation techniques on six axes and maps them to regulatory documentation requirements.

  3. LLM-Based Config Synthesis requires Disambiguation

    cs.NI 2025-07 conditional novelty 6.0 of 10

    LLM-based incremental config synthesis needs user disambiguation of insertion placement; Clarify uses differential questions and binary search to resolve it.

  4. Calibrating Pre-trained Language Classifiers on LLM-generated Noisy Labels via Iterative Refinement

    cs.CL 2025-05 conditional novelty 6.0 of 10

    SiDyP improves classifiers trained on LLM-generated noisy labels by retrieving likely true labels from embedding-space neighbors and iteratively refining them with a simplex diffusion model, reporting average gains of...

  5. HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices

    cs.CL 2025-05 conditional novelty 6.0 of 10

    HomeBench is a new smart home benchmark that exposes near-zero success rates for top LLMs on invalid multi-device instructions.

  6. TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments

    cs.HC 2025-05 conditional novelty 6.0 of 10

    TransBench is a new benchmark of 1,459 screenshots and 22,000 grounding instructions for measuring how well GUI agents transfer across app versions, platforms, and applications.

  7. A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data

    cs.AI 2026-01 conditional novelty 5.0 of 10

    A metric-oriented survey that classifies intrinsic quality and trustworthiness metrics for LLM-generated data across six modalities and documents systematic evaluation gaps in the current literature.

  8. MCTS-Refined CoT: High-Quality Fine-Tuning Data for LLM-Based Repository Issue Resolution

    cs.SE 2025-06 conditional novelty 5.0 of 10

    MCTS-REFINE uses tree search plus strict ground-truth matching to build chain-of-thought training data that lifts open-source LLM issue-resolution scores on SWE-bench.

  9. Explainable AI: XAI-Guided Context-Aware Data Augmentation

    cs.CL 2025-06 conditional novelty 4.0 of 10

    XAI-guided augmentation that replaces the least important words, identified by Integrated Gradients, with back-translated synonyms or paraphrases improves hate speech and sentiment classification accuracy by up to 8 p...

  10. Infinite-Instruct: Synthesizing Scaling Code instruction Data with Bidirectional Synthesis and Static Verification

    cs.CL 2025-05 reject novelty 4.0 of 10

    A synthetic data pipeline for code instruction tuning reports large benchmark gains, but the headline improvements are internally inconsistent.

Pith tools