Pith. sign in

REVIEW 8 cited by

A Survey on Data Augmentation in Large Model Era

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.15422 v2 pith:5EQXW344 submitted 2024-01-27 cs.LG cs.CLcs.CV

classification cs.LGcs.CLcs.CV
keywords dataaugmentationlargemodelsmethodshigh-qualitylanguagemodel-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large models, encompassing large language and diffusion models, have shown exceptional promise in approximating human-level intelligence, garnering significant interest from both academic and industrial spheres. However, the training of these large models necessitates vast quantities of high-quality data, and with continuous updates to these models, the existing reservoir of high-quality data may soon be depleted. This challenge has catalyzed a surge in research focused on data augmentation methods. Leveraging large models, these data augmentation techniques have outperformed traditional approaches. This paper offers an exhaustive review of large model-driven data augmentation methods, adopting a comprehensive perspective. We begin by establishing a classification of relevant studies into three main categories: image augmentation, text augmentation, and paired data augmentation. Following this, we delve into various data post-processing techniques pertinent to large model-based data augmentation. Our discussion then expands to encompass the array of applications for these data augmentation methods within natural language processing, computer vision, and audio signal processing. We proceed to evaluate the successes and limitations of large model-based data augmentation across different scenarios. Concluding our review, we highlight prospective challenges and avenues for future exploration in the field of data augmentation. Our objective is to furnish researchers with critical insights, ultimately contributing to the advancement of more sophisticated large models. We consistently maintain the related open-source materials at: https://github.com/MLGroup-JLU/LLM-data-aug-survey.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI Governance

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A new taxonomy characterizes 48 post-training AI adaptation techniques on six axes and maps them to regulatory documentation requirements.

  2. Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models

    cs.CL 2026-01 conditional novelty 6.0 of 10

    A new benchmark shows multilingual tool-calling errors in LLMs are mostly parameter-language mismatches at the execution boundary, not failures of intent understanding.

  3. Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AutoCaption uses MCTS to generate fine-grained video key points, forming the MCTS-VCB benchmark that ranks MLLMs and yields training data improving a fine-tuned model's captioning.

  4. Debunk and Infer: Multimodal Fake News Detection via Diffusion-Generated Evidence and LLM Reasoning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A framework called DIFND generates debunking evidence via conditional diffusion and uses multi-agent MLLM reasoning to detect fake news videos, outperforming baselines on FakeSV and FVC.

  5. NoiseCutMix: A Novel Data Augmentation Approach by Mixing Estimated Noise in Diffusion Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Mixing the estimated noise of two class prompts at each denoising step of Stable Diffusion generates natural augmented images that improve fine-grained classification over CutMix on some datasets.

  6. Multi-turn Natural Language to Graph Query Language Translation

    cs.AI 2025-08 conditional novelty 5.0 of 10

    MTGQL is an LLM-generated 4,500-dialogue Chinese financial benchmark for multi-turn natural language to graph query translation, with a dependency-aware baseline reaching 40.60% overall exact match.

  7. Separation Logic of Generic Resources via Sheafeology

    cs.LO 2025-08 unverdicted novelty 5.0 of 10

    Sheafeology uses sheaf categories to make first-order logic resource-aware, yielding separation logics for generic resources such as memory and random variables.

  8. OAT-Rephrase: Optimization-Aware Training Data Rephrasing for Zeroth-Order LLM Fine-Tuning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Rephrasing training data with an LLM that has read the MeZO paper gives small and inconsistent accuracy gains for zeroth-order LLM fine-tuning, with no error bars reported.

Pith tools