Pith. sign in

REVIEW 5 cited by

A Survey of Data Augmentation Approaches for NLP

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.03075 v5 pith:IM6CHVOU submitted 2021-05-07 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords dataaugmentationapproachesareachallengesgithubliteraturemotivate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data augmentation has recently seen increased interest in NLP due to more work in low-resource domains, new tasks, and the popularity of large-scale neural networks that require large amounts of training data. Despite this recent upsurge, this area is still relatively underexplored, perhaps due to the challenges posed by the discrete nature of language data. In this paper, we present a comprehensive and unifying survey of data augmentation for NLP by summarizing the literature in a structured manner. We first introduce and motivate data augmentation for NLP, and then discuss major methodologically representative approaches. Next, we highlight techniques that are used for popular NLP applications and tasks. We conclude by outlining current challenges and directions for future research. Overall, our paper aims to clarify the landscape of existing literature in data augmentation for NLP and motivate additional work in this area. We also present a GitHub repository with a paper list that will be continuously updated at https://github.com/styfeng/DataAug4NLP

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models

    cs.CL 2026-01 conditional novelty 6.0 of 10

    A new benchmark shows multilingual tool-calling errors in LLMs are mostly parameter-language mismatches at the execution boundary, not failures of intent understanding.

  2. Text Reinforcement for Multimodal Time Series Forecasting

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Reinforcement learning trains an LLM to generate improved text from time series, improving multimodal forecasting on Time-MMD.

  3. Using Sign Language Production as Data Augmentation to enhance Sign Language Translation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Adding synthetic sign-language data produced by stitching, a GAN, or Gaussian splatting to the training set improves sign-language translation, with the largest gains for skeleton-pose models.

  4. AI-in-the-Loop Sensing and Communication Joint Design for Edge Intelligence

    cs.LG 2025-02 conditional novelty 4.0 of 10

    An AI-in-the-loop JSAC framework tunes sensing and transmission to lower both resource costs and validation loss in federated edge learning.

  5. Multi-Agent Language Models: Advancing Cooperation, Coordination, and Adaptation

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A thesis proposal repurposing two prior papers on LM agents for text games, framed as a path to theory-of-mind AI, with no new theory-of-mind evidence.

Pith tools