Pith. sign in

REVIEW 8 cited by

EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1901.11196 v2 pith:HFT5ZELH submitted 2019-01-31 cs.CL

classification cs.CL
keywords classificationdataperformancerandomtaskstexttrainingaugmentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present EDA: easy data augmentation techniques for boosting performance on text classification tasks. EDA consists of four simple but powerful operations: synonym replacement, random insertion, random swap, and random deletion. On five text classification tasks, we show that EDA improves performance for both convolutional and recurrent neural networks. EDA demonstrates particularly strong results for smaller datasets; on average, across five datasets, training with EDA while using only 50% of the available training set achieved the same accuracy as normal training with all available data. We also performed extensive ablation studies and suggest parameters for practical use.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 198 citations worldwide. Full citation record

  1. Exact Reformulation and Optimization for Direct Metric Optimization in Binary Imbalanced Classification

    cs.LG 2025-07 conditional novelty 8.0 of 10

    A continuous exact reformulation lets precision, recall, and F-beta metrics be optimized with gradient methods, avoiding smooth surrogate losses.

  2. inversedMixup: Data Augmentation via Inverting Mixed Embeddings

    cs.CL 2026-01 conditional novelty 6.0 of 10

    Mixing BERT embeddings and inverting them into text with LLaMA produces interpretable augmented sentences, improves few-shot classification on some datasets, and exposes 'manifold intrusion' in text Mixup.

  3. Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual Grounding

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Unit-robust text augmentation (3DTE) plus text-guided geometry attention (TGE) lifts monocular 3D visual grounding accuracy on Mono3DRefer, notably in far-range scenes.

  4. Iterative Augmentation with Summarization Refinement (IASR) Evaluation for Unstructured Survey data Modeling and Analysis

    cs.CL 2025-07 reject novelty 5.0 of 10

    The paper evaluates four LLMs as text augmenters and reports GPT-3.5 Turbo as the best, and that combining augmentation with GPT topic labels increases BERTopic's discovered topics from 5 to 20 with zero overlap.

  5. Explainable AI: XAI-Guided Context-Aware Data Augmentation

    cs.CL 2025-06 conditional novelty 4.0 of 10

    XAI-guided augmentation that replaces the least important words, identified by Integrated Gradients, with back-translated synonyms or paraphrases improves hate speech and sentiment classification accuracy by up to 8 p...

  6. MASTER: Enhancing Large Language Model via Multi-Agent Simulated Teaching

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A multi-agent simulated teaching pipeline creates BOOST-QA, and fine-tuning on it lifts reported LLM benchmark scores by up to 31 points over the original data.

  7. 3D Skeleton-Based Action Recognition: A Review

    cs.CV 2025-06 reject novelty 3.0 of 10

    A task-oriented review of skeleton-based action recognition that reorganizes known methods along a data processing pipeline and contains no new experimental result.

  8. Improving QA Efficiency with DistilBERT: Fine-Tuning and Inference on mobile Intel CPUs

    cs.CL 2025-05 conditional novelty 2.0 of 10

    Fine-tuned DistilBERT on 45,000 SQuAD examples achieves validation F1 0.6536 and 0.1208 s average inference per question on an Intel i7-1355U CPU.

Pith tools