REVIEW 7 cited by
Data Augmentation using Pre-trained Transformer Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Language model based pre-trained models such as BERT have provided significant gains across different NLP tasks. In this paper, we study different types of transformer based pre-trained models such as auto-regressive models (GPT-2), auto-encoder models (BERT), and seq2seq models (BART) for conditional data augmentation. We show that prepending the class labels to text sequences provides a simple yet effective way to condition the pre-trained models for data augmentation. Additionally, on three classification benchmarks, pre-trained Seq2Seq model outperforms other data augmentation methods in a low-resource setting. Further, we explore how different pre-trained model based data augmentation differs in-terms of data diversity, and how well such methods preserve the class-label information.
Forward citations
Cited by 7 Pith papers
-
Backtranslation and paraphrasing in the LLM era? Comparing data augmentation methods for emotion classification
Backtranslation and paraphrasing produce competitive or better classification gains than zero-shot and few-shot generation when augmenting a low-resource emotion dataset.
-
CDS: Knowledge Component-Driven Data Synthesis Guided by Cognitive Diagnosis Theory
A knowledge-component diagnostic pipeline, inspired by cognitive diagnosis theory, generates weakness-targeted synthetic data that improves small LLMs on math, code, and exam benchmarks by up to 13.1 percentage points.
-
Machine Learning Information Retrieval and Summarisation to Support Systematic Review on Outcomes Based Contracting
A proof-of-concept showing that synthetic data augmentation improves passage retrieval for full-text systematic review of social science literature, based on six test papers.
-
The Synthetic Mirror -- Synthetic Data at the Age of Agentic AI
A position paper claiming existing data and AI regulations are unprepared for synthetic data generated by and for agentic AI, recommending targeted legal amendments and new standards.
-
Explainable AI: XAI-Guided Context-Aware Data Augmentation
XAI-guided augmentation that replaces the least important words, identified by Integrated Gradients, with back-translated synonyms or paraphrases improves hate speech and sentiment classification accuracy by up to 8 p...
-
Su-RoBERTa: A Semi-supervised Approach to Predicting Suicide Risk through Social Media using Base Language Models
Su-RoBERTa, a semi-supervised RoBERTa classifier with GPT-2 data augmentation, achieves 69.84% weighted F1 on a suicide risk prediction benchmark.
-
Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities
A literature review that classifies LLM text data augmentation into simple, prompt-based, retrieval-based, and hybrid techniques, with post-processing and evaluation notes.
Discussion (0). Continue with ORCID to comment.