REVIEW 14 cited by
Synthetic Data Generation with Large Language Models for Text Classification: Potential and Limitations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The collection and curation of high-quality training data is crucial for developing text classification models with superior performance, but it is often associated with significant costs and time investment. Researchers have recently explored using large language models (LLMs) to generate synthetic datasets as an alternative approach. However, the effectiveness of the LLM-generated synthetic data in supporting model training is inconsistent across different classification tasks. To better understand factors that moderate the effectiveness of the LLM-generated synthetic data, in this study, we look into how the performance of models trained on these synthetic data may vary with the subjectivity of classification. Our results indicate that subjectivity, at both the task level and instance level, is negatively associated with the performance of the model trained on synthetic data. We conclude by discussing the implications of our work on the potential and limitations of leveraging LLM for synthetic data generation.
Forward citations
Cited by 14 Pith papers
-
Optimizing the Preconditioner: A Black-box Online-to-Nonconvex Conversion with Static Regret Minimization Oracles
An OCO algorithm with only O(√T) static regret, pluggable as a preconditioner selector, recovers the classical O(1/√T) stationarity rate on smooth stochastic nonconvex problems and the O(T^{-2/7}) rate on nonsmooth ones.
-
Risk In Context: Benchmarking Privacy Leakage of Foundation Models in Synthetic Tabular Data Generation
LLM-based tabular generators reproduce seed rows often enough that membership-inference attacks succeed more against them than against GAN, VAE, or diffusion baselines.
-
What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning
Moderately diverse LLM-generated data can improve fine-tuned model performance in low-data settings when distribution shift is minimal, while high diversity or large distribution shift hurts.
-
Calibrating Pre-trained Language Classifiers on LLM-generated Noisy Labels via Iterative Refinement
SiDyP improves classifiers trained on LLM-generated noisy labels by retrieving likely true labels from embedding-space neighbors and iteratively refining them with a simplex diffusion model, reporting average gains of...
-
Measuring Diversity in Synthetic Datasets
DCScore measures dataset diversity as the sum of self-classification probabilities under a softmax similarity matrix, and the paper shows it tracks generation temperature, human judgment, and LLM rankings.
-
Less is More: Adaptive Coverage for Synthetic Training Data
A max-coverage graph algorithm with an adaptive similarity threshold selects 10-30% of synthetic data that trains classifiers as well as or better than the full dataset on sentiment, relation extraction, and NER tasks.
-
Few-shot LLM Synthetic Data with Distribution Matching
SynAlign generates LLM synthetic text from diversity-guided demonstrations, then reweights it by MMD-based distribution matching, improving downstream classification accuracy.
-
Data Quality Enhancement on the Basis of Diversity with Large Language Models for Text Classification: Uncovered, Difficult, and Noisy
DQE selects about half of a training set via greedy sampling plus similarity-based categorization of uncovered, difficult, and noisy samples, and reports improved text classification accuracy over full-data fine-tuning.
-
Large Language Models as Robust Data Generators in Software Analytics: Are We There Yet?
Pre-trained models fine-tuned on LLM-generated data are less robust to adversarial attacks than models fine-tuned on human-written data in three software analytics tasks.
-
Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration
The paper proposes a systematic classification and two roadmaps for using foundation models (LLMs and wireless foundation models) to design Synesthesia of Machines systems for 6G, with preliminary case-study evidence ...
-
Analyzing Patient Daily Movement Behavior Dynamics Using Two-Stage Encoding Model
A two-stage pipeline encodes daily home-activity text with MiniLM, clusters the embeddings, and applies PageRank to derive per-patient behavior vectors.
-
Two-Stage Representation Learning for Analyzing Movement Behavior Dynamics in People Living with Dementia
A two-stage pipeline of LLM text encoding and PageRank over transition clusters yields 5-element state vectors that modestly predict MMSE and ADAS-Cog scores in 50 dementia patients.
-
The Application of MATEC (Multi-AI Agent Team Care) Framework in Sepsis Care
A 10-physician pilot rated a multi-agent LLM team for sepsis care as useful and accurate, but accuracy was self-reported without ground truth.
-
Mitigating Sycophancy in Decoder-Only Transformer Architectures: Synthetic Data Intervention
Synthetic data intervention is reported to reduce sycophancy in GPT-4o on 100 true-false questions, but the experiment does not establish that the model was actually trained.
Discussion (0). Continue with ORCID to comment.