Pith. sign in

Data Generation Using Large Language Models for Text Classification: An Empirical Case Study

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Using Large Language Models (LLMs) to generate synthetic data for model training has become increasingly popular in recent years. While LLMs are capable of producing realistic training data, the effectiveness of data generation is influenced by various factors, including the choice of prompt, task complexity, and the quality, quantity, and diversity of the generated data. In this work, we focus exclusively on using synthetic data for text classification tasks. Specifically, we use natural language understanding (NLU) models trained on synthetic data to assess the quality of synthetic data from different generation approaches. This work provides an empirical analysis of the impact of these factors and offers recommendations for better data generation practices.

fields

cs.CL 1

years

2025 1

verdicts

UNVERDICTED 1

representative citing papers

Do Biased Models Have Biased Thoughts?

cs.CL · 2025-08-08 · unverdicted · novelty 3.0

The manuscript is internally inconsistent: the abstract describes an LLM fairness experiment while the body is a different paper on pilot-wave quantum mechanics, so no coherent result can be assessed.

citing papers explorer

Showing 1 of 1 citing paper.

  • Do Biased Models Have Biased Thoughts? cs.CL · 2025-08-08 · unverdicted · none · ref 31 · internal anchor

    The manuscript is internally inconsistent: the abstract describes an LLM fairness experiment while the body is a different paper on pilot-wave quantum mechanics, so no coherent result can be assessed.