Pith. sign in

REVIEW 2 cited by

Data Generation Using Large Language Models for Text Classification: An Empirical Case Study

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12813 v2 pith:WJGC7A3P submitted 2024-06-27 cs.CL cs.AI

classification cs.CLcs.AI
keywords datagenerationsyntheticlanguagemodelsclassificationempiricalfactors
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Using Large Language Models (LLMs) to generate synthetic data for model training has become increasingly popular in recent years. While LLMs are capable of producing realistic training data, the effectiveness of data generation is influenced by various factors, including the choice of prompt, task complexity, and the quality, quantity, and diversity of the generated data. In this work, we focus exclusively on using synthetic data for text classification tasks. Specifically, we use natural language understanding (NLU) models trained on synthetic data to assess the quality of synthetic data from different generation approaches. This work provides an empirical analysis of the impact of these factors and offers recommendations for better data generation practices.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Does Prompt Design Impact Quality of Data Imputation by LLMs?

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Group-wise CSV prompts with correlation-based column pruning reduce LLM imputation prompt size while roughly maintaining or slightly improving classifier-based imputation quality on two imbalanced datasets.

  2. Do Biased Models Have Biased Thoughts?

    cs.CL 2025-08 unverdicted novelty 3.0 of 10

    The manuscript is internally inconsistent: the abstract describes an LLM fairness experiment while the body is a different paper on pilot-wave quantum mechanics, so no coherent result can be assessed.

Pith tools