Pith. sign in

REVIEW 2 cited by

Evaluating Language Models as Synthetic Data Generators

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.03679 v2 pith:4BK5RD6T submitted 2024-12-04 cs.CL

Evaluating Language Models as Synthetic Data Generators

classification cs.CL
keywords datagenerationabilitybettergeneratorslanguagemodelmodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Given the increasing use of synthetic data in language model (LM) post-training, an LM's ability to generate high-quality data has become nearly as crucial as its ability to solve problems directly. While prior works have focused on developing effective data generation methods, they lack systematic comparison of different LMs as data generators in a unified setting. To address this gap, we propose AgoraBench, a benchmark that provides standardized settings and metrics to evaluate LMs' data generation abilities. Through synthesizing 1.26 million training instances using 6 LMs and training 99 student models, we uncover key insights about LMs' data generation capabilities. First, we observe that LMs exhibit distinct strengths. For instance, GPT-4o excels at generating new problems, while Claude-3.5-Sonnet performs better at enhancing existing ones. Furthermore, our analysis reveals that an LM's data generation ability doesn't necessarily correlate with its problem-solving ability. Instead, multiple intrinsic features of data quality-including response quality, perplexity, and instruction difficulty-collectively serve as better indicators. Finally, we demonstrate that strategic choices in output format and cost-conscious model selection significantly impact data generation effectiveness.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Modular Techniques for Synthetic Long-Context Data Generation in Language Model Training and Evaluation

    cs.CL 2025-09 reject novelty 3.0

    A synthetic long-context data generation framework is described, but with no empirical evaluation or comparison to existing methods.

  2. Generative Models for Synthetic Data: Transforming Data Mining in the GenAI Era

    cs.LG 2025-08 unverdicted novelty 1.0

    A tutorial proposal outlining how generative models can synthesize data across modalities for data mining, with no new research results.