The authors generate and publicly release the first large-scale open dataset of three million structured moral fables produced by small open language models together with a reproducible LLM-judge evaluation pipeline.
Synthetic Data Generation Using Large Language Models: Advances in Text and Code, March 2025
3 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
A 21-configuration LLM judge ensemble (7 models × 3 prompts) validates synthetic e-commerce attribute labels at 95.2% agreement with human experts across 4 languages and 12,726 products.
S^3-R1 generates synthetic multi-hop questions and uses combined intermediate and final rewards to train RL models for retrieval and answering, reporting up to 10% better out-of-domain generalization.
citing papers explorer
-
TF1-EN-3M: Three Million Synthetic Moral Fables for Training Small, Open Language Models
The authors generate and publicly release the first large-scale open dataset of three million structured moral fables produced by small open language models together with a reproducible LLM-judge evaluation pipeline.
-
SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation
A 21-configuration LLM judge ensemble (7 models × 3 prompts) validates synthetic e-commerce attribute labels at 95.2% agreement with human experts across 4 languages and 12,726 products.
-
$S^3$-R1: Learning to Retrieve and Answer Step-by-Step with Synthetic Data
S^3-R1 generates synthetic multi-hop questions and uses combined intermediate and final rewards to train RL models for retrieval and answering, reporting up to 10% better out-of-domain generalization.