Pith. sign in

REVIEW 7 cited by

How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.05090 v1 pith:62JL6FYC submitted 2024-04-07 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords modelcollapsedatasyntheticmodelstrainingwhenavoided
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The phenomenon of model collapse, introduced in (Shumailov et al., 2023), refers to the deterioration in performance that occurs when new models are trained on synthetic data generated from previously trained models. This recursive training loop makes the tails of the original distribution disappear, thereby making future-generation models forget about the initial (real) distribution. With the aim of rigorously understanding model collapse in language models, we consider in this paper a statistical model that allows us to characterize the impact of various recursive training scenarios. Specifically, we demonstrate that model collapse cannot be avoided when training solely on synthetic data. However, when mixing both real and synthetic data, we provide an estimate of a maximal amount of synthetic data below which model collapse can eventually be avoided. Our theoretical conclusions are further supported by empirical validations.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Generative Foundation Model for Chest Radiography

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A diffusion-based generative model for chest X-rays, trained on 960k image-report pairs, improves downstream classification, segmentation, detection, and fairness when its outputs are used for data augmentation or pre...

  2. What Matters in LLM-generated Data: Diversity and Its Effect on Model Fine-Tuning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Moderately diverse LLM-generated data can improve fine-tuned model performance in low-data settings when distribution shift is minimal, while high diversity or large distribution shift hurts.

  3. Unlocking Speech Instruction Data Potential with Query Rewriting

    cs.AI 2025-07 conditional novelty 5.0 of 10

    A multi-LLM rewriting and multi-agent validation pipeline makes text-to-speech synthesized speech instruction data far more usable and improves downstream speech instruction following.

  4. Using Sign Language Production as Data Augmentation to enhance Sign Language Translation

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Adding synthetic sign-language data produced by stitching, a GAN, or Gaussian splatting to the training set improves sign-language translation, with the largest gains for skeleton-pose models.

  5. A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations

    cs.CL 2025-07 conditional novelty 4.0 of 10

    PATTR adds a target-length penalty to the Type-Token Ratio, producing a lexical diversity score with tunable, reduced short-text bias for LLM synthetic data.

  6. LLM Web Dynamics: Tracing Model Collapse in a Network of LLMs

    cs.LG 2025-05 conditional novelty 4.0 of 10

    Under a shared retrieval-augmented memory, multiple LLMs' outputs converge to near-identical semantic answers, and the analogous Gaussian mixture system is proven to collapse.

  7. The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback

    cs.LG 2025-09 reject novelty 3.0 of 10

    A recursive fine-tuning study claims quality filtering reverses model collapse, but the paper's own data show the filtered model only matched its starting score.

Pith tools