Pith. sign in

REVIEW 8 cited by

A Tale of Tails: Model Collapse as a Change of Scaling Laws

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.07043 v2 pith:JYMI52FY submitted 2024-02-10 cs.LG cs.AIcs.CL

A Tale of Tails: Model Collapse as a Change of Scaling Laws

classification cs.LG cs.AIcs.CL
keywords scalingdatamodellawscollapsemodelswillchange
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

As AI model size grows, neural scaling laws have become a crucial tool to predict the improvements of large models when increasing capacity and the size of original (human or natural) training data. Yet, the widespread use of popular models means that the ecosystem of online data and text will co-evolve to progressively contain increased amounts of synthesized data. In this paper we ask: How will the scaling laws change in the inevitable regime where synthetic data makes its way into the training corpus? Will future models, still improve, or be doomed to degenerate up to total (model) collapse? We develop a theoretical framework of model collapse through the lens of scaling laws. We discover a wide range of decay phenomena, analyzing loss of scaling, shifted scaling with number of generations, the ''un-learning" of skills, and grokking when mixing human and synthesized data. Our theory is validated by large-scale experiments with a transformer on an arithmetic task and text generation using the large language model Llama2.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs

    cs.SE 2026-06 unverdicted novelty 6.0

    Experiments across code LLMs show no-review collapses fastest, human-gated filters slow collapse, and AI self-gates lose effect over time, degenerating to ungated self-training under self-confirming acceptance as prov...

  2. Generative artificial intelligence reduces social welfare through model collapse

    physics.soc-ph 2026-04 unverdicted novelty 6.0

    A game-theoretic model shows that individually rational adoption of generative AI causes model collapse that reduces collective social welfare for important tasks, with habit formation creating spillovers from low-sta...

  3. Scaling Synthetic Data Creation with 1,000,000,000 Personas

    cs.CL 2024-06 unverdicted novelty 6.0

    A curated set of one billion personas enables scalable, diverse synthetic data generation for LLM training across reasoning, instructions, knowledge, NPCs, and tools.

  4. The Crowded Embedding Space: A Mean-Field Mechanism for Emergent Marginalization in Retrieval-Augmented Agents

    cs.IR 2026-06 unverdicted novelty 5.0

    A mean-field analysis of embedding-space crowding shows a phase transition and Fokker-Planck dynamics that drive retrieval-augmented agents to self-organize toward exclusive service of majority interests.

  5. Epistemic diversity across language models mitigates knowledge collapse

    cs.LG 2025-12 reject novelty 5.0

    In repeated self-training loops on Wikitext2, ecosystems of four small language models show lower average perplexity than one, two, or sixteen models, but the paper's broader claims about monotonic optima, robustness,...

  6. Statistical Properties of Training & Generalization

    stat.ML 2026-06 unverdicted novelty 2.0

    Neural scaling laws in deep learning interact with physics constraints and inductive biases beyond classical statistics.

  7. Securing High-Concurrency Ticket Sales: A Framework Based on Microservice

    cs.SE 2025-12 reject novelty 2.0

    A railway ticketing system built from standard Spring Cloud components is reported to reach 817 req/s on a train-query interface, but only under a 100-thread local VM test with inconsistent purchase-interface data.

  8. Statistical Properties of Training & Generalization

    stat.ML 2026-06 unverdicted novelty 1.0

    Review of neural scaling laws and their relation to constraints and inductive biases when applying machine learning to physics problems.