REVIEW 12 cited by
Self-Consuming Generative Models Go MAD
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Seismic advances in generative AI algorithms for imagery, text, and other data types has led to the temptation to use synthetic data to train next-generation models. Repeating this process creates an autophagous (self-consuming) loop whose properties are poorly understood. We conduct a thorough analytical and empirical analysis using state-of-the-art generative image models of three families of autophagous loops that differ in how fixed or fresh real training data is available through the generations of training and in whether the samples from previous generation models have been biased to trade off data quality versus diversity. Our primary conclusion across all scenarios is that without enough fresh real data in each generation of an autophagous loop, future generative models are doomed to have their quality (precision) or diversity (recall) progressively decrease. We term this condition Model Autophagy Disorder (MAD), making analogy to mad cow disease.
Forward citations
Cited by 12 Pith papers
-
Black-Box Detection of LLM-Generated Text Using Generalized Jensen-Shannon Divergence
SurpMark detects machine-generated text by estimating state-transition matrices from discretized surprisals and scoring them with generalized Jensen-Shannon divergence to human versus machine references.
-
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
A new modern Chinese poetry detection benchmark shows most current AI-text detectors are unreliable, particularly when LLMs imitate a human style.
-
Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English
After ChatGPT's release, science and tech podcast speakers used AI-associated words like 'surpass' and 'align' more often, while control synonyms showed no average shift.
-
Exploring the Structure of AI-Induced Language Change in Scientific English
In PubMed abstracts, AI-associated 'spiking' words rise together with their synonyms rather than replacing them, and declining words show less systematic, more organic patterns.
-
Ambient Diffusion Omni: Training Good Models with Bad Data
Ambient Diffusion Omni trains diffusion models on mixed-quality data by learning when corrupted images can be treated as clean, improving generation quality and diversity.
-
Model Collapse: On Recursion, Noise, and Uncharted Machine Visions
Model collapse, like video feedback, turns machine vision inward to generate self-referential worlds, revealing AI’s reliance on human content and inviting an aesthetics of noise and recursion against transhumanism.
-
Epistemic diversity across language models mitigates knowledge collapse
In repeated self-training loops on Wikitext2, ecosystems of four small language models show lower average perplexity than one, two, or sixteen models, but the paper's broader claims about monotonic optima, robustness,...
-
From trigonal to triclinic: Symmetry-tuned Rashba effects in buckled honeycomb SrHfO$_{3}$-based heterostructures
DFT calculations predict a Rashba coefficient of 0.34 eV-Angstrom and 29 meV band splitting in the non-centrosymmetric P1 phase of a (SrHfO3)2/(LaAlO3)4(111) superlattice.
-
Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
ML conferences should create an official peer-reviewed track dedicated to refuting and critiquing previously published work.
-
Kaleidoscope Gallery: Exploring Ethics and Generative AI Through Art
Ethics experts' definitions of five ethical theories, rendered as DALL-E 3 images and re-evaluated by the same experts, yield eight themes showing how morality, society, and learned associations shape and bias the mod...
-
Optimal Self-Distillation for Rectified Flow via Linear Probing
For linear rectified flow with ridge regression on fixed interpolants, optimally mixed self-distillation strictly improves velocity risk whenever the teacher is off the ridge stationary point, with a closed-form mixin...
-
LLM Web Dynamics: Tracing Model Collapse in a Network of LLMs
Under a shared retrieval-augmented memory, multiple LLMs' outputs converge to near-identical semantic answers, and the analogous Gaussian mixture system is proven to collapse.
Discussion (0). Sign in to comment.