REVIEW 2 major objections 5 minor 33 references
Model Collapse: On Recursion, Noise, and Uncharted Machine Visions
T0 review · 2 major / 5 minor · reviewed 2026-07-14 · grok-4.5
Pith's one-line read Model collapse is not only AI failure but a recursive mirror that, like video feedback, generates inward machine visions while proving systems still need human noise.
desk verdict Solid media-archaeology essay that reframes documented model collapse via video feedback; interpretive, not technical, and fine for its venue. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The formal and operational parallel between analog video feedback (optical and mixer loops that autonomously generate self-perpetuating abstractions) and recursive AI training that induces model collapse; this mechanism reframes positive-feedback noise as generative for art and as diagnostic of AI’s irreducible data dependency.
What would settle it
A series of controlled recursive-training runs on open diffusion or language models that produce only sterile, non-patterned degradation with no morphogenetic or self-sustaining visual/linguistic structures comparable to historical video feedback, while simultaneously showing that the same models maintain long-term coherence indefinitely on purely synthetic data.
Extended reading notes
Core claim
Recursive training that produces model collapse challenges transhumanist ideals of self-sustaining AI while inviting an aesthetic perspective: noise and recursion become key concepts for both artmaking and the AI ecosystem. Collapse turns so-called machine vision inward so that it generates worlds from within rather than transmitting external reality, yet the entire system still depends on continual human-generated content, especially for foundation models trained on massive datasets.
Load-bearing premise
The visual and conceptual resemblances between early video-feedback abstractions and the noisy, repetitive outputs of recursively trained models are strong enough to treat collapse as a creative and epistemological opportunity rather than pure technical failure.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reframes model collapse—the progressive degradation of generative models when recursively trained on their own synthetic outputs—as both an engineering failure (word repetition, pixel noise, loss of coherence) and an aesthetic/epistemological opportunity. Through media archaeology, it draws formal and conceptual parallels between analog video synthesis techniques (optical feedback and mixer/system feedback from the 1960s–80s, illustrated by Lodge, Donebauer, Crutchfield, and contemporary Phase Shift works) and recursive AI training (citing Shumailov et al. 2024, Boháček & Farid 2023, and artistic experiments by Kraft, Chatonsky, Liu/Kühr). Anchored in the author’s prior definition of art as noise magnification (versus engineering’s noise suppression), it argues that collapse functions as a recursive mirror in which machine vision generates worlds from within rather than transmitting the external world. This reveals foundation models’ ongoing dependency on fresh human-produced data, thereby challenging transhumanist ideals of autonomous machine intelligence while positioning noise and recursion as central to both artmaking and the AI ecosystem.
Significance. If the interpretive reframing holds, the essay offers a timely, non-alarmist contribution to the ethics and aesthetics of AI by recuperating documented failure modes of foundation models as sites of creative and critical potential. It productively links cybernetic notions of positive feedback, historical video art, and current technical literature on model collapse, while underscoring the material data dependency that undercuts claims of self-sustaining AI. Strengths include accurate citation of engineering results (Shumailov, Boháček/Farid, Alemohammad), concrete case studies of both historical and contemporary works, and a clear challenge to extractive/transhumanist narratives. The piece is well-suited to a volume on ethics and aesthetics of AI and advances a networked, multi-scalar view of agency without overclaiming causal or quantitative equivalence between video feedback and recursive training.
major comments (2)
- [§2 Model Collapse: Positive Feedback, Spiraling Entropy] §2 (Model Collapse) and the bridge from §1: The central analogy between mixer feedback (self-referential loops with no external camera input, producing autonomous RGB flows) and recursive training on synthetic data is conceptually productive and explicitly framed as aesthetic rather than causal. However, the load-bearing claim that collapse “generates worlds from within” would be strengthened by a more explicit discussion of the differences in temporality, scale, and agency (real-time embodied knob-turning vs. offline statistical iteration across massive datasets). Without this, the media-archaeological parallel risks remaining suggestive rather than fully diagnostic of what recursive training uniquely reveals about AI data dependency.
- [§3 Data Dependency: A Transhumanist Dystopia] §3 and Conclusion: The challenge to transhumanist ideals rests on the well-supported observation that foundation models require continual human-generated content (Meta 2025 scraping request; Gibney 2024; Shumailov et al.). Chatonsky’s Disnovation is an effective speculative foil, yet the leap from this dystopian installation (and the space-race rhetoric of Bezos/Musk) to a general critique of “transhumanist ideals” could be tightened. Clarifying which specific claims of autonomy or transcendence are empirically undermined by data dependency—versus which remain speculative—would make the political-aesthetic argument more precise without diluting its force.
minor comments (5)
- [§2] §2, paragraph on GANs: typographical slip “one the one hand” should be “on the one hand”.
- [Figures 1–7] Figures 1–7 are well-chosen and described, but the manuscript would benefit from brief captions that more explicitly flag the recursive iteration count or technical setup (already partially present) so readers can track the visual degradation without returning to the body text.
- [Introduction] The self-citation to the author’s 2023 NECSUS article and 2024 PhD is appropriate as background for the “art as noise magnification” definition, yet a single clarifying sentence early in the Introduction stating that this is a prior working definition (not derived here) would reduce any appearance of circularity for readers unfamiliar with the earlier work.
- [Conclusion] Conclusion’s closing speculation on recursion, error-correcting codes, supersymmetry, and “the fabric of the cosmos” is intriguing but abrupt; a shorter, more tightly linked final paragraph would better preserve the essay’s focus on AI aesthetics and data dependency.
- [§2 and References] Minor orthographic consistency: “Ègor Kraft” appears with grave accent; confirm preferred artist spelling across text and references.
Circularity Check
No significant circularity: interpretive media-archaeology essay uses author's prior aesthetic definition as framing background, not as a load-bearing derivation that reduces claims by construction.
-
self citation load bearing
[Introduction (pp. 1–2) and §2 (p. 6)]
"art strives to magnify the noise within and via any given medium, as opposed to engineering that seeks to suppress it (Boutet de Monvel, 2023, p. 121). ... I have even argued, in my PhD, that it constitutes the condition sine qua non for the renewal and diversification of art forms against the engineering imperative to minimize uncertainty, formalized as: regenerated art / message (output) = original art / signal (input) + noise (Boutet de Monvel, 2024, p. 60)."
The paper's aesthetic reframing of model collapse as noise magnification and positive-feedback opportunity rests on the author's own prior definition and formula, cited from her 2023 NECSUS article and 2024 PhD. This is ordinary self-citation of background framing rather than a circular derivation: the definition is not claimed to be proved by the present case studies, nor does any empirical or formal result of the paper reduce to it by construction. The engineering facts about collapse and data dependency remain externally sourced.
full rationale
This is an aesthetics and media-archaeology essay, not a technical derivation paper. It has no equations, fitted parameters, uniqueness theorems, or quantitative predictions that could reduce to inputs by construction. Model-collapse phenomena (word repetition, pixel noise, performance degradation) and data-dependency claims are taken from independent external sources (Shumailov et al. 2024, Boháček & Farid 2023, Alemohammad et al. 2023, Gibney 2024, etc.). Historical video-feedback cases (Lodge, Donebauer, Crutchfield, Kop/Jay) are likewise external. The sole self-citations are to the author's own 2021 seminar, 2023 NECSUS article, and 2024 PhD for the prior definition of art as noise magnification and the related formula regenerated art = signal + noise; these supply the interpretive lens that reframes collapse as aesthetic opportunity, but they are presented as background rather than derived from or forced by the present paper's claims. The central argument (collapse as recursive mirror challenging transhumanism while inviting aesthetic attention to noise/recursion) therefore remains independent of any circular reduction. Score 1 reflects only the minor, non-load-bearing self-citation of the framing definition; no pattern 1–6 is present at a level that forces the result.
Assumptions & free parameters
assumptions (4)
- ad hoc to paper Art strives to magnify the noise within and via any given medium, as opposed to engineering that seeks to suppress it (author’s 2021/2023 definition).
- domain assumption Shannon’s linear communication model and Wiener’s distinction between negative (homeostatic) and positive (amplifying) feedback apply to both analog video loops and recursive AI training.
- domain assumption Foundation models cannot sustain meaningful variation indefinitely without regular access to substantial amounts of fresh human-generated content.
- domain assumption Optical and system feedback in analog video produce self-sustaining machine visions that no longer transmit an external world.
Cite this review
Pith. "Pith review of Model Collapse: On Recursion, Noise, and Uncharted Machine Visions." pith.science (2026). https://pith.science/paper/YEXTS7UF
@misc{pith2026260709705,
author = {Pith},
title = {Pith review of: Model Collapse: On Recursion, Noise, and Uncharted Machine Visions},
year = {2026},
howpublished = {\url{https://pith.science/paper/YEXTS7UF}},
note = {Machine review of arXiv:2607.09705}
}
read the original abstract
Since 2023, computer scientists have warned against model collapse -- the contamination of training sets with AI-generated outputs that progressively degrade model performance. Exemplifying a positive-feedback-driven failure, it produces effects such as word repetition or pixel noise, ultimately leading to a loss of meaning and coherence -- at least from an engineering standpoint. From a creative one, however, collapse is not merely a breakdown: it also functions as a recursive mirror that recalls early analog video feedback experiments, raising once again the question of what happens when a system turns inward and sees itself. In such cases, so-called machine vision no longer transmits the world (as in tele-vision) but increasingly generates worlds from within. Drawing on media archaeology through case studies of both historical video synthesis techniques and contemporary artistic uses of machine learning, this paper examines what recursive training reveals about the dependent nature of AI-generated data. It argues that the potential effects of collapse challenge transhumanist ideals while inviting an aesthetic perspective, positioning noise and recursion as key concepts for understanding both artmaking and the AI ecosystem. Distributing agency across scales and networks, the latter currently remains reliant on new human-produced content, particularly within foundation models trained on massive datasets.
Figures
Reference graph
Works this paper leans on
-
[1]
1 Model Collapse: On Recursion, Noise, and Uncharted Machine Visions Violaine Boutet de Monvel violaine.boutetdemonvel@biblhertz.it1 Bibliotheca Hertziana – Max Planck Institute for Art History MPRG Impett – Machine Visual Culture: Artificial Intelligence and the History of Seeing Abstract: Since 2023, computer scientists have warned against model collaps...
2023
-
[2]
My aim was to reframe real-time analog synthesis – particularly video feedback techniques – as a precursor to current machine-learning-based art
and soon after diffusion models (Sohl-Dickstein, et al., 2015). My aim was to reframe real-time analog synthesis – particularly video feedback techniques – as a precursor to current machine-learning-based art. To this end, I compared the room left for human mastery in the recursive operations of both closed-circuit video setups and generative AI models, w...
2015
-
[3]
2 via any given medium, as opposed to engineering that seeks to suppress it (Boutet de Monvel, 2023, p. 121). In this context, noise refers to any agency that may introduce interference into the signal path beyond the artist’s design – for example, the randomness of audience behavior in happenings, environmental contingencies like the weather, not to ment...
2023
-
[4]
These operations are now distributed across scales and networks, from training sets to latent spaces, and they unfold at speeds that far exceed human perception (Hansen, 2014, p
Video Synthesis: Optical and Mixer Feedback From a human-machine interaction standpoint, comparing past signal-based image production with today’s generative AI is valuable because the former offers a situated, embodied perspective on synthetic processes that are otherwise difficult to grasp. These operations are now distributed across scales and networks...
2014
-
[5]
Yet the analogy ends here: in video, the outcome is far more hypnotic
3 its own monitor – the visual equivalent of bringing a microphone close to its amplifier, whose resulting audio feedback is harsh to the ear and conventionally regarded as a malfunction. Yet the analogy ends here: in video, the outcome is far more hypnotic. When no tangible object intervenes, the continuous, abstracted flow of the apparatus looped back o...
1963
-
[6]
the experience of a foetus inside a mother’s womb
The images show optical feedback effects produced using standard broadcast equipment. Such a basic video feedback setup – achievable, in practice, with any consumer device – can autonomously generate self-perpetuating inward abstractions, while further zooming in and pivoting the camera along the monitor’s X and Y axes give rise to ever more intricate vis...
1974
-
[7]
3 See Peter Donebauer, Entering, video, BBC, London, UK, 6 min 50 s,
Retrieved October 10, 2025, from https://youtu.be/sjDFvoRNpOM. 3 See Peter Donebauer, Entering, video, BBC, London, UK, 6 min 50 s,
2025
-
[8]
4 See James P
Retrieved October 10, 2025, from https://archive.org/details/1974PeterDONEBAUERENTERING720. 4 See James P. Crutchfield, Space-Time Dynamics in Video Feedback, video, Entropy Productions, Santa Cruz, USA, 15 min 33 s,
2025
Show all 33 references
-
[9]
5 See Bruce Gowers, Bohemian Rhapsody, promotional video, Elstree Studios, Borehamwood, UK, 6 min,
Retrieved October 10, 2025, from https://youtu.be/B4Kn3djJMCE. 5 See Bruce Gowers, Bohemian Rhapsody, promotional video, Elstree Studios, Borehamwood, UK, 6 min,
2025
-
[10]
Preprint for Ethics and Aesthetics of Artificial Intelligence (Naples: Orthotes Edizioni), October
Retrieved October 10, 2025, from https://youtu.be/fJ9rUzIMcZQ. Preprint for Ethics and Aesthetics of Artificial Intelligence (Naples: Orthotes Edizioni), October
2025
-
[11]
The images show optical feedback effects produced using standard broadcast equipment. This introduces the other major technique in video synthesis: system feedback, which – unlike optical feedback – requires no camera, only a monitor connected to a real-time image processing t...
2019
-
[12]
Preprint for Ethics and Aesthetics of Artificial Intelligence (Naples: Orthotes Edizioni), October
Retrieved October 10, 2025, from https://youtu.be/WgbscQ34rBk. Preprint for Ethics and Aesthetics of Artificial Intelligence (Naples: Orthotes Edizioni), October
2025
-
[13]
5 Both optical and system feedback techniques – which can, of course, be combined (Gwin, 1972, p
1972
-
[14]
– have the power to generate, in real-time, an infinite, complex, and ever-changing mise en abyme of the external setup and/or its internal circuitry. They produce self-sustaining machine visions that, paradoxically, appear life-like precisely because they maintain acute sensi...
1973
-
[15]
My definition of art as noise magnification naturally led to wonder whether an equivalent of mixer feedback might exist in diffusion-based text-to-image models
Model Collapse: Positive Feedback, Spiraling Entropy In my PhD, I adopted a media-archaeological approach to show how earlier techniques of video or even vector synthesis could shed light on the recursive operations of generative AI. My definition of art as noise magnification...
2023
-
[16]
Yet, as noted earlier, positive feedback may also be harnessed creatively to produce difference within repetition
6 By contrast, positive feedback amplifies noise, often with destructive consequences – death, for example, when the body can no longer regulate itself. Yet, as noted earlier, positive feedback may also be harnessed creatively to produce difference within repetition. I have ev...
2024
-
[17]
older hispanic man
The images show Stable Diffusion outputs after the first and sixth iterations of recursive training, respectively, starting from a 50% corrupted dataset and the prompt “older hispanic man.” One of the most comprehensive studies on model collapse, led by Ilia Shumailov at Oxfor...
2024
-
[18]
older hispanic man
At Stanford, Matyás Boháček and Hany Farid used Stable Diffusion to generate successive training sets from a single prompt: “older hispanic man” (Boháček, Farid, 2023, p. 2). With each iteration, the resulting images visibly degraded, accumulating pixel noise (Figure 4). Anoth...
2023
-
[19]
Viewed through the lens of my definition of art as noise magnification, however, phenomena deemed problematic from an engineering standpoint may instead present creative potential
– while ethology offers an even starker metaphor: data cannibalism (Pedersen, 2023). Viewed through the lens of my definition of art as noise magnification, however, phenomena deemed problematic from an engineering standpoint may instead present creative potential. Preprint fo...
2023
-
[20]
a single chair on a white background
7 Ègor Kraft’s 2023 recursive video One and Infinite Chairs provides an early example of artistic engagement with model collapse. It was inspired by Joseph Kosuth’s 1965 conceptual assemblage One and Three Chairs – in which a wooden chair, its photograph, and a textual definit...
2023
-
[21]
Preprint for Ethics and Aesthetics of Artificial Intelligence (Naples: Orthotes Edizioni), October
Retrieved October 10, 2025, from https://kraft.studio/chair/. Preprint for Ethics and Aesthetics of Artificial Intelligence (Naples: Orthotes Edizioni), October
2025
-
[22]
the richness and variety that comes with human-generated content
8 ideally shift discussions about creativity toward broader scalar dynamics beyond sole human agency (Krysztoforska, Kenny, 2024). In practice, however, the phenomenon demonstrates the opposite: current frontier models cannot sustain meaningful variation indefinitely without r...
2024
-
[23]
What, then, does this data dependency reveal about the broader AI ecosystem? A transhumanist dystopia imagined by Grégory Chatonsky as early as 2022 offers a partial answer
– in other words, external noise. What, then, does this data dependency reveal about the broader AI ecosystem? A transhumanist dystopia imagined by Grégory Chatonsky as early as 2022 offers a partial answer. Titled Disnovation, the work took the form of a multimedia installati...
2022
-
[24]
Preprint for Ethics and Aesthetics of Artificial Intelligence (Naples: Orthotes Edizioni), October
Retrieved October 10, 2025, from https://youtu.be/tYZOXVimcN0. Preprint for Ethics and Aesthetics of Artificial Intelligence (Naples: Orthotes Edizioni), October
2025
-
[25]
Their appearances had been drastically aged
9 Beyond the artist’s prophetic figure, the installation also included ten silent portraits representing GAFAM leaders – including Mark Zuckerberg, Jeff Bezos, and Elon Musk, the respective heads of Meta, Amazon, and Tesla (Figure 6). Their appearances had been drastically age...
2022
-
[26]
Both aim to develop space tourism for a billionaire clientele and plan inhabited missions to the Moon by the end of the 2020s, as part of NASA’s Commercial Lunar Payload Services (CLPS) and Artemis programs – potentially offering, if not a solution, at least an escape from Ear...
2025
-
[27]
Sputnik Moment for AI
10 Like video feedback before it, this territory promises to be fertile yet calls for further artistic exploration. Meanwhile, one point remains clear: deep learning, for all its scale, still depends on continual access to vast quantities of high-quality human data, inherently...
1958
-
[28]
Alemohammad S., et al
Retrieved October 10, 2025, from https://www.project-syndicate.org/commentary/china-ai-deepseek-raises-difficult-questions-for-united-states-by-daron-acemoglu-2025-02. Alemohammad S., et al. (2023), ‘Self-Consuming Generative Models Go MAD’. DOI: https://doi.org/10.48550/arXiv...
-
[29]
Chatonsky G
11 — (2021, November), Court-circuiter de concert le fil de l’information: quand le bruit fait œuvre, Conference presentation at the Séminaire du CREAViS, Paris: Université Sorbonne Nouvelle. Chatonsky G. (2022), Disnovation. Retrieved October 10, 2025, from https://chatonsky....
2021
-
[30]
Donebauer P
Retrieved October 10, 2025, from https://www.euronews.com/next/2025/05/13/meta-is-about-to-use-europeans-social-posts-to-train-its-ai-heres-how-you-can-prevent-it. Donebauer P. (1975), ‘Electronic Painting’, Video and Audio-Visual Review, vol. 1, pp. 30-33. Doran C.F., et al. ...
-
[31]
(2024), ‘The infinite as paradigm: Reframing the limits of AI art’, NECSUS European Journal of Media Studies
12 Krysztoforska M., Kenny O. (2024), ‘The infinite as paradigm: Reframing the limits of AI art’, NECSUS European Journal of Media Studies. #Enough, vol. 13 (2), pp. 132-155. DOI: http://dx.doi.org/10.25969/mediarep/23670. Liu T.C., Kühr L.E. (2025), Arafed Futures. Retrieved ...
2024 doi
-
[32]
Shumailov I., et al
Retrieved October 10, 2025, from https://medium.com/artificial-corner/the-consequence-of-data-cannibalism-gpt-4-and-the-jpeg-effect-1ea38faf9e82. Shumailov I., et al. (2024), ‘AI models collapse when trained on recursively generated data’, Nature, vol. 631, pp. 755-59. Sohl-Di...
-
[33]
(2023), ‘Algorithmic Images: Artificial Intelligence and Visual Culture’, Grey Room, no
Somaini A. (2023), ‘Algorithmic Images: Artificial Intelligence and Visual Culture’, Grey Room, no. 93, pp. 74-115. Wenger E. (2024), ‘AI returns gibberish when trained on generated data’, Nature, vol. 631, pp. 742-743. Wiener N. (1948), Cybernetics: Or Control and Communicati...
2023
Reviewed July 14, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.