REVIEW 4 major objections 5 minor 18 references
Next Token Prediction Is a Dead End for Creativity
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper argues that next-token prediction is a dead end for genuine improvisational creativity, and that battle rap exposes why.
desk verdict The battle-rap framing is genuinely interesting, but the 'dead end' claim is a philosophical assertion dressed up as an empirical result, and even the authors' own proposed fixes rely on token predictors. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the autoregressive next-token prediction objective—maximizing the probability of the next token given all previous tokens—which the paper argues structurally optimizes for plausible continuation and cannot represent turn-taking, anticipation, rebuttal, rhythmic timing, or lived temporal experience. Against this, the paper sets a proposed co-creative system whose components are voice cloning, cadence modelling, and rhyme scoring operating in a real-time feedback loop, with flow reframed as an emergent property of the human-machine interaction rather than an internal state of the model.
What would settle it
A live battle-rap evaluation in which an LLM-based system with real-time cadence and adversarial feedback is judged by expert MCs and audiences as indistinguishable from a human freestyler on timing, rebuttal, and emotional impact would falsify the claim that token prediction cannot support improvisational co-creation.
Extended reading notes
Core claim
The authors claim that token prediction is a dead end for creativity in live, performative domains. They argue that next-token models optimize for surface-level coherence and plausible continuation, whereas improvisational co-creation such as battle rap requires anticipation, rebuttal, stylistic divergence, rhythmic alignment, and the capacity to lose oneself in the moment—capacities that presuppose agency, embodiment, and subjective temporality. They use battle rap to expose these gaps, propose a real-time improvisation system with voice cloning, cadence modelling, and rhyme scoring, and reframe creative authenticity as something that can be co-constituted between a human and a personalised AI rather than an exclusively human trait.
Load-bearing premise
The argument collapses if creativity is defined solely by the novelty and value of outputs rather than by the internal experience of the creator: the dead-end claim depends on authenticity requiring subjective experience, intentionality, and embodiment.
Editorial extensions
If this is right
- Scaling next-token prediction will not yield authentic improvisational co-creation in performative domains; performance on interactivity will plateau even as output fluency grows.
- Evaluation of machine creativity must shift from static output benchmarks to real-time measures of multimodal alignment, cadence, and adversarial responsiveness.
- AI's role in creative settings should be reframed as co-creative infrastructure that facilitates human flow, not as a standalone originator of creative acts.
- Model architectures should integrate planning, rhythm-awareness, and feedback sensitivity into generative loops rather than relying on unidirectional token streams.
- Dialogue itself should be reimagined as improvisation and co-authored action rather than as statistical completion of text.
Reading between the lines
- The battle-rap case is a compressed instance of a broader class of real-time performative domains—debate, negotiation, live comedy, therapy—where timing and turn-taking dominate; if the paper's architectural critique holds, those domains face the same limitation.
- A testable extension would compare an LLM-based battle system augmented with cadence and adversarial components against a plain next-token baseline in expert-judged live rounds; the paper proposes the system but does not run such an experiment, so the empirical gap between the philosophical claim and the proposed machinery remains open.
- The paper's co-creative framing implies that future benchmarks should score the human-AI dyad rather than the model's output alone, since authenticity is claimed to emerge from intentional, expressive alignment between user and system.
- If audience studies continue to show that people cannot distinguish AI-generated from human rap, the dead-end thesis becomes a normative claim about process rather than product, and an output-based definition of creativity would undercut it.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that next-token prediction is a dead end for creativity, particularly for real-time, interactive, and performative domains such as battle rap. The authors contend that autoregressive language models optimize for surface-level coherence and cannot genuinely engage in adversarial or emotionally resonant exchanges because they lack intrinsic motivation, intentionality, and self-aware agency. They propose a vision of co-creative human–AI systems built on interaction-first paradigms, multimodal feedback, cadence modeling, and hybrid generative loops, and they offer battle rap as a test bed for this agenda. The paper includes background on creativity research, a description of a proposed real-time improvisation system, and a discussion of alternative perspectives that both support and challenge the main thesis.
Significance. If the categorical claim were established, it would be a significant negative result for LLM-based creativity and a strong argument for reorienting AI research toward interactive, embodied, and goal-directed architectures. The paper has clear strengths: it engages with both supportive and contrary evidence, identifies a concrete and demanding domain (battle rap) that combines linguistic, rhythmic, and adversarial constraints, and proposes specific research challenges that could guide future work. It also cites recent empirical findings on LLM creativity, including human–AI indistinguishability in poetry, and it acknowledges the philosophical contestability of its core premise. However, the central conclusion is far stronger than the arguments and evidence provided, and the paper's own proposals are in tension with its 'dead end' framing. The contribution as written is a manifesto that would benefit from substantially revised claims and a clearer separation of philosophical assertion from demonstrated fact.
major comments (4)
- [Artificial vs Real Creativity] The section adopts Runco's premise that authentic creativity requires intrinsic motivation, intentionality, and self-aware agency, and it uses this premise to conclude that AI cannot be authentically creative. The paper never defends this premise against the output-based definitions it cites in 'Alternative Perspectives' ('creativity is commonly defined by the novelty and value of the outputs... regardless of how or why it was produced'). The categorical claims in the abstract and introduction, such as 'cannot truly engage in adversarial or emotionally resonant exchanges,' are therefore restatements of a contested philosophical position rather than established results. To support the 'dead end' thesis, the authors must either provide an argument for their preferred definition of creativity or explicitly restrict their impossibility claim to that definition.
- [Towards Real-Time Improvisation] The proposed system (Figure 3 and the surrounding text) integrates voice cloning, cadence modeling, rhyme scoring, Creative Beam Search, and feedback loops into a generative framework. All of these components are described as augmenting generators that still predict tokens. If these augmentations address the limitations identified in the paper, then next-token prediction is not a dead end but merely an insufficient objective; if they do not, the paper needs to explain how a non-predictive architecture would differ. The 'Conclusion' research challenges, which call for 'model architectures that integrate planning, rhythm-awareness, and feedback sensitivity into generative loops,' similarly presuppose a role for token-predicting components. This tension between the title's 'dead end' and the body's proposed improvements must be resolved.
- [Alternative Perspectives (footnote 2)] Footnote 2 concedes that most rap battles involve pre-written verses delivered with improvised timing or emphasis, and that fully improvised 'off-the-dome' freestyling is considerably harder and less common in formal competitions. This concession significantly weakens the paper's use of battle rap as a case study of real-time, spontaneous improvisation, which is the basis for the claim that next-token models cannot handle such domains. The argument may apply only to a narrow subset of battle rap, and the paper does not address this scope limitation when drawing its broad conclusions about the inadequacy of token prediction in 'live, creative exchanges.'
- [Abstract / Introduction] The abstract claims the paper 'demonstrates' that predictive models 'cannot truly engage' in adversarial or emotionally resonant exchanges, but the manuscript contains no empirical demonstration, no controlled comparison, and no formal proof of impossibility. The 'Alternative Perspectives' section itself cites evidence that humans cannot distinguish AI-generated poetry from human-written poetry and that LLMs can rival or surpass humans on divergent-thinking tasks; these findings are never refuted. As a position paper, the manuscript can argue for a research agenda, but the strong 'cannot' claims require either an argument that the cited evidence is irrelevant under the paper's definition of creativity or an explicit acknowledgment that the claim is a philosophical stance rather than a demonstrated result.
minor comments (5)
- [Background] There is a typo in 'Large Language Language Models' (LLMs) and a missing apostrophe in 'despite LLMs successes'; these should be corrected.
- [Modelling Human Responses] The sentence 'Vernon, Hocking, and Farahar 2016; and 2024; Marrone, Cropley, and and 2024' contains broken and garbled references; please repair these citations.
- [Conclusion] There is a double comma in 'As Csikszentmihalyi (Csikszentmihalyi 1997) and Runco (Runco 2023) state,, creativity is more than output'; remove the extra comma.
- [Alternative Perspectives] The citation 'Newkeen 2024' refers to a news article rather than a peer-reviewed source; the evidence for 'collective creativity' would be better supported by the underlying primary study or a more established reference.
- [General] The manuscript references Tables 1–4 and Figures 1–3, but these do not appear in the provided text; please ensure all tables and figures are embedded and clearly captioned in the submitted version.
Circularity Check
No circular derivation: the thesis is a stipulated definitional argument, and the only self-citation is not load-bearing.
full rationale
This is a position paper, not a derivational or empirical study. There are no equations, no fitted parameters, and no benchmark predictions, so the failure mode of a fitted input being renamed as a prediction is absent. The central claim that next-token prediction cannot 'truly engage' in creativity is obtained by adopting Runco's definition of authentic creativity as requiring intrinsic motivation, intentionality, and self-aware agency; the paper then observes that autoregressive models lack these. That is a valid syllogism from a stipulated premise, but the premise itself is contested, and the paper's own 'Alternative Perspectives' section concedes output-based definitions of creativity under which AI outputs qualify. Consequently, the conclusion is a restatement of a definitional commitment rather than a circular equation; it is an argumentative weakness, not a derivation-level circularity. The only same-author citation (Ọlátúnjí et al. 2025) is used to motivate battle rap as a case study and to support the claim that LLMs cannot model rhythmic constraints or beat alignment. This self-citation is background and supporting evidence; it is not the sole or primary justification of the dead-end thesis, which rests on the intentionality argument and on independent philosophical and empirical references. No uniqueness theorem, no imported ansatz, and no renamed known result appear. Score 1 reflects the stipulative character of the definitional premise, not a circular derivation.
Assumptions & free parameters
assumptions (3)
- domain assumption Authentic creativity requires subjective experience, intentionality, and embodiment.
- domain assumption Battle rap is a representative test bed for real-time interactive creativity such that failure there implies failure generally.
- domain assumption Current token prediction cannot incorporate planning, rhythmic alignment, or turn-taking by architectural construction, not merely by training configuration.
Cite this review
Pith. "Pith review of Next Token Prediction Is a Dead End for Creativity." pith.science (2026). https://pith.science/paper/ODSNDKQR
@misc{pith2026250519277,
author = {Pith},
title = {Pith review of: Next Token Prediction Is a Dead End for Creativity},
year = {2026},
howpublished = {\url{https://pith.science/paper/ODSNDKQR}},
note = {Machine review of arXiv:2505.19277}
}
read the original abstract
This paper argues that token prediction is fundamentally misaligned with real creativity. While next-token models have enabled impressive advances in language generation, their architecture favours surface-level coherence over spontaneity, originality, and improvisational risk. We use battle rap as a case study to expose the limitations of predictive systems, demonstrating that they cannot truly engage in adversarial or emotionally resonant exchanges. By reframing creativity as an interactive process rather than a predictive output, we offer a vision for AI systems that are more expressive, responsive, and aligned with human creative practice.
Reference graph
Works this paper leans on
-
[1]
Next Token Prediction Is a Dead End for Creativity: Why It’s Impossible to Lose Yourself in the Moment Ìbùkún Ọlátúnjí¹ Mark Sheppard²* ¹Computational Foundry, Swansea University, Crymlyn Burrows, Skewen, Swansea SA10 6JW, UK ²University of Kent, Canterbury, Kent CT2 7NZ, U Abstract This position paper argues that token prediction is fundamentally misalig...
work page 2020
-
[2]
As model architectures and training techniques advance we might expect even richer real-time interactions. For example, an AI MC might build on its opponent’s last verse with wit and coherence in real time. 2 However, from an interactional standpoint, systems can be designed to emulate the external signs of losing oneself through delay, stylistic deviatio...
work page 2024
-
[6]
Next-token predictors are adaptive and evolving in their creative capabilities. Techniques like prompt iteration and ensemble querying can elicit a diversity of ideas analogous to a human brainstorming session. A recent study showed that when an LLM is queried multiple times for a task, its “collective creativity” can rival that of a group of 8–10 humans ...
work page 2024
-
[7]
With the correct interaction strategies, next-token models exhibit a breadth of ideas and inventiveness comparable to, and often surpassing, human collaborators. Empirical research indicates that LLMs not only produce creative artifacts, but may do so using strategies reminiscent of human creativity. For instance, Nath et al. (Nath, Dayan, and Stevenson 2...
work page 2024
-
[10]
[emphasis added]. Immersion in Human–AI Systems Losing oneself in the creative act suggests immersion in a psychological state of flow, marked by spontaneity, deep absorption, and a blurring of self–other boundaries (Csikszentmihalyi 1997). Such states often emerge from perceived lack of control (Mulatti and Treccani
work page 1997
-
[11]
or non goal-directed cognition such as mind-wandering (Gable, Hopper, and Schooler 2019). Crucially, such states depend on agency (the capacity for self-directed action), embodiment (a physically situated perspective), and subjective temporality (a lived experience of time). Current AI systems lack these prerequisites. Without intrinsic motivation, self-a...
work page 2019
-
[12]
In Joint Proceedings of the ACM IUI 2021 Workshops
Generative models can help writers without writing for them. In Joint Proceedings of the ACM IUI 2021 Workshops . Bell, A., and Valley, T
work page 2021
-
[14]
arXiv preprint arXiv:2005.14165
Language models are few-shot learners. arXiv preprint arXiv:2005.14165. Caffeine
arXiv 2005
Show all 18 references
-
[15]
A call for clarity in beam search: How it works and when it stops. In Calzolari, N.; Kan, M.-Y.; Hoste, V.; Lenci, A.; Sakti, S.; and Xue, N., eds., Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation ( LREC-CO...
2024
-
[16]
arXiv preprint arXiv:2405.00899v2
Characterising the creative process in humans and large language models. arXiv preprint arXiv:2405.00899v2. Newkeen, R
-
[2009]
we define flow as ’all of the rhythmical and articulative features of a rapper’s delivery of the lyrics.’ In contrast to this, most LLMs optimise for semantic coherence and are unable to model rhythmic constraints or beat alignment (Ọlátúnjí et al. 2025). To capture cadence, w...
2025
-
[2014]
arXiv preprint arXiv:1410.6142
The Lovelace 2.0 Test of Artificial Creativity and Intelligence. arXiv preprint arXiv:1410.6142. Riedl, M
-
[2017]
arXiv preprint arXiv:1706.03762
Attention is all you need. arXiv preprint arXiv:1706.03762. Vernon, D.; Hocking, I.; and Farahar, C
-
[2020]
Negotiation Journal 36(1):57–72
The art of negotiation exercise design: Five basic principles to produce powerful learning experiences. Negotiation Journal 36(1):57–72. Bellemare-Pepin, A.; Lespinasse, F.; Thölke, P.; Harel, Y.; Mathewson, K.; Olson, J. A.; Bengio, Y.; and Jerbi, K. 2024.Divergent creativity...
2024
-
[2021]
with at high quality (Franceschelli and Musolesi 2024b), challenging the notion that they merely regurgitate training data. In creative writing tasks, for example, these models can generate original work that surprises human users; in many cases human-users are not able to dis...
2024
-
[2023]
co-creative spatiality,
describe “co-creative spatiality,” where human and non-human agencies, including AI, intertwine within spatial, material, and symbolic contexts. Similarly, work in computational creativity argues for recognising AI systems as collaborators within hybrid creative ecosystems (Co...
2012
-
[2024]
2023), CBS approximates a simple yet effective generate-and-test loop that may be adapted for interactive domains like freestyle rap
with LLM-as-a-Judge self-evaluation (Zheng et al. 2023), CBS approximates a simple yet effective generate-and-test loop that may be adapted for interactive domains like freestyle rap. These strategies represent a shift from purely predictive models toward architectures capable...
2023
-
[2025]
Nijstad and Baas 2010)
found that LLMs tackle creative tasks (e.g., inventing new uses for everyday objects in the Alternative Uses Test) in similar ways to humans, employing flexible and persistent approaches (Bernard A. Nijstad and Baas 2010). Such parallels hint that next-token models can emulate...
2010
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.