REVIEW 3 major objections 4 minor 1 cited by
"I made this (sort of)": Negotiating authorship, confronting fraudulence, and exploring new musical spaces with prompt-based AI music generation
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper claims that state-of-the-art prompt-based AI music platforms, trained on commercially released recordings, cannot generate music that is not practiced, polished, and produced, and that human imperfection is a musical space they…
desk verdict A candid practice-based paper whose central generalization about AI music platforms outruns its undocumented evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is A Difficult Christmas's first track: a prompt-based attempt to simulate novice violin practice, contrasted with the author's own unpolished violin playing in the following tracks. The mechanism invoked to explain the ceiling is the training-data distribution: commercial platforms are fitted to released, professionally produced music, so unpracticed, anxious, unproduced performance sits outside their typical output. A secondary machinery is the author's term "pseudoplunderphonics"—plundering at one remove through a machine-learning pipeline trained on plundered recordings—which frames his reuse of generated audio as a distinctive authorial act.
What would settle it
Generate a large set of outputs from Udio or Suno using systematically varied prompts that describe a novice, unpracticed, anxious performer (ideally including audio conditioning with an actual beginner's practice session), and ask experienced listeners to identify convincing unpolished performances; if a substantial share of outputs are mistaken for real beginner recordings, the claimed expressive ceiling is not a platform limitation.
Extended reading notes
Core claim
On the author's own terms, the discovery is an expressive ceiling: despite prompt flexibility, Udio and Suno are trained on commercially released, professionally recorded music, so their outputs reliably land in a practiced, polished, produced register. The first track of A Difficult Christmas is offered as evidence—each prompt was an attempt to make Udio generate something sounding like a novice violin student practicing, and the generated audio did not sound unskilled. The ten tracks that follow are the author playing solo violin unpolished, presented as sounds not possible with Udio. Working from this limitation, the author reframes authorship: prompt authoring alone is weak authorship, curation is editing, and remixing, sampling, or live performance can justify a claim to authorship; the album also invites machines to listen to non-expert music so future training can include it.
Load-bearing premise
The central claim rests on the assumption that the prompts and settings used in A Difficult Christmas are a fair test of what Udio can do; the paper does not report the exact prompt text, platform version, parameters, or number of trials.
Editorial extensions
If this is right
- If the polished-produced ceiling is real, prompting alone cannot reach the expressive register of amateur, beginner, or deliberately unpolished music, so artists seeking that register must play, record, or process audio themselves.
- Human imperfection becomes a distinguishing resource: unpolished performance is an authenticity signal that AI-generated music currently lacks.
- Platforms trained on future datasets that include unpolished music, such as A Difficult Christmas, could lose this limitation, so the ceiling is a contingent property of training data rather than a permanent law.
- Authorship claims based only on writing prompts are weak under this account; curation, remixing, sampling, and live performance carry the authorial weight.
Reading between the lines
- If the limitation is real, it suggests a testable asymmetry: commercial generators may model "good" performance far better than "bad" performance, and the same could hold for other aesthetic negatives such as out-of-tune, clumsy, or amateurish output, which could be probed with controlled perceptual studies.
- The author's LLM-mediated interview method could be extended to other creative practitioners as a structured self-reflection protocol, with the caveat that the LLM's flattering, uncited responses shape what gets articulated.
- An unstated corollary is that the unpolished register is itself a scarce and increasingly valuable resource in the AI era; as synthetic music floods distribution channels, deliberately imperfect human performance may gain cultural and economic value.
- The paper implicitly predicts that if such albums enter training data, platforms will begin to imitate non-expert performance, which would erase the current boundary; one could monitor releases for tell-tale amateurish artifacts.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a first-person, practice-based reflection on two albums made with the commercial prompt-based music generation platforms Suno and Udio. The first album, "Music from the Spam Folder", turns junk mail into prompts and curates the outputs; the second, "A Difficult Christmas", is built around the author's claim that these platforms cannot generate music that is unpracticed, unpolished, and unproduced. The middle of the paper is a transcript of an interview conducted by ChatGPT 4o, seeded with the album liner notes and paper text, on authorship, fraudulence, and new creative spaces; the author then edits his answers and adds a reflection on the method.
Significance. Within the creative-practice-research genre, the paper has real strengths: it is unusually candid about the limits of authorship and about the author's own musical limitations; it connects the work to relevant literature on authorship, plunderphonics, and music-performance modeling; and it is transparent about the LLM's role, including its tendency to flatter and its failure to cite literature. The central empirical claim—that current prompt-based AI music generators cannot produce unpolished, unpracticed, unproduced music—would be a significant observation about the training-data bias of commercial music-generation systems if it were established. As it stands, however, it is a single undocumented negative result generalized to a class of systems, and the paper does not provide the evidence needed to support that generalization.
major comments (3)
- [Abstract; §2.3; §1] The paper's load-bearing empirical claim is stated as a general platform limitation in the Abstract and in §2.3 ('beyond today's most powerful prompt-based AI music generation technology'), but the evidence is limited to the author's own trial-and-error with Udio, described in §1 only as "an attempt to make Udio generate something that sounds like a novice violin student practicing." No prompt texts, platform versions, model parameters, number of attempts, or criteria for judging 'convincing' are reported. Because prompt-conditioned generative systems are highly sensitive to wording and stochastic sampling, a small set of undocumented failures cannot establish a capability ceiling; the observed failures could be a prompt-engineering artifact. The claim should either be reframed as a first-person experiential report or be backed by a systematic, reproducible search protocol with released prompt logs.
- [Abstract; §2.3] The claim is about "today's most powerful prompt-based AI music generation technology" and the paper names both Suno and Udio, but the §2.3 evidence concerns Udio only. No dates, service versions, or comparative tests with Suno or other state-of-the-art systems are given. If the assertion is meant to cover a class of systems, it needs comparative evidence or a narrowed scope; otherwise the reader cannot tell whether the limitation is specific to one platform, one version, or one prompting style.
- [§1; §3] The LLM-mediated interview is presented as the paper's method, but the protocol is underspecified: §1 says the author "explored different LLM model versions, prompts and trajectories of thought" and settled on the final one, yet none of those alternatives or the selection criteria are described, and §3 acknowledges the LLM "was not able to cite any literature" and "bordered on flattering." These admissions are honest, but they undermine the final section's methodological claims unless the protocol and its limitations are documented, for example in an appendix, or the conclusions about the method are tempered.
minor comments (4)
- [§1] There is a missing space in "Music from the Spam Foldermarvels" in the second paragraph of §1.
- [§2.3] In the list of absent descriptors, "uskilled" should be "unskilled".
- [References] The in-text citation "Agarwal and Greer (2023)" does not match the reference entry "Agarwal, M. and Geer, R."; please align the spelling.
- [§2.3] The claim that studies of non-expert musicians "number far less" than studies of experts would benefit from a systematic count or at least a clearer scope, since the sentence currently reads as an impression rather than a finding.
Circularity Check
No significant circularity: the platform-limitation claim is an empirical, trial-and-error observation, not a derivation from its own inputs or from load-bearing self-citations.
full rationale
This is a reflective practice-based paper, not a derivation chain with equations or fitted parameters. The central assertion, that state-of-the-art prompt-based AI music generation platforms cannot produce music that is unpracticed, unpolished, and unproduced, is presented as an empirical observation from the author's own trial-and-error work with Udio and Suno (Sections 1 and 2.3). The claim is supported by the creation of A Difficult Christmas, whose first track was an attempt to prompt Udio to generate a novice violin student practicing; the paper reports that this attempt failed. Nothing in the paper defines the target property into existence or fits a parameter and then calls the fit a prediction. The self-citations to Sturm (2024a, 2024b) document the albums and their liner notes, but they do not carry the logical weight of the inability claim; the claim stands or falls on the reader's willingness to accept the author's undocumented negative prompting results. The LLM interview is seeded with the author's own framing, abstract, and introduction, but the paper transparently identifies this as a method for self-reflection and does not use the interview as evidence for the platform-limitation claim. The lack of prompt texts, platform versions, and number of attempts is a genuine evidential weakness relevant to correctness and falsifiability, but it is not circularity. No self-citation chain, imported uniqueness theorem, or ansatz-via-citation is used to force a conclusion. The paper therefore contains no circular step by the standards of this review.
Assumptions & free parameters
assumptions (3)
- domain assumption Suno and Udio are representative of state-of-the-art prompt-based AI music generation platforms.
- ad hoc to paper The specific prompts used in A Difficult Christmas were reasonable attempts to generate novice violin music.
- domain assumption The LLM interview responses are treated as a valid basis for self-reflection.
Cite this review
Pith. "Pith review of "I made this (sort of)": Negotiating authorship, confronting fraudulence, and exploring new musical spaces with prompt-based AI music generation." pith.science (2026). https://pith.science/paper/LDFDY2UA
@misc{pith2026250723365,
author = {Pith},
title = {Pith review of: "I made this (sort of)": Negotiating authorship, confronting fraudulence, and exploring new musical spaces with prompt-based AI music generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/LDFDY2UA}},
note = {Machine review of arXiv:2507.23365}
}
read the original abstract
I reflect on my experience creating two music albums centered on state-of-the-art prompt-based AI music generation platforms. The first album explicitly poses the question: What happens when I collide my junk mail with these platforms? The second album is a direct response to the first, and toys with the inability of state-of-the-art prompt-based AI music generation platforms to generate music that is not ``practiced'', ``polished'', and ``produced''. I seed a large language model (LLM) with information about these albums and have it interview me, which results in the exploration of several deeper questions: To what extent am I the author? Where am I in the resulting music? How is my musical identity changing as I am faced with machines that are in some ways far more talented than I? What new musical spaces does my work open, for me or anyone/thing else? I conclude by reflecting on my reflections, as well as LLM-mediated self-reflection as method.
Figures
Forward citations
Cited by 1 Pith paper
-
Finding the noise: Zero-shot AI Music Detection
A zero-shot method based on fakeprints, NMF and a blur-based reconstruction error detects unknown AI-music generators in one-class and clustering setups, working for most services but missing Mubert and pre-v9 Mureka.
Reference graph
Works this paper leans on
-
[1]
Agarwal, M. and Geer, R. (2023). Spectrogram-based deep learning for flute audition assessment and intelligent feedback. In Proc. IEEE Int Symp. Multimedia
work page 2023
-
[2]
Anscomb, C. (2025). Who Authors AI Art? (And why does it matter?) . AI & Society: Knowledge, Culture and Communication
work page 2025
-
[3]
Baudrillard, J. (1994). Simulacra and simulation , volume S. F. Glaser (trans.). The University of Michigan Press
work page 1994
-
[4]
Bozkurt, B., Baysal, O., and Yuret, D. (2017). A dataset and baseline system for singing voice assessment. In Proc. Int. Symp. Computer Music Multidisciplinary Research
work page 2017
-
[5]
Bretan, M. and Weinberg, G. (2016). A survey of robotic musicianship. Communications of the ACM , 59(5):100--109
work page 2016
-
[6]
Caldwell, M. (2023). What Is an ``Author''?---Copyright Authorship of AI Art Through a Philosophical Lens . Houston Law Review , 61(2)
work page 2023
-
[7]
E., Grachten, M., Goebl, W., and Widmer, G
Cancino-Chac\' o n, C. E., Grachten, M., Goebl, W., and Widmer, G. (2018). Computational models of expressive music performance: A comprehensive and critical review. Front. Digit. Humanities , 5
work page 2018
-
[8]
Caporusso, N. (2023). Generative artificial intelligence and the emergence of creative displacement anxiety. Research in Psychology and Behavior , 3(1)
work page 2023
Show all 61 references
-
[9]
Collins, N. (2007). Composing to subvert content retrieval engines. ARRAY ICMA Online Journal
2007
-
[10]
Dang, H., Mecke, L., Lehmann, F., Goller, S., and Buschek, D. (2022). How to Prompt? Opportunities and Challenges of Zero- and Few-Shot Learning for Human-AI Interaction in Creative Applications of Generative Models . In Proc. Generative AI and HCI Workshop at CHI
2022
-
[11]
Friberg, A., Bresin, R., and Sundberg, J. (2006). Overview of the KTH rule system for musical performance . Advances in Cognitive Psychology , 2(2-3):145--161
2006
-
[12]
Gabrielsson, A. (2003). Music performance research at the millennium. Psychology of Music , 31(3):221--272
2003
-
[13]
Giraldo, S., Waddell, G., Nou, I., Ortega, A., Mayor, O., Perez, A., Williamon, A., and Ramirez, R. (2019). Automatic assessment of tone quality in violin music performance. Frontiers in Psychology , 10
2019
-
[14]
and Lee, K
Han, Y. and Lee, K. (2014). Hierarchical approach to detect common mistakes of beginner flute players. In Proc. Int. Symp. Music Info. Retrieval
2014
-
[15]
(October 20 1974)
Hentoff, N. (October 20 1974). With his ears on right Charles Ives Remembered . In The New York Times , page 3
1974
-
[16]
Hiraga, R., Bresin, R., Hirata, K., and Katayose, H. (2004). Rencon 2004: Turing test for musical expression. In Proc. Int. Conf. New Interfaces Musical Expression
2004
-
[17]
Holm-Hudson, K. (1997). Quotation and Context: Sampling and John Oswald's Plunderphonics . Leonardo Music J. , 7:17--25
1997
-
[18]
Huang, J., Hung, Y.-N., Pati, A., Gururani, S., and Lerch, A. (2020). Score-informed networks for music performance assessment. In Proc. Int. Symp. Music Info. Retrieval
2020
-
[19]
Juslin, P. N. and Sloboda, J., editors (2010). Handbook of Music and Emotion: Theory, Research, Applications . Oxford University Press
2010
-
[20]
L., and Larsen, J
Kereliuk, C., Sturm, B. L., and Larsen, J. (2015). Deep learning and music adversaries. IEEE Trans. Multimedia , 17(11):2059--2071
2015
-
[21]
A., Lee, S.-W., Hong, H., Yang, C., and Kim, Y.-H
Kim, T., Bae, S., Kim, H. A., Lee, S.-W., Hong, H., Yang, C., and Kim, Y.-H. (2024). Mindfuldiary: Harnessing large language model to support psychiatric patients' journaling. In Proc. CHI
2024
-
[22]
J., Rafferty, A
Kumar, H., Xiao, R., Lawson, B., Musabirov, I., Shi, J., Wang, X., Luo, H., Williams, J. J., Rafferty, A. N., Stamper, J., and Liut, M. (2024). Supporting self-reflection at scale with large language models: Insights from randomized field experiments in classrooms. In Proc. AC...
2024
-
[23]
and Hayhoe, B
Lewis, M. and Hayhoe, B. (2024). The digital Balint: using AI in reflective practice . Education for Primary Care , 35(6):198--202
2024
-
[24]
and Liu, M
Li, X. and Liu, M. (2024). Rethinking ChatGPT's Success: Usability and Cognitive Behaviors Enabled by Auto-regressive LLMs' Prompting . arXiv , (2405.10474)
2024 arXiv
-
[25]
Lubar, S. (2014). Curator as auteur. The Public Historian , 36(1):71--76
2014
-
[26]
Luo, Y.-J., Su, L., Yang, Y.-H., and Chi, T.-S. (2015). Detection of common mistakes in novice violin playing. In Proc. Int. Symp. Music Info. Retrieval
2015
-
[27]
McCormack, J., Gifford, T., and Hutchings, P. (2019). Autonomy, authenticity, authorship and intention in computer generated art. In Proc. EvoMUSART 2019: Computational Intelligence in Music, Sound, Art and Design
2019
-
[28]
L., and Dixon, S
Mishra, S., Benetos, E., Sturm, B. L., and Dixon, S. (2020). Reliable local explanations for machine listening. In Proc. Int. Joint Conf. Neural Networks
2020
-
[29]
L., and Dixon, S
Mishra, S., Stoller, D., Benetos , E., Sturm , B. L., and Dixon, S. (2019). Gan-based generation and automatic selection of explanations for neural networks. In Proc. SafeML workshop at ICLR 2019
2019
-
[30]
L., and Dixon, S
Mishra, S., Sturm, B. L., and Dixon, S. (2017). Local interpretable model-agnostic explanations for music content analysis. In Proc. Int. Symp. Music Info. Retrieval
2017
-
[31]
L., and Dixon, S
Mishra, S., Sturm, B. L., and Dixon, S. (2018a). Understanding a deep machine listening model through feature inversion. In Proc. Int. Symp. Music Info. Retrieval
2018
-
[32]
L., and Dixon, S
Mishra, S., Sturm, B. L., and Dixon, S. (2018b). `What are You Listening to?' Explaining Predictions of Deep Machine Listening Systems . In Proc. EUSIPCO
2018
-
[33]
Morris, M. R. (2024). Prompting considered harmful. Communications of the ACM , 67(12):28--30
2024
-
[34]
Morsi, A., Zhang, H., Maezawa, A., Dixon, S., and Serra, X. (2024). Simulating piano performance mistakes for music learning. In Proc. Sound and Music Computing Conf
2024
-
[35]
Oppenlaender, J., Linder, R., and Silvennoinen, J. (2024). Prompting AI Art: An Investigation into the Creative Skill of Prompt Engineering . Human-Computer Interaction
2024
-
[36]
Oswald, J. (1985). Plunderphonics, or audio piracy as a compositional prerogative. In Wired Society Electro-Acoustic Conference
1985
-
[37]
Pachet, F. (2003). The continuator: Musical interaction with style. Journal of New Music Research , 32(3):333--341
2003
-
[38]
L., and Maruri-Aguilar, H
Rodr\'iguez-Algarra, F., Sturm, B. L., and Maruri-Aguilar, H. (2016). Analysing scattering-based music content analysis systems: W here's the music? In Proc. Int. Symp. Music Info. Retrieval , pages 344--350
2016
-
[39]
Scherzinger, M. (2014). Musical property: Widening or withering? J. Popular Music Studies
2014
-
[40]
Shroff, L. (2025). ChatGPT Gave Instructions for Murder, Self-Mutilation, and Devil Worship . The Atlantic
2025
-
[41]
R., Schleider, J
Song, I., Park, S., Pendse, S. R., Schleider, J. L., Choudhury, M. D., and Kim, Y.-H. (2025). Exploreself: Fostering user-driven exploration and reflection on personal challenges with adaptive guidance by large language models. arXiv , 2409.09662
2025 arXiv
-
[42]
Sturm, B. L. (2012). Two systems for automatic music genre recognition: What are they really recognizing? In Proc. ACM MIRUM Workshop , pages 69--74
2012
-
[43]
Sturm, B. L. (2013a). Classification accuracy is not enough: On the evaluation of music genre recognition systems. J. Intell. Info. Systems , 41(3):371--406
2013
-
[44]
Sturm, B. L. (2013b). Evaluating music emotion recognition: Lessons from music genre recognition? In Proc. Int. Conf. Multimedia Expo
2013
-
[45]
Sturm, B. L. (2014a). A simple method to determine if a music information retrieval system is a ``horse''. IEEE Trans. Multimedia , 16(6):1636--1644
2014
-
[46]
Sturm, B. L. (2014b). The state of the art ten years after a state of the art: Future research in music information retrieval. J. New Music Research , 43(2):147--172
2014
-
[47]
Sturm, B. L. (2016). The ``horse'' inside: Seeking causes behind the behaviors of music content analysis systems. ACM Computers in Entertainment , 14(2)
2016
-
[48]
Sturm, B. L. (2018a). How stuff works: LSTM model of folk music transcriptions. In Proc. Joint Workshop on Machine Learning for Music, ICML
2018
-
[49]
Sturm, B. L. (2018b). What do these 5,599,881 parameters mean? An analysis of a specific LSTM music transcription model, starting with the 70,281 parameters of its softmax layer . In Proc. Music Metacreation workshop of ICCC
2018
-
[50]
L., Kereliuk, C., and Pikrakis, A
Sturm, B. L., Kereliuk, C., and Pikrakis, A. (2014). A closer look at deep learning neural networks with low-level spectral periodicity features. In Proc. Int. Workshop on Cognitive Info. Process. , pages 1--6
2014
-
[51]
Sturm, B. L. T. (2024a). A Difficult Christmas (liner notes). Technical report, KTH Royal Institute of Technology
2024
-
[52]
Sturm, B. L. T. (2024b). Music From the Spam Folder (liner notes). Technical report, KTH Royal Institute of Technology
2024
-
[53]
Studies of Music Performance (Papers given at a seminar organized by the Music Acoustics Committee of the Royal Swedish Academy of Music) , number 39
Sundberg, J., editor (1983). Studies of Music Performance (Papers given at a seminar organized by the Music Acoustics Committee of the Royal Swedish Academy of Music) , number 39. Royal Swedish Academy of Music
1983
-
[54]
R., S avelka, J., Mee \`u s, S., Godet, M., and Benyekhlef, K
Tan, J., Westermann, H., Pottanigari, N. R., S avelka, J., Mee \`u s, S., Godet, M., and Benyekhlef, K. (2024). Robots in the Middle: Evaluating LLMs in Dispute Resolution . arXiv , 2410.07053
2024 arXiv
-
[55]
Tan, S. (2024). Are We All Musicians Now? Authenticity, Musicianship, and AI Music Generator Suno . SocArXiv
2024
-
[56]
Torres, A. J. R., Alberto, J. M. C., Guieb, A. P. J., and Villarama, J. A. (2025). Language, Identity, and Ethics in AI-Driven Art: Perspectives from Human Artists in Digital Environments . Lang. Tech. and Social Media , 3(1):17--29
2025
-
[57]
Turino, T. (2008). Music as social life: The politics of participation . The University of Chicago Press
2008
-
[58]
UMG Recordings, Inc
US District Court for the District of Massachusetts (2024a). UMG Recordings, Inc. et al. v. Suno, Inc. et al 1:24-cv-11611
2024
-
[59]
UMG Recordings, Inc
US District Court for the District of Massachusetts (2024b). UMG Recordings, Inc. et al. v. Uncharted Labs, Inc. et al 1:24-cv-04777
2024
-
[60]
Yao, H., Lou, J., Qin, Z., and Ren, K. (2024). Promptcare: Prompt copyright protection by watermark injection and verification. In Proc. IEEE Symp. Security and Privacy , pages 845--861
2024
-
[61]
D., Wong, R
Zamfirescu-Pereira, J. D., Wong, R. Y., Hartmann, B., and Yang, Q. (2023). Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts . In Proc. CHI
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.