{"id":"098c1bf5-f69c-4351-be00-23fbb43f48dc","arxiv_id":"2507.23365","paper_version":1,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Prompt-based AI music platforms produce polished, professional-sounding output, but the author argues they cannot reproduce unpolished novice performance, a gap he explores with two albums and an LLM-mediated self-interview.","lead":"A researcher describes making two albums with AI music generators, one from spam emails and one built around amateur violin playing. He uses a chatbot to interview himself about authorship, fraudulence, and the things AI music platforms cannot do.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The universal platform-limitation claim rests on undocumented, non-systematic prompt attempts; without prompt texts, versions, and trials, the claimed inability is unfalsifiable.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing point: the claim that Udio cannot generate unpolished music depends on the prompts used being fair tests, and those prompts are not documented. My stress-test concurs. The paper's central assertion about platform limitations is empirical and general, yet the supporting evidence is a single, anecdotal, non-reproducible set of attempts. The manuscript itself acknowledges only 'trial and error' without reporting the search strategy. Therefore the claim is unfalsifiable as presented; it could be a prompt-engineering failure rather than a fundamental platform limitation. This does not change the reader's verdict of UNVERDICTED, since the reader already flagged the same insufficiency. The proposed concrete test—a systematic prompt-space search with documented prompts and blinded evaluation—would settle whether the concern actually lands. I do not see grounds for REJECT because the paper is primarily a reflective, practice-based contribution, not a rigorous empirical study; its value as autoethnography is independent of the generalizing claim. Nor is ACCEPT appropriate without the missing evidence. Thus UNCHANGED is the correct verdict recommendation.","tokens_in":13347,"tokens_out":3275,"duration_ms":32704,"concrete_test":"Obtain from the author the exact text of every prompt and all platform/parameter settings used in A Difficult Christmas (the prompts are reportedly embedded as speech in track 1). Then run a systematic search: for each prompt, generate N=50 outputs on current Udio and Suno, plus variations such as 'novice violin student practicing', 'bad amateur violinist playing out of tune', and audio-input conditioning with a real beginner recording. Have blinded raters judge whether any output convincingly simulates unpracticed/unpolished performance. If any condition yields convincing outputs, the universal inability claim is refuted; if none do across diverse prompts, the claim gains support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—that state-of-the-art prompt-based AI music generation platforms cannot produce unpracticed, unpolished, unproduced music—is stated as a general limitation (Abstract; §2.3: 'it seems that to simulate an unpolished and anxious music performance by an unskilled performer is, ironically, beyond today's most powerful prompt-based AI music generation technology'). The only evidence is the author's own trial-and-error with Udio on A Difficult Christmas, described as 'an attempt to make Udio generate something that sounds like a novice violin student practicing' (§1). The manuscript does not report the prompt texts, platform versions, model parameters, number of attempts, or criteria for 'convincing'. Because prompt-based generators are highly sensitive to prompt wording, conditioning, and stochastic sampling, and because the author did not perform a systematic search over prompt space, the observed failure could be a prompt-engineering artifact rather than a platform capability ceiling. The claim is further framed as about 'today's most powerful' platforms, but no comparative or temporal evidence is provided. Thus the paper's most general, falsifiable assertion is not adequately supported; it is a single, undocumented negative result generalized to a class of systems.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a first-person, practice-based reflection on two albums made with the commercial prompt-based music generation platforms Suno and Udio. The first album, \"Music from the Spam Folder\", turns junk mail into prompts and curates the outputs; the second, \"A Difficult Christmas\", is built around the author's claim that these platforms cannot generate music that is unpracticed, unpolished, and unproduced. The middle of the paper is a transcript of an interview conducted by ChatGPT 4o, seeded with the album liner notes and paper text, on authorship, fraudulence, and new creative spaces; the author then edits his answers and adds a reflection on the method.","tokens_in":13561,"tokens_out":4315,"duration_ms":42354,"significance":"Within the creative-practice-research genre, the paper has real strengths: it is unusually candid about the limits of authorship and about the author's own musical limitations; it connects the work to relevant literature on authorship, plunderphonics, and music-performance modeling; and it is transparent about the LLM's role, including its tendency to flatter and its failure to cite literature. The central empirical claim—that current prompt-based AI music generators cannot produce unpolished, unpracticed, unproduced music—would be a significant observation about the training-data bias of commercial music-generation systems if it were established. As it stands, however, it is a single undocumented negative result generalized to a class of systems, and the paper does not provide the evidence needed to support that generalization.","major_comments":[{"comment":"The paper's load-bearing empirical claim is stated as a general platform limitation in the Abstract and in §2.3 ('beyond today's most powerful prompt-based AI music generation technology'), but the evidence is limited to the author's own trial-and-error with Udio, described in §1 only as \"an attempt to make Udio generate something that sounds like a novice violin student practicing.\" No prompt texts, platform versions, model parameters, number of attempts, or criteria for judging 'convincing' are reported. Because prompt-conditioned generative systems are highly sensitive to wording and stochastic sampling, a small set of undocumented failures cannot establish a capability ceiling; the observed failures could be a prompt-engineering artifact. The claim should either be reframed as a first-person experiential report or be backed by a systematic, reproducible search protocol with released prompt logs.","section":"Abstract; §2.3; §1"},{"comment":"The claim is about \"today's most powerful prompt-based AI music generation technology\" and the paper names both Suno and Udio, but the §2.3 evidence concerns Udio only. No dates, service versions, or comparative tests with Suno or other state-of-the-art systems are given. If the assertion is meant to cover a class of systems, it needs comparative evidence or a narrowed scope; otherwise the reader cannot tell whether the limitation is specific to one platform, one version, or one prompting style.","section":"Abstract; §2.3"},{"comment":"The LLM-mediated interview is presented as the paper's method, but the protocol is underspecified: §1 says the author \"explored different LLM model versions, prompts and trajectories of thought\" and settled on the final one, yet none of those alternatives or the selection criteria are described, and §3 acknowledges the LLM \"was not able to cite any literature\" and \"bordered on flattering.\" These admissions are honest, but they undermine the final section's methodological claims unless the protocol and its limitations are documented, for example in an appendix, or the conclusions about the method are tempered.","section":"§1; §3"}],"minor_comments":[{"comment":"There is a missing space in \"Music from the Spam Foldermarvels\" in the second paragraph of §1.","section":"§1"},{"comment":"In the list of absent descriptors, \"uskilled\" should be \"unskilled\".","section":"§2.3"},{"comment":"The in-text citation \"Agarwal and Greer (2023)\" does not match the reference entry \"Agarwal, M. and Geer, R.\"; please align the spelling.","section":"References"},{"comment":"The claim that studies of non-expert musicians \"number far less\" than studies of experts would benefit from a systematic count or at least a clearer scope, since the sentence currently reads as an impression rather than a finding.","section":"§2.3"}],"recommendation":"major_revision","confidential_remarks":"This is a genuine creative-practice paper with a transparent method and an interesting central question, but for a scientific venue the main claim is under-evidenced rather than wrong. I would encourage the editor to treat this as a major revision with the option of reframing the paper as autoethnographic; if the author instead wants to keep the general platform-limitation claim, the revision must include prompt logs, versions, trials, and a protocol. The footnote \"This paper was not accepted to AIMC 2025\" is unusual and should probably be removed or integrated into the narrative."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bob—you should read this one. It's a rare thing: a practice-based paper that is genuinely reflective and not padded with methodology theater. The two albums are real artifacts, and the idea of 'pseudoplunderphonics'—sampling the outputs of models trained on already-plundered music—is a fresh and useful term. The LLM-mediated self-interview is an honest experiment, and the author openly says the LLM was flattering, couldn't cite literature, and at times felt like therapy. That kind of transparency is worth crediting. The citation list is thoughtful, including the Suno/Udio lawsuits, Oswald, Collins, and a handful of works on modeling novice performance.\n\nThe soft spot is the one the stress-test names, and it's real. The central claim—that state-of-the-art prompt-based AI music platforms cannot generate unpracticed, unpolished, unproduced music—is stated as a general limitation (Abstract, §2.3), but the evidence is an undocumented set of attempts with Udio. No prompt texts, no platform versions, no parameter settings, no number of trials, no criteria for 'convincing.' So the failure could be a prompt-engineering artifact rather than a platform ceiling. The paper does hedge with 'it seems' and 'probably', but the generalization still outruns the evidence.\n\nThat said, this is an autoethnographic paper, not a systematic evaluation. If the claim were softened to 'in my experience' or accompanied by the actual prompts and a short description of the search over prompt space, the paper would be fine on its own terms. The circularity concern about the LLM being seeded with the author's framing is minor: the author acknowledges it, and the central empirical claim doesn't depend on that interview.\n\nWho's this for? People working on human-AI co-creativity, music information retrieval, and creative practice. It won't change the field, but it's a valuable first-person data point. I'd send it to a serious referee, not desk reject it, with the advice to ask for prompt transparency or a reframed claim. I'd also bring it to a reading group: there's enough here to argue about.","headline":"A candid practice-based paper whose central generalization about AI music platforms outruns its undocumented evidence.","tokens_in":14050,"tokens_out":3060,"would_cite":true,"duration_ms":31685,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that state-of-the-art prompt-based AI music platforms, trained on commercially released recordings, cannot generate music that is not practiced, polished, and produced, and that human imperfection is a musical space they…","keywords":["AI music generation","authorship","fraudulence","prompt engineering","unpolished music","Suno","Udio","LLM-mediated self-reflection"],"falsifier":"Generate a large set of outputs from Udio or Suno using systematically varied prompts that describe a novice, unpracticed, anxious performer (ideally including audio conditioning with an actual beginner's practice session), and ask experienced listeners to identify convincing unpolished performances; if a substantial share of outputs are mistaken for real beginner recordings, the claimed expressive ceiling is not a platform limitation.","tokens_in":13139,"feed_emoji":"🎻","tokens_out":6366,"duration_ms":63946,"temperature":0.7,"pith_summary":"This paper is an artist-researcher's account of two albums made with Suno and Udio, prompt-based AI music generators. The first album, Music from the Spam Folder, treats junk mail as lyrics and celebrates what the machines can do; the second, A Difficult Christmas, is built around what the author claims they cannot do. The paper's central claim is that current state-of-the-art prompt-based AI music platforms cannot generate music that is not practiced, polished, and produced: the novice-violin practicer he prompted Udio to produce came out sounding competent. That claim matters because it locates a genuine expressive limit in the technology, and it gives human imperfection a positive role as a space AI has not yet colonized. The paper also develops a method, an LLM conducting an interview of the author, as a way of reflecting on authorship and fraudulence.","feed_headline":"AI music bots can't fake a beginner practicing","feed_subtitle":"Two albums built with Suno and Udio expose the polish ceiling—and why human imperfection still matters.","key_machinery":"The load-bearing object is A Difficult Christmas's first track: a prompt-based attempt to simulate novice violin practice, contrasted with the author's own unpolished violin playing in the following tracks. The mechanism invoked to explain the ceiling is the training-data distribution: commercial platforms are fitted to released, professionally produced music, so unpracticed, anxious, unproduced performance sits outside their typical output. A secondary machinery is the author's term \"pseudoplunderphonics\"—plundering at one remove through a machine-learning pipeline trained on plundered recordings—which frames his reuse of generated audio as a distinctive authorial act.","core_discovery":"On the author's own terms, the discovery is an expressive ceiling: despite prompt flexibility, Udio and Suno are trained on commercially released, professionally recorded music, so their outputs reliably land in a practiced, polished, produced register. The first track of A Difficult Christmas is offered as evidence—each prompt was an attempt to make Udio generate something sounding like a novice violin student practicing, and the generated audio did not sound unskilled. The ten tracks that follow are the author playing solo violin unpolished, presented as sounds not possible with Udio. Working from this limitation, the author reframes authorship: prompt authoring alone is weak authorship, curation is editing, and remixing, sampling, or live performance can justify a claim to authorship; the album also invites machines to listen to non-expert music so future training can include it.","pith_inferences":["If the limitation is real, it suggests a testable asymmetry: commercial generators may model \"good\" performance far better than \"bad\" performance, and the same could hold for other aesthetic negatives such as out-of-tune, clumsy, or amateurish output, which could be probed with controlled perceptual studies.","The author's LLM-mediated interview method could be extended to other creative practitioners as a structured self-reflection protocol, with the caveat that the LLM's flattering, uncited responses shape what gets articulated.","An unstated corollary is that the unpolished register is itself a scarce and increasingly valuable resource in the AI era; as synthetic music floods distribution channels, deliberately imperfect human performance may gain cultural and economic value.","The paper implicitly predicts that if such albums enter training data, platforms will begin to imitate non-expert performance, which would erase the current boundary; one could monitor releases for tell-tale amateurish artifacts."],"forward_implications":["If the polished-produced ceiling is real, prompting alone cannot reach the expressive register of amateur, beginner, or deliberately unpolished music, so artists seeking that register must play, record, or process audio themselves.","Human imperfection becomes a distinguishing resource: unpolished performance is an authenticity signal that AI-generated music currently lacks.","Platforms trained on future datasets that include unpolished music, such as A Difficult Christmas, could lose this limitation, so the ceiling is a contingent property of training data rather than a permanent law.","Authorship claims based only on writing prompts are weak under this account; curation, remixing, sampling, and live performance carry the authorial weight."],"supporting_citations":[{"why":"The album under analysis; its liner notes document the novice-violin prompts and the invitation for machines to listen to non-expert music.","marker":"(Sturm, 2024a)"},{"why":"The companion album; its liner notes and process description define the workflow of prompting, curating, and remixing that the paper interrogates.","marker":"(Sturm, 2024b)"},{"why":"Supplies the \"musicalization of everyday life\" framing and a scholarly account of authenticity in AI music.","marker":"(Tan, 2024)"},{"why":"Alleges Suno's training data include copyrighted commercial recordings, the basis for the paper's explanation of the polished-produced output register.","marker":"(US District Court for the District of Massachusetts, 2024a)"},{"why":"Alleges Udio's training data include copyrighted commercial recordings, grounding the same explanation for Udio.","marker":"(US District Court for the District of Massachusetts, 2024b)"},{"why":"Provides the call to compose music that subverts content-retrieval and classification systems, which the album's machine-listening invitation answers.","marker":"(Collins, 2007)"},{"why":"Quoted epigraph about Ives's father, used to frame imperfection as music worth hearing.","marker":"(Hentoff, 1974)"}],"fun_headline_variants":["AI music bots can't fake a beginner practicing","Prompt-based AI music has a polish problem","Suno and Udio can't sound unskilled, so I played","Curation over prompting: AI music authorship","AI music's polished output challenges authenticity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the assumption that the prompts and settings used in A Difficult Christmas are a fair test of what Udio can do; the paper does not report the exact prompt text, platform version, parameters, or number of trials.","fun_headline_variants_meta":{"raw":{"variants":["AI music bots can't fake a beginner practicing","Prompt-based AI music has a polish problem","Suno and Udio can't sound unskilled, so I played","Curation over prompting: AI music authorship","AI music's polished output challenges authenticity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000332,"raw_usage":{"total_tokens":1819,"prompt_tokens":889,"completion_tokens":930,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":857}},"tokens_in":505,"tokens_out":930,"duration_ms":10977,"temperature":1.0,"reasoning_tokens":857,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T10:48:02.385137+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate a large set of outputs from Udio or Suno using systematically varied prompts that describe a novice, unpracticed, anxious performer (ideally including audio conditioning with an actual beginner's practice session), and ask experienced listeners to identify convincing unpolished performances; if a substantial share of outputs are mistaken for real beginner recordings, the claimed expressive ceiling is not a platform limitation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the call to compose music that subverts content-retrieval and classification systems, which the album's machine-listening invitation answers."},{"cited_title":"(October 20 1974)","cited_arxiv_id":null,"evidence_quote":"Quoted epigraph about Ives's father, used to frame imperfection as music worth hearing."}],"review_version":1}