{"id":"a9cc94c9-1d99-4e0c-b73b-83f44c66d1ed","arxiv_id":"2506.02856","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Preference and perceived emotional efficacy dissociate for AI-generated versus human-composed music, with listeners preferring AI tracks but crediting human tracks with stronger functional emotion elicitation.","lead":"Listeners judged short human-made and AI-generated tracks in calm and upbeat moods, and while they often preferred the AI tracks, they more often said the human tracks were the ones that actually produced the intended emotion. The results suggest that preference alone is not a reliable measure of whether generative music works for emotional regulation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The preference–efficacy dissociation is confounded by track identity: with one AI and one human track per emotion case, authorship and musical features are perfectly collinear, so the central claim about AI vs. human music as classes is not supported.","rationale":"The reader's verdict identifies the single-stimulus confound as the weakest assumption, and I agree that it is the most load-bearing issue. The paper's central contribution is a dissociation between preference and perceived efficacy attributed to musical origin (Section 4.1.3). For this dissociation to generalize to 'AI-generated music' and 'human-composed music' as classes, the two tracks in each emotion case must be representative of their respective classes. Appendix A shows they are not: the AI Calm track is 72 BPM with synthesizer pads; the human Calm track is 53 BPM rubato acoustic piano. The AI Upbeat track is a loop-based electronic piece at 91 BPM; the human Upbeat track is live guitar and percussion at 105 BPM. These differences in tempo, instrumentation, and production style are exactly the kinds of features that plausibly drive both preference and perceived efficacy, independent of authorship. Because the random effect for condition was singular (Section 3.1.1), the analysis could not estimate stimulus-level variance, and the Poisson GLMs on aggregate counts cannot separate the 'origin' effect from the 'track' effect. The result is that the headline claim is underdetermined by the data. I do not think this warrants REJECT: the study is explicitly a pilot, the authors acknowledge limited stimulus variety in Section 5.4, and the Incorrect-labeling condition provides some evidence that explicit origin labels did not drive the effect. However, the claim as stated in the abstract—that human music was more effective 'regardless of labeling' and that preference dissociates from efficacy for AI versus human music—needs to be qualified to 'these particular tracks' or supported with a multi-stimulus replication. The reader's CONDITIONAL verdict captures this appropriately, so I recommend no change.","tokens_in":18988,"tokens_out":5692,"duration_ms":60297,"concrete_test":"Conduct a follow-up with a stimulus set of at least 6–10 tracks per origin per emotion case, matched on tempo, instrumentation, and loudness (or with these features statistically controlled), and fit a mixed-effects model with random intercepts for both participant and track. Test the origin × response-type interaction (preference vs. efficacy). If the dissociation persists after accounting for track-level variance, the central claim is supported; if it attenuates or disappears, the original result is an artifact of the single, acoustically distinct tracks used.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that listeners dissociate preference from perceived efficacy as a function of musical origin—rests on comparing exactly one AI-generated and one human-composed track per emotion case. Appendix A documents large acoustic differences between the paired tracks: 72 vs 53 BPM in Calm (sustained synth pads vs rubato acoustic piano) and 91 vs 105 BPM in Upbeat (loop-based electronic vs live guitar and percussion). Origin is therefore perfectly confounded with track identity: any observed difference in preference or efficacy between 'AI' and 'human' could be caused by tempo, instrumentation, production style, or other musical features rather than authorship. The authors attempted a GLMM with a random effect for condition but report a singular fit (Section 3.1.1) and fell back to Poisson GLMs on aggregate counts, which cannot separate stimulus-level variation from origin. Because there is only one stimulus per cell, no statistical model can disentangle the two. Consequently, the dissociation reported in Section 4.1.3 (human music more effective than preferred; AI music more preferred than effective) is a statement about these four specific tracks, not about AI-generated versus human-composed music as classes. The abstract's claim that human music was 'significantly more likely' to be rated effective overall is also driven by the Upbeat condition; in Calm, efficacy did not differ between origins (Section 4.1.2). This does not invalidate the study as a pilot, but it means the headline conclusion—that preference alone is not a valid proxy for functional efficacy in generative music—is not yet established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a mixed-methods user study (N = 152) comparing one AI-generated (SunoAI) and one human-composed (Myndstream) one-minute instrumental track per emotion case (Calm, Upbeat), under Correctly labeled, Incorrectly labeled, and Unlabeled conditions. Participants rated each track with GEMIAC, indicated preference and functional efficacy, and provided free-text justifications. Quantitative analyses use Poisson GLMs on aggregate counts after a singular GLMM fit; qualitative coding follows Braun and Clarke. The paper claims that listeners dissociate preference from perceived efficacy by musical origin: AI music is preferred more, human music is judged more effective, and this pattern persists regardless of labeling. The conclusion argues that preference alone is insufficient to evaluate generative music for functional emotional applications.","tokens_in":19232,"tokens_out":6110,"duration_ms":63385,"significance":"The research question is timely and the mixed-methods design is thorough: attention checks, multiple-comparison corrections, transparent reporting of a singular model fit, and a named thematic-analysis procedure are all strengths. The preference–efficacy dissociation is an interesting and practically relevant construct for evaluating generative music in wellness contexts. However, the central class-level claim is not supported because musical origin is perfectly confounded with the specific track and its acoustic features; the study is best interpreted as a hypothesis-generating pilot. If reframed accordingly, the qualitative insights and the demonstrated dissociation for these particular stimuli could still be a useful contribution.","major_comments":[{"comment":"Origin is perfectly confounded with track identity. For each emotion case there is exactly one AI and one human track, and Appendix A documents systematic acoustic differences: the Calm pair differs in tempo (72 vs 53 BPM) and instrumentation (layered synth pads vs rubato acoustic piano), and the Upbeat pair differs in tempo (91 vs 105 BPM) and production (loop-based electronic vs live guitar and percussion). The GLMM with a random effect for condition produced a singular fit (Section 3.1.1), and the fallback Poisson GLMs on aggregate counts cannot separate authorship from stimulus identity. Therefore the dissociation reported in Section 4.1.3 (human music more effective than preferred; AI music more preferred than effective) is a statement about these four tracks, not about AI-generated versus human-composed music as classes. The abstract's and conclusion's class-level claims should be qualified accordingly.","section":"3.1.1, Appendix A"},{"comment":"The abstract's claim that 'participants were significantly more likely to rate human-composed music, regardless of labeling, as more effective at eliciting target emotional states' is not supported in the Calm condition. Table 2 shows equal total efficacy selections for AI and human music in Calm (69 vs 69), and Section 4.1.2 reports β ≈ 0, z = 0.00, p = 1.00 for this contrast. The significant human advantage appears only in the Upbeat condition (β = 1.17, z = −5.98). The global claim should be restricted to the Upbeat case or reported as an interaction.","section":"4.1.2, Table 2, Abstract"},{"comment":"The pooled preference–efficacy dissociation masks opposite patterns by emotion case. In Calm, AI music was preferred more often than human music (92 vs 49 total selections in Table 1) while efficacy was equal; in Upbeat, human music was both preferred more (83 vs 64) and more often judged effective (110 vs 34). Averaging across emotion cases therefore conflates origin with the specific tracks and emotion-case context. Reporting the dissociation separately for each emotion case, or with a three-way interaction among origin, outcome type, and emotion case, would be necessary to support the claim that listeners systematically dissociate preference from efficacy as a function of origin.","section":"4.1.3"}],"minor_comments":[{"comment":"The sentence 'These results suggests that participants were more likely...' contains a subject–verb agreement error; it should read 'These results suggest...'.","section":"4.1.1"},{"comment":"The column labels 'Neither (Neg.)', 'Neither (Pos.)', and 'Neither (Neutral)' are not defined in the captions; the hand-coding of open-ended 'neither' responses should be described in the main text or caption.","section":"Tables 1 and 2"},{"comment":"The paper notes that a GLMM with a random effect for condition had a singular fit and therefore uses Poisson GLMs on aggregate counts, but each participant contributes two preference and two efficacy judgments (one per emotion case), so the observations are not independent; this non-independence is not addressed by the aggregate Poisson models and should be acknowledged as a limitation.","section":"3.1.1"},{"comment":"The limitations paragraph acknowledges 'stimulus variety and sample size' but does not explicitly state that the one-stimulus-per-cell design prevents any class-level inference about AI-generated versus human-composed music; this consequence should be stated directly.","section":"5.4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central claim is undermined by the stimulus confound, but the paper already positions itself as a pilot study and discloses the limitation qualitatively. A major revision that reframes the conclusions to be track-specific and explicitly advises against class-level generalization would be credible. The second author's affiliation with Myndstream, the source of the human-composed stimuli, is disclosed in the byline, but the editor may wish to consider whether the stimulus selection should be independently audited in a revision. The paper is longer than typical for the reported scope; condensing Sections 2 and 5 would improve readability."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a promising pilot on a question that matters—whether preference ratings are a good proxy for whether generative music actually does the emotional job—but the headline dissociation is about four specific tracks, not about AI vs human music as classes, and the abstract overstates the efficacy result. That said, it deserves referee time, not a desk reject.\n\nWhat's genuinely new: the explicit preference-efficacy dissociation measured in a functional emotion setting, plus deceptive-labeling data with GEMIAC. Prior work (Shank et al.) showed labels change liking; this paper goes further and shows people can prefer AI music while still crediting human music with stronger functional efficacy, and that labels mattered little. The qualitative material is rich, the thematic analysis follows a named method, and the labeling-accuracy sanity check in Section 4.2 is a nice touch. Multiple-comparison corrections are reported.\n\nWhere it's soft, in proportion: the one-stimulus-per-condition design means authorship is perfectly confounded with track identity. Appendix A gives the evidence: 72 vs 53 BPM in Calm, synth pads vs acoustic piano; 91 vs 105 BPM in Upbeat, loop-based vs live guitar. No statistical model can separate \"AI\" from \"this particular Suno output\" because the random-effect term collapsed. The paper acknowledges limited stimulus variety but still states class-level claims in the abstract. The abstract also says human music was significantly more likely to be rated effective overall; in the Calm condition, efficacy selections were exactly tied (69 vs 69), so this holds only for Upbeat. The preference-efficacy dissociation itself appears in both emotion cases in the raw counts, but the significance tests are on aggregates and need reanalysis with more stimuli. There is also no IRB/consent statement in the text (with deception present, that's a must), no data or audio release, and I'd want a conflict-of-interest statement given the second author's company supplied the human tracks.\n\nMy take: the core idea is worth pursuing and the direction of the effect is plausible. As reported, it is a pilot with a load-bearing confound, not a demonstrated difference between human and AI music as classes. It is fixable: more stimuli per cell, preregistered mixed models with stimulus random effects, accurate abstract, and full materials. I'd recommend sending to peer review with a major-revision expectation rather than rejecting, because the question and the dissociation are genuinely useful.","headline":"Promising pilot on a real gap, but the central dissociation is confounded by one track per condition and the abstract overstates the efficacy effect.","tokens_in":19780,"tokens_out":4031,"would_cite":false,"duration_ms":44754,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that listeners separate liking from emotional effectiveness: human-composed music wins on efficacy, AI-generated music wins on preference.","keywords":["AI-generated music","human-composed music","music emotion regulation","preference versus efficacy","perceived authenticity","generative music evaluation","GEMIAC","listener perceptions"],"falsifier":"Replicate the preference-versus-efficacy comparison with a corpus of many AI-generated and many human-composed tracks per emotion case, matched on tempo, instrumentation, and production polish, and fit a model with random effects for individual tracks. If the dissociation disappears or reverses once origin is varied within matched tracks, the paper's central claim is an artifact of the two specific stimuli rather than a general origin effect.","tokens_in":18752,"feed_emoji":"🎧","tokens_out":4783,"duration_ms":46871,"temperature":0.7,"pith_summary":"This paper tries to establish that listeners separate liking a piece of music from trusting it to do an emotional job, and that the separation depends on perceived authorship. In a 152-participant study with Calm and Upbeat music, human-composed tracks were significantly more likely than AI-generated tracks to be chosen as the most effective at eliciting the target emotion, even though measured emotional responses were largely similar. At the same time, AI-generated tracks were significantly more likely to be preferred. None of this shifted with whether the music was labeled, mislabeled, or unlabeled, which the authors read as evidence that authenticity judgments are tied to perceived authorship rather than actual origin. The paper concludes that preference ratings alone are not a valid proxy for functional success in generative music systems built for emotion regulation.","feed_headline":"People prefer AI music, yet call human tracks more effective","feed_subtitle":"A 152-listener study finds liking and emotional function split, so preference alone cannot judge generative music","key_machinery":"The load-bearing design is a three-way labeling manipulation, correctly labeled, incorrectly labeled, and unlabeled music, applied to paired one-minute instrumental tracks in two emotion cases, with each participant reporting preference, efficacy, and emotion-intensity ratings. The central object is the preference-efficacy pair: for each emotion case, participants chose which of two songs they preferred and which more effectively conveyed the target emotion. The paper then compares those two choice distributions with Poisson generalized linear models, and the dissociation is quantified as opposite signs in the origin coefficients for preference versus efficacy. The qualitative arm, coded by thematic analysis, supplies the mechanism listeners themselves report: perceived humanness, expressed as imperfection, flow, and soul, anchors efficacy judgments even when preference goes the other way.","core_discovery":"The central discovery is a preference-efficacy dissociation: human-composed music was significantly more likely to be selected as most effective than as most preferred, while AI-generated music was significantly more likely to be preferred than effective. This dissociation appeared against a background of null results: participants' emotional responses did not differ significantly by label or, in most comparisons, by origin, and labeling condition had no meaningful effect on perceived efficacy. The authors interpret this as showing that 'what works' emotionally and 'what I like' are distinct judgments, and that generative music evaluation schemes built on preference alone can miss the functional dimension. Qualitative responses reinforce the interpretation: listeners associated humanness with imperfection, flow, organic quality, and soul, and often assigned human characteristics to AI music when they believed it was human-composed.","pith_inferences":["If the preference-efficacy dissociation generalizes, then human-annotation pipelines that collect only 'do you like this?' will systematically misalign with therapeutic or functional goals; adding a separate 'would this work for the intended state?' question would be a cheap, testable fix.","The confusion between actual and perceived authorship raises the possibility that the dissociation is driven by a mental model of what a human composer sounds like, rather than by audible differences; a study with matched production style could separate those.","The finding that listeners who preferred AI music often mislabeled it as human suggests that preference may partly be preference for perceived humanness; an implicit-association design could quantify that overlap.","For generative wellness music, the practical design consequence may be to preserve audible traces of human performance, such as rubato, dynamic shaping, and small imperfections, rather than chasing perceptual indistinguishability."],"forward_implications":["Preference cannot serve as the sole optimization target for music generation systems aimed at mood regulation; systems trained on preference feedback may optimize for liking while missing functional efficacy.","Labeling a piece as human or AI does not change how effective listeners find it, so perceived efficacy appears robust to framing, at least for these ambient and upbeat styles.","Listeners may describe AI music as preferred while still choosing human-composed music for emotional work, so wellness and therapeutic applications should evaluate functional outcomes, not just appeal.","The qualities listeners associate with humanness, micro-expressive variation, idiosyncratic phrasing, and imperfection, become concrete design targets for generative systems seeking emotional resonance.","Because emotion-intensity responses were similar across origins, the dissociation is not explained by large differences in felt emotion; it lives in the comparison judgments, not in raw affect."],"supporting_citations":[{"why":"Supplies the GEMIAC instrument used to measure music-induced emotions across the experimental conditions.","marker":"[2]"},{"why":"Supplies the STAI-S state anxiety measure used as the exploratory anxiety outcome.","marker":"[1]"},{"why":"Provides the prior result that listeners enjoy music less when they believe it was AI-composed, which motivates the labeling hypotheses.","marker":"[37]"},{"why":"Describes MusicRL, a preference-feedback alignment system whose simplified rating scheme is the target of the paper's critique of preference-only evaluation.","marker":"[41]"},{"why":"Describes EmotionBox, an emotion-sensitive generation system whose less intensive preference evaluation this study aims to extend.","marker":"[40]"},{"why":"Supplies the Braun and Clarke thematic analysis guidelines that structure the qualitative coding of free-text responses.","marker":"[45]"},{"why":"Provides the SpecTTTra/SONICS synthetic-song detection framework used to frame authenticity and source detection in the discussion.","marker":"[23]"},{"why":"Provides the LLark multimodal music model, referenced as an open-source approach to music analysis and authenticity reasoning.","marker":"[24]"}],"fun_headline_variants":["AI music wins preference, humans win efficacy","Liking vs. effectiveness: listeners split AI and human music","Preference favors AI, but human music feels more effective","AI beats human on preference, loses on emotional function","Listeners prefer AI music yet rate human tracks more effective"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results assume that the single AI-generated track and the single human-composed track used in each emotion case represent their entire classes; since the paired tracks differ systematically in tempo, instrumentation, and production style, authorship is completely confounded with the specific piece, so the preference-efficacy dissociation could be about the tracks rather than about human versus AI origin.","fun_headline_variants_meta":{"raw":{"variants":["AI music wins preference, humans win efficacy","Liking vs. effectiveness: listeners split AI and human music","Preference favors AI, but human music feels more effective","AI beats human on preference, loses on emotional function","Listeners prefer AI music yet rate human tracks more effective"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000262,"raw_usage":{"total_tokens":1588,"prompt_tokens":930,"completion_tokens":658,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":580}},"tokens_in":546,"tokens_out":658,"duration_ms":6680,"temperature":1.0,"reasoning_tokens":580,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:14:02.012428+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replicate the preference-versus-efficacy comparison with a corpus of many AI-generated and many human-composed tracks per emotion case, matched on tempo, instrumentation, and production polish, and fit a model with random effects for individual tracks. If the dissociation disappears or reverses once origin is varied within matched tracks, the paper's central claim is an artifact of the two specific stimuli rather than a general origin effect.","supporting_citations":[{"cited_title":"Introducing the geneva music-induced affect checklist (gemiac) a brief instrument for the rapid assessment of musically induced emotions","cited_arxiv_id":null,"evidence_quote":"Supplies the GEMIAC instrument used to measure music-induced emotions across the experimental conditions."},{"cited_title":"The state-trait anxiety inventory","cited_arxiv_id":null,"evidence_quote":"Supplies the STAI-S state anxiety measure used as the exploratory anxiety outcome."},{"cited_title":"Ai composer bias: Listeners like music less when they think it was composed by an ai","cited_arxiv_id":null,"evidence_quote":"Provides the prior result that listeners enjoy music less when they believe it was AI-composed, which motivates the labeling hypotheses."},{"cited_title":"Emotionbox: A music-element-driven emotional music generation system based on music psychology","cited_arxiv_id":null,"evidence_quote":"Describes EmotionBox, an emotion-sensitive generation system whose less intensive preference evaluation this study aims to extend."},{"cited_title":"Thematic analysis","cited_arxiv_id":null,"evidence_quote":"Supplies the Braun and Clarke thematic analysis guidelines that structure the qualitative coding of free-text responses."}],"review_version":1}