{"id":"3bb8615d-a3f4-41f8-bb85-00501f8097c0","arxiv_id":"2509.00813","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"AImoclips is a new open benchmark showing that text-to-music systems convey high-arousal emotions better than low-arousal ones and that all models converge toward emotionally neutral music.","lead":"This paper introduces AImoclips, a benchmark of 991 AI-generated music clips from six text-to-music systems, rated by 111 human listeners for emotional valence and arousal. It finds that commercial models tend to sound more pleasant than intended, while all models bias toward emotional neutrality, especially for fine-grained valence.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground-truth emotion norms are English (Warriner) while all raters are Korean; the sign of model-specific valence deviations and the quadrant grouping may be artifacts of this mismatch.","rationale":"The central claim consists of three parts: (i) all systems centralize toward neutrality; (ii) commercial systems are more pleasant than intended, open-source less; (iii) these deviations are model-specific and reliable. Part (ii) is the most distinctive and load-bearing because it frames an absolute evaluative judgment about model families. The only anchor for 'intended' is Warriner's English norms. The raters are Korean; without evidence that these norms match Korean affective judgments, the sign of the deviation is an artifact of the norm source. The relative ordering between models would survive a constant shift, but the paper does not report whether the post-exclusion dataset has equal clip counts per system-intent, and the absolute directions in the abstract do not. The quadrant analysis (Section 4.3) is also dependent on norm-based quadrant assignments; if a word's arousal falls on the other side of 5.0 for Korean raters, the 'high-arousal advantage' conclusion could change. This is not a mere cross-cultural nicety: GlobalMood is cited in Related Work, indicating the authors are aware of cross-cultural variation, yet no calibration is performed. The proposed calibration study directly measures the population-specific ground truth and re-runs the same statistics, settling whether the headline directions are robust. The paper still has strong merits: a large open dataset, six systems, 111 raters, and significant ANOVAs; the concern is correctable and does not invalidate the benchmark.","tokens_in":7420,"tokens_out":11603,"duration_ms":154794,"concrete_test":"Recruit at least 30 new Korean-speaking participants (not from the original rating pool) to rate the 12 English emotion words on the same 9-point valence/arousal scales used in the survey. Use the mean of these ratings as population-matched ground truth, then recompute the model-wise signed deviations (replicating Figure 3a and the ANOVA in Section 4.2) and the per-quadrant absolute-deviation analysis (Section 4.3). If any model's mean valence deviation changes sign, or if the quadrant with the smallest absolute deviation changes, the headline result is not robust to the English-norm assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's signed-deviation analysis (Section 4.2, Figure 3a) subtracts Warriner et al. (2013) English word norms from clip ratings, yet all raters are fluent Korean speakers (Section 3.3) and no Korean affective norms for the 12 intent words were collected. The benchmark therefore equates 'intended emotion' with English-language norms for a Korean-listener population. If the Korean valence/arousal norms for words like 'scared' or 'dull' differ in mean or extremity, every clip's deviation shifts by an intent-specific constant. Because the headline claims are about the sign of deviation ('commercial more pleasant than intended; open-source less pleasant'), these signs are not identified without population-matched ground truth. The 'high-arousal advantage' analysis (Section 4.3) is also vulnerable because intents were assigned to quadrants using English norms; a word that is high-arousal for American English speakers may cross the arousal boundary for Korean raters, changing the quadrant grouping. The paper mentions GlobalMood [22] but does not use it to calibrate or acknowledge this cultural mismatch. Inter-rater reliability is also unreported, but the norm mismatch is the more direct threat to the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents AImoclips, a benchmark for evaluating emotion conveyance in text-to-music (TTM) generation. The authors select 12 English emotion words spanning four valence–arousal quadrants, generate 1,008 clips with six TTM systems (four open-source, two commercial), and collect continuous valence/arousal ratings from 111 Korean-speaking participants. After excluding 17 clips with few ratings, 991 clips remain. Using Warriner et al.'s English affective norms as ground truth, the paper reports that all systems show a centralizing tendency toward neutrality, commercial models produce higher valence than intended while open-source models produce lower valence, and high-arousal intents are conveyed more accurately. Statistical significance is assessed with two-way ANOVAs and pairwise comparisons.","tokens_in":7670,"tokens_out":6344,"duration_ms":78685,"significance":"The dataset is a useful new resource: it provides publicly available AI-generated clips with dense valence/arousal annotations, covers a broader model set than prior work (cf. Gao et al. [23]), and addresses an underexplored evaluation dimension. The ANOVA results are reported with effect sizes, and the paper is generally transparent about clip generation and survey design. If the ground-truth norm issue is resolved, the benchmark could support future affective-controllability research. However, the headline signed-deviation claims are conditional on an unexamined cross-cultural assumption, and reliability evidence is missing; these issues must be addressed before the benchmark's conclusions can be taken as established.","major_comments":[{"comment":"The signed deviations in Fig. 3a are computed as clip ratings (from 111 fluent Korean speakers, §3.3) minus Warriner et al. [26] English word norms (§4.2). If Korean valence/arousal norms for the 12 intent words differ from English norms, each clip's deviation shifts by an intent-specific constant, so the sign of per-model mean deviation—the basis for the claim that commercial systems are 'more pleasant than intended' and open-source systems are 'less pleasant'—can change even though the model main effect in the ANOVA is unchanged. The quadrant grouping in §4.3 also uses English norms; words such as 'scared' or 'dull' may cross valence/arousal boundaries for Korean raters. The authors should collect Korean norms from the same participant population, or provide a sensitivity analysis showing which conclusions survive plausible intent-level norm offsets, and discuss the limitation explicit","section":"§3.1, §3.3, Fig. 3a"},{"comment":"No inter-rater reliability statistic (e.g., ICC or Krippendorff's alpha) is reported. With only 4–9 ratings per clip, the benchmark's claim to measure 'conveyed emotion' per clip requires evidence of agreement; without it, model-specific deviations may partly reflect rater noise. Please report reliability per model and quadrant, and discuss the minimum number of ratings needed.","section":"§3.3, §4.1"},{"comment":"Seventeen clips with ≤3 ratings were excluded, but the per-model and per-intent distributions of excluded clips are not reported. If exclusions concentrate in one system (e.g., generation failures or extreme content), the reported means and ANOVAs could be biased. Please report the exclusion table and confirm the main results are stable when all 1,008 clips are analyzed (e.g., with appropriate weighting).","section":"§3.3, §4.1"}],"minor_comments":[{"comment":"Typos: 'activites' should be 'activities' (§3.3); 'such ashappy' should be 'such as happy' (§4.3).","section":"§3.3, §4.3"},{"comment":"Figure 2 is referenced as 'presented in 2'; should be 'presented in Figure 2'.","section":"§4.1"},{"comment":"The corresponding author email contains a corrupted sequence ('envel⌢pe-⌢penrotation@kaist.ac.kr'); please fix.","section":"Author block"},{"comment":"Please state whether the 12 intent words were presented to participants in English or Korean during the rating task; this is relevant to interpreting the ground-truth comparison.","section":"§3.3"},{"comment":"The sample-rate explanation for valence differences is speculative; consider citing supporting evidence or phrasing it as a hypothesis.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"The central issue is the ground-truth norms. I believe the paper can be repaired by collecting Korean norms or reframing claims as relative to English lexical norms, and by adding reliability and exclusion analyses. The dataset contribution is solid, so I would not reject if these points are addressed convincingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, this is one of the first benchmarks to look specifically at how well TTM systems convey intended emotion, using continuous valence-arousal ratings from 111 listeners across 991 clips from six systems (four open-source, two commercial). The dataset is open, the analysis is straightforward ANOVA with clear figures, and the comparison to Gao et al. is honest. For that reason it's a useful contribution to TTM evaluation.\n\nSecond, the central claim about model-specific valence deviations—commercial models rated more pleasant than intended, open-source less pleasant—is built on a ground-truth assumption that isn't tested. The authors use Warriner et al.'s English word norms as the intended emotion scores, but all raters are fluent Korean speakers. No Korean affective norms were collected for the 12 intent words. If Korean valence/arousal norms for words like 'scared' or 'dull' differ from English means, every model's deviation shifts by an intent-specific constant. That shifts the signs of the deviations. The commercial/open-source split could survive, but you can't tell from the current analysis. The same problem affects the quadrant grouping in Section 4.3, since quadrants are assigned using the English norms. This isn't a killer—the centralizing tendency (all models drift toward neutral) is likely robust, and the absolute-deviation analysis is less sensitive to the shift—but the sign claims in Section 4.2 and Figure 3a are genuinely unidentified.\n\nOther issues are smaller. No inter-rater reliability is reported, which matters when you're averaging 4–9 ratings per clip. Seventeen clips with few ratings were excluded, and the paper doesn't show whether they cluster by model or intent. Sample rates differ across models, and the authors note this, which is fine.\n\nBottom line: the benchmark itself is worth having, and the analysis is competent. But before I'd take the headline findings as reliable, the authors need to either collect or borrow Korean affective norms for the intent words, or rephrase the claims as relative differences that don't depend on the norm scale. That's doable in revision. The paper deserves peer review; I'd send it to a workshop or conference reviewer with that requirement.","headline":"A genuinely useful new benchmark for emotion conveyance in text-to-music, but the headline commercial-vs-open-source valence claim is hostage to an unexamined English-norm/Korean-rater mismatch.","tokens_in":8155,"tokens_out":3241,"would_cite":true,"duration_ms":37602,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper builds a benchmark of 991 AI music clips and 6,162 human ratings to test whether text-to-music systems deliver the emotions they are prompted with, and finds that all systems drift toward neutrality while commercial and open-sourc","keywords":["text-to-music generation","emotion conveyance","valence-arousal model","human evaluation","affective controllability","music generation benchmark","open-source vs commercial models","emotional neutrality bias"],"falsifier":"Recompute every model deviation using valence and arousal norms for the 12 emotion words collected from Korean-speaking raters. If the commercial-versus-open-source split or the universal pull toward neutrality disappears or reverses, the paper's central claim is an artifact of using English norms as ground truth rather than a stable property of the systems.","tokens_in":7329,"feed_emoji":"🎵","tokens_out":6156,"duration_ms":65561,"temperature":0.7,"pith_summary":"Text-to-music systems promise to turn a prompt like \"anxious, instrumental\" into music that actually sounds anxious. AImoclips tests that promise by collecting 991 clips from six current systems and continuous valence–arousal ratings from 111 listeners. The benchmark's central result is that none of the systems reliably lands on the intended emotion: every model compresses perceived emotion toward the neutral center of the valence–arousal plane. Commercial systems overshoot toward pleasant, open-source systems undershoot, and high-arousal emotions such as angry or excited are conveyed noticeably better than low-arousal ones. The authors argue these deviations are model-specific and stable, making emotion prompts a currently unreliable control mechanism.","feed_headline":"All six text-to-music models drift toward neutral emotion","feed_subtitle":"A 991-clip, 111-listener benchmark shows emotional prompts land off-target, with commercial systems skewing pleasant.","key_machinery":"The load-bearing object is AImoclips itself: an open dataset of 991 ten-second clips, each generated from one of 12 emotion words chosen to cover the four quadrants of the valence–arousal plane, with each clip rated on valence and arousal by 4 to 9 of the 111 participants. The analytic mechanism is the deviation score, the difference between average listener ratings and the emotion word's English normative valence/arousal score, aggregated per model, per quadrant, and per emotion intent, then tested with two-way ANOVA and pairwise comparisons. This turns \"does the music sound like the emotion word?\" into a numeric quantity that can be compared across systems.","core_discovery":"On its own terms, the paper's discovery is a reproducible, model-specific gap between the emotion a text prompt names and the emotion listeners actually hear. Averaging human ratings per clip and subtracting the emotion word's normative scores shows that all six systems pull perceived valence and arousal toward the center: generated music sounds emotionally blander than the word that prompted it. The pull is not symmetric. Suno and Udio, the two commercial systems, produce music rated as more pleasant than the intent, while the four open-source systems produce music rated as less pleasant; in arousal, AudioLDM 2 and Mustango skew low while the rest skew high. A two-way ANOVA and pairwise com","pith_inferences":["Editorial inference: because ground-truth scores come from English word norms while all 111 raters are fluent Korean speakers, the reported deviations probably mix true model bias with cross-linguistic differences in what emotion words mean; collecting Korean norms for the same 12 words would separate the two.","Editorial inference: the commercial pleasantness advantage could be explained by audio quality or production style rather than semantic emotion fidelity; a matched experiment controlling loudness, sample rate, and production would test this.","Editorial inference: the tendency toward neutrality may be partly a measurement effect of averaging across raters or of cropping random 10-second segments; per-rater distributions or whole-clip ratings would show whether the center bias is in the models or the metric."],"forward_implications":["If the centralizing tendency is general, emotion words alone are not a dependable control interface for TTM systems; expressive extremes need additional conditioning or post-generation editing.","The reliable split between commercial and open-source valence biases gives model developers and auditors a concrete target: commercial systems appear to carry a positivity bias, open-source systems a negativity bias.","Better conveyance of high-arousal intents implies that low-arousal affect is the harder control problem and should get focused attention in model training and evaluation.","AImoclips can be reused as a training set for automatic emotion predictors or as a fine-tuning signal to align TTM models with perceived rather than intended emotion."],"supporting_citations":[{"why":"Supplies the English valence and arousal norms used as ground-truth scores for the 12 emotion words.","marker":"[26]"},{"why":"Supplies the circumplex-model sample words from which the emotion intents were selected.","marker":"[24]"},{"why":"One of the four open-source TTM systems that generated the benchmark clips.","marker":"[28]"},{"why":"One of the four open-source TTM systems that generated the benchmark clips.","marker":"[29]"},{"why":"One of the four open-source TTM systems that generated the benchmark clips.","marker":"[30]"},{"why":"One of the four open-source TTM systems that generated the benchmark clips.","marker":"[31]"},{"why":"One of the two commercial TTM systems whose output shows the pleasantness overshoot.","marker":"[32]"},{"why":"One of the two commercial TTM systems whose output shows the pleasantness overshoot.","marker":"[33]"},{"why":"Human-preference benchmark invoked to explain why commercial outputs may be perceived as more pleasant.","marker":"[8]"},{"why":"Earlier emotion-conveyance study whose findings on high-arousal emotional expression this benchmark extends.","marker":"[23]"}],"fun_headline_variants":["All text-to-music AI drift to neutral, commercial skew pleasant","Emotional prompts in music AI: every model goes bland","Music AI misses emotional cues, all six models flatten affect","Six TTM models all pull toward neutral emotion, study finds","Text-to-music AI: even commercial models overshoot pleasantness"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The benchmark treats English word norms as the true valence and arousal of each emotion intent, even though all 111 raters were fluent Korean speakers; if affective word meanings differ across languages, the measured deviations shift by that difference.","fun_headline_variants_meta":{"raw":{"variants":["All text-to-music AI drift to neutral, commercial skew pleasant","Emotional prompts in music AI: every model goes bland","Music AI misses emotional cues, all six models flatten affect","Six TTM models all pull toward neutral emotion, study finds","Text-to-music AI: even commercial models overshoot pleasantness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000124,"raw_usage":{"total_tokens":934,"prompt_tokens":733,"completion_tokens":201,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":115}},"tokens_in":477,"tokens_out":201,"duration_ms":3425,"temperature":1.0,"reasoning_tokens":115,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T13:09:52.354551+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute every model deviation using valence and arousal norms for the 12 emotion words collected from Korean-speaking raters. If the commercial-versus-open-source split or the universal pull toward neutrality disappears or reverses, the paper's central claim is an artifact of using English norms as ground truth rather than a stable property of the systems.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the English valence and arousal norms used as ground-truth scores for the 12 emotion words."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the circumplex-model sample words from which the emotion intents were selected."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"One of the four open-source TTM systems that generated the benchmark clips."},{"cited_title":"Copet, F","cited_arxiv_id":null,"evidence_quote":"One of the four open-source TTM systems that generated the benchmark clips."},{"cited_title":"Evans, J","cited_arxiv_id":null,"evidence_quote":"One of the four open-source TTM systems that generated the benchmark clips."},{"cited_title":"Accessed: May-June, 2025","cited_arxiv_id":null,"evidence_quote":"One of the two commercial TTM systems whose output shows the pleasantness overshoot."},{"cited_title":"Accessed: May-June, 2025","cited_arxiv_id":null,"evidence_quote":"One of the two commercial TTM systems whose output shows the pleasantness overshoot."},{"cited_title":"Grötschla, A","cited_arxiv_id":null,"evidence_quote":"Human-preference benchmark invoked to explain why commercial outputs may be perceived as more pleasant."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier emotion-conveyance study whose findings on high-arousal emotional expression this benchmark extends."}],"review_version":1}