{"id":"af1f58fc-be65-4624-8e88-972fdb183c1e","arxiv_id":"2509.00051","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A taxonomy and critical review of evaluation metrics for music generation, identifying gaps such as weak correlation with human perception and lack of standardization.","lead":"This paper surveys the metrics used to evaluate computer-generated music, grouping them into a taxonomy for audio and symbolic formats. It argues that current evaluation is inconsistent, culturally biased, and poorly connected to what human listeners actually enjoy.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'poor correlation' premise rests on a single citation and an uncited assertion; if that evidence is weaker than claimed, the central motivation for a new framework is under-supported.","rationale":"Reader's weakest_assumption is essentially the same: representativeness and accuracy of cited findings. I agree. The survey is useful as a taxonomy, and many metrics are listed, but the central claim of a research gap is largely a negative/empirical claim about the literature. The paper does not describe a systematic review method, and the key negative claim about objective metrics is supported by one citation plus an uncited assertion. This does not make the paper fraudulent or worthless; it means the central motivation is conditionally acceptable. The taxonomy and future directions can stand, but the critical review should either cite the primary correlation studies with numbers or qualify the claims. Since the reader already issued CONDITIONAL, I do not recommend changing the verdict; I only sharpen the specific empirical check that would settle it.","tokens_in":21779,"tokens_out":5414,"duration_ms":65262,"concrete_test":"Extract the actual human-correlation results from Yuan et al. 2025 (Yue, arXiv:2503.08638) for CLAP-score, FAD, and KLD: locate the table/figure reporting correlation with human preference judgments and record the coefficients and the evaluation setup (which human ratings, which prompts, which embedding backbones). Do the same for KAD (Chung et al. 2025) and MAD (Huang et al. 2025) to verify the claimed 'better correlation than FAD.' If the coefficients are uniformly low (e.g., Spearman <0.3) and the conditions match the survey's broad claim, the concern is resolved. If any coefficient is moderate/high, or if the comparisons were run under conditions that do not generalize to the survey's blanket statement, the critical premise fails and the 'no comprehensive framework' motivation needs to be softened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that current objective metrics suffer from 'poor correlation between objective metrics and human perception' is load-bearing for the survey's motivation, but its support in the manuscript is thin. Section 4.1 attributes the claim to a single study: 'Yuan et al. (2025) showed that many widely used objective metrics, such as CLAP-score, FAD, and KLD, often align poorly with human preferences.' No correlation coefficients, experimental conditions, or comparison baselines are reported. In Section 3.1.1, the survey states that 'Both KAD and MAD metrics have shown better correlation with human preferences than FAD' without citation or numbers. The taxonomy's usefulness does not depend on every metric being flawed, but the 'critical review' and the 'research gap' do. If Yuan et al. actually measured something narrower (e.g., correlation on a specific TTM benchmark or with a particular prompt set), or if KAD/MAD's reported advantage is limited to one embedding/feature configuration, then the survey's generalization to 'many widely used objective metrics' is not established. This is an internal-evidence gap, not a disagreement with consensus: the survey itself does not supply the evidence needed to support its strongest critical claim. Compounding this, the survey describes no systematic search protocol, so it cannot demonstrate that the cited findings are representative.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper surveys evaluation metrics for music generation, organizing them into a taxonomy (Figure 2) spanning reference-based and reference-free objective metrics for audio and symbolic music, human evaluation criteria, and benchmarks. It argues that current evaluation methodologies suffer from poor correlation between objective metrics and human perception, cross-cultural bias, and lack of standardization, and it proposes future directions including a human-in-the-loop RL-trained aesthetic predictor. Appendices provide metric definitions and toolkits.","tokens_in":22086,"tokens_out":5192,"duration_ms":52725,"significance":"If the taxonomy and critical review are accurate, the paper would be a useful reference for the music-generation community. The explicit organization of symbolic and audio metrics, the inclusion of recent human-preference datasets, and the attention to cross-cultural bias are strengths. The main value, however, rests on two load-bearing premises: that the surveyed metric set is representative, and that the empirical evidence for the 'poor correlation' limitation is reliable. These premises currently receive insufficient support, which limits the confidence one can place in the survey's central motivation.","major_comments":[{"comment":"The central critical claim that objective metrics such as CLAP-score, FAD, and KLD 'often align poorly with human preferences' is attributed to Yuan et al. (2025) with no correlation coefficients, task scope, or comparison baselines; the earlier assertion that KAD and MAD have 'shown better correlation with human preferences than FAD' is cited to no source at all. Because this claim motivates the entire review and the proposed framework, please report the quantitative evidence (and its domain/generalizability) or soften the claim to match the evidence.","section":"§4.1 and §3.1.1"},{"comment":"The survey describes itself as a comprehensive overview, but it does not describe a systematic search protocol, inclusion/exclusion criteria, or coverage analysis. Without such a protocol, the representativeness of Figure 2 and the critical review cannot be assessed, and the reader cannot tell whether important metrics or benchmarks were omitted. Please add a methodology subsection detailing databases, years, search terms, screening, and the number of papers considered, or qualify the claims as representative rather than comprehensive.","section":"Figure 2 and §3"},{"comment":"Section 3.3 describes MusicPrefs as having crowdsourced pairwise ratings for fidelity and musicality, while Section 4.1 groups it with preference datasets that rely 'solely on overall impression' (Huang et al., 2025; Liu et al., 2025a). These statements are in tension. If MusicPrefs uses dimensional ratings, the limitation claim should be revised; if not, the Section 3.3 description should be corrected.","section":"§3.3 vs §4.1"}],"minor_comments":[{"comment":"'With MIDI datasets being the most popular for example- Lakh MIDI Dataset (Raffel, 2016), Popular examples include' is a grammatical duplication; clean up.","section":"§2.3"},{"comment":"'distuned' should be 'out-of-tune' or 'dissonant'.","section":"§4.1"},{"comment":"'There is Figure 2 lists the metrics' is ungrammatical; remove 'There is'.","section":"§3.1.2"},{"comment":"The formula for MOA is missing a closing parenthesis after b_i^(y).","section":"Appendix B, Eq. (3)"},{"comment":"'Adherance' should be 'Adherence'.","section":"Appendix A.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a legitimate survey with a coherent taxonomy. The main risk is overgeneralizing from a single citation; the authors should address the three major comments before publication. I saw no red flags regarding circularity or self-citation. The paper fits the journal's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the survey. It's a competent organizing effort, probably the most complete taxonomy of music generation evaluation metrics I've seen, including the newer KAD, MAD, FAD∞, and aesthetic predictors. Figure 2 alone is worth a cite. The critical review is mostly fair: the points about interpretability, Western-centric bias, and lack of standardization are real and well-supported. The proposed three-component framework and RL-based scorer are sensible as a research agenda, not as a result.\n\nThe soft spot is exactly where the stress-test note points. The claim that 'many widely used objective metrics… align poorly with human preferences' rests on a single citation (Yuan et al., 2025) with no numbers, and the KAD/MAD-better-than-FAD sentence in Section 3.1.1 is uncited. That's a weakness in the survey's critical apparatus, not a fatal flaw. The taxonomy doesn't depend on every metric being broken; it's useful regardless. But if the paper wants to make the gap argument stick, it should report the actual correlation coefficients and contexts.\n\nAlso, there's no description of how the survey was put together (search strategy, inclusion criteria). For a survey claiming comprehensiveness that's a legitimate reviewer ask, but it's correctable.\n\nI'd send this to peer review. A good referee will ask for the missing numbers and a methodology paragraph, and possibly a table summarizing human-correlation evidence across metrics. The reading group might appreciate it, especially the tables and figures, but it's more reference than discussion piece.\n\nVerdict: competent survey, worth citing, needs minor-to-moderate revisions.","headline":"Useful taxonomy, uneven evidence for the headline critique; worth a serious referee but needs the correlation claims backed by numbers.","tokens_in":22509,"tokens_out":1853,"would_cite":true,"duration_ms":22032,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Music evaluation has not kept pace with generation: no comprehensive framework exists for judging generated scores, and common metrics often disagree with human perception.","keywords":["music generation evaluation","evaluation metrics","taxonomy","text-to-music generation","symbolic music","audio generation","human evaluation","cross-cultural bias"],"falsifier":"Settle the central gap claim by checking a fixed slice of recent music-generation papers (say, 50 papers from 2024–2025) for a shared evaluation protocol—same metric battery, same reference sets, same thresholds, same reporting format. If such a protocol is already in common use, the no-comprehensive-framework claim collapses; if every paper improvises its own criteria, the survey's diagnosis is confirmed.","tokens_in":21712,"feed_emoji":"🎵","tokens_out":8141,"duration_ms":90938,"temperature":0.7,"pith_summary":"This survey argues that music generation has advanced faster than its evaluation: there is still no comprehensive, standardized framework for judging generated music, and many widely used objective metrics correlate poorly with what human listeners actually prefer. To make the gap visible, the paper organizes existing evaluation methods into a taxonomy spanning audio and symbolic representations, reference-based and reference-free metrics, and human listening tests. Its critical review identifies three structural problems—objective metrics that lack interpretation and thresholds, a Western-centric bias in datasets and metrics, and a proliferation of ad hoc criteria that block cross-model comparison—and proposes a future framework in which quality, adherence to instruction, and similarity to reference are scored along interpretable dimensions.","feed_headline":"No standard yardstick yet for grading AI music","feed_subtitle":"A taxonomy of audio and symbolic metrics exposes weak ties to human perception and a Western bias.","key_machinery":"The organizing object is the taxonomy in Figure 2, which classifies surveyed metrics into human versus automatic evaluation; within automatic, reference-based versus reference-free; within each, audio versus symbolic representations; and, at the lowest level, by musical feature (pitch, chord, rhythm, structure, originality) and purpose (quality, similarity, adherence). The taxonomy does the argumentative work: by making the whole metric landscape visible at once, it exposes the three structural problems—weak correlation with human perception, Western-centric bias, and lack of standardization—that motivate the paper's proposed response. That response is a three-component scorer (quality/struc","core_discovery":"Music evaluation has not kept pace with generation: no comprehensive framework exists for judging generated scores, and common metrics often disagree with human perception. The paper substantiates this with a taxonomy (Figure 2) covering human versus automatic evaluation, reference-based versus reference-free metrics, and audio versus symbolic representations. Its critical review isolates load-bearing flaws: FAD, KLD, CLAP-score, and Overlapped Area rank models but lack thresholds and correlate poorly with listener preference; symbolic evaluation is fragmented and perceptually ungrounded; datasets and metrics are Western-centric; ad hoc criteria block comparison. The remedy is a three-compon","pith_inferences":["A direct test of the paper's core diagnosis: build a small, culturally diverse expert-rated set of generated songs and measure how well existing objective metrics rank them; if FAD/KLD/CLAP orderings disagree with expert ratings, the critical claim holds for that sample.","The proposed human-in-the-loop framework implies that future generation papers could report scores on interpretable dimensions (structure, coherence, expressiveness, prompt fidelity) rather than an overall mean, following the paper's own food-critic analogy.","If aesthetic predictors are trained to emulate expert raters across dimensions, they could serve as proxies in large-scale evaluation, but only when the underlying preference dataset includes non-Western genres and multi-dimensional, expert annotations.","The taxonomy suggests a missing piece: a shared evaluation registry where teams register their metrics, reference sets, and participant profiles, making standardization failures visible rather than accidental."],"forward_implications":["If the taxonomy is right, no single metric such as FAD or CLAP-score can stand alone as evidence of musical quality; models should be reported on a multi-dimensional battery.","Cross-model comparisons that rely on self-defined criteria are not comparable; the field needs shared reference sets, thresholds, and reporting formats.","Aesthetic predictors trained on current human preference datasets inherit biases toward Western, well-resourced genres, keeping low-resource genre evaluation unreliable until datasets diversify.","Symbolic music evaluation will continue to lag audio evaluation unless it adopts perceptually grounded, temporally sensitive features.","Standardized, expert-designed listening test criteria and diverse participant pools are needed for human evaluation to yield generalizable conclusions."],"supporting_citations":[{"why":"Supplies the foundational OA/KLD symbolic evaluation framework and the critique that evaluation lacks perceptual grounding.","marker":"(Yang and Lerch, 2020)"},{"why":"Introduces FAD, the widely used audio distribution distance whose Gaussian assumption and classifier dependence the paper criticizes.","marker":"(Kilgour et al., 2018)"},{"why":"Provides the empirical finding that CLAP-score, FAD, and KLD align poorly with human preferences, supporting the interpretation critique.","marker":"(Yuan et al., 2025)"},{"why":"Quantifies Western-centric bias in 152 dataset papers (5.7% non-Western), grounding the cross-cultural critique.","marker":"(Mehta et al., 2025)"},{"why":"Defines Audiobox Aesthetics, the aesthetic predictor representing the current trend the paper evaluates and finds misaligned with human preference datasets.","marker":"(Tjandra et al., 2025)"},{"why":"Contributes SongEval, a multi-dimensional expert-rated song benchmark that supports the call for interpretable dimensions.","marker":"(Yao et al., 2025)"},{"why":"Supplies the MAD metric and MusicPrefs dataset; MAD avoids Gaussian assumptions, showing progress toward perceptually aligned metrics.","marker":"(Huang et al., 2025)"},{"why":"Proposes KAD, using MMD without Gaussian assumptions, supporting the claim that newer metrics correlate better with human preference.","marker":"(Chung et al., 2025)"},{"why":"Introduces MusicLM, MusicCaps, and MuLan score, giving representative generation models and evaluation benchmarks used across the field.","marker":"(Agostinelli et al., 2023)"},{"why":"Provides MusicBench/Mustango, a benchmark and model with control-specific evaluation criteria, illustrating non-standardized symbolic criteria.","marker":"(Melechovsky et al., 2023)"}],"fun_headline_variants":["AI music is easy to make, hard to judge","Music AI outruns its own scorecard","Why AI music still lacks a fair scorecard","Evaluating AI music: metrics miss human ears"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The survey's conclusions stand on the assumption that the metrics it catalogues and the studies it cites are representative of the whole field; if major metrics or counter-evidence were missed, the critical review would lose its force.","fun_headline_variants_meta":{"raw":{"variants":["AI music is easy to make, hard to judge","Music AI outruns its own scorecard","Why AI music still lacks a fair scorecard","Evaluating AI music: metrics miss human ears"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000242,"raw_usage":{"total_tokens":1289,"prompt_tokens":599,"completion_tokens":690,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":343,"completion_tokens_details":{"reasoning_tokens":630}},"tokens_in":343,"tokens_out":690,"duration_ms":7367,"temperature":1.0,"reasoning_tokens":630,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T16:51:14.183140+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Settle the central gap claim by checking a fixed slice of recent music-generation papers (say, 50 papers from 2024–2025) for a shared evaluation protocol—same metric battery, same reference sets, same thresholds, same reporting format. If such a protocol is already in common use, the no-comprehensive-framework claim collapses; if every paper improvises its own criteria, the survey's diagnosis is confirmed.","supporting_citations":[],"review_version":1}