REVIEW 3 major objections 5 minor 15 references
A Survey on Evaluation Metrics for Music Generation
T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Music evaluation has not kept pace with generation: no comprehensive framework exists for judging generated scores, and common metrics often disagree with human perception.
desk verdict Useful taxonomy, uneven evidence for the headline critique; worth a serious referee but needs the correlation claims backed by numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing object is the taxonomy in Figure 2, which classifies surveyed metrics into human versus automatic evaluation; within automatic, reference-based versus reference-free; within each, audio versus symbolic representations; and, at the lowest level, by musical feature (pitch, chord, rhythm, structure, originality) and purpose (quality, similarity, adherence). The taxonomy does the argumentative work: by making the whole metric landscape visible at once, it exposes the three structural problems—weak correlation with human perception, Western-centric bias, and lack of standardization—that motivate the paper's proposed response. That response is a three-component scorer (quality/struc
What would settle it
Settle the central gap claim by checking a fixed slice of recent music-generation papers (say, 50 papers from 2024–2025) for a shared evaluation protocol—same metric battery, same reference sets, same thresholds, same reporting format. If such a protocol is already in common use, the no-comprehensive-framework claim collapses; if every paper improvises its own criteria, the survey's diagnosis is confirmed.
Extended reading notes
Core claim
Music evaluation has not kept pace with generation: no comprehensive framework exists for judging generated scores, and common metrics often disagree with human perception. The paper substantiates this with a taxonomy (Figure 2) covering human versus automatic evaluation, reference-based versus reference-free metrics, and audio versus symbolic representations. Its critical review isolates load-bearing flaws: FAD, KLD, CLAP-score, and Overlapped Area rank models but lack thresholds and correlate poorly with listener preference; symbolic evaluation is fragmented and perceptually ungrounded; datasets and metrics are Western-centric; ad hoc criteria block comparison. The remedy is a three-compon
Load-bearing premise
The survey's conclusions stand on the assumption that the metrics it catalogues and the studies it cites are representative of the whole field; if major metrics or counter-evidence were missed, the critical review would lose its force.
Editorial extensions
If this is right
- If the taxonomy is right, no single metric such as FAD or CLAP-score can stand alone as evidence of musical quality; models should be reported on a multi-dimensional battery.
- Cross-model comparisons that rely on self-defined criteria are not comparable; the field needs shared reference sets, thresholds, and reporting formats.
- Aesthetic predictors trained on current human preference datasets inherit biases toward Western, well-resourced genres, keeping low-resource genre evaluation unreliable until datasets diversify.
- Symbolic music evaluation will continue to lag audio evaluation unless it adopts perceptually grounded, temporally sensitive features.
- Standardized, expert-designed listening test criteria and diverse participant pools are needed for human evaluation to yield generalizable conclusions.
Reading between the lines
- A direct test of the paper's core diagnosis: build a small, culturally diverse expert-rated set of generated songs and measure how well existing objective metrics rank them; if FAD/KLD/CLAP orderings disagree with expert ratings, the critical claim holds for that sample.
- The proposed human-in-the-loop framework implies that future generation papers could report scores on interpretable dimensions (structure, coherence, expressiveness, prompt fidelity) rather than an overall mean, following the paper's own food-critic analogy.
- If aesthetic predictors are trained to emulate expert raters across dimensions, they could serve as proxies in large-scale evaluation, but only when the underlying preference dataset includes non-Western genres and multi-dimensional, expert annotations.
- The taxonomy suggests a missing piece: a shared evaluation registry where teams register their metrics, reference sets, and participant profiles, making standardization failures visible rather than accidental.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper surveys evaluation metrics for music generation, organizing them into a taxonomy (Figure 2) spanning reference-based and reference-free objective metrics for audio and symbolic music, human evaluation criteria, and benchmarks. It argues that current evaluation methodologies suffer from poor correlation between objective metrics and human perception, cross-cultural bias, and lack of standardization, and it proposes future directions including a human-in-the-loop RL-trained aesthetic predictor. Appendices provide metric definitions and toolkits.
Significance. If the taxonomy and critical review are accurate, the paper would be a useful reference for the music-generation community. The explicit organization of symbolic and audio metrics, the inclusion of recent human-preference datasets, and the attention to cross-cultural bias are strengths. The main value, however, rests on two load-bearing premises: that the surveyed metric set is representative, and that the empirical evidence for the 'poor correlation' limitation is reliable. These premises currently receive insufficient support, which limits the confidence one can place in the survey's central motivation.
major comments (3)
- [§4.1 and §3.1.1] The central critical claim that objective metrics such as CLAP-score, FAD, and KLD 'often align poorly with human preferences' is attributed to Yuan et al. (2025) with no correlation coefficients, task scope, or comparison baselines; the earlier assertion that KAD and MAD have 'shown better correlation with human preferences than FAD' is cited to no source at all. Because this claim motivates the entire review and the proposed framework, please report the quantitative evidence (and its domain/generalizability) or soften the claim to match the evidence.
- [Figure 2 and §3] The survey describes itself as a comprehensive overview, but it does not describe a systematic search protocol, inclusion/exclusion criteria, or coverage analysis. Without such a protocol, the representativeness of Figure 2 and the critical review cannot be assessed, and the reader cannot tell whether important metrics or benchmarks were omitted. Please add a methodology subsection detailing databases, years, search terms, screening, and the number of papers considered, or qualify the claims as representative rather than comprehensive.
- [§3.3 vs §4.1] Section 3.3 describes MusicPrefs as having crowdsourced pairwise ratings for fidelity and musicality, while Section 4.1 groups it with preference datasets that rely 'solely on overall impression' (Huang et al., 2025; Liu et al., 2025a). These statements are in tension. If MusicPrefs uses dimensional ratings, the limitation claim should be revised; if not, the Section 3.3 description should be corrected.
minor comments (5)
- [§2.3] 'With MIDI datasets being the most popular for example- Lakh MIDI Dataset (Raffel, 2016), Popular examples include' is a grammatical duplication; clean up.
- [§4.1] 'distuned' should be 'out-of-tune' or 'dissonant'.
- [§3.1.2] 'There is Figure 2 lists the metrics' is ungrammatical; remove 'There is'.
- [Appendix B, Eq. (3)] The formula for MOA is missing a closing parenthesis after b_i^(y).
- [Appendix A.3] 'Adherance' should be 'Adherence'.
Circularity Check
No significant circularity; the survey is a self-contained review of external literature.
full rationale
This paper is a survey: it derives no metrics, fits no parameters, and makes no formal predictions from its own equations. Its taxonomy is an organizational categorization of metrics drawn from the cited literature, and its critical claims are attributed to external studies (e.g., Yuan et al., 2025; Mehta et al., 2025; Yao et al., 2025; Zhang et al., 2025). There are no instances where a definition presupposes a result the paper claims to establish, and no fitted input is relabeled as a prediction. The authors do not cite their own prior work, so there is no self-citation chain and no imported uniqueness theorem. Unsupported or thinly sourced assertions, such as the uncited statement that KAD and MAD correlate better with human preference than FAD, are evidence-quality concerns rather than circularity under the specified criteria. The central gap claim ('no comprehensive framework exists for evaluating generated music scores') is an empirical literature claim that does not reduce to the paper's own definitions. The survey's value may depend on the representativeness and accuracy of the cited evidence, but that is a correctness risk, not a circular derivation.
Assumptions & free parameters
assumptions (2)
- domain assumption The set of surveyed metrics is representative of the field, and the cited empirical findings about metric limitations are accurate and generalizable.
- domain assumption The taxonomy's binary oppositions (audio vs symbolic, reference-based vs reference-free) capture the essential structure of music generation evaluation.
Cite this review
Pith. "Pith review of A Survey on Evaluation Metrics for Music Generation." pith.science (2026). https://pith.science/paper/7XODTNZ5
@misc{pith2026250900051,
author = {Pith},
title = {Pith review of: A Survey on Evaluation Metrics for Music Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/7XODTNZ5}},
note = {Machine review of arXiv:2509.00051}
}
read the original abstract
Despite significant advancements in music generation systems, the methodologies for evaluating generated music have not progressed as expected due to the complex nature of music, with aspects such as structure, coherence, creativity, and emotional expressiveness. In this paper, we shed light on this research gap, introducing a detailed taxonomy for evaluation metrics for both audio and symbolic music representations. We include a critical review identifying major limitations in current evaluation methodologies which includes poor correlation between objective metrics and human perception, cross-cultural bias, and lack of standardization that hinders cross-model comparisons. Addressing these gaps, we further propose future research directions towards building a comprehensive evaluation framework for music generation evaluation.
Figures
Reference graph
Works this paper leans on
-
[6]
Jsymbolic 2.2: Extracting features from sym- bolic music for use in musicological and mir re- search. In ISMIR, pages 348–354. Atharva Mehta, Shivam Chauhan, Amirbek Djanibekov, Atharva Kulkarni, Gus Xia, and Monojit Choudhury. 2025. Music for all: Represen- tational bias and cross-cultural adaptability of music generation models. In Findings of the Assoc...
arXiv 2025
-
[10]
arXiv preprint arXiv:2302.04456
Ernie-music: Text-to-waveform music gen- eration with diffusion models. arXiv preprint arXiv:2302.04456. A Human Evaluation A.1 Musical Structure according to Music Theory Structureness: If the music is structured nicely or not (Liu et al., 2022). More fine-grained structural aspects were used by (Yu et al., 2022). Short- term structure: Whether the gener...
arXiv 2022
-
[11]
Used Pitch Class (UPC) (Dong et al., 2018) is number of used pitch classes per bar
measures the number of unique pitches and note durations within a sequence respetively. Used Pitch Class (UPC) (Dong et al., 2018) is number of used pitch classes per bar. Qualified Note (QN) (Dong et al., 2018) is the proportion of notes that are at least three time steps long (equivalent to a 32nd note or longer). This metric indicates whether the music...
work page 2018
-
[12]
extends (Dong et al., 2018)’s Qualified Note metric (which excluded notes shorter than a 32nd note) by measuring how often note durations match standard values (1, 1/2, 1/4, 1/8, 1/16) in- cluding dotted, triplet, and tied forms. Consecutive Pitch Repetitions (CPR)(Trieu and Keller, 2018) measures the frequency of occur- rences of some number of consecuti...
work page 2018
-
[13]
A higher CTD means there are more abrupt changes in the chord progression
between each pair of adjacent chords in a sequence. A higher CTD means there are more abrupt changes in the chord progression. Chord Tone to Non-Chord Tone Ratio (CT- nCTR) (Yeh et al., 2021) is the ratio of notes that match the underlying chord (chord tones) to those that don’t (non-chord tones). A higher CTnCTR indicates that most notes fit well with th...
work page 2021
-
[15]
is computed by first measuring the propor- tion of correctly predicted attributes for each sam- ple, then averaging these values across the entire test set. Dynamics correlation (Wu et al., 2024) measures how well a generated audio score matches the dy- namic variations (smoothed frame wise loudness) of a reference performance by calculating Pear- son’s c...
work page 2024
-
[1908]
Hang Chu, Raquel Urtasun, and Sanja Fidler
PMLR. Hang Chu, Raquel Urtasun, and Sanja Fidler
-
[2006]
In Proceedings of the 1st ACM workshop on Audio and music computing multimedia, pages 21–26
Detecting harmonic change in musical audio. In Proceedings of the 1st ACM workshop on Audio and music computing multimedia, pages 21–26. Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck. 2018. Enabling factorized piano music modeling and gen- eration with the maestr...
arXiv 2018
Show all 15 references
-
[2015]
a good algorithm does not steal–it imitates
Music pattern discovery with variable markov oracle: A unified approach to symbolic and audio representations. In ISMIR, pages 176–182. Yutian Wang, Wanyin Yang, Zhenrong Dai, Yilong Zhang, Kun Zhao, and Hui Wang. 2024. Melo- trans: A text to symbolic music generation model fo...
2024 arXiv
-
[2016]
arXiv preprint arXiv:1611.03477
Song from pi: A musically plausible net- work for pop music generation. arXiv preprint arXiv:1611.03477. Ching-Hua Chuan and Dorien Herremans. 2018. Mod- eling temporal tonal relations in polyphonic music through deep networks with a novel image-based representation. In Procee...
2018 arXiv
-
[2018]
In 2018 IEEE 12th international conference on seman- tic computing (ICSC), pages 377–382
Deepj: Style-specific music generation. In 2018 IEEE 12th international conference on seman- tic computing (ICSC), pages 377–382. IEEE. Cory McKay, Julie Cumming, and Ichiro Fujinaga
2018
-
[2020]
H = − 12X i=1 pi log2 pi (1) where H is the Pitch Class Entropy
To calculate pitch histogram entropy, we can create a 12-dimensional pitch class histogram with the notes that appear in a certain period of the music score and calculate the entropy of that histogram. H = − 12X i=1 pi log2 pi (1) where H is the Pitch Class Entropy. pi is the ...
2021
-
[2022]
Ad- vances in neural information processing systems , 35:1376–1388
Museformer: Transformer with fine-and coarse-grained attention for music generation. Ad- vances in neural information processing systems , 35:1376–1388. Dingyao Yu, Kaitao Song, Peiling Lu, Tianyu He, Xu Tan, Wei Ye, Shikun Zhang, and Jiang Bian
-
[2023]
arXiv preprint arXiv:2310.11954
Musicagent: An ai agent for music under- standing and generation with large language models. arXiv preprint arXiv:2310.11954. R Yuan et al. 2024. Chatmusician: understanding and generating music intrinsically with llm, arxiv (cor- nell university)(2024). Ruibin Yuan, Hanfeng L...
2024 arXiv
-
[2024]
arXiv preprint arXiv:2402.00282
Pam: Prompting audio-language models for audio quality assessment. arXiv preprint arXiv:2402.00282. Chris Donahue, Huanru Henry Mao, Yiting Ethan Li, Garrison W Cottrell, and Julian McAuley. 2019. Lakhnes: Improving multi-instrumental music gen- eration with cross-domain pre-t...
2019 arXiv
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.