Pith. sign in

REVIEW 3 major objections 5 minor 15 references

A Survey on Evaluation Metrics for Music Generation

T0 review · 3 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Music evaluation has not kept pace with generation: no comprehensive framework exists for judging generated scores, and common metrics often disagree with human perception.

desk verdict Useful taxonomy, uneven evidence for the headline critique; worth a serious referee but needs the correlation claims backed by numbers. read the letter →

arxiv 2509.00051 v1 pith:7XODTNZ5 submitted 2025-08-24 cs.SD cs.MMeess.AS

classification cs.SDcs.MMeess.AS
keywords musicgenerationevaluationmetricstaxonomytext-to-musicsymbolicaudiohumancross-culturalbias
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that music generation has advanced faster than its evaluation: there is still no comprehensive, standardized framework for judging generated music, and many widely used objective metrics correlate poorly with what human listeners actually prefer. To make the gap visible, the paper organizes existing evaluation methods into a taxonomy spanning audio and symbolic representations, reference-based and reference-free metrics, and human listening tests. Its critical review identifies three structural problems—objective metrics that lack interpretation and thresholds, a Western-centric bias in datasets and metrics, and a proliferation of ad hoc criteria that block cross-model comparison—and proposes a future framework in which quality, adherence to instruction, and similarity to reference are scored along interpretable dimensions.

What carries the argument

The organizing object is the taxonomy in Figure 2, which classifies surveyed metrics into human versus automatic evaluation; within automatic, reference-based versus reference-free; within each, audio versus symbolic representations; and, at the lowest level, by musical feature (pitch, chord, rhythm, structure, originality) and purpose (quality, similarity, adherence). The taxonomy does the argumentative work: by making the whole metric landscape visible at once, it exposes the three structural problems—weak correlation with human perception, Western-centric bias, and lack of standardization—that motivate the paper's proposed response. That response is a three-component scorer (quality/struc

What would settle it

Settle the central gap claim by checking a fixed slice of recent music-generation papers (say, 50 papers from 2024–2025) for a shared evaluation protocol—same metric battery, same reference sets, same thresholds, same reporting format. If such a protocol is already in common use, the no-comprehensive-framework claim collapses; if every paper improvises its own criteria, the survey's diagnosis is confirmed.

Watch

Extended reading notes

Core claim

Music evaluation has not kept pace with generation: no comprehensive framework exists for judging generated scores, and common metrics often disagree with human perception. The paper substantiates this with a taxonomy (Figure 2) covering human versus automatic evaluation, reference-based versus reference-free metrics, and audio versus symbolic representations. Its critical review isolates load-bearing flaws: FAD, KLD, CLAP-score, and Overlapped Area rank models but lack thresholds and correlate poorly with listener preference; symbolic evaluation is fragmented and perceptually ungrounded; datasets and metrics are Western-centric; ad hoc criteria block comparison. The remedy is a three-compon

Load-bearing premise

The survey's conclusions stand on the assumption that the metrics it catalogues and the studies it cites are representative of the whole field; if major metrics or counter-evidence were missed, the critical review would lose its force.

Editorial extensions

If this is right

  • If the taxonomy is right, no single metric such as FAD or CLAP-score can stand alone as evidence of musical quality; models should be reported on a multi-dimensional battery.
  • Cross-model comparisons that rely on self-defined criteria are not comparable; the field needs shared reference sets, thresholds, and reporting formats.
  • Aesthetic predictors trained on current human preference datasets inherit biases toward Western, well-resourced genres, keeping low-resource genre evaluation unreliable until datasets diversify.
  • Symbolic music evaluation will continue to lag audio evaluation unless it adopts perceptually grounded, temporally sensitive features.
  • Standardized, expert-designed listening test criteria and diverse participant pools are needed for human evaluation to yield generalizable conclusions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the paper's core diagnosis: build a small, culturally diverse expert-rated set of generated songs and measure how well existing objective metrics rank them; if FAD/KLD/CLAP orderings disagree with expert ratings, the critical claim holds for that sample.
  • The proposed human-in-the-loop framework implies that future generation papers could report scores on interpretable dimensions (structure, coherence, expressiveness, prompt fidelity) rather than an overall mean, following the paper's own food-critic analogy.
  • If aesthetic predictors are trained to emulate expert raters across dimensions, they could serve as proxies in large-scale evaluation, but only when the underlying preference dataset includes non-Western genres and multi-dimensional, expert annotations.
  • The taxonomy suggests a missing piece: a shared evaluation registry where teams register their metrics, reference sets, and participant profiles, making standardization failures visible rather than accidental.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper surveys evaluation metrics for music generation, organizing them into a taxonomy (Figure 2) spanning reference-based and reference-free objective metrics for audio and symbolic music, human evaluation criteria, and benchmarks. It argues that current evaluation methodologies suffer from poor correlation between objective metrics and human perception, cross-cultural bias, and lack of standardization, and it proposes future directions including a human-in-the-loop RL-trained aesthetic predictor. Appendices provide metric definitions and toolkits.

Significance. If the taxonomy and critical review are accurate, the paper would be a useful reference for the music-generation community. The explicit organization of symbolic and audio metrics, the inclusion of recent human-preference datasets, and the attention to cross-cultural bias are strengths. The main value, however, rests on two load-bearing premises: that the surveyed metric set is representative, and that the empirical evidence for the 'poor correlation' limitation is reliable. These premises currently receive insufficient support, which limits the confidence one can place in the survey's central motivation.

major comments (3)
  1. [§4.1 and §3.1.1] The central critical claim that objective metrics such as CLAP-score, FAD, and KLD 'often align poorly with human preferences' is attributed to Yuan et al. (2025) with no correlation coefficients, task scope, or comparison baselines; the earlier assertion that KAD and MAD have 'shown better correlation with human preferences than FAD' is cited to no source at all. Because this claim motivates the entire review and the proposed framework, please report the quantitative evidence (and its domain/generalizability) or soften the claim to match the evidence.
  2. [Figure 2 and §3] The survey describes itself as a comprehensive overview, but it does not describe a systematic search protocol, inclusion/exclusion criteria, or coverage analysis. Without such a protocol, the representativeness of Figure 2 and the critical review cannot be assessed, and the reader cannot tell whether important metrics or benchmarks were omitted. Please add a methodology subsection detailing databases, years, search terms, screening, and the number of papers considered, or qualify the claims as representative rather than comprehensive.
  3. [§3.3 vs §4.1] Section 3.3 describes MusicPrefs as having crowdsourced pairwise ratings for fidelity and musicality, while Section 4.1 groups it with preference datasets that rely 'solely on overall impression' (Huang et al., 2025; Liu et al., 2025a). These statements are in tension. If MusicPrefs uses dimensional ratings, the limitation claim should be revised; if not, the Section 3.3 description should be corrected.
minor comments (5)
  1. [§2.3] 'With MIDI datasets being the most popular for example- Lakh MIDI Dataset (Raffel, 2016), Popular examples include' is a grammatical duplication; clean up.
  2. [§4.1] 'distuned' should be 'out-of-tune' or 'dissonant'.
  3. [§3.1.2] 'There is Figure 2 lists the metrics' is ungrammatical; remove 'There is'.
  4. [Appendix B, Eq. (3)] The formula for MOA is missing a closing parenthesis after b_i^(y).
  5. [Appendix A.3] 'Adherance' should be 'Adherence'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the survey is a self-contained review of external literature.

full rationale

This paper is a survey: it derives no metrics, fits no parameters, and makes no formal predictions from its own equations. Its taxonomy is an organizational categorization of metrics drawn from the cited literature, and its critical claims are attributed to external studies (e.g., Yuan et al., 2025; Mehta et al., 2025; Yao et al., 2025; Zhang et al., 2025). There are no instances where a definition presupposes a result the paper claims to establish, and no fitted input is relabeled as a prediction. The authors do not cite their own prior work, so there is no self-citation chain and no imported uniqueness theorem. Unsupported or thinly sourced assertions, such as the uncited statement that KAD and MAD correlate better with human preference than FAD, are evidence-quality concerns rather than circularity under the specified criteria. The central gap claim ('no comprehensive framework exists for evaluating generated music scores') is an empirical literature claim that does not reduce to the paper's own definitions. The survey's value may depend on the representativeness and accuracy of the cited evidence, but that is a correctness risk, not a circular derivation.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

As a survey, the paper introduces no free parameters or invented entities. Its central claim rests on two domain assumptions: that the surveyed literature is representative, and that its organizing taxonomy is a valid way to structure the field.

assumptions (2)
  • domain assumption The set of surveyed metrics is representative of the field, and the cited empirical findings about metric limitations are accurate and generalizable.
    The critical review in Section 4 relies on individual citations (e.g., Yuan et al. 2025, Zhang et al. 2025) for strong claims about objective metrics' poor correlation with human perception and aesthetic predictor biases. If these findings are not generalizable, the survey's conclusions weaken.
  • domain assumption The taxonomy's binary oppositions (audio vs symbolic, reference-based vs reference-free) capture the essential structure of music generation evaluation.
    The entire paper is organized around Figure 2's taxonomy; if this categorization is artificial or misses key dimensions, the survey's contribution is diminished.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on Evaluation Metrics for Music Generation." pith.science (2026). https://pith.science/paper/7XODTNZ5

@misc{pith2026250900051,
  author       = {Pith},
  title        = {Pith review of: A Survey on Evaluation Metrics for Music Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7XODTNZ5}},
  note         = {Machine review of arXiv:2509.00051}
}
read the original abstract

Despite significant advancements in music generation systems, the methodologies for evaluating generated music have not progressed as expected due to the complex nature of music, with aspects such as structure, coherence, creativity, and emotional expressiveness. In this paper, we shed light on this research gap, introducing a detailed taxonomy for evaluation metrics for both audio and symbolic music representations. We include a critical review identifying major limitations in current evaluation methodologies which includes poor correlation between objective metrics and human perception, cross-cultural bias, and lack of standardization that hinders cross-model comparisons. Addressing these gaps, we further propose future research directions towards building a comprehensive evaluation framework for music generation evaluation.

Figures

Figures reproduced from arXiv: 2509.00051 by the authors.

Figure 1
Figure 1. An illustration of music-related tasks with their corresponding evaluation metrics. Unlike other tasks, [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Music Generation Evaluation Taxonomy [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. A reliable Music Quality Scorer Model can elevate the current music generation evaluation scenario. [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

15 extracted references · 8 canonical work pages

  1. [6]

    In ISMIR, pages 348–354

    Jsymbolic 2.2: Extracting features from sym- bolic music for use in musicological and mir re- search. In ISMIR, pages 348–354. Atharva Mehta, Shivam Chauhan, Amirbek Djanibekov, Atharva Kulkarni, Gus Xia, and Monojit Choudhury. 2025. Music for all: Represen- tational bias and cross-cultural adaptability of music generation models. In Findings of the Assoc...

  2. [10]

    arXiv preprint arXiv:2302.04456

    Ernie-music: Text-to-waveform music gen- eration with diffusion models. arXiv preprint arXiv:2302.04456. A Human Evaluation A.1 Musical Structure according to Music Theory Structureness: If the music is structured nicely or not (Liu et al., 2022). More fine-grained structural aspects were used by (Yu et al., 2022). Short- term structure: Whether the gener...

  3. [11]

    Used Pitch Class (UPC) (Dong et al., 2018) is number of used pitch classes per bar

    measures the number of unique pitches and note durations within a sequence respetively. Used Pitch Class (UPC) (Dong et al., 2018) is number of used pitch classes per bar. Qualified Note (QN) (Dong et al., 2018) is the proportion of notes that are at least three time steps long (equivalent to a 32nd note or longer). This metric indicates whether the music...

  4. [12]

    Consecutive Pitch Repetitions (CPR)(Trieu and Keller, 2018) measures the frequency of occur- rences of some number of consecutive pitch rep- etitions

    extends (Dong et al., 2018)’s Qualified Note metric (which excluded notes shorter than a 32nd note) by measuring how often note durations match standard values (1, 1/2, 1/4, 1/8, 1/16) in- cluding dotted, triplet, and tied forms. Consecutive Pitch Repetitions (CPR)(Trieu and Keller, 2018) measures the frequency of occur- rences of some number of consecuti...

  5. [13]

    A higher CTD means there are more abrupt changes in the chord progression

    between each pair of adjacent chords in a sequence. A higher CTD means there are more abrupt changes in the chord progression. Chord Tone to Non-Chord Tone Ratio (CT- nCTR) (Yeh et al., 2021) is the ratio of notes that match the underlying chord (chord tones) to those that don’t (non-chord tones). A higher CTnCTR indicates that most notes fit well with th...

  6. [15]

    is computed by first measuring the propor- tion of correctly predicted attributes for each sam- ple, then averaging these values across the entire test set. Dynamics correlation (Wu et al., 2024) measures how well a generated audio score matches the dy- namic variations (smoothed frame wise loudness) of a reference performance by calculating Pear- son’s c...

  7. [1908]

    Hang Chu, Raquel Urtasun, and Sanja Fidler

    PMLR. Hang Chu, Raquel Urtasun, and Sanja Fidler

  8. [2006]

    In Proceedings of the 1st ACM workshop on Audio and music computing multimedia, pages 21–26

    Detecting harmonic change in musical audio. In Proceedings of the 1st ACM workshop on Audio and music computing multimedia, pages 21–26. Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck. 2018. Enabling factorized piano music modeling and gen- eration with the maestr...

Show all 15 references
  1. [2015]

    a good algorithm does not steal–it imitates

    Music pattern discovery with variable markov oracle: A unified approach to symbolic and audio representations. In ISMIR, pages 176–182. Yutian Wang, Wanyin Yang, Zhenrong Dai, Yilong Zhang, Kun Zhao, and Hui Wang. 2024. Melo- trans: A text to symbolic music generation model fo...

  2. [2016]

    arXiv preprint arXiv:1611.03477

    Song from pi: A musically plausible net- work for pop music generation. arXiv preprint arXiv:1611.03477. Ching-Hua Chuan and Dorien Herremans. 2018. Mod- eling temporal tonal relations in polyphonic music through deep networks with a novel image-based representation. In Procee...

  3. [2018]

    In 2018 IEEE 12th international conference on seman- tic computing (ICSC), pages 377–382

    Deepj: Style-specific music generation. In 2018 IEEE 12th international conference on seman- tic computing (ICSC), pages 377–382. IEEE. Cory McKay, Julie Cumming, and Ichiro Fujinaga

  4. [2020]

    H = − 12X i=1 pi log2 pi (1) where H is the Pitch Class Entropy

    To calculate pitch histogram entropy, we can create a 12-dimensional pitch class histogram with the notes that appear in a certain period of the music score and calculate the entropy of that histogram. H = − 12X i=1 pi log2 pi (1) where H is the Pitch Class Entropy. pi is the ...

  5. [2022]

    Ad- vances in neural information processing systems , 35:1376–1388

    Museformer: Transformer with fine-and coarse-grained attention for music generation. Ad- vances in neural information processing systems , 35:1376–1388. Dingyao Yu, Kaitao Song, Peiling Lu, Tianyu He, Xu Tan, Wei Ye, Shikun Zhang, and Jiang Bian

  6. [2023]

    arXiv preprint arXiv:2310.11954

    Musicagent: An ai agent for music under- standing and generation with large language models. arXiv preprint arXiv:2310.11954. R Yuan et al. 2024. Chatmusician: understanding and generating music intrinsically with llm, arxiv (cor- nell university)(2024). Ruibin Yuan, Hanfeng L...

  7. [2024]

    arXiv preprint arXiv:2402.00282

    Pam: Prompting audio-language models for audio quality assessment. arXiv preprint arXiv:2402.00282. Chris Donahue, Huanru Henry Mao, Yiting Ethan Li, Garrison W Cottrell, and Julian McAuley. 2019. Lakhnes: Improving multi-instrumental music gen- eration with cross-domain pre-t...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.