Pith. sign in

REVIEW 5 major objections 10 minor 43 references

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems

T0 review · 5 major / 10 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Automatic metrics disagree so strongly with human preference that text-to-music model rankings depend on the chosen evaluation.

desk verdict The headline inconsistency is plausible but under-supported because one side of the comparison is a pretrained aesthetics predictor, not fresh human ratings; still, the paper's empirical scaffold is worth engaging. read the letter →

arxiv 2504.21815 v1 pith:CPPOWC4H submitted 2025-04-30 eess.AS

classification eess.AS
keywords Text-to-MusicGenerationHumanPreferenceAlignmentEvaluationMetricsMusicQualityAssessmentGenerativeAudioModelsMADKADAestheticpredictors
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper studies whether automatic evaluation metrics can stand in for human preference when judging text-to-music systems. Using five recent generation models, it compares a learned aesthetics predictor (scoring content enjoyment, usefulness, production complexity, and production quality) against human pairwise judgments, and against reference-based distribution distances (MAD and KAD). The results show poor agreement: the best aesthetics dimension predicts human choices only about 62 percent of the time, with Spearman correlations below 0.26, and the aesthetics leaderboard disagrees with the distribution-metric leaderboard. The paper concludes that current evaluation practice is limited and calls for human-centered evaluation, releasing a benchmark dataset of generated samples to support it.

What carries the argument

The central machinery is a set of three evaluation perspectives applied to the same five text-to-music systems: (1) pairwise human preference judgments from the MusicPref dataset, treated as ground truth; (2) the AudioBox-Aesthetics neural predictor, which outputs scalar scores for content enjoyment, content usefulness, production complexity, and production quality; and (3) reference-based distribution distances, MAD and KAD, computed on audio embeddings (PANNs features) between generated sets and the human-composed LP-MusicCaps reference. MAD quantifies divergence using a KL-style measure, KAD uses a kernelized maximum mean discrepancy style distance. Comparing these perspectives on the same outputs is what exposes the inconsistency.

What would settle it

Collect fresh human preference ratings for a random sample of the released benchmark clips, then compare the resulting model ranking with the rankings from AudioBox-Aesthetics scores and from MAD/KAD; if the automatic metrics reproduce the human ranking or agree strongly with each other on the same clips, the paper's central claim of inconsistency would not hold.

Watch

Extended reading notes

Core claim

The paper's central discovery is that different evaluation perspectives give different answers about which text-to-music model is best, and none of the automatic proxies closely tracks human preference. On MusicPref pairwise comparisons, the AudioBox-Aesthetics model's score differences agree with human preferences at best around 62.3% for musicality (Content Usefulness) and 59.6% for fidelity (Production Quality), with Spearman correlations at best 0.258, far from a reliable predictor. The aesthetics-based leaderboard ranks JASCO highest on content usefulness and production quality, while reference-based MAD and KAD rank MusicGen-Large and JASCO as closest to human-composed MusicCaps recordings, with DiffRhythm, Stable-Audio-Open, and YuE farther away. The paper interprets these inconsistencies as evidence that no single learned or distributional metric currently captures human aesthetic judgment for music generation.

Load-bearing premise

The claim that automatic metrics diverge from human preference depends on treating the AudioBox-Aesthetics predictor's scores as a valid measure of human aesthetic judgment for the MusicPref pairs and the newly generated benchmark clips; the paper does not test that predictor against fresh human ratings of those exact outputs.

Editorial extensions

If this is right

  • Model rankings from text-to-music evaluations are metric-dependent; changing from aesthetics scores to MAD or KAD can change which system looks best.
  • Learned aesthetics scores cannot currently replace human listening tests: at roughly 62% best-case accuracy and Spearman correlations below 0.26, they capture only a weak signal of human pairwise preference.
  • Reference-based distribution metrics and aesthetic predictors measure different properties and should be reported side by side rather than treated as interchangeable.
  • The released benchmark of generated clips and evaluation scores gives other researchers a fixed corpus for testing new metrics against the same text-to-music outputs.
  • Because aesthetics scores vary with musical content (for example, rhythmic and electronic tags score higher on production complexity), model comparisons should control for genre and prompt content.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the AudioBox-Aesthetics predictor does not transfer to the MusicPref pairs and the newly generated clips (the paper does not validate it against fresh human ratings of these exact outputs), the reported agreement numbers may reflect cross-corpus transfer rather than intrinsic human-predictor agreement; a listening test on the released benchmark would separate the two.
  • The genre-dependent patterning of production-complexity scores suggests the aesthetics model may encode training-data or annotator biases; a testable extension is to re-rank models within matched genre clusters.
  • The same three-perspective comparison could be applied to other generative domains, where learned preference models and distribution metrics are also used as substitutes for human judgment without cross-validation against fresh human ratings.
  • If the weak agreement holds generally, current preference-optimized text-to-music systems may be optimizing a reward-model proxy that real listeners would only partly endorse; measuring the reward model's agreement with fresh human preference on generated outputs would show how much optimization signal survives.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 10 minor

Summary. The paper investigates how different evaluation procedures rank text-to-music systems, using five recent models (JASCO, Stable-Audio-Open, MusicGen-Large, YuE, DiffRhythm) and LP-MusicCaps prompts. It compares pairwise human preference labels from MusicPref against the Meta AudioBox-Aesthetics predictor across four dimensions (CE, CU, PC, PQ), reports accuracy and Spearman correlations in Section 3, and then uses the same predictor to produce a leaderboard in Section 4, including a tag-cluster bias analysis. Section 5 computes MAD and KAD distributional scores against MusicCaps ground truth. The paper reports that the rankings differ sharply across these evaluation perspectives, concluding that current evaluation practice has significant inconsistencies and advocating for more human-centered evaluation.

Significance. If substantiated, the descriptive pattern is a useful cautionary result for the audio generation community: model rankings differ depending on whether one uses learned aesthetic scores, MAD, or KAD, and the paper provides a public benchmark of generated samples. The cross-metric comparison and the tag-cluster analysis are constructive and go beyond a single evaluation axis. However, the central inference about human preferences rests on treating a pretrained predictor as a human surrogate, and the headline claim of 'significant inconsistencies' is not backed by significance tests or confidence intervals. The paper is therefore a suggestive comparative study rather than a rigorous demonstration of misalignment with human preference.

major comments (5)
  1. [Abstract and §3, Tables 2–3] The abstract claims 'significant inconsistencies' across metrics, but the paper reports no significance tests, confidence intervals, or effect sizes for the accuracies and Spearman correlations. For example, with 2,049 non-tie musicality pairs, the 62.3% accuracy for CU is roughly 12 percentage points above the 50% random baseline, so deviation from chance is likely real, but the differences among CE (0.604), CU (0.623), and PQ (0.596) need paired tests or confidence intervals before the paper can claim that the metrics are inconsistent. The negative correlations for PC (-0.002 for musicality, -0.087 for fidelity) also need an explicit test against zero and against the other metrics.
  2. [§3 and §4, Table 4] The paper labels Section 3 'Human vs. Human' and the abstract frames the findings as revealing inconsistencies with human preference, but no fresh human ratings are collected on the exact test stimuli. The Section 3 comparison is between MusicPref human pairwise labels and the pretrained Meta AudioBox-Aesthetics predictor, not between two human annotation sources. Section 4 then uses the same predictor as the judge for the leaderboard, so Table 4 is a proxy ranking unless the predictor is validated on those exact generated clips. The AudioBox-Aesthetics training corpus, described as one-third music, is not shown to cover MusicPref clips or the 10–95 second TTM outputs, so the observed poor agreement could be cross-corpus transfer error rather than evidence about human-human or human-machine disagreement.
  3. [§4.2] The dataset description is internally inconsistent. The paper says it uses the full LP-MusicCaps prompt set of 5,521 prompts, but then states that DiffRhythm and YuE use only 50 generated lyrics and that JASCO uses 50 drum tracks and 100 chord progressions. It is not clear how many clips were generated per model, whether all models saw the same prompt set, or how the conditioning inputs were paired with prompts. Without this information, Table 4 and the released benchmark cannot be reproduced or meaningfully interpreted as a controlled comparison.
  4. [§4.3 and §5] The model outputs differ in duration (JASCO 10 seconds, MusicGen-Large 20 seconds, Stable-Audio-Open 47 seconds, YuE around 50 seconds, DiffRhythm 95 seconds, with MusicCaps ground truth at 10 seconds), and the models differ in conditioning inputs such as lyrics, chords, and drum tracks. Because both the aesthetics predictor and the PANNs-wavegram-logmel embeddings used for MAD/KAD can be sensitive to duration and content composition, the cross-model comparisons in Table 4 and Figures 2–3 may partly reflect input differences rather than intrinsic generation quality. The paper acknowledges the conditioning confound in one sentence, but offers no control, sensitivity analysis, or duration-matched comparison.
  5. [§5] The central claim of inconsistency between human preference and distributional metrics is never tested on the same items. Section 5 contains no human judgments at all; it compares MAD/KAD values between generated sets and MusicCaps ground truth, while the 'human preference' side comes from the AudioBox-Aesthetics predictor in Section 4. The mismatch between the two rankings is therefore inferred across datasets rather than measured on the same clips. Additionally, the KAD sentence is ambiguous: 'MusicGen-Large and JASCO achieve the lowest distances relative to MusicCaps GT (7.65 and 5.51)' does not state which model has which value; the surrounding text and Figure 3 suggest JASCO has 5.51 and MusicGen-Large has 7.65, and this should be stated explicitly.
minor comments (10)
  1. [Abstract] The word 'significant' in the abstract should be replaced with a precise statistical statement or qualified with confidence intervals, since no significance tests are reported.
  2. [§4.3] There are typos in this section: 'synthezised' should be 'synthesized', 'aethetics' should be 'aesthetics', and 'teh' in §4.1 should be 'the'.
  3. [Table 4] The model name is written 'YUE' in the table but 'Yue' and 'YuE' elsewhere; please standardize the capitalization.
  4. [§2.2 and §3] The dataset name appears as 'MusicPrefs' in §2.2 but as 'MusicPref' in Section 3 and in reference [12]; the spelling should be made consistent.
  5. [Table 1] The row for 'Meta-AudioBox-Aesthetics' marks 'Human Involvement' with a check mark, but this is a pretrained automatic predictor, not a human annotation process; the table should distinguish direct human involvement from models trained on human annotations.
  6. [§4.3] The table caption '(no low)*' is not explained in the caption itself, and the text says the subset removes recordings 'under the low quality tag or captions' while the caption says it removes recordings with the 'low quality' tag; please clarify the exact filtering criterion.
  7. [§1 and Conclusion] The claim of being 'the first systematic study of human preference alignment in music generation' is too strong given existing work such as MusicEval and the KAD/MAD papers; please soften the novelty claim.
  8. [§4.4] The number of KMeans clusters (15) appears hand-chosen, and no robustness check is reported for the tag clustering; a short sensitivity analysis or a reference to the choice would help.
  9. [Throughout] The paper states that a benchmark dataset is released, but no URL, repository link, license, or download instructions are provided; a data-availability statement is needed.
  10. [Introduction] The sentence 'The work does not relate to Huy Phan's work at Meta' appears in the main text; this kind of statement belongs in an acknowledgments or conflict-of-interest section, not in the introduction.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper measures agreement among independent external metrics and human labels; no fitted parameter is renamed as a prediction.

full rationale

The claimed result is an empirical comparison of independent external resources, not a derivation from an internal fit. Section 3 compares MusicPref pairwise human labels (Huang et al., [12]) with the Meta AudioBox-Aesthetics predictor ([25]); neither resource is constructed from the other, and the paper explicitly separates 'human annotated pairs' from 'learnt preference scores from neural network.' The low accuracy (62.3% best CU) is measured, not forced. Section 4 reports pretrained aesthetic scores on generated clips; the paper does not fit any parameter to these clips and then call the fit a prediction. Section 5 computes MAD and KAD using the original implementations and PANNs embeddings [37]; these are external metrics. No equation in the paper turns an input into the claimed output by definition. There are self-citations ([14]-[18], [26], [28], [30]) but they appear as examples or background motivation in the related-work section, not as load-bearing premises of the three experiments; none invokes a uniqueness theorem or an ansatz. The substantive weakness—using the aesthetics predictor as a surrogate for human judgment without fresh validation on the target outputs—is a validity/correctness concern, not a circularity concern, because the predictor's output is not equivalent to the paper's conclusion by construction. Therefore no circularity step can be exhibited, and the score is 0.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claim does not rest on fitted parameters in the paper itself, but it does rest on transfer assumptions for the pretrained aesthetics predictor, the MusicPref labels, and the reference-based metrics. The conditioning input sizes are hand-chosen and affect comparability.

free parameters (2)
  • Hand-chosen conditioning set sizes for generated benchmark = 50 lyrics (GPT-4o), 50 drum tracks (The Drum Tamer), 100 chord progressions (GPT-4o)
    Section 4.2 fixes these sizes for lyric-conditioned and chord/drum-conditioned models; results for DiffRhythm, YuE, and JASCO depend on this small hand-selected set, so the comparison may not reflect the full prompt distribution.
  • Number of KMeans clusters for tag analysis = 15
    Section 4.4 fixes 15 semantic clusters for interpretability; the observed content bias pattern can depend on this choice.
assumptions (4)
  • domain assumption Meta-AudioBox-Aesthetics predictor scores are valid proxies for human aesthetic judgment when applied to MusicPref pairs and to generated TTM outputs.
    Sections 3 and 4 use CE/CU/PC/PQ scores as the human side of the comparison; without transfer validity, low agreement and leaderboard results carry no direct human-preference meaning.
  • domain assumption MusicPref pairwise human judgments correctly capture the preference target for musicality and fidelity.
    Section 3 uses 2,515 pairs (non-tie subsets of 2,049 and 1,889) as ground truth; the paper imports this dataset from [12] without re-validation.
  • domain assumption MAD and KAD computed in the PANNs-wavegram-logmel embedding space reflect perceptual/distributional closeness to human-composed music.
    Section 5 treats lower MAD/KAD as better distributional alignment; this assumes the embeddings and kernel choices preserve human-relevant audio similarity.
  • domain assumption Generated outputs from the five models are comparable despite different durations, conditioning modalities, and prompt subsets.
    Section 4.3 compares means and standard deviations across models with outputs ranging from 10 to 95 seconds and different additional inputs (lyrics, drum tracks, chord progressions); comparability is assumed rather than controlled.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems." pith.science (2026). https://pith.science/paper/CPPOWC4H

@misc{pith2026250421815,
  author       = {Pith},
  title        = {Pith review of: From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CPPOWC4H}},
  note         = {Machine review of arXiv:2504.21815}
}
read the original abstract

Evaluating generative models remains a fundamental challenge, particularly when the goal is to reflect human preferences. In this paper, we use music generation as a case study to investigate the gap between automatic evaluation metrics and human preferences. We conduct comparative experiments across five state-of-the-art music generation approaches, assessing both perceptual quality and distributional similarity to human-composed music. Specifically, we evaluate synthesis music from various perceptual dimensions and examine reference-based metrics such as Mauve Audio Divergence (MAD) and Kernel Audio Distance (KAD). Our findings reveal significant inconsistencies across the different metrics, highlighting the limitation of the current evaluation practice. To support further research, we release a benchmark dataset comprising samples from multiple models. This study provides a broader perspective on the alignment of human preference in generative modeling, advocating for more human-centered evaluation strategies across domains.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

43 extracted references · 29 canonical work pages

  1. [1]

    INTRODUCTION With the rapid advances in generative models, a fundamental question arises: How well do these models truly perform, and how can we evaluate them more thoroughly and systematically? While early efforts often relied on proxy objectives or handcrafted metrics [1, 2, 3], the growing complexity of generative systems demands more human- centered a...

  2. [2]

    From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems

    RELA TED WORKS In this work, we propose a categorization of performance evaluations for music generation systems based on three dimensions: objectives, human involvement, and modality, as summarized in Table 1. Since our focus is on aligning music generation with human preferences, this section is organized into objective and subjective evaluations, depen...

  3. [3]

    HUMAN VS. HUMAN: DO MUSICPREF ANNOTA TIONS AGREE WITH AESTHETICS SCORES? Can different sources of human annotation reach a common ground? In this experiment, our aim is to understand the relationship be- tween human-annotated pairwise preferences and the independently trained aesthetics predictor, conducting a detailed comparative analy- sis across their ...

  4. [4]

    low quality

    TTM LEADERBOARD: WHO DOES BEST, AND UNDER WHOSE JUDGE? In this section, we benchmark five recent text-to-music (TTM) gen- eration models by evaluating their outputs on a standardized set of prompts, and reporting the four AudioBox-aesthetics metrics over synthezised music clips. 4.1. Models • JASCO: A large-scale model optimized for musical generation wit...

  5. [5]

    We examine reference-based evaluation metrics on the generated dataset computed in Section 4, offering a complementary perspective to subjective assessments

    DIVERSITY OR DRIFT? DA TASET-LEVEL ALIGNMENT USING REFERENCE-BASED SCORES While the previous experiments centered on aesthetics scores and pairwise human preference alignment, this section investigates the distributional alignment among each generated music sets and human- composed set. We examine reference-based evaluation metrics on the generated datase...

  6. [6]

    Our results reveal substantial incon- sistencies between different evaluation perspectives, highlighting the challenges of fully capturing human judgment through automated proxies

    CONCLUSION AND FUTURE WORK In this work, we presented a cross-referenced discussion of text-to- music generation evaluations. Our results reveal substantial incon- sistencies between different evaluation perspectives, highlighting the challenges of fully capturing human judgment through automated proxies. By benchmarking five recent models and releasing t...

  7. [7]

    Acoustic scene generation with condi- tional SampleRNN,

    Qiuqiang Kong, Yong Xu, Turab Iqbal, Yin Cao, Wenwu Wang, and Mark D. Plumbley, “Acoustic scene generation with condi- tional SampleRNN,” in 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2019, pp. 925–929. Fig. 3 . Pairwise KAD scores between the evaluated models and the MusicCaps ground-truth dataset. Lower score...

  8. [8]

    AudioGen: Textually guided audio genera- tion,

    Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre D´efossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi, “AudioGen: Textually guided audio genera- tion,” in The Eleventh International Conference on Learning Representations, 2023

Show all 43 references
  1. [9]

    DExter: Learning and Controlling Performance Expression with Diffusion Models,

    Huan Zhang, Shreyan Chowdhury, Carlos Eduardo Cancino- Chac´on, Jinhua Liang, Simon Dixon, and Gerhard Widmer, “DExter: Learning and Controlling Performance Expression with Diffusion Models,” Applied Sciences, vol. 14, no. 15, 2024

  2. [10]

    DeepSeek-R1: Incentivizing reasoning capa- bility in llms via reinforcement learning,

    DeepSeek-AI, “DeepSeek-R1: Incentivizing reasoning capa- bility in llms via reinforcement learning,” arXiv:2501.12948, 2025

  3. [11]

    Qwen2.5 technical report,

    Qwen Team, “Qwen2.5 technical report,” arXiv:2412.15115, 2025

  4. [12]

    DSPO: Direct score preference optimization for diffusion model align- ment,

    Huaisheng Zhu, Teng Xiao, and Vasant G Honavar, “DSPO: Direct score preference optimization for diffusion model align- ment,” in The Thirteenth International Conference on Learning Representations, 2025

  5. [13]

    Seed-TTS: A family of high-quality versatile speech generation models,

    Seed Team, “Seed-TTS: A family of high-quality versatile speech generation models,” arXiv:2406.02430, 2024

  6. [14]

    1, Association for Computing Machinery, 2024

    Navonil Majumder, Chia-Yu Hung, Deepanway Ghosal, Wei- Ning Hsu, Rada Mihalcea, and Soujanya Poria, Tango 2: Align- ing Diffusion-based Text-to-Audio Generations through Direct Preference Optimization, vol. 1, Association for Computing Machinery, 2024

  7. [15]

    BATON: Aligning text-to-audio model using human preference feedback,

    Huan Liao, Haonan Han, Kai Yang, Tianjiao Du, Rui Yang, Qinmei Xu, Zunnan Xu, Jingquan Liu, Jiasheng Lu, and Xiu Li, “BATON: Aligning text-to-audio model using human preference feedback,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intellige...

  8. [16]

    DRAGON: Distributional rewards optimize diffusion generative models,

    Yatong Bai, Jonah Casebeer, Somayeh Sojoudi, and Nicholas J. Bryan, “DRAGON: Distributional rewards optimize diffusion generative models,” arXiv:2504.15217, 2025

  9. [17]

    SMART: Tuning a symbolic music generation system with an audio do- main aesthetic reward,

    Nicolas Jonason, Luca Casini, and Bob L. T. Sturm, “SMART: Tuning a symbolic music generation system with an audio do- main aesthetic reward,” arXiv:2504.16839, 2025

  10. [18]

    Aligning text-to-music evaluation with human pref- erences,

    Yichen Huang, Zachary Novack, Koichi Saito, Jiatong Shi, Shinji Watanabe, Yuki Mitsufuji, John Thickstun, and Chris Donahue, “Aligning text-to-music evaluation with human pref- erences,” arXiv:2503.16669, 2025

  11. [19]

    KAD: No more FAD! An effective and efficient evaluation metric for audio generation,

    Yoonjin Chung, Pilsun Eu, Junwon Lee, Keunwoo Choi, Juhan Nam, and Ben Sangbae Chon, “KAD: No more FAD! An effective and efficient evaluation metric for audio generation,” arXiv:2502.15602, 2025

  12. [20]

    WavCraft: Audio editing and generation with large language models,

    Jinhua Liang, Huan Zhang, Haohe Liu, and Yin Cao, “WavCraft: Audio editing and generation with large language models,” in ICLR 2024 Workshop on Large Language Model (LLM) Agents, 2024

  13. [21]

    WavJourney: Compositional audio creation with large language models,

    Xubo Liu, Zhongkai Zhu, Haohe Liu, Yi Yuan, Meng Cui, Qiushi Huang, Jinhua Liang, Yin Cao, Qiuqiang Kong, Mark D Plumbley, et al., “WavJourney: Compositional audio creation with large language models,” IEEE Transactions on Audio, Speech and Language Processing, 2025

  14. [22]

    Bridging paintings and music – exploring emotion based music genera- tion through paintings,

    Tanisha Hisariya, Huan Zhang, and Jinhua Liang, “Bridging paintings and music – exploring emotion based music genera- tion through paintings,” arXiv:2409.07827, 2024

  15. [23]

    Hierarchical symbolic pop music generation with graph neural networks,

    Wen Qing Lim, Jinhua Liang, and Huan Zhang, “Hierarchical symbolic pop music generation with graph neural networks,” arXiv:2409.08155, 2024

  16. [24]

    Render- box: Expressive performance rendering with text control,

    Huan Zhang, Akira Maezawa, and Simon Dixon, “Render- box: Expressive performance rendering with text control,” arXiv:2502.07711, 2025

  17. [25]

    Leveraging pre-trained audioldm for sound generation: A benchmark study,

    Yi Yuan, Haohe Liu, Jinhua Liang, Xubo Liu, Mark D. Plumb- ley, and Wenwu Wang, “Leveraging pre-trained audioldm for sound generation: A benchmark study,” in 2023 31st European Signal Processing Conference (EUSIPCO), 2023, pp. 765–769

  18. [26]

    Diffsound: Discrete diffu- sion model for text-to-sound generation,

    Dongchao Yang, Jianwei Yu, Helin Wang, Wen Wang, Chao Weng, Yuexian Zou, and Dong Yu, “Diffsound: Discrete diffu- sion model for text-to-sound generation,” IEEE/ACM Transac- tions on Audio, Speech, and Language Processing, vol. 31, pp. 1720–1733, 2023

  19. [27]

    Adapting frechet audio distance for generative music evaluation,

    Azalea Gui, Hannes Gamper, Sebastian Braun, and Dimitra Em- manouilidou, “Adapting frechet audio distance for generative music evaluation,” in 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024

  20. [28]

    Zero-shot unsupervised and text-based audio editing using DDPM inversion,

    Hila Manor and Tomer Michaeli, “Zero-shot unsupervised and text-based audio editing using DDPM inversion,” inICML 2024 Workshop on Structured Probabilistic Inference & Generative Modeling, 2024

  21. [29]

    AudioMorphix: Training-free audio editing with diffu- sion probabilistic models,

    Jinhua Liang, Yi Yuan, Dongya Jia, Xiaobin Zhuang, Zhengxi Liu, Yuanzhe Chen, Zhuo Chen, Yuping Wang, and Yuxuan Wang, “AudioMorphix: Training-free audio editing with diffu- sion probabilistic models,” arxiv, 2024

  22. [30]

    A comparison of deep learning MOS pre- dictors for speech synthesis quality,

    Alessandro Ragano, Emmanouil Benetos, Michael Chinen, Helard Becerra Martinez, Chandan K A Reddy, Jan Skoglund, and Andrew Hines, “A comparison of deep learning MOS pre- dictors for speech synthesis quality,” in 2023 34th Irish Signals and Systems Conference (ISSC), 2023, pp. 1–6

  23. [31]

    Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound,

    Andros Tjandra, Yi-Chiao Wu, Baishan Guo, John Hoffman, Brian Ellis, Apoorv Vyas, Bowen Shi, Sanyuan Chen, Matt Le, Nick Zacharov, Carleigh Wood, Ann Lee, and Wei-Ning Hsu, “Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound,” arXiv:250...

  24. [32]

    From Audio Encoders to Piano Judges: Benchmarking Performance Under- standing for Solo Piano,

    Huan Zhang, Jinhua Liang, and Simon Dixon, “From Audio Encoders to Piano Judges: Benchmarking Performance Under- standing for Solo Piano,” in Proceeding of the 25t International Society on Music Information Retrieval (ISMIR), 2024

  25. [33]

    Piano Skills Assessment,

    Paritosh Parmar, Jaiden Reddy, and Brendan Morris, “Piano Skills Assessment,” in IEEE 23th International Workshop on Multimedia Signal Processing (MMSP), 2021

  26. [34]

    LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment,

    Huan Zhang, Vincent Cheung, Hayato Nishioka, Simon Dixon, and Shinichi Furuya, “LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment,” in In Proceed- ings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025

  27. [35]

    MusicEval: A generative music dataset with expert ratings for automatic text-to-music evaluation,

    Cheng Liu, Hui Wang, Jinghua Zhao, Shiwan Zhao, Hui Bu, Xin Xu, Jiaming Zhou, Haoqin Sun, and Yong Qin, “MusicEval: A generative music dataset with expert ratings for automatic text-to-music evaluation,” arXiv:2501.10811, 2025

  28. [36]

    How does the teacher rate? Observations from the NeuroPiano dataset,

    Huan Zhang, Vincent Cheung, Hayato Nishioka, Simon Dixon, and Shinichi Furuya, “How does the teacher rate? Observations from the NeuroPiano dataset,” in Late-Breaking and Demo Ses- sion of the International Society on Music Information Retrieval (ISMIR), 2024

  29. [37]

    Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation,

    Or Tal, Alon Ziv, Itai Gat, Felix Kreuk, and Yossi Adi, “Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation,” in Proceeding of the 25t Interna- tional Society on Music Information Retrieval (ISMIR), 2024

  30. [38]

    Stable Audio Open,

    Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons, “Stable Audio Open,” Arxiv preprint arXiv:2407.14358, jul 2024

  31. [39]

    Simple and controllable music generation,

    Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre D´efossez, “Simple and controllable music generation,” in Thirty-seventh Confer- ence on Neural Information Processing Systems, 2023

  32. [40]

    Yue: Scaling open foundation models for long-form music generation,

    HKUST and MAP, “Yue: Scaling open foundation models for long-form music generation,” arXiv:2503.08638, 2025

  33. [41]

    DiffRhythm: Blazingly fast and embarrassingly simple end-to-end full- length song generation with latent diffusion,

    Ning Ziqian, Chen Huakang, Jiang Yuepeng, Hao Chunbo, Ma Guobin, Wang Shuai, Yao Jixun, and Xie Lei, “DiffRhythm: Blazingly fast and embarrassingly simple end-to-end full- length song generation with latent diffusion,” arXiv preprint arXiv:2503.01183, 2025

  34. [42]

    LP-MusicCaps: LLM-based pseudo music captioning,

    SeungHeon Doh, Keunwoo Choi, Jongpil Lee, and Juhan Nam, “LP-MusicCaps: LLM-based pseudo music captioning,” in Pro- ceeding of the 24th International Society on Music Information Retrieval (ISMIR), 2023

  35. [43]

    PANNs: Large-scale pre- trained audio neural networks for audio pattern recognition,

    Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D. Plumbley, “PANNs: Large-scale pre- trained audio neural networks for audio pattern recognition,” IEEE/ACM Trans. Audio, Speech and Lang. Proc., vol. 28, pp. 2880–2894, Nov. 2020

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.