Pith. sign in

REVIEW 2 cited by

Aligning Text-to-Music Evaluation with Human Preferences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.16669 v1 pith:KC7YK55V submitted 2025-03-20 cs.SD cs.AIeess.AS

Aligning Text-to-Music Evaluation with Human Preferences

classification cs.SD cs.AIeess.AS
keywords humanaudiodesideratadivergenceeffectivelyevaluatingevaluationfind
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Despite significant recent advances in generative acoustic text-to-music (TTM) modeling, robust evaluation of these models lags behind, relying in particular on the popular Fr\'echet Audio Distance (FAD). In this work, we rigorously study the design space of reference-based divergence metrics for evaluating TTM models through (1) designing four synthetic meta-evaluations to measure sensitivity to particular musical desiderata, and (2) collecting and evaluating on MusicPrefs, the first open-source dataset of human preferences for TTM systems. We find that not only is the standard FAD setup inconsistent on both synthetic and human preference data, but that nearly all existing metrics fail to effectively capture desiderata, and are only weakly correlated with human perception. We propose a new metric, the MAUVE Audio Divergence (MAD), computed on representations from a self-supervised audio embedding model. We find that this metric effectively captures diverse musical desiderata (average rank correlation 0.84 for MAD vs. 0.49 for FAD and also correlates more strongly with MusicPrefs (0.62 vs. 0.14).

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AImoclips: A Benchmark for Evaluating Emotion Conveyance in Text-to-Music Generation

    cs.SD 2025-08 conditional novelty 6.0

    AImoclips is a new open benchmark showing that text-to-music systems convey high-arousal emotions better than low-arousal ones and that all models converge toward emotionally neutral music.

  2. The Name-Free Gap: Policy-Aware Stylistic Control in Music Generation

    cs.SD 2025-08 conditional novelty 5.0

    Word-based style descriptors generated by an LLM can shift MusicGen outputs toward a target artist's sound almost as much as using the artist's name, defining a name-free gap.