{"id":"a8bb5893-8e8d-40dc-8f0d-11f68256f771","arxiv_id":"2607.06929","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":3,"one_line_summary":"MADB is a 9,999-track music aesthetics benchmark with multi-dimensional professional annotations revealing that current pretrained audio models capture only partial aesthetic information.","lead":"The paper introduces MADB, a dataset of 9,999 music tracks annotated by 30 trained musicians across 10 aesthetic dimensions and one overall score, plus textual comments. It provides a new benchmark for training AI to evaluate music quality the way humans do, which is increasingly needed for AI music generation.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"No significant objection identified. The evaluation methodology has minor gaps (unspecified split ratio, single-seed MuQ), but none undermine the central claim that current models show moderate-at-best aesthetic prediction performance.","rationale":"The paper's central claim — that current models capture only partial aesthetic information — is supported by converging evidence: moderate correlations from supervised models (Table 2) and near-zero correlations from zero-shot audio-only LLM probing (Table 3). The annotation framework is professionally grounded with ICC~0.8 reliability. The methodological gaps (unspecified split ratio, single-seed MuQ evaluation) are worth noting but do not rise to the level of a load-bearing concern because: (a) the LLM experiment provides independent corroboration that doesn't depend on the supervised training setup; (b) even with optimistic variance estimates, LCC~0.72 is clearly below saturation; (c) the paper's primary contribution is the dataset itself, whose value depends on annotation quality and coverage, both of which are adequately demonstrated. The reader's ACCEPT verdict with MODERATE confidence is appropriate. The dimension-completeness concern the reader raises is a legitimate observation but is explicitly acknowledged in Section 5 and does not undermine the dataset's utility as a benchmark.","tokens_in":12943,"tokens_out":2404,"duration_ms":146024,"concrete_test":"Re-run the MuQ experiment with 4 seeds (matching the CLAP protocol) and report the split ratio used. If MuQ's LCC on the overall score drops below 0.65 or rises above 0.78 across seeds, the 'substantial gap' framing would need recalibration. Also confirm the validation set size — if fewer than 1000 tracks, report bootstrap confidence intervals on the correlation metrics to verify the gap claim is not an artifact of small validation samples.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader identifies dimension completeness as the weakest assumption, but this is acknowledged as a limitation rather than a load-bearing flaw — the paper does not claim the 10 dimensions are exhaustive, only that they provide structured coverage aligned with the music production pipeline. The more concrete concern is methodological: (1) Section 4.1 does not specify the train/validation split ratio, making it impossible to assess the variance of reported metrics; (2) MuQ and MERT are evaluated under a single seed (42) while CLAP variants use 4 seeds — the best-performing model (MuQ, LCC=0.718) thus has no reported variance, making the 'substantial gap' framing less well-supported than it appears. However, these are second-order issues. The core evidence is robust: ICC~0.8 supports annotation reliability, and the LLM audio-only experiment (Table 3, near-zero correlations across all dimensions) provides independent confirmation that raw audio representations fail to capture aesthetic judgments. Even if MuQ's true LCC shifted by ±0.05 with multiple seeds, the qualitative conclusion — models are far from saturating human-level aesthetic assessment — would hold. The paper is primarily a dataset contribution, and the dataset's value (scale, multi-dimensional annotations, professional annotators, ICC~0.8) stands regardless of the exact model performance numbers.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"The manuscript introduces MADB, a large-scale dataset for music aesthetic assessment comprising 9,999 tracks annotated by 30 trained musicians across 10 perceptual dimensions and one overall score, along with textual comments and tags. The authors establish a benchmark by evaluating several pretrained audio models (MuQ, MERT, CLAP variants) and a multimodal LLM (Qwen2-Audio) on the task of regressing aesthetic scores. The results indicate that current models capture only partial aesthetic information, with a substantial gap remaining to human-level judgment. The dataset is a valuable contribution, offering multi-dimensional, professional annotations that are currently lacking in the field.","tokens_in":13155,"tokens_out":1703,"duration_ms":96001,"significance":"The primary significance of this work lies in the dataset itself. The scale (9,999 tracks), the expertise of the annotators (30 trained musicians), and the multi-dimensional annotation framework aligned with the music production pipeline are substantial assets. The inclusion of textual comments and semantic tags further enriches the dataset, enabling multimodal analysis. The benchmarking results, particularly the near-zero correlation of audio-only LLM probing, provide a clear and falsifiable baseline for future research. The release of the dataset and code is a strong positive for reproducibility and community engagement.","major_comments":[{"comment":"Section 4.1 states that the dataset is randomly split into training and validation sets with a 'fixed ratio', but the ratio itself is not specified. This is a load-bearing methodological detail because the size of the validation set directly affects the variance of the reported metrics (MSE, LCC, SRCC, KRCC) in Table 2 and Table 3. Please specify the exact split ratio and, ideally, the number of samples in each set.","section":null},{"comment":"Section 4.1 and Table 2: The evaluation uses 4 fixed random seeds for CLAP-based experiments, but only a single seed (42) for MuQ and MERT. Since MuQ is the best-performing model and is central to the claim that models capture 'only partial aesthetic information' (LCC=0.718 on the overall score), the absence of variance estimates for MuQ and MERT makes it difficult to assess the statistical significance of the performance gap between MuQ and the CLAP variants. Please run MuQ and MERT with multiple seeds and report the standard deviation, consistent with the CLAP experiments.","section":null}],"minor_comments":[{"comment":"Section 3.5 uses the abbreviation 'ICCK', while Appendix A.1 uses 'ICCk' and 'ICC'. Please standardize the abbreviation for Intraclass Correlation Coefficient.","section":null},{"comment":"Section 3.2: 'Comments those less than 10 words will be ignored.' should be rephrased for grammatical correctness, e.g., 'Comments that are less than 10 words are ignored.'","section":null},{"comment":"Table 2 is quite dense and difficult to read. Consider splitting it into two tables (one for MSE/LCC, one for SRCC/KRCC) or using a clearer layout to improve readability.","section":null},{"comment":"Section 4.5 claims that the use of English comments 'proves the reliability of translations.' This is an overstatement; it merely suggests that the translated comments retain useful signal. Please soften this claim.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The dataset contribution is strong and clearly within the scope of the journal. The methodological gaps (missing split ratio, single-seed evaluation for the best model) are significant enough to require clarification but do not undermine the core value of the dataset. The self-citations to Jin et al. are frequent but appear relevant to the context of aesthetic evaluation. The paper should be accepted after these minor revisions are addressed."},"author_rebuttal":null,"desk_editor":{"model":"glm-5.2","letter":"Short version: MADB is a genuinely useful dataset paper. 9,999 tracks, 30 professional annotators, ~10 ratings per track across 10 perceptual dimensions plus textual comments, ICC values around 0.8. That's the contribution, and it's a real one. The benchmark experiments are secondary but do their job — they show current models are far from saturating human aesthetic judgment, which is the paper's central claim and is well-supported by the evidence.","headline":"Solid dataset contribution with real annotation rigor; methodological gaps are second-order and don't undermine the core finding.","tokens_in":13647,"tokens_out":696,"would_cite":true,"duration_ms":37904,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Audio models fail at music aesthetics; text comments carry the signal","keywords":["music aesthetic assessment","multi-dimensional annotation","audio representation learning","multimodal evaluation","benchmark dataset","human-aligned perception","contrastive learning"],"falsifier":"The strongest potential falsifier is the zero-shot audio-only LLM result: if a future model, trained on this dataset or a similar one, could predict multi-dimensional aesthetic scores from raw audio alone at correlation levels matching or exceeding text-based prediction, the paper's central claim about the audio representation bottleneck would be overturned.","tokens_in":13273,"feed_emoji":"","tokens_out":1100,"duration_ms":100680,"temperature":0.7,"pith_summary":"The paper introduces a dataset of 9,999 music tracks rated by 30 professionally trained annotators across 10 perceptual dimensions (such as melody perception, arrangement emotion, and singing skill) plus an overall aesthetic score, with short textual comments for each track. The authors use this dataset to benchmark a range of pretrained audio encoders and find that even the best models capture only partial aesthetic information: the top-performing audio encoder reaches a linear correlation of about 0.72 with human overall-score judgments, while a large multimodal language model given audio alone produces near-zero or negative correlation. The paper's central claim is that current audio representations are insufficient for human-aligned aesthetic assessment, but that human-written textual comments provide a strong supervisory signal that substantially closes the gap. When the language model is given comments and tags (with or without audio), correlation jumps to around 0.73, suggesting that aesthetic reasoning is currently carried more by language than by audio features.","feed_headline":"","feed_subtitle":"","key_machinery":"The benchmark has three moving parts: (1) a 10-dimension annotation framework covering composition, arrangement, performance, and post-production stages, each rated on a 0–5 scale by about 10 trained annotators per track; (2) a CLAP-based audio-text alignment pipeline that fuses annotator comments and structured tags via a learnable gating mechanism before contrastive pre-adaptation, tested against frozen encoders (MuQ, MERT, CLAP variants) with a Transformer regression head; (3) a zero-shot probing setup using a multimodal LLM (Qwen2-Audio-7B) with three input configurations (audio-only, text-only, audio+text) to isolate modality contributions without dataset-specific training.","core_discovery":"The paper establishes that the bottleneck for automatic music aesthetic assessment lies in audio representation quality, not in the availability of annotation dimensions or annotator agreement. Human comments encode compressed aesthetic reasoning that current audio models cannot extract from waveforms alone, as demonstrated by the sharp contrast between near-zero audio-only correlation and strong text-based correlation in zero-shot language-model probing. The dataset itself, with its multi-dimensional ratings and inter-annotator ICC values around 0.8, demonstrates that trained musicians can produce consistent multi-dimensional aesthetic judgments, making the gap a model limitation rather a标注","pith_inferences":["If textual comments are the primary carrier of aesthetic signal, then the quality and granularity of comment collection may matter more than the number of annotators per track for future dataset construction; a dataset with fewer annotators but richer per-track reasoning text might outperform one with more ratings but terse comments.","The 10-dimension framework is organized around music production stages, but the paper does not test whether these dimensions are independent or whether some are redundant; a factor analysis of the rating matrix could reveal whether 10 dimensions collapse to fewer latent factors, which would simplify both annotation and modeling.","The near-zero audio-only LLM result is obtained without any task-specific fine-tuning, so it does not rule out the possibility that a multimodal LLM fine-tuned on aesthetic ratings could learn to extract aesthetic features from audio; the paper's zero-shot design isolates intrinsic modality informativeness but leaves the fine-tuned ceiling unmeasured."],"forward_implications":["If the finding holds, future music aesthetic models will need to either develop audio encoders that internalize the perceptual reasoning currently captured only in text, or rely on hybrid pipelines where language models interpret audio features through aesthetic prompts.","The dataset's inclusion of AI-generated tracks (from Suno and Levo) alongside human-composed music means the benchmark can directly measure whether generative models are closing the aesthetic gap over time, providing a longitudinal evaluation tool for music generation research.","The strong performance of text-based supervision without aesthetic-score labels suggests that large-scale weakly-supervised pretraining on music reviews and comments could improve audio aesthetic models without requiring expensive multi-annotator rating data.","The finding that audio-only LLM probing yields near-zero correlation implies that current multimodal LLMs do not internally develop aesthetic representations from audio, challenging the assumption that scale alone will solve audio understanding."],"fun_headline_variants":["Audio representations, not human agreement, bottleneck music AI aesthetics","Music aesthetics gap: audio models fail where trained annotators agree","Trained musicians agree on music aesthetics, but audio models fall short","Text models capture music aesthetics better than audio models in zero-shot"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The paper assumes that its 10 chosen perceptual dimensions comprehensively decompose music aesthetics into stable, generalizable criteria. If key aesthetic factors fall outside these dimensions, or if the dimensions overlap substantially rather than capturing independent aspects of perception, the benchmark's validity as a holistic evaluation tool would be compromised.","fun_headline_variants_meta":{"raw":{"variants":["Audio representations, not human agreement, bottleneck music AI aesthetics","Music aesthetics gap: audio models fail where trained annotators agree","Trained musicians agree on music aesthetics, but audio models fall short","Text models capture music aesthetics better than audio models in zero-shot"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":1423,"prompt_tokens":395,"completion_tokens":1028,"prompt_tokens_details":null},"tokens_in":395,"tokens_out":1028,"duration_ms":32400,"temperature":1.0,"reasoning_tokens":1042,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T22:43:04.208791+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"The strongest potential falsifier is the zero-shot audio-only LLM result: if a future model, trained on this dataset or a similar one, could predict multi-dimensional aesthetic scores from raw audio alone at correlation levels matching or exceeding text-based prediction, the paper's central claim about the audio representation bottleneck would be overturned.","supporting_citations":[],"review_version":1}