{"id":"ddbadd87-492c-4421-985c-0a545c944a4b","arxiv_id":"2509.01336","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The first AudioMOS challenge compared automatic predictors of human quality scores for synthetic audio across three tracks, and most of the 24 participating teams outperformed the organizers' baselines.","lead":"The AudioMOS Challenge 2025 is a new competition where 24 teams built systems to automatically predict human quality ratings for synthetic music, speech, and audio. The paper reports the challenge design, the winning approaches, and evidence that most teams beat the provided baselines.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"System-level SRCC rankings lack uncertainty quantification; with only 20 Track-3 conditions and noisy subjective labels, the baseline-vs-top gap may not be statistically significant.","rationale":"The reader's weakest assumption concerns the reliability/stability of human-rated ground-truth labels, which is closely related to my concern. I agree that no inter-annotator agreement or confidence intervals are reported. My concern is more specific: it focuses on the small number of systems/conditions (especially 20 in Track 3) and the absence of any uncertainty quantification around the system-level SRCC values, which are the sole basis for the 'improvements over baselines' claim. This is a gap in statistical support, not an internal inconsistency. The paper is otherwise well-structured and the challenge logistics are described clearly. Because the reader already issued a conditional verdict, and my concern substantiates that conditionality without shifting the verdict, I recommend leaving the verdict unchanged. The concrete test proposed would either confirm or resolve this concern; if the bootstrap intervals show clear separation, the claim is likely solid, but the paper should still report these numbers to be reproducible.","tokens_in":11983,"tokens_out":3872,"duration_ms":45526,"concrete_test":"Obtain the released test-set ratings (ideally per-rating or per-clip data) from the challenge website. For Track 3, compute cluster-bootstrap 95% confidence intervals for each team's system-level SRCC by resampling audio clips (or conditions, preserving the 10 ratings per clip) with replacement, using the published evaluation script. Determine whether the baseline's confidence interval overlaps the top-performing team's interval. If they overlap significantly, or if bootstrap coverage indicates the ranking is unstable, the conclusion that 'improvements over the baselines were confirmed' lacks statistical support. Additionally compute Krippendorff's alpha on the per-rating test-set labels for Track 3; a value below 0.5 would indicate unreliable ground truth.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 'improvements over the baselines were confirmed' rests entirely on system-level Spearman rank correlations (SRCC) reported in Section IV. For Track 3, this metric is computed over just 20 conditions (Table I), where each condition's MOS is derived from only ~20 clips × 10 ratings. With such a small sample, SRCC estimates have wide confidence intervals. The paper reports no confidence intervals, no significance tests, and no inter-annotator agreement (e.g., Krippendorff's alpha) for the test-set labels. The gap between the baseline (0.749) and the top team (0.955) in Track 3, or the rank differences in Track 2 (baseline 9th/10th of 10 on some axes), could be within sampling variability if listener ratings are noisy. Unlike previous VoiceMOS challenges, the test-set labels for Tracks 1 and 3 are not described as having been collected with explicit quality control in this paper (Track 1 cites a separate dataset paper; Track 3 does not describe rating reliability at all). Since the paper directs readers to the website for raw scores, the published results cannot be independently assessed for stability. This is a load-bearing gap: if the rankings are not stable, the primary evidence for the challenge's success—that teams beat the baselines—is unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper summarizes the AudioMOS Challenge 2025, a three-track benchmark for automatic quality assessment of synthetic audio: Track 1 covers text-to-music MOS and textual alignment (MusicEval), Track 2 covers the four Meta Audiobox Aesthetics axes on TTS/TTA/TTM samples, and Track 3 covers MOS prediction for synthetic speech at different sampling rates. It describes the datasets, baseline systems, the 24 participating teams, the main system-level Spearman rank correlation (SRCC) results, and lessons from system description forms. The central claim is that participating systems consistently outperformed the challenge baselines, with the phrase 'improvements over the baselines were confirmed' appearing in the abstract and Section IV.","tokens_in":12269,"tokens_out":7501,"duration_ms":86767,"significance":"If the reported results are robust, the challenge is a valuable community asset: it introduces new annotated test sets, provides public baseline code, documents top-system designs (SSL features, ensembling, specialized losses), and extends the VoiceMOS paradigm to music and general audio. The organizers' choice to make data and baselines available and to collect structured system descriptions is a concrete contribution. The main reservation is statistical: the evidence for the central claim consists of point estimates of system-level SRCC with no confidence intervals, significance tests, or reliability metrics for the test labels. Because the primary conclusion is exactly that the baselines were outperformed, this gap needs to be addressed before the claim can be accepted as confirmed.","major_comments":[{"comment":"The paper's core conclusion—'improvements over the baselines were confirmed'—rests entirely on system-level SRCC point estimates. No confidence intervals, bootstrap resampling, or significance tests are reported. In Track 3, the system-level SRCC is computed over only 20 conditions (Table I), with each condition's MOS based on roughly 20 audio clips × 10 ratings. With N=20, the difference between B03 (0.749) and the best system (0.955) may well be within sampling variability; a formal test is needed. Since raw scores are only on the website, readers cannot assess the stability of the rankings from the paper. Please add confidence intervals or permutation tests for the primary metric (at minimum for baseline-vs-best comparisons), or weaken the 'confirmed' wording accordingly.","section":"Section IV-A and Figs. 1–3"},{"comment":"The test-set labels for Tracks 2 and 3 are not documented with reliability evidence. For Track 2, annotator qualification (Pearson > 0.7 on a golden set) is described for the AES-Natural training data, but it is not stated whether the 3,060 test-set samples were annotated under the same protocol or with any inter-annotator agreement check. For Track 3, the paper reports 10 ratings per sample and 20 listeners but no agreement metric. If the test labels are noisy, the system-level SRCC rankings—and hence the central claim that baselines were outperformed—become unstable. Please document the test-set labeling protocol and report inter-annotator agreement where available.","section":"Section II-C / Table I"}],"minor_comments":[{"comment":"There are LaTeX spacing artifacts in 'V oiceMOS' and 'F r ´echet'; these should be fixed in the camera-ready version.","section":"Throughout"},{"comment":"The table is difficult to parse, especially the Track 3 row. Please clarify the numbers for train/dev/test samples and systems/conditions, and reconcile the table with the text stating that the Track 3 test phase contains 400 samples.","section":"Table I"},{"comment":"The paper states that raw scores and rankings are on the challenge website. For a self-contained archival summary, please include a supplementary table of all system-level SRCC values (and ideally the other metrics) in an appendix.","section":"Section IV-A"},{"comment":"The statement that scaling up training data 'is not essentially effective' is an uncontrolled comparison between the baseline and participant systems that differ in architecture, ensembling, and training objectives. This should be phrased as a hypothesis rather than a conclusion.","section":"Section IV-C"},{"comment":"The phrase '20 listeners participated in total' is ambiguous. Please clarify whether this is per listening test, across all four tests, or across all parts, and whether the same listeners rated all conditions.","section":"Section II-C"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is the first standardized benchmark for MOS prediction across music, audio aesthetics, and sampling-rate-specific speech. That's a real contribution: three new tracks, public datasets (MusicEval, Audiobox Aesthetics-derived, and a new mixed-sampling-rate test set), and leaderboard results from 24 teams. The system-description analysis is also useful—the finding that most top teams used ensembling, and that SSL features plus explicit sampling-rate IDs mattered in Track 3, is actionable.\n\nThe main claim, that most teams beat the baselines, is directionally supported. B01 ranked lowest in Track 1, B02 ranked 9/7/9/10 across the four axes in Track 2, and B03 was worst in Track 3. But 'improvements over the baselines were confirmed' is strong. Everything rests on system-level SRCC computed over 30, 36, and 20 conditions, respectively, and the paper reports no confidence intervals or significance tests. In Track 3, with n=20, a Spearman rho of 0.955 vs 0.749 is not clearly outside sampling error. The stress-test note is fair. Track 3 also gives no information on annotator reliability, and raw scores are only on a website, so a reader cannot check the stability of rankings from the paper itself.\n\nThat said, this is not a reason to dismiss the paper. Challenge summaries in this area routinely omit uncertainty quantification, and the pattern here is consistent: the baselines lose on nearly every metric and axis. The more interesting analyses—which conditions are hardest (super-resolutioned natural speech, 16kHz conditions) and the finding that 16kHz becomes harder when mixed with other rates—are robust to the significance question. The authors also honestly note that scaling training data didn't help much in Track 2.\n\nFor peer review, I'd send it through. It deserves a serious referee, mainly to push the authors to temper the 'confirmed' language and either provide bootstrap CIs for the system-level SRCCs or make raw predictions available for external analysis. At a venue like Interspeech or SLT, this is acceptable with minor revision. I'd cite it if I work on audio quality prediction or TTM evaluation.\n\nRecommendation: don't desk reject; accept with revisions.","headline":"First multi-modal audio MOS benchmark with useful results, but 'confirmed' is overstated; the system-level rank correlations lack uncertainty quantification.","tokens_in":12762,"tokens_out":3020,"would_cite":true,"duration_ms":34114,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The AudioMOS Challenge 2025 establishes that automatic prediction of human quality scores for synthetic audio is a tractable benchmark task across music, general audio, and multi-rate speech.","keywords":["mean opinion score prediction","audio quality assessment","text-to-music evaluation","Audiobox Aesthetics","sampling rate","self-supervised learning","model ensemble","challenge benchmark"],"falsifier":"Recompute Track 3 system-level SRCC using only half of the 10 ratings per sample, or bootstrap over the 20 conditions, and count how often team rankings change; if the ordering of teams flips substantially, the claim that teams improved over baselines becomes unstable.","tokens_in":11895,"feed_emoji":"🎧","tokens_out":5454,"duration_ms":62005,"temperature":0.7,"pith_summary":"The paper reports on the first challenge devoted to automatic prediction of human quality ratings for generated audio. It claims that across three tasks—expert-rated text-to-music quality, the four Audiobox Aesthetics axes for speech/music/sound, and synthetic speech quality across sampling rates—most participating systems outperformed the provided baseline in system-level Spearman rank correlation. If these results hold, the challenge provides reusable benchmarks and confirms that self-supervised audio representations plus ensemble methods can predict perceived quality of synthetic audio without new listening tests. The intended consequence is that automatic evaluation can replace or augment expensive human listening tests for audio generation systems.","feed_headline":"Baselines fall in first audio-MOS challenge","feed_subtitle":"24 teams raced to predict human quality scores for music, speech, and sound; most outperformed official baselines.","key_machinery":"The carrying objects are three baselines—CLAP with two MLP heads for Track 1, WavLM with MLP blocks for Track 2, and fine-tuned SSL-MOS for Track 3—plus the challenge's primary metric, system-level Spearman rank correlation between predicted and human scores across systems or conditions. The baselines define the bar to beat; the metric makes ranking the target rather than absolute score accuracy. The top systems' shared machinery is self-supervised audio representations, often music-specific, combined with specialized training objectives and model ensembling.","core_discovery":"In the paper's terms, the discovery is that a shared challenge with standardized human-labeled data makes automatic MOS prediction for synthetic audio tractable across previously unaddressed modalities. Track 1 shows the CLAP-based baseline ranking last on system-level SRCC for both overall musical quality and textual alignment. Track 2 shows the WavLM-based baseline ranking at or near the bottom across the four aesthetic axes. Track 3 shows every participating system exceeding the baseline's system-level SRCC of 0.749. The top systems relied on self-supervised representations, task-specific losses, and model ensembling. The paper interprets this as evidence that the community can advance au","pith_inferences":["If the Track 2 result generalizes, scaling training data matters less than choosing representations that match the target domain; a direct test would be re-training the baseline on gradually larger public subsets.","The systematic underprediction of 16 kHz conditions suggests predictors latch onto sampling-rate artifacts; sampling-rate-aware augmentation is a testable remedy.","Future editions could report per-rater agreement or bootstrap intervals on system-level SRCC to show whether 10 ratings per sample suffice for stable rankings."],"forward_implications":["Track 1 establishes that automatic predictors can rank text-to-music systems by both musical quality and textual alignment, with textual alignment the harder dimension.","Track 2 shows the four-axis aesthetics scheme can be predicted for speech, music, and sound, and that a small public training set can beat a baseline trained on 500 hours of proprietary data.","Track 3 shows mixing sampling rates changes the MOS task: 16 kHz conditions become the hardest to rank and predictors systematically under-rank them.","Successful systems in all three tracks consistently combine self-supervised representations with ensemble learning.","The released datasets give the research community reusable benchmarks for automatic evaluation of generated audio."],"supporting_citations":[{"why":"Establishes the challenge format and system-level SRCC evaluation protocol used throughout.","marker":"[1]"},{"why":"Defines the primary metric and the ranking-error analysis method reused in Track 3.","marker":"[3]"},{"why":"Supplies the BVCC human-rated speech dataset used by several teams for pretraining and semi-supervised training.","marker":"[4]"},{"why":"Defines the four Audiobox Aesthetics axes and provides the AES-Natural training labels for Track 2.","marker":"[9]"},{"why":"Supplies the expert-rated text-to-music corpus with quality and alignment labels for Track 1.","marker":"[10]"},{"why":"Provides the CLAP model used as the Track 1 baseline encoder.","marker":"[21]"},{"why":"Provides the WavLM pretrained encoder used in the Track 2 baseline.","marker":"[23]"},{"why":"Provides the SSL-MOS model that the Track 3 baseline fine-tunes.","marker":"[27]"},{"why":"Supplies natural speech utterances used in the Track 3 listening tests.","marker":"[14]"}],"fun_headline_variants":["First audio-MOS challenge: 24 teams beat baselines","Inaugural AudioMOS Challenge: baselines outperformed","24 teams outperform baselines in first audio-MOS challenge","Three tracks, 24 teams, baselines beaten in AudioMOS","AudioMOS 2025: baselines fall to 24 teams"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The human ratings used as ground truth are stable enough that ranking a small set of systems or conditions by those ratings is meaningful.","fun_headline_variants_meta":{"raw":{"variants":["First audio-MOS challenge: 24 teams beat baselines","Inaugural AudioMOS Challenge: baselines outperformed","24 teams outperform baselines in first audio-MOS challenge","Three tracks, 24 teams, baselines beaten in AudioMOS","AudioMOS 2025: baselines fall to 24 teams"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000914,"raw_usage":{"total_tokens":3710,"prompt_tokens":637,"completion_tokens":3073,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":381,"completion_tokens_details":{"reasoning_tokens":2987}},"tokens_in":381,"tokens_out":3073,"duration_ms":24655,"temperature":1.0,"reasoning_tokens":2987,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:36:46.518539+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute Track 3 system-level SRCC using only half of the 10 ratings per sample, or bootstrap over the 20 conditions, and count how often team rankings change; if the ordering of teams flips substantially, the claim that teams improved over baselines becomes unstable.","supporting_citations":[{"cited_title":"The V oiceMOS Challenge 2024: Beyond Speech Quality Prediction,","cited_arxiv_id":null,"evidence_quote":"Defines the primary metric and the ranking-error analysis method reused in Track 3."},{"cited_title":"How do voices from past speech synthesis challenges compare today?","cited_arxiv_id":null,"evidence_quote":"Supplies the BVCC human-rated speech dataset used by several teams for pretraining and semi-supervised training."},{"cited_title":"MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation,","cited_arxiv_id":null,"evidence_quote":"Supplies the expert-rated text-to-music corpus with quality and alignment labels for Track 1."},{"cited_title":"Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,","cited_arxiv_id":null,"evidence_quote":"Provides the CLAP model used as the Track 1 baseline encoder."},{"cited_title":"LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,","cited_arxiv_id":null,"evidence_quote":"Supplies natural speech utterances used in the Track 3 listening tests."}],"review_version":1}