Pith. sign in

REVIEW 2 major objections 5 minor 5 cited by

The AudioMOS Challenge 2025

T0 review · 2 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read The AudioMOS Challenge 2025 establishes that automatic prediction of human quality scores for synthetic audio is a tractable benchmark task across music, general audio, and multi-rate speech.

desk verdict First multi-modal audio MOS benchmark with useful results, but 'confirmed' is overstated; the system-level rank correlations lack uncertainty quantification. read the letter →

arxiv 2509.01336 v1 pith:XJ7BSIKY submitted 2025-09-01 cs.SD eess.AS

classification cs.SDeess.AS
keywords meanopinionscorepredictionaudioqualityassessmenttext-to-musicevaluationAudioboxAestheticssamplingrateself-supervisedlearningmodelensemblechallengebenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper reports on the first challenge devoted to automatic prediction of human quality ratings for generated audio. It claims that across three tasks—expert-rated text-to-music quality, the four Audiobox Aesthetics axes for speech/music/sound, and synthetic speech quality across sampling rates—most participating systems outperformed the provided baseline in system-level Spearman rank correlation. If these results hold, the challenge provides reusable benchmarks and confirms that self-supervised audio representations plus ensemble methods can predict perceived quality of synthetic audio without new listening tests. The intended consequence is that automatic evaluation can replace or augment expensive human listening tests for audio generation systems.

What carries the argument

The carrying objects are three baselines—CLAP with two MLP heads for Track 1, WavLM with MLP blocks for Track 2, and fine-tuned SSL-MOS for Track 3—plus the challenge's primary metric, system-level Spearman rank correlation between predicted and human scores across systems or conditions. The baselines define the bar to beat; the metric makes ranking the target rather than absolute score accuracy. The top systems' shared machinery is self-supervised audio representations, often music-specific, combined with specialized training objectives and model ensembling.

What would settle it

Recompute Track 3 system-level SRCC using only half of the 10 ratings per sample, or bootstrap over the 20 conditions, and count how often team rankings change; if the ordering of teams flips substantially, the claim that teams improved over baselines becomes unstable.

Watch

Extended reading notes

Core claim

In the paper's terms, the discovery is that a shared challenge with standardized human-labeled data makes automatic MOS prediction for synthetic audio tractable across previously unaddressed modalities. Track 1 shows the CLAP-based baseline ranking last on system-level SRCC for both overall musical quality and textual alignment. Track 2 shows the WavLM-based baseline ranking at or near the bottom across the four aesthetic axes. Track 3 shows every participating system exceeding the baseline's system-level SRCC of 0.749. The top systems relied on self-supervised representations, task-specific losses, and model ensembling. The paper interprets this as evidence that the community can advance au

Load-bearing premise

The human ratings used as ground truth are stable enough that ranking a small set of systems or conditions by those ratings is meaningful.

Editorial extensions

If this is right

  • Track 1 establishes that automatic predictors can rank text-to-music systems by both musical quality and textual alignment, with textual alignment the harder dimension.
  • Track 2 shows the four-axis aesthetics scheme can be predicted for speech, music, and sound, and that a small public training set can beat a baseline trained on 500 hours of proprietary data.
  • Track 3 shows mixing sampling rates changes the MOS task: 16 kHz conditions become the hardest to rank and predictors systematically under-rank them.
  • Successful systems in all three tracks consistently combine self-supervised representations with ensemble learning.
  • The released datasets give the research community reusable benchmarks for automatic evaluation of generated audio.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the Track 2 result generalizes, scaling training data matters less than choosing representations that match the target domain; a direct test would be re-training the baseline on gradually larger public subsets.
  • The systematic underprediction of 16 kHz conditions suggests predictors latch onto sampling-rate artifacts; sampling-rate-aware augmentation is a testable remedy.
  • Future editions could report per-rater agreement or bootstrap intervals on system-level SRCC to show whether 10 ratings per sample suffice for stable rankings.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper summarizes the AudioMOS Challenge 2025, a three-track benchmark for automatic quality assessment of synthetic audio: Track 1 covers text-to-music MOS and textual alignment (MusicEval), Track 2 covers the four Meta Audiobox Aesthetics axes on TTS/TTA/TTM samples, and Track 3 covers MOS prediction for synthetic speech at different sampling rates. It describes the datasets, baseline systems, the 24 participating teams, the main system-level Spearman rank correlation (SRCC) results, and lessons from system description forms. The central claim is that participating systems consistently outperformed the challenge baselines, with the phrase 'improvements over the baselines were confirmed' appearing in the abstract and Section IV.

Significance. If the reported results are robust, the challenge is a valuable community asset: it introduces new annotated test sets, provides public baseline code, documents top-system designs (SSL features, ensembling, specialized losses), and extends the VoiceMOS paradigm to music and general audio. The organizers' choice to make data and baselines available and to collect structured system descriptions is a concrete contribution. The main reservation is statistical: the evidence for the central claim consists of point estimates of system-level SRCC with no confidence intervals, significance tests, or reliability metrics for the test labels. Because the primary conclusion is exactly that the baselines were outperformed, this gap needs to be addressed before the claim can be accepted as confirmed.

major comments (2)
  1. [Section IV-A and Figs. 1–3] The paper's core conclusion—'improvements over the baselines were confirmed'—rests entirely on system-level SRCC point estimates. No confidence intervals, bootstrap resampling, or significance tests are reported. In Track 3, the system-level SRCC is computed over only 20 conditions (Table I), with each condition's MOS based on roughly 20 audio clips × 10 ratings. With N=20, the difference between B03 (0.749) and the best system (0.955) may well be within sampling variability; a formal test is needed. Since raw scores are only on the website, readers cannot assess the stability of the rankings from the paper. Please add confidence intervals or permutation tests for the primary metric (at minimum for baseline-vs-best comparisons), or weaken the 'confirmed' wording accordingly.
  2. [Section II-C / Table I] The test-set labels for Tracks 2 and 3 are not documented with reliability evidence. For Track 2, annotator qualification (Pearson > 0.7 on a golden set) is described for the AES-Natural training data, but it is not stated whether the 3,060 test-set samples were annotated under the same protocol or with any inter-annotator agreement check. For Track 3, the paper reports 10 ratings per sample and 20 listeners but no agreement metric. If the test labels are noisy, the system-level SRCC rankings—and hence the central claim that baselines were outperformed—become unstable. Please document the test-set labeling protocol and report inter-annotator agreement where available.
minor comments (5)
  1. [Throughout] There are LaTeX spacing artifacts in 'V oiceMOS' and 'F r ´echet'; these should be fixed in the camera-ready version.
  2. [Table I] The table is difficult to parse, especially the Track 3 row. Please clarify the numbers for train/dev/test samples and systems/conditions, and reconcile the table with the text stating that the Track 3 test phase contains 400 samples.
  3. [Section IV-A] The paper states that raw scores and rankings are on the challenge website. For a self-contained archival summary, please include a supplementary table of all system-level SRCC values (and ideally the other metrics) in an appendix.
  4. [Section IV-C] The statement that scaling up training data 'is not essentially effective' is an uncontrolled comparison between the baseline and participant systems that differ in architecture, ensembling, and training objectives. This should be phrased as a hypothesis rather than a conclusion.
  5. [Section II-C] The phrase '20 listeners participated in total' is ambiguous. Please clarify whether this is per listening test, across all four tests, or across all parts, and whether the same listeners rated all conditions.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: challenge results are externally produced by 24 independent teams; baselines are standard benchmarks, not derivations of the outcomes.

full rationale

The paper's central claim—that teams improved over baselines—rests on system-level SRCC values computed from participant submissions evaluated against human-rated ground truth. No step in this chain defines a prediction in terms of the quantity it claims to predict. The baselines (CLAP-based B01, WavLM-based B02, SSL-MOS B03) are trained models evaluated on held-out splits; they are not fitted to the test labels. The cited resources (MusicEval, Audiobox Aesthetics, SSL-MOS) include works by the authors, but they serve as datasets or pretrained feature extractors, and the participants' systems are external. No uniqueness theorem or ansatz is imported from self-citations to force a conclusion. The skeptical concern about small-sample SRCC and missing confidence intervals is a statistical robustness issue, not a circularity of derivation. Hence no circular step is identifiable by the paper's own equations or citations.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim depends on the reliability of human ratings, the representativeness of the test sets, and the chosen evaluation metric. No free parameters are fitted in the paper itself; the baselines are described but not used as derivational inputs.

assumptions (3)
  • domain assumption Human MOS ratings from expert or qualified annotators are a valid ground truth for synthetic audio quality.
    The challenge uses subjective ratings as the target labels for all three tracks; the paper provides some quality control (e.g., probe trials in Track 1, annotator qualification in Track 2) but no inter-annotator agreement metrics. This is the foundation of all benchmark results.
  • domain assumption The test sets are representative and sufficiently large to produce stable system-level rankings.
    Track 2 has 3060 test samples from multiple systems, while Track 3 has only 20 conditions; the paper does not analyze the statistical precision of the system-level SRCC values. This is invoked in Section II and Section IV.
  • domain assumption System-level Spearman rank correlation is an appropriate primary evaluation metric.
    The paper defines system-level SRCC as the primary metric in Section IV-A, following prior VoiceMOS challenges. This choice determines how all results are interpreted.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The AudioMOS Challenge 2025." pith.science (2026). https://pith.science/paper/XJ7BSIKY

@misc{pith2026250901336,
  author       = {Pith},
  title        = {Pith review of: The AudioMOS Challenge 2025},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XJ7BSIKY}},
  note         = {Machine review of arXiv:2509.01336}
}
read the original abstract

This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three tracks. The first track aims to assess text-to-music samples in terms of overall quality and textual alignment. The second track is based on the four evaluation dimensions of Meta Audiobox Aesthetics, and the test set consists of text-to-speech, text-to-audio, and text-to-music samples. The third track focuses on synthetic speech quality assessment in different sampling rates. The challenge attracted 24 unique teams from both academia and industry, and improvements over the baselines were confirmed. The outcome of this challenge is expected to facilitate development and progress in the field of automatic evaluation for audio generation systems.

Figures

Figures reproduced from arXiv: 2509.01336 by the authors.

Figure 2
Figure 2. Bar plot of system-level SRCC values of all participants [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Bar plot of system-level SRCC values of all participants [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves

    cs.SD 2026-07 conditional novelty 6.0 of 10

    SongSQA predicts both an overall singing-quality score and a temporal segment-level score curve for full-length songs, using teacher-generated pseudo labels and a learnable attention aggregator.

  2. JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions

    eess.AS 2026-05 unverdicted novelty 6.0 of 10

    JASTIN is an instruction-driven audio evaluation system that achieves state-of-the-art correlation with human ratings on speech, sound, music, and out-of-domain tasks without task-specific retraining.

  3. Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling

    cs.SD 2026-01 conditional novelty 6.0 of 10

    A song-aesthetics model with multi-stem cross-attention and hierarchical interval regression beats two adapted MOS baselines on average, with some dimensions showing ties or losses.

  4. Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Frozen SSL-Transformer embeddings generalize better than fine-tuned SSL or ViViT for cross-corpus MOS prediction, matching specialized SOTA on URGENT 2024 with MSE 0.36.

  5. Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents

    cs.CL 2026-05 unverdicted novelty 4.0 of 10

    Audio language models are benchmarked on five semantic and paralinguistic reasoning tasks to reveal limitations in handling spoken audio evidence, accent variation, and domain shifts.

Reference graph

Works this paper leans on

52 extracted references · 39 canonical work pages · cited by 5 Pith papers

  1. [1]

    The V oiceMOS Challenge 2022,

    W.-C. Huang, E. Cooper, Y . Tsao, H.-M. Wang, T. Toda, and J. Yamag- ishi, “The V oiceMOS Challenge 2022,” in Proc. Interspeech, 2022, pp. 4536–4540

  2. [2]

    The V oiceMOS Challenge 2023: Zero-Shot Subjective Speech Quality Prediction for Multiple Domains,

    E. Cooper, W.-C. Huang, Y . Tsao, H.-M. Wang, T. Toda, and J. Yam- agishi, “The V oiceMOS Challenge 2023: Zero-Shot Subjective Speech Quality Prediction for Multiple Domains,” in Proc. ASRU , 2023, pp. 1–7

  3. [3]

    The V oiceMOS Challenge 2024: Beyond Speech Quality Prediction,

    W.-C. Huang, S.-W. Fu, E. Cooper, R. E. Zezario, T. Toda, H.-M. Wang, J. Yamagishi, and Y . Tsao, “The V oiceMOS Challenge 2024: Beyond Speech Quality Prediction,” in Proc. SLT, 2024, pp. 803–810

  4. [4]

    How do voices from past speech synthesis challenges compare today?

    E. Cooper and J. Yamagishi, “How do voices from past speech synthesis challenges compare today?” in Proc. 11th ISCA Speech Synthesis Workshop (SSW 11) , 2021, pp. 183–188

  5. [5]

    Fr ´echet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms,

    K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, “Fr ´echet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms,” in Proc. Interspeech, 2019, pp. 2350–2354

  6. [6]

    Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models,

    R. Huang, J. Huang, D. Yang, Y . Ren, L. Liu, M. Li, Z. Ye, J. Liu, X. Yin, and Z. Zhao, “Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models,” in Proc. ICML , vol. 202, 23–29 Jul 2023, pp. 13 916–13 932

  7. [7]

    Evaluating generative audio systems and their metrics,

    A. Vinay and A. Lerch, “Evaluating generative audio systems and their metrics,” in Proc. ISMIR, 2022

  8. [8]

    Correlation of Fr ´echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependent,

    M. Tailleur, J. Lee, M. Lagrange, K. Choi, L. M. Heller, K. Imoto, and Y . Okamoto, “Correlation of Fr ´echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependent,” in Proc. EUSIPCO, 2024, pp. 56–60

Show all 52 references
  1. [9]

    Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound,

    A. Tjandra, Y .-C. Wu, B. Guo, J. Hoffman, B. Ellis, A. Vyas, B. Shi, S. Chen, M. Le, N. Zacharov, C. Wood, A. Lee, and W.-N. Hsu, “Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound,” 2025. [Online]. Available: https://arxiv.org/abs/2502.05139

  2. [10]

    MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation,

    C. Liu, H. Wang, J. Zhao, S. Zhao, H. Bu, X. Xu, J. Zhou, H. Sun, and Y . Qin, “MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation,” in Proc. ICASSP, 2025

  3. [11]

    LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning,

    M. Kawamura, R. Yamamoto, Y . Shirahata, T. Hasumi, and K. Tachibana, “LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning,” in Proc. Interspeech, 2024, pp. 1850–1854

  4. [12]

    AudioCaps: Generating captions for audios in the wild,

    C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proc. NAACL-HLT, Jun. 2019, pp. 119–132

  5. [13]

    Musiclm: Generating music from text,

    A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi et al. , “Musiclm: Generating music from text,” arXiv preprint arXiv:2301.11325 , 2023

  6. [14]

    LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,

    Y . Koizumi, H. Zen, S. Karita, Y . Ding, K. Yatabe, N. Morioka, M. Bacchiani, Y . Zhang, W. Han, and A. Bapna, “LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,” in Proc. Interspeech , 2023, pp. 5496–5500

  7. [15]

    Hi-Fi-CAPTAIN: High-fidelity and high-capacity conversational speech synthesis corpus developed by NICT,

    T. Okamoto, Y . Shiga, and H. Kawai, “Hi-Fi-CAPTAIN: High-fidelity and high-capacity conversational speech synthesis corpus developed by NICT,” https://ast-astrec.nict.go.jp/en/release/hi-fi-captain/, 2023

  8. [16]

    World: a vocoder-based high-quality speech synthesis system for real-time applications,

    M. Morise, F. Yokomori, and K. Ozawa, “World: a vocoder-based high-quality speech synthesis system for real-time applications,” IEICE TRANSACTIONS on Information and Systems , vol. 99, no. 7, pp. 1877– 1884, 2016

  9. [17]

    Fast Neural Speech Waveform Generative Models With Fully-Connected Layer-Based Upsampling,

    H. Yamashita, T. Okamoto, R. Takashima, Y . Ohtani, T. Takiguchi, T. Toda, and H. Kawai, “Fast Neural Speech Waveform Generative Models With Fully-Connected Layer-Based Upsampling,” IEEE Access, vol. 12, pp. 31 409–31 421, 2024

  10. [18]

    Speech masking system based on spatially separated multiple TTS maskers with a compact circular loudspeaker array,

    T. Okamoto, “Speech masking system based on spatially separated multiple TTS maskers with a compact circular loudspeaker array,” in Proc. ASRU, 2025

  11. [19]

    AudioSR: Versatile audio super-resolution at scale,

    H. Liu, K. Chen, Q. Tian, W. Wang, and M. D. Plumbley, “AudioSR: Versatile audio super-resolution at scale,” in Proc. ICASSP , 2024, pp. 1076–1080

  12. [20]

    pyloudnorm: A simple yet flexible loudness meter in python,

    C. J. Steinmetz and J. Reiss, “pyloudnorm: A simple yet flexible loudness meter in python,” in Audio Engineering Society Convention

  13. [21]

    Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,

    Y . Wu, K. Chen, T. Zhang, Y . Hui, T. Berg-Kirkpatrick, and S. Dub- nov, “Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,” in Proc. ICASSP, 2023

  14. [22]

    HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection,

    K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection,” in Proc. ICASSP, 2022

  15. [23]

    Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

    S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre- training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022

  16. [24]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017

  17. [25]

    Layer normalization,

    J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” arXiv preprint arXiv:1607.06450, 2016

  18. [26]

    Gaussian error linear units (gelus),

    D. Hendrycks and K. Gimpel, “Gaussian error linear units (gelus),” arXiv preprint arXiv:1606.08415, 2016

  19. [27]

    Generalization ability of MOS prediction networks,

    E. Cooper, W.-C. Huang, T. Toda, and J. Yamagishi, “Generalization ability of MOS prediction networks,” in Proc. ICASSP, 2022, pp. 8442– 8446

  20. [28]

    PAM: Prompting Audio-Language Models for Audio Quality Assessment,

    S. Deshmukh, D. Alharthi, B. Elizalde, H. Gamper, M. Al Ismail, R. Singh, B. Raj, and H. Wang, “PAM: Prompting Audio-Language Models for Audio Quality Assessment,” in Proc. Interspeech, 2024, pp. 3320–3324

  21. [29]

    EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,

    J. Richter, Y .-C. Wu, S. Krenn, S. Welker, B. Lay, S. Watanabe, A. Richard, and T. Gerkmann, “EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,” in Proc. Interspeech, 2024, pp. 4873–4877

  22. [30]

    Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models,

    Y . Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou, “Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models,” arXiv preprint arXiv:2311.07919 , 2023

  23. [31]

    BEATs: Audio Pre-Training with Acoustic Tokenizers,

    S. Chen, Y . Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “BEATs: Audio Pre-Training with Acoustic Tokenizers,” in Proc. ICML, vol. 202, 23–29 Jul 2023, pp. 5178–5193

  24. [32]

    Masked Modeling Duo: Towards a Universal Audio Pre-training Frame- work,

    D. Niizumi, D. Takeuchi, Y . Ohishi, N. Harada, and K. Kashino, “Masked Modeling Duo: Towards a Universal Audio Pre-training Frame- work,” IEEE/ACM TASLP, vol. 32, pp. 2391–2406, 2024

  25. [33]

    High Fidelity Neural Audio Compression,

    A. D ´efossez, J. Copet, G. Synnaeve, and Y . Adi, “High Fidelity Neural Audio Compression,” TMLR, 2023

  26. [34]

    Scaling up masked audio encoder learning for general audio classification,

    H. Dinkel, Z. Yan, Y . Wang, J. Zhang, Y . Wang, and B. Wang, “Scaling up masked audio encoder learning for general audio classification,” in Proc. Interspeech, 2024, pp. 547–551

  27. [35]

    MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training,

    Y . LI, R. Yuan, G. Zhang, Y . Ma, X. Chen, H. Yin, C. Xiao, C. Lin, A. Ragni, E. Benetos, N. Gyenge, R. Dannenberg, R. Liu, W. Chen, G. Xia, Y . Shi, W. Huang, Z. Wang, Y . Guo, and J. Fu, “MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training,” i...

  28. [36]

    MuQ: Self-supervised music representation learning with mel residual vector quantization,

    H. Zhu, Y . Zhou, H. Chen, J. Yu, Z. Ma, R. Gu, Y . Luo, W. Tan, and X. Chen, “MuQ: Self-supervised music representation learning with mel residual vector quantization,” arXiv preprint arXiv:2501.01108 , 2025

  29. [37]

    BERT: Pre-training of deep bidirectional transformers for language understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, J. Burstein, C. Doran, and T. Solorio, Eds., Jun. 2019, pp. 4171–4186

  30. [38]

    RoBERTa: A robustly optimized bert pretraining approach,

    Y . Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V . Stoyanov, “RoBERTa: A robustly optimized bert pretraining approach,” arXiv preprint arXiv:1907.11692 , 2019

  31. [39]

    Qwen3 Technical Report,

    A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. L...

  32. [40]

    Robust Speech Recognition via Large-Scale Weak Super- vision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. Mcleavey, and I. Sutskever, “Robust Speech Recognition via Large-Scale Weak Super- vision,” in Proc. ICML, vol. 202, 23–29 Jul 2023, pp. 28 492–28 518

  33. [41]

    HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,

    W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” IEEE/ACM TASLP, vol. 29, pp. 3451–3460, 2021

  34. [42]

    Scaling Speech Technology to 1,000+ Languages,

    V . Pratap, A. Tjandra, B. Shi, P. Tomasello, A. Babu, S. Kundu, A. Elkahky, Z. Ni, A. Vyas, M. Fazel-Zarandi, A. Baevski, Y . Adi, X. Zhang, W.-N. Hsu, A. Conneau, and M. Auli, “Scaling Speech Technology to 1,000+ Languages,” JMLR, vol. 25, no. 97, pp. 1–52, 2024

  35. [43]

    EAT: Self- Supervised Pre-Training with Efficient Audio Transformer,

    W. Chen, Y . Liang, Z. Ma, Z. Zheng, and X. Chen, “EAT: Self- Supervised Pre-Training with Efficient Audio Transformer,” in Proc. IJCAI, 8 2024, pp. 3807–3815

  36. [44]

    LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech,

    W.-C. Huang, E. Cooper, J. Yamagishi, and T. Toda, “LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech,” in Proc. ICASSP, 2022, pp. 896–900

  37. [45]

    KAN: Kolmogorov–Arnold networks,

    Z. Liu, Y . Wang, S. Vaidya, F. Ruehle, J. Halverson, M. Soljacic, T. Y . Hou, and M. Tegmark, “KAN: Kolmogorov–Arnold networks,” in Proc. ICLR, 2025

  38. [46]

    Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality,

    T. Dao and A. Gu, “Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality,” in Proc. ICML, vol. 235, 21–27 Jul 2024, pp. 10 041–10 071

  39. [47]

    Sampling- Frequency-Independent Convolutional Layer and its Application to Audio Source Separation,

    K. Saito, T. Nakamura, K. Yatabe, and H. Saruwatari, “Sampling- Frequency-Independent Convolutional Layer and its Application to Audio Source Separation,” IEEE/ACM TASLP, vol. 30, pp. 2928–2943, 2022

  40. [48]

    Kolmogorov-Arnold transformer,

    X. Yang and X. Wang, “Kolmogorov-Arnold transformer,” arXiv preprint arXiv:2409.10594, 2024

  41. [49]

    VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music,

    J. Shi, H. jin Shim, J. Tian, S. Arora, H. Wu, D. Petermann, J. Q. Yip, Y . Zhang, Y . Tang, W. Zhang, D. S. Alharthi, Y . Huang, K. Saito, J. Han, Y . Zhao, C. Donahue, and S. Watanabe, “VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music,” in Proc. NAACL – Sys...

  42. [50]

    XGBoost: A Scalable Tree Boosting Sys- tem,

    T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting Sys- tem,” in Proc. ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , ser. KDD ’16, 2016, p. 785–794

  43. [51]

    Pseudo Label Is Better Than Human Label,

    D. Hwang, K. C. Sim, Z. Huo, and T. Strohman, “Pseudo Label Is Better Than Human Label,” in Proc. Interspeech, 2022, pp. 1421–1425

  44. [150]

    Audio Engineering Society, 2021

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.