REVIEW 2 major objections 5 minor 5 cited by
The AudioMOS Challenge 2025
T0 review · 2 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The AudioMOS Challenge 2025 establishes that automatic prediction of human quality scores for synthetic audio is a tractable benchmark task across music, general audio, and multi-rate speech.
desk verdict First multi-modal audio MOS benchmark with useful results, but 'confirmed' is overstated; the system-level rank correlations lack uncertainty quantification. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying objects are three baselines—CLAP with two MLP heads for Track 1, WavLM with MLP blocks for Track 2, and fine-tuned SSL-MOS for Track 3—plus the challenge's primary metric, system-level Spearman rank correlation between predicted and human scores across systems or conditions. The baselines define the bar to beat; the metric makes ranking the target rather than absolute score accuracy. The top systems' shared machinery is self-supervised audio representations, often music-specific, combined with specialized training objectives and model ensembling.
What would settle it
Recompute Track 3 system-level SRCC using only half of the 10 ratings per sample, or bootstrap over the 20 conditions, and count how often team rankings change; if the ordering of teams flips substantially, the claim that teams improved over baselines becomes unstable.
Extended reading notes
Core claim
In the paper's terms, the discovery is that a shared challenge with standardized human-labeled data makes automatic MOS prediction for synthetic audio tractable across previously unaddressed modalities. Track 1 shows the CLAP-based baseline ranking last on system-level SRCC for both overall musical quality and textual alignment. Track 2 shows the WavLM-based baseline ranking at or near the bottom across the four aesthetic axes. Track 3 shows every participating system exceeding the baseline's system-level SRCC of 0.749. The top systems relied on self-supervised representations, task-specific losses, and model ensembling. The paper interprets this as evidence that the community can advance au
Load-bearing premise
The human ratings used as ground truth are stable enough that ranking a small set of systems or conditions by those ratings is meaningful.
Editorial extensions
If this is right
- Track 1 establishes that automatic predictors can rank text-to-music systems by both musical quality and textual alignment, with textual alignment the harder dimension.
- Track 2 shows the four-axis aesthetics scheme can be predicted for speech, music, and sound, and that a small public training set can beat a baseline trained on 500 hours of proprietary data.
- Track 3 shows mixing sampling rates changes the MOS task: 16 kHz conditions become the hardest to rank and predictors systematically under-rank them.
- Successful systems in all three tracks consistently combine self-supervised representations with ensemble learning.
- The released datasets give the research community reusable benchmarks for automatic evaluation of generated audio.
Reading between the lines
- If the Track 2 result generalizes, scaling training data matters less than choosing representations that match the target domain; a direct test would be re-training the baseline on gradually larger public subsets.
- The systematic underprediction of 16 kHz conditions suggests predictors latch onto sampling-rate artifacts; sampling-rate-aware augmentation is a testable remedy.
- Future editions could report per-rater agreement or bootstrap intervals on system-level SRCC to show whether 10 ratings per sample suffice for stable rankings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper summarizes the AudioMOS Challenge 2025, a three-track benchmark for automatic quality assessment of synthetic audio: Track 1 covers text-to-music MOS and textual alignment (MusicEval), Track 2 covers the four Meta Audiobox Aesthetics axes on TTS/TTA/TTM samples, and Track 3 covers MOS prediction for synthetic speech at different sampling rates. It describes the datasets, baseline systems, the 24 participating teams, the main system-level Spearman rank correlation (SRCC) results, and lessons from system description forms. The central claim is that participating systems consistently outperformed the challenge baselines, with the phrase 'improvements over the baselines were confirmed' appearing in the abstract and Section IV.
Significance. If the reported results are robust, the challenge is a valuable community asset: it introduces new annotated test sets, provides public baseline code, documents top-system designs (SSL features, ensembling, specialized losses), and extends the VoiceMOS paradigm to music and general audio. The organizers' choice to make data and baselines available and to collect structured system descriptions is a concrete contribution. The main reservation is statistical: the evidence for the central claim consists of point estimates of system-level SRCC with no confidence intervals, significance tests, or reliability metrics for the test labels. Because the primary conclusion is exactly that the baselines were outperformed, this gap needs to be addressed before the claim can be accepted as confirmed.
major comments (2)
- [Section IV-A and Figs. 1–3] The paper's core conclusion—'improvements over the baselines were confirmed'—rests entirely on system-level SRCC point estimates. No confidence intervals, bootstrap resampling, or significance tests are reported. In Track 3, the system-level SRCC is computed over only 20 conditions (Table I), with each condition's MOS based on roughly 20 audio clips × 10 ratings. With N=20, the difference between B03 (0.749) and the best system (0.955) may well be within sampling variability; a formal test is needed. Since raw scores are only on the website, readers cannot assess the stability of the rankings from the paper. Please add confidence intervals or permutation tests for the primary metric (at minimum for baseline-vs-best comparisons), or weaken the 'confirmed' wording accordingly.
- [Section II-C / Table I] The test-set labels for Tracks 2 and 3 are not documented with reliability evidence. For Track 2, annotator qualification (Pearson > 0.7 on a golden set) is described for the AES-Natural training data, but it is not stated whether the 3,060 test-set samples were annotated under the same protocol or with any inter-annotator agreement check. For Track 3, the paper reports 10 ratings per sample and 20 listeners but no agreement metric. If the test labels are noisy, the system-level SRCC rankings—and hence the central claim that baselines were outperformed—become unstable. Please document the test-set labeling protocol and report inter-annotator agreement where available.
minor comments (5)
- [Throughout] There are LaTeX spacing artifacts in 'V oiceMOS' and 'F r ´echet'; these should be fixed in the camera-ready version.
- [Table I] The table is difficult to parse, especially the Track 3 row. Please clarify the numbers for train/dev/test samples and systems/conditions, and reconcile the table with the text stating that the Track 3 test phase contains 400 samples.
- [Section IV-A] The paper states that raw scores and rankings are on the challenge website. For a self-contained archival summary, please include a supplementary table of all system-level SRCC values (and ideally the other metrics) in an appendix.
- [Section IV-C] The statement that scaling up training data 'is not essentially effective' is an uncontrolled comparison between the baseline and participant systems that differ in architecture, ensembling, and training objectives. This should be phrased as a hypothesis rather than a conclusion.
- [Section II-C] The phrase '20 listeners participated in total' is ambiguous. Please clarify whether this is per listening test, across all four tests, or across all parts, and whether the same listeners rated all conditions.
Circularity Check
No circularity: challenge results are externally produced by 24 independent teams; baselines are standard benchmarks, not derivations of the outcomes.
full rationale
The paper's central claim—that teams improved over baselines—rests on system-level SRCC values computed from participant submissions evaluated against human-rated ground truth. No step in this chain defines a prediction in terms of the quantity it claims to predict. The baselines (CLAP-based B01, WavLM-based B02, SSL-MOS B03) are trained models evaluated on held-out splits; they are not fitted to the test labels. The cited resources (MusicEval, Audiobox Aesthetics, SSL-MOS) include works by the authors, but they serve as datasets or pretrained feature extractors, and the participants' systems are external. No uniqueness theorem or ansatz is imported from self-citations to force a conclusion. The skeptical concern about small-sample SRCC and missing confidence intervals is a statistical robustness issue, not a circularity of derivation. Hence no circular step is identifiable by the paper's own equations or citations.
Assumptions & free parameters
assumptions (3)
- domain assumption Human MOS ratings from expert or qualified annotators are a valid ground truth for synthetic audio quality.
- domain assumption The test sets are representative and sufficiently large to produce stable system-level rankings.
- domain assumption System-level Spearman rank correlation is an appropriate primary evaluation metric.
Cite this review
Pith. "Pith review of The AudioMOS Challenge 2025." pith.science (2026). https://pith.science/paper/XJ7BSIKY
@misc{pith2026250901336,
author = {Pith},
title = {Pith review of: The AudioMOS Challenge 2025},
year = {2026},
howpublished = {\url{https://pith.science/paper/XJ7BSIKY}},
note = {Machine review of arXiv:2509.01336}
}
read the original abstract
This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three tracks. The first track aims to assess text-to-music samples in terms of overall quality and textual alignment. The second track is based on the four evaluation dimensions of Meta Audiobox Aesthetics, and the test set consists of text-to-speech, text-to-audio, and text-to-music samples. The third track focuses on synthetic speech quality assessment in different sampling rates. The challenge attracted 24 unique teams from both academia and industry, and improvements over the baselines were confirmed. The outcome of this challenge is expected to facilitate development and progress in the field of automatic evaluation for audio generation systems.
Figures
Forward citations
Cited by 5 Pith papers
-
Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves
SongSQA predicts both an overall singing-quality score and a temporal segment-level score curve for full-length songs, using teacher-generated pseudo labels and a learnable attention aggregator.
-
JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions
JASTIN is an instruction-driven audio evaluation system that achieves state-of-the-art correlation with human ratings on speech, sound, music, and out-of-domain tasks without task-specific retraining.
-
Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling
A song-aesthetics model with multi-stem cross-attention and hierarchical interval regression beats two adapted MOS baselines on average, with some dimensions showing ties or losses.
-
Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation
Frozen SSL-Transformer embeddings generalize better than fine-tuned SSL or ViViT for cross-corpus MOS prediction, matching specialized SOTA on URGENT 2024 with MSE 0.36.
-
Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents
Audio language models are benchmarked on five semantic and paralinguistic reasoning tasks to reveal limitations in handling spoken audio evidence, accent variation, and domain shifts.
Reference graph
Works this paper leans on
-
[1]
The V oiceMOS Challenge 2022,
W.-C. Huang, E. Cooper, Y . Tsao, H.-M. Wang, T. Toda, and J. Yamag- ishi, “The V oiceMOS Challenge 2022,” in Proc. Interspeech, 2022, pp. 4536–4540
2022
-
[2]
The V oiceMOS Challenge 2023: Zero-Shot Subjective Speech Quality Prediction for Multiple Domains,
E. Cooper, W.-C. Huang, Y . Tsao, H.-M. Wang, T. Toda, and J. Yam- agishi, “The V oiceMOS Challenge 2023: Zero-Shot Subjective Speech Quality Prediction for Multiple Domains,” in Proc. ASRU , 2023, pp. 1–7
work page 2023
-
[3]
The V oiceMOS Challenge 2024: Beyond Speech Quality Prediction,
W.-C. Huang, S.-W. Fu, E. Cooper, R. E. Zezario, T. Toda, H.-M. Wang, J. Yamagishi, and Y . Tsao, “The V oiceMOS Challenge 2024: Beyond Speech Quality Prediction,” in Proc. SLT, 2024, pp. 803–810
work page 2024
-
[4]
How do voices from past speech synthesis challenges compare today?
E. Cooper and J. Yamagishi, “How do voices from past speech synthesis challenges compare today?” in Proc. 11th ISCA Speech Synthesis Workshop (SSW 11) , 2021, pp. 183–188
work page 2021
-
[5]
Fr ´echet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms,
K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, “Fr ´echet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms,” in Proc. Interspeech, 2019, pp. 2350–2354
work page 2019
-
[6]
Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models,
R. Huang, J. Huang, D. Yang, Y . Ren, L. Liu, M. Li, Z. Ye, J. Liu, X. Yin, and Z. Zhao, “Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models,” in Proc. ICML , vol. 202, 23–29 Jul 2023, pp. 13 916–13 932
work page 2023
-
[7]
Evaluating generative audio systems and their metrics,
A. Vinay and A. Lerch, “Evaluating generative audio systems and their metrics,” in Proc. ISMIR, 2022
work page 2022
-
[8]
M. Tailleur, J. Lee, M. Lagrange, K. Choi, L. M. Heller, K. Imoto, and Y . Okamoto, “Correlation of Fr ´echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependent,” in Proc. EUSIPCO, 2024, pp. 56–60
work page 2024
Show all 52 references
-
[9]
Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound,
A. Tjandra, Y .-C. Wu, B. Guo, J. Hoffman, B. Ellis, A. Vyas, B. Shi, S. Chen, M. Le, N. Zacharov, C. Wood, A. Lee, and W.-N. Hsu, “Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound,” 2025. [Online]. Available: https://arxiv.org/abs/2502.05139
2025 arXiv
-
[10]
MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation,
C. Liu, H. Wang, J. Zhao, S. Zhao, H. Bu, X. Xu, J. Zhou, H. Sun, and Y . Qin, “MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation,” in Proc. ICASSP, 2025
2025
-
[11]
LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning,
M. Kawamura, R. Yamamoto, Y . Shirahata, T. Hasumi, and K. Tachibana, “LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning,” in Proc. Interspeech, 2024, pp. 1850–1854
2024
-
[12]
AudioCaps: Generating captions for audios in the wild,
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proc. NAACL-HLT, Jun. 2019, pp. 119–132
2019
-
[13]
Musiclm: Generating music from text,
A. Agostinelli, T. I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi et al. , “Musiclm: Generating music from text,” arXiv preprint arXiv:2301.11325 , 2023
2023 arXiv
-
[14]
LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,
Y . Koizumi, H. Zen, S. Karita, Y . Ding, K. Yatabe, N. Morioka, M. Bacchiani, Y . Zhang, W. Han, and A. Bapna, “LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,” in Proc. Interspeech , 2023, pp. 5496–5500
2023
-
[15]
Hi-Fi-CAPTAIN: High-fidelity and high-capacity conversational speech synthesis corpus developed by NICT,
T. Okamoto, Y . Shiga, and H. Kawai, “Hi-Fi-CAPTAIN: High-fidelity and high-capacity conversational speech synthesis corpus developed by NICT,” https://ast-astrec.nict.go.jp/en/release/hi-fi-captain/, 2023
2023
-
[16]
World: a vocoder-based high-quality speech synthesis system for real-time applications,
M. Morise, F. Yokomori, and K. Ozawa, “World: a vocoder-based high-quality speech synthesis system for real-time applications,” IEICE TRANSACTIONS on Information and Systems , vol. 99, no. 7, pp. 1877– 1884, 2016
2016
-
[17]
Fast Neural Speech Waveform Generative Models With Fully-Connected Layer-Based Upsampling,
H. Yamashita, T. Okamoto, R. Takashima, Y . Ohtani, T. Takiguchi, T. Toda, and H. Kawai, “Fast Neural Speech Waveform Generative Models With Fully-Connected Layer-Based Upsampling,” IEEE Access, vol. 12, pp. 31 409–31 421, 2024
2024
-
[18]
Speech masking system based on spatially separated multiple TTS maskers with a compact circular loudspeaker array,
T. Okamoto, “Speech masking system based on spatially separated multiple TTS maskers with a compact circular loudspeaker array,” in Proc. ASRU, 2025
2025
-
[19]
AudioSR: Versatile audio super-resolution at scale,
H. Liu, K. Chen, Q. Tian, W. Wang, and M. D. Plumbley, “AudioSR: Versatile audio super-resolution at scale,” in Proc. ICASSP , 2024, pp. 1076–1080
2024
-
[20]
pyloudnorm: A simple yet flexible loudness meter in python,
C. J. Steinmetz and J. Reiss, “pyloudnorm: A simple yet flexible loudness meter in python,” in Audio Engineering Society Convention
-
[21]
Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,
Y . Wu, K. Chen, T. Zhang, Y . Hui, T. Berg-Kirkpatrick, and S. Dub- nov, “Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,” in Proc. ICASSP, 2023
2023
-
[22]
HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection,
K. Chen, X. Du, B. Zhu, Z. Ma, T. Berg-Kirkpatrick, and S. Dubnov, “HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection,” in Proc. ICASSP, 2022
2022
-
[23]
Wavlm: Large-scale self-supervised pre- training for full stack speech processing,
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre- training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
-
[24]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
-
[25]
Layer normalization,
J. L. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” arXiv preprint arXiv:1607.06450, 2016
2016 arXiv
-
[26]
Gaussian error linear units (gelus),
D. Hendrycks and K. Gimpel, “Gaussian error linear units (gelus),” arXiv preprint arXiv:1606.08415, 2016
2016 arXiv
-
[27]
Generalization ability of MOS prediction networks,
E. Cooper, W.-C. Huang, T. Toda, and J. Yamagishi, “Generalization ability of MOS prediction networks,” in Proc. ICASSP, 2022, pp. 8442– 8446
2022
-
[28]
PAM: Prompting Audio-Language Models for Audio Quality Assessment,
S. Deshmukh, D. Alharthi, B. Elizalde, H. Gamper, M. Al Ismail, R. Singh, B. Raj, and H. Wang, “PAM: Prompting Audio-Language Models for Audio Quality Assessment,” in Proc. Interspeech, 2024, pp. 3320–3324
2024
-
[29]
EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,
J. Richter, Y .-C. Wu, S. Krenn, S. Welker, B. Lay, S. Watanabe, A. Richard, and T. Gerkmann, “EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,” in Proc. Interspeech, 2024, pp. 4873–4877
2024
-
[30]
Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models,
Y . Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou, “Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models,” arXiv preprint arXiv:2311.07919 , 2023
2023 arXiv
-
[31]
BEATs: Audio Pre-Training with Acoustic Tokenizers,
S. Chen, Y . Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “BEATs: Audio Pre-Training with Acoustic Tokenizers,” in Proc. ICML, vol. 202, 23–29 Jul 2023, pp. 5178–5193
2023
-
[32]
Masked Modeling Duo: Towards a Universal Audio Pre-training Frame- work,
D. Niizumi, D. Takeuchi, Y . Ohishi, N. Harada, and K. Kashino, “Masked Modeling Duo: Towards a Universal Audio Pre-training Frame- work,” IEEE/ACM TASLP, vol. 32, pp. 2391–2406, 2024
2024
-
[33]
High Fidelity Neural Audio Compression,
A. D ´efossez, J. Copet, G. Synnaeve, and Y . Adi, “High Fidelity Neural Audio Compression,” TMLR, 2023
2023
-
[34]
Scaling up masked audio encoder learning for general audio classification,
H. Dinkel, Z. Yan, Y . Wang, J. Zhang, Y . Wang, and B. Wang, “Scaling up masked audio encoder learning for general audio classification,” in Proc. Interspeech, 2024, pp. 547–551
2024
-
[35]
MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training,
Y . LI, R. Yuan, G. Zhang, Y . Ma, X. Chen, H. Yin, C. Xiao, C. Lin, A. Ragni, E. Benetos, N. Gyenge, R. Dannenberg, R. Liu, W. Chen, G. Xia, Y . Shi, W. Huang, Z. Wang, Y . Guo, and J. Fu, “MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training,” i...
2024
-
[36]
MuQ: Self-supervised music representation learning with mel residual vector quantization,
H. Zhu, Y . Zhou, H. Chen, J. Yu, Z. Ma, R. Gu, Y . Luo, W. Tan, and X. Chen, “MuQ: Self-supervised music representation learning with mel residual vector quantization,” arXiv preprint arXiv:2501.01108 , 2025
2025 arXiv
-
[37]
BERT: Pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, J. Burstein, C. Doran, and T. Solorio, Eds., Jun. 2019, pp. 4171–4186
2019
-
[38]
RoBERTa: A robustly optimized bert pretraining approach,
Y . Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V . Stoyanov, “RoBERTa: A robustly optimized bert pretraining approach,” arXiv preprint arXiv:1907.11692 , 2019
1907 arXiv
-
[39]
Qwen3 Technical Report,
A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. L...
2025 arXiv
-
[40]
Robust Speech Recognition via Large-Scale Weak Super- vision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. Mcleavey, and I. Sutskever, “Robust Speech Recognition via Large-Scale Weak Super- vision,” in Proc. ICML, vol. 202, 23–29 Jul 2023, pp. 28 492–28 518
2023
-
[41]
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,
W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” IEEE/ACM TASLP, vol. 29, pp. 3451–3460, 2021
2021
-
[42]
Scaling Speech Technology to 1,000+ Languages,
V . Pratap, A. Tjandra, B. Shi, P. Tomasello, A. Babu, S. Kundu, A. Elkahky, Z. Ni, A. Vyas, M. Fazel-Zarandi, A. Baevski, Y . Adi, X. Zhang, W.-N. Hsu, A. Conneau, and M. Auli, “Scaling Speech Technology to 1,000+ Languages,” JMLR, vol. 25, no. 97, pp. 1–52, 2024
2024
-
[43]
EAT: Self- Supervised Pre-Training with Efficient Audio Transformer,
W. Chen, Y . Liang, Z. Ma, Z. Zheng, and X. Chen, “EAT: Self- Supervised Pre-Training with Efficient Audio Transformer,” in Proc. IJCAI, 8 2024, pp. 3807–3815
2024
-
[44]
LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech,
W.-C. Huang, E. Cooper, J. Yamagishi, and T. Toda, “LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech,” in Proc. ICASSP, 2022, pp. 896–900
2022
-
[45]
KAN: Kolmogorov–Arnold networks,
Z. Liu, Y . Wang, S. Vaidya, F. Ruehle, J. Halverson, M. Soljacic, T. Y . Hou, and M. Tegmark, “KAN: Kolmogorov–Arnold networks,” in Proc. ICLR, 2025
2025
-
[46]
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality,
T. Dao and A. Gu, “Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality,” in Proc. ICML, vol. 235, 21–27 Jul 2024, pp. 10 041–10 071
2024
-
[47]
Sampling- Frequency-Independent Convolutional Layer and its Application to Audio Source Separation,
K. Saito, T. Nakamura, K. Yatabe, and H. Saruwatari, “Sampling- Frequency-Independent Convolutional Layer and its Application to Audio Source Separation,” IEEE/ACM TASLP, vol. 30, pp. 2928–2943, 2022
2022
-
[48]
Kolmogorov-Arnold transformer,
X. Yang and X. Wang, “Kolmogorov-Arnold transformer,” arXiv preprint arXiv:2409.10594, 2024
2024 arXiv
-
[49]
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music,
J. Shi, H. jin Shim, J. Tian, S. Arora, H. Wu, D. Petermann, J. Q. Yip, Y . Zhang, Y . Tang, W. Zhang, D. S. Alharthi, Y . Huang, K. Saito, J. Han, Y . Zhao, C. Donahue, and S. Watanabe, “VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music,” in Proc. NAACL – Sys...
2025
-
[50]
XGBoost: A Scalable Tree Boosting Sys- tem,
T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting Sys- tem,” in Proc. ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , ser. KDD ’16, 2016, p. 785–794
2016
-
[51]
Pseudo Label Is Better Than Human Label,
D. Hwang, K. C. Sim, Z. Huo, and T. Strohman, “Pseudo Label Is Better Than Human Label,” in Proc. Interspeech, 2022, pp. 1421–1425
2022
-
[150]
Audio Engineering Society, 2021
2021
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.