REVIEW 4 major objections 3 minor 49 references
P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge
T0 review · 4 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper reports that P.808 ACR subjective listening tests, the gold standard for speech quality, can miss intelligibility loss and hallucinations, giving generative enhancers high scores even when phone sequences are wrong.
desk verdict Useful practical contribution with a cautionary anecdote; the 'first time' claim outruns the evidence, but the paper deserves a serious referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument runs on the contrast between reference-free and reference-based metrics. The reference-free side includes the P.808 ACR MOS itself plus DNSMOS and NISQA; the reference-based side uses ESTOI for intelligibility and LPS (Levenshtein phone similarity, $\mathrm{LPS} = 1 - \mathrm{LPD}$) for phone fidelity, computed between enhanced and clean phone sequences. LPS is the hallucination detector because it compares actual phone strings, so a low LPS with high MOS is read as evidence that listeners scored fluency without registering wrong words. A second mechanism is the localization pipeline, which puts every instruction and test utterance in the target language using TTS-generated audio instructions and a target-language speech database disjoint from the rating clips.
What would settle it
Take the model-13 outputs that scored high MOS but low LPS and ask native listeners to either transcribe them or choose between the actual and hallucinated transcript; if listeners can reliably recover the correct phones when asked directly, then ACR MOS is not ignoring phone fidelity but merely not being asked about it, and the claimed subjective blind spot would be weakened.
Extended reading notes
Core claim
The central claim is that P.808 ACR listening tests, despite being subjective and treated as the gold standard, are by design reference-free quality ratings and can be insensitive to hallucinations and phone substitutions in generative speech enhancement. Evidence comes from the URGENT 2025 multilingual blind test: the same language ordering appears for MOS and LPS, the unseen Japanese language has the worst MOS and LPS (0.53) despite the best DNSMOS and NISQA scores, and one generative model received the second-highest MOS (3.34) yet the lowest ESTOI (0.53) and LPS (0.58). The authors state that this is, to their knowledge, the first report that ACR listening tests may ignore intelligibility and phone fidelity, and they recommend pairing MOS with ESTOI or LPS so that a high reference-free score together with a drop in LPS is read as a sign of hallucinations.
Load-bearing premise
The conclusion that ACR listening tests can miss hallucinations depends on LPS being a valid measure of perceptually relevant phone fidelity; if LPS is not a reliable hallucination detector, the gap between high MOS and low LPS for model #13 would indict LPS rather than the subjective test.
Editorial extensions
If this is right
- Reporting P.808 MOS alone for generative or hybrid enhancement systems is insufficient; MOS should be accompanied by ESTOI or LPS so that hallucinations are visible.
- Reference-free objective metrics such as DNSMOS and NISQA may rank systems highly even when intelligibility and phone fidelity collapse, so the URGENT multi-metric average-ranking approach is justified.
- A high reference-free subjective or objective score together with a drop in LPS should be treated as a warning sign of hallucination.
- The localization results show that non-English listening tests have lower worker acceptance rates (62.8% for German, 41.1% for Chinese, 28.9% for Japanese versus 82.3% for English), so multilingual P.808 deployment needs resubmission and reliability checking.
- The unseen Japanese language scores best on reference-free metrics but lowest on MOS and LPS, indicating that multilingual evaluation needs special attention to languages absent from training data.
Reading between the lines
- A direct extension of this paper would be an automatic hallucination alarm: compute MOS, DNSMOS, and NISQA alongside LPS, and flag any output whose reference-free scores are high while LPS falls below a language-specific threshold; the URGENT 2025 data could be used to tune that threshold.
- Because LPS only measures phone-string agreement, the same blind spot could affect downstream tasks such as automatic speech recognition; a generative enhancer that passes ACR MOS may still break ASR, so downstream-task metrics should probably be reported too.
- The lower acceptance rates in non-English languages suggest that crowdsourced multilingual ACR tests may be noisier precisely where hallucination detection is hardest; a useful test would be to compare MOS from professional native listeners against the crowd MOS for the same hallucinated outputs.
- A plausible testable prediction is that adding a forced-choice intelligibility or transcription question to the same listening session would reveal hallucinations that the ACR rating scale hides.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a localization process for ITU-T P.808 crowdsourced ACR listening tests from English into German, Chinese, and Japanese, and applies it to the URGENT 2025 speech enhancement challenge. It reports MOS results alongside objective metrics (PESQ, DNSMOS, NISQA, ESTOI, LPS) for a selection of challenge models. The central analytical claim is that P.808 ACR MOS, although reference-free and generally treated as a gold standard, can be insensitive to hallucinations produced by generative speech enhancement methods, and that MOS should therefore be accompanied by intelligibility or phone-fidelity metrics such as ESTOI or LPS. Evidence for this claim is drawn from model #13, a generative model with high MOS and DNSMOS/NISQA scores but low ESTOI and LPS, and from the low LPS scores observed for the unseen Japanese language.
Significance. If the central claim is upheld, the paper is significant: it challenges the uncritical use of subjective ACR MOS as a gold standard for generative speech enhancement, offers practical multilingual localization guidance, and provides a concrete recommendation (pair MOS with ESTOI/LPS) that is falsifiable and actionable for future challenges. The paper ships real challenge data, MOS confidence intervals, and a documented test protocol, which are strengths. However, the central claim currently rests on a single model and on the implicit validity of LPS as a perceptual hallucination detector; the paper would be much stronger with a direct subjective intelligibility or content-error test, an independent validation of LPS, and a statistical treatment of the objective metrics.
major comments (4)
- [Section IV, Table 3] The claim that P.808 ACR listening tests may ignore hallucination is inferred from a single generative model (ID 13), whose MOS of 3.34 contrasts with ESTOI 0.53 and LPS 0.58. This is an n=1 observation; no analysis is provided across all models (e.g., a scatterplot of MOS vs. LPS by model type, or a correlation conditional on model category), and the phrase "for the first time" overstates the strength of the evidence. To make the claim load-bearing, please add a direct subjective intelligibility or content-error detection test on the same stimuli, or at minimum a formal listening protocol with native expert annotators, to establish that listeners actually failed to perceive the hallucinations rather than that LPS fails as a perceptual proxy.
- [Section 3.2 and LPS definition] The paper assumes, without independent validation, that LPS validly measures perceptual phone fidelity and hallucination. LPS is defined as 1 minus the Levenshtein phone distance between enhanced and clean phone sequences, and it is supported only by the authors' own prior work [3]. A discrepancy between MOS and LPS could therefore be an artifact of LPS (e.g., sensitivity to phone alignment or ASR errors, or lack of perceptual calibration), not evidence that ACR is blind. Please provide (a) an independent validation of LPS against perceptual phone-fidelity judgments, especially for generative outputs, and (b) a robustness check using a different phone aligner or ASR system for the Japanese results and for model #13.
- [Sections 2.2 and 4.1] The subjective test validity is not established for non-English languages. The localization process asks crowdworkers only for self-reported fluency, not nativeness, and the acceptance rates after reliability checks are 82.3% (EN), 62.8% (DE), 41.1% (ZH), and 28.9% (JP). The high MOS of model #13 could thus reflect a listener pool that was insufficiently native or attentive to notice content errors, which is consistent with the authors' own recommendation to "ensure native listeners." To support the claim that ACR as a method is blind to hallucinations, the authors need to demonstrate that qualified native listeners also rate #13 highly despite hallucinations, or to analyze the ratings separately by listener language proficiency.
- [Section 4.2, Tables 1 and 3] Only MOS values are reported with 95% confidence intervals; ESTOI, LPS, PESQ, DNSMOS, and NISQA are given without confidence intervals or significance tests. Statements such as "by far poorest overall intelligibility and phone fidelity" for model #13 and the "dramatic drop" of LPS for Japanese are therefore not statistically supported. Please provide bootstrap confidence intervals or per-condition significance tests for the objective metrics, particularly for the language-level comparisons in Table 1 and the model-level comparisons in Table 3.
minor comments (3)
- [Section 4.2] The "informal inspection" of model #13 in DE and EN is mentioned as confirmation of hallucination, but no details are given (number of listeners, stimuli, judgment criteria). Please either formalize this inspection or remove it as evidence.
- [References] Reference [10] is listed as "Anonymous"; please update it to the actual challenge paper with authors and venue once available.
- [Section 5 and Abstract] The paper repeatedly states that scripts will be released "soon" but gives no repository URL or release timeline; for reproducibility, please provide a persistent link or supplement at publication time.
Circularity Check
No circular reduction is present: the central claim is an empirical MOS-versus-LPS discrepancy, and the only self-citation concerns the LPS metric's provenance without being load-bearing.
full rationale
The paper's central claim—that P.808 ACR listening tests may ignore intelligibility, phone fidelity, or hallucination—is an empirical observation from URGENT 2025 results, not a derivation. It rests on the measured discrepancy for model #13 between subjective MOS (3.34) and objective ESTOI (0.53) and LPS (0.58). These quantities are computed independently, so the observation is not forced by construction; a different pattern could have undermined the conclusion. The only notable self-citation is the use of LPS, attributed to the authors' own prior work [3], and the URGENT challenge infrastructure [5]. LPS is a defined metric (1 minus Levenshtein phone distance) and its numerical values are externally computable; the paper does not fit LPS to the MOS data or rename an input as a prediction. The inference that high MOS with low LPS means listeners missed hallucinations does rely on LPS validity and on the listener pool being sufficiently native, but those are validity and evidence concerns, not circularity. Accordingly, no circular step is identified; the score of 2 reflects only the presence of minor self-citations that are not load-bearing.
Assumptions & free parameters
assumptions (4)
- domain assumption The LPS metric measures phone fidelity and reliably detects hallucinations.
- domain assumption P.808 ACR crowdsourced listening tests remain the gold standard for speech quality.
- domain assumption The localized test versions are equivalent to the English original despite different acceptance rates.
- domain assumption Eight ratings per utterance are sufficient for MOS reliability.
Cite this review
Pith. "Pith review of P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge." pith.science (2026). https://pith.science/paper/QSEV4CVW
@misc{pith2026250711306,
author = {Pith},
title = {Pith review of: P.808 Multilingual Speech Enhancement Testing: Approach and Results of URGENT 2025 Challenge},
year = {2026},
howpublished = {\url{https://pith.science/paper/QSEV4CVW}},
note = {Machine review of arXiv:2507.11306}
}
read the original abstract
In speech quality estimation for speech enhancement (SE) systems, subjective listening tests so far are considered as the gold standard. This should be even more true considering the large influx of new generative or hybrid methods into the field, revealing issues of some objective metrics. Efforts such as the Interspeech 2025 URGENT Speech Enhancement Challenge also involving non-English datasets add the aspect of multilinguality to the testing procedure. In this paper, we provide a brief recap of the ITU-T P.808 crowdsourced subjective listening test method. A first novel contribution is our proposed process of localizing both text and audio components of Naderi and Cutler's implementation of crowdsourced subjective absolute category rating (ACR) listening tests involving text-to-speech (TTS). Further, we provide surprising analyses of and insights into URGENT Challenge results, tackling the reliability of (P.808) ACR subjective testing as gold standard in the age of generative AI. Particularly, it seems that for generative SE methods, subjective (ACR MOS) and objective (DNSMOS, NISQA) reference-free metrics should be accompanied by objective phone fidelity metrics to reliably detect hallucinations. Finally, we will soon release our localization scripts and methods for easy deployment for new multilingual speech enhancement subjective evaluations according to ITU-T P.808.
Reference graph
Works this paper leans on
-
[2]
Hallucination in Perceptual Metric- Driven Speech Enhancement Networks,
G. Close, T. Hain, and S. Goetze, “Hallucination in Perceptual Metric- Driven Speech Enhancement Networks,” inProc. of EUSIPCO, Lyon, France, Aug. 2024, pp. 21–25
work page 2024
-
[4]
D. de Oliveira, J. Richter, J.-M. Lemercier, T. Peer, and T. Gerkmann, “On the Behavior of Intrusive and Non-Intrusive Speech Enhancement Metrics in Predictive and Generative Settings,” inProc. of 15th ITG Conference on Speech Communication, Aachen, Germany, Sep 2023, pp. 260–264
work page 2023
-
[3]
Evaluation Metrics for Generative Speech Enhancement Methods: Issues and Perspectives,
J. Pirklbauer, M. Sach, K. Fluyt, W. Tirry, W. Wardah, S. M ¨oller, and T. Fingscheidt, “Evaluation Metrics for Generative Speech Enhancement Methods: Issues and Perspectives,” inProc. of 15th ITG Conference on Speech Communication, Aachen, Germany, Sep 2023, pp. 265–269
work page 2023
-
[1]
Speech Enhancement and Dereverberation With Diffusion-Based Generative Models,
J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, “Speech Enhancement and Dereverberation With Diffusion-Based Generative Models,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 2351–2364, 2023
work page 2023
-
[5]
URGENT Challenge: Universality, Robustness, and Generalizability for Speech Enhancement,
W. Zhang, R. Scheibler, K. Saijo, S. Cornell, C. Li, Z. Ni, J. Pirklbauer, M. Sach, S. Watanabe, T. Fingscheidt, and Y . Qian, “URGENT Challenge: Universality, Robustness, and Generalizability for Speech Enhancement,” inProc. of Interspeech, Kos, Greece, Sep. 2024, pp. 4868–4872
work page 2024
-
[6]
ITU,Rec. P .800: Methods for Subjective Determination of Transmission Quality, International Telecommunication Union, Telecommunication Standardization Sector (ITU-T), Aug. 1996
work page 1996
-
[7]
An Open Source Implementation of ITU- T Recommendation P.808 with Validation,
B. Naderi and R. Cutler, “An Open Source Implementation of ITU- T Recommendation P.808 with Validation,” inProc. of Interspeech, Shanghai, China, Oct. 2020, pp. 2862–2866
work page 2020
-
[8]
ITU,Rec. P .808: Subjective Evaluation of Speech Quality with a Crowd- sourcing Approach, International Telecommunication Union, Telecom- munication Standardization Sector (ITU-T), Jun. 2001
work page 2001
Show all 49 references
-
[9]
Subjective Evaluation of Noise Suppression Algorithms in Crowdsourcing,
B. Naderi and R. Cutler, “Subjective Evaluation of Noise Suppression Algorithms in Crowdsourcing,” inProc. of Interspeech, Brno, Czech Republic, Aug. 2021, pp. 2132–2136
2021
-
[10]
Interspeech 2025 URGENT Speech Enhancement Chal- lenge,
Anonymous, “Interspeech 2025 URGENT Speech Enhancement Chal- lenge,” submitted to Interspeech, 2025
2025
-
[11]
Disentangling the Impacts of Language and Channel Variability on Speech Separation Networks,
F.-L. Wang, H.-S. Lee, Y . Tsao, and H.-M. Wang, “Disentangling the Impacts of Language and Channel Variability on Speech Separation Networks,” inProc. of Interspeech, Incheon, Korea, Sep. 2022, pp. 5343– 5347
2022
-
[12]
Crowdsourced Multilingual Speech Intelligibility Testing,
L. Lechler and K. Wojcicki, “Crowdsourced Multilingual Speech Intelligibility Testing,” inProc. of ICASSP, Seoul, Korea, Apr. 2024, pp. 1441–1445
2024
-
[13]
A Subjective Listening Test of Six Different Artificial Bandwidth Extension Approaches in English, Chinese, German, and Korean,
J. Abel, M. Kaniewska, C. Guillaume, W. Tirry, H. Pulakka, V . Myllyl ¨a, J. Sj ¨oberg, P. Alku, I. Katsir, D. Malah, I. Cohen, M. A. Tugtekin Turan, E. Erzin, T. Schlien, P. Vary, A. Nour-Eldin, P. Kabal, and T. Fingscheidt, “A Subjective Listening Test of Six Different Artif...
2016
-
[14]
Language barriers: Evaluating cross-lingual performance of CNN and transformer architectures for speech quality estimation,
W. Wardah, T. M. K. B ¨uy¨uktas, K. Shchegelskiy, S. M ¨oller, and R. P. Spang, “Language barriers: Evaluating cross-lingual performance of CNN and transformer architectures for speech quality estimation,” inProc. of ISCA/ITG Workshop on Diversity in Large Speech and Language ...
2025
-
[15]
NESC: Robust Neural End-2-End Speech Coding with GANs,
N. Pia, K. Gupta, S. Korse, M. Multrus, and G. Fuchs, “NESC: Robust Neural End-2-End Speech Coding with GANs,” inProc. of Interspeech, Incheon, Korea, Sep. 2022, pp. 4212–4216
2022
-
[16]
Enhancing Multilingual TTS with V oice Conversion Based Data Augmentation and Posterior Embedding,
H.-W. Yoon, J.-S. Kim, R. Yamamoto, R. Terashima, C.-H. Song, J.-M. Kim, and E. Song, “Enhancing Multilingual TTS with V oice Conversion Based Data Augmentation and Posterior Embedding,” inProc. of ICASSP, Seoul, Korea, Apr. 2024, pp. 12 186–12 190
2024
-
[17]
P .807: Subjective Test Methodology for Assessing Speech In- telligibility, International Telecommunication Union, Telecommunication Standardization Sector (ITU-T), Feb
ITU,Rec. P .807: Subjective Test Methodology for Assessing Speech In- telligibility, International Telecommunication Union, Telecommunication Standardization Sector (ITU-T), Feb. 2016
2016
-
[18]
Amazon Mechanical Turk,
“Amazon Mechanical Turk,” 2005, accessed: 2025-04-04. [Online]. Available: https://www.mturk.com/
2005
-
[19]
Application of Just-Noticeable Difference in Quality as Environment Suitability Test for Crowdsourcing Speech Quality Assessment Task,
B. Naderi and S. M ¨oller, “Application of Just-Noticeable Difference in Quality as Environment Suitability Test for Crowdsourcing Speech Quality Assessment Task,” inProc. of QoMEX, Athlone, Ireland, May 2020, pp. 1–6
2020
-
[20]
Amazon Polly,
“Amazon Polly,” 2016, accessed: 2025-04-04. [Online]. Available: https://aws.amazon.com/polly/
2016
-
[21]
ICASSP 2023 Deep Noise Suppression Challenge,
H. Dubey, A. Aazami, V . Gopal, B. Naderi, S. Braun, R. Cutler, A. Ju, M. Zohourian, M. Tang, M. Golestaneh, and R. Aichner, “ICASSP 2023 Deep Noise Suppression Challenge,”IEEE Open Journal of Signal Processing, vol. 5, pp. 725–737, 2024
2023
-
[22]
LibriTTS: A Corpus Derived from LibriSpeech for Text-to- Speech,
H. Zen, V . Dang, R. Clark, Y . Zhang, R. J. Weiss, Y . Jia, Z. Chen, and Y . Wu, “LibriTTS: A Corpus Derived from LibriSpeech for Text-to- Speech,” inProc. of Interspeech, Graz, Austria, Sep. 2019, pp. 1526–1530
2019
-
[23]
The V oice Bank Corpus: Design, Collection and Data Analysis of a Large Regional Accent Speech Database,
C. Veaux, J. Yamagishi, and S. King, “The V oice Bank Corpus: Design, Collection and Data Analysis of a Large Regional Accent Speech Database,” inProc. of O-COCOSDA/CASLRE, Gurgaon, India, Nov. 2013, pp. 1–4
2013
-
[24]
CSR-I (WSJ0) Complete,
J. Garofolo, D. Graff, D. Paul, and D. Pallett, “CSR-I (WSJ0) Complete,” Linguistic Data Consortium, Philadelphia, 2007
2007
-
[25]
EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,
J. Richter, Y .-C. Wu, S. Krenn, S. Welker, B. Lay, S. Watanabe, A. Richard, and T. Gerkmann, “EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,” in Proc. of Interspeech, Kos, Greece, Sep. 2024, pp. 4873–4877
2024
-
[26]
MLS: A Large-Scale Multilingual Dataset for Speech Research,
V . Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A Large-Scale Multilingual Dataset for Speech Research,” inProc. of Interspeech, Shanghai, China, Oct. 2020, pp. 2757–2761
2020
-
[27]
Common V oice: A Massively-Multilingual Speech Corpus,
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common V oice: A Massively-Multilingual Speech Corpus,” inProc. of LREC, Marseille, France, May 2020, pp. 4218–4222
2020
-
[28]
Open-Source Multi-Speaker Corpora of the English Accents in the British Isles,
I. Demirsahin, O. Kjartansson, A. Gutkin, and C. Rivera, “Open-Source Multi-Speaker Corpora of the English Accents in the British Isles,” in Proc. of LREC, Marseille, France, May 2020, pp. 6532–6541
2020
-
[29]
Crowdsourcing Latin American Spanish for Low-Resource Text-to-Speech,
A. Guevara-Rukoz, I. Demirsahin, F. He, S.-H. C. Chu, S. Sarin, K. Pi- patsrisawat, A. Gutkin, A. Butryna, and O. Kjartansson, “Crowdsourcing Latin American Spanish for Low-Resource Text-to-Speech,” inProc. of LREC, Marseille, France, May 2020, pp. 6504–6513
2020
-
[30]
Vibravox: A Dataset of French Speech Captured with Body-Conduction Audio Sensors,
J. Hauret, M. Olivier, T. Joubaud, C. Langrenne, S. Poir ´ee, V . Zimpfer, and ´Eric Bavu, “Vibravox: A Dataset of French Speech Captured with Body-Conduction Audio Sensors,”arXiv:2407.11828, Mar. 2025
2025 arXiv
-
[31]
Mining the Spoken Wikipedia for Speech Data and Beyond,
A. K ¨ohn, F. Stegen, and T. Baumann, “Mining the Spoken Wikipedia for Speech Data and Beyond,” inProc. of LREC, Portoro ˇz, Slovenia, May 2016, pp. 4644–4647
2016
-
[32]
AISHELL-3: A Multi- Speaker Mandarin TTS Corpus,
Y . Shi, H. Bu, X. Xu, S. Zhang, and M. Li, “AISHELL-3: A Multi- Speaker Mandarin TTS Corpus,” inProc. of Interspeech, Brno, Czechia, Aug. 2021, pp. 2756–2760
2021
-
[33]
Hi-Fi-CAPTAIN: High-Fidelity and High-Capacity Conversational Speech Synthesis Corpus Developed by NICT,
T. Okamoto, Y . Shiga, and H. Kawai, “Hi-Fi-CAPTAIN: High-Fidelity and High-Capacity Conversational Speech Synthesis Corpus Developed by NICT,” https://ast-astrec.nict.go.jp/en/release/hi-fi-captain/, 2023
2023
-
[34]
JVNV: A Corpus of Japanese Emotional Speech With Verbal Content and Nonverbal Expressions,
D. Xin, J. Jiang, S. Takamichi, Y . Saito, A. Aizawa, and H. Saruwatari, “JVNV: A Corpus of Japanese Emotional Speech With Verbal Content and Nonverbal Expressions,”IEEE Access, vol. 12, pp. 19 752–19 764, 2024
2024
-
[35]
CHiME-6 Challenge: Tackling Multispeaker Speech Recognition for Unsegmented Recordings,
S. Watanabe, M. Mandel, J. Barker, E. Vincent, A. Arora, X. Chang, S. Khudanpur, V . Manohar, D. Povey, D. Rajet al., “CHiME-6 Challenge: Tackling Multispeaker Speech Recognition for Unsegmented Recordings,” inProc. of CHiME Workshop, virtual, May 2020
2020
-
[36]
Introducing the COVID-19 YouTube (COVYT) Speech Dataset Featuring the Same Speakers with and Without Infection,
A. Triantafyllopoulos, A. Semertzidou, M. Song, F. B. Pokorny, and B. W. Schuller, “Introducing the COVID-19 YouTube (COVYT) Speech Dataset Featuring the Same Speakers with and Without Infection,”Biomedical Signal Processing and Control, vol. 88, p. 105642, 2024
2024
-
[37]
FLEURS: FEW-Shot Learning Evaluation of Universal Representations of Speech,
A. Conneau, M. Ma, S. Khanuja, Y . Zhang, V . Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna, “FLEURS: FEW-Shot Learning Evaluation of Universal Representations of Speech,” inProc. of SLT, Doha, Qatar, Jan. 2023, pp. 798–805
2023
-
[38]
V oiceHome-2, an Extended Corpus for Multichannel Speech Processing in Real Homes,
N. Bertin, E. Camberlein, R. Lebarbenchon, E. Vincent, S. Sivasankaran, I. Illina, and F. Bimbot, “V oiceHome-2, an Extended Corpus for Multichannel Speech Processing in Real Homes,”Speech Communication, vol. 106, pp. 68–78, 2019
2019
-
[39]
AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario,
Y . Fu, L. Cheng, S. Lv, Y . Jv, Y . Kong, Z. Chen, Y . Hu, L. Xie, J. Wu, H. Bu, X. Xu, J. Du, and J. Chen, “AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario,” inProc. of Interspeech, Brno, Czechi...
2021
-
[40]
Yodas: Youtube-Oriented Dataset for Audio and Speech,
X. Li, S. Takamichi, T. Saeki, W. Chen, S. Shiota, and S. Watanabe, “Yodas: Youtube-Oriented Dataset for Audio and Speech,” inProc. of ASRU, Taipei, Taiwan, Dec. 2023
2023
-
[41]
ICASSP 2024 Speech Signal Improvement Challenge,
N. C. Ristea, A. Saabas, R. Cutler, B. Naderi, S. Braun, and S. Branets, “ICASSP 2024 Speech Signal Improvement Challenge,” inICASSP, Seoul, Korea, Apr. 2024
2024
-
[42]
P .862: Perceptual Evaluation of Speech Quality (PESQ), Inter- national Telecommunication Union, Telecommunication Standardization Sector (ITU-T), Feb
ITU,Rec. P .862: Perceptual Evaluation of Speech Quality (PESQ), Inter- national Telecommunication Union, Telecommunication Standardization Sector (ITU-T), Feb. 2001
2001
-
[43]
DNSMOS: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Sup- pressors,
C. K. A. Reddy, V . Gopal, and R. Cutler, “DNSMOS: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Sup- pressors,” inProc. of ICASSP, Toronto, ON, Canada, Jun. 2021, pp. 6493–6497
2021
-
[44]
NISQA: A Deep CNN- Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets,
G. Mittag, B. Naderi, A. Chehadi, and S. M ¨oller, “NISQA: A Deep CNN- Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets,” inProc. of Interspeech, Brno, Czech Republic, Aug. 2021, pp. 2127–2131
2021
-
[45]
An Algorithm for Predicting the Intelligibility of Speech Masked by Modulated Noise Maskers,
J. Jensen and C. H. Taal, “An Algorithm for Predicting the Intelligibility of Speech Masked by Modulated Noise Maskers,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 24, no. 11, pp. 2009– 2022, 2016
2009
-
[46]
URGENT 2025 Speech Enhancement Challenge Scoreboard,
“URGENT 2025 Speech Enhancement Challenge Scoreboard,” 2025, accessed: 2025-04-04. [Online]. Available: https://urgent-challenge.com/ competitions/13#final results
2025
-
[47]
A Comparison of Speech Intelligibility and Subjective Quality with Hearing-Aid Processing in Older Adults with Hearing Loss,
K. H. Arehart, S. H. Chon, E. M. H. Lundberg, L. O. H. Jr., J. M. Kates, M. C. Anderson, V . H. Rallapalli, and P. E. Souza, “A Comparison of Speech Intelligibility and Subjective Quality with Hearing-Aid Processing in Older Adults with Hearing Loss,”International Journal of A...
2022
-
[48]
On the Impact of Speech Intelligibility on Speech Quality in the Context of V oice over IP Telephony,
F. Schiffner, J. Skowronek, and A. Raake, “On the Impact of Speech Intelligibility on Speech Quality in the Context of V oice over IP Telephony,” inProc. of QoMEX, Singapore, Singapore, Sep. 2014, pp. 59–60
2014
-
[49]
Study on the Correlation Between Objective Evaluations and Subjective Speech Quality and Intelligibility,
H.-T. Chiang, K.-H. Hung, S.-W. Fu, H.-C. Kuo, M.-H. Tsai, and Y . Tsao, “Study on the Correlation Between Objective Evaluations and Subjective Speech Quality and Intelligibility,” inProc. of ASRU, Taipei, Taiwan, Dec. 2023, pp. 1–7
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.