REVIEW 2 major objections 6 minor 79 references
Modern single-channel speech enhancement does not improve SOTA ASR on real Dutch noisy speech; five of eight models still stay under 22% WER without it.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-15 00:02 UTC pith:ZHJNZEQ2
load-bearing objection Solid resource paper: new Dutch realistic test set plus a clean negative result that single-channel SE does not help modern ASR on it. the 2 major comments →
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On real Dutch semi-spontaneous speech recorded in public indoor noise, none of five single-channel speech-enhancement algorithms improves the word-error rate of eight state-of-the-art ASR systems; most significantly degrade it. At the same time, five of those ASR systems already achieve average WER below 22 percent on the raw recordings, showing that modern recognizers can be surprisingly robust without enhancement.
What carries the argument
DRES itself: a 1.5-hour multi-channel Dutch corpus of elicited (free-speech, picture-card, prompt-card) speech from 80 speakers in four real public buildings, used as a realistic test set that exposes the gap between synthetic noise mixtures and genuine acoustic conditions.
Load-bearing premise
That the five chosen single-channel enhancers and the eight off-the-shelf ASR models, run only on one microphone of a four-channel array at uncontrolled speaker distance, are representative enough of modern practice for the null enhancement result to generalise beyond this Dutch indoor setting.
What would settle it
A controlled re-run in which a multi-channel or jointly-trained enhancement front-end, or a larger set of SE algorithms, yields a statistically significant WER reduction on the same DRES utterances for at least one of the eight ASR models.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces DRES, a 1.5-hour Dutch semi-spontaneous speech corpus from 80 speakers recorded with a four-channel linear array in four public indoor environments with real background talkers and noise. Orthographic transcripts follow the Jasmin-CGN protocol. The authors evaluate five single-channel SE algorithms (spectral subtraction, spectral noise gating, MetricGAN-OKD, and two SGMSE+ checkpoints) via DNSMOS and eight off-the-shelf SOTA ASR systems (Google Chirp 3/Telephony, Azure, MMS, Whisper-large-v3/turbo, NeMo-nl, CGN-Conformer) via WER on the unenhanced baseline and after each SE method. On the baseline, five of eight ASR models achieve average WER below 22% (best: Google Chirp 3 at 11.2%). None of the SE methods improves WER for any ASR model; most produce statistically significant degradations (paired speaker-based bootstrap). The authors conclude that modern single-channel SE does not help SOTA ASR on this realistic Dutch data, in contrast to recent results on artificial English mixtures, and motivate multi-channel work.
Significance. Real noisy, semi-spontaneous Dutch speech with multi-person-checked transcripts is scarce; DRES fills a clear evaluation gap for both ASR and SE. The central negative result—that none of five widely used single-channel SE algorithms improves WER of eight contemporary ASR systems, while most degrade it—is carefully scoped, supported by Table 2, DNSMOS (Fig. 3), and speaker-based bootstrap tests with 95% CIs, and usefully contrasts with recent English artificial-mixture findings. Strengths include transparent corpus design, multi-location recording, multi-person transcription checking, and explicit future-work plans for multi-channel SE. If the null SE result holds under broader conditions, it has practical implications for whether single-channel SE should be inserted before modern E2E ASR in real indoor babble.
major comments (2)
- The central claim that modern single-channel SE fails to help SOTA ASR rests on a specific sample of five SE algorithms and eight ASR models evaluated only on channel 2 with uncontrolled 1–1.5 m distance (§3.1–3.2, Table 2, §4.2). While the measured null result on this corpus is solid, the manuscript should more explicitly bound the generalization claim (already partially acknowledged in §5) and, if space permits, add at least one multi-channel or alternative SE baseline, or a short analysis of residual noise/artifact types that may explain the WER degradation despite DNSMOS gains. Without that, the contrast with [40,41] remains suggestive rather than fully diagnostic of language vs. realism vs. algorithm class.
- Speaker-to-array distance is uncontrolled and only subjectively estimated (§2.2). Because distance affects SNR, reverberation, and the relative benefit of SE, the paper should report at least a coarse distance or level distribution (or a sensitivity check) so that readers can judge how representative the acoustic conditions are of the claimed 1–1.5 m scenario. This is load-bearing for interpreting both the absolute WERs and the SE null result.
minor comments (6)
- Table 2: report per-location WERs after SE (or at least note that the additional analysis found no location-wise improvements) so the reader can verify the claim that SE never helped at any site.
- Fig. 2 / Fig. 3: DNSMOS distributions are clear, but adding mean ± std or median values in the caption or a small table would aid quantitative comparison.
- §2.4: vocabulary size (2,842) and total speech duration after silence removal are useful; a short note on speaking rate or average utterance length would further characterize the semi-spontaneous material.
- Clarify whether amplitude normalization to 0.707 is applied before or after SE and whether it interacts with any of the SE implementations.
- Minor typographical consistency: “SotA” vs “SOTA”, and ensure all model names (e.g., Whisper-large-V3 vs Whisper-large-v3) match the Hugging Face / API identifiers used.
- Release statement: upon acceptance the corpus will be released; a short note on license and expected metadata (speaker demographics already in Table 1) would strengthen the contribution.
Circularity Check
No circularity: empirical evaluation of off-the-shelf ASR/SE models on a newly collected held-out Dutch test set; outcomes are measured, not defined by construction.
full rationale
DRES is introduced as a new 1.5-hour multi-speaker Dutch semi-spontaneous noisy-speech test corpus. The paper's central claims are (1) five of eight off-the-shelf SOTA ASR models achieve average WER < 22% on the unenhanced baseline and (2) none of five single-channel SE algorithms (SS, SNG, MetricGAN-OKD, SGMSE+ WSJ0-CHiME3, SGMSE+ Voicebank-Demand) improve WER, with most significantly degrading it (Table 2, bootstrap tests §3.3/§4.2). These quantities are obtained by running pre-trained/public checkpoints and classical algorithms on held-out DRES audio and scoring against independently created orthographic transcriptions with external metrics (WER, DNSMOS P.835). No parameters are fitted to DRES and then re-presented as predictions; no uniqueness theorem or ansatz is imported from the authors' prior work to force the null SE result; self-citations are limited to ordinary dataset/protocol references (Jasmin-CGN, CGN-Conformer) that do not define the reported WERs. The derivation chain is therefore self-contained empirical measurement, not a reduction of outputs to inputs by construction. Score 0 is the correct honest finding.
Axiom & Free-Parameter Ledger
free parameters (3)
- Amplitude normalization target (0.707) =
0.707
- Beam size for Whisper/NeMo/CGN-Conformer decoding =
10
- Bootstrap samples (10,000 speaker-based) =
10000
axioms (4)
- domain assumption Word error rate on orthographic transcripts (ignoring non-linguistic symbols) is a valid primary measure of ASR performance on semi-spontaneous Dutch.
- domain assumption DNSMOS P.835 is a sufficient no-reference proxy for speech quality of enhanced signals.
- ad hoc to paper The eight selected commercial and open ASR models and five SE algorithms adequately represent current SOTA single-channel practice for the claimed contrast with recent English work.
- ad hoc to paper Channel 2 of the four-channel array is representative for single-channel analysis of the recorded conditions.
invented entities (1)
-
DRES corpus
no independent evidence
read the original abstract
We present DRES: a 1.5-hour Dutch realistic elicited (semi-spontaneous) speech dataset from 80 speakers recorded in noisy, public indoor environments. DRES was designed as a test set for the evaluation of state-of-the-art (SotA) automatic speech recognition (ASR) and speech enhancement (SE) models in a real-world scenario: a person speaking in a public indoor space with background talkers and noise. The speech was recorded with a four-channel linear microphone array. In this work we evaluate the speech quality of five well-known single-channel SE algorithms and the recognition performance of eight SotA off-the-shelf ASR models before and after applying SE on the speech of DRES. We found that five out of the eight ASR models have WERs lower than 22\% on DRES, despite the challenging conditions. In contrast to recent work, we did not find a positive effect of modern single-channel SE on ASR performance, emphasizing the importance of evaluating in realistic conditions.
Reference graph
Works this paper leans on
-
[1]
Introduction Automatic speech recognition (ASR) systems are widely used in diverse scenarios, including emergency centres [1], health care services [2], and education [3]. To ensure reliable performance in practice, it is crucial that ASR systems are able to recognize speech from diverse speakers and under diverse acoustic conditions, including noisy envi...
Pith/arXiv arXiv 2026
-
[2]
Corpus design To elicit spontaneous speech and to ensure both lexical and pho- netic diversity, three tasks were designed:
The DRES corpus 2.1. Corpus design To elicit spontaneous speech and to ensure both lexical and pho- netic diversity, three tasks were designed:
-
[3]
Free speech:speakers were encouraged to talk freely on a topic of their own choice or chosen from a preselected list
-
[4]
Picture card:speakers were instructed to randomly select a picture from 26 pre-prepared picture cards and describe their card or tell a short story fitting that card
-
[5]
Tell something about [topic 1] or [topic 2]
Prompt card:speakers were instructed to randomly select a topic from 26 pre-prepared prompt cards and talk about it. The prompt and picture cards were created by the first author (D.G.). Picture cards were designed to have a dreamlike aesthetic and were made with GPT-4o using the Bing image generation interface [46]. An example is shown in Figure 1a. Figu...
-
[6]
The speech quality is particularly low for the recordings made at location Ahoy, highlighting the difficult acoustic conditions
(see Section 3.3), of the speech signals per recording loca- tion and averaged over the four locations. The speech quality is particularly low for the recordings made at location Ahoy, highlighting the difficult acoustic conditions. The quality is similar across the other three buildings. A Kruskal-Wallis test indicated that the DNSMOS differed over recor...
-
[7]
Noisereduce
Experiments 3.1. Speech enhancement algorithms We selected five single-channel SE algorithms based on ease-of-use, computational complexity, SE performance, and popularity. The baseline speech signal and the SE algorithms are described below. Asbaseline signal(Base) we use the recordings from micro- phone 2, which is at the center of the microphone-array ...
2026
-
[8]
Base” is identical to “Overall
Results 4.1. Speech quality The left panel of Figure 3 shows the DNSMOS score of the base signal (note that “Base” is identical to “Overall” in Figure 2), and for each of the SE algorithms. As the figure shows, the median MOS-score of each SE algorithm improves compared to the baseline signal. The largest improvement is achieved for both versions of SGMSE...
-
[9]
We evaluated several SOTA ASR models on speech from DRES and evaluated the effect of several classical and modern single-channel speech enhancement (SE) al- gorithms
Discussion and conclusions In this work we presented DRES: a dataset of Dutch elicited speech in realistic conditions. We evaluated several SOTA ASR models on speech from DRES and evaluated the effect of several classical and modern single-channel speech enhancement (SE) al- gorithms. We found that, without SE, two out of eight ASR mod- els, Google Chirp ...
-
[10]
Acknowledgments The authors thank Ilse Huisman and Ansen Weng for their help and support with constructing the ground-truth transcriptions
-
[11]
After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication
Generative AI Use Disclosure During the preparation of this work the authors used ChatGPT to improve language and readability. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication
-
[12]
Development of speech recognition systems in emergency call centers,
A. Valizada, N. Akhundova, and S. Rustamov, “Development of speech recognition systems in emergency call centers,”Symmetry, vol. 13, no. 4, p. 634, 2021
2021
-
[13]
A systematic review of speech recognition technology in health care,
M. Johnsonet al., “A systematic review of speech recognition technology in health care,”BMC medical informatics and decision making, vol. 14, no. 1, p. 94, 2014
2014
-
[14]
Evaluating automatic speech recognition-based language learning systems: A case study,
J. Van Doremalen, L. Boves, J. Colpaert, C. Cucchiarini, and H. Strik, “Evaluating automatic speech recognition-based language learning systems: A case study,”Computer Assisted Language Learning, vol. 29, no. 4, pp. 833–851, 2016
2016
-
[15]
Investigating RNN-based speech enhancement methods for noise- robust Text-to-Speech,
C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Investigating RNN-based speech enhancement methods for noise- robust Text-to-Speech,” in9th ISCA Workshop on Speech Synthesis Workshop (SSW 9), 2016, pp. 146–152
2016
-
[16]
Subjective evaluation and comparison of speech enhance- ment algorithms,
Y . Hu, “Subjective evaluation and comparison of speech enhance- ment algorithms,”Speech communication, vol. 49, pp. 588–601, 2007
2007
-
[17]
EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,
J. Richteret al., “EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,” in Interspeech 2024. ISCA, 2024, pp. 4873–4877
2024
-
[18]
The AURORA experimental frame- work for the performance evaluation of speech recognition systems under noisy conditions,
H.-G. Hirsch and D. Pearce, “The AURORA experimental frame- work for the performance evaluation of speech recognition systems under noisy conditions,” inAutomatic Speech Recognition: Chal- lenges for the New Millenium, 2000, pp. 181–188
2000
-
[19]
The PASCAL CHiME speech separation and recognition challenge,
J. Barker, E. Vincent, N. Ma, H. Christensen, and P. Green, “The PASCAL CHiME speech separation and recognition challenge,” Computer Speech & Language, vol. 27, no. 3, pp. 621–633, 2013
2013
-
[20]
The HIWIRE database, a noisy and non- native English speech corpus for cockpit communication,
J. Seguraet al., “The HIWIRE database, a noisy and non- native English speech corpus for cockpit communication,”Online. http://www. hiwire. org, vol. 8, 2007
2007
-
[21]
The evolution of the lombard effect: 100 years of psychoacoustic research,
H. Brumm and S. A. Zollinger, “The evolution of the lombard effect: 100 years of psychoacoustic research,”Behaviour, vol. 148, no. 11–13, pp. 1173–1198, 2011
2011
-
[22]
The lombard effect: a reflex to better communicate with others in noise,
J.-C. Junqua, S. Fincke, and K. Field, “The lombard effect: a reflex to better communicate with others in noise,” inIEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP). IEEE, 1999, pp. 2083–2086 vol.4
1999
-
[23]
Speech production modifications produced by competing talkers, babble, and stationary noise,
Y . Lu and M. Cooke, “Speech production modifications produced by competing talkers, babble, and stationary noise,”The Journal of the Acoustical Society of America, vol. 124, no. 5, pp. 3261–3275, Nov. 2008
2008
-
[24]
Influence of Sound Im- mersion and Communicative Interaction on the Lombard Effect,
M. Garnier, N. Henrich, and D. Dubois, “Influence of Sound Im- mersion and Communicative Interaction on the Lombard Effect,” Journal of Speech, Language, and Hearing Research, vol. 53, no. 3, pp. 588–608, Jun. 2010
2010
-
[25]
Under- standing Lombard speech: a review of compensation techniques towards improving speech based recognition systems,
S. Uma Maheswari, A. Shahina, and A. Nayeemulla Khan, “Under- standing Lombard speech: a review of compensation techniques towards improving speech based recognition systems,”Artificial Intelligence Review, vol. 54, pp. 2495–2523, Sep. 2020
2020
-
[26]
The third ‘chime’speech separation and recognition challenge: Dataset, task and baselines,
J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The third ‘chime’speech separation and recognition challenge: Dataset, task and baselines,” inIEEE workshop on automatic speech recognition and understanding (ASRU), 2015, pp. 504–511
2015
-
[27]
MISP-meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,
C. Hang, C.-H. H. Yang, J.-C. Gu, S. M. Siniscalchi, and J. Du, “MISP-meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,” inProc. of the 63rd Annual Meeting of the Association for Computational Linguistics, Jul. 2025, pp. 15 479–15 492
2025
-
[28]
One-pass single-channel noisy speech recognition using a combination of noisy and enhanced features,
M. Fujimoto and H. Kawai, “One-pass single-channel noisy speech recognition using a combination of noisy and enhanced features,” inInterspeech 2019. ISCA, Sep. 2019, pp. 486–490
2019
-
[29]
How bad are artifacts?: Analyzing the impact of speech enhancement errors on ASR,
K. Iwamotoet al., “How bad are artifacts?: Analyzing the impact of speech enhancement errors on ASR,” inInterspeech 2022. ISCA, Sep. 2022, pp. 5418–5422
2022
-
[30]
Noise robust automatic speech recognition: review and analysis,
M. Dua, Akanksha, and S. Dua, “Noise robust automatic speech recognition: review and analysis,”International Journal of Speech Technology, vol. 26, no. 2, pp. 475–519, Jun. 2023
2023
-
[31]
The Spoken Dutch Corpus. Overview and first evalu- ation,
N. Oostdijk, “The Spoken Dutch Corpus. Overview and first evalu- ation,” inProc. of the 2nd International Conference on Language Resources and Evaluation. Athens, Greece: European Language Resources Association, May 2000
2000
-
[32]
JASMIN-CGN: Extension of the spoken Dutch corpus with speech of elderly people, children and non-natives in the human-machine interaction modality,
C. Cucchiarini, H. Van hamme, O. van Herwijnen, and F. Smits, “JASMIN-CGN: Extension of the spoken Dutch corpus with speech of elderly people, children and non-natives in the human-machine interaction modality,” inProc. of the Fifth International Conference on Language Resources and Evaluation (LREC’06). Genoa, Italy: ELRA, May 2006
2006
-
[33]
Common voice: A massively-multilingual speech corpus,
R. Ardilaet al., “Common voice: A massively-multilingual speech corpus,”arXiv preprint arXiv:1912.06670, 2019
Pith/arXiv arXiv 1912
-
[34]
Robust speech recognition via large-scale weak super- vision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak super- vision,” inInternational conference on machine learning. PMLR, 2023, pp. 28 492–28 518
2023
-
[35]
Uncovering bias in asr systems: Evaluating wav2vec2 and whisper for dutch speakers,
M. Fuckner, S. Horsman, P. Wiggers, and I. Janssen, “Uncovering bias in asr systems: Evaluating wav2vec2 and whisper for dutch speakers,” inInternational Conference on Speech Technology and Human-Computer Dialogue. IEEE, 2023, pp. 146–151
2023
-
[36]
R. Raes, S. Lensink, and M. Pechenizkiy, “Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data,”arXiv preprint arXiv:2411.09431, 2024
Pith/arXiv arXiv 2024
-
[37]
wav2vec 2.0: A framework for self-supervised learning of speech representations,
A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems, vol. 33, pp. 12 449–12 460, 2020
2020
-
[38]
To- wards inclusive automatic speech recognition,
S. Feng, B. M. Halpern, O. Kudina, and O. Scharenborg, “To- wards inclusive automatic speech recognition,”Computer Speech & Language, vol. 84, p. 101567, 2024
2024
-
[39]
Google USM: Scaling Automatic Speech Recog- nition Beyond 100 Languages,
Y . Zhanget al., “Google USM: Scaling Automatic Speech Recog- nition Beyond 100 Languages,”arXiv preprint arXiv:2303.01037, 2023
Pith/arXiv arXiv 2023
-
[40]
Speech-to-text overview,
Microsoft, “Speech-to-text overview,” https://learn.microsoft.com/ en-us/azure/ai-services/speech-service/speech-to-text, 2026, ac- cessed: 2026-01-04
2026
-
[41]
Scaling speech technology to 1,000+ languages,
V . Pratapet al., “Scaling speech technology to 1,000+ languages,” Journal of Machine Learning Research, vol. 25, no. 97, pp. 1–52, 2024
2024
-
[42]
stt nl fastconformer hybrid large pc,
NVIDIA, “stt nl fastconformer hybrid large pc,” 2023. [Online]. Available: https://huggingface.co/nvidia/stt nl fastconformer hybrid large pc
2023
-
[43]
A new standard for speech recognition and translation from the NVIDIA nemo canary model,
——, “A new standard for speech recognition and translation from the NVIDIA nemo canary model,” 2024, accessed: 2026-01-04. [Online]. Available: https: //developer.nvidia.com/blog/new-standard-for-speech-\protect\ penalty-\@Mrecognition-and-translation-from-the-nvidia- \ protect\penalty-\@Mnemo-canary-model/
2024
-
[44]
YuanyuanZhang/DutchCGNConformerFBank: ES- Pnet2 ASR model,
Y . Zhang, “YuanyuanZhang/DutchCGNConformerFBank: ES- Pnet2 ASR model,” 2024. [Online]. Available: https: //huggingface.co/YuanyuanZhang/DutchCGNConformerFBank
2024
-
[45]
Far-field automatic speech recognition,
R. Haeb-Umbach, J. Heymann, L. Drude, S. Watanabe, M. Del- croix, and T. Nakatani, “Far-field automatic speech recognition,” Proceedings of the IEEE, vol. 109, no. 2, pp. 124–148, Feb. 2021
2021
-
[46]
Advances in microphone array processing and multichannel speech enhancement,
G. Huanget al., “Advances in microphone array processing and multichannel speech enhancement,” inIEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, Apr. 2025, pp. 1–5
2025
-
[47]
Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups,
G. Hintonet al., “Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups,” IEEE Signal Processing Magazine, vol. 29, no. 6, pp. 82–97, 2012
2012
-
[48]
Bridging the gap between monaural speech enhancement and recognition with distortion- independent acoustic modeling,
P. Wang, K. Tan, and D. L. Wang, “Bridging the gap between monaural speech enhancement and recognition with distortion- independent acoustic modeling,”IEEE/ACM Transactions on Au- dio, Speech, and Language Processing, vol. 28, pp. 39–48, 2020
2020
-
[49]
NAaLOSS: Re- thinking the Objective of Speech Enhancement,
K.-H. Ho, E.-L. Yu, J.-W. Hung, and B. Chen, “NAaLOSS: Re- thinking the Objective of Speech Enhancement,” inIEEE 33rd International Workshop on Machine Learning for Signal Process- ing (MLSP). IEEE, Sep. 2023, pp. 1–6
2023
-
[50]
Speech enhancement and dereverberation with diffusion-based generative models,
J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, “Speech enhancement and dereverberation with diffusion-based generative models,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 2351–2364, 2023
2023
-
[51]
StoRM: A Diffusion-Based Stochastic Regeneration Model for Speech En- hancement and Dereverberation,
J.-M. Lemercier, J. Richter, S. Welker, and T. Gerkmann, “StoRM: A Diffusion-Based Stochastic Regeneration Model for Speech En- hancement and Dereverberation,”IEEE/ACM Transactions on Au- dio, Speech, and Language Processing, vol. 31, pp. 2724–2737, 2023
2023
-
[52]
Latent-level enhancement with flow matching for robust automatic speech recognition,
D.-H. Yang and J.-H. Chang, “Latent-level enhancement with flow matching for robust automatic speech recognition,” 2026
2026
-
[53]
Evaluating speech enhancement performance across demographics and language,
J. Giraldo, A. Peir ´o-Lilja, C. Armentano-Oller, R. Zevallos, and C. Espa˜na-Bonet, “Evaluating speech enhancement performance across demographics and language,” inInterspeech 2025. ISCA, Aug. 2025, pp. 1353–1357
2025
-
[54]
Enhancement of speech corrupted by acoustic noise,
M. Berouti, R. Schwartz, and J. Makhoul, “Enhancement of speech corrupted by acoustic noise,” inIEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), vol. 4. IEEE, 1979, pp. 208–211
1979
-
[55]
Finding, visualizing, and quantifying latent structure across diverse animal vocal reper- toires,
T. Sainburg, M. Thielk, and T. Q. Gentner, “Finding, visualizing, and quantifying latent structure across diverse animal vocal reper- toires,”PLOS Computational Biology, vol. 16, no. 10, p. e1008228, Oct. 2020
2020
-
[56]
MetricGAN-OKD: Multi-metric optimization of MetricGAN via online knowledge distillation for speech enhancement,
W. Shin, B. H. Lee, J. S. Kim, H. J. Park, and S. W. Han, “MetricGAN-OKD: Multi-metric optimization of MetricGAN via online knowledge distillation for speech enhancement,” inProc. of the 40th International Conference on Machine Learning, vol. 202. PMLR, 23–29 Jul 2023, pp. 31 521–31 538
2023
-
[57]
[Online]
OpenAI, “Gpt-4o,” 2024, multimodal generative model, ac- cessed via Bing. [Online]. Available: https://openai.com/index/ gpt-4o-and-more-tools-to-chatgpt-free/
2024
-
[58]
Dnsmos p.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,
C. K. A. Reddy, V . Gopal, and R. Cutler, “Dnsmos p.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, May 2022, pp. 886–890
2022
-
[59]
Suppression of acoustic noise in speech using spectral subtraction,
S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,”IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 27, no. 2, pp. 113–120, Apr. 1979
1979
-
[60]
Speech enhancement—a review of modern methods,
D. O’Shaughnessy, “Speech enhancement—a review of modern methods,”IEEE Transactions on Human-Machine Systems, vol. 54, no. 1, pp. 110–120, Feb. 2024
2024
-
[61]
A perceptual masking approach for noise robust speech recognition,
H. K. Maganti and M. Matassoni, “A perceptual masking approach for noise robust speech recognition,”EURASIP Journal on Audio, Speech, and Music Processing, vol. 2012, no. 1, Dec. 2012
2012
-
[62]
Pyroomacoustics: A python package for audio room simulation and array processing al- gorithms,
R. Scheibler, E. Bezzam, and I. Dokmanic, “Pyroomacoustics: A python package for audio room simulation and array processing al- gorithms,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, Apr. 2018, pp. 351–355
2018
-
[63]
Toward a computational neu- roethology of vocal communication: From bioacoustics to neu- rophysiology, emerging tools and future directions,
T. Sainburg and T. Q. Gentner, “Toward a computational neu- roethology of vocal communication: From bioacoustics to neu- rophysiology, emerging tools and future directions,”Frontiers in Behavioral Neuroscience, vol. 15, Dec. 2021
2021
-
[64]
Domain general noise reduction for time series signals with noisereduce,
T. Sainburg and A. Zorea, “Domain general noise reduction for time series signals with noisereduce,”Scientific Reports, vol. 15, no. 1, Aug. 2025
2025
-
[65]
Objective and subjective evaluation of diffusion-based speech en- hancement for dysarthric speech,
D. de Groot, T. Patel, D. Kayande, O. Scharenborg, and Z. Yue, “Objective and subjective evaluation of diffusion-based speech en- hancement for dysarthric speech,” inInterspeech 2025. ISCA, Aug. 2025, pp. 2740–2744
2025
-
[66]
Tim Sainburg et al., “timsainb/noisereduce: 3.0.3,” 2024. [Online]. Available: doi.org/10.5281/ZENODO.12688680
-
[67]
Speech Enhancement and Dereverberation with Diffusion-based Generative Models,
S. Welker, J.-M. Lemercier, and J. Richter, “Speech Enhancement and Dereverberation with Diffusion-based Generative Models,”
-
[68]
Available: https://github.com/sp-uhh/sgmse
[Online]. Available: https://github.com/sp-uhh/sgmse
-
[69]
Speech-to-text: Transcription models,
Google Cloud, “Speech-to-text: Transcription models,” https:// cloud.google.com/speech-to-text/docs/transcription-model, 2026, accessed: 2026-02-18
2026
-
[70]
facebook/mms-1b-all: Massively Multilingual Speech (MMS),
Meta AI, “facebook/mms-1b-all: Massively Multilingual Speech (MMS),” 2023. [Online]. Available: https://huggingface.co/ facebook/mms-1b-all
2023
-
[71]
Whisper-large-v3,
OpenAI, “Whisper-large-v3,” 2023. [Online]. Available: https: //huggingface.co/openai/whisper-large-v3
2023
-
[72]
Fleurs: Few-shot learning evaluation of universal representations of speech,
A. Conneau, M. Ma, S. Khanuja, Y . Zhang, V . Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna, “Fleurs: Few-shot learning evaluation of universal representations of speech,” arXiv preprint arXiv:2205.12446, 2022. [Online]. Available: https://arxiv.org/abs/2205.12446
Pith/arXiv arXiv 2022
-
[73]
Whisper github discussion #2363,
OpenAI, “Whisper github discussion #2363,” 2024. [Online]. Available: https://github.com/openai/whisper/discussions/2363
2024
-
[74]
Mls: A large-scale multilingual dataset for speech research,
V . Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “Mls: A large-scale multilingual dataset for speech research,”arXiv preprint arXiv:2012.03411, 2020
Pith/arXiv arXiv 2012
-
[75]
C. Wanget al., “V oxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,”arXiv preprint arXiv:2101.00390, 2021
Pith/arXiv arXiv 2021
-
[76]
Bootstrap estimates for confidence intervals in asr performance evaluation,
M. Bisani and H. Ney, “Bootstrap estimates for confidence intervals in asr performance evaluation,” inIEEE ICASSP, vol. 1, 2004, pp. I–409
2004
-
[77]
Good practices for evaluation of machine learning systems,
L. Ferrer, O. Scharenborg, and T. B ¨ackstr¨om, “Good practices for evaluation of machine learning systems,”arXiv preprint arXiv:2412.03700, 2024
Pith/arXiv arXiv 2024
-
[78]
D. D. Boos and L. A. Stefanski,Essential Statistical Inference: Theory and Methods. Springer, 2013, section 11.6: Bootstrap Resampling for Hypothesis Tests
2013
-
[79]
ITU-T Recommenda- tion P.835: Subjective test methodology for evaluating speech communication systems that include noise suppression algorithm,
International Telecommunication Union, “ITU-T Recommenda- tion P.835: Subjective test methodology for evaluating speech communication systems that include noise suppression algorithm,” https://www.itu.int/rec/T-REC-P.835/en, June 2021
2021
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.