Pith. sign in

REVIEW 2 major objections 6 minor 79 references

Modern single-channel speech enhancement does not improve SOTA ASR on real Dutch noisy speech; five of eight models still stay under 22% WER without it.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-15 00:02 UTC pith:ZHJNZEQ2

load-bearing objection Solid resource paper: new Dutch realistic test set plus a clean negative result that single-channel SE does not help modern ASR on it. the 2 major comments →

arxiv 2603.09725 v2 pith:ZHJNZEQ2 submitted 2026-03-10 eess.AS

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition

classification eess.AS
keywords Dutch speechrealistic noisy speechsemi-spontaneous speechspeech recognitionspeech enhancementword error rateDNSMOSmicrophone array
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Most speech-recognition and enhancement research still relies on clean speech mixed with artificial noise. This paper asks what happens when the same state-of-the-art systems are tested on real people speaking Dutch in busy public indoor spaces with genuine background talkers. The authors release DRES, a 1.5-hour semi-spontaneous corpus from 80 speakers recorded with a four-channel microphone array in four different buildings. On the unenhanced recordings, five of eight off-the-shelf ASR systems already achieve average word-error rates below 22 percent. When five well-known single-channel enhancement algorithms (classical and modern) are applied first, none of them improves recognition accuracy; most make it worse even though they raise objective speech-quality scores. The central message is that results obtained on synthetic English mixtures do not automatically transfer to realistic multilingual conditions, so evaluation on genuine noisy speech remains essential.

Core claim

On real Dutch semi-spontaneous speech recorded in public indoor noise, none of five single-channel speech-enhancement algorithms improves the word-error rate of eight state-of-the-art ASR systems; most significantly degrade it. At the same time, five of those ASR systems already achieve average WER below 22 percent on the raw recordings, showing that modern recognizers can be surprisingly robust without enhancement.

What carries the argument

DRES itself: a 1.5-hour multi-channel Dutch corpus of elicited (free-speech, picture-card, prompt-card) speech from 80 speakers in four real public buildings, used as a realistic test set that exposes the gap between synthetic noise mixtures and genuine acoustic conditions.

Load-bearing premise

That the five chosen single-channel enhancers and the eight off-the-shelf ASR models, run only on one microphone of a four-channel array at uncontrolled speaker distance, are representative enough of modern practice for the null enhancement result to generalise beyond this Dutch indoor setting.

What would settle it

A controlled re-run in which a multi-channel or jointly-trained enhancement front-end, or a larger set of SE algorithms, yields a statistically significant WER reduction on the same DRES utterances for at least one of the eight ASR models.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper introduces DRES, a 1.5-hour Dutch semi-spontaneous speech corpus from 80 speakers recorded with a four-channel linear array in four public indoor environments with real background talkers and noise. Orthographic transcripts follow the Jasmin-CGN protocol. The authors evaluate five single-channel SE algorithms (spectral subtraction, spectral noise gating, MetricGAN-OKD, and two SGMSE+ checkpoints) via DNSMOS and eight off-the-shelf SOTA ASR systems (Google Chirp 3/Telephony, Azure, MMS, Whisper-large-v3/turbo, NeMo-nl, CGN-Conformer) via WER on the unenhanced baseline and after each SE method. On the baseline, five of eight ASR models achieve average WER below 22% (best: Google Chirp 3 at 11.2%). None of the SE methods improves WER for any ASR model; most produce statistically significant degradations (paired speaker-based bootstrap). The authors conclude that modern single-channel SE does not help SOTA ASR on this realistic Dutch data, in contrast to recent results on artificial English mixtures, and motivate multi-channel work.

Significance. Real noisy, semi-spontaneous Dutch speech with multi-person-checked transcripts is scarce; DRES fills a clear evaluation gap for both ASR and SE. The central negative result—that none of five widely used single-channel SE algorithms improves WER of eight contemporary ASR systems, while most degrade it—is carefully scoped, supported by Table 2, DNSMOS (Fig. 3), and speaker-based bootstrap tests with 95% CIs, and usefully contrasts with recent English artificial-mixture findings. Strengths include transparent corpus design, multi-location recording, multi-person transcription checking, and explicit future-work plans for multi-channel SE. If the null SE result holds under broader conditions, it has practical implications for whether single-channel SE should be inserted before modern E2E ASR in real indoor babble.

major comments (2)
  1. The central claim that modern single-channel SE fails to help SOTA ASR rests on a specific sample of five SE algorithms and eight ASR models evaluated only on channel 2 with uncontrolled 1–1.5 m distance (§3.1–3.2, Table 2, §4.2). While the measured null result on this corpus is solid, the manuscript should more explicitly bound the generalization claim (already partially acknowledged in §5) and, if space permits, add at least one multi-channel or alternative SE baseline, or a short analysis of residual noise/artifact types that may explain the WER degradation despite DNSMOS gains. Without that, the contrast with [40,41] remains suggestive rather than fully diagnostic of language vs. realism vs. algorithm class.
  2. Speaker-to-array distance is uncontrolled and only subjectively estimated (§2.2). Because distance affects SNR, reverberation, and the relative benefit of SE, the paper should report at least a coarse distance or level distribution (or a sensitivity check) so that readers can judge how representative the acoustic conditions are of the claimed 1–1.5 m scenario. This is load-bearing for interpreting both the absolute WERs and the SE null result.
minor comments (6)
  1. Table 2: report per-location WERs after SE (or at least note that the additional analysis found no location-wise improvements) so the reader can verify the claim that SE never helped at any site.
  2. Fig. 2 / Fig. 3: DNSMOS distributions are clear, but adding mean ± std or median values in the caption or a small table would aid quantitative comparison.
  3. §2.4: vocabulary size (2,842) and total speech duration after silence removal are useful; a short note on speaking rate or average utterance length would further characterize the semi-spontaneous material.
  4. Clarify whether amplitude normalization to 0.707 is applied before or after SE and whether it interacts with any of the SE implementations.
  5. Minor typographical consistency: “SotA” vs “SOTA”, and ensure all model names (e.g., Whisper-large-V3 vs Whisper-large-v3) match the Hugging Face / API identifiers used.
  6. Release statement: upon acceptance the corpus will be released; a short note on license and expected metadata (speaker demographics already in Table 1) would strengthen the contribution.

Circularity Check

0 steps flagged

No circularity: empirical evaluation of off-the-shelf ASR/SE models on a newly collected held-out Dutch test set; outcomes are measured, not defined by construction.

full rationale

DRES is introduced as a new 1.5-hour multi-speaker Dutch semi-spontaneous noisy-speech test corpus. The paper's central claims are (1) five of eight off-the-shelf SOTA ASR models achieve average WER < 22% on the unenhanced baseline and (2) none of five single-channel SE algorithms (SS, SNG, MetricGAN-OKD, SGMSE+ WSJ0-CHiME3, SGMSE+ Voicebank-Demand) improve WER, with most significantly degrading it (Table 2, bootstrap tests §3.3/§4.2). These quantities are obtained by running pre-trained/public checkpoints and classical algorithms on held-out DRES audio and scoring against independently created orthographic transcriptions with external metrics (WER, DNSMOS P.835). No parameters are fitted to DRES and then re-presented as predictions; no uniqueness theorem or ansatz is imported from the authors' prior work to force the null SE result; self-citations are limited to ordinary dataset/protocol references (Jasmin-CGN, CGN-Conformer) that do not define the reported WERs. The derivation chain is therefore self-contained empirical measurement, not a reduction of outputs to inputs by construction. Score 0 is the correct honest finding.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 1 invented entities

The paper is an empirical corpus-and-benchmark study. It rests on standard speech-processing assumptions (WER as ASR metric, DNSMOS as quality proxy, off-the-shelf models as SOTA representatives) and on design choices for elicitation and recording. No free parameters are fitted to produce the central WER claims; no new physical entities are postulated.

free parameters (3)
  • Amplitude normalization target (0.707) = 0.707
    Hand-chosen peak level applied to all utterances before SE/ASR to avoid clipping; affects absolute levels but is conventional.
  • Beam size for Whisper/NeMo/CGN-Conformer decoding = 10
    Fixed to 10 for three models; decoding hyperparameter that can move WER slightly.
  • Bootstrap samples (10,000 speaker-based) = 10000
    Number of resamples for significance tests; conventional but chosen by authors.
axioms (4)
  • domain assumption Word error rate on orthographic transcripts (ignoring non-linguistic symbols) is a valid primary measure of ASR performance on semi-spontaneous Dutch.
    Standard ASR evaluation practice invoked throughout Section 3.3 and Table 2.
  • domain assumption DNSMOS P.835 is a sufficient no-reference proxy for speech quality of enhanced signals.
    Used as the sole objective quality metric in Sections 2.5 and 4.1; trained on ITU-T P.835 ratings but not validated with new human listening on DRES.
  • ad hoc to paper The eight selected commercial and open ASR models and five SE algorithms adequately represent current SOTA single-channel practice for the claimed contrast with recent English work.
    Selection justified by popularity and prior use (Section 3) but is a finite convenience sample that underpins the generalization in the abstract and Section 5.
  • ad hoc to paper Channel 2 of the four-channel array is representative for single-channel analysis of the recorded conditions.
    Explicit choice in Section 2.5; multi-channel information is recorded but unused for the reported SE/ASR results.
invented entities (1)
  • DRES corpus no independent evidence
    purpose: Provide a realistic Dutch semi-spontaneous multi-channel test set for ASR and SE evaluation under public indoor noise.
    New dataset constructed by the authors; independent evidence will exist once released and used by others, but at submission it is defined by this paper.

pith-pipeline@v1.1.0-grok45 · 17226 in / 3104 out tokens · 27586 ms · 2026-07-15T00:02:48.881195+00:00 · methodology

0 comments
read the original abstract

We present DRES: a 1.5-hour Dutch realistic elicited (semi-spontaneous) speech dataset from 80 speakers recorded in noisy, public indoor environments. DRES was designed as a test set for the evaluation of state-of-the-art (SotA) automatic speech recognition (ASR) and speech enhancement (SE) models in a real-world scenario: a person speaking in a public indoor space with background talkers and noise. The speech was recorded with a four-channel linear microphone array. In this work we evaluate the speech quality of five well-known single-channel SE algorithms and the recognition performance of eight SotA off-the-shelf ASR models before and after applying SE on the speech of DRES. We found that five out of the eight ASR models have WERs lower than 22\% on DRES, despite the challenging conditions. In contrast to recent work, we did not find a positive effect of modern single-channel SE on ASR performance, emphasizing the importance of evaluating in realistic conditions.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

79 extracted references · 1 canonical work pages

  1. [1]

    Introduction Automatic speech recognition (ASR) systems are widely used in diverse scenarios, including emergency centres [1], health care services [2], and education [3]. To ensure reliable performance in practice, it is crucial that ASR systems are able to recognize speech from diverse speakers and under diverse acoustic conditions, including noisy envi...

  2. [2]

    Corpus design To elicit spontaneous speech and to ensure both lexical and pho- netic diversity, three tasks were designed:

    The DRES corpus 2.1. Corpus design To elicit spontaneous speech and to ensure both lexical and pho- netic diversity, three tasks were designed:

  3. [3]

    Free speech:speakers were encouraged to talk freely on a topic of their own choice or chosen from a preselected list

  4. [4]

    Picture card:speakers were instructed to randomly select a picture from 26 pre-prepared picture cards and describe their card or tell a short story fitting that card

  5. [5]

    Tell something about [topic 1] or [topic 2]

    Prompt card:speakers were instructed to randomly select a topic from 26 pre-prepared prompt cards and talk about it. The prompt and picture cards were created by the first author (D.G.). Picture cards were designed to have a dreamlike aesthetic and were made with GPT-4o using the Bing image generation interface [46]. An example is shown in Figure 1a. Figu...

  6. [6]

    The speech quality is particularly low for the recordings made at location Ahoy, highlighting the difficult acoustic conditions

    (see Section 3.3), of the speech signals per recording loca- tion and averaged over the four locations. The speech quality is particularly low for the recordings made at location Ahoy, highlighting the difficult acoustic conditions. The quality is similar across the other three buildings. A Kruskal-Wallis test indicated that the DNSMOS differed over recor...

  7. [7]

    Noisereduce

    Experiments 3.1. Speech enhancement algorithms We selected five single-channel SE algorithms based on ease-of-use, computational complexity, SE performance, and popularity. The baseline speech signal and the SE algorithms are described below. Asbaseline signal(Base) we use the recordings from micro- phone 2, which is at the center of the microphone-array ...

  8. [8]

    Base” is identical to “Overall

    Results 4.1. Speech quality The left panel of Figure 3 shows the DNSMOS score of the base signal (note that “Base” is identical to “Overall” in Figure 2), and for each of the SE algorithms. As the figure shows, the median MOS-score of each SE algorithm improves compared to the baseline signal. The largest improvement is achieved for both versions of SGMSE...

  9. [9]

    We evaluated several SOTA ASR models on speech from DRES and evaluated the effect of several classical and modern single-channel speech enhancement (SE) al- gorithms

    Discussion and conclusions In this work we presented DRES: a dataset of Dutch elicited speech in realistic conditions. We evaluated several SOTA ASR models on speech from DRES and evaluated the effect of several classical and modern single-channel speech enhancement (SE) al- gorithms. We found that, without SE, two out of eight ASR mod- els, Google Chirp ...

  10. [10]

    Acknowledgments The authors thank Ilse Huisman and Ansen Weng for their help and support with constructing the ground-truth transcriptions

  11. [11]

    After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication

    Generative AI Use Disclosure During the preparation of this work the authors used ChatGPT to improve language and readability. After using this tool, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication

  12. [12]

    Development of speech recognition systems in emergency call centers,

    A. Valizada, N. Akhundova, and S. Rustamov, “Development of speech recognition systems in emergency call centers,”Symmetry, vol. 13, no. 4, p. 634, 2021

  13. [13]

    A systematic review of speech recognition technology in health care,

    M. Johnsonet al., “A systematic review of speech recognition technology in health care,”BMC medical informatics and decision making, vol. 14, no. 1, p. 94, 2014

  14. [14]

    Evaluating automatic speech recognition-based language learning systems: A case study,

    J. Van Doremalen, L. Boves, J. Colpaert, C. Cucchiarini, and H. Strik, “Evaluating automatic speech recognition-based language learning systems: A case study,”Computer Assisted Language Learning, vol. 29, no. 4, pp. 833–851, 2016

  15. [15]

    Investigating RNN-based speech enhancement methods for noise- robust Text-to-Speech,

    C. Valentini-Botinhao, X. Wang, S. Takaki, and J. Yamagishi, “Investigating RNN-based speech enhancement methods for noise- robust Text-to-Speech,” in9th ISCA Workshop on Speech Synthesis Workshop (SSW 9), 2016, pp. 146–152

  16. [16]

    Subjective evaluation and comparison of speech enhance- ment algorithms,

    Y . Hu, “Subjective evaluation and comparison of speech enhance- ment algorithms,”Speech communication, vol. 49, pp. 588–601, 2007

  17. [17]

    EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,

    J. Richteret al., “EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,” in Interspeech 2024. ISCA, 2024, pp. 4873–4877

  18. [18]

    The AURORA experimental frame- work for the performance evaluation of speech recognition systems under noisy conditions,

    H.-G. Hirsch and D. Pearce, “The AURORA experimental frame- work for the performance evaluation of speech recognition systems under noisy conditions,” inAutomatic Speech Recognition: Chal- lenges for the New Millenium, 2000, pp. 181–188

  19. [19]

    The PASCAL CHiME speech separation and recognition challenge,

    J. Barker, E. Vincent, N. Ma, H. Christensen, and P. Green, “The PASCAL CHiME speech separation and recognition challenge,” Computer Speech & Language, vol. 27, no. 3, pp. 621–633, 2013

  20. [20]

    The HIWIRE database, a noisy and non- native English speech corpus for cockpit communication,

    J. Seguraet al., “The HIWIRE database, a noisy and non- native English speech corpus for cockpit communication,”Online. http://www. hiwire. org, vol. 8, 2007

  21. [21]

    The evolution of the lombard effect: 100 years of psychoacoustic research,

    H. Brumm and S. A. Zollinger, “The evolution of the lombard effect: 100 years of psychoacoustic research,”Behaviour, vol. 148, no. 11–13, pp. 1173–1198, 2011

  22. [22]

    The lombard effect: a reflex to better communicate with others in noise,

    J.-C. Junqua, S. Fincke, and K. Field, “The lombard effect: a reflex to better communicate with others in noise,” inIEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP). IEEE, 1999, pp. 2083–2086 vol.4

  23. [23]

    Speech production modifications produced by competing talkers, babble, and stationary noise,

    Y . Lu and M. Cooke, “Speech production modifications produced by competing talkers, babble, and stationary noise,”The Journal of the Acoustical Society of America, vol. 124, no. 5, pp. 3261–3275, Nov. 2008

  24. [24]

    Influence of Sound Im- mersion and Communicative Interaction on the Lombard Effect,

    M. Garnier, N. Henrich, and D. Dubois, “Influence of Sound Im- mersion and Communicative Interaction on the Lombard Effect,” Journal of Speech, Language, and Hearing Research, vol. 53, no. 3, pp. 588–608, Jun. 2010

  25. [25]

    Under- standing Lombard speech: a review of compensation techniques towards improving speech based recognition systems,

    S. Uma Maheswari, A. Shahina, and A. Nayeemulla Khan, “Under- standing Lombard speech: a review of compensation techniques towards improving speech based recognition systems,”Artificial Intelligence Review, vol. 54, pp. 2495–2523, Sep. 2020

  26. [26]

    The third ‘chime’speech separation and recognition challenge: Dataset, task and baselines,

    J. Barker, R. Marxer, E. Vincent, and S. Watanabe, “The third ‘chime’speech separation and recognition challenge: Dataset, task and baselines,” inIEEE workshop on automatic speech recognition and understanding (ASRU), 2015, pp. 504–511

  27. [27]

    MISP-meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,

    C. Hang, C.-H. H. Yang, J.-C. Gu, S. M. Siniscalchi, and J. Du, “MISP-meeting: A real-world dataset with multimodal cues for long-form meeting transcription and summarization,” inProc. of the 63rd Annual Meeting of the Association for Computational Linguistics, Jul. 2025, pp. 15 479–15 492

  28. [28]

    One-pass single-channel noisy speech recognition using a combination of noisy and enhanced features,

    M. Fujimoto and H. Kawai, “One-pass single-channel noisy speech recognition using a combination of noisy and enhanced features,” inInterspeech 2019. ISCA, Sep. 2019, pp. 486–490

  29. [29]

    How bad are artifacts?: Analyzing the impact of speech enhancement errors on ASR,

    K. Iwamotoet al., “How bad are artifacts?: Analyzing the impact of speech enhancement errors on ASR,” inInterspeech 2022. ISCA, Sep. 2022, pp. 5418–5422

  30. [30]

    Noise robust automatic speech recognition: review and analysis,

    M. Dua, Akanksha, and S. Dua, “Noise robust automatic speech recognition: review and analysis,”International Journal of Speech Technology, vol. 26, no. 2, pp. 475–519, Jun. 2023

  31. [31]

    The Spoken Dutch Corpus. Overview and first evalu- ation,

    N. Oostdijk, “The Spoken Dutch Corpus. Overview and first evalu- ation,” inProc. of the 2nd International Conference on Language Resources and Evaluation. Athens, Greece: European Language Resources Association, May 2000

  32. [32]

    JASMIN-CGN: Extension of the spoken Dutch corpus with speech of elderly people, children and non-natives in the human-machine interaction modality,

    C. Cucchiarini, H. Van hamme, O. van Herwijnen, and F. Smits, “JASMIN-CGN: Extension of the spoken Dutch corpus with speech of elderly people, children and non-natives in the human-machine interaction modality,” inProc. of the Fifth International Conference on Language Resources and Evaluation (LREC’06). Genoa, Italy: ELRA, May 2006

  33. [33]

    Common voice: A massively-multilingual speech corpus,

    R. Ardilaet al., “Common voice: A massively-multilingual speech corpus,”arXiv preprint arXiv:1912.06670, 2019

  34. [34]

    Robust speech recognition via large-scale weak super- vision,

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak super- vision,” inInternational conference on machine learning. PMLR, 2023, pp. 28 492–28 518

  35. [35]

    Uncovering bias in asr systems: Evaluating wav2vec2 and whisper for dutch speakers,

    M. Fuckner, S. Horsman, P. Wiggers, and I. Janssen, “Uncovering bias in asr systems: Evaluating wav2vec2 and whisper for dutch speakers,” inInternational Conference on Speech Technology and Human-Computer Dialogue. IEEE, 2023, pp. 146–151

  36. [36]

    Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data,

    R. Raes, S. Lensink, and M. Pechenizkiy, “Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data,”arXiv preprint arXiv:2411.09431, 2024

  37. [37]

    wav2vec 2.0: A framework for self-supervised learning of speech representations,

    A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Advances in neural information processing systems, vol. 33, pp. 12 449–12 460, 2020

  38. [38]

    To- wards inclusive automatic speech recognition,

    S. Feng, B. M. Halpern, O. Kudina, and O. Scharenborg, “To- wards inclusive automatic speech recognition,”Computer Speech & Language, vol. 84, p. 101567, 2024

  39. [39]

    Google USM: Scaling Automatic Speech Recog- nition Beyond 100 Languages,

    Y . Zhanget al., “Google USM: Scaling Automatic Speech Recog- nition Beyond 100 Languages,”arXiv preprint arXiv:2303.01037, 2023

  40. [40]

    Speech-to-text overview,

    Microsoft, “Speech-to-text overview,” https://learn.microsoft.com/ en-us/azure/ai-services/speech-service/speech-to-text, 2026, ac- cessed: 2026-01-04

  41. [41]

    Scaling speech technology to 1,000+ languages,

    V . Pratapet al., “Scaling speech technology to 1,000+ languages,” Journal of Machine Learning Research, vol. 25, no. 97, pp. 1–52, 2024

  42. [42]

    stt nl fastconformer hybrid large pc,

    NVIDIA, “stt nl fastconformer hybrid large pc,” 2023. [Online]. Available: https://huggingface.co/nvidia/stt nl fastconformer hybrid large pc

  43. [43]

    A new standard for speech recognition and translation from the NVIDIA nemo canary model,

    ——, “A new standard for speech recognition and translation from the NVIDIA nemo canary model,” 2024, accessed: 2026-01-04. [Online]. Available: https: //developer.nvidia.com/blog/new-standard-for-speech-\protect\ penalty-\@Mrecognition-and-translation-from-the-nvidia- \ protect\penalty-\@Mnemo-canary-model/

  44. [44]

    YuanyuanZhang/DutchCGNConformerFBank: ES- Pnet2 ASR model,

    Y . Zhang, “YuanyuanZhang/DutchCGNConformerFBank: ES- Pnet2 ASR model,” 2024. [Online]. Available: https: //huggingface.co/YuanyuanZhang/DutchCGNConformerFBank

  45. [45]

    Far-field automatic speech recognition,

    R. Haeb-Umbach, J. Heymann, L. Drude, S. Watanabe, M. Del- croix, and T. Nakatani, “Far-field automatic speech recognition,” Proceedings of the IEEE, vol. 109, no. 2, pp. 124–148, Feb. 2021

  46. [46]

    Advances in microphone array processing and multichannel speech enhancement,

    G. Huanget al., “Advances in microphone array processing and multichannel speech enhancement,” inIEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, Apr. 2025, pp. 1–5

  47. [47]

    Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups,

    G. Hintonet al., “Deep Neural Networks for Acoustic Modeling in Speech Recognition: The Shared Views of Four Research Groups,” IEEE Signal Processing Magazine, vol. 29, no. 6, pp. 82–97, 2012

  48. [48]

    Bridging the gap between monaural speech enhancement and recognition with distortion- independent acoustic modeling,

    P. Wang, K. Tan, and D. L. Wang, “Bridging the gap between monaural speech enhancement and recognition with distortion- independent acoustic modeling,”IEEE/ACM Transactions on Au- dio, Speech, and Language Processing, vol. 28, pp. 39–48, 2020

  49. [49]

    NAaLOSS: Re- thinking the Objective of Speech Enhancement,

    K.-H. Ho, E.-L. Yu, J.-W. Hung, and B. Chen, “NAaLOSS: Re- thinking the Objective of Speech Enhancement,” inIEEE 33rd International Workshop on Machine Learning for Signal Process- ing (MLSP). IEEE, Sep. 2023, pp. 1–6

  50. [50]

    Speech enhancement and dereverberation with diffusion-based generative models,

    J. Richter, S. Welker, J.-M. Lemercier, B. Lay, and T. Gerkmann, “Speech enhancement and dereverberation with diffusion-based generative models,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 31, pp. 2351–2364, 2023

  51. [51]

    StoRM: A Diffusion-Based Stochastic Regeneration Model for Speech En- hancement and Dereverberation,

    J.-M. Lemercier, J. Richter, S. Welker, and T. Gerkmann, “StoRM: A Diffusion-Based Stochastic Regeneration Model for Speech En- hancement and Dereverberation,”IEEE/ACM Transactions on Au- dio, Speech, and Language Processing, vol. 31, pp. 2724–2737, 2023

  52. [52]

    Latent-level enhancement with flow matching for robust automatic speech recognition,

    D.-H. Yang and J.-H. Chang, “Latent-level enhancement with flow matching for robust automatic speech recognition,” 2026

  53. [53]

    Evaluating speech enhancement performance across demographics and language,

    J. Giraldo, A. Peir ´o-Lilja, C. Armentano-Oller, R. Zevallos, and C. Espa˜na-Bonet, “Evaluating speech enhancement performance across demographics and language,” inInterspeech 2025. ISCA, Aug. 2025, pp. 1353–1357

  54. [54]

    Enhancement of speech corrupted by acoustic noise,

    M. Berouti, R. Schwartz, and J. Makhoul, “Enhancement of speech corrupted by acoustic noise,” inIEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), vol. 4. IEEE, 1979, pp. 208–211

  55. [55]

    Finding, visualizing, and quantifying latent structure across diverse animal vocal reper- toires,

    T. Sainburg, M. Thielk, and T. Q. Gentner, “Finding, visualizing, and quantifying latent structure across diverse animal vocal reper- toires,”PLOS Computational Biology, vol. 16, no. 10, p. e1008228, Oct. 2020

  56. [56]

    MetricGAN-OKD: Multi-metric optimization of MetricGAN via online knowledge distillation for speech enhancement,

    W. Shin, B. H. Lee, J. S. Kim, H. J. Park, and S. W. Han, “MetricGAN-OKD: Multi-metric optimization of MetricGAN via online knowledge distillation for speech enhancement,” inProc. of the 40th International Conference on Machine Learning, vol. 202. PMLR, 23–29 Jul 2023, pp. 31 521–31 538

  57. [57]

    [Online]

    OpenAI, “Gpt-4o,” 2024, multimodal generative model, ac- cessed via Bing. [Online]. Available: https://openai.com/index/ gpt-4o-and-more-tools-to-chatgpt-free/

  58. [58]

    Dnsmos p.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,

    C. K. A. Reddy, V . Gopal, and R. Cutler, “Dnsmos p.835: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, May 2022, pp. 886–890

  59. [59]

    Suppression of acoustic noise in speech using spectral subtraction,

    S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,”IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 27, no. 2, pp. 113–120, Apr. 1979

  60. [60]

    Speech enhancement—a review of modern methods,

    D. O’Shaughnessy, “Speech enhancement—a review of modern methods,”IEEE Transactions on Human-Machine Systems, vol. 54, no. 1, pp. 110–120, Feb. 2024

  61. [61]

    A perceptual masking approach for noise robust speech recognition,

    H. K. Maganti and M. Matassoni, “A perceptual masking approach for noise robust speech recognition,”EURASIP Journal on Audio, Speech, and Music Processing, vol. 2012, no. 1, Dec. 2012

  62. [62]

    Pyroomacoustics: A python package for audio room simulation and array processing al- gorithms,

    R. Scheibler, E. Bezzam, and I. Dokmanic, “Pyroomacoustics: A python package for audio room simulation and array processing al- gorithms,” inIEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, Apr. 2018, pp. 351–355

  63. [63]

    Toward a computational neu- roethology of vocal communication: From bioacoustics to neu- rophysiology, emerging tools and future directions,

    T. Sainburg and T. Q. Gentner, “Toward a computational neu- roethology of vocal communication: From bioacoustics to neu- rophysiology, emerging tools and future directions,”Frontiers in Behavioral Neuroscience, vol. 15, Dec. 2021

  64. [64]

    Domain general noise reduction for time series signals with noisereduce,

    T. Sainburg and A. Zorea, “Domain general noise reduction for time series signals with noisereduce,”Scientific Reports, vol. 15, no. 1, Aug. 2025

  65. [65]

    Objective and subjective evaluation of diffusion-based speech en- hancement for dysarthric speech,

    D. de Groot, T. Patel, D. Kayande, O. Scharenborg, and Z. Yue, “Objective and subjective evaluation of diffusion-based speech en- hancement for dysarthric speech,” inInterspeech 2025. ISCA, Aug. 2025, pp. 2740–2744

  66. [66]

    timsainb/noisereduce: 3.0.3,

    Tim Sainburg et al., “timsainb/noisereduce: 3.0.3,” 2024. [Online]. Available: doi.org/10.5281/ZENODO.12688680

  67. [67]

    Speech Enhancement and Dereverberation with Diffusion-based Generative Models,

    S. Welker, J.-M. Lemercier, and J. Richter, “Speech Enhancement and Dereverberation with Diffusion-based Generative Models,”

  68. [68]

    Available: https://github.com/sp-uhh/sgmse

    [Online]. Available: https://github.com/sp-uhh/sgmse

  69. [69]

    Speech-to-text: Transcription models,

    Google Cloud, “Speech-to-text: Transcription models,” https:// cloud.google.com/speech-to-text/docs/transcription-model, 2026, accessed: 2026-02-18

  70. [70]

    facebook/mms-1b-all: Massively Multilingual Speech (MMS),

    Meta AI, “facebook/mms-1b-all: Massively Multilingual Speech (MMS),” 2023. [Online]. Available: https://huggingface.co/ facebook/mms-1b-all

  71. [71]

    Whisper-large-v3,

    OpenAI, “Whisper-large-v3,” 2023. [Online]. Available: https: //huggingface.co/openai/whisper-large-v3

  72. [72]

    Fleurs: Few-shot learning evaluation of universal representations of speech,

    A. Conneau, M. Ma, S. Khanuja, Y . Zhang, V . Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna, “Fleurs: Few-shot learning evaluation of universal representations of speech,” arXiv preprint arXiv:2205.12446, 2022. [Online]. Available: https://arxiv.org/abs/2205.12446

  73. [73]

    Whisper github discussion #2363,

    OpenAI, “Whisper github discussion #2363,” 2024. [Online]. Available: https://github.com/openai/whisper/discussions/2363

  74. [74]

    Mls: A large-scale multilingual dataset for speech research,

    V . Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “Mls: A large-scale multilingual dataset for speech research,”arXiv preprint arXiv:2012.03411, 2020

  75. [75]

    V oxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,

    C. Wanget al., “V oxpopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,”arXiv preprint arXiv:2101.00390, 2021

  76. [76]

    Bootstrap estimates for confidence intervals in asr performance evaluation,

    M. Bisani and H. Ney, “Bootstrap estimates for confidence intervals in asr performance evaluation,” inIEEE ICASSP, vol. 1, 2004, pp. I–409

  77. [77]

    Good practices for evaluation of machine learning systems,

    L. Ferrer, O. Scharenborg, and T. B ¨ackstr¨om, “Good practices for evaluation of machine learning systems,”arXiv preprint arXiv:2412.03700, 2024

  78. [78]

    D. D. Boos and L. A. Stefanski,Essential Statistical Inference: Theory and Methods. Springer, 2013, section 11.6: Bootstrap Resampling for Hypothesis Tests

  79. [79]

    ITU-T Recommenda- tion P.835: Subjective test methodology for evaluating speech communication systems that include noise suppression algorithm,

    International Telecommunication Union, “ITU-T Recommenda- tion P.835: Subjective test methodology for evaluating speech communication systems that include noise suppression algorithm,” https://www.itu.int/rec/T-REC-P.835/en, June 2021