REVIEW 2 major objections 5 minor 44 references
ASR matches or beats human listeners on diverse Dutch speech
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
State-of-the-art speech recognizers match human word-error rates on Dutch child speech and outperform native listeners on older-adults and Flemish-teenager speech in a 120-utterance pilot.
T0 review reviewed 2026-08-01 challenge →
load-bearing objection A genuinely new but small pilot benchmark: ASR parity with human listeners on selected diverse Dutch speech is plausible but limited to the 120 hand-picked utterances; the paper says so itself. the 2 major comments →
Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
On the same 40-utterance test sets per speaker group, Google Telephony achieved word error rates of 12.8% for child speech, 17.1% for older adults, and 10.3% on Flemish, compared with human listener rates of 17.6%, 26.8%, and 19.3%. Paired bootstrap tests show no significant human-vs-ASR difference for child speech; Google Telephony and Whisper significantly beat the human listeners for older-adult speech; and Google Telephony significantly beats humans and the other systems for Flemish teen speech. The paper's stated conclusion is that ASR reaches parity with human performance for child speech and that some systems exceed human performance for older-adult and Flemish speech, extending earli
What carries the argument
The core instrument is a matched-stimulus benchmarking protocol: identical utterances are transcribed by lay native listeners and by three ASR systems, all transcriptions are normalized (lowercasing, punctuation removal, unambiguous spelling correction, stripping of fillers), and performance is scored as word error rate. Statistical comparisons use a paired bootstrap with 10,000 speaker-level resamples and 95% confidence intervals, supplemented by a multi-factor linear model examining speaker age, reported gender, and regional accent, plus an analysis of error types and utterance length.
Load-bearing premise
That the 40 utterances per speaker group, chosen to balance age/gender/region, and the convenience listeners who transcribed them are representative enough that the measured human-vs-ASR differences generalize to all Dutch child, older-adult, and Flemish speech.
What would settle it
Transcribe the full corpus test sets (thousands of utterances per group) with a large demographically balanced panel of native Dutch listeners, using the same normalization and bootstrap. If the human word error rate on the full older-adult set drops below Google Telephony's full-set WER of 24.1%, the claimed superiority is refuted; if human WER rises above the ASR values, the claim is strengthened.
If this is right
- The human-listener ceiling assumption fails for at least some diverse Dutch speech; ASR can be a reliable substitute for human transcription in these conditions.
- Child speech, once a known weak spot, is now recognized with human-level accuracy by off-the-shelf systems.
- Aging voices and regional accents are the remaining weak points, pointing ASR research toward acoustic variability due to age and dialect.
- Benchmark conclusions depend heavily on test-set selection; the full-corpus WERs run 8.4% higher than the selected 40-item subsets, so reported human-vs-ASR gaps should be read as test-specific.
- Inclusive speech technology for Dutch appears closer than the typical narrative suggests, at least for telephony-style human-machine interaction.
Where Pith is reading between the lines
- If the same pattern holds for other languages, human-parity benchmarking on diverse speech may need to shift from 'can machines catch humans' to 'which acoustic variants still defeat both'. A direct multilingual replication would test this.
- The listener sample was drawn from the experimenters' social and work circles and included no payment; a representative population sample might produce lower or higher human WERs, which could flip the older-adult and Flemish superiority results.
- The 8.4% gap between selected stimuli and full test sets implies that testing human listeners on the full corpus test sets (rather than 40 items) is the natural next step; the paper's own data suggest the parity claim may not survive a harder test set.
- Human deletion errors (5-9% WER) vs ASR substitution-dominant errors hint at a qualitative difference—humans say 'didn't understand' while machines guess—which matters when ASR is used as an accessibility tool.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Huisman et al. report a preliminary benchmark of three off-the-shelf ASR systems (Google Telephony, Whisper-large-v3, and a custom Conformer model) against native Dutch listeners on 40 utterances per speaker group from the Jasmin-CGN corpus: child speech (7–11 years), older adults (59–96 years), and Flemish (teenagers and older adults). Word error rates are compared using paired-bootstrap confidence intervals. On the selected stimuli, Google Telephony achieves the best ASR performance and matches or outperforms human listeners on several groups; the authors interpret this as evidence that ASR systems reach human parity on child speech and exceed humans on older-adult and Flemish speech. The paper also compares ASR WERs on the experimental stimuli with the full Jasmin-CGN test sets, where ASR WERs are substantially higher.
Significance. If the result holds, it challenges the common assumption that human listeners set an upper bound on ASR performance, extending previous demonstrations on standard speech to diverse Dutch speech. The study is transparent about its exploratory nature, provides a useful comparative benchmark of three systems, and includes full-corpus ASR results, which is a strength. However, the headline claim is based on only 120 hand-selected stimuli, and the paper's own full-set comparison shows that these subsets are substantially easier for ASR than the broader test sets. The parity/superiority conclusion therefore currently applies to the specific investigated stimuli, not to diverse speech in general.
major comments (2)
- [II-A, II-B, III-A] The 40 stimuli per group are hand-selected to balance demographic labels and utterance length, not randomly sampled, and no power analysis is reported. With 40 items per group, the paired-bootstrap CIs are wide; for example, on child speech Google Telephony's WER is 12.8% versus humans' 17.6%, yet the difference is reported as nonsignificant, suggesting the design may lack power to detect meaningful gaps. In addition, the Flemish experiment allowed listeners to hear each stimulus twice (II-B), which plausibly improves human performance relative to the single-listening condition used in the other experiments. This does not invalidate the comparison but should be discussed as a potential confound and limits cross-group claims.
- [II-D, III-C] The statistical analysis for demographic effects uses a multi-factor linear model but reports uncorrected per-model tests and p-values. Given the number of tests (3 models × 3 speaker groups × gender/regional-accent effects), some significant results (e.g., Whisper's gender effect on Flemish, t(40)=−2.149, p=0.0378) are expected by chance. The paper should either use multiple-comparison corrections or explicitly label these demographic analyses as exploratory. This does not affect the primary parity claim but weakens the reliability of the variability findings.
minor comments (5)
- [Table I] ASR WERs are given as point estimates without confidence intervals; add CIs or state that ASR is deterministic on the given test set.
- [Figures 1–4] The figures would benefit from error bars or per-bin sample counts; the text notes 'no statistical tests were carried out' for the age and length splits, but the captions do not state this.
- [II-B] The sentence 'Each listener participated only in one experiment, except for two participants who participated in both the child speech and older adult speech experiments' is self-contradictory; rephrase to clarify the sample.
- [IV, footnote 2] The footnote about Google Telephony possibly being trained on CGN-Jasmin data is important for interpreting the main result and should be moved to the methodology or limitations section rather than appearing only in the discussion.
- [IV] Minor typo: 'extend the results of [7]' should be 'extend the results in [7]' or a clearer phrasing.
Circularity Check
No circularity: empirical benchmark with measured WERs; self-citations are contextual, not load-bearing.
full rationale
The paper does not derive predictions from fitted parameters or from a mathematical chain that reduces to its own inputs. The central claim—that ASR can reach human parity on child speech and exceed humans on some older-adult and Flemish speech—is a direct reading of measured WERs in Table I, obtained from matched human listeners and three off-the-shelf ASR systems on the same 120 stimuli from the external Jasmin-CGN corpus. No parameter is fitted to the human results and then 'predicted' back; no uniqueness theorem or author-imported ansatz forces the outcome. The self-citations ([15], [37]) are used for model selection and prior context, not to establish the present benchmark numbers. The acknowledged possibility that Google Telephony was trained on Jasmin-CGN data (Section IV, footnote 2) and the explicit caveat that different or more stimuli could change the conclusions (Section IV) are experimental limitations and generalizability risks, not circularity: they do not make any result equivalent to its input by construction. The comparison of ASR on full test sets with human results on 40-item subsets is methodologically imperfect, but again it is not a circular derivation. The study is self-contained as an empirical benchmark, so no significant circularity is present.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption The 40-stimuli selection is representative enough to benchmark human-vs-ASR performance for each speaker group.
- domain assumption Human listeners recruited from the experimenters' social circles and mostly from the West region represent native Dutch listening ability.
- domain assumption The off-the-shelf ASR systems were not trained on the Jasmin-CGN evaluation data.
- domain assumption WER computed after manual spelling correction and token normalization measures human and ASR recognition comparably.
Cite this review
Pith. "Pith review of Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results." pith.science (2026). https://pith.science/paper/74637HXZ
@misc{pith2026260719049,
author = {Pith},
title = {Pith review of: Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results},
year = {2026},
howpublished = {\url{https://pith.science/paper/74637HXZ}},
note = {Machine review of arXiv:2607.19049}
}
read the original abstract
Humans are often considered to be the best listeners and seen as the upper-bound performance of automatic speech recognition (ASR) systems. We present a preliminary comparison of the performances of state-of-the-art ASR systems and Dutch native listeners on the recognition of "diverse" speech, specifically Dutch child and older adults' speech and Flemish. Google Telephony outperformed the other ASR systems. Importantly, the ASR systems showed similar performance to the listeners, and in specific cases even outperformed them. Slight performance differences between the listeners and ASR systems were found related to speaker's age and regional accents and utterance length. Future research should focus on making ASR systems more robust to acoustic variability related to aging and regional accents. A comparison of ASR recognition performances on the test stimuli and the full Jasmin-CGN test sets showed the influence of the specific test sets on the conclusions regarding benchmarking human and ASR performance.
Figures
Reference graph
Works this paper leans on
-
[1]
Speech recognition by machines and humans,
R. P. Lippmann, “Speech recognition by machines and humans,”Speech Communication, no. 22, pp. 1–15, 1997
1997
-
[2]
The Interspeech 2008 Consonant Challenge
M. Cooke and O. Scharenborg, “The Interspeech 2008 Consonant Challenge.” inProceedings of Interspeech, Brisbane, Australia, 2008, pp. 1765–1768
2008
-
[3]
Robustness of spectro-temporal features against intrinsic and extrinsic variations in automatic speech recogni- tion,
B. T. Meyer and B. Kollmeier, “Robustness of spectro-temporal features against intrinsic and extrinsic variations in automatic speech recogni- tion,”Speech Communication, vol. 53, pp. 753–767, 2011
2011
-
[4]
What’s the difference? Comparing humans and machines on the aurora2 speech recognition task,
B. T. Meyer, “What’s the difference? Comparing humans and machines on the aurora2 speech recognition task,” inProceedings of Interspeech, Lyon, France, 2013, pp. 2634–2638
2013
-
[5]
Comparing human and automatic speech recognition in simple and complex acoustic scenes,
C. Spille, B. Kollmeier, and B. T. Meyer, “Comparing human and automatic speech recognition in simple and complex acoustic scenes,” Computer Speech & Language, vol. 52, pp. 123–140, 2018
2018
-
[6]
A comparison of automatic and human speech recognition in null grammar,
Amit Juneja, “A comparison of automatic and human speech recognition in null grammar,”Journal of the Acoustical Society of America, vol. 3, no. 131, pp. EL256–61, 2012
2012
-
[7]
Toward human parity in conversational speech recognition,
W. Xiong, J. Droppo, X. Huang, F. Seide, M. L. Seltzer, A. Stolcke, D. Yu, and G. Zweig, “Toward human parity in conversational speech recognition,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 25, no. 12, pp. 2410–2423, 2017
2017
-
[8]
Transcription methods for consistency, volume and efficiency,
M. L. Glenn, S. M. Strassel, H. Lee, K. Maeda, R. Zakhary, and X. Li, “Transcription methods for consistency, volume and efficiency,” inProceedings of the Seventh International Conference on Language Resources and Evaluation (LREC’10), Valletta, Malta, May 2010
2010
-
[9]
Speech recognition in adverse conditions by humans and machines
C. Patman and E. Chodroff, “Speech recognition in adverse conditions by humans and machines.”Journal of the Acoustical Society of America - Express Letters, vol. 4 11, 2024
2024
-
[10]
Wav2Vec 2.0: A framework for self-supervised learning of speech representations,
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “Wav2Vec 2.0: A framework for self-supervised learning of speech representations,” Neural Inf. Process. Syst., no. 33, p. 12449–12460, 2020
2020
-
[11]
Whisper-large-v3,
OpenAI, “Whisper-large-v3,” 2023. [Online]. Available: https: //huggingface.co/openai/whisper-large-v3
2023
-
[12]
Evaluation of automatic speech recognition for conversational speech in Dutch, English and German: What goes missing?
A. Lopez, A. Liesenfeld, and M. Dingemanse, “Evaluation of automatic speech recognition for conversational speech in Dutch, English and German: What goes missing?” inProceedings of the 18th Conference on Natural Language Processing (KONVENS 2022), Potsdam, Germany, 12–15 Sep. 2022, pp. 135–143
2022
-
[13]
Revisiting parity of human vs. machine conversational speech tran- scription,
C. Mansfield, S. Ng, G.-A. Levow, R. A. Wright, and M. Ostendorf, “Revisiting parity of human vs. machine conversational speech tran- scription,” inProceedings Interspeech 2021 – Annual Conference of the International Speech Communication Association, Brno, Czechia, 2021, pp. 1997–2001
2021
-
[14]
Bidirec- tional LSTM-RNN for improving automated assessment of non-native children’s speech,
Y . Qian, K. Evanini, X. Wang, C. M. Lee, and M. Mulholland, “Bidirec- tional LSTM-RNN for improving automated assessment of non-native children’s speech,” inProceedings of Interspeech, Stockholm, Sweden, 2017
2017
-
[15]
Towards inclusive automatic speech recognition,
S. Feng, B. M. Halpern, O. Kudina, and O. Scharenborg, “Towards inclusive automatic speech recognition,”Computer Speech & Language, vol. 84, p. 101567, 2024
2024
-
[16]
Bias in Flemish automatic speech recognition,
A. Herygers, V . Verkhodanova, M. Coler, O. Scharenborg, and M. Georges, “Bias in Flemish automatic speech recognition,” inESSV Konferenz Elektronische Sprachsignalverarbeitung, Germany, 2023
2023
-
[17]
End-to-end neural systems for automatic children speech recognition: An empirical study,
P. Gurunath Shivakumar and S. Narayanan, “End-to-end neural systems for automatic children speech recognition: An empirical study,”Com- puter Speech & Language, vol. 72, p. 101289, 2022
2022
-
[18]
Automatic recognition of Dutch dysarthric speech, a pilot study,
E. Sanders, M. B. Ruiter, L. Beijer, and H. Strik, “Automatic recognition of Dutch dysarthric speech, a pilot study,” inProceedings of the International Conference on Spoken Language Processing (ICSLP), 2002, pp. 661–664
2002
-
[19]
Phonetic analysis of dysarthric speech tempo and applications to robust personalised dysarthric speech recognition,
F. Xiong, J. Barker, and H. Christensen, “Phonetic analysis of dysarthric speech tempo and applications to robust personalised dysarthric speech recognition,” inInternational Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019, pp. 5836–5840
2019
-
[20]
Recent progress in the CUHK dysarthric speech recognition system,
S. Liu, M. Geng, S. Hu, X. Xie, M. Cui, J. Yu, X. Liu, and H. Meng, “Recent progress in the CUHK dysarthric speech recognition system,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 2267–2281, 2021
2021
-
[21]
Cross-lingual self-supervised speech repre- sentations for improved dysarthric speech recognition,
A. Hernandez, P. A. P ´erez-Toro, E. Noeth, J. R. Orozco-Arroyave, A. Maier, and S. H. Yang, “Cross-lingual self-supervised speech repre- sentations for improved dysarthric speech recognition,” inInterspeech, 2022, pp. 51–55
2022
-
[22]
A. Alsayegh and T. Masood, “Zero-shot recognition of dysarthric speech using commercial automatic speech recognition and multimodal large language models,”arXiv preprint arXiv:2512.17474, 2025
arXiv 2025
-
[23]
On the impact of dysarthric speech on contemporary ASR cloud platforms,
L. De Russis and F. Corno, “On the impact of dysarthric speech on contemporary ASR cloud platforms,”Journal of Reliable Intelligent Environments, vol. 5, no. 3, pp. 163–172, 2019
2019
-
[24]
Low-resource automatic speech recognition and error analyses of oral cancer speech,
B. M. Halpern, S. Feng, R. van Son, M. van den Brekel, and O. Scharen- borg, “Low-resource automatic speech recognition and error analyses of oral cancer speech,”Speech Communication, vol. 141, pp. 14–27, 2022
2022
-
[25]
Evaluation of speech intelligibility for children with cleft lip and palate by means of automatic speech recognition,
M. Schuster, A. Maier, T. Haderlein, E. Nkenke, U. Wohlleben, F. Rosanowski, U. Eysholdt, and E. N ¨oth, “Evaluation of speech intelligibility for children with cleft lip and palate by means of automatic speech recognition,”International Journal of Pediatric Otorhinolaryn- gology, vol. 70, no. 10, pp. 1741–1747, 2006
2006
-
[26]
See what I’m saying? Comparing intelligent personal assistant use for native and non-native language speakers,
Y . Wu, D. Rough, A. Bleakley, J. Edwards, O. Cooney, P. R. Doyle, L. Clark, and B. R. Cowan, “See what I’m saying? Comparing intelligent personal assistant use for native and non-native language speakers,” in 22nd International Conference on Human Computer Interaction with Mobile Devices and Services, 2020, pp. 1–9
2020
-
[27]
Do you understand the words that are comin outta my mouth? V oice assistant comprehension of medication names,
A. Palanica, A. Thommandram, A. Lee, M. Li, and Y . Fossat, “Do you understand the words that are comin outta my mouth? V oice assistant comprehension of medication names,”NPJ Digital Medicine, no. 2(1), pp. 1–6, 2019
2019
-
[28]
Mit- igating bias against non-native accents,
Y . Zhang, Y . Zhang, B. M. Halpern, T. Patel, and O. Scharenborg, “Mit- igating bias against non-native accents,” inProceedings of Interspeech, Incheon, Korea, 2022, p. 3168–3172
2022
-
[29]
Racial disparities in automated speech recognition,
A. Koenecke, A. Nam, E. Lake, J. Nudell, M. Quartey, Z. Mengesha, C. Toups, J. R. Rickford, D. Jurafsky, and S. Goel, “Racial disparities in automated speech recognition,”Proceedings of the National Academy of Sciences, vol. 117, no. 14, pp. 7684–7689, 2020
2020
-
[30]
Gender and dialect bias in YouTube’s automatic captions,
R. Tatman, “Gender and dialect bias in YouTube’s automatic captions,” inProceedings of the First ACL Workshop on Ethics in Natural Language Processing, D. Hovy, S. Spruit, M. Mitchell, E. M. Bender, M. Strube, and H. Wallach, Eds., Valencia, Spain, Apr. 2017, pp. 53–59
2017
-
[31]
Effects of talker dialect, gender & race on accuracy of Bing speech and YouTube automatic captions
R. Tatman and C. Kasten, “Effects of talker dialect, gender & race on accuracy of Bing speech and YouTube automatic captions.” in Proceedings of Interspeech, Stockholm, Sweden, 2017, pp. 934–938
2017
-
[32]
M. A. Shariah and M. Sawalha, “The effects of speakers’ gender, age, and region on overall performance of Arabic automatic speech recognition systems using the phonetically rich and balanced modern standard Arabic speech corpus,” inProceedings of the 2nd Workshop of Arabic Corpus Linguistics WACL-2, 2013
2013
-
[33]
Speech-to-text: Transcription models,
Google Cloud, “Speech-to-text: Transcription models,” https://cloud. google.com/speech-to-text/docs/transcription-model, 2025, accessed: 2025-09-08
2025
-
[34]
JASMIN- CGN: Extension of the Spoken Dutch Corpus with speech of elderly people, children and non-natives in the human-machine interaction modality,
C. Cucchiarini, H. Van hamme, O. Herwijnen, and F. Smits, “JASMIN- CGN: Extension of the Spoken Dutch Corpus with speech of elderly people, children and non-natives in the human-machine interaction modality,” inLREC 2006, Genoa, Italy, 2006
2006
-
[35]
FFmpeg loudnorm
The FFmpeg developers, “FFmpeg loudnorm.” [Online]. Available: https://ffmpeg.org/ffmpeg-filters.html
-
[36]
Lexical information drives perceptual learning of distorted speech: evidence from the comprehension of noise-vocoded sentences
M. H. Davis, I. S. Johnsrude, A. Hervais-Adelman, K. Taylor, and C. McGettigan, “Lexical information drives perceptual learning of distorted speech: evidence from the comprehension of noise-vocoded sentences.”Journal of Experimental Psychology: General, vol. 134, no. 2, p. 222, 2005
2005
-
[37]
Speech recognition performance disparities between Dutch diverse speaker groups
Y . Zhang, T. Valck, and O. Scharenborg, “Speech recognition performance disparities between Dutch diverse speaker groups.” 2026. [Online]. Available: https://doi.org/10.1515/phon-2025-0061
-
[38]
Un- supervised cross-lingual representation learning for speech recognition
A. Conneau, A. Baevski, R. Collobert, A. Mohamed, and M. Auli, “Un- supervised cross-lingual representation learning for speech recognition.” inProceedings of Interspeech, Brno, Czechia, 2021, pp. 2426–2430
2021
-
[39]
The Spoken Dutch Corpus. Overview and first evalua- tion,
N. Oostdijk, “The Spoken Dutch Corpus. Overview and first evalua- tion,” inProceedings of the 2nd International Conference on Language Resources and Evaluation (LREC’00). Athens, Greece: European Language Resources Association (ELRA), May 2000
2000
-
[40]
Bootstrap estimates for confidence intervals in ASR performance evaluation,
M. Bisani and H. Ney, “Bootstrap estimates for confidence intervals in ASR performance evaluation,” inIEEE ICASSP, vol. 1, 2004, pp. I–409
2004
-
[41]
Good practices for eval- uation of machine learning systems,
L. Ferrer, O. Scharenborg, and T. B ¨ackstr¨om, “Good practices for eval- uation of machine learning systems,”arXiv preprint arXiv:2412.03700, 2024
Pith/arXiv arXiv 2024
-
[42]
Emmeans: Estimated marginal means, aka least-squares means
R. Lenth, “Emmeans: Estimated marginal means, aka least-squares means .”R package version 2.0. 1, 2023
2023
-
[43]
Ageing voices: The effect of changes in voice parameters on ASR Performance,
R. Vipperla, S. Renals, and J. Frankel, “Ageing voices: The effect of changes in voice parameters on ASR Performance,”EURASIP Journal on Audio, Speech, and Music Processing, vol. 2010, pp. 1–10, 2010
2010
-
[44]
Uncovering bias in asr systems: Evaluating Wav2vec2 and Whisper for dutch speakers,
M. Fuckner, S. Horsman, P. Wiggers, and I. Janssen, “Uncovering bias in asr systems: Evaluating Wav2vec2 and Whisper for dutch speakers,” in2023 International Conference on Speech Technology and Human- Computer Dialogue (SpeD). IEEE, 2023, pp. 146–151
2023
This paper was first reviewed by deepseek-v4-flash on August 1, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.