REVIEW 4 major objections 4 minor 8 references
Automatic Speech Recognition Services: Deaf and Hard-of-Hearing Usability
T0 review · 4 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Current speech recognition services are largely unusable for Deaf and Hard-of-Hearing voices, and custom vocabulary models do not change that.
desk verdict A small, candid usability study showing commercial ASR still fails on DHH speech; the headline result is credible, the category-level analysis is on shakier ground. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument runs on Word Error Rate (WER), the standard measure in which a recognized transcript is aligned to a reference and substitutions, deletions, and insertions are counted as a fraction of reference words. WER is computed for each recording under three ASR configurations and compared across three audio-quality categories ('bad', 'fine', 'good') assigned by a single naive listener, with the custom-vocabulary configuration tested against its matching base configuration. The 'custom vocabulary' mechanism is the paper's key intervention: it provides the ASR with a keyword list to give it context awareness, and the paper tests whether that intervention changes the WER pattern for DHH speech.
What would settle it
Have several naive listeners independently label the same 45 recordings as 'good', 'fine', or 'bad' and compute agreement; if the labels do not reproduce, the group-level WER comparisons and the 'bad equals fine' equivalence in Table 1 are not supported.
Extended reading notes
Core claim
The central claim is that DHH speech, as represented by this dataset, is not usable with current commercial ASR services. The 95% confidence intervals for WER are (91.338, 97.443) for the 'bad' audio category, (82.109, 91.316) for 'fine', and (51.288, 66.068) for 'good'. For 'bad' and 'fine' audio, the differences between categories were not statistically significant after Bonferroni correction, and all three services behaved similarly. Adding a custom vocabulary model to one service shifted the median for 'good' audio by a little more than 10% but with a standard deviation above 20%, and the difference was not significant; for 'bad' and 'fine' audio there was essentially no improvement. The paper concludes that DHH users with 'bad' or 'fine' voices cannot achieve word error rates comparable to the 5-6% reported for non-DHH speech, and that even 'good' voices will give unpredictable results.
Load-bearing premise
The entire grouping of recordings into 'good', 'fine', and 'bad' comes from one listener's judgment, with no independent check; if another listener would sort them differently, the statistical comparisons lose their anchor.
Editorial extensions
If this is right
- A DHH speaker whose voice sounds 'bad' or 'fine' to a naive listener can expect word error rates above 80%, so spoken output would need extensive correction to be usable.
- Even 'good'-sounding DHH speech produced median WERs above 45% with large spread, so voice interfaces cannot rely on human clarity judgments to predict ASR success.
- Adding a custom vocabulary list to a commercial ASR service did not produce a significant WER improvement in any audio category in this dataset.
- Interface designers should not treat ASR as an accessible input method for DHH users unless per-speaker or acoustic-model adaptation is shown to work.
Reading between the lines
- A direct extension would be a per-speaker calibration study: the large variance in the 'good' category suggests some DHH voices may already be near usable WER, so testing each speaker's repeated utterances could identify who can rely on current ASR.
- The failure of the vocabulary list to help implies the bottleneck is acoustic, not lexical; a natural next test is custom acoustic models or speaker adaptation rather than keyword lists.
- Because the categories rest on one listener's labels, a replication using several independent listeners and a pre-registered definition of 'good', 'fine', and 'bad' would show whether the bad/fine equivalence is a property of DHH speech or of the labeling scheme.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper evaluates the performance of three commercial automatic speech recognition (ASR) services—IBM Watson, Microsoft Translator, and Microsoft Presentation Translator with a custom vocabulary—on recordings of Deaf and Hard-of-Hearing (DHH) speech. The authors use 45 audio files from the NTID Clarke Sentences intelligibility corpus, categorized into 'good', 'fine', and 'bad' by a single naive listener, and compute word error rates (WER) with the NIST SCTK tool. They report high WER across all categories, with 95% confidence intervals of (91.338, 97.443) for 'bad', (82.109, 91.316) for 'fine', and (51.288, 66.068) for 'good', and find no statistically significant improvement from the custom vocabulary model. The conclusion states that DHH individuals cannot achieve WERs comparable to non-DHH users if their voice falls in the 'bad' or 'fine' categories, and that even 'good' voices yield unpredictable ASR results.
Significance. If the results hold, the paper provides valuable empirical evidence on the accessibility gap in commercial ASR for DHH users, addressing an important and under-studied problem. The use of standard WER scoring, real public ASR services, and a well-known intelligibility corpus are strengths. The qualitative finding that these particular DHH voices are poorly recognized is likely robust. However, the paper's quantitative claims and category-level generalizations rest on a single listener's subjective categorization, a small convenience sample, and an external non-DHH baseline, which limit the strength of the conclusions. The paper is a useful preliminary study but not a fully controlled evaluation.
major comments (4)
- [Methodology: Audio Dataset] The partition of the 45 audio files into 'good', 'fine', and 'bad' categories rests entirely on the judgment of a single naive listener, with no category definitions, no selection procedure, and no inter-rater reliability assessment. Because Table 1 and the conclusion ("if their voice fell within the 'bad' or 'fine' audio categories") stratify all analyses by these categories, the category-level claims are not reproducible or independently verifiable. The paper should either provide a second or third rater and report agreement (e.g., Cohen's kappa), or define the categories using objective acoustic or per-file intelligibility criteria.
- [Results] The claim that DHH speech yields WERs far above the 5-6% range for non-DHH speech is supported only by external newspaper/blog citations and not by an in-study non-DHH baseline. The one-sample t-test confidence intervals (91.338, 97.443) for 'bad', (82.109, 91.316) for 'fine', and (51.288, 66.068) for 'good' describe the DHH sample alone; they do not statistically test against a matched control group. Without a control group recorded and scored under the same protocol, the comparison to 5-6% is informal. The paper should either add a non-DHH control condition or explicitly frame the study as descriptive and avoid comparative language such as 'would not be able to achieve equal WERs'.
- [Improvements in WER for Deaf and Hard-of-Hearing Speech with Context Awareness] The null result for customization is based on only 14-16 speakers per category with very high variance, and the paper reports no power analysis or effect-size confidence intervals. A non-significant p-value (the lowest reported is .5472) does not support the conclusion that customization provides no improvement; it only indicates that the study had insufficient power to detect an effect. Additionally, the MSPPT custom model was pre-seeded with the entire list of Clarke Sentences—the same sentences used in the evaluation—so the customization condition is not representative of a real context-aware deployment. This contamination could bias results in either direction and should be acknowledged or remedied with a held-out keyword list.
- [Conclusion] The sentence "DHH individuals would not be able to achieve equal WERs as the non-DHH population if their voice fell within the 'bad' or 'fine' audio categories" overgeneralizes from a sample of 45 speakers, of whom only 30 fall in the 'bad' and 'fine' groups. The reported confidence intervals are intervals for the sample mean, not prediction intervals for future DHH speakers, and the sample is a convenience subset of a larger corpus. The conclusion should be tempered to the sample or supported by a mixed-effects model treating speaker as a random effect, so that population-level inferences are properly quantified.
minor comments (4)
- [Table 1] The header 'Bonferroni correlation results' is imprecise; the paper appears to apply a Bonferroni correction for multiple comparisons, not a correlation. Please rename it 'Bonferroni-corrected p-values'.
- [WER Analysis] The explanation of WER states that the total of substitutions, deletions, and insertions is divided by the total number of words in the reference transcript. This is correct, but the sentence could be tightened to explicitly note that insertions increase the numerator without affecting the denominator, as this is a common point of confusion.
- [Conclusion] There is a typo: 'Even if their voice is "good", is still likely that they will get unpredictable results' should read 'Even if their voice is "good", it is still likely that they will get unpredictable results'.
- [General] The paper would benefit from a reproducibility statement indicating whether the audio subset, reference transcripts, and ASR outputs are available, since the WER calculations can be exactly reproduced only if these materials are shared.
Circularity Check
No circularity: empirical WER measurements against external ASR services; the only self-citation supplies dataset strata, not the predicted outcomes.
full rationale
The paper's derivation chain is an empirical measurement, not a mathematical derivation. WERs are computed by NIST SCTK from reference transcripts and ASR outputs, and all comparisons are between observed WER distributions across three fixed audio categories. The claim that DHH speech yields high WERs is a measured result against external black-box services (IBM, MS, MSPPT), not an implication of the categorization; the categories are independent inputs from prior work [6]. The custom-vocabulary comparison is a controlled variation of the same ASR, and the custom model was seeded with the full Clarke sentence list used in the test material. This is a methodological contamination that could bias toward finding an improvement, yet the paper reports no significant improvement; it cannot be a circularity that explains the null result. The one-listener good/fine/bad categorization raises validity and reliability concerns, but those are empirical-design questions, not circular reasoning: no fitted parameter is renamed as a prediction, and no equation reduces to its own input. The self-citation [6] is used only to source the audio files and their pre-existing labels, not to justify the WER conclusions, so it is not load-bearing in the circularity sense.
Assumptions & free parameters
assumptions (4)
- domain assumption The Clarke Sentences intelligibility score (0-50) assigned by a speech pathologist is a valid measure of how understandable DHH speech is to humans.
- domain assumption The WER computed by NIST SCTK is an appropriate metric for ASR accuracy in this usability context.
- domain assumption The 45 audio files selected in prior work [6] by a naive listener are representative of DHH speech.
- domain assumption The public web demos of IBM Watson and Microsoft Translator as accessed in 2019 reflect the production behavior of these commercial ASR services.
Cite this review
Pith. "Pith review of Automatic Speech Recognition Services: Deaf and Hard-of-Hearing Usability." pith.science (2026). https://pith.science/paper/PZNCHECJ
@misc{pith2026190902853,
author = {Pith},
title = {Pith review of: Automatic Speech Recognition Services: Deaf and Hard-of-Hearing Usability},
year = {2026},
howpublished = {\url{https://pith.science/paper/PZNCHECJ}},
note = {Machine review of arXiv:1909.02853}
}
read the original abstract
Nowadays, speech is becoming a more common, if not standard, interface to technology. This can be seen in the trend of technology changes over the years. Increasingly, voice is used to control programs, appliances and personal devices within homes, cars, workplaces, and public spaces through smartphones and home assistant devices using Amazon's Alexa, Google's Assistant and Apple's Siri, and other proliferating technologies. However, most speech interfaces are not accessible for Deaf and Hard-of-Hearing (DHH) people. In this paper, performances of current Automatic Speech Recognition (ASR) with voices of DHH speakers are evaluated. ASR has improved over the years, and is able to reach Word Error Rates (WER) as low as 5-6% [1][2][3], with the help of cloud-computing and machine learning algorithms that take in custom vocabulary models. In this paper, a custom vocabulary model is used, and the significance of the improvement is evaluated when using DHH speech.
Figures
Reference graph
Works this paper leans on
-
[1]
Making sense of Google CEO Sundar Pichai’s plan to move every direction at once
2017. Making sense of Google CEO Sundar Pichai’s plan to move every direction at once. https://www.cnbc.com/2017/05/ 18/google-ceo-sundar-pichai-machine-learning-big-data.html
work page 2017
-
[2]
Microsoft researchers achieve new conversational speech recognition milestone
2017. Microsoft researchers achieve new conversational speech recognition milestone. https://www.microsoft.com/en-us/ research/blog/microsoft-researchers-achieve-new-conversational-speech-recognition-milestone/
work page 2017
-
[3]
Reaching new records in speech recognition
2017. Reaching new records in speech recognition. https://www.ibm.com/blogs/watson/2017/03/ reaching-new-records-in-speech-recognition/
work page 2017
- [5]
-
[6]
Abraham T. Glasser, Kesavan R. Kushalnagar, and Raja S. Kushalnagar. 2017. Feasibility of Using Automatic Speech Recognition with Voices of Deaf and Hard-of-Hearing Individuals. In Proceedings of the 19th International ACM SIGACCESS Conference on Computers and Accessibility - ASSETS '17. ACM Press. https://doi.org/10.1145/3132525.3134819
arXiv 2017
-
[7]
Marjorie E. Magner. 1972. A speech intelligibility test for deaf children . Technical Report. Clarke School for the Deaf, Northampton, MA
work page 1972
-
[8]
Nancy S. McGarr. 1983. The Intelligibility of Deaf Speech to Experienced and Inexperienced Listeners. Journal of Speech Language and Hearing Research 26, 3 (sep 1983), 451. https://doi.org/10.1044/jshr.2603.451
-
[9]
National Institute of Standards and Technology. [n.d.]. SCTK. https://github.com/usnistgov/SCTK
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.