Pith. sign in

REVIEW 4 major objections 4 minor 8 references

Automatic Speech Recognition Services: Deaf and Hard-of-Hearing Usability

T0 review · 4 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Current speech recognition services are largely unusable for Deaf and Hard-of-Hearing voices, and custom vocabulary models do not change that.

desk verdict A small, candid usability study showing commercial ASR still fails on DHH speech; the headline result is credible, the category-level analysis is on shakier ground. read the letter →

arxiv 1909.02853 v1 pith:PZNCHECJ submitted 2019-09-03 cs.HC cs.SDeess.AS

classification cs.HCcs.SDeess.AS
keywords AutomaticSpeechRecognitionDeafandHard-of-HearingWordErrorRateCustomLanguageModelUsabilityAccessibilityVoiceInterfaces
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether today's commercial speech recognition services can understand Deaf and Hard-of-Hearing (DHH) speakers, and whether giving the service a custom vocabulary list closes the gap. Using sentence-list recordings from 45 DHH speakers sorted by a naive listener into 'good', 'fine', or 'bad' audio, it measures word error rates (WER) for three ASR configurations. The paper finds very high WERs for all categories, with 95% confidence intervals of roughly 91-97% for 'bad', 82-91% for 'fine', and 51-66% for 'good'. The custom vocabulary model did not significantly improve any category. The practical stakes are direct: voice interfaces in phones, cars, and home assistants will not serve DHH users at the error levels reported here.

What carries the argument

The argument runs on Word Error Rate (WER), the standard measure in which a recognized transcript is aligned to a reference and substitutions, deletions, and insertions are counted as a fraction of reference words. WER is computed for each recording under three ASR configurations and compared across three audio-quality categories ('bad', 'fine', 'good') assigned by a single naive listener, with the custom-vocabulary configuration tested against its matching base configuration. The 'custom vocabulary' mechanism is the paper's key intervention: it provides the ASR with a keyword list to give it context awareness, and the paper tests whether that intervention changes the WER pattern for DHH speech.

What would settle it

Have several naive listeners independently label the same 45 recordings as 'good', 'fine', or 'bad' and compute agreement; if the labels do not reproduce, the group-level WER comparisons and the 'bad equals fine' equivalence in Table 1 are not supported.

Watch

Extended reading notes

Core claim

The central claim is that DHH speech, as represented by this dataset, is not usable with current commercial ASR services. The 95% confidence intervals for WER are (91.338, 97.443) for the 'bad' audio category, (82.109, 91.316) for 'fine', and (51.288, 66.068) for 'good'. For 'bad' and 'fine' audio, the differences between categories were not statistically significant after Bonferroni correction, and all three services behaved similarly. Adding a custom vocabulary model to one service shifted the median for 'good' audio by a little more than 10% but with a standard deviation above 20%, and the difference was not significant; for 'bad' and 'fine' audio there was essentially no improvement. The paper concludes that DHH users with 'bad' or 'fine' voices cannot achieve word error rates comparable to the 5-6% reported for non-DHH speech, and that even 'good' voices will give unpredictable results.

Load-bearing premise

The entire grouping of recordings into 'good', 'fine', and 'bad' comes from one listener's judgment, with no independent check; if another listener would sort them differently, the statistical comparisons lose their anchor.

Editorial extensions

If this is right

  • A DHH speaker whose voice sounds 'bad' or 'fine' to a naive listener can expect word error rates above 80%, so spoken output would need extensive correction to be usable.
  • Even 'good'-sounding DHH speech produced median WERs above 45% with large spread, so voice interfaces cannot rely on human clarity judgments to predict ASR success.
  • Adding a custom vocabulary list to a commercial ASR service did not produce a significant WER improvement in any audio category in this dataset.
  • Interface designers should not treat ASR as an accessible input method for DHH users unless per-speaker or acoustic-model adaptation is shown to work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct extension would be a per-speaker calibration study: the large variance in the 'good' category suggests some DHH voices may already be near usable WER, so testing each speaker's repeated utterances could identify who can rely on current ASR.
  • The failure of the vocabulary list to help implies the bottleneck is acoustic, not lexical; a natural next test is custom acoustic models or speaker adaptation rather than keyword lists.
  • Because the categories rest on one listener's labels, a replication using several independent listeners and a pre-registered definition of 'good', 'fine', and 'bad' would show whether the bad/fine equivalence is a property of DHH speech or of the labeling scheme.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper evaluates the performance of three commercial automatic speech recognition (ASR) services—IBM Watson, Microsoft Translator, and Microsoft Presentation Translator with a custom vocabulary—on recordings of Deaf and Hard-of-Hearing (DHH) speech. The authors use 45 audio files from the NTID Clarke Sentences intelligibility corpus, categorized into 'good', 'fine', and 'bad' by a single naive listener, and compute word error rates (WER) with the NIST SCTK tool. They report high WER across all categories, with 95% confidence intervals of (91.338, 97.443) for 'bad', (82.109, 91.316) for 'fine', and (51.288, 66.068) for 'good', and find no statistically significant improvement from the custom vocabulary model. The conclusion states that DHH individuals cannot achieve WERs comparable to non-DHH users if their voice falls in the 'bad' or 'fine' categories, and that even 'good' voices yield unpredictable ASR results.

Significance. If the results hold, the paper provides valuable empirical evidence on the accessibility gap in commercial ASR for DHH users, addressing an important and under-studied problem. The use of standard WER scoring, real public ASR services, and a well-known intelligibility corpus are strengths. The qualitative finding that these particular DHH voices are poorly recognized is likely robust. However, the paper's quantitative claims and category-level generalizations rest on a single listener's subjective categorization, a small convenience sample, and an external non-DHH baseline, which limit the strength of the conclusions. The paper is a useful preliminary study but not a fully controlled evaluation.

major comments (4)
  1. [Methodology: Audio Dataset] The partition of the 45 audio files into 'good', 'fine', and 'bad' categories rests entirely on the judgment of a single naive listener, with no category definitions, no selection procedure, and no inter-rater reliability assessment. Because Table 1 and the conclusion ("if their voice fell within the 'bad' or 'fine' audio categories") stratify all analyses by these categories, the category-level claims are not reproducible or independently verifiable. The paper should either provide a second or third rater and report agreement (e.g., Cohen's kappa), or define the categories using objective acoustic or per-file intelligibility criteria.
  2. [Results] The claim that DHH speech yields WERs far above the 5-6% range for non-DHH speech is supported only by external newspaper/blog citations and not by an in-study non-DHH baseline. The one-sample t-test confidence intervals (91.338, 97.443) for 'bad', (82.109, 91.316) for 'fine', and (51.288, 66.068) for 'good' describe the DHH sample alone; they do not statistically test against a matched control group. Without a control group recorded and scored under the same protocol, the comparison to 5-6% is informal. The paper should either add a non-DHH control condition or explicitly frame the study as descriptive and avoid comparative language such as 'would not be able to achieve equal WERs'.
  3. [Improvements in WER for Deaf and Hard-of-Hearing Speech with Context Awareness] The null result for customization is based on only 14-16 speakers per category with very high variance, and the paper reports no power analysis or effect-size confidence intervals. A non-significant p-value (the lowest reported is .5472) does not support the conclusion that customization provides no improvement; it only indicates that the study had insufficient power to detect an effect. Additionally, the MSPPT custom model was pre-seeded with the entire list of Clarke Sentences—the same sentences used in the evaluation—so the customization condition is not representative of a real context-aware deployment. This contamination could bias results in either direction and should be acknowledged or remedied with a held-out keyword list.
  4. [Conclusion] The sentence "DHH individuals would not be able to achieve equal WERs as the non-DHH population if their voice fell within the 'bad' or 'fine' audio categories" overgeneralizes from a sample of 45 speakers, of whom only 30 fall in the 'bad' and 'fine' groups. The reported confidence intervals are intervals for the sample mean, not prediction intervals for future DHH speakers, and the sample is a convenience subset of a larger corpus. The conclusion should be tempered to the sample or supported by a mixed-effects model treating speaker as a random effect, so that population-level inferences are properly quantified.
minor comments (4)
  1. [Table 1] The header 'Bonferroni correlation results' is imprecise; the paper appears to apply a Bonferroni correction for multiple comparisons, not a correlation. Please rename it 'Bonferroni-corrected p-values'.
  2. [WER Analysis] The explanation of WER states that the total of substitutions, deletions, and insertions is divided by the total number of words in the reference transcript. This is correct, but the sentence could be tightened to explicitly note that insertions increase the numerator without affecting the denominator, as this is a common point of confusion.
  3. [Conclusion] There is a typo: 'Even if their voice is "good", is still likely that they will get unpredictable results' should read 'Even if their voice is "good", it is still likely that they will get unpredictable results'.
  4. [General] The paper would benefit from a reproducibility statement indicating whether the audio subset, reference transcripts, and ASR outputs are available, since the WER calculations can be exactly reproduced only if these materials are shared.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical WER measurements against external ASR services; the only self-citation supplies dataset strata, not the predicted outcomes.

full rationale

The paper's derivation chain is an empirical measurement, not a mathematical derivation. WERs are computed by NIST SCTK from reference transcripts and ASR outputs, and all comparisons are between observed WER distributions across three fixed audio categories. The claim that DHH speech yields high WERs is a measured result against external black-box services (IBM, MS, MSPPT), not an implication of the categorization; the categories are independent inputs from prior work [6]. The custom-vocabulary comparison is a controlled variation of the same ASR, and the custom model was seeded with the full Clarke sentence list used in the test material. This is a methodological contamination that could bias toward finding an improvement, yet the paper reports no significant improvement; it cannot be a circularity that explains the null result. The one-listener good/fine/bad categorization raises validity and reliability concerns, but those are empirical-design questions, not circular reasoning: no fitted parameter is renamed as a prediction, and no equation reduces to its own input. The self-citation [6] is used only to source the audio files and their pre-existing labels, not to justify the WER conclusions, so it is not load-bearing in the circularity sense.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters or invented entities; the reported CIs and p-values are standard statistics. The main assumptions are about dataset representativeness and the validity of the categories and WER metric.

assumptions (4)
  • domain assumption The Clarke Sentences intelligibility score (0-50) assigned by a speech pathologist is a valid measure of how understandable DHH speech is to humans.
    The study uses these scores to characterize the dataset (averages 25/43/48 for bad/fine/good), and the figures plot them, though the correlation with WER is not reported.
  • domain assumption The WER computed by NIST SCTK is an appropriate metric for ASR accuracy in this usability context.
    WER counts substitutions, deletions, and insertions; the paper relies on it as the sole outcome measure, including for very high error rates where insertions can make WER exceed 100%.
  • domain assumption The 45 audio files selected in prior work [6] by a naive listener are representative of DHH speech.
    No sampling procedure or selection criteria are described; the paper generalizes from 45 files (out of 650) to the DHH population.
  • domain assumption The public web demos of IBM Watson and Microsoft Translator as accessed in 2019 reflect the production behavior of these commercial ASR services.
    Web demos may differ from paid APIs, may have different models, and have changed since the study.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Automatic Speech Recognition Services: Deaf and Hard-of-Hearing Usability." pith.science (2026). https://pith.science/paper/PZNCHECJ

@misc{pith2026190902853,
  author       = {Pith},
  title        = {Pith review of: Automatic Speech Recognition Services: Deaf and Hard-of-Hearing Usability},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PZNCHECJ}},
  note         = {Machine review of arXiv:1909.02853}
}
read the original abstract

Nowadays, speech is becoming a more common, if not standard, interface to technology. This can be seen in the trend of technology changes over the years. Increasingly, voice is used to control programs, appliances and personal devices within homes, cars, workplaces, and public spaces through smartphones and home assistant devices using Amazon's Alexa, Google's Assistant and Apple's Siri, and other proliferating technologies. However, most speech interfaces are not accessible for Deaf and Hard-of-Hearing (DHH) people. In this paper, performances of current Automatic Speech Recognition (ASR) with voices of DHH speakers are evaluated. ASR has improved over the years, and is able to reach Word Error Rates (WER) as low as 5-6% [1][2][3], with the help of cloud-computing and machine learning algorithms that take in custom vocabulary models. In this paper, a custom vocabulary model is used, and the significance of the improvement is evaluated when using DHH speech.

Figures

Figures reproduced from arXiv: 1909.02853 by the authors.

Figure 1
Figure 1. Scatterplot of intelligibility scores and WER for the audio database [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Histogram showing frequency of intelligibility scores Even though ASR technology has improved dramatically over the past few years, and is now being incorporated in everyday technologies, it still has an usability challenge when it comes to the DHH population. DHH speech generally sounds different from hearing speech, and varies greatly between DHH individuals. DHH speech is often so variable that there is difficult… view at source ↗
Figure 3
Figure 3. Side by side boxplots of WER for all ASRs and audio categories Asterisks (*) denote outliers Three different ASRs ("IBM", "MS", and "MSPPT") and three different audio categories are used ("bad", "fine", and "good"). See the respective sections for in-depth explanations. WER for Deaf and Hard-of-Hearing Speech [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Comparison of WER results boxplots for each ASR service and each audio category [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

8 extracted references · 7 canonical work pages

  1. [1]

    Making sense of Google CEO Sundar Pichai’s plan to move every direction at once

    2017. Making sense of Google CEO Sundar Pichai’s plan to move every direction at once. https://www.cnbc.com/2017/05/ 18/google-ceo-sundar-pichai-machine-learning-big-data.html

  2. [2]

    Microsoft researchers achieve new conversational speech recognition milestone

    2017. Microsoft researchers achieve new conversational speech recognition milestone. https://www.microsoft.com/en-us/ research/blog/microsoft-researchers-achieve-new-conversational-speech-recognition-milestone/

  3. [3]

    Reaching new records in speech recognition

    2017. Reaching new records in speech recognition. https://www.ibm.com/blogs/watson/2017/03/ reaching-new-records-in-speech-recognition/

  4. [5]

    G. E. Dahl, Dong Yu, Li Deng, and A. Acero. 2012. Context-Dependent Pre-Trained Deep Neural Networks for Large- Vocabulary Speech Recognition. IEEE Transactions on Audio, Speech, and Language Processing 20, 1 (jan 2012), 30–42. https://doi.org/10.1109/tasl.2011.2134090

  5. [6]

    Glasser, Kesavan R

    Abraham T. Glasser, Kesavan R. Kushalnagar, and Raja S. Kushalnagar. 2017. Feasibility of Using Automatic Speech Recognition with Voices of Deaf and Hard-of-Hearing Individuals. In Proceedings of the 19th International ACM SIGACCESS Conference on Computers and Accessibility - ASSETS '17. ACM Press. https://doi.org/10.1145/3132525.3134819

  6. [7]

    Marjorie E. Magner. 1972. A speech intelligibility test for deaf children . Technical Report. Clarke School for the Deaf, Northampton, MA

  7. [8]

    Nancy S. McGarr. 1983. The Intelligibility of Deaf Speech to Experienced and Inexperienced Listeners. Journal of Speech Language and Hearing Research 26, 3 (sep 1983), 451. https://doi.org/10.1044/jshr.2603.451

  8. [9]

    National Institute of Standards and Technology. [n.d.]. SCTK. https://github.com/usnistgov/SCTK

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.