REVIEW 3 major objections 4 minor 30 references
Deaf, Hard of Hearing, and Hearing Perspectives on using Automatic Speech Recognition in Conversation
T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Commercial speech recognizers fail on deaf users' speech, with a 78% word error rate.
desk verdict The abstract's headline WER numbers don't appear in the body; what's left is an honest but thin experience report. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is word error rate (WER), the metric by which the paper judges commercial ASR engines, paired with an experience report of seven no-cost apps (DEAFCOM, Dragon Dictation, Siri, Virtual Voice, Ava, Google Assistant, and Amazon Alexa) used in everyday settings. The mechanism behind the high WER is that third-generation deep-learning recognizers, trained on large corpora of typical hearing speech, generalize to mild accents but not to the wider variation in pitch, formants, segmental articulation, and prosody found in deaf speech; the paper cites evidence that pitch and formant variation in deaf children is too large to serve as reliable recognition features. This combination lets the paper separate the core recognition failure from usability problems such as lag, noise, and interface design, and argue that fixing the interface alone is not enough.
What would settle it
Run a controlled study in which a diverse group of deaf, hard of hearing, and hearing speakers read the same sentences to the same set of commercial ASR apps in both quiet and noisy conditions, and compare word error rates and user ratings. If median deaf-speech WER falls below the roughly 15% implied by the 85% accuracy threshold, or if a substantial share of DHH participants find the apps usable in real tasks, the paper's central claim would be contradicted.
Extended reading notes
Core claim
The central claim is that automatic speech recognition built into popular commercial products performs so poorly on deaf and hard of hearing speakers that voice-controlled interfaces cannot serve them as they are. Measured on the authors' own speech, word error rate was roughly 78% for deaf speech against roughly 18% for hearing speech, far above the 85–90% accuracy that prior work says is needed for a transcript to be useful, and above the 98% accuracy argued to preserve meaning and intent. The authors attribute the failure to a mismatch between acoustic models trained mainly on hearing speech and the segmental and prosodic characteristics of deaf speech, such as pitch, formants, rate, pausing, volume, intonation, and stress. They also document secondary barriers: latency that grows in noisy settings, random inserted text, lack of feedback about speech volume and microphone placement, and interfaces that do not readily support text input or on-the-fly correction. The paper concludes that significant advances in speech recognition software or alternative input approaches are needed before these systems are accessible to DHH users.
Load-bearing premise
The conclusion rests on the premise that the authors' own informal evaluation of seven apps in everyday settings represents the wider deaf and hard of hearing population; if their experiences do not generalize, the claim that current speech-controlled interfaces are not usable by DHH people does not follow.
Editorial extensions
If this is right
- Current voice assistants and dictation apps cannot be considered accessible to deaf and hard of hearing users, because their WER on deaf speech exceeds published usability thresholds.
- Access to speech-controlled devices for DHH users requires either retraining acoustic models on deaf speech or a new recognition paradigm; simple interface patches will not close the gap.
- Even deaf speakers whose voices are intelligible to hearing people may still get unusable ASR results, since the paper found large WER variance among speakers rated high on speech intelligibility.
- The gap between the reported WERs means that deaf users' speech is misrecognized more than four times as often as hearing users' speech, a difference large enough to preclude hands-free interaction in most settings.
- Real-time captioning from ASR in noisy, multi-speaker settings remains below the accuracy and latency needed for classroom or workplace use, so DHH users will continue to depend on human captioners, interpreters, or text-to-text alternatives.
Reading between the lines
- Beyond the paper, a controlled study with a larger, diverse DHH sample and the same apps would probably show that WER varies widely by speaker, app, and noise level; the 78% figure is an anchor from the authors' own use, not a measured population statistic.
- The paper implies but does not say that the 'trained on majority speech' failure mode extends to other atypical speech populations, so a public benchmark built from deaf speech could double as a general test of inclusive ASR.
- The reported success of one deaf user with Amazon Alexa suggests that speaker-adaptive systems may reach some deaf users before full general-purpose ASR does, making per-speaker adaptation a promising testable extension.
- A direct comparison of deaf and hearing speech under identical noise and multi-speaker conditions would separate acoustic-model failure from interface lag and microphone issues; the paper's informal design does not allow that separation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is an experience report on the accessibility barriers that deaf, hard of hearing, and hearing people face when using automatic speech recognition (ASR) applications in mixed-group conversation. It describes deaf speech intelligibility ratings from a figure of about 650 deaf people, reports a word error rate of about 53% for deaf speakers rated 5.0 in personal testing, and recounts the authors' use of seven commercial ASR apps in real-world settings such as classrooms, interviews, and conversations. The abstract makes a stronger quantitative claim not present in the body: that deaf speech had approximately a 78% WER compared to 18% for hearing speech, and that current speech-controlled interfaces are not usable by deaf and hard of hearing people.
Significance. If the headline numbers were supported, they would represent an important and actionable accessibility finding with direct design implications for commercial speech interfaces. The qualitative lessons and recommendations in Section 6, such as the need for visible text output, external microphone support, and control of lag, may be useful to practitioners building accessible ASR tools. The paper does not, however, provide a verifiable basis for its central quantitative claim: the 78%/18% comparison is absent from the body, the only reported deaf-speech WER is a single unprotocoled 53% figure, and the user-experience evaluation in Section 5 is informal self-use with no sample size or controls. The paper's value is as an experience report, not as evidence for the strong, general conclusion stated in the abstract.
major comments (3)
- [Abstract and Section 2.3] The abstract's headline quantities—deaf speech WER approximately 78% versus hearing speech WER approximately 18%—never appear in the body; the only deaf-speech WER reported is about 53% in Section 2.3, attributed to 'personal testing' with no recognizer names or versions, no utterance corpus, no sample size, and no scoring definition. This omission makes the paper's central quantitative finding unverifiable and requires either removing the numbers or adding a full methods description and data.
- [Section 5] The evaluation in Section 5 consists of the authors using seven apps on their personal devices in unspecified everyday settings, with no participant count, no per-app accuracy or latency statistics, no measure of variability across speakers, and no matched hearing controls. From this informal self-evaluation one cannot infer the abstract's universal conclusion that 'current speech-controlled interfaces are not usable by DHH people.'
- [Section 7] Section 7 states that 'it would be very beneficial to have an experiment along with a survey' and to 'recruit people with different backgrounds,' which directly acknowledges that the systematic evidence needed for the abstract's quantitative and universal usability claim has not been collected. The conclusion also refers to 'one of our studies' for the WER variance, but that study is never named or described in the manuscript, so the paper cites its own unreported data as evidence.
minor comments (4)
- [Title page] The affiliation line contains a typo, 'Rochester Institute of T echnology', which should be corrected.
- [Section 5] There are several wording slips, for example 'Ava is is an eponymous product' and later 'they does not mean they will actually use it'; these should be copyedited.
- [References] Reference [2] lists the authors as 'undefined undefined undefined' and reference [4] lacks publisher and year; both entries need to be completed.
- [Figure 2] Figure 2 is described as showing the distribution of about 650 deaf speakers' intelligibility ratings, but no source, axis labels, or methodology for these ratings is provided; if the figure is retained, it needs a proper caption and citation.
Circularity Check
No equation-level circularity; the only self-referential element is an unnamed 'one of our studies' WER figure, which is minor because the qualitative usability claim is independently supported.
-
other
[Section 7 (Conclusions), with Section 2.3]
"As found in one of our studies, there is a big variance in WER for those whom have voices that are understandable by a hearing person. This means that even though a deaf person might seem to speak well, ASR still won't always have good results."
The quantitative WER-variance claim is attributed to an unnamed study by the same authors. The abstract's 78% versus 18% comparison is absent from the body; the only in-body WER figure, about 53% for deaf speakers rated 5.0, is described as coming from 'personal testing' with no corpus, recognizer names or versions, or scoring definition. The support is therefore self-referential: the reader is asked to accept the authors' unpublished testing as evidence, and the paper itself later says an experiment would be beneficial. This is not a derivation by construction, and it is not load-bearing for the qualitative usability conclusion, which is independently supported by direct app-use observations and external citations; hence the score is low.
full rationale
This is an experience report, not a derivation: there are no fitted parameters, no equations, and no first-principles result whose output could equal its input by construction. The central qualitative claim that current ASR applications are not usable for many DHH speakers is supported by the authors' direct use of seven apps (Section 5) and by external citations (Venkatagiri et al., Jeyalakshmi et al., Kushalnagar et al.) that independently document ASR difficulty with deaf speech. The abstract's quantitative headline ('Deaf speech had approximately a 78% word error rate compared to a hearing speech 18% WER') does not appear in the body, and the only in-body WER figure (about 53% for deaf speakers rated 5.0, Section 2.3) is attributed to 'personal testing' and then to 'one of our studies' (Section 7) without naming the study or providing a corpus, recognizer versions, or scoring protocol. That is a verifiability problem, not a circular derivation: no quantity is defined in terms of another, and no fitted input is relabeled as a prediction. Section 7 itself concedes 'It would be very beneficial to have an experiment along with a survey,' which confirms the report is not presenting a measured experiment. Because the qualitative conclusion has independent support, the unnamed self-study is a minor self-referential element rather than a load-bearing circular premise. Score 2 reflects that minor self-citation; the substance of the paper is not circular.
Assumptions & free parameters
assumptions (3)
- domain assumption Word error rate thresholds from captioning studies (85-90% accuracy needed for classroom use) transfer unchanged to voice-controlled interfaces and conversational ASR.
- domain assumption The commercial ASR systems evaluated in 2016-2017 are representative of current speech-controlled interfaces at the time of the arXiv submission in 2019.
- ad hoc to paper An unreported measurement of deaf and hearing speech WER exists and supports the abstract's 78% and 18% figures.
Cite this review
Pith. "Pith review of Deaf, Hard of Hearing, and Hearing Perspectives on using Automatic Speech Recognition in Conversation." pith.science (2026). https://pith.science/paper/VJOT7IEK
@misc{pith2026190901176,
author = {Pith},
title = {Pith review of: Deaf, Hard of Hearing, and Hearing Perspectives on using Automatic Speech Recognition in Conversation},
year = {2026},
howpublished = {\url{https://pith.science/paper/VJOT7IEK}},
note = {Machine review of arXiv:1909.01176}
}
read the original abstract
Many personal devices have transitioned from visual-controlled interfaces to speech-controlled interfaces to reduce costs and interactive friction, supported by the rapid growth in capabilities of speech-controlled interfaces, e.g., Amazon Echo or Apple's Siri. A consequence is that people who are deaf or hard of hearing (DHH) may be unable to use these speech-controlled devices. We show that deaf speech has a high error rate compared to hearing speech, in commercial speech-controlled interfaces. Deaf speech had approximately a 78% word error rate (WER) compared to a hearing speech 18% WER. Our findings show that current speech-controlled interfaces are not usable by DHH people. Based on our findings, significant advances in speech recognition software or alternative approaches will be needed for deaf use of speech-controlled interfaces. We show that current speech-controlled interfaces are not usable by DHH people.
Figures
Reference graph
Works this paper leans on
-
[1]
INTRODUCTION Deaf or hard of hearing (DHH) people usually cannot under- stand speech unaided, and usually depend on additional sup- port such as hearing aids or speech-to-text technology, espe- cially in multi-speaker environments. Simple low-technology aids such as using paper and pen to write back and forth or to text back and forth can work, but are ab...
work page Pith review arXiv 2017
-
[2]
PARTICIPANTS Deaf, hard of hearing and hearing speakers and listeners have different challenges and accessibility needs in mixed group conversation in most settings, including academic and workplace settings. 2.1 Deaf Participants Deaf participants have challenges in both accessing and fol- lowing spoken information and in conveying information ef- ficientl...
-
[3]
ASR EVOLUTION Speech recognition is still a very difficult task for applica- tions. For the last 50 years, researchers and inventors have iteratively implemented and improved applications that can understand speech, including conversation. First-generation systems adopted a pattern-matching ap- proach in which speech waveforms were matched with spe- cific wo...
-
[4]
ASR FOR CONVERSA TIONAL USE Major obstacles that limit ASR use in group conversation for deaf and hard of hearing includes text accuracy and lag time. Additionally, ASR accuracy for any application can be affected by other variables including auditory factors such as speech fidelity, ambient noise and microphone quality, and computing factors such as availa...
-
[5]
USER EXPERIENCE To investigate the capabilities of current ASR applications, from Fall 2016 through Summer 2017, the authors used a variety of applications on personal devices in everyday, real- world settings. The purpose of using the speech recogni- tion applications was to facilitate face-to-face spoken lan- guage interactions by providing a visible te...
work page 2016
-
[6]
The following lessons were learned from daily expe- rience with the technology
LESSONS LEARNED When using ASR technology, deaf, hard of hearing and hear- ing people report different ways of interacting with the tech- nology. The following lessons were learned from daily expe- rience with the technology
-
[7]
They do not generally consider speech to text translations to be particularly important
Hearing people often use the technology when their hands are occupied as a way to do basic tasks such as looking up simple queries on the web. They do not generally consider speech to text translations to be particularly important
-
[8]
Deaf people most often reported that their experience with the technology for their own use was restricted to messing around with the technology in order to see what the technology might interpret from their speech. Some report that ASR is starting to work better for them now, to the point where some deaf speech can be comfortably understandable
Show all 30 references
-
[9]
Generally most deaf people we met have not had this level of success with the technology yet
One deaf person reported that she commonly would use Amazon Alexa for its intended purposes and said that she was satisfied with its use. Generally most deaf people we met have not had this level of success with the technology yet
-
[10]
Several reported that the technology does not work well enough for them or that they felt uncomfort- able using the technology in a public setting
So far, only a small portion of the hearing population seems to often use the voice recognition feature of their phone. Several reported that the technology does not work well enough for them or that they felt uncomfort- able using the technology in a public setting. This led ...
-
[11]
Phones have various con- nection strengths and reliability
ASR services generally use Internet connections to send audio data to their service. Phones have various con- nection strengths and reliability. Some have difficulty with wireless networks and connections, so performing the actual analysis can take erratic amounts of time. Also,...
-
[12]
Many people often men- tion this when discussing why they do not use ASR as much as texting
When using ASR to communicate with another per- son, the ability to change the text on the fly has been limited and/or not feasible. Many people often men- tion this when discussing why they do not use ASR as much as texting
-
[13]
The two most popu- lar personal assistants for smart phones, Google Assis- tant and Siri display their commands, but sometimes only voice their responses
Interaction with ASR devices such as Alexa and the Google Home has been limited since deaf people do not have access to the verbal responses from said devices after commands are spoken to it. The two most popu- lar personal assistants for smart phones, Google Assis- tant and S...
-
[14]
They reported that this has been mostly a result of the user interface design of the ASR app itself and the fact that they were not intuitive to use
A few deaf people have reported that they experience obstacles when they try to converse with a hearing person who have never used ASR technology before. They reported that this has been mostly a result of the user interface design of the ASR app itself and the fact that they ...
-
[15]
Using a system like this might increase the interaction of users with the device
Users said that it would be ideal if ASR systems could tell the user if repetition of a specific word was needed rather than the whole word. Using a system like this might increase the interaction of users with the device
-
[16]
Heavy accents or unique accents still confuse the software too much to be usable on a daily basis
Mild accents do not affect ASR much anymore due to advances in the technology. Heavy accents or unique accents still confuse the software too much to be usable on a daily basis. While improvements in speech recognition are being reported with third generation processing strateg...
-
[17]
Whenever there are errors, the errors are time consuming to fix and the interfaces are not customizable
CONCLUSIONS Although many people use ASR systems such as Siri or Google for recreational use and every once in a while to send a text, they are not comfortable using these systems for sus- tained conversational use, as the systems have higher than tolerable error rates, especi...
-
[18]
Learning via direct and mediated instruction by deaf students
Marc Marschark, Patricia Sapere, Carol Convertino, and Jeff Pelz. Learning via direct and mediated instruction by deaf students. Journal of deaf studies and deaf education , 13(4):546–561, 2008
2008
-
[19]
Walter, and undefined undefined undefined
Susan Bannerman Foster, Gerard G. Walter, and undefined undefined undefined. Deaf students in postsecondary education. Routledge, 1992
1992
-
[20]
Professional jobs and hearing loss: A comparison of deaf and hard of hearing consumers
Daniel L Boutin. Professional jobs and hearing loss: A comparison of deaf and hard of hearing consumers. Journal of Rehabilitation , 75(1):36–40, Mar 2009
2009
-
[21]
Deaf workers: Educated and employed, but limited in career growth
Ronald R Kelly. Deaf workers: Educated and employed, but limited in career growth
-
[22]
Speech recognition technology applications in communication disorders
HS Venkatagiri. Speech recognition technology applications in communication disorders. American Journal of Speech-Language Pathology , 11(4):323–332, 2002
2002
-
[23]
Deaf speech assessment using digital processing techniques
C Jeyalakshmi, V Krishnamurthi, and A Revathy. Deaf speech assessment using digital processing techniques. Signal & Image Processing: An International Journal (SIPIJ) , 1(1):14–25, 2010
2010
-
[24]
Raja. S. Kushalnagar, Walter S. Lasecki, and Jeffrey P. Bigham. Accessibility Evaluation of Classroom Captions. ACM Transactions on Accessible Computing, 5(3):1–24, 2014
2014
-
[25]
Anderson
Karen L. Anderson. Access is the issue, not hearing loss: New policy clarification requires schools to ensure effective communication access. SIG 9 Perspectives on Hearing and Hearing Disorders in Childhood, 25(1):24–36, 2015
2015
-
[26]
Automatic speech recognition: A deep learning approach
Dong Yu and Li Deng. Automatic speech recognition: A deep learning approach . Springer, 2014
2014
-
[27]
Using automatic speech recognition to enhance education for all students: Turning a vision into reality
Mike Wald. Using automatic speech recognition to enhance education for all students: Turning a vision into reality. In Frontiers in Education, 2005. FIE’05. Proceedings 35th Annual Conference, pages S3G–S3G. IEEE, 2005
2005
-
[28]
Inclusion of deaf students in computer science classes using real-time speech transcription
Richard Kheir and Thomas Way. Inclusion of deaf students in computer science classes using real-time speech transcription. ACM Sigcse Bulletin , 39(3):261–265, 2007
2007
-
[29]
Recognition means more than just getting the words right: Beyond accuracy to readability
Ross Stuckless. Recognition means more than just getting the words right: Beyond accuracy to readability. Speech Technology, 1:30–35, 1999
1999
-
[30]
A. Hannun. Deep speech: Lessons from deep learning, 2015
2015
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.