{"id":"4648ee0c-3ea8-41ed-8474-ac4383b1149f","arxiv_id":"1909.01176","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"The abstract claims commercial ASR reaches about 78% word error rate on deaf speech versus 18% on hearing speech, but the body of the paper reports no protocol for that comparison and only cites 53% WER for the most intelligible deaf speakers.","lead":"This paper reports on using commercial speech recognition apps in face-to-face conversation, and claims that deaf and hard of hearing speakers get far higher word error rates than hearing speakers. It matters because voice-controlled assistants and dictation tools are becoming default interfaces, and the paper argues these tools currently exclude deaf and hard of hearing users.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 78% vs 18% WER comparison is absent from the body; the central usability claim rests on informal author self-use with no protocol, so the headline quantitative finding is unverifiable.","rationale":"The body is an honest qualitative experience report, and those observations are consistent with prior work. The problem is that the abstract adds a precise quantitative comparison that the body cannot support. Section 2.3 gives a 53% WER for the most intelligible deaf speakers from personal testing but no method; the abstract's 78% WER is not even the same number and has no corresponding section. Section 5's evaluation is informal, by the authors themselves, across seven apps with no per-app metrics or controls. Section 7's call for a future experiment and survey is an in-text admission that the present work does not supply the evidence needed for a population-level usability claim. Under the review rule that statements of missing support count as evidence, this acknowledgement strengthens the REJECT outcome. The qualitative content should be preserved in a revision, with the quantitative headline either removed or fully supported. Because the reader already reached REJECT on essentially these grounds, no change in verdict is needed.","tokens_in":8170,"tokens_out":4756,"duration_ms":47026,"concrete_test":"Ask the authors for the raw data behind the abstract's 78% and 18% WER numbers, specifically the audio corpus, speaker intelligibility ratings, recognizer names and versions, reference transcripts, and scoring scripts. Independently recompute the WERs with a standard aligner such as SCLITE. If the data cannot be produced or the recomputed values do not match the reported rates, the quantitative headline fails and should be removed from the abstract.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, as stated in the abstract, is a quantitative one: deaf speech had approximately a 78% WER versus 18% for hearing speech, and therefore current speech-controlled interfaces are not usable by DHH people. That exact comparison does not appear in the body. The only deaf-speech WER reported is in Section 2.3: \"amongst deaf people rated 5.0, ASR technology had a Word Error Rate of about 53%,\" attributed to personal testing with no sample size, no recognizer names or versions, no utterance corpus, and no scoring definition. Section 5 describes the authors' own use of seven apps in everyday settings but provides no per-app accuracy data, no matched hearing controls, and no measure of variability across speakers. Section 7 states that \"it would be very beneficial to have an experiment along with a survey,\" which explicitly marks the current report as lacking the experimental basis needed for the abstract's claim. The 78%/18% headline is therefore a claim without derivation: if the numbers are wrong or unrepresentative, the usability conclusion is unsupported. The qualitative observations may be valuable, but they cannot carry the quantitative conclusion as written.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is an experience report on the accessibility barriers that deaf, hard of hearing, and hearing people face when using automatic speech recognition (ASR) applications in mixed-group conversation. It describes deaf speech intelligibility ratings from a figure of about 650 deaf people, reports a word error rate of about 53% for deaf speakers rated 5.0 in personal testing, and recounts the authors' use of seven commercial ASR apps in real-world settings such as classrooms, interviews, and conversations. The abstract makes a stronger quantitative claim not present in the body: that deaf speech had approximately a 78% WER compared to 18% for hearing speech, and that current speech-controlled interfaces are not usable by deaf and hard of hearing people.","tokens_in":8310,"tokens_out":3510,"duration_ms":37888,"significance":"If the headline numbers were supported, they would represent an important and actionable accessibility finding with direct design implications for commercial speech interfaces. The qualitative lessons and recommendations in Section 6, such as the need for visible text output, external microphone support, and control of lag, may be useful to practitioners building accessible ASR tools. The paper does not, however, provide a verifiable basis for its central quantitative claim: the 78%/18% comparison is absent from the body, the only reported deaf-speech WER is a single unprotocoled 53% figure, and the user-experience evaluation in Section 5 is informal self-use with no sample size or controls. The paper's value is as an experience report, not as evidence for the strong, general conclusion stated in the abstract.","major_comments":[{"comment":"The abstract's headline quantities—deaf speech WER approximately 78% versus hearing speech WER approximately 18%—never appear in the body; the only deaf-speech WER reported is about 53% in Section 2.3, attributed to 'personal testing' with no recognizer names or versions, no utterance corpus, no sample size, and no scoring definition. This omission makes the paper's central quantitative finding unverifiable and requires either removing the numbers or adding a full methods description and data.","section":"Abstract and Section 2.3"},{"comment":"The evaluation in Section 5 consists of the authors using seven apps on their personal devices in unspecified everyday settings, with no participant count, no per-app accuracy or latency statistics, no measure of variability across speakers, and no matched hearing controls. From this informal self-evaluation one cannot infer the abstract's universal conclusion that 'current speech-controlled interfaces are not usable by DHH people.'","section":"Section 5"},{"comment":"Section 7 states that 'it would be very beneficial to have an experiment along with a survey' and to 'recruit people with different backgrounds,' which directly acknowledges that the systematic evidence needed for the abstract's quantitative and universal usability claim has not been collected. The conclusion also refers to 'one of our studies' for the WER variance, but that study is never named or described in the manuscript, so the paper cites its own unreported data as evidence.","section":"Section 7"}],"minor_comments":[{"comment":"The affiliation line contains a typo, 'Rochester Institute of T echnology', which should be corrected.","section":"Title page"},{"comment":"There are several wording slips, for example 'Ava is is an eponymous product' and later 'they does not mean they will actually use it'; these should be copyedited.","section":"Section 5"},{"comment":"Reference [2] lists the authors as 'undeﬁned undeﬁned undeﬁned' and reference [4] lacks publisher and year; both entries need to be completed.","section":"References"},{"comment":"Figure 2 is described as showing the distribution of about 650 deaf speakers' intelligibility ratings, but no source, axis labels, or methodology for these ratings is provided; if the figure is retained, it needs a proper caption and citation.","section":"Figure 2"}],"recommendation":"reject","confidential_remarks":"The gap between the abstract's claim and the evidence in the body is not a presentation issue; it is a missing empirical basis. The qualitative experience report could be publishable if reframed as such and stripped of the unverified WER numbers, but the manuscript in its current form overclaims in a way that cannot be repaired by local edits. I would encourage the authors to conduct the experiment and survey they themselves call for, and to resubmit with actual data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nYou should know this one is two papers stapled together: the abstract advertises a quantitative result—78% WER for deaf speech versus 18% for hearing speech—that never appears in the body. The body itself is a modest, honest experience report about using seven ASR apps in everyday settings. Those two things do not match, and the mismatch is the whole story.\n\nWhat's genuinely there: the authors describe real use of DEAFCOM, Dragon Dictation, Siri, Virtual Voice, Ava, Google Assistant, and Alexa in classrooms, conversations, job interviews, and speech practice. The lessons learned in Section 6 are the kind of practitioner observations that can be useful. The paper also correctly leans on prior work (Venkatagiri, Jeyalakshmi, Kushalnagar et al.) for the claim that deaf speech is harder for ASR. If this were framed as an experience report with no quantitative pretensions, it would be a reasonable short paper.\n\nThe soft spot is the abstract and the Section 7 conclusions. The 78% vs 18% WER is not derived anywhere. The only number in the body is a 53% WER for deaf speakers rated 5.0 on a clinician scale, attributed to 'personal testing,' with no sample size, no recognizer, no corpus, no scoring. Section 5 is self-use by the authors with no per-app data and no hearing controls. The conclusion even admits 'it would be very beneficial to have an experiment along with a survey,' which says plainly that the experimental basis is missing. So the headline claim is a claim without derivation.\n\nIs the qualitative part sufficient to carry the paper? Not as written, because the abstract and title sell the quantitative finding. The fix is simple: rewrite the abstract to say 'we report our experiences with seven ASR apps and note that deaf speakers we tested had high error rates,' and either remove the 78%/18% numbers or include the full protocol and data. Until then, the paper cannot be accepted.\n\nCite? Not for the numbers. The lessons learned might be citable in an accessibility survey, but I wouldn't build on this. Reading group? Maybe, as a case study in abstract–body mismatch. Peer review? I would desk reject this version and invite a resubmission with an honest abstract. The underlying experience report is worth preserving, but the current one is not trustable as a quantitative study.","headline":"The abstract's headline WER numbers don't appear in the body; what's left is an honest but thin experience report.","tokens_in":8949,"tokens_out":3403,"would_cite":false,"duration_ms":34802,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Commercial speech recognizers fail on deaf users' speech, with a 78% word error rate.","keywords":["automatic speech recognition","deaf and hard of hearing","accessibility","word error rate","voice-controlled interfaces","speech-controlled devices","mixed-group conversation","usability evaluation"],"falsifier":"Run a controlled study in which a diverse group of deaf, hard of hearing, and hearing speakers read the same sentences to the same set of commercial ASR apps in both quiet and noisy conditions, and compare word error rates and user ratings. If median deaf-speech WER falls below the roughly 15% implied by the 85% accuracy threshold, or if a substantial share of DHH participants find the apps usable in real tasks, the paper's central claim would be contradicted.","tokens_in":7854,"feed_emoji":"🗣️","tokens_out":6948,"duration_ms":63886,"temperature":0.7,"pith_summary":"This paper argues that current commercial speech-controlled interfaces—voice assistants, dictation apps, and speech-to-text apps—are not usable by deaf and hard of hearing users. The authors report that deaf speech produced roughly a 78% word error rate in these systems, compared with roughly 18% for hearing speech. Drawing on their own real-world use of seven off-the-shelf apps in classrooms, conversations, job interviews, and speech practice, they contend that accuracy and lag fall below the threshold needed for communication, especially in noisy multi-speaker settings. If this is right, the shift toward voice-first devices leaves a large population unable to use a growing class of mainstream technology.","feed_headline":"Voice interfaces fail deaf users at 78% word error rate","feed_subtitle":"Deaf speech is misheard far more than hearing speech, making voice assistants unusable today.","key_machinery":"The carrying object is word error rate (WER), the metric by which the paper judges commercial ASR engines, paired with an experience report of seven no-cost apps (DEAFCOM, Dragon Dictation, Siri, Virtual Voice, Ava, Google Assistant, and Amazon Alexa) used in everyday settings. The mechanism behind the high WER is that third-generation deep-learning recognizers, trained on large corpora of typical hearing speech, generalize to mild accents but not to the wider variation in pitch, formants, segmental articulation, and prosody found in deaf speech; the paper cites evidence that pitch and formant variation in deaf children is too large to serve as reliable recognition features. This combination lets the paper separate the core recognition failure from usability problems such as lag, noise, and interface design, and argue that fixing the interface alone is not enough.","core_discovery":"The central claim is that automatic speech recognition built into popular commercial products performs so poorly on deaf and hard of hearing speakers that voice-controlled interfaces cannot serve them as they are. Measured on the authors' own speech, word error rate was roughly 78% for deaf speech against roughly 18% for hearing speech, far above the 85–90% accuracy that prior work says is needed for a transcript to be useful, and above the 98% accuracy argued to preserve meaning and intent. The authors attribute the failure to a mismatch between acoustic models trained mainly on hearing speech and the segmental and prosodic characteristics of deaf speech, such as pitch, formants, rate, pausing, volume, intonation, and stress. They also document secondary barriers: latency that grows in noisy settings, random inserted text, lack of feedback about speech volume and microphone placement, and interfaces that do not readily support text input or on-the-fly correction. The paper concludes that significant advances in speech recognition software or alternative input approaches are needed before these systems are accessible to DHH users.","pith_inferences":["Beyond the paper, a controlled study with a larger, diverse DHH sample and the same apps would probably show that WER varies widely by speaker, app, and noise level; the 78% figure is an anchor from the authors' own use, not a measured population statistic.","The paper implies but does not say that the 'trained on majority speech' failure mode extends to other atypical speech populations, so a public benchmark built from deaf speech could double as a general test of inclusive ASR.","The reported success of one deaf user with Amazon Alexa suggests that speaker-adaptive systems may reach some deaf users before full general-purpose ASR does, making per-speaker adaptation a promising testable extension.","A direct comparison of deaf and hearing speech under identical noise and multi-speaker conditions would separate acoustic-model failure from interface lag and microphone issues; the paper's informal design does not allow that separation."],"forward_implications":["Current voice assistants and dictation apps cannot be considered accessible to deaf and hard of hearing users, because their WER on deaf speech exceeds published usability thresholds.","Access to speech-controlled devices for DHH users requires either retraining acoustic models on deaf speech or a new recognition paradigm; simple interface patches will not close the gap.","Even deaf speakers whose voices are intelligible to hearing people may still get unusable ASR results, since the paper found large WER variance among speakers rated high on speech intelligibility.","The gap between the reported WERs means that deaf users' speech is misrecognized more than four times as often as hearing users' speech, a difference large enough to preclude hands-free interaction in most settings.","Real-time captioning from ASR in noisy, multi-speaker settings remains below the accuracy and latency needed for classroom or workplace use, so DHH users will continue to depend on human captioners, interpreters, or text-to-text alternatives."],"supporting_citations":[{"why":"Quantifies that deaf viewers receive only 50–80% of information even with accurate captions, establishing the access gap that ASR is meant to narrow.","marker":"[1]"},{"why":"Supplies evidence that ASR failed to recognize voice commands from speakers with low speech intelligibility and that those users could not correct dictation errors.","marker":"[5]"},{"why":"Documents larger than normal pitch and formant variation in deaf children's speech, explaining why standard recognition features are unreliable for deaf voices.","marker":"[6]"},{"why":"Earlier classroom captioning study showing that ASR captions are choppy and lag behind, framing the latency problems the authors observed.","marker":"[7]"},{"why":"Provides the deep-learning ASR context and reports about 20% WER under far-field, noise, accent, and multi-talker conditions, the baseline the paper compares with deaf speech.","marker":"[9]"},{"why":"Establishes the 85% accuracy and five-second lag thresholds beyond which students cannot effectively use transcripts.","marker":"[10]"},{"why":"Sets the classroom-usefulness standard at roughly 90% transcript accuracy.","marker":"[11]"},{"why":"Argues that 98% accuracy may be needed to preserve meaning and intent, the stricter bar the paper invokes.","marker":"[12]"}],"fun_headline_variants":["Deaf speech: 78% word error rate on voice assistants","Voice assistants mishear deaf speakers 4x more than hearing","Why Siri can't understand deaf users: 78% error rate","Deaf users: voice control has 78% error rate, unusable","Voice recognition: 78% errors for deaf speech vs 18%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion rests on the premise that the authors' own informal evaluation of seven apps in everyday settings represents the wider deaf and hard of hearing population; if their experiences do not generalize, the claim that current speech-controlled interfaces are not usable by DHH people does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Deaf speech: 78% word error rate on voice assistants","Voice assistants mishear deaf speakers 4x more than hearing","Why Siri can't understand deaf users: 78% error rate","Deaf users: voice control has 78% error rate, unusable","Voice recognition: 78% errors for deaf speech vs 18%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000309,"raw_usage":{"total_tokens":1743,"prompt_tokens":901,"completion_tokens":842,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":517,"completion_tokens_details":{"reasoning_tokens":748}},"tokens_in":517,"tokens_out":842,"duration_ms":7821,"temperature":1.0,"reasoning_tokens":748,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:26:35.592309+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled study in which a diverse group of deaf, hard of hearing, and hearing speakers read the same sentences to the same set of commercial ASR apps in both quiet and noisy conditions, and compare word error rates and user ratings. If median deaf-speech WER falls below the roughly 15% implied by the 85% accuracy threshold, or if a substantial share of DHH participants find the apps usable in real tasks, the paper's central claim would be contradicted.","supporting_citations":[{"cited_title":"Deaf, Hard of Hearing, and Hearing Perspectives on using Automatic Speech Recognition in Conversation","cited_arxiv_id":"1909.01176","evidence_quote":"Quantifies that deaf viewers receive only 50–80% of information even with accurate captions, establishing the access gap that ASR is meant to narrow."},{"cited_title":"application","cited_arxiv_id":null,"evidence_quote":"Supplies evidence that ASR failed to recognize voice commands from speakers with low speech intelligibility and that those users could not correct dictation errors."},{"cited_title":"The following lessons were learned from daily expe- rience with the technology","cited_arxiv_id":null,"evidence_quote":"Documents larger than normal pitch and formant variation in deaf children's speech, explaining why standard recognition features are unreliable for deaf voices."},{"cited_title":"They do not generally consider speech to text translations to be particularly important","cited_arxiv_id":null,"evidence_quote":"Earlier classroom captioning study showing that ASR captions are choppy and lag behind, framing the latency problems the authors observed."},{"cited_title":"Generally most deaf people we met have not had this level of success with the technology yet","cited_arxiv_id":null,"evidence_quote":"Provides the deep-learning ASR context and reports about 20% WER under far-field, noise, accent, and multi-talker conditions, the baseline the paper compares with deaf speech."},{"cited_title":"Several reported that the technology does not work well enough for them or that they felt uncomfort- able using the technology in a public setting","cited_arxiv_id":null,"evidence_quote":"Establishes the 85% accuracy and five-second lag thresholds beyond which students cannot effectively use transcripts."},{"cited_title":"Phones have various con- nection strengths and reliability","cited_arxiv_id":null,"evidence_quote":"Sets the classroom-usefulness standard at roughly 90% transcript accuracy."},{"cited_title":"Many people often men- tion this when discussing why they do not use ASR as much as texting","cited_arxiv_id":null,"evidence_quote":"Argues that 98% accuracy may be needed to preserve meaning and intent, the stricter bar the paper invokes."}],"review_version":1}