{"id":"3f420096-b7ae-4f2a-bccd-cd84cd27db9b","arxiv_id":"1909.01167","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Commercial ASR transcribes deaf and hard-of-hearing speech with about a 77% word error rate, versus 18% for hearing speech.","lead":"The authors measured how often a commercial speech recognition service mis-transcribes sentences spoken by deaf and hard-of-hearing people, finding about 77% of words wrong versus 18% for hearing speakers. The result suggests current voice-controlled devices are effectively unusable for many deaf users, unless ASR is retrained on deaf speech.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim leaps from 78% sentence-level WER on one ASR API to 'speech-controlled interfaces are unusable'; the unvalidated proxy is the weakest link.","rationale":"In good faith, the paper establishes a plausible measurement: Microsoft Translator transcribes a small, deliberately stratified sample of deaf speech with roughly 78% WER and hearing speech with 18% WER. The numbers are internally consistent, no hidden parameters or circular reasoning are apparent, and the direction of the gap is consistent with prior work. The reader's identified weakness (non-representative sampling, unmatched recording conditions) is real but does not threaten the qualitative result: the sample includes a 'good' intelligibility group at 53% WER, and the hearing group's noisy recording would, if anything, raise hearing WER and reduce the gap. The more load-bearing weakness is the inferential leap from sentence-level WER on one ASR API to the blanket conclusion that 'current speech-controlled interfaces are not usable by DHH people.' Consumer voice assistants are command-oriented, with constrained vocabularies, prompt-based flows, and confirmation/repair mechanisms; the cited single-digit result ([3], 13% WER for poor-intelligibility deaf speakers) shows constrained tasks can be much less error-prone than open sentence dictation. Because the manuscript provides no command-level usability data and tests only Microsoft Translator, the central claim is broader than the evidence. This is a correctness-risk issue that a revision can address by narrowing the claim or adding the direct command-task experiment; it does not invalidate the reported WER measurement, so I keep the reader's CONDITIONAL verdict and would not reject or accept the paper as-is.","tokens_in":2840,"tokens_out":7816,"duration_ms":82125,"concrete_test":"Conduct a within-subjects command-task study: the same 45 deaf and 5 hearing speakers speak 15 standard voice-interface commands (e.g., 'set a timer for five minutes', 'call Alex', 'play the news', 'turn on the lights') directly to Amazon Echo and Google Home in a quiet room; score per-command success from device logs. If deaf command-success rates are high (say >80%) while sentence WER remains near 78%, the conclusion that current speech-controlled interfaces are unusable by DHH people is unsupported. If command success is also near zero, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The manuscript's empirical result (Sections 2.2, 3.1, 3.2) is a word error rate of Microsoft Translator on Clarke Sentences. The abstract and Section 4 generalize this to all 'current speech-controlled interfaces' (Alexa, Siri, etc.). For that conclusion to hold, sentence-transcription WER on one cloud API must predict command-level success on consumer voice interfaces. This is load-bearing and untested. Voice interfaces use short, constrained command grammars, command-biased language models, and confirmation/repair loops; they are not pure dictation. The authors themselves cite [3] showing a constrained single-digit task yielded only 13% WER for deaf speakers with poor intelligibility, which demonstrates that the same speakers can perform far better when the recognition task is narrow. Therefore, the observed 77% WER on 10-syllable sentences does not directly establish that a short command would fail, especially if the interface supports confirmation or adaptation. The experiment also tests only Microsoft Translator, not the engines inside the consumer devices (Amazon, Google, Apple); references [6-8] are press announcements, not comparative evaluations. Thus the broad usability claim is an overreach relative to the data. The reader's sample-selection and recording-mismatch concerns are real but less central, because the stratified sample deliberately includes good speakers and the noisy hearing-condition recording biases the hearing WER upward, shrinking the gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a single empirical measurement of automatic speech recognition (ASR) performance on deaf and hard-of-hearing (DHH) speech. The authors sampled 45 Clarke Sentences recordings from a 650-speaker DHH dataset, stratified by a naive listener into good, mediocre, and poor intelligibility groups, and transcribed them with the Microsoft Translator Speech API. They also transcribed recordings from five hearing speakers reading the same sentence lists, recorded in a noisy lab. Word error rates were computed with the NIST SCTK toolkit. The reported results are an average 77% WER for the DHH samples (53% for the 'good' subgroup) versus 18% for the hearing baseline. The authors conclude that current speech-controlled interfaces such as Amazon Echo and Apple's Siri are not usable by DHH people and that alternative input methods or substantial ASR improvements are needed.","tokens_in":3004,"tokens_out":2683,"duration_ms":29579,"significance":"If the reported measurement is valid, the finding that a mainstream commercial ASR system mis-transcribes about three out of four words of DHH speech is an important accessibility data point. The paper's strengths are its use of a real DHH speech corpus with clinician-assigned intelligibility scores, a standard scoring tool (SCTK), and a stratified attempt to include good, mediocre, and poor speakers. The comparison between intelligibility ratings and WER is a useful contribution. However, the paper's central conclusion is broader than the experiment supports: only one specific API (Microsoft Translator) is tested on sentence-length read speech, while the conclusion is about all current speech-controlled interfaces. The sampling of 45 speakers by a single naive listener and the unmatched hearing baseline also limit the quantitative claim. The main value of the paper is as a preliminary negative result that motivates more careful evaluation of ASR accessibility for DHH users.","major_comments":[{"comment":"The conclusion that 'current speech-controlled interfaces are not usable by DHH people' is not supported by the data. The experiment tests only the Microsoft Translator Speech API on 10-syllable Clarke Sentences, not commercial voice interfaces such as Amazon Echo or Siri, which typically use short command grammars, command-specific language models, and confirmation/repair loops. The authors' own cited reference [3] reports a 13% WER for a constrained single-digit task with deaf speakers who have poor intelligibility, demonstrating that narrow recognition tasks can perform far better than free-form sentence transcription. The paper should either test actual voice-command interfaces or explicitly limit the claim to sentence-level dictation with the tested API.","section":"Abstract and Section 4"},{"comment":"The 45 DHH samples were chosen by one naive listener rather than randomly sampled from the 650-speaker dataset, with 15 samples labeled good (intelligibility 40+), 16 mediocre (30-40), and 14 poor (10-30). This selection protocol may bias the sample toward or away from speakers who are typical of the broader DHH population, and no inter-rater reliability or selection criteria are reported. Because the paper's headline 77% WER depends on the composition of these 45 samples, the authors should describe the selection procedure in detail, provide a random or systematic sampling strategy, or at least give the intelligibility distribution of the full dataset to allow readers to judge representativeness.","section":"Section 3.2"},{"comment":"The hearing baseline is based on only five speakers recorded with a cell phone in a noisy lab, whereas the DHH recordings come from an existing corpus recorded under different, unspecified conditions. The authors state that if the hearing speakers were recorded in the same setting as the DHH samples, they would expect an even lower WER, but this does not remove the recording-condition mismatch from the central 77% versus 18% comparison. The paper should report the recording conditions of both sets, ideally use matched recordings, and provide confidence intervals or a statistical comparison for the main difference, given the small n and large variance among DHH speakers.","section":"Section 3.1"},{"comment":"No measures of statistical uncertainty are given for the central estimates. The paper reports only the average WER for each group and a t-test between the good and mediocre DHH subgroups; it does not report confidence intervals for the 77% overall DHH WER or the 18% hearing WER. Given the small sample sizes and the high variability evident in Figure 2, the authors should provide confidence intervals or standard errors for all reported means, ideally with a bootstrap or similar resampling analysis.","section":"Sections 3.1 and 3.2"}],"minor_comments":[{"comment":"The statement that the Microsoft Translator API 'matches other similar transcription software' is supported only by press announcements in references [6-8], not by comparative evaluations. The authors should either cite peer-reviewed comparisons or soften the claim.","section":"Section 2.2"},{"comment":"The text says 'as shown in Figure 3.2' but the figure is labeled 'Figure 3'; the cross-reference should be corrected.","section":"Section 3.2"},{"comment":"The sentence 'We show that current speech-controlled interfaces are not usable by DHH people' appears twice in the abstract; one occurrence should be removed.","section":"Abstract"},{"comment":"References [6], [7], and [8] are web/press items with incomplete bibliographic details; full access dates and publication venues should be added.","section":"Section 5"},{"comment":"The axes and point markers in Figure 2 would benefit from explicit labels and a legend, since the reader must infer which markers correspond to 'good', 'mediocre', and 'poor' groups from the text.","section":"Figure 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript appears to be a short paper based on work presented at ASSETS 2017, but the arXiv version does not clearly state its provenance; the authors should clarify the venue and any differences from the published version. The main technical concern is the gap between the narrow experiment (one API, sentence transcription) and the broad usability conclusion; this gap can be closed by rewriting the conclusion and adding a limitations subsection, rather than by additional experiments, so major revision seems appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful thing here is a clean, reproducible measurement: the authors took 45 deaf/hard-of-hearing recordings from an existing Clarke Sentences corpus, ran them through Microsoft Translator, scored with SCTK, and got a 77% WER, against 18% for 5 hearing controls recorded in a noisy lab. Those specific numbers for a current commercial API are new, and the stratified breakdown (53% for the 'good' intelligibility group) is informative. They also avoided curve-fitting entirely; this is just an honest probe of a named system.\n\nWhere I part ways is the leap from that measurement to the conclusion that 'current speech-controlled interfaces are not usable by DHH people.' The experiment tests sentence-length dictation through one cloud API. It does not test the command grammars, small vocabularies, or confirmation/repair loops in Alexa, Siri, or Google Assistant. And the authors themselves cite a prior digit-recognition task where deaf speakers with poor intelligibility got only 13% WER—which is exactly the kind of constrained condition that voice commands approximate. So the data support 'a current dictation API has high WER on this deaf speech corpus,' not 'all current speech-controlled interfaces are unusable.' That is the paper's main soft spot, and it is load-bearing.\n\nThe other issues are smaller but real. The 45 samples were chosen by one naive listener, not randomly; the hearing baseline is five speakers in a noisier environment; and there are no confidence intervals. The recording mismatch actually biases against the hearing group, so it shrinks the gap rather than inflating it, but it still bothers me for comparing absolute rates. The sample-selection issue could matter if the listener's choices skew toward less-intelligible speakers, though the deliberate inclusion of good speakers partially covers that.\n\nI'd send this to a serious referee. The measurement is worth having as a data point, and the paper is short enough that the overgeneralization can be fixed by rewording and by either adding task-level data or tempering the claims. It is not a wasted submission.","headline":"A useful benchmark of one ASR API on deaf speech, undermined by an overbroad 'interfaces are unusable' claim that the data don't support.","tokens_in":3604,"tokens_out":1949,"would_cite":true,"duration_ms":19896,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper finds that commercial speech recognition mis-transcribes roughly three of every four words spoken by deaf and hard-of-hearing users, about 78% word error rate, versus 18% for hearing speech, making current voice interfaces…","keywords":["deaf speech","automatic speech recognition","word error rate","speech-controlled interfaces","accessibility","Clarke Sentences","deaf and hard-of-hearing"],"falsifier":"Run a larger random sample of deaf speakers through the same commercial ASR under identical recording conditions as hearing speakers and compare word error rates. If the average deaf WER drops below roughly 40% for intelligible speakers, or if any large deaf cohort matches hearing-level WER, the claim that voice interfaces are unusable by DHH people would be weakened; if the 78% versus 18% gap persists across services and newer models, the claim is strengthened.","tokens_in":2565,"feed_emoji":"🗣️","tokens_out":6111,"duration_ms":55886,"temperature":0.7,"pith_summary":"This paper asks whether the shift to voice-controlled devices has left deaf and hard-of-hearing (DHH) users behind. Testing a commercial speech-recognition service on recordings of deaf speakers reading standardized sentences, it finds an average word error rate of about 78% for deaf speech, compared to about 18% for hearing speech. Even deaf speakers judged highly intelligible by a human listener still received a 53% error rate. The authors conclude that current speech-controlled interfaces are not usable by DHH people and that either major advances in speech recognition or alternative input methods are needed.","feed_headline":"Deaf speech gets 78% word error rate on commercial ASR","feed_subtitle":"Hearing speech scored 18%, so current voice interfaces are effectively unusable by deaf and hard-of-hearing users.","key_machinery":"The machinery is a measurement pipeline: recordings of deaf speakers reading Clarke Sentences, a standardized intelligibility test, are fed into a commercial speech-recognition API, and the resulting transcripts are scored with a standard word error rate toolkit. The Clarke score provides a human intelligibility baseline that separates 'good,' 'mediocre,' and 'poor' deaf speech, and the WER comparison quantifies how far commercial ASR lags behind a human listener.","core_discovery":"The central discovery is that commercial automatic speech recognition, trained mostly on hearing voices, does not transfer to deaf speech. Across 45 deaf speech samples the word error rate averaged 77% (which the paper reports as approximately 78%), while five hearing speakers in a noisier environment had an 18% average. Performance varied widely among deaf speakers, and even the 'good' group with intelligibility scores of 40 and above averaged 53% WER. The authors interpret this as showing that current voice interfaces are effectively unusable for DHH users, and that a much larger and more varied database of deaf speech would be needed to train ASR to an acceptable level.","pith_inferences":["The 78% WER probably understates the functional problem for command interfaces, because a mis-recognized command may still be a recognized word but the wrong one; task-completion rates might be even lower than word accuracy suggests.","Testing the same deaf speakers across several commercial ASR services (phone assistants, smart speakers, dictation) could reveal whether the gap is intrinsic to deaf speech or exaggerated by one vendor's training data.","A plausible path forward, not explored in the paper, is few-shot personalization: modern ASR models could adapt to an individual deaf speaker with a small amount of their own speech during device setup."],"forward_implications":["Voice-only devices such as smart speakers and phone assistants will continue to mishear deaf users, so they cannot serve as a primary interface without dedicated training on deaf speech.","Producing deaf-inclusive ASR training data is a hard practical problem because the DHH population is small and hugely varied, so per-user adaptation or alternative input pathways may be necessary.","Human intelligibility ratings do not predict ASR accuracy at the top end (53% WER for 'good' speech), so usability testing for voice interfaces must use actual ASR output, not human ratings.","An automated WER-based feedback tool could replace the Clarke test for practical purposes, giving DHH users a realistic measure of how well voice interfaces understand them."],"supporting_citations":[{"why":"Supplies the Clarke Sentences intelligibility test and shows its validity as a measure of deaf speech intelligibility, the basis for selecting and rating the samples.","marker":"[5]"},{"why":"Provides the standard word error rate scoring toolkit used to compute every WER in the study.","marker":"[9]"},{"why":"Earlier evidence that ASR on deaf speech with limited vocabulary has high error rates compared to hearing speech, motivating the current comparison.","marker":"[3]"},{"why":"Establishes that the commercial API used in the study represents state-of-the-art speech recognition, so its poor deaf-speech performance is not due to an outdated system.","marker":"[6]"},{"why":"Also establishes the commercial recognition system's state-of-the-art status, supporting the claim that the gap reflects current technology's limitation.","marker":"[8]"}],"fun_headline_variants":["Voice interfaces fail deaf users: 78% word error","Commercial ASR stumbles on deaf speech: 78% WER","Deaf speech: 78% error on voice assistants","Voice assistants unusable for deaf: 78% word error","ASR for deaf speech: 78% error vs 18% hearing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion depends on the 45 deaf speech samples selected by one naive listener being representative of DHH speakers generally, and on the deaf and hearing recordings being comparable even though the hearing group was recorded in a noisy lab with only five subjects.","fun_headline_variants_meta":{"raw":{"variants":["Voice interfaces fail deaf users: 78% word error","Commercial ASR stumbles on deaf speech: 78% WER","Deaf speech: 78% error on voice assistants","Voice assistants unusable for deaf: 78% word error","ASR for deaf speech: 78% error vs 18% hearing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000458,"raw_usage":{"total_tokens":2231,"prompt_tokens":817,"completion_tokens":1414,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":433,"completion_tokens_details":{"reasoning_tokens":1325}},"tokens_in":433,"tokens_out":1414,"duration_ms":9440,"temperature":1.0,"reasoning_tokens":1325,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:25:03.496460+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a larger random sample of deaf speakers through the same commercial ASR under identical recording conditions as hearing speakers and compare word error rates. If the average deaf WER drops below roughly 40% for intelligible speakers, or if any large deaf cohort matches hearing-level WER, the claim that voice interfaces are unusable by DHH people would be weakened; if the 78% versus 18% gap persists across services and newer models, the claim is strengthened.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Clarke Sentences intelligibility test and shows its validity as a measure of deaf speech intelligibility, the basis for selecting and rating the samples."},{"cited_title":"Samar and Dale Evan Metz","cited_arxiv_id":null,"evidence_quote":"Provides the standard word error rate scoring toolkit used to compute every WER in the study."},{"cited_title":"The average WER was 18%","cited_arxiv_id":null,"evidence_quote":"Earlier evidence that ASR on deaf speech with limited vocabulary has high error rates compared to hearing speech, motivating the current comparison."},{"cited_title":"Osberger and Nancy S","cited_arxiv_id":null,"evidence_quote":"Establishes that the commercial API used in the study represents state-of-the-art speech recognition, so its poor deaf-speech performance is not due to an outdated system."},{"cited_title":"Phonetically sensitive discriminants for improved speech recognition","cited_arxiv_id":null,"evidence_quote":"Also establishes the commercial recognition system's state-of-the-art status, supporting the claim that the gap reflects current technology's limitation."}],"review_version":1}