{"id":"331f7953-4396-4682-9251-79996e5a02e7","arxiv_id":"1909.02853","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Commercial ASR services produce high word error rates for deaf and hard-of-hearing speech even for the most intelligible speakers, and a custom vocabulary model does not significantly improve recognition.","lead":"The paper measured how well three commercial speech-recognition services understand recordings of deaf and hard-of-hearing speakers, using 45 recordings grouped by how easily a listener could understand them. It found that word error rates were high in all groups, that even the clearest speakers got unpredictable results, and that adding a custom vocabulary list did not significantly improve recognition for this population.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing assumption is the one-listener good/fine/bad categorization: with no inter-rater reliability or reported per-file intelligibility scores, every stratified WER comparison and the population-level conclusion rests on an unvalidated grouping.","rationale":"The reader's weakest_assumption correctly identifies the one-listener categorization as the load-bearing point, and my reading agrees. The central claim is not merely that some DHH voices get high WERs; it is that voices falling into 'bad' or 'fine' categories will not achieve parity and that 'good' voices give unpredictable results. All of this generalizes from a grouping created by one listener with no reliability evidence. The custom-vocabulary leakage (the MSPPT model was seeded with the full Clarke Sentences list, which is the test material) is a real flaw, but it biases toward improvement, so the null customization result is not the fragile part of the argument; the category system is. The statistical treatment also treats repeated WER measurements from the same audio across three ASRs as independent, which likely narrows confidence intervals, but the qualitative high-WER finding would survive a more conservative analysis. Since the manuscript already warrants a conditional verdict due to these limitations, my stress-test does not move the verdict; it sharpens the specific check that would either validate or undermine the category-level conclusion.","tokens_in":4875,"tokens_out":3732,"duration_ms":41119,"concrete_test":"Have at least three naive listeners independently re-categorize the same 45 recordings (or a fresh stratified sample from the NTID Clarke Sentences corpus) into good/fine/bad and compute Fleiss' kappa. If kappa is below about 0.6, re-run the Table 1 between-category t-tests using only recordings with unanimous labels, or use the speech pathologist's per-file intelligibility scores as a continuous predictor of WER. If the unanimity-based or continuous analyses change which category comparisons are significant, the category-driven conclusion is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 'Methodology: Audio Dataset' states that 45 files were categorized by a single naive listener into good/fine/bad, with average intelligibility scores 25/43/48, but gives no selection procedure, no category definitions, no reliability check, and no per-file intelligibility scores. Table 1 and the one-sample t-test CIs then compare WER across these categories, and the Conclusion generalizes to 'if their voice fell within the bad or fine audio categories.' If the categories are not stable across listeners or do not reflect separable acoustic/intelligibility groups, the non-significant bad-fine tests are uninterpretable and the 'good' versus 'bad' contrast is not a reproducible basis for predicting ASR performance. The external comparison to 5-6% WER is also unsupported by an in-study non-DHH baseline, but the subjective category system is more load-bearing because it is the variable that organizes the entire analysis. Even with this concern, the raw WER values are high, so the qualitative conclusion that these particular DHH voices are poorly recognized is likely robust; what is jeopardized is the category-level generalization.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper evaluates the performance of three commercial automatic speech recognition (ASR) services—IBM Watson, Microsoft Translator, and Microsoft Presentation Translator with a custom vocabulary—on recordings of Deaf and Hard-of-Hearing (DHH) speech. The authors use 45 audio files from the NTID Clarke Sentences intelligibility corpus, categorized into 'good', 'fine', and 'bad' by a single naive listener, and compute word error rates (WER) with the NIST SCTK tool. They report high WER across all categories, with 95% confidence intervals of (91.338, 97.443) for 'bad', (82.109, 91.316) for 'fine', and (51.288, 66.068) for 'good', and find no statistically significant improvement from the custom vocabulary model. The conclusion states that DHH individuals cannot achieve WERs comparable to non-DHH users if their voice falls in the 'bad' or 'fine' categories, and that even 'good' voices yield unpredictable ASR results.","tokens_in":4960,"tokens_out":3529,"duration_ms":36016,"significance":"If the results hold, the paper provides valuable empirical evidence on the accessibility gap in commercial ASR for DHH users, addressing an important and under-studied problem. The use of standard WER scoring, real public ASR services, and a well-known intelligibility corpus are strengths. The qualitative finding that these particular DHH voices are poorly recognized is likely robust. However, the paper's quantitative claims and category-level generalizations rest on a single listener's subjective categorization, a small convenience sample, and an external non-DHH baseline, which limit the strength of the conclusions. The paper is a useful preliminary study but not a fully controlled evaluation.","major_comments":[{"comment":"The partition of the 45 audio files into 'good', 'fine', and 'bad' categories rests entirely on the judgment of a single naive listener, with no category definitions, no selection procedure, and no inter-rater reliability assessment. Because Table 1 and the conclusion (\"if their voice fell within the 'bad' or 'fine' audio categories\") stratify all analyses by these categories, the category-level claims are not reproducible or independently verifiable. The paper should either provide a second or third rater and report agreement (e.g., Cohen's kappa), or define the categories using objective acoustic or per-file intelligibility criteria.","section":"Methodology: Audio Dataset"},{"comment":"The claim that DHH speech yields WERs far above the 5-6% range for non-DHH speech is supported only by external newspaper/blog citations and not by an in-study non-DHH baseline. The one-sample t-test confidence intervals (91.338, 97.443) for 'bad', (82.109, 91.316) for 'fine', and (51.288, 66.068) for 'good' describe the DHH sample alone; they do not statistically test against a matched control group. Without a control group recorded and scored under the same protocol, the comparison to 5-6% is informal. The paper should either add a non-DHH control condition or explicitly frame the study as descriptive and avoid comparative language such as 'would not be able to achieve equal WERs'.","section":"Results"},{"comment":"The null result for customization is based on only 14-16 speakers per category with very high variance, and the paper reports no power analysis or effect-size confidence intervals. A non-significant p-value (the lowest reported is .5472) does not support the conclusion that customization provides no improvement; it only indicates that the study had insufficient power to detect an effect. Additionally, the MSPPT custom model was pre-seeded with the entire list of Clarke Sentences—the same sentences used in the evaluation—so the customization condition is not representative of a real context-aware deployment. This contamination could bias results in either direction and should be acknowledged or remedied with a held-out keyword list.","section":"Improvements in WER for Deaf and Hard-of-Hearing Speech with Context Awareness"},{"comment":"The sentence \"DHH individuals would not be able to achieve equal WERs as the non-DHH population if their voice fell within the 'bad' or 'fine' audio categories\" overgeneralizes from a sample of 45 speakers, of whom only 30 fall in the 'bad' and 'fine' groups. The reported confidence intervals are intervals for the sample mean, not prediction intervals for future DHH speakers, and the sample is a convenience subset of a larger corpus. The conclusion should be tempered to the sample or supported by a mixed-effects model treating speaker as a random effect, so that population-level inferences are properly quantified.","section":"Conclusion"}],"minor_comments":[{"comment":"The header 'Bonferroni correlation results' is imprecise; the paper appears to apply a Bonferroni correction for multiple comparisons, not a correlation. Please rename it 'Bonferroni-corrected p-values'.","section":"Table 1"},{"comment":"The explanation of WER states that the total of substitutions, deletions, and insertions is divided by the total number of words in the reference transcript. This is correct, but the sentence could be tightened to explicitly note that insertions increase the numerator without affecting the denominator, as this is a common point of confusion.","section":"WER Analysis"},{"comment":"There is a typo: 'Even if their voice is \"good\", is still likely that they will get unpredictable results' should read 'Even if their voice is \"good\", it is still likely that they will get unpredictable results'.","section":"Conclusion"},{"comment":"The paper would benefit from a reproducibility statement indicating whether the audio subset, reference transcripts, and ASR outputs are available, since the WER calculations can be exactly reproduced only if these materials are shared.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is a short extended abstract and should be judged as such. The core measurement—high WER for these DHH voices—is credible and useful to the accessibility community, so I do not recommend rejection. However, the category-level analysis and the customization null result need methodological support before the claims can be accepted. The authors may be able to address the major comments with additional analyses of existing data (e.g., per-file intelligibility scores, multiple raters) rather than new data collection, so I see this as a feasible major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing to know: this is a straightforward, honest measurement of three ASR engines on recorded DHH speech, and its core claim—WER is very high, between about 50% and 97% depending on how intelligible a voice sounds—almost certainly holds for this dataset. The genuinely new bit relative to the author's 2017 ASSETS paper is the custom-vocabulary comparison (MSPPT vs MS) and the Bonferroni-corrected category contrasts. That part is a legitimate extension.\n\nWhat it does well: it uses a standard scoring tool (SCTK), explains transcript cleaning, reports confidence intervals, and does not oversell the custom model. The null result for customization is actually meaningful because the custom model was given the exact Clarke sentence list—the most favorable configuration—and still did not help. The paper is also short and readable, which counts for something.\n\nThe soft spots are real, though. The load-bearing weakness is the one-listener good/fine/bad categorization. No inter-rater reliability, no selection procedure, no per-file intelligibility scores. All the stratified comparisons and the conclusion that bad and fine are equivalent rest on that subjective grouping. A second listener could easily shift category boundaries and change which t-tests are significant. There is also no in-study non-DHH baseline; the 5-6% WER comes from press releases that used different tasks and conditions. The null result for customization is interpreted as \"no improvement,\" but with n=14-16 per group and no power analysis, the test may simply be underpowered. And the custom model was seeded with the evaluation sentence list, which biases toward finding improvement; that the result was still null makes the no-benefit conclusion stronger for the optimistic case, but weaker for realistic deployments. Data and code are not available, and the cloud services have changed, so exact replication is impossible.\n\nNone of this sinks the main message. Even the \"good\" voices landed in the 51-66% WER band with wide variance, and \"bad\"/\"fine\" are in the 80s-90s. Anyone building speech interfaces should not assume DHH users will get the advertised near-human accuracy.\n\nBottom line: this is a useful short paper for accessibility and HCI researchers and for ASR product teams. It deserves a serious referee because the population is underserved and the measurement is transparent. A revision should add a second rater or use the pathologist scores directly, run a non-DHH baseline through the same pipeline, and soften the \"bad equals fine\" claim to match the evidence.","headline":"A small, candid usability study showing commercial ASR still fails on DHH speech; the headline result is credible, the category-level analysis is on shakier ground.","tokens_in":5589,"tokens_out":2043,"would_cite":false,"duration_ms":21442,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Current speech recognition services are largely unusable for Deaf and Hard-of-Hearing voices, and custom vocabulary models do not change that.","keywords":["Automatic Speech Recognition","Deaf and Hard-of-Hearing","Word Error Rate","Custom Language Model","Speech Usability","Accessibility","Voice Interfaces"],"falsifier":"Have several naive listeners independently label the same 45 recordings as 'good', 'fine', or 'bad' and compute agreement; if the labels do not reproduce, the group-level WER comparisons and the 'bad equals fine' equivalence in Table 1 are not supported.","tokens_in":4553,"feed_emoji":"🗣️","tokens_out":7179,"duration_ms":67892,"temperature":0.7,"pith_summary":"This paper asks whether today's commercial speech recognition services can understand Deaf and Hard-of-Hearing (DHH) speakers, and whether giving the service a custom vocabulary list closes the gap. Using sentence-list recordings from 45 DHH speakers sorted by a naive listener into 'good', 'fine', or 'bad' audio, it measures word error rates (WER) for three ASR configurations. The paper finds very high WERs for all categories, with 95% confidence intervals of roughly 91-97% for 'bad', 82-91% for 'fine', and 51-66% for 'good'. The custom vocabulary model did not significantly improve any category. The practical stakes are direct: voice interfaces in phones, cars, and home assistants will not serve DHH users at the error levels reported here.","feed_headline":"Speech recognition gets over 90% of words wrong for many Deaf voices","feed_subtitle":"Even 'good'-sounding recordings kept median word errors above 45%, and custom vocabulary did not help.","key_machinery":"The argument runs on Word Error Rate (WER), the standard measure in which a recognized transcript is aligned to a reference and substitutions, deletions, and insertions are counted as a fraction of reference words. WER is computed for each recording under three ASR configurations and compared across three audio-quality categories ('bad', 'fine', 'good') assigned by a single naive listener, with the custom-vocabulary configuration tested against its matching base configuration. The 'custom vocabulary' mechanism is the paper's key intervention: it provides the ASR with a keyword list to give it context awareness, and the paper tests whether that intervention changes the WER pattern for DHH speech.","core_discovery":"The central claim is that DHH speech, as represented by this dataset, is not usable with current commercial ASR services. The 95% confidence intervals for WER are (91.338, 97.443) for the 'bad' audio category, (82.109, 91.316) for 'fine', and (51.288, 66.068) for 'good'. For 'bad' and 'fine' audio, the differences between categories were not statistically significant after Bonferroni correction, and all three services behaved similarly. Adding a custom vocabulary model to one service shifted the median for 'good' audio by a little more than 10% but with a standard deviation above 20%, and the difference was not significant; for 'bad' and 'fine' audio there was essentially no improvement. The paper concludes that DHH users with 'bad' or 'fine' voices cannot achieve word error rates comparable to the 5-6% reported for non-DHH speech, and that even 'good' voices will give unpredictable results.","pith_inferences":["A direct extension would be a per-speaker calibration study: the large variance in the 'good' category suggests some DHH voices may already be near usable WER, so testing each speaker's repeated utterances could identify who can rely on current ASR.","The failure of the vocabulary list to help implies the bottleneck is acoustic, not lexical; a natural next test is custom acoustic models or speaker adaptation rather than keyword lists.","Because the categories rest on one listener's labels, a replication using several independent listeners and a pre-registered definition of 'good', 'fine', and 'bad' would show whether the bad/fine equivalence is a property of DHH speech or of the labeling scheme."],"forward_implications":["A DHH speaker whose voice sounds 'bad' or 'fine' to a naive listener can expect word error rates above 80%, so spoken output would need extensive correction to be usable.","Even 'good'-sounding DHH speech produced median WERs above 45% with large spread, so voice interfaces cannot rely on human clarity judgments to predict ASR success.","Adding a custom vocabulary list to a commercial ASR service did not produce a significant WER improvement in any audio category in this dataset.","Interface designers should not treat ASR as an accessible input method for DHH users unless per-speaker or acoustic-model adaptation is shown to work."],"supporting_citations":[{"why":"Reports the 5-6% word error rate for non-DHH speech that frames the parity goal.","marker":"[1]"},{"why":"Another source for the same low-WER baseline used in the comparison.","marker":"[2]"},{"why":"A third source for the low-WER baseline against which DHH results are judged.","marker":"[3]"},{"why":"Documents the variability of DHH speech patterns, used in the conclusion to explain unpredictable ASR results.","marker":"[4]"},{"why":"Provides the context-aware deep-neural-network mechanism behind the custom vocabulary feature the paper tests.","marker":"[5]"},{"why":"The preliminary study this work extends; also the source of the 45 recordings and their good/fine/bad labels.","marker":"[6]"},{"why":"The sentence intelligibility test from which the dataset was drawn.","marker":"[7]"},{"why":"Shows DHH speech is hard for human listeners too, motivating the expectation of ASR difficulty.","marker":"[8]"},{"why":"The scoring toolkit used to compute Word Error Rates in the WER analysis.","marker":"[9]"}],"fun_headline_variants":["ASR fails Deaf users: word error rates top 90%","Deaf speech still breaks speech recognition: WER above 90%","Custom vocab fails to fix ASR for Deaf voices","Speech recognition unusable for many Deaf speakers","Even good Deaf speech trips up ASR, WER 51-66%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire grouping of recordings into 'good', 'fine', and 'bad' comes from one listener's judgment, with no independent check; if another listener would sort them differently, the statistical comparisons lose their anchor.","fun_headline_variants_meta":{"raw":{"variants":["ASR fails Deaf users: word error rates top 90%","Deaf speech still breaks speech recognition: WER above 90%","Custom vocab fails to fix ASR for Deaf voices","Speech recognition unusable for many Deaf speakers","Even good Deaf speech trips up ASR, WER 51-66%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000163,"raw_usage":{"total_tokens":1238,"prompt_tokens":935,"completion_tokens":303,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":216}},"tokens_in":551,"tokens_out":303,"duration_ms":3409,"temperature":1.0,"reasoning_tokens":216,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:25:46.405656+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have several naive listeners independently label the same 45 recordings as 'good', 'fine', or 'bad' and compute agreement; if the labels do not reproduce, the group-level WER comparisons and the 'bad equals fine' equivalence in Table 1 are not supported.","supporting_citations":[{"cited_title":"Microsoft researchers achieve new conversational speech recognition milestone","cited_arxiv_id":null,"evidence_quote":"Another source for the same low-WER baseline used in the comparison."},{"cited_title":"Reaching new records in speech recognition","cited_arxiv_id":null,"evidence_quote":"A third source for the low-WER baseline against which DHH results are judged."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the context-aware deep-neural-network mechanism behind the custom vocabulary feature the paper tests."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The sentence intelligibility test from which the dataset was drawn."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows DHH speech is hard for human listeners too, motivating the expectation of ASR difficulty."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The scoring toolkit used to compute Word Error Rates in the WER analysis."},{"cited_title":"Making sense of Google CEO Sundar Pichai’s plan to move every direction at once","cited_arxiv_id":null,"evidence_quote":"Reports the 5-6% word error rate for non-DHH speech that frames the parity goal."}],"review_version":1}