{"id":"be3b1153-9cf5-4a8c-ae46-eca8d55429a0","arxiv_id":"2505.10596","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A decade-long audit finds persistent inclusivity gaps in speech AI for healthcare: English-heavy datasets, little demographic metadata, no speech-impaired samples, and limited bias research.","lead":"A review of speech datasets and healthcare AI papers from 2015 to 2024 finds they mostly cover English, standard accents, and narrow demographics, with almost no speech-impaired voices. It argues these gaps can make AI speech tools less reliable for marginalized patients and calls for more inclusive data and policy.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline '10%' is contradicted by the paper's own Table 3: 479 inclusive vs 676 control is ~71%, not 10%; the central quantitative claim has no reported basis.","rationale":"Good-faith reading: the paper is a descriptive audit, and its qualitative direction, that speech AI resources skew toward English, standard accents, and narrow demographics, is plausible and consistent with prior work. The load-bearing problem is narrower: the paper's central quantitative claim, the '10%' figure, is contradicted by the counts in its own Table 3. The reader's weakest assumption (unvalidated, low-precision OpenAlex queries) is related but less direct: even if the queries were exactly as described, 479/676 is about 71%, not 10%, so the claim cannot be checked. The discrepancy is not cosmetic because the abstract and Section 3.2 use 'just 10%' to argue inclusivity remains a niche area; a figure of 71% would change the qualitative emphasis, and it would also suggest the full-text queries are too broad. Additional internal inconsistencies on the dataset side (38 vs 41 datasets, conflicting CoVoST2/VoxPopuli language counts, and absence of obvious candidates like Common Voice) reinforce the need for a careful reproduction pass. The qualitative conclusion is not overturned, so I do not move the verdict; it remains CONDITIONAL pending correction and transparent reproduction of the bibliometric queries.","tokens_in":9992,"tokens_out":11378,"duration_ms":100714,"concrete_test":"Reconstruct the exact OpenAlex queries (full-text term, domain filter, 2015-2024), export the 2024 inclusive and control result lists, compute the intersection and ratio, and verify whether any reported subset (e.g., 68/676) yields the '10%' figure; if not, correct Table 3 and the abstract. This single reproduction check settles whether the headline statistic is a reporting error or an unsupported claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.2 states that inclusivity-focused research 'accounted for just 10% of all speech technology papers in 2024.' Table 3 reports the inclusive full-text search as 479 papers and the control as 676 papers for 2024, which is 70.9%, not 10%. The same paragraph describes a 'narrow down' to 120 papers in 2023 and 97 in 2024, but no column of Table 3 contains those numbers, and the gender-specific column shows 0 gender papers in 2024, an implausible result for a broad full-text query that includes the word 'gender.' Thus the central numeric anchor of the abstract cannot be derived from the reported data; it may reflect a typo, a column misalignment, or an unreported query variant. The paper's own Limitations 3.3 mentions keyword and database coverage but not this internal inconsistency. This does not overturn the qualitative direction of the argument, but the specific '10%' figure is load-bearing for the claim that inclusivity is a niche area, and as reported it is unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper surveys inclusivity in AI speech technologies for healthcare over 2015-2024, covering both speech datasets and research literature. For datasets, the author searched PapersWithCode and evaluated each dataset on four dimensions: language coverage, accent diversity, demographic representation (gender, age, ethnicity), and inclusion of speech impairments. For research trends, the author used OpenAlex full-text queries with a Health Sciences domain filter, comparing a general 'AI speech technology healthcare' control search with inclusivity-focused searches and with title/abstract keyword filters for gender, age, speech impairment, and ethnicity. The paper concludes that datasets and research disproportionately favor high-resource languages, standardized accents, and narrow demographic groups, and that no surveyed dataset explicitly includes impaired speech samples.","tokens_in":10153,"tokens_out":2508,"duration_ms":23809,"significance":"If the quantitative results were reliable, this paper would provide a useful, compact evidence base for a widely discussed problem: speech technology for healthcare risks encoding demographic and linguistic biases. The paper's strengths are its explicit four-dimensional inclusivity coding scheme for datasets, its transparent description of both search strategies, and its candid limitations sections. It also makes falsifiable claims (e.g., 'none of the datasets explicitly include impaired speech samples') that are directly checkable against Table 1. However, several load-bearing numbers in the manuscript are internally inconsistent with the tables, and the bibliometric analysis is not validated for precision or recall. The qualitative direction of the argument is plausible and broadly supported by the dataset inventory, but the quantitative anchors need correction and re-verification before the paper can be considered reliable.","major_comments":[{"comment":"The claim that inclusivity-focused research 'accounted for just 10% of all speech technology papers in 2024' is directly contradicted by the paper's own Table 3, which reports 479 inclusive papers and 676 control papers for 2024, i.e., 70.9%. No column in Table 3 contains the value 10%, and the same paragraph's 'narrow down' figures of 120 papers in 2023 and 97 in 2024 do not appear anywhere in the table. The '10%' figure is the central quantitative anchor of the inclusivity-is-niche claim, so the authors must either correct the calculation, clarify the exact query variant that produced it, or remove it and rephrase the claim to match what the data actually show.","section":"Section 3.2, Table 3"},{"comment":"The text in Section 2.2 states that the search 'yielded 38 datasets,' but Table 1 lists 41 rows. Additionally, the text says CoVoST2 covers 21 languages while Table 1 says 16, and VoxPopuli is reported as 23 languages in Table 1 but as 22 in the text. These inconsistencies mean the dataset inventory, which is the qualitative core of the paper, cannot currently be used as a reliable reference without re-counting. Please reconcile the row count and the language counts, and state explicitly whether Table 1 is exhaustive or a sample.","section":"Section 2.2, Table 1"},{"comment":"The gender search column reports 0 papers in 2024, despite the stated title/abstract keywords including 'gender,' 'women,' 'men,' 'transgender,' 'non-binary,' and 'gender-diverse.' A broad full-text query that includes the word 'gender' returning zero gender-tagged papers in 2024 is implausible and suggests either a column misalignment, a different query variant, or an error in the table. Since the paper uses this result to claim that gender-focused research is 'strikingly limited,' the authors need to provide the exact query used for each column and verify that the counts correspond to the correct search criteria.","section":"Section 3.2, Table 3 gender column"},{"comment":"The bibliometric conclusions rest on OpenAlex full-text searches such as 'inclusive AI speech technology healthcare bias,' but no precision or recall validation is reported, and the searches are likely to retrieve papers that merely mention the query words rather than papers actually focused on inclusivity. The authors should report the exact query strings, the date of retrieval, and ideally a manual precision check on a random sample of hits, or at minimum replace exact counts and percentages with ranges or directional claims that are robust to this uncertainty.","section":"Section 3.1, Section 3.3, Table 3"}],"minor_comments":[{"comment":"There is a typo in the phrase 'such as Najavo or Maori'; it should be 'Navajo.'","section":"Section 2.2"},{"comment":"The word 'inclusiivity' should be 'inclusivity.'","section":"Section 3.1"},{"comment":"The phrase 'future researcher should expand datasets' should be 'future researchers should expand datasets.'","section":"Section 4"},{"comment":"The table headers are visually run together (e.g., 'Language(s) Accent(s) Demographic...'), which makes the table hard to parse; adding explicit vertical separators or an improved layout would help.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper's scope and contribution are modest for a full research article, but the topic is timely and the qualitative dataset review is useful. The main concern is that the quantitative claims in Section 3.2 are internally inconsistent with Table 3, and the OpenAlex search counts are unvalidated; these are fixable with a careful revision, so I do not recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful but rough audit. The dataset inventory (Table 1) is a decent quick reference, and the qualitative conclusion — speech AI resources for healthcare skew English, standard accent, narrow demographics — is well supported and consistent with prior literature. But the paper's own numbers do not hold together. The abstract's '10%' claim is contradicted by Table 3, where 479/676 is about 71%; the stress-test note is right. The 'narrow down' figures of 120 and 97 appear in that same table but are discussed as if from a different search, and the table header is mangled in ways that make the column mapping guesswork. Add to that: 38 datasets in the text vs 41 rows in the table, CoVoST2 as 21 vs 16 languages, VoxPopuli as 22 vs 23. These are not fatal to the qualitative thesis, but they are exactly the kind of errors that need a referee.\n\nWhat's good: the paper is clearly written, honest about search limitations, and cites the standard bias-in-ASR literature. The compiled metadata table is a reasonable starting inventory for someone new to the area. The OpenAlex trend counts are at least directionally informative.\n\nWeak spots beyond the arithmetic: exact queries are not given, so the bibliometric counts are not reproducible; no precision/recall check on full-text searches; and the 'no impaired speech data' claim is scoped to PapersWithCode but the abstract states it more broadly, ignoring that known dysarthric speech datasets (TORGO, etc.) exist outside that platform. The paper would be stronger if the abstract and conclusion were rephrased to match the actual scope.\n\nWho should read it: dataset designers or policy folks wanting a checklist of inclusivity gaps; not someone looking for methods. I'd send it to a serious referee — the flaws are fixable and the inventory is worth preserving — but I would not cite it without checking the numbers against primary sources.","headline":"A clearly written audit with a useful dataset inventory, but its headline 10% figure is contradicted by its own Table 3 and several internal numbers don't reconcile.","tokens_in":10708,"tokens_out":6319,"would_cite":false,"duration_ms":54557,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"None of the 38 AI speech datasets examined in this decade-long review includes impaired speech.","keywords":["AI speech recognition","inclusive AI","healthcare","speech datasets","speech impairment","bias","health equity","text-to-speech"],"falsifier":"If a re-run of the dataset search finds even one public speech dataset from 2015 to 2024 with explicitly labeled impaired-speech samples, the paper's 'none' claim is false; and if hand-labeling a random sample of the inclusivity-search hits shows most are not about inclusive speech technology, the 10 percent share needs revising.","tokens_in":9742,"feed_emoji":"🗣️","tokens_out":9174,"duration_ms":78196,"temperature":0.7,"pith_summary":"This paper argues that after a decade of growth, AI speech technology for healthcare is not inclusive enough for the populations that need it. Reviewing 38 public speech datasets and bibliometric records of health-sciences publications from 2015 to 2024, it finds that datasets skew toward English and other high-resource languages, standard accents, and narrow demographic categories, and that none of the datasets explicitly includes speech from people with speech impairments. It also finds that research explicitly on inclusivity or bias was only about 10 percent of speech-technology papers in 2024. The stakes are clinical: speech recognition trained on such data risks mishearing older adults, non-native speakers, and people with speech disorders, which can corrupt clinical documentation or lead to misdiagnosis.","feed_headline":"None of 38 AI speech datasets includes impaired speech","feed_subtitle":"A decade of speech corpora skews toward English and standard accents; healthcare AI pays the price.","key_machinery":"The argument runs on two measurement instruments. First, a four-metric inclusivity audit -- language coverage, accent diversity, demographic representation (gender, age, ethnicity), and speech-impairment inclusion -- applied to each of the 38 datasets found through a paper-dataset indexing platform. Second, a bibliometric comparison of a broad full-text search for AI speech technology in healthcare against a narrower full-text search for inclusive speech AI with bias, both filtered to health sciences and to title/abstract keywords for gender, age, speech impairment, and ethnicity. Outliers such as the EdAcc corpus (nine English accents plus gender, age, and ethnicity metadata) and CoVoST2 (multilingual, with three genders and eight age groups) serve as existence proofs that more inclusive design is feasible.","core_discovery":"The central claim is that the data foundation of healthcare speech AI is systematically narrow. Across 38 open speech-recognition and speech-synthesis datasets published between 2015 and 2024, English dominates, non-standard accents are rarely labeled, gender is often unspecified, non-binary genders appear in only one dataset, age diversity is sparse, and no dataset explicitly includes impaired speech samples. On the research side, a control search for speech technology in healthcare grew from 21 papers in 2015 to 676 in 2024, while the inclusivity-focused search reached 479 papers in 2024, about 10 percent of the field. The paper reads this gap as a risk to healthcare equity, since biased training data can make AI misinterpret speech from marginalized groups.","pith_inferences":["The paper does not validate the search precision, so the 10 percent figure is likely an upper bound on true inclusivity-focused research; a manual labeling study of sampled hits would settle it.","Because the dataset audit relies on published metadata, the reported absence of impaired speech and demographic detail is a documented-absence claim; some corpora may contain such speakers without saying so.","A direct extension of the paper's argument would benchmark commercial and open speech recognizers on held-out dysarthric, elderly, child, and non-native speech, converting the dataset gaps into measurable word-error-rate disparities.","Regulatory or procurement criteria for healthcare speech AI could reasonably require per-language, per-accent, and per-disability accuracy reporting, which the paper's equity framing implicitly calls for."],"forward_implications":["If these gaps persist, clinical speech systems will keep performing worst for patients whose voices are least represented, including older adults, non-native speakers, children, and people with dysarthria or aphasia.","The absence of impaired-speech samples means models cannot be expected to recognize disordered speech; dedicated datasets for dysarthria, aphasia, and similar conditions are a prerequisite, not an option.","Because research on gender, age, speech impairment, and ethnicity is a small fraction of the field, with near-zero counts in several years, funding and evaluation criteria that reward inclusive dataset design would change the trajectory.","The positive outliers show that intentional dataset design can include multiple accents, ages, and genders, so the current narrowness is a choice rather than a technical limit."],"supporting_citations":[{"why":"Anchors the English-dominant baseline; it is the most-used speech-recognition corpus in the reviewed set.","marker":"[31]"},{"why":"Supplies the paper's main positive example of accent, gender, age, and ethnicity metadata in one corpus.","marker":"[35]"},{"why":"Provides the outlier case with three gender categories and eight age groups across many languages.","marker":"[42]"},{"why":"Documents Mandarin coverage with regional accents, the paper's main non-English high-resource example.","marker":"[4]"},{"why":"Provides 19 Indonesian languages, supporting the claim that Southeast Asian multilingual resources are emerging but limited.","marker":"[5]"},{"why":"Supports the observation that European languages dominate multilingual coverage.","marker":"[41]"},{"why":"Shows a medical-domain Vietnamese speech dataset, evidence that medical speech data exist but remain narrow.","marker":"[23]"},{"why":"Grounds the clinical risk claim by documenting racial disparities in commercial speech recognition.","marker":"[19]"}],"fun_headline_variants":["Decade of AI speech data excludes impaired speech entirely","Healthcare AI speech datasets skew English, miss diversity","38 datasets, no impaired speech: AI healthcare gap exposed","AI speech in healthcare lacks inclusivity, decade review shows"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The bibliometric counts assume full-text keyword searches in a scholarly database accurately measure how much research actually focuses on inclusive AI speech; no precision or recall check is reported.","fun_headline_variants_meta":{"raw":{"variants":["Decade of AI speech data excludes impaired speech entirely","Healthcare AI speech datasets skew English, miss diversity","38 datasets, no impaired speech: AI healthcare gap exposed","AI speech in healthcare lacks inclusivity, decade review shows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000397,"raw_usage":{"total_tokens":1986,"prompt_tokens":761,"completion_tokens":1225,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":377,"completion_tokens_details":{"reasoning_tokens":1162}},"tokens_in":377,"tokens_out":1225,"duration_ms":9873,"temperature":1.0,"reasoning_tokens":1162,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:15:39.893839+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If a re-run of the dataset search finds even one public speech dataset from 2015 to 2024 with explicitly labeled impaired-speech samples, the paper's 'none' claim is false; and if hand-labeling a random sample of the inclusivity-search hits shows most are not about inclusive speech technology, the 10 percent share needs revising.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Anchors the English-dominant baseline; it is the most-used speech-recognition corpus in the reviewed set."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the paper's main positive example of accent, gender, age, and ethnicity metadata in one corpus."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents Mandarin coverage with regional accents, the paper's main non-English high-resource example."},{"cited_title":"VietMed: A Dataset and Benchmark for Automatic Speech Recognition of Vietnamese in the Medical Domain","cited_arxiv_id":"2404.05659","evidence_quote":"Shows a medical-domain Vietnamese speech dataset, evidence that medical speech data exist but remain narrow."}],"review_version":1}