Pith. sign in

REVIEW 4 major objections 4 minor 52 references

Inclusivity of AI Speech in Healthcare: A Decade Look Back

T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read None of the 38 AI speech datasets examined in this decade-long review includes impaired speech.

desk verdict A clearly written audit with a useful dataset inventory, but its headline 10% figure is contradicted by its own Table 3 and several internal numbers don't reconcile. read the letter →

arxiv 2505.10596 v1 pith:36GCQUWQ submitted 2025-05-15 cs.CY cs.AI

classification cs.CYcs.AI
keywords AIspeechrecognitioninclusivehealthcaredatasetsimpairmentbiashealthequitytext-to-speech
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that after a decade of growth, AI speech technology for healthcare is not inclusive enough for the populations that need it. Reviewing 38 public speech datasets and bibliometric records of health-sciences publications from 2015 to 2024, it finds that datasets skew toward English and other high-resource languages, standard accents, and narrow demographic categories, and that none of the datasets explicitly includes speech from people with speech impairments. It also finds that research explicitly on inclusivity or bias was only about 10 percent of speech-technology papers in 2024. The stakes are clinical: speech recognition trained on such data risks mishearing older adults, non-native speakers, and people with speech disorders, which can corrupt clinical documentation or lead to misdiagnosis.

What carries the argument

The argument runs on two measurement instruments. First, a four-metric inclusivity audit -- language coverage, accent diversity, demographic representation (gender, age, ethnicity), and speech-impairment inclusion -- applied to each of the 38 datasets found through a paper-dataset indexing platform. Second, a bibliometric comparison of a broad full-text search for AI speech technology in healthcare against a narrower full-text search for inclusive speech AI with bias, both filtered to health sciences and to title/abstract keywords for gender, age, speech impairment, and ethnicity. Outliers such as the EdAcc corpus (nine English accents plus gender, age, and ethnicity metadata) and CoVoST2 (multilingual, with three genders and eight age groups) serve as existence proofs that more inclusive design is feasible.

What would settle it

If a re-run of the dataset search finds even one public speech dataset from 2015 to 2024 with explicitly labeled impaired-speech samples, the paper's 'none' claim is false; and if hand-labeling a random sample of the inclusivity-search hits shows most are not about inclusive speech technology, the 10 percent share needs revising.

Watch

Extended reading notes

Core claim

The central claim is that the data foundation of healthcare speech AI is systematically narrow. Across 38 open speech-recognition and speech-synthesis datasets published between 2015 and 2024, English dominates, non-standard accents are rarely labeled, gender is often unspecified, non-binary genders appear in only one dataset, age diversity is sparse, and no dataset explicitly includes impaired speech samples. On the research side, a control search for speech technology in healthcare grew from 21 papers in 2015 to 676 in 2024, while the inclusivity-focused search reached 479 papers in 2024, about 10 percent of the field. The paper reads this gap as a risk to healthcare equity, since biased training data can make AI misinterpret speech from marginalized groups.

Load-bearing premise

The bibliometric counts assume full-text keyword searches in a scholarly database accurately measure how much research actually focuses on inclusive AI speech; no precision or recall check is reported.

Editorial extensions

If this is right

  • If these gaps persist, clinical speech systems will keep performing worst for patients whose voices are least represented, including older adults, non-native speakers, children, and people with dysarthria or aphasia.
  • The absence of impaired-speech samples means models cannot be expected to recognize disordered speech; dedicated datasets for dysarthria, aphasia, and similar conditions are a prerequisite, not an option.
  • Because research on gender, age, speech impairment, and ethnicity is a small fraction of the field, with near-zero counts in several years, funding and evaluation criteria that reward inclusive dataset design would change the trajectory.
  • The positive outliers show that intentional dataset design can include multiple accents, ages, and genders, so the current narrowness is a choice rather than a technical limit.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not validate the search precision, so the 10 percent figure is likely an upper bound on true inclusivity-focused research; a manual labeling study of sampled hits would settle it.
  • Because the dataset audit relies on published metadata, the reported absence of impaired speech and demographic detail is a documented-absence claim; some corpora may contain such speakers without saying so.
  • A direct extension of the paper's argument would benchmark commercial and open speech recognizers on held-out dysarthric, elderly, child, and non-native speech, converting the dataset gaps into measurable word-error-rate disparities.
  • Regulatory or procurement criteria for healthcare speech AI could reasonably require per-language, per-accent, and per-disability accuracy reporting, which the paper's equity framing implicitly calls for.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper surveys inclusivity in AI speech technologies for healthcare over 2015-2024, covering both speech datasets and research literature. For datasets, the author searched PapersWithCode and evaluated each dataset on four dimensions: language coverage, accent diversity, demographic representation (gender, age, ethnicity), and inclusion of speech impairments. For research trends, the author used OpenAlex full-text queries with a Health Sciences domain filter, comparing a general 'AI speech technology healthcare' control search with inclusivity-focused searches and with title/abstract keyword filters for gender, age, speech impairment, and ethnicity. The paper concludes that datasets and research disproportionately favor high-resource languages, standardized accents, and narrow demographic groups, and that no surveyed dataset explicitly includes impaired speech samples.

Significance. If the quantitative results were reliable, this paper would provide a useful, compact evidence base for a widely discussed problem: speech technology for healthcare risks encoding demographic and linguistic biases. The paper's strengths are its explicit four-dimensional inclusivity coding scheme for datasets, its transparent description of both search strategies, and its candid limitations sections. It also makes falsifiable claims (e.g., 'none of the datasets explicitly include impaired speech samples') that are directly checkable against Table 1. However, several load-bearing numbers in the manuscript are internally inconsistent with the tables, and the bibliometric analysis is not validated for precision or recall. The qualitative direction of the argument is plausible and broadly supported by the dataset inventory, but the quantitative anchors need correction and re-verification before the paper can be considered reliable.

major comments (4)
  1. [Section 3.2, Table 3] The claim that inclusivity-focused research 'accounted for just 10% of all speech technology papers in 2024' is directly contradicted by the paper's own Table 3, which reports 479 inclusive papers and 676 control papers for 2024, i.e., 70.9%. No column in Table 3 contains the value 10%, and the same paragraph's 'narrow down' figures of 120 papers in 2023 and 97 in 2024 do not appear anywhere in the table. The '10%' figure is the central quantitative anchor of the inclusivity-is-niche claim, so the authors must either correct the calculation, clarify the exact query variant that produced it, or remove it and rephrase the claim to match what the data actually show.
  2. [Section 2.2, Table 1] The text in Section 2.2 states that the search 'yielded 38 datasets,' but Table 1 lists 41 rows. Additionally, the text says CoVoST2 covers 21 languages while Table 1 says 16, and VoxPopuli is reported as 23 languages in Table 1 but as 22 in the text. These inconsistencies mean the dataset inventory, which is the qualitative core of the paper, cannot currently be used as a reliable reference without re-counting. Please reconcile the row count and the language counts, and state explicitly whether Table 1 is exhaustive or a sample.
  3. [Section 3.2, Table 3 gender column] The gender search column reports 0 papers in 2024, despite the stated title/abstract keywords including 'gender,' 'women,' 'men,' 'transgender,' 'non-binary,' and 'gender-diverse.' A broad full-text query that includes the word 'gender' returning zero gender-tagged papers in 2024 is implausible and suggests either a column misalignment, a different query variant, or an error in the table. Since the paper uses this result to claim that gender-focused research is 'strikingly limited,' the authors need to provide the exact query used for each column and verify that the counts correspond to the correct search criteria.
  4. [Section 3.1, Section 3.3, Table 3] The bibliometric conclusions rest on OpenAlex full-text searches such as 'inclusive AI speech technology healthcare bias,' but no precision or recall validation is reported, and the searches are likely to retrieve papers that merely mention the query words rather than papers actually focused on inclusivity. The authors should report the exact query strings, the date of retrieval, and ideally a manual precision check on a random sample of hits, or at minimum replace exact counts and percentages with ranges or directional claims that are robust to this uncertainty.
minor comments (4)
  1. [Section 2.2] There is a typo in the phrase 'such as Najavo or Maori'; it should be 'Navajo.'
  2. [Section 3.1] The word 'inclusiivity' should be 'inclusivity.'
  3. [Section 4] The phrase 'future researcher should expand datasets' should be 'future researchers should expand datasets.'
  4. [Table 1] The table headers are visually run together (e.g., 'Language(s) Accent(s) Demographic...'), which makes the table hard to parse; adding explicit vertical separators or an improved layout would help.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper's claims are descriptive summaries of dataset annotations and OpenAlex search counts; the internal inconsistency in the 10% figure is a soundness issue, not an input-output loop.

full rationale

The paper makes no fitted-parameter-then-prediction move and contains no derivation chain that reduces to its own inputs. The dataset claims (e.g., 'none of the datasets we found explicitly include impaired speech samples') are direct annotations of the 38 datasets listed in Table 1, coded against criteria stated in Section 2.1. The bibliometric claims are summaries of raw OpenAlex full-text searches reported in Table 3; the conclusion that inclusivity research is a 'small fraction' or 'niche area' is a descriptive reading of those counts, not a quantity derived from a fitted parameter. The paper's own Limitations sections (2.3 and 3.3) explicitly acknowledge the dependence on PapersWithCode coverage, OpenAlex indexing, and keyword choice, which are validity threats rather than circular steps. The skeptic's observation that the abstract's '10%' figure conflicts with Table 3 (479 inclusive vs 676 control is about 71%) is an internal-consistency and evidentiary problem, not circularity: the claim is unsupported or misreported, but it does not follow from the inputs by construction. There are no load-bearing self-citations, no imported uniqueness theorems, and no ansatz smuggled in via citation. The analysis is self-contained as a descriptive survey, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The audit makes no fitted-parameter or invented-entity assumptions. Its conclusions rest on platform representativeness, on the validity of full-text search counts, and on treating missing metadata as evidence of absence.

assumptions (4)
  • domain assumption ASR and TTS datasets are the most relevant speech data for healthcare contexts.
    Section 2.1 states this assumption to justify searching only ASR and TTS datasets on PapersWithCode.
  • domain assumption PapersWithCode and OpenAlex provide sufficient coverage to characterize the field.
    The entire audit is bounded by these two platforms; the author acknowledges this limitation in Sections 2.3 and 3.3.
  • domain assumption Full-text search counts in OpenAlex accurately measure research volume and focus.
    Section 3.1 uses full-text terms such as "AI speech technology healthcare" as counts without validating precision or recall.
  • domain assumption Datasets without stated demographic or accent metadata are treated as not inclusive.
    Section 2.2 repeatedly equates lack of specification with lack of representation, for example when stating that LibriSpeech does not specify age or gender distribution and therefore scores No on those metrics.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Inclusivity of AI Speech in Healthcare: A Decade Look Back." pith.science (2026). https://pith.science/paper/36GCQUWQ

@misc{pith2026250510596,
  author       = {Pith},
  title        = {Pith review of: Inclusivity of AI Speech in Healthcare: A Decade Look Back},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/36GCQUWQ}},
  note         = {Machine review of arXiv:2505.10596}
}
read the original abstract

The integration of AI speech recognition technologies into healthcare has the potential to revolutionize clinical workflows and patient-provider communication. However, this study reveals significant gaps in inclusivity, with datasets and research disproportionately favouring high-resource languages, standardized accents, and narrow demographic groups. These biases risk perpetuating healthcare disparities, as AI systems may misinterpret speech from marginalized groups. This paper highlights the urgent need for inclusive dataset design, bias mitigation research, and policy frameworks to ensure equitable access to AI speech technologies in healthcare.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

52 extracted references · 30 canonical work pages

  1. [1]

    Sourav Banerjee, Ayushi Agarwal, and Promila Ghosh. 2024. High-precision medical speech recognition through synthetic data and semantic correction: UNITED-MEDASR. arXiv preprint arXiv:2412.00055 (2024)

  2. [2]

    Benjamin Beilharz, Xin Sun, Sariya Karimova, and Stefan Riezler. 2019. LibriVoxDeEn: A corpus for German-to-English speech translation and German speech recognition. arXiv preprint arXiv:1910.07924 (2019)

  3. [3]

    Marcely Zanon Boito, William N Havard, Mahault Garnerin, Éric Le Ferrand, and Laurent Besacier. 2019. Mass: A large and clean multilingual corpus of sentence-aligned spoken utterances extracted from the bible. arXiv preprint arXiv:1907.12895 (2019)

  4. [4]

    Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng. 2017. Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline. In 2017 20th conference of the oriental chapter of the international coordinating committee on speech databases and speech I/O systems and assessment (O-COCOSDA) . IEEE, 1–5

  5. [5]

    Samuel Cahyawijaya, Holy Lovenia, Alham Fikri Aji, Genta Indra Winata, Bryan Wilie, Rahmad Mahendra, Christian Wibisono, Ade Romadhony, Karissa Vincentio, Fajri Koto, et al. 2022. NusaCrowd: Open source initiative for Indonesian NLP resources. arXiv preprint arXiv:2212.09648 (2022)

  6. [6]

    Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du, Wei-Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, et al. 2021. Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio. arXiv preprint arXiv:2106.06909 (2021)

  7. [7]

    Chenye Cui, Yi Ren, Jinglin Liu, Feiyang Chen, Rongjie Huang, Ming Lei, and Zhou Zhao. 2021. Emovie: A mandarin emotion speech dataset with a simple emotional text-to-speech model. arXiv preprint arXiv:2106.09317 (2021)

  8. [8]

    Jacob Donley, Vladimir Tourbabin, Jung-Suk Lee, Mark Broyles, Hao Jiang, Jie Shen, Maja Pantic, Vamsi Krishna Ithapu, and Ravish Mehra. 2021. Easycom: An augmented reality dataset to support algorithms for easy communication in noisy environments. arXiv preprint arXiv:2107.04174 (2021)

Show all 52 references
  1. [9]

    Jiayu Du, Xingyu Na, Xuechen Liu, and Hui Bu. 2018. Aishell-2: Transforming mandarin asr research into industrial scale. arXiv preprint arXiv:1808.10583 (2018). , Vol. 1, No. 1, Article . Publication date: June 2025. 8 • Larasati et al

  2. [10]

    Yihui Fu, Luyao Cheng, Shubo Lv, Yukai Jv, Yuxiang Kong, Zhuo Chen, Yanxin Hu, Lei Xie, Jian Wu, Hui Bu, et al. 2021. Aishell-4: An open source dataset for speech enhancement, separation, recognition and speaker diarization in conference scenario. arXiv preprint arXiv:2104.036...

  3. [11]

    Gonçal Garcés Díaz-Munío, Joan Albert Silvestre Cerdà, Javier Jorge-Cano, Adrián Giménez Pastor, Javier Iranzo-Sánchez, Pau Baquero- Arnal, Nahuel Roselló, Alejandro Manuel Pérez-González de Martos, Jorge Civera Saiz, José Alberto Sanchis Navarro, et al . 2021. Europarl-ASR: A...

  4. [12]

    Vikram Gupta, Rini Sharon, Ramit Sawhney, and Debdoot Mukherjee. 2022. Adima: Abuse detection in multilingual audio. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 6172–6176

  5. [13]

    Jung-Woo Ha, Kihyun Nam, Jingu Kang, Sang-Woo Lee, Sohee Yang, Hyunhoon Jung, Eunmi Kim, Hyeji Kim, Soojin Kim, Hyun Ah Kim, et al. 2020. Clovacall: Korean goal-oriented dialog speech corpus for automatic speech recognition of contact centers. arXiv preprint arXiv:2004.09367 (2020)

  6. [14]

    Javier Iranzo-Sánchez, Joan Albert Silvestre-Cerda, Javier Jorge, Nahuel Roselló, Adria Giménez, Albert Sanchis, Jorge Civera, and Alfons Juan. 2020. Europarl-st: A multilingual corpus for speech translation of parliamentary debates. In ICASSP 2020-2020 IEEE International Conf...

  7. [15]

    Andrzej Izworski, Ryszard Tadeusiewicz, and Wieslaw Wszolek. 2004. Artificial intelligence methods in diagnostics of the pathological speech signals. In Knowledge-Based Intelligent Information and Engineering Systems: 8th International Conference, KES 2004, Wellington, New Zea...

  8. [16]

    Ye Jia, Michelle Tadmor Ramanovich, Quan Wang, and Heiga Zen. 2022. CVSS Corpus and Massively Multilingual Speech-to-Speech Translation. In Proceedings of the Thirteenth Language Resources and Evaluation Conference . 6691–6703

  9. [17]

    Maree Johnson, Samuel Lapkin, Vanessa Long, Paula Sanchez, Hanna Suominen, Jim Basilakis, and Linda Dawson. 2014. A systematic review of speech recognition technology in health care. BMC medical informatics and decision making 14 (2014), 1–14

  10. [18]

    Willemijn Klaassen, Bram van Dijk, and Marco Spruit. 2024. A Review of Challenges in Speech-based Conversational AI for Elderly Care. arXiv preprint arXiv:2412.07388 (2024)

  11. [19]

    Allison Koenecke, Andrew Nam, Emily Lake, Joe Nudell, Minnie Quartey, Zion Mengesha, Connor Toups, John R Rickford, Dan Jurafsky, and Sharad Goel. 2020. Racial disparities in automated speech recognition. Proceedings of the national academy of sciences 117, 14 (2020), 7684–7689

  12. [20]

    Rostislav Kolobov, Olga Okhapkina, Olga Omelchishina, Andrey Platunov, Roman Bedyakin, Vyacheslav Moshkin, Dmitry Menshikov, and Nikolay Mikhaylovskiy. 2021. Mediaspeech: Multilanguage asr benchmark and dataset. arXiv preprint arXiv:2103.16193 (2021)

  13. [21]

    Ginny Kwong and Tom Stafford. 2020. Empowering physicians with a digital workflow and AI-based clinical documentation programmes at Halifax Health. Management in Healthcare 4, 4 (2020), 294–299

  14. [22]

    Siddique Latif, Junaid Qadir, Adnan Qayyum, Muhammad Usama, and Shahzad Younis. 2020. Speech technology for healthcare: Opportunities, challenges, and state of the art. IEEE Reviews in Biomedical Engineering 14 (2020), 342–356

  15. [23]

    Khai Le-Duc. 2024. Vietmed: A dataset and benchmark for automatic speech recognition of vietnamese in the medical domain. arXiv preprint arXiv:2404.05659 (2024)

  16. [24]

    Georgia Maniati, Alexandra Vioni, Nikolaos Ellinas, Karolos Nikitaras, Konstantinos Klapsas, June Sig Sung, Gunu Jho, Aimilios Chalamandaris, and Pirros Tsiakoulis. 2022. SOMOS: The samsung open mos dataset for the evaluation of neural text-to-speech synthesis. arXiv preprint ...

  17. [25]

    Sonia Marrone. 2007. Understanding barriers to health care: a review of disparities in health care services among indigenous populations. International Journal of Circumpolar Health 66, 3 (2007), 188–198

  18. [26]

    Joshua L Martin and Kelly Elizabeth Wright. 2023. Bias in automatic speech recognition: The case of African American language. Applied Linguistics 44, 4 (2023), 613–630

  19. [27]

    Quentin Meeus, Marie-Francine Moens, and Hugo Van hamme. 2024. MSNER: A Multilingual Speech Dataset for Named Entity Recognition. In 20th Joint ACL-ISO Workshop on Interoperable Semantic Annotation at LREC-COLING

  20. [28]

    Saida Mussakhojayeva, Aigerim Janaliyeva, Almas Mirzakhmetov, Yerbolat Khassanov, and Huseyin Atakan Varol. 2021. KazakhTTS: An open-source Kazakh text-to-speech synthesis dataset. arXiv preprint arXiv:2104.08459 (2021)

  21. [29]

    Patrick K O’Neill, Vitaly Lavrukhin, Somshubra Majumdar, Vahid Noroozi, Yuekai Zhang, Oleksii Kuchaiev, Jagadeesh Balam, Yuliya Dovzhenko, Keenan Freyberg, Michael D Shulman, et al. 2021. Spgispeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end spe...

  22. [30]

    Vasile Păiş, Radu Ion, Andrei-Marius Avram, Elena Irimia, Verginica Barbu Mititelu, and Maria Mitrofan. 2021. Human-machine interaction speech corpus from the robin project. In 2021 International Conference on Speech Technology and Human-Computer Dialogue (SpeD). IEEE, 91–96

  23. [31]

    Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 5206–5210. , Vol. 1, No. 1, Article . Publ...

  24. [32]

    Nikita Pavlichenko, Ivan Stelmakh, and Dmitry Ustalov. 2021. Crowdspeech and voxdiy: Benchmark datasets for crowdsourced audio transcription. arXiv preprint arXiv:2107.01091 (2021)

  25. [33]

    Akam Qader and Hossein Hassani. 2019. Kurdish (sorani) speech to text: Presenting an experimental dataset. arXiv preprint arXiv:1911.13087 (2019)

  26. [34]

    Colleen Richey, Maria A Barrios, Zeb Armstrong, Chris Bartels, Horacio Franco, Martin Graciarena, Aaron Lawson, Mahesh Kumar Nandwana, Allen Stauffer, Julien van Hout, et al . 2018. Voices obscured in complex environmental settings (voices) corpus. arXiv preprint arXiv:1804.05...

  27. [35]

    Ramon Sanabria, Nikolay Bogoychev, Nina Markl, Andrea Carmantini, Ondrej Klejch, and Peter Bell. 2023. The edinburgh international accents of english corpus: Towards the democratization of english asr. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and ...

  28. [36]

    Yao Shi, Hui Bu, Xin Xu, Shaoji Zhang, and Ming Li. 2020. Aishell-3: A multi-speaker mandarin tts corpus and the baselines. arXiv preprint arXiv:2010.11567 (2020)

  29. [37]

    Per Erik Solberg and Pablo Ortiz. 2022. The Norwegian parliamentary speech corpus. arXiv preprint arXiv:2201.10881 (2022)

  30. [38]

    Gemma Turon, Dhanshree Arora, and Miquel Duran-Frigola. 2024. Infectious Disease Research Laboratories in Africa Are Not Using AI Yet– Large Language Models May Facilitate Adoption. ACS Infectious Diseases 10, 9 (2024), 3083–3085

  31. [39]

    Binling Wang, Wenxuan Hu, Jing Li, Yiming Zhi, Zheng Li, Qingyang Hong, Lin Li, Dong Wang, Liming Song, and Cheng Yang. 2021. Olr 2021 challenge: Datasets, rules and baselines. In 2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APS...

  32. [40]

    Changhan Wang, Juan Pino, Anne Wu, and Jiatao Gu. 2020. Covost: A diverse multilingual speech-to-text translation corpus. arXiv preprint arXiv:2002.01320 (2020)

  33. [41]

    Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Em- manuel Dupoux. 2021. VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation. arXiv pre...

  34. [42]

    Changhan Wang, Anne Wu, and Juan Pino. 2020. Covost 2 and massively multilingual speech-to-text translation. arXiv preprint arXiv:2007.10310 (2020)

  35. [43]

    Dong Wang and Xuewei Zhang. 2015. Thchs-30: A free chinese speech corpus. arXiv preprint arXiv:1512.01882 (2015)

  36. [44]

    Pete Warden. 2018. Speech commands: A dataset for limited-vocabulary speech recognition. arXiv preprint arXiv:1804.03209 (2018)

  37. [45]

    Zehui Yang, Yifan Chen, Lei Luo, Runyan Yang, Lingxuan Ye, Gaofeng Cheng, Ji Xu, Yaohui Jin, Qingqing Zhang, Pengyuan Zhang, et al

  38. [46]

    Fan Yu, Zhuoyuan Yao, Xiong Wang, Keyu An, Lei Xie, Zhijian Ou, Bo Liu, Xiulin Li, and Guanqiong Miao. 2021. The SLT 2021 children speech recognition challenge: Open datasets, rules and baselines. In 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 1117–1123

  39. [47]

    Rohola Zandie, Mohammad H Mahoor, Julia Madsen, and Eshrat S Emamian. 2021. Ryanspeech: A corpus for conversational text-to- speech synthesis. arXiv preprint arXiv:2106.08468 (2021)

  40. [48]

    Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. 2019. Libritts: A corpus derived from librispeech for text-to-speech. arXiv preprint arXiv:1904.02882 (2019)

  41. [49]

    Binbin Zhang, Hang Lv, Pengcheng Guo, Qijie Shao, Chao Yang, Lei Xie, Xin Xu, Hui Bu, Xiaoyu Chen, Chenchen Zeng, et al. 2022. Wenetspeech: A 10000+ hours multi-domain mandarin corpus for speech recognition. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speec...

  42. [50]

    Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023. SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities. In The 2023 Conference on Empirical Methods in Natural Language Processing

  43. [51]

    Jun Zhang, Jingyue Wu, Yiyi Qiu, Aiguo Song, Weifeng Li, Xin Li, and Yecheng Liu. 2023. Intelligent speech technologies for transcription, disease diagnosis, and medical equipment interactive control in smart hospitals: A review. Computers in Biology and Medicine 153 (2023), 1...

  44. [2022]

    arXiv preprint arXiv:2203.16844 (2022)

    Open source magicdata-ramc: A rich annotated mandarin conversational (ramc) speech dataset. arXiv preprint arXiv:2203.16844 (2022)

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.