REVIEW 4 major objections 7 minor 1 cited by
Voice of a Continent: Mapping Africa's Speech Technology Frontier
T0 review · 4 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A unified benchmark and model family for African speech moves beyond data scarcity.
desk verdict Useful aggregation and benchmark, but completeness and SoTA claims outrun the evidence; needs claim-scoping and artifact release. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is SimbaBench, a unified evaluation suite that aggregates and standardizes 26 public audio corpora into a common 16-kHz mono WAV format with a unified JSON schema, then partitions data into training, development, and test splits (official splits when available, otherwise 90/10). The load-bearing mechanism is multilingual fine-tuning: for ASR, five baseline models (AfriHuBERT, MMS-1B-all, SeamlessM4T-v2, Wav2Vec2-XLS-R, Whisper-v3-large) are fine-tuned with a shared 215-hour training set (5 hours per language, 30 minutes dev per language) plus a CTC layer for the non-Whisper models; for TTS, the MMS-TTS model is extended to seven previously unsupported languages by starting from checkpoints of linguistically related languages; for SLID, AfriHuBERT is fine-tuned on the same 215-hour split. This machinery turns abundant but fragmented public data into a reproducible benchmark and a set of competitive models.
What would settle it
A concrete check would be to compile an independent list of African speech datasets from sources such as national language archives, university repositories, and radio-station corpora, and compare it against the 26 sources in Table 1; finding even one medium-size public corpus (say, over 50 hours) not covered would falsify the claim that SimbaBench unifies all publicly available African speech data.
Extended reading notes
Core claim
The paper claims that the first comprehensive, harmonized benchmark for African speech—SimbaBench—can be built entirely from publicly available sources, and that models adapted to it (Simba-H, Simba-M, Simba-S, Simba-X, Simba-W for ASR; Simba-TTS; Simba-SLID) achieve state-of-the-art results across multiple African languages and three speech tasks. The data analysis reveals severe imbalance: five languages (Kinyarwanda, Hausa, Yoruba, Swahili, Igbo) dominate the corpus, while most of the 61 languages have under an hour of audio. The model evaluations show that dataset quality and domain diversity strongly affect performance, and that closely related Niger-Congo languages benefit from transfer, whereas Afro-Asiatic languages like Amharic can excel even with moderate data. The paper concludes that broad coverage, multilingual adaptation, and language-family-aware training are the keys to making speech technology work for Africa's low-resource languages.
Load-bearing premise
The paper assumes that its curated set of 26 public sources includes every meaningful public African speech corpus, which is the basis for calling SimbaBench comprehensive and the map systematic.
Editorial extensions
If this is right
- If SimbaBench is correct as claimed, a researcher can evaluate any new African speech model on a single consistent harness, making results across papers directly comparable.
- The finding that fine-tuning with just five hours per language yields large gains implies that modest additional data collection could push many low-resource languages into usable ASR range.
- The language-family transfer results suggest that focusing data collection on well-represented families (e.g., Niger-Congo) could automatically improve their under-resourced relatives.
- The dataset-quality and domain-diversity effects imply that benchmark results will vary widely, so future African speech claims should report per-dataset and per-domain breakdowns rather than a single averaged score.
- The benchmark's release on Hugging Face with standardized test configurations should make it the default evaluation protocol for African speech work, assuming the paper's coverage claims hold.
Reading between the lines
- A natural extension the authors did not test: whether the multilingual fine-tuning recipe transfers to other low-resource regions (e.g., South Asian or indigenous American languages) with similarly sparse data, which would suggest the recipe is general rather than Africa-specific.
- The observed mismatch between speaker population and data volume (e.g., Oromo with 45M speakers and only 34 hours) implies that a population-weighted data collection strategy would produce far larger fairness gains than simply adding more hours to already-rich languages.
- The TTS evaluation relies on ASR error rates as a proxy for synthesis intelligibility, which the authors do not validate against human listeners; a human-perception study could settle whether the reported Simba-TTS gains are real.
- The 61-language ceiling is a function of public data availability, so a testable consequence is that a deeper search (including national archives, radio archives, and non-English portals) could double or triple the language count, either strengthening or undermining the claim of comprehensiveness.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper curates 8,605 hours of publicly available African speech audio from 26 sources, organizes it into a unified evaluation suite called SimbaBench covering ASR, TTS, and SLID across 61 languages, and fine-tunes a family of models, Simba, derived from existing multilingual speech models. It also presents a quantitative analysis of resource distribution across languages, language families, and speaker populations. The headline claims are that SimbaBench is a comprehensive, systematic map of all publicly available African speech data and that the Simba models achieve state-of-the-art performance across multiple African languages and speech tasks.
Significance. If the paper's claims were fully supported, the contribution would be valuable: a standardized, reproducible benchmark for African speech tasks, an explicit multilingual training/development split, and a transparent preprocessing pipeline would be useful community resources. The data-curation effort is large and the descriptive analysis of resource imbalances is informative. However, the completeness and state-of-the-art claims are not currently established by the evidence presented; the paper compares only against generic multilingual baselines and its own curated split, and the curation is not shown to be exhaustive. The resource contribution is therefore real, but the significance as argued is overstated.
major comments (4)
- [Abstract and §3.1/Table 1] The abstract and §3.1 claim that SimbaBench unifies 'all publicly available African speech datasets' and systematically maps the continent's speech space, but the curation is not exhaustive by the paper's own evidence. Ogun et al. (2024), '1000 African Voices,' is discussed in §2 as an African speech resource, yet no row for it appears in Table 1 or Appendix A, and no exclusion rationale is given. Similarly, Table 1 uses only 'Common Voice (CV-19)' (Mozilla Foundation, 2023), despite newer Common Voice releases with additional African languages and hours; no reason is provided for freezing at CV-19. The completeness claim is therefore unsupported as stated. The paper should either include these resources or document a well-defined inclusion/exclusion protocol and replace 'all publicly available' with a scoped claim.
- [§6 and Tables 2/F.1] The 'state-of-the-art' claim in the abstract, §1, and §6 is not supported because all baselines are generic multilingual models (Whisper, Seamless, MMS, AfriHuBERT, XLS-R) evaluated in the paper's own zero-shot setting; no comparison is made to previously published task-specific African ASR systems that report WER/CER on the same datasets (e.g., BembaSpeech, Lwazi, NCHLT, ALFFA). Table F.1 shows that Simba-ASR improves over these zero-shot baselines, but that establishes an adaptation gain, not state-of-the-art status relative to prior fine-tuned systems. Please add explicit comparisons to prior published per-dataset results and qualify the claim accordingly.
- [§5.3 and Table 3] TTS quality is reported only as WER/CER of an unspecified 'best available ASR model' (Section 5.3), and Table 3 shows WER values as high as 91.84 for Southern Sotho and 78.31 for Afrikaans. Without stating which ASR model was used and with no human/listening evaluation, the claim that these results are 'reasonable' or state-of-the-art is not established. At minimum, the paper should specify the evaluation ASR model and its language support, and it should temper the TTS claims or add a subjective evaluation component.
- [Table F.2] The SLID results do not support a global state-of-the-art claim. Simba-SLID improves macro-F1 on some low-resource languages (e.g., Edo 6.25 to 80.12, Tonga 31.48 to 56.47), but it dramatically regresses on several high-resource languages (Lingala 96.29 to 15.86, Malagasy 98.55 to 70.12, Swahili VoxLingua 99.96 to 94.29). The statement that 'Simba-SLID shows notable gains on low-resource languages' is therefore only a partial description; the overall average improvement is 1.38 points. Please report the per-language trade-offs honestly and avoid overgeneralizing the result.
minor comments (7)
- [§2] The word 'intenstive' should be 'intensive'.
- [§3.3] The text reads 'Kernel Density Estimate (KDN)'; the usual abbreviation is KDE, and this should be corrected.
- [§5.3] 'We evaluateSimbaBench' is missing a space; it should read 'We evaluate SimbaBench'.
- [Table 1] The row label 'AfriSpeech (Accented-African))' has an extra closing parenthesis; it should be 'AfriSpeech (Accented African)'.
- [Table F.1] The footnote contains the typo 'Westren' and the header contains 'our’sSimba'; both should be corrected to 'Western' and 'our Simba'.
- [§4] The paper says SimbaBench 'will be hosted on the Hugging Face Datasets platform' and that training/dev splits will be released, but no repository URL or availability date is given; please add the link or state that the resource is forthcoming.
- [§5.2 footnote and §6] The number of languages is reported inconsistently: the footnote says 43 African languages are used for ASR finetuning, Table 1 reports 42 ASR languages, and §6 says 56 language-specific test sets representing 46 languages. Please clarify these counts.
Circularity Check
No significant circularity: benchmark construction, held-out evaluation, and model fine-tuning are independent; self-citations are not load-bearing.
full rationale
The paper's central claims are (1) that SimbaBench aggregates and harmonizes publicly available African speech data and (2) that the Simba models achieve strong results on held-out test splits of that data. Neither claim reduces to its own inputs by construction. Simba-ASR and Simba-SLID are trained on a multilingual training split (5 hours per language for ASR) and evaluated on per-dataset test splits, using official splits when available and otherwise a 90%-10% partition as stated in Section 4. This is standard supervised evaluation, not a fitted parameter renamed as a prediction. The SLID data marked 'Ours' (OlongoAfrica, UDHR, VOA) are newly curated raw audio resources; Simba-SLID is fine-tuned on the ASR training split, not on the SLID test sets, so the reported SLID results are not fitted to the evaluation data. The TTS comparison in Table 3 mixes MMS-TTS results on supported languages with Simba-TTS results on previously unsupported languages, which is a presentation weakness, but it is not a circular derivation. The paper contains several self-citations (e.g., Adebara and Abdul-Mageed 2022; Adebara et al. 2025; Elmadany et al. 2024; Toyin et al. 2023), but none is load-bearing for the benchmark's completeness or the models' performance; no uniqueness theorem or ansatz is imported from prior work. The skeptical concern about excluded corpora such as Ogun et al. (2024) or newer Common Voice releases is a completeness and correctness issue with the empirical claim that SimbaBench unifies 'all publicly available' datasets, not a circularity issue, because that claim is falsifiable and not true by definition. Overall, the derivation chain is self-contained and externally evaluated against MMS, Seamless, Whisper, and other baselines.
Assumptions & free parameters
free parameters (3)
- ASR per-language training sample size =
5 hours per language
- TTS per-language training sample size =
12 hours per language
- Per-language development set size =
30 minutes per language
assumptions (3)
- domain assumption Ground-truth transcriptions and language IDs in the 26 public datasets are accurate.
- domain assumption The selected 26 public sources cover all publicly available African speech datasets.
- domain assumption ASR-based WER/CER is a valid proxy for TTS intelligibility, and macro-F1 is the right SLID metric.
Cite this review
Pith. "Pith review of Voice of a Continent: Mapping Africa's Speech Technology Frontier." pith.science (2026). https://pith.science/paper/QDQGV4A7
@misc{pith2026250518436,
author = {Pith},
title = {Pith review of: Voice of a Continent: Mapping Africa's Speech Technology Frontier},
year = {2026},
howpublished = {\url{https://pith.science/paper/QDQGV4A7}},
note = {Machine review of arXiv:2505.18436}
}
read the original abstract
Africa's rich linguistic diversity remains significantly underrepresented in speech technologies, creating barriers to digital inclusion. To alleviate this challenge, we systematically map the continent's speech space of datasets and technologies, leading to a new comprehensive benchmark SimbaBench for downstream African speech tasks. Using SimbaBench, we introduce the Simba family of models, achieving state-of-the-art performance across multiple African languages and speech tasks. Our benchmark analysis reveals critical patterns in resource availability, while our model evaluation demonstrates how dataset quality, domain diversity, and language family relationships influence performance across languages. Our work highlights the need for expanded speech technology resources that better reflect Africa's linguistic diversity and provides a solid foundation for future research and development efforts toward more inclusive speech technologies.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 1 Pith paper
-
AfroScope: A Framework for Studying the Linguistic Landscape of Africa
A new framework combines a 713-language African LID dataset, strong baselines, and a contrastive-embedding hierarchical step that improves macro-F1 by 4.55 on a 29-language confusable subset.
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Ife Adebara and Muhammad Abdul-Mageed. 2022. https://doi.org/10.18653/v1/2022.acl-long.265 Towards afrocentric NLP for A frican languages: Where we are and where we can go . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3814--3841, Dublin, Ireland. Association for Computational Li...
-
[4]
Ife Adebara, AbdelRahim Elmadany, and Muhammad Abdul-Mageed. 2024. Cheetah: Natural language generation for 517 african languages. arXiv preprint arXiv:2401.01053
work page Pith review arXiv 2024
-
[5]
Ife Adebara, AbdelRahim Elmadany, Muhammad Abdul-Mageed, and Alcides Inciarte. 2022 a . https://doi.org/10.18653/v1/2022.emnlp-main.128 A fro LID : A neural language identification tool for A frican languages . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 1958--1981, Abu Dhabi, United Arab Emirates. Asso...
-
[6]
Ife Adebara, AbdelRahim Elmadany, Muhammad Abdul-Mageed, and Alcides Alcoba Inciarte. 2022 b . Serengeti: Massively multilingual language models for africa. arXiv preprint arXiv:2212.10785
work page Pith review arXiv 2022
-
[7]
Ife Adebara, Hawau Olamide Toyin, Nahom Tesfu Ghebremichael, AbdelRahim Elmadany, and Muhammad Abdul-Mageed. 2025. https://arxiv.org/abs/2502.19582 Where are we? evaluating llm performance on african languages . Preprint, arXiv:2502.19582
work page Pith review arXiv 2025
-
[8]
David Ifeoluwa Adelani, Jessica Ojo, Israel Abebe Azime, Jian Yun Zhuang, Jesujoba O Alabi, Xuanli He, Millicent Ochieng, Sara Hooker, Andiswa Bukula, En-Shiun Annie Lee, et al. 2024. Irokobench: A new benchmark for african languages in the age of large language models. arXiv preprint arXiv:2406.03368
arXiv 2024
Show all 56 references
-
[9]
Jesujoba O Alabi, Xuechen Liu, Dietrich Klakow, and Junichi Yamagishi. 2024. Afrihubert: A self-supervised speech representation model for african languages. arXiv preprint arXiv:2409.20201
2024 arXiv
-
[10]
Antonios Anastasopoulos, Angela Fan, Dani Haziza, et al. 2023. Seamlessm4t: Massively multilingual & multimodal machine translation. arXiv preprint arXiv:2308.04760
2023 arXiv
-
[11]
Asamoah Owusu, A
D. Asamoah Owusu, A. Korsah, B. Quartey, S. Nwolley Jnr., D. Sampah, D. Adjepon-Yamoah, and L. Omane Boateng. 2022. Github - ashesi-org/financial-inclusion-speech-dataset: A speech dataset to support financial inclusion created by ashesi university and nokwary technologies wit...
2022
-
[12]
Arun Babu, Atma Tjandra, Kushal Lakhotia, Apoorv Chauhan, Qiantong Wang, Naman Goyal, Vineel Pratap Jain, Vitaliy Liptchinsky, Ahmed El-Kishky, Juan Pino, Abdelrahman Mohamed, and et al. 2021. Xls-r: Self-supervised cross-lingual speech representation learning at scale. In Pro...
2021
-
[13]
Davel, Charl van Heerden, Febe de Wet, and Jaco Badenhorst
Etienne Barnard, Marelie H. Davel, Charl van Heerden, Febe de Wet, and Jaco Badenhorst. 2014. https://www.researchgate.net/publication/301858320_The_nchlt_speech_corpus_of_the_south_african_languages The nchlt speech corpus of the south african languages . In Proceedings of th...
2014
-
[14]
Tadesse Destaw Belay, Israel Abebe Azime, Ibrahim Said Ahmad, Idris Abdulmumin, Abinew Ali Ayele, Shamsuddeen Hassan Muhammad, and Seid Muhie Yimam. 2025. Afroxlmr-social: Adapting pre-trained language models for african languages social media text. arXiv preprint arXiv:2503.18247
2025
-
[15]
Emily M. Bender. 2011. https://doi.org/10.33011/lilt.v6i.1239 On achieving and evaluating language-independence in nlp . Linguistic Issues in Language Technology, 6
2011 doi
-
[16]
Laurent Besacier and Elodie Gauthier. 2023. Alffa\_public: African languages factored lattices for automatic speech recognition. https://github.com/getalp/ALFFA_PUBLIC
2023
-
[17]
Herv \'e Bredin and Antoine Laurent . 2021. End-to-end speaker segmentation for overlap-aware resegmentation . In Proc. Interspeech 2021, Brno, Czech Republic
2021
-
[18]
Herv \'e Bredin , Ruiqing Yin , Juan Manuel Coria , Gregory Gelly , Pavel Korshunov , Marvin Lavechin , Diego Fustes , Hadrien Titeux , Wassim Bouaziz , and Marie-Philippe Gill . 2020. pyannote.audio: neural building blocks for speaker diarization . In ICASSP 2020, IEEE Intern...
2020
-
[19]
Hervé Bredin. 2023. pyannote.audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe . In Proc. INTERSPEECH 2023
2023
-
[20]
Wenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng, Guangyan Zhang, Qichao Wang, Yiwen Guo, and Irwin King. 2024. Recent advances in speech language models: A survey. arXiv preprint arXiv:2410.03751
2024 arXiv
-
[21]
Ewald Van der westhuizen and Thomas Niesler. 2018. A First South African Corpus of Multilingual Code-switched Soap Opera Speech . In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resour...
2018
-
[22]
Digital Umuganda . 2023. Afrispeech kinyarwanda male and female tts datasets. https://huggingface.co/datasets/DigitalUmuganda/afrispeak_kinyarwanda_male_tts_dataset
2023
-
[23]
Moussa Doumbouya, Lisa Einstein, and Chris Piech. 2021. Using radio archives for low-resource speech recognition: Towards an intelligent virtual assistant for illiterate users. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35
2021
-
[24]
AbdelRahim Elmadany, Ife Adebara, and Muhammad Abdul-Mageed. 2024. Toucan: Many-to-many translation for 150 african language pairs. arXiv preprint arXiv:2407.04796
2024 arXiv
- [25]
-
[26]
Alexander Gutkin, I s n Demir s ahin, Oddur Kjartansson, Clara Rivera, and K \'o lá Túb \`o sún. 2020. https://doi.org/10.21437/Interspeech.2020-1096 Developing an Open-Source Corpus of Yoruba Speech . In Proceedings of Interspeech 2020, pages 404--408, Shanghai, China. Intern...
2020 doi
-
[27]
Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al. 2024. Audiogpt: Understanding and generating speech, music, sound, and talking head. In Proceedings of the AAAI Conference on Artificial ...
2024
-
[28]
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020. The state and fate of linguistic diversity and inclusion in the nlp world. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, page 6282. Association...
2020
-
[29]
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Victor Sanh, Thomas Wolf, Lysandre Mouillet, Teven Le Scao, and Alexander M. Rush. 2021. Datasets: A community library for ...
2021
-
[30]
Josh Meyer, David Adelani, Edresson Casanova, Alp \"O ktem, Daniel Whitenack, Julian Weber, Salomon Kabongo Kabenamualu, Elizabeth Salesky, Iroro Orife, Colin Leong, Perez Ogayo, Chris Chinenye Emezue, Jonathan Mukiibi, Salomey Osei, Apelete Agbolo, Victor Akinode, Bernard Opo...
2022 arXiv
-
[31]
T. I. Modipa, M. H. Davel, and F. De Wet. 2015. Implications of sepedi/english code switching for asr systems. In Proceedings of the Pattern Recognition Association of South Africa (PRASA), pages 112--117
2015
-
[32]
Mozilla Foundation . 2023. Mozilla common voice: A massively multilingual open dataset for voice technologies. https://commonvoice.mozilla.org
2023
-
[33]
NaijaVoices . 2024. Naijavoices dataset: A multilingual speech corpus for nigerian languages. https://naijavoices.com/
2024
-
[34]
Tu Anh Nguyen, Eugene Kharitonov, Jade Copet, Yossi Adi, Wei-Ning Hsu, Ali Elkahky, Paden Tomasello, Robin Algayres, Benoit Sagot, Abdelrahman Mohamed, et al. 2023. Generative spoken dialogue language modeling. Transactions of the Association for Computational Linguistics, 11:250--266
2023
-
[35]
Tu Anh Nguyen, Benjamin Muller, Bokai Yu, Marta R Costa-Jussa, Maha Elbayad, Sravya Popuri, Christophe Ropers, Paul-Ambroise Duquenne, Robin Algayres, Ruslan Mavlyutov, et al. 2025. Spirit-lm: Interleaved spoken and written language model. Transactions of the Association for C...
2025
-
[36]
Sewade Ogun, Abraham T Owodunni, Tobi Olatunji, Eniola Alese, Babatunde Oladimeji, Tejumade Afonja, Kayode Olaleye, Naome A Etori, and Tosin Adewumi. 2024. https://www.isca-archive.org/interspeech_2024/ogun24_interspeech.pdf 1000 african voices: Advancing inclusive multi-speak...
2024
-
[37]
Jessica Ojo, Kelechi Ogueji, Pontus Stenetorp, and David Ifeoluwa Adelani. 2023. How good are large language models on african languages? arXiv e-prints, pages arXiv--2311
2023
-
[38]
Akintunde Oladipo, Mofetoluwa Adeyemi, Orevaoghene Ahia, Abraham Toluwalase Owodunni, Odunayo Ogundepo, David Ifeoluwa Adelani, and Jimmy Lin. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.11 Better quality pre-training data and t5 models for A frican languages . In Procee...
2023 doi
-
[39]
Tobi Olatunji, Tejumade Afonja, Aditya Yadavalli, Chris Chinenye Emezue, Sahib Singh, Bonaventure FP Dossou, Joanne Osuchukwu, Salomey Osei, Atnafu Lambebo Tonja, Naome Etori, et al. 2023. Afrispeech-200: Pan-african accented speech dataset for clinical and general domain asr....
2023 arXiv
-
[40]
Alexis Plaquet and Hervé Bredin. 2023. Powerset multi-class cross entropy loss for neural speaker diarization . In Proc. INTERSPEECH 2023
2023
-
[41]
Edoardo Maria Ponti, Helen O ' Horan, Yevgeni Berzak, Ivan Vuli \'c , Roi Reichart, Thierry Poibeau, Ekaterina Shutova, and Anna Korhonen. 2019. https://doi.org/10.1162/coli_a_00357 Modeling language variation and universals: A survey on typological linguistics for natural lan...
2019 doi
-
[42]
Vineel Pratap, Ann Lee, Qiantong Xu, Anuroop Sriram, Tatiana Likhomanenko, Brandon Sottile, et al. 2023. Scaling speech technology to 1,000+ languages. In Proc. Interspeech
2023
-
[43]
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022. https://arxiv.org/abs/2212.04356 Robust speech recognition via large-scale weak supervision . Preprint, arXiv:2212.04356
2022 arXiv
-
[44]
Machel Reid, Junjie Hu, Graham Neubig, and Yutaka Matsuo. 2021. Afromt: Pretraining strategies and reproducible benchmarks for translation of 8 african languages. arXiv preprint arXiv:2109.04715
2021 arXiv
-
[45]
Claytone Sikasote and Antonios Anastasopoulos. 2022. https://aclanthology.org/2022.lrec-1.790 Bembaspeech: A speech recognition corpus for the bemba language . In Proceedings of the Language Resources and Evaluation Conference, pages 7277--7283, Marseille, France. European Lan...
2022
-
[46]
Claytone Sikasote, Kalinda Siaminwe, Stanly Mwape, Bangiwe Zulu, Mofya Phiri, Martin Phiri, David Zulu, Mayumbo Nyirenda, and Antonios Anastasopoulos. 2023. https://doi.org/10.21437/Interspeech.2023-1979 Zambezi Voice: A Multilingual Speech Corpus for Zambian Languages . In Pr...
2023 doi
-
[47]
The Brick House Cooperative . 2024. Olongoafrica multilingual anthology. https://lingua.olongoafrica.com/. A collection of translated and narrated short stories in various African languages, including Edo, Tamazight, Yoruba, Swahili, Hausa, Tiv, Shona, Ibibio, Igbo, and Nigeri...
2024
-
[48]
Hawau Olamide Toyin, Amirbek Djanibekov, Ajinkya Kulkarni, and Hanan Aldarmaki. 2023. https://doi.org/10.18653/v1/2023.arabicnlp-1.5 A r TST : A rabic text and speech transformer . In Proceedings of ArabicNLP 2023, pages 41--51, Singapore (Hybrid). Association for Computationa...
2023 doi
-
[49]
Universal Declaration of Human Rights Audio . 2025. Universal declaration of human rights audio project. https://udhr.audio/. A project providing audio recordings of the Universal Declaration of Human Rights in multiple languages to promote accessibility and linguistic diversity
2025
- [50]
-
[51]
Charl Van Heerden, Neil Kleynhans, and Marelie H. Davel. 2016. http://www.isca-speech.org/archive/Interspeech_2016/pdfs/1412.PDF Improving the lwazi asr baseline . In Proceedings of Interspeech 2016
2016
-
[52]
Daniel van Niekerk, Charl van Heerden, Marelie Davel, Neil Kleynhans, Oddur Kjartansson, Martin Jansche, and Linne Ha. 2017. http://dx.doi.org/10.21437/Interspeech.2017-1139 Rapid development of TTS corpora for four South African languages . In Proc. Interspeech 2017, pages 21...
2017 doi
-
[53]
Voice of Africa . 2025. Voice of africa. https://thevoiceofafrica.com/about/. A multilingual platform delivering news and stories from across the African continent
2025
-
[54]
Jiayi Wang, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, et al. 2023. Afrimte and africomet: Enhancing comet to embrace under-resourced african languages. arXiv preprint arXiv:2311.09828
2023 arXiv
-
[55]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020
-
[56]
Marcely Zanon Boito , Vivek Iyer, Nikolaos Lagos, Laurent Besacier, and Ioan Calapodescu. 2024. https://doi.org/10.21437/Interspeech.2024-938 mhubert-147: A compact multilingual hubert model . In Interspeech 2024, pages 3939--3943
2024 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.