REVIEW 4 major objections 6 minor 1 cited by
The State of Large Language Models for African Languages: Progress and Challenges
T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that roughly 42 of Africa's 2,000+ languages have any support in current LLMs, with four languages — Amharic, Swahili, Afrikaans, Malagasy — appearing in every multilingual SLM and only three of 23 active scripts handled.
desk verdict Useful survey, but the signature 42-language count is built on SLM/SSLM tables only and overclaims LLM support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is a language-by-model support matrix: a 42-language by ten-model grid of yes/no entries built by reviewing each model's official documentation, technical reports, and publications. This matrix produces the headline counts, while a companion 37-script inventory (23 of which are still in use) generates the script coverage result. The paper also uses the distinction between script-agnostic, partially script-agnostic, and non-script-agnostic tokenizers to explain why models that handle Latin and Ge'ez well can still fail on scripts they never saw during training.
What would settle it
Run a standardized probe — language identification, generation fluency, translation accuracy, and basic comprehension — on all 42 claimed languages plus a random sample of unsupported African languages across the reviewed models; if substantially more than 42 languages are usable, or if more than four languages pass in every multilingual SLM, the paper's central count is wrong.
Extended reading notes
Core claim
The paper's central claim is that current language-model ecosystems leave Africa almost entirely out of the loop. Out of more than 2,000 African languages, only about 42 appear in any of the reviewed LLMs, SLMs, or SSLMs, and just four of those — Amharic, Swahili, Afrikaans, and Malagasy — are supported by every multilingual SLM examined. Script support is equally narrow, with only Latin, Arabic, and Ge'ez among 23 actively used scripts being widely handled. The study also identifies 23 publicly available African-language datasets and argues that these resources are sparse, unevenly distributed across tasks, and concentrated in relatively simple classification problems rather than translation, named entity recognition, or question answering.
Load-bearing premise
The 42-language count assumes each model's official documentation or technical report is an accurate and complete list of the languages the model can actually handle; several major LLMs have no such documentation, so their counted languages are uncertain.
Editorial extensions
If this is right
- Scaling up generic multilingual models will not close the gap by itself: most of the 42 supported languages appear in only one or two models, so future coverage depends on targeted data collection and tokenizer design.
- A practical benchmark for progress is clear: the next milestone would be moving more than a handful of languages into the 'supported by all multilingual SLMs' column.
- The script finding implies that byte-level or character-level tokenizers, which avoid script-specific vocabularies, are a concrete route to improving support for the roughly 20 neglected active scripts.
- The 23 available datasets are concentrated in classification and sentiment tasks, so building translation, NER, and question-answering corpora is a direct prerequisite for broader model evaluation.
- The proposed roadmap — standardisation and normalisation, then quality datasets, then specialised small models, then general SLMs, then LLMs — makes data and benchmark creation the true bottleneck rather than model size.
Reading between the lines
- Beyond the paper, the 42-language figure probably underestimates what undocumented LLMs can produce in practice, but it also overestimates reliable support because 'documented coverage' is not the same as verified fluency; probing each claimed language would settle the real number.
- The script-agnostic finding suggests a testable extension: adding script normalisation and a small set of active scripts (for instance N'Ko, Tifinagh, Adlam, or Vai) to existing tokenizers could raise coverage without training from scratch.
- A useful extension of the review would be to weight the support matrix by speaker population, since four well-supported languages are not necessarily the most spoken across the continent, and such a weighted view could redirect resource allocation.
- The scaling-law discussion implies that African language models may not need to reach trillion-token scales to be useful; the binding constraint is corpus quality and evaluation, so a compute-optimal path for low-resource languages may diverge from the usual parameter-count rules.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a review of language coverage, script support, training data, benchmarks, and technical challenges for African languages across three model categories: six large language models (LLMs), eight foundational small language models (SLMs), and six specialized African-centric small language models (SSLMs). The headline findings are that only about 42 of Africa's 2,000+ languages receive any support in existing LLMs, SLMs, or SSLMs; that four languages (Amharic, Swahili, Afrikaans, Malagasy) are supported by all multilingual SLMs; that only Latin, Arabic, and Ge'ez scripts are widely handled among 23 active scripts; and that 23 public datasets are available for African-language NLP. The paper concludes with a roadmap from linguistic standardization and corpus development to SSLMs, larger SLMs, and eventually LLMs.
Significance. The review addresses an important and timely gap: African languages are severely underrepresented in NLP, and a systematic, checkable account of current coverage would be valuable to researchers, funders, and model developers. The paper's strength is that it compiles detailed tables (Tables 4-10) with model architectures, training data, tokenization, language support, and datasets, and it makes its central counts falsifiable. The roadmap in Figure 3 is a reasonable synthesis of the field's priorities. However, the significance is conditional: the headline statistic of '42 supported languages' conflates LLMs with SLMs/SSLMs in a way the tables do not justify, and several counts in the tables are internally inconsistent. These issues must be resolved before the paper's main quantitative claims can be relied upon.
major comments (4)
- [Section 4 (Question One), Conclusion, Appendix B Table 8] The 'about 42 supported languages' claim is not supported by the presented evidence for LLMs. Appendix B Table 8, which is the only support for the count, has columns only for the six SSLMs and four multilingual SLMs (AfriBERTa, AfriTeVa, AfroLM, AfroXLMR, EthioLLM, EthioMT, mBERT, mT5, XLM-R, NLLB-200); none of the six reviewed LLMs from Table 1 appears. Section 4 (Question One) further states that GPT-4, Gemini 1.5, PaLM 2, and DeepSeek 'have no clear documentation about the languages they support.' The abstract and conclusion nevertheless say 'only about 42 have any support in existing LLMs, SLMs, or SSLMs.' This is a scope mismatch: at most the 42-language figure is the union over the ten tabulated SLMs/SSLMs. Since Table 3 attributes '101-language focused' coverage to Aya 23, the paper's own data suggest that LLM coverage could be larger than 42, so the claim is not a conservative lower bound. Please either add an empirically grounded LLM coverage column, or revise the central claim to refer only to SLMs and SSLMs.
- [Section 4 (Question One), Tables 6-8] The numerical backbone of the review is internally inconsistent. The text says 'a total of 38 African languages are supported across six SLMs. Collectively, these models support approximately 42 African languages.' Table 7 lists 38 foundation-model languages, while Table 8 lists 42 rows for the combined SLM/SSLM set, but the tables count different sets, and Table 8 double-counts 'Afaan Oromo' and 'Oromo' (rows 12-13), which are the same language. It also treats 'Fulah' (row 21) and 'Luganda' (row 9) separately from 'Fulfulde' and 'Ganda' used in Table 7, so the distinct-language union is below 42. In addition, 'Luo' has support count 1 in Table 7 (row 30) but support count 5 in Table 8 (row 10); this may be defensible if the scopes differ, but the paper never explains the scope change. Please provide one deduplicated language list and a single, clearly defined support-counting method.
- [Section 4 (Question Four), Table 9] The '23 public datasets' figure is not a count of African-language benchmark datasets. Table 9 includes CoNLL 2003 NER (row 16), ANERCorp (row 17), and AG News (row 19), which are English or Arabic datasets, and the 'Languages' column aggregates non-African languages as well. The text's wording, 'around 23 publicly available datasets are used by models solely SSLMs,' therefore overstates the African-language resource base. Please filter the table to African-language datasets, or redefine the statistic clearly (e.g., '23 datasets used in the reviewed SSLMs, of which X target African languages').
- [Section 4 (Question Two), Table 10] The script-coverage claim is not backed by a model-script matrix. The text asserts that 'from 23 actively working scripts, only 3 are used in large language and small language models,' but Table 10 merely lists scripts with usage status; no table or mapping shows which of the reviewed models support Tifinagh, Vai, Bamum, N'Ko, Adlam, or the other active scripts. Please add a script-support matrix for the reviewed models, or limit the claim to 'the reviewed models,' and note where script support is inferred rather than documented.
minor comments (6)
- [Section 3.1, Table 1] The paper does not disclose that authors Abinew Ali Ayele and Seid Muhie Yimam are co-authors of EthioLLM [Tonja et al., 2024b] and EthioMT [Tonja et al., 2024c], both reviewed favorably in Section 4 and Table 6. Please add a conflict-of-interest or author-contribution disclosure.
- [Author affiliations] In the author block, 'Bayero University Kano, India' should presumably read 'Bayero University Kano, Nigeria.'
- [Table 3 and Section 2] The paper's parameter-based SLM/LLM threshold (<7B) is inconsistent with placing Aya 23 (8B, Table 3) in the LLM category; explain the categorization, since Aya 23 is later absent from Table 8.
- [Table 10] Row 27 lists 'Luo' as a script, but Luo (Dholuo) is a language, not a script; this appears to be an error in a table about writing systems.
- [References] Several citations are incomplete: the Google Research references (2019, 2021) contain '[Insert Date]' placeholders, and the Meta AI Research NLLB reference lacks full author and venue details.
- [Figure 3 caption] The caption begins 'Figure 3: Figure 3: Roadmap...' with a duplicated label; please remove the duplicate.
Circularity Check
No significant circularity: the paper is a review whose headline counts are compiled from its own tables, and its self-citations are not load-bearing.
full rationale
This is a survey/review, not a derivation, so most circularity patterns do not apply. The central quantities (42 supported languages, 23 datasets, three scripts) are compiled counts from Tables 6-10, and the abstract and conclusion merely restate those counts rather than predicting them. The only self-citations (EthioLLM, EthioMT, EthioBenchMarks) point to peer-reviewed external publications by members of the same group; these models are reviewed alongside other SSLMs, and they are not used as evidence for the underrepresentation claim. If anything, removing them would lower the reported 42-language count and strengthen the paper's conclusion, so the central claim is not forced by self-citation. The paper itself notes that GPT-4, Gemini 1.5, PaLM 2, and DeepSeek lack documented language lists (Section 4, Question One), and Table 8 has no LLM columns, so the 'LLMs' portion of the conclusion is under-supported by the presented evidence. That is a scope and accuracy limitation, not circularity: the count remains the union of the explicitly tabulated SLMs/SSLMs and is not derived from the conclusion it is used to support. No circular step can be quoted from the text.
Assumptions & free parameters
assumptions (4)
- domain assumption Africa has over 2,000 languages.
- domain assumption 23 of 37 African scripts are in active use.
- domain assumption Models with fewer than 7B parameters are SLMs, and models with emergent abilities are LLMs.
- domain assumption Chinchilla scaling N approximately equals 20 times D, where N is parameters and D is tokens.
Cite this review
Pith. "Pith review of The State of Large Language Models for African Languages: Progress and Challenges." pith.science (2026). https://pith.science/paper/5HTHKJXY
@misc{pith2026250602280,
author = {Pith},
title = {Pith review of: The State of Large Language Models for African Languages: Progress and Challenges},
year = {2026},
howpublished = {\url{https://pith.science/paper/5HTHKJXY}},
note = {Machine review of arXiv:2506.02280}
}
read the original abstract
Large Language Models (LLMs) are transforming Natural Language Processing (NLP), but their benefits are largely absent for Africa's 2,000 low-resource languages. This paper comparatively analyzes African language coverage across six LLMs, eight Small Language Models (SLMs), and six Specialized SLMs (SSLMs). The evaluation covers language coverage, training sets, technical limitations, script problems, and language modelling roadmaps. The work identifies 42 supported African languages and 23 available public data sets, and it shows a big gap where four languages (Amharic, Swahili, Afrikaans, and Malagasy) are always treated while there is over 98\% of unsupported African languages. Moreover, the review shows that just Latin, Arabic, and Ge'ez scripts are identified while 20 active scripts are neglected. Some of the primary challenges are lack of data, tokenization biases, computational costs being very high, and evaluation issues. These issues demand language standardization, corpus development by the community, and effective adaptation methods for African languages.
Figures
Forward citations
Cited by 1 Pith paper
-
AfroScope: A Framework for Studying the Linguistic Landscape of Africa
A new framework combines a 713-language African LID dataset, strong baselines, and a contrastive-embedding hierarchical step that improves macro-F1 by 4.55 on a 29-language confusable subset.
Reference graph
Works this paper leans on
-
[1]
Masakhaner: Named entity recognition for african languages
David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, et al. Masakhaner: Named entity recognition for african languages. Transactions of the Association for Computational Linguistics, 9: 0 1116--1131, 2021
2021
-
[2]
Alabi, David Ifeoluwa Adelani, Marius Mosbach, and Dietrich Klakow
Jesujoba O. Alabi, David Ifeoluwa Adelani, Marius Mosbach, and Dietrich Klakow. Adapting pre-trained language models to A frican languages via multilingual adaptive fine-tuning. In Nicoletta Calzolari, Chu-Ren Huang, Hansaem Kim, James Pustejovsky, Leo Wanner, Key-Sun Choi, Pum-Mo Ryu, Hsin-Hsi Chen, Lucia Donatelli, Heng Ji, Sadao Kurohashi, Patrizia Pag...
work page 2022
-
[3]
A ra BERT : Transformer-based model for A rabic language understanding
Wissam Antoun, Fady Baly, and Hazem Hajj. A ra BERT : Transformer-based model for A rabic language understanding. In Hend Al-Khalifa, Walid Magdy, Kareem Darwish, Tamer Elsayed, and Hamdy Mubarak, editors, Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection, pages 9--15, Ma...
work page 2020
-
[4]
A ra GPT 2: Pre-trained transformer for A rabic language generation
Wissam Antoun, Fady Baly, and Hazem Hajj. A ra GPT 2: Pre-trained transformer for A rabic language generation. In Nizar Habash, Houda Bouamor, Hazem Hajj, Walid Magdy, Wajdi Zaghouani, Fethi Bougares, Nadi Tomeh, Ibrahim Abu Farha, and Samia Touileb, editors, Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 196--207, Kyiv, Ukrai...
work page 2021
-
[5]
The world's writing systems, 2024
Atelier National de Recherche Typographique (ANRT) , Institut Designlabor Gutenberg (IDG) , and Script Encoding Initiative (SEI) . The world's writing systems, 2024. URL https://www.worldswritingsystems.org/. Accessed: 2025-05-30
work page 2024
-
[6]
Clark, Dan Garrette, Iulia Turc, and John Wieting
Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting. Canine: Pre-training an efficient tokenization-free encoder for language representation. Transactions of the Association for Computational Linguistics, 10: 0 73--91, 2022. doi:10.1162/tacl_a_00448. URL https://aclanthology.org/2022.tacl-1.5/
-
[7]
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross-lingual representation learning at scale. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault, editors, Proceedings of the 58th Annual Meeting of the ...
doi:10.18653/v1/2020 2020
-
[8]
Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Yan He, Elahe Kalbassi, Linlu Wang, Long Wang, Shuo Zhang, Angela Fan, et al. No language left behind: Scaling human-centered machine translation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), Abu Dhabi, UAE, 2022. Associati...
work page 2022
Show all 48 references
-
[9]
I ndic BART : A pre-trained model for indic natural language generation
Raj Dabre, Himani Shrotriya, Anoop Kunchukuttan, Ratish Puduppully, Mitesh Khapra, and Pratyush Kumar. I ndic BART : A pre-trained model for indic natural language generation. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Findings of the Association for...
2022 doi
-
[10]
Dai, Jonathan H
Andrew M. Dai, Jonathan H. Clark, Kevin Robinson, Maysam Moussalem, Sebastian Ruder, Siamak Shakeri, Jacob Austin, et al. Palm 2 technical report. Technical report, Google, 2023
2023
-
[11]
Deepseek-v3 technical report
DeepSeek-AI. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, December 2024
2024 arXiv
-
[12]
BERT : Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North A merican Chapter ...
2019
-
[13]
Bonaventure F. P. Dossou, Atnafu Lambebo Tonja, Oreen Yousuf, Salomey Osei, Abigail Oppong, Iyanuoluwa Shode, Oluwabusayo Olufunke Awoyomi, and Chris Emezue. A fro LM : A self-active learning-based multilingual pretrained language model for 23 A frican languages. In Angela Fan...
2022
-
[14]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zha...
2024
-
[15]
Ethnologue: Languages of the world, 2025
Ethnologue . Ethnologue: Languages of the world, 2025. Accessed: 2025-05-20
2025
-
[16]
Language-agnostic BERT sentence embedding
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. Language-agnostic BERT sentence embedding. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics ...
2022 doi
-
[17]
Guerreiro, Duarte M
Nuno M. Guerreiro, Duarte M. Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, and Andr \'e F. T. Martins. Hallucinations in large multilingual translation models. Transactions of the Association for Computational Linguistics, 11: 0 1500--1517, 2023. doi:...
2023 doi
-
[18]
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022
2022 arXiv
-
[19]
Glot500: Scaling multilingual corpora and language models to 500 languages
Ayyoob Imani, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, Andr \'e Martins, Fran c ois Yvon, and Hinrich Sch \"u tze. Glot500: Scaling multilingual corpora and language models to 500 languages. In Anna Roger...
2023
-
[20]
A fri T e VA : Extending ?small data? pretraining approaches to sequence-to-sequence models
Odunayo Jude Ogundepo, Akintunde Oladipo, Mofetoluwa Adeyemi, Kelechi Ogueji, and Jimmy Lin. A fri T e VA : Extending ?small data? pretraining approaches to sequence-to-sequence models. In Colin Cherry, Angela Fan, George Foster, Gholamreza (Reza) Haffari, Shahram Khadivi, Nan...
2022
-
[21]
I ndo LEM and I ndo BERT : A benchmark dataset and pre-trained language model for I ndonesian NLP
Fajri Koto, Afshin Rahimi, Jey Han Lau, and Timothy Baldwin. I ndo LEM and I ndo BERT : A benchmark dataset and pre-trained language model for I ndonesian NLP . In Donia Scott, Nuria Bel, and Chengqing Zong, editors, Proceedings of the 28th International Conference on Computat...
2020 doi
-
[22]
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. Cross-lingual language model pretraining. Advances in Neural Information Processing Systems (NeurIPS), 32, 2019
2019
-
[23]
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461, 2019
1910 arXiv
-
[24]
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019
1907 arXiv
-
[25]
Scaling laws for fact memorization of large language models
Xingyu Lu, Xiaonan Li, Qinyuan Cheng, Kai Ding, Xuanjing Huang, and Xipeng Qiu. Scaling laws for fact memorization of large language models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, pa...
2024 doi
-
[26]
F in GPT : Large generative models for a small language
Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, Thomas Wang, Nouamane Tazi, Teven Scao, Thomas Wolf, Osma Suominen, Samuli Sairanen, Mikko Merioksa, Jyrki ...
2023
-
[27]
A ra T 5: Text-to-text transformers for A rabic language generation
El Moatez Billah Nagoudi, AbdelRahim Elmadany, and Muhammad Abdul-Mageed. A ra T 5: Text-to-text transformers for A rabic language generation. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Co...
2022 doi
-
[28]
Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced african languages
Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced african languages. In Proceedings of the First Workshop on Natural Language Processing for African Languages (AfriNLP 2021), p...
2021 doi
-
[29]
Gpt-4 technical report, 2023
OpenAI. Gpt-4 technical report, 2023. Technical Report
2023
-
[30]
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21 0 (140): 0 1--67, 2020
2020
-
[31]
Bert multilingual model readme, 2019
Google Research. Bert multilingual model readme, 2019. Accessed: [Insert Date]
2019
-
[32]
Multilingual t5: Massively multilingual pre-trained text-to-text transformer, 2021
Google Research. Multilingual t5: Massively multilingual pre-trained text-to-text transformer, 2021. Accessed:
2021
-
[33]
NLLB-200 : Distilled 600m-parameter model of No Language Left Behind ( NLLB ), 2022
Meta AI Research. NLLB-200 : Distilled 600m-parameter model of No Language Left Behind ( NLLB ), 2022. Accessed:
2022
-
[34]
How good is your tokenizer? on the monolingual performance of multilingual language models
Phillip Rust, Jonas Pfeiffer, Ivan Vuli \'c , Sebastian Ruder, and Iryna Gurevych. How good is your tokenizer? on the monolingual performance of multilingual language models. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Me...
2021
-
[35]
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019
1910 arXiv
-
[36]
The language barrier: Dissecting safety challenges of LLM s in multilingual contexts
Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, and Daniel Khashabi. The language barrier: Dissecting safety challenges of LLM s in multilingual contexts. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findi...
2024 doi
-
[37]
Gemini: A family of highly capable multimodal models, December 2023
Gemini Team and Google DeepMind. Gemini: A family of highly capable multimodal models, December 2023. Technical Report
2023
-
[38]
Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al
NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. No language left behind: Scaling human-centered machine translation. In Proceedings of the 2022 Conferenc...
2022
-
[39]
Ethiollm: Multilingual large language models for ethiopian languages with task evaluation
Atnafu Lambebo Tonja, Israel Abebe Azime, Tadesse Destaw Belay, Mesay Gemeda Yigezu, Moges Ahmed Mehamed, Abinew Ali Ayele, Ebrahim Chekol Jibril, Michael Melese Woldeyohannis, Olga Kolesnikova, Philipp Slusallek, et al. Ethiollm: Multilingual large language models for ethiopi...
2024 arXiv
-
[40]
E thio LLM : Multilingual large language models for E thiopian languages with task evaluation
Atnafu Lambebo Tonja, Israel Abebe Azime, Tadesse Destaw Belay, Mesay Gemeda Yigezu, Moges Ahmed Ah Mehamed, Abinew Ali Ayele, Ebrahim Chekol Jibril, Michael Melese Woldeyohannis, Olga Kolesnikova, Philipp Slusallek, Dietrich Klakow, and Seid Muhie Yimam. E thio LLM : Multilin...
2024
-
[41]
E thio MT : Parallel corpus for low-resource E thiopian languages
Atnafu Lambebo Tonja, Olga Kolesnikova, Alexander Gelbukh, and Jugal Kalita. E thio MT : Parallel corpus for low-resource E thiopian languages. In Rooweither Mabuya, Muzi Matfunjwa, Mmasibidi Setaka, and Menno van Zaanen, editors, Proceedings of the Fifth Workshop on Resources...
2024
-
[42]
Aya model: An instruction finetuned open-access multilingual language model
Ahmet \"U st \"u n, Viraat Aryabumi, Zheng Yong, Wei-Yin Ko, Daniel D ' souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker. Aya mode...
2024
-
[43]
A comprehensive survey of small language models in the era of large language models: Techniques, enhancements, applications, collaboration with llms, and trustworthiness
Fali Wang, Zhiwei Zhang, Xianren Zhang, Zongyu Wu, TzuHao Mo, Qiuhao Lu, Wanjing Wang, Rui Li, Junjie Xu, Xianfeng Tang, Qi He, Yao Ma, Ming Huang, and Suhang Wang. A comprehensive survey of small language models in the era of large language models: Techniques, enhancements, a...
2024 arXiv
-
[44]
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35: 0 24824--24837, 2022
2022
-
[45]
Bloom: A 176b-parameter open-access multilingual language model
BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Saso Luccioni, François Yvon, et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint, arXiv:2211.05100, 2022...
2022 arXiv
-
[46]
Shijie Wu and Mark Dredze. Are all languages created equal in multilingual BERT ? In Spandana Gella, Johannes Welbl, Marek Rei, Fabio Petroni, Patrick Lewis, Emma Strubell, Minjoon Seo, and Hannaneh Hajishirzi, editors, Proceedings of the 5th Workshop on Representation Learnin...
2020 doi
-
[47]
m T 5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. m T 5: A massively multilingual pre-trained text-to-text transformer. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, St...
2021
-
[48]
B y T 5: Towards a token-free future with pre-trained byte-to-byte models
Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. B y T 5: Towards a token-free future with pre-trained byte-to-byte models. Transactions of the Association for Computational Linguistics, 10: 0 291--306, 2022. do...
2022 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.