Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

The State of Large Language Models for African Languages: Progress and Challenges

T0 review · 4 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper claims that roughly 42 of Africa's 2,000+ languages have any support in current LLMs, with four languages — Amharic, Swahili, Afrikaans, Malagasy — appearing in every multilingual SLM and only three of 23 active scripts handled.

desk verdict Useful survey, but the signature 42-language count is built on SLM/SSLM tables only and overclaims LLM support. read the letter →

arxiv 2506.02280 v3 pith:5HTHKJXY submitted 2025-06-02 cs.AI

classification cs.AI
keywords Africanlanguageslargelanguagemodelssmalllow-resourceNLPcoveragescripttokenizationmultilingualbenchmarkdatasets
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review compares six large language models, eight foundational small language models, and six African-centred specialised models to measure how much of the continent's linguistic diversity is actually served. The authors' central finding is that only about 42 of more than 2,000 African languages are supported by any of these models, and that four languages — Amharic, Swahili, Afrikaans, and Malagasy — receive support in every multilingual small language model. Script coverage is similarly thin: only Latin, Arabic, and Ge'ez are widely handled, leaving roughly 20 active scripts unsupported. The paper cares because the gap is not just about a few languages; it means more than 98% of African languages are effectively absent from modern language technology, and the obstacles are documented as missing data, tokenization bias, high compute costs, and weak evaluation infrastructure.

What carries the argument

The load-bearing device is a language-by-model support matrix: a 42-language by ten-model grid of yes/no entries built by reviewing each model's official documentation, technical reports, and publications. This matrix produces the headline counts, while a companion 37-script inventory (23 of which are still in use) generates the script coverage result. The paper also uses the distinction between script-agnostic, partially script-agnostic, and non-script-agnostic tokenizers to explain why models that handle Latin and Ge'ez well can still fail on scripts they never saw during training.

What would settle it

Run a standardized probe — language identification, generation fluency, translation accuracy, and basic comprehension — on all 42 claimed languages plus a random sample of unsupported African languages across the reviewed models; if substantially more than 42 languages are usable, or if more than four languages pass in every multilingual SLM, the paper's central count is wrong.

Watch

Extended reading notes

Core claim

The paper's central claim is that current language-model ecosystems leave Africa almost entirely out of the loop. Out of more than 2,000 African languages, only about 42 appear in any of the reviewed LLMs, SLMs, or SSLMs, and just four of those — Amharic, Swahili, Afrikaans, and Malagasy — are supported by every multilingual SLM examined. Script support is equally narrow, with only Latin, Arabic, and Ge'ez among 23 actively used scripts being widely handled. The study also identifies 23 publicly available African-language datasets and argues that these resources are sparse, unevenly distributed across tasks, and concentrated in relatively simple classification problems rather than translation, named entity recognition, or question answering.

Load-bearing premise

The 42-language count assumes each model's official documentation or technical report is an accurate and complete list of the languages the model can actually handle; several major LLMs have no such documentation, so their counted languages are uncertain.

Editorial extensions

If this is right

  • Scaling up generic multilingual models will not close the gap by itself: most of the 42 supported languages appear in only one or two models, so future coverage depends on targeted data collection and tokenizer design.
  • A practical benchmark for progress is clear: the next milestone would be moving more than a handful of languages into the 'supported by all multilingual SLMs' column.
  • The script finding implies that byte-level or character-level tokenizers, which avoid script-specific vocabularies, are a concrete route to improving support for the roughly 20 neglected active scripts.
  • The 23 available datasets are concentrated in classification and sentiment tasks, so building translation, NER, and question-answering corpora is a direct prerequisite for broader model evaluation.
  • The proposed roadmap — standardisation and normalisation, then quality datasets, then specialised small models, then general SLMs, then LLMs — makes data and benchmark creation the true bottleneck rather than model size.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the 42-language figure probably underestimates what undocumented LLMs can produce in practice, but it also overestimates reliable support because 'documented coverage' is not the same as verified fluency; probing each claimed language would settle the real number.
  • The script-agnostic finding suggests a testable extension: adding script normalisation and a small set of active scripts (for instance N'Ko, Tifinagh, Adlam, or Vai) to existing tokenizers could raise coverage without training from scratch.
  • A useful extension of the review would be to weight the support matrix by speaker population, since four well-supported languages are not necessarily the most spoken across the continent, and such a weighted view could redirect resource allocation.
  • The scaling-law discussion implies that African language models may not need to reach trillion-token scales to be useful; the binding constraint is corpus quality and evaluation, so a compute-optimal path for low-resource languages may diverge from the usual parameter-count rules.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This paper is a review of language coverage, script support, training data, benchmarks, and technical challenges for African languages across three model categories: six large language models (LLMs), eight foundational small language models (SLMs), and six specialized African-centric small language models (SSLMs). The headline findings are that only about 42 of Africa's 2,000+ languages receive any support in existing LLMs, SLMs, or SSLMs; that four languages (Amharic, Swahili, Afrikaans, Malagasy) are supported by all multilingual SLMs; that only Latin, Arabic, and Ge'ez scripts are widely handled among 23 active scripts; and that 23 public datasets are available for African-language NLP. The paper concludes with a roadmap from linguistic standardization and corpus development to SSLMs, larger SLMs, and eventually LLMs.

Significance. The review addresses an important and timely gap: African languages are severely underrepresented in NLP, and a systematic, checkable account of current coverage would be valuable to researchers, funders, and model developers. The paper's strength is that it compiles detailed tables (Tables 4-10) with model architectures, training data, tokenization, language support, and datasets, and it makes its central counts falsifiable. The roadmap in Figure 3 is a reasonable synthesis of the field's priorities. However, the significance is conditional: the headline statistic of '42 supported languages' conflates LLMs with SLMs/SSLMs in a way the tables do not justify, and several counts in the tables are internally inconsistent. These issues must be resolved before the paper's main quantitative claims can be relied upon.

major comments (4)
  1. [Section 4 (Question One), Conclusion, Appendix B Table 8] The 'about 42 supported languages' claim is not supported by the presented evidence for LLMs. Appendix B Table 8, which is the only support for the count, has columns only for the six SSLMs and four multilingual SLMs (AfriBERTa, AfriTeVa, AfroLM, AfroXLMR, EthioLLM, EthioMT, mBERT, mT5, XLM-R, NLLB-200); none of the six reviewed LLMs from Table 1 appears. Section 4 (Question One) further states that GPT-4, Gemini 1.5, PaLM 2, and DeepSeek 'have no clear documentation about the languages they support.' The abstract and conclusion nevertheless say 'only about 42 have any support in existing LLMs, SLMs, or SSLMs.' This is a scope mismatch: at most the 42-language figure is the union over the ten tabulated SLMs/SSLMs. Since Table 3 attributes '101-language focused' coverage to Aya 23, the paper's own data suggest that LLM coverage could be larger than 42, so the claim is not a conservative lower bound. Please either add an empirically grounded LLM coverage column, or revise the central claim to refer only to SLMs and SSLMs.
  2. [Section 4 (Question One), Tables 6-8] The numerical backbone of the review is internally inconsistent. The text says 'a total of 38 African languages are supported across six SLMs. Collectively, these models support approximately 42 African languages.' Table 7 lists 38 foundation-model languages, while Table 8 lists 42 rows for the combined SLM/SSLM set, but the tables count different sets, and Table 8 double-counts 'Afaan Oromo' and 'Oromo' (rows 12-13), which are the same language. It also treats 'Fulah' (row 21) and 'Luganda' (row 9) separately from 'Fulfulde' and 'Ganda' used in Table 7, so the distinct-language union is below 42. In addition, 'Luo' has support count 1 in Table 7 (row 30) but support count 5 in Table 8 (row 10); this may be defensible if the scopes differ, but the paper never explains the scope change. Please provide one deduplicated language list and a single, clearly defined support-counting method.
  3. [Section 4 (Question Four), Table 9] The '23 public datasets' figure is not a count of African-language benchmark datasets. Table 9 includes CoNLL 2003 NER (row 16), ANERCorp (row 17), and AG News (row 19), which are English or Arabic datasets, and the 'Languages' column aggregates non-African languages as well. The text's wording, 'around 23 publicly available datasets are used by models solely SSLMs,' therefore overstates the African-language resource base. Please filter the table to African-language datasets, or redefine the statistic clearly (e.g., '23 datasets used in the reviewed SSLMs, of which X target African languages').
  4. [Section 4 (Question Two), Table 10] The script-coverage claim is not backed by a model-script matrix. The text asserts that 'from 23 actively working scripts, only 3 are used in large language and small language models,' but Table 10 merely lists scripts with usage status; no table or mapping shows which of the reviewed models support Tifinagh, Vai, Bamum, N'Ko, Adlam, or the other active scripts. Please add a script-support matrix for the reviewed models, or limit the claim to 'the reviewed models,' and note where script support is inferred rather than documented.
minor comments (6)
  1. [Section 3.1, Table 1] The paper does not disclose that authors Abinew Ali Ayele and Seid Muhie Yimam are co-authors of EthioLLM [Tonja et al., 2024b] and EthioMT [Tonja et al., 2024c], both reviewed favorably in Section 4 and Table 6. Please add a conflict-of-interest or author-contribution disclosure.
  2. [Author affiliations] In the author block, 'Bayero University Kano, India' should presumably read 'Bayero University Kano, Nigeria.'
  3. [Table 3 and Section 2] The paper's parameter-based SLM/LLM threshold (<7B) is inconsistent with placing Aya 23 (8B, Table 3) in the LLM category; explain the categorization, since Aya 23 is later absent from Table 8.
  4. [Table 10] Row 27 lists 'Luo' as a script, but Luo (Dholuo) is a language, not a script; this appears to be an error in a table about writing systems.
  5. [References] Several citations are incomplete: the Google Research references (2019, 2021) contain '[Insert Date]' placeholders, and the Meta AI Research NLLB reference lacks full author and venue details.
  6. [Figure 3 caption] The caption begins 'Figure 3: Figure 3: Roadmap...' with a duplicated label; please remove the duplicate.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a review whose headline counts are compiled from its own tables, and its self-citations are not load-bearing.

full rationale

This is a survey/review, not a derivation, so most circularity patterns do not apply. The central quantities (42 supported languages, 23 datasets, three scripts) are compiled counts from Tables 6-10, and the abstract and conclusion merely restate those counts rather than predicting them. The only self-citations (EthioLLM, EthioMT, EthioBenchMarks) point to peer-reviewed external publications by members of the same group; these models are reviewed alongside other SSLMs, and they are not used as evidence for the underrepresentation claim. If anything, removing them would lower the reported 42-language count and strengthen the paper's conclusion, so the central claim is not forced by self-citation. The paper itself notes that GPT-4, Gemini 1.5, PaLM 2, and DeepSeek lack documented language lists (Section 4, Question One), and Table 8 has no LLM columns, so the 'LLMs' portion of the conclusion is under-supported by the presented evidence. That is a scope and accuracy limitation, not circularity: the count remains the union of the explicitly tabulated SLMs/SSLMs and is not derived from the conclusion it is used to support. No circular step can be quoted from the text.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no new fitted parameters or invented entities. It relies on background facts and classifications from cited sources, which are listed as axioms.

assumptions (4)
  • domain assumption Africa has over 2,000 languages.
    Used as denominator for the 98% unsupported claim; sourced from Ethnologue [2025], not independently verified.
  • domain assumption 23 of 37 African scripts are in active use.
    Taken from a typography source (Atelier National de Recherche Typographique et al., 2024); the 23-script number drives the script-neglect claim.
  • domain assumption Models with fewer than 7B parameters are SLMs, and models with emergent abilities are LLMs.
    The paper adopts Wang et al. [2024] classification to categorize models into LLMs/SLMs/SSLMs.
  • domain assumption Chinchilla scaling N approximately equals 20 times D, where N is parameters and D is tokens.
    Used in the roadmap discussion to reason about model sizes versus training tokens; note the relation is stated inverted relative to Hoffmann et al.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The State of Large Language Models for African Languages: Progress and Challenges." pith.science (2026). https://pith.science/paper/5HTHKJXY

@misc{pith2026250602280,
  author       = {Pith},
  title        = {Pith review of: The State of Large Language Models for African Languages: Progress and Challenges},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5HTHKJXY}},
  note         = {Machine review of arXiv:2506.02280}
}
read the original abstract

Large Language Models (LLMs) are transforming Natural Language Processing (NLP), but their benefits are largely absent for Africa's 2,000 low-resource languages. This paper comparatively analyzes African language coverage across six LLMs, eight Small Language Models (SLMs), and six Specialized SLMs (SSLMs). The evaluation covers language coverage, training sets, technical limitations, script problems, and language modelling roadmaps. The work identifies 42 supported African languages and 23 available public data sets, and it shows a big gap where four languages (Amharic, Swahili, Afrikaans, and Malagasy) are always treated while there is over 98\% of unsupported African languages. Moreover, the review shows that just Latin, Arabic, and Ge'ez scripts are identified while 20 active scripts are neglected. Some of the primary challenges are lack of data, tokenization biases, computational costs being very high, and evaluation issues. These issues demand language standardization, corpus development by the community, and effective adaptation methods for African languages.

Figures

Figures reproduced from arXiv: 2506.02280 by the authors.

Figure 1
Figure 1. Languages across SLMs. All monolingual foundational Small Language Models included in the study such as T5, BERT, RoBERTa do not support African languages di￾rectly, although some have been used as base mod￾els for further adaptations. Among the 38 African languages analyzed, only four Afrikaans, Amharic, Swahili, and Malagasy are fully supported by all multilingual SLMs.Appendix B and [PITH_FULL_IMAGE:figures/full… view at source ↗
Figure 2
Figure 2. Different NLP tasks and the number of datasets prepared for the task Question Five. Exploring the prospective of LLMs for Africa. Answer. Currently, artificial intelligence shows emergent ability which are arithmetic reasoning, agentic behaviour, common sense reasoning and symbolic reasoning. The path for this destination is very clear. Constructing LLMs for African lan￾guages is both challenge and an integral oppor… view at source ↗
Figure 3
Figure 3. Figure 3: Roadmap for African Language-Model Development. The diagram proceeds bottom [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Model parameter size versus number of training tokens. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AfroScope: A Framework for Studying the Linguistic Landscape of Africa

    cs.CL 2026-01 conditional novelty 6.0 of 10

    A new framework combines a 713-language African LID dataset, strong baselines, and a contrastive-embedding hierarchical step that improves macro-F1 by 4.55 on a 29-language confusable subset.

Reference graph

Works this paper leans on

48 extracted references · 20 canonical work pages · cited by 1 Pith paper

  1. [1]

    Masakhaner: Named entity recognition for african languages

    David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D’souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, et al. Masakhaner: Named entity recognition for african languages. Transactions of the Association for Computational Linguistics, 9: 0 1116--1131, 2021

  2. [2]

    Alabi, David Ifeoluwa Adelani, Marius Mosbach, and Dietrich Klakow

    Jesujoba O. Alabi, David Ifeoluwa Adelani, Marius Mosbach, and Dietrich Klakow. Adapting pre-trained language models to A frican languages via multilingual adaptive fine-tuning. In Nicoletta Calzolari, Chu-Ren Huang, Hansaem Kim, James Pustejovsky, Leo Wanner, Key-Sun Choi, Pum-Mo Ryu, Hsin-Hsi Chen, Lucia Donatelli, Heng Ji, Sadao Kurohashi, Patrizia Pag...

  3. [3]

    A ra BERT : Transformer-based model for A rabic language understanding

    Wissam Antoun, Fady Baly, and Hazem Hajj. A ra BERT : Transformer-based model for A rabic language understanding. In Hend Al-Khalifa, Walid Magdy, Kareem Darwish, Tamer Elsayed, and Hamdy Mubarak, editors, Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection, pages 9--15, Ma...

  4. [4]

    A ra GPT 2: Pre-trained transformer for A rabic language generation

    Wissam Antoun, Fady Baly, and Hazem Hajj. A ra GPT 2: Pre-trained transformer for A rabic language generation. In Nizar Habash, Houda Bouamor, Hazem Hajj, Walid Magdy, Wajdi Zaghouani, Fethi Bougares, Nadi Tomeh, Ibrahim Abu Farha, and Samia Touileb, editors, Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 196--207, Kyiv, Ukrai...

  5. [5]

    The world's writing systems, 2024

    Atelier National de Recherche Typographique (ANRT) , Institut Designlabor Gutenberg (IDG) , and Script Encoding Initiative (SEI) . The world's writing systems, 2024. URL https://www.worldswritingsystems.org/. Accessed: 2025-05-30

  6. [6]

    Clark, Dan Garrette, Iulia Turc, and John Wieting

    Jonathan H. Clark, Dan Garrette, Iulia Turc, and John Wieting. Canine: Pre-training an efficient tokenization-free encoder for language representation. Transactions of the Association for Computational Linguistics, 10: 0 73--91, 2022. doi:10.1162/tacl_a_00448. URL https://aclanthology.org/2022.tacl-1.5/

  7. [7]

    Unsupervised cross-lingual representation learning at scale

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross-lingual representation learning at scale. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault, editors, Proceedings of the 58th Annual Meeting of the ...

  8. [8]

    Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Yan He, Elahe Kalbassi, Linlu Wang, Long Wang, Shuo Zhang, Angela Fan, et al

    Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Yan He, Elahe Kalbassi, Linlu Wang, Long Wang, Shuo Zhang, Angela Fan, et al. No language left behind: Scaling human-centered machine translation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), Abu Dhabi, UAE, 2022. Associati...

Show all 48 references
  1. [9]

    I ndic BART : A pre-trained model for indic natural language generation

    Raj Dabre, Himani Shrotriya, Anoop Kunchukuttan, Ratish Puduppully, Mitesh Khapra, and Pratyush Kumar. I ndic BART : A pre-trained model for indic natural language generation. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Findings of the Association for...

  2. [10]

    Dai, Jonathan H

    Andrew M. Dai, Jonathan H. Clark, Kevin Robinson, Maysam Moussalem, Sebastian Ruder, Siamak Shakeri, Jacob Austin, et al. Palm 2 technical report. Technical report, Google, 2023

  3. [11]

    Deepseek-v3 technical report

    DeepSeek-AI. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, December 2024

  4. [12]

    BERT : Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors, Proceedings of the 2019 Conference of the North A merican Chapter ...

  5. [13]

    Bonaventure F. P. Dossou, Atnafu Lambebo Tonja, Oreen Yousuf, Salomey Osei, Abigail Oppong, Iyanuoluwa Shode, Oluwabusayo Olufunke Awoyomi, and Chris Emezue. A fro LM : A self-active learning-based multilingual pretrained language model for 23 A frican languages. In Angela Fan...

  6. [14]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zha...

  7. [15]

    Ethnologue: Languages of the world, 2025

    Ethnologue . Ethnologue: Languages of the world, 2025. Accessed: 2025-05-20

  8. [16]

    Language-agnostic BERT sentence embedding

    Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. Language-agnostic BERT sentence embedding. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics ...

  9. [17]

    Guerreiro, Duarte M

    Nuno M. Guerreiro, Duarte M. Alves, Jonas Waldendorf, Barry Haddow, Alexandra Birch, Pierre Colombo, and Andr \'e F. T. Martins. Hallucinations in large multilingual translation models. Transactions of the Association for Computational Linguistics, 11: 0 1500--1517, 2023. doi:...

  10. [18]

    Training compute-optimal large language models

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022

  11. [19]

    Glot500: Scaling multilingual corpora and language models to 500 languages

    Ayyoob Imani, Peiqin Lin, Amir Hossein Kargaran, Silvia Severini, Masoud Jalili Sabet, Nora Kassner, Chunlan Ma, Helmut Schmid, Andr \'e Martins, Fran c ois Yvon, and Hinrich Sch \"u tze. Glot500: Scaling multilingual corpora and language models to 500 languages. In Anna Roger...

  12. [20]

    A fri T e VA : Extending ?small data? pretraining approaches to sequence-to-sequence models

    Odunayo Jude Ogundepo, Akintunde Oladipo, Mofetoluwa Adeyemi, Kelechi Ogueji, and Jimmy Lin. A fri T e VA : Extending ?small data? pretraining approaches to sequence-to-sequence models. In Colin Cherry, Angela Fan, George Foster, Gholamreza (Reza) Haffari, Shahram Khadivi, Nan...

  13. [21]

    I ndo LEM and I ndo BERT : A benchmark dataset and pre-trained language model for I ndonesian NLP

    Fajri Koto, Afshin Rahimi, Jey Han Lau, and Timothy Baldwin. I ndo LEM and I ndo BERT : A benchmark dataset and pre-trained language model for I ndonesian NLP . In Donia Scott, Nuria Bel, and Chengqing Zong, editors, Proceedings of the 28th International Conference on Computat...

  14. [22]

    Cross-lingual language model pretraining

    Guillaume Lample and Alexis Conneau. Cross-lingual language model pretraining. Advances in Neural Information Processing Systems (NeurIPS), 32, 2019

  15. [23]

    Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension

    Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461, 2019

  16. [24]

    Roberta: A robustly optimized bert pretraining approach

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019

  17. [25]

    Scaling laws for fact memorization of large language models

    Xingyu Lu, Xiaonan Li, Qinyuan Cheng, Kai Ding, Xuanjing Huang, and Xipeng Qiu. Scaling laws for fact memorization of large language models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, pa...

  18. [26]

    F in GPT : Large generative models for a small language

    Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, Thomas Wang, Nouamane Tazi, Teven Scao, Thomas Wolf, Osma Suominen, Samuli Sairanen, Mikko Merioksa, Jyrki ...

  19. [27]

    A ra T 5: Text-to-text transformers for A rabic language generation

    El Moatez Billah Nagoudi, AbdelRahim Elmadany, and Muhammad Abdul-Mageed. A ra T 5: Text-to-text transformers for A rabic language generation. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Co...

  20. [28]

    Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced african languages

    Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced african languages. In Proceedings of the First Workshop on Natural Language Processing for African Languages (AfriNLP 2021), p...

  21. [29]

    Gpt-4 technical report, 2023

    OpenAI. Gpt-4 technical report, 2023. Technical Report

  22. [30]

    Exploring the limits of transfer learning with a unified text-to-text transformer

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21 0 (140): 0 1--67, 2020

  23. [31]

    Bert multilingual model readme, 2019

    Google Research. Bert multilingual model readme, 2019. Accessed: [Insert Date]

  24. [32]

    Multilingual t5: Massively multilingual pre-trained text-to-text transformer, 2021

    Google Research. Multilingual t5: Massively multilingual pre-trained text-to-text transformer, 2021. Accessed:

  25. [33]

    NLLB-200 : Distilled 600m-parameter model of No Language Left Behind ( NLLB ), 2022

    Meta AI Research. NLLB-200 : Distilled 600m-parameter model of No Language Left Behind ( NLLB ), 2022. Accessed:

  26. [34]

    How good is your tokenizer? on the monolingual performance of multilingual language models

    Phillip Rust, Jonas Pfeiffer, Ivan Vuli \'c , Sebastian Ruder, and Iryna Gurevych. How good is your tokenizer? on the monolingual performance of multilingual language models. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors, Proceedings of the 59th Annual Me...

  27. [35]

    Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter

    Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter. arXiv preprint arXiv:1910.01108, 2019

  28. [36]

    The language barrier: Dissecting safety challenges of LLM s in multilingual contexts

    Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, and Daniel Khashabi. The language barrier: Dissecting safety challenges of LLM s in multilingual contexts. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findi...

  29. [37]

    Gemini: A family of highly capable multimodal models, December 2023

    Gemini Team and Google DeepMind. Gemini: A family of highly capable multimodal models, December 2023. Technical Report

  30. [38]

    Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al

    NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. No language left behind: Scaling human-centered machine translation. In Proceedings of the 2022 Conferenc...

  31. [39]

    Ethiollm: Multilingual large language models for ethiopian languages with task evaluation

    Atnafu Lambebo Tonja, Israel Abebe Azime, Tadesse Destaw Belay, Mesay Gemeda Yigezu, Moges Ahmed Mehamed, Abinew Ali Ayele, Ebrahim Chekol Jibril, Michael Melese Woldeyohannis, Olga Kolesnikova, Philipp Slusallek, et al. Ethiollm: Multilingual large language models for ethiopi...

  32. [40]

    E thio LLM : Multilingual large language models for E thiopian languages with task evaluation

    Atnafu Lambebo Tonja, Israel Abebe Azime, Tadesse Destaw Belay, Mesay Gemeda Yigezu, Moges Ahmed Ah Mehamed, Abinew Ali Ayele, Ebrahim Chekol Jibril, Michael Melese Woldeyohannis, Olga Kolesnikova, Philipp Slusallek, Dietrich Klakow, and Seid Muhie Yimam. E thio LLM : Multilin...

  33. [41]

    E thio MT : Parallel corpus for low-resource E thiopian languages

    Atnafu Lambebo Tonja, Olga Kolesnikova, Alexander Gelbukh, and Jugal Kalita. E thio MT : Parallel corpus for low-resource E thiopian languages. In Rooweither Mabuya, Muzi Matfunjwa, Mmasibidi Setaka, and Menno van Zaanen, editors, Proceedings of the Fifth Workshop on Resources...

  34. [42]

    Aya model: An instruction finetuned open-access multilingual language model

    Ahmet \"U st \"u n, Viraat Aryabumi, Zheng Yong, Wei-Yin Ko, Daniel D ' souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker. Aya mode...

  35. [43]

    A comprehensive survey of small language models in the era of large language models: Techniques, enhancements, applications, collaboration with llms, and trustworthiness

    Fali Wang, Zhiwei Zhang, Xianren Zhang, Zongyu Wu, TzuHao Mo, Qiuhao Lu, Wanjing Wang, Rui Li, Junjie Xu, Xianfeng Tang, Qi He, Yao Ma, Ming Huang, and Suhang Wang. A comprehensive survey of small language models in the era of large language models: Techniques, enhancements, a...

  36. [44]

    Chain-of-thought prompting elicits reasoning in large language models

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35: 0 24824--24837, 2022

  37. [45]

    Bloom: A 176b-parameter open-access multilingual language model

    BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Saso Luccioni, François Yvon, et al. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint, arXiv:2211.05100, 2022...

  38. [46]

    Shijie Wu and Mark Dredze. Are all languages created equal in multilingual BERT ? In Spandana Gella, Johannes Welbl, Marek Rei, Fabio Petroni, Patrick Lewis, Emma Strubell, Minjoon Seo, and Hannaneh Hajishirzi, editors, Proceedings of the 5th Workshop on Representation Learnin...

  39. [47]

    m T 5: A massively multilingual pre-trained text-to-text transformer

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. m T 5: A massively multilingual pre-trained text-to-text transformer. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, St...

  40. [48]

    B y T 5: Towards a token-free future with pre-trained byte-to-byte models

    Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, and Colin Raffel. B y T 5: Towards a token-free future with pre-trained byte-to-byte models. Transactions of the Association for Computational Linguistics, 10: 0 291--306, 2022. do...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.