Pith. sign in

REVIEW 4 major objections 3 minor 51 references

Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments

T0 review · 4 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Newly translated Winogrande and clinical MMLU benchmarks in eight African languages reveal a 12.0–19.9 percentage-point performance gap between English and the average African language for the best LLM (GPT-4o), and fine-tuning on the…

desk verdict A genuinely useful benchmark and fine-tuning resource for eight low-resource African languages, with translation-quality caveats that need to be stated more carefully. read the letter →

arxiv 2412.12417 v1 pith:CFTZHTAJ submitted 2024-12-16 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords LLMperformancegaplow-resourceAfricanlanguagesbenchmarktranslationWinograndeMMLUclinicalsectionsfine-tuningcross-lingualtransferculturalappropriateness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper attempts to make LLM performance measurable in eight low-resource African languages by creating about one million human-translated words of established reasoning and medical benchmarks, and then uses those benchmarks to quantify how much worse state-of-the-art models perform compared with English. The headline finding is that the best model, GPT-4o, shows a performance gap of 12.0 to 19.9 percentage points between English and the average of the 11 African languages covered. The paper also argues that the gap is partly closable: fine-tuning on the translated data yields average mono-lingual gains of 5.6%, high-quality data beats low-quality data by 5.4%, cross-lingual transfer adds 2.9%, and GPT-4o scores about 3.0% higher out-of-the-box on questions that native speakers judge culturally appropriate. If true, the work provides reusable evaluation resources and evidence for data-centric strategies to reduce language inequity in LLMs.

What carries the argument

The central object is the newly released parallel benchmark: Winogrande and three clinical sections of MMLU (college medicine, clinical knowledge, virology) translated into eight languages — Amharic, Bambara, Igbo, Sepedi, Shona, Sesotho, Setswana, Tsonga — and used together with existing translations for Afrikaans, Zulu, and Xhosa. These translated QA sets carry the argument because they make possible the first direct English-versus-African-language comparison on the same multiple-choice questions; the fine-tuning machinery (mono- vs cross-lingual data, GPT-4o quality scoring, and data-volume sampling) then tests whether the gap can be reduced and which data characteristics matter.

What would settle it

Take a random sample of 100 MMLU clinical-knowledge questions, have two independent professional medical translators translate them into Bambara and Igbo, then have a third expert resolve disagreements; if GPT-4o's accuracy on these adjudicated translations exceeds that on the paper's translations by more than a few points, translation quality is a substantial confound of the reported gap.

Watch

Extended reading notes

Core claim

The central discovery, on the paper's own terms, is that the performance gap between LLMs in English and in native African languages is large and measurable, not a side effect of missing benchmarks: GPT-4o drops 12.0% to 19.9% absolute in accuracy when evaluated on human-translated MMLU clinical sections and Winogrande in 11 African languages, with the largest single-language deficits for Bambara. The paper's second discovery is that fine-tuning on modest amounts of translated data narrows that gap, with average mono-lingual improvements of 5.6%, and that data quality and cultural suitability matter: using LLM-selected high-quality training samples improves performance by 5.4% over low-quality samples, and GPT-4o performs 3.0% better out-of-the-box on questions labeled culturally appropriate by native speakers.

Load-bearing premise

The central claim depends on the translated benchmarks being accurate measures of meaning; the MMLU medical translations were never reviewed by human medical experts, and in Bambara up to 79% of MMLU rows are near-duplicates of automatic machine-translation output, so systematic translation errors could be inflating (or masking) part of the 12–20 point gap.

Editorial extensions

If this is right

  • If the gap is real, state-of-the-art LLMs are substantially less trustworthy in these languages on medical and reasoning questions, which matters for any attempt to deploy AI-assisted health information for the over 160 million speakers covered.
  • Fine-tuning on hundreds of translated examples can recover a meaningful part of the gap (5.6% average mono-lingual gain), so collecting small, high-quality translated datasets is a viable improvement path.
  • Aligning fine-tuning domain with target domain gives the largest boosts (up to 17.4% for college-medicine tuning evaluated on clinical knowledge), so resource creators should prioritize domain-matched data.
  • Quality filtering with an LLM annotator is worth about 5.4% over unfiltered data, meaning data selection can substitute for larger data volume in low-resource settings.
  • Questions that native speakers judge culturally appropriate are easier for GPT-4o (3.0% boost), so cultural adaptation of evaluation content affects measured capability.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because model performance tracks language family and resource level (Afrikaans outscores Bambara by 22.9–56.1 points on every benchmark), the dominant driver of the gap is likely pre-training data exposure rather than task difficulty, which argues for directing translation and fine-tuning resources to Mande and Semitic family languages first.
  • The paper's machine-translation comparisons show that for several models, machine-translated queries can match or beat human-translated ones on MMLU, hinting that models have learned a 'translationese' distribution; a direct test would be to fine-tune the same model separately on human vs machine translations and compare generalization to new human text.
  • Because the cultural-appropriateness annotations show very low inter-annotator agreement (the highest kappa value is 0.211), the reported 3.0% cultural boost is likely an underestimate of the true effect; a replication using more annotators per item or a culturally grounded taxonomy could reveal a larger and more consistent effect.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. The paper introduces human-translated versions of Winogrande and three clinical MMLU sections into eight low-resource African languages (plus three prior languages), evaluates a range of LLMs on these benchmarks, and reports a consistent 12.0%-19.9% absolute performance gap between English and the average of eleven African languages. It then fine-tunes Llama 3 70B over 400 configurations, reporting mono-lingual gains of 5.6%, cross-lingual gains of 2.9%, and a 3.0% out-of-the-box lift on culturally appropriate Winogrande items. The benchmarks, translations, and code are released publicly.

Significance. If the measurements are valid, this is a useful contribution: it expands reasoning-benchmark coverage to eight underrepresented languages, provides a large public translation resource, and gives practical evidence about fine-tuning with limited data. The paper is unusually transparent about reproducibility (seeds, prompts, hyperparameters, and full appendix tables) and about its own limitations, including the absence of statistical tests and variability in translation quality. However, the headline gap and the fine-tuning conclusions rest on translation fidelity and annotation reliability, both of which are not yet adequately established.

major comments (4)
  1. [Appendix Section A and Tables A.12-A.13] The MMLU translations received no human quality review; the only fidelity check is ROUGE-1 similarity to Google Translate, which is extreme and bimodal (up to 79% of Bambara rows have ROUGE-1 > 0.95, while up to 60% of Sesotho rows have ROUGE-1 < 0.5). ROUGE-1 similarity to a machine translation cannot distinguish a correct translation from one that merely resembles MT output, so the medical knowledge scores in Table 2 may be depressed by systematic translation errors. This is not hypothetical: Tables A.12-A.13 show that GPT-4o on Igbo college medicine scores 58.4 on the human translation but 68.2 on the machine translation, a 9.8-point gap in the opposite direction of the headline claim. Because these MMLU scores feed directly into the 12.0-19.9% gap and the fine-tuning experiments, the authors need to add a human validation sample of the MMLU translations (or an equivalent diagnostic) before the central claim can be taken at face value.
  2. [Section C and Table A.16] The Winogrande quality and cultural-appropriateness annotations have very low inter-annotator agreement: the maximum Cohen's kappa for translation quality is 0.211 (Shona), and for appropriateness it is near zero for all languages (e.g., 0.074 for Bambara). The paper nonetheless uses these labels to define "good/understandable" translations and to split data into "culturally appropriate" vs "culturally inappropriate" subsets, from which it concludes a 3.0% average performance boost. With near-zero agreement, the appropriateness split is largely noise, and the measured lift could reflect translation quality or other confounds rather than a cultural construct. The authors acknowledge the low kappa but do not provide an adjudicated or consensus-based set of labels to show that the split is reliable. Please report results on an adjudicated subset or demonstrate that the effect is robust to alternative annotation aggregation schemes.
  3. [Results, Tables 1 and 2] The headline performance gap is reported as point estimates without confidence intervals or significance tests. For MMLU virology, the test set has only 166 questions per language; a 12 percentage-point gap on that sample has a standard error of roughly 3 points, so the lower end of the claimed 12.0-19.9% range is not statistically distinguishable from chance-level differences for some language-benchmark combinations. The reproducibility checklist explicitly states that no statistical tests were used. This does not invalidate the direction of the result, but the central quantitative claim should be accompanied by uncertainty estimates (e.g., bootstrap confidence intervals) or at least by the per-benchmark sample sizes needed to assess the precision of the gap.
  4. [Fine-tuning with Varying Data Quality and Quantity] The high/low-quality split is based on GPT-4o LLM-as-an-Annotator scores, with no human validation that these scores correlate with actual fine-tuning usefulness. The 5.4% average advantage of the high-quality tertile could be driven by surface features (e.g., question length, topic balance, or translationese patterns) rather than by the construct of "quality." Since the abstract and discussion present high-quality dataset fine-tuning as a key finding, the authors should either validate the GPT-4o ratings against a human-rated sample or show that the effect persists after controlling for simple confounds such as length and lexical diversity.
minor comments (3)
  1. [Fine-tuning with Varying Languages and Domains] There is a typo in the phrase "Lllama 3 70B" (extra 'l'); please correct it.
  2. [Figure A.14] The color bar for Cohen's kappa ranges from 0.0 to 1.0, but several cells contain negative kappa values; this makes the visualization misleading. Please extend the color scale to negative values or recode the affected cells explicitly.
  3. [Abstract and Methods] The abstract states "approximately 1 million human-translated words," but summing the per-language counts in the Methods (73,742 words for Winogrande plus 27,107 words for the three MMLU sections, times 8 languages) gives approximately 807,000 words, which is closer to 0.8 million. Please either recompute the total or revise the wording to be accurate.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the benchmark measurements and fine-tuning results are empirical and not forced by construction.

full rationale

The paper's derivation chain is empirical rather than definitional. The translated benchmarks are new human translations, with Afrikaans/Zulu/Xhosa transparently sourced from the authors' prior BMGF dataset and labeled as such; the headline 12.0-19.9 percentage-point gap is an out-of-the-box accuracy difference on those translated prompts, not an identity. Fine-tuning gains are computed on held-out benchmarks (e.g., tuning on Winogrande train or MMLU college medicine, then evaluating on Winogrande test, clinical knowledge, virology, and Belebele), so the reported 5.6% mono-lingual gain is not a refit of training data. The high-versus-low quality comparison uses GPT-4o-generated tertile scores to select fine-tuning data for a different model (Llama 3 70B), making it an empirical test rather than a construction. The cultural-appropriateness lifts are compared against an English baseline on the same splits precisely to avoid defining the effect into existence. The lack of human quality review for MMLU translations and the low inter-annotator agreement on Winogrande appropriateness are validity threats to the measured gap, but they are not circularity: nothing in the paper forces the reported numbers to equal its inputs.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The paper's empirical claims rest on several background assumptions: that translated multiple-choice items measure the same construct in each language; that MMLU translations, which were not human-validated, are accurate enough to measure medical knowledge; and that GPT-4o's quality scores identify genuinely better fine-tuning data. The central quantities are not derived from equations, so no fitted parameters are required beyond the data-dependent tertile split. The 'first known translations' claim also depends on an unstated assumption that IrokoBench does not already cover the same benchmarks in these languages.

free parameters (1)
  • High/low quality tertile split = Top and bottom thirds of GPT-4o quality scores per dataset
    The definition of high-quality vs low-quality fine-tuning data is determined by ordering each dataset by GPT-4o quality scores and taking the top and bottom thirds. The headline 5.4% gain from high- over low-quality data depends on this split and is not derived from an independent quality standard.
assumptions (4)
  • domain assumption Translated multiple-choice items preserve the reasoning construct of the original English benchmarks
    All accuracy numbers in Table 1, Table 2, and the fine-tuning results are computed from accuracy on translated items. If translation changes the meaning or difficulty, the measured gap and gains reflect artifacts. This assumption enters in Aim 1 and is relied upon throughout.
  • domain assumption Unrated languages in the Joshi et al. (2020) resource taxonomy are low-resource
    The Introduction's footnote 1 states this assumption to support the claim that all native African languages are low-resource. It frames the paper's motivation but does not enter the measurements.
  • ad hoc to paper GPT-4o quality scores correlate with fine-tuning usefulness
    In Aim 3, the high-quality set is defined by GPT-4o LLM-as-an-Annotator scores. The claim that high-quality data improves fine-tuning more than low-quality data depends on this assumption, since no human quality labels are used for the fine-tuning data.
  • ad hoc to paper MMLU translations are accurate enough for medical knowledge evaluation
    The paper reports no human quality evaluation for the Translated.com MMLU translations, only ROUGE-1 similarity to machine translation. The clinical knowledge gap measurements in Table 1 and Table 2 assume these translations are valid. This is flagged because up to 79% of Bambara MMLU rows had ROUGE-1 >0.95 against Google Translate.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments." pith.science (2026). https://pith.science/paper/CFTZHTAJ

@misc{pith2026241212417,
  author       = {Pith},
  title        = {Pith review of: Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CFTZHTAJ}},
  note         = {Machine review of arXiv:2412.12417}
}
read the original abstract

Large Language Models (LLMs) have shown remarkable performance across various tasks, yet significant disparities remain for non-English languages, and especially native African languages. This paper addresses these disparities by creating approximately 1 million human-translated words of new benchmark data in 8 low-resource African languages, covering a population of over 160 million speakers of: Amharic, Bambara, Igbo, Sepedi (Northern Sotho), Shona, Sesotho (Southern Sotho), Setswana, and Tsonga. Our benchmarks are translations of Winogrande and three sections of MMLU: college medicine, clinical knowledge, and virology. Using the translated benchmarks, we report previously unknown performance gaps between state-of-the-art (SOTA) LLMs in English and African languages. Finally, using results from over 400 fine-tuned models, we explore several methods to reduce the LLM performance gap, including high-quality dataset fine-tuning (using an LLM-as-an-Annotator), cross-lingual transfer, and cultural appropriateness adjustments. Key findings include average mono-lingual improvements of 5.6% with fine-tuning (with 5.4% average mono-lingual improvements when using high-quality data over low-quality data), 2.9% average gains from cross-lingual transfer, and a 3.0% out-of-the-box performance boost on culturally appropriate questions. The publicly available benchmarks, translations, and code from this study support further research and development aimed at creating more inclusive and effective language technologies.

Figures

Figures reproduced from arXiv: 2412.12417 by the authors.

Figure 2
Figure 2. Mono- and Cross-lingual LLM Performance Gains. The figure displays boxplots of performance gains when fine-tuning with either the translated Winogrande train set (left) or MMLU college medicine section (right). The fine-tuned models were evaluated across 4 datasests (x-axis) for mono-lingual gains (blue) across 11 African languages, and cross-lingual gains (green) across 110 African lan￾guage pairs. The most signifi… view at source ↗
Figure 1
Figure 1. GPT-4o Winogrande Performance on “appro￾priate” vs. “inappropriate” Data. GPT-4o was evaluated on Winogrande (test set) out-of-the-box in each target lan￾guage and in English. Top plot: the absolute performance on QA pairs considered culturally “appropriate” and “inap￾propriate” according to native speakers. Bottom plot: per￾formance lifts for each language (green) and in English (grey), using the same annotations. … view at source ↗
Figure 3
Figure 3. LLM Performance Across Quality and Quan￾tity Combinations. The figure displays LLM performance when fine-tuning by data quality and quantity, using MMLU college medicine and evaluating on MMLU clinical knowl￾edge (which had the greatest mono-lingual gains from Fig￾ure 2). The quality of samples was rated using GPT-4o LLM￾as-an-Annotator scores. The lowest tertile and highest ter￾tile were defined as low (yellow) and… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

51 extracted references · 33 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    A.; Awan, A

    Abdin, M.; Jacobs, S. A.; Awan, A. A.; Aneja, J.; Awadallah, A.; Awadalla, H.; Bach, N.; Bahree, A.; Bakhtiari, A.; Behl, H.; et al. 2024. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219

  4. [4]

    L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al

    Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  5. [5]

    Adebara, I.; Elmadany, A.; Abdul-Mageed, M.; and Inciarte, A. 2022. AfroLID: A Neural Language Identification Tool for African Languages. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 1958--1981

  6. [6]

    I.; Ojo, J.; Azime, I

    Adelani, D. I.; Ojo, J.; Azime, I. A.; Zhuang, J. Y.; Alabi, J. O.; He, X.; Ochieng, M.; Hooker, S.; Bukula, A.; Lee, E.-S. A.; et al. 2024. IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models. arXiv preprint arXiv:2406.03368

  7. [7]

    Arora, A.; Kaffee, L.-a.; and Augenstein, I. 2023. Probing Pre-Trained Language Models for Cross-Cultural Differences in Values. In Dev, S.; Prabhakaran, V.; Adelani, D.; Hovy, D.; and Benotti, L., eds., Proceedings of the First Workshop on Cross-Cultural Considerations in NLP (C3NLP), 114--130. Dubrovnik, Croatia: Association for Computational Linguistics

  8. [8]

    Aryabumi, V.; Dang, J.; Talupuru, D.; Dash, S.; Cairuz, D.; Lin, H.; Venkitesh, B.; Smith, M.; Marchisio, K.; Ruder, S.; et al. 2024. Aya 23: Open weight releases to further multilingual progress. arXiv preprint arXiv:2405.15032

Show all 51 references
  1. [9]

    N.; Husa, D.; Goyal, N.; Krishnan, A.; Zettlemoyer, L.; and Khabsa, M

    Bandarkar, L.; Liang, D.; Muller, B.; Artetxe, M.; Shukla, S. N.; Husa, D.; Goyal, N.; Krishnan, A.; Zettlemoyer, L.; and Khabsa, M. 2024. The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants. In Proceedings of the 62nd Annual Meeting of th...

  2. [10]

    Bapna, A.; Caswell, I.; Kreutzer, J.; Firat, O.; van Esch, D.; Siddhant, A.; Niu, M.; Baljekar, P.; Garcia, X.; Macherey, W.; et al. 2022. Building machine translation systems for the next thousand languages. arXiv preprint arXiv:2205.03983

  3. [11]

    Benjamin, M. 2019. Empirical Evaluation of Google Translate across 107 Languages. https://www.teachyoubackwards.com/empirical-evaluation/. Accessed: 2024-07-25

  4. [12]

    Beukman, M.; and Fokam, M. 2023. Analysing Cross-Lingual Transfer in Low-Resourced African Named Entity Recognition. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association f...

  5. [13]

    BMGF. 2024. Expanding Reasoning Benchmarks in Low-Resourced African Languages: Winogrande and Clinical MMLU in Afrikaans, Xhosa, and Zulu. https://github.com/InstituteforDiseaseModeling/winogrande-mmlu-clinical-za

  6. [14]

    Borin, L.; Comrie, B.; and Saxena, A. 2013. The Intercontinental Dictionary Series: A rich and principled database for language comparison

  7. [15]

    Chen, J.; and Mueller, J. 2024. Automated data curation for robust language model fine-tuning. arXiv preprint arXiv:2403.12776

  8. [16]

    CIA. 2024. The World Factbook — World. https://www.cia.gov/the-world-factbook/countries/world/#people-and-society. Accessed: 2024-07-29

  9. [17]

    Cole, J.; Zhang, M.; Gillick, D.; Eisenschlos, J.; Dhingra, B.; and Eisenstein, J. 2023. Selectively Answering Ambiguous Questions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 530--543

  10. [18]

    Conneau, A.; Rinott, R.; Lample, G.; Williams, A.; Bowman, S.; Schwenk, H.; and Stoyanov, V. 2018. XNLI: Evaluating Cross-lingual Sentence Representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2475. Association for Computat...

  11. [19]

    De Melo, G.; Imaizumi, V.; and Cozman, F. 2019. Winograd schemas in portuguese. In Anais do XVI Encontro Nacional de Intelig \^e ncia Artificial e Computacional , 787--798. SBC

  12. [20]

    Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024. The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783

  13. [21]

    Galal, S. 2024. Africa: Share of global poverty by country 2024. https://www.statista.com/statistics/1228553/extreme-poverty-as-share-of-global-population-in-africa-by-country. Accessed: 2024-07-29

  14. [22]

    Hammerstrom, H. 2015. Ethnologue 16/17/18th editions: A comprehensive review . LANGUAGE, 91(3)

  15. [23]

    Hartvigsen, T.; Gabriel, S.; Palangi, H.; Sap, M.; Ray, D.; and Kamar, E. 2022. ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volu...

  16. [24]

    Hendrycks, D.; Burns, C.; Basart, S.; Critch, A.; Li, J.; Song, D.; and Steinhardt, J. 2021 a . Aligning AI With Shared Human Values. Proceedings of the International Conference on Learning Representations (ICLR)

  17. [25]

    Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021 b . Measuring Massive Multitask Language Understanding. Proceedings of the International Conference on Learning Representations (ICLR)

  18. [26]

    Hu, J.; Ruder, S.; Siddhant, A.; Neubig, G.; Firat, O.; and Johnson, M. 2020. Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation. In International Conference on Machine Learning, 4411--4421. PMLR

  19. [27]

    ITU. 2023. Facts and Figures 2023. https://www.itu.int/itu-d/reports/statistics/facts-figures-2023/. Accessed: 2024-07-29

  20. [28]

    Joshi, P.; Santy, S.; Budhiraja, A.; Bali, K.; and Choudhury, M. 2020. The State and Fate of Linguistic Diversity and Inclusion in the NLP World. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 6282. Association for Computational Linguistics

  21. [29]

    Liang, Y.; Duan, N.; Gong, Y.; Wu, N.; Guo, F.; Qi, W.; Gong, M.; Shou, L.; Jiang, D.; Cao, G.; et al. 2020. XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Langu...

  22. [30]

    V.; Mihaylov, T.; Artetxe, M.; Wang, T.; Chen, S.; Simig, D.; Ott, M.; Goyal, N.; Bhosale, S.; Du, J.; et al

    Lin, X. V.; Mihaylov, T.; Artetxe, M.; Wang, T.; Chen, S.; Simig, D.; Ott, M.; Goyal, N.; Bhosale, S.; Du, J.; et al. 2021. Few-shot learning with multilingual language models. arXiv preprint arXiv:2112.10668

  23. [31]

    M.; Reddy, S.; Collier, N.; and Elliott, D

    Liu, F.; Bugliarello, E.; Ponti, E. M.; Reddy, S.; Collier, N.; and Elliott, D. 2021. Visually Grounded Reasoning across Languages and Cultures. arXiv:2109.13238

  24. [32]

    Longpre, S.; Yauney, G.; Reif, E.; Lee, K.; Roberts, A.; Zoph, B.; Zhou, D.; Wei, J.; Robinson, K.; Mimno, D.; et al. 2024. A Pretrainer’s Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, and Toxicity. In Proceedings of the 2024 Conference o...

  25. [33]

    S.; Shen, S.; Yong, Z

    Muennighoff, N.; Wang, T.; Sutawika, L.; Roberts, A.; Biderman, S.; Le Scao, T.; Bari, M. S.; Shen, S.; Yong, Z. X.; Schoelkopf, H.; et al. 2023. Crosslingual Generalization through Multitask Finetuning. In Proceedings of the 61st Annual Meeting of the Association for Computat...

  26. [34]

    NLLB Team ; et al. 2024. Scaling neural machine translation to 200 languages. Nature, 630(8018): 841

  27. [35]

    OpenAI. 2023. GPT-3.5 Turbo fine-tuning and API updates. https://openai.com/index/gpt-3-5-turbo-fine-tuning-and-api-updates/. Accessed: 2024-07-26

  28. [36]

    OpenAI. 2024. Hello GPT-4o. https://openai.com/index/hello-gpt-4o/. Accessed: 2024-07-26

  29. [37]

    Polignano, M.; Basile, P.; and Semeraro, G. 2024. Advanced Natural-based interaction for the ITAlian language: LLaMAntino-3-ANITA. arXiv:2405.07101

  30. [38]

    A.; Haznitrama, F

    Putri, R. A.; Haznitrama, F. G.; Adhista, D.; and Oh, A. 2024. Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese. arXiv:2402.17302

  31. [39]

    L.; Bhagavatula, C.; and Choi, Y

    Sakaguchi, K.; Bras, R. L.; Bhagavatula, C.; and Choi, Y. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9): 99--106

  32. [40]

    A.; and Choi, Y

    Sap, M.; Gabriel, S.; Qin, L.; Jurafsky, D.; Smith, N. A.; and Choi, Y. 2020. Social Bias Frames: Reasoning about Social and Power Implications of Language. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5477--5490

  33. [41]

    Shaham, U.; Herzig, J.; Aharoni, R.; Szpektor, I.; Tsarfaty, R.; and Eyal, M. 2024. Multilingual instruction tuning with just a pinch of multilinguality. arXiv preprint arXiv:2401.01854

  34. [42]

    Statistics South Africa . 2020. Labour Market Dynamics in South Africa . Technical report, Statistics South Africa, Private Bag X44, Pretoria 0001

  35. [43]

    Tan, Z.; Beigi, A.; Wang, S.; Guo, R.; Bhattacharjee, A.; Jiang, B.; Karami, M.; Li, J.; Cheng, L.; and Liu, H. 2024. Large language models for data annotation: A survey. arXiv preprint arXiv:2402.13446

  36. [44]

    \"U st \"u n, A.; Aryabumi, V.; Yong, Z.-X.; Ko, W.-Y.; D'souza, D.; Onilude, G.; Bhandari, N.; Singh, S.; Ooi, H.-L.; Kayid, A.; et al. 2024. Aya model: An instruction finetuned open-access multilingual language model. arXiv preprint arXiv:2402.07827

  37. [45]

    Villalobos, P.; Ho, A.; Sevilla, J.; Besiroglu, T.; Heim, L.; and Hobbhahn, M. 2024. Position: Will we run out of data? Limits of LLM scaling based on human-generated data. In Forty-first International Conference on Machine Learning

  38. [46]

    Wen, J.; Ke, P.; Sun, H.; Zhang, Z.; Li, C.; Bai, J.; and Huang, M. 2023. Unveiling the Implicit Toxicity in Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1322--1338

  39. [47]

    Whitehouse, C.; Choudhury, M.; and Aji, A. 2023. LLM-powered Data Augmentation for Enhanced Cross-lingual Performance. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 671--686

  40. [48]

    Xu, J.; Ju, D.; Li, M.; Boureau, Y.-L.; Weston, J.; and Dinan, E. 2021. Bot-adversarial dialogue for safe conversational agents. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2...

  41. [49]

    Yao, B.; Jiang, M.; Yang, D.; and Hu, J. 2023. Benchmarking LLM-based Machine Translation on Cultural Awareness

  42. [50]

    P.; et al

    Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li13, D.; Xing35, E. P.; et al. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv preprint arXiv:2306.05685

  43. [51]

    Zhou, C.; Liu, P.; Xu, P.; Iyer, S.; Sun, J.; Mao, Y.; Ma, X.; Efrat, A.; Yu, P.; Yu, L.; et al. 2024. Lima: Less is more for alignment. Advances in Neural Information Processing Systems, 36

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.