REVIEW 4 major objections 3 minor 51 references
Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments
T0 review · 4 major / 3 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Newly translated Winogrande and clinical MMLU benchmarks in eight African languages reveal a 12.0–19.9 percentage-point performance gap between English and the average African language for the best LLM (GPT-4o), and fine-tuning on the…
desk verdict A genuinely useful benchmark and fine-tuning resource for eight low-resource African languages, with translation-quality caveats that need to be stated more carefully. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the newly released parallel benchmark: Winogrande and three clinical sections of MMLU (college medicine, clinical knowledge, virology) translated into eight languages — Amharic, Bambara, Igbo, Sepedi, Shona, Sesotho, Setswana, Tsonga — and used together with existing translations for Afrikaans, Zulu, and Xhosa. These translated QA sets carry the argument because they make possible the first direct English-versus-African-language comparison on the same multiple-choice questions; the fine-tuning machinery (mono- vs cross-lingual data, GPT-4o quality scoring, and data-volume sampling) then tests whether the gap can be reduced and which data characteristics matter.
What would settle it
Take a random sample of 100 MMLU clinical-knowledge questions, have two independent professional medical translators translate them into Bambara and Igbo, then have a third expert resolve disagreements; if GPT-4o's accuracy on these adjudicated translations exceeds that on the paper's translations by more than a few points, translation quality is a substantial confound of the reported gap.
Extended reading notes
Core claim
The central discovery, on the paper's own terms, is that the performance gap between LLMs in English and in native African languages is large and measurable, not a side effect of missing benchmarks: GPT-4o drops 12.0% to 19.9% absolute in accuracy when evaluated on human-translated MMLU clinical sections and Winogrande in 11 African languages, with the largest single-language deficits for Bambara. The paper's second discovery is that fine-tuning on modest amounts of translated data narrows that gap, with average mono-lingual improvements of 5.6%, and that data quality and cultural suitability matter: using LLM-selected high-quality training samples improves performance by 5.4% over low-quality samples, and GPT-4o performs 3.0% better out-of-the-box on questions labeled culturally appropriate by native speakers.
Load-bearing premise
The central claim depends on the translated benchmarks being accurate measures of meaning; the MMLU medical translations were never reviewed by human medical experts, and in Bambara up to 79% of MMLU rows are near-duplicates of automatic machine-translation output, so systematic translation errors could be inflating (or masking) part of the 12–20 point gap.
Editorial extensions
If this is right
- If the gap is real, state-of-the-art LLMs are substantially less trustworthy in these languages on medical and reasoning questions, which matters for any attempt to deploy AI-assisted health information for the over 160 million speakers covered.
- Fine-tuning on hundreds of translated examples can recover a meaningful part of the gap (5.6% average mono-lingual gain), so collecting small, high-quality translated datasets is a viable improvement path.
- Aligning fine-tuning domain with target domain gives the largest boosts (up to 17.4% for college-medicine tuning evaluated on clinical knowledge), so resource creators should prioritize domain-matched data.
- Quality filtering with an LLM annotator is worth about 5.4% over unfiltered data, meaning data selection can substitute for larger data volume in low-resource settings.
- Questions that native speakers judge culturally appropriate are easier for GPT-4o (3.0% boost), so cultural adaptation of evaluation content affects measured capability.
Reading between the lines
- Because model performance tracks language family and resource level (Afrikaans outscores Bambara by 22.9–56.1 points on every benchmark), the dominant driver of the gap is likely pre-training data exposure rather than task difficulty, which argues for directing translation and fine-tuning resources to Mande and Semitic family languages first.
- The paper's machine-translation comparisons show that for several models, machine-translated queries can match or beat human-translated ones on MMLU, hinting that models have learned a 'translationese' distribution; a direct test would be to fine-tune the same model separately on human vs machine translations and compare generalization to new human text.
- Because the cultural-appropriateness annotations show very low inter-annotator agreement (the highest kappa value is 0.211), the reported 3.0% cultural boost is likely an underestimate of the true effect; a replication using more annotators per item or a culturally grounded taxonomy could reveal a larger and more consistent effect.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces human-translated versions of Winogrande and three clinical MMLU sections into eight low-resource African languages (plus three prior languages), evaluates a range of LLMs on these benchmarks, and reports a consistent 12.0%-19.9% absolute performance gap between English and the average of eleven African languages. It then fine-tunes Llama 3 70B over 400 configurations, reporting mono-lingual gains of 5.6%, cross-lingual gains of 2.9%, and a 3.0% out-of-the-box lift on culturally appropriate Winogrande items. The benchmarks, translations, and code are released publicly.
Significance. If the measurements are valid, this is a useful contribution: it expands reasoning-benchmark coverage to eight underrepresented languages, provides a large public translation resource, and gives practical evidence about fine-tuning with limited data. The paper is unusually transparent about reproducibility (seeds, prompts, hyperparameters, and full appendix tables) and about its own limitations, including the absence of statistical tests and variability in translation quality. However, the headline gap and the fine-tuning conclusions rest on translation fidelity and annotation reliability, both of which are not yet adequately established.
major comments (4)
- [Appendix Section A and Tables A.12-A.13] The MMLU translations received no human quality review; the only fidelity check is ROUGE-1 similarity to Google Translate, which is extreme and bimodal (up to 79% of Bambara rows have ROUGE-1 > 0.95, while up to 60% of Sesotho rows have ROUGE-1 < 0.5). ROUGE-1 similarity to a machine translation cannot distinguish a correct translation from one that merely resembles MT output, so the medical knowledge scores in Table 2 may be depressed by systematic translation errors. This is not hypothetical: Tables A.12-A.13 show that GPT-4o on Igbo college medicine scores 58.4 on the human translation but 68.2 on the machine translation, a 9.8-point gap in the opposite direction of the headline claim. Because these MMLU scores feed directly into the 12.0-19.9% gap and the fine-tuning experiments, the authors need to add a human validation sample of the MMLU translations (or an equivalent diagnostic) before the central claim can be taken at face value.
- [Section C and Table A.16] The Winogrande quality and cultural-appropriateness annotations have very low inter-annotator agreement: the maximum Cohen's kappa for translation quality is 0.211 (Shona), and for appropriateness it is near zero for all languages (e.g., 0.074 for Bambara). The paper nonetheless uses these labels to define "good/understandable" translations and to split data into "culturally appropriate" vs "culturally inappropriate" subsets, from which it concludes a 3.0% average performance boost. With near-zero agreement, the appropriateness split is largely noise, and the measured lift could reflect translation quality or other confounds rather than a cultural construct. The authors acknowledge the low kappa but do not provide an adjudicated or consensus-based set of labels to show that the split is reliable. Please report results on an adjudicated subset or demonstrate that the effect is robust to alternative annotation aggregation schemes.
- [Results, Tables 1 and 2] The headline performance gap is reported as point estimates without confidence intervals or significance tests. For MMLU virology, the test set has only 166 questions per language; a 12 percentage-point gap on that sample has a standard error of roughly 3 points, so the lower end of the claimed 12.0-19.9% range is not statistically distinguishable from chance-level differences for some language-benchmark combinations. The reproducibility checklist explicitly states that no statistical tests were used. This does not invalidate the direction of the result, but the central quantitative claim should be accompanied by uncertainty estimates (e.g., bootstrap confidence intervals) or at least by the per-benchmark sample sizes needed to assess the precision of the gap.
- [Fine-tuning with Varying Data Quality and Quantity] The high/low-quality split is based on GPT-4o LLM-as-an-Annotator scores, with no human validation that these scores correlate with actual fine-tuning usefulness. The 5.4% average advantage of the high-quality tertile could be driven by surface features (e.g., question length, topic balance, or translationese patterns) rather than by the construct of "quality." Since the abstract and discussion present high-quality dataset fine-tuning as a key finding, the authors should either validate the GPT-4o ratings against a human-rated sample or show that the effect persists after controlling for simple confounds such as length and lexical diversity.
minor comments (3)
- [Fine-tuning with Varying Languages and Domains] There is a typo in the phrase "Lllama 3 70B" (extra 'l'); please correct it.
- [Figure A.14] The color bar for Cohen's kappa ranges from 0.0 to 1.0, but several cells contain negative kappa values; this makes the visualization misleading. Please extend the color scale to negative values or recode the affected cells explicitly.
- [Abstract and Methods] The abstract states "approximately 1 million human-translated words," but summing the per-language counts in the Methods (73,742 words for Winogrande plus 27,107 words for the three MMLU sections, times 8 languages) gives approximately 807,000 words, which is closer to 0.8 million. Please either recompute the total or revise the wording to be accurate.
Circularity Check
No significant circularity; the benchmark measurements and fine-tuning results are empirical and not forced by construction.
full rationale
The paper's derivation chain is empirical rather than definitional. The translated benchmarks are new human translations, with Afrikaans/Zulu/Xhosa transparently sourced from the authors' prior BMGF dataset and labeled as such; the headline 12.0-19.9 percentage-point gap is an out-of-the-box accuracy difference on those translated prompts, not an identity. Fine-tuning gains are computed on held-out benchmarks (e.g., tuning on Winogrande train or MMLU college medicine, then evaluating on Winogrande test, clinical knowledge, virology, and Belebele), so the reported 5.6% mono-lingual gain is not a refit of training data. The high-versus-low quality comparison uses GPT-4o-generated tertile scores to select fine-tuning data for a different model (Llama 3 70B), making it an empirical test rather than a construction. The cultural-appropriateness lifts are compared against an English baseline on the same splits precisely to avoid defining the effect into existence. The lack of human quality review for MMLU translations and the low inter-annotator agreement on Winogrande appropriateness are validity threats to the measured gap, but they are not circularity: nothing in the paper forces the reported numbers to equal its inputs.
Assumptions & free parameters
free parameters (1)
- High/low quality tertile split =
Top and bottom thirds of GPT-4o quality scores per dataset
assumptions (4)
- domain assumption Translated multiple-choice items preserve the reasoning construct of the original English benchmarks
- domain assumption Unrated languages in the Joshi et al. (2020) resource taxonomy are low-resource
- ad hoc to paper GPT-4o quality scores correlate with fine-tuning usefulness
- ad hoc to paper MMLU translations are accurate enough for medical knowledge evaluation
Cite this review
Pith. "Pith review of Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments." pith.science (2026). https://pith.science/paper/CFTZHTAJ
@misc{pith2026241212417,
author = {Pith},
title = {Pith review of: Bridging the Gap: Enhancing LLM Performance for Low-Resource African Languages with New Benchmarks, Fine-Tuning, and Cultural Adjustments},
year = {2026},
howpublished = {\url{https://pith.science/paper/CFTZHTAJ}},
note = {Machine review of arXiv:2412.12417}
}
read the original abstract
Large Language Models (LLMs) have shown remarkable performance across various tasks, yet significant disparities remain for non-English languages, and especially native African languages. This paper addresses these disparities by creating approximately 1 million human-translated words of new benchmark data in 8 low-resource African languages, covering a population of over 160 million speakers of: Amharic, Bambara, Igbo, Sepedi (Northern Sotho), Shona, Sesotho (Southern Sotho), Setswana, and Tsonga. Our benchmarks are translations of Winogrande and three sections of MMLU: college medicine, clinical knowledge, and virology. Using the translated benchmarks, we report previously unknown performance gaps between state-of-the-art (SOTA) LLMs in English and African languages. Finally, using results from over 400 fine-tuned models, we explore several methods to reduce the LLM performance gap, including high-quality dataset fine-tuning (using an LLM-as-an-Annotator), cross-lingual transfer, and cultural appropriateness adjustments. Key findings include average mono-lingual improvements of 5.6% with fine-tuning (with 5.4% average mono-lingual improvements when using high-quality data over low-quality data), 2.9% average gains from cross-lingual transfer, and a 3.0% out-of-the-box performance boost on culturally appropriate questions. The publicly available benchmarks, translations, and code from this study support further research and development aimed at creating more inclusive and effective language technologies.
Figures
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Abdin, M.; Jacobs, S. A.; Awan, A. A.; Aneja, J.; Awadallah, A.; Awadalla, H.; Bach, N.; Bahree, A.; Bakhtiari, A.; Behl, H.; et al. 2024. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219
arXiv 2024
-
[4]
L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
arXiv 2023
-
[5]
Adebara, I.; Elmadany, A.; Abdul-Mageed, M.; and Inciarte, A. 2022. AfroLID: A Neural Language Identification Tool for African Languages. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 1958--1981
work page 2022
-
[6]
Adelani, D. I.; Ojo, J.; Azime, I. A.; Zhuang, J. Y.; Alabi, J. O.; He, X.; Ochieng, M.; Hooker, S.; Bukula, A.; Lee, E.-S. A.; et al. 2024. IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models. arXiv preprint arXiv:2406.03368
arXiv 2024
-
[7]
Arora, A.; Kaffee, L.-a.; and Augenstein, I. 2023. Probing Pre-Trained Language Models for Cross-Cultural Differences in Values. In Dev, S.; Prabhakaran, V.; Adelani, D.; Hovy, D.; and Benotti, L., eds., Proceedings of the First Workshop on Cross-Cultural Considerations in NLP (C3NLP), 114--130. Dubrovnik, Croatia: Association for Computational Linguistics
work page 2023
-
[8]
Aryabumi, V.; Dang, J.; Talupuru, D.; Dash, S.; Cairuz, D.; Lin, H.; Venkitesh, B.; Smith, M.; Marchisio, K.; Ruder, S.; et al. 2024. Aya 23: Open weight releases to further multilingual progress. arXiv preprint arXiv:2405.15032
arXiv 2024
Show all 51 references
-
[9]
N.; Husa, D.; Goyal, N.; Krishnan, A.; Zettlemoyer, L.; and Khabsa, M
Bandarkar, L.; Liang, D.; Muller, B.; Artetxe, M.; Shukla, S. N.; Husa, D.; Goyal, N.; Krishnan, A.; Zettlemoyer, L.; and Khabsa, M. 2024. The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants. In Proceedings of the 62nd Annual Meeting of th...
2024
-
[10]
Bapna, A.; Caswell, I.; Kreutzer, J.; Firat, O.; van Esch, D.; Siddhant, A.; Niu, M.; Baljekar, P.; Garcia, X.; Macherey, W.; et al. 2022. Building machine translation systems for the next thousand languages. arXiv preprint arXiv:2205.03983
2022 arXiv
-
[11]
Benjamin, M. 2019. Empirical Evaluation of Google Translate across 107 Languages. https://www.teachyoubackwards.com/empirical-evaluation/. Accessed: 2024-07-25
2019
-
[12]
Beukman, M.; and Fokam, M. 2023. Analysing Cross-Lingual Transfer in Low-Resourced African Named Entity Recognition. In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association f...
2023
-
[13]
BMGF. 2024. Expanding Reasoning Benchmarks in Low-Resourced African Languages: Winogrande and Clinical MMLU in Afrikaans, Xhosa, and Zulu. https://github.com/InstituteforDiseaseModeling/winogrande-mmlu-clinical-za
2024
-
[14]
Borin, L.; Comrie, B.; and Saxena, A. 2013. The Intercontinental Dictionary Series: A rich and principled database for language comparison
2013
-
[15]
Chen, J.; and Mueller, J. 2024. Automated data curation for robust language model fine-tuning. arXiv preprint arXiv:2403.12776
2024 arXiv
-
[16]
CIA. 2024. The World Factbook — World. https://www.cia.gov/the-world-factbook/countries/world/#people-and-society. Accessed: 2024-07-29
2024
-
[17]
Cole, J.; Zhang, M.; Gillick, D.; Eisenschlos, J.; Dhingra, B.; and Eisenstein, J. 2023. Selectively Answering Ambiguous Questions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 530--543
2023
-
[18]
Conneau, A.; Rinott, R.; Lample, G.; Williams, A.; Bowman, S.; Schwenk, H.; and Stoyanov, V. 2018. XNLI: Evaluating Cross-lingual Sentence Representations. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2475. Association for Computat...
2018
-
[19]
De Melo, G.; Imaizumi, V.; and Cozman, F. 2019. Winograd schemas in portuguese. In Anais do XVI Encontro Nacional de Intelig \^e ncia Artificial e Computacional , 787--798. SBC
2019
-
[20]
Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024. The Llama 3 Herd of Models. arXiv preprint arXiv:2407.21783
2024 arXiv
-
[21]
Galal, S. 2024. Africa: Share of global poverty by country 2024. https://www.statista.com/statistics/1228553/extreme-poverty-as-share-of-global-population-in-africa-by-country. Accessed: 2024-07-29
2024
-
[22]
Hammerstrom, H. 2015. Ethnologue 16/17/18th editions: A comprehensive review . LANGUAGE, 91(3)
2015
-
[23]
Hartvigsen, T.; Gabriel, S.; Palangi, H.; Sap, M.; Ray, D.; and Kamar, E. 2022. ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volu...
2022
-
[24]
Hendrycks, D.; Burns, C.; Basart, S.; Critch, A.; Li, J.; Song, D.; and Steinhardt, J. 2021 a . Aligning AI With Shared Human Values. Proceedings of the International Conference on Learning Representations (ICLR)
2021
-
[25]
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021 b . Measuring Massive Multitask Language Understanding. Proceedings of the International Conference on Learning Representations (ICLR)
2021
-
[26]
Hu, J.; Ruder, S.; Siddhant, A.; Neubig, G.; Firat, O.; and Johnson, M. 2020. Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation. In International Conference on Machine Learning, 4411--4421. PMLR
2020
-
[27]
ITU. 2023. Facts and Figures 2023. https://www.itu.int/itu-d/reports/statistics/facts-figures-2023/. Accessed: 2024-07-29
2023
-
[28]
Joshi, P.; Santy, S.; Budhiraja, A.; Bali, K.; and Choudhury, M. 2020. The State and Fate of Linguistic Diversity and Inclusion in the NLP World. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 6282. Association for Computational Linguistics
2020
-
[29]
Liang, Y.; Duan, N.; Gong, Y.; Wu, N.; Guo, F.; Qi, W.; Gong, M.; Shou, L.; Jiang, D.; Cao, G.; et al. 2020. XGLUE: A New Benchmark Dataset for Cross-lingual Pre-training, Understanding and Generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Langu...
2020
-
[30]
V.; Mihaylov, T.; Artetxe, M.; Wang, T.; Chen, S.; Simig, D.; Ott, M.; Goyal, N.; Bhosale, S.; Du, J.; et al
Lin, X. V.; Mihaylov, T.; Artetxe, M.; Wang, T.; Chen, S.; Simig, D.; Ott, M.; Goyal, N.; Bhosale, S.; Du, J.; et al. 2021. Few-shot learning with multilingual language models. arXiv preprint arXiv:2112.10668
2021 arXiv
-
[31]
M.; Reddy, S.; Collier, N.; and Elliott, D
Liu, F.; Bugliarello, E.; Ponti, E. M.; Reddy, S.; Collier, N.; and Elliott, D. 2021. Visually Grounded Reasoning across Languages and Cultures. arXiv:2109.13238
2021 arXiv
-
[32]
Longpre, S.; Yauney, G.; Reif, E.; Lee, K.; Roberts, A.; Zoph, B.; Zhou, D.; Wei, J.; Robinson, K.; Mimno, D.; et al. 2024. A Pretrainer’s Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, and Toxicity. In Proceedings of the 2024 Conference o...
2024
-
[33]
S.; Shen, S.; Yong, Z
Muennighoff, N.; Wang, T.; Sutawika, L.; Roberts, A.; Biderman, S.; Le Scao, T.; Bari, M. S.; Shen, S.; Yong, Z. X.; Schoelkopf, H.; et al. 2023. Crosslingual Generalization through Multitask Finetuning. In Proceedings of the 61st Annual Meeting of the Association for Computat...
2023
-
[34]
NLLB Team ; et al. 2024. Scaling neural machine translation to 200 languages. Nature, 630(8018): 841
2024
-
[35]
OpenAI. 2023. GPT-3.5 Turbo fine-tuning and API updates. https://openai.com/index/gpt-3-5-turbo-fine-tuning-and-api-updates/. Accessed: 2024-07-26
2023
-
[36]
OpenAI. 2024. Hello GPT-4o. https://openai.com/index/hello-gpt-4o/. Accessed: 2024-07-26
2024
-
[37]
Polignano, M.; Basile, P.; and Semeraro, G. 2024. Advanced Natural-based interaction for the ITAlian language: LLaMAntino-3-ANITA. arXiv:2405.07101
2024 arXiv
-
[38]
A.; Haznitrama, F
Putri, R. A.; Haznitrama, F. G.; Adhista, D.; and Oh, A. 2024. Can LLM Generate Culturally Relevant Commonsense QA Data? Case Study in Indonesian and Sundanese. arXiv:2402.17302
2024 arXiv
-
[39]
L.; Bhagavatula, C.; and Choi, Y
Sakaguchi, K.; Bras, R. L.; Bhagavatula, C.; and Choi, Y. 2021. Winogrande: An adversarial winograd schema challenge at scale. Communications of the ACM, 64(9): 99--106
2021
-
[40]
A.; and Choi, Y
Sap, M.; Gabriel, S.; Qin, L.; Jurafsky, D.; Smith, N. A.; and Choi, Y. 2020. Social Bias Frames: Reasoning about Social and Power Implications of Language. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5477--5490
2020
-
[41]
Shaham, U.; Herzig, J.; Aharoni, R.; Szpektor, I.; Tsarfaty, R.; and Eyal, M. 2024. Multilingual instruction tuning with just a pinch of multilinguality. arXiv preprint arXiv:2401.01854
2024 arXiv
-
[42]
Statistics South Africa . 2020. Labour Market Dynamics in South Africa . Technical report, Statistics South Africa, Private Bag X44, Pretoria 0001
2020
-
[43]
Tan, Z.; Beigi, A.; Wang, S.; Guo, R.; Bhattacharjee, A.; Jiang, B.; Karami, M.; Li, J.; Cheng, L.; and Liu, H. 2024. Large language models for data annotation: A survey. arXiv preprint arXiv:2402.13446
2024 arXiv
-
[44]
\"U st \"u n, A.; Aryabumi, V.; Yong, Z.-X.; Ko, W.-Y.; D'souza, D.; Onilude, G.; Bhandari, N.; Singh, S.; Ooi, H.-L.; Kayid, A.; et al. 2024. Aya model: An instruction finetuned open-access multilingual language model. arXiv preprint arXiv:2402.07827
2024 arXiv
-
[45]
Villalobos, P.; Ho, A.; Sevilla, J.; Besiroglu, T.; Heim, L.; and Hobbhahn, M. 2024. Position: Will we run out of data? Limits of LLM scaling based on human-generated data. In Forty-first International Conference on Machine Learning
2024
-
[46]
Wen, J.; Ke, P.; Sun, H.; Zhang, Z.; Li, C.; Bai, J.; and Huang, M. 2023. Unveiling the Implicit Toxicity in Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1322--1338
2023
-
[47]
Whitehouse, C.; Choudhury, M.; and Aji, A. 2023. LLM-powered Data Augmentation for Enhanced Cross-lingual Performance. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 671--686
2023
-
[48]
Xu, J.; Ju, D.; Li, M.; Boureau, Y.-L.; Weston, J.; and Dinan, E. 2021. Bot-adversarial dialogue for safe conversational agents. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2...
2021
-
[49]
Yao, B.; Jiang, M.; Yang, D.; and Hu, J. 2023. Benchmarking LLM-based Machine Translation on Cultural Awareness
2023
-
[50]
P.; et al
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li13, D.; Xing35, E. P.; et al. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv preprint arXiv:2306.05685
2023 arXiv
-
[51]
Zhou, C.; Liu, P.; Xu, P.; Iyer, S.; Sun, J.; Mao, Y.; Ma, X.; Efrat, A.; Yu, P.; Yu, L.; et al. 2024. Lima: Less is more for alignment. Advances in Neural Information Processing Systems, 36
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.