REVIEW 3 major objections 6 minor 81 references
Analysis of Indic Language Capabilities in LLMs
T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Hindi, Bengali, Marathi, Telugu, and Tamil are the Indic languages most ready for AI safety benchmarks, this review of 28 large language models argues.
desk verdict Useful synthesis of Indic LLM capability evidence with a defensible top-five safety benchmark recommendation, but the capability-to-safety transfer is assumed, not tested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing device is Table 4, a synthesized ranking of twelve Indic languages into high, medium, and low performance on natural language understanding and generation tasks, built from fourteen evaluation studies. The report treats this ranking, combined with Census speaker counts, as the basis for its recommendation. The survey of twenty-eight models and their attributes provides the context but the recommendation itself rests on the performance synthesis.
What would settle it
Run a safety evaluation, for example refusal of harmful prompts or detection of toxic content, in the five recommended languages plus a low-rated language with many speakers such as Oriya or Punjabi. If the low-rated language matches or exceeds the recommended five on safety-relevant behaviour, then the capability-based priority list would be wrong even if Table 4's general-performance ranking is correct.
Extended reading notes
Core claim
The report's central claim is that the strongest candidates for safety benchmarking are the five most widely spoken Indic languages: Hindi, Bengali, Marathi, Telugu, and Tamil. This is based on a synthesis of existing evaluation results, summarized in Table 4, which buckets twelve languages into high, medium, and low performance on natural language understanding and generation tasks. The five recommended languages combine consistently high performance across a majority of models and subtasks with large speaker populations. The report explicitly warns that these rankings are relative, not absolute, and that the gap between English and all Indic languages remains large.
Load-bearing premise
The recommendation assumes that strong performance on translated general-understanding and generation benchmarks shows that a language is ready for safety benchmarking, even though the report never tests safety-specific behaviour.
Editorial extensions
If this is right
- Safety benchmark builders should begin with Hindi, Bengali, Marathi, Telugu, and Tamil, since only these combine consistently high model performance with large speaker populations.
- For languages beyond these five, the report recommends against unequivocal inclusion; model performance is more uneven and speaker counts do not align with capability.
- The alignment of performance and speaker count means the top five are also where unsafe model behaviour would affect the most people, making them the natural starting point for safety testing.
- The priority list is time-sensitive: as new training corpora and fine-tuned models appear, a language currently rated low could become benchmark-ready, so the ranking should be revisited.
Reading between the lines
- Because most existing benchmarks are direct English translations, the high ratings in Table 4 may overstate real-world cultural fluency; safety evaluations in these five languages should include locally authored prompts, not only translated ones.
- The report's exclusion of safety datasets leaves the recommendation as a capability proxy. Running a shared safety benchmark in the five recommended and several non-recommended languages would directly test whether general performance predicts safety behaviour.
- A language such as Oriya has a large speaker base but low current model performance; with focused corpus-building it could leapfrog into benchmark readiness faster than the current ranking suggests.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript, produced with MLCommons funding, surveys the landscape of LLM support for Indic languages and proposes a prioritization of languages for inclusion in AI safety benchmarks. It reviews 28 LLMs with Indic capabilities and 14 evaluation papers, categorizes 12 major Indic languages into high/medium/low understanding and generation performance buckets (Table 4), and concludes in Section 5 that Hindi, Bengali, Marathi, Telugu, and Tamil are the natural candidates for safety benchmarks. The paper is a desk-review: it aggregates existing evidence on model inventories, training corpus statistics (Wikipedia, Common Crawl, mC4), and evaluation studies rather than conducting new experiments.
Significance. If its central recommendation is accepted, the paper would directly influence which Indic languages get prioritized in MLCommons' safety benchmark design, with real-world consequences for the evaluation of AI safety in multilingual settings. The paper's descriptive contributions are solid: it provides a systematic inventory of 28 models with license/access/training data attributes, a useful taxonomy of NLU/NLG tasks, a clear compilation of corpus statistics, and a candid acknowledgment that most existing benchmarks are translated English datasets. These strengths make the survey a valuable reference for practitioners. However, the paper's central policy claim rests on a capability-to-safety transfer assumption that is neither argued nor tested, and the ranking that drives the recommendation is not reproducible from the information provided. The paper is best viewed as a high-quality descriptive survey whose prescriptive conclusion is significantly stronger than the evidence supports.
major comments (3)
- [Section 5 (heading 'Prioritizing Languages for Inclusion in Benchmarks') and Section 4] The central claim, 'Based on Table 4 Hindi, Bengali, Marathi, Telugu and Tamil emerge as natural candidates for inclusion in safety benchmarks,' is load-bearing and rests on an untested proxy. The paper explicitly states in Section 4 that safety-related datasets are precluded, and Section 4.1 acknowledges that most evaluations are direct translations of English datasets and may not capture socio-cultural understanding. No evidence is provided that high NLU/NLG performance on translated benchmarks transfers to meaningful safety evaluation (e.g., refusal behavior, culturally appropriate responses, handling of harmful content in Indic languages). Given that the stated purpose of the report is to inform safety benchmark design, the authors should either (i) substantially hedge the recommendation to 'languages with currently strongest general capabilities, pending safety-specific evaluation,' or (ii) provide a concrete argument and supporting evidence for why general capability ranks serve as a valid proxy for safety readiness.
- [Table 4 and Section 5] Table 4 lists Urdu as 'high' for both understanding and generation, with 50,772,631 speakers, yet Section 5 excludes Urdu from the recommended five without explanation. If the selection criterion is high capability plus large speaker count, Urdu's exclusion is internally inconsistent; if a different criterion is used (e.g., script, data availability, or a speaker threshold), it is not stated. This inconsistency undermines the reproducibility of the prioritization and suggests the selection rule is applied post hoc. Please clarify the exact algorithm by which the five languages are selected and explain explicitly why Urdu — and, for the same reason, Gujarati — is not recommended despite ranking medium/high.
- [Section 4.1, Table 4] The high/medium/low labels in Table 4 are produced through a qualitative synthesis with no reproducible aggregation protocol. The definitions ('consistently ranked at the top,' 'average performance,' 'poor scores,' 'across a majority of models') do not specify which datasets, which models, which tasks, or how 'majority' is counted across heterogeneous studies. Because the entire policy recommendation depends on these labels, the absence of a transparent scoring rubric or a supplementary table with per-dataset/per-model evidence is a major limitation. Please provide either a detailed rubric with explicit thresholds or release the underlying synthesis table so the assignment to 'high/medium/low' can be independently verified.
minor comments (6)
- [Section 3.1] In the Airavata description, 'LymSys-Chat' appears to be a typo for 'LMSYS-Chat'; the reference list entry [77] supports the correct spelling.
- [Section 3.3] The sentence 'most models do not release their data under and an open data license' contains a grammatical error; it should be 'under an open data license.'
- [Table 4] The column header 'Langauge' should be 'Language.' Similar typos occur elsewhere (e.g., 'unkown' in Table 5).
- [Section 1, references] Reference [11] (Ali et al., 'Taking Stock of Concept Inventories in Computing Education') appears unrelated to the Bing Chat Spanish footwear example cited in the text. Please verify and replace with the correct source.
- [Section 5, footnote 7] The footnote states that Hindi is already included in MLCommons' v1 AI safety benchmark; if so, the phrase 'natural candidates for inclusion' should be clarified to indicate that the recommendation aims at future benchmark versions or expansion beyond v1.
- [Appendix B] The appendix describes MILU as covering 11 Indic languages and evaluating 45 LLMs, but the main text says 'we found fourteen papers' and describes a maximum of 12 languages in most datasets; consider harmonizing the counts in a summary table to avoid confusion.
Circularity Check
No circularity: the report is a desk-review synthesis of external evaluations; its five-language recommendation is a bounded interpretive judgment, not a quantity derived from its own definitions or fitted parameters.
full rationale
This is a survey-style report, not a derivation. It contains no equations, no fitted parameters, and no claimed prediction that could reduce by construction to its inputs. Section 4 summarizes externally published evaluation studies, and Table 4 is an explicitly qualitative categorization ('a high-level approximation') of those cited results, with the caveat that the rankings are relative and not universal. Section 5 recommends Hindi, Bengali, Marathi, Telugu, and Tamil for safety benchmarks by combining Table 4 performance with speaker counts; that is an interpretation of external evidence, not a self-validating construction. The paper explicitly limits its scope: 'Since the goal of the report is to provide suggestions on which Indic language should be included in future safety benchmark datasets, we preclude an analysis of trust and safety datasets relevant to Indic languages.' It also flags that most datasets are 'a direct translation of existing English datasets which is a serious limitation.' Those are acknowledged assumptions about external validity, not circular reasoning. There are no load-bearing self-citations: the authors cite external benchmark papers, and their own prior work is not invoked to justify the central claim. The recommendation's dependence on the untested transfer from general NLU/NLG performance to safety readiness is a weakness in the argument's evidence base, but the claim is not definitionally equivalent to its inputs. Thus the appropriate finding is no significant circularity (score 0).
Assumptions & free parameters
assumptions (5)
- domain assumption Performance labels in Table 4 are comparable across heterogeneous evaluation studies, datasets, models, and tasks.
- domain assumption Developer claims of Indic language support, from technical reports, blogs, and model cards, reflect actual usable capability.
- domain assumption General NLU/NLG benchmark performance is a valid proxy for readiness for safety benchmarking.
- domain assumption Speaker counts from the Census of India 2011 are an appropriate measure of real-world language exposure for prioritization.
- domain assumption Public corpus distributions from Wikipedia, Common Crawl, and C4 accurately indicate the Indic language data available to LLM developers.
Cite this review
Pith. "Pith review of Analysis of Indic Language Capabilities in LLMs." pith.science (2026). https://pith.science/paper/BRKD2IMH
@misc{pith2026250113912,
author = {Pith},
title = {Pith review of: Analysis of Indic Language Capabilities in LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/BRKD2IMH}},
note = {Machine review of arXiv:2501.13912}
}
read the original abstract
This report evaluates the performance of text-in text-out Large Language Models (LLMs) to understand and generate Indic languages. This evaluation is used to identify and prioritize Indic languages suited for inclusion in safety benchmarks. We conduct this study by reviewing existing evaluation studies and datasets; and a set of twenty-eight LLMs that support Indic languages. We analyze the LLMs on the basis of the training data, license for model and data, type of access and model developers. We also compare Indic language performance across evaluation datasets and find that significant performance disparities in performance across Indic languages. Hindi is the most widely represented language in models. While model performance roughly correlates with number of speakers for the top five languages, the assessment after that varies.
Figures
Reference graph
Works this paper leans on
-
[1]
[n. d.]. Introducing Llama 3.1: Our most capable models to date — ai.meta.com. https://ai.meta.com/blog/meta-llama-3-1/
-
[2]
Paper 1 of 2018 on Language, Table C-16, Census of India 2011
2011. Paper 1 of 2018 on Language, Table C-16, Census of India 2011. https://censusindia.gov.in/nada/index.php/catalog/42458/ download/46089/C-16_25062018.pdf
work page 2011
-
[3]
2022. https://community.openai.com/t/chatgpt-spanish-support/24706/3 7Hindi is already included in v1 AI safety benchmark released by MLCommons 8This gap between language speakers and model capabilities can be addressed by model developers but is out of scope for the MLCommons’ AI Safety Working Group Analysis of Indic Language Capabilities in LLMs • 9
work page 2022
-
[4]
Rajbhasha Vibhag, Ministry of Home Affairs
2024. Rajbhasha Vibhag, Ministry of Home Affairs. https://rajbhasha.gov.in/en/languages-included-eighth-schedule-indian-constitution
work page 2024
-
[5]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[6]
Parul Agarwal, Aisha Asif, Shantipriya Parida, Sambit Sekhar, Satya Ranjan Dash, and Subhadarshi Panda. 2023. Generative Chatbot Adaptation for Odia Language: A Critical Evaluation. In 2023 1st International Conference on Circuits, Power and Intelligent Systems (CCPIS). 1–7. https://doi.org/10.1109/CCPIS59145.2023.10291329
-
[7]
Divyanshu Aggarwal, Vivek Gupta, and Anoop Kunchukuttan. 2022. IndicXNLI: Evaluating Multilingual Inference for Indian Languages. arXiv:2204.08776 [cs.CL]
arXiv 2022
-
[8]
Divyanshu Aggarwal, Ashutosh Sathe, and Sunayana Sitaram. 2024. Maple: Multilingual evaluation of parameter efficient finetuning of large language models. arXiv preprint arXiv:2401.07598 (2024)
work page Pith review arXiv 2024
Show all 81 references
-
[9]
Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Maxamed Axmed, et al. 2023. Mega: Multilingual evaluation of generative ai. arXiv preprint arXiv:2303.12528 (2023)
2023 arXiv
-
[10]
Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Prachi Jain, Maxamed Axmed, Kalika Bali, et al. 2023. Megaverse: Benchmarking large language models across languages, modalities, models and tasks. arXiv preprint arXi...
2023 arXiv
-
[11]
Murtaza Ali, Sourojit Ghosh, Prerna Rao, Raveena Dhegaskar, Sophia Jawort, Alix Medler, Mengqi Shi, and Sayamindu Dasgupta
-
[12]
Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Jon Ander Campos, Yi Chern Tan, Kelly Marchisio, Max Bartolo, Sebastian Ruder, Acyr Locatelli, Julia Kreutzer, Nick Frosst, Aidan Gomez, Phil Blunsom, Marzieh...
2024 arXiv
-
[13]
Abhinand Balachandran. 2023. Tamil-llama: A new tamil language model based on llama 2. arXiv preprint arXiv:2311.05845 (2023)
2023 arXiv
-
[14]
Abhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, and Rifat Shahriyar. 2023. BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla. In Findings of the Association for Computational Linguistics: EACL 2023, Andreas ...
2023 doi
-
[15]
Khapra, and Pratyush Kumar
Kaushal Santosh Bhogale, Sai Sundaresan, Abhigyan Raman, Tahir Javed, Mitesh M. Khapra, and Pratyush Kumar. 2023. Vistaar: Diverse Benchmarks and Training Sets for Indian Language ASR. arXiv:2305.15386 [cs.CL]
2023 arXiv
-
[16]
Rishi Bommasani, Dilara Soylu, Thomas I Liao, Kathleen A Creel, and Percy Liang. 2023. Ecosystem graphs: The social footprint of foundation models. arXiv preprint arXiv:2303.15772 (2023)
2023 arXiv
-
[17]
Govind Choudhary. 2023. Chatgpt now speaks Hindi, Assamese, Bengali and other Indian languages! https: //www.livemint.com/technology/apps/chatgpt-now-speaks-hindi-assamese-bengali-and-other-indian-languages-heres-how- to-get-replies-in-local-languages-11687948693940.html
2023
-
[18]
Monojit Choudhury and Amit Deshpande. 2021. How Linguistically Fair are Multilingual Pre-Trained Language Models?. In AAAI-21. AAAI, AAAI. https://www.microsoft.com/en-us/research/publication/how-linguistically-fair-are-multilingual-pre-trained-language- models/
2021
-
[19]
A Conneau. 2019. Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116 (2019)
2019 arXiv
-
[20]
Raj Dabre, Himani Shrotriya, Anoop Kunchukuttan, Ratish Puduppully, Mitesh M Khapra, and Pratyush Kumar. 2021. IndicBART: A pre-trained model for indic natural language generation. arXiv preprint arXiv:2109.02903 (2021)
2021 arXiv
-
[21]
Mithun Das, Punyajoy Saha, Binny Mathew, and Animesh Mukherjee. 2022. HateCheckHIn: Evaluating Hindi Hate Speech Detection Models. arXiv:2205.00328 [cs.CL]
2022 arXiv
-
[22]
Paresh Dave. 2023. ChatGPT Is Cutting Non-English Languages Out of the AI Revolution. Wired (2023). https://www.wired.com/story/ chatgpt-non-english-languages-ai-revolution/
2023
-
[23]
Andrew Deck. 2023. We tested ChatGPT in Bengali, Kurdish, and Tamil. it failed. https://restofworld.org/2023/chatgpt-problems- global-language-testing/
2023
-
[24]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 [cs.CL] https://arxiv.org/abs/1810.04805
2019 arXiv
-
[25]
Martin Dittus and Mark Graham. 2019. The Language Geography of Wikipedia. https://internetlanguages.org/en/numbers/wikipedia- language-geography/
2019
-
[27]
Khapra, Anoop Kunchukuttan, and Pratyush Kumar
Sumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M. Khapra, Anoop Kunchukuttan, and Pratyush Kumar. 2023. Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages. 10 • Aatman Vaidya, Tarunima P...
2023 arXiv
-
[28]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)
2024 arXiv
-
[29]
Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, et al. 2022. Massive: A 1m-example multilingual natural language understanding dataset with 51 typologically- diverse languages. ...
2022 arXiv
-
[30]
Common Crawl Foundation. [n. d.]. Statistics of Common Crawl–Distribution of Languages. https://commoncrawl.github.io/cc-crawl- statistics/plots/languages
-
[31]
Jay Gala, Thanmay Jayakumar, Jaavid Aktar Husain, Mohammed Safi Ur Rahman Khan, Diptesh Kanojia, Ratish Puduppully, Mitesh M Khapra, Raj Dabre, Rudra Murthy, Anoop Kunchukuttan, et al. 2024. Airavata: Introducing hindi instruction-tuned llm. arXiv preprint arXiv:2401.15006 (2024)
2024 arXiv
-
[32]
Rishav Hada, Varun Gumma, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2024. METAL: Towards Multilingual Meta-Evaluation. arXiv preprint arXiv:2404.01667 (2024)
2024 arXiv
-
[33]
Barry Haddow and Faheem Kirefu. 2020. PMIndia – A Collection of Parallel Corpora of Languages of India. arXiv:2001.09907 [cs.CL]
2020 arXiv
-
[34]
Sohel Rahman, and Rifat Shahriyar
Tahmid Hasan, Abhik Bhattacharjee, Md Saiful Islam, Kazi Samin, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar
-
[35]
Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023. How good are GPT models at machine translation? A comprehensive evaluation. arXiv.arXiv preprint arXiv:2302.09210 (2023)
2023 arXiv
-
[36]
Carolin Holtermann, Paul Röttger, Timm Dill, and Anne Lauscher. 2024. Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ. arXiv:2403.03814 [cs.CL]
2024 arXiv
-
[37]
Kushal Jain, Adwait Deshpande, Kumar Shridhar, Felix Laumann, and Ayushman Dash. 2020. Indic-Transformers: An Analysis of Transformer Language Models for Indian Languages. arXiv:2011.02323 [cs.CL] https://arxiv.org/abs/2011.02323
2020 arXiv
-
[38]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023)
2023 arXiv
-
[39]
Zhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding, and Graham Neubig. 2020. X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language Models. arXiv:2010.06189 [cs.CL]
2020 arXiv
-
[40]
Mohsinul Kabir, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Mir Tafseer Nayeem, M Saiful Bari, and Enamul Hoque
-
[41]
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, NC Gokul, Avik Bhattacharyya, Mitesh M Khapra, and Pratyush Kumar. 2020. IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages. In Findings of the Associa...
2020
-
[42]
Nora Kassner, Philipp Dufter, and Hinrich Schütze. 2021. Multilingual LAMA: Investigating Knowledge in Multilingual Pretrained Language Models. arXiv:2102.00894 [cs.CL]
2021 arXiv
-
[43]
Simran Khanuja, Diksha Bansal, Sarvesh Mehtani, Savya Khosla, Atreyee Dey, Balaji Gopalan, Dilip Kumar Margam, Pooja Aggarwal, Rajiv Teja Nagipogu, Shachi Dave, et al. 2021. Muril: Multilingual representations for indian languages. arXiv preprint arXiv:2103.10730 (2021)
2021 arXiv
-
[44]
Simran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram, and Monojit Choudhury. 2020. GLUECoS: An evaluation benchmark for code-switched NLP. arXiv preprint arXiv:2004.12376 (2020)
2020 arXiv
-
[45]
Sankalp KJ, Vinija Jain, Sreyoshi Bhaduri, Tamoghna Roy, and Aman Chadha. 2024. Decoding the Diversity: A Review of the Indic AI Research Landscape. arXiv:2406.09559 [cs.CL] https://arxiv.org/abs/2406.09559
2024 arXiv
-
[46]
Guneet S Kohli, Shantipriya Parida, Sambit Sekhar, Samirit Saha, Nipun B Nair, Parul Agarwal, Sonal Khosla, Kusumlata Patiyal, and Debasish Dhal. 2023. Building a llama2-finetuned llm for odia language utilizing domain knowledge instruction set. In Proceedings of the Third Int...
2023
-
[47]
Khapra, and Pratyush Kumar
Aman Kumar, Himani Shrotriya, Prachi Sahu, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, Amogh Mishra, Mitesh M. Khapra, and Pratyush Kumar. 2022. IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages. arXiv:2203.05437 [cs.CL]
2022 arXiv
-
[48]
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2023. Bloom: A 176b-parameter open-access multilingual language model. (2023)
2023
-
[49]
Wei Qi Leong, Jian Gang Ngui, Yosephine Susanto, Hamsawardhini Rengarajan, Kengatharaiyer Sarveswaran, and William Chandra Tjhi
-
[50]
David McCandless. 2024. The rise of Generative AI large language models (llms) like chatgpt. https://informationisbeautiful.net/ visualizations/the-rise-of-generative-ai-large-language-models-llms-like-chatgpt/
2024
-
[51]
Arnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M Khapra, Pratyush Kumar, Rudra Murthy V, and Anoop Kunchukuttan
-
[52]
Shubham Mittal, Megha Sundriyal, and Preslav Nakov. 2023. Lost in Translation, Found in Spans: Identifying Claims in Multilingual Social Media. arXiv:2310.18205 [cs.CL]
2023 arXiv
-
[53]
arXiv:2309.06085 [cs.CL]
BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models. arXiv:2309.06085 [cs.CL]
-
[54]
Gabriel Nicholas and Aliya Bhatia. 2023. Lost in Translation: Large Language Models in Non-English Content Analysis. arXiv:2306.07377 [cs.CL]
2023 arXiv
-
[55]
Mitodru Niyogi and Arnab Bhattacharya. 2024. Paramanu: A Family of Novel Efficient Indic Generative Foundation Language Models. arXiv preprint arXiv:2401.18034 (2024)
2024
-
[56]
Millicent Ochieng, Varun Gumma, Sunayana Sitaram, Jindong Wang, Keshet Ronen, Kalika Bali, and Jacki O’Neill. 2024. Beyond Metrics: Evaluating LLMs’ Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios. arXiv preprint arXiv:2406.00343 (2024)
2024 arXiv
-
[57]
OpenAI. 2024. Hello GPT-4o. https://openai.com/index/hello-gpt-4o/
2024
-
[58]
Rossi, and Thien Huu Nguyen
Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen
-
[59]
arXiv:2309.09400 [cs.CL]
CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages. arXiv:2309.09400 [cs.CL]
-
[60]
Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan Ak, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Divyanshu Kakwani, Navneet Kumar, et al. 2022. Samanantar: The largest publicly available parallel corpora collection for 11 indic languages. ...
2022
-
[61]
Mielke, Cibu Johny, Isin Demirsahin, and Keith Hall
Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov, Sabrina J. Mielke, Cibu Johny, Isin Demirsahin, and Keith Hall. 2020. Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset. InProceedings of the Twelfth Language Resources and Evaluation Conference...
2020
-
[62]
Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021. XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation. In Proceedings of the 2021 Conference on E...
2021
-
[63]
Sarvam. 2024. Sarvam 1. https://www.sarvam.ai/blogs/sarvam-1
2024
-
[64]
Shantipriya Parida, Satya Ranjan Dash, Ondřej Bojar, Petr Motlicek, Priyanka Pattnaik, and Debasish Kumar Mallick. 2020. OdiEnCorp 2.0: Odia-English Parallel Corpus for Machine Translation. In Proceedings of the WILDRE5– 5th Workshop on Indian Language Data: Resources and Eval...
2020
-
[65]
Ojha, Saraswati Sahoo, Satya Ranjan Dash, and Bijayalaxmi Dash
Shantipriya Parida, Kalyanamalini Sahoo, Atul Kr. Ojha, Saraswati Sahoo, Satya Ranjan Dash, and Bijayalaxmi Dash. 2022. Universal Dependency Treebank for Odia Language. arXiv:2205.11976 [cs.CL]
2022 arXiv
-
[66]
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023)
2023 arXiv
-
[67]
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. 2024. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118 (2024)
2024 arXiv
-
[68]
Mistral Team. 2023. Mixtral of experts. https://mistral.ai/news/mixtral-of-experts/
2023
-
[69]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023)
2023 arXiv
-
[70]
Harman Singh, Nitish Gupta, Shikhar Bharadwaj, Dinesh Tewari, and Partha Talukdar. 2024. IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages. arXiv preprint arXiv:2404.16816 (2024)
2024 arXiv
-
[71]
Anirudh Srinivasan, Sunayana Sitaram, Tanuja Ganu, Sandipan Dandapat, Kalika Bali, and Monojit Choudhury. 2021. Predicting the Performance of Multilingual NLP Models. arXiv:2110.08875 [cs.CL]
2021 arXiv
-
[72]
Ishaan Watts, Varun Gumma, Aditya Yadavalli, Vivek Seshadri, Manohar Swaminathan, and Sunayana Sitaram. 2024. PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data.arXiv preprint arXiv:2406.15053 (2024)
2024 arXiv
-
[73]
Shijie Wu and Mark Dredze. 2020. Are All Languages Created Equal in Multilingual BERT?. In Proceedings of the 5th Workshop on Representation Learning for NLP , Spandana Gella, Johannes Welbl, Marek Rei, Fabio Petroni, Patrick Lewis, Emma Strubell, Minjoon Seo, and Hannaneh Haj...
2020 doi
-
[75]
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. arXiv:2010.11934 [cs.CL] https://arxiv.org/abs/2010.11934
2021 arXiv
-
[76]
Aniket Vashishtha, Kabir Ahuja, and Sunayana Sitaram. 2023. On evaluating and mitigating gender biases in multilingual settings. arXiv preprint arXiv:2307.01503 (2023)
2023 arXiv
-
[77]
Sshubam Verma, Mohammed Safi Ur Rahman Khan, Vishwajeet Kumar, Rudra Murthy, and Jaydeep Sen. 2024. MILU: A Multi-task Indic Language Understanding Benchmark. arXiv:2411.02538 [cs.CL] https://arxiv.org/abs/2411.02538
2024 arXiv
-
[82]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong W...
2024 arXiv
-
[83]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P Xing, et al. 2023. Lmsys-chat-1m: A large-scale real-world llm conversation dataset. arXiv preprint arXiv:2309.11998 (2023). A Indic LLM Table Data Tabl...
2023 arXiv
-
[2021]
arXiv:2106.13822 [cs.CL]
XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages. arXiv:2106.13822 [cs.CL]
-
[2022]
arXiv preprint arXiv:2212.10168 (2022)
Naamapadam: a large-scale named entity annotated data for Indic languages. arXiv preprint arXiv:2212.10168 (2022). Analysis of Indic Language Capabilities in LLMs • 11
2022 arXiv
-
[2023]
In Proceedings of the 2023 ACM Conference on International Computing Education Research V.1 (ICER 2023)
Taking Stock of Concept Inventories in Computing Education: A Systematic Literature Review. In Proceedings of the 2023 ACM Conference on International Computing Education Research V.1 (ICER 2023) . ACM, 397–415. https://doi.org/10.1145/3568813.3600120
2023
-
[2024]
arXiv:2309.13173 [cs.CL]
BenLLMEval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP. arXiv:2309.13173 [cs.CL]
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.