Pith. sign in

REVIEW 3 major objections 6 minor 81 references

Analysis of Indic Language Capabilities in LLMs

T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Hindi, Bengali, Marathi, Telugu, and Tamil are the Indic languages most ready for AI safety benchmarks, this review of 28 large language models argues.

desk verdict Useful synthesis of Indic LLM capability evidence with a defensible top-five safety benchmark recommendation, but the capability-to-safety transfer is assumed, not tested. read the letter →

arxiv 2501.13912 v1 pith:BRKD2IMH submitted 2025-01-23 cs.CL

classification cs.CL
keywords IndiclanguageslargelanguagemodelsmultilingualevaluationsafetybenchmarksnaturalunderstandinggenerationHindiBengali
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review asks which Indian languages are strong enough in current large language models to justify inclusion in AI safety benchmarks. The authors synthesize fourteen evaluation studies and survey twenty-eight models that claim Indic language support. They conclude that Hindi, Bengali, Marathi, Telugu, and Tamil are the natural candidates, because models show consistently high understanding and generation performance in those languages. For other languages, either model performance is mediocre or poor, or speaker counts and model abilities diverge. The stakes are practical: safety testing in a language is only meaningful if the model can actually understand and produce that language.

What carries the argument

The load-bearing device is Table 4, a synthesized ranking of twelve Indic languages into high, medium, and low performance on natural language understanding and generation tasks, built from fourteen evaluation studies. The report treats this ranking, combined with Census speaker counts, as the basis for its recommendation. The survey of twenty-eight models and their attributes provides the context but the recommendation itself rests on the performance synthesis.

What would settle it

Run a safety evaluation, for example refusal of harmful prompts or detection of toxic content, in the five recommended languages plus a low-rated language with many speakers such as Oriya or Punjabi. If the low-rated language matches or exceeds the recommended five on safety-relevant behaviour, then the capability-based priority list would be wrong even if Table 4's general-performance ranking is correct.

Watch

Extended reading notes

Core claim

The report's central claim is that the strongest candidates for safety benchmarking are the five most widely spoken Indic languages: Hindi, Bengali, Marathi, Telugu, and Tamil. This is based on a synthesis of existing evaluation results, summarized in Table 4, which buckets twelve languages into high, medium, and low performance on natural language understanding and generation tasks. The five recommended languages combine consistently high performance across a majority of models and subtasks with large speaker populations. The report explicitly warns that these rankings are relative, not absolute, and that the gap between English and all Indic languages remains large.

Load-bearing premise

The recommendation assumes that strong performance on translated general-understanding and generation benchmarks shows that a language is ready for safety benchmarking, even though the report never tests safety-specific behaviour.

Editorial extensions

If this is right

  • Safety benchmark builders should begin with Hindi, Bengali, Marathi, Telugu, and Tamil, since only these combine consistently high model performance with large speaker populations.
  • For languages beyond these five, the report recommends against unequivocal inclusion; model performance is more uneven and speaker counts do not align with capability.
  • The alignment of performance and speaker count means the top five are also where unsafe model behaviour would affect the most people, making them the natural starting point for safety testing.
  • The priority list is time-sensitive: as new training corpora and fine-tuned models appear, a language currently rated low could become benchmark-ready, so the ranking should be revisited.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because most existing benchmarks are direct English translations, the high ratings in Table 4 may overstate real-world cultural fluency; safety evaluations in these five languages should include locally authored prompts, not only translated ones.
  • The report's exclusion of safety datasets leaves the recommendation as a capability proxy. Running a shared safety benchmark in the five recommended and several non-recommended languages would directly test whether general performance predicts safety behaviour.
  • A language such as Oriya has a large speaker base but low current model performance; with focused corpus-building it could leapfrog into benchmark readiness faster than the current ranking suggests.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This manuscript, produced with MLCommons funding, surveys the landscape of LLM support for Indic languages and proposes a prioritization of languages for inclusion in AI safety benchmarks. It reviews 28 LLMs with Indic capabilities and 14 evaluation papers, categorizes 12 major Indic languages into high/medium/low understanding and generation performance buckets (Table 4), and concludes in Section 5 that Hindi, Bengali, Marathi, Telugu, and Tamil are the natural candidates for safety benchmarks. The paper is a desk-review: it aggregates existing evidence on model inventories, training corpus statistics (Wikipedia, Common Crawl, mC4), and evaluation studies rather than conducting new experiments.

Significance. If its central recommendation is accepted, the paper would directly influence which Indic languages get prioritized in MLCommons' safety benchmark design, with real-world consequences for the evaluation of AI safety in multilingual settings. The paper's descriptive contributions are solid: it provides a systematic inventory of 28 models with license/access/training data attributes, a useful taxonomy of NLU/NLG tasks, a clear compilation of corpus statistics, and a candid acknowledgment that most existing benchmarks are translated English datasets. These strengths make the survey a valuable reference for practitioners. However, the paper's central policy claim rests on a capability-to-safety transfer assumption that is neither argued nor tested, and the ranking that drives the recommendation is not reproducible from the information provided. The paper is best viewed as a high-quality descriptive survey whose prescriptive conclusion is significantly stronger than the evidence supports.

major comments (3)
  1. [Section 5 (heading 'Prioritizing Languages for Inclusion in Benchmarks') and Section 4] The central claim, 'Based on Table 4 Hindi, Bengali, Marathi, Telugu and Tamil emerge as natural candidates for inclusion in safety benchmarks,' is load-bearing and rests on an untested proxy. The paper explicitly states in Section 4 that safety-related datasets are precluded, and Section 4.1 acknowledges that most evaluations are direct translations of English datasets and may not capture socio-cultural understanding. No evidence is provided that high NLU/NLG performance on translated benchmarks transfers to meaningful safety evaluation (e.g., refusal behavior, culturally appropriate responses, handling of harmful content in Indic languages). Given that the stated purpose of the report is to inform safety benchmark design, the authors should either (i) substantially hedge the recommendation to 'languages with currently strongest general capabilities, pending safety-specific evaluation,' or (ii) provide a concrete argument and supporting evidence for why general capability ranks serve as a valid proxy for safety readiness.
  2. [Table 4 and Section 5] Table 4 lists Urdu as 'high' for both understanding and generation, with 50,772,631 speakers, yet Section 5 excludes Urdu from the recommended five without explanation. If the selection criterion is high capability plus large speaker count, Urdu's exclusion is internally inconsistent; if a different criterion is used (e.g., script, data availability, or a speaker threshold), it is not stated. This inconsistency undermines the reproducibility of the prioritization and suggests the selection rule is applied post hoc. Please clarify the exact algorithm by which the five languages are selected and explain explicitly why Urdu — and, for the same reason, Gujarati — is not recommended despite ranking medium/high.
  3. [Section 4.1, Table 4] The high/medium/low labels in Table 4 are produced through a qualitative synthesis with no reproducible aggregation protocol. The definitions ('consistently ranked at the top,' 'average performance,' 'poor scores,' 'across a majority of models') do not specify which datasets, which models, which tasks, or how 'majority' is counted across heterogeneous studies. Because the entire policy recommendation depends on these labels, the absence of a transparent scoring rubric or a supplementary table with per-dataset/per-model evidence is a major limitation. Please provide either a detailed rubric with explicit thresholds or release the underlying synthesis table so the assignment to 'high/medium/low' can be independently verified.
minor comments (6)
  1. [Section 3.1] In the Airavata description, 'LymSys-Chat' appears to be a typo for 'LMSYS-Chat'; the reference list entry [77] supports the correct spelling.
  2. [Section 3.3] The sentence 'most models do not release their data under and an open data license' contains a grammatical error; it should be 'under an open data license.'
  3. [Table 4] The column header 'Langauge' should be 'Language.' Similar typos occur elsewhere (e.g., 'unkown' in Table 5).
  4. [Section 1, references] Reference [11] (Ali et al., 'Taking Stock of Concept Inventories in Computing Education') appears unrelated to the Bing Chat Spanish footwear example cited in the text. Please verify and replace with the correct source.
  5. [Section 5, footnote 7] The footnote states that Hindi is already included in MLCommons' v1 AI safety benchmark; if so, the phrase 'natural candidates for inclusion' should be clarified to indicate that the recommendation aims at future benchmark versions or expansion beyond v1.
  6. [Appendix B] The appendix describes MILU as covering 11 Indic languages and evaluating 45 LLMs, but the main text says 'we found fourteen papers' and describes a maximum of 12 languages in most datasets; consider harmonizing the counts in a summary table to avoid confusion.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the report is a desk-review synthesis of external evaluations; its five-language recommendation is a bounded interpretive judgment, not a quantity derived from its own definitions or fitted parameters.

full rationale

This is a survey-style report, not a derivation. It contains no equations, no fitted parameters, and no claimed prediction that could reduce by construction to its inputs. Section 4 summarizes externally published evaluation studies, and Table 4 is an explicitly qualitative categorization ('a high-level approximation') of those cited results, with the caveat that the rankings are relative and not universal. Section 5 recommends Hindi, Bengali, Marathi, Telugu, and Tamil for safety benchmarks by combining Table 4 performance with speaker counts; that is an interpretation of external evidence, not a self-validating construction. The paper explicitly limits its scope: 'Since the goal of the report is to provide suggestions on which Indic language should be included in future safety benchmark datasets, we preclude an analysis of trust and safety datasets relevant to Indic languages.' It also flags that most datasets are 'a direct translation of existing English datasets which is a serious limitation.' Those are acknowledged assumptions about external validity, not circular reasoning. There are no load-bearing self-citations: the authors cite external benchmark papers, and their own prior work is not invoked to justify the central claim. The recommendation's dependence on the untested transfer from general NLU/NLG performance to safety readiness is a weakness in the argument's evidence base, but the claim is not definitionally equivalent to its inputs. Thus the appropriate finding is no significant circularity (score 0).

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central recommendation rests on several unverified domain assumptions: comparability of heterogeneous evaluation results, reliability of developer capability claims, and the proxy from general performance to safety readiness. No free parameters or invented entities are introduced.

assumptions (5)
  • domain assumption Performance labels in Table 4 are comparable across heterogeneous evaluation studies, datasets, models, and tasks.
    Section 4.1 aggregates 14 papers into ordinal high/medium/low buckets without adjusting for differences in task difficulty, dataset composition, or model version.
  • domain assumption Developer claims of Indic language support, from technical reports, blogs, and model cards, reflect actual usable capability.
    Section 2 states that support was assessed from descriptions and benchmarks rather than from independent behavior testing by the authors.
  • domain assumption General NLU/NLG benchmark performance is a valid proxy for readiness for safety benchmarking.
    Section 4 explicitly excludes safety datasets; Section 5 uses Table 4's understanding/generation rankings to recommend languages for a safety benchmark.
  • domain assumption Speaker counts from the Census of India 2011 are an appropriate measure of real-world language exposure for prioritization.
    Section 5 compares model performance against speaker counts to argue for including the top five languages, treating speaker count as a proxy for potential safety impact.
  • domain assumption Public corpus distributions from Wikipedia, Common Crawl, and C4 accurately indicate the Indic language data available to LLM developers.
    Section 3.2 uses these distributions to infer relative model capabilities, but many model training datasets are proprietary or undisclosed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Analysis of Indic Language Capabilities in LLMs." pith.science (2026). https://pith.science/paper/BRKD2IMH

@misc{pith2026250113912,
  author       = {Pith},
  title        = {Pith review of: Analysis of Indic Language Capabilities in LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BRKD2IMH}},
  note         = {Machine review of arXiv:2501.13912}
}
read the original abstract

This report evaluates the performance of text-in text-out Large Language Models (LLMs) to understand and generate Indic languages. This evaluation is used to identify and prioritize Indic languages suited for inclusion in safety benchmarks. We conduct this study by reviewing existing evaluation studies and datasets; and a set of twenty-eight LLMs that support Indic languages. We analyze the LLMs on the basis of the training data, license for model and data, type of access and model developers. We also compare Indic language performance across evaluation datasets and find that significant performance disparities in performance across Indic languages. Hindi is the most widely represented language in models. While model performance roughly correlates with number of speakers for the top five languages, the assessment after that varies.

Figures

Figures reproduced from arXiv: 2501.13912 by the authors.

Figure 1
Figure 1. presents a structured overview of the models, datasets, and evaluation methods considered in our analysis. We conclude by contrasting LLM capabilities in Indic languages with real-world Indic language usage to provide suggestions on how to prioritize Indian languages for inclusion in future benchmarks. Indic LLM Research LLMs Pre-Trained LLMs GPT Family (4, 4-o, 3.5-turbo) [5], Mistral [38], Gemma Family [67] Llama … view at source ↗
Figure 2
Figure 2. Evaluation Tasks the evaluation papers may differ from the 28 models analyzed in the previous section. To evaluate LLMs effectively, it is essential to assess their core abilities, specifically their capacity to understand and generate natural language text. This ability is assessed through a number of smaller tasks. In [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

81 extracted references · 33 canonical work pages

  1. [1]

    [n. d.]. Introducing Llama 3.1: Our most capable models to date — ai.meta.com. https://ai.meta.com/blog/meta-llama-3-1/

  2. [2]

    Paper 1 of 2018 on Language, Table C-16, Census of India 2011

    2011. Paper 1 of 2018 on Language, Table C-16, Census of India 2011. https://censusindia.gov.in/nada/index.php/catalog/42458/ download/46089/C-16_25062018.pdf

  3. [3]

    2022. https://community.openai.com/t/chatgpt-spanish-support/24706/3 7Hindi is already included in v1 AI safety benchmark released by MLCommons 8This gap between language speakers and model capabilities can be addressed by model developers but is out of scope for the MLCommons’ AI Safety Working Group Analysis of Indic Language Capabilities in LLMs • 9

  4. [4]

    Rajbhasha Vibhag, Ministry of Home Affairs

    2024. Rajbhasha Vibhag, Ministry of Home Affairs. https://rajbhasha.gov.in/en/languages-included-eighth-schedule-indian-constitution

  5. [5]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  6. [6]

    Parul Agarwal, Aisha Asif, Shantipriya Parida, Sambit Sekhar, Satya Ranjan Dash, and Subhadarshi Panda. 2023. Generative Chatbot Adaptation for Odia Language: A Critical Evaluation. In 2023 1st International Conference on Circuits, Power and Intelligent Systems (CCPIS). 1–7. https://doi.org/10.1109/CCPIS59145.2023.10291329

  7. [7]

    Divyanshu Aggarwal, Vivek Gupta, and Anoop Kunchukuttan. 2022. IndicXNLI: Evaluating Multilingual Inference for Indian Languages. arXiv:2204.08776 [cs.CL]

  8. [8]

    Divyanshu Aggarwal, Ashutosh Sathe, and Sunayana Sitaram. 2024. Maple: Multilingual evaluation of parameter efficient finetuning of large language models. arXiv preprint arXiv:2401.07598 (2024)

Show all 81 references
  1. [9]

    Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Maxamed Axmed, et al. 2023. Mega: Multilingual evaluation of generative ai. arXiv preprint arXiv:2303.12528 (2023)

  2. [10]

    Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Prachi Jain, Maxamed Axmed, Kalika Bali, et al. 2023. Megaverse: Benchmarking large language models across languages, modalities, models and tasks. arXiv preprint arXi...

  3. [11]

    Murtaza Ali, Sourojit Ghosh, Prerna Rao, Raveena Dhegaskar, Sophia Jawort, Alix Medler, Mengqi Shi, and Sayamindu Dasgupta

  4. [12]

    Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Jon Ander Campos, Yi Chern Tan, Kelly Marchisio, Max Bartolo, Sebastian Ruder, Acyr Locatelli, Julia Kreutzer, Nick Frosst, Aidan Gomez, Phil Blunsom, Marzieh...

  5. [13]

    Abhinand Balachandran. 2023. Tamil-llama: A new tamil language model based on llama 2. arXiv preprint arXiv:2311.05845 (2023)

  6. [14]

    Abhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, and Rifat Shahriyar. 2023. BanglaNLG and BanglaT5: Benchmarks and Resources for Evaluating Low-Resource Natural Language Generation in Bangla. In Findings of the Association for Computational Linguistics: EACL 2023, Andreas ...

  7. [15]

    Khapra, and Pratyush Kumar

    Kaushal Santosh Bhogale, Sai Sundaresan, Abhigyan Raman, Tahir Javed, Mitesh M. Khapra, and Pratyush Kumar. 2023. Vistaar: Diverse Benchmarks and Training Sets for Indian Language ASR. arXiv:2305.15386 [cs.CL]

  8. [16]

    Rishi Bommasani, Dilara Soylu, Thomas I Liao, Kathleen A Creel, and Percy Liang. 2023. Ecosystem graphs: The social footprint of foundation models. arXiv preprint arXiv:2303.15772 (2023)

  9. [17]

    Govind Choudhary. 2023. Chatgpt now speaks Hindi, Assamese, Bengali and other Indian languages! https: //www.livemint.com/technology/apps/chatgpt-now-speaks-hindi-assamese-bengali-and-other-indian-languages-heres-how- to-get-replies-in-local-languages-11687948693940.html

  10. [18]

    Monojit Choudhury and Amit Deshpande. 2021. How Linguistically Fair are Multilingual Pre-Trained Language Models?. In AAAI-21. AAAI, AAAI. https://www.microsoft.com/en-us/research/publication/how-linguistically-fair-are-multilingual-pre-trained-language- models/

  11. [19]

    A Conneau. 2019. Unsupervised cross-lingual representation learning at scale. arXiv preprint arXiv:1911.02116 (2019)

  12. [20]

    Raj Dabre, Himani Shrotriya, Anoop Kunchukuttan, Ratish Puduppully, Mitesh M Khapra, and Pratyush Kumar. 2021. IndicBART: A pre-trained model for indic natural language generation. arXiv preprint arXiv:2109.02903 (2021)

  13. [21]

    Mithun Das, Punyajoy Saha, Binny Mathew, and Animesh Mukherjee. 2022. HateCheckHIn: Evaluating Hindi Hate Speech Detection Models. arXiv:2205.00328 [cs.CL]

  14. [22]

    Paresh Dave. 2023. ChatGPT Is Cutting Non-English Languages Out of the AI Revolution. Wired (2023). https://www.wired.com/story/ chatgpt-non-english-languages-ai-revolution/

  15. [23]

    Andrew Deck. 2023. We tested ChatGPT in Bengali, Kurdish, and Tamil. it failed. https://restofworld.org/2023/chatgpt-problems- global-language-testing/

  16. [24]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 [cs.CL] https://arxiv.org/abs/1810.04805

  17. [25]

    Martin Dittus and Mark Graham. 2019. The Language Geography of Wikipedia. https://internetlanguages.org/en/numbers/wikipedia- language-geography/

  18. [27]

    Khapra, Anoop Kunchukuttan, and Pratyush Kumar

    Sumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M. Khapra, Anoop Kunchukuttan, and Pratyush Kumar. 2023. Towards Leaving No Indic Language Behind: Building Monolingual Corpora, Benchmark and Models for Indic Languages. 10 • Aatman Vaidya, Tarunima P...

  19. [28]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)

  20. [29]

    Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, et al. 2022. Massive: A 1m-example multilingual natural language understanding dataset with 51 typologically- diverse languages. ...

  21. [30]

    Common Crawl Foundation. [n. d.]. Statistics of Common Crawl–Distribution of Languages. https://commoncrawl.github.io/cc-crawl- statistics/plots/languages

  22. [31]

    Jay Gala, Thanmay Jayakumar, Jaavid Aktar Husain, Mohammed Safi Ur Rahman Khan, Diptesh Kanojia, Ratish Puduppully, Mitesh M Khapra, Raj Dabre, Rudra Murthy, Anoop Kunchukuttan, et al. 2024. Airavata: Introducing hindi instruction-tuned llm. arXiv preprint arXiv:2401.15006 (2024)

  23. [32]

    Rishav Hada, Varun Gumma, Mohamed Ahmed, Kalika Bali, and Sunayana Sitaram. 2024. METAL: Towards Multilingual Meta-Evaluation. arXiv preprint arXiv:2404.01667 (2024)

  24. [33]

    Barry Haddow and Faheem Kirefu. 2020. PMIndia – A Collection of Parallel Corpora of Languages of India. arXiv:2001.09907 [cs.CL]

  25. [34]

    Sohel Rahman, and Rifat Shahriyar

    Tahmid Hasan, Abhik Bhattacharjee, Md Saiful Islam, Kazi Samin, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar

  26. [35]

    Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023. How good are GPT models at machine translation? A comprehensive evaluation. arXiv.arXiv preprint arXiv:2302.09210 (2023)

  27. [36]

    Carolin Holtermann, Paul Röttger, Timm Dill, and Anne Lauscher. 2024. Evaluating the Elementary Multilingual Capabilities of Large Language Models with MultiQ. arXiv:2403.03814 [cs.CL]

  28. [37]

    Kushal Jain, Adwait Deshpande, Kumar Shridhar, Felix Laumann, and Ayushman Dash. 2020. Indic-Transformers: An Analysis of Transformer Language Models for Indian Languages. arXiv:2011.02323 [cs.CL] https://arxiv.org/abs/2011.02323

  29. [38]

    Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7B. arXiv preprint arXiv:2310.06825 (2023)

  30. [39]

    Zhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding, and Graham Neubig. 2020. X-FACTR: Multilingual Factual Knowledge Retrieval from Pretrained Language Models. arXiv:2010.06189 [cs.CL]

  31. [40]

    Mohsinul Kabir, Mohammed Saidul Islam, Md Tahmid Rahman Laskar, Mir Tafseer Nayeem, M Saiful Bari, and Enamul Hoque

  32. [41]

    Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, NC Gokul, Avik Bhattacharyya, Mitesh M Khapra, and Pratyush Kumar. 2020. IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages. In Findings of the Associa...

  33. [42]

    Nora Kassner, Philipp Dufter, and Hinrich Schütze. 2021. Multilingual LAMA: Investigating Knowledge in Multilingual Pretrained Language Models. arXiv:2102.00894 [cs.CL]

  34. [43]

    Simran Khanuja, Diksha Bansal, Sarvesh Mehtani, Savya Khosla, Atreyee Dey, Balaji Gopalan, Dilip Kumar Margam, Pooja Aggarwal, Rajiv Teja Nagipogu, Shachi Dave, et al. 2021. Muril: Multilingual representations for indian languages. arXiv preprint arXiv:2103.10730 (2021)

  35. [44]

    Simran Khanuja, Sandipan Dandapat, Anirudh Srinivasan, Sunayana Sitaram, and Monojit Choudhury. 2020. GLUECoS: An evaluation benchmark for code-switched NLP. arXiv preprint arXiv:2004.12376 (2020)

  36. [45]

    Sankalp KJ, Vinija Jain, Sreyoshi Bhaduri, Tamoghna Roy, and Aman Chadha. 2024. Decoding the Diversity: A Review of the Indic AI Research Landscape. arXiv:2406.09559 [cs.CL] https://arxiv.org/abs/2406.09559

  37. [46]

    Guneet S Kohli, Shantipriya Parida, Sambit Sekhar, Samirit Saha, Nipun B Nair, Parul Agarwal, Sonal Khosla, Kusumlata Patiyal, and Debasish Dhal. 2023. Building a llama2-finetuned llm for odia language utilizing domain knowledge instruction set. In Proceedings of the Third Int...

  38. [47]

    Khapra, and Pratyush Kumar

    Aman Kumar, Himani Shrotriya, Prachi Sahu, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, Amogh Mishra, Mitesh M. Khapra, and Pratyush Kumar. 2022. IndicNLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in Indic Languages. arXiv:2203.05437 [cs.CL]

  39. [48]

    Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al. 2023. Bloom: A 176b-parameter open-access multilingual language model. (2023)

  40. [49]

    Wei Qi Leong, Jian Gang Ngui, Yosephine Susanto, Hamsawardhini Rengarajan, Kengatharaiyer Sarveswaran, and William Chandra Tjhi

  41. [50]

    David McCandless. 2024. The rise of Generative AI large language models (llms) like chatgpt. https://informationisbeautiful.net/ visualizations/the-rise-of-generative-ai-large-language-models-llms-like-chatgpt/

  42. [51]

    Arnav Mhaske, Harshit Kedia, Sumanth Doddapaneni, Mitesh M Khapra, Pratyush Kumar, Rudra Murthy V, and Anoop Kunchukuttan

  43. [52]

    Shubham Mittal, Megha Sundriyal, and Preslav Nakov. 2023. Lost in Translation, Found in Spans: Identifying Claims in Multilingual Social Media. arXiv:2310.18205 [cs.CL]

  44. [53]

    arXiv:2309.06085 [cs.CL]

    BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models. arXiv:2309.06085 [cs.CL]

  45. [54]

    Gabriel Nicholas and Aliya Bhatia. 2023. Lost in Translation: Large Language Models in Non-English Content Analysis. arXiv:2306.07377 [cs.CL]

  46. [55]

    Mitodru Niyogi and Arnab Bhattacharya. 2024. Paramanu: A Family of Novel Efficient Indic Generative Foundation Language Models. arXiv preprint arXiv:2401.18034 (2024)

  47. [56]

    Millicent Ochieng, Varun Gumma, Sunayana Sitaram, Jindong Wang, Keshet Ronen, Kalika Bali, and Jacki O’Neill. 2024. Beyond Metrics: Evaluating LLMs’ Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios. arXiv preprint arXiv:2406.00343 (2024)

  48. [57]

    OpenAI. 2024. Hello GPT-4o. https://openai.com/index/hello-gpt-4o/

  49. [58]

    Rossi, and Thien Huu Nguyen

    Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen

  50. [59]

    arXiv:2309.09400 [cs.CL]

    CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages. arXiv:2309.09400 [cs.CL]

  51. [60]

    Gowtham Ramesh, Sumanth Doddapaneni, Aravinth Bheemaraj, Mayank Jobanputra, Raghavan Ak, Ajitesh Sharma, Sujit Sahoo, Harshita Diddee, Divyanshu Kakwani, Navneet Kumar, et al. 2022. Samanantar: The largest publicly available parallel corpora collection for 11 indic languages. ...

  52. [61]

    Mielke, Cibu Johny, Isin Demirsahin, and Keith Hall

    Brian Roark, Lawrence Wolf-Sonkin, Christo Kirov, Sabrina J. Mielke, Cibu Johny, Isin Demirsahin, and Keith Hall. 2020. Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset. InProceedings of the Twelfth Language Resources and Evaluation Conference...

  53. [62]

    Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021. XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation. In Proceedings of the 2021 Conference on E...

  54. [63]

    Sarvam. 2024. Sarvam 1. https://www.sarvam.ai/blogs/sarvam-1

  55. [64]

    Shantipriya Parida, Satya Ranjan Dash, Ondřej Bojar, Petr Motlicek, Priyanka Pattnaik, and Debasish Kumar Mallick. 2020. OdiEnCorp 2.0: Odia-English Parallel Corpus for Machine Translation. In Proceedings of the WILDRE5– 5th Workshop on Indian Language Data: Resources and Eval...

  56. [65]

    Ojha, Saraswati Sahoo, Satya Ranjan Dash, and Bijayalaxmi Dash

    Shantipriya Parida, Kalyanamalini Sahoo, Atul Kr. Ojha, Saraswati Sahoo, Satya Ranjan Dash, and Bijayalaxmi Dash. 2022. Universal Dependency Treebank for Odia Language. arXiv:2205.11976 [cs.CL]

  57. [66]

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023)

  58. [67]

    Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. 2024. Gemma 2: Improving open language models at a practical size. arXiv preprint arXiv:2408.00118 (2024)

  59. [68]

    Mistral Team. 2023. Mixtral of experts. https://mistral.ai/news/mixtral-of-experts/

  60. [69]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288 (2023)

  61. [70]

    Harman Singh, Nitish Gupta, Shikhar Bharadwaj, Dinesh Tewari, and Partha Talukdar. 2024. IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages. arXiv preprint arXiv:2404.16816 (2024)

  62. [71]

    Anirudh Srinivasan, Sunayana Sitaram, Tanuja Ganu, Sandipan Dandapat, Kalika Bali, and Monojit Choudhury. 2021. Predicting the Performance of Multilingual NLP Models. arXiv:2110.08875 [cs.CL]

  63. [72]

    Ishaan Watts, Varun Gumma, Aditya Yadavalli, Vivek Seshadri, Manohar Swaminathan, and Sunayana Sitaram. 2024. PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data.arXiv preprint arXiv:2406.15053 (2024)

  64. [73]

    Shijie Wu and Mark Dredze. 2020. Are All Languages Created Equal in Multilingual BERT?. In Proceedings of the 5th Workshop on Representation Learning for NLP , Spandana Gella, Johannes Welbl, Marek Rei, Fabio Petroni, Patrick Lewis, Emma Strubell, Minjoon Seo, and Hannaneh Haj...

  65. [75]

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A massively multilingual pre-trained text-to-text transformer. arXiv:2010.11934 [cs.CL] https://arxiv.org/abs/2010.11934

  66. [76]

    Aniket Vashishtha, Kabir Ahuja, and Sunayana Sitaram. 2023. On evaluating and mitigating gender biases in multilingual settings. arXiv preprint arXiv:2307.01503 (2023)

  67. [77]

    Sshubam Verma, Mohammed Safi Ur Rahman Khan, Vishwajeet Kumar, Rudra Murthy, and Jaydeep Sen. 2024. MILU: A Multi-task Indic Language Understanding Benchmark. arXiv:2411.02538 [cs.CL] https://arxiv.org/abs/2411.02538

  68. [82]

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong W...

  69. [83]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P Xing, et al. 2023. Lmsys-chat-1m: A large-scale real-world llm conversation dataset. arXiv preprint arXiv:2309.11998 (2023). A Indic LLM Table Data Tabl...

  70. [2021]

    arXiv:2106.13822 [cs.CL]

    XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages. arXiv:2106.13822 [cs.CL]

  71. [2022]

    arXiv preprint arXiv:2212.10168 (2022)

    Naamapadam: a large-scale named entity annotated data for Indic languages. arXiv preprint arXiv:2212.10168 (2022). Analysis of Indic Language Capabilities in LLMs • 11

  72. [2023]

    In Proceedings of the 2023 ACM Conference on International Computing Education Research V.1 (ICER 2023)

    Taking Stock of Concept Inventories in Computing Education: A Systematic Literature Review. In Proceedings of the 2023 ACM Conference on International Computing Education Research V.1 (ICER 2023) . ACM, 397–415. https://doi.org/10.1145/3568813.3600120

  73. [2024]

    arXiv:2309.13173 [cs.CL]

    BenLLMEval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP. arXiv:2309.13173 [cs.CL]

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.