Pith. sign in

REVIEW 3 cited by

IndicXNLI: Evaluating Multilingual Inference for Indian Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.08776 v1 pith:34CHV6KL submitted 2022-04-19 cs.CL cs.AI

classification cs.CLcs.AI
keywords indicxnlilanguagesmodelspre-traineddatasetindicadvancesanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While Indic NLP has made rapid advances recently in terms of the availability of corpora and pre-trained models, benchmark datasets on standard NLU tasks are limited. To this end, we introduce IndicXNLI, an NLI dataset for 11 Indic languages. It has been created by high-quality machine translation of the original English XNLI dataset and our analysis attests to the quality of IndicXNLI. By finetuning different pre-trained LMs on this IndicXNLI, we analyze various cross-lingual transfer techniques with respect to the impact of the choice of language models, languages, multi-linguality, mix-language input, etc. These experiments provide us with useful insights into the behaviour of pre-trained models for a diverse set of languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A machine-translated version of MMLU-Pro in nine Indic languages is released as a benchmark, with baseline accuracy scores for multilingual LLMs.

  2. Analysis of Indic Language Capabilities in LLMs

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A desk-research review finds that LLM performance is strongest for Hindi, Bengali, Marathi, Telugu, and Tamil, and recommends prioritizing these five languages for safety benchmarks.

  3. On Importance of Layer Pruning for Smaller BERT Models and Low Resource Languages

    cs.CL 2025-01 reject novelty 3.0 of 10

    Layer-pruned MahaBERT-v2 and Google-Muril models roughly match full models on Marathi headline and paragraph classification but lose ground on document classification, and they do not always beat same-size scratch-tra...

Pith tools