REVIEW 16 cited by
IndicTrans2: Towards High-Quality and Accessible Machine Translation Models for all 22 Scheduled Indian Languages
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
India has a rich linguistic landscape with languages from 4 major language families spoken by over a billion people. 22 of these languages are listed in the Constitution of India (referred to as scheduled languages) are the focus of this work. Given the linguistic diversity, high-quality and accessible Machine Translation (MT) systems are essential in a country like India. Prior to this work, there was (i) no parallel training data spanning all 22 languages, (ii) no robust benchmarks covering all these languages and containing content relevant to India, and (iii) no existing translation models which support all the 22 scheduled languages of India. In this work, we aim to address this gap by focusing on the missing pieces required for enabling wide, easy, and open access to good machine translation systems for all 22 scheduled Indian languages. We identify four key areas of improvement: curating and creating larger training datasets, creating diverse and high-quality benchmarks, training multilingual models, and releasing models with open access. Our first contribution is the release of the Bharat Parallel Corpus Collection (BPCC), the largest publicly available parallel corpora for Indic languages. BPCC contains a total of 230M bitext pairs, of which a total of 126M were newly added, including 644K manually translated sentence pairs created as part of this work. Our second contribution is the release of the first n-way parallel benchmark covering all 22 Indian languages, featuring diverse domains, Indian-origin content, and source-original test sets. Next, we present IndicTrans2, the first model to support all 22 languages, surpassing existing models on multiple existing and new benchmarks created as a part of this work. Lastly, to promote accessibility and collaboration, we release our models and associated data with permissive licenses at https://github.com/AI4Bharat/IndicTrans2.
Forward citations
Cited by 16 Pith papers
-
Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages
Dedicated monolingual models and tokenizers for Tamil, Telugu, Kannada, and Malayalam outperform a shared multilingual model and mGPT on tokenizer efficiency and most fine-tuned tasks, but the evaluation is single-run...
-
Andha-Dhun: A First Look at Audio Descriptions in Hindi
The paper introduces Andha-Dhun, the first Hindi audio description dataset, and shows that direct generation from dense captions outperforms translation of English ADs, while machine translation fails to resolve cultu...
-
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
LLM-as-judge metrics (DeepSeek-V3 most of all) correlate best with human ratings across six Indian languages, though all segment-level correlations are low and many differences lack confidence intervals.
-
Analyzing the Effect of Linguistic Similarity on Cross-Lingual Transfer: Tasks and Experimental Setups Matter
Cross-lingual transfer success is best predicted by syntactic similarity for POS tagging and parsing, by trigram overlap for n-gram topic models, and by mBERT pretraining coverage for mBERT-based topic models.
-
PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from Related Example Banks
PromptRefine uses alternating minimization over language-specific retrievers plus diversity-aware DPP fine-tuning to select cross-lingual in-context examples, improving few-shot generation in low-resource Indic languages.
-
Neural Machine Translation for Low-Resource Tangkhul--English
Fine-tuning ByT5-large on 38,336 Tangkhul–English sentence pairs yields BLEU 39.97, the first reported MT system for Tangkhul.
-
Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women
Using 25-30 seconds of audio per speaker from 100 rural Bhojpuri women, synthetic speech augmentation cuts ASR word error on the new SRUTI benchmark by 4.7 points.
-
On the effective transfer of knowledge from English to Hindi Wikipedia
A retrieval, neutralization, and machine-translation pipeline can add relevant factual text to Hindi Wikipedia biography sections, but the claimed 65% and 62% gains are not fully supported by the reported evaluation.
-
HITSZ's End-To-End Speech Translation Systems Combining Sequence-to-Sequence Auto Speech Recognition Model and Indic Large Language Model for IWSLT 2025 in Indic Track
An end-to-end Whisper-plus-Krutrim speech translation system for English and Hindi, Bengali, and Tamil is evaluated on IWSLT 2025, with chain-of-thought fine-tuning showing large but selectively measured BLEU gains.
-
CycleDistill: Bootstrapping Machine Translation using LLMs with Cyclical Distillation
CycleDistill improves low-resource Indic-to-English translation by iteratively fine-tuning an LLM on synthetic parallel data it generated itself from monolingual sources, gaining 20-30 chrF points over a few-shot base...
-
Evaluating Machine Translation Models for English-Hindi Language Pairs: A Comparative Analysis
Google Translate outperforms NLLB-200, OPUS-MT, and IndicTrans2 on English-Hindi translation on 18,000+ parallel sentences and a 400-question FAQ corpus, with all models degrading as sentence length grows.
-
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
A machine-translated version of MMLU-Pro in nine Indic languages is released as a benchmark, with baseline accuracy scores for multilingual LLMs.
-
Domain-adaptative Continual Learning for Low-resource Tasks: Evaluation on Nepali
Continual pretraining of Llama 3 8B on synthetic Nepali-English data improves its Nepali generation but causes English forgetting, with limited evidence for latent retention.
-
Enabling Low-Resource Language Retrieval: Establishing Baselines for Urdu MS MARCO
A machine-translated Urdu MS MARCO dataset and fine-tuned mT5 reranker achieve MRR@10 0.248 and Recall@10 0.438, outperforming zero-shot baselines and providing first baselines for Urdu IR.
-
BhashaVerse : Translation Ecosystem for Indian Subcontinent Languages
A claimed 2B-parameter multi-task translation model for 36 Indian languages, built from pivoted and synthetic corpora, evaluated without baselines and with inconsistent reported numbers.
-
A Review of the Marathi Natural Language Processing
A literature review of Marathi NLP resources, models, and evaluation metrics, with no new experiments or data.
Discussion (0). Continue with ORCID to comment.