Pith. sign in

REVIEW 16 cited by

IndicTrans2: Towards High-Quality and Accessible Machine Translation Models for all 22 Scheduled Indian Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.16307 v3 pith:RBBNZTDF submitted 2023-05-25 cs.CL

classification cs.CL
keywords languagesmodelsindiaworkparallelscheduledtranslationbenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

India has a rich linguistic landscape with languages from 4 major language families spoken by over a billion people. 22 of these languages are listed in the Constitution of India (referred to as scheduled languages) are the focus of this work. Given the linguistic diversity, high-quality and accessible Machine Translation (MT) systems are essential in a country like India. Prior to this work, there was (i) no parallel training data spanning all 22 languages, (ii) no robust benchmarks covering all these languages and containing content relevant to India, and (iii) no existing translation models which support all the 22 scheduled languages of India. In this work, we aim to address this gap by focusing on the missing pieces required for enabling wide, easy, and open access to good machine translation systems for all 22 scheduled Indian languages. We identify four key areas of improvement: curating and creating larger training datasets, creating diverse and high-quality benchmarks, training multilingual models, and releasing models with open access. Our first contribution is the release of the Bharat Parallel Corpus Collection (BPCC), the largest publicly available parallel corpora for Indic languages. BPCC contains a total of 230M bitext pairs, of which a total of 126M were newly added, including 644K manually translated sentence pairs created as part of this work. Our second contribution is the release of the first n-way parallel benchmark covering all 22 Indian languages, featuring diverse domains, Indian-origin content, and source-original test sets. Next, we present IndicTrans2, the first model to support all 22 languages, surpassing existing models on multiple existing and new benchmarks created as a part of this work. Lastly, to promote accessibility and collaboration, we release our models and associated data with permissive licenses at https://github.com/AI4Bharat/IndicTrans2.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 37 citations worldwide. Full citation record

  1. Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Dedicated monolingual models and tokenizers for Tamil, Telugu, Kannada, and Malayalam outperform a shared multilingual model and mGPT on tokenizer efficiency and most fine-tuned tasks, but the evaluation is single-run...

  2. Andha-Dhun: A First Look at Audio Descriptions in Hindi

    cs.CV 2026-07 conditional novelty 6.0 of 10

    The paper introduces Andha-Dhun, the first Hindi audio description dataset, and shows that direct generation from dense captions outperforms translation of English ADs, while machine translation fails to resolve cultu...

  3. Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages

    cs.CL 2025-10 conditional novelty 6.0 of 10

    LLM-as-judge metrics (DeepSeek-V3 most of all) correlate best with human ratings across six Indian languages, though all segment-level correlations are low and many differences lack confidence intervals.

  4. Analyzing the Effect of Linguistic Similarity on Cross-Lingual Transfer: Tasks and Experimental Setups Matter

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Cross-lingual transfer success is best predicted by syntactic similarity for POS tagging and parsing, by trigram overlap for n-gram topic models, and by mBERT pretraining coverage for mBERT-based topic models.

  5. PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from Related Example Banks

    cs.CL 2024-12 conditional novelty 6.0 of 10

    PromptRefine uses alternating minimization over language-specific retrievers plus diversity-aware DPP fine-tuning to select cross-lingual in-context examples, improving few-shot generation in low-resource Indic languages.

  6. Neural Machine Translation for Low-Resource Tangkhul--English

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    Fine-tuning ByT5-large on 38,336 Tangkhul–English sentence pairs yields BLEU 39.97, the first reported MT system for Tangkhul.

  7. Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women

    eess.AS 2025-06 conditional novelty 5.0 of 10

    Using 25-30 seconds of audio per speaker from 100 rural Bhojpuri women, synthetic speech augmentation cuts ASR word error on the new SRUTI benchmark by 4.7 points.

  8. On the effective transfer of knowledge from English to Hindi Wikipedia

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A retrieval, neutralization, and machine-translation pipeline can add relevant factual text to Hindi Wikipedia biography sections, but the claimed 65% and 62% gains are not fully supported by the reported evaluation.

  9. HITSZ's End-To-End Speech Translation Systems Combining Sequence-to-Sequence Auto Speech Recognition Model and Indic Large Language Model for IWSLT 2025 in Indic Track

    cs.CL 2025-07 conditional novelty 4.0 of 10

    An end-to-end Whisper-plus-Krutrim speech translation system for English and Hindi, Bengali, and Tamil is evaluated on IWSLT 2025, with chain-of-thought fine-tuning showing large but selectively measured BLEU gains.

  10. CycleDistill: Bootstrapping Machine Translation using LLMs with Cyclical Distillation

    cs.CL 2025-06 conditional novelty 4.0 of 10

    CycleDistill improves low-resource Indic-to-English translation by iteratively fine-tuning an LLM on synthetic parallel data it generated itself from monolingual sources, gaining 20-30 chrF points over a few-shot base...

  11. Evaluating Machine Translation Models for English-Hindi Language Pairs: A Comparative Analysis

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Google Translate outperforms NLLB-200, OPUS-MT, and IndicTrans2 on English-Hindi translation on 18,000+ parallel sentences and a 400-question FAQ corpus, with all models degrading as sentence length grows.

  12. IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding

    cs.CL 2025-01 conditional novelty 4.0 of 10

    A machine-translated version of MMLU-Pro in nine Indic languages is released as a benchmark, with baseline accuracy scores for multilingual LLMs.

  13. Domain-adaptative Continual Learning for Low-resource Tasks: Evaluation on Nepali

    cs.CL 2024-12 conditional novelty 4.0 of 10

    Continual pretraining of Llama 3 8B on synthetic Nepali-English data improves its Nepali generation but causes English forgetting, with limited evidence for latent retention.

  14. Enabling Low-Resource Language Retrieval: Establishing Baselines for Urdu MS MARCO

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A machine-translated Urdu MS MARCO dataset and fine-tuned mT5 reranker achieve MRR@10 0.248 and Recall@10 0.438, outperforming zero-shot baselines and providing first baselines for Urdu IR.

  15. BhashaVerse : Translation Ecosystem for Indian Subcontinent Languages

    cs.CL 2024-12 reject novelty 4.0 of 10

    A claimed 2B-parameter multi-task translation model for 36 Indian languages, built from pivoted and synthetic corpora, evaluated without baselines and with inconsistent reported numbers.

  16. A Review of the Marathi Natural Language Processing

    cs.CL 2024-12 unverdicted

    A literature review of Marathi NLP resources, models, and evaluation metrics, with no new experiments or data.

Pith tools