Dedicated monolingual models and tokenizers for Tamil, Telugu, Kannada, and Malayalam outperform a shared multilingual model and mGPT on tokenizer efficiency and most fine-tuned tasks, but the evaluation is single-run and partly unequal-budget.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages
Dedicated monolingual models and tokenizers for Tamil, Telugu, Kannada, and Malayalam outperform a shared multilingual model and mGPT on tokenizer efficiency and most fine-tuned tasks, but the evaluation is single-run and partly unequal-budget.