Systematic experiments show that training a vision-language model on 100 languages with only 25 to 50 percent non-English data yields strong multilingual gains, and synthetic OCR data is key for non-Latin scripts.
For each plot type, we define 50 configurations, so we have 100 plots/images in total per language
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Centurio: On Drivers of Multilingual Ability of Large Vision-Language Model
Systematic experiments show that training a vision-language model on 100 languages with only 25 to 50 percent non-English data yields strong multilingual gains, and synthetic OCR data is key for non-Latin scripts.