REVIEW 3 cited by
Breaking the Curse of Multilinguality with Cross-lingual Expert Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Despite their popularity in non-English NLP, multilingual language models often underperform monolingual ones due to inter-language competition for model parameters. We propose Cross-lingual Expert Language Models (X-ELM), which mitigate this competition by independently training language models on subsets of the multilingual corpus. This process specializes X-ELMs to different languages while remaining effective as a multilingual ensemble. Our experiments show that when given the same compute budget, X-ELM outperforms jointly trained multilingual models across all considered languages and that these gains transfer to downstream tasks. X-ELM provides additional benefits over performance improvements: new experts can be iteratively added, adapting X-ELM to new languages without catastrophic forgetting. Furthermore, training is asynchronous, reducing the hardware requirements for multilingual training and democratizing multilingual modeling.
Forward citations
Cited by 3 Pith papers
-
On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
Arab cultural entities that double as everyday Arabic words are harder for language models to recognize, especially when tokenized as single tokens.
-
Soro: A Lightweight Foundation Model and Chatbot for Tajik
Tajik-specialized Gemma 3 derivatives (12B/27B) beat same-size baselines by ~6–8 points on new Tajik exams after 1.9B-token continual pretraining, with FP8/INT4 still usable on edge GPUs.
-
Salamandra Technical Report
Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.
Discussion (0). Continue with ORCID to comment.