Entropy2Vec turns the cross-lingual surprise of monolingual language models into dense language embeddings that resemble typological families and match curated vectors in downstream tasks.
LinguAlchemy: Fusing Typological and Geographical Elements for Unseen Language Generalization
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Pretrained language models (PLMs) have become remarkably adept at task and language generalization. Nonetheless, they often fail when faced with unseen languages. In this work, we present LinguAlchemy, a regularization method that incorporates various linguistic information covering typological, geographical, and phylogenetic features to align PLMs representation to the corresponding linguistic information on each language. Our LinguAlchemy significantly improves the performance of mBERT and XLM-R on low-resource languages in multiple downstream tasks such as intent classification, news classification, and semantic relatedness compared to fully finetuned models and displaying a high degree of unseen language generalization. We further introduce AlchemyScale and AlchemyTune, extension of LinguAlchemy which adjusts the linguistic regularization weights automatically, alleviating the need for hyperparameter search.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations
Entropy2Vec turns the cross-lingual surprise of monolingual language models into dense language embeddings that resemble typological families and match curated vectors in downstream tasks.