Aya-23-8B appears to activate multiple related languages internally and concentrate code-mixing neurons in final layers, but the paper's own limitations undercut the claim that these are language-specific neurons.
Language-specific Neurons Do Not Facilitate Cross-Lingual Transfer
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Multilingual large language models (LLMs) aim towards robust natural language understanding across diverse languages, yet their performance significantly degrades on low-resource languages. This work explores whether existing techniques to identify language-specific neurons can be leveraged to enhance cross-lingual task performance of lowresource languages. We conduct detailed experiments covering existing language-specific neuron identification techniques (such as Language Activation Probability Entropy and activation probability-based thresholding) and neuron-specific LoRA fine-tuning with models like Llama 3.1 and Mistral Nemo. We find that such neuron-specific interventions are insufficient to yield cross-lingual improvements on downstream tasks (XNLI, XQuAD) in lowresource languages. This study highlights the challenges in achieving cross-lingual generalization and provides critical insights for multilingual LLMs.
fields
cs.CL 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
What Language(s) Does Aya-23 Think In? How Multilinguality Affects Internal Language Representations
Aya-23-8B appears to activate multiple related languages internally and concentrate code-mixing neurons in final layers, but the paper's own limitations undercut the claim that these are language-specific neurons.