Plasticity loss in GPT-style transformers on multilingual tasks persists from 5M to 314M parameters, follows a sublinear scaling law with model size, and occurs in both continual and stationary settings.
arXiv preprint arXiv:2109.00267 , year=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Can Scale Save Us From Plasticity Loss in Large Language Models?
Plasticity loss in GPT-style transformers on multilingual tasks persists from 5M to 314M parameters, follows a sublinear scaling law with model size, and occurs in both continual and stationary settings.