A multilevel scheme that alternates fine transformer training with two half-depth coarse models reaches the single-level training loss with 44 percent fewer FLOPs on one small language-model setup.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A multilevel approach to accelerate the training of Transformers
A multilevel scheme that alternates fine transformer training with two half-depth coarse models reaches the single-level training loss with 44 percent fewer FLOPs on one small language-model setup.