Incremental layer-wise training of a 124M-parameter GPT-2 model underperforms standard full-layer training at equal computational cost and catches up only after extra continual training.
Greedy layer- wise training of deep networks,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
On the Effectiveness of Incremental Training of Large Language Models
Incremental layer-wise training of a 124M-parameter GPT-2 model underperforms standard full-layer training at equal computational cost and catches up only after extra continual training.