Peri-LN, which normalizes both the input and output of each sublayer, reduces activation-variance growth and gradient spikes during LLM pretraining, outperforming Pre-LN and Post-LN at scales up to 3.2B parameters.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
Peri-LN, which normalizes both the input and output of each sublayer, reduces activation-variance growth and gradient spikes during LLM pretraining, outperforming Pre-LN and Post-LN at scales up to 3.2B parameters.