µnit Scaling combines unit-variance initialization, rearranged LayerNorm, a square-root softmax analysis, and µ-Parametrization-style learning-rate rules to train 1B-13B LLMs in FP8 with no dynamic scaling and zero-shot hyperparameter transfer.
Unit scaling: Out-of-the-box low-precision training
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
$\mu$nit Scaling: Simple and Scalable FP8 LLM Training
µnit Scaling combines unit-variance initialization, rearranged LayerNorm, a square-root softmax analysis, and µ-Parametrization-style learning-rate rules to train 1B-13B LLMs in FP8 with no dynamic scaling and zero-shot hyperparameter transfer.