An nGPT training recipe with new optimizer and normalization components cuts pretraining tokens roughly in half for a 14B hybrid mixture-of-experts language model compared with an AdamW baseline.
F or each coordinate with dt,i > 0, γt,i = 1 1 + exp ( −a log dt,i ǫgate ) (25) = 1 1 + (ǫgate/dt,i)a (26) = d a t,i d a t,i +ǫ a gate
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Training nGPT
An nGPT training recipe with new optimizer and normalization components cuts pretraining tokens roughly in half for a 14B hybrid mixture-of-experts language model compared with an AdamW baseline.