Pith. sign in

REVIEW 6 cited by

Transformers without Tears: Improving the Normalization of Self-Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.05895 v2 pith:DVLKKOYY submitted 2019-10-14 cs.CL cs.LGstat.ML

classification cs.CLcs.LGstat.ML
keywords performancetrainingbleuchangesfixnormnormalizationprenormscalenorm
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We evaluate three simple, normalization-centric changes to improve Transformer training. First, we show that pre-norm residual connections (PreNorm) and smaller initializations enable warmup-free, validation-based training with large learning rates. Second, we propose $\ell_2$ normalization with a single scale parameter (ScaleNorm) for faster training and better performance. Finally, we reaffirm the effectiveness of normalizing word embeddings to a fixed length (FixNorm). On five low-resource translation pairs from TED Talks-based corpora, these changes always converge, giving an average +1.1 BLEU over state-of-the-art bilingual baselines and a new 32.8 BLEU on IWSLT'15 English-Vietnamese. We observe sharper performance curves, more consistent gradient norms, and a linear relationship between activation scaling and decoder depth. Surprisingly, in the high-resource setting (WMT'14 English-German), ScaleNorm and FixNorm remain competitive but PreNorm degrades performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Long-Term Embeddings for Balanced Personalization

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    Long-Term Embeddings anchor sequential recommendation models to fixed content-based item representations to capture stable preferences and ensure version compatibility, resulting in uplifts in user engagement and fina...

  2. GPT-NeoX-20B: An Open-Source Autoregressive Language Model

    cs.CL 2022-04 accept novelty 6.0 of 10

    GPT-NeoX-20B is a publicly released 20B parameter autoregressive language model trained on the Pile that shows strong gains in five-shot reasoning over similarly sized prior models.

  3. Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks

    cs.LG 2025-12 reject novelty 5.0 of 10

    A small complex-valued spectral model (PRISM) and a hybrid Wave-Particle Transformer are claimed to show that phase-based interference is a sufficient reasoning primitive, with a 4.94 vs 5.28 perplexity win on WikiTex...

  4. Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

    cs.CL 2022-01 unverdicted novelty 5.0 of 10

    Trained the largest monolithic 530B-parameter transformer language model to date and reported new state-of-the-art zero- and few-shot results on multiple NLP benchmarks.

  5. UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning

    cs.LG 2025-08 conditional novelty 4.0 of 10

    A redesigned memory-layer architecture with five engineering improvements reaches performance parity with 8-expert MoE at similar compute, with lower memory access and stronger long-context memorization.

  6. A Comprehensive Overview of Large Language Models

    cs.CL 2023-07 unverdicted novelty 2.0 of 10

    A survey paper providing an overview of Large Language Models, their background, and recent advances in the field.

Pith tools