Pith. sign in

REVIEW 19 cited by

On Layer Normalization in the Transformer Architecture

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.04745 v2 pith:SAPHDR3R submitted 2020-02-12 cs.LG cs.CLstat.ML

classification cs.LGcs.CLstat.ML
keywords layerstagetransformerwarm-upnormalizationgradientslearningpre-ln
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The Transformer is widely used in natural language processing tasks. To train a Transformer however, one usually needs a carefully designed learning rate warm-up stage, which is shown to be crucial to the final performance but will slow down the optimization and bring more hyper-parameter tunings. In this paper, we first study theoretically why the learning rate warm-up stage is essential and show that the location of layer normalization matters. Specifically, we prove with mean field theory that at initialization, for the original-designed Post-LN Transformer, which places the layer normalization between the residual blocks, the expected gradients of the parameters near the output layer are large. Therefore, using a large learning rate on those gradients makes the training unstable. The warm-up stage is practically helpful for avoiding this problem. On the other hand, our theory also shows that if the layer normalization is put inside the residual blocks (recently proposed as Pre-LN Transformer), the gradients are well-behaved at initialization. This motivates us to remove the warm-up stage for the training of Pre-LN Transformers. We show in our experiments that Pre-LN Transformers without the warm-up stage can reach comparable results with baselines while requiring significantly less training time and hyper-parameter tuning on a wide range of applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 110 citations worldwide. Full citation record

  1. Detangled: A Framework for Creating, Editing, and Inferencing Feature Rich Hair Strands

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A 5D texture parameterization plus centerline-based canonical space and supervised diffusion enables generation and texture transfer of feature-rich hair strands independent of style.

  2. NAE: Normalizing AutoEncoder

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A conditional surrogate loss that always picks the gradient estimate aligned with the reconstruction loss improves flow autoencoder training and reaches state-of-the-art generative performance on molecules, tabular da...

  3. TauPolaris: reconstructing tau lepton polarimetric vectors with conditional normalizing flows

    hep-ph 2026-08 conditional novelty 6.0 of 10

    A conditional normalizing flow reconstructs tau polarimetric vectors from simulated LHC events, improving spin-observable resolution by about 40% over a regression baseline and projecting >4.3 sigma entanglement separ...

  4. AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model

    astro-ph.GA 2026-07 conditional novelty 6.0 of 10

    An uncertainty-aware transformer reconstructs masked AGN broad lines and spectral halves with 4-16% flux errors and beats eleven purpose-built Lyα-reconstruction algorithms on a blind benchmark.

  5. D$e^+e^-$ffusion: Capturing the Beam-Beam Physics of $e^+e^-$ Collisions with Diffusion Models

    hep-ph 2026-07 conditional novelty 6.0 of 10

    A diffusion model trained on GuineaPig++ reproduces FCC-ee beam-induced pair-production distributions at particle and detector level, about 10^4 times faster.

  6. Transformer-based machine learning using low-level calorimeter signals for collimated photon identification at collider experiments

    hep-ph 2026-07 accept novelty 6.0 of 10

    Cell-level Transformers classify collimated ALP photon-jets versus single photons with AUC 0.98 and regress diphoton mass to ~64 MeV, beating shower-shape and other ML baselines in an ATLAS-like GEANT4 simulation.

  7. Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Splitting weight matrices into a fixed-norm direction and learnable per-row/column magnitudes improves LLM training over AdamW/Muon, removes weight decay and warmup, and transfers the optimal LR across width.

  8. Numerical Fragility in Transformers: A Layer-wise Theory for Risk Estimation and Selective Stabilization

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Layer-wise first-order diagnostics (κsoftmax, κscore, κ(V), ρLN) forecast and localize low-precision forward mismatch in a Tiny-ViT, with a modest LayerNorm-ε stabilization effect.

  9. Guard Vector: Beyond English LLM Guardrails with Task-Vector Composition and Streaming-Aware Prefix SFT

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A task vector from an English guard model transfers safety classification to Korean, Chinese, and Japanese models, and a prefix-SFT variant maintains accuracy under streaming with a single-token classifier.

  10. Layerwise Importance Analysis of Feed-Forward Networks in Transformer-based Language Models

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Concentrating all feed-forward-network capacity into the middle 70% of a Transformer's layers, at fixed total parameter count, outperforms the standard uniform layout across model sizes and language tasks.

  11. A CNN-Transformer for Classification of Longitudinal 3D MRI Images -- A Case Study on Hepatocellular Carcinoma Prediction

    cs.CV 2025-01 reject novelty 6.0 of 10

    A CNN-Transformer trained on longitudinal 3D MRIs claims high accuracy for predicting next-scan hepatocellular carcinoma, but its time-aware positional encoding reveals the future diagnosis date to the model.

  12. STST-JEPA: Shallow-Target Spatio-Temporal Joint Embedding Prediction Architecture For EEG Self-Supervised Learning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A JEPA-style EEG foundation model with shallow EMA targets plus light reconstruction reaches strong multi-task transfer and 3.06-year validation age MAE on a large multi-site corpus.

  13. CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis

    eess.AS 2025-08 conditional novelty 5.0 of 10

    CLEAR is a zero-shot TTS model that autoregressively predicts compact continuous audio latents with a per-token rectified flow head, reaching 1.88% WER on LibriSpeech Subset-B with an RTF of 0.29 and a 96 ms streaming delay.

  14. A Multimodal Architecture for Endpoint Position Prediction in Team-based Multiplayer Games

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A U-Net++ with multimodal encoders and agent attention improves World of Tanks endpoint prediction, with KL divergence loss and rendered icons giving the best relative FDE at 1.78.

  15. PyViT-FUSE: A Foundation Model for Multi-Sensor Earth Observation Data

    cs.CV 2025-04 conditional novelty 5.0 of 10

    PyViT-FUSE is a self-supervised vision transformer that fuses an arbitrary set of heterogeneous-resolution satellite bands via attention, and shows promise on a solar-panel segmentation task.

  16. CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling

    cs.CV 2025-02 conditional novelty 5.0 of 10

    CLIP-UP converts a pre-trained dense CLIP into an MoE model and improves zero-shot text-image retrieval beyond dense baselines at lower inference cost.

  17. Vision Transformers for Weakly-Supervised Microorganism Enumeration

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Vision transformers are competitive but not superior to ResNets for weakly-supervised microorganism counting when trained from scratch.

  18. Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization

    cs.LG 2025-06 reject novelty 4.0 of 10

    Wavelet-based positional encodings are claimed to improve how transformers extrapolate to longer sequences, with a toy experiment supporting the claim but with weak theory.

  19. Technical Report: Small Language Model for Japanese Clinical and Medicine

    cs.CL 2024-12 conditional novelty 4.0 of 10

    Fine-tuned 1.2B Japanese medical SLM tops 6 of 8 JMED-LLM tasks against larger models, but the comparison is confounded by benchmark-specific fine-tuning.

Pith tools