Pith. sign in

REVIEW 14 cited by

ReZero is All You Need: Fast Convergence at Large Depth

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2003.04887 v2 pith:F3ZUAXMO submitted 2020-03-10 cs.LG cs.CLstat.ML

classification cs.LGcs.CLstat.ML
keywords convergencedeeparchitecturedynamicalfastisometrylayernetworks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep networks often suffer from vanishing or exploding gradients due to inefficient signal propagation, leading to long training times or convergence difficulties. Various architecture designs, sophisticated residual-style networks, and initialization schemes have been shown to improve deep signal propagation. Recently, Pennington et al. used free probability theory to show that dynamical isometry plays an integral role in efficient deep learning. We show that the simplest architecture change of gating each residual connection using a single zero-initialized parameter satisfies initial dynamical isometry and outperforms more complex approaches. Although much simpler than its predecessors, this gate enables training thousands of fully connected layers with fast convergence and better test performance for ResNets trained on CIFAR-10. We apply this technique to language modeling and find that we can easily train 120-layer Transformers. When applied to 12 layer Transformers, it converges 56% faster on enwiki8.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Controlled Study of Attention-Only Transformers

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Attention-only transformers match standard transformers within 0.006 nats of loss at matched parameter count, with the residual gap localized to low-context parametric recall.

  2. TFM-Retouche: A Lightweight Input-Space Adapter for Tabular Foundation Models

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    TFM-Retouche is an architecture-agnostic input-space residual adapter that improves tabular foundation model accuracy on 51 datasets by learning input corrections through the frozen backbone, with an identity guard to...

  3. PermaFrost-Attack: Stealth Pretraining Seeding(SPS) for planting Logic Landmines During LLM Training

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    Stealth Pretraining Seeding plants persistent unsafe behaviors in LLMs via diffuse poisoned web content that activates on precise triggers and evades standard evaluation.

  4. Deep learning-based phase-field modelling of brittle fracture in anisotropic media

    physics.comp-ph 2026-03 unverdicted novelty 7.0 of 10

    A variational physics-informed neural network solves higher-order anisotropic phase-field fracture models by minimizing total energy with B-spline enriched trial functions.

  5. Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization

    cs.LG 2026-07 conditional novelty 6.0 of 10

    SEL weight reparameterization reaches matched OpenWebText validation loss in 1.32–1.49× fewer transformer steps via a sign-aware exponential-linear map and mismatched initialization.

  6. NNNN: Neural Networks for Newtonian Noise Mitigation at the Einstein Telescope

    astro-ph.IM 2026-06 unverdicted novelty 6.0 of 10

    Convolutional and graph neural networks outperform the Wiener filter by factors of 15-80 in predicting Newtonian noise from single seismic events on synthetic seismometer array data.

  7. TFM-Retouche: A Lightweight Input-Space Adapter for Tabular Foundation Models

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    TFM-Retouche is an input-space residual adapter that lifts TabICLv2 performance by 56 Elo points on 51 tabular datasets while remaining architecture-agnostic and computationally light.

  8. Prototype Transformer: Towards Language Model Architectures Interpretable by Design

    cs.AI 2026-02 conditional novelty 6.0 of 10

    ProtoT is an autoregressive language model whose attention is replaced by learned prototype channels that are claimed to capture nameable concepts and allow targeted edits, at linear sequence cost but with slightly lo...

  9. Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers

    cs.LG 2026-02 unverdicted novelty 6.0 of 10

    TaperNorm gradually removes internal normalization in pre-norm transformers via learned gates that reach zero, revealing final norm as a scale anchor and enabling up to 1.18x faster KV-cached decoding with small loss ...

  10. Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization

    cs.LG 2026-07 unverdicted novelty 5.0 of 10

    Exponential-linear weight reparameterization (ELWR) reaches matched OpenWebText transformer validation loss in 1.32–1.49× fewer steps than linear weights.

  11. Review Residuals: Update-Conditioned Residual Gating for Transformers

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    Review Residuals add an update-conditioned gate to transformer residual connections, yielding depth-stable training and performance gains that emerge and grow with model size from 590M parameters upward.

  12. HAARES Half-Split Residual Basis Routing for Deep Transformers

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    HAARES is a lightweight residual basis router that augments block summaries with an RMS-matched half-split detail vector and reports consistent gains over Block AttnRes in 48-layer 201M models across language modeling...

  13. Attention Residuals

    cs.CL 2026-03 unverdicted novelty 5.0 of 10

    Attention Residuals replaces fixed residual summation with input-dependent softmax attention over preceding layers, and a blocked variant is shown to improve uniformity and downstream performance in a 48B-parameter mo...

  14. NRR-Core: Non-Resolution Reasoning as a Computational Framework for Contextual Identity and Ambiguity Preservation

    cs.CL 2025-12 conditional novelty 3.0 of 10

    A gated two-embedding toy model can output near-maximum uncertainty before context and resolve perfectly afterward, but the uncertainty is enforced by a hand-set gate.

Pith tools