Pith. sign in

REVIEW 1 cited by

The exploding gradient problem demystified - definition, prevalence, impact, origin, tradeoffs, and solutions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1712.05577 v4 pith:IIFOKVCD submitted 2017-12-15 cs.LG cs.CV

classification cs.LGcs.CV
keywords explodinggradientsproblemgradientnetworkarchitecturesnetworksresidual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Whereas it is believed that techniques such as Adam, batch normalization and, more recently, SeLU nonlinearities "solve" the exploding gradient problem, we show that this is not the case in general and that in a range of popular MLP architectures, exploding gradients exist and that they limit the depth to which networks can be effectively trained, both in theory and in practice. We explain why exploding gradients occur and highlight the *collapsing domain problem*, which can arise in architectures that avoid exploding gradients. ResNets have significantly lower gradients and thus can circumvent the exploding gradient problem, enabling the effective training of much deeper networks. We show this is a direct consequence of the Pythagorean equation. By noticing that *any neural network is a residual network*, we devise the *residual trick*, which reveals that introducing skip connections simplifies the network mathematically, and that this simplicity may be the major cause for their success.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Mean-Field Model for Two-Layer Neural Networks Trained with Consensus-Based Optimization

    cs.LG 2025-11 conditional novelty 5.0 of 10

    CBO can train small two-layer networks, a hybrid CBO-Adam method improves convergence and stability, and a Wasserstein mean-field model of CBO has monotonically decreasing variance.

Pith tools