As ResNets grow deep and wide with fixed dropout rate, dropout training and random-gradient-masking training converge to the same limiting dynamics, and the common masking variants collapse to one limit.
Title resolution pending
1 Pith paper cite this work, alongside 36 external citations. Polarity classification is still indexing.
1
Pith paper citing it
36
external citations · OpenAlex
fields
stat.ML 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Dropout and Random Gradient Masking Are Asymptotically Equivalent in Large ResNets
As ResNets grow deep and wide with fixed dropout rate, dropout training and random-gradient-masking training converge to the same limiting dynamics, and the common masking variants collapse to one limit.