REVIEW 1 cited by
Depth Dependence of μP Learning Rates in ReLU MLPs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Depth Dependence of μP Learning Rates in ReLU MLPs
read the original abstract
In this short note we consider random fully connected ReLU networks of width $n$ and depth $L$ equipped with a mean-field weight initialization. Our purpose is to study the dependence on $n$ and $L$ of the maximal update ($\mu$P) learning rate, the largest learning rate for which the mean squared change in pre-activations after one step of gradient descent remains uniformly bounded at large $n,L$. As in prior work on $\mu$P of Yang et. al., we find that this maximal update learning rate is independent of $n$ for all but the first and last layer weights. However, we find that it has a non-trivial dependence of $L$, scaling like $L^{-3/2}.$
Forward citations
Cited by 1 Pith paper
-
Dropout and Random Gradient Masking Are Asymptotically Equivalent in Large ResNets
As ResNets grow deep and wide with fixed dropout rate, dropout training and random-gradient-masking training converge to the same limiting dynamics, and the common masking variants collapse to one limit.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.