Optimal depth-wise learning-rate scaling in deep scalar linear networks is data-dependent, so data-agnostic rules fail to transfer while the data-aware rule yields depth-independent linear convergence.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2representative citing papers
A mechanics of the learning process is emerging in deep learning theory, characterized by dynamics, coarse statistics, and falsifiable predictions across idealized settings, limits, laws, hyperparameters, and universal behaviors.
citing papers explorer
-
Optimal Learning Rate Scaling Depends on Data in Deep Scalar Linear Networks
Optimal depth-wise learning-rate scaling in deep scalar linear networks is data-dependent, so data-agnostic rules fail to transfer while the data-aware rule yields depth-independent linear convergence.
-
There Will Be a Scientific Theory of Deep Learning
A mechanics of the learning process is emerging in deep learning theory, characterized by dynamics, coarse statistics, and falsifiable predictions across idealized settings, limits, laws, hyperparameters, and universal behaviors.