Gradient descent from moderately small random initialization jointly trains both layers of a ReLU network and converges linearly to the global minimizer for a linear target at order-wise optimal sample complexity by progressing through alignment, growth, and refinement phases.
Feature averaging: An implicit bias of gradient descent leading to non-robustness in neural networks.arXiv preprint arXiv:2410.10322, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks
Gradient descent from moderately small random initialization jointly trains both layers of a ReLU network and converges linearly to the global minimizer for a linear target at order-wise optimal sample complexity by progressing through alignment, growth, and refinement phases.