REVIEW 4 minor 300 references
Noisy SGD on wide two-layer ReLU nets collapses to at most 2P-1 effective directions set by the training data's linear dichotomies.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 01:11 UTC pith:LKR7LU7Y
load-bearing objection Solid multivariate extension of Shevchenko et al. that actually characterizes the stationary measure (width collapse + non-redundancy), with limitations already scoped correctly by the authors.
Implicit Bias of SGD in Multivariate ReLU Networks: Effective Width Collapse
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Despite infinite overparameterization, noisy SGD (vanishing temperature, positive weight decay) selects continuous piecewise-affine predictors whose kink hyperplanes number at most 2P(X)-1 and whose input-weight/bias marginal is supported on at most that many rays; each ray induces a distinct, incomparable ternary activation pattern on the training inputs.
What carries the argument
The Boltzmann fixed-point characterization of the unique free-energy minimizer of the mean-field PDE, analyzed after a bounded smooth approximation of ReLU; the Quadratic Positivity Condition and the subsequent shrinkage of a data-induced cluster set on the sphere of input weights and biases determine both the vanishing Hessian outside finitely many hyperplanes and the support of the limiting measure.
Load-bearing premise
The joint infinite-width, vanishing-noise, ReLU limit is studied only through accumulation points of approximating sequences, whose existence is not proved for the measures themselves, and the limiting weight-decay strength must stay strictly positive.
What would settle it
Train a wide two-layer ReLU net by noisy SGD with vanishing noise and positive weight decay on a fixed multivariate data set; if the trained input weights/biases fail to concentrate on far fewer than 2P-1 rays, or if the learned function has more than 2P-1 kink hyperplanes that violate the ternary-pattern non-redundancy conditions, the claim is false.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the implicit bias of noisy SGD for wide two-layer ReLU networks in multivariate regression under the mean-field regime of Mei et al. (2018). Training is approximated by a Wasserstein gradient flow of a free energy that admits a unique stationary measure satisfying a Boltzmann fixed-point condition. Working with bounded, smoothed approximations of ReLU and then taking limits, the authors characterize accumulation points of the stationary measures and the associated predictors as temperature vanishes and the truncation is removed. Theorem 4 shows that any such limiting predictor is continuous piecewise affine, with kink hyperplanes numbering at most 2P(X)-1, where P(X) is the number of linear dichotomies of the training inputs. Theorem 5 shows that the marginal of input weights and biases is supported on at most that many rays, yielding effective width collapse. Theorem 6 establishes a non-redundancy property: distinct rays induce incomparable ternary activation patterns on the training data. The proofs shift the analysis from input space to parameter space, introduce a Quadratic Positivity Condition and a cluster set on the sphere, and prove shrinkage of that set to finitely many points. Experiments illustrate alignment and piecewise affinity.
Significance. If correct, the results give the first derivation of effective width collapse as an implicit bias of noisy SGD for multivariate ReLU regression with essentially arbitrary finite training data, together with a data-dependent combinatorial bound controlled by P(X). This extends the univariate analysis of Shevchenko et al. (2022) both to higher input dimension and to a structural description of the stationary measure itself (not only the represented function). The bound is shown to improve on related static TV-norm bounds of de Dios–Bruna and Del Grande et al. for generic data. The geometric argument (cluster-set shrinkage on the sphere, non-redundancy of ternary patterns) is reusable and the appendices supply a detailed derivation chain. Code is provided. The main caveats—analysis of accumulation points rather than proven joint limits, and the requirement that limiting weight decay stay positive—are stated explicitly by the authors and do not undermine the claims as scoped.
minor comments (4)
- The existence of 2-Wasserstein convergent subsequences for the measures (needed for Theorem 5) is left more open than the pointwise existence for the predictors (Proposition 19). A short clarifying sentence in §4.2 or Appendix G would help readers see the precise gap.
- Figures 1–3 and 5–6 are informative but would benefit from a brief caption note on how the learned directions were extracted (DBSCAN thresholds, norm cutoff) so that the experimental protocol is fully self-contained.
- In the comparison with Del Grande et al. (Appendix F), the improvement factor 2^d/(d+1)^2 is stated for the regime 2d+1≤M; a one-line numerical example for small (d,M) would make the improvement more tangible.
- Notation for the cluster set (Ω̂^α vs. Ω̂*) and the sectors O_k is introduced gradually across §6; a short summary table or paragraph at the start of §6 would improve readability of the proof outline.
Circularity Check
No significant circularity: structural claims follow from the Boltzmann fixed-point of Mei et al. via independent geometric analysis of the cluster set; accumulation-point scoping is explicit and non-circular.
full rationale
The derivation chain begins from the free-energy minimizer and Boltzmann fixed-point condition of Mei et al. (2018, Theorem 1 / eq. 5), which is an external result under stated regularity assumptions on the (smoothed/truncated) activation. From there the paper derives a Hessian bound (Prop. 8), introduces the Quadratic Positivity Condition on the parameter-space sectors induced by the training inputs (Prop. 9), proves shrinkage of the associated cluster set ˆΩ* to at most 2P(X)-1 points on the sphere (Prop. 10, using only the classical count of 2P sectors from Cover 1965 and a merging argument that excludes the all-negative sector), and obtains vanishing Hessian outside the corresponding hyperplanes (Thm. 11) together with support of the input-weight/bias marginal on the rays (Thm. 5) and non-redundancy of ternary patterns (Prop. 13 / Thm. 6). None of these steps defines a quantity in terms of the final bound and recovers it, fits a free parameter to data and re-labels the fit as a prediction, or imports a uniqueness theorem from overlapping authors as an external fact that forces the claim. The restriction to accumulation points of sequences (β_n, λ_n, m_n) (rather than a proven joint limit m, β → ∞) is stated explicitly (end of §3.2, Assumption 2, Prop. 10, §5 comparison with Shevchenko et al.) and is accompanied by existence of such points (Prop. 19, Lemma 15); the geometric conclusions hold for every such accumulation point and do not rely on uniqueness of a joint limit. Self-citations are to technical lemmas of the univariate predecessor or to the authors’ own related work on other regimes; none is load-bearing for the central support/non-redundancy statements. P(X) is the classical combinatorial quantity of Cover (1965). Consequently the claimed effective-width collapse is a genuine structural consequence of the mean-field stationary condition, not a restatement of its inputs.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Noisy SGD with step-size scheme s_{k,N}=ε ξ(kε) converges, in the joint limit N→∞, ε→0, to the Wasserstein gradient flow of the free energy F_{β,λ}^σ whose unique minimizer satisfies the Boltzmann fixed-point equation (Mei et al. 2018, Thm. 1 / Lemma 10.3).
- ad hoc to paper The m-truncated, τ-smoothed ReLU neuron converges pointwise to the true ReLU as m,τ→∞ and satisfies the C^4 boundedness hypotheses of the mean-field theorem.
- ad hoc to paper Any sequence (β_n,λ_n,m_n) obeying Assumption 2 with m_n,β_n→∞ and λ_n→λ̄>0 admits a subsequence along which the network functions converge pointwise and (for Thm. 5) the measures converge in 2-Wasserstein distance.
- standard math Standard facts on hyperplane arrangements, Cover’s bound on linear dichotomies, and the geometry of quadratic cones on the sphere.
invented entities (2)
-
Cluster set Ω̂^α (and its limit Ω̂*)
no independent evidence
-
Quadratic Positivity Condition
no independent evidence
read the original abstract
We study the implicit bias of noisy stochastic gradient descent in training wide two-layer ReLU networks for multivariate regression. In a mean-field regime, the training dynamics are approximated by a Wasserstein gradient flow that converges to a unique stationary measure. We characterize the structure of this stationary measure and the predictor it represents. We show that, despite the network being infinitely overparameterized, the learned predictor admits an effectively finite representation: the input weights and biases align along finitely many directions, leading to an effective width collapse. In particular, the solution function is continuous piecewise affine, with affine regions determined by the cells of a finite hyperplane arrangement. The number of learned directions, and hence hyperplanes, is bounded above by $2\mathcal{P}-1$, where $\mathcal{P}$ denotes the number of linear dichotomies realizable on the training inputs. We further establish a non-redundancy property of the learned representation by proving that each learned direction induces a unique ternary activation pattern on the training data. Consequently, the complexity of the learned predictor is governed by the combinatorial geometry of the training data.
Figures
Reference graph
Works this paper leans on
-
[1]
Advances in neural information processing systems , volume=
The rich and the simple: On the implicit bias of adam and sgd , author=. Advances in neural information processing systems , volume=
-
[2]
Proceedings of the Second International Conference on Knowledge Discovery and Data Mining , pages=
A density-based algorithm for discovering clusters in large spatial databases with noise , author=. Proceedings of the Second International Conference on Knowledge Discovery and Data Mining , pages=
-
[3]
The Fourteenth International Conference on Learning Representations , year=
Never Saddle for Reparameterized Steepest Descent as Mirror Flow , author=. The Fourteenth International Conference on Learning Representations , year=
-
[4]
International Conference on Learning Representations , volume=
Flavors of margin: Implicit bias of steepest descent in homogeneous neural networks , author=. International Conference on Learning Representations , volume=
-
[5]
2026 , eprint=
Implicit Bias of Mirror Flow in Homogeneous Neural Networks: Sparse and Dense Feature Learning , author=. 2026 , eprint=
2026
-
[6]
On sparsity in overparametrised shallow
de Dios, Jaume and Bruna, Joan , journal=. On sparsity in overparametrised shallow
-
[7]
Global Minimizers of _p -Regularized Objectives Yield the Sparsest
Julia B Nakhleh and Robert D Nowak , booktitle=. Global Minimizers of _p -Regularized Objectives Yield the Sparsest. 2026 , url=
2026
-
[8]
Journal of Machine Learning Research , volume=
Early alignment in two-layer networks training is a two-edged sword , author=. Journal of Machine Learning Research , volume=
-
[9]
Advances in Neural Information Processing Systems , volume=
Towards understanding the condensation of neural networks at initial training , author=. Advances in Neural Information Processing Systems , volume=
-
[10]
arXiv preprint arXiv:2009.10713 , year=
Towards a mathematical understanding of neural network-based machine learning: what we know and what we don't , author=. arXiv preprint arXiv:2009.10713 , year=
Pith/arXiv arXiv 2009
-
[11]
Tensor programs
Yang, Greg and Hu, Edward J , booktitle=. Tensor programs. 2021 , organization=
2021
-
[12]
arXiv preprint arXiv:2011.14522 , year=
Feature learning in infinite-width neural networks , author=. arXiv preprint arXiv:2011.14522 , year=
Pith/arXiv arXiv 2011
-
[13]
Implicit Compressibility of Overparametrized Neural Networks Trained with Heavy-Tailed
Yijun Wan and Melih Barsbey and Abdellatif Zaidi and Umut Simsekli , booktitle=. Implicit Compressibility of Overparametrized Neural Networks Trained with Heavy-Tailed. 2024 , url=
2024
-
[14]
Journal of Machine Learning Research , volume=
Law of large numbers and central limit theorem for wide two-layer neural networks: the mini-batch and noisy case , author=. Journal of Machine Learning Research , volume=
-
[15]
SIAM Journal on Applied Mathematics , volume=
Mean field analysis of neural networks: A law of large numbers , author=. SIAM Journal on Applied Mathematics , volume=. 2020 , publisher=
2020
-
[16]
Conference on learning theory , pages=
Benefits of depth in neural networks , author=. Conference on learning theory , pages=. 2016 , organization=
2016
-
[17]
Rademacher and
Bartlett, Peter L and Mendelson, Shahar , journal=. Rademacher and
-
[18]
2026 , eprint=
Better Neural Network Expressivity: Subdividing the Simplex , author=. 2026 , eprint=
2026
-
[19]
SIAM Journal on Discrete Mathematics , volume =
Hertrich, Christoph and Basu, Amitabh and Di Summa, Marco and Skutella, Martin , title =. SIAM Journal on Discrete Mathematics , volume =. 2023 , nodoi =. https://doi.org/10.1137/22M1489332 , abstract =
-
[20]
2020 , cdate=
Henning Petzka and Martin Trimmel and Cristian Sminchisescu , title=. 2020 , cdate=
2020
-
[21]
A Complete Symmetry Classification of Shallow
Pranavkrishnan Ramakrishnan , year=. A Complete Symmetry Classification of Shallow. 2604.14037 , archivePrefix=
-
[22]
Johanna Marie Gegenfurtner and Moritz Grillo and Guido Montúfar , year=. The Symmetries of Three-Layer. 2605.18319 , archivePrefix=
-
[23]
Lampert , booktitle=
Mary Phuong and Christoph H. Lampert , booktitle=. Functional vs.\ parametric equivalence of. 2020 , url=
2020
-
[24]
Hidden Symmetries of
Grigsby, Elisenda and Lindsey, Kathryn and Rolnick, David , booktitle =. Hidden Symmetries of. 2023 , noeditor =
2023
-
[25]
Moritz Grillo and Guido Montúfar , year=. Most. 2605.03601 , archivePrefix=
-
[26]
2005 , publisher=
Gradient flows: in metric spaces and in the space of probability measures , author=. 2005 , publisher=
2005
-
[27]
Journal of Machine Learning Research , year =
Tolga Ergen and Mert Pilanci , title =. Journal of Machine Learning Research , year =
-
[28]
Proceedings of the 37th International Conference on Machine Learning , pages =
Neural Networks are Convex Regularizers: Exact Polynomial-time Convex Optimization Formulations for Two-layer Networks , author =. Proceedings of the 37th International Conference on Machine Learning , pages =. 2020 , editor =
2020
-
[29]
Proceedings of the 58th Annual ACM Symposium on Theory of Computing , pages =
Bakaev, Egor and Brunck, Florestan and Hertrich, Christoph and Stade, Jack and Yehudayoff, Amir , title =. Proceedings of the 58th Annual ACM Symposium on Theory of Computing , pages =. 2026 , noisbn =
2026
-
[30]
Journal of Machine Learning Research , year =
Tao Luo and Zhi-Qin John Xu and Zheng Ma and Yaoyu Zhang , title =. Journal of Machine Learning Research , year =
-
[31]
1975 , publisher=
Facing up to Arrangements: Face-Count Formulas for Partitions of Space by Hyperplanes , author=. 1975 , publisher=
1975
-
[32]
How Does the
Lai, Kuo-Wei and Wang, Guanghui and Tao, Molei and Muthukumar, Vidya , journal=. How Does the
-
[33]
The Fourteenth International Conference on Learning Representations , year=
Gradient Descent with Large Step Sizes: Chaos and Fractal Convergence Region , author=. The Fourteenth International Conference on Learning Representations , year=
-
[34]
Advances in neural information processing systems , volume=
Implicit regularization in deep learning may not be explainable by norms , author=. Advances in neural information processing systems , volume=
-
[35]
Advances in neural information processing systems , volume=
Convex neural networks , author=. Advances in neural information processing systems , volume=
-
[36]
Journal of Machine Learning Research , volume=
Breaking the curse of dimensionality with convex neural networks , author=. Journal of Machine Learning Research , volume=
-
[37]
SIAM journal on mathematical analysis , volume=
The variational formulation of the Fokker--Planck equation , author=. SIAM journal on mathematical analysis , volume=. 1998 , publisher=
1998
-
[38]
Annales de l'Institut Henri Poincare (B) Probabilites et statistiques , volume=
Uniform-in-time propagation of chaos for mean field Langevin dynamics , author=. Annales de l'Institut Henri Poincare (B) Probabilites et statistiques , volume=. 2025 , organization=
2025
-
[39]
arXiv preprint arXiv:2603.01842 , year=
Uniform-in-time concentration in two-layer neural networks via transportation inequalities , author=. arXiv preprint arXiv:2603.01842 , year=
-
[40]
Nearly-tight
Bartlett, Peter L and Harvey, Nick and Liaw, Christopher and Mehrabian, Abbas , journal=. Nearly-tight
-
[41]
1993 , publisher=
Topologies on closed and closed convex sets , author=. 1993 , publisher=
1993
-
[42]
Geometric combinatorics , volume=
An introduction to hyperplane arrangements , author=. Geometric combinatorics , volume=
-
[43]
The annals of mathematical statistics , pages=
A stochastic approximation method , author=. The annals of mathematical statistics , pages=. 1951 , publisher=
1951
-
[44]
Annales de l'Institut Henri Poincare (B) Probabilites et statistiques , volume=
Mean-field Langevin dynamics and energy landscape of neural networks , author=. Annales de l'Institut Henri Poincare (B) Probabilites et statistiques , volume=. 2021 , organization=
2021
-
[45]
arXiv preprint arXiv:2603.17785 , year=
A Dual Certificate Approach to Sparsity in Infinite-Width Shallow Neural Networks , author=. arXiv preprint arXiv:2603.17785 , year=
-
[46]
IEEE Transactions on Electronic Computers , volume=
Geometrical and Statistical Properties of Systems of Linear Inequalities with Applications in Pattern Recognition , author=. IEEE Transactions on Electronic Computers , volume=
-
[47]
2013 , publisher=
Convergence of probability measures , author=. 2013 , publisher=
2013
-
[48]
Handbook of differential equations: evolutionary equations , volume=
Gradient flows of probability measures , author=. Handbook of differential equations: evolutionary equations , volume=. 2007 , publisher=
2007
-
[49]
The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=
Deep linear networks for regression are implicitly regularized towards flat minima , author=. The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=
-
[50]
The Thirteenth International Conference on Learning Representations , year=
Universal Sharpness Dynamics in Neural Network Training: Fixed Point Analysis, Edge of Stability, and Route to Chaos , author=. The Thirteenth International Conference on Learning Representations , year=
-
[51]
Ussr computational mathematics and mathematical physics , volume=
Some methods of speeding up the convergence of iteration methods , author=. Ussr computational mathematics and mathematical physics , volume=. 1964 , publisher=
1964
-
[52]
Learning multiple layers of features from tiny images.(2009) , author=
2009
-
[53]
Advances in neural information processing systems , volume=
Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function , author=. Advances in neural information processing systems , volume=
-
[54]
2003 , publisher=
The symmetry perspective: from equilibrium to chaos in phase space and physical space , author=. 2003 , publisher=
2003
-
[55]
Physics Letters A , volume=
Final state sensitivity: an obstruction to predictability , author=. Physics Letters A , volume=. 1983 , publisher=
1983
-
[56]
Mathematische Zeitschrift , volume=
Geometry of real polynomial mappings , author=. Mathematische Zeitschrift , volume=. 2002 , publisher=
2002
-
[57]
2000 , publisher=
Introduction to topological manifolds , author=. 2000 , publisher=
2000
-
[58]
Physical Review E , volume=
Mixed basin boundary structures of chaotic systems , author=. Physical Review E , volume=. 1999 , publisher=
1999
-
[59]
Physical review letters , volume=
Sporadically fractal basin boundaries of chaotic systems , author=. Physical review letters , volume=. 1999 , publisher=
1999
-
[60]
Reviews of Modern Physics , volume=
Fractal structures in nonlinear dynamics , author=. Reviews of Modern Physics , volume=. 2009 , publisher=
2009
-
[61]
Physical Review Letters , volume=
Fractal basin boundaries, long-lived chaotic transients, and unstable-unstable pair bifurcation , author=. Physical Review Letters , volume=. 1983 , publisher=
1983
-
[62]
Physica D: Nonlinear Phenomena , volume=
Wada basin boundaries and basin cells , author=. Physica D: Nonlinear Phenomena , volume=. 1996 , publisher=
1996
-
[63]
Physica D: Nonlinear Phenomena , volume=
Basins of wada , author=. Physica D: Nonlinear Phenomena , volume=. 1991 , publisher=
1991
-
[64]
Wada basins and chaotic invariant sets in the H
Aguirre, Jacobo and Vallejo, Juan C and Sanju. Wada basins and chaotic invariant sets in the H. Physical Review E , volume=. 2001 , publisher=
2001
-
[65]
Proceedings of the IEEE international conference on computer vision , pages=
Unifying nuclear norm and bilinear factorization approaches for low-rank matrix decomposition , author=. Proceedings of the IEEE international conference on computer vision , pages=
-
[66]
Information and Inference: A Journal of the IMA , volume=
The non-convex geometry of low-rank matrix optimization , author=. Information and Inference: A Journal of the IMA , volume=. 2019 , publisher=
2019
-
[67]
Sur les transformations ponctuelles , url =
Hadamard , journal =. Sur les transformations ponctuelles , url =
-
[68]
The American Mathematical Monthly , volume=
Period three implies chaos , author=. The American Mathematical Monthly , volume=. 1975 , publisher=
1975
-
[69]
1995 , publisher=
Introduction to the modern theory of dynamical systems , author=. 1995 , publisher=
1995
-
[70]
Pure and Applied Geophysics , volume=
The essence of chaos , author=. Pure and Applied Geophysics , volume=. 1996 , publisher=
1996
-
[71]
A condition for the properness of polynomial maps , journal =
Ngu. A condition for the properness of polynomial maps , journal =. 2009 , pages =
2009
-
[72]
Geometry of real polynomial mappings , Ty =
Jelonek, Zbigniew , Da =. Geometry of real polynomial mappings , Ty =. Mathematische Zeitschrift , Number =. 2002 , Bdsk-Url-1 =. doi:10.1007/s002090100298 , Id =
-
[73]
Properness of Polynomial Maps with
Fukui, Toshizumi and Tsuchiya, Takeki , Da =. Properness of Polynomial Maps with. Arnold Mathematical Journal , Number =. 2023 , Bdsk-Url-1 =. doi:10.1007/s40598-022-00205-2 , Id =
-
[74]
Testing sets for properness of polynomial mappings , Ty =
Jelonek, Zbigniew , Da =. Testing sets for properness of polynomial mappings , Ty =. Mathematische Annalen , Number =. 1999 , Bdsk-Url-1 =. doi:10.1007/s002080050316 , Id =
-
[75]
2008 , publisher=
Equilibrium States and the Ergodic Theory of Anosov Diffeomorphisms , author=. 2008 , publisher=
2008
-
[76]
, number=
Anosov, D.V. , number=. Geodesic Flows on Closed. 1969 , publisher=
1969
-
[77]
Smale , title =
S. Smale , title =. Bulletin of the American Mathematical Society , number =
-
[78]
2014 , publisher=
Topological dynamical systems: an introduction to the dynamics of continuous mappings , author=. 2014 , publisher=
2014
-
[79]
2007 , publisher=
Discrete chaos: with applications in science and engineering , author=. 2007 , publisher=
2007
-
[80]
2012 , publisher=
One-dimensional dynamics , author=. 2012 , publisher=
2012
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.