REVIEW 2 major objections 1 cited by
Local minima of high-dimensional two-layer ReLU networks admit an exact low-dimensional description in summary statistics and coincide with the attractive fixed points of one-pass SGD.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 23:15 UTC pith:5BLQ2JLI
load-bearing objection Abstract promises an exact low-dimensional taxonomy of all local minima (and their SGD accessibility) for a minimal two-layer ReLU teacher-student model, but the supplied manuscript is pure mojibake so nothing can be checked. the 2 major comments →
Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networks
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
In the realisable Gaussian teacher-student setting, the local minima of the population loss of the two-layer ReLU network admit an exact representation by a finite collection of summary statistics of the student and teacher weights. These critical points are identical to the attractive fixed points of the one-pass SGD dynamics written in the same summary-statistic coordinates. Consequently the landscape is organised into discrete hierarchical families of minima whose stability and reachability under gradient-based training are controlled by the degree of overparameterisation.
What carries the argument
The low-dimensional summary statistics of the student and teacher weight vectors. They furnish an exact coordinate chart for every local minimum of the population loss and simultaneously serve as the state space in which one-pass SGD becomes an autonomous dynamical system whose attractive fixed points are those minima.
Load-bearing premise
The analysis is performed on the infinite-data population loss with isotropic Gaussian covariates and a teacher of exactly the same two-layer ReLU form; if any of those idealisations fails, both the location and the stability of the reported families of minima can change.
What would settle it
For moderate dimension (e.g., d around 100 and a handful of neurons) compute population-loss critical points by direct high-precision optimisation and test whether every critical point lies exactly on the predicted summary-statistic manifold; any genuine critical point off that manifold would falsify the exact low-dimensional representation.
If this is right
- Overparameterisation systematically alters the stability of discrete families of minima, rendering many spurious solutions unstable or unreachable under gradient dynamics.
- Global minima become the dominant attractors of one-pass SGD once the student is sufficiently wide, reducing the probability of convergence to suboptimal fixed points.
- Any landscape analysis that ignores the summary-statistic reduction or the fixed-point link to SGD will miss entire hierarchical families of critical points even in this minimal model.
- The hierarchical organisation of minima is an intrinsic geometric feature of the ReLU teacher-student landscape rather than a finite-sample artefact.
Where Pith is reading between the lines
- Analogous summary-statistic reductions may exist for deeper or mildly non-realisable architectures and could explain why overparameterisation helps in practice beyond the two-layer case.
- Finite-sample noise is likely to split or blur the discrete families, so empirical loss surfaces on moderate data sets may appear smoother or more rugged than the exact population picture predicts.
- The fixed-point correspondence supplies a concrete design principle for schedules or regularisers that preferentially steer trajectories toward the global rather than the spurious families.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the population loss of two-layer ReLU networks of the form ∑_k ReLU(w_k^ op x) in a realisable Gaussian teacher-student setting. It claims that every local minimum admits an exact finite-dimensional representation in a set of summary statistics independent of ambient dimension, that these critical points coincide with attractive fixed points of the one-pass SGD flow written in the same coordinates, and that the resulting landscape organises into discrete hierarchical families whose stability and reachability change with over-parameterisation, making global minima increasingly accessible. The abstract further asserts that common simplifying assumptions miss essential features of even this minimal model.
Significance. If the exact low-dimensional characterisation and the SGD fixed-point correspondence hold rigorously, the work would supply a rare, fully interpretable description of the non-convex landscape of a minimal neural network and a concrete dynamical explanation of how over-parameterisation improves accessibility of global minima. Such a result would be of clear interest to the theoretical machine-learning community and would serve as a useful benchmark against which more approximate mean-field or landscape analyses could be tested. Because the manuscript text supplied for review is entirely corrupted, however, none of these claims can be verified from definitions, theorems or proofs.
major comments (2)
- The entire body of the manuscript (all sections after the abstract) consists solely of mojibake / encoding garbage. No definitions of the summary statistics, no statement of the main theorem, no reduced loss, no Jacobian analysis of the SGD flow, and no treatment of ReLU non-differentiability are recoverable. Consequently the central claims of exact low-dimensional representation and of correspondence with attractive fixed points cannot be checked at all.
- Without a readable statement of the precise algebraic form of the reduced loss or of the conditions under which the correspondence is claimed to be exhaustive, it is impossible to assess whether the hierarchical organisation of minima or the over-parameterisation accessibility statements are rigorously established or merely conjectural.
Circularity Check
No circularity can be exhibited: readable abstract is a first-principles landscape/SGD analysis; body is mojibake so no derivation steps are quotable.
full rationale
The only recoverable text is the abstract, which states a theoretical programme: population loss of a two-layer ReLU student in a realisable Gaussian teacher-student setting admits an exact low-dimensional description of local minima via summary statistics, and those minima coincide with attractive fixed points of one-pass SGD written in the same coordinates. That claim is framed as a derivation under explicit modelling assumptions (isotropic Gaussians, realisable teacher, infinite-data loss), not as a fit to data or a renaming of an empirical pattern. No fitted constants, no self-definitional identities, and no load-bearing uniqueness theorems imported from the authors appear in the abstract. The supplied full-manuscript body is pure encoding garbage (mojibake) and yields no equations, definitions of the summary statistics, Jacobian analysis, or theorem statements that could be checked for circular reduction. Per the hard rules, circularity may be asserted only when a specific reduction can be quoted; none can. Residual concern that the summary-statistic map might be chosen so the fixed-point correspondence is tautological remains speculative without readable text and does not raise the score. Honest finding: no significant circularity identified.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Covariates x are i.i.d. isotropic Gaussian; the population loss is the expectation under this measure.
- domain assumption The problem is realisable: the teacher is itself a finite sum of ReLUs of the same form, so zero population loss is attainable.
- ad hoc to paper Student network is exactly ∑_{k=1}^K ReLU(w_k^⊤ x) (output weights fixed to +1).
- domain assumption One-pass SGD dynamics admit a closed description in the same summary statistics that label the critical points.
invented entities (2)
-
Summary-statistic coordinates that exactly label all local minima
no independent evidence
-
Discrete hierarchical families of local minima
no independent evidence
read the original abstract
We study the population loss landscape of two-layer ReLU networks of the form $\sum_{k=1}^K \mathrm{ReLU}(w_k^\top x)$ in a realisable teacher-student setting with Gaussian covariates. We show that local minima admit an exact low-dimensional representation in terms of summary statistics, yielding a sharp and interpretable characterisation of the landscape. We further establish a direct link with one-pass SGD: local minima correspond to attractive fixed points of the dynamics in summary statistics space. This perspective reveals a hierarchical organisation of minima into discrete families and shows how overparameterisation changes their stability and reachability under gradient-based dynamics. In this overparameterised regime, global minima become increasingly accessible, attracting the dynamics and reducing convergence to spurious solutions. Overall, our results reveal intrinsic limitations of common simplifying assumptions, which may miss essential features of the loss landscape even in minimal neural network models.
Figures
Forward citations
Cited by 1 Pith paper
-
Spectral phase transitions and trainability in neural network learning dynamics
SGD on neural network weights induces a BBP phase transition that detaches signal eigenvalues from the random bulk, yielding an analytically solvable phase diagram for trainability in a linear teacher-student model.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.