Pith. sign in

REVIEW 2 major objections 3 minor 73 references

Cascading Through the Hierarchy: Regularizer-Induced Feature Detection as Phase Transitions in Deep Linear Neural Networks

T0 review · 2 major / 3 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Tuning L2 regularization strength in deep linear networks produces a predictable cascade of phase transitions, one per learnable singular direction, with explicit critical strengths and exponents.

desk verdict The aligned-case results are real and reusable—closed-form transition strengths, Hessian spectra, and critical dynamics that match numerics—but the abstract oversells the generality and the order-parameter language needs a careful pass. read the letter →

arxiv 2608.06597 v1 pith:PORDWYME submitted 2026-08-06 cond-mat.stat-mech cond-mat.dis-nncs.LG

classification cond-mat.stat-mechcond-mat.dis-nncs.LG
keywords deeplinearnetworksL2regularizationphasetransitionslosslandscapeHessianspectrumorderparameterssingularvalueshrinkagegradientflow
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Under an alignment condition on the data covariances, this paper establishes that tuning the L2 regularization strength $\beta$ in deep linear networks induces a cascade of phase transitions, one for each learnable singular direction of the regression matrix. For one hidden layer, the transitions are continuous: the $j$-th direction turns on at $\beta^*_j = \eta_j$ and its singular value grows like $\sqrt{\eta_j - \beta}$. For networks with three or more layers, the same cascade is first-order, with explicit thresholds for the appearance of a nontrivial minimum and for it to become globally optimal. This matters because it turns regularization strength into a tunable probe of the hierarchical feature structure of the loss landscape, with the singular values of the learned map as measurable order parameters.

What carries the argument

The central object is the macrostate $W = W^{(n)} \cdots W^{(1)}$ and its singular-value spectrum $\sigma$. In the 0-balanced subspace, where adjacent weight matrices share the same nonzero singular values, regularization closes the gradient-flow dynamics on the three macro-quantities $W_L$, $W$, $W_R$. Under the alignment condition, the singular values evolve independently under the single-variable potentials $L_j$, so every critical point of the full loss landscape is labelled by which singular values are finite. The Hessian at those points splits into three countable sectors: massive data-dependent directions, $\beta$-dependent directions from explicitly broken symmetries, and exact zero modes from the degeneracy of microstates belonging to the same macrostate. For one hidden layer, an exact matrix differential equation solution supplies the algebraic $t^{-1/2}$ relaxation at the transition.

What would settle it

Train a one-hidden-layer linear network on Gaussian data with $\Sigma_{xx}=I$ and $\Sigma_{yx}=\mathrm{diag}(0.95, 0.55)$, using gradient flow with weight decay at $\beta=0.60$; the predicted minimal macrostate has only one nonzero singular value, $\sigma_1=\sqrt{0.35}\approx 0.59$, and $\sigma_2=0$ exactly. Measuring a nonzero second singular value at that $\beta$, or observing $\sigma_2$ to jump discontinuously when $\beta$ crosses $0.55$, would contradict the second-order cascade.

Watch

Extended reading notes

Core claim

At the core is an effective loss on the singular values $\sigma_j$ of the weight matrices, valid when the input covariance $\Sigma_{xx}$ and the cross-covariance $\Sigma_{yx}$ share the same right singular vectors. The loss separates into per-direction potentials $L_j(\sigma_j) = \kappa_j \sigma_j^{2n} - 2\eta_j \sigma_j^n + n\beta \sigma_j^2$. For $n=2$ this is a quartic potential with a unique minimum that moves continuously away from zero as $\beta$ crosses $\eta_j$, giving second-order transitions and a $t^{-1/2}$ critical slowing down. For $n\ge 3$ it is a higher-order polynomial whose nontrivial minimum appears by bifurcation and becomes globally preferred by level crossing, giving first-order transitions with explicit thresholds. The paper also derives the complete Hessian spectrum at every rank-$r$ minimum, decomposing it into massive data-dependent modes, $\beta$-dependent modes, and zero modes from macrostate degeneracy, and proves that only the ordered set of largest singular values yields local minima.

Load-bearing premise

The load-bearing premise is that the input covariance $\Sigma_{xx}$ and the cross-covariance $\Sigma_{yx}$ share the same right singular vectors, so each singular direction of the regression matrix evolves independently; if that simultaneous diagonalization fails, the singular vectors of the learned map become $\beta$-dependent and the per-direction transition picture no longer holds.

Editorial extensions

If this is right

  • For one-hidden-layer networks, the trained model's rank is a step function of $\beta$: direction $j$ is learned only below $\beta = \eta_j$, and its singular value grows continuously from zero as $\sqrt{\eta_j - \beta}$.
  • For networks with three or more layers, each direction appears through a first-order transition: a nonzero minimum appears at $\beta^{(\mathrm{det})}_j$ and only becomes the global minimum at $\beta^*_j$, producing hysteresis when $\beta$ is annealed up and down.
  • At any global minimum of rank $r$, the Hessian spectrum has three sharply separated sectors: massive data-dependent eigenvalues, $\beta$-dependent eigenvalues, and exact zero modes, whose multiplicities can be counted from $r$ and the network dimensions.
  • Only when the finite singular values are the $r$ largest ones is the critical point a local minimum; any other subset gives a saddle with at least one negative direction.
  • Near a second-order transition the slowest singular-value mode decays as $t^{-1/2}$, so convergence time diverges at the onset, giving an operational signature of the transition in finite-time training runs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension: for data whose covariances nearly satisfy alignment, the paper's perturbative result predicts that the singular vectors of the learned map rotate continuously with $\beta$; tracking this rotation should reveal transition points without assuming exact alignment.
  • The effective single-variable potential form suggests an operational definition of a learned feature in generic networks: a singular direction is learned when its potential changes from one minimum to two, so one could try to recover the potentials $L_j$ empirically from measured singular-value trajectories under different $\beta$.
  • The paper's analogy with information-bottleneck compression is more than formal: if $\beta$ acts as a bottleneck parameter, the rank-reduced solutions should lie on a relevance-compression frontier for Gaussian data, which is a direct numerical check.
  • For non-linear networks, the paper's numerics indicate the same transitions appear as jumps in effective rank; an extension would be to test whether the $t^{-1/2}$ critical slowing down also appears near those jumps in tanh networks.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper analyzes deep linear networks with L2 regularization, treating the regularization strength β as an external field. Under an alignment condition on the input and cross covariances (Eq. 17), it derives an effective loss L(σ)=Σ_j(κ_j σ_j^{2n} − 2η_j σ_j^n + nβ σ_j^2), which acts as a Landau free energy for the singular values of the weight matrices. For one hidden layer this yields a cascade of second-order transitions at β*_j=η_j with order-parameter scaling σ_j∼(η_j−β)^{1/2}; for n≥3 layers it yields first-order transitions with bifurcation and level-crossing strengths (25)–(26). The paper also derives complete Hessian spectra at the critical points, analyzes degeneracies and symmetries, gives an exact Riccati solution for the 1-hidden-layer dynamics with t^{−1} critical slowing down, and presents numerical support plus tentative extensions to non-linear networks.

Significance. If the alignment condition is accepted, the paper provides one of the most complete exactly solvable models in this line of work: parameter-free predictions for transition points, order-parameter scaling, Hessian spectra, and critical dynamics, all derived from the loss rather than fitted. The effective Landau description and the connection between microscopic Hessian geometry and macroscopic order parameters are genuine strengths. The main weakness is that the exact results rest on a nongeneric simultaneous-diagonalizability premise, and the paper's abstract and Key Findings present these as predictions for the general model. The generic case is only treated perturbatively and numerically, so the quantitative content of the central claim is conditional.

major comments (2)
  1. [Abstract; §I (Key Findings); §III.A–B; Eq. (17); §IV.B] The quantitative claims — β*_j=η_j and σ_j∼(η_j−β)^{1/2} for one hidden layer, the first-order strengths (25)–(26), and the Hessian spectra in Tables II, III, and VII — are all derived under the alignment condition (17). For generic covariance matrices, and in particular for any finite number of training samples, this condition is violated; Section IV.B concedes that the singular vectors of W* become β-dependent, that the decoupled Landau picture (15) no longer holds, and that only the onset β_onset=η_max is established in the generic case. The abstract and Key Findings nevertheless present the cascade as an analytic prediction of the model without this qualification, and the claim of providing a rigorous underpinning of the numerical cascades in [21,22] is therefore not established for generic data. The authors should either explicitly restrict the central claims to the aligned case or supply a rigorous perturbation or continuity argument showing that the transition strengths are unchanged in the generic case.
  2. [§III.A.1; Tables II, III, VII; App. C] The Hessian spectra are presented in the main text as general results for aligned covariances, but Appendix C derives them only for Σxx=1. Specifically, App. C1 states 'For Σxx=1 for simplicity' before Table VI, and the captions of Tables III and VII carry the same restriction, while Table II does not. For general aligned Σxx with κ_j≠1, the eigenvalues will depend on the κ_j (already visible in the one-hidden-layer optimum σ*_j^2=(η_j−β)/κ_j), and no closed-form spectrum is provided for this case. The paper should either state the Σxx=1 restriction wherever these tables appear or supply the general-aligned-case spectrum.
minor comments (3)
  1. [§III.B heading and Eq. (25)] The section title 'n−1-hidden Layer Networks' and the introductory sentence 'n>3 layers' conflict with Eq. (25), which states 'n≥3 layers', and with Fig. 7, which treats the two-hidden-layer case n=3. The class of architectures should be stated consistently as n≥3.
  2. [Table II and §III.A.2] The main-text sentence after Fig. 5 says there are 'another nine eigenvalues λ with 0<λ≤4β' for the full-rank example; according to Table II with r=3, the β-sector contains six 4β eigenvalues and the zero sector contains three zero eigenvalues, so the sentence appears to misstate the count.
  3. [Table II caption] Table II should explicitly state that the spectrum is computed for Σxx=1, as Tables III and VII do; without this note the table reads as valid for arbitrary aligned covariances.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the analytic cascade is derived from the regularized loss under the stated alignment condition, and the self-citations to the authors' prior phenomenology are not load-bearing.

full rationale

The paper's central predictions—β*_j=η_j with σ_j∼√(η_j−β) for one hidden layer, the first-order bifurcation and level-crossing strengths (25)–(26) for deeper networks, and the Hessian spectra of Tables II, III, VII—are obtained by minimizing the effective potential L(σ)=Σ_j L_j(σ_j), which is itself derived from the original MSE-plus-L2 loss (2) and the 0-balanced macro dynamics. No parameter is fitted to the target quantities: the transition points are algebraic functions of the input covariance singular values η_j and κ_j, and the order-parameter scaling follows from solving ∂L_j/∂σ_j=0. The alignment condition (17) is an explicitly stated premise that decouples the singular-value dynamics; its failure for generic or finite-sample covariances is acknowledged in Section IV.B, where the exact analysis is withdrawn and only perturbative and numerical support is offered. That is a scope restriction rather than circularity. The self-citations [21,22] document prior phenomenological observations that motivated the study; they are not used to justify any analytic result, and the derivation stands independently of them. Hence no circular step can be identified.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The model introduces no free parameters fitted to data and no new physical entities. The central results depend on the alignment condition and the infinite-data, zero-noise gradient flow limit, which are stated, not hidden. The regularization strength β is a tunable control parameter, not a fitted constant, and the transition strengths are explicit functions of the data covariance singular values.

assumptions (6)
  • domain assumption Alignment condition: Σxx = R S_xx R^T and Σyx = L S_yx R^T share the same right singular vectors R (Eq. 17)
    This condition makes the SVD of Σyx Σxx^{-1} inherited from the separate covariances and decouples the singular-value dynamics into independent scalar potentials. All exact transition strengths and Hessian spectra depend on it. The authors acknowledge that it is violated by finite-sample empirical covariances (Section IV.B).
  • domain assumption Ndata → ∞ so that empirical covariances equal population covariances (Section I.A)
    The exact analysis replaces Σ(e) with Σ and ignores finite-sample fluctuations. Section IV.B discusses finite-N smearing only numerically.
  • domain assumption Zero temperature, noiseless gradient flow with infinitesimal learning rate, starting from zero and following the decay directions (Section I.C, Appendix B)
    The dynamics is deterministic gradient flow; the macro-dynamics equations (13)-(14) and the 0-balanced condition (10)-(11) hold in this limit. Noise is only invoked in the discussion of first-order transition escape and hysteresis.
  • domain assumption Σxx invertible and Σyx full rank (Section I.A)
    This ensures the optimal macrostate W* = Σyx Σxx^{-1} exists and that the number of transitions corresponds to the rank of Σyx.
  • standard math Existence and uniqueness of solutions of the matrix Riccati equation, and the matrix identity (E9) used for the exact dynamics (Appendix E)
    The solution formula follows from standard results on matrix Riccati equations, cited to [61]. No ad hoc assumptions are introduced beyond the known theory.
  • standard math The 0-balanced condition (11) characterizes the long-time critical point subspace of gradient flow with L2 regularization
    Equation (10) shows exponential convergence to the 0-balanced manifold; this is a derived consequence of the gradient flow equations, not an assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Cascading Through the Hierarchy: Regularizer-Induced Feature Detection as Phase Transitions in Deep Linear Neural Networks." pith.science (2026). https://pith.science/paper/PORDWYME

@misc{pith2026260806597,
  author       = {Pith},
  title        = {Pith review of: Cascading Through the Hierarchy: Regularizer-Induced Feature Detection as Phase Transitions in Deep Linear Neural Networks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PORDWYME}},
  note         = {Machine review of arXiv:2608.06597}
}
read the original abstract

A scientific theory of deep learning, comprising learning dynamics and statistical properties of learned models, is rapidly gaining attention. One of the corner stones of this development are analytically solvable toy models, allowing for the fully tractable analysis of the learning dynamics. Here we analytically investigate such a toy model using the regularization strength as a tunable external parameter - akin to external fields in statistical physics. In previous studies, (i) an onset of learning transition was predicted analytically and (ii) it was phenomenologically/numerically established that tuning the regularization strength can result in a cascade of phase transitions. The number of those transitions was linked to the geometry of the loss landscape determined by the model complexity. Setting up a rigorous framework underpinning the previous numerical observations, our investigation reveals a precise connection between those cascades of phase transitions, learnable features and the underlying geometry. We provide analytic predictions of these phase transitions as well as tractable order parameters related to learned features. At the level of the minimal model, we connect this macroscopic perspective (that can be condensed into an effective description) to the microscopic perspective in terms of the geometry of the loss landscape characterized by the Hessian spectrum. Thus, the presented model provides a platform to explore and sharpen advances made in the scientific theory of deep learning rooted in statistical physics concepts.

Figures

Figures reproduced from arXiv: 2608.06597 by the authors.

Figure 1
Figure 1. FIG. 1 [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2: Path through the loss landscape, based on Hessian eigendirections revealing a basin structure in the [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. (d) (for more hidden layers). A. Cascade of Second Order Transitions: 1 hidden Layer Networks For 1 hidden layer, decreasing the regularization strength results in a cascade of (continuous) second or￾der transitions. Under the alignment assumption of the data distribution (17)), the order parameter are effec￾tively described by L(σ) = Prmax j=1 Lj (σj ) with the indi￾vidual quartic potentials: Lj (σj ) = κjσ 4 j − 2… view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: FIG. 4: Effective loss landscapes (1 hidden layer) for [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: FIG. 5: Hessian spectra for 1 hidden (a,b) and 2 hidden layers (c,d) with [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: FIG. 6 [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 7
Figure 7. Figure 7: FIG. 7 [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: FIG. 8 [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: FIG. 9 [PITH_FULL_IMAGE:figures/full_fig_p017_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

73 extracted references · 57 canonical work pages

  1. [1]

    M. K. Transtrum, B. B. Machta, K. S. Brown, B. C. Daniels, C. R. Myers, and J. P. Sethna, Perspective: Sloppiness and emergent theories in physics, biology, and beyond, The Journal of Chemical Physics143, 010901 (2015)

  2. [2]

    Mehta, M

    P. Mehta, M. Bukov, C.-H. Wang, A. G. Day, C. Richard- son, C. K. Fisher, and D. J. Schwab, A high-bias, low- variance introduction to Machine Learning for physicists, Physics Reports810, 1 (2019), a high-bias, low-variance introduction to Machine Learning for physicists

  3. [3]

    Bahri, J

    Y. Bahri, J. Kadmon, J. Pennington, S. S. Schoenholz, J. Sohl-Dickstein, and S. Ganguli, Statistical Mechanics of Deep Learning, Annual Review of Condensed Matter Physics11, 501 (2020)

  4. [4]

    D. A. Roberts, S. Yaida, and B. Hanin,The Principles of DeepLearningTheory: AnEffectiveTheoryApproachto Understanding Neural Networks(Cambridge University Press, 2022)

  5. [5]

    Seroussi, G

    I. Seroussi, G. Naveh, and Z. Ringel, Separation of scales and a thermodynamic description of feature learning in some CNNs, Nature Communications14, 908 (2023)

  6. [6]

    A. Saxe, J. McClelland, and S. Ganguli, Exact solutions to the nonlinear dynamics of learning in deep linear neu- ral networks, inProceedings of the International Confer- ence on Learning Represenatations 2014(2014)

  7. [7]

    Jacot, F

    A. Jacot, F. Ged, B. Şimşek, C. Hongler, and F. Gabriel, Saddle-to-saddle dynamics in deep linear networks: Small initialization training, symmetry, and sparsity (2022), arXiv:2106.15933 [stat.ML]

  8. [8]

    Menon, The geometry of the deep linear network, inXIV Symposium on Probability and Stochastic Pro- cesses, edited by C

    G. Menon, The geometry of the deep linear network, inXIV Symposium on Probability and Stochastic Pro- cesses, edited by C. G. Higuera Chan, J. A. López Mim- bela, S. I. López, and C. G. Pacheco (Springer Nature Switzerland, Cham, 2025) pp. 1–47

Show all 73 references
  1. [9]

    Zhang, A

    Y. Zhang, A. M. Saxe, and P. E. Latham, Saddle-to- saddle dynamics explains a simplicity bias across neural network architectures, inThe Fourteenth International Conference on Learning Representations(2026)

  2. [10]

    Kunin, G

    D. Kunin, G. L. Marchetti, F. Chen, D. Karkada, J. B. Simon, M. R. DeWeese, S. Ganguli, and N. Miolane, Al- ternating gradient flows: A theory of feature learning in two-layer neural networks, inThe Thirty-ninth Annual Conference on Neural Information Processing Systems (2026)

  3. [11]

    K.FukumizuandS.Amari,Localminimaandplateausin hierarchical structures of multilayer perceptrons, Neural Networks13, 317 (2000)

  4. [12]

    Zhang, Z

    Y. Zhang, Z. Zhang, T. Luo, and Z. J. Xu, Embed- ding principle of loss landscape of deep neural net- works, inAdvances in Neural Information Processing Systems, Vol. 34, edited by M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (Curran As- sociates, Inc., ...

  5. [13]

    B.Zhao, R.Walters,andR.Yu,SymmetryinNeuralNet- work Parameter Spaces, Transactions on Machine Learn- ing Research (2026)

  6. [14]

    Ziyin, Y

    L. Ziyin, Y. Xu, T. Poggio, and I. Chuang, Parameter symmetrypotentiallyunifies deep learningtheory(2025), arXiv:2502.05300 [cs.LG]

  7. [15]

    Goldenfeld,Lectures On Phase Transitions And The Renormalization Group(CRC Press, 1992)

    N. Goldenfeld,Lectures On Phase Transitions And The Renormalization Group(CRC Press, 1992)

  8. [16]

    S. Watanabe, Review and prospect of algebraic research in equivalent framework between statistical mechanics and machine learning theory, Reviews in Mathematical Physics0, 2461009 (2025)

  9. [17]

    C. C. J. Dominé, N. Anguita, A. M. Proca, L. Braun, D. Kunin, P. A. M. Mediano, and A. M. Saxe, From lazy to rich: Exact learning dynamics in deep linear networks, inThe Thirteenth International Conference on Learning Representations(2025)

  10. [18]

    Ziyin, B

    L. Ziyin, B. Li, and X. Meng, Exact solutions of a deep linear network, inAdvances in Neural Information Pro- cessing Systems, Vol. 35, edited by S. Koyejo, S. Mo- hamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Curran Associates, Inc., 2022) pp. 24446–24458

  11. [19]

    Ziyin and M

    L. Ziyin and M. Ueda, Zeroth, first, and second-order phase transitions in deep neural networks, Phys. Rev. Res.5, 043243 (2023)

  12. [20]

    Ziyin, Symmetry induces structure and constraint of learning, inForty-first International Conference on Ma- chine Learning(2024)

    L. Ziyin, Symmetry induces structure and constraint of learning, inForty-first International Conference on Ma- chine Learning(2024)

  13. [21]

    I. T. Ersoy and K. Wiesner, Exploring L2-Phase Transi- tions on Error Landscapes, inHigh-dimensional Learning Dynamics 2025(2025)

  14. [22]

    I. T. Ersoy, A. F. C. Licha, and K. Wiesner, Phase tran- sitions reveal hierarchical structure in deep neural net- works (2025), arXiv:2512.11866 [cs.LG]

  15. [23]

    M. K. Winter and L. M. C. Janssen, Glassy dynamics in deep neural networks: A structural comparison, Phys. Rev. Res.7, 023010 (2025)

  16. [24]

    Ringel, N

    Z. Ringel, N. Rubin, E. Mor, M. Helias, and I. Seroussi, Applications of Statistical Field Theory in Deep Learning (2025), arXiv:2502.18553 [stat.ML]

  17. [25]

    Simon, D

    J. Simon, D. Kunin, A. Atanasov, E. Boix-Adserà, B. Bordelon, J. Cohen, N. Ghosh, F. Guth, A. Jacot, M. Kamb, D. Karkada, E. J. Michaud, B. Ottlik, and J. Turnbull, There Will Be a Scientific Theory of Deep Learning (2026), arXiv:2604.21691 [stat.ML]

  18. [26]

    Baldi and K

    P. Baldi and K. Hornik, Neural networks and principal component analysis: Learning from examples without lo- cal minima, Neural Networks2, 53 (1989)

  19. [27]

    Fukumizu, Dynamics of Batch Learning in Multilayer Neural Networks, inICANN 98, edited by L

    K. Fukumizu, Dynamics of Batch Learning in Multilayer Neural Networks, inICANN 98, edited by L. Niklasson, M. Bodén, and T. Ziemke (Springer London, London,

  20. [28]

    Kawaguchi, Deep Learning without Poor Local Min- ima, inAdvances in Neural Information Processing Systems, Vol

    K. Kawaguchi, Deep Learning without Poor Local Min- ima, inAdvances in Neural Information Processing Systems, Vol. 29, edited by D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Curran Asso- ciates, Inc., 2016)

  21. [29]

    E. M. Achour, F. Malgouyres, and S. Gerchinovitz, The Loss Landscape of Deep Linear Neural Networks: a Second-order Analysis, Journal of Machine Learning Re- search25, 1 (2024)

  22. [30]

    Wendin and C

    J. Wendin and C. Altafini, Gradient Flow Equations for Deep Linear Neural Networks: A Survey from a Network Perspective, SIAM Review68, 293 (2026)

  23. [31]

    S. P. Singh, G. Bachmann, and T. Hofmann, Analytic Insights into Structure and Rank of Neural Network Hes- sian Maps, inAdvances in Neural Information Process- ing Systems, edited by A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (2021). 19

  24. [32]

    S. P. Singh and T. Hofmann, Closed form of the Hessian spectrum for some Neural Networks, inHigh-dimensional Learning Dynamics 2024: The Emergence of Structure and Reasoning(2024)

  25. [33]

    Lindsey and G

    K. Lindsey and G. Menon, Regularization implies balancedness in the deep linear network (2025), arXiv:2511.01137 [cs.LG]

  26. [34]

    A. Chen, T. Kotwal, and G. Menon, Entropic Regularization in the Deep Linear Network (2025), arXiv:2512.06137 [cs.NE]

  27. [35]

    Menon and T

    G. Menon and T. Yu, An entropy formula for the Deep Linear Network (2026), arXiv:2509.09088 [cs.LG]

  28. [36]

    J.-F. Cai, E. J. Candès, and Z. Shen, A Singular Value Thresholding Algorithm for Matrix Completion, SIAM Journal on Optimization20, 1956 (2010)

  29. [37]

    J. D. M. Rennie and N. Srebro, Fast maximum margin matrix factorization for collaborative prediction, inPro- ceedings of the 22nd International Conference on Ma- chine Learning, ICML ’05 (Association for Computing Machinery, New York, NY, USA, 2005) p. 713–719

  30. [38]

    Scarvelis and J

    C. Scarvelis and J. Solomon, Nuclear Norm Regulariza- tion for Deep Learning, inThe Thirty-eighth Annual Conference on Neural Information Processing Systems (2024)

  31. [39]

    Duan and G

    H. Duan and G. Montúfar, Understanding Learn- ing Invariance in Deep Linear Networks (2025), arXiv:2506.13714 [stat.ML]

  32. [40]

    Wang and A

    Z. Wang and A. Jacot, Implicit bias of SGD inL 2- regularized linear DNNs: One-way jumps from high to low rank, inThe Twelfth International Conference on Learning Representations(2024)

  33. [41]

    Chechik, A

    G. Chechik, A. Globerson, N. Tishby, and Y. Weiss, In- formationBottleneckforGaussianVariables,inAdvances in Neural Information Processing Systems, Vol. 16, edited by S. Thrun, L. Saul, and B. Schölkopf (MIT Press, 2003)

  34. [42]

    Arora, N

    S. Arora, N. Cohen, and E. Hazan, On the Optimiza- tion of Deep Networks: Implicit Acceleration by Overpa- rameterization, inProceedings of the 35th International Conference on Machine Learning, Proceedings of Ma- chine Learning Research, Vol. 80, edited by J. Dy and A. Krause...

  35. [43]

    B. Zhao, I. Ganev, R. Walters, R. Yu, and N. Dehmamy, Symmetries, Flat Minima, and the Conserved Quantities of Gradient Flow, inThe Eleventh International Confer- ence on Learning Representations(2023)

  36. [44]

    B.LiandD.Saad,ExploringtheFunctionSpaceofDeep- Learning Machines, Phys. Rev. Lett.120, 248301 (2018)

  37. [45]

    Trager, K

    M. Trager, K. Kohn, and J. Bruna, Pure and Spurious Critical Points: a Geometric Study of Linear Networks, inInternational Conference on Learning Representations (2020)

  38. [46]

    B. Bah, H. Rauhut, U. Terstiege, and M. Westdicken- berg, Learning deep linear neural networks: Riemannian gradient flows and convergence to global minimizers, In- formation and Inference: A Journal of the IMA11, 307 (2022)

  39. [47]

    Liang, L

    K. Liang, L. Chen, B. Liu, and Q. Liu, Cautious Optimiz- ers: Improving Training with One Line of Code, inThe Fourteenth International Conference on Learning Repre- sentations(2026)

  40. [48]

    L. Chen, J. Li, K. Liang, B. Su, C. Xie, N. W. Pierse, C. Liang, N. Lao, and Q. Liu, Cautious Weight Decay (2026), arXiv:2510.12402 [cs.LG]

  41. [49]

    Altland and B

    A. Altland and B. Simons,Condensed Matter Field The- ory, 3rd ed. (Cambridge University Press, 2023)

  42. [50]

    Braun, C

    L. Braun, C. Dominé, J. Fitzgerald, and A. Saxe, Ex- act learning dynamics of deep linear networks with prior knowledge, inAdvances in Neural Information Process- ing Systems, Vol. 35, edited by S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Curran Associat...

  43. [51]

    I. T. Ersoy and K. Wiesner, Noise-driven escape from metastable phases explains grokking in deep neural net- works, inHigh-dimensional Learning Dynamics 2026 (2026)

  44. [52]

    T. Mori, L. Ziyin, K. Liu, and M. Ueda, Power-Law Escape Rate of SGD, inProceedings of the 39th In- ternational Conference on Machine Learning, Proceed- ings of Machine Learning Research, Vol. 162, edited by K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Saba...

  45. [53]

    M. S. Advani, A. M. Saxe, and H. Sompolinsky, High- dimensional dynamics of generalization error in neural networks, Neural Networks132, 428 (2020)

  46. [54]

    Roy and M

    O. Roy and M. Vetterli, The effective rank: A measure of effective dimensionality, in2007 15th European Signal Processing Conference(2007) pp. 606–610

  47. [55]

    Yunis, K

    D. Yunis, K. K. Patel, S. Wheeler, P. H. P. Savarese, G. Vardi, J. Frankle, K. Livescu, M. Maire, and M. Wal- ter, Rank Minimization, Alignment and Weight Decay in Neural Networks, inHigh-dimensional Learning Dynam- ics 2024: The Emergence of Structure and Reasoning (2024)

  48. [56]

    L.Arola-FernándezandL.Lacasa,Effectivetheoryofcol- lective deep learning, Phys. Rev. Res.6, L042040 (2024)

  49. [57]

    Li and H

    Q. Li and H. Sompolinsky, Statistical Mechanics of Deep Linear Neural Networks: The Backpropagating Kernel Renormalization, Phys. Rev. X11, 031059 (2021)

  50. [58]

    Bachtis, G

    D. Bachtis, G. Biroli, A. Decelle, and B. Seoane, Cas- cade of phase transitions in the training of energy-based models*, Journal of Statistical Mechanics: Theory and Experiment2025, 074004 (2025)

  51. [59]

    Kuzborskij and Y

    I. Kuzborskij and Y. A. Yadkori, Low-rank bias, weight decay, and model merging in neural networks (2025), arXiv:2502.17340 [cs.LG]

  52. [60]

    A. Pal, F. Holtorf, A. Larsson, T. Loman, Utkarsh, F. Schäfer, Q. Qu, E. Alan, and C. Rackauckas, Non- linearSolve.jl: High-Performance and Robust Solvers for Systems of Nonlinear Equations in Julia, ACM Trans. Math. Softw.52, 1 (2026)

  53. [61]

    Sasagawa, On the finite escape phenomena for ma- trix Riccati equations, IEEE Transactions on Automatic Control27, 977 (1982)

    T. Sasagawa, On the finite escape phenomena for ma- trix Riccati equations, IEEE Transactions on Automatic Control27, 977 (1982). 20 Appendix A: Loss Landscape, Equations of Motion and Critical Points For a deep forward network withnlayers (n−1hidden layers), the equations of ...

  54. [62]

    Critical Points The critical points of the loss landscape are determined by∇W (i)L= 0. This directly gives rise to the0-balanced condition for any finiteβsince: 2βW (i)W (i)T ! =−(W (n)...W (i+1))TN(W (i)...W (1))T 2βW (i+1)TW (i+1) ! =−(W (n)...W (i+1))TN(W (i)...W (1))T   ...

  55. [63]

    Aligned (Co)Variances If the alignment condition (17) holds,Σxx andΣ yx can be ’diagonalized’ from the onset. In this frame of reference, (A4) can be solved by matricesW∗ =S W with singular valuessi≥0: n−1∑ k=0 S 2k n W (SW ˜Σxx− ˜Σyx)S 2(n−1−k) n W +nβS W ! = 0⇔s 3− 2 n i κi−...

  56. [64]

    Formal Connection to unregularized Case The critical point structure and (A4) can also be cast into a form reminiscent of theβ= 0case. Assuming that Wr has rankr, we can rewrite the condition (A4) formally by introducing a shifted variance matrix:   n−1∑ k=0 (WrWT r ) k n...

  57. [65]

    Here,Dis a diagonal matrix with entriesκi

    Non-Aligned (Co)Variances: Perturbatively (1 hidden layer) Based on criticality conditions (A4), we can perturbatively study the situation whereΣ xx andΣ yx = diag(η1,...,η rmax)are not aligned with, e.g.:Σ xx =D+ϵ ∑ iαi(eij +e ji)withϵ≪1. Here,Dis a diagonal matrix with entri...

  58. [66]

    For 1 hidden layer (n= 2):The first path corresponds to the direction dictated by the leading singular value ofΣyx, which isη 1 with (right) singular vector⃗ r1 =⃗ e1

    Paths through Loss Landscape & Symmetry Example Unregularized Case:Due to the simplicity of the chosen (co)variances, the geometry of the loss landscape can be visualized based on straight paths through the loss landscape in the directions of the Hessian eigendirections (where...

  59. [67]

    To analyze the Hessian, we are working with column vectorization, where vecc(W) := (W :,1,W :,2,...), such that the Hessian becomes a matrix

    1 hidden Layer Starting point of the analysis are the gradients of the loss (which are matrix valued;ˆNW :=W (2)W (1)Σxx−Σyx): ∂L ∂W (1) = 2 ( W (2)T ˆNW +βW (1) ) , ∂L ∂W (2) = 2 ( ˆNWW (1)T +βW (2) ) . To analyze the Hessian, we are working with column vectorization, where v...

  60. [68]

    The rankrof the criticalW (i)’s is crucial to understand the Hessian spectrum

    2 hidden Layers In the following, we derive the Hessian spectrum for the 2 hidden layer case under the alignment condition and Σxx =1 i,i for simplicity. The rankrof the criticalW (i)’s is crucial to understand the Hessian spectrum. It is convenient to first identify the (near...

  61. [69]

    The main point is that the flat directions andβ-directions do not depend on the model details as they are associated with (broken) symmetry transformations

    Deeper Networks:n>2Layers The strategy forn= 3(2hidden layers) can be extended to arbitrary depth. The main point is that the flat directions andβ-directions do not depend on the model details as they are associated with (broken) symmetry transformations. In the general case, ...

  62. [70]

    However, we should get access to some of the eigenvectors and eigenvalues (namely those that do not break the0-balance condition)

    Jacobian Due to the restriction to this subspace, we will not be able to extract all information about, e.g., the Hessian from this dynamics. However, we should get access to some of the eigenvectors and eigenvalues (namely those that do not break the0-balance condition). The ...

  63. [71]

    The special case of 1 hidden layer (n= 2) instead can also be approached differently by making use of the Riccati equation

    1 hidden Layer In principle, (D1) can be further evaluated for arbitrary depth. The special case of 1 hidden layer (n= 2) instead can also be approached differently by making use of the Riccati equation. Fordin =d out =d hidden =dandΣ xx =1, the macro dynamics can be written i...

  64. [72]

    2 hidden Layers (n= 3) In case of two hidden layers, the Jacobian is constructed from the derivatives, evaluated at the critical points: − ∂⃗ gW ∂WR =W R T⊗ ˆN+1 i,i⊗ ˆNWR +1i,i⊗WL ˆN=RS 2 i,iRT⊗(LS 3 o,iRT−Σyx)−2β1 i,i⊗LS 2 o,iR, =−βRS 2RT⊗LS−1R−RS 2RT⊗ ˆPo,⊥Σyx ˆPi,⊥−2β1 i,i...

  65. [73]

    Generalization to deeper Networks For networks with arbitrary depthn>2, the Jacobian can be analyzed in a fashion similar to then= 3case. For a finite rankrcritical point, the Jacobian in ther-sector reads (forΣxx =1andΣ yx diagonalized from the onset): 1 2Jr = ( 2β1S n r ⊗1−β...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.