Pith. sign in

REVIEW 3 major objections 5 minor 68 references

The bulk of near-zero Hessian eigenvalues in neural nets are weakly broken continuous symmetries of the architecture.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-10 16:49 UTC pith:OA4P2EIV

load-bearing objection Clean eigenvector-level account of the Hessian bulk as weakly broken architectural symmetries; the math and diagnostics hold up. the 3 major comments →

arxiv 2607.07845 v1 pith:OA4P2EIV submitted 2026-07-08 cs.LG cond-mat.dis-nn

Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks

classification cs.LG cond-mat.dis-nn
keywords Hessian spectrumpseudo-Goldstone modesnetwork symmetriesloss landscape geometryFisher information matrixReLU networksoverparametrization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Neural-network loss landscapes are famously anisotropic: a few large-curvature directions sit above a dense bulk of near-zero Hessian eigenvalues. This paper argues that the bulk is not mysterious but is the spectral signature of continuous reparametrization symmetries that leave the network function unchanged. In deep linear networks those symmetries are exact, generate flat directions, and produce exact zero modes whose eigenvectors can be written down explicitly from the singular vectors of the weight matrices. Switching on a ReLU-type nonlinearity breaks the symmetries only weakly, lifting the zero modes into parametrically light pseudo-Goldstone modes that still live almost entirely inside the original symmetry subspace. High-curvature directions, by contrast, remain orthogonal to that subspace. The same two-tier structure appears in a trained three-layer network on CIFAR-10 and in a convolutional example, linking the Hessian bulk to approximate architectural symmetries.

Core claim

The bulk of near-zero Hessian eigenvalues consists of the weakly lifted pseudo-Goldstone modes of the continuous symmetries of the network parametrization. In linear networks the symmetries are exact and their generators form an explicit orthogonal basis of the null space; a ReLU nonlinearity breaks them weakly, so that high-curvature eigenvectors stay orthogonal to the symmetry subspace while the bulk eigenvectors remain almost entirely inside it.

What carries the argument

Explicit orthogonal generators of the GL-type interlayer symmetries (built from singular vectors of consecutive weight matrices) together with the eigenvector-overlap diagnostic that measures how much each Hessian mode of a nonlinear network lies inside that linear symmetry subspace.

Load-bearing premise

That the null space of a linear comparison model (nonlinearities removed, low-variance input directions zeroed) remains a faithful diagnostic of the bulk even after training on real data, without a general guarantee that training or non-Gaussian inputs do not mix the subspaces beyond the residual already measured.

What would settle it

Train a multilayer ReLU network to a genuine critical point, compute the leading Hessian eigenvectors, and check whether their overlap with the linear-symmetry subspace remains near zero for the high-curvature modes and near one for the bulk; a clear failure of that two-tier structure would falsify the claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • The bulk of near-zero modes is largely architectural and therefore persists across data sets and training algorithms that preserve the same continuous symmetries.
  • Because the same directions remain zero modes of the Fisher matrix at arbitrary parameters, natural-gradient and second-order methods automatically ignore or treat them specially.
  • Any architecture containing fully connected or convolutional blocks inherits an analogous bulk whose size is fixed by layer widths and over-parametrization count.
  • The adiabatic connection from linear to ReLU spectra supplies a controlled starting point for analytic approximations of the bulk eigenvalues.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the bulk is mostly architectural, pruning or regularization that deliberately targets the symmetry subspace may remove far more parameters than curvature-based pruning alone suggests.
  • The same diagnostic should apply to the fully connected blocks inside transformers; verifying the two-tier overlap there would test whether the mechanism survives attention and residual pathways.
  • Because the residual term in the Hessian can reintroduce curvature off critical points, the pseudo-Goldstone picture may degrade late in training when gradients no longer vanish.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper argues that the bulk of near-zero Hessian eigenvalues of neural-network training losses consists of weakly lifted pseudo-Goldstone modes of continuous architectural symmetries of the parametrization. In multilayer linear networks these symmetries are exact; the authors construct an explicit orthogonal basis of generators (via SVDs of the weight matrices) that spans the null space of the Hessian/Fisher and matches the known finite-eigenvalue count. A Leaky-ReLU deformation is treated as an explicit symmetry-breaking perturbation: the two-layer Gaussian student–teacher Fisher splits exactly into linear and absolute-value blocks, bulk eigenvalues rise as ε², and Kato perturbation theory supplies a confinement criterion (vanishing ε⁴ correction iff eigenvectors remain in the symmetry subspace). Eigenvector overlaps confirm that high-curvature modes are orthogonal to the symmetry subspace while the bulk lies inside it. The same diagnostic is applied to a three-layer student–teacher model, a trained three-layer ReLU MLP on CIFAR-10 (with whitening and cumulative coverage Ok), and a minimal convolutional network with nonlinear (circulant) generators.

Significance. If the eigenvector-level claim holds, it supplies a single architectural origin for the Hessian bulk that has been missing from the literature on outliers, Fisher spectra, and flat minima. The work goes beyond counting zero modes: it constructs the generators explicitly, derives the ε² lifting and confinement criterion analytically for the two-layer Gaussian case, and measures overlaps and cumulative coverage on both idealized and trained models. Code and data are released. The mechanism is expected to organize Fisher/Gauss–Newton spectra as well, with direct consequences for natural-gradient and second-order methods, and the convolutional example shows the diagnostic is not limited to fully connected layers. These are concrete, falsifiable contributions at the level of eigenvectors rather than eigenvalue counts alone.

major comments (3)
  1. End Matter C and SM §II establish the clean ε² law and confinement criterion only for the two-layer Gaussian student–teacher Fisher (exact block split, vanishing cross term by parity). In the three-layer SM case (Fig. S1) the ε² trend already bends before ε=1 and leading-mode overlaps level off near oi≈0.2 rather than near zero. The main-text claim that the bulk “lies almost entirely within” the symmetry subspace therefore needs an explicit scope statement: for deeper nets the adiabatic connection is only approximate and residual leakage is O(1). A short quantitative bound or additional depth-controlled experiment would make the generality claim load-bearing rather than extrapolative.
  2. CIFAR-10 section and SM §IV: the comparison subspace P mixes architectural GL-type symmetries with data-covariance flat directions obtained by zeroing low-variance input components. Dimension counting separates the two contributions (383 232 vs 26 748), and Ok≈0.95 bounds residual leakage of the unmeasured tail, but the non-whitened run (Fig. S2 right) shows that covariance-induced spread entangles the two mechanisms and smooths the overlap transition. The paper should state more sharply which fraction of the observed bulk is architectural versus data-driven, and whether the dramatic suppression of the leading Neff Ceff modes survives when the linear comparison Fisher is built without artificially zeroing variances.
  3. Eq. (1) and the Mexican-hat appendix correctly note that symmetry generators are exact Hessian zero modes only at critical points; off criticality the residual term can produce finite curvature. The CIFAR experiment evaluates the training-loss Hessian at a non-critical endpoint and compares to the Fisher null space of the linear model. While the observed two-tier structure is still striking, a brief check that the residual contribution along the measured bulk directions remains small (or a comparison to the Gauss–Newton matrix itself) would close the gap between the critical-point theory and the trained-network diagnostic.
minor comments (5)
  1. Fig. 2 caption and main text: the N rescaling modes that remain exactly zero for all ε are stated to lie “below the plotted range”; a short inset or explicit note that they are omitted would avoid the impression that the bulk starts above zero.
  2. Notation for the projector P and the overlap oi (Eq. 8) is introduced cleanly, but the random baseline orandom = rank(P)/d is quoted with different numerical values in different figures; a single consistent formula and the associated standard deviation for the CIFAR case would help the reader.
  3. The convolutional construction (SM §V) uses nonlinear generators involving C(w(1))⁻¹; a one-sentence remark in the main text that these are still continuous symmetries of the function (hence still produce Fisher zeros when exact) would clarify why the more general ϕ is needed.
  4. References [46,47] already count finite eigenvalues of deep linear networks; the novelty claim is correctly placed on the eigenvectors and the nonlinear extension, but a slightly sharper sentence distinguishing the present work from those counts would help.
  5. Typographical: “parametrization” is used consistently in the abstract/title; a few places in the SM switch to “parameterization.” Standardize.

Circularity Check

0 steps flagged

No circularity: symmetry generators defined independently of Hessian; overlaps and coverage are measured diagnostics, not fitted predictions.

full rationale

The derivation chain is self-contained and non-circular. Continuous symmetries of the linear network are defined by function-preserving transformations (GL+ action inserting M and M^{-1} between layers, Eq. 4), yielding generators φ'_A (Eq. 5) constructed via SVD of the weight matrices; these are shown to span the exact null space of H (and of F at arbitrary parameters) by direct verification of mutual orthogonality and dimension counting. The nonlinear case treats ReLU as an explicit perturbation of the linear network (Leaky-ReLU family, Eq. 6); the ε^{2} bulk scaling and confinement follow from the exact block decomposition of the Fisher (End Matter C, Eqs. C4–C6) under Gaussian parity, not from any fit. Overlaps o_i (Eq. 8) and cumulative coverage O_k (Eq. 9) are post-hoc measurements of the already-computed Hessian eigenvectors against the independently constructed linear symmetry projector P; the linear comparison model is obtained simply by deleting nonlinearities from the trained weights (or zeroing low-variance input directions), never by optimizing parameters to match the spectrum. The convolutional generators are likewise derived from the requirement that the transformed circulant remain circulant (SM V). No self-definitional loop, no fitted parameter re-labeled as prediction, and no load-bearing self-citation of an unverified uniqueness claim appear. The GitHub reference is solely for data availability.

Axiom & Free-Parameter Ledger

2 free parameters · 4 axioms · 1 invented entities

The central claim rests on standard linear-algebra and perturbation theory plus the architectural definition of continuous symmetries; no free parameters are fitted to force the bulk explanation, and the only invented language is the pseudo-Goldstone analogy (already standard in physics).

free parameters (2)
  • Neff (whitening cutoff) = 100
    Chosen as the top 100 covariance directions that capture 90 % of CIFAR variance; used to define the comparison subspace but not fitted to the overlap curves.
  • Leaky-ReLU leakage ε
    Control parameter swept from 0 to 1; not fitted, used only to demonstrate continuous lifting.
axioms (4)
  • standard math Continuous symmetries of the network function generate exact zero modes of the Fisher (and of the Hessian at critical points) via the identity H ϕ' + [∂θ(ϕ')ᵀ] ∂θ L = 0.
    Standard consequence of differentiating an invariance; cited from Kunin et al. and used throughout.
  • standard math Kato analytic perturbation theory for eigenvalues leaving a degenerate level applies to the Fisher family F(ε).
    Invoked in End Matter to obtain the ε² leading term and the non-negative ε⁴ leakage correction.
  • domain assumption For positively homogeneous activations the diagonal rescaling subgroup remains an exact symmetry for all ε.
    Used to keep N modes pinned at zero; standard for ReLU-type networks.
  • domain assumption At an interpolating student-teacher minimum the residual term in the Gauss-Newton decomposition vanishes, so H = F.
    Allows the exact spectral analysis of the two- and three-layer models.
invented entities (1)
  • pseudo-Goldstone modes of architectural symmetries independent evidence
    purpose: Name the weakly lifted bulk eigenvalues that remain confined to the linear symmetry subspace.
    Direct analogy to explicit symmetry breaking in physics; the modes themselves are ordinary Hessian eigenvectors, not new physical objects.

pith-pipeline@v1.1.0-grok45 · 28394 in / 2520 out tokens · 30318 ms · 2026-07-10T16:49:59.962808+00:00 · methodology

0 comments
read the original abstract

The Hessian of the training loss governs the local geometry of the loss landscape, yet despite existing explanations for its largest eigenvalues, the origin of the vast multitude of vanishingly small eigenvalues remains elusive. We argue that the bulk consists of the weakly lifted pseudo-Goldstone modes of the continuous symmetries of the network parametrization. In deep linear networks these symmetries are exact: they generate flat directions and hence exact zero modes, whose eigenvectors we construct explicitly. Introducing a ReLU nonlinearity as a perturbation, we show that it breaks these symmetries weakly and explicitly. Resolving the spectrum at the level of eigenvectors, we find that the high-curvature directions are orthogonal to the symmetry subspace, while the bulk lies almost entirely within it. We demonstrate the mechanism in a two-layer ReLU student--teacher model and in a network trained on CIFAR-10. A convolutional example demonstrates that the same diagnostic extends beyond fully connected layers. Together, these results link the Hessian bulk to weakly broken symmetries and clarify the origin of near-zero modes.

Figures

Figures reproduced from arXiv: 2607.07845 by Bernd Rosenow, Marcel K\"uhn.

Figure 1
Figure 1. Figure 1: FIG. 1. Symmetries in fully connected layers. Left: original [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2. Ranked Hessian eigenvalues of the two-layer Leaky [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 4
Figure 4. Figure 4: FIG. 4. Three-layer ReLU network trained on CIFAR-10, [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: FIG. 5. Convolutional student–teacher network (10), input [PITH_FULL_IMAGE:figures/full_fig_p005_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: FIG. 6. Zero Hessian eigenvalue from rotational symmetry, [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: FIG. 7. Two-layer student–teacher model ( [PITH_FULL_IMAGE:figures/full_fig_p008_7.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

68 extracted references · 68 canonical work pages · 7 internal anchors

  1. [1]

    LeCun, I

    Y. LeCun, I. Kanter, and S. A. Solla, Eigenvalues of co- variance matrices: Application to neural-network learn- ing, Phys. Rev. Lett.66, 2396 (1991)

  2. [2]

    Geiger, S

    M. Geiger, S. Spigler, S. d’Ascoli, L. Sagun, M. Baity- Jesi, G. Biroli, and M. Wyart, Jamming transition as a paradigm to understand the loss landscape of deep neural networks, Phys. Rev. E100, 012115 (2019)

  3. [3]

    Becker, Y

    S. Becker, Y. Zhang, and A. A. Lee, Geometry of En- ergy Landscapes and the Optimizability of Deep Neural Networks, Phys. Rev. Lett.124, 108301 (2020)

  4. [4]

    Atanasov, A

    A. Atanasov, A. Meterez, J. Simon, and C. Pehlevan, The Optimization Landscape of SGD Across the Fea- ture Learning Strength, inInternational Conference on Learning Representations(2025)

  5. [5]

    Cohen, S

    J. Cohen, S. Kaur, Y. Li, J. Z. Kolter, and A. Talwalkar, Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability, inInternational Conference on Learning Representations(2021)

  6. [6]

    Amari, Natural Gradient Works Efficiently in Learn- ing, Neural Computation10, 251 (1998)

    S. Amari, Natural Gradient Works Efficiently in Learn- ing, Neural Computation10, 251 (1998)

  7. [7]

    Martens, Deep learning via Hessian-free optimization, inProceedings of the 27th International Conference on International Conference on Machine Learning(Omni- press, 2010) p

    J. Martens, Deep learning via Hessian-free optimization, inProceedings of the 27th International Conference on International Conference on Machine Learning(Omni- press, 2010) p. 735–742

  8. [8]

    Martens and R

    J. Martens and R. Grosse, Optimizing Neural Networks with Kronecker-factored Approximate Curvature, inPro- ceedings of the 32nd International Conference on Ma- chine Learning, Vol. 37 (PMLR, 2015) pp. 2408–2417

  9. [9]

    Z. Yao, A. Gholami, S. Shen, M. Mustafa, K. Keutzer, and M. Mahoney, Adahessian: An adaptive second or- der optimizer for machine learning, inproceedings of the AAAI conference on artificial intelligence, Vol. 35 (2021) pp. 10665–10673

  10. [10]

    Hochreiter and J

    S. Hochreiter and J. Schmidhuber, Flat Minima, Neural Computation9, 1 (1997)

  11. [11]

    N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyan- skiy, and P. T. P. Tang, On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima, inInternational Conference on Learning Representations (2017)

  12. [12]

    Foret, A

    P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur, Sharpness-aware Minimization for Efficiently Improving Generalization, inInternational Conference on Learning Representations(2021)

  13. [13]

    Chaudhari, A

    P. Chaudhari, A. Choromanska, S. Soatto, Y. Le- Cun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina, Entropy-SGD: biasing gradient descent into wide valleys, Journal of Statistical Mechanics: Theory and Experiment , 124018 (2019)

  14. [14]

    Three Factors Influencing Minima in SGD

    S. Jastrzębski, Z. Kenton, D. Arpit, N. Ballas, A. Fis- cher, Y. Bengio, and A. Storkey, Three Factors Influenc- ing Minima in SGD (2018), arXiv:1711.04623

  15. [15]

    LeCun, J

    Y. LeCun, J. Denker, and S. Solla, Optimal Brain Dam- age, inAdvances in Neural Information Processing Sys- 6 tems, Vol. 2 (Morgan-Kaufmann, 1989)

  16. [16]

    Hassibi and D

    B. Hassibi and D. Stork, Second order derivatives for network pruning: Optimal Brain Surgeon, inAdvances in Neural Information Processing Systems, Vol. 5 (Morgan- Kaufmann, 1992)

  17. [17]

    S. P. Singh and D. Alistarh, WoodFisher: Efficient Second-Order Approximation for Neural Network Com- pression, inAdvances in Neural Information Process- ing Systems, Vol. 33 (Curran Associates, Inc., 2020) pp. 18098–18109

  18. [18]

    Frantar, S

    E. Frantar, S. Ashkboos, T. Hoefler, and D. Alis- tarh, OPTQ: Accurate Quantization for Generative Pre- trained Transformers, inThe Eleventh International Conference on Learning Representations(2023)

  19. [19]

    Z. Yao, A. Gholami, Q. Lei, K. Keutzer, and M. W. Ma- honey, Hessian-based Analysis of Large Batch Training and Robustness to Adversaries, inAdvances in Neural Information Processing Systems, Vol. 31 (Curran Asso- ciates, Inc., 2018)

  20. [20]

    Moosavi-Dezfooli, A

    S.-M. Moosavi-Dezfooli, A. Fawzi, J. Uesato, and P. Frossard, Robustness via Curvature Regularization, and Vice Versa, in2019 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR)(2019)pp. 9070–9078

  21. [21]

    Singla and S

    S. Singla and S. Feizi, Second-Order Provable Defenses against Adversarial Attacks, inProceedings of the 37th International Conference on Machine Learning, Proceed- ings of Machine Learning Research, Vol. 119 (PMLR,

  22. [22]

    Garipov, P

    T. Garipov, P. Izmailov, D. Podoprikhin, D. P. Vetrov, and A. G. Wilson, Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs, inAdvances in Neural Infor- mation Processing Systems, Vol. 31, edited by S. Bengio, H.Wallach, H.Larochelle, K.Grauman, N.Cesa-Bianchi, and R. Garnett (Curran Associates, Inc., 2018)

  23. [23]

    J. Brea, B. Simsek, B. Illing, and W. Gerstner, Weight- space symmetry in deep networks gives rise to permuta- tion saddles, connected by equal-loss valleys across the loss landscape (2019), arXiv:1907.02911

  24. [24]

    Ainsworth, J

    S. Ainsworth, J. Hayase, and S. Srinivasa, Git Re-Basin: Merging Models modulo Permutation Symmetries, in The Eleventh International Conference on Learning Rep- resentations(2023)

  25. [25]

    A. Ito, M. Yamada, and A. Kumagai, Linear Mode Con- nectivity between Multiple Models modulo Permutation Symmetries, inProceedings of the 42nd International Conference on Machine Learning, Proceedings of Ma- chine Learning Research, Vol. 267 (PMLR, 2025) pp. 26611–26626

  26. [26]

    Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond

    L. Sagun, L. Bottou, and Y. LeCun, Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond (2017), arXiv:1611.07476

  27. [27]

    Empirical Analysis of the Hessian of Over-Parametrized Neural Networks

    L. Sagun, U. Evci, V. U. Guney, Y. Dauphin, and L. Bottou, Empirical Analysis of the Hessian of Over- ParametrizedNeuralNetworks(2018),arXiv:1706.04454

  28. [28]

    Ghorbani, S

    B. Ghorbani, S. Krishnan, and Y. Xiao, An Investiga- tionintoNeuralNetOptimizationviaHessianEigenvalue Density, inProceedings of the 36th International Confer- ence on Machine Learning, Vol. 97 (PMLR, 2019) pp. 2232–2241

  29. [29]

    Papyan, Measurements of Three-Level Hierarchical StructureintheOutliersintheSpectrumofDeepnetHes- sians, inProceedings of the 36th International Conference on Machine Learning, Vol

    V. Papyan, Measurements of Three-Level Hierarchical StructureintheOutliersintheSpectrumofDeepnetHes- sians, inProceedings of the 36th International Conference on Machine Learning, Vol. 97 (2019) pp. 5012–5021

  30. [30]

    Papyan, Traces of Class/Cross-Class Structure Per- vade Deep Learning Spectra, Journal of Machine Learn- ing Research21, 1 (2020)

    V. Papyan, Traces of Class/Cross-Class Structure Per- vade Deep Learning Spectra, Journal of Machine Learn- ing Research21, 1 (2020)

  31. [31]

    Arjevani and M

    Y. Arjevani and M. Field, Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of Symme- try, inAdvances in Neural Information Processing Sys- tems, Vol. 33 (Curran Associates, Inc., 2020) pp. 5441– 5452

  32. [32]

    Arjevani and M

    Y. Arjevani and M. Field, Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II, inAdvances in Neural Informa- tion Processing Systems,Vol.34(CurranAssociates, Inc.,

  33. [33]

    Pennington and P

    J. Pennington and P. Worah, The Spectrum of the Fisher Information Matrix of a Single-Hidden-Layer Neural Net- work, inAdvances in Neural Information Processing Sys- tems, Vol. 31 (Curran Associates, Inc., 2018)

  34. [34]

    Karakida, S

    R. Karakida, S. Akaho, and S.-i. Amari, Universal Statis- tics of Fisher Information in Deep Neural Networks: Mean Field Approach, inProceedings of the Twenty- Second International Conference on Artificial Intelli- gence and Statistics, Proceedings of Machine Learning Research, Vol. 89 (PMLR, 2019) pp. 1032–1041

  35. [35]

    Karakida, S

    R. Karakida, S. Akaho, and S.-i. Amari, Pathological Spectra of the Fisher Information Metric and Its Vari- ants in Deep Neural Networks, Neural Computation33, 2274 (2021)

  36. [36]

    Emergent properties of the local geometry of neural loss landscapes

    S. Fort and S. Ganguli, Emergent properties of the local geometry of neural loss landscapes (2019), arXiv:1910.05929

  37. [37]

    Gradient Descent Happens in a Tiny Subspace

    G. Gur-Ari, D. A. Roberts, and E. Dyer, Gradi- ent Descent Happens in a Tiny Subspace (2018), arXiv:1812.04754

  38. [38]

    S. S. Du, W. Hu, and J. D. Lee, Algorithmic regulariza- tion in learning deep homogeneous models: Layers are automatically balanced, Advances in neural information processing systems31(2018)

  39. [39]

    Simsek, F

    B. Simsek, F. Ged, A. Jacot, F. Spadaro, C. Hongler, W. Gerstner, and J. Brea, Geometry of the Loss Land- scape in Overparameterized Neural Networks: Symme- tries and Invariances, inProceedings of the 38th Interna- tional Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 139 (PMLR, 2021) pp. 9722–9732

  40. [40]

    Kunin, J

    D. Kunin, J. Sagastuy-Brena, S. Ganguli, D. L. Yamins, and H. Tanaka, Neural Mechanics: Symmetry and Bro- ken Conservation Laws in Deep Learning Dynamics, in International Conference on Learning Representations (2021)

  41. [41]

    Tanaka and D

    H. Tanaka and D. Kunin, Noether’s Learning Dynamics: Role of Symmetry Breaking in Neural Networks, inAd- vances in Neural Information Processing Systems, edited by A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (2021)

  42. [42]

    B. Zhao, N. Dehmamy, R. Walters, and R. Yu, Symmetry Teleportation for Accelerated Optimization, inAdvances in Neural Information Processing Systems, Vol. 35 (Cur- ran Associates, Inc., 2022) pp. 16679–16690

  43. [43]

    Marcotte, R

    S. Marcotte, R. Gribonval, and G. Peyré, Abide by the law and follow the flow: conservation laws for gradient flows, inThirty-seventh Conference on Neural Informa- tion Processing Systems(2023)

  44. [44]

    Watanabe,Algebraic geometry and statistical learning theory, Vol

    S. Watanabe,Algebraic geometry and statistical learning theory, Vol. 25 (Cambridge university press, 2009). 7

  45. [45]

    K.FukumizuandS.Amari,Localminimaandplateausin hierarchical structures of multilayer perceptrons, Neural Networks13, 317 (2000)

  46. [46]

    S. P. Singh, G. Bachmann, and T. Hofmann, Analytic Insights into Structure and Rank of Neural Network Hes- sian Maps, inAdvances in Neural Information Process- ing Systems, Vol. 34 (Curran Associates, Inc., 2021) pp. 23914–23927

  47. [47]

    Bernacchia, M

    A. Bernacchia, M. Lengyel, and G. Hennequin, Exact natural gradient in deep linear networks and its applica- tion to the nonlinear case, inAdvances in Neural Infor- mation Processing Systems, Vol. 31 (Curran Associates, Inc., 2018)

  48. [48]

    C. M. Bishop and N. M. Nasrabadi,Pattern recognition and machine learning, Vol. 4 (Springer, 2006)

  49. [49]

    Heskes, On “natural” learning and pruning in multi- layered perceptrons, Neural Computation12, 881 (2000)

    T. Heskes, On “natural” learning and pruning in multi- layered perceptrons, Neural Computation12, 881 (2000)

  50. [50]

    J.Martens,NewInsightsandPerspectivesontheNatural Gradient Method, Journal of Machine Learning Research 21, 1 (2020)

  51. [51]

    See Supplemental Material for derivations, model exten- sions, and numerical details

  52. [52]

    Kawaguchi, Deep Learning without Poor Local Min- ima, inAdvances in Neural Information Processing Sys- tems, Vol

    K. Kawaguchi, Deep Learning without Poor Local Min- ima, inAdvances in Neural Information Processing Sys- tems, Vol. 29 (Curran Associates, Inc., 2016)

  53. [53]

    A. L. Maas, A. Y. Hannun, A. Y. Ng,et al., Rectifier nonlinearities improve neural network acoustic models, inProceedings of the 30th International Conference on Machine Learning, Vol. 28 (2013)

  54. [54]

    Karakida, S

    R. Karakida, S. Akaho, and S.-i. Amari, The Normal- ization Method for Alleviating Pathological Sharpness in Wide Neural Networks, inAdvances in Neural Informa- tion Processing Systems,Vol.32(CurranAssociates, Inc., 2019)

  55. [55]

    Krizhevsky,Learning multiple layers of features from tiny images, Tech

    A. Krizhevsky,Learning multiple layers of features from tiny images, Tech. Rep. (University of Toronto, Toronto, Ontario, 2009)

  56. [56]

    Zhang, S

    C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals, Understanding deep learning (still) requires rethinking generalization, Communications of the ACM64(2021)

  57. [57]

    independent compo- nents

    A. J. Bell and T. J. Sejnowski, The “independent compo- nents” of natural scenes are edge filters, Vision Research 37, 3327 (1997)

  58. [58]

    J. Lee, S. Schoenholz, J. Pennington, B. Adlam, L. Xiao, R. Novak, and J. Sohl-Dickstein, Finite Versus Infinite Neural Networks: an Empirical Study, inAdvances in Neural Information Processing Systems, Vol. 33 (Curran Associates, Inc., 2020) pp. 15156–15172

  59. [59]

    Fukushima, Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position, Biological cybernetics36, 193 (1980)

    K. Fukushima, Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position, Biological cybernetics36, 193 (1980)

  60. [60]

    LeCun, L

    Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, Gradient-based learning applied to document recogni- tion, Proceedings of the IEEE86, 2278 (1998)

  61. [61]

    Goodfellow, Y

    I. Goodfellow, Y. Bengio, and A. Courville,Deep Learn- ing(MIT press Cambridge, MA, USA, 2016)

  62. [62]

    R.PascanuandY.Bengio,Revisitingnaturalgradientfor deep networks, arXiv preprint arXiv:1301.3584 (2013)

  63. [63]

    Kühn and B

    M. Kühn and B. Rosenow, Github repos- itory,https://github.com/RosenowGroup/ approximate-symmetries-hessian(2026)

  64. [64]

    Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks

    T. Kato,Perturbation theory for linear operators, 2nd ed. (Springer, 1995). END MA TTER Appendix A: Curvature despite symmetry in the Mexican-hat picture—The mechanism of Eq. (1) is seen most simply in the rotation-invariant potential L(x, y) = (x2 +y 2 −1) 2 (Fig. 6). At the minimum(1,0) the symmetry direction(0,1)is a flat zero mode. Away from the minim...

  65. [65]

    For all (k,n )∈{1,...,M}2, at least one of the two blocks in Eq

    M≤min(N,C ). For all (k,n )∈{1,...,M}2, at least one of the two blocks in Eq. (S6) is nonzero. Hence all M2 generators are finite and linearly independent, yielding exactly M2 symmetries. This holds in particular in bottleneck configurations, where the naive parameter-counting argument (d−NC ) would underestimate the number of symmetry directions (M2)

  66. [66]

    In this case one would naively obtainM2 generators

    M≥max(N,C ). In this case one would naively obtainM2 generators. However, onlyM2−(M−N)(M−C) = MN +MC−NC of them are independent. Indeed, for (k,n ) such that boths(1) n = 0 ands(2) k = 0, the vector Eq. (S6) vanishes. Removing these (M−N)(M−C) zero vectors leaves precisely the expected number of symmetry generators

  67. [67]

    independent components

    C < M < N (and analogously N < M < C). Here the construction above yields M2 generators, while parameter counting predicts an additional ( N−M)(M−C) symmetry directions. These arise from additional generators that emerge from our definition via the singular value decomposition. For indices n > Mand k > C, additional generators are given by ϕ′(k>C,n>M) C<M...

  68. [68]

    15156–15172

    pp. 15156–15172. [S6] M. K¨ uhn and B. Rosenow, Github repository, https://github.com/RosenowGroup/approximate-symmetries-hessian (2026). 10 140 150 160 170 180 190 Index (i) 0 5 Eigenvalue ( i) 140 150 160 170 180 190 Index (i) 0.0 0.5 1.0 Overlap (oi) d K random baseline FIG. S3. Convolutional student–teacher network with input width 200, hidden width N...