REVIEW 3 major objections 5 minor 68 references
The bulk of near-zero Hessian eigenvalues in neural nets are weakly broken continuous symmetries of the architecture.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-10 16:49 UTC pith:OA4P2EIV
load-bearing objection Clean eigenvector-level account of the Hessian bulk as weakly broken architectural symmetries; the math and diagnostics hold up. the 3 major comments →
Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The bulk of near-zero Hessian eigenvalues consists of the weakly lifted pseudo-Goldstone modes of the continuous symmetries of the network parametrization. In linear networks the symmetries are exact and their generators form an explicit orthogonal basis of the null space; a ReLU nonlinearity breaks them weakly, so that high-curvature eigenvectors stay orthogonal to the symmetry subspace while the bulk eigenvectors remain almost entirely inside it.
What carries the argument
Explicit orthogonal generators of the GL-type interlayer symmetries (built from singular vectors of consecutive weight matrices) together with the eigenvector-overlap diagnostic that measures how much each Hessian mode of a nonlinear network lies inside that linear symmetry subspace.
Load-bearing premise
That the null space of a linear comparison model (nonlinearities removed, low-variance input directions zeroed) remains a faithful diagnostic of the bulk even after training on real data, without a general guarantee that training or non-Gaussian inputs do not mix the subspaces beyond the residual already measured.
What would settle it
Train a multilayer ReLU network to a genuine critical point, compute the leading Hessian eigenvectors, and check whether their overlap with the linear-symmetry subspace remains near zero for the high-curvature modes and near one for the bulk; a clear failure of that two-tier structure would falsify the claim.
If this is right
- The bulk of near-zero modes is largely architectural and therefore persists across data sets and training algorithms that preserve the same continuous symmetries.
- Because the same directions remain zero modes of the Fisher matrix at arbitrary parameters, natural-gradient and second-order methods automatically ignore or treat them specially.
- Any architecture containing fully connected or convolutional blocks inherits an analogous bulk whose size is fixed by layer widths and over-parametrization count.
- The adiabatic connection from linear to ReLU spectra supplies a controlled starting point for analytic approximations of the bulk eigenvalues.
Where Pith is reading between the lines
- If the bulk is mostly architectural, pruning or regularization that deliberately targets the symmetry subspace may remove far more parameters than curvature-based pruning alone suggests.
- The same diagnostic should apply to the fully connected blocks inside transformers; verifying the two-tier overlap there would test whether the mechanism survives attention and residual pathways.
- Because the residual term in the Hessian can reintroduce curvature off critical points, the pseudo-Goldstone picture may degrade late in training when gradients no longer vanish.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that the bulk of near-zero Hessian eigenvalues of neural-network training losses consists of weakly lifted pseudo-Goldstone modes of continuous architectural symmetries of the parametrization. In multilayer linear networks these symmetries are exact; the authors construct an explicit orthogonal basis of generators (via SVDs of the weight matrices) that spans the null space of the Hessian/Fisher and matches the known finite-eigenvalue count. A Leaky-ReLU deformation is treated as an explicit symmetry-breaking perturbation: the two-layer Gaussian student–teacher Fisher splits exactly into linear and absolute-value blocks, bulk eigenvalues rise as ε², and Kato perturbation theory supplies a confinement criterion (vanishing ε⁴ correction iff eigenvectors remain in the symmetry subspace). Eigenvector overlaps confirm that high-curvature modes are orthogonal to the symmetry subspace while the bulk lies inside it. The same diagnostic is applied to a three-layer student–teacher model, a trained three-layer ReLU MLP on CIFAR-10 (with whitening and cumulative coverage Ok), and a minimal convolutional network with nonlinear (circulant) generators.
Significance. If the eigenvector-level claim holds, it supplies a single architectural origin for the Hessian bulk that has been missing from the literature on outliers, Fisher spectra, and flat minima. The work goes beyond counting zero modes: it constructs the generators explicitly, derives the ε² lifting and confinement criterion analytically for the two-layer Gaussian case, and measures overlaps and cumulative coverage on both idealized and trained models. Code and data are released. The mechanism is expected to organize Fisher/Gauss–Newton spectra as well, with direct consequences for natural-gradient and second-order methods, and the convolutional example shows the diagnostic is not limited to fully connected layers. These are concrete, falsifiable contributions at the level of eigenvectors rather than eigenvalue counts alone.
major comments (3)
- End Matter C and SM §II establish the clean ε² law and confinement criterion only for the two-layer Gaussian student–teacher Fisher (exact block split, vanishing cross term by parity). In the three-layer SM case (Fig. S1) the ε² trend already bends before ε=1 and leading-mode overlaps level off near oi≈0.2 rather than near zero. The main-text claim that the bulk “lies almost entirely within” the symmetry subspace therefore needs an explicit scope statement: for deeper nets the adiabatic connection is only approximate and residual leakage is O(1). A short quantitative bound or additional depth-controlled experiment would make the generality claim load-bearing rather than extrapolative.
- CIFAR-10 section and SM §IV: the comparison subspace P mixes architectural GL-type symmetries with data-covariance flat directions obtained by zeroing low-variance input components. Dimension counting separates the two contributions (383 232 vs 26 748), and Ok≈0.95 bounds residual leakage of the unmeasured tail, but the non-whitened run (Fig. S2 right) shows that covariance-induced spread entangles the two mechanisms and smooths the overlap transition. The paper should state more sharply which fraction of the observed bulk is architectural versus data-driven, and whether the dramatic suppression of the leading Neff Ceff modes survives when the linear comparison Fisher is built without artificially zeroing variances.
- Eq. (1) and the Mexican-hat appendix correctly note that symmetry generators are exact Hessian zero modes only at critical points; off criticality the residual term can produce finite curvature. The CIFAR experiment evaluates the training-loss Hessian at a non-critical endpoint and compares to the Fisher null space of the linear model. While the observed two-tier structure is still striking, a brief check that the residual contribution along the measured bulk directions remains small (or a comparison to the Gauss–Newton matrix itself) would close the gap between the critical-point theory and the trained-network diagnostic.
minor comments (5)
- Fig. 2 caption and main text: the N rescaling modes that remain exactly zero for all ε are stated to lie “below the plotted range”; a short inset or explicit note that they are omitted would avoid the impression that the bulk starts above zero.
- Notation for the projector P and the overlap oi (Eq. 8) is introduced cleanly, but the random baseline orandom = rank(P)/d is quoted with different numerical values in different figures; a single consistent formula and the associated standard deviation for the CIFAR case would help the reader.
- The convolutional construction (SM §V) uses nonlinear generators involving C(w(1))⁻¹; a one-sentence remark in the main text that these are still continuous symmetries of the function (hence still produce Fisher zeros when exact) would clarify why the more general ϕ is needed.
- References [46,47] already count finite eigenvalues of deep linear networks; the novelty claim is correctly placed on the eigenvectors and the nonlinear extension, but a slightly sharper sentence distinguishing the present work from those counts would help.
- Typographical: “parametrization” is used consistently in the abstract/title; a few places in the SM switch to “parameterization.” Standardize.
Circularity Check
No circularity: symmetry generators defined independently of Hessian; overlaps and coverage are measured diagnostics, not fitted predictions.
full rationale
The derivation chain is self-contained and non-circular. Continuous symmetries of the linear network are defined by function-preserving transformations (GL+ action inserting M and M^{-1} between layers, Eq. 4), yielding generators φ'_A (Eq. 5) constructed via SVD of the weight matrices; these are shown to span the exact null space of H (and of F at arbitrary parameters) by direct verification of mutual orthogonality and dimension counting. The nonlinear case treats ReLU as an explicit perturbation of the linear network (Leaky-ReLU family, Eq. 6); the ε^{2} bulk scaling and confinement follow from the exact block decomposition of the Fisher (End Matter C, Eqs. C4–C6) under Gaussian parity, not from any fit. Overlaps o_i (Eq. 8) and cumulative coverage O_k (Eq. 9) are post-hoc measurements of the already-computed Hessian eigenvectors against the independently constructed linear symmetry projector P; the linear comparison model is obtained simply by deleting nonlinearities from the trained weights (or zeroing low-variance input directions), never by optimizing parameters to match the spectrum. The convolutional generators are likewise derived from the requirement that the transformed circulant remain circulant (SM V). No self-definitional loop, no fitted parameter re-labeled as prediction, and no load-bearing self-citation of an unverified uniqueness claim appear. The GitHub reference is solely for data availability.
Axiom & Free-Parameter Ledger
free parameters (2)
- Neff (whitening cutoff) =
100
- Leaky-ReLU leakage ε
axioms (4)
- standard math Continuous symmetries of the network function generate exact zero modes of the Fisher (and of the Hessian at critical points) via the identity H ϕ' + [∂θ(ϕ')ᵀ] ∂θ L = 0.
- standard math Kato analytic perturbation theory for eigenvalues leaving a degenerate level applies to the Fisher family F(ε).
- domain assumption For positively homogeneous activations the diagonal rescaling subgroup remains an exact symmetry for all ε.
- domain assumption At an interpolating student-teacher minimum the residual term in the Gauss-Newton decomposition vanishes, so H = F.
invented entities (1)
-
pseudo-Goldstone modes of architectural symmetries
independent evidence
read the original abstract
The Hessian of the training loss governs the local geometry of the loss landscape, yet despite existing explanations for its largest eigenvalues, the origin of the vast multitude of vanishingly small eigenvalues remains elusive. We argue that the bulk consists of the weakly lifted pseudo-Goldstone modes of the continuous symmetries of the network parametrization. In deep linear networks these symmetries are exact: they generate flat directions and hence exact zero modes, whose eigenvectors we construct explicitly. Introducing a ReLU nonlinearity as a perturbation, we show that it breaks these symmetries weakly and explicitly. Resolving the spectrum at the level of eigenvectors, we find that the high-curvature directions are orthogonal to the symmetry subspace, while the bulk lies almost entirely within it. We demonstrate the mechanism in a two-layer ReLU student--teacher model and in a network trained on CIFAR-10. A convolutional example demonstrates that the same diagnostic extends beyond fully connected layers. Together, these results link the Hessian bulk to weakly broken symmetries and clarify the origin of near-zero modes.
Figures
Reference graph
Works this paper leans on
- [1]
- [2]
- [3]
-
[4]
A. Atanasov, A. Meterez, J. Simon, and C. Pehlevan, The Optimization Landscape of SGD Across the Fea- ture Learning Strength, inInternational Conference on Learning Representations(2025)
work page 2025
- [5]
-
[6]
Amari, Natural Gradient Works Efficiently in Learn- ing, Neural Computation10, 251 (1998)
S. Amari, Natural Gradient Works Efficiently in Learn- ing, Neural Computation10, 251 (1998)
work page 1998
-
[7]
J. Martens, Deep learning via Hessian-free optimization, inProceedings of the 27th International Conference on International Conference on Machine Learning(Omni- press, 2010) p. 735–742
work page 2010
-
[8]
J. Martens and R. Grosse, Optimizing Neural Networks with Kronecker-factored Approximate Curvature, inPro- ceedings of the 32nd International Conference on Ma- chine Learning, Vol. 37 (PMLR, 2015) pp. 2408–2417
work page 2015
-
[9]
Z. Yao, A. Gholami, S. Shen, M. Mustafa, K. Keutzer, and M. Mahoney, Adahessian: An adaptive second or- der optimizer for machine learning, inproceedings of the AAAI conference on artificial intelligence, Vol. 35 (2021) pp. 10665–10673
work page 2021
-
[10]
S. Hochreiter and J. Schmidhuber, Flat Minima, Neural Computation9, 1 (1997)
work page 1997
-
[11]
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyan- skiy, and P. T. P. Tang, On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima, inInternational Conference on Learning Representations (2017)
work page 2017
- [12]
-
[13]
P. Chaudhari, A. Choromanska, S. Soatto, Y. Le- Cun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina, Entropy-SGD: biasing gradient descent into wide valleys, Journal of Statistical Mechanics: Theory and Experiment , 124018 (2019)
work page 2019
-
[14]
Three Factors Influencing Minima in SGD
S. Jastrzębski, Z. Kenton, D. Arpit, N. Ballas, A. Fis- cher, Y. Bengio, and A. Storkey, Three Factors Influenc- ing Minima in SGD (2018), arXiv:1711.04623
work page internal anchor Pith review Pith/arXiv arXiv 2018
- [15]
-
[16]
B. Hassibi and D. Stork, Second order derivatives for network pruning: Optimal Brain Surgeon, inAdvances in Neural Information Processing Systems, Vol. 5 (Morgan- Kaufmann, 1992)
work page 1992
-
[17]
S. P. Singh and D. Alistarh, WoodFisher: Efficient Second-Order Approximation for Neural Network Com- pression, inAdvances in Neural Information Process- ing Systems, Vol. 33 (Curran Associates, Inc., 2020) pp. 18098–18109
work page 2020
-
[18]
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alis- tarh, OPTQ: Accurate Quantization for Generative Pre- trained Transformers, inThe Eleventh International Conference on Learning Representations(2023)
work page 2023
-
[19]
Z. Yao, A. Gholami, Q. Lei, K. Keutzer, and M. W. Ma- honey, Hessian-based Analysis of Large Batch Training and Robustness to Adversaries, inAdvances in Neural Information Processing Systems, Vol. 31 (Curran Asso- ciates, Inc., 2018)
work page 2018
-
[20]
S.-M. Moosavi-Dezfooli, A. Fawzi, J. Uesato, and P. Frossard, Robustness via Curvature Regularization, and Vice Versa, in2019 IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR)(2019)pp. 9070–9078
work page 2019
-
[21]
S. Singla and S. Feizi, Second-Order Provable Defenses against Adversarial Attacks, inProceedings of the 37th International Conference on Machine Learning, Proceed- ings of Machine Learning Research, Vol. 119 (PMLR,
-
[22]
T. Garipov, P. Izmailov, D. Podoprikhin, D. P. Vetrov, and A. G. Wilson, Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs, inAdvances in Neural Infor- mation Processing Systems, Vol. 31, edited by S. Bengio, H.Wallach, H.Larochelle, K.Grauman, N.Cesa-Bianchi, and R. Garnett (Curran Associates, Inc., 2018)
work page 2018
-
[23]
J. Brea, B. Simsek, B. Illing, and W. Gerstner, Weight- space symmetry in deep networks gives rise to permuta- tion saddles, connected by equal-loss valleys across the loss landscape (2019), arXiv:1907.02911
work page internal anchor Pith review Pith/arXiv arXiv 2019
-
[24]
S. Ainsworth, J. Hayase, and S. Srinivasa, Git Re-Basin: Merging Models modulo Permutation Symmetries, in The Eleventh International Conference on Learning Rep- resentations(2023)
work page 2023
-
[25]
A. Ito, M. Yamada, and A. Kumagai, Linear Mode Con- nectivity between Multiple Models modulo Permutation Symmetries, inProceedings of the 42nd International Conference on Machine Learning, Proceedings of Ma- chine Learning Research, Vol. 267 (PMLR, 2025) pp. 26611–26626
work page 2025
-
[26]
Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond
L. Sagun, L. Bottou, and Y. LeCun, Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond (2017), arXiv:1611.07476
work page internal anchor Pith review Pith/arXiv arXiv 2017
-
[27]
Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
L. Sagun, U. Evci, V. U. Guney, Y. Dauphin, and L. Bottou, Empirical Analysis of the Hessian of Over- ParametrizedNeuralNetworks(2018),arXiv:1706.04454
work page internal anchor Pith review Pith/arXiv arXiv 2018
-
[28]
B. Ghorbani, S. Krishnan, and Y. Xiao, An Investiga- tionintoNeuralNetOptimizationviaHessianEigenvalue Density, inProceedings of the 36th International Confer- ence on Machine Learning, Vol. 97 (PMLR, 2019) pp. 2232–2241
work page 2019
-
[29]
V. Papyan, Measurements of Three-Level Hierarchical StructureintheOutliersintheSpectrumofDeepnetHes- sians, inProceedings of the 36th International Conference on Machine Learning, Vol. 97 (2019) pp. 5012–5021
work page 2019
-
[30]
V. Papyan, Traces of Class/Cross-Class Structure Per- vade Deep Learning Spectra, Journal of Machine Learn- ing Research21, 1 (2020)
work page 2020
-
[31]
Y. Arjevani and M. Field, Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of Symme- try, inAdvances in Neural Information Processing Sys- tems, Vol. 33 (Curran Associates, Inc., 2020) pp. 5441– 5452
work page 2020
-
[32]
Y. Arjevani and M. Field, Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II, inAdvances in Neural Informa- tion Processing Systems,Vol.34(CurranAssociates, Inc.,
-
[33]
J. Pennington and P. Worah, The Spectrum of the Fisher Information Matrix of a Single-Hidden-Layer Neural Net- work, inAdvances in Neural Information Processing Sys- tems, Vol. 31 (Curran Associates, Inc., 2018)
work page 2018
-
[34]
R. Karakida, S. Akaho, and S.-i. Amari, Universal Statis- tics of Fisher Information in Deep Neural Networks: Mean Field Approach, inProceedings of the Twenty- Second International Conference on Artificial Intelli- gence and Statistics, Proceedings of Machine Learning Research, Vol. 89 (PMLR, 2019) pp. 1032–1041
work page 2019
-
[35]
R. Karakida, S. Akaho, and S.-i. Amari, Pathological Spectra of the Fisher Information Metric and Its Vari- ants in Deep Neural Networks, Neural Computation33, 2274 (2021)
work page 2021
-
[36]
Emergent properties of the local geometry of neural loss landscapes
S. Fort and S. Ganguli, Emergent properties of the local geometry of neural loss landscapes (2019), arXiv:1910.05929
work page internal anchor Pith review Pith/arXiv arXiv 2019
-
[37]
Gradient Descent Happens in a Tiny Subspace
G. Gur-Ari, D. A. Roberts, and E. Dyer, Gradi- ent Descent Happens in a Tiny Subspace (2018), arXiv:1812.04754
work page internal anchor Pith review Pith/arXiv arXiv 2018
-
[38]
S. S. Du, W. Hu, and J. D. Lee, Algorithmic regulariza- tion in learning deep homogeneous models: Layers are automatically balanced, Advances in neural information processing systems31(2018)
work page 2018
-
[39]
B. Simsek, F. Ged, A. Jacot, F. Spadaro, C. Hongler, W. Gerstner, and J. Brea, Geometry of the Loss Land- scape in Overparameterized Neural Networks: Symme- tries and Invariances, inProceedings of the 38th Interna- tional Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 139 (PMLR, 2021) pp. 9722–9732
work page 2021
- [40]
-
[41]
H. Tanaka and D. Kunin, Noether’s Learning Dynamics: Role of Symmetry Breaking in Neural Networks, inAd- vances in Neural Information Processing Systems, edited by A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan (2021)
work page 2021
-
[42]
B. Zhao, N. Dehmamy, R. Walters, and R. Yu, Symmetry Teleportation for Accelerated Optimization, inAdvances in Neural Information Processing Systems, Vol. 35 (Cur- ran Associates, Inc., 2022) pp. 16679–16690
work page 2022
-
[43]
S. Marcotte, R. Gribonval, and G. Peyré, Abide by the law and follow the flow: conservation laws for gradient flows, inThirty-seventh Conference on Neural Informa- tion Processing Systems(2023)
work page 2023
-
[44]
Watanabe,Algebraic geometry and statistical learning theory, Vol
S. Watanabe,Algebraic geometry and statistical learning theory, Vol. 25 (Cambridge university press, 2009). 7
work page 2009
-
[45]
K.FukumizuandS.Amari,Localminimaandplateausin hierarchical structures of multilayer perceptrons, Neural Networks13, 317 (2000)
work page 2000
-
[46]
S. P. Singh, G. Bachmann, and T. Hofmann, Analytic Insights into Structure and Rank of Neural Network Hes- sian Maps, inAdvances in Neural Information Process- ing Systems, Vol. 34 (Curran Associates, Inc., 2021) pp. 23914–23927
work page 2021
-
[47]
A. Bernacchia, M. Lengyel, and G. Hennequin, Exact natural gradient in deep linear networks and its applica- tion to the nonlinear case, inAdvances in Neural Infor- mation Processing Systems, Vol. 31 (Curran Associates, Inc., 2018)
work page 2018
-
[48]
C. M. Bishop and N. M. Nasrabadi,Pattern recognition and machine learning, Vol. 4 (Springer, 2006)
work page 2006
-
[49]
T. Heskes, On “natural” learning and pruning in multi- layered perceptrons, Neural Computation12, 881 (2000)
work page 2000
-
[50]
J.Martens,NewInsightsandPerspectivesontheNatural Gradient Method, Journal of Machine Learning Research 21, 1 (2020)
work page 2020
-
[51]
See Supplemental Material for derivations, model exten- sions, and numerical details
-
[52]
K. Kawaguchi, Deep Learning without Poor Local Min- ima, inAdvances in Neural Information Processing Sys- tems, Vol. 29 (Curran Associates, Inc., 2016)
work page 2016
-
[53]
A. L. Maas, A. Y. Hannun, A. Y. Ng,et al., Rectifier nonlinearities improve neural network acoustic models, inProceedings of the 30th International Conference on Machine Learning, Vol. 28 (2013)
work page 2013
-
[54]
R. Karakida, S. Akaho, and S.-i. Amari, The Normal- ization Method for Alleviating Pathological Sharpness in Wide Neural Networks, inAdvances in Neural Informa- tion Processing Systems,Vol.32(CurranAssociates, Inc., 2019)
work page 2019
-
[55]
Krizhevsky,Learning multiple layers of features from tiny images, Tech
A. Krizhevsky,Learning multiple layers of features from tiny images, Tech. Rep. (University of Toronto, Toronto, Ontario, 2009)
work page 2009
- [56]
-
[57]
A. J. Bell and T. J. Sejnowski, The “independent compo- nents” of natural scenes are edge filters, Vision Research 37, 3327 (1997)
work page 1997
-
[58]
J. Lee, S. Schoenholz, J. Pennington, B. Adlam, L. Xiao, R. Novak, and J. Sohl-Dickstein, Finite Versus Infinite Neural Networks: an Empirical Study, inAdvances in Neural Information Processing Systems, Vol. 33 (Curran Associates, Inc., 2020) pp. 15156–15172
work page 2020
-
[59]
K. Fukushima, Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position, Biological cybernetics36, 193 (1980)
work page 1980
- [60]
-
[61]
I. Goodfellow, Y. Bengio, and A. Courville,Deep Learn- ing(MIT press Cambridge, MA, USA, 2016)
work page 2016
-
[62]
R.PascanuandY.Bengio,Revisitingnaturalgradientfor deep networks, arXiv preprint arXiv:1301.3584 (2013)
work page internal anchor Pith review Pith/arXiv arXiv 2013
-
[63]
M. Kühn and B. Rosenow, Github repos- itory,https://github.com/RosenowGroup/ approximate-symmetries-hessian(2026)
work page 2026
-
[64]
Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks
T. Kato,Perturbation theory for linear operators, 2nd ed. (Springer, 1995). END MA TTER Appendix A: Curvature despite symmetry in the Mexican-hat picture—The mechanism of Eq. (1) is seen most simply in the rotation-invariant potential L(x, y) = (x2 +y 2 −1) 2 (Fig. 6). At the minimum(1,0) the symmetry direction(0,1)is a flat zero mode. Away from the minim...
work page 1995
-
[65]
For all (k,n )∈{1,...,M}2, at least one of the two blocks in Eq
M≤min(N,C ). For all (k,n )∈{1,...,M}2, at least one of the two blocks in Eq. (S6) is nonzero. Hence all M2 generators are finite and linearly independent, yielding exactly M2 symmetries. This holds in particular in bottleneck configurations, where the naive parameter-counting argument (d−NC ) would underestimate the number of symmetry directions (M2)
-
[66]
In this case one would naively obtainM2 generators
M≥max(N,C ). In this case one would naively obtainM2 generators. However, onlyM2−(M−N)(M−C) = MN +MC−NC of them are independent. Indeed, for (k,n ) such that boths(1) n = 0 ands(2) k = 0, the vector Eq. (S6) vanishes. Removing these (M−N)(M−C) zero vectors leaves precisely the expected number of symmetry generators
-
[67]
C < M < N (and analogously N < M < C). Here the construction above yields M2 generators, while parameter counting predicts an additional ( N−M)(M−C) symmetry directions. These arise from additional generators that emerge from our definition via the singular value decomposition. For indices n > Mand k > C, additional generators are given by ϕ′(k>C,n>M) C<M...
work page 2000
-
[68]
pp. 15156–15172. [S6] M. K¨ uhn and B. Rosenow, Github repository, https://github.com/RosenowGroup/approximate-symmetries-hessian (2026). 10 140 150 160 170 180 190 Index (i) 0 5 Eigenvalue ( i) 140 150 160 170 180 190 Index (i) 0.0 0.5 1.0 Overlap (oi) d K random baseline FIG. S3. Convolutional student–teacher network with input width 200, hidden width N...
work page 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.