Pith. sign in

REVIEW 9 cited by

Emergent properties of the local geometry of neural loss landscapes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.05929 v1 pith:K4FQ62ZX submitted 2019-10-14 cs.LG cs.NEstat.ML

Emergent properties of the local geometry of neural loss landscapes

classification cs.LG cs.NEstat.ML
keywords neuralcurvaturepropertiesdimensionaldirectionslandscapeslocalloss
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The local geometry of high dimensional neural network loss landscapes can both challenge our cherished theoretical intuitions as well as dramatically impact the practical success of neural network training. Indeed recent works have observed 4 striking local properties of neural loss landscapes on classification tasks: (1) the landscape exhibits exactly $C$ directions of high positive curvature, where $C$ is the number of classes; (2) gradient directions are largely confined to this extremely low dimensional subspace of positive Hessian curvature, leaving the vast majority of directions in weight space unexplored; (3) gradient descent transiently explores intermediate regions of higher positive curvature before eventually finding flatter minima; (4) training can be successful even when confined to low dimensional {\it random} affine hyperplanes, as long as these hyperplanes intersect a Goldilocks zone of higher than average curvature. We develop a simple theoretical model of gradients and Hessians, justified by numerical experiments on architectures and datasets used in practice, that {\it simultaneously} accounts for all $4$ of these surprising and seemingly unrelated properties. Our unified model provides conceptual insights into the emergence of these properties and makes connections with diverse topics in neural networks, random matrix theory, and spin glasses, including the neural tangent kernel, BBP phase transitions, and Derrida's random energy model.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks

    cs.LG 2026-07 accept novelty 7.0

    The Hessian bulk consists of weakly broken continuous symmetries of the network parametrization, exact zeros in linear nets that ReLU lifts as pseudo-Goldstone modes whose eigenvectors stay in the symmetry subspace.

  2. Why Muon Outperforms Adam: A Curvature Perspective

    cs.LG 2026-06 conditional novelty 7.0

    Muon outperforms Adam by reducing curvature penalty via lower Normalized Directional Sharpness, as shown via Taylor approximation on LLM training and proven on stylized quadratic problems with heterogeneous curvature.

  3. Hessian Surgery: Class-Targeted Post-Hoc Rebalancing via Hessian Spike Perturbation

    cs.LG 2026-05 unverdicted novelty 7.0

    Spectral Surgery uses a sensitivity matrix and constrained optimization to perturb weights along Hessian spike eigenvectors, rebalancing per-class accuracy on CIFAR-10 and ISIC-2019 without retraining.

  4. Hessian Surgery: Class-Targeted Post-Hoc Rebalancing via Hessian Spike Perturbation

    cs.LG 2026-05 unverdicted novelty 7.0

    Hessian Surgery perturbs trained model weights along Hessian spike eigenvectors via a sensitivity matrix and constrained optimization to rebalance per-class accuracy on CIFAR-10 and ISIC-2019 without retraining.

  5. Dimensional Criticality at Grokking Across MLPs and Transformers

    cs.LG 2026-04 unverdicted novelty 7.0

    Effective cascade dimension D(t) crosses D=1 at the grokking transition in MLPs and Transformers, with opposite directions for modular addition versus XOR, consistent with attraction to a shared critical manifold.

  6. Characterizing Optimizer-Dependent Training Dynamics Through Hessian Eigenvector Displacement and Localization

    cs.LG 2026-06 unverdicted novelty 6.0

    Hessian eigenvector displacement and inverse participation ratio metrics show SGD stabilizing leading curvature directions while Adam causes more reorganization and parameter localization in MLP training.

  7. Grokking as Dimensional Phase Transition in Neural Networks

    cs.LG 2026-04 unverdicted novelty 6.0

    Grokking occurs as the effective dimensionality of the gradient field transitions from sub-diffusive to super-diffusive at the onset of generalization, exhibiting self-organized criticality.

  8. The Stable Recovery Manifold: Geometric Principles Governing Recoverability in Continual Learning

    cs.LG 2026-06 unverdicted novelty 5.0

    Empirical analysis of sequential ResNet-18 training on Split CIFAR-100 finds stable recovery subspace dimensionality supporting the Stable Recovery Manifold hypothesis that forgotten knowledge remains compactly decodable.

  9. Comparing Classical Simulation and Sample-Based Learning of Quantum Systems

    quant-ph 2026-05 unverdicted novelty 5.0

    Empirical study finds neural-network learning difficulty (via Hessian eigenvalue and random subspace optimization) correlates with classical simulation hardness parameterized by MPS bond dimension and T-gate count.