Pith. sign in

hub

Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond

21 Pith papers cite this work, alongside 121 external citations. Polarity classification is still indexing.

21 Pith papers citing it
121 external citations · Pith
abstract

We look at the eigenvalues of the Hessian of a loss function before and after training. The eigenvalue distribution is seen to be composed of two parts, the bulk which is concentrated around zero, and the edges which are scattered away from zero. We present empirical evidence for the bulk indicating how over-parametrized the system is, and for the edges that depend on the input data.

hub tools

citation-role summary

background 3

citation-polarity summary

roles

background 3

polarities

background 3

representative citing papers

The Implicit Bias of Depth: From Neural Collapse to Softmax Codes

cs.LG · 2026-05-21 · unverdicted · novelty 7.0

Depth induces an implicit low-rank bias in deep unconstrained feature models trained with unregularized multiclass cross-entropy, promoting softmax codes over neural collapse via more efficient norm propagation.

AMUSE: Anytime Muon with Stable Gradient Evaluation

cs.LG · 2026-05-21 · accept · novelty 6.0

AMUSE stabilizes Muon with time-varying schedule-free gradient evaluation, improving the performance-iteration Pareto frontier without learning-rate schedules.

On the Convergence Analysis of Muon

stat.ML · 2025-05-29 · conditional · novelty 6.0

Muon's convergence rate depends on an average Hessian curvature along its update directions, which can be much smaller than the worst-case Lipschitz constant when Hessians are low-rank.

citing papers explorer

Showing 21 of 21 citing papers.