Pith. sign in

REVIEW 6 cited by

PyHessian: Neural Networks Through the Lens of the Hessian

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.07145 v3 pith:UK2MV5ZJ submitted 2019-12-16 cs.LG cs.NAmath.NA

classification cs.LGcs.NAmath.NA
keywords hessiannetworksneuralbatchlandscapelossnormalizationpyhessian
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present PYHESSIAN, a new scalable framework that enables fast computation of Hessian (i.e., second-order derivative) information for deep neural networks. PYHESSIAN enables fast computations of the top Hessian eigenvalues, the Hessian trace, and the full Hessian eigenvalue/spectral density, and it supports distributed-memory execution on cloud/supercomputer systems and is available as open source. This general framework can be used to analyze neural network models, including the topology of the loss landscape (i.e., curvature information) to gain insight into the behavior of different models/optimizers. To illustrate this, we analyze the effect of residual connections and Batch Normalization layers on the trainability of neural networks. One recent claim, based on simpler first-order analysis, is that residual connections and Batch Normalization make the loss landscape smoother, thus making it easier for Stochastic Gradient Descent to converge to a good solution. Our extensive analysis shows new finer-scale insights, demonstrating that, while conventional wisdom is sometimes validated, in other cases it is simply incorrect. In particular, we find that Batch Normalization does not necessarily make the loss landscape smoother, especially for shallower networks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Modes of Sequence Models and Learning Coefficients

    cs.LG 2025-04 conditional novelty 7.0 of 10

    Under gradient and log-probability insensitivity assumptions, SGLD-based local learning coefficient estimates cannot distinguish a sequence distribution from its mode-truncated effective version.

  2. A Geometric Perspective on Stabilizing Value Conflict Resolution

    cs.LG 2026-07 conditional novelty 6.0 of 10

    An annealing-inspired chain-of-thought prompt lowers the sharpest loss-landscape curvature and improves moral reasoning benchmark scores in small LLMs.

  3. On the Superlinear Relationship between SGD Noise Covariance and Loss Landscape Curvature

    cs.LG 2026-02 reject novelty 5.0 of 10

    SGD noise covariance is claimed to follow the second moment of per-sample Hessians, giving a superlinear power law C_ii ∝ H_ii^γ with 1 ≤ γ ≤ 2.

  4. Investigating generalization capabilities of neural networks by means of loss landscapes and Hessian analysis

    cs.LG 2024-12 conditional novelty 5.0 of 10

    A Hessian-based criterion, KH05, increases when a model generalizes worse to a new dataset, offering a cheap generalization estimate, but the evidence is limited and the criterion's exponent was chosen post hoc.

  5. C-Flat++: Towards a More Efficient and Powerful Framework for Continual Learning

    cs.LG 2025-08 conditional novelty 4.0 of 10

    Adding zeroth- and first-order flatness penalties to continual learning losses yields small consistent accuracy gains across seven methods, with the gated C-Flat++ variant at roughly 30% of the update cost.

  6. The effects of Hessian eigenvalue spectral density type on the applicability of Hessian analysis to generalization capability assessment of neural networks

    cs.LG 2025-04 conditional novelty 3.0 of 10

    Hessian spectral density type, mainly positive versus mainly negative, determines whether Hessian-based generalization criteria apply, and mainly negative spectra are attributed to external gradient manipulation.

Pith tools