Pith. sign in

REVIEW 1 cited by

Charting the Topography of the Neural Network Landscape with Thermal-Like Noise

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.01335 v2 pith:PYVAOP2L submitted 2023-04-03 cond-mat.stat-mech cs.LG

classification cond-mat.stat-mechcs.LG
keywords dynamicslandscapestatisticsboundaryclassificationdatadecisiondimension
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The training of neural networks is a complex, high-dimensional, non-convex and noisy optimization problem whose theoretical understanding is interesting both from an applicative perspective and for fundamental reasons. A core challenge is to understand the geometry and topography of the landscape that guides the optimization. In this work, we employ standard Statistical Mechanics methods, namely, phase-space exploration using Langevin dynamics, to study this landscape for an over-parameterized fully connected network performing a classification task on random data. Analyzing the fluctuation statistics, in analogy to thermal dynamics at a constant temperature, we infer a clear geometric description of the low-loss region. We find that it is a low-dimensional manifold whose dimension can be readily obtained from the fluctuations. Furthermore, this dimension is controlled by the number of data points that reside near the classification decision boundary. Importantly, we find that a quadratic approximation of the loss near the minimum is fundamentally inadequate due to the exponential nature of the decision boundary and the flatness of the low-loss region. This causes the dynamics to sample regions with higher curvature at higher temperatures, while producing quadratic-like statistics at any given temperature. We explain this behavior by a simplified loss model which is analytically tractable and reproduces the observed fluctuation statistics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. First-Passage Approach to Optimizing Perturbations for Improved Training of Machine Learning Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    The authors show that when unperturbed neural network training reaches a quasi-steady state, the mean time to a target test accuracy under periodic perturbations can be predicted from a single perturbation experiment.

Pith tools