Pith. sign in

REVIEW 1 cited by

Accelerating Neural Network Training Along Sharp and Flat Directions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.11972 v1 pith:GYUKM7ZY submitted 2025-05-17 cs.LG stat.ML

classification cs.LGstat.ML
keywords subspacebulk-sgddirectionsdominanthessianalongflatnetwork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work has highlighted a surprising alignment between gradients and the top eigenspace of the Hessian -- termed the Dominant subspace -- during neural network training. Concurrently, there has been growing interest in the distinct roles of sharp and flat directions in the Hessian spectrum. In this work, we study Bulk-SGD, a variant of SGD that restricts updates to the orthogonal complement of the Dominant subspace. Through ablation studies, we characterize the stability properties of Bulk-SGD and identify critical hyperparameters that govern its behavior. We show that updates along the Bulk subspace, corresponding to flatter directions in the loss landscape, can accelerate convergence but may compromise stability. To balance these effects, we introduce interpolated gradient methods that unify SGD, Dom-SGD, and Bulk-SGD. Finally, we empirically connect this subspace decomposition to the Generalized Gauss-Newton and Functional Hessian terms, showing that curvature energy is largely concentrated in the Dominant subspace. Our findings suggest a principled approach to designing curvature-aware optimizers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mini-batch Noise Lowers Sharpness via Dominant-Subspace Fluctuations

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Mini-batch noise lowers top-k Hessian sharpness through fluctuations along the dominant (top-eigenvector) subspace, and adding the derived covariance correction to GD reproduces SGD's sharpness dynamics.

Pith tools