REVIEW 13 cited by
New insights and perspectives on the natural gradient method
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
New insights and perspectives on the natural gradient method
read the original abstract
Natural gradient descent is an optimization method traditionally motivated from the perspective of information geometry, and works well for many applications as an alternative to stochastic gradient descent. In this paper we critically analyze this method and its properties, and show how it can be viewed as a type of 2nd-order optimization method, with the Fisher information matrix acting as a substitute for the Hessian. In many important cases, the Fisher information matrix is shown to be equivalent to the Generalized Gauss-Newton matrix, which both approximates the Hessian, but also has certain properties that favor its use over the Hessian. This perspective turns out to have significant implications for the design of a practical and robust natural gradient optimizer, as it motivates the use of techniques like trust regions and Tikhonov regularization. Additionally, we make a series of contributions to the understanding of natural gradient and 2nd-order methods, including: a thorough analysis of the convergence speed of stochastic natural gradient descent (and more general stochastic 2nd-order methods) as applied to convex quadratics, a critical examination of the oft-used "empirical" approximation of the Fisher matrix, and an analysis of the (approximate) parameterization invariance property possessed by natural gradient methods (which we show also holds for certain other curvature, but notably not the Hessian).
Forward citations
Cited by 13 Pith papers
-
AMIGO: a Data-Driven Calibration of the JWST Interferometer
AMIGO is an end-to-end differentiable forward model of JWST AMI that corrects detector systematics to recover high-precision astrometry and detect close high-contrast companions.
-
How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection
Meta-learning training concentrates loss-relevant LoRA updates in query/key projections and spreads them in output projections, relative to standard empirical-risk training.
-
Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior
Data-similarity and data-influence produce significantly overlapping rankings of training documents for LLM outputs, with asymmetry allowing a favorable cost-accuracy trade-off.
-
Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
ADAS improves low-NFE performance by 9-10 percentage points on math and code tasks by greedily discounting attention-strong candidates during subset construction in masked diffusion decoding.
-
Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
Soft attention-discounted greedy reranking improves low-NFE parallel decoding of masked diffusion LMs by ~9–10 points when plugged into Top-k, Fast-dLLM, and EB-Sampler.
-
Approximating Hartree-Fock theory via an efficiently local reformulation
A reorganized Hartree-Fock framework imposes tunable orbital locality by pairing local degrees of freedom with local solution conditions, maintaining efficient SCF optimization and competitive reaction-energy accuracy.
-
Rotation-Preserving Supervised Fine-Tuning
RPSFT improves the in-domain versus out-of-domain performance trade-off during LLM supervised fine-tuning by penalizing rotations in pretrained singular subspaces as a proxy for loss-sensitive directions.
-
Loss-aware state space geometry for quantum variational algorithms
Loss-aware natural gradient variants are introduced by embedding the loss hypersurface in a statistical manifold or using quantum state overlaps, yielding conformal updates that adjust effective step size.
-
Low Rank Based Subspace Inference for the Laplace Approximation of Bayesian Neural Networks
Derives optimal low-rank subspace for Laplace approx in BNNs, provides scalable outperforming version, and new comparison metric.
-
Natural gradient descent with momentum
Introduces natural-gradient versions of Heavy-Ball and Nesterov momentum methods for function approximation on differentiable nonlinear manifolds.
-
Learnability Window in Gated Recurrent Neural Networks
The learnability window of a gated RNN grows with dataset size at a rate fixed by the decay of an effective learning-rate envelope and by the tail index of gradient noise.
-
General Uncertainty Estimation with Delta Variances
Delta Variances provide computationally efficient epistemic uncertainty quantification for neural networks by recovering popular techniques as special cases and demonstrating competitive performance in a weather simul...
-
Hessian based analysis of SGD for Deep Nets: Dynamics and Generalization
Provides Hessian-based theoretical characterizations of SGD dynamics and a scale-invariant generalization bound for deep nets, backed by experiments on synthetic data, MNIST, and CIFAR-10.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.