Pith. sign in

REVIEW 13 cited by

New insights and perspectives on the natural gradient method

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1412.1193 v11 pith:LXVBIMHL submitted 2014-12-03 cs.LG stat.ML

New insights and perspectives on the natural gradient method

classification cs.LG stat.ML
keywords gradientnaturalhessianmatrixmethoddescentfisherinformation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Natural gradient descent is an optimization method traditionally motivated from the perspective of information geometry, and works well for many applications as an alternative to stochastic gradient descent. In this paper we critically analyze this method and its properties, and show how it can be viewed as a type of 2nd-order optimization method, with the Fisher information matrix acting as a substitute for the Hessian. In many important cases, the Fisher information matrix is shown to be equivalent to the Generalized Gauss-Newton matrix, which both approximates the Hessian, but also has certain properties that favor its use over the Hessian. This perspective turns out to have significant implications for the design of a practical and robust natural gradient optimizer, as it motivates the use of techniques like trust regions and Tikhonov regularization. Additionally, we make a series of contributions to the understanding of natural gradient and 2nd-order methods, including: a thorough analysis of the convergence speed of stochastic natural gradient descent (and more general stochastic 2nd-order methods) as applied to convex quadratics, a critical examination of the oft-used "empirical" approximation of the Fisher matrix, and an analysis of the (approximate) parameterization invariance property possessed by natural gradient methods (which we show also holds for certain other curvature, but notably not the Hessian).

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AMIGO: a Data-Driven Calibration of the JWST Interferometer

    astro-ph.IM 2025-10 unverdicted novelty 7.0

    AMIGO is an end-to-end differentiable forward model of JWST AMI that corrects detector systematics to recover high-precision astrometry and detect close high-contrast companions.

  2. How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection

    eess.AS 2026-07 conditional novelty 6.0

    Meta-learning training concentrates loss-relevant LoRA updates in query/key projections and spreads them in output projections, relative to standard empirical-risk training.

  3. Quantifying the Agreement Between Data-Influence and Data-Similarity to Understand LLM Behavior

    cs.LG 2026-06 unverdicted novelty 6.0

    Data-similarity and data-influence produce significantly overlapping rankings of training documents for LLM outputs, with asymmetry allowing a favorable cost-accuracy trade-off.

  4. Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models

    cs.CL 2026-06 unverdicted novelty 6.0

    ADAS improves low-NFE performance by 9-10 percentage points on math and code tasks by greedily discounting attention-strong candidates during subset construction in masked diffusion decoding.

  5. Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models

    cs.CL 2026-06 conditional novelty 6.0

    Soft attention-discounted greedy reranking improves low-NFE parallel decoding of masked diffusion LMs by ~9–10 points when plugged into Top-k, Fast-dLLM, and EB-Sampler.

  6. Approximating Hartree-Fock theory via an efficiently local reformulation

    physics.chem-ph 2026-06 unverdicted novelty 6.0

    A reorganized Hartree-Fock framework imposes tunable orbital locality by pairing local degrees of freedom with local solution conditions, maintaining efficient SCF optimization and competitive reaction-energy accuracy.

  7. Rotation-Preserving Supervised Fine-Tuning

    cs.LG 2026-05 unverdicted novelty 6.0

    RPSFT improves the in-domain versus out-of-domain performance trade-off during LLM supervised fine-tuning by penalizing rotations in pretrained singular subspaces as a proxy for loss-sensitive directions.

  8. Loss-aware state space geometry for quantum variational algorithms

    quant-ph 2026-04 unverdicted novelty 6.0

    Loss-aware natural gradient variants are introduced by embedding the loss hypersurface in a statistical manifold or using quantum state overlaps, yielding conformal updates that adjust effective step size.

  9. Low Rank Based Subspace Inference for the Laplace Approximation of Bayesian Neural Networks

    cs.LG 2025-02 unverdicted novelty 6.0

    Derives optimal low-rank subspace for Laplace approx in BNNs, provides scalable outperforming version, and new comparison metric.

  10. Natural gradient descent with momentum

    cs.LG 2026-04 unverdicted novelty 5.0

    Introduces natural-gradient versions of Heavy-Ball and Nesterov momentum methods for function approximation on differentiable nonlinear manifolds.

  11. Learnability Window in Gated Recurrent Neural Networks

    cs.LG 2025-12 conditional novelty 5.0

    The learnability window of a gated RNN grows with dataset size at a rate fixed by the decay of an effective learning-rate envelope and by the tail index of gradient noise.

  12. General Uncertainty Estimation with Delta Variances

    cs.LG 2025-02 unverdicted novelty 5.0

    Delta Variances provide computationally efficient epistemic uncertainty quantification for neural networks by recovering popular techniques as special cases and demonstrating competitive performance in a weather simul...

  13. Hessian based analysis of SGD for Deep Nets: Dynamics and Generalization

    cs.LG 2019-07 unverdicted novelty 4.0

    Provides Hessian-based theoretical characterizations of SGD dynamics and a scale-invariant generalization bound for deep nets, backed by experiments on synthetic data, MNIST, and CIFAR-10.