Pith. sign in

REVIEW 3 cited by

Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1807.07540 v5 pith:QG47Q6I5 submitted 2018-07-19 stat.ML cs.LG

classification stat.MLcs.LG
keywords bayesiannetworkneuralnormalizeroptimizationadamfilteringgradient
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We formulate the problem of neural network optimization as Bayesian filtering, where the observations are the backpropagated gradients. While neural network optimization has previously been studied using natural gradient methods which are closely related to Bayesian inference, they were unable to recover standard optimizers such as Adam and RMSprop with a root-mean-square gradient normalizer, instead getting a mean-square normalizer. To recover the root-mean-square normalizer, we find it necessary to account for the temporal dynamics of all the other parameters as they are geing optimized. The resulting optimizer, AdaBayes, adaptively transitions between SGD-like and Adam-like behaviour, automatically recovers AdamW, a state of the art variant of Adam with decoupled weight decay, and has generalisation performance competitive with SGD.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Importance of Encoder Choice:A Tabular-Image Study

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Tabular encoder choice reorders multimodal rankings, can erase apparent fusion gains, and requires non-vanilla extraction for in-context learning models to avoid train-test representation shift.

  2. NeuralChaos: Optimal Adapted Approximation of Square Integrable Predictable Processes

    math.PR 2026-07 conditional novelty 5.0 of 10

    A finite-sampling neural architecture is dense in the Hilbert space of square-integrable predictable processes and attains best-N-term chaoslet rates for compressible or Malliavin-regular processes.

  3. Incomplete Data Multi-Source Static Computed Tomography Reconstruction with Diffusion Priors and Implicit Neural Representation

    physics.med-ph 2025-01 conditional novelty 4.0 of 10

    A diffusion-prior algorithm with affine projection and implicit neural representation improves sparse-view volume reconstruction for a multi-source static CT system.

Pith tools