Pith. sign in

REVIEW 1 cited by

Randomised Splitting Methods and Stochastic Gradient Descent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.04274 v1 pith:VUM5VXBW submitted 2025-04-05 math.OC cs.NAmath.NAstat.ML

classification math.OCcs.NAmath.NAstat.ML
keywords gradientstochasticbiasstrategydescentmethodsminibatchingreduced
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We explore an explicit link between stochastic gradient descent using common batching strategies and splitting methods for ordinary differential equations. From this perspective, we introduce a new minibatching strategy (called Symmetric Minibatching Strategy) for stochastic gradient optimisation which shows greatly reduced stochastic gradient bias (from $\mathcal{O}(h^2)$ to $\mathcal{O}(h^4)$ in the optimiser stepsize $h$), when combined with momentum-based optimisers. We justify why momentum is needed to obtain the improved performance using the theory of backward analysis for splitting integrators and provide a detailed analytic computation of the stochastic gradient bias on a simple example. Further, we provide improved convergence guarantees for this new minibatching strategy using Lyapunov techniques that show reduced stochastic gradient bias for a fixed stepsize (or learning rate) over the class of strongly-convex and smooth objective functions. Via the same techniques we also improve the known results for the Random Reshuffling strategy for stochastic gradient descent methods with momentum. We argue that this also leads to a faster convergence rate when considering a decreasing stepsize schedule. Both the reduced bias and efficacy of decreasing stepsizes are demonstrated numerically on several motivating examples.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adaptive Momentum and Nonlinear Damping for Neural Network Training

    cs.LG 2026-01 conditional novelty 6.0 of 10

    Per-parameter kinetic-energy friction and cubic damping in momentum optimizers close much of the Adam–mSGD gap on transformer training, with deterministic exponential-convergence guarantees for strongly convex losses.

Pith tools