Pith. sign in

REVIEW 7 cited by

On the Convergence of SGD with Biased Gradients

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.00051 v2 pith:VGO5SVEV submitted 2020-07-31 cs.LG math.OCstat.ML

classification cs.LGmath.OCstat.ML
keywords biasedgradientconvergencecompressiononlyratesupdatesaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We analyze the complexity of biased stochastic gradient methods (SGD), where individual updates are corrupted by deterministic, i.e. biased error terms. We derive convergence results for smooth (non-convex) functions and give improved rates under the Polyak-Lojasiewicz condition. We quantify how the magnitude of the bias impacts the attainable accuracy and the convergence rates (sometimes leading to divergence). Our framework covers many applications where either only biased gradient updates are available, or preferred, over unbiased ones for performance reasons. For instance, in the domain of distributed learning, biased gradient compression techniques such as top-k compression have been proposed as a tool to alleviate the communication bottleneck and in derivative-free optimization, only biased gradient estimators can be queried. We discuss a few guiding examples that show the broad applicability of our analysis.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 16 citations worldwide. Full citation record

  1. A Gradient Flow Perspective on Minimum MMD Estimation

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Preconditioned gradient descent for min-MMD estimation converges asymptotically to a global minimizer under gradient-dominance and projection-residual conditions, and empirically beats standard GD.

  2. CaliMatch: Adaptive Calibration for Improving Safe Semi-supervised Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    CaliMatch adaptively calibrates classifier and OOD detector confidences in safe semi-supervised learning, improving classification accuracy on CIFAR-10/100, SVHN, TinyImageNet and ImageNet.

  3. Optimization over Sparse Support-Preserving Sets: Two-Step Projection with Global Optimality Guarantees

    math.OC 2025-06 conditional novelty 6.0 of 10

    An iterative hard-thresholding variant with a two-step projection offers global objective-value guarantees for sparse optimization with support-preserving convex constraints, including the first zeroth-order hard-thre...

  4. Discrete State Diffusion Models: A Sample Complexity Perspective

    cs.LG 2025-10 reject novelty 5.0 of 10

    Claims the first Õ(ε⁻²) sample-complexity bound for discrete-state diffusion, but the zero-approximation-error, optimization-error, and hardness lemmas carrying the proof are internally broken.

  5. Neighbor-Sampling Based Momentum Stochastic Methods for Training Graph Neural Networks

    math.OC 2025-08 unverdicted novelty 5.0 of 10

    The paper creates Adam-style optimizers that combine neighbor sampling and control variates for graph neural networks, with optimal convergence rates and better node-classification performance than control-variate SGD.

  6. Stacey: Promoting Stochastic Steepest Descent via Accelerated $\ell_p$-Smooth Nonconvex Optimization

    cs.LG 2025-06 reject novelty 5.0 of 10

    STACEY is a new ℓ_p steepest descent optimizer with primal-dual interpolation; its convergence theory covers only the unaccelerated base algorithm, and its empirical gains rely on grid-searched hyperparameters.

  7. Stochastic Optimization and Data Science

    math.OC 2026-05 unverdicted novelty 2.0 of 10

    The paper motivates stochastic optimization problems from statistical perspectives and describes offline and online approaches to solve expectation minimization problems.

Pith tools