Pith. sign in

REVIEW 1 cited by

High Probability Convergence of Clipped-SGD Under Heavy-tailed Noise

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.05437 v2 pith:B3AWM4PV submitted 2023-02-10 math.OC

classification math.OC
keywords convergenceemphprobabilityhighnoisestochasticboundsgradient
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

While the convergence behaviors of stochastic gradient methods are well understood \emph{in expectation}, there still exist many gaps in the understanding of their convergence with \emph{high probability}, where the convergence rate has a logarithmic dependency on the desired success probability parameter. In the \emph{heavy-tailed noise} setting, where the stochastic gradient noise only has bounded $p$-th moments for some $p\in(1,2]$, existing works could only show bounds \emph{in expectation} for a variant of stochastic gradient descent (SGD) with clipped gradients, or high probability bounds in special cases (such as $p=2$) or with extra assumptions (such as the stochastic gradients having bounded non-central moments). In this work, using a novel analysis framework, we present new and time-optimal (up to logarithmic factors) \emph{high probability} convergence bounds for SGD with clipping under heavy-tailed noise for both convex and non-convex smooth objectives using only minimal assumptions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. High Probability Convergence of Distributed Clipped Stochastic Gradient Descent with Heavy-tailed Noise

    math.OC 2025-06 conditional novelty 4.0 of 10

    Distributed clipped stochastic gradient descent over time-varying directed graphs is proved to converge with high probability under heavy-tailed gradient noise.

Pith tools