From gradient clipping to normalization for heavy tailed sgd

Florian H¨ ubler, Ilyas Fatkhullin, Niao He · 2024 · arXiv 2410.13849

7 Pith papers cite this work. Polarity classification is still indexing.

7 Pith papers citing it

read on arXiv browse 7 citing papers

citation-role summary

background 1

citation-polarity summary

background 1

representative citing papers

Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters

cs.LG · 2026-05-12 · unverdicted · novelty 7.0

Spectral clipping of leading singular values in gradient matrices stabilizes SGD for non-convex problems with heavy-tailed noise and achieves the optimal convergence rate O(K^{(2-2α)/(3α-2)}).

Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition

math.OC · 2026-05-07 · unverdicted · novelty 7.0

Muon with Nesterov momentum and inexact polar decomposition achieves optimal convergence rates of O(ε^(-(3α-2)/(α-1))) under heavy-tailed noise for ε-stationary points in non-convex settings.

Decentralized Nonconvex Optimization under Heavy-Tailed Noise: Normalization and Optimal Convergence

math.OC · 2025-05-06 · conditional · novelty 7.0

GT-NSGDm achieves the optimal non-asymptotic convergence rate O(1/T^{(p-1)/(3p-2)}) for decentralized nonconvex stochastic optimization under zero-mean heavy-tailed noise with p-th moment.

LionMuon: Alternating Spectral and Sign Descent for Efficient Training

cs.LG · 2026-05-19 · unverdicted · novelty 6.0 · 2 refs

LionMuon alternates Lion and Muon steps with shared dual-EMA buffer to Pareto-dominate existing optimizers in loss and compute on models up to 720M parameters.

Stochastic Zeroth-Order Optimization Under Heavy-Tailed Noise

math.OC · 2026-05-17 · unverdicted · novelty 6.0

RSC-ZO achieves high-probability ε-stationary points for stochastic ZO optimization under weak-L_p heavy-tailed noise with Õ(d^{p/2(p-1)} ε^{-(3p-2)/(p-1)}) function queries.

Can Entry-Wise Clipping Give Spectral Control of Stochastic Gradients?

cs.LG · 2026-05-26 · unverdicted · novelty 5.0

Entry-wise clipping achieves spectral control of gradients via localization under heavy-tailed contamination, with O(ε^{-4}) convergence and empirical savings on NanoGPT pretraining.

Zeroth-Order Nonconvex Nonsmooth Optimization with Heavy-Tailed Noise

cs.LG · 2026-05-23 · unverdicted · novelty 5.0

Proposes a clipped two-point zeroth-order algorithm achieving O(d^{p/2(p-1)} δ^{-1} ε^{-(2p-1)/p-1}) complexity for (δ, ε)-Goldstein stationary points in nonconvex nonsmooth problems with heavy-tailed noise.

citing papers explorer

Showing 6 of 6 citing papers after filters.

Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters cs.LG · 2026-05-12 · unverdicted · none · ref 28
Spectral clipping of leading singular values in gradient matrices stabilizes SGD for non-convex problems with heavy-tailed noise and achieves the optimal convergence rate O(K^{(2-2α)/(3α-2)}).
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition math.OC · 2026-05-07 · unverdicted · none · ref 19
Muon with Nesterov momentum and inexact polar decomposition achieves optimal convergence rates of O(ε^(-(3α-2)/(α-1))) under heavy-tailed noise for ε-stationary points in non-convex settings.
LionMuon: Alternating Spectral and Sign Descent for Efficient Training cs.LG · 2026-05-19 · unverdicted · none · ref 17 · 2 links
LionMuon alternates Lion and Muon steps with shared dual-EMA buffer to Pareto-dominate existing optimizers in loss and compute on models up to 720M parameters.
Stochastic Zeroth-Order Optimization Under Heavy-Tailed Noise math.OC · 2026-05-17 · unverdicted · none · ref 11
RSC-ZO achieves high-probability ε-stationary points for stochastic ZO optimization under weak-L_p heavy-tailed noise with Õ(d^{p/2(p-1)} ε^{-(3p-2)/(p-1)}) function queries.
Can Entry-Wise Clipping Give Spectral Control of Stochastic Gradients? cs.LG · 2026-05-26 · unverdicted · none · ref 37
Entry-wise clipping achieves spectral control of gradients via localization under heavy-tailed contamination, with O(ε^{-4}) convergence and empirical savings on NanoGPT pretraining.
Zeroth-Order Nonconvex Nonsmooth Optimization with Heavy-Tailed Noise cs.LG · 2026-05-23 · unverdicted · none · ref 20
Proposes a clipped two-point zeroth-order algorithm achieving O(d^{p/2(p-1)} δ^{-1} ε^{-(2p-1)/p-1}) complexity for (δ, ε)-Goldstein stationary points in nonconvex nonsmooth problems with heavy-tailed noise.

From gradient clipping to normalization for heavy tailed sgd

citation-role summary

citation-polarity summary

fields

years

verdicts

roles

polarities

representative citing papers

citing papers explorer