Random scaling and momentum for non-smooth non-convex optimization

Qinzi Zhang, Ashok Cutkosky · arXiv 2405.09742

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

read on arXiv browse 3 citing papers

citation-role summary

method 1

citation-polarity summary

use method 1

representative citing papers

Stochastic Non-Smooth Convex Optimization with Unbounded Gradients

math.OC · 2026-05-15 · unverdicted · novelty 7.0

Clipped AdamW with exponentially weighted accumulation achieves superior global convergence rates for convex stochastic generalized Lipschitz optimization compared to SGD and AdaGrad.

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models

cs.LG · 2026-04-16 · unverdicted · novelty 6.0

StoSignSGD resolves SignSGD divergence on non-smooth objectives via structural stochasticity, matching optimal convex rates and improving non-convex bounds while delivering 1.44-2.14x speedups in FP8 LLM pretraining.

Optimal Projection-Free Adaptive SGD for Matrix Optimization

math.OC · 2026-04-02 · unverdicted · novelty 6.0

Proving stability of Leon's preconditioner enables the first tuning-free Nesterov-accelerated projection-free adaptive SGD variant with improved non-smooth non-convex rates.

citing papers explorer

Showing 3 of 3 citing papers after filters.

Stochastic Non-Smooth Convex Optimization with Unbounded Gradients math.OC · 2026-05-15 · unverdicted · none · ref 12
Clipped AdamW with exponentially weighted accumulation achieves superior global convergence rates for convex stochastic generalized Lipschitz optimization compared to SGD and AdaGrad.
StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models cs.LG · 2026-04-16 · unverdicted · none · ref 48
StoSignSGD resolves SignSGD divergence on non-smooth objectives via structural stochasticity, matching optimal convex rates and improving non-convex bounds while delivering 1.44-2.14x speedups in FP8 LLM pretraining.
Optimal Projection-Free Adaptive SGD for Matrix Optimization math.OC · 2026-04-02 · unverdicted · none · ref 14
Proving stability of Leon's preconditioner enables the first tuning-free Nesterov-accelerated projection-free adaptive SGD variant with improved non-smooth non-convex rates.

Random scaling and momentum for non-smooth non-convex optimization

citation-role summary

citation-polarity summary

fields

years

verdicts

roles

polarities

representative citing papers

citing papers explorer