LE-SAM inverts SAM by fixing the loss budget instead of the parameter-space radius, yielding better generalization across benchmarks.
arXiv preprint arXiv:1905.00313 , title =
9 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
A direct Lyapunov argument establishes near-linear convergence of a more adaptive gradient descent method for convex functions satisfying fourth-order growth, bypassing an earlier intricate ravine-monitoring proof.
Polyak-style step sizes for Schedule-Free SGD and Adam achieve O(1/sqrt(t)) anytime last-iterate rates for convex Lipschitz problems using per-iteration loss and gradient information.
Stein Diffusion Guidance corrects approximate posteriors in diffusion sampling via a Stein variational mechanism and surrogate SOC objective to enable effective guidance beyond high-density regimes.
Parameter-free first-order methods attain optimal oracle complexity O(ε^{-2/(1+3ρ)}) for convex function-constrained optimization under Hölder smoothness by combining modified Polyak steps, Nesterov momentum, and APL level-set methods.
Restart schemes for SGD on KL-satisfying non-smooth weakly convex problems deliver accelerated convergence robust to exponent misspecification, with optimal schedules resembling Polyak steps.
Large empirical study of 56 optimizers on 1092 BBVI tasks finds no single winner but a selection of five suffices for near-best performance.
Proposes Polyak schedulers for SAM with convergence proofs in deterministic and stochastic settings and empirical results showing reduced tuning needs.
Lecture notes on convergence theory for deterministic gradient descent and stochastic gradient methods under standard assumptions.
citing papers explorer
-
Fix the Loss, Not the Radius: Rethinking the Adversarial Perturbation of Sharpness-Aware Minimization
LE-SAM inverts SAM by fixing the loss budget instead of the parameter-space radius, yielding better generalization across benchmarks.
-
A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
A direct Lyapunov argument establishes near-linear convergence of a more adaptive gradient descent method for convex functions satisfying fourth-order growth, bypassing an earlier intricate ravine-monitoring proof.
-
Taking the Road Less Scheduled with Adaptive Polyak Steps
Polyak-style step sizes for Schedule-Free SGD and Adam achieve O(1/sqrt(t)) anytime last-iterate rates for convex Lipschitz problems using per-iteration loss and gradient information.
-
Stein Diffusion Guidance: Training-Free Posterior Correction for Sampling Beyond High-Density Regions
Stein Diffusion Guidance corrects approximate posteriors in diffusion sampling via a Stein variational mechanism and surrogate SOC objective to enable effective guidance beyond high-density regimes.
-
Uniformly Optimal and Parameter-free First-order Methods for Convex and Function-constrained Optimization
Parameter-free first-order methods attain optimal oracle complexity O(ε^{-2/(1+3ρ)}) for convex function-constrained optimization under Hölder smoothness by combining modified Polyak steps, Nesterov momentum, and APL level-set methods.
-
Restart and Adaptive Acceleration in Stochastic Gradient Methods
Restart schemes for SGD on KL-satisfying non-smooth weakly convex problems deliver accelerated convergence robust to exponent misspecification, with optimal schedules resembling Polyak steps.
-
Large-scale empirical tuning and comparison of default optimizers for variational inference
Large empirical study of 56 optimizers on 1092 BBVI tasks finds no single winner but a selection of five suffices for near-best performance.
-
Adaptive Sharpness-Aware Minimization with a Polyak-type Step size: A Theory-Grounded Scheduler
Proposes Polyak schedulers for SAM with convergence proofs in deterministic and stochastic settings and empirical results showing reduced tuning needs.
-
Introduction to stochastic gradient methods
Lecture notes on convergence theory for deterministic gradient descent and stochastic gradient methods under standard assumptions.