REVIEW 3 major objections 4 minor 51 references
Heavy-ball dynamics with Hessian-driven damping for non-convex optimization under the {\L}ojasiewicz condition
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper proves explicit, worst-case optimal exponential convergence rates for heavy-ball dynamics with Hessian damping under the order-2 Łojasiewicz condition, with a closed-form rate formula and a new recommended parameter tuning.
desk verdict A genuinely new explicit rate for DIN under L(2), but the main theorem as stated doesn't deliver the claimed optimal tuning because of a boundary/case-analysis bug; worth a serious referee after a repair. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Lyapunov energy $V(t)=a(F(x(t))-F(\bar{x}))+\frac12\|\beta\nabla F(x(t))+\dot{x}(t)\|^2$, with a free weight $a\ge0$. Along (DIN), its time derivative reduces to $\dot{V}(t)\le -RV(t)$ plus signed gradient and velocity terms, provided $a$ satisfies the algebraic pair (H1): $1-\alpha\beta\le a\le 1+\alpha\beta$ and $a^2-((\alpha-\mu\beta)\beta+1)a+(1-\alpha\beta)\mu\beta^2\ge0$, where $R=(\alpha\beta+1-a)/\beta$. Condition (H1) is exactly what makes the energy decay exponentially, and Lemmas 2 and 3 supply an exhaustive case analysis—based on the roots of the associated quadratic and quartic in $\beta$—showing which $a$ can be chosen for each friction pair and computing the maximal $R$. The same energy, written for the first-order formulation as $U(t)=a(F(x(t))-F(\bar{x}))+\frac12\|(\alpha-1/\beta)x(t)+y(t)/\beta\|^2$, transfers the rates to (g-DIN) and to the perturbed systems. The avoidance result uses the center-manifold theorem to show that initial conditions attracted to strict saddles form a measure-zero set.
What would settle it
For the one-dimensional quadratic $F(x)=\mu x^2/2$, solve $\ddot{x}+(\alpha+\mu\beta)\dot{x}+\mu x=0$ to obtain the exact rate $r(\alpha,\beta)=\alpha+\mu\beta-\sqrt{\max\{0,(\alpha+\mu\beta)^2-4\mu\}}$, and check whether the paper's formula $R(\alpha,\beta)$ ever exceeds it for an admissible pair—if it does, the worst-case optimality claim is false; a direct numerical scan of condition (H1) over the parameter regions of Lemma 2, starting near $\beta=\alpha/\mu$, would expose any missing branch of the case analysis.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a complete Lyapunov-based characterization of exponential convergence for the DIN system under the order-2 Łojasiewicz condition. For any $C^2$ function satisfying $|F(x)-F(\bar{x})|\le \frac{1}{2\mu}\|\nabla F(x)\|^2$ near a critical limit point $\bar{x}$, the trajectory satisfies $F(x(t))-F(\bar{x})\le \bigl(F(x(t_0))-F(\bar{x})+\frac{1}{2C(\alpha,\beta)}\|\beta\nabla F(x(t_0))+\dot{x}(t_0)\|^2\bigr)e^{-R(\alpha,\beta)(t-t_0)}$, with $R$ and $C$ given explicitly in (3)–(5). Maximizing the geometric factor over admissible Lyapunov weights yields $R(\alpha,\beta)=2\alpha$ in the region $\{\alpha\le\sqrt{\mu}\}\cap\{\beta\in[\alpha/\mu,1/\alpha]\}$ and an algebraic expression elsewhere; choosing $\alpha=\sqrt{\mu}-\varepsilon/2$ and $\beta\in[(2\sqrt{\mu}-\varepsilon)/(2\mu),2/(2\sqrt{\mu}-\varepsilon)]$ achieves the rate $e^{-(2\sqrt{\mu}-\varepsilon)t}$ and forces $\|\nabla F(x(t))\|=o(e^{-(\sqrt{\mu}-\varepsilon)t})$. Because the one-dimensional quadratic gives $\sup_{\alpha,\beta>0}r(\alpha,\beta)=2\sqrt{\mu}$, the $e^{-2\sqrt{\mu}t}$ rate is worst-case optimal. The same analysis extends to the perturbed system with an additive error $g$, preserving the exponential rate under the integrability condition $e^{Rs/2}\|g(s)\|\in L^1$, and to the first-order equivalent system (g-DIN); a separate center-manifold argument shows strict saddle points are avoided from almost every initial condition.
Load-bearing premise
The explicit rate formula and the optimal-tuning claim rest on an exhaustive case analysis (Lemmas 2, 3 and A.4) asserting that for every friction pair $(\alpha,\beta)$ a Lyapunov weight $a\ge0$ exists satisfying the algebraic inequalities (H1); if any branch of that case analysis is wrong, the closed-form rate and the optimal tuning collapse, even though the underlying convergence claim could in principle survive.
Editorial extensions
If this is right
- The objective gap of (DIN) decays like $e^{-(2\sqrt{\mu}-\varepsilon)t}$ for every $C^2$ function satisfying the order-2 Łojasiewicz condition, not only for strongly convex objectives, so accelerated linear convergence is available under a much weaker geometric assumption.
- The gradient norm satisfies $\|\nabla F(x(t))\|=o(e^{-(\sqrt{\mu}-\varepsilon)t})$, meaning both the function value and the stationarity measure converge at the accelerated square-root-of-$\mu$ rate.
- The optimal tuning is $\alpha=\sqrt{\mu}$, $\beta=1/\sqrt{\mu}$—less isotropic friction and more Hessian-driven damping than the earlier $\alpha=2\sqrt{\mu}$ choice—and numerical experiments on quadratics and the Rosenbrock function support this recommendation.
- The $e^{-2\sqrt{\mu}t}$ rate is worst-case optimal for the whole dynamics family, since the quadratic example $F(x)=\mu\|x\|^2/2$ cannot be improved upon.
- With additive perturbation $g$, the exponential rate is preserved whenever $e^{R(\alpha,\beta)s/2}\|g(s)\|\in L^1$; under the Łojasiewicz condition of order $q\in(1,2)$, the rate degrades to $t^{-q/(2-q)}$, and almost every initial condition avoids strict saddle points.
Reading between the lines
- If the Lyapunov argument survives discretization with matched rates, inertial Newton algorithms with Hessian damping should inherit the $e^{-2\sqrt{\mu}k}$ linear rate under the Polyak–Łojasiewicz condition in finite-dimensional smooth optimization; the paper lists discretization as a future direction, but the continuous analysis here is the natural template.
- The worst-case optimality at the quadratic suggests that, within this second-order flow family, the order-2 Łojasiewicz condition alone buys the same worst-case constant as strong convexity; obtaining faster worst-case rates would require exploiting extra structure such as error bounds or tame geometry.
- The recommended tuning $\alpha\approx\sqrt{\mu}$, $\beta\approx 1/\sqrt{\mu}$ has a robustness reading: curvature information allows one to reduce velocity damping, which should mitigate oscillation without sacrificing worst-case speed; this is testable on high-dimensional overparameterized losses where Hessian-vector products are inexpensive.
- The perturbed-system estimates give a concrete noise budget—$e^{R(\alpha,\beta)s/2}\|g(s)\|$ integrable—that stochastic or inexact variants of DIN must respect, which is a testable design criterion for stochastic inertial Newton methods.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the heavy-ball dynamics with Hessian-driven damping (DIN) for C^2 non-convex objectives satisfying a Łojasiewicz condition of order 2. The main results are explicit exponential convergence rates for the objective gap and the gradient norm as functions of the friction parameters (α,β), an optimal worst-case tuning achieving rate O(e^{-2√μ t}), robustness estimates under additive gradient perturbations for both the second-order system and its first-order equivalent (g-DIN), and an almost-sure avoidance result for strict saddle points. The proofs rely on a Lyapunov energy with a parameter a whose feasibility conditions are analyzed through a case study in Lemma 2 and Lemma 3. The optimality claim is justified by an independent quadratic lower bound in Remark 3. The paper is self-contained in its main Lyapunov derivation, with the central question being the correctness of the algebraic case analysis that converts the Lyapunov inequalities into the closed-form rate R(α,β).
Significance. If the rate formula and its optimal tuning are correct, the paper makes a substantial contribution: it establishes worst-case optimal accelerated linear rates for the DIN system under Łojasiewicz condition L(2), improving over previous results for strongly convex objectives and providing explicit parameter-dependent rates in a non-convex setting. The perturbation estimates and the saddle-avoidance theorem extend the scope of earlier work, and the quadratic lower bound in Remark 3 is a sound independent check of the optimal rate. The Lyapunov approach is coherent and mostly self-contained. However, the significance is conditional on fixing the boundary/overlap misclassifications in Lemma 3, since the closed-form rate formula and the optimal-tuning claim rest on that case analysis.
major comments (3)
- [Lemma 3, Eq. (30)-(31) and its proof (Section B)] The proof partitions the admissible parameter set into F− (inf a = 1−αβ) and F+ (inf a = y+), but the condition defining F+, namely y+ ≥ max{0,1−αβ}, does not imply that the first feasible interval [max{0,1−αβ}, y−] is empty. Whenever max{0,1−αβ} ≤ y−, the infimum is 1−αβ even if y+ ≥ 1−αβ. At the proposed optimal tuning (α,β) = (√μ,1/√μ) one has 1−αβ = 0, y− = 0, and y+ = 1; hence a = 0 is feasible and gives R = 2√μ, yet the proof assigns this point to F+ and reports R = √μ. The classification therefore underestimates the rate on an overlap region that includes the claimed optimal parameters.
- [Theorem 1, Eq. (3) and Corollary 1, Eq. (7)] Theorem 1 restricts the first branch to α < √μ and β ∈ (α/μ, 1/α) with open endpoints, while Corollary 1 claims the rate 2√μ−ε for α = √μ−ε/2 and β in the closed interval [α/μ, 1/α]. Under the displayed formula (3), choosing β = α/μ places the point in the second branch and gives a slower rate (for α > √μ/2, the second branch evaluates to μ/α rather than 2α). Thus the closed-interval rate stated in Corollary 1 is not a consequence of Theorem 1 as written, and Remark 4's statement that the optimal rate is obtained at α = √μ, β = 1/√μ is not supported by the theorem's formula. These boundary cases are exactly where the defective F+/F− partition in Lemma 3 misleads the statement.
- [Lemma 2, Eqs. (25)-(26)] The definitions of β1 and β2 in the statement of Lemma 2 involve the Lyapunov parameter a, since the expressions contain (a − 2√(2μ))^2. This makes the conditions self-referential: a must be known before the admissible range for β can be checked, which is circular. The correct expressions, given in Lemma A.4 in terms of α, should replace them. This is a substantive error in a lemma that is cited as guaranteeing the existence of the Lyapunov parameter.
minor comments (4)
- [Theorems 3 and 5, condition i)] The statement says 'F satisfies L(2) with μ > 0' in the context of q ∈ (1,2); it should read L(q).
- [Section 1.1.2, displayed (g-DIN) system] The two equations in the displayed (g-DIN) system appear identical; by comparison with Section 3.2, the first equation should contain the term β∇F(x(t)).
- [Appendix C, display (C.11)] The set S is defined using ∇²F(¯z), but for the vector field f of (65) the relevant object is ∇f(¯z); as written it conflicts with Definition 2 and with the proof of Lemma 4.
- [Figure 1 and Corollary 1] If the branch conditions in Theorem 1 are corrected to include the closed interval and the endpoint α = √μ, the plot in Figure 1 and the associated discussion should be updated to reflect the rate behavior on the boundary.
Circularity Check
No circular derivation; one self-referential definition in Lemma 2 noted, plus correctness concerns outside circularity.
-
self definitional
[Lemma 2, equations (25)–(26), Section 4 (p. 12–13)]
"β1 = 1/(2µ)(2√(2µ) − α − sqrt((a − 2√(2µ))^2 − 4µ)) (25) β2 = 1/(2µ)(2√(2µ) − α + sqrt((a − 2√(2µ))^2 − 4µ)) (26) ... If 0 < α≤ 2(√2 − 1)√µ • If β ∈ [β1, β2], then max{0, 1 − αβ} ≤a ≤ 1 + αβ"
The admissible set of the Lyapunov parameter a in Lemma 2 is declared through intervals whose endpoints β1 and β2 are themselves functions of a. Thus 'a satisfies Lemma 2' is a fixed-point condition of the form a ∈ I(β1(a), β2(a)); no fixed-point or monotonicity argument is supplied. The proof of Lemma 2 in Appendix B actually replaces these with the a-independent β± from Lemma A.4, so the intended argument is non-circular and the central rate formula does have independent content. But as stated, Lemma 2 is self-referential, and Lemma 3's case analysis uses β1, β2 as if they were constants. This is a contingent defect in the written derivation chain, not a forced prediction.
full rationale
The central rate derivation is essentially self-contained. Lemma 1 derives an exponential Lyapunov inequality under the algebraic condition (H1); Lemma A.4 gives an a-independent sign characterization of the quadratic P2(y); Lemma 3 maximizes the geometric exponent over the feasible set; and Remark 3 checks worst-case optimality against an explicit scalar quadratic example where the DIN ODE reduces to a linear second-order equation and the rate r(α,β) is computed directly. No rate constant is fitted to a subset of data and then reported as a prediction: the Lyapunov parameter a is optimized analytically over an explicitly characterized feasible set, and the optimal exponent O(e^{-2√μ t}) is corroborated by an independent example, not by construction. Citations to [2] are used only for existence, equivalence of (DIN)/(g-DIN), and convergence under the Łojasiewicz condition; these are standard and are not the source of the new rate formula. The saddle-avoidance argument relies on Carr's center manifold theorem and on Lemma 5, whose eigenvalue content is stated and used transparently. The one circular-looking object is Lemma 2's definition of β1 and β2 in terms of the very parameter a they are meant to constrain; this is a self-referential statement, but the appendix proof of Lemma 2 bypasses it using the a-independent β± of Lemma A.4, so the main theorem does not reduce to its own inputs. The skeptical concerns about boundary/overlap cases in Lemma 3 (for example at α=√μ, β=1/√μ) are correctness issues in the case analysis, not episodes of prediction-by-fitting or definitional equivalence, and therefore they do not raise the circularity score above 2.
Assumptions & free parameters
assumptions (5)
- domain assumption Łojasiewicz condition of order q ∈ (1,2] as in Assumption 2
- domain assumption Existence, global well-posedness and convergence of DIN trajectories from Proposition 1 (cited from [2])
- domain assumption Equivalence of DIN and g-DIN (cited from [2, Theorem 6.1])
- standard math Carr's center manifold theorem (Theorem 7, [42])
- standard math Grönwall and Bihari integral inequalities (Lemmas A.1 and A.3)
Cite this review
Pith. "Pith review of Heavy-ball dynamics with Hessian-driven damping for non-convex optimization under the {\L}ojasiewicz condition." pith.science (2026). https://pith.science/paper/AWA33KAL
@misc{pith2026250611705,
author = {Pith},
title = {Pith review of: Heavy-ball dynamics with Hessian-driven damping for non-convex optimization under the \Lojasiewicz condition},
year = {2026},
howpublished = {\url{https://pith.science/paper/AWA33KAL}},
note = {Machine review of arXiv:2506.11705}
}
read the original abstract
In this paper, we examine the convergence properties of heavy-ball dynamics with Hessian-driven damping in smooth non-convex optimization problems satisfying a {\L}ojasiewicz condition. In this general setting, we provide a series of tight, worst-case optimal convergence rate guarantees as a function of the dynamics' friction coefficients and the {\L}ojasiewicz exponent of the problem's objective function. Importantly, the linear rates that we obtain improve on previous available rates and they suggest a different tuning of the dynamics' damping terms, even in the strongly convex regime. We complement our analysis with a range of stability estimates in the presence of perturbation errors and inexact gradient input, as well as an avoidance result showing that the dynamics under study avoid strict saddle points from almost every initial condition,
Figures
Reference graph
Works this paper leans on
-
[1]
F. Al v arez, On the minimizing property of a second order dissipative system in hilbert spaces, SIAM Journal on Control and Optimization, 38 (2000), pp. 1102–1119
work page 2000
-
[2]
F. Al v arez, H. Attouch, J. Bolte, and P. Redont , A second-order gradient-like dissipative dynamical system with hessian-driven damping.: Application to optimization and mechanics, Journal de mathématiques pures et appliquées, 81 (2002), pp. 747–779
work page 2002
-
[3]
M. Anitescu, Degenerate nonlinear programming with a quadratic growth condition, SIAM Journal on Optimization, 10 (2000), pp. 1116–1135
work page 2000
-
[4]
A. S. Antipin , Minimization of convex functions on convex sets by means of differential equations, Differential equations, 30 (1994), pp. 1365–1375
work page 1994
-
[5]
V. Apidopoulos, N. Gina tt a, and S. Villa,Convergence rates for the heavy-ball continuous dynamics for non-convex optimization, under polyak–Łojasiewicz condition, Journal of Global Optimization, 84 (2022), pp. 563–589
work page 2022
-
[6]
H. Attouch, J. Bolte, P. Redont, and A. Soubeyran , Proximal alternating minimization and projection methods for nonconvex problems: An approach based on the kurdyka-łojasiewicz inequality, Mathematics of Operations Research, 35 (2010), pp. 438–457. NON-CONVEX HEA VY-BALL DYNAMICS WITH HESSIAN-DRIVEN DAMPING 33
work page 2010
-
[7]
H. Attouch, J. Bol te, and B. F. Sv aiter , Convergence of descent methods for semi-algebraic and tame problems: proximal algorithms, forward–backward splitting, and regularized Gauss–Seidel methods, Mathematical Programming, 137 (2013), pp. 91–129
work page 2013
-
[8]
H. Attouch, Z. Chbani, J. F adili, and H. Riahi , First-order optimization algorithms via inertial systems with hessian driven damping, Mathematical Programming, (2020), pp. 1–43
work page 2020
Show all 51 references
-
[9]
A ttouch, J
H. A ttouch, J. F adili, and V. Kungur tsev, On the effect of perturbations in first-order optimization methods with inertia and hessian driven damping, arXiv preprint arXiv:2106.16159, (2021)
2021 arXiv
-
[10]
Attouch, X
H. Attouch, X. Goudou, and P. Redont , The heavy ball with friction method, i. the continuous dynamical system: global exploration of the local minima of a real-valued function by asymptotic analysis of a dissipative dynamical system, Communications in Contemporary Mathematics...
2000
-
[11]
Attouch, P.-E
H. Attouch, P.-E. Maingé, and P. Redont , A second-order differential system with hessian- driven damping; application to non-elastic shock laws, Differential Equations & Applications, 4 (2012), pp. 27–65
2012
-
[12]
Aubin and J.-P
J.-P. Aubin and J.-P. Aubin , Set-valued analysis, Springer, 1999
1999
-
[13]
Aujol and C
J.-F. Aujol and C. Dossal , Convergence rates of the heavy ball method for quasi-strongly convex optimization, SIAM Journal on Optimization, 32 (2022), pp. 1817–1842
2022
-
[14]
, Convergence rates of the heavy-ball method under the Łojasiewicz property, Mathematical Programming, 198 (2023), pp. 195–254
2023
-
[15]
Aujol, C
J.-F. Aujol, C. Dossal, V. H. Hoàng, H. Labarrière, and A. Rondepierre , Fast convergence of inertial dynamics with hessian-driven damping under geometry assumptions, Applied Mathematics & Optimization, 88 (2023), p. 81
2023
-
[16]
Aujol, C
J.-F. Aujol, C. Dossal, H. Labarrière, and A. Rondepierre , Heavy ball momentum for non- strongly convex optimization, arXiv preprint arXiv:2403.06930, (2024)
2024 arXiv
-
[17]
Bégout, J
P. Bégout, J. Bolte, and M. A. Jendoubi , On damped second-order gradient systems, Journal of Differential Equations, 259 (2015), pp. 3115–3143
2015
-
[18]
Bihari , A generalization of a lemma of bellman and its application to uniqueness problems of differential equations, Acta Mathematica Hungarica, 7 (1956), pp
I. Bihari , A generalization of a lemma of bellman and its application to uniqueness problems of differential equations, Acta Mathematica Hungarica, 7 (1956), pp. 81–94
1956
-
[19]
Bolte, A
J. Bolte, A. Daniilidis, and A. Lewis , The Łojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems, SIAM Journal on Optimization, 17 (2006), pp. 1205–1223 (electronic)
2006
-
[20]
Bolte, A
J. Bolte, A. Daniilidis, and A. Lewis , The łojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems, SIAM Journal on Optimization, 17 (2007), pp. 1205–1223
2007
-
[21]
Bolte, A
J. Bolte, A. Daniilidis, A. Lewis, and M. Shiota , Clarke subgradients of stratifiable functions, SIAM Journal on Optimization, 18 (2007), pp. 556–572
2007
-
[22]
Bolte, A
J. Bolte, A. Daniilidis, O. Ley, and L. Mazet , Characterizations of łojasiewicz inequalities: subgradient flows, talweg, convexity, Transactions of the American Mathematical Society, 362 (2010), pp. 3319–3363
2010
-
[23]
Bol te, T
J. Bol te, T. P. Nguyen, J. Peypouquet, and B. W. Suter , From error bounds to the complexity of first-order descent methods for convex functions, Mathematical Programming, 165 (2017), pp. 471–507
2017
-
[24]
Castera, Inertial newton algorithms avoiding strict saddle points, Journal of Optimization Theory and Applications, 199 (2023), pp
C. Castera, Inertial newton algorithms avoiding strict saddle points, Journal of Optimization Theory and Applications, 199 (2023), pp. 881–903
2023
-
[25]
Castera, H
C. Castera, H. Attouch, J. F adili, and P. Ochs , Continuous newton-like methods featuring inertia and variable mass, SIAM Journal on Optimization, 34 (2024), pp. 251–277
2024
-
[26]
Castera, J
C. Castera, J. Bolte, C. Févotte, and E. Pauwels , An inertial newton algorithm for deep learning, Journal of Machine Learning Research, 22 (2021), pp. 1–31
2021
-
[27]
Clarke, Functional analysis, calculus of variations and optimal control, vol
F. Clarke, Functional analysis, calculus of variations and optimal control, vol. 264, Springer Science & Business Media, 2013
2013
-
[28]
Drusvyatskiy and A
D. Drusvyatskiy and A. S. Lewis , Error bounds, quadratic growth, and linear convergence of proximal methods, Mathematics of Operations Research, 43 (2018), pp. 919–948
2018
-
[29]
Garrigos, L
G. Garrigos, L. Rosasco, and S. Villa , Convergence of the forward-backward algorithm: Beyond the worst case with the help of geometry, arXiv preprint arXiv:1703.09477, (2017)
2017 arXiv
-
[30]
Goudou and J
X. Goudou and J. Munier , The gradient and heavy ball with friction dynamical systems: the quasiconvex case, Mathematical Programming, 116 (2009), pp. 173–191. 34 V. APIDOPOULOS, V. MA VROGEORGOU, AND T. G. TSIRONIS
2009
-
[31]
Haraux and M.-A
A. Haraux and M.-A. Jendoubi , Convergence of solutions of second-order gradient-like systems with analytic nonlinearities, journal of differential equations, 144 (1998), pp. 313–320
1998
-
[32]
Karimi, J
H. Karimi, J. Nutini, and M. Schmidt , Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition, in Machine Learning and Knowledge Discovery in Databases, P. Frasconi, N. Landwehr, G. Manco, and J. Vreeken, eds., Cham, 2016, Springer ...
2016
-
[33]
Kurdyka, On gradients of functions definable in o-minimal structures, Annales de l’institut Fourier, 48 (1998), pp
K. Kurdyka, On gradients of functions definable in o-minimal structures, Annales de l’institut Fourier, 48 (1998), pp. 769–783
1998
-
[34]
S. Łojasiewicz, Une propriété topologique des sous-ensembles analytiques réels, in Les Équations aux Dérivées Partielles (Paris, 1962), Éditions du Centre National de la Recherche Scientifique, Paris, 1963, pp. 87–89
1962
-
[35]
Université de Grenoble., 43 (1993), pp
, Sur la géométrie semi- et sous-analytique, Annales de l’Institut Fourier. Université de Grenoble., 43 (1993), pp. 1575–1595
1993
-
[36]
Luo and P
Z.-Q. Luo and P. Tseng , Error bounds and convergence analysis of feasible descent methods: a general approach, Annals of Operations Research, 46 (1993), pp. 157–178
1993
-
[37]
Maulen-Soto, J
R. Maulen-Soto, J. F adili, and P. Ochs , Inertial methods with viscous and hessian driven damping for non-convex optimization, arXiv preprint arXiv:2407.12518, (2024)
2024 arXiv
-
[38]
Necoara, Y
I. Necoara, Y. Nesterov, and F. Glineur , Linear convergence of first order methods for non-strongly convex optimization, Mathematical Programming, 175 (2019), pp. 69–107
2019
-
[39]
Nesterov, Introductory lectures on convex optimization: A basic course, 2013
Y. Nesterov, Introductory lectures on convex optimization: A basic course, 2013
2013
-
[40]
Palis Jr
J. Palis Jr. and W. de Melo , Geometric Theory of Dynamical Systems, Springer-Verlag, 1982
1982
-
[41]
J. P. Penot , Conditioning convex and nonconvex problems, Journal of Optimization Theory and Applications, 90 (1996), pp. 535–554
1996
-
[42]
Perko, Differential equations and dynamical systems, vol
L. Perko, Differential equations and dynamical systems, vol. 7, Springer Science & Business Media, 2013
2013
-
[43]
Polyak and P
B. Polyak and P. Shcherbakov , Lyapunov functions: An optimization theory perspective, IFAC- PapersOnLine, 50 (2017), pp. 7456–7461
2017
-
[44]
B. T. Pol y ak,Gradient methods for the minimisation of functionals, USSRComputationalMathematics and Mathematical Physics, 3 (1963), pp. 864–878
1963
-
[45]
, Some methods of speeding up the convergence of iteration methods, USSR Computational Mathematics and Mathematical Physics, 4 (1964), pp. 1–17
1964
-
[46]
J. W. Siegel , Accelerated first-order methods: Differential equations and lyapunov functions, arXiv preprint arXiv:1903.05671, (2019)
2019 arXiv
-
[47]
Simsek, F
B. Simsek, F. Ged, A. Jacot, F. Sp adaro, C. Hongler, W. Gerstner, and J. Brea , Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances, in International Conference on Machine Learning, PMLR, 2021, pp. 9722–9732
2021
-
[48]
Teschl, Ordinary Differential Equations and Dynamical Systems, vol
G. Teschl, Ordinary Differential Equations and Dynamical Systems, vol. 140, American Mathematical Soc., 2012
2012
-
[49]
W ang and J
Z. W ang and J. Peypouquet , Accelerated gradient methods via inertial systems with hessian-driven damping, arXiv preprint arXiv:2502.16953, (2025)
2025 arXiv
-
[50]
Zhang , New analysis of linear convergence of gradient-type methods via unifying error bound conditions, Mathematical Programming, 180 (2020), pp
H. Zhang , New analysis of linear convergence of gradient-type methods via unifying error bound conditions, Mathematical Programming, 180 (2020), pp. 371–416
2020
-
[51]
Zhang and W
H. Zhang and W. Yin , Gradient methods for convex minimization: better rates under weaker conditions, arXiv preprint arXiv:1303.4645, (2013)
2013 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.