Pith. sign in

REVIEW 3 major objections 4 minor 51 references

Heavy-ball dynamics with Hessian-driven damping for non-convex optimization under the {\L}ojasiewicz condition

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The paper proves explicit, worst-case optimal exponential convergence rates for heavy-ball dynamics with Hessian damping under the order-2 Łojasiewicz condition, with a closed-form rate formula and a new recommended parameter tuning.

desk verdict A genuinely new explicit rate for DIN under L(2), but the main theorem as stated doesn't deliver the claimed optimal tuning because of a boundary/case-analysis bug; worth a serious referee after a repair. read the letter →

arxiv 2506.11705 v1 pith:AWA33KAL submitted 2025-06-13 math.OC

classification math.OC MSC 90C2634D0546N1065K0565B99
keywords Non-convexoptimizationinertialmethodsHessian-drivendampingŁojasiewiczconditionPolyak–ŁojasiewiczconvergencerateLyapunovanalysisstrictsaddleavoidance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper studies the second-order inertial Newton-like system $\ddot{x}(t)+\alpha\dot{x}(t)+\beta\nabla^2F(x(t))\dot{x}(t)+\nabla F(x(t))=0$ for minimizing a smooth, not necessarily convex function $F$ on a Hilbert space. Its central claim is that when $F$ satisfies the Łojasiewicz condition of order 2, the objective gap $F(x(t))-F(\bar{x})$ decays exponentially with an explicit, closed-form rate $R(\alpha,\beta)$ depending on the damping coefficients, and the best rate over all tunings is $e^{-(2\sqrt{\mu}-\varepsilon)t}$. This rate is worst-case optimal, because the quadratic $F(x)=\frac{\mu}{2}\|x\|^2$ cannot converge faster under any member of this dynamics family. If the paper is right, the classical accelerated heavy-ball rate extends beyond strong convexity, improves on earlier DIN rates by a square-root factor, and recommends a different tuning ($\alpha\approx\sqrt{\mu}$, $\beta\approx 1/\sqrt{\mu}$) than previously used. The same Lyapunov machinery also yields stability estimates under perturbation errors and almost-sure avoidance of strict saddle points.

What carries the argument

The central object is the Lyapunov energy $V(t)=a(F(x(t))-F(\bar{x}))+\frac12\|\beta\nabla F(x(t))+\dot{x}(t)\|^2$, with a free weight $a\ge0$. Along (DIN), its time derivative reduces to $\dot{V}(t)\le -RV(t)$ plus signed gradient and velocity terms, provided $a$ satisfies the algebraic pair (H1): $1-\alpha\beta\le a\le 1+\alpha\beta$ and $a^2-((\alpha-\mu\beta)\beta+1)a+(1-\alpha\beta)\mu\beta^2\ge0$, where $R=(\alpha\beta+1-a)/\beta$. Condition (H1) is exactly what makes the energy decay exponentially, and Lemmas 2 and 3 supply an exhaustive case analysis—based on the roots of the associated quadratic and quartic in $\beta$—showing which $a$ can be chosen for each friction pair and computing the maximal $R$. The same energy, written for the first-order formulation as $U(t)=a(F(x(t))-F(\bar{x}))+\frac12\|(\alpha-1/\beta)x(t)+y(t)/\beta\|^2$, transfers the rates to (g-DIN) and to the perturbed systems. The avoidance result uses the center-manifold theorem to show that initial conditions attracted to strict saddles form a measure-zero set.

What would settle it

For the one-dimensional quadratic $F(x)=\mu x^2/2$, solve $\ddot{x}+(\alpha+\mu\beta)\dot{x}+\mu x=0$ to obtain the exact rate $r(\alpha,\beta)=\alpha+\mu\beta-\sqrt{\max\{0,(\alpha+\mu\beta)^2-4\mu\}}$, and check whether the paper's formula $R(\alpha,\beta)$ ever exceeds it for an admissible pair—if it does, the worst-case optimality claim is false; a direct numerical scan of condition (H1) over the parameter regions of Lemma 2, starting near $\beta=\alpha/\mu$, would expose any missing branch of the case analysis.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is a complete Lyapunov-based characterization of exponential convergence for the DIN system under the order-2 Łojasiewicz condition. For any $C^2$ function satisfying $|F(x)-F(\bar{x})|\le \frac{1}{2\mu}\|\nabla F(x)\|^2$ near a critical limit point $\bar{x}$, the trajectory satisfies $F(x(t))-F(\bar{x})\le \bigl(F(x(t_0))-F(\bar{x})+\frac{1}{2C(\alpha,\beta)}\|\beta\nabla F(x(t_0))+\dot{x}(t_0)\|^2\bigr)e^{-R(\alpha,\beta)(t-t_0)}$, with $R$ and $C$ given explicitly in (3)–(5). Maximizing the geometric factor over admissible Lyapunov weights yields $R(\alpha,\beta)=2\alpha$ in the region $\{\alpha\le\sqrt{\mu}\}\cap\{\beta\in[\alpha/\mu,1/\alpha]\}$ and an algebraic expression elsewhere; choosing $\alpha=\sqrt{\mu}-\varepsilon/2$ and $\beta\in[(2\sqrt{\mu}-\varepsilon)/(2\mu),2/(2\sqrt{\mu}-\varepsilon)]$ achieves the rate $e^{-(2\sqrt{\mu}-\varepsilon)t}$ and forces $\|\nabla F(x(t))\|=o(e^{-(\sqrt{\mu}-\varepsilon)t})$. Because the one-dimensional quadratic gives $\sup_{\alpha,\beta>0}r(\alpha,\beta)=2\sqrt{\mu}$, the $e^{-2\sqrt{\mu}t}$ rate is worst-case optimal. The same analysis extends to the perturbed system with an additive error $g$, preserving the exponential rate under the integrability condition $e^{Rs/2}\|g(s)\|\in L^1$, and to the first-order equivalent system (g-DIN); a separate center-manifold argument shows strict saddle points are avoided from almost every initial condition.

Load-bearing premise

The explicit rate formula and the optimal-tuning claim rest on an exhaustive case analysis (Lemmas 2, 3 and A.4) asserting that for every friction pair $(\alpha,\beta)$ a Lyapunov weight $a\ge0$ exists satisfying the algebraic inequalities (H1); if any branch of that case analysis is wrong, the closed-form rate and the optimal tuning collapse, even though the underlying convergence claim could in principle survive.

Editorial extensions

If this is right

  • The objective gap of (DIN) decays like $e^{-(2\sqrt{\mu}-\varepsilon)t}$ for every $C^2$ function satisfying the order-2 Łojasiewicz condition, not only for strongly convex objectives, so accelerated linear convergence is available under a much weaker geometric assumption.
  • The gradient norm satisfies $\|\nabla F(x(t))\|=o(e^{-(\sqrt{\mu}-\varepsilon)t})$, meaning both the function value and the stationarity measure converge at the accelerated square-root-of-$\mu$ rate.
  • The optimal tuning is $\alpha=\sqrt{\mu}$, $\beta=1/\sqrt{\mu}$—less isotropic friction and more Hessian-driven damping than the earlier $\alpha=2\sqrt{\mu}$ choice—and numerical experiments on quadratics and the Rosenbrock function support this recommendation.
  • The $e^{-2\sqrt{\mu}t}$ rate is worst-case optimal for the whole dynamics family, since the quadratic example $F(x)=\mu\|x\|^2/2$ cannot be improved upon.
  • With additive perturbation $g$, the exponential rate is preserved whenever $e^{R(\alpha,\beta)s/2}\|g(s)\|\in L^1$; under the Łojasiewicz condition of order $q\in(1,2)$, the rate degrades to $t^{-q/(2-q)}$, and almost every initial condition avoids strict saddle points.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the Lyapunov argument survives discretization with matched rates, inertial Newton algorithms with Hessian damping should inherit the $e^{-2\sqrt{\mu}k}$ linear rate under the Polyak–Łojasiewicz condition in finite-dimensional smooth optimization; the paper lists discretization as a future direction, but the continuous analysis here is the natural template.
  • The worst-case optimality at the quadratic suggests that, within this second-order flow family, the order-2 Łojasiewicz condition alone buys the same worst-case constant as strong convexity; obtaining faster worst-case rates would require exploiting extra structure such as error bounds or tame geometry.
  • The recommended tuning $\alpha\approx\sqrt{\mu}$, $\beta\approx 1/\sqrt{\mu}$ has a robustness reading: curvature information allows one to reduce velocity damping, which should mitigate oscillation without sacrificing worst-case speed; this is testable on high-dimensional overparameterized losses where Hessian-vector products are inexpensive.
  • The perturbed-system estimates give a concrete noise budget—$e^{R(\alpha,\beta)s/2}\|g(s)\|$ integrable—that stochastic or inexact variants of DIN must respect, which is a testable design criterion for stochastic inertial Newton methods.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies the heavy-ball dynamics with Hessian-driven damping (DIN) for C^2 non-convex objectives satisfying a Łojasiewicz condition of order 2. The main results are explicit exponential convergence rates for the objective gap and the gradient norm as functions of the friction parameters (α,β), an optimal worst-case tuning achieving rate O(e^{-2√μ t}), robustness estimates under additive gradient perturbations for both the second-order system and its first-order equivalent (g-DIN), and an almost-sure avoidance result for strict saddle points. The proofs rely on a Lyapunov energy with a parameter a whose feasibility conditions are analyzed through a case study in Lemma 2 and Lemma 3. The optimality claim is justified by an independent quadratic lower bound in Remark 3. The paper is self-contained in its main Lyapunov derivation, with the central question being the correctness of the algebraic case analysis that converts the Lyapunov inequalities into the closed-form rate R(α,β).

Significance. If the rate formula and its optimal tuning are correct, the paper makes a substantial contribution: it establishes worst-case optimal accelerated linear rates for the DIN system under Łojasiewicz condition L(2), improving over previous results for strongly convex objectives and providing explicit parameter-dependent rates in a non-convex setting. The perturbation estimates and the saddle-avoidance theorem extend the scope of earlier work, and the quadratic lower bound in Remark 3 is a sound independent check of the optimal rate. The Lyapunov approach is coherent and mostly self-contained. However, the significance is conditional on fixing the boundary/overlap misclassifications in Lemma 3, since the closed-form rate formula and the optimal-tuning claim rest on that case analysis.

major comments (3)
  1. [Lemma 3, Eq. (30)-(31) and its proof (Section B)] The proof partitions the admissible parameter set into F− (inf a = 1−αβ) and F+ (inf a = y+), but the condition defining F+, namely y+ ≥ max{0,1−αβ}, does not imply that the first feasible interval [max{0,1−αβ}, y−] is empty. Whenever max{0,1−αβ} ≤ y−, the infimum is 1−αβ even if y+ ≥ 1−αβ. At the proposed optimal tuning (α,β) = (√μ,1/√μ) one has 1−αβ = 0, y− = 0, and y+ = 1; hence a = 0 is feasible and gives R = 2√μ, yet the proof assigns this point to F+ and reports R = √μ. The classification therefore underestimates the rate on an overlap region that includes the claimed optimal parameters.
  2. [Theorem 1, Eq. (3) and Corollary 1, Eq. (7)] Theorem 1 restricts the first branch to α < √μ and β ∈ (α/μ, 1/α) with open endpoints, while Corollary 1 claims the rate 2√μ−ε for α = √μ−ε/2 and β in the closed interval [α/μ, 1/α]. Under the displayed formula (3), choosing β = α/μ places the point in the second branch and gives a slower rate (for α > √μ/2, the second branch evaluates to μ/α rather than 2α). Thus the closed-interval rate stated in Corollary 1 is not a consequence of Theorem 1 as written, and Remark 4's statement that the optimal rate is obtained at α = √μ, β = 1/√μ is not supported by the theorem's formula. These boundary cases are exactly where the defective F+/F− partition in Lemma 3 misleads the statement.
  3. [Lemma 2, Eqs. (25)-(26)] The definitions of β1 and β2 in the statement of Lemma 2 involve the Lyapunov parameter a, since the expressions contain (a − 2√(2μ))^2. This makes the conditions self-referential: a must be known before the admissible range for β can be checked, which is circular. The correct expressions, given in Lemma A.4 in terms of α, should replace them. This is a substantive error in a lemma that is cited as guaranteeing the existence of the Lyapunov parameter.
minor comments (4)
  1. [Theorems 3 and 5, condition i)] The statement says 'F satisfies L(2) with μ > 0' in the context of q ∈ (1,2); it should read L(q).
  2. [Section 1.1.2, displayed (g-DIN) system] The two equations in the displayed (g-DIN) system appear identical; by comparison with Section 3.2, the first equation should contain the term β∇F(x(t)).
  3. [Appendix C, display (C.11)] The set S is defined using ∇²F(¯z), but for the vector field f of (65) the relevant object is ∇f(¯z); as written it conflicts with Definition 2 and with the proof of Lemma 4.
  4. [Figure 1 and Corollary 1] If the branch conditions in Theorem 1 are corrected to include the closed interval and the endpoint α = √μ, the plot in Figure 1 and the associated discussion should be updated to reflect the rate behavior on the boundary.

Circularity Check

1 steps flagged · score 2.0 of 10

No circular derivation; one self-referential definition in Lemma 2 noted, plus correctness concerns outside circularity.

  1. self definitional [Lemma 2, equations (25)–(26), Section 4 (p. 12–13)]
    "β1 = 1/(2µ)(2√(2µ) − α − sqrt((a − 2√(2µ))^2 − 4µ)) (25) β2 = 1/(2µ)(2√(2µ) − α + sqrt((a − 2√(2µ))^2 − 4µ)) (26) ... If 0 < α≤ 2(√2 − 1)√µ • If β ∈ [β1, β2], then max{0, 1 − αβ} ≤a ≤ 1 + αβ"

    The admissible set of the Lyapunov parameter a in Lemma 2 is declared through intervals whose endpoints β1 and β2 are themselves functions of a. Thus 'a satisfies Lemma 2' is a fixed-point condition of the form a ∈ I(β1(a), β2(a)); no fixed-point or monotonicity argument is supplied. The proof of Lemma 2 in Appendix B actually replaces these with the a-independent β± from Lemma A.4, so the intended argument is non-circular and the central rate formula does have independent content. But as stated, Lemma 2 is self-referential, and Lemma 3's case analysis uses β1, β2 as if they were constants. This is a contingent defect in the written derivation chain, not a forced prediction.

full rationale

The central rate derivation is essentially self-contained. Lemma 1 derives an exponential Lyapunov inequality under the algebraic condition (H1); Lemma A.4 gives an a-independent sign characterization of the quadratic P2(y); Lemma 3 maximizes the geometric exponent over the feasible set; and Remark 3 checks worst-case optimality against an explicit scalar quadratic example where the DIN ODE reduces to a linear second-order equation and the rate r(α,β) is computed directly. No rate constant is fitted to a subset of data and then reported as a prediction: the Lyapunov parameter a is optimized analytically over an explicitly characterized feasible set, and the optimal exponent O(e^{-2√μ t}) is corroborated by an independent example, not by construction. Citations to [2] are used only for existence, equivalence of (DIN)/(g-DIN), and convergence under the Łojasiewicz condition; these are standard and are not the source of the new rate formula. The saddle-avoidance argument relies on Carr's center manifold theorem and on Lemma 5, whose eigenvalue content is stated and used transparently. The one circular-looking object is Lemma 2's definition of β1 and β2 in terms of the very parameter a they are meant to constrain; this is a self-referential statement, but the appendix proof of Lemma 2 bypasses it using the a-independent β± of Lemma A.4, so the main theorem does not reduce to its own inputs. The skeptical concerns about boundary/overlap cases in Lemma 3 (for example at α=√μ, β=1/√μ) are correctness issues in the case analysis, not episodes of prediction-by-fitting or definitional equivalence, and therefore they do not raise the circularity score above 2.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claim rests on the standard assumption of a Łojasiewicz-type inequality, on a few cited structural results from [2] (equivalence and convergence), and on center-manifold theory. There are no fitted numerical constants and no newly postulated physical or mathematical entities. The Lyapunov parameter a is an internal proof variable, not a free parameter of the result.

assumptions (5)
  • domain assumption Łojasiewicz condition of order q ∈ (1,2] as in Assumption 2
    Defines the function class for which all rates are proven; the constant µ and exponent q drive the rate formulas in Theorems 1-5.
  • domain assumption Existence, global well-posedness and convergence of DIN trajectories from Proposition 1 (cited from [2])
    Used to ensure the trajectory eventually stays in the neighborhood Ω where the local L(2) or L(q) inequality applies (Remark 2).
  • domain assumption Equivalence of DIN and g-DIN (cited from [2, Theorem 6.1])
    Theorems 4 and 5 transfer DIN results to the first-order coupled system, and Theorem 6 uses the equivalence to map saddle avoidance of g-DIN back to DIN.
  • standard math Carr's center manifold theorem (Theorem 7, [42])
    The foundation of Lemma 4, which asserts that initial conditions attracted to fixed points with a positive eigenvalue form a measure-zero set.
  • standard math Grönwall and Bihari integral inequalities (Lemmas A.1 and A.3)
    Standard tools used throughout the Lyapunov proofs to convert differential inequalities into exponential or sublinear bounds.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Heavy-ball dynamics with Hessian-driven damping for non-convex optimization under the {\L}ojasiewicz condition." pith.science (2026). https://pith.science/paper/AWA33KAL

@misc{pith2026250611705,
  author       = {Pith},
  title        = {Pith review of: Heavy-ball dynamics with Hessian-driven damping for non-convex optimization under the \Lojasiewicz condition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AWA33KAL}},
  note         = {Machine review of arXiv:2506.11705}
}
read the original abstract

In this paper, we examine the convergence properties of heavy-ball dynamics with Hessian-driven damping in smooth non-convex optimization problems satisfying a {\L}ojasiewicz condition. In this general setting, we provide a series of tight, worst-case optimal convergence rate guarantees as a function of the dynamics' friction coefficients and the {\L}ojasiewicz exponent of the problem's objective function. Importantly, the linear rates that we obtain improve on previous available rates and they suggest a different tuning of the dynamics' damping terms, even in the strongly convex regime. We complement our analysis with a range of stability estimates in the presence of perturbation errors and inexact gradient input, as well as an avoidance result showing that the dynamics under study avoid strict saddle points from almost every initial condition,

Figures

Figures reproduced from arXiv: 2506.11705 by the authors.

Figure 1
Figure 1. Illustration of the convergence exponent R(α, β) expressed in (3) as a function of α and β. where R(α, β) =    2α if {α < √ µ} &  β ∈  α µ , 1 α  1 + αβ − Q(α, β) 2β if {α > 0} &  β ∈  0, α µ  ∪  1 α , +∞  (3) and C(α, β) =    1 − αβ if {α < √ µ} &  β ∈  α µ , 1 α  (1 + αβ + Q(α, β)) 2 if {α > 0} &  β ∈  0, α µ  ∪  1 α , +∞  (4) and Q(α, β) is given by Q(α, β) := q ((α − µβ)β + 1)2 … view at source ↗
Figure 2
Figure 2. Performance of the trajectories generated by system (DIN), for minimiz￾ing a quadratic function F(x) = ⟨Ax, x⟩, in terms of objective function values in logarithmic scale (ln (F(x(t)) − F∗)). Each column corresponds to the selection of a different condition number, starting from κ = 100 (left) to κ = 1000 (right). We compare four different choices of friction parameters α > 0 for the system (DIN): α = 2√µ (green), α… view at source ↗
Figure 3
Figure 3. Comparison of (DIN) and the plain heavy ball with friction (HBF) on the Rosenbrock’s function. The first column illustrates the values of the objective function (in logarithmic scale: ln (F(x(t)) − F∗)), while the second one the generated trajectories (x(t), y(t)) in R 2 . In green and magenta the (DIN) system with (α, β) =  2 √µ, 1 2 √µ  and (α, β) = √µ − ε, 1√µ  (respectively). In blue and yellow lines the (HB… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

51 extracted references · 47 canonical work pages

  1. [1]

    Al v arez, On the minimizing property of a second order dissipative system in hilbert spaces, SIAM Journal on Control and Optimization, 38 (2000), pp

    F. Al v arez, On the minimizing property of a second order dissipative system in hilbert spaces, SIAM Journal on Control and Optimization, 38 (2000), pp. 1102–1119

  2. [2]

    Al v arez, H

    F. Al v arez, H. Attouch, J. Bolte, and P. Redont , A second-order gradient-like dissipative dynamical system with hessian-driven damping.: Application to optimization and mechanics, Journal de mathématiques pures et appliquées, 81 (2002), pp. 747–779

  3. [3]

    Anitescu, Degenerate nonlinear programming with a quadratic growth condition, SIAM Journal on Optimization, 10 (2000), pp

    M. Anitescu, Degenerate nonlinear programming with a quadratic growth condition, SIAM Journal on Optimization, 10 (2000), pp. 1116–1135

  4. [4]

    A. S. Antipin , Minimization of convex functions on convex sets by means of differential equations, Differential equations, 30 (1994), pp. 1365–1375

  5. [5]

    Apidopoulos, N

    V. Apidopoulos, N. Gina tt a, and S. Villa,Convergence rates for the heavy-ball continuous dynamics for non-convex optimization, under polyak–Łojasiewicz condition, Journal of Global Optimization, 84 (2022), pp. 563–589

  6. [6]

    Attouch, J

    H. Attouch, J. Bolte, P. Redont, and A. Soubeyran , Proximal alternating minimization and projection methods for nonconvex problems: An approach based on the kurdyka-łojasiewicz inequality, Mathematics of Operations Research, 35 (2010), pp. 438–457. NON-CONVEX HEA VY-BALL DYNAMICS WITH HESSIAN-DRIVEN DAMPING 33

  7. [7]

    Attouch, J

    H. Attouch, J. Bol te, and B. F. Sv aiter , Convergence of descent methods for semi-algebraic and tame problems: proximal algorithms, forward–backward splitting, and regularized Gauss–Seidel methods, Mathematical Programming, 137 (2013), pp. 91–129

  8. [8]

    Attouch, Z

    H. Attouch, Z. Chbani, J. F adili, and H. Riahi , First-order optimization algorithms via inertial systems with hessian driven damping, Mathematical Programming, (2020), pp. 1–43

Show all 51 references
  1. [9]

    A ttouch, J

    H. A ttouch, J. F adili, and V. Kungur tsev, On the effect of perturbations in first-order optimization methods with inertia and hessian driven damping, arXiv preprint arXiv:2106.16159, (2021)

  2. [10]

    Attouch, X

    H. Attouch, X. Goudou, and P. Redont , The heavy ball with friction method, i. the continuous dynamical system: global exploration of the local minima of a real-valued function by asymptotic analysis of a dissipative dynamical system, Communications in Contemporary Mathematics...

  3. [11]

    Attouch, P.-E

    H. Attouch, P.-E. Maingé, and P. Redont , A second-order differential system with hessian- driven damping; application to non-elastic shock laws, Differential Equations & Applications, 4 (2012), pp. 27–65

  4. [12]

    Aubin and J.-P

    J.-P. Aubin and J.-P. Aubin , Set-valued analysis, Springer, 1999

  5. [13]

    Aujol and C

    J.-F. Aujol and C. Dossal , Convergence rates of the heavy ball method for quasi-strongly convex optimization, SIAM Journal on Optimization, 32 (2022), pp. 1817–1842

  6. [14]

    , Convergence rates of the heavy-ball method under the Łojasiewicz property, Mathematical Programming, 198 (2023), pp. 195–254

  7. [15]

    Aujol, C

    J.-F. Aujol, C. Dossal, V. H. Hoàng, H. Labarrière, and A. Rondepierre , Fast convergence of inertial dynamics with hessian-driven damping under geometry assumptions, Applied Mathematics & Optimization, 88 (2023), p. 81

  8. [16]

    Aujol, C

    J.-F. Aujol, C. Dossal, H. Labarrière, and A. Rondepierre , Heavy ball momentum for non- strongly convex optimization, arXiv preprint arXiv:2403.06930, (2024)

  9. [17]

    Bégout, J

    P. Bégout, J. Bolte, and M. A. Jendoubi , On damped second-order gradient systems, Journal of Differential Equations, 259 (2015), pp. 3115–3143

  10. [18]

    Bihari , A generalization of a lemma of bellman and its application to uniqueness problems of differential equations, Acta Mathematica Hungarica, 7 (1956), pp

    I. Bihari , A generalization of a lemma of bellman and its application to uniqueness problems of differential equations, Acta Mathematica Hungarica, 7 (1956), pp. 81–94

  11. [19]

    Bolte, A

    J. Bolte, A. Daniilidis, and A. Lewis , The Łojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems, SIAM Journal on Optimization, 17 (2006), pp. 1205–1223 (electronic)

  12. [20]

    Bolte, A

    J. Bolte, A. Daniilidis, and A. Lewis , The łojasiewicz inequality for nonsmooth subanalytic functions with applications to subgradient dynamical systems, SIAM Journal on Optimization, 17 (2007), pp. 1205–1223

  13. [21]

    Bolte, A

    J. Bolte, A. Daniilidis, A. Lewis, and M. Shiota , Clarke subgradients of stratifiable functions, SIAM Journal on Optimization, 18 (2007), pp. 556–572

  14. [22]

    Bolte, A

    J. Bolte, A. Daniilidis, O. Ley, and L. Mazet , Characterizations of łojasiewicz inequalities: subgradient flows, talweg, convexity, Transactions of the American Mathematical Society, 362 (2010), pp. 3319–3363

  15. [23]

    Bol te, T

    J. Bol te, T. P. Nguyen, J. Peypouquet, and B. W. Suter , From error bounds to the complexity of first-order descent methods for convex functions, Mathematical Programming, 165 (2017), pp. 471–507

  16. [24]

    Castera, Inertial newton algorithms avoiding strict saddle points, Journal of Optimization Theory and Applications, 199 (2023), pp

    C. Castera, Inertial newton algorithms avoiding strict saddle points, Journal of Optimization Theory and Applications, 199 (2023), pp. 881–903

  17. [25]

    Castera, H

    C. Castera, H. Attouch, J. F adili, and P. Ochs , Continuous newton-like methods featuring inertia and variable mass, SIAM Journal on Optimization, 34 (2024), pp. 251–277

  18. [26]

    Castera, J

    C. Castera, J. Bolte, C. Févotte, and E. Pauwels , An inertial newton algorithm for deep learning, Journal of Machine Learning Research, 22 (2021), pp. 1–31

  19. [27]

    Clarke, Functional analysis, calculus of variations and optimal control, vol

    F. Clarke, Functional analysis, calculus of variations and optimal control, vol. 264, Springer Science & Business Media, 2013

  20. [28]

    Drusvyatskiy and A

    D. Drusvyatskiy and A. S. Lewis , Error bounds, quadratic growth, and linear convergence of proximal methods, Mathematics of Operations Research, 43 (2018), pp. 919–948

  21. [29]

    Garrigos, L

    G. Garrigos, L. Rosasco, and S. Villa , Convergence of the forward-backward algorithm: Beyond the worst case with the help of geometry, arXiv preprint arXiv:1703.09477, (2017)

  22. [30]

    Goudou and J

    X. Goudou and J. Munier , The gradient and heavy ball with friction dynamical systems: the quasiconvex case, Mathematical Programming, 116 (2009), pp. 173–191. 34 V. APIDOPOULOS, V. MA VROGEORGOU, AND T. G. TSIRONIS

  23. [31]

    Haraux and M.-A

    A. Haraux and M.-A. Jendoubi , Convergence of solutions of second-order gradient-like systems with analytic nonlinearities, journal of differential equations, 144 (1998), pp. 313–320

  24. [32]

    Karimi, J

    H. Karimi, J. Nutini, and M. Schmidt , Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition, in Machine Learning and Knowledge Discovery in Databases, P. Frasconi, N. Landwehr, G. Manco, and J. Vreeken, eds., Cham, 2016, Springer ...

  25. [33]

    Kurdyka, On gradients of functions definable in o-minimal structures, Annales de l’institut Fourier, 48 (1998), pp

    K. Kurdyka, On gradients of functions definable in o-minimal structures, Annales de l’institut Fourier, 48 (1998), pp. 769–783

  26. [34]

    S. Łojasiewicz, Une propriété topologique des sous-ensembles analytiques réels, in Les Équations aux Dérivées Partielles (Paris, 1962), Éditions du Centre National de la Recherche Scientifique, Paris, 1963, pp. 87–89

  27. [35]

    Université de Grenoble., 43 (1993), pp

    , Sur la géométrie semi- et sous-analytique, Annales de l’Institut Fourier. Université de Grenoble., 43 (1993), pp. 1575–1595

  28. [36]

    Luo and P

    Z.-Q. Luo and P. Tseng , Error bounds and convergence analysis of feasible descent methods: a general approach, Annals of Operations Research, 46 (1993), pp. 157–178

  29. [37]

    Maulen-Soto, J

    R. Maulen-Soto, J. F adili, and P. Ochs , Inertial methods with viscous and hessian driven damping for non-convex optimization, arXiv preprint arXiv:2407.12518, (2024)

  30. [38]

    Necoara, Y

    I. Necoara, Y. Nesterov, and F. Glineur , Linear convergence of first order methods for non-strongly convex optimization, Mathematical Programming, 175 (2019), pp. 69–107

  31. [39]

    Nesterov, Introductory lectures on convex optimization: A basic course, 2013

    Y. Nesterov, Introductory lectures on convex optimization: A basic course, 2013

  32. [40]

    Palis Jr

    J. Palis Jr. and W. de Melo , Geometric Theory of Dynamical Systems, Springer-Verlag, 1982

  33. [41]

    J. P. Penot , Conditioning convex and nonconvex problems, Journal of Optimization Theory and Applications, 90 (1996), pp. 535–554

  34. [42]

    Perko, Differential equations and dynamical systems, vol

    L. Perko, Differential equations and dynamical systems, vol. 7, Springer Science & Business Media, 2013

  35. [43]

    Polyak and P

    B. Polyak and P. Shcherbakov , Lyapunov functions: An optimization theory perspective, IFAC- PapersOnLine, 50 (2017), pp. 7456–7461

  36. [44]

    B. T. Pol y ak,Gradient methods for the minimisation of functionals, USSRComputationalMathematics and Mathematical Physics, 3 (1963), pp. 864–878

  37. [45]

    , Some methods of speeding up the convergence of iteration methods, USSR Computational Mathematics and Mathematical Physics, 4 (1964), pp. 1–17

  38. [46]

    J. W. Siegel , Accelerated first-order methods: Differential equations and lyapunov functions, arXiv preprint arXiv:1903.05671, (2019)

  39. [47]

    Simsek, F

    B. Simsek, F. Ged, A. Jacot, F. Sp adaro, C. Hongler, W. Gerstner, and J. Brea , Geometry of the loss landscape in overparameterized neural networks: Symmetries and invariances, in International Conference on Machine Learning, PMLR, 2021, pp. 9722–9732

  40. [48]

    Teschl, Ordinary Differential Equations and Dynamical Systems, vol

    G. Teschl, Ordinary Differential Equations and Dynamical Systems, vol. 140, American Mathematical Soc., 2012

  41. [49]

    W ang and J

    Z. W ang and J. Peypouquet , Accelerated gradient methods via inertial systems with hessian-driven damping, arXiv preprint arXiv:2502.16953, (2025)

  42. [50]

    Zhang , New analysis of linear convergence of gradient-type methods via unifying error bound conditions, Mathematical Programming, 180 (2020), pp

    H. Zhang , New analysis of linear convergence of gradient-type methods via unifying error bound conditions, Mathematical Programming, 180 (2020), pp. 371–416

  43. [51]

    Zhang and W

    H. Zhang and W. Yin , Gradient methods for convex minimization: better rates under weaker conditions, arXiv preprint arXiv:1303.4645, (2013)

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.