Pith. sign in

REVIEW 2 major objections 6 minor 2 cited by

On the Iterate Convergence of Bregman Projected Gradient Method

T0 review · 2 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper proves that the Bregman projected gradient method with the Shannon entropy kernel converges to a critical point for continuous subanalytic objectives on polyhedra, under a new scaled Kurdyka–Łojasiewicz property.

desk verdict The SKŁ framework is a real contribution and Proposition 3 survives the stress test; send it to a serious optimization journal. read the letter →

arxiv 2608.05035 v1 pith:DRGOYA4H submitted 2026-08-05 math.OC

classification math.OC MSC 90C2690C30
keywords BregmanprojectedgradientmethodShannonentropykernelscaledKurdyka-Łojasiewiczpropertyiterateconvergencesubanalyticfunctionsrelativesmoothnesslinearnonconvexoptimization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper targets a long-standing open question: do the iterates of the Bregman projected gradient method (BPGM) converge when the kernel is the Shannon entropy and the feasible set is a polyhedron? It answers yes for a broad class of objectives by introducing a new regularity condition, the scaled Kurdyka–Łojasiewicz (SKŁ) property, which adapts the classical KŁ geometry to the multiplicative Bregman geometry induced by the Shannon kernel. The main theorem states that, under relative smoothness and bounded iterates, an SKŁ objective forces the BPGM sequence to converge to a critical point; a companion result shows that every continuous subanalytic function is SKŁ, so the condition is satisfiable without Lipschitz-gradient assumptions. Prior results either needed Lipschitz kernels, strong convexity-type conditions, or local uniqueness near the limit, and a simple example shows that the Euclidean relative-error condition can fail. If correct, the paper gives an essentially nonrestrictive convergence guarantee for BPGM with entropy kernels.

What carries the argument

The central object is the scaled Kurdyka–Łojasiewicz (SKŁ) property. At a point $x^*$ in the domain of the limiting subdifferential, a function $H$ is SKŁ with desingularizing function $\varphi$ if, in a neighborhood of $x^*$ and for values $H(x^*)$ below $H(x)$ below $H(x^*)+\eta$, it holds that $\varphi'(H(x)-H(x^*))\,\operatorname{dist}(0,\operatorname{Diag}(\sqrt{x})\partial H(x))\ge 1$. This replaces the Euclidean distance in the classical KŁ inequality with the norm induced by the Hessian of the Shannon kernel, which is the geometry in which BPGM actually makes progress. The supporting machinery is Proposition 4, which derives scaled sufficient decrease and scaled relative error conditions with constants $\kappa_1=(1/\bar\alpha-L)e^{-\bar\alpha\beta}$ and $\kappa_2=\alpha e^{-\alpha\beta}$; these scaled conditions combine with the SKŁ inequality to yield a finite-length estimate and the convergence conclusion.

What would settle it

Search for a Bregman projected-gradient trajectory satisfying (A1)–(A3) with continuous subanalytic $f$ whose bounded iterates have two distinct accumulation points; Theorem 1 asserts that this cannot happen. A sharper check is to compute $\operatorname{dist}(0,\operatorname{Diag}(\sqrt{x_k})\partial F(x_k))$ along a candidate trajectory and test whether the SKŁ inequality fails at the limit point.

Watch

Extended reading notes

Core claim

The paper's central claim is Theorem 1: suppose $f$ is $L$-relatively smooth with respect to the Shannon entropy kernel $h_S$ over a polyhedron $D_P$ with nonempty relative interior, the step sizes satisfy $\alpha\le\alpha_k\le\bar\alpha$ with $\bar\alpha<1/L$, and the generated BPGM sequence is bounded; if $F=f+\delta_{D_P}$ is an SKŁ function, then the whole sequence converges to a critical point of $F$. Proposition 3 supplies the broad applicability: every proper subanalytic function continuous on its domain is SKŁ, so continuous subanalytic objectives satisfy the hypothesis. Theorem 2 adds that an SKŁ exponent of $1/2$ yields $R$-linear convergence of the iterates and $Q$-linear convergence of the function values, and Theorem 3 transfers the classical KŁ exponent $1/2$ to the SKŁ setting under strict complementarity and local Lipschitz continuity of $\nabla f$. The proof mechanism is a pair of scaled descent and relative-error inequalities in which the diagonal scaling $\operatorname{Diag}(\sqrt{x_k})$ converts the multiplicative Shannon update into a form that supports a KŁ-style convergence argument.

Load-bearing premise

The load-bearing premise is that a finite relative-smoothness constant $L$ is known and all step sizes stay strictly below $1/L$; the proof also takes the boundedness of the iterates as a given, without characterizing when that boundedness holds.

Editorial extensions

If this is right

  • For any continuous subanalytic objective that is relatively smooth with respect to the Shannon entropy over a polyhedron, a bounded BPGM run is guaranteed to converge to a critical point, with no Lipschitz-gradient or convexity assumption on $f$.
  • When the SKŁ exponent is $1/2$, the BPGM iterates converge $R$-linearly and the function values $Q$-linearly, extending linear convergence to problems whose Euclidean KŁ exponent is known to be $1/2$ and which satisfy strict complementarity.
  • The known failure of the Euclidean relative-error condition is bypassed, so convergence proofs no longer require local uniqueness of the limit or additional Bregman growth conditions.
  • Subanalyticity gives a ready-made supply of SKŁ functions, including polynomial, semialgebraic, and other tame objectives, making the key assumption checkable in typical applications.
  • If the SKŁ exponent is strictly larger than $1/2$, the convergence rate is sublinear, so the exponent value directly controls the speed guarantee for BPGM.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension, not claimed in the paper, is that the same SKŁ framework may apply to other Bregman kernels whose Hessian induces a compatible diagonal scaling, not just the Shannon entropy kernel.
  • A testable extension is to compute the SKŁ exponent numerically on practical entropy-kernel problems such as Poisson image deblurring or optimal transport, since only the exponent $1/2$ case yields linear convergence.
  • Example 2 suggests that, without strict complementarity, boundary critical points can have a larger SKŁ exponent than the Euclidean KŁ exponent; a reasonable conjecture is that the worst-case SKŁ exponent is governed by the order of vanishing of the objective along the boundary.
  • Since the proof of Proposition 3 uses only the Łojasiewicz inequality and continuity, the subanalyticity assumption is likely replaceable by any o-minimally definable class of functions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This paper studies the iterate convergence of the Bregman projected gradient method (BPGM) with the Shannon entropy kernel over polyhedral feasible sets. The authors introduce a scaled Kurdyka-Łojasiewicz (SKŁ) property adapted to the Bregman geometry, prove that bounded BPGM sequences converge to critical points when the objective plus indicator satisfies SKŁ (Theorem 1), and show that every continuous subanalytic function is SKŁ (Proposition 3). They further establish linear convergence under an SKŁ exponent of 1/2 (Theorem 2) and prove that a KŁ exponent of 1/2 together with strict complementarity implies an SKŁ exponent of 1/2 (Theorem 3). The paper is technically ambitious, with detailed proofs and extensive appendix material, and it presents explicit examples illustrating the failure of the standard relative-error condition.

Significance. If the results are correct, the SKŁ framework is a genuine advance: it provides the first iterate-convergence result for BPGM with the Shannon kernel for nonconvex objectives without requiring Lipschitz-continuous gradients, and it applies to a broad class of subanalytic objectives. The paper is honest about its assumptions (bounded iterates, known relative-smoothness constant) and includes counterexamples (Examples 1 and 2) that clarify why the standard KŁ framework fails. The main theorems are proved in detail, with the hard multiplier bounds deferred to the appendix. The main obstacle to acceptance is the proof of Proposition 3, where a key chain-rule inclusion is asserted with an incorrect citation and no proof; this is repairable but currently load-bearing.

major comments (2)
  1. [Section 3.2, Eq. (24)] The proof of Proposition 3 asserts, citing [31, Theorem 10.6], that ∇G(y)^⊤∂̂H(G(y)) ⊆ ∂̂g(y) for the coordinatewise square G. The cited theorem in standard variational analysis gives the opposite chain-rule inclusion, ∂(H∘G)(y) ⊆ ∇G(y)^⊤∂H(G(y)), under a qualification condition, so the citation does not justify the assertion. Since this inclusion is the only mechanism connecting the KŁ inequality for g with the scaled subdifferential of H, Proposition 3 is not established as written. Because Proposition 3 is the bridge that turns Theorem 1 into a statement about all continuous subanalytic objectives, this is a load-bearing gap. The missing inclusion is nevertheless true for the specific G(y)=y^2: one can prove the Fréchet version directly from the definition of subgradient and then pass to the outer limit. I therefore expect the gap to be repairable by supplying a short proof and a correct reference.
  2. [Section 3.2, Eq. (25)] The step from the outer limit of a product to the product of outer limits requires a common sequence argument; in general, limsup A(y') · limsup B(y') need not be contained in limsup (A(y')B(y')). In the present setting the step is valid because ∇G is continuous, but the manuscript should state this explicitly. As written, this is another unproven step in the central proof of Proposition 3.
minor comments (6)
  1. [Abstract] The abstract uses the spelling 'subanalytical' while the body consistently uses 'subanalytic'; please unify the terminology.
  2. [Definition 1] Please spell out that dist(0, Diag(√x)∂H(x)) denotes inf_{w∈∂H(x)} ||Diag(√x)w||, since the scaling of an unbounded subdifferential set at boundary points can be ambiguous.
  3. [Section 3.1.2, Corollary 1] The proof of Corollary 1 does not explicitly choose d>0 such that B(x*,d)⊆U and d<ρ; this choice should be stated at the beginning of the proof because condition (14) in Proposition 5 depends on it.
  4. [Lemma 1] The statement of Lemma 1 is for all α>0, but the proof uses sequences α_k; please clarify that the contradiction argument applies to arbitrary sequences α_k>0, so the claimed uniformity in α follows.
  5. [Section 4.1, Theorem 2] The argument that γ∈[0,1) is valid but terse; please spell out that if γ<0 then the right-hand side of (29) is negative while the left-hand side is nonnegative, which is impossible for any k with F(xk)>F(x*).
  6. [Assumptions (A2)-(A3)] The paper assumes the relative-smoothness constant L is known and that ᾱ<1/L, but gives no guidance on how to verify or estimate L for concrete nonconvex objectives. A brief remark on this point would help readers apply the results.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the SKL results are derived from the external Lojasiewicz inequality; the paper's self-citations are motivational only.

full rationale

The central claims are Theorem 1 and Proposition 3, along with Theorems 2–3. Proposition 3 is not circular: it obtains the SKL inequality for H by applying the external nonsmooth Lojasiewicz inequality [10, Theorem 3.1] to the composite g(y)=H(y^2), then converts dist(0,∂g(√x)) into a bound on dist(0,Diag(√x)∂H(x)) via [31, Theorem 10.6]. The cited KL inequality is an independent external input; the paper does not assume that H is already SKL. If the chain-rule conversion is invalid, the proof would be incomplete, but that is a correctness question, not a circularity question. Theorem 1's proof uses the SKL inequality together with the scaled sufficient decrease and scaled relative error conditions derived in Proposition 4 from the BPGM update and relative smoothness. The SKL definition is designed to match those scaled conditions, but it is not defined as 'BPGM converges' and its use in the proof is a genuine derivation. No parameters are fitted to data and then reported as predictions. The self-citation [13] is cited only to motivate the difficulty of spurious stationary points; it does not enter any derivation, and Example 1 independently reproduces the failure of the Euclidean relative error condition. The assumptions (A1)–(A3), including the uncomputed L-relative-smoothness constant and the boundedness of the iterate sequence, are assumption gaps rather than circularity, because they are stated inputs rather than conclusions. Overall, the claimed derivation chain is self-contained against external theorems and does not reduce to its own inputs.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

The central claim rests on standard optimization and subanalytic-geometry machinery plus explicit domain assumptions. No free parameters are fitted to data, and the SKŁ property is a formal definition rather than an extra postulated entity.

assumptions (7)
  • domain assumption A1: Feasible set is DP={Ax=b,x≥0} with A full row rank and DP∩R^n_{++} nonempty.
    Used throughout; enables the explicit multiplier formula in Proposition 1 and the nonempty interior needed for the Shannon kernel.
  • domain assumption A2: f is C^1 and L-relatively smooth with respect to the Shannon entropy kernel on DP.
    Yields Lemma 2's descent inequality; this is the main modeling assumption.
  • domain assumption A3: Step sizes satisfy α_k∈[α,ᾱ] with 0<α<ᾱ<1/L.
    Needed for strict descent and for the uniform constants κ1,κ2 in the scaled conditions.
  • domain assumption BPGM sequence is bounded.
    Needed to select an accumulation point in Corollary 1; no boundedness criterion is provided.
  • domain assumption Strict complementarity holds at x* in Theorem 3.
    Used to prove the SKŁ exponent 1/2 from KŁ exponent 1/2; Example 2 shows the implication can fail without it.
  • standard math Every continuous subanalytic function satisfies the KŁ property with some exponent ([10, Theorem 3.1]).
    Basis of Proposition 3, invoked to transfer the KŁ growth of g(y)=F((y_i^2)) to the scaled SKŁ inequality.
  • standard math Hoffman error bound holds for the polyhedron S in Fact 1.
    Used to compare the projected point x̂ to x in Theorem 3's proof.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the Iterate Convergence of Bregman Projected Gradient Method." pith.science (2026). https://pith.science/paper/DRGOYA4H

@misc{pith2026260805035,
  author       = {Pith},
  title        = {Pith review of: On the Iterate Convergence of Bregman Projected Gradient Method},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DRGOYA4H}},
  note         = {Machine review of arXiv:2608.05035}
}
abstract

The iterate convergence of \textit{Bregman projected gradient method} (BPGM) has remained a long-standing open problem, especially for the widely adopted Shannon entropy kernel. Existing convergence results are often limited, relying on Lipschitz continuity of the kernel's gradient or restrictive conditions on the objective function. In this paper, we develop a novel convergence analysis framework for the BPGM with the Shannon entropy kernel, yielding strong convergence results for a broad class of objective functions under linear constraints. The cornerstone of our framework is a new concept called \textit{scaled Kurdyka-\L{}ojasiewicz} (SK\L{}) property, which captures the local growth behavior of a function under the Bregman geometry. We show that the SK\L{} property ensures the iterate convergence of BPGM and holds for all continuous subanalytical functions. Furthermore, we prove that the BPGM sequence exhibits linear convergence if the problem possesses an SK\L\ exponent of $1/2$. We then furnish the examples of functions with the SK\L\ exponent $1/2$ by proving that the SK\L\ exponent $1/2$ is implied by the K\L{} exponent $1/2$ under strict complementarity and local Lipschitz continuity of the objective's gradient. Building on these novel results, our work takes a first step towards resolving the open problem of BPGM iterate convergence.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Unified Framework for Iterate Convergence of Bregman Proximal Methods

    math.OC 2026-08 conditional novelty 8.0 of 10

    A unified framework using scaled Kurdyka-Lojasiewicz inequalities shows that Bregman proximal point and gradient methods, and mirror flow, converge for closed-domain separable kernels and subanalytic or definable objectives.

  2. Establishing Boundary KKT Convergence of Mirror Descent through Reparameterization

    math.OC 2026-08 conditional novelty 7.0 of 10

    Under verifiable joint conditions on the objective, the Legendre kernel, and the feasible geometry, mirror descent converges to a boundary KKT point with explicit rates.

Reference graph

Works this paper leans on

43 extracted references · 42 canonical work pages · cited by 2 Pith papers

  1. [13]

    Spurious Stationarity and Hardness Results for Bregman Proximal-Type Algorithms

    He Chen, Jiajin Li, and Anthony Man-Cho So. Spurious stationarity and hardness results for mirror descent.arXiv preprint arXiv:2404.08073, version 3, 2026

  2. [1]

    Hessian Riemannian gradient flows in convex programming.SIAM journal on control and optimization, 43(2):477–501, 2004

    Felipe Alvarez, Jérôme Bolte, and Olivier Brahic. Hessian Riemannian gradient flows in convex programming.SIAM journal on control and optimization, 43(2):477–501, 2004

  3. [2]

    On the convergence of the proximal algorithm for nonsmooth functions involving analytic features.Mathematical Programming, Series B, 116:5–16, 2009

    Hedy Attouch and Jérôme Bolte. On the convergence of the proximal algorithm for nonsmooth functions involving analytic features.Mathematical Programming, Series B, 116:5–16, 2009

  4. [3]

    Hedy Attouch, Jérôme Bolte, and Benar Fux Svaiter. Convergence of descent methods for semi-algebraic and tame problems: Proximal algorithms, forward–backward splitting, and regularized Gauss–Seidel methods.Mathematical Programming, Series A, 137(1): 91–129, 2013

  5. [4]

    The rate of convergence of Bregman proximal methods: Local geometry versus regularity versus sharpness.SIAM Journal on Optimization, 34(3):2440–2471, 2024

    Waïss Azizian, Franck Iutzeler, Jérôme Malick, and Panayotis Mertikopoulos. The rate of convergence of Bregman proximal methods: Local geometry versus regularity versus sharpness.SIAM Journal on Optimization, 34(3):2440–2471, 2024

  6. [5]

    A descent lemma beyond Lipschitz gradient continuity: First-order methods revisited and applications.Mathematics of Operations Research, 42(2):330–348, 2017

    Heinz H Bauschke, Jérôme Bolte, and Marc Teboulle. A descent lemma beyond Lipschitz gradient continuity: First-order methods revisited and applications.Mathematics of Operations Research, 42(2):330–348, 2017

  7. [6]

    Heinz H Bauschke, Jérôme Bolte, Jiawei Chen, Marc Teboulle, and Xianfu Wang. On linear convergence of non-Euclidean gradient methods without strong convexity and Lipschitz gradient continuity.Journal of Optimization Theory and Applications, 182(3): 1068–1087, 2019

  8. [7]

    MOS-SIAM Series on Optimization

    Amir Beck.First-Order Methods in Optimization. MOS-SIAM Series on Optimization. Society for Industrial and Applied Mathematics, Philadelphia, Pennsylvania, 2017

Show all 43 references
  1. [8]

    Image deblurring with Poisson data: From cells to galaxies.Inverse Problems, 25(12):123006, 2009

    Mario Bertero, Patrizia Boccacci, Gabriele Desidera, and Giuseppe Vicidomini. Image deblurring with Poisson data: From cells to galaxies.Inverse Problems, 25(12):123006, 2009

  2. [9]

    Barrier operators and associated gradient-like dy- namical systems for constrained minimization problems.SIAM journal on control and optimization, 42(4):1266–1292, 2003

    Jérôme Bolte and Marc Teboulle. Barrier operators and associated gradient-like dy- namical systems for constrained minimization problems.SIAM journal on control and optimization, 42(4):1266–1292, 2003

  3. [10]

    The Łojasiewicz inequality for nons- mooth subanalytic functions with applications to subgradient dynamical systems.SIAM Journal on Optimization, 17(4):1205–1223, 2007

    Jérôme Bolte, Aris Daniilidis, and Adrian Lewis. The Łojasiewicz inequality for nons- mooth subanalytic functions with applications to subgradient dynamical systems.SIAM Journal on Optimization, 17(4):1205–1223, 2007

  4. [11]

    Proximal alternating linearized minimization for nonconvex and nonsmooth problems.Mathematical Programming, Series A, 146(1):459–494, 2014

    Jérôme Bolte, Shoham Sabach, and Marc Teboulle. Proximal alternating linearized minimization for nonconvex and nonsmooth problems.Mathematical Programming, Series A, 146(1):459–494, 2014. 36

  5. [12]

    First order methods beyond convexity and Lipschitz gradient continuity with applications to quadratic inverse problems.SIAM Journal on Optimization, 28(3):2131–2151, 2018

    Jérôme Bolte, Shoham Sabach, Marc Teboulle, and Yakov Vaisbourd. First order methods beyond convexity and Lipschitz gradient continuity with applications to quadratic inverse problems.SIAM Journal on Optimization, 28(3):2131–2151, 2018

  6. [14]

    Springer, New York, 2003

    Francisco Facchinei and Jong-Shi Pang.Finite-Dimensional Variational Inequalities and Complementarity Problems, Volume I. Springer, New York, 2003

  7. [15]

    On approximate solutions of systems of linear inequalities

    Alan J Hoffman. On approximate solutions of systems of linear inequalities. InSelected Papers of Alan J Hoffman: With Commentary, pages 174–176. World Scientific, 2003

  8. [16]

    Riemannian proximal gradient methods.Mathematical Programming, 194(1):371–413, 2022

    Wen Huang and Ke Wei. Riemannian proximal gradient methods.Mathematical Programming, 194(1):371–413, 2022

  9. [17]

    Central paths, generalized proximal point methods, and Cauchy trajectories in Riemannian manifolds.SIAM Journal on Control and Optimization, 37(2):566–588, 1999

    Alfredo N Iusem, BF Svaiter, and João Xavier da Cruz Neto. Central paths, generalized proximal point methods, and Cauchy trajectories in Riemannian manifolds.SIAM Journal on Control and Optimization, 37(2):566–588, 1999

  10. [18]

    Non-Convex Optimization for Machine Learning

    Prateek Jain and Purushottam Kar. Non-Convex Optimization for Machine Learning. Foundations and Trends®in Machine Learning, 10(3–4):142–336, 2017

  11. [19]

    Unifying mirror descent and dual averaging.Mathematical Programming, Series A, 199(1-2):793–830, 2023

    Anatoli Juditsky, Joon Kwon, and Éric Moulines. Unifying mirror descent and dual averaging.Mathematical Programming, Series A, 199(1-2):793–830, 2023

  12. [20]

    Bregman Finito/MISO for nonconvex regularized finite sum minimization without Lipschitz gradient continuity.SIAM Journal on Optimization, 32(3):2230–2262, 2022

    Puya Latafat, Andreas Themelis, Masoud Ahookhosh, and Panagiotis Patrinos. Bregman Finito/MISO for nonconvex regularized finite sum minimization without Lipschitz gradient continuity.SIAM Journal on Optimization, 32(3):2230–2262, 2022

  13. [21]

    Guoyin Li and Ting Kei Pong. Calculus of the exponent of Kurdyka-Łojasiewicz inequality and its applications to linear convergence of first-order methods.Foundations of Computational Mathematics, 18(5):1199–1232, 2018

  14. [22]

    A convergent single-loop algorithm for relaxation of Gromov-Wasserstein in graph data

    Jiajin Li, Jianheng Tang, Lemin Kong, Huikang Liu, Jia Li, Anthony Man-Cho So, and Jose Blanchet. A convergent single-loop algorithm for relaxation of Gromov-Wasserstein in graph data. InThe Eleventh International Conference on Learning Representations, 2023

  15. [23]

    Relatively smooth convex optimization by first-order methods, and applications.SIAM Journal on Optimization, 28(1):333–354, 2018

    Haihao Lu, Robert M Freund, and Yurii Nesterov. Relatively smooth convex optimization by first-order methods, and applications.SIAM Journal on Optimization, 28(1):333–354, 2018

  16. [24]

    Errorbounds forquadraticsystems

    Zhi-QuanLuoandJosFSturm. Errorbounds forquadraticsystems. InHigh Performance Optimization, pages 383–404. Springer, Boston, 2000. 37

  17. [25]

    Error bounds and convergence analysis of feasible descent methods: A general approach.Annals of Operations Research, 46(1):157–178, 1993

    Zhi-Quan Luo and Paul Tseng. Error bounds and convergence analysis of feasible descent methods: A general approach.Annals of Operations Research, 46(1):157–178, 1993

  18. [26]

    Pearson, South Carolina, 1999

    Arthur Mattuck.Introduction to Analysis. Pearson, South Carolina, 1999

  19. [27]

    Self-scaled barriers and interior-point methods for convex programming.Mathematics of Operations research, 22(1):1–42, 1997

    Yu E Nesterov and Michael J Todd. Self-scaled barriers and interior-point methods for convex programming.Mathematics of Operations research, 22(1):1–42, 1997

  20. [28]

    Springer, Berlin, Heidelberg, 2018

    Yurii Nesterov.Lectures on Convex Optimization, volume 137 ofSpringer Optimization and Its Applications. Springer, Berlin, Heidelberg, 2018

  21. [29]

    Springer Series in Operations Re- search and Financial Engineering, 2nd edn

    Stephen Wright Nocedal.Numerical Optimization. Springer Series in Operations Re- search and Financial Engineering, 2nd edn. Springer, New York, 2006

  22. [30]

    Computational Optimal Transport: With Applica- tions to Data Science.Foundations and Trends®in Machine Learning, 11(5-6):355–607, 2019

    Gabriel Peyré, Marco Cuturi, et al. Computational Optimal Transport: With Applica- tions to Data Science.Foundations and Trends®in Machine Learning, 11(5-6):355–607, 2019

  23. [31]

    Springer Science & Business Media, Berlin, Heidelberg, 2009

    R Tyrrell Rockafellar and Roger J-B Wets.Variational Analysis, volume 317. Springer Science & Business Media, Berlin, Heidelberg, 2009

  24. [32]

    Nonconvex optimization for signal processing and machine learning [from the guest editors].IEEE Signal Processing Magazine, 37(5):15–17, 2020

    Anthony Man-Cho So, Prateek Jain, Wing-Kin Ma, and Gesualdo Scutari. Nonconvex optimization for signal processing and machine learning [from the guest editors].IEEE Signal Processing Magazine, 37(5):15–17, 2020

  25. [33]

    A new non- monotonic algorithm for pet image reconstruction

    Suvrit Sra, Dongmin Kim, Inderjit Dhillon, and Bernhard Schölkopf. A new non- monotonic algorithm for pet image reconstruction. In2009 IEEE Nuclear Science Symposium Conference Record (NSS/MIC), pages 2500–2502, 2009

  26. [34]

    Neural Information Processing Series

    Suvrit Sra, Sebastian Nowozin, and Stephen J Wright.Optimization for Machine Learning. Neural Information Processing Series. MIT Press, Cambridge, Massachusetts, 2012

  27. [35]

    Approximate Bregman proximal gradient algo- rithm for relatively smooth nonconvex optimization.Computational Optimization and Applications, 90(1):227–256, 2025

    Shota Takahashi and Akiko Takeda. Approximate Bregman proximal gradient algo- rithm for relatively smooth nonconvex optimization.Computational Optimization and Applications, 90(1):227–256, 2025

  28. [36]

    New Bregman proximal type algorithms for solving DC optimization problems.Computational Optimization and Applications, 83(3):893–931, 2022

    Shota Takahashi, Mituhiro Fukuda, and Mirai Tanaka. New Bregman proximal type algorithms for solving DC optimization problems.Computational Optimization and Applications, 83(3):893–931, 2022

  29. [37]

    A simplified view of first order methods for optimization.Mathematical Programming, Series B, 170(1):67–96, 2018

    Marc Teboulle. A simplified view of first order methods for optimization.Mathematical Programming, Series B, 170(1):67–96, 2018

  30. [38]

    Approximation accuracy, gradient methods, and error bound for structured convex optimization.Mathematical Programming, Series B, 125(2):263–295, 2010

    Paul Tseng. Approximation accuracy, gradient methods, and error bound for structured convex optimization.Mathematical Programming, Series B, 125(2):263–295, 2010. 38

  31. [39]

    Inertial proximal gradient methods with Bregman regularization for a class of nonconvex optimization problems

    Zhongming Wu, Chongshou Li, Min Li, and Andrew Lim. Inertial proximal gradient methods with Bregman regularization for a class of nonconvex optimization problems. Journal of Global Optimization, 79:617–644, 2021

  32. [40]

    Gromov-Wasserstein learning for graph matching and node embedding

    HongtengXu, DixinLuo, HongyuanZha, andLawrenceCarinDuke. Gromov-Wasserstein learning for graph matching and node embedding. InProceedings of the 36th International Conference on Machine Learning, pages 6932–6941. PMLR, 2019

  33. [41]

    Proximal-like incremental aggregated gradient method with linear convergence under Bregman distance growth conditions

    Hui Zhang, Yu-Hong Dai, Lei Guo, and Wei Peng. Proximal-like incremental aggregated gradient method with linear convergence under Bregman distance growth conditions. Mathematics of Operations Research, 46(1):61–81, 2021

  34. [42]

    A unified approach to error bounds for structured convex optimization problems.Mathematical Programming, Series A, 165(2):689–728, 2017

    Zirui Zhou and Anthony Man-Cho So. A unified approach to error bounds for structured convex optimization problems.Mathematical Programming, Series A, 165(2):689–728, 2017

  35. [43]

    Level-set subdifferential error bounds and linear convergence of Bregman proximal gradient method.Journal of Optimization Theory and Applications, 189(3):889–918, 2021

    Daoli Zhu, Sien Deng, Minghua Li, and Lei Zhao. Level-set subdifferential error bounds and linear convergence of Bregman proximal gradient method.Journal of Optimization Theory and Applications, 189(3):889–918, 2021. 39

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.