Pith. sign in

REVIEW 4 minor 38 references

The ballistic limit of the log-Sobolev constant equals the Polyak-{\L}ojasiewicz constant

T0 review · 0 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The low-temperature limit of the log-Sobolev constant equals the Polyak–Łojasiewicz constant.

desk verdict Exact low-temperature bridge between log-Sobolev and PL constants, with a solid proof and an honest discussion of the necessary unique-minimizer assumption. read the letter →

arxiv 2411.11415 v1 pith:ET3F45YM submitted 2024-11-18 math.PR math.FAmath.OC

classification math.PRmath.FAmath.OC MSC 60J6060E15
keywords ballisticlog-SobolevconstantPolyak-Lojasiewiczlow-temperatureGibbsmeasuresLangevindynamicsinequalityPoincaregradientflowlimit
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proves an exact identity linking two previously separate rates: the speed at which the Langevin diffusion converges to the Gibbs measure proportional to $e^{-f/t}$ at low temperature, measured by the log-Sobolev constant $C_{\mathsf{LS}}(\mu_t)$, and the speed at which gradient flow converges for $f$, measured by the Polyak–Łojasiewicz constant $C_{\mathsf{PL}}(f)$. Under the assumptions that $f$ is twice continuously differentiable, has a unique global minimizer, and satisfies $\Delta f \le L(1+\|\nabla f\|^2)$, the paper shows that the limit $\lim_{t\to 0^+} C_{\mathsf{LS}}(\mu_t)/t$ exists exactly when $C_{\mathsf{PL}}(f)$ is finite, and the two numbers are equal. The same argument yields the companion formula $\lim_{t\to 0^+} C_{\mathsf{P}}(\mu_t)/t = 1/\lambda_{\min}(\nabla^2 f(x^\star))$ for the Poincaré constant. A sympathetic reader would care because the analogy between sampling and optimization becomes a theorem with an exact constant, and the Polyak–Łojasiewicz constant acquires a new meaning as a low-temperature spectral quantity of Gibbs measures.

What carries the argument

The machinery views the log-Sobolev inequality as a Polyak–Łojasiewicz inequality for the Kullback–Leibler functional on the space of probability measures, then proves matching lower and upper bounds. The lower bound localizes: test measures concentrated near any point $x$ force the log-Sobolev inequality to imply the pointwise gradient inequality defining $C_{\mathsf{PL}}(f)$. The upper bound compares $\mu_t$ with a Gaussian of covariance $t[\nabla^2 f(x^\star)]^{-1}$, splits space into a small ball around the minimizer and its complement, and uses quadratic growth of PL functions to control the tail. The final constant is assembled by first proving a defective log-Sobolev inequality and then applying an improved tightening lemma that converts it into a full log-Sobolev inequality without losing a factor of two. The companion Poincaré result uses a Lyapunov-function argument to show $C_{\mathsf{P}}(\mu_t)=O(t)$, with the exact prefactor identified by a Gaussian rescaling near the minimizer.

What would settle it

Compute the ballistic log-Sobolev constant for $f(x)=\frac{\alpha}{2}\,d(x,K)^2$ with a convex set $K$ of nonempty interior: $C_{\mathsf{PL}}(f)$ is finite, yet $\mu_t$ converges to the uniform measure on $K$, making $C_{\mathsf{bLS}}(f)$ infinite and showing that removing the unique-minimizer assumption breaks the theorem.

Watch

Extended reading notes

Core claim

The central claim is an equality of constants. Define $C_{\mathsf{PL}}(f)$ as the least $C$ such that $f(x)-f^\star \le \frac{C}{2}\|\nabla f(x)\|^2$ for all $x$; this is the constant that controls uniform exponential convergence of gradient flow. For $\mu_t \propto e^{-f/t}$, let $C_{\mathsf{LS}}(\mu_t)$ be the smallest constant such that $\mathrm{KL}(\nu\|\mu_t)\le \frac{C}{2}\mathrm{FI}(\nu\|\mu_t)$ for all smooth compactly supported $\nu$. Theorem 1 states that when $f\in C^2(\mathbb{R}^d)$ has a unique global minimizer and $\Delta f\le L(1+\|\nabla f\|^2)$, the ballistic log-Sobolev constant $C_{\mathsf{bLS}}(f):=\lim_{t\to0^+} C_{\mathsf{LS}}(\mu_t)/t$ exists if and only if $C_{\mathsf{PL}}(f)<\infty$, and in that case $C_{\mathsf{bLS}}(f)=C_{\mathsf{PL}}(f)$. Theorem 2 states that $C_{\mathsf{bP}}(f):=\lim_{t\to0^+} C_{\mathsf{P}}(\mu_t)/t$ equals $1/\lambda_{\min}(\nabla^2 f(x^\star))$. The authors also show the uniqueness assumption is not removable: for $f(x)=\frac{\alpha}{2}d(x,K)^2$ with a convex set $K$ of nonempty interior, $C_{\mathsf{PL}}(f)$ is finite but the Gibbs measures converge to the uniform measure on $K$, so the ballistic log-Sobolev constant is infinite.

Load-bearing premise

The load-bearing premise is that the function has exactly one global minimum and that its curvature does not grow faster than a constant times $1+\|\nabla f\|^2$; if the set of minima has any interior, the Gibbs measures spread out instead of concentrating, and the equality can fail.

Editorial extensions

If this is right

  • If $f$ satisfies the assumptions, then $C_{\mathsf{LS}}(\mu_t)=t\,C_{\mathsf{PL}}(f)+o(t)$, so the exponential rate of convergence of the Langevin dynamics in Kullback–Leibler divergence is dictated, in the low-temperature limit, by the same constant that governs gradient flow.
  • To leading order, the Poincaré constant of the same Gibbs measures is $C_{\mathsf{P}}(\mu_t)\sim t/\lambda_{\min}(\nabla^2 f(x^\star))$; this ballistic limit sees only the local Hessian at the minimizer, not the global landscape.
  • The Polyak–Łojasiewicz constant can be read off from spectral data of Gibbs measures, giving a new characterization of the PL condition in terms of sampling.
  • The paper's non-asymptotic estimates imply $C_{\mathsf{LS}}(\mu_t)\le C_{\mathsf{PL}}(f)t + \text{lower-order terms}$, an improvement by a factor of $t$ over known constant-order bounds for unique-minimizer PL landscapes.
  • When the global minimizer is not unique, the exact equality fails; the paper leaves the refined conjecture $C_{\mathsf{LS}}(\mu_t)-C_{\mathsf{LS}}(\mu_0)\sim t\,C_{\mathsf{PL}}(f)$ as an open problem.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to use numerical estimates of the spectral gap of discretized Gibbs measures at small $t$ as an estimator of $C_{\mathsf{PL}}(f)$; the paper does not address dimension dependence or discretization error for such an estimator.
  • The identity suggests a finite-temperature notion of PL constant, $C_{\mathsf{PL},t}(f):=C_{\mathsf{LS}}(\mu_t)/t$, whose limit is the usual PL constant; studying its approach to the limit could inform annealing schedules, but this direction is not in the paper.
  • For functions whose minimizer set has positive dimension, the failure of the equality indicates that the ballistic limit should depend on the shape of the minimizer set rather than only on $C_{\mathsf{PL}}(f)$; the paper does not propose a general formula for that case.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. The paper proves that for a C^2 function f on R^d with a unique global minimizer and with Δf ≤ L(1+||∇f||^2), the low-temperature ('ballistic') limit C_bLS(f)=lim_{t→0+} C_LS(μ_t)/t of the log-Sobolev constant of μ_t ∝ e^{-f/t} equals the Polyak–Łojasiewicz constant C_PL(f) (Theorem 1). The lower bound (Theorem 11) is obtained by testing the LSI against smooth measures supported near arbitrary points; the upper bound (Theorem 12) uses a Gaussian LSI comparison near the minimizer, a small/large scale split with radius r0=A√t, tail control via quadratic growth and the PL inequality, and an improved Rothaus tightening (Lemma 3). The paper also proves an exact formula for the corresponding Poincaré constant, C_bP(f)=1/λ_min(∇^2f(x*)) (Theorem 2), and shows through a distance-to-convex-set example that the uniqueness assumption is necessary; the multiple-minimizer case is left as open problem (1.6).

Significance. If correct, these results give a clean and quantitatively sharp bridge between optimization and sampling: the normalized low-temperature log-Sobolev constant captures the global PL constant, while the Poincaré constant only sees the Hessian at the minimizer. The paper's proofs are largely self-contained, the main hypotheses are explicit, and the authors honestly delineate the necessity of the unique-minimizer assumption and the open general formula. The non-asymptotic bounds in Remarks 9 and 13 and in Proposition 14 are useful byproducts. The only external input that carries the exact constant is the improved tightening lemma [Wan24, Prop. 5]; if that lemma is correct as quoted, the main argument is coherent and the result is likely to become a standard reference.

minor comments (4)
  1. [Section 4, Step 1] In the KL-decomposition displayed at the start of Step 1, the term '-1/t' inside the integral appears to be spurious, and the quadratic term should be written as (1/2)||x||^2_{Σ_t^{-1}} to match the density of N(0,Σ_t) under the notation fixed at the end of the introduction; the subsequent cancellation with the Gaussian LSI is correct once this notation is adjusted.
  2. [Section 4, Step 4] The denominator in the displayed bound for (1/t^2)∫||∇f||^2 e^{-g} should be 1 - 2L_1 t, not 1 - 2t/L_1; the stated condition t < 1/(2L_1) is consistent with the former, and the subsequent order estimates are unaffected.
  3. [Section 2.1, Lemma 3] Please add a precise pointer to [Wan24, Prop. 5] and, if space permits, a short proof sketch, because this lemma is the one imported ingredient that fixes the exact leading-order constant in Theorem 12.
  4. [Throughout] There are several OCR/encoding artifacts in the text (for example, the inserted '/suppress' in Polyak–Łojasiewicz and the malformed '{ f /greaterorequalslant0 }'); these should be cleaned before the final version.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the main theorem is proved by independent lower and upper bounds, with no fitted inputs or load-bearing self-citations.

full rationale

The paper's central claim, C_bLS(f) = C_PL(f), is established by two independent arguments: Theorem 11 derives the PL inequality directly from the log-Sobolev inequality, and Theorem 12 upper-bounds the log-Sobolev constant using the PL condition as a structural input. Neither bound assumes the conclusion; the two constants are defined independently, and no parameter is fitted to data or to the target quantity. The only imported result bearing on the exact constant is Lemma 3, the improved tightening lemma from [Wan24], which is a separate work by a different author and is not equivalent to the theorem being proved. The paper's self-citations are confined to related-work remarks and a standard integration-by-parts trick credited to [CEL+24, Lemma 20]; these are not load-bearing. The unique-minimizer assumption is explicitly shown to be necessary via a counterexample, and the paper openly states the multiple-minimizer case as an open problem rather than hiding it. No circular step, fitted-input-as-prediction, ansatz-smuggling, or renamed known result was found.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claim is supported by the explicit theorem hypotheses and by external functional inequalities from the literature. No fitted parameters and no new postulated entities appear. The most important external input is the improved tightening lemma from Wang (2024), which is critical for obtaining the exact constant in the upper bound.

assumptions (5)
  • domain assumption Improved tightening lemma: from a (C,D)-defective LSI with finite Poincare constant, C_LS(mu) <= C + D/2 C_P(mu) (Lemma 3, from Wang 2024 Prop 5).
    Recent external result, not proved in this paper, used in the final step of Theorem 12 to obtain the exact constant without a factor of 2.
  • standard math Quadratic growth inequality for PL functions: if C_PL(h) is finite, then h(x)-h* >= d(x,S)^2/(2 C_PL(h)) (Proposition 6, attributed to Otto-Villani and KNS16).
    Used throughout the upper bound and in Lemma 15 to control tails and the Hessian at the minimizer.
  • standard math Poincare Lyapunov criterion (Lemma 10, from Bakry-Gentil-Ledoux Theorem 4.6.2).
    Used in Lemma 8 to prove the O(t) Poincare bound C_P(mu_t) required for the tightening step.
  • standard math Gaussian log-Sobolev inequality and Bakry-Emery criterion for strongly log-concave measures.
    Used in the local small-scale approximation in Theorem 12 and in the background section.
  • domain assumption Explicit hypotheses of the main theorem: f in C^2(R^d), f has a unique global minimizer, Delta f <= L0 + L1 ||grad f||^2, and mu_t = exp(-f/t)/Z_t is a probability measure.
    These are the stated conditions under which Theorems 1, 2, and 12 are proved. The unique minimizer condition is shown to be necessary for the equality.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The ballistic limit of the log-Sobolev constant equals the Polyak-{\L}ojasiewicz constant." pith.science (2026). https://pith.science/paper/ET3F45YM

@misc{pith2026241111415,
  author       = {Pith},
  title        = {Pith review of: The ballistic limit of the log-Sobolev constant equals the Polyak-\Lojasiewicz constant},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ET3F45YM}},
  note         = {Machine review of arXiv:2411.11415}
}
abstract

The Polyak-Lojasiewicz (PL) constant of a function $f \colon \mathbb{R}^d \to \mathbb{R}$ characterizes the best exponential rate of convergence of gradient flow for $f$, uniformly over initializations. Meanwhile, in the theory of Markov diffusions, the log-Sobolev (LS) constant plays an analogous role, governing the exponential rate of convergence for the Langevin dynamics from arbitrary initialization in the Kullback-Leibler divergence. We establish a new connection between optimization and sampling by showing that the low temperature limit $\lim_{t\to 0^+} t^{-1} C_{\mathsf{LS}}(\mu_t)$ of the LS constant of $\mu_t \propto \exp(-f/t)$ is exactly the PL constant of $f$, under mild assumptions. In contrast, we show that the corresponding limit for the Poincar\'e constant is the inverse of the smallest eigenvalue of $\nabla^2 f$ at the minimizer.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

38 extracted references · 33 canonical work pages

  1. [1]

    Jason Altschuler, Sinho Chewi, Patrik R Gerber, and Austin J. Stromme. Averaging on the B ures-- W asserstein manifold: dimension-free convergence of gradient descent. Advances in Neural Information Processing Systems , 2021

  2. [2]

    A simple proof of the P oincar \'e inequality for a large class of probability measures including the log-concave case

    Dominique Bakry, Franck Barthe, Patrick Cattiaux, and Arnaud Guillin. A simple proof of the P oincar \'e inequality for a large class of probability measures including the log-concave case. Electron. Commun. Probab. , 13:60--66, 2008

  3. [3]

    Metastability in reversible diffusion processes

    Anton Bovier, Michael Eckhoff, V \'e ronique Gayrard, and Markus Klein. Metastability in reversible diffusion processes. I. S harp asymptotics for capacities and exit times. J. Eur. Math. Soc. (JEMS) , 6(4):399--424, 2004

  4. [4]

    Kramers' law: validity, derivations and generalisations

    Nils Berglund. Kramers' law: validity, derivations and generalisations. Markov Process. Relat. Fields , 19(3):459--490, 2013

  5. [5]

    Analysis and geometry of Markov diffusion operators , volume 103

    Dominique Bakry, Ivan Gentil, and Michel Ledoux. Analysis and geometry of Markov diffusion operators , volume 103. Springer, 2014

  6. [6]

    Erdogdu, Mufan (B.) Li, Ruoqi Shen, and Matthew S

    Sinho Chewi, Murat A. Erdogdu, Mufan (B.) Li, Ruoqi Shen, and Matthew S. Zhang. Analysis of L angevin M onte C arlo from P oincar\' e to log- S obolev. Found. Comput. Math. , 24(4), 2024

  7. [7]

    Gerber, Holden Lee, and Chen Lu

    Sinho Chewi, Patrik R. Gerber, Holden Lee, and Chen Lu. Fisher information lower bounds for sampling. In Shipra Agrawal and Francesco Orabona, editors, Proceedings of the 34th International Conference on Algorithmic Learning Theory , volume 201 of Proceedings of Machine Learning Research , pages 375--410. PMLR, 2 2023

  8. [8]

    A note on T alagrand’s transportation inequality and logarithmic S obolev inequality

    Patrick Cattiaux, Arnaud Guillin, and Li-Ming Wu. A note on T alagrand’s transportation inequality and logarithmic S obolev inequality. Probab. Theory Relat. Fields , 148:285--304, 2010

Show all 38 references
  1. [9]

    Log-concave sampling

    Sinho Chewi. Log-concave sampling. Book draft available at https://chewisinho.github.io , 2024

  2. [10]

    Colding and William P

    Tobias H. Colding and William P. Minicozzi II . ojasiewicz inequalities and applications. arXiv preprint arXiv:1402.5087 , 2014

  3. [11]

    Gradient descent algorithms for B ures-- W asserstein barycenters

    Sinho Chewi, Tyler Maunu, Philippe Rigollet, and Austin J Stromme. Gradient descent algorithms for B ures-- W asserstein barycenters. In Conference on Learning Theory , pages 1276--1304. PMLR, 2020

  4. [12]

    Chen and Karthik Sridharan

    August Y. Chen and Karthik Sridharan. From optimization to sampling via L yapunov potentials. arXiv preprint arXiv:2410.02979 , 2024

  5. [13]

    Further and stronger analogy between sampling and optimization: L angevin M onte C arlo and gradient descent

    Arnak Dalalyan. Further and stronger analogy between sampling and optimization: L angevin M onte C arlo and gradient descent. In Conference on Learning Theory , pages 678--689. PMLR, 2017

  6. [14]

    The activated complex in chemical reactions

    Henry Eyring. The activated complex in chemical reactions. J. Chem. Phys. , 3(2):107--115, 1935

  7. [15]

    Metastability in reversible diffusion processes II : precise asymptotics for small eigenvalues

    V \'e ronique Gayrard, Anton Bovier, and Markus Klein. Metastability in reversible diffusion processes II : precise asymptotics for small eigenvalues. J. Eur. Math. Soc. (JEMS) , 7(1):69--99, 2005

  8. [16]

    Gelfand and Sanjoy K

    Saul B. Gelfand and Sanjoy K. Mitter. Recursive stochastic algorithms for global optimization in R^d . SIAM J. Control Optim. , 29(5):999--1018, 1991

  9. [17]

    Holley, Shigeo Kusuoka, and Daniel W

    Richard A. Holley, Shigeo Kusuoka, and Daniel W. Stroock. Asymptotics of the spectral gap with applications to the theory of simulated annealing. J. Funct. Anal , 83(2):333--347, 1989

  10. [18]

    Linear convergence of gradient and proximal-gradient methods under the P olyak-- ojasiewicz condition

    Hamed Karimi, Julie Nutini, and Mark Schmidt. Linear convergence of gradient and proximal-gradient methods under the P olyak-- ojasiewicz condition . Joint European Conference on Machine Learning and Knowledge Discovery in Databases , 2016

  11. [19]

    Hendrik A. Kramers. Brownian motion in a field of force and the diffusion model of chemical reactions. Physica , 7(4):284--304, 1940

  12. [20]

    Improved convergence rate of stochastic gradient L angevin dynamics with variance reduction and its application to optimization

    Yuri Kinoshita and Taiji Suzuki. Improved convergence rate of stochastic gradient L angevin dynamics with variance reduction and its application to optimization. Advances in Neural Information Processing Systems , 35:19022--19034, 2022

  13. [21]

    Mufan (B.) Li and Murat A. Erdogdu. Riemannian L angevin algorithm for solving semidefinite programs. Bernoulli , 29(4):3093--3113, 2023

  14. [22]

    On escape time, L yapunov function, P oincar\' e inequality, and the KLS conjecture beyond convexity

    Mufan (B.) Li. On escape time, L yapunov function, P oincar\' e inequality, and the KLS conjecture beyond convexity. https://mufan-li.github.io/lyapunov_escape/, 2021

  15. [23]

    A topological property of real analytic subsets (in F rench)

    Stanislaw ojasiewicz. A topological property of real analytic subsets (in F rench). Coll. du CNRS, Les \'e quations aux d \'e riv \'e es partielles , 117(87-89):2, 1963

  16. [24]

    Loss landscapes and optimization in over-parameterized non-linear systems and neural networks

    Chaoyue Liu, Libin Zhu, and Mikhail Belkin. Loss landscapes and optimization in over-parameterized non-linear systems and neural networks. Appl. Comput. Harmon. Anal. , 59:85--116, 2022

  17. [25]

    Poincar \'e and logarithmic S obolev inequalities by decomposition of the energy landscape

    Georg Menz and Andr \'e Schlichting. Poincar \'e and logarithmic S obolev inequalities by decomposition of the energy landscape. Ann. Probab. , pages 1809--1884, 2014

  18. [26]

    The geometry of dissipative evolution equations: the porous medium equation

    Felix Otto. The geometry of dissipative evolution equations: the porous medium equation. Commun. Partial Differ. Equ. , 26(1-2):101--174, 2001

  19. [27]

    Generalization of an inequality by T alagrand and links with the logarithmic S obolev inequality

    Felix Otto and C \'e dric Villani. Generalization of an inequality by T alagrand and links with the logarithmic S obolev inequality. J. Funct. Anal. , 173(2):361--400, 2000

  20. [28]

    Boris T. Polyak. Gradient methods for solving equations and inequalities (in R ussian). USSR Computational Mathematics and Mathematical Physics , 4(6):17--32, 1964

  21. [29]

    Oscar S. Rothaus. Analytic inequalities, isoperimetric inequalities and logarithmic S obolev inequalities. J. Funct. Anal. , 64(2):296--313, 1985

  22. [30]

    Non-convex learning via stochastic gradient L angevin dynamics: a nonasymptotic analysis

    Maxim Raginsky, Alexander Rakhlin, and Matus Telgarsky. Non-convex learning via stochastic gradient L angevin dynamics: a nonasymptotic analysis. In Conference on Learning Theory , pages 1674--1703. PMLR, 2017

  23. [31]

    Philippe Rigollet and Austin J. Stromme. On the sample complexity of entropic optimal transport. arXiv preprint 2206.13472 , 2022

  24. [32]

    Austin J. Stromme. Minimum intrinsic dimension scaling for entropic optimal transport. arXiv preprint 2306.03398 , 2023

  25. [33]

    Local optimality and generalization guarantees for the L angevin algorithm via empirical metastability

    Belinda Tzen, Tengyuan Liang, and Maxim Raginsky. Local optimality and generalization guarantees for the L angevin algorithm via empirical metastability. In Conference On Learning Theory , pages 857--875. PMLR, 2018

  26. [34]

    Topics in optimal transportation , volume 58 of Graduate Studies in Mathematics

    C\' e dric Villani. Topics in optimal transportation , volume 58 of Graduate Studies in Mathematics . American Mathematical Society, Providence, RI, 2003

  27. [35]

    Analysis for diffusion processes on R iemannian manifolds , volume 18 of Advanced Series on Statistical Science & Applied Probability

    Feng-Yu Wang. Analysis for diffusion processes on R iemannian manifolds , volume 18 of Advanced Series on Statistical Science & Applied Probability . World Scientific Publishing Co. Pte. Ltd., Hackensack, NJ, 2014

  28. [36]

    Uniform log- S obolev inequalites for mean field particles with flat-convex energy

    Songbo Wang. Uniform log- S obolev inequalites for mean field particles with flat-convex energy. arXiv preprint arXiv:2408.03283 , 2024

  29. [37]

    Global convergence of L angevin dynamics based algorithms for nonconvex optimization

    Pan Xu, Jinghui Chen, Difan Zou, and Quanquan Gu. Global convergence of L angevin dynamics based algorithms for nonconvex optimization. arXiv preprint arXiv:1707.06618 , 2017

  30. [38]

    A hitting time analysis of stochastic gradient L angevin dynamics

    Yuchen Zhang, Percy Liang, and Moses Charikar. A hitting time analysis of stochastic gradient L angevin dynamics. In Conference on Learning Theory , pages 1980--2022. PMLR, 2017

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.