Pith. sign in

REVIEW 5 major objections 5 minor 67 references

Poincare Inequality for Local Log-Polyak-\L ojasiewicz Measures: Non-asymptotic Analysis in Low-temperature Regime

T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read At low temperatures, the Poincaré constant of a Gibbs measure is bounded below by the first Laplace–Beltrami eigenvalue of its minimizer manifold.

desk verdict New and plausible reduction of the low-temperature Poincare constant to the Laplace-Beltrami eigenvalue of the minimizer manifold, but the no-saddle assumption is doing heavy lifting and a few technical inconsistencies need fixing. read the letter →

arxiv 2501.00429 v2 pith:LY3SRTPB submitted 2024-12-31 math.PR math.FA

classification math.PRmath.FA MSC 60J6058J5035P15
keywords Poincaréinequalitynon-log-concavemeasureLangevindynamicsPolyak–ŁojasiewiczLaplace–Beltramieigenvaluelow-temperatureregimeGibbsnon-isolatedminima
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proves that low-temperature Gibbs measures built from a class of non-convex potentials with non-isolated minima satisfy a Poincaré inequality whose constant stays bounded away from zero as the temperature $\epsilon$ tends to zero. This class, called Log-PL$^\circ$ measures, consists of potentials with a local Polyak–Łojasiewicz inequality near their minimizers, with the minimizer set forming a connected compact smooth submanifold $S$ of the ambient space. The central estimate is $\rho_{\mu_\epsilon} \ge C_P \lambda_1(S)$ for all sufficiently small $\epsilon$, where $\lambda_1(S)$ is the first non-zero eigenvalue of the Laplace–Beltrami operator on $S$. Because the Poincaré constant controls the spectral gap of Langevin dynamics, the paper concludes that Langevin diffusion and its discretizations converge to equilibrium in time $\tilde{O}(1/\epsilon)$, a sub-exponential rate normally associated with log-concave measures. The result matters because over-parameterized learning problems empirically have landscapes with connected, degenerate minimizer sets, which previous theory did not cover.

What carries the argument

The load-bearing object is the Log-PL$^\circ$ measure, defined by a potential that satisfies the local Polyak–Łojasiewicz inequality $|\nabla V|^2 \ge \nu (V - \min V)$ near each connected component of its minimizer set, together with the no-saddle condition that every critical point outside those neighborhoods is a strict local maximum. These assumptions force all local minima to lie in one connected component, and force the global minimizer set $S$ to be a compact $C^2$ embedded submanifold without boundary. The proof's machinery is a two-step reduction: first a Lyapunov-function criterion (following Menz and Schlichting) reduces the Poincaré constant of $\mu_\epsilon$ to the Neumann eigenvalue $\lambda^n_1(U)$ of the Laplacian on the thin tube $U = S_{\sqrt{C\epsilon}}$; then, because a thin tube around an embedded submanifold is a tubular neighborhood, a stability analysis using the Weyl tube formula, tensorization of the Poincaré inequality, and boundedness of the second fundamental form shows $\lambda^n_1(U)$ differs from $\lambda_1(S)$ only by $O(\sqrt{\epsilon})$. Assembling the two steps gives the temperature-independent constant.

What would settle it

Compute or simulate the spectral gap of Langevin dynamics for the paper's own example $V(x)=\|x\|^3/3-\|x\|^2/2$ in $\mathbb{R}^2$, whose minimizer set is the unit circle with $\lambda_1(S)=1$; the theorem predicts $\rho_{\mu_\epsilon}\ge C_P$ for all small $\epsilon$. More decisively, modify this potential to create one saddle point on the circle's complement while preserving the local PL condition; if the measured Poincaré constant drops to an exponentially small value as $\epsilon\to 0$, the no-saddle assumption is essential.

Watch

Extended reading notes

Core claim

The paper's central claim is Theorem 4: for a Log-PL$^\circ$ Gibbs measure $\mu_\epsilon \propto \exp(-V/\epsilon)$ with a non-singleton optimal set $S$, the Poincaré constant satisfies $\rho_{\mu_\epsilon} \ge C_P \lambda_1(S)$ once $\epsilon$ is below an explicit threshold. Here $S$ is a compact $C^2$ embedded submanifold without boundary, and $\lambda_1(S)>0$ is the first non-trivial eigenvalue of its Laplace–Beltrami operator. The lower bound is independent of $\epsilon$, so the paper establishes that a far-from-log-concave, possibly non-contractible landscape can still have a temperature-independent spectral gap, and hence $\tilde{O}(1/\epsilon)$ mixing. The proof is non-asymptotic: the constants $C_P$ depend on the potential's smoothness constants, the local PL constants, the second fundamental form of $S$, and the tubular-neighborhood radius, but not on $\epsilon$.

Load-bearing premise

The load-bearing premise is the no-saddle condition (Assumption 2): every critical point outside the minimizer neighborhoods is a strict local maximum, and if a saddle point exists the proof's connectedness argument and tube reduction can fail.

Editorial extensions

If this is right

  • For any potential in the claimed class with a non-singleton minimizer manifold, the Langevin SDE converges to $\mu_\epsilon$ in $\chi^2$-divergence in time $\tilde{O}(1/\epsilon)$ at all sufficiently small temperatures.
  • The same $\tilde{O}(1/\epsilon)$ rate transfers to the discrete-time Langevin Monte Carlo algorithm in Rényi divergence, via the paper's combination with existing LMC analysis.
  • The lower bound holds even though $\mu_\epsilon$ is not log-concave and the potential may have local maxima; non-contractibility of the minimizer set (for instance a circle) does not produce an exponential bottleneck.
  • The paper frames the result as a step toward establishing the stronger logarithmic Sobolev inequality for Log-PL$^\circ$ measures.
  • In the complementary singleton case, the paper notes the global-PL setting gives $\rho_{\mu_\epsilon}=\Omega(1/\epsilon)$ and therefore even faster $\tilde{O}(1)$ mixing.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The proof's two-step structure suggests that the no-saddle condition, not the local PL condition alone, is what forces uni-modality: a saddle point whose energy lies below the mountain pass could create a second basin even when PL holds locally. One testable extension is to allow saddles that are maxima in all but one direction and to check whether the Poincaré constant then acquires an extra $1/\
  • In the multi-modal setting, the same tube argument could be run on each basin of attraction with reflecting boundary conditions, making the per-basin mixing time $\tilde{O}(1/\epsilon)$ before the exponential metastability time; the paper notes the partitioning idea but does not develop it.
  • The eigenvalue $\lambda_1(S)$ may serve as a practical landscape diagnostic: wide, flat minimizer manifolds have small $\lambda_1(S)$, so the bound predicts slow sub-exponential mixing even in the absence of energy barriers.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper studies the Poincaré constant of low-temperature Gibbs measures whose potential satisfies a local Polyak-Łojasiewicz inequality and whose set of global minima is a non-singleton compact manifold. Under a no-saddle assumption on all critical points outside a neighborhood of the minimizers, the authors prove that the set of minima is connected, that it is a C² embedding submanifold without boundary, and that, for sufficiently small temperature, the Poincaré constant is bounded below by a temperature-independent multiple of the first nonzero Laplace-Beltrami eigenvalue of the manifold. The proof first reduces the global Poincaré inequality to a Neumann eigenvalue problem on a thin tube around the manifold via a Lyapunov argument, then derives a stability estimate comparing the tube's Neumann eigenvalue to the manifold's spectral gap, and finally combines the two steps. As a consequence, the Langevin dynamics converges in chi-square divergence at rate O-tilde(1/epsilon).

Significance. If the main theorem is correct, the paper provides a nontrivial extension of Poincaré inequality results beyond log-concave and strongly convex settings to potentials whose minimizers form a manifold. The connection between the Poincaré constant of the Gibbs measure and the spectral gap of the Laplace-Beltrami operator on the optimal set is conceptually appealing and could serve as a template for further work. The proof is driven by well-known tools (Mountain Pass theorem, tubular neighborhood theorem, tensorization of Poincaré inequalities, Lyapunov conditions), and the final bound is expressed through the intrinsic geometric quantity lambda_1(S), with no fitted constants. The machine-checkable structure of the arguments is not present, but the reliance on standard external theorems makes the main line verifiable. The main limitation is that the central no-saddle assumption (Assumption 2) is strong; the paper's stated motivation from neural-network landscapes, which typically have saddle points, is therefore broader than the actual theorem supports.

major comments (5)
  1. [Assumption 2; Proposition 3] Assumption 2 is load-bearing and restricts the result to a no-saddle class. The Mountain Pass argument in Proposition 3 (Appendix C.1) uses Assumption 2 to rule out 'global mountain passing points', and Lemma 1 uses it to isolate local maxima as strict. If Assumption 2 is violated, the connectivity of the optimal set can fail: a potential with a circle of global minima, a higher local minimum, and a barrier between them still satisfies Assumptions 1, 3 and 4, but the Gibbs measure has two metastable wells and its Poincaré constant decays exponentially in 1/epsilon. The paper should state clearly that the theorem applies only to this no-saddle class, and the introduction's claims about relevance to over-parameterized neural networks should be tempered accordingly. This is a scope issue, not an internal inconsistency.
  2. [Proposition 4 and Proposition 6] There is an inconsistency between the product Poincaré constant and the claimed tube eigenvalue bound. Proposition 4 (Bakry et al.) states that the product of two spaces with Poincaré constants C1 and C2 has Poincaré constant at least max{C1,C2}, whereas Definition 1 and the appendix proof of Proposition 6 use the convention that the Poincaré constant is the reciprocal of the best constant in the variance inequality. The min-max calculation in the proof of Proposition 6 (Appendix F.3) computes lambda_1(S x B(epsilon)) = min{lambda_1(S), lambda_1(B)}, which is consistent with a max convention for the constants. The text should resolve this notational mismatch explicitly, because the direction of the inequality in Proposition 6 depends on which convention is used.
  3. [Theorem 2 and Lemma 4] The step from the truncated Gibbs measure to the Neumann eigenvalue uses the Holley-Stroock perturbation principle (Proposition 2) with the potential difference V - tilde V on U. Lemma 4 states that exp{C_bar} rho_{epsilon,U} >= rho_U = lambda_1^n(U), but the proof is only sketched. Since U has diameter of order sqrt(epsilon) and V is C² on a neighborhood of S, V varies by O(epsilon) on U, so the oscillation term is O(1); but the constant C_bar = 4LC should be derived explicitly. As written, the dependence of the final constant on L, C, and the geometry of S is not fully quantified. This is a presentation issue in a chain of estimates whose main qualitative conclusion does not depend on the exact constants, but the derivation should be spelled out for the non-asymptotic claim to be complete.
  4. [Lemma 6 and Proposition 6, gradient formula] The gradient transformation in Lemma 6 appears to have a block-diagonal inconsistency. In the displayed equation, the matrix on the right is written as [I_k + sum r_l G_tilde(l), 0; 0, I_{d-k}], but the text immediately below the equation typesets the same matrix with the blocks in a different order. More importantly, the inverse of the matrix (I_k + sum r_l G_tilde(l)) should appear in the expression for |nabla_y phi|² in the proof of Proposition 6; the displayed formula in Appendix F.3 includes (I + sum r_l G_tilde)^{-1} on both sides, which is correct, but the presentation in Lemma 6 should be corrected to match. This is a notational and typesetting issue that does not affect the final stability bound, but it should be fixed for the reader to verify the argument.
  5. [Assumption 3, eq. (13)] Assumption 3 imposes an exponential error bound |nabla V(x)| >= nu e^{b dist(x,S)} outside a compact set. This is stronger than the coercivity and Assumption 3' that precede it. The claim in Remark 3 that eq. (13) implies Assumption 3' is correct, but the reverse is not true, and the text uses the stronger assumption throughout without noting that the exponential growth rate b enters the constant C and the allowable range of epsilon in Lemma 3. The dependence on b is not further discussed; the authors should state whether the final epsilon-regime degrades as b becomes small.
minor comments (5)
  1. [Abstract and Section 1] The phrase 'local Polyak-Lojasiewicz (PL) inequality' could be confused with the standard global PL condition; the paper should clarify in a footnote that the local condition is used only in neighborhoods of the local minima.
  2. [Section 2.5, Remark 2] The tractrix example is interesting but the displayed curvature formula has a typo: the numerator should be |x'(t)y''(t)-x''(t)y'(t)|, not |x''(t)y'(t)-x''(t)y'(t)|.
  3. [Section 4.2, Proposition 5] The statement 'the PI constant rho_{mu_B} >= 1/(C_tilde epsilon)' uses a lower bound on rho, but with the convention in Definition 1, this means the Dirichlet-to-variance ratio is at least 1/(C_tilde epsilon). The wording 'PI constant' should be aligned with Definition 1 to avoid the reversal that appears in Proposition 4.
  4. [Appendix F.3, proof of Proposition 6] The notation (g)^{-1} in the gradient energy formula is ambiguous: it should be the inverse of the metric tensor g_S, and the expression should be written with indices to make clear that it is a contraction, not a matrix inverse in the ambient coordinates.
  5. [Throughout] There are several typos, including 'quantative' for 'quantitative', 'Neumman' for 'Neumann', 'Rebjock-Boumal' spelled inconsistently, and equation (8) written with a missing factor in the Lyapunov inequality. These do not affect the mathematics but should be corrected in a revision.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the Poincaré bound derives from external Lyapunov, perturbation, and tube-stability theorems, not from its own conclusion.

full rationale

The paper's central claim (Theorem 4) is not circular. The asserted lower bound ρ_{μ_ε} ≥ C_P λ_1(S) expresses the Poincaré constant through the intrinsic first nonzero Laplace-Beltrami eigenvalue of the optimal set S, which is not used to define the measure, the Lyapunov function, or any fitted parameter. The derivation chain is transparent: (i) the Menz-Schlichting Lyapunov criterion (external, their Theorem 3.8) reduces PI on R^d to a Lyapunov drift bound plus PI on the truncated tube U = S_{√(Cε)}; (ii) the Holley-Stroock perturbation principle (external) controls the density variation inside U up to an ε-independent exponential factor; (iii) the tubular neighborhood theorem (external, Guillemin), the Weyl tube formula, and the Bakry et al. tensorization property (external) bound the Neumann eigenvalue of U by λ_1(S) up to O(√ε) errors; (iv) Assumptions 1-4 are used only to supply the ingredients for these external theorems. The paper's own new step—that the optimal set is a compact smooth submanifold—is proved via external results of Katriel and Rebjock-Boumal, not via the conclusion. The few self-references (Li et al. 2024, Yang et al. 2020) are background citations for SGD modeling and PL optimization, and are not load-bearing for the main theorem. Assumption 2 is indeed strong and scope-limiting: the no-saddle condition is essential for the uni-modality and Lyapunov arguments, so the result is conditional on that assumption. But a restrictive or debatable assumption is a correctness/scope concern, not circularity, because the theorem is explicitly stated under that assumption. No fitted value is renamed as a prediction and no load-bearing step reduces to the paper's own earlier work. Hence the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

The central claim rests on four domain assumptions (local PL, no saddles, tail error bound, bounded second fundamental form) and on standard external theorems (Mountain Pass, tubular neighborhood, Rebjock-Boumal). No free parameters are fitted and no new entities are invented.

assumptions (7)
  • domain assumption Assumption 1: local Polyak-Lojasiewicz inequality holds in neighborhoods of each connected component of the minima set.
    Defines the Log-PL^circ measure class; used throughout, especially for quadratic growth near S.
  • domain assumption Assumption 2: every critical point outside N(S) has negative definite Hessian, so no saddle points exist.
    Central to the Mountain Pass proof of uni-modality in Proposition 3 and to Lemma 1.
  • domain assumption Assumption 3: tail error bound and polynomial growth of Delta V beyond a compact set.
    Needed for the Lyapunov function verification in Lemma 3.
  • domain assumption Assumption 4: the second fundamental form of S is bounded.
    Used in the spectral stability estimates of Proposition 6 via the tube change-of-variables formula.
  • standard math Katriel's Mountain Pass Theorem.
    Invoked in Proposition 3 to force connectedness of local minima.
  • standard math Tubular neighborhood theorem for compact C2 embedding submanifolds.
    Used in Section 4 to define the coordinate system (u,r) for the tube around S.
  • standard math Rebjock-Boumal characterization of local PL functions as local manifolds of minimizers.
    Used in Corollary 1 and Lemma 7 to establish the global manifold structure of S.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Poincare Inequality for Local Log-Polyak-\L ojasiewicz Measures: Non-asymptotic Analysis in Low-temperature Regime." pith.science (2026). https://pith.science/paper/LY3SRTPB

@misc{pith2026250100429,
  author       = {Pith},
  title        = {Pith review of: Poincare Inequality for Local Log-Polyak-\L ojasiewicz Measures: Non-asymptotic Analysis in Low-temperature Regime},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LY3SRTPB}},
  note         = {Machine review of arXiv:2501.00429}
}
abstract

Potential functions in highly pertinent applications, such as deep learning in over-parameterized regime, are empirically observed to admit non-isolated minima. To understand the convergence behavior of stochastic dynamics in such landscapes, we propose to study the class of log-P{\L}$^\circ$ measures $\mu_\epsilon \propto \exp(-V/\epsilon)$, where the potential $V$ satisfies a local Polyak-{\L}ojasiewicz (P{\L}) inequality, and its set of local minima is provably connected. Notably, potentials in this class can exhibit local maxima and we characterize its optimal set $S$ to be a compact ${C}^2$ embedding submanifold of ${R}^d$ without boundary. The non-contractibility of $S$ distinguishes our function class from the classical convex setting topologically. Moreover, the embedding structure induces a naturally defined Laplacian-Beltrami operator on $S$, and we show that its first non-trivial eigenvalue provides an $\epsilon$-independent lower bound for the Poincar\'e constant in the Poincar\'e inequality of $\mu_\epsilon$. As a direct consequence, Langevin dynamics with such non-convex potential $V$ and diffusion coefficient $\epsilon$ converges to its equilibrium $\mu_\epsilon$ at a rate of $\tilde{O}(1/\epsilon)$, provided $\epsilon$ is sufficiently small. Here $\tilde{O}$ hides logarithmic terms.

Figures

Figures reproduced from arXiv: 2501.00429 by the authors.

Figure 1
Figure 1. The circle in (a) can be represented using two local charts (blue and green). Using the tubular neighborhood theorem, in a local region of 𝑈 (outlined with the red dashed line) we transform the uniform measure to a pair of decoupled measures on the tangent and normal directions. (a) Uniform measure 𝜇𝑈 (over (𝑥, 𝑦)) under the Cartisian coordinate (𝑥, 𝑦); (b) Uniform measure 𝜇𝑈 (over (𝑥, 𝑦)) under the local coordinate… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

67 extracted references · 51 canonical work pages

  1. [1]

    Arrhenius

    Svante. Arrhenius. On the reaction velocity of the inversion of cane sugar by acids. In Selected readings in chemical kinetics, pages 31--35. Elsevier, 1967

  2. [2]

    Diffusions hypercontractives

    Dominique Bakry and Michel \'E mery. Diffusions hypercontractives. In S \'e minaire de Probabilit \'e s XIX 1983/84: Proceedings , pages 177--206. Springer, 2006

  3. [3]

    A simple proof of the Poincaré inequality for a large class of probability measures

    Dominique Bakry, Franck Barthe, Patrick Cattiaux, and Arnaud Guillin. A simple proof of the Poincaré inequality for a large class of probability measures . Electronic Communications in Probability, 13 0 (none): 0 60 -- 66, 2008 a . doi:10.1214/ECP.v13-1352. URL https://doi.org/10.1214/ECP.v13-1352

  4. [4]

    Rate of convergence for ergodic continuous markov processes: Lyapunov versus poincar \'e

    Dominique Bakry, Patrick Cattiaux, and Arnaud Guillin. Rate of convergence for ergodic continuous markov processes: Lyapunov versus poincar \'e . Journal of Functional Analysis, 254 0 (3): 0 727--759, 2008 b

  5. [5]

    Analysis and geometry of Markov diffusion operators, volume 103

    Dominique Bakry, Ivan Gentil, Michel Ledoux, et al. Analysis and geometry of Markov diffusion operators, volume 103. Springer, 2014

  6. [6]

    High-dimensional limit theorems for sgd: Effective dynamics and critical scaling

    Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath. High-dimensional limit theorems for sgd: Effective dynamics and critical scaling. Advances in Neural Information Processing Systems, 35: 0 25349--25362, 2022

  7. [7]

    B \'e rard

    Pierre H. B \'e rard. Spectral Geometry: Direct and Inverse Problems. 1986. URL https://api.semanticscholar.org/CorpusID:117992829

  8. [8]

    An Introduction to Optimization on Smooth Manifolds

    Nicolas Boumal. An Introduction to Optimization on Smooth Manifolds. Cambridge University Press, 2023

Show all 67 references
  1. [9]

    Metastability in reversible diffusion processes

    Anton Bovier, Michael Eckhoff, V \'e ronique Gayrard, and Markus Klein. Metastability in reversible diffusion processes. i. sharp asymptotics for capacities and exit times. J. Eur. Math. Soc.(JEMS), 6 0 (4): 0 399--424, 2004

  2. [10]

    Hitting times, functional inequalities, lyapunov conditions and uniform ergodicity

    Patrick Cattiaux and Arnaud Guillin. Hitting times, functional inequalities, lyapunov conditions and uniform ergodicity. Journal of Functional Analysis, 272 0 (6): 0 2361--2391, 2017. ISSN 0022-1236. doi:https://doi.org/10.1016/j.jfa.2016.10.003. URL https://www.sciencedirect....

  3. [11]

    A note on talagrand’s transportation inequality and logarithmic sobolev inequality

    Patrick Cattiaux, Arnaud Guillin, and Li-Ming Wu. A note on talagrand’s transportation inequality and logarithmic sobolev inequality. Probability theory and related fields, 148: 0 285--304, 2010

  4. [12]

    A lower bound for the smallest eigenvalue of the laplacian

    Jeff Cheeger. A lower bound for the smallest eigenvalue of the laplacian. Problems in analysis, 625 0 (195-199): 0 110, 1970

  5. [13]

    An almost constant lower bound of the isoperimetric coefficient in the kls conjecture

    Yuansi Chen. An almost constant lower bound of the isoperimetric coefficient in the kls conjecture. Geometric and Functional Analysis, 31: 0 34--61, 2020. URL https://api.semanticscholar.org/CorpusID:227209486

  6. [14]

    An almost constant lower bound of the isoperimetric coefficient in the kls conjecture

    Yuansi Chen. An almost constant lower bound of the isoperimetric coefficient in the kls conjecture. Geometric and Functional Analysis, 31: 0 34--61, 2021

  7. [15]

    The ballistic limit of the log-sobolev constant equals the Polyak- ojasiewicz constant

    Sinho Chewi and Austin J Stromme. The ballistic limit of the log-sobolev constant equals the Polyak- ojasiewicz constant. November 2024

  8. [16]

    Analysis of langevin monte carlo from poincare to log-sobolev

    Sinho Chewi, Murat A Erdogdu, Mufan Li, Ruoqi Shen, and Matthew S Zhang. Analysis of langevin monte carlo from poincare to log-sobolev. Foundations of Computational Mathematics, pages 1--51, 2024

  9. [17]

    The critical locus of overparameterized neural networks

    Y Cooper. The critical locus of overparameterized neural networks. arXiv preprint arXiv:2005.04210, 2020

  10. [18]

    The loss landscape of overparameterized neural networks

    Yaim Cooper. The loss landscape of overparameterized neural networks. arXiv preprint arXiv:1804.10200, 2018

  11. [19]

    Brian Davies

    E. Brian Davies. Spectral Theory and Differential Operators. Cambridge Studies in Advanced Mathematics. Cambridge University Press, 1995

  12. [20]

    Essentially no barriers in neural network energy landscape

    Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht. Essentially no barriers in neural network energy landscape. In International conference on machine learning, pages 1309--1318. PMLR, 2018

  13. [21]

    Thin shell implies spectral gap up to polylog via a stochastic localization scheme

    Ronen Eldan. Thin shell implies spectral gap up to polylog via a stochastic localization scheme. Geometric and Functional Analysis, 23: 0 532 -- 569, 2013. URL https://api.semanticscholar.org/CorpusID:253637768

  14. [22]

    On the convergence of langevin monte carlo: The interplay between tail growth and smoothness

    Murat A Erdogdu and Rasa Hosseinzadeh. On the convergence of langevin monte carlo: The interplay between tail growth and smoothness. In Conference on Learning Theory, pages 1776--1822. PMLR, 2021

  15. [23]

    Lawrence C. Evans. Partial differential equations. American Mathematical Society, 2010

  16. [24]

    The activated complex in chemical reactions

    Henry Eyring. The activated complex in chemical reactions. The Journal of Chemical Physics, 3 0 (2): 0 107--115, 1935

  17. [25]

    Convergence rates for the stochastic gradient descent method for non-convex objective functions

    Benjamin Fehrman, Benjamin Gess, and Arnulf Jentzen. Convergence rates for the stochastic gradient descent method for non-convex objective functions. Journal of Machine Learning Research, 21 0 (136): 0 1--48, 2020

  18. [26]

    Topology and geometry of half-rectified network optimization

    C Daniel Freeman and Joan Bruna. Topology and geometry of half-rectified network optimization. In 5th International Conference on Learning Representations, ICLR 2017, 2017

  19. [27]

    Random Perturbations of Dynamical Systems, volume 260

    Mark I Freidlin and Alexander D Wentzell. Random Perturbations of Dynamical Systems, volume 260. Springer Science & Business Media, 2012

  20. [28]

    Loss surfaces, mode connectivity, and fast ensembling of dnns

    Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson. Loss surfaces, mode connectivity, and fast ensembling of dnns. Advances in neural information processing systems, 31, 2018

  21. [29]

    Metastability in reversible diffusion processes ii: Precise asymptotics for small eigenvalues

    V \'e ronique Gayrard, Anton Bovier, and Markus Klein. Metastability in reversible diffusion processes ii: Precise asymptotics for small eigenvalues. Journal of the European Mathematical Society, 7 0 (1): 0 69--99, 2005

  22. [30]

    Differential Topology

    Alan Pollack Victor Guillemin. Differential Topology. 1974

  23. [31]

    Nonlinear analysis on manifolds: Sobolev spaces and inequalities

    Emmanuel Hebey. Nonlinear analysis on manifolds: Sobolev spaces and inequalities. 1999. URL https://api.semanticscholar.org/CorpusID:118747316

  24. [32]

    Isoperimetric problems for convex bodies and a localization lemma

    Ravi Kannan, L \'a szl \'o Mikl \'o s Lov \'a sz, and Mikl \'o s Simonovits. Isoperimetric problems for convex bodies and a localization lemma. Discrete & Computational Geometry, 13: 0 541--559, 1995. URL https://api.semanticscholar.org/CorpusID:14881695

  25. [33]

    Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition

    Hamed Karimi, Julie Nutini, and Mark Schmidt. Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2016, Riva del Garda, Italy, September 19...

  26. [34]

    Mountain pass theorems and global homeomorphism theorems

    Guy Katriel. Mountain pass theorems and global homeomorphism theorems. In Annales de l'Institut Henri Poincar \'e C, Analyse non lin \'e aire , volume 11, pages 189--209. Elsevier, 1994

  27. [35]

    Elimination of all bad local minima in deep learning

    Kenji Kawaguchi and Leslie Kaelbling. Elimination of all bad local minima in deep learning. In International Conference on Artificial Intelligence and Statistics, pages 853--863. PMLR, 2020

  28. [36]

    Brownian motion in a field of force and the diffusion model of chemical reactions

    Hendrik Anthony Kramers. Brownian motion in a field of force and the diffusion model of chemical reactions. physica, 7 0 (4): 0 284--304, 1940

  29. [37]

    Explaining landscape connectivity of low-cost solutions for multilayer nets

    Rohith Kuditipudi, Xiang Wang, Holden Lee, Yi Zhang, Zhiyuan Li, Wei Hu, Rong Ge, and Sanjeev Arora. Explaining landscape connectivity of low-cost solutions for multilayer nets. Advances in neural information processing systems, 32, 2019

  30. [38]

    John M. Lee. Introduction to Smooth Manifolds. 2020

  31. [39]

    Yin Tat Lee and Santosh S. Vempala. The kannan-lov\'asz-simonovits conjecture, 2018. URL https://arxiv.org/abs/1807.03465

  32. [40]

    Eldan's stochastic localization and the kls conjecture: Isoperimetry, concentration and mixing

    Yin Tat Lee and Santosh S Vempala. Eldan's stochastic localization and the kls conjecture: Isoperimetry, concentration and mixing. Annals of Mathematics, 199 0 (3): 0 1043--1092, 2024

  33. [41]

    The effect of smooth parametrizations on nonconvex optimization landscapes

    Eitan Levin, Joe Kileel, and Nicolas Boumal. The effect of smooth parametrizations on nonconvex optimization landscapes. Mathematical Programming, pages 1--49, 2024

  34. [42]

    Stochastic modified equations and adaptive stochastic gradient algorithms

    Qianxiao Li, Cheng Tai, and E Weinan. Stochastic modified equations and adaptive stochastic gradient algorithms. In International Conference on Machine Learning, pages 2101--2110. PMLR, 2017

  35. [43]

    A hessian-aware stochastic differential equation for modelling sgd

    Xiang Li, Zebang Shen, Liang Zhang, and Niao He. A hessian-aware stochastic differential equation for modelling sgd. arXiv preprint arXiv:2405.18373, 2024

  36. [44]

    What happens after sgd reaches zero loss?-a mathematical framework

    Zhiyuan Li, Tianhao Wang, and Sanjeev Arora. What happens after sgd reaches zero loss?-a mathematical framework. In 10th International Conference on Learning Representations, ICLR 2022, 2022

  37. [45]

    Adding one neuron can eliminate all bad local minima

    Shiyu Liang, Ruoyu Sun, and Jason D Lee. Adding one neuron can eliminate all bad local minima. Advances in Neural Information Processing Systems, 31, 2018 a

  38. [46]

    Understanding the loss surface of neural networks for binary classification

    Shiyu Liang, Ruoyu Sun, Yixuan Li, and Rayadurgam Srikant. Understanding the loss surface of neural networks for binary classification. In International Conference on Machine Learning, pages 2835--2843. PMLR, 2018 b

  39. [47]

    Exploring neural network landscapes: Star-shaped and geodesic connectivity

    Zhanran Lin, Puheng Li, and Lei Wu. Exploring neural network landscapes: Star-shaped and geodesic connectivity. arXiv preprint arXiv:2404.06391, 2024

  40. [48]

    Loss landscapes and optimization in over-parameterized non-linear systems and neural networks

    Chaoyue Liu, Libin Zhu, and Mikhail Belkin. Loss landscapes and optimization in over-parameterized non-linear systems and neural networks. Applied and Computational Harmonic Analysis, 59: 0 85--116, 2022

  41. [49]

    A topological property of real analytic subsets

    Stanislaw Lojasiewicz. A topological property of real analytic subsets. Coll. du CNRS, Les \'e quations aux d \'e riv \'e es partielles , 117 0 (87-89): 0 2, 1963

  42. [50]

    Poincaré and logarithmic Sobolev inequalities by decomposition of the energy landscape

    Georg Menz and Andr \'e Schlichting. Poincaré and logarithmic Sobolev inequalities by decomposition of the energy landscape . The Annals of Probability, 42 0 (5): 0 1809 -- 1884, 2014. doi:10.1214/14-AOP908. URL https://doi.org/10.1214/14-AOP908

  43. [51]

    Stasheff

    John Milnor and James D. Stasheff. Characteristic Classes. (AM-76), Volume 76. Princeton University Press, Princeton, 1974. ISBN 9781400881826. doi:doi:10.1515/9781400881826. URL https://doi.org/10.1515/9781400881826

  44. [52]

    On connected sublevel sets in deep learning

    Quynh Nguyen. On connected sublevel sets in deep learning. In International conference on machine learning, pages 4790--4799. PMLR, 2019

  45. [53]

    Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks

    Samet Oymak and Mahdi Soltanolkotabi. Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks. IEEE Journal on Selected Areas in Information Theory, 1 0 (1): 0 84--105, 2020

  46. [54]

    Homogenization of sgd in high-dimensions: Exact dynamics and generalization properties

    Courtney Paquette, Elliot Paquette, Ben Adlam, and Jeffrey Pennington. Homogenization of sgd in high-dimensions: Exact dynamics and generalization properties. arXiv preprint arXiv:2205.07069, 2022

  47. [55]

    Gradient methods for minimizing functionals

    Boris Teodorovich Polyak. Gradient methods for minimizing functionals. Zhurnal vychislitel'noi matematiki i matematicheskoi fiziki, 3 0 (4): 0 643--653, 1963

  48. [56]

    Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis

    Maxim Raginsky, Alexander Rakhlin, and Matus Telgarsky. Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis. In Conference on Learning Theory, pages 1674--1703. PMLR, 2017

  49. [57]

    Fast convergence to non-isolated minima: four equivalent conditions for c 2 functions

    Quentin Rebjock and Nicolas Boumal. Fast convergence to non-isolated minima: four equivalent conditions for c 2 functions. Mathematical Programming, pages 1--49, 2024

  50. [58]

    On the quality of the initial basin in overspecified neural networks

    Itay Safran and Ohad Shamir. On the quality of the initial basin in overspecified neural networks. In International Conference on Machine Learning, pages 774--782. PMLR, 2016

  51. [59]

    Eigenvalues of the hessian in deep learning: Singularity and beyond

    Levent Sagun, Leon Bottou, and Yann LeCun. Eigenvalues of the hessian in deep learning: Singularity and beyond. arXiv preprint arXiv:1611.07476, 2016

  52. [60]

    Empirical analysis of the hessian of over-parametrized neural networks

    Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou. Empirical analysis of the hessian of over-parametrized neural networks. arXiv preprint arXiv:1706.04454, 2017

  53. [61]

    L.W. Tu. An Introduction to Manifolds. Universitext. Springer New York, 2010. ISBN 9781441973993. URL https://books.google.com.hk/books?id=br1KngEACAAJ

  54. [62]

    Spurious valleys in two-layer neural network optimization landscapes

    Luca Venturi, Afonso S Bandeira, and Joan Bruna. Spurious valleys in two-layer neural network optimization landscapes. arXiv preprint arXiv:1802.06384, 2018

  55. [63]

    On the volume of tubes

    Hermann Weyl. On the volume of tubes. American Journal of Mathematics, 61: 0 461, 1939. URL https://api.semanticscholar.org/CorpusID:124362885

  56. [64]

    Sampling as optimization in the space of measures: The langevin dynamics as a composite optimization problem

    Andre Wibisono. Sampling as optimization in the space of measures: The langevin dynamics as a composite optimization problem. In Conference on Learning Theory, pages 2093--3027. PMLR, 2018

  57. [65]

    Stochastic gradient descent with noise of machine learning type part ii: Continuous time analysis

    Stephan Wojtowytsch. Stochastic gradient descent with noise of machine learning type part ii: Continuous time analysis. Journal of Nonlinear Science, 34 0 (1): 0 16, 2024

  58. [66]

    Global convergence and variance reduction for a class of nonconvex-nonconcave minimax problems

    Junchi Yang, Negar Kiyavash, and Niao He. Global convergence and variance reduction for a class of nonconvex-nonconcave minimax problems. Advances in Neural Information Processing Systems, 33: 0 1153--1165, 2020

  59. [67]

    A hitting time analysis of stochastic gradient langevin dynamics

    Yuchen Zhang, Percy Liang, and Moses Charikar. A hitting time analysis of stochastic gradient langevin dynamics. In Conference on Learning Theory, pages 1980--2022. PMLR, 2017

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.