REVIEW 5 major objections 5 minor 67 references
Poincare Inequality for Local Log-Polyak-\L ojasiewicz Measures: Non-asymptotic Analysis in Low-temperature Regime
T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read At low temperatures, the Poincaré constant of a Gibbs measure is bounded below by the first Laplace–Beltrami eigenvalue of its minimizer manifold.
desk verdict New and plausible reduction of the low-temperature Poincare constant to the Laplace-Beltrami eigenvalue of the minimizer manifold, but the no-saddle assumption is doing heavy lifting and a few technical inconsistencies need fixing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Log-PL$^\circ$ measure, defined by a potential that satisfies the local Polyak–Łojasiewicz inequality $|\nabla V|^2 \ge \nu (V - \min V)$ near each connected component of its minimizer set, together with the no-saddle condition that every critical point outside those neighborhoods is a strict local maximum. These assumptions force all local minima to lie in one connected component, and force the global minimizer set $S$ to be a compact $C^2$ embedded submanifold without boundary. The proof's machinery is a two-step reduction: first a Lyapunov-function criterion (following Menz and Schlichting) reduces the Poincaré constant of $\mu_\epsilon$ to the Neumann eigenvalue $\lambda^n_1(U)$ of the Laplacian on the thin tube $U = S_{\sqrt{C\epsilon}}$; then, because a thin tube around an embedded submanifold is a tubular neighborhood, a stability analysis using the Weyl tube formula, tensorization of the Poincaré inequality, and boundedness of the second fundamental form shows $\lambda^n_1(U)$ differs from $\lambda_1(S)$ only by $O(\sqrt{\epsilon})$. Assembling the two steps gives the temperature-independent constant.
What would settle it
Compute or simulate the spectral gap of Langevin dynamics for the paper's own example $V(x)=\|x\|^3/3-\|x\|^2/2$ in $\mathbb{R}^2$, whose minimizer set is the unit circle with $\lambda_1(S)=1$; the theorem predicts $\rho_{\mu_\epsilon}\ge C_P$ for all small $\epsilon$. More decisively, modify this potential to create one saddle point on the circle's complement while preserving the local PL condition; if the measured Poincaré constant drops to an exponentially small value as $\epsilon\to 0$, the no-saddle assumption is essential.
Extended reading notes
Core claim
The paper's central claim is Theorem 4: for a Log-PL$^\circ$ Gibbs measure $\mu_\epsilon \propto \exp(-V/\epsilon)$ with a non-singleton optimal set $S$, the Poincaré constant satisfies $\rho_{\mu_\epsilon} \ge C_P \lambda_1(S)$ once $\epsilon$ is below an explicit threshold. Here $S$ is a compact $C^2$ embedded submanifold without boundary, and $\lambda_1(S)>0$ is the first non-trivial eigenvalue of its Laplace–Beltrami operator. The lower bound is independent of $\epsilon$, so the paper establishes that a far-from-log-concave, possibly non-contractible landscape can still have a temperature-independent spectral gap, and hence $\tilde{O}(1/\epsilon)$ mixing. The proof is non-asymptotic: the constants $C_P$ depend on the potential's smoothness constants, the local PL constants, the second fundamental form of $S$, and the tubular-neighborhood radius, but not on $\epsilon$.
Load-bearing premise
The load-bearing premise is the no-saddle condition (Assumption 2): every critical point outside the minimizer neighborhoods is a strict local maximum, and if a saddle point exists the proof's connectedness argument and tube reduction can fail.
Editorial extensions
If this is right
- For any potential in the claimed class with a non-singleton minimizer manifold, the Langevin SDE converges to $\mu_\epsilon$ in $\chi^2$-divergence in time $\tilde{O}(1/\epsilon)$ at all sufficiently small temperatures.
- The same $\tilde{O}(1/\epsilon)$ rate transfers to the discrete-time Langevin Monte Carlo algorithm in Rényi divergence, via the paper's combination with existing LMC analysis.
- The lower bound holds even though $\mu_\epsilon$ is not log-concave and the potential may have local maxima; non-contractibility of the minimizer set (for instance a circle) does not produce an exponential bottleneck.
- The paper frames the result as a step toward establishing the stronger logarithmic Sobolev inequality for Log-PL$^\circ$ measures.
- In the complementary singleton case, the paper notes the global-PL setting gives $\rho_{\mu_\epsilon}=\Omega(1/\epsilon)$ and therefore even faster $\tilde{O}(1)$ mixing.
Reading between the lines
- The proof's two-step structure suggests that the no-saddle condition, not the local PL condition alone, is what forces uni-modality: a saddle point whose energy lies below the mountain pass could create a second basin even when PL holds locally. One testable extension is to allow saddles that are maxima in all but one direction and to check whether the Poincaré constant then acquires an extra $1/\
- In the multi-modal setting, the same tube argument could be run on each basin of attraction with reflecting boundary conditions, making the per-basin mixing time $\tilde{O}(1/\epsilon)$ before the exponential metastability time; the paper notes the partitioning idea but does not develop it.
- The eigenvalue $\lambda_1(S)$ may serve as a practical landscape diagnostic: wide, flat minimizer manifolds have small $\lambda_1(S)$, so the bound predicts slow sub-exponential mixing even in the absence of energy barriers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the Poincaré constant of low-temperature Gibbs measures whose potential satisfies a local Polyak-Łojasiewicz inequality and whose set of global minima is a non-singleton compact manifold. Under a no-saddle assumption on all critical points outside a neighborhood of the minimizers, the authors prove that the set of minima is connected, that it is a C² embedding submanifold without boundary, and that, for sufficiently small temperature, the Poincaré constant is bounded below by a temperature-independent multiple of the first nonzero Laplace-Beltrami eigenvalue of the manifold. The proof first reduces the global Poincaré inequality to a Neumann eigenvalue problem on a thin tube around the manifold via a Lyapunov argument, then derives a stability estimate comparing the tube's Neumann eigenvalue to the manifold's spectral gap, and finally combines the two steps. As a consequence, the Langevin dynamics converges in chi-square divergence at rate O-tilde(1/epsilon).
Significance. If the main theorem is correct, the paper provides a nontrivial extension of Poincaré inequality results beyond log-concave and strongly convex settings to potentials whose minimizers form a manifold. The connection between the Poincaré constant of the Gibbs measure and the spectral gap of the Laplace-Beltrami operator on the optimal set is conceptually appealing and could serve as a template for further work. The proof is driven by well-known tools (Mountain Pass theorem, tubular neighborhood theorem, tensorization of Poincaré inequalities, Lyapunov conditions), and the final bound is expressed through the intrinsic geometric quantity lambda_1(S), with no fitted constants. The machine-checkable structure of the arguments is not present, but the reliance on standard external theorems makes the main line verifiable. The main limitation is that the central no-saddle assumption (Assumption 2) is strong; the paper's stated motivation from neural-network landscapes, which typically have saddle points, is therefore broader than the actual theorem supports.
major comments (5)
- [Assumption 2; Proposition 3] Assumption 2 is load-bearing and restricts the result to a no-saddle class. The Mountain Pass argument in Proposition 3 (Appendix C.1) uses Assumption 2 to rule out 'global mountain passing points', and Lemma 1 uses it to isolate local maxima as strict. If Assumption 2 is violated, the connectivity of the optimal set can fail: a potential with a circle of global minima, a higher local minimum, and a barrier between them still satisfies Assumptions 1, 3 and 4, but the Gibbs measure has two metastable wells and its Poincaré constant decays exponentially in 1/epsilon. The paper should state clearly that the theorem applies only to this no-saddle class, and the introduction's claims about relevance to over-parameterized neural networks should be tempered accordingly. This is a scope issue, not an internal inconsistency.
- [Proposition 4 and Proposition 6] There is an inconsistency between the product Poincaré constant and the claimed tube eigenvalue bound. Proposition 4 (Bakry et al.) states that the product of two spaces with Poincaré constants C1 and C2 has Poincaré constant at least max{C1,C2}, whereas Definition 1 and the appendix proof of Proposition 6 use the convention that the Poincaré constant is the reciprocal of the best constant in the variance inequality. The min-max calculation in the proof of Proposition 6 (Appendix F.3) computes lambda_1(S x B(epsilon)) = min{lambda_1(S), lambda_1(B)}, which is consistent with a max convention for the constants. The text should resolve this notational mismatch explicitly, because the direction of the inequality in Proposition 6 depends on which convention is used.
- [Theorem 2 and Lemma 4] The step from the truncated Gibbs measure to the Neumann eigenvalue uses the Holley-Stroock perturbation principle (Proposition 2) with the potential difference V - tilde V on U. Lemma 4 states that exp{C_bar} rho_{epsilon,U} >= rho_U = lambda_1^n(U), but the proof is only sketched. Since U has diameter of order sqrt(epsilon) and V is C² on a neighborhood of S, V varies by O(epsilon) on U, so the oscillation term is O(1); but the constant C_bar = 4LC should be derived explicitly. As written, the dependence of the final constant on L, C, and the geometry of S is not fully quantified. This is a presentation issue in a chain of estimates whose main qualitative conclusion does not depend on the exact constants, but the derivation should be spelled out for the non-asymptotic claim to be complete.
- [Lemma 6 and Proposition 6, gradient formula] The gradient transformation in Lemma 6 appears to have a block-diagonal inconsistency. In the displayed equation, the matrix on the right is written as [I_k + sum r_l G_tilde(l), 0; 0, I_{d-k}], but the text immediately below the equation typesets the same matrix with the blocks in a different order. More importantly, the inverse of the matrix (I_k + sum r_l G_tilde(l)) should appear in the expression for |nabla_y phi|² in the proof of Proposition 6; the displayed formula in Appendix F.3 includes (I + sum r_l G_tilde)^{-1} on both sides, which is correct, but the presentation in Lemma 6 should be corrected to match. This is a notational and typesetting issue that does not affect the final stability bound, but it should be fixed for the reader to verify the argument.
- [Assumption 3, eq. (13)] Assumption 3 imposes an exponential error bound |nabla V(x)| >= nu e^{b dist(x,S)} outside a compact set. This is stronger than the coercivity and Assumption 3' that precede it. The claim in Remark 3 that eq. (13) implies Assumption 3' is correct, but the reverse is not true, and the text uses the stronger assumption throughout without noting that the exponential growth rate b enters the constant C and the allowable range of epsilon in Lemma 3. The dependence on b is not further discussed; the authors should state whether the final epsilon-regime degrades as b becomes small.
minor comments (5)
- [Abstract and Section 1] The phrase 'local Polyak-Lojasiewicz (PL) inequality' could be confused with the standard global PL condition; the paper should clarify in a footnote that the local condition is used only in neighborhoods of the local minima.
- [Section 2.5, Remark 2] The tractrix example is interesting but the displayed curvature formula has a typo: the numerator should be |x'(t)y''(t)-x''(t)y'(t)|, not |x''(t)y'(t)-x''(t)y'(t)|.
- [Section 4.2, Proposition 5] The statement 'the PI constant rho_{mu_B} >= 1/(C_tilde epsilon)' uses a lower bound on rho, but with the convention in Definition 1, this means the Dirichlet-to-variance ratio is at least 1/(C_tilde epsilon). The wording 'PI constant' should be aligned with Definition 1 to avoid the reversal that appears in Proposition 4.
- [Appendix F.3, proof of Proposition 6] The notation (g)^{-1} in the gradient energy formula is ambiguous: it should be the inverse of the metric tensor g_S, and the expression should be written with indices to make clear that it is a contraction, not a matrix inverse in the ambient coordinates.
- [Throughout] There are several typos, including 'quantative' for 'quantitative', 'Neumman' for 'Neumann', 'Rebjock-Boumal' spelled inconsistently, and equation (8) written with a missing factor in the Lyapunov inequality. These do not affect the mathematics but should be corrected in a revision.
Circularity Check
No significant circularity: the Poincaré bound derives from external Lyapunov, perturbation, and tube-stability theorems, not from its own conclusion.
full rationale
The paper's central claim (Theorem 4) is not circular. The asserted lower bound ρ_{μ_ε} ≥ C_P λ_1(S) expresses the Poincaré constant through the intrinsic first nonzero Laplace-Beltrami eigenvalue of the optimal set S, which is not used to define the measure, the Lyapunov function, or any fitted parameter. The derivation chain is transparent: (i) the Menz-Schlichting Lyapunov criterion (external, their Theorem 3.8) reduces PI on R^d to a Lyapunov drift bound plus PI on the truncated tube U = S_{√(Cε)}; (ii) the Holley-Stroock perturbation principle (external) controls the density variation inside U up to an ε-independent exponential factor; (iii) the tubular neighborhood theorem (external, Guillemin), the Weyl tube formula, and the Bakry et al. tensorization property (external) bound the Neumann eigenvalue of U by λ_1(S) up to O(√ε) errors; (iv) Assumptions 1-4 are used only to supply the ingredients for these external theorems. The paper's own new step—that the optimal set is a compact smooth submanifold—is proved via external results of Katriel and Rebjock-Boumal, not via the conclusion. The few self-references (Li et al. 2024, Yang et al. 2020) are background citations for SGD modeling and PL optimization, and are not load-bearing for the main theorem. Assumption 2 is indeed strong and scope-limiting: the no-saddle condition is essential for the uni-modality and Lyapunov arguments, so the result is conditional on that assumption. But a restrictive or debatable assumption is a correctness/scope concern, not circularity, because the theorem is explicitly stated under that assumption. No fitted value is renamed as a prediction and no load-bearing step reduces to the paper's own earlier work. Hence the appropriate score is 0.
Assumptions & free parameters
assumptions (7)
- domain assumption Assumption 1: local Polyak-Lojasiewicz inequality holds in neighborhoods of each connected component of the minima set.
- domain assumption Assumption 2: every critical point outside N(S) has negative definite Hessian, so no saddle points exist.
- domain assumption Assumption 3: tail error bound and polynomial growth of Delta V beyond a compact set.
- domain assumption Assumption 4: the second fundamental form of S is bounded.
- standard math Katriel's Mountain Pass Theorem.
- standard math Tubular neighborhood theorem for compact C2 embedding submanifolds.
- standard math Rebjock-Boumal characterization of local PL functions as local manifolds of minimizers.
Cite this review
Pith. "Pith review of Poincare Inequality for Local Log-Polyak-\L ojasiewicz Measures: Non-asymptotic Analysis in Low-temperature Regime." pith.science (2026). https://pith.science/paper/LY3SRTPB
@misc{pith2026250100429,
author = {Pith},
title = {Pith review of: Poincare Inequality for Local Log-Polyak-\L ojasiewicz Measures: Non-asymptotic Analysis in Low-temperature Regime},
year = {2026},
howpublished = {\url{https://pith.science/paper/LY3SRTPB}},
note = {Machine review of arXiv:2501.00429}
}
abstract
Potential functions in highly pertinent applications, such as deep learning in over-parameterized regime, are empirically observed to admit non-isolated minima. To understand the convergence behavior of stochastic dynamics in such landscapes, we propose to study the class of log-P{\L}$^\circ$ measures $\mu_\epsilon \propto \exp(-V/\epsilon)$, where the potential $V$ satisfies a local Polyak-{\L}ojasiewicz (P{\L}) inequality, and its set of local minima is provably connected. Notably, potentials in this class can exhibit local maxima and we characterize its optimal set $S$ to be a compact ${C}^2$ embedding submanifold of ${R}^d$ without boundary. The non-contractibility of $S$ distinguishes our function class from the classical convex setting topologically. Moreover, the embedding structure induces a naturally defined Laplacian-Beltrami operator on $S$, and we show that its first non-trivial eigenvalue provides an $\epsilon$-independent lower bound for the Poincar\'e constant in the Poincar\'e inequality of $\mu_\epsilon$. As a direct consequence, Langevin dynamics with such non-convex potential $V$ and diffusion coefficient $\epsilon$ converges to its equilibrium $\mu_\epsilon$ at a rate of $\tilde{O}(1/\epsilon)$, provided $\epsilon$ is sufficiently small. Here $\tilde{O}$ hides logarithmic terms.
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
Dominique Bakry and Michel \'E mery. Diffusions hypercontractives. In S \'e minaire de Probabilit \'e s XIX 1983/84: Proceedings , pages 177--206. Springer, 2006
work page 1983
-
[3]
A simple proof of the Poincaré inequality for a large class of probability measures
Dominique Bakry, Franck Barthe, Patrick Cattiaux, and Arnaud Guillin. A simple proof of the Poincaré inequality for a large class of probability measures . Electronic Communications in Probability, 13 0 (none): 0 60 -- 66, 2008 a . doi:10.1214/ECP.v13-1352. URL https://doi.org/10.1214/ECP.v13-1352
-
[4]
Rate of convergence for ergodic continuous markov processes: Lyapunov versus poincar \'e
Dominique Bakry, Patrick Cattiaux, and Arnaud Guillin. Rate of convergence for ergodic continuous markov processes: Lyapunov versus poincar \'e . Journal of Functional Analysis, 254 0 (3): 0 727--759, 2008 b
work page 2008
-
[5]
Analysis and geometry of Markov diffusion operators, volume 103
Dominique Bakry, Ivan Gentil, Michel Ledoux, et al. Analysis and geometry of Markov diffusion operators, volume 103. Springer, 2014
work page 2014
-
[6]
High-dimensional limit theorems for sgd: Effective dynamics and critical scaling
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath. High-dimensional limit theorems for sgd: Effective dynamics and critical scaling. Advances in Neural Information Processing Systems, 35: 0 25349--25362, 2022
work page 2022
-
[7]
Pierre H. B \'e rard. Spectral Geometry: Direct and Inverse Problems. 1986. URL https://api.semanticscholar.org/CorpusID:117992829
work page 1986
-
[8]
An Introduction to Optimization on Smooth Manifolds
Nicolas Boumal. An Introduction to Optimization on Smooth Manifolds. Cambridge University Press, 2023
2023
Show all 67 references
-
[9]
Metastability in reversible diffusion processes
Anton Bovier, Michael Eckhoff, V \'e ronique Gayrard, and Markus Klein. Metastability in reversible diffusion processes. i. sharp asymptotics for capacities and exit times. J. Eur. Math. Soc.(JEMS), 6 0 (4): 0 399--424, 2004
2004
-
[10]
Hitting times, functional inequalities, lyapunov conditions and uniform ergodicity
Patrick Cattiaux and Arnaud Guillin. Hitting times, functional inequalities, lyapunov conditions and uniform ergodicity. Journal of Functional Analysis, 272 0 (6): 0 2361--2391, 2017. ISSN 0022-1236. doi:https://doi.org/10.1016/j.jfa.2016.10.003. URL https://www.sciencedirect....
2017 doi
-
[11]
A note on talagrand’s transportation inequality and logarithmic sobolev inequality
Patrick Cattiaux, Arnaud Guillin, and Li-Ming Wu. A note on talagrand’s transportation inequality and logarithmic sobolev inequality. Probability theory and related fields, 148: 0 285--304, 2010
2010
-
[12]
A lower bound for the smallest eigenvalue of the laplacian
Jeff Cheeger. A lower bound for the smallest eigenvalue of the laplacian. Problems in analysis, 625 0 (195-199): 0 110, 1970
1970
-
[13]
An almost constant lower bound of the isoperimetric coefficient in the kls conjecture
Yuansi Chen. An almost constant lower bound of the isoperimetric coefficient in the kls conjecture. Geometric and Functional Analysis, 31: 0 34--61, 2020. URL https://api.semanticscholar.org/CorpusID:227209486
2020
-
[14]
An almost constant lower bound of the isoperimetric coefficient in the kls conjecture
Yuansi Chen. An almost constant lower bound of the isoperimetric coefficient in the kls conjecture. Geometric and Functional Analysis, 31: 0 34--61, 2021
2021
-
[15]
The ballistic limit of the log-sobolev constant equals the Polyak- ojasiewicz constant
Sinho Chewi and Austin J Stromme. The ballistic limit of the log-sobolev constant equals the Polyak- ojasiewicz constant. November 2024
2024
-
[16]
Analysis of langevin monte carlo from poincare to log-sobolev
Sinho Chewi, Murat A Erdogdu, Mufan Li, Ruoqi Shen, and Matthew S Zhang. Analysis of langevin monte carlo from poincare to log-sobolev. Foundations of Computational Mathematics, pages 1--51, 2024
2024
-
[17]
The critical locus of overparameterized neural networks
Y Cooper. The critical locus of overparameterized neural networks. arXiv preprint arXiv:2005.04210, 2020
2005 arXiv
-
[18]
The loss landscape of overparameterized neural networks
Yaim Cooper. The loss landscape of overparameterized neural networks. arXiv preprint arXiv:1804.10200, 2018
2018 arXiv
-
[19]
Brian Davies
E. Brian Davies. Spectral Theory and Differential Operators. Cambridge Studies in Advanced Mathematics. Cambridge University Press, 1995
1995
-
[20]
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht. Essentially no barriers in neural network energy landscape. In International conference on machine learning, pages 1309--1318. PMLR, 2018
2018
-
[21]
Thin shell implies spectral gap up to polylog via a stochastic localization scheme
Ronen Eldan. Thin shell implies spectral gap up to polylog via a stochastic localization scheme. Geometric and Functional Analysis, 23: 0 532 -- 569, 2013. URL https://api.semanticscholar.org/CorpusID:253637768
2013
-
[22]
On the convergence of langevin monte carlo: The interplay between tail growth and smoothness
Murat A Erdogdu and Rasa Hosseinzadeh. On the convergence of langevin monte carlo: The interplay between tail growth and smoothness. In Conference on Learning Theory, pages 1776--1822. PMLR, 2021
2021
-
[23]
Lawrence C. Evans. Partial differential equations. American Mathematical Society, 2010
2010
-
[24]
The activated complex in chemical reactions
Henry Eyring. The activated complex in chemical reactions. The Journal of Chemical Physics, 3 0 (2): 0 107--115, 1935
1935
-
[25]
Convergence rates for the stochastic gradient descent method for non-convex objective functions
Benjamin Fehrman, Benjamin Gess, and Arnulf Jentzen. Convergence rates for the stochastic gradient descent method for non-convex objective functions. Journal of Machine Learning Research, 21 0 (136): 0 1--48, 2020
2020
-
[26]
Topology and geometry of half-rectified network optimization
C Daniel Freeman and Joan Bruna. Topology and geometry of half-rectified network optimization. In 5th International Conference on Learning Representations, ICLR 2017, 2017
2017
-
[27]
Random Perturbations of Dynamical Systems, volume 260
Mark I Freidlin and Alexander D Wentzell. Random Perturbations of Dynamical Systems, volume 260. Springer Science & Business Media, 2012
2012
-
[28]
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson. Loss surfaces, mode connectivity, and fast ensembling of dnns. Advances in neural information processing systems, 31, 2018
2018
-
[29]
Metastability in reversible diffusion processes ii: Precise asymptotics for small eigenvalues
V \'e ronique Gayrard, Anton Bovier, and Markus Klein. Metastability in reversible diffusion processes ii: Precise asymptotics for small eigenvalues. Journal of the European Mathematical Society, 7 0 (1): 0 69--99, 2005
2005
-
[30]
Differential Topology
Alan Pollack Victor Guillemin. Differential Topology. 1974
1974
-
[31]
Nonlinear analysis on manifolds: Sobolev spaces and inequalities
Emmanuel Hebey. Nonlinear analysis on manifolds: Sobolev spaces and inequalities. 1999. URL https://api.semanticscholar.org/CorpusID:118747316
1999
-
[32]
Isoperimetric problems for convex bodies and a localization lemma
Ravi Kannan, L \'a szl \'o Mikl \'o s Lov \'a sz, and Mikl \'o s Simonovits. Isoperimetric problems for convex bodies and a localization lemma. Discrete & Computational Geometry, 13: 0 541--559, 1995. URL https://api.semanticscholar.org/CorpusID:14881695
1995
-
[33]
Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt. Linear convergence of gradient and proximal-gradient methods under the polyak- ojasiewicz condition. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2016, Riva del Garda, Italy, September 19...
2016
-
[34]
Mountain pass theorems and global homeomorphism theorems
Guy Katriel. Mountain pass theorems and global homeomorphism theorems. In Annales de l'Institut Henri Poincar \'e C, Analyse non lin \'e aire , volume 11, pages 189--209. Elsevier, 1994
1994
-
[35]
Elimination of all bad local minima in deep learning
Kenji Kawaguchi and Leslie Kaelbling. Elimination of all bad local minima in deep learning. In International Conference on Artificial Intelligence and Statistics, pages 853--863. PMLR, 2020
2020
-
[36]
Brownian motion in a field of force and the diffusion model of chemical reactions
Hendrik Anthony Kramers. Brownian motion in a field of force and the diffusion model of chemical reactions. physica, 7 0 (4): 0 284--304, 1940
1940
-
[37]
Explaining landscape connectivity of low-cost solutions for multilayer nets
Rohith Kuditipudi, Xiang Wang, Holden Lee, Yi Zhang, Zhiyuan Li, Wei Hu, Rong Ge, and Sanjeev Arora. Explaining landscape connectivity of low-cost solutions for multilayer nets. Advances in neural information processing systems, 32, 2019
2019
-
[38]
John M. Lee. Introduction to Smooth Manifolds. 2020
2020
-
[39]
Yin Tat Lee and Santosh S. Vempala. The kannan-lov\'asz-simonovits conjecture, 2018. URL https://arxiv.org/abs/1807.03465
2018 arXiv
-
[40]
Eldan's stochastic localization and the kls conjecture: Isoperimetry, concentration and mixing
Yin Tat Lee and Santosh S Vempala. Eldan's stochastic localization and the kls conjecture: Isoperimetry, concentration and mixing. Annals of Mathematics, 199 0 (3): 0 1043--1092, 2024
2024
-
[41]
The effect of smooth parametrizations on nonconvex optimization landscapes
Eitan Levin, Joe Kileel, and Nicolas Boumal. The effect of smooth parametrizations on nonconvex optimization landscapes. Mathematical Programming, pages 1--49, 2024
2024
-
[42]
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and E Weinan. Stochastic modified equations and adaptive stochastic gradient algorithms. In International Conference on Machine Learning, pages 2101--2110. PMLR, 2017
2017
-
[43]
A hessian-aware stochastic differential equation for modelling sgd
Xiang Li, Zebang Shen, Liang Zhang, and Niao He. A hessian-aware stochastic differential equation for modelling sgd. arXiv preprint arXiv:2405.18373, 2024
2024 arXiv
-
[44]
What happens after sgd reaches zero loss?-a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora. What happens after sgd reaches zero loss?-a mathematical framework. In 10th International Conference on Learning Representations, ICLR 2022, 2022
2022
-
[45]
Adding one neuron can eliminate all bad local minima
Shiyu Liang, Ruoyu Sun, and Jason D Lee. Adding one neuron can eliminate all bad local minima. Advances in Neural Information Processing Systems, 31, 2018 a
2018
-
[46]
Understanding the loss surface of neural networks for binary classification
Shiyu Liang, Ruoyu Sun, Yixuan Li, and Rayadurgam Srikant. Understanding the loss surface of neural networks for binary classification. In International Conference on Machine Learning, pages 2835--2843. PMLR, 2018 b
2018
-
[47]
Exploring neural network landscapes: Star-shaped and geodesic connectivity
Zhanran Lin, Puheng Li, and Lei Wu. Exploring neural network landscapes: Star-shaped and geodesic connectivity. arXiv preprint arXiv:2404.06391, 2024
2024 arXiv
-
[48]
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks
Chaoyue Liu, Libin Zhu, and Mikhail Belkin. Loss landscapes and optimization in over-parameterized non-linear systems and neural networks. Applied and Computational Harmonic Analysis, 59: 0 85--116, 2022
2022
-
[49]
A topological property of real analytic subsets
Stanislaw Lojasiewicz. A topological property of real analytic subsets. Coll. du CNRS, Les \'e quations aux d \'e riv \'e es partielles , 117 0 (87-89): 0 2, 1963
1963
-
[50]
Poincaré and logarithmic Sobolev inequalities by decomposition of the energy landscape
Georg Menz and Andr \'e Schlichting. Poincaré and logarithmic Sobolev inequalities by decomposition of the energy landscape . The Annals of Probability, 42 0 (5): 0 1809 -- 1884, 2014. doi:10.1214/14-AOP908. URL https://doi.org/10.1214/14-AOP908
2014 doi
-
[51]
Stasheff
John Milnor and James D. Stasheff. Characteristic Classes. (AM-76), Volume 76. Princeton University Press, Princeton, 1974. ISBN 9781400881826. doi:doi:10.1515/9781400881826. URL https://doi.org/10.1515/9781400881826
1974 doi
-
[52]
On connected sublevel sets in deep learning
Quynh Nguyen. On connected sublevel sets in deep learning. In International conference on machine learning, pages 4790--4799. PMLR, 2019
2019
-
[53]
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi. Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks. IEEE Journal on Selected Areas in Information Theory, 1 0 (1): 0 84--105, 2020
2020
-
[54]
Homogenization of sgd in high-dimensions: Exact dynamics and generalization properties
Courtney Paquette, Elliot Paquette, Ben Adlam, and Jeffrey Pennington. Homogenization of sgd in high-dimensions: Exact dynamics and generalization properties. arXiv preprint arXiv:2205.07069, 2022
2022 arXiv
-
[55]
Gradient methods for minimizing functionals
Boris Teodorovich Polyak. Gradient methods for minimizing functionals. Zhurnal vychislitel'noi matematiki i matematicheskoi fiziki, 3 0 (4): 0 643--653, 1963
1963
-
[56]
Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis
Maxim Raginsky, Alexander Rakhlin, and Matus Telgarsky. Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis. In Conference on Learning Theory, pages 1674--1703. PMLR, 2017
2017
-
[57]
Fast convergence to non-isolated minima: four equivalent conditions for c 2 functions
Quentin Rebjock and Nicolas Boumal. Fast convergence to non-isolated minima: four equivalent conditions for c 2 functions. Mathematical Programming, pages 1--49, 2024
2024
-
[58]
On the quality of the initial basin in overspecified neural networks
Itay Safran and Ohad Shamir. On the quality of the initial basin in overspecified neural networks. In International Conference on Machine Learning, pages 774--782. PMLR, 2016
2016
-
[59]
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun. Eigenvalues of the hessian in deep learning: Singularity and beyond. arXiv preprint arXiv:1611.07476, 2016
2016 arXiv
-
[60]
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou. Empirical analysis of the hessian of over-parametrized neural networks. arXiv preprint arXiv:1706.04454, 2017
2017 arXiv
-
[61]
L.W. Tu. An Introduction to Manifolds. Universitext. Springer New York, 2010. ISBN 9781441973993. URL https://books.google.com.hk/books?id=br1KngEACAAJ
2010
-
[62]
Spurious valleys in two-layer neural network optimization landscapes
Luca Venturi, Afonso S Bandeira, and Joan Bruna. Spurious valleys in two-layer neural network optimization landscapes. arXiv preprint arXiv:1802.06384, 2018
2018 arXiv
-
[63]
On the volume of tubes
Hermann Weyl. On the volume of tubes. American Journal of Mathematics, 61: 0 461, 1939. URL https://api.semanticscholar.org/CorpusID:124362885
1939
-
[64]
Sampling as optimization in the space of measures: The langevin dynamics as a composite optimization problem
Andre Wibisono. Sampling as optimization in the space of measures: The langevin dynamics as a composite optimization problem. In Conference on Learning Theory, pages 2093--3027. PMLR, 2018
2018
-
[65]
Stochastic gradient descent with noise of machine learning type part ii: Continuous time analysis
Stephan Wojtowytsch. Stochastic gradient descent with noise of machine learning type part ii: Continuous time analysis. Journal of Nonlinear Science, 34 0 (1): 0 16, 2024
2024
-
[66]
Global convergence and variance reduction for a class of nonconvex-nonconcave minimax problems
Junchi Yang, Negar Kiyavash, and Niao He. Global convergence and variance reduction for a class of nonconvex-nonconcave minimax problems. Advances in Neural Information Processing Systems, 33: 0 1153--1165, 2020
2020
-
[67]
A hitting time analysis of stochastic gradient langevin dynamics
Yuchen Zhang, Percy Liang, and Moses Charikar. A hitting time analysis of stochastic gradient langevin dynamics. In Conference on Learning Theory, pages 1980--2022. PMLR, 2017
1980
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.