REVIEW 1 major objections 4 minor 38 references
Mirror Langevin diffusions: Convergence rates and Markov chain approximations
T0 review · 1 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper proves that a two-step Gibbs sampler with exact stationary law $\mu$ approximates the Mirror Langevin diffusion and contracts in $\chi^2$ at rate $1-c_0\epsilon$ per step, giving a mixing time consistent with the diffusion.
desk verdict Solid diffusion-side theory with a genuinely new Markov-chain construction, but the advertised O(1/ε) chi-square rate is conditional on an external Poincaré inequality for the pushed-forward measure that the paper never verifies. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Hessian manifold with metric $\nabla^2 u$ and its dual coordinate $x^*=\nabla u(x)$, together with the two-step Gibbs kernel $r_\epsilon(z|x)$ built from the Gaussian conditional $q_\epsilon(y|x)=N(x^*, \epsilon \nabla^2 u(x))$ and its Bayes reverse $\hat q_\epsilon(z|y)$. The identity $F(x^*)=V(x)+\log\det\nabla^2 u(x)$ ties the primal and dual densities and underlies the coordinate-invariance argument. The rate proof proceeds by comparing the discrete Dirichlet energy of $r_\epsilon$ with the energy of the chain $(A, A+\sigma_0\sqrt{\epsilon}Z)$ where $A\sim e^{-F}$, whose maximal correlation is controlled by a strong data processing inequality; the comparison constant comes from the $L$-smoothness of $F$ and the uniform Hessian bounds of Assumption 1. A key technical observation is that the Laplace phase $\Psi(w)=(w^*-y)^T(\partial w/\partial w^*)(w^*-y)$ has vanishing third derivatives at its minimizer $w_0=y^*$, which makes the Gaussian Laplace approximation accurate to order $\epsilon^2$.
What would settle it
Construct a smooth, strongly convex $u$ with bounded Hessian and a smooth $V$ satisfying the paper's Assumption 1 for which the push-forward $\nu=(\nabla u)_\#\mu$ fails to satisfy a Poincar\'e inequality (for example, a density with a sufficiently heavy tail), then compute the spectral gap or chi-square contraction coefficient of the two-step kernel $r_\epsilon$; if the gap is not $\Theta(\epsilon)$ or the bound $(1-c_0\epsilon)$ fails, the theorem's reachable scope is smaller than stated.
Extended reading notes
Core claim
The central discovery is a transfer principle: the primal MLD with stationary density $\mu=e^{-V}$ and the dual MLD with stationary density $\nu=e^{-F}$, where $F(\nabla u(x))=V(x)+\log\det\nabla^2 u(x)$, are the same Langevin diffusion on the Hessian manifold $(\mathbb{R}^d,\nabla^2 u)$ written in two coordinate charts. This lets the paper prove functional inequalities for whichever side is easier, then transfer them to the other with the same constant. Using Lyapunov functions of the form $e^{\alpha u}$ and $e^{\alpha u^*}$, it obtains sufficient tail conditions for a Poincar\'e or log-Sobolev inequality for the MLD without log-concavity of either $\mu$ or $\nu$. For the Markov chain, the paper proves that the Gibbs sampler's discrete Dirichlet energy is comparable to that of an isotropic Gaussian two-step chain, and that a strong data processing inequality bounds the chi-square contraction coefficient by $(1+\sigma_0^2\epsilon/c_F)^{-2}$, yielding a spectral gap of order $\epsilon$.
Load-bearing premise
The rate guarantee depends on the pushed-forward density $\nu=(\nabla u)_\#\mu$ being well behaved in the sense that it satisfies a Poincar\'e inequality; the paper does not derive this from its other assumptions, so the user must choose $u$ to make it true.
Editorial extensions
If this is right
- For any target $\mu$ for which a mirror map $u$ makes $e^{-F}$ Poincar\'e with $F$ $L$-smooth, the two-step Gibbs chain is an unbiased sampler: every iterate has a density, the stationary law is exactly $\mu$, and the $\chi^2$ error contracts by $1-c_0\epsilon$ per step.
- Mixing time of the chain is $O(\epsilon^{-1}\log(1/\delta))$ steps to come within $\delta$ in $\chi^2$, matching the continuous MLD's $O(\log(1/\delta))$ time under the $\epsilon$ time rescaling.
- The Euclidean choice $u(x)=\|x\|^2/2$ reduces to the proximal sampler with $N(x,\epsilon I)$, so the paper supplies an unbiased, rate-guaranteed discretization of classical Langevin diffusion.
- The Lyapunov conditions of Theorems 1 and 2 are checkable from first-order information on $V$ and the Hessian quadratic form of $u$, so they offer a concrete route to certify exponential mixing for targets that are not strongly log-concave.
- Because primal and dual representations share the same Dirichlet energy, any Poincar\'e or log-Sobolev constant proved for one side transfers to the other with the same constant.
Reading between the lines
- The mirror map $u$ can be read as a user-chosen preconditioner: the theorems suggest a design rule of thumb, pick $u$ so the Brenier push-forward $\nu=(\nabla u)_\#\mu$ is as log-concave or Poincar\'e-friendly as possible, since the dual condition is the bottleneck; this is not stated as an algorithm in the paper.
- The comparison technique may apply beyond this chain: any Gibbs sampler whose forward conditional is close to an isotropic Gaussian in the duality coordinates, and whose reverse conditional has a uniform Gaussian lower bound, would inherit a $1-c\epsilon$ chi-square contraction whenever the target side satisfies Poincar\'e; this is an extension the paper leaves implicit.
- The vanishing-third-derivative property of $\Psi$ suggests the Gaussian approximation to the Schr\"odinger bridge is accurate to $O(\epsilon^2)$, so one could build higher-order unbiased discretizations of the MLD or accelerated Sinkhorn schemes by adding $\epsilon^2$ corrections; the paper does not pursue this.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies Mirror Langevin diffusions (MLD) on Hessian manifolds, i.e., Langevin diffusions intrinsic to the metric g=∇²u with stationary density μ=e^{-V}. The main theoretical results are: (i) Lyapunov-function sufficient conditions for the MLD to satisfy a Poincaré inequality (Theorem 1) or a logarithmic Sobolev inequality (Theorem 2), intended to cover targets that are not strongly log-concave; (ii) a two-step Gibbs-sampler Markov chain with transition density r_ε that is exactly stationary with respect to μ (Theorem 3); (iii) weak convergence of the continuous-time interpolation of this chain to the MLD as ε↓0 (Theorem 4); and (iv) a χ²-divergence contraction rate of the form (1-c₀ε)^k for the discrete chain, under the additional hypothesis that the push-forward measure ν=(∇u)_#μ satisfies a Poincaré inequality and F is L-smooth (Theorem 5). The proofs use Laplace-expansion computations, comparison of discrete Dirichlet energies, and the strong data processing inequality of Klartag and Ordentlich.
Significance. If the results hold, the paper makes a useful contribution. The Lyapunov conditions in Theorems 1–2 are checkable and avoid the often-untractable curvature-dimension verification; the construction of an exactly stationary Markov chain approximation to the MLD is elegant and improves on biased Euler-type discretizations. The Laplace computations in Section 4 are careful, and the observation that the relevant third derivatives vanish at the minimizer is a genuinely useful simplification. The SDPI-based comparison in Theorem 5 gives the correct diffusion-consistent time scale and connects the Gibbs sampler to entropic optimal transport. The main limitation is that Theorem 5's rate is conditional on a Poincaré inequality for the push-forward ν, a condition that is not derived from the paper's other assumptions and can fail in natural examples; this restricts the advertised 'guaranteed convergence rate' more than the abstract and introduction suggest.
major comments (1)
- [Theorem 5(i) and its proof, Step 4] The Poincaré inequality for ν=(∇u)_#μ is an external hypothesis that is not implied by Assumption 1 or by L-smoothness of F. For example, take d≥1, V(x)=(d+1)log(1+‖x‖²), and u(x)=c‖x‖²/2. Then u and u* satisfy the uniform Hessian and derivative bounds of Assumption 1, and F is L-smooth, but ν has density proportional to (1+‖y/c‖²)^{-(d+1)}, which has polynomial tails. Since a Poincaré inequality implies exponential concentration of 1-Lipschitz functions, ν has no Poincaré inequality, i.e., c_F=∞. In Step 4 the Klartag–Ordentlich bound S²≤(1+σ₀²ε/c_F)^{-1} then becomes vacuous, and the comparison argument yields no positive spectral gap. The theorem is correctly stated as a conditional result, but the paper should either supply checkable sufficient conditions for Theorem 5(i) — for example, by applying the dual version of Theorem 1 to the dual MLD, which would give a Lyapunov condition of the form y·∇F(y)-αyᵀ∇²u*(y)y→∞ — or explicitly delimit Theorem 5 as a transfer theorem whose applicability must be verified case by case.
minor comments (4)
- [Example 1] The classical Langevin diffusion for a 1-strongly convex potential converges in KL at rate 2, not 1, under the Bakry–Émery criterion. The comparison in Example 1 should therefore be 2γ versus 2, rather than 2γ versus 1. The qualitative conclusion that the MLD can be arbitrarily faster is unaffected.
- [Theorem 5, Step 4] In the sentence 'ηχ²≤(1+σ₀²ε/c_F)^{-2}=1+Θ(ε)', the final equality has the wrong sign: (1+aε)^{-2}=1-2aε+O(ε²), so it should read 1-Θ(ε). As written, the spectral gap 1-ηχ² would be negative. The subsequent conclusion is correct once the sign is fixed.
- [Example 3] The potential u(x)=log∑ᵢ e^{xᵢ} is not strictly convex: its Hessian diag(w)-wwᵀ has the all-ones vector in its kernel. Hence Example 3 is not covered by Theorem 1, which assumes a global diffeomorphism and locally uniformly elliptic Hessian, nor by Assumption 1 later. The example should either be replaced by a uniformly convex approximation or explicitly labelled as formal.
- [Theorem 5, Step 3] The lower bound e^{F(y)-\tilde F(y)}≥e^{-Lσ₀²d/2} is valid; the pointwise inequality from L-smoothness has the correct direction, and the Gaussian expectation is in fact (1+Lσ₀²ε)^{-d/2} exp(σ₀²ε‖∇F(y)‖²/(2(1+Lσ₀²ε))), which is bounded below by the displayed constant. No correction is needed here.
Circularity Check
No significant circularity: Theorem 5's chi-square rate is explicitly conditional on an assumed Poincaré inequality for the push-forward ν, and all central proofs use external, independently stated results rather than self-referential predictions.
full rationale
The paper's derivations are conditional statements with explicit hypotheses, and no load-bearing step reduces to its own input by construction. Theorem 1 proves a Poincaré inequality for the MLD from the Lyapunov condition (18) by constructing the Lyapunov function ξ_α = e^{αu} and applying external criteria from [4] and [6]; the function and hypotheses are not derived from the target inequality. Theorem 4 proves the diffusion approximation by verifying the three conditions of Durrett [17, Theorem 8.7.1] through Lemmas 2–5, which are proved from Assumption 1 inside the paper; no diffusion limit is assumed or imported. Theorem 5 states, as an explicit hypothesis, that ν = e^{-F} = (∇u)_#μ satisfies a Poincaré inequality, and then transfers that assumed inequality to the discrete chain via the external SDPI theorem of Klartag–Ordentlich [22] and the framework of Raginsky [34]. The fact that Assumption 1 alone does not imply this Poincaré inequality for ν is a limitation on the theorem's applicability, not circularity, because the theorem declares the condition as an input rather than claiming to derive it. Self-citations to [14], [31], and [32] are motivational or used for standard coordinate-change facts, such as 'See [14, Theorem 3.5]' for the dual SDE; no central conclusion is justified solely by an unverified self-citation. The paper itself describes the Sinkhorn-chain connection as conjectural ('conjectured to converge to a time-inhomogeneous generalization of the MLD'), which further confirms that this connection is not load-bearing. No fitted constants, renamed empirical patterns, or definitional equivalences appear in the main results.
Assumptions & free parameters
assumptions (4)
- domain assumption mu = e^{-V} and nu = e^{-F} = (nabla u)_#mu are fully supported probability densities and weak solutions of the primal/dual MLD exist.
- domain assumption u and u^* are strictly convex and C^2; under Assumption 1, u, u^* are C^6 with derivatives bounded up to order 6 and c0 I <= nabla^2 u <= C0 I.
- domain assumption V has bounded second derivatives (Assumption 1 (iii)); in Theorem 5, F is L-smooth and nu satisfies a Poincaré inequality.
- standard math Standard external theorems hold: [4, Theorem 4.6.2] (Lyapunov to Poincaré), [4, Theorem 5.2.1] (CD to KL), [22, Theorem 1.1] (SDPI for Gaussian channels), [17, Theorem 8.7.1] (diffusion approximation).
Cite this review
Pith. "Pith review of Mirror Langevin diffusions: Convergence rates and Markov chain approximations." pith.science (2026). https://pith.science/paper/EDIBR7NZ
@misc{pith2026260722892,
author = {Pith},
title = {Pith review of: Mirror Langevin diffusions: Convergence rates and Markov chain approximations},
year = {2026},
howpublished = {\url{https://pith.science/paper/EDIBR7NZ}},
note = {Machine review of arXiv:2607.22892}
}
abstract
Given a strongly convex function $u$, equip $R^d$ with a Riemannian metric given by the Hessian $\nabla^2 u$. This is a so-called Hessian manifold. Given a probability density $\mu$ one may run a Langevin diffusion intrinsic to the manifold with stationary distribution $\mu$. Such (Hessian) manifold-valued Langevin diffusions are called Mirror Langevin diffusions (MLD) which have recently become popular. One of the questions we explore is whether, given $\mu$, one can choose $u$ to get an exponential convergence to equilibrium for the MLD, especially if $\mu$ is not strongly log-concave. Our results are based on Lyapunov function methods and give sufficient conditions for a Poincar\'e or a log-Sobolev inequality to hold for the MLD. These, in turn, imply exponential convergence. We also introduce a Markov chain approximation to the MLD given by a two step Gibbs sampler with stationary distribution $\mu$. This Markov chain is a variant of the Sinkhorn Markov chain introduced in arXiv:2307.16421 that is conjectured to converge to a time-inhomogeneous generalization of the MLD. Under suitable assumptions, we prove that the Markov chain has a guaranteed convergence rate in $\chi^2$ that is consistent with the diffusion time scale. Our proofs are based on ideas from entropic optimal transport and strong data processing inequalities.
Reference graph
Works this paper leans on
-
[1]
K. Ahn and S. Chewi. Efficient constrained sampling via the mirror-Langevin algo- rithm. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, editors,Advances in Neural Information Processing Systems, volume 34, pages 28405– 28418. Curran Associates, Inc., 2021
work page 2021
- [2]
- [3]
- [4]
-
[5]
P. Cattiaux and A. Guillin. Functional inequalities via Lyapunov conditions. InOpti- mal Transportation: Theory and Applications, pages 155–186. Cambridge University Press, 2011
work page 2011
-
[6]
P. Cattiaux and A. Guillin. Functional inequalities, Lyapunov conditions and uniform ergodicity.Journal of Functional Analysis, 272(6):2361–2391, 2017
work page 2017
-
[7]
P. Cattiaux, A. Guillin, W. F.-Y., and L. Wu. Lyapunov conditions for super Poincar´ e inequalities.Journal of Functional Analysis, 256:1821–1841, 2009. 36 BENJAMIN CAPDEVILLE, YOUNG-HEON KIM, AND SOUMIK PAL
work page 2009
-
[8]
Y. Chen, S. Chewi, A. Salim, and A. Wibisono. Improved analysis for a proximal algorithm for sampling. In P.-L. Loh and M. Raginsky, editors,Proceedings of Thirty Fifth Conference on Learning Theory, volume 178 ofProceedings of Machine Learning Research, pages 2984–3014. PMLR, 02–05 Jul 2022
work page 2022
Show all 38 references
-
[9]
Chen and S
Z. Chen and S. S. Vempala. Optimal convergence rate of Hamiltonian Monte Carlo for strongly logconcave distributions.Theory of Computing, 18(9):1–18, 2022
2022
-
[10]
Cheng, N
X. Cheng, N. S. Chatterji, Y. Abbasi-Yadkori, P. L. Bartlett, and M. I. Jordan. Sharp convergence rates for Langevin dynamics in the nonconvex setting.arXiv e-prints, page arXiv:1805.01648, May 2018
2018 arXiv
-
[11]
S. Chewi. Log-concave sampling. Available online at chewisinho.github.io/main.pdf, 2026
2026
-
[12]
Chewi, T
S. Chewi, T. Le Gouic, C. Lu, T. Maunu, P. Rigollet, and A. Stromme. Exponential ergodicity of mirror-Langevin diffusions. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors,Advances in Neural Information Processing Systems, volume 33, pages 19573–19585. ...
2020
-
[13]
Chiarini, G
A. Chiarini, G. Conforti, G. Greco, and L. Tamanini. A semiconcavity approach to stability of entropic plans and exponential convergence of Sinkhorn’s algorithm. arxiv preprint [math.PR], 2025
2025
-
[14]
Deb, Y.-H
N. Deb, Y.-H. Kim, S. Pal, and G. Schiebinger. Wasserstein mirror gradient flow as the limit of the Sinkhorn algorithm.The Annals of Probability, 2026
2026
-
[15]
Diaconis, K
P. Diaconis, K. Khare, and L. Saloff-Coste. Gibbs sampling, exponential families and orthogonal polynomials.Statistical Science, 23(2):151–178, 2008
2008
-
[16]
Durmus and E
A. Durmus and E. Moulines. Nonasymptotic convergence analysis for the unadjusted Langevin algorithm.The Annals of Applied Probability, 27(3):1551–1587, 2017
2017
-
[17]
Durrett.Stochastic Calculus: A Practical Introduction
R. Durrett.Stochastic Calculus: A Practical Introduction. Probability and Stochastics Series. Taylor & Francis, 1996
1996
-
[18]
Hille.Ordinary Differential Equations in the Complex Domain
E. Hille.Ordinary Differential Equations in the Complex Domain. Dover Books on Mathematics. Dover Publications, 1997
1997
-
[19]
Hsieh, A
Y.-P. Hsieh, A. Kavis, P. Rolland, and V. Cevher. Mirrored Langevin dynamics. Advances in Neural Information Processing Systems, 31, 2018
2018
-
[20]
E. P. Hsu.Stochastic analysis on manifolds, volume 38 ofGraduate Studies in Math- ematics. American Mathematical Society, Providence, RI, 2002
2002
-
[21]
B. Klartag. Logarithmically-Concave Moment Measures I. In B. Klartag and E. Mil- man, editors,Geometric Aspects of Functional Analysis, volume 2116, pages 231–260. Springer International Publishing, 2014
2014
-
[22]
Klartag and O
B. Klartag and O. Ordentlich. The strong data processing inequality under the heat flow.IEEE Transactions on Information Theory, 71(5):3317–3333, 2025
2025
-
[23]
A. V. Kolesnikov. Hessian metrics, CD(K,N)-spaces, and optimal transportation of log-concave measures.Discrete and Continuous Dynamical Systems, 34(4):1511–1532, 2014
2014
-
[24]
Y. T. Lee, R. Shen, and K. Tian. Structured logconcave sampling with a restricted Gaussian oracle. In M. Belkin and S. Kpotufe, editors,Proceedings of Thirty Fourth Conference on Learning Theory, volume 134 ofProceedings of Machine Learning Re- search, pages 2993–3050. PMLR, 1...
2021
-
[25]
L´ eonard
C. L´ eonard. A survey of the Schr¨ odinger problem and some of its connections with optimal transport.Discrete Contin. Dyn. Syst., 34(4):1533–1574, 2014
2014
-
[26]
R. Li, M. Tao, S. S. Vempala, and A. Wibisono. The mirror Langevin algorithm converges with vanishing bias. In S. Dasgupta and N. Haghtalab, editors,Proceedings of The 33rd International Conference on Algorithmic Learning Theory, volume 167 ofProceedings of Machine Learning Re...
2022
-
[27]
S. Lisini. Nonlinear diffusion equations with variable coefficients as gradient flows in Wasserstein spaces.ESAIM: Control, Optimisation and Calculus of Variations, 15(3):712–740, 2009
2009
-
[28]
J. S. Liu, W. H. Wong, and A. Kong. Covariance structure and convergence rate of the Gibbs sampler with various scans.Journal of the Royal Statistical Society. Series B (Methodological), 57(1):157–169, 1995
1995
-
[29]
Mangoubi and A
O. Mangoubi and A. Smith. Mixing of Hamiltonian Monte Carlo on strongly log- concave distributions 2: Numerical integrators. In K. Chaudhuri and M. Sugiyama, editors,Proceedings of the Twenty-Second International Conference on Artificial In- telligence and Statistics, volume 8...
2019
-
[30]
Meyn and R
S. Meyn and R. L. Tweedie.Markov chains and stochastic stability. Second Edition. Cambridge University Press, 2009
2009
-
[31]
Mulcahy and S
G. Mulcahy and S. Pal. Diffusion approximations to Schr¨ odinger bridges on mani- folds”. Arxiv preprint 2512.18867 [math.PR], 2025
2025
-
[32]
S. Pal. On the difference between entropic cost and the optimal transport cost.The Annals of Applied Probability, 34(1B):1003–1028, 2024
2024
-
[33]
G. Parisi. Correlation functions and computer simulations.Nuclear Physics B, 180:378–384, 1981
1981
-
[34]
Raginsky
M. Raginsky. Strong data processing inequalities and Φ-Sobolev inequalities for dis- crete channels.IEEE Transactions on Information Theory, 62:3355–3389, 2014
2014
-
[35]
G. O. Roberts and R. L. Tweedie. Exponential convergence of Langevin distributions and their discrete approximations.Bernoulli, 2(4):341–363, 1996
1996
-
[36]
Shima.The Geometry of Hessian Structures
H. Shima.The Geometry of Hessian Structures. G - Reference,Information and In- terdisciplinary Subjects Series. World Scientific, 2007
2007
-
[37]
Vempala and A
S. Vempala and A. Wibisono. Rapid convergence of the unadjusted Langevin al- gorithm: Isoperimetry suffices. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alch´ e-Buc, E. Fox, and R. Garnett, editors,Advances in Neural Information Processing Systems, volume 32. Curran Ass...
2019
-
[38]
K. S. Zhang, G. Peyr´ e, J. Fadili, and M. Pereyra. Wasserstein control of mirror Langevin Monte Carlo. InConference on Learning Theory, pages 3814–3841. PMLR, 2020. Laboratoire Math´ematiques d’Orsay, Universit´e Paris-Saclay, Inria ParMA, 91405, Orsay, France, Email: benjami...
2020
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.