Pith. sign in

REVIEW 1 major objections 4 minor 89 references

On the convergence of graph Laplacians with a symmetric divergence

T0 review · 1 major / 4 minor · reviewed 2026-07-08 · grok-4.5

Pith's one-line read Graph Laplacians built from a smooth symmetric divergence converge pointwise to the Laplace–Beltrami operator of the metric induced by its Hessian.

desk verdict Clean fourth-order comparison for smooth symmetric divergences that justifies graph-Laplacian convergence; the only soft spot is Sinkhorn non-degeneracy, and it only hits the example. read the letter →

arxiv 2607.05892 v1 pith:STQR7YKK submitted 2026-07-07 stat.ML cs.LG

classification stat.MLcs.LG
keywords graphLaplacianmanifoldlearningsymmetricdivergenceLaplace-BeltramiSinkhornpointwiseconvergenceRiemannianmetricgeodesicdistance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper extends a classical comparison between Euclidean distance and geodesic distance on a manifold to a broad class of smooth symmetric divergences. Given a smooth compact connected Riemannian submanifold of Euclidean space equipped with such a divergence D that is non-degenerate, the Hessian of D at each point defines a Riemannian metric g, and D itself differs from the squared geodesic distance of g by at most a fourth-order term. That single estimate is enough to transfer the usual pointwise convergence proofs for graph Laplacians from ordinary distances to the divergence D, so the resulting discrete operators converge to the Laplace–Beltrami operator of g. The result therefore licenses the use of more flexible similarity measures—most notably Sinkhorn divergences between probability measures that themselves lie on a manifold—inside manifold-learning pipelines while still recovering the correct continuum geometry. A reader who works with data that naturally come with divergences rather than Euclidean distances obtains a rigorous justification for treating those divergences as if they were squared geodesic distances.

What carries the argument

The fourth-order approximation |D(p,q) − d_g(p,q)^2| ≤ K d_g(p,q)^4 obtained by Taylor expansion of the smooth divergence about the diagonal; it replaces the classical Euclidean-to-geodesic estimate and is the sole geometric input needed for the convergence argument.

What would settle it

On a concrete low-dimensional compact manifold of probability measures, compute the Hessian of the Sinkhorn divergence; if it fails to be positive definite at some point, or if numerical graph Laplacians built from that divergence fail to converge to the known Laplace–Beltrami operator while the fourth-order bound is violated, the claim is false.

Watch

Extended reading notes

Core claim

If M is a smooth compact connected submanifold of R^d carrying a smooth symmetric divergence D whose Hessian is positive definite at every point, the bilinear form g_p = (1/2) Hess_p D(p,·) is a Riemannian metric, and there exists K > 0 such that |D(p,q) − d_g(p,q)^2| ≤ K d_g(p,q)^4 for all p,q in M. This bound alone guarantees pointwise convergence of the graph Laplacians constructed from D to the Laplace–Beltrami operator of g.

Load-bearing premise

The divergence must have a positive-definite Hessian at every point of the compact manifold so that the induced bilinear form is a genuine Riemannian metric everywhere.

Editorial extensions

If this is right

  • Graph Laplacians that use any smooth non-degenerate symmetric divergence (including Sinkhorn) as edge weights still converge pointwise to the Laplace–Beltrami operator of the induced metric.
  • Manifold-learning algorithms may replace Euclidean distances by divergences that better match the data geometry without losing the asymptotic link to Riemannian geometry.
  • The same fourth-order control applies to other spectral methods whose proofs rely only on a squared-distance approximation of this quality.
  • When data points are themselves probability measures lying on a manifold, the Sinkhorn divergence yields a graph Laplacian that recovers the geometry of that manifold of measures.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The argument should extend immediately to other divergences from optimal transport or information geometry once their Hessians are verified to be positive definite.
  • If non-degeneracy holds only on a compact subset, local pointwise convergence of the graph Laplacian should still be obtainable by the same expansion.
  • The explicit fourth-order remainder suggests a systematic way to derive higher-order corrections to the discrete Laplacian from further terms in the Taylor expansion of D.
  • Numerical estimation of the constant K on low-dimensional Sinkhorn examples would give concrete sample-size guidance for the asymptotic regime.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. The manuscript establishes that if M is a smooth compact connected Riemannian submanifold of R^d equipped with a smooth symmetric divergence D that is non-degenerate (Hess_p(D(p,·)) positive definite for every p), and if the Riemannian metric is defined by g_p := (1/2) Hess_p(D(p,·)), then there exists K > 0 such that |D(p,q) - d_g(p,q)^2| ≤ K d_g(p,q)^4 for all p,q in M. This fourth-order comparison is shown to be sufficient for pointwise convergence of graph Laplacians built from D to the Laplace–Beltrami operator of g. The argument proceeds by Taylor expansion in normal coordinates (odd orders vanish by symmetry and criticality on the diagonal) together with uniformity on the compact unit sphere bundle, followed by standard integral-operator arguments for the Laplacian limit. Sinkhorn divergences on a manifold-parametrized family of probability measures are discussed as a motivating example.

Significance. If the result holds, it cleanly extends the classical Euclidean-distance comparison that underpins manifold-learning analyses of graph Laplacians to a broader class of smooth symmetric divergences. The derivation is essentially parameter-free once non-degeneracy supplies the metric g, and the fourth-order bound is shown to be the only local asymptotic needed for the usual pointwise convergence theorems. This is a useful technical contribution for settings (e.g., Sinkhorn-based kernels on probability measures) where the ambient Euclidean metric is unnatural. The paper is self-contained once standard Riemannian geometry and known graph-Laplacian convergence results are granted, and it correctly separates the abstract theorem from the illustrative Sinkhorn family.

major comments (1)
  1. [Sinkhorn examples section / statement of the main estimate] The non-degeneracy hypothesis (that (1/2)Hess_p(D(p,·)) is positive definite at every p, so that g is a genuine Riemannian metric) is left as a global assumption on the compact manifold. This is legitimate for the abstract theorem, but the Sinkhorn examples section only verifies the condition locally or in low-dimensional regimes. If the paper wishes to claim that the convergence result applies to Sinkhorn-based graph Laplacians on the full family under consideration, a clearer delimitation of the known regimes (or an additional argument establishing global non-degeneracy) is needed; otherwise the example should be more carefully caveated so that it does not appear to inherit the theorem unconditionally.
minor comments (4)
  1. [Abstract and Introduction] The abstract and introduction both state the classical Euclidean comparison with a one-sided inequality 0 ≤ d_g^{2} - ∥p-q∥^{2} ≤ K d_g^{4}, while the general result is two-sided |D - d_g^{2}| ≤ K d_g^{4}. A brief remark clarifying that the absolute-value form is the natural one once D is no longer required to dominate the squared geodesic distance would improve readability.
  2. [Proof of the main estimate] Notation for the Hessian Hess_p(D(p,·)) and the induced metric g_p is introduced cleanly, but the subsequent appearance of the exponential map and normal coordinates in the Taylor argument would benefit from an explicit sentence recalling that the odd-order terms vanish by symmetry of D and the fact that p is a critical point of D(p,·).
  3. [Graph Laplacian convergence section] When invoking the “standard integral-operator arguments” for pointwise Laplacian convergence, a precise pointer to the reference (or a short self-contained sketch of the near-diagonal contribution) would help readers who are not already experts in the graph-Laplacian literature.
  4. [Throughout] A few minor typographical inconsistencies appear in the placement of subscripts on Hess and in the spacing around the absolute-value bars in the main estimate; these are easily corrected in copy-editing.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for a careful reading and for the constructive minor-revision recommendation. The single major comment concerns the presentation of the Sinkhorn examples relative to the global non-degeneracy hypothesis of the main theorem. We agree that the current wording can leave the impression that the examples inherit the theorem unconditionally, and we will revise the Sinkhorn section to delimit the known regimes more carefully and to add explicit caveats. The abstract theorem itself is unchanged. Detailed responses follow.

read point-by-point responses
  1. Referee: The non-degeneracy hypothesis (that (1/2)Hess_p(D(p,·)) is positive definite at every p, so that g is a genuine Riemannian metric) is left as a global assumption on the compact manifold. This is legitimate for the abstract theorem, but the Sinkhorn examples section only verifies the condition locally or in low-dimensional regimes. If the paper wishes to claim that the convergence result applies to Sinkhorn-based graph Laplacians on the full family under consideration, a clearer delimitation of the known regimes (or an additional argument establishing global non-degeneracy) is needed; otherwise the example should be more carefully caveated so that it does not appear to inherit the theorem unconditionally.

    Authors: We agree with the referee. The main theorem correctly takes global non-degeneracy of D (i.e., positive-definiteness of (1/2)Hess_p(D(p,·)) at every p) as a standing hypothesis that supplies the Riemannian metric g; that hypothesis is not claimed to be automatic for every smooth symmetric divergence. In the Sinkhorn section we presently verify the Hessian condition only in local charts or in low-dimensional model families, and we do not supply a global non-degeneracy argument for the full manifold-parametrized family under consideration. Consequently the examples cannot be said to inherit the convergence theorem unconditionally. We will revise the Sinkhorn discussion to (i) state explicitly that non-degeneracy is verified only locally / in the regimes already treated, (ii) list those regimes clearly, and (iii) caveat that the pointwise Laplacian convergence result applies to Sinkhorn-based graph Laplacians only on the open sets (or subfamilies) where the Hessian remains positive definite. We will also adjust any surrounding prose that might suggest an unconditional transfer of the theorem. We do not attempt a new global non-degeneracy proof in this revision, as that lies outside the paper’s present scope; the abstract theorem remains unchanged and continues to separate cleanly from the illustrative examples. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: g is induced by D by definition, but the fourth-order comparison and Laplacian limit are derived, not assumed.

full rationale

The paper defines a Riemannian metric g from a smooth symmetric divergence D via g_p := (1/2) Hess_p(D(p,·)) under a non-degeneracy hypothesis, then proves by Taylor expansion in normal coordinates (odd orders vanish by symmetry and criticality on the diagonal, with uniformity on the compact unit sphere bundle) that |D(p,q) − d_g(p,q)^2| ≤ K d_g(p,q)^4. That estimate is then fed into standard integral-operator arguments for pointwise graph-Laplacian convergence as bandwidth tends to zero. Defining g from D and comparing D to the squared geodesic distance of the induced metric is a legitimate differential-geometric theorem, not a self-definitional loop: the comparison and the Laplacian limit are derived rather than built into the definition of D or g. There are no fitted parameters renamed as predictions, no load-bearing self-citations of uniqueness theorems, no ansatz smuggled in via prior author work, and no renaming of a known empirical pattern. Non-degeneracy is left as an explicit hypothesis (including for the Sinkhorn illustrative family), which is legitimate for an if-then theorem and does not create circularity. The derivation is self-contained once standard Riemannian geometry and known graph-Laplacian convergence results are granted.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper rests on standard smooth Riemannian geometry, compactness of M, and the classical pointwise convergence theory for graph Laplacians under Euclidean (or geodesic) kernels. The only paper-specific ingredients are the definition of g from the Hessian of D and the non-degeneracy hypothesis that makes g a Riemannian metric. No free parameters are fitted; no new physical entities are invented. The Sinkhorn family is taken from the existing optimal-transport literature.

assumptions (4)
  • domain assumption M is a smooth, compact, connected Riemannian submanifold of R^d (or an abstract compact manifold).
    Stated in the abstract and used throughout to obtain uniform constants K and to apply standard manifold-learning sampling arguments.
  • ad hoc to paper D is a smooth symmetric divergence that is nondegenerate: Hess_p(D(p,·)) is positive definite for every p, so g_p := (1/2)Hess_p(D(p,·)) is a Riemannian metric.
    This is the central structural hypothesis that replaces the Euclidean embedding; without it the fourth-order comparison fails.
  • domain assumption Standard pointwise convergence theorems for graph Laplacians (bandwidth →0, n→∞ at compatible rates, kernel regularity, sampling density bounded away from 0 and ∞) continue to hold when the kernel is built from a distance that is fourth-order close to geodesic distance.
    The paper invokes and adapts existing results from the manifold-learning literature rather than re-proving them from scratch.
  • standard math Taylor expansion with remainder on a compact Riemannian manifold controls the difference between D and d_g^{2} up to order 4.
    Ordinary multivariable calculus / Riemannian normal-coordinate expansion; used to prove the key estimate.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On the convergence of graph Laplacians with a symmetric divergence." pith.science (2026). https://pith.science/paper/STQR7YKK

@misc{pith2026260705892,
  author       = {Pith},
  title        = {Pith review of: On the convergence of graph Laplacians with a symmetric divergence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/STQR7YKK}},
  note         = {Machine review of arXiv:2607.05892}
}
abstract

When analyzing a manifold learning algorithm for data lying on a smooth, compact, connected Riemannian submanifold $(\mathcal{M}, g)$ of $\mathbb{R}^d$, a key estimate for the geodesic distance $d_g$ is that there exists $K > 0$ such that $0 \leq d_g(p, q)^2 - \|p-q\|^2 \leq K d_g(p, q)^4$ for all $p, q \in \mathcal{M}$. We observe that more generally, when $\mathcal{M}$ is equipped with a smooth symmetric divergence $D$ satisfying a non-degeneracy condition and $g$ is given by $g_p := \frac{1}{2}\mathrm{Hess}_p(D(p, \cdot))$ for all $p \in \mathcal{M}$, there exists $K > 0$ such that $\left| D(p, q) - d_g(p, q)^2 \right| \leq K d_g(p, q)^4$ for all $p, q \in \mathcal{M}$. We demonstrate that this is sufficient for the pointwise convergence of graph Laplacians constructed with $D$ and discuss examples where $D$ is given by the Sinkhorn divergence on a family of probability measures parametrized by a manifold.

Figures

Figures reproduced from arXiv: 2607.05892 by the authors.

Figure 1
Figure 1. Visualization of the approximation error in Proposition [PITH_FULL_IMAGE:figures/full_fig_p021_1.png] view at source ↗
Figure 2
Figure 2. Example of pointwise convergence of the discrete Laplacian for the example in Section [PITH_FULL_IMAGE:figures/full_fig_p021_2.png] view at source ↗
Figure 3
Figure 3. Embedding of samples into R 2 using the first two non-constant eigenvectors v (1), v(2) of (D(εN ,N) ) −1L (εN ,N) for the example in Section 5.1. Each point is colored according to the value of θi . the weight matrix, set εN = 4N −1/3.01 and test using the function f(θ) = sin(θ)+ 1 2 cos(2θ). For this choice of h, m2 as defined in (24) can be computed to be m2 = (2π) m/2 = √ 2π. For convenience, we plot θ 7→ 2volg(… view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Visualization of the approximation error in ( [PITH_FULL_IMAGE:figures/full_fig_p025_4.png]
Figure 5
Figure 5. Figure 5: Examples of pointwise convergence of the discrete Laplacian for the example in Sec [PITH_FULL_IMAGE:figures/full_fig_p026_5.png]
Figure 6
Figure 6. Figure 6: Embedding of samples into R 2 using the first two non-constant eigenvectors v (1), v(2) of (D(εN ,N) ) −1L (εN ,N) (constructed with D = S1) for the example in Section 5.2. Each point is colored according to the value of xi . 26 [PITH_FULL_IMAGE:figures/full_fig_p026_6.png]
Figure 7
Figure 7. Figure 7: Plots of θ 7→ Sbβ(µ0,µθ)−Sβ(µ0,µθ) θ 4 , where Sbβ is the Sinkhorn divergence approximated with Sinkhorn’s algorithm and Sβ(µ0, µθ) is given by the closed-form formula for the example in Section 5.1. 48 [PITH_FULL_IMAGE:figures/full_fig_p048_7.png]
Figure 8
Figure 8. Figure 8: Plots of x 7→ Sbβ(µ1,µx)−Sβ(µ1,µx) (x−1)4 , where Sbβ is the Sinkhorn divergence approximated with Sinkhorn’s algorithm and Sβ(µ1, µx) is given by the closed-form formula for the example in Section 5.2. 49 [PITH_FULL_IMAGE:figures/full_fig_p049_8.png]
Figure 9
Figure 9. Figure 9: Same as Fig [PITH_FULL_IMAGE:figures/full_fig_p050_9.png]
Figure 10
Figure 10. Figure 10: Subtracting the right-hand column from the left-hand column of Fig. [PITH_FULL_IMAGE:figures/full_fig_p051_10.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

89 extracted references · 89 canonical work pages

  1. [1]

    S.-I. Amari. Natural gradient works efficiently in learning.Neural Computation, 10(2):251–276, 1998

  2. [2]

    Amari.Information geometry and its applications, volume 194

    S.-I. Amari.Information geometry and its applications, volume 194. Springer, 2016

  3. [3]

    Ambrosio, N

    L. Ambrosio, N. Gigli, and G. Savaré.Gradient flows: in metric spaces and in the space of probability measures. Springer Science & Business Media, 2008

  4. [4]

    Arias-Castro and W

    E. Arias-Castro and W. Qiao. Embedding distributional data.The Annals of Statistics, 53(2):615–646, 2025

  5. [5]

    Aronszajn

    N. Aronszajn. Theory of Reproducing Kernels.Transactions of the American Mathematical Society, 68(3):337–404, 1950. 28

  6. [6]

    U. M. Ascher and C. Greif.A First Course in Numerical Methods. Society for Industrial and Applied Mathematics, Philadelphia, PA, 2011

  7. [7]

    J. E. Atkins, E. G. Boman, and B. Hendrickson. A spectral algorithm for seriation and the consecutive ones problem.SIAM Journal on Computing, 28(1):297–310, 1998

  8. [8]

    N. Ay, J. Jost, H. Vân Lê, and L. Schwachhöfer.Information geometry, volume 64. Springer, 2017

Show all 89 references
  1. [9]

    J. Bates. The embedding dimension of Laplacian eigenfunction maps.Applied and Computa- tional Harmonic Analysis, 37(3):516–530, 2014

  2. [10]

    Belkin and P

    M. Belkin and P. Niyogi. Laplacian eigenmaps and spectral techniques for embedding and clustering. In T. Dietterich, S. Becker, and Z. Ghahramani, editors,Advances in Neural Infor- mation Processing Systems, volume 14. MIT Press, 2001

  3. [11]

    Belkin and P

    M. Belkin and P. Niyogi. Laplacian eigenmaps for dimensionality reduction and data repre- sentation.Neural computation, 15(6):1373–1396, 2003

  4. [12]

    Belkin and P

    M. Belkin and P. Niyogi. Towards a theoretical foundation for Laplacian-based manifold methods. InInternational Conference on Computational Learning Theory, pages 486–500. Springer, 2005

  5. [13]

    Belkin and P

    M. Belkin and P. Niyogi. Convergence of Laplacian eigenmaps.Advances in Neural Information Processing Systems, 19, 2006

  6. [14]

    Belkin and P

    M. Belkin and P. Niyogi. Towards a theoretical foundation for Laplacian-based manifold methods.Journal of Computer and System Sciences, 74(8):1289–1308, 2008

  7. [15]

    Benamou and Y

    J.-D. Benamou and Y. Brenier. A computational fluid mechanics solution to the Monge- Kantorovich mass transfer problem.Numerische Mathematik, 84(3):375–393, 2000

  8. [16]

    Bérard, G

    P. Bérard, G. Besson, and S. Gallot. Embedding Riemannian manifolds by their heat kernel. Geometric & Functional Analysis GAFA, 4:373–398, 1994

  9. [17]

    Bernstein, V

    M. Bernstein, V. De Silva, J. C. Langford, and J. B. Tenenbaum. Graph approximations to geodesics on embedded manifolds. Technical report, Citeseer, 2000

  10. [18]

    Bigot, R

    J. Bigot, R. Gouet, T. Klein, and A. Lopez. Geodesic PCA in the Wasserstein space by convex PCA. InAnnales de l’Institut Henri Poincaré (B) Probabilités et Statistiques, volume 53, pages 1–26, 2017

  11. [19]

    Calder, N

    J. Calder, N. García Trillos, and M. Lewicka. Lipschitz regularity of graph Laplacians on random data clouds.SIAM Journal on Mathematical Analysis, 54(1):1169–1222, 2022

  12. [20]

    Calder and N

    J. Calder and N. G. Trillos. Improved spectral convergence rates for graph Laplacians on ε-graphs and k-NN graphs.Applied and Computational Harmonic Analysis, 60:123–175, 2022

  13. [21]

    Carlier, L

    G. Carlier, L. Chizat, and M. Laborde. Displacement smoothness of entropic optimal transport. ESAIM: Control, Optimisation and Calculus of Variations, 30:25, 2024. 29

  14. [22]

    K. M. Carter, R. Raich, W. G. Finn, and A. O. Hero III. FINE: Fisher information non- parametric embedding.IEEE Transactions on Pattern Analysis and Machine Intelligence, 31(11):2093–2098, 2009

  15. [23]

    Chen and W

    Y. Chen and W. Li. Optimal transport natural gradient for statistical manifolds with contin- uous sample space.Information Geometry, 3(1):1–32, 2020

  16. [24]

    Cheng and H.-T

    X. Cheng and H.-T. Wu. Convergence of graph Laplacian with kNN self-tuned kernels.Infor- mation and Inference: A Journal of the IMA, 11(3):889–957, 2022

  17. [25]

    Cheng and N

    X. Cheng and N. Wu. Eigen-convergence of Gaussian kernelized graph Laplacian by manifold heat interpolation.Applied and Computational Harmonic Analysis, 61:132–190, 2022

  18. [26]

    Chewi, J

    S. Chewi, J. Niles-Weed, and P. Rigollet.Statistical optimal transport, volume 2364 ofLecture Notes in Mathematics. Springer, Cham, [2025]©2025. École d’Été de Probabilités de Saint- Flour XLIX – 2019

  19. [27]

    R. R. Coifman and S. Lafon. Diffusion maps.Applied and Computational Harmonic Analysis, 21(1):5–30, 2006

  20. [28]

    R. R. Coifman, S. Lafon, A. B. Lee, M. Maggioni, B. Nadler, F. Warner, and S. W. Zucker. Geometric diffusions as a tool for harmonic analysis and structure definition of data: Diffusion maps.Proceedings of the National Academy of Sciences, 102(21):7426–7431, 2005

  21. [29]

    M. Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. In C. Burges, L. Bottou, M. Welling, Z. Ghahramani, and K. Weinberger, editors,Advances in Neural In- formation Processing Systems, volume 26. Curran Associates, Inc., 2013

  22. [30]

    Diepeveen, C

    W. Diepeveen, C. Esteve-Yagüe, J. Lellmann, O. Öktem, and C.-B. Schönlieb. Riemannian geometry for efficient analysis of protein dynamics data.Proceedings of the National Academy of Sciences, 121(33):e2318951121, 2024

  23. [31]

    M. P. Do Carmo and J. Flaherty Francis.Riemannian geometry, volume 2. Springer, 1992

  24. [32]

    D. B. Dunson, H.-T. Wu, and N. Wu. Spectral convergence of graph Laplacian and heat kernel reconstruction inL ∞ from random samples.Applied and Computational Harmonic Analysis, 55:282–336, 2021

  25. [33]

    Fefferman, S

    C. Fefferman, S. Mitter, and H. Narayanan. Testing the manifold hypothesis.Journal of the American Mathematical Society, 29(4):983–1049, 2016

  26. [34]

    Feydy, T

    J. Feydy, T. Séjourné, F.-X. Vialard, S.-i. Amari, A. Trouvé, and G. Peyré. Interpolating between Optimal Transport and MMD using Sinkhorn Divergences. InThe 22nd International Conference on Artificial Intelligence and Statistics, pages 2681–2690. PMLR, 2019

  27. [35]

    Flamary, N

    R. Flamary, N. Courty, A. Gramfort, M. Z. Alaya, A. Boisbunon, S. Chambon, L. Chapel, A. Corenflos, K. Fatras, N. Fournier, L. Gautheron, N. T. Gayraud, H. Janati, A. Rakotoma- monjy, I. Redko, A. Rolet, A. Schutz, V. Seguy, D. J. Sutherland, R. Tavenard, A. Tong, and T. Vayer...

  28. [36]

    Flamary, C

    R. Flamary, C. Vincent-Cuaz, N. Courty, A. Gramfort, O. Kachaiev, H. Quang Tran, L. David, C. Bonet, N. Cassereau, T. Gnassounou, E. Tanguy, J. Delon, A. Collas, S. Mazelet, L. Chapel, T. Kerdoncuff, X. Yu, M. Feickert, P. Krzakala, T. Liu, and E. Fernandes Montesuma. POT Pyth...

  29. [37]

    Errorestimatesforspectralconvergence of the graph Laplacian on random geometric graphs toward the Laplace–Beltrami operator

    N.GarcíaTrillos, M.Gerlach, M.Hein, andD.Slepčev. Errorestimatesforspectralconvergence of the graph Laplacian on random geometric graphs toward the Laplace–Beltrami operator. Foundations of Computational Mathematics, 20(4):827–887, 2020

  30. [38]

    Genevay, L

    A. Genevay, L. Chizat, F. Bach, M. Cuturi, and G. Peyré. Sample Complexity of Sinkhorn Divergences. InThe 22nd international conference on artificial intelligence and statistics, pages 1574–1583. PMLR, 2019

  31. [39]

    Genevay, G

    A. Genevay, G. Peyré, and M. Cuturi. Learning Generative Models with Sinkhorn Divergences. InInternational Conference on Artificial Intelligence and Statistics, pages 1608–1617. PMLR, 2018

  32. [40]

    Gonzalez-Sanz, J.-M

    A. Gonzalez-Sanz, J.-M. Loubes, and J. Niles-Weed. Weak limits of entropy regularized optimal transport; potentials, plans and divergences.arXiv preprint arXiv:2207.07427, 2024

  33. [41]

    K. Hamm, N. Henscheid, and S. Kang. Wassmap: Wasserstein isometric mapping for image manifold learning.SIAM Journal on Mathematics of Data Science, 5(2):475–501, 2023

  34. [42]

    K. Hamm, C. Moosmüller, B. Schmitzer, and M. Thorpe. Manifold Learning in Wasserstein Space.SIAM Journal on Mathematical Analysis, 57(3):2983–3029, 2025

  35. [43]

    Hardion and H

    M. Hardion and H. Lavenant. Gradient flows of potential energies in the geometry of sinkhorn divergences.arXiv preprint arXiv:2511.14278, 2025

  36. [44]

    Hein, J.-Y

    M. Hein, J.-Y. Audibert, and U. Von Luxburg. From graphs to manifolds–weak and strong pointwise consistency of graph Laplacians. InInternational Conference on Computational Learning Theory, pages 470–485. Springer, 2005

  37. [45]

    Hörmander.The analysis of linear partial differential operators

    L. Hörmander.The analysis of linear partial differential operators. I. Classics in Mathematics. Springer-Verlag, Berlin, 2003. Distribution theory and Fourier analysis, Reprint of the second (1990) edition [Springer, Berlin; MR1065993 (91m:35001a)]

  38. [46]

    Kileel, A

    J. Kileel, A. Moscovich, N. Zelesko, and A. Singer. Manifold learning with arbitrary norms. Journal of Fourier Analysis and Applications, 27(5):82, 2021

  39. [47]

    Kullback.Information theory and statistics

    S. Kullback.Information theory and statistics. Courier Corporation, 1997

  40. [48]

    Lang.Fundamentals of differential geometry, volume 191 ofGraduate Texts in Mathematics

    S. Lang.Fundamentals of differential geometry, volume 191 ofGraduate Texts in Mathematics. Springer-Verlag, New York, 1999

  41. [49]

    Lavenant, J

    H. Lavenant, J. Luckhardt, G. Mordant, B. Schmitzer, and L. Tamanini. The Riemannian ge- ometry of Sinkhorn divergences.Annales de l’Institut Henri Poincaré C, Analyse non linéaire, Oct. 2025

  42. [50]

    Li and G

    W. Li and G. Montúfar. Natural gradient via optimal transport.Information Geometry, 1(2):181–214, 2018. 31

  43. [51]

    Li and J

    W. Li and J. Zhao. Wasserstein information matrix.Information Geometry, 6(1):203–255, 2023

  44. [52]

    J. Lott. Some Geometric Calculations on Wasserstein Space.Commun. Math. Phys, 277:423– 437, 2008

  45. [53]

    P. Y. Lu, R. Dangovski, and M. Soljačić. Discovering conservation laws using optimal transport and manifold learning.Nature Communications, 14(1):4744, 2023

  46. [54]

    Luise, A

    G. Luise, A. Rudi, M. Pontil, and C. Ciliberto. Differential properties of Sinkhorn approx- imation for learning with Wasserstein distance.Advances in Neural Information Processing Systems, 31, 2018

  47. [55]

    Mena and J

    G. Mena and J. Niles-Weed. Statistical bounds for entropic optimal transport: sample com- plexity and the central limit theorem.Advances in Neural Information Processing Systems, 32, 2019

  48. [56]

    Universalkernels.Journal of Machine Learning Research, 7(12), 2006

    C.A.Micchelli, Y.Xu, andH.Zhang. Universalkernels.Journal of Machine Learning Research, 7(12), 2006

  49. [57]

    Mishne, R

    G. Mishne, R. Talmon, R. Meir, J. Schiller, M. Lavzin, U. Dubin, and R. R. Coifman. Hierar- chical coupled-geometry analysis for neuronal structure and activity pattern discovery.IEEE Journal of Selected Topics in Signal Processing, 10(7):1238–1253, 2016

  50. [58]

    Muandet, K

    K. Muandet, K. Fukumizu, B. Sriperumbudur, and B. Schölkopf. Kernel mean embedding of distributions: A review and beyond.Foundations and Trends®in Machine Learning, 10(1- 2):1–141, 2017

  51. [59]

    Nadler, S

    B. Nadler, S. Lafon, R. R. Coifman, and I. G. Kevrekidis. Diffusion maps, spectral clustering andreactioncoordinatesofdynamicalsystems.Applied and Computational Harmonic Analysis, 21(1):113–127, 2006

  52. [60]

    M. Nutz. Introduction to entropic optimal transport.Lecture notes, Columbia University, 2022

  53. [61]

    M. C. A. Oliver, M. Roberts, C.-B. Schönlieb, and M. Thorpe. Laplace learning in Wasserstein space.arXiv preprint arXiv:2511.13229, 2025

  54. [62]

    F. Otto. The geometry of dissipative evolution equations: the porous medium equation.Com- munications in Partial Differential Equations, 26(1-2):101–174, 2001

  55. [63]

    Peyré and M

    G. Peyré and M. Cuturi. Computational optimal transport: With applications to data science. Foundations and Trends®in Machine Learning, 11(5-6):355–607, 2019

  56. [64]

    C. R. Rao. Differential metrics in probability spaces.Differential geometry in statistical infer- ence, 10:217–240, 1987

  57. [65]

    Rumpf and B

    M. Rumpf and B. Wirth. Variational time discretization of geodesic calculus.IMA Journal of Numerical Analysis, 35(3):1011–1046, 2015

  58. [66]

    Seguy and M

    V. Seguy and M. Cuturi. Principal geodesic analysis for probability measures under the optimal transport metric.Advances in Neural Information Processing Systems, 28, 2015. 32

  59. [67]

    Z. Shen, Z. Wang, A. Ribeiro, and H. Hassani. Sinkhorn natural gradient for generative models. Advances in Neural Information Processing Systems, 33:1646–1656, 2020

  60. [68]

    Shirdhonkar and D

    S. Shirdhonkar and D. W. Jacobs. Approximate earth mover’s distance in linear time. In2008 IEEE Conference on Computer Vision and Pattern Recognition, pages 1–8. IEEE, 2008

  61. [69]

    Simon-Gabriel and B

    C.-J. Simon-Gabriel and B. Schölkopf. Kernel distribution embeddings: Universal kernels, characteristic kernels and kernel metrics on distributions.Journal of Machine Learning Re- search, 19(44):1–29, 2018

  62. [70]

    Simon-Gabriel and B

    C.-J. Simon-Gabriel and B. Schölkopf. Kernel distribution embeddings: Universal kernels, characteristic kernels and kernel metrics on distributions.arXiv preprint arXiv:1604.05251, 2019

  63. [71]

    A. Singer. From graph to manifold Laplacian: The convergence rate.Applied and Computa- tional Harmonic Analysis, 21(1):128–134, 2006

  64. [72]

    A. Singer. Wilson statistics: derivation, generalization and applications to electron cryomi- croscopy.Foundations of Crystallography, 77(5):472–479, 2021

  65. [73]

    Singer and H.-T

    A. Singer and H.-T. Wu. Spectral convergence of the connection Laplacian from random samples.Information and Inference: A Journal of the IMA, 6(1):58–123, 2017

  66. [74]

    Arelationshipbetweenarbitrarypositivematricesanddoublystochasticmatrices

    R.Sinkhorn. Arelationshipbetweenarbitrarypositivematricesanddoublystochasticmatrices. The Annals of Mathematical Statistics, 35(2):876–879, 1964

  67. [75]

    Smola, A

    A. Smola, A. Gretton, L. Song, and B. Schölkopf. A Hilbert space embedding for distributions. InInternational Conference on Algorithmic Learning Theory, pages 13–31. Springer, 2007

  68. [76]

    Sriperumbudur, K

    B. Sriperumbudur, K. Fukumizu, and G. Lanckriet. On the relation between universality, characteristic kernels and RKHS embedding of measures. InProceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, pages 773–780. JMLR Work- shop and ...

  69. [77]

    Steinwart

    I. Steinwart. On the influence of the kernel on the consistency of support vector machines. Journal of Machine Learning Research, 2(Nov):67–93, 2001

  70. [78]

    Steinwart and A

    I. Steinwart and A. Christmann.Support vector machines. Springer Science & Business Media, 2008

  71. [79]

    J. B. Tenenbaum, V. d. Silva, and J. C. Langford. A global geometric framework for nonlinear dimensionality reduction.Science, 290(5500):2319–2323, 2000

  72. [80]

    N. G. Trillos, C. Li, and R. Venkatraman. Minimax Rates for the Estimation of Eigenpairs of Weighted Laplace-Beltrami Operators on Manifolds.arXiv preprint arXiv:2506.00171, 2025

  73. [81]

    N. G. Trillos, A. Little, D. McKenzie, and J. M. Murphy. Fermat distances: Metric approxima- tion, spectral convergence, and clustering algorithms.Journal of Machine Learning Research, 25(176):1–65, 2024

  74. [82]

    Villani et al.Optimal transport: old and new, volume 338

    C. Villani et al.Optimal transport: old and new, volume 338. Springer, 2009. 33

  75. [83]

    M. Wahl. A kernel-based analysis of Laplacian Eigenmaps.arXiv preprint arXiv:2402.16481, 2024

  76. [84]

    Warren, A

    A. Warren, A. Afanassiev, F. Kobayashi, Y.-H. Kim, and G. Schiebinger. Principal curves in metric spaces and the space of probability measures.arXiv preprint arXiv:2505.04168, 2025

  77. [85]

    H. Whitney. Functions differentiable on the boundaries of regions.Annals of Mathematics, 35(3):482–485, 1934

  78. [86]

    Xu and A

    L. Xu and A. Singer. Manifold learning in metric spaces.Applied and Computational Harmonic Analysis, page 101813, 2025

  79. [87]

    Zelesko, A

    N. Zelesko, A. Moscovich, J. Kileel, and A. Singer. Earthmover-based manifold learning for analyzing molecular conformation spaces. In2020 IEEE 17th International Symposium on Biomedical Imaging (ISBI), pages 1715–1719. IEEE, 2020. A Proofs for Section 3 Proof of Proposition 3...

  80. [88]

    Since we assumed thatDis continuous andD(p, q) = 0if and only ifp=q, D∗ := inf p,q∈M:d g(p,q)≥r D(p, q) is strictly greater than 0

    Approximation ofG εfwith an integral over a geodesic ball.For anyp∈ M, ε∈(0, r)andf∈C 3(M),we can approximateG εf(p)with an integral over the geodesic ball B(p, r)of radiusr: Gεf(p)− 1 εm/2 Z B(p,r) Kε(p, q)f(q)dV g(q) ≤ volg(M \B(p, r)) sup q∈M\B(p,r) |Kε(p, q)f(q)| εm/2 . Si...

  81. [89]

    We now approximateε −m/2R B(p,r) Kε(p, q)f(q)dV g(q)withm 0f(p) + εm2 2 (ω(p)f(p)−∆ gf(p)), where ω:M →Ris to be determined

    Taylor expansions and integration in normal coordinates.Fix anyp∈ M. We now approximateε −m/2R B(p,r) Kε(p, q)f(q)dV g(q)withm 0f(p) + εm2 2 (ω(p)f(p)−∆ gf(p)), where ω:M →Ris to be determined. For this, we use normal coordinates aroundpand Taylor expan- sions. For concretenes...

Pith tools

Reviewed July 8, 2026 · model on record in the stance chip above.