REVIEW 2 major objections 5 minor 38 references
Dimension-free Convergence Rate in Sliced Wasserstein Distance for Empirical Measures of Markov Processes
T0 review · 2 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read The paper claims that the sliced Wasserstein distance W_{p,λ}, averaging one-dimensional projections over a full-support measure on the dual sphere, is topologically stronger than finite-dimensional-distribution convergence on Banach spaces
desk verdict Sliced-Wasserstein rates on Banach spaces hit a real snag: the central topological theorem rests on a false uniform estimate, though the quantitative dimension-free bounds may survive. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The sliced distance W_{p,λ}(μ,ν):=(∫_S W_p(μ^θ,ν^θ)^p λ(dθ))^{1/p}, with μ^θ the push-forward of μ under the bounded linear functional θ. It turns a difficult infinite-dimensional transport problem into a family of one-dimensional transport problems that can be attacked with empirical-process and spectral-gap methods. The proof of the topological characterization relies on Fourier transforms and Kantorovich duality to relate W_{p,λ} to finite-dimensional distributions; the rate theorems use ergodicity (exponential in Assumption (A1), subexponential in (4.1)) together with Poincaré and heat-kernel bounds for the projected one-dimensional laws.
What would settle it
Evaluate the sup in inequality (2.6) for B=ℓ^2 and ξ a nonzero coordinate functional: for θ≠θ_ξ, sup_{x∈B} |e^{i∥ξ∥θ(x)} - e^{iξ(x)}| = 2, not ≤ε/4, because ∥ξ∥(θ−θ_ξ) is an unbounded linear functional. Exhibiting a sequence θ_k→θ_ξ with λ(D_ε)>0 but W_p distributions converging would make the claimed implication concrete to test.
Extended reading notes
Core claim
The central claim is Theorem 2.1(1): for λ with full support on the dual unit sphere, W_{p,λ}(μ_n,μ)→0 holds if and only if μ_n converges to μ in finite-dimensional distributions and the sliced p-th moments satisfy the uniform integrability condition (2.1). The paper further proves that in finite dimensions W_{p,λ} is topologically equivalent to the usual Wasserstein distance, that W_{p,λ} is complete under a moment bound on a separable Hilbert space, and that it gives dimension-free convergence rates for empirical measures of ergodic Markov processes, including sharp t^{-1}/n^{-1} rates for W_{2,λ} under explicit spectral assumptions.
Load-bearing premise
The theorem presupposes that a tiny change in the projection direction changes the exponential e^{iξ(x)} by at most ε/4 uniformly over all x∈B; for unbounded linear functionals on an infinite-dimensional Banach space this uniform bound is unavailable once the directions differ at all.
Editorial extensions
If this is right
- If W_{p,λ} defines a metric equivalent to W_p on R^d and dimension-free on infinite-dimensional spaces, empirical-measure convergence rates no longer degrade with the ambient dimension.
- For exponentially ergodic Markov processes with invariant measure having enough moments, the bound (3.11) gives E_μ[W_{p,λ}(μ_t,μ)^{2p}] ≤ C/t, and the analogous discrete-time bound with C/n.
- Under the stronger assumptions (A2), the W_{2,λ} rate is sharp: t E_μ[W_{2,λ}(μ_t,μ)^2] stays bounded away from zero and above, giving the t^{-1} order as the exact rate.
- For Markov processes with subexponential ergodicity, the rates remain t^{-1} and n^{-1} provided the decay profile ξ(t) is integrable, covering heavy-tailed potentials.
- The SPDE application shows the framework handles infinite-dimensional state spaces, so the empirical measures of solutions to partially dissipative semilinear SPDEs converge to the invariant measure at the same dimension-free rate.
Reading between the lines
- The rate theorems in Sections 3–5 rely on one-dimensional projection bounds and would remain meaningful even if the topological equivalence theorem requires modification; the equivalence is used for motivation and for identifying the limit, not for the rate estimates themselves.
- The non-completeness example in Remark 1.1 suggests that for actual simulation or inference, one would likely work inside a completion of the Banach space or restrict attention to measures supported on a smaller subspace, so the practical use of W_{p,λ} may not require a Polish space.
- A numerical comparison of W_{p,λ} versus W_p for the empirical measure of a dissipative SPDE projected onto high-dimensional subspaces would test whether the claimed dimension-free rates translate into observable accuracy gains in practice.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a sliced Wasserstein distance W_{p,λ} on the space of probability measures with finite p-th moment on a Banach space B, using a full-support probability λ on the unit sphere of the dual space B*. Theorem 2.1 claims that W_{p,λ}-convergence is equivalent to convergence in finite-dimensional distributions plus a uniform p-th integrability condition on the projected tail masses, and that in Hilbert spaces W_{p,λ}-Cauchy sequences with bounded p-th moments converge. The main body uses this framework to derive dimension-free convergence rates for empirical measures of Markov processes: exponential ergodicity is treated in Theorems 3.2 and 3.4, a subexponential ergodicity extension in Theorem 4.1, and an application to partially dissipative SPDEs in Theorem 5.1. The paper also claims sharpness for some of these rates.
Significance. If the topological characterization in Theorem 2.1 were correct, the paper would provide a genuine infinite-dimensional extension of sliced Wasserstein distances and useful dimension-free bounds for Markov-chain simulation. The quantitative estimates are nontrivial and appear to be independent of dimension; the SPDE application is concrete. However, the entire advertised framework rests on Theorem 2.1(1), which is false. The paper's central claim that W_{p,λ} convergence is stronger than finite-dimensional-distribution convergence is contradicted by a counterexample, and the proof relies on an invalid uniform estimate. The rate theorems may still be interesting as estimates for the pseudometric W_{p,λ}, but the paper's stated interpretation and topological conclusions cannot be accepted as they stand.
major comments (2)
- [Theorem 2.1(1), Eq. (2.6)] The proof of Theorem 2.1(1)(a) is invalid. For θ ∈ D_ε with θ ≠ θ_ξ, the linear functional ξ − ∥ξ∥_* θ is nonzero and hence unbounded on the infinite-dimensional Banach space B. Therefore the left-hand side of (2.6) is not uniformly bounded by ε/4 in x ∈ B; for any such θ one can choose x making the exponential difference arbitrarily close to 2. Consequently the lower bound in (2.7) does not follow. The theorem is not merely unproved but false. Let B = ℓ^2, let {v_m} be a dense sequence of finite-support unit vectors enumerated so that v_{m,n} ≠ 0 implies m ≥ 2n, and set λ = Σ_m 2^{-m} δ_{v_m}. This λ has full support on the unit sphere, and ∫ θ_n^2 λ(dθ) ≤ 2^{-2n+1}. Take μ_n = δ_{n e_n} and μ = δ_0. Then W_{2,λ}(μ_n,μ)^2 = n^2 ∫ θ_n^2 λ(dθ) → 0. But for ξ = (n^{-3/4}) ∈ ℓ^2, the characteristic function of μ_n at ξ is exp(i n ξ_n) = exp(i n^{1/4}), which does not converge to 1. Hence μ_
- [Theorem 2.1(2)(3)] The failure of (2.6) propagates to the rest of Theorem 2.1. The Fourier-Cauchy argument in the proof of Theorem 2.1(2) explicitly invokes 'the argument leading to (2.7)', and the proof of Theorem 2.1(3) uses Theorem 2.1(1) to obtain weak convergence from W_{p,λ} convergence. Both steps are therefore unsupported. The counterexample above also shows that W_{p,λ} does not control convergence of cylindrical Fourier transforms uniformly over B*, so the Bochner-Minlos route cannot work as written. While the finite-dimensional part Theorem 2.1(3) may be true by other arguments, it is not established by the supplied proof. More importantly, the advertised property that W_{p,λ} is topologically stronger than finite-dimensional-distribution convergence is false: the distance can tend to zero while the measures are point masses moving macroscopically far apart in B. The motivation that this distanc
minor comments (5)
- [Introduction] Typos: 'Lispchitz' should be 'Lipschitz'; 'machining learning' should be 'machine learning'.
- [Proof of Theorem 2.1] The notation S^d and Sd appears in several places where S is clearly intended, e.g., in equations (2.8) and the surrounding text. Since S already denotes the unit sphere, the exponent d is confusing.
- [Remark 1.1] Remark 1.1 appears after Theorem 2.1 and should be renumbered as Remark 2.1 or similar.
- [Theorem 3.2] Near the end of the proof, 'Then (5.9) holds' should refer to the displayed equation of the theorem, presumably (3.5).
- [References] Reference [34] is missing a comma between the author's name and the title; the reference list should be checked for formatting consistency.
Circularity Check
No circular derivation: the dimension-free rates follow from mixing/spectral assumptions; only minor non-load-bearing self-citations are present.
full rationale
The paper's main estimates (Theorems 3.2, 3.4, 4.1) are upper bounds on E[W_{p,λ}(μ_t,μ)^{2p}]. Their constants come from the explicit ergodicity assumptions (A1)/(4.1) and from the spectral/heat-kernel quantities α(θ), γ(θ,t) in (A2); they are not fitted to the t^{-1} or n^{-1} rates, and the sharpness claim in Remark 3.1 is anchored by an external CLT lower bound [38]. Theorem 2.1 is a topological characterization; even if the reader's objection to estimate (2.6) is correct, that is a proof error, not a reduction of the conclusion to its own inputs. The paper cites several works by the same author ([27], [32], [33], [36]), but these are used for parameter-free supporting lemmas (heat-kernel bounds, a coupling construction) or contextual asymptotics whose assumptions do not include the target rates. Thus there is no fitted-input-called-prediction, no ansatz-smuggling, and no self-citation chain forcing the central claims. The score of 2 reflects the presence of these minor self-citations, not an actual circular step.
Assumptions & free parameters
assumptions (4)
- domain assumption (A1): Exponential convergence of the Markov semigroup in L^2(μ): ||P_t − μ||_{L^2(μ)} ≤ c0 e^{-κ0 t}.
- ad hoc to paper Existence of a probability measure λ on the unit sphere S of B* with full support.
- domain assumption (A2): For each θ, the projected one-dimensional diffusion has a Poincaré inequality, heat-kernel moment bound, and heat-kernel trace bound (3.15)-(3.17).
- standard math The finite-dimensional result [3, Proposition 7.4] bounding 1D empirical Wasserstein distance by integrated sup-differences of CDFs.
Cite this review
Pith. "Pith review of Dimension-free Convergence Rate in Sliced Wasserstein Distance for Empirical Measures of Markov Processes." pith.science (2026). https://pith.science/paper/WURQZTD7
@misc{pith2026260720150,
author = {Pith},
title = {Pith review of: Dimension-free Convergence Rate in Sliced Wasserstein Distance for Empirical Measures of Markov Processes},
year = {2026},
howpublished = {\url{https://pith.science/paper/WURQZTD7}},
note = {Machine review of arXiv:2607.20150}
}
abstract
To derive dimension-free convergence rates of empirical measures for Markov processes on a Banach space, we adopt the sliced Wasserstein distance (SW distance) induced by a probability measure with full support on the unit ball of the dual space. This distance is topologically stronger than the convergence in finite-dimensional distributions, and is topologically equivalent to the Wasserstein distance when the Banach space is finite-dimensional. Under this distance, we derive dimension-free convergence rates for the empirical measures of ergodic Markov processes on $\BB$, which can be sharp as illustrated by concrete examples. The study provides an efficient way to simulate infinite-dimensional distributions using sample trajectories of Markov processes, so that the $``$curse of dimensionality" appearing to the classical Wasserstein distance is avoided. The main results apply to a broad class of infinite-dimensional models, and are illustrated by partially dissipative SPDEs in the end of the paper.
Reference graph
Works this paper leans on
-
[1]
Ambrosio, F
L. Ambrosio, F. Stra, D. Trevisan,A PDE approach to a 2-dimensional matching prob- lem,Probab. Theory Related Fields 173(2019), 433-477
2019
-
[2]
Bakry, I
D. Bakry, I. Gentil, M. Ledoux,Analysis and Geometry of Markov Diffusion Operators, Springer, Berlin, 2014
2014
-
[3]
Bobkov, M
S. Bobkov, M. Ledoux,One-dimensional empirical measures, order statistics, and Kan- torovich transport distances,Mem. Amer. Math. Soc. 261(2019), no. 1259, v+126 pp
2019
-
[4]
Bonneel, J
N. Bonneel, J. Rabin, G. Peyr´ e, H. Pfister,Sliced and radon Wasserstein barycenters of measures,Journal of Mathematical Imaging and Vision 51(2015), 22-45
2015
-
[5]
Bonnotte,Unidimensional and Evolution Methods for Optimal Transportation,PhD thesis, Paris-Sud University, 2013
N. Bonnotte,Unidimensional and Evolution Methods for Optimal Transportation,PhD thesis, Paris-Sud University, 2013
2013
-
[6]
Chen,Explicit bounds of the first eigenvalue,Sci
M.-F. Chen,Explicit bounds of the first eigenvalue,Sci. Chin. A. 43(2000), 1051-1059
2000
-
[7]
J. R. Correa, M. Romero,On the asymptotic behavior of the expectation of the maximum of i.i.d. random variables,Operat. Research Letters 49(2021), 785-786
2021
-
[8]
Da Prato, J
G. Da Prato, J. Zabczyk,Stochastic Equations In Infinite Dimensions,Cambridge Uni- versity Press, Cambridge, 1992
1992
Show all 38 references
-
[9]
Deshpande, Y.-T
I. Deshpande, Y.-T. Hu, R. Sun, A. Pyrros, N. Siddiqui, S. Koyejo, Z. Zhao, D. Forsyth, A.G. Schwing,Max-sliced Wasserstein distance and its use for GANs,Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 10648- 10656
2019
-
[10]
Fournier, A
N. Fournier, A. Guillin,On the rate of convergence in Wasserstein distance of the empir- ical measure,Probab. Theory Relat. Fields 162(2015), 707-738
2015
-
[11]
R. Han, C. Rush, J. Wiesel,Max-sliced Wasserstein concneration and uniform ratio bounds of empirical measures on RKHS,arXiv:2405.13153. 38
-
[12]
Huesmann, M.Goldman, D
M. Huesmann, M.Goldman, D. Trevisan,Asymptotics for random quadratic transporta- tion costs,arXiv:2409.08612
-
[13]
Huesmann, F
M. Huesmann, F. Mattesini, D. Trevisan,Wasserstein asympototics for the empirical measure of fractional Brownian motion on a flat torus,Stoch. Proc. Appl. 155(2023), 1-26
2023
-
[14]
Kolouri, K
S. Kolouri, K. Nadjahi, U. Simsekli, R. Badeau, G. Rohde,Generalized sliced Wasserstein distances,Advances in Neural Information Processing Systems 32(2019), 261-272
2019
-
[15]
Ledoux,On optimal matching of Gaussian samples,In: Zap
M. Ledoux,On optimal matching of Gaussian samples,In: Zap. Nauchn. Sem. S.- Peterburg. Otdel. Mat. Inst. Steklov. (POMI) vol. 457. In: Veroyatnost Stat. 25(2017), 226-264
2017
-
[16]
Ledoux,Optimal matching of random samples and rates of convergence of empirical measures,Lecture Notes in Math., 2313
M. Ledoux,Optimal matching of random samples and rates of convergence of empirical measures,Lecture Notes in Math., 2313. Springer, Cham, 2023, 615-627
2023
-
[17]
Ledoux, J.-X
M. Ledoux, J.-X. Zhu,On optimal matching of Gaussian samples III,Probab. Math. Statist. 41(2021), 237-265
2021
-
[18]
H. Li, B. Wu,Wasserstein convergence rates for empirical measures of subordinated processes on noncompact manifolds,J. Theor. Probab. 36(2023), 1243-1268
2023
-
[19]
Manole, S
T. Manole, S. Balakrishnan, L. Wasserman,Minimax confidence intervals for the Sliced Wasserstein distance,Electronic J. Statistics 16(2022), 2252-2345
2022
-
[20]
J. L. M. Olea, C. Rush, A. Velez, J. Wiesel,The out-of-sample prediction error of the square-root-LASSO and related estimators,arXiv:2211.07608
-
[21]
Cuturi,Subspace robust Wasserstein distance,International conference on machine learning, PMLR, 2019, pp
F.-P., Paty, M. Cuturi,Subspace robust Wasserstein distance,International conference on machine learning, PMLR, 2019, pp. 5072-5081
2019
-
[22]
International Conference on Scale Space and Variational Methods in Computer Vision
J. Rabin, G. Peyr´ e, J. Delon, M. Bernot,Wasserstein barycenter and its application to texture mixing,In “International Conference on Scale Space and Variational Methods in Computer Vision”, 2011, 435-446. Springer
2011
-
[23]
R¨ ockner, F.-Y
M. R¨ ockner, F.-Y. Wang,Weak Poincar´ e inequalities and convergence rates of Markov semigroups,J. Funct. Anal. 185(2001), 564-603
2001
-
[24]
R¨ uschendorf,The Wasserstein distance and approximation theorems,Z
L. R¨ uschendorf,The Wasserstein distance and approximation theorems,Z. Wahrsch. Verw. Gebiete 70(1985), 117-129
1985
-
[25]
Talagrand,Scaling and non-standard matching theorems,Comptes Rendus Acad
M. Talagrand,Scaling and non-standard matching theorems,Comptes Rendus Acad. Sci. Paris, Math. 356(2018), 692-695
2018
-
[26]
Trevisan, F.-Y
D. Trevisan, F.-Y. Wang,J.-X. Zhu,Wasserstein asymptotics for empirical measures of diffusions on four dimensional closed manifolds,Electron. Commun. Probab. 30 (2025), Paper No. 68, 13 pp. 39
2025
-
[27]
Wang,Logarithmic Sobolev inequalities on noncompact Riemannian manifolds, Probability Theory Relat
F.-Y. Wang,Logarithmic Sobolev inequalities on noncompact Riemannian manifolds, Probability Theory Relat. Fields 109(1997), 417-424
1997
-
[28]
Wang,Harnack Inequality for Stochastic Partial Differential Equations,Math
F.-Y. Wang,Harnack Inequality for Stochastic Partial Differential Equations,Math. Brief. Springer, 2013
2013
-
[29]
Wang,Precise limit in Wasserstein distance for conditional empirical measures of Dirichlet diffusion processes,J
F.-Y. Wang,Precise limit in Wasserstein distance for conditional empirical measures of Dirichlet diffusion processes,J. Funct. Anal. 280(2021), 108998, 23pp
2021
-
[30]
Wang,Wasserstein convergence rate for empirical measures on noncompact mani- folds,Stoch
F.-Y. Wang,Wasserstein convergence rate for empirical measures on noncompact mani- folds,Stoch. Proc. Appl. 144(2022), 271–287
2022
-
[31]
Wang,Convergence in Wasserstein distance for empirical measures of Dirichlet diffusion processes on manifolds,J
F.-Y. Wang,Convergence in Wasserstein distance for empirical measures of Dirichlet diffusion processes on manifolds,J. Eur. Math. Soc. 25(2023), 3695-3725
2023
-
[32]
Wang,Convergence in Wasserstein distance for empirical measures of semilinear SPDEs,Ann
F.-Y. Wang,Convergence in Wasserstein distance for empirical measures of semilinear SPDEs,Ann. Appl. Probab. 33(2023), 70–84
2023
-
[33]
Wang,Wasserstein convergence rate for empirical measures of Markov processes, Appl
F.-Y. Wang,Wasserstein convergence rate for empirical measures of Markov processes, Appl. Math. Opt. 92(2025), Paper No. 4, 41 pp
2025
-
[34]
WangConvergence in Wasserstein distance for empirical measures of non- symmetric subordinated diffusion processes,Stoch
F.-Y. WangConvergence in Wasserstein distance for empirical measures of non- symmetric subordinated diffusion processes,Stoch. Proc. Appl. (2026)
2026
-
[35]
F.-Y. Wang, P. Ren,Distribution Dependent Stochastic Differential Equations,World Scientific, 2025
2025
-
[36]
F.-Y. Wang, B. Wu, J.-X. Zhu,SharpL q-convergence Rate inp-Wasserstein distance for empirical measures of diffusion processes,Stoch. Proc. Appl. 195(2026), 104869
2026
-
[37]
Wang, J.-X
F.-Y. Wang, J.-X. Zhu,Limit Theorems in Wasserstein Distance for Empirical Measures of Diffusion Processes on Riemannian Manifolds,Ann. l’Inst. H. Poinc. Probab. Statist. 59(2023), 437-475
2023
-
[38]
Wu,Moderate deviations of dependent random variables related to CLT,Ann
L. Wu,Moderate deviations of dependent random variables related to CLT,Ann. Probab. 23(1995), 420-445. 40
1995
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.