REVIEW 7 minor 30 references
Wasserstein Gradient Flows of MMD Functionals with Distance Kernels under Sobolev Regularization
T0 review · 0 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read By adding a Sobolev H1 regularizer to the quantile formulation, this paper proves that one-dimensional MMD gradient flows with both positive and negative distance kernels exist—uniquely for the negative kernel and as generalized…
desk verdict Solid existence theory for regularized 1D MMD flows, with a repairable gap in the negative-kernel reduction that does not touch the main theorems. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying device is the isometric embedding of $P_2(\mathbb{R})$ into the cone $C(0,1)\subset L^2(0,1)$ of quantile functions, combined with the H1 Sobolev regularizer. The regularizer is $F_H(u)=\frac{\lambda}{2}\int_0^1 |u'(s)|^2\,ds$ (infinite outside $H^1(0,1)$), and for the negative kernel it is shifted to $F_{H,\nu}(u)=F_H(u-Q_\nu)$ so that the Neumann boundary conditions are prescribed by the target quantile; its subdifferential is $-\lambda u''$ with those boundary conditions. Adding $I_{C(0,1)}$ forces the flow to stay in the quantile cone. For the negative kernel the regularized functional is convex, so maximal monotone operator theory applies; for the positive kernel the compact Sobolev embedding $H^1(0,1)\hookrightarrow L^2(0,1)$ gives compact sublevels in the minimizing-movement scheme, and convexity of the pieces supplies the chain rule needed for the limiting-subdifferential formulation.
What would settle it
Run the negative-kernel Dirac-to-Dirac example (initial $\delta_{-1}$, target $\delta_0$) with the implicit Euler scheme (25) and compare the pre-contact phase with the explicit heat-equation solution (19); then track the free boundary $\beta(t)$ in the Stefan phase. If, for some $\lambda>0$ and small step size $\tau$, the numerical quantile $g(t,\cdot)$ leaves the cone $C(0,1)$ or the free boundary $\beta(t)$ decreases, the H1 regularization would fail to preserve the Wasserstein/quantile structure, contradicting Proposition 3.6 and the Wasserstein interpretation of Theorem 3.4.
Extended reading notes
Core claim
The paper's central claim is that the Sobolev-regularized functionals $\widetilde F^\pm_\nu = F^\pm_\nu + F_H^{(\nu)} + I_{C(0,1)}$ have well-defined Wasserstein-type gradient flows in one dimension. For the negative kernel, $\widetilde F^-_\nu$ is proper, convex, and lower semicontinuous, and Theorem 3.4 asserts a unique strong solution $g\in H^1_{\mathrm{loc}}([0,\infty);L^2(0,1))$ of $\partial_t g\in -\partial \widetilde F^-_\nu(g)$, whose push-forward is the unique Wasserstein gradient flow of $\widetilde F^-_\nu$. For the positive kernel, $\widetilde F^+_\nu$ is not $\lambda$-convex, but the H1 term supplies compact sublevels and the required chain rule, so Theorem 3.10 asserts the existence of generalized minimizing movements $g\in \mathrm{GMM}(\widetilde F^+_\nu,g_0)$ that are strong $H^1(0,T;L^2(0,1))$ solutions of $\partial_t g\in -\partial_l \widetilde F^+_\nu(g)$ with the limiting subdifferential, and the associated curves in $P_2(\mathbb{R})$ are Wasserstein absolutely continuous flows. The positive-kernel Dirac example gives $g(t,s)=-1-t$, so the regularized positive flow is not the time reversal of the negative one.
Load-bearing premise
Both main theorems assume the initial quantile $g_0$ (and, for the negative kernel, the target quantile $Q_\nu$) lies in $C(0,1)\cap H^1(0,1)$, which forces the initial and target measures to have compact and convex support; the paper's fractional-Sobolev relaxation of this restriction is only sketched.
Editorial extensions
If this is right
- For the negative distance kernel, the regularized flow exists uniquely, can be approximated by an implicit Euler scheme that solves a Neumann boundary-value problem at each step, and converges to the unregularized flow as $\lambda\downarrow 0$.
- For the positive distance kernel, existence of a flow is restored despite the lack of $\lambda$-convexity: generalized minimizing movements exist on every finite time horizon and yield Wasserstein absolutely continuous curves.
- The regularization gives a stronger long-time bound, $\mathrm{MMD}_K^2(\mu(t),\nu)+|Q_{\mu(t)}-Q_\nu|^2_{H^1} \le W_2^2(\mu_0,\nu)/(2t)$, so the flow approaches the target at least as fast as the unregularized one.
- In the Dirac-to-Dirac example, the negative-kernel flow is explicit before it touches the target (a heat equation with forcing $2s$) and then becomes a Stefan free-boundary problem, showing the regularization turns a dissipating mass tail into a moving support.
- The regularized positive-kernel flow from $\delta_{-1}$ away from $\delta_0$ is the translation $\delta_{-1-t}$, so positive and negative kernel flows are not time reversals of one another.
Reading between the lines
- The paper does not state this, but the 'horns' at the support boundary, whose height is set by the target density, suggest that in higher dimensions a Laplacian-type penalty could similarly prevent mass from spreading to infinity; a testable extension is to run the same construction with a radial Laplacian term on quantile-like radial maps.
- The fractional Sobolev relaxation in Remark 3.13 is only sketched; making it precise for $\sigma<1/2$ would likely extend the theorems to initial and target measures with unbounded or disconnected support, and the natural first check is whether the compactness argument survives without the absolute continuity of $g_0$.
- Theorem 3.10 asserts existence but not uniqueness of the generalized minimizing movement, so different subsequences of step sizes or different $\lambda$ values may select different positive-kernel flows; this is a concrete numerical question for non-constant initial data.
- The Stefan problem in Example 3.7 is a well-defined free-boundary problem; a rigorous analysis of $\beta(t)$ would turn the numerically observed formation of a Dirac at the target into a theorem and give a benchmark for Wasserstein numerical schemes.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies Wasserstein gradient flows on P2(R) of MMD functionals with distance kernel K(x,y)=±|x-y|. Using the isometric embedding of P2(R) into the cone C(0,1)⊂L2(0,1) via quantile functions, it reformulates the flow as an L2 gradient flow for an associated functional, and adds a Sobolev H1 regularization λ/2∫|u'|^2 with Neumann boundary conditions (shifted by the target quantile for the negative kernel). Main results: Theorem 3.4 proves existence and uniqueness of a strong solution of the Cauchy problem for the regularized negative-kernel functional, yielding a unique Wasserstein gradient flow; Theorem 3.10 proves existence of generalized minimizing movements for the regularized positive-kernel functional, solving a Cauchy problem with the limiting subdifferential. Proposition 3.6 gives conditions under which the cone constraint is automatically preserved, reducing the flow to the unconstrained Neumann problem. The paper also derives explicit Dirac-to-Dirac examples, a Stefan-type free boundary problem for the negative kernel after contact, convergence as λ↓0 and t↑∞ remarks, and numerical experiments showing that the regularization removes a 'dissipation-of-mass' defect.
Significance. The existence theory for the positive distance kernel appears to be new; the verification of the Rossi–Savaré compactness and chain-rule hypotheses via the compact Sobolev embedding is natural and sound. The negative-kernel result is a regularized extension of the existing convex-flow theory, and the explicit formulas and numerics illustrate a practically relevant improvement. The paper is generally careful: the central theorems are proved from stated assumptions via standard semigroup theory, the regularization parameter λ is explicit and not fitted, and the text is candid about parts that remain numerical (the Stefan problem after contact) or sketched (fractional Sobolev relaxation). The referee checked the disputed monotonicity argument in Proposition 3.6; it is correct, because the term gn(s2)-gn(s1)≥0, the strict decrease of gn+1, and the inequality R^-_ν(gn+1(s1))≥R^+_ν(gn+1(s2)) make the central strict inequality true independently of τ and of the target density. Overall this is a solid contribution; the issues below are local and presentation-level.
minor comments (7)
- [Section 2.2 (after Eq. (5))] The statement that for a proper convex and lsc functional F one has dom(∂F)=dom(F) is false in general; the paper's own FH in Lemma 3.2 is a counterexample, since dom(FH)=H1(0,1) while dom(∂FH) is a proper dense subset. Please replace this with the correct statement that dom(∂F) is dense in dom(F).
- [Theorem 3.4] The claim that the convergence in (3) is 'globally uniform in t∈[0,∞)' is stronger than the standard compact-interval convergence result for convex lsc semigroups; please either supply a proof or weaken the statement to uniform convergence on compact intervals.
- [Proposition 3.6 (proof)] The proof chooses an interval (a,b)⊂(0,1) with g'_{n+1}(a)=g'_{n+1}(b)=0 around a point where g'_{n+1}<0; if the negative region touches the boundary, one of the endpoints may be 0 or 1. The monotonicity contradiction is unaffected, but the proof should allow one-sided intervals or explicitly justify that boundary endpoints can be handled by approximation.
- [Example 3.7] The Stefan problem for t≥t* is introduced formally, and the paper does not prove well-posedness for it or the claimed emergence of a Dirac point at 0. Since this is presented as a numerical observation, this is acceptable, but the formal nature of the Stefan formulation should be stated more explicitly.
- [Section 4.1 / Algorithm 1] The equivalence of (25) and (26) is stated only for continuous CDFs, yet the Dirac-to-Dirac example has a discontinuous Rν; the text should make explicit that in that case the BVP solver is applied only to the absolutely continuous part, with the singular part treated by the separate procedure described in Section 4.2.
- [Remark 3.13] The claim that Theorems 3.4 and 3.10 remain valid with fractional Sobolev spaces is only a sketch; please label it as an outlook/conjecture rather than a statement of established results.
- [Throughout (Section 4.2)] There are minor typographical issues: 'Figues' should be 'Figures', 'T echnical' should be 'Technical', and 'absolute continuous' should be 'absolutely continuous'.
Circularity Check
No circularity: the existence theorems reduce to standard semigroup/compactness results plus an explicit functional identity, and no fitted quantity is presented as a prediction.
full rationale
The paper's central claims are Theorem 3.4 (unique strong solution and minimizing movement for ~F−ν) and Theorem 3.10 (generalized minimizing movement and strong solution for ~F+ν). Both are proven by verifying convexity, lower semicontinuity, boundedness from below, and the compactness/chain-rule hypotheses, then applying standard semigroup theory ([10], [1]) and the external Rossi–Savaré result [26]. The hypotheses are checked from Lemma 3.1 and Lemma 3.2, not from the theorems' conclusions. Lemma 3.1 is quoted from the authors' prior work [13], but it is an explicit subdifferential formula and association identity rather than the target existence statement, and its proof is a direct calculation; Theorem 2.4 is reproved in the appendix. The parameter λ is an explicit diffusion constant chosen by the user, never fitted to data, so no fitted input is renamed as a prediction. The numerical section solves the same Neumann problem (25) derived from the convex functional (16), and the Dirac-to-Dirac example uses an explicit Fourier solution (19), again not assuming the conclusion. Remark 3.13 honestly limits the H1 assumptions, and the possible gap in Proposition 3.6's strict-inequality argument is a correctness/mathematical-risk issue rather than a circular reduction, since Theorems 3.4 and 3.10 establish existence for the full functional (14) including the indicator IC(0,1). No equation is equal to another by construction, and no load-bearing argument reduces to an unverified self-citation.
Assumptions & free parameters
free parameters (1)
- λ =
various (10^-2, 1, 10^-4 in numerics)
assumptions (4)
- standard math Standard properties of Wasserstein space P2(R), the quantile isometry, and subdifferential theory for convex functionals.
- domain assumption Initial and target quantiles lie in H1(0,1) (g0∈C(0,1)∩H1(0,1) for both kernels; Qν∈C(0,1)∩H1(0,1) for negative kernel).
- domain assumption Condition (15): s↦2s−λQν''(s) non-decreasing, for Proposition 3.6.
- standard math Rossi-Savaré generalized minimizing movement theory, including the compactness lemma and chain rule (Appendix A.2).
Cite this review
Pith. "Pith review of Wasserstein Gradient Flows of MMD Functionals with Distance Kernels under Sobolev Regularization." pith.science (2026). https://pith.science/paper/NCUCRR6G
@misc{pith2026241109848,
author = {Pith},
title = {Pith review of: Wasserstein Gradient Flows of MMD Functionals with Distance Kernels under Sobolev Regularization},
year = {2026},
howpublished = {\url{https://pith.science/paper/NCUCRR6G}},
note = {Machine review of arXiv:2411.09848}
}
abstract
We consider Wasserstein gradient flows of maximum mean discrepancy (MMD) functionals $\text{MMD}_K^2(\cdot, \nu)$ for positive and negative distance kernels $K(x,y) := \pm |x-y|$ and given target measures $\nu$ on $\mathbb{R}$. Since in one dimension the Wasserstein space can be isometrically embedded into the cone $\mathcal C(0,1) \subset L_2(0,1)$ of quantile functions, Wasserstein gradient flows can be characterized by the solution of an associated Cauchy problem on $L_2(0,1)$. While for the negative kernel, the MMD functional is geodesically convex, this is not the case for the positive kernel, which needs to be handled to ensure the existence of the flow. We propose to add a regularizing Sobolev term $|\cdot|^2_{H^1(0,1)}$ corresponding to the Laplacian with Neumann boundary conditions to the Cauchy problem of quantile functions. Indeed, this ensures the existence of a generalized minimizing movement for the positive kernel. Furthermore, for the negative kernel, we demonstrate by numerical examples how the Laplacian rectifies a "dissipation-of-mass" defect of the MMD gradient flow.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
L. Ambrosio, N. Gigli, and G. Savare.Gradient Flows. Lectures in Mathematics ETH Zürich. Birkhäuser, Basel, 2nd edition, 2008
work page 2008
-
[2]
Arbel, A
M. Arbel, A. Korba, A. Salim, and A. Gretton. Maximum mean discrepancy gradient flow. In Advances in Neural Information Processing Systems, volume 32, 2019
2019
-
[3]
H. Attouch. Familles d’opérateurs maximaux monotones et mesurabilité.Annali di Matematica Pura ed Applicata, 120(4):35–111, 1979
work page 1979
-
[4]
V. Barbu. Nonlinear Differential Equations of Monotone Types in Banach Spaces. Springer Science & Business Media, 2010
work page 2010
-
[5]
M. Binkowski, D. J. Sutherland, M. Arbel, and A. Gretton. Demystifying MMD GANs. In Proceedings ICLR 2018. OpenReview, 2018
work page 2018
-
[6]
A. Bobrowski and R. Rudnicki. On convergence and asymptotic behaviour of semigroups of operators. Philosophical Transactions of the Royal Society A, 378, 2020
work page 2020
-
[7]
G. A. Bonaschi, J. A. Carrillo, M. Di Francesco, and M. A. Peletier. Equivalence of gradient flows and entropy solutions for singular nonlocal interaction equations in 1d.ESAIM: Control, Optimisation and Calculus of Variations, 21(2):414–441, 2015
work page 2015
-
[8]
K. M. Borgwardt, A. Gretton, M. J. Rasch, H.-P. Kriegel, B. Schölkopf, and A. J. Smola. Integrating structured biological data by kernel maximum mean discrepancy.Bioinformatics, 22(14):e49–e57, 2006
work page 2006
Show all 30 references
-
[9]
Boufadène and F.-X
S. Boufadène and F.-X. Vialard. On the global convergence of Wasserstein gradient flow of the Coulomb discrepancy.arXiv preprint arXiv:2312.00800, 2023
2023
-
[10]
H. Brezis. Operateurs Maximaux Monotones. North-Holland Mathematics Studies, 1973
1973
-
[11]
Brezis and A
H. Brezis and A. Pazy. Convergence and approximation of semigroups of nonlinear operators in Banach spaces.Journal of Functional Analysis, 9:63–74, 1972
1972
-
[12]
Y. Chen, D. Z. Huang, J. Huang, S. Reich, and A. M. Stuart. Sampling via gradient flows in the space of probability measures.arXiv preprint arXiv:2310.03597, 2023
2023 arXiv
-
[13]
Duong, V
R. Duong, V. Stein, R. Beinert, J. Hertrich, and G. Steidl. Wasserstein gradient flows of MMD functionals with distance kernel and Cauchy problems on quantile functions.arXiv preprint arXiv:2408.07498, 2024
2024 arXiv
-
[14]
G. K. Dziugaite, D. M. Roy, and Z. Ghahramani. Training generative neural networks via maximum mean discrepancy optimization. InProceedings UAI 2015. UAI, 2015. 25
2015
-
[15]
Ehler, M
M. Ehler, M. Gräf, S. Neumayer, and G. Steidl. Curve based approximation of measures on manifolds by discrepancy minimization. Foundations of Computational Mathematics, 21(6):1595–1642, 2021
2021
-
[16]
Gretton, K
A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola. A kernel two-sample test. J. Mach. Learn. Res., 13(25):723–773, 2012
2012
-
[17]
Hagemann, J
P. Hagemann, J. Hertrich, F. Altekrüger, R. Beinert, J. Chemseddine, and G. Steidl. Posterior sampling based on gradient flows of the MMD with negative distance kernel. InInternational Conference on Learning Representations, 2024
2024
-
[18]
Hertrich, R
J. Hertrich, R. Beinert, M. Gräf, and G. Steidl. Wasserstein gradient flows of the discrepancy with distance kernel on the line. InInternational Conference on Scale Space and Variational Methods in Computer Vision, pages 431–443. Springer, 2023
2023
-
[19]
Hertrich, M
J. Hertrich, M. Gräf, R. Beinert, and G. Steidl. Wasserstein steepest descent flows of discrep- ancies with Riesz kernels.Journal of Mathematical Analysis and Applications, 531(1):127829, 2024
2024
-
[20]
Laumont, V
R. Laumont, V. Bortoli, A. Almansa, J. Delon, A. Durmus, and M. Pereyra. Bayesian imaging using plug & play priors: when Langevin meets Tweedie.SIAM Journal on Imaging Sciences, 15(2):701–737, 2022
2022
-
[21]
F. Maggi. Optimal Mass Transport on Euclidean Spaces. Cambridge Studies in Advanced Mathematics. Cambridge University Press, 2023
2023
-
[22]
Modeste and C
T. Modeste and C. Dombry. Characterization of translation invariant MMD on R d and connections with Wasserstein distances.Journal of Machine Learning Research, 25(237):1–39, 2024
2024
-
[23]
Neumayer and G
S. Neumayer and G. Steidl. From optimal transport to discrepancy. In K. Chen, C.-B. Schön- lieb, X.-C. Tai, and L. Younes, editors,Handbook of Mathematical Models and Algorithms in Computer Vision and Imaging: Mathematical Imaging and Vision, pages 1–36. Springer, 2023
2023
-
[24]
R. R. Phelps. Convex functions, Monotone Operators and Differentiability (2nd edition). Springer, Berlin, Heidelberg, New York, 1993
1993
-
[25]
Randomvariables, monotonerelations, andconvexanalysis
R.T.RockafellarandJ.O.Royset. Randomvariables, monotonerelations, andconvexanalysis. Mathematical Programming, 148:297–331, 2014
2014
-
[26]
Rossi and G
R. Rossi and G. Savaré. Gradient flows of non convex functionals in Hilbert spaces and applications. ESAIM: Control, Optimisation and Calculus of Variations, 12:564–614, 2006
2006
-
[27]
Sejdinovic, B
D. Sejdinovic, B. Sriperumbudur, A. Gretton, and K. Fukumizu. Equivalence of distance-based and RKHS-based statistics in hypothesis testing.The Annals of Statistics, 41(5):2263 – 2291, 2013
2013
-
[28]
Teuber, G
T. Teuber, G. Steidl, P. Gwosdek, C. Schmaltz, and J. Weickert. Dithering by differences of convex functions.SIAM Journal on Imaging Sciences, 4(1):79–108, 2011
2011
-
[29]
Villani.Topics in Optimal Transportation
C. Villani.Topics in Optimal Transportation. Number 58 in Graduate Studies in Mathematics. American Mathematical Society, Providence, 2003
2003
-
[30]
C. Villani. Optimal Transport. Springer, Berlin, 2009. 26 A Appendix A.1 Proof of Theorem 2.4 The proof is taken from [13, Theorem 3.5] with slight adjustments to our setting. Let γt := (g(t))#Λ(0,1). Then, it is easy to check thatγt are probability measures inP2(R) with quant...
2009
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.