Pith. sign in

REVIEW 3 minor 4 cited by

For Ahlfors-regular measures, empirical approximations achieve the sharp rate N to the minus one-half times one plus q over beta in energy distance.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-30 18:29 UTC pith:KIKAYEKN

load-bearing objection The paper claims the first sharp two-sided rate N^{-1/2(1+q/β)} for empirical energy distance under explicit Ahlfors regularity, turning a 2014 qualitative result quantitative.

arxiv 2605.18497 v2 pith:KIKAYEKN submitted 2026-05-18 math.PR math.OCmath.STstat.TH

Sharp Rates of MMD Empirical Estimation with Power Kernels

classification math.PR math.OCmath.STstat.TH
keywords maximum mean discrepancyenergy distanceempirical approximationAhlfors regularitypower kernelsconvergence ratesprobability measures
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper proves sharp two-sided bounds on how fast N-point empirical measures can approximate a target probability measure using the energy distance from power kernels. For targets whose compact support satisfies an Ahlfors regularity condition with exponent beta, both every possible empirical measure and the best one achieve exactly the rate N to the power of minus one-half times one plus q over beta. This supplies explicit quantitative speeds for a prior result that only established consistency without rates.

Core claim

Given a probability measure ω on R^d with compact support satisfying an Ahlfors regularity condition of exponent β ∈ (0,d], the sharp two-sided bound E_q(μ_N, ω) ≍ N^{-½(1 + q/β)} holds both for the worst-case empirical measure μ_N (lower bound) and for an optimally chosen empirical measure μ_N (upper bound).

What carries the argument

The energy distance E_q induced by the power kernel K_q(x,y) = -|x-y|^q, together with the Ahlfors regularity condition on the support of ω.

Load-bearing premise

The target measure ω has compact support satisfying an Ahlfors regularity condition of exponent β in (0,d].

What would settle it

Observe an Ahlfors-regular ω for which the energy distance of some sequence of empirical measures converges at a rate other than N to the power of minus one-half times one plus q over beta.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • The exponent depends on the regularity parameter β rather than ambient dimension d.
  • The lower bound applies uniformly to every configuration of N points.
  • The upper bound is attained by at least one choice of N points.
  • The same rate governs the classical energy distance when q equals 1.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The result may extend to other integral probability metrics whose kernels have comparable singularity.
  • Sampling algorithms could be tuned to achieve this specific rate rather than generic Monte Carlo scaling.
  • Relaxing compactness or the Ahlfors condition would likely produce a different exponent or a one-sided bound.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 3 minor

Summary. The paper establishes quantitative rates for empirical estimation of probability measures via the MMD (energy distance) with power kernels K_q(x,y) = -|x-y|^q for q in (0,2). Under the assumption that the target measure ω has compact support satisfying an Ahlfors regularity condition of exponent β ∈ (0,d], it proves the sharp two-sided bound E_q(μ_N, ω) ≍ N^{-1/2 (1 + q/β)} that holds both for arbitrary (worst-case) N-point empirical measures (lower bound) and for optimally chosen ones (upper bound). This complements the qualitative consistency result of Fornasier and Hütter.

Significance. If the result holds, it supplies the first sharp quantitative rates for this class of MMD empirical problems, turning a qualitative consistency statement into a precise asymptotic with matching upper and lower bounds. The explicit dependence on the Ahlfors exponent β and the kernel parameter q is a clear strength, as is the two-sided nature of the claim.

minor comments (3)
  1. The abstract and introduction state the main theorem clearly, but the manuscript should include a short remark on whether the Ahlfors condition is also necessary for the lower bound or only sufficient.
  2. Notation for the energy distance is introduced as E_q^2 in the abstract but then used as E_q in the displayed rate; a single consistent symbol throughout would improve readability.
  3. The reference to Fornasier and Hütter is cited for the qualitative result; adding one sentence on how the new quantitative bound improves upon or relates to other known rates in the MMD literature would help situate the contribution.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for their positive summary and recommendation of minor revision. No specific major comments were raised in the report, so we have no points to address or revisions to propose.

Circularity Check

0 steps flagged

No significant circularity detected

full rationale

The paper derives sharp two-sided rates E_q(μ_N, ω) ≍ N^{-½(1 + q/β)} directly from the Ahlfors regularity hypothesis on the compact support of ω. This is a self-contained mathematical proof establishing both the lower bound (for arbitrary N-point measures) and matching upper bound (for optimal choice), without any reduction to fitted parameters, self-definitional loops, or load-bearing self-citations. The cited 2014 consistency result (Fornasier-Hütter) supplies only the qualitative narrow convergence fact and is independent of the quantitative exponent derived here; the present work explicitly complements it by adding rates. No enumerated circularity pattern applies.

Axiom & Free-Parameter Ledger

0 free parameters · 1 axioms · 0 invented entities

The central claim rests on the Ahlfors-regularity assumption (a domain assumption) and standard facts from geometric measure theory and potential theory. No free parameters are fitted inside the derivation; q and β are given parameters of the kernel and the measure class. No new entities are postulated.

axioms (1)
  • domain assumption ω has compact support satisfying Ahlfors regularity of exponent β ∈ (0,d]
    Invoked in the statement of the main theorem to obtain the precise exponent; without it the claimed rate does not hold.

pith-pipeline@v0.9.1-grok · 5851 in / 1529 out tokens · 75635 ms · 2026-06-30T18:29:01.097428+00:00 · methodology

0 comments
read the original abstract

We establish quantitative rates of convergence for the empirical estimation of probability measures by means of the Maximum Mean Discrepancy (MMD) with power kernel $K_q(x,y) = -|x-y|^q$, $q \in (0,2)$. The resulting discrepancy is the classical \emph{energy distance} $$\mathcal E_q^2(\mu, \omega) = -\frac{1}{2}\iint_{\mathbb{R}^d \times \mathbb{R}^d} |x-y|^q \, d(\mu - \omega)(x)\, d(\mu - \omega)(y),$$ and we ask how fast the best $N$-point empirical approximation $\inf_{\mu_N \in \mathcal{P}^N}\mathcal{E}_q(\mu_N,\omega)$ decays as $N \to \infty$. Given a probability measure $\omega$ on $\mathbb{R}^d$ with compact support satisfying an Ahlfors regularity condition of exponent $\beta \in (0,d]$, we prove that the sharp two-sided bound $$\mathcal E_q(\mu_N, \omega) \asymp N^{-\frac{1}{2}\left(1 + \frac{q}{\beta}\right)}$$ holds both for the worst-case empirical measure $\mu_N$ (lower bound, holding for every configuration of $N$ points) and for an optimally chosen empirical measure $\mu_N$ (upper bound). This complements the qualitative consistency result of Fornasier and H\"utter \cite{fornasier2014consistency}, who proved narrow convergence of the minimizers of $\mathcal E_q^2(\cdot, \omega)$ over empirical measures without quantitative rates.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Wasserstein gradient flows for Coulomb discrepancies

    math.AP 2026-07 accept novelty 7.0

    Under Coulomb MMD energy, the Wasserstein gradient flow converges exponentially to uniformly positive targets under PL-type coercivity, exhibits rigidity of critical points, and admits no uniform whole-space convergence rate.

  2. The nonlocal attraction-repulsion transport equation with power kernels

    math.AP 2026-07 conditional novelty 7.0

    A proof of global well-posedness, uniform support confinement, a fractional-Laplacian free-boundary characterization of stationary states, and explicit stationary profiles for nonlocal attraction-repulsion continuity ...

  3. Wasserstein gradient flows for Coulomb discrepancies

    math.AP 2026-07 conditional novelty 6.5

    Wasserstein gradient flows of Coulomb MMD exist globally, become instantly bounded, decay exponentially on the torus via a defective PL inequality, but face spatial-infinity obstructions on R^d.

  4. The nonlocal attraction-repulsion transport equation with power kernels

    math.AP 2026-07 accept novelty 6.5

    Global Lagrangian well-posedness, uniform support bounds under attraction dominance, free-boundary characterization of zero-flux stationary states, and subsequential convergence are established for the power-kernel at...

Reference graph

Works this paper leans on

42 extracted references · 42 canonical work pages · cited by 2 Pith papers · 2 internal anchors

  1. [1]

    Oxford Mathematical Monographs

    Luigi Ambrosio, Nicola Fusco, and Diego Pallara.Functions of bounded variation and free dis- continuity problems. Oxford Mathematical Monographs. The Clarendon Press, Oxford University Press, New York, 2000

  2. [2]

    Lectures in Mathematics ETH Z¨ urich

    Luigi Ambrosio, Nicola Gigli, and Giuseppe Savar´ e.Gradient Flows in Metric Spaces and in the Space of Probability Measures. Lectures in Mathematics ETH Z¨ urich. Birkh¨ auser, Basel, 2nd edition, 2008

  3. [3]

    From kinetic theory to AI: A rediscovery of high-dimensional divergences and their properties.Mathematical Models and Methods in Applied Sciences, 2026

    Gennaro Auricchio, Giovanni Brigati, Paolo Giudici, and Giuseppe Toscani. From kinetic theory to AI: A rediscovery of high-dimensional divergences and their properties.Mathematical Models and Methods in Applied Sciences, 2026. Preprint arXiv:2507.11387

  4. [4]

    The equivalence of Fourier-based and Wasserstein metrics on imaging problems.Atti Accademia Nazionale dei Lincei

    Gennaro Auricchio, Andrea Codegoni, Stefano Gualandi, Giuseppe Toscani, and Marco Veneroni. The equivalence of Fourier-based and Wasserstein metrics on imaging problems.Atti Accademia Nazionale dei Lincei. Rendiconti Lincei. Matematica e Applicazioni, 31(3):627–649, 2020

  5. [5]

    Luca Brandolini, William W. L. Chen, Leonardo Colzani, Giacomo Gigante, and Giancarlo Travaglini. Discrepancy and numerical integration on metric measure spaces.Journal of Geo- metric Analysis, 29(1):328–369, 2019

  6. [6]

    A projection algorithm on measures sets

    Nicolas Chauffert, Philippe Ciuciu, Jonas Kahn, and Pierre Weiss. A projection method on mea- sures sets.Constructive Approximation, 45(1):83–111, February 2017. Preprint arXiv:1509.00229, 2015

  7. [7]

    Kernel two-sample tests for manifold data.Bernoulli, 30(4):2572– 2597, 2024

    Xiuyuan Cheng and Yao Xie. Kernel two-sample tests for manifold data.Bernoulli, 30(4):2572– 2597, 2024

  8. [8]

    Quantita- tive convergence of Wasserstein gradient flows of kernel mean discrepancies, 2026

    L´ ena ¨ ıc Chizat, Maria Colombo, Roberto Colombo, and Xavier Fern´ andez-Real. Quantita- tive convergence of Wasserstein gradient flows of kernel mean discrepancies, 2026. Preprint, arXiv:2603.01977. SHARP RATES OF MMD EMPIRICAL ESTIMATION WITH POWER KERNELS 33

  9. [9]

    Sinkhorn distances: Lightspeed computation of optimal transport

    Marco Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport. In Christopher J. C. Burges, L´ eon Bottou, Zoubin Ghahramani, and Kilian Q. Weinberger, editors,Advances in Neural Information Processing Systems, volume 26, pages 2292–2300, 2013

  10. [10]

    Sugli estremi dei momenti delle funzioni di ripartizione doppia.Annali della Scuola Normale Superiore di Pisa - Scienze Fisiche e Matematiche, Ser

    Giorgio Dall’Aglio. Sugli estremi dei momenti delle funzioni di ripartizione doppia.Annali della Scuola Normale Superiore di Pisa - Scienze Fisiche e Matematiche, Ser. 3, 10(1-2):35–74, 1956

  11. [11]

    Constructive quantization: approxi- mation by empirical measures.Annales de l’I.H.P

    Steffen Dereich, Michael Scheutzow, and Reik Schottstedt. Constructive quantization: approxi- mation by empirical measures.Annales de l’I.H.P. Probabilit´ es et statistiques, 49(4):1183–1203, 2013

  12. [12]

    Asymptotic behavior of gradient flows driven by nonlocal power repulsion and attraction potentials in one dimension.SIAM Journal on Mathematical Analysis, 46(6):3814–3837, 2014

    Marco Di Francesco, Massimo Fornasier, Jan-Christian H¨ utter, and Daniel Matthes. Asymptotic behavior of gradient flows driven by nonlocal power repulsion and attraction potentials in one dimension.SIAM Journal on Mathematical Analysis, 46(6):3814–3837, 2014

  13. [13]

    Springer Monographs in Mathematics

    Irene Fonseca and Giovanni Leoni.Modern Methods in the Calculus of Variations: Lp Spaces. Springer Monographs in Mathematics. Springer New York, 2007

  14. [14]

    Consistency of variational continuous- domain quantization via kinetic theory.Applicable Analysis, 92(6):1283–1298, 2013

    Massimo Fornasier, Jan Haˇ skovec, and Gabriele Steidl. Consistency of variational continuous- domain quantization via kinetic theory.Applicable Analysis, 92(6):1283–1298, 2013

  15. [15]

    Consistency of Probability Measure Quantization by Means of Power Repulsion-Attraction Potentials

    Massimo Fornasier and Jan-Christian H¨ utter. Consistency of probability measure quantization by means of power repulsion–attraction potentials.Journal of Fourier Analysis and Applications, 22(3):694–749, 2016. Preprint arXiv:1310.1120, 2013

  16. [16]

    On the rate of convergence in Wasserstein distance of the empirical measure.Probability Theory and Related Fields, 162(3–4):707–738, 2015

    Nicolas Fournier and Arnaud Guillin. On the rate of convergence in Wasserstein distance of the empirical measure.Probability Theory and Related Fields, 162(3–4):707–738, 2015

  17. [17]

    Birkh¨ auser, Boston, 1997

    Bert Fristedt and Lawrence Gray.A Modern Approach to Probability Theory. Birkh¨ auser, Boston, 1997

  18. [18]

    Diameter bounded equal measure partitions of Ahlfors regular metric measure spaces.Discrete Comput

    Giacomo Gigante and Paul Leopardi. Diameter bounded equal measure partitions of Ahlfors regular metric measure spaces.Discrete Comput. Geom., 57(2):419–430, 2017

  19. [19]

    Springer, Berlin, 2000

    Siegfried Graf and Harald Luschgy.Foundations of Quantization for Probability Distributions, volume 1730 ofLecture Notes in Mathematics. Springer, Berlin, 2000

  20. [20]

    Borgwardt, Malte J

    Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Sch¨ olkopf, and Alexander Smola. A kernel two-sample test.Journal of Machine Learning Research, 13(25):723–773, 2012

  21. [21]

    Borgwardt, Malte J

    Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Sch¨ olkopf, and Alexander J. Smola. A kernel method for the two-sample-problem. In Bernhard Sch¨ olkopf, John Platt, and Thomas Hofmann, editors,Advances in Neural Information Processing Systems 19 (NIPS 2006), pages 513–520. MIT Press, December 2006

  22. [22]

    Posterior sampling based on gradient flows of the MMD with negative distance kernel

    Paul Hagemann, Johannes Hertrich, Fabian Altekr¨ uger, Robert Beinert, Jannis Chemseddine, and Gabriele Steidl. Posterior sampling based on gradient flows of the MMD with negative distance kernel. InThe Twelfth International Conference on Learning Representations, 2024

  23. [23]

    Generative sliced MMD flows with riesz kernels

    Johannes Hertrich, Christian Wald, Fabian Altekr¨ uger, and Paul Hagemann. Generative sliced MMD flows with riesz kernels. InThe Twelfth International Conference on Learning Represen- tations, 2024

  24. [24]

    Hutchinson

    John E. Hutchinson. Fractals and self-similarity.Indiana Univ. Math. J., 30(5):713–747, 1981

  25. [25]

    Distance covariance in metric spaces.The Annals of Probability, 41(5):3284–3305, 2013

    Russell Lyons. Distance covariance in metric spaces.The Annals of Probability, 41(5):3284–3305, 2013

  26. [26]

    Characterization of translation invariant MMD on Rd and connections with Wasserstein distances.Journal of Machine Learning Research, 25:1–39, 2024

    Thibault Modeste and Cl´ ement Dombry. Characterization of translation invariant MMD on Rd and connections with Wasserstein distances.Journal of Machine Learning Research, 25:1–39, 2024

  27. [27]

    Computational optimal transport.Foundations and Trends in Machine Learning, 11(5–6):355–607, 2019

    Gabriel Peyr´ e and Marco Cuturi. Computational optimal transport.Foundations and Trends in Machine Learning, 11(5–6):355–607, 2019

  28. [28]

    Electrostatic halftoning.Computer Graphics Forum, 29(8):2313–2327, December 2010

    Christian Schmaltz, Pascal Gwosdek, Andr´ es Bruhn, and Joachim Weickert. Electrostatic halftoning.Computer Graphics Forum, 29(8):2313–2327, December 2010

  29. [29]

    I. J. Schoenberg. Metric spaces and completely monotone functions.Annals of Mathematics, 39(4):811–841, 1938

  30. [30]

    I. J. Schoenberg. Metric spaces and positive definite functions.Transactions of the American Mathematical Society, 44(3):522–536, November 1938

  31. [31]

    Equivalence of distance-based and RKHS-based statistics in hypothesis testing.The Annals of Statistics, 41(5):2263–2291, October 2013

    Dino Sejdinovic, Bharath Sriperumbudur, Arthur Gretton, and Kenji Fukumizu. Equivalence of distance-based and RKHS-based statistics in hypothesis testing.The Annals of Statistics, 41(5):2263–2291, October 2013. SHARP RATES OF MMD EMPIRICAL ESTIMATION WITH POWER KERNELS 34

  32. [32]

    Vogelstein

    Cencheng Shen and Joshua T. Vogelstein. The exact equivalence of distance and kernel methods in hypothesis testing.AStA Advances in Statistical Analysis, 105(3):385–403, 2021

  33. [33]

    Sz´ ekely

    G´ abor J. Sz´ ekely. Potential and kinetic energy in statistics. Lecture Notes, Budapest Institute of Technology (Technical University of Budapest), 1989

  34. [34]

    Sz´ ekely

    G´ abor J. Sz´ ekely. E-statistics: The Energy of Statistical Samples. Technical Report 02-16, Department of Mathematics and Statistics, Bowling Green State University, 2002

  35. [35]

    Sz´ ekely and Maria L

    G´ abor J. Sz´ ekely and Maria L. Rizzo. A new test for multivariate normality.Journal of Multivariate Analysis, 93(1):58–80, 2005

  36. [36]

    Sz´ ekely and Maria L

    G´ abor J. Sz´ ekely and Maria L. Rizzo. Energy statistics: A class of statistics based on distances. Journal of Statistical Planning and Inference, 143(8):1249–1272, August 2013

  37. [37]

    Sz´ ekely and Maria L

    G´ abor J. Sz´ ekely and Maria L. Rizzo.The Energy of Data and Distance Correlation, volume 171 ofChapman & Hall/CRC Monographs on Statistics and Applied Probability. Chapman and Hall/CRC Press, Boca Raton, 2023

  38. [38]

    Sz´ ekely, Maria L

    G´ abor J. Sz´ ekely, Maria L. Rizzo, and Nail K. Bakirov. Measuring and testing dependence by correlation of distances.The Annals of Statistics, 35(6):2769–2794, December 2007

  39. [39]

    Dithering by differences of convex functions.SIAM Journal on Imaging Sciences, 4(1):79–108, 2011

    Tanja Teuber, Gabriele Steidl, Pascal Gwosdek, Christian Schmaltz, and Joachim Weickert. Dithering by differences of convex functions.SIAM Journal on Imaging Sciences, 4(1):79–108, 2011

  40. [40]

    Cambridge University Press, Cambridge, 2005

    Holger Wendland.Scattered Data Approximation, volume 17 ofCambridge Monographs on Applied and Computational Mathematics. Cambridge University Press, Cambridge, 2005

  41. [41]

    Kernel two-sample tests in high dimensions: interplay between moment discrepancy and dimension-and-sample orders.Biometrika, 110(2):411–430, 2023

    Jian Yan and Xianyang Zhang. Kernel two-sample tests in high dimensions: interplay between moment discrepancy and dimension-and-sample orders.Biometrika, 110(2):411–430, 2023

  42. [42]

    Paul L. Zador. Topics in the asymptotic quantization of continuous random variables. Technical report, Bell Laboratories, Murray Hill, NJ, 1966. (Francesco Colasanto)Department of Mathematics, CIT School, Technical University of Munich, Munich, Germany Email address:francesco.colasanto@tum.de (Matteo Focardi)DiMaI U. Dini, Universit `a di Firenze, Florenc...