Pith. sign in

REVIEW 2 major objections 6 minor 145 references

Minimax Lower Bounds of Kernel Discrepancy Estimation: MMD, HSIC, KSD

T0 review · 2 major / 6 minor · reviewed 2026-07-31 · grok-4.5

Pith's one-line read Estimating MMD, HSIC and KSD cannot beat the parametric n^{-1/2} rate on general spaces under mild kernel assumptions.

desk verdict Clean Le Cam lower bounds that close the topological-space gap for MMD/HSIC/KSD, with one overstated tightness claim for unbounded kernels. read the letter →

arxiv 2607.24235 v1 pith:2WBBHXPP submitted 2026-07-27 stat.ML cs.LGmath.STstat.TH

classification stat.MLcs.LGmath.STstat.TH
keywords maximummeandiscrepancyHilbert-SchmidtindependencecriterionkernelSteinminimaxlowerboundreproducingHilbertspaceembeddingperturbationsLeCammethod
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Kernel discrepancies such as MMD, HSIC and KSD are standard tools for comparing distributions, testing independence and checking goodness of fit. Existing estimators already achieve the parametric rate n^{-1/2} (or n^{-1/2}+m^{-1/2} for two-sample MMD) under mild conditions, even with unbounded kernels. This paper proves matching minimax lower bounds of the same order, holding on arbitrary topological spaces rather than only Euclidean space with restrictive kernels. The same lower bounds transfer immediately to estimation of the mean embedding and the centered cross-covariance operator. The result closes the optimality question for these widely used discrepancy measures.

What carries the argument

Le Cam’s two-point method applied to a carefully chosen adversarial pair of measures obtained by continuous bounded perturbations of a base measure; the construction forces the functional values to separate at rate n^{-1/2} while keeping the KL divergence between the product measures bounded.

What would settle it

Exhibit a topological space, a characteristic kernel, and a sequence of estimators whose risk is o(n^{-1/2}) uniformly over the stated class of measures, or prove that every continuous bounded function is almost surely constant for every measure in that class.

Watch

Extended reading notes

Core claim

Under the assumptions that the kernel is characteristic (or I-characteristic) and that there exists at least one probability measure admitting a non-almost-surely-constant continuous bounded perturbation, the minimax risk of estimating MMD is of exact order n^{-1/2}+m^{-1/2} and the risks of estimating HSIC and KSD are of exact order n^{-1/2}, on general topological spaces. The identical rates hold for the mean embedding and the centered cross-covariance operator.

Load-bearing premise

There must exist at least one probability measure in the class that is not almost-surely constant under every continuous bounded real function; without such a non-constant perturbation the two-point argument cannot be built.

Editorial extensions

If this is right

  • Existing U-statistic, V-statistic and accelerated estimators of MMD, HSIC and KSD are minimax optimal on far more general domains than previously known.
  • The same optimality statement holds for mean-embedding estimation and for estimation of the centered cross-covariance operator.
  • No estimator of these discrepancies can improve on the parametric rate under the stated mild conditions, even when kernels are unbounded.
  • The lower-bound technique applies uniformly across two-sample, independence and goodness-of-fit settings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same perturbation-plus-Le-Cam template should extend without change to other integral-probability-metric discrepancies once a characteristic property and a non-constant continuous function are available.
  • On spaces where every continuous function is constant almost everywhere (highly pathological topologies), the lower bound may fail and faster rates could become possible.
  • Practical kernel choice on non-Euclidean data (graphs, manifolds, sequences) can now safely target the parametric rate without fear that a cleverer estimator exists.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper establishes minimax lower bounds for estimating three kernel discrepancies — MMD (rate n^{-1/2} + m^{-1/2}), HSIC (n^{-1/2}), and KSD (n^{-1/2}) — on general topological spaces, plus corollaries for the mean embedding and the centered cross-covariance operator. The proofs use Le Cam's two-point method with adversarial measures built by multiplicative perturbations P^{(n)}(A) = ∫_A (1 + ε_n φ) dP_0, where ε_n = c n^{-1/2} and φ is a bounded continuous mean-zero perturbation whose existence is guaranteed by a non-degeneracy assumption (Assumptions 1(ii)/6(ii)/10(ii)). Separation of the functional values follows from injectivity of the mean embedding (characteristic / I-characteristic kernels), and the KL bound KL(P^{(n)} || P_0) ≤ α ε_n^2 is proved in Lemma B.6. The KSD result (Theorem 11) is recalled from the authors' prior AISTATS paper; the MMD, HSIC, and the two corollaries are new. Prior lower bounds were restricted to R^d with translation-invariant or radial kernels, so the generality here is the main contribution.

Significance. If the results hold — and I believe they do, modulo the framing issue in Major Comment 1 — the paper provides the first minimax lower bounds for MMD, HSIC, KSD, mean-embedding, and cross-covariance estimation that apply on general topological spaces with unbounded kernels, substantially relaxing the R^d / translation-invariant / radial assumptions of Tolstikhin et al. (2016, 2017), Chamakh and Szabó (2024), and Kalinke and Szabó (2024). The lower bounds are parameter-free in the relevant sense (the constant c in eps_n = c n^{-1/2} is arbitrary, and B > 0 is exhibited, not fitted) and rely only on characteristicness plus a mild non-degeneracy assumption. The proofs are short, checkable, and the assumptions are close to necessary for a nontrivial lower bound. This is a useful reference result for the kernel-discrepancy literature.

major comments (2)
  1. [§1, contribution (i); §3, Theorems 3, 7, 11; class P1(K;X) defined in §2] The claim in contribution (i) that the lower bounds 'match the upper ones of available estimators, hence settle their minimax optimality' for unbounded kernels is not fully supported as stated. The lower bounds are proved over P1(K;X) = {P : E_P sqrt(K(X,X)) < infty}. For unbounded K, the cited n^{-1/2} upper bounds (Kalinke et al. 2025b, 2026) require subexponentiality of sqrt(K(X,X)) under P, and even n^{-1/2} convergence of the empirical mean embedding requires the second moment E_P[K(X,X)] < infty. P1(K;X) contains measures with E[K(X,X)] = infty, over which no uniform n^{-1/2} upper bound exists. Two consequences: (a) over the stated class P1(K;X) with unbounded K, tightness at n^{-1/2} cannot be certified — the true minimax rate there may be slower; (b) a lower bound over the larger class P1 does not automatically transfer to the smaller subexponential classes where the upper bound
  2. [Abstract; §1, paragraph on convergence rates; §3.1–3.3] Related to the previous point: the abstract and §3 should state explicitly which class the optimality claim refers to. As written, Theorem 3 (and similarly 7, 11) lower-bounds the minimax risk over [P1(K;X)]^2, while the matching upper bounds cited in §1 hold only over moment/tail-restricted subclasses. A reader can currently read the paper as claiming optimality over P1(K;X) itself for unbounded kernels, which the arguments do not establish. Please align the statement of the class in the theorems, the upper-bound citations in §1, and the 'settle the question' sentence in the abstract; bounded-kernel cases (where P1 = M_1^+ and the upper bounds of Smola et al. 2007 apply) are fine as is.
minor comments (6)
  1. [§A.5] In the proof of Corollary 8 the map F is defined as F : P1(K;X) -> R, P -> C_K(P), but C_K(P) is an element of H_K; the codomain should be H_K.
  2. [§A.3] The equation label (A.9) is used both in the proof of Theorem 3 and again in the proof of Corollary 4; the second occurrence should be renumbered.
  3. [§3.1, Lemma 2] Lemma 2 (sufficient conditions for Assumption 1(ii)) assumes K in C_b(X^2), so it does not cover the unbounded-kernel regime that is the paper's main selling point. Assumption 1(ii) can in fact be verified much more cheaply whenever C_b(X) contains a non-constant function (e.g., X Tychonoff with at least two points): take phi_0 non-constant in C_b(X) and P0 a two-point mixture of Diracs at x, y with phi_0(x) != phi_0(y); finitely supported measures are always in P1(K;X). A remark to this effect would strengthen the 'mild assumptions' claim for unbounded kernels.
  4. [§1, paragraph on estimator convergence rates] In §1 the upper bounds for unbounded kernels are attributed to Kalinke et al. (2025b, 2026), which are KSD papers; please clarify which results provide the n^{-1/2} + m^{-1/2} (resp. n^{-1/2}) upper bounds for MMD (resp. HSIC) with unbounded kernels, or restrict that sentence to KSD.
  5. [§A.4] In (A.12), step (g) (C_phi > 0) uses that P^{(n)} != P0 via Lemma C.2; a pointer to Lemma C.2 at that step would parallel the treatment in (A.9) and help the reader.
  6. [Appendix A, Figure 1] The dependency chart is helpful, but 'C.5 C.3' in the header row is easy to misread; consider formatting the figure caption or layout so that external results (Appendix C) are visually distinguished from auxiliary ones (Appendix B).

Circularity Check

1 steps flagged · score 1.0 of 10

No significant circularity: standard Le Cam two-point construction; minor self-citation of prior KSD result is fully reproved here.

  1. self citation load bearing [p. 3 (contributions) and Theorem 11 / A.6]
    "This article expands the work of Cribeiro-Ramallo et al. (2026), which settled the minimax lower bound of KSD estimation on general topological spaces (recalled in Theorem 11, with proof for completeness)"

    The KSD lower bound is attributed to the authors' own prior paper. This is only a minor self-citation: the complete Le Cam argument is written out in A.6 under Assumption 10, so the present manuscript does not rely on an external unverified claim. Score contribution is therefore minimal.

full rationale

The claimed minimax rates are obtained by the classical Le Cam method: adversarial pairs are built by continuous bounded perturbations of a base measure (A.8, A.16) with free scale ε_n = c n^{-1/2} (c > 0 arbitrary), separation of the functional is shown to be Θ(ε_n) via injectivity/characteristicness plus the mean-embedding identity (Lemma B.5), and the product KL is O(1) by the second-order expansion of KL under small perturbations (Lemma B.6). Balancing separation against the KL budget forces the n^{-1/2} (resp. n^{-1/2}+m^{-1/2}) lower bound; nothing is fitted to data, and no equation equates the target rate to an input by definition. The only self-reference is that Theorem 11 restates the authors' earlier KSD lower bound; the full proof is reproduced in A.6 under the same mild assumptions, so the citation is not load-bearing. Corollaries for the mean embedding and centered cross-covariance follow by the reverse triangle inequality from the same pairs and are likewise self-contained. No self-definitional loop, fitted-as-prediction step, uniqueness import, or renamed empirical pattern appears.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claims rest on standard measure-theoretic and RKHS facts plus three domain assumptions (characteristic / I-characteristic kernels and existence of a non-constant continuous perturbation). No free parameters are fitted to data; the constants c, c1, c2 that scale the perturbations are arbitrary positive numbers chosen for the Le Cam argument. No new physical or statistical entities are postulated.

assumptions (5)
  • standard math Le Cam's two-point method (Theorem C.5 / Cam 1973; Tsybakov 2009): if two parameters are 2s-separated and their laws have KL ≤ α, then any estimator has risk at least f(α)s with positive probability.
    Invoked as the sole lower-bound engine in the proofs of Theorems 3, 7 and 11.
  • domain assumption Kernel K is characteristic (injective mean embedding) on P1(K;X) for MMD; product kernel is I-characteristic for HSIC; Stein kernel is characteristic w.r.t. P0 for KSD.
    Assumptions 1(i), 6(i), 10(i). Converts positive embedding distance of the adversarial pair into positive discrepancy; without it the distance control step fails.
  • domain assumption Existence of (P0, φ0) ∈ P1 × Cb(X) with φ0 not P0-a.s. constant (and the analogous product-space version).
    Assumptions 1(ii), 6(ii), 10(ii). Supplies the continuous perturbation used to build the adversarial measures via Lemma C.1.
  • standard math Bochner integrability of the canonical feature map under P ∈ P1(K;X); separability of the Stein RKHS when needed for measurability.
    Background RKHS/measure theory used to define mean embeddings and KSD (Section 2).
  • standard math KL of product measures factors as the sum of KLs (Lemma C.3 / Tsybakov).
    Used to bound KL(P_n^{(n)} ∥ P_0^n) by n times the single-copy KL.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Minimax Lower Bounds of Kernel Discrepancy Estimation: MMD, HSIC, KSD." pith.science (2026). https://pith.science/paper/2WBBHXPP

@misc{pith2026260724235,
  author       = {Pith},
  title        = {Pith review of: Minimax Lower Bounds of Kernel Discrepancy Estimation: MMD, HSIC, KSD},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2WBBHXPP}},
  note         = {Machine review of arXiv:2607.24235}
}
abstract

Over the past 20 years, kernel discrepancies have been leveraged as a highly powerful tool for quantifying the disagreement of distributions, with numerous successful applications in two-sample, goodness-of-fit, and independence testing, among others. Their fastest estimators are known to converge at a parametric rate---$n^{-1/2}$---under mild conditions. While this rate is known to be minimax optimal on $\mathbb R^d$ under strict assumptions with bounded kernels, little is known about its optimality beyond the finite-dimensional Euclidean setting with unbounded kernels. In this work, we prove that the minimax lower bound of estimation of the most popular kernel discrepancies (maximum mean discrepancy, Hilbert-Schmidt independence criterion and kernel Stein discrepancy; MMD, HSIC, KSD) is $n^{-1/2}$ on general topological spaces, and under mild assumptions on the kernel; the same rates are shown (as corollaries) to hold for the estimation of the mean embedding and the centered cross-covariance operator. Our results settle the question of optimal estimation of these kernel discrepancies.

Figures

Figures reproduced from arXiv: 2607.24235 by the authors.

Figure 1
Figure 1. Summary of the dependencies of our results. R1 ← R2 means that “result R1 depends on R2”. B.3, B.6, C.5 and C.3 are all used in Theorems 3, 7 and 11. A.1 Proof of Lemma 2 Let φ0 ∈ HK ⊆ Cb(X ) be the assumed non-constant element, where the inclusion HK ⊆ Cb(X ) holds by K ∈ Cb(X 2 ) (Steinwart and Christmann, 2008, Lemma 4.28). Further, as X is separable, there exists a countable dense subset {xi}∞ i=1 ⊆ X . Let P0 =… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

145 extracted references · 5 linked inside Pith

  1. [1]

    Adaptive test of independence based on HSIC measures

    M\' e lisande Albert, B\' e atrice Laurent, Amandine Marrel, and Anouar Meynaoui. Adaptive test of independence based on HSIC measures. Annals of Statistics, 50 0 (2): 0 858--879, 2022

  2. [2]

    Universal robust regression via maximum mean discrepancy

    Pierre Alquier and Mathieu Gerber. Universal robust regression via maximum mean discrepancy. Biometrika, 111 0 (1): 0 71--92, 2024

  3. [3]

    Estimation of copulas via maximum mean discrepancy

    Pierre Alquier, Badr-Eddine Chérief-Abdellatif, Alexis Derumigny, and Jean-David Fermanian. Estimation of copulas via maximum mean discrepancy. Journal of the American Statistical Association, 118 0 (543): 0 1997--2012, 2023

  4. [4]

    Gaunt, Fatemeh Ghaderinezhad, Jackson Gorham, Arthur Gretton, Christophe Ley, Qiang Liu, Lester Mackey, Chris J

    Andreas Anastasiou, Alessandro Barp, Fran c ois-Xavier Briol, Bruno Ebner, Robert E. Gaunt, Fatemeh Ghaderinezhad, Jackson Gorham, Arthur Gretton, Christophe Ley, Qiang Liu, Lester Mackey, Chris J. Oates, Gesine Reinert, and Yvik Swan. S tein's method meets computational statistics: a review of some recent developments. Statistical Science, 38 0 (1): 0 12...

  5. [5]

    Anderson, Peter Hall, and Donald M

    Niall H. Anderson, Peter Hall, and Donald M. Titterington. Two-sample test statistics for measuring discrepancies between two multivariate probability density functions using kernel-based density estimates. Journal of Multivariate Analysis, 50: 0 41--54, 1994

  6. [6]

    Maximum mean discrepancy gradient flow

    Michael Arbel, Anna Korba, Adil Salim, and Arthur Gretton. Maximum mean discrepancy gradient flow. In Advances in Neural Information Processing Systems (NeurIPS), pages 6484--6494, 2019

  7. [7]

    Theory of reproducing kernels

    Nachman Aronszajn. Theory of reproducing kernels. Transactions of the American Mathematical Society, 68: 0 337--404, 1950

  8. [8]

    Hilbertian Kernels and Spline Functions

    Marc Atteia. Hilbertian Kernels and Spline Functions. North-Holland Publishing Company, Amsterdam, 1992

Show all 145 references
  1. [9]

    On the optimality of kernel-embedding based goodness-of-fit tests

    Krishnakumar Balasubramanian, Tong Li, and Ming Yuan. On the optimality of kernel-embedding based goodness-of-fit tests. Journal of Machine Learning Research, 22 0 (1): 0 1--45, 2021

  2. [10]

    LeJEPA : Provable and scalable self-supervised learning without the heuristics

    Randall Balestriero and Yann LeCun. LeJEPA : Provable and scalable self-supervised learning without the heuristics. Technical report, 2025. (https://arxiv.org/abs/2511.08544)

  3. [11]

    Ludwig Baringhaus and C. Franz. On a new multivariate two-sample test. Journal of Multivariate Analysis, 88: 0 190--206, 2004

  4. [12]

    Alessandro Barp, Chris. J. Oates, Emilio Porcu, and Mark Girolami. A R iemann– S tein kernel method. Bernoulli, 28 0 (4): 0 2181 -- 2208, 2022

  5. [13]

    Targeted separation and convergence with kernel discrepancies

    Alessandro Barp, Carl-Johann Simon-Gabriel, Mark Girolami, and Lester Mackey. Targeted separation and convergence with kernel discrepancies. Journal of Machine Learning Research, 25 0 (378): 0 1--50, 2024

  6. [14]

    A kernel stein test of goodness of fit for sequential models

    Jerome Baum, Heishiro Kanagawa, and Arthur Gretton. A kernel stein test of goodness of fit for sequential models. In International Conference on Machine Learning (ICML), pages 1936--1953, 2023

  7. [15]

    Weighted quantization using MMD : From mean field to mean shift using gradient flows

    Ayoub Belhadji, Daniel Sharp, and Youssef Marzouk. Weighted quantization using MMD : From mean field to mean shift using gradient flows. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2026

  8. [16]

    Reproducing Kernel Hilbert Spaces in Probability and Statistics

    Alain Berlinet and Christine Thomas-Agnan. Reproducing Kernel Hilbert Spaces in Probability and Statistics. Kluwer, 2004

  9. [17]

    Tests of mutual or serial independence of random vectors with applications

    Martin Bilodeau and Aur \'e lien Guetsop Nangue. Tests of mutual or serial independence of random vectors with applications. Journal of Machine Learning Research, 18: 0 1--40, 2017

  10. [18]

    Graph kernels: State-of-the-art and future challenges

    Karsten Borgwardt, Elisabetta Ghisu, Felipe Llinares-L \'o pez, Leslie O'Bray, and Bastian Riec. Graph kernels: State-of-the-art and future challenges. Foundations and Trends in Machine Learning, 13 0 (5-6): 0 531--712, 2020

  11. [19]

    Borgwardt and Hans-Peter Kriegel

    Karsten M. Borgwardt and Hans-Peter Kriegel. Shortest-path kernels on graphs. In International Conference on Data Mining (ICDM), pages 74--81, 2005

  12. [20]

    Distribution free tests for model selection based on maximum mean discrepancy with estimated parameters

    Florian Br \"u ck, Jean-David Fermanian, and Aleksey Min. Distribution free tests for model selection based on maximum mean discrepancy with estimated parameters. Journal of Machine Learning Research, 26 0 (100): 0 1--52, 2025

  13. [21]

    Convergence of estimates under dimensionality restrictions

    Lucien Le Cam. Convergence of estimates under dimensionality restrictions. Annals of Statistics, 1: 0 38--53, 1973

  14. [22]

    Vector valued reproducing kernel H ilbert spaces and universality

    Claudio Carmeli, Ernesto De Vito, Alessandro Toigo, and Veronica Umanit \'a . Vector valued reproducing kernel H ilbert spaces and universality. Analysis and Applications, 8: 0 19--61, 2010

  15. [23]

    Distance metrics for measuring joint dependence with application to causal inference

    Shubhadeep Chakraborty and Xianyang Zhang. Distance metrics for measuring joint dependence with application to causal inference. Journal of the American Statistical Association, 114 0 (528): 0 1638--1650, 2019

  16. [24]

    Keep it tighter -- a story on analytical mean embeddings

    Linda Chamakh and Zolt \'a n Szab \'o . Keep it tighter -- a story on analytical mean embeddings. Technical report, 2024. (https://arxiv.org/abs/2110.09516)

  17. [25]

    Nystr \"o m kernel mean embeddings

    Antoine Chatalic, Nicolas Schreuder, Alessandro Rudi, and Lorenzo Rosasco. Nystr \"o m kernel mean embeddings. In International Conference on Machine Learning (ICML), pages 3006--3024, 2022

  18. [26]

    A scalable N ystr \"o m-based kernel two-sample test with permutations

    Antoine Chatalic, Marco Letizia, Nicolas Schreuder, and Lorenzo Rosasco. A scalable N ystr \"o m-based kernel two-sample test with permutations. Electronic Journal of Statistics, 20 0 (1): 0 2608--2642, 2026

  19. [27]

    Louis H. Y. Chen. S tein's method of normal approximation: Some recollections and reflections. Annals of Statistics, 49 0 (4): 0 1850--1863, 2021

  20. [28]

    Wilson Ye Chen, Lester Mackey, Jackson Gorham, Fran c ois-Xavier Briol, and Chris J. Oates. Stein points. In International Conference on Machine Learning (ICML), pages 844--853, 2018

  21. [29]

    Stein point M arkov chain M onte C arlo

    Wilson Ye Chen, Alessandro Barp, Fran c ois-Xavier Briol, Jackson Gorham, Mark Girolami, Lester Mackey, and Chris Oates. Stein point M arkov chain M onte C arlo. In International Conference on Machine Learning (ICML), pages 1011--1021, 2019

  22. [30]

    Kernel two-sample tests for manifold data

    Xiuyuan Cheng and Yao Xie. Kernel two-sample tests for manifold data. Bernoulli, 30 0 (4): 0 2572--2597, 2024

  23. [31]

    Signature moments to characterize laws of stochastic processes

    Ilya Chevyrev and Harald Oberhauser. Signature moments to characterize laws of stochastic processes. Journal of Machine Learning Research, 23 0 (176): 0 1--42, 2022

  24. [32]

    A kernel independence test for random processes

    Kacper Chwialkowski and Arthur Gretton. A kernel independence test for random processes. In International Conference on Machine Learning (ICML), pages 1422--1430, 2014

  25. [33]

    A kernel test of goodness of fit

    Kacper Chwialkowski, Heiko Strathmann, and Arthur Gretton. A kernel test of goodness of fit. In International Conference on Machine Learning (ICML), pages 2606--2615, 2016

  26. [34]

    Block HSIC L asso: Model-free biomarker detection for ultra-high dimensional data

    Héctor Climente-Gonz \'a lez, Chlo \'e -Agathe Azencott, Samuel Kaski, and Makoto Yamada. Block HSIC L asso: Model-free biomarker detection for ultra-high dimensional data. Bioinformatics, 35 0 (14): 0 i427--i435, 2019

  27. [35]

    Adversarial subspace generation for outlier detection in high-dimensional data

    Jose Cribeiro-Ramallo, Federico Matteucci, Paul Enciu, Alexander Jenke, Vadim Arzamasov, Thorsten Strufe, and Klemens B \"o hm. Adversarial subspace generation for outlier detection in high-dimensional data. Transactions on Machine Learning Research (TMLR), 2025

  28. [36]

    The minimax lower bound of kernel S tein discrepancy estimation

    Jose Cribeiro-Ramallo, Agnideep Aich, Florian Kalinke, Ashit Baran Aich, and Zolt \'a n Szab \'o . The minimax lower bound of kernel S tein discrepancy estimation. In International Conference on Artificial Intelligence and Statistics (AISTATS), 2026

  29. [37]

    Real Analysis and Probability

    Richard Dudley. Real Analysis and Probability. Cambridge University Press, 2004

  30. [38]

    Roy, and Zoubin Ghahramani

    Gintare Karolina Dziugaite, Daniel M. Roy, and Zoubin Ghahramani. Training generative neural networks via maximum mean discrepancy optimization. In Conference on Uncertainty in Artificial Intelligence (UAI), page 258–267, 2015

  31. [39]

    A maximum-mean-discrepancy goodness-of-fit test for censored data

    Tamara Fernandez and Arthur Gretton. A maximum-mean-discrepancy goodness-of-fit test for censored data. In International Conference on Artificial Intelligence and Statistics (AISTATS), pages 2966--2975, 2019

  32. [40]

    Kernelized Stein discrepancy tests of goodness-of-fit for time-to-event data

    Tamara Fernandez, Nicolas Rivera, Wenkai Xu, and Arthur Gretton. Kernelized Stein discrepancy tests of goodness-of-fit for time-to-event data. In International Conference on Machine Learning (ICML), pages 3112--3122, 2020

  33. [41]

    Gerald B. Folland. Real Analysis -- Modern Techniques and Their Applications. John Wiley & Sons, 1999

  34. [42]

    Kernel measures of conditional dependence

    Kenji Fukumizu, Arthur Gretton, Xiaohai Sun, and Bernhard Sch \"o lkopf. Kernel measures of conditional dependence. In Advances in Neural Information Processing Systems (NeurIPS), pages 498--496, 2008

  35. [43]

    Bayesian posterior approximation via greedy particle optimization

    Futoshi Futami, Zhenghang Cui, Issei Sato, and Masashi Sugiyama. Bayesian posterior approximation via greedy particle optimization. In AAAI Conference on Artificial Intelligence ( AAAI ) , pages 3606--3613, 2019

  36. [44]

    Multi-instance kernels

    Thomas G \"a rtner, Peter Flach, Adam Kowalczyk, and Alexander Smola. Multi-instance kernels. In International Conference on Machine Learning (ICML), pages 179--186, 2002

  37. [45]

    Interaction-force transport gradient flows

    Egor Gladin, Pavel Dvurechensky, Alexander Mielke, and Jia-Jie Zhu. Interaction-force transport gradient flows. In Advances in Neural Information Processing Systems (NeurIPS), pages 14484--14508, 2024

  38. [46]

    Minimax estimation of kernel S tein discrepancy: Trace versus H ilbert- S chmidt scales

    Davit Gogolashvili. Minimax estimation of kernel S tein discrepancy: Trace versus H ilbert- S chmidt scales. Technical report, 2026. (https://arxiv.org/abs/2607.03367)

  39. [47]

    Measuring sample quality with kernels

    Jackson Gorham and Lester Mackey. Measuring sample quality with kernels. In International Conference on Machine Learning (ICML), pages 1292--1301, 2017

  40. [48]

    Measuring statistical dependence with H ilbert- S chmidt norms

    Arthur Gretton, Olivier Bousquet, Alex Smola, and Bernhard Sch \"o lkopf. Measuring statistical dependence with H ilbert- S chmidt norms. In Algorithmic Learning Theory (ALT), pages 63--78, 2005 a

  41. [49]

    Kernel methods for measuring independence

    Arthur Gretton, Ralf Herbrich, Alexander Smola, Olivier Bousquet, and Bernhard Sch \"o lkopf. Kernel methods for measuring independence. Journal of Machine Learning Research, 6 0 (70): 0 2075--2129, 2005 b

  42. [50]

    A kernel statistical test of independence

    Arthur Gretton, Kenji Fukumizu, Choon Hui Teo, Le Song, Bernhard Sch \"o lkopf, and Alexander Smola. A kernel statistical test of independence. In Advances in Neural Information Processing Systems (NeurIPS), pages 585--592, 2008

  43. [51]

    A kernel two-sample test

    Arthur Gretton, Karsten Borgwardt, Malte Rasch, Bernhard Sch \"o lkopf, and Alexander Smola. A kernel two-sample test. Journal of Machine Learning Research, 13 0 (25): 0 723--773, 2012

  44. [52]

    Cross product kernels for fuzzy set similarity

    Jorge Guevara, Roberto Hirata, and St \'e phane Canu. Cross product kernels for fuzzy set similarity. In International Conference on Fuzzy Systems (FUZZ-IEEE), pages 1--6, 2017

  45. [53]

    Minimax optimal goodness-of-fit testing with kernel S tein discrepancy

    Omar Hagrass, Bharath Sriperumbudur, and Krishnakumar Balasubramanian. Minimax optimal goodness-of-fit testing with kernel S tein discrepancy. Bernoulli, 32 0 (1): 0 299--324, 2026

  46. [54]

    Convolution kernels on discrete structures

    David Haussler. Convolution kernels on discrete structures. Technical report, University of California at Santa Cruz, 1999. (http://cbse.soe.ucsc.edu/sites/default/files/convolutions.pdf)

  47. [55]

    Hilbertian metrics and positive definite kernels on probability measures

    Matthias Hein and Olivier Bousquet. Hilbertian metrics and positive definite kernels on probability measures. In International Conference on Artificial Intelligence and Statistics (AISTATS), pages 136--143, 2005

  48. [56]

    The reproducing S tein kernel approach for post-hoc corrected sampling

    Liam Hodgkinson, Robert Salomone, and Fred Roosta. The reproducing S tein kernel approach for post-hoc corrected sampling. Technical report, 2021. (https://arxiv.org/abs/2001.09266)

  49. [57]

    The K endall and M allows kernels for permutations

    Yunlong Jiao and Jean-Philippe Vert. The K endall and M allows kernels for permutations. In International Conference on Machine Learning (ICML), volume 37, pages 2982--2990, 2016

  50. [58]

    Optimal online change detection via random F ourier features

    Florian Kalinke and Shakeel Gavioli-Akilagun. Optimal online change detection via random F ourier features. In Advances in Neural Information Processing Systems, pages 9866--9901, 2025

  51. [59]

    Nystr \"o m M - H ilbert- S chmidt independence criterion

    Florian Kalinke and Zolt \'a n Szab \'o . Nystr \"o m M - H ilbert- S chmidt independence criterion. In Conference on Uncertainty in Artificial Intelligence (UAI), pages 1005--1015, 2023

  52. [60]

    The minimax rate of HSIC estimation for translation-invariant kernels

    Florian Kalinke and Zolt \' a n Szab \' o . The minimax rate of HSIC estimation for translation-invariant kernels. In Advances in Neural Information Processing Systems (NeurIPS), pages 108468--108489, 2024

  53. [61]

    Maximum mean discrepancy on exponential windows for online change detection

    Florian Kalinke, Marco Heyden, Georg Gntuni, Edouard Fouch \'e , and Klemens B \"o hm. Maximum mean discrepancy on exponential windows for online change detection. Transactions on Machine Learning Research, 2025 a

  54. [62]

    Sriperumbudur

    Florian Kalinke, Zolt \' a n Szab \' o , and Bharath K. Sriperumbudur. Nyström kernel S tein discrepancy. In International Conference on Artificial Intelligence and Statistics (AISTATS), pages 388--396, 2025 b

  55. [63]

    Sriperumbudur

    Florian Kalinke, Zolt \' a n Szab \' o , and Bharath K. Sriperumbudur. Nyström kernel S tein discrepancy tests. Technical report, 2026. (https://arxiv.org/abs/2605.25173)

  56. [64]

    A kernel S tein test for comparing latent variable models

    Heishiro Kanagawa, Wittawat Jitkrittum, Lester Mackey, Kenji Fukumizu, and Arthur Gretton. A kernel S tein test for comparing latent variable models. Journal of the Royal Statistical Society Series B: Statistical Methodology, 85 0 (3): 0 986--1011, 2023

  57. [65]

    Deep signature transforms

    Patrick Kidger, Patric Bonnier, Imanol Perez Arribas, Cristopher Salvi, and Terry Lyons. Deep signature transforms. In Advances in Neural Information Processing Systems (NeurIPS), pages 3105---3115, 2019

  58. [66]

    Kir \'a ly and Harald Oberhauser

    Franz J. Kir \'a ly and Harald Oberhauser. Kernels for sequentially ordered data. Journal of Machine Learning Research, 20: 0 1--45, 2019

  59. [67]

    N-Distances and Their Applications

    Lev Klebanov. N-Distances and Their Applications. Charles University, Prague, 2005

  60. [68]

    The multiscale L aplacian graph kernel

    Risi Kondor and Horace Pan. The multiscale L aplacian graph kernel. In Advances in Neural Information Processing Systems (NeurIPS), pages 2982--2990, 2016

  61. [69]

    Kondor and John Lafferty

    Risi I. Kondor and John Lafferty. Diffusion kernels on graphs and other discrete input. In International Conference on Machine Learning (ICML), pages 315--322, 2002

  62. [70]

    A non-asymptotic analysis for S tein variational gradient descent

    Anna Korba, Adil Salim, Michael Arbel, Giulia Luise, and Arthur Gretton. A non-asymptotic analysis for S tein variational gradient descent. In Advances in Neural Information Processing Systems (NeurIPS), pages 4672--4682, 2020

  63. [71]

    Kernel S tein discrepancy descent

    Anna Korba, Pierre-Cyril Aubin-Frankowski, Szymon Majewski, and Pierre Ablin. Kernel S tein discrepancy descent. In International Conference on Machine Learning (ICML), pages 5719--5730, 2021

  64. [72]

    Kozarzewski

    Piotr A. Kozarzewski. On existence of the support of a B orel measure. Demonstratio Mathematica, 51 0 (1): 0 76--84, 2018

  65. [73]

    Kernel two-sample and independence tests for nonstationary random processes

    Felix Laumann, Julius von Kügelgen, and Mauricio Barahona. Kernel two-sample and independence tests for nonstationary random processes. Engineering Proceedings, 5 0 (1), 2021

  66. [74]

    MMD GAN : Towards Deeper Understanding of Moment Matching Network

    Chun-Liang Li, Wei-Cheng Chang, Yu Cheng, Yiming Yang, and Barnab \'a s P \'o czos. MMD GAN : Towards Deeper Understanding of Moment Matching Network . In Advances in Neural Information Processing Systems (NeurIPS), page 2200–2210, 2017

  67. [75]

    Debiased distribution compression

    Lingxiao Li, Raaz Dwivedi, and Lester Mackey. Debiased distribution compression. In International Conference on Machine Learning (ICML), pages 27675--27731, 2024

  68. [76]

    M-statistic for kernel change-point detection

    Shuang Li, Yao Xie, Hanjun Dai, and Le Song. M-statistic for kernel change-point detection. In Advances in Neural Information Processing Systems (NeurIPS), page 3366–3374, 2015 a

  69. [77]

    Scan B -statistic for kernel change-point detection

    Shuang Li, Yao Xie, Hanjun Dai, and Le Song. Scan B -statistic for kernel change-point detection. Sequential Analysis, 38: 0 503--544, 2019

  70. [78]

    Generative moment matching networks

    Yujia Li, Kevin Swersky, and Richard Zemel. Generative moment matching networks. In International Conference on Machine Learning (ICML), pages 1718--1727, 2015 b

  71. [79]

    Kernel S tein tests for multiple model comparison

    Jen Ning Lim, Makoto Yamada, Bernhard Sch \"o lkopf, and Wittawat Jitkrittum. Kernel S tein tests for multiple model comparison. In Advances in Neural Information Processing Systems (NeurIPS), pages 2243--2253, 2019

  72. [80]

    Sutherland

    Feng Liu, Wenkai Xu, Jie Liu, Guangquan Zhang, Arthur Gretton, and Danica J. Sutherland. Learning deep kernels for non-parametric two-sample tests. In International Conference on Machine Learning (ICML), pages 6316--6326, 2020

  73. [81]

    Stein variational gradient descent: A general purpose B ayesian inference algorithm

    Qiang Liu and Dilin Wang. Stein variational gradient descent: A general purpose B ayesian inference algorithm. In Advances in Neural Information Processing Systems (NeurIPS), pages 2378--2386, 2016

  74. [82]

    Stein variational gradient descent as moment matching

    Qiang Liu and Dilin Wang. Stein variational gradient descent as moment matching. In Advances in Neural Information Processing Systems (NeurIPS), pages 8854--8863, 2018

  75. [83]

    A kernelized Stein discrepancy for goodness-of-fit tests

    Qiang Liu, Jason Lee, and Michael Jordan. A kernelized Stein discrepancy for goodness-of-fit tests. In International Conference on Machine Learning (ICML), pages 276--284, 2016

  76. [84]

    On the robustness of kernel goodness-of-fit tests

    Xing Liu and Fran c ois-Xavier Briol. On the robustness of kernel goodness-of-fit tests. Journal of Machine Learning Research, 26 0 (262): 0 1--72, 2025

  77. [85]

    Text classification using string kernels

    Huma Lodhi, Craig Saunders, John Shawe-Taylor, Nello Cristianini, and Chris Watkins. Text classification using string kernels. Journal of Machine Learning Research, 2: 0 419--444, 2002

  78. [86]

    Distance covariance in metric spaces

    Russell Lyons. Distance covariance in metric spaces. The Annals of Probability, 41: 0 3284--3305, 2013

  79. [87]

    Wainwright, Michael I

    Horia Mania, Aaditya Ramdas, Martin J. Wainwright, Michael I. Jordan, and Benjamin Recht. On kernel methods for covariates that are rankings. Electronic Journal of Statistics, 12 0 (2): 0 2537--2577, 2018

  80. [88]

    High-dimensional multi-task averaging and application to kernel mean embedding

    Hannah Marienwald, Jean-Baptiste Fermanian, and Gilles Blanchard. High-dimensional multi-task averaging and application to kernel mean embedding. In Proceedings of The 24th International Conference on Artificial Intelligence and Statistics, pages 1963--1971, 2021

  81. [89]

    Sequential kernelized S tein discrepancy

    Diego Martinez-Taboada and Aaditya Ramdas. Sequential kernelized S tein discrepancy. In International Conference on Artificial Intelligence and Statistics (AISTATS), pages 1288--1296, 2025

  82. [90]

    H-SPLID : HSIC -based saliency preserving latent information decomposition

    Lukas Miklautz, Chengzhi Shi, Andrii Shkabrii, Theodoros Thirimachos Davarakis, Prudence Lam, Claudia Plant, Jennifer Dy, and Stratis Ioannidis. H-SPLID : HSIC -based saliency preserving latent information decomposition. In Advances in Neural Information Processing Systems (Ne...

  83. [91]

    Distinguishing cause from effect using observational data: Methods and benchmarks

    Joris Mooij, Jonas Peters, Dominik Janzing, Jakob Zscheischler, and Bernhard Sch \"o lkopf. Distinguishing cause from effect using observational data: Methods and benchmarks. Journal of Machine Learning Research, 17: 0 1--102, 2016

  84. [92]

    Integral probability metrics and their generating classes of functions

    Alfred M \"u ller. Integral probability metrics and their generating classes of functions. Advances in Applied Probability, 29: 0 429--443, 1997

  85. [93]

    James R. Munkres. Topology. Prentice Hall, Inc., Upper Saddle River, NJ, 2000

  86. [94]

    Graph alignment kernels using W eisfeiler and L eman hierarchies

    Giannis Nikolentzos and Michalis Vazirgiannis. Graph alignment kernels using W eisfeiler and L eman hierarchies. In International Conference on Artificial Intelligence and Statistics (AISTATS), pages 2019--2034, 2023

  87. [95]

    Oates, Mark Girolami, and Nicolas Chopin

    Chris J. Oates, Mark Girolami, and Nicolas Chopin. Control functionals for M onte C arlo integration. Journal of the Royal Statistical Society Series B: Statistical Methodology, 79 0 (3): 0 695--718, 2017

  88. [96]

    u hlmann, Bernhard Sch \

    Niklas Pfister, Peter B \"u hlmann, Bernhard Sch \"o lkopf, and Jonas Peters. Kernel-based tests for joint independence. Journal of the Royal Statistical Society Series B: Statistical Methodology, 80 0 (1): 0 5--31, 2018

  89. [97]

    Sequential kernelized independence testing

    Aleksandr Podkopaev, Patrick Bl\" o baum, Shiva Kasiviswanathan, and Aaditya Ramdas. Sequential kernelized independence testing. In International Conference on Machine Learning ( ICML ) , pages 27957--27993, 2023

  90. [98]

    Kernelized sorting

    Novi Quadrianto, Le Song, and Alex Smola. Kernelized sorting. In Advances in Neural Information Processing Systems (NeurIPS), pages 1289--1296, 2009

  91. [99]

    Efficiently learning significant Fourier feature pairs for statistical independence testing

    Yixin Ren, Yewei Xia, Hao Zhang, Jihong Guan, and Shuigeng Zhou. Efficiently learning significant Fourier feature pairs for statistical independence testing. In Advances in Neural Information Processing Systems (NeurIPS), pages 99800--99835, 2024

  92. [100]

    Regression-based conditional independence test with adaptive kernels

    Yixin Ren, Juncai Zhang, Yewei Xia, Ruxin Wang, Feng Xie, Jihong Guan, Hao Zhang, and Shuigeng Zhou. Regression-based conditional independence test with adaptive kernels. Artificial Intelligence, 347 0 (C), 2025

  93. [101]

    Theory of Reproducing Kernels and Applications

    Saburou Saitoh and Yoshihiro Sawano. Theory of Reproducing Kernels and Applications. Springer Singapore, 2016

  94. [102]

    Data-centric prediction explanation via kernelized S tein discrepancy

    Mahtab Sarvmaili, Hassan Sajjad, and Ga Wu. Data-centric prediction explanation via kernelized S tein discrepancy. In International Conference on Learning Representations (ICLR), 2025

  95. [103]

    KSD aggregated goodness-of-fit test

    Antonin Schrab, Benjamin Guedj, and Arthur Gretton. KSD aggregated goodness-of-fit test. In Advances in Neural Information Processing Systems (NeurIPS), pages 32624--32638, 2022

  96. [104]

    M MD aggregated two-sample test

    Antonin Schrab, Ilmun Kim, M\' e lisande Albert, B\' e atrice Laurent, Benjamin Guedj, and Arthur Gretton. M MD aggregated two-sample test. Journal of Machine Learning Research, 24 0 (194): 0 1--81, 2023

  97. [105]

    Graph filtration kernels

    Till Hendrik Schulz, Pascal Welke, and Stefan Wrobel. Graph filtration kernels. In AAAI Conference on Artifical Intelligence (AAAI), pages 8196--8203, 2022

  98. [106]

    A kernel test for three-variable interactions

    Dino Sejdinovic, Arthur Gretton, and Wicher Bergsma. A kernel test for three-variable interactions. In Advances in Neural Information Processing Systems (NeurIPS), pages 1124--1132, 2013 a

  99. [107]

    Equivalence of distance-based and RKHS -based statistics in hypothesis testing

    Dino Sejdinovic, Bharath Sriperumbudur, Arthur Gretton, and Kenji Fukumizu. Equivalence of distance-based and RKHS -based statistics in hypothesis testing. Annals of Statistics, 41: 0 2263--2291, 2013 b

  100. [108]

    A permutation-free kernel independence test

    Shubhanshu Shekhar, Ilmun Kim, and Aaditya Ramdas. A permutation-free kernel independence test. Journal of Machine Learning Research, 24 0 (369): 0 1--68, 2023

  101. [109]

    Kernel distribution embeddings: Universal kernels, characteristic kernels and kernel metrics on distributions

    Carl-Johann Simon-Gabriel and Bernhard Sch \"o lkopf. Kernel distribution embeddings: Universal kernels, characteristic kernels and kernel metrics on distributions. Journal of Machine Learning Research, 19 0 (44): 0 1--29, 2018

  102. [110]

    A H ilbert space embedding for distributions

    Alexander Smola, Arthur Gretton, Le Song, and Bernhard Sch \"o lkopf. A H ilbert space embedding for distributions. In Algorithmic Learning Theory (ALT), pages 13--31, 2007

  103. [111]

    Smola, Arthur Gretton, and Karsten M

    Le Song, Alexander J. Smola, Arthur Gretton, and Karsten M. Borgwardt. A dependence maximization view of clustering. In International Conference on Machine Learning (ICML), pages 815--822, 2007

  104. [112]

    Feature selection via dependence maximization

    Le Song, Alex Smola, Arthur Gretton, Justin Bedo, and Karsten Borgwardt. Feature selection via dependence maximization. Journal of Machine Learning Research, 13 0 (1): 0 1393--1434, 2012

  105. [113]

    Hilbert space embeddings and metrics on probability measures

    Bharath Sriperumbudur, Arthur Gretton, Kenji Fukumizu, Bernhard Sch \"o lkopf, and Gert Lanckriet. Hilbert space embeddings and metrics on probability measures. Journal of Machine Learning Research, 11: 0 1517--1561, 2010

  106. [114]

    A bound for the error in the normal approximation to the distribution of a sum of dependent random variables

    Charles Stein. A bound for the error in the normal approximation to the distribution of a sum of dependent random variables. In Berkeley Symposium on Mathematical Statistics and Probability, pages 583--602, 1972

  107. [115]

    Support Vector Machines

    Ingo Steinwart and Andreas Christmann. Support Vector Machines. Springer, 2008

  108. [116]

    Strictly proper kernel scores and characteristic kernels on compact spaces

    Ingo Steinwart and Johanna Ziegel. Strictly proper kernel scores and characteristic kernels on compact spaces. Applied and Computational Harmonic Analysis, 51: 0 510--542, 2021

  109. [117]

    Sutherland, Hsiao-Yu Tung, Heiko Strathmann, Soumyajit De, Aaditya Ramdas, Alex Smola, and Arthur Gretton

    Danica J. Sutherland, Hsiao-Yu Tung, Heiko Strathmann, Soumyajit De, Aaditya Ramdas, Alex Smola, and Arthur Gretton. Generative models and model criticism via optimized maximum mean discrepancy. In International Conference on Learning Representations (ICLR), 2017

  110. [118]

    Sutherland

    Wilson A. Sutherland. Introduction to M etric and T opological S paces . Oxford University Press, Oxford, second edition, 2009

  111. [119]

    Sriperumbudur

    Zolt \'a n Szab \'o and Bharath K. Sriperumbudur. Characteristic and universal tensor product kernels. Journal of Machine Learning Research, 18 0 (233): 0 1--29, 2018

  112. [120]

    Testing for equal distributions in high dimension

    G \'a bor Sz \'e kely and Maria Rizzo. Testing for equal distributions in high dimension. InterStat, 5: 0 1249--1272, 2004

  113. [121]

    A new test for multivariate normality

    G \'a bor Sz \'e kely and Maria Rizzo. A new test for multivariate normality. Journal of Multivariate Analysis, 93: 0 58--80, 2005

  114. [122]

    Minimax estimation of maximal mean discrepancy with radial kernels

    Ilya Tolstikhin, Bharath Sriperumbudur, and Bernhard Sch \"o lkopf. Minimax estimation of maximal mean discrepancy with radial kernels. In Advances in Neural Information Processing Systems (NeurIPS), pages 1930--1938, 2016

  115. [123]

    Minimax estimation of kernel mean embeddings

    Ilya Tolstikhin, Bharath Sriperumbudur, and Krikamol Muandet. Minimax estimation of kernel mean embeddings. Journal of Machine Learning Research, 18: 0 1--47, 2017

  116. [124]

    Random F ourier signature features

    Csaba T\' o th, Harald Oberhauser, and Zolt\' a n Szab\' o . Random F ourier signature features. SIAM Journal on Mathematics of Data Science , 7 0 (1): 0 329--354, 2025

  117. [125]

    Tsybakov

    Alexandre B. Tsybakov. Introduction to Nonparametric Estimation. Springer, 2009

  118. [126]

    Spline Models for Observational Data

    Grace Wahba. Spline Models for Observational Data. SIAM, CBMS-NSF Regional Conference Series in Applied Mathematics, 1990

  119. [127]

    Congye Wang, Wilson Ye Chen, Heishiro Kanagawa, and Chris J. Oates. Stein -importance sampling. In Advances in Neural Information Processing Systems (NeurIPS), pages 71948--71994, 2023

  120. [128]

    Dynamic alignment kernels

    Chris Watkins. Dynamic alignment kernels. In Advances in Neural Information Processing Systems (NeurIPS), pages 39--50, 1999

  121. [129]

    Nonparametric independence testing for small sample sizes

    Leila Wehbe and Aaditya Ramdas. Nonparametric independence testing for small sample sizes. In International Joint Conference on Artificial Intelligence ( IJCAI ) , pages 3777--3783, 2015

  122. [130]

    George Wynne and Andrew B. Duncan. A kernel two-sample test for functional data. Journal of Machine Learning Research, 23: 0 1--51, 2022

  123. [131]

    Statistical depth meets machine learning: Kernel mean embeddings and depth in functional data analysis

    George Wynne and Stanislav Nagy. Statistical depth meets machine learning: Kernel mean embeddings and depth in functional data analysis. International Statistical Review, 93 0 (2): 0 317--348, 2025

  124. [132]

    Kasprzak, and Andrew B

    George Wynne, Miko aj J. Kasprzak, and Andrew B. Duncan. A F ourier representation of kernel S tein discrepancy with application to goodness-of-fit tests for measures on infinite dimensional H ilbert spaces. Bernoulli, 31 0 (2): 0 868--893, 2025

  125. [133]

    A S tein goodness-of-fit test for directional distributions

    Wenkai Xu and Takeru Matsuda. A S tein goodness-of-fit test for directional distributions. In International Conference on Artificial Intelligence and Statistics (AISTATS), pages 320--330, 2020

  126. [134]

    Interpretable S tein goodness-of-fit tests on R iemannian manifold

    Wenkai Xu and Takeru Matsuda. Interpretable S tein goodness-of-fit tests on R iemannian manifold. In International Conference on Machine Learning (ICML), pages 11502--11513, 2021

  127. [135]

    A S tein goodness-of-test for exponential random graph models

    Wenkai Xu and Gesine Reinert. A S tein goodness-of-test for exponential random graph models. In International Conference on Artificial Intelligence and Statistics ( AISTATS ) , pages 415--423, 2021

  128. [136]

    Goodness-of-fit testing for discrete distributions via S tein discrepancy

    Jiasen Yang, Qiang Liu, Vinayak Rao, and Jennifer Neville. Goodness-of-fit testing for discrete distributions via S tein discrepancy. In International Conference on Machine Learning ( ICML ) , pages 5561--5570, 2018

  129. [137]

    Rao, and Jennifer Neville

    Jiasen Yang, Vinayak A. Rao, and Jennifer Neville. A Stein - Papangelou goodness-of-fit test for point processes. In International Conference on Artificial Intelligence and Statistics ( AISTATS ) , pages 226--235, 2019

  130. [138]

    A fast and accurate kernel-based independence test with applications to high-dimensional and functional data

    Jin-Ting Zhang and Tianming Zhu. A fast and accurate kernel-based independence test with applications to high-dimensional and functional data. Journal of Multivariate Analysis, 202: 0 105320, 2024

  131. [139]

    Robust object detection in adverse weather with feature decorrelation via independence learning

    Yutao Zhang, Shiyu Xuan, and Zechao Li. Robust object detection in adverse weather with feature decorrelation via independence learning. Pattern Recognition, 169: 0 111790, 2026

  132. [140]

    Inductive moment matching

    Linqi Zhou, Stefano Ermon, and Jiaming Song. Inductive moment matching. In International Conference on Machine Learning (ICML), pages 78651--78686, 2025 a

  133. [141]

    A class of optimal estimators for the covariance operator in reproducing kernel H ilbert spaces

    Yang Zhou, Di-Rong Chen, and Wei Huang. A class of optimal estimators for the covariance operator in reproducing kernel H ilbert spaces. Journal of Multivariate Analysis, 169: 0 166--178, 2019

  134. [142]

    Sutherland, and Feng Liu

    Zhijian Zhou, Xunye Tian, Liuhua Peng, Chao Lei, Antonin Schrab, Danica J. Sutherland, and Feng Liu. DUAL : Learning diverse kernels for aggregated two-sample and independence testing. In Advances in Neural Information Processing Systems (NeurIPS), pages 130661--130695, 2025 b

  135. [143]

    KerJEPA : Kernel discrepancies for Euclidean self-supervised learning

    Eric Zimmermann, Harley Wiltzer, Justin Szeto, David Alvarez-Melis, and Lester Mackey. KerJEPA : Kernel discrepancies for Euclidean self-supervised learning. Technical report, 2025. (https://arxiv.org/abs/2512.19605)

  136. [144]

    Zinger, Ashot Kakosyan, and Lev Klebanov

    Abram A. Zinger, Ashot Kakosyan, and Lev Klebanov. A characterization of distributions by mean values of statistics and certain probabilistic metrics. Journal of Soviet Mathematics, 1992

  137. [145]

    Probability metrics

    Vladimir Zolotarev. Probability metrics. Theory of Probability and its Applications, 28: 0 278--302, 1983

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.