Pith. sign in

REVIEW 2 major objections 4 minor 2 cited by

A Dictionary of Closed-Form Kernel Mean Embeddings

T0 review · 2 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper compiles known closed-form kernel mean embeddings for common kernel–distribution pairs, adds new entries for Wendland and fractional Brownian motion kernels, and supplies derivation rules and a unit-tested Python library.

desk verdict Useful reference dictionary of known kernel mean embeddings, but the new Wendland/uniform entry is wrong as typeset, a real problem for a paper whose purpose is to be copied from. read the letter →

arxiv 2504.18830 v1 pith:J74ZNQOW submitted 2025-04-26 stat.ML cs.LGcs.NAmath.NAstat.CO

classification stat.MLcs.LGcs.NAmath.NAstat.CO
keywords kernelmeanembeddingBayesianquadraturemaximumdiscrepancyMatérnWendlandfractionalBrownianmotionSteinreproducingclosed-formintegration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is trying to establish that the practical bottleneck in kernel-based integration and two-sample inference—the need for closed-form kernel mean embeddings—can be removed for a wide set of standard kernel and distribution pairs, and that further embeddings can be built from these by composition rules. It collects previously scattered formulas into one dictionary, computes new entries for Wendland and fractional Brownian motion kernels, and packages the results in a Python library with unit tests. A sympathetic reader would care because Bayesian quadrature, worst-case error bounds, and maximum mean discrepancy inference become directly implementable for the covered pairs without rederiving integrals or resorting to sampling approximations.

What carries the argument

The central object is the kernel mean embedding $K_P(x)=\int_\Omega K(x,y)\,dP(y)$ together with its double integral $K_{PP}$; these are the two quantities that Bayesian quadrature, kernel quadrature, and MMD-based tests need in closed form. The paper's machinery is a table of explicit formulas for these objects, supplemented by four composition rules: product kernels with product distributions multiply, sum kernels with mixture distributions add, a change of measure converts an intractable embedding into a known one with weights $p/q$, and a change of variable $K^\phi(x,y)=K(\phi(x),\phi(y))$ moves embeddings across pushforward distributions. Stein kernels work in the opposite direction: instead of deriving an embedding for a given kernel, they define a kernel whose embedding is identically zero, giving closed forms for unnormalised densities through automatic differentiation of the score function.

What would settle it

Take Equation (35), the Wendland order-zero embedding for a centered Gaussian measure, evaluate it at a few points such as $x=0$ with $\ell=\sigma=1$, and compare against direct high-precision numerical integration of $\int (1-|x-y|/\ell)_+\,dP(y)$; any mismatch beyond integration tolerance would refute the claim that the dictionary entry is correct.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that the pair of objects $K_P(x)=\int_\Omega K(x,y)\,dP(y)$ and $K_{PP}=\int_\Omega\int_\Omega K(x,y)\,dP(x)\,dP(y)$ have tractable closed forms for a substantial collection of kernel and distribution pairs, and that these forms can be organized by kernel family. The dictionary covers Gaussian, Matérn, Wendland, fractional Brownian motion, power-series, sphere, and periodic Sobolev kernels against uniform, Gaussian, and spherical measures, with explicit formulas for both the embedding and its double integral where known. The paper also shows that kernels can be designed so that the embedding becomes trivial: Langevin Stein reproducing kernels satisfy $\widetilde{K}_P(x)=\widetilde{K}_{PP}=0$ for any sufficiently regular distribution with an available score function. The paper's claim is that the listed formulas are correct and that the accompanying implementation mirrors them, with the library's unit tests serving as numerical checks of the identities.

Load-bearing premise

The dictionary's correctness stands on the exactness of every listed formula; for the newly added Wendland and fractional Brownian motion entries, the paper's own support is numerical agreement with integration routines rather than a written proof, so a single transcription error in a formula would falsify that entry.

Editorial extensions

If this is right

  • Bayesian quadrature and kernel quadrature can be applied directly to uniform and Gaussian targets from the listed formulas, without Monte Carlo approximation of the embedding or the variance term.
  • MMD-based two-sample and goodness-of-fit tests gain exact computable values for the covered kernel–distribution pairs, removing the sampling noise that enters when embeddings are estimated.
  • The product, mixture, change-of-measure, and change-of-variable rules let users assemble embeddings for distributions not explicitly listed, so the dictionary extends beyond its table of entries.
  • Stein reproducing kernels provide closed-form embeddings for any distribution with an available score function, including Bayesian posteriors known only up to a normalising constant, at the cost of using a kernel tailored to the target distribution.
  • The Python library's unit tests double as numerical checks of every formula, letting practitioners move from identity to implementation with less risk of transcription error.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, the change-of-variable rule points to a cheap way to obtain embeddings for distributions defined only by samplers or generative models: pair a known uniform or Gaussian embedding with the inverse cumulative distribution function or a learned bijection, then check the result against independent quadrature.
  • Beyond the paper, the missing $K_{PP}$ values for Matérn–Gaussian and Wendland–Gaussian pairs are a natural completion target, since the listed $K_P$ formulas appear integrable in terms of error functions and elementary functions; adding them would complete the variance term needed for Bayesian quadrature.
  • Beyond the paper, the constant-embedding phenomenon for periodic and sphere kernels suggests a general criterion: any stationary kernel whose covariance function integrates to zero over the sphere will have constant $K_P$ and $K_{PP}$, which would let researchers generate new dictionary entries without symbolic integration.
  • Beyond the paper, the dictionary's practical impact depends on whether practitioners treat it as a living collection; a community-contributed extension covering conditional distributions and kernel products would directly serve the Bayesian quadrature variants mentioned in the conclusion.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper collects closed-form expressions for kernel mean embeddings K_P and their integrals K_PP for a range of kernel/distribution pairs (Gaussian, Matérn, Wendland, fractional Brownian motion, power-series, and sphere kernels), reviews generic techniques for constructing new embeddings (product and mixture rules, change of measure/variable, matrix-valued kernels, Stein kernels), and releases an MIT-licensed Python library implementing the formulas. The central claim is that the typeset dictionary and the accompanying code are a reliable, comprehensive reference that users can copy from directly.

Significance. If the dictionary is correct, it fills a practical gap: kernel mean embeddings are needed in Bayesian quadrature, kernel quadrature, MMD-based inference, and related areas, and the formulas are currently scattered across several literatures. The paper's strengths are its scope, the clear transformation rules in Section 4, and the companion open-source library with unit tests. The library and tests are reproducible artifacts and are a genuine asset. However, the value of the paper depends entirely on the accuracy of formulas that are explicitly intended to be copied, and the newly computed Wendland entry contains a wrong formula and wrong branch conditions. For this reason the current version cannot be accepted as a reliable dictionary.

major comments (2)
  1. [3.3, Eq. (34)] The Wendland/uniform embedding K_PP is incorrect as printed. For K0(x,y)=(1-|x-y|/ell)_+ on [a,b] with r=b-a, the integral depends only on r/ell. The second branch, printed as (3r-ell)/(3 r^2 ell), is not scale invariant and has the wrong dimension; direct integration of Eq. (33) in the ell<r regime gives ell(3r-ell)/(3r^2). The third branch prints the value 1 - r/(3ell) under the condition r>ell, but that value is the correct result for r<ell, so the case logic sends users to the wrong value in the large-support regime. Because this is one of the newly computed entries in a paper whose stated purpose is to provide formulas to be copied, this is a load-bearing correctness failure rather than a cosmetic typo and must be corrected.
  2. [3, paragraph after Table 1] The statement that the library and its tests 'can be thought of as numerical proofs of the identities here' is not supported. Unit tests against numerical integration are useful verification, but they are not proofs and cannot validate the formulas as typeset unless the tests are shown to cover every branch and every scaling regime. The error in Eq. (34) is concrete evidence that the current test coverage is insufficient. The authors should replace this language with a precise account of what is tested, mark which entries are newly computed, and state the verification method used for each such entry.
minor comments (4)
  1. [3.1, Eq. (12)] In Eq. (12) the argument of the remaining error function, erf(r_i/(ell sqrt(2))), uses a bare ell where ell_i is meant; this should be fixed for consistency with the product notation.
  2. [3.2, Eq. (31)] In Eq. (31), the prefactor of the Gaussian term in the second square bracket appears to have unbalanced parentheses; please check and re-set the formula.
  3. [Table 1] The table uses '?' for several K_PP entries; a short note explaining that these are not currently known (or not included) would prevent readers from interpreting them as open computational challenges.
  4. [3.3, Eq. (34)] The separate case r=2ell in Eq. (34) is unnecessary once the correct scale-invariant formula is used, since the value 5/12 is obtained by continuity; simplifying the branch structure would make the formula easier to verify.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: dictionary entries are cited independent results or direct integrals of kernel and measure definitions, with numerical tests used as external checks.

full rationale

The paper is a reference collection rather than a fitted predictive model, so the main circularity patterns do not apply. The tabulated embeddings are either quoted from the cited literature (e.g., Matérn embeddings from Ming and Guillas 2021 and Ginsbourger et al. 2016; spherical embeddings from Gräf 2013 and Ehler et al. 2019) or obtained by direct integration of the stated kernel and density. Section 3.3 states "We have computed these embeddings with Mathematica" for the Wendland entries, and Section 3 states "Our Python library and its tests can be thought of as numerical proofs of the identities here": both are checks against the defining integrals, not refits of a target quantity. The Stein-kernel identities in Section 4.2 are explicitly "by construction" from the definition of \tilde K, and the paper labels them as such rather than presenting them as independently derived predictions. The few self-references (Briol et al. 2019b; ProbNum contributors) are pointers to prior formulations or software, and no load-bearing argument reduces to an unverified self-citation. An apparent typographical defect in the piecewise Wendland/uniform formula in Eq. (34) would be a correctness problem in a copied formula, not a circular dependency: the claimed embedding is still defined from Eq. (33), and numerical integration provides an independent benchmark. Hence the derivation chain is self-contained and no circular step is exhibited.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities appear. The central claim rests on standard mathematical background plus the paper-specific assumption that numerical tests establish correctness.

assumptions (3)
  • domain assumption The kernel K and measure P satisfy the integrability condition ∫ K(x,x) dP(x) < ∞ on the stated domains.
    Stated at the start of Section 1; all formulas in the dictionary rely on the existence of the integrals in Eqs. (1)-(2).
  • standard math Standard Gaussian integral identities (completion of the square, error function properties, Isserlis' theorem) are valid for the parameter ranges used.
    Used implicitly throughout Section 3, e.g., in Eqs. (13)-(14) and (43)-(44).
  • ad hoc to paper The 'numerical proofs' provided by the Python unit tests are sufficient evidence of correctness for all claimed identities.
    The paper states tests can be thought of as numerical proofs, but this is a methodological assumption that numerical agreement on finitely many points guarantees correctness.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Dictionary of Closed-Form Kernel Mean Embeddings." pith.science (2026). https://pith.science/paper/J74ZNQOW

@misc{pith2026250418830,
  author       = {Pith},
  title        = {Pith review of: A Dictionary of Closed-Form Kernel Mean Embeddings},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/J74ZNQOW}},
  note         = {Machine review of arXiv:2504.18830}
}
read the original abstract

Kernel mean embeddings -- integrals of a kernel with respect to a probability distribution -- are essential in Bayesian quadrature, but also widely used in other computational tools for numerical integration or for statistical inference based on the maximum mean discrepancy. These methods often require, or are enhanced by, the availability of a closed-form expression for the kernel mean embedding. However, deriving such expressions can be challenging, limiting the applicability of kernel-based techniques when practitioners do not have access to a closed-form embedding. This paper addresses this limitation by providing a comprehensive dictionary of known kernel mean embeddings, along with practical tools for deriving new embeddings from known ones. We also provide a Python library that includes minimal implementations of the embeddings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Kernel Quantile Embeddings and Associated Probability Metrics

    stat.ML 2025-05 conditional novelty 6.0 of 10

    Kernel quantile embeddings produce a new family of distribution distances that are probability metrics under separating kernels and can be estimated in O(n log^2 n) time.

  2. Gaussian Processes and Reproducing Kernels: Connections and Equivalences

    stat.ML 2025-06 accept novelty 4.0 of 10

    Gaussian process methods and reproducing kernel Hilbert space methods are two faces of one isometry, shown here across interpolation, regression, numerical integration, and dependence testing.

Reference graph

Works this paper leans on

88 extracted references · 71 canonical work pages · cited by 2 Pith papers

  1. [1]

    Alquier and M

    P. Alquier and M. Gerber. Universal robust regression via maximum mean discrepancy . Biometrika, 111 0 (1): 0 71--92, 2023

  2. [2]

    Alquier and M

    P. Alquier and M. Gerber. regMMD: Robust regression and estimation through maximum mean discrepancy minimization, 2024. URL https://cran.r-project.org/package=regMMD. R package version 0.0.1

  3. [3]

    M. A. \' A lvarez, L. Rosasco, and N. D. Lawrence. Kernels for vector-valued functions: A review . Foundations and Trends in Machine Learning , 4 0 (3): 0 195--266, 2012

  4. [4]

    Anastasiou, A

    A. Anastasiou, A. Barp, F.-X. Briol, B. Ebner, R. E. Gaunt, F. Ghaderinezhad, J. Gorham, A. Gretton, C. Ley, Q. Liu, L. Mackey, C. J. Oates, G. Reinert, and Y. Swan. Stein's method meets computational statistics: A review of some recent developments . Statistical Science, 38 0 (1): 0 120--139, 2023

  5. [5]

    Arbel, A

    M. Arbel, A. Korba, A. Salim, and A. Gretton. Maximum mean discrepancy gradient flow . In Advances in Neural Information Processing Systems, volume 32, 2019

  6. [6]

    F. Bach, S. Lacoste-Julien, and G. Obozinski. On the equivalence between herding and conditional gradient algorithms . In Proceedings of the International Conference on Machine Learning, pages 1355--1362, 2012

  7. [7]

    Barp, F.-X

    A. Barp, F.-X. Briol, A. B. Duncan, M. Girolami, and L. Mackey. Minimum Stein discrepancy estimators . In Advances in Neural Information Processing Systems, volume 32, pages 12964--12976, 2019

  8. [8]

    A. Barp, C. J. Oates, E. Porcu, and M. Girolami. A Riemannian–Stein kernel method . Bernoulli, 28 0 (4): 0 2181--2208, 2022

Show all 88 references
  1. [9]

    Belhadji, R

    A. Belhadji, R. Bardenet, and P. Chainais. Kernel quadrature with DPPs . In Advances in Neural Information Processing Systems, volume 32, pages 12927--12937, 2019

  2. [10]

    Belhadji, D

    A. Belhadji, D. Sharp, and Y. Marzouk. Weighted quantization using MMD: From mean field to mean shift via gradient flows . arXiv:2502.10600, 2025

  3. [11]

    Berlinet and C

    A. Berlinet and C. Thomas-Agnan. Reproducing Kernel Hilbert Spaces in Probability and Statistics . Springer Science+Business Media, New York, 2004

  4. [12]

    Bharti, M

    A. Bharti, M. Naslidnyk, O. Key, S. Kaski, and F.-X. Briol. Optimally-weighted estimators of the maximum mean discrepancy for likelihood-free inference . In International Conference on Machine Learning, pages 2289--2312, 2023

  5. [13]

    Briol, A

    F.-X. Briol, A. Barp, A. B. Duncan, and M. Girolami. Statistical inference for generative models with maximum mean discrepancy . arXiv:1906.05944v1, 2019 a

  6. [14]

    Briol, C

    F.-X. Briol, C. J. Oates, M. Girolami, M. A. Osborne, and D. Sejdinovic. Probabilistic integration: A role in statistical computation? (with discussion and rejoinder). Statistical Science, 34 0 (1): 0 1--22, 2019 b

  7. [15]

    Chamakh and Z

    L. Chamakh and Z. Szabó. Keep it tighter - A story on analytical mean embeddings . arXiv:2110.09516v3, 2024

  8. [16]

    Chatalic, N

    A. Chatalic, N. Schreuder, E. De Vito , and L. Rosasco. Efficient numerical integration in reproducing kernel Hilbert spaces via leverage scores sampling . arXiv:2311.13548v1, 2023

  9. [17]

    W. Y. Chen, L. Mackey, J. Gorham, F.-X. Briol, and C. J. Oates. Stein points . In Proceedings of the International Conference on Machine Learning, pages 843--852, 2018

  10. [18]

    W. Y. Chen, A. Barp, F.-X. Briol, J. Gorham, M. Girolami, L. Mackey, and C. J. Oates. Stein point Markov chain Monte Carlo . In Proceedings of the International Conference on Machine Learning, pages 1011--1021, 2019

  11. [19]

    Y. Chen, M. Welling, and A. Smola. Super-samples from kernel herding . In Proceedings of the Conference on Uncertainty in Artificial Intelligence, pages 109--116, 2010

  12. [20]

    Z. Chen, A. Mustafi, P. Glaser, A. Korba, A. Gretton, and B. K. Sriperumbudur. (De)-regularized maximum mean discrepancy gradient flow . arXiv:2409.14980v1, 2024 a

  13. [21]

    Z. Chen, M. Naslidnyk, A. Gretton, and F.-X. Briol. Conditional Bayesian quadrature . Uncertainty in Artificial Intelligence, pages 648--684, 2024 b

  14. [22]

    Ch \' e rief-Abdellatif and P

    B.-E. Ch \' e rief-Abdellatif and P. Alquier. MMD-Bayes: Robust Bayesian estimation via maximum mean discrepancy . In Proceedings of The 2nd Symposium on Advances in Approximate Bayesian Inference (AABI), pages 1--21, 2020

  15. [23]

    Ch \' e rief-Abdellatif and P

    B.-E. Ch \' e rief-Abdellatif and P. Alquier. Finite sample properties of parametric MMD estimation: robustness to misspecification and dependence . Bernoulli, 28 0 (1): 0 181--213, 2022

  16. [24]

    Chwialkowski, H

    K. Chwialkowski, H. Strathmann, and A. Gretton. A kernel test of goodness of fit . In Proceedings of the International Conference on Machine Learning, pages 2606--2615, 2016

  17. [25]

    M. P. Deisenroth, M. F. Huber, and U. D. Hanebeck. Analytic moment-based G aussian process filtering. In Proceedings of the International Conference on Machine Learning, pages 225--232, 2009

  18. [26]

    Dellaporta, J

    C. Dellaporta, J. Knoblauch, T. Damoulas, and F.-X. Briol. Robust Bayesian inference for simulator-based models via the MMD posterior bootstrap . In Proceedings of the International Conference on Artificial Intelligence and Statistics, pages 943--970, 2022

  19. [27]

    Dick and F

    J. Dick and F. Pillichshammer. Digital Nets and Sequences: Discrepancy Theory and Quasi-Monte Carlo Integration. Cambridge University Press, 2010

  20. [28]

    J. Dick, F. Y. Kuo, and I. H. Sloan. High-dimensional integration: the quasi- M onte C arlo way. Acta Numerica, 22: 0 133--288, 2013

  21. [29]

    Durrande, D

    N. Durrande, D. Ginsbourger, O. Roustant, and L. Carraro. ANOVA kernels and RKHS of zero mean functions for model-based sensitivity analysis. Journal of Multivariate Analysis, 115: 0 57--67, 2013

  22. [30]

    Dwivedi and L

    R. Dwivedi and L. Mackey. Kernel thinning . Journal of Machine Learning Research, 25, 2024

  23. [31]

    G. K. Dziugaite, D. M. Roy, and Z. Ghahramani. Training generative neural networks via maximum mean discrepancy optimization . In Uncertainty in Artificial Intelligence, 2015

  24. [32]

    Ehler, M

    M. Ehler, M. Graef, and C. J. Oates. Optimal Monte Carlo integration on closed manifolds . Statistics and Computing, 29: 0 1204--1214, 2019

  25. [33]

    E. N. Epperly and E. Moreno. Kernel quadrature with randomly pivoted cholesky . In Advances in Neural Information Processing Systems, volume 36, pages 65850--65868, 2023

  26. [34]

    Flaxman, D

    S. Flaxman, D. Sejdinovic, J. P. Cunningham, and S. Filippi. Bayesian learning of kernel embeddings . In Uncertainty in Artificial Intelligence, pages 182--191, 2016

  27. [35]

    Fukumizu, L

    K. Fukumizu, L. Song, and A. Gretton. Kernel Bayes' rule: Bayesian inference with positive definite kernels . Journal of Machine Learning Research, 14: 0 3753--3783, 2013

  28. [36]

    Fuselier, T

    E. Fuselier, T. Hangelbroek, F. J. Narcowich, J. D. Ward, and G. B. Wright. Kernel based quadratures on spheres and other homogeneous spaces . Numerische Mathematik, 127 0 (1): 0 57--92, 2014

  29. [37]

    Gessner, J

    A. Gessner, J. Gonzalez, and M. Mahsereci. Active multi-information source Bayesian quadrature . In Uncertainty in Artificial Intelligence, pages 712--721, 2020

  30. [38]

    Ginsbourger, O

    D. Ginsbourger, O. Roustant, D. Schuhmacher, N. Durrande, and N. Lenz. On ANOVA decompositions of kernels and G aussian random field paths. In Monte Carlo and Quasi-Monte Carlo Methods, volume 163 of Springer Proceedings in Mathematics & Statistics, pages 315--330, 2016

  31. [39]

    Gr \"a f

    M. Gr \"a f. Efficient Algorithms for the Computation of Optimal Quadrature Points on R iemannian Manifolds . PhD thesis, Chemnitz University of Technology, 2013

  32. [40]

    Gretton, K

    A. Gretton, K. M. Borgwardt, M. J. Rasch, B. Schölkopf, and A. Smola. A kernel two-sample test. Journal of Machine Learning Research, 13: 0 723--773, 2012

  33. [41]

    Gunter, M

    T. Gunter, M. A. Osborne, R. Garnett, P. Hennig, and S. J. Roberts. Sampling for inference in probabilistic models with fast B ayesian quadrature. In Advances in Neural Information Processing Systems, volume 24, pages 2789--2797, 2014

  34. [42]

    Hennig, M

    P. Hennig, M. A. Osborne, and H. P. Kersting. Probabilistic Numerics: Computation as Machine Learning. Cambridge University Press, 2022

  35. [43]

    Hertrich, C

    J. Hertrich, C. Wald, F. Altekr \" u ger, and P. Hagemann. Generative sliced MMD flows with Riesz kernels . In International Conference on Learning Representations, 2024

  36. [44]

    Huang, A

    D. Huang, A. Bharti, A. Souza, L. Acerbi, and S. Kaski. Learning robust statistics for simulation-based inference under model misspecification . In Advances in Neural Information Processing Systems, pages 7289--7310, 2023

  37. [45]

    Kanagawa, B

    M. Kanagawa, B. K. Sriperumbudur, and K. Fukumizu. Convergence analysis of deterministic kernel-based quadrature rules in misspecified settings . Foundations of Computational Mathematics, 20: 0 155--194, 2020

  38. [46]

    Karvonen, C

    T. Karvonen, C. J. Oates, and S. Särkkä. A B ayes-- S ard cubature method. In Advances in Neural Information Processing Systems, volume 31, pages 5882--5893, 2018

  39. [47]

    a rkk \

    T. Karvonen, S. S \" a rkk \" a , and C. J. Oates. Symmetry exploits for Bayesian cubature methods . Statistics and Computing, 29: 0 1231--1248, 2019

  40. [48]

    Kellner and A

    J. Kellner and A. Celisse. A one-sample test for normality with kernel methods . Bernoulli, 25 0 (3): 0 1816--1837, 2019

  41. [49]

    Korba, P

    A. Korba, P. C. Aubin-Frankowski, S. Majewski, and P. Ablin. Kernel Stein discrepancy descent . In Proceedings of the International Conference on Machine Learning, pages 5719--5730, 2021

  42. [50]

    Lacoste-Julien, F

    S. Lacoste-Julien, F. Lindsten, and F. Bach. Sequential kernel herding: Frank-Wolfe optimization for particle filtering . In Proceedings of the International Conference on Artificial Intelligence and Statistics, pages 544--552, 2015

  43. [51]

    W. Li, Z. Zeng, A. Vergari, and G. van den Broeck. Tractable computation of expected kernels . In Uncertainty in Artificial Intelligence, pages 1163--1173, 2021

  44. [52]

    Q. Liu, J. D. Lee, and M. I. Jordan. A kernelized Stein discrepancy for goodness-of-fit tests and model evaluation . In International Conference on Machine Learning, pages 276--284, 2016

  45. [53]

    J. R. Lloyd and Z. Ghahramani. Statistical model criticism using kernel two sample tests . Advances in Neural Information Processing Systems, 28: 0 829--837, 2015

  46. [54]

    Matsubara, J

    T. Matsubara, J. Knoblauch, F.-X. Briol, and C. J. Oates. Robust generalised Bayesian inference for intractable likelihoods . Journal of the Royal Statistical Society. Series B (Statistical Methodology), 84 0 (3): 0 997--1022, 2022

  47. [55]

    Ming and S

    D. Ming and S. Guillas. Linked G aussian process emulation for systems of computer models using M atérn kernels and adaptive design. SIAM/ASA Journal on Uncertainty Quantification, 9 0 (4): 0 1615--1642, 2021

  48. [56]

    Muandet, K

    K. Muandet, K. Fukumizu, B. K. Sriperumbudur, and B. Sch \" o lkopf. Kernel mean embedding of distributions: A review and beyonds . Foundations and Trends in Machine Learning , 10 0 (1-2): 0 1--141, 2016 a

  49. [57]

    Muandet, B

    K. Muandet, B. Sriperumbudur, K. Fukumizu, A. Gretton, and B. Sch \" o lkopf. Kernel mean shrinkage estimators . Journal of Machine Learning Research, 17 0 (48): 0 1--41, 2016 b

  50. [58]

    Muandet, M

    K. Muandet, M. Kanagawa, S. Saengkyongam, and S. Marukatat. Counterfactual mean embeddings . Journal of Machine Learning Research, 22 0 (162): 0 7322--7392, 2021

  51. [59]

    Niederreiter

    H. Niederreiter. Random Number Generation and Quasi-Monte Carlo Methods . Society for Industrial and Applied Mathematics, 1992

  52. [60]

    Nishiyama and K

    Y. Nishiyama and K. Fukumizu. Characteristic kernels and infinitely divisible distributions . Journal of Machine Learning Research, 17 0 (180): 0 1--28, 2016

  53. [61]

    Nishiyama, M

    Y. Nishiyama, M. Kanagawa, A. Gretton, and K. Fukumizu. Model-based kernel sum rule: kernel Bayesian inference with probabilistic models . Machine Learning, 109 0 (5): 0 939--972, 2020. ISSN 15730565

  54. [62]

    Z. Niu, J. Meier, and F.-X. Briol. Discrepancy-based inference for intractable generative models using quasi-Monte Carlo . Electronic Journal of Statistics, 17 0 (1): 0 1411--1456, 2023

  55. [63]

    C. J. Oates, M. Girolami, and N. Chopin. Control functionals for Monte Carlo integration . Journal of the Royal Statistical Society. Series B (Statistical Methodology), 79 0 (3): 0 695--718, 2017

  56. [64]

    C. J. Oates, J. Cockayne, F.-X. Briol, and M. Girolami. Convergence rates for a class of estimators based on Stein's identity . Bernoulli, 25 0 (2): 0 1141--1159, 2019

  57. [65]

    A. O'Hagan. Bayes-- H ermite quadrature. Journal of Statistical Planning and Inference, 29 0 (3): 0 245--260, 1991

  58. [66]

    Pacchiardi, S

    L. Pacchiardi, S. Khoo, and R. Dutta. Generalized Bayesian likelihood-free inference . Electronic Journal of Statistics, 18: 0 3628--3686, 2024

  59. [67]

    Paleyes, M

    A. Paleyes, M. Mahsereci, and N. D. Lawrence. Emukit: A P ython toolkit for decision making under uncertainty. Proceedings of the Python in Science Conference, 2023

  60. [68]

    Prüher and O

    J. Prüher and O. Straka. Gaussian process quadrature moment transform. IEEE Transactions on Automatic Control, 63 0 (9): 0 2844--2854, 2018

  61. [69]

    Rathinavel and F

    J. Rathinavel and F. J. Hickernell. Fast automatic B ayesian cubature using lattice sampling. Statistics and Computing, 29 0 (6): 0 1215--1229, 2019

  62. [70]

    R. M. Rustamov. Closed-form expressions for maximum mean discrepancy with applications to Wasserstein auto-encoders . Stat, 10 0 (1): 0 1--12, 2021

  63. [71]

    Sejdinovic

    D. Sejdinovic. An overview of causal inference using kernel embeddings . arXiv:2410.22754, 2024

  64. [72]

    S. Si, C. J. Oates, A. B. Duncan, L. Carin, and F.-X. Briol. Scalable control variates for Monte Carlo methods via stochastic optimization . In Monte Carlo and Quasi-Monte Carlo Methods. MCQMC 2020, pages 205--221. Springer, 2022

  65. [73]

    Singh, M

    R. Singh, M. Sahani, and A. Gretton. Kernel instrumental variable regression . In Advances in Neural Information Processing Systems, volume 32, pages 4593--4605, 2019

  66. [74]

    Singh, L

    R. Singh, L. Xu, and A. Gretton. Kernel methods for causal functions: dose, heterogeneous and incremental response curves . Biometrika, 111 0 (2): 0 497--516, 2024

  67. [75]

    Sommariva and M

    A. Sommariva and M. Vianello. Numerical cubature on scattered data by radial basis functions . Computing, 76 0 (3-4): 0 295--310, 2006

  68. [76]

    L. Song, X. Zhang, A. Smola, A. Gretton, and B. Sch \" o lkopf. Tailoring density estimation via reproducing kernel moment matching . In Proceedings of the International Conference on Machine Learning, pages 992--999, 2008

  69. [77]

    L. F. South, T. Karvonen, C. Nemeth, and C. J. Oates. Semi-exact control functionals from S ard's method. Biometrika, 109 0 (2): 0 351--367, 2022

  70. [78]

    B. K. Sriperumbudur. On the optimal estimation of probability measures in weak and strong topologies . Bernoulli, 22 0 (3): 0 1839--1893, 2016

  71. [79]

    B. K. Sriperumbudur, A. Gretton, K. Fukumizu, B. Sch \" o lkopf, and G. R. G. Lanckriet. Hilbert space embeddings and metrics on probability measures . Journal of Machine Learning Research, 11, 2010

  72. [80]

    Z. Sun, A. Barp, and F.-X. Briol. Vector-valued control variates . In Proceedings of the International Conference on Machine Learning, pages 32819--32846, 2023

  73. [81]

    Tolstikhin, B

    I. Tolstikhin, B. Sriperumbudur, and K. Muandet. Minimax estimation of kernel mean embeddings . Journal of Machine Learning Research, 18 0 (1): 0 3002--3048, 2017

  74. [82]

    Tolstikhin, O

    I. Tolstikhin, O. Bousquet, S. Gelly, and B. Schoelkopf. Wasserstein auto-encoders . In International Conference on Learning Representations, 2018

  75. [83]

    G. Wahba. Spline Models for Observational Data. Society for Industrial and Applied Mathematics, 1990

  76. [84]

    Wendland

    H. Wendland. Piecewise polynomial, positive definite and compactly supported radial functions of minimal degree. Advances in Computational Mathematics, 4: 0 389--396, 1995

  77. [85]

    a mer, M. Pf \

    J. Wenger, N. Kr \" a mer, M. Pf \" o rtner, J. Schmidt, N. Bosch, N. Effenberger, J. Zenn, A. Gessner, T. Karvonen, F.-X. Briol, M. Mahsereci, and P. Hennig. ProbNum: Probabilistic numerics in Python . arXiv:2112.02100v1, 2021

  78. [86]

    Wolfer and P

    G. Wolfer and P. Alquier. Variance-aware estimation of kernel mean embedding . arXiv:2210.06672v2, 2024

  79. [87]

    Xi, F.-X

    X. Xi, F.-X. Briol, and M. Girolami. Bayesian quadrature for multiple related integrals . In Proceedings of the International Conference on Machine Learning, pages 8533--8564, 2018

  80. [88]

    L. Xu, A. Korba, and D. Slep c ev. Accurate quantization of measures via interacting particle-based optimization . In Proceedings of the International Conference of Machine Learning, pages 24576--24595, 2022

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.