REVIEW 4 minor 1 cited by
Minimax Estimation of Kernel Stein Discrepancy: Trace versus Hilbert-Schmidt Scales
T0 review · 0 major / 4 minor · reviewed 2026-07-12 · grok-4.5
Pith's one-line read The sharp minimax scale for estimating kernel Stein discrepancy is set by the Hilbert–Schmidt norm of the Stein covariance, and the debiased U-statistic attains it while the plug-in V-statistic stays at the larger trace scale.
desk verdict Clean fixed-target minimax theory that pins KSD estimation to the HS scale of C⋆ and shows the usual V-statistic is strictly worse by reff^{1/4}. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Stein covariance operator C⋆ = E_{P0}[ξ_{P0}(X) ⊗ ξ_{P0}(X)], whose Hilbert–Schmidt norm (equivalently tr(C⋆²)^{1/2}) sets the minimax constant, together with the degeneracy of the associated U-statistic at the target forced by Stein’s identity.
What would settle it
At a high-dimensional standard Gaussian target with fixed-bandwidth Gaussian kernel, compute Monte-Carlo risks of both the V-statistic and the square-root U-statistic for growing d at fixed n; if the observed gap fails to track (1+8γ)^{d/8} while the U-statistic risk tracks the explicit Hilbert–Schmidt formula, the spectral claim is false.
Extended reading notes
Core claim
When the target P0 is fixed, the minimax risk of estimating KSD(P0,P) over a natural off-diagonal moment class is of order the square root of the Hilbert–Schmidt norm of the Stein covariance C⋆ divided by n. The positive-part square-root U-statistic attains this scale; the plug-in V-statistic cannot, remaining pinned to the larger trace scale √(tr(C⋆)/n) and therefore suboptimal by the fourth root of the effective rank of C⋆.
Load-bearing premise
The target score function is treated as known exactly; any estimation error in the score is outside the present bounds and could erase the claimed Hilbert–Schmidt advantage.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies minimax estimation of the scalar Kernel Stein Discrepancy KSD(P0,P) for a fixed target P0 known only through its score, from n i.i.d. samples of P. It identifies the sharp spectral constant as the Hilbert–Schmidt norm of the Stein covariance operator C⋆, so that the minimax scale is √(∥C⋆∥HS/n). This scale is attained by the positive-part square-root U-statistic that removes diagonal terms; the usual plug-in V-statistic remains at the larger trace scale √(tr(C⋆)/n) and is therefore suboptimal by reff(C⋆)^{1/4}. Matching upper bounds (Theorems 1–2), a fuzzy-hypothesis lower bound over the target-normalized class P(A) (Theorem 5), and an explicit Gaussian calculation showing an exponential-in-dimension gap for fixed bandwidth (Section 3.4) are provided.
Significance. The contribution is a clean refinement of existing rate-optimal results for KSD estimation: once the target is fixed, the relevant constant is the HS norm of C⋆ rather than its trace, and the debiased U-statistic is minimax optimal while the V-statistic is not. The argument is standard and transparent (reverse triangle inequality, U-statistic variance with degeneracy at the target, Tsybakov two fuzzy hypotheses along Stein eigen-coordinates), the class P(A) is well-motivated by the off-diagonal representation of KSD², and the Gaussian spectral formulas make the V/U gap concrete and dimension-dependent. The known-score premise is already flagged by the authors as outside the present scope. The result is of clear interest for KSD-based diagnostics and goodness-of-fit, especially in moderate-to-high dimension.
minor comments (4)
- In the abstract and introduction the minimax scale is written √(∥C⋆∥HS/n); later (Theorem 2, Corollary 6) it is equivalently tr(C⋆²)^{1/4}/√n. A single consistent notation for the HS norm would reduce minor ambiguity for readers who do not immediately equate the two.
- Figure 1 caption and the surrounding text in §3.4 refer to Monte Carlo risks of [KSD_V and [KSD_U; the notation for the estimators is slightly inconsistent with the definitions (4)–(5) (hat vs. bracket). Aligning the symbols would help.
- Appendix D: the truncation argument that removes the boundedness assumption on the Stein coordinates g_j is correct but dense; a short remark that the same construction works under a uniform L^{2+δ} moment on the first m coordinates would make the scope of the lower bound clearer without lengthening the proof.
- Typographical: several displayed equations have line-break artifacts (e.g., the definition of reff(C⋆) and the Gaussian tr(C⋆²) formula). These are purely presentational.
Circularity Check
No significant circularity: standard self-contained minimax derivation with independent upper/lower bounds.
full rationale
The paper identifies the minimax scale of KSD estimation as √(∥C⋆∥HS/n) via classical tools: reverse-triangle/Jensen for the V-statistic (Theorem 1), Hoeffding U-statistic variance plus Hölder for the square-root U-statistic (Theorem 2, Appendix A), and a Tsybakov two-fuzzy-hypotheses construction along the eigenbasis of C⋆ for the matching lower bound (Theorem 5, Appendix D). The class P(A) normalizes the off-diagonal moment by tr(C⋆²) so A is a pure regularity parameter, not a fitted constant. Spectral quantities for the Gaussian target are obtained by direct MGF calculation (Appendix C), not by fitting. Citations (Stein, Liu/Chwialkowski, Tolstikhin, Cribeiro-Ramallo, Tsybakov, Hoeffding) supply background and comparison; none is a self-citation that forces the present spectral claim by construction. The known-score premise is an explicit modeling assumption, not a circular definition of the estimand. No step reduces the claimed rate to its inputs by definition or by a fitted parameter renamed as prediction.
Assumptions & free parameters
free parameters (2)
- A (regularity constant in class P(A)) =
≥4
- Gaussian kernel bandwidth γ =
example values γ=1/2 and γ=α/d
assumptions (5)
- domain assumption Stein’s identity: E_{P0} ξ_{P0}(X) = 0 (Assumption 1)
- domain assumption Finite second moments of the Stein kernel under P0 and under every P considered (Assumption 1)
- standard math Hoeffding’s variance formula for second-order U-statistics
- standard math Tsybakov’s two fuzzy hypotheses method (Theorems 2.14–2.15 of Tsybakov 2009)
- domain assumption Target density p0 continuously differentiable and everywhere positive; kernel k in C^{1,1}
Cite this review
Pith. "Pith review of Minimax Estimation of Kernel Stein Discrepancy: Trace versus Hilbert-Schmidt Scales." pith.science (2026). https://pith.science/paper/PYAOGSFA
@misc{pith2026260703367,
author = {Pith},
title = {Pith review of: Minimax Estimation of Kernel Stein Discrepancy: Trace versus Hilbert-Schmidt Scales},
year = {2026},
howpublished = {\url{https://pith.science/paper/PYAOGSFA}},
note = {Machine review of arXiv:2607.03367}
}
abstract
Kernel Stein Discrepancy (KSD) compares a sample to a fixed target distribution known only through its score, and is widely used for goodness-of-fit testing, sample quality assessment, and approximate inference. We study the estimation of $\operatorname{KSD}(P_0,P)$ from $n$ independent observations and identify the sharp spectral constant governing the minimax risk: it is the Hilbert-Schmidt norm of the Stein covariance operator $C_\star$, giving the minimax scale $\sqrt{\|C_\star\|_{\mathrm{HS}}/n}$. This scale is attained by the positive-part square-root U-statistic, whereas the standard plug-in V-statistic remains at the trace scale $\sqrt{\operatorname{tr}(C_\star)/n}$ and is therefore suboptimal by the fourth root of the effective rank of $C_\star$; for a Gaussian target with a fixed-bandwidth Gaussian kernel this factor is exponential in the dimension.
Figures
Forward citations
Cited by 1 Pith paper
-
Minimax Lower Bounds of Kernel Discrepancy Estimation: MMD, HSIC, KSD
Minimax lower bounds for MMD, HSIC and KSD estimation are n^{-1/2} on general topological spaces under mild kernel assumptions, matching existing estimators.
Reference graph
Works this paper leans on
-
[1]
A bound for the error in the normal approximation to the distribution of a sum of dependent random variables
Charles Stein. A bound for the error in the normal approximation to the distribution of a sum of dependent random variables. InProceedings of the Sixth Berkeley Symposium on Mathematical Statistics and Probability, Volume 2: Probability Theory, pages 583–602. University of California Press, 1972
1972
-
[2]
Institute of Mathematical Statistics, 1986
Charles Stein.Approximate Computation of Expectations. Institute of Mathematical Statistics, 1986
1986
-
[3]
Louis H. Y . Chen, Larry Goldstein, and Qi-Man Shao.Normal Approximation by Stein’s Method. Springer, 2011
2011
-
[4]
Measuring sample quality with Stein’s method
Jackson Gorham and Lester Mackey. Measuring sample quality with Stein’s method. InAdvances in Neural Information Processing Systems (NeurIPS), volume 28, 2015
2015
-
[5]
Lee, and Michael I
Qiang Liu, Jason D. Lee, and Michael I. Jordan. A kernelized Stein discrepancy for goodness-of-fit tests. In International Conference on Machine Learning (ICML), pages 276–284, 2016
2016
-
[6]
A kernel test of goodness of fit
Kacper Chwialkowski, Heiko Strathmann, and Arthur Gretton. A kernel test of goodness of fit. InInternational Conference on Machine Learning (ICML), pages 2606–2615, 2016
2016
-
[7]
Measuring sample quality with kernels
Jackson Gorham and Lester Mackey. Measuring sample quality with kernels. InInternational Conference on Machine Learning (ICML), pages 1292–1301, 2017
2017
-
[8]
Probabilistic inference and learning with stein’s method.arXiv preprint arXiv:2603.07467, 2026
Qiang Liu, Lester Mackey, and Chris Oates. Probabilistic inference and learning with stein’s method.arXiv preprint arXiv:2603.07467, 2026
arXiv 2026
Show all 29 references
-
[9]
Stein variational gradient descent: A general purpose Bayesian inference algorithm
Qiang Liu and Dilin Wang. Stein variational gradient descent: A general purpose Bayesian inference algorithm. InAdvances in Neural Information Processing Systems (NeurIPS), 2016
2016
-
[10]
Improved finite-particle convergence rates for stein variational gradient descent
Krishna Balasubramanian, Sayan Banerjee, and Promit Ghosal. Improved finite-particle convergence rates for stein variational gradient descent. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[11]
Huggins and Lester Mackey
Jonathan H. Huggins and Lester Mackey. Random feature Stein discrepancies. InAdvances in Neural Information Processing Systems (NeurIPS), volume 31, 2018. 9 Minimax Estimation of Kernel Stein Discrepancy
2018
-
[12]
Nyström kernel stein discrepancy.arXiv preprint arXiv:2406.08401, 2024
Florian Kalinke, Zoltán Szabó, and Bharath K Sriperumbudur. Nyström kernel stein discrepancy.arXiv preprint arXiv:2406.08401, 2024
2024 arXiv
-
[13]
A linear-time kernel goodness-of-fit test
Wittawat Jitkrittum, Wenkai Xu, Zoltán Szabó, Kenji Fukumizu, and Arthur Gretton. A linear-time kernel goodness-of-fit test. InAdvances in Neural Information Processing Systems (NeurIPS), volume 30, 2017
2017
-
[14]
A kernel Stein test for comparing latent variable models.Journal of the Royal Statistical Society Series B: Statistical Methodology, 85(3):986–1011, 2023
Heishiro Kanagawa, Wittawat Jitkrittum, Lester Mackey, Kenji Fukumizu, and Arthur Gretton. A kernel Stein test for comparing latent variable models.Journal of the Royal Statistical Society Series B: Statistical Methodology, 85(3):986–1011, 2023
2023
-
[15]
Wilson Ye Chen, Lester Mackey, Jackson Gorham, François-Xavier Briol, and Chris J. Oates. Stein points. In International Conference on Machine Learning (ICML), pages 844–853, 2018
2018
-
[16]
Duncan, Mark Girolami, and Lester Mackey
Alessandro Barp, François-Xavier Briol, Andrew B. Duncan, Mark Girolami, and Lester Mackey. Minimum Stein discrepancy estimators. InAdvances in Neural Information Processing Systems (NeurIPS), volume 32, 2019
2019
-
[17]
Minimax optimal goodness-of-fit testing with kernel stein discrepancy.Bernoulli, 32(1):299–324, 2026
Omar Hagrass, Bharath Sriperumbudur, and Krishnakumar Balasubramanian. Minimax optimal goodness-of-fit testing with kernel stein discrepancy.Bernoulli, 32(1):299–324, 2026
2026
-
[18]
Nyström kernel stein discrepancy tests.arXiv preprint arXiv:2605.25173, 2026
Florian Kalinke, Zoltán Szabó, and Bharath K Sriperumbudur. Nyström kernel stein discrepancy tests.arXiv preprint arXiv:2605.25173, 2026
2026 arXiv
-
[19]
Borgwardt, Malte J
Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, and Alexander Smola. A kernel two-sample test.Journal of Machine Learning Research, 13:723–773, 2012
2012
-
[20]
Kernel mean embedding of distributions: A review and beyond.Foundations and Trends in Machine Learning, 10(1–2):1–141, 2017
Krikamol Muandet, Kenji Fukumizu, Bharath Sriperumbudur, and Bernhard Schölkopf. Kernel mean embedding of distributions: A review and beyond.Foundations and Trends in Machine Learning, 10(1–2):1–141, 2017
2017
-
[21]
Minimax estimation of kernel mean embed- dings.Journal of Machine Learning Research, 18(86):1–47, 2017
Ilya Tolstikhin, Bharath K Sriperumbudur, and Krikamol Muandet. Minimax estimation of kernel mean embed- dings.Journal of Machine Learning Research, 18(86):1–47, 2017
2017
-
[22]
Tolstikhin, Bharath K
Ilya O. Tolstikhin, Bharath K. Sriperumbudur, and Bernhard Schölkopf. Minimax estimation of maximum mean discrepancy with radial kernels. InAdvances in Neural Information Processing Systems (NeurIPS), volume 29, pages 1930–1938, 2016
1930
-
[23]
The minimax lower bound of kernel stein discrepancy estimation
Jose Cribeiro-Ramallo, Agnideep Aich, Florian Kalinke, Ashit Baran Aich, and Zoltán Szabó. The minimax lower bound of kernel stein discrepancy estimation. InThe 29th International Conference on Artificial Intelligence and Statistics, 2026
2026
-
[24]
Estimation of integral functionals of a density.The Annals of Statistics, 23(1): 11–29, 1995
Lucien Birgé and Pascal Massart. Estimation of integral functionals of a density.The Annals of Statistics, 23(1): 11–29, 1995
1995
-
[25]
A simple adaptive estimator of the integrated square of a density.Bernoulli, 14 (1):47–61, 2008
Evarist Giné and Richard Nickl. A simple adaptive estimator of the integrated square of a density.Bernoulli, 14 (1):47–61, 2008
2008
-
[26]
A class of statistics with asymptotically normal distribution.The Annals of Mathematical Statistics, 19(3):293–325, 1948
Wassily Hoeffding. A class of statistics with asymptotically normal distribution.The Annals of Mathematical Statistics, 19(3):293–325, 1948
1948
-
[27]
Tsybakov.Introduction to Nonparametric Estimation
Alexandre B. Tsybakov.Introduction to Nonparametric Estimation. Springer, 2009. A Proof of Theorem 2 Unbiasedness follows from the KSD identity: EP KP0(X1, X2) =∥µ P ∥2 H = KSD2(P0, P). Let h(x, y) :=K P0(x, y). The variance formula for a second-order U-statistic [26, Eq. 5.13...
2009
-
[28]
(Fixed bandwidth.)For fixed γ >0 , tr(C2 ⋆)≍ γ d2(1+8γ) −d/2, the optimal scale is≍γ d1/2(1+8γ) −d/8/√n, and the gapr 1/4 eff ≍γ (1 + 8γ)d/8 isexponential ind
-
[29]
tr(C2 ⋆)≍ α d, the optimal scale is ≍α d1/4/√n, and the gap r1/4 eff ≍α d1/4 is polynomial ind
(Scaled bandwidth γ=α/d ). tr(C2 ⋆)≍ α d, the optimal scale is ≍α d1/4/√n, and the gap r1/4 eff ≍α d1/4 is polynomial ind. Proof. The bracket in Proposition 10 is c1(γ)d+c 0(γ) with c1(γ) = 4γ 2(4γ+ 1)2.(i)The c1d term dominates, giving tr(C2 ⋆)≍d 2(1+8γ) −d/2 and r1/4 eff = (...
Reviewed July 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.