REVIEW 7 minor 34 references
Schoenberg characterization of continuous non-stationary isotropic positive definite kernels
T0 review · 0 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Every continuous rotation-invariant positive definite kernel on $\mathbb{R}^d$ is a unique Gegenbauer series.
desk verdict Solid Schoenberg-type characterization for non-stationary isotropic kernels on R^d (d≥2); minor presentational flaws, core proof sound. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the series representation with normalized Gegenbauer polynomials $\widehat{C}_n^{(\lambda)}$ evaluated at the inner product of the normalized vectors. The argument carries by the homeomorphism $T(r,v)=rv$ between $(0,\infty)\times S^{d-1}$ and $\mathbb{R}^d\setminus\{0\}$, which turns an isotropic kernel on $\mathbb{R}^d$ into a kernel on the product space of the form $f_S(r,s,\rho)=\kappa(r,s,rs\rho)$. On that product space the cited characterization supplies the series with positive definite coefficient kernels, and continuity plus the condition $\alpha_n(r,0)=\alpha_n(0,r)=0$ for $n\ge 1$ extends the representation to the origin. For $d=\infty$ the Gegenbauer polynomials are replaced by monomials $\rho^n$, giving a power series in the cosine.
What would settle it
Compute the Gegenbauer coefficients in (5) for a concrete continuous $O(d)$-invariant kernel such as $K(x,y)=\exp(-\|x-y\|^2)$. If any coefficient $\alpha_n$ fails to be positive definite on $[0,\infty)^2$, or fails to vanish at the origin, the representation (3) is false. Conversely, a continuous $O(d)$-invariant positive definite kernel whose Gegenbauer projections are not summable on the diagonal would also disprove the theorem.
Extended reading notes
Core claim
Theorem 2.1 states that for $d\in\{2,3,\ldots,\infty\}$, a continuous $K:\mathbb{R}^d\times\mathbb{R}^d\to\mathbb{C}$ is an isotropic positive definite kernel—meaning $K(Ux,Uy)=K(x,y)$ for every orthogonal $U$—if and only if it admits the representation $$K(x,y)=\sum_{n=0}^{\infty}\$alpha_n^{{(d)}}$(\|x\|,\|y\|)\,\widehat{C}$_n^{{(\lambda)}}$\!\left(\left\langle \frac{x}{\|x\|},\frac{y}{\|y\|}\right\rangle\right)$$ with $\lambda=(d-2)/2$, where the $\alpha_n^{(d)}$ are unique continuous positive definite kernels on $[0,\infty)^2$, summable on the diagonal, and zero at the origin for $n\ge 1$. For finite $d$ the coefficients are obtained by Gegenbauer projection of $\kappa(r,s,rs\rho)$ with weight $(1-\rho^2)^{\lambda-1/2}$. The same framework yields Theorem 3.1 and Theorem 3.2: strict positive definiteness holds essentially when the even and odd parts of the coefficient sequence are themselves strictly positive definite on $(0,\infty)$, with the $d=2$ case requiring intersection with every full arithmetic progression in $\mathbb{Z}$.
Load-bearing premise
The whole theorem rests on the external characterization of positive definite kernels on $X\times S^{d-1}$ (Theorem A.1); if that result were false, misstated, or inapplicable, the series representation and every strict-positive-definiteness condition derived from it would collapse.
Editorial extensions
If this is right
- Every continuous $O(d)$-invariant positive definite kernel, including non-stationary ones, has a unique explicit series expansion, so kernel design can be done by choosing the coefficient kernels $\alpha_n$ instead of the full function $K$.
- Stationary isotropic kernels and dot product kernels are recovered as special cases, giving a single framework for both classical classes.
- The infinite-width neural network Gaussian process (NNGP) kernel and the neural tangent kernel are isotropic and therefore fall under the characterization, with the NNGP kernel expressed as $\sum_m \alpha_m(\|x\|,\|x'\|)\langle x/\|x\|, x'/\|x'\|\rangle^m$ in $\ell_2$.
- Strict positive definiteness can be read off from the parity of the active coefficients: infinitely many even and odd $n$ must contribute, with $\alpha_0(0,0)>0$.
- In infinite dimension the representation reduces to a power series in the cosine, which is the natural Schoenberg form on the infinite-dimensional sphere.
Reading between the lines
- A natural next step, hinted at in the paper's remark on universal kernels, is a characterization of universal isotropic kernels on $\mathbb{R}^d$: universality should correspond to a density condition on the coefficient kernels $\alpha_n$ in addition to their positive definiteness.
- The coefficient formula (5) suggests a practical numerical certificate: for any candidate isotropic kernel, compute the Gegenbauer projections and check positivity and zero-at-origin; this could be used in kernel learning or Gaussian process design.
- If the load-bearing external product-space characterization (Theorem A.1) were to fail for some edge case, the zero-at-origin extension and therefore the full representation would break down; the $d=2$ arithmetic-progression condition in Theorem 3.2 shows where such edge effects concentrate.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper characterizes continuous positive definite kernels on R^d that are invariant under the orthogonal group O(d), i.e. isotropic but not necessarily stationary kernels. The main result, Theorem 2.1, states that such kernels are exactly those of the form K(x,y)=sum_{n=0}^\infty \alpha_n^{(d)}(||x||,||y||) \hat C_n^{(\lambda)}(<x/||x||, y/||y||>), with unique continuous positive definite coefficient kernels \alpha_n on [0,\infty)^2 that are summable on the diagonal and vanish at the origin for n\ge 1; the case d=\infty uses monomials in place of Gegenbauer polynomials. Section 3 gives necessary and sufficient conditions for strict positive definiteness, with the case d=2 treated separately because of its arithmetic-progression structure. Section 4 shows that stationary isotropic kernels and dot product kernels are special cases and that infinite-width neural network kernels (NNGP and, in a remark, NTK) are isotropic non-stationary kernels covered by the characterization. The proofs reduce the main theorem to a restated theorem of Guella and Menegatto for kernels on (0,\infty)\times S^{d-1}, with a careful extension to the origin via Lemma A.3 and the zero-at-origin condition.
Significance. If the result holds, it fills a natural gap in the Schoenberg-type literature: it unifies the classical characterizations of stationary isotropic kernels and dot product kernels, and it applies to an important class of neural network kernels that are isotropic but non-stationary. The boundary treatment at the origin, through the zero-at-origin condition (4) and Lemma A.3, is careful and is exactly what makes the extension from (0,\infty)\times S^{d-1} to R^d work. The paper is transparent about its external dependencies: Theorem A.1 is restated from [10], and one Gram-matrix equivalence is cited from the authors' prior work [2, Appendix F]; neither is circular. The main proof is clear, and the strict positive definiteness criteria are concrete enough to be applied. The paper does not ship code or machine-checked proofs, but the analytic arguments are reproducible from the cited sources.
minor comments (7)
- [A.1, proof of Theorem 2.1, d=2 case] The displayed bound |\hat C_n^{(0)}(\rho)(1-\rho^2)^{-1/2}|\le 1 used for dominated convergence is false near \rho=\pm 1. A valid integrable majorant is C(1-\rho^2)^{-1/2}, which follows from the uniform bound |\hat C_n^{(0)}(\rho)|\le 1 and integrability of the weight; with this replacement the continuity-extension argument goes through unchanged.
- [Theorem 3.1 and Eq. (13)] The expressions c^T [\alpha_n(r_i,r_j)] c and \sum c_i c_j K(x_i,x_j) should use the conjugate transpose or \bar c, since the quadratic form for a complex positive definite kernel is Hermitian; as written, c^T A c need not be real and the condition is not equivalent to strict positive definiteness for complex-valued kernels.
- [Example 4.7] The arccosine kernel formula K_NNGP(x,y)=(1/\pi)||x||||y|| J_1(\cos^{-1}(\langle x,y\rangle)) is missing the normalization of the inner product; the argument should be \cos^{-1}(\langle x,y\rangle/(||x||||y||)) unless the formula is intended only for unit-norm inputs.
- [Section A.2, proof of Theorem 3.1] The proof relies on Theorems 3.5, 3.7, and 3.9 of [10] without stating them, and the translation from S^d to S^{d-1} is only mentioned in passing; a short statement of the exact external conditions used would make the claimed equivalences (ii) and (iii) verifiable without consulting the cited paper.
- [Theorem 2.1] Since (iii) alone should imply that K is continuous, it would be helpful to state explicitly that the summability of the \alpha_n on the diagonal together with the Cauchy-Schwarz inequality gives uniform convergence of the series on bounded sets; this is implicit but not spelled out.
- [Abstract and Theorem 2.1] The characterization excludes d=1, but the abstract says 'on R^d' without this restriction; the domain d\in\{2,3,\ldots,\infty\} should be stated in the abstract to avoid overclaiming.
- [Section 2] There is a typo in 'independently discoverd' (should be 'discovered').
Circularity Check
No significant circularity: the central Gegenbauer representation is imported from an external theorem, and the authors' self-citations supply only elementary intermediate equivalences.
full rationale
The main Theorem 2.1 is not circular by construction. Assertion (iii)'s expansion in normalized Gegenbauer polynomials is obtained by applying the external Theorem A.1 (Guella and Menegatto) to f_S(r,s,rho)=kappa(r,s,rs rho) on (0,infinity)xS^{d-1}, then extending the resulting coefficient kernels to the origin. The coefficient kernels alpha_n^{(d)} are uniquely determined by Gegenbauer orthogonality in finite dimension and by power-series coefficient uniqueness in d=infinity, so no fitted parameter is relabelled as a prediction and no assertion is equivalent to its inputs by definition. The authors' prior work is cited twice in the proof chain: [2, Appendix F] for the elementary equivalence (i) iff (ii), namely that an O(d)-invariant continuous kernel depends only on (||x||,||y||,<x,y>), and [2, Prop. F.4] for the standard fact that point configurations with the same Gram matrix are related by an orthogonal transformation. Both are parameter-free supporting facts that do not contain or presuppose the target series representation. The core extension from X x S^{d-1} to R^d, including the zero-at-origin condition (4), is proved within the paper. Section 3 similarly transfers strict-positive-definiteness conditions from the independent Guella-Menegatto framework rather than assuming them. The only caveat is a normal, non-circular self-citation for a preliminary equivalence; hence the low score.
Assumptions & free parameters
assumptions (5)
- domain assumption Guella-Menegatto characterization of positive definite kernels on X times S^{d-1} (Theorem A.1)
- domain assumption Benning-Doring [2, Appendix F] equivalence of O(d)-invariance and dependence on (||x||,||y||,<x,y>)
- standard math Standard orthogonality, normalization, and uniform bounds for Gegenbauer polynomials (Reimer [18])
- standard math Lemma A.2: any two finite configurations with identical Gram matrices are related by an orthogonal transformation
- domain assumption Hanin's recursion for NNGP kernels (Hanin [11], Eq. (1.7) and (1.8))
Cite this review
Pith. "Pith review of Schoenberg characterization of continuous non-stationary isotropic positive definite kernels." pith.science (2026). https://pith.science/paper/ENEPDE7F
@misc{pith2026250622048,
author = {Pith},
title = {Pith review of: Schoenberg characterization of continuous non-stationary isotropic positive definite kernels},
year = {2026},
howpublished = {\url{https://pith.science/paper/ENEPDE7F}},
note = {Machine review of arXiv:2506.22048}
}
abstract
We characterize the continuous isotropic positive definite kernels on $\mathbb{R}^d$, where isotropy refers to invariance under the orthogonal group $O(d)$ but not necessarily stationarity. Furthermore, we characterize strict positive definiteness for such kernels. The class of isotropic kernels is fairly general as it unifies stationary isotropic and dot product kernels, and includes neural network kernels that arise from infinite-width limits of neural networks. As an application, we further characterize the continuous isotropic Gaussian random functions in terms of a series representation.
Reference graph
Works this paper leans on
-
[2]
F. Benning and L. D¨ oring. Random Function Descent. In Advances in Neural Information Processing Systems , volume 37, pages 111248– 111298, Vancouver, Canada, Dec. 2024. Curran Associates, Inc. URL https://proceedings.neurips.cc/paper_files/paper/2024/hash/ c980ad0fb46f0cfb0faabcd42b30a67a-Abstract-Conference.html
work page 2024
-
[10]
J. C. Guella and V. A. Menegatto. Schoenberg’s Theorem for Positive Definite Functions on Products: A Unifying Framework. Journal of Fourier Analysis and Applications , 25(4):1424–1446, Aug. 2019. ISSN 1531-5851. doi: 10.1007/s00041-018-9631-5
-
[1]
S. Ament and C. Gomes. Scalable First-Order Bayesian Optimization via Structured Automatic Differentiation. In Proceedings of the 39th Interna- tional Conference on Machine Learning, Baltimore, Maryland, USA, 2022
work page 2022
-
[3]
C. Berg and E. Porcu. From Schoenberg Coefficients to Schoenberg Func- tions. Constructive Approximation, 45(2):217–241, Apr. 2017. ISSN 1432-
work page 2017
-
[4]
S. Bochner. Monotone Funktionen, Stieltjessche Integrale und harmonische Analyse. Mathematische Annalen, 108(1):378–410, Dec. 1933. ISSN 1432-
work page 1933
-
[5]
Positive definite matrix must be Hermitian
chesslad. Positive definite matrix must be Hermitian. Mathematics Stack Exchange, 2019. URL https://math.stackexchange.com/q/3434192
- [6]
-
[7]
F. de Roos, A. Gessner, and P. Hennig. High-Dimensional Gaussian Process Inference with Derivatives. In Proceedings of the 38th International Con- ference on Machine Learning , pages 2535–2545. PMLR, July 2021. URL https://proceedings.mlr.press/v139/de-roos21a.html
work page 2021
Show all 34 references
-
[8]
Estrade, A
A. Estrade, A. Fari˜ nas, and E. Porcu. Covariance functions on spheres cross time: Beyond spatial isotropy and temporal stationarity. Statistics & Probability Letters, 151:1–7, Aug. 2019. ISSN 0167-7152. doi: 10.1016/j. spl.2019.03.011
2019 doi
-
[9]
Galy-Fajou, D
T. Galy-Fajou, D. Widmann, S. Yalburgi, W. Tebbutt, st–, I. Falk, S. C. Surace, S. Ridderbusch, T. Wright, H. Ge, S. Khan, P. Monti- cone, L. Mones, david-vicente, A. R. Gnadt, J. Giersdorf, J. TagBot, R. Viljoen, S. Sch¨ olly, T. E. Fjelde, and K. ¨Ocal. JuliaGaussianPro- ces...
-
[11]
B. Hanin. Random neural networks in the infinite width limit as Gaussian processes. The Annals of Applied Probability, 33(6A):4798–4819, Dec. 2023. ISSN 1050-5164, 2168-8737. doi: 10.1214/23-AAP1933
2023 doi
-
[12]
Jacot, F
A. Jacot, F. Gabriel, and C. Hongler. Neural Tangent Kernel: Convergence and Generalization in Neural Networks. In Advances in Neural Informa- tion Processing Systems, volume 31, Montr´ eal, Canada, 2018. Curran Asso- ciates, Inc. URL https://proceedings.neurips.cc/paper/2018/...
2018
-
[13]
J. Lee, Y. Bahri, R. Novak, S. S. Schoenholz, J. Pennington, and J. Sohl- Dickstein. Deep Neural Networks as Gaussian Processes. In International Conference on Learning Representations, Vancouver, Canada, 2018. URL https://openreview.net/forum?id=B1EA-M-0Z
2018
-
[14]
C. A. Micchelli, Y. Xu, and H. Zhang. Universal Kernels. Journal of Machine Learning Research , 7(12), 2006. URL http://www.jmlr.org/ papers/volume7/micchelli06a/micchelli06a.pdf
2006
-
[15]
Panchenko
D. Panchenko. The Sherrington-Kirkpatrick Model. Springer Monographs in Mathematics. Springer, New York, NY, 2013. ISBN 978-1-4614-6288-0 978-1-4614-6289-7. doi: 10.1007/978-1-4614-6289-7. 11
2013 doi
-
[16]
A. Pinkus. Strictly Positive Definite Functions on a Real Inner Product Space. Advances in Computational Mathematics, 20(4):263–271, May 2004. ISSN 1572-9044. doi: 10.1023/A:1027362918283
2004 doi
-
[17]
C. E. Rasmussen and C. K. Williams. Gaussian Processes for Machine Learning. Number 3 in Adaptive Computation and Machine Learning. MIT Press, Cambridge, Massachusetts, 2 edition, 2006. ISBN 0-262-18253- X. URL http://gaussianprocess.org/gpml/chapters/RW.pdf
2006
-
[18]
M. Reimer. Gegenbauer Polynomials. In Multivariate Polynomial Approx- imation, pages 19–38. Birkh¨ auser, Basel, 2003. ISBN 978-3-0348-8095-4. doi: 10.1007/978-3-0348-8095-4 2
2003 doi
-
[19]
Sasv´ ari
Z. Sasv´ ari. Multivariate Characteristic and Correlation Functions . Num- ber 50 in De Gruyter Studies in Mathematics. Walter de Gruyter, Berlin/Boston, Mar. 2013. ISBN 978-3-11-022399-6
2013
-
[20]
I. J. Schoenberg. Metric spaces and positive definite functions. Transactions of the American Mathematical Society , 44(3):522–536,
-
[21]
I. J. Schoenberg. Positive definite functions on spheres. Duke Mathematical Journal, 9(1):96–108, Mar. 1942. ISSN 0012-7094, 1547-7398. doi: 10.1215/ S0012-7094-42-00908-6
1942
-
[22]
Sch¨ olkopf and A
B. Sch¨ olkopf and A. J. Smola.Learning with Kernels: Support Vector Ma- chines, Regularization, Optimization, and Beyond . MIT Press, Cambridge, Mass., 1 edition, 2002. ISBN 978-0-262-53657-8
2002
-
[23]
J. B. Simon, S. Anand, and M. Deweese. Reverse Engineering the Neural Tangent Kernel. In Proceedings of the 39th International Conference on Machine Learning, pages 20215–20231. PMLR, June 2022. URL https: //proceedings.mlr.press/v162/simon22a.html
2022
-
[24]
Williams
C. Williams. Computing with Infinite Networks. In Advances in Neural Information Processing Systems , volume 9. MIT Press,
-
[25]
G. Yang. Tensor Programs I: Wide Feedforward or Recur- rent Neural Networks of Any Architecture are Gaussian Pro- cesses. In Advances in Neural Information Processing Systems , vol- ume 32, Vancouver, Canada, 2019. Curran Associates, Inc. URL https://proceedings.neurips.cc/pap...
2019
-
[26]
⇐”: This follows directly from the fact that U is a linear isometry that pre- serves norms and inner products. “⇒
G. Yang. Tensor Programs II: Neural Tangent Kernel for Any Architecture, Nov. 2020. URL http://arxiv.org/abs/2006.14548. 12 A Appendix A.1 Proofs for Section 2 Since we are interested in characterizing isotropic kernels on Rd we consider Sd−1, whereas the source we build our p...
2020 arXiv
-
[31]
Case d< ∞: For all n∈ N and allr,s∈ (0,∞) we have by Theorem A.1 α(d) n (r,s ) =Z ∫1 −1 =fS(r,s,ρ) /bracehtipdownleft/bracehtipupright/bracehtipupleft/bracehtipdownright κ(r,s,rsρ ) ˆC(λ) n (ρ)w(ρ)dρ. To extendα(d) n to a continuous function on [0,∞)2, which yields (5) by defi...
-
[32]
(ii), (iii), (iv)⇒ (i)
Case d =∞: First note that the α(∞) n are continuous positive definite kernels on (0,∞) [10, Example 2.4]. For the extension we select two orthonormal vectorse1,e 2 and observe K(re1,re 1)−K(re1,re 2) = ∞∑ n=0 α(∞) n (r,r )⟨e1,e 1⟩n− ∞∑ n=0 α(∞) n (r,r )⟨e1,e 2⟩n = ∞∑ n=1 α(∞)...
-
[33]
(i)⇒ (ii), (iii)
on (0,∞)× Sd−1. Since these conditions are satisfied for K>0 if they are satisfied for K we thereby see that K>0 is strictly positive definite on Rd\{ 0}. 17 If m> 1 the second term in (13) is therefore non-zero. If m = 1 then the first term is m∑ i,j=1 cicjK0(xi,xj) =|c1|2α(d...
-
[34]
Remark A.4 (Universal kernels)
replaced by Theorems 3.6, 3.8 and 3.9 respectively. Remark A.4 (Universal kernels) . For future work characterizing universal isotropic kernels we point out their characterization on the sphere by Micchelli et al. [14, Theorem 10]. It might be possible to extend this result to...
-
[940]
doi: 10.1007/s00365-016-9323-9
-
[1807]
doi: 10.1007/BF01452844. 10
-
[1938]
URL https://community.ams.org/journals/tran/1938-044-03/ S0002-9947-1938-1501980-0/S0002-9947-1938-1501980-0.pdf
1938
-
[1996]
URL https://proceedings.neurips.cc/paper/1996/hash/ ae5e3ce40e0404a45ecacaaf05e5f735-Abstract.html
1996
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.