REVIEW 2 minor 12 references
How Deep Are Deep GPs, Really? A Sharp Threshold and a Non-Gaussian Limit for Compositional GPs
T0 review · 0 major / 2 minor · reviewed 2026-06-27 · grok-4.3
Pith's one-line read Compositional Gaussian process priors converge to non-Gaussian limits below a bandwidth threshold of order sqrt(d).
desk verdict The paper sharpens the degeneracy threshold for deep RBF GP priors to Θ(√d) and proves non-Gaussian non-degenerate limits with coordinate dependence below it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The sharp bandwidth threshold r_c(d) = Θ(√d) that separates the degenerate constant limit from the non-Gaussian limit distribution π_{ar{Z}} of the depth-asymptotic compositional GP prior.
What would settle it
Sample many realizations of a deep GP at increasing depths for a fixed bandwidth slightly below sqrt(d) and check whether the joint distribution of function values at several fixed input points stabilizes to a non-constant non-Gaussian form with persistent coordinate correlations.
Extended reading notes
Core claim
For compositional vector-valued Gaussian processes with the RBF kernel in the wide-network limit, when the bandwidth r lies below the sharp threshold r_c(d) = Θ(√d) the prior converges to a limit distribution π_{ar{Z}}. These limit distributions are non-degenerate and non-Gaussian, with non-vanishing dependence between coordinates. Above the threshold the prior degenerates to constants. The threshold sharpens earlier bounds on the degenerate regime.
Load-bearing premise
Each layer is an independent vector-valued Gaussian process with the RBF kernel in the infinite-width limit and with the same fixed bandwidth r used across all layers.
Editorial extensions
If this is right
- Below the threshold deep GP priors admit non-trivial limiting distributions that can serve as probabilistic models for layered functions.
- The limiting distributions are non-Gaussian and exhibit non-vanishing dependence between coordinates.
- The transition between degenerate and non-degenerate regimes occurs sharply at bandwidth order sqrt(d).
- Empirical samples of the limits display complex multimodal structure whose width narrows with dimension.
Reading between the lines
- The narrow window of usable bandwidths around sqrt(d) may require careful tuning when applying deep GPs in high dimensions.
- Similar thresholds could exist for other kernels or for finite-width networks, offering a route to test the wide-limit predictions.
- Non-Gaussian limits suggest that uncertainty estimates from deep GPs may differ qualitatively from those of single-layer GPs.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript analyzes the depth-n recursion of vector-valued RBF Gaussian process priors in the wide-network limit. It identifies a sharp threshold r_c(d) = Θ(√d) separating a degenerate regime (convergence to constant functions) from a non-degenerate regime, proves that for r < r_c(d) the prior converges to a non-degenerate non-Gaussian fixed-point measure π_{ar Z} with non-vanishing coordinate dependence, and supplies empirical verification of the threshold together with illustrations of multimodal structure in the limit distributions.
Significance. If the stated proofs hold, the work supplies a precise, dimension-dependent characterization of when compositional GP priors remain non-trivial at large depth, extending earlier degeneracy results for the RBF kernel. The demonstration that non-Gaussian limits with persistent dependence exist below the threshold is a substantive advance for the theory of deep Bayesian models. Credit is given for the mathematical analysis establishing the sharp threshold and the non-Gaussian character of the limit, as well as for the empirical confirmation across dimensions.
minor comments (2)
- Abstract and §1: the notation π_{ar Z} is introduced without an immediate informal description of the measure; adding one sentence would improve accessibility for readers outside the immediate subfield.
- Empirical section: the figures illustrating multimodal behavior of π_{ar Z} would benefit from explicit axis labels or captions stating the precise values of d and r used in each panel.
Simulated Author's Rebuttal
We thank the referee for their positive summary, recognition of the significance of the sharp threshold result, and recommendation of minor revision. The report raises no specific major comments requiring response.
Circularity Check
No significant circularity detected
full rationale
The paper's central results consist of mathematical proofs identifying a sharp bandwidth threshold r_c(d) = Θ(√d) and establishing convergence of the depth-n recursion of vector-valued RBF GPs (in the wide limit) to a non-degenerate non-Gaussian fixed-point measure π_{ar{Z}} with persistent coordinate dependence for r below the threshold. These statements are derived directly from the stated modeling assumptions (independent layers, fixed r, RBF kernel, wide-network limit taken first) via analysis of the recursion; no step reduces by construction to a fitted input, self-definition, or load-bearing self-citation. The derivation is therefore self-contained against external benchmarks.
Assumptions & free parameters
assumptions (2)
- domain assumption Each layer is a vector-valued Gaussian process with RBF kernel.
- domain assumption The wide-network limit applies to the prior.
Cite this review
Pith. "Pith review of How Deep Are Deep GPs, Really? A Sharp Threshold and a Non-Gaussian Limit for Compositional GPs." pith.science (2026). https://pith.science/paper/HHSKPRTF
@misc{pith2026260608218,
author = {Pith},
title = {Pith review of: How Deep Are Deep GPs, Really? A Sharp Threshold and a Non-Gaussian Limit for Compositional GPs},
year = {2026},
howpublished = {\url{https://pith.science/paper/HHSKPRTF}},
note = {Machine review of arXiv:2606.08218}
}
abstract
Compositional priors describe the generic properties of layered functions in deep Bayesian models, where deep neural networks with random weights are a canonical example.In the wide-network limit, the prior is a Gaussian process with a depth-dependent kernel, and its behaviour as depth grows has been extensively studied through this kernel. Here, we study another case, where each layer itself is a vector valued Gaussian process, and our aim is similarly to understand the limiting behaviour of the prior as depth grows. Previous GP work has established that for the RBF kernel and a certain range of bandwidths $r$, the prior degenerates in the limit, converging to the set of constant functions -- which is not useful as a probabilistic model. In this paper we establish several new results. First, we identify a sharp bandwidth threshold $r_c(d) = \Theta(\sqrt{d})$ above which the limit is degenerate, strengthening the earlier bounds. Second, and more importantly, we show that for $r$ below the threshold $r_c(d)$ the prior converges to a limit distribution $\pi_{\bar{Z}}$. We also prove that these distributions are non-degenerate and non-Gaussian, with non-vanishing dependence between coordinates. In contrast to the previously known degenerate regime, deep Gaussian process priors can therefore admit non-trivial limits. Empirically, we verify the threshold across a range of dimensions $d$, and demonstrate a complex multimodal behaviour of the limit distributions $\pi_{\bar{Z}}$ -- a regime that becomes increasingly narrow with $d$ and would be hard to identify without knowing the threshold.
Figures
Figures from the paper (10 more)
Reference graph
Works this paper leans on
-
[1]
Soufiane Hayou, Arnaud Doucet, and Judith Rousseau
arXiv:2210.00688. Soufiane Hayou, Arnaud Doucet, and Judith Rousseau. On the impact of the activation function on deep neural networks training. InProceedings of the 36th International Conference on Machine Learning (ICML),
-
[2]
Interpretable deep Gaussian processes with moments.arXiv preprint arXiv:1905.10963,
Chi-Ken Lu, Scott Cheng-Hsin Yang, Xiaoran Hao, and Patrick Shafto. Interpretable deep Gaussian processes with moments.arXiv preprint arXiv:1905.10963,
-
[3]
Greg Yang. Scaling limits of wide neural networks with weight sharing: Gaussian process behavior, gradient independence, and neural tangent kernel derivation. InarXiv preprint arXiv:1902.04760,
-
[4]
Then (11) reads ui+1 =F(u i)g 2 i .(12) Ifu 1 = 0thenu i ≡0; otherwiseu i >0a.s
Proof of Theorem A.1.Setu i :=v 2 i ≥0,L i := logu i, andF(u) := 2(1−e −u/(2r2)). Then (11) reads ui+1 =F(u i)g 2 i .(12) Ifu 1 = 0thenu i ≡0; otherwiseu i >0a.s. for alli(sinceg i ̸= 0a.s.), and (12) gives Li+1 =L i +H(L i) + logg2 i ,(13) where H(L) := log F(e L)/eL . The functionH is continuous onR, satisfies H(L) → −2 logr as L→ −∞ (sinceF(u)/u→1/r 2 ...
2019
-
[5]
By the strong Markov property applied atτk + 1 (from which point the chain restarts from the positive stateuτk+1), the same argument givesτk+1 <∞ a.s
Iterate: set τ0 := 0and τk+1 := inf{i > τ k :u i > δ}. By the strong Markov property applied atτk + 1 (from which point the chain restarts from the positive stateuτk+1), the same argument givesτk+1 <∞ a.s. on{τ k <∞}. By induction everyτ k is a.s. finite, soui > δfor infinitely manyia.s.; in particular lim sup i→∞ ui ≥δ >0a.s., sov i =± √ui doesnotconverg...
1993
-
[6]
distance
combine to give positive recurrence and convergence to equilibrium under a Lyapunov drift condition. First,Foster’s drift criterion(Theorem 11.3.4 of Meyn and Tweedie [1993]): if the chain isψ-irreducible and there exist a measurable functionV : X → [0,∞ ), a small setC, and constantsc >0,b <∞, such that the drift condition EV(X i+1)|X i =x≤V(x)−c+b1 C(x)...
1993
-
[7]
Now we use thisκ to delimit the small set
The constant κ is fixed by the chain and the choices ofL0, L0. Now we use thisκ to delimit the small set. We want a strictly negative drift≤ −κ in Region 2, i.e. EV (Li+1) |L i = L≤V (L) −κ, which (using the uniform bound above) holds wheneverV (L) ≥ 2κ. Since V (L) = L−L 0 in Region 2, this requiresL≥L 0 + 2κ — and this is precisely why the small set is ...
1993
-
[8]
The SLLN argument of part (i) goes through verbatim with logg 2 j replaced bylogX j and Elogg 2 replaced byElogX , givinglim supi Li/i≤ρ d a.s
The pointwise bound H(L) ≤ − 2 logr from (14) is unchanged (it uses only1 −e −x ≤x ); the limit H(L) → −2 logr as L→ −∞ is unchanged. The SLLN argument of part (i) goes through verbatim with logg 2 j replaced bylogX j and Elogg 2 replaced byElogX , givinglim supi Li/i≤ρ d a.s. The non-convergence step in part (ii) is identical (zero-mean random walkSn :=P...
1993
Show all 12 references
-
[9]
: maxs<t |logu st| ≤M o , a compact subset of(0 ,∞ )( n 2). Foru /∈C there exists at least one pair( s0, t0)with |Ls0t0 |> M ; that pair contributes ∆(Ls0t0)≤ −B n 2 −1by the choice ofM, while every other pair contributes at mostB, giving X s<t ∆(Lst)≤ −B n 2 −1 +B n 2 −1 =−B−...
1993
-
[10]
that is strictly positive jointly on any compact subset of the interior timesC. This givesψ-irreducibility of u(i) with respect to Lebesgue measure on(0,∞ )( n 2), the small-set minorisationP (u,· ) ≥η ν (·)on C for someη >0and probability measureν, and strong aperiodicity (m=...
1993
-
[11]
Hence(¯Z∞,1, ¯Z∞,2)is not jointly Gaussian. Proof of Theorem 4.6.Finiteness of Eπv |L|: the Lyapunov functionV (L) = (L0 −L )+ + (L−L 0)+ from the proof of Theorem A.1 (ii) satisfiesEπv V <∞ (standard for positive Harris recurrent chains with a Foster– Lyapunov drift function;...
1993
-
[12]
trajectories starting at v1 = 1; the grey dashed reference line has slope −2 logλ predicted by Theorem 4.1 (i), and is hidden behind the simulated curve in every case
Each curve is the mean oflogv 2 i over300i.i.d. trajectories starting at v1 = 1; the grey dashed reference line has slope −2 logλ predicted by Theorem 4.1 (i), and is hidden behind the simulated curve in every case. At matched λ, the decay rate is the same across dimensions, c...
Reviewed June 27, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.