Pith. sign in

REVIEW 2 minor 12 references

How Deep Are Deep GPs, Really? A Sharp Threshold and a Non-Gaussian Limit for Compositional GPs

T0 review · 0 major / 2 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read Compositional Gaussian process priors converge to non-Gaussian limits below a bandwidth threshold of order sqrt(d).

desk verdict The paper sharpens the degeneracy threshold for deep RBF GP priors to Θ(√d) and proves non-Gaussian non-degenerate limits with coordinate dependence below it. read the letter →

arxiv 2606.08218 v1 pith:HHSKPRTF submitted 2026-06-06 cs.LG cs.AImath.STstat.MLstat.TH

classification cs.LGcs.AImath.STstat.MLstat.TH
keywords deepGaussianprocessescompositionalpriorsbandwidththresholdnon-GaussianlimitsRBFkernelwidelimitdegeneracy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper studies the limiting behavior of deep compositional Gaussian process priors as the number of layers grows. It identifies a sharp threshold on the RBF kernel bandwidth that separates a regime where the prior collapses to constant functions from one where it approaches a different distribution. Below the threshold the limit is shown to be non-degenerate and non-Gaussian while retaining dependence between coordinates. This matters because it demonstrates that deep GP priors can remain useful non-trivial models rather than always degenerating with depth.

What carries the argument

The sharp bandwidth threshold r_c(d) = Θ(√d) that separates the degenerate constant limit from the non-Gaussian limit distribution π_{ar{Z}} of the depth-asymptotic compositional GP prior.

What would settle it

Sample many realizations of a deep GP at increasing depths for a fixed bandwidth slightly below sqrt(d) and check whether the joint distribution of function values at several fixed input points stabilizes to a non-constant non-Gaussian form with persistent coordinate correlations.

Watch

Extended reading notes

Core claim

For compositional vector-valued Gaussian processes with the RBF kernel in the wide-network limit, when the bandwidth r lies below the sharp threshold r_c(d) = Θ(√d) the prior converges to a limit distribution π_{ar{Z}}. These limit distributions are non-degenerate and non-Gaussian, with non-vanishing dependence between coordinates. Above the threshold the prior degenerates to constants. The threshold sharpens earlier bounds on the degenerate regime.

Load-bearing premise

Each layer is an independent vector-valued Gaussian process with the RBF kernel in the infinite-width limit and with the same fixed bandwidth r used across all layers.

Editorial extensions

If this is right

  • Below the threshold deep GP priors admit non-trivial limiting distributions that can serve as probabilistic models for layered functions.
  • The limiting distributions are non-Gaussian and exhibit non-vanishing dependence between coordinates.
  • The transition between degenerate and non-degenerate regimes occurs sharply at bandwidth order sqrt(d).
  • Empirical samples of the limits display complex multimodal structure whose width narrows with dimension.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The narrow window of usable bandwidths around sqrt(d) may require careful tuning when applying deep GPs in high dimensions.
  • Similar thresholds could exist for other kernels or for finite-width networks, offering a route to test the wide-limit predictions.
  • Non-Gaussian limits suggest that uncertainty estimates from deep GPs may differ qualitatively from those of single-layer GPs.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 2 minor

Summary. The manuscript analyzes the depth-n recursion of vector-valued RBF Gaussian process priors in the wide-network limit. It identifies a sharp threshold r_c(d) = Θ(√d) separating a degenerate regime (convergence to constant functions) from a non-degenerate regime, proves that for r < r_c(d) the prior converges to a non-degenerate non-Gaussian fixed-point measure π_{ar Z} with non-vanishing coordinate dependence, and supplies empirical verification of the threshold together with illustrations of multimodal structure in the limit distributions.

Significance. If the stated proofs hold, the work supplies a precise, dimension-dependent characterization of when compositional GP priors remain non-trivial at large depth, extending earlier degeneracy results for the RBF kernel. The demonstration that non-Gaussian limits with persistent dependence exist below the threshold is a substantive advance for the theory of deep Bayesian models. Credit is given for the mathematical analysis establishing the sharp threshold and the non-Gaussian character of the limit, as well as for the empirical confirmation across dimensions.

minor comments (2)
  1. Abstract and §1: the notation π_{ar Z} is introduced without an immediate informal description of the measure; adding one sentence would improve accessibility for readers outside the immediate subfield.
  2. Empirical section: the figures illustrating multimodal behavior of π_{ar Z} would benefit from explicit axis labels or captions stating the precise values of d and r used in each panel.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for their positive summary, recognition of the significance of the sharp threshold result, and recommendation of minor revision. The report raises no specific major comments requiring response.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity detected

full rationale

The paper's central results consist of mathematical proofs identifying a sharp bandwidth threshold r_c(d) = Θ(√d) and establishing convergence of the depth-n recursion of vector-valued RBF GPs (in the wide limit) to a non-degenerate non-Gaussian fixed-point measure π_{ar{Z}} with persistent coordinate dependence for r below the threshold. These statements are derived directly from the stated modeling assumptions (independent layers, fixed r, RBF kernel, wide-network limit taken first) via analysis of the recursion; no step reduces by construction to a fitted input, self-definition, or load-bearing self-citation. The derivation is therefore self-contained against external benchmarks.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central claim rests on standard properties of Gaussian processes and the compositional structure; no free parameters are fitted as this is a theoretical limit analysis.

assumptions (2)
  • domain assumption Each layer is a vector-valued Gaussian process with RBF kernel.
    This is the setup for the compositional prior as described in the abstract.
  • domain assumption The wide-network limit applies to the prior.
    Mentioned as the context for the GP prior.

how reviews work

0 comments
Cite this review

Pith. "Pith review of How Deep Are Deep GPs, Really? A Sharp Threshold and a Non-Gaussian Limit for Compositional GPs." pith.science (2026). https://pith.science/paper/HHSKPRTF

@misc{pith2026260608218,
  author       = {Pith},
  title        = {Pith review of: How Deep Are Deep GPs, Really? A Sharp Threshold and a Non-Gaussian Limit for Compositional GPs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HHSKPRTF}},
  note         = {Machine review of arXiv:2606.08218}
}
abstract

Compositional priors describe the generic properties of layered functions in deep Bayesian models, where deep neural networks with random weights are a canonical example.In the wide-network limit, the prior is a Gaussian process with a depth-dependent kernel, and its behaviour as depth grows has been extensively studied through this kernel. Here, we study another case, where each layer itself is a vector valued Gaussian process, and our aim is similarly to understand the limiting behaviour of the prior as depth grows. Previous GP work has established that for the RBF kernel and a certain range of bandwidths $r$, the prior degenerates in the limit, converging to the set of constant functions -- which is not useful as a probabilistic model. In this paper we establish several new results. First, we identify a sharp bandwidth threshold $r_c(d) = \Theta(\sqrt{d})$ above which the limit is degenerate, strengthening the earlier bounds. Second, and more importantly, we show that for $r$ below the threshold $r_c(d)$ the prior converges to a limit distribution $\pi_{\bar{Z}}$. We also prove that these distributions are non-degenerate and non-Gaussian, with non-vanishing dependence between coordinates. In contrast to the previously known degenerate regime, deep Gaussian process priors can therefore admit non-trivial limits. Empirically, we verify the threshold across a range of dimensions $d$, and demonstrate a complex multimodal behaviour of the limit distributions $\pi_{\bar{Z}}$ -- a regime that becomes increasingly narrow with $d$ and would be hard to identify without knowing the threshold.

Figures

Figures reproduced from arXiv: 2606.08218 by the authors.

Figure 1
Figure 1. The compositional chain (1) on n input points. At depth i the same random Gaussian function Gi is applied to every one of the n points, coupling all trajectories through the shared randomness. Convergence and sharp thresholds. We give a sharp almost-sure threshold separating a supercrit￾ical regime, in which the chain synchronises, from a subcritical regime, in which it admits a nontrivial stationary law. The critic… view at source ↗
Figure 2
Figure 2. The dichotomy of Theorem 4.1 at d = 100. Left (supercritical): trajectories of mean log v 2 i over 300 i.i.d. copies starting at v1 = 1, for λ ∈ {1.1, 1.5, 2.5}; the grey dashed reference line has slope −2 log λ predicted by Theorem 4.1 (i) and is hidden behind the simulated curve in every case. Middle (subcritical): the same kind of trajectories for ten λ values spanning 0.30 to 0.999 (purple to yellow). All but th… view at source ↗
Figure 3
Figure 3. Per-chain t-SNE embeddings of Z¯ i,t ∈ R 100 at depth i = 1000, for three subcritical λ (rows) and five i.i.d. chains (columns); n = 1000 points per panel, perplexity 30, PCA initialisation. The visible-structure window in d = 100 is narrow: λ = 0.90 (top row) embeds as a featureless isotropic blob in every chain, indistinguishable across chains; λ = 0.99 and λ = 0.995 (lower rows) produce chain-specific multimodal … view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Supercritical regime r > rc(d): trajectories of log v 2 i for λ ∈ {1.1, 1.5, 2.5} across d = 1, 10, 100. Each curve is the mean of log v 2 i over 300 i.i.d. trajectories starting at v1 = 1; the grey dashed reference line has slope −2 log λ predicted by Theorem 4.1 (i),…
Figure 5
Figure 5. Figure 5: Subcritical regime r < rc(d): trajectories of log v 2 i for ten λ values spanning 0.30 to 0.999, across d = 1, 10, 100, all panels at uniform depth 2000. Each curve is the mean of log v 2 i over 300 i.i.d. trajectories starting at v1 = 1, colour-coded from low (purple)…
Figure 6
Figure 6. Figure 6: Empirical stationary law of log v 2 i at λ ∈ {0.40, 0.70, 0.94}, overlaid per dimension (d = 1, 10, 100 left to right). Each histogram is from a chain of 2 ·105 samples after a 2·104 -step burn-in; the d = 1 pane trims the leftmost 8% of each sample to keep the bulk vi…
Figure 7
Figure 7. Figure 7: Sample std of log v 2 i across the 300 i.i.d. copies used in Figures 4 and 5, vs. depth i. Top row: supercritical regime, λ ∈ {1.1, 1.5, 2.5}, depth 200. Bottom row: subcritical regime, λ ∈ {0.30, . . . , 0.999}, depth 2000; the std saturates as the chain approaches it…
Figure 8
Figure 8. Figure 8: Per-chain histograms of Z¯ i,t ∈ R at depth i = 300, d = 1. Rows: λ ∈ {0.10, 0.30, 0.60, 0.85}. Dashed: N (0, 1). The marginal stays exactly N (0, 1) at every depth, so any departure from the dashed curve is a within-chain dependence effect. The top row (λ = 0.10) is e…
Figure 9
Figure 9. Figure 9: Per-chain PC1–PC2 scatters of Z¯ i,t ∈ R 10 at depth i = 600, five i.i.d. chains (columns); n = 1000 points per panel, centred and projected onto the top two PCs. Rows: λ ∈ {0.66, 0.91, 0.95, 0.97}. The top two rows are still close to a Gaussian cloud; the lower rows s…
Figure 10
Figure 10. Figure 10: Empirical PC1–PC2 density (Gaussian KDE, Scott bandwidth) of the same per-chain projections [PITH_FULL_IMAGE:figures/full_fig_p024_10.png]
Figure 11
Figure 11. Figure 11: Per-chain t-SNE embeddings of Z¯ i,t at d = 10, depth i = 600, perplexity 30, PCA initialisation. Same λ and chains as [PITH_FULL_IMAGE:figures/full_fig_p025_11.png]
Figure 12
Figure 12. Figure 12: Per-chain PC1–PC2 scatters of Z¯ i,t ∈ R 100 at depth i = 1000, the linear-PCA companion to [PITH_FULL_IMAGE:figures/full_fig_p026_12.png]
Figure 13
Figure 13. Figure 13: Empirical PC1–PC2 density (Gaussian KDE, Scott bandwidth) of the same per-chain projections [PITH_FULL_IMAGE:figures/full_fig_p027_13.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

12 extracted references · 3 canonical work pages

  1. [1]

    Soufiane Hayou, Arnaud Doucet, and Judith Rousseau

    arXiv:2210.00688. Soufiane Hayou, Arnaud Doucet, and Judith Rousseau. On the impact of the activation function on deep neural networks training. InProceedings of the 36th International Conference on Machine Learning (ICML),

  2. [2]

    Interpretable deep Gaussian processes with moments.arXiv preprint arXiv:1905.10963,

    Chi-Ken Lu, Scott Cheng-Hsin Yang, Xiaoran Hao, and Patrick Shafto. Interpretable deep Gaussian processes with moments.arXiv preprint arXiv:1905.10963,

  3. [3]

    Scaling limits of wide neural networks with weight sharing: Gaussian process behavior, gradient independence, and neural tangent kernel derivation

    Greg Yang. Scaling limits of wide neural networks with weight sharing: Gaussian process behavior, gradient independence, and neural tangent kernel derivation. InarXiv preprint arXiv:1902.04760,

  4. [4]

    Then (11) reads ui+1 =F(u i)g 2 i .(12) Ifu 1 = 0thenu i ≡0; otherwiseu i >0a.s

    Proof of Theorem A.1.Setu i :=v 2 i ≥0,L i := logu i, andF(u) := 2(1−e −u/(2r2)). Then (11) reads ui+1 =F(u i)g 2 i .(12) Ifu 1 = 0thenu i ≡0; otherwiseu i >0a.s. for alli(sinceg i ̸= 0a.s.), and (12) gives Li+1 =L i +H(L i) + logg2 i ,(13) where H(L) := log F(e L)/eL . The functionH is continuous onR, satisfies H(L) → −2 logr as L→ −∞ (sinceF(u)/u→1/r 2 ...

  5. [5]

    By the strong Markov property applied atτk + 1 (from which point the chain restarts from the positive stateuτk+1), the same argument givesτk+1 <∞ a.s

    Iterate: set τ0 := 0and τk+1 := inf{i > τ k :u i > δ}. By the strong Markov property applied atτk + 1 (from which point the chain restarts from the positive stateuτk+1), the same argument givesτk+1 <∞ a.s. on{τ k <∞}. By induction everyτ k is a.s. finite, soui > δfor infinitely manyia.s.; in particular lim sup i→∞ ui ≥δ >0a.s., sov i =± √ui doesnotconverg...

  6. [6]

    distance

    combine to give positive recurrence and convergence to equilibrium under a Lyapunov drift condition. First,Foster’s drift criterion(Theorem 11.3.4 of Meyn and Tweedie [1993]): if the chain isψ-irreducible and there exist a measurable functionV : X → [0,∞ ), a small setC, and constantsc >0,b <∞, such that the drift condition EV(X i+1)|X i =x≤V(x)−c+b1 C(x)...

  7. [7]

    Now we use thisκ to delimit the small set

    The constant κ is fixed by the chain and the choices ofL0, L0. Now we use thisκ to delimit the small set. We want a strictly negative drift≤ −κ in Region 2, i.e. EV (Li+1) |L i = L≤V (L) −κ, which (using the uniform bound above) holds wheneverV (L) ≥ 2κ. Since V (L) = L−L 0 in Region 2, this requiresL≥L 0 + 2κ — and this is precisely why the small set is ...

  8. [8]

    The SLLN argument of part (i) goes through verbatim with logg 2 j replaced bylogX j and Elogg 2 replaced byElogX , givinglim supi Li/i≤ρ d a.s

    The pointwise bound H(L) ≤ − 2 logr from (14) is unchanged (it uses only1 −e −x ≤x ); the limit H(L) → −2 logr as L→ −∞ is unchanged. The SLLN argument of part (i) goes through verbatim with logg 2 j replaced bylogX j and Elogg 2 replaced byElogX , givinglim supi Li/i≤ρ d a.s. The non-convergence step in part (ii) is identical (zero-mean random walkSn :=P...

Show all 12 references
  1. [9]

    : maxs<t |logu st| ≤M o , a compact subset of(0 ,∞ )( n 2). Foru /∈C there exists at least one pair( s0, t0)with |Ls0t0 |> M ; that pair contributes ∆(Ls0t0)≤ −B n 2 −1by the choice ofM, while every other pair contributes at mostB, giving X s<t ∆(Lst)≤ −B n 2 −1 +B n 2 −1 =−B−...

  2. [10]

    that is strictly positive jointly on any compact subset of the interior timesC. This givesψ-irreducibility of u(i) with respect to Lebesgue measure on(0,∞ )( n 2), the small-set minorisationP (u,· ) ≥η ν (·)on C for someη >0and probability measureν, and strong aperiodicity (m=...

  3. [11]

    Hence(¯Z∞,1, ¯Z∞,2)is not jointly Gaussian. Proof of Theorem 4.6.Finiteness of Eπv |L|: the Lyapunov functionV (L) = (L0 −L )+ + (L−L 0)+ from the proof of Theorem A.1 (ii) satisfiesEπv V <∞ (standard for positive Harris recurrent chains with a Foster– Lyapunov drift function;...

  4. [12]

    trajectories starting at v1 = 1; the grey dashed reference line has slope −2 logλ predicted by Theorem 4.1 (i), and is hidden behind the simulated curve in every case

    Each curve is the mean oflogv 2 i over300i.i.d. trajectories starting at v1 = 1; the grey dashed reference line has slope −2 logλ predicted by Theorem 4.1 (i), and is hidden behind the simulated curve in every case. At matched λ, the decay rate is the same across dimensions, c...

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.