Pith. sign in

REVIEW 5 minor 24 references

Large deviation principles for convolutional Bayesian neural networks

T0 review · 0 major / 5 minor · reviewed 2026-07-15 · grok-4.5

Pith's one-line read The first large-deviation principle for convolutional Bayesian networks shows how their random covariances deviate from the Gaussian-process limit as channels grow.

desk verdict First clean LDP for infinite-channel CNNs via general patch extractors; solid extension of known FCNN results with no real holes. read the letter →

arxiv 2603.06023 v2 pith:2ZQFQTSQ submitted 2026-03-06 math.PR stat.ML

classification math.PRstat.ML MSC 60F1060B2068T07
keywords largedeviationsconvolutionalneuralnetworksBayesianinfinite-channellimitGaussianprocessescovarianceconcentrationpatchextractors
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Wide convolutional networks with Gaussian weights are already known to look like Gaussian processes once the number of channels becomes large. This paper asks what happens just beyond that limit: how rare are the atypical covariance configurations that survive after the law of large numbers has kicked in? The authors encode a broad family of multi-dimensional CNN architectures by a general patch-extractor map and prove that the sequence of random conditional covariance matrices obeys a large-deviation principle whose rate function is an explicit Legendre transform built from those extractors. The same rate governs the posterior after finitely many observations, and a by-product is a short proof that the covariances concentrate and the network becomes Gaussian. A sympathetic reader cares because the result gives the first precise exponential-cost description of atypical behaviour for convolutional architectures, not only fully connected ones.

What carries the argument

The Markov chain of conditional covariance matrices whose transition kernels are empirical averages of the continuous patch-extractor map G^(ℓ); its large-deviation rate is obtained by verifying the conditional LDP continuity condition and then upgrading the resulting weak LDP by exponential tightness.

What would settle it

Construct a continuous activation and patch extractor that violate the little-o remainder condition, compute the empirical moment-generating function of the resulting G^(ℓ) for large but finite channel counts, and check whether the observed large-deviation rate still coincides with the Legendre transform claimed in Theorem 3.3.

Watch

Extended reading notes

Core claim

Under a Gaussian prior, linear channel growth, and an asymptotic Lipschitz condition on the activation and patch extractors, the joint law of the random covariance tensors of a multi-layer CNN satisfies a large-deviation principle with good rate function equal to a weighted sum of layer-wise Legendre transforms of the log-moment generating functions of the patch-extractor maps.

Load-bearing premise

The activation and every receptive-field map must be almost Lipschitz at infinity (their difference is controlled by a linear term plus a remainder that grows slower than the size of the input); without that control the exponential-equivalence step that identifies the rate function can fail.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. The paper establishes large-deviation principles for a broad class of multidimensional convolutional Bayesian neural networks in the infinite-channel regime. Under a Gaussian prior on the weights (A1), linear growth of the channel counts (A2), and an asymptotic Lipschitz condition on the activation and patch extractors (A4), the sequence of random conditional covariance tensors (K^(2,n),…,K^(L+1,n)) satisfies an LDP on the product of positive-semidefinite cones with good rate function I_{2,…,L+1}(Q_2,…,Q_{L+1})=α_1 I_1(Q_2|K^(1))+∑_{ℓ=2}^L α_ℓ I_ℓ(Q_{ℓ+1}|Q_ℓ), where each I_ℓ is the Legendre transform of the log-moment generating function of the patch-extractor map G^(ℓ) (Theorem 3.3). The same rate governs the posterior of the final covariance after conditioning on finitely many Gaussian observations (Proposition 3.5), and a rescaled network-output LDP follows by contraction (Proposition 3.6). As corollaries under the weaker growth condition (A3) the authors recover covariance concentration (Theorem 3.1) and Gaussian-process convergence (Theorem 3.2). The argument proceeds by identifying the Markov structure of the covariances (Proposition 4.1), verifying conditional LDP continuity of the transition kernels via Cramér plus exponential equivalence (Lemmas 6.3–6.6, Proposition 6.7), establishing joint exponential tightness (Proposition 6.9), and inducting to the full LDP (Proposition 6.10).

Significance. This appears to be the first large-deviation principle for convolutional architectures. It extends the known Gaussian-process limit of wide CNNs (Novak et al., Yang et al.) to a full large-deviation description of the covariance process, covers general receptive fields via a mild patch-extractor formalism, and simultaneously yields a streamlined proof of concentration and Gaussian equivalence. The rate function is derived from first principles as an explicit Legendre transform; no free parameters are fitted. The technical pipeline (Markov structure, conditional LDP continuity, exponential tightness, induction) is standard and carefully executed, and Assumption (A4) is weaker than the corresponding hypotheses used for fully-connected networks in the cited literature. The results therefore constitute a genuine and useful advance for the probabilistic theory of Bayesian CNNs.

minor comments (5)
  1. Throughout the manuscript the word “convolutional” is misspelled as “CONVOULUTIONAL” in several running headers and page titles (e.g., pages 3, 5, 7, …). A global search-and-replace is needed.
  2. In the definition of the patch extractor for the 2-D zero-padding example (Eq. (4) and the surrounding display), the ordering of the nine mask offsets is listed inconsistently with the set M defined in (3); while the mathematics is unaffected, a uniform ordering would improve readability.
  3. Lemma 3.4 and Proposition 3.5 are stated without proofs, the authors referring the reader to the analogous FCNN arguments in [2]. A short sketch of the Gaussian-manipulation steps that produce the posterior density of K^(L+1,n) would make the paper more self-contained.
  4. The notation for the flattened versus tensor forms of G^(ℓ) and K^(ℓ) is introduced in Remark 2 but then used interchangeably; a single sentence reminding the reader that all matrix norms and traces are understood after the standard (i,μ) o k(i,μ) flattening would avoid occasional ambiguity.
  5. In Assumption (A2) the phrase “increase linearly as a function of n” is slightly informal; writing C_ℓ(n)/n oα_ℓ∈(0,∞) already appears later and could be placed in the assumption itself for precision.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the LDP rate function is the Legendre transform of an explicit log-mgf of the patch-extractor maps, obtained via Cramér + exponential equivalence + conditional LDP continuity + tightness, all proved in-place.

full rationale

The derivation chain for the central claim (Theorem 3.3) is self-contained. The covariance sequence is a Markov chain with kernels ν_{ℓ+1,n} built from the explicit maps G^{(ℓ)} (Definition 4.2, Proposition 4.1). Under (A6) the kernels satisfy the conditional LDP continuity condition of Chaganty with rate α_ℓ I_ℓ (Proposition 6.7), where I_ℓ is the Legendre transform of the log-moment generating function M_ℓ of G^{(ℓ)} (eqs. (9),(19)). Exponential equivalence of the empirical averages S_n and ˜S_n follows from the asymptotic Lipschitz bound (Lemma 6.3–6.6); exponential tightness is proved by induction with Markov’s inequality (Proposition 6.9). The full LDP is then obtained by induction + upgrading the weak LDP via tightness (Proposition 6.10). Lemma 4.3 verifies that the CNN patch-extractor maps satisfy (A5)–(A6) under the paper’s (A3)–(A4). No parameters are fitted, no target quantity is inserted by definition, and the self-citations to FCNN LDP papers ([2],[16],[17],[22]) supply only reusable technical lemmas or omitted routine arguments for the posterior (Proposition 3.5) and rescaled-output (Proposition 3.6) corollaries; they do not encode the CNN rate function. The result is therefore not circular.

Assumptions & free parameters 0 free parameters · 4 assumptions · 2 invented entities

The central LDP rests on four modelling assumptions (Gaussian weights, linear growth of channel counts, continuity/growth of activation and patch extractors, asymptotic Lipschitz condition) plus standard large-deviation machinery (Cramér’s theorem, exponential tightness, contraction). No free parameters are fitted; the only numerical constants that appear are the asymptotic channel ratios α_ℓ which are part of the scaling regime, not fitted quantities.

assumptions (4)
  • domain assumption Weights W^(ℓ)_{m,c',c} are i.i.d. N(0,λ_ℓ^{-1}) (Assumption A1).
    Standard Gaussian prior for Bayesian neural nets; used from Proposition 2.1 onward to obtain the conditional Gaussian structure.
  • domain assumption Channel counts satisfy C_ℓ(n)/n o α_ℓ ∈ (0,∞) while depth L, spatial sizes N_ℓ and number of inputs P remain fixed (Assumption A2).
    Defines the infinite-channel scaling regime in which the LDP is stated.
  • domain assumption Activation σ and patch extractors R^(i,ℓ) are continuous and satisfy the exponential growth bound of degree r_σ < 2 (Assumption A3) or the stronger asymptotic Lipschitz condition (Assumption A4).
    A3 yields concentration/Gaussian equivalence; A4 is required for the exponential-equivalence step that identifies the LDP rate function.
  • standard math Cramér’s theorem in Banach spaces and the conditional LDP continuity criterion of Chaganty (1997) hold for the empirical means of the maps G^(ℓ).
    Invoked in Lemmas 6.4–6.6 and Proposition 6.2 to obtain the layer-wise rate functions.
invented entities (2)
  • Patch-extractor operators R^(i,ℓ) and the associated maps G^(ℓ)
    purpose: Encode arbitrary convolutional receptive fields (stride, padding, pooling) so that the covariance recursion becomes a Markov chain on positive-semidefinite matrices.
    These are definitional devices that unify many practical CNN architectures; they are not new physical objects and possess no independent empirical content beyond the architectures they represent.
  • Layer-wise rate functions I_ℓ(Q_2|Q_1) defined by the Legendre transform of log M_ℓ(Q_0|Q_1)
    purpose: Provide the explicit good rate function of the joint LDP for the covariance sequence.
    Derived objects; their existence follows from the abstract large-deviation theory once the moment-generating function is shown to be finite in a neighbourhood of the origin.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large deviation principles for convolutional Bayesian neural networks." pith.science (2026). https://pith.science/paper/2ZQFQTSQ

@misc{pith2026260306023,
  author       = {Pith},
  title        = {Pith review of: Large deviation principles for convolutional Bayesian neural networks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2ZQFQTSQ}},
  note         = {Machine review of arXiv:2603.06023}
}
read the original abstract

While suitably scaled CNNs with Gaussian initialization are known to converge to Gaussian processes as the number of channels diverges, little is known beyond this Gaussian limit. We establish a large deviation principle (LDP) for convolutional neural networks in the infinite-channel regime. We consider a broad class of multidimensional CNN architectures characterized by general receptive fields encoded through a patch-extractor function satisfying mild structural assumptions. Our main result establishes a large deviation principle (LDP) for the sequence of conditional covariance matrices under Gaussian prior distribution on the weights. We further derive an LDP for the posterior distribution obtained by conditioning on a finite number of observations. In addition, we provide a streamlined proof of the concentration of the conditional covariances and of the Gaussian equivalence of the network. To the best of our knowledge, this is the first large deviation principle established for convolutional neural networks.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

24 extracted references · 3 linked inside Pith

  1. [1]

    Aiudi, R

    R. Aiudi, R. Pacelli, P. Baglioni, A. Vezzani, R. Burioni, and P. Rotondo. Local kernel renormalization as a mechanism for feature learning in overparametrized convolutional neural networks.Nat. Commun., 16(1):568, 2025

  2. [2]

    Andreis, F

    L. Andreis, F. Bassetti, and C. Hirsch. LDP for the covariance process in fully connected Gaussian neural networks.Electron. J. Probab., 31:Paper No. 22, 35, 2026

  3. [3]

    Baglioni, R

    P. Baglioni, R. Pacelli, R. Aiudi, F. Di Renzo, A. Vezzani, R. Burioni, and P. Rotondo. Predictive power of a Bayesian effective action for fully connected one hidden layer neural networks in the proportional limit.Phys. Rev. Lett., 133(2):027301, 2024

  4. [4]

    Bassetti, L

    F. Bassetti, L. Ladelli, and P. Rotondo. Proportional infinite-width infinite-depth limit for deep linear neural networks.arXiv:2411.15267, 2024

  5. [5]

    Bhatia.Matrix Analysis, volume 169 ofGraduate Texts in Mathematics

    R. Bhatia.Matrix Analysis, volume 169 ofGraduate Texts in Mathematics. Springer, 1997

  6. [6]

    Celli and G

    L. Celli and G. Peccati. Entropic bounds for conditionally gaussian vectors and applications to neural networks.arXiv:2504.08335, 2025

  7. [7]

    N. R. Chaganty. Large deviations for joint distributions and statistical applications.Sankhya A, pages 147–166, 1997

  8. [8]

    A. G. de G. Matthews, J. Hron, M. Rowland, R. E. Turner, and Z. Ghahramani. Gaussian process behaviour in wide deep neural networks. InInternational Conference on Learning Representations, 2018

Show all 24 references
  1. [9]

    Dembo and O

    A. Dembo and O. Zeitouni.Large Deviations Techniques and Applications, volume 38 of Stochastic Modelling and Applied Probability. Springer-Verlag, Berlin, 2010

  2. [10]

    Favaro, B

    S. Favaro, B. Hanin, D. Marinucci, I. Nourdin, and G. Peccati. Quantitative clts in deep neural networks.Probability Theory and Related Fields, pages 1–45, 2025

  3. [11]

    Goodfellow, Y

    I. Goodfellow, Y. Bengio, and A. Courville.Deep Learning. MIT Press, 2016.http://www. deeplearningbook.org

  4. [12]

    B. Hanin. Random neural networks in the infinite width limit as Gaussian processes.Ann. Appl. Probab., 33(6A):4798–4819, 2023

  5. [13]

    B. Hanin. Random fully connected neural networks as perturbatively solvable hierarchies.J. Mach. Learn., 25(267):1–58, 2024

  6. [14]

    J. Lee, J. Sohl-Dickstein, J. Pennington, R. Novak, S. Schoenholz, and Y. Bahri. Deep neural networks as Gaussian processes. InInternational Conference on Learning Representations, 2018

  7. [15]

    M. Li, M. Nica, and D. Roy. The neural covariance sde: Shaped infinite depth-and-width networks at initialization.Advances in Neural Information Processing Systems, 35:10795– 10808, 2022

  8. [16]

    Macci, B

    C. Macci, B. Pacchiarotti, K. Papagiannouli, G. Torrisi, and D. Trevisan. Functional large deviations for wide deep neural networks with gaussian initialization and lipschitz activations. arXiv:2601.18276, 2026

  9. [17]

    Macci, B

    C. Macci, B. Pacchiarotti, and G. L. Torrisi. Large and moderate deviations for gaussian neural networks.Journal of Applied Probability, 2025

  10. [18]

    Novak, L

    R. Novak, L. Xiao, Y. Bahri, J. Lee, G. Yang, D. A. Abolafia, J. Pennington, and J. Sohl- dickstein. Bayesian deep convolutional networks with many channels are Gaussian processes. InInternational Conference on Learning Representations, 2019

  11. [19]

    Pacelli, S

    R. Pacelli, S. Ariosto, M. Pastore, F. Ginelli, M. Gherardi, and P. Rotondo. A statistical mechanics framework for Bayesian deep neural networks beyond the infinite-width limit.Nat. Mach. Intell., 5:1497–1507, 2023

  12. [20]

    Rezakhanlou

    F. Rezakhanlou. Lectures on the large deviation principle, Mar. 2017. Lecture notes, March 25, 2017

  13. [21]

    Trevisan

    D. Trevisan. Wide deep neural networks with Gaussian weights are very close to Gaussian processes.arXiv:2312.11737, 2023

  14. [22]

    Q. Vogel. Large deviations of Gaussian neural networks with ReLU activation.Statist. Probab. Lett., 230:110611, 2026. CONVOULUTIONAL NN 23

  15. [23]

    G. Yang. Tensor programs I: Wide feedforward or recurrent neural networks of any architec- ture are gaussian processes. volume 32, 2019

  16. [24]

    Yang and E

    G. Yang and E. J. Hu. Tensor programs IV: Feature learning in infinite-width neural networks. In M. Meila and T. Zhang, editors,Proceedings of the 38th International Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, pages 11727– 11737. PMLR...

Pith tools

Reviewed July 15, 2026 · model on record in the stance chip above.