REVIEW 5 minor 24 references
Large deviation principles for convolutional Bayesian neural networks
T0 review · 0 major / 5 minor · reviewed 2026-07-15 · grok-4.5
Pith's one-line read The first large-deviation principle for convolutional Bayesian networks shows how their random covariances deviate from the Gaussian-process limit as channels grow.
desk verdict First clean LDP for infinite-channel CNNs via general patch extractors; solid extension of known FCNN results with no real holes. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Markov chain of conditional covariance matrices whose transition kernels are empirical averages of the continuous patch-extractor map G^(ℓ); its large-deviation rate is obtained by verifying the conditional LDP continuity condition and then upgrading the resulting weak LDP by exponential tightness.
What would settle it
Construct a continuous activation and patch extractor that violate the little-o remainder condition, compute the empirical moment-generating function of the resulting G^(ℓ) for large but finite channel counts, and check whether the observed large-deviation rate still coincides with the Legendre transform claimed in Theorem 3.3.
Extended reading notes
Core claim
Under a Gaussian prior, linear channel growth, and an asymptotic Lipschitz condition on the activation and patch extractors, the joint law of the random covariance tensors of a multi-layer CNN satisfies a large-deviation principle with good rate function equal to a weighted sum of layer-wise Legendre transforms of the log-moment generating functions of the patch-extractor maps.
Load-bearing premise
The activation and every receptive-field map must be almost Lipschitz at infinity (their difference is controlled by a linear term plus a remainder that grows slower than the size of the input); without that control the exponential-equivalence step that identifies the rate function can fail.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper establishes large-deviation principles for a broad class of multidimensional convolutional Bayesian neural networks in the infinite-channel regime. Under a Gaussian prior on the weights (A1), linear growth of the channel counts (A2), and an asymptotic Lipschitz condition on the activation and patch extractors (A4), the sequence of random conditional covariance tensors (K^(2,n),…,K^(L+1,n)) satisfies an LDP on the product of positive-semidefinite cones with good rate function I_{2,…,L+1}(Q_2,…,Q_{L+1})=α_1 I_1(Q_2|K^(1))+∑_{ℓ=2}^L α_ℓ I_ℓ(Q_{ℓ+1}|Q_ℓ), where each I_ℓ is the Legendre transform of the log-moment generating function of the patch-extractor map G^(ℓ) (Theorem 3.3). The same rate governs the posterior of the final covariance after conditioning on finitely many Gaussian observations (Proposition 3.5), and a rescaled network-output LDP follows by contraction (Proposition 3.6). As corollaries under the weaker growth condition (A3) the authors recover covariance concentration (Theorem 3.1) and Gaussian-process convergence (Theorem 3.2). The argument proceeds by identifying the Markov structure of the covariances (Proposition 4.1), verifying conditional LDP continuity of the transition kernels via Cramér plus exponential equivalence (Lemmas 6.3–6.6, Proposition 6.7), establishing joint exponential tightness (Proposition 6.9), and inducting to the full LDP (Proposition 6.10).
Significance. This appears to be the first large-deviation principle for convolutional architectures. It extends the known Gaussian-process limit of wide CNNs (Novak et al., Yang et al.) to a full large-deviation description of the covariance process, covers general receptive fields via a mild patch-extractor formalism, and simultaneously yields a streamlined proof of concentration and Gaussian equivalence. The rate function is derived from first principles as an explicit Legendre transform; no free parameters are fitted. The technical pipeline (Markov structure, conditional LDP continuity, exponential tightness, induction) is standard and carefully executed, and Assumption (A4) is weaker than the corresponding hypotheses used for fully-connected networks in the cited literature. The results therefore constitute a genuine and useful advance for the probabilistic theory of Bayesian CNNs.
minor comments (5)
- Throughout the manuscript the word “convolutional” is misspelled as “CONVOULUTIONAL” in several running headers and page titles (e.g., pages 3, 5, 7, …). A global search-and-replace is needed.
- In the definition of the patch extractor for the 2-D zero-padding example (Eq. (4) and the surrounding display), the ordering of the nine mask offsets is listed inconsistently with the set M defined in (3); while the mathematics is unaffected, a uniform ordering would improve readability.
- Lemma 3.4 and Proposition 3.5 are stated without proofs, the authors referring the reader to the analogous FCNN arguments in [2]. A short sketch of the Gaussian-manipulation steps that produce the posterior density of K^(L+1,n) would make the paper more self-contained.
- The notation for the flattened versus tensor forms of G^(ℓ) and K^(ℓ) is introduced in Remark 2 but then used interchangeably; a single sentence reminding the reader that all matrix norms and traces are understood after the standard (i,μ) o k(i,μ) flattening would avoid occasional ambiguity.
- In Assumption (A2) the phrase “increase linearly as a function of n” is slightly informal; writing C_ℓ(n)/n oα_ℓ∈(0,∞) already appears later and could be placed in the assumption itself for precision.
Circularity Check
No significant circularity: the LDP rate function is the Legendre transform of an explicit log-mgf of the patch-extractor maps, obtained via Cramér + exponential equivalence + conditional LDP continuity + tightness, all proved in-place.
full rationale
The derivation chain for the central claim (Theorem 3.3) is self-contained. The covariance sequence is a Markov chain with kernels ν_{ℓ+1,n} built from the explicit maps G^{(ℓ)} (Definition 4.2, Proposition 4.1). Under (A6) the kernels satisfy the conditional LDP continuity condition of Chaganty with rate α_ℓ I_ℓ (Proposition 6.7), where I_ℓ is the Legendre transform of the log-moment generating function M_ℓ of G^{(ℓ)} (eqs. (9),(19)). Exponential equivalence of the empirical averages S_n and ˜S_n follows from the asymptotic Lipschitz bound (Lemma 6.3–6.6); exponential tightness is proved by induction with Markov’s inequality (Proposition 6.9). The full LDP is then obtained by induction + upgrading the weak LDP via tightness (Proposition 6.10). Lemma 4.3 verifies that the CNN patch-extractor maps satisfy (A5)–(A6) under the paper’s (A3)–(A4). No parameters are fitted, no target quantity is inserted by definition, and the self-citations to FCNN LDP papers ([2],[16],[17],[22]) supply only reusable technical lemmas or omitted routine arguments for the posterior (Proposition 3.5) and rescaled-output (Proposition 3.6) corollaries; they do not encode the CNN rate function. The result is therefore not circular.
Assumptions & free parameters
assumptions (4)
- domain assumption Weights W^(ℓ)_{m,c',c} are i.i.d. N(0,λ_ℓ^{-1}) (Assumption A1).
- domain assumption Channel counts satisfy C_ℓ(n)/n o α_ℓ ∈ (0,∞) while depth L, spatial sizes N_ℓ and number of inputs P remain fixed (Assumption A2).
- domain assumption Activation σ and patch extractors R^(i,ℓ) are continuous and satisfy the exponential growth bound of degree r_σ < 2 (Assumption A3) or the stronger asymptotic Lipschitz condition (Assumption A4).
- standard math Cramér’s theorem in Banach spaces and the conditional LDP continuity criterion of Chaganty (1997) hold for the empirical means of the maps G^(ℓ).
invented entities (2)
-
Patch-extractor operators R^(i,ℓ) and the associated maps G^(ℓ)
-
Layer-wise rate functions I_ℓ(Q_2|Q_1) defined by the Legendre transform of log M_ℓ(Q_0|Q_1)
Cite this review
Pith. "Pith review of Large deviation principles for convolutional Bayesian neural networks." pith.science (2026). https://pith.science/paper/2ZQFQTSQ
@misc{pith2026260306023,
author = {Pith},
title = {Pith review of: Large deviation principles for convolutional Bayesian neural networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/2ZQFQTSQ}},
note = {Machine review of arXiv:2603.06023}
}
read the original abstract
While suitably scaled CNNs with Gaussian initialization are known to converge to Gaussian processes as the number of channels diverges, little is known beyond this Gaussian limit. We establish a large deviation principle (LDP) for convolutional neural networks in the infinite-channel regime. We consider a broad class of multidimensional CNN architectures characterized by general receptive fields encoded through a patch-extractor function satisfying mild structural assumptions. Our main result establishes a large deviation principle (LDP) for the sequence of conditional covariance matrices under Gaussian prior distribution on the weights. We further derive an LDP for the posterior distribution obtained by conditioning on a finite number of observations. In addition, we provide a streamlined proof of the concentration of the conditional covariances and of the Gaussian equivalence of the network. To the best of our knowledge, this is the first large deviation principle established for convolutional neural networks.
Reference graph
Works this paper leans on
-
[1]
Aiudi, R
R. Aiudi, R. Pacelli, P. Baglioni, A. Vezzani, R. Burioni, and P. Rotondo. Local kernel renormalization as a mechanism for feature learning in overparametrized convolutional neural networks.Nat. Commun., 16(1):568, 2025
2025
-
[2]
Andreis, F
L. Andreis, F. Bassetti, and C. Hirsch. LDP for the covariance process in fully connected Gaussian neural networks.Electron. J. Probab., 31:Paper No. 22, 35, 2026
2026
-
[3]
Baglioni, R
P. Baglioni, R. Pacelli, R. Aiudi, F. Di Renzo, A. Vezzani, R. Burioni, and P. Rotondo. Predictive power of a Bayesian effective action for fully connected one hidden layer neural networks in the proportional limit.Phys. Rev. Lett., 133(2):027301, 2024
2024
-
[4]
F. Bassetti, L. Ladelli, and P. Rotondo. Proportional infinite-width infinite-depth limit for deep linear neural networks.arXiv:2411.15267, 2024
arXiv 2024
-
[5]
Bhatia.Matrix Analysis, volume 169 ofGraduate Texts in Mathematics
R. Bhatia.Matrix Analysis, volume 169 ofGraduate Texts in Mathematics. Springer, 1997
1997
-
[6]
L. Celli and G. Peccati. Entropic bounds for conditionally gaussian vectors and applications to neural networks.arXiv:2504.08335, 2025
arXiv 2025
-
[7]
N. R. Chaganty. Large deviations for joint distributions and statistical applications.Sankhya A, pages 147–166, 1997
1997
-
[8]
A. G. de G. Matthews, J. Hron, M. Rowland, R. E. Turner, and Z. Ghahramani. Gaussian process behaviour in wide deep neural networks. InInternational Conference on Learning Representations, 2018
2018
Show all 24 references
-
[9]
Dembo and O
A. Dembo and O. Zeitouni.Large Deviations Techniques and Applications, volume 38 of Stochastic Modelling and Applied Probability. Springer-Verlag, Berlin, 2010
2010
-
[10]
Favaro, B
S. Favaro, B. Hanin, D. Marinucci, I. Nourdin, and G. Peccati. Quantitative clts in deep neural networks.Probability Theory and Related Fields, pages 1–45, 2025
2025
-
[11]
Goodfellow, Y
I. Goodfellow, Y. Bengio, and A. Courville.Deep Learning. MIT Press, 2016.http://www. deeplearningbook.org
2016
-
[12]
B. Hanin. Random neural networks in the infinite width limit as Gaussian processes.Ann. Appl. Probab., 33(6A):4798–4819, 2023
2023
-
[13]
B. Hanin. Random fully connected neural networks as perturbatively solvable hierarchies.J. Mach. Learn., 25(267):1–58, 2024
2024
-
[14]
J. Lee, J. Sohl-Dickstein, J. Pennington, R. Novak, S. Schoenholz, and Y. Bahri. Deep neural networks as Gaussian processes. InInternational Conference on Learning Representations, 2018
2018
-
[15]
M. Li, M. Nica, and D. Roy. The neural covariance sde: Shaped infinite depth-and-width networks at initialization.Advances in Neural Information Processing Systems, 35:10795– 10808, 2022
2022
-
[16]
Macci, B
C. Macci, B. Pacchiarotti, K. Papagiannouli, G. Torrisi, and D. Trevisan. Functional large deviations for wide deep neural networks with gaussian initialization and lipschitz activations. arXiv:2601.18276, 2026
2026
-
[17]
Macci, B
C. Macci, B. Pacchiarotti, and G. L. Torrisi. Large and moderate deviations for gaussian neural networks.Journal of Applied Probability, 2025
2025
-
[18]
Novak, L
R. Novak, L. Xiao, Y. Bahri, J. Lee, G. Yang, D. A. Abolafia, J. Pennington, and J. Sohl- dickstein. Bayesian deep convolutional networks with many channels are Gaussian processes. InInternational Conference on Learning Representations, 2019
2019
-
[19]
Pacelli, S
R. Pacelli, S. Ariosto, M. Pastore, F. Ginelli, M. Gherardi, and P. Rotondo. A statistical mechanics framework for Bayesian deep neural networks beyond the infinite-width limit.Nat. Mach. Intell., 5:1497–1507, 2023
2023
-
[20]
Rezakhanlou
F. Rezakhanlou. Lectures on the large deviation principle, Mar. 2017. Lecture notes, March 25, 2017
2017
-
[21]
Trevisan
D. Trevisan. Wide deep neural networks with Gaussian weights are very close to Gaussian processes.arXiv:2312.11737, 2023
2023 arXiv
-
[22]
Q. Vogel. Large deviations of Gaussian neural networks with ReLU activation.Statist. Probab. Lett., 230:110611, 2026. CONVOULUTIONAL NN 23
2026
-
[23]
G. Yang. Tensor programs I: Wide feedforward or recurrent neural networks of any architec- ture are gaussian processes. volume 32, 2019
2019
-
[24]
Yang and E
G. Yang and E. J. Hu. Tensor programs IV: Feature learning in infinite-width neural networks. In M. Meila and T. Zhang, editors,Proceedings of the 38th International Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, pages 11727– 11737. PMLR...
2021
Reviewed July 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.