{"id":"2d2695fa-d087-4a52-9a49-dbbc5e2f5db9","arxiv_id":"2603.06023","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Convolutional Bayesian NNs with Gaussian weights satisfy an LDP for their conditional covariance matrices (and posterior) in the infinite-channel limit, with an explicit good rate function built from layer-wise cumulant generators.","lead":"The paper proves large-deviation principles for the random covariance tensors of multidimensional convolutional Bayesian neural networks as the number of channels goes to infinity. This quantifies rare-event probabilities beyond the known Gaussian-process limit and covers general receptive fields used in practice.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The reader correctly isolates (A4) as the most delicate hypothesis, yet the manuscript supplies a complete, self-contained proof chain from the patch-extractor formalism through the general covariance structure (Def. 4.2) to the full LDP (Prop. 6.10). The technical steps (Lemmas 6.3–6.5, exponential tightness via Markov inequality and Gaussian moments) are standard and appear free of circularity or hidden assumptions that would collapse the rate-function identification. Because the paper already weakens the FCNN hypotheses and the remainder of the argument is robust, no adjustment of the ACCEPT verdict is warranted.","tokens_in":27185,"tokens_out":450,"duration_ms":4242,"concrete_test":"Independently re-derive the exponential-equivalence estimate of Lemma 6.6 under the weaker growth condition of (A3) alone (i.e., without the o(|x|) remainder of (A4)); if the limsup of (1/n)log P{\\|S_n-\\tilde S_n\\|_F>\\delta} remains -\\infty for every \\delta>0, then (A4) is not load-bearing and the rate function is robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Theorem 3.3) rests on the Markov structure of the covariance sequence (Prop. 4.1), the conditional LDP continuity of the kernels \nu_{ℓ+1,n} (Prop. 6.7), and exponential tightness (Prop. 6.9). Assumption (A4) is used only to obtain the asymptotic Lipschitz bound (15) that yields exponential equivalence of S_n and \tilde S_n (Lemma 6.6). The paper already notes that (A4) is weaker than the corresponding hypotheses in the FCNN literature it cites, and the remainder of the argument (Cramér, contraction, induction) is standard and free of gaps that would invalidate the rate-function identification. No internal inconsistency or missing step that would overturn the LDP appears.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper establishes large-deviation principles for a broad class of multidimensional convolutional Bayesian neural networks in the infinite-channel regime. Under a Gaussian prior on the weights (A1), linear growth of the channel counts (A2), and an asymptotic Lipschitz condition on the activation and patch extractors (A4), the sequence of random conditional covariance tensors (K^(2,n),…,K^(L+1,n)) satisfies an LDP on the product of positive-semidefinite cones with good rate function I_{2,…,L+1}(Q_2,…,Q_{L+1})=α_1 I_1(Q_2|K^(1))+∑_{ℓ=2}^L α_ℓ I_ℓ(Q_{ℓ+1}|Q_ℓ), where each I_ℓ is the Legendre transform of the log-moment generating function of the patch-extractor map G^(ℓ) (Theorem 3.3). The same rate governs the posterior of the final covariance after conditioning on finitely many Gaussian observations (Proposition 3.5), and a rescaled network-output LDP follows by contraction (Proposition 3.6). As corollaries under the weaker growth condition (A3) the authors recover covariance concentration (Theorem 3.1) and Gaussian-process convergence (Theorem 3.2). The argument proceeds by identifying the Markov structure of the covariances (Proposition 4.1), verifying conditional LDP continuity of the transition kernels via Cramér plus exponential equivalence (Lemmas 6.3–6.6, Proposition 6.7), establishing joint exponential tightness (Proposition 6.9), and inducting to the full LDP (Proposition 6.10).","tokens_in":27432,"tokens_out":1055,"duration_ms":9421,"significance":"This appears to be the first large-deviation principle for convolutional architectures. It extends the known Gaussian-process limit of wide CNNs (Novak et al., Yang et al.) to a full large-deviation description of the covariance process, covers general receptive fields via a mild patch-extractor formalism, and simultaneously yields a streamlined proof of concentration and Gaussian equivalence. The rate function is derived from first principles as an explicit Legendre transform; no free parameters are fitted. The technical pipeline (Markov structure, conditional LDP continuity, exponential tightness, induction) is standard and carefully executed, and Assumption (A4) is weaker than the corresponding hypotheses used for fully-connected networks in the cited literature. The results therefore constitute a genuine and useful advance for the probabilistic theory of Bayesian CNNs.","major_comments":[],"minor_comments":[{"comment":"Throughout the manuscript the word “convolutional” is misspelled as “CONVOULUTIONAL” in several running headers and page titles (e.g., pages 3, 5, 7, …). A global search-and-replace is needed.","section":null},{"comment":"In the definition of the patch extractor for the 2-D zero-padding example (Eq. (4) and the surrounding display), the ordering of the nine mask offsets is listed inconsistently with the set M defined in (3); while the mathematics is unaffected, a uniform ordering would improve readability.","section":null},{"comment":"Lemma 3.4 and Proposition 3.5 are stated without proofs, the authors referring the reader to the analogous FCNN arguments in [2]. A short sketch of the Gaussian-manipulation steps that produce the posterior density of K^(L+1,n) would make the paper more self-contained.","section":null},{"comment":"The notation for the flattened versus tensor forms of G^(ℓ) and K^(ℓ) is introduced in Remark 2 but then used interchangeably; a single sentence reminding the reader that all matrix norms and traces are understood after the standard (i,μ)\to k(i,μ) flattening would avoid occasional ambiguity.","section":null},{"comment":"In Assumption (A2) the phrase “increase linearly as a function of n” is slightly informal; writing C_ℓ(n)/n\toα_ℓ∈(0,∞) already appears later and could be placed in the assumption itself for precision.","section":null}],"recommendation":"accept","confidential_remarks":"The manuscript is a clean, technically solid contribution that fits well in a probability journal. The self-citations to the authors’ earlier FCNN LDP papers are appropriate and supply independent technical lemmas; I see no novelty or citation issues. I recommend acceptance with only the listed minor corrections."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This paper gives the first large-deviation principle for the covariance sequence of convolutional Bayesian nets in the infinite-channel limit. That is the real news: the GP limit for CNNs was already known, but nothing beyond it. They cover a broad class of multi-dimensional architectures by encoding receptive fields with a mild patch-extractor map, then prove an LDP for the random covariances under Gaussian weights, plus the induced posterior LDP and a rescaled-output LDP. As a bonus they give a short proof of covariance concentration and Gaussian equivalence that improves on the earlier CNN literature.\n\nThe technical route is standard and carefully written: Markov structure of the layer-wise covariances, conditional LDP continuity via Cramér plus exponential equivalence, exponential tightness, then induction. The rate function is the expected Legendre transform of the log-moment generating function of the patch map; no free parameters, no circularity. Assumption (A4) (asymptotic Lipschitz on activation and extractors) is the load-bearing one for the equivalence step, but it is weaker than the corresponding hypotheses in the FCNN LDP papers they cite, and the rest of the argument does not depend on anything exotic.\n\nSoft spots are minor and proportional. Depth, spatial sizes and number of observations stay fixed while only channels grow; that is the usual infinite-width regime, not a flaw. The significance is real inside asymptotic NN theory (rare-event rates, posterior concentration) but does not reorganize the broader field. Citations are appropriate; the self-cites to their own FCNN LDP work supply reusable lemmas rather than padding.\n\nThis is for people who already care about LDPs or Gaussian-process limits for neural nets. The math is self-contained and free of gaps that would invalidate the main theorem. I would send it to a serious referee without hesitation; it is a clean theoretical contribution that fills a concrete hole.","headline":"First clean LDP for infinite-channel CNNs via general patch extractors; solid extension of known FCNN results with no real holes.","tokens_in":28011,"tokens_out":476,"would_cite":true,"duration_ms":8964,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60F10","60B20","68T07"],"pacs":[],"model":"grok-4.5","headline":"The first large-deviation principle for convolutional Bayesian networks shows how their random covariances deviate from the Gaussian-process limit as channels grow.","keywords":["large deviations","convolutional neural networks","Bayesian neural networks","infinite-channel limit","Gaussian processes","covariance concentration","patch extractors"],"falsifier":"Construct a continuous activation and patch extractor that violate the little-o remainder condition, compute the empirical moment-generating function of the resulting G^(ℓ) for large but finite channel counts, and check whether the observed large-deviation rate still coincides with the Legendre transform claimed in Theorem 3.3.","tokens_in":28100,"feed_emoji":"📐","tokens_out":555,"duration_ms":4866,"temperature":0.7,"pith_summary":"Wide convolutional networks with Gaussian weights are already known to look like Gaussian processes once the number of channels becomes large. This paper asks what happens just beyond that limit: how rare are the atypical covariance configurations that survive after the law of large numbers has kicked in? The authors encode a broad family of multi-dimensional CNN architectures by a general patch-extractor map and prove that the sequence of random conditional covariance matrices obeys a large-deviation principle whose rate function is an explicit Legendre transform built from those extractors. The same rate governs the posterior after finitely many observations, and a by-product is a short proof that the covariances concentrate and the network becomes Gaussian. A sympathetic reader cares because the result gives the first precise exponential-cost description of atypical behaviour for convolutional architectures, not only fully connected ones.","feed_headline":"First large-deviation principle for convolutional nets","feed_subtitle":"Random covariances obey an explicit exponential cost as channels grow to infinity","key_machinery":"The Markov chain of conditional covariance matrices whose transition kernels are empirical averages of the continuous patch-extractor map G^(ℓ); its large-deviation rate is obtained by verifying the conditional LDP continuity condition and then upgrading the resulting weak LDP by exponential tightness.","core_discovery":"Under a Gaussian prior, linear channel growth, and an asymptotic Lipschitz condition on the activation and patch extractors, the joint law of the random covariance tensors of a multi-layer CNN satisfies a large-deviation principle with good rate function equal to a weighted sum of layer-wise Legendre transforms of the log-moment generating functions of the patch-extractor maps.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["CNN covariances obey LDP beyond Gaussian limit as channels grow","Large deviations for random CNN covariances under Gaussian priors","Explicit rate functions guide multi-layer CNN paths at infinite width","Conditional CNN covariances satisfy layer-wise large deviation principle","First LDP for convolutional nets: covariances pay exponential cost"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The activation and every receptive-field map must be almost Lipschitz at infinity (their difference is controlled by a linear term plus a remainder that grows slower than the size of the input); without that control the exponential-equivalence step that identifies the rate function can fail.","fun_headline_variants_meta":{"raw":{"variants":["CNN covariances obey LDP beyond Gaussian limit as channels grow","Large deviations for random CNN covariances under Gaussian priors","Explicit rate functions guide multi-layer CNN paths at infinite width","Conditional CNN covariances satisfy layer-wise large deviation principle","First LDP for convolutional nets: covariances pay exponential cost"]},"model":"grok-4.5","effort":"low","cost_usd":0.005942,"raw_usage":{"total_tokens":1505,"prompt_tokens":676,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":59420000,"prompt_tokens_details":{"text_tokens":676,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":744,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":676,"tokens_out":85,"duration_ms":5859,"temperature":1.0,"reasoning_tokens":744,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-15T14:02:52.238034+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Construct a continuous activation and patch extractor that violate the little-o remainder condition, compute the empirical moment-generating function of the resulting G^(ℓ) for large but finite channel counts, and check whether the observed large-deviation rate still coincides with the Legendre transform claimed in Theorem 3.3.","supporting_citations":[],"review_version":1}