{"id":"3fb16b7d-f143-4757-b644-97816e8bda1d","arxiv_id":"2412.15911","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Output-output correlations in finite Bayesian one-hidden-layer networks follow the kernel shape renormalization order parameter, with readout weight overlap equal to Q*_ab/λ1.","lead":"Finite Bayesian one-hidden-layer networks with several outputs show readout correlations that vanish in the infinite-width limit, and this paper predicts them from a renormalized kernel shape in the proportional limit. Monte Carlo experiments on synthetic data, MNIST, and CIFAR10 match the predictions, giving a concrete mechanism for finite-width feature learning.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (4) inherits SM Eq. (11), a Gaussian equivalence stated in the P,N1→∞ limit; but the quoted CLT also requires the input dimension N0 to diverge. For fixed N0, the h-covariance C has rank ≤N0, so the per-hidden-unit q variables cannot become Gaussian.","rationale":"The reader's weakest assumption is exactly the Gaussian equivalence in SM Eq. (11), including the note that finite-N0 violations are not bounded. My stress-test sharpens this into a concrete failure mode: the CLT is indexed by the sample count P, but the underlying Gaussian field h has rank at most N0. When N0 is fixed while P grows, q_a depends on only N0 Gaussian variables, so the distribution cannot converge to a Gaussian. This is not a matter of finite-size corrections; it is a mismatch between the stated limit and the theorem being invoked. The paper's own supplementary investigation (Fig. 4c) shows discrepancies shrinking as N0 increases, which is the expected signature of this missing limit. Because the action (1) and therefore Q* are built on Eq. (11), the central correlation formula (4) inherits the issue. This does not overturn the paper: the numerical agreement at N0=784 suggests the approximation may be quantitatively useful, and the reader's CONDITIONAL verdict already accounts for this. I would keep the conditional status, with the additional explicit requirement that the N0 dependence be stated and tested.","tokens_in":18245,"tokens_out":16236,"duration_ms":157297,"concrete_test":"Run the Langevin protocol of Fig. 2 on synthetic Gaussian data with P=N1=1000, α=1, λ0=λ1=1, T=0.01, D=10 and vary N0 over {20,50,100,200,500,1000}; plot the maximum off-diagonal relative error |⟨v_a·v_b⟩/N1 − Q*_ab/λ1|/max|Q*_ab/λ1| versus N0. If the error does not decrease toward zero as N0 grows, or remains O(1) when P,N1 are increased at fixed small N0, then the Gaussian equivalence in SM Eq. (11) is not valid in the stated proportional limit.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The weakest point is the Gaussian equivalence used to obtain the effective action (1) and hence Eq. (4). SM Eq. (10)-(12) defines q_a = (1/sqrt{N1 λ1}) Σ_μ \\bar{s}_{μ,a} σ(h_μ), with h ∼ N(0,C), C_{μν}=x_μ·x_ν/(λ0 N0). Eq. (11) asserts P({q}) → Gaussian when P,N1→∞ at fixed α=P/N1, citing Breuer-Major CLTs. However, for fixed input dimension N0, C has rank at most N0, so h lies in an N0-dimensional subspace; q_a is a function of only N0 independent Gaussian coordinates. The sum has P terms but no CLT over μ applies, and Eq. (12)'s det C and C^{-1} are not even defined when N0<P+1, which is the case in the experiments (N0=122,144,784 with P=1000). Breuer-Major requires the Gaussian field to have a decaying covariance structure; a fixed-rank C does not satisfy it. A valid statement would require N0→∞ jointly with P,N1, with a rate condition and an error bound. The paper's proportional-limit claim specifies only P,N1→∞ and 'α=P/N1 fixed,' omitting N0. The SM Fig. 4(c) itself shows that the theory-data discrepancy decreases as N0 increases at fixed P,N1, consistent with N0 being an essential control parameter. Because Q* in Eq. (4) is the minimizer of the same action built on Eq. (11), the numerical agreement in Fig. 2 is currently evidence for an N0-dependent approximation, not for the stated limit.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies finite-width Bayesian one-hidden-layer networks with multiple outputs in the proportional limit P,N1→∞ at fixed α=P/N1. Starting from an effective action with a D×D order parameter Q (Eq. (1)), the authors obtain a renormalized NNGP kernel K_Q = Q ⊗ K_NNGP and derive closed-form expressions for the posterior predictive mean and covariance (SM Eqs. (17)–(18)). The central result is Eq. (4), which identifies the normalized readout-overlap matrix ⟨vv^T⟩/N1 with Q*/λ1, where Q* minimizes the effective action. The paper validates this prediction against Langevin-dynamics simulations on synthetic data, MNIST, and CIFAR10 for several values of α and D, and additionally compares infinite-width Bayesian predictions with Adam-trained networks. The numerical agreement reported in Figs. 1, 2, and SM Fig. 5 is the main empirical support for the claim that kernel shape renormalization quantitatively explains output-output correlations.","tokens_in":18659,"tokens_out":4851,"duration_ms":50161,"significance":"If the central derivation is sound, Eq. (4) is a valuable quantitative bridge between statistical-mechanics order parameters and observable weight correlations in finite Bayesian networks. The paper provides a simple physical interpretation of negative off-diagonal overlaps (e.g., confusable CIFAR10 classes) and extends the recent proportional-limit program for Bayesian deep learning to multi-output architectures. The Monte Carlo validation is carried out carefully, with blocking error estimates, thermalization checks, and multiple datasets, and the comparison with Adam is a useful practical benchmark. However, the theoretical result rests on a Gaussian equivalence whose asymptotic regime is not fully specified; in particular, the role of the input dimension N0 is not included in the stated proportional limit, and the manuscript itself provides evidence (SM Fig. 4(c)) that N0 is a controlling parameter. Until this gap is closed or explicitly acknowledged as a controlled approximation, the quantitative agreement in Figs. 2 and 5 is evidence for a useful heuristic but not for the theorem stated in the text.","major_comments":[{"comment":"The Gaussian equivalence asserted in SM Eq. (11) is not justified in the stated limit. The variables h_μ have covariance C, and for fixed input dimension N0 this covariance has rank at most N0. When N0 < P+1, which is the case in the experiments (P=1000 with N0=784, 144, or 122), det C and C^{-1} in SM Eq. (12) are undefined, and the density entering the Gaussian expectation is not well-defined. The cited Breuer-Major type theorems do not apply to sums over a fixed-rank Gaussian vector. The main text defines the proportional limit as P,N1→∞ at fixed α, without any growth condition on N0; consequently the effective action (1) and the key result (4) are not derived in the stated limit. Please specify a joint limit N0,P,N1→∞ with a rate condition and provide error bounds, or explicitly reformulate the results as a finite-N0 heuristic approximation whose accuracy must be tested separately.","section":"SM Eqs. (10)-(12) and main-text proportional limit"},{"comment":"SM Fig. 4(c) reports that the relative discrepancy Δ between theory and simulation decreases when N0 increases at fixed P and N1. This is direct evidence that N0 is an essential control parameter of the Gaussian equivalence, and it corroborates the concern raised above. The main text currently presents the agreement in Figs. 1 and 2 as supporting the stated proportional-limit theory, but all experiments use a single fixed N0 per dataset. To make the claim quantitative, the manuscript should include a systematic study varying N0 (e.g., through random projections of MNIST at several N0 values) showing convergence toward the predictions as N0 grows, or explicitly limit the claims to a heuristic regime. As written, the experimental section does not test the limit that the theory invokes.","section":"SM Fig. 4(c) and Fig. 1(b)"},{"comment":"The overlap formula ⟨va·vb⟩/N1 = Q*_ab/λ1 is derived from the modified partition function in SM Eq. (22), whose saddle-point reduction again relies on the same Gaussian equivalence used for Eq. (1). Thus the central prediction (main-text Eq. (4)) inherits all of the limitations of SM Eq. (11). The paper should state this dependency explicitly; it is not merely a numerical nuisance but a logical link in the derivation.","section":"SM Sec. II, derivation of Eq. (29)"}],"minor_comments":[{"comment":"The caption says 'N0 = 122' for the random Gaussian data, but SM Sec. III A specifies N0 = 12^2 = 144 for the synthetic dataset. Please correct the typo and make the notation consistent throughout.","section":"Fig. 1 caption"},{"comment":"The dataset name 'CIF AR10' and 'Imagenette' appear with inconsistent spacing; also, Ref. [43] ('A smaller subset of 10 easily classified classes from imagenet') lacks author and year information. These should be cleaned up for a final version.","section":"References and text"},{"comment":"The perturbative expansion Q = Q(0) + α1 Q(1) + ... is used to produce the dashed 'heuristic proportional th.' curves in Fig. 1, but the expansion parameter α1 is not small in the presented experiments (α up to 1 for N1=1000). The text should note that this is a formal expansion and that its range of validity is not assessed.","section":"SM Sec. IV, Eq. (41)"}],"recommendation":"major_revision","confidential_remarks":"The paper is a natural continuation of the authors' prior work (Refs. [26, 29, 53]) and cites it prominently, but the new numerical experiments are independent of that. The main issue is not circularity; it is the missing mathematical control of the Gaussian equivalence when N0 is fixed. If the authors can either prove a joint-limit statement with N0→∞ or clearly reposition the paper as an approximate/empirical theory with a documented range of validity, the result would be publishable. I would not reject outright because the central formula is clean, the numerics are careful, and the concern is a gap in the asymptotic justification rather than a demonstrated internal contradiction."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things before reading this one. The main formula, <v_a·v_b>/N1 = Q*_ab/λ1, is not new in this paper—the SM points back to the authors' own COMPENG paper [53]—and the effective action it is built on comes from [26]. What this paper adds is a careful multi-output version of the story: closed-form generalization loss, the physical reading of off-diagonal Q as weight overlaps, and extensive Langevin validation across synthetic data, MNIST, and CIFAR10, including α=0.1 and α=1, with proper error bars. That validation is the real value. The numerics match, and the Adam comparison in Fig. 3 is honest: Bayesian training is at least competitive on these FC models, which is a useful data point.\n\nThe weak spot is the Gaussian equivalence that underlies everything. SM Eq. (11) is invoked as a Breuer-Major type CLT in the limit P,N1→∞ at fixed α, but the paper never includes the input dimension N0 in that limit. For fixed N0, C has rank at most N0, so when N0<P+1—which is the case in the experiments (N0=784 or 122, P=1000)—Eq. (12) is formal, with det C and C^{-1} not defined in the usual sense. More importantly, the q variables are mixtures over the hidden weights; their conditional variance over the data may not self-average unless N0 also diverges. The SM's own Fig. 4(c) shows the theory-data discrepancy shrinking as N0 increases at fixed N1, which is evidence that N0 is a control parameter the stated limit is missing. This does not sink the paper: the numerical agreement is suggestive that the approximation is good for these N0 values, and the heuristic is standard in this literature. But the paper should state the N0→∞ condition, discuss why Breuer-Major applies (or why the approximation is valid at finite N0), and ideally give a bound or at least a scaling argument.\n\nMinor: no code or data artifact, which makes the quantitative match harder to audit, though the Langevin details are specific enough to replicate. Self-citation is not a problem here; the prior results are real.\n\nBottom line: this is a paper for readers who work on finite-width Bayesian nets and the proportional limit, and it deserves a serious referee. I would ask the authors to clarify the N0 issue and release code/data. I would probably cite it for the validation, not for the formula.","headline":"A well-tested application of an existing proportional-limit effective action; the headline correlation formula was already in the authors' own conference paper, but the validation and generalization-loss analysis make it a useful reference.","tokens_in":19178,"tokens_out":7273,"would_cite":true,"duration_ms":73598,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single matrix Q* predicts how outputs correlate in finite Bayesian networks.","keywords":["kernel shape renormalization","output-output correlations","proportional limit","Bayesian neural networks","one-hidden-layer networks","NNGP kernel","readout weight overlaps","generalization loss"],"falsifier":"Train a fully thermalized one-hidden-layer Bayesian network with $D=3$ outputs on synthetic Gaussian inputs where two classes are built to share a hidden feature and the third is independent, with $P=N_1=1000$, and measure $\\langle v_a \\cdot v_b\\rangle/N_1$ from Langevin samples; Eq. (4) predicts this equals $Q^*_{ab}/\\lambda_1$ with $Q^*$ from minimizing Eq. (1), up to thermodynamic-limit corrections. If the off-diagonal overlap stays near zero, or has the wrong sign, while $\\alpha$ is held fixed and the sampler has converged, the Gaussian equivalence behind $Q^*$ fails; alternatively, checking that the discrepancy shrinks as $N_1$ grows at fixed $\\alpha$ tests the predicted scaling.","tokens_in":18078,"feed_emoji":"🧠","tokens_out":12370,"duration_ms":94424,"temperature":0.7,"pith_summary":"Finite-width Bayesian networks with several outputs show correlations between outputs that the infinite-width description cannot produce; this paper claims those correlations are a genuine finite-load effect, controlled by the same object that renormalizes the kernel: a $D \\times D$ matrix $Q^*$ that minimizes a free-energy action in the proportional limit. The central formula is $\\langle v_a \\cdot v_b\\rangle/N_1 = Q^*_{ab}/\\lambda_1$, which says the expected overlap of readout weight vectors for outputs $a$ and $b$ is exactly the corresponding entry of the saddle-point matrix divided by the readout prior scale. If the claim is right, a practitioner can compute from the data and the architecture how strongly the outputs of a finite one-hidden-layer Bayesian network are coupled and can watch those correlations vanish as the width grows, since $Q^* \\to 1$ in the infinite-width limit. The paper also shows that the same effective theory predicts the generalization loss and that Bayesian training matches gradient-descent training in generalization on standard image datasets, so the correlations are not a sampling artifact.","feed_headline":"A single matrix predicts output-output correlations in finite nets","feed_subtitle":"In the proportional limit, readout weight overlaps become a testable prediction of shared features.","key_machinery":"The central object is the $D \\times D$ order-parameter matrix $Q$ in the effective action $S_{\\mathrm{FC}}(Q)$, together with the renormalized kernel $K^{(\\mathrm{R})}_Q = Q \\otimes K_{\\mathrm{NNGP}}$ formed by its tensor product with the infinite-width NNGP kernel. $Q$ is integrated out by a saddle point, so its minimizing value $Q^*$ enters every prediction; it renormalizes the shape of the kernel by mixing the $D$ outputs, something a scalar renormalization cannot do. The load-bearing identity is $\\langle v v^T\\rangle/N_1 = Q^*/\\lambda_1$, obtained by differentiating the partition function with respect to a matrix-valued readout prior; it turns an abstract saddle-point matrix into a direct observable, the overlap of readout weight vectors, and provides the mechanism by which shared hidden features become output-output correlations.","core_discovery":"The paper establishes, in the proportional thermodynamic limit where $P$ and $N_1$ grow together at fixed $\\alpha = P/N_1$, that the posterior statistics of a fully connected one-hidden-layer Bayesian network with $D$ outputs are governed by a renormalized kernel $K^{(\\mathrm{R})}_Q(X,X) = Q \\otimes K_{\\mathrm{NNGP}}(X,X)$, with $Q$ a $D \\times D$ matrix order parameter and $K_{\\mathrm{NNGP}}$ the infinite-width Gaussian process kernel. The minimizing value $Q^*$ of the effective action in Eq. (1) simultaneously fixes the generalization loss and, through Eq. (4), the output-output correlations: after training, the normalized overlap of the readout weight vectors $v_a$ and $v_b$ equals $Q^*_{ab}/\\lambda_1$. Off-diagonal entries of $Q^*$ are therefore not fitting noise but a data-dependent prediction: at finite $\\alpha$, outputs whose classes share hidden features acquire systematically negative overlaps (for example, vehicle classes in CIFAR10), and in the infinite-width limit $\\alpha \\to 0$, $Q^* \\to 1$ and the correlations vanish, reproducing the classic result that $D$ outputs decouple into $D$ independent networks. The paper validates Eq. (4) by Langevin sampling on synthetic, MNIST, and CIFAR10 data at $\\alpha = 0.1$ and $\\alpha = 1$, finding quantitative agreement.","pith_inferences":["If Eq. (4) also holds for networks trained by gradient descent rather than Langevin sampling, the off-diagonal matrix $Q^*/\\lambda_1$ would provide a direct, weight-based diagnostic of feature sharing in practical finite networks; the paper does not test readout correlations under gradient-descent training, only aggregate generalization.","The same Gaussian-equivalence route should predict output-output correlations in other finite architectures with collective hidden variables, such as convolutional networks; measuring readout-overlap matrices in finite convolutional networks with shared filters would be a direct test.","The paper's observation that theory-data discrepancies shrink as the input dimension $N_0$ grows suggests finite-input corrections beyond $\\alpha$; a refined effective theory including $1/N_0$ corrections could make Eq. (4) quantitative at smaller $N_0$."],"forward_implications":["At any fixed load $\\alpha = P/N_1$, output-output correlations are a thermodynamic prediction of the model, not a finite-size accident; they survive in the proportional limit and vanish only as $Q^*$ approaches the identity.","The readout-overlap matrix $\\langle v v^T\\rangle/N_1$ is a direct experimental estimator of $Q^*/\\lambda_1$, so measuring trained readout weights is enough to read off the kernel renormalization.","Predictions for new data must use the renormalized kernel $K^{(\\mathrm{R})}_{Q^*} = Q^* \\otimes K_{\\mathrm{NNGP}}$; using the bare infinite-width kernel at finite $\\alpha$ misestimates both the generalization loss and the covariance of predictions.","Classes that share visual or semantic features are predicted to show systematically negative off-diagonal overlaps, giving a concrete signature of feature sharing that can be checked class by class.","Single-output networks cannot exhibit kernel shape renormalization, since with $D=1$ the matrix $Q$ is a scalar and cannot change the kernel's shape; the effect is intrinsically multi-output."],"supporting_citations":[{"why":"Supplies the classic infinite-width result that with $D$ readout units the network behaves as $D$ independent single-output networks, the baseline against which the paper's correlations are measured.","marker":"[1]"},{"why":"Flags the failure of the infinite-width description to account for output-output correlations in finite networks, the phenomenon the paper sets out to explain.","marker":"[2]"},{"why":"Defines the NNGP kernel $K_{\\mathrm{NNGP}}(X,X)$ whose shape is renormalized by $Q$ in Eq. (2).","marker":"[5]"},{"why":"Derives the proportional-limit effective action for deep linear networks, the starting point for the kernel renormalization formalism used here.","marker":"[25]"},{"why":"Provides the effective action for Bayesian deep networks beyond the infinite-width limit, from which Eq. (1) and the multi-output renormalized kernel are taken.","marker":"[26]"},{"why":"Demonstrates the predictive power of the Bayesian effective action for one-hidden-layer networks in the proportional limit, supporting the numerical validation strategy.","marker":"[29]"},{"why":"Supplies the generalized central limit theorem used to justify the Gaussian equivalence for the hidden-layer collective variables.","marker":"[41]"}],"fun_headline_variants":["Renormalized kernel predicts finite-net output correlations","Single matrix predicts finite-net output correlations","Output correlations in finite nets from kernel renormalization","Finite-width nets: output correlations from kernel renormalization","Kernel renormalization explains output correlations in finite nets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole calculation assumes the hidden-layer sums that couple the outputs are jointly Gaussian once $P$ and $N_1$ are large in fixed proportion, and the paper does not bound how far finite networks can stray from that approximation.","fun_headline_variants_meta":{"raw":{"variants":["Renormalized kernel predicts finite-net output correlations","Single matrix predicts finite-net output correlations","Output correlations in finite nets from kernel renormalization","Finite-width nets: output correlations from kernel renormalization","Kernel renormalization explains output correlations in finite nets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000273,"raw_usage":{"total_tokens":1675,"prompt_tokens":1027,"completion_tokens":648,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":643,"completion_tokens_details":{"reasoning_tokens":573}},"tokens_in":643,"tokens_out":648,"duration_ms":5624,"temperature":1.0,"reasoning_tokens":573,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:58:39.677818+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a fully thermalized one-hidden-layer Bayesian network with $D=3$ outputs on synthetic Gaussian inputs where two classes are built to share a hidden feature and the third is independent, with $P=N_1=1000$, and measure $\\langle v_a \\cdot v_b\\rangle/N_1$ from Langevin samples; Eq. (4) predicts this equals $Q^*_{ab}/\\lambda_1$ with $Q^*$ from minimizing Eq. (1), up to thermodynamic-limit corrections. If the off-diagonal overlap stays near zero, or has the wrong sign, while $\\alpha$ is held fixed and the sampler has converged, the Gaussian equivalence behind $Q^*$ fails; alternatively, checking that the discrepancy shrinks as $N_1$ grows at fixed $\\alpha$ tests the predicted scaling.","supporting_citations":[{"cited_title":"Priors for infinite networks,","cited_arxiv_id":null,"evidence_quote":"Supplies the classic infinite-width result that with $D$ readout units the network behaves as $D$ independent single-output networks, the baseline against which the paper's correlations are measured."},{"cited_title":"Introduction to gaussian processes,","cited_arxiv_id":null,"evidence_quote":"Flags the failure of the infinite-width description to account for output-output correlations in finite networks, the phenomenon the paper sets out to explain."},{"cited_title":"Deep neural networks as gaussian processes,","cited_arxiv_id":null,"evidence_quote":"Defines the NNGP kernel $K_{\\mathrm{NNGP}}(X,X)$ whose shape is renormalized by $Q$ in Eq. (2)."},{"cited_title":"Statistical mechanics of deep linear neural networks: The backpropagating kernel renormalization,","cited_arxiv_id":null,"evidence_quote":"Derives the proportional-limit effective action for deep linear networks, the starting point for the kernel renormalization formalism used here."},{"cited_title":"A statistical mechanics framework for bayesian deep neural networks beyond the infinite-width limit,","cited_arxiv_id":null,"evidence_quote":"Provides the effective action for Bayesian deep networks beyond the infinite-width limit, from which Eq. (1) and the multi-output renormalized kernel are taken."},{"cited_title":"Predictive power of a bayesian effective action for fully connected one hidden layer neural networks in the proportional limit,","cited_arxiv_id":null,"evidence_quote":"Demonstrates the predictive power of the Bayesian effective action for one-hidden-layer networks in the proportional limit, supporting the numerical validation strategy."},{"cited_title":"Central limit theorems for non-linear functionals of gaussian fields,","cited_arxiv_id":null,"evidence_quote":"Supplies the generalized central limit theorem used to justify the Gaussian equivalence for the hidden-layer collective variables."}],"review_version":1}