Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Statistical mechanics of extensive-width Bayesian neural networks near interpolation

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A replica-symmetric free entropy predicts the Bayes-optimal test error and ordered feature-specialisation transitions of a two-layer Bayesian network near interpolation, for generic activations and weight priors.

desk verdict A genuinely new effective theory for feature learning in extensive-width two-layer BNNs near interpolation, but the load-bearing step is an uncontrolled second-moment closure; treat Results 2.1/2.2 as conjectures, and peer review should ask for a CGF check. read the letter →

arxiv 2505.24849 v2 pith:XCNB3NNM submitted 2025-05-30 stat.ML cond-mat.dis-nncond-mat.stat-mechcs.ITcs.LGmath.IT

classification stat.MLcond-mat.dis-nncond-mat.stat-mechcs.ITcs.LGmath.IT MSC 82B4460B2068T07
keywords Bayesianneuralnetworksstatisticalmechanicsoflearningreplicamethodextensive-widthinterpolationregimespecialisationtransitionsteacher-studentmodelHermiteexpansion
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims to give a closed-form statistical-mechanics description of Bayes-optimal learning in a two-layer neural network whose hidden layer is large — width proportional to the input dimension — and which is trained on quadratically many samples, the near-interpolation regime in which genuine feature learning occurs. The central object is a replica-symmetric free entropy whose extremization yields predictions for the Bayes-optimal generalization error, and a phase diagram of learning transitions, valid for essentially any activation function with a Hermite expansion and any i.i.d. prior on the weights. The predicted phenomenology is that with scarce data the student learns only non-linear combinations of the teacher's weights, in a phase universal across weight priors, while beyond a threshold it 'specializes', aligning its hidden neurons with the teacher's, with features tied to larger readout amplitudes learned first. If this is right, the theory supplies an information-theoretic benchmark for any training algorithm on this architecture, and it predicts that reaching the specialized solutions may require time exponential in the input dimension.

What carries the argument

The argument rests on a Gaussian ansatz: the replicated post-activations $\lambda^a$ are assumed jointly Gaussian with covariance given by the Hermite expansion of the activation, which makes the energetic term in the replicated partition function a low-dimensional integral. The entropic term is carried by the symmetric tensors $S^a_2 = W^{a\top}\mathrm{diag}(v)W^a/\sqrt{k}$ (Eq. (12)) and by the functional overlap $Q_W(v)$ measuring how much student neurons with readout value $v$ align with the corresponding teacher neurons. The load-bearing simplification is Eq. (15): the true conditional law of the matrices $S^a_2$ is replaced by a tilted generalised-Wishart measure matched only to its second moment, with Lagrange multiplier $\tau(Q_W) = \operatorname{mmse}_S^{-1}(1 - \mathbb{E}_{v\sim P_v}[v^2 Q_W(v)^2])$ enforcing that moment. This reduction is what makes the remaining extensive-rank matrix integrals tractable through spherical (HCIZ) integration, yielding the two matrix-denoising free entropies $\iota(\cdot)$ that close the variational formula.

What would settle it

Recompute the leading-order free entropy for a case outside the exactly solvable quadratic-activation class — for instance binary (Rademacher) inner weights with a cubic term in the activation — by an independent method that tracks fourth moments of the tensors $S^a_2$ (or by an ansatz matching the third and fourth moments of their entries), and compare with Eq. (6): any $O(n)$ mismatch would falsify Result 2.1. A cheaper, directly observable test is to check at numerically accessible sizes ($d \approx 150$–$250$) that the ordering of specialisation thresholds $\alpha_{\mathrm{sp},v}(\gamma)$ by readout magnitude and the exact prior-independence of the generalisation error below $\alpha_{\mathrm{sp}}$ both hold; visible prior dependence in the universal phase would contradict the theory.

Watch

Extended reading notes

Core claim

Stated on the paper's own terms: in the joint limit $d, k, n \to \infty$ with $k/d \to \gamma$ and $n/d^2 \to \alpha$, the averaged log-partition function (free entropy) of the teacher–student network equals the extremum of the variational functional $f_{\mathrm{RS}}^{\alpha,\gamma}$ in Eq. (6) over the order parameters $(q_2, \hat{q}_2, Q_W, \hat{Q}_W)$ that solve the saddle-point equations (76). The extremizer feeds Result 2.2, an approximate formula for the Bayes-optimal mean-square generalisation error (Eq. (7), simplified to Eq. (79) for linear readout with Gaussian label noise), and defines the specialisation transitions $\alpha_{\mathrm{sp},v}(\gamma)$ as the thresholds at which the per-readout overlap $Q_W(v)$ leaves zero. The paper claims this picture holds for generic Hermite-expandable activations and generic i.i.d. inner-weight priors, featuring a universal phase in which the free entropy is independent of the weight prior, an ordered cascade of specialisation transitions driven by the readout amplitudes $|v|$ (a continuum of transitions for continuous readout laws), and recovery of the previously solved quadratic-activation, Gaussian-weight case on the universal branch $Q_W(v) = 0$. It further reports that the specialising solution, though Bayes-optimal, can be hard to reach for practical algorithms — empirically exponential in $d$ for ADAM and Hamiltonian Monte Carlo with homogeneous or Rademacher readouts — while an adapted GAMP-RIE matches the universal branch in polynomial time.

Load-bearing premise

The load-bearing premise is that the replicated post-activations are jointly Gaussian and that the conditional law of the second-order tensors contributes to the free entropy only through its second moment, so the tilted Wishart replacement (15) is exact at leading order; the paper validates this numerically but gives no error bound.

Editorial extensions

If this is right

  • Any Hermite-expandable activation and i.i.d. weight prior now yields a numerically evaluable prediction of the Bayes-optimal test error at $n = \Theta(d^2)$, providing an information-theoretic benchmark that all learners on this architecture can be compared against.
  • The data-hunger of a teacher feature is set by the corresponding readout amplitude: features with larger $|v|$ specialise at smaller $\alpha$, so heterogeneous readouts produce a whole sequence of transitions, reducing to a continuum when the readout law is continuous.
  • In the universal phase the optimal performance is independent of the inner-weight prior; the prior only begins to matter once specialisation occurs and the higher-order Hermite terms of the activation become informative rather than acting as effective noise.
  • The Bayes-optimal specialised solution can be computationally hard to find — empirically requiring time exponential in $d$ for ADAM and HMC on homogeneous or Rademacher readouts — while a polynomial-time GAMP-RIE extension attains only the universal branch, signalling a statistical-to-computational gap.
  • The formalism reduces to the exactly known quadratic-activation Gaussian solution on the universal branch $Q_W = 0$, and matches the previously studied $n = \Theta(d)$ linear regime as $\alpha \to 0$, so it connects the known tractable limits.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the second-moment matching is exact at leading order, the same tilted-Wishart reduction should apply to any extensive-rank inference problem with exchangeable rows; a testable extension would replace $\mathrm{diag}(v)$ by a general covariance and predict specialisation thresholds ordered by the covariance spectrum rather than by readout amplitudes.
  • The paper's mechanism — higher Hermite terms of the activation are unlearnable noise until weights align — suggests that in deeper networks the same cascade would repeat layer by layer, with each layer's second-order tensor statistics having to be mastered before its higher-order features become useful.
  • A practical consequence the authors leave implicit: for targets with homogeneous readouts, the exponential-time evidence implies that polynomial-time learners are confined to the universal branch at $n = \Theta(d^2)$, whereas heterogeneous readouts may be the only route for polynomial-time learners to approach the Bayes-optimal error.
  • The continuum of transitions predicted for continuous readouts is a sharp, falsifiable signature: finite-size simulations should show successive plateaus in the overlap profile $Q_W(v)$ whose locations across $\alpha$ track the readout density, with spacing that narrows where the support of $P_v$ is dense.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies Bayes-optimal learning of a two-layer teacher-student network in the extensive-width regime where input dimension d, hidden width k, and sample size n diverge with k/d -> gamma and n/d^2 -> alpha. The main contribution is a replica-based free-entropy functional (Result 2.1, Eq. (6)) whose saddle-point equations (76) determine overlap order parameters, and a formula for the Bayes-optimal generalization error (Result 2.2, Eq. (7)). The derivation combines a joint-Gaussian ansatz for replicated post-activations, a localization approximation (10) for higher Hadamard powers of the overlap matrix, and a second-moment-matched tilted generalized Wishart replacement (15) for the conditional law of the second-order tensors. The paper validates the theory numerically against HMC, ADAM, Metropolis-Hastings, and an extended GAMP-RIE, and reports a sequence of specialization transitions as alpha increases. The authors state explicitly that the formulas are not expected to be exact and that the derivation relies on uncontrolled approximations.

Significance. If the theory is correct, it provides a tractable effective description of feature learning and specialization in extensively wide Bayesian neural networks near interpolation, going beyond the exactly solvable quadratic-activation/Gaussian-weight case. The numerical validation is extensive and the predictions are genuinely parameter-free: the order parameters solve self-consistent equations rather than being fitted to simulations. The paper also contributes a methodological blend of spherical integrals and replica techniques that may be useful for related extensive-rank problems. However, the central approximation at Eq. (15) is uncontrolled, and the known failure on quadratic activation with Gaussian weights (footnote 1) shows that the variational closure can select the wrong overlap branch. The strength of the numerical evidence is therefore tempered by the lack of any bound or universality argument for the moment-matched measure replacement.

major comments (3)
  1. [Sec. 4.3, Eq. (15)] The replacement of the conditional law P((S^a_2)|Q_W) in (14) by the tilted generalized Wishart measure (15) is the load-bearing uncontrolled step. After the Fourier representation, the entropic potential F_S in (65) involves the integral of exp((tau + hat_q2)/2 sum_a<b Tr S^a_2 S^b_2) under this conditional law; what matters is the entire scaled cumulant generating function of Tr S^a_2 S^b_2 at O(1) tilt, not just its first cumulant. Matching the second moment (63) controls only the linear term in the tilt, and higher cumulants can shift an O(d^2) contribution to the free entropy. The paper gives no bound or universality proof, and explicitly disclaims exactness in Sec. 2.2 and Sec. 5. Footnote 1 provides a concrete instance where the closure fails qualitatively: for sigma(x)=x^2 with Gaussian weights, the variational potential is flat enough that the theory selects Q_W(v)>0 while the rigorous solution has Q_W(v)=0. A targeted numerical check of the cumulant generating function under the true and simplified measures, or a restriction of the claims away from the quadratic/Gaussian case, is needed to support the specialization thresholds in Result 2.1.
  2. [Sec. 2.2, Result 2.1 and following paragraph] The paper labels (6) as a 'Result' and claims a generic activation scope, but the derivation relies on at least three unproved hypotheses: joint Gaussianity of replicated post-activations, diagonal localization (10) for Hadamard powers ell>=3, and the moment-matched replacement (15). The manuscript's own text states that the formulas are not expected to be exact and that finer order parameters would require unsolvable multi-matrix models. Given the known counterexample in footnote 1, the language 'Result 2.1' and the claim of generic activation overstate the status of the formulas. I recommend either stating these hypotheses explicitly as assumptions in the main results (and excluding the quadratic-activation/Gaussian-prior case from the claimed scope), or reframing the central claims as a quantitatively validated heuristic rather than a result.
  3. [Sec. 4.1, Eq. (10)] The localization approximation (10) for ell>=3 is asserted without proof and justified by analogy with Wishart matrices, where higher Hadamard powers have localized eigenvectors. The posterior measure over W^a is not rotationally invariant when the prior P_W is not Gaussian, so the Wishart analogy is not directly applicable. The numerical validation in Fig. 4 covers selected activations and parameters and cannot establish the approximation for the full claimed scope. The status of (10) as an ansatz should be made explicit in the statement of the main results, and its domain of validity should be discussed.
minor comments (4)
  1. [Eq. (7)] The displayed formula for the Bayes-optimal mean-square generalization error is typeset with an awkward line break in the expectation E_{lambda_0,lambda_1}[...]; using an explicit double integral over lambda_0 and lambda_1 would improve readability.
  2. [Fig. 1 caption] The caption uses 'divided by two' and 'half-generalisation error' for ADAM and HMC; the relationship between the Gibbs error and the Bayes error from App. C should be recalled in the caption so the factor of two is not confusing.
  3. [Sec. 4.2, text after Eq. (11)] The sentence 'This is is verified numerically a posteriori' contains a duplicated 'is'; please correct the typo.
  4. [App. G, Eq. (102)] The sentence 'it is not difficult to show' and the phrase 'realated' in the surrounding text should be rewritten; the dependence of mmse_S(tau) on gamma would benefit from an explicit statement of the conjectured constant C(gamma).

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the replica free-entropy and generalization-error predictions are solved from model primitives and checked against independent numerics; the uncontrolled moment closure is an accuracy limitation, not a circular reduction.

full rationale

The paper's central claim is a variational free entropy obtained by replica method combined with spherical integrals. The only reductions in the derivation are explicit approximations stated as assumptions: the joint Gaussian ansatz on replicated post-activations ('an ansatz we cannot prove but that we validate a posteriori', Sec. 4.1) and the replacement of the conditional second-order tensor law by a tilted generalized Wishart measure with matching second moment (Eq. 15). No parameter of the theory is fitted to simulated labels or test errors; the order parameters extremize the RS potential computed from activation Hermite coefficients, weight priors, alpha and gamma. The claimed specialization thresholds and generalization-error curves therefore do not reduce to their inputs by construction. The GAMP-RIE match to the universal branch is an explicitly designed algorithmic consistency check, not independent evidence used to fix the theory. Self-citations are present but are not load-bearing: the Gaussian ansatz is not imported as proved from Camilli et al.; it is flagged as unproven and validated numerically, and the replica-symmetry justification cites rigorous results whose assumptions do not include the target formula. The paper itself disclaims exactness and exposes the uncontrolled second-moment closure (footnote 1 and Sec. 5), which is a correctness/validity risk rather than circularity. No step was found where a predicted quantity equals a fitted parameter or an imported self-citation by definition.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central formulas rest on four unproved ansatze: joint Gaussianity of post-activations, localization of higher Hadamard powers, replacement of the conditional tensor measure by a tilted generalized Wishart with second-moment matching, and replica symmetry. These are standard heuristics in the field but are not proved for this model; the paper validates them numerically. No free parameters are fitted to data and no new physical entities are introduced.

assumptions (5)
  • domain assumption Joint Gaussianity of the replicated post-activations {lambda_a} for a=0..s, conditionally on the weights.
    Section 4.1: 'our key hypothesis is that {lambda_a} is jointly Gaussian, an ansatz we cannot prove but that we validate a posteriori'. It is needed to derive the energetic potential FE in Eq. (41).
  • ad hoc to paper Localization of higher Hadamard powers of the overlap matrix: for l>=3, ((1/d) W_i^a^T W_j^b)^l is approximately delta_{ij} Q_W^ab(v)^l up to od(1) norm.
    Section 4.2, Eq. (10). Reduces Q^ab_l to functions of QW; justified by analogy with Wishart matrices and numerical checks, not proved.
  • ad hoc to paper Replacement of the conditional measure P((S2^a)|QW) by the tilted generalized Wishart measure in Eq. (15), matching only the second moment (63).
    Section 4.3 and 4.4; the authors state the equality holds at leading exponential order assuming validity of the measure simplification. Higher moments could change the free entropy if they contribute at leading order.
  • domain assumption Replica trick limits commute and replica symmetry holds at the saddle point.
    Section 4 assumes the limits commute, and the RS ansatz is adopted following Bayes-optimal inference results (Barbier and Panchenko). Not proved for this extensive-rank matrix model.
  • domain assumption The readout vector v is quenched; treating it as learnable changes only subleading terms.
    Section 2.1 argues this by counting k unknowns versus kd inner weights; the assumption is used throughout the replica computation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Statistical mechanics of extensive-width Bayesian neural networks near interpolation." pith.science (2026). https://pith.science/paper/XCNB3NNM

@misc{pith2026250524849,
  author       = {Pith},
  title        = {Pith review of: Statistical mechanics of extensive-width Bayesian neural networks near interpolation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XCNB3NNM}},
  note         = {Machine review of arXiv:2505.24849}
}
read the original abstract

For three decades statistical mechanics has been providing a framework to analyse neural networks. However, the theoretically tractable models, e.g., perceptrons, random features models and kernel machines, or multi-index models and committee machines with few neurons, remained simple compared to those used in applications. In this paper we help reducing the gap between practical networks and their theoretical understanding through a statistical physics analysis of the supervised learning of a two-layer fully connected network with generic weight distribution and activation function, whose hidden layer is large but remains proportional to the inputs dimension. This makes it more realistic than infinitely wide networks where no feature learning occurs, but also more expressive than narrow ones or with fixed inner weights. We focus on the Bayes-optimal learning in the teacher-student scenario, i.e., with a dataset generated by another network with the same architecture. We operate around interpolation, where the number of trainable parameters and of data are comparable and feature learning emerges. Our analysis uncovers a rich phenomenology with various learning transitions as the number of data increases. In particular, the more strongly the features (i.e., hidden neurons of the target) contribute to the observed responses, the less data is needed to learn them. Moreover, when the data is scarce, the model only learns non-linear combinations of the teacher weights, rather than "specialising" by aligning its weights with the teacher's. Specialisation occurs only when enough data becomes available, but it can be hard to find for practical training algorithms, possibly due to statistical-to-computational~gaps.

Figures

Figures reproduced from arXiv: 2505.24849 by the authors.

Figure 1
Figure 1. Theoretical prediction (solid curves) of the Bayes-optimal mean-square generalisation error for Gaussian inner weights with ReLU(x) activation (blue curves) and Tanh(2x) activation (red curves), d = 150, γ = 0.5, with linear readout with Gaussian label noise of variance ∆ = 0.1 and different Pv laws. The dashed lines are the theoretical predictions associated with the universal solution, obtained by plugging QW (v) … view at source ↗
Figure 2
Figure 2. Theoretical prediction (solid curves) of the Bayes-optimal mean￾square generalisation error for binary inner weights and polynomial activa￾tions: σ1 = He2/ √ 2, σ2 = He3/ √ 6, σ3 = He2/ √ 2 + He3/6, with γ = 0.5, d = 150, linear readout with Gaussian label noise with ∆ = 1.25, and homogeneous readouts v = 1. Dots are optimal errors computed via Gibbs errors (see [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 4
Figure 4. Hamiltonian Monte Carlo dynamics of the overlaps Qℓ = Q01 ℓ between student and teacher weights for ℓ ∈ [5], with activation function ReLU(x), d = 200, γ = 0.5, linear readout with ∆ = 0.1 and two choices of sample rates and readouts: α = 1.0 with Pv = δ1 (Left) and α = 3.0 with Pv = N (0, 1) (Right). The teacher weights W0 are Gaussian. The dynamics is initialised informatively, i.e., on W0 . The overlap Q1 always … view at source ↗
Figures from the paper (8 more)
Figure 5
Figure 5. Figure 5: Different theoretical curves and numerical results for ReLU(x) activation, Pv = 1 4 (δ−3/ √ 5 + δ−1/ √ 5 + δ 1/ √ 5 + δ 3/ √ 5 ) , d = 150, γ = 0.5, with linear readout with Gaussian noise of variance ∆ = 0.1 Top left: Optimal mean-square generalisation error predicted…
Figure 6
Figure 6. Figure 6: Generalisation error for ReLU activation and Rademacher readout prior Pv of the theory reported in the main text (solid blue) versus the branch obtained from the simplified ansatz (83) (solid red); the green solid line shows QW ≡ 0 (universal branch), and empty circles…
Figure 7
Figure 7. Figure 7: Theoretical prediction (solid curves) of the Bayes-optimal mean-square generalisation error for binary inner weights and ReLU, eLU activations, with γ = 0.5, d = 150, Gaussian label noise with ∆ = 0.1, and fixed readouts v = 1. Dashed lines are obtained from the soluti…
Figure 8
Figure 8. Figure 8: Semilog (Left) and log-log (Right) plots of the number of gradient updates needed to achieve a test loss below the threshold ε ∗ < εuni. Student network trained with ADAM with optimised batch size for each point. The dataset was generated from a teacher network with Re…
Figure 9
Figure 9. Figure 9: Same as in [PITH_FULL_IMAGE:figures/full_fig_p032_9.png]
Figure 10
Figure 10. Figure 10: Trajectories of the generalisation error of neural networks trained with ADAM at fixed batch size B = ⌊n/4⌋, learning rate 0.05, for ReLU activation with parameters ∆ = 10−4 for the linear readout, γ = 0.5 and α = 5.0 > αsp (= 0.22, 0.12, 0.02 for homogeneous, Rademac…
Figure 11
Figure 11. Figure 11: Trajectories of the overlap q2 in HMC runs initialised uninformatively for the polynomial activation σ3 = He2/ √ 2 + He3/6 with parameters ∆ = 0.1 for the linear readout, γ = 0.5 and α = 1.0. Left: Homogeneous readouts. Centre: Rademacher readouts. Right: Gaussian rea…
Figure 12
Figure 12. Figure 12: Semilog (Left) and log-log (Right) plots of the number of Hamiltonian Monte Carlo steps needed to achieve an overlap q ∗ 2 > quni 2 , that certifies the universal solution is outperformed. The dataset was generated from a teacher with polynomial activation σ3 = He2/ √…

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Microscopic and collective signatures of feature learning in neural networks

    cond-mat.dis-nn 2025-08 conditional novelty 6.0 of 10

    In over-parameterized Bayesian one-hidden-layer networks, class-manifold separation becomes nonmonotonic in temperature and hidden weights develop data-dependent correlations, signatures of feature learning despite Ga...

Reference graph

Works this paper leans on

31 extracted references · 27 canonical work pages · cited by 1 Pith paper

  1. [2]

    | QW ) For the reader’s convenience we report here the measure P ((Sa

  2. [5]

    universal

    , d = 150, γ = 0.5, with linear readout with Gaussian noise of variance ∆ = 0.1 Top left: Optimal mean-square generalisation error predicted by the theory reported in the main text (solid blue) versus the branch obtained from the simplified ansatz(83) (solid red); the green solid line shows the universal branch corresponding toQW ≡ 0, and empty circles ar...

  3. [8]

    (57) Recall V is the support of Pv (assumed discrete for the moment)

    | QW ) = V kd W (QW )−1 Z 0,sY a dPW (Wa)δ(Sa 2 − Wa⊺(v)Wa/ √ k) 0,sY a≤b Y v∈V Y i∈Iv δ(d Qab W (v) − Wa⊺ i Wb i ). (57) Recall V is the support of Pv (assumed discrete for the moment). Recall also that we have quenched the readout weights to the ground truth. Indeed, as discussed in the main, considering them learnable or fixed to the truth does not cha...

  4. [9]

    (58) The measure is coupled only through the latter δ’s

    | QW ) 1 d2 TrSa 2Sb 2 = V kd W (QW )−1 Z 0,sY a dPW (Wa) 1 kd2 Tr[Wa⊺(v)WaWb⊺(v)Wb] × 0,sY a≤b Y v∈V Y i∈Iv δ(d Qab W (v) − Wa⊺ i Wb i ). (58) The measure is coupled only through the latter δ’s. We can decouple the measure at the cost of introducing Fourier conjugates whose values will then be fixed by a saddle point computation. The second moment comput...

  5. [10]

    (61) We have used the fact that⟨ · ⟩ˆB(v) is symmetric if the prior PW is, thus forcing us to match j with p if i ̸= l

    | QW ) 1 d2 TrSa 2Sb 2 = 1 kd2 kX i,l=1 dX j,p=1 ⟨W a ijviW a ipW b ljvlW b lp⟩{ ˆB(v)}v∈V = 1 k X v∈V v2 X i∈Iv D 1 d dX j=1 W a ijW b ij 2E ˆB(v) + 1 k dX j=1 D kX i=1 vi(W a ij)2 d kX l̸=i,1 vl(W b lj)2 d E ˆB(v) . (61) We have used the fact that⟨ · ⟩ˆB(v) is symmetric if the prior PW is, thus forcing us to match j with p if i ̸= l. Considering that by...

  6. [11]

    (63) Notice that the effective law ˜P ((Sa

    | QW ) 1 d2 TrSa 2Sb 2 = X v∈V Pv(v)v2Qab W (v)2 + γ¯v2 = Ev∼Pv v2Qab W (v)2 + γ¯v2. (63) Notice that the effective law ˜P ((Sa

  7. [12]

    | QW ) in (15) is the least restrictive choice among the Wishart-type distributions with a trace moment fixed precisely to the one above. In more specific terms, it is the solution of the following maximum entropy problem: inf P,τ n DKL(P ∥ P ⊗s+1 S ) + sX a≤b,0 τ ab EP 1 d2 TrSa 2Sb 2 − γ¯v2 − Ev∼Pv v2Qab W (v)2 o , (64) where PS is a generalised Wishart...

  8. [13]

    (65) The factor V kd W (QW ) was already treated in the previous section

    | QW ) 0,sY a≤b δ(d2Qab 2 − TrSa 2Sb 2). (65) The factor V kd W (QW ) was already treated in the previous section. However, here it will contribute as a tilt of the overall entropic contribution, and the Fourier conjugates ˆQab W (v) will appear in the final variational principle. Let us now proceed with the relaxation of the measure P ((Sa

Show all 31 references
  1. [14]

    | QW ) by replacing it with ˜P ((Sa

  2. [15]

    | QW ) given by (15): eFS = V kd W (QW ) Z d ˆQ2 exp − d2 2 sX a≤b,0 ˆQab 2 Qab 2 1 ˜V kd W (QW ) Z sY a=0 dPS(Sa

  3. [16]

    As usual, the Nishimori identities impose Qaa 2 = r2 = 1 + γ¯v2 without the need of any Fourier conjugate

    exp sX a≤b,0 τab + ˆQab 2 2 TrSa 2Sb 2 (66) where we have introduced another set of Fourier conjugates ˆQ2 for Q2. As usual, the Nishimori identities impose Qaa 2 = r2 = 1 + γ¯v2 without the need of any Fourier conjugate. Hence, similarly to τ aa, ˆQaa 2 = 0 too. Furthermore, ...

  4. [17]

    Y (or S0, ξ)

    exp τ + ˆq2 2 sX a<b,0 TrSa 2Sb 2 = 1 n E ln Z dPS(S2) exp 1 2 Tr p τ + ˆq2YS2 − (τ + ˆq2) S2 2 2 , (67) 20 Statistical mechanics of extensive-width Bayesian neural networks near interpolation where Y = Y(τ + ˆq2) = √τ + ˆq2S0 2 + ξ with ξ/ √ d a standard GOE matrix, and the o...

  5. [18]

    | QW ). Simplifying using the value of r2 = 1 + γ¯v2 according to the Nishimori identities, and using the I-MMSE relation between ι(τ ) and mmseS(τ ), we get mmseS(τ ) = 1 − Ev∼Pv v2QW (v)2 ⇐ ⇒ τ = mmse−1 S 1 − Ev∼Pv v2QW (v)2 . (74) Since mmseS is a monotonic decreasing funct...

  6. [19]

    | QW ) through moment matching A crucial step that allowed us to obtain a closed-form expression for the model’s free entropy is the relaxation ˜P ((Sa

  7. [20]

    | QW ) (15) of the true measure P ((Sa

  8. [21]

    | QW ) (14) entering the replicated partition function, as explained in Sec. 4. The specific form we chose (tilted Wishart distribution with a matching second moment) has the advantage of capturing crucial features of the true measure, such as the fact that the matrices Sa 2 a...

  9. [22]

    In this case, inspired by (Sakata & Kabashima, 2013; Kabashima et al., 2016), in order to relax (14) we can propose the Gaussian ansatz d ¯P ((Sa

    | QW ) (14) is the coupling between different replicas, becoming more and more relevant as α increases. In this case, inspired by (Sakata & Kabashima, 2013; Kabashima et al., 2016), in order to relax (14) we can propose the Gaussian ansatz d ¯P ((Sa

  10. [23]

    | QW ) = sY a=0 dSa 2 dY α=1 δ(Sa 2;αα − √ k¯v) × dY α1<α2 e− 1 2 Ps a,b=0 Sa 2;α1 α2 ¯τ ab(QW )Sb 2;α1 α2 p (2π)s+1 det( ¯τ (QW )−1) , (83) where ¯v is the mean of the readout prior Pv, and ¯τ (QW ) := (¯τ ab(QW ))a,b is fixed by [ ¯τ (QW )−1]ab = Ev∼Pv v2Qab W (v)2. In words...

  11. [25]

    sp” and “uni

    | QW ) (Gaussian with coupled replicas, leading to fsp, vs. pure generalised Wishart with independent replicas, leading to funi), rather than within a unified theory as in the main text; (ii) the predicted critical value ¯αsp(γ) seems to be systematically larger than the one o...

  12. [26]

    components

    | QW ) we treated the tensor Sa 2 as a whole, without considering the possibility that its “components” Sa 2;α1α2 (v) := vp |Iv| X i∈Iv W a iα1 W a iα2 (92) could follow different laws for different v ∈ V. To do so, let us define Qab 2 = 1 k X v,v′ v v′ X i∈Iv,j∈Iv′ (Ωab ij )2...

  13. [27]

    the true distribution P ((Sa

    | QW ) 1 d2 TrSa 2(v)Sb 2(v′)⊺ = δvv′ v2Qab W (v)2 + γ vv′p Pv(v)Pv(v′) (94) 25 Statistical mechanics of extensive-width Bayesian neural networks near interpolation w.r.t. the true distribution P ((Sa

  14. [28]

    | QW ) reported in (14). Despite the already good match of the theory in the main with the numerics, taking into account this additional level of structure thanks to a refined simplified measure could potentially lead to further improvements. The simplified measure able to mat...

  15. [29]

    standard Gaussian entries

    | QW ) ∝ Y v∈V Y a dP v S(Sa 2(v)) × Y v∈V Y a<b e 1 2 ¯τ ab v (QW )T rSa 2 (v)Sb 2(v), (95) where P v S is the law of a random matrix v ¯W ¯W⊺|Iv|−1/2 with ¯W ∈ Rd×|Iv| having i.i.d. standard Gaussian entries. For properly chosen (¯τ ab v ), (94) is verified for this simplifi...

  16. [30]

    This entails dropping entirely the dependencies among matrix entries, induced by their Wishart-like form (92), for each Sa 2(v)

    | QW ) (14) while taking into account the additional structure induced by (93), (94) and maintaining a solvable model, is to consider a generalisation of the relaxation (83). This entails dropping entirely the dependencies among matrix entries, induced by their Wishart-like fo...

  17. [31]

    universal

    | QW ) = Y v∈V sY a=0 dSa 2(v) dY α=1 δ(Sa 2;αα(v) − v p |Iv|) × Y v∈V dY α1<α2 e− 1 2 Ps a,b=0 Sa 2;α1 α2 (v)¯τ ab v (QW )Sb 2;α1 α2 (v) p (2π)s+1 det( ¯τv(QW )−1) . (96) The parameters (¯τ ab v (QW )) are then properly chosen to enforce (94) for all 0 ≤ a ≤ b ≤ s and v, v′ ∈...

  18. [916]

    URL https://doi

    doi: 10.1007/s00220-022-04387-w. URL https://doi. org/10.1007/s00220-022-04387-w . Barbier, J., Krzakala, F., Macris, N., Miolane, L., and Zdeborov ´a, L. Optimal errors and phase transitions in high-dimensional gen- eralized linear models. Proceedings of the National Academy ...

  19. [2016]

    URL https: //doi.org/10.1080/00018732.2016.1211393

    doi: 10.1080/00018732.2016.1211393. URL https: //doi.org/10.1080/00018732.2016.1211393. 14 Statistical mechanics of extensive-width Bayesian neural networks near interpolation A. Hermite basis and Mehler’s formula Recall the Hermite expansion of the activation: σ(x) = ∞X ℓ=0 µ...

  20. [2017]

    Krzakala, F., M ´ezard, M., and Zdeborov ´a, L

    URL https://arxiv.org/abs/1412.6980. Krzakala, F., M ´ezard, M., and Zdeborov ´a, L. Phase diagram and approximate message passing for blind calibration and dic- tionary learning. In 2013 IEEE International Symposium on Information Theory , pp. 659–663, 2013. doi: 10.1109/ISIT...

  21. [2019]

    doi: 10.1007/s00440-018-0879-0

    ISSN 1432-2064. doi: 10.1007/s00440-018-0879-0. URL https://doi.org/10.1007/s00440-018-0879-0 . Barbier, J. and Macris, N. Statistical limits of dictionary learning: Random matrix theory and the spectral replica method. Phys. Rev. E, 106:024136, Aug 2022. doi: 10.1103/PhysRevE...

  22. [2022]

    Guionnet, A

    URL https://proceedings.mlr.press/v145/ goldt22a.html. Guionnet, A. and Zeitouni, O. Large deviations asymptotics for spherical integrals. Journal of Functional Analysis , 188 (2):461–515, 2002. ISSN 0022-1236. doi: 10.1006/jfan. 2001.3833. URL https://www.sciencedirect.com/ s...

  23. [2023]

    Camilli, F., Tieplova, D., Bergamin, E., and Barbier, J

    URL https://arxiv.org/abs/2307.05635. Camilli, F., Tieplova, D., Bergamin, E., and Barbier, J. Information- theoretic reduction of deep neural networks to linear models in the overparametrized proportional regime. The 38th Annual Confer- ence on Learning Theory (to appear), 20...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.