REVIEW 3 major objections 4 minor 1 cited by
Statistical mechanics of extensive-width Bayesian neural networks near interpolation
T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A replica-symmetric free entropy predicts the Bayes-optimal test error and ordered feature-specialisation transitions of a two-layer Bayesian network near interpolation, for generic activations and weight priors.
desk verdict A genuinely new effective theory for feature learning in extensive-width two-layer BNNs near interpolation, but the load-bearing step is an uncontrolled second-moment closure; treat Results 2.1/2.2 as conjectures, and peer review should ask for a CGF check. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument rests on a Gaussian ansatz: the replicated post-activations $\lambda^a$ are assumed jointly Gaussian with covariance given by the Hermite expansion of the activation, which makes the energetic term in the replicated partition function a low-dimensional integral. The entropic term is carried by the symmetric tensors $S^a_2 = W^{a\top}\mathrm{diag}(v)W^a/\sqrt{k}$ (Eq. (12)) and by the functional overlap $Q_W(v)$ measuring how much student neurons with readout value $v$ align with the corresponding teacher neurons. The load-bearing simplification is Eq. (15): the true conditional law of the matrices $S^a_2$ is replaced by a tilted generalised-Wishart measure matched only to its second moment, with Lagrange multiplier $\tau(Q_W) = \operatorname{mmse}_S^{-1}(1 - \mathbb{E}_{v\sim P_v}[v^2 Q_W(v)^2])$ enforcing that moment. This reduction is what makes the remaining extensive-rank matrix integrals tractable through spherical (HCIZ) integration, yielding the two matrix-denoising free entropies $\iota(\cdot)$ that close the variational formula.
What would settle it
Recompute the leading-order free entropy for a case outside the exactly solvable quadratic-activation class — for instance binary (Rademacher) inner weights with a cubic term in the activation — by an independent method that tracks fourth moments of the tensors $S^a_2$ (or by an ansatz matching the third and fourth moments of their entries), and compare with Eq. (6): any $O(n)$ mismatch would falsify Result 2.1. A cheaper, directly observable test is to check at numerically accessible sizes ($d \approx 150$–$250$) that the ordering of specialisation thresholds $\alpha_{\mathrm{sp},v}(\gamma)$ by readout magnitude and the exact prior-independence of the generalisation error below $\alpha_{\mathrm{sp}}$ both hold; visible prior dependence in the universal phase would contradict the theory.
Extended reading notes
Core claim
Stated on the paper's own terms: in the joint limit $d, k, n \to \infty$ with $k/d \to \gamma$ and $n/d^2 \to \alpha$, the averaged log-partition function (free entropy) of the teacher–student network equals the extremum of the variational functional $f_{\mathrm{RS}}^{\alpha,\gamma}$ in Eq. (6) over the order parameters $(q_2, \hat{q}_2, Q_W, \hat{Q}_W)$ that solve the saddle-point equations (76). The extremizer feeds Result 2.2, an approximate formula for the Bayes-optimal mean-square generalisation error (Eq. (7), simplified to Eq. (79) for linear readout with Gaussian label noise), and defines the specialisation transitions $\alpha_{\mathrm{sp},v}(\gamma)$ as the thresholds at which the per-readout overlap $Q_W(v)$ leaves zero. The paper claims this picture holds for generic Hermite-expandable activations and generic i.i.d. inner-weight priors, featuring a universal phase in which the free entropy is independent of the weight prior, an ordered cascade of specialisation transitions driven by the readout amplitudes $|v|$ (a continuum of transitions for continuous readout laws), and recovery of the previously solved quadratic-activation, Gaussian-weight case on the universal branch $Q_W(v) = 0$. It further reports that the specialising solution, though Bayes-optimal, can be hard to reach for practical algorithms — empirically exponential in $d$ for ADAM and Hamiltonian Monte Carlo with homogeneous or Rademacher readouts — while an adapted GAMP-RIE matches the universal branch in polynomial time.
Load-bearing premise
The load-bearing premise is that the replicated post-activations are jointly Gaussian and that the conditional law of the second-order tensors contributes to the free entropy only through its second moment, so the tilted Wishart replacement (15) is exact at leading order; the paper validates this numerically but gives no error bound.
Editorial extensions
If this is right
- Any Hermite-expandable activation and i.i.d. weight prior now yields a numerically evaluable prediction of the Bayes-optimal test error at $n = \Theta(d^2)$, providing an information-theoretic benchmark that all learners on this architecture can be compared against.
- The data-hunger of a teacher feature is set by the corresponding readout amplitude: features with larger $|v|$ specialise at smaller $\alpha$, so heterogeneous readouts produce a whole sequence of transitions, reducing to a continuum when the readout law is continuous.
- In the universal phase the optimal performance is independent of the inner-weight prior; the prior only begins to matter once specialisation occurs and the higher-order Hermite terms of the activation become informative rather than acting as effective noise.
- The Bayes-optimal specialised solution can be computationally hard to find — empirically requiring time exponential in $d$ for ADAM and HMC on homogeneous or Rademacher readouts — while a polynomial-time GAMP-RIE extension attains only the universal branch, signalling a statistical-to-computational gap.
- The formalism reduces to the exactly known quadratic-activation Gaussian solution on the universal branch $Q_W = 0$, and matches the previously studied $n = \Theta(d)$ linear regime as $\alpha \to 0$, so it connects the known tractable limits.
Reading between the lines
- If the second-moment matching is exact at leading order, the same tilted-Wishart reduction should apply to any extensive-rank inference problem with exchangeable rows; a testable extension would replace $\mathrm{diag}(v)$ by a general covariance and predict specialisation thresholds ordered by the covariance spectrum rather than by readout amplitudes.
- The paper's mechanism — higher Hermite terms of the activation are unlearnable noise until weights align — suggests that in deeper networks the same cascade would repeat layer by layer, with each layer's second-order tensor statistics having to be mastered before its higher-order features become useful.
- A practical consequence the authors leave implicit: for targets with homogeneous readouts, the exponential-time evidence implies that polynomial-time learners are confined to the universal branch at $n = \Theta(d^2)$, whereas heterogeneous readouts may be the only route for polynomial-time learners to approach the Bayes-optimal error.
- The continuum of transitions predicted for continuous readouts is a sharp, falsifiable signature: finite-size simulations should show successive plateaus in the overlap profile $Q_W(v)$ whose locations across $\alpha$ track the readout density, with spacing that narrows where the support of $P_v$ is dense.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies Bayes-optimal learning of a two-layer teacher-student network in the extensive-width regime where input dimension d, hidden width k, and sample size n diverge with k/d -> gamma and n/d^2 -> alpha. The main contribution is a replica-based free-entropy functional (Result 2.1, Eq. (6)) whose saddle-point equations (76) determine overlap order parameters, and a formula for the Bayes-optimal generalization error (Result 2.2, Eq. (7)). The derivation combines a joint-Gaussian ansatz for replicated post-activations, a localization approximation (10) for higher Hadamard powers of the overlap matrix, and a second-moment-matched tilted generalized Wishart replacement (15) for the conditional law of the second-order tensors. The paper validates the theory numerically against HMC, ADAM, Metropolis-Hastings, and an extended GAMP-RIE, and reports a sequence of specialization transitions as alpha increases. The authors state explicitly that the formulas are not expected to be exact and that the derivation relies on uncontrolled approximations.
Significance. If the theory is correct, it provides a tractable effective description of feature learning and specialization in extensively wide Bayesian neural networks near interpolation, going beyond the exactly solvable quadratic-activation/Gaussian-weight case. The numerical validation is extensive and the predictions are genuinely parameter-free: the order parameters solve self-consistent equations rather than being fitted to simulations. The paper also contributes a methodological blend of spherical integrals and replica techniques that may be useful for related extensive-rank problems. However, the central approximation at Eq. (15) is uncontrolled, and the known failure on quadratic activation with Gaussian weights (footnote 1) shows that the variational closure can select the wrong overlap branch. The strength of the numerical evidence is therefore tempered by the lack of any bound or universality argument for the moment-matched measure replacement.
major comments (3)
- [Sec. 4.3, Eq. (15)] The replacement of the conditional law P((S^a_2)|Q_W) in (14) by the tilted generalized Wishart measure (15) is the load-bearing uncontrolled step. After the Fourier representation, the entropic potential F_S in (65) involves the integral of exp((tau + hat_q2)/2 sum_a<b Tr S^a_2 S^b_2) under this conditional law; what matters is the entire scaled cumulant generating function of Tr S^a_2 S^b_2 at O(1) tilt, not just its first cumulant. Matching the second moment (63) controls only the linear term in the tilt, and higher cumulants can shift an O(d^2) contribution to the free entropy. The paper gives no bound or universality proof, and explicitly disclaims exactness in Sec. 2.2 and Sec. 5. Footnote 1 provides a concrete instance where the closure fails qualitatively: for sigma(x)=x^2 with Gaussian weights, the variational potential is flat enough that the theory selects Q_W(v)>0 while the rigorous solution has Q_W(v)=0. A targeted numerical check of the cumulant generating function under the true and simplified measures, or a restriction of the claims away from the quadratic/Gaussian case, is needed to support the specialization thresholds in Result 2.1.
- [Sec. 2.2, Result 2.1 and following paragraph] The paper labels (6) as a 'Result' and claims a generic activation scope, but the derivation relies on at least three unproved hypotheses: joint Gaussianity of replicated post-activations, diagonal localization (10) for Hadamard powers ell>=3, and the moment-matched replacement (15). The manuscript's own text states that the formulas are not expected to be exact and that finer order parameters would require unsolvable multi-matrix models. Given the known counterexample in footnote 1, the language 'Result 2.1' and the claim of generic activation overstate the status of the formulas. I recommend either stating these hypotheses explicitly as assumptions in the main results (and excluding the quadratic-activation/Gaussian-prior case from the claimed scope), or reframing the central claims as a quantitatively validated heuristic rather than a result.
- [Sec. 4.1, Eq. (10)] The localization approximation (10) for ell>=3 is asserted without proof and justified by analogy with Wishart matrices, where higher Hadamard powers have localized eigenvectors. The posterior measure over W^a is not rotationally invariant when the prior P_W is not Gaussian, so the Wishart analogy is not directly applicable. The numerical validation in Fig. 4 covers selected activations and parameters and cannot establish the approximation for the full claimed scope. The status of (10) as an ansatz should be made explicit in the statement of the main results, and its domain of validity should be discussed.
minor comments (4)
- [Eq. (7)] The displayed formula for the Bayes-optimal mean-square generalization error is typeset with an awkward line break in the expectation E_{lambda_0,lambda_1}[...]; using an explicit double integral over lambda_0 and lambda_1 would improve readability.
- [Fig. 1 caption] The caption uses 'divided by two' and 'half-generalisation error' for ADAM and HMC; the relationship between the Gibbs error and the Bayes error from App. C should be recalled in the caption so the factor of two is not confusing.
- [Sec. 4.2, text after Eq. (11)] The sentence 'This is is verified numerically a posteriori' contains a duplicated 'is'; please correct the typo.
- [App. G, Eq. (102)] The sentence 'it is not difficult to show' and the phrase 'realated' in the surrounding text should be rewritten; the dependence of mmse_S(tau) on gamma would benefit from an explicit statement of the conjectured constant C(gamma).
Circularity Check
No significant circularity: the replica free-entropy and generalization-error predictions are solved from model primitives and checked against independent numerics; the uncontrolled moment closure is an accuracy limitation, not a circular reduction.
full rationale
The paper's central claim is a variational free entropy obtained by replica method combined with spherical integrals. The only reductions in the derivation are explicit approximations stated as assumptions: the joint Gaussian ansatz on replicated post-activations ('an ansatz we cannot prove but that we validate a posteriori', Sec. 4.1) and the replacement of the conditional second-order tensor law by a tilted generalized Wishart measure with matching second moment (Eq. 15). No parameter of the theory is fitted to simulated labels or test errors; the order parameters extremize the RS potential computed from activation Hermite coefficients, weight priors, alpha and gamma. The claimed specialization thresholds and generalization-error curves therefore do not reduce to their inputs by construction. The GAMP-RIE match to the universal branch is an explicitly designed algorithmic consistency check, not independent evidence used to fix the theory. Self-citations are present but are not load-bearing: the Gaussian ansatz is not imported as proved from Camilli et al.; it is flagged as unproven and validated numerically, and the replica-symmetry justification cites rigorous results whose assumptions do not include the target formula. The paper itself disclaims exactness and exposes the uncontrolled second-moment closure (footnote 1 and Sec. 5), which is a correctness/validity risk rather than circularity. No step was found where a predicted quantity equals a fitted parameter or an imported self-citation by definition.
Assumptions & free parameters
assumptions (5)
- domain assumption Joint Gaussianity of the replicated post-activations {lambda_a} for a=0..s, conditionally on the weights.
- ad hoc to paper Localization of higher Hadamard powers of the overlap matrix: for l>=3, ((1/d) W_i^a^T W_j^b)^l is approximately delta_{ij} Q_W^ab(v)^l up to od(1) norm.
- ad hoc to paper Replacement of the conditional measure P((S2^a)|QW) by the tilted generalized Wishart measure in Eq. (15), matching only the second moment (63).
- domain assumption Replica trick limits commute and replica symmetry holds at the saddle point.
- domain assumption The readout vector v is quenched; treating it as learnable changes only subleading terms.
Cite this review
Pith. "Pith review of Statistical mechanics of extensive-width Bayesian neural networks near interpolation." pith.science (2026). https://pith.science/paper/XCNB3NNM
@misc{pith2026250524849,
author = {Pith},
title = {Pith review of: Statistical mechanics of extensive-width Bayesian neural networks near interpolation},
year = {2026},
howpublished = {\url{https://pith.science/paper/XCNB3NNM}},
note = {Machine review of arXiv:2505.24849}
}
read the original abstract
For three decades statistical mechanics has been providing a framework to analyse neural networks. However, the theoretically tractable models, e.g., perceptrons, random features models and kernel machines, or multi-index models and committee machines with few neurons, remained simple compared to those used in applications. In this paper we help reducing the gap between practical networks and their theoretical understanding through a statistical physics analysis of the supervised learning of a two-layer fully connected network with generic weight distribution and activation function, whose hidden layer is large but remains proportional to the inputs dimension. This makes it more realistic than infinitely wide networks where no feature learning occurs, but also more expressive than narrow ones or with fixed inner weights. We focus on the Bayes-optimal learning in the teacher-student scenario, i.e., with a dataset generated by another network with the same architecture. We operate around interpolation, where the number of trainable parameters and of data are comparable and feature learning emerges. Our analysis uncovers a rich phenomenology with various learning transitions as the number of data increases. In particular, the more strongly the features (i.e., hidden neurons of the target) contribute to the observed responses, the less data is needed to learn them. Moreover, when the data is scarce, the model only learns non-linear combinations of the teacher weights, rather than "specialising" by aligning its weights with the teacher's. Specialisation occurs only when enough data becomes available, but it can be hard to find for practical training algorithms, possibly due to statistical-to-computational~gaps.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 1 Pith paper
-
Microscopic and collective signatures of feature learning in neural networks
In over-parameterized Bayesian one-hidden-layer networks, class-manifold separation becomes nonmonotonic in temperature and hidden weights develop data-dependent correlations, signatures of feature learning despite Ga...
Reference graph
Works this paper leans on
-
[2]
| QW ) For the reader’s convenience we report here the measure P ((Sa
-
[5]
, d = 150, γ = 0.5, with linear readout with Gaussian noise of variance ∆ = 0.1 Top left: Optimal mean-square generalisation error predicted by the theory reported in the main text (solid blue) versus the branch obtained from the simplified ansatz(83) (solid red); the green solid line shows the universal branch corresponding toQW ≡ 0, and empty circles ar...
work page 2025
-
[8]
(57) Recall V is the support of Pv (assumed discrete for the moment)
| QW ) = V kd W (QW )−1 Z 0,sY a dPW (Wa)δ(Sa 2 − Wa⊺(v)Wa/ √ k) 0,sY a≤b Y v∈V Y i∈Iv δ(d Qab W (v) − Wa⊺ i Wb i ). (57) Recall V is the support of Pv (assumed discrete for the moment). Recall also that we have quenched the readout weights to the ground truth. Indeed, as discussed in the main, considering them learnable or fixed to the truth does not cha...
-
[9]
(58) The measure is coupled only through the latter δ’s
| QW ) 1 d2 TrSa 2Sb 2 = V kd W (QW )−1 Z 0,sY a dPW (Wa) 1 kd2 Tr[Wa⊺(v)WaWb⊺(v)Wb] × 0,sY a≤b Y v∈V Y i∈Iv δ(d Qab W (v) − Wa⊺ i Wb i ). (58) The measure is coupled only through the latter δ’s. We can decouple the measure at the cost of introducing Fourier conjugates whose values will then be fixed by a saddle point computation. The second moment comput...
-
[10]
| QW ) 1 d2 TrSa 2Sb 2 = 1 kd2 kX i,l=1 dX j,p=1 ⟨W a ijviW a ipW b ljvlW b lp⟩{ ˆB(v)}v∈V = 1 k X v∈V v2 X i∈Iv D 1 d dX j=1 W a ijW b ij 2E ˆB(v) + 1 k dX j=1 D kX i=1 vi(W a ij)2 d kX l̸=i,1 vl(W b lj)2 d E ˆB(v) . (61) We have used the fact that⟨ · ⟩ˆB(v) is symmetric if the prior PW is, thus forcing us to match j with p if i ̸= l. Considering that by...
-
[11]
(63) Notice that the effective law ˜P ((Sa
| QW ) 1 d2 TrSa 2Sb 2 = X v∈V Pv(v)v2Qab W (v)2 + γ¯v2 = Ev∼Pv v2Qab W (v)2 + γ¯v2. (63) Notice that the effective law ˜P ((Sa
-
[12]
| QW ) in (15) is the least restrictive choice among the Wishart-type distributions with a trace moment fixed precisely to the one above. In more specific terms, it is the solution of the following maximum entropy problem: inf P,τ n DKL(P ∥ P ⊗s+1 S ) + sX a≤b,0 τ ab EP 1 d2 TrSa 2Sb 2 − γ¯v2 − Ev∼Pv v2Qab W (v)2 o , (64) where PS is a generalised Wishart...
-
[13]
(65) The factor V kd W (QW ) was already treated in the previous section
| QW ) 0,sY a≤b δ(d2Qab 2 − TrSa 2Sb 2). (65) The factor V kd W (QW ) was already treated in the previous section. However, here it will contribute as a tilt of the overall entropic contribution, and the Fourier conjugates ˆQab W (v) will appear in the final variational principle. Let us now proceed with the relaxation of the measure P ((Sa
Show all 31 references
-
[14]
| QW ) by replacing it with ˜P ((Sa
-
[15]
| QW ) given by (15): eFS = V kd W (QW ) Z d ˆQ2 exp − d2 2 sX a≤b,0 ˆQab 2 Qab 2 1 ˜V kd W (QW ) Z sY a=0 dPS(Sa
-
[16]
As usual, the Nishimori identities impose Qaa 2 = r2 = 1 + γ¯v2 without the need of any Fourier conjugate
exp sX a≤b,0 τab + ˆQab 2 2 TrSa 2Sb 2 (66) where we have introduced another set of Fourier conjugates ˆQ2 for Q2. As usual, the Nishimori identities impose Qaa 2 = r2 = 1 + γ¯v2 without the need of any Fourier conjugate. Hence, similarly to τ aa, ˆQaa 2 = 0 too. Furthermore, ...
-
[17]
Y (or S0, ξ)
exp τ + ˆq2 2 sX a<b,0 TrSa 2Sb 2 = 1 n E ln Z dPS(S2) exp 1 2 Tr p τ + ˆq2YS2 − (τ + ˆq2) S2 2 2 , (67) 20 Statistical mechanics of extensive-width Bayesian neural networks near interpolation where Y = Y(τ + ˆq2) = √τ + ˆq2S0 2 + ξ with ξ/ √ d a standard GOE matrix, and the o...
2024
-
[18]
| QW ). Simplifying using the value of r2 = 1 + γ¯v2 according to the Nishimori identities, and using the I-MMSE relation between ι(τ ) and mmseS(τ ), we get mmseS(τ ) = 1 − Ev∼Pv v2QW (v)2 ⇐ ⇒ τ = mmse−1 S 1 − Ev∼Pv v2QW (v)2 . (74) Since mmseS is a monotonic decreasing funct...
-
[19]
| QW ) through moment matching A crucial step that allowed us to obtain a closed-form expression for the model’s free entropy is the relaxation ˜P ((Sa
-
[20]
| QW ) (15) of the true measure P ((Sa
-
[21]
| QW ) (14) entering the replicated partition function, as explained in Sec. 4. The specific form we chose (tilted Wishart distribution with a matching second moment) has the advantage of capturing crucial features of the true measure, such as the fact that the matrices Sa 2 a...
-
[22]
In this case, inspired by (Sakata & Kabashima, 2013; Kabashima et al., 2016), in order to relax (14) we can propose the Gaussian ansatz d ¯P ((Sa
| QW ) (14) is the coupling between different replicas, becoming more and more relevant as α increases. In this case, inspired by (Sakata & Kabashima, 2013; Kabashima et al., 2016), in order to relax (14) we can propose the Gaussian ansatz d ¯P ((Sa
2013
-
[23]
| QW ) = sY a=0 dSa 2 dY α=1 δ(Sa 2;αα − √ k¯v) × dY α1<α2 e− 1 2 Ps a,b=0 Sa 2;α1 α2 ¯τ ab(QW )Sb 2;α1 α2 p (2π)s+1 det( ¯τ (QW )−1) , (83) where ¯v is the mean of the readout prior Pv, and ¯τ (QW ) := (¯τ ab(QW ))a,b is fixed by [ ¯τ (QW )−1]ab = Ev∼Pv v2Qab W (v)2. In words...
-
[25]
sp” and “uni
| QW ) (Gaussian with coupled replicas, leading to fsp, vs. pure generalised Wishart with independent replicas, leading to funi), rather than within a unified theory as in the main text; (ii) the predicted critical value ¯αsp(γ) seems to be systematically larger than the one o...
-
[26]
components
| QW ) we treated the tensor Sa 2 as a whole, without considering the possibility that its “components” Sa 2;α1α2 (v) := vp |Iv| X i∈Iv W a iα1 W a iα2 (92) could follow different laws for different v ∈ V. To do so, let us define Qab 2 = 1 k X v,v′ v v′ X i∈Iv,j∈Iv′ (Ωab ij )2...
-
[27]
the true distribution P ((Sa
| QW ) 1 d2 TrSa 2(v)Sb 2(v′)⊺ = δvv′ v2Qab W (v)2 + γ vv′p Pv(v)Pv(v′) (94) 25 Statistical mechanics of extensive-width Bayesian neural networks near interpolation w.r.t. the true distribution P ((Sa
-
[28]
| QW ) reported in (14). Despite the already good match of the theory in the main with the numerics, taking into account this additional level of structure thanks to a refined simplified measure could potentially lead to further improvements. The simplified measure able to mat...
-
[29]
standard Gaussian entries
| QW ) ∝ Y v∈V Y a dP v S(Sa 2(v)) × Y v∈V Y a<b e 1 2 ¯τ ab v (QW )T rSa 2 (v)Sb 2(v), (95) where P v S is the law of a random matrix v ¯W ¯W⊺|Iv|−1/2 with ¯W ∈ Rd×|Iv| having i.i.d. standard Gaussian entries. For properly chosen (¯τ ab v ), (94) is verified for this simplifi...
2000
-
[30]
This entails dropping entirely the dependencies among matrix entries, induced by their Wishart-like form (92), for each Sa 2(v)
| QW ) (14) while taking into account the additional structure induced by (93), (94) and maintaining a solvable model, is to consider a generalisation of the relaxation (83). This entails dropping entirely the dependencies among matrix entries, induced by their Wishart-like fo...
-
[31]
universal
| QW ) = Y v∈V sY a=0 dSa 2(v) dY α=1 δ(Sa 2;αα(v) − v p |Iv|) × Y v∈V dY α1<α2 e− 1 2 Ps a,b=0 Sa 2;α1 α2 (v)¯τ ab v (QW )Sb 2;α1 α2 (v) p (2π)s+1 det( ¯τv(QW )−1) . (96) The parameters (¯τ ab v (QW )) are then properly chosen to enforce (94) for all 0 ≤ a ≤ b ≤ s and v, v′ ∈...
2025
-
[916]
URL https://doi
doi: 10.1007/s00220-022-04387-w. URL https://doi. org/10.1007/s00220-022-04387-w . Barbier, J., Krzakala, F., Macris, N., Miolane, L., and Zdeborov ´a, L. Optimal errors and phase transitions in high-dimensional gen- eralized linear models. Proceedings of the National Academy ...
-
[2016]
URL https: //doi.org/10.1080/00018732.2016.1211393
doi: 10.1080/00018732.2016.1211393. URL https: //doi.org/10.1080/00018732.2016.1211393. 14 Statistical mechanics of extensive-width Bayesian neural networks near interpolation A. Hermite basis and Mehler’s formula Recall the Hermite expansion of the activation: σ(x) = ∞X ℓ=0 µ...
2016
-
[2017]
Krzakala, F., M ´ezard, M., and Zdeborov ´a, L
URL https://arxiv.org/abs/1412.6980. Krzakala, F., M ´ezard, M., and Zdeborov ´a, L. Phase diagram and approximate message passing for blind calibration and dic- tionary learning. In 2013 IEEE International Symposium on Information Theory , pp. 659–663, 2013. doi: 10.1109/ISIT...
2013 arXiv
-
[2019]
doi: 10.1007/s00440-018-0879-0
ISSN 1432-2064. doi: 10.1007/s00440-018-0879-0. URL https://doi.org/10.1007/s00440-018-0879-0 . Barbier, J. and Macris, N. Statistical limits of dictionary learning: Random matrix theory and the spectral replica method. Phys. Rev. E, 106:024136, Aug 2022. doi: 10.1103/PhysRevE...
-
[2022]
Guionnet, A
URL https://proceedings.mlr.press/v145/ goldt22a.html. Guionnet, A. and Zeitouni, O. Large deviations asymptotics for spherical integrals. Journal of Functional Analysis , 188 (2):461–515, 2002. ISSN 0022-1236. doi: 10.1006/jfan. 2001.3833. URL https://www.sciencedirect.com/ s...
2002
-
[2023]
Camilli, F., Tieplova, D., Bergamin, E., and Barbier, J
URL https://arxiv.org/abs/2307.05635. Camilli, F., Tieplova, D., Bergamin, E., and Barbier, J. Information- theoretic reduction of deep neural networks to linear models in the overparametrized proportional regime. The 38th Annual Confer- ence on Learning Theory (to appear), 20...
2025 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.