{"id":"7bd565fa-1fae-4f08-8bd3-eab1f4e6a925","arxiv_id":"2501.18530","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Near n ~ d^2, Bayes-optimal two-layer networks with generic activations and weights have a universal phase, a specialisation phase, and a first-order transition between them.","lead":"This paper derives a statistical physics formula for the best possible prediction error of a Bayesian two-layer neural network when the number of training examples scales like the square of the input dimension. It shows that near interpolation the network has two regimes, a universal regime where the prior on the weights does not matter and a specialisation regime where the student aligns with the teacher, separated by a sharp transition.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The specialisation free entropy f_sp rests on the untested entropic Gaussian ansatz (22) for S2 given Q_W, not just on the post-activation Gaussian ansatz; α_sp and Result 3.2 shift if this conditional law is non-Gaussian. Direct MCMC test required.","rationale":"The reader correctly flags the Gaussian ansatz on replicated post-activations and Eq. (15) as load-bearing. My independent stress-test agrees with that general concern but isolates a more specific and, I think, more consequential assumption: the entropic Gaussian ansatz (22) for the conditional law of the tensor S2 given Q_W. This ansatz is what makes the specialisation free entropy f_sp tractable and is the source of the determinant term in Eq. (94). It is not the same assumption as joint Gaussianity of the λ_a: even if the post-activations were exactly Gaussian, the entropic term that controls the overlap q_W and the transition point could be wrong if the conditional law of S2 is non-Gaussian. The paper's own validation is indirect: Fig. 3 checks Eq. (16), which is a statement about the overlaps Q_ℓ, not about the joint distribution of S2 entries, and the reported deviations near transitions and the 1% miss for Gaussian read-outs in App. I are exactly where this ansatz would show up. I therefore recommend keeping the reader's CONDITIONAL verdict: the central claim is plausible and well-supported by numerics, but the specialisation branch, and hence the predicted transition and the post-transition error, remains conditional on an unverified distributional ansatz. The proposed HMC test is concrete and would settle whether the concern lands: it directly measures the object assumed in (22) in the regime where the specialisation solution is thermodynamically dominant and, at least with informative initialisation, algorithmically accessible. If the Gaussian conditional law is confirmed, the paper's central claim is substantially strengthened. If not, the free entropy comparison (7) is not reliable and the specialisation phase prediction would need to be revised or supplemented with a more general entropic ansatz.","tokens_in":41022,"tokens_out":6398,"duration_ms":77547,"concrete_test":"Run HMC for the Gaussian-weight ReLU setting of Fig. 2 (d=200, γ=0.5, Δ=0.1, α just above α_sp≈5.54, informative initialisation) and collect stationary samples {W^a} with a fixed empirical off-diagonal overlap q_W=Q_W^{ab}. For the d(d−1)/2 off-diagonal entries of S2=k^{−1/2}W^T diag(v)W, estimate the joint law of (S2_{αβ}^a, S2_{αβ}^b) across replicas a,b, in particular the fourth-order cumulant normalised by the square of the covariance (Q_W^{ab})^2. Compare with the bivariate Gaussian implied by (22). If the normalised cumulants vanish as d grows, the concern is settled. If they are O(1), re-derive Eq. (94) using the empirical conditional law and check whether the α_sp from Eq. (7) shifts by more than the width of the reported deviations around α≈4 in Fig. 2; if it does, the specialisation transition and the associated Result 3.2 predictions are not yet established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim selects the Bayes-optimal branch by the free-entropy comparison α_sp in Eq. (7), then reads off q_K and the generalisation error from Result 3.2. The universal free entropy f_uni is computed under ansatz (21), while the specialisation free entropy f_sp is computed under the separate entropic ansatz (22): conditioned on the overlap matrix Q_W, the off-diagonal entries of the hidden-layer tensor S2 are assumed zero-mean jointly Gaussian with covariance (Q_W^{◦2}), and the diagonal entries are taken to concentrate. This is an independent distributional assumption, distinct from the Gaussian ansatz on the replicated post-activations λ_a and from the Hadamard-power approximation (15). It enters f_sp through the Gaussian integral leading to the determinant term in Eq. (94); the saddle-point system (Ssp) and therefore the predicted α_sp are direct consequences of that functional form. The paper validates Eq. (16) in Fig. 3, but no experiment tests the conditional law of S2 itself. The only supporting evidence is the a-posteriori match of generalisation curves, which is a weak test in the narrow window where the specialisation branch is selected, and the authors themselves report small but non-negligible deviations near transitions (Sec. 6) and an approximately 1% miss of the predicted specialisation overlap for Gaussian read-outs (App. I). If the conditional law of S2 has non-Gaussian fourth-order cumulants at O(1) after the Gaussian scaling, the determinant in Eq. (94) is replaced by a different functional, and both the location of the transition and the specialisation-phase error formula would change.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies Bayes-optimal learning in a teacher-student two-layer network with i.i.d. Gaussian inputs, extensive width k=Theta(d), and n=Theta(d^2) samples. The authors use the replica method under a Gaussian ansatz for the replicated post-activations and a novel ansatz for the conditional law of the second-order tensor S2 to derive free-entropy expressions for two candidate solutions: a 'universal' phase with trivial weight overlap and a 'specialisation' phase in which the hidden weights align with the teacher. They predict a first-order phase transition at alpha_sp(gamma) via free-entropy comparison, provide saddle-point equations for the overlaps, and derive the limiting Bayes-optimal generalisation error. Numerical experiments with MCMC, HMC, and an extended GAMP-RIE are compared with the theory for binary and Gaussian weights, polynomial and ReLU/ELU activations, and the algorithmic hardness of reaching the specialisation solution is investigated.","tokens_in":41359,"tokens_out":7937,"duration_ms":91112,"significance":"The paper is a serious and detailed replica-method analysis. Its strongest contribution is a coherent effective theory for the Bayes-optimal generalisation error of two-layer networks in the quadratic-sampling regime, extending previous work that was restricted to quadratic activation and Gaussian weights. The derivation is self-contained, the saddle-point systems (Suni) and (Ssp) are explicit and reproducible, and the numerical experiments cover several activations, weight priors, and sampling rates, including a publicly available code. The theory makes falsifiable quantitative predictions, in particular the transition criterion (7) and the universal/specialisation error curves, and the paper is honest about the heuristic status of its key ansaetze. The main concern is that the entropic ansatz (22) for the specialisation phase is a second, untested distributional assumption on which the predicted transition and errors rest; the a-posteriori match of generalisation curves is a necessary but not sufficient check.","major_comments":[{"comment":"The specialisation phase free entropy f_sp relies on the entropic ansatz (22), which postulates that, conditionally on Q_W, the off-diagonal entries of the tensors S^a_2 are zero-mean jointly Gaussian with covariance (Q_W^{circ 2}) and that diagonal entries concentrate. This is a separate distributional assumption from the Gaussian ansatz on the post-activations lambda_a and from the Hadamard-power approximation (15). It enters f_sp through the Gaussian integral leading to the determinant term in Eq. (94), and the saddle-point system (Ssp), hence alpha_sp in Eq. (7) and the overlaps used in Result 3.2, are direct consequences of this functional form. The paper validates Eq. (16) in Fig. 3 and checks the generalisation curves, but no experiment tests the conditional law of S_2 itself. The authors report small but non-negligible deviations near transitions (Sec. 6) and a ~1% miss of the predicted specialisation overlap for Gaussian read-outs (App. I). I request a direct test: from HMC samples, compute the empirical conditional distribution of the off-diagonal S^a_2 entries given the empirical Q_W, compare its covariance and fourth-order cumulants with the Gaussian (22), and verify the determinant expression in Eq. (94) by Monte Carlo evaluation of both sides. This is currently the main unverified load-bearing input to the predicted transition.","section":"Section 5, Eq. (22)"},{"comment":"The simplification (Omega^ab)^ell approximately delta_ij (Q^ab_W)^ell for ell >= 3 is used to write the covariance K in Eq. (57) and hence the g(q_W) terms in Results 3.1 and 3.2. It is only checked numerically in Fig. 3 for a single polynomial activation and for the few overlaps appearing in that simulation; it is not verified for ReLU, ELU, or for off-equilibrium MCMC trajectories. Since g(q_W) contributes directly to q_K(q2, q_W) and to the free entropies f_uni and f_sp, I ask the authors to include a quantitative test of (15), for example by comparing both sides of the quadratic form (10) for a set of activations and for intermediate values of q_W, or to prove a bound that justifies the approximation in the relevant limit. Without such a test, the quantitative predictions of alpha_sp and of the specialisation-phase generalisation error rest on an assumption that is only spot-checked.","section":"Eq. (15), Section 5"}],"minor_comments":[{"comment":"The sentence after Eq. (60), 'combined with the Nishimori identities , combined with the Nishimori identities mW = qW, ...', contains a duplicated phrase; please remove the repetition.","section":"Appendix E.3"},{"comment":"The notation S^a_{2;alpha alpha} (with a semicolon) is not defined; please introduce a consistent notation for the entries of S^a_2, for example S^a_2(alpha, beta).","section":"Eq. (22)"},{"comment":"The abstract claims the theory applies to 'any activation function'; the body requires the activation to admit a Hermite expansion and the main text assumes mu_0 = 0 (App. F relaxes this to non-centred activations). Please qualify the abstract to reflect these conditions.","section":"Abstract"},{"comment":"The sentence about MCMC points following the specialisation curve before the transition should state clearly that these points are obtained with informative initialisation and therefore sample the metastable specialisation branch, not the equilibrium Gibbs distribution.","section":"Fig. 1 caption"},{"comment":"The statement 'we observe small but non-negligible deviations ... see for instance Fig. 2 around alpha = 4' should specify whether the deviation is within the finite-size error bars and, if not, discuss its possible origin, such as lack of equilibration or a systematic effect of the ansaetze.","section":"Section 6"},{"comment":"Please note in the caption that the Hermite coefficients are defined with respect to probabilist's Hermite polynomials, since the numerical values depend on the normalization convention.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is well-suited to the journal and the numerical study is substantial. I would be willing to accept after the requested direct test of the entropic ansatz (22) and an extension of the check of (15); without these, the specialisation phase transition remains a conjecture supported only by indirect evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. This paper generalises the Maillard et al. theory to arbitrary Hermite-expandable activations and generic iid weight priors in the n~d^2 regime, and it predicts a specialisation phase for Gaussian inner weights when the activation has Hermite degree above two. That is genuinely new. The second thing is less comforting: the entire specialisation branch, including the transition location alpha_sp, depends on a Gaussian ansatz on the conditional law of the hidden-layer tensor S2 given the overlap Q_W (ansatz (22)). That ansatz is distinct from the post-activation Gaussian ansatz and from the Hadamard-power approximation (15), and the paper does not test it directly. If it fails, the determinant in Eq. (94), the saddle-point system, and hence alpha_sp all shift.\n\nWhat the paper does well: the replica calculation is careful and internally consistent; the authors are transparent about what is ansatz and what is computed. The universal phase is on firmer ground: it follows from Gaussian universality and matches extensive MCMC and GAMP-RIE numerics. The Hadamard-power approximation (15) is checked directly in Fig. 3, which is good. The algorithmic side—extending GAMP-RIE to generic activations and showing it tracks the universal branch—is a useful contribution. The treatment of the mutual information and the link to matrix denoising is also coherent.\n\nWhere it is soft. The stress-test concern holds up. Ansatz (22) is chosen as the simplest law matching the asymptotic second-moment property of S2; it is not derived and no experiment probes the fourth-order cumulants of the conditional law. The a-posteriori match of generalisation curves is a weak test in the narrow specialisation window, and the paper itself reports a ~1% miss for Gaussian readouts (App. I) and small deviations near transitions. The post-activation Gaussian ansatz is also unproven, but that is a known, widely used heuristic in this literature; the specialisation ansatz is the one that needs stronger evidence, because the main new quantitative claim—the existence and location of the transition—sits on it.\n\nBottom line: this is a plausible and potentially important paper, more heuristic than the authors' confidence suggests. Send it to a serious referee; the referee should ask for either a derivation of (22) from a more controlled approximation or a direct MCMC test of the conditional law of S2, plus a commit-hash-pinned code release. With that, a conditional accept.","headline":"A solid heuristic extension of the quadratic-activation theory to generic activations and priors, but the specialisation transition rests on an unvalidated entropic ansatz and needs direct testing.","tokens_in":41876,"tokens_out":3255,"would_cite":true,"duration_ms":35929,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["82B44","60B20","68T07","62F15"],"pacs":[],"model":"deepseek-v4-flash","headline":"For a two-layer Bayesian network trained with $n=\\Theta(d^2)$ samples, the paper derives an effective theory of the Bayes-optimal generalisation error and predicts a discontinuous transition at a computable sampling rate $\\alpha_{\\rm sp}$…","keywords":["two-layer neural networks","teacher-student model","Bayes-optimal generalisation error","interpolation regime","replica method","Gaussian equivalence","phase transition","Hermite expansion"],"falsifier":"Compute the fourth-order cumulant $\\kappa_4 = \\mathbb{E}[\\lambda_0\\lambda_1^3] - 3K^{01}K^{11}$ of posterior-sampled post-activations on a fresh test input in the universal phase at, say, $d=1000$; the Gaussian ansatz predicts $\\kappa_4\\to0$, so a clearly nonzero value would falsify Result 3.2 and criterion (7) with it.","tokens_in":40758,"feed_emoji":"🧠","tokens_out":14429,"duration_ms":140837,"temperature":0.7,"pith_summary":"The paper considers a two-layer teacher–student network in which the input dimension $d$ and hidden width $k$ are both large with $k/d\\to\\gamma$, and the number of training samples $n$ is of order $d^2$, i.e. the regime where the number of trainable parameters matches the dataset size. Its central claim is that the Bayes-optimal generalisation error has a sharp, discontinuous phase transition at a computable sampling rate $\\alpha_{\\rm sp}(\\gamma)$: below it the student learns only certain nonlinear combinations of the teacher weights and the error is universal in the teacher weight prior; above it the student's hidden weights align with the teacher's and the error drops quickly but becomes prior-dependent. For activations with Hermite coefficients beyond the quadratic one, the transition occurs even for Gaussian weights; for a purely quadratic activation with Gaussian weights it does not. The same effective theory predicts that the superior specialisation solution is often algorithmically hard to find, so near interpolation the best performance reachable by practical algorithms can lie strictly above the Bayes-optimal value.","feed_headline":"Bayes-optimal error drops abruptly at a computable data threshold","feed_subtitle":"Theory predicts exactly where the optimal error stops being universal and the student locks onto the teacher.","key_machinery":"The engine is a replica calculation of the quenched free entropy. Its core object is the covariance $K^{ab}=\\sum_{\\ell\\ge1}\\mu_\\ell^2 Q^{ab}_\\ell/\\ell!$ of the replicated post-activations $(\\lambda_a)_{a\\ge0}$, which under the replica-symmetric ansatz concentrates to $q_K^*+(r_K-q_K^*)\\delta_{ab}$. The calculation rests on two substitutions: joint Gaussianity of the post-activations, and Eq. (15), which replaces Hadamard powers $(\\Omega^{ab}_{ij})^\\ell$ of the overlap matrix by the diagonal $\\delta_{ij}(Q^{ab}_W)^\\ell$ for $\\ell\\ge3$. The two phases correspond to two ansatze for the law of the order-2 tensor $S_2^a$: a rotationally invariant one (universal) and a Gaussian one aligned with $Q_W$ (specialisation). From these follow the free entropies $f_{\\rm uni}$ and $f_{\\rm sp}$, and transition criterion (7) compares the two scalars to locate $\\alpha_{\\rm sp}$.","core_discovery":"On the paper's own terms: in the limit $d,k,n\\to\\infty$ with $k/d\\to\\gamma$ and $n/d^2\\to\\alpha$, the Bayes-optimal generalisation error of a fully trained two-layer student with a centred activation is given by Result 3.2. The single input needed is the extremal overlap $q_K$: for $\\alpha<\\alpha_{\\rm sp}(\\gamma)$ it comes from the universal free entropy $f_{\\rm uni}$, and for $\\alpha>\\alpha_{\\rm sp}(\\gamma)$ from the specialisation free entropy $f_{\\rm sp}$; the transition point $\\alpha_{\\rm sp}(\\gamma)$ is the first $\\alpha$ for which $f_{\\rm sp}\\ge f_{\\rm uni}$, Eq. (7). In the universal phase the overlap $q_W$ between student and teacher inner weights is zero, so the student recovers only the combinations $S_1^0=k^{-1/2}v^{0\\top}W^0$ and $S_2^0=k^{-1/2}W^{0\\top}{\\rm diag}(v^0)W^0$, and the generalisation error is independent of the weight law. In the specialisation phase $q_W$ is positive, the student synchronises with the teacher, and the error decreases faster with explicit prior dependence, including a contribution $g(1)-g(q_W)$ from Hermite degrees $\\ell\\ge3$. The paper validates both branches numerically and reports that the specialisation branch is metastable or exponentially slow to reach for standard algorithms in many cases.","pith_inferences":["If the central claim is right, adding a small third-degree Hermite term to a quadratic activation should create an $\\alpha_{\\rm sp}$ for Gaussian weights, with the threshold decreasing as the third coefficient grows; this can be checked with the paper's own saddle-point equations.","If the central claim is right, the same universal-to-specialisation transition should appear in deeper or structured architectures whenever the sample count matches the parameter count, with layer-wise overlaps replacing the single matrix $Q_W$.","If the central claim is right, any estimator that is effectively rotationally invariant on the order-2 tensor is confined to the universal branch, which would explain a statistical-to-computational gap for certain teacher targets near interpolation.","If the central claim is right, the sharp error drop at $\\alpha_{\\rm sp}$ is a quantitative candidate mechanism for grokking-like sudden generalisation improvements in networks trained near interpolation."],"forward_implications":["For any activation with a nonzero Hermite coefficient beyond degree two, the paper predicts a threshold $\\alpha_{\\rm sp}(\\gamma)$ where the Bayes-optimal generalisation error drops discontinuously; the value is computable from Eq. (7) alone.","Below the transition the generalisation error is universal: only the mean and variance of the teacher's inner-weight prior matter, so a student with binary or Gaussian prior performs identically.","Above the transition the error is prior-dependent and decays faster; for Rademacher weights the overlap $q_W$ saturates exponentially fast in $\\alpha$, while for Gaussian weights the saturation is algebraic.","The polynomial-time approximate message-passing predictor adapted in the paper tracks the universal branch for all $\\alpha$, both before and after the transition, and therefore provides a tractable learner that may be strictly suboptimal.","For odd activations with $\\mu_2=0$ the error is flat in the universal phase and drops only at the transition, because the order-2 tensor is absent and higher-order tensors are learned only through specialisation."],"supporting_citations":[{"why":"sets the interpolation-regime problem for quadratic activation and Gaussian weights, maps it to a generalised linear model, and supplies the approximate message-passing algorithm that this work extends to arbitrary activation.","marker":"Maillard et al. (2024b)"},{"why":"supplies the Gaussian ansatz on replicated post-activations (their Conjecture 3.1) that this paper generalises from linearly many to quadratically many samples.","marker":"Cui et al. (2023)"},{"why":"provides the dictionary-learning ansatz generalised here to build the specialisation-phase branch of the free entropy.","marker":"Sakata & Kabashima (2013)"},{"why":"establishes the analogous universal/specialisation phases in extensive-rank matrix denoising and the mutual-information bound used for Rademacher inner weights.","marker":"Barbier et al. (2024)"},{"why":"supplies the generalised-linear-model free entropy and phase-transition results that fix the energetic potential and the universal-phase equations.","marker":"Barbier et al., 2019"},{"why":"introduces the Hermite-scalar order parameters used to write the post-activation covariance as an infinite series of overlap matrices.","marker":"Aguirre-López et al. (2025)"}],"fun_headline_variants":["Abrupt drop in optimal error at a computable threshold","Sharp transition: universal to specialised in wide nets","Near interpolation, optimal learning jumps discontinuously","Bayes-optimal error breaks: a sharp phase transition in nets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the replicated post-activations $(\\lambda_a)_{a\\ge0}$ are jointly Gaussian, an ansatz the paper states it cannot prove; if that fails for some activation or weight prior, the free entropies, the transition point $\\alpha_{\\rm sp}$, and the generalisation-error formula (Result 3.2) all collapse with it.","fun_headline_variants_meta":{"raw":{"variants":["Abrupt drop in optimal error at a computable threshold","Sharp transition: universal to specialised in wide nets","Near interpolation, optimal learning jumps discontinuously","Bayes-optimal error breaks: a sharp phase transition in nets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1503,"prompt_tokens":1060,"completion_tokens":443,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":676,"completion_tokens_details":{"reasoning_tokens":379}},"tokens_in":676,"tokens_out":443,"duration_ms":6077,"temperature":1.0,"reasoning_tokens":379,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T23:06:05.630030+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the fourth-order cumulant $\\kappa_4 = \\mathbb{E}[\\lambda_0\\lambda_1^3] - 3K^{01}K^{11}$ of posterior-sampled post-activations on a fresh test input in the universal phase at, say, $d=1000$; the Gaussian ansatz predicts $\\kappa_4\\to0$, so a clearly nonzero value would falsify Result 3.2 and criterion (7) with it.","supporting_citations":[],"review_version":1}