{"id":"14cdda2f-6409-474c-a518-b4d3b332ffaa","arxiv_id":"2510.24616","paper_version":5,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A replica/HCIZ theory predicts the Bayes-optimal generalization error of proportional-width MLPs near interpolation and discovers layer-wise specialization transitions that make deeper targets harder to learn.","lead":"This paper uses statistical physics to predict the best possible performance of deep multi-layer perceptrons trained exactly at the point where data and parameters are comparable. It finds sharp phase transitions in which hidden neurons gradually specialize toward the teacher, with deeper layers learning later and harder.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Result 1's exactness for generic activations rests on the moment-matched Wishart ansatz (31), which is in concrete tension with the rigorous quadratic-activation result [95].","rationale":"I read the paper as aiming to give exact large-N limits for Bayes-optimal learning of linear-width MLPs near interpolation, with Results 1, 3, and 4 as the mathematical core. The manuscript is honest about its main hypothesis and even about the tension with [95], which is exactly why the measure replacement (31) is the most load-bearing concern: it is the one step that is both unproved and for which a concrete disagreement with a rigorous result is acknowledged in the paper itself. The Gaussian ansatz (11) is also central, but it is extensively validated and is not the place where a specific contradiction is reported. My concern does not overturn the paper: the deep-L results with μ1=μ2=0 do not use matrix integrals, the numerics for tanh/ReLU are broad and convincing, and the paper already marks the μ2≠0 case as special. But it does mean the headline 'exact for generic activations' should be qualified, and the conditional verdict is the right one. I keep the reader's verdict unchanged because the same cautious verdict already captures this risk; the concern sharpens the condition rather than flipping the assessment.","tokens_in":71350,"tokens_out":5910,"duration_ms":58761,"concrete_test":"Recompute the RS potential f_RS^(1) in Eq. (15) for L=1, σ(x)=x2, PW=N(0,1), Pv=δ1 at γ=0.5 and α∈{1,2,3,4,5}, using high-precision evaluation of ι(·) (e.g., 1000-point quadrature and the exact MMSE formula) and of τ(Q). Then check whether the global maximizer has Q>0, and compare the Bayes-optimal error from the Q=0 branch with the rigorous result of [95] for the same setting. If [95] gives Q=0 while the RS potential selects Q>0, then Eq. (31) is not an asymptotically exact simplification and Result 1 must be restricted to activations with μ2=0 or explicitly qualified as conjectural.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Result 1 gives the exact Bayes-optimal free entropy and order parameters is made for arbitrary shallow activations with μ0=0. But when μ2≠0, the derivation relies on replacing the true conditional law of S2 by a generalized Wishart measure with an exponential tilt fixed only by a single moment-matching condition (Eq. 31). This measure replacement is not derived from the original prior; its asymptotic equivalence is the key unproved step. The paper itself flags a concrete tension in Remark 4: for L=1, σ(x)=x2, PW=N(0,1), numerical maximization of the RS potential (15) appears to select Q(v)>0 when γ≲1, whereas the rigorous equations of [95] give Q(v)=0 for all (α,γ). Since σ=x2 satisfies the hypotheses of Result 1, this is not an exotic outside-the-domain case; if [95] is correct, the moment-matched measure is not equivalent to the true measure at the accuracy needed to select the equilibrium OPs. The response that the free-entropy difference is ≤1% and the potential is flat does not resolve the issue: a small free-energy error can still shift the location of the maximizing Q and change predicted specialization transitions (e.g., jumps in R2 and Q(v)). Thus the exactness of Eq. (15) for generic shallow MLPs is not established; the strongest defensible claim is the paper's own conjecture that exactness holds for μ2=0, where matrix integrals drop out and a partial proof exists (App. B6).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a replica-symmetric statistical physics theory for Bayes-optimal learning of a teacher MLP by a matched student MLP in the proportional-width, interpolation scaling d,k_l,n → ∞ with k_l/d → γ_l and n/d^2 → α. The central results are formulas for the limiting free entropy and order parameters for shallow MLPs (Result 1), two-hidden-layer MLPs (Result 3), and arbitrary-depth MLPs under restrictive hypotheses (Result 4). These formulas determine the Bayes-optimal generalization error (Result 2) and predict layer-wise and neuron-wise specialization transitions. The theory is tested extensively against HMC, Metropolis, GAMP-RIE, and ADAM, and the paper identifies the Gaussian ansatz on replicated post-activations (Eq. (11)) and a moment-matched generalized-Wishart replacement (Eq. (31)) as the main unproved ingredients, with partial proofs for special cases.","tokens_in":71739,"tokens_out":3356,"duration_ms":39321,"significance":"If correct, the paper would be a significant step toward a quantitative theory of feature learning in fully trained, finite-width MLPs in the interpolation regime, going beyond kernel, random-feature, and mean-field limits. The identification of functional order parameters indexed by readout amplitudes and effective readouts gives a concrete, falsifiable picture of how specialization propagates across layers and neurons. Strengths include the absence of fitted constants, the breadth of numerical validation with multiple algorithm families, the direct test of the Gaussian hypothesis in Fig. 4, and the partial proof in App. B6 for μ2 = 0. The phenomenological predictions — e.g., shallow-to-deep propagation of specialization and the difficulty of reaching the specialized state — are interesting and well supported by the simulations. However, the exactness claim for generic shallow activations is not established, and the paper itself documents a concrete tension with a rigorous quadratic-activation result. The safest and most defensible core is the μ2 = 0 shallow case and the deep results under (H2)/(H3), where matrix-integral approximations are absent or less central.","major_comments":[{"comment":"Result 1 is stated for arbitrary shallow activations with μ0 = 0, but for μ2 ≠ 0 its derivation relies on replacing the true conditional measure (30) by the moment-matched generalized-Wishart measure (31). This replacement is not derived, and a single moment condition does not determine the large-deviation rate function needed to select the equilibrium order parameters. The paper itself, in Remark 4, reports that for σ(x)=x², which satisfies the hypotheses of Result 1, numerical maximization of the RS potential selects Q(v)>0 for γ≲1 whereas the rigorous equations of [95] give Q(v)=0 for all (α,γ). The response that the free-entropy difference is ≤1% and the potential is flat does not resolve the issue: a small free-energy error can shift the location of the maximizing Q and change the predicted specialization transitions. Since σ=x² is inside the stated domain, Result 1 is not exact as","section":"Result 1 and Remark 4"},{"comment":"The diagonal-concentration assumptions on Hadamard powers (27)–(28) and the measure simplification (31) are load-bearing for the entropic potential, not merely technical. The paper states these are assumptions and validates them only a posteriori through the same learning curves the theory is meant to predict. This circularity is particularly acute for the μ2 ≠ 0 shallow case, where the HCIZ integral is evaluated under the simplified measure. The manuscript should either provide a direct test of (27)–(28) at the level of the large-deviation functional (not just of the resulting generalization error), or clearly mark the μ2 ≠ 0 formula as conjectural. Without this, Results 1 and 2 for generic activations such as ReLU cannot be regarded as established.","section":"Eqs. (27)–(28) and (31)"},{"comment":"The deep-layer results also rely on unproved simplifications, although matrix integrals are absent for L≥3. For L=2, the entropic contribution is evaluated using a relaxation of the conditional law of W^(2:1) with an exponential tilt fixed by moment matching (App. C1). The rectangular spherical integral then gives the result. This is a further instance of the same moment-matching issue: a single overlap moment is matched, but the full measure is replaced by a Gaussian-product base measure. The numerical agreement is good, but the claims of exactness in Remarks 4 and the text for L≥2 should be softened unless a proof strategy or a rigorous check of the measure equivalence is supplied. The paper has a partial proof only for the shallow μ2 = 0 case (App. B6).","section":"Results 3 and 4 / App. C1"}],"minor_comments":[{"comment":"The symbol K* is used both for the asymptotic off-diagonal covariance in the Gaussian hypothesis and for the evaluated function K(R2*,Q*). This is a potential source of confusion; consider using K∞ for the object in (10).","section":"Notation around Eq. (10) and Eq. (16)"},{"comment":"The definition of τ(Q) via mmse^{-1}_S is terse; the reader must consult App. B1 to see that this is the Lagrange multiplier enforcing the moment condition. A one-sentence intuitive explanation would help.","section":"Section II A, τ(Q) in Eq. (14)"},{"comment":"The claim that 'the free-entropy difference never exceeds ≈1%' is not documented with a figure or table. Given that the paper makes a quantitative claim about the size of the error, this statement should be backed by a plot of the RS potential versus Q in the problematic γ≲1 regime.","section":"Remark 4"},{"comment":"The partial proof for μ2 = 0 is a strength, but the precise hypotheses under which it applies (e.g., bounded activation, finite Hermite support) are not stated in the main text. Please state them explicitly.","section":"App. B6"}],"recommendation":"major_revision","confidential_remarks":"The paper is ambitious and contains a large amount of correct-looking physics and valuable numerics, but the exactness claims for the shallow generic-σ case go beyond what the derivation supports, as the authors themselves acknowledge in Remark 4. The contradiction with [95] for σ=x² is the key issue: it means the current Result 1 cannot be accepted as stated. The authors can likely fix this by re-scoping the main claims, marking the μ2≠0 formulas as conjectures, and strengthening the discussion of when the moment-matched measure is expected to be accurate. I would not reject, because the μ2=0 case and the deep-layer phenomenology remain substantial and defensible contributions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The sharpest thing to know: this is a serious attempt at a long-open problem, and it is the first framework I know that puts all four properties P1-P4 together. But the exactness claim for generic shallow activations is not established. Read it as a well-validated conjecture plus explicit hypotheses, not as a theorem, and note that the mu_2 != 0 case has a concrete open question the authors themselves identify.\n\nWhat is actually new is real. The replica/HCIZ combination is a genuine step beyond matrix-denoising and quadratic-activation settings, and the order parameters they introduce — functional overlaps indexed by readout amplitudes and effective readouts, including the two-argument Q_2(v,v^(2)) — are a useful way to organize the problem. No constants are fitted to data; the saddle-point equations derive from the priors. The numerical validation is extensive and mostly convincing: HMC, Metropolis, GAMP-RIE, and ADAM all line up with the theory across a wide parameter range, and the layer-wise specialisation transitions are a nice, interpretable payoff. For activations with mu_2 = 0, where the matrix integrals drop out, the case is stronger and there is even a partial proof in App. B6. The paper is also unusually honest about its own assumptions.\n\nThe soft spots are proportionate to how load-bearing they are. The Gaussian hypothesis (11) is unproved but plausible and directly tested; the diagonal-concentration assumptions (27)-(28) are in the same family as standard Wishart facts, though not proven for posterior samples. The real issue is the measure replacement in Eq. (31): the conditional law of S_2 is replaced by a generalized Wishart measure with one moment-matching condition. That is not derived, and if it fails, Result 1 fails. Remark 4 makes the tension concrete: for sigma(x)=x^2 with PW=N(0,1), which is inside the domain of Result 1, numerical maximization of the RS potential appears to select Q(v)>0 for gamma <= 1, while the rigorous equations of [95] give Q(v)=0 for all parameters. The authors' reply — free-entropy difference under 1% and a flat potential — does not dissolve the problem, because a small free-energy error can shift the maximizer and change predicted specialisation transitions. So the strongest defensible claim is the one the paper itself makes for mu_2 = 0: conjectured exact with strong evidence. For generic activations, exactness should be flagged as open rather than asserted.\n\nWho is this for: statistical physicists working on learning, and ML theorists interested in feature learning at interpolation. It deserves a serious referee — send it out — but the referees should require the generic-activation statements to be softened and the mu_2 != 0 question to be left explicitly open. I would cite it.","headline":"A serious, mostly honest attack on a hard open problem, with an exactness claim that currently outruns the proof — especially for generic activations with a second Hermite component.","tokens_in":72192,"tokens_out":2275,"would_cite":true,"duration_ms":27098,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["82B44","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"Statistical physics can now compute the Bayes-optimal learning limit of a deep neural network with width proportional to its input, in the interpolation regime.","keywords":["multi-layer perceptron","interpolation regime","Bayes-optimal generalization","replica method","HCIZ integral","feature learning","specialisation transition","statistical physics of learning"],"falsifier":"Run a large-d Bayesian sampling (e.g., HMC) of a shallow MLP in the interpolation regime with a generic activation (µ2 ≠ 0) and measure the moment generating function of the student's post-activations on test inputs, as in the paper's FIG. 4. If the relative error between the empirical and theoretical Gaussian MGF does not vanish as d grows — e.g., stays above O(1/√d) — the Gaussian hypothesis (11) is violated and Result 1 cannot hold exactly.","tokens_in":71253,"feed_emoji":"🧠","tokens_out":5649,"duration_ms":59115,"temperature":0.7,"pith_summary":"The paper claims that the optimal (Bayes-optimal) learning of a multi-layer perceptron in the interpolation regime — where widths are proportional to the input dimension and the number of samples is proportional to the square of it — is governed by a replica-symmetric free entropy formula built from a small set of functional overlaps between teacher and student weights. If correct, the theory exactly predicts the minimum achievable generalization error, the amount of data needed to reach it, and the sequence of 'specialisation' transitions through which hidden units align with target features. The analysis covers shallow networks with generic activation functions, two-layer networks with odd activations, and deeper networks under a restricted activation class, and it explains why feature learning outperforms kernels and random features: non-linear terms beyond the linear and quadratic components can only be exploited once the student specialises. A rich phenomenology follows, including layer-wise specialisation propagating from inner to outer layers, neuron-wise inhomogeneity driven by readout amplitudes, and metastable states that trap practical training algorithms.","feed_headline":"Replica theory delivers exact learning limits for deep networks","feed_subtitle":"Overlap order parameters locate the specialisation transitions where feature learning begins, layer by layer.","key_machinery":"The central objects are (i) the Gaussian ansatz on replicated post-activations, which reduces the energetic part to a low-dimensional covariance K*; (ii) the replacement of the conditional law of the quadratic composite S_2 = W^⊤ diag(v0) W by a generalized Wishart prior with an exponential tilt, whose Lagrange multiplier is fixed by matching the moment E[v² Q(v)²] + γ v̄² — this step is the crux that allows the theory to handle matrices that lack rotational invariance; and (iii) the use of HCIZ spherical integrals (and their rectangular counterpart for L=2) to evaluate the entropy of the matrix order parameters. The order parameters themselves — functional overlaps Q(v), Q1(v^(2)), Q2(v, v^","core_discovery":"Under the Gaussian hypothesis (11) — that the replicated post-activations of teacher and student converge to a jointly Gaussian law with covariance K* — combined with a measure simplification (31) that replaces the true conditional law of the quadratic sufficient statistics by a generalized Wishart prior with exponential tilt fixed by moment matching, the paper derives replica-symmetric formulas (Results 1, 3, 4) for the limiting free entropy of an MLP with L hidden layers in the proportional-width, quadratic-sample regime. The formulas express the free entropy as a variational problem over a few functional order parameters: overlaps labelled by readout amplitudes (and, for L=2, by effective","pith_inferences":["If the Gaussian ansatz extends to mismatched teacher-student settings (as the authors suggest), the same variational formulas could predict how much data a network needs to learn from a different function class, a step toward quantitative scaling laws.","The shallow-to-deep specialisation ordering implies a testable transfer-learning prediction: representations from early layers of a trained network should transfer to new tasks with smaller data budgets than those from deeper layers, because deeper layers need more data to specialise.","The formalism's success suggests that other extensive-rank matrix inference problems lacking rotational invariance, beyond the matrix-sensing problems the paper explicitly names, might be treated by the same replica-plus-HCIZ blend."],"forward_implications":["The Bayes-optimal generalization error for proportional-width MLPs in the interpolation regime is computable by maximizing a low-dimensional RS potential (Results 1–4), giving sharp limits that any algorithm trained on the same data cannot beat.","Feature learning beats kernels and random features because higher-order components of the teacher can only be exploited once the student's weights align (specialise) with those of the target; kernels never specialise, which explains the performance gap shown in FIG. 2.","Specialisation transitions are generically present and can be partial: sub-populations of neurons connected to larger readout amplitudes specialise first, and for L≥2 the transitions are layer-wise, propagating from inner to outer layers.","Deeper targets are harder: the overlap of the l-th layer decreases with layer index, and more data per layer is needed as L grows (FIG. 19).","Algorithms such as HMC, GAMP-RIE and ADAM get trapped in metastable states predicted by the theory; over-parameterisation (wider students) can recover part of the gap but the specialised equilibrium remains exponentially hard to reach in some cases."],"fun_headline_variants":["Statistical physics sets exact deep learning limits","Replica theory reveals layer-by-layer specialisation","Exact limits for deep network learning near interpolation","Bayes-optimal MLP: specialisation and learning transitions"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is the Gaussian ansatz (11) — that the replicated post-activations of teacher and student converge to a jointly Gaussian vector with covariance K* — supplemented by the moment-matching replacement of the conditional law of S_2 (31); if either fails, the replica formulas do not follow.","fun_headline_variants_meta":{"raw":{"variants":["Statistical physics sets exact deep learning limits","Replica theory reveals layer-by-layer specialisation","Exact limits for deep network learning near interpolation","Bayes-optimal MLP: specialisation and learning transitions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000332,"raw_usage":{"total_tokens":1710,"prompt_tokens":799,"completion_tokens":911,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":851}},"tokens_in":543,"tokens_out":911,"duration_ms":8622,"temperature":1.0,"reasoning_tokens":851,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T07:39:29.541622+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a large-d Bayesian sampling (e.g., HMC) of a shallow MLP in the interpolation regime with a generic activation (µ2 ≠ 0) and measure the moment generating function of the student's post-activations on test inputs, as in the paper's FIG. 4. If the relative error between the empirical and theoretical Gaussian MGF does not vanish as d grows — e.g., stays above O(1/√d) — the Gaussian hypothesis (11) is violated and Result 1 cannot hold exactly.","supporting_citations":[],"review_version":1}