{"id":"5e0c0f2f-e684-497a-b6eb-0ba22e661bc2","arxiv_id":"2602.23039","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A single dynamical mean-field theory unifies Bayesian, gradient-flow, and Langevin training of random-feature regression and explains finite-time generalization error on power-law spectra.","lead":"The paper derives one mathematical framework that describes how the test error changes over time for random-feature regression trained by Bayesian inference, gradient flow, or stochastic Langevin dynamics. It explains spectral bias — easy modes are learned first — and why early stopping can improve generalization on power-law data.","discovery_kind":"unification","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Variational Gaussian approximation is asserted exact without finite-N error control; the quantitative claim rests on an unproven N→∞ limit.","rationale":"The reader identified the variational Gaussian approximation as the weakest assumption, and I agree that it is the single most load-bearing concern for the central claim that the effective equations quantitatively explain the generalization dynamics. The asserted N→∞ exactness is not proved; the only support is a CLT statement in Section 5, and the simulations are restricted to N=P=100 with no error bars. The proposed N-scaling check directly probes whether the approximation is under control at the sizes used to claim quantitative agreement. I also note the abstract promises analytical scaling laws that are not delivered, but that is a completeness issue secondary to the correctness of the central reduction. The reader's CONDITIONAL verdict remains appropriate: the framework is plausible and well supported, but the quantitative claim should be conditional on a controlled finite-N verification.","tokens_in":17946,"tokens_out":23136,"duration_ms":233437,"concrete_test":"For the same spectral profile η_i=i^{-3/2}, β=10, gβ=10^3, and teacher weights as in Fig. 2, simulate the full coupled system (Eq. 4) for N=P=50, 100, 200, 400, averaging over ≥10^5 disorder realizations. Compute the maximum relative deviation between the numerical average L_test(t) and the theoretical curve (29) solved from Eqs. (25)-(28) with matching discretization. If the deviation does not decrease at least as N^{-1/2}, the VGA's finite-N error is uncontrolled at the reported sizes; if it does, the exactness concern is mitigated for the demonstrated regime.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central reduction to the N independent effective equations (19)-(20) is obtained by averaging the MSRDJ generating functional over the Gaussian feature disorder and then approximating the resulting non-quadratic action (35) by a variational Gaussian (Appendix C). The exactness of the VGA is justified only by the statement in Section 5 that 'the central limit theorem guarantees the Gaussianity of the process' when N→∞. No bound is given for finite-N corrections, and the simulations in Figs. 2-4 use N=P=100, where the order parameters R̄ and C̄ still fluctuate at O(N^{-1/2}). These fluctuations feed back into the effective noise (Eq. 20) and response K (Eq. 21), so the true single-mode process is only asymptotically Gaussian. Because the disorder-averaged action contains the nonlinear ln det term in Eq. (35), the VGA is not exact at finite N. The figures show no error bars and only one value of N, so the agreement with theory cannot discriminate a correct N→∞ theory from one with uncontrolled O(N^{-1/2}) bias. If the VGA is not effectively exact at N=100, the claimed quantitative explanation of the generalization-error dynamics is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper derives a dynamical mean-field theory for random-feature regression with power-law kernel eigenvalue spectra, starting from the MSRDJ path-integral representation and a variational Gaussian approximation (VGA). The effective theory reduces the disorder-averaged training dynamics to N independent effective Langevin equations, coupled only through a collective memory kernel K and a self-consistently determined noise correlator built from order parameters R, C, and C̃. The test error is expressed as 1/2 C̄(t,t). The framework is claimed to unify Bayesian inference, gradient flow, and Langevin training, and the authors compare the resulting predictions with simulations at N=P=100 for spectral bias, the bias-variance trade-off, regularization effects, and early stopping.","tokens_in":18244,"tokens_out":8848,"duration_ms":90392,"significance":"If the large-N assumptions are valid, the paper provides a compact, parameter-free reduction of high-dimensional random-feature learning dynamics to scalar order parameters, unifying prior static and deterministic treatments (Canatar et al., Advani et al., Maloney et al., Bordelon et al.) and giving a mechanistic explanation of early stopping. The derivation in Appendices A–C is detailed, and the simulations in Figures 1–4 serve as independent consistency checks with no fitted parameters; these are genuine strengths. The main reservations are that one advertised analytical result is absent and that the VGA/self-averaging argument lacks finite-N control, so the quantitative claim is not yet fully supported.","major_comments":[{"comment":"The Introduction promises \"analytical results that relate the power-law exponent of the feature kernel, regularization, and early stopping time to obtain a minimal generalization gap.\" I could not find such a result in Sections 4–5. The paper provides a self-consistent integral-equation description and numerical evidence of early-stopping minima (Figs. 2–4), but no closed-form or asymptotic expression connects the exponent γ, the regularization gβ, and the stopping time to the minimal test error. Either supply the derivation (even in a special limit) or revise the advertised contribution; as written, one of the three central claims is not delivered.","section":"Section 1 (Introduction), Sections 4–5"},{"comment":"The derivation of the effective equations (19)–(25) rests on the variational Gaussian approximation applied to the action (35), whose nonlinear ln det term is non-quadratic. The only justification for exactness is the sentence in Section 5 that \"the central limit theorem guarantees the Gaussianity of the process\" as N→∞. No finite-N error bound or CLT statement (what is summed, in what sense) is given, and the term is not VGA-exact at finite N. All simulations use N=P=100 (Figs. 2–4) without error bars or a second value of N, so the agreement cannot separate a correct N→∞ theory from one with O(N^{-1/2}) bias. Please add finite-N scaling (e.g., N=200, 400 at fixed P/N), error bars, and a more precise large-N argument.","section":"Appendix C; Section 5; Eq. (35)"},{"comment":"The manuscript states that \"the training process is indeed self-averaging\" and that the order parameters R, C, C̃ \"concentrate around their expectation values as N→∞,\" but no evidence or variance calculation is provided. Since the goal is typical-case behavior, the relevance of ⟨Z⟩_Q rather than ⟨ln Z⟩_Q depends on this concentration. The simulations average over 10^5 disorder realizations, but no single-realization trajectory or standard deviation of L_test is shown. A demonstration of vanishing fluctuations with N is needed to support the self-averaging claim.","section":"Section 4.1 and Section 5"}],"minor_comments":[{"comment":"The expression L_test = 1/2 C̄(t,s) should be written as 1/2 C̄(t,t) (or with an explicit evaluation at s=t).","section":"Section 4.3, Eq. (29)"},{"comment":"The teacher weight ar{w}_i is sometimes typeset as w_i in the equations, which makes the teacher-student notation ambiguous. Please ensure overbars are consistently used.","section":"Eqs. (14), (19), (25)"},{"comment":"K is first defined as a two-time kernel K(t,s), but it is then written as K(t−s) and solved in the Fourier domain. Clarify the causal/one-sided Fourier convention and how the t=0 initial condition is encoded in the Fourier solution.","section":"Section 4.2, Eqs. (20)–(24)"},{"comment":"Please define η_1 and the normalization of Λ. The caption states Λ_ij = i^{-3/2} δ_ij, which corresponds to γ=1/2 in Eq. (9), but the value of η_1 is not given.","section":"Figure captions, Figs. 2–4"},{"comment":"The reference entry for Naveh et al. is malformed (\"journal = arXiv\"). Please correct it.","section":"References"},{"comment":"The phrase \"neural scaling laws\" is used broadly. The numerical experiments use a single value of P and N, so the paper does not directly demonstrate L_test ∼ N^{-α} or P^{-α} scaling. Please either add experiments varying P and N or clarify in the abstract that the paper studies the dynamics that underlie such scaling laws without extracting the scaling exponents here.","section":"Title and Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper is a serious physics-style contribution and the core derivation appears self-consistent. The main risks are over-claiming a missing analytical result and uncontrolled asymptotic justification of the VGA. If the authors can add finite-N scaling simulations and either derive or remove the minimal-gap claim, the paper would be publishable; I would not reject it at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe paper is a serious dynamical mean-field treatment of random-feature regression with power-law kernel spectra. The genuinely new thing is that a single effective equation set, Eqs. (19)–(25), covers Bayesian GP inference, gradient flow with/without weight decay, and Langevin dynamics with L2 regularization built in. That is a real unification of results that previously lived in separate papers. The derivation is detailed and self-consistent, and the simulations in Figs. 1–4 track the theory without fitted parameters, including the bias-variance tradeoff that explains early stopping. I think the core framework is defensible.\n\nThe main soft spot is the variational Gaussian approximation. The paper claims exactness in the N→∞ limit by saying the central limit theorem guarantees Gaussianity of the process (Discussion, Sec. 5). That is a plausible DMFT statement, but it is asserted, not proved; no finite-N error bound is given. The numerics use N=P=100, with 10^5 disorder samples but no error bars, and only one spectral exponent, so the agreement cannot distinguish a correct asymptotic theory from one with O(N^{-1/2}) bias. The stress-test note is right that this is load-bearing.\n\nSecond, the Introduction promises analytical results relating the power-law exponent, regularization, and early-stopping time to a minimal generalization gap. I could not find a closed-form relation in the text; what is delivered is a qualitative picture from the bias-variance decomposition. That gap should be fixed, either by deriving the relation or by toning down the promise.\n\nMinor: the paper provides no code, which slows reproducibility, and the figures would be stronger with error bars and a second value of N.\n\nOverall, the central argument holds up; the VGA concern is real but not disqualifying, because the effective equations are derived from first principles and the simulations are independent checks. I would send this to peer review, with a request to address the VGA exactness, add error bars and finite-N scaling, and either derive or retract the promised early-stopping formula. It deserves referee time.","headline":"A credible DMFT unifying Bayesian, gradient-flow, and Langevin regression with power-law spectra; main caveats are the unproven VGA exactness and a missing promised early-stopping relation.","tokens_in":18670,"tokens_out":2669,"would_cite":true,"duration_ms":26980,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In the large-N limit, kernel-regression training splits into N independent effective modes, and the test error is exactly one of the order parameters.","keywords":["neural scaling laws","random feature regression","dynamical mean-field theory","generalization dynamics","early stopping","power-law spectra","kernel regression","Langevin training"],"falsifier":"Compute, from simulations at fixed finite N and P, the connected third- or fourth-order cumulant of the order parameters C(t,s) and R(t,s) over disorder realizations. If these cumulants are not small compared with the Gaussian predictions, the central-limit justification fails and the theory's error curves lose quantitative control. A sharper test: measure the actual noise correlator of the effective single-mode equations and compare it with the predicted P η_i K*C̄*K^T; a mismatch at accessible N directly falsifies the self-consistency.","tokens_in":17856,"feed_emoji":"🧠","tokens_out":6128,"duration_ms":61189,"temperature":0.7,"pith_summary":"Random-feature regression on power-law data looks high-dimensional and mode-coupled, but this paper argues that after averaging over training sets the dynamics collapse to a tractable mean-field form. The central claim is that the training process is self-averaging and that, as N grows, the coupled system becomes N statistically independent effective Langevin equations, one per kernel eigenmode, tied together only by a collective memory kernel and a shared colored noise. The generalization error is exactly the diagonal of the correlation order parameter, L_test = 1/2 C̄(t,t), so the theory produces the full time course of bias and variance. From that single object the paper derives neural scaling laws, the slowdown of learning due to an Onsager-like reaction term, and a quantitative account of early stopping. Because Bayesian GP inference, gradient flow, and Langevin training are all limits of the same effective equations, the result unifies previously separate lines of work.","feed_headline":"Training dynamics collapse to N independent modes","feed_subtitle":"For power-law data, one self-consistent equation per eigenmode explains neural scaling laws and early stopping.","key_machinery":"The load-bearing object is the dynamical mean-field reduction: after disorder-averaging the stochastic training dynamics, the partition function is approximated by the best Gaussian process (variational Gaussian approximation), characterized by three order parameters: the response R(t,s), the correlation C(t,s), and the conjugate correlation C̃(t,s). These order parameters self-consistently determine a memory kernel K = (1−R)^{-1} and a colored noise with correlator 2β^{-1}δ + P η_i K*C̄*K^T. The resulting effective equation per mode i is (∂t + 1/(gβ))v_i + P η_i ∫K v_i = w̄_i/(gβ) + ξ_i, so the whole collective dynamics is encoded in how the eigenvalue η_i enters this single-mode equation.","core_discovery":"The paper establishes a dynamical mean-field theory for the typical learning dynamics of random-feature regression with power-law-distributed kernel eigenvalues. Its core result is a set of N decoupled effective equations, one for each eigenmode of the feature kernel, whose only mutual coupling is through a collective memory kernel K(t−s) and a common, self-consistently determined, time-correlated noise. The stationary point of the variational Gaussian approximation yields explicit closed equations for the mode means and Green's functions, and the disorder-averaged autocorrelation C̄(t,s) satisfies a self-consistent equation whose diagonal is the test error. The theory therefore yields not j","pith_inferences":["Editorial inference: the self-consistency relation between the memory kernel and the noise means a single measured trajectory of one mode's variance, combined with the known spectrum η_i, pins down the collective kernel K; this gives a practical way to verify the theory on real training runs without averaging over many datasets.","Editorial inference: if the variational Gaussian step has a slow, power-law approach to Gaussianity, then extrapolating scaling laws from small models to large ones—the practical motivation in the introduction—would inherit the same exponent-dependent error; measuring the fourth-order cumulant of C̄ at finite N would quantify how reliable such extrapolation is.","Editorial inference: the same DMFT structure should apply beyond Gaussian feature statistics to any rotationally invariant random feature ensemble in which the empirical kernel is Wishart-like, so the theory's predictions for scaling-law exponents may be more universal than the specific Gaussian-feature model used to derive them."],"forward_implications":["Scaling laws in this model are not free-fitting curves: the power-law exponent γ of the kernel spectrum, the regularization strength, and the stopping time determine the generalization-error dynamics through one self-consistent equation.","Early stopping emerges from the bias–variance trade-off in time: the bias falls monotonically as fast modes learn, while a self-reinforcing effective noise steadily grows the variance; the optimal stopping time is where the two cross.","Bayesian GP inference (t→∞), gradient flow (β→∞), and finite-temperature Langevin dynamics are limits of one effective single-mode equation, so results transfer between these training regimes.","Stronger L2 regularization has little effect on the location or depth of the early-stopping minimum, but it suppresses the late-time variance divergence, so it primarily helps when training noise is significant.","Because large eigenmodes are learned earlier and more accurately, the model gives a dynamical account of spectral bias: the order in which modes are learned is governed by η_i, with small-eigenvalue modes slowing down all others through the collective kernel."],"fun_headline_variants":["Eigenmode decoupling: one equation per mode drives neural scaling","Dynamical mean-field theory unifies Bayesian, gradient flow, Langevin","Power-law spectra: training splits into N independent eigenmodes","A single self-consistent equation per eigenmode predicts early stopping","Random feature regression: dynamics reduce to N decoupled modes"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Everything quantitative rests on the claim that, for large N, the variational Gaussian approximation becomes exact because the disorder-averaged dynamics are Gaussian; no error bound is given, and if non-Gaussianity is still sizable at the finite N and P used in practice, the theory's predicted error curves are uncontrolled.","fun_headline_variants_meta":{"raw":{"variants":["Eigenmode decoupling: one equation per mode drives neural scaling","Dynamical mean-field theory unifies Bayesian, gradient flow, Langevin","Power-law spectra: training splits into N independent eigenmodes","A single self-consistent equation per eigenmode predicts early stopping","Random feature regression: dynamics reduce to N decoupled modes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000389,"raw_usage":{"total_tokens":1887,"prompt_tokens":741,"completion_tokens":1146,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":1057}},"tokens_in":485,"tokens_out":1146,"duration_ms":8920,"temperature":1.0,"reasoning_tokens":1057,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T20:29:04.482988+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute, from simulations at fixed finite N and P, the connected third- or fourth-order cumulant of the order parameters C(t,s) and R(t,s) over disorder realizations. If these cumulants are not small compared with the Gaussian predictions, the central-limit justification fails and the theory's error curves lose quantitative control. A sharper test: measure the actual noise correlator of the effective single-mode equations and compare it with the predicted P η_i K*C̄*K^T; a mismatch at accessible N directly falsifies the self-consistency.","supporting_citations":[],"review_version":1}