{"id":"b512ffce-c76f-4c1e-8417-fafae4aa2983","arxiv_id":"2501.19158","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Training dynamics of Gaussian energy-based models decouple across eigenmodes of the empirical covariance, yielding exact early-stopping times, random-matrix finite-sample corrections, and a GCV-type train-test relation for energy-based models.","lead":"A theory paper shows that overfitting in Gaussian energy-based models can be understood through the spectral decomposition of the data covariance, with each eigenvalue learning at its own speed. This yields exact predictions for early stopping, finite-sample random-matrix corrections, and a cross-validation style relation for energy-based models.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Generic-EBM/NTK extension (Sec. 7) does not reproduce its stated GEBM limit: the ODE dJ/dt = I - (ĈJ + JĈ) has scalar solution (1-e^{-2ĉt})/(2ĉ), not Eq. (14)'s (1-e^{-ĉt})/ĉ, so the proposed NTK dynamics need correction before the abstract's 'deriving' claim stands.","rationale":"I read the paper as claiming (i) exact GEBM training dynamics, (ii) RMT finite-sample corrections and the GCV analog, (iii) an approximate BM extension, and (iv) a generic NTK extension. Claims (i)-(ii) are well supported: Eqs. (5)-(6) follow from the projected gradient equations, Fig. 4 shows finite-N convergence to RMT, and Eq. (9) is derived in Appendix G.2. Claim (iii) is explicitly approximate and the paper says so, so the abstract's 'minimal variations' wording is an overstatement but not a hidden flaw. Claim (iv) is the most load-bearing because the title promises a framework for energy-based modeling generally, and Section 7 contains a concrete factor-of-two inconsistency: the stated matrix ODE does not yield Eq. (14). The reader's weakest assumption (eigenvector noise) is not where I find the problem: RMT deterministic equivalents already account for eigenvector misalignment in the proportional limit, and Appendix B plus Fig. 4 provide supporting evidence for the eigenvalue-driven picture. I agree with the CONDITIONAL verdict, but for a different reason: the core GEBM results stand, while Section 7 should be corrected or explicitly demoted; if the concrete test confirms the inconsistency, the abstract's final sentence should be revised.","tokens_in":31663,"tokens_out":22338,"duration_ms":238760,"concrete_test":"Re-derive the GEBM limit of Section 7 in vectorized form. Start from L_SM = 1/2 E[||Jx||^2] - Tr J (Eq. 13 for a Gaussian EBM) and compute dJ/dt under the same symmetrization convention used to write dJ/dt = -(ĈJ + JĈ) + I. Then check whether the scalar evolution is dj/dt = 1 - xj or dj/dt = 1 - 2xj. If the latter, Eq. (14) is wrong by a factor of 2 and Eqs. (17)-(21) need to be recomputed; if the former, specify the convention that removes the factor and verify Eqs. (15)-(16) are sign-consistent with Eq. (13). This single check settles whether the generic-EBM extension is a derivation or a conjecture.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The GEBM/RMT/GCV core (Sections 2-5, Appendices G-H) is internally consistent and numerically supported; I do not find a load-bearing flaw there. The load-bearing problem is the claimed generalization in Section 7. After Eq. (13) the paper states that score matching on a GEBM gives dJ/dt = -(ĈJ + JĈ) + I and that this 'leads to' j_t(x) = (1-e^{-xt})/x (Eq. 14). For an initial condition commuting with Ĉ, the scalarized ODE is dj/dt = 1 - 2xj, whose solution is j_t(x) = (1-e^{-2xt})/(2x), not Eq. (14). The factor is not a harmless convention choice: it propagates into Eqs. (17)-(21), the proposed empirical-RKHS solution used to justify a generic-EBM framework. The manuscript itself calls the section a postulate and requests experimental investigation, but the abstract's final sentence says the NTK dynamics are 'derived'. As written, that derivation is internally inconsistent. The consequence is not that the GEBM results fail; the title-level extension to arbitrary EBMs rests on Section 7 and needs correction or clear demotion to a conjecture.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies overfitting in Gaussian energy-based models (GEBMs) by diagonalizing the training dynamics in the eigenbasis of the empirical covariance. It derives an explicit Lambert-W solution for eigenvalue dynamics, uses random matrix theory to obtain asymptotic train/test energies and coupling errors, derives a GCV-type relation for spectral-L1 regularization, and extends the analysis to Boltzmann machines. The core GEBM analysis is exact and is supported by extensive numerical tests. The final section attempts to generalize the framework to arbitrary energy-based models via neural tangent kernel (NTK) dynamics of score matching, and this extension is the main source of concern.","tokens_in":32054,"tokens_out":7503,"duration_ms":77932,"significance":"If the core results stand, which the numerical evidence strongly supports, the paper is a valuable contribution: the Lambert-W trajectory (Eq. 6) is verified exactly, the RMT predictions are parameter-free once the population spectrum and aspect ratio are fixed, the GCV-type relation (Eq. 9) is derived rather than fitted, and the experiments cover realistic covariance spectra as well as finite-N convergence. The Boltzmann-machine section is honestly presented as qualitative and provides useful insight. The NTK extension in Section 7 is not sound as written, but the paper's central GEBM/RMT/GCV contribution is independent of it and remains publishable after the extension is corrected or explicitly demoted to a conjecture.","major_comments":[{"comment":"The claimed derivation of j_t(x) = (1-e^{-xt})/x is incorrect. For a commuting initial condition, the scalarized version of dJ/dt = I - (\\hat CJ + J\\hat C) is dj/dt = 1 - 2xj, whose solution is (1-e^{-2xt})/(2x), not Eq. (14). The factor of two is not a harmless convention: it propagates into Eqs. (17)-(21) and into the closing statement that 'in the GEBM case we recover (14)'. Note also that Eq. (14) describes score-matching dynamics, which is a different training algorithm from the likelihood gradient ascent used in Sections 2-5; if it is intended only as a proxy, this must be stated explicitly. This issue is load-bearing for the abstract's final claim of 'deriving the neural tangent kernel dynamics'.","section":"Section 7, Eq. (14)"},{"comment":"The transition from the RKHS dynamics of the score function to the parameter update (21) is asserted rather than derived. The application of the function j_t to the empirical kernel matrix \\hat K in Eqs. (17) and (19) requires a spectral definition that is not given, and the 'parameter-sample duality' that produces Eq. (21) is not spelled out. As written, the section has the status of a heuristic proposal, not a derivation. The manuscript itself uses the word 'postulate' in Section 7 and the Discussion calls for further experimental investigation, so the authors should either supply the missing derivation or explicitly label this section as conjecture.","section":"Section 7, Eqs. (17)-(21)"},{"comment":"The abstract's final sentence ('deriving the neural tangent kernel dynamics of the score function') is stronger than what the body supports: Section 7 states that the score-matching approach is postulated, and Section 8 says the NTK extension 'deserves further experimental investigations'. Given the factor-of-two error in Eq. (14), the abstract overstates the result. The title-level claim of a framework for arbitrary EBMs should be scaled back to the pairwise/Gaussian cases, with the NTK part presented as a conjectural extension.","section":"Abstract vs. Section 7/Discussion"}],"minor_comments":[{"comment":"The notation \\tau v_\\alpha dv_\\beta/dt in Eq. (5) is ambiguous; it should be written as \\tau v_\\alpha^\\top dv_\\beta/dt to make clear that the left-hand side is the scalar projection of the eigenvector rotation.","section":"Section 2, Eq. (5)"},{"comment":"The derivation assumes non-symmetric perturbations of J, whereas the numerical training uses the symmetrized gradient (28). The authors note in Fig. 7 that the difference is small, but a short sentence in the main text indicating that this is a controlled approximation would help readers who only consult the main text.","section":"Appendix A, Eq. (28)"},{"comment":"The indicator notation 11_{(a,b)}^x is nonstandard and easy to misread; please use \\mathbf{1}_{(a,b)}(x).","section":"Section 3, Eq. (10)"},{"comment":"The term 'Wasserstrein distance' is misspelled and should be 'Wasserstein distance'.","section":"Appendix E"},{"comment":"There is a typo: 'reguavlarization' should be 'regularization'.","section":"Appendix H.1.1"},{"comment":"The treatment of eigenvector fluctuations is empirically convincing, but the conditions under which eigenvectors of \\hat C_M stay close to population eigenvectors (e.g., absence of level repulsion or eigenvalue crossings) are not stated. For trace-level observables this is not a blocker, yet a precise remark would strengthen the asymptotic claims at low \\rho.","section":"Section 4 / Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The GEBM/RMT/GCV core is strong and should be published after revision. The Section 7 factor-of-two error is significant but local; requiring either a corrected derivation or an explicit demotion of Section 7 to a conjecture is a reasonable condition for acceptance. I do not regard the eigenvector-misalignment concern as a blocker because the predicted observables are trace-level and are validated by finite-N numerics in Fig. 4."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the Gaussian-EBM core is the real content, and it is good. Exact spectral training dynamics (Lambert-W solution, relaxation times ∝ c^{-2}), the RMT finite-sample corrections, and the eL1 GCV analog E_test = E_train/(1 − E_train/ρ) (Eq. 9) are clean, parameter-free, and backed by numerics that actually match as N grows. The leave-one-out concentration argument for Eq. (9) is a genuine derivation, not a fit. The 'val-pop' experiment in Fig. 5a is a nice falsifiable check that the overfitting bump is eigenvalue-driven. This is the first exact spectral theory of overfitting for a canonical generative model that I know of, and the BM section honestly labels itself as qualitative.\n\nThe soft spots, in order:\n\n1. Section 7 has a real internal inconsistency. The stated score-matching ODE dJ/dt = I − (ĈJ + JĈ), scalarized for commuting initial conditions, solves to (1 − e^{−2ct})/(2c), not Eq. (14)'s (1 − e^{−ct})/c. That factor of two propagates into Eqs. (17)–(21), so the RKHS/NTK machinery built on it does not follow from the stated dynamics. The paper does call this section a postulate, but the abstract says the NTK dynamics are 'derived' — that overstates what is actually shown. Either fix the factor or publicly demote the section to conjecture.\n\n2. The BM extension (Eq. 12) is approximate by design — the authors drop the diagonal constraint and say so. Fine, but the abstract's phrase 'extends with minimal variations' is weaker than it sounds once you read the appendix.\n\n3. The eigenvector-noise assumption is the load-bearing idealization for the RMT corrections. Appendix B gives real support (eigenvector preservation plots, val-pop check), so I do not consider it a flaw, but it is empirically supported rather than proven — at low ρ or with closely spaced modes it could bite.\n\n4. No code or seed-level error bars. Minor for a theory paper, but the eigenvalue trajectories in Fig. 2b would benefit from it.\n\nWho this is for: people working on EBM training dynamics, inverse Ising, and RMT-based covariance cleaning. The core GEBM/RMT/GCV content is solid and does not have a load-bearing flaw. It deserves a serious referee — I would send it out, and I would require Section 7 to be fixed or demoted before publication.","headline":"The Gaussian-EBM/RMT/GCV core is a genuinely clean result that deserves a serious referee; the Section 7 NTK extension has a factor-of-2 inconsistency that must be fixed or demoted before the abstract's 'derived' claim can stand.","tokens_in":32557,"tokens_out":2952,"would_cite":true,"duration_ms":29699,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60B20","62H25","68T05"],"pacs":[],"model":"deepseek-v4-flash","headline":"For Gaussian energy-based models, finite-data overfitting is a predictable spectral effect: weak covariance modes are learned last, and their underestimated eigenvalues set an optimal early-stopping time that random matrix theory locates…","keywords":["energy-based models","overfitting","Gaussian model","random matrix theory","spectral bias","early stopping","generalized cross-validation","inverse Ising"],"falsifier":"Train a Gaussian energy-based model on data drawn from a population covariance with a degenerate or near-degenerate block of weak eigenvalues, or at aspect ratio $\\rho$ close to 1, and compare the measured non-monotonic reconstruction error and optimal stopping time against the random-matrix-theory prediction; a systematic shift or disappearance of the overfitting bump would show that eigenvector misalignment, not eigenvalue distortion alone, sets the overfitting timescale.","tokens_in":31440,"feed_emoji":"📉","tokens_out":6174,"duration_ms":61449,"temperature":0.7,"pith_summary":"This paper argues that overfitting in energy-based models is not a diffuse pathology but a predictable spectral phenomenon: with finite data, the learned coupling matrix relaxes mode by mode along the eigenvectors of the empirical covariance, strong modes converge quickly and weak modes slowly, and the slow modes are exactly the ones whose empirical eigenvalues are biased downward. In the Gaussian testbed this yields an analytic training trajectory in Lambert-W form, a closed-form relation between train and test energy under spectral L1 regularization, and random-matrix-theory formulas that locate the early-stopping optimum. If correct, this gives practitioners a principled way to set early-stopping times or shrinkage corrections for pairwise energy-based models, and it extends qualitatively to Boltzmann machines for inverse Ising problems.","feed_headline":"Weak data modes drive overfitting in energy models","feed_subtitle":"Exact early-stopping times and a test-error formula follow from the covariance spectrum.","key_machinery":"The load-bearing object is the eigendecomposition of the empirical covariance matrix $\\hat{C}_M$. Projecting gradient ascent onto that basis decouples the training dynamics into independent equations per mode; the exact solution uses the Lambert $W_0$ function for the Gaussian model, and the asymptotic form of the empirical spectrum is supplied by random matrix theory in the proportional limit $M/N = \\rho$. The relation $E_{\\text{test}} = E_{\\text{train}}/(1 - E_{\\text{train}}/\\rho)$ is derived by a leave-one-out argument and serves as the EBM analogue of generalized cross-validation. For Boltzmann machines the same decomposition is used with the mean-field correlation approximation $C = (I - J)^{-1}$, giving an approximate eigenvalue evolution with the same timescale separation.","core_discovery":"On the paper's own terms, the central discovery is that finite-sample overfitting in a Gaussian energy-based model is driven by the spectrum of the empirical covariance matrix, not by eigenvector corruption: each coupling eigenvalue evolves independently as $J_\\alpha(t) = \\frac{1}{\\hat{c}_\\alpha} + \\frac{1}{\\hat{c}_\\alpha} W_0\\left[B_\\alpha e^{-(\\hat{c}_\\alpha)^2 t/\\tau}\\right]$ with relaxation time proportional to $(\\hat{c}_\\alpha)^{-2}$, so stronger data modes are learned first and weaker, noisier modes later. Because finite $M$ underestimates the weak eigenvalues, the limiting couplings $1/\\hat{c}_\\alpha$ overshoot the true $1/c^*_\\alpha$; early in training the eigenvalues cross their ground-truth values, producing a non-monotonic reconstruction error and a well-defined optimal stopping time $t_{\\min}(\\rho)$ that matches asymptotic random matrix theory. For spectral-L1 regularization the same mechanism yields $E_{\\text{test}} = E_{\\text{train}} / (1 - E_{\\text{train}}/\\rho)$, a generalized-cross-validation analogue for EBMs. The paper further claims the same timescale structure appears in binary Boltzmann machines under a mean-field approximation, and sketches a score-matching neural-tangent-kernel route through which the eigenvalue picture should extend to general energy-based models.","pith_inferences":["Beyond the paper: monitoring the bulk edge of the empirical spectral density could serve as a practical early-warning signal for when weak-mode fitting begins in any pairwise energy-based model.","Beyond the paper: the train-test energy relation could be tested as a model-selection score on non-Gaussian pairwise models; if it holds approximately, it would give a held-out-free estimate of test log-likelihood for inverse problems.","Beyond the paper: a stress test with degenerate or nearly degenerate population eigenvalues would separate the eigenvalue-distortion mechanism from eigenvector-noise effects, since the idealized val-pop experiment removes only the eigenvalue distortion."],"forward_implications":["Early stopping can be selected from the data alone: random-matrix-theory formulas give $t_{\\min}(\\rho)$ and match finite-size training, although this time does not coincide with the peak of the test log-likelihood.","Shrinkage corrections based on eigenvalue cleaning, including a simple downsampling polynomial extrapolation, reduce the overfitting bump without requiring ground truth; the polynomial fit also works for the inverse Ising Boltzmann machine where rotationally invariant shrinkage is not applicable.","Generation quality stabilizes before the reconstruction-error minimum, so metrics based on generated samples can miss ongoing degradation of the inferred couplings.","The train-test energy relation for spectral-L1 regularization offers a way to estimate test log-likelihood without a test set, in the same spirit as generalized cross-validation for ridge regression.","The score-matching neural-tangent-kernel formulation predicts that generic energy-based models in the kernel or lazy regime follow a linear empirical-RKHS dynamics of the same form, making the finite-$M$ spectral mechanism the default explanation of overfitting there as well."],"supporting_citations":[{"why":"Supplies the asymptotic spectral density of the empirical covariance matrix used for the random-matrix-theory predictions of training and test energies.","marker":"Marčenko & Pastur, 1967"},{"why":"Gives the finite-$M$ eigenvalue fluctuations and eigenvector consistency results underpinning the eigenvalue corrections and the downsampling extrapolation.","marker":"Ledoit & Péché, 2011"},{"why":"Defines generalized cross-validation, whose train-test relation the paper adapts into its EBM analogue for spectral-L1 regularization.","marker":"Golub et al., 1979"},{"why":"Provides the rotationally invariant shrinkage estimators used to clean the empirical covariance matrix before training the Gaussian model.","marker":"Bun et al., 2017"},{"why":"Justifies the linear-in-$1/m$ scaling of downsampled eigenvalues on which the empirical polynomial-fit correction is based.","marker":"Baik & Silverstein, 2006"},{"why":"Documents the same qualitative non-monotonic reconstruction behavior in more complex energy-based models, motivating the Gaussian testbed as a faithful minimal case.","marker":"Agoritsas et al., 2023"}],"fun_headline_variants":["Spectral eigenmodes set early-stop time in energy models","Covariance spectrum predicts overfitting in energy-based learning","Weak data modes cause overfitting: early stopping from spectrum","Energy model overfitting traced to empirical covariance eigenvalues","Early-stopping formula from spectral decomposition of data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole analysis treats overfitting as caused by distortions in the eigenvalues of the empirical covariance, assuming its eigenvectors stay close enough to the true ones that eigenvector noise can be ignored; if modes are too close together or the sample size is too small, that assumption gives way.","fun_headline_variants_meta":{"raw":{"variants":["Spectral eigenmodes set early-stop time in energy models","Covariance spectrum predicts overfitting in energy-based learning","Weak data modes cause overfitting: early stopping from spectrum","Energy model overfitting traced to empirical covariance eigenvalues","Early-stopping formula from spectral decomposition of data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000515,"raw_usage":{"total_tokens":2528,"prompt_tokens":1002,"completion_tokens":1526,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":618,"completion_tokens_details":{"reasoning_tokens":1448}},"tokens_in":618,"tokens_out":1526,"duration_ms":9729,"temperature":1.0,"reasoning_tokens":1448,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T21:06:26.780522+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a Gaussian energy-based model on data drawn from a population covariance with a degenerate or near-degenerate block of weak eigenvalues, or at aspect ratio $\\rho$ close to 1, and compare the measured non-monotonic reconstruction error and optimal stopping time against the random-matrix-theory prediction; a systematic shift or disappearance of the overfitting bump would show that eigenvector misalignment, not eigenvalue distortion alone, sets the overfitting timescale.","supporting_citations":[{"cited_title":"and P \\'e ch \\'e , S","cited_arxiv_id":null,"evidence_quote":"Gives the finite-$M$ eigenvalue fluctuations and eigenvector consistency results underpinning the eigenvalue corrections and the downsampling extrapolation."},{"cited_title":"Generalized cross-validation as a method for choosing a good ridge parameter","cited_arxiv_id":null,"evidence_quote":"Defines generalized cross-validation, whose train-test relation the paper adapts into its EBM analogue for spectral-L1 regularization."},{"cited_title":"and Silverstein, J","cited_arxiv_id":null,"evidence_quote":"Justifies the linear-in-$1/m$ scaling of downsampled eigenvalues on which the empirical polynomial-fit correction is based."},{"cited_title":"Explaining the effects of non-convergent MCMC in the training of energy-based models","cited_arxiv_id":null,"evidence_quote":"Documents the same qualitative non-monotonic reconstruction behavior in more complex energy-based models, motivating the Gaussian testbed as a faithful minimal case."}],"review_version":1}