{"id":"84879df4-65da-41e6-a8a7-1add385f34cb","arxiv_id":"2505.04898","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Gradient descent iterates of finite-width multi-layer networks on single-index data obey a state evolution law, giving exact training/test error formulas and a data-driven test error estimator.","lead":"This paper proves exact statistical formulas for how gradient descent trains multi-layer neural networks with small width but large, high-dimensional data. It also provides an algorithm that estimates test error during training without knowing the true signal, useful for early stopping.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The main theorem depends on matrix-variate GFOM state evolution borrowed from an unreviewed preprint and stated without proof; an unverified hypothesis there would invalidate the central claim.","rationale":"I read the paper as an honest conditional contribution: under Assumptions A1-A5 it claims a precise state-evolution characterization, and the readable portions of the proof appear internally coherent. The simulations support the estimator within the stated regime and the authors openly document the C^4 restriction, the ReLU instability, and the existential constants. The reader's weakest assumption, (A4), is a genuine scope limitation, but it is explicitly stated and does not threaten the correctness of the conditional theorem. My concern is different: the proof of Theorem 3.2 leans on matrix-variate GFOM state evolution imported from an unreviewed preprint and stated in Appendix B without proof. A single unverified hypothesis in that imported machinery would collapse the main theorem and its corollaries. The reader listed this as a conditional point but did not make it the weakest assumption; I regard it as more load-bearing. Because the concern is a verification gap rather than a demonstrated error, the correct disposition is to keep the reader's CONDITIONAL verdict while requiring independent verification of Theorems B.6-B.7 before the central claim can be regarded as secure.","tokens_in":85683,"tokens_out":4963,"duration_ms":56981,"concrete_test":"Independently prove Theorem B.6/B.7 in the minimal case q=1 (scalar rows), with polynomial F,G and m=n, using a self-contained leave-one-out or Lindeberg argument that does not cite [Han24]; then check that the truncated functions S_t and T_{alpha;t} defined in (9.1) satisfy the C^3 bounds and that the delocalization Proposition B.8 supplies the l-infinity control used in Lemma 9.5. If the q=1 proof fails for sub-Gaussian, non-Gaussian entries, or requires a fourth-moment condition absent from Assumption (A2), the central theorem is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of Theorem 3.2 is obtained by converting gradient descent into an auxiliary GFOM (Section 7.2) and then applying the entrywise matrix-variate GFOM state evolution Theorems B.6-B.7. These two theorems are stated in Appendix B without proofs and are imported from [Han24] 'in its full strength.' Proposition 7.1, the bridge between the auxiliary GFOM and state evolution, is exactly an application of these theorems, and Proposition 7.2 then controls the GD-to-auxiliary error. If Theorems B.6-B.7 have an unstated hypothesis that the update functions in (7.5) do not satisfy, or if their proof requires Gaussian features rather than the claimed sub-Gaussian entries, the chain from Proposition 7.1 through Theorem 3.2, Theorem 4.2, Theorem 4.3, and Theorem 5.2 does not hold. The paper's own Section 7 flags that the auxiliary GFOM functions G<2>_t are not globally Lipschitz and must be truncated and smoothed in (9.1); Lemmas 9.3-9.5 supply the needed controls, but only assuming the imported theorems are correct. This is not a scope caveat; it is a correctness risk at the single most load-bearing juncture.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a non-asymptotic state-evolution theory for the iterates of vanilla gradient descent on bounded-width, bounded-depth multilayer networks in the proportional regime m/n ≍ 1, under single-index data with sub-Gaussian features and smooth activations. The main result (Theorem 3.2) states that, in an averaged empirical-row sense, the first-layer weights n^{1/2}W_1^{(t)} behave like Gaussian vectors whose law is defined recursively by a state-evolution procedure, while deeper-layer weights concentrate around deterministic matrices; the error is (KΛκ*)^{c_t} n^{-1/c_t}. The framework yields: characterizations of training and test errors as low-dimensional Gaussian integrals (Theorem 4.2); a link- and signal-agnostic online estimator of the test error (Algorithm 1, proved consistent in Theorem 4.3); and a structural result showing the learned function remains effectively single-index with an effective signal combining the true signal and the initialization (Theorem 5.2). The proof (Sections 7-13) reduces GD to an auxiliary GFOM whose state evolution follows from matrix-variate GFOM theorems imported from the preprint [Han24] and stated in Appendix B. Simulations validate the estimator and, unusually, document its breakdown for wide networks and ReLU activations (Section 6.4).","tokens_in":85844,"tokens_out":23636,"duration_ms":226866,"significance":"If correct, this is a substantial contribution relative to the NTK, mean-field, and tensor-program lines: it provides the first precise distributional characterization of GD in the finite-width proportional regime, allows weight movement away from initialization, covers arbitrary depth at bounded width, and yields a test-error estimator requiring no knowledge of the link function or signal. The manuscript's strengths include complete-looking proofs for the main theorems with explicit non-asymptotic error rates, an estimator whose construction is genuinely agnostic to φ* and μ* (Section 4.3), a structural single-index representation of the learned model (Theorem 5.2), and unusually honest empirical reporting: Section 6.4 documents the instability for ReLU and the deviations for wide networks, and Remarks 4(2) and 6 state the regularity conjectures explicitly.","major_comments":[{"comment":"The central proof chain rests on Theorems B.6 and B.7, the matrix-variate GFOM state-evolution results stated in Appendix B and imported 'in its full strength' from the preprint [Han24] (Section 1.5). Proposition 7.1 is exactly an application of these theorems to the auxiliary GFOM (7.7), and Proposition 7.2 then transfers the resulting characterization to the GD iterates (7.3); Theorem 3.2 and its consequences (Theorems 4.2, 4.3, 5.2) therefore inherit any gap in B.6-B.7. Since B.6-B.7 are stated without proofs and [Han24] is an unreviewed preprint, the referee cannot verify from the manuscript that the hypotheses of these theorems hold for the application at hand. Please either include complete proofs of Theorems B.2, B.3, B.6, B.7 and Propositions B.4, B.8, or state explicitly which peer-reviewed results cover the asymmetric entrywise case needed here and verify their hypotheses line by line.","section":"Appendix B (Theorems B.6-B.7); Sections 7.2-7.3; Section 1.5"},{"comment":"Even granting Theorems B.6-B.7, the manuscript does not verify that the truncated GFOM (9.2) satisfies hypothesis (D*2). The update functions in (9.2) contain ˇG^{(2)}_t ≡ −φ^{-1}η^{(t−1)}_1·χ_{b_t}(S(·,·,·)) with the truncation (9.1). Value truncation of S does not make the derivatives of the composition bounded: by Proposition 8.3-(7), ∂_{u_kℓ}S grows polynomially in ∥u∥ and ∥φ_*(w)∥, so ∂(χ_{b_t}∘S) = χ'_{b_t}(S)·∂S is unbounded as a function on R^{q[0:t]}, and the same holds a fortiori for the higher derivatives. Concretely, for L = 2, q = 1, φ* ≡ 0, ξ = 0, and σ_1(x) = x + 2 sin x (which satisfies (A4) with a sufficiently large Λ), S(u) = σ(u)σ'(u) vanishes at points where σ'(u) = 0 while S'(u) grows linearly, so ∂(χ_b∘S) is unbounded. Hence the global C^3 bound required in (D*2) (and similarly (D*2)') is not established, and Lemma 9.3's appeal to the meta GFOM theory is not justified as written. The gap is repairable, for instance by also truncating the arguments u, w, or by proving a version of the GFOM theorem under polynomial-growth derivative bounds whose log n factors would be absorbed by the (KΛκ* log n)^{ct} pre-factors already present in Lemmas 9.3-9.5, but as it stands the step from 'pseudo-Lipschitz plus delocalization' (Section 7.3.1) to 'application of B.6' is missing. This is the single most load-bearing juncture of the paper and needs to be closed.","section":"Section 9.2, Eq. (9.1), Lemma 9.3"},{"comment":"Assumption (A4), which requires activations in C^4 with bounded derivatives up to order 4, is load-bearing in three structurally different places: the derivative bounds for the theoretical gradient maps that define the Onsager corrections (Propositions 8.2-8.3), the Lindeberg-principle step replacing test inputs by Gaussians (Lemma 11.1 and Lemma 11.2 in Section 11.1, which need third derivatives of the network map), and Algorithm 1 itself, whose step (3) computes σ''_1. The assumption excludes ReLU; the paper's own simulation (Figure 4, left) shows that setting σ'' ≡ 0 makes the estimator unstable and drifting, and the minimal-regularity conjecture in Remarks 4(2) and 6 (Lipschitz σ') is not proved. The paper is admirably candid about this, but the abstract and introduction advertise 'general multi-layer neural networks' consistent with 'practical architectures' (Section 1.1, Table 1); the reader should be told at the abstract level that the main results require smooth activations and that the algorithmic estimator needs user-supplied second derivatives.","section":"Assumption (A4), Section 3.1; Remarks 4(2) and 6; Section 6.4"}],"minor_comments":[{"comment":"The statement of Theorem 3.4 lists only (A3)-(A5) and ϕ^{-1} ≤ K, but the proof (Step 5 of Section 10.3) invokes Theorem 3.2, which requires the full Assumption A; please state the full hypotheses and the intended limiting order (first m ∧ n → ∞ at fixed ϕ, then ϕ → ∞).","section":"Theorem 3.4; Section 10.3"},{"comment":"The assertion that the generalization gap satisfies E(0) Gap^{(t)}(X, Y) = Θ(1) is made without proof; since both terms in the gap are expectations over state-evolution variables that depend on t in a coupled way, either supply a one-line justification or explicitly label the claim as a heuristic.","section":"Section 4.1, after Theorem 4.2"},{"comment":"Because the dependence of c_t on t, q, L is not tracked, the rates (KΛκ*)^{ct} n^{-1/ct} are not quantitative at the values of n used in the simulations; the numerical validation in Section 6 is consequently the only quantitative evidence at practical sample sizes, and the paper should say so where the rates are advertised.","section":"Remark 4(1); Section 6"},{"comment":"There are two obvious typos: 'converegence' and 'arises ubiquitously is this regime'.","section":"Section 1.2"},{"comment":"The explanation 'the apparently smaller update observed when L = 5 is likely due to instability in the fifth-layer updates' is unclear; if the intended meaning is that the fifth-layer relative update is small because ∥W_5^{(0)}∥ is large and the update is noisy, please say so explicitly.","section":"Section 6.2, bullet (3)"},{"comment":"The claim that the estimator requires only 'a small number of closed-form computations' (Section 1.3) should be quantified: the forward-derivative tensors ∂_ℓ Ĥ^{(t-1)}_α and ∂^{(b)}_ℓ P̂^{(t-1)}_α add an O(q) factor per layer over backpropagation, which is material for larger widths.","section":"Algorithm 1, Steps (1)-(3)"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to have a high impact if the proof chain is closed. The principal risk is the dependence on [Han24], an unreviewed preprint by the first author; I would encourage the editor to require the proofs of the GFOM meta-theorems in Appendix B as part of the final version, so that the published record does not depend on an unverified preprint. The honest reporting in Section 6.4 and the explicit limitations in Remarks 4 and 6 are strengths that should be preserved. One point not to lose in revision: the asymmetry between the theorem's smoothness assumption (C^4) and the estimator's practical requirement (σ'') should be stated in the abstract."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a serious, honest, unusually complete theory paper — over 80 pages of proof, explicit assumptions, and simulations that report their own breakdowns. If the central theorem is right, it is a genuine within-field advance. The thing I'd want resolved before fully trusting it is the missing proof of the matrix-variate GFOM state evolution theorems (Appendix B.6-B.7), imported without proof from the same author's unreviewed preprint [Han24] and load-bearing for the whole chain.\n\nWhat's actually new: the first precise distributional characterization of gradient descent iterates for multi-layer networks in the finite-width proportional regime, with O(1) learning rates and non-Gaussian features, where weights move non-trivially from initialization. Prior precise characterizations were two-layer or infinite-width. The iterative reduction of GD to auxiliary GFOM iterates is a real technical contribution, and the derivative machinery (Propositions 8.2-8.3) driving the Onsager corrections is substantial. The payoffs are genuine: train/test error formulas and a test-error estimator that never needs the link function or signal, validated numerically for tanh at L=2,3,5 with honest failure reports for wide networks and ReLU.\n\nSoft spots, in proportion. First, the imported theorems. The stress-test worry — that B.6-B.7 might have silent hypotheses the updates (7.5) violate — does not land hard on reading the paper: Section 9 does real truncation and smoothing work (Lemmas 9.3-9.5) that puts the non-Lipschitz functions in the stated framework. What does land is the verification gap: an unreviewed same-author preprint carries the weight, and the paper says so ('in its full strength'). A referee should ask for those proofs or independent verification.\n\nSecond, C^4 activations. Load-bearing, excludes ReLU, and Algorithm 1 explicitly computes σ''. The paper is honest about this — Remark 6 conjectures the right minimal regularity, and Figure 4 shows the estimator destabilizing on ReLU. It limits practical reach but is not hidden.\n\nThird, the constants c_t are existential, so no concrete rate is available for fixed (t, q, L). The authors flag this too.\n\nI don't see an internal inconsistency in the readable portions, and the central argument holds up within its stated regime. The reader's conditional verdict is about right, and this paper deserves serious peer review — the B.6-B.7 question is exactly what referees should resolve. Recommended: accept for review, not desk reject.","headline":"A serious, honest, complete-looking theory paper, but the central claim rests on unproved GFOM state evolution theorems imported from the same author's unreviewed preprint; send it to referees and make that the first question.","tokens_in":86479,"tokens_out":7239,"would_cite":true,"duration_ms":69232,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60E15","60G15"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper establishes a per-iteration state-evolution law for vanilla gradient descent on finite-width multilayer networks, showing first-layer weights become Gaussian-fluctuating while deeper layers concentrate, and gives a consistent…","keywords":["gradient descent dynamics","multi-layer neural networks","state evolution","finite-width proportional regime","single-index regression","generalization error estimation","general first order methods","non-asymptotic theory"],"falsifier":"Fix a smooth activation such as sigmoid, width q=10, m/n≈1/2, and run vanilla gradient descent with learning rate 2 on single-index data with sub-gaussian features; compare the empirical row distribution of $n^{{1/2}}$$W_1^{{(t)}}$ with the state-evolution Gaussian for t=2 and t=5 across n=600, 2400, and 9600. If the pseudo-Lipschitz discrepancy does not decay to zero at the stated polynomial rate, the central claim is false; separately, for ReLU with the second derivative set to zero, the paper's Figure 4 already shows the test-error estimate drifting from the truth.","tokens_in":85351,"feed_emoji":"🧠","tokens_out":4651,"duration_ms":48461,"temperature":0.7,"pith_summary":"This paper tries to establish a precise, per-iteration law for what vanilla gradient descent does to a multilayer neural network when sample size and feature dimension grow together while width and depth stay bounded. If true, it replaces vague narratives about feature learning with a quantitative description: the first-layer weights fluctuate as Gaussians whose covariance is dictated by a state-evolution recursion, the deeper-layer weights concentrate on deterministic matrices, and both training and test error are given by low-dimensional Gaussian integrals. Because the description holds for non-Gaussian features and for individual initializations, it covers behavior that infinite-width theories cannot see, including nontrivial movement away from initialization. The practical payoff is an augmented gradient-descent routine that, at every iteration, outputs a consistent estimate of the model's test error without knowing the signal or link function, which could guide early stopping.","feed_headline":"Gradient descent dynamics pinned down for finite-width networks","feed_subtitle":"A state-evolution law tracks weights, training and test error at every step, with a consistent generalization estimator.","key_machinery":"The engine is an iterative reduction scheme: gradient descent is shown, step by step, to be indistinguishable from a sequence of matrix-variate 'general first order methods' (GFOMs), iterative schemes in which each update is a row-wise nonlinear function of previous iterates plus a matrix multiplication by the data matrix. The GFOM iterates admit an entrywise, non-asymptotic Gaussian state evolution, and the debiasing coefficients connecting the two are matrix-valued Onsager correction matrices that absorb the correlations between successive pre-activations and restore approximate Gaussianity.","core_discovery":"For every iteration t of vanilla gradient descent under proportional sample size and feature dimension, the row-wise empirical distribution of the scaled first-layer weights $n^{{1/2}}$$W_1^{{(t)}}$ is approximated, in a pseudo-Lipschitz mean with error (KΛκ*)^{c_t} $n^{{-1/c_t}}$, by a Gaussian law produced by a state-evolution recursion; the deeper-layer weights concentrate around deterministic matrices $V_α^{{(t)}}$ with the same order of error. The state evolution combines Onsager correction matrices that debias the first-layer pre-activations, deterministic updates depending smoothly on the initialization, and a Gaussian component encoding the high-dimensional noise. From this, the paper derives that both training error and test error are, up to the same error rate, averages of low-dimensional Gaussian integrals over deterministic residual maps; that these integrals reveal a non-vanishing generalization gap in the proportional regime; and that the trained network remains a single-index function whose effective signal is a linear combination of the true signal and the initialization. The paper further shows that these Gaussian integrals can be computed online by augmenting gradient descent with closed-form correction estimates, yielding a consistent estimator of test error that requires no knowledge of the link function or signal.","pith_inferences":["Editorial inference: the machinery, though demonstrated for single-index data and full-batch gradient descent, is structurally modular; the same reduction should extend to multi-index signals and to stochastic gradient descent with proportionally large mini-batches, as the paper notes are omitted formal extensions.","Editorial inference: the algorithmic estimator should be checked on activations that are only once or twice differentiable; the paper's simulations with a smoothed ReLU already support the conjecture that Lipschitz first derivatives may suffice, while ReLU with a hard kink requires genuinely new correction terms rather than simply setting the second derivative to zero.","Editorial inference: the single-index representation of the learned model gives a concrete falsifiable prediction: after training on single-index data, the network's output should depend on a test input only through its projection onto the span of the true signal and the initialization rows, up to the stated Gaussian error.","Editorial inference: the state-evolution error bound worsens with iteration count t, so practical guarantees are best for early stopping; whether the accumulated error can be controlled uniformly over all t is a natural next question."],"forward_implications":["Training and test error can be evaluated to leading order as means over low-dimensional Gaussian vectors involving deterministic residual maps, so no simulation of the full network is needed to predict either quantity.","The generalization gap is generically non-vanishing in the proportional regime and shrinks as the sample-to-dimension ratio becomes large.","A consistent estimator of test error can be computed at every gradient-descent iteration without algorithmic convergence and without knowing the link function or the underlying signal, enabling principled early stopping and hyperparameter tuning.","The learned model retains the structure of a single-index function, with an effective signal that combines the true signal and the initialization and a Gaussian component that vanishes as the sample size grows relative to dimension.","The distributional law applies to non-Gaussian feature designs satisfying sub-gaussian tail conditions, not only to Gaussian data."],"supporting_citations":[{"why":"Supplies the matrix-variate, entrywise, non-asymptotic GFOM state evolution theory that the iterative reduction invokes.","marker":"[Han24]"},{"why":"Provides the conceptual blueprint for augmenting iterative algorithms with consistent online estimates of generalization error.","marker":"[HX24]"},{"why":"Supplies the square-root estimate for covariance matrices used when comparing Gaussian laws in the test-error and structural proofs.","marker":"[BHX23]"},{"why":"Supplies the Lindeberg principle that replaces independent sub-gaussian test inputs by Gaussian ones in the test-error argument.","marker":"[Cha06]"}],"fun_headline_variants":["Finite-width GD: Gaussian law per step","Precise GD iterates for finite-width nets","State evolution pins GD dynamics at finite width","GD yields consistent test-error estimator"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument, including the algorithm, requires every activation to be four times differentiable with bounded derivatives and the link function to be three times differentiable; ordinary ReLU networks are therefore excluded, and the paper's own simulations show the estimator becoming unstable for ReLU.","fun_headline_variants_meta":{"raw":{"variants":["Finite-width GD: Gaussian law per step","Precise GD iterates for finite-width nets","State evolution pins GD dynamics at finite width","GD yields consistent test-error estimator"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000878,"raw_usage":{"total_tokens":3857,"prompt_tokens":1063,"completion_tokens":2794,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":679,"completion_tokens_details":{"reasoning_tokens":2738}},"tokens_in":679,"tokens_out":2794,"duration_ms":19155,"temperature":1.0,"reasoning_tokens":2738,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:20:44.198921+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fix a smooth activation such as sigmoid, width q=10, m/n≈1/2, and run vanilla gradient descent with learning rate 2 on single-index data with sub-gaussian features; compare the empirical row distribution of $n^{{1/2}}$$W_1^{{(t)}}$ with the state-evolution Gaussian for t=2 and t=5 across n=600, 2400, and 9600. If the pseudo-Lipschitz discrepancy does not decay to zero at the stated polynomial rate, the central claim is false; separately, for ReLU with the second derivative set to zero, the paper's Figure 4 already shows the test-error estimate drifting from the truth.","supporting_citations":[],"review_version":1}