{"id":"6a93a93b-1639-4557-bfc0-596f9e3d6c95","arxiv_id":"1908.10030","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"For finite single-hidden-layer networks with symmetric weight initialization, the output distribution is a Gaussian with an O(1/N) fourth-Hermite correction, an instance of the classical Edgeworth expansion.","lead":"This paper shows that the outputs of a large but finite neural network at initialization are not exactly Gaussian, but close: a Gaussian plus a small correction that shrinks as the network gets wider. The finding offers a concrete way to quantify when the popular infinite-width Gaussian process model of neural networks is a good approximation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The marginal Edgeworth correction in Eq. 18 is plausible, but the paper's process-level NNGP claim rests on an unproved multivariate extension (footnote 3) that is never tested; the function-space conclusion is unsupported.","rationale":"The reader's weakest-assumption analysis identifies footnote 3, and I agree that this is the load-bearing gap for the paper's broader claim. Eq. 18, read as a purely marginal statement about the density of the output at a fixed input, is a standard Edgeworth expansion for a sum of i.i.d. terms and is not the source of difficulty: the RG derivation reproduces the known CLT/Edgeworth structure, and the experiments confirm the H4 shape and the approximate 1/N scaling. What the paper does not establish is the step from marginals to a stochastic process. A Gaussian process is characterized by joint distributions, and the paper's own footnote concedes that the multivariate Edgeworth extension is merely expected, not derived. Consequently, the conclusions in §4 about the usefulness of the Gaussian-process framework for finite networks, and the interpretation of how the perturbation behaves under training, go beyond what the mathematics and the experiments (which use a single input, x = 1) can support. I considered whether the fitted c4 rather than a predicted coefficient is the stronger weakness; it is a real limitation, but it does not threaten the mathematical claim of Eq. 18, which only determines the shape and scaling, leaving c4 as an input-dependent constant. The multivariate gap is therefore the single most load-bearing concern, and the conditional verdict should stand unchanged.","tokens_in":7326,"tokens_out":15431,"duration_ms":171570,"concrete_test":"Generate an ensemble of 10^8 single-hidden-layer networks (N = 128, ReLU, Glorot, as in §3.1) and record outputs at two inputs, x1 = 1 and x2 = 2. Estimate the empirical bivariate CDF or characteristic function and compare it with a bivariate Gaussian using the NNGP kernel plus the O(1/N) multivariate Edgeworth correction constructed from the single-unit joint fourth cumulants, including cross terms among the two outputs. If the observed joint correction deviates from this prediction at the level of the marginal correction shown in Fig. 1, the unproven footnote-3 step fails and the function-space claim is not established; if it matches, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. 18 and the §2 renormalization-group derivation concern the marginal density of the output at one fixed input. The abstract and §4 frame the result as a correction to a neural-network Gaussian process, i.e., a statement about a distribution over functions. A Gaussian process is determined by all finite-dimensional joint distributions, not by its one-point marginals. The only bridge from the marginal calculation to the process-level claim is footnote 3, which states that the joint distribution of outputs at multiple inputs is 'expected' to be a multivariate Edgeworth expansion, but explicitly says it is not shown there. Because a Gaussian-plus-H4 marginal at each x does not imply the joint output distribution is Gaussian-plus-multivariate-Edgeworth, the function-space conclusion in §4 is unsupported by the derivation or by the experiments in §3, which evaluate only x = 1 and fit c4 rather than predict it. Cross-input cumulants could in principle enter at the same 1/N order as the marginal H4 correction, so the multivariate step is load-bearing for the paper's broader NNGP framing. The marginal Eq. 18 itself is a standard i.i.d. Edgeworth result and is likely correct; the conditional verdict is appropriate, but the condition should be the unproven joint behavior, not the single-input formula.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper derives a finite-width correction to the Gaussian-process limit of a single-hidden-layer neural network at initialization. Using a renormalization-group argument for the single-input marginal distribution of the output y = Σ_i v_i f(u_i x), it obtains an Edgeworth expansion whose leading term for symmetric weight distributions is a fourth-Hermite perturbation of size c4(x)/N (Eqs. 16–18). It then reports empirical support: a comparison of empirical and predicted CDF differences at N = 128 (Sec. 3.1) and a power-law fit α ∝ N^{−1.07} (Sec. 3.2), and discusses consequences for when the NNGP framework remains valid during training.","tokens_in":7558,"tokens_out":10448,"duration_ms":110393,"significance":"The paper has a clear strength: the renormalization-group derivation transparently yields the scaling N^{1−n/2} for the Edgeworth corrections, and the predicted 1/N scaling of the leading non-Gaussianity is falsifiable. The observation that the Gaussian covariance is width-independent is also useful context. However, the process-level NNGP interpretation is not established: footnote 3 concedes that the multivariate Edgeworth expansion is not shown, and the numerical validation fits rather than predicts the only amplitude c4. If the paper is read as a claim about single-input marginals, it is a modest but plausible contribution; if it is read as a claim about Gaussian-process corrections, the current evidence is incomplete.","major_comments":[{"comment":"The displayed linearized eigen-equation is not the correct linearization of the renormalization operator R. For the fixed-point characteristic function g(k) = p*~(k), linearizing R gives λ_n φ~_n(k) = 2 g(k/√2) φ~_n(k/√2), and substituting φ~_n(k) = (ik)^n g(k) yields λ_n = 2^{1−n/2}. As written, Eq. (11) has an incorrect prefactor (1/(λ_n σ))√(2/π) and an exponential e^{−k²/(4σ)} that is inconsistent with the Gaussian fixed point of Eq. (9); a reader cannot reproduce Eq. (14) from the displayed equation. Please correct the equation and clarify whether σ denotes the standard deviation or the variance.","section":"2.4, Eq. (11)"},{"comment":"The paper repeatedly frames Eq. (17) as an Edgeworth expansion of a stochastic process, for instance by calling it 'the expansion of a stochastic process rather than a probability distribution.' But Eq. (17) is a statement about the marginal density of the output at one fixed input x. A stochastic process is determined by all finite-dimensional joint distributions, and footnote 3 explicitly says that the multivariate Edgeworth expansion for two or more inputs is expected but not shown. Cross-input joint cumulants could in principle enter at the same O(1/N) order as the marginal H4 term, so the process-level conclusions in Sec. 4 are not supported by the derivation or by the experiments, which evaluate only a single input. This gap is load-bearing for the NNGP framing; the authors should either prove or test the multivariate extension, or restrict their claims to single-input marginals.","section":"2.4, footnote 3; Sec. 4"},{"comment":"The constant c4 ≈ 9.405 is fitted from the same empirical CDF that is then compared with the predicted CDF in Fig. 1. Consequently the 'excellent agreement' in Fig. 1 is partly by construction: it validates the functional form but not the amplitude of the correction. To make the confirmation non-circular, the authors should either compute c4 from the weight distribution and moments, or fit c4 on one subset of the ensemble and compare with an independent subset. Without this, Sec. 3.1 does not provide independent evidence for Eq. (18).","section":"3.1, Eq. (19)"},{"comment":"The scaling test fits α separately for each N and then fits a power law α ∝ N^{−1.07}, but no error bars or confidence intervals are reported, and the fitting procedure is not specified. Moreover, under Glorot initialization the input and output weight variances change as N varies, so c4(x) may itself depend on N; the observed exponent is therefore not a direct test of the N^{−1} prediction of Eq. (18) unless c4 is shown to be N-independent or its N-dependence is accounted for. Please provide uncertainty estimates and, if possible, a direct prediction of α(N) for the exact initialization used.","section":"3.2, Eq. (20)"}],"minor_comments":[{"comment":"The notation for σ is inconsistent: the text calls σ the standard deviation, but Eq. (9) writes the Gaussian Fourier transform with e^{−k²/(2σ²)}, which is not the characteristic function of N(0,σ²) under the usual convention. Please reconcile the notation and the exponential factors.","section":"2.3, Eq. (9)"},{"comment":"The authors note that the small-N behavior is 'slightly steeper' than N^{−1}, but there is no discussion of where the asymptotic Edgeworth expansion is expected to break down. Reporting the residuals or a goodness-of-fit measure would help the reader assess whether the deviations are consistent with finite-N corrections.","section":"3.2"},{"comment":"The text does not specify how α is fitted to each empirical CDF difference, for example whether the fit is by least squares on a grid of y values and over what range of y. This information is needed for reproducibility.","section":"3.2, Eq. (20)"},{"comment":"Minor typographical issues include 'Tensorﬂow' with a ligature in the references; the manuscript would benefit from a final proofreading pass.","section":"1"}],"recommendation":"major_revision","confidential_remarks":"The single-input Edgeworth result is likely correct and the RG derivation, once Eq. (11) is fixed, gives the right scaling. The main reservations are the unproved multivariate extension, which is load-bearing for the paper's NNGP framing, and the partly circular empirical validation of c4. These are addressable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Joe — quick read of Antognini 1908.10030. The useful core: for a single hidden layer at init, symmetric weights, the output marginal is Gaussian plus an O(1/N) fourth-Hermite correction, and the paper shows this via a clean RG argument that recovers the classical Edgeworth expansion. That is a nice, explicit finite-width result and worth having. The RG eigenvalues are correct, Eq. 17/18 are standard, and the paper says so honestly. Credit where due: no overclaiming of novelty, the derivation is readable, and the asymptotic statement is properly qualified.\n\nWhere it gets soft. First, the abstract and Section 4 frame this as a correction to a neural network Gaussian process, i.e. a distribution over functions. The derivation only gives the one-point marginal. Footnote 3 punts on the joint distribution, but the joint is what defines a Gaussian process. Without it, the function-space conclusion is unsupported. Cross-input cumulants could enter at the same 1/N order, so this is not a minor gap in exposition; it is the difference between a marginal finite-width statement and an NNGP statement. Second, the empirics are weaker than the derivation. In section 3.1, c4 is fitted from the same CDF it is then compared to; section 3.2 fits alpha separately per width and also fits the exponent, reporting N^-1.07 rather than testing N^-1. No error bars anywhere, and only one architecture/initializer/input x=1. These are addressable, but they mean the experiments don't independently confirm the theory; they just show consistency with a fitted amplitude.\n\nOverall: the marginal result is likely correct, the process-level claim is not established, and the experiments overstate. The paper deserves a serious referee — it is clearly written and the asymptotic structure is worth checking — but the right verdict is conditional, with the condition being joint behavior, not the one-point formula. If you cite it, cite it for the marginal Edgeworth correction, not for a rigorous NNGP finite-width theory.","headline":"Marginal Edgeworth correction is right and clearly derived; process-level NNGP claim rests on an unproven multivariate extension, so treat as a marginal result.","tokens_in":8121,"tokens_out":1694,"would_cite":false,"duration_ms":17381,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60F05","62E20","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"For finite-width neural networks at initialization, the output distribution is a Gaussian plus a fourth-Hermite correction that shrinks as one over the hidden-layer width.","keywords":["neural network Gaussian process","finite width corrections","Edgeworth expansion","Hermite polynomials","renormalization group","single hidden layer","output distribution at initialization","Gaussian process limit"],"falsifier":"Measure the joint statistics of outputs at two or more inputs for a large ensemble of finite-width networks, for example the connected four-point cumulant $\\langle y(x_1)y(x_2)y(x_3)y(x_4)\\rangle_c$, and compare it with the multivariate Edgeworth prediction whose leading term is set by the same $c_4(x)/N$ amplitudes. If the joint distribution deviates from the multivariate Gaussian in a way not captured by that expansion, or with a different dependence on $N$, the function-space Gaussian process approximation would fail even though the single-input marginal is well fitted.","tokens_in":7070,"feed_emoji":"🧠","tokens_out":6596,"duration_ms":56265,"temperature":0.7,"pith_summary":"This paper attempts to characterize how far a finite-width neural network at initialization is from the Gaussian process it becomes at infinite width. For a single-hidden-layer network whose weights are drawn from a symmetric distribution, the paper derives that the output distribution is approximately a Gaussian perturbed by the fourth Hermite polynomial, with the perturbation amplitude proportional to $c_4(x)/N$. The derivation uses a renormalization-group linearization around the Gaussian fixed point and recovers an Edgeworth expansion whose higher-order terms decay faster with $N$. If correct, this gives a quantitative handle on when the neural network Gaussian process framework is a faithful model for practical, finite networks, and it identifies where to look for deviations during training.","feed_headline":"Finite neural nets deviate from Gaussian at order 1/N","feed_subtitle":"For symmetric weight initializations, the deviation is a fourth-Hermite polynomial whose size is set by hidden width.","key_machinery":"The engine of the derivation is a renormalization-group transformation acting on the distribution of the products $h_i v_i$ of hidden activations and output weights. Coarse-graining pairs hidden units and averages them, which is a convolution in Fourier space, and renormalization rescales the result to hold variance fixed; the Gaussian is the fixed point of this transformation. Linearizing around the fixed point yields eigenfunctions $\\varphi_n(y) = H_n(y)\\,\\mathcal{N}(y;0,\\sigma^2)$ with eigenvalues $\\lambda_n = 2^{1-n/2}$, so after $\\log_2 N$ coarse-grainings the $n$-th eigenfunction is suppressed by $N^{1-n/2}$. The fourth Hermite polynomial, $H_4(y) = (y/\\sigma)^4 - 6(y/\\sigma)^2 + 3$, is the first surviving irrelevant direction when the symmetric initial distribution kills the skew term, and it carries the $1/N$ correction.","core_discovery":"The central claim is stated in Eq. 18: for an ensemble of large finite single-hidden-layer networks with symmetric weight initialization, the probability density of outputs at a fixed input is $p_y(y) \\simeq \\mathcal{N}(y;0,\\sigma^2)\\bigl[1 - \\frac{c_4(x)}{N}(3 - 6(y/\\sigma)^2 + (y/\\sigma)^4)\\bigr]$. The skew term that would appear at order $N^{-1/2}$ vanishes because the symmetric initialization forces $c_3=0$, so the leading deviation from Gaussianity is the fourth-Hermite term at order $1/N$. The paper verifies this against ensembles of up to $10^8$ networks with ReLU activations, obtaining good agreement and measuring the perturbation amplitude to scale roughly as $N^{-1.07}$ rather than the predicted $N^{-1}$. This is an asymptotic Edgeworth expansion, so for fixed order the correction decreases as $N\\to\\infty$, while the full series does not converge.","pith_inferences":["An immediate, testable consequence of the paper's own single-input result is that the joint distribution at several inputs should show $O(1/N)$ fourth-order cross-cumulants whose coefficients are products and sums of the same $c_4(x_i)$; measuring those would test the unproven multivariate step.","Because the renormalization-group argument only uses finite variance and symmetry of the weight distribution, the same leading $H_4$ structure should appear for other activations and other symmetric initializations, with only $c_4(x)$ changing.","The input-dependent $c_4(x)$ suggests that the finite-width correction acts like a small non-Gaussian field over input space, which could matter for uncertainty estimates and Bayesian model comparison even when predictions are accurate.","Tracking $c_4$ through training would turn the paper's qualitative question into a quantitative one: a measurement of whether the perturbation amplitude grows, shrinks, or stays fixed would determine whether the infinite-width limit is an attractor for training dynamics."],"forward_implications":["For symmetric initializations, the output distribution of a single-hidden-layer network is never exactly Gaussian at finite width; the deviation is $O(1/N)$ and explicitly given by the fourth Hermite polynomial.","The infinite-width Gaussian process description of such networks is accurate only up to corrections of order $1/N$, with higher-order terms decaying as $N^{-n/2+1}$.","The size of the correction depends on the input through $c_4(x)$, so some inputs are more Gaussian than others at the same width.","If the $1/N$ perturbation shrinks under gradient descent, the Gaussian process and neural tangent kernel descriptions remain valid throughout training; if it grows, there may be a point where they break down."],"supporting_citations":[{"why":"Establishes the infinite-width single-hidden-layer network as a Gaussian process, the baseline this paper corrects.","marker":"Neal (1996)"},{"why":"Independently proves the same Gaussian process limit for infinite networks.","marker":"Williams (1997)"},{"why":"Supplies the renormalization-group fixed-point and linearization method the derivation follows.","marker":"Sethna (2006)"},{"why":"Justifies identifying the Gaussian as the stable fixed point for finite-variance distributions.","marker":"Feller (1966)"},{"why":"Defines the symmetric uniform initialization used in the experiments.","marker":"Glorot & Bengio (2010)"},{"why":"Provides the rigorous Edgeworth expansion used to identify the expansion coefficients.","marker":"Hansen (2006)"},{"why":"Extends the Gaussian process view to training via the tangent kernel, motivating the paper's discussion of how the perturbation behaves under training.","marker":"Jacot et al. (2018)"},{"why":"Reports empirical agreement between finite networks and Gaussian process or tangent-kernel predictions, cited as evidence that the perturbation stays small.","marker":"Lee et al. (2019)"},{"why":"Derives the multivariate Edgeworth expansion that the paper assumes governs joint output distributions at multiple inputs.","marker":"Sellentin et al. (2017)"}],"fun_headline_variants":["Finite net deviation is fourth Hermite at 1/N","Neural nets Edgeworth: leading correction at 1/N","Symmetric init kills skew, leaves 1/N kurtosis","Finite width: non-Gaussianity scales as 1/N","Edgeworth expansion reveals neural net finite-size"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper derives the output density at a single input but assumes, without proof, that the joint distribution of outputs at multiple inputs is a multivariate Edgeworth expansion, and the Gaussian process picture as a distribution over functions depends on that joint behavior.","fun_headline_variants_meta":{"raw":{"variants":["Finite net deviation is fourth Hermite at 1/N","Neural nets Edgeworth: leading correction at 1/N","Symmetric init kills skew, leaves 1/N kurtosis","Finite width: non-Gaussianity scales as 1/N","Edgeworth expansion reveals neural net finite-size"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000216,"raw_usage":{"total_tokens":1397,"prompt_tokens":875,"completion_tokens":522,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":438}},"tokens_in":491,"tokens_out":522,"duration_ms":5813,"temperature":1.0,"reasoning_tokens":438,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:55:34.254467+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the joint statistics of outputs at two or more inputs for a large ensemble of finite-width networks, for example the connected four-point cumulant $\\langle y(x_1)y(x_2)y(x_3)y(x_4)\\rangle_c$, and compare it with the multivariate Edgeworth prediction whose leading term is set by the same $c_4(x)/N$ amplitudes. If the joint distribution deviates from the multivariate Gaussian in a way not captured by that expansion, or with a different dependence on $N$, the function-space Gaussian process approximation would fail even though the single-input marginal is well fitted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Independently proves the same Gaussian process limit for infinite networks."}],"review_version":1}