{"id":"b33d5388-28f4-498e-b876-4141e7903c72","arxiv_id":"2412.06943","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"For rotationally invariant random matrices, entrywise nonlinearities are asymptotically equivalent to a linear combination of the original matrix and an independent Gaussian matrix, with coefficients computed from the function and the matrix variance.","lead":"This paper proves that applying a nonlinear function entry by entry to a symmetric random matrix with rotational symmetry changes the eigenvalue spectrum in a simple asymptotic way: it is the same as adding an independent Gaussian matrix to a scaled copy of the original matrix. This 'Gaussian equivalence' gives a general explanation for a phenomenon previously seen only in special matrix models used in machine learning.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Centeredness (Assumption 2) is load-bearing: the showcased ReLU and max examples are not centered, so their matrices carry a sqrt(N) outlier absent from the Gaussian equivalent; the abstract's unrestricted GEP claim is not proven.","rationale":"The reader's weakest assumption is exactly the point where the central GEP claim is least secure. For centered polynomials the combinatorial argument is coherent: Proposition 1 plus Leonov-Shiryaev reduce the calculation to cyclic cumulants, and the graph/cactus analysis plausibly controls the subleading contributions. I found no internal contradiction inside the centered-polynomial theorem. The problem is the distance between Theorem 4/Corollary 11 and the advertised conclusions. Section 5 applies ReLU without centering and without flagging the predictable sqrt(N) outlier; Corollary 11 as written predicts no such outlier. Since the proof of Theorem 4 starts from c1(y_ij)=0 and builds the moment expansion on that, the non-centered case is not a minor extension: a rank-one shift changes outlier/strong-convergence behavior and breaks the moment method, even though the bulk weak limit may still match after subtracting the mean. The polynomial restriction is similarly load-bearing for the abstract's unqualified 'non-linear function' claim; the ReLU/max sections explicitly call the extension 'feasible' and provide numerics only. That is honest, but it means the showcased examples are not consequences of the theorem. A largest-eigenvalue check settles whether the full-spectrum claim fails; I expect it does. This supports the reader's CONDITIONAL verdict, so no adjustment is needed.","tokens_in":26723,"tokens_out":11895,"duration_ms":146230,"concrete_test":"Compute the largest eigenvalue of Y_N in the Section 5 ReLU example for N = 1000, 2000, 5000 and compare it with the largest eigenvalue of Y_hat_N. If Y_N's leading eigenvalue grows like sqrt(kappa2/(2 pi)) sqrt(N) while Y_hat_N's stays O(1), the correctness of Corollary 11 as a full-spectrum statement for non-centered f is disproved; equivalently, subtract mu N^{-1/2} J from Y_N and check that the bulk spectra then agree. A one-line check with f(x)=x+1 on a GOE gives the same conclusion analytically.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Under Assumption 2, c1(y_ij)=0 is essential to the moment-cumulant proof of Theorem 4. If E[f(g)] = mu != 0, then for i != j we have y_ij = (1/sqrt(N)) f(sqrt(N) x_ij) = y_ij^c + mu/sqrt(N), so Y_N = Y_N^c + mu N^{-1/2} J up to a diagonal correction. The rank-one term has an eigenvalue of order mu sqrt(N), whereas the linear model Y_hat_N = theta1 X_N + theta2 Z_N has no such outlier under the stated hypotheses. Hence Corollary 11 is false as a full-spectrum statement for non-centered f; it can hold only as a bulk/vague limit, a qualification not present in the abstract. Section 5's ReLU example has mu = sqrt(kappa2/(2 pi)) > 0 and should exhibit an eigenvalue of order sqrt(N), yet the outlier is acknowledged only in the max example (Section 7). The proof never controls this perturbation: centeredness is what kills c1(y_ij) before the graph/cumulant expansion, and the moment method used cannot see an outlier whose contribution to moments grows. The polynomial restriction is a second facet: ReLU and max are not polynomials, so the numerical agreement in Figures 1, 3, and 4 is evidence, not proof, for the abstract's general formulation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies the entrywise application of nonlinear functions to symmetric orthogonally invariant random matrices, proposing a Gaussian equivalence principle: the asymptotic eigenvalue distribution of Y_N (defined by y_ij = N^{-1/2} f(N^{1/2} x_ij) off the diagonal, zero on the diagonal) should equal that of a linear combination theta_1 X_N + theta_2 Z_N, where Z_N is an independent GOE. The main theorems (Theorems 4, 12, and 14) derive the free cumulants of Y_N in the case where f is a polynomial and Assumption 2 (centeredness of f with respect to the off-diagonal entry distribution) holds, and the corollaries give explicit linear models. Numerical examples for the ReLU function (Section 5) and the max function (Sections 7 and 8) are presented as evidence for the general principle. The abstract claims the result for general non-linear functions, including multivariate and correlated cases, but the proofs are restricted to centered polynomials.","tokens_in":27003,"tokens_out":3756,"duration_ms":40569,"significance":"If fully established, the result would provide a broad Gaussian equivalence principle for orthogonally invariant ensembles, unifying and extending existing results for Wigner-type and kernel random matrices, and would give explicit coefficients computable from the activation function and the second free cumulant. The use of classical cumulants and the Leonov--Shiryaev formula to pass from entrywise nonlinearities to free cumulants is conceptually appealing and yields clean formulas that agree with the bulk numerical spectra shown in the paper. However, as it stands, the advertised general claim is not proven: the theorems require f to be a centered polynomial, while the showcased ReLU and max functions are neither centered nor polynomial, and the proof of the crucial subcycle/disconnected-cycle cancellations is informal. The paper does not supply machine-checked proofs or reproducible code, so the numerical agreement should be regarded as illustrative evidence only. With the claims properly calibrated, the paper would make a useful contribution.","major_comments":[{"comment":"The abstract states that a Gaussian equivalence principle holds for entrywise application of non-linear functions in general, but Theorem 4 is proved only for polynomials satisfying Assumption 2. The examples in Section 5 (ReLU) and Section 7 (max) are not centered; their Gaussian means are sqrt(5/(2 pi)) and 1/sqrt(pi), respectively. Consequently, the numerical agreement in Figures 1 and 3 cannot be taken as evidence for the stated general theorem; at most it supports a bulk-limit statement. The authors should either restrict the abstract and theorem statements to the centered polynomial case or provide the missing argument for general f, and the outlier phenomenon should be acknowledged in Section 5 just as it is in Section 7.","section":"Abstract and Section 3 (Assumption 2)"},{"comment":"The treatment of subcycles and disconnected cycles is not rigorous. The proof relies on phrases such as \"Careful analysis also reveals\" and \"It is clear that this construction can lead to partitions pi which are crossing\" without giving the counting of contributing multi-indices, and the estimate with the parameters d and d_j is introduced informally. The later conclusion that crossing partitions contribute O(N^{-1}) is asserted rather than proven. Since the free moment-cumulant formula for Y_N depends on exactly these cancellations, this is a load-bearing gap in the proof of Theorem 4.","section":"Section 3, continuation of the proof after Definition 8"},{"comment":"Corollary 11 is stated as an equality of asymptotic eigenvalue distributions, but under Assumption 2 it is at most a statement about the bulk spectrum. For a non-centered f, the matrix Y_N contains a rank-one term of size mu N^{-1/2} up to diagonal corrections, producing an eigenvalue of order sqrt(N) that is not present in theta_1 X_N + theta_2 Z_N; the proof of Theorem 4 uses Assumption 2 to eliminate c_1(y_ij) before the Leonov--Shiryaev expansion. The manuscript should explicitly state that Corollary 11 and the related multivariate corollaries hold only for centered f, and it should describe the outlier separately, as it does in the max example.","section":"Corollary 11 and Assumption 2"},{"comment":"Theorem 12 is stated without the polynomial restriction, but the proof of Corollary 13 begins \"Let us assume that our function f is a (multivariate) polynomial.\" This mismatch between statement and proof leaves the theorem unproved for arbitrary f. Similarly, the proof of Theorem 14 treats only the full-cycle case and refers to the univariate argument for multiple cycles without addressing the correlated-matrix complications in detail. The statements should be aligned with what is actually proven.","section":"Theorem 12 and proof of Corollary 13"}],"minor_comments":[{"comment":"The name \"Cauchy-Schwartz\" should be \"Cauchy-Schwarz.\"","section":"Section 4"},{"comment":"\"It is feasible that the statements ... are also true\" should be \"It is plausible that ... are also true.\"","section":"Section 5"},{"comment":"There is a typo: \"cyclity\" should be \"cyclicity.\"","section":"Proposition 3 proof"},{"comment":"The word \"similiar\" should be \"similar.\"","section":"Section 2"},{"comment":"The sentence about the eigenvalue of order 40 would be more informative if it gave the predicted scaling, e.g., sqrt(N) times the Gaussian mean of max(g1,g2).","section":"Section 7"}],"recommendation":"major_revision","confidential_remarks":"The gap between the abstract's general formulation and the actual theorems is substantial and should be addressed before publication. The numerical sections are illustrative but do not substitute for the missing proof for non-polynomial or non-centered activations. I would recommend that the authors either restrict the paper's title and abstract to the centered polynomial case or provide a rigorous treatment of the general case. If the claims are calibrated accordingly, the paper fits the journal's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth taking seriously. Its real contribution is structural: for a symmetric orthogonally invariant ensemble with a limit distribution of all orders, entrywise application of a centered polynomial f at the usual 1/sqrt(N) scaling gives the same asymptotic eigenvalue distribution as theta1*X_N + theta2*GOE, with theta1 = E[f'(g)] and theta2 determined by the variance of f(g) and the second free cumulant of X_N. The multivariate versions, including the correlated case with mixed free cumulants, are genuinely new and unify several special-case results from the literature. The proof route through classical cumulants and Leonov-Shiryaev is appropriate, and the appendix contains real work, including a careful treatment of the orthogonal analogue of the entry-cumulant/free-cumulant relation. The numerics use the predicted closed-form coefficients, not fits; that is to the authors' credit.\n\nThe soft spots are real, but they are mainly gaps between what is proven and what the abstract advertises.\n\nFirst, the formal theorem is for centered polynomials. Assumption 2 is load-bearing: it removes the first cumulant of the entries, and that is essential to the argument. The abstract's unrestricted claim for general non-linear functions is not proven. ReLU and max are neither polynomials nor centered. The max example does at least acknowledge the sqrt(N) outlier from the nonzero mean; the ReLU example does not, even though the same mechanism applies and should produce a comparable outlier.\n\nSecond, the proof's treatment of subcycles and disconnected cycles is informal. Phrases like 'Careful analysis also reveals' and 'It is clear' cover exactly the part that separates a proof from a heuristic. Appendix C helps for disjoint cycles without subcycles, but the full suppression argument is not written out.\n\nThird, the moment computation gives convergence of expected traces, not concentration in probability. If the intended statement is the usual GEP convergence in probability or almost surely, that step is missing.\n\nNone of this makes me think the main claim is false under the stated hypotheses; the flaws are in scope control and in proof completeness. The paper deserves a serious referee. My recommendation: send it to peer review, but require the authors to state the polynomial and centeredness hypotheses in the abstract, acknowledge the ReLU outlier, and either complete or explicitly defer the subleading-contribution and concentration arguments.","headline":"A worthwhile structural generalization of Gaussian equivalence to orthogonally invariant matrices, with a solid centered-polynomial core but an abstract that oversells non-polynomial and non-centered cases.","tokens_in":27521,"tokens_out":2919,"would_cite":true,"duration_ms":35544,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["60B20","46L54"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that applying a centered nonlinear function entrywise to an orthogonally invariant random matrix leaves the same asymptotic spectrum as a linear combination of the matrix with an independent GOE.","keywords":["random matrices","orthogonal invariance","Gaussian equivalence principle","free cumulants","entrywise nonlinear functions","ReLU","GOE","eigenvalue distribution"],"falsifier":"Take $X_N$ to be a single GOE and apply the centered cubic $f(x)=x^3-3x$ entrywise. The theorem predicts $\\theta_1=0$ and that $Y_N$ converges to a semicircle with second moment $\\mathbb{E}[f(g)^2]=6$; simulating large $N$ and comparing the empirical spectrum to that semicircle would falsify the equivalence if a systematic deviation beyond finite-size fluctuations appeared. A second check targets the centeredness assumption directly: applying the uncentered ReLU to any of the paper's ensembles should produce an eigenvalue of order $\\sqrt{N}$ outside the bulk, while the bulk alone should still match the linear model.","tokens_in":26477,"feed_emoji":"🎲","tokens_out":10521,"duration_ms":103087,"temperature":0.7,"pith_summary":"This paper claims that for a symmetric random matrix ensemble whose joint entry distribution is unchanged by conjugation with any orthogonal matrix, the entrywise application of a nonlinear function has a surprisingly simple spectral effect: in the large-size limit it is equivalent to a linear combination of the original matrix and an independent Gaussian orthogonal ensemble (GOE). The precise statement is that, for a centered nonlinearity $f$, the matrix $Y_N$ with off-diagonal entries $y_{ij} = N^{-1/2} f(\\sqrt{N} x_{ij})$ and zero diagonal has the same asymptotic eigenvalue distribution as $\\theta_1 X_N + \\theta_2 Z_N$, where $Z_N$ is a standard GOE and the two coefficients are computed from $f$ and from one Gaussian test variable of variance $\\kappa_2$, the second free cumulant of the limiting spectrum of $X_N$. The same equivalence is proved for multivariate functions applied entrywise to several matrices, including matrices that are correlated with each other. The paper would matter because it reduces a nonlinear, high-dimensional transformation to a linear model whose spectrum is known, and it explains why such reductions appear across machine-learning settings.","feed_headline":"A Gaussian equivalence holds for orthogonally invariant matrices","feed_subtitle":"ReLU and max entrywise maps change a matrix's spectrum only through a linear term plus Gaussian noise.","key_machinery":"The engine of the proof is the combinatorial relation between classical cumulants of the matrix entries and free cumulants of the limiting eigenvalue distribution, processed through the Leonov-Shiryaev formula for cumulants of products. Orthogonal invariance forces almost all entry cumulants of $X_N$ to vanish; the only ones surviving asymptotically have a cyclic index structure, and their scaled limits are the free cumulants $\\kappa_n$. When the nonlinearity is inserted, Leonov-Shiryaev reorganizes the expanded moments, and the leading cyclic cumulants of $Y_N$ come from one full cycle of the original entries together with Gaussian pairings of the remaining factors. That pairing count evaluates to $\\mathbb{E}[f'(g)]^n$ per full cycle, which is the mechanism that produces the linear replacement model.","core_discovery":"The central claim, stated on the paper's own terms, is a cumulant-level equality. For an orthogonally invariant symmetric ensemble $X_N$ with a limit distribution of all orders and a centered polynomial $f$, the asymptotic free cumulants $\\kappa_n^f$ of the entrywise-transformed matrix $Y_N$ are $\\kappa_1^f=0$, $\\kappa_2^f = \\mathbb{E}[f(g)^2]-\\mathbb{E}[f(g)]^2$, and $\\kappa_n^f = \\kappa_n \\,\\mathbb{E}[f'(g)]^n$ for $n\\ge 3$, where $g$ is Gaussian with mean zero and variance $\\kappa_2$. These are exactly the free cumulants of $\\theta_1 X_N + \\theta_2 Z_N$ with $\\theta_1=\\mathbb{E}[f'(g)]$ and $\\theta_2 = \\sqrt{\\mathbb{E}[f(g)^2]-\\mathbb{E}[f(g)]^2-\\kappa_2 \\mathbb{E}[f'(g)]^2}$, so the nonlinear operation never creates spectral structure beyond a rescaling of $X_N$ plus an independent GOE. In the multivariate case the formula replaces $f'$ by partial derivatives and $\\kappa_n$ by mixed free cumulants, with the second-order mixed cumulants entering the coefficient of the GOE term.","pith_inferences":["A natural extrapolation the paper leaves open is that not only the eigenvalue distribution but also finer spectral quantities, such as eigenvector overlaps or linear-statistics fluctuations, of $Y_N$ are inherited from the linear model $\\theta_1 X_N + \\theta_2 Z_N$; this would be a testable conjecture, not a proven consequence.","Because the theorem is proved only for polynomials while the numerical examples use ReLU and max, a quantitative finite-$N$ check of the rate of convergence for non-polynomial functions would indicate whether the polynomial restriction is a proof artifact or a genuine boundary.","In applications where entrywise activations are applied to random matrices with nonzero mean, subtracting the mean before the nonlinearity should remove the outlier and restore the clean Gaussian-equivalence description; this is an actionable design rule suggested by the paper's outlier remark.","The result suggests a general heuristic: rotational invariance, not the particular entry distribution, is what makes nonlinear preprocessing 'linearize' in the spectrum; ensembles that are only approximately invariant may show the same equivalence up to corrections controlled by the violation of invariance."],"forward_implications":["For every centered polynomial nonlinearity, the bulk spectrum of the transformed matrix is computable as the free convolution of a rescaled copy of the original spectrum with a semicircle.","If $\\theta_1 = \\mathbb{E}[f'(g)] = 0$, which happens for even functions, the original matrix's spectral shape disappears entirely and $Y_N$ becomes asymptotically semicircular.","For multivariate functions, the asymptotic effect of functions such as $\\max(x_1,x_2)$ on two ensembles is the same as $\\theta_1 X_N^{(1)} + \\theta_2 X_N^{(2)} + \\theta Z_N$, with coefficients given by the Gaussian expectations of the corresponding partial derivatives.","The equivalence also holds when the input matrices are correlated, as long as the whole tuple is jointly orthogonally invariant; the mixed free cumulants then determine the coefficients.","Uncentered functions are a genuine boundary: the bulk still follows the linear model, but the nonzero mean produces an additional outlier eigenvalue, so the centeredness assumption is not a formality."],"supporting_citations":[{"why":"Provides the cumulant-to-free-cumulant relation for orthogonally invariant ensembles used in Proposition 1 to identify cyclic cumulant limits with free cumulants.","marker":"[CC07]"},{"why":"Supplies the higher-order free cumulant framework whose arguments are adapted in the appendix to prove the cyclic scaling for the orthogonal case.","marker":"[CMSS07]"},{"why":"Gives the graph and cactus decomposition and the vanishing of non-Eulerian graph contributions used to identify which index structures survive in the moment calculation.","marker":"[MFC+19]"},{"why":"The Leonov-Shiryaev formula is the combinatorial identity that transfers cumulants of products of transformed entries back to cumulants of the original entries.","marker":"[LS59]"},{"why":"Provides the free moment-cumulant formula that converts the computed cumulants $\\kappa_n^f$ into the limiting eigenvalue distribution.","marker":"[MS17]"}],"fun_headline_variants":["Entrywise nonlinearities on invariant matrices: linear plus GOE","Gaussian equivalence for entrywise nonlinear matrix transforms","ReLU and max maps on random matrices: only linear plus GOE","Entrywise nonlinearity: spectrum becomes linear combination plus GOE","Nonlinear entrywise maps on invariant matrices: Gaussian equivalence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the nonlinear function is centered, so the transformed off-diagonal entries have mean zero; the paper's own max-function example shows that without this condition an extra large eigenvalue appears and the full spectral claim fails.","fun_headline_variants_meta":{"raw":{"variants":["Entrywise nonlinearities on invariant matrices: linear plus GOE","Gaussian equivalence for entrywise nonlinear matrix transforms","ReLU and max maps on random matrices: only linear plus GOE","Entrywise nonlinearity: spectrum becomes linear combination plus GOE","Nonlinear entrywise maps on invariant matrices: Gaussian equivalence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001197,"raw_usage":{"total_tokens":4925,"prompt_tokens":925,"completion_tokens":4000,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":3916}},"tokens_in":541,"tokens_out":4000,"duration_ms":27019,"temperature":1.0,"reasoning_tokens":3916,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:19:58.215943+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take $X_N$ to be a single GOE and apply the centered cubic $f(x)=x^3-3x$ entrywise. The theorem predicts $\\theta_1=0$ and that $Y_N$ converges to a semicircle with second moment $\\mathbb{E}[f(g)^2]=6$; simulating large $N$ and comparing the empirical spectrum to that semicircle would falsify the equivalence if a systematic deviation beyond finite-size fluctuations appeared. A second check targets the centeredness assumption directly: applying the uncentered ReLU to any of the paper's ensembles should produce an eigenvalue of order $\\sqrt{N}$ outside the bulk, while the bulk alone should still match the linear model.","supporting_citations":[],"review_version":1}