{"id":"19694962-27cb-4439-be12-66fc8fd4cd88","arxiv_id":"2607.07845","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"The Hessian bulk consists of weakly broken continuous symmetries of the network parametrization, exact zeros in linear nets that ReLU lifts as pseudo-Goldstone modes whose eigenvectors stay in the symmetry subspace.","lead":"Near-zero Hessian eigenvalues in neural nets are weakly lifted pseudo-Goldstone modes of continuous architectural symmetries. This gives a concrete origin for the bulk of the loss-landscape spectrum that governs optimization and flatness.","discovery_kind":"first_principles","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The central claim is eigenvector-level: high-curvature directions are orthogonal to the architectural symmetry subspace while the bulk lies inside it. The paper supplies (i) an explicit orthogonal basis of generators for linear networks (Appendix B and SM I), (ii) an exact ε^{2} scaling plus confinement criterion for the two-layer Gaussian student-teacher (End Matter C), (iii) direct overlap measurements that realize the predicted two-tier structure on both idealized and trained models, and (iv) a cumulative-coverage argument that closes the top-k gap. The reader's weakest assumption correctly flags the absence of a general mixing theorem, yet that absence is already controlled by the reported residual (5% on whitened CIFAR, ~12% on original-variance data) and by the analytic confinement criterion. Because the residual is small and the two-tier pattern is reproducible across architectures (including the convolutional case), the assumption does not undermine the claim. Code and data are public, so the concrete test above can be run immediately. Consequently the ACCEPT verdict stands without adjustment.","tokens_in":24257,"tokens_out":607,"duration_ms":7686,"concrete_test":"Recompute the CIFAR-10 cumulative coverage O_k (Eq. 9) after replacing the linear-comparison Fisher nullspace by the exact nullspace of the Gauss-Newton matrix of the trained ReLU network itself (evaluated at the same endpoint, residual term dropped). If O_5000 remains ≥0.9 the diagnostic is robust; a drop below ~0.7 would indicate that training-induced mixing is larger than the paper's residual bound.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption (that the linear-comparison Fisher nullspace remains a faithful bulk diagnostic after training and for non-Gaussian data) is real but already quantified and bounded by the paper itself. In the student-teacher setting the adiabatic connection is analytic: H(ε)=(1-ε/2)^{2}H_lin+(ε/2)^{2}H_|·| with vanishing cross term (End Matter C), so bulk eigenvalues track ε^{2} with no ε^{4} correction precisely when eigenvectors stay inside the symmetry subspace; the measured overlaps confirm this. On CIFAR-10 the cumulative coverage O_5000≈0.95 (Eq. 9) shows that the unmeasured tail can contribute at most 5% residual leakage into the high-curvature subspace. The three-layer SM case and the non-whitened CIFAR run exhibit partial mixing, yet the two-tier structure and the dramatic suppression of the leading modes (thousands of standard deviations below the random baseline) survive. No internal inconsistency or hidden assumption that would overturn the eigenvector-level claim is present.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that the bulk of near-zero Hessian eigenvalues of neural-network training losses consists of weakly lifted pseudo-Goldstone modes of continuous architectural symmetries of the parametrization. In multilayer linear networks these symmetries are exact; the authors construct an explicit orthogonal basis of generators (via SVDs of the weight matrices) that spans the null space of the Hessian/Fisher and matches the known finite-eigenvalue count. A Leaky-ReLU deformation is treated as an explicit symmetry-breaking perturbation: the two-layer Gaussian student–teacher Fisher splits exactly into linear and absolute-value blocks, bulk eigenvalues rise as ε², and Kato perturbation theory supplies a confinement criterion (vanishing ε⁴ correction iff eigenvectors remain in the symmetry subspace). Eigenvector overlaps confirm that high-curvature modes are orthogonal to the symmetry subspace while the bulk lies inside it. The same diagnostic is applied to a three-layer student–teacher model, a trained three-layer ReLU MLP on CIFAR-10 (with whitening and cumulative coverage Ok), and a minimal convolutional network with nonlinear (circulant) generators.","tokens_in":24509,"tokens_out":1283,"duration_ms":23438,"significance":"If the eigenvector-level claim holds, it supplies a single architectural origin for the Hessian bulk that has been missing from the literature on outliers, Fisher spectra, and flat minima. The work goes beyond counting zero modes: it constructs the generators explicitly, derives the ε² lifting and confinement criterion analytically for the two-layer Gaussian case, and measures overlaps and cumulative coverage on both idealized and trained models. Code and data are released. The mechanism is expected to organize Fisher/Gauss–Newton spectra as well, with direct consequences for natural-gradient and second-order methods, and the convolutional example shows the diagnostic is not limited to fully connected layers. These are concrete, falsifiable contributions at the level of eigenvectors rather than eigenvalue counts alone.","major_comments":[{"comment":"End Matter C and SM §II establish the clean ε² law and confinement criterion only for the two-layer Gaussian student–teacher Fisher (exact block split, vanishing cross term by parity). In the three-layer SM case (Fig. S1) the ε² trend already bends before ε=1 and leading-mode overlaps level off near oi≈0.2 rather than near zero. The main-text claim that the bulk “lies almost entirely within” the symmetry subspace therefore needs an explicit scope statement: for deeper nets the adiabatic connection is only approximate and residual leakage is O(1). A short quantitative bound or additional depth-controlled experiment would make the generality claim load-bearing rather than extrapolative.","section":null},{"comment":"CIFAR-10 section and SM §IV: the comparison subspace P mixes architectural GL-type symmetries with data-covariance flat directions obtained by zeroing low-variance input components. Dimension counting separates the two contributions (383 232 vs 26 748), and Ok≈0.95 bounds residual leakage of the unmeasured tail, but the non-whitened run (Fig. S2 right) shows that covariance-induced spread entangles the two mechanisms and smooths the overlap transition. The paper should state more sharply which fraction of the observed bulk is architectural versus data-driven, and whether the dramatic suppression of the leading Neff Ceff modes survives when the linear comparison Fisher is built without artificially zeroing variances.","section":null},{"comment":"Eq. (1) and the Mexican-hat appendix correctly note that symmetry generators are exact Hessian zero modes only at critical points; off criticality the residual term can produce finite curvature. The CIFAR experiment evaluates the training-loss Hessian at a non-critical endpoint and compares to the Fisher null space of the linear model. While the observed two-tier structure is still striking, a brief check that the residual contribution along the measured bulk directions remains small (or a comparison to the Gauss–Newton matrix itself) would close the gap between the critical-point theory and the trained-network diagnostic.","section":null}],"minor_comments":[{"comment":"Fig. 2 caption and main text: the N rescaling modes that remain exactly zero for all ε are stated to lie “below the plotted range”; a short inset or explicit note that they are omitted would avoid the impression that the bulk starts above zero.","section":null},{"comment":"Notation for the projector P and the overlap oi (Eq. 8) is introduced cleanly, but the random baseline orandom = rank(P)/d is quoted with different numerical values in different figures; a single consistent formula and the associated standard deviation for the CIFAR case would help the reader.","section":null},{"comment":"The convolutional construction (SM §V) uses nonlinear generators involving C(w(1))⁻¹; a one-sentence remark in the main text that these are still continuous symmetries of the function (hence still produce Fisher zeros when exact) would clarify why the more general ϕ is needed.","section":null},{"comment":"References [46,47] already count finite eigenvalues of deep linear networks; the novelty claim is correctly placed on the eigenvectors and the nonlinear extension, but a slightly sharper sentence distinguishing the present work from those counts would help.","section":null},{"comment":"Typographical: “parametrization” is used consistently in the abstract/title; a few places in the SM switch to “parameterization.” Standardize.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The central eigenvector-level claim is sound and the analytic two-layer control is unusually clean for this literature. The three major points above are clarifications of scope and residual terms rather than contradictions; they can be addressed with modest text and, if desired, one extra panel. I would not block acceptance on them. Fit for a Physical Review Letters / PRX-style venue is good given the Goldstone-mode framing and the explicit constructions; for a pure ML venue the CIFAR MLP (53 % accuracy) may look thin, but the diagnostic itself is architecture-level and does not require SOTA accuracy."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This paper finally gives a concrete origin for the near-zero bulk of neural Hessians: they are the weakly lifted continuous symmetries of the parametrization (exact in linear nets, approximately preserved under ReLU). That is the one thing worth knowing.\n\nWhat is new is not the existence of the symmetries—Kunin, Singh, Bernacchia and others already counted zeros and linked generators to flat directions—but the explicit orthogonal GL basis, the controlled Leaky-ReLU expansion that produces the ε² lift with a confinement criterion from Kato, and the direct eigenvector overlaps that place the bulk inside the linear symmetry subspace. The two-layer Gaussian student–teacher case is especially clean: the Fisher splits exactly, the cross term vanishes by parity, and the measured oi and ε² tracks match the prediction. The CIFAR MLP and the convolutional example show the same two-tier structure survives training and non-FC layers. Code and data are public.\n\nSoft spots are real but bounded. The linear-comparison nullspace is a diagnostic, not a theorem that training or non-Gaussian data cannot mix subspaces; the three-layer SM and non-whitened CIFAR runs already show partial leakage. Cumulative coverage O5000≈0.95 still caps the unmeasured tail residual at 5%, and the leading-mode suppression is thousands of standard deviations below random. The CIFAR model is a plain MLP at ~53%, and only the top 5k eigenvectors are computed—limitations the authors flag. Free parameters (whitening cutoff, ε) are transparent. Citation pattern is appropriate; no circularity.\n\nThis is for people who care about Hessian geometry, natural gradient, second-order methods, or flatness arguments. The central claim holds at the eigenvector level. I would send it to referees and I would cite the generator construction and the overlap diagnostic.","headline":"Clean eigenvector-level account of the Hessian bulk as weakly broken architectural symmetries; the math and diagnostics hold up.","tokens_in":25084,"tokens_out":458,"would_cite":true,"duration_ms":6236,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"The bulk of near-zero Hessian eigenvalues in neural nets are weakly broken continuous symmetries of the architecture.","keywords":["Hessian spectrum","pseudo-Goldstone modes","network symmetries","loss landscape geometry","Fisher information matrix","ReLU networks","overparametrization"],"falsifier":"Train a multilayer ReLU network to a genuine critical point, compute the leading Hessian eigenvectors, and check whether their overlap with the linear-symmetry subspace remains near zero for the high-curvature modes and near one for the bulk; a clear failure of that two-tier structure would falsify the claim.","tokens_in":25177,"feed_emoji":"📐","tokens_out":856,"duration_ms":10162,"temperature":0.7,"pith_summary":"Neural-network loss landscapes are famously anisotropic: a few large-curvature directions sit above a dense bulk of near-zero Hessian eigenvalues. This paper argues that the bulk is not mysterious but is the spectral signature of continuous reparametrization symmetries that leave the network function unchanged. In deep linear networks those symmetries are exact, generate flat directions, and produce exact zero modes whose eigenvectors can be written down explicitly from the singular vectors of the weight matrices. Switching on a ReLU-type nonlinearity breaks the symmetries only weakly, lifting the zero modes into parametrically light pseudo-Goldstone modes that still live almost entirely inside the original symmetry subspace. High-curvature directions, by contrast, remain orthogonal to that subspace. The same two-tier structure appears in a trained three-layer network on CIFAR-10 and in a convolutional example, linking the Hessian bulk to approximate architectural symmetries.","feed_headline":"Near-zero Hessian modes are weakly broken network symmetries","feed_subtitle":"High-curvature directions sit outside the symmetry subspace; the bulk lives almost entirely inside it","key_machinery":"Explicit orthogonal generators of the GL-type interlayer symmetries (built from singular vectors of consecutive weight matrices) together with the eigenvector-overlap diagnostic that measures how much each Hessian mode of a nonlinear network lies inside that linear symmetry subspace.","core_discovery":"The bulk of near-zero Hessian eigenvalues consists of the weakly lifted pseudo-Goldstone modes of the continuous symmetries of the network parametrization. In linear networks the symmetries are exact and their generators form an explicit orthogonal basis of the null space; a ReLU nonlinearity breaks them weakly, so that high-curvature eigenvectors stay orthogonal to the symmetry subspace while the bulk eigenvectors remain almost entirely inside it.","pith_inferences":["If the bulk is mostly architectural, pruning or regularization that deliberately targets the symmetry subspace may remove far more parameters than curvature-based pruning alone suggests.","The same diagnostic should apply to the fully connected blocks inside transformers; verifying the two-tier overlap there would test whether the mechanism survives attention and residual pathways.","Because the residual term in the Hessian can reintroduce curvature off critical points, the pseudo-Goldstone picture may degrade late in training when gradients no longer vanish."],"forward_implications":["The bulk of near-zero modes is largely architectural and therefore persists across data sets and training algorithms that preserve the same continuous symmetries.","Because the same directions remain zero modes of the Fisher matrix at arbitrary parameters, natural-gradient and second-order methods automatically ignore or treat them specially.","Any architecture containing fully connected or convolutional blocks inherits an analogous bulk whose size is fixed by layer widths and over-parametrization count.","The adiabatic connection from linear to ReLU spectra supplies a controlled starting point for analytic approximations of the bulk eigenvalues."],"fun_headline_variants":["Near-zero Hessian bulk from weakly broken network symmetries","Pseudo-Goldstone modes explain Hessian near-zero eigenvalues","Weakly lifted symmetries form the Hessian vanishing bulk","ReLU breaks exact symmetries into near-zero Hessian modes","High-curvature axes lie outside network symmetry subspace"],"cache_read_input_tokens":13824,"weakest_assumption_plain":"That the null space of a linear comparison model (nonlinearities removed, low-variance input directions zeroed) remains a faithful diagnostic of the bulk even after training on real data, without a general guarantee that training or non-Gaussian inputs do not mix the subspaces beyond the residual already measured.","fun_headline_variants_meta":{"raw":{"variants":["Near-zero Hessian bulk from weakly broken network symmetries","Pseudo-Goldstone modes explain Hessian near-zero eigenvalues","Weakly lifted symmetries form the Hessian vanishing bulk","ReLU breaks exact symmetries into near-zero Hessian modes","High-curvature axes lie outside network symmetry subspace"]},"model":"grok-4.5","effort":"low","cost_usd":0.005272,"raw_usage":{"total_tokens":1398,"prompt_tokens":731,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":52720000,"prompt_tokens_details":{"text_tokens":731,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":603,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":731,"tokens_out":64,"duration_ms":5865,"temperature":1.0,"reasoning_tokens":603,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T16:49:59.962808+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Train a multilayer ReLU network to a genuine critical point, compute the leading Hessian eigenvectors, and check whether their overlap with the linear-symmetry subspace remains near zero for the high-curvature modes and near one for the bulk; a clear failure of that two-tier structure would falsify the claim.","supporting_citations":[],"review_version":1}