{"id":"ecf7a11c-b2fe-4cd0-a38f-803eefecf87a","arxiv_id":"2502.05300","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"This position paper argues that parameter symmetry breaking and restoration unify three hierarchies in deep learning: learning dynamics, model complexity, and representation formation.","lead":"Parameter symmetries, changes to a network's weights that do not alter its output, are proposed as the common cause behind several distinct deep learning phenomena. The paper argues that symmetry breaking and restoration drive when networks learn abruptly, how their capacity adapts, and how abstract representations form.","discovery_kind":"paradigm_shift","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The empirical support for all three hypotheses rests on thresholded pairwise-distance counts that are not shown to be threshold-stable or causal; the central position needs an intervention test.","rationale":"The reader's weakest assumption is the same one I would defend: the operational symmetry metric and the causal interpretation are the least secure links in an otherwise coherent argument. I agree with the CONDITIONAL verdict. The paper is explicitly a position paper; its value is in assembling evidence and proposing a research program, not in delivering one falsifiable theorem. There is original content (Theorem 4) and an honest statement that the formal mechanism for entropy-driven restoration is still lacking (Section 6). I also note the missing reference for Conjecture 1, marked 'Removed for anonymity. In Preparation.'; that is a real gap, but it is less load-bearing because Section B.2 proves a special case and the central empirical claims do not depend on the full conjecture. The threshold/causality issue does not refute the position, but it means the current figures cannot distinguish the symmetry mechanism from a generic correlate of training dynamics. The syre intervention is already used elsewhere in the paper, so the proposed test is available at low cost. If it passes, the central claim is substantially strengthened; if it fails, the empirical support would need to be reconsidered. No change to the reader's verdict is needed.","tokens_in":19716,"tokens_out":6212,"duration_ms":66015,"concrete_test":"Re-run the Figure 2 teacher-student experiment (five-layer tanh FCN, 300 samples, SGD, the four initialization scales) with permutation symmetry removed by the syre intervention of Ref. [88], keeping all other hyperparameters fixed. If the loss plateaus and abrupt jumps persist with unchanged timing, then ΔG-threshold crossings are not causal and the Dynamics Hypothesis loses its empirical support; if they disappear, causality is supported. As a control, sweep ΔG_th over {0.01, 0.05, 0.1, 0.2, 0.5, 1.0} and require that detected event times align with loss changepoints no worse than a control metric such as the parameter update norm, with its threshold chosen to match the event count.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central position is an empirical claim: the observed hierarchies are 'primarily determined by' symmetry breaking and restoration. Every direct test of that claim uses the same operational protocol: a symmetric state is defined by PGθ = θ (Definition 2), ΔG is measured as ||θ − PGθ||² (Eq. 3), and then a hand-chosen threshold (ΔG_th = 0.05–0.2 in Figure 2, 1.0 in Figure 4) converts this continuous residual into binary 'symmetry-breaking events' by counting adjacent-neuron pairs sorted by norm (Section A.1). This protocol has two unsecured links. First, the threshold choice can change which events are detected; the reported coincidence between these events and loss jumps in Figure 2 or rank changes in Figure 4 is not shown to be stable under the threshold, and for the double-rotation symmetry in Figure 3 the measured quantity is not Eq. (3) at all but the eigenvalue-difference proxy in Eq. (13). Second, alignment is not causation: ΔG growth, loss drops, and representation-rank changes are all functions of the same parameter trajectory, so a common third factor such as gradient scale, effective learning rate, or noise level could produce the same figures without symmetry playing a driving role. Since the Dynamics, Complexity, and Representation hypotheses are all supported by this same measurement style, this is the load-bearing weakness of the paper.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that parameter symmetry breaking and restoration is the unifying mechanism behind three empirically observed hierarchies in deep learning: the temporal hierarchy of learning dynamics (e.g., saddle-to-saddle jumps), the complexity hierarchy (e.g., width-independent generalization), and the representational hierarchy (e.g., neural collapse and representation alignment). The paper formalizes parameter symmetries with Definition 1 and Definition 2, defines the symmetry-breaking distance Delta_G in Eq. (3), and presents experimental evidence from MLPs, transformers, ViTs, and ResNets in which symmetry-breaking events are aligned with loss jumps, complexity changes, and representation-rank changes. It also appeals to several theoretical results, including Theorem 1 from the authors' prior work, Theorem 2 from a companion preprint, and Conjecture 1 from an in-preparation manuscript. The stated position is that these three hierarchies are primarily determined by symmetry dynamics, with implications for deliberately engineering symmetries to control learning. The paper itself acknowledges in Section 7 that validating or falsifying its three hypotheses remains the most important next step.","tokens_in":19952,"tokens_out":4704,"duration_ms":48277,"significance":"If the central position is correct, it offers a genuinely unifying research program: distinct phenomena currently explained by separate theories could be understood as different manifestations of a single symmetry-breaking/restoration dynamic, and it would provide a design principle for introducing or removing symmetries in models. The paper is valuable as a synthesis: it connects extensive prior literature on symmetry-induced saddles, stochastic collapse, neural collapse, and representation alignment, and it states three concrete, individually testable hypotheses. Its strengths include explicit operational definitions of symmetry-breaking distance, reproducible experimental protocols in Section A, and honest acknowledgment of alternative views in Section 7. However, the current empirical support is largely correlational and depends on hand-chosen thresholds in the symmetry-breaking measure, and the main theoretical props are informal or unpublished self-references.","major_comments":[{"comment":"The operational definition of a symmetry-breaking event is threshold-dependent, and no threshold-stability analysis is reported. The text sets Delta_G_th between 0.05 and 0.2 in most experiments but uses 1.0 in Figure 4, and for the double-rotation symmetry in Figure 3 the measured quantity is the eigenvalue-difference proxy in Eq. (13) rather than Eq. (3). The alignment between symmetry-breaking events and loss jumps or rank changes could therefore be an artifact of threshold choice. The authors should report the detected event times, or the coincidence statistics, over a sweep of Delta_G_th, or use a scale-invariant statistic such as a relative increase in Delta_G. This is load-bearing because all three hypotheses are evaluated with this same measurement style.","section":"Section 3, Eq. (3), Figure 2"},{"comment":"The experiments demonstrate temporal coincidence between Delta_G growth and loss drops or representation-rank changes, but coincidence along the same optimization trajectory does not establish that symmetry breaking 'primarily determines' these hierarchies. A common third factor, such as effective learning rate, gradient scale, or stochastic noise, could drive both the symmetry measure and the observed leaps. Figure 5 provides an intervention for the representation hypothesis (removing permutation symmetry suppresses neural collapse), but no analogous intervention is provided for the dynamics and complexity hypotheses. The authors should either soften the 'primarily determined' claim to a correlational one or add intervention experiments, e.g., maintaining the parameters on the symmetric subspace during training or removing/applying individual symmetries and measuring whether the loss jumps and complexity adaptation disappear.","section":"Sections 3-5, central Position statement"},{"comment":"The main text claims that at a G-symmetric state the effective model dimension decreases by exactly rank(PG) 'throughout training.' The formal statement in Appendix B.1 shows that the reduced-parameter representation f' exists for all GD/SGD iterations, but the NTK feature-masking statement is proved only in the lazy-training limit (Eqs. (21)-(25) with kappa < 1/(2 lambda_max(A))). The text conflates the exact reduction of parameter count with the stronger kernel-regime claim. This should be flagged explicitly, since Section 4 uses the NTK statement to argue that symmetric states are low-capacity states from which gradient methods cannot escape.","section":"Section 4 and Appendix B.1, Theorem 1/Theorem 3"},{"comment":"The space-quantization result supporting the Occam's-razor argument relies on the unverified regularity condition in Eq. (26): ||nabla_theta ell_0(...)|| <= K||theta||^q with K = K0 m^{-alpha}. This assumption does the real work in the proof, as Eq. (37) makes clear; without independent evidence that K decays with the number of active neurons for realistic permutable losses, the bound in Eq. (27) is conditional. The authors should prove or at least numerically verify this scaling condition for concrete losses such as MSE with ReLU or attention layers, or otherwise state the result as conditional.","section":"Section 4, Conjecture 1 and Appendix B.2, Theorem 4"},{"comment":"Several load-bearing theoretical items are informal self-references or unpublished: Theorem 1 generalizes the authors' ICLR 2025 paper [88], Theorem 2 is taken from the authors' companion preprint [87], and Conjecture 1 is attributed to an in-preparation paper by two of the authors, cited as 'Removed for anonymity.' For a journal submission, these should be replaced with peer-reviewed versions, precise statements included in the paper, or proofs supplied in the appendix. Reference [73] in particular cannot be verified in its current form and should not be treated as a citable source.","section":"Sections 4, 5, Appendix B; references [73], [87], [88]"}],"minor_comments":[{"comment":"Equation (7) and the surrounding text contain severe Unicode artifacts (e.g., '⌟⟨⟨⟪rl⟫l⟩⟩⟪...') that make the displayed formula unreadable; this must be fixed before publication.","section":"Section 6, Eq. (7)"},{"comment":"The notation Ndos is defined as the number of small-Delta_G pairs, but then Ndosb = h - Ndos is called the degree of symmetry breaking; the relationship of these counts to the number of broken generators discussed in Section 3 should be stated explicitly, since h is the number of neurons only for the sorted-neighbor procedure.","section":"Appendix A.1, Eq. (12)-(14)"},{"comment":"Minor typos: 'intialized' should be 'initialized', 'iteraiton' should be 'iteration', and 'kernalized' should be 'kernelized'.","section":"Appendix B.1"},{"comment":"The caption does not state the threshold Delta_G_th used for the black dotted lines; the reader must infer it from Section A.1. Please report the exact threshold in the caption or in the main text.","section":"Figure 2 caption"},{"comment":"The Taylor expansion in Eq. (4) is stated for small Delta_G but the validity condition on Delta_G relative to the Hessian is not discussed; a sentence clarifying the assumed smoothness and the order of the remainder would improve rigor.","section":"Section 6 and Eq. (4)"}],"recommendation":"major_revision","confidential_remarks":"This is a position paper, and I have judged it by that standard: clear central claim, testable hypotheses, and breadth of synthesis count for credit, and the absence of full proofs is less damning than in a theory paper. However, the threshold-dependence and causality concerns are not merely presentational; they are the empirical load-bearing wall of the paper, and the unpublished self-references need to be resolved. I also note that the citation '[73] Removed for anonymity' is inappropriate in a submission and may confuse readers and reviewers; the editor should require the authors to supply the full citation or remove the reference. The paper is not ready for acceptance, but it is not a reject: the three hypotheses are well-posed and could be strengthened with threshold sweeps and intervention experiments within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a synthesis paper with a real thesis: parameter symmetry breaking and restoration could unify three hierarchies (learning dynamics, model complexity, representation formation). I think the position is worth taking seriously, and the paper says it well. What is actually new is the framing itself, plus an original proof (Theorem 4) of a space-quantization bound that gives a concrete mechanism for width-independent complexity. The proof is sound given its assumptions, though the key gradient regularity condition (Eq. 26, with K = K0 m^{-\\alpha}) is unverified and doing a lot of work. Credit where due: the taxonomy of symmetries is useful, and the physics analogy is thought-provoking rather than hand-wavy.\n\nThe soft spots are real and concentrated in two places. First, the three hypotheses lean heavily on the authors' own earlier and in-preparation results: Theorem 1 generalizes [88], Theorem 2 is from [87], and Conjecture 1 is from a paper cited as \"Removed for anonymity\" [73]. That placeholder is a red flag: a reader cannot check the conjecture that underlies the complexity argument. Even if the companion results are legitimate, the paper as it stands does not give the community enough to verify its central claims.\n\nSecond, the experiments do not establish that symmetry breaking is causal, or even that the measured events are threshold-stable. The protocol counts neuron pairs whose L2 distance exceeds a hand-chosen threshold (0.05–0.2, or 1.0 in Figure 4). The paper does not show that the reported alignments with loss drops or rank changes survive under reasonable threshold variation, and for double-rotation symmetry the measure switches to an eigenvalue-difference proxy (Eq. 13). More fundamentally, symmetry distance, loss, and rank are all readouts of the same parameter trajectory; the figures show correlation, not that symmetry is the driving mechanism. An intervention test—deliberately breaking or restoring a symmetry and seeing whether the hierarchy changes—would address the causation gap. Without it, the \"primarily determined by\" claim is not established.\n\nThese are load-bearing weaknesses for the position, but not for the paper as a position paper. The authors are explicit that this is a call for a research direction, and they acknowledge the lack of a formal treatment of the restoration mechanism. I'd send this to peer review, but I'd ask the authors to make the companion results public, test threshold stability, and add at least one intervention experiment. It is a useful paper for researchers thinking about unifying principles in deep learning theory, and it deserves referee time despite the gaps.","headline":"A plausible, clearly written position paper that is worth refereeing, but its empirical support rests on thresholded correlations rather than causal tests, and its key theoretical props are mostly the authors' own informal or in-preparation results.","tokens_in":20540,"tokens_out":2388,"would_cite":false,"duration_ms":25672,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"Parameter symmetry could unify deep learning theory.","keywords":["parameter symmetry","symmetry breaking and restoration","learning dynamics","model complexity","neural representations","neural collapse","deep learning theory","position paper"],"falsifier":"A decisive check is to train the same teacher-student networks with all parameter symmetries removed and ask whether the sharp loss plateaus, complexity jumps, and collapsed representations still appear with unchanged timing; if they do, the claim that symmetry breaking drives the three hierarchies is falsified. A second check is to vary the threshold $\\Delta_G^{\\mathrm{th}}$ across a wide range and show that the reported coincidence between symmetry-breaking events and loss jumps disappears at alternative reasonable thresholds.","tokens_in":19459,"feed_emoji":"⚛️","tokens_out":8023,"duration_ms":70289,"temperature":0.7,"pith_summary":"The paper argues that the fragmented explanations of deep learning can be unified by one mechanism: parameter symmetry breaking and restoration. It claims that three observed hierarchies of neural networks — the abrupt, phase-transition-like jumps in learning dynamics, the adaptive complexity of trained models, and the progressive abstraction of representations — are different manifestations of the same symmetry dynamics. If the position is right, symmetries are not a technical nuisance but a central principle for understanding and intentionally engineering how AI systems learn. The paper is a position piece: it synthesizes existing results, states three concrete hypotheses, and supports each with experiments and theorems drawn from prior work.","feed_headline":"Parameter symmetry could unify deep learning's three hierarchies","feed_subtitle":"Loss jumps, complexity adaptation, and representations may all come from one symmetry mechanism.","key_machinery":"The central object is the parameter symmetry group and the projection matrix $P_G$ onto its invariant subspace. A parameter $\\theta$ is in a $G$-symmetric state when $P_G \\theta = \\theta$; the symmetry-breaking distance $\\Delta_G = \\|\\theta - P_G\\theta\\|_2^2$ measures how far training has moved the parameters from that state, and the degree of symmetry counts how many pairwise neuron distances exceed a chosen threshold. This gives an operational definition of symmetry breaking and restoration, and it connects to a theorem: at a $G$-symmetric state the effective parameter dimension drops by $\\mathrm{rank}(P_G)$, which in the lazy-training regime is equivalent to masking the neural tangent kernel features. The same projection yields a decomposition of the model near a symmetric state into a quadratic kernel model plus a smaller residual model, which is why feature learning is tied to symmetry.","core_discovery":"The central claim is a unifying hypothesis: the hierarchies of learning dynamics, model complexity, and representation formation are primarily determined by parameter symmetry breaking and restoration. Concretely, the paper argues that the loss landscape of a network is organized into symmetry classes, that gradient-based training moves between these classes, and that each transition changes the effective number of parameters and therefore the model's capacity. It further argues that invariant, hierarchical, and universal representations require parameter symmetry: removing permutation symmetry eliminates neural collapse, while the double rotation symmetry of deep linear networks provably forces universally aligned representations across different models. All three hierarchies are presented as corollaries of the same symmetry mechanism rather than as separate phenomena.","pith_inferences":["Editorial inference: If the symmetry-to-symmetry picture transfers from deep linear to nonlinear networks, representation alignment between models may become predictable from the symmetry groups of their architectures, offering a design rule for cross-model compatibility.","Editorial inference: The framework suggests a clean experiment: add a single scalar multiplier to a layer's weights (creating one new symmetry) and vary only the regularization strength; if hierarchy-like loss plateaus appear and disappear with that symmetry, the causal story is directly tested.","Editorial inference: The threshold-based symmetry metric is a practical proxy; replacing it with a scale-invariant or topologically defined order parameter would make the theory robust to the choice of $\\Delta_G^{\\mathrm{th}}$.","Editorial inference: The space-quantization argument could be extended from neuron weights to attention heads and vocabulary embeddings, predicting that representation rank saturates with width in transformers under weight decay, consistent with the ViT observations in the paper."],"forward_implications":["If symmetry dynamics drive learning, then sharp loss drops during training should coincide with symmetry-breaking events, and removing symmetries should eliminate the corresponding plateaus.","The space-quantization conjecture implies that under weight decay a layer contains at most a regularizer-dependent number of non-identical neurons, making effective model complexity and generalization essentially width-independent.","Neural collapse should disappear when permutation symmetry is removed and reappear when it is restored.","The double-rotation symmetry of deep linear networks forces universal alignment of representations across arbitrarily different models, giving a proof of a Platonic-representation-style statement in that setting.","Practitioners should be able to engineer desired hierarchies by deliberately introducing or removing symmetries in models, loss functions, and data."],"supporting_citations":[{"why":"Supplies the definition of the symmetric state and the theorem that a $G$-symmetric solution reduces effective parameter dimension, and gives the symmetry-removal method used in the representation experiments.","marker":"[88]"},{"why":"Shows that weight decay makes symmetric states energetically favorable and offers the mechanism for introducing new symmetries into a layer.","marker":"[83]"},{"why":"Establishes saddle-to-saddle dynamics in deep linear networks, the learning-dynamics picture the paper's first hypothesis extends.","marker":"[33]"},{"why":"Shows that gradient noise attracts SGD to simpler subnetworks, supplying the implicit-regularization mechanism for symmetry restoration.","marker":"[12]"},{"why":"Proves that double rotation symmetry in deep linear networks yields universally aligned representations across different models, the basis of Theorem 2.","marker":"[87]"},{"why":"States the space-quantization conjecture bounding the number of distinct neurons under regularization, which underlies the width-independence argument.","marker":"[73]"},{"why":"Demonstrates loss leaps and complexity jumps during SGD training, the phenomenon the dynamics hypothesis is built to explain.","marker":"[1]"},{"why":"Documents neural collapse as a prevalent terminal-phase phenomenon, the representation effect the paper ties to permutation symmetry.","marker":"[54]"}],"fun_headline_variants":["Symmetry breaking may unify deep learning's three hierarchies","Deep learning's three hierarchies reduce to one symmetry mechanism","Parameter symmetry: the key to deep learning's three hierarchies","Symmetry breaking and restoration unify deep learning's hierarchies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the hand-chosen threshold $\\Delta_G^{\\mathrm{th}}$ (0.05–0.2 in most experiments, 1 in one figure) faithfully tracks whether a true group-theoretic symmetry is broken, and that the measured symmetry transitions actually cause the loss jumps and representation changes rather than merely tracking a third factor such as learning-rate or gradient-noise dynamics.","fun_headline_variants_meta":{"raw":{"variants":["Symmetry breaking may unify deep learning's three hierarchies","Deep learning's three hierarchies reduce to one symmetry mechanism","Parameter symmetry: the key to deep learning's three hierarchies","Symmetry breaking and restoration unify deep learning's hierarchies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001621,"raw_usage":{"total_tokens":6383,"prompt_tokens":811,"completion_tokens":5572,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":427,"completion_tokens_details":{"reasoning_tokens":5506}},"tokens_in":427,"tokens_out":5572,"duration_ms":36083,"temperature":1.0,"reasoning_tokens":5506,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T19:51:55.396936+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check is to train the same teacher-student networks with all parameter symmetries removed and ask whether the sharp loss plateaus, complexity jumps, and collapsed representations still appear with unchanged timing; if they do, the claim that symmetry breaking drives the three hierarchies is falsified. A second check is to vary the threshold $\\Delta_G^{\\mathrm{th}}$ across a wide range and show that the reported coincidence between symmetry-breaking events and loss jumps disappears at alternative reasonable thresholds.","supporting_citations":[{"cited_title":"Remove symmetries to control model expressivity","cited_arxiv_id":null,"evidence_quote":"Supplies the definition of the symmetric state and the theorem that a $G$-symmetric solution reduces effective parameter dimension, and gives the symmetry-removal method used in the representation experiments."},{"cited_title":"Symmetry induces structure and constraint of learning","cited_arxiv_id":null,"evidence_quote":"Shows that weight decay makes symmetric states energetically favorable and offers the mechanism for introducing new symmetries into a layer."},{"cited_title":"Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks","cited_arxiv_id":"2306.04251","evidence_quote":"Shows that gradient noise attracts SGD to simpler subnetworks, supplying the implicit-regularization mechanism for symmetry restoration."},{"cited_title":"Neural thermodynamics i: Entropic forces in deep and universal representation learning","cited_arxiv_id":null,"evidence_quote":"Proves that double rotation symmetry in deep linear networks yields universally aligned representations across different models, the basis of Theorem 2."},{"cited_title":"Removed for anonymity","cited_arxiv_id":null,"evidence_quote":"States the space-quantization conjecture bounding the number of distinct neurons under regularization, which underlies the width-independence argument."},{"cited_title":"Prevalence of neural collapse during the terminal phase of deep learning training","cited_arxiv_id":null,"evidence_quote":"Documents neural collapse as a prevalent terminal-phase phenomenon, the representation effect the paper ties to permutation symmetry."}],"review_version":1}