{"id":"96a27fff-00e2-45cc-acd0-2db6a49fb260","arxiv_id":"2501.05550","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Training deep neural networks from low-variance random weights drives a morphological instability that forms periodic channel structures in the weight distribution, independent of the data.","lead":"This paper derives a theory predicting that deep neural network weights spontaneously organize into channel-like structures during training, and shows these structures form independently of the data. It matters because it connects microscopic training dynamics to macroscopic weight morphology, with potential implications for how we understand and design deep learning systems.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The connectivity dynamics used to predict channel formation are derived from a weight update that silently drops the label-dependent error term ΔY, so the 'from first principles' instability is not tied to actual gradient training.","rationale":"The paper's central claim is that a 'from first principles' derivation shows gradient-based training destabilizes the homogeneous weight state, producing data-independent channels and periodic amplitude modulations. The empirical observations (Figs. 2e, 3e, 4) are consistent with the phenomenon and provide some independent support. However, the derivation's pivot is the connectivity equation Eq. (3)/(S94). Tracing the supplement, the single-sample gradient update in Eq. (S40) contains the training error ΔY^(m); the subsequent derivation of d r_j/dt in Eqs. (S79)-(S94) omits this factor entirely. Unless the U symbol was silently redefined to absorb ΔY (it was not, per Eq. S36), the theory describes a data-independent update rule, not gradient descent. This is more fundamental than the constant-U approximation flagged by the reader: even with constant U, the effective coupling is multiplied by a training-error factor that varies across samples and can change sign. Thus the instability mechanism, and the claim that it arises from gradient-based training independently of data, is not established. The test is to re-derive the connectivity dynamics with ΔY retained; if the c_j become data-dependent, the first-principles claim fails. The empirical channel formation may still be a real phenomenon, and the paper should be accepted conditionally on a corrected derivation or on additional experiments that isolate the mechanism.","tokens_in":30223,"tokens_out":10330,"duration_ms":108555,"concrete_test":"Independently re-derive Eqs. (S79)-(S94) starting from Eq. (S40), carrying the factor η ΔY^(m) through the sums over samples and the definition of c_j, or equivalently derive the connectivity dynamics directly from ∂L/∂w in Eq. (S37). If the resulting c_j depend on ΔY in a way that is not constant and sign-definite across nodes and samples, then Eq. (3) does not follow from gradient descent and the instability derivation must be revised before the central claim can be accepted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing problem is that the coarse-grained connectivity dynamics, Eq. (3) / (S94), are not actually derived from the gradient update used to train the networks. In the supplement, the single-sample weight increment is written in Eq. (S40) as Δw = η ΔY^(m) Σ_{n,n'} U_n^{n'} w w, where ΔY^(m) = Y^(m) − Ŷ^(m) is the training error. In Section 3.2, this expression is used to derive d r_j/dt, but the factor η ΔY^(m) is silently absent from Eq. (S79) onward. The U defined in Eq. (S36) contains only inputs and path weights, not the label-dependent error. Consequently, the effective couplings c_j in Eq. (S94) are never shown to have the sign or magnitude properties claimed; for a generic mini-batch or a target distribution with both positive and negative residuals, ΔY^(m) varies across samples and in sign, so the effective coupling ΔY U is neither constant nor sign-definite. The homogeneous-state instability is therefore not a derived property of gradient-based training; it is a property of a different, data-independent update. This is upstream of the reader's constant-U concern (S105): even granting constant U, ΔY U is not constant. The empirical channel observations remain, but the central 'from first principles' claim is unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript claims that training deep feedforward networks from a low-variance, near-homogeneous initialization induces a morphological instability in the weights, producing channel-like structures that are independent of the training data and whose width then oscillates periodically from layer to layer. The authors derive effective dynamics for nodal connectivities (Eq. 3) and for channel amplitudes (Eq. 4) using a path-activity formalism, solve these equations numerically, and compare the predicted channel formation and amplitude correlations with experiments on synthetic cluster data, MNIST, white wine quality, and California Housing. They also report that large initial weight variance suppresses channel formation and that increments in the embedding dimension of hidden representations correlate with increments in channel amplitude.","tokens_in":30568,"tokens_out":7236,"duration_ms":77527,"significance":"If the central claim is established, the paper would provide a data-independent, self-organizing mechanism for nontrivial weight structure in deep networks, with possible consequences for pruning, the lottery-ticket hypothesis, and the debate on emergent capabilities. The empirical side is a real strength: the channel structures are clearly visualized across several datasets, the accessible-node and variance-sweep analyses give quantitative support, and the theory makes a falsifiable prediction (loss of structure at high initial variance) that the authors test. The paper is also clearly written and the path-activity framework is an interesting bridge between neural-network training and pattern-formation theory. However, the \"from first principles\" derivation is not currently controlled: the label-dependent error term is dropped, the constant-U approximation is acknowledged but unchecked, the sqrt(r) closure is justified by the same correlation that is later presented as validation, and the instability is seeded by unmodeled fluctuations in the growth-rate constants.","major_comments":[{"comment":"The derivation of the connectivity dynamics drops the label-dependent training error. Eq. (S41) defines the single-sample weight increment with the factor ηΔY^(m), but from Eq. (S79) onward this factor is absent, and the effective growth rates c_j in Eqs. (S91)–(S94) contain only U, sums of weights, and normalization factors. The sign and magnitude of the effective coupling are therefore not derived from gradient descent; the positivity of c_j is assumed in §3.4 rather than demonstrated, and the claimed independence from training data follows only because the data dependence has been discarded. This is the load-bearing step for the main theoretical claim. The authors should either carry ΔY through the coarse-graining, explicitly allowing for the fact that the batch-averaged ΔY U is not sign-definite for generic targets, or restrict the claim to a regime in which ΔY is a known positive common prefactor and verify that regime empirically.","section":"Supplementary Theory, §3.2 (Eqs. S39–S94)"},{"comment":"The approximation U_t^s ≈ |U| with a fixed sign is stated to be an important assumption, and it is justified only by the heuristic that weights grow from very small initial values so that the output reaches order one. The argument does not quantify the variation of U across node pairs as training proceeds, does not control the sample dependence in Eq. (S36), and does not address the effect of partially inactive paths. Since this assumption enters the derivation of Eq. (S94) and hence Eq. (3) of the main text, the morphological instability is not derived under controlled conditions. A quantitative check of the distribution of U_t^s during early training would be needed before the instability can be called generic.","section":"Supplementary Theory, §3.4, Eq. (S105)"},{"comment":"The substitution Ω_in/out ≈ sqrt(r_j) is justified in the supplement by the empirical correlation between Ω_in and Ω_out, and the same correlation is later presented in Fig. 2e as evidence for the theory. This creates a self-consistency loop: the connectivity equation whose instability predicts channel formation already encodes the correlation that the theory is supposed to explain. The empirical observations are not invalidated by this, but the claim that the instability is predicted from first principles is weakened. The authors should test the sensitivity of the instability to this closure approximation, for example by solving Eq. (3) with Ω_in and Ω_out kept as separate variables.","section":"Supplementary Theory, §3.2, before Eq. (S93); main-text Fig. 2e"},{"comment":"In the linear stability analysis, the perturbation δc_j of the growth-rate constants is introduced by hand, and the instability criterion is δc_j > ⟨δc⟩. No dynamics for c_j are derived, so the instability is seeded by unmodeled fluctuations rather than by the weight dynamics themselves. The normalization argument ⟨δr⟩ ≈ 0 does not specify how δc_j emerges from the initial weight statistics or from finite-size effects. At minimum, the authors should show that the distribution of δc_j generated by their initialization has support on the required side of the instability threshold, or provide a model for the evolution of c_j.","section":"Supplementary Theory, §4.3, Eq. (S117)"}],"minor_comments":[{"comment":"The right-hand side max_m {N_active(l)} does not contain m; either define N_active(l,m) or write the maximum over samples explicitly.","section":"Methods, 'Calculation of the embedding dimension' (Eqs. 5–6)"},{"comment":"The amplitude variable is introduced as a_l ≡ N Σ_j r_j^{(l)}, while the supplement defines R^{(l)} = Σ_j r_j^{(l)} and then a = N R; the two notations should be reconciled to avoid confusion.","section":"Main text, 'Periodic channel amplitudes'"},{"comment":"The phrase 'The inlay shows' should be 'The inset shows', and 'the neural networks was trained' should be 'the neural network was trained'.","section":"Fig. 3 caption"},{"comment":"The theory assumes zero biases and plain gradient descent, while the experiments use biases and the Adam optimizer; please state explicitly whether and why the theory is expected to carry over to Adam and mini-batch training.","section":"Methods, 'Training of deep neural networks'"},{"comment":"There is a stray period in the definition of c_left ('.cL'); additionally, the symbols cR_j and cL_j are reused with different meanings in Eqs. (S91)–(S92) and Eqs. (S97)–(S100), which should be flagged to prevent confusion.","section":"Supplementary Eq. (S178)"},{"comment":"The statement that 'weights in adjacent layers are positively coupled' is presented before the sign of ΔY and U has been controlled; it should be rephrased as merely 'coupled' or qualified by the assumptions of the derivation.","section":"Introduction, 'Channel morphologies'"}],"recommendation":"major_revision","confidential_remarks":"The paper's empirical observations are interesting and likely salvageable, but the theoretical claim as stated in the abstract and introduction is stronger than what the derivation supports. A revision that either repairs the ΔY/U steps or reframes the contribution as a phenomenological coarse-grained model with explicit assumptions would be more honest and would make the contribution much easier to evaluate. The security-related framing in the abstract is speculative relative to the technical content and could be toned down."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, the empirical observation is real: low-variance-initialized ReLU networks do self-organize into channel-like structures whose widths oscillate across layers, and this happens on multiple datasets and disappears at high initialization variance. That is a new and potentially useful finding, and the figures are convincing. Second, the theory that is sold as the main event does not hold up. The stress-test note is right. The single-sample weight increment in Eq. (S40) contains the label-dependent error ΔY, but when the supplement derives the connectivity dynamics in Section 3.2, that factor vanishes without comment. From Eq. (S79) onward, the U-terms contain only inputs and path weights, not the residual. So the effective couplings c_j in Eq. (S94) are never shown to have the sign or magnitude that the instability requires. The homogeneous-state instability is a property of a data-independent update, not of gradient descent on the actual loss. That is not a small technical gap; it breaks the \"derived from first principles\" claim.\n\nWhat the paper does well beyond the experiments: the coarse-grained connectivity variable r_j = Ω_in Ω_out is a sensible object, and the intralayer repression term in Eq. (3) is a plausible phenomenological model once the ΔY issue is set aside. The comparison with lottery tickets and pruning is also a fair and useful discussion, not overclaimed.\n\nThe soft spots beyond the ΔY drop: the constant-U approximation in S105 is a heuristic that the authors admit is an assumption; the Ω≈√r approximation in the derivation uses the very correlation in Fig. 2e that the paper presents as evidence, which is a mild but real self-consistency loop; the experiments use Adam while the theory is written for SGD; and the amplitude-oscillation equation (4) is a highest-order truncation whose rigorous status is unclear. None of these individually are fatal if the ΔY issue were fixed, but together they mean the theoretical contribution is not first-principles derivation.\n\nThis paper is for a reader who wants to see that neural networks can spontaneously develop data-independent weight morphology, and who will treat the theory as a phenomenological model until the derivation is repaired. The empirical part deserves a serious referee, and the paper should not be desk-rejected. But the authors need to either include the ΔY dependence explicitly in the derivation or reframe the theory as a data-independent instability that competes with the gradient signal. As submitted, I would not rely on the theory, and I would not cite the derivation in my own work.","headline":"Genuine empirical finding, but the 'first principles' derivation silently drops the training error term, so the predicted instability is not actually derived from gradient training.","tokens_in":31025,"tokens_out":2025,"would_cite":false,"duration_ms":22527,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that gradient descent training destabilizes the homogeneous weight state of deep networks, driving weights into data-independent channel morphologies that oscillate in width with a two-layer period.","keywords":["emergence","morphogenesis","machine learning","deep neural networks","weight morphology","morphological instability","self-organization","channel structures"],"falsifier":"Compute the active-path sums $U^t_s$ from the path-activity formula on a small trained network; if their magnitudes or signs differ substantially across node pairs, the uniform-growth-rate approximation collapses and the predicted channel instability should not appear in a network trained from the paper's initialisation.","tokens_in":29987,"feed_emoji":"🧠","tokens_out":15929,"duration_ms":133613,"temperature":0.7,"pith_summary":"The paper argues that ordinary gradient-based training makes deep neural networks self-organize into patterned weight structures, and that this pattern formation is independent of the training data. Treating the network as a many-particle system, the authors define a nodal connectivity $r^{(l)}_n$, the product of a node's in- and outgoing weight fractions, and derive its effective time evolution from a path-activity representation of the network output. The homogeneous, near-random initial state is a fixed point of that dynamics, but it is unstable: small perturbations grow into layers split into strongly and weakly connected nodes, and these line up into channel-like structures running from input to output. At later times, higher-order coupling between neighbouring layers modulates the channel width, producing an oscillatory pattern with a two-layer period. The paper verifies these predictions numerically on four datasets and reports that the oscillating channel width tracks the effective dimensionality of the hidden data representations.","feed_headline":"Deep networks spontaneously grow periodic weight channels","feed_subtitle":"A condensed-matter calculation shows the uniform weight state is unstable — independent of the training data.","key_machinery":"The nodal connectivity $r^{(l)}_n = \\Omega_{\\mathrm{in}}(n,l)\\cdot\\Omega_{\\mathrm{out}}(n,l)$, the product of a node's incoming and outgoing absolute-weight fractions, is the morphological unit of the theory. Its effective dynamics are derived from the path-activity formalism, which rewrites the network output as a sum over all input-to-output paths of the product of weights along each path, with ReLU activations absorbed into binary path activities; this makes weight dynamics tractable as coupled path statistics. A linear perturbation analysis of the resulting connectivity equation yields the channel-forming instability, while a second-order amplitude equation for $a_l = N\\sum_j r^{(l)}_j$ captures the interlayer inhibition that produces the period-two modulation.","core_discovery":"The central claim is that the homogeneous state of a deep feedforward network — the near-uniform, low-variance weight configuration at initialisation — is morphologically unstable under gradient-based training, and that this instability produces periodic weight morphologies. Starting from the path-activity representation, in which the output is a sum over all input-to-output paths of the product of weights along each path with ReLU nonlinearities absorbed into binary activities, the authors coarse-grain to the connectivity $r^{(l)}_j = \\Omega_{\\mathrm{in}}(j,l)\\,\\Omega_{\\mathrm{out}}(j,l)$ and derive $$\\frac{$dr^{{(l)}}$_j}{dt} \\approx c_j $r^{{(l)}}$_j\\bigl(1-\\sqrt{$r^{{(l)}}$_j}\\bigr) - $r^{{(l)}}$_j \\sum_{i\\neq j} c_i\\sqrt{$r^{{(l)}}$_i},$$ where the first term is bounded growth with rate $c_j$ and the second is a repulsive interaction with other nodes in the layer. The homogeneous state $r_j = 1/N^2$ with equal growth rates is an exact fixed point, but a node whose rate exceeds the connectivity-weighted average grows while others shrink, giving a bimodal distribution of connectivities in each layer. Because strongly connected nodes in one layer attach to strongly connected nodes in adjacent layers, this produces a channel of highly connected nodes spanning the network. Including the higher-order dependence of the growth rates on neighbouring layers yields amplitude dynamics for $a_l = N\\sum_j r^{(l)}_j$ whose dominant term, $a_l(1-\\sqrt{a_l})(c_R\\sqrt{a_{l+1}}+c_L\\sqrt{a_{l-1}})$, is always negative, so each layer's channel width is suppressed by the widths of its neighbours, producing a periodic, anticorrelated modulation with a two-layer period. Numerical experiments on synthetic cluster data, MNIST, wine quality, and California Housing confirm channel formation and the oscillation, and show that channel formation is absent when initial weights have large variance and in poorly trained networks.","pith_inferences":["An immediate test the paper does not run: train the same architectures on randomly permuted labels. If the instability is truly data-independent, the same channels and two-layer oscillations should appear, cleanly separating morphology from data-driven structure.","If channel amplitude tracks embedding dimension generically, then width scheduling — starting wide, narrowing, then widening — should either reinforce or fight the emergent oscillation; comparing learning curves under such schedules would turn the observed correlation into a design principle.","The derivation is anchored to squared-error loss and gradient descent, so a natural stress test is whether the instability survives cross-entropy losses and other optimisers; the mechanism of nearest-neighbour weight feedback through shared paths suggests it should, but the paper does not claim this.","The paper's interpretation of emergent structures as potential pruning targets suggests that pruning by self-organized channel membership, rather than by per-weight magnitude, might isolate trainable subnetworks more directly than magnitude-based lottery-ticket searches."],"forward_implications":["Networks trained from small-variance initialisations will generically develop a channel of strongly connected nodes, so the structure of a trained network is partly a self-organized by-product of training rather than a direct image of the data.","The predicted period-two oscillation means the effective width of the network's computational backbone alternates layer by layer; measuring layer-wise activity counts should reveal the same alternation in any sufficiently deep feedforward network.","Channel formation is absent for large initial-weight variance and in poorly trained networks with accuracy below 20 percent, so the morphology serves as a structural marker of successful learning.","Because channel amplitude and the embedding dimension of hidden representations oscillate together, the emergent morphology is tied to periodic expansion and compression of the data representation, connecting self-organization to the function of kernel-like and autoencoder-like transformations.","The mechanism transfers to any feedforward architecture whose output is a sum over paths, including convolutional networks and, via an analogous path framework, transformers, and to sigmoidal activations approximated piecewise-linearly."],"supporting_citations":[{"why":"Provides the path-activity representation of network output as a sum over paths, the starting point for the derivation of weight and connectivity dynamics.","marker":"[12]"},{"why":"The cited source for the effective repressive interaction between neurons in the same layer that shapes the connectivity dynamics in Eq. (3).","marker":"[41]"},{"why":"The MNIST handwritten-digit benchmark, one of the four datasets used to verify channel formation and amplitude oscillations.","marker":"[42]"},{"why":"The white-wine-quality dataset used as a nontrivial benchmark for the predicted weight morphologies.","marker":"[43]"},{"why":"The California Housing regression dataset used as a further benchmark for the predicted weight morphologies.","marker":"[44]"},{"why":"Supplies the method for computing the embedding dimension of hidden data representations whose layer-to-layer increments are correlated with channel amplitude.","marker":"[45]"},{"why":"Cover's theorem, invoked to explain why the oscillating embedding dimension could aid learning through repeated expansion and compression of representations.","marker":"[48]"},{"why":"Establishes an analogous path framework for transformers, grounding the paper's claim that the instability extends beyond fully connected feedforward networks.","marker":"[51]"}],"fun_headline_variants":["Deep nets spontaneously grow periodic weight channels","Training instability births periodic weight structures","Neural nets self-organize into periodic weight channels","Data-independent periodic weights emerge in deep nets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The model's central approximation is that, early in training, the summed contributions of all hidden paths connecting two nodes ($U^t_s$) have the same magnitude and sign for every pair of nodes, so the growth rates in the connectivity equation are uniform across nodes; this uniformity is justified only by the heuristic that weights start tiny and then grow, and if those path sums actually vary with the data or with weight evolution, the predicted instability may not occur in real training.","fun_headline_variants_meta":{"raw":{"variants":["Deep nets spontaneously grow periodic weight channels","Training instability births periodic weight structures","Neural nets self-organize into periodic weight channels","Data-independent periodic weights emerge in deep nets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000201,"raw_usage":{"total_tokens":1438,"prompt_tokens":1061,"completion_tokens":377,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":677,"completion_tokens_details":{"reasoning_tokens":323}},"tokens_in":677,"tokens_out":377,"duration_ms":4463,"temperature":1.0,"reasoning_tokens":323,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:13:12.335307+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the active-path sums $U^t_s$ from the path-activity formula on a small trained network; if their magnitudes or signs differ substantially across node pairs, the uniform-growth-rate approximation collapses and the predicted channel instability should not appear in a network trained from the paper's initialisation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the path-activity representation of network output as a sum over paths, the starting point for the derivation of weight and connectivity dynamics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The cited source for the effective repressive interaction between neurons in the same layer that shapes the connectivity dynamics in Eq. (3)."},{"cited_title":"The MNIST database of handwritten digits","cited_arxiv_id":null,"evidence_quote":"The MNIST handwritten-digit benchmark, one of the four datasets used to verify channel formation and amplitude oscillations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The California Housing regression dataset used as a further benchmark for the predicted weight morphologies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the method for computing the embedding dimension of hidden data representations whose layer-to-layer increments are correlated with channel amplitude."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cover's theorem, invoked to explain why the oscillating embedding dimension could aid learning through repeated expansion and compression of representations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes an analogous path framework for transformers, grounding the paper's claim that the instability extends beyond fully connected feedforward networks."}],"review_version":1}