{"id":"9383a0b7-3476-4fdf-9d8d-f925f249ea03","arxiv_id":"2511.02003","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"The paper reframes SGD training of deep networks as a local Lagrangian with data confined to the boundaries, but the advertised energy continuity equation is absent from the body.","lead":"This paper reorganizes the training dynamics of deep networks into a data-independent 'bulk' and a data-dependent 'boundary' term. It claims this exposes local structure and even an energy conservation law for learning, but the conservation law is never derived in the text.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract promises an energy continuity equation that the paper never derives; this missing advertised consequence is the most load-bearing gap.","rationale":"The reader's weakest_assumption concerns the deterministic gradient-flow limit of SGD, which is a legitimate modelling concern. However, I identify a more direct and unequivocal problem: the paper explicitly claims to derive an energy continuity equation, yet no such equation is present in the manuscript. This is an internal inconsistency between the abstract and the body, and it sacrifices the paper's central advertised physical consequence. The algebraic bulk–boundary decomposition (Eq. 7) appears correct as a formal rearrangement, but it is insufficient to support the paper's claims without the missing physical derivation. Therefore the reader's REJECT verdict remains appropriate, but for the reason of the missing energy equation rather than the stochasticity issue.","tokens_in":7917,"tokens_out":7307,"duration_ms":82504,"concrete_test":"Run a full-text search for 'energy', 'continuity', 'conservation', and 'Noether'. If no derivation appears outside the abstract, the claim is unsupported. Additionally, attempt to derive a continuity equation from the bulk Lagrangian (Eq. 10) via Noether's theorem under the discrete symmetry m→m+1; if no conserved current exists or its conservation depends on the boundary term, the advertised consequence cannot hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract states: 'As a physical consequence of locality and homogeneity, we derive the energy continuity equation within a deep neural network.' However, the full text never derives such an equation. Sections 'Bulk–Boundary Decomposition', 'Field Description', and 'Discussion and Outlook' contain no equation, no Noether current, and no conservation law that could be interpreted as an energy continuity equation. The closest candidate is the discrete translational symmetry m→m+1 of L_bulk, but no associated current or continuity equation is constructed. Even granting the algebraic identity (Eq. 7) and the deterministic gradient-flow limit (Eq. 3), the paper fails to deliver its central advertised physical consequence. This is not a minor omission: the energy continuity equation is presented as a key result, and its absence leaves the framework's physical significance unsubstantiated.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a \"bulk–boundary decomposition\" of the training Lagrangian of a deep neural network. Starting from a stochastic-gradient-descent update (Eq. 2) and a continuous-time deterministic gradient-flow approximation (Eq. 3), the authors change variables from weights and biases to weights and pre-activations (Eq. 6), obtaining a Lagrangian (Eq. 7) split into a data-independent bulk term and a data-dependent boundary term. They then sketch continuum field-theory Lagrangians for the boundary and bulk sectors (Eqs. 9–10) and claim, in the abstract, that locality and homogeneity imply an energy continuity equation. The paper is short, programmatic, and explicitly leaves the complete field-theoretic construction to future work.","tokens_in":8176,"tokens_out":6735,"duration_ms":73209,"significance":"The algebraic identity in Eq. (7) is checkable and appears valid; promoting pre-activations to dynamical variables to expose adjacent-layer locality is a reasonable idea. If the advertised energy continuity equation were actually derived and the continuum limit made rigorous, the framework could offer a useful physical perspective on deep-learning dynamics. However, as submitted, the central advertised consequence is absent from the text, the field-theory section is explicitly illustrative rather than derived, and the connection to stochastic gradient descent is not justified. The paper does not currently establish the new framework it claims.","major_comments":[{"comment":"The abstract states: 'As a physical consequence of locality and homogeneity, we derive the energy continuity equation within a deep neural network.' The main text contains no such equation, no Noether current, and no conservation law. The only symmetry discussed is the discrete translational symmetry m→m+1 of L_bulk, but no continuity equation is constructed from it. This is the advertised physical consequence, so the paper's central claim is unsupported.","section":"Abstract; Sections 'Bulk–Boundary Decomposition' and 'Field Description'"},{"comment":"The continuum Lagrangians in Eqs. (9)–(10) are not derived. Footnote 5 states that the complete Lagrangian depends on a specific coordinate choice 'which we do not detail here' and that only representative terms are shown. No lattice-spacing expansion, no coordinate assignment for X, Y, z, w, and no limit is specified. The reader cannot reproduce or verify Eqs. (9)–(10); they are at most a sketch, not a field-theoretic formulation.","section":"Field Description, Eqs. (9)–(10), footnote 5"},{"comment":"The continuous-time limit Eq. (3) is not the SGD dynamics defined in Eq. (2). In SGD the pair (X, Y) is drawn randomly at each step, so the loss ℓ is a stochastic function of time; replacing it by a fixed-sample deterministic gradient flow changes the process. Consequently, the Lagrangian L(t) in Eq. (5) and any conservation law derived from it do not describe stochastic gradient descent. This is load-bearing because the paper motivates the entire framework from the SGD formulation.","section":"Fundamentals, Eq. (3)"},{"comment":"The claimed bulk–boundary separation is a definitional regrouping: terms explicitly containing X, Y, or ℓ are assigned to L_boundary, and the remaining terms to L_bulk. The identity is valid, but the paper does not demonstrate that this regrouping yields new analytic power. No observable, conserved quantity, or training-related quantity is computed from the decomposition to show that it 'reveals' or 'governs' the claimed structure beyond the algebra.","section":"Bulk–Boundary Decomposition, Eq. (7)"}],"minor_comments":[{"comment":"The first term in L_boundary appears to be missing the factor 1/2 that follows from expanding (1/2)(b_i^(0))^2 in Eq. (5) for m=0; please check the coefficient.","section":"Eq. (7)"},{"comment":"Footnote 5 contradicts the presentation of Eqs. (9)–(10) as derived results; the text should either label these as a proposed expansion or provide the missing coordinate assignment and expansion.","section":"Footnote 5 and Eqs. (9)–(10)"},{"comment":"The integral kernel W(x,x') is not related explicitly to the discrete weights W_ij^(m); the continuum limit that gives Eq. (8) should be stated.","section":"Eq. (8)"},{"comment":"The term 'homogeneous' is used loosely: L_bulk in Eq. (10) contains terms of different orders in a_x and a_y, so translational symmetry holds only in the continuum limit and for a repeated layer structure, which should be clarified.","section":"Field Description, Eq. (10)"}],"recommendation":"reject","confidential_remarks":"This is a brief programmatic note. The main advertised result, the energy continuity equation, is absent from the manuscript, and the field-theoretic construction is explicitly incomplete. The gap is not a local presentation issue; it is the core of the paper's claimed contribution. I recommend rejection, though the authors may wish to pursue a full derivation and rigorous continuum limit in a future submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things. First, the algebraic decomposition in Eq. (7) is real and checkable: substituting Eq. (6) into the Lagrangian does separate a data-independent bulk from a data-dependent boundary. Second, the abstract's claim that the paper derives an energy continuity equation is false: no continuity equation, Noether current, or conservation law appears anywhere in the text. That is a major gap, not a stylistic one.\n\nWhat the paper does well: the change of variables from biases to pre-activations is carried over from the authors' earlier Phys. Rev. Research paper (Ref. [6]), but the explicit bulk-boundary partition is a clean rearrangement that makes the locality along depth manifest. The writing is clear, and the authors are honest that the field-theoretic part is only a sketch—they explicitly say they show 'representative terms' and that exploring implications is future work. The discrete Lagrangian identity is a nice pedagogical object.\n\nWhere it falls short: the missing energy continuity equation is the advertised physical consequence and is simply absent. The closest candidate is the discrete translational symmetry of the bulk, but no current is constructed. The continuum field theory is partial; Eqs. (9)-(10) are a few expanded terms, not a complete local action. The continuous-time limit (Eq. 3) treats SGD as deterministic gradient flow for a fixed sample, which is a standard idealization but not the stochastic process named in the title. And the novelty is incremental: the core re-coordinatization is identical to Ref. [6], so the new content is the regrouping, which is definitional.\n\nThis paper is for readers interested in physics-inspired reformulations of neural network training. It is not ready for publication as is. But it deserves a serious referee: the identity is checkable, the authors are clear about what they have and haven't done, and the missing derivation could plausibly be supplied. I would send it out, with the clear expectation that the referee will ask for the continuity equation to be derived or the abstract rewritten.","headline":"A checkable algebraic decomposition and an honest sketch of a field theory, but the promised energy continuity equation is missing from the paper.","tokens_in":8610,"tokens_out":3515,"would_cite":false,"duration_ms":36096,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A change of variables reorganizes the training Lagrangian into a data-independent bulk and a data-dependent boundary.","keywords":["bulk-boundary decomposition","neural network training dynamics","Lagrangian mechanics","locality","field theory","stochastic gradient descent","translational symmetry","continuum limit"],"falsifier":"Simulate the discrete SGD updates of Eq. (2) on a small network with a fixed sample and compare the resulting trajectories to the continuous-time flow of Eq. (3) at decreasing learning rates; if the flow does not converge to the discrete dynamics, the continuous-time Lagrangian premise fails. (Alternatively, the paper's own text can be checked: the abstract announces an energy continuity equation that the main text never actually writes down.)","tokens_in":7825,"feed_emoji":"🧠","tokens_out":5665,"duration_ms":61646,"temperature":0.7,"pith_summary":"This paper tries to show that the training dynamics of a deep neural network can be rewritten so that architecture and data act in separate sectors: an interior 'bulk' that depends only on the network's structure and activation functions, and a 'boundary' at the input and output layers where training samples act. The separation is achieved by promoting the network's pre-activations to dynamical variables and eliminating the biases through the network's own recursive relation. Making the depth-direction locality explicit lets the authors translate the discrete network into a field theory, in which the interior has translational symmetry and the boundary is a stochastic interface. If the framework holds, standard tools from statistical physics—continuum limits, symmetry arguments, possibly conservation laws—become available for studying how deep networks learn.","feed_headline":"Training dynamics decompose into architecture bulk and data boundary","feed_subtitle":"Separating architecture from data opens field-theoretic tools for understanding deep learning.","key_machinery":"The key mechanism is a change of variables: the recursive relation z_i^{(m+1)} = sum_j W_ij^{(m)} sigma(z_j^{(m)}) + b_i^{(m)} is solved for the bias, b_i^{(m)} = z_i^{(m+1)} - sum_j W_ij^{(m)} sigma(z_j^{(m)}), and the biases are eliminated as independent parameters. This promotes the pre-activations z to degrees of freedom, rewriting the kinetic and potential terms so that adjacent layers couple only locally, and the loss depends only on the top layer z^{(M)}. The resulting bulk-boundary decomposition (Eq. 7) is then the object from which locality, symmetry, and a field-theoretic continuum limit are extracted.","core_discovery":"The central discovery is an algebraic identity: in the basis where the weights W and the neuron pre-activations z are the independent degrees of freedom, the training Lagrangian of Eq. (5) reorganizes into L = L_bulk + L_boundary. L_bulk couples only adjacent layers and contains no reference to the training data, while L_boundary contains all dependence on the input X and target Y through the loss. This makes manifest that information flow along depth is local, that repeated layer structures give a discrete translational symmetry, and that data enter only at the two ends of the network.","pith_inferences":["If the bulk term is truly data-independent, then two networks with identical architecture and different training tasks share the same interior dynamics; a concrete test is to measure bulk-term statistics across tasks.","The boundary/bulk split suggests that architectural design could be viewed as engineering the bulk Lagrangian, while data curation tunes the boundary, potentially giving a new design principle.","The missing energy continuity equation, promised in the abstract but absent from the text, is the most direct next step: deriving it from Noether's theorem would give a testable conservation law for training.","One could test the field-theoretic limit by training narrow-but-deep local architectures (like the illustrated chain) and checking whether coarse-grained observables obey the continuum equations."],"forward_implications":["One can study architecture-dependent dynamics and data-dependent dynamics separately, isolating the role of the loss from the role of network structure.","The explicit depth-direction locality permits a continuum limit in which the network becomes a field theory with an emergent coordinate along depth.","Repeated layer structures give a discrete translational symmetry in depth, opening the door to momentum/energy conservation arguments and renormalization-group-style analyses.","Boundary stochasticity is naturally a candidate for a statistical-mechanics treatment, potentially linking generalization to effective thermal ensembles.","The decomposition suggests long-range order may emerge from local training interactions, possibly characterizing successfully trained networks."],"fun_headline_variants":["Training dynamics: architecture bulk, data boundary","Bulk-boundary split reveals neural net structure","Deep learning Lagrangian: data only at the edges","Neural nets: local bulk, data-driven boundary","Separating architecture from data in training"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the discrete stochastic SGD updates of Eq. (2) can be replaced by the deterministic continuous-time gradient flow of Eq. (3) on a fixed training sample; if that replacement misrepresents the stochastic sampling, all subsequent Lagrangian and field-theoretic claims inherit the error.","fun_headline_variants_meta":{"raw":{"variants":["Training dynamics: architecture bulk, data boundary","Bulk-boundary split reveals neural net structure","Deep learning Lagrangian: data only at the edges","Neural nets: local bulk, data-driven boundary","Separating architecture from data in training"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00016,"raw_usage":{"total_tokens":986,"prompt_tokens":577,"completion_tokens":409,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":321,"completion_tokens_details":{"reasoning_tokens":340}},"tokens_in":321,"tokens_out":409,"duration_ms":5123,"temperature":1.0,"reasoning_tokens":340,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T00:15:53.265506+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate the discrete SGD updates of Eq. (2) on a small network with a fixed sample and compare the resulting trajectories to the continuous-time flow of Eq. (3) at decreasing learning rates; if the flow does not converge to the discrete dynamics, the continuous-time Lagrangian premise fails. (Alternatively, the paper's own text can be checked: the abstract announces an energy continuity equation that the main text never actually writes down.)","supporting_citations":[],"review_version":1}