{"id":"4f694fbd-ed7c-4342-88fb-b3c4affa7753","arxiv_id":"2511.10362","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The adjacency-matrix reformulation of deep-linear gradient flow reveals a quotient-space structure of the loss landscape, but the proof that arcs between critical points determine stable and unstable manifolds is incomplete.","lead":"This paper rewrites the training equations of deep linear networks as matrix flows using the network's adjacency matrix, giving a compact view of a loss landscape made only of global minima and saddles. It is worth reading for a new structural perspective on deep linear training, though its most novel claim about unstable directions rests on an unproven assumption.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 24's arc-based stability classification is unsupported and self-contradicted by Example 6.2.1.","rationale":"The Reader's verdict correctly identifies the arc-existence and transfer-of-monotonicity gap in Theorem 24. My stress-test confirms this is the load-bearing weak point, and adds a concrete internal inconsistency: the theorem's stable-submanifold count for A=0 contradicts the paper's own Example 6.2.1, where the origin has a stable direction w₁=−w₂. This is not a mere gap in exposition; it shows that the theorem cannot be interpreted as a statement about the usual stable/unstable manifolds of the gradient flow. The survey portions of the paper—adjacency-matrix reformulation, conservation laws, Hessian quadratic form, known critical-point classifications—are largely standard and appear reliable; the original contribution around quotient-space arcs and stability classification is what needs correction. Because the flaw is in the newly claimed stability classification rather than in the survey content, a conditional acceptance is appropriate: the central claim should be revised to say that the arcs provide a monotone path in loss space, not that they determine dynamical stable/unstable submanifolds, or a proof of the transfer from arcs to flow-invariant manifolds should be supplied.","tokens_in":49892,"tokens_out":5240,"duration_ms":59319,"concrete_test":"Analytically settle the contradiction by applying Theorem 24 item 3 to the example of §6.2.1. For dx=dy=1, h=2, the critical point A=0 has signature S=∅, so the theorem gives 2^1−1=1 unstable submanifold and 2^0−1=0 stable submanifolds. But linearizing (62) at w=0 gives J=[[0,σ],[σ,0]], whose stable eigenvector is (1,−1); the line w₁=−w₂ is forward invariant and every point on it converges to the origin. This is a stable submanifold, directly contradicting the theorem's count. If the authors intend a different definition of 'stable submanifold,' it must be stated; otherwise Theorem 24 cannot be used to justify the headline claim.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's headline claim is that the adjacency-matrix construction 'allows to easily determine stable and unstable submanifolds at the saddle points, even when the Hessian fails to obtain them.' This rests on Theorem 24 (§4.6). The theorem considers an arc A(α) in the fiber φ⁻¹(Φ(α)) connecting two critical points, observes that L(A(α)) is monotone in α when the signatures are nested (S₁⊂S₂), and then states that 'A₁ has an unstable submanifold, while A₂ has a stable one.' But α is not time; the authors explicitly concede that these arcs 'cannot be trajectories of the gradient flow.' Monotonicity of a Lyapunov function along an arbitrary curve connecting two equilibria does not, by itself, imply the existence of invariant manifolds of the flow. No argument is given that the curve can be deformed into, or is tangent to, actual trajectories, nor that the convergence along it is governed by the gradient-flow dynamics. Moreover, the claimed counts in item 3 are internally contradicted by Section 6.2.1. For the single-hidden-node example, dy=1 and the origin has signature S=∅, so item 3 predicts 2^k−1 = 0 stable submanifolds. Yet the same paper shows that the line w₁=−w₂ is a stable submanifold of the origin, with trajectories converging to it. Thus either 'stable submanifold' is being used with a nonstandard meaning that is never defined, or the theorem's counts are wrong. Since this theorem is the main original basis for the abstract's stability claim, the central claim is not currently supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper surveys the gradient flow dynamics of deep linear neural networks with quadratic loss, reformulating the standard layer-wise equations (9) as a single matrix ODE (12) for the block-shift adjacency matrix A = blk1(W1,...,Wh). Sections 3–5 derive structural properties (nilpotency, conservation laws, Hessian quadratic form, aligned/decoupled/balanced reductions) and review known critical-point and convergence results from [1,10,24,40] in this notation. The claimed original contribution is Theorem 24, which uses monotonicity of the loss along arcs in the fibers of the quotient map φ(A)=A^h to classify stable and unstable submanifolds of critical points, advertised in the abstract as allowing one to determine stable and unstable submanifolds 'even when the Hessian fails to obtain them.'","tokens_in":50305,"tokens_out":8483,"duration_ms":83478,"significance":"The adjacency-matrix formulation is a useful organizing tool: it compresses the multi-layer equations, makes conservation laws transparent (Prop. 6), and simplifies several second-order computations (Props. 20–23). As a survey, the paper collects and reformulates a substantial body of recent results, and the derivations in Sections 3–5 are mostly clean. The claimed stability classification, however, is not established. Theorem 24 infers dynamical invariant manifolds from monotonicity of a Lyapunov function along an arbitrary connecting arc, and the authors explicitly note that the arc need not be a gradient trajectory. The theorem is also contradicted by the paper's own Example 6.2.1. Since this theorem is the main original basis for the abstract's central claim, the significance of the paper as a research contribution is conditional: if Theorem 24 is removed or downgraded to a statement about loss monotonicity, the survey remains valuable but the advertised new tool is lost.","major_comments":[{"comment":"The theorem asserts that monotonicity of L(A(α)) along an arc A(α) connecting critical points implies that 'A1 has an unstable submanifold, while A2 has a stable one.' The text immediately after the proof states that these arcs 'cannot be trajectories of the gradient flow' and that 'we do not know how the arcs A(α) relate to the trajectories of (12).' Monotonicity of a Lyapunov function along an arbitrary curve between equilibria does not, by itself, imply the existence of invariant manifolds of the flow. A proof that the arc can be realized by—or is tangent to—actual gradient trajectories is needed. As written the inference is a non sequitur, and it is the load-bearing step for the abstract's claim about determining stable/unstable submanifolds.","section":"§4.6, Theorem 24(2) and following paragraph"},{"comment":"For the single-hidden-node example, dy=1 and the origin has signature S=∅. Item 3 of Theorem 24 predicts 2^0−1=0 stable submanifolds for this critical point. Section 6.2.1, however, explicitly identifies the line w1=−w2 as the stable submanifold of the origin, with trajectories converging to it. Either 'stable submanifold' is being used in a nonstandard, undefined sense, or the count in Theorem 24(3) is wrong. This internal contradiction is decisive for the advertised stability classification.","section":"§4.6, Theorem 24(3) vs. §6.2.1"},{"comment":"The theorem assumes an arc A(α) ∈ φ^{-1}(Φ(α)) with A(0)=A1 and A(1)=A2. No construction or existence proof is supplied, and the parametrization (32)–(33) does not make such a continuous lift with prescribed endpoints evident. Existence must be proved, or explicitly stated as a hypothesis with its limitations analyzed; otherwise the classification applies only to arcs whose existence is itself unverified.","section":"§4.6, statement of Theorem 24"}],"minor_comments":[{"comment":"'Level sets are unbounded invariant sets of critical points' is imprecise: the full level sets of L contain noncritical points. Only the critical level sets M_c of Prop. 15 are invariant and composed of critical points. Suggest rephrasing.","section":"Abstract and §4.3"},{"comment":"The symbol S is used both for the indicator set {i : s_i=1} and for the diagonal matrix S=diag(s), sometimes in the same passage (e.g., Prop. 14). A separate symbol for the diagonal matrix would reduce ambiguity.","section":"§2.1 and §4.2"},{"comment":"The claim that the graph of ˙z = z^{h-1}(σ_j − z^h) has '3 critical points ±σ_j and 0' is not correct in general. The nonzero critical points are the real roots of z^h=σ_j; there are two of them only when h is even, and they are not ±σ_j unless h=1 or the notation is otherwise defined.","section":"§5.4, Eq. (61)"}],"recommendation":"major_revision","confidential_remarks":"The survey part is solid and the adjacency-matrix reformulation is genuinely convenient. The original Theorem 24, however, is the advertised contribution and it is currently unsupported: the arc argument does not establish dynamical invariant manifolds, and the theorem's counts are contradicted by the paper's own example. I would ask the authors either to supply a dynamical proof of the stability classification or to revise the abstract and Theorem 24 to state only monotonicity of L along arcs. With that downgrade the paper is a competent survey, but its original contribution becomes much smaller."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: this is a well-organized survey of deep linear network gradient flows, reformulated in terms of the adjacency matrix. The reformulation is clean and the quotient-space viewpoint is genuinely useful. But the paper's headline claim—that the adjacency construction easily determines stable and unstable submanifolds at saddle points—rests on Theorem 24, and that theorem does not hold up.\n\nWhat's new: the flow (12), Prop. 17 (isolated critical points in the quotient), and Thm. 24 (arc classification). The first two are solid as far as I can tell; the derivations in Sections 3–5 are careful, and the survey itself covers Kawaguchi, Achour, Chitour et al. accurately. The conservation-law interpretation via JQ(t)=C is neat. No self-citations, no fitted parameters, no circularity. This is a genuine contribution to the deep-linear-networks subfield.\n\nThe problem is Theorem 24. It builds an arc A(α) between two critical points in a fiber, shows the loss is monotone in α when signatures are nested, and then concludes that one critical point has an unstable submanifold and the other a stable one. But α is not time. The arc is explicitly not a trajectory of the gradient flow, and the authors concede they don't know how it relates to the flow. Monotonicity of a Lyapunov function along an arbitrary curve doesn't imply invariant manifolds. That's a missing argument, not a minor gap.\n\nWorse, item 3 of the theorem contradicts the paper's own Example 6.2.1. For dy=1 and the origin (S=∅, k=0), the theorem gives 2^0−1=0 stable submanifolds. The example then shows the line w1=−w2 is a stable submanifold of the origin. So either 'stable submanifold' is being used in a nonstandard way (never defined) or the count is wrong. Either way, the central stability claim is currently unsupported.\n\nWhere does that leave us? The survey is reliable and worth reading as a reference. The new structural results in Prop. 17 and the matrix-ODE formulation are useful. But the abstract's promise about stable/unstable submanifolds overreaches. If the authors revise Theorem 24—either by proving the arc-to-flow transfer or by withdrawing the stability-count claim—I'd be happy to engage with it. As it stands, I'd send it to review rather than desk-reject, because the survey is solid and the flaw is isolated, but the referee should push hard on this theorem.","headline":"The survey is solid, but the headline stability theorem (Thm. 24) is unsupported and contradicted by the paper's own example.","tokens_in":50732,"tokens_out":4094,"would_cite":false,"duration_ms":39006,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["34D05","15A18","37C10","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that the entire gradient-flow training of a deep linear network can be rewritten as a single matrix ODE for the network's block-shift adjacency matrix, and that from this representation one can read off stable and unstable","keywords":["gradient flow","deep linear neural networks","adjacency matrix","loss landscape","saddle points","quotient space","isospectral ODE","conservation laws"],"falsifier":"For a small two- or three-layer network, compute the image of the map A ↦ Aʰ and check whether the linear interpolation Φ(α) = (1−α)Φ₁ + αΦ₂ between two critical products stays in that image for all α ∈ [0,1]; exhibit any nested pair for which some α is not the h-th power of any block-shift matrix. That would sever the arc on which Theorem 24's monotonicity argument rests.","tokens_in":49813,"feed_emoji":"🧮","tokens_out":4756,"duration_ms":52674,"temperature":0.7,"pith_summary":"The paper argues that the gradient flow of a deep linear network is naturally one polynomial, nilpotent, isospectral matrix ODE written on the adjacency matrix of the network, whose h-th power is the input-output map. In this representation the loss is a quadratic function of that power, and every critical point is labeled by which singular values of the data it has learned. The central claim is that the quotient structure induced by the power map lets one determine stable and unstable submanifolds of every critical point by drawing arcs between critical points and checking monotonicity of the loss, even where the Hessian is singular and silent. If right, this gives a global, non-local map of the loss landscape and explains why gradient descent navigates infinitely many degenerate saddles successfully.","feed_headline":"One adjacency matrix exposes all saddles of deep linear training","feed_subtitle":"A block-shift matrix rewrites training so that critical points form 2^d levels and saddle directions are read off without the Hessian.","key_machinery":"The key object is the block-shift adjacency matrix A = blk₁(W₁,…,Wₕ), a matrix whose only nonzero blocks sit one diagonal below the main block diagonal, so that the input-output map equals Aʰ. Two derived structures carry the argument: the quotient map φ(A)=Aʰ, whose fibers collect all networks with the same critical value, and the arc A(α) ∈ φ⁻¹((1−α)Φ₁+αΦ₂) connecting two critical points, along which monotonicity of L signals stable and unstable submanifolds. The block-shift SVD alignment and the conserved quantities JQ=C justify restricting to tractable invariant submanifolds such as balanced and decoupled dynamics.","core_discovery":"The central discovery is that the gradient flow of a deep linear network is a matrix ODE Ȧ = Σⱼ (A^{h−j})ᵀ (E−Aʰ)(A^{j−1})ᵀ on the block-shift adjacency matrix A = blk₁(W₁,…,Wₕ), with E encoding the data singular values. Because A is nilpotent, products of layer matrices become powers of A, and the loss becomes ½‖E−Aʰ‖²_F. The critical points are exhausted by binary signatures s ∈ {0,1}^{d_y}: the signature s=1 gives global minima, every other signature gives a saddle point, strict or non-strict depending on the parametrization A = P(A₁+Z)P⁻¹. The paper proves that through the quotient map φ(A)=Aʰ and its fibers φ⁻¹(Φ), one can classify stable and unstable submanifolds of all critical points","pith_inferences":["If the existence of the arc A(α) could be established in general, the monotonicity criterion would provide a global partial order on critical points and a topological proof of saddle-to-saddle learning, potentially extending beyond quadratic losses where Hessian methods fail.","The quotient/fiber picture suggests a testable bridge to matrix sensing and matrix completion: the greedy low-rank principle should be expressible as a statement about which arcs between critical points are realizable in the fiber, not merely which critical points are attractive.","A concrete empirical check: in a three-layer near-balanced network, trajectories should approximately follow the monotone arcs of Theorem 24; if they do, the arcs are effective descriptors of learning despite not being actual flow trajectories.","The paper's claim that implicit low-rank regularization is confined to near-balanced initialization, rather than universal, invites an experiment sweeping initialization scale and measuring the rank of A along trajectories to delineate where the greedy low-rank story holds."],"forward_implications":["Every trajectory converges to a single critical point, and the only attractors are global minima; every non-global critical point is a saddle, including non-strict saddles where the Hessian is positive semidefinite.","In the quotient space Aʰ there are exactly 2^{d_y} isolated critical points, one per critical value, with critical values given by distinct partial sums of the squared data singular values.","For a critical point of signature S, the paper counts 2^{d_y−|S|}−1 unstable submanifolds and 2^{|S|}−1 stable ones, without computing eigenvalues of the Hessian.","Under 0-balanced initialization, all layers synchronize their singular values and the critical point is maximally tightened; near-balanced initialization generically yields sequential learning from the largest to the smallest singular value when the scalar channels start equally.","Restricting to block-shift diagonal or decoupled dynamics reduces the gradient flow to d_y independent scalar generalized logistic ODEs, reproducing the plateau-and-discovery behavior seen in simulations."],"fun_headline_variants":["Block-shift matrix exposes every saddle in deep linear flow","Deep linear training reduced to one nilpotent matrix ODE","Adjacency matrix reveals all critical points, no Hessian needed","Quotient map sorts deep linear saddles into binary signatures","One block matrix maps the full loss landscape of deep linear nets"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The argument's load-bearing premise is that between any two critical points with nested signatures there exists a continuous arc of adjacency matrices whose input-output power interpolates linearly between the two critical values, and that monotonicity of the loss along that arc transfers to the actual training flow; the paper assumes this arc rather than proving it exists.","fun_headline_variants_meta":{"raw":{"variants":["Block-shift matrix exposes every saddle in deep linear flow","Deep linear training reduced to one nilpotent matrix ODE","Adjacency matrix reveals all critical points, no Hessian needed","Quotient map sorts deep linear saddles into binary signatures","One block matrix maps the full loss landscape of deep linear nets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000159,"raw_usage":{"total_tokens":1119,"prompt_tokens":850,"completion_tokens":269,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":198}},"tokens_in":594,"tokens_out":269,"duration_ms":3650,"temperature":1.0,"reasoning_tokens":198,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T22:26:29.518340+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For a small two- or three-layer network, compute the image of the map A ↦ Aʰ and check whether the linear interpolation Φ(α) = (1−α)Φ₁ + αΦ₂ between two critical products stays in that image for all α ∈ [0,1]; exhibit any nested pair for which some α is not the h-th power of any block-shift matrix. That would sever the arc on which Theorem 24's monotonicity argument rests.","supporting_citations":[],"review_version":1}