REVIEW 3 major objections 6 minor 1 cited by
Geometry of fibers of the multiplication map of deep linear neural networks
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper determines, for tuples of composable matrices of fixed shapes whose product is zero (or a fixed matrix), the codimension $C$ and the number $\theta$ of largest-dimensional irreducible pieces of that variety, and proves both are…
desk verdict Solid Poincaré-series computation of C and θ for zero-product matrix tuples, with a real but probably fixable proof gap in the structural lemma behind the explicit formulas. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument runs through the orbit stratification of the representation space: the group $G = \prod_i \mathrm{GL}_{d_i}$ acts on tuples by change of basis, and for the equioriented type A quiver the $G$-orbits are classified by Kostant partitions, i.e. multiplicities of interval modules $M_{ij}$ in the decomposition of a representation. Lace diagrams, arrays of dots in $N+1$ columns partitioned into intervals, are the combinatorial pictures of these partitions. Three facts make the machinery work: orbit closure is controlled by entrywise inequality of rank patterns; the codimension of an orbit is a bilinear form in the partition coming from $\mathrm{Ext}(M,M)$; and the longest interval module $M_{0N}$ is both injective and projective, so adding copies of it leaves normal slices unchanged. This last fact is what ultimately makes $C$ and $\theta$ independent of the ordering of the layer widths, and it reduces the top-dimensional components to the solutions of a quadratic integer program that can be solved by closest-vector rounding.
What would settle it
Compute the top-dimensional components of the zero-product variety $\Sigma^0_{(2,3,5,7)}$ by direct elimination, and check each component's Kostant partition for a lace diagram with exactly one missing step in each of the top two rows and all horizontal intervals present below; a single component without such a diagram would show the quadratic integer program does not give the true codimension.
Extended reading notes
Core claim
The central discovery is that the zero-product variety $\Sigma^0_d = \{ (A_1,\ldots,A_N) : A_N \cdots A_1 = 0 \}$ has codimension $C$ equal to the minimum of a quadratic integer program in the layer widths, and the number $\theta$ of its top-dimensional irreducible components equals the number of minimizers. The same two numbers for the fiber $\mathrm{mult}^{-1}(B)$ over a rank-$r$ matrix $B$ follow by reduction (Lemmas 4.5 and 4.6). The paper proves the statement three ways: Theorem 5.5 expresses the relevant generating series as $Q^r_d = P_r \sum_{s=0}^{\min d - r} (-1)^s q^{\binom{s}{2}} P_s P_{d-r-s}$, whose lowest term is $\theta q^C$; Theorem 6.1 derives the quadratic program from the combinatorics of lace diagrams; Theorem 7.10 solves the program explicitly using the weakly increasing rearrangement of $d$, the closest-vector problem in a type A root lattice, and the quantity $S = \sum_{i=0}^m d'_i$. Corollary 5.10 records that $C$ and $\theta$ are permutation-invariant, and Theorem 8.6 identifies the real log-canonical threshold of $K_B^{\mathrm{DLN}}(A) = \|\mathrm{mult}(A)-B\|_2^2$ with half the codimension of $\mathrm{mult}^{-1}(B)$.
Load-bearing premise
The explicit computation rests on the assertion, argued by analogy and a sketched induction, that for weakly increasing widths every largest zero-product component can be drawn as a lace diagram with exactly one missing adjacent interval in each of the top rows and all horizontal intervals present in the rows below; if that assertion failed, the quadratic program and closed formulas would not describe the true $C$ and $\theta$.
Editorial extensions
If this is right
- For any fixed layer widths, the codimension and the number of top components of the zero-product variety, and of every fiber $\mathrm{mult}^{-1}(B)$, can be read off from a quadratic integer program or a closed formula instead of from resolving the variety.
- Permuting the layer widths leaves $C$ and $\theta$ unchanged: a network with widths $(d_0,d_1,d_2)$ and its mirror image have the same zero-product codimension and component count, though not the same ambient dimension or total component count.
- The real log-canonical threshold of the squared-Frobenius loss equals $C/2$, saturating the general inequality $\mathrm{rlct}(F) \le \mathrm{codim}\,F^{-1}(0)/2$; for Bayesian inference on deep linear networks this makes the leading asymptotic term explicitly computable from the widths.
- For equal widths $d$ and depth $N$, the explicit formula gives codimension $d(d+1)/2$ whenever $N>d$, independent of depth, and for fixed $N$ the codimension grows like $\frac{N+1}{2N} d^2$ as $d \to \infty$, showing that the zero-product locus is controlled mainly by width.
Reading between the lines
- Beyond the paper: the permutation invariance suggests there should be a bijective, purely combinatorial involution on Kostant partitions that preserves codimension and the number of minimal orbits; finding one could give an elementary proof and possibly extend the formulas to other quivers.
- Beyond the paper: because the global threshold equals half the fiber codimension, the formulas give a practical route to the local real log-canonical threshold for deep linear networks if the local equality $\mathrm{rlct}_A(K_B)=\mathrm{codim}_A\,\mathrm{mult}^{-1}(B)/2$, which the paper leaves to future work, holds.
- Beyond the paper: the closest-vector reformulation ties the counting problem for $\theta$ to the geometry of type A root lattices, so $\theta$ can be read as the number of lattice points at a fixed minimal distance; this may connect component counting to the complexity of integer programming at large depth.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies the variety of tuples of composable matrices whose product has a fixed rank r, and in particular the zero-product locus Σ^0_d and the fibers mult^{-1}(B) of the multiplication map. The main results are three descriptions of the codimension C of the top-dimensional components and their number θ: a Poincaré-series formula in equivariant cohomology (Theorem 5.5), a quadratic integer program (Theorem 6.1), and an explicit closed formula (Theorem 7.10). The paper also proves that C and θ are invariant under permutations of the dimension vector (Corollary 5.10) and reformulates Aoyagi's computation of the real log-canonical threshold of the squared Frobenius loss of deep linear networks as rlct(K_B^DLN) = codim mult^{-1}(B)/2 (Theorem 8.6). The exposition is generally clear, and the quiver-representation framework is used effectively to organize the orbit structure.
Significance. If fully established, the results solve a natural problem about the geometry of matrix multiplication and give a clean, explicit description of the singularities relevant to deep linear networks. The Poincaré-series method in Section 5 is elegant, and the paper provides several worked examples that make the formulas concrete. The main obstruction is that the proof of Theorem 6.1 rests on a structural classification of lace diagrams that is only sketched; the explicit formulas in Theorem 7.10 inherit this gap. Because the missing step is localized and appears fillable, the paper is promising but needs a completed proof before it can be accepted.
major comments (3)
- [Section 6.2, proof of Theorem 6.1] The classification of the lower rows in the lace diagram is asserted by the sentence 'An analoguous argument shows that in the rows below the top d'_0 rows all possible horizontal intervals are laces of L' and is not proved. This assertion is load-bearing: the displayed decomposition M = e_1(I_{00}+I_{1N}) + ... + e_N(I_{0,N-1}+I_{NN}) + f_1(I_{1N}) + ... + f_N(I_{NN}), the quadratic integer program (QIP), and hence Theorems 7.5 and 7.10 all depend on it. Please provide a proof that a lower row with a missing [i-1,i] interval can be extended to a strictly larger orbit closure that still lies in Σ^0_{d'}, for instance using the rank-pattern order of Theorem 3.8, or supply a reference for this structural fact.
- [Lemma 6.4] The proof of Lemma 6.4 is given only as 'Induction on N'. This lemma is used to represent every orbit in the weakly increasing case by a lace diagram with only horizontal laces, and it is not immediate from the definition of Kostant partitions. Please give the full induction argument or an explicit reference; as written, this is a gap in the proof of Theorem 6.1.
- [Theorem 7.5] The step in which the projection of s to the hyperplane has 'possibly a few last components negative' and the problem is reduced to the first m coordinates is not justified. This truncation step is the bridge from the quadratic integer program to the explicit formula, and it needs a proof, for example a water-filling or convexity argument showing that the optimal e_i vanish for i > m and that the remaining problem is exactly the closest-vector problem in the smaller simplex.
minor comments (6)
- [Lemma 5.9] The product (1-q^{s-r+1})...(1-q^r) should read (1-q^{s-r+1})...(1-q^s); the same typo appears in the displayed computation after Eq. (12). The subsequent use in Theorem 5.5 is with r = s, so the conclusion is unaffected, but the statement as written is false for r < s.
- [Proof of Theorem 7.5] In the displayed chain of equalities, the middle term is missing a factor of 2 in front of the first sum; the final identity is correct, but the intermediate display should be corrected.
- [Lemma 7.1] The last sentence says 'since d' is weakly decreasing', but d' is weakly increasing; this should be corrected.
- [Throughout] The word 'analoguous' appears in Section 6.2; it should be 'analogous'.
- [Remark 8.7] Remark 8.7 refers to 'the formula of Theorem' without a number; the intended theorem number should be supplied.
- [Appendix A] The proof of Theorem A.1 is only a sketch. Since this appendix is presented as a bonus and is not used in the main theorems, this is acceptable, but a precise reference to the details in [FR04] and [FR02] would be helpful.
Circularity Check
No substantive circularity: C and theta are derived from quiver geometry; self-citations are non-load-bearing and the main risk is an omitted proof, not circularity.
full rationale
The derivation chain for the main results (Theorems 5.5, 6.1, 7.10) is structurally self-contained: orbits of Rep_d are classified by Kostant partitions via Gabriel's theorem; codimensions are computed from dim Ext(M,M) (Lemma 3.4 and Corollary 3.5); the Poincaré series identity (Theorem 5.7) is proved from a spectral sequence; and the QIP is reduced to a closest-lattice-point problem in Section 7. None of these steps defines C or theta in terms of themselves or fits any quantity to the target data. The reduction to weakly increasing dimension vectors in Theorem 6.1 uses Corollary 5.10, which is established earlier from Theorem 5.5, so it is not circular. The structural claim about lace diagrams in Section 6.2 (Lemma 6.4 and the 'analoguous argument') is abbreviated: Lemma 6.4 is proved only as 'Induction on N' and the lower-row structure is asserted rather than demonstrated. I flag this as an omitted proof and a correctness risk, but it is not a circular reduction: the claim is an independent combinatorial statement about lace diagrams, not an input to the QIP. Theorem 8.6 is obtained by comparing the independently computed codimension formula (20) with Aoyagi's independently computed rlct formula [Aoy24, Theorem 1]; this is a translation/identification of two external computations, not the renaming of one result as another. The self-citations [FR02], [Rim13], and [RWY18] are used for standard lemmas (Ext dimension, stabilizer homotopy type, spectral sequence degeneration) that do not assume the paper's target results and are externally checkable; they therefore do not raise the circularity score. Overall the paper exhibits no circular reasoning; the score of 2 reflects only the presence of minor, non-load-bearing self-citations.
Assumptions & free parameters
assumptions (5)
- standard math Gabriel's theorem for equioriented type A quivers classifies indecomposable representations (Theorem 2.5)
- standard math Dimension formula for Ext between indecomposable representations: dim Ext(M_ij,M_uv)=1 iff i+1<=u<=j+1<=v (Equation 5)
- standard math Degeneration of the spectral sequence in equivariant cohomology for the orbit stratification (Lemma 5.8)
- domain assumption Properties of the real log-canonical threshold, including the bound rlct(F) <= codim(F^{-1}(0))/2 (Proposition 8.4)
- domain assumption Aoyagi's Theorem 1 computing the RLCT of the deep linear network loss (cited as [Aoy24])
Cite this review
Pith. "Pith review of Geometry of fibers of the multiplication map of deep linear neural networks." pith.science (2026). https://pith.science/paper/ZKLTSVNY
@misc{pith2026241119920,
author = {Pith},
title = {Pith review of: Geometry of fibers of the multiplication map of deep linear neural networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZKLTSVNY}},
note = {Machine review of arXiv:2411.19920}
}
abstract
We study the geometry of the algebraic set of tuples of composable matrices which multiply to a fixed matrix, using tools from the theory of quiver representations. In particular, we determine its codimension $C$ and the number $\theta$ of its top-dimensional irreducible components. Our solution is presented in three forms: a Poincar\'e series in equivariant cohomology, a quadratic integer program, and an explicit formula. In the course of the proof, we establish a surprising property: $C$ and $\theta$ are invariant under arbitrary permutations of the dimension vector. We also show that the real log-canonical threshold of the function taking a tuple to the square Frobenius norm of its product is $C/2$. These results are motivated by the study of deep linear neural networks in machine learning and Bayesian statistics (singular learning theory) and show that deep linear networks are in a certain sense ``mildly singular".
Figures
Forward citations
Cited by 1 Pith paper
-
Degenerations of multisingularities and Artin algebras
A singularity-theoretic 'stable hierarchy' on Artin algebras is shown to be computable from automorphism-group data and to extend the classical degeneration order beyond fixed rank.
Reference graph
Works this paper leans on
-
[1]
G. E. Andrews, R. Askey, and R. Roy. Special Functions . Number 71 in Enc. of Math. and Appl. CUP, 1999
work page 1999
-
[2]
M. Atiyah and R. Bott. The Y ang- M ills equation over R iemann surfaces. Phil. Trans. of the Royal Soc. London , 308(1505):523--615, 1983
work page 1983
-
[3]
A convergence analysis of gradient descent for deep linear neural networks, 2019
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu. A convergence analysis of gradient descent for deep linear neural networks, 2019
work page 2019
-
[4]
S. Abeasis and A. Del Fra . Degenerations for the representations of an equioriented quiver of type \(A_m\) . Boll. Unione Mat. Ital., Suppl. , 2:157--171, 1980
work page 1980
-
[5]
S. Abeasis and A. Del Fra . Degenerations for the representations of an equioriented quiver of type D^m . Adv. Math. , 52(2):81--172, 1984
work page 1984
-
[6]
S. Abeasis and A. Del Fra . Degenerations for the representations of a quiver of type \( A _ m\) . Journal of Algebra , 93:376--412, 1985
work page 1985
-
[7]
S. Abeasis, A. Del Fra , and H. Kraft. The geometry of representations of \(A_m\) . Math. Ann. , 256:401--418, 1981
work page 1981
-
[8]
E. Mehdi Achour, F. Malgouyres, and S. Gerchinovitz. The loss landscape of deep linear neural networks: a second-order analysis, 2024
work page 2024
Show all 48 references
-
[9]
M. Aoyagi. Consideration on the learning efficiency of multiple-layered neural networks with linear units. Neural Networks , 172(106132):1--11, 2024
2024
-
[10]
M. Atiyah. Resolution of singularities and division of distributions. Communications on pure and applied mathematics , 23(2):145--150, 1970
1970
-
[11]
A. S. Buch, L. Feh\'er, and R. Rim\'anyi. Positivity of quiver coefficients through T hom polynomials. Adv. Math. , 197:306--320, 2005
2005
-
[12]
Z. Chen, E. Lau, J. Mendel, S. Wei, and D. Murfet. Dynamical versus bayesian phase transitions in a toy model of superposition, 2023
2023
-
[13]
De Concini and E
C. De Concini and E. Strickland. On the variety of complexes. Advances in Mathematics , 41(1):57--77, 1981
1981
-
[14]
Conway and N
J. Conway and N. Sloane. Sphere Packings, Lattices and Groups , volume 290 of GL . Springer, 1988
1988
-
[15]
C. C. J. Dominé, N. Anguita, A. M. Proca, L. Braun, D. Kunin, P. A. M. Mediano, and A. M. Saxe. From lazy to rich: Exact learning dynamics in deep linear networks, 2024
2024
-
[16]
Feh\'er and R
L. Feh\'er and R. Rim\'anyi. Classes of degeneracy loci for quivers---the T hom polynomial point of view. Duke Math. J. , 114(2):193--213, 2002
2002
-
[17]
L. M. Feh\'er and R. Rim\'anyi. Calculation of T hom polynomials and other cohomological obstructions for group actions. In T. Gaffney and M. Ruas, editors, Real and Complex Singularities (Sao Carlos, 2002) , number 354 in Contemp. Math., pages 69--93. AMS, 2004
2002
-
[18]
Hoogland, G
J. Hoogland, G. Wang, M. Farrugia-Roberts, L. Carroll, S. Wei, and D. Murfet. The developmental landscape of in-context learning, 2024
2024
-
[19]
Jacot, F
A. Jacot, F. Ged, B. S im s ek, C. Hongler, and F. Gabriel. Saddle-to-saddle dynamics in deep linear networks: Small initialization training, symmetry, and sparsity, 2022
2022
-
[20]
Ji and M
Z. Ji and M. Telgarsky. Gradient descent aligns the layers of deep linear networks, 2019
2019
-
[21]
Kawaguchi
K. Kawaguchi. Deep learning without poor local minima. Advances in neural information processing systems , 29, 2016
2016
-
[22]
M. E. Kazarian. Characteristic classes of singularity theory. In The Arnold-Gelfand mathematical seminars , pages 325--340. Birkhauser, 1997
1997
-
[23]
Kirillov
A. Kirillov. Quiver representations and quiver varieties . Number 174 in GSM. AMS, 2016
2016
-
[24]
Knutson, E
A. Knutson, E. Miller, and M. Shimozono. Four positive formulae for type a quiver polynomials. Inv. Math. , 166(2):229--325, 2006
2006
-
[25]
Koll \'a r
J. Koll \'a r. Singularities of pairs. In Proceedings of Symposia in Pure Mathematics , volume 62, pages 221--288. American Mathematical Society, 1997
1997
-
[26]
Kinser and J
R. Kinser and J. Rajschot. Type a quiver loci and schubert varieties. Journal of Commutative Algebra , 7(2):265--301, 2015
2015
-
[27]
Koncki and R
J. Koncki and R. Rim\'anyi. The main reasons for matrices multiplying to zero. In preparation, 2024
2024
-
[28]
S. P. Lehalleur. Real jet schemes, real contact loci and the real log-canonical threshold. to appear, 2024
2024
-
[29]
E. Lau, Z. Furman, G. Wang, D. Murfet, and S. Wei. The local learning coefficient: A singularity-aware complexity measure, 2024
2024
-
[30]
Sh. Lin. Algebraic methods for evaluating integrals in Bayesian statistics . University of California, Berkeley, 2011
2011
-
[31]
Lu and K
H. Lu and K. Kawaguchi. Depth creates no bad local minima, 2017
2017
-
[32]
Lakshmibai and Peter Magyar
V. Lakshmibai and Peter Magyar. Degeneracy schemes, quiver schemes, and schubert varieties. International Mathematics Research Notices , 1998(12):627--640, 01 1998
1998
-
[33]
Marion and Ch
P. Marion and Ch. Lénaïc. Deep linear networks for regression are implicitly regularized towards flat minima, 2024
2024
-
[34]
Musili and C
Ch. Musili and C. S. Seshadri. Schubert varieties and the variety of complexes. In Arithmetic and Geometry: Papers Dedicated to IR Shafarevich on the Occasion of His Sixtieth Birthday. Volume II: Geometry , pages 329--359. Springer, 1983
1983
-
[35]
M. Mustata. Impanga lecture notes on log canonical thresholds, 2011
2011
-
[36]
M. Reineke. Poisson automorphisms and quiver moduli. Journal of the Institute of Mathematics of Jussieu , 9(3):653–667, 2010
2010
-
[37]
R. Rimanyi. On the cohomological hall algebra of dynkin quivers, 2013
2013
-
[38]
Rim\'anyi, A
R. Rim\'anyi, A. Weigandt, and A. Yong. Partition identities and quiver representations. J. of Alg. Comb. , 47:129--169, 2018
2018
-
[39]
J. R. Shewchuk and S. Bhattacharya. The geometry of the set of equivalent linear neural networks, 2024. arXiv:2404.14855
2024 arXiv
-
[40]
A. M. Saxe, J. L. McClelland, and S. Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks, 2014
2014
-
[41]
Trager, K
M. Trager, K. Kohn, and J. Bruna. Pure and spurious critical points: a geometric study of linear networks, 2020
2020
-
[42]
D. Voigt. Endliche algebraische Gruppen . Number 592 in Lecture Notes in Math. Springer, 1977
1977
-
[43]
Watanabe
S. Watanabe. Algebraic Geometry and Statistical Learning Theory . CUP, 2009
2009
-
[44]
Watanabe
S. Watanabe. Mathematical theory of B ayesian statistics . Chapman& Hall, 2018
2018
-
[45]
Recent advances in algebraic geometry and bayesian statistics
Sumio Watanabe. Recent advances in algebraic geometry and bayesian statistics. Information Geometry , 7(Suppl 1):187--209, 2024
2024
-
[46]
G. Wang, J. Hoogland, S. van Wingerden, Z. Furman, and D. Murfet. Differentiation and specialization of attention heads via the refined local learning coefficient, 2024
2024
-
[47]
A. V. Zelevinskii. Two remarks on graded nilpotent classes. Russian Mathematical Surveys , 40(1):249, 1985
1985
-
[48]
Ziyin, B
L. Ziyin, B. Li, and X. Meng. Exact solutions of a deep linear network. arXiv, 2022
2022
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.