REVIEW 2 major objections 7 minor 28 references
The Maximum Likelihood Degree of Toric Models is Monotonic
T0 review · 2 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The ML degree of a toric model cannot increase on any facial submodel.
desk verdict Clean proof of a genuinely open conjecture; the main theorem is correct, with one terse generic-avoidance step that needs a slightly more careful exposition but no real gap. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the scaled toric variety $X_{A,c}$, defined as the closure of the monomial parametrization $\theta \mapsto (c_i \theta_0 \theta^{a_i})$, together with its facial submodels $X_{A_F,c_F}$ obtained by restricting to lattice points of a face $F$ of $\operatorname{conv}(A)$. The ML degree counts complex solutions of the likelihood equations for generic data. The proof mechanism is the parameter continuation theorem, which says that specializing parameters in a polynomial system cannot increase the number of isolated solutions, combined with a block-triangular Jacobian computation that certifies the lifted facial solutions are isolated in the deformed full system.
What would settle it
Compute both ML degrees for a specific scaled toric variety and one of its faces—for instance, the all-ones three-dimensional cube (ML degree 8) and a facet (ML degree 4). Finding any face $F$ with $\operatorname{MLdeg}(X_{A_F,c_F}) > \operatorname{MLdeg}(X_{A,c})$, or generic facial data whose lifted point has singular Jacobian, would refute Theorem 6.
Extended reading notes
Core claim
The central claim is Theorem 6: for a scaled toric variety $X_{A,c}$ and any face $F$ of $\operatorname{conv}(A)$, the facial submodel $X_{A_F,c_F}$ satisfies $\operatorname{MLdeg}(X_{A_F,c_F}) \leq \operatorname{MLdeg}(X_{A,c})$. The proof works by induction on dimension, reduces to the case where $F$ is a facet, and studies the likelihood equations after extending facial data by zeros. Each generic facial critical point lifts to a solution $\hat{\theta} = (\hat{\theta}_F, 0)$ of the full system, and this point is isolated because the Jacobian is block triangular with lower-right entry $\theta_0 \partial_{\theta_d}(\theta_d g)|_{\theta_d=0}$, which is nonzero for generic facial data by a dimension count. The parameter continuation theorem then implies the number of isolated solutions can only decrease under data specialization, giving the inequality.
Load-bearing premise
The proof assumes that a certain nonzero polynomial built from the face direction does not happen to vanish at all the solution points of a generic smaller problem, because the face has fewer coordinates than the whole model; if it did vanish, the key step showing the lifted solutions stay isolated would fail.
Editorial extensions
If this is right
- For an undirected graphical model, the ML degree of any induced-subgraph model is at most the ML degree of the original graph model (Corollary 13).
- For quasi-independence models, restricting to an induced subgraph of the associated bipartite graph can only lower or preserve the ML degree (Corollary 16).
- Whenever the data linear space meets the scaled toric variety transversally, the number of complex likelihood solutions is at most the ML degree, even for non-generic data with zeros (Corollary 9).
- Monotonicity fails for arbitrary non-facial submodels: deleting one column of the design matrix can raise the ML degree from 1 to 3 (Example 7).
- The tropical likelihood degeneration of Section 5 lets one watch which solutions survive the face limit, and a tropical basis would turn this into a precise combinatorial refinement, which the paper leaves as Problem 12.
Reading between the lines
- If the block-triangular Jacobian structure is the only ingredient needed, the same monotonicity should hold for other families of very affine varieties with facial submodels; testing it on Gaussian graphical models or other non-toric likelihood varieties would be a direct next step.
- Example 8's zero-data computations suggest that the number of isolated solutions is governed by which zeros fall inside or outside the face support; a data-plus-scaling discriminant would give a complete answer to when the count drops, stays, or becomes infinite.
- A tropical basis for the Puiseux likelihood equations would refine the inequality into an explicit bookkeeping of which critical points belong to each face, potentially yielding a combinatorial formula for the ML-degree drop along flags of faces.
- Because the theorem holds for arbitrary scalings, it constrains the Euler stratification of the parameter space: the ML-degree strata over any face cannot exceed the corresponding stratum of the full polytope.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proves that the maximum likelihood (ML) degree of a toric model is monotone with respect to the face poset of the defining polytope: for any scaled toric variety X_{A,c} and any face F of conv(A), the ML degree of the facial submodel X_{A_F,c_F} is at most that of the full model. This settles a conjecture of Coons and Sullivant. The proof uses a parameter continuation argument, extending generic data to data with zeros on the complement of the face, and shows that the lifted critical points of the facial submodel are isolated solutions of the full model's likelihood system. The paper also discusses implications for data zeros, connects the result to tropical likelihood degenerations, and applies the main theorem to discrete graphical models and quasi-independence models.
Significance. If the main theorem is correct, it resolves a natural and previously open conjecture in algebraic statistics, providing a clean structural statement about ML degrees under taking facial submodels. The proof strategy, based on Morgan–Sommese parameter continuation and a block-triangular Jacobian computation, is potentially reusable beyond the specific setting. The applications to graphical and quasi-independence models are immediate and strengthen results that previously required the full model to have ML degree one. The paper also contains reproducible computational experiments (e.g., Section 4, Table 2) and proposes a tropical degeneration framework that, while not fully rigorous, points to a promising direction for refining the monotonicity statement. The central claim is falsifiable and the proof, after a repair described below, is sound.
major comments (2)
- [Section 3, proof of Theorem 6] The generic-avoidance step is not justified as written. The sentence 'Its vanishing set has dimension d − 1 while dim(X_{A_F,c_F}) = d' contains two errors: the vanishing set of a nonzero polynomial in d−1 variables has dimension at most d−2 (or d−1 if one works in the full (θ_0,...,θ_{d-1})-space), and the facial variety has dimension d−1, not d. Moreover, a dimension comparison alone cannot prove the avoidance claim, since a hypersurface can contain a variety of the same dimension. The needed argument is to consider the incidence variety I = {(u_F, θ) : L_{A_F}(u_F; θ) = 0}, show it is irreducible and projects dominantly onto the θ-torus, and then note that the nonzero Laurent polynomial h = ∂_{θ_d}(θ_d g)|_{θ_d=0} cannot vanish identically on I; hence the projection of I ∩ Z(h) to the data space is a proper closed subset. This yields a Zariski-open set of u_F for which no critical point lies in Z(h). The authors should incorporate this argument to make the proof complete.
- [Section 3, proof of Theorem 6] After defining α, the assertion that every solution θ̂_F is isolated in V(L2) requires that the lower-right Jacobian entry θ0 ∂_{θ_d}(θ_d g)|_{θ_d=0} is nonzero at θ̂_F. This is exactly the avoidance condition discussed in the previous comment. The paper's one-line justification is insufficient; the proof should explicitly state that the chosen generic u_F simultaneously avoids the finitely many critical points of the facial model and the hypersurface Z(h), and that this is possible because both conditions hold on a Zariski-open set of data.
minor comments (7)
- [Section 3, proof of Theorem 6] The notation ∂θdθdg is terse and ambiguous; it should be written as ∂_{θ_d}(θ_d g)|_{θ_d=0} throughout.
- [Section 3, proof of Theorem 6] The inequality α ≥ 0 holds because every lattice point outside the facet F has positive last coordinate after the unimodular transformation, but this justification is omitted.
- [Section 4, Example 8] Table 1 displays scalings as 3×3 matrices while the design matrix A in the same example is 5×9; the authors should clarify that the matrices are flattened into length-9 scaling vectors.
- [Section 4, Table 2] The entries '∞' in Table 2 are not defined in the caption; state explicitly that they denote positive-dimensional solution components of the likelihood system.
- [Section 5, Equation (5)] The notation δ_F(a) with a ∈ A is ambiguous because A is a matrix; use lattice points a_i ∈ Z^d instead.
- [Section 5, Example 11] The phrase 'This induces a regular subdivision of the Cayley polytope' would benefit from a brief definition or reference for the Cayley configuration, since it is central to the claimed failure of the sufficient criterion.
- [Section 6, Corollary 13] The proof relies on [GMS06, Lemma A.2] for the face property; stating the content of that lemma would make the proof more self-contained.
Circularity Check
No significant circularity: the monotonicity theorem is proved from parameter continuation and standard algebraic geometry; self-citations are contextual only.
full rationale
The central theorem is proved by a parameter homotopy from generic data to data with zeros supported on the facet. The facial ML degree is not an input to the full model's computation; instead, generic facial critical points are shown to lift to isolated critical points of the deformed full system, and Morgan–Sommese continuation bounds the number of isolated solutions by the generic count. Birch's theorem, the GKZ normalization facts, and the likelihood-equation formulation from [ABB+19, Definition 6] are standard external and verifiable inputs. Self-citations such as [ABB+19], [AO24], [ADMV24], and [TW24] appear in examples and contextual remarks and do not carry the proof of Theorem 6. The generic-avoidance claim about the polynomial h is a Zariski-openness statement rather than a fitted or renamed prediction; it can be justified independently through an incidence-variety argument, so the terse dimension wording in the proof is at most a rigor or exposition issue, not a circular one. No step reduces by definition to its own input, and no fitted parameter is renamed as a prediction.
Assumptions & free parameters
assumptions (4)
- standard math Parameter Continuation Theorem (Morgan and Sommese 1989): the number of isolated solutions of a square polynomial system is finite and upper semicontinuous in parameters.
- standard math Birch's Theorem for log-affine models, giving the likelihood equations as A p = A u / u_+ and the MLE as the unique nonnegative solution.
- standard math Affine unimodular transformations preserve the ML degree of toric varieties, and any facet can be moved to a coordinate hyperplane with the polytope in the nonnegative orthant.
- domain assumption For generic data the critical points of the likelihood equations are nonsingular isolated points with nonzero coordinates, away from the hyperplane arrangement H.
Cite this review
Pith. "Pith review of The Maximum Likelihood Degree of Toric Models is Monotonic." pith.science (2026). https://pith.science/paper/3CY4FYUP
@misc{pith2026250702719,
author = {Pith},
title = {Pith review of: The Maximum Likelihood Degree of Toric Models is Monotonic},
year = {2026},
howpublished = {\url{https://pith.science/paper/3CY4FYUP}},
note = {Machine review of arXiv:2507.02719}
}
read the original abstract
We settle a conjecture by Coons and Sullivant stating that the maximum likelihood (ML) degree of a facial submodel of a toric model is at most the ML degree of the model itself. We discuss the impact on the ML degree from observing zeros in the data. Moreover, we connect this problem to tropical likelihood degenerations, and show how the results can be applied to discrete graphical and quasi-independence models.
Figures
Reference graph
Works this paper leans on
-
[1]
The maximum likelihood degree of toric varieties
Carlos Am \'e ndola, Nathan Bliss, Isaac Burke, Courtney R.\ Gibbons, Martin Helmer, Serkan Ho s ten, Evan D.\ Nash, Jose I.\ Rodriguez, and Daniel Smolkin. The maximum likelihood degree of toric varieties. Journal of Symbolic Computation , 92:222--242, 2019
work page 2019
-
[2]
Daniele Agostini, Taylor Brysiewicz, Claudia Fevola, Lukas K \"u hne, Bernd Sturmfels, Simon Telen, and Thomas Lam. Likelihood degenerations. Advances in Mathematics , 414:108863, 2023
work page 2023
-
[3]
On the maximum likelihood degree for Gaussian graphical models
Carlos Am \'e ndola, Rodica Andreea Dinu, Mateusz Micha ek, and Martin Vodi c ka. On the maximum likelihood degree for G aussian graphical models. arXiv:2410.07007 https://arxiv.org/abs/2410.07007 , 2024
work page Pith review arXiv 2024
-
[4]
Likelihood geometry of reflexive polytopes
Carlos Am \'e ndola and Janike Oldekop. Likelihood geometry of reflexive polytopes. Algebraic Statistics , 15(1):113--143, 2024
work page 2024
-
[5]
Bates, Paul Breiding, Tianran Chen, Jonathan D
Daniel J. Bates, Paul Breiding, Tianran Chen, Jonathan D. Hauenstein, Anton Leykin, and Frank Sottile. Numerical nonlinear algebra. arXiv:2302.08585 https://arxiv.org/abs/2302.08585 , 2024
arXiv 2024
-
[6]
Tropical toric maximum likelihood estimation
Eric Boniface, Karel Devriendt, and Serkan Ho s ten. Tropical toric maximum likelihood estimation. arXiv:2404.10567 https://arxiv.org/abs/2404.10567 , 2024
work page Pith review arXiv 2024
-
[7]
Matroid stratification of ML degrees of independence models
Oliver Clarke, Serkan Hoşten, Nataliia Kushnerchuk, and Janike Oldekop. Matroid stratification of ML degrees of independence models. Algebraic Statistics , 15(2):199--223, 2024
work page 2024
-
[8]
Fabrizio Catanese, Serkan Ho s ten, Amit Khetan, and Bernd Sturmfels. The maximum likelihood degree. American Journal of Mathematics , 128(3):671--697, 2006
work page 2006
Show all 28 references
-
[9]
Quasi-independence models with rational maximum likelihood estimator
Jane I.\ Coons and Seth Sullivant. Quasi-independence models with rational maximum likelihood estimator. Journal of Symbolic Computation , 104:917--941, 2021
2021
-
[10]
Lectures on Algebraic Statistics , volume 39 of Oberwolfach Seminars
Mathias Drton, Bernd Sturmfels, and Seth Sullivant. Lectures on Algebraic Statistics , volume 39 of Oberwolfach Seminars . Springer, 2009
2009
-
[11]
Engineered complete intersections: slightly degenerate B ernstein-- K ouchnirenko-- K hovanskii
Alexander Esterov. Engineered complete intersections: slightly degenerate B ernstein-- K ouchnirenko-- K hovanskii. arXiv:2401.12099 https://arxiv.org/abs/2401.12099 , 2024
2024 arXiv
-
[12]
Zelevinsky
Izrail M.\ Gel'fand, Mikhail M.\ Kapranov, and Andrei V. Zelevinsky. Discriminants, resultants, and multidimensional determinants . Mathematics: Theory & Applications, Birkh\"auser Boston, Inc., Boston, MA, 1994
1994
-
[13]
On the toric algebra of graphical models
Dan Geiger, Christopher Meek, and Bernd Sturmfels. On the toric algebra of graphical models. The Annals of Statistics , 11(3):1463--1492, 2006
2006
-
[14]
Maximum likelihood geometry in the presence of data zeros
Elizabeth Gross and Jose Rodriguez. Maximum likelihood geometry in the presence of data zeros. Proceedings of the International Symposium on Symbolic and Algebraic Computation, ISSAC , 2013
2013
-
[15]
Performing the exact test of H ardy-- W einberg proportion for multiple alleles
Sun Wei Guo and Elizabeth A.\ Thompson. Performing the exact test of H ardy-- W einberg proportion for multiple alleles. Biometrics , 48(2):361--372, 1992
1992
-
[16]
The maximum likelihood data singular locus
Emil Horobeţ and Jose I.\ Rodriguez. The maximum likelihood data singular locus. Journal of Symbolic Computation , 79:99--107, 2017
2017
-
[17]
Likelihood Geometry , pages 63--117
June Huh and Bernd Sturmfels. Likelihood Geometry , pages 63--117. Lecture Notes in Mathematics. Springer International Publishing, 2014
2014
-
[18]
Logarithmic discriminants of hyperplane arrangements
Leonie Kayser, Andreas Kretschmer, and Simon Telen. Logarithmic discriminants of hyperplane arrangements. Le Matematiche , 80(1):325--346, 2025
2025
-
[19]
Morgan and Andrew J
Alexander P. Morgan and Andrew J. Sommese. Coefficient-parameter polynomial continuation. Applied Mathematics and Computation , 29(2):123--160, 1989
1989
-
[20]
Introduction to Tropical Geometry , volume 161
Diane Maclagan and Bernd Sturmfels. Introduction to Tropical Geometry , volume 161. American Mathematical Society, 2015
2015
-
[21]
O SCAR -- O pen S ource C omputer A lgebra R esearch system, V ersion 1.4.1, 2025
2025
-
[22]
Likelihood equations and scattering amplitudes
Bernd Sturmfels and Simon Telen. Likelihood equations and scattering amplitudes. Algebraic Statistics , 12(2):167--186, 2021
2021
-
[23]
A monotonicity property of h -vectors and h^* -vectors
Richard P.\ Stanley. A monotonicity property of h -vectors and h^* -vectors. European Journal of Combinatorics , 14(3):251--258, 1993
1993
-
[24]
Algebraic Statistics
Seth Sullivant. Algebraic Statistics . Graduate Studies in Mathematics. American Mathematical Society, 2018
2018
-
[25]
Maximum likelihood estimation from a tropical and a bernstein--sato perspective
Anna-Laura Sattelberger and Robin van der Veer. Maximum likelihood estimation from a tropical and a bernstein--sato perspective. International Mathematics Research Notices , 2023(6):5263--5292, 2023
2023
-
[26]
Andrew J.\ Sommese and Charles W. Wampler. The Numerical Solution of Systems of Polynomials Arising in Engineering and Science . World Scientific, 2005
2005
-
[27]
Toric amplitudes and universal adjoints
Simon Telen. Toric amplitudes and universal adjoints. arXiv:2504.00897 https://arxiv.org/abs/2504.00897 , 2025
2025
-
[28]
Euler stratifications of hypersurface families
Simon Telen and Maximilian Wiesmann. Euler stratifications of hypersurface families. arXiv:2407.18176 https://arxiv.org/abs/2407.18176 , 2024
2024 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.