Pith. sign in

REVIEW 2 major objections 7 minor 28 references

The Maximum Likelihood Degree of Toric Models is Monotonic

T0 review · 2 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The ML degree of a toric model cannot increase on any facial submodel.

desk verdict Clean proof of a genuinely open conjecture; the main theorem is correct, with one terse generic-avoidance step that needs a slightly more careful exposition but no real gap. read the letter →

arxiv 2507.02719 v1 pith:3CY4FYUP submitted 2025-07-03 math.AG math.STstat.TH

classification math.AGmath.STstat.TH MSC 14M2562R01
keywords maximumlikelihooddegreetoricvarietiesfacialsubmodelsmonotonicityconjectureparametercontinuationequationsdiscretegraphicalmodelsquasi-independence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proves that the maximum likelihood degree—the number of complex critical points of the likelihood equations for generic data—never increases when a toric model is restricted to a face of its defining polytope. The result settles a conjecture from earlier work on quasi-independence models and covers every scaling of the toric variety, not only generic scalings. Because the face poset of the polytope organizes the model's submodels, the inequality gives a uniform bound: any facial submodel has likelihood-estimation complexity at most that of the full model. The proof deforms generic data to data with zeros outside the face and uses parameter continuation to show the full system retains at least as many isolated solutions as the facial system.

What carries the argument

The load-bearing object is the scaled toric variety $X_{A,c}$, defined as the closure of the monomial parametrization $\theta \mapsto (c_i \theta_0 \theta^{a_i})$, together with its facial submodels $X_{A_F,c_F}$ obtained by restricting to lattice points of a face $F$ of $\operatorname{conv}(A)$. The ML degree counts complex solutions of the likelihood equations for generic data. The proof mechanism is the parameter continuation theorem, which says that specializing parameters in a polynomial system cannot increase the number of isolated solutions, combined with a block-triangular Jacobian computation that certifies the lifted facial solutions are isolated in the deformed full system.

What would settle it

Compute both ML degrees for a specific scaled toric variety and one of its faces—for instance, the all-ones three-dimensional cube (ML degree 8) and a facet (ML degree 4). Finding any face $F$ with $\operatorname{MLdeg}(X_{A_F,c_F}) > \operatorname{MLdeg}(X_{A,c})$, or generic facial data whose lifted point has singular Jacobian, would refute Theorem 6.

Watch

Extended reading notes

Core claim

The central claim is Theorem 6: for a scaled toric variety $X_{A,c}$ and any face $F$ of $\operatorname{conv}(A)$, the facial submodel $X_{A_F,c_F}$ satisfies $\operatorname{MLdeg}(X_{A_F,c_F}) \leq \operatorname{MLdeg}(X_{A,c})$. The proof works by induction on dimension, reduces to the case where $F$ is a facet, and studies the likelihood equations after extending facial data by zeros. Each generic facial critical point lifts to a solution $\hat{\theta} = (\hat{\theta}_F, 0)$ of the full system, and this point is isolated because the Jacobian is block triangular with lower-right entry $\theta_0 \partial_{\theta_d}(\theta_d g)|_{\theta_d=0}$, which is nonzero for generic facial data by a dimension count. The parameter continuation theorem then implies the number of isolated solutions can only decrease under data specialization, giving the inequality.

Load-bearing premise

The proof assumes that a certain nonzero polynomial built from the face direction does not happen to vanish at all the solution points of a generic smaller problem, because the face has fewer coordinates than the whole model; if it did vanish, the key step showing the lifted solutions stay isolated would fail.

Editorial extensions

If this is right

  • For an undirected graphical model, the ML degree of any induced-subgraph model is at most the ML degree of the original graph model (Corollary 13).
  • For quasi-independence models, restricting to an induced subgraph of the associated bipartite graph can only lower or preserve the ML degree (Corollary 16).
  • Whenever the data linear space meets the scaled toric variety transversally, the number of complex likelihood solutions is at most the ML degree, even for non-generic data with zeros (Corollary 9).
  • Monotonicity fails for arbitrary non-facial submodels: deleting one column of the design matrix can raise the ML degree from 1 to 3 (Example 7).
  • The tropical likelihood degeneration of Section 5 lets one watch which solutions survive the face limit, and a tropical basis would turn this into a precise combinatorial refinement, which the paper leaves as Problem 12.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the block-triangular Jacobian structure is the only ingredient needed, the same monotonicity should hold for other families of very affine varieties with facial submodels; testing it on Gaussian graphical models or other non-toric likelihood varieties would be a direct next step.
  • Example 8's zero-data computations suggest that the number of isolated solutions is governed by which zeros fall inside or outside the face support; a data-plus-scaling discriminant would give a complete answer to when the count drops, stays, or becomes infinite.
  • A tropical basis for the Puiseux likelihood equations would refine the inequality into an explicit bookkeeping of which critical points belong to each face, potentially yielding a combinatorial formula for the ML-degree drop along flags of faces.
  • Because the theorem holds for arbitrary scalings, it constrains the Euler stratification of the parameter space: the ML-degree strata over any face cannot exceed the corresponding stratum of the full polytope.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 7 minor

Summary. The paper proves that the maximum likelihood (ML) degree of a toric model is monotone with respect to the face poset of the defining polytope: for any scaled toric variety X_{A,c} and any face F of conv(A), the ML degree of the facial submodel X_{A_F,c_F} is at most that of the full model. This settles a conjecture of Coons and Sullivant. The proof uses a parameter continuation argument, extending generic data to data with zeros on the complement of the face, and shows that the lifted critical points of the facial submodel are isolated solutions of the full model's likelihood system. The paper also discusses implications for data zeros, connects the result to tropical likelihood degenerations, and applies the main theorem to discrete graphical models and quasi-independence models.

Significance. If the main theorem is correct, it resolves a natural and previously open conjecture in algebraic statistics, providing a clean structural statement about ML degrees under taking facial submodels. The proof strategy, based on Morgan–Sommese parameter continuation and a block-triangular Jacobian computation, is potentially reusable beyond the specific setting. The applications to graphical and quasi-independence models are immediate and strengthen results that previously required the full model to have ML degree one. The paper also contains reproducible computational experiments (e.g., Section 4, Table 2) and proposes a tropical degeneration framework that, while not fully rigorous, points to a promising direction for refining the monotonicity statement. The central claim is falsifiable and the proof, after a repair described below, is sound.

major comments (2)
  1. [Section 3, proof of Theorem 6] The generic-avoidance step is not justified as written. The sentence 'Its vanishing set has dimension d − 1 while dim(X_{A_F,c_F}) = d' contains two errors: the vanishing set of a nonzero polynomial in d−1 variables has dimension at most d−2 (or d−1 if one works in the full (θ_0,...,θ_{d-1})-space), and the facial variety has dimension d−1, not d. Moreover, a dimension comparison alone cannot prove the avoidance claim, since a hypersurface can contain a variety of the same dimension. The needed argument is to consider the incidence variety I = {(u_F, θ) : L_{A_F}(u_F; θ) = 0}, show it is irreducible and projects dominantly onto the θ-torus, and then note that the nonzero Laurent polynomial h = ∂_{θ_d}(θ_d g)|_{θ_d=0} cannot vanish identically on I; hence the projection of I ∩ Z(h) to the data space is a proper closed subset. This yields a Zariski-open set of u_F for which no critical point lies in Z(h). The authors should incorporate this argument to make the proof complete.
  2. [Section 3, proof of Theorem 6] After defining α, the assertion that every solution θ̂_F is isolated in V(L2) requires that the lower-right Jacobian entry θ0 ∂_{θ_d}(θ_d g)|_{θ_d=0} is nonzero at θ̂_F. This is exactly the avoidance condition discussed in the previous comment. The paper's one-line justification is insufficient; the proof should explicitly state that the chosen generic u_F simultaneously avoids the finitely many critical points of the facial model and the hypersurface Z(h), and that this is possible because both conditions hold on a Zariski-open set of data.
minor comments (7)
  1. [Section 3, proof of Theorem 6] The notation ∂θdθdg is terse and ambiguous; it should be written as ∂_{θ_d}(θ_d g)|_{θ_d=0} throughout.
  2. [Section 3, proof of Theorem 6] The inequality α ≥ 0 holds because every lattice point outside the facet F has positive last coordinate after the unimodular transformation, but this justification is omitted.
  3. [Section 4, Example 8] Table 1 displays scalings as 3×3 matrices while the design matrix A in the same example is 5×9; the authors should clarify that the matrices are flattened into length-9 scaling vectors.
  4. [Section 4, Table 2] The entries '∞' in Table 2 are not defined in the caption; state explicitly that they denote positive-dimensional solution components of the likelihood system.
  5. [Section 5, Equation (5)] The notation δ_F(a) with a ∈ A is ambiguous because A is a matrix; use lattice points a_i ∈ Z^d instead.
  6. [Section 5, Example 11] The phrase 'This induces a regular subdivision of the Cayley polytope' would benefit from a brief definition or reference for the Cayley configuration, since it is central to the claimed failure of the sufficient criterion.
  7. [Section 6, Corollary 13] The proof relies on [GMS06, Lemma A.2] for the face property; stating the content of that lemma would make the proof more self-contained.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the monotonicity theorem is proved from parameter continuation and standard algebraic geometry; self-citations are contextual only.

full rationale

The central theorem is proved by a parameter homotopy from generic data to data with zeros supported on the facet. The facial ML degree is not an input to the full model's computation; instead, generic facial critical points are shown to lift to isolated critical points of the deformed full system, and Morgan–Sommese continuation bounds the number of isolated solutions by the generic count. Birch's theorem, the GKZ normalization facts, and the likelihood-equation formulation from [ABB+19, Definition 6] are standard external and verifiable inputs. Self-citations such as [ABB+19], [AO24], [ADMV24], and [TW24] appear in examples and contextual remarks and do not carry the proof of Theorem 6. The generic-avoidance claim about the polynomial h is a Zariski-openness statement rather than a fitted or renamed prediction; it can be justified independently through an incidence-variety argument, so the terse dimension wording in the proof is at most a rigor or exposition issue, not a circular one. No step reduces by definition to its own input, and no fitted parameter is renamed as a prediction.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The proof uses classical theorems from algebraic and numerical algebraic geometry, not fitted values. The only subtle premise is the generic-data avoidance of the hypersurface h=0, justified by a dimension count. No new entities or free parameters are introduced.

assumptions (4)
  • standard math Parameter Continuation Theorem (Morgan and Sommese 1989): the number of isolated solutions of a square polynomial system is finite and upper semicontinuous in parameters.
    Applied in the proof of Theorem 6 to conclude that the number of isolated solutions at specialized data with zeros is at most the ML degree.
  • standard math Birch's Theorem for log-affine models, giving the likelihood equations as A p = A u / u_+ and the MLE as the unique nonnegative solution.
    Defines the polynomial system L_A whose isolated solution count is the ML degree (Section 2).
  • standard math Affine unimodular transformations preserve the ML degree of toric varieties, and any facet can be moved to a coordinate hyperplane with the polytope in the nonnegative orthant.
    Used in the proof of Theorem 6 to put the facet F in the hyperplane {last coordinate = 0}.
  • domain assumption For generic data the critical points of the likelihood equations are nonsingular isolated points with nonzero coordinates, away from the hyperplane arrangement H.
    Needed so that the parameter continuation count matches the ML degree and the lifted points are nonsingular; cited to [GR13].

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Maximum Likelihood Degree of Toric Models is Monotonic." pith.science (2026). https://pith.science/paper/3CY4FYUP

@misc{pith2026250702719,
  author       = {Pith},
  title        = {Pith review of: The Maximum Likelihood Degree of Toric Models is Monotonic},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3CY4FYUP}},
  note         = {Machine review of arXiv:2507.02719}
}
read the original abstract

We settle a conjecture by Coons and Sullivant stating that the maximum likelihood (ML) degree of a facial submodel of a toric model is at most the ML degree of the model itself. We discuss the impact on the ML degree from observing zeros in the data. Moreover, we connect this problem to tropical likelihood degenerations, and show how the results can be applied to discrete graphical and quasi-independence models.

Figures

Figures reproduced from arXiv: 2507.02719 by the authors.

Figure 1
Figure 1. Examples of three-dimensional polytopes. In (a), the cube with side-length two is [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

28 extracted references · 25 canonical work pages

  1. [1]

    The maximum likelihood degree of toric varieties

    Carlos Am \'e ndola, Nathan Bliss, Isaac Burke, Courtney R.\ Gibbons, Martin Helmer, Serkan Ho s ten, Evan D.\ Nash, Jose I.\ Rodriguez, and Daniel Smolkin. The maximum likelihood degree of toric varieties. Journal of Symbolic Computation , 92:222--242, 2019

  2. [2]

    Likelihood degenerations

    Daniele Agostini, Taylor Brysiewicz, Claudia Fevola, Lukas K \"u hne, Bernd Sturmfels, Simon Telen, and Thomas Lam. Likelihood degenerations. Advances in Mathematics , 414:108863, 2023

  3. [3]

    On the maximum likelihood degree for Gaussian graphical models

    Carlos Am \'e ndola, Rodica Andreea Dinu, Mateusz Micha ek, and Martin Vodi c ka. On the maximum likelihood degree for G aussian graphical models. arXiv:2410.07007 https://arxiv.org/abs/2410.07007 , 2024

  4. [4]

    Likelihood geometry of reflexive polytopes

    Carlos Am \'e ndola and Janike Oldekop. Likelihood geometry of reflexive polytopes. Algebraic Statistics , 15(1):113--143, 2024

  5. [5]

    Bates, Paul Breiding, Tianran Chen, Jonathan D

    Daniel J. Bates, Paul Breiding, Tianran Chen, Jonathan D. Hauenstein, Anton Leykin, and Frank Sottile. Numerical nonlinear algebra. arXiv:2302.08585 https://arxiv.org/abs/2302.08585 , 2024

  6. [6]

    Tropical toric maximum likelihood estimation

    Eric Boniface, Karel Devriendt, and Serkan Ho s ten. Tropical toric maximum likelihood estimation. arXiv:2404.10567 https://arxiv.org/abs/2404.10567 , 2024

  7. [7]

    Matroid stratification of ML degrees of independence models

    Oliver Clarke, Serkan Hoşten, Nataliia Kushnerchuk, and Janike Oldekop. Matroid stratification of ML degrees of independence models. Algebraic Statistics , 15(2):199--223, 2024

  8. [8]

    The maximum likelihood degree

    Fabrizio Catanese, Serkan Ho s ten, Amit Khetan, and Bernd Sturmfels. The maximum likelihood degree. American Journal of Mathematics , 128(3):671--697, 2006

Show all 28 references
  1. [9]

    Quasi-independence models with rational maximum likelihood estimator

    Jane I.\ Coons and Seth Sullivant. Quasi-independence models with rational maximum likelihood estimator. Journal of Symbolic Computation , 104:917--941, 2021

  2. [10]

    Lectures on Algebraic Statistics , volume 39 of Oberwolfach Seminars

    Mathias Drton, Bernd Sturmfels, and Seth Sullivant. Lectures on Algebraic Statistics , volume 39 of Oberwolfach Seminars . Springer, 2009

  3. [11]

    Engineered complete intersections: slightly degenerate B ernstein-- K ouchnirenko-- K hovanskii

    Alexander Esterov. Engineered complete intersections: slightly degenerate B ernstein-- K ouchnirenko-- K hovanskii. arXiv:2401.12099 https://arxiv.org/abs/2401.12099 , 2024

  4. [12]

    Zelevinsky

    Izrail M.\ Gel'fand, Mikhail M.\ Kapranov, and Andrei V. Zelevinsky. Discriminants, resultants, and multidimensional determinants . Mathematics: Theory & Applications, Birkh\"auser Boston, Inc., Boston, MA, 1994

  5. [13]

    On the toric algebra of graphical models

    Dan Geiger, Christopher Meek, and Bernd Sturmfels. On the toric algebra of graphical models. The Annals of Statistics , 11(3):1463--1492, 2006

  6. [14]

    Maximum likelihood geometry in the presence of data zeros

    Elizabeth Gross and Jose Rodriguez. Maximum likelihood geometry in the presence of data zeros. Proceedings of the International Symposium on Symbolic and Algebraic Computation, ISSAC , 2013

  7. [15]

    Performing the exact test of H ardy-- W einberg proportion for multiple alleles

    Sun Wei Guo and Elizabeth A.\ Thompson. Performing the exact test of H ardy-- W einberg proportion for multiple alleles. Biometrics , 48(2):361--372, 1992

  8. [16]

    The maximum likelihood data singular locus

    Emil Horobeţ and Jose I.\ Rodriguez. The maximum likelihood data singular locus. Journal of Symbolic Computation , 79:99--107, 2017

  9. [17]

    Likelihood Geometry , pages 63--117

    June Huh and Bernd Sturmfels. Likelihood Geometry , pages 63--117. Lecture Notes in Mathematics. Springer International Publishing, 2014

  10. [18]

    Logarithmic discriminants of hyperplane arrangements

    Leonie Kayser, Andreas Kretschmer, and Simon Telen. Logarithmic discriminants of hyperplane arrangements. Le Matematiche , 80(1):325--346, 2025

  11. [19]

    Morgan and Andrew J

    Alexander P. Morgan and Andrew J. Sommese. Coefficient-parameter polynomial continuation. Applied Mathematics and Computation , 29(2):123--160, 1989

  12. [20]

    Introduction to Tropical Geometry , volume 161

    Diane Maclagan and Bernd Sturmfels. Introduction to Tropical Geometry , volume 161. American Mathematical Society, 2015

  13. [21]

    O SCAR -- O pen S ource C omputer A lgebra R esearch system, V ersion 1.4.1, 2025

  14. [22]

    Likelihood equations and scattering amplitudes

    Bernd Sturmfels and Simon Telen. Likelihood equations and scattering amplitudes. Algebraic Statistics , 12(2):167--186, 2021

  15. [23]

    A monotonicity property of h -vectors and h^* -vectors

    Richard P.\ Stanley. A monotonicity property of h -vectors and h^* -vectors. European Journal of Combinatorics , 14(3):251--258, 1993

  16. [24]

    Algebraic Statistics

    Seth Sullivant. Algebraic Statistics . Graduate Studies in Mathematics. American Mathematical Society, 2018

  17. [25]

    Maximum likelihood estimation from a tropical and a bernstein--sato perspective

    Anna-Laura Sattelberger and Robin van der Veer. Maximum likelihood estimation from a tropical and a bernstein--sato perspective. International Mathematics Research Notices , 2023(6):5263--5292, 2023

  18. [26]

    Andrew J.\ Sommese and Charles W. Wampler. The Numerical Solution of Systems of Polynomials Arising in Engineering and Science . World Scientific, 2005

  19. [27]

    Toric amplitudes and universal adjoints

    Simon Telen. Toric amplitudes and universal adjoints. arXiv:2504.00897 https://arxiv.org/abs/2504.00897 , 2025

  20. [28]

    Euler stratifications of hypersurface families

    Simon Telen and Maximilian Wiesmann. Euler stratifications of hypersurface families. arXiv:2407.18176 https://arxiv.org/abs/2407.18176 , 2024

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.