{"id":"fd6f9bc2-8157-458d-9b25-bef4a737e390","arxiv_id":"2507.02719","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The ML degree is monotone with respect to faces of the defining polytope for scaled toric models, settling a conjecture of Coons and Sullivant.","lead":"This paper proves a conjecture in algebraic statistics: the maximum likelihood (ML) degree of a toric model is at least the ML degree of any facial submodel. The proof relies on a perturbation that introduces zero entries into the data, and the result has consequences for discrete graphical and quasi-independence models.","discovery_kind":"first_principles","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the generic-avoidance step in Theorem 6 is terse but can be made rigorous; the parameter-continuation proof is sound.","rationale":"I read the proof in good faith. The central claim is Theorem 6: ML degree is monotone with respect to faces. The proof reduces to the facet case by induction, specializes generic data to data with zeros on the complement of the face, and exhibits MLdeg(X_{A_F,c_F}) many isolated lifted solutions of the full likelihood system. The only truly delicate point is the generic avoidance of the hypersurface Z(h) at all facial critical points. The reader flagged this same step, and I agree it is the least detailed part of the argument. However, reconstructing the incidence dimension count shows the step is valid: the incidence variety for the facial likelihood equations is irreducible and projects dominantly to the facial parameter torus, so a nonzero h cannot vanish on the whole incidence; consequently the bad data set is a proper closed subvariety. The paper's phrasing 'dim(X_{A_F,c_F}) = d' appears to be a typo for d−1, and the L1/L2 decomposition is slightly redundant because V(L1) ⊂ V(L2), but neither issue affects correctness. The tropical section is explicitly incomplete and does not bear on the main theorem. Therefore the reader's ACCEPT verdict remains appropriate.","tokens_in":12543,"tokens_out":49085,"duration_ms":603722,"concrete_test":"For a nontrivial instance (e.g., the 3-cube of Example 1 with a facet F), construct L_{A_F}, h = g|θ_d=0, and for several random generic data vectors u_F compute all solutions θ̂_F of L_{A_F}; verify that h(θ̂_F) ≠ 0 for each and that the Jacobian of L2 at the lifted point (θ̂_F, 0) is nonsingular. This directly checks the generic-avoidance and isolation step of Theorem 6.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3, proof of Theorem 6. The most delicate step is the choice of generic u_F so that no facial critical point θ̂_F lies in Z(h), where h = ∂_{θ_d}(θ_d g)|_{θ_d=0}. The paper's one-line dimension justification is imprecise: the vanishing set of h in the (θ0,...,θ_{d-1})-space has dimension d−1, while the facial variety has dimension d−1, not d as printed. This wording does not, however, conceal a real gap. The incidence variety I = {(u_F, θ) : L_{A_F}(u_F; θ) = 0} is irreducible, being the Zariski closure of the parameterization by θ and the affine data solutions. Its projection to the θ-torus is dominant, since for any θ with f_F(θ) ≠ 0 the data u_F = p_F(θ) makes θ a critical point after setting θ0 = 1/f_F(θ). A nonzero Laurent polynomial h therefore cannot vanish on all of I; hence dim(I ∩ Z(h)) = dim I − 1 and the projection of this intersection to the data space lies in a proper closed subset. Thus a Zariski-open set of u_F avoids Z(h) at every critical point. The subsequent block-triangular Jacobian computation then correctly shows the lifted points are isolated in V(L_A(θ; u˜)). No load-bearing flaw is identified.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proves that the maximum likelihood (ML) degree of a toric model is monotone with respect to the face poset of the defining polytope: for any scaled toric variety X_{A,c} and any face F of conv(A), the ML degree of the facial submodel X_{A_F,c_F} is at most that of the full model. This settles a conjecture of Coons and Sullivant. The proof uses a parameter continuation argument, extending generic data to data with zeros on the complement of the face, and shows that the lifted critical points of the facial submodel are isolated solutions of the full model's likelihood system. The paper also discusses implications for data zeros, connects the result to tropical likelihood degenerations, and applies the main theorem to discrete graphical models and quasi-independence models.","tokens_in":12803,"tokens_out":13433,"duration_ms":148410,"significance":"If the main theorem is correct, it resolves a natural and previously open conjecture in algebraic statistics, providing a clean structural statement about ML degrees under taking facial submodels. The proof strategy, based on Morgan–Sommese parameter continuation and a block-triangular Jacobian computation, is potentially reusable beyond the specific setting. The applications to graphical and quasi-independence models are immediate and strengthen results that previously required the full model to have ML degree one. The paper also contains reproducible computational experiments (e.g., Section 4, Table 2) and proposes a tropical degeneration framework that, while not fully rigorous, points to a promising direction for refining the monotonicity statement. The central claim is falsifiable and the proof, after a repair described below, is sound.","major_comments":[{"comment":"The generic-avoidance step is not justified as written. The sentence 'Its vanishing set has dimension d − 1 while dim(X_{A_F,c_F}) = d' contains two errors: the vanishing set of a nonzero polynomial in d−1 variables has dimension at most d−2 (or d−1 if one works in the full (θ_0,...,θ_{d-1})-space), and the facial variety has dimension d−1, not d. Moreover, a dimension comparison alone cannot prove the avoidance claim, since a hypersurface can contain a variety of the same dimension. The needed argument is to consider the incidence variety I = {(u_F, θ) : L_{A_F}(u_F; θ) = 0}, show it is irreducible and projects dominantly onto the θ-torus, and then note that the nonzero Laurent polynomial h = ∂_{θ_d}(θ_d g)|_{θ_d=0} cannot vanish identically on I; hence the projection of I ∩ Z(h) to the data space is a proper closed subset. This yields a Zariski-open set of u_F for which no critical point lies in Z(h). The authors should incorporate this argument to make the proof complete.","section":"Section 3, proof of Theorem 6"},{"comment":"After defining α, the assertion that every solution θ̂_F is isolated in V(L2) requires that the lower-right Jacobian entry θ0 ∂_{θ_d}(θ_d g)|_{θ_d=0} is nonzero at θ̂_F. This is exactly the avoidance condition discussed in the previous comment. The paper's one-line justification is insufficient; the proof should explicitly state that the chosen generic u_F simultaneously avoids the finitely many critical points of the facial model and the hypersurface Z(h), and that this is possible because both conditions hold on a Zariski-open set of data.","section":"Section 3, proof of Theorem 6"}],"minor_comments":[{"comment":"The notation ∂θdθdg is terse and ambiguous; it should be written as ∂_{θ_d}(θ_d g)|_{θ_d=0} throughout.","section":"Section 3, proof of Theorem 6"},{"comment":"The inequality α ≥ 0 holds because every lattice point outside the facet F has positive last coordinate after the unimodular transformation, but this justification is omitted.","section":"Section 3, proof of Theorem 6"},{"comment":"Table 1 displays scalings as 3×3 matrices while the design matrix A in the same example is 5×9; the authors should clarify that the matrices are flattened into length-9 scaling vectors.","section":"Section 4, Example 8"},{"comment":"The entries '∞' in Table 2 are not defined in the caption; state explicitly that they denote positive-dimensional solution components of the likelihood system.","section":"Section 4, Table 2"},{"comment":"The notation δ_F(a) with a ∈ A is ambiguous because A is a matrix; use lattice points a_i ∈ Z^d instead.","section":"Section 5, Equation (5)"},{"comment":"The phrase 'This induces a regular subdivision of the Cayley polytope' would benefit from a brief definition or reference for the Cayley configuration, since it is central to the claimed failure of the sufficient criterion.","section":"Section 5, Example 11"},{"comment":"The proof relies on [GMS06, Lemma A.2] for the face property; stating the content of that lemma would make the proof more self-contained.","section":"Section 6, Corollary 13"}],"recommendation":"major_revision","confidential_remarks":"The main theorem is correct and the proof can be repaired with the incidence-variety argument described in the major comments. The paper is within the scope of the journal; the self-citations are appropriate and not excessive. I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper settles a genuine open problem: Coons and Sullivant's conjecture that the ML degree of a facial submodel of a toric model never exceeds the ML degree of the full model. I read the main proof carefully and it holds up. The new ideas are the use of data zeros to lift facial critical points and the parameter-continuation bound from Morgan and Sommese. That is a clean way in, and it gives the theorem for all scalings, not just generic ones. The corollaries for undirected graphical models and quasi-independence models are real, and they fill a gap that earlier ML-degree-one results left open.\n\nStrengths: the proof is honest about what it does not do. The tropical section is explicitly exploratory; Example 11 shows a naive tropical-basis criterion fails, and Problem 12 is stated as open. That is the right way to write a section like that. The examples are useful, and Table 2 gives concrete symbolic computations.\n\nSoft spots: the most delicate step in Theorem 6 — choosing generic facial data so that no critical point lies on the hypersurface h = 0 — is dispatched too quickly. The printed dimension count is wrong as written: h lives in variables θ1,...,θ_{d-1}, so its zero set as a cylinder in (θ0,...,θ_{d-1}) has dimension d-1, and the facial variety has projective dimension d-1, not d. Fortunately the argument survives. The incidence variety I is irreducible and dominates the θ-torus, so a nonzero h cannot vanish on all of I; the bad data lie in a proper closed set. That is a one-paragraph fix, not a hole. There are also a few typos and the 'code in Maple' promised in Example 10 is not in the text, so the computations there are not independently checkable from the paper alone.\n\nBottom line: this is a solid paper with a correct main theorem, a repairable gap in exposition, and useful applications. Anyone working on likelihood geometry or algebraic statistics will want to know it. It deserves a serious referee, probably with a request for a clearer treatment of the generic-avoidance step.","headline":"Clean proof of a genuinely open conjecture; the main theorem is correct, with one terse generic-avoidance step that needs a slightly more careful exposition but no real gap.","tokens_in":13338,"tokens_out":3725,"would_cite":true,"duration_ms":42550,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["14M25","62R01"],"pacs":[],"model":"deepseek-v4-flash","headline":"The ML degree of a toric model cannot increase on any facial submodel.","keywords":["maximum likelihood degree","toric varieties","facial submodels","monotonicity conjecture","parameter continuation","likelihood equations","discrete graphical models","quasi-independence models"],"falsifier":"Compute both ML degrees for a specific scaled toric variety and one of its faces—for instance, the all-ones three-dimensional cube (ML degree 8) and a facet (ML degree 4). Finding any face $F$ with $\\operatorname{MLdeg}(X_{A_F,c_F}) > \\operatorname{MLdeg}(X_{A,c})$, or generic facial data whose lifted point has singular Jacobian, would refute Theorem 6.","tokens_in":12344,"feed_emoji":"📐","tokens_out":10285,"duration_ms":108600,"temperature":0.7,"pith_summary":"This paper proves that the maximum likelihood degree—the number of complex critical points of the likelihood equations for generic data—never increases when a toric model is restricted to a face of its defining polytope. The result settles a conjecture from earlier work on quasi-independence models and covers every scaling of the toric variety, not only generic scalings. Because the face poset of the polytope organizes the model's submodels, the inequality gives a uniform bound: any facial submodel has likelihood-estimation complexity at most that of the full model. The proof deforms generic data to data with zeros outside the face and uses parameter continuation to show the full system retains at least as many isolated solutions as the facial system.","feed_headline":"Restricting a toric model to a face never increases its ML degree","feed_subtitle":"The theorem proves facial submodels' likelihood complexity is bounded by the original model.","key_machinery":"The load-bearing object is the scaled toric variety $X_{A,c}$, defined as the closure of the monomial parametrization $\\theta \\mapsto (c_i \\theta_0 \\theta^{a_i})$, together with its facial submodels $X_{A_F,c_F}$ obtained by restricting to lattice points of a face $F$ of $\\operatorname{conv}(A)$. The ML degree counts complex solutions of the likelihood equations for generic data. The proof mechanism is the parameter continuation theorem, which says that specializing parameters in a polynomial system cannot increase the number of isolated solutions, combined with a block-triangular Jacobian computation that certifies the lifted facial solutions are isolated in the deformed full system.","core_discovery":"The central claim is Theorem 6: for a scaled toric variety $X_{A,c}$ and any face $F$ of $\\operatorname{conv}(A)$, the facial submodel $X_{A_F,c_F}$ satisfies $\\operatorname{MLdeg}(X_{A_F,c_F}) \\leq \\operatorname{MLdeg}(X_{A,c})$. The proof works by induction on dimension, reduces to the case where $F$ is a facet, and studies the likelihood equations after extending facial data by zeros. Each generic facial critical point lifts to a solution $\\hat{\\theta} = (\\hat{\\theta}_F, 0)$ of the full system, and this point is isolated because the Jacobian is block triangular with lower-right entry $\\theta_0 \\partial_{\\theta_d}(\\theta_d g)|_{\\theta_d=0}$, which is nonzero for generic facial data by a dimension count. The parameter continuation theorem then implies the number of isolated solutions can only decrease under data specialization, giving the inequality.","pith_inferences":["If the block-triangular Jacobian structure is the only ingredient needed, the same monotonicity should hold for other families of very affine varieties with facial submodels; testing it on Gaussian graphical models or other non-toric likelihood varieties would be a direct next step.","Example 8's zero-data computations suggest that the number of isolated solutions is governed by which zeros fall inside or outside the face support; a data-plus-scaling discriminant would give a complete answer to when the count drops, stays, or becomes infinite.","A tropical basis for the Puiseux likelihood equations would refine the inequality into an explicit bookkeeping of which critical points belong to each face, potentially yielding a combinatorial formula for the ML-degree drop along flags of faces.","Because the theorem holds for arbitrary scalings, it constrains the Euler stratification of the parameter space: the ML-degree strata over any face cannot exceed the corresponding stratum of the full polytope."],"forward_implications":["For an undirected graphical model, the ML degree of any induced-subgraph model is at most the ML degree of the original graph model (Corollary 13).","For quasi-independence models, restricting to an induced subgraph of the associated bipartite graph can only lower or preserve the ML degree (Corollary 16).","Whenever the data linear space meets the scaled toric variety transversally, the number of complex likelihood solutions is at most the ML degree, even for non-generic data with zeros (Corollary 9).","Monotonicity fails for arbitrary non-facial submodels: deleting one column of the design matrix can raise the ML degree from 1 to 3 (Example 7).","The tropical likelihood degeneration of Section 5 lets one watch which solutions survive the face limit, and a tropical basis would turn this into a precise combinatorial refinement, which the paper leaves as Problem 12."],"supporting_citations":[{"why":"Poses the monotonicity conjecture and proves the ML-degree-one base case that the present proof extends.","marker":"[CS21]"},{"why":"Provides the parameter continuation theorem, the engine that turns data specialization into an inequality of isolated-solution counts.","marker":"[MS89]"},{"why":"Gives the likelihood equations for scaled toric varieties and the generic-scaling identity between ML degree and normalized volume.","marker":"[ABB+19]"},{"why":"Provides the lemma that the polytope of an induced graphical submodel is a face of the full model's polytope, which makes Corollary 13 an application of Theorem 6.","marker":"[GMS06]"},{"why":"Supplies the h*-vector monotonicity used for the generic-scaling case, where ML degree equals normalized volume.","marker":"[Sta93]"},{"why":"Supplies the fact that generic-data critical points have nonzero coordinates and motivates the data-zero discussion of Section 4.","marker":"[GR13]"}],"fun_headline_variants":["Facial submodels cap ML degree at original model","ML degree drops when toric model is restricted to a face","Toric model faces never raise maximum likelihood degree","ML degree monotone under facial restriction for toric models","Submodels on faces have ML degree no higher than full model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proof assumes that a certain nonzero polynomial built from the face direction does not happen to vanish at all the solution points of a generic smaller problem, because the face has fewer coordinates than the whole model; if it did vanish, the key step showing the lifted solutions stay isolated would fail.","fun_headline_variants_meta":{"raw":{"variants":["Facial submodels cap ML degree at original model","ML degree drops when toric model is restricted to a face","Toric model faces never raise maximum likelihood degree","ML degree monotone under facial restriction for toric models","Submodels on faces have ML degree no higher than full model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000724,"raw_usage":{"total_tokens":3174,"prompt_tokens":799,"completion_tokens":2375,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":415,"completion_tokens_details":{"reasoning_tokens":2295}},"tokens_in":415,"tokens_out":2375,"duration_ms":16808,"temperature":1.0,"reasoning_tokens":2295,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:24:03.651033+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute both ML degrees for a specific scaled toric variety and one of its faces—for instance, the all-ones three-dimensional cube (ML degree 8) and a facet (ML degree 4). Finding any face $F$ with $\\operatorname{MLdeg}(X_{A_F,c_F}) > \\operatorname{MLdeg}(X_{A,c})$, or generic facial data whose lifted point has singular Jacobian, would refute Theorem 6.","supporting_citations":[],"review_version":1}