{"id":"321170ee-a6e4-42a3-b145-22796f6bd504","arxiv_id":"2607.12919","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Under JC, K2P, and K3P substitution models, the topology of a level-1 phylogenetic network is fully identifiable from leaf-pattern distributions, and trees can be distinguished from networks unless the network is a tree with 2-blobs.","lead":"This paper proves that the branching structure of certain evolutionary networks (level-1 phylogenetic networks) can be recovered exactly from DNA sequence patterns under three standard mutation models, at every point in the parameter space, not just generically. It also shows that a non-tree network almost always produces a detectable signal that no ordinary tree can mimic, so reticulate evolution like hybridization leaves a statistical footprint.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unproved Lemma 5.1 (2-sub-blob suppression) is the load-bearing step for tree-network distinguishability; if the inherited [37] argument does not generalize, Corollary 5.8 does not follow.","rationale":"The reader's weakest-assumption diagnosis is correct: Lemma 5.1 is the least secure load-bearing step. It is used before the induction in Lemmas 5.5 and 5.6, so a failure there would invalidate the tree-network distinguishability result, one of the paper's two headline contributions. The paper is otherwise strong: the small-case invariants are explicit and checkable, the K3P level-1 computations appear internally consistent, and the main reduction logic is plausible. I found no fatal internal contradiction, but the omitted proof of Lemma 5.1 (and to a lesser extent the dependence on [17]) justifies keeping the verdict conditional rather than accepting without reservation.","tokens_in":33704,"tokens_out":17167,"duration_ms":149442,"concrete_test":"Enumerate and symbolically test the smallest non-trivial case: a trinet (or 4-leaf network) whose non-trivial 3-blob contains a 2-sub-blob with two reticulation vertices. Write the JC/K2P Fourier parameterizations of N and of N' after suppression; use polynomial elimination (e.g., in Sage) to check whether every point of M_N lies in M_N'—equivalently, whether for generic parameters of N there exist parameters of N' satisfying the coordinate equations. Repeat for root-trapping and non-root-trapping 2-sub-blobs. If the elimination ideal does not contain the parameterization of N, Lemma 5.1 is false; if it does, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5's main conclusion (Cor 5.8) depends on reducing arbitrary networks to trinets with no 2-sub-blobs via Cor 5.2. The engine is Lemma 5.1, stated without proof: suppressing a 2-sub-blob B gives M_N ⊆ M_N' (equality if B does not trap the root). The paper says this follows by a 'straightforward generalization' of [37, Sec. 5] from 2-blobs to 2-sub-blobs. That is not self-evident: a 2-sub-blob inside a larger k-blob can contain multiple reticulations and cycles, and after contraction the induced transition between the two attachment points must be exactly representable by a single edge whose transition matrix lies in the JC/K2P/K3P model. [37] treats blobs, not arbitrary sub-blobs; the needed closure properties are asserted but not verified for the contraction. If containment fails in any edge case, Cor 5.2 fails and the trinet inequalities (Lemmas 5.5/5.6) cannot be applied to restrictions containing such substructures, so the claim that a network with a non-trivial m-blob (m≥3) is never tree-like loses its proof. A secondary but related incompleteness is that Theorem 4.9 also relies on [17, Thm 14(b)] without reproducing it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies full identifiability of phylogenetic network topologies from the leaf-pattern distribution under the JC, K2P, and K3P substitution models. In Section 4 it proves that the semi-directed parameter of a level-1 network is identifiable at every point of the biologically restricted parameter space Θ0, up to the placement of reticulation vertices in triangles; the proof proceeds by establishing trinet/quarnet inequalities and then reducing arbitrary level-1 networks to 3- and 4-leaf restrictions via marginalization. In Section 5 it turns to tree-network distinguishability: for JC and K2P it claims that any trinet with a non-trivial 3-blob is distinguishable from a 3-star tree, and via a suppression lemma for 2-sub-blobs it derives the broader statement (Corollary 5.8) that a network containing a non-trivial m-blob with m≥3 cannot share a leaf-pattern distribution with a network without such a blob, e.g. a tree. The main results are thus: full level-1 identifiability for all three models, and a positive resolution of Conjecture 2.16 of [14] for arbitrary level under JC and K2P, conditional on an unproved structural lemma.","tokens_in":34064,"tokens_out":10256,"duration_ms":102061,"significance":"If the results hold, they constitute a substantial advance over prior generic identifiability results: the level-1 theorem is non-generic and covers K3P as well as JC and K2P, and the tree-network distinguishability result gives a strong, falsifiable signature of reticulation. The paper also explicitly draws consequences for statistical consistency of network inference and for coalescent-based models. The proofs are largely explicit: the invariants are written down as concrete polynomials and the positivity arguments on Θ0 are checkable. However, the Section 5 results currently rest on Lemma 5.1, which is stated without proof and is load-bearing for Corollary 5.2, Lemmas 5.5/5.6, and Corollary 5.8; until that lemma is supplied, the tree-network distinguishability theorem is not fully established.","major_comments":[{"comment":"Lemma 5.1 is stated without proof and is load-bearing for the main tree-network distinguishability claim. The sentence that it 'follows by a straightforward generalization of the results in Section 5 of [37]' is not an argument. The generalization from 2-blobs to 2-sub-blobs inside a larger k-blob is nontrivial: one must show that marginalizing over all internal states of the sub-blob yields, between the two attachment points, an effective transition matrix that still lies in the JC/K2P/K3P model, including when the sub-blob contains multiple reticulation cycles and when it traps the root. Corollary 5.2, the no-2-sub-blob reductions in Lemmas 5.5 and 5.6, and hence Corollary 5.8 all depend on this. Please provide a full proof or a precise reduction to [37] that covers arbitrary 2-sub-blobs, not just 2-blobs.","section":"§5.1, Lemma 5.1"},{"comment":"The proof of Theorem 4.9 delegates the key combinatorial reduction to [17, Theorem 14(b)] without stating that theorem. This external result is the sole justification for the assertion that any two distinct level-1 networks (modulo triangles and contracting triangles) differ on a 3- or 4-leaf restriction in one of the enumerated ways. The theorem's hypotheses and conclusion should be reproduced or stated precisely, and the exact meaning of 'modulo the placement of the reticulation vertices in any triangles and contracting triangles to single vertices' should be made explicit, so that the reader can verify the theorem applies to every case used in the proof.","section":"§4.3, Theorem 4.9"},{"comment":"The proof of Lemma 5.4 is too compressed for a lemma that carries the induction in Lemmas 5.5 and 5.6. In Case 1, the claim that a non-trivial 2-blob in the basin has at least one attachment point (for otherwise it would be a 2-sub-blob in N) needs a precise boundary-count argument; as written it is not obvious why exactly two vertices of the blob are adjacent to V\\W. In Case 2, the notion of a 'lowest' reticulation presupposes a partial order that is not defined for semi-directed networks, and the final step that the blob containing u' or r* is a non-trivial 3-blob requires checking the number of incident edges. Please expand this proof.","section":"§5.2, Lemma 5.4"}],"minor_comments":[{"comment":"The lemma says the models are 'on the parameter set Θ0(N)' for both M and M'; presumably M' is on Θ0(N'), the parameter set of the suppressed network. Please clarify.","section":"§5.1, Lemma 5.1"},{"comment":"In the base case k=1, the text says 'we identify edges d and e', and edges e and e''' using the edge labelling from Figure 4, but Figure 4 does not label e' and e''. Please align the notation.","section":"§5.2, Lemma 5.5"},{"comment":"The discussion of restriction mentions that 2-blobs can be more complex at higher levels and refers to Section 5; it may be helpful to state explicitly there that 2-sub-blobs will not be suppressed in Definition 2.2.","section":"§2.1, Definition 2.2"},{"comment":"Several positivity computations are summarized with expressions such as 'the reader can check'; these are checkable, but a small comment on how the invariant multisets are compared would improve readability.","section":"Appendix A, Lemma 4.6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong contribution to algebraic phylogenetics, and I expect the main claims are correct. The decisive issue is not the algebraic invariant arguments but the unproved structural Lemma 5.1, which is central to Section 5. I would not reject the paper, but I would require the authors to supply a complete proof or a genuinely detailed reduction to [37] before publication. The reliance on [17, Thm 14(b)] should also be made transparent by stating the theorem. The level-1 part may be acceptable once the external theorem is stated; the Section 5 part needs the missing lemma."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Full identifiability for level-1 networks on Θ0 is a genuine step change, and it holds up. The tree-network result is close, but Lemma 5.1 is an unproved load-bearing step that needs to be closed before I'd call it a theorem.\n\nWhat's new: previous work only gave generic identifiability (off a measure-zero set). This paper proves full identifiability at every point of Θ0 for level-1 networks under JC, K2P, and K3P, modulo triangle reticulation placement. It also extends the trinet inequality to arbitrary level under JC and K2P, positively answering Conjecture 2.16 of [14]. The restriction framework (Theorem 3.6) is clean and useful: if restricted models are disjoint, the full models are disjoint. The small-case invariant polynomials in Section 4 are explicit; I spot-checked a couple of substitutions and they work. Corollary 4.10, describing the intersection of two level-1 models as the union of models of their maximal shared displayed networks, is a strong statement and well-argued by induction.\n\nThe soft spot is exactly where the stress-test points. Lemma 5.1 asserts that suppressing a 2-sub-blob gives a model containment (equality when the blob does not trap the root), and the proof is \"a straightforward generalization of [37, Section 5]\". That is not good enough for a lemma that carries Section 5. The generalization from 2-blobs to 2-sub-blobs inside a larger k-blob requires the contraction to be realizable as a single edge in the model, and the paper does not show the closure properties hold in that setting. The claim may well be true, but it is not self-evident, and if it fails in a corner case then Corollary 5.2 and Corollary 5.8 fall. The referee should demand a complete proof of Lemma 5.1, or a precise citation with the proof. The use of [17, Thm 14(b)] in Theorem 4.9 is a smaller gap: if that theorem is published, restating it would help; if not, the proof is incomplete there too.\n\nI would send this to peer review. The level-1 full identifiability result alone is a major contribution and deserves a serious referee. The tree-network part is promising and likely right, but incomplete as written. I would cite the level-1 result now and wait on the tree-network one. It's a good paper for a reading group — there's a real chance the 2-sub-blob generalization is subtle.","headline":"Full identifiability for level-1 networks on Θ0 is a genuine step change; the tree-network result is close but Lemma 5.1 needs a real proof before I'd call it a theorem.","tokens_in":90,"tokens_out":3722,"would_cite":true,"duration_ms":64360,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["92D15","62R01"],"pacs":[],"model":"deepseek-v4-flash","headline":"For level-1 phylogenetic networks, the leaf-pattern distribution determines the semi-directed topology under JC, K2P, and K3P, and most networks are distinguishable from all trees.","keywords":["phylogenetic networks","identifiability","Jukes-Cantor model","Kimura models","site-pattern distribution","trinet inequality","displayed trees","level-1 networks"],"falsifier":"Find a trinet with a non-trivial 3-blob and a 3-star tree plus edge parameters in (0,1) for which the JC invariant Q = q_ACC q_CAC q_CCA − q_AAA q_TCG^2 equals zero at a point in the network model; the paper proves this cannot happen. More directly, if any network with a 3-blob and a tree share a point in their leaf-pattern distributions with positive parameters, Corollary 5.8 is false.","tokens_in":33603,"feed_emoji":"🧬","tokens_out":5116,"duration_ms":51918,"temperature":0.7,"pith_summary":"The paper proves that under the Jukes-Cantor, Kimura 2-parameter, and Kimura 3-parameter substitution models, the semi-directed topology of a level-1 phylogenetic network is fully identifiable from the distribution of leaf patterns, at every parameter point in a biologically reasonable space. It also proves that under JC and K2P, no phylogenetic tree can produce the same leaf-pattern distribution as a network containing a non-trivial blob with at least three incident edges. The only networks that can be mistaken for trees are trees augmented with small 2-blobs. If correct, these results make network inference statistically consistent for level-1 networks and give a formal sense in which reticulate evolution leaves a detectable signature in DNA sequence data.","feed_headline":"Level-1 networks are fully identifiable from DNA data","feed_subtitle":"Under JC, K2P, and K3P substitution models, network topology is recoverable; most networks also can't be mistaken for a tree.","key_machinery":"The argument uses displayed-tree mixture models and their Fourier parametrization: under group-based models, the network distribution is a convex combination of tree distributions. The load-bearing device is the trinet inequality. For JC, the polynomial Q = q_ACC q_CAC q_CCA − q_AAA q_TCG^2 is zero on every trinet without a non-trivial 3-blob (in particular on a 3-star tree) and strictly positive on every trinet with a non-trivial 3-blob; a K2P analogue uses Q = q_AGG q_GAG q_CCA^2 − q_AAA q_GGA q_TCG^2. Marginalization lemmas transfer these inequalities from 3-leaf subnetworks to the full network, and a separate lemma, stated without proof via a generalization of earlier work, shows that su","core_discovery":"The central claim is Theorem 4.9: two distinct level-1 semi-directed phylogenetic networks, compared modulo the placement of reticulation vertices inside triangles, have disjoint leaf-pattern models on the restricted parameter space for JC, K2P, and K3P. So the network topology can be recovered from exact site-pattern probabilities without generic-position caveats. The second claim, Corollary 5.8, says that under JC and K2P, if one network contains a non-trivial m-blob with m at least 3 and another does not, their models have empty intersection; in particular, no such network is statistically indistinguishable from a tree. The proof works by restriction: explicit polynomial inequalities call","pith_inferences":["If the same reasoning extends to equivariant substitution models, similar full identifiability may hold for models beyond group-based ones, such as the strand-symmetric model; the paper leaves this open.","The 2-blob caveat suggests that 2-blobs are statistically invisible under these models—a testable claim: simulate a tree and a 2-blob-augmented tree with the same displayed trees and compare their site-pattern distributions.","The K3P gap at arbitrary level looks bridgeable by replacing the single invariant with a simultaneous set of invariants, since the level-1 K3P case already requires four invariants.","A practical consequence not discussed in the paper: the trinet inequality could serve directly as a test statistic for detecting reticulation from quartet site patterns."],"forward_implications":["Level-1 network reconstruction from sequence data is statistically consistent under JC, K2P, and K3P, up to the semi-directed network and triangle placement.","Under JC and K2P, a network with a non-trivial 3-blob can never be inferred as a tree by any method based exactly on leaf-pattern distributions, so reticulation becomes statistically testable.","The intersection of two level-1 network models is exactly the union of the models of their maximal shared displayed networks, so residual ambiguity is fully described.","The trinet inequality answers a previously open conjecture and extends tree-network distinguishability from low-level networks to arbitrary level under JC and K2P.","The combinatorial results transfer to several coalescent-based models, giving identifiability results for classes of galled tree-child networks of arbitrary level."],"fun_headline_variants":["Level-1 networks fully identifiable from DNA, no caveats","DNA data recovers level-1 network topology exactly","Full identifiability for level-1 networks under JC, K2P, K3P","No generic-position limits: level-1 networks identifiable"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The proof that suppressing a 2-sub-blob leaves the model of the original network contained in the model of the suppressed network is stated without proof, citing a straightforward generalization of a previous section; if that containment fails in some edge case, the lift from trinets to all networks collapses.","fun_headline_variants_meta":{"raw":{"variants":["Level-1 networks fully identifiable from DNA, no caveats","DNA data recovers level-1 network topology exactly","Full identifiability for level-1 networks under JC, K2P, K3P","No generic-position limits: level-1 networks identifiable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000381,"raw_usage":{"total_tokens":1893,"prompt_tokens":814,"completion_tokens":1079,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":1005}},"tokens_in":558,"tokens_out":1079,"duration_ms":11206,"temperature":1.0,"reasoning_tokens":1005,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T06:11:26.463742+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Find a trinet with a non-trivial 3-blob and a 3-star tree plus edge parameters in (0,1) for which the JC invariant Q = q_ACC q_CAC q_CCA − q_AAA q_TCG^2 equals zero at a point in the network model; the paper proves this cannot happen. More directly, if any network with a 3-blob and a tree share a point in their leaf-pattern distributions with positive parameters, Corollary 5.8 is false.","supporting_citations":[],"review_version":2}