Pith. sign in

REVIEW 3 major objections 4 minor 61 references

Methodological considerations for semialgebraic hypothesis testing with incomplete U-statistics

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The SDL test for semialgebraic models works in phylogenetics, but only with careful tuning.

desk verdict A genuinely useful, honest empirical evaluation of the SDL test on phylogenetic models, with a real but disclosed tuning-on-evaluation issue that tempers the headline performance claims. read the letter →

arxiv 2507.13531 v1 pith:6DD3WDBU submitted 2025-07-17 q-bio.PE stat.ME

classification q-bio.PEstat.ME MSC 62F0362F4092D15
keywords semialgebraicmodelshypothesistestingincompleteU-statisticsGaussianmultiplierbootstrapphylogeneticsquartetconcordancefactorCavender-Farris-Neymanmodelintersection-uniontest
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper evaluates whether the SDL test—a stochastic, bootstrap-based method for testing hypotheses defined by polynomial equality and inequality constraints—can be used reliably on biological semialgebraic models. Working through trinomial models from phylogenomics and a Cavender-Farris-Neyman four-taxon model, it finds that the method can achieve strong performance, competitive with deterministic tests, but only when the user makes good implementation choices. The core message is that naive application is unlikely to work well: kernel order, the choice of redundant constraints, the amount of symmetrization, and whether a reducible model is tested component-wise all materially change the rejection region and error rates. The paper therefore proposes practical safeguards, including simulation-guided tuning and automatic constraint augmentation, while explicitly offering no new theory.

What carries the argument

The machinery is the randomized incomplete U-statistic: a user-specified symmetric kernel $h$ on $m$ data points estimates the constraint vector $f(\theta)$, and the test statistic is the studentized maximum of the average of such kernel evaluations over random subsamples. A Gaussian multiplier bootstrap calibrated with a divide-and-conquer estimator of the H\'ajek projection (the conditional expectation of the kernel given one data point) supplies the null distribution, which is why the method remains valid at singularities and boundaries. The paper's contribution is to show how the knobs of this machinery—kernel order $m$, sampling budget $N$, bootstrap size $A$ and $n_1$, the constraint list, and the symmetrization step—act as levers on the rejection region.

What would settle it

Run the SDL test with the paper's recommended choices on a different semialgebraic model in higher dimensions—for example a five-taxon CFN model or a larger phylogenetic network—and check whether the empirical type I error at nominal 0.05 stays at or below 0.05 across several model parameters, including singular points; any model where the test becomes anti-conservative at the suggested settings would show the guidance is not generic.

Watch

Extended reading notes

Core claim

The SDL test, built from randomized incomplete U-statistics and a Gaussian multiplier bootstrap, works on semialgebraic phylogenetic models, but its finite-sample behaviour is governed by user choices that the original proposal left under-specified. With a moderately increased kernel order, redundant constraints generated by random convex combinations, and partial rather than full symmetrization of the kernel, the test matches or approaches the power of deterministic tests on the trinomial quartet-concordance models while remaining conservative or valid at nominal levels. On a four-taxon CFN model, the same procedure yields topology tests without likelihood computation. The authors' central claim is that the method is practically usable across all models they considered, yet no general rules for choosing parameters exist; simulation at several model points is the most informative guide.

Load-bearing premise

The practical guidance rests on a handful of low-dimensional examples ($n=300$ trinomial models and one four-taxon CFN setting), and the paper says no general rules could be extracted; if those settings are not representative, the recommended tuning steps may not transfer to other semialgebraic models.

Editorial extensions

If this is right

  • Users should simulate from the model before trusting SDL p-values, testing multiple parameter points including singularities; no one-size-fits-all choice of kernel order $m$ exists.
  • Redundant constraints, e.g., random convex combinations of the defining inequalities, can make the rejection region less dependent on an arbitrary semialgebraic description.
  • Partial symmetrization with roughly ten random permutations can substitute for full symmetrization, which is computationally prohibitive when constraint degrees are high.
  • For reducible models, an intersection-union test over irreducible components can improve power and speed relative to testing the whole model directly.
  • The CFN topology test shows that SDL can identify gene-tree topology from semialgebraic descriptions alone, without likelihood computation or optimization.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves implicit that the same tuning discipline would be needed for any new semialgebraic model, not just phylogenetics; the exact settings found here are not portable.
  • A natural next step, not pursued here, is a theoretical account of random partial symmetrization, whose third source of randomness falls outside the existing asymptotic justification.
  • The random-convex-combination constraint augmentation could be viewed as averaging over semialgebraic descriptions; formalizing that average might yield a deterministic recipe for choosing how many redundant constraints to add.
  • The CFN topology-testing idea may scale to larger trees if component decompositions and semialgebraic descriptions are computed automatically.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper evaluates the SDL stochastic hypothesis test for semialgebraic models on four trinomial submodels and a four-taxon CFN model, investigating how kernel order m, choice of constraint description, redundant constraints via random convex combinations, partial symmetrization, and decomposition into irreducible components affect Type I error, power, and p-value stability. The authors report that the method performs well after careful tuning, but they emphasize that they found no general rules for parameter choice and that simulation at a number of model points is the most informative approach.

Significance. If the empirical findings hold, this is a useful practical guide for applying SDL tests to semialgebraic phylogenetic models, and the paper makes several concrete, actionable recommendations: use redundant constraints, consider intersection-union tests for reducible models, and use partial symmetrization with only a few permutations. The manuscript is honest and careful in several respects: the code is publicly available on GitHub, simulations use n = 300 with 1000 datasets for size estimates, and the authors explicitly flag the lack of theory for partial symmetrization and the absence of general rules. The main weaknesses are that the quantitative performance claims are weakened by the selection-on-evaluation-set issue for m, and the breadth of the 'across different settings' claim in the abstract is not matched by the narrow set of simulation settings.

major comments (3)
  1. [§3.4.1] The kernel order m is selected using simulations at the same null points that are later used to report the test's empirical size, so the reported Type I error for the chosen m is an optimistically biased estimate of the procedure's actual Type I error. The text states, 'In the following subsections, we use the largest m which simulations suggest gives a valid test size at a number of model points,' and the subsequent discussion of Model 2 and Figure 2 reports sizes at exactly those points. Because the authors' central recommendation is simulation-guided parameter choice, the quantitative support for the headline claim of 'excellent performance' is weaker than it appears. I recommend splitting the simulations into a tuning set and an evaluation set, or reporting sensitivity analyses over a grid of m values at several different θ points, so that the reported size is not conditioned on the same simulations used to select m.
  2. [§3.4.3] The random partial symmetrization with s permutations is presented as an effective substitute for full symmetrization, but the authors correctly note that 'theory justifying its use is currently lacking' and that it introduces a third source of randomness not covered by the asymptotic justification in [53]. Since this device is later used in the CFN application in Section 4, the validity of the results reported there depends on an unproven approximation. The paper should either provide additional empirical evidence that the unbiasedness of the kernel and the bootstrap approximation are preserved (for example, by comparing partial symmetrization with full symmetrization on a model where full symmetrization is computationally feasible, and by checking the null distribution of p-values), or explicitly restrict the claims about partial symmetrization to empirically observed behavior and mark the CFN inference procedure as exploratory.
  3. [Section 3 and Section 4] The paper's general guidance rests on a small set of low-dimensional simulation settings: four trinomial submodels at n = 300 and one four-taxon CFN model. The authors acknowledge this in §3.4.1 when they say they were 'unable to develop any general rules to apply.' This is an honest limitation, but the abstract's statement that the method 'performs remarkably well across different settings' overstates the evidence, since 'different settings' are a handful of closely related low-dimensional models. I suggest tempering the abstract and the introductory emphasis to say that excellent performance was observed after tuning in the specific models studied, rather than suggesting a general empirical guarantee.
minor comments (4)
  1. [Introduction, paragraph 2] There is a typo: 'ignore the the challenges' should read 'ignore the challenges.'
  2. [§3.4.1] The phrase 'For η = m = 1' is slightly confusing because η and m are conceptually distinct (m = η · max deg(fi)); for Model 1 the degree is 1 so they coincide, but this equivalence should be stated explicitly to avoid confusion.
  3. [§3.4.2] When introducing random convex combinations of inequalities, it would be clearer to state explicitly that a convex combination of polynomials that are non-positive on the model is again non-positive on the model, so the added constraints do not change the null set.
  4. [§3.4.4] The intersection-union test is described as taking the maximum of component-test p-values; it may be worth noting explicitly that this is valid only when each component test is level α, which is the case if the SDL test's conservatism holds for each component.

Circularity Check

1 steps flagged · score 3.0 of 10

Reported Type I error at the chosen kernel order m is partly self-confirming, since m is selected from those same null simulations; no equation-level circularity or load-bearing self-citation is present.

  1. fitted input called prediction [Section 3.4.1 (applied in Sections 3.4.2-3.4.5)]
    "In the following subsections, we use the largest m which simulations suggest gives a valid test size at a number of model points, including singularities and boundary points. For instance, we find that for Model 2 (discussed in the next subsection) m = 5 gave good performance for the boundary parameter point (1/3,1/3,1/3), with empirical test size closely tracking the nominal level (plot not shown)."

    The kernel order m is chosen by inspecting empirical Type I error at null model points, and then the empirical sizes obtained at that chosen m are reported as evidence that the test has valid size. The reported 'valid test size' is therefore the selection criterion rather than an independent evaluation of the procedure: at the selected m, matching the nominal level is partly guaranteed by the selection loop. This makes the favorable Type I error numbers for the headline 'excellent performance' claim less informative than they appear. The paper's own caveat that no general rules could be developed and that simulation at model points is the most informative approach mitigates the concern, but the quantitative support for correct size at the chosen m is still partially self-confirming.

full rationale

This is a methods-evaluation paper, not a derivation, so there is no equation-level circularity in the sense of a claimed result reducing by construction to its inputs. The SDL test itself is imported from [53] (Sturma, Drton, Leung), whose authors are distinct from the present authors, and its asymptotic validity is cited as an external, independently stateable result; no self-citation chain is load-bearing. The main circularity burden is procedural: Section 3.4.1 explicitly selects the kernel order m as the largest value that simulations suggest gives a valid test size at a set of null model points, and then the empirical sizes at that same m are used to document correct size behavior, making the reported Type I error partly an artifact of tuning on the evaluation set. The paper is transparent about this limitation, explicitly states that no general rules were found, and recommends simulation at model points as the most informative approach. The other contributions—constraint augmentation via random convex combinations, partial symmetrization, and intersection-union decomposition—are exploratory illustrations rather than predictive claims, and they are compared against deterministic benchmarks rather than being claimed as derived results. Thus the circularity score is low: one selection-on-evaluation loop weakens the quantitative support for the positive performance claim, but the paper's central content remains independent simulation-based evidence rather than a circular derivation.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The SDL method itself contributes a fixed set of user parameters (m, N, n1, A). This paper adds two more (r and s) and, for its performance claims, tunes m by simulation. No new physical or mathematical entities are postulated; the central assumptions are the semialgebraic representation of the null models and the validity of the bootstrap approximation inherited from the SDL theory.

free parameters (6)
  • m (kernel order) = 1, 5, 15, 25, 45 depending on model; 5 for Models 1-2, 15 for Model 3, 25 for direct Model 4, 5 for IUT
    User-chosen and partly tuned via simulation to keep empirical size near nominal; affects power and p-value stochasticity (Section 3.4.1).
  • N (computational budget) = 1000
    Set to 1000 for n=300; increasing reduces p-value randomness but can affect validity (Sections 3.3 and 3.4).
  • n1 (bootstrap subset size) = 300 (n)
    Follows the [53] suggestion; number of terms used in the divide-and-conquer estimator for g.
  • A (bootstrap draws) = 1000
    Number of W samples used to estimate the p-value; follows the [53] suggestion.
  • r (random convex combinations) = 10 or 100
    Introduced by this paper to reduce dependence on the choice of inequality constraints (Section 3.4.2); no supporting theory is given.
  • s (random permutations for partial symmetrization) = 1, 10, 100
    Introduced to approximate full kernel symmetrization; effective in the examples shown but explicitly lacking theoretical justification (Section 3.4.3).
assumptions (5)
  • domain assumption Data are i.i.d. samples Xi ~ Pθ for θ in Θ
    Assumed in Section 2.1 as the hypothesis-testing setting; all simulations generate i.i.d. multinomial or CFN data.
  • domain assumption The null model Θ0 is a basic semialgebraic set defined by finitely many polynomial inequalities fi ≤ 0 (Eq. 2.1)
    The SDL method requires this representation; the phylogenetic models are presented this way in Section 3.2.
  • domain assumption Unbiased estimators bθj of θj exist and are used to build the kernel h, so that E[h] = f(θ)
    Kernel construction in Section 2.3 assumes this for every constraint polynomial.
  • standard math The Gaussian multiplier bootstrap approximation of the test statistic distribution is valid at the finite sample sizes used
    Inherited from [53, Corollary 2.10 and Theorem 2.4]; the paper does not re-prove it and relies on it for defining p-values.
  • ad hoc to paper Random partial symmetrization with s permutations does not break the unbiasedness or bootstrap approximation
    Used in Sections 3.4.3 and 4; the authors note that theory justifying its use is currently lacking.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Methodological considerations for semialgebraic hypothesis testing with incomplete U-statistics." pith.science (2026). https://pith.science/paper/6DD3WDBU

@misc{pith2026250713531,
  author       = {Pith},
  title        = {Pith review of: Methodological considerations for semialgebraic hypothesis testing with incomplete U-statistics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6DD3WDBU}},
  note         = {Machine review of arXiv:2507.13531}
}
read the original abstract

Recently, Sturma, Drton, and Leung proposed a general-purpose stochastic method for hypothesis testing in models defined by polynomial equality and inequality constraints. Notably, the method remains theoretically valid even near irregular points, such as singularities and boundaries, where traditional testing approaches often break down. In this paper, we evaluate its practical performance on a collection of biologically motivated models from phylogenetics. While the method performs remarkably well across different settings, we catalogue a number of issues that should be considered for effective application.

Figures

Figures reproduced from arXiv: 2507.13531 by the authors.

Figure 1
Figure 1. Each characterizes the frequencies of the three possible quartet gene tree topologies [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 1
Figure 1. Parameter spaces (blue line segments) of four submodels of the trinomial model, with parameter space ∆2 . The submodels capture the form of the quartet Concordance Factor if the species relationships have specific features, as described in the text. a more complete explanation, but knowledge of the application is not necessary for a reader primarily interested in the SDL test for other uses. While each model is comp… view at source ↗
Figure 2
Figure 2. SDL test behaviour for Model 1, with m = 1, 5, and 15 (top row to bottom). The left column shows nominal vs. empirical sizes for the SDL and LR tests; the middle, histograms of p-value differences; and the right, SDL rejection regions. SDL test and the p-values computed with LR, for the same 1000 datasets. The right column depicts the SDL rejection region for all datasets of size n = 300. Importantly, [PITH_FULL_IM… view at source ↗
Figures from the paper (18 more)
Figure 3
Figure 3. Figure 3: Rejection regions for Model 2 under the SDL test using (L to R) a) the constraints y −z ≤ 0, z −y ≤ 0, and 1/3−x ≤ 0; b) replacing the last inequality by 2/3−x−y ≤ 0; c) including r = 10 random convex combinations of the inequalities of (a) ; and d) including r = 100 r…
Figure 4
Figure 4. Figure 4: Rejection regions for Model 3 under the SDL test using (L to R) s = 1, 10, and 100 random permutations to partially symmetrize h˘. For all, m = 15. In [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: (L) Rejection region for Model 4 obtained from SDL test using semial￾gebraic description given above. (R) Rejection region for an Intersection-Union test using the SDL tests for the 3 irreducible components of Model 4 (each essentially Model 2) [PITH_FULL_IMAGE:figure…
Figure 6
Figure 6. Figure 6: Rejection regions for SDL tests of (L-R) (a) the Hardy-Weinberg 2-allele model defined by y 2−4xz = 0, (b) a nodal cubic model defined by (y−1/3)2−6(x− 2/5)2 (x−1/9) = 0, (c) a cuspidal cubic model, defined by (y−1/3)2−(x−1/3)3 = 0. 14 [PITH_FULL_IMAGE:figures/full_fi…
Figure 7
Figure 7. Figure 7: Rejection regions for SDL tests of the cuspidal cubic (L-R) with (a) constraints supplemented by 1/3 − x ≤ 0; (b) constraints supplemented by the inequality from (a) plus r = 10 random convex combinations of inequalities, and (c) constraints supplemented by 3 linear in…
Figure 8
Figure 8. Figure 8: The 4-leaf binary tree topologies, with edge lengths ti . The names Txy|zw indicate the partition of leaves induced by the central edge. Let T be one of the leaf-labeled trees of [PITH_FULL_IMAGE:figures/full_fig_p016_8.png]
Figure 9
Figure 9. Figure 9: Left: The tree T12|34 with edge lengths t1 = t3 = a and t2 = t4 = t5 = b, in units of expected number of substitutions per site. Right: The tree space, with a, b varying from 0 to 1.2. In red, nine parameter pairs with a, b ∈ {0.05, 0.2, 0.8}. The dashed blue curve is …
Figure 10
Figure 10. Figure 10: Aggregated p-values for a test of the true null hypothesis H12|34 (left) and a false null hypothesis H13|24 (right) for datasets in Collection 1. Constraints sets CDM and PDM, and number of convex combinations r = 0 and 20 are varied. For r = 0 and the true H12|34, th…
Figure 11
Figure 11. Figure 11: p-values obtained from the SDL test on Collection 2 for different con￾straint sets: CDM (top 3 rows) and PDM (bottom 3 rows). The hypotheses tested are H12|34 (left 3 columns) and H13|23 (right 3 columns), with r = 0. The SDL test performed quite poorly when testing t…
Figure 12
Figure 12. Figure 12: (right) comparing the amalgamated distribution of W with and without the internal branch inequality for aggregate data from 1000 trees drawn randomly from the treespace shown in [PITH_FULL_IMAGE:figures/full_fig_p022_12.png]
Figure 13
Figure 13. Figure 13: Performance of the SDL test for inferring the tree topology T12|34 using different constraint sets and values of m. Left: CDM and PDM constraints with m = 12. Right: CDD constraints with m = 12, CDD with the inequality of Eq. (4.3) and m = 12, CDD with the inequality …
Figure 14
Figure 14. Figure 14: Performance of 3 methods of topological tree inference on data from Collection 1: (left) the SDL-based inference method using the CDD constraint set with the internal edge inequality, with m = 30 and r = 20, (middle) Maximum Likelihood [35],(right) the SVD method. All…
Figure 15
Figure 15. Figure 15: Gene trees (in red) form within a species tree and network (black ‘tubes’) Considering only trees or networks relating four species, a quartet Concordance Factor (CF) for a fixed network is the vector of probabilities of the 3 possible unrooted topological gene trees …
Figure 16
Figure 16. Figure 16: Rejection regions for Models 1, 2, 3, 4, and Hardy-Weinberg 2-alleles, using deterministic tests, as described in text, with sample size n = 300. Appendix C. Additional Details on the CFN model C.1. Generating sets for the CFN ideal. This section details the derivatio…
Figure 17
Figure 17. Figure 17: Aggregated p-values for a test of the true null hypothesis H12|34 from datasets in Collection 1. Columns correspond to choices of defining polynomials. Rows correspond to the value of r. 32 [PITH_FULL_IMAGE:figures/full_fig_p032_17.png]
Figure 18
Figure 18. Figure 18: Aggregated p-values for a a test of H13|24 (a false null hypothesis) from datatsets in Collection 1. Columns correspond to choices of defining polynomials. Rows correspond to the value of r. The test of H14|23 produced similar results. In the r = 0 case, [PITH_FULL_I…
Figure 19
Figure 19. Figure 19: Performance of the SDL test for inferring the tree topology T12|34. Columns correspond to different CFN model constraints (CDD, CDM, CDR, PDM, PDR), and rows represent the number of convex combinations used, r = 0 and r = 20. Grey levels represent the frequency of cor…
Figure 20
Figure 20. Figure 20: Histogram of p-values for CDM (left) and PDM (right) for a tree in the Felsenstein zone (a = 0.8 and b = 0.05) with n = 10000 bp and m = 12. The reduced long branch attraction bias for SDL is not unique to the parameters used for [PITH_FULL_IMAGE:figures/full_fig_p03…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

61 extracted references · 57 canonical work pages

  1. [53]

    Testing many constraints in possibly irregular models using incomplete U-statistics

    Nils Sturma, Mathias Drton, and Dennis Leung. “Testing many constraints in possibly irregular models using incomplete U-statistics”. In: Journal of the Royal Statistical Society Series B: Statistical Methodology (Mar. 2024), qkae022. issn: 1369-7412. doi: 10.1093/jrsssb/qkae022

  2. [1]

    Identifying the rooted species tree from the distribution of unrooted gene trees under the coalescent

    E.S. Allman, J.H. Degnan, and J.A. Rhodes. “Identifying the rooted species tree from the distribution of unrooted gene trees under the coalescent”. In: Journal of Mathe- matical Biology 62.6 (2011), pp. 833–862

  3. [2]

    Identifiability of parameters in latent struc- ture models with many observed variables

    E.S. Allman, C. Matias, and J.A. Rhodes. “Identifiability of parameters in latent struc- ture models with many observed variables”. In: The Annals of Statistics 37.6A (2009), pp. 3099–3132. doi: 10 . 1214 / 09 - AOS689. url: https : / / doi . org / 10 . 1214 / 09 - AOS689

  4. [3]

    TINNiK: Inference of the tree of blobs of a species network under the coalescent

    E.S. Allman et al. “TINNiK: Inference of the tree of blobs of a species network under the coalescent”. In: Algorithms in Molecular Biology 19.1 (2024), p. 23. doi: 10.1186/ s13015-024-00266-2

  5. [4]

    Quartets and parameter recovery for the gen- eral Markov model of sequence mutation

    Elizabeth Allman and John Rhodes. “Quartets and parameter recovery for the gen- eral Markov model of sequence mutation”. In: AMRX Applied Mathematics Research eXpress 2004 (Jan. 2004). doi: 10.1155/S1687120004020283

  6. [5]

    The identifiability of tree topology for phyloge- netic models, including covarion and mixture models

    Elizabeth Allman and John Rhodes. “The identifiability of tree topology for phyloge- netic models, including covarion and mixture models”. In: Journal of Computational Biology 13 (July 2006), pp. 1101–13. doi: 10.1089/cmb.2006.13.1101

  7. [6]

    The tree of blobs of a species network: identifiability under the coalescent

    Elizabeth S Allman et al. “The tree of blobs of a species network: identifiability under the coalescent”. In: Journal of Mathematical Biology 86.1 (2023), p. 10. 37

  8. [7]

    Split scores: a tool to quantify phylogenetic signal in genome-scale data

    Elizabeth S. Allman, Laura S. Kubatko, and John A. Rhodes. “Split scores: a tool to quantify phylogenetic signal in genome-scale data”. In: Systematic Biology 66.4 (Jan. 2017), pp. 620–636. issn: 1063-5157. doi: 10.1093/sysbio/syw103 . eprint: https: //academic.oup.com/sysbio/article-pdf/66/4/620/25423838/syw103.pdf . url: https://doi.org/10.1093/sysbio/syw103

Show all 61 references
  1. [8]

    Phylogenetic ideals and varieties for the general Markov model

    Elizabeth S. Allman and John A. Rhodes. “Phylogenetic ideals and varieties for the general Markov model”. In:Advances in Applied Mathematics 40 (2 Feb. 2008), pp. 127–

  2. [9]

    Identifying species network features from gene tree quartets under the coalescent model

    Hector Ba˜ nos. “Identifying species network features from gene tree quartets under the coalescent model”. In: Bulletin of Mathematical Biology 81 (2019), pp. 494–534

  3. [10]

    Code repository for ”Methodological considerations for semi- algebraic hypothesis testing with incomplete U-statistics”

    David Barnhill et al. Code repository for ”Methodological considerations for semi- algebraic hypothesis testing with incomplete U-statistics” . Version 1.0.0. July 2025. url: https : / / github . com / marinagarrote / Semialg - Hypothesis - Test - with - Incomplete-U-Stats

  4. [11]

    Detectability of varied hybridization scenarios using genome- scale hybrid detection methods

    Marianne B. Bjorner et al. “Detectability of varied hybridization scenarios using genome- scale hybrid detection methods”. In: Bulletin of the Society of Systematic Biologists 3.1 (Oct. 2024). doi: 10.18061/bssb.v3i1.9284

  5. [12]

    Some properties of incomplete U-statistics

    Gunnar Blom. “Some properties of incomplete U-statistics”. In: Biometrika (1976), pp. 573–580

  6. [13]

    Colored Gaussian DAG models

    Tobias Boege et al. “Colored Gaussian DAG models”. In: arXiv preprint arXiv:2404.04024 (2024)

  7. [14]

    Reduced U-statistics and the Hodges-Lehmann estima- tor

    BM Brown and DG Kildea. “Reduced U-statistics and the Hodges-Lehmann estima- tor”. In: The Annals of Statistics (1978), pp. 828–835

  8. [15]

    Causal Discovery with Latent Confounders Based on Higher-Order Cumulants

    Ruichu Cai et al. “Causal Discovery with Latent Confounders Based on Higher-Order Cumulants”. In: Proceedings of the 40th International Conference on Machine Learning. ICML’23. Honolulu, Hawaii, USA: JMLR.org, 2023

  9. [16]

    Performance of a new invariants method on homogeneous and nonhomogeneous quartet trees

    M Casanellas and J Fern´ andez-S´ anchez. “Performance of a new invariants method on homogeneous and nonhomogeneous quartet trees”. In: Molecular Biology and Evolution 24.1 (Oct. 2006), pp. 288–293. issn: 0737-4038. doi: 10.1093/molbev/msl153. eprint: https://academic.oup.com/...

  10. [17]

    SAQ: Semi- algebraic quartet reconstruction

    Marta Casanellas, Jesus Fernandez-Sanchez, and Marina Garrote-Lopez. “SAQ: Semi- algebraic quartet reconstruction”. In: IEEE/ACM Transactions on Computational Bi- ology and Bioinformatics 18 (6 Nov. 2021), pp. 2855–2861. issn: 1545-5963. doi: 10. 1109/TCBB.2021.3101278

  11. [18]

    Geometry of the Kimura 3-parameter model

    Marta Casanellas and Jes´ us Fern´ andez-S´ anchez. “Geometry of the Kimura 3-parameter model”. In: Advances in Applied Mathematics 41.3 (2008), pp. 265–292

  12. [19]

    Distance to the stochastic part of phylogenetic varieties

    Marta Casanellas, Jes´ us Fern´ andez-S´ anchez, and Marina Garrote-L´ opez. “Distance to the stochastic part of phylogenetic varieties”. In: Journal of Symbolic Computation 104 (May 2021), pp. 653–682. issn: 07477171. doi: 10.1016/j.jsc.2020.09.003

  13. [20]

    Phylogenetic mixtures and linear invariants for equal input models

    Marta Casanellas and Mike Steel. “Phylogenetic mixtures and linear invariants for equal input models”. In: Journal of Mathematical Biology 74 (5 Apr. 2017), pp. 1107–

  14. [21]

    Designing weights for quartet-based methods when data are heterogeneous across lineages

    Marta Casanellas et al. “Designing weights for quartet-based methods when data are heterogeneous across lineages”. In: Bulletin of Mathematical Biology 85 (7 July 2023), p. 68. issn: 0092-8240. doi: 10.1007/s11538-023-01167-y

  15. [22]

    Invariants of phylogenies in a simple case with discrete states

    James A. Cavender and Joseph Felsenstein. “Invariants of phylogenies in a simple case with discrete states”. In: Journal of Classification 4.1 (1987), pp. 57–71. doi: 10.1007/BF01890075. url: https://doi.org/10.1007/BF01890075

  16. [23]

    Gaussian and bootstrap approximations for high-dimensional U-statistics and their applications

    Xiaohui Chen. “Gaussian and bootstrap approximations for high-dimensional U-statistics and their applications”. In: The Annals of Statistics 46.2 (2018)

  17. [24]

    Randomized incomplete U -statistics in high dimen- sions

    Xiaohui Chen and Kengo Kato. “Randomized incomplete U -statistics in high dimen- sions”. In: The Annals of Statistics 47.6 (2019), pp. 3127–3156

  18. [25]

    Quartet inference from SNP data under the coa- lescent model

    Julia Chifman and Laura Kubatko. “Quartet inference from SNP data under the coa- lescent model”. In: Bioinformatics 30 (23 Dec. 2014), pp. 3317–3324. issn: 1367-4811. doi: 10.1093/bioinformatics/btu530

  19. [26]

    Algebraic Statistics and Contingency Table Problems: Log-Linear Models, Likelihood Estimation, and Disclosure Limitation

    Adrian Dobra et al. “Algebraic Statistics and Contingency Table Problems: Log-Linear Models, Likelihood Estimation, and Disclosure Limitation”. In: Emerging Applications of Algebraic Geometry . Ed. by Mihai Putinar and Seth Sullivant. New York, NY: Springer New York, 2009, pp....

  20. [27]

    On the ideals of equivariant tree models

    Jan Draisma and Jochen Kuttler. “On the ideals of equivariant tree models”. In: Math- ematische Annalen 344.3 (Dec. 2008), pp. 619–644. issn: 1432-1807. doi: 10.1007/ s00208-008-0320-6 . url: http://dx.doi.org/10.1007/s00208-008-0320-6

  21. [28]

    Likelihood ratio tests and singularities

    Mathias Drton. “Likelihood ratio tests and singularities”. In: Ann. Statist. 37.1 (2009), pp. 979–1012

  22. [29]

    Model selection and local geometry

    Robin J Evans. “Model selection and local geometry”. In: The Annals of Statistics 48.6 (2020), pp. 3513–3544

  23. [30]

    Cases in which parsimony or compatibility methods Will be pos- itively misleading

    Joseph Felsenstein. “Cases in which parsimony or compatibility methods Will be pos- itively misleading”. In: Systematic Zoology 27 (4 Dec. 1978), p. 401. issn: 00397989. doi: 10.2307/2412923

  24. [31]

    Invariant versus classical quartet in- ference when evolution is heterogeneous across sites and lineages

    Jes´ us Fern´ andez-S´ anchez and Marta Casanellas. “Invariant versus classical quartet in- ference when evolution is heterogeneous across sites and lineages”. In: Systematic Bi- ology 65 (2 Mar. 2016), pp. 280–291. issn: 1063-5157. doi: 10.1093/sysbio/syv086

  25. [32]

    Grayson and Michael E

    Daniel R. Grayson and Michael E. Stillman. Macaulay2, a software system for research in algebraic geometry . Available at http://www2.macaulay2.com

  26. [33]

    PARTIAL IDENTIFIABILITY OF RESTRICTED LA- TENT CLASS MODELS

    Yuqi Gu and Gongjun Xu. “PARTIAL IDENTIFIABILITY OF RESTRICTED LA- TENT CLASS MODELS”. In: The Annals of Statistics 48.4 (2020), pp. 2082–2107. issn: 00905364, 21688966. url: https://www.jstor.org/stable/26931550 (visited on 06/15/2025)

  27. [34]

    A framework for the quantitative study of evo- lutionary trees

    Michael D. Hendy and David Penny. “A framework for the quantitative study of evo- lutionary trees”. In: Systematic Zoology 38 (4 Dec. 1989), p. 297. issn: 00397989. doi: 10.2307/2992396

  28. [35]

    A maximum likelihood estimator for quartets under the Cavender-Farris-Neyman model

    Max Hill and Jose Israel Rodriguez. “A maximum likelihood estimator for quartets under the Cavender-Farris-Neyman model”. In: ACM Communications in Computer Algebra 58 (2 June 2024), pp. 35–38. issn: 1932-2232. doi: 10.1145/3712023.3712028

  29. [36]

    Performance of phylogenetic methods in simulation

    John P. Huelsenbeck. “Performance of phylogenetic methods in simulation”. In: Sys- tematic Biology 44.1 (Mar. 1995), pp. 17–48. issn: 1063-5157. doi: 10.1093/sysbio/ 39 44 . 1 . 17. eprint: https : / / academic . oup . com / sysbio / article - pdf / 44 / 1 / 17 / 19501493/44-1...

  30. [37]

    The asymptotic distributions of incomplete U-statistics

    Svante Janson. “The asymptotic distributions of incomplete U-statistics”. In: Zeitschrift f¨ ur Wahrscheinlichkeitstheorie und Verwandte Gebiete66.4 (1984), pp. 495–505

  31. [38]

    Maximum Likelihood Estimation of Symmetric Group-Based Models via Numerical Algebraic Geometry

    Dimitra Kosta and Kaie Kubjas. “Maximum Likelihood Estimation of Symmetric Group-Based Models via Numerical Algebraic Geometry”. In: Bulletin of Mathematical Biology 81.2 (Oct. 2018), pp. 337–360. issn: 1522-9602. doi: 10.1007/s11538-018- 0523-2. url: http://dx.doi.org/10.1007...

  32. [39]

    A rate-independent technique for analysis of nucleic acid sequences: evolutionary parsimony

    James A Lake. “A rate-independent technique for analysis of nucleic acid sequences: evolutionary parsimony.” In: Molecular biology and evolution 4.2 (1987), pp. 167–191

  33. [40]

    Lauritzen

    Steffen L. Lauritzen. Graphical Model. Oxford University Press, 1996

  34. [41]

    Fourier transform inequalities for phylogenetic trees

    Frederick A Matsen. “Fourier transform inequalities for phylogenetic trees”. In:IEEE/ACM transactions on computational biology and bioinformatics 6.1 (2008), pp. 89–95

  35. [42]

    Detecting hybrid speciation in the presence of incomplete lineage sorting using gene tree incongruence: a model

    C. Meng and L.S. Kubatko. “Detecting hybrid speciation in the presence of incomplete lineage sorting using gene tree incongruence: a model”. In: Theoretical Population Bi- ology 75.1 (2009), pp. 35–45. issn: 00405809. doi: 10.1016/j.tpb.2008.10.004

  36. [43]

    Hypothesis testing near singularities and boundaries

    Jonathan D Mitchell, Elizabeth S Allman, and John A Rhodes. “Hypothesis testing near singularities and boundaries”. In: Electronic Journal of Statistics 13.1 (2019), p. 2150

  37. [44]

    Relationships between gene trees and species trees

    P. Pamilo and M. Nei. “Relationships between gene trees and species trees.” In: Mol. Biol. Evol. 5.5 (1988), pp. 568–583

  38. [45]

    Maximum Likelihood Inference of Small Trees in the Presence of Long Branches

    Sarah L. Parks and Nick Goldman. “Maximum Likelihood Inference of Small Trees in the Presence of Long Branches”. In: Systematic Biology 63.5 (July 2014), pp. 798–

  39. [46]

    R: A Language and Environment for Statistical Computing

    R Core Team. R: A Language and Environment for Statistical Computing . R Foun- dation for Statistical Computing. Vienna, Austria, 2023. url: https : / / www . R - project.org/

  40. [47]

    MSCquartets 1.0: quartet methods for species trees and net- works under the multispecies coalescent model in R

    John A Rhodes et al. “MSCquartets 1.0: quartet methods for species trees and net- works under the multispecies coalescent model in R”. In: Bioinformatics 37.12 (Oct. 2020), pp. 1766–1768. issn: 1367-4803. doi: 10 . 1093 / bioinformatics / btaa868. eprint: https : / / academic ...

  41. [48]

    Invariant based quartet puzzling

    Joseph P Rusinko and Brian Hipp. “Invariant based quartet puzzling”. In: Algorithms for Molecular Biology 7 (2012), pp. 1–9. doi: 10.1186/1748-7188-7-35

  42. [49]

    Causal Discovery of Linear Non- Gaussian Causal Models with Unobserved Confounding

    Daniela Schkoda, Elina Robeva, and Mathias Drton. “Causal Discovery of Linear Non- Gaussian Causal Models with Unobserved Confounding”. In: arXiv:2408.04907 (2024)

  43. [50]

    Phylogenetics

    Charles Semple and Mike Steel. Phylogenetics. Vol. 24. Oxford University Press on Demand, 2003

  44. [51]

    Approximating high-dimensional infinite- order U -statistics: Statistical and computational guarantees

    Yanglei Song, Xiaohui Chen, and Kengo Kato. “Approximating high-dimensional infinite- order U -statistics: Statistical and computational guarantees”. In: Electronic Journal of Statistics 13.2 (2019), pp. 4794–4848. 40

  45. [52]

    TestGGM: Testing Gaussian Graphical Models

    Nils Sturma. TestGGM: Testing Gaussian Graphical Models . R package version 1.0

  46. [54]

    Toric ideals of phylogenetic invariants

    Bernd Sturmfels and Seth Sullivant. “Toric ideals of phylogenetic invariants”. In: Jour- nal of Computational Biology 12 (4 May 2005), pp. 457–481. issn: 1066-5277. doi: 10.1089/cmb.2005.12.457

  47. [55]

    Algebraic Statistics

    Seth Sullivant. Algebraic Statistics. Vol. 194. American Mathematical Soc., 2018

  48. [56]

    Long Branch Attraction Biases in Phylogenetics

    Edward Susko and Andrew J Roger. “Long Branch Attraction Biases in Phylogenetics”. In: Systematic Biology 70.4 (Feb. 2021), pp. 838–843. issn: 1063-5157. doi: 10.1093/ sysbio/syab001. eprint: https://academic.oup.com/sysbio/article-pdf/70/4/ 838/38663996/syab001.pdf. url: http...

  49. [57]

    High-dimensional causal discovery under Non- Gaussianity

    Y. Samuel Wang and Mathias Drton. “High-dimensional causal discovery under Non- Gaussianity”. In: Biometrika 107.1 (2019), pp. 41–59. eprint: https://academic.oup. com/biomet/article-pdf/107/1/41/32450889/asz055.pdf. Department of Mathematics, United States Naval Academy Depar...

  50. [148]

    doi: 10.1016/j.aam.2006.10.002

    issn: 01968858. doi: 10.1016/j.aam.2006.10.002

  51. [811]

    issn: 1063-5157. doi: 10 . 1093 / sysbio / syu044. eprint: https : / / academic . oup . com / sysbio / article - pdf / 63 / 5 / 798 / 24585598 / syu044 . pdf. url: https : //doi.org/10.1093/sysbio/syu044

  52. [1138]

    doi: 10.1007/s00285-016-1055-8

    issn: 0303-6812. doi: 10.1007/s00285-016-1055-8 . 38

  53. [2021]

    url: https://github.com/NilsSturma/TestGGM/blob/main/DESCRIPTION

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.