{"id":"2c35ab45-1552-45b3-ba17-a13d75160e62","arxiv_id":"1908.01424","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A new R package uses summary statistics to test multispecies coalescent simulators, exposing flaws in Mesquite and Hybrid-lambda while confirming SimPhy and Phybase.","lead":"The authors built statistical tests that check whether computer simulators of the multispecies coalescent process produce valid gene trees, and they found that two of four widely used simulators fail the checks. Because these simulators are used to validate new methods for inferring species trees, the discovery of errors affects the reliability of published phylogenomic studies.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1 row ((A,B),D) has expected counts inconsistent with the paper's own formula and with its p-values; qualitative conclusions survive, but the reported data need correction.","rationale":"I read the paper as a practical validation of two summary-statistic tests and an application to four simulators. The theoretical derivations in §4 are sound, and the pairwise-distance failures for Hybrid-lambda and Mesquite are visually and statistically compelling; the topology tests also strongly implicate Mesquite. My review found the central claim well supported. However, the printed expected triple counts in Table 1 for S3 ((A,B),D) are inconsistent with the paper's own formulas and with the p-values in the same table. Resolving this is necessary for the paper's quantitative results to be fully reproducible, though the qualitative conclusions survive. I therefore recommend conditional acceptance rather than a change in the science.","tokens_in":13159,"tokens_out":21186,"duration_ms":209999,"concrete_test":"Recompute the row ((A,B),D) of Table 1 from §4.2: for S3, x = 1000/2000 + 1000/3000 = 0.8333, so expected counts are 71028, 14486, 14486; then recompute the chi-square statistic and p-value for Phybase (71072, 14450, 14478) using both the printed and the corrected expectations. The reported p=0.940 should match only the corrected expectation. Independently run MSCsimtester's rootedTriple function on S3 and compare its output with the table.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"In Table 1 (and Table S3), the row for the rooted triple ((A,B),D) on species tree S3 lists Expected counts 70044, 14977, 14977 and internal branch 0.83̄3. The formula in §4.2 gives P(discordant) = (2/3)exp(-x), where x is the path length in coalescent units from MRCA(A,B) to MRCA(A,B,D); for S3 this path has edges of 1000 gen/N=2000 and 1000 gen/N=3000, so x=0.8333 and E(concordant)=71028, E(each discordant)=14486. The printed expectation is thus off by about 1%. More tellingly, the p-value reported for Phybase in that row (0.940) cannot be obtained from the printed expectation: χ²≈50 against (70044,14977,14977) gives p≈1e-11, while χ²≈0.12 against the corrected (71028,14486,14486) gives p≈0.94. The p-values for Hybrid-lambda and SimPhy in the same row similarly match only the corrected expectation. So the 'Expected' column for this row is wrong, although the p-values themselves appear to be computed with the correct formula. This is a concrete internal inconsistency in one of the two headline summary-statistic tests, and it should be corrected and the MSCsimtester output re-verified. It does not alter the qualitative verdicts: Mesquite fails catastrophically under either expectation, and the other simulators still pass.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops statistical tools to test whether gene tree samples produced by multispecies coalescent (MSC) simulators conform to the MSC model. Two summary statistics are used: the distribution of pairwise distances between taxa on gene trees, and the frequencies of rooted triple topologies. Theoretical distributions for both statistics are derived from first principles in Sections 4.1 and 4.2, and the tools are implemented in an R package called MSCsimtester. The tests are applied to 100,000 gene tree samples from four published simulators. The authors conclude that SimPhy and correctly parameterized Phybase produce samples consistent with the MSC, Hybrid-lambda fails the metric tests but passes the topological tests, and Mesquite fails both metric and topological tests. The paper explicitly acknowledges that the chosen summary statistics cannot give an ironclad guarantee of correctness but argues they are likely to uncover most problems.","tokens_in":13433,"tokens_out":11901,"duration_ms":104622,"significance":"If the results hold, the paper provides a practical validation toolkit for a widely used class of phylogenomic simulation software. The theoretical derivations in Sections 4.1 and 4.2 are transparent, self-contained, and correct, and the statistical procedures are applied sensibly, including subsampling to avoid over-rejection on very large samples. The finding that two popular simulators (Mesquite and, in the metric sense, Hybrid-lambda) produce invalid MSC samples is an important community service. The authors also provide an R package, which strengthens the paper's utility. The main limitation, that the two selected summary statistics may not detect all possible deviations, is explicitly acknowledged by the authors and is reasonable for a practical testing framework.","major_comments":[{"comment":"The row for rooted triple ((A,B),D) in Table 1 (and Table S3) reports Expected counts 70044, 14977, 14977, but the formula in §4.2 gives P(discordant) = (1/3)exp(-x) with x = 1000/2000 + 1000/3000 = 0.8333, yielding expected counts 71028, 14486, 14486. The printed expectation is inconsistent with the paper's own formula and does not sum to 100,000 (it sums to 99,998). The p-values in that row (e.g., Phybase 0.940) are consistent with the corrected expectation, so the error appears to be in the printed Expected column rather than in the computation. Please correct the table and re-verify that all rows' p-values match the reported expectations.","section":"Table 1 / §4.2"},{"comment":"The SimPhy counts in the ((A,B),C) row of Table 1 are 59120, 20397, 20483, but the reported p-value 0.554 is inconsistent with these counts under the stated expectation (59564, 20217, 20217); the chi-squared statistic is approximately 8.41, which with 2 degrees of freedom yields p ≈ 0.015. Moreover, the SimPhy counts in Table S3 for the same species tree and rooted triple are 59764, 20091, 20145, and these counts likewise do not yield the reported p-value 0.504 (they give p ≈ 0.425). The SimPhy rows for ((A,C),D) and ((B,C),D) also differ between Table 1 and Table S3. These discrepancies suggest typographical errors or results from different simulation runs; the authors must reconcile all tables and ensure the p-values correspond to the reported count vectors.","section":"Table 1 and Table S3, SimPhy row for ((A,B),C)"}],"minor_comments":[{"comment":"There is a typo in the text: 'SymPhy' should be 'SimPhy'.","section":"Section 3"},{"comment":"The word 'exchangability' should be 'exchangeability'.","section":"Section 4.2"},{"comment":"The notation used in the species trees, e.g., '#2000', is not defined in the main text; please add a sentence explaining that it denotes the population size assigned to the preceding edge.","section":"Supplementary methods"},{"comment":"The caveat that the two summary statistics may not detect all simulator errors is important and is stated appropriately; consider adding a short paragraph in the Discussion on the sensitivity of the tests and how users might combine them with additional statistics.","section":"Section 1"},{"comment":"The rooted triple p-values are computed from a single sample of 100,000 gene trees, with the authors noting that this gives preliminary results. It would be helpful to report the variability of p-values across repeated simulations, as is done for the Anderson-Darling test in Figure 5.","section":"Section 3 / Table 1"}],"recommendation":"major_revision","confidential_remarks":"The two table inconsistencies described in the major comments are the main obstacles to publication. The qualitative conclusions appear robust—Mesquite clearly fails and the other simulators pass—so I expect a carefully corrected revision to be publishable. I would also encourage the authors to deposit the simulation outputs or a script that reproduces Table 1, because the discrepancies I found could not be resolved without access to the raw data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time. Allman, Baños, and Rhodes have written a practical paper: they derive two summary-statistic distributions for gene trees under the multispecies coalescent (pairwise distances, rooted triple frequencies) and use them to test four simulators. The headline result is that Mesquite fails on both metrics and Hybrid-lambda fails on metric distances; Phybase and SimPhy pass. This matters because simulation studies are only as good as the simulator, and these two packages have been used a lot.\n\nThe derivations in Section 4 are standard coalescent theory, correctly applied. The statistical machinery (Anderson-Darling with subsampling to avoid over-rejection, chi-square tests with 2 df) is appropriate. The authors are honest about the initial confusion with Phybase's input format, and they are appropriately cautious in claiming that no finite set of summary statistics gives an ironclad guarantee. That caveat is minor because the tests found real problems.\n\nI checked the stress-test note on Table 1, and it is correct. For species tree S3, the (A,B,D) triple should have expected counts 71028 / 14486 / 14486, not the printed 70044 / 14977 / 14977. The p-values in that row (0.940, 0.207, 0.934) agree with the corrected expectation, so the error is in the printed 'Expected' column, not in the p-value computation. It looks like a rounding/transcription slip, possibly from using x=0.8 instead of x=0.8333. This is a clear typo in a headline table and should be fixed before publication, with the MSCsimtester output re-verified. The qualitative conclusions do not change: Mesquite is still catastrophically off, and the other simulators still pass.\n\nA second soft spot is that the R package is announced but not included as an archive or thoroughly documented in the preprint. The code would be nice to see, but again it doesn't affect the mathematical results.\n\nOverall this is a solid, citable contribution. I would send it to a competent referee with a request for minor revision. The table fix and a pointer to the code would make it clean.","headline":"A practical, mostly correct MSC simulator testing paper with a real (but non-fatal) typo in its headline Table 1; the qualitative findings about Mesquite and Hybrid-lambda hold up.","tokens_in":13962,"tokens_out":3694,"would_cite":true,"duration_ms":31657,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A simple pair of summary statistics can certify which gene-tree simulators sample the multispecies coalescent correctly — and two popular ones fail.","keywords":["multispecies coalescent","gene tree simulation","summary statistics","pairwise distance distribution","rooted triple frequencies","simulator validation","Anderson-Darling test","MSCsimtester"],"falsifier":"Take a clearly incorrect simulator that preserves the marginal pairwise-distance distribution and rooted triple frequencies of the MSC but breaks other features, for example by drawing each gene tree's topology and coalescence times independently rather than jointly. If the MSCsimtester tests return uniformly distributed p-values on its output, that would refute the paper's assertion that these statistics are likely to uncover most problems; if the tests reject, the assertion is supported.","tokens_in":12941,"feed_emoji":"🧬","tokens_out":8102,"duration_ms":70643,"temperature":0.7,"pith_summary":"Simulators of the multispecies coalescent are widely used to generate test datasets for species-tree inference, but a simulator can be subtly wrong in ways users do not notice. This paper establishes that two analytically tractable summaries of a gene-tree sample — the distribution of pairwise distances between taxa and the frequencies of rooted triple topologies — can be compared against exact theoretical values to judge whether a sample truly comes from the MSC. Applying these tests to four published simulators, the paper finds that SimPhy and correctly configured Phybase pass, while Hybrid-lambda produces incorrect metric gene trees and Mesquite fails on both topological and metric properties. The checks are packaged as an R package so that developers and users can routinely test simulators. If this holds, many published simulation-based evaluations built on the flawed simulators would need critical re-reading.","feed_headline":"Two gene-tree simulators fail new coalescent checks","feed_subtitle":"Simple summary statistics catch wrong samples that users might otherwise trust.","key_machinery":"The load-bearing objects are two closed-form theoretical distributions. First, the pairwise distance density: if two lineages enter the same population at node $v$ and then traverse edges $e_1,\\dots,e_k$ to the root, with population-size functions $N_{e_i}(t)$ and escape probabilities $\\eta_i = \\exp\\left(-\\int_0^{\\ell_i} 1/N_{e_i}(\\tau)\\,d\\tau\\right)$, the density of the time to coalescence is a piecewise (shifted, scaled) exponential with discontinuities at population boundaries. Second, the rooted triple probabilities: for three taxa whose two shallowest lineages enter the shared population with internal branch length $x$ in coalescent units, $P(((a,b),c)) = 1 - (2/3)e^{-x}$ and the two discordant topologies each have probability $(1/3)e^{-x}$. These give exact expected histograms and counts that a simulator sample can be tested against.","core_discovery":"The paper's central discovery is that the MSC imposes exact, computable distributions on two summary statistics of a gene-tree sample: for any pair of taxa, the coalescence times follow a piecewise exponential density whose pieces are determined by the population sizes along the path from the pair's most recent common ancestor to the root; and for any three taxa, the three rooted gene-tree topologies occur with frequencies $P(((a,b),c)) = 1 - (2/3)e^{-x}$ and $P(((a,c),b)) = P(((b,c),a)) = (1/3)e^{-x}$, where $x$ is the internal branch length in coalescent units. By comparing a simulator's output on these two statistics to the theory — using an Anderson-Darling test for the distance distribution and a chi-squared test for the triple counts — the paper shows that valid and invalid simulators separate cleanly. In particular, among four published simulators, SimPhy and correctly parameterized Phybase produce samples in accord with the MSC, Hybrid-$\\lambda$ samples correctly in topology but wrong in metric gene trees, and Mesquite fails on both. The authors argue these tools should be standard for simulator validation.","pith_inferences":["Inference: The same two-statistic approach could be extended to the multispecies network coalescent, where rooted triple frequencies and pairwise distance distributions have analogous closed forms, allowing detection of errors in hybridization simulators.","Inference: Because the tests check marginal distributions only, a simulator that draws each gene tree's topology and coalescence times independently—while preserving the correct marginals—would pass; testing the joint distribution (for example, via the covariance of pairwise distances) would be the next strengthening.","Inference: Re-analysis of simulation studies that relied on Mesquite or Hybrid-lambda as ground truth could reveal which published performance comparisons of species-tree inference methods are affected."],"forward_implications":["Simulation studies that used Mesquite or Hybrid-lambda to generate MSC gene trees may have drawn invalid samples, so their conclusions should be re-examined cautiously.","The R package MSCsimtester gives developers a way to detect errors in new simulators and users a way to verify that input parameters are interpreted correctly.","Correctly configuring Phybase requires supplying branch lengths as $\\mu t$ and population sizes as $\\theta = 4\\mu N$; with that, its samples pass the tests.","Because the tests rely on large samples and subsampling to avoid misleadingly small p-values, they provide a practical standard for routine simulator validation."],"supporting_citations":[{"why":"Describes SimPhy, the simulator whose samples pass all tests and serve as the positive control.","marker":"[Mallo et al., 2016]"},{"why":"Describes Phybase, the R package whose samples pass only when its parameters are set as $\\mu t$ and $\\theta = 4\\mu N$.","marker":"[Liu and Yu, 2010]"},{"why":"Describes Hybrid-lambda, the simulator the tests show samples topologies correctly but metric gene trees incorrectly.","marker":"[Zhu et al., 2015]"},{"why":"Describes Mesquite, whose samples the tests show match neither topological nor metric MSC predictions.","marker":"[Maddison and Maddison, 2018]"},{"why":"Supplies the rooted triple probability formula that the topology test compares against.","marker":"[Pamilo and Nei, 1988]"},{"why":"Provides the goodness-of-fit test the paper applies to the pairwise distance distributions.","marker":"[Anderson and Darling, 1952]"},{"why":"Gives the piecewise exponential density for pairwise coalescence times on a species tree, the theoretical curve used in the metric test.","marker":"[Allman et al., 2019]"}],"fun_headline_variants":["Two simulators flunk new coalescent checks","Summary stats reveal flawed gene-tree simulators","New tests catch bad MSC simulator output","Coalescent simulators get a statistical reality check","SimPhy passes, Mesquite fails new MSC tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The tests assume the two chosen summary statistics—pairwise distances and rooted triple frequencies—are sensitive enough to reveal any meaningful departure from the MSC, so a simulator error that leaves these two margins unchanged would go undetected.","fun_headline_variants_meta":{"raw":{"variants":["Two simulators flunk new coalescent checks","Summary stats reveal flawed gene-tree simulators","New tests catch bad MSC simulator output","Coalescent simulators get a statistical reality check","SimPhy passes, Mesquite fails new MSC tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000363,"raw_usage":{"total_tokens":1923,"prompt_tokens":874,"completion_tokens":1049,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":490,"completion_tokens_details":{"reasoning_tokens":978}},"tokens_in":490,"tokens_out":1049,"duration_ms":11447,"temperature":1.0,"reasoning_tokens":978,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:13:25.581746+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a clearly incorrect simulator that preserves the marginal pairwise-distance distribution and rooted triple frequencies of the MSC but breaks other features, for example by drawing each gene tree's topology and coalescence times independently rather than jointly. If the MSCsimtester tests return uniformly distributed p-values on its output, that would refute the paper's assertion that these statistics are likely to uncover most problems; if the tests reject, the assertion is supported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes SimPhy, the simulator whose samples pass all tests and serve as the positive control."},{"cited_title":"and Yu, L","cited_arxiv_id":null,"evidence_quote":"Describes Phybase, the R package whose samples pass only when its parameters are set as $\\mu t$ and $\\theta = 4\\mu N$."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes Hybrid-lambda, the simulator the tests show samples topologies correctly but metric gene trees incorrectly."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes Mesquite, whose samples the tests show match neither topological nor metric MSC predictions."},{"cited_title":"and Nei, M","cited_arxiv_id":null,"evidence_quote":"Supplies the rooted triple probability formula that the topology test compares against."},{"cited_title":"goodness-of-fit","cited_arxiv_id":null,"evidence_quote":"Provides the goodness-of-fit test the paper applies to the pairwise distance distributions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the piecewise exponential density for pairwise coalescence times on a species tree, the theoretical curve used in the metric test."}],"review_version":1}