{"id":"a5d7f2de-7648-4e97-8f49-64f95b21a3f3","arxiv_id":"2412.06042","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"The authors implement and compare three infinite mixture models for across-site evolutionary variation in BEAST X, finding that the best model type differs across three viral data sets.","lead":"This paper extends Bayesian nonparametric mixture models for molecular evolution, adding hierarchical Dirichlet process and infinite hidden Markov model variants to the BEAST X package, and tests them on three viral data sets. A smart generalist might read it because the choice of model for variation across sequence sites can change inferred evolutionary trees and rates, and the paper suggests the best model depends on the data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Model rankings rest on single stabilized-harmonic-mean estimates with no error bars and reported MCMC mixing problems; the 2-4 log-unit margins that choose the 'best' infinite mixture prior are within plausible estimator error.","rationale":"I read the paper as an applied methodological study whose central empirical claim is that infinite mixture models, and different nonparametric priors in different settings, improve model fit over standard across-site models. The implementation in BEAST X, the availability of XML files, and the posterior tree and partition summaries are genuine contributions. The weakest link is the numerical evidence: a single stabilized-harmonic-mean marginal likelihood per model, with the authors themselves noting that better estimators exist and that MCMC convergence and mixing were inconsistent across replicates. This is exactly the reader's weakest assumption. I do not see an independent internal inconsistency; the fixed-clock parameterization difference between infinite and standard models is a modeling choice that a sensitivity analysis could address, but it is not a logical flaw. Because the reader's conditional verdict already requires stronger model-comparison evidence, my assessment does not move the verdict. The requested change is therefore no change: keep the paper conditional on re-estimation or error-quantified replication of the marginal likelihoods that determine the close rankings.","tokens_in":28658,"tokens_out":8699,"duration_ms":97506,"concrete_test":"Re-estimate Table A1 for the six GTR analyses on RSV A and HCV (DP, HDP-Codon, IHMM per data set) from the supplied BEAST XMLs using 10 independent MCMC chains per model, and compute both the stabilized harmonic mean per chain and, where implementable, a path-sampling or generalized stepping-stone estimate on the collapsed mixture likelihood. Report mean, standard deviation, and the sign of every pairwise difference with claimed margins under 10 log units. If any such sign changes across chains or estimators, the qualitative 'best prior varies' conclusion is unsupported; if all small-margin orderings are stable, the measurement concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table A1 reports one log marginal likelihood per model, computed with the stabilized harmonic mean estimator described in Section 3. The Discussion explicitly concedes that path sampling, stepping-stone, and generalized stepping-stone outperform this estimator, and that the authors 'observed inconsistent performance in MCMC convergence and mixing across replicates.' The paper's ranking claims are load-bearing. Most margins are large, but the qualitative conclusion that 'different types of infinite mixture models emerge as the best choices in different scenarios' depends on small margins: on RSV A with GTR, HDP-Codon beats IHMM by 4 log units and IHMM beats DP by 3; on HCV with GTR, DP beats IHMM by 2 log units. A harmonic mean estimator with known bias and high variance, applied to chains with acknowledged mixing failures and reported as a single point estimate, can plausibly flip those signs. The enormous IHMM margins on RABV may be real, but they do not validate the close comparisons that establish which prior wins. Without a more reliable estimator or replication-based uncertainty, the strongest claim is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops Bayesian nonparametric infinite mixture models for across-site evolutionary variation, extending Dirichlet process mixtures to hierarchical Dirichlet processes with codon-position groups (HDP-Codon) and infinite hidden Markov models (IHMM). The models are implemented in BEAST X, using a data-squashing MCMC strategy and BEAGLE parallelization. The authors compare the fit of these models with standard approaches (No Variation, Gamma, Codon, Codon+Gamma) under HKY and GTR substitution models on three viral data sets (RSV A, HCV subtype 4, RABV), using log marginal likelihood estimates obtained from a stabilized harmonic mean estimator. They report that infinite mixture models generally outperform standard models, that IHMM and HDP can substantially outperform DP models, and that the best-performing infinite mixture prior varies across data sets and substitution models. They also compare posterior phylogenetic trees, split frequencies, and treespace heatmaps. The main quantitative conclusions rest on a single set of marginal likelihood estimates without reported Monte Carlo uncertainty, and the authors explicitly acknowledge limitations of the estimator and MCMC mixing.","tokens_in":28902,"tokens_out":5533,"duration_ms":54590,"significance":"If the empirical rankings are reliable, the paper makes a useful contribution: it introduces two underused Bayesian nonparametric priors for site-heterogeneity modeling, provides a working implementation in the widely used BEAST X platform, combines data-squashing MCMC with BEAGLE parallel likelihood evaluation, and makes XML input files publicly available. The comparison across three viral data sets and several standard models is a sensible first evaluation, and the appendix on posterior tree differences adds practical information. However, the central claim that 'different types of infinite mixture models emerge as the best choices in different scenarios' is only as strong as the marginal likelihood estimates, and the current estimator plus acknowledged mixing problems leave the quantitative rankings vulnerable. The significance is therefore conditional on strengthening the model-comparison evidence; the methodological framework itself is a solid basis for a revised manuscript.","major_comments":[{"comment":"The paper's central ranking claims rest on single stabilized harmonic mean estimates of log marginal likelihood, with no replication or uncertainty quantification. The Discussion explicitly concedes that path sampling, stepping-stone, and generalized stepping-stone outperform this estimator, and that 'we have observed inconsistent performance in MCMC convergence and mixing across replicates.' Many of the qualitative conclusions depend on small margins: for RSV A with GTR, HDP-Codon beats IHMM by 4 log units and IHMM beats DP by 3 log units; for HCV with GTR, DP beats IHMM by 2 log units (Table A1). The stabilized harmonic mean is known to be biased and high-variance, especially under poor mixing, so these margins may not be robust. The authors should either provide replicated marginal likelihood estimates with standard errors, implement a more reliable estimator for HDP/IHMM, or demonstrate that the within-model Monte Carlo error is much smaller than the reported differences. Without this, the claim that different infinite mixture priors win in different scenarios is not established.","section":"Section 3, Table A1, and Section 4"},{"comment":"The authors set the HKY transition/transversion ratio base distribution hyperparameters using estimates from HKY + No Variation analyses of the same data: 'we take an empirical Bayes approach and adopt the estimated mean and ten times the estimated standard deviation of transition/transversion rate estimates from HKY + No Variation models.' This uses the data twice and is asymmetric: HKY infinite mixture models receive data-informed priors, while GTR infinite mixture models and standard HKY models do not. Since the paper's conclusion that HKY-based mixtures outperform GTR-based mixtures is central to its 'best model varies' narrative, this prior specification may systematically favor HKY mixtures. The authors should either infer these hyperparameters jointly, use a fully hierarchical prior, or at a minimum assess sensitivity of the rankings to this empirical Bayes choice.","section":"Section 3, empirical Bayes HKY base distribution"},{"comment":"The paper provides no convergence diagnostics, effective sample sizes, or replicate-level summaries for the MCMC runs, despite relying on an approximate data-squashing sampler and acknowledging inconsistent mixing. This matters not only for marginal likelihoods but also for the posterior quantities summarized in Table 1 and Figures A16-A20, including posterior medians, credibility intervals, and site-specific posterior means. The authors should report convergence diagnostics (e.g., ESS, Gelman-Rubin across replicates) for at least the main analyses, or state explicitly which results come from unreplicated single chains and how the acknowledged mixing problems affect those estimates.","section":"Section 2.4 and Section 4, MCMC convergence"}],"minor_comments":[{"comment":"The model is referred to as 'HKY + HDP' in the text but 'HKY + HDP-Codon' in Table A1 and elsewhere; please standardize the naming throughout.","section":"Section 3.1 and Table A1"},{"comment":"The claim of 'substantial differences' in phylogenetic posterior distributions should be reconciled with the reported split-frequency correlations of 0.97-1.00 and ASDSF values below 0.01; the heatmaps may show differences, but the split-frequency evidence suggests mostly similar posteriors, so the language should be tempered or the basis for 'substantial' should be clarified.","section":"Appendix A.3"},{"comment":"The notation in the data-squashing description, such as 'n(t)zj * P(Xj|...)' and the subscript on zj, is confusing; using a clearer index for the category (e.g., k) and defining the star would improve readability.","section":"Section 2.4"},{"comment":"There are several typographical errors in the reference list, including 'Jounal of Molecular Evolution', 'annd', 'orocess', 'Deparment', 'nucelotide', and 'Bayess'; a careful copyedit is needed.","section":"References"},{"comment":"In the RABV panel, the bar for HKY + IHMM is so much larger than the others that it compresses the visual differences among the remaining models; consider using a broken axis or a separate inset for the large values.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid methods contribution with a valuable implementation, but the central empirical claim currently rests on marginal likelihood estimates whose reliability is explicitly questioned by the authors themselves. The path to acceptance is clear: add uncertainty quantification or a more reliable estimator for the infinite mixture models, and address the empirical Bayes asymmetry in the HKY base distribution. I do not see grounds for rejection, but the paper is not yet ready without these revisions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is a real methods contribution, not just talk. The authors implement three Bayesian nonparametric priors—Dirichlet process, hierarchical Dirichlet process, and infinite hidden Markov model—for across-site evolutionary variation in BEAST X, with data squashing and BEAGLE parallelization, and they make the XML files public. That is concrete, reproducible work. The exposition of the three priors is clear and the appendix comparisons of tree posteriors are a nice addition.\n\nThe broad qualitative finding—infinite mixtures often beat standard fixed-partition and gamma models—is plausible and probably true. The margins are sometimes enormous, and the authors are honest that the best infinite mixture varies across data sets and substitution models.\n\nNow the soft spot, and it is load-bearing. All quantitative rankings come from a single stabilized harmonic mean marginal likelihood estimate per model. No error bars, no replication. The authors themselves note in the Discussion that path sampling, stepping-stone, and generalized stepping-stone outperform this estimator, and that they observed inconsistent MCMC convergence and mixing across replicates. That matters because some of the key comparisons determining which prior wins are close: on RSV A with GTR, HDP-Codon beats IHMM by 4 log units and IHMM beats DP by 3; on HCV with GTR, DP beats IHMM by 2. Those margins are inside the known noise level of a harmonic mean estimator. The huge IHMM margins on RABV might be real, but they do not rescue the small-margin claims. The conclusion that “different types of infinite mixtures emerge as best in different scenarios” depends at least partly on those small margins.\n\nThe mild empirical Bayes step—fixing HKY base-distribution hyperparameters from HKY+No Variation runs—is transparent and, in practice, a minor concern. It does not change my overall read.\n\nWho should read this: anyone working on across-site heterogeneity in phylogenetics, and anyone teaching model comparison pitfalls. I would not desk-reject it. It deserves peer review, but the referees should push for a more reliable marginal likelihood estimate on at least a subset of models, or replication-based uncertainty. Without that, the specific rankings are not established, even though the general message probably holds.","headline":"Genuine implementation and honest discussion, but the model rankings rest on a marginal likelihood estimator the authors themselves flag as unreliable, so the quantitative claims need validation.","tokens_in":29405,"tokens_out":3208,"would_cite":true,"duration_ms":35046,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","92D15"],"pacs":[],"model":"deepseek-v4-flash","headline":"Bayesian infinite mixture models—Dirichlet process mixtures, hierarchical Dirichlet processes, and infinite hidden Markov models—fit viral alignments better than standard models, and the best prior type varies by data set.","keywords":["infinite mixture models","Dirichlet process mixtures","hierarchical Dirichlet processes","infinite hidden Markov models","across-site evolutionary variation","Bayesian phylogenetics","marginal likelihood estimation","data squashing MCMC"],"falsifier":"Re-estimate the log marginal likelihood for every model on all three data sets with generalized stepping-stone or path sampling and check whether the rankings survive—specifically, whether HKY+IHMM keeps its roughly 100-unit lead over HKY+HDP on RSV A, HKY+DP keeps its lead on HCV, and HKY+IHMM keeps its more than 1000-unit lead on rabies; if the ordering changes materially, the paper's comparative claim fails.","tokens_in":28459,"feed_emoji":"🧬","tokens_out":14182,"duration_ms":121238,"temperature":0.7,"pith_summary":"The paper argues that Bayesian infinite mixture models—models in which the number of hidden evolutionary classes is unbounded and inferred from the data—should replace fixed a priori partitions when modeling how evolution varies across alignment sites. It goes beyond the standard Dirichlet process mixture by adding two Bayesian nonparametric priors: a hierarchical Dirichlet process that pools information across predefined groups of sites, and an infinite hidden Markov model that lets the evolutionary class of one site depend on the class of its neighbor. Analyzing respiratory syncytial virus, hepatitis C subtype 4, and rabies virus data, the paper finds that the best infinite mixture beats the best standard model in every scenario, and that with one exception every infinite mixture beats every standard model. The winning prior type differs by data set, so the practical claim is that all three priors belong in the modeling toolkit. The framework is implemented in a widely used Bayesian phylogenetic software package and scales through data-squashing MCMC and parallel likelihood evaluation.","feed_headline":"Infinite mixture models beat fixed partitions across viral data sets","feed_subtitle":"Bayesian priors infer how many evolutionary classes an alignment needs, and where each site belongs.","key_machinery":"The load-bearing objects are three Bayesian nonparametric priors on the site-specific evolutionary parameters $\\theta_i=(\\rho_i,Q_i)$—the overall substitution rate and the relative exchange-rate matrix for site $i$. A Dirichlet process (built from a stick-breaking representation or Chinese restaurant process) yields a discrete random measure whose atoms are the distinct evolutionary categories, making the number of categories $K$ a random variable with countably infinite support. A hierarchical Dirichlet process couples group-specific Dirichlet processes through a shared discrete base distribution, so groups such as codon positions share the same set of categories while assigning them different weights. An infinite hidden Markov model uses the same construction with an unbounded number of states and makes the category at site $i$ depend on the category at site $i-1$, capturing spatial correlation along the alignment. Posterior sampling couples a data-squashing MCMC scheme that jointly proposes category updates for blocks of similar sites with parallel likelihood evaluations across the $s\\times c$ site-category grid, and all models are compared by estimating the log marginal likelihood with the adjusted stabilized harmonic mean estimator.","core_discovery":"On its own terms, the paper's central discovery is that the three Bayesian nonparametric priors are not interchangeable: Dirichlet process mixtures, hierarchical Dirichlet processes, and infinite hidden Markov models each capture a different structure in the data, and each is the best-fitting choice in at least one of the six data-set and substitution-model combinations analyzed. The best infinite mixture outperforms the best standard model in all scenarios, and with the exception of the Dirichlet process paired with a GTR substitution model on rabies data, infinite mixtures always outperform standard models. The largest gains come from the infinite hidden Markov model, which beats the next-best model by roughly 100 log marginal likelihood units on respiratory syncytial virus data and by more than 1000 units on rabies data. These fit gains are not cosmetic: the appendix shows that different across-site variation models produce different maximum clade credibility trees, split frequencies, and posterior distributions over tree space. Under infinite mixtures, the simpler HKY substitution model often fits better than GTR, the reverse of what happens under standard models.","pith_inferences":["We infer that the infinite hidden Markov model's very large margins on rabies virus data provide indirect evidence for spatial autocorrelation in evolutionary parameters along real alignments, a mechanism that could be tested by comparing IHMM against a fixed codon-partition model with within-partition spatial ordering on additional protein-coding data sets.","We infer that the qualitative conclusion—infinite mixtures beat standard models—is more secure than the fine-grained ordering of HDP, IHMM, and DP, because the harmonic mean estimator the paper uses is one the authors themselves describe as less reliable than path sampling and stepping-stone methods.","We infer a natural extension: define HDP groups by gene or genomic region in whole-genome data, and couple site-level HDPs across tree branches to model branch-specific and site-specific variation jointly, both directions the paper lists as future work but does not test.","We infer that combining IHMM's spatial dependence with decoupled clustering of substitution rates versus exchange-rate parameters is a concrete model variant that could outperform any prior considered here."],"forward_implications":["Practicing phylogeneticists can replace fixed partitions and discretized-gamma rate models with infinite mixtures that infer the number of evolutionary categories and the category of each site from the data, removing a major a priori modeling choice.","Because the best-fitting prior changes with data set and substitution model, model comparison for across-site variation should include all three priors rather than defaulting to Dirichlet process mixtures.","The infinite hidden Markov model's large wins on some alignments imply that adjacent sites often evolve under correlated categories that exchangeable mixtures cannot capture.","Choice of across-site variation model changes posterior phylogenetic inferences, including maximum clade credibility trees, split frequencies, and treespace occupation.","Data-squashing MCMC and parallel likelihood evaluation make these analyses computationally feasible on alignments with thousands of sites."],"supporting_citations":[{"why":"Supplies the hierarchical Dirichlet process construction and the Chinese restaurant franchise sampler used for HDP and IHMM.","marker":"Teh et al. (2006)"},{"why":"Introduces the infinite hidden Markov model with a countably infinite state space that the paper adapts.","marker":"Beal et al. (2002)"},{"why":"Establishes Dirichlet process mixtures for across-site rate variation, the baseline infinite mixture the paper extends.","marker":"Huelsenbeck and Suchard (2007)"},{"why":"Provides an earlier DP mixture framework and the hepatitis C analysis setup with root-height prior that the paper adopts.","marker":"Wu et al. (2013)"},{"why":"Supplies the data-squashing MCMC strategy for jointly updating site-to-category assignments.","marker":"Guha (2010)"},{"why":"Provides the stabilized harmonic mean marginal likelihood estimator on which all model comparisons rest.","marker":"Newton and Raftery (1994)"},{"why":"Provides the adjusted version of the harmonic mean estimator that the paper actually uses.","marker":"Redelings and Suchard (2005)"},{"why":"Supplies the respiratory syncytial virus subgroup A data set used in one of the three empirical comparisons.","marker":"Zlateva et al. (2005)"},{"why":"Supplies the hepatitis C subtype 4 data set used in one of the three empirical comparisons.","marker":"Ray et al. (2000)"},{"why":"Supplies the rabies virus data set used in one of the three empirical comparisons.","marker":"Biek et al. (2007)"}],"fun_headline_variants":["Infinite mixtures beat fixed partitions in viral phylogenetics","Infinite hidden Markov models yield biggest phylogenetic fit gains","No single prior best for across-site evolutionary variation","Bayesian nonparametric priors: each wins on different data","Infinite mixtures adapt to each dataset: no one-size-fits-all prior"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the stabilized harmonic mean estimator gives comparably accurate log marginal likelihoods across all models compared; the paper itself notes that path sampling and stepping-stone estimators are more accurate and that MCMC convergence and mixing were inconsistent across replicates, so if harmonic mean estimates are biased differently by model, the claimed rankings could change.","fun_headline_variants_meta":{"raw":{"variants":["Infinite mixtures beat fixed partitions in viral phylogenetics","Infinite hidden Markov models yield biggest phylogenetic fit gains","No single prior best for across-site evolutionary variation","Bayesian nonparametric priors: each wins on different data","Infinite mixtures adapt to each dataset: no one-size-fits-all prior"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001125,"raw_usage":{"total_tokens":4703,"prompt_tokens":993,"completion_tokens":3710,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":3628}},"tokens_in":609,"tokens_out":3710,"duration_ms":25103,"temperature":1.0,"reasoning_tokens":3628,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:04:44.885791+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-estimate the log marginal likelihood for every model on all three data sets with generalized stepping-stone or path sampling and check whether the rankings survive—specifically, whether HKY+IHMM keeps its roughly 100-unit lead over HKY+HDP on RSV A, HKY+DP keeps its lead on HCV, and HKY+IHMM keeps its more than 1000-unit lead on rabies; if the ordering changes materially, the paper's comparative claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the hierarchical Dirichlet process construction and the Chinese restaurant franchise sampler used for HDP and IHMM."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes Dirichlet process mixtures for across-site rate variation, the baseline infinite mixture the paper extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the data-squashing MCMC strategy for jointly updating site-to-category assignments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the stabilized harmonic mean marginal likelihood estimator on which all model comparisons rest."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the adjusted version of the harmonic mean estimator that the paper actually uses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the respiratory syncytial virus subgroup A data set used in one of the three empirical comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the hepatitis C subtype 4 data set used in one of the three empirical comparisons."}],"review_version":1}