Pith. sign in

REVIEW 2 major objections 5 minor 36 references

Posterior bounds on divergence time of two sequences under dependent-site evolutionary models

T0 review · 2 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A new theorem places the posterior distribution of divergence time within a logarithmic factor of the p-distance, even under dependent-site evolutionary models.

desk verdict Real concentration result for a single-lineage branch time under dependent-site CTMCs, but Eq. (2) is not the two-sequence divergence-time likelihood in general, and Corollary 1's proof has an unabsorbed log n. read the letter →

arxiv 2507.19659 v1 pith:EMY5VYMR submitted 2025-07-25 q-bio.PE math.PR

classification q-bio.PEmath.PR MSC 60J2792D1562F15
keywords divergencetimeestimationposteriorconcentrationp-distancesite-dependentsubstitutionmodelscontext-dependentratescontinuous-timeMarkovchainbirth-deathcouplingphylogenetics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

For two observed DNA sequences with Hamming distance $r$, the paper proves that when $r\log(n)/n$ is small, the posterior distribution of their divergence time $T$ concentrates below $(r/n)\log(n)$ times a constant, with the probability of being outside this interval decaying like $\exp(-O(r\log n))$. The result holds for a broad class of continuous-time Markov models in which sites evolve dependently, provided the prior does not place exponentially small mass near $r/n$. For models with constant mutation rates, a corollary removes the log factor and shows that $T$ exceeds a constant multiple of the p-distance with vanishingly small posterior probability. This matters because under site-dependent models the likelihood is expensive or intractable, so knowing in advance that high-posterior $T$ values live near the cheap p-distance makes MCMC and maximum-likelihood searches much more efficient.

What carries the argument

The proof works on the uniformized jump-chain representation of the continuous-time Markov chain, writing the likelihood as $e^{-\lambda T} \sum_m (\lambda T)^m/m!\, \tilde{R}^m_{x,y}$, a Poisson mixture over the number of mutations $m$. The sum is split into three regimes: few jumps, an intermediate range, and very many jumps. The few-jump terms are controlled by a Poisson tail bound; the intermediate and many-jump terms are bounded by coupling the Hamming distance from $x$ to a stochastically minorizing birth-death chain on $\{0,\dots,n\}$, whose stationary distribution is binomial. Hitting-time estimates control the intermediate regime, and a spectral-gap bound via canonical paths controls the many-jump regime once the chain is near stationarity. A crude lower bound, Lemma 1, $p(T,\tilde{Q})(y|x) \ge \exp(-\lambda T + r\log(T\tilde{\gamma}_{\min}))$, combines with the upper bounds to give the exponential likelihood ratio.

What would settle it

Simulate a reversible, stationary dependent-site continuous-time Markov chain with known divergence time $T^*$ and a non-pathological prior, for several large $n$ with $r \approx T^* n$ and $r\log(n)/n$ small; if the observed posterior repeatedly puts non-negligible mass above $(r/n)\log(n)\, c(\epsilon)$ for fixed $c(\epsilon)$, the bound fails. Equivalently, for a constant-rate model, check numerically whether the probability that $T$ exceeds $c r/n$ decays like $e^{-c''r}$; a polynomial or flat tail would contradict Corollary 1.

Watch

Extended reading notes

Core claim

At the center is Theorem 1: under the low-divergence condition $r\log(n)/n = o(1)$, there is a constant $c(\epsilon)$ such that, for all sufficiently large $n$, the posterior probability that $T$ lies in $(0, (r/n)\log(n)\, c(\epsilon))$ is at least $1 - \exp(-c'(\epsilon)\, r\log n)$, whenever the prior gives $J = ((1-\epsilon)r/n, (1+\epsilon)r/n)$ mass that does not decay exponentially in $r$. The proof obtains a sharper two-sided likelihood statement, Theorem 2: for any $T^*$ in $J$ and any $T$ outside the longer interval, the likelihood ratio $p(T,\tilde{Q})(y|x)/p(T^*,\tilde{Q})(y|x)$ is at most $\exp(-O(r\log n))$. For symmetric constant-rate models, Corollary 1 strengthens this to an interval $(0, c r/n)$ with posterior tail $\exp(-c'' r)$, removing the log factor. The paper's interpretation is that the simple p-distance is, up to a logarithmic factor, a genuine upper limit on where the posterior of divergence time can sit under a large class of dependent-site models.

Load-bearing premise

The proof treats the probability that one observed sequence evolves into the other along a single lineage as the likelihood of their divergence time, which is the full likelihood only when the model is reversible and stationary; for general non-reversible dependent-site models, the object being bounded is a different quantity.

Editorial extensions

If this is right

  • For low-to-moderately diverged sequences, posterior sampling and maximum-likelihood routines only need to evaluate the likelihood for $T$ in $(0, O((r/n)\log n))$, avoiding expensive likelihood approximations at large $T$ under site-dependent models.
  • For constant-rate models, the posterior credible region shrinks to $(0, c r/n)$, recovering and refining the earlier logarithmic bound for the two-state symmetric model.
  • The condition $r\log(n)/n = o(1)$ marks the limits of the result: once the sequences are too diverged, the likelihood flattens and a bound of this type cannot hold, consistent with the JC69 maximum-likelihood estimator failing for $r/n \ge 3/4$.
  • The posterior tail bound is uniform over a broad class of dependent-site models, so the same interval can be used to initialize or constrain iterative optimization in models where exact likelihoods are intractable.
  • The bound also implies that, under a non-pathological prior, the posterior cannot be dominated by large divergence times even though the likelihood ratio favors small $T$ exponentially strongly.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors suspect that the $\log n$ factor is an artifact of ignoring the conditional distribution over sequences at distance $r$; a natural test is whether the $c r/n$ bound holds under dependent-site models once $P(X_m = y \mid d_H = r)$ is accounted for, as done in the constant-rate corollary.
  • The theorem is stated for the single-lineage transition probability $p(T)(y|x)$; if the true two-sequence likelihood under a non-reversible model differs, the bound should be rechecked against a symmetrized likelihood before being used in practice.
  • For finite $n$ the constants $c$ and $c'$ are implicit; simulations could calibrate how large $n$ must be for the exponential tail to dominate, and whether the bound is practically useful for typical sequence lengths.
  • A related testable extension is that the same coupling and spectral-gap machinery might yield concentration bounds on estimated branch lengths for three or more sequences, or on whole phylogenetic trees, rather than only pairwise divergence times.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This paper studies posterior concentration of the divergence time T between two length-n DNA sequences under continuous-time Markov models with site dependence. The main result (Theorem 1) states that when the Hamming distance r satisfies r log n / n = o(1), the posterior probability that T lies outside (0, (r/n) log n c(epsilon)) is at most e^{-c'(epsilon) r log n}, provided the prior gives non-exponentially small mass to a small interval around r/n. A corollary claims an improved bound T = O(r/n) with probability 1 - e^{-O(r)} for symmetric constant-rate models. The proofs use uniformization, a stochastic dominance coupling of the Hamming-distance process to a birth-death chain, hitting-time and spectral-gap estimates, and explicit tail bounds.

Significance. If correct, the result would provide a rigorous, model-rich generalization of Mihaescu and Steel's logarithmic posterior bound, connecting the simple p-distance to high-posterior credible regions for divergence times under dependent-site models. The paper is largely self-contained: the supporting lemmas are stated with explicit constants, the coupling argument is detailed, and the spectral-gap bound is constructive. The main concern is that the likelihood used in Eq. (2) is a single-lineage transition probability rather than the two-sequence divergence likelihood, so the theorem as stated does not apply to the biological target named in the title. The proof of Corollary 1 also leaves a 10 log n term that invalidates the e^{-O(r)} claim for bounded r.

major comments (2)
  1. [Section 2, Eq. (2)] The likelihood is defined as the single-lineage transition probability p(T,\tilde Q)(y|x), but the paper's stated target is the divergence time of two sequences descended from an unknown common ancestor. For a stationary, reversible model the pairwise likelihood obtained by integrating over the ancestor is proportional to p_{2T}(y|x), not p_T(y|x); for general non-reversible dependent-site models it is not expressible as a single transition probability at all. Consequently Theorem 1 and Corollary 1 bound the posterior of a different quantity (an ancestor-to-descendant time, or a rescaled half-branch length) unless the model class and the meaning of T are restricted and the statements are revised. This is load-bearing because the abstract and introduction claim concentration for the divergence time of two sequences.
  2. [Section 4 (Theorem 2 proof) and Section 5.6 (Corollary 1 proof)] The upper bound in Lemma 11 carries a +10 log n term. In the proof of Theorem 2, after choosing h(n) = 2r log(n c2 gamma_min)/(n c0), the ratio bound has exponent -(r-10) log n + O(r), which is positive for r < 10; since the assumption r log n / n = o(1) allows r constant, the conclusion exp(-O(r log n)) is not justified for small r. In Corollary 1, substituting t1 = O(r/n) into the displayed ratio bound leaves an exponent -O(r) + 10 log n + O(1), so the claimed e^{-O(r)} bound fails for bounded r. Both gaps can likely be repaired by choosing h(n) with a larger constant (or, for Corollary 1, by accepting a slower rate), but as written the proofs do not establish the stated rates.
minor comments (5)
  1. [Section 2, notation] The definition of \tilde gamma_max is identical to \tilde gamma_min (both written as min over b and \tilde x); it should be max over b and \tilde x, since later arguments use \tilde gamma_max as the maximum rate.
  2. [Theorem 1 statement] The constant c'(epsilon) appears in the probability bound but is not introduced in the statement; please define it alongside c(epsilon).
  3. [Lemma 4 proof] The final Hoeffding step is omitted; as written, the bound exp(-m* gamma_min^2 / 8) appears to require lambda = 1 or a missing lambda^2 factor, and the inequality "2 delta / (lambda t q) - rho <= -gamma_min / 2" is dimensionally inconsistent and should instead read -gamma_min / (2 lambda).
  4. [Theorem 1 and Section 3] The phrase "\eta_0(J) does not decay exponentially in r" should be formalized (for example, as \eta_0(J) >= e^{-o(r)}), since the proof uses this condition to control \eta_0(I^c)/\eta_0(J).
  5. [Theorem 2 statement] The interval notation "I := (0, r c' / n log(n))" is ambiguous; it should read I := (0, (r/n) c' log n).

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the derivation is self-contained under its stated model assumptions.

full rationale

The paper derives posterior concentration bounds for divergence time from explicit upper and lower bounds on sequence transition probabilities (Lemmas 1-11), assembled in Theorem 2 and then converted into the posterior statement in Theorem 1. No parameter is fitted to data and then renamed as a prediction: the constants c(epsilon) and c'(epsilon) are either explicitly defined in the proof or existentially quantified, and the interval (0, p_hat log(n) c(epsilon)) is constructed from the observed p-distance and the stated condition r log(n)/n = o(1). The prior condition that eta0(J) does not decay exponentially is a genuine assumption about the prior, not an import of the conclusion. The self-citations in the paper (refs. [20] and [31]) are contextual or motivational and are not load-bearing for Theorem 1; no uniqueness theorem is invoked from prior work, and no ansatz is smuggled in via citation. The possible objection that Eq. (2) defines the posterior using the single-lineage transition probability p(T,Q~)(y|x) rather than a likelihood that marginalizes over an unknown ancestral sequence is a modeling or correctness concern about whether the theorem addresses the intended biological quantity, not a circularity: the theorem's claims concern the posterior generated by the likelihood as explicitly defined. Nothing in the derivation reduces by construction to its own inputs, so the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper is a pure proof paper. It fits no data; all numerical constants arise from proof choices. The central claim rests on the CTMC model assumptions, on identifying the transition probability as the divergence-time likelihood, and on the prior mass condition. No new entities are introduced.

assumptions (4)
  • domain assumption The evolutionary process is a finite-state CTMC on sequence space with generator Q~ given by (1), with no simultaneous multiple substitutions and all single-site mutation rates bounded below by gamma_min > 0.
    Stated in Section 2; the lower bound in Lemma 1 and all mixing bounds use gamma_min > 0.
  • domain assumption The likelihood of divergence time T is the transition probability p(T,Q~)(y|x) from observed sequence x to y, rather than a joint two-leaf likelihood with unknown ancestor.
    Equation (2) defines L(T|x,y) = p(T,Q~)(y|x). This is only the full divergence-time likelihood under a reversible, stationary interpretation, which is not stated.
  • domain assumption The prior eta0 has non-pathological mass: eta0(J) does not decay exponentially in r for J around r/n.
    Required in Theorem 1 and Corollary 1; standard priors satisfy this, but the posterior tail bound is conditional on it.
  • standard math Standard Markov chain and coupling results: stochastic dominance coupling existence, Poisson tail bounds, Hoeffding's inequality, the Aldous-Fill hitting bound in (32), and Diaconis-Stroock canonical paths.
    Used throughout Section 5; these are textbook results.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Posterior bounds on divergence time of two sequences under dependent-site evolutionary models." pith.science (2026). https://pith.science/paper/EMY5VYMR

@misc{pith2026250719659,
  author       = {Pith},
  title        = {Pith review of: Posterior bounds on divergence time of two sequences under dependent-site evolutionary models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EMY5VYMR}},
  note         = {Machine review of arXiv:2507.19659}
}
read the original abstract

Let x and y be two length n DNA sequences, and suppose we would like to estimate the divergence time T. A well known simple but crude estimate of T is p := d(x,y)/n, the fraction of mutated sites (the p-distance). We establish a posterior concentration bound on T, showing that the posterior distribution of T concentrates within a logarithmic factor of p when d(x,y)log(n)/n = o(1). Our bounds hold under a large class of evolutionary models, including many standard models that incorporate site dependence. As a special case, we show that T exceeds p with vanishingly small posterior probability as n increases under models with constant mutation rates, complementing the result of Mihaescu and Steel (Appl Math Lett 23(9):975--979, 2010). Our approach is based on bounding sequence transition probabilities in various convergence regimes of the underlying evolutionary process. Our result may be useful for improving the efficiency of iterative optimization and sampling schemes for estimating divergence times in phylogenetic inference.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

36 extracted references · 36 canonical work pages

  1. [1]

    Molecular Biology and Evolution 12(3), 391–404 (1995)

    Russo, C.A., Takezaki, N., Nei, M.: Molecular phylogeny and divergence times of drosophilid species. Molecular Biology and Evolution 12(3), 391–404 (1995)

  2. [2]

    Theoretical Population Biology 48(2), 198– 221 (1995)

    Takahata, N., Satta, Y., Klein, J.: Divergence time and population size in the lineage leading to modern humans. Theoretical Population Biology 48(2), 198– 221 (1995)

  3. [3]

    Mol Biol Evol 20(3), 424–434 (2003)

    Glazko, G.V., Nei, M.: Estimation of divergence times for major lineages of primate species. Mol Biol Evol 20(3), 424–434 (2003)

  4. [4]

    Nature 392, 917–920 (1998)

    Kumar, S., Hedges, S.B.: A molecular timescale for vertebrate evolution. Nature 392, 917–920 (1998)

  5. [5]

    Journal of Molecular Evolution 22(2), 160–174 (1985) 21

    Hasegawa, M., Kishino, H., Yano, T.: Dating of the human-ape splitting by a molecular clock of mitochondrial DNA. Journal of Molecular Evolution 22(2), 160–174 (1985) 21

  6. [6]

    Science 155(3760), 279–284 (1967)

    Fitch, W.M., Margoliash, E.: Construction of phylogenetic trees. Science 155(3760), 279–284 (1967)

  7. [7]

    Systematic Biology 20(4), 406–416 (1971)

    Fitch, W.M.: Toward defining the course of evolution: Minimum change for a specific tree topology. Systematic Biology 20(4), 406–416 (1971)

  8. [8]

    Molecular Biology and Evolution 4(4), 406–425 (1987)

    Saitou, N., Nei, M.: The neighbor-joining method: a new method for recon- structing phylogenetic trees. Molecular Biology and Evolution 4(4), 406–425 (1987)

Show all 36 references
  1. [9]

    Molecular Biology and Evolution 9(5), 945–967 (1992)

    Rzhetsky, A., Nei, M.: A simple method for estimating and testing minimum- evolution trees. Molecular Biology and Evolution 9(5), 945–967 (1992)

  2. [10]

    In: Munro, H.N

    Jukes, T.H., Cantor, C.R.: Evolution of protein molecules. In: Munro, H.N. (ed.) Mammalian Protein Metabolism, pp. 121–132. Academic Press, New York (1969)

  3. [11]

    Journal of Molecular Evolution 16(2), 111–120 (1980)

    Kimura, M.: A Simple Method for Estimating Evolutionary Rates of Base Sub- stitutions through Comparative Studies of Nucleotide Sequences. Journal of Molecular Evolution 16(2), 111–120 (1980)

  4. [12]

    Molecular Biology and Evolution 25(3), 568–579 (2008)

    Yang, Z., Nielsen, R.: Mutation-selection models of codon substitution and their use to estimate selective strengths on codon usage. Molecular Biology and Evolution 25(3), 568–579 (2008)

  5. [13]

    Systematic Zoology 22(3), 240–249 (1973)

    Felsenstein, J.: Maximum likelihood and minimum-steps methods for estimating evolutionary trees from data on discrete characters. Systematic Zoology 22(3), 240–249 (1973)

  6. [14]

    Journal of Molecular Evolution 17, 368–376 (1981)

    Felenstein, J.: Evolutionary trees from DNA sequences: A maximum likelihood approach. Journal of Molecular Evolution 17, 368–376 (1981)

  7. [15]

    Molecular Biology and Evolution 14(7), 717–724 (1997)

    Yang, Z., Rannala, B.: Bayesian phylogenetic inference using DNA sequences: a Markov chain Monte Carlo method. Molecular Biology and Evolution 14(7), 717–724 (1997)

  8. [16]

    Molecular Biology and Evolution 15, 1647–1657 (1998)

    Thorne, J.L., Kishino, H., Painter., I.S.: Estimating the rate of evolution of the rate of molecular evolution. Molecular Biology and Evolution 15, 1647–1657 (1998)

  9. [17]

    Molecular Biology and Evolution 15, 1069–1081 (1998)

    Pedersen, A.K., Wiuf, C., Christiansen, F.B.: A codon-based model designed to describe lentiviral evolution. Molecular Biology and Evolution 15, 1069–1081 (1998)

  10. [18]

    Molecular Biology and Evolution 20, 1692–1704 (2003)

    Robinson, D., Jones, D., Kishino, H., Goldman, N., Thorne, J.: Protein evolution with dependence among codons due to tertiary structure. Molecular Biology and Evolution 20, 1692–1704 (2003)

  11. [19]

    Cell Host Microbe 23(6), 759–765 (2018)

    Wiehe, K., Bradley, T., Meyerhoff, R., Hart, C., Williams, W., Easterhoff, D., 22 Faison, W., Kepler, T., Saunders, K., Alam, S., Bonsignori, M., Haynes, B.: Func- tional relevance of improbable antibody mutations for HIV broadly neutralizing antibody development. Cell Host Mi...

  12. [20]

    The Journal of Immunology (2023)

    Mathews, J., Itallie, E.V., Li, Y., Wiehe, K., Schmidler, S.C.: Computing the Inducibility of B Cell Lineages Under a Context-Dependent Model of Affin- ity Maturation: Applications to Sequential Vaccine Design. The Journal of Immunology (2023). (provisionally accepted)

  13. [21]

    Molecular Biology and Evolution 21(3), 468–488 (2004)

    Siepel, A., Haussler, D.: Phylogenetic estimation of context-dependent substi- tution rates by maximum likelihood. Molecular Biology and Evolution 21(3), 468–488 (2004)

  14. [22]

    Advances in Applied Probability 32(2), 499–517 (2000)

    Jensen, J., Pedersen, A.-M.: Probabilistic models of DNA sequence evolution with context dependent rates of substitution. Advances in Applied Probability 32(2), 499–517 (2000)

  15. [23]

    Journal of Computational Biology 27, 361–375 (2020)

    Larson, G., Thorne, J.L., Schmidler, S.C.: Incorporating nearest-neighbor site dependence into protein evolution models. Journal of Computational Biology 27, 361–375 (2020)

  16. [24]

    Proceedings of the National Academy of Science 101, 13994–14001 (2004)

    Hwang, D., Green, P.: Bayesian Markov chain Monte Carlo sequence anal- ysis reveals varying neutral substitution patterns in mammalian evolution. Proceedings of the National Academy of Science 101, 13994–14001 (2004)

  17. [25]

    Journal of Computational Biology 5, 149–163 (1998)

    Haeseler, A., Sch¨ oniger, M.: Evolution of DNA or amino acid sequences with dependent sites. Journal of Computational Biology 5, 149–163 (1998)

  18. [26]

    Journal of Computational Biology 12, 1166–1182 (2005)

    Christensen, O.F., Hobolth, A., Jensen, J.L.: Pseudo-likelihood analysis of context-dependent codon substitution models. Journal of Computational Biology 12, 1166–1182 (2005)

  19. [27]

    Bioinformatics 21, 2322–2328 (2005)

    Arndt, P.F., Hwa, T.: Identification and measurement of neighbour-dependent nucleotide substitution processes. Bioinformatics 21, 2322–2328 (2005)

  20. [28]

    Bioinformatics 20 Suppl 1, 216–223 (2004)

    Lunter, G., Hein, J.: A nucleotide substitution model with nearest-neighbour interactions. Bioinformatics 20 Suppl 1, 216–223 (2004)

  21. [29]

    Molecular Biology and Evolution 18, 763–776 (2001)

    Pederson, A.-M., Jensen, J.: A dependent rates model and MCMC based method- ology for the maximum likelihood analysis of sequences with overlapping reading frames. Molecular Biology and Evolution 18, 763–776 (2001)

  22. [30]

    Applied Mathematics Letters 23(9), 975–979 (2010)

    Mihaescu, R., Steel, M.: Logarithmic bounds on the posterior divergence time of two sequences. Applied Mathematics Letters 23(9), 975–979 (2010)

  23. [31]

    Unpublished Manuscript 23

    Mathews, J., Schmidler, S.C.: Approximating marginal likelihoods in evolutionary models under site-dependence (2025). Unpublished Manuscript 23

  24. [32]

    John Wiley & Sons, Hoboken, New Jersey (1995)

    Ross, S.M.: Stochastic Processes, 2nd edn. John Wiley & Sons, Hoboken, New Jersey (1995)

  25. [33]

    Unfinished monograph (2002)

    Aldous, D., Fill, J.A.: Reversible Markov Chains and Random Walks on Graphs. Unfinished monograph (2002)

  26. [34]

    American Mathematical Society, Providence, RI (2009)

    Levin, D.A., Peres, Y., Wilmer, E.L.: Markov Chains and Mixing Times, 2nd edn. American Mathematical Society, Providence, RI (2009)

  27. [35]

    Annals of Applied Probability 1(1), 36–61 (1991)

    Diaconis, P., Stroock, D.: Geometric bounds for eigenvalues of Markov chains. Annals of Applied Probability 1(1), 36–61 (1991)

  28. [36]

    ALEA, Latin American Journal of Probability and Mathematical Statistics 10(1), 293–321 (2013) 24

    Guan-Yu, C., Saloff-Coste, L.: On the mixing time and spectral gap for birth and death chains. ALEA, Latin American Journal of Probability and Mathematical Statistics 10(1), 293–321 (2013) 24

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.