REVIEW 2 major objections 3 minor 14 references
Allele trees for the mother-dependent neutral mutations model and their scaling limits in the rare mutations regime
T0 review · 2 major / 3 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Finite-allele populations with rare neutral mutations converge to a single universal allele tree.
desk verdict The main theorem's scaling is wrong for every non-root node: level-1 allele family sizes are O(n), so n^{-2}A_u cannot converge to a positive limit; the paper's intermediate clone-mutant chain results are solid, but Theorem 1.1 as stated is false. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrier of the argument is the clone–mutant Markov chain $((T_k,M_{k+1}))_{k\ge 0}$: $T_k(i)$ is the number of type-$i$ individuals in allelic generation $k$ (the clone families), and $M_{k+1}(i)$ counts the type-$i$ mutants that start the next allelic generation. Its one-step transition is governed by the pair $(T_0,M_1)$ associated with a single ancestor. The proof converts this pair into a random-walk object: $T_0(i)$ is the first hitting time $\tau_0$ of level zero by the breadth-first random walk that codes the type-$i$ subtree, and $M_1$ is the total mutant offspring accumulated along that walk; the transition law then follows from a generating-function inversion formula. In the scaling limit the random walk converges to a Brownian motion with drift, so the hitting time becomes the inverse Gaussian $\theta_1$ and the offspring structure becomes the ranked atoms of a Poisson random measure with intensity $Y_u\nu(dz)$. The allele tree itself is assembled by a recursive priority rule: rank mothers by decreasing number of mutant children, then by increasing type, then by decreasing subtree size.
What would settle it
The ordering equivalence can be settled by simulation or direct construction: for $d=2$, large $n$, and several values of $c$ and $\sigma^2$, record the first allelic generation and compare the rank of a subfamily under the Section 4 priority rule with its rank by decreasing size. If the two orders differ on a set of positive limiting probability, the limit is not the ranked Poisson tree of Definition 5.3. A simpler margin to check is the one-dimensional law: if $n^{-2}T_0$ does not converge to the inverse Gaussian $\mathrm{IG}(1/c,1/\sigma^2)$, or $n^{-1}M_1$ does not converge to the constant allocation $\frac{c}{d-1}\theta_1$ on each other type, Theorem 1.1 is false.
Extended reading notes
Core claim
The central claim is Theorem 1.1: starting from $n$ individuals of a single type $j$, rescale the allele tree $A^{(n)}$ as $(n^{-2}A^{(n)}_u,\,C^{(n)}_u,\,n^{-1}d^{(n)}_u)_{u\in\mathbb{U}}$. Under hypotheses (H2) and (H3), namely mutation rate $r(n)\sim c/n$ and a critical offspring law with finite variance $\sigma^2$, this rescaled tree converges in the sense of finite-dimensional distributions to $\big(Y_u,\,C_u,\,\frac{c}{d-1}Y_u\sum_{i\neq C_u}e_i\big)_{u\in\mathbb{U}}$. Here $(Y_u)_{u\in\mathbb{U}}$ is a tree-indexed CSBP with reproduction measure $\nu(dz)=\frac{c}{\sqrt{2\pi\sigma^2 z^3}}\exp\!\big(-\frac{c^2 z}{2\sigma^2}\big)dz$ and random root value $\theta_1\sim\mathrm{IG}(1/c,1/\sigma^2)$, the first-passage time of a Brownian motion with drift $c$. The limiting object is the universal allele tree introduced in the infinite-allele setting; the finite-allele model changes only the allocation of mutants across types. If the initial population contains all types in proportions $y(i)$, the result extends to a forest of such tree-indexed CSBPs, one for each initial type.
Load-bearing premise
The load-bearing premise is that the deterministic bookkeeping order used to list allelic subfamilies—largest mutant families first, ties broken by type and then by subtree size—coincides after rescaling with the decreasing-size ranking of the limiting Poisson atoms, and that the induction over levels of the genealogical index tree is legitimate for finite-dimensional convergence; Section 6.6 assumes this without a separate proof.
Editorial extensions
If this is right
- Large allelic subfamilies in the rare-mutation limit have sizes governed by the stable $1/2$ reproduction measure $\nu(dz)$, so the infinite-allele universality class persists with finitely many alleles.
- The limiting mutant counts of a subfamily of size $Y_u$ and type $C_u$ are exactly $(c/(d-1))Y_u$ for each other type; the allele set only rescales the mutant vector, it does not alter the branching structure.
- The whole clone–mutant chain converges to a continuous-state Markov chain whose transition cumulant is $\kappa_j(x,z)=\kappa(x(j)+\frac{c}{d-1}\sum_{i\ne j}z(i))$, so allelic generations remain Markovian in the limit.
- With all types present initially, the limit is a forest of independent tree-indexed CSBPs, one per initial type, with independent inverse Gaussian root sizes.
Reading between the lines
- The linear factor $c/(d-1)$ is a concrete finite-allele correction that could be looked for in population-genetic data: under rare neutral mutations, the relative mutant counts on the other $d-1$ alleles should be exchangeable with ratio fixed by $c$, independent of $\sigma^2$.
- A natural extension the paper does not treat lets the number of alleles grow with $n$; if $d(n)\to\infty$, the factor $c/(d-1)$ suggests convergence to a continuum-of-types object, possibly the infinite-allele universal tree again.
- The unproved ordering equivalence in Section 6.6 could be tested by an explicit coupling of the recursive allele-tree order with the decreasing-size ranking of Poisson atoms; if it fails for one tie-breaking rule, the theorem may still hold after modifying the priority rule.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the mother-dependent neutral mutations model with d alleles, where each mutant child acquires a type different from its mother, uniformly at random. It defines a multitype allele tree whose nodes record the sizes of clone subfamilies, their types, and their mutant offspring vectors. The main result, Theorem 1.1, claims that when the initial population consists of n individuals of one type, the mutation rate is r(n) ~ c/n, and the offspring distribution is critical with variance σ^2, the rescaled tree (n^{-2}A_u, C_u, n^{-1}d_u) converges in finite-dimensional distributions to Bertoin's universal tree-indexed CSBP with reproduction measure ν(dz) = c(2πσ^2 z^3)^{-1/2} exp(-c^2 z/(2σ^2)) dz and random initial population θ_1 ~ IG(1/c, 1/σ^2). The proof strategy is to analyze a clone-mutant Markov chain (Section 3), prove its scaling limit (Lemma 5.1, Proposition 5.2), and then argue that the allele tree inherits this limit (Section 6.6).
Significance. If the main theorem were correct, it would provide a finite-allele extension of Bertoin's universal allele tree and would be a valuable contribution to the scaling limits of branching structures with neutral mutations. The paper's auxiliary analysis is largely sound: the random-walk representation in Lemma 3.5 and the weak convergence proofs of Lemma 5.1 and Proposition 5.2 are carried out in detail, and the clone-mutant Markov chain is a natural and potentially reusable tool. However, the central theorem is internally inconsistent with the model's own scaling: non-root clone subfamilies have size of order n, not n^2, so the claimed convergence cannot hold as stated.
major comments (2)
- [Theorem 1.1, Sections 1.2 and 6.6] For any non-root node u, A_u^{(n)} is defined in Section 4 as the number of individuals in the clone subfamily rooted at a single mutant individual. Under P^{r(n)}_{e_i}, Lemma 3.5 identifies this size with the first hitting time τ_0 of 0 by a random walk with step mean -r(n) = -c/n + o(1/n) and variance σ^2 + o(1). Wald's identity gives E[A_u^{(n)}] = n/c + o(n), and the classical near-critical limit gives n^{-1} A_u^{(n)} ⇒ a nontrivial law, so n^{-2} A_u^{(n)} → 0 in probability. Similarly, d_u^{(n)} for a level-1 node has mean of order 1, so n^{-1} d_u^{(n)} → 0. However, the tree-indexed CSBP of Definition 5.3, with the intensity ν in (9) of infinite total mass, has positive atoms at every level almost surely. Hence the claimed finite-dimensional convergence to (Y_u) with Y_u > 0 for |u|=1 is false. The proof in Section 6.6 derives the theorem from Lemma 5.1, but that lemma concerns the aggregate root family T_0^{(n)}, which is of order n^2; it does not control the individual subfamily sizes, which are of order n. The assertion that 'Lemma 5.1 also implies' the convergence of the ordered atoms is therefore not justified and is in fact inconsistent with the model's own scaling.
- [Section 6.6 and Definition 5.3] The proof of Theorem 1.1 asserts without proof that the ordering of allele-tree nodes defined in Section 4 (rank mothers by decreasing number of mutant children, then by type and subtree size) coincides in the limit with ranking the atoms of the Poisson random measure in Definition 5.3 in decreasing order, and that this order is almost surely well-defined with no ties. This is a load-bearing step for finite-dimensional convergence, because the ranking operation is not continuous in the product topology and the asserted tie-free property of the limiting Poisson atoms is not established. Section 6.6 provides no argument beyond citing 'arguments similar to those in [2, Theorem 1]'; given the additional type-ordering mechanism in Section 4, this is an essential gap.
minor comments (3)
- [Theorem 1.1, statement] The statement says the model follows the law P^{r(n)}_{n^{-1} e_j} and 'starting with one indivivual of type j'; this should be P^{r(n)}_{n e_j} and 'starting with n individuals of type j'.
- [Equation (9)] The exponential in the reproduction measure uses the variable y, which is undefined; it should be z, as in Definition 5.3.
- [Section 6.4, around equation (33)] The text says 'the following convergence holds a.s.' followed by a weak convergence arrow; this should be convergence in distribution (functional CLT), not almost sure convergence.
Circularity Check
No circularity: the scaling limit is computed from the random-walk CLT and the external Bertoin framework; self-citations are attribution only.
full rationale
The derivation is self-contained with respect to its main claim. Theorem 1.1 is reduced to Lemma 5.1, and Lemma 5.1 is obtained by a direct functional CLT on the random walk S^{(n,j)} defined in Section 6.4, with mean and variance computed from the model's offspring law; the inverse Gaussian variable theta_1 and the reproduction measure nu(dz) emerge from that computation rather than being inserted to match the tree-indexed CSBP of Definition 5.3. The recursive extension to all levels invokes the Markov property of the clone-mutant chain (Lemma 3.2) and follows the argument of [2, Theorem 1], an external source. The only self-citations are [3] for the definition of the mother-dependent model and [4] as a related branching-property reference; neither supplies a load-bearing premise, and the proof does not cite a self-authored uniqueness theorem or fit any parameter to the limiting object. The possible gap in Section 6.6 concerning equality of the two orderings in the limit is an unproved justification step, not a circular reduction: even if it failed, the claimed limit would be wrong, not identical to the input by construction.
Assumptions & free parameters
assumptions (4)
- standard math General branching property for multitype BGW processes (Jagers, Theorem 2.2): subtrees rooted at a stopping line are independent and distributed as the original process with the root's type.
- standard math Donsker invariance principle for mean-zero finite-variance random walks; first-passage times converge to hitting times of Brownian motion with drift.
- domain assumption The tree-indexed CSBP with reproduction measure nu and initial population a exists and is characterized by the Poissonian construction in Definition 5.3, as established in Bertoin [2].
- ad hoc to paper Under the scaling regime, ranking mothers by decreasing number of mutant children coincides with ranking the limiting clone subfamily sizes (Poisson atoms) in decreasing order, with no ties almost surely.
Cite this review
Pith. "Pith review of Allele trees for the mother-dependent neutral mutations model and their scaling limits in the rare mutations regime." pith.science (2026). https://pith.science/paper/PNENHWCP
@misc{pith2026250420030,
author = {Pith},
title = {Pith review of: Allele trees for the mother-dependent neutral mutations model and their scaling limits in the rare mutations regime},
year = {2026},
howpublished = {\url{https://pith.science/paper/PNENHWCP}},
note = {Machine review of arXiv:2504.20030}
}
read the original abstract
The mother-dependent neutral mutations model describes the evolution of a population across discrete generations, where neutral mutations occur among a finite set of possible alleles. In this model, each mutant child acquires a type different from that of its mother, chosen uniformly at random. In this work, we define a multitype allele tree associated with this model and analyze its scaling limit through a Markov chain that tracks the sizes of allelic subfamilies and their mutant descendants. We show that this Markov chain converges to a continuous-state Markov process, whose transition probabilities depend on the sizes of the initial allelic populations and those of their mutant offspring in the first allelic generation. As a result, the allele tree converges to a multidimensional limiting object, which can be described in terms of the universal allele tree introduced by Bertoin (2010).
Figures
Reference graph
Works this paper leans on
-
[3]
Crossing bridges between percolation models and Bienaym\'e-Galton-Watson trees
A. Blancas, M.C. Fittipaldi, and S Hern´ andez-Torres. Crossing bridges between percolation models and Bienaym´ e-Galton-Watson trees. To appear in XIV Symposium on Probability and Stochastic Processes. Preprint available at arXiv:2411.09621. 2The order by clone descendants is defined above [2, Lemma 2] and considered in the proof of [2, Theorem 1]. 29
-
[1]
J. Bertoin. The structure of the allelic partition of the total population for galton–watson processes with neutral mutations.Ann. Probab., 37(4):1502–1523, 2009
work page 2009
-
[2]
J. Bertoin. A limit theorem for trees of alleles in branching processes with rare neutral mutations.Stoch. Process. Their Appl., 120(5):678–697, 2010
work page 2010
-
[4]
A. Blancas and S. Palau. Coalescent point process of branching trees in a varying environment. Electron. Commun. Probab., 29:1 – 15, 2024
work page 2024
-
[5]
L. Chaumont and R. Liu. Coding multitype forests: application to the law of the total population of branching forests.Trans. Amer. Math. Soc., 368(4):2723–2747, 2016
work page 2016
-
[6]
B. Chauvin. Sur la propri´ et´ e de branchement. InAnnales de l’IHP Probabilit´ es et statistiques, volume 22, pages 233–236, 1986
work page 1986
-
[7]
M. Dwass. The total progeny in a branching process and a related random walk.J. Appl. Probab., 6(3):682–686, 1969
1969
-
[8]
P. Jagers. General branching processes as markov fields.Stoch. Process. Their Appl., 32(2):183–212, 1989
work page 1989
Show all 14 references
-
[9]
Kuipers, K
J. Kuipers, K. Jahn, B. J Raphael, and N. Beerenwinkel. Single-cell sequencing data reveal widespread recurrence and loss of mutational hits in the life histories of tumors.Genome Res., 27(11):1885–1894, 2017
2017
-
[10]
A. E. Kyprianou. Martingale convergence and the stopped branching random walk.Probab. Theory Relat. Fields., 116(3):405–419, 2000
2000
-
[11]
L. A. Mathew, P. R. Staab, L. E. Rose, and D. Metzler. Why to account for finite sites in population genetic studies and how to do this with Jaatha 2.0.Ecol. Evol., 3(11):3647–3662, 2013
2013
-
[12]
R. Otter. The multiplicative process.Ann. Math. Stat., pages 206–224, 1949
1949
-
[13]
Pitman.Combinatorial stochastic processes: Ecole d’et´ e de probabilit´ es de Saint-Flour XXXII-2002
J. Pitman.Combinatorial stochastic processes: Ecole d’et´ e de probabilit´ es de Saint-Flour XXXII-2002. Springer, 2006
2002
-
[14]
Zafar, A
H. Zafar, A. Tzen, N. Navin, K. Chen, and L. Nakhleh. SiFit: inferring tumor trees from single-cell sequencing data under finite-sites models.Genome Biol., 18:1–20, 2017. 30
2017
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.