REVIEW 2 major objections 3 minor 1 cited by
Stick-breaking Pitman-Yor processes given the species sampling size
T0 review · 2 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A Pitman-Yor process conditioned on the species-sampling count $N_{S_{\alpha,\theta}}(\lambda)=m$ still has an explicit stick-breaking representation, with the normalized generalized gamma process as the $m=0$ special case.
desk verdict The paper has a nice idea and a repairable but load-bearing bug: the Gamma(m) factor in Omega_m breaks the general-m stick-breaking construction. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery has three parts. First, the mixed-Poisson conditioning identity for the total abundance $A$: $P(A\in da\,|\,N_A(\lambda)=m)=a^m e^{-\lambda a}P(A\in da)/E[A^m e^{-\lambda A}]$, which converts conditioning on the species count into a size-biased exponential tilt of the stable density. Second, the standard size-biased deletion construction for the ranked jumps of a stable subordinator, giving a product-form joint density for the first $n$ stick weights and the remaining stable mass. Third, a $\beta$-gamma integral identity (Proposition 3.1) that rewrites the tilted density so that, conditionally on auxiliary variables $R_k(\lambda)=((\tilde G_{k-1}+\lambda^\alpha)/(\tilde G_k+\lambda^\alpha))^{1/\alpha}$, the sticks become conditionally independent with densities supported on $(r,1)$. The $R_k$ variables are the load-bearing mechanism: they carry the dependence induced by conditioning and reduce to independent $\beta$ variables in the classical GEM$(\alpha,\theta)$ recovery.
What would settle it
Take $\alpha=1/2$ and $m=1$, simulate a 1/2-stable subordinator until total time $\lambda$, reject all paths that do not have exactly one arrival before $\lambda$, form the size-biased permutation of the normalized jumps, and compare the empirical density of the first stick $W_1$ with the explicit density (3.9) or with Theorem 4.1 specialized to $\alpha=1/2$. Since 1/2-stable jumps can be generated exactly, this is a direct Monte Carlo check; systematic mismatch at any $\lambda>0$ would refute the theorem.
Extended reading notes
Core claim
The central claim is Theorem 4.1. For $(P^{[m]}_\ell(\lambda))$---the law of the ranked masses $(P_\ell) \sim \mathrm{PD}(\alpha,0)$, the two-parameter Poisson-Dirichlet law with $\theta=0$, conditioned on $N_{S_\alpha}(\lambda)=m$---the size-biased permutation $(\tilde P_k(\lambda))$ takes the stick-breaking form $\tilde P_k=(1-W_k)\prod_{l<k}W_l$, where $W_\ell$ is given by $W_\ell := [\beta^{(\ell)}_{N_{\ell,m}(\lambda)-\alpha,\alpha}((1-R_{\ell,m}(\lambda))/R_{\ell,m}(\lambda))+1]^{-1}$. Here the $\beta^{(\ell)}$ are independent $\mathrm{Beta}(1-\alpha,\alpha)$ variables independent of the $R$'s, the $R_{\ell,m}(\lambda)$ are ratio variables built from gamma partial sums and a tilted stable variable, and $N_{\ell,m}(\lambda)$ is a discrete random variable whose probability mass function is explicit. Given the $R$'s, the $W$'s are conditionally independent; unconditionally they are dependent, which is exactly how the conditioning on the count $m$ breaks the classical GEM independence. The paper also proves the $m=0$ specialization $W_k=1-\beta^{(k)}_{1-\alpha,\alpha}(1-R_k(\lambda))$, identifying the normalized generalized gamma process, and shows that randomizing $\lambda$ by a gamma variable recovers the classical independent-$\beta$ GEM$(\alpha,\theta)$ stick weights.
Load-bearing premise
The paper assumes the accuracy of an unpublished predecessor's claim that conditioning on the mixed-Poisson count $N_A(\lambda)=m$ is fully captured by the tilted density $a^m e^{-\lambda a}$ times the law of $A$, and that the conditional law of the ranked masses depends on $A$ only through its value; if that premise fails, the stick-breaking formulas describe tilted mixtures rather than the conditional Pitman-Yor process as claimed.
Editorial extensions
If this is right
- For $m=0$, the normalized generalized gamma process has an explicit stick-breaking representation whose sticks are simple functions of beta variables and the gamma-ratio variables $R_k(\lambda)$.
- For $m\geq 1$, exact simulation of the conditional process is possible without rejection sampling: draw $N_{\ell,m}(\lambda)$ and $R_{\ell,m}(\lambda)$, then draw independent beta variables and form $W_\ell$ by the formula in Theorem 4.1.
- At $\alpha=1/2$ the results specialize to explicit stick-breaking for the normalized inverse Gaussian process and for the conditional laws $P^{[m]}_{1/2}(\lambda)$, expressed through inverse Gaussian and Brownian variables.
- Randomizing $\lambda$ via a gamma variable recovers GEM$(\alpha,\theta)$, and in that case the auxiliary $R$-variables become independent beta variables while the $W_\ell$ reduce to the classical independent $\mathrm{Beta}(\theta+\ell\alpha,1-\alpha)$ sticks.
Reading between the lines
- An implicit consequence: the same tilted-density route should yield explicit conditional stick-breaking for any Poisson-Kingman partition whose L\'evy density admits a tractable exponential tilt, not only the stable case; the beta-gamma integral would need a tailored analogue.
- The simple binomial description of $N_{\ell,m}(\lambda)$ in the GEM recovery hints at a sequential-arrival-time reading, connecting the conditional sticks to Chinese-restaurant-type occupancy counts.
- A testable extension: Theorem 4.1 could serve as the sampling core of a Gibbs scheme for normalized generalized gamma mixture models that condition on the observed number of clusters at a latent time, possibly removing auxiliary-variable layers from current posterior algorithms.
- Because the $R$ variables encode the tilted law of the total mass, the representation also opens a route to large-time asymptotics of the first stick, such as the behaviour of $W_1$ as $\lambda$ grows.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper derives explicit stick-breaking representations for Pitman-Yor random probability measures conditioned on the value m of a mixed Poisson process with random rate equal to the α-diversity. It first translates the conditioning into a tilted Poisson-Kingman mixing distribution (Section 1.2), then proves a randomization identity (Proposition 1.1 and Corollary 2.1) that reduces the problem to the stable case P^{[m]}_α(λ). The main results are Theorem 3.1 (m=0, the normalized generalized gamma process), Theorem 3.2 (m=1), and Theorem 4.1 (general m), in which the size-biased permutation has stick-breaking weights W_ℓ = [β_{N_{\ell,m}-α,α}((1-R_{\ell,m})/R_{\ell,m})+1]^{-1}, with R_{\ell,m} and N_{\ell,m} defined in Sections 4.2-4.3. A Brownian/local-time specialization for α=1/2 is given in Corollary 4.4.
Significance. The results are nontrivial and, if the displayed typos are corrected, constitute a solid contribution: Theorem 4.1 gives an explicit constructive algorithm for sampling conditional Pitman-Yor processes, and the m=0 case gives a stick-breaking representation for the normalized generalized gamma process, recovering and placing in context an earlier unpublished result. I checked the algebraic core of the construction: the factor Γ(m) in the definition of Ω_m is not an error (for m=2, Γ(2)=1 and Eq. (2.4) matches direct differentiation), and the probabilities q_{k,m} in Proposition 4.1(ii) sum to 1. The main issue is that several intermediate displayed densities contain typographical errors in normalizing constants and exponents; these are local but must be fixed.
major comments (2)
- [Section 3.3, Eq. (3.6), Lemma 3.3, and Theorem 3.1 proof] As printed, Eq. (3.6) has the wrong normalizing coefficient: for n=1 the coefficient should be α/Γ(1−α), and in general α^n/Γ(1−α)^n, rather than 1/(α^n Γ(1−α)^n). Correspondingly, in Lemma 3.3 and in the augmentation argument in the proof of Theorem 3.1, the factor (1−w_k)^{α−1} should be (1−w_k)^{−α}. With the printed exponent the indicated density does not integrate to 1 and is incompatible with the stated representation W_k=1−β_{1−α,α}(1−R_k). The same normalizing error propagates to Eq. (4.1) in the general-m case. These are local corrections, but they occur in the proof of the central theorem.
- [Section 2.2, Eq. (2.6); Section 4.2, definition of Y^{(j)}_{m,ℓ}(x_j)] The formulas defining S_{α,m}(λ) and Y^{(j)}_{m,ℓ}(x_j) are not well-formed as printed: the argument of τ_α in Eq. (2.6) reads 'λ^α + G_m^α − K_m(λ)' and the second display has a syntactically garbled form, and the display in Section 4.2 has the same problem. Since these identities are used to generate the variables in the general-m stick-breaking representation, they need to be restated precisely with correct superscripts and parentheses.
minor comments (3)
- [Throughout] There are numerous typographical errors: 'thoeory' (page 3), 'statisics' (abstract), 'porposition' (page 14), and inconsistent use of multiplication in Eq. (3.7).
- [Section 4.5, Eq. (4.9)] In Eq. (4.9), θ appears on the right-hand side in the definition of f^{[m]}_{1/2}(t|λ), although the left-hand side is the θ-free density from (2.5). Please clarify whether θ is a free parameter in this display or whether the display is intended as the conditional density of S_{1/2,θ}.
- [References] The paper relies on the unpublished manuscript [30] and the author's unpublished manuscript [18]; since the conditioning formulas in Section 1.2 are elementary and stated explicitly, this is not a correctness issue, but it would help readers if the status of [30] were clarified.
Circularity Check
No circular derivation: the conditioning and stick-breaking results are derived from external Poisson–Kingman theory and in-paper calculations, not from their own conclusions.
full rationale
I find no significant circularity. The conditioning identities (1.2) and (1.3) are not imported as black boxes: the paper explicitly states they are obtained by elementary conditioning on a mixed Poisson process, and the tilting kernel is a direct consequence of the Poisson likelihood. The reliance on Pitman's unpublished manuscript [30] is an external-source dependence, not a self-citation chain: Pitman is not an author of this paper, and the formulas are re-derived in the text. The central m=0 normalized generalized gamma stick-breaking representation is proved in Theorem 3.1 through the paper's own integral identity and augmentation argument, even though the abstract notes that the representation first appeared in the author's unpublished work [18]; that self-citation is historical and not load-bearing. The general-m construction in Sections 4.1–4.3 is built from the Perman–Pitman–Yor size-biased decomposition [24], the stable-subordinator Laplace transform identity (2.4), and internal calculations for the variables R_{l,m} and N_{l,m}; no parameter is fitted to data and no fitted value is later renamed as a prediction. Self-citations such as [18] and [19] are contextual or provide external supporting results, but the central derivation does not reduce to them. A separate reviewer concern about a possible missing Gamma factor in (2.4) and Proposition 4.1 would be a mathematical consistency issue, not a circularity issue, and does not affect this assessment.
Assumptions & free parameters
assumptions (4)
- domain assumption Conditional on A=a, the mixed Poisson process N_A has N_A(λ) ~ Poisson(λa), and the waiting times T_r are G_r/a.
- domain assumption The law of (P_ℓ) given A=t and N_A(λ)=m coincides with PD(α|t), independent of N, so P^{[m]}_α(λ) is the mixture over the tilted density (2.2).
- standard math Size-biased deletion and ranking: the GEM(α,θ) stick weights satisfy W_k ~ Beta(θ+kα, 1-α), and the Markov chain structure in (3.1)-(3.2) from Perman-Pitman-Yor [24] and Pitman-Yor [32].
- standard math The identity E[S^m_α e^{-λS_α}] = α e^{-λ^α} λ^{α-m} Ω_m(λ^α) in Eq. (2.4), relating moments of stable variables to generalized Stirling numbers.
Cite this review
Pith. "Pith review of Stick-breaking Pitman-Yor processes given the species sampling size." pith.science (2026). https://pith.science/paper/D7YXNQNB
@misc{pith2026190807186,
author = {Pith},
title = {Pith review of: Stick-breaking Pitman-Yor processes given the species sampling size},
year = {2026},
howpublished = {\url{https://pith.science/paper/D7YXNQNB}},
note = {Machine review of arXiv:1908.07186}
}
abstract
Random discrete distributions, say $F,$ known as species sampling models, represent a rich class of models for classification and clustering, in Bayesian statistics and machine learning. They also arise in various areas of probability and its applications. Jim Pitman, within the species sampling context, shows that mixed Poisson processes may be interpreted as the sample size up till a given time or in terms of waiting times of appearance of individuals to be classified. He notes connections to some recent work in the Bayesian statistic/machine learning literature, with some more classical results. We let $F:=F_{\alpha,\theta},$ be a Pitman-Yor process for $\alpha\in (0,1),$ and $\theta>-\alpha,$ with $\alpha$-diversity equivalent in distribution to $S^{-\alpha}_{\alpha,\theta},$ and let $(N_{S_{\alpha,\theta}}(\lambda),\lambda\ge 0)$ denote a mixed Poisson process with rate $S_{\alpha,\theta}.$ In this paper we derive explicit stick-breaking representations of $F_{\alpha,\theta}$ given $N_{S_{\alpha,\theta}}(\lambda)=m.$ More precisely, if $(P_{\ell})\sim \mathrm{PD}(\alpha,\theta)$, denotes a ranked sequence following the two parameter Poisson-Dirichlet distribution, we obtain explicit representations of the sized biased permutation of $(P_{\ell})|N_{S_{\alpha,\theta}}(\lambda)=m.$ Due to distributional results we shall develop in a more general context, it suffices to consider the stable case $F_{\alpha,0}|N_{S_{\alpha}}(\lambda)=m.$ Notably, it follows that $F_{\alpha,0}|N_{S_{\alpha}}(\lambda)=0,$ is equivalent in distribution to the popular normalized generalized gamma process. Hence, we obtain explicit stick-breaking representations for the generalized gamma class as a special case.
Forward citations
Cited by 1 Pith paper
-
Poisson Hierarchical Indian Buffet Processes-With Indications for Microbiome Species Sampling Models
A new hierarchical Bayesian nonparametric model for sparse count data provides exact generative sampling, tractable posterior inference, and predictive rules for unseen species in microbiome studies.
Reference graph
Works this paper leans on
-
[30]
Pitman, J. (2017). Mixed Poisson and negative binomial models for clus- tering and species sampling. Manuscript in preparation
work page 2017
-
[1]
Aldous, D. and Pitman, J. (1998). The standard additive coalescent Ann. Probab. 26, 1703-1726
work page 1998
-
[2]
Barlow, M., Pitman, J. and Yor, M. (1989). Une extension multidimen- sionnelle de la loi de l’arc sinus. In S´ eminaire de Probabilit´ es XXIII(Azema, J., Meyer, P.-A. and Yor, M., Eds.), 294–314, Lecture Notes in Math ematics 1372, Springer, Berlin
work page 1989
-
[3]
Bertoin, J. (2006). Random fragmentation and coagulation processes, Cam- bridge University Press
work page 2006
-
[4]
(2013) Clusters and fea- tures from combinatorial stochastic processes
Broderick, T, Jordan, M.I., and Pitman, J. (2013) Clusters and fea- tures from combinatorial stochastic processes. Statist. Sci. 28, 289-312
work page 2013
-
[5]
Devroye, L. (2009). Random variate generation for exponentially and poly- nomially tilted stable distributions. ACM Transactions on Modeling and Computer Simulation (TOMACS) 19, Issue 4, Article No. 18
work page 2009
-
[6]
Doksum, K. (1974). Tailfree and neutral random probabilities and their posterior distributions. Ann. Probab. 2 183–201
work page 1974
-
[7]
(1988) Population genetics theory - the past and the future
Ewens, W.J. (1988) Population genetics theory - the past and the future. In S. Lessard, editor, Mathematical and Statistical Problems in Evolution. University of Montreal Press, Montreal, 1988
work page 1988
Show all 35 references
-
[8]
and Pitman, J
Gnedin, A. and Pitman, J. (2005). Regenerative composition structures Ann. Probab. 33, 445-479
2005
-
[9]
and Teh, Y.W
F avaro, S, Lijoi, A., Nipoti,B.,Pr¨unster, I. and Teh, Y.W. (2016). On the stick-breaking representation for homogeneous NRMIs. Bayesian Anal- ysis, 11 697-724
2016
-
[10]
and Teh, Y.W
F avaro, S, Lomelli, M, Nipoti,B. and Teh, Y.W. (2014). On the stick- breaking representation of sigma-stable Poisson-Kingman models. Electronic Journal of Statistics , 8 1063-1085
2014
-
[11]
and Pr¨unster, I
F avaro, S., Ljoi, A. and Pr¨unster, I. (2012). On the stick-breaking representation of normalized inverse Gaussian priors. Biometrika, 99 663- 674
2012
-
[12]
Ferguson, T.S. (1973). A Bayesian analysis of some nonparametric prob- lems. Ann. Statist. 1, 209–230
1973
-
[13]
A., Corbet, S
Fisher, R. A., Corbet, S. A., and Williams, C.B. (1943) The relation between the number of species and the number of individuals in a rand om sample of an animal population. The Journal of Animal Ecology 42-58
1943
-
[14]
and James, L.F
Ishwaran, H. and James, L.F. (2001). Gibbs sampling methods for stick- breaking priors. J. Amer. Statist. Assoc. 96, 161–173
2001
-
[15]
and James, L.F
Ishwaran, H. and James, L.F. (2003). Generalized weighted Chinese restaurant processes for species sampling mixture models. Statist. Sinica , 13 1211-1235
2003
-
[16]
James, L.F. (2006). Poisson calculus for spatial neutral to the right pro- cesses. Ann, Stat 34, 416-440
2006
-
[17]
James, L. F. (2010). Lamperti type laws. Ann. Appl. Probab. 20, 1303– 1340
2010
-
[18]
James, L.F. (2013). Stick-breaking PG( α,ζ )-Generalized Gamma Pro- imsart-generic ver. 2007/12/10 file: Arxivstickbreaking 2019.tex date: August 21, 2019 Lancelot F. James/conditional GEM(α, θ ) 21 cesses. Unpublished manuscript. arXiv:1308.6570[math.PR]
2013 arXiv
-
[19]
F., Lijoi, A
James, L. F., Lijoi, A. and Pr¨unster, I. (2009). Posterior analysis for normalized random measures with independent increments. Scand. J. Stat. 36 76–97
2009
-
[20]
Kingman, J. F. C. (1975). Random discrete distributions J. R. Stat. Soc. Ser. B 37 1-22
1975
-
[21]
Lau, J. W. and Cripps, E. (2015). Stick-Breaking Representation and Computation for Normalized Generalized Gamma Processes. Sankhya A , 77 300–329
2015
-
[22]
and Pr¨unster, I
Lijoi, A. and Pr¨unster, I. (2010), Models Beyond the Dirichlet Process, in Bayesian Nonparametrics , eds. N. L. Hjort, C. Holmes, P.Mller, and S. G. Walker, Cambridge, UK: Cambridge University Press, pp. 80136
2010
-
[23]
McCloskey, J. W. (1965). A model for the distribution of individuals by species in an environment. Ph.D. thesis, Michigan State Univ
1965
-
[24]
and Yor, M
Perman, M., Pitman, J. and Yor, M. (1992). Size-biased sampling of Poisson point processes and excursions. Probab. Theory Related Fields. 92, 21-39
1992
-
[25]
Pitman, J. (1995). Exchangeable and partially exchangeable random par- titions. Probab. Theory Related Fields. 102, 145-158
1995
-
[26]
Pitman, J. (1996). Some developments of the Blackwell-MacQueen urn scheme. Statistics, probability and game theory, 245–267, IMS Lecture No tes Monogr. Ser., 30, Inst. Math. Statist., Hayward, CA
1996
-
[27]
Pitman, J. (1999). Brownian motion, bridge, excursion, and meander char- acterized by sampling at independent uniform times. Electron. J. Probab. 4, paper no. 11, 1–33
1999
-
[28]
Pitman, J. (2003). Poisson-Kingman partitions. In Science and Statis- tics: A Festschrift for Terry Speed. (D.R. Goldstein, Ed.), 1–34, Institute of Mathematical Statistics Hayward, California
2003
-
[29]
Pitman, J. (2006). Combinatorial stochastic processes. Lectures from the 32nd Summer School on Probability Theory held in Saint-Flour, July 7– 24,
2006
-
[31]
and Yor, M
Pitman, J. and Yor, M. (1992). Arcsine laws and interval partitions derived from a stable subordinator. Proc. London Math. Soc. 65, 326–356
1992
-
[32]
and Yor, M
Pitman, J. and Yor, M. (1997). The two-parameter Poisson-Dirichlet distribution derived from a stable subordinator. Ann. Probab. 25, 855–900
1997
-
[33]
Sethuraman, J. (1994). A constructive definition of Dirichlet priors. Statist. Sinica 4, 639-650
1994
-
[34]
(2017) Frequency of Frequen- cies Distributions and Size-Dependent Exchangeable Random Partit ions, J
Zhou,M., F avaro,S.and W alker, S.G. (2017) Frequency of Frequen- cies Distributions and Size-Dependent Exchangeable Random Partit ions, J. Amer. Statist. Assoc. 112, 1623-1635, imsart-generic ver. 2007/12/10 file: Arxivstickbreaking 2019.tex date: August 21, 2019
2017
-
[2002]
Lecture Notes in Mathematic s, 1875
With a foreword by Jean Picard. Lecture Notes in Mathematic s, 1875. Springer-Verlag, Berlin
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.